Cohere's SVP Technology - Saurabh Baji

Topics covered
Popular Clips
Questions from this episode
- Asked by 100 people
- Asked by 40 people
- Asked by 34 people
- Asked by 29 people
- Asked by 19 people
- Asked by 15 people
- Asked by 11 people
- Asked by 8 people
- Asked by 6 people
- Asked by 4 people
- Asked by 3 people
- Asked by 2 people
- Asked by 1 person
Episode Highlights
Resource Optimization
Cohere's approach to resource optimization in AI model deployment is both innovative and pragmatic. highlights the importance of utilizing cloud infrastructure efficiently, drawing parallels with past experiences in serverless computing 1. By leveraging techniques like running multiple fine-tunes on a single GPU, Cohere maximizes hardware utilization while minimizing costs 2. This strategy ensures that customers can achieve high performance without the need for massive computational resources.
We have literally customers who are running 50 fine tunes on one single GPU.
---
Additionally, the use of serverless AI systems allows for flexible and efficient deployment, adapting to the varying needs of enterprises 3.
  Â
Cost-Effective AI
Cohere's focus on cost-effective AI solutions makes advanced technology accessible to a broader range of enterprises. emphasizes the balance between performance and affordability, ensuring that AI models are not only powerful but also economically viable 4. By offering full fine-tuning capabilities, Cohere allows businesses to customize models to their specific needs without incurring prohibitive costs 2.
My job as SVP of engineering is to really make sure that we are applying this amazing technology in the way that customers find useful.
---
This approach fosters independence among customers, enabling them to leverage AI technology effectively while maintaining control over their resources 5.
  Â
Innovative Techniques
Cohere employs innovative techniques to enhance the efficiency and performance of AI models. The ability to run multiple fine-tunes on a single GPU is a testament to their commitment to maximizing resource use 6. discusses the pragmatic implementation of retrieval-augmented generation (RAG), which integrates enterprise data with AI models to improve output quality 7.
The retrieval part almost doesn't get enough credit with retrieval augmented generation.
---
Additionally, the use of synthetic data solutions allows for effective model training, even with limited data, by focusing on imparting new abilities to the model rather than overwhelming it with unnecessary information 8.
Related Episodes


Bold AI Predictions From Cohere Co-founder
Answers 383 questions

Cohere co-founder Nick Frosst on building LLM apps for business
Answers 383 questions

#80 AIDAN GOMEZ [CEO Cohere] - Language as Software
Answers 383 questions

Aiden Gomez - CEO of Cohere (AI's 'Inner Monologue' – Crucial for Reasoning)
Answers 383 questions

Kaggle, ML Community / Engineering (Sanyam Bhutani)
Answers 383 questions

$450M AI Startup In 3 Years | Chai AI
Answers 383 questions

Speechmatics CTO - Next-Generation Speech Recognition
Answers 383 questions

Jay Alammar on LLMs, RAG, and AI Engineering
Answers 383 questions

Sara Hooker - Why US AI Act Compute Thresholds Are Misguided
Answers 383 questions

Joscha Bach and Connor Leahy on AI risk
Answers 383 questions

#92 - SARA HOOKER - Fairness, Interpretability, Language Models
Answers 383 questions

[SPONSORED] The Digitized Self: AI, Identity and the Human Psyche (YouAi)
Answers 383 questions

Connor Leahy - e/acc, AGI and the future.
Answers 383 questions

#97 SREEJAN KUMAR - Human Inductive Biases in Machines from Language
Answers 383 questions

CAN MACHINES REPLACE US? (AI vs Humanity) - Maria Santacaterina
Answers 383 questions
