# Paying for it

65 sessions · follows the tags cost
Page: https://mlopstalks.com/threads/paying-for-it

Cost becomes a much more common topic in the archive in 2023. The sessions that follow ask where the money goes and what teams can afford to run.

## 2021

1 session.

- [Law of Diminishing Returns for Running AI Proof-of-Concepts](https://mlopstalks.com/talks/law-of-diminishing-returns-for-running-ai-proof-of-concepts) (Oguzhan Gencoglu, Top Data Science). Claim: Before writing code, teams should ask problem owners about production expectations, budget, maintenance, support, speed, fairness, regulation, and privacy. [21:01](https://www.youtube.com/watch?v=j09xbtudJgs&t=1261s)

## 2022

3 sessions.

- [Building ML/Data Platform on Top of Kubernetes](https://mlopstalks.com/talks/building-ml-data-platform-on-top-of-kubernetes) (Julien Bisconti). Claim: Companies often underestimate the total cost of building and maintaining their own platform, including finding the right abstractions, training people, staffing on-call rotations, and keeping the system running. [10:26](https://www.youtube.com/watch?v=u1ggSj0OwMU&t=626s)
- [Bringing DevOps Agility to ML](https://mlopstalks.com/talks/bringing-devops-agility-to-ml) (Luis Ceze, OctoML). Claim: Luis Ceze argues that efficient inference matters because successful models may be used an extremely large number of times, making inference costs and energy use significant. [43:24](https://www.youtube.com/watch?v=kPr9lxVn-N0&t=2604s)

1 more from 2022 on this thread: https://mlopstalks.com/threads/paying-for-it/2022

## 2023

27 sessions.

- [Cost Optimization and Performance](https://mlopstalks.com/talks/cost-optimization-and-performance) (Lina Weichbrodt & Luis Ceze, OctoML & Jared Zoneraich, Prompt Layer & Daniel Campos, Neeva & Mario Kostelac, Intercom). Pushed back: The best cost-reduction approach was disputed between selecting and optimizing the right model and hardware, versus pruning and distilling a larger model. [8:45](https://www.youtube.com/watch?v=wxq1ZeAM9fc&t=525s)
- [Efficiently Scaling and Deploying LLMs](https://mlopstalks.com/talks/efficiently-scaling-and-deploying-llms) (Hanlin Tang, MosaicML). Claim: The cost and difficulty of training large language models are lower than commonly assumed when appropriate tooling is used. [12:07](https://www.youtube.com/watch?v=AVccFl8-5-8&t=727s)
- [Using LLMs to Punch Above Your Weight!](https://mlopstalks.com/talks/using-llms-to-punch-above-your-weight) (Cameron Feenstra, Anzen). Claim: Cameron Feenstra says relying on an API can make iteration on business logic, prompts, and model inputs expensive or even cost-prohibitive. [19:07](https://www.youtube.com/watch?v=1_NTxx3CJXg&t=1147s)
- [Why is MLOps Hard in an Enterprise?](https://mlopstalks.com/talks/why-is-mlops-hard-in-an-enterprise) (Maria Vechtomova & Basak Eskili, Ahold Delhaize). Claim: Maria Vechtomova says Ahold Delhaize reduced infrastructure costs by standardizing MLOps processes and changing how clusters were created, used, and monitored. [50:10](https://www.youtube.com/watch?v=RKbMww5kxHE&t=3010s)

23 more from 2023 on this thread: https://mlopstalks.com/threads/paying-for-it/2023

## 2024

14 sessions.

- [Anatomy of a Software 3.0 Company](https://mlopstalks.com/talks/anatomy-of-a-software-3-0-company) (Sarah Guo, Conviction). Pushed back: Sarah Guo disputed the idea that all AI companies must immediately be gross-margin positive, saying a company can accept negative gross margins if inference costs are expected to decline. [27:41](https://www.youtube.com/watch?v=UFqHtIls7sM&t=1661s)
- [Reliable Hallucination Detection in Large Language Models](https://mlopstalks.com/talks/reliable-hallucination-detection-in-large-language-models) (Jiaxin Zhang, Intuit AI Research). Claim: Three to five samples can provide competitive performance while reducing sampling cost compared with larger sample sizes. [17:35](https://www.youtube.com/watch?v=G5tBnPEyjLc&t=1055s)
- [How to Actually Use Cost Effective AI in Your Business](https://mlopstalks.com/talks/how-to-actually-use-cost-effective-ai-in-your-business) (Eddie Mattia, Outerbounds & Scott Perry, AWS). Pushed back: The speakers distinguish AWS Trainium and Inferentia from GPUs rather than treating them as AWS versions of GPUs. [22:03](https://www.youtube.com/watch?v=nWrgiwsLgO8&t=1323s)
- [LLMs to agents: The Beauty & Perils of Investing in GenAI](https://mlopstalks.com/talks/llms-to-agents-the-beauty-perils-of-investing-in-genai) (Sandeep Bakshi, Prosus & Meera Clark, Redpoint Ventures & George Robson, Sequoia Capital). Pushed back: Meera Clark argues that many AI businesses are not currently viable because of cost, while George Robson and Sandeep Bakshi say proprietary data and relative competition can still support viable businesses. [20:31](https://www.youtube.com/watch?v=PtVp6zLYcoY&t=1231s)

10 more from 2024 on this thread: https://mlopstalks.com/threads/paying-for-it/2024

## 2025

14 sessions.

- [The Agent Landscape - Lessons Learned Putting Agents Into Production](https://mlopstalks.com/talks/the-agent-landscape-lessons-learned-putting-agents-into-production) (Paul van der Boor & Floris Fok, Prosus Group). Pushed back: The speakers reject the idea that declining cost per token automatically means that agent systems are becoming cheaper overall, because cost per answer can rise as agents use more tokens and calls. [21:33](https://www.youtube.com/watch?v=lRGldru7ohU&t=1293s)
- [How Agents Changed Vibe Coding Forever](https://mlopstalks.com/talks/how-agents-changed-vibe-coding-forever) (Beyang Liu, Sourcegraph). Pushed back: Beyang Liu argues against keeping coding-agent costs low as the primary goal when higher costs save substantial human labor. [33:37](https://www.youtube.com/watch?v=ekHpNKtb9oY&t=2017s)
- [Small Language Models are the Future of Agentic AI](https://mlopstalks.com/talks/small-language-models-are-the-future-of-agentic-ai) (Nehil Jain, Stealth AI Startup & Sonam Gupta, AICamp). Pushed back: Nehil Jain disputes the idea that smaller models are always cheaper by pointing to endpoint utilization, talent, and management costs. [29:26](https://www.youtube.com/watch?v=UKYJjJg_91k&t=1766s)
- [How to Optimize AI Agents in Production](https://mlopstalks.com/talks/how-to-optimize-ai-agents-in-production) (). Pushed back: Nimrod disputes the idea that the model alone determines accuracy, arguing that configuration choices can produce better accuracy-cost tradeoffs. [10:54](https://www.youtube.com/watch?v=--le-yBdVPk&t=654s)

10 more from 2025 on this thread: https://mlopstalks.com/threads/paying-for-it/2025

## 2026

6 sessions.

- [The Shadow AI Problem Nobody's Talking About](https://mlopstalks.com/talks/the-shadow-ai-problem-nobodys-talking-about) (Euro Beinat, Prosus Group). Claim: People will use AI tools only when the tools are immediately helpful to them, so organizations should lower the barriers rather than force employees to spend separate time on experimentation. [23:31](https://www.youtube.com/watch?v=WYBX6fqQyFo&t=1411s)
- [The Coding Agent Multiverse of Madness](https://mlopstalks.com/talks/the-coding-agent-multiverse-of-madness) (Ankit Mathur, Databricks). Claim: The coding agent gateway is intended to give developers freedom to use their preferred tools while giving enterprise administrators centralized security, procurement, observability, and cost controls. [7:27](https://www.youtube.com/watch?v=A3GubVxcNHQ&t=447s)
- [How We Cut LLM Latency 70% With TensorRT in Production](https://mlopstalks.com/talks/how-we-cut-llm-latency-70-with-tensorrt-in-production) (Maher Hanafi, Betterworks). Pushed back: Maher Hanafi rejected the initial assumption that smaller GPUs and multiple models per GPU would be the most efficient configuration. [12:08](https://www.youtube.com/watch?v=wTrv1hMQbVg&t=728s)
- [Agents & the $40M Bet on Multiplayer AI](https://mlopstalks.com/talks/agents-the-40m-bet-on-multiplayer-ai) (Stanislas Polu, Dust). Pushed back: Stanislas Polu rejects the assumption that a model-performance plateau would immediately make inference tokens nearly free because demand and limited inference capacity could preserve high margins. [42:32](https://www.youtube.com/watch?v=NsLPju6TZVc&t=2552s)

2 more from 2026 on this thread: https://mlopstalks.com/threads/paying-for-it/2026
