Threads / Paying for it

Paying for it

Cost becomes a much more common topic in the archive in 2023. The sessions that follow ask where the money goes and what teams can afford to run.

Follows the tags cost · 65 sessions · 2020 to 2026
20211 session
20223 sessions
Podcast · MLOps Coffee Sessions #86

Building ML/Data Platform on Top of Kubernetes

Julien Bisconti

ClaimCompanies often underestimate the total cost of building and maintaining their own platform, including finding the right abstractions, training people, staffing on-call rotations, and keeping the system running.10:26

Podcast · MLOps Coffee Sessions #121

Bringing DevOps Agility to ML

Luis Ceze, OctoML

ClaimLuis Ceze argues that efficient inference matters because successful models may be used an extremely large number of times, making inference costs and energy use significant.43:24

1 more from 2022 on this thread
202327 sessions
Talk · LLMs in Production 2023

Cost Optimization and Performance

Lina Weichbrodt & Luis Ceze, OctoML & Jared Zoneraich, Prompt Layer & Daniel Campos, Neeva & Mario Kostelac, Intercom

Pushed backThe best cost-reduction approach was disputed between selecting and optimizing the right model and hardware, versus pruning and distilling a larger model.8:45

Talk · LLMs in Production 2023

Efficiently Scaling and Deploying LLMs

Hanlin Tang, MosaicML

ClaimThe cost and difficulty of training large language models are lower than commonly assumed when appropriate tooling is used.12:07

Talk · LLMs in Production 2023

Using LLMs to Punch Above Your Weight!

Cameron Feenstra, Anzen

ClaimCameron Feenstra says relying on an API can make iteration on business logic, prompts, and model inputs expensive or even cost-prohibitive.19:07

Podcast · MLOps Podcast #159

Why is MLOps Hard in an Enterprise?

Maria Vechtomova & Basak Eskili, Ahold Delhaize

ClaimMaria Vechtomova says Ahold Delhaize reduced infrastructure costs by standardizing MLOps processes and changing how clusters were created, used, and monitored.50:10

23 more from 2023 on this thread
202414 sessions
Talk · AI in Production 2024

Anatomy of a Software 3.0 Company

Sarah Guo, Conviction

Pushed backSarah Guo disputed the idea that all AI companies must immediately be gross-margin positive, saying a company can accept negative gross margins if inference costs are expected to decline.27:41

Talk · Agents in Production 2024

LLMs to agents: The Beauty & Perils of Investing in GenAI

Sandeep Bakshi, Prosus & Meera Clark, Redpoint Ventures & George Robson, Sequoia Capital

Pushed backMeera Clark argues that many AI businesses are not currently viable because of cost, while George Robson and Sandeep Bakshi say proprietary data and relative competition can still support viable businesses.20:31

10 more from 2024 on this thread
202514 sessions
Reading group · MLOps Reading Group

Small Language Models are the Future of Agentic AI

Nehil Jain, Stealth AI Startup & Sonam Gupta, AICamp

Pushed backNehil Jain disputes the idea that smaller models are always cheaper by pointing to endpoint utilization, talent, and management costs.29:26

Talk · Agents in Production 2025

How to Optimize AI Agents in Production

Pushed backNimrod disputes the idea that the model alone determines accuracy, arguing that configuration choices can produce better accuracy-cost tradeoffs.10:54

10 more from 2025 on this thread
20266 sessions
Talk

The Shadow AI Problem Nobody's Talking About

Euro Beinat, Prosus Group

ClaimPeople will use AI tools only when the tools are immediately helpful to them, so organizations should lower the barriers rather than force employees to spend separate time on experimentation.23:31

Talk · Coding Agents Conference 2026

The Coding Agent Multiverse of Madness

Ankit Mathur, Databricks

ClaimThe coding agent gateway is intended to give developers freedom to use their preferred tools while giving enterprise administrators centralized security, procurement, observability, and cost controls.7:27

Podcast · MLOps Podcast #384

Agents & the $40M Bet on Multiplayer AI

Stanislas Polu, Dust

Pushed backStanislas Polu rejects the assumption that a model-performance plateau would immediately make inference tokens nearly free because demand and limited inference capacity could preserve high margins.42:32

2 more from 2026 on this thread