Paying for it
Cost becomes a much more common topic in the archive in 2023. The sessions that follow ask where the money goes and what teams can afford to run.
Law of Diminishing Returns for Running AI Proof-of-Concepts
ClaimBefore writing code, teams should ask problem owners about production expectations, budget, maintenance, support, speed, fairness, regulation, and privacy.21:01
Building ML/Data Platform on Top of Kubernetes
ClaimCompanies often underestimate the total cost of building and maintaining their own platform, including finding the right abstractions, training people, staffing on-call rotations, and keeping the system running.10:26
Bringing DevOps Agility to ML
ClaimLuis Ceze argues that efficient inference matters because successful models may be used an extremely large number of times, making inference costs and energy use significant.43:24
Cost Optimization and Performance
Pushed backThe best cost-reduction approach was disputed between selecting and optimizing the right model and hardware, versus pruning and distilling a larger model.8:45
Efficiently Scaling and Deploying LLMs
ClaimThe cost and difficulty of training large language models are lower than commonly assumed when appropriate tooling is used.12:07
Using LLMs to Punch Above Your Weight!
ClaimCameron Feenstra says relying on an API can make iteration on business logic, prompts, and model inputs expensive or even cost-prohibitive.19:07
Why is MLOps Hard in an Enterprise?
ClaimMaria Vechtomova says Ahold Delhaize reduced infrastructure costs by standardizing MLOps processes and changing how clusters were created, used, and monitored.50:10
Anatomy of a Software 3.0 Company
Pushed backSarah Guo disputed the idea that all AI companies must immediately be gross-margin positive, saying a company can accept negative gross margins if inference costs are expected to decline.27:41
Reliable Hallucination Detection in Large Language Models
ClaimThree to five samples can provide competitive performance while reducing sampling cost compared with larger sample sizes.17:35
How to Actually Use Cost Effective AI in Your Business
Pushed backThe speakers distinguish AWS Trainium and Inferentia from GPUs rather than treating them as AWS versions of GPUs.22:03
LLMs to agents: The Beauty & Perils of Investing in GenAI
Pushed backMeera Clark argues that many AI businesses are not currently viable because of cost, while George Robson and Sandeep Bakshi say proprietary data and relative competition can still support viable businesses.20:31
The Agent Landscape - Lessons Learned Putting Agents Into Production
Pushed backThe speakers reject the idea that declining cost per token automatically means that agent systems are becoming cheaper overall, because cost per answer can rise as agents use more tokens and calls.21:33
How Agents Changed Vibe Coding Forever
Pushed backBeyang Liu argues against keeping coding-agent costs low as the primary goal when higher costs save substantial human labor.33:37
Small Language Models are the Future of Agentic AI
Pushed backNehil Jain disputes the idea that smaller models are always cheaper by pointing to endpoint utilization, talent, and management costs.29:26
How to Optimize AI Agents in Production
Pushed backNimrod disputes the idea that the model alone determines accuracy, arguing that configuration choices can produce better accuracy-cost tradeoffs.10:54
The Shadow AI Problem Nobody's Talking About
ClaimPeople will use AI tools only when the tools are immediately helpful to them, so organizations should lower the barriers rather than force employees to spend separate time on experimentation.23:31
The Coding Agent Multiverse of Madness
ClaimThe coding agent gateway is intended to give developers freedom to use their preferred tools while giving enterprise administrators centralized security, procurement, observability, and cost controls.7:27
How We Cut LLM Latency 70% With TensorRT in Production
Pushed backMaher Hanafi rejected the initial assumption that smaller GPUs and multiple models per GPU would be the most efficient configuration.12:08
Agents & the $40M Bet on Multiplayer AI
Pushed backStanislas Polu rejects the assumption that a model-performance plateau would immediately make inference tokens nearly free because demand and limited inference capacity could preserve high margins.42:32