Paying for it in 2024

14 sessions

Anatomy of a Software 3.0 CompanySarah Guo, Conviction · 35:21 · Mar 2024 · 1,564 views · AI in Production 2024

Pushed backSarah Guo disputed the idea that all AI companies must immediately be gross-margin positive, saying a company can accept negative gross margins if inference costs are expected to decline.27:41

Reliable Hallucination Detection in Large Language ModelsJiaxin Zhang, Intuit AI Research · 35:24 · Apr 2024 · 1,095 views · AI in Production 2024

ClaimThree to five samples can provide competitive performance while reducing sampling cost compared with larger sample sizes.17:35

LLMOps and GenAI at Enterprise Scale - Challenges and OpportunitiesAndy McMahon, NatWest Group · 13:04 · May 2024 · 564 views · AI in Production 2024

ClaimFor enterprise foundation-model selection, cost per query, cost per user, latency, throughput and whether the model solves the business problem matter more than leaderboard position.2:30

Streamlining Model DeploymentDaniel Lenton, Unify · 21:39 · May 2024 · 334 views · AI in Production 2024

ClaimDaniel Lenton says the number of AI-as-a-service endpoints has grown rapidly, with different endpoints offering different cost, performance, and specialization profiles.1:04

Productionizing AI: How to Think From the EndAnnie Condon · 11:11 · May 2024 · 348 views · AI in Production 2024

ClaimDeploying LLM-based features requires considering API rate limits, cost, processing time, tokenization, embeddings, compute, evaluation, and the quality of user-provided data.6:09

No GPU Before PMFStanislas Polu, Dust · 12:44 · May 2024 · 196 views · AI in Production 2024

ClaimFor B2C AI products, the current priority order is cost, speed, and performance.2:40

Navigating the Emerging LLMOps StackHien Luu, DoorDash · 12:51 · May 2024 · 534 views · AI in Production 2024

ClaimHien Luu says that organizations supporting high-scale QPS use cases need to consider cost and latency.2:30

PodcastAWS Trainium and InferentiaKamran Khan, Annapurna ML & Matthew McClean, AWS, Annapurna Labs · 45:23 · Jun 2024 · 826 views · MLOps Podcast

ClaimAWS developed Trainium and Inferentia as purpose-built AI accelerators to offer customers more choice, higher performance, lower cost, and easier access to compute.2:48

AI-Powered Data Unification for Data PlatformsShelby Heinecke, Salesforce · 13:16 · Oct 2024 · 118 views · DE4AI 2024
How To Cut Your Data Infrastructure Costs in HalfJose Navaro, Cleo · 12:33 · Oct 2024 · 162 views

ClaimJose Navaro says that tracking data infrastructure costs is part of the culture at Cleo because the company helps users manage their finances.1:27

How to Actually Use Cost Effective AI in Your BusinessEddie Mattia, Outerbounds & Scott Perry, AWS · 49:05 · Nov 2024 · 262 views · MLOps Community Mini Summit #9

Pushed backThe speakers distinguish AWS Trainium and Inferentia from GPUs rather than treating them as AWS versions of GPUs.22:03

LLMs to agents: The Beauty & Perils of Investing in GenAISandeep Bakshi, Prosus & Meera Clark, Redpoint Ventures & George Robson, Sequoia Capital · 33:25 · Nov 2024 · 320 views · Agents in Production 2024

Pushed backMeera Clark argues that many AI businesses are not currently viable because of cost, while George Robson and Sandeep Bakshi say proprietary data and relative competition can still support viable businesses.20:31

PodcastAI-Driven Code: Navigating Due Diligence & Transparency in MLOpsMatt van Itallie, Sema · 57:02 · Nov 2024 · 237 views · MLOps Podcast

ClaimSema’s cloud-cost assessment can estimate spending trends, spending maturity and potential savings after about an hour of setup.11:21

We're Using AI Agents at Work (and it's amazing)Euro Beinat, Prosus Group · 27:37 · Dec 2024 · 2,987 views