Paying for it in 2023

27 sessions

Cost Optimization and PerformanceLina Weichbrodt & Luis Ceze, OctoML & Jared Zoneraich, Prompt Layer & Daniel Campos, Neeva & Mario Kostelac, Intercom · 36:06 · May 2023 · 930 views · LLMs in Production 2023

Pushed backThe best cost-reduction approach was disputed between selecting and optimizing the right model and hardware, versus pruning and distilling a larger model.8:45

Efficiently Scaling and Deploying LLMsHanlin Tang, MosaicML · 25:14 · May 2023 · 13K views · LLMs in Production 2023

ClaimThe cost and difficulty of training large language models are lower than commonly assumed when appropriate tooling is used.12:07

Using LLMs to Punch Above Your Weight!Cameron Feenstra, Anzen · 35:49 · May 2023 · 696 views · LLMs in Production 2023

ClaimCameron Feenstra says relying on an API can make iteration on business logic, prompts, and model inputs expensive or even cost-prohibitive.19:07

PodcastWhy is MLOps Hard in an Enterprise?Maria Vechtomova & Basak Eskili, Ahold Delhaize · 55:06 · May 2023 · 756 views · MLOps Podcast

ClaimMaria Vechtomova says Ahold Delhaize reduced infrastructure costs by standardizing MLOps processes and changing how clusters were created, used, and monitored.50:10

PodcastThe Long Tail of ML DeploymentTuhin Srivastava, Baseten · 50:37 · Jun 2023 · 618 views · MLOps Podcast

ClaimTuhin Srivastava believes smaller, targeted models are preferable when a product requires concentrated value, speed, cost control and performance.28:27

PodcastFrom Arduinos to LLMs: Exploring the Spectrum of MLSoham Chatterjee, Sleek · 44:50 · Jun 2023 · 560 views · MLOps Podcast

ClaimBuilding an application around an LLM API is initially cheap and easy, but complexity, reliability problems, and costs increase as the product grows.21:58

The Emerging Toolkit for Reliable, High-quality LLM ApplicationsMatei Zaharia, Databricks · 31:01 · Jun 2023 · 4,509 views · LLMs in Production 2023

ClaimReliable LLM applications must address operational issues such as cost, performance, availability, model drift, timeliness, and privacy.5:14

Taking ImgFlip's 'This Meme Does Not Exist' to the Next Level with a LLMStefan Ojanen, Genesis Cloud · 14:49 · Jul 2023 · 307 views
Understanding the LLM EconomicsNikunj Bajaj, TrueFoundry · 31:23 · Jul 2023 · 1,036 views · LLMs in Production 2023

ClaimNikunj Bajaj says GPT-4 pricing has separate charges for prompt tokens and response tokens.6:12

It Worked When I Prompted ItSoham Chatterjee, Sleek · 14:29 · Jul 2023 · 310 views · LLMs in Production 2023

ClaimAs an LLM application becomes more complex, longer prompts can cause API costs to rise quickly.5:51

PodcastTreating Prompt Engineering More Like CodeMaxime Beauchemin, Preset · 1:14:18 · Jul 2023 · 674 views · MLOps Podcast

ClaimMaxime Beauchemin says Promptimize can compare different prompts, models, and parameter settings and report differences in accuracy, speed, and cost.17:52

Building ProductsSam Charrington, TWIML AI Podcast & George Mathew, Insight Partners & Asmitha Rathis, PromptOps & Natalia Burina, Meta & Sahar Mor, Stripe · 45:18 · Jul 2023 · 300 views · LLMs in Production 2023

ClaimSahar Mor says prompt chaining, asking a model to critique its own answer, and requiring citations can help mitigate hallucinations, although some of these methods increase cost and latency.23:39

Making LLM Inference AffordableDaniel Campos, Snowflake · 32:07 · Jul 2023 · 1,782 views · LLMs in Production 2023

ClaimThe reduction in summarization cost allowed Neeva to summarize its entire index offline instead of doing the work online.9:15

The Confidence Checklist for LLMs in ProductionRohit Agarwal, Portkey.ai · 32:34 · Aug 2023 · 839 views · LLMs in Production 2023

ClaimRohit Agarwal recommends captchas, rate limits, and monitoring to limit the cost of abusive traffic and DDoS attacks.7:47

LLM XGBoost: Can a Fine-Tuned LLM Beat XGBoost on Tabular Data?Sebastian Cattes, iwt · 10:51 · Aug 2023 · 4,235 views
Preemption Chaos and Optimizing Server StartupBradley Heilbrun, Replit · 12:42 · Aug 2023 · 215 views · LLMs in Production 2023

ClaimUsing preemptible GPUs can cut cloud costs by two-thirds while maintaining uptime for users.2:05

PodcastFrugalGPT: Better Quality and Lower Cost for LLM ApplicationsLingjiao Chen, Stanford University · 1:02:59 · Aug 2023 · 973 views · MLOps Podcast
PodcastTecton Round-table // Get your ML Application Into ProductionKevin Stumpf, Derek Salama, Eddie Esquivel & Isaac Cameron, Tecton · 55:42 · Sept 2023 · 405 views · MLOps Coffee Sessions

ClaimDerek Salama describes an Uber pricing model that cost 50 cents in additional compute per ride and was not deployed in markets where the profit margin was about 10 cents per ride.11:06

Finetuning Open-Source LLMsSebastian Raschka, Lightning AI · 29:04 · Oct 2023 · 3,662 views · LLMs in Production 2023

ClaimIn Sebastian Raschka's speed test, low-rank adaptation trained a 7-billion-parameter model on one GPU in about one hour, while full fine-tuning took about nine hours on six GPUs.14:49

Fireside Chat with LLM StartupsPaul van der Boor & Sandeep Bakshi, Prosus Group & Shriyash Upadhyay, Martian & Lars Maaløe, Corti & Pietro Gagliano, Transitional Forms · 30:46 · Oct 2023 · 584 views · LLMs in Production 2023

ClaimMartian routes each user request to the model that provides the highest performance at the lowest cost.8:22

What Drives GenAI Development in the Next 3 YearsEuro Beinat, Prosus · 18:54 · Oct 2023 · 612 views · LLMs in Production 2023

ClaimThe cost of achieving a given level of model quality tends to decrease over time, while model quality continues to improve.5:38

AI in Education Fireside ChatKlinton Bicknell, Duolingo & Bill Salak, Brainly & Yeva Hyusyan, SoloLearn · 31:01 · Oct 2023 · 446 views · LLMs in Production 2023

ClaimKlinton Bicknell says running language models naively at large scale can be cost-prohibitive, so Duolingo uses them to pre-generate options or create inexpensive rules and detectors for real-time use.21:26

Amplifying Impact with Generative AI: Insights from 10,000 ColleaguesPaul van der Boor, Prosus · 32:12 · Oct 2023 · 258 views

ClaimPrompt economics are affected by language because tokenization can require substantially more tokens for some non-English languages and for code than for English.16:40

Exploring the Latency/Throughput & Cost Space for LLM InferenceTimothée Lacroix, Mistral · 30:25 · Oct 2023 · 29K views

ClaimTimothée Lacroix says the talk focuses on the cost of inference, throughput, latency, and deployment of open-source large language models.1:13

Efficient Serving of LLMs for Experimentation and Production with Fireworks.aiDmytro Dzhulgakov, Fireworks.ai · 11:43 · Oct 2023 · 1,085 views

ClaimFine-tuning can reduce serving costs because it enables shorter prompts and sometimes smaller models for the same quality.1:59

Current State of LLMs in ProductionApurva Misra, Truckstop · 11:46 · Nov 2023 · 686 views · LLMs in Production 2023

ClaimQuery routing can reduce cost by sending complex queries to more capable models and simpler queries to less expensive models, and caching can reduce the number of LLM calls.9:14

Building RAG-based LLM Applications for ProductionPhilipp Moritz & Yifei Feng, Anyscale · 30:23 · Nov 2023 · 3,003 views · LLMs in Production 2023

ClaimRay Data was used to scale document processing and embedding because it supports different data sources and can use CPU and GPU resources together.13:55