inference

53 talks

PodcastAI Is Fast. AI Projects Are Slow. Let's Fix That.JRocketRide's Joe Maionchi · 56:48 · Jun 2026 · 293 views · MLOps PodcastPodcastHow We Cut LLM Latency 70% With TensorRT in ProductionMaher Hanafi, Betterworks · 1:05:20 · Apr 2026 · 489 views · MLOps PodcastPodcastFixing GPU Starvation in Large-Scale Distributed TrainingKashish Mittal, Uber · 52:49 · Apr 2026 · 316 views · MLOps PodcastShip Agents: A Virtual Conference Track 2Adam Boaz Becker & Sarmad Absil, Trial Cyber & Divia Mahajan, Amazon Alexa · 1:38:26 · Apr 2026 · 261 views · Ship Agents 2026Quantized LLM Training at Scale with ZeRO++Guanhua Wang, Microsoft · 24:35 · Sept 2025 · 157 views · AI in Production 2025Reading groupSmall Language Models are the Future of Agentic AIAdam Becker, MLOps Community & Nehil Jain, Stealth AI Startup & Sonam Gupta, AI Camp · 58:13 · Sept 2025 · 807 views · MLOps Reading GroupVoice model performance optimizationMadison Kanna, Baseten · 16:32 · Aug 2025 · 282 views · Agents in Production 2025ML Engineers Who Ignore LLMs Are Voluntarily Retiring EarlyKostas Pardalis & Yoni Michael, Typedef · 1:37:23 · Jun 2025 · 997 viewsTesting AI Intelligence: The Benchmarking BattleGreg Kamradt, Arc Prize · 48:31 · Jun 2025 · 358 viewsBuilding Out GPU CloudsMohan Atreya, Rafay Systems · 47:58 · May 2025 · 326 viewsPodcastEfficient GPU infrastructure at LinkedInAnimesh Singh, LinkedIn · 59:14 · Mar 2025 · 619 views · MLOps PodcastPodcastEfficient Deployment of Models at the EdgeKrishna Sridhar, Qualcomm · 51:34 · Jan 2025 · 739 views · MLOps PodcastState of AI Report 2024Nathan Benaich, Air Street Capital · 27:00 · Dec 2024 · 633 views · Agents in Production 2024PodcastLLM Distillation and CompressionGuanhua "Alex" Wang, Microsoft · 49:48 · Dec 2024 · 579 views · MLOps PodcastPodcastAI's Next FrontierAditya Naganath, Kleiner Perkins · 56:04 · Dec 2024 · 264 views · MLOps PodcastReading groupThe Future of AI: Long-Context RAGValdimar Eggertsson, Snjallgögn (Smart Data inc.) & Sophia Skowronski, Breckinridge Capital Advisors & Adam Becker & Binoy Perera, MLOps Community · 49:19 · Dec 2024 · 259 views · MLOps Reading GroupPodcastHow to Optimize Large AI Models with PyTorchMichael Gschwind, Meta Platforms · 57:44 · Nov 2024 · 519 views · MLOps PodcastReading groupSmall Models, Big Ideas: The Next Frontier in AIKorri Jones, Chick-fil-A Corporate Support Center & Valdimar Eggertsson, Snjallgögn (Smart Data inc.) & Sophia Skowronski, Breckinridge Capital Advisors & Lihu Chen, Imperial College London & Binoy Perera, MLOps Community · 58:44 · Nov 2024 · 275 views · MLOps Reading GroupPodcastComposable Memory for GPU OptimizationBernie Wu, MemVerge · 55:19 · Oct 2024 · 381 views · MLOps PodcastAI-Powered Data Unification for Data PlatformsShelby Heinecke, Salesforce · 13:16 · Oct 2024 · 118 views · DE4AI 2024DuckDB is fast for analytics, but what can it do for AI?Mehdi Ouazza, MotherDuck · 12:49 · Oct 2024 · 651 views · DE4AI 2024Boosting LLMs: Performance, Scaling, and Structured OutputsTom Sabo, SAS & Matt Squire, Fuzzy Labs & Vaibhav Gupta, Boundary ML · 1:01:24 · Oct 2024 · 403 views · MLOps Mini Summit 2024PodcastAI For Good - Detecting Harmful Content at ScaleMatar Haller, ActiveFence · 51:28 · Jul 2024 · 402 views · MLOps PodcastPodcastAWS Trainium and InferentiaKamran Khan, Annapurna ML & Matthew McClean, AWS, Annapurna Labs · 45:23 · Jun 2024 · 826 views · MLOps PodcastEnabling Efficient Trillion Parameter Scale Training for Deep Learning ModelsTunji Ruwase, Microsoft · 27:36 · Apr 2024 · 577 views · AI in Production 2024Fine Tuning LlamasKai Davenport · 13:37 · Apr 2024 · 258 views · AI in Production 2024Graduating from Proprietary to Open Source Models in ProductionPhilip Kiely, Baseten · 23:16 · Apr 2024 · 147 views · AI in Production 2024Vision Pipelines in Production: Serving & OptimisationsBiswaroop Bhattacharjee, Prem AI · 13:24 · Apr 2024 · 281 views · AI in Production 2024Anatomy of a Software 3.0 CompanySarah Guo, Conviction · 35:21 · Mar 2024 · 1,564 views · AI in Production 2024Model Merging and Mixtures of ExpertsMaxime Labonne, J.P. Morgan · 11:17 · Mar 2024 · 2,020 views · AI in Production 2024PodcastFounding, Funding, and the Future of MLOpsMihail Eric, Storia AI · 57:31 · Jan 2024 · 377 views · MLOps PodcastPodcastLLMs in Focus: From One-Size Fits All to Verticalized SolutionsVenky Ganti & Laurel Orr, Numbers Station · 55:15 · Dec 2023 · 434 views · MLOps PodcastPodcastBuilding the Future of AI in Software DevelopmentVarun Mohan, Codeium · 1:04:35 · Dec 2023 · 1,085 views · MLOps PodcastThe State of Open Source AI: Deployment Engines, Licences, & HardwareCasper da Costa-Luis, Premai · 10:56 · Nov 2023 · 253 viewsPodcastImpact of LLMs on the Tech Stack and Product DevelopmentAnand Das, Bito · 55:31 · Nov 2023 · 431 views · MLOps PodcastExploring the Latency/Throughput & Cost Space for LLM InferenceTimothée Lacroix, Mistral · 30:25 · Oct 2023 · 29K viewsPodcastFrugalGPT: Better Quality and Lower Cost for LLM ApplicationsLingjiao Chen, Stanford University · 1:02:59 · Aug 2023 · 973 views · MLOps PodcastBuilding RedPajamaVipul Ved Prakash, Together · 27:52 · Aug 2023 · 478 views · LLMs in Production 2023End-to-end Modern Machine Learning in ProductionOmar Sanseviero, Hugging Face · 10:30 · Aug 2023 · 753 views · LLMs in Production 2023Making LLM Inference AffordableDaniel Campos, Snowflake · 32:07 · Jul 2023 · 1,782 views · LLMs in Production 2023Challenges in Providing LLMs as a ServiceHemant Jain, Cohere AI · 11:43 · Jul 2023 · 466 views · LLMs in Production 2023LLM on KubernetesShrinand Javadekar, Outerbounds & Manjot Pahwa, Lightspeed India & Rahul Parundekar, A.I. Hero & Patrick Barker · 36:24 · Jul 2023 · 1,123 views · Conference in Production 2023Create a Contextual Chatbot with LLM and a Vector Database in 10 MinutesRaahul Dutta, Elsevier · 10:07 · Jul 2023 · 1,858 views · MLOps Community AmsterdamPodcastPython Power: How Daft Embeds Models and Revolutionizes Data ProcessingSammy Sidhu, Eventual · 51:30 · Jul 2023 · 577 views · MLOps PodcastUnderstanding the LLM EconomicsNikunj Bajaj, TrueFoundry · 31:23 · Jul 2023 · 1,036 views · LLMs in Production 2023Large Model Training and Inference with DeepSpeedSamyam Rajbhandari, Microsoft DeepSpeed · 36:23 · Jun 2023 · 9,517 views · LLMs in Production 2023PodcastFrom Arduinos to LLMs: Exploring the Spectrum of MLSoham Chatterjee, Sleek · 44:50 · Jun 2023 · 560 views · MLOps PodcastPodcastML Scalability ChallengesWaleed Kadous, Anyscale · 1:00:03 · Apr 2023 · 838 views · MLOps PodcastLarge Language Models in Production Round-table ConversationDiego Oppenheimer, Factory HQ & David Hershey, Unusual Ventures & Hannes Hapke, Digits & James Richards, Bountiful & Rebecca Qian, Facebook AI Research · 57:21 · Mar 2023 · 14K views · LLMs in Production 2023MeetupThe 7 Lines of Code You Need to Run Faster Real-time InferenceAdrian Boguszewski, Intel · 49:21 · Mar 2023 · 608 views · MLOps MeetupPodcastBringing DevOps Agility to MLLuis Ceze, OctoML · 1:04:27 · Sept 2022 · 1,217 views · MLOps Coffee SessionsMeetupThe Role of Resource Management in MLOpsRonen Dar & Gijsbert Janssen van Doorn, Run:AI · 53:38 · May 2022 · 276 views · MLOps MeetupPodcastDon't Listen Unless You Are Going to Do ML in ProductionKyle Morris, banana.dev · 51:30 · Mar 2022 · 768 views · MLOps Coffee Sessions