gpus

41 talks

PodcastHow We Cut LLM Latency 70% With TensorRT in ProductionMaher Hanafi, Betterworks · 1:05:20 · Apr 2026 · 489 views · MLOps PodcastPodcastFixing GPU Starvation in Large-Scale Distributed TrainingKashish Mittal, Uber · 52:49 · Apr 2026 · 316 views · MLOps PodcastPodcastPerformance Optimization and Software/Hardware Co-design across PyTorch, CUDA, and NVIDIA GPUsChris Fregly, AI performance engineer, startup founder, and investor · 1:25:50 · Mar 2026 · 675 views · MLOps PodcastFast & Asynchronous: Drift Your AI, Not Your GPU BillArtem Yushkovskiy, Delivery Hero · 33:30 · Feb 2026 · 99 views · Coding Agents Conference 2026PodcastSpeed and Scale: How Today's AI Datacenters Are Operating Through HypergrowthKris Beevers, NetBox Labs · 1:07:17 · Feb 2026 · 171 views · MLOps PodcastAccelerating Growth Through Optimizing GPU UsageSahil Khanna, Adobe · 23:53 · Jan 2026 · 129 views · AI in Production 2025PodcastThe GPU Uptime BattleAndy Pernsteiner, VAST Data · 1:33:46 · Nov 2025 · 491 views · MLOps PodcastBuilding Data Centers for GPU CloudsCraig Tavares, Buzz HPC · 46:00 · Oct 2025 · 457 viewsQuantized LLM Training at Scale with ZeRO++Guanhua Wang, Microsoft · 24:35 · Sept 2025 · 157 views · AI in Production 2025Voice model performance optimizationMadison Kanna, Baseten · 16:32 · Aug 2025 · 282 views · Agents in Production 2025PodcastThe Truth About LLM TrainingPaul van der Boor & Zulkuf Genc, Prosus Group · 55:47 · Aug 2025 · 822 views · Agents in Production SeriesBuilding Out GPU CloudsMohan Atreya, Rafay Systems · 47:58 · May 2025 · 326 viewsPodcastEfficient GPU infrastructure at LinkedInAnimesh Singh, LinkedIn · 59:14 · Mar 2025 · 619 views · MLOps PodcastPodcastLLM Distillation and CompressionGuanhua "Alex" Wang, Microsoft · 49:48 · Dec 2024 · 579 views · MLOps PodcastPodcastAI's Next FrontierAditya Naganath, Kleiner Perkins · 56:04 · Dec 2024 · 264 views · MLOps PodcastHow to Actually Use Cost Effective AI in Your BusinessEddie Mattia, Outerbounds & Scott Perry, AWS · 49:05 · Nov 2024 · 262 views · MLOps Community Mini Summit #9PodcastComposable Memory for GPU OptimizationBernie Wu, MemVerge · 55:19 · Oct 2024 · 381 views · MLOps PodcastHow GPUs are Revolutionizing AI Data Management · 30:02 · Oct 2024 · 181 viewsPodcastAWS Trainium and InferentiaKamran Khan, Annapurna ML & Matthew McClean, AWS, Annapurna Labs · 45:23 · Jun 2024 · 826 views · MLOps PodcastPodcastHandling Multi-Terabyte LLM CheckpointsSimon Karasik, Nebius AI · 55:37 · Apr 2024 · 654 views · MLOps PodcastEnabling Efficient Trillion Parameter Scale Training for Deep Learning ModelsTunji Ruwase, Microsoft · 27:36 · Apr 2024 · 577 views · AI in Production 2024Innovative Gen AI Applications: Beyond TextDiana C. Montañes Mondragon & Nick Schenone, QuantumBlack · 54:45 · Apr 2024 · 852 views · MLOps Mini Summit 2024Introducing DBRX: The Future of Language ModelsDavis Blalock, Bandish Shah, Abhi Venigalla & Ajay Saini, Databricks · 48:36 · Apr 2024 · 496 views · MLOps Coffee SessionsFine Tuning LlamasKai Davenport · 13:37 · Apr 2024 · 258 views · AI in Production 2024Productionizing Health Insurance Appeal GenerationHolden Karau, Netflix · 27:11 · Apr 2024 · 594 views · AI in Production 2024PodcastBuilding the Future of AI in Software DevelopmentVarun Mohan, Codeium · 1:04:35 · Dec 2023 · 1,085 views · MLOps PodcastThe State of Open Source AI: Deployment Engines, Licences, & HardwareCasper da Costa-Luis, Premai · 10:56 · Nov 2023 · 253 viewsEfficient Serving of LLMs for Experimentation and Production with Fireworks.aiDmytro Dzhulgakov, Fireworks.ai · 11:43 · Oct 2023 · 1,085 viewsExploring the Latency/Throughput & Cost Space for LLM InferenceTimothée Lacroix, Mistral · 30:25 · Oct 2023 · 29K viewsPreemption Chaos and Optimizing Server StartupBradley Heilbrun, Replit · 12:42 · Aug 2023 · 215 views · LLMs in Production 2023Considerations and Optimizations for Deploying Open Source LLMs at Your CompanyOscar Rovira, Mystic AI · 11:31 · Aug 2023 · 474 viewsMaking LLM Inference AffordableDaniel Campos, Snowflake · 32:07 · Jul 2023 · 1,782 views · LLMs in Production 2023The Next Million AI AppsMark Huang, Preemo · 1:03:23 · Jul 2023 · 992 views · LLMs in Pod Con Part 2 Workshop Day 1, 2023LLM on KubernetesShrinand Javadekar, Outerbounds & Manjot Pahwa, Lightspeed India & Rahul Parundekar, A.I. Hero & Patrick Barker · 36:24 · Jul 2023 · 1,123 views · Conference in Production 2023Scalable Evaluation and Serving of Open Source LLMsWaleed Kadous, Anyscale · 34:57 · Jul 2023 · 1,336 views · LLMs in Production 2023Large Model Training and Inference with DeepSpeedSamyam Rajbhandari, Microsoft DeepSpeed · 36:23 · Jun 2023 · 9,517 views · LLMs in Production 2023Efficiently Scaling and Deploying LLMsHanlin Tang, MosaicML · 25:14 · May 2023 · 13K views · LLMs in Production 2023PodcastGPU For Machine LearningRonen Dar & Gijsbert Janssen van Doorn, Run:ai · 1:03:33 · May 2022 · 440 views · MLOps Coffee SessionsMeetupThe Role of Resource Management in MLOpsRonen Dar & Gijsbert Janssen van Doorn, Run:AI · 53:38 · May 2022 · 276 views · MLOps MeetupPodcastDon't Listen Unless You Are Going to Do ML in ProductionKyle Morris, banana.dev · 51:30 · Mar 2022 · 768 views · MLOps Coffee SessionsMeetupDeep Dive on Paperspace ToolingMisha Kutsovsky, Paperspace · 1:07:15 · Jul 2020 · 263 views · MLOps Meetup