inference
53 talks
PodcastAI Is Fast. AI Projects Are Slow. Let's Fix That.
PodcastHow We Cut LLM Latency 70% With TensorRT in Production
PodcastFixing GPU Starvation in Large-Scale Distributed Training
Ship Agents: A Virtual Conference Track 2
Quantized LLM Training at Scale with ZeRO++
Reading groupSmall Language Models are the Future of Agentic AI
Voice model performance optimization
ML Engineers Who Ignore LLMs Are Voluntarily Retiring Early
Testing AI Intelligence: The Benchmarking Battle
Building Out GPU Clouds
PodcastEfficient GPU infrastructure at LinkedIn
PodcastEfficient Deployment of Models at the Edge
State of AI Report 2024
PodcastLLM Distillation and Compression
PodcastAI's Next Frontier
Reading groupThe Future of AI: Long-Context RAG
PodcastHow to Optimize Large AI Models with PyTorch
Reading groupSmall Models, Big Ideas: The Next Frontier in AI
PodcastComposable Memory for GPU Optimization
AI-Powered Data Unification for Data Platforms
DuckDB is fast for analytics, but what can it do for AI?
Boosting LLMs: Performance, Scaling, and Structured Outputs
PodcastAI For Good - Detecting Harmful Content at Scale
PodcastAWS Trainium and Inferentia
Enabling Efficient Trillion Parameter Scale Training for Deep Learning Models
Fine Tuning Llamas
Graduating from Proprietary to Open Source Models in Production
Vision Pipelines in Production: Serving & Optimisations
Anatomy of a Software 3.0 Company
Model Merging and Mixtures of Experts
PodcastFounding, Funding, and the Future of MLOps
PodcastLLMs in Focus: From One-Size Fits All to Verticalized Solutions
PodcastBuilding the Future of AI in Software Development
The State of Open Source AI: Deployment Engines, Licences, & Hardware
PodcastImpact of LLMs on the Tech Stack and Product Development
Exploring the Latency/Throughput & Cost Space for LLM Inference
PodcastFrugalGPT: Better Quality and Lower Cost for LLM Applications
Building RedPajama
End-to-end Modern Machine Learning in Production
Making LLM Inference Affordable
Challenges in Providing LLMs as a Service
LLM on Kubernetes
Create a Contextual Chatbot with LLM and a Vector Database in 10 Minutes
PodcastPython Power: How Daft Embeds Models and Revolutionizes Data Processing
Understanding the LLM Economics
Large Model Training and Inference with DeepSpeed
PodcastFrom Arduinos to LLMs: Exploring the Spectrum of ML
PodcastML Scalability Challenges
Large Language Models in Production Round-table Conversation
MeetupThe 7 Lines of Code You Need to Run Faster Real-time Inference
PodcastBringing DevOps Agility to ML
MeetupThe Role of Resource Management in MLOps
PodcastDon't Listen Unless You Are Going to Do ML in Production