gpus
41 talks
PodcastHow We Cut LLM Latency 70% With TensorRT in Production
PodcastFixing GPU Starvation in Large-Scale Distributed Training
PodcastPerformance Optimization and Software/Hardware Co-design across PyTorch, CUDA, and NVIDIA GPUs
Fast & Asynchronous: Drift Your AI, Not Your GPU Bill
PodcastSpeed and Scale: How Today's AI Datacenters Are Operating Through Hypergrowth
Accelerating Growth Through Optimizing GPU Usage
PodcastThe GPU Uptime Battle
Building Data Centers for GPU Clouds
Quantized LLM Training at Scale with ZeRO++
Voice model performance optimization
PodcastThe Truth About LLM Training
Building Out GPU Clouds
PodcastEfficient GPU infrastructure at LinkedIn
PodcastLLM Distillation and Compression
PodcastAI's Next Frontier
How to Actually Use Cost Effective AI in Your Business
PodcastComposable Memory for GPU Optimization
How GPUs are Revolutionizing AI Data Management
PodcastAWS Trainium and Inferentia
PodcastHandling Multi-Terabyte LLM Checkpoints
Enabling Efficient Trillion Parameter Scale Training for Deep Learning Models
Innovative Gen AI Applications: Beyond Text
Introducing DBRX: The Future of Language Models
Fine Tuning Llamas
Productionizing Health Insurance Appeal Generation
PodcastBuilding the Future of AI in Software Development
The State of Open Source AI: Deployment Engines, Licences, & Hardware
Efficient Serving of LLMs for Experimentation and Production with Fireworks.ai
Exploring the Latency/Throughput & Cost Space for LLM Inference
Preemption Chaos and Optimizing Server Startup
Considerations and Optimizations for Deploying Open Source LLMs at Your Company
Making LLM Inference Affordable
The Next Million AI Apps
LLM on Kubernetes
Scalable Evaluation and Serving of Open Source LLMs
Large Model Training and Inference with DeepSpeed
Efficiently Scaling and Deploying LLMs
PodcastGPU For Machine Learning
MeetupThe Role of Resource Management in MLOps
PodcastDon't Listen Unless You Are Going to Do ML in Production
MeetupDeep Dive on Paperspace Tooling