evals
214 talks
Reading groupLoop Engineering
Coding Agents Are Secretly General Agents
PodcastLogs Are All You Need: Rethinking Observability with AI Agents
Architecting Modern AI Systems
Architecting Modern AI Systems: Platforms, Agents, and Integration
PodcastIt's 2026, and We're Still Talking Evals
Stop Shipping on Vibes: How to Build Real Evals for Coding Agents
Using Agents in Production: Past Present and Future
A Playground for AI Engineers
PodcastMLflow Leading Open Source
Simulate to Scale: How realistic simulations power reliable agents in production
Agents as Search Engineers
Building an Orchestration Layer for Agentic Commerce at Loblaws
How AI covered a human's paternity leave
The Future of Coding: AI Agents & the Next Tech Revolution
Beyond the Gold Standard: Evaluating and Trusting Agents in the Wild
Stop Building AI Like Traditional Software
Co-Engineering: The New Era of Human-AI Collaboration
Context Engineering pitfalls for our e-commerce agent
Multi-Agent Systems for the Misinformation Lifecycle
Real-Time Voice Agents in Production
Structured Dissent Patterns for Agentic Production Reliability
Inside OpenAI's AI Agent Collaboration System
Tool definitions are the new Prompt Engineering
PodcastEnterprise AI Operations: The Missing Piece
Building Agentic Tools for Production
Reading groupAI REWIND 2025 - MLOps Reading Group Year-end Special
PodcastThe Future of AI Agents are Sandboxes
PodcastDoes AgenticRAG Really Work?
PodcastHow Sierra AI Does Context Engineering
PodcastVoice AI's Biggest Weakness Exposed
MCP-Enabled Agents
Sub-Agent Architectures: What You Can Leverage
Meta-Prompting: The Hack That's Changing Production AI
Big updates to MLflow 3.0
How to build agents that take ACTION
Evals Aren't Useful? Really?
Beyond Chatbots: How to build Agentic AI systems with Google Gemini
Evaluating AI Agents: Why It Matters and How We Do It
How to Optimize AI Agents in Production
Designing AI Agents for the Complex Realities of Healthcare
Before Building AI Agents Watch These Hard Earned Lessons
Catastrophic agent failure and how to avoid it
Advancing the Cost-Quality Frontier in Agentic AI
Cutting Costs with Artificial Intelligence
A Deep Discussion with the Author of "Context Rot"
Building Real-Time, Reliable Voice AI: From Simulation to Production
Smart Agents Start with Smart LLM Choices
Fast, Trustworthy, Reliable Voice Agents: MLOps That Blend LLM Annotation with Human QA
From Spikes to Stories: AI-Augmented Troubleshooting in the Network Wild
APICA: The Digital Colleague at the Port of Antwerp-Bruges
Iterating on Your AI Evals
Evaluation-Driven Development with MLflow 3.0
Why Language Models Need a Lesson in Education
Advanced Context Engineering
The Science of Improving AI Agents
PodcastThe Truth About LLM Training
The Hidden Bottlenecks Slowing Down AI Agents
MLflow 3.0: The Future of AI Agents
AI Agent Development Tradeoffs You NEED to Know
PodcastInside Uber's AI Revolution: Everything About How They Use AI/ML
ML Engineers Who Ignore LLMs Are Voluntarily Retiring Early
Testing AI Intelligence: The Benchmarking Battle
Everything Hard About Building AI Agents Today
How Product Metrics Become LLM Evaluations
RAG from Scratch with Best Practices
MCP is not going to change everything (yet)
Evaluation of Agentic System
PodcastMaking AI Reliable is the Greatest Challenge of the 2020s
Reading groupA-MEM: Agentic Memory for LLM Agents
PodcastAI-Powered Product Ideation with Synthetic Consumer Testing
PodcastWe're All Finetuning Incorrectly
PodcastBuilding Trust Through Technology: Responsible AI in Practice
AI in Production 2025 | Keynote
PodcastI Let An AI Play Pokémon! - Claude plays Pokémon Creator
PodcastFuture of Software, Agents in the Enterprise, and Inception Stage Company Building
PodcastAI SQL Data Analyst
PodcastWeb Agents: The Cutting Edge of AI is Here?
PodcastThe Challenge with AI Voice Agents
PodcastThe Agent Landscape - Lessons Learned Putting Agents Into Production
PodcastLook At Your ****ing Data 👀
Reading groupAI Agents: The Future of ML Engineering?
PodcastAutonomous AI SRE: The Future of Site Reliability Engineering
PodcastReal LLM Success Stories: How They Actually Work
PodcastAI Careers Insights from Ex Meta Staff Eng
PodcastReal World AI Agent Stories
Reading groupHow AgentOps Enables Observability
PodcastHolistic Evaluation of Generative AI Systems
Is More Really Better: Delve Into Document Strategy
Simulation Techniques for AI Agents from Self-Driving
Building Replit Agent - Hard Lessons Learned
Maximize Your Productivity with LLMs: Task Utility Explained
How to Make AI Agents that ACTUALLY WORK
The Open Source AI Coding Revolution
How AI Agents Will Change Customer Support
AI Agents: The Future of Productivity, or Just a Fad?
Few Shot Code Generation to Autonomous Software Engineering Agents
PodcastWe Can All Be AI Engineers and We Can Do It with Open Source Models
PodcastThe EU AI Act: Navigating New Legislation
PodcastSystematically Test and Evaluate Your LLMs Apps
GenAI in production with MLflow
Supercharging Your RAG System: Techniques and Challenges
Turn Data Chaos into AI Strategy with Programmatic AI Data Development
PodcastMaking Your Company LLM-native
Reading groupIntegrating Knowledge Graphs & Vector RAG for Efficient Information Extraction
PodcastRAG Quality Starts with Data Quality
PodcastVisualize - Bringing Structure to Unstructured Data
PodcastDesign and Development Principles for LLMOps
PodcastHarnessing AI APIs for Safer, Accurate, & Reliable Applications
Balancing Speed and Safety
PodcastReliable LLM Products, Fueled by Feedback
A Blueprint for Scalable & Reliable Enterprise AI/ML Systems
PodcastAI in Healthcare
PodcastEvaluating the Effectiveness of Large Language Models
PodcastExtending AI: From Industry to Innovation
PodcastAI For Good - Detecting Harmful Content at Scale
PodcastAI Agents for Consumers
PodcastNavigating the AI Frontier: The Power of Synthetic Data and Agent Evaluations in LLM Development
How to Build Production-Ready AI Models for Manufacturing
PodcastUber's Michelangelo: Strategic AI Overhaul and Impact
Evaluating Quality and Improving LLM Products at Scale
Beyond Guess-and-Check: Towards AI-assisted Prompt Engineering
Evaluating Language Models
PodcastRetrieval Augmented Generation
Building AI Products across Multiple Domains: Commonalities & Non-Commonalities
Productionizing AI: How to Think From the End
PodcastRecSys at Spotify
PodcastFedML Nexus AI: Your Generative AI Platform at Scale
PodcastWhat is AI Quality?
LLMOps and GenAI at Enterprise Scale - Challenges and Opportunities
Reliable Hallucination Detection in Large Language Models
Shipping LLMs: Buckle Up & Enjoy the Ride
Lessons from Building LLM-based Social Media Products
Making Sense of LLMOps
From MVP to Production
Charting LLMOps Odyssey
Navigating via Retrieval Evaluation to Demystify LLM Wonderland
PodcastThe Art and Science of Training LLMs
Security and Privacy
Model Merging and Mixtures of Experts
PodcastA Decade of AI Safety and Trust
PodcastThe Real E2E RAG Stack
PodcastBecoming an AI Evangelist
LLM Use Cases in Production
PodcastEvaluating and Integrating ML Models
PodcastLLM Evaluation with Arize AI's Aparna Dhinakaran
PodcastRAG Has Been Oversimplified
PodcastThe Myth of AI Breakthroughs
PodcastPioneering AI Models for Regional Languages
PodcastLanguage, Graphs, and AI in Industry
PodcastLLMs in Focus: From One-Size Fits All to Verticalized Solutions
PodcastBuilding the Future of AI in Software Development
MeetupScaling MLOps for Computer Vision
PodcastEnterprises Using MLOps, the Changing LLM Landscape, MLOps Pipelines
Model Blind Spot Discovery for Better Models
Authoring Interactive, Shareable AI Evaluation Reports with Zeno
Evaluating LLMs for AI Risk
GenAI: An Unreliable Information Store
PodcastImpact of LLMs on the Tech Stack and Product Development
Product Engineering for LLMs
Building RAG-based LLM Applications for Production
PodcastBuilding Effective Products with GenAI
Current State of LLMs in Production
From Building Self-driving Cars to Building LLM Applications
LLMs in Production at GetYourGuide
The Truth About AI Agents
Observability for LLMs
PodcastMLOps vs ML Orchestration
PodcastAll About Evaluating LLM Applications
Fireside Chat - The Future of LLMs
Taming AI Product Development Through Test-driven Prompt Engineering
Using LLMs to Power Consumer Search at Scale
LIMA: Less is More for Alignment
MLOps vs LLMOps
PodcastAll the Hard Stuff with LLMs in Product Development
Everything We've Been Taught About ML is Wrong
Incorporating LLMs in High-stake Use Cases
PodcastMLOps at the Age of Generative AI
Combining LLMs with Knowledge Bases to Prevent Hallucinations
Unleashing Code Completion with LLMs
PodcastExperiment Tracking in the Age of LLMs
Lessons Learned Productionising LLMs for Stripe Support
Building Products
LLMs For the Rest of Us
Building Reliable AI Agents
PodcastTreating Prompt Engineering More Like Code
Evaluating LLM-based Applications
Evaluation
Building and Curating Datasets for RLHF and LLM Fine-tuning
Stopping Hallucinations From Hurting Your LLMs
Building Production Copilots
Embeddings and Retrieval for LLMs: Techniques and Challenges
Pitfalls and Best Practices: 5 Lessons from LLMs in Production
Scalable Evaluation and Serving of Open Source LLMs
The Emerging Toolkit for Reliable, High-quality LLM Applications
LLMs in Production Conference - Part II
Using LLMs to Punch Above Your Weight!
Agentic Relationship Management
Building Defensible Products with LLMs
DevTools for Language Models: Unlocking the Future of AI-Driven Applications
Challenges and Opportunities in Building Data Science Solutions with LLMs
Want High Performing LLMs? Hint: It Is All About Your Data
PodcastManaging Machine Learning Projects
MeetupFrom Expectations to Synthetic Data Generation
PodcastModel Monitoring in Practice: Top Trends
PodcastMLOps at Stripe
PodcastBetter Use Cases for Text Embeddings
PodcastCalibration for ML at Etsy
PodcastBuild a Culture of ML Testing and Model Quality
PodcastML Stepping Stones: Challenges & Opportunities for Companies
PodcastLinkedIn Job Recommendations
MeetupA Missing Link in the ML Infrastructure Stack
PodcastThe Godfather Of MLOps
PodcastContinuous Integration for ML