Watching models in production
How do you tell whether a model still works once people depend on it? Speakers look for answers in drift statistics, production traces and evals, including the disputed practice of asking one model to judge another.
High Stakes ML: Active Failures, Latent Factors
Pushed backFlavio Clesio rejects using large technology companies as the main reliability benchmark for high-stakes systems.14:06
MLOps - The Blind Men and the Elephant
Pushed backSaurav Chakravorty argued that continuous retraining is not necessary in every setup, while continuous monitoring of scoring data is necessary.29:49
Venture Capital in Machine Learning Startups
Pushed backJohn Spindler disputes the common preference for deep learning by saying that linear regression is often the better choice when it fits the problem.18:07
Monitoring the Machine Learning Stack
Pushed backLina Weichbrodt said that real-time response monitoring is needed in addition to offline data-quality checks such as Great Expectations or TensorFlow Data Validation.21:00
Model Watching: Keeping Your Project in Production
Pushed backBen Wilson argues that the tools used for drift monitoring matter less than knowing which kinds of drift and statistical behavior to monitor.18:57
MLOps Investments
Pushed backSarah Catanzaro says industry and academia both contribute to the gap between research and practical ML because industry rarely provides realistic structured-data benchmarks and context.43:06
Building ML Blocks with Kubeflow Orchestration with Feature Store
Pushed backAniruddha Choudhury distinguishes a feature store from a SQL database by emphasizing low-latency online retrieval, feature consistency, and support for batch and streaming ingestion.53:23
mlctl and Hydrosphere Open Source MLOps Libraries Demo
Pushed backAlex Chung questioned whether Hydrosphere should continue serving models or focus on monitoring and integration with existing tools, and recommended the latter focus.31:41
Towards Observability for ML Pipelines
Pushed backShreya Shankar rejects the common practice of monitoring thousands of feature-level KL divergences as the primary way to operate ML systems.18:14
Platform Thinking: A Lemonade Case Study
Pushed backAutomatic model monitoring was not considered suitable for Lemonade's process, so data scientists had to configure monitors manually.14:09
Lessons from Studying FAANG ML Systems
Pushed backErnest Chan pushed back on the idea that shadow mode would necessarily solve the problem of a model facing major COVID-related data drift, saying it might not help the existing model unless that model were turned off.36:39
DataOps is a Software Engineering Challenge
Pushed backMicha Kunze says commercial data-observability tooling was not valuable enough for his team's use case because they needed integrated checks that could stop pipelines, rather than only post hoc metrics.46:09
Want High Performing LLMs? Hint: It Is All About Your Data
Pushed backVikram Chatterji says there is not yet a broadly adequate metric for evaluating prompts across models, while practitioners often default to BLEU scores and similar measures.31:56
Embeddings and Retrieval for LLMs: Techniques and Challenges
Pushed backAnton Troynikov argued that heuristics in robotics and retrieval systems are brittle and that models should eventually handle more of the required tasks.32:54
Evaluating LLM-based Applications
Pushed backJosh Tobin disputes the common view that human evaluation is always the best way to evaluate language models, arguing that automated model evaluation can be useful and that the best approach combines automated and human evaluation.13:34
Evaluation
Pushed backJosh Tobin disputes relying on public benchmarks for application development, saying they are nearly useless when they do not measure a company's user data and outcomes.18:17
Language, Graphs, and AI in Industry
Pushed backPaco Nathan disputes the model-centric focus on benchmark scores without equal attention to data quality, cost, security and domain-specific evaluations.33:10
Pioneering AI Models for Regional Languages
Pushed backDemetrios Brinkmann suggested that large language models may already beat humans on many exams, while Aleksa Gordić disputed the interpretation because evaluation data may have appeared in training data.22:54
LLM Evaluation with Arize AI's Aparna Dhinakaran
Pushed backAparna Dhinakaran argues that binary or multiclass evaluations are more useful than numeric score evaluations because LLM scores often have no reliable meaning on a spectrum.36:50
The Real E2E RAG Stack
Pushed backSam Bean disputes the idea that neural networks should be added to search or evaluation because they automatically make systems simpler.44:18
AI Careers Insights from Ex Meta Staff Eng
Pushed backIlya Reznik disputes the assumption that a high benchmark score demonstrates real-world usefulness.18:18
AI Agents: The Future of ML Engineering?
Pushed backThe speakers questioned whether the benchmark's contamination analysis adequately established that model performance was not inflated by memorization.27:50
Web Agents: The Cutting Edge of AI is Here?
Pushed backThe team found that performance on the WebArena benchmark did not translate reliably to the web-agent tasks they cared about.17:49
AI in Production 2025 | Keynote
Pushed backThe speaker argued that graph RAG does not have a single established best practice and that retrieval strategies require experimentation.44:10
Enterprise AI Operations: The Missing Piece
Pushed backThe cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included.28:48
Real-Time Voice Agents in Production
Pushed backPanos Stravopodis rejected the assumption that human agents necessarily perform better than AI agents and recommended head-to-head benchmarking.14:10
Structured Dissent Patterns for Agentic Production Reliability
Pushed backThe swarm's task-aligned evaluation rated the system highly, while the DPFL benchmark treated the approach as overengineered roleplay because it expected ground truth.15:31
How AI covered a human's paternity leave
Pushed backQuinten Rosseel disputes the common emphasis on text-to-SQL benchmarks as the main measure of agent success, arguing that business context is the real challenge.4:36