Threads / Watching models in production

Watching models in production

How do you tell whether a model still works once people depend on it? Speakers look for answers in drift statistics, production traces and evals, including the disputed practice of asking one model to judge another.

Follows the tags monitoringobservabilityevals · 352 sessions · 2020 to 2026
202027 sessions
Meetup · MLOps Meetup #10

MLOps - The Blind Men and the Elephant

Saurav Chakravorty, Brillo

Pushed backSaurav Chakravorty argued that continuous retraining is not necessary in every setup, while continuous monitoring of scoring data is necessary.29:49

Meetup · MLOps Meetup #16

Venture Capital in Machine Learning Startups

John Spindler, Capital Enterprise

Pushed backJohn Spindler disputes the common preference for deep learning by saying that linear regression is often the better choice when it fits the problem.18:07

Meetup · MLOps Meetup #23

Monitoring the Machine Learning Stack

Lina Weichbrodt, DKB

Pushed backLina Weichbrodt said that real-time response monitoring is needed in addition to offline data-quality checks such as Great Expectations or TensorFlow Data Validation.21:00

23 more from 2020 on this thread
202138 sessions
Podcast · MLOps Coffee Sessions #33

MLOps Investments

Sarah Catanzaro, Amplify Partners

Pushed backSarah Catanzaro says industry and academia both contribute to the gap between research and practical ML because industry rarely provides realistic structured-data benchmarks and context.43:06

Talk · Social Good Tech Working Group 2021

mlctl and Hydrosphere Open Source MLOps Libraries Demo

Alex Chung, Intuit

Pushed backAlex Chung questioned whether Hydrosphere should continue serving models or focus on monitoring and integration with existing tools, and recommended the latter focus.31:41

34 more from 2021 on this thread
202231 sessions
Podcast · MLOps Coffee Sessions #75

Towards Observability for ML Pipelines

Shreya Shankar, UC Berkeley

Pushed backShreya Shankar rejects the common practice of monitoring thousands of feature-level KL divergences as the primary way to operate ML systems.18:14

Podcast · MLOps Coffee Sessions #79

Platform Thinking: A Lemonade Case Study

Orr Shilon, Lemonade

Pushed backAutomatic model monitoring was not considered suitable for Lemonade's process, so data scientists had to configure monitors manually.14:09

Podcast · MLOps Coffee Sessions #84

Lessons from Studying FAANG ML Systems

Ernest Chan, Duo Security

Pushed backErnest Chan pushed back on the idea that shadow mode would necessarily solve the problem of a model facing major COVID-related data drift, saying it might not help the existing model unless that model were turned off.36:39

Meetup · MLOps Meetup #100

DataOps is a Software Engineering Challenge

Micha Kunze, Maersk

Pushed backMicha Kunze says commercial data-observability tooling was not valuable enough for his team's use case because they needed integrated checks that could stop pipelines, rather than only post hoc metrics.46:09

27 more from 2022 on this thread
202365 sessions
Talk · LLMs in Production 2023

Evaluating LLM-based Applications

Josh Tobin, Gantry

Pushed backJosh Tobin disputes the common view that human evaluation is always the best way to evaluate language models, arguing that automated model evaluation can be useful and that the best approach combines automated and human evaluation.13:34

Talk · LLMs in Production 2023

Evaluation

Amrutha Gujjar, Structured & Josh Tobin, Gantry & Sohini Roy, NVIDIA

Pushed backJosh Tobin disputes relying on public benchmarks for application development, saying they are nearly useless when they do not measure a company's user data and outcomes.18:17

61 more from 2023 on this thread
202480 sessions
Podcast · MLOps Podcast #201

Language, Graphs, and AI in Industry

Paco Nathan, Derwen, Inc.

Pushed backPaco Nathan disputes the model-centric focus on benchmark scores without equal attention to data quality, cost, security and domain-specific evaluations.33:10

Podcast · MLOps Podcast #203

Pioneering AI Models for Regional Languages

Aleksa Gordić, OrtusAI

Pushed backDemetrios Brinkmann suggested that large language models may already beat humans on many exams, while Aleksa Gordić disputed the interpretation because evaluation data may have appeared in training data.22:54

Podcast · MLOps Podcast #210

LLM Evaluation with Arize AI's Aparna Dhinakaran

Arize AI's Aparna Dhinakaran

Pushed backAparna Dhinakaran argues that binary or multiclass evaluations are more useful than numeric score evaluations because LLM scores often have no reliable meaning on a spectrum.36:50

Podcast · MLOps Podcast #217

The Real E2E RAG Stack

Sam Bean, Rewind.ai

Pushed backSam Bean disputes the idea that neural networks should be added to search or evaluation because they automatically make systems simpler.44:18

76 more from 2024 on this thread
202569 sessions
Reading group · MLOps Reading Group

AI Agents: The Future of ML Engineering?

Matt Squire, Fuzzy Labs

Pushed backThe speakers questioned whether the benchmark's contamination analysis adequately established that model performance was not inflated by memorization.27:50

Podcast · Agents in Production Series

Web Agents: The Cutting Edge of AI is Here?

Paul van der Boor & Chiara Caratelli, Prosus Group

Pushed backThe team found that performance on the WebArena benchmark did not translate reliably to the web-agent tasks they cared about.17:49

Talk · AI in Production 2025

AI in Production 2025 | Keynote

Pushed backThe speaker argued that graph RAG does not have a single established best practice and that retrieval strategies require experimentation.44:10

65 more from 2025 on this thread
202642 sessions
Podcast · MLOps Podcast #345

Enterprise AI Operations: The Missing Piece

Rani Radhakrishnan, PwC US

Pushed backThe cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included.28:48

Talk · Agents in Production 2026

Real-Time Voice Agents in Production

Panos Stravopodis, Elyos AI

Pushed backPanos Stravopodis rejected the assumption that human agents necessarily perform better than AI agents and recommended head-to-head benchmarking.14:10

Talk · Coding Agents Conference 2026

How AI covered a human's paternity leave

Quinten Rosseel, Wobby

Pushed backQuinten Rosseel disputes the common emphasis on text-to-SQL benchmarks as the main measure of agent success, arguing that business context is the real challenge.4:36

38 more from 2026 on this thread