Watching models in production in 2026

42 sessions

Reading groupAI REWIND 2025 - MLOps Reading Group Year-end SpecialSophia Skowronski, Breckinridge Capital Advisors & Adam Becker, MLOps Community & Rohan Prasad, EvolutionIQ & Nehil Jain, Stealth AI Startup & Sonam Gupta, AICamp & Lucas Pavanelli, Stone · 2:06:08 · Jan 2026 · 206 views · MLOps Reading Group

ClaimStatic benchmarks can become outdated because models may memorize them and because they do not represent dynamic real-world environments.42:39

Building Agentic Tools for ProductionSam Partee, Arcade AI · 23:55 · Jan 2026 · 415 views · Agents in Production 2025

ClaimProduction agentic tools need generated and carefully maintained tool schemas, per-tool authorization requirements, and tool evaluations.1:12

PodcastEnterprise AI Operations: The Missing PieceRani Radhakrishnan, PwC US · 41:28 · Jan 2026 · 279 views · MLOps Podcast

Pushed backThe cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included.28:48

Inside OpenAI's AI Agent Collaboration System · 19:22 · Jan 2026 · 201 views

ClaimEvals are structured, objective ways to measure an agent's or model's performance, while graders evaluate the model's outputs, reasoning, and justification.1:06

Tool definitions are the new Prompt EngineeringChiara Caratelli, Prosus Group & Alex Salazar, Arcade.dev · 57:12 · Jan 2026 · 379 views

ClaimAgent evaluations should connect to business value and include user-experience elements rather than relying only on generic helpfulness or faithfulness metrics.21:03

MCP Security: What Happens When Your Agents Talk to Everything? · 24:26 · Jan 2026 · 155 views · Agents in Production 2025

ClaimContext-aware permissions evaluate an action against the request context and policy rules instead of deciding only whether an agent has access to a tool.11:21

Multi-Agent Systems for the Misinformation Lifecycle · 28:27 · Jan 2026 · 190 views · Agents in Production 2026

ClaimA single retrieval-augmented generation system is a useful baseline, but user-generated and generative-AI content with difficult ground truth requires a more sophisticated system.6:12

Real-Time Voice Agents in ProductionPanos Stravopodis, Elyos AI · 42:16 · Jan 2026 · 439 views · Agents in Production 2026

Pushed backPanos Stravopodis rejected the assumption that human agents necessarily perform better than AI agents and recommended head-to-head benchmarking.14:10

Structured Dissent Patterns for Agentic Production ReliabilityPhil Stafford, MLOps Community · 28:13 · Jan 2026 · 181 views · Agents in Production 2025

Pushed backThe swarm's task-aligned evaluation rated the system highly, while the DPFL benchmark treated the approach as overengineered roleplay because it expected ground truth.15:31

Context Engineering pitfalls for our e-commerce agentNishikant Dhanuka & Chiara Carateli, Prosus · 28:06 · Jan 2026 · 401 views · Agents in Production 2025

ClaimChiara Carateli says improving tool and variable names helped the team remove many edge cases from the system prompt while preserving evaluations for those cases.14:59

Feedback Loops for Agentic Workflows · 21:16 · Jan 2026 · 827 views · AI Agent World Tour 2025

ClaimGood feedback loops improve agents because they make behavior observable, allow mistakes to be caught early, and let agents learn from their outputs.7:48

From Notebooks to Production FASTERShahd Alghrsi, Virgin Media · 13:43 · Jan 2026 · 266 views

ClaimThe team uses Evidently AI to monitor models and data quality, including drift and performance degradation, through reports and dashboards.6:46

Co-Engineering: The New Era of Human-AI CollaborationKiriti Badam, OpenAI · 29:25 · Jan 2026 · 541 views

ClaimThe main principles of co-engineering are to build context, stay opinionated, improve iteratively, make agents observable, and make them accountable.6:05

Stop Building AI Like Traditional SoftwareAishwarya Naresh Reganti, LevelUp Labs · 20:47 · Jan 2026 · 425 views

ClaimAI evaluations should use task-specific measures and account for non-deterministic behavior rather than relying only on traditional tests.8:15

Why Emotion Matters More Than SoundAnoop Dawar, Deepgram & Ajeet Grewal, Sierra · 26:44 · Feb 2026 · 109 views

ClaimA speech-to-speech system should remain observable and interpretable so developers can determine whether a failure came from recognition, reasoning, or synthesis.19:49

PodcastSpeed and Scale: How Today's AI Datacenters Are Operating Through HypergrowthKris Beevers, NetBox Labs · 1:07:17 · Feb 2026 · 171 views · MLOps Podcast

ClaimOperational automation requires combining real-time observability of physical and logical systems with the intended design so that deviations can be diagnosed and corrected.57:20

Beyond the Gold Standard: Evaluating and Trusting Agents in the WildSanjana Sharma, Prosus · 24:45 · Feb 2026 · 199 views · Agents in Production 2025

ClaimSanjana Sharma says organizations should shift from model-first thinking to system-first thinking, with explicit context, rules and evaluation practices.3:33

The Future of Coding: AI Agents & the Next Tech RevolutionRicky Doar, Cursor · 26:45 · Feb 2026 · 321 views · Coding Agents Conference 2026
How AI covered a human's paternity leaveQuinten Rosseel, Wobby · 51:59 · Feb 2026 · 110 views · Coding Agents Conference 2026

Pushed backQuinten Rosseel disputes the common emphasis on text-to-SQL benchmarks as the main measure of agent success, arguing that business context is the real challenge.4:36

Agents as Search EngineersSantoshkalyan Rayadhurgam, Meta · 29:38 · Feb 2026 · 169 views · Coding Agents Conference 2026

ClaimAgentic search treats retrieval as part of a stateful control loop involving reasoning strategies, tool calls, and memory rather than as a final endpoint.3:27

Building an Orchestration Layer for Agentic Commerce at LoblawsMefta Sadat, Loblaw Digital · 25:15 · Feb 2026 · 366 views · Agents in Production 2025

ClaimAlfred manages privacy by removing personally identifiable information from messages sent to model providers and masking it in observability data.6:16

PodcastCracking the Black Box: Real-Time Neuron Monitoring & Causality TracesMike Oaten, TIKOS · 47:26 · Feb 2026 · 76 views · MLOps Podcast

ClaimMike Oaten says regulatory observability focuses on risk, including identifying prohibited and high-risk AI systems and deciding what risks must be tracked.6:07

PodcastMLflow Leading Open SourceDatabricks' Corey Zumar · 58:24 · Feb 2026 · 324 views · MLOps Podcast

ClaimMulti-turn evaluation is important because issues such as repetition, escalation to human support, coherence, and context retention often cannot be detected from a single input-output pair.7:18

Simulate to Scale: How realistic simulations power reliable agents in productionSachi Shah, Sierra · 21:10 · Feb 2026 · 201 views · Agents in Production 2026

ClaimAgent tests should evaluate whether the user achieved their goal rather than require one specific sequence or response.3:23

A Playground for AI EngineersPaulo Vasconcellos, Hotmart · 54:42 · Feb 2026 · 191 views · Coding Agents Conference 2026

ClaimClassical models require continuous training, evaluation, and monitoring for data drift and concept drift.7:33

Open vs Closed Source Agent Infra?Adel El Hallak, NVIDIA · 30:45 · Feb 2026 · 108 views · Coding Agents Conference 2026

Pushed backAdel rejected the idea that NVIDIA's toolkit locks teams into one framework, saying it emits standard traces and allows teams to keep their existing observability platforms.18:11

Using Agents in Production: Past Present and FutureEuro Beinat, Prosus · 23:12 · Mar 2026 · 498 views · Agents in Production 2026

ClaimProsus evaluates agents across productivity, quality, agility, and independence, with implementation depending on the specific use case.22:31

Stop Shipping on Vibes: How to Build Real Evals for Coding AgentsJessica Wang, Braintrust · 29:09 · Mar 2026 · 26K views · Coding Agents Conference 2026

Pushed backJessica Wang challenged the conclusion that the evaluation proved agentic search was universally better than vector search.22:34

MCP Dev Summit [Day 1]Shannon Williams, Obot AI & Jim Zemlin, Linux Foundation & David Soria Parra, Anthropic & David Nalley, AWS & James Hood, Amazon Web Services & Magna Sumasandra & Rash Tini, Uber & Sheng Liang, Obot AI & Aaron Wang, Duolingo & Adam Seligman & Zayn Turner, Workato & Diamond Bishop, Datadog & Nick Aldridge, Mousetrap & Alex Salazar, Arcade.dev & Jake Diamond Arivich, Jupyter & Kiierra Dodson, Further & Daniel Abdel Samid, Apollo & Juan Antonio Oz, Stacklok & Alharith Hussin, Alterion & Rick Nucci, Guru & Jonathan Rochelle, Lutely & Harshul Jain, Audible & Abhishek Khanna, Blueflame AI & Du'An Lightfoot, AWS & Lin Sun, Solo.io & Saurabh Yergattikar, eBay & Sanjay Vakil, DirectBooker & Jonathan Freeland, Shashank Khanna & Hillary Curran & Cecilia Liu, Docker & Diamond Bishop, Datadog & Paul Carleton, Anthropic · 7:14:47 · Apr 2026 · 8,135 views · MCP Dev Summit 2026

ClaimDiamond Bishop argues that agents should be evaluated offline and online with a continuously maintained evaluation system before they are launched.2:32:58

The Coding Agent Multiverse of MadnessAnkit Mathur, Databricks · 27:52 · Apr 2026 · 3,194 views · Coding Agents Conference 2026

ClaimThe coding agent gateway is intended to give developers freedom to use their preferred tools while giving enterprise administrators centralized security, procurement, observability, and cost controls.7:27

Ship Agents: A Virtual Conference Track 2Adam Boaz Becker & Sarmad Absil, Trial Cyber & Divia Mahajan, Amazon Alexa · 1:38:26 · Apr 2026 · 261 views · Ship Agents 2026

ClaimAgents can silently lose important context, drift away from their original goal, produce confident but unsupported answers, and compound errors across steps.43:26

PodcastIt's 2026, and We're Still Talking EvalsMaggie Konstanty, Prosus · 40:57 · Apr 2026 · 389 views · MLOps Podcast

ClaimPre-production evals protect a team from shipping a product that performs badly, while production evals measure whether the product delivers quality with real users.2:00

PodcastWhy Agents are Driving Software Development to the CloudZach Lloyd, Warp · 51:08 · Apr 2026 · 841 views · MLOps Podcast

ClaimA scalable agent system needs audit traces, handoff, granular access control, memory, evaluations, flexible deployment, programmability, and observability.28:04

PodcastGetting Humans Out of the Way: How to Work with Teams of AgentsRob Ennals, Broomy · 50:31 · May 2026 · 395 views · MLOps Podcast
Architecting Modern AI Systems · 56:58 · May 2026 · 261 views

Pushed backAlan argued that manual evaluation and expert review are needed rather than relying only on LLM-based judging.25:37

Architecting Modern AI Systems: Platforms, Agents, and IntegrationAllen Roush, BuzzHPC & Frédéric Bénard, Mila & Shuo Wang, Bell Canada · 57:00 · May 2026 · 343 views · BuzzHPC Roundtable 2026

ClaimThe hackathon evaluation infrastructure ran on Kubernetes in the BuzzHPC environment and provided access to hosted LLMs, GPUs, and CPUs.3:07

PodcastAI Is Fast. AI Projects Are Slow. Let's Fix That.JRocketRide's Joe Maionchi · 56:48 · Jun 2026 · 293 views · MLOps Podcast

ClaimRocketRide provides full traces of agent requests, tool inputs, tool results, and intermediate steps so developers can identify where a pipeline fails.16:54

PodcastLogs Are All You Need: Rethinking Observability with AI AgentsSherwood Callaway, Sazabi · 46:40 · Jun 2026 · 660 views · MLOps Podcast

Pushed backSherwood Callaway disputed the traditional view that metrics, logs, and traces are all required for observability.5:23

Coding Agents Are Secretly General AgentsJay Hack, ClickUp · 1:12:03 · Jul 2026 · 243 views
Reading groupLoop EngineeringDavid DeStefano & Sam Christensen, EvolutionIQ & Valdimar Eggertsson, Snjallgögn (Smart Data inc.) & Sparsh Jain, CentralAgent AI · 56:10 · Aug 2026 · 384 views · MLOps Reading Group

ClaimArthur Coleman found that Claude could break existing functionality while implementing a new change, so he added regression testing early in the process.11:47

MCPs for Observability StacksDiana Todea, VictoriaMetrics · 24:28 · Aug 2026 · 101 views

ClaimDevelopers want metrics, logs, traces, and events brought together so they can get complete system context.1:31

Responsible Autonomy: Building Governance Frameworks for AI That Act in the Real World via MCPSaurabh Mishra, Optum · 27:55 · Aug 2026 · 115 views

ClaimSaurabh Mishra says agent observability can track token usage, latency, cost, error rate, and business key performance indicators.23:59