Reading groupAI REWIND 2025 - MLOps Reading Group Year-end SpecialClaimStatic benchmarks can become outdated because models may memorize them and because they do not represent dynamic real-world environments.42:39
42 sessions
Reading groupAI REWIND 2025 - MLOps Reading Group Year-end SpecialClaimStatic benchmarks can become outdated because models may memorize them and because they do not represent dynamic real-world environments.42:39
Building Agentic Tools for ProductionClaimProduction agentic tools need generated and carefully maintained tool schemas, per-tool authorization requirements, and tool evaluations.1:12
PodcastEnterprise AI Operations: The Missing PiecePushed backThe cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included.28:48
Inside OpenAI's AI Agent Collaboration SystemClaimEvals are structured, objective ways to measure an agent's or model's performance, while graders evaluate the model's outputs, reasoning, and justification.1:06
Tool definitions are the new Prompt EngineeringClaimAgent evaluations should connect to business value and include user-experience elements rather than relying only on generic helpfulness or faithfulness metrics.21:03
MCP Security: What Happens When Your Agents Talk to Everything?ClaimContext-aware permissions evaluate an action against the request context and policy rules instead of deciding only whether an agent has access to a tool.11:21
Multi-Agent Systems for the Misinformation LifecycleClaimA single retrieval-augmented generation system is a useful baseline, but user-generated and generative-AI content with difficult ground truth requires a more sophisticated system.6:12
Real-Time Voice Agents in ProductionPushed backPanos Stravopodis rejected the assumption that human agents necessarily perform better than AI agents and recommended head-to-head benchmarking.14:10
Structured Dissent Patterns for Agentic Production ReliabilityPushed backThe swarm's task-aligned evaluation rated the system highly, while the DPFL benchmark treated the approach as overengineered roleplay because it expected ground truth.15:31
Context Engineering pitfalls for our e-commerce agentClaimChiara Carateli says improving tool and variable names helped the team remove many edge cases from the system prompt while preserving evaluations for those cases.14:59
Feedback Loops for Agentic WorkflowsClaimGood feedback loops improve agents because they make behavior observable, allow mistakes to be caught early, and let agents learn from their outputs.7:48
From Notebooks to Production FASTERClaimThe team uses Evidently AI to monitor models and data quality, including drift and performance degradation, through reports and dashboards.6:46
Co-Engineering: The New Era of Human-AI CollaborationClaimThe main principles of co-engineering are to build context, stay opinionated, improve iteratively, make agents observable, and make them accountable.6:05
Stop Building AI Like Traditional SoftwareClaimAI evaluations should use task-specific measures and account for non-deterministic behavior rather than relying only on traditional tests.8:15
Why Emotion Matters More Than SoundClaimA speech-to-speech system should remain observable and interpretable so developers can determine whether a failure came from recognition, reasoning, or synthesis.19:49
PodcastSpeed and Scale: How Today's AI Datacenters Are Operating Through HypergrowthClaimOperational automation requires combining real-time observability of physical and logical systems with the intended design so that deviations can be diagnosed and corrected.57:20
Beyond the Gold Standard: Evaluating and Trusting Agents in the WildClaimSanjana Sharma says organizations should shift from model-first thinking to system-first thinking, with explicit context, rules and evaluation practices.3:33
How AI covered a human's paternity leavePushed backQuinten Rosseel disputes the common emphasis on text-to-SQL benchmarks as the main measure of agent success, arguing that business context is the real challenge.4:36
Agents as Search EngineersClaimAgentic search treats retrieval as part of a stateful control loop involving reasoning strategies, tool calls, and memory rather than as a final endpoint.3:27
Building an Orchestration Layer for Agentic Commerce at LoblawsClaimAlfred manages privacy by removing personally identifiable information from messages sent to model providers and masking it in observability data.6:16
PodcastCracking the Black Box: Real-Time Neuron Monitoring & Causality TracesClaimMike Oaten says regulatory observability focuses on risk, including identifying prohibited and high-risk AI systems and deciding what risks must be tracked.6:07
PodcastMLflow Leading Open SourceClaimMulti-turn evaluation is important because issues such as repetition, escalation to human support, coherence, and context retention often cannot be detected from a single input-output pair.7:18
Simulate to Scale: How realistic simulations power reliable agents in productionClaimAgent tests should evaluate whether the user achieved their goal rather than require one specific sequence or response.3:23
A Playground for AI EngineersClaimClassical models require continuous training, evaluation, and monitoring for data drift and concept drift.7:33
Open vs Closed Source Agent Infra?Pushed backAdel rejected the idea that NVIDIA's toolkit locks teams into one framework, saying it emits standard traces and allows teams to keep their existing observability platforms.18:11
Using Agents in Production: Past Present and FutureClaimProsus evaluates agents across productivity, quality, agility, and independence, with implementation depending on the specific use case.22:31
Stop Shipping on Vibes: How to Build Real Evals for Coding AgentsPushed backJessica Wang challenged the conclusion that the evaluation proved agentic search was universally better than vector search.22:34
MCP Dev Summit [Day 1]ClaimDiamond Bishop argues that agents should be evaluated offline and online with a continuously maintained evaluation system before they are launched.2:32:58
The Coding Agent Multiverse of MadnessClaimThe coding agent gateway is intended to give developers freedom to use their preferred tools while giving enterprise administrators centralized security, procurement, observability, and cost controls.7:27
Ship Agents: A Virtual Conference Track 2ClaimAgents can silently lose important context, drift away from their original goal, produce confident but unsupported answers, and compound errors across steps.43:26
PodcastIt's 2026, and We're Still Talking EvalsClaimPre-production evals protect a team from shipping a product that performs badly, while production evals measure whether the product delivers quality with real users.2:00
PodcastWhy Agents are Driving Software Development to the CloudClaimA scalable agent system needs audit traces, handoff, granular access control, memory, evaluations, flexible deployment, programmability, and observability.28:04
Architecting Modern AI SystemsPushed backAlan argued that manual evaluation and expert review are needed rather than relying only on LLM-based judging.25:37
Architecting Modern AI Systems: Platforms, Agents, and IntegrationClaimThe hackathon evaluation infrastructure ran on Kubernetes in the BuzzHPC environment and provided access to hosted LLMs, GPUs, and CPUs.3:07
PodcastAI Is Fast. AI Projects Are Slow. Let's Fix That.ClaimRocketRide provides full traces of agent requests, tool inputs, tool results, and intermediate steps so developers can identify where a pipeline fails.16:54
PodcastLogs Are All You Need: Rethinking Observability with AI AgentsPushed backSherwood Callaway disputed the traditional view that metrics, logs, and traces are all required for observability.5:23
Reading groupLoop EngineeringClaimArthur Coleman found that Claude could break existing functionality while implementing a new change, so he added regression testing early in the process.11:47
MCPs for Observability StacksClaimDevelopers want metrics, logs, traces, and events brought together so they can get complete system context.1:31
Responsible Autonomy: Building Governance Frameworks for AI That Act in the Real World via MCPClaimSaurabh Mishra says agent observability can track token usage, latency, cost, error rate, and business key performance indicators.23:59