Humans in the loop in 2025

24 sessions

PodcastEvolving Workflow OrchestrationAlex Milowski, Entrepreneur and Computer Scientist 路 1:14:35 路 Feb 2025 路 425 views 路 MLOps Podcast

ClaimThe current trend in data science and machine learning is moving away from separate workflow-definition languages toward Python code and annotations.31:41

PodcastLook At Your ****ing Data 馃憖Kenny Daniel, Hyperparam 路 1:05:26 路 Feb 2025 路 292 views 路 MLOps Podcast

Pushed backKenny argued that specialized expert data work should not generally be outsourced to generic labeling companies when domain expertise is required.19:20

Which Economic Tasks are Performed with AI? Evidence from Millions of Claude ConversationsValdimar Eggertsson, Snjallg枚gn (Smart Data inc.) & Sophia Skowronski, Breckinridge Capital Advisors 路 55:09 路 Apr 2025 路 234 views 路 MLOps Community Reading Group 2025

Pushed backAdam Becker disputes treating debugging an error as fully automating a person's job, because the task can still require human context and judgment.26:46

PodcastAI-Powered Product Ideation with Synthetic Consumer TestingLuca Fiaschi, PyMC Labs 路 1:00:44 路 Apr 2025 路 281 views 路 MLOps Podcast

ClaimThe workflow still requires a human with sufficient business context to determine whether the data and outputs are sensible.17:09

PodcastHow Sama is Improving ML Models to Make AVs SaferDuncan Curtis, Sama 路 45:35 路 Apr 2025 路 282 views 路 MLOps Podcast

ClaimSama combines its technology platform with human intelligence to capture the information that AI models need while reducing unnecessary annotation work.0:50

PodcastAI Data Engineers: Data Engineering After AIVikram Chennai, Ardent AI 路 48:07 路 Apr 2025 路 345 views 路 MLOps Podcast

ClaimThe agent is not currently fully autonomous because human approval remains important when changes could affect production systems.25:29

PodcastMaking AI Reliable is the Greatest Challenge of the 2020sAlon Bochman, RagMetrics 路 1:01:38 路 May 2025 路 213 views 路 MLOps Podcast

Pushed backAlon Bochman argues that a jury of multiple LLM judges should be adopted only if it produces a higher human agreement rate on the team's task.48:55

Building an AI agent with LangGraph, step by step tutorial 路 10:39 路 May 2025 路 1,643 views

ClaimThe free Intro to LangGraph course from LangGraph Academy covers agent state, memory, and human-in-the-loop integration.1:18

MCP is not going to change everything (yet)Sam Partee, Arcade AI & Rahul Parundekar, AI Hero 路 1:04:43 路 May 2025 路 552 views
How Product Metrics Become LLM EvaluationsRaza Habib, Humanloop 路 53:07 路 Jun 2025 路 529 views

ClaimHumanloop has three core pillars: production tracing and observability, development datasets and evaluators, and tools for managing prompts and collaboration.16:48

PodcastKnowledge is Eventually ConsistentDevin Stein, Dosu 路 55:15 路 Aug 2025 路 390 views 路 MLOps Podcast

Pushed backDemetrios Brinkmann challenges whether asking experts to explicitly approve and save facts is too much work.18:04

How AI Will Transform The Energy SectorAdam Sroka, Hypercube 路 23:13 路 Aug 2025 路 198 views 路 Agents in Production 2025

ClaimAdam Sroka says Jellyfish keeps human approval gates in place until its change-classification systems reach a high level of confidence.11:28

I Built A Trustworthy Voice Assistant 路 12:17 路 Aug 2025 路 142 views 路 Agents in Production 2025

ClaimA voice agent listens, understands, decides, and responds in natural language, turning human speech into structured, actionable interactions.0:48

Iterating on Your AI EvalsMariana Prazeres 路 13:47 路 Aug 2025 路 162 views 路 Agents in Production 2025

Pushed backHuman-in-the-loop evaluation may be introduced too late in the advanced-system example.12:22

Fast, Trustworthy, Reliable Voice Agents: MLOps That Blend LLM Annotation with Human QAErik Goron, HappyRobot 路 17:08 路 Aug 2025 路 145 views

ClaimHappyRobot combines human annotations with LLM-generated annotations and aligns LLM annotator prompts with human signals to create accurate datasets across use cases.6:46

Too much lock-in for too little gain: agent frameworks are a dead-endValliappa Lakshmanan 路 35:37 路 Aug 2025 路 259 views

ClaimAgentic systems should use human input as a learning path toward greater autonomy rather than beginning as fully autonomous systems.2:29

If There's Free Compute, There's Abuse: Fighting Fraud with Lightweight LLM AgentsJonas Scholz, Sliplane 路 15:07 路 Aug 2025 路 89 views

ClaimA human remains responsible for the final decision because the agents still produce too many false positives to ban users automatically.6:32

PodcastThe Era of AI Agents in MarketingJoel Horwitz, Neoteric3D 路 48:57 路 Sept 2025 路 195 views 路 MLOps Podcast

Pushed backAI-generated thought leadership can replace authentic human thought leadership and relationship building.40:05

Catastrophic agent failure and how to avoid itEdward Upton, Asteroid 路 25:21 路 Sept 2025 路 223 views 路 Agents in Production 2025

Pushed backHumans cannot be assumed to provide a perfect solution for avoiding false positives, because human labeling itself can contain mistakes.19:44

Designing AI Agents for the Complex Realities of HealthcareDr. Sarah Gebauer, Validara Health 路 16:03 路 Sept 2025 路 434 views

ClaimAI agents in healthcare need human oversight, escalation at critical junctures, and appropriate communication with team members.2:55

Underwriting Assist: A Multi-Agent SystemSomya Rai, EXL 路 33:27 路 Sept 2025 路 198 views 路 Agents in Production 2025

ClaimHuman review is included throughout the underwriting workflow because quoting and decisioning require accurate information, compliance, and auditability.5:28

Why You Should Care About Observability in LLM WorkflowsColin McNamera, AlwaysCool.ai 路 15:03 路 Sept 2025 路 188 views

ClaimThe team used custom GPTs, a Python adapter, and FDA and USDA interfaces to generate nutrition facts labels more accurately.2:52

Evals Aren't Useful? Really?Chiara Caratelli, Prosus Group 路 25:25 路 Oct 2025 路 482 views

ClaimWhen an evaluation exposes a problem, the fix may involve prompting, code changes for security issues, or a reviewer that checks the output before returning it.5:23

Fine-Tuned Models Are Getting Out of HandJaipal Singh Goud, Prem AI 路 36:48 路 Nov 2025 路 586 views

ClaimAI systems are low-trust systems operating in high-trust environments, so interaction design and human-like behavior matter alongside model quality.4:46