# Humans in the loop

82 sessions · follows the tags human-in-the-loop
Page: https://mlopstalks.com/threads/humans-in-the-loop

People label the training data and review LLM output. As agents take on more work, teams have to decide which actions still need a person's approval.

## 2020

2 sessions.

- [Scaling Human-in-the-Loop Machine Learning](https://mlopstalks.com/talks/scaling-human-in-the-loop-machine-learning) (Robert Munro). Pushed back: Robert Munro rejected the idea that data annotation is merely the uninteresting part of machine learning work. [10:56](https://www.youtube.com/watch?v=LwbbGsuNpao&t=656s)
- [Creating Beautiful Ambient Music with Google Brain's Music Transformer](https://mlopstalks.com/talks/creating-beautiful-ambient-music-with-google-brains-music-transformer) (Daniel Jeffries, Pachyderm). Pushed back: Daniel Jeffries disputes the idea that artificial intelligence must either replace humans or be replaced by humans, arguing instead for human-machine collaboration. [7:17](https://www.youtube.com/watch?v=z95ciIqMuRo&t=437s)

## 2021

6 sessions.

- [Product Management in Machine Learning](https://mlopstalks.com/talks/product-management-in-machine-learning) (Laszlo Sragner, Hypergolic). Claim: Laszlo Sragner says monitoring, evaluation, and labeling are tightly coupled and should happen as one continuous process for machine learning models. [9:51](https://www.youtube.com/watch?v=Sl7WrlbXf9E&t=591s)
- [Data Selection for Data-Centric AI: Data Quality Over Quantity](https://mlopstalks.com/talks/data-selection-for-data-centric-ai-data-quality-over-quantity) (Cody Coleman). Pushed back: Cody Coleman argues against using all available data by default, because noisy data and labels increase cost and difficulty. [41:11](https://www.youtube.com/watch?v=v7Pj7a6KXSU&t=2471s)
- [Data-Centric AI Means Centralizing Training Data](https://mlopstalks.com/talks/data-centric-ai-means-centralizing-training-data) (Alberto Rizzoli, V7). Pushed back: Alberto Rizzoli disputes the usefulness of treating ImageNet as a fully reliable benchmark because some classes contain substantial labeling errors. [13:47](https://www.youtube.com/watch?v=DFXVIE8GRF8&t=827s)
- [The Future of AI and ML in Process Automation](https://mlopstalks.com/talks/the-future-of-ai-and-ml-in-process-automation) (Slater Victoroff, Indico Data). Claim: Slater Victoroff says changing an OCR engine can invalidate labels when labels are stored only as positions in extracted text. [23:25](https://www.youtube.com/watch?v=w40RyIDYzkA&t=1405s)

2 more from 2021 on this thread: https://mlopstalks.com/threads/humans-in-the-loop/2021

## 2022

6 sessions.

- [Applications of Data Science](https://mlopstalks.com/talks/applications-of-data-science) (Connie Yang, Pallet). Claim: Pallet built a multi-label, multi-class machine-learning classifier to automate job-data classification and labeling in its backend. [13:58](https://www.youtube.com/watch?v=rc8zLY15WZU&t=838s)
- [MLOps Critiques](https://mlopstalks.com/talks/mlops-critiques) (Matthijs Brouns, Xccelerated.io). Claim: Matthijs Brouns says teams can automate model retraining and produce a report for human approval without automating every model-release decision. [40:56](https://www.youtube.com/watch?v=SS2_jQN3sG0&t=2456s)
- [Labeled Datasets that Correct Themselves Automatically](https://mlopstalks.com/talks/labeled-datasets-that-correct-themselves-automatically) (Curtis Northcutt, Cleanlab). Pushed back: Automated data cleaning should not be treated as a completely human-free process, because high-quality corrections still require human involvement. [37:46](https://www.youtube.com/watch?v=IwDGDAHgzAY&t=2266s)
- [Creative AI: Using ML to Create Art, Music, and Jokes](https://mlopstalks.com/talks/creative-ai-using-ml-to-create-art-music-and-jokes) (Suyash Joshi, MLOps Community). Pushed back: Suyash Joshi says the copyright claim for an autonomously AI-generated image was rejected because human authorship was required. [17:32](https://www.youtube.com/watch?v=IIAfK9a4l6c&t=1052s)

2 more from 2022 on this thread: https://mlopstalks.com/threads/humans-in-the-loop/2022

## 2023

21 sessions.

- [LLMs as Intelligent Assistants](https://mlopstalks.com/talks/llms-as-intelligent-assistants) (Sarah Aerni, Salesforce). Pushed back: Sarah Aerni rejects replacing human review with fully autonomous generated output and argues that human review remains critical. [23:40](https://www.youtube.com/watch?v=E0929WqB72k&t=1420s)
- [MLOps at the Age of Generative AI](https://mlopstalks.com/talks/mlops-at-the-age-of-generative-ai) (Barak Turovsky, Scale Venture Partners). Pushed back: Barak Turovsky disputes the idea that companies can simply add a large language model on top of an existing tool or replace staff without redesigning processes, handling hallucinations, and adding human exception handling. [46:58](https://www.youtube.com/watch?v=lrf0V4X2dzM&t=2818s)
- [Automating Data Annotation with LLMs](https://mlopstalks.com/talks/automating-data-annotation-with-llms) (Nikolai Liubimov, Michael Malyuk & Chris Hoge, HumanSignal). Pushed back: Human annotation is described as the gold standard, while large foundation models are largely created from unlabeled data, creating a tension between the two approaches. [7:12](https://www.youtube.com/watch?v=mTVNE0Sw5vI&t=432s)
- [LLMs in Production at GetYourGuide](https://mlopstalks.com/talks/llms-in-production-at-getyourguide) (Meghana Satish & Tina Treimane, GetYourGuide). Pushed back: The GPT-4 evaluator was not consistently better than human evaluators, although it sometimes detected missed information and false positives. [20:58](https://www.youtube.com/watch?v=3GEz_0ddFIo&t=1258s)

17 more from 2023 on this thread: https://mlopstalks.com/threads/humans-in-the-loop/2023

## 2024

16 sessions.

- [Language, Graphs, and AI in Industry](https://mlopstalks.com/talks/language-graphs-and-ai-in-industry) (Paco Nathan, Derwen, Inc.). Claim: Paco Nathan says AI applications in regulated or industrial settings need software engineering, operations, security, legal review and domain expertise rather than only an API call. [1:00:42](https://www.youtube.com/watch?v=zDSGctGdB2A&t=3642s)
- [Turn Data Chaos into AI Strategy with Programmatic AI Data Development](https://mlopstalks.com/talks/turn-data-chaos-into-ai-strategy-with-programmatic-ai-data-development) (Elena Boiarskaia, Snorkel AI). Pushed back: Programmatic labeling functions should not be used directly as the final inference rules because a rule-based system is not robust or scalable enough for new data. [22:35](https://www.youtube.com/watch?v=mqHK6DuDrxE&t=1355s)
- [The Future of Healthcare: AI is Here](https://mlopstalks.com/talks/the-future-of-healthcare-ai-is-here) (). Pushed back: Shaun disputes the idea that AI should merely match human performance, saying HeyRevia's results show its agents can outperform humans in comparable healthcare phone-call scenarios. [23:14](https://www.youtube.com/watch?v=GhJZrq1GycM&t=1394s)
- [Why Planning is the New Search](https://mlopstalks.com/talks/why-planning-is-the-new-search) (). Pushed back: Fabian disputes the assumption that agentic workflows should immediately be fully self-driving, arguing that customers generally want human oversight and gradual automation first. [26:11](https://www.youtube.com/watch?v=_U5ovDn50mw&t=1571s)

12 more from 2024 on this thread: https://mlopstalks.com/threads/humans-in-the-loop/2024

## 2025

24 sessions.

- [Look At Your ****ing Data 👀](https://mlopstalks.com/talks/look-at-your-ing-data) (Kenny Daniel, Hyperparam). Pushed back: Kenny argued that specialized expert data work should not generally be outsourced to generic labeling companies when domain expertise is required. [19:20](https://www.youtube.com/watch?v=6EMnkAHmoag&t=1160s)
- [Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations](https://mlopstalks.com/talks/which-economic-tasks-are-performed-with-ai-evidence-from-millions-of-claude) (Valdimar Eggertsson, Snjallgögn (Smart Data inc.) & Sophia Skowronski, Breckinridge Capital Advisors). Pushed back: Adam Becker disputes treating debugging an error as fully automating a person's job, because the task can still require human context and judgment. [26:46](https://www.youtube.com/watch?v=DKqocE5JHfU&t=1606s)
- [Making AI Reliable is the Greatest Challenge of the 2020s](https://mlopstalks.com/talks/making-ai-reliable-is-the-greatest-challenge-of-the-2020s) (Alon Bochman, RagMetrics). Pushed back: Alon Bochman argues that a jury of multiple LLM judges should be adopted only if it produces a higher human agreement rate on the team's task. [48:55](https://www.youtube.com/watch?v=d4PGxNM3Iis&t=2935s)
- [Knowledge is Eventually Consistent](https://mlopstalks.com/talks/knowledge-is-eventually-consistent) (Devin Stein, Dosu). Pushed back: Demetrios Brinkmann challenges whether asking experts to explicitly approve and save facts is too much work. [18:04](https://www.youtube.com/watch?v=HvtzIx1vgmc&t=1084s)

20 more from 2025 on this thread: https://mlopstalks.com/threads/humans-in-the-loop/2025

## 2026

7 sessions.

- [Enterprise AI Operations: The Missing Piece](https://mlopstalks.com/talks/enterprise-ai-operations-the-missing-piece) (Rani Radhakrishnan, PwC US). Pushed back: The cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included. [28:48](https://www.youtube.com/watch?v=jTmV_jlob5I&t=1728s)
- [Structured Dissent Patterns for Agentic Production Reliability](https://mlopstalks.com/talks/structured-dissent-patterns-for-agentic-production-reliability) (Phil Stafford, MLOps Community). Claim: Phil Stafford says unresolved conflicts are reported as matters requiring more information, more time, or human decision-making. [8:11](https://www.youtube.com/watch?v=blOifXIJLe4&t=491s)
- [Stop Building AI Like Traditional Software](https://mlopstalks.com/talks/stop-building-ai-like-traditional-software) (Aishwarya Naresh Reganti, LevelUp Labs). Claim: A customer-support system can progress from routing tickets, to suggesting resolutions for human agents, to autonomously drafting and resolving common tickets. [6:46](https://www.youtube.com/watch?v=_CToYjO18J4&t=406s)
- [Everything We Got Wrong About Research-Plan-Implement](https://mlopstalks.com/talks/everything-we-got-wrong-about-research-plan-implement) (Dexter Horthy, HumanLayer). Pushed back: Dexter Horthy rejected the claim that software factories should never have humans read either the generated plan or code. [24:18](https://www.youtube.com/watch?v=YwZR6tc7qYg&t=1458s)

3 more from 2026 on this thread: https://mlopstalks.com/threads/humans-in-the-loop/2026
