Threads / Humans in the loop

Humans in the loop

People label the training data and review LLM output. As agents take on more work, teams have to decide which actions still need a person's approval.

Follows the tags human-in-the-loop · 82 sessions · 2020 to 2026
20202 sessions
20216 sessions
Meetup · MLOps Meetup #54

Product Management in Machine Learning

Laszlo Sragner, Hypergolic

ClaimLaszlo Sragner says monitoring, evaluation, and labeling are tightly coupled and should happen as one continuous process for machine learning models.9:51

2 more from 2021 on this thread
20226 sessions
Meetup · MLOps Meetup #95

Applications of Data Science

Connie Yang, Pallet

ClaimPallet built a multi-label, multi-class machine-learning classifier to automate job-data classification and labeling in its backend.13:58

Podcast · MLOps Coffee Sessions #100

MLOps Critiques

Matthijs Brouns, Xccelerated.io

ClaimMatthijs Brouns says teams can automate model retraining and produce a report for human approval without automating every model-release decision.40:56

2 more from 2022 on this thread
202321 sessions
Talk · LLMs in Production 2023

LLMs as Intelligent Assistants

Sarah Aerni, Salesforce

Pushed backSarah Aerni rejects replacing human review with fully autonomous generated output and argues that human review remains critical.23:40

Podcast · MLOps Podcast #169

MLOps at the Age of Generative AI

Barak Turovsky, Scale Venture Partners

Pushed backBarak Turovsky disputes the idea that companies can simply add a large language model on top of an existing tool or replace staff without redesigning processes, handling hallucinations, and adding human exception handling.46:58

Talk · LLMs in Production 2023

Automating Data Annotation with LLMs

Nikolai Liubimov, Michael Malyuk & Chris Hoge, HumanSignal

Pushed backHuman annotation is described as the gold standard, while large foundation models are largely created from unlabeled data, creating a tension between the two approaches.7:12

Talk · LLMs in Production 2023

LLMs in Production at GetYourGuide

Meghana Satish & Tina Treimane, GetYourGuide

Pushed backThe GPT-4 evaluator was not consistently better than human evaluators, although it sometimes detected missed information and false positives.20:58

17 more from 2023 on this thread
202416 sessions
Podcast · MLOps Podcast #201

Language, Graphs, and AI in Industry

Paco Nathan, Derwen, Inc.

ClaimPaco Nathan says AI applications in regulated or industrial settings need software engineering, operations, security, legal review and domain expertise rather than only an API call.1:00:42

Talk

The Future of Healthcare: AI is Here

Pushed backShaun disputes the idea that AI should merely match human performance, saying HeyRevia's results show its agents can outperform humans in comparable healthcare phone-call scenarios.23:14

Talk

Why Planning is the New Search

Pushed backFabian disputes the assumption that agentic workflows should immediately be fully self-driving, arguing that customers generally want human oversight and gradual automation first.26:11

12 more from 2024 on this thread
202524 sessions
Podcast · MLOps Podcast #292

Look At Your ****ing Data 👀

Kenny Daniel, Hyperparam

Pushed backKenny argued that specialized expert data work should not generally be outsourced to generic labeling companies when domain expertise is required.19:20

Podcast · MLOps Podcast #335

Knowledge is Eventually Consistent

Devin Stein, Dosu

Pushed backDemetrios Brinkmann challenges whether asking experts to explicitly approve and save facts is too much work.18:04

20 more from 2025 on this thread
20267 sessions
Podcast · MLOps Podcast #345

Enterprise AI Operations: The Missing Piece

Rani Radhakrishnan, PwC US

Pushed backThe cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included.28:48

Talk

Stop Building AI Like Traditional Software

Aishwarya Naresh Reganti, LevelUp Labs

ClaimA customer-support system can progress from routing tickets, to suggesting resolutions for human agents, to autonomously drafting and resolving common tickets.6:46

3 more from 2026 on this thread