Humans in the loop
People label the training data and review LLM output. As agents take on more work, teams have to decide which actions still need a person's approval.
Scaling Human-in-the-Loop Machine Learning
Pushed backRobert Munro rejected the idea that data annotation is merely the uninteresting part of machine learning work.10:56
Creating Beautiful Ambient Music with Google Brain's Music Transformer
Pushed backDaniel Jeffries disputes the idea that artificial intelligence must either replace humans or be replaced by humans, arguing instead for human-machine collaboration.7:17
Product Management in Machine Learning
ClaimLaszlo Sragner says monitoring, evaluation, and labeling are tightly coupled and should happen as one continuous process for machine learning models.9:51
Data Selection for Data-Centric AI: Data Quality Over Quantity
Pushed backCody Coleman argues against using all available data by default, because noisy data and labels increase cost and difficulty.41:11
Data-Centric AI Means Centralizing Training Data
Pushed backAlberto Rizzoli disputes the usefulness of treating ImageNet as a fully reliable benchmark because some classes contain substantial labeling errors.13:47
The Future of AI and ML in Process Automation
ClaimSlater Victoroff says changing an OCR engine can invalidate labels when labels are stored only as positions in extracted text.23:25
Applications of Data Science
ClaimPallet built a multi-label, multi-class machine-learning classifier to automate job-data classification and labeling in its backend.13:58
MLOps Critiques
ClaimMatthijs Brouns says teams can automate model retraining and produce a report for human approval without automating every model-release decision.40:56
Labeled Datasets that Correct Themselves Automatically
Pushed backAutomated data cleaning should not be treated as a completely human-free process, because high-quality corrections still require human involvement.37:46
Creative AI: Using ML to Create Art, Music, and Jokes
Pushed backSuyash Joshi says the copyright claim for an autonomously AI-generated image was rejected because human authorship was required.17:32
LLMs as Intelligent Assistants
Pushed backSarah Aerni rejects replacing human review with fully autonomous generated output and argues that human review remains critical.23:40
MLOps at the Age of Generative AI
Pushed backBarak Turovsky disputes the idea that companies can simply add a large language model on top of an existing tool or replace staff without redesigning processes, handling hallucinations, and adding human exception handling.46:58
Automating Data Annotation with LLMs
Pushed backHuman annotation is described as the gold standard, while large foundation models are largely created from unlabeled data, creating a tension between the two approaches.7:12
LLMs in Production at GetYourGuide
Pushed backThe GPT-4 evaluator was not consistently better than human evaluators, although it sometimes detected missed information and false positives.20:58
Language, Graphs, and AI in Industry
ClaimPaco Nathan says AI applications in regulated or industrial settings need software engineering, operations, security, legal review and domain expertise rather than only an API call.1:00:42
Turn Data Chaos into AI Strategy with Programmatic AI Data Development
Pushed backProgrammatic labeling functions should not be used directly as the final inference rules because a rule-based system is not robust or scalable enough for new data.22:35
The Future of Healthcare: AI is Here
Pushed backShaun disputes the idea that AI should merely match human performance, saying HeyRevia's results show its agents can outperform humans in comparable healthcare phone-call scenarios.23:14
Why Planning is the New Search
Pushed backFabian disputes the assumption that agentic workflows should immediately be fully self-driving, arguing that customers generally want human oversight and gradual automation first.26:11
Look At Your ****ing Data 👀
Pushed backKenny argued that specialized expert data work should not generally be outsourced to generic labeling companies when domain expertise is required.19:20
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
Pushed backAdam Becker disputes treating debugging an error as fully automating a person's job, because the task can still require human context and judgment.26:46
Making AI Reliable is the Greatest Challenge of the 2020s
Pushed backAlon Bochman argues that a jury of multiple LLM judges should be adopted only if it produces a higher human agreement rate on the team's task.48:55
Knowledge is Eventually Consistent
Pushed backDemetrios Brinkmann challenges whether asking experts to explicitly approve and save facts is too much work.18:04
Enterprise AI Operations: The Missing Piece
Pushed backThe cost of AI should not be calculated simply as replacing human workers, because review, storage, retrieval, and other operating costs must also be included.28:48
Structured Dissent Patterns for Agentic Production Reliability
ClaimPhil Stafford says unresolved conflicts are reported as matters requiring more information, more time, or human decision-making.8:11
Stop Building AI Like Traditional Software
ClaimA customer-support system can progress from routing tickets, to suggesting resolutions for human agents, to autonomously drafting and resolving common tickets.6:46
Everything We Got Wrong About Research-Plan-Implement
Pushed backDexter Horthy rejected the claim that software factories should never have humans read either the generated plan or code.24:18