Humans in the loop in 2024

16 sessions

PodcastLanguage, Graphs, and AI in IndustryPaco Nathan, Derwen, Inc. · 1:18:29 · Jan 2024 · 534 views · MLOps Podcast

ClaimPaco Nathan says AI applications in regulated or industrial settings need software engineering, operations, security, legal review and domain expertise rather than only an API call.1:00:42

From Research to Production: Fine-Tuning & Aligning LLMsPhilipp Schmid, Hugging Face · 38:03 · Apr 2024 · 1,283 views · AI in Production 2024

ClaimA good instruction dataset needs clear instructions, diverse tasks and topics, consistent formatting, and high-quality human feedback.12:57

Graphs and LanguageLouis Guitton · 11:28 · Apr 2024 · 510 views · AI in Production 2024

ClaimGraph visualization can help humans understand black-box models and support expert semi-supervision.1:54

Data Labeling Best PracticesCharles Brecque, TextMine · 12:59 · May 2024 · 682 views · AI in Production 2024

ClaimTextMine built its own data labeling team and fine-tuned its own models, and this talk presents its lessons rather than prescribing one correct labeling method.1:07

Ghostwriter - AI Writing That Learns From YouJonny Dimond, Shortwave · 29:53 · May 2024 · 654 views · AI in Production 2024
PodcastReliable LLM Products, Fueled by FeedbackChinar Movsisyan, Feedback Intelligence · 49:17 · Jul 2024 · 368 views · MLOps Podcast

ClaimThe drone project used more than 10K high-resolution images that had to be manually annotated before training a custom YOLO 3 model.4:25

Balancing Speed and SafetyRemy Thellier, Vectice & Erica Greene, Yahoo & Shreya Rajpal, Guardrails AI · 35:40 · Aug 2024 · 159 views · AIQCON 2024

ClaimSuccessfully deployed large language model applications are often constrained by human review or limited to internal question answering.20:28

Vision and Strategies for Attracting & Driving AI Talents in High GrowthAshley Antonides, Two Six Technologies & Olga Beregovaya, Smartling & Shailvi Wakhlu, Shailvi Ventures LLC · 30:34 · Aug 2024 · 151 views · AIQCON 2024

ClaimOlga Beregovaya says human reviewers remain necessary for tasks such as assessment, ranking, validation, post-editing, and fact checking because AI models can hallucinate and data can be biased.15:27

The Next Revolution in AI: LLMs and Beyond · 13:47 · Oct 2024 · 187 views

ClaimFor qualitative applications, human review and quick sanity checks are useful for improving prompts and content.8:46

Turn Data Chaos into AI Strategy with Programmatic AI Data DevelopmentElena Boiarskaia, Snorkel AI · 27:14 · Oct 2024 · 67 views · DE4AI 2024

Pushed backProgrammatic labeling functions should not be used directly as the final inference rules because a rule-based system is not robust or scalable enough for new data.22:35

PodcastHow Agentic Workflows Will Change EverythingRaj Rikhy, Microsoft · 49:13 · Oct 2024 · 705 views · MLOps Podcast

ClaimFor an initial MVP, developers should keep a human in the loop, inspect the agent's planned actions, and constrain its environment.17:48

Cleric AI SRE: Towards Self-healing Autonomous SoftwareWillem Pienaar, Cleric · 29:01 · Nov 2024 · 2,379 views · Agents in Production 2024

ClaimProduction infrastructure is complex, dynamic, and difficult for humans to keep in mind as systems and connections grow.2:57

The Future of Healthcare: AI is Here · 24:56 · Dec 2024 · 402 views

Pushed backShaun disputes the idea that AI should merely match human performance, saying HeyRevia's results show its agents can outperform humans in comparable healthcare phone-call scenarios.23:14

Hundreds of Users Love Our Data Analyst AI AgentIoannis Zempekakis & Donné Stevenson · 29:00 · Dec 2024 · 1,116 views
Why Planning is the New Search · 27:48 · Dec 2024 · 466 views

Pushed backFabian disputes the assumption that agentic workflows should immediately be fully self-driving, arguing that customers generally want human oversight and gradual automation first.26:11

Building Reliable AgentsEno Reyes, Factory · 24:45 · Dec 2024 · 553 views

ClaimHuman input should remain part of agentic-system design because these systems are unlikely to complete large tasks reliably every time.19:36