Humans in the loop in 2023

21 sessions

Large Language Models in Production Round-table ConversationDiego Oppenheimer, Factory HQ & David Hershey, Unusual Ventures & Hannes Hapke, Digits & James Richards, Bountiful & Rebecca Qian, Facebook AI Research · 57:21 · Mar 2023 · 14K views · LLMs in Production 2023

ClaimHannes Hapke says that generated text at Digits is filtered for unsafe content and hallucination patterns and is reviewed by an accountant before being sent to a client.13:20

Guiding LLMs While Staying in the Driver's SeatJacob van Gogh, Adept AI · 10:02 · May 2023 · 350 views · LLMs in Production 2023
LLMs as Intelligent AssistantsSarah Aerni, Salesforce · 28:45 · Jun 2023 · 967 views · LLMs in Production 2023

Pushed backSarah Aerni rejects replacing human review with fully autonomous generated output and argues that human review remains critical.23:40

Building and Curating Datasets for RLHF and LLM Fine-tuningDaniel Vila Suero, Argilla · 58:51 · Jul 2023 · 3,363 views · LLMs in Production 2023

ClaimHuman feedback can concern outputs, inputs, prompts, and other parts of the data pipeline.4:21

EvaluationAbi Aryan, Independent Consultant & Amrutha Gujjar, Structured & Josh Tobin, Gantry & Sohini Roy, NVIDIA · 38:19 · Jul 2023 · 732 views · LLMs in Production 2023

ClaimJosh Tobin says LLM application evaluation should combine automated evaluation with human evaluation, with nontechnical stakeholders helping produce evaluations and technical teams consuming them.13:16

Building Reliable AI AgentsTravis Fischer · 17:43 · Jul 2023 · 541 views · LLMs in Production 2023

ClaimTravis Fischer views agents as a spectrum ranging from deterministic, human-directed programs to fully self-driving programs such as AutoGPT and BabyAGI.7:17

Designing Human in the Loop Experiences for LLMsAlberto Rizzoli, V7 · 11:40 · Aug 2023 · 1,462 views · LLMs in Production 2023

ClaimHuman-in-the-loop interaction for LLMs includes labeling, teaching models, and getting information into a model’s knowledge.1:48

PodcastMLOps at the Age of Generative AIBarak Turovsky, Scale Venture Partners · 56:56 · Aug 2023 · 743 views · MLOps Podcast

Pushed backBarak Turovsky disputes the idea that companies can simply add a large language model on top of an existing tool or replace staff without redesigning processes, handling hallucinations, and adding human exception handling.46:58

Incorporating LLMs in High-stake Use CasesYada Pruksachatkun, Moonhub · 11:01 · Aug 2023 · 870 views · LLMs in Production 2023

ClaimHigh-stakes applications should include humans in the loop, either through domain experts using the product or through background experts checking alerts and potentially fatal cases.3:57

RLHF Data Collection in PracticeAndrew Mauboussin, Surge AI · 12:10 · Aug 2023 · 684 views · LLMs in Production 2023

ClaimSurge AI is a full-stack human feedback company that hires contractors, builds its own labeling platform, and delivers data to clients.1:34

UX of an LLM UserMisty Free, Jasper & Davis Treybig, Innovation Endeavors & Dina Yerlan, Adobe Firefly & Artem Harutyunyan, Bardeen AI · 31:48 · Aug 2023 · 742 views · LLMs in Production 2023

ClaimBardeen AI shows users a generated automation in preview mode while keeping the description box available for refinement.9:08

Evolving AI Governance for an LLM WorldDiego Oppenheimer, Factory · 14:47 · Aug 2023 · 517 views · LLMs in Production 2023

ClaimLLM-powered workflows should have documented fallback paths to human experts when they cannot provide required quality guarantees.12:36

PodcastUsing Large Language Models at AngelListThibaut Labarre, AngelList · 51:42 · Aug 2023 · 931 views · MLOps Podcast
Fireside Chat - The Future of LLMsDavid Hershey, Unusual Ventures & Daniel Jeffries, AI Infrastructure Alliance · 36:07 · Aug 2023 · 324 views · LLMs in Production 2023
Automating Data Annotation with LLMsNikolai Liubimov, Michael Malyuk & Chris Hoge, HumanSignal · 1:03:04 · Oct 2023 · 4,358 views · LLMs in Production 2023

Pushed backHuman annotation is described as the gold standard, while large foundation models are largely created from unlabeled data, creating a tension between the two approaches.7:12

Amplifying Impact with Generative AI: Insights from 10,000 ColleaguesPaul van der Boor, Prosus · 32:12 · Oct 2023 · 258 views

ClaimFor user-facing applications at scale, the team generally expects a human to check quality continuously or to provide social quality control.31:05

LLMs in Production at GetYourGuideMeghana Satish & Tina Treimane, GetYourGuide · 29:39 · Oct 2023 · 1,478 views · LLMs in Production 2023

Pushed backThe GPT-4 evaluator was not consistently better than human evaluators, although it sometimes detected missed information and false positives.20:58

From Building Self-driving Cars to Building LLM ApplicationsEffy Zhang, Baserun · 10:45 · Nov 2023 · 578 views · LLMs in Production 2023

ClaimLLM development should be collaborative by enabling nontechnical users to annotate monitoring and test results, review tests, and influence prompts or model experiments.10:02

Assess the Value and Feasibility of LLM Use Cases with a ChecklistRens Dimmendaal & Eva Bosma, Xebia Data · 10:39 · Nov 2023 · 415 views

ClaimThe required level of human involvement depends on reliability needs and risk tolerance.6:47

AI Squared: Breaking LLMs out of the Chat ApplicationBenjamin Harvey, AI Squared · 52:20 · Nov 2023 · 566 views · LLMs in Production 2023

ClaimBenjamin Harvey says AI Squared uses human feedback and analytics to improve model accuracy and performance over time.5:19

PodcastModel Management in a Regulated EnvironmentDarek Kłeczek, Weights & Biases & Mark Huang, Gradient & Oliver Chipperfield, M-KOPA & Michelle Marie Conway, Lloyds Banking Group · 58:30 · Dec 2023 · 198 views · MLOps Coffee Sessions

ClaimOliver Chipperfield's team uses automated retraining and validation processes while retaining a human approval step before models are released to stakeholders.18:02