Large Language Models in Production Round-table ConversationClaimHannes Hapke says that generated text at Digits is filtered for unsafe content and hallucination patterns and is reviewed by an accountant before being sent to a client.13:20
21 sessions
Large Language Models in Production Round-table ConversationClaimHannes Hapke says that generated text at Digits is filtered for unsafe content and hallucination patterns and is reviewed by an accountant before being sent to a client.13:20
LLMs as Intelligent AssistantsPushed backSarah Aerni rejects replacing human review with fully autonomous generated output and argues that human review remains critical.23:40
Building and Curating Datasets for RLHF and LLM Fine-tuningClaimHuman feedback can concern outputs, inputs, prompts, and other parts of the data pipeline.4:21
EvaluationClaimJosh Tobin says LLM application evaluation should combine automated evaluation with human evaluation, with nontechnical stakeholders helping produce evaluations and technical teams consuming them.13:16
Building Reliable AI AgentsClaimTravis Fischer views agents as a spectrum ranging from deterministic, human-directed programs to fully self-driving programs such as AutoGPT and BabyAGI.7:17
Designing Human in the Loop Experiences for LLMsClaimHuman-in-the-loop interaction for LLMs includes labeling, teaching models, and getting information into a model’s knowledge.1:48
PodcastMLOps at the Age of Generative AIPushed backBarak Turovsky disputes the idea that companies can simply add a large language model on top of an existing tool or replace staff without redesigning processes, handling hallucinations, and adding human exception handling.46:58
Incorporating LLMs in High-stake Use CasesClaimHigh-stakes applications should include humans in the loop, either through domain experts using the product or through background experts checking alerts and potentially fatal cases.3:57
RLHF Data Collection in PracticeClaimSurge AI is a full-stack human feedback company that hires contractors, builds its own labeling platform, and delivers data to clients.1:34
UX of an LLM UserClaimBardeen AI shows users a generated automation in preview mode while keeping the description box available for refinement.9:08
Evolving AI Governance for an LLM WorldClaimLLM-powered workflows should have documented fallback paths to human experts when they cannot provide required quality guarantees.12:36
Automating Data Annotation with LLMsPushed backHuman annotation is described as the gold standard, while large foundation models are largely created from unlabeled data, creating a tension between the two approaches.7:12
Amplifying Impact with Generative AI: Insights from 10,000 ColleaguesClaimFor user-facing applications at scale, the team generally expects a human to check quality continuously or to provide social quality control.31:05
LLMs in Production at GetYourGuidePushed backThe GPT-4 evaluator was not consistently better than human evaluators, although it sometimes detected missed information and false positives.20:58
From Building Self-driving Cars to Building LLM ApplicationsClaimLLM development should be collaborative by enabling nontechnical users to annotate monitoring and test results, review tests, and influence prompts or model experiments.10:02
Assess the Value and Feasibility of LLM Use Cases with a ChecklistClaimThe required level of human involvement depends on reliability needs and risk tolerance.6:47
AI Squared: Breaking LLMs out of the Chat ApplicationClaimBenjamin Harvey says AI Squared uses human feedback and analytics to improve model accuracy and performance over time.5:19
PodcastModel Management in a Regulated EnvironmentClaimOliver Chipperfield's team uses automated retraining and validation processes while retaining a human approval step before models are released to stakeholders.18:02