PodcastEvolving Workflow OrchestrationClaimThe current trend in data science and machine learning is moving away from separate workflow-definition languages toward Python code and annotations.31:41
24 sessions
PodcastEvolving Workflow OrchestrationClaimThe current trend in data science and machine learning is moving away from separate workflow-definition languages toward Python code and annotations.31:41
PodcastLook At Your ****ing Data 馃憖Pushed backKenny argued that specialized expert data work should not generally be outsourced to generic labeling companies when domain expertise is required.19:20
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude ConversationsPushed backAdam Becker disputes treating debugging an error as fully automating a person's job, because the task can still require human context and judgment.26:46
PodcastAI-Powered Product Ideation with Synthetic Consumer TestingClaimThe workflow still requires a human with sufficient business context to determine whether the data and outputs are sensible.17:09
PodcastHow Sama is Improving ML Models to Make AVs SaferClaimSama combines its technology platform with human intelligence to capture the information that AI models need while reducing unnecessary annotation work.0:50
PodcastAI Data Engineers: Data Engineering After AIClaimThe agent is not currently fully autonomous because human approval remains important when changes could affect production systems.25:29
PodcastMaking AI Reliable is the Greatest Challenge of the 2020sPushed backAlon Bochman argues that a jury of multiple LLM judges should be adopted only if it produces a higher human agreement rate on the team's task.48:55
Building an AI agent with LangGraph, step by step tutorialClaimThe free Intro to LangGraph course from LangGraph Academy covers agent state, memory, and human-in-the-loop integration.1:18
How Product Metrics Become LLM EvaluationsClaimHumanloop has three core pillars: production tracing and observability, development datasets and evaluators, and tools for managing prompts and collaboration.16:48
PodcastKnowledge is Eventually ConsistentPushed backDemetrios Brinkmann challenges whether asking experts to explicitly approve and save facts is too much work.18:04
How AI Will Transform The Energy SectorClaimAdam Sroka says Jellyfish keeps human approval gates in place until its change-classification systems reach a high level of confidence.11:28
I Built A Trustworthy Voice AssistantClaimA voice agent listens, understands, decides, and responds in natural language, turning human speech into structured, actionable interactions.0:48
Iterating on Your AI EvalsPushed backHuman-in-the-loop evaluation may be introduced too late in the advanced-system example.12:22
Fast, Trustworthy, Reliable Voice Agents: MLOps That Blend LLM Annotation with Human QAClaimHappyRobot combines human annotations with LLM-generated annotations and aligns LLM annotator prompts with human signals to create accurate datasets across use cases.6:46
Too much lock-in for too little gain: agent frameworks are a dead-endClaimAgentic systems should use human input as a learning path toward greater autonomy rather than beginning as fully autonomous systems.2:29
If There's Free Compute, There's Abuse: Fighting Fraud with Lightweight LLM AgentsClaimA human remains responsible for the final decision because the agents still produce too many false positives to ban users automatically.6:32
PodcastThe Era of AI Agents in MarketingPushed backAI-generated thought leadership can replace authentic human thought leadership and relationship building.40:05
Catastrophic agent failure and how to avoid itPushed backHumans cannot be assumed to provide a perfect solution for avoiding false positives, because human labeling itself can contain mistakes.19:44
Designing AI Agents for the Complex Realities of HealthcareClaimAI agents in healthcare need human oversight, escalation at critical junctures, and appropriate communication with team members.2:55
Underwriting Assist: A Multi-Agent SystemClaimHuman review is included throughout the underwriting workflow because quoting and decisioning require accurate information, compliance, and auditability.5:28
Why You Should Care About Observability in LLM WorkflowsClaimThe team used custom GPTs, a Python adapter, and FDA and USDA interfaces to generate nutrition facts labels more accurately.2:52
Evals Aren't Useful? Really?ClaimWhen an evaluation exposes a problem, the fix may involve prompting, code changes for security issues, or a reviewer that checks the output before returning it.5:23
Fine-Tuned Models Are Getting Out of HandClaimAI systems are low-trust systems operating in high-trust environments, so interaction design and human-like behavior matter alongside model quality.4:46