Agents in production
What can an agent do on its own, and which tools should it get? Once it starts acting, someone needs a way to find out when it goes wrong.
Machine Learning in Cybersecurity
Pushed backMonika presented government use of open source intelligence as requiring special permissions and serving security purposes, while Demetrios Brinkmann raised concerns that governments could abuse the tools for surveillance.41:28
LLMs as Intelligent Assistants
Pushed backSarah Aerni rejects replacing human review with fully autonomous generated output and argues that human review remains critical.23:40
Building Production Copilots
Pushed backTristan Zajonc says autonomous agents do not work reliably yet, while the host expresses doubt based on what has been implemented so far.17:18
Building Reliable AI Agents
Pushed backFully autonomous agents are not yet the right default for production; more deterministic, code-driven agents should be used instead.9:12
Using LLMs to Power Consumer Search at Scale
Pushed backAravind Srinivas rejects describing Perplexity as merely an LLM wrapper, arguing that the product requires substantial orchestration and software engineering around the models.18:18
MLOps at the Crossroads
Pushed backLLMOps should be treated as a distinct specialization with specialized tools rather than only as an extension of existing MLOps.19:55
Becoming an AI Evangelist
Pushed backDemetrios Brinkmann says agents are not yet ready, while Alex Volkov argues that they are improving and should not be expected to remain at their current level forever.1:21
Managing Small Knowledge Graphs for Multi-agent Systems
Pushed backTom Smoker disputes the idea that one natural-language instruction should be trusted to make a multi-agent system complete a task correctly every time.36:41
Navigating the AI Frontier: The Power of Synthetic Data and Agent Evaluations in LLM Development
Pushed backBoris Selitser argued that online intervention is often undesirable for complex agents because they may recover or find useful workarounds that developers did not anticipate.13:42
Machine Learning, AI Agents, and Autonomy
Pushed backEgor Kraev disputes the idea that LLMs should be judged as complete solutions or agents, arguing that they are additional components in a larger system.9:33
AI Agents: The Future of ML Engineering?
Pushed backSuccess on Kaggle competitions was disputed as evidence that an agent can perform general machine learning engineering or automate scientific discovery.7:38
The Agent Landscape - Lessons Learned Putting Agents Into Production
Pushed backThe speakers reject the idea that declining cost per token automatically means that agent systems are becoming cheaper overall, because cost per answer can rise as agents use more tokens and calls.21:33
Web Agents: The Cutting Edge of AI is Here?
Pushed backThe team found that performance on the WebArena benchmark did not translate reliably to the web-agent tasks they cared about.17:49
AI REWIND 2025 - MLOps Reading Group Year-end Special
Pushed backRohan Prasad challenged the assumption that adding all available information to ever-larger context windows reliably improves LLM results.15:10
Building Agentic Tools for Production
Pushed backSam Partee rejected the idea that all tools should be put in one MCP gateway and said agents should not have too many tools.19:26
Enterprise AI Operations: The Missing Piece
Pushed backStarting with one agent and scaling sequentially is not the only valid approach; parallel development can work when processes are standardized and repeated across departments.9:42
Tool definitions are the new Prompt Engineering
Pushed backAlex Salazar argued that most people should currently use carefully selected, specific tools, while the long-term goal is to let agents choose from many tools.11:49