benchmarks
16 talks
MCP-Enabled Agents
Testing AI Intelligence: The Benchmarking Battle
Which Economic Tasks are Performed with AI? Evidence from Millions of Claude Conversations
Reading groupAI Agents: The Future of ML Engineering?
PodcastThe Battle for AI: Robustness vs Privacy
Few Shot Code Generation to Autonomous Software Engineering Agents
PodcastHow to Optimize Large AI Models with PyTorch
PodcastThe EU AI Act: Navigating New Legislation
Boosting LLMs: Performance, Scaling, and Structured Outputs
Reading groupExploring Long Context Language Models
Evaluating Language Models
Streamlining Model Deployment
PodcastEvaluating and Integrating ML Models
The Truth About AI Agents
PodcastAll About Evaluating LLM Applications
PodcastData Selection for Data-Centric AI: Data Quality Over Quantity