PodcastRe-Platforming Your Tech StackClaimProduction teams need to consider monitoring, risk, longevity and ongoing maintenance rather than treating a model as finished when it is deployed.7:44
19 sessions
PodcastRe-Platforming Your Tech StackClaimProduction teams need to consider monitoring, risk, longevity and ongoing maintenance rather than treating a model as finished when it is deployed.7:44
PodcastEfficient Deployment of Models at the EdgeClaimKrishna Sridhar said that Qualcomm AI Hub can deploy one model across multiple generations of devices and adapt execution to the available hardware.33:32
Building AI That Remembers YouClaimLetta provides an agent service that manages memory, agent state, tool execution, and additional data sources for client applications.2:50
PodcastKubernetes, AI Gateways, and the Future of MLOpsPushed backDemetrios Brinkmann suggested that KServe is built on KNative and helps provision model resources, while Alexa Griffith clarified that KNative can run serverless services generally and KServe is specifically focused on AI and ML models.19:39
PodcastGenAI Traffic: Why API Infrastructure Must Evolve... AgainClaimThe move from monoliths to microservices made scaling more resource-efficient but created a new networking problem because services moved between addresses.9:02
PodcastStreaming Ecosystem Complexities and Cost ManagementClaimRohit Agrawal joined Tecton as an individual contributor and now manages teams working on streaming data, batch data, online inference, and offline inference.2:10
Building Robust AI Systems with Battle-tested FrameworksClaimCharles Frye describes Modal as serverless Python infrastructure for running code on remote CPUs and GPUs with minimal configuration.30:01
PodcastMLOps with DatabricksPushed backMaria Vechtomova disagrees with the idea that Databricks should be used for every model-serving situation.26:50
Building Out GPU CloudsClaimCustomers often cannot experiment easily because cloud providers may require long commitments or charge very high reserved pricing.1:53
PodcastThe Creator of FastAPI's Next ChapterClaimSebastián Ramírez says FastAPI will remain fully open source with all features available regardless of where users deploy it.24:36
PodcastInside Uber's AI Revolution: Everything About How They Use AI/MLClaimUber measures developer velocity with anecdotal team feedback and proxy metrics such as training runs, model deployments, evaluation pipelines and generated reports.8:03
PodcastThe Rise of Sovereign AI and Global AI Innovation in a World of US ProtectionismPushed backFrank Meehan responds that anonymized research data can be shared, but identifiable medical data used for patient care and internal health-system inference requires local protection.19:02
PodcastThe Truth About LLM TrainingClaimProsus evaluates models for domain understanding, language capabilities, cost, fine-tuning and distillation potential, and inference performance.1:50
Integration of AI into Traditional SystemsPushed backHakan Tek rejects the idea that using a third-party service can provide complete certainty about data security.14:30
The Hidden Infrastructure Behind Every AI AgentClaimAgent activity can generate calls to OpenAI, Bedrock, Anthropic, Gemini, internal tools, databases, and embedding services.3:10
How to Self-Host an AI AgentPushed backInference operations should not remain isolated in separate teams because centralized infrastructure and knowledge can coexist with decentralized business-specific knowledge.21:52
PodcastIs Open Source Software Actually Secure?Pushed backDemetrios Brinkmann frames the path to production as having been shortened, while Hudson Buzby says it is mainly easy to deploy because standards are still immature and teams have not yet developed consistent practices.8:23
What It Takes to Run Multi-Agent SystemsClaimReducing latency requires bringing AI closer to where the data resides instead of constantly moving data between locations.7:55
Building Data Centers for GPU CloudsClaimBuzz High Performance Compute is a Canadian cloud service provider with data-center facilities in Sweden and Canada that operates GPU clouds at scale.0:27