PodcastMLOps at the CrossroadsClaimPatrick Barker says his MLOps experience with traditional NLP, anomaly detection, tabular models, Kubernetes, MLflow, and model serving did not transfer substantially to building applications with LLMs.13:10
32 sessions
PodcastMLOps at the CrossroadsClaimPatrick Barker says his MLOps experience with traditional NLP, anomaly detection, tabular models, Kubernetes, MLflow, and model serving did not transfer substantially to building applications with LLMs.13:10
PodcastAds Ranking Evolution at PinterestClaimPinterest's early ads-ranking stack combined XGBoost-based GBDTs, TensorFlow-based models, and a C++ serving library.14:32
PodcastDesigning ML Infra for ML & LLM Use CasesClaimEngineers scaling AI systems need to consider data scalability, model bias and outliers, model serving performance, and compliance and privacy.49:07
Charting LLMOps OdysseyClaimYinxi Zhang says CI/CD pipelines should test, stage, monitor, and deploy LLM applications, including unit tests, integration tests, and end-to-end tests.26:27
Vision Pipelines in Production: Serving & OptimisationsClaimBiswaroop Bhattacharjee is discussing how to serve and optimize vision pipelines in production.1:01
Building a Python-Centric Feature Platform to Power Production AI ApplicationsClaimA feature platform connects raw data to transformed features for model training and online inference while bridging offline and online environments.5:19
Graduating from Proprietary to Open Source Models in ProductionPushed backPhilip Kiely does not give a definite position on whether companies will serve models with alternative hardware such as Groq.22:02
Productionizing Health Insurance Appeal GenerationPushed backHolden Karau rejects using the latest container version as a good deployment practice and says the version should be pinned.11:50
Introducing DBRX: The Future of Language ModelsClaimDBRX inference was tested by deploying an optimized inference web server, sending it large numbers of requests, measuring throughput, and optimizing the system based on those results.27:18
PodcastGenAI in Production - Challenges and TrendsClaimAt Amazon Alexa, Verena Weber worked on natural-language-understanding models and handled model retraining, deployment, release updates, and research projects.18:18
LLMOps and GenAI at Enterprise Scale - Challenges and OpportunitiesClaimDeployment of LLMs and generative AI is difficult, and relatively few organizations are successfully deploying generative AI at scale.0:44
Streamlining Model DeploymentPushed backThe discussion disputes the assumption that selecting the same model across endpoint providers necessarily produces the same output quality.17:35
PodcastFedML Nexus AI: Your Generative AI Platform at ScaleClaimSalman Avestimehr says ownership, scalability, observability, privacy, and safety are major challenges when building generative AI applications.5:25
Productionizing AI: How to Think From the EndPushed backThe idea that deploying LLMs has a quick and dirty approach was challenged because production deployment is not actually simple.10:17
PodcastRetrieval Augmented GenerationClaimSyed Asad says OpenAI raises concerns about cost and data handling, while local tools such as Ollama can be difficult to deploy in production because of remote-connection errors.23:18
Navigating the Emerging LLMOps StackClaimHien Luu identifies inference and serving as especially important challenges in building and deploying LLM applications.2:05
PodcastAWS Trainium and InferentiaClaimInferentia launched in 2019 for inference acceleration, while Trainium launched in 2022 with a focus on training and distributing large language models and generative AI workloads.4:19
PodcastUber's Michelangelo: Strategic AI Overhaul and ImpactClaimMichelangelo 1.0 treated every project the same, regardless of its business impact, so a high-revenue model could receive the same support and service levels as a model without clear return on investment.10:27
How to Build Production-Ready AI Models for ManufacturingClaimThe latency, accuracy and safety requirements for an AI system depend heavily on its application.12:44
PodcastAll Data Scientists Should Learn Software Engineering PrinciplesClaimCatherine Nelson says data scientists should retain enough ownership of their models to work with software engineers on production code and preserve the reasoning behind model choices.25:01
PodcastExtending AI: From Industry to InnovationClaimMoving machine-learning systems from batch processing into application-focused, real-time environments requires data scientists to work with application developers and software interfaces.6:30
PodcastMLOps for GenAI ApplicationsPushed backHarcharan Kabbay says Kubernetes resource requests should represent a minimum starting allocation rather than reserving the full possible limit.48:10
Boosting LLMs: Performance, Scaling, and Structured OutputsClaimMatt Squire says the team measured latency, output throughput in tokens per second, and the rate of successfully served requests.12:06
11 lessons learned from doing deploymentsPushed backSol Rashidi disputes the idea that AI deployment problems are mainly technical, arguing that 70% of the hurdles are non-technical.10:40
Common ML Serving Architectures ExplainedClaimRebecca Taylor says that data scientists may hand over notebooks or model artifacts to other teams that productionize and deploy them.2:38
From Notebook to Kubernetes: Scaling GenAI Pipelines with ZenMLClaimFlux produced better cat images than the Stable Diffusion model in the demonstrated inference comparison.9:07
Real-Time Event Processing for AI/ML with NumaflowClaimNumaflow provides out-of-the-box sources and sinks so teams can focus on processing or inference logic rather than connecting to messaging systems.9:13
The Next Revolution in AI: LLMs and BeyondClaimApplications using LLMs need to reduce hallucinations, fit prompts and runtime within the required latency, and incorporate fresh information such as search results and enterprise data.4:05
Why DuckDB is the Future of DataClaimDuckDB runs in process instead of requiring a separate database server.5:54
PodcastComposable Memory for GPU OptimizationClaimBernie Wu says memory-level checkpointing combined with framework-level checkpointing can improve resilience for long-running training and inference workloads.8:09
How to Actually Use Cost Effective AI in Your BusinessPushed backThe speakers argue that large machine-learning workloads do not necessarily require Kubernetes because AWS Batch can run them without users managing Kubernetes infrastructure.44:25
AI Agents: The Future of Productivity, or Just a Fad?ClaimSam Partee says Arcade's tool SDK creates a Python package for writing and serving a tool function to a large language model.20:30