Serving and deployment in 2024

32 sessions

PodcastMLOps at the CrossroadsPatrick Barker, Kentauros AI & Farhood Etaati, AIMedic · 49:02 · Jan 2024 · 345 views · MLOps Podcast

ClaimPatrick Barker says his MLOps experience with traditional NLP, anomaly detection, tabular models, Kubernetes, MLflow, and model serving did not transfer substantially to building applications with LLMs.13:10

PodcastAds Ranking Evolution at PinterestAayush Mudgal, Pinterest · 52:38 · Feb 2024 · 601 views · MLOps Podcast

ClaimPinterest's early ads-ranking stack combined XGBoost-based GBDTs, TensorFlow-based models, and a C++ serving library.14:32

PodcastDesigning ML Infra for ML & LLM Use CasesAmritha Arun Babu, Klaviyo & Abhik Choudhury, IBM · 1:00:18 · Mar 2024 · 876 views · MLOps Podcast

ClaimEngineers scaling AI systems need to consider data scalability, model bias and outliers, model serving performance, and compliance and privacy.49:07

Charting LLMOps OdysseyYinxi Zhang, Databricks · 38:53 · Apr 2024 · 405 views · AI in Production 2024

ClaimYinxi Zhang says CI/CD pipelines should test, stage, monitor, and deploy LLM applications, including unit tests, integration tests, and end-to-end tests.26:27

Vision Pipelines in Production: Serving & OptimisationsBiswaroop Bhattacharjee, Prem AI · 13:24 · Apr 2024 · 281 views · AI in Production 2024

ClaimBiswaroop Bhattacharjee is discussing how to serve and optimize vision pipelines in production.1:01

Building a Python-Centric Feature Platform to Power Production AI ApplicationsMatt Bleifer, Tecton · 27:11 · Apr 2024 · 296 views

ClaimA feature platform connects raw data to transformed features for model training and online inference while bridging offline and online environments.5:19

Graduating from Proprietary to Open Source Models in ProductionPhilip Kiely, Baseten · 23:16 · Apr 2024 · 147 views · AI in Production 2024

Pushed backPhilip Kiely does not give a definite position on whether companies will serve models with alternative hardware such as Groq.22:02

Productionizing Health Insurance Appeal GenerationHolden Karau, Netflix · 27:11 · Apr 2024 · 594 views · AI in Production 2024

Pushed backHolden Karau rejects using the latest container version as a good deployment practice and says the version should be pinned.11:50

Introducing DBRX: The Future of Language ModelsDavis Blalock, Bandish Shah, Abhi Venigalla & Ajay Saini, Databricks · 48:36 · Apr 2024 · 496 views · MLOps Coffee Sessions

ClaimDBRX inference was tested by deploying an optimized inference web server, sending it large numbers of requests, measuring throughput, and optimizing the system based on those results.27:18

PodcastGenAI in Production - Challenges and TrendsVerena Weber, Verena Weber · 48:43 · Apr 2024 · 671 views · MLOps Podcast

ClaimAt Amazon Alexa, Verena Weber worked on natural-language-understanding models and handled model retraining, deployment, release updates, and research projects.18:18

LLMOps and GenAI at Enterprise Scale - Challenges and OpportunitiesAndy McMahon, NatWest Group · 13:04 · May 2024 · 564 views · AI in Production 2024

ClaimDeployment of LLMs and generative AI is difficult, and relatively few organizations are successfully deploying generative AI at scale.0:44

Streamlining Model DeploymentDaniel Lenton, Unify · 21:39 · May 2024 · 334 views · AI in Production 2024

Pushed backThe discussion disputes the assumption that selecting the same model across endpoint providers necessarily produces the same output quality.17:35

PodcastFedML Nexus AI: Your Generative AI Platform at ScaleSalman Avestimehr, FedML · 52:34 · May 2024 · 451 views · MLOps Podcast

ClaimSalman Avestimehr says ownership, scalability, observability, privacy, and safety are major challenges when building generative AI applications.5:25

Productionizing AI: How to Think From the EndAnnie Condon · 11:11 · May 2024 · 348 views · AI in Production 2024

Pushed backThe idea that deploying LLMs has a quick and dirty approach was challenged because production deployment is not actually simple.10:17

PodcastRetrieval Augmented GenerationSyed Asad, KiwiTech · 44:10 · May 2024 · 1,070 views · MLOps Podcast

ClaimSyed Asad says OpenAI raises concerns about cost and data handling, while local tools such as Ollama can be difficult to deploy in production because of remote-connection errors.23:18

Navigating the Emerging LLMOps StackHien Luu, DoorDash · 12:51 · May 2024 · 534 views · AI in Production 2024

ClaimHien Luu identifies inference and serving as especially important challenges in building and deploying LLM applications.2:05

PodcastAWS Trainium and InferentiaKamran Khan, Annapurna ML & Matthew McClean, AWS, Annapurna Labs · 45:23 · Jun 2024 · 826 views · MLOps Podcast

ClaimInferentia launched in 2019 for inference acceleration, while Trainium launched in 2022 with a focus on training and distributing large language models and generative AI workloads.4:19

PodcastUber's Michelangelo: Strategic AI Overhaul and Impact · 35:36 · Jun 2024 · 840 views · MLOps Podcast

ClaimMichelangelo 1.0 treated every project the same, regardless of its business impact, so a high-revenue model could receive the same support and service levels as a model without clear return on investment.10:27

How to Build Production-Ready AI Models for ManufacturingPavol Bielik, LatticeFlow AI & Aniket Singh & Mohan Mahadevan & Jürgen Weichenberger, Schneider Electric · 56:38 · Jun 2024 · 861 views

ClaimThe latency, accuracy and safety requirements for an AI system depend heavily on its application.12:44

PodcastAll Data Scientists Should Learn Software Engineering PrinciplesCatherine Nelson, Freelance Data Scientist · 52:55 · Jul 2024 · 1,059 views · MLOps Podcast

ClaimCatherine Nelson says data scientists should retain enough ownership of their models to work with software engineers on production code and preserve the reasoning behind model choices.25:01

PodcastExtending AI: From Industry to InnovationSophia Rowland & David Weik, SAS · 1:01:37 · Jul 2024 · 250 views · MLOps Podcast

ClaimMoving machine-learning systems from batch processing into application-focused, real-time environments requires data scientists to work with application developers and software interfaces.6:30

PodcastMLOps for GenAI ApplicationsHarcharan Kabbay, World Wide Technology · 1:05:02 · Aug 2024 · 743 views · MLOps Podcast

Pushed backHarcharan Kabbay says Kubernetes resource requests should represent a minimum starting allocation rather than reserving the full possible limit.48:10

Boosting LLMs: Performance, Scaling, and Structured OutputsTom Sabo, SAS & Matt Squire, Fuzzy Labs & Vaibhav Gupta, BoundaryML · 1:01:24 · Oct 2024 · 403 views · MLOps Mini Summit 2024

ClaimMatt Squire says the team measured latency, output throughput in tokens per second, and the rate of successfully served requests.12:06

11 lessons learned from doing deploymentsSol Rashidi, ExecutiveAI LLC · 35:59 · Oct 2024 · 275 views · DE4AI 2024

Pushed backSol Rashidi disputes the idea that AI deployment problems are mainly technical, arguing that 70% of the hurdles are non-technical.10:40

Common ML Serving Architectures ExplainedRebecca Taylor, Lidl e-commerce · 17:33 · Oct 2024 · 443 views

ClaimRebecca Taylor says that data scientists may hand over notebooks or model artifacts to other teams that productionize and deploy them.2:38

From Notebook to Kubernetes: Scaling GenAI Pipelines with ZenMLAlex Strick van Linschoten, ZenML · 12:42 · Oct 2024 · 256 views

ClaimFlux produced better cat images than the Stable Diffusion model in the demonstrated inference comparison.9:07

Real-Time Event Processing for AI/ML with NumaflowSri Harsha Yayi, Intuit · 22:38 · Oct 2024 · 601 views · DE4AI 2024

ClaimNumaflow provides out-of-the-box sources and sinks so teams can focus on processing or inference logic rather than connecting to messaging systems.9:13

The Next Revolution in AI: LLMs and Beyond · 13:47 · Oct 2024 · 187 views

ClaimApplications using LLMs need to reduce hallucinations, fit prompts and runtime within the required latency, and incorporate fresh information such as search results and enterprise data.4:05

Why DuckDB is the Future of DataProf. Dr. Hannes Mühleisen, DuckDB Labs · 31:56 · Oct 2024 · 2,698 views

ClaimDuckDB runs in process instead of requiring a separate database server.5:54

PodcastComposable Memory for GPU OptimizationBernie Wu, MemVerge · 55:19 · Oct 2024 · 381 views · MLOps Podcast

ClaimBernie Wu says memory-level checkpointing combined with framework-level checkpointing can improve resilience for long-running training and inference workloads.8:09

How to Actually Use Cost Effective AI in Your BusinessEddie Mattia, Outerbounds & Scott Perry, AWS · 49:05 · Nov 2024 · 262 views · MLOps Community Mini Summit #9

Pushed backThe speakers argue that large machine-learning workloads do not necessarily require Kubernetes because AWS Batch can run them without users managing Kubernetes infrastructure.44:25

AI Agents: The Future of Productivity, or Just a Fad?Sam Partee, Arcade AI · 35:18 · Dec 2024 · 728 views

ClaimSam Partee says Arcade's tool SDK creates a Python package for writing and serving a tool function to a large language model.20:30