Serving and deployment in 2023

33 sessions

PodcastHow A Manager Became a Believer in DevOps for Machine LearningKeith Trnka, 98.6 · 55:49 · Apr 2023 · 361 views · MLOps Podcast

ClaimKeith Trnka says projects usually fail because they do not address user or business needs, or because the surrounding deployment, reliability, monitoring, and rollback work is inadequate.5:31

PodcastML Scalability ChallengesWaleed Kadous, Anyscale · 1:00:03 · Apr 2023 · 838 views · MLOps Podcast

ClaimWaleed Kadous says machine-learning infrastructure should let developers move from development to scalable serving while retaining their preferred Python libraries.25:18

MeetupDeclarative MLOps: Streamlining Model Serving on KubernetesRahul Parundekar, AI Hero · 58:58 · Apr 2023 · 2,634 views · MLOps Meetup

ClaimRahul Parundekar says Kubernetes lets an ML engineer declare a target layout, including model-server and backend replicas, and then handles scheduling and orchestration.7:24

Cost Optimization and PerformanceLina Weichbrodt & Luis Ceze, OctoML & Jared Zoneraich, Prompt Layer & Daniel Campos, Neeva & Mario Kostelac, Intercom · 36:06 · May 2023 · 930 views · LLMs in Production 2023

ClaimBringing models in-house gives a company more control over latency optimization, API rate limits, and service availability.6:40

Data Privacy and SecurityDiego Oppenheimer, Factory & Gevorg Karapetyan, ZERO Systems & Vin Vashishta, V Squared & Saahil Jain, U.com & Shreya Rajpal · 25:44 · May 2023 · 651 views · LLMs in Production 2023
No Rose Without a Thorn - Obstacles to Successful LLM DeploymentsTanmay Chopra, Neeva · 10:24 · May 2023 · 606 views · LLMs in Production 2023

ClaimTanmay Chopra divides the obstacles to deploying LLMs into production into infrastructure challenges and output-related challenges.0:46

PodcastWhy is MLOps Hard in an Enterprise?Maria Vechtomova & Basak Eskili, Ahold Delhaize · 55:06 · May 2023 · 756 views · MLOps Podcast

ClaimMaria Vechtomova and Basak Eskili say their standardized deployment process makes it possible to deploy a model to another Ahold Delhaize brand in a few minutes, although the model still needs to be adapted to the brand's data.13:30

PodcastClean Code for Data ScientistsMatt Sharp, Shopify · 46:13 · Jun 2023 · 1,301 views · MLOps Podcast

ClaimMatt Sharp says Merlin provides command-line tools, dependency management, image creation, boilerplate code, and automatic scaling for machine learning services.31:43

PodcastThe Long Tail of ML DeploymentTuhin Srivastava, Baseten · 50:37 · Jun 2023 · 618 views · MLOps Podcast

ClaimQuantumBlack is technology agnostic and helps organizations improve the process of building and deploying AI while reducing friction between teams.2:03

Large Model Training and Inference with DeepSpeedSamyam Rajbhandari, Microsoft DeepSpeed · 36:23 · Jun 2023 · 9,517 views · LLMs in Production 2023

ClaimDeepSpeed is a library for training, compression, and inference of large models.1:59

Scalable Evaluation and Serving of Open Source LLMsWaleed Kadous, Anyscale · 34:57 · Jul 2023 · 1,336 views · LLMs in Production 2023

Pushed backThe displayed model cost and performance estimates do not fully account for batching, which can change results substantially.10:19

Taking ImgFlip's 'This Meme Does Not Exist' to the Next Level with a LLMStefan Ojanen, Genesis Cloud · 14:49 · Jul 2023 · 307 views

ClaimThe project uses ImgFlip's meme dataset to build a content-aware LLM that can create quality memes and cover more templates than the existing service.4:55

PodcastOpen Source and Fast Decision MakingRob Hirschfeld, RackN · 1:00:02 · Jul 2023 · 254 views · MLOps Podcast

ClaimGiving narrowly focused services access to entire repositories, drives, or communication systems creates a serious security risk.11:26

Build and Customize LLMs in Less than 10 Lines of YAMLTravis Addair, Predibase · 34:05 · Jul 2023 · 820 views · LLMs in Production 2023

Pushed backTravis argued that fine-tuning a smaller model can match or outperform a much larger model for a sufficiently bounded task at lower latency and cost.17:17

LLMs For the Rest of UsVikram Sreekanti, Aqueduct & Joseph Gonzalez, UC Berkeley and Aqueduct · 24:33 · Jul 2023 · 349 views · LLMs in Production 2023
Challenges in Providing LLMs as a ServiceHemant Jain, Cohere · 11:43 · Jul 2023 · 466 views · LLMs in Production 2023

ClaimServing large language models can require splitting them across multiple accelerators, which creates communication and computation trade-offs.1:52

LLM on KubernetesShrinand Javadekar, Outerbounds & Manjot Pahwa, Lightspeed India & Rahul Parundekar, AI Hero & Patrick Barker · 36:24 · Jul 2023 · 1,123 views · Conference in Production 2023

Pushed backRahul Parundekar argues that companies should prioritize a repeatable deployment platform and model iteration over over-optimizing Kubernetes autoscaling while GPUs are scarce.16:42

Unleashing Code Completion with LLMsMonmayuri Ray, GitLab · 17:41 · Aug 2023 · 602 views · LLMs in Production 2023

ClaimModel selection for code completion should consider the objective, training data, evaluation benchmarks, model weights, tuning frameworks, cost, and latency.5:54

End-to-end Modern Machine Learning in ProductionOmar Sanseviero, Hugging Face · 10:30 · Aug 2023 · 753 views · LLMs in Production 2023

ClaimOmar Sanseviero aims to increase awareness of tools that make it easier to use state-of-the-art machine learning models in products and services.0:47

PodcastUsing Large Language Models at AngelListThibaut Labarre, AngelList · 51:42 · Aug 2023 · 931 views · MLOps Podcast

Pushed backThe idea that off-the-shelf machine-learning services were sufficient for AngelList was rejected because they were too expensive, could not scale as needed, and lacked the required extraction capabilities.12:57

Considerations and Optimizations for Deploying Open Source LLMs at Your CompanyOscar Rovira, Mystic AI · 11:31 · Aug 2023 · 474 views

ClaimDeploying an open-source LLM as a fast, secure, scalable API endpoint is broadly a software engineering problem similar to deploying other machine-learning models, but it may require much more memory.1:03

Enabling Defense Missions with Local LLMsGerred Dillon, Defense Unicorns · 23:07 · Aug 2023 · 408 views · LLMs in Production 2023

ClaimRegulated environments are often restricted, isolated, and controlled for ingress and egress, and some are completely air-gapped or deployed at the edge.2:59

Preemption Chaos and Optimizing Server StartupBradley Heilbrun, Replit · 12:42 · Aug 2023 · 215 views · LLMs in Production 2023

ClaimServing large language models at low latency requires powerful GPUs, and larger models require better and more expensive GPUs.2:30

LLMs vs LMs in ProductionDenys Linkov, Voiceflow · 24:44 · Aug 2023 · 1,337 views · LLMs in Production 2023

ClaimVoiceflow's custom NLU model outperformed GPT-4 on both cost and accuracy in one test because GPT-4 inference cost much more.22:43

PodcastHarnessing MLOps in FinanceMichelle Marie Conway, Lloyds Banking Group · 1:05:16 · Sept 2023 · 503 views · MLOps Podcast
PodcastBuilding an ML Platform: Insights, Community, and AdvocacyStephen Batifol, Wolt · 45:49 · Oct 2023 · 568 views · MLOps Podcast

Pushed backStephen Batifol says deployed machine learning models should be treated as ordinary software for operations and on-call rather than as a separate special category.43:20

PodcastMLOps at GetYourGuideJean Machado, Meghana Satish, Olivia Houghton & Theodore Meynard, GetYourGuide · 1:03:53 · Oct 2023 · 327 views · MLOps Podcast

Pushed backJean Machado argued that people should not be blocked from using LLMs because they do not use the platform's Python-centered templates, although the platform will provide managed tooling for use cases that need observability and controls.54:11

Fireside Chat with LLM StartupsPaul van der Boor & Sandeep Bakshi, Prosus Group & Shriyash Upadhyay, Martian & Lars Maaløe, Corti & Pietro Gagliano, Transitional Forms · 30:46 · Oct 2023 · 584 views · LLMs in Production 2023
Efficient Serving of LLMs for Experimentation and Production with Fireworks.aiDmytro Dzhulgakov, Fireworks.ai · 11:43 · Oct 2023 · 1,085 views

ClaimFine-tuning can reduce serving costs because it enables shorter prompts and sometimes smaller models for the same quality.1:59

Building RAG-based LLM Applications for ProductionPhilipp Moritz & Yifei Feng, Anyscale · 30:23 · Nov 2023 · 3,003 views · LLMs in Production 2023

ClaimRay Serve was used to deploy the application because it supports composing embedding models, ranking models, and language models, as well as end-to-end streaming.15:59

The State of Open Source AI: Deployment Engines, Licences, & HardwareCasper da Costa-Luis, Prem AI · 10:56 · Nov 2023 · 253 views

ClaimCasper da Costa-Luis says international web-service providers may need to handle different legal requirements across countries, or avoid serving some countries.4:53

PodcastDesigning for Forward Compatibility in Gen AIRohit Agarwal, Portkey.ai · 1:00:18 · Nov 2023 · 382 views · MLOps Podcast

ClaimRohit Agarwal says open-source model deployments are difficult because compute demand is hard to manage when traffic has spikes and troughs.13:56

PodcastChallenges Operationalizing ML (And Some Solutions)Nathan Ryan Frank, WW Grainger · 52:28 · Dec 2023 · 571 views · MLOps Podcast

Pushed backNathan Ryan Frank presents notebooks as useful for exploration but argues that their work should be moved into shareable, testable, deployable structures rather than left as an unstructured production artifact.21:47