Threads / Serving and deployment

Serving and deployment

The mechanics of getting a model to answer requests, from Flask apps and Kubernetes to inference servers and hosted model APIs.

Follows the tags model-servingdeployment · 204 sessions · 2020 to 2026
202032 sessions
Meetup · MLOps Meetup #7

TrueLayer's MLOps Pipeline

Alex Spanos, TrueLayer

Pushed backAlex Spanos favors simple interpretable algorithms in regulated financial services, rather than moving to more complex models when interpretability is reduced.19:45

Meetup · MLOps Meetup #11

Machine Learning at Scale in Mercado Libre

Carlos de la Torre, Mercado Libre

Pushed backCarlos de la Torre argued that Fury Data Apps should not automate model deployment because responsibility for production deployment must remain explicit.45:30

28 more from 2020 on this thread
202144 sessions
Podcast · MLOps Coffee Sessions #30

MLOps Engineering Labs Recap, Part 1

John Savage, Overstock & Alexey Naiden & Varuna Jayasiri & Michel Vasconcelos, Bank of Nordeste

Pushed backMichel Vasconcelos argues that MLflow is a strong experimentation tool but is not sufficient by itself for large-scale, containerized deployments.34:54

40 more from 2021 on this thread
202237 sessions
Reading group · MLOps Reading Group #3

Federated Learning: Machine Learning on the Edge

Varun Kumar Khare, Nimble Edge

Pushed backThe speaker disputed the idea that federated learning is not production-ready by describing existing large-scale deployments in products.40:02

Podcast · MLOps Coffee Sessions #79

Platform Thinking: A Lemonade Case Study

Orr Shilon, Lemonade

Pushed backAutomatic model testing was not considered sufficient to make the team comfortable deploying a model, so people still had to take responsibility for testing.41:17

Podcast · MLOps Coffee Sessions #100

MLOps Critiques

Matthijs Brouns, Xccelerated.io

Pushed backMatthijs Brouns rejects the idea that model deployment is simply an already solved problem, arguing that important gaps remain in MLOps.17:08

Podcast · MLOps Coffee Sessions #108

MLflow vs Kubeflow 2022

Byron Allen, Contino

Pushed backGeorge Pearse argues that Kubeflow's complexity makes it less suitable for small shops, while Byron Allen says Kubeflow has a valid role through managed services and in larger organizations.45:30

33 more from 2022 on this thread
202333 sessions
Talk · Conference in Production 2023

LLM on Kubernetes

Shrinand Javadekar, Outerbounds & Manjot Pahwa, Lightspeed India & Rahul Parundekar, AI Hero & Patrick Barker

Pushed backRahul Parundekar argues that companies should prioritize a repeatable deployment platform and model iteration over over-optimizing Kubernetes autoscaling while GPUs are scarce.16:42

Podcast · MLOps Podcast #171

Using Large Language Models at AngelList

Thibaut Labarre, AngelList

Pushed backThe idea that off-the-shelf machine-learning services were sufficient for AngelList was rejected because they were too expensive, could not scale as needed, and lacked the required extraction capabilities.12:57

29 more from 2023 on this thread
202432 sessions
Talk · AI in Production 2024

Streamlining Model Deployment

Daniel Lenton, Unify

Pushed backThe discussion disputes the assumption that selecting the same model across endpoint providers necessarily produces the same output quality.17:35

28 more from 2024 on this thread
202519 sessions
Podcast · MLOps Podcast #294

Kubernetes, AI Gateways, and the Future of MLOps

Alexa Griffith, Bloomberg

Pushed backDemetrios Brinkmann suggested that KServe is built on KNative and helps provision model resources, while Alexa Griffith clarified that KNative can run serverless services generally and KServe is specifically focused on AI and ML models.19:39

Podcast · MLOps Podcast #314

MLOps with Databricks

Maria Vechtomova, Ahold Delhaize | Marvelous MLOps

Pushed backMaria Vechtomova disagrees with the idea that Databricks should be used for every model-serving situation.26:50

15 more from 2025 on this thread
20267 sessions
Talk · Coding Agents Conference 2026

Fast & Asynchronous: Drift Your AI, Not Your GPU Bill

Artem Yushkovskiy, Delivery Hero

ClaimData scientists use plain Python functions that take dictionaries and return dictionaries, while platform teams manage the deployment and autoscaling configuration.12:01

Talk

Why AI Agents Shouldn't Replace Your Fraud Models

Varant Zanoyan, Zipline AI

ClaimHigh-stakes systems need auditability, high query throughput, and low latency, so they are better suited to rules engines and machine learning models than to full agentic decisioning.6:49

3 more from 2026 on this thread