Serving and deployment
The mechanics of getting a model to answer requests, from Flask apps and Kubernetes to inference servers and hosted model APIs.
Building an ML Platform at SurveyMonkey
Pushed backShubhi Jain said SurveyMonkey's inference architecture did not run on Kubernetes.34:48
TrueLayer's MLOps Pipeline
Pushed backAlex Spanos favors simple interpretable algorithms in regulated financial services, rather than moving to more complex models when interpretability is reduced.19:45
10 Years Deploying ML in the Enterprise: The Inside Scoop!
Pushed backCharles Martin disputes the idea that running a machine learning model inside a database such as SQL Server is a simple solution.27:20
Machine Learning at Scale in Mercado Libre
Pushed backCarlos de la Torre argued that Fury Data Apps should not automate model deployment because responsibility for production deployment must remain explicit.45:30
Lessons Learned From Hosting the ML Engineered Podcast
Pushed backCharlie You argues that the hardest or most differentiating part of production machine learning is often not the algorithm but data, deployment and maintenance.58:57
MLOps Engineering Labs Recap, Part 1
Pushed backMichel Vasconcelos argues that MLflow is a strong experimentation tool but is not sufficient by itself for large-scale, containerized deployments.34:54
Operationalizing Machine Learning at a Large Financial Institution
Pushed backDaniel Stahl disputes the assumption that models should be deployed as microservices by default, arguing that batch processing was simpler and sufficient for Regions' current needs.28:08
Law of Diminishing Returns for Running AI Proof-of-Concepts
Pushed backOguzhan Gencoglu says the most crucial role skill is problem translation, not model training or deployment, which may be controversial.45:06
Federated Learning: Machine Learning on the Edge
Pushed backThe speaker disputed the idea that federated learning is not production-ready by describing existing large-scale deployments in products.40:02
Platform Thinking: A Lemonade Case Study
Pushed backAutomatic model testing was not considered sufficient to make the team comfortable deploying a model, so people still had to take responsibility for testing.41:17
MLOps Critiques
Pushed backMatthijs Brouns rejects the idea that model deployment is simply an already solved problem, arguing that important gaps remain in MLOps.17:08
MLflow vs Kubeflow 2022
Pushed backGeorge Pearse argues that Kubeflow's complexity makes it less suitable for small shops, while Byron Allen says Kubeflow has a valid role through managed services and in larger organizations.45:30
Scalable Evaluation and Serving of Open Source LLMs
Pushed backThe displayed model cost and performance estimates do not fully account for batching, which can change results substantially.10:19
Build and Customize LLMs in Less than 10 Lines of YAML
Pushed backTravis argued that fine-tuning a smaller model can match or outperform a much larger model for a sufficiently bounded task at lower latency and cost.17:17
LLM on Kubernetes
Pushed backRahul Parundekar argues that companies should prioritize a repeatable deployment platform and model iteration over over-optimizing Kubernetes autoscaling while GPUs are scarce.16:42
Using Large Language Models at AngelList
Pushed backThe idea that off-the-shelf machine-learning services were sufficient for AngelList was rejected because they were too expensive, could not scale as needed, and lacked the required extraction capabilities.12:57
Graduating from Proprietary to Open Source Models in Production
Pushed backPhilip Kiely does not give a definite position on whether companies will serve models with alternative hardware such as Groq.22:02
Productionizing Health Insurance Appeal Generation
Pushed backHolden Karau rejects using the latest container version as a good deployment practice and says the version should be pinned.11:50
Streamlining Model Deployment
Pushed backThe discussion disputes the assumption that selecting the same model across endpoint providers necessarily produces the same output quality.17:35
Productionizing AI: How to Think From the End
Pushed backThe idea that deploying LLMs has a quick and dirty approach was challenged because production deployment is not actually simple.10:17
Kubernetes, AI Gateways, and the Future of MLOps
Pushed backDemetrios Brinkmann suggested that KServe is built on KNative and helps provision model resources, while Alexa Griffith clarified that KNative can run serverless services generally and KServe is specifically focused on AI and ML models.19:39
MLOps with Databricks
Pushed backMaria Vechtomova disagrees with the idea that Databricks should be used for every model-serving situation.26:50
The Rise of Sovereign AI and Global AI Innovation in a World of US Protectionism
Pushed backFrank Meehan responds that anonymized research data can be shared, but identifiable medical data used for patient care and internal health-system inference requires local protection.19:02
Integration of AI into Traditional Systems
Pushed backHakan Tek rejects the idea that using a third-party service can provide complete certainty about data security.14:30
From Notebooks to Production FASTER
ClaimThe team serves about 200 data scientists and develops platform tools with their feedback.2:01
Speed and Scale: How Today's AI Datacenters Are Operating Through Hypergrowth
ClaimNetBox is the system of record for infrastructure, modeling everything from space, power, and cooling through racks, GPU servers, switching, cabling, IP addresses, configurations, and automation.1:39
Fast & Asynchronous: Drift Your AI, Not Your GPU Bill
ClaimData scientists use plain Python functions that take dictionaries and return dictionaries, while platform teams manage the deployment and autoscaling configuration.12:01
Why AI Agents Shouldn't Replace Your Fraud Models
ClaimHigh-stakes systems need auditability, high query throughput, and low latency, so they are better suited to rules engines and machine learning models than to full agentic decisioning.6:49