PodcastMost Underrated MLOps TopicsClaimMarian Ignev says that simple deployment examples often omit monitoring, data drift, and concept drift.18:28
38 sessions
PodcastMost Underrated MLOps TopicsClaimMarian Ignev says that simple deployment examples often omit monitoring, data drift, and concept drift.18:28
PodcastLessons Learned From Hosting the ML Engineered PodcastClaimMachine learning projects require managing data, model life cycles, data drift and monitoring in addition to training and shipping models.7:23
MeetupHow Explainable AI is Critical to Building Responsible AIClaimResponsible AI requires processes covering training-data assessment, model validation, monitoring, and debugging.13:01
MeetupProduct Management in Machine LearningClaimLaszlo Sragner says monitoring, evaluation, and labeling are tightly coupled and should happen as one continuous process for machine learning models.9:51
MeetupHow to Avoid Suffering in MLOps/Data Engineering RoleClaimIgor Lushchyk says model monitoring includes both infrastructure and service metrics and monitoring the model's inference performance.46:12
MeetupOperationalizing Machine Learning at a Large Financial InstitutionClaimDaniel Stahl identifies training, scoring, and monitoring as the three core pipelines needed across model types and architectures.10:52
MeetupA Missing Link in the ML Infrastructure StackClaimFull Stack Deep Learning was created to teach practitioners about production machine learning topics that typical machine learning classes did not cover, including scoping, testing, debugging, deployment, and monitoring.6:26
MeetupModel Watching: Keeping Your Project in ProductionPushed backBen Wilson argues that the tools used for drift monitoring matter less than knowing which kinds of drift and statistical behavior to monitor.18:57
PodcastMLOps InvestmentsPushed backSarah Catanzaro says industry and academia both contribute to the gap between research and practical ML because industry rarely provides realistic structured-data benchmarks and context.43:06
MeetupDeploying Machine Learning Models at Scale in CloudClaimThe team uses SageMaker and AWS as building blocks for a pipeline that covers feature engineering, cleaning, analysis, continuous retraining, CI/CD promotion, checks between environments, and production monitoring.10:50
PodcastLuigi in Production Part 2ClaimLuigi Patruno's team moved several models from development notebooks into production processes with data validation, drift monitoring, and ongoing monitoring.8:02
MeetupFrom Idea to Production MLClaimMichael Munn says teams should define business goals and evaluation metrics early so model development does not become an unproductive exercise in complexity.10:51
PodcastScaling AI in ProductionClaimSrivatsan Srinivasan says machine-learning algorithms make up only a small part of the total machine-learning work, which also includes data collection, deployment, monitoring, pipelines, feature engineering, and feature stores.1:45
MeetupOperationalize Machine Learning at Scale with MLOpsClaimChristopher Bergh says source code should be versioned, production systems should be monitored, development systems should be tested, and infrastructure should be scriptable, testable, and deployable.14:06
PodcastModel Performance Monitoring and Why You Need it YesterdayClaimModel performance management covers visibility across the entire model life cycle, including training, validation, deployment, monitoring, and analysis.31:52
MeetupPractical MLOps Part 2ClaimAlfredo Deza learned Bash scripting after Carlos Cole encouraged him to write basic server-monitoring code and offered to answer questions for 15 minutes each day.6:42
PodcastMaturing Machine Learning in EnterpriseClaimKyle Gallatin says machine learning has moved from isolated proof-of-concept data science projects toward MLOps, and that governance, observability, and visibility are the next stage.10:13
MeetupEngineering MLOpsClaimA robust CI/CD pipeline should treat the pipeline rather than the model as the final product and should include quality assurance, testing, release strategies, governance, and monitoring.12:35
How Pinterest Powers Image SimilarityClaimPinterest focuses on building machine-learning systems that can be productionized, monitored, debugged, explained, and rolled back when signals or models cause problems.4:25
PodcastLearning from 150 Successful ML-enabled Products at Booking.comClaimPablo Estevez says model monitoring should consider feature drift, concept drift, delayed feedback, and response distributions rather than only comparing predictions with outcomes.26:20
MeetupBuilding ML Blocks with Kubeflow Orchestration with Feature StorePushed backAniruddha Choudhury distinguishes a feature store from a SQL database by emphasizing low-latency online retrieval, feature consistency, and support for batch and streaming ingestion.53:23
MeetupBuilding an ML Platform from Scratch: Live Coding Session - Part 2ClaimAlon Gubkin says the session will cover training orchestration and model monitoring.4:23
mlctl and Hydrosphere Open Source MLOps Libraries DemoPushed backAlex Chung questioned whether Hydrosphere should continue serving models or focus on monitoring and integration with existing tools, and recommended the latter focus.31:41
PodcastMachine Learning SREClaimSLOs and production monitoring for machine learning are still developing because the SRE community lacks mature ways to compare different systems and define what good means across them.6:36
PodcastA Few Learnings from Building a Bootstrapped MLOps Services StartupClaimFor companies moving from one to ten, a strong engineering culture, testing, monitoring and a proper MLOps process become increasingly important.10:17
MeetupDoing MLOpsClaimMLOps is a feedback loop involving source control, testing, infrastructure as code, deployment, monitoring, and model retraining.12:05
MeetupEnd to End MLOps BasicsClaimThe MLOps lifecycle includes model development, training operationalization, continuous training, model deployment, prediction serving, continuous monitoring, and data and model management.5:57
PodcastLinkedIn Job RecommendationsClaimLinkedIn formed a team of linguists to refine the definition of a bad job recommendation and evaluate recommendation quality across product surfaces.23:23
MeetupMLOps at Volvo CarsClaimLeonard Aukea says machine-learning monitoring needs to be easier for data scientists because configuring custom metrics through Prometheus and Grafana requires too many awkward steps.19:26
Reading groupImpact of SWE in ML ProjectsPushed backLaszlo Sragner emphasized that monitoring and model-performance analysis cannot generally be fully codified, while another participant argued for integrating software, statistical, behavioral, and acceptance tests into one workflow.40:24
MeetupModel Monitoring: The Million Dollar ProblemPushed backThe presenters did not identify one universally best monitoring tool; they said the choice depends largely on the use case, ecosystem, integrations, and personal preference.49:40
MeetupThe Not So Talked About Reasons Model Monitoring FailsPushed backOren Razon disputes treating model observability as only a data scientist's responsibility.32:58
PodcastML Stepping Stones: Challenges & Opportunities for CompaniesClaimJohn Crousse says not every feature containing machine learning has a business case for constant monitoring or incremental improvement.15:47
Reading groupThe ML Test ScorePushed backSkylar Payne disputed the idea that monitoring products can generally infer useful thresholds automatically without iteration.29:00
MeetupML Drift: How to Identify Issues Before They Become ProblemsPushed backAmy Hodler separates concept drift from data drift, while noting that some people consider concept drift a type of data drift.11:11
Podcast2022 Predictions for MLOps and the IndustryClaimReah Miyara says machine-learning observability can proactively identify problems, support root-cause analysis, and help organizations troubleshoot and improve models.9:05
MeetupSetting up an ML Platform on GCP: Lessons LearnedClaimCloud Composer provided scheduled pipeline execution, retries, alerting, and improved pipeline resilience and observability.19:49