MeetupHierarchy of Machine Learning NeedsData quality in 2020
19 sessions
MeetupMid-Scale Production Feature EngineeringClaimDr. Venkata Pingali says model reproducibility is insufficient without reproducibility and lineage for the data used by the model.13:13
MeetupTrueLayer's MLOps PipelinePushed backAlex Spanos says the machine learning pipeline should eventually resemble mature DevOps practice, while noting that machine learning has additional moving parts such as data versioning, parameters, and metrics.51:03
Meetup10 Years Deploying ML in the Enterprise: The Inside Scoop!ClaimCharles Martin says that unvalidated and changing data inputs make it difficult to automate machine learning systems reliably.29:06
MeetupWhy Data Scientists Should Know Data EngineeringClaimData engineering can help data scientists clean and explore data more efficiently by providing tools and techniques for working at larger scales.52:26
MeetupScaling Human-in-the-Loop Machine LearningClaimHuman disagreement can be useful when multiple responses are valid.20:01
MeetupMonitoring the Machine Learning StackPushed backLina Weichbrodt said that real-time response monitoring is needed in addition to offline data-quality checks such as Great Expectations or TensorFlow Data Validation.21:00
MeetupFeature Stores: An Essential Part of the ML Stack to Build Great DataClaimKevin Stumpf identifies data scientists, production training pipelines and low-latency production models as the main consumers of feature-store data.8:41
MeetupML ObservabilityClaimAparna Dhinakaran says production monitoring should address issues including drift, distribution changes, black-box behavior, data quality problems, and previously unseen feature values.15:39
PodcastA Conversation Around Feature StoresClaimA feature store is a data platform for AI that gives data scientists an API for obtaining and transforming data without requiring them to build data pipelines themselves.4:54
MeetupCreating Beautiful Ambient Music with Google Brain's Music TransformerClaimDaniel Jeffries says the project encountered outdated libraries, hard-coded variables, broken code, failed MIDI-generation approaches, dependency problems, pipeline freezes, corrupted data, and bugs.41:41
PodcastAnalyzing "Continuous Delivery and Automation Pipelines in ML", Part 3ClaimThe session examines the Google paper “Continuous Delivery and Automation Pipelines in ML,” focusing on data science steps and machine learning maturity levels.1:01
MeetupMLOps #37 When You Say Data Scientist Do You Mean Data Engineer? Lessons Learned From Startup LifeClaimCompanies should bring in an MLOps or data engineering person when their product needs a maintained and monitored machine learning pipeline.16:21
PodcastData Engineering + ML + Software EngineeringClaimSatish Chandra Gupta said a data pipeline should distinguish needs that require real-time processing from needs that can tolerate latency, because real-time processing costs more.27:20
PodcastIntroducing Data Downtime: From Firefighting to WinningClaimBarr Moses’s five pillars of data observability are freshness, volume, schema, distribution, and lineage.31:01
MeetupThe Current MLOps LandscapeClaimNathan Benaich says the most active MLOps areas are data labeling and annotation, data quality, monitoring, and explainability.8:20
PodcastDeep in the Heart of DataClaimCarl Steinbach says user-defined functions are black boxes to query optimizers, which limits optimization and makes data lineage and impact analysis less precise.17:25

