Data quality in 2020

19 sessions

MeetupHierarchy of Machine Learning NeedsPhil Winder, Winder Research · 58:26 · Apr 2020 · 710 views · MLOps Meetup
MeetupMid-Scale Production Feature EngineeringDr. Venkata Pingali, Scribble Data · 1:01:35 · Apr 2020 · 341 views · MLOps Meetup

ClaimDr. Venkata Pingali says model reproducibility is insufficient without reproducibility and lineage for the data used by the model.13:13

MeetupTrueLayer's MLOps PipelineAlex Spanos, TrueLayer · 56:17 · Apr 2020 · 217 views · MLOps Meetup

Pushed backAlex Spanos says the machine learning pipeline should eventually resemble mature DevOps practice, while noting that machine learning has additional moving parts such as data versioning, parameters, and metrics.51:03

Meetup10 Years Deploying ML in the Enterprise: The Inside Scoop!Charles Martin, MLOps Community · 1:02:48 · May 2020 · 135 views · MLOps Meetup

ClaimCharles Martin says that unvalidated and changing data inputs make it difficult to automate machine learning systems reliably.29:06

MeetupWhy Data Scientists Should Know Data EngineeringDan Sullivan · 58:28 · May 2020 · 232 views · MLOps Meetup

ClaimData engineering can help data scientists clean and explore data more efficiently by providing tools and techniques for working at larger scales.52:26

MeetupScaling Human-in-the-Loop Machine LearningRobert Munro · 55:04 · Jun 2020 · 688 views · MLOps Meetup

ClaimHuman disagreement can be useful when multiple responses are valid.20:01

MeetupMonitoring the Machine Learning StackLina Weichbrodt, DKB · 55:32 · Jul 2020 · 894 views · MLOps Meetup

Pushed backLina Weichbrodt said that real-time response monitoring is needed in addition to offline data-quality checks such as Great Expectations or TensorFlow Data Validation.21:00

MeetupFeature Stores: An Essential Part of the ML Stack to Build Great DataKevin Stumpf, Tecton · 1:05:46 · Jul 2020 · 1,511 views · MLOps Meetup

ClaimKevin Stumpf identifies data scientists, production training pipelines and low-latency production models as the main consumers of feature-store data.8:41

MeetupML ObservabilityAparna Dhinakaran, Arize AI · 55:04 · Jul 2020 · 1,766 views · MLOps Meetup

ClaimAparna Dhinakaran says production monitoring should address issues including drift, distribution changes, black-box behavior, data quality problems, and previously unseen feature values.15:39

PodcastA Conversation Around Feature StoresVenkata Pingali, Scribble Data · 1:03:18 · Jul 2020 · 256 views · MLOps Coffee Sessions

ClaimA feature store is a data platform for AI that gives data scientists an API for obtaining and transforming data without requiring them to build data pipelines themselves.4:54

MeetupCreating Beautiful Ambient Music with Google Brain's Music TransformerDaniel Jeffries, Pachyderm · 55:52 · Aug 2020 · 291 views · MLOps Meetup

ClaimDaniel Jeffries says the project encountered outdated libraries, hard-coded variables, broken code, failed MIDI-generation approaches, dependency problems, pipeline freezes, corrupted data, and bugs.41:41

MeetupStreaming Machine Learning with Apache Kafka and Tiered StorageKai Waehner, Confluent · 52:50 · Sept 2020 · 487 views · MLOps Meetup
PodcastAnalyzing "Continuous Delivery and Automation Pipelines in ML", Part 3David Ponte, Benevolent AI · 1:06:28 · Oct 2020 · 294 views · MLOps Coffee Sessions

ClaimThe session examines the Google paper “Continuous Delivery and Automation Pipelines in ML,” focusing on data science steps and machine learning maturity levels.1:01

MeetupMLOps #37 When You Say Data Scientist Do You Mean Data Engineer? Lessons Learned From Startup LifeElizabeth Chabot, Deloitte · 1:00:43 · Oct 2020 · 404 views · MLOps Meetup

ClaimCompanies should bring in an MLOps or data engineering person when their product needs a maintained and monitored machine learning pipeline.16:21

PodcastData Engineering + ML + Software EngineeringSatish Chandra Gupta, Slang Labs · 57:05 · Oct 2020 · 429 views · MLOps Coffee Sessions

ClaimSatish Chandra Gupta said a data pipeline should distinguish needs that require real-time processing from needs that can tolerate latency, because real-time processing costs more.27:20

MeetupUN Global PlatformMark Craddock, Global Certification and Training Ltd (GCATI) · 58:48 · Nov 2020 · 240 views · MLOps Meetup
PodcastIntroducing Data Downtime: From Firefighting to WinningBarr Moses, Monte Carlo · 1:00:51 · Nov 2020 · 429 views · MLOps Coffee Sessions

ClaimBarr Moses’s five pillars of data observability are freshness, volume, schema, distribution, and lineage.31:01

MeetupThe Current MLOps LandscapeNathan Benaich, Air Street Capital & Timothy Chen, Essence VC · 58:31 · Nov 2020 · 1,510 views · MLOps Meetup

ClaimNathan Benaich says the most active MLOps areas are data labeling and annotation, data quality, monitoring, and explainability.8:20

PodcastDeep in the Heart of DataCarl Steinbach, LinkedIn · 55:27 · Dec 2020 · 509 views · MLOps Coffee Sessions

ClaimCarl Steinbach says user-defined functions are black boxes to query optimizers, which limits optimization and makes data lineage and impact analysis less precise.17:25