Data quality in 2022

17 sessions

PodcastData Mesh: Data Quality Control Mechanism for MLOps?Scott Hirleman, DataStax · 57:03 · Jan 2022 · 632 views · MLOps Coffee Sessions
MeetupTrustworthy Data for Machine LearningChad Sanderson, Convoy · 51:04 · Feb 2022 · 516 views · MLOps Meetup

Pushed backChad Sanderson initially argued that data scientists should own the quality of the modeling code they write, but later concluded that data scientists are not software engineers and that engineers should own data quality.42:26

PodcastBetter Use Cases for Text EmbeddingsVincent Warmerdam, Explosion · 48:20 · Feb 2022 · 413 views · MLOps Coffee Sessions
MeetupOrchestrating Machine Learning Workflows with PrefectKevin Kho, Prefect · 1:04:18 · Mar 2022 · 4,216 views · MLOps Meetup

ClaimWorkflow orchestration helps define how pipelines respond to failures such as failed API calls, unavailable databases, nonconverging models, and malformed data.6:24

MeetupBuilding a Modern Data Analytics StackJeff Katz, Jigsaw Labs · 55:21 · Mar 2022 · 773 views · MLOps Meetup

ClaimFivetran's Mixpanel schema includes an event identifier, a distinct user identifier, the event name, and the event time.14:17

PodcastMLOps as Tool to Shape Team and CultureCiro Greco, Coveo · 43:02 · Apr 2022 · 439 views · MLOps Coffee Sessions

ClaimCiro Greco considers data quality more important than marginal improvements to model quality in Coveo's situation.17:07

PodcastReal-Time Processing with Apache Flink, Kafka, and PinotJacob Tsafatinos, Elemy · 53:41 · May 2022 · 1,986 views · MLOps Coffee Sessions
MeetupDataOps is a Software Engineering ChallengeMicha Kunze, Maersk · 57:57 · May 2022 · 770 views · MLOps Meetup

Pushed backMicha Kunze argues that full-job or pipeline tests are often more valuable and stable than testing every individual function, although he still uses unit tests for some complicated transformations.51:23

PodcastFixing Your ML Data Blind SpotsYash Sheth, Galileo · 51:41 · Jun 2022 · 530 views · MLOps Coffee Sessions

ClaimYash Sheth says more advanced teams use Galileo to build active-learning pipelines that identify new production samples to annotate and add to their datasets.36:37

MeetupThe Post Modern StackJacopo Tagliabue, Coveo · 1:04:58 · Jun 2022 · 857 views · MLOps Meetup

Pushed backJacopo Tagliabue disputes the view that building an ML pipeline requires a very large team or a million people.9:20

PodcastLabeled Datasets that Correct Themselves AutomaticallyCurtis Northcutt, Cleanlab · 1:06:03 · Jul 2022 · 1,395 views · MLOps Coffee Sessions

Pushed backCleanlab's results are not based only on model mistakes, because the model also reflects errors in the data it was trained on.29:55

MeetupSo Fresh and So Data CleanTommy Dang, Mage · 49:58 · Jul 2022 · 568 views · MLOps Meetup

ClaimMage is an open-source code editor that helps users transform data and build machine learning pipelines.3:24

PodcastData Engineering for MLChad Sanderson, Convoy · 57:54 · Aug 2022 · 773 views · MLOps Coffee Sessions

Pushed backJosh Wills argued that data quality and contracts should not be left entirely to downstream data teams because upstream producers also need responsibility for quality.39:40

MeetupFrom Expectations to Synthetic Data GenerationFabiana Clemente, YData · 55:39 · Aug 2022 · 527 views · MLOps Meetup

ClaimThe workflow uses YData Synthetic for generation, pandas profiling for exploratory analysis, and Great Expectations for validation.3:42

Monitoring Unstructured DataAparna Dhinakaran & Jason Lopatecki, Arize AI · 13:12 · Sept 2022 · 551 views · MLOps Lightning Sessions
MeetupDriving ML Data Quality with Data ContractsAndrew Jones, GoCardless · 34:30 · Nov 2022 · 744 views · MLOps Meetup

ClaimGoCardless uses data contracts to improve data quality.4:05

Podcast"Real-Time" ML: Features and InferenceSasha Ovsankin & Rupesh Gupta, LinkedIn · 51:55 · Dec 2022 · 610 views · MLOps Podcast