PodcastData Mesh: Data Quality Control Mechanism for MLOps?Data quality in 2022
17 sessions
MeetupTrustworthy Data for Machine LearningPushed backChad Sanderson initially argued that data scientists should own the quality of the modeling code they write, but later concluded that data scientists are not software engineers and that engineers should own data quality.42:26
MeetupOrchestrating Machine Learning Workflows with PrefectClaimWorkflow orchestration helps define how pipelines respond to failures such as failed API calls, unavailable databases, nonconverging models, and malformed data.6:24
MeetupBuilding a Modern Data Analytics StackClaimFivetran's Mixpanel schema includes an event identifier, a distinct user identifier, the event name, and the event time.14:17
PodcastMLOps as Tool to Shape Team and CultureClaimCiro Greco considers data quality more important than marginal improvements to model quality in Coveo's situation.17:07
MeetupDataOps is a Software Engineering ChallengePushed backMicha Kunze argues that full-job or pipeline tests are often more valuable and stable than testing every individual function, although he still uses unit tests for some complicated transformations.51:23
PodcastFixing Your ML Data Blind SpotsClaimYash Sheth says more advanced teams use Galileo to build active-learning pipelines that identify new production samples to annotate and add to their datasets.36:37
MeetupThe Post Modern StackPushed backJacopo Tagliabue disputes the view that building an ML pipeline requires a very large team or a million people.9:20
PodcastLabeled Datasets that Correct Themselves AutomaticallyPushed backCleanlab's results are not based only on model mistakes, because the model also reflects errors in the data it was trained on.29:55
MeetupSo Fresh and So Data CleanClaimMage is an open-source code editor that helps users transform data and build machine learning pipelines.3:24
PodcastData Engineering for MLPushed backJosh Wills argued that data quality and contracts should not be left entirely to downstream data teams because upstream producers also need responsibility for quality.39:40
MeetupFrom Expectations to Synthetic Data GenerationClaimThe workflow uses YData Synthetic for generation, pandas profiling for exploratory analysis, and Great Expectations for validation.3:42
MeetupDriving ML Data Quality with Data ContractsClaimGoCardless uses data contracts to improve data quality.4:05



