Threads / Data quality

Data quality

The oldest complaint in the archive, restated for every new kind of model.

Follows the tags data-qualitydata-pipelines · 128 sessions · 2020 to 2026
202019 sessions
Meetup · MLOps Meetup #6

Mid-Scale Production Feature Engineering

Dr. Venkata Pingali, Scribble Data

ClaimDr. Venkata Pingali says model reproducibility is insufficient without reproducibility and lineage for the data used by the model.13:13

Meetup · MLOps Meetup #7

TrueLayer's MLOps Pipeline

Alex Spanos, TrueLayer

Pushed backAlex Spanos says the machine learning pipeline should eventually resemble mature DevOps practice, while noting that machine learning has additional moving parts such as data versioning, parameters, and metrics.51:03

Meetup · MLOps Meetup #23

Monitoring the Machine Learning Stack

Lina Weichbrodt, DKB

Pushed backLina Weichbrodt said that real-time response monitoring is needed in addition to offline data-quality checks such as Great Expectations or TensorFlow Data Validation.21:00

15 more from 2020 on this thread
202123 sessions
Talk #8

The revolution of Federated Learning

Fabiana Clemente, MLOps Community & Ramen Dutta, TensoAI

ClaimSuccessful adoption of federated learning requires high-quality data, talented data scientists, and the ability to adapt to new data sets and quickly validate and release models.15:46

Podcast · MLOps Coffee Sessions #26

Machine Learning Feature Store Panel Discussion

Daniel Galinkin, iFood & Matias Dominguez, Rappi & Simarpal Khaira, Intuit

ClaimMatias Dominguez says a small company without a market-validated product may not need to buy or build a full feature store.9:29

Podcast · MLOps Coffee Sessions #61

ML Tests

Svet Penkov, Efemarai

Pushed backSvet Penkov takes the view that measuring data quality beyond basic validity is not always meaningful and that model performance in the intended domain is the more useful measure.29:34

Meetup · MLOps Meetup #90

Building 12-Factor Data Apps with Kedro

Ivan Danov, QuantumBlack

Pushed backIvan Danov says Kedro is not another orchestrator like Airflow or Kubeflow because it focuses on pipeline authoring rather than workflow execution and monitoring.39:06

19 more from 2021 on this thread
202217 sessions
Meetup · MLOps Meetup #93

Trustworthy Data for Machine Learning

Chad Sanderson, Convoy

Pushed backChad Sanderson initially argued that data scientists should own the quality of the modeling code they write, but later concluded that data scientists are not software engineers and that engineers should own data quality.42:26

Meetup · MLOps Meetup #100

DataOps is a Software Engineering Challenge

Micha Kunze, Maersk

Pushed backMicha Kunze argues that full-job or pipeline tests are often more valuable and stable than testing every individual function, although he still uses unit tests for some complicated transformations.51:23

Meetup · MLOps Meetup #103

The Post Modern Stack

Jacopo Tagliabue, Coveo

Pushed backJacopo Tagliabue disputes the view that building an ML pipeline requires a very large team or a million people.9:20

13 more from 2022 on this thread
202321 sessions
Meetup · MLOps Meetup #124

Dataframes Are All You Need: MLOps on Easy Mode

Jay Chia, Eventual

ClaimJay Chia says dataframes can support much of the MLOps workflow, including reading and writing data, exploring and processing it, feeding training pipelines, evaluating models, and running batch predictions.15:07

Talk · LLMs in Production 2023

Building LLM Applications for Production

Chip Huyen, Claypot AI

ClaimChip Huyen says LLMs do not reliably follow a required output schema, making it difficult for applications to parse their responses.5:19

17 more from 2023 on this thread
202437 sessions
Podcast · MLOps Podcast #226

The Rise of Modern Data Management

Chad Sanderson, Gable

Pushed backChad Sanderson says the main problem with data lineage is upstream of analytical systems, rather than downstream inside tools such as Snowflake or Databricks.29:07

Talk · DE4AI 2024

How Feature Stores Work

Simba Khadder, Featureform

Pushed backSimba Khadder argues that data scientists should not have to become expert data engineers to build production-grade feature pipelines.8:27

Talk

Building an ML Platform from scratch

Pushed backBen says feature transformations should ideally be centralized in SQL mesh or a similar modeling system, while Eric favors monolithic pipelines early and sees decoupling as a later maturity step.1:26:27

33 more from 2024 on this thread
20259 sessions
Podcast · MLOps Podcast #292

Look At Your ****ing Data 👀

Kenny Daniel, Hyperparam

Pushed backKenny argued that the lack of discussion about LLM data quality is not explained only by data sources being secret; researchers also rarely discuss how they clean and filter data.4:13

5 more from 2025 on this thread
20262 sessions