Data quality
The oldest complaint in the archive, restated for every new kind of model.
Mid-Scale Production Feature Engineering
ClaimDr. Venkata Pingali says model reproducibility is insufficient without reproducibility and lineage for the data used by the model.13:13
TrueLayer's MLOps Pipeline
Pushed backAlex Spanos says the machine learning pipeline should eventually resemble mature DevOps practice, while noting that machine learning has additional moving parts such as data versioning, parameters, and metrics.51:03
10 Years Deploying ML in the Enterprise: The Inside Scoop!
ClaimCharles Martin says that unvalidated and changing data inputs make it difficult to automate machine learning systems reliably.29:06
Monitoring the Machine Learning Stack
Pushed backLina Weichbrodt said that real-time response monitoring is needed in addition to offline data-quality checks such as Great Expectations or TensorFlow Data Validation.21:00
The revolution of Federated Learning
ClaimSuccessful adoption of federated learning requires high-quality data, talented data scientists, and the ability to adapt to new data sets and quickly validate and release models.15:46
Machine Learning Feature Store Panel Discussion
ClaimMatias Dominguez says a small company without a market-validated product may not need to buy or build a full feature store.9:29
ML Tests
Pushed backSvet Penkov takes the view that measuring data quality beyond basic validity is not always meaningful and that model performance in the intended domain is the more useful measure.29:34
Building 12-Factor Data Apps with Kedro
Pushed backIvan Danov says Kedro is not another orchestrator like Airflow or Kubeflow because it focuses on pipeline authoring rather than workflow execution and monitoring.39:06
Trustworthy Data for Machine Learning
Pushed backChad Sanderson initially argued that data scientists should own the quality of the modeling code they write, but later concluded that data scientists are not software engineers and that engineers should own data quality.42:26
DataOps is a Software Engineering Challenge
Pushed backMicha Kunze argues that full-job or pipeline tests are often more valuable and stable than testing every individual function, although he still uses unit tests for some complicated transformations.51:23
The Post Modern Stack
Pushed backJacopo Tagliabue disputes the view that building an ML pipeline requires a very large team or a million people.9:20
Labeled Datasets that Correct Themselves Automatically
Pushed backCleanlab's results are not based only on model mistakes, because the model also reflects errors in the data it was trained on.29:55
ML in Production: A DS from Ubisoft Perspective
ClaimJean-Michel Daignan says scalability testing can be more important for his pipelines than unit-testing every individual function.16:06
Multilingual Programming and a Project Structure to Enable It
ClaimDevelopment pipelines can call scripts in sequence or in parallel, record timings and logs, and give production engineers a simple command to run the required workflow.43:03
Dataframes Are All You Need: MLOps on Easy Mode
ClaimJay Chia says dataframes can support much of the MLOps workflow, including reading and writing data, exploring and processing it, feeding training pipelines, evaluating models, and running batch predictions.15:07
Building LLM Applications for Production
ClaimChip Huyen says LLMs do not reliably follow a required output schema, making it difficult for applications to parse their responses.5:19
The Rise of Modern Data Management
Pushed backChad Sanderson says the main problem with data lineage is upstream of analytical systems, rather than downstream inside tools such as Snowflake or Databricks.29:07
ML and AI as Distinct Control Systems in Heavy Industrial Settings
Pushed backRichard Howes rejects the idea that subject-matter professionals should be replaced in large-scale ML analysis and says they are needed to validate the outputs.38:35
How Feature Stores Work
Pushed backSimba Khadder argues that data scientists should not have to become expert data engineers to build production-grade feature pipelines.8:27
Building an ML Platform from scratch
Pushed backBen says feature transformations should ideally be centralized in SQL mesh or a similar modeling system, while Eric favors monolithic pipelines early and sees decoupling as a later maturity step.1:26:27
Look At Your ****ing Data 👀
Pushed backKenny argued that the lack of discussion about LLM data quality is not explained only by data sources being secret; researchers also rarely discuss how they clean and filter data.4:13
Streaming Ecosystem Complexities and Cost Management
ClaimA typical streaming pipeline connects Kafka to a processor such as Spark or Flink, then to storage and a serving layer, with each part requiring different skills.5:11
How Sama is Improving ML Models to Make AVs Safer
ClaimHuman work is usually the most expensive part of an AI data pipeline, except in specialized cases where data collection involves expensive equipment.4:35
GraphBI: Expanding Analytics to All Data Through the Combination of GenAI, Graph, & Visual Analytics
Pushed backWeidong Yang disputes treating ontology as universally valid and says its truth must be limited to a defined domain.21:00
Fixing GPU Starvation in Large-Scale Distributed Training
Pushed backKashish Mittal rejects the idea that low GPU utilization is mainly caused by an expensive model forward pass and says the data pipeline is the problem.9:08
Why AI Agents Shouldn't Replace Your Fraud Models
ClaimIn agentic experimentation, an agent can create features, add them to a model, train it, evaluate the results, and deploy an end-to-end pipeline to development.7:38