Data quality in 2024

37 sessions

PodcastFounding, Funding, and the Future of MLOpsMihail Eric, Storia AI · 57:31 · Jan 2024 · 377 views · MLOps Podcast

ClaimMihail Eric says data quality needs to be addressed end to end, from labeling through training and serving, with model predictions fed back into the data and labeling process.9:56

PodcastHow Data Platforms Affect ML & AIJake Watson, The Oakland Group · 39:12 · Jan 2024 · 468 views · MLOps Podcast

ClaimThe balance between modeling, data quality, and speed should depend on user requirements, business criticality, and the need to scale.24:32

PodcastLightweight Feature PlatformMatt Bleifer & Mike Eastham, Tecton · 1:03:58 · Feb 2024 · 251 views · MLOps Podcast

ClaimTecton is a feature platform that decouples feature engineering code from models and helps manage feature pipelines.2:57

PodcastAds Ranking Evolution at PinterestAayush Mudgal, Pinterest · 52:38 · Feb 2024 · 601 views · MLOps Podcast

ClaimConversion optimization required third-party integrations to collect, transform, identify, and pass conversion data into model-training pipelines.16:46

PodcastInformation Retrieval & RelevanceDaniel Svonava, Superlinked · 56:05 · Feb 2024 · 672 views · MLOps Podcast

ClaimBuilding a vector system involves combining data engineering and machine learning engineering, including deciding where models run and how inference services are accessed from data pipelines.24:44

PodcastManaging Data for Effective GenAI ApplicationAnu Arora & Anass Bensrhir, QuantumBlack, AI by McKinsey · 51:01 · Mar 2024 · 520 views · MLOps Podcast

ClaimData quality for generative AI must be checked at the input, prompt, and output stages.18:59

PodcastA Decade of AI Safety and TrustPetar Tsankov, LatticeFlow AI · 58:05 · Mar 2024 · 301 views · MLOps Podcast

ClaimData quality, model validation, and model improvement are tightly connected and should be addressed across the AI life cycle.27:09

PodcastThe Art and Science of Training LLMsBandish Shah & Davis Blalock, MosaicML/Databricks · 1:15:12 · Mar 2024 · 850 views · MLOps Podcast

ClaimDavis Blalock says tokenization can create invalid bytes, unexpected whitespace behavior, and evaluation failures.1:03:15

Building a Python-Centric Feature Platform to Power Production AI ApplicationsMatt Bleifer, Tecton · 27:11 · Apr 2024 · 296 views

ClaimTecton initially used Spark across batch pipelines, streaming pipelines, and training-data generation jobs, while managing orchestration, backfills, and retries for users.13:03

From MVP to ProductionEric Peter, Databricks & Donné Stevenson & Phillip Carter, Honeycomb & Andrew Hoh, Last Mile AI · 32:53 · Apr 2024 · 399 views · AI in Production 2024

ClaimAndrew Hoh says keeping knowledge current is fundamentally a version-control problem involving retrieval infrastructure, data sources, processing pipelines, and the underlying language model.25:56

Introducing DBRX: The Future of Language ModelsDavis Blalock, Bandish Shah, Abhi Venigalla & Ajay Saini, Databricks · 48:36 · Apr 2024 · 496 views · MLOps Coffee Sessions

ClaimThe team used PyTorch FSDP because it is flexible across model architectures and can distribute training across large clusters without requiring three-dimensional or pipeline parallelism.22:54

Innovative Gen AI Applications: Beyond TextDiana C. Montañes Mondragon & Nick Schenone, QuantumBlack · 54:45 · Apr 2024 · 852 views · MLOps Mini Summit 2024

ClaimThe call-center pipeline uses diarization, transcription and translation, personally identifiable information recognition, and large-language-model analysis.29:29

PodcastThe Rise of Modern Data ManagementChad Sanderson, Gable · 57:53 · Apr 2024 · 540 views · MLOps Podcast

Pushed backChad Sanderson says the main problem with data lineage is upstream of analytical systems, rather than downstream inside tools such as Snowflake or Databricks.29:07

Data Labeling Best PracticesCharles Brecque, TextMine · 12:59 · May 2024 · 682 views · AI in Production 2024
PodcastML and AI as Distinct Control Systems in Heavy Industrial SettingsRichard Howes, Metaformed · 56:31 · Jun 2024 · 313 views · MLOps Podcast

Pushed backRichard Howes rejects the idea that subject-matter professionals should be replaced in large-scale ML analysis and says they are needed to validate the outputs.38:35

Data Quality = Quality AISamuel Partee, Redis & Chad Sanderson, Gable & Joe Reis, Ternary Data & Maria Zhang, Proactive AI Lab Inc & Pushkar Garg, Clari · 27:15 · Aug 2024 · 335 views · AIQCON 2024

ClaimMaria Zhang says data quality means monitoring metrics such as completeness, accuracy, validity, and timeliness.0:50

PodcastBigQuery Feature StoreNicolas Mauti, Malt · 50:39 · Aug 2024 · 461 views · MLOps Podcast

ClaimMalt manages feature-table updates through SQL files in GitLab, with validation through a pull request and scheduled execution by Airflow.20:10

PodcastRAG Quality Starts with Data QualityAdam Kamor, Tonic.ai · 59:34 · Sept 2024 · 450 views · MLOps Podcast

ClaimTonic Textual is designed to build data pipelines for people creating retrieval-augmented generation systems.4:41

PodcastGlobal Feature Store: Optimizing Locally and Scaling Globally at Delivery HeroGottam Sai Bharath & Cole Bailey, Delivery Hero · 50:19 · Sept 2024 · 484 views · MLOps Podcast

ClaimCole Bailey said Logistics had built a real-time feature store using streaming pipelines, while Pandora had built a batch feature store using warehouse queries and caches.10:30

11 lessons learned from doing deploymentsSol Rashidi, ExecutiveAI LLC · 35:59 · Oct 2024 · 275 views · DE4AI 2024

ClaimData does not need to be perfect, but teams should prioritize use cases whose data is at least reasonably clean and accessible.27:19

Chronon: Airbnb's Open-Source Data Platform · 12:35 · Oct 2024 · 978 views

ClaimBuilding production machine-learning or prompt systems requires multiple pipelines and infrastructure components, including batch processing, streaming, key-value storage, serving, orchestration, and monitoring.5:19

Data Contracts: The Missing Piece of the Data PuzzleMark Freeman, Humu · 13:40 · Oct 2024 · 463 views

ClaimData contracts and data observability are both needed because they address different parts of data quality.1:14

Data Engineering: The Missing Piece of Your Data Science Puzzle · 12:10 · Oct 2024 · 184 views

ClaimThe infrastructure needs an online feature store for low-latency model serving and offline computation for batch models and training pipelines.3:24

Data Quality Management Techniques - The Complete Guide · 28:07 · Oct 2024 · 1,186 views

ClaimData quality problems can create legal liability, especially when biased or inaccurate data causes harm or violates regulations.4:35

How Data Capture Transforms ML ObservabilityPushkar Gar, Clari · 23:49 · Oct 2024 · 104 views

ClaimPushkar says the serving endpoint is the last place in the pipeline where data issues that were missed earlier can still be identified.8:02

How Feature Stores WorkSimba Khadder, Featureform · 30:33 · Oct 2024 · 95 views · DE4AI 2024

Pushed backSimba Khadder argues that data scientists should not have to become expert data engineers to build production-grade feature pipelines.8:27

How to Make Your Data Science Reproducible (and Why You Should Care)Ciro Greco, Bauplan · 11:59 · Oct 2024 · 123 views

ClaimA broken pipeline can be caused by changes in the code, data, environment, or cloud architecture.2:18

LLMs in Financial Services: Personalized Portfolio Recommendation EnginesAkmal Chaudhri · 15:56 · Oct 2024 · 81 views · DE4AI 2024
Real-Time Event Processing for AI/ML with NumaflowSri Harsha Yayi, Intuit · 22:38 · Oct 2024 · 601 views · DE4AI 2024

ClaimA Numaflow pipeline consists of vertices connected by edges, with each vertex running as a container.8:54

Scaling Data Reliably: A Journey in Growing Through Data Pain PointsMiriah Peterson · 16:04 · Oct 2024 · 138 views · DE4AI 2024

ClaimData reliability engineering treats data quality as an engineering problem and applies engineering practices to improve data and the systems built on it.5:07

The Daft distributed Python data engine: multimodal data curation at any scaleJay Chia, Eventual · 26:34 · Oct 2024 · 384 views · DE4AI 2024

ClaimDaft provides an interface for selecting and filtering rows and then streaming the selected data directly into training pipelines.15:50

The Evolution of Lyft's Feature StoreDevon Mittow, Lyft · 13:13 · Oct 2024 · 109 views · DE4AI 2024
The Future of Software Architecture for GenAI: Real-Time Data Streaming · 12:19 · Oct 2024 · 182 views
The Only Constant is (Data) ChangeBenjamin Rogojan, Seattle Data Guy & Chad Sanderson, Gable & Christophe Blefari, NAO & Maggie Hays, Acryl Data · 40:50 · Oct 2024 · 59 views · DE4AI 2024

ClaimData teams need to control the transition from one-off data pipelines to production assets so that unmanaged dependencies do not cause later failures.20:47

Turn Data Chaos into AI Strategy with Programmatic AI Data DevelopmentElena Boiarskaia, Snorkel AI · 27:14 · Oct 2024 · 67 views · DE4AI 2024

ClaimData labeling is needed throughout a generative AI pipeline, including retrieval-augmented generation, fine-tuning, alignment, prompt engineering, and evaluation.5:10

Knowledge as a ServicePrashanth Chandrasekar, Stack Overflow · 26:42 · Dec 2024 · 217 views · Agents in Production 2024
Building an ML Platform from scratch · 1:46:08 · Dec 2024 · 1,055 views

Pushed backBen says feature transformations should ideally be centralized in SQL mesh or a similar modeling system, while Eric favors monolithic pipelines early and sees decoupling as a later maturity step.1:26:27