The revolution of Federated LearningClaimSuccessful adoption of federated learning requires high-quality data, talented data scientists, and the ability to adapt to new data sets and quickly validate and release models.15:46
23 sessions
The revolution of Federated LearningClaimSuccessful adoption of federated learning requires high-quality data, talented data scientists, and the ability to adapt to new data sets and quickly validate and release models.15:46
PodcastMachine Learning Feature Store Panel DiscussionClaimMatias Dominguez says a small company without a market-validated product may not need to buy or build a full feature store.9:29
Meetup'Git for Data' - Who, What, How and Why?ClaimThe Git-for-data landscape includes tools for data versioning, data catalogs, data-pipeline versioning, and version-control databases.9:07
MeetupMLOps Community 1 Year Anniversary!ClaimMore research is needed in MLOps to explore new tools, processes, and methods for validating machine learning pipelines.46:06
MeetupDeploying Machine Learning Models at Scale in CloudClaimVishnu Prathish brings software engineering practices to MLOps because his deployment patterns and pipelines combine data science workflows with traditional software engineering.0:29
PodcastLuigi in Production Part 2ClaimLuigi Patruno's team moved several models from development notebooks into production processes with data validation, drift monitoring, and ongoing monitoring.8:02
PodcastScaling AI in ProductionClaimSrivatsan Srinivasan says machine-learning algorithms make up only a small part of the total machine-learning work, which also includes data collection, deployment, monitoring, pipelines, feature engineering, and feature stores.1:45
How Pinterest Powers Image SimilarityClaimPinterest wanted to make the transition from its mature batch pipeline to near-real-time processing as seamless as possible for consumers.9:56
MeetupWhat MLOps Has Taught MeClaimEwan Nicolson says data validation tools such as Great Expectations make him more confident that unusual data or results are not entering the system.17:07
PodcastMLOps InsightsClaimTesting in machine learning involves more than software unit tests; it can include data quality tests, training checks, serving infrastructure tests, and end-to-end tests.1:12
PodcastData Selection for Data-Centric AI: Data Quality Over QuantityClaimCody Coleman recommends starting with a small, clean data source and a simple model, then adding other sources incrementally when working with poor or undocumented organizational data.1:05:24
PodcastML TestsPushed backSvet Penkov takes the view that measuring data quality beyond basic validity is not always meaningful and that model performance in the intended domain is the more useful measure.29:34
MeetupData-Centric AI Means Centralizing Training DataClaimAlberto Rizzoli argues that training-data quality and labeling should be treated as an iterative process alongside model development.11:37
MeetupDurable Data Discovery: Making Exploratory Analysis StickClaimJames Campbell says Great Expectations focuses on helping people understand and communicate about data without making data quality a black box.54:39
PodcastThe Future of AI and ML in Process AutomationClaimSlater Victoroff says changing an OCR engine can invalidate labels when labels are stored only as positions in extracted text.23:25
Reading groupImpact of SWE in ML ProjectsClaimLaszlo Sragner said that schema tests should stop a pipeline when an upstream schema changes, while distribution changes and model-performance changes require analysis rather than being fully automated as software tests.33:59
MeetupModel Monitoring: The Million Dollar ProblemClaimFunctional monitoring checks data quality, data drift, model performance, and prediction behavior.2:15
PodcastMachine Learning at Reasonable ScaleClaimMachine-learning teams should work backward from their goals and constraints, avoid maintaining infrastructure when a service can do it better, and focus their time on data quality and model iteration.12:54
Reading groupThe ML Test ScoreClaimFor a pipeline jungle, Skylar Payne recommends first establishing confidence in the data, creating a simple baseline, and gradually refactoring toward the full system.26:27
MeetupML Drift: How to Identify Issues Before They Become ProblemsClaimData integrity problems such as swapped fields, incorrect units, missing values, and schema mismatches can look like model drift or cause real performance problems.18:33
MeetupBuilding 12-Factor Data Apps with KedroPushed backIvan Danov says Kedro is not another orchestrator like Airflow or Kubeflow because it focuses on pipeline authoring rather than workflow execution and monitoring.39:06