PodcastLook At Your ****ing Data 馃憖Pushed backKenny argued that the lack of discussion about LLM data quality is not explained only by data sources being secret; researchers also rarely discuss how they clean and filter data.4:13
9 sessions
PodcastLook At Your ****ing Data 馃憖Pushed backKenny argued that the lack of discussion about LLM data quality is not explained only by data sources being secret; researchers also rarely discuss how they clean and filter data.4:13
PodcastStreaming Ecosystem Complexities and Cost ManagementClaimA typical streaming pipeline connects Kafka to a processor such as Spark or Flink, then to storage and a serving layer, with each part requiring different skills.5:11
PodcastHow Sama is Improving ML Models to Make AVs SaferClaimHuman work is usually the most expensive part of an AI data pipeline, except in specialized cases where data collection involves expensive equipment.4:35
PodcastAI Data Engineers: Data Engineering After AIClaimAn AI data engineer is an AI agent connected to a company's stack that can perform data engineering tasks such as building pipelines and schema migrations.2:50
PodcastGraphBI: Expanding Analytics to All Data Through the Combination of GenAI, Graph, & Visual AnalyticsPushed backWeidong Yang disputes treating ontology as universally valid and says its truth must be limited to a defined domain.21:00
ML Engineers Who Ignore LLMs Are Voluntarily Retiring EarlyClaimYoni Michael says interactive AI is where many companies start, but scaling non-deterministic models into reliable production products requires deterministic pipelines and new engineering practices.16:27
PodcastReal-time Feature Generation at LyftClaimLyft's original cron-based pipeline introduced latency because each step waited for the next scheduled job.9:56
Building Advanced Agents Over Complex DataClaimImproving data quality is a major focus for raising response quality in RAG systems.3:22