PodcastFoundational Models are the Future but...Data quality in 2023
21 sessions
PodcastML in Production: A DS from Ubisoft PerspectiveClaimJean-Michel Daignan says scalability testing can be more important for his pipelines than unit-testing every individual function.16:06
PodcastMultilingual Programming and a Project Structure to Enable ItClaimDevelopment pipelines can call scripts in sequence or in parallel, record timings and logs, and give production engineers a simple command to run the required workflow.43:03
MeetupDataframes Are All You Need: MLOps on Easy ModeClaimJay Chia says dataframes can support much of the MLOps workflow, including reading and writing data, exploring and processing it, feeding training pipelines, evaluating models, and running batch predictions.15:07
Building LLM Applications for ProductionClaimChip Huyen says LLMs do not reliably follow a required output schema, making it difficult for applications to parse their responses.5:19
PodcastEliminating Garbage In/Garbage Out for Analytics and MLClaimRoy Hasson says data quality checks that run only after data reaches a warehouse are reactive because downstream users and models may already depend on bad data.35:29
Lessons Learned Productionising LLMs for Stripe SupportClaimStripe addressed the problem by splitting it into question validation, topic identification, context-based answer generation, and tone adjustment.2:55
LIMA: Less is More for AlignmentClaimIncreasing the amount of Stack Exchange data did not improve generation quality when the data quality was held constant, because it did not add more tasks.5:30
Data Quality's Impact on Large Language ModelsClaimData quality is the foundation for analytics, data science, machine learning, and large language model initiatives because input data quality influences the output and return on investment.2:01
False Starts and Dead Ends: Building a Retrieval Augmented Generation SystemClaimFiles that are called PDFs may be invalid PDFs or poorly photocopied documents, so PDF handling is a data-quality problem.2:05
Model Blind Spot Discovery for Better ModelsClaimVectorFlow is an open-source vector embedding pipeline that extracts, chunks, embeds, and uploads raw data to vector databases.1:36
PodcastThe Role of Infrastructure in ML Leveraging Open SourceClaimNiels Bantilan built Pandera to catch data-frame type, formatting, range, and null-value errors and to let users encode data schemas in their code.5:01








