PodcastLanguage, Graphs, and AI in IndustryClaimPaco Nathan says enterprise teams generally want to keep confidential data inside their own cloud perimeter or on-premises systems and prefer open source that their security teams can inspect.28:37
25 sessions
PodcastLanguage, Graphs, and AI in IndustryClaimPaco Nathan says enterprise teams generally want to keep confidential data inside their own cloud perimeter or on-premises systems and prefer open source that their security teams can inspect.28:37
PodcastSmall Data, Big Impact: The Story Behind DuckDBClaimMotherDuck is a separate company that provides a managed service, while DuckDB is the open-source project.31:38
Model Merging and Mixtures of ExpertsClaimMost of the 7B models on the Open LLM leaderboard were model merges at the time of the talk.0:47
From Research to Production: Fine-Tuning & Aligning LLMsClaimPhilipp Schmid says open source is necessary to use generative AI responsibly and reduce harm.2:08
Vision Pipelines in Production: Serving & OptimisationsPushed backBasic generation with more configuration was considered insufficient for the required consistency and control, so fine-tuning was chosen instead.3:31
Productionizing Health Insurance Appeal GenerationClaimCalifornia's independent medical review board data is publicly downloadable and can be used to generate synthetic data for fine-tuning.6:42
Fine Tuning LlamasPushed backKai Davenport rejects the idea that fine-tuning is always better than retrieval-augmented generation and argues that the two approaches can be combined.4:57
Enabling Efficient Trillion Parameter Scale Training for Deep Learning ModelsClaimDeepSpeed has been used to train open-source large models ranging from 5 billion parameters to half a trillion parameters.19:20
PodcastBeyond AGI, Can AI Help Save the Planet?ClaimAI2 open-sources its environmental AI data, training processes, weights, and models so people can inspect, criticize, improve, and find bugs.29:44
Data Labeling Best PracticesClaimTextMine built its own data labeling team and fine-tuned its own models, and this talk presents its lessons rather than prescribing one correct labeling method.1:07
Explaining ChatGPT to Anyone in 10 MinutesClaimSupervised fine-tuning uses examples of desired outputs, while RLHF uses human rankings of alternative outputs.7:51
PodcastFedML Nexus AI: Your Generative AI Platform at ScaleClaimSalman Avestimehr says FedML can provide fully on-premise deployment inside an enterprise's VPN so proprietary data stays in a trusted environment.8:27
PodcastOpen Standards Make MLOps Easier and Silos HarderClaimIbis separates a data frame API from the execution engine by compiling data frame code into backend-native code.7:17
Ghostwriter - AI Writing That Learns From YouPushed backThe speaker disputes the assumption that autocomplete can be made reliable through system-prompt instructions alone and says fine-tuning was needed.10:08
No GPU Before PMFClaimFine-tuning large language models is not yet well understood scientifically, including because researchers still do not know how to perform online training.4:27
PodcastAlignment is RealPushed backDemetrios Brinkmann questioned whether DSPy was too much of a research project to trust in production, while Shiva Bhattacharjee defended using a modified, self-hosted version of it.4:57
Building a Data Infrastructure for AI/MLClaimOrganizations should favor open-source or cloud-agnostic components and open data and table formats to avoid lock-in.10:10
Chronon: Airbnb's Open-Source Data PlatformClaimChronon was open sourced after being battle tested at Airbnb and Stripe.3:07
The Daft distributed Python data engine: multimodal data curation at any scaleClaimDaft can resize images, upload them as JPEGs, remove heavyweight image columns and write the remaining data to Parquet.14:06
Unified Data + AI Governance with Unity CatalogClaimThe open lakehouse approach is intended to let customers own their data and assets while using an open catalog to provide a unified governance view.14:05
Why DuckDB is the Future of DataClaimHannes Mühleisen says DuckDB is a relational analytical data management system that is fast, free, and open source.2:52
PodcastWe Can All Be AI Engineers and We Can Do It with Open Source ModelsPushed backAI-spec tests are not intended to replace general-purpose evaluation tools; they focus on the knowledge and API-calling features defined by the spec.37:00
PodcastHow to Optimize Large AI Models with PyTorchClaimPyTorch's open-source and community-based model lets developers build on existing work instead of first creating an entire framework.10:07
PodcastUnleashing Unconstrained News Knowledge Graphs to Combat MisinformationPushed backRobert Caulk disputes the idea that the system can freely scrape any online source, saying it only uses sources where crawling is allowed or licensed.46:51