# The Myth of AI Breakthroughs

Jonathan Frankle, Databricks | MLOps Podcast | Episode 205 | 1:10:03
Hosted by Denny Lee

Source: https://www.youtube.com/watch?v=pPntxWpfNEA
Channel: MLOps Community, now AAIF Live (https://www.youtube.com/@AAIFLive-x1r). Summarised by MLOps Talks.
Page: https://mlopstalks.com/talks/the-myth-of-ai-breakthroughs
Published: 2024-01-19
Tags: debugging, evals, orchestration

## TL;DR
- AI research needs empirical testing at the scale and cost that real users care about, since many highly promoted ideas fail to hold up in those settings.
- Policy decisions about systems such as facial recognition should focus on whether automated decision-making belongs in the process, rather than relying on accuracy or a human reviewer as an easy answer.
- The field should wait for ideas to survive replication and consolidation before treating them as breakthroughs, while investing in practical improvements to training efficiency, evaluation, and debugging.

## Summary
Jonathan Frankle argues that AI progress is usually slower and less tidy than online hype suggests. His research team tests papers and training methods against the needs of customers who may spend millions of dollars on a model, and he says reproduction often takes at least a month while serious decisions can take six months. He applies the same standard to his own work, including the Lottery Ticket Hypothesis, whose early results were less compelling at larger scale. The conversation also covers his work on police facial recognition, where he says accuracy and human review do not settle the policy question. His role is to provide technical information to policy professionals, not to replace their judgment. Frankle recommends waiting two or three months before judging new papers, trusting researchers close to the experiments, and treating incremental improvements as more useful than constantly searching for the next revolution.

## Key ideas
### Facial recognition policy is about delegated decisions, not just model accuracy
[10:02](https://www.youtube.com/watch?v=pPntxWpfNEA&t=602s)
Jonathan Frankle says poor accuracy is not the right basis for deciding whether AI should be used, because facial recognition can become highly accurate with enough data and good data. The harder question is whether society should delegate authority to an automated system or use its output inside another decision process. A hiring model that filters resumes delegates authority, while a model that scores resumes still influences a human decision. Frankle also rejects human review as an automatic solution. People have documented weaknesses in face recognition, including the other-race effect, and training does not remove every human failure. The policy issue is whether automated decisions should be part of the process and whether they make consequential actions too frictionless.

### Technical people should support policy professionals instead of trying to redesign policy
[21:00](https://www.youtube.com/watch?v=pPntxWpfNEA&t=1260s)
Frankle describes his work with Georgetown Law's Center for Privacy and Technology as technical service to people trained to make policy decisions. He helped investigate police use of facial recognition and contributed technical information to a report based on Freedom of Information Act requests to more than a hundred state and local police departments. He says technical workers are not equipped to decide how society should be run. Their role is to explain how systems work, identify risks, and answer questions when policy professionals ask. He objects to the instinct in technology circles to disrupt or reinvent policy. The report also tried to help small police departments by providing procurement questions alongside information about risks and safeguards.

### The Lottery Ticket Hypothesis began as a question about wasted neural-network capacity
[26:25](https://www.youtube.com/watch?v=pPntxWpfNEA&t=1585s)
Frankle explains that his dissertation asked how large a neural network needs to be and whether every part of a network is needed for learning. The Lottery Ticket Hypothesis examined whether networks could contain smaller, trainable subnetworks and whether training could be made more efficient. The early work was done with limited hardware, including a laptop, K80 GPUs, and borrowed resources. Frankle says the topic was influential because it encouraged people to ask whether neural networks were being trained intelligently and whether costs could be reduced. He is candid that the initial results worked in the settings he tested and were less compelling at larger scales, which led him to redo the work and rewrite the paper.

### MosaicML built research findings into a managed training product
[29:43](https://www.youtube.com/watch?v=pPntxWpfNEA&t=1783s)
Frankle says MosaicML's final product became a platform that handles many parts of model training, including GPU failures, loss spikes, optimization problems, and hyperparameter choices. His research team tests its own tools by training models before customers use them. The team then turns validated decisions into recipes that customers can run without reproducing all the research themselves. Frankle describes maintaining large amounts of evidence about choices such as positional encodings, token-to-parameter ratios, optimization strategies, and model architecture. He says MosaicML's approach is based on making training cheaper and more reliable for people who do not have abundant GPU access. The product includes the team's tested practices rather than direct consulting.

### Infrastructure choices depend on budget and workload
[36:32](https://www.youtube.com/watch?v=pPntxWpfNEA&t=2192s)
Frankle has used Slurm extensively, including in academic settings, but says it was not designed for machine-learning workloads. He points to issues such as gang scheduling, higher-level primitives, and the amount of supporting work users must handle themselves. MosaicML built its own orchestration stack to include secret management, logging to tools such as MLflow and Weights & Biases, and automatic hard-failure detection. He does not call Slurm a bad tool. When budgets are tight, it can be the right choice, just as TensorBoard or open-source Spark can be sensible options. Paid or managed systems can improve productivity when an organization can afford them.

### The field should treat Qstar-style claims as hypotheses requiring evidence
[44:20](https://www.youtube.com/watch?v=pPntxWpfNEA&t=2660s)
Frankle jokes that Qstar will revolutionize AGI, bring about the singularity within months, and eliminate money, then explicitly marks the claim as sarcasm. He uses the episode to criticize the habit of treating a tweet, paper, or rumor as the next event that will overturn the field. His lab tests prominent ideas and, in the course of that work, often finds that they do not hold up in the settings customers care about. He avoids public takedowns because researchers may have done good work at the scale available to them, and public criticism can damage junior scientists' careers. Instead, he shares results privately or describes what worked and failed in constructive technical writing.

### Waiting and listening to people close to the experiments are ways to reduce hype
[55:50](https://www.youtube.com/watch?v=pPntxWpfNEA&t=3350s)
Frankle waits roughly two or three months before studying a new paper in depth, unless it still appears relevant after the initial excitement fades. He says early ideas are often followed by papers that test, refine, and consolidate them. He also relies on researchers on his team who are doing the experiments every day. As a manager, he believes he is less informed about practical details than the person building a system. He prefers hiring people directly from PhD programs over relying on seniority or fame, because hands-on researchers have read the papers and run the experiments. When a team member has strong evidence for a direction, Frankle pushes on the reasoning and then usually trusts their judgment.

### Model research needs time, especially when mistakes can cost millions
[1:00:30](https://www.youtube.com/watch?v=pPntxWpfNEA&t=3630s)
Frankle says a careful paper reproduction takes at least a month, while ideas worth serious investment may require six months. Questions such as choosing between RoPE and ALiBi, assessing LoRA, or changing a model architecture can affect many future training runs. The cost of skipping that work becomes clear when a team trains a model costing millions of dollars and discovers that the setup was poor. His team keeps fact sheets for technical decisions so it can explain why each choice was made. He is more interested in incremental efficiency, evaluation for long-context models, synthetic data, and neural-network debuggability than in predicting the next revolutionary architecture.

## Notable quotes
- Jonathan Frankle: "We need to be empirical." (28:06)
- Jonathan Frankle: "Being useful doesn't mean that we tell people what to do." (20:36)
- Jonathan Frankle: "I don't tend to make big changes to model architectures until they've been out and popular for a few months." (56:18)
- Jonathan Frankle: "It's months, it's always months." (1:00:30)
- Jonathan Frankle: "Do science and only say what you can back up with data." (1:09:08)

## Tools & references mentioned
- Databricks
- MosaicML
- Denny Lee
- Demetrios Brinkmann
- Georgetown Law
- Center for Privacy and Technology
- Claire Garvey
- Alvaro Bedoya
- Joy Buolamwini
- Anil Jain
- Michael Carbin
- Naina Raikar
- IBM
- The Lottery Ticket Hypothesis
- Perpetual Lineup
- Qstar
- GPT-4
- ChatGPT
- BERT
- RoPE
- ALiBi
- LoRA
- Slurm
- MLflow
- Weights & Biases
- TensorBoard
- Spark
- NFNets
- ResNets
- Vision Transformers
- State Space Models
- Anthropic
- Hugging Face
- OpenAI
- Scale AI
- Julia Adibi

## Who should watch
- You are deciding whether a new paper, architecture, or training method deserves engineering time and want a slower test than social-media excitement.
- Your team trains large models with limited GPU access and needs to understand which infrastructure and research decisions are worth validating before spending heavily.
- You work at the boundary between machine learning and public policy, especially where automated systems influence policing, hiring, or other consequential decisions.

## Editor's note

Jonathan Frankle says Slurm was not designed for machine-learning workloads, leaving users to handle supporting work themselves. ZenML lets teams define workflows as Python pipelines and run the same code on different infrastructure through configuration. Each run records its steps, inputs, outputs, and code version, so teams can inspect why a training result was produced.

Written by the MLOps Talks editors (the ZenML team), not by the speaker.

## Related talks

- [Making AI Reliable is the Greatest Challenge of the 2020s](https://mlopstalks.com/talks/making-ai-reliable-is-the-greatest-challenge-of-the-2020s) (Alon Bochman, RagMetrics, 1:01:38)
- [Everything Hard About Building AI Agents Today](https://mlopstalks.com/talks/everything-hard-about-building-ai-agents-today) (Shreya Shankar & Willem Pienaar, Cleric, 47:03)
- [AI Agents: The Future of ML Engineering?](https://mlopstalks.com/talks/ai-agents-the-future-of-ml-engineering) (Matt Squire, Fuzzy Labs & Adam Becker, MLOps Community, 53:10)
- [A Blueprint for Scalable & Reliable Enterprise AI/ML Systems](https://mlopstalks.com/talks/a-blueprint-for-scalable-reliable-enterprise-ai-ml-systems) (Hira Dangol, Bank of America & Rama Akkiraju, NVIDIA & Nitin Aggarwal, Google & Steven Eliuk, IBM, 35:39)
- [Beyond the Matrix: AI and the Future of Human Creativity](https://mlopstalks.com/talks/beyond-the-matrix-ai-and-the-future-of-human-creativity) (Fausto Albers, AI Builders Club, 55:09)
