State of AI Report 2024

Nathan Benaich, Air Street Capital27:00 · Dec 2024 · 633 views
Thumbnail for State of AI Report 2024 Watch on YouTube
TL;DR
  1. 1

    Frontier model performance is becoming more competitive as Anthropic, Google, xAI, Meta, and Chinese companies release capable systems alongside OpenAI.

  2. 2

    AI adoption is growing quickly, with lower inference costs, higher enterprise retention, and AI companies reaching significant revenue faster than comparable SaaS companies.

  3. 3

    Safety concerns have shifted from fears of autonomous takeover toward immediate problems such as deepfakes, jailbreaking, copyright disputes, data use, and the energy demands of compute infrastructure.

Summary

Nathan Benaich presents an editor's cut of the 2024 State of AI Report, covering research, industry, politics, safety, and predictions. He describes OpenAI's early frontier lead, the rise of inference scaling and smaller distilled models, Meta's Llama ecosystem, Chinese model progress, and AI applications in biology and robotics. Nvidia remains dominant in research compute, while model prices have fallen sharply. Enterprise use is expanding, and AI-first companies are reaching revenue milestones faster than comparable SaaS companies. Benaich also discusses copyright lawsuits, self-driving, AI hardware, nuclear power deals driven by compute demand, and the changing safety debate. He expects more AI for science, AI-generated software, and greater government scrutiny as frontier labs seek state-backed funding. In questions, he argues that model economics will depend on the value of each task and recommends viewing AI as a coach or thought partner for some uses.

Key ideas
01:41

Frontier model leadership is spreading beyond OpenAI

Benaich says OpenAI was clearly ahead in producing highly capable models during the previous year, but that gap is shrinking as Anthropic, Google, xAI, and others release frontier systems. OpenAI's o1 model focused on complex reasoning by pausing, understanding a problem, and working through it step by step. Benaich connects this to inference scaling and says the model performed well on mathematics, science, and coding problems, including tasks that resembled the strategies of PhD students. He presents this as a response to the criticism that language models only reproduce learned statistics.

03:04

Open models and distillation are changing how capable systems are built

Meta's Llama 3 releases helped drive many forks and derivatives on Hugging Face. Benaich says the field is moving toward training very large models first, then using them to create training data and improve smaller models. Gemini Flash is one example, using a larger model to help train a smaller one through distillation. He also points to Chinese models from DeepSeek, Alibaba, and Qwen, particularly in vision-language and coding, as increasingly capable and remarkably open source. The result is more competition with large US vendors and more models that can be adapted for specific uses.

05:34

Large models are moving into biology and robotics

Benaich describes how models are spreading beyond text and images into scientific domains. Profluent trained models on protein and amino acid sequences to learn how sequences produce functional proteins. The company then used those models to create new genome editors with sequences unlike naturally occurring editors, while showing that they functioned in human cells. Robotics has also returned to investor attention after moving in and out of fashion. Benaich says foundation models for robotics and general-purpose embodied AI are among the industries investors would consider especially exciting.

06:56

Nvidia's compute position remains unusually strong

The report tracks the size of compute clusters built by governments and private and public cloud providers. Benaich says systems are now about an order of magnitude larger than those built roughly a year earlier, and cites xAI's 100,000-GPU cluster as having been brought online in about 122 days. In AI research papers, Nvidia chips remain far ahead of other hardware. He says about 35,000 papers used an Nvidia chip during the year, around 11 times the combined total for the other major groups in his comparison. The A100 remains the most popular Nvidia chip, and chips can remain useful for research for almost a decade.

10:18

AI products are gaining adoption while their economics are still unsettled

Benaich says generative AI companies have attracted very large funding rounds, although monetization and margins remain uncertain. At the same time, inference and training are becoming more efficient, and the cost of systems with equivalent intelligence has fallen by one or two orders of magnitude across providers. Ramp data in the presentation shows retention for a basket of AI products rising from 41% to 63% over a year. Stripe's comparison of promising companies found that AI companies reached more than $30 million in annual revenue in a little over a year and a half, while comparable SaaS companies took almost five years.

14:29

Applications are becoming visible in speech, biology, video, and autonomous driving

Benaich says speech and speech recognition have crossed the uncanny valley, with multilingual and lifelike voices in different styles and personas. He discusses the combination of Recursion and Exscientia, which joined biology-based drug discovery with chemistry and created a business with a large GPU cluster in biopharma. Video generation is improving in long-form coherence and physical consistency. He also says Waymo provides a striking consumer experience in San Francisco, Los Angeles, and Phoenix, after years in which self-driving companies were often criticized for delays.

17:37

The political and safety debate has moved toward immediate operational problems

Benaich describes limited US frontier-model rules, state-level regulation, and European restrictions affecting products from companies such as Anthropic, Apple, Meta, and xAI. He also discusses disputes over training data and GDPR. In safety, he contrasts the 2023 focus on catastrophic AI risk with a 2024 environment in which companies are racing to raise money and acquire users. Red-teamers continue to defeat jailbreak safeguards, while the harms Benaich sees today are more ordinary, including deepfake fraud, copyright disputes, and misuse of personal data. He also says rising GPU demand is putting earlier net-zero commitments under pressure.

21:48

Benaich expects AI for science and state involvement to grow

For the coming year, Benaich predicts that an AI-generated research paper could be accepted at a major machine learning conference and that someone with limited coding ability could use generative tools to create a widely shared app or website. He also expects frontier labs seeking tens of billions of dollars from sovereign states to face serious national-security review. In the questions, he argues that model prices will depend on the value of the task, with some work handled by a model and other work by a model combined with a human. He says AI for science remains underdeveloped and points to research showing improved materials discovery, patents, and productivity when workers use AI tools.

"AI companies took just a little over a year and a half to get to over $30 million in annual revenue compared to their SaaS companies which took almost five years."14:28
Who should watch
  • You need a compact view of the major AI research, product, investment, policy, and safety developments covered in the 2024 State of AI Report.
  • You are assessing where model competition, compute spending, inference costs, or enterprise adoption may move next.
  • You work on AI for science, robotics, autonomous driving, or generative software and want Benaich's account of which applications are becoming concrete.