Ad
Skip to content

KI-Pioneer Sutton calls synthetic data a "big mistake" in the face of an infinitely complex world

Image description
Nano Banana Pro prompted by THE DECODER

AI researcher Richard Sutton pushes back against a core strategy at the leading AI labs. Synthetic data won't solve the scaling problem for large language models, he says, and blames the sheer complexity of the world.

Turing Award winner Richard Sutton is considered one of the founders of reinforcement learning. He wrote the field's standard textbook, mentored researchers like David Silver, who later worked on AlphaGo, and penned the influential 2019 essay "The Bitter Lesson." In it, Sutton argues that over the long run, only the AI methods that scale with compute win out, things like search and learning, not the ones that rely on built-in human knowledge.

In a recent conversation, he introduced his new company, Oak Lab, which he founded with his former student Khurram Javeed. He and his cofounder also talked about the limits of current training methods.

Why LLMs are only half a win

For Sutton, large language models are both a good and a bad example of the Bitter Lesson. Good because they scaled enormously with compute and you could simply "drink in the internet." Bad because they hit a wall at exactly that point. The internet is finite, and the real world is "massively bigger than everything we stored on the internet." At that point, Sutton says, you lean too heavily on human knowledge, and that ultimately holds the systems back.

Asked whether synthetic data could break through this bottleneck, Sutton doesn't mince words. "No, that's that's just a big mistake." The reasoning comes from the "Big World Hypothesis" that Javeed formulated and the group in Alberta has been working on for years.

The world is too big for any simulation

The core assumption is that the world is infinitely complex and "massively more complex than your mind than any agents any agent." Any simulation of it is tiny, "microscopic." A small program can only ever produce a small world that doesn't match reality, with wrong friction values or an inaccurate model of a robot's motor behavior. Sutton also points out that the world contains many other agents whose inner workings simply can't be generated as synthetic data. "There's no way we can have synthetic data for other people's minds."

A second objection is the human bottleneck. Who decides which synthetic data is good or bad? By Javeed's argument, that takes human experts. "You need human experts who know what's a good data set and what's a bad data set for that approach to scale. So it is bottlenecked by humans." Say you wanted to train a drone that moves like a bat using echolocation. You'd first have to hire domain experts. That makes the approach limited by human expertise, and it doesn't scale. Even with self-driving cars trained in simulation, humans end up fixing the gap between simulation and reality.

Sutton's alternative is learning from your own experience

Sutton's fix is to take humans out of the loop and let agents learn from their own experience. An agent should learn its own world model and keep correcting it, instead of relying on a frozen simulation model that humans built. "Simulators they make themselves."

Sutton also criticizes the fact that today's language models stop learning after training. "Their weights never change." What's needed instead is real continual learning, which in Sutton's view is just learning, since "all learning is continual," without wiping out old knowledge, the so-called catastrophic forgetting.

Sutton thinks this problem can be solved, in part with a method called "Continual Backprop" that his team published in Nature. He calls language models an "amazing scientific breakthrough," but only "like 20% or a quarter of intelligence."

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Read on for the full picture.
Subscribe for hype-free coverage.

  • Full access to every article on THE DECODER
  • No ads
  • Join the comments and community discussions
  • A weekly AI news recap via mail
  • 6x/year: "AI Radar" — deep dives on the AI topics that matter most
  • Daily AI news, always up to date
  • Our full ten-year archive
  • Covered by a team with 10+ years in AI
Subscribe to The Decoder