Using synthetic data for AI training is 'a big mistake,' says AI pioneer Rich Sutton
AI pioneer Rich Sutton argues that the industry's reliance on synthetic data for training large language models is a significant mistake. He advocates for experiential, real-world data as the only viable path for future AI development.
Why it matters
This debate challenges the current scaling strategies of major AI companies and highlights a fundamental disagreement on the future of machine learning.
Richard Sutton said synthetic data is not the solution for scaling AI. Business Wire/AP Rich Sutton criticized tech's reliance on synthetic data, urging real-world experiential learning. Tech giants like Google and OpenAI are going out of their way for real-world data. Sutton's Oak Lab focuses on teaching AI agents from experiences, not made-up datasets. Rich Sutton helped pioneer the technology behind today's AI boom. Now he thinks Big Tech's push to keep AI going is all wrong. On an episode of Sequoia's podcast released Tuesday, the Canadian computer scientist and Turing Award winner said that the AI industry is headed in the wrong direction because of its reliance on synthetic training data. "That's just a big mistake," he said when asked about using synthetic data as a way to keep scaling large language models. "Maybe it's the next big lesson."
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in