Toddlers are Still Smarter Than ChatGPT, and Scientists Want to Know Why

For at least 100,000 years, there has been exactly one thing on Earth capable of learning a human language to full fluency: a human child.

That changed four years ago when large language models began producing text so natural it could pass for human. But look behind the curtain, and something striking becomes apparent.

The AI gets there by consuming an almost incomprehensible mountain of data. The toddler does it on a shoestring.

That gap — and what it might reveal about both artificial intelligence and the developing human mind — is one of the most fascinating open questions in science right now.

The Scale of the Problem

A modern large language model might consume 15 trillion word-like tokens during training. Meta’s Llama 3.1 did exactly that. Frontier models could be training on ten times more. Meanwhile, a child raised in a language-rich home hears around 100 million words before their teens. Add reading and you get to perhaps 300 million by age 20.

To make the difference tangible: if you printed out the training data for a modern LLM, the stack of paper would reach past the International Space Station. A child’s equivalent exposure would stack up about 20 metres. And children are working with far less than even that when they take their first linguistic steps.

“We still have to burn down a forest and scrape the entire sum of all human knowledge,” says cognitive scientist Michael Frank at Stanford, “to re-create this milestone that happens in our living rooms over the course of a year.”

Researchers call this the data efficiency gap. And closing it — or even just understanding it — could reshape both AI and our understanding of childhood development.

Why This Gap Matters

For a decade, the dominant strategy for improving language models has been simple: make them bigger and feed them more data. But there are limits to this approach. The freely available text on the internet may be largely exhausted for training purposes within a decade. And scaling is expensive — computationally, financially, and environmentally.

Children demonstrate that something far more efficient is theoretically possible. A competition called BabyLM, now in its fourth year, challenges researchers to train language models on child-scale datasets of around 100 million words. Results have been genuinely surprising: the 2024 winner, a model called GPT-BERT, outperformed Meta’s Llama 2 — trained on roughly 15,000 times as much data — on certain language benchmarks.

But BabyLM models are nowhere near the fluency of commercial AI. They often can’t generate text at all. Something is still missing — and identifying what that something is has become one of the central puzzles of cognitive science and AI research.

The Embodiment Problem

One obvious difference between a child and a language model is that a child is not just processing text. They are moving through a physical world, watching faces, hearing tone of voice, handling objects, and making social inferences. Language arrives embedded in a rich sensory and emotional context that a model reading from a text file simply doesn’t have.

Researchers have tried to address this by training models on video footage of children’s actual lives. The SAYCam project wired three babies with headcams, recording two hours a week of their first two and a half years. Models trained on this footage were able to learn simple words — ball, cat — without any of the built-in biases many theories assumed were necessary. It was a meaningful result, but nowhere close to a two-year-old.

A more ambitious project at Princeton spent five years recording the first 1,000 days of 17 children’s lives — 12 hours a day, wired homes, cameras in every room except bedrooms and bathrooms. The scale of that dataset would have been impossible to work with until recently. “For the first time, we have the input,” says neuroscientist Uri Hasson. “It’s really only the beginning.”

The Missing Ingredient: Active Curiosity

Even vast amounts of video may not be sufficient, because children aren’t passive observers. They are relentlessly active experimenters. They reach, shake, drop, babble, and demand attention — constantly testing cause and effect, maximising their ability to make a predictable impact on the world. Research by developmental psychologist Alison Gopnik shows that what looks like play is in fact a highly effective learning strategy.

Children also have a social intelligence that models entirely lack. They don’t just process information — they reason about the person giving it to them, inferring why an adult is saying something particular, what the adult knows, and what they might be trying to teach. This layered social cognition gives children an efficiency advantage that passive statistical learning cannot replicate. When models have been built to learn through interaction with other models — a rough approximation of social learning — they haven’t outperformed standard ones.

The uncomfortable conclusion is that the gap between children and AI isn’t simply a matter of more data or better architecture. It may require rethinking what learning actually is.

A Tool for Understanding Ourselves

Beyond the engineering challenge, there’s a deeper reason to care about this gap. Children have been the only language-learning entities in existence for the entirety of human history. Now there’s something else — imperfect, alien, but genuinely linguistic. That creates, for the first time, the possibility of comparative study.

When an LLM learns syntax from statistics alone — something linguists confidently declared impossible for decades — it forces a reckoning with assumptions that have shaped the field since the 1950s. When a model fails at something a two-year-old handles effortlessly, it points to something real and important about the human mind.

“For the last 100,000 years, humans have been the only entities in the universe that use language,” says Alex Warstadt, one of BabyLM’s founders. “Now there’s this other linguistic entity. Finally we have a model organism.”

The child still wins — easily, and on a fraction of the resources. Understanding exactly why may be the most consequential question at the intersection of AI and human development, with implications for everything from how we build the next generation of AI models to what it means to be a thinking, language-using creature in the first place.