Why Brains Beat Transformers (For Now)
The human brain runs on roughly 20 watts of power. It weighs about three pounds. It can learn a new concept from a handful of examples, generalize effortlessly across domains, and operate for decades without a firmware update.
Today's best LLMs require megawatts of compute, training on trillions of tokens, and still can't reliably do things a teenager finds trivial—like understanding when they're being sarcastic, or figuring out that a chair tipped on its side is still a chair.
This isn't a complaint about AI. It's an observation about how much we still don't understand.
The Efficiency Gap
A single A100 GPU draws about 400 watts. Training GPT-4 is estimated to have consumed somewhere north of 50 GWh. That's roughly the annual electricity consumption of 5,000 US homes — all to produce what amounts to a very sophisticated autocomplete system with a shaky grasp of arithmetic.
The neocortex, by comparison, performs somewhere around 1015 synaptic operations per second on 20W. That's a 50,000x efficiency advantage. Even accounting for the fact that brains and GPUs do different kinds of computation, the gap is staggering.
What the Brain Does Differently
Three things stand out:
- Architecture is learned, not fixed. Your brain rewires itself constantly. Synaptic pruning, neurogenesis, and Hebbian plasticity mean the physical structure of your brain is a continuous function of your experience. Transformers have one forward pass shape that's fixed at training time.
- Feedback loops are everywhere. The cortex has massive recurrent connectivity. Information doesn't just flow forward — it loops back, corrects itself, and integrates context at every level. Transformers have attention, but it's a pale imitation of recurrent cortical processing.
- Embodiment matters. Brains evolved to control bodies moving through physical space. Our concepts of object permanence, causality, and even mathematics are grounded in sensorimotor experience. LLMs have no bodies and no grounding — just statistics over text.
What This Means for AGI
I don't think scale alone will bridge this gap. The returns to scale are real — we've seen that — but they're also diminishing at the frontier. The next leap probably doesn't come from a bigger model. It comes from understanding what the brain figured out and translating those principles into engineered systems.
We're not going to build AGI by throwing compute at transformers. We're going to build it by understanding intelligence itself — and the brain is still the best map we have.
Some of the most exciting work right now is in predictive coding networks, biologically plausible learning rules, and hybrid systems that combine symbolic reasoning with neural nets. None of it is production-ready yet. But the direction is right.