Brandt Welker

← Essays

AI

Moravec Inverted: 500 Million Years of Pretraining

The abilities that feel effortless- seeing, walking, knowing who's in the room- feel easy because they're the most optimized thing you own. Evolution is pretraining, and the machine inherited the wrong five thousand years.

·3 min read

In 1988, a Carnegie Mellon roboticist named Hans Moravec wrote down something that still sounds backward. He said it’s comparatively easy to make a computer perform like an adult on an intelligence test or at checkers, and difficult or impossible to give it the skills of a one-year-old when it comes to perception and mobility. That’s Moravec’s paradox, and almost forty years of AI have proven him right.

The stuff we call hard (chess, calculus, proving a theorem) turned out to be cheap for a machine. The stuff we take for granted as easy, seeing a face, catching a ball nobody warned you about, is brutal. A four-year-old does the second without thinking and can’t do the first at all. We’re still fighting to make a robot pick up a cup it has no experience with.

The effort you feel doing something is inversely proportional to how long the system was optimized to do it. Seeing feels like nothing because it’s the most optimized thing you own. Math feels like work because it’s the newest layer, bolted on last.

Evolution is pretraining. The genome is the checkpoint. Your lifetime is fine-tuning on top of it. Vision, balance, object permanence, knowing who’s in the room and where their hands are, all of it got something like five hundred million years of optimization, run on the most expensive cluster ever built, the one where models that ran badly died. Symbol manipulation got maybe five thousand. Elizabeth Spelke’s work at Harvard shows infants arrive already holding the physics. They know objects are solid and don’t wink out when a hand covers them, before they have a word for any of it. It shipped in the checkpoint.

Which is why AI is where it is. We trained it on the one corpus we bothered to write down, five thousand years of symbols, and it went superhuman at the part that’s newest for us. It never got the five hundred million years. Nobody wrote the motor cortex down. So it’s brilliant at the layer we find hard and clumsy at the layer we find trivial, and people call that a paradox. It’s a training log.

The tempting move is to skip the training and copy the biology straight out. Read the weights out of a brain. You can’t. We’ve had the full wiring diagram of C. elegans, three hundred and two neurons, for decades, and still can’t run it forward and get the worm back. The connectome isn’t the program. Weights don’t port. The learning rule does: the procedure that grew them. Steal the rule, not the checkpoint.

And that buries the comforting story that an AI is a baby learning from scratch, a mind early on the same climb we made. It’s the opposite of a baby. A baby is a full physical prior on an empty symbolic slate, everything about bodies and space baked in, not one spoken word yet. The LLM is the negative of that, an empty physical prior on a full symbolic slate, fluent in everything we ever wrote and unable to feel the floor. Not the same system at two points on one line. Mirror images. So when the machine fumbles what your toddler nails, don’t read it as stupid; read the training log. The ease you feel is old optimization you never paid for.