aillmopiniontools
◈ AI · Opinion

Two Karpathys

Split arcane-circuit sigil, a violet speech-bubble with a hype spark on the left and gold-and-cyan code braces on the right, the Two Karpathys mark

On February 12, 2026, Andrej Karpathy posted microgpt. One file, 200 lines of pure Python, no dependencies. It handles the dataset, the tokenizer, the autograd engine, the attention and MLP blocks, the Adam training loop, and inference, all in code short enough to read end to end in one sitting and small enough to run on a laptop. He called it the culmination of several earlier projects. Set that against the churn of the AI news cycle that same week, model launches, benchmark leaderboards, threads arguing about what some demo means for where this is all headed, and the small file said more about who Karpathy is than any of the noise around it.

Karpathy is, for practical purposes, two people, and only one of them is good for you. One sits down and writes minimal, legible, reproducible code that runs exactly as advertised. The other coins a phrase, watches it detonate across the industry, and rarely does much to contain the blast. Both are real. The first is one of the most valuable engineers and teachers working in AI. The second is a problem, and in a market this overheated, a bigger one than his fans want to admit.

The engineer

Start with what’s real, because it’s the part worth protecting. nanoGPT, from 2022, is still the repository people clone when they want to understand a GPT training loop without wading through a framework. llm.c goes the other direction entirely, GPT-2 and GPT-3 pretraining written in raw C and CUDA with no 245MB of PyTorch sitting between you and the hardware. In May 2024 it reproduced GPT-2 124M in 90 minutes for $20, a number Simon Willison flagged and wrote up at the time, and by July it had scaled to GPT-2 1.6B on a single 8xH100 node in 24 hours for $672. Those aren’t demo numbers. They’re a receipt you can reproduce.

nanochat, from October 2025, is the fullest version of the instinct. The tagline is exact. It calls nanochat “the best ChatGPT that $100 can buy,” and the repository delivers, a minimal, from-scratch, full-stack pipeline of roughly 8,000 lines, from a Rust tokenizer through FineWeb pretraining, SmolTalk midtraining, supervised fine-tuning, optional RL, and a web UI to serve the result. The $100 tier runs about four hours on one node. A $1,000 tier runs about 41.6 hours for something better. It is honest about exactly what each tier buys, and it never pretends to be more than a working clone you can build yourself. microgpt is that same instinct taken to its limit, a GPT small enough to hold in your head.

The through-line is a harness philosophy, minimal and legible scaffolding you can read start to finish, with no framework magic hiding what happens at each step. In a field that mostly ships behind a wall of abstraction, that is a rare and generous choice, and it is the part of Karpathy’s output other engineers build habits around instead of merely citing. This is the work that will still matter in ten years.

The hype-man

Andrej Karpathy speaking, 2019
Gladwin Analytics · CC BY 3.0

Now the other one. Karpathy has a gift for the phrase that travels, and he uses it constantly. In January 2023 he tweeted that “the hottest new programming language is English.” It became a talk title and a conference refrain, repeated long after the tweet scrolled away. It is directionally interesting, prompting really is an interface, but it is also an absolute with the caveats sanded off, and Gary Marcus was right to push back that deterministic languages are not going anywhere. Karpathy never qualified it. He rarely does.

Software 2.0, from 2017, is the serious version of the same move. Neural networks as software “compiled” from data, the programmer’s job shifting from writing logic to curating datasets. Prescient, genuinely. It also quietly oversells itself, because Software 1.0 never leaves the building. Every model in production still sits inside hand-written glue, pipelines, deployment, and monitoring. By June 2025 he had extended the taxonomy again at YC’s AI Startup School, casting LLMs as utilities, fabs, and operating systems, and coding agents as an Iron Man suit you pilot. Clean framework, useful vocabulary, and exactly the kind of grand narrative that becomes doctrine the moment ten thousand people repeat it without the footnotes.

Then there is “vibe coding.” On February 2, 2025 he described “a new kind of coding where you fully give in to the vibes, embrace exponentials, and forget that the code even exists.” It went viral, Collins named it the 2025 Word of the Year, and within months it had hardened in common use into permission to ship code nobody read or understood. He later called it a “shower of thoughts throwaway tweet,” and about the origin he’s right. But here is the part his defenders skip. When you are one of the most-followed engineers in the field, and your last several coinages have each turned into an industry, a throwaway tweet is not a private aside. It is a match dropped in a dry field, and by 2025 Karpathy had watched enough of his matches catch to know what happens next. “Forget that the code even exists” is a thrilling line and a terrible instruction, and he had every reason to know it would be read as an instruction.

It would be easier to forgive if he couldn’t tell the difference. He can. In an October 2025 conversation with Dwarkesh Patel he pushed hard against the industry’s agent hype: “there’s some over-prediction going on in the industry. In my mind, this is more accurately described as the decade of agents,” and, bluntly, “the industry is making too big of a jump and is trying to pretend like this is amazing, and it’s not.” So the caution is available to him. He spends it selectively, cooling the hype cycles he didn’t start and stoking the ones he did. A man who can say “it’s not amazing” about agents in one breath and “forget that the code even exists” about vibe coding in the next is hard to read as unaware of which lever he is pulling.

That is the real objection. Not that his ideas are wrong, most are half-right in the way good provocations are, but that he keeps lobbing them into an already-manic market and then shrugs at the fire. The fervor has a cost. It shows up as juniors who skipped the fundamentals because English was supposedly the only language left, as capital chasing whatever narrative he minted last quarter, as expectations no shipping product can meet. He didn’t cause the AI bubble. He is one of its most effective, and most quotable, accelerants.

Every bubble has a middle ground

We are in a bubble, and that is not the insult people take it for. Railway mania in the 1840s ruined thousands of investors and left Britain a national rail network it still runs on. The dot-com crash vaporized pets.com, and the fiber laid in the frenzy still carries traffic while Amazon came out the far side larger than anything the mania promised. Every real bubble is a true story told at a volume the story can’t support. The value is real. The froth is also real. Then the froth clears, and what’s left is the quiet, unglamorous middle, the rails and the fiber and the few who were building something the whole time.

AI has that middle, and Karpathy is standing on both sides of it at once. The durable part of this era will look like his code, small, legible, reproducible, honest about what it does and doesn’t do. The part that evaporates will look like his slogans, the frames that felt like prophecy and turn out to be marketing. The same person produced both, and the slogans are louder. When the correction comes, and it will, the people who took “English is the hottest programming language” as a career plan will be worse off than the ones who cloned nanoGPT and read every line.

None of this makes the engineer less good. It makes the hype-man harder to excuse. Karpathy’s real contribution is a shelf of programs that show their work and a way of teaching that assumes you can handle the truth. His sensational side is a bet that attention is free, and in an overheated market attention is the most expensive thing there is. The tweets age badly. The code still compiles. A few years from now, when the bubble has done what bubbles do, that second sentence is the only one still worth quoting.