Artificial intelligence has a seasonal disorder. Every few decades, the field discovers something real and decides the rest is mostly scheduling—until the world arrives, wearing muddy shoes.
I The Summer That Named Itself
In the summer of 1956, ten researchers gathered at Dartmouth College for what John McCarthy had described, in a funding proposal, as a two-month investigation into artificial intelligence. It was McCarthy's second attempt at a name — four years earlier, a collection he'd edited with Claude Shannon, Automata Studies, drew essays on formal automata theory, not the thinking machines he was after. The phrase was new. The confidence was not. The proposal stated that "every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it."
There are cautious grant proposals. This was not one of them.
The summer ended. Intelligence, rudely, continued to require more than a summer.
But the Dartmouth meeting did something enormous. It gave the field a name, a club, a set of ambitions, and a proof of concept with excellent timing. Logic Theorist, built by Newell, Simon, and Shaw, proved theorems from Principia Mathematica through symbolic search. It was not a toy in the dismissive sense. It solved real formal problems. It made the old philosophical question feel suddenly mechanical: if a machine can prove theorems, what exactly is missing?
Herbert Simon, never a man to bring a candle when a floodlight was available, told the Operations Research Society in 1957 that "there are now in the world machines that think, that learn, and that create."
The room, broadly speaking, was willing to be convinced.
That same year, General Problem Solver made the ambition explicit. The name sounds almost comic now, but it was not meant as marketing fluff. It was a research hypothesis: if a problem could be described precisely, a machine could search its way through the space of possible solutions.
The trouble was hidden in the word precisely.
II The Prophets Were Not Fools
Flush with Logic Theorist's success, Simon made two specific predictions in 1956. Within ten years, a computer would be world chess champion. Within ten years, a computer would discover and prove an important new mathematical theorem.
The theorem prediction was basically correct before the ink dried. Logic Theorist had already produced the result. The chess prediction missed by four decades, and the machine that finally defeated Garry Kasparov in 1997 looked nothing like the general reasoning systems Simon had imagined. Deep Blue was specialized hardware, brute-force search, and an astonishing amount of chess-specific engineering. It did not so much "think like a human" as calculate through a blizzard faster than any human could breathe.
That distinction matters. Simon was not a crank. He was one of the sharpest minds of the twentieth century — a man who would, two decades later, win the Nobel Prize in Economics for arguing that real decision-makers don't optimize, they satisfice, settling for good enough instead of computing their way to perfect. He understood bounded rationality better than almost anyone alive. He just couldn't see it in his own forecast.
McCarthy made the same bet on the same board, twelve years on — literally. In 1968 he wagered that a computer would beat the chess master David Levy within ten years. He lost, narrowly, just before the clock ran out.
From the inside, a breakthrough rarely looks like a local maximum. It looks like a door.
III The First Winter
The early programs were impressive because they did things computers were not supposed to do. They proved theorems, solved puzzles, played games, and manipulated symbols in ways that seemed, if you looked at them in the right way, like reasoning.
The angle was important.
These systems lived in clean rooms of the mind. The rules were known. The inputs were tidy. The goal was explicit. The environment sat politely still while the program reasoned about it. Real life, as usual, had not read the proposal.
Machine translation became the public face of the problem. In 1954, Georgetown and IBM demonstrated a system that translated sixty Russian sentences into English automatically and correctly. The press loved it. Researchers estimated that fully automatic translation was three to five years away.
Ten years later, the official verdict was brutal: too expensive, too little progress, not nearly useful enough. Russian, inconveniently, contained more than sixty sentences.
Terry Winograd's SHRDLU, arriving in 1971, made the same point more elegantly. Inside a simulated tabletop of blocks and pyramids, it held something like a conversation — following multi-step instructions, tracking pronouns across turns, explaining its own reasoning when asked why. For a moment, it looked like a machine that understood English. What made the demonstration work was also what sealed it shut: add one new kind of object to the tabletop, and the illusion needed as much hand-built machinery all over again. Winograd built the demonstration, then years later argued that the understanding on that screen had never really been there.
The pattern repeated. A robot could navigate a maze and then lose its dignity in an ordinary room. A medical program could reason from clean inputs and then meet a patient, which is to say a source of contradictory symptoms, missing information, and human imprecision. A theorem prover could operate beautifully in a formal system and then discover that most of the world is not written in formal notation.
Claude Shannon had already put a date on it. Closing out The Thinking Machine, a CBS documentary made with MIT, he told the country in 1960 that something close to the robot of science fiction fame would emerge from the laboratories within ten or fifteen years. Nobody had scheduled the winter in between.
By the early 1970s, the people with chequebooks had noticed. In Britain, the government asked mathematician James Lighthill to review the field. The Lighthill Report did not arrive carrying flowers. Lighthill's case rested on a single technical villain he called the combinatorial explosion: widen the "universe of discourse" a machine had to reason about, even modestly, and the computation required did not grow with it. It detonated. The BBC turned the verdict into theatre: Lighthill, elevated on a lectern above the stage, faced down three AI researchers — John McCarthy among them — in a televised 1973 debate that played out less like a panel discussion and more like a public sentencing. Funding was cut. In the United States, DARPA also pulled back.
Neural networks did not even get to wait for Lighthill. Marvin Minsky and Seymour Papert had already gone after Frank Rosenblatt's perceptron in a 1969 critique, proving with unforgiving rigor that a single layer of the things could not learn XOR — exclusive or, a distinction a child grasps without ever being told it has a name. Funding for neural networks dried up years before Lighthill made the freeze official — and the answer to Minsky and Papert kept arriving and going unheard for another seventeen years.
The first AI Winter did not arrive with a bell tolling. It arrived as fewer approvals, smaller grants, delayed renewals, and conversations in which "promising" began to mean "not this year."
Winters in AI are rarely announced. They are administered.
IV The Second Spring
The next boom did not begin with a philosophical breakthrough. It began with useful software.
Expert systems in the late 1970s and 1980s worked well enough to embarrass sceptics. MYCIN diagnosed bacterial infections at the level of senior physicians. XCON configured computer systems for Digital Equipment Corporation and saved the company tens of millions of dollars a year. These were not paper promises. They were deployed systems doing valuable work.
That success made the next mistake almost irresistible.
If some rules could help diagnose infections, and some more rules could configure computers, then perhaps intelligence was finally ready to be industrialized. More rules, more experts, more knowledge engineers, more specialized hardware. The problem was not conceptual anymore. It was logistical. Hire enough people to extract enough knowledge from enough experts, and the machines would become steadily more capable.
For a while, this was a perfectly reasonable thing to believe. This is the uncomfortable part of AI history: the optimistic story is often locally correct. The systems really do work. The money really does arrive. The charts really do go up.
Then, in 1982, Japan poured accelerant on the whole thing.
The Ministry of International Trade and Industry announced the Fifth Generation Computer Systems project, a ten-year national effort to build a new class of computers based on logic programming and parallel inference. The machines would reason. They would understand natural language. They would support a knowledge-based society.
The West reacted with the calm confidence for which geopolitical technology races are famous. Which is to say: not at all.
Edward Feigenbaum and Pamela McCorduck published The Fifth Generation, part technical argument, part siren. DARPA launched the Strategic Computing Initiative. Europe launched ESPRIT. Britain launched Alvey. A Japanese government program had turned AI into an international contest with budgets, committees, acronyms, and the faint smell of panic.
Edsger Dijkstra, watching from the sidelines, suggested that asking whether machines can think was about as relevant as asking whether submarines can swim. This was a beautiful line, and therefore had little effect on funding decisions.
The spring was on.
V The Second Winter
Expert systems had a problem that looked small until it filled the room: experts know things they cannot easily say.
A senior physician does not diagnose by walking down a complete decision tree. An engineer configuring a system does not consciously enumerate every condition that shaped their judgment. Expertise is full of shortcuts, exceptions, muscle memory, domain folklore, and tiny acts of recognition that vanish when someone asks, "Could you write that as an if-then rule?"
XCON eventually grew to about 10,000 rules and needed a dedicated team to keep it current as DEC's product line changed. This was still useful. It also revealed the trap. The system did not understand computer configuration. It was a very large, very valuable, very fragile instruction manual that had to be rewritten whenever somebody shipped a new product.
Every edge case demanded a new rule. Every new rule could conflict with another rule. Every update required experts, engineers, meetings, tests, and the quiet fear that fixing one corner had broken another. The systems grew more capable and less graceful at the same time.
Then the hardware story collapsed.
Much of the AI boom had been tied to specialized Lisp machines. They were elegant, expensive, and suddenly competing with cheap desktop computers that improved at a terrifying pace. When general-purpose hardware became good enough, the specialized machines lost their reason to exist. Companies stopped buying. Vendors folded. Procurement departments developed selective amnesia about last year's revolution.
Japan's Fifth Generation project ended after a decade and 54 billion yen. The internet was emerging, commodity microprocessors were winning, and the underlying bet had been overtaken by a different future.
The second AI Winter did not refute all of expert systems. MYCIN was still impressive. XCON still saved money. The problem was harsher than failure: success had been real, useful, and insufficient.
That is the line AI keeps tripping over.
VI The Third Spring, For Now
The quiet lasted the better part of two decades. Neural networks stayed unfashionable and the funding stayed thin, and the handful of people who kept training them anyway mostly did so without an audience.
In 2012, Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton entered AlexNet into the ImageNet competition. It cut the previous best error rate so dramatically that some researchers initially wondered whether the scoring system had broken.
The scoring system was fine. The assumptions were broken.
Deep learning, trained on GPUs, fed enough labeled data, and allowed to learn its own internal features, worked at a scale the field had not quite believed in. The result did not look like symbolic reasoning. It did not look like expert rules. It did not ask people to write down what a cat was. It absorbed examples until something inside the network became useful.
This spring has a feature the earlier ones mostly lacked: scaling keeps paying rent. More data, more compute, larger models, better training runs. Not every improvement is smooth, not every benchmark matters, and the economics are ferocious, but the basic pattern has held long enough to change the world. ChatGPT made AI a consumer habit almost overnight. Models that once felt like demos now write code, pass exams, generate images, search documents, tutor students, and draft contracts.
They also confidently invent sources that do not exist.
That is not a footnote. It is the old gap wearing new clothes.
The current systems are not brittle in the way expert systems were brittle. They are flexible, fluent, and weirdly broad. They fail differently: hallucinations, hidden reasoning errors, shallow robustness, benchmark contamination, reward hacking, security vulnerabilities, sycophancy, and a talent for making wrong answers sound like they have tenure.
The field knows this better than it used to. Modern AI researchers hedge. They say "probably" and "may" and "we do not yet understand" with the careful tone of people who have read the archives. Safety teams measure failure modes before product teams finish naming the feature. Labs publish model cards, risk evaluations, preparedness frameworks, red-team reports.
Mostly.
There is still a familiar pressure in the room: the product demo works, the investor call is tomorrow, and the future has once again been pencilled in as inevitable.
VII The Pattern
Look across the winters and the shape becomes hard to miss.
Logic Theorist really proved theorems. The Georgetown-IBM system really translated its sixty sentences. MYCIN really diagnosed infections. XCON really saved DEC a fortune. AlexNet really changed computer vision. ChatGPT really changed what millions of people expect computers to do.
The demonstrations were real.
The mistake came next, when reality was asked to resemble the demonstration. It rarely agreed. The lab problem had clean inputs; the world had ambiguity. The benchmark had fixed rules; deployment had shifting incentives. The expert had explained the rule; the actual expertise lived partly outside language. The model scored well on the test; the user needed it to be right in a situation nobody had thought to test.
AI springs are powered by genuine breakthroughs. AI winters are powered by the discovery that a breakthrough is not the same thing as an ending.
So is this spring different?
Yes, in important ways. The systems are broader. The scaling behavior is more durable. The commercial demand is not imaginary. The tools are already useful across writing, coding, design, research, customer support, education, biology, and a hundred places where "narrow demo" no longer quite fits.
And no, in the one way that matters most. The gap between demonstration and deployment is still here. It has not vanished. It has become more fluent.
The question is not whether today's AI is real. It is obviously real. The question is whether the field can resist its oldest reflex: mistaking a working demo for a finished world.
History does not say winter is inevitable. It says spring is intoxicating. It says the hard part usually begins after the demo works, when the machine leaves the clean room and meets institutions, incentives, liability, boredom, edge cases, bad data, clever users, tired operators, and the full messy weather of the world.
AI keeps learning this lesson.
Then, somehow, it has to learn it again.
Want to see the moments that shaped each cycle? Explore the timeline.