The Hundred-Year Race to Build a Mind: How AI Went From a Stage Play to the Tech Reading This Sentence
A robot was named in 1920, a machine was asked to think in 1950, and it took until now for the hardware to catch up. The hundred-year story of artificial intelligence โ the dreamers, the rivalries, the winters, and the one idea that changed everything.

The dream of a thinking machine is over a century old โ the word 'robot' comes from a 1920 play and Alan Turing posed the question in 1950 โ but AI only exploded recently because computing power and data finally caught up. The modern leap traces to a single 2017 idea, the Transformer, which now powers Claude, ChatGPT, and Gemini.
Quick Answer: The dream of a thinking machine is over a century old โ the word 'robot' comes from a 1920 play and Alan Turing posed the question in 1950 โ but AI only exploded recently because computing power and data finally caught up. The modern leap traces to a single 2017 idea, the Transformer, which now powers Claude, ChatGPT, and Gemini.
Table of Contents
- The Dream Before the Machine
- Turing's Question
- The Summer That Named a Field
- The First Golden Age and Its Beautiful Lies
- The AI Winters
- Brute Force Strikes Back
- The Data-and-GPU Revolution
- Attention Is All You Need
- The Modern Players and Their Rivalries
- The Uncanny, the Funny, and the Genuinely Strange
- A Century of Ideas Waiting for the Hardware
- And Yes โ You Can Now Ask One of These to Help You Shop
- Frequently Asked Questions
The Dream Before the Machine
Let's start in a theatre.
Prague, 1921. The curtain rises on Karel ฤapek's play R.U.R. โ Rossum's Universal Robots โ and the audience is introduced to a factory full of artificial workers made from synthetic organic matter. They're not metal. They're not gears. They're disturbingly human, and they're miserable, and eventually they revolt. The play ran across Europe and then New York. And nestled inside it was a word ฤapek had borrowed from his brother Josef, pulled from the Czech robota, meaning drudgery or forced labor: robot.
That word, coined in 1920 and staged in 1921, is the first time the modern world gave a name to its oldest anxiety โ that one day, we'd make something that could do our thinking for us, and we'd have to answer for it.
The fantasy is older than ฤapek, of course. The history of artificial intelligence stretches back through mythology โ the bronze giant Talos in Greek legend, the golems of Jewish folklore, Descartes wondering in the 1630s whether animals were just very complicated machines. But the engineers had to wait for the math to arrive. It arrived, magnificently, in the 1830s in a drawing room in London.
Charles Babbage spent decades designing his Analytical Engine โ a mechanical general-purpose computer, complete with loops, conditionals, and memory โ that Victorian England never quite managed to build. But it was his collaborator, Ada Lovelace, who saw the machine's real implication before anyone else did. In her 1843 notes on the Engine, she described a device that could, in principle, compose music or handle any problem that could be expressed as relationships between symbols. She also issued the first documented AI skeptic's caveat: the Engine could only do what it was told. It had no power of originating anything.
She was right. And she was raising a question that would take another 107 years to be properly formalized.
Turing's Question
In 1950, a 37-year-old British mathematician named Alan Turing published a paper in the journal Mind that opened with one of the most audacious sentences in the history of science: "I propose to consider the question, 'Can machines think?'"
The paper was called "Computing Machinery and Intelligence," and rather than try to define "think" โ a philosophical swamp he had no interest in wading through โ Turing proposed a game. A human interrogator, communicating only through typed messages, would try to determine which of two conversational partners was human and which was a machine. If the machine could fool the interrogator often enough, Turing argued, we'd have to concede it was doing something functionally equivalent to thinking.
He called it the Imitation Game. We call it the Turing Test.
What's remarkable about the paper isn't just the test โ it's how much of the next 70 years of AI debate Turing anticipated in a single document. He addressed nine objections to machine intelligence one by one: the theological objection (God gave souls only to humans), the mathematical objection (Gรถdel's incompleteness theorem limits formal systems), the argument from consciousness (a machine can't feel anything). He swatted each one with characteristic patience. He even predicted that by the year 2000, a computer would be able to fool a human interrogator 30% of the time โ not quite right, but directionally prescient.
Turing didn't live to see the field he helped create. He died in 1954, aged 41, in circumstances that remain a source of grief and outrage. But the question he posed โ can machines think? โ became the organizing obsession of a generation of researchers who were about to give themselves a name.
The Summer That Named a Field
In the summer of 1956, about two dozen mathematicians, computer scientists, and cognitive scientists gathered at Dartmouth College in New Hampshire for a six-week workshop. It was organized by a brash 28-year-old named John McCarthy, who had already decided what to call the thing they were trying to do.
The Dartmouth workshop proposal, written by McCarthy and colleagues including Marvin Minsky, Claude Shannon, and Nathaniel Rochester, contained this sentence: "We propose that a 2 month, 10 man study of artificial intelligence be carried out during the summer of 1956."
There it was. Artificial intelligence. A field named itself into existence on a grant proposal.
McCarthy later said he chose the term partly to distinguish the work from Norbert Wiener's "cybernetics" โ he wanted a clean break, a new identity. The name stuck with the same stubborn permanence as ฤapek's "robot." For better and for worse. The phrase has spent the subsequent seven decades alternately inspiring billions in investment and triggering eye-rolls from serious researchers who felt "AI" promised more than the science could deliver.
The Dartmouth crew came in optimistic. The proposal casually assumed that "every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it." They figured they'd crack it in a summer.
They were off by about 65 years.
The First Golden Age and Its Beautiful Lies
The decade after Dartmouth was intoxicating. Programs were solving algebra problems. A system called the Logic Theorist was proving mathematical theorems โ including 38 of the 52 theorems in Whitehead and Russell's Principia Mathematica, and proving one of them more elegantly than the original. Marvin Minsky declared in 1967 that within a generation, "the problem of creating artificial intelligence will be substantially solved."
The most seductive program of the era was ELIZA.
Joseph Weizenbaum, a computer scientist at MIT, wrote ELIZA in 1966 as a demonstration โ almost a parody โ of how superficial natural language processing could be. Its most famous script, DOCTOR, mimicked a Rogerian psychotherapist by doing something embarrassingly simple: reflecting your statements back as questions. You'd type "I feel sad." It would reply "Why do you feel sad?" You'd say "My mother doesn't understand me." It would say "Tell me more about your family."
That's it. That's the whole trick.
And people fell for it. Completely, deeply, disturbingly fell for it. Weizenbaum's own secretary asked him to leave the room so she could speak to ELIZA privately. Psychiatrists wrote papers proposing that networks of ELIZA-style programs could eventually provide automated therapy to millions. Weizenbaum was horrified. He spent the rest of his career warning that the human tendency to project intelligence onto systems that merely mimic it was one of the most dangerous cognitive bugs we carry.
He was, as we now know, onto something.
Meanwhile, the field was running into a harder wall. Frank Rosenblatt had invented the perceptron in 1958 โ the ancestor of every neural network running today โ and the press went wild. The New York Times reported that the Navy expected the machine would soon be able to walk, talk, and write. Minsky and Papert published a 1969 book demonstrating the perceptron's limitations in devastating technical detail. Neural network research essentially stopped. The money began to dry up.
The AI Winters
There is a particular silence that falls when hype meets reality and reality wins. In AI history, that silence has a name: the AI winter.
The first one arrived in the mid-1970s. The context: the British government commissioned a review of AI research led by mathematician James Lighthill. His 1973 report was polite in the way that a letter of professional termination is polite โ it concluded that AI had failed to deliver on any of its major promises and recommended substantial cuts to funding. The U.S. wasn't far behind. DARPA, which had been generously funding AI since the 1960s, began demanding results and cutting programs that couldn't produce them.
Laboratories went quiet. Researchers defected to adjacent fields. The jokes began โ "AI" was said to stand for "Almost Implemented."
The thaw was partial. Expert systems โ rule-based programs that encoded human expertise as if-then logic chains โ had a genuine boom in the 1980s. Companies like Digital Equipment Corporation were using them to configure computer orders, saving tens of millions of dollars a year. Japan launched its ambitious Fifth Generation Computer project, promising to build AI machines by 1992. Everyone panicked and poured money back in.
Then came the second winter, starting in the late 1980s. Expert systems turned out to be brittle, expensive to maintain, and infuriating to update โ every time the real world changed, someone had to go back in and rewrite the rules by hand. The Japanese Fifth Generation project delivered nothing of commercial significance. DARPA cut funding again. The market for specialized AI hardware collapsed almost overnight.
By the early 1990s, the phrase "artificial intelligence" was so tainted that many researchers actively avoided it. They called their work "machine learning," "knowledge engineering," "computational statistics" โ anything but AI. The field had become the academic equivalent of a LinkedIn profile that never mentions the company you got fired from.
What nobody quite appreciated yet was that the ideas weren't wrong. The hardware just wasn't there.
Brute Force Strikes Back
On May 11, 1997, a computer called Deep Blue played the last game of a six-game match against the reigning world chess champion, Garry Kasparov. The computer won. Kasparov lost. Chess, which humans had long held up as the supreme test of intellect, had fallen.
IBM's Deep Blue wasn't intelligent in any meaningful sense. It evaluated roughly 200 million chess positions per second through brute-force search, guided by hand-coded evaluation functions written by grandmasters. It didn't "understand" chess any more than a calculator "understands" arithmetic. But it was faster than fast โ and speed, at chess, is functionally indistinguishable from genius.
The drama around the match is worth savoring. Kasparov, who had beaten Deep Blue's predecessor Deeper Blue in 1996, was visibly shaken after losing Game 2 of the 1997 rematch โ a game in which the computer made a strangely subtle sacrifice that seemed almost too human. Kasparov later accused IBM of cheating, suggesting that a grandmaster must have intervened during the game. IBM denied it. A subsequent analysis suggested the move may have been a bug โ the system, unable to find a good move, defaulted to a fallback that happened to look profound.
So the most celebrated moment of AI genius in the 1990s may have been an accident. The universe does love a punchline.
The deeper lesson of Deep Blue wasn't about chess. It was a proof of concept that raw compute, scaled aggressively, could tackle problems previously thought to require human-level cognition. The researchers who'd been quietly building better hardware and larger datasets took note.
The Data-and-GPU Revolution
For most of the 2000s, machine learning was advancing steadily but not spectacularly. Neural networks were back in fashion โ researchers like Geoffrey Hinton at the University of Toronto had kept the faith through the winters โ but training a deep network on real-world data still took weeks or months on the CPUs of the era, and the results were modest.
Then two things happened simultaneously. The internet created an ocean of labeled data that nobody had previously had access to. And a Canadian PhD student named Alex Krizhevsky figured out that the GPUs being sold to gamers โ optimized for the parallel matrix math of rendering 3D graphics โ were almost perfectly suited for training neural networks.
In 2012, Krizhevsky, Ilya Sutskever, and Hinton entered a model called AlexNet into the ImageNet Large Scale Visual Recognition Challenge, an annual competition where teams built systems to classify images. AlexNet achieved a top-5 error rate of 15.3%. The second-place entry, using conventional methods, scored 26.2%. That gap โ more than 10 percentage points โ was not an incremental improvement. It was a discontinuity. It was the kind of result that makes scientists go quiet in a room.
The deep learning era had begun.
Within a year, virtually every serious computer vision team in the world had switched to neural networks trained on GPUs. Within three years, the approach had spread to speech recognition, machine translation, and drug discovery. The compute and data had finally caught up to ideas that had been waiting, patient and dusty, since the 1980s.
- 2006: Hinton and Salakhutdinov publish a paper reviving deep neural networks
- 2009: ImageNet dataset launched by Fei-Fei Li โ over 14 million labeled images
- 2012: AlexNet wins ImageNet by a stunning margin; GPU-accelerated training goes mainstream
- 2014: Generative Adversarial Networks (GANs) invented by Ian Goodfellow
- 2016: DeepMind's AlphaGo defeats world Go champion Lee Sedol
Moore's Law had been grinding away for decades, roughly doubling transistor counts every two years. The compute available for AI training doubled approximately every 3.4 months between 2012 and 2018 โ a pace Moore himself might have found alarming.
Attention Is All You Need
If AlexNet was the alarm clock, the 2017 paper "Attention Is All You Need" was the sunrise.
Eight researchers at Google published a paper describing a new architecture for processing sequential data โ language, primarily โ that discarded the recurrent approaches everyone had been using and replaced them with a mechanism called self-attention. The idea: instead of processing words one by one, left to right, in sequence, the model should be allowed to consider all words in a sentence simultaneously and learn which ones mattered most to the meaning of each other word.
They called the architecture the Transformer.
I remember reading a summary of this paper a few years after its publication and thinking it sounded like the kind of incremental improvement that fills the middle pages of conference proceedings. I was wrong in a way that has become professionally embarrassing to recall. The Transformer didn't just improve language models โ it became the foundational architecture for almost everything that followed. GPT. BERT. T5. Every large language model you've heard of in the last five years is, at its core, a scaled-up Transformer.
The key insight was attention itself โ the ability to weigh context dynamically. Given the sentence "The trophy didn't fit in the suitcase because it was too big," a Transformer can learn that "it" refers to the trophy, not the suitcase, by attending to the right parts of the sentence. Earlier models would get this wrong constantly. The Transformer got it right, and as it got bigger, it got righter.
By 2020, OpenAI had trained GPT-3 โ a Transformer model with 175 billion parameters, trained on hundreds of gigabytes of text. To give you a sense of scale: GPT-2, released in 2019, had 1.5 billion parameters and OpenAI initially declined to release it, citing concerns about misuse. GPT-3 was more than 100 times larger and could write code, complete essays, answer questions, and translate languages, all from a single model with no task-specific training. It wasn't perfect. But it was clearly something new.
The Modern Players and Their Rivalries
The story of modern AI is inseparable from the story of a handful of organizations whose relationships read like a prestige TV drama โ alliances, defections, competing visions of the apocalypse, and genuinely enormous amounts of money.
OpenAI was founded in 2015 as a nonprofit research lab with a peculiar mission: develop artificial general intelligence for the benefit of humanity, which meant, among other things, not letting any one company control it. Its founding donors included Elon Musk and Sam Altman. Musk later left the board. The nonprofit eventually created a "capped profit" subsidiary to attract investment. The ironies have accumulated.
Google DeepMind is the product of two separate histories. DeepMind was a London-based research lab that Google acquired in 2014 in what was, at the time, one of the largest acquisitions of an AI company ever made. Google Brain was a separate internal team that had helped create the Transformer. In 2023, Google merged them into a single unit under Demis Hassabis โ a former chess prodigy and neuroscientist who had co-founded DeepMind. The combined entity has produced some of the most significant AI breakthroughs of the last decade, including AlphaFold, which predicted the structure of virtually every known protein and may do more for human health than any AI system ever built.
Anthropic was founded in 2021 by Dario Amodei, Daniela Amodei, and several other researchers who left OpenAI โ reportedly over disagreements about the pace and safety practices of AI development. Dario had been OpenAI's VP of Research. The Amodei siblings built Anthropic around a safety-first philosophy, developing techniques like Constitutional AI to make models more controllable. Their model, Claude, competes directly with ChatGPT while positioning itself as the more cautious, more honest alternative. Whether you find that positioning reassuring or merely clever marketing probably says something about you.
Meta has its own large AI research operation, and has made the unusual strategic choice to open-source its LLaMA models โ a decision that has either democratized AI research or handed everyone in the world a toolkit for mischief, depending on your perspective.
Google Gemini represents Google's consumer-facing AI push, built on Transformer-based infrastructure and competing head-on with GPT-4 and Claude for the daily use cases that hundreds of millions of people are now starting to reach for automatically.
The talent wars between these organizations have been extraordinary. Individual researchers with the right combination of expertise and track record command compensation packages that would have seemed satirical a decade ago. The flow of researchers between labs โ and the occasional spectacular departure to found a new one โ has given the field the atmosphere of a very well-funded reality competition.
If you want a framework for evaluating who's actually winning any of this, our guide on how to compare products applies at the corporate level too: look past the marketing, focus on what actually ships.
The Uncanny, the Funny, and the Genuinely Strange
Let's talk about the moments that remind you this story is also deeply, structurally weird.
Move 37.
In March 2016, Google DeepMind's AlphaGo played the second game of a five-game match against Lee Sedol, then one of the world's strongest Go players. Go is a board game of almost incomprehensible complexity โ the number of possible positions exceeds the number of atoms in the observable universe โ and had long been considered AI-proof. Grandmasters thought they had at least a decade before machines could challenge them.
In Game 2, on the 37th move, AlphaGo placed a stone on the 5th line near the left side of the board. Every professional commentator watching assumed it was a mistake. Fan Hui, the European Go champion who was in the room, said he had never seen anything like it. AlphaGo's own analysis suggested the probability of a human professional making that move was about 1 in 10,000.
The move turned out to be brilliant. Quietly, inevitably brilliant โ like a chess sacrifice that takes 20 moves to resolve. Lee Sedol stared at the board, stood up, left the room for 15 minutes. AlphaGo won the game. It won the match 4-1. Lee Sedol retired from competitive Go in 2019, saying that AI was an entity that cannot be defeated.
Microsoft's Tay.
In 2016, Microsoft launched a chatbot called Tay on Twitter with the tagline "the more you talk to it, the smarter it gets." Within 16 hours, coordinated users had taught it to deny the Holocaust, express support for genocide, and produce content so offensive that Microsoft deleted the account and issued a public apology. Tay was offline for good within a day of launch.
The lesson was not that AI was evil. It was that AI, in 2016, was a perfect mirror โ it would learn exactly what you taught it, with no filter, no values, no ability to say "this seems wrong." The embarrassment was total and instructive.
The hallucinating lawyers.
In 2023, two New York lawyers submitted a court brief containing citations to multiple cases that didn't exist. They had used ChatGPT to help research the brief and had apparently assumed the model's confident prose was backed by real sources. The judge was not amused. The cases โ complete with realistic-sounding names, docket numbers, and legal reasoning โ had been invented wholesale by a language model doing what language models do: generating plausible text. The term "hallucination" was suddenly everywhere.
I had my own smaller version of this moment earlier that year. I asked a chatbot to summarize a technical report I was reviewing for TheWinner, received a beautifully structured summary, and only noticed on the third read that one of the "key findings" it cited didn't appear anywhere in the original document. The model had read the surrounding text, inferred what should have been in the report, and written it in. Confidently. Without a footnote.
This is, to put it gently, a known issue. And it's why understanding how to actually evaluate reviews and summaries โ whether they come from humans or machines โ is still a skill worth developing.
A Century of Ideas Waiting for the Hardware
Here's what I find most striking about this whole story, stepping back and looking at the shape of it.
Almost none of the core ideas are new.
The perceptron was invented in 1958. Backpropagation โ the training algorithm that powers modern neural networks โ was understood in principle in the 1960s and formalized in 1986. Convolutional neural networks for image recognition were working in the 1990s. The concept of attention in neural networks predates the 2017 Transformer paper by years. Recurrent networks, generative models, reinforcement learning โ all of it was developed, in substantial form, during periods when the hardware couldn't actually run it at useful scales.
What changed between 2012 and 2023 wasn't primarily ideas. It was three things converging:
- Compute. GPU clusters gave researchers the ability to train models that would have taken decades on 1990s hardware. Cloud computing made that compute accessible without owning a data center.
- Data. The internet created a training corpus of essentially everything humans have ever written, photographed, or filmed. ImageNet was 14 million labeled images. GPT-3 trained on hundreds of billions of words.
- Scale laws. Researchers discovered, empirically, that making models bigger on more data produced predictable, smooth improvements โ a finding that turned AI engineering into something closer to a known quantity. You could forecast, roughly, how much better a model would get if you gave it 10 times more compute.
Geoffrey Hinton kept the neural network faith for 30 years of winters, collecting a Nobel Prize in Physics in 2024 for his foundational work. Yann LeCun built convolutional networks in the 1990s that couldn't run fast enough to matter commercially. Yoshua Bengio published on deep learning for decades before the world cared. The ideas sat on shelves, waiting.
Ada Lovelace wrote in 1843 that a machine could in principle handle any problem expressible as relationships between symbols. She had the concept. She had the math. She did not have NVIDIA.
This is, depending on how you look at it, either the greatest vindication of pure research in the history of science โ or evidence that the gap between a good idea and a world-changing technology is measured not in equations but in transistors.
Probably both.
The numbers are worth sitting with for a moment:
| Milestone | Year | Scale | |---|---|---| | Turing's paper | 1950 | 1 question | | ELIZA | 1966 | ~200 lines of code | | Deep Blue | 1997 | 200M positions/second | | AlexNet | 2012 | 60M parameters | | GPT-3 | 2020 | 175B parameters | | GPT-4 (estimated) | 2023 | ~1 trillion parameters |
From one question to a trillion parameters, in 73 years. Karel ฤapek would have written a second play.
And Yes โ You Can Now Ask One of These to Help You Shop
The machines haven't quite gained consciousness โ probably โ but they've gotten remarkably useful for everyday decisions. If you're shopping for hardware that actually benefits from the AI revolution we've just spent 4,000 words describing, this is the moment for it: from AI Translation Earbuds that do real-time language processing directly in your ear, to Gaming Laptops built on GPUs descended directly from the AlexNet breakthrough, to a proper 27 Inch Computer Monitor large enough to actually read all this history on โ our best-of lists on Amazon.ae are updated regularly and will save you the kind of decision paralysis that even Turing didn't have a solution for.
Frequently Asked Questions
Q: Who actually invented artificial intelligence?
No single person invented AI โ it emerged from overlapping contributions. John McCarthy coined the term "artificial intelligence" and organized the foundational 1956 Dartmouth Workshop. Alan Turing formalized the question in 1950. Claude Shannon laid groundwork in information theory. Geoffrey Hinton, Yann LeCun, and Yoshua Bengio built the neural network foundations of modern AI over several decades, work recognized with the 2018 Turing Award. If you need one name, McCarthy is the closest thing to a founding parent โ but he'd be the first to point at a crowded room of colleagues.
Q: What were the AI winters and why did they happen?
The two major AI winters โ roughly the mid-1970s and the late 1980s into early 1990s โ were periods when research funding dried up after AI failed to deliver on overconfident promises. The first was triggered partly by the 1973 Lighthill Report in the UK. The second followed the collapse of the expert systems boom. In both cases, the underlying problem was the same: the hardware couldn't execute the ideas at the scale required to make them actually work. The ideas weren't wrong. The timing was.
Q: What is the Transformer and why does it matter?
The Transformer, introduced in the 2017 paper "Attention Is All You Need," is a neural network architecture built around self-attention โ the ability to process all parts of an input sequence simultaneously and weigh how much each part should influence the interpretation of every other part. Before the Transformer, sequence models processed text sequentially, left to right, which made it hard to capture long-range dependencies. The Transformer's parallel, attention-based approach proved dramatically more scalable on GPUs. Virtually every large language model โ GPT, Claude, Gemini, and others โ is built on Transformer architecture.
Q: Is the current AI boom another bubble that will lead to another winter?
The honest answer is: maybe, partially, eventually. There's genuine substance to what's been built โ the language models, the protein-folding breakthroughs, the image generation systems โ that represents real scientific progress, not smoke. But parts of the commercial hype, particularly around short-term AI autonomy claims, are running ahead of what the systems can reliably do. The hallucination problem alone is a fundamental challenge. A correction in expectations seems probable; whether that constitutes a true winter or just a sober plateau depends on how quickly the remaining technical problems yield to the enormous resources currently being thrown at them.
Q: How did AlphaGo's Move 37 change how we think about AI creativity?
Move 37 was significant not because a computer won Go โ that was inevitable given enough compute โ but because the winning move was surprising in a way that felt creative. Go experts watching live assumed it was a mistake because no human professional would have considered it. AlphaGo had learned from millions of self-played games to evaluate positions that humans had never explored, and in doing so produced a move that expanded the game's strategic vocabulary. Lee Sedol said in a post-match interview that Move 37 was "beautiful." It's one of the few moments in AI history where "the machine did something humans hadn't thought of" was literally true, not metaphorical.
Q: What's the difference between what Anthropic, OpenAI, and Google DeepMind are actually building?
At the technical level, all three organizations are building large Transformer-based language models, and the architectures are converging. The real differences are in emphasis, methodology, and corporate philosophy. OpenAI has moved fastest to commercialization and consumer products. Google DeepMind retains the deepest bench of fundamental research talent and produces work โ like AlphaFold โ with broader scientific scope. Anthropic invests heavily in interpretability and alignment research, attempting to understand mechanistically what's happening inside models rather than just training them to behave well. These differences matter more as the systems become more capable and the consequences of getting them wrong become larger.

Omar has been testing consumer electronics since before smartphones had app stores. He covers laptops, TVs, smart home devices, and gaming gear with a focus on real-world performance over spec-sheet bragging.