Educatifu
Open menu

A Short History of Artificial Intelligence

From Turing's question "can machines think?" to ChatGPT — the seventy-year story of AI's booms, winters, and the quiet shift from hand-written rules to learning from data.

Introductoryby Educatifuaihistory

Artificial intelligence can feel like it appeared overnight, with ChatGPT. In fact it is one of the oldest ambitions in computing, older than the field has a right to be — and its history is a cycle of soaring promises, crushing disappointments, and one deep idea that quietly won in the end: instead of telling a computer exactly what to do, let it learn from examples. This short book traces that arc, from a philosophical question to the technology reshaping the world.

1. Can machines think?

In 1950, Alan Turing asked the question that started it all in a paper called Computing Machinery and Intelligence. Rather than argue about what "thinking" means, he proposed a test: if a person conversing with a hidden machine could not tell it apart from a human, on what grounds could we deny it thinks? This became the famous Turing test.

Six years later, in the summer of 1956, a small group of researchers gathered at Dartmouth College for a workshop where John McCarthy coined the term "artificial intelligence". Their proposal was breathtakingly confident: that every aspect of learning and intelligence could "in principle be so precisely described that a machine can be made to simulate it." The field was born in optimism.

2. The age of rules

The first two decades of AI were dominated by symbolic approaches — the idea that intelligence is the manipulation of symbols according to logical rules. Early programs proved mathematical theorems, played checkers, and solved puzzles. ELIZA (1966) imitated a psychotherapist convincingly enough to unnerve its creator, and the perceptron, an early neural network, hinted at machines that could learn.

By the 1970s and 80s this matured into expert systems: programs that captured the knowledge of human specialists as thousands of hand-written if–then rules, and were sold to diagnose diseases or configure computers. For a while, it looked like the path to intelligence was simply to write down enough rules.

3. The AI winters

It wasn't. Twice, the grand promises ran far ahead of what the technology could deliver, and the funding collapsed — periods now called the AI winters.

The first, in the mid-1970s, followed influential reports concluding that AI had badly overpromised. The second, in the late 1980s, came when the expert-system boom went bust: the systems were brittle, expensive to maintain, and unable to handle anything outside their narrow rules. Each winter taught the same hard lesson. Hand-coding intelligence, rule by rule, does not scale to the messiness of the real world.

4. Learning from data

The idea that eventually broke the deadlock had been around all along: machine learning — don't program the answer, program the ability to learn it from examples. Through the 1990s and 2000s, as digital data and computing power grew, statistical methods that learned patterns from data quietly began to outperform hand-crafted rules at tasks like spam filtering, search and recommendation.

Neural networks — loose mathematical caricatures of the brain — were part of this revival. A method called backpropagation, popularised in 1986, gave a practical way to train networks with many layers. But they remained hungry for data and computing power that did not yet exist.

5. The deep learning revolution

Then, in 2012, everything accelerated. A neural network called AlexNet, trained on Nvidia GPUs, crushed the field at the ImageNet image-recognition contest. It was a public, undeniable demonstration that deep learning — big neural networks, big data, big compute — worked, and worked dramatically better than anything before.

The results came fast: speech recognition that finally worked, translation that approached fluency, and in 2016 AlphaGo defeating a world champion at Go, a game long thought to be decades beyond machines. The old debate about rules was over. The future was learning.

6. Attention, and the language machines

The final leap came from a 2017 paper with a bold title: Attention Is All You Need. It introduced the Transformer, an architecture that processes whole sequences at once and learns which parts to "pay attention" to. Transformers scaled astonishingly well: make them bigger, feed them more text, and they kept improving.

Trained on much of the public internet, these became large language models. GPT-3 (2020) could write essays and code from a prompt; ChatGPT (late 2022) put that ability in front of hundreds of millions of people and made "generative AI" a household phrase almost overnight.

7. What comes next

Today's AI is genuinely powerful and genuinely limited. It can write, draw, code and converse, yet it has no understanding in the human sense, confidently invents falsehoods ("hallucinations"), and reflects the biases of its training data. Whether scaling these models leads toward general intelligence, or toward a new plateau and another winter, is the field's great open question.

What is clear is the shape of the story. Seventy years of trying to write down intelligence failed; a few decades of letting machines learn it succeeded beyond expectation. Understanding that arc — its hype, its winters, and the ideas that endured — is the best preparation for thinking clearly about whatever the machines do next.

← All books

Bring us the problem, not a perfect specification

Tell us what needs to change, who it affects and any important deadline. We will review the context and reply with useful next questions.

  1. 01Share contextDescribe the workflow, constraint or risk.
  2. 02Clarify togetherWe identify missing facts and useful options.
  3. 03Choose a startAgree a focused assessment or delivery step.
Start a conversation