Artificial intelligence can feel as if it arrived overnight. In only a few months, conversational systems moved from research demonstrations into homes, classrooms, and workplaces. People who had never studied machine learning were suddenly asking a computer to explain ideas, draft plans, write code, and continue a conversation.
But AI did not begin with ChatGPT. It did not begin with deep learning, the internet, or even the modern computer. Its history is a long chain of questions: Can reasoning be described? Can a machine follow those descriptions? Can it learn from experience? Can it understand language well enough to become useful to regular people?
The short answer: artificial intelligence officially became a scientific field in 1956.
The more meaningful answer begins decades earlier and continues through several cycles of ambition, disappointment, and reinvention.
What do I mean by artificial intelligence?
I use artificial intelligence to describe computer systems that perform tasks associated with human intelligence. These tasks include recognizing patterns, understanding language, learning from examples, making predictions, solving problems, and choosing an action.
Most AI available in April 2023 is narrow AI. It may be extremely capable, but it is built or trained for a limited range of purposes. A system that recognizes images does not automatically understand medicine. A chess system does not become a financial adviser. A language model can work across many subjects, yet that breadth should not be confused with human understanding.
Artificial general intelligence, or AGI, is the idea of a system able to learn and reason across intellectual tasks with broad, flexible capability. It remains an aspiration, not an established achievement.
This distinction is important. AI is not one machine and not one level of intelligence. It is a field containing many methods, goals, and degrees of capability.
Before AI had a name
The scientific story starts with the foundations of computation.
In 1936, Alan Turing described an abstract machine that could carry out a sequence of precise instructions. A Turing machine is not a physical computer sitting on a desk. It is a mathematical model—this is like describing the rules of cooking before building the kitchen. The model helped define what a general-purpose computer could calculate and where calculation has limits.
In 1943, Warren McCulloch and Walter Pitts published a mathematical model of an artificial neuron. Their work suggested that networks of simple units could perform logical operations. The model was elementary compared with today’s neural networks, but the underlying idea was powerful: complex behavior might emerge from many simple connections.
In 1949, Donald Hebb described how connections between neurons could strengthen through repeated activity. In everyday terms, this is like a path through grass becoming clearer each time someone walks across it.
Then, in 1950, Turing published Computing Machinery and Intelligence. Instead of becoming trapped in the definition of “thinking,” he proposed an operational question: could a machine converse in a way that made its responses difficult to distinguish from those of a person? His imitation game later became known as the Turing Test.
These ideas did not yet form a unified field. They were pieces on a table: computation, logic, learning, language, and models of the brain. In 1956, researchers gave the puzzle a name.
The summer when AI officially began
The Dartmouth Summer Research Project on Artificial Intelligence, held in New Hampshire in 1956, is widely recognized as the formal birth of AI as a discipline.
John McCarthy organized the workshop and helped coin the term artificial intelligence. The 1955 proposal made an ambitious conjecture: aspects of learning and intelligence could be described precisely enough for a machine to simulate them. Participants and contributors connected with the project included Marvin Minsky, Claude Shannon, Nathaniel Rochester, Allen Newell, Herbert Simon, and others who would influence computing for decades.
Why does this gathering matter? Not because every problem was solved during one summer. It matters because a community formed around a shared research agenda. Naming a field gives people a place to gather, disagree, build tools, train students, and measure progress.
The optimism was enormous. Early researchers believed major features of intelligence might be reproduced quickly. That confidence generated important work, but it also created expectations that the available computers, data, and methods could not meet.
The first approach: intelligence as rules
Early AI relied heavily on symbolic AI. Engineers represented knowledge using symbols, rules, and logical relationships. This is like giving a machine a very large instruction manual: if a particular condition is true, follow a particular rule.
This approach produced important milestones. Allen Newell, Herbert Simon, and J. C. Shaw developed Logic Theorist, an early program that proved mathematical theorems. John McCarthy created LISP, a programming language that became central to AI research. Joseph Weizenbaum’s ELIZA, introduced in the 1960s, used pattern matching to imitate a conversation with a psychotherapist. Shakey the Robot combined sensing, planning, and movement in a mobile system.
These systems showed that computers could manipulate symbols and perform tasks that appeared intelligent. They also exposed the weakness of hand-written knowledge. The real world contains too many exceptions. A rule system can appear impressive inside its designed boundary and become fragile when the situation changes.
When promises outran results, enthusiasm and funding declined. The periods of reduced confidence became known as AI winters. I see these winters as an essential part of the story. Innovation does not travel in a straight line. A powerful idea can arrive before the infrastructure needed to make it practical.
Expert systems: knowledge becomes a product
During the 1970s and 1980s, expert systems created another wave of adoption. These programs captured specialist knowledge as rules and used it to recommend conclusions in areas such as chemistry, medicine, and equipment configuration.
For organizations, the appeal was clear: preserve valuable expertise and make it repeatable. Yet the systems were expensive to build and maintain. When knowledge changed, people had to update the rules. When an unexpected case appeared, the software could fail in ways that were difficult to explain or repair.
The limitation was not that rules were useless. Rules remain valuable today. The limitation was expecting hand-authored rules to represent the full complexity of the world.
The decisive shift: machines learn from data
The next major change was from programming every answer to training systems from examples.
Machine learning allows a system to identify patterns in data and use those patterns to make a prediction. This is like teaching someone to recognize a dog by showing many examples rather than writing a perfect checklist for every breed, angle, and lighting condition.
Statistical methods became increasingly important through the 1990s and 2000s. In 1997, IBM’s Deep Blue defeated reigning world chess champion Garry Kasparov in a match under standard tournament controls. Deep Blue relied on massive search, specialized hardware, evaluation methods, and chess knowledge. It was a landmark demonstration of focused machine capability, not general intelligence.
Deep learning changes the scale
Neural networks had existed for decades, but three conditions finally came together: much larger datasets, improved algorithms, and powerful graphics processing units.
In 2012, a deep neural network commonly known as AlexNet achieved a breakthrough result in the ImageNet image-recognition competition. Deep learning uses neural networks with many processing layers—this is like passing an image through a sequence of specialists, where early layers detect simple edges and later layers combine them into more meaningful patterns.
In 2016, DeepMind’s AlphaGo defeated Lee Sedol four games to one. Go had long been considered especially difficult because the number of possible positions is enormous and good play depends on patterns that are difficult to express as simple rules. AlphaGo combined deep neural networks with search and reinforcement learning, where a system improves through feedback from its actions.
Transformers make language the interface
In 2017, Google researchers published Attention Is All You Need and introduced the Transformer architecture.
The key idea is attention. Instead of processing every word only in a fixed sequence, a model can learn which parts of the input matter most to one another. This is like reading a paragraph and mentally connecting a pronoun to the person named several sentences earlier. Transformers could train efficiently in parallel and became the foundation for a new generation of language models.
Large language models learn statistical patterns from enormous collections of text. They generate an answer by predicting useful sequences of tokens, which are small pieces of language. A token is like a puzzle piece: it may be a word, part of a word, or punctuation. By repeatedly choosing the next fitting piece, the model builds a response.
This mechanism can produce remarkably fluent language, but fluency is not the same as truth. A model may generate incorrect information with confidence. Its output needs judgment, verification, and responsible experience design.
ChatGPT brings AI to the conversation
OpenAI released ChatGPT as a research preview on November 30, 2022. The breakthrough for regular users was not simply that a language model could generate text. It was the conversational experience. People could ask follow-up questions, refine an instruction, challenge an answer, and continue working in ordinary language.
That interface reduced the distance between advanced AI and human intent. A prompt became a new kind of control surface. Instead of learning a programming language or navigating a complex menu, a person could describe the desired outcome.
In March 2023, OpenAI introduced GPT‑4, a multimodal model designed to accept text and image inputs and produce text outputs, although image input was still not broadly available at launch. This showed the direction of travel: AI systems were beginning to connect multiple forms of information within a more general interaction model.
From my perspective in Experience Engineering, this is as important as the model architecture. Technology becomes transformative when capability and usability meet. The engine matters, but so do the steering wheel, dashboard, road rules, and trust of the driver.
So, when did AI begin?
If I must choose one date, I choose 1956, when the Dartmouth workshop formally established artificial intelligence as a field.
But AI has several beginnings. Turing helped define computation in 1936. McCulloch and Pitts modeled an artificial neuron in 1943. Turing reframed machine intelligence in 1950. Dartmouth named the field in 1956. Machine learning shifted attention from rules to data. Deep learning expanded what pattern recognition could do. Transformers made language a powerful interface. ChatGPT brought that interface to millions of everyday conversations.
The history is not a race toward one final machine. It is a record of people repeatedly changing the question. First: Can a machine calculate? Then: Can it reason? Can it learn? Can it perceive? Can it communicate? And now: How should people and intelligent systems work together?
That final question may define the next era. The quality of AI will not be measured only by what a model can generate. It will also be measured by whether the experience is understandable, trustworthy, inclusive, safe, and genuinely valuable to the person using it.
Takeaways
What I keep
- AI officially became a scientific field at the Dartmouth workshop in 1956.
- Its foundations began earlier with computation, artificial neurons, and the question of machine intelligence.
- Symbolic AI used explicit rules; machine learning shifted the field toward patterns learned from data.
- Deep learning accelerated when data, algorithms, and GPU computing came together.
- The 2017 Transformer architecture created the foundation for modern large language models.
- ChatGPT’s November 2022 launch made advanced language AI accessible through ordinary conversation.
- As of April 2023, today’s AI remains narrow AI; AGI is still a goal, not an established reality.