The T in GPT: The 2017 Transformer Paper
Attention Is All You Need, explained. A 2017 Google paper called Attention Is All You Need introduced the Transformer, the design behind GPT and most modern chatbots. Here is the idea.

Attention links between words, from the 2017 paper: Google, CC BY-SA 4.0
A short paper with a big afterlife
In June 2017 a team of eight researchers at Google posted a paper with a catchy title, Attention Is All You Need. It was presented at the NeurIPS conference later that year. The paper proposed a new neural network design called the Transformer and showed that it beat earlier systems at translating English into German and French while training faster.
Few papers have aged so well. The T in GPT, the family of models behind ChatGPT, stands for Transformer. So do the T in BERT and the core of nearly every large language model in use today.
Video: Transformers, the tech behind LLMs | Deep Learning Chapter 5 (3Blue1Brown), embedded from YouTube.
The T in ChatGPT comes from one 2017 paper
The 2017 Google paper 'Attention Is All You Need' introduced the transformer architecture behind GPT and most modern language models.
YouTube @StrawberryLemonadAI (launching soon). Photos in this Short: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg - DeKostia (CC BY-SA 4.0) via Wikimedia Commons | PDC server room.jpg - Esquilo (CC BY-SA 3.0) via W
The problem it solved
Before 2017, the leading language models were recurrent networks. They read text one word at a time, passing a running summary from step to step. That made them slow to train, because each step had to wait for the last, and they tended to forget details from far back in a long sentence.
The Transformer dropped that one-at-a-time reading. It looks at every word in a passage at once and lets each word decide which other words matter to it. That mechanism is called self-attention.
How attention works, in plain words
Picture the sentence: the trophy did not fit in the suitcase because it was too big. To understand it, a reader must link it to trophy. In a Transformer, each word builds a small query, roughly what am I looking for, and every word also offers a key, roughly what do I contain. Matching queries against keys produces attention weights, and each word then blends in information from the words it attends to most.
The model runs many of these attention heads side by side, each free to track a different kind of relationship, and stacks many layers of them. Because the calculations happen in parallel, they map neatly onto graphics chips, which made it practical to train far bigger models on far more text.
From translation to chatbots
Researchers quickly realized the design was general. In 2018 OpenAI released its first GPT model, trained to predict the next word on a large body of text, and Google released BERT. Each new generation grew larger, and Transformers spread to images, audio, code and even protein biology.
The authors went on to work across the industry, several founding their own companies. Their idea, that attention alone could carry the load, became the foundation of the current AI boom.

- arXiv: Attention Is All You Need
- Wikipedia: Attention Is All You Need
- CNBC: Aidan Gomez helped invent the transformer at Google
Facts on this page were checked against these sources.
- Attention links between words, from the 2017 paper: Google, CC BY-SA 4.0
- The Transformer architecture: dvgodoy, CC BY 4.0
- Short photos: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg - DeKostia (CC BY-SA 4.0) via Wikimedia Commons | PDC server room.jpg - Esquilo (CC BY-SA 3.0) via Wikimedia Commons | Backlit keyboard.jpg - Colin (CC BY-SA 4.0) via Wikimedia Commons | Wikimedia Foundation Servers-8055 13.jpg - Victorgrigas (CC BY-SA 3.0) via Wikimedia Commons | Code on computer monitor (Unsplash).jpg - Markus Spiske markusspiske (CC0) via Wikimedia Commons
Text written by Strawberry Lemonadai.
← Deep Blue vs Kasparov, 1997ChatGPT's Launch and Record Growth →







