The T in GPT: The 2017 Transformer Paper

Attention Is All You Need, explained. A 2017 Google paper called Attention Is All You Need introduced the Transformer, the design behind GPT and most modern chatbots. Here is the idea.

Figure from the Transformer paper showing attention lines linking words in a sentence

Attention links between words, from the 2017 paper: Google, CC BY-SA 4.0

A short paper with a big afterlife

In June 2017 a team of eight researchers at Google posted a paper with a catchy title, Attention Is All You Need. It was presented at the NeurIPS conference later that year. The paper proposed a new neural network design called the Transformer and showed that it beat earlier systems at translating English into German and French while training faster.

Few papers have aged so well. The T in GPT, the family of models behind ChatGPT, stands for Transformer. So do the T in BERT and the core of nearly every large language model in use today.

Video: Transformers, the tech behind LLMs | Deep Learning Chapter 5 (3Blue1Brown), embedded from YouTube.

Our Short

The T in ChatGPT comes from one 2017 paper

The 2017 Google paper 'Attention Is All You Need' introduced the transformer architecture behind GPT and most modern language models.

YouTube @StrawberryLemonadAI (launching soon). Photos in this Short: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg - DeKostia (CC BY-SA 4.0) via Wikimedia Commons | PDC server room.jpg - Esquilo (CC BY-SA 3.0) via W

The problem it solved

Before 2017, the leading language models were recurrent networks. They read text one word at a time, passing a running summary from step to step. That made them slow to train, because each step had to wait for the last, and they tended to forget details from far back in a long sentence.

The Transformer dropped that one-at-a-time reading. It looks at every word in a passage at once and lets each word decide which other words matter to it. That mechanism is called self-attention.

How attention works, in plain words

Picture the sentence: the trophy did not fit in the suitcase because it was too big. To understand it, a reader must link it to trophy. In a Transformer, each word builds a small query, roughly what am I looking for, and every word also offers a key, roughly what do I contain. Matching queries against keys produces attention weights, and each word then blends in information from the words it attends to most.

The model runs many of these attention heads side by side, each free to track a different kind of relationship, and stacks many layers of them. Because the calculations happen in parallel, they map neatly onto graphics chips, which made it practical to train far bigger models on far more text.

From translation to chatbots

Researchers quickly realized the design was general. In 2018 OpenAI released its first GPT model, trained to predict the next word on a large body of text, and Google released BERT. Each new generation grew larger, and Transformers spread to images, audio, code and even protein biology.

The authors went on to work across the industry, several founding their own companies. Their idea, that attention alone could carry the load, became the foundation of the current AI boom.

Media credits
  • Attention links between words, from the 2017 paper: Google, CC BY-SA 4.0
  • The Transformer architecture: dvgodoy, CC BY 4.0
  • Short photos: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg - DeKostia (CC BY-SA 4.0) via Wikimedia Commons | PDC server room.jpg - Esquilo (CC BY-SA 3.0) via Wikimedia Commons | Backlit keyboard.jpg - Colin (CC BY-SA 4.0) via Wikimedia Commons | Wikimedia Foundation Servers-8055 13.jpg - Victorgrigas (CC BY-SA 3.0) via Wikimedia Commons | Code on computer monitor (Unsplash).jpg - Markus Spiske markusspiske (CC0) via Wikimedia Commons

Text written by Strawberry Lemonadai.

← Deep Blue vs Kasparov, 1997ChatGPT's Launch and Record Growth →

Shop

The Strawberry Lemonadai collection

Tees and hoodies in the Strawberry Lemonadai colors, printed to order in the USA. Use code FIRST15 from the newsletter for 15% off your first order.

Watch

Strawberry Lemonadai Shorts

Quick, accurate, under a minute. Every Short is original and made only for this channel.

YouTube @StrawberryLemonadAI
Channel launching soon. Subscribe on the newsletter to hear first.
Get notified
A computer beat the world chess champion in 1997Photos: IBM Deep Blue at Computer History Museum (9361685537).jpg - Anton Chiang from Cupertino, CA, USA (CC BY 2.0) via Wikimedia Commons | Garry Kasparov (37097592314).jpg - Gage Skidmore from Peoria, AZ, United States of America (CC BY-SA 2.0) via Wikimedia Commons | Chess game Staunton No. 6 perfil view 8.jpg - Wilfredor (CC0) via Wikimedia Commons | Chess game Staunton No. 6.jpg - Wilfredor (CC0) via
An AI solved a 50 year biology puzzle and won a Nobel PrizePhotos: Human Oxy-Hemoglobin Protein.jpg - PDB code 2DN1 1.25 a resolution crystal structures of human (CC0) via Wikimedia Commons | Ribbon diagram of the DED.jpg - BQUB16-Oibanez (CC BY-SA 4.0) via Wikimedia Commons | Laboratory pipettes.jpg - J.N. Eskra (CC BY-SA 4.0) via Wikimedia Commons | Use of a Multichannel Pipette for High-Throughput Liquid Handling in a Biosafety Cabinet.jpg - Siduduziwe Nxumal
ChatGPT hit an estimated 100 million users in 2 monthsPhotos: BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Diverse people using phones.jpeg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Crowd of people with phones.jpg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Sam Altman CropEdit James Tamim.jpg - TechCrunch (CC BY 2.0) via Wikimedia Commons | Ilya Sutskever and Sam Altman in TAU.jpg - Eladkarmel (CC B
The T in ChatGPT comes from one 2017 paperPhotos: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg
The word robot was invented for a 1920 playPhotos: Karel Čapek podepisuje první výtisky Povětroně, Pestrý týden 27.1.1934.jpg - Unknown authorUnknown author (Public domain) via Wikimedia Commons | Karel Čapek 30.léta.jpg - re-photo by David Sedlecký (Public domain) via Wikimedia Commons | Plakat za predstavo R.U.R Rossmus Universal Robots v Narodnem gledališču v Mariboru 28. oktobra 1933.jpg - Unknown authorUnknown author (Public domain) via Wikim
Newsletter

The weekly AI digest

The AI stories that actually matter this week, explained in plain English, plus one useful tool tip. No hype, no jargon, unsubscribe anytime.