Neural Networks Explained

How machines learn from examples. Neural networks learn by adjusting millions of tiny numbers until their guesses improve. A plain-English guide to the engine behind modern AI.

Diagram of a simple neural network linking image features to labels for a starfish and a sea urchin

A simple network linking features to answers: Mikael Häggström, M.D., CC0

Many small calculators wired together

A neural network is a large collection of simple units, often called neurons, arranged in layers. Each unit takes in some numbers, multiplies each by a weight, adds them up and passes the result through a simple rule that decides how strongly it fires. The output of one layer becomes the input of the next.

On its own, one unit is not smart. The power comes from scale. With enough units and layers, the network can represent very complicated patterns, such as which arrangements of pixels look like a cat or which words tend to follow a phrase.

Video: But what is a neural network? | Deep learning chapter 1 (3Blue1Brown), embedded from YouTube.

Learning by being wrong, slightly less each time

A new network starts with random weights and makes terrible guesses. Training fixes that. You show it an example, compare its answer with the correct one, and measure the error. An algorithm called backpropagation then works out how much each weight contributed to that error, and gradient descent nudges every weight a tiny step in the direction that would have helped.

Repeat that millions of times across a large dataset and the weights settle into values that work. Nobody writes the rules by hand. The network discovers useful features on its own, edges and textures in early layers, shapes and objects in later ones.

A long road to success

The idea is old. Warren McCulloch and Walter Pitts described artificial neurons in 1943, Frank Rosenblatt described the perceptron in 1958, and his Mark I Perceptron machine, shown publicly in 1960, learned to recognize simple images. Enthusiasm faded when early networks hit hard limits. In 1986 a paper by David Rumelhart, Geoffrey Hinton and Ronald Williams helped popularize backpropagation for multi-layer networks.

The real breakthrough came in 2012, when a deep network called AlexNet won a major image recognition contest by a wide margin. It was trained on graphics chips using a large labeled photo collection. Data, computing power and better methods had finally lined up.

Where you meet them

Neural networks now power speech recognition, photo search, translation, spam filters, recommendation feeds and chatbots. Different shapes suit different jobs: convolutional networks for images, and Transformers for language and much more. All share the same basic recipe of weighted connections tuned by example, which is why understanding this one idea unlocks most of modern AI.

Media credits
  • A simple network linking features to answers: Mikael Häggström, M.D., CC0
  • The Mark I Perceptron, an early learning machine: National Museum of the U.S. Navy, Public domain

Text written by Strawberry Lemonadai.

← AlphaFold and the 2024 Chemistry NobelLarge Language Models Explained →

Shop

The Strawberry Lemonadai collection

Tees and hoodies in the Strawberry Lemonadai colors, printed to order in the USA. Use code FIRST15 from the newsletter for 15% off your first order.

Watch

Strawberry Lemonadai Shorts

Quick, accurate, under a minute. Every Short is original and made only for this channel.

YouTube @StrawberryLemonadAI
Channel launching soon. Subscribe on the newsletter to hear first.
Get notified
A computer beat the world chess champion in 1997Photos: IBM Deep Blue at Computer History Museum (9361685537).jpg - Anton Chiang from Cupertino, CA, USA (CC BY 2.0) via Wikimedia Commons | Garry Kasparov (37097592314).jpg - Gage Skidmore from Peoria, AZ, United States of America (CC BY-SA 2.0) via Wikimedia Commons | Chess game Staunton No. 6 perfil view 8.jpg - Wilfredor (CC0) via Wikimedia Commons | Chess game Staunton No. 6.jpg - Wilfredor (CC0) via
An AI solved a 50 year biology puzzle and won a Nobel PrizePhotos: Human Oxy-Hemoglobin Protein.jpg - PDB code 2DN1 1.25 a resolution crystal structures of human (CC0) via Wikimedia Commons | Ribbon diagram of the DED.jpg - BQUB16-Oibanez (CC BY-SA 4.0) via Wikimedia Commons | Laboratory pipettes.jpg - J.N. Eskra (CC BY-SA 4.0) via Wikimedia Commons | Use of a Multichannel Pipette for High-Throughput Liquid Handling in a Biosafety Cabinet.jpg - Siduduziwe Nxumal
ChatGPT hit an estimated 100 million users in 2 monthsPhotos: BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Diverse people using phones.jpeg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Crowd of people with phones.jpg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Sam Altman CropEdit James Tamim.jpg - TechCrunch (CC BY 2.0) via Wikimedia Commons | Ilya Sutskever and Sam Altman in TAU.jpg - Eladkarmel (CC B
The T in ChatGPT comes from one 2017 paperPhotos: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg
The word robot was invented for a 1920 playPhotos: Karel Čapek podepisuje první výtisky Povětroně, Pestrý týden 27.1.1934.jpg - Unknown authorUnknown author (Public domain) via Wikimedia Commons | Karel Čapek 30.léta.jpg - re-photo by David Sedlecký (Public domain) via Wikimedia Commons | Plakat za predstavo R.U.R Rossmus Universal Robots v Narodnem gledališču v Mariboru 28. oktobra 1933.jpg - Unknown authorUnknown author (Public domain) via Wikim
Newsletter

The weekly AI digest

The AI stories that actually matter this week, explained in plain English, plus one useful tool tip. No hype, no jargon, unsubscribe anytime.