From neuron to network
SCALE UP

15 pixels > 10 digits

The network is trained on 3x5 digit patterns and their noisy variants.

The output is a probability distribution over ten classes.

It contains 634 trainable parameters: 600 weights on connections and 34 bias.

For scale: GPT‑3 has 175 billion parameters [3].

The network sees894% confidence
8
94%
9
3%
0
1%
6
1%
2
0%
634trainees parameter

600 weights + 34 bias

GPT‑3 175B ≈ ×276 million
Interactive
04/ 31