01 · Embedding Real vectors, trained on Shakespeare

Meaning becomes distance.

Every token gets a vector — a point in space. Training pulls words that behave alike close together, with no dictionary and no labels. Below is the real embedding table from a model that read only Shakespeare, flattened from 128 dimensions to 2. These are the very sub-word tokens from the last station — the · still marks a leading space. Hover or tap a word to see its nearest neighbours — they are measured across all 128 numbers, not on this flat picture, so a line can reach right across the map.

try king · love · lord · he
hover or tap a word

What you're seeing is real, and honestly limited. Pronouns drift together — open ·he and its closest are He, ·we, ·she. Titles keep company too: ·lord sits beside ·sir and ·king, because they turn up in the same slots. But 2D is a shadow of 128 dimensions, and Shakespeare is small — so this is contextual similarity, not full meaning. Models built specifically for word vectors (word2vec, GloVe), trained on far more text, sharpen this until analogies like king − man + woman ≈ queen almost line up — though even that famous result is more fragile and cherry-picked than it's usually shown.