0 XP

Attention · The cost of looking

Everything, all at once

Every word looks at every word

Attention lets a word pull meaning from the others around it. But a model doesn’t know in advance which words will matter, so it doesn’t guess — every single token looks at every single token, and the useful connections get high weights while the rest get low ones.

The attention patterns here are hand-authored illustrations of what trained models reliably do — not weights read out of a running model. the original paper