Traditional Transformer path
- 01 / Tokenize — slice text into discrete sub-word vocabulary IDs.
- 02 / Embed — map IDs into high-dimensional vectors and retain them in accelerator memory.
- 03 / Attend — calculate pairwise self-attention across the sequence, producing O(N2) relationship work.
- 04 / Scale — grow memory and compute pressure as context length increases.