writing
Two threads. A GPT in pure Python builds a working transformer out of lists and loops, then adds back the machinery real inference needs — batching, ragged batching, a KV cache, sampling — and proves at each step that the faster version still agrees with the slow one. Evaluation is the day job: benchmarks, judge models, assertion design, and what failure analysis actually looks like at scale.
There is an RSS feed.
No matching items