Naive-N0.5-Flash: 48 layers, layers 0, 5, 11, 17, 23, 29, 35, 41 and 47 use DSA with 4 KV heads, key 192 and value 128 wide, top 2,048; the other 39 use a 128 token window with 8 KV heads. Index keys are 128 numbers in FP8.
GB means 10⁹ bytes. 90% of GPU memory counts as usable, the rest goes to the runtime. Weights = parameters × bytes per weight; real checkpoints differ by a few percent.
Linear attention state is stored in BF16. Sliding window layers keep exactly the window; some servers keep a little more.
Speed ceiling = total memory bandwidth ÷ bytes read per new word, for one person, with perfect scaling across GPUs. It ignores compute and communication, so real speed is lower. Quality at long context is not measured here.