A single long-context prompt can stall every other user's token stream through head-of-line blocking. Prefill/decode disaggregation splits serving into two pools and cuts p99 inter-token latency. A tiny runnable Go simulation.
One Long Prompt Shouldn't Freeze Everyone's Tokens: Prefill/Decode Disaggregation