A side-by-side of all five Gemma-4 variants on Inferentia2 — PLE, KV-sharing, MatFormer, mixed attention, and a 128-expert MoE — the recipe that evolved to carry all of them, and the one bug that shows up in every single one.
Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me