Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me

A side-by-side of all five Gemma-4 variants on Inferentia2 — PLE, KV-sharing, MatFormer, mixed attention, and a 128-expert MoE — the recipe that evolved to carry all of them, and the one bug that shows up in every single one.

Read Original

Related

Dev.to tutorial 5m ago

Outline Wiki 自架教學(三):Codex 串接 MCP

本篇要解決的問題 上一篇是 Outline + Claude,不一定每個人都有訂閱 Claude 方案,所以也提供 Codex 的方式。 跟著本篇一步步走,完成設定後,Codex...