NVIDIA has released Gated DeltaNet-2, a linear attention layer that decouples erasing old content from writing new conte...

NVIDIA has released Gated DeltaNet-2, a linear attention layer that decouples erasing old content from writing new content via separate channel-wise gates. At 1.3B parameters trained on 100B tokens, it outperforms Mamba-2, Gated DeltaNet and KDA on language modelling and long-context retrieval. https://www.marktechpost.com/2026/05/24/nvidia-ai-releases-gated-deltanet-2-a-linear-attention-layer-that-decouples-erase-and-write-in-the-delta-rule/ #AIagent #AI #GenAI #ML

Read Original

Related