Sakana AI and NVIDIA have introduced TwELL, a new approach using CUDA kernels that achieves 20.5% inference and 21.9% tr...

Sakana AI and NVIDIA have introduced TwELL, a new approach using CUDA kernels that achieves 20.5% inference and 21.9% training speedup in large language models. The technique targets feedforward layers, which account for over two-thirds of model parameters and 80% of FLOPs, by inducing over 99% sparsity and translating it into real GPU throughput gains through new sparse data formats. https://www.marktechpost.com/2026/05/11/sakana-ai-and-nvidia-introduce-twell-with-cuda-kernels-for-20-5-inference-and-21-9-training-speedup-in-llms/ #AIagent #AI #GenAI #Infrastructure

Read Original

Related

Mastodon discussion 12m ago

プロンプトエンジニアリング技術の基礎と応用プロンプトエンジニアリング技術は、AIモデルの精度と効率を高めるために不可欠です。この記事では、プロンプトエンジニアリング技術の基礎と応用について詳しく解説します。https://ai-blog-s...

プロンプトエンジニアリング技術の基礎と応用プロンプトエンジニアリング技術は、AIモデルの精度と効率を高めるために不可欠です。この記事では、プロンプトエンジニアリング技術の基礎と応用について詳しく解説します。https://ai-blog-seven-wine.vercel.app/ja/posts/2026-08-05-am-x2xvp#プロンプトエンジニア...

Mastodon discussion 14m ago

この間公開のことでバルトさんがシグルドさんとケンカしてましたOpenAI、アップル訴訟に反論 弁護士の誤送信メールも公開 https://ascii.jp/elem/000/004/424/4424766/?rss#Apple #LLM #...

この間公開のことでバルトさんがシグルドさんとケンカしてましたOpenAI、アップル訴訟に反論 弁護士の誤送信メールも公開 https://ascii.jp/elem/000/004/424/4424766/?rss#Apple #LLM #news #bot