mrciffa (@davideciffa)Luce PFlash가 prefix caching, cold-start 튜닝, CUDA VMM 수정과 block sparse attention autotune 개선으로 크게 빨라졌습니다. warm 상태는 약 10배, cold 상태는 약 2.5배 향상됐으며 Qwen3.6 27B에서 사용 가능하다고 밝혔습니다.https://x.com/davideciffa/status/2050906786232689033#caching #cuda #sparseattention #qwen #llm
Related
What happens when an AI joins a team as a teammate rather than a tool? A new study from researchers at UC Irvine and the...
What happens when an AI joins a team as a teammate rather than a tool? A new study from researchers at UC Irvine and the University of Tübingen explores that question.https://scott...
Stock futures climb amid hopes of U.S.-Iran deal; SpaceX falls 10%: Live updatesStocks rallied on Tuesday as the S&P 500...
Stock futures climb amid hopes of U.S.-Iran deal; SpaceX falls 10%: Live updatesStocks rallied on Tuesday as the S&P 500 and the Dow Industrials closed at new records.https://www.c...
This Plant Has Been Hiding a Killer Secret. Darwin Called It 150 Years AgoNew research appears to confirm that some Saxi...
This Plant Has Been Hiding a Killer Secret. Darwin Called It 150 Years AgoNew research appears to confirm that some Saxifraga plants do indeed munch on insects for subsistence.http...