NVIDIA has introduced a 4-bit pretraining methodology using NVFP4, validated on a 12 billion parameter hybrid Mamba-Transformer trained on 10 trillion tokens - the longest publicly documented 4-bit pretraining run. Accuracy closely matches the FP8 baseline at 62.58% versus 62.62% on MMLU-Pro. https://www.marktechpost.com/2026/05/18/nvidia-introduces-a-4-bit-pretraining-methodology-using-nvfp4-validated-on-a-12b-hybrid-mamba-transformer-at-10t-token-horizon/ #AIagent #AI #GenAI #AIResearch
Related
South Korea: Massive margin calls on retail investors, the country becomes one big casino. That's the future. A generati...
South Korea: Massive margin calls on retail investors, the country becomes one big casino. That's the future. A generation of gambling addict, optimizing for nothing else than payi...
Would you rather have:• unlimited API creditsor• unlimited GPU hours? #AI #MachineLearning #LLM
Would you rather have:• unlimited API creditsor• unlimited GPU hours? #AI #MachineLearning #LLM
I don't work with chatbots, but this is a good article about how Airbnb used Eval Driven Development (EDD - kind of like...
I don't work with chatbots, but this is a good article about how Airbnb used Eval Driven Development (EDD - kind of like Test Driven Development (TDD)) to get a handle on testing n...