An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode.
Related
LuthandoYekani/AWS-PartyRock-Projects: AWS PartyRock AI applications and case studies demonstrating generative AI, prompt engineering, and AI-assisted data analysis.
AWS PartyRock AI applications and case studies demonstrating generative AI, prompt engineering, and AI-assisted data analysis.
me-anurag/Fundamentals-of-Generative-AI-and-LLMs-NPTEL: Notes, slides, and assignment solutions for the NPTEL course Fundamentals of Generative AI and Large Language Models, IISc Bangalore. Covers VAEs, GANs, diffusion models, Transformers, and LLMs.
Notes, slides, and assignment solutions for the NPTEL course Fundamentals of Generative AI and Large Language Models, IISc Bangalore. Covers VAEs, GANs, diffusion models, Transform...