DeepSeek, a Chinese AI lab, recently announced it would open-source parts of its technology as part of an event called “Open Source Week.” The company shared plans to release five code repositories, all of which have already been tested and used in real-world AI applications. Today, the event kicked off with the launch of FlashMLA, a highly optimized Multi-head Latent Attention (MLA) decoding kernel built for NVIDIA Hopper GPUs, including the H100. This can help make large-scale AI models run faster and more efficiently without wasting computing power.
DeepSeek FlashMLA – The Efficient MLA Decoding Kernel for Hopper GPUs
FlashMLA is designed to provide efficient decoding for variable-length sequences, which is a common requirement in many machine-learning applications. Training and running AI requires a ton of computing power, and this tool is built to squeeze the most performance out of Hopper GPUs, cutting down processing time while keeping things smooth.
Key Features of FlashMLA
The key highlights of FlashMLA include:
1. BF16 Support
FlashMLA takes advantage of Hopper GPUs’ special BF16 (Brain Floating-Point) format, helping AI models run faster while using less memory.
2. Paged KV Cache
It uses a paged key-value (KV) cache with a 64-block size, making memory access quicker and more efficient. This leads to exceptional memory bandwidth performance, which helps AI models process data without slowing down.
3. Remarkable Performance Gains
In testing, FlashMLA hit speeds of 3,000 GB/s for memory-heavy tasks and 580 TFLOPS (trillions of calculations per second) for complex computations on the H800 SXM5 GPU using CUDA 12.6.
Installation and Setup of FlashMLA
Installing FlashMLA is straightforward, provided that users meet the necessary hardware requirements. The kernel requires an NVIDIA Hopper GPU, specifically the H100, as well as CUDA version 12.3 or above and PyTorch version 2.0 or above. Users can install the kernel by executing the installation script included in the repository. The simplicity of the setup process allows developers to quickly integrate FlashMLA into their machine-learning workflows.
FlashMLA Real-World Performance Testing
To evaluate the real-world performance of FlashMLA, the YouTuber Bijan Bowen conducted a series of tests using an H100 GPU. In his assessment, he compared the kernel against existing solutions, such as FlashAttention 2.
1. Benchmarking Results
During the benchmarking process, FlashMLA demonstrated remarkable performance metrics. The kernel achieved a maximum speed of 2,975 GB/s in memory-bound configurations, with compute-bound scenarios reaching 532 TFLOPS. These figures illustrate the kernel’s capability to manage GPU resources to accelerate machine learning tasks efficiently.
2. FlashMLA vs. FlashAttention 2
Bowen also checked how FlashMLA compared to FlashAttention 2, and FlashMLA came out on top. It used GPU resources more efficiently, which means it wasted less time and energy processing data. When he tested FlashAttention 2 on a consumer-level RTX 4060 GPU, it didn’t even come close to the performance of FlashMLA on the H100. This comparison proved that it is better suited for high-end GPUs that handle massive AI workloads.
FlashMLA’s Impact on AI Acceleration
The implications of this MLU decoding kernel extend beyond its immediate performance metrics. As a pioneering project in the open-source domain, it sets a precedent for future developments in AI acceleration. The kernel’s design and efficiency optimizations could inspire further enhancements in consumer GPUs. This might make high-end machine learning more accessible, bringing powerful AI tools to more people, not just those with expensive hardware.
More to Come From DeepSeek
FlashMLA is just the beginning of DeepSeek AI. With four more releases still to come, the AI company is making advanced AI tools more accessible to developers everywhere. By sharing its work, the company is helping others build smarter, faster AI systems without starting from scratch. Right now, this MLU decoding kernel is setting the stage for a big leap in GPU performance and machine learning. Stay tuned—there’s more on the way.
| Latest From Us
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure








One Response
Keep up the great work! Thank you so much for sharing a great post.