From CUDA to MLX
IBM Research extends K-Search with a CUDA-to-MLX translation layer, transferring expert kernel knowledge to Apple Silicon. This achieves 97% of FlashAttention performance and a ~20x faster Mamba SSM prefill. The K-Search framework brings decades of kernel expertise to Apple Silicon.
- K-Search framework extended with CUDA-to-MLX translation layer
- Achieves 97% of FlashAttention performance on Apple Silicon
- ~20x faster Mamba SSM prefill
