From CUDA to MLX
IBM Research has extended K-Search, a kernel search framework, to transfer expert kernel knowledge from CUDA to Apple Silicon's MLX, achieving 97% of FlashAttention performance. This breakthrough enables faster kernel optimization on Apple Silicon. The K-Search framework utilizes evolutionary search to optimize kernel performance.
- K-Search transfers kernel expertise from CUDA to MLX
- Achieves 97% of FlashAttention performance on Apple Silicon
- Enables ~20x faster Mamba SSM prefill
