From CUDA to MLX
IBM Research extends K-Search with a CUDA-to-MLX translation layer, transferring expert kernel knowledge to Apple Silicon. This achieves 97% of FlashAttention performance and a ~20x faster Mamba SSM prefill. The K-Search framework leverages evolutionary search to optimize kernel performance.
- K-Search transfers kernel expertise from CUDA to MLX
- Achieves 97% of FlashAttention performance on Apple Silicon
- Mamba SSM prefill is ~20x faster with K-Search optimization
