For most developers, writing custom CUDA kernels is like handling nuclear waste: everyone knows it’s the secret to raw speed, but almost nobody wants to touch it. It’s a grueling process of manual tuning, memory alignment, and fighting with hardware constraints. But a new trend in 'auto-research' is changing the game, turning a week of expert engineering into an overnight AI task.
The Magic of Auto-Research
Recent breakthroughs, such as those highlighted by developer Sankalp and the autokernel project, demonstrate a shift from using LLMs as simple autocomplete tools to using them as autonomous optimization agents. By combining Codex or Claude with a verification loop, developers are achieving staggering results—including one instance of a 232x speedup over a baseline kernel for the qr_v2 problem.
The secret isn't just the AI's ability to write code, but its ability to iterate. By treating the LLM like a linear programming solver—giving it strict constraints, a clear goal, and a way to verify correctness—the AI can course-correct until it finds a solution that maximizes Tensor Core utilization.
From Triton to Megakernels
This isn't just about one-off wins. Tools like autokernel allow users to feed in a PyTorch model and wake up to optimized Triton kernels. We're seeing this ripple through the ecosystem, from Cursor's multi-agent systems achieving a 38% geomean speedup across 235 kernels on Blackwell GPUs, to the 'Lucebox Megakernel' pushing decode speeds to a buttery-smooth 413 tok/s on Qwen 3.5.
The Future of Performance
We are entering an era where the 'human-in-the-loop' is no longer the one writing the assembly-level logic, but the one defining the constraints. As these agents get better at navigating the intricacies of GPU architecture, the barrier between high-level PyTorch code and bare-metal performance is effectively disappearing.
Sources
Media



