KernelBench measures correctness and speed of generated GPU kernels
KernelBench evaluates LLMs on 250 PyTorch tasks to assess their ability to generate correct and fast GPU kernels, using fast_p—the percentage of generated kernels that are both correct and faster than the baseline—as its metric.
Stützende Belege
Originalauszug
KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Kontext
The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.