用于自我改进的 Harness 工程
KernelBench:评估所生成核函数的正确性与执行速度。包含250个PyTorch任务,用于评测大语言模型(LLM)能否编写出既快速又正确的核函数。
原始摘录
KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
上下文
评估指标 fast_p = 生成核函数中既正确、又快于基线版本的比例。
原始上下文
The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.