Tools & products

KernelBench

1 sources · 1 mentions

Views belong to a particular speaker and passage. Counts describe this reviewed collection, not product ratings or market consensus.

Mention only

Lilian Weng presents KernelBench as a benchmark of 250 PyTorch tasks assessing LLMs’ ability to generate correct and fast GPU kernels, using fast_p as its primary metric.

Lilian Weng ·

Supporting evidence

Harness Engineering for Self-Improvement

Original excerpt

KernelBench : evaluate correctness and speed for generated GPU kernels. 250 PyTorch tasks to evaluate whether LLM can write fast and correct kernels.
Context

The evaluation metric fast_p = the percentage of generated kernels that are correct and faster than baseline.

About this interpretation

The object-specific interpretation and Chinese translation were checked independently against the source. This is an AI semantic review, not playback verification. Reviewed Oct 5, 2026 · qwen3.8-max-0902

Report an issue
← All mentioned things