BigHugger
sk Skill · fmh66

kernel-benchmark

Standalone kernel benchmarking skill for cuda-cpp, cutlass, cute-dsl, and triton implementations. Use when the user wants to compare a custom CUDA/CUTLASS .cu kernel or CuTe DSL/Triton .py kernel against selectable PyTorch eager, torch.compile, or FlashInfer baselines, validate correctness, measure execution time with KernelBench-style CUDA event timing, or generate benchmark.md for kernel optimization results.

installs 8w
0
30-day movement
starts with the next reading
Related entries
1
Connections
0
pythoncppbashPython
Host repository
fmh66/kernel-opt-agent
Host stars
16
Host language
Python