Skip to content

Repository files navigation

hexkernels

538 Hexagon NSP kernels that actually use the accelerator

Every entry carries HVX or HMX intrinsics, a scalar reference to check it against,
and a harness to run both. Plus the machinery that proves it.

License Python Release

kernels HMX speedup no deps

Quickstart · The library · Anti-cheat · Docs · Contributing


Important

A scalar C file is never a library entry here. It ships as reference.c, which is a different job. If the source has no vector or matrix intrinsic, it is not a kernel — and that rule is enforced by a test, not by taste.

Quickstart

pip install -e .          # no SDK, no hardware, no third-party dependencies
from hexkernels.library import find

find(hmx=True)                       # 73 kernels that reach the matrix engine
find(origin="expert", dtype="fp16")  # hand-written fp16, with measured speedups
find(tier="T2", buildable=True)      # complete bundles at one working-set tier
find(elf_confirmed=True)             # the disassembler confirms it, not just the source

k = find(hmx=True)[0]
k.source         # the accelerated kernel
k.reference      # scalar ground truth
k.speedup        # measured, against that reference
k.elf_confirmed  # True / False / None — None means never scanned, not "no"

Browsing and querying the corpus needs nothing but Python. You only need the Hexagon SDK to build, simulate, or disassemble.

The library

originnwhat it attests
expert350 Hand-written, each measured against its own scalar baseline.
Median 3.39× · p90 18.2× · best 124.7×. Ships 421 near-miss kernels.
mined171 Written against operators mined from the PyTorch registry.
⚠️ 126 carry a tier mismatch — recorded in tier_match, not hidden.
model17 Written by a language model. The survivors of 1,920 attempts.

73 reach HMX, the matrix engine. Spread across fp32 (168), int8 (96), fp16 (36), uint8 (29), int16 (18) and mixed-precision variants.

Every entry has the same shape, whatever its origin
kernels/<origin>/<name>/
├── kernel.c       the accelerated kernel        ← the library entry
├── reference.c    scalar ground truth
├── harness.c      correctness + cycle harness
├── kernel_api.h   entry-point declaration        (expert only)
├── nearmiss_*.c   plausible WRONG kernels the harness must reject
├── PROMPT.md      the ask that produced it
└── spec.json      normalised metadata

So a consumer never branches on origin. The entry point is always candidate_kernel. kernels/index.json is the whole corpus in one file; tools/assemble_library.py rebuilds it.

spec.json.bundle answers "can I build this?" in one field:

bundle n meaning
complete 467 reference and harness both present
harness-regenerable 71 harness too large to ship — HARNESS.md has the command

A harness embeds base64 golden vectors, so its size tracks the working set, not complexity. The largest is 93 MB; those 71 alone would be 675 MB of an otherwise 16 MB corpus. They regenerate in ~14 s per batch. Nothing irreproducible was dropped.

The three things worth reading

1. Anti-cheat — the reason this exists

A working scalar loop passes a correctness test and often beats a mediocre vectorised kernel on cycles.

So a score based on speed does not merely fail to reward accelerator use — it actively punishes it. Correctness and speed both fail as evidence. The only thing that settles it is the disassembled ELF.

This is not hypothetical. All 17 model-written kernels have HVX intrinsics in the source. Of the 12 that were scanned:

ELF verdict n
✅ confirms the mechanism 9
refutes it 3
⬜ never scanned 5

Writing the intrinsic is not the same as the binary containing it. → docs/ANTI-CHEAT.md

2. The agentic flow — hexkernels/forge/

PyTorch op → FX graph → Linalg IR → MLIR fusion → affine loops → scalar C
                                                                     │
                                           the QUESTION, not the answer
                                                                     ▼
                         a model reconstructs it as accelerated Hexagon C
                                                                     │
                             judged against a golden harness, then the ELF

The pipeline refuses to skip a stage, and stage (g) is the gate: if the emitted reference cannot pass the harness generated beside it, no prompt is written at all. hexkernels/gym/ closes the loop — a profiler names the bottleneck and the model is told which mechanism is missing. → docs/AGENTIC-FLOW.md

3. Simulator and silicon — core/, device/qdc/

core/ holds the toolchain wrapper, target detection, the simulator fleet and the reward. device/qdc/ runs the same kernels on real hardware — an adb-server port forward, not a shell, which bills a session for its full timeout if you don't release it.

verdict.py is vendored in full because it is what keeps a false pass off the money path: a job that ran zero tests once reported passing, and cycles_total=0 satisfied a check that only tested for the substring. → docs/QDC.md

Two ladders, both called T0–T3

Warning

T0T3 names two unrelated things. They have been confused for each other before. Always say which ladder you mean.

outcome ladder (per attempt) task tiers (per task, by size)
T0 doesn't compile fits L1D (≤16 KB) — hvx only
T1 compiles, wrong up to L2 (≤1 MB) — + l2fetch
T2 correct, but scalar inside VTCM — + dma + vtcm
T3 correct and uses the mechanism exceeds VTCM — must stream tiles

The mechanisms a task is entitled to are derived from its size — the corpus's central claim. → docs/TIERS.md

The study behind it

Can a model be got to use the accelerator at all? Three rungs, 1,920 attempts.

rung the ask reached a mechanism
0 bare, unaided 3 / 640
1 + a turn loop reporting failures 0 / 640
2 + the mechanism named explicitly 20 / 640 in source · 2 confirmed in ELF

Rung 1 drove correctness to 100% and mechanism use to zero. A feedback loop that optimises what you measure will optimise away what you didn't. That is only visible because the detector reads the binary. → docs/study/

Documentation

LIBRARY.md The corpus: schema, the three origins, what each attests
ANTI-CHEAT.md The detectors, and the used_vtcm asymmetry you must disclose
AGENTIC-FLOW.md The pipeline, stage by stage
QDC.md On-silicon measurement, and why completion isn't a verdict
TIERS.md The two ladders, and the entitlement gate
MEASUREMENT.md ⚠️ The traps. Read before quoting any cycle number

Caution

Never parallelise hexagon-sim. It starves the machine, orphans processes, and makes a contention timeout indistinguishable from a wrong answer — 352 verdicts had to be discarded and re-taken over exactly that.

Layout

kernels/              the library — one dir per kernel, one schema, + index.json
hexkernels/
├── library/          load and query the corpus
├── anticheat/        the ELF disassembly detectors
├── forge/            the agentic pipeline (+ frontend: trace, graph, emit, schedule)
├── gym/              profile-and-tune loop
├── core/             toolchain, target detection, simulator fleet, reward
└── device/qdc/       on-silicon measurement
benchmark/            the frozen corpus definition
env/harness/          harness headers every build includes
tools/                the script that assembles kernels/

Contributing

Kernels especially — but also detectors, docs, and anything that makes a claim here more falsifiable. Start with CONTRIBUTING.md; there are issue templates for contributing a kernel and for disputing a verdict.

A wrong verdict is the most serious kind of bug here — more serious than a crash. Correcting a result is always welcome, including this project's own.

License

Apache-2.0. See CODE_OF_CONDUCT.md for community expectations.

About

538 Hexagon NSP kernels that actually use the accelerator (HVX/HMX), plus the agentic pipeline, simulator/silicon measurement, and the ELF-disassembly anti-cheat detector that proves it

Resources

Code of conduct

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages