GPU systems, compilers, and mathematics. Based in Vancouver.
I'm building FindTensor, an experimental compiler and runtime for LLM inference.
- KernelIndex — GPU kernels indexed by operation, shape, dtype, and hardware.
- B200 kernels — CUDA C++ and CuTe DSL kernels for SOL-ExecBench.
- H100 serving estimator — GPU time per request across 91 published vLLM runs.
- SmolLM2 conformance — CPU checks for full-sequence and KV-cache inference.
- Tensor parallel reference — A CPU reference for process-isolated decoder execution.
- FlashInfer — My fork of FlashInfer.
Previously worked with Tor Aamodt at UBC on GPU architecture simulation.
Also working on formalizations in Lean.
FindTensor · LinkedIn · Email



