xpu
host ≡ deviceA header-only C++23 library that lets the same allocation, layout, and math code compile for either CPU or CUDA. Forget about #ifdef wrapping.
Computer Engineering @ University of Waterloo
I enjoy writing low-level systems and software, especially with C++. To foster a community that surrounds my passion, I founded the UWHPC design team. My research experience focuses on distributed GPU algorithms for maximal independent sets. I'm particularly interested in HFT, GPU programming, and ML infrastructure. Open to discussing new opportunities and collaborations. Feel free to reach out.
0x0000 Experience
Shopify
UWHPC, Sedra Design Team
University of Waterloo, ECE
University of Waterloo, IQOL
0x0020 Projects
A header-only C++23 library that lets the same allocation, layout, and math code compile for either CPU or CUDA. Forget about #ifdef wrapping.
Stochastically computes ground-state energies no analytic solution can reach. Halving the error costs four times the samples, so accuracy is a compute budget.
Hits 91% of theoretical memory bandwidth on GPU, 68× faster than the multithreaded, vectorized CPU at 134M cells. Validated with analytical solutions.
11× faster with CUDA than the OpenMP path on 28M-cell grids. Cache-aligned SoA and SIMD cut CPU kernel time 44% over the AoS baseline.
Barnes-Hut gives 53× over a direct OpenMP kernel at 131K bodies. A 4th-order Yoshida integrator holds energy error to 2e-12 across 249 simulated years.
0x0048 Writing