consteval auto greet() const noexcept -> void {

Karl Keshavarzi

Computer Engineering @ University of Waterloo

I enjoy writing low-level systems and software, especially with C++. To foster a community that surrounds my passion, I founded the UWHPC design team. My research experience focuses on distributed GPU algorithms for maximal independent sets. I'm particularly interested in HFT, GPU programming, and ML infrastructure. Open to discussing new opportunities and collaborations. Feel free to reach out.

}

0x0000 Experience

Software Engineering Intern

May 2026 — Aug 2026

Shopify

Founder, HPC Engineer

Jan 2026 — Present

UWHPC, Sedra Design Team

Undergraduate Research Assistant

Sep 2026 — Present

University of Waterloo, ECE

Undergraduate Research Assistant

Sep 2026 — Present

University of Waterloo, IQOL

0x0020 Projects

xpu

host ≡ device

A header-only C++23 library that lets the same allocation, layout, and math code compile for either CPU or CUDA. Forget about #ifdef wrapping.

C++23CUDAUWHPC
Source

Variational Monte Carlo

E[Ψ] = ⟨Ψ|H|Ψ⟩

Stochastically computes ground-state energies no analytic solution can reach. Halving the error costs four times the samples, so accuracy is a compute budget.

C++23CUDAPythonUWHPC
Source

Heat Solver

∂T/∂t = α∇²T

Hits 91% of theoretical memory bandwidth on GPU, 68× faster than the multithreaded, vectorized CPU at 134M cells. Validated with analytical solutions.

C++20CUDA
Paper Source

Maxwell Solver

∇×E = −∂B/∂t

11× faster with CUDA than the OpenMP path on 28M-cell grids. Cache-aligned SoA and SIMD cut CPU kernel time 44% over the AoS baseline.

C++17CUDAPython
Source

N-Body Gravity Engine

F = Gm₁m₂/r²

Barnes-Hut gives 53× over a direct OpenMP kernel at 131K bodies. A 4th-order Yoshida integrator holds energy error to 2e-12 across 249 simulated years.

C++17OpenMPPythonMatplotlib
Paper Source

0x0048 Writing