← All projects
/04 SYSTEMS 2025 Systems · 2025

CUDA Kernels — transformers, close to the metal.

End-to-end CUDA implementation of scaled dot-product attention and FFN layers on NVIDIA A100. Profiled compute- vs memory-bound regimes against cuBLAS and PyTorch native ops.

Hardware
A100
80GB
Attention impl.
3
Naive / Tiled / Fused
Speedup vs PT
1.4×
Long-context
Peak FLOPs
42
% of A100 theoretical
§ 01
PROBLEM

What it had to solve.

PyTorch's default attention is easy to reach for but leaves performance on the table — especially at long context, where the quadratic memory access dominates latency.

§ 02
APPROACH

How it works.

Three kernel variants were implemented: a naive baseline, a tiled shared-memory version, and a fused softmax-matmul. Each was profiled with Nsight Compute to characterize the compute-vs-memory split across sequence lengths.

§ 03
RESULTS

What it delivered.

Fused kernel hit 42% of A100 peak FLOPs at 4K context, outperforming PyTorch native by 1.4× and landing within 15% of cuBLAS for representative workloads.