2026Systems
CUDA Blackwell Labs
Author, public repository
CUDA Blackwell Labs is a 22-project plan I worked through on an NVIDIA DGX Spark, whose GB10 chip shares 128 GB of memory between the CPU and the GPU. It runs from a hardware probe and memory microbenchmarks, through the CUDA-to-PTX-to-SASS pipeline, streams and CUDA Graphs, to hand-written FP16 and FP4 tensor-core instructions, TMA tile copies and a FlashAttention-style softmax. Sustained reads reached 231 GB/s, 85% of the 273 GB/s peak. Bandwidth fell from about 900 GB/s to DRAM speed once the working set outgrew the 24 MB L2 cache. cuBLAS ran a 4096 × 4096 GEMM at about 91 TFLOP/s in FP16, against 18 TFLOP/s in FP32.
