Instructions to use AethronPhantom/pyc-kernels with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use AethronPhantom/pyc-kernels with Kernels:
# !pip install kernels from kernels import get_kernel kernel = get_kernel("AethronPhantom/pyc-kernels") - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| tags: | |
| - cuda | |
| - gpu-optimization | |
| - kernels | |
| - gemm | |
| - hpc | |
| # PyC CUDA Kernels | |
| Twenty CUDA `.cu` prototype files from the PyC kernel lab. The collection | |
| covers baseline and experimental matmul, Ada FP32/Tensor Core variants, Hopper | |
| WMMA BF16/FP16 variants, asynchronous staging, CTA/warp/K-stage experiments, | |
| and cuBLASLt control lanes. | |
| These are source artifacts for inspection, compilation, and performance | |
| engineering—not precompiled model weights and not a claim of universal | |
| benchmark superiority. The 19 documented lab lanes are: | |
| | Kernel | Short description | Source | | |
| |---|---|---| | |
| | `matrix_mult` | Compile-only CUDA matmul baseline. | `kernels/prototypes/baseline/matmul/kernel.cu` | | |
| | `tokenizer_kernel` | Experimental tokenizer/matrix kernel. | `kernels/prototypes/experimental/tokenizer_matmul/kernel.cu` | | |
| | `ada_gemm` | Ada FP32 shared-memory GEMM. | `kernels/prototypes/ada/gemm/kernel.cu` | | |
| | `ada_gemm_k64_warp32_async` | Ada FP32 GEMM with `cp.async` double buffering. | `kernels/prototypes/ada/gemm_k64_warp32_async/kernel.cu` | | |
| | `ada_tensor_core_fp16` | Ada WMMA FP16 GEMM. | `kernels/prototypes/ada/tensor_core/kernel.cu` | | |
| | `ada_tensor_core_bf16` | Ada WMMA BF16 GEMM. | `kernels/prototypes/ada/tensor_core/kernel.cu` | | |
| | `hopper_tensor_core_fp16` | Hopper 64×64×16 WMMA FP16 GEMM. | `kernels/prototypes/hopper/tensor_core/kernel.cu` | | |
| | `hopper_tensor_core_bf16` | Hopper 64×64×16 WMMA BF16 GEMM. | `kernels/prototypes/hopper/tensor_core/kernel.cu` | | |
| | `hopper_tensor_core_bf16_perf` | Large-shape Hopper BF16 throughput lane. | `kernels/prototypes/hopper/tensor_core/kernel.cu` | | |
| | `hopper_tensor_core_bf16_warp2n` | Hopper lane with two N fragments per warp. | `kernels/prototypes/hopper/tensor_core/kernel.cu` | | |
| | `hopper_tensor_core_bf16_k32` | Hopper lane with 32-wide K staging. | `kernels/prototypes/hopper/tensor_core/kernel.cu` | | |
| | `hopper_tensor_core_bf16_warp2n_k32` | Combined warp2n and K32 experiment. | `kernels/prototypes/hopper/tensor_core/kernel.cu` | | |
| | `hopper_cublaslt_bf16` | cuBLASLt BF16 hardware-ceiling control. | `kernels/prototypes/hopper/cublaslt_bf16/kernel.cu` | | |
| | `hopper_cublaslt_bf16_check` | cuBLASLt BF16 correctness lane. | `kernels/prototypes/hopper/cublaslt_bf16/kernel.cu` | | |
| | `hopper_tensor_core_bf16_async` | Hopper BF16 WMMA with async double buffering. | `kernels/prototypes/hopper/tensor_core_async/kernel.cu` | | |
| | `hopper_tensor_core_bf16_async_wide` | Async WMMA with a wider 64×128 CTA. | `kernels/prototypes/hopper/tensor_core_async/kernel.cu` | | |
| | `hopper_tensor_core_bf16_async_wide_k64` | Wide async lane with K64 stages. | `kernels/prototypes/hopper/tensor_core_async/kernel.cu` | | |
| | `hopper_tensor_core_bf16_async_square` | Async WMMA with a 128×128 CTA. | `kernels/prototypes/hopper/tensor_core_async/kernel.cu` | | |
| | `hopper_tensor_core_bf16_async_square_k64` | Square async WMMA with 128×128×64 tiles. | `kernels/prototypes/hopper/tensor_core_async/kernel.cu` | | |
| The source files are organized by prototype family; several documented lanes | |
| share one parameterized `.cu` implementation. See the companion | |
| [performance Space](https://huggingface.co/spaces/AethronPhantom/pyc-kernel-lab) | |
| for the optimization story and selected measurements. | |