# Speed Test Guide: HRM SRAM/DRAM Benchmarking Use this guide to measure the **hardware performance** (latency, throughput, memory overhead) of the memory-tiered architecture. > [!NOTE] > This guide uses **synthetic data**. No external datasets or prior training are required. You can run this immediately after setup. --- ### 1. Environment Setup Create a clean virtual environment and install dependencies: ```bash python -m venv venv source venv/bin/activate pip install --upgrade pip pip install -r requirements.txt pip install triton matplotlib ``` --- ### 2. Run Hardware Comparison This command compares the **Baseline (Original)** model vs. the **Tiered (Optimized)** model across multiple configurations. ```bash python run_benchmark.py --mode compare --plot ``` **What this does:** 1. Generates random tensors to simulate a reasoning workload. 2. Compiles Triton kernels for the L-level (SRAM) and H-level (DRAM) paths. 3. Records precise GPU timing using `torch.cuda.Event`. 4. Saves charts to `benchmark_results/`. --- ### 3. Deep Dive into Tiered Metrics To get a detailed breakdown of memory tier behavior: ```bash python run_benchmark.py --mode tiered --iterations 50 --batch-sizes 8 --seq-lens 128 ``` **Key Metrics to watch:** - **L_lat(μs):** Speed of the fast-updating L-level module in SRAM. - **H_lat(μs):** Speed of the planning H-level module in DRAM. - **H/L Ratio:** Shows the latency multiplier between memory tiers. - **Memory Efficiency:** Percentage of time spent on math vs. memory transfer. --- ### 4. Custom Hardware Stress Test Test the limits of your GPU by increasing batch sizes or sequence lengths: ```bash python run_benchmark.py \ --mode compare \ --batch-sizes 64,128 \ --seq-lens 256,512 \ --hidden-size 1024 ```