|
Download bit_accelerator/docs/DATAPATH.md from Snapkitty/rust-opencl-gpu: direct link, hf CLI and curl.
- Browser
- Download file 13.1 kB
-
https://huggingface.co/Snapkitty/rust-opencl-gpu/resolve/main/bit_accelerator/docs/DATAPATH.md
- Command line
-
hf download hf://Snapkitty/rust-opencl-gpu/bit_accelerator/docs/DATAPATH.md
-
curl -L -o DATAPATH.md https://huggingface.co/Snapkitty/rust-opencl-gpu/resolve/main/bit_accelerator/docs/DATAPATH.md
13.1 kB
| # Bit-Accelerator Datapath Architecture | |
| ## 1. High-Level Datapath Diagram | |
| ``` | |
| βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | |
| β BIT ACCELERATOR DATAPATH β | |
| βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ | |
| Stage 0: DECODE & OPERAND FETCH | |
| ββββββββββββββββ | |
| β Opcode ββββββββββββΊ Opcode Decoder (6-to-12 decoder) | |
| β Operands ββββββββββββΊ Register File (3R1W) | |
| β β | |
| ββββββββββββββββ | |
| β | |
| βΌ | |
| Stage 1: ADDRESS GENERATION | |
| ββββββββββββββββββββββββββββββββββββββββββββ | |
| β Adder (Rs1 + Rs2) ββββ 64-bit Carry β | |
| β eff_bit_address = base + offset β | |
| β β | |
| β Output: 64-bit effective bit address β | |
| ββββββββββββββββββββββββββββββββββββββββββββ | |
| β | |
| βΌ | |
| Stage 2: WORD & BIT INDEX CALCULATION | |
| ββββββββββββββββββββββββββββββββββββββββ | |
| β Divider / Shifter (Γ·64 & mod 64) β | |
| β word_index = eff_addr >> 6 β Combinational or pipelined | |
| β bit_index = eff_addr & 0x3F β | |
| β β | |
| β Output: (word_index[57:0], bit_idx) β | |
| ββββββββββββββββββββββββββββββββββββββββ | |
| β | |
| βΌ | |
| Stage 3: MEMORY ACCESS & TLB LOOKUP | |
| ββββββββββββββββββββββββββββββββββββββββββββ | |
| β TLB Lookup (word_index β phys_address) β | |
| β L1 Cache Tag/Index Match β | |
| β Memory Request to Load/Store Queue β | |
| β β | |
| β Output: 64-bit word from L1/L2/Memory β | |
| ββββββββββββββββββββββββββββββββββββββββββββ | |
| β | |
| βΌ | |
| Stage 4: BIT EXTRACTION | |
| ββββββββββββββββββββββββββββββββββββββββ | |
| β Barrel Shifter β | |
| β (word >> bit_index) β | |
| β β | |
| β AND Gate (extract LSB) β | |
| β result = shifted_word & 1 β | |
| β β | |
| β Output: 1-bit result β | |
| ββββββββββββββββββββββββββββββββββββββββ | |
| β | |
| βΌ | |
| Stage 5: WRITEBACK | |
| ββββββββββββββββββββββββββββββββββββββββ | |
| β Register File Write β | |
| β result βββΊ Rd β | |
| ββββββββββββββββββββββββββββββββββββββββ | |
| ``` | |
| --- | |
| ## 2. Component Specifications | |
| ### 2.1 Address Generator (Stage 1) | |
| **Component**: 64-bit Ripple-Carry Adder or Kogge-Stone Parallel-Prefix Adder | |
| **Inputs**: | |
| - `base_address[63:0]` from Rs1 register | |
| - `bit_offset[63:0]` from Rs2 register | |
| - `carry_in` = 0 | |
| **Outputs**: | |
| - `effective_bit_address[63:0]` | |
| **Latency**: 1 cycle (Kogge-Stone) or 2 cycles (simpler adder) | |
| **Control Signals**: | |
| - `enable_add`: 1 if instruction requires address computation | |
| - `select_operand`: 0 = Rs1+Rs2, 1 = Rs3+Rs4 (for multi-register ops) | |
| --- | |
| ### 2.2 Word and Bit Index Calculator (Stage 2) | |
| **Component**: Fixed-Point Divider (Γ· 64) and Modulo-64 Unit | |
| **Inputs**: | |
| - `effective_bit_address[63:0]` from Address Generator | |
| **Outputs**: | |
| - `word_index[57:0]` = `effective_bit_address >> 6` | |
| - `bit_index[5:0]` = `effective_bit_address[5:0]` | |
| **Hardware Implementation**: | |
| ``` | |
| word_index[i] = effective_bit_address[i+6], for i = 0 to 57 | |
| bit_index[i] = effective_bit_address[i], for i = 0 to 5 | |
| ``` | |
| (This is purely combinational wiring - no logic!) | |
| **Latency**: 0 cycles (combinational) | |
| --- | |
| ### 2.3 TLB & Cache Lookup (Stage 3) | |
| **Component**: TLB (Translation Lookaside Buffer) + L1 Cache Controller | |
| **Inputs**: | |
| - `word_index[57:0]` (virtual word address, 6-bit aligned = 6 bits from effective address) | |
| - Control: `is_load`, `is_store` | |
| **Outputs**: | |
| - `cache_hit` (1-bit flag) | |
| - `data_word[63:0]` (from L1 cache on hit) | |
| - `phys_address[40:0]` (for miss β L2/Memory) | |
| **Behavior**: | |
| **On Hit** (L1 cache, TLB translate): | |
| - Return data directly from cache (fastest path) | |
| - Latency: 1 cycle | |
| **On Miss** (TLB or L1 cache miss): | |
| - Initiate memory request to L2 | |
| - Stall pipeline until data returns | |
| - Latency: variable (10+ cycles to L2, 50+ to main memory) | |
| **Special Case: Page Boundary**: | |
| - If `bit_address` and `(bit_address + 63)` span different pages: | |
| - Hardware may need two memory accesses | |
| - OR: Load both words and merge results (dual-load architecture) | |
| --- | |
| ### 2.4 Barrel Shifter (Stage 4) | |
| **Component**: 64-bit Logarithmic Shifter | |
| **Inputs**: | |
| - `word[63:0]` from memory/cache | |
| - `bit_index[5:0]` from Index Calculator | |
| - `shift_direction` (0=right, 1=left, from decoder) | |
| - `rotate_mode` (0=shift, 1=rotate) | |
| **Implementation** (Rotating 64-bit barrel shifter): | |
| ``` | |
| Level 0: Shift by 32: word[i] β (shift_by_32 ? word[(i+32)%64] : word[i]) | |
| Level 1: Shift by 16: word[i] β (shift_by_16 ? word[(i+16)%64] : word[i]) | |
| Level 2: Shift by 8: word[i] β (shift_by_8 ? word[(i+8)%64] : word[i]) | |
| Level 3: Shift by 4: word[i] β (shift_by_4 ? word[(i+4)%64] : word[i]) | |
| Level 4: Shift by 2: word[i] β (shift_by_2 ? word[(i+2)%64] : word[i]) | |
| Level 5: Shift by 1: word[i] β (shift_by_1 ? word[(i+1)%64] : word[i]) | |
| bit_index[5:0] is decoded into 6 control signals: {shift_by_32, ..., shift_by_1} | |
| ``` | |
| **Outputs**: | |
| - `shifted_word[63:0]` | |
| **Latency**: 1 cycle (combinational through 6 levels of muxes) | |
| --- | |
| ### 2.5 Bit Extractor (Stage 4, after shifter) | |
| **Component**: AND gate + optional sign-extension logic | |
| **Inputs**: | |
| - `shifted_word[63:0]` from Barrel Shifter | |
| - `mask[63:0]` (computed from field width) | |
| - Single bit: `mask = 64'h0000_0000_0000_0001` | |
| - Multi-bit (width w): `mask = (1 << w) - 1` | |
| **Outputs**: | |
| - `extracted_field[63:0]` = `shifted_word & mask` | |
| **For BITLOAD** (single-bit extract): | |
| - Output: `result[63:0]` with `result[0] = extracted_field[0]`, `result[63:1] = 0` | |
| **For BITFIELD** (multi-bit extract): | |
| - Output: `result[63:0]` = extracted field, zero-extended | |
| **Latency**: 1 cycle (combinational AND + logic) | |
| --- | |
| ### 2.6 Control Unit (Combinational + Sequential) | |
| **Combinational Decoder**: | |
| ``` | |
| Opcode[5:0] β 12 control signals: | |
| - enable_addr_gen | |
| - enable_mem_load | |
| - enable_mem_store | |
| - enable_shift | |
| - enable_extract | |
| - enable_writeback | |
| - shift_direction (L/R) | |
| - rotate_mode | |
| - field_width[5:0] | |
| - is_signed_extend | |
| ``` | |
| **Sequential State Machine** (for cache misses, exceptions): | |
| ``` | |
| State 0: IDLE β await instruction | |
| State 1: EXECUTE β run datapath | |
| State 2: MEMORY_WAIT β stall on L1 miss | |
| State 3: EXCEPTION β handle fault | |
| State 4: WRITEBACK β commit result | |
| ``` | |
| --- | |
| ### 2.7 Register File (3-Read, 1-Write Port) | |
| **Capacity**: 32 Γ 64-bit registers | |
| **Read Ports**: | |
| - Port A: Rs1 (base address) | |
| - Port B: Rs2 (offset/secondary operand) | |
| - Port C: Rs3 (value to write for BITSTORE) | |
| **Write Port**: | |
| - Port W: Rd (destination register) | |
| **Access Time**: 1 cycle (combinational read, synchronous write) | |
| --- | |
| ## 3. Data Flow for BITLOAD Instruction | |
| **Instruction**: `BITLOAD R5, R10, R15` | |
| ``` | |
| Clock | Stage | Operation | |
| ------|-------|------------------------------------------- | |
| 0 | 0 | Decode BITLOAD, fetch R10=0x1000, R15=0x42 | |
| | | | |
| 1 | 1 | Address generator: eff_addr = 0x1000 + 0x42 = 0x1042 | |
| | | | |
| 2 | 2 | Index calc (combinational): | |
| | | word_idx = 0x1042 >> 6 = 0x41 | |
| | | bit_idx = 0x1042 & 0x3F = 0x02 | |
| | | | |
| 3 | 3 | TLB lookup: word_index β phys_addr = 0x41000 | |
| | | L1 cache hit: word = 0xABCD_EF01_2345_6789 | |
| | | | |
| 4 | 4 | Shifter: shifted = word >> 2 = 0x2AF37_BC04_8D15_9E | |
| | | Extractor: result = shifted & 1 = 0 | |
| | | | |
| 5 | 5 | Writeback: R5 β 0 | |
| ``` | |
| --- | |
| ## 4. Pipeline Depth Analysis | |
| **Best Case** (L1 hit): | |
| - Stage 0: Decode (1 cycle) | |
| - Stage 1: Address gen (1 cycle) | |
| - Stage 2: Index calc (0 cycles, combinational) | |
| - Stage 3: Memory access (1 cycle hit) | |
| - Stage 4: Extract (1 cycle) | |
| - Stage 5: Writeback (committed) | |
| - **Total: 4 cycles latency** | |
| **Worst Case** (L2 miss, then main memory): | |
| - Stages 0-3: same as above | |
| - Stage 3: Memory miss, initiate L2 request (10+ cycles) | |
| - Stage 4: Extract (1 cycle) | |
| - **Total: 15+ cycles latency** | |
| --- | |
| ## 5. Throughput Analysis | |
| **Instruction Issue Rate**: 1 instruction per cycle (assuming no structural hazards) | |
| **Bottlenecks**: | |
| - Memory bandwidth: If high % of instructions are BITLOAD/BITSTORE | |
| - Cache behavior: Working set fit in L1 cache is critical | |
| - Register contention: Low (3 reads, 1 write per cycle is feasible) | |
| **Optimization**: Dual-issue for independent bit operations: | |
| - Issue BITLOAD + BITCOUNT in same cycle (different functional units) | |
| --- | |
| ## 6. Critical Path (for clock frequency) | |
| The critical path is the longest combinational delay in the pipeline: | |
| ``` | |
| TDC = T_mux(operand select) + T_adder(64-bit) + T_setup(register) | |
| β 0.3ns + 1.2ns + 0.2ns = 1.7ns | |
| β 588 MHz (f_clock β€ 1/1.7ns) | |
| ``` | |
| For 1+ GHz target, require pipelined adder (Kogge-Stone): | |
| ``` | |
| T_KS_adder β 0.8ns β allows ~1.2 GHz | |
| ``` | |
| --- | |
| ## 7. Physical Implementation Details | |
| ### 7.1 Gate Counts (Estimates) | |
| | Component | Gates | | |
| |-----------|-------| | |
| | 64-bit Adder (Kogge-Stone) | 1,500 | | |
| | Barrel Shifter (6 levels) | 4,000 | | |
| | TLB (16-32 entries) | 5,000 | | |
| | L1 Cache Interface | 10,000 | | |
| | Register File (32Γ64) | 20,000 | | |
| | Control Decoder + FSM | 3,000 | | |
| | Data Muxes, AND gates, misc | 5,000 | | |
| | **Total estimate** | **~50k gates** | | |
| ### 7.2 Area Estimate | |
| - Technology: 7nm FinFET | |
| - Gate density: ~80M gates/mmΒ² | |
| - **Area β 50k / 80M β 0.6 mmΒ²** (sub-mmΒ² functional unit) | |
| ### 7.3 Power Estimate | |
| - Dynamic power: ~2mW (1 GHz, 1.2V, 50k gates, 80% switching activity) | |
| - Static power: ~0.5mW (leakage) | |
| - **Total: ~2.5mW** (idle or active, instruction-dependent) | |
| --- | |
| ## 8. Memory System Integration | |
| ### 8.1 Memory Hierarchy | |
| ``` | |
| CPU Pipeline | |
| β | |
| L1 I-Cache (32 KB) βββ 1 cycle (instructions) | |
| L1 D-Cache (32 KB) βββ 1 cycle (data for BITLOAD) | |
| β (miss) | |
| L2 Cache (256 KB) βββ 10 cycles | |
| β (miss) | |
| L3 Cache (8 MB) βββ 30 cycles | |
| β (miss) | |
| Main Memory (DDR4) βββ 50+ cycles | |
| ``` | |
| ### 8.2 TLB Integration | |
| ``` | |
| Virtual Bit Address (64-bit) | |
| β | |
| Extract word_index[57:0] = bit_address[63:6] | |
| β | |
| TLB lookup (16-entry, 4-way associative) | |
| βββΊ Hit: phys_word_addr = phys_base[40:0] | word_index[13:0] | |
| βββΊ Miss: page walk (25+ cycles) | |
| β | |
| L1 Cache tag lookup | |
| ``` | |
| ### 8.3 Boundary Conditions | |
| **Case 1**: Bit 0-63 fit in single 64-bit word | |
| - Single memory access (normal path) | |
| **Case 2**: Extract bits spanning two words (e.g., bits 60-65) | |
| - Hardware loads both words | |
| - Shifts, masks, combines | |
| - Transparent to ISA | |
| **Case 3**: Bits span two pages (e.g., word 0x1FFFF to 0x20000) | |
| - May require two separate TLB lookups + memory accesses | |
| - Hardware handles by stalling until both words available | |
| - Or: Software constraint to avoid (less common) | |
| --- | |
| ## 9. Write-Path (BITSTORE) | |
| ``` | |
| BITSTORE Rs1, Rs2, Rs3 | |
| Stage 1: Address gen β eff_addr = Rs1 + Rs2 | |
| Stage 2: Index calc β word_idx, bit_idx | |
| Stage 3: TLB lookup + load current word | |
| Stage 4: Bit mask & merge: | |
| mask = 1 << bit_idx | |
| new_word = (old_word & ~mask) | ((Rs3 & 1) << bit_idx) | |
| Stage 5: Store new_word back to memory (writeback cache) | |
| ``` | |
| **Read-Modify-Write**: | |
| - Load old word: 3 cycles | |
| - Compute new value: 1 cycle | |
| - Store: pipelined (doesn't wait for commit) | |
| - **Total latency: 4-5 cycles** | |
| --- | |
| ## 10. Register Transfer Language (RTL) Skeleton | |
| ```verilog | |
| // Stage 1: Address Generation | |
| reg [63:0] eff_addr; | |
| always @(posedge clk) | |
| eff_addr <= rs1_data + rs2_data; | |
| // Stage 2: Index Calculation (combinational) | |
| wire [57:0] word_index = eff_addr[63:6]; | |
| wire [5:0] bit_index = eff_addr[5:0]; | |
| // Stage 3: Memory Access | |
| wire [63:0] memory_word = l1_cache_read(word_index); | |
| // Stage 4: Bit Extraction | |
| wire [63:0] shifted = memory_word >> bit_index; | |
| wire [63:0] extracted = shifted & 64'h1; | |
| // Stage 5: Writeback | |
| always @(posedge clk) | |
| rd_data <= extracted; | |
| ``` | |