Download bit_accelerator/docs/DATAPATH.md from Snapkitty/rust-opencl-gpu: direct link, hf CLI and curl.
- Browser
- Download file 13.1 kB
-
https://huggingface.co/Snapkitty/rust-opencl-gpu/resolve/main/bit_accelerator/docs/DATAPATH.md
- Command line
-
hf download hf://Snapkitty/rust-opencl-gpu/bit_accelerator/docs/DATAPATH.md
-
curl -L -o DATAPATH.md https://huggingface.co/Snapkitty/rust-opencl-gpu/resolve/main/bit_accelerator/docs/DATAPATH.md
Bit-Accelerator Datapath Architecture
1. High-Level Datapath Diagram
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β BIT ACCELERATOR DATAPATH β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Stage 0: DECODE & OPERAND FETCH
ββββββββββββββββ
β Opcode ββββββββββββΊ Opcode Decoder (6-to-12 decoder)
β Operands ββββββββββββΊ Register File (3R1W)
β β
ββββββββββββββββ
β
βΌ
Stage 1: ADDRESS GENERATION
ββββββββββββββββββββββββββββββββββββββββββββ
β Adder (Rs1 + Rs2) ββββ 64-bit Carry β
β eff_bit_address = base + offset β
β β
β Output: 64-bit effective bit address β
ββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
Stage 2: WORD & BIT INDEX CALCULATION
ββββββββββββββββββββββββββββββββββββββββ
β Divider / Shifter (Γ·64 & mod 64) β
β word_index = eff_addr >> 6 β Combinational or pipelined
β bit_index = eff_addr & 0x3F β
β β
β Output: (word_index[57:0], bit_idx) β
ββββββββββββββββββββββββββββββββββββββββ
β
βΌ
Stage 3: MEMORY ACCESS & TLB LOOKUP
ββββββββββββββββββββββββββββββββββββββββββββ
β TLB Lookup (word_index β phys_address) β
β L1 Cache Tag/Index Match β
β Memory Request to Load/Store Queue β
β β
β Output: 64-bit word from L1/L2/Memory β
ββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
Stage 4: BIT EXTRACTION
ββββββββββββββββββββββββββββββββββββββββ
β Barrel Shifter β
β (word >> bit_index) β
β β
β AND Gate (extract LSB) β
β result = shifted_word & 1 β
β β
β Output: 1-bit result β
ββββββββββββββββββββββββββββββββββββββββ
β
βΌ
Stage 5: WRITEBACK
ββββββββββββββββββββββββββββββββββββββββ
β Register File Write β
β result βββΊ Rd β
ββββββββββββββββββββββββββββββββββββββββ
2. Component Specifications
2.1 Address Generator (Stage 1)
Component: 64-bit Ripple-Carry Adder or Kogge-Stone Parallel-Prefix Adder
Inputs:
base_address[63:0]from Rs1 registerbit_offset[63:0]from Rs2 registercarry_in= 0
Outputs:
effective_bit_address[63:0]
Latency: 1 cycle (Kogge-Stone) or 2 cycles (simpler adder)
Control Signals:
enable_add: 1 if instruction requires address computationselect_operand: 0 = Rs1+Rs2, 1 = Rs3+Rs4 (for multi-register ops)
2.2 Word and Bit Index Calculator (Stage 2)
Component: Fixed-Point Divider (Γ· 64) and Modulo-64 Unit
Inputs:
effective_bit_address[63:0]from Address Generator
Outputs:
word_index[57:0]=effective_bit_address >> 6bit_index[5:0]=effective_bit_address[5:0]
Hardware Implementation:
word_index[i] = effective_bit_address[i+6], for i = 0 to 57
bit_index[i] = effective_bit_address[i], for i = 0 to 5
(This is purely combinational wiring - no logic!)
Latency: 0 cycles (combinational)
2.3 TLB & Cache Lookup (Stage 3)
Component: TLB (Translation Lookaside Buffer) + L1 Cache Controller
Inputs:
word_index[57:0](virtual word address, 6-bit aligned = 6 bits from effective address)- Control:
is_load,is_store
Outputs:
cache_hit(1-bit flag)data_word[63:0](from L1 cache on hit)phys_address[40:0](for miss β L2/Memory)
Behavior:
On Hit (L1 cache, TLB translate):
- Return data directly from cache (fastest path)
- Latency: 1 cycle
On Miss (TLB or L1 cache miss):
- Initiate memory request to L2
- Stall pipeline until data returns
- Latency: variable (10+ cycles to L2, 50+ to main memory)
Special Case: Page Boundary:
- If
bit_addressand(bit_address + 63)span different pages:- Hardware may need two memory accesses
- OR: Load both words and merge results (dual-load architecture)
2.4 Barrel Shifter (Stage 4)
Component: 64-bit Logarithmic Shifter
Inputs:
word[63:0]from memory/cachebit_index[5:0]from Index Calculatorshift_direction(0=right, 1=left, from decoder)rotate_mode(0=shift, 1=rotate)
Implementation (Rotating 64-bit barrel shifter):
Level 0: Shift by 32: word[i] β (shift_by_32 ? word[(i+32)%64] : word[i])
Level 1: Shift by 16: word[i] β (shift_by_16 ? word[(i+16)%64] : word[i])
Level 2: Shift by 8: word[i] β (shift_by_8 ? word[(i+8)%64] : word[i])
Level 3: Shift by 4: word[i] β (shift_by_4 ? word[(i+4)%64] : word[i])
Level 4: Shift by 2: word[i] β (shift_by_2 ? word[(i+2)%64] : word[i])
Level 5: Shift by 1: word[i] β (shift_by_1 ? word[(i+1)%64] : word[i])
bit_index[5:0] is decoded into 6 control signals: {shift_by_32, ..., shift_by_1}
Outputs:
shifted_word[63:0]
Latency: 1 cycle (combinational through 6 levels of muxes)
2.5 Bit Extractor (Stage 4, after shifter)
Component: AND gate + optional sign-extension logic
Inputs:
shifted_word[63:0]from Barrel Shiftermask[63:0](computed from field width)- Single bit:
mask = 64'h0000_0000_0000_0001 - Multi-bit (width w):
mask = (1 << w) - 1
- Single bit:
Outputs:
extracted_field[63:0]=shifted_word & mask
For BITLOAD (single-bit extract):
- Output:
result[63:0]withresult[0] = extracted_field[0],result[63:1] = 0
For BITFIELD (multi-bit extract):
- Output:
result[63:0]= extracted field, zero-extended
Latency: 1 cycle (combinational AND + logic)
2.6 Control Unit (Combinational + Sequential)
Combinational Decoder:
Opcode[5:0] β 12 control signals:
- enable_addr_gen
- enable_mem_load
- enable_mem_store
- enable_shift
- enable_extract
- enable_writeback
- shift_direction (L/R)
- rotate_mode
- field_width[5:0]
- is_signed_extend
Sequential State Machine (for cache misses, exceptions):
State 0: IDLE β await instruction
State 1: EXECUTE β run datapath
State 2: MEMORY_WAIT β stall on L1 miss
State 3: EXCEPTION β handle fault
State 4: WRITEBACK β commit result
2.7 Register File (3-Read, 1-Write Port)
Capacity: 32 Γ 64-bit registers
Read Ports:
- Port A: Rs1 (base address)
- Port B: Rs2 (offset/secondary operand)
- Port C: Rs3 (value to write for BITSTORE)
Write Port:
- Port W: Rd (destination register)
Access Time: 1 cycle (combinational read, synchronous write)
3. Data Flow for BITLOAD Instruction
Instruction: BITLOAD R5, R10, R15
Clock | Stage | Operation
------|-------|-------------------------------------------
0 | 0 | Decode BITLOAD, fetch R10=0x1000, R15=0x42
| |
1 | 1 | Address generator: eff_addr = 0x1000 + 0x42 = 0x1042
| |
2 | 2 | Index calc (combinational):
| | word_idx = 0x1042 >> 6 = 0x41
| | bit_idx = 0x1042 & 0x3F = 0x02
| |
3 | 3 | TLB lookup: word_index β phys_addr = 0x41000
| | L1 cache hit: word = 0xABCD_EF01_2345_6789
| |
4 | 4 | Shifter: shifted = word >> 2 = 0x2AF37_BC04_8D15_9E
| | Extractor: result = shifted & 1 = 0
| |
5 | 5 | Writeback: R5 β 0
4. Pipeline Depth Analysis
Best Case (L1 hit):
- Stage 0: Decode (1 cycle)
- Stage 1: Address gen (1 cycle)
- Stage 2: Index calc (0 cycles, combinational)
- Stage 3: Memory access (1 cycle hit)
- Stage 4: Extract (1 cycle)
- Stage 5: Writeback (committed)
- Total: 4 cycles latency
Worst Case (L2 miss, then main memory):
- Stages 0-3: same as above
- Stage 3: Memory miss, initiate L2 request (10+ cycles)
- Stage 4: Extract (1 cycle)
- Total: 15+ cycles latency
5. Throughput Analysis
Instruction Issue Rate: 1 instruction per cycle (assuming no structural hazards)
Bottlenecks:
- Memory bandwidth: If high % of instructions are BITLOAD/BITSTORE
- Cache behavior: Working set fit in L1 cache is critical
- Register contention: Low (3 reads, 1 write per cycle is feasible)
Optimization: Dual-issue for independent bit operations:
- Issue BITLOAD + BITCOUNT in same cycle (different functional units)
6. Critical Path (for clock frequency)
The critical path is the longest combinational delay in the pipeline:
TDC = T_mux(operand select) + T_adder(64-bit) + T_setup(register)
β 0.3ns + 1.2ns + 0.2ns = 1.7ns
β 588 MHz (f_clock β€ 1/1.7ns)
For 1+ GHz target, require pipelined adder (Kogge-Stone):
T_KS_adder β 0.8ns β allows ~1.2 GHz
7. Physical Implementation Details
7.1 Gate Counts (Estimates)
| Component | Gates |
|---|---|
| 64-bit Adder (Kogge-Stone) | 1,500 |
| Barrel Shifter (6 levels) | 4,000 |
| TLB (16-32 entries) | 5,000 |
| L1 Cache Interface | 10,000 |
| Register File (32Γ64) | 20,000 |
| Control Decoder + FSM | 3,000 |
| Data Muxes, AND gates, misc | 5,000 |
| Total estimate | ~50k gates |
7.2 Area Estimate
- Technology: 7nm FinFET
- Gate density: ~80M gates/mmΒ²
- Area β 50k / 80M β 0.6 mmΒ² (sub-mmΒ² functional unit)
7.3 Power Estimate
- Dynamic power: ~2mW (1 GHz, 1.2V, 50k gates, 80% switching activity)
- Static power: ~0.5mW (leakage)
- Total: ~2.5mW (idle or active, instruction-dependent)
8. Memory System Integration
8.1 Memory Hierarchy
CPU Pipeline
β
L1 I-Cache (32 KB) βββ 1 cycle (instructions)
L1 D-Cache (32 KB) βββ 1 cycle (data for BITLOAD)
β (miss)
L2 Cache (256 KB) βββ 10 cycles
β (miss)
L3 Cache (8 MB) βββ 30 cycles
β (miss)
Main Memory (DDR4) βββ 50+ cycles
8.2 TLB Integration
Virtual Bit Address (64-bit)
β
Extract word_index[57:0] = bit_address[63:6]
β
TLB lookup (16-entry, 4-way associative)
βββΊ Hit: phys_word_addr = phys_base[40:0] | word_index[13:0]
βββΊ Miss: page walk (25+ cycles)
β
L1 Cache tag lookup
8.3 Boundary Conditions
Case 1: Bit 0-63 fit in single 64-bit word
- Single memory access (normal path)
Case 2: Extract bits spanning two words (e.g., bits 60-65)
- Hardware loads both words
- Shifts, masks, combines
- Transparent to ISA
Case 3: Bits span two pages (e.g., word 0x1FFFF to 0x20000)
- May require two separate TLB lookups + memory accesses
- Hardware handles by stalling until both words available
- Or: Software constraint to avoid (less common)
9. Write-Path (BITSTORE)
BITSTORE Rs1, Rs2, Rs3
Stage 1: Address gen β eff_addr = Rs1 + Rs2
Stage 2: Index calc β word_idx, bit_idx
Stage 3: TLB lookup + load current word
Stage 4: Bit mask & merge:
mask = 1 << bit_idx
new_word = (old_word & ~mask) | ((Rs3 & 1) << bit_idx)
Stage 5: Store new_word back to memory (writeback cache)
Read-Modify-Write:
- Load old word: 3 cycles
- Compute new value: 1 cycle
- Store: pipelined (doesn't wait for commit)
- Total latency: 4-5 cycles
10. Register Transfer Language (RTL) Skeleton
// Stage 1: Address Generation
reg [63:0] eff_addr;
always @(posedge clk)
eff_addr <= rs1_data + rs2_data;
// Stage 2: Index Calculation (combinational)
wire [57:0] word_index = eff_addr[63:6];
wire [5:0] bit_index = eff_addr[5:0];
// Stage 3: Memory Access
wire [63:0] memory_word = l1_cache_read(word_index);
// Stage 4: Bit Extraction
wire [63:0] shifted = memory_word >> bit_index;
wire [63:0] extracted = shifted & 64'h1;
// Stage 5: Writeback
always @(posedge clk)
rd_data <= extracted;