SNAPKITTYWEST's picture
October 2026 main drop: mirror from GitHub
a8baeed verified
|
Raw History Blame Contribute Delete
13.1 kB
# Bit-Accelerator Datapath Architecture
## 1. High-Level Datapath Diagram
```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ BIT ACCELERATOR DATAPATH β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Stage 0: DECODE & OPERAND FETCH
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Opcode │──────────► Opcode Decoder (6-to-12 decoder)
β”‚ Operands │──────────► Register File (3R1W)
β”‚ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
Stage 1: ADDRESS GENERATION
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Adder (Rs1 + Rs2) ◄─── 64-bit Carry β”‚
β”‚ eff_bit_address = base + offset β”‚
β”‚ β”‚
β”‚ Output: 64-bit effective bit address β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
Stage 2: WORD & BIT INDEX CALCULATION
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Divider / Shifter (Γ·64 & mod 64) β”‚
β”‚ word_index = eff_addr >> 6 β”‚ Combinational or pipelined
β”‚ bit_index = eff_addr & 0x3F β”‚
β”‚ β”‚
β”‚ Output: (word_index[57:0], bit_idx) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
Stage 3: MEMORY ACCESS & TLB LOOKUP
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ TLB Lookup (word_index β†’ phys_address) β”‚
β”‚ L1 Cache Tag/Index Match β”‚
β”‚ Memory Request to Load/Store Queue β”‚
β”‚ β”‚
β”‚ Output: 64-bit word from L1/L2/Memory β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
Stage 4: BIT EXTRACTION
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Barrel Shifter β”‚
β”‚ (word >> bit_index) β”‚
β”‚ β”‚
β”‚ AND Gate (extract LSB) β”‚
β”‚ result = shifted_word & 1 β”‚
β”‚ β”‚
β”‚ Output: 1-bit result β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚
β–Ό
Stage 5: WRITEBACK
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Register File Write β”‚
β”‚ result ──► Rd β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
---
## 2. Component Specifications
### 2.1 Address Generator (Stage 1)
**Component**: 64-bit Ripple-Carry Adder or Kogge-Stone Parallel-Prefix Adder
**Inputs**:
- `base_address[63:0]` from Rs1 register
- `bit_offset[63:0]` from Rs2 register
- `carry_in` = 0
**Outputs**:
- `effective_bit_address[63:0]`
**Latency**: 1 cycle (Kogge-Stone) or 2 cycles (simpler adder)
**Control Signals**:
- `enable_add`: 1 if instruction requires address computation
- `select_operand`: 0 = Rs1+Rs2, 1 = Rs3+Rs4 (for multi-register ops)
---
### 2.2 Word and Bit Index Calculator (Stage 2)
**Component**: Fixed-Point Divider (Γ· 64) and Modulo-64 Unit
**Inputs**:
- `effective_bit_address[63:0]` from Address Generator
**Outputs**:
- `word_index[57:0]` = `effective_bit_address >> 6`
- `bit_index[5:0]` = `effective_bit_address[5:0]`
**Hardware Implementation**:
```
word_index[i] = effective_bit_address[i+6], for i = 0 to 57
bit_index[i] = effective_bit_address[i], for i = 0 to 5
```
(This is purely combinational wiring - no logic!)
**Latency**: 0 cycles (combinational)
---
### 2.3 TLB & Cache Lookup (Stage 3)
**Component**: TLB (Translation Lookaside Buffer) + L1 Cache Controller
**Inputs**:
- `word_index[57:0]` (virtual word address, 6-bit aligned = 6 bits from effective address)
- Control: `is_load`, `is_store`
**Outputs**:
- `cache_hit` (1-bit flag)
- `data_word[63:0]` (from L1 cache on hit)
- `phys_address[40:0]` (for miss β†’ L2/Memory)
**Behavior**:
**On Hit** (L1 cache, TLB translate):
- Return data directly from cache (fastest path)
- Latency: 1 cycle
**On Miss** (TLB or L1 cache miss):
- Initiate memory request to L2
- Stall pipeline until data returns
- Latency: variable (10+ cycles to L2, 50+ to main memory)
**Special Case: Page Boundary**:
- If `bit_address` and `(bit_address + 63)` span different pages:
- Hardware may need two memory accesses
- OR: Load both words and merge results (dual-load architecture)
---
### 2.4 Barrel Shifter (Stage 4)
**Component**: 64-bit Logarithmic Shifter
**Inputs**:
- `word[63:0]` from memory/cache
- `bit_index[5:0]` from Index Calculator
- `shift_direction` (0=right, 1=left, from decoder)
- `rotate_mode` (0=shift, 1=rotate)
**Implementation** (Rotating 64-bit barrel shifter):
```
Level 0: Shift by 32: word[i] ← (shift_by_32 ? word[(i+32)%64] : word[i])
Level 1: Shift by 16: word[i] ← (shift_by_16 ? word[(i+16)%64] : word[i])
Level 2: Shift by 8: word[i] ← (shift_by_8 ? word[(i+8)%64] : word[i])
Level 3: Shift by 4: word[i] ← (shift_by_4 ? word[(i+4)%64] : word[i])
Level 4: Shift by 2: word[i] ← (shift_by_2 ? word[(i+2)%64] : word[i])
Level 5: Shift by 1: word[i] ← (shift_by_1 ? word[(i+1)%64] : word[i])
bit_index[5:0] is decoded into 6 control signals: {shift_by_32, ..., shift_by_1}
```
**Outputs**:
- `shifted_word[63:0]`
**Latency**: 1 cycle (combinational through 6 levels of muxes)
---
### 2.5 Bit Extractor (Stage 4, after shifter)
**Component**: AND gate + optional sign-extension logic
**Inputs**:
- `shifted_word[63:0]` from Barrel Shifter
- `mask[63:0]` (computed from field width)
- Single bit: `mask = 64'h0000_0000_0000_0001`
- Multi-bit (width w): `mask = (1 << w) - 1`
**Outputs**:
- `extracted_field[63:0]` = `shifted_word & mask`
**For BITLOAD** (single-bit extract):
- Output: `result[63:0]` with `result[0] = extracted_field[0]`, `result[63:1] = 0`
**For BITFIELD** (multi-bit extract):
- Output: `result[63:0]` = extracted field, zero-extended
**Latency**: 1 cycle (combinational AND + logic)
---
### 2.6 Control Unit (Combinational + Sequential)
**Combinational Decoder**:
```
Opcode[5:0] β†’ 12 control signals:
- enable_addr_gen
- enable_mem_load
- enable_mem_store
- enable_shift
- enable_extract
- enable_writeback
- shift_direction (L/R)
- rotate_mode
- field_width[5:0]
- is_signed_extend
```
**Sequential State Machine** (for cache misses, exceptions):
```
State 0: IDLE β†’ await instruction
State 1: EXECUTE β†’ run datapath
State 2: MEMORY_WAIT β†’ stall on L1 miss
State 3: EXCEPTION β†’ handle fault
State 4: WRITEBACK β†’ commit result
```
---
### 2.7 Register File (3-Read, 1-Write Port)
**Capacity**: 32 Γ— 64-bit registers
**Read Ports**:
- Port A: Rs1 (base address)
- Port B: Rs2 (offset/secondary operand)
- Port C: Rs3 (value to write for BITSTORE)
**Write Port**:
- Port W: Rd (destination register)
**Access Time**: 1 cycle (combinational read, synchronous write)
---
## 3. Data Flow for BITLOAD Instruction
**Instruction**: `BITLOAD R5, R10, R15`
```
Clock | Stage | Operation
------|-------|-------------------------------------------
0 | 0 | Decode BITLOAD, fetch R10=0x1000, R15=0x42
| |
1 | 1 | Address generator: eff_addr = 0x1000 + 0x42 = 0x1042
| |
2 | 2 | Index calc (combinational):
| | word_idx = 0x1042 >> 6 = 0x41
| | bit_idx = 0x1042 & 0x3F = 0x02
| |
3 | 3 | TLB lookup: word_index β†’ phys_addr = 0x41000
| | L1 cache hit: word = 0xABCD_EF01_2345_6789
| |
4 | 4 | Shifter: shifted = word >> 2 = 0x2AF37_BC04_8D15_9E
| | Extractor: result = shifted & 1 = 0
| |
5 | 5 | Writeback: R5 ← 0
```
---
## 4. Pipeline Depth Analysis
**Best Case** (L1 hit):
- Stage 0: Decode (1 cycle)
- Stage 1: Address gen (1 cycle)
- Stage 2: Index calc (0 cycles, combinational)
- Stage 3: Memory access (1 cycle hit)
- Stage 4: Extract (1 cycle)
- Stage 5: Writeback (committed)
- **Total: 4 cycles latency**
**Worst Case** (L2 miss, then main memory):
- Stages 0-3: same as above
- Stage 3: Memory miss, initiate L2 request (10+ cycles)
- Stage 4: Extract (1 cycle)
- **Total: 15+ cycles latency**
---
## 5. Throughput Analysis
**Instruction Issue Rate**: 1 instruction per cycle (assuming no structural hazards)
**Bottlenecks**:
- Memory bandwidth: If high % of instructions are BITLOAD/BITSTORE
- Cache behavior: Working set fit in L1 cache is critical
- Register contention: Low (3 reads, 1 write per cycle is feasible)
**Optimization**: Dual-issue for independent bit operations:
- Issue BITLOAD + BITCOUNT in same cycle (different functional units)
---
## 6. Critical Path (for clock frequency)
The critical path is the longest combinational delay in the pipeline:
```
TDC = T_mux(operand select) + T_adder(64-bit) + T_setup(register)
β‰ˆ 0.3ns + 1.2ns + 0.2ns = 1.7ns
β‰ˆ 588 MHz (f_clock ≀ 1/1.7ns)
```
For 1+ GHz target, require pipelined adder (Kogge-Stone):
```
T_KS_adder β‰ˆ 0.8ns β†’ allows ~1.2 GHz
```
---
## 7. Physical Implementation Details
### 7.1 Gate Counts (Estimates)
| Component | Gates |
|-----------|-------|
| 64-bit Adder (Kogge-Stone) | 1,500 |
| Barrel Shifter (6 levels) | 4,000 |
| TLB (16-32 entries) | 5,000 |
| L1 Cache Interface | 10,000 |
| Register File (32Γ—64) | 20,000 |
| Control Decoder + FSM | 3,000 |
| Data Muxes, AND gates, misc | 5,000 |
| **Total estimate** | **~50k gates** |
### 7.2 Area Estimate
- Technology: 7nm FinFET
- Gate density: ~80M gates/mmΒ²
- **Area β‰ˆ 50k / 80M β‰ˆ 0.6 mmΒ²** (sub-mmΒ² functional unit)
### 7.3 Power Estimate
- Dynamic power: ~2mW (1 GHz, 1.2V, 50k gates, 80% switching activity)
- Static power: ~0.5mW (leakage)
- **Total: ~2.5mW** (idle or active, instruction-dependent)
---
## 8. Memory System Integration
### 8.1 Memory Hierarchy
```
CPU Pipeline
↓
L1 I-Cache (32 KB) ←── 1 cycle (instructions)
L1 D-Cache (32 KB) ←── 1 cycle (data for BITLOAD)
↓ (miss)
L2 Cache (256 KB) ←── 10 cycles
↓ (miss)
L3 Cache (8 MB) ←── 30 cycles
↓ (miss)
Main Memory (DDR4) ←── 50+ cycles
```
### 8.2 TLB Integration
```
Virtual Bit Address (64-bit)
↓
Extract word_index[57:0] = bit_address[63:6]
↓
TLB lookup (16-entry, 4-way associative)
β”œβ”€β–Ί Hit: phys_word_addr = phys_base[40:0] | word_index[13:0]
β”œβ”€β–Ί Miss: page walk (25+ cycles)
↓
L1 Cache tag lookup
```
### 8.3 Boundary Conditions
**Case 1**: Bit 0-63 fit in single 64-bit word
- Single memory access (normal path)
**Case 2**: Extract bits spanning two words (e.g., bits 60-65)
- Hardware loads both words
- Shifts, masks, combines
- Transparent to ISA
**Case 3**: Bits span two pages (e.g., word 0x1FFFF to 0x20000)
- May require two separate TLB lookups + memory accesses
- Hardware handles by stalling until both words available
- Or: Software constraint to avoid (less common)
---
## 9. Write-Path (BITSTORE)
```
BITSTORE Rs1, Rs2, Rs3
Stage 1: Address gen β†’ eff_addr = Rs1 + Rs2
Stage 2: Index calc β†’ word_idx, bit_idx
Stage 3: TLB lookup + load current word
Stage 4: Bit mask & merge:
mask = 1 << bit_idx
new_word = (old_word & ~mask) | ((Rs3 & 1) << bit_idx)
Stage 5: Store new_word back to memory (writeback cache)
```
**Read-Modify-Write**:
- Load old word: 3 cycles
- Compute new value: 1 cycle
- Store: pipelined (doesn't wait for commit)
- **Total latency: 4-5 cycles**
---
## 10. Register Transfer Language (RTL) Skeleton
```verilog
// Stage 1: Address Generation
reg [63:0] eff_addr;
always @(posedge clk)
eff_addr <= rs1_data + rs2_data;
// Stage 2: Index Calculation (combinational)
wire [57:0] word_index = eff_addr[63:6];
wire [5:0] bit_index = eff_addr[5:0];
// Stage 3: Memory Access
wire [63:0] memory_word = l1_cache_read(word_index);
// Stage 4: Bit Extraction
wire [63:0] shifted = memory_word >> bit_index;
wire [63:0] extracted = shifted & 64'h1;
// Stage 5: Writeback
always @(posedge clk)
rd_data <= extracted;
```