metadata
language:
- en
- zh
license: mit
library_name: pytorch
tags:
- text-generation
- bilingual
- BPE
- transformer
- reasoning
- custom-architecture
- GQA
- flash-attention
datasets:
- eyanchao/echoic-data
Echoic
Bilingual (Chinese/English) language model β custom Transformer architecture trained from scratch by eyanchao.
No HuggingFace Transformers dependency. Pure PyTorch with SDPA, GQA, QK-Norm, Z-loss.
Current Model
| Version | Params | Dim | Layers | Heads | KV Heads | Tokenizer | Status |
|---|---|---|---|---|---|---|---|
| v30 | ~1.1B | 1408 | 32 | 32 | 8 (GQA) | BPE 32k | π Training |
| β | β Retired | ||||||
| 8 (GQA) | β Retired |
Architecture (v30)
| Technique | Description |
|---|---|
| BPE Tokenizer | SentencePiece 32k vocab β compresses Chinese 3:1 |
| GQA | Grouped Query Attention (32Q / 8KV heads) |
| QK-Norm | RMSNorm on Q/K projections |
| Z-loss | Logit L2 regularization |
| SDPA | PyTorch fused scaled_dot_product_attention |
| RoPE + SwiGLU + RMSNorm | Llama-style pre-norm architecture |
| seq_len 512 | Long context for CoT reasoning |
| Gradient checkpointing | Per-layer activation recompute |
| torch.compile | JIT with TF32 precision + inductor cache |
| Layer-wise LR decay | Higher layers learn slower (0.8Γ) |
| Dynamic LR | Auto-half on consecutive loss spikes |
Training
- Platform: Modal (serverless A100 40GB)
- Data: eyanchao/echoic-data β 18 files, multi-task mixture
- Math 40% | Code 20% | Encyclopedia 25% | Stories 15% | Identity SFT
- Checkpointing: Volume + HF Hub dual backup, auto-resume
Usage
import torch, sentencepiece as spm
from huggingface_hub import hf_hub_download
# Load BPE tokenizer
sp = spm.SentencePieceProcessor()
sp.load(hf_hub_download("eyanchao/echoic-lite", "bpe_tokenizer.model"))
# Model: EchoicLM_1B (dim=1408, layers=32, heads=32, kv_heads=8, seq_len=512)
# See echoic_v30_bpe.ipynb for full class definition
def generate(prompt, model, max_tokens=200, temp=0.8, top_k=50):
x = torch.tensor([sp.encode(prompt)], dtype=torch.long)
out = model.generate(x, max_new_tokens=max_tokens, temperature=temp, top_k=top_k)
return sp.decode(out[0].tolist())
Files
| File | Description |
|---|---|
bpe_tokenizer.model |
BPE SentencePiece model (32k vocab) |
bpe_tokenizer.vocab |
BPE vocabulary |
v30-best-e*.pt |
Training checkpoints |
v30-final.pt |
Final weights (after training) |
Evolution
v26 60M β v27 200M β v28 480M β v29 1B β v30 BPE 1B β ...
β current
v26-v29 were character-level (98 vocab). v30 switches to BPE 32k vocab with seq_len 512 for serious reasoning capability.