Phonon-2 Inference Model Pack (ipsilondev/Phonom2-fermion)
Stand-alone inference package and Git submodule containing weights, manifests, and offline tokenizer for Phonon-2 (2.1-bit ternary FastConformer 0.6B + TDT ASR).
Files Included
| File | Size | Purpose |
|---|---|---|
model.fermion |
169.2 MB | 2.1-bit quint5 packed model container (24-layer FastConformer + 2-layer LSTM decoder + joint) |
config.json |
277 KB | Architecture hyperparameters, mel filterbank settings, and model metadata |
packed_manifest.json |
821 B | Manifest indexing packed ternary tensor offsets, trit planes, and scale vectors |
bps_manifest.json |
799 B | Fermion package manifest |
tokenizer.json |
1.1 MB | Offline vocabulary and SentencePiece BPE tokenizer (8,193 tokens) |
tokenizer_config.json |
352 B | Tokenizer configuration |
processor_config.json |
408 B | Audio preprocessor configuration (128 mel bins, 16kHz) |
reference_transformers.py |
6.8 KB | Stand-alone PyTorch / Transformers model loader |
fermion_container.py |
3.9 KB | Stand-alone container decompressor and trit dequantizer |
Submodule Usage
Add this repository as a Git submodule in your downstream projects:
Via Hugging Face (Recommended for large model weights)
git submodule add https://huggingface.co/ipsilondev/Phonom2-fermion models/phonon2
git submodule update --init --recursive
Via GitHub
git submodule add https://github.com/ipsilondev/Phonom2-fermion.git models/phonon2
git submodule update --init --recursive
git lfs pull
Python Inference Quickstart
import torch
import reference_transformers as rt
# Load model using local container and tokenizer (completely offline)
model, processor, _ = rt.load_model(
container_path="models/phonon2/model.fermion",
base_dir="models/phonon2",
dtype=torch.float32
)
model.eval()
# Process audio
import soundfile as sf
audio, sr = sf.read("test.wav")
inputs = processor([audio], sampling_rate=16000, return_tensors="pt")
with torch.no_grad():
out = model.generate(input_features=inputs["input_features"])
text = processor.batch_decode(out.sequences, skip_special_tokens=True)[0]
print("Transcript:", text)
License
Model weights: CC-BY-4.0 (derived from NVIDIA Parakeet TDT 0.6B). Code & tools: Apache-2.0.
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support