You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Prism v2.0

A test model I made.

Character-level coding language model trained from scratch in PyTorch on a single consumer GPU. Built to learn how language models actually work โ€” not to compete with production code LMs.

Quick facts (v2)

  • Parameters: 147,133,154
  • Architecture: CharMLP with 12-layer residual MLP stack (GELU + LayerNorm), single self-attention layer
  • Context window: 384 characters
  • Vocab: 98 characters (printable ASCII + tab/newline/CR)
  • Precision: BF16 mixed
  • Training: ~6 hours on RTX 3070, final loss 1.04
  • Weights format: safetensors (no pickle, no code-exec risk on load)

v2 fixed the contaminated vocab from v1 (which had Chinese / Arabic / emoji chars in there and produced garbage output). v2 is strict ASCII-only.

Intended use

  • Learning how LMs are trained
  • Generating short code snippets (Python, JS, C/C++, Rust, Go)
  • Educational demonstrations

Out of scope

  • Production code generation
  • Commercial deployment
  • Safety-critical software
  • Anything requiring correct output

How to use

Prism is not an HF Transformers model. Load weights directly via PyTorch + safetensors. See inference.py for a working example.

import json, torch
from safetensors.torch import load_file

cfg  = json.loads(open("config.json").read())
vocab = json.loads(open("vocab.json").read())
state = load_file("model.safetensors")
# Build CharMLP(...) and load_state_dict(state). See inference.py.

Sampling

Recommended sampling: temperature 0.7, top-k 40, repetition penalty 1.1. Plain greedy or low-temperature sampling tends to get stuck in whitespace loops because the model's single-head attention has no positional encoding (a known architecture limitation).

Limitations

  • Tiny context window โ€” Prism forgets things said more than ~3 lines ago
  • Single small attention layer (no positional encoding) โ€” limited contextual reasoning
  • Character-level tokenizer โ€” slow on long inputs
  • Cold-start prompts often collapse into whitespace; prime with real code
  • English + code only; will not produce coherent natural language
  • No safety / RLHF training โ€” can produce nonsense or replicate bias from training data

License

MIT โ€” see LICENSE.

Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support