Vega v1

Vega v1 is a 1B-parameter GPT model trained from scratch. It was pretrained on 13B tokens using 8 Intel XPUs for 2 weeks, then SFT’d on 2B tokens on 8 XPUs for another 2 days.

It can have basic conversations and recall well known facts. Of course, it hallucinates very often.

Vega v1 is a baseline for my future adventures into LLM training; it is very useful, but it's fun to play with.

This repository contains the model checkpoint and tokenizer. The included vega_v1_inference.py entry point is the reference inference program.

Model architecture: vega_v1 Layers: 14 Hidden size: 2560 Vocabulary size: 65536

Quickstart

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained('oriyonay/vega-v1', trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained('oriyonay/vega-v1', trust_remote_code=True)

Because Vega v1 uses custom model code, trust_remote_code=True is required.

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support