Safetensors
GGUF
English
gemma3_text
base
sml
void
pretrained-from-scratch
void.0 / README.md
appvoid's picture
Update README.md
7ed9b6d verified
|
Raw
History Blame Contribute Delete
1.43 kB
---
datasets:
- CEAMFA/rewrite
- HuggingFaceFW/fineweb-edu
- appvoid/no-prompt-15k
language:
- en
tags:
- base
- sml
- void
- pretrained-from-scratch
---
<style>
img {
display: block;
position: static;
width: min(76%, 256px);
height: auto;
max-width: 100%;
margin: 3rem auto 2.5rem;
object-fit: contain;
border: 2px solid rgba(255, 255, 255, 0.16);
border-radius: 1rem;
outline: none;
user-select: none;
-webkit-user-select: none;
-moz-user-select: none;
-webkit-user-drag: none;
filter: none !important;
transform: scale(1) !important;
animation: none !important;
box-shadow: none !important;
position: relative;
z-index: 1;
transform-origin: center;
}
</style>
<img src="https://huggingface.co/appvoid/void.0-preview/resolve/main/logo.png"/>
Introducing **void**: our first ever language model, trained from scratch with a novel hybrid tokenizer on 300M high-quality tokens (total of 2 epochs on a B300) using 4096 as context window. Total cost was $23 dollars. It has around 140m parameters. Future releases are expected to be published in the following weeks.
**Disclaimer:** Even though the model is based on gemma 3 architecture, the tokenizer is different so you might need to wait until this model can be added to llama.cpp
If you want to sponsor future model releases, you can get information on how to make contributions here: [CEAMFA](https://huggingface.co/CEAMFA)