---
pipeline_tag: text-to-image
library_name: pytorch
license: apache-2.0
tags:
- solintelligence
- solpix
- flow-matching
- latent-image-generation
- final-release
---
# SolPix
SolPix turns a text prompt into a 512x512 image. Its roughly 49M-parameter flow transformer predicts image latents; Flan-T5 Base encodes the prompt, and the SANA 1.1 DC-AE decodes the result. Both external models stay frozen.
## Components
| Setting | Value |
|---|---|
| Generator | Approximately 49M parameters |
| Architecture | 15-block U-shaped joint text/image transformer |
| Hidden width | 512 |
| Attention | 8 heads, head dimension 64 |
| FFN | SwiGLU, width 1,152 |
| Long skip connections | 7 |
| Conditioning | Shared adaptive layer normalization |
| Local mixing | Depthwise 3x3 convolution |
| Objective | Rectified flow matching with logit-normal time sampling |
| Image latents | 32 channels, 32x spatial compression |
| 512x512 latent grid | 16x16 |
| Text encoder | Frozen `google/flan-t5-base`, 768 features, up to 96 tokens |
| Decoder | Frozen SANA 1.1 DC-AE F32C32 |
| Saved optimizer step | 210,000 |
| Configured training schedule | 5,000,000 steps |
The parameter count covers the generator only. Downloading the frozen text encoder and autoencoder adds separate dependencies. `SolPixTransformer2D` predicts latent velocity. `AutoencoderDCSol` loads the corresponding Diffusers `AutoencoderDC` for decoding.
## Generate an image
Download the repository, install its requirements, and give `generate.py` a prompt and output path:
```bash
python -m pip install -r requirements.txt
python generate.py --prompt "A glass greenhouse in a quiet garden after rain" --output solpix.png
```
To choose a local checkpoint, seed, or sampling settings:
```bash
python generate.py \
--checkpoint ./step_00210000.pt \
--prompt "A small red sailboat on a misty lake at sunrise" \
--seed 1234 --steps 40 --guidance-scale 3.5 \
--output ./solpix.png
```
On its first run, the helper downloads Flan-T5 Base and the pinned SANA DC-AE revision. It uses Euler integration with classifier-free guidance, running on CUDA when available. CPU inference is supported but slow.
The flow path is `x_t = (1 - t) x_clean + t noise`. Sampling runs from `t=1` down to `t=0`; encoder and decoder identifiers and revision pins are in `config.json`.
## Data and checkpoint history
We used [MONET v1.2.0](https://huggingface.co/datasets/jasperai/monet) with curation seed `20260924`. The split contains 174,603 training examples and a 9,300-example validation holdout. SANA F32C32 image latents and Flan-T5 Base caption states were encoded before training.
MONET draws from CC12M, CommonCatalog-CC-BY, COYO, Diffusion-Aesthetic-4K, and LAION. Flux Klein, Flux Schnell, and Z-Image supply synthetic captions. Curation checks resolution, aesthetics, NSFW content, watermarks, and near duplicates.
The source records include CC BY 4.0, Apache 2.0, Google permissive, and MIT license labels. A label on a record doesn't grant a new license to its contents. Images and dataset shards aren't redistributed in this repository.
The Windows v1.0 continuation used BF16 on one RTX 3080 Ti, with batch size 4 and gradient accumulation 16. We released optimizer step 210,000. The documented 1.1 continuation retains that split and targets step 300,000.
## Reading the samples
We haven't run a formal image-quality or prompt-following benchmark on this checkpoint. The gallery shows generated examples, without supplying a held-out quality estimate. Composition errors, artifacts, and weak text or fine-detail rendering remain limitations.
There is no built-in safety classifier. Dataset filtering doesn't remove every bias or unwanted association. Flan-T5 and SANA DC-AE have separate licenses and usage terms.
All 15 samples below use the released checkpoint at 512x512, with 32 Euler steps, guidance scale 3.5, and the pinned SANA DC-AE decoder. Their files, prompts, seeds, and SHA-256 values are recorded in `samples/`.
### Sample 01

Prompt: Three Black men sharing french fries at a neighborhood diner, candid documentary photography.
Seed: 260926
### Sample 02

Prompt: A red fox standing in fresh snow beneath pine trees at winter dawn, wildlife photography.
Seed: 260927
### Sample 03

Prompt: A glass greenhouse filled with ferns after rain, soft natural light, botanical photograph.
Seed: 260928
### Sample 04

Prompt: A handmade cobalt blue teapot on a pale stone table, clean studio product photograph.
Seed: 260929
### Sample 05

Prompt: A white sailboat crossing a calm blue bay at golden hour, fine art landscape photograph.
Seed: 260930
### Sample 06

Prompt: An orange cat curled on a wooden chair in a sunlit bookshop, cozy editorial photograph.
Seed: 260931
### Sample 07

Prompt: A small street cafe reflected in wet pavement at night, warm window light, city photograph.
Seed: 260932
### Sample 08

Prompt: A wooden lighthouse on a rocky coast under a cloudy sky, atmospheric landscape photograph.
Seed: 260933
### Sample 09

Prompt: A bowl of ripe peaches on a kitchen counter, morning light, natural still life photograph.
Seed: 260934
### Sample 10

Prompt: A snow-covered cabin among tall pine trees at blue hour, quiet winter landscape photograph.
Seed: 260935
### Sample 11

Prompt: A baker placing fresh bread on a cooling rack in a bright kitchen, documentary photograph.
Seed: 260936
### Sample 12

Prompt: A goldfinch perched on a thin branch among spring blossoms, close-up wildlife photograph.
Seed: 260937
### Sample 13

Prompt: A red bicycle leaning against a brick wall on a leafy neighborhood street, lifestyle photograph.
Seed: 260938
### Sample 14

Prompt: A lemon cake with a slice cut out on a ceramic plate, bright tabletop food photograph.
Seed: 260939
### Sample 15

Prompt: A small observatory beneath a clear star-filled sky, distant mountains, night landscape photograph.
Seed: 260940
## Files
- `step_00210000.pt`: EMA and raw weights, optimizer state, configuration, and training arguments.
- `solpix/`: the transformer, decoder adapter, configuration, data, and training components.
- `generate.py`: prompt-to-image generation. `train.py` starts training; `sample_latents.py` samples latents.
- `samples/`: the 15 PNGs and their metadata. `config.json` records the architecture and external-model manifest.
## License
The code, checkpoint weights, configuration, model card, and supplied banner use [Apache 2.0](LICENSE). Attribution is in [NOTICE](NOTICE). Dataset, Flan-T5, and SANA DC-AE licenses apply separately.