Asilarkness's picture
Upload README.md with huggingface_hub
ec06204 verified
|
Raw
History Blame Contribute Delete
1.3 kB
---
language:
- en
license: mit
pipeline_tag: text-to-image
tags:
- anime
- diffusion
- text-to-image
- pytorch
- latent-diffusion
- ddim
- classifier-free-guidance
---
# AnimeGen - Lightweight Latent Diffusion Text-to-Image Model
Compact text-to-image model trained from scratch on **65,675** real anime images with Danbooru tags. Uses **latent diffusion** + **DDIM** + **classifier-free guidance (CFG)**.
## Architecture
1. **Autoencoder** - 96x96 -> 64-channel latent. Reconstruction **PSNR 27.8 dB**.
2. **GRU tag encoder** - comma-separated Danbooru tags -> conditioning vector.
3. **Conditional Latent Diffusion U-Net** (base=256) + **CFG** at sampling.
## Key Metrics
- Trained on **65,675 unique** anime images (96x96)
- AE reconstruction: **PSNR 27.8 dB**
- **Classifier-free guidance** (cfg=4.0) at sampling
- **Latent normalization** (per-channel mean/std) prevents out-of-distribution latents
- **DDIM 30 steps** + CFG: fast generation
## Usage
```python
python infer.py --ckpt model.pt --prompt "1girl, halo, cowboy shot, long hair" --out results -n 4 --cfg 4.0
```
Use comma-separated Danbooru-style tags: `1girl, halo, cowboy shot, long hair`, `2girls, school uniform`, etc.
## Files
- `model.pt` - full trained model (15 MB)
- `infer.py` - inference script