Instructions to use pynk17/sd-naruto-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use pynk17/sd-naruto-fp16 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("pynk17/sd-naruto-fp16", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
sd-naruto-fp16
sd-naruto-fp16 is an experimental Stable Diffusion v1.5 fine-tune on the lambdalabs/naruto-blip-captions dataset using FP16 mixed-precision training.
This repository is Experiment B in a three-run engineering comparison:
- FP32 baseline
- FP16 mixed precision
- FP16 mixed precision with gradient checkpointing
The purpose of this run is to provide a controlled FP16 comparison against the FP32 baseline while keeping the main fine-tuning configuration unchanged.
Model Details
Model Description
The model was fine-tuned from stable-diffusion-v1-5/stable-diffusion-v1-5 using the Hugging Face Diffusers train_text_to_image.py workflow.
The run uses the standard Stable Diffusion text-to-image pipeline components, including the VAE, text encoder, tokenizer, conditional UNet, scheduler, safety checker, and feature extractor.
- Developed by: pynk17
- Model type: Stable Diffusion text-to-image latent diffusion model
- Training precision: FP16 mixed precision
- Base model:
stable-diffusion-v1-5/stable-diffusion-v1-5 - Dataset:
lambdalabs/naruto-blip-captions - Framework: PyTorch / Hugging Face Diffusers
- Pipeline:
StableDiffusionPipeline - Model repository:
pynk17/sd-naruto-fp16
Repository Contents
The experiment output contains the standard Stable Diffusion pipeline components:
unetvaetext_encodertokenizerschedulersafety_checkerfeature_extractormodel_index.json- TensorBoard
logs
Uses
Direct Use
The model can be used for experimental Naruto-style text-to-image generation.
It is especially useful as the FP16 arm of a controlled systems comparison examining:
- FP32 versus FP16 mixed-precision fine-tuning
- training precision as a systems optimization
- inference behavior across the resulting pipelines
- memory-oriented optimization strategies
- comparison with FP16 plus gradient checkpointing
Out-of-Scope Use
This model is an experimental fine-tune and has not been validated for production deployment.
It should not be assumed to provide:
- robust prompt following outside the fine-tuning distribution
- consistent photorealism
- unbiased outputs
- production-level reliability
- complete safety guarantees
- stable anatomical or scene-level correctness
How to Get Started with the Model
import torch
from diffusers import StableDiffusionPipeline
model_id = "pynk17/sd-naruto-fp16"
pipe = StableDiffusionPipeline.from_pretrained(
model_id,
torch_dtype=torch.float16,
)
pipe = pipe.to("cuda")
generator = torch.Generator(device="cuda").manual_seed(42)
image = pipe(
prompt="a ninja character with blue armor and spiky hair",
negative_prompt="blurry, low quality, distorted",
num_inference_steps=30,
guidance_scale=7.5,
height=256,
width=256,
generator=generator,
).images[0]
image.save("sd_naruto_fp16_sample.png")
Training Details
Training Data
The model was fine-tuned on:
lambdalabs/naruto-blip-captions
The Diffusers training script used:
image_column = "image"
caption_column = "text"
Training Procedure
Training was performed with the Hugging Face Diffusers train_text_to_image.py example.
This experiment used the same principal fine-tuning configuration as the FP32 baseline, with the key change:
mixed_precision = "fp16"
The run was launched through accelerate on one CUDA process.
Training Hyperparameters
| Hyperparameter | Value |
|---|---|
| Base model | stable-diffusion-v1-5/stable-diffusion-v1-5 |
| Dataset | lambdalabs/naruto-blip-captions |
| Image column | image |
| Caption column | text |
| Resolution | 512 × 512 |
| Train batch size | 1 |
| Gradient accumulation steps | 4 |
| Maximum training steps | 100 |
| Learning rate | 1e-5 |
| LR scheduler | constant |
| LR warmup steps | 0 |
| Random seed | 42 |
| Mixed precision | FP16 |
| Accelerator processes | 1 |
| Device | CUDA |
The effective batch size implied by batch size 1 and four gradient-accumulation steps is approximately 4 samples per optimizer update for a single process.
Evaluation
Evaluation Procedure
The experiment notebook compares the FP32, FP16, and FP16 plus gradient-checkpointing models under a shared inference configuration.
The comparison uses:
seed = 42
num_inference_steps = 30
guidance_scale = 7.5
height = 256
width = 256
negative_prompt = "blurry, low quality, distorted"
The shared prompt set is:
a ninja character with blue armor and spiky haira female ninja character standing in a foresta ninja character using a fire attacka ninja character wearing black clothes at nighta detailed portrait of a ninja character
Results
The FP16 model successfully loaded through StableDiffusionPipeline and participated in the same five-prompt inference comparison as the FP32 and FP16 plus gradient-checkpointing models.
The notebook reports the following aggregate inference-time result:
B_FP16: 1.39 ± 0.02 sec/image
For reference, the same notebook reports:
A_FP32: 1.37 ± 0.02 sec/image
C_FP16_GC: 1.40 ± 0.02 sec/image
These measurements reflect the specific notebook runtime and inference configuration. They should not be interpreted as hardware-independent performance benchmarks.
No FID, KID, CLIP score, or other formal generative-quality metric was reported.
Bias, Risks, and Limitations
The model inherits limitations from both Stable Diffusion v1.5 and the Naruto BLIP caption fine-tuning dataset.
The experiment did not include a dedicated bias, memorization, or safety evaluation.
The saved pipeline includes the Stable Diffusion safety checker. In the broader comparison notebook, the safety checker is shown replacing a generated result with a black image when potential NSFW content is detected.
Generated outputs should therefore be treated as experimental samples rather than validated production outputs.
Technical Specifications
Model Architecture and Objective
This model uses the Stable Diffusion v1.5 latent diffusion architecture.
The training workflow initializes the standard pretrained components, including:
AutoencoderKLUNet2DConditionModel- text encoder
- tokenizer
- diffusion scheduler
The conditional UNet is optimized on image-caption pairs from the Naruto BLIP captions dataset.
The distinguishing systems configuration for this repository is FP16 mixed-precision fine-tuning.
Compute Infrastructure
The FP16 training run was launched using:
accelerate launch --mixed_precision="fp16"
The notebook records:
Num processes: 1
Process index: 0
Local process index: 0
Device: cuda
Mixed precision type: fp16
Reproducibility Notes
The FP16 training configuration is:
accelerate launch \
--mixed_precision="fp16" \
train_text_to_image.py \
--pretrained_model_name_or_path="stable-diffusion-v1-5/stable-diffusion-v1-5" \
--dataset_name="lambdalabs/naruto-blip-captions" \
--image_column="image" \
--caption_column="text" \
--resolution=512 \
--train_batch_size=1 \
--gradient_accumulation_steps=4 \
--max_train_steps=100 \
--learning_rate=1e-5 \
--lr_scheduler="constant" \
--lr_warmup_steps=0 \
--seed=42
This repository should be interpreted together with:
pynk17/sd-naruto-fp32-baselinepynk17/sd-naruto-fp16-grad-checkpointing
The three runs form a controlled experiment around training precision and memory-oriented optimization.
Model Card Author
Priyanka / pynk17
- Downloads last month
- 14
Model tree for pynk17/sd-naruto-fp16
Base model
stable-diffusion-v1-5/stable-diffusion-v1-5