inference.py

Model Overview

A large-scale implementation of the blip architecture, built for generation tasks.

Architecture

  • Architecture: blip
  • Scale: large
  • Attention: flash
  • Fusion strategy: tucker
  • Task head: generation
  • Activation: gelu
  • Normalization: batchnorm
  • Initialization: orthogonal

Training

  • Optimizer: lion
  • LR scheduler: polynomial

Files

  • inference.py — main artifact of this repository

License

See the license field above.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support