Qlip / README.md
Qarvexium's picture
Update README.md
f2a8149 verified
|
Raw
History Blame Contribute Delete
1.83 kB
---
language:
- en
license: mit
tags:
- text-encoder
- text-to-image
- diffusion
- stable-diffusion
- qlip
- contrastive-learning
library_name: pytorch
pipeline_tag: feature-extraction
---
# Qlip
Qlip is a novel text encoder architecture designed for text-to-image (Qanvas) generation tasks sharing similar objectives with CLIP, Qlip is optimized for maximum quality while taking little compute power.
## Model Details
- **Model Type:** Text Encoder
- **Architecture:** Custom Transformer-based encoder
- **Parameters:** ~50M
- **Vocab Size:** 48,000 tokens
- **Embedding Dimension:** 512
- **Max Sequence Length:** 256 tokens
- **Hidden Size:** 512
- **Layers:** 8 transformer encoder layers
- **Attention Heads:** 8
- **FFN Size:** 2048
## Installation
```bash
pip install torch sentencepiece
```
## Usage
### Basic Text Encoding
```python
from Qlip import load_text_encoder, get_tokenizer, encode_text
# Load model and tokenizer
model = load_text_encoder(device='cuda')
tokenizer = get_tokenizer()
# Encode text
embeddings = encode_text("a beautiful sunset over the ocean", device='cuda')
print(embeddings.shape) # torch.Size([1, 512])
```
### Compute Text Similarity
```python
from Qlip import encode_text, similarity
emb1 = encode_text("a cat sitting on a mat", device='cuda')
emb2 = encode_text("a dog lying on a rug", device='cuda')
sim = similarity(emb1, emb2)
print(f"Similarity: {sim:.2f}%")
```
## Intended Use
Qlip is designed for:
- Text-to-image generation (as the text encoder component)
- Text-image similarity computation
- Feature extraction for downstream tasks
- Semantic search and retrieval
## Limitations
- Maximum sequence length of 256 tokens
- English-only
- Trained specifically for image generation tasks
- May not generalize well to other text encoding tasks
## License
MIT License