tinye5 / README.md
rizwan3d's picture
Update README.md
d71ca62 verified
|
Raw
History Blame Contribute Delete
5.89 kB
---
language:
- en
license: apache-2.0
library_name: transformers
pipeline_tag: feature-extraction
base_model: sentence-transformers/all-MiniLM-L6-v2
tags:
- sentence-transformers
- embeddings
- text-embeddings
- semantic-search
- information-retrieval
- onnx
- int8
- minilm
- e5
---
# TinyE5-L6-384
**TinyE5-L6-384** is a compact **384-dimensional text embedding model** built from
`sentence-transformers/all-MiniLM-L6-v2` and fine-tuned for semantic search and
information retrieval using E5-style query/passage prefixes.
The model uses **mean pooling + L2 normalization** and is designed for applications
where model size, CPU latency, and deployment efficiency matter.
## Highlights
- **6 Transformer layers**
- **384-dimensional embeddings**
- **22.7M parameters**
- Only the **last 2 Transformer layers** were fine-tuned
- **Mean pooling**
- **L2-normalized embeddings**
- E5-style `query:` and `passage:` prefixes
- FP32 and dynamically quantized **INT8 ONNX** deployment
- INT8 ONNX size: **21.8 MB**
- Measured INT8 CPU throughput: **377.7 texts/s**
---
## Model Architecture
| Property | Value |
|---|---|
| Base model | `sentence-transformers/all-MiniLM-L6-v2` |
| Transformer layers | 6 |
| Hidden / embedding size | 384 |
| Total parameters | 22,713,216 |
| Fine-tuned layers | Last 2 Transformer layers |
| Trainable parameters | 3,548,928 |
| Trainable percentage | 15.62% |
| Pooling | Mean pooling |
| Normalization | L2 normalization |
| Query prefix | `query: ` |
| Passage prefix | `passage: ` |
| Training max length | 128 tokens |
---
## Benchmark Results
### Quality
| Model | Size | STS Spearman ↑ | SciFact Recall@10 ↑ | SciFact nDCG@10 ↑ |
| ------------------------------- | ----------: | -------------: | ------------------: | ----------------: |
| **TinyE5-L6-384 (Safetensors)** | 86.7 MB | 0.8138 | 0.7342 | 0.5897 |
| **TinyE5-L6-384 (FP32 ONNX)** | 86.2 MB | 0.8138 | 0.7342 | 0.5897 |
| **TinyE5-L6-384 (INT8 ONNX)** | **21.8 MB** | 0.8058 | 0.7599 | 0.6020 |
### CPU Inference Performance
| Model | Precision | Size | Latency ↓ | Throughput ↑ | Session RAM ↓ |
| ---------------------- | --------- | ----------: | ---------------: | -----------------: | ------------: |
| TinyE5-L6-384 ONNX | FP32 | 86.2 MB | 2.35 ms/text | 424.7 texts/s | 90.4 MB |
| **TinyE5-L6-384 ONNX** | **INT8** | **21.8 MB** | **0.77 ms/text** | **1299.4 texts/s** | **11.2 MB** |
### INT8 vs FP32 ONNX
| Metric | FP32 | INT8 | Change |
| ----------------- | ------------: | -----------------: | -----------------: |
| Model size | 86.2 MB | **21.8 MB** | **~74.7% smaller** |
| CPU latency | 2.35 ms/text | **0.77 ms/text** | **~67.2% lower** |
| CPU throughput | 424.7 texts/s | **1299.4 texts/s** | **~205.9% higher** |
| Session RAM | 90.4 MB | **11.2 MB** | **~87.6% lower** |
| STS Spearman | **0.8138** | 0.8058 | -0.0080 |
| SciFact Recall@10 | 0.7342 | **0.7599** | **+0.0257** |
| SciFact nDCG@10 | 0.5897 | **0.6020** | **+0.0123** |
### End-to-End Benchmark Time
| Model | STS Time ↓ | SciFact Time ↓ |
| ------------------------------- | ---------: | -------------: |
| **TinyE5-L6-384 (Safetensors)** | 0.72s | 3.71s |
| TinyE5-L6-384 (FP32 ONNX) | 8.82s | 265.58s |
| TinyE5-L6-384 (INT8 ONNX) | 5.46s | 179.11s |
> **Summary:** INT8 quantization reduces TinyE5-L6-384 from **86.2 MB to 21.8 MB (~74.7% smaller)**, lowers CPU inference latency from **2.35 ms/text to 0.77 ms/text (~67.2% lower)**, and increases CPU throughput from **424.7 to 1299.4 texts/s (~3.06× throughput)**. Session RAM also drops from **90.4 MB to 11.2 MB (~87.6% lower)**. STS Spearman decreases slightly from **0.8138 to 0.8058**, while SciFact Recall@10 improves from **0.7342 to 0.7599** and nDCG@10 improves from **0.5897 to 0.6020**. The all-MiniLM-L6-v2 baseline still achieves the highest overall quality, with **0.8194 STS Spearman, 0.7923 Recall@10, and 0.6494 nDCG@10**.
## Docker Example
### INT8 ONNX
```yaml
services:
embedding-server:
image: ghcr.io/huggingface/text-embeddings-inference:cpu-1.9
ports:
- "80:80"
volumes:
- ./data:/data
command:
- --model-id
- GrowBitLabs/tinye5
- --revision
- int8-onnx
- --pooling
- mean
```
### Available Variants
All variants are published under the same Hugging Face repository:
```text
GrowBitLabs/tinye5
```
| Variant | Revision | Model |
| ---------------- | ----------- | ------------------- |
| Safetensors FP32 | `main` | `model.safetensors` |
| ONNX FP32 | `fp32-onnx` | `onnx/model.onnx` |
| ONNX INT8 | `int8-onnx` | `onnx/model.onnx` |
### Safetensors
```yaml
command:
- --model-id
- GrowBitLabs/tinye5
- --revision
- main
- --pooling
- mean
```
### FP32 ONNX
```yaml
command:
- --model-id
- GrowBitLabs/tinye5
- --revision
- fp32-onnx
- --pooling
- mean
```
### INT8 ONNX
```yaml
command:
- --model-id
- GrowBitLabs/tinye5
- --revision
- int8-onnx
- --pooling
- mean
```
For CPU deployment, the **INT8 ONNX** revision is recommended when model size and inference throughput are the priority.
## Attribution
Developed by **[GrowBit Labs](https://growbitlabs.com)**.