File size: 14,477 Bytes
7ca7c4c
1c28a94
 
 
7ca7c4c
 
1c28a94
 
7ca7c4c
1c28a94
1e111ae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1c28a94
7ca7c4c
 
1c28a94
 
 
1e111ae
1c28a94
 
1e111ae
1c28a94
 
 
 
 
 
 
 
 
 
1e111ae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1c28a94
 
 
 
 
 
 
 
 
1e111ae
 
 
 
 
 
 
 
 
 
 
 
 
1c28a94
 
 
1e111ae
 
1c28a94
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1e111ae
 
 
1c28a94
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1e111ae
1c28a94
 
 
1e111ae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1c28a94
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1e111ae
 
 
1c28a94
1e111ae
 
 
 
1c28a94
1e111ae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1c28a94
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1e111ae
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1c28a94
 
 
 
 
 
 
 
 
 
 
 
1e111ae
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
---
title: LFM2.5 Embedding API
emoji: ๐Ÿง 
colorFrom: indigo
colorTo: red
sdk: docker
app_port: 7860
pinned: true
license: mit
tags:
  - embeddings
  - openai-compatible
  - llama-cpp
  - cpu-only
  - rag
  - vector-search
  - semantic-search
  - retrieval
  - text-embedding
  - gguf
  - quantization
  - liquid-ai
  - lfm2.5
  - fastapi
  - docker
  - inference
  - ai
  - machine-learning
  - nlp
  - search
  - similarity
  - cosine-similarity
  - knowledge-base
  - document-indexing
  - multilingual
short_description: OpenAI-compatible embeddings API running LFM2.5 on CPU
---

# ๐Ÿง  LFM2.5 Embedding API (CPU-Optimized)

![Docker](https://img.shields.io/badge/Docker-Ready-blue)
![FastAPI](https://img.shields.io/badge/FastAPI-green)
![llama.cpp](https://img.shields.io/badge/llama.cpp-CPU%20Optimized-orange)
![OpenAI Compatible](https://img.shields.io/badge/OpenAI-Compatible-success)
![License](https://img.shields.io/badge/License-MIT%20%2B%20LFM--1.0-lightgrey)

An **ultra-lightweight** and **100% OpenAI-compatible** REST API for generating embeddings using the [LiquidAI/LFM2.5-Embedding-350M](https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M) model, optimized to run on **pure CPU** (no GPU required) on Hugging Face Spaces Free tier.

## โšก Key Features

- ๐Ÿš€ **Fast cold start** (~3-5s) thanks to pre-loaded Q8_0 GGUF model
- ๐Ÿ’พ **Minimal RAM usage** (~500MB) โ€” runs comfortably within HF Free's 16GB limit
- ๐ŸŽฏ **Asymmetric embeddings** โ€” supports `query:` and `document:` prefixes for maximum RAG precision
- ๐Ÿ” **Bearer Token authentication** via HF Secrets
- ๐Ÿค– **Strict OpenAI standard** โ€” works with any OpenAI-compatible client
- ๐ŸŒ **Multilingual support** โ€” LFM2.5 handles 100+ languages
- ๐Ÿ“ฆ **Easy deployment** โ€” clone and run in 5 minutes

## โš ๏ธ Performance & Limitations (IMPORTANT)

### Current Performance on HF Free Tier

| Metric | Value | Notes |
|--------|-------|-------|
| **Cold Start** | 3-5 seconds | First request after inactivity |
| **Inference Time** | 10-15 seconds | Per embedding request (50-500 words) |
| **Throughput** | ~6-8 requests/minute | Sustained rate |
| **Max Context** | 512 tokens | Optimal for embeddings |
| **Dimensions** | 1024 floats | Per embedding vector |

### Why 10-15 Seconds?

This API runs on **Hugging Face Spaces Free tier**, which uses **shared CPU resources**:

- **CPU Throttling**: The HF hypervisor dynamically limits CPU cycles when multiple containers compete for resources
- **No AVX2 Optimization**: While the code is compiled with AVX2 support, the free tier's virtualization layer doesn't fully expose these CPU instructions
- **Shared Infrastructure**: Your container shares physical CPU cores with other users' Spaces

### When to Use This API

โœ… **Perfect for:**
- Prototyping and development
- Small-scale RAG applications (<1000 documents)
- Personal projects and experimentation
- Batch processing with async queues
- Learning and education
- Backup/fallback embedding service

โŒ **Not ideal for:**
- Real-time user-facing search (latency too high)
- High-throughput production systems (>100 req/min)
- Applications requiring <1s response times
- Critical infrastructure with SLA requirements

### ๐Ÿš€ Need Better Performance?

If you need faster inference, consider these alternatives:

| Option | Latency | Cost | Setup Complexity |
|--------|---------|------|------------------|
| **This Space (HF Free)** | 10-15s | $0 | โญ Minimal |
| **HF Space Paid (Basic)** | 2-5s | ~$0.60/h | โญ Minimal |
| **Cloudflare Workers AI** | 50-200ms | $0 (10k neurons/day) | โญโญ Low |
| **OpenAI Embeddings** | 100-300ms | $0.0001/1K tokens | โญโญ Low |
| **Self-hosted (GPU)** | 10-50ms | Hardware cost | โญโญโญโญ High |

**Recommended Alternative**: [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/) with `@cf/qwen/qwen3-embedding-0.6b` offers 50-200ms latency on their global edge network, with 10,000 free neurons/day (~18k requests/day for 500-token embeddings).

## ๐Ÿ”Œ Universal Compatibility

This API can be used as an embedding backend for **any tool** that supports OpenAI-compatible endpoints:

| Tool | Works? | Notes |
|------|--------|-------|
| **OpenClaw** | โœ… Full | Use `queryInputType: "query"` and `documentInputType: "document"` |
| **Open WebUI** | โœ… Full | Configure as "OpenAI API" in RAG Settings |
| **LangChain** | โœ… Full | Use `OpenAIEmbeddings(openai_api_base=...)` |
| **LlamaIndex** | โœ… Full | Use `OpenAIEmbedding(api_base=...)` |
| **Cursor** | โœ… Full | Point `api_base` to this Space |
| **Continue** | โœ… Full | Configure in `config.json` |
| **Dify** | โœ… Full | Use "OpenAI Embeddings" node |
| **Flowise** | โœ… Full | Use OpenAI Embeddings component |
| **n8n** | โœ… Full | Use OpenAI node with custom base URL |
| **Haystack** | โœ… Full | Use `OpenAIEmbedder` with `api_base_url` |
| **Semantic Kernel** | โœ… Full | Configure OpenAI connector |
| **AutoGen** | โœ… Full | Use OpenAI-compatible embedding model |
| **CrewAI** | โœ… Full | Configure embedding provider |
| **OpenAI SDK (Python)** | โœ… Full | Override `base_url` and `api_key` |
| **OpenAI SDK (Node.js)** | โœ… Full | Override `baseURL` and `apiKey` |

## ๐Ÿ› ๏ธ Quick Start (How to Clone and Use)

### 1. Duplicate this Space
Click the **three dots** (โ‹ฎ) in the top-right corner โ†’ **"Duplicate this Space"** โ†’ choose visibility (Public/Private).

### 2. Configure the Secret
In the duplicated Space, go to **Settings โ†’ Variables and Secrets** and add:

| Name | Value |
|------|-------|
| `API_KEY` | `your-secret-key-here` |

*(Use a strong string with 32+ characters. Tip: Use [1Password's Password Generator](https://1password.com/password-generator) for cryptographically secure random passwords)*

### 3. Wait for Build
The Dockerfile will compile `llama.cpp` and download the Q8_0 model automatically (~5-8 minutes on first build).

### 4. Test with cURL
```bash
curl -X POST "https://YOUR-USERNAME-YOUR-SPACE.hf.space/v1/embeddings" \
     -H "Authorization: Bearer YOUR_SECRET_KEY" \
     -H "Content-Type: application/json" \
     -d '{
           "input": "Text to generate embedding for",
           "model": "LiquidAI/LFM2.5-Embedding-350M",
           "input_type": "document"
         }'
```

## ๐Ÿ“ Usage Examples

### Python with OpenAI SDK

```python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_SECRET_KEY",
    base_url="https://YOUR-USERNAME-YOUR-SPACE.hf.space/v1/"
)

# For indexing documents (store in vector DB)
doc_response = client.embeddings.create(
    model="LiquidAI/LFM2.5-Embedding-350M",
    input="OpenClaw is a multilingual RAG tool.",
    extra_body={"input_type": "document"}
)

# For searching (user query)
query_response = client.embeddings.create(
    model="LiquidAI/LFM2.5-Embedding-350M",
    input="How does vector search work?",
    extra_body={"input_type": "query"}
)

print(query_response.data[0].embedding[:5])  # First 5 floats
```

### LangChain Integration

```python
from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(
    model="LiquidAI/LFM2.5-Embedding-350M",
    openai_api_key="YOUR_SECRET_KEY",
    openai_api_base="https://YOUR-USERNAME-YOUR-SPACE.hf.space/v1/"
)

# Generate embeddings
vectors = embeddings.embed_documents(["Text 1", "Text 2"])
query_vector = embeddings.embed_query("Search query")
```

### LlamaIndex Integration

```python
from llama_index.embeddings.openai import OpenAIEmbedding

embed_model = OpenAIEmbedding(
    model="LiquidAI/LFM2.5-Embedding-350M",
    api_key="YOUR_SECRET_KEY",
    api_base="https://YOUR-USERNAME-YOUR-SPACE.hf.space/v1/"
)

embeddings = embed_model.get_text_embedding("Your text here")
```

### JavaScript/Node.js

```javascript
import OpenAI from 'openai';

const client = new OpenAI({
  apiKey: 'YOUR_SECRET_KEY',
  baseURL: 'https://YOUR-USERNAME-YOUR-SPACE.hf.space/v1/'
});

const response = await client.embeddings.create({
  model: 'LiquidAI/LFM2.5-Embedding-350M',
  input: 'Your text here'
});

console.log(response.data[0].embedding.slice(0, 5));
```

## ๐Ÿ—๏ธ Architecture

```
+-------------------------------------------------------+
|  Hugging Face Space (2 vCPU / 16GB RAM)               |
|                                                       |
|   +------------------+         +------------------+   |
|   |                  |         | llama.cpp (C++)  |   |
|   |      Docker      +-------->|   + Q8_0 GGUF    |   |
|   |    Container     |         |  (~380MB RAM)    |   |
|   |                  |         +------------------+   |
|   +--------+---------+                                |
|            |                                          |
|            v                                          |
|   +------------------+         +------------------+   |
|   |     FastAPI      |         |  /v1/embeddings  |   |
|   |    + Uvicorn     +-------->| (OpenAI-compat.) |   |
|   +------------------+         +------------------+   |
+-------------------------------------------------------+
```

- **Engine**: `llama-cpp-python` (Python wrapper for llama.cpp)
- **Model**: `LFM2.5-Embedding-350M-Q8_0.gguf` (8-bit quantization for maximum precision)
- **Framework**: FastAPI + Uvicorn (1 worker)
- **Pooling**: Native CLS Token (LFM2.5 standard)
- **Dimensions**: 1024 floats per embedding
- **Context**: 512 tokens (optimal for embeddings)

## ๐ŸŽฏ Best Practices

### 1. Use Asymmetric Embeddings
Always specify `input_type` for better RAG performance:
- `"document"` when indexing/storing text
- `"query"` when searching

### 2. Implement Chunking
Break long documents into ~400-token chunks with 50-token overlap for optimal retrieval.

### 3. Batch Requests
Send multiple texts in a single request when possible:
```python
response = client.embeddings.create(
    input=["Text 1", "Text 2", "Text 3"],
    model="LiquidAI/LFM2.5-Embedding-350M"
)
```

### 4. Set Appropriate Timeouts
Configure your HTTP client with 30-second timeouts to handle cold starts.

### 5. Use Async Queues for Indexing
For large document collections, implement background processing to avoid blocking user interactions.

## โš–๏ธ Licensing

This repository contains **two distinct layers** with different licenses:

### 1. API Code (Infrastructure) โ€” MIT License
All Python code (FastAPI, Dockerfile, scripts) is licensed under the **MIT License**. You are free to clone, modify, use commercially, and distribute the API infrastructure.

### 2. AI Model (Weights and GGUF) โ€” LFM Open License v1.0
The model weights are owned by **Liquid AI, Inc.** and licensed under the **[LFM Open License v1.0](https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M/blob/main/LICENSE)**.

#### โš ๏ธ CRITICAL: Commercial Use Threshold

**Commercial use is PERMITTED only if your Legal Entity's total annual revenue does NOT exceed $10,000,000 USD (ten million US dollars).**

If your entity exceeds this threshold, you **must obtain a separate commercial license** from Liquid AI, Inc.

**What this means for you:**

โœ… **PERMITTED:**
- Using this API to generate embeddings for RAG applications
- Integration in commercial products (under the $10M threshold)
- Academic research and non-commercial projects
- Personal projects and experimentation

โš ๏ธ **REQUIRES COMMERCIAL LICENSE:**
- Entities with >$10M annual revenue must contact [Liquid AI](https://www.liquid.ai)
- Exceeding the threshold without a license results in **automatic termination** (Section 11)

๐Ÿšซ **PROHIBITED:**
- Using Liquid AI trademarks ("Liquid AI", "LFM", etc.) to promote your product
- Patent litigation against Liquid AI (triggers license termination)
- Use for training competing foundation models

๐Ÿ“„ **Full License Text**: See the [`LICENSE`](./LICENSE) file in this repository or the [official license](https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M/blob/main/LICENSE) on Hugging Face.

## ๐Ÿ™ Credits

This project was architected in collaboration with:

- ๐Ÿค– **Qwen3.7 Max** โ€” DevOps architecture, Docker/CPU optimization, llama.cpp integration, and build troubleshooting
- ๐Ÿค– **Gemini 3.1 Pro** โ€” Code review, asymmetric embedding validation (query/document), and OpenClaw compliance analysis

Base model by [Liquid AI](https://huggingface.co/LiquidAI).

## ๐Ÿ“‹ Compliance Checklist

Before deploying this project in production, verify:

- [ ] Your entity's annual revenue is under $10M USD, OR you have obtained a commercial license from Liquid AI
- [ ] You have read and understood the full LFM Open License v1.0
- [ ] If redistributing, you have included the LICENSE file and preserved all copyright notices
- [ ] You are not using Liquid AI trademarks to promote your product
- [ ] You understand that violations result in automatic license termination
- [ ] You have configured appropriate timeouts (30s+) in your client
- [ ] You have implemented chunking for long documents
- [ ] You understand the performance characteristics (10-15s per request)

## ๐Ÿ› Troubleshooting

### Issue: Cold Start Takes Too Long
**Solution**: This is normal for HF Free tier. The first request after inactivity takes 3-5 seconds to load the model. Subsequent requests are faster.

### Issue: Requests Timeout
**Solution**: Increase your HTTP client timeout to 30 seconds. The HF Free tier can be slow under load.

### Issue: Build Fails with OOMKilled
**Solution**: This is a known HF builder limitation. The current Dockerfile uses `CMAKE_BUILD_PARALLEL_LEVEL=1` to avoid this. If it still fails, try a Factory Reboot.

### Issue: API Returns 401 Unauthorized
**Solution**: Verify your `API_KEY` secret is correctly configured in Settings โ†’ Variables and Secrets. The name must be exactly `API_KEY` (uppercase).

### Issue: Poor Search Results
**Solution**: Ensure you're using `input_type: "query"` for searches and `input_type: "document"` for indexing. This asymmetric approach significantly improves RAG quality.

## ๐Ÿค Contributing

Contributions are welcome! Please open an issue or pull request.

## ๐Ÿ“„ License Summary

- **API Code**: MIT License (see [`LICENSE`](./LICENSE))
- **Model Weights**: LFM Open License v1.0 by Liquid AI, Inc.

---

**Built with โค๏ธ for the RAG and vector search open-source community.**

**Need faster performance?** Consider [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/) with `@cf/qwen/qwen3-embedding-0.6b` for 50-200ms latency on their global edge network.