Desearch Embedding 4B

Desearch Embedding 4B is a text embedding model for web and news search, developed by Desearch. It is fine-tuned from Qwen/Qwen3-Embedding-4B and maps search queries and web documents into a shared vector space for first-stage retrieval.

Highlights

  • Built for news and the open web. Fine-tuned on recent news coverage, reference pages and encyclopedic articles, the content a web search engine has to rank every day.
  • Real queries in every form. Short keyword searches, natural-language questions, noisy queries, questions about dated news events, and multi-hop questions that combine facts from linked pages.
  • One model for chunks and whole pages. Documents are seen as paragraph chunks, page openings and full pages, the units a search index stores.
  • Binary first-stage ready. Trained with a binary objective on the leading 256 dimensions, for search stacks that shortlist with compact binary vectors before rescoring with full vectors.

Model details

Base model Qwen/Qwen3-Embedding-4B
Parameters 4B
Embedding dimension 2560
Max sequence length 32K tokens
Pooling Last token, L2-normalized
Query instruction Built-in query prompt
Language English
Training method LoRA fine-tuning
License Apache 2.0

Usage

The weights are a LoRA adapter on Qwen/Qwen3-Embedding-4B; the base model downloads automatically.

pip install -U sentence-transformers peft

Using Sentence Transformers

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("desearch/Desearch-Embedding-4B")

queries = [
    "Who became Apple’s chief executive and will lead the Sept. 9, 2026 “Surprise and Shine” product event?",
    "How much does the base Mac Mini M6 cost in US dollars as of its August 2026 announcement?",
]
documents = [
    "John Ternus, Apple's new chief executive, takes the stage at the September 9, 2026 “Surprise and Shine” event to unveil the latest iPhones and Apple Watches.",
    "Apple surprised buyers in late August 2026 with the Mac Mini M6 at $899 and the Mac Mini M5 Pro at $1,699, both shipping on September 22.",
]

query_embeddings = model.encode(queries, prompt_name="query")
document_embeddings = model.encode(documents)
print(model.similarity(query_embeddings, document_embeddings))

Queries use the built-in query prompt; documents are encoded as they are.

Using Transformers

import torch
import torch.nn.functional as F
from peft import PeftModel
from transformers import AutoModel, AutoTokenizer

task = "Given a web search query, retrieve relevant passages that answer the query"
questions = [
    "Who became Apple’s chief executive and will lead the Sept. 9, 2026 “Surprise and Shine” product event?",
    "How much does the base Mac Mini M6 cost in US dollars as of its August 2026 announcement?",
]
queries = [f"Instruct: {task}\nQuery:{q}" for q in questions]
documents = [
    "John Ternus, Apple's new chief executive, takes the stage at the September 9, 2026 “Surprise and Shine” event to unveil the latest iPhones and Apple Watches.",
    "Apple surprised buyers in late August 2026 with the Mac Mini M6 at $899 and the Mac Mini M5 Pro at $1,699, both shipping on September 22.",
]

tokenizer = AutoTokenizer.from_pretrained("desearch/Desearch-Embedding-4B", padding_side="left")
model = AutoModel.from_pretrained("Qwen/Qwen3-Embedding-4B", dtype=torch.bfloat16)
model = PeftModel.from_pretrained(model, "desearch/Desearch-Embedding-4B").eval()


def encode(texts):
    batch = tokenizer(texts, padding=True, truncation=True, max_length=8192, return_tensors="pt")
    with torch.no_grad():
        hidden = model(**batch).last_hidden_state
    return F.normalize(hidden[:, -1], p=2, dim=1)


print(encode(queries) @ encode(documents).T)

Recommended use cases

  • First-stage retrieval for web search, news search and retrieval-augmented generation
  • Semantic search over news archives, documentation and reference content
  • Candidate generation ahead of a reranker

Limitations

  • Tuned on English text; other languages are not a focus of this release.
  • Queries need the query prompt; documents are encoded without one.
  • Trained on documents up to 2,048 tokens; split longer pages into passages.
  • Embeddings differ from Qwen/Qwen3-Embedding-4B; re-embed an existing corpus when switching.

License

This model is licensed under the Apache License 2.0. It is derived from Qwen/Qwen3-Embedding-4B, which is also licensed under the Apache License 2.0.

Citation

@misc{desearch2026embedding,
  title  = {Desearch Embedding 4B},
  author = {Desearch},
  year   = {2026},
  url    = {https://huggingface.co/desearch/Desearch-Embedding-4B}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for desearch/Desearch-Embedding-4B

Finetuned
(71)
this model