pplx-embed-v2-context-9b-preview

pplx-embed-v2-context-9b-preview is a contextual embedding model for document chunks in RAG systems. A document is passed as a list of chunks; the chunks are encoded together, so each chunk's embedding reflects its surrounding context, and one embedding is returned per chunk.

This is a preview release, not a final model. Weights, embeddings, and the interface may change in later versions without backward compatibility, so embeddings produced with this preview should not be mixed with embeddings from a future release.

Queries and documents are encoded with different methods: use encode_queries for queries and encode for document chunks. The model is trained with separate query and document prefixes, and encoding queries with encode silently degrades retrieval quality.

Like pplx-embed-context-v1, the model natively produces unnormalized int8-quantized embeddings. Compare embeddings with cosine similarity, or pass normalize_embeddings=True and use the dot product.

Model

Model Dimensions MRL Quantization Instruction Pooling
pplx-embed-v2-context-9b-preview 2048 1024, 2048 INT8 No (fixed query/document prefixes) Mean

Usage

Requires transformers>=5.4.0, torch, numpy, safetensors, and tqdm. The model uses custom code, so load it with trust_remote_code=True.

from transformers import AutoModel

model = AutoModel.from_pretrained(
    "perplexity-ai/pplx-embed-v2-context-9b-preview",
    trust_remote_code=True,
).to("cuda")

doc_chunks = [
    [
        "Curiosity begins in childhood with endless questions about the world.",
        "As we grow, curiosity drives us to explore new ideas.",
        "Scientific breakthroughs often start with a curious question.",
    ],
    [
        "The curiosity rover explores Mars searching for ancient life.",
        "Each discovery on Mars sparks new questions about the universe.",
    ],
]

# One (chunk_count, 2048) array per document:
# doc_embeddings[0].shape == (3, 2048), doc_embeddings[1].shape == (2, 2048)
doc_embeddings = model.encode(doc_chunks, normalize_embeddings=True)

# Each query is a single-chunk row.
queries = [["What drives scientific breakthroughs?"]]
query_embeddings = model.encode_queries(queries, normalize_embeddings=True)

scores = doc_embeddings[0] @ query_embeddings[0][0]

Options

encode(documents, ...) and encode_queries(queries, ...) accept:

Argument Default Description
batch_size 32 Documents (or queries) per forward pass
normalize_embeddings False L2-normalize outputs
convert_to_numpy True Return NumPy arrays; False returns CPU tensors
show_progress_bar False Show a progress bar
device None Move the model to this device before encoding

Matryoshka dimensions

The model was trained with Matryoshka losses at 1024 and 2048 dimensions. To use 1024-dimensional embeddings, take the first 1024 values of each unnormalized embedding and normalize afterwards:

import numpy as np

emb = model.encode(doc_chunks)  # unnormalized int8 values
emb_1024 = [e[:, :1024] / np.linalg.norm(e[:, :1024], axis=-1, keepdims=True) for e in emb]

Other truncation sizes were not trained.

Downloads last month
145
Safetensors
Model size
8B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support