Instructions to use mixedbread-ai/mxbai-embed-large-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use mixedbread-ai/mxbai-embed-large-v1 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers.js
How to use mixedbread-ai/mxbai-embed-large-v1 with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('feature-extraction', 'mixedbread-ai/mxbai-embed-large-v1'); - Transformers
How to use mixedbread-ai/mxbai-embed-large-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="mixedbread-ai/mxbai-embed-large-v1")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("mixedbread-ai/mxbai-embed-large-v1") model = AutoModel.from_pretrained("mixedbread-ai/mxbai-embed-large-v1", device_map="auto") - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use mixedbread-ai/mxbai-embed-large-v1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf mixedbread-ai/mxbai-embed-large-v1:F16 # Run inference directly in the terminal: llama cli -hf mixedbread-ai/mxbai-embed-large-v1:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf mixedbread-ai/mxbai-embed-large-v1:F16 # Run inference directly in the terminal: llama cli -hf mixedbread-ai/mxbai-embed-large-v1:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf mixedbread-ai/mxbai-embed-large-v1:F16 # Run inference directly in the terminal: ./llama-cli -hf mixedbread-ai/mxbai-embed-large-v1:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf mixedbread-ai/mxbai-embed-large-v1:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf mixedbread-ai/mxbai-embed-large-v1:F16
Use Docker
docker model run hf.co/mixedbread-ai/mxbai-embed-large-v1:F16
- LM Studio
- Jan
- Ollama
How to use mixedbread-ai/mxbai-embed-large-v1 with Ollama:
ollama run hf.co/mixedbread-ai/mxbai-embed-large-v1:F16
- Unsloth Studio
How to use mixedbread-ai/mxbai-embed-large-v1 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for mixedbread-ai/mxbai-embed-large-v1 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for mixedbread-ai/mxbai-embed-large-v1 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for mixedbread-ai/mxbai-embed-large-v1 to start chatting
- Atomic Chat new
- Docker Model Runner
How to use mixedbread-ai/mxbai-embed-large-v1 with Docker Model Runner:
docker model run hf.co/mixedbread-ai/mxbai-embed-large-v1:F16
- Lemonade
How to use mixedbread-ai/mxbai-embed-large-v1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull mixedbread-ai/mxbai-embed-large-v1:F16
Run and chat with the model
lemonade run user.mxbai-embed-large-v1-F16
List all available models
lemonade list
Add exported onnx model 'model_O3.onnx'
Hello!
This pull request adds an exported onnx model (model_O3.onnx).
Config
OptimizationConfig(
optimization_level=2,
enable_transformers_specific_optimizations=True,
optimize_for_gpu=False,
fp16=False,
disable_gelu_fusion=False,
disable_attention_fusion=False,
disable_bias_gelu_fusion=False,
disable_layer_norm_fusion=False,
disable_rotary_embeddings=False,
disable_skip_layer_norm_fusion=False,
disable_bias_skip_layer_norm_fusion=False,
disable_skip_group_norm_fusion=False,
disable_bias_splitgelu_fusion=False,
disable_bias_add_fusion=False,
disable_group_norm_fusion=True,
disable_embed_layer_norm_fusion=True,
enable_gemm_fast_gelu_fusion=False,
use_mask_index=False,
disable_packed_kv=True,
no_attention_mask=False,
use_raw_attention_mask=False,
disable_shape_inference=False,
use_multi_head_attention=False,
enable_gelu_approximation=True,
use_group_norm_channels_first=False,
disable_packed_qkv=False,
disable_nhwc_conv=False
)
Testing this pull request
You can test this pull request before merging by loading the model from this PR with the revision argument:
from sentence_transformers import SentenceTransformer
# NOTE: Update this to the number of your pull request
pr_number = 2
model = SentenceTransformer(
"mixedbread-ai/mxbai-embed-large-v1",
revision=f"refs/pr/{pr_number}",
backend="onnx",
model_kwargs={"file_name": "model_O3.onnx"},
)
# Verify that everything works as expected
embeddings = model.encode(["The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium."])
print(embeddings.shape)
similarities = model.similarity(embeddings, embeddings)
print(similarities)
This PR was auto-generated with export_optimized_onnx_model.