Instructions to use nvidia/llama-nemotron-embed-1b-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nvidia/llama-nemotron-embed-1b-v2 with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("nvidia/llama-nemotron-embed-1b-v2", trust_remote_code=True) sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Transformers
How to use nvidia/llama-nemotron-embed-1b-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="nvidia/llama-nemotron-embed-1b-v2", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nvidia/llama-nemotron-embed-1b-v2", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Dismatch of this model and the model from nvidia/llama-3_2-nv-embedqa-1b-v2
I just test the embedding of this model and compare the result from nvidia/llama-3_2-nv-embedqa-1b-v2. The vectors are different. What's the reason?
@mzl Please can you provide more details on the reference you're comparing to? Are you referring to the hosted version running here https://build.nvidia.com/nvidia/llama-3_2-nv-embedqa-1b-v2 or something else?
I suspect this may have been related to the version of transformers used to load the model . We have made a change in #16 to improve support for more transformers versioins which should mitigate this issue. Closing this issue now.
Feel free to open a new one with more details if there is still an issue with latest version.