Instructions to use AtlaAI/Selene-1-Mini-Llama-3.1-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AtlaAI/Selene-1-Mini-Llama-3.1-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AtlaAI/Selene-1-Mini-Llama-3.1-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AtlaAI/Selene-1-Mini-Llama-3.1-8B") model = AutoModelForCausalLM.from_pretrained("AtlaAI/Selene-1-Mini-Llama-3.1-8B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AtlaAI/Selene-1-Mini-Llama-3.1-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AtlaAI/Selene-1-Mini-Llama-3.1-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtlaAI/Selene-1-Mini-Llama-3.1-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AtlaAI/Selene-1-Mini-Llama-3.1-8B
- SGLang
How to use AtlaAI/Selene-1-Mini-Llama-3.1-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AtlaAI/Selene-1-Mini-Llama-3.1-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtlaAI/Selene-1-Mini-Llama-3.1-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AtlaAI/Selene-1-Mini-Llama-3.1-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AtlaAI/Selene-1-Mini-Llama-3.1-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AtlaAI/Selene-1-Mini-Llama-3.1-8B with Docker Model Runner:
docker model run hf.co/AtlaAI/Selene-1-Mini-Llama-3.1-8B
Limited Evaluation Capabilities
The model seems to be a good small LLM-as-a-Judge model overall from what I can say so far. Yet, it only seems to be good at evaluating question-answer pairs or generated answer to ground truth answer, which is an important, but a very limited aspect of evaluation.
The model struggles when, e.g. tasked to evaluate whether the retrieved context is relevant for answering the question or whether the answer contains the retrieved context correctly. The model fails completely when tasked to evaluate whether the information contained in the retieved chunks is similar to the information contained in the ground truth chunks.
From what I can tell it seems to me that the model wasn't trained on these types of evaluation tasks, hence the prompts being OOD, which is sad, as it limits its capabilities and its usage.
The model also fails at comparing and evaluating ground truth chunks with retrieved entities/relationships and their descriptions in GraphRAG scenarios.
All these aspects are very important, as shipping a system that was evaluated only on the question and its answer is way to inconsistent in reality.
Hi @h4rz3rk4s3 we appreciate for the thoughtful feedback!
Would love to know more about the exact use case you tried our model with. We're looking to learn more about how to improve the next generation of our models. Would you be open to sharing more? Feel free to email me about this as well: maurice@atla-ai.com
You are correct that we have not trained the model for the retrieved chunks vs. ground truth chunks scenario you describe; so your assessment does not surprise me there. I am more puzzled about the struggles you have found when trying to evaluate whether the answer contains the retrieved context correctly. This is a task that we do specifically train for and have seen users have really good results. Would love to understand better where we failed you there.
Excited to hear from you!
Maurice