Instructions to use tiiuae/falcon-mamba-7b-instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tiiuae/falcon-mamba-7b-instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tiiuae/falcon-mamba-7b-instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tiiuae/falcon-mamba-7b-instruct") model = AutoModelForCausalLM.from_pretrained("tiiuae/falcon-mamba-7b-instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tiiuae/falcon-mamba-7b-instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tiiuae/falcon-mamba-7b-instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tiiuae/falcon-mamba-7b-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tiiuae/falcon-mamba-7b-instruct
- SGLang
How to use tiiuae/falcon-mamba-7b-instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tiiuae/falcon-mamba-7b-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tiiuae/falcon-mamba-7b-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tiiuae/falcon-mamba-7b-instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tiiuae/falcon-mamba-7b-instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tiiuae/falcon-mamba-7b-instruct with Docker Model Runner:
docker model run hf.co/tiiuae/falcon-mamba-7b-instruct
Trained on ChatGPT Responses?
I told falcon-mamba its own name, but falcon-mamba told me it was not falcon-mamba. Rather, it said it was a model trained by OpenAI. Were ChatGPT responses used to train this model? If so, is this data part of RefinedWeb? If not, what’s the issue?
Hi @astrologos
Thanks for your message. I can confirm we did not trained the model explicitly on ChatGPT prompts. However note that the common web datasets (including instruction datasets) that are present nowadays to train and fine tune LLMs are quite contaminated with synthetic data generated from LLMs such as ChatGPT (see for example the screenshot below, taken from: https://huggingface.co/spaces/HuggingFaceFW/blogpost-fineweb-v1).
One technique to overcome this is to perform further model alignment by doing identification tuning so that the model "knows" he is not ChatGPT but another LLM. I suggest to stay up to date in this organization as we will release more models in the future which are better in all sense, including identification and alignment
Hi @ybelkada , many thanks for your detailed response. I have a lot of hope for SSMs. Hopefully someday we'll see constant-time Deep Koopman LLMs ;)