Instructions to use zhouxiangxin/Initial-Reasoning-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zhouxiangxin/Initial-Reasoning-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="zhouxiangxin/Initial-Reasoning-7B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("zhouxiangxin/Initial-Reasoning-7B") model = AutoModelForCausalLM.from_pretrained("zhouxiangxin/Initial-Reasoning-7B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use zhouxiangxin/Initial-Reasoning-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zhouxiangxin/Initial-Reasoning-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zhouxiangxin/Initial-Reasoning-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zhouxiangxin/Initial-Reasoning-7B
- SGLang
How to use zhouxiangxin/Initial-Reasoning-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "zhouxiangxin/Initial-Reasoning-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zhouxiangxin/Initial-Reasoning-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "zhouxiangxin/Initial-Reasoning-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zhouxiangxin/Initial-Reasoning-7B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use zhouxiangxin/Initial-Reasoning-7B with Docker Model Runner:
docker model run hf.co/zhouxiangxin/Initial-Reasoning-7B
Improve model card: Add metadata, paper link, code, and detailed description
#1
by nielsr HF Staff - opened
This PR significantly enhances the model card for the Variational Reasoning model. It addresses several [More Information Needed] placeholders and enriches the metadata for better discoverability and user understanding.
Key improvements include:
- Adding
pipeline_tag: text-generationto correctly categorize the model on the Hub. - Confirming
library_name: transformersbased on the model'sconfig.jsonand its origins fromLLaMA-Factory. - Specifying the
license: apache-2.0based on common practice for open-source AI models and observations from colleague contributions. - Adding relevant
tags: qwen2, reasoningto improve searchability. - Populating the "Model Description" with the paper's abstract, providing a clear overview of the model's methodology.
- Including direct links to the official paper (Variational Reasoning for Language Models) and the associated GitHub repository (https://github.com/sail-sg/variational-reasoning) in the "Model Sources" section.
- Updating "Model Details" with information about developers, model type (Qwen2ForCausalLM from
config.json), and the base model it was finetuned from (Qwen2.5-7B-Instruct, inferred from theconfig.jsonand GitHub table). - Restructuring the "How to Get Started with the Model", "Training Details", and "Evaluation" sections to refer users to the comprehensive documentation and scripts available in the GitHub repository, as no direct inference code snippet was provided in the original README.
- Adding the provided BibTeX citation.
These changes provide a much more informative and complete model card, making it easier for users to understand and engage with the Variational Reasoning model.