Instructions to use Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S") model = AutoModelForCausalLM.from_pretrained("Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S
- SGLang
How to use Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S with Docker Model Runner:
docker model run hf.co/Clinical-Reasoning-Hub/Diagnostic-Reasoning-Q3X1S
Diagnostic-Reasoning-Q3 โ bfloat16 serialisation
The same trained model as Diagnostic-Reasoning-Q3X1, stored in bfloat16 rather than float16.
The published results were produced from the Q3X1 float16 release, whose weights are cast to bfloat16 at load by the evaluation harness. Use Q3X1 to reproduce the published figures, and see its model card for evaluation results, overlap analysis and limitations.
- Downloads last month
- -