Image-Text-to-Text
Transformers
Safetensors
GGUF
English
qwen3_5
nvidia
nemotron
tool-calling
function-calling
vision-language
multimodal
vllm
endpoints-template
custom_code
conversational
Instructions to use wdrones/nemo-qwen3_5_4b_base_jellyfish with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="wdrones/nemo-qwen3_5_4b_base_jellyfish", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("wdrones/nemo-qwen3_5_4b_base_jellyfish", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("wdrones/nemo-qwen3_5_4b_base_jellyfish", trust_remote_code=True, device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M # Run inference directly in the terminal: llama cli -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M # Run inference directly in the terminal: llama cli -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Use Docker
docker model run hf.co/wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "wdrones/nemo-qwen3_5_4b_base_jellyfish" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wdrones/nemo-qwen3_5_4b_base_jellyfish", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
- SGLang
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "wdrones/nemo-qwen3_5_4b_base_jellyfish" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wdrones/nemo-qwen3_5_4b_base_jellyfish", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "wdrones/nemo-qwen3_5_4b_base_jellyfish" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wdrones/nemo-qwen3_5_4b_base_jellyfish", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Ollama
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Ollama:
ollama run hf.co/wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
- Unsloth Studio
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for wdrones/nemo-qwen3_5_4b_base_jellyfish to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for wdrones/nemo-qwen3_5_4b_base_jellyfish to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for wdrones/nemo-qwen3_5_4b_base_jellyfish to start chatting
- Pi
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Docker Model Runner:
docker model run hf.co/wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
- Lemonade
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Run and chat with the model
lemonade run user.nemo-qwen3_5_4b_base_jellyfish-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use wdrones/nemo-qwen3_5_4b_base_jellyfish with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "wdrones/nemo-qwen3_5_4b_base_jellyfish:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| """Custom Inference Handler for Hugging Face Inference Endpoints. | |
| This checkpoint is a Qwen3.5 vision-language, tool-calling SFT model | |
| (``Qwen3_5ForConditionalGeneration``). The handler wraps the model in the | |
| ``EndpointHandler`` interface expected by the HF Inference Toolkit: | |
| https://huggingface.co/docs/inference-endpoints/main/en/engines/toolkit | |
| Request payload (``data``) accepts the following keys: | |
| inputs (required) Either: | |
| * a plain string prompt, or | |
| * a list of OpenAI/Qwen-style chat messages, e.g. | |
| [{"role": "user", "content": "Hello"}] | |
| Message ``content`` may be a string or a list of | |
| {"type": "text"|"image"|"video", ...} parts. Image | |
| parts accept "image" / "image_url" as a URL, | |
| data URI, or base64 string. | |
| tools (optional) List of tool/function JSON schemas. When provided they | |
| are injected via the model's chat template so the model | |
| can emit <tool_call> blocks. | |
| parameters (optional) Dict of generation kwargs (max_new_tokens, | |
| temperature, top_p, top_k, do_sample, ...). | |
| Response: a list with a single dict ``[{"generated_text": "..."}]``. | |
| """ | |
| from __future__ import annotations | |
| import base64 | |
| import io | |
| from typing import Any, Dict, List, Optional, Union | |
| import torch | |
| from PIL import Image | |
| from transformers import AutoModelForImageTextToText, AutoProcessor | |
| # Defaults applied when the caller does not override them via ``parameters``. | |
| DEFAULT_GENERATION_KWARGS: Dict[str, Any] = { | |
| "max_new_tokens": 1024, | |
| "do_sample": False, | |
| "temperature": 0.7, | |
| "top_p": 0.9, | |
| } | |
| class EndpointHandler: | |
| def __init__(self, path: str = "") -> None: | |
| # ``path`` is the directory with the model weights provided by the | |
| # Inference Endpoint runtime (the repository root). | |
| self.processor = AutoProcessor.from_pretrained(path, trust_remote_code=True) | |
| dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32 | |
| self.model = AutoModelForImageTextToText.from_pretrained( | |
| path, | |
| dtype=dtype, | |
| device_map="auto" if torch.cuda.is_available() else None, | |
| trust_remote_code=True, | |
| ) | |
| self.model.eval() | |
| self.tokenizer = getattr(self.processor, "tokenizer", self.processor) | |
| # ------------------------------------------------------------------ # | |
| # Helpers | |
| # ------------------------------------------------------------------ # | |
| def _load_image(ref: str) -> Image.Image: | |
| """Load a PIL image from a URL, data URI, or raw base64 string.""" | |
| if ref.startswith("http://") or ref.startswith("https://"): | |
| import requests | |
| resp = requests.get(ref, timeout=30) | |
| resp.raise_for_status() | |
| return Image.open(io.BytesIO(resp.content)).convert("RGB") | |
| if ref.startswith("data:"): | |
| # data:image/png;base64,XXXX | |
| ref = ref.split(",", 1)[1] | |
| return Image.open(io.BytesIO(base64.b64decode(ref))).convert("RGB") | |
| def _collect_images(self, messages: List[Dict[str, Any]]) -> List[Image.Image]: | |
| images: List[Image.Image] = [] | |
| for message in messages: | |
| content = message.get("content") | |
| if not isinstance(content, list): | |
| continue | |
| for item in content: | |
| if not isinstance(item, dict): | |
| continue | |
| if item.get("type") == "image" or "image" in item or "image_url" in item: | |
| ref = item.get("image") or item.get("image_url") | |
| if isinstance(ref, dict): | |
| ref = ref.get("url") | |
| if isinstance(ref, str): | |
| images.append(self._load_image(ref)) | |
| return images | |
| def _build_messages( | |
| self, inputs: Union[str, List[Dict[str, Any]]] | |
| ) -> List[Dict[str, Any]]: | |
| if isinstance(inputs, str): | |
| return [{"role": "user", "content": inputs}] | |
| return inputs | |
| # ------------------------------------------------------------------ # | |
| # Inference | |
| # ------------------------------------------------------------------ # | |
| def __call__(self, data: Dict[str, Any]) -> List[Dict[str, Any]]: | |
| inputs = data.pop("inputs", data) | |
| tools: Optional[List[Dict[str, Any]]] = data.pop("tools", None) | |
| parameters: Dict[str, Any] = data.pop("parameters", {}) or {} | |
| messages = self._build_messages(inputs) | |
| images = self._collect_images(messages) | |
| prompt = self.processor.apply_chat_template( | |
| messages, | |
| tools=tools, | |
| tokenize=False, | |
| add_generation_prompt=True, | |
| ) | |
| processor_kwargs: Dict[str, Any] = {"text": [prompt], "return_tensors": "pt"} | |
| if images: | |
| processor_kwargs["images"] = images | |
| model_inputs = self.processor(**processor_kwargs) | |
| model_inputs = {k: v.to(self.model.device) for k, v in model_inputs.items()} | |
| generation_kwargs = {**DEFAULT_GENERATION_KWARGS, **parameters} | |
| # When sampling is disabled, drop sampling-only knobs to avoid warnings. | |
| if not generation_kwargs.get("do_sample", False): | |
| for key in ("temperature", "top_p", "top_k"): | |
| generation_kwargs.pop(key, None) | |
| generated_ids = self.model.generate(**model_inputs, **generation_kwargs) | |
| # Strip the prompt tokens so only the completion is decoded. | |
| input_len = model_inputs["input_ids"].shape[1] | |
| new_tokens = generated_ids[:, input_len:] | |
| text = self.processor.batch_decode( | |
| new_tokens, | |
| skip_special_tokens=True, | |
| clean_up_tokenization_spaces=False, | |
| )[0] | |
| return [{"generated_text": text}] | |