Post
58
Maverick-4B-Unity-XR-Agent is now on Hugging Face.
It's a 4B model that turns spoken or typed English into actions in Unity scenes. Say "put the red mug on the table" or "turn on the lamp", and it returns the tool call your app executes. If a command could mean two objects, it asks which one. If it can't do something, it says so instead of guessing.
Everything runs on the user's machine through llama.cpp: no API key, no internet connection. The Q4_K_M GGUF is 2.5 GB and needs about 3 GB of GPU memory, so it fits on a 4 GB laptop GPU and usually answers in one to three seconds. It is fine-tuned from Qwen3-4B with QLoRA on about 20,000 English conversations.
Results:
- 83.8% on 499 human-written ALFRED instructions (right action on the right object). The base model, Qwen3-4B, scores 57.1%. The strongest of the five other models we tested, from 1.7B to 120B parameters, was Ministral 3 14B at 67.1%.
- 91.7% on object types it never saw in training.
- 97.3% on 440 commands run through a live Unity scene.
There is also a Unity package that starts the model, describes the scene to it and carries out its tool calls. You install it from the Package Manager with a Git URL.
Model: ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Unity package: ErenAta00/Maverick-Unity
Full write-up: https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent
Built at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University:
ExtendedRealityLabMCBU
Released under Apache-2.0. Feedback and bug reports are welcome in the Community tab.
It's a 4B model that turns spoken or typed English into actions in Unity scenes. Say "put the red mug on the table" or "turn on the lamp", and it returns the tool call your app executes. If a command could mean two objects, it asks which one. If it can't do something, it says so instead of guessing.
Everything runs on the user's machine through llama.cpp: no API key, no internet connection. The Q4_K_M GGUF is 2.5 GB and needs about 3 GB of GPU memory, so it fits on a 4 GB laptop GPU and usually answers in one to three seconds. It is fine-tuned from Qwen3-4B with QLoRA on about 20,000 English conversations.
Results:
- 83.8% on 499 human-written ALFRED instructions (right action on the right object). The base model, Qwen3-4B, scores 57.1%. The strongest of the five other models we tested, from 1.7B to 120B parameters, was Ministral 3 14B at 67.1%.
- 91.7% on object types it never saw in training.
- 97.3% on 440 commands run through a live Unity scene.
There is also a Unity package that starts the model, describes the scene to it and carries out its tool calls. You install it from the Package Manager with a Git URL.
Model: ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF
Unity package: ErenAta00/Maverick-Unity
Full write-up: https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent
Built at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University:
Released under Apache-2.0. Feedback and bug reports are welcome in the Community tab.