FunctionGemma 270M: phone assistant (20 tools, English + Hinglish)

This is a full-parameter fine-tune of google/functiongemma-270m-it. It is a Siri-style assistant that turns a user command into calls to 20 phone tools: alarms, timers, date/time, flashlight, Wi-Fi, Bluetooth, volume, brightness, weather, calendar, reminders, messages, calls, apps, music, media control, navigation, web search and notes. For small talk or unsupported requests it replies without calling a tool.

Results (800 held-out examples, all 20 tools in every prompt, greedy decoding)

Metric Base This model
Exact match (all functions + arguments) 22.3% 75.4%
Tool selection 31.6% 93.0%
Parse errors 10.6% 0.25%
No-call recall (stays silent when it should) 90.4% 79.8%

Exact match by slice: single call 75.3%, parallel (2–3 calls) 70.8%, English 79.3%, Hinglish 53.3%. The confidence interval on the exact-match gain is [+49.2, +56.9] pp (95%, cluster bootstrap). Known weaknesses: it over-calls on Hinglish small talk, garbles some names, misconverts some times, and phrases free-text arguments differently from the gold labels.

Usage

Use the same tool schemas and developer prompt that the model was trained with. They are in tools.json and prompting.py in the GitHub repo. The model emits FunctionGemma's call syntax.

import json
from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "ankit-pn/functiongemma-270m-phone-assistant"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)
tools = json.load(open("tools.json"))          # from the GitHub repo
from prompting import SYSTEM                  # from the GitHub repo

msgs = [{"role": "developer", "content": SYSTEM},
        {"role": "user", "content": "kal subah 6:30 ka alarm laga do aur wifi band kar do"}]
ids = tok.apply_chat_template(msgs, tools=tools, add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = model.generate(**ids, max_new_tokens=128, do_sample=False)
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=False))
# <start_function_call>call:set_alarm{time:<escape>06:30<escape>}<end_function_call><start_function_call>call:toggle_wifi{state:<escape>off<escape>}<end_function_call>...

score.py in the GitHub repo has a strict parser for this output format.

Training

  • 3,200 examples, 3 epochs, 600 steps, effective batch 16, LR 5e-5 with cosine schedule and 5% warmup. Loss on assistant tokens only. Two Kaggle T4s, FP32 weights with BF16 autocast, 85 minutes.
  • The weights are saved in FP32. The final checkpoint is used as-is; it was not selected on test data.
  • Note: the chat template of the Kaggle copy of FunctionGemma v1 duplicates }<end_function_call> in tool-call turns. This model was trained with the Hugging Face hub template, which is included here.

Limitations

Results come from a single seed. The dataset was written by LLMs following a spec, not collected from real users. This is a research artifact: validate the calls before executing them on a real device. Use is subject to the Gemma Terms of Use.

Downloads last month
417
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ankit-pn/functiongemma-270m-phone-assistant

Finetuned
(463)
this model

Dataset used to train ankit-pn/functiongemma-270m-phone-assistant