tharunpranavsakthivel's picture
Refresh model card and checksums
8b981f3 verified
|
Raw
History Blame Contribute Delete
3.87 kB
---
library_name: transformers
pipeline_tag: text-generation
language: en
license: gemma
base_model: google/functiongemma-270m-it
base_model_relation: finetune
tags:
- tinyshell
- structured-generation
- json
- tool-use
- function-calling
- text-generation
---
# FunctionGemma TinyShell
Fine-tuned compact language model for the **TinyShell ShellIntent**
natural-language-to-structured-IR task.
## Base model
`google/functiongemma-270m-it`
This is a supervised fine-tune of the upstream FunctionGemma instruction model.
The upstream model is distributed under Google’s Gemma Terms of Use. See
`NOTICE` and `LICENSE` for the required notice and the authoritative terms.
## Task
The model converts natural-language instructions into structured TinyShell
ShellIntent JSON.
Supported high-level decisions include:
- `compile`
- `clarify`
- `unsupported`
## Held-out evaluation
| Metric | Result |
|---|---:|
| JSON parse rate | 99.00% |
| Schema validity | 95.50% |
| IR exact match | 33.00% |
| Decision accuracy | 94.50% |
| Operation accuracy | 61.00% |
| Slot precision | 66.64% |
| Slot recall | 57.87% |
| Slot F1 | 61.94% |
| Risk accuracy | 93.50% |
| Confirmation accuracy | 94.00% |
| Clarify accuracy | 100.00% |
| Unsupported accuracy | 30.00% |
| Multi-operation accuracy | 22.22% |
| Median inference latency | 2049.0688229992884 ms |
## Generation policy
**JSON-completion stopping criterion**
FunctionGemma and Falcon-H1 initially produced a valid first JSON object but
frequently continued generating additional content. Their corrected final
evaluation uses a generation-time stopping criterion that terminates once the
first complete top-level JSON object is generated. This is generation control,
not post-hoc JSON repair.
LFM2.5 terminated correctly under the original inference configuration.
## Training
The model was fine-tuned with supervised causal language modeling.
- Seed: `42`
- Best validation loss: `0.12864468747895444`
- Training time: `1370.75630064` seconds
- Peak GPU memory: `5.382477760314941` GB
Prompt tokens were masked from the language-model loss and the assistant JSON
response was used as the supervised target.
Training used 1,600 examples, with 200 validation examples and 200 held-out
test examples. The random seed was `42`. The frozen source hashes and complete
training metadata are included in `evaluation/training_result.json`.
## Included files
- Fine-tuned model weights
- Model configuration
- Tokenizer / processor files
- Chat template when saved
- Generation configuration when saved
- `evaluation/final_metrics.json`
- `evaluation/test_predictions.jsonl`
- `evaluation/training_result.json`
- `inference_example.py`
- `requirements.txt`
- `LICENSE` and `NOTICE`
- `SHA256SUMS.txt`
## Limitations
This pilot used one training seed. Test-set bootstrap intervals quantify
held-out sample uncertainty but do not replace independent repeated training.
Exact ShellIntent matching is intentionally strict: one incorrect operation,
argument, or structured field makes the complete IR prediction incorrect.
This model emits untrusted structured intent. Do not execute model output
directly. Validate the JSON against the TinyShell schema, compile it through a
deterministic platform-aware compiler, apply safety checks, and require user
confirmation where appropriate.
## License
The model weights are a derivative of FunctionGemma and are subject to the
[Gemma Terms of Use](https://ai.google.dev/gemma/terms), including the
incorporated [Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy).
The required Gemma distribution notice is in `NOTICE`.
The TinyShell training data contribution is attributed under [CC BY
4.0](https://creativecommons.org/licenses/by/4.0/). Upstream source material
may have separate terms; see the TinyShell dataset documentation for details.