File size: 5,134 Bytes
ab083a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
---
language:
- en
license: apache-2.0
base_model: sapientinc/HRM-Text-1B
library_name: keras
pipeline_tag: text-generation
tags:
- hrm
- prefix-lm
- code
- tool-use
- keras
- keras-hub
- sft
---

# HRM-Text-1B Code & Tool-Use SFT

This page tracks successive code and tool-use SFT stages of
[`sapientinc/HRM-Text-1B`](https://huggingface.co/sapientinc/HRM-Text-1B),
saved as native Keras/KerasHub presets. The current default revision and the
immutable `stage-a-v1` tag are a full-parameter Stage A pilot. They adapt the
base PrefixLM to the study's coding and tool-use transcript syntax.

> **Pilot release.** This model has not yet undergone downstream benchmark
> evaluation. Training loss is not a measure of coding-agent or tool-use task
> performance. Do not treat it as a production-ready agent.

## Release lineage

| Tag | State | Availability |
| --- | --- | --- |
| `stage-a-v1` | Completed 10M-response-token Stage A pilot, seed 17 | Current default revision |
| `stage-b-v1` | Stage B continuation | Not trained or published yet |

Stage B will be published as `stage-b-v1` on this same page only after its
training and release checks complete. `stage-a-v1` will remain immutable for
reproducibility.

## Training result

![Training loss](training-loss.png)

| Metric | Value |
| --- | ---: |
| Training budget | 10,000,147 response tokens |
| Serialized tokens | 19,650,833 |
| Microbatches / optimizer updates | 38,248 / 4,781 planned |
| Final training loss | 0.03527 |
| Validation loss | 0.02808 |
| Final validation token accuracy | 0.90152 |
| Hardware | One NVIDIA H100 80 GB |
| Wall time | 14 h 44 min |

The chart uses the canonical completed-run telemetry segment. The thin trace is
the loss logged every ten microbatches; the dark trace is a 25-point trailing
mean.

## Training data and protocol

- Dataset: [`pzarzycki/hrm-text-code-tools-sft`](https://huggingface.co/datasets/pzarzycki/hrm-text-code-tools-sft), canonical v2 Stage A.
- Source: `nvidia/OpenCodeInstruct` pinned to revision
  `8f3ba5bafe4d6e8db46082cf7ae6741bc370604d` (CC-BY-4.0).
- The deterministic pilot selection contains 38,248 rows from the sealed
  Stage A training split. It is not a full pass over all 1.13M Stage A rows.
- Context cap: 4,096 serialized tokens. Oversized rows were rejected during
  data preparation; no example was silently truncated.
- The only condition used was `direct`.

The learned SFT envelope is:

```text
<|im_start|><|object_ref_start|>instruction<|im_end|>response<|box_end|>
```

`<user>`, `<assistant>`, `<tools>`, `<tool_call>`, and `<tool_result>` inside
the content are ordinary learned transcript markup, not additional pretrained
HRM control tokens.

## Optimization

Full-parameter BF16 fine-tuning used KerasHub HRM-Text support, AdamW
(`lr=3e-5`, `beta_1=0.9`, `beta_2=0.95`, `weight_decay=0.1`), global gradient
clip norm 1.0, EMA (`0.999`), batch size 1 with gradient accumulation 8,
and 3% warmup. Loss is applied only to response tokens, including
`<|box_end|>`; prefix tokens receive zero loss weight.

## Files

| Path | Purpose |
| --- | --- |
| `preset/` | Final Keras preset: model configuration, tokenizer, preprocessor, and final weights. |
| `checkpoint-00020000.keras` | Full Keras checkpoint saved at microbatch 20,000. |
| `training/pilot-completion.json` | Immutable completion manifest and metrics SHA-256. |
| `training/pilot-completed-metrics.jsonl` | Canonical first-run loss, learning-rate, token, and throughput telemetry. |
| `training/gpu.jsonl` | GPU utilization, VRAM, power, and temperature telemetry. |
| `training/platform.jsonl` | RunPod GPU/CPU utilization samples. |
| `training/history.json` | Final Keras training and validation summary. |
| `training/tensorboard/` | TensorBoard events for the completed training run. |
| `ARTIFACTS.tsv` | Sizes and release-integrity checksums for the key artifacts. |

## Loading and prompting

The preset requires a KerasHub build with HRM-Text support. With the study's
HRM-enabled KerasHub branch installed, load the preset from a local clone or
using the Hugging Face handle:

```python
from keras_hub.models import HrmTextCausalLM

model = HrmTextCausalLM.from_preset(
    "hf://pzarzycki/hrm-text-1b-code-tools-sft"
)
prompt = (
    "<|im_start|><|object_ref_start|>"
    "Write a Python function that returns the larger of two integers."
    "<|im_end|>"
)
print(model.generate(prompt, max_length=256))
```

HRM-Text is a PrefixLM base model, not a ChatML/Qwen chat model. Keep the
`direct` condition and the `<|im_end|>` instruction boundary at inference.

## Limitations and intended use

This is a research artifact for continued evaluation of code generation and a
fixed tool-transcript protocol. It does not itself execute tools, validate tool
calls, sandbox generated code, or establish benchmark performance. Use
independent task-level evaluation before deployment.

## Provenance

The base model is Apache-2.0. Stage A training data is CC-BY-4.0; source pins,
dataset checksums, serializer details, preflight configuration, and completed
telemetry are included in this repository's `training/` artifacts.