Spaces:
Running
Running
Publish Codex Benchmark evaluation evidence
Browse files- DATA_FORMAT.md +1 -1
- README.md +11 -5
- REPORT.md +11 -5
- THIRD_PARTY_NOTICES.md +12 -1
- app.js +1 -1
- comparison.json +2218 -271
- data.json +0 -0
- episodes.csv +20 -0
- episodes.json +0 -0
- index.html +3 -3
- licenses/LIBERO-LICENSE.txt +21 -0
- licenses/robocasa-LICENSE.txt +28 -0
- licenses/robosuite-LICENSE.txt +28 -0
- manifest.json +16 -13
- protocol.json +62 -3
- publication-manifest.json +39 -27
- summary.json +48 -8
- task-index.json +1463 -0
DATA_FORMAT.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
| 2 |
|
| 3 |
`data.json` contains the task catalog, per-family summary and completed episode records. `episodes.json` and `episodes.csv` contain the same measured results. Pending tasks have no invented verdict or usage. `attempt-history.json` indexes preserved provider interruptions separately.
|
| 4 |
|
| 5 |
-
Each `episodes/
|
| 6 |
|
| 7 |
Codex native sessions use response_item records. `provider-usage.jsonl` contains native token_usage_record events, not reconstructed HTTP requests. Input includes cached input; output includes reasoning. Failed requests without reported usage add no tokens. Missing usage remains unknown.
|
| 8 |
|
|
|
|
| 2 |
|
| 3 |
`data.json` contains the task catalog, per-family summary and completed episode records. `episodes.json` and `episodes.csv` contain the same measured results. Pending tasks have no invented verdict or usage. `attempt-history.json` indexes preserved provider interruptions separately.
|
| 4 |
|
| 5 |
+
Each `episodes/FAMILY/SLOT/seed-0/` directory contains the native verdict, protocol, owner journals, full Codex session, structured trajectory, provider usage records, transcript, observed images, saved resources, video and measured provenance.
|
| 6 |
|
| 7 |
Codex native sessions use response_item records. `provider-usage.jsonl` contains native token_usage_record events, not reconstructed HTTP requests. Input includes cached input; output includes reasoning. Failed requests without reported usage add no tokens. Missing usage remains unknown.
|
| 8 |
|
README.md
CHANGED
|
@@ -10,12 +10,18 @@ pinned: false
|
|
| 10 |
|
| 11 |
# Codex Benchmark
|
| 12 |
|
| 13 |
-
|
| 14 |
|
| 15 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
|
| 17 |
-
|
| 18 |
|
| 19 |
-
|
| 20 |
|
| 21 |
-
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
# Codex Benchmark
|
| 12 |
|
| 13 |
+
62 completed ordinary Codex tasks across RoboCasa, LIBERO Long, and RoboDojo. Results are reported separately by simulator family.
|
| 14 |
|
| 15 |
+
| Family | Native successes | Tasks | Success rate |
|
| 16 |
+
| --- | ---: | ---: | ---: |
|
| 17 |
+
| RoboCasa | 2 | 10 | 20.00% |
|
| 18 |
+
| LIBERO Long | 9 | 10 | 90.00% |
|
| 19 |
+
| RoboDojo | 31 | 42 | 73.81% |
|
| 20 |
|
| 21 |
+
Stock Codex CLI 0.160.0, GPT-6 Astra / high / default, seed 0, fresh independent workspace. RoboCasa and LIBERO use 6,000 steps at 20 Hz; LIBERO uses initial state 0. RoboDojo retains 7,500 frames at 25 Hz. All use stepped simulation and an 8-hour safety limit. Each task has one formal episode.
|
| 22 |
|
| 23 |
+
RoboCasa and LIBERO use the latest reviewed catalog instructions and the same native simulator source, assets, and task settings as their Kinex references. LIBERO Kinex was fully rerun at caac19a; RoboCasa retains its current a1eaece results. The intervening Kinex commits affect session timing and footer display. See [comparison.json](comparison.json) for paired results, usage, timing, and evidence boundaries.
|
| 24 |
|
| 25 |
+
Native spectator videos are presented at 4x speed and 20 FPS. Per-task pages retain full sessions, observed images, saved tools and memos, native verdicts, instructions, protocol, usage, and source identities. Native failures are valid results; detailed failure mechanisms are unreviewed unless supported by a separate review. The four historical RoboDojo provider interruptions remain in [attempt-history.json](attempt-history.json).
|
| 26 |
+
|
| 27 |
+
[Results](episodes.csv) · [Protocol](protocol.json) · [Comparison](comparison.json)
|
REPORT.md
CHANGED
|
@@ -1,11 +1,17 @@
|
|
| 1 |
# Codex Benchmark
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
|
| 7 |
-
|
| 8 |
|
| 9 |
-
|
| 10 |
|
| 11 |
-
|
|
|
|
|
|
|
|
|
| 1 |
# Codex Benchmark
|
| 2 |
|
| 3 |
+
62 completed ordinary Codex tasks across RoboCasa, LIBERO Long, and RoboDojo. Results are reported separately by simulator family.
|
| 4 |
|
| 5 |
+
| Family | Native successes | Tasks | Success rate |
|
| 6 |
+
| --- | ---: | ---: | ---: |
|
| 7 |
+
| RoboCasa | 2 | 10 | 20.00% |
|
| 8 |
+
| LIBERO Long | 9 | 10 | 90.00% |
|
| 9 |
+
| RoboDojo | 31 | 42 | 73.81% |
|
| 10 |
|
| 11 |
+
Stock Codex CLI 0.160.0, GPT-6 Astra / high / default, seed 0, fresh independent workspace. RoboCasa and LIBERO use 6,000 steps at 20 Hz; LIBERO uses initial state 0. RoboDojo retains 7,500 frames at 25 Hz. All use stepped simulation and an 8-hour safety limit. Each task has one formal episode.
|
| 12 |
|
| 13 |
+
RoboCasa and LIBERO use the latest reviewed catalog instructions and the same native simulator source, assets, and task settings as their Kinex references. LIBERO Kinex was fully rerun at caac19a; RoboCasa retains its current a1eaece results. The intervening Kinex commits affect session timing and footer display. See [comparison.json](comparison.json) for paired results, usage, timing, and evidence boundaries.
|
| 14 |
|
| 15 |
+
Native spectator videos are presented at 4x speed and 20 FPS. Per-task pages retain full sessions, observed images, saved tools and memos, native verdicts, instructions, protocol, usage, and source identities. Native failures are valid results; detailed failure mechanisms are unreviewed unless supported by a separate review. The four historical RoboDojo provider interruptions remain in [attempt-history.json](attempt-history.json).
|
| 16 |
+
|
| 17 |
+
[Results](episodes.csv) · [Protocol](protocol.json) · [Comparison](comparison.json)
|
THIRD_PARTY_NOTICES.md
CHANGED
|
@@ -1,5 +1,16 @@
|
|
| 1 |
# Third-party notices
|
| 2 |
|
| 3 |
-
This edition presents
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 4 |
|
| 5 |
RoboDojo source: https://github.com/RoboDojo-Benchmark/RoboDojo/tree/726e9aabfaa642203722eb126f5eaf0f37f3e1ad . The pinned source carries the RoboDojo Non-Commercial Research License; see licenses/RoboDojo-LICENSE.txt. Attribution: the RoboDojo authors and the Isaac Sim, robotics and scene-asset contributors. This release presents non-commercial evaluation results, rendered views and agent evidence. No simulator asset packs, saved-layout archives or Isaac Sim software are redistributed. Underlying components retain their own terms.
|
|
|
|
| 1 |
# Third-party notices
|
| 2 |
|
| 3 |
+
This edition presents rendered episodes from RoboCasa, LIBERO Long and RoboDojo. No simulator asset packs are redistributed.
|
| 4 |
+
|
| 5 |
+
- robocasa: pinned source `456174f62b89b8fca99eaaf33949c29fec9cfc2a`; see [license](licenses/robocasa-LICENSE.txt).
|
| 6 |
+
- robosuite: pinned source `5ce6643f3092639d08f7b0f90ed1c6a84f50552c`; see [license](licenses/robosuite-LICENSE.txt).
|
| 7 |
+
- LIBERO: pinned source `8f1084e3132a39270c3a13ebe37270a43ece2a01`; see [license](licenses/LIBERO-LICENSE.txt).
|
| 8 |
+
- RoboDojo: pinned source `726e9aabfaa642203722eb126f5eaf0f37f3e1ad`; see [license](licenses/RoboDojo-LICENSE.txt).
|
| 9 |
+
|
| 10 |
+
RoboCasa code is MIT-licensed. Its pinned README identifies assets and datasets as CC BY 4.0: https://github.com/robocasa/robocasa/tree/456174f62b89b8fca99eaaf33949c29fec9cfc2a#license . Video and observation images derive from that scene; the published video combines two native camera streams and accelerates simulation time by 4x. Attribution: the RoboCasa Team and the robosuite contributors, with component notices retained above. Underlying third-party assets retain their own terms. This publication makes no blanket license claim for all model outputs, recordings or bundled generated resources.
|
| 11 |
+
|
| 12 |
+
LIBERO source: https://github.com/Lifelong-Robot-Learning/LIBERO/tree/8f1084e3132a39270c3a13ebe37270a43ece2a01 . Its README identifies code as MIT and datasets as CC BY 4.0. Attribution: the LIBERO authors and robosuite contributors; the original notices are preserved in licenses/.
|
| 13 |
+
|
| 14 |
+
Instruction clarifications for three LIBERO Long tasks draw on [BenchMend](https://benchmend-gallery.static.hf.space/gallery/index.html). Attribution: the BenchMend contributors. Additional current catalog refinements are recorded in each task instruction and source provenance.
|
| 15 |
|
| 16 |
RoboDojo source: https://github.com/RoboDojo-Benchmark/RoboDojo/tree/726e9aabfaa642203722eb126f5eaf0f37f3e1ad . The pinned source carries the RoboDojo Non-Commercial Research License; see licenses/RoboDojo-LICENSE.txt. Attribution: the RoboDojo authors and the Isaac Sim, robotics and scene-asset contributors. This release presents non-commercial evaluation results, rendered views and agent evidence. No simulator asset packs, saved-layout archives or Isaac Sim software are redistributed. Underlying components retain their own terms.
|
app.js
CHANGED
|
@@ -31,7 +31,7 @@ function failureCategories(review){
|
|
| 31 |
}
|
| 32 |
function renderFailureOverview(){
|
| 33 |
const failures=data.episodes.filter(e=>!e.success);
|
| 34 |
-
$('#failure-overview').innerHTML=`<div class="section-top"><h2>
|
| 35 |
$('#failure-filter').innerHTML='<option value="all">All results</option><option value="failed">All unsuccessful results</option>';
|
| 36 |
}
|
| 37 |
const successPercent=value=>value==null?'—':(value*100).toFixed(2)+'%';
|
|
|
|
| 31 |
}
|
| 32 |
function renderFailureOverview(){
|
| 33 |
const failures=data.episodes.filter(e=>!e.success);
|
| 34 |
+
$('#failure-overview').innerHTML=`<div class="section-top"><h2>当前结果与失败分析</h2><a href="attempt-history.json" target="_blank" rel="noopener">Attempt history ↗</a></div><p class="section-note">${data.tasks.length} 个任务 · ${data.episodes.filter(e=>e.success).length} 个成功 · ${failures.length} 个原生失败 · ${data.tasks.length-data.episodes.length} 个待发布。原生判定与完整 session 已保留;具体失败机制尚未逐项定位。服务端中断及其补跑记录见尝试历史。</p>`;
|
| 35 |
$('#failure-filter').innerHTML='<option value="all">All results</option><option value="failed">All unsuccessful results</option>';
|
| 36 |
}
|
| 37 |
const successPercent=value=>value==null?'—':(value*100).toFixed(2)+'%';
|
comparison.json
CHANGED
|
@@ -1,354 +1,2301 @@
|
|
| 1 |
{
|
| 2 |
-
"
|
| 3 |
-
"
|
| 4 |
-
|
| 5 |
-
"
|
| 6 |
-
"
|
| 7 |
-
"kinex_successes_same_subset": 35,
|
| 8 |
-
"kinex_successes_all42": 35,
|
| 9 |
-
"all42_comparison_complete": true
|
| 10 |
},
|
| 11 |
-
"
|
| 12 |
{
|
| 13 |
-
"
|
| 14 |
-
"
|
| 15 |
-
"
|
| 16 |
-
"
|
| 17 |
-
"
|
| 18 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 19 |
},
|
| 20 |
{
|
| 21 |
-
"
|
| 22 |
-
"
|
| 23 |
-
"
|
| 24 |
-
"
|
| 25 |
-
"
|
| 26 |
-
|
| 27 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 28 |
{
|
| 29 |
-
"task_key": "
|
| 30 |
-
"
|
| 31 |
-
"
|
| 32 |
"codex_success": false,
|
| 33 |
"kinex_success": false,
|
| 34 |
-
"
|
| 35 |
-
|
| 36 |
-
|
| 37 |
-
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 43 |
},
|
| 44 |
{
|
| 45 |
-
"task_key": "
|
| 46 |
-
"
|
| 47 |
-
"
|
| 48 |
-
"codex_success": true,
|
| 49 |
-
"kinex_success": true,
|
| 50 |
-
"kinex_version": "0.10.0"
|
| 51 |
-
},
|
| 52 |
-
{
|
| 53 |
-
"task_key": "task04/06",
|
| 54 |
-
"native_id": "robodojo/store-tools-in-toolbox",
|
| 55 |
-
"status": "completed",
|
| 56 |
"codex_success": false,
|
| 57 |
"kinex_success": false,
|
| 58 |
-
"
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
|
| 63 |
-
|
| 64 |
-
|
| 65 |
-
|
| 66 |
-
|
| 67 |
-
|
| 68 |
-
|
| 69 |
-
|
| 70 |
-
|
| 71 |
-
|
| 72 |
-
|
| 73 |
-
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
|
| 77 |
-
|
| 78 |
-
|
| 79 |
-
|
| 80 |
-
|
| 81 |
-
|
| 82 |
-
|
| 83 |
-
|
| 84 |
-
|
| 85 |
-
|
| 86 |
-
|
| 87 |
-
|
| 88 |
-
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
|
| 92 |
-
|
| 93 |
-
|
| 94 |
-
|
| 95 |
-
|
| 96 |
-
|
| 97 |
-
"
|
| 98 |
-
|
| 99 |
-
|
| 100 |
-
|
| 101 |
-
|
| 102 |
-
|
| 103 |
-
|
| 104 |
-
|
| 105 |
-
|
| 106 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 107 |
},
|
| 108 |
{
|
| 109 |
-
"task_key": "
|
| 110 |
-
"
|
| 111 |
-
"
|
| 112 |
-
"codex_success":
|
| 113 |
-
"kinex_success": true,
|
| 114 |
-
"kinex_version": "0.10.3"
|
| 115 |
-
},
|
| 116 |
-
{
|
| 117 |
-
"task_key": "task04/15",
|
| 118 |
-
"native_id": "robodojo/classify-objects",
|
| 119 |
-
"status": "completed",
|
| 120 |
-
"codex_success": true,
|
| 121 |
-
"kinex_success": true,
|
| 122 |
-
"kinex_version": "0.10.2"
|
| 123 |
-
},
|
| 124 |
-
{
|
| 125 |
-
"task_key": "task04/16",
|
| 126 |
-
"native_id": "robodojo/fill-egg-holder",
|
| 127 |
-
"status": "completed",
|
| 128 |
-
"codex_success": true,
|
| 129 |
"kinex_success": false,
|
| 130 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
},
|
| 132 |
{
|
| 133 |
-
"task_key": "
|
| 134 |
-
"
|
| 135 |
-
"
|
| 136 |
"codex_success": false,
|
| 137 |
"kinex_success": false,
|
| 138 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 139 |
},
|
| 140 |
{
|
| 141 |
-
"task_key": "
|
| 142 |
-
"
|
| 143 |
-
"
|
| 144 |
"codex_success": true,
|
| 145 |
-
"kinex_success":
|
| 146 |
-
"
|
| 147 |
-
|
| 148 |
-
|
| 149 |
-
|
| 150 |
-
|
| 151 |
-
|
| 152 |
-
|
| 153 |
-
|
| 154 |
-
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 179 |
},
|
| 180 |
{
|
| 181 |
-
"task_key": "
|
| 182 |
-
"
|
| 183 |
-
"
|
| 184 |
"codex_success": false,
|
| 185 |
-
"kinex_success":
|
| 186 |
-
"
|
| 187 |
-
|
| 188 |
-
|
| 189 |
-
|
| 190 |
-
|
| 191 |
-
|
| 192 |
-
|
| 193 |
-
|
| 194 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 195 |
},
|
| 196 |
{
|
| 197 |
-
"task_key": "
|
| 198 |
-
"
|
| 199 |
-
"
|
| 200 |
"codex_success": false,
|
| 201 |
-
"kinex_success":
|
| 202 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 203 |
},
|
| 204 |
{
|
| 205 |
-
"task_key": "
|
| 206 |
-
"
|
| 207 |
-
"
|
| 208 |
-
"codex_success":
|
| 209 |
"kinex_success": true,
|
| 210 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 211 |
},
|
| 212 |
{
|
| 213 |
-
"task_key": "
|
| 214 |
-
"
|
| 215 |
-
"
|
| 216 |
"codex_success": false,
|
| 217 |
"kinex_success": false,
|
| 218 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 219 |
},
|
| 220 |
{
|
| 221 |
-
"task_key": "
|
| 222 |
-
"
|
| 223 |
-
"
|
| 224 |
"codex_success": true,
|
| 225 |
-
"kinex_success":
|
| 226 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 227 |
},
|
| 228 |
{
|
| 229 |
-
"task_key": "
|
| 230 |
-
"
|
| 231 |
-
"
|
| 232 |
"codex_success": true,
|
| 233 |
"kinex_success": true,
|
| 234 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 235 |
},
|
| 236 |
{
|
| 237 |
-
"task_key": "
|
| 238 |
-
"
|
| 239 |
-
"
|
| 240 |
"codex_success": true,
|
| 241 |
"kinex_success": true,
|
| 242 |
-
"
|
| 243 |
-
|
| 244 |
-
|
| 245 |
-
|
| 246 |
-
|
| 247 |
-
|
| 248 |
-
|
| 249 |
-
|
| 250 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 251 |
},
|
| 252 |
{
|
| 253 |
-
"task_key": "
|
| 254 |
-
"
|
| 255 |
-
"
|
| 256 |
"codex_success": true,
|
| 257 |
"kinex_success": true,
|
| 258 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 259 |
},
|
| 260 |
{
|
| 261 |
-
"task_key": "
|
| 262 |
-
"
|
| 263 |
-
"
|
| 264 |
-
"codex_success": true,
|
| 265 |
-
"kinex_success": true,
|
| 266 |
-
"kinex_version": "0.10.3"
|
| 267 |
-
},
|
| 268 |
-
{
|
| 269 |
-
"task_key": "task04/39",
|
| 270 |
-
"native_id": "robodojo/push-t",
|
| 271 |
-
"status": "completed",
|
| 272 |
"codex_success": false,
|
| 273 |
"kinex_success": true,
|
| 274 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 275 |
},
|
| 276 |
{
|
| 277 |
-
"task_key": "
|
| 278 |
-
"
|
| 279 |
-
"
|
| 280 |
"codex_success": true,
|
| 281 |
"kinex_success": true,
|
| 282 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 283 |
},
|
| 284 |
{
|
| 285 |
-
"task_key": "
|
| 286 |
-
"
|
| 287 |
-
"
|
| 288 |
"codex_success": true,
|
| 289 |
"kinex_success": true,
|
| 290 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 291 |
},
|
| 292 |
{
|
| 293 |
-
"task_key": "
|
| 294 |
-
"
|
| 295 |
-
"
|
| 296 |
"codex_success": true,
|
| 297 |
"kinex_success": true,
|
| 298 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 299 |
},
|
| 300 |
{
|
| 301 |
-
"task_key": "
|
| 302 |
-
"
|
| 303 |
-
"
|
| 304 |
"codex_success": true,
|
| 305 |
"kinex_success": true,
|
| 306 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 307 |
},
|
| 308 |
{
|
| 309 |
-
"task_key": "
|
| 310 |
-
"
|
| 311 |
-
"
|
| 312 |
"codex_success": true,
|
| 313 |
"kinex_success": true,
|
| 314 |
-
"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 315 |
},
|
| 316 |
{
|
| 317 |
-
"task_key": "
|
| 318 |
-
"
|
| 319 |
-
"
|
| 320 |
"codex_success": true,
|
| 321 |
"kinex_success": true,
|
| 322 |
-
"
|
| 323 |
-
|
| 324 |
-
|
| 325 |
-
|
| 326 |
-
|
| 327 |
-
|
| 328 |
-
|
| 329 |
-
|
| 330 |
-
|
| 331 |
-
|
| 332 |
-
|
| 333 |
-
|
| 334 |
-
|
| 335 |
-
|
| 336 |
-
|
| 337 |
-
|
| 338 |
-
|
| 339 |
-
|
| 340 |
-
|
| 341 |
-
|
| 342 |
-
|
| 343 |
-
|
| 344 |
-
|
| 345 |
-
|
| 346 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 347 |
}
|
| 348 |
],
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 349 |
"limitations": [
|
| 350 |
-
"
|
| 351 |
-
"
|
| 352 |
-
"
|
|
|
|
|
|
|
| 353 |
]
|
| 354 |
}
|
|
|
|
| 1 |
{
|
| 2 |
+
"reference_repo": "RLE-Bench/kinex-v0.10.0-benchmark",
|
| 3 |
+
"reference_delivery_time": "2026-10-09T17:52:51.393664+00:00",
|
| 4 |
+
"reference_revisions": {
|
| 5 |
+
"dataset": "07dd924831cf10515ccda0a05d4a8ce169df7965",
|
| 6 |
+
"space": "83483dadd2a7140a4f23bd247ff91183444b1da5"
|
|
|
|
|
|
|
|
|
|
| 7 |
},
|
| 8 |
+
"families": [
|
| 9 |
{
|
| 10 |
+
"family": "task01",
|
| 11 |
+
"tasks": 10,
|
| 12 |
+
"codex_successes": 2,
|
| 13 |
+
"kinex_successes": 1,
|
| 14 |
+
"codex_usage": {
|
| 15 |
+
"input_tokens": 63285354,
|
| 16 |
+
"cached_input_tokens": 62436736,
|
| 17 |
+
"uncached_input_tokens": 848618,
|
| 18 |
+
"output_tokens": 204733,
|
| 19 |
+
"reasoning_output_tokens": 78505,
|
| 20 |
+
"response_count": 1112,
|
| 21 |
+
"cost_usd": null,
|
| 22 |
+
"attempts": 10,
|
| 23 |
+
"reported_attempts": {
|
| 24 |
+
"input_tokens": 10,
|
| 25 |
+
"cached_input_tokens": 10,
|
| 26 |
+
"uncached_input_tokens": 10,
|
| 27 |
+
"output_tokens": 10,
|
| 28 |
+
"reasoning_output_tokens": 10,
|
| 29 |
+
"response_count": 10,
|
| 30 |
+
"cost_usd": 0
|
| 31 |
+
},
|
| 32 |
+
"audit_complete_attempts": 10,
|
| 33 |
+
"cache_hit_rate": 0.9865906098905601,
|
| 34 |
+
"tokens_per_task": {
|
| 35 |
+
"input_tokens": 6328535.4,
|
| 36 |
+
"cached_input_tokens": 6243673.6,
|
| 37 |
+
"uncached_input_tokens": 84861.8,
|
| 38 |
+
"output_tokens": 20473.3,
|
| 39 |
+
"reasoning_output_tokens": 7850.5
|
| 40 |
+
}
|
| 41 |
+
},
|
| 42 |
+
"kinex_usage": {
|
| 43 |
+
"input_tokens": 120363624,
|
| 44 |
+
"cached_input_tokens": 118975104,
|
| 45 |
+
"uncached_input_tokens": 1388520,
|
| 46 |
+
"output_tokens": 247504,
|
| 47 |
+
"reasoning_output_tokens": 98713,
|
| 48 |
+
"response_count": 2003,
|
| 49 |
+
"cost_usd": null,
|
| 50 |
+
"attempts": 10,
|
| 51 |
+
"reported_attempts": {
|
| 52 |
+
"input_tokens": 10,
|
| 53 |
+
"cached_input_tokens": 10,
|
| 54 |
+
"uncached_input_tokens": 10,
|
| 55 |
+
"output_tokens": 10,
|
| 56 |
+
"reasoning_output_tokens": 10,
|
| 57 |
+
"response_count": 10,
|
| 58 |
+
"cost_usd": 0
|
| 59 |
+
},
|
| 60 |
+
"audit_complete_attempts": 10,
|
| 61 |
+
"cache_hit_rate": 0.9884639565189562,
|
| 62 |
+
"tokens_per_task": {
|
| 63 |
+
"input_tokens": 12036362.4,
|
| 64 |
+
"cached_input_tokens": 11897510.4,
|
| 65 |
+
"uncached_input_tokens": 138852.0,
|
| 66 |
+
"output_tokens": 24750.4,
|
| 67 |
+
"reasoning_output_tokens": 9871.3
|
| 68 |
+
}
|
| 69 |
+
},
|
| 70 |
+
"codex_wall_time": {
|
| 71 |
+
"reported_attempts": 10,
|
| 72 |
+
"sum_s": 11720.851507,
|
| 73 |
+
"mean_s": 1172.0851507,
|
| 74 |
+
"median_s": 1239.5881319999999
|
| 75 |
+
},
|
| 76 |
+
"kinex_wall_time": {
|
| 77 |
+
"reported_attempts": 10,
|
| 78 |
+
"sum_s": 21798.805297,
|
| 79 |
+
"mean_s": 2179.8805297,
|
| 80 |
+
"median_s": 2131.8586855000003
|
| 81 |
+
}
|
| 82 |
},
|
| 83 |
{
|
| 84 |
+
"family": "task02",
|
| 85 |
+
"tasks": 10,
|
| 86 |
+
"codex_successes": 9,
|
| 87 |
+
"kinex_successes": 10,
|
| 88 |
+
"codex_usage": {
|
| 89 |
+
"input_tokens": 12586499,
|
| 90 |
+
"cached_input_tokens": 12106752,
|
| 91 |
+
"uncached_input_tokens": 479747,
|
| 92 |
+
"output_tokens": 93969,
|
| 93 |
+
"reasoning_output_tokens": 34987,
|
| 94 |
+
"response_count": 344,
|
| 95 |
+
"cost_usd": null,
|
| 96 |
+
"attempts": 10,
|
| 97 |
+
"reported_attempts": {
|
| 98 |
+
"input_tokens": 10,
|
| 99 |
+
"cached_input_tokens": 10,
|
| 100 |
+
"uncached_input_tokens": 10,
|
| 101 |
+
"output_tokens": 10,
|
| 102 |
+
"reasoning_output_tokens": 10,
|
| 103 |
+
"response_count": 10,
|
| 104 |
+
"cost_usd": 0
|
| 105 |
+
},
|
| 106 |
+
"audit_complete_attempts": 10,
|
| 107 |
+
"cache_hit_rate": 0.9618839996729829,
|
| 108 |
+
"tokens_per_task": {
|
| 109 |
+
"input_tokens": 1258649.9,
|
| 110 |
+
"cached_input_tokens": 1210675.2,
|
| 111 |
+
"uncached_input_tokens": 47974.7,
|
| 112 |
+
"output_tokens": 9396.9,
|
| 113 |
+
"reasoning_output_tokens": 3498.7
|
| 114 |
+
}
|
| 115 |
+
},
|
| 116 |
+
"kinex_usage": {
|
| 117 |
+
"input_tokens": 17966387,
|
| 118 |
+
"cached_input_tokens": 17024128,
|
| 119 |
+
"uncached_input_tokens": 942259,
|
| 120 |
+
"output_tokens": 83439,
|
| 121 |
+
"reasoning_output_tokens": 26298,
|
| 122 |
+
"response_count": 599,
|
| 123 |
+
"cost_usd": null,
|
| 124 |
+
"attempts": 10,
|
| 125 |
+
"reported_attempts": {
|
| 126 |
+
"input_tokens": 10,
|
| 127 |
+
"cached_input_tokens": 10,
|
| 128 |
+
"uncached_input_tokens": 10,
|
| 129 |
+
"output_tokens": 10,
|
| 130 |
+
"reasoning_output_tokens": 10,
|
| 131 |
+
"response_count": 10,
|
| 132 |
+
"cost_usd": 0
|
| 133 |
+
},
|
| 134 |
+
"audit_complete_attempts": 10,
|
| 135 |
+
"cache_hit_rate": 0.947554341337521,
|
| 136 |
+
"tokens_per_task": {
|
| 137 |
+
"input_tokens": 1796638.7,
|
| 138 |
+
"cached_input_tokens": 1702412.8,
|
| 139 |
+
"uncached_input_tokens": 94225.9,
|
| 140 |
+
"output_tokens": 8343.9,
|
| 141 |
+
"reasoning_output_tokens": 2629.8
|
| 142 |
+
}
|
| 143 |
+
},
|
| 144 |
+
"codex_wall_time": {
|
| 145 |
+
"reported_attempts": 10,
|
| 146 |
+
"sum_s": 3681.418654,
|
| 147 |
+
"mean_s": 368.14186540000003,
|
| 148 |
+
"median_s": 266.88804849999997
|
| 149 |
+
},
|
| 150 |
+
"kinex_wall_time": {
|
| 151 |
+
"reported_attempts": 10,
|
| 152 |
+
"sum_s": 4811.317263,
|
| 153 |
+
"mean_s": 481.1317263,
|
| 154 |
+
"median_s": 409.560245
|
| 155 |
+
}
|
| 156 |
+
}
|
| 157 |
+
],
|
| 158 |
+
"tasks": [
|
| 159 |
{
|
| 160 |
+
"task_key": "task01/01",
|
| 161 |
+
"instruction": "Place the specified condiments from the counter on the top shelf of the fridge. Move any existing top-shelf items to other shelves. Release the items and move the gripper clear of the stored items.",
|
| 162 |
+
"instruction_policy": "modified",
|
| 163 |
"codex_success": false,
|
| 164 |
"kinex_success": false,
|
| 165 |
+
"codex_steps": 5496,
|
| 166 |
+
"kinex_steps": 3704,
|
| 167 |
+
"codex_usage": {
|
| 168 |
+
"accounting": "reported-responses",
|
| 169 |
+
"audit_complete": true,
|
| 170 |
+
"cache_hit_rate": 0.9875774826559943,
|
| 171 |
+
"cache_reported_input_tokens": 7077229,
|
| 172 |
+
"cache_write_input_tokens": 0,
|
| 173 |
+
"cache_write_reported_input_tokens": 7077229,
|
| 174 |
+
"cached_input_tokens": 6989312,
|
| 175 |
+
"completed_turns": 1,
|
| 176 |
+
"cost_usd": null,
|
| 177 |
+
"failed_turns": 0,
|
| 178 |
+
"input_tokens": 7077229,
|
| 179 |
+
"known_cache_write_input_tokens": 0,
|
| 180 |
+
"known_cached_input_tokens": 6989312,
|
| 181 |
+
"known_input_tokens": 7077229,
|
| 182 |
+
"known_output_tokens": 20496,
|
| 183 |
+
"known_reasoning_output_tokens": 5873,
|
| 184 |
+
"output_tokens": 20496,
|
| 185 |
+
"reasoning_output_tokens": 5873,
|
| 186 |
+
"reasoning_reported_output_tokens": 20496,
|
| 187 |
+
"reported_responses": {
|
| 188 |
+
"cache_reported_input_tokens": 140,
|
| 189 |
+
"cache_write_input_tokens": 140,
|
| 190 |
+
"cache_write_reported_input_tokens": 140,
|
| 191 |
+
"cached_input_tokens": 140,
|
| 192 |
+
"input_tokens": 140,
|
| 193 |
+
"output_tokens": 140,
|
| 194 |
+
"reasoning_output_tokens": 140,
|
| 195 |
+
"reasoning_reported_output_tokens": 140
|
| 196 |
+
},
|
| 197 |
+
"response_count": 140,
|
| 198 |
+
"response_ids_complete": true,
|
| 199 |
+
"schema": "rlebench/token-usage/1",
|
| 200 |
+
"source": "Codex token_usage_record per response",
|
| 201 |
+
"uncached_input_tokens": 87917,
|
| 202 |
+
"unidentified_usage_records": 0
|
| 203 |
+
},
|
| 204 |
+
"kinex_usage": {
|
| 205 |
+
"input_tokens": 9503123,
|
| 206 |
+
"output_tokens": 23563,
|
| 207 |
+
"cached_input_tokens": 9397760,
|
| 208 |
+
"cache_write_input_tokens": 0,
|
| 209 |
+
"reasoning_output_tokens": 8032,
|
| 210 |
+
"cache_reported_input_tokens": 9503123,
|
| 211 |
+
"cache_write_reported_input_tokens": 9503123,
|
| 212 |
+
"reasoning_reported_output_tokens": 23563,
|
| 213 |
+
"schema": "rlebench/token-usage/1",
|
| 214 |
+
"source": "Provider terminal response usage",
|
| 215 |
+
"response_count": 191,
|
| 216 |
+
"reported_responses": {
|
| 217 |
+
"input_tokens": 191,
|
| 218 |
+
"output_tokens": 191,
|
| 219 |
+
"cached_input_tokens": 191,
|
| 220 |
+
"cache_write_input_tokens": 191,
|
| 221 |
+
"reasoning_output_tokens": 191,
|
| 222 |
+
"cache_reported_input_tokens": 191,
|
| 223 |
+
"cache_write_reported_input_tokens": 191,
|
| 224 |
+
"reasoning_reported_output_tokens": 191
|
| 225 |
+
},
|
| 226 |
+
"known_cached_input_tokens": 9397760,
|
| 227 |
+
"known_input_tokens": 9503123,
|
| 228 |
+
"known_output_tokens": 23563,
|
| 229 |
+
"known_cache_write_input_tokens": 0,
|
| 230 |
+
"known_reasoning_output_tokens": 8032,
|
| 231 |
+
"cache_hit_rate": 0.9889128026649766,
|
| 232 |
+
"uncached_input_tokens": 105363,
|
| 233 |
+
"cost_usd": null,
|
| 234 |
+
"response_ids_complete": true,
|
| 235 |
+
"usage_source": "provider-raw",
|
| 236 |
+
"request_attempts": 192,
|
| 237 |
+
"rejected_attempts": 0,
|
| 238 |
+
"unresolved_attempts": [],
|
| 239 |
+
"audit_complete": true,
|
| 240 |
+
"accounting": "reported-responses"
|
| 241 |
+
},
|
| 242 |
+
"codex_wall_time_s": 1267.271401,
|
| 243 |
+
"kinex_wall_time_s": 1805.038681,
|
| 244 |
+
"kinex_version": "0.10.3",
|
| 245 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 246 |
+
"same_roboenv_source": true
|
| 247 |
},
|
| 248 |
{
|
| 249 |
+
"task_key": "task01/02",
|
| 250 |
+
"instruction": "Remove the specified fruit from the bowl and place it on the small plate. Then place the bowl with only the specified meat in the microwave, close the door, and press the start button to microwave the meat. Release the bowl and move the gripper clear of it.",
|
| 251 |
+
"instruction_policy": "modified",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 252 |
"codex_success": false,
|
| 253 |
"kinex_success": false,
|
| 254 |
+
"codex_steps": 5903,
|
| 255 |
+
"kinex_steps": 5048,
|
| 256 |
+
"codex_usage": {
|
| 257 |
+
"accounting": "reported-responses",
|
| 258 |
+
"audit_complete": true,
|
| 259 |
+
"cache_hit_rate": 0.9810887797632594,
|
| 260 |
+
"cache_reported_input_tokens": 4438106,
|
| 261 |
+
"cache_write_input_tokens": 0,
|
| 262 |
+
"cache_write_reported_input_tokens": 4438106,
|
| 263 |
+
"cached_input_tokens": 4354176,
|
| 264 |
+
"completed_turns": 1,
|
| 265 |
+
"cost_usd": null,
|
| 266 |
+
"failed_turns": 0,
|
| 267 |
+
"input_tokens": 4438106,
|
| 268 |
+
"known_cache_write_input_tokens": 0,
|
| 269 |
+
"known_cached_input_tokens": 4354176,
|
| 270 |
+
"known_input_tokens": 4438106,
|
| 271 |
+
"known_output_tokens": 22048,
|
| 272 |
+
"known_reasoning_output_tokens": 9003,
|
| 273 |
+
"output_tokens": 22048,
|
| 274 |
+
"reasoning_output_tokens": 9003,
|
| 275 |
+
"reasoning_reported_output_tokens": 22048,
|
| 276 |
+
"reported_responses": {
|
| 277 |
+
"cache_reported_input_tokens": 86,
|
| 278 |
+
"cache_write_input_tokens": 86,
|
| 279 |
+
"cache_write_reported_input_tokens": 86,
|
| 280 |
+
"cached_input_tokens": 86,
|
| 281 |
+
"input_tokens": 86,
|
| 282 |
+
"output_tokens": 86,
|
| 283 |
+
"reasoning_output_tokens": 86,
|
| 284 |
+
"reasoning_reported_output_tokens": 86
|
| 285 |
+
},
|
| 286 |
+
"response_count": 86,
|
| 287 |
+
"response_ids_complete": true,
|
| 288 |
+
"schema": "rlebench/token-usage/1",
|
| 289 |
+
"source": "Codex token_usage_record per response",
|
| 290 |
+
"uncached_input_tokens": 83930,
|
| 291 |
+
"unidentified_usage_records": 0
|
| 292 |
+
},
|
| 293 |
+
"kinex_usage": {
|
| 294 |
+
"input_tokens": 13113774,
|
| 295 |
+
"output_tokens": 31080,
|
| 296 |
+
"cached_input_tokens": 12981632,
|
| 297 |
+
"cache_write_input_tokens": 0,
|
| 298 |
+
"reasoning_output_tokens": 13344,
|
| 299 |
+
"cache_reported_input_tokens": 13113774,
|
| 300 |
+
"cache_write_reported_input_tokens": 13113774,
|
| 301 |
+
"reasoning_reported_output_tokens": 31080,
|
| 302 |
+
"schema": "rlebench/token-usage/1",
|
| 303 |
+
"source": "Provider terminal response usage",
|
| 304 |
+
"response_count": 214,
|
| 305 |
+
"reported_responses": {
|
| 306 |
+
"input_tokens": 214,
|
| 307 |
+
"output_tokens": 214,
|
| 308 |
+
"cached_input_tokens": 214,
|
| 309 |
+
"cache_write_input_tokens": 214,
|
| 310 |
+
"reasoning_output_tokens": 214,
|
| 311 |
+
"cache_reported_input_tokens": 214,
|
| 312 |
+
"cache_write_reported_input_tokens": 214,
|
| 313 |
+
"reasoning_reported_output_tokens": 214
|
| 314 |
+
},
|
| 315 |
+
"known_cached_input_tokens": 12981632,
|
| 316 |
+
"known_input_tokens": 13113774,
|
| 317 |
+
"known_output_tokens": 31080,
|
| 318 |
+
"known_cache_write_input_tokens": 0,
|
| 319 |
+
"known_reasoning_output_tokens": 13344,
|
| 320 |
+
"cache_hit_rate": 0.9899234194519442,
|
| 321 |
+
"uncached_input_tokens": 132142,
|
| 322 |
+
"cost_usd": null,
|
| 323 |
+
"response_ids_complete": true,
|
| 324 |
+
"usage_source": "provider-raw",
|
| 325 |
+
"request_attempts": 215,
|
| 326 |
+
"rejected_attempts": 0,
|
| 327 |
+
"unresolved_attempts": [],
|
| 328 |
+
"audit_complete": true,
|
| 329 |
+
"accounting": "reported-responses"
|
| 330 |
+
},
|
| 331 |
+
"codex_wall_time_s": 1211.904863,
|
| 332 |
+
"kinex_wall_time_s": 2458.67869,
|
| 333 |
+
"kinex_version": "0.10.3",
|
| 334 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 335 |
+
"same_roboenv_source": true
|
| 336 |
},
|
| 337 |
{
|
| 338 |
+
"task_key": "task01/03",
|
| 339 |
+
"instruction": "Place two dumplings into each of the tupperware containers and then place both containers on a shelf in the fridge. Release the dumplings and containers and move the gripper clear of them.",
|
| 340 |
+
"instruction_policy": "modified",
|
| 341 |
+
"codex_success": false,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 342 |
"kinex_success": false,
|
| 343 |
+
"codex_steps": 5053,
|
| 344 |
+
"kinex_steps": 2809,
|
| 345 |
+
"codex_usage": {
|
| 346 |
+
"accounting": "reported-responses",
|
| 347 |
+
"audit_complete": true,
|
| 348 |
+
"cache_hit_rate": 0.9899699443798293,
|
| 349 |
+
"cache_reported_input_tokens": 11115392,
|
| 350 |
+
"cache_write_input_tokens": 0,
|
| 351 |
+
"cache_write_reported_input_tokens": 11115392,
|
| 352 |
+
"cached_input_tokens": 11003904,
|
| 353 |
+
"completed_turns": 1,
|
| 354 |
+
"cost_usd": null,
|
| 355 |
+
"failed_turns": 0,
|
| 356 |
+
"input_tokens": 11115392,
|
| 357 |
+
"known_cache_write_input_tokens": 0,
|
| 358 |
+
"known_cached_input_tokens": 11003904,
|
| 359 |
+
"known_input_tokens": 11115392,
|
| 360 |
+
"known_output_tokens": 29598,
|
| 361 |
+
"known_reasoning_output_tokens": 14166,
|
| 362 |
+
"output_tokens": 29598,
|
| 363 |
+
"reasoning_output_tokens": 14166,
|
| 364 |
+
"reasoning_reported_output_tokens": 29598,
|
| 365 |
+
"reported_responses": {
|
| 366 |
+
"cache_reported_input_tokens": 171,
|
| 367 |
+
"cache_write_input_tokens": 171,
|
| 368 |
+
"cache_write_reported_input_tokens": 171,
|
| 369 |
+
"cached_input_tokens": 171,
|
| 370 |
+
"input_tokens": 171,
|
| 371 |
+
"output_tokens": 171,
|
| 372 |
+
"reasoning_output_tokens": 171,
|
| 373 |
+
"reasoning_reported_output_tokens": 171
|
| 374 |
+
},
|
| 375 |
+
"response_count": 171,
|
| 376 |
+
"response_ids_complete": true,
|
| 377 |
+
"schema": "rlebench/token-usage/1",
|
| 378 |
+
"source": "Codex token_usage_record per response",
|
| 379 |
+
"uncached_input_tokens": 111488,
|
| 380 |
+
"unidentified_usage_records": 0
|
| 381 |
+
},
|
| 382 |
+
"kinex_usage": {
|
| 383 |
+
"input_tokens": 6179963,
|
| 384 |
+
"output_tokens": 17710,
|
| 385 |
+
"cached_input_tokens": 6033792,
|
| 386 |
+
"cache_write_input_tokens": 0,
|
| 387 |
+
"reasoning_output_tokens": 5948,
|
| 388 |
+
"cache_reported_input_tokens": 6179963,
|
| 389 |
+
"cache_write_reported_input_tokens": 6179963,
|
| 390 |
+
"reasoning_reported_output_tokens": 17710,
|
| 391 |
+
"schema": "rlebench/token-usage/1",
|
| 392 |
+
"source": "Provider terminal response usage",
|
| 393 |
+
"response_count": 135,
|
| 394 |
+
"reported_responses": {
|
| 395 |
+
"input_tokens": 135,
|
| 396 |
+
"output_tokens": 135,
|
| 397 |
+
"cached_input_tokens": 135,
|
| 398 |
+
"cache_write_input_tokens": 135,
|
| 399 |
+
"reasoning_output_tokens": 135,
|
| 400 |
+
"cache_reported_input_tokens": 135,
|
| 401 |
+
"cache_write_reported_input_tokens": 135,
|
| 402 |
+
"reasoning_reported_output_tokens": 135
|
| 403 |
+
},
|
| 404 |
+
"known_cached_input_tokens": 6033792,
|
| 405 |
+
"known_input_tokens": 6179963,
|
| 406 |
+
"known_output_tokens": 17710,
|
| 407 |
+
"known_cache_write_input_tokens": 0,
|
| 408 |
+
"known_reasoning_output_tokens": 5948,
|
| 409 |
+
"cache_hit_rate": 0.9763475930195699,
|
| 410 |
+
"uncached_input_tokens": 146171,
|
| 411 |
+
"cost_usd": null,
|
| 412 |
+
"response_ids_complete": true,
|
| 413 |
+
"usage_source": "provider-raw",
|
| 414 |
+
"request_attempts": 135,
|
| 415 |
+
"rejected_attempts": 0,
|
| 416 |
+
"unresolved_attempts": [],
|
| 417 |
+
"audit_complete": true,
|
| 418 |
+
"accounting": "reported-responses"
|
| 419 |
+
},
|
| 420 |
+
"codex_wall_time_s": 1749.701956,
|
| 421 |
+
"kinex_wall_time_s": 1444.988461,
|
| 422 |
+
"kinex_version": "0.10.3",
|
| 423 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 424 |
+
"same_roboenv_source": true
|
| 425 |
},
|
| 426 |
{
|
| 427 |
+
"task_key": "task01/04",
|
| 428 |
+
"instruction": "Gather the specified vegetables from the fridge and place them on a tray on the dining counter. Then gather the specified meats from the fridge and place them on the other tray. Release the food and move the gripper clear of the food and both trays.",
|
| 429 |
+
"instruction_policy": "modified",
|
| 430 |
"codex_success": false,
|
| 431 |
"kinex_success": false,
|
| 432 |
+
"codex_steps": 5435,
|
| 433 |
+
"kinex_steps": 5449,
|
| 434 |
+
"codex_usage": {
|
| 435 |
+
"accounting": "reported-responses",
|
| 436 |
+
"audit_complete": true,
|
| 437 |
+
"cache_hit_rate": 0.9899239913597869,
|
| 438 |
+
"cache_reported_input_tokens": 11714063,
|
| 439 |
+
"cache_write_input_tokens": 0,
|
| 440 |
+
"cache_write_reported_input_tokens": 11714063,
|
| 441 |
+
"cached_input_tokens": 11596032,
|
| 442 |
+
"completed_turns": 1,
|
| 443 |
+
"cost_usd": null,
|
| 444 |
+
"failed_turns": 0,
|
| 445 |
+
"input_tokens": 11714063,
|
| 446 |
+
"known_cache_write_input_tokens": 0,
|
| 447 |
+
"known_cached_input_tokens": 11596032,
|
| 448 |
+
"known_input_tokens": 11714063,
|
| 449 |
+
"known_output_tokens": 32111,
|
| 450 |
+
"known_reasoning_output_tokens": 11814,
|
| 451 |
+
"output_tokens": 32111,
|
| 452 |
+
"reasoning_output_tokens": 11814,
|
| 453 |
+
"reasoning_reported_output_tokens": 32111,
|
| 454 |
+
"reported_responses": {
|
| 455 |
+
"cache_reported_input_tokens": 190,
|
| 456 |
+
"cache_write_input_tokens": 190,
|
| 457 |
+
"cache_write_reported_input_tokens": 190,
|
| 458 |
+
"cached_input_tokens": 190,
|
| 459 |
+
"input_tokens": 190,
|
| 460 |
+
"output_tokens": 190,
|
| 461 |
+
"reasoning_output_tokens": 190,
|
| 462 |
+
"reasoning_reported_output_tokens": 190
|
| 463 |
+
},
|
| 464 |
+
"response_count": 190,
|
| 465 |
+
"response_ids_complete": true,
|
| 466 |
+
"schema": "rlebench/token-usage/1",
|
| 467 |
+
"source": "Codex token_usage_record per response",
|
| 468 |
+
"uncached_input_tokens": 118031,
|
| 469 |
+
"unidentified_usage_records": 0
|
| 470 |
+
},
|
| 471 |
+
"kinex_usage": {
|
| 472 |
+
"input_tokens": 18456926,
|
| 473 |
+
"output_tokens": 44833,
|
| 474 |
+
"cached_input_tokens": 18169984,
|
| 475 |
+
"cache_write_input_tokens": 0,
|
| 476 |
+
"reasoning_output_tokens": 22099,
|
| 477 |
+
"cache_reported_input_tokens": 18456926,
|
| 478 |
+
"cache_write_reported_input_tokens": 18456926,
|
| 479 |
+
"reasoning_reported_output_tokens": 44833,
|
| 480 |
+
"schema": "rlebench/token-usage/1",
|
| 481 |
+
"source": "Provider terminal response usage",
|
| 482 |
+
"response_count": 236,
|
| 483 |
+
"reported_responses": {
|
| 484 |
+
"input_tokens": 236,
|
| 485 |
+
"output_tokens": 236,
|
| 486 |
+
"cached_input_tokens": 236,
|
| 487 |
+
"cache_write_input_tokens": 236,
|
| 488 |
+
"reasoning_output_tokens": 236,
|
| 489 |
+
"cache_reported_input_tokens": 236,
|
| 490 |
+
"cache_write_reported_input_tokens": 236,
|
| 491 |
+
"reasoning_reported_output_tokens": 236
|
| 492 |
+
},
|
| 493 |
+
"known_cached_input_tokens": 18169984,
|
| 494 |
+
"known_input_tokens": 18456926,
|
| 495 |
+
"known_output_tokens": 44833,
|
| 496 |
+
"known_cache_write_input_tokens": 0,
|
| 497 |
+
"known_reasoning_output_tokens": 22099,
|
| 498 |
+
"cache_hit_rate": 0.9844534241509122,
|
| 499 |
+
"uncached_input_tokens": 286942,
|
| 500 |
+
"cost_usd": null,
|
| 501 |
+
"response_ids_complete": true,
|
| 502 |
+
"usage_source": "provider-raw",
|
| 503 |
+
"request_attempts": 237,
|
| 504 |
+
"rejected_attempts": 0,
|
| 505 |
+
"unresolved_attempts": [],
|
| 506 |
+
"audit_complete": true,
|
| 507 |
+
"accounting": "reported-responses"
|
| 508 |
+
},
|
| 509 |
+
"codex_wall_time_s": 1973.915146,
|
| 510 |
+
"kinex_wall_time_s": 3231.211091,
|
| 511 |
+
"kinex_version": "0.10.3",
|
| 512 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 513 |
+
"same_roboenv_source": true
|
| 514 |
},
|
| 515 |
{
|
| 516 |
+
"task_key": "task01/05",
|
| 517 |
+
"instruction": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer bowl, lower the mixer head fully, and turn the speed knob to begin making cheesecake filling.",
|
| 518 |
+
"instruction_policy": "modified",
|
| 519 |
"codex_success": true,
|
| 520 |
+
"kinex_success": false,
|
| 521 |
+
"codex_steps": 5224,
|
| 522 |
+
"kinex_steps": 5930,
|
| 523 |
+
"codex_usage": {
|
| 524 |
+
"accounting": "reported-responses",
|
| 525 |
+
"audit_complete": true,
|
| 526 |
+
"cache_hit_rate": 0.9885663091343345,
|
| 527 |
+
"cache_reported_input_tokens": 12355153,
|
| 528 |
+
"cache_write_input_tokens": 0,
|
| 529 |
+
"cache_write_reported_input_tokens": 12355153,
|
| 530 |
+
"cached_input_tokens": 12213888,
|
| 531 |
+
"completed_turns": 1,
|
| 532 |
+
"cost_usd": null,
|
| 533 |
+
"failed_turns": 0,
|
| 534 |
+
"input_tokens": 12355153,
|
| 535 |
+
"known_cache_write_input_tokens": 0,
|
| 536 |
+
"known_cached_input_tokens": 12213888,
|
| 537 |
+
"known_input_tokens": 12355153,
|
| 538 |
+
"known_output_tokens": 35229,
|
| 539 |
+
"known_reasoning_output_tokens": 13787,
|
| 540 |
+
"output_tokens": 35229,
|
| 541 |
+
"reasoning_output_tokens": 13787,
|
| 542 |
+
"reasoning_reported_output_tokens": 35229,
|
| 543 |
+
"reported_responses": {
|
| 544 |
+
"cache_reported_input_tokens": 202,
|
| 545 |
+
"cache_write_input_tokens": 202,
|
| 546 |
+
"cache_write_reported_input_tokens": 202,
|
| 547 |
+
"cached_input_tokens": 202,
|
| 548 |
+
"input_tokens": 202,
|
| 549 |
+
"output_tokens": 202,
|
| 550 |
+
"reasoning_output_tokens": 202,
|
| 551 |
+
"reasoning_reported_output_tokens": 202
|
| 552 |
+
},
|
| 553 |
+
"response_count": 202,
|
| 554 |
+
"response_ids_complete": true,
|
| 555 |
+
"schema": "rlebench/token-usage/1",
|
| 556 |
+
"source": "Codex token_usage_record per response",
|
| 557 |
+
"uncached_input_tokens": 141265,
|
| 558 |
+
"unidentified_usage_records": 0
|
| 559 |
+
},
|
| 560 |
+
"kinex_usage": {
|
| 561 |
+
"input_tokens": 10985080,
|
| 562 |
+
"output_tokens": 19054,
|
| 563 |
+
"cached_input_tokens": 10880128,
|
| 564 |
+
"cache_write_input_tokens": 0,
|
| 565 |
+
"reasoning_output_tokens": 6692,
|
| 566 |
+
"cache_reported_input_tokens": 10985080,
|
| 567 |
+
"cache_write_reported_input_tokens": 10985080,
|
| 568 |
+
"reasoning_reported_output_tokens": 19054,
|
| 569 |
+
"schema": "rlebench/token-usage/1",
|
| 570 |
+
"source": "Provider terminal response usage",
|
| 571 |
+
"response_count": 217,
|
| 572 |
+
"reported_responses": {
|
| 573 |
+
"input_tokens": 217,
|
| 574 |
+
"output_tokens": 217,
|
| 575 |
+
"cached_input_tokens": 217,
|
| 576 |
+
"cache_write_input_tokens": 217,
|
| 577 |
+
"reasoning_output_tokens": 217,
|
| 578 |
+
"cache_reported_input_tokens": 217,
|
| 579 |
+
"cache_write_reported_input_tokens": 217,
|
| 580 |
+
"reasoning_reported_output_tokens": 217
|
| 581 |
+
},
|
| 582 |
+
"known_cached_input_tokens": 10880128,
|
| 583 |
+
"known_input_tokens": 10985080,
|
| 584 |
+
"known_output_tokens": 19054,
|
| 585 |
+
"known_cache_write_input_tokens": 0,
|
| 586 |
+
"known_reasoning_output_tokens": 6692,
|
| 587 |
+
"cache_hit_rate": 0.9904459503253504,
|
| 588 |
+
"uncached_input_tokens": 104952,
|
| 589 |
+
"cost_usd": null,
|
| 590 |
+
"response_ids_complete": true,
|
| 591 |
+
"usage_source": "provider-raw",
|
| 592 |
+
"request_attempts": 217,
|
| 593 |
+
"rejected_attempts": 0,
|
| 594 |
+
"unresolved_attempts": [],
|
| 595 |
+
"audit_complete": true,
|
| 596 |
+
"accounting": "reported-responses"
|
| 597 |
+
},
|
| 598 |
+
"codex_wall_time_s": 1964.925029,
|
| 599 |
+
"kinex_wall_time_s": 2570.336196,
|
| 600 |
+
"kinex_version": "0.10.3",
|
| 601 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 602 |
+
"same_roboenv_source": true
|
| 603 |
},
|
| 604 |
{
|
| 605 |
+
"task_key": "task01/06",
|
| 606 |
+
"instruction": "Turn on the sink faucet. Move the specified vegetable from the counter into the sink while the water is running. Turn off the faucet, move the vegetable into the pot next to the stove, and move the pot to the specified burner.",
|
| 607 |
+
"instruction_policy": "modified",
|
| 608 |
"codex_success": false,
|
| 609 |
+
"kinex_success": false,
|
| 610 |
+
"codex_steps": 5416,
|
| 611 |
+
"kinex_steps": 5856,
|
| 612 |
+
"codex_usage": {
|
| 613 |
+
"accounting": "reported-responses",
|
| 614 |
+
"audit_complete": true,
|
| 615 |
+
"cache_hit_rate": 0.9880091550479709,
|
| 616 |
+
"cache_reported_input_tokens": 9630097,
|
| 617 |
+
"cache_write_input_tokens": 0,
|
| 618 |
+
"cache_write_reported_input_tokens": 9630097,
|
| 619 |
+
"cached_input_tokens": 9514624,
|
| 620 |
+
"completed_turns": 1,
|
| 621 |
+
"cost_usd": null,
|
| 622 |
+
"failed_turns": 0,
|
| 623 |
+
"input_tokens": 9630097,
|
| 624 |
+
"known_cache_write_input_tokens": 0,
|
| 625 |
+
"known_cached_input_tokens": 9514624,
|
| 626 |
+
"known_input_tokens": 9630097,
|
| 627 |
+
"known_output_tokens": 24897,
|
| 628 |
+
"known_reasoning_output_tokens": 10825,
|
| 629 |
+
"output_tokens": 24897,
|
| 630 |
+
"reasoning_output_tokens": 10825,
|
| 631 |
+
"reasoning_reported_output_tokens": 24897,
|
| 632 |
+
"reported_responses": {
|
| 633 |
+
"cache_reported_input_tokens": 145,
|
| 634 |
+
"cache_write_input_tokens": 145,
|
| 635 |
+
"cache_write_reported_input_tokens": 145,
|
| 636 |
+
"cached_input_tokens": 145,
|
| 637 |
+
"input_tokens": 145,
|
| 638 |
+
"output_tokens": 145,
|
| 639 |
+
"reasoning_output_tokens": 145,
|
| 640 |
+
"reasoning_reported_output_tokens": 145
|
| 641 |
+
},
|
| 642 |
+
"response_count": 145,
|
| 643 |
+
"response_ids_complete": true,
|
| 644 |
+
"schema": "rlebench/token-usage/1",
|
| 645 |
+
"source": "Codex token_usage_record per response",
|
| 646 |
+
"uncached_input_tokens": 115473,
|
| 647 |
+
"unidentified_usage_records": 0
|
| 648 |
+
},
|
| 649 |
+
"kinex_usage": {
|
| 650 |
+
"input_tokens": 26424116,
|
| 651 |
+
"output_tokens": 28463,
|
| 652 |
+
"cached_input_tokens": 26255616,
|
| 653 |
+
"cache_write_input_tokens": 0,
|
| 654 |
+
"reasoning_output_tokens": 11086,
|
| 655 |
+
"cache_reported_input_tokens": 26424116,
|
| 656 |
+
"cache_write_reported_input_tokens": 26424116,
|
| 657 |
+
"reasoning_reported_output_tokens": 28463,
|
| 658 |
+
"schema": "rlebench/token-usage/1",
|
| 659 |
+
"source": "Provider terminal response usage",
|
| 660 |
+
"response_count": 347,
|
| 661 |
+
"reported_responses": {
|
| 662 |
+
"input_tokens": 347,
|
| 663 |
+
"output_tokens": 347,
|
| 664 |
+
"cached_input_tokens": 347,
|
| 665 |
+
"cache_write_input_tokens": 347,
|
| 666 |
+
"reasoning_output_tokens": 347,
|
| 667 |
+
"cache_reported_input_tokens": 347,
|
| 668 |
+
"cache_write_reported_input_tokens": 347,
|
| 669 |
+
"reasoning_reported_output_tokens": 347
|
| 670 |
+
},
|
| 671 |
+
"known_cached_input_tokens": 26255616,
|
| 672 |
+
"known_input_tokens": 26424116,
|
| 673 |
+
"known_output_tokens": 28463,
|
| 674 |
+
"known_cache_write_input_tokens": 0,
|
| 675 |
+
"known_reasoning_output_tokens": 11086,
|
| 676 |
+
"cache_hit_rate": 0.9936232493075643,
|
| 677 |
+
"uncached_input_tokens": 168500,
|
| 678 |
+
"cost_usd": null,
|
| 679 |
+
"response_ids_complete": true,
|
| 680 |
+
"usage_source": "provider-raw",
|
| 681 |
+
"request_attempts": 347,
|
| 682 |
+
"rejected_attempts": 0,
|
| 683 |
+
"unresolved_attempts": [],
|
| 684 |
+
"audit_complete": true,
|
| 685 |
+
"accounting": "reported-responses"
|
| 686 |
+
},
|
| 687 |
+
"codex_wall_time_s": 1456.598048,
|
| 688 |
+
"kinex_wall_time_s": 3438.331724,
|
| 689 |
+
"kinex_version": "0.10.3",
|
| 690 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 691 |
+
"same_roboenv_source": true
|
| 692 |
},
|
| 693 |
{
|
| 694 |
+
"task_key": "task01/07",
|
| 695 |
+
"instruction": "Take the specified meat from the fridge and place it on the digital scale on the counter by the fridge. Release it and move the gripper clear while waiting a few seconds for a reading. Then move it to the plate on the dining counter, release it, and move the gripper clear.",
|
| 696 |
+
"instruction_policy": "modified",
|
| 697 |
"codex_success": false,
|
| 698 |
+
"kinex_success": false,
|
| 699 |
+
"codex_steps": 1426,
|
| 700 |
+
"kinex_steps": 2427,
|
| 701 |
+
"codex_usage": {
|
| 702 |
+
"accounting": "reported-responses",
|
| 703 |
+
"audit_complete": true,
|
| 704 |
+
"cache_hit_rate": 0.9663948735854094,
|
| 705 |
+
"cache_reported_input_tokens": 1217225,
|
| 706 |
+
"cache_write_input_tokens": 0,
|
| 707 |
+
"cache_write_reported_input_tokens": 1217225,
|
| 708 |
+
"cached_input_tokens": 1176320,
|
| 709 |
+
"completed_turns": 1,
|
| 710 |
+
"cost_usd": null,
|
| 711 |
+
"failed_turns": 0,
|
| 712 |
+
"input_tokens": 1217225,
|
| 713 |
+
"known_cache_write_input_tokens": 0,
|
| 714 |
+
"known_cached_input_tokens": 1176320,
|
| 715 |
+
"known_input_tokens": 1217225,
|
| 716 |
+
"known_output_tokens": 8600,
|
| 717 |
+
"known_reasoning_output_tokens": 2758,
|
| 718 |
+
"output_tokens": 8600,
|
| 719 |
+
"reasoning_output_tokens": 2758,
|
| 720 |
+
"reasoning_reported_output_tokens": 8600,
|
| 721 |
+
"reported_responses": {
|
| 722 |
+
"cache_reported_input_tokens": 36,
|
| 723 |
+
"cache_write_input_tokens": 36,
|
| 724 |
+
"cache_write_reported_input_tokens": 36,
|
| 725 |
+
"cached_input_tokens": 36,
|
| 726 |
+
"input_tokens": 36,
|
| 727 |
+
"output_tokens": 36,
|
| 728 |
+
"reasoning_output_tokens": 36,
|
| 729 |
+
"reasoning_reported_output_tokens": 36
|
| 730 |
+
},
|
| 731 |
+
"response_count": 36,
|
| 732 |
+
"response_ids_complete": true,
|
| 733 |
+
"schema": "rlebench/token-usage/1",
|
| 734 |
+
"source": "Codex token_usage_record per response",
|
| 735 |
+
"uncached_input_tokens": 40905,
|
| 736 |
+
"unidentified_usage_records": 0
|
| 737 |
+
},
|
| 738 |
+
"kinex_usage": {
|
| 739 |
+
"input_tokens": 7005688,
|
| 740 |
+
"output_tokens": 18192,
|
| 741 |
+
"cached_input_tokens": 6921216,
|
| 742 |
+
"cache_write_input_tokens": 0,
|
| 743 |
+
"reasoning_output_tokens": 6177,
|
| 744 |
+
"cache_reported_input_tokens": 7005688,
|
| 745 |
+
"cache_write_reported_input_tokens": 7005688,
|
| 746 |
+
"reasoning_reported_output_tokens": 18192,
|
| 747 |
+
"schema": "rlebench/token-usage/1",
|
| 748 |
+
"source": "Provider terminal response usage",
|
| 749 |
+
"response_count": 153,
|
| 750 |
+
"reported_responses": {
|
| 751 |
+
"input_tokens": 153,
|
| 752 |
+
"output_tokens": 153,
|
| 753 |
+
"cached_input_tokens": 153,
|
| 754 |
+
"cache_write_input_tokens": 153,
|
| 755 |
+
"reasoning_output_tokens": 153,
|
| 756 |
+
"cache_reported_input_tokens": 153,
|
| 757 |
+
"cache_write_reported_input_tokens": 153,
|
| 758 |
+
"reasoning_reported_output_tokens": 153
|
| 759 |
+
},
|
| 760 |
+
"known_cached_input_tokens": 6921216,
|
| 761 |
+
"known_input_tokens": 7005688,
|
| 762 |
+
"known_output_tokens": 18192,
|
| 763 |
+
"known_cache_write_input_tokens": 0,
|
| 764 |
+
"known_reasoning_output_tokens": 6177,
|
| 765 |
+
"cache_hit_rate": 0.987942369114925,
|
| 766 |
+
"uncached_input_tokens": 84472,
|
| 767 |
+
"cost_usd": null,
|
| 768 |
+
"response_ids_complete": true,
|
| 769 |
+
"usage_source": "provider-raw",
|
| 770 |
+
"request_attempts": 155,
|
| 771 |
+
"rejected_attempts": 0,
|
| 772 |
+
"unresolved_attempts": [],
|
| 773 |
+
"audit_complete": true,
|
| 774 |
+
"accounting": "reported-responses"
|
| 775 |
+
},
|
| 776 |
+
"codex_wall_time_s": 394.373333,
|
| 777 |
+
"kinex_wall_time_s": 1396.812595,
|
| 778 |
+
"kinex_version": "0.10.3",
|
| 779 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 780 |
+
"same_roboenv_source": true
|
| 781 |
},
|
| 782 |
{
|
| 783 |
+
"task_key": "task01/08",
|
| 784 |
+
"instruction": "Pick up the sponge from the counter and scrub across a broad area of the cutting board, keeping the sponge grasped and in contact with the board throughout the scrubbing motion. Once finished, release the sponge and retract the gripper well away from it.",
|
| 785 |
+
"instruction_policy": "modified",
|
| 786 |
+
"codex_success": false,
|
| 787 |
"kinex_success": true,
|
| 788 |
+
"codex_steps": 1212,
|
| 789 |
+
"kinex_steps": 1007,
|
| 790 |
+
"codex_usage": {
|
| 791 |
+
"accounting": "reported-responses",
|
| 792 |
+
"audit_complete": true,
|
| 793 |
+
"cache_hit_rate": 0.9606465096549688,
|
| 794 |
+
"cache_reported_input_tokens": 775611,
|
| 795 |
+
"cache_write_input_tokens": 0,
|
| 796 |
+
"cache_write_reported_input_tokens": 775611,
|
| 797 |
+
"cached_input_tokens": 745088,
|
| 798 |
+
"completed_turns": 1,
|
| 799 |
+
"cost_usd": null,
|
| 800 |
+
"failed_turns": 0,
|
| 801 |
+
"input_tokens": 775611,
|
| 802 |
+
"known_cache_write_input_tokens": 0,
|
| 803 |
+
"known_cached_input_tokens": 745088,
|
| 804 |
+
"known_input_tokens": 775611,
|
| 805 |
+
"known_output_tokens": 6607,
|
| 806 |
+
"known_reasoning_output_tokens": 1428,
|
| 807 |
+
"output_tokens": 6607,
|
| 808 |
+
"reasoning_output_tokens": 1428,
|
| 809 |
+
"reasoning_reported_output_tokens": 6607,
|
| 810 |
+
"reported_responses": {
|
| 811 |
+
"cache_reported_input_tokens": 26,
|
| 812 |
+
"cache_write_input_tokens": 26,
|
| 813 |
+
"cache_write_reported_input_tokens": 26,
|
| 814 |
+
"cached_input_tokens": 26,
|
| 815 |
+
"input_tokens": 26,
|
| 816 |
+
"output_tokens": 26,
|
| 817 |
+
"reasoning_output_tokens": 26,
|
| 818 |
+
"reasoning_reported_output_tokens": 26
|
| 819 |
+
},
|
| 820 |
+
"response_count": 26,
|
| 821 |
+
"response_ids_complete": true,
|
| 822 |
+
"schema": "rlebench/token-usage/1",
|
| 823 |
+
"source": "Codex token_usage_record per response",
|
| 824 |
+
"uncached_input_tokens": 30523,
|
| 825 |
+
"unidentified_usage_records": 0
|
| 826 |
+
},
|
| 827 |
+
"kinex_usage": {
|
| 828 |
+
"input_tokens": 1225659,
|
| 829 |
+
"output_tokens": 6356,
|
| 830 |
+
"cached_input_tokens": 1182720,
|
| 831 |
+
"cache_write_input_tokens": 0,
|
| 832 |
+
"reasoning_output_tokens": 1012,
|
| 833 |
+
"cache_reported_input_tokens": 1225659,
|
| 834 |
+
"cache_write_reported_input_tokens": 1225659,
|
| 835 |
+
"reasoning_reported_output_tokens": 6356,
|
| 836 |
+
"schema": "rlebench/token-usage/1",
|
| 837 |
+
"source": "Provider terminal response usage",
|
| 838 |
+
"response_count": 42,
|
| 839 |
+
"reported_responses": {
|
| 840 |
+
"input_tokens": 42,
|
| 841 |
+
"output_tokens": 42,
|
| 842 |
+
"cached_input_tokens": 42,
|
| 843 |
+
"cache_write_input_tokens": 42,
|
| 844 |
+
"reasoning_output_tokens": 42,
|
| 845 |
+
"cache_reported_input_tokens": 42,
|
| 846 |
+
"cache_write_reported_input_tokens": 42,
|
| 847 |
+
"reasoning_reported_output_tokens": 42
|
| 848 |
+
},
|
| 849 |
+
"known_cached_input_tokens": 1182720,
|
| 850 |
+
"known_input_tokens": 1225659,
|
| 851 |
+
"known_output_tokens": 6356,
|
| 852 |
+
"known_cache_write_input_tokens": 0,
|
| 853 |
+
"known_reasoning_output_tokens": 1012,
|
| 854 |
+
"cache_hit_rate": 0.9649666016404237,
|
| 855 |
+
"uncached_input_tokens": 42939,
|
| 856 |
+
"cost_usd": null,
|
| 857 |
+
"response_ids_complete": true,
|
| 858 |
+
"usage_source": "provider-raw",
|
| 859 |
+
"request_attempts": 42,
|
| 860 |
+
"rejected_attempts": 0,
|
| 861 |
+
"unresolved_attempts": [],
|
| 862 |
+
"audit_complete": true,
|
| 863 |
+
"accounting": "reported-responses"
|
| 864 |
+
},
|
| 865 |
+
"codex_wall_time_s": 276.421696,
|
| 866 |
+
"kinex_wall_time_s": 418.12241,
|
| 867 |
+
"kinex_version": "0.10.3",
|
| 868 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 869 |
+
"same_roboenv_source": true
|
| 870 |
},
|
| 871 |
{
|
| 872 |
+
"task_key": "task01/09",
|
| 873 |
+
"instruction": "Pick the specified vegetable and the cream cheese from the fridge, place both fully inside the blender, and turn it on.",
|
| 874 |
+
"instruction_policy": "modified",
|
| 875 |
"codex_success": false,
|
| 876 |
"kinex_success": false,
|
| 877 |
+
"codex_steps": 2216,
|
| 878 |
+
"kinex_steps": 4628,
|
| 879 |
+
"codex_usage": {
|
| 880 |
+
"accounting": "reported-responses",
|
| 881 |
+
"audit_complete": true,
|
| 882 |
+
"cache_hit_rate": 0.9733556695336862,
|
| 883 |
+
"cache_reported_input_tokens": 1973478,
|
| 884 |
+
"cache_write_input_tokens": 0,
|
| 885 |
+
"cache_write_reported_input_tokens": 1973478,
|
| 886 |
+
"cached_input_tokens": 1920896,
|
| 887 |
+
"completed_turns": 1,
|
| 888 |
+
"cost_usd": null,
|
| 889 |
+
"failed_turns": 0,
|
| 890 |
+
"input_tokens": 1973478,
|
| 891 |
+
"known_cache_write_input_tokens": 0,
|
| 892 |
+
"known_cached_input_tokens": 1920896,
|
| 893 |
+
"known_input_tokens": 1973478,
|
| 894 |
+
"known_output_tokens": 11695,
|
| 895 |
+
"known_reasoning_output_tokens": 4578,
|
| 896 |
+
"output_tokens": 11695,
|
| 897 |
+
"reasoning_output_tokens": 4578,
|
| 898 |
+
"reasoning_reported_output_tokens": 11695,
|
| 899 |
+
"reported_responses": {
|
| 900 |
+
"cache_reported_input_tokens": 48,
|
| 901 |
+
"cache_write_input_tokens": 48,
|
| 902 |
+
"cache_write_reported_input_tokens": 48,
|
| 903 |
+
"cached_input_tokens": 48,
|
| 904 |
+
"input_tokens": 48,
|
| 905 |
+
"output_tokens": 48,
|
| 906 |
+
"reasoning_output_tokens": 48,
|
| 907 |
+
"reasoning_reported_output_tokens": 48
|
| 908 |
+
},
|
| 909 |
+
"response_count": 48,
|
| 910 |
+
"response_ids_complete": true,
|
| 911 |
+
"schema": "rlebench/token-usage/1",
|
| 912 |
+
"source": "Codex token_usage_record per response",
|
| 913 |
+
"uncached_input_tokens": 52582,
|
| 914 |
+
"unidentified_usage_records": 0
|
| 915 |
+
},
|
| 916 |
+
"kinex_usage": {
|
| 917 |
+
"input_tokens": 19377448,
|
| 918 |
+
"output_tokens": 36860,
|
| 919 |
+
"cached_input_tokens": 19235968,
|
| 920 |
+
"cache_write_input_tokens": 0,
|
| 921 |
+
"reasoning_output_tokens": 16258,
|
| 922 |
+
"cache_reported_input_tokens": 19377448,
|
| 923 |
+
"cache_write_reported_input_tokens": 19377448,
|
| 924 |
+
"reasoning_reported_output_tokens": 36860,
|
| 925 |
+
"schema": "rlebench/token-usage/1",
|
| 926 |
+
"source": "Provider terminal response usage",
|
| 927 |
+
"response_count": 308,
|
| 928 |
+
"reported_responses": {
|
| 929 |
+
"input_tokens": 308,
|
| 930 |
+
"output_tokens": 308,
|
| 931 |
+
"cached_input_tokens": 308,
|
| 932 |
+
"cache_write_input_tokens": 308,
|
| 933 |
+
"reasoning_output_tokens": 308,
|
| 934 |
+
"cache_reported_input_tokens": 308,
|
| 935 |
+
"cache_write_reported_input_tokens": 308,
|
| 936 |
+
"reasoning_reported_output_tokens": 308
|
| 937 |
+
},
|
| 938 |
+
"known_cached_input_tokens": 19235968,
|
| 939 |
+
"known_input_tokens": 19377448,
|
| 940 |
+
"known_output_tokens": 36860,
|
| 941 |
+
"known_cache_write_input_tokens": 0,
|
| 942 |
+
"known_reasoning_output_tokens": 16258,
|
| 943 |
+
"cache_hit_rate": 0.9926987289554331,
|
| 944 |
+
"uncached_input_tokens": 141480,
|
| 945 |
+
"cost_usd": null,
|
| 946 |
+
"response_ids_complete": true,
|
| 947 |
+
"usage_source": "provider-raw",
|
| 948 |
+
"request_attempts": 308,
|
| 949 |
+
"rejected_attempts": 0,
|
| 950 |
+
"unresolved_attempts": [],
|
| 951 |
+
"audit_complete": true,
|
| 952 |
+
"accounting": "reported-responses"
|
| 953 |
+
},
|
| 954 |
+
"codex_wall_time_s": 625.032726,
|
| 955 |
+
"kinex_wall_time_s": 3327.021406,
|
| 956 |
+
"kinex_version": "0.10.3",
|
| 957 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 958 |
+
"same_roboenv_source": true
|
| 959 |
},
|
| 960 |
{
|
| 961 |
+
"task_key": "task01/10",
|
| 962 |
+
"instruction": "Pick the specified vegetable from the fridge and hold it under running water from the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting. Release it and move the gripper clear.",
|
| 963 |
+
"instruction_policy": "modified",
|
| 964 |
"codex_success": true,
|
| 965 |
+
"kinex_success": false,
|
| 966 |
+
"codex_steps": 3969,
|
| 967 |
+
"kinex_steps": 4505,
|
| 968 |
+
"codex_usage": {
|
| 969 |
+
"accounting": "reported-responses",
|
| 970 |
+
"audit_complete": true,
|
| 971 |
+
"cache_hit_rate": 0.9777504182000669,
|
| 972 |
+
"cache_reported_input_tokens": 2989000,
|
| 973 |
+
"cache_write_input_tokens": 0,
|
| 974 |
+
"cache_write_reported_input_tokens": 2989000,
|
| 975 |
+
"cached_input_tokens": 2922496,
|
| 976 |
+
"completed_turns": 1,
|
| 977 |
+
"cost_usd": null,
|
| 978 |
+
"failed_turns": 0,
|
| 979 |
+
"input_tokens": 2989000,
|
| 980 |
+
"known_cache_write_input_tokens": 0,
|
| 981 |
+
"known_cached_input_tokens": 2922496,
|
| 982 |
+
"known_input_tokens": 2989000,
|
| 983 |
+
"known_output_tokens": 13452,
|
| 984 |
+
"known_reasoning_output_tokens": 4273,
|
| 985 |
+
"output_tokens": 13452,
|
| 986 |
+
"reasoning_output_tokens": 4273,
|
| 987 |
+
"reasoning_reported_output_tokens": 13452,
|
| 988 |
+
"reported_responses": {
|
| 989 |
+
"cache_reported_input_tokens": 68,
|
| 990 |
+
"cache_write_input_tokens": 68,
|
| 991 |
+
"cache_write_reported_input_tokens": 68,
|
| 992 |
+
"cached_input_tokens": 68,
|
| 993 |
+
"input_tokens": 68,
|
| 994 |
+
"output_tokens": 68,
|
| 995 |
+
"reasoning_output_tokens": 68,
|
| 996 |
+
"reasoning_reported_output_tokens": 68
|
| 997 |
+
},
|
| 998 |
+
"response_count": 68,
|
| 999 |
+
"response_ids_complete": true,
|
| 1000 |
+
"schema": "rlebench/token-usage/1",
|
| 1001 |
+
"source": "Codex token_usage_record per response",
|
| 1002 |
+
"uncached_input_tokens": 66504,
|
| 1003 |
+
"unidentified_usage_records": 0
|
| 1004 |
+
},
|
| 1005 |
+
"kinex_usage": {
|
| 1006 |
+
"input_tokens": 8091847,
|
| 1007 |
+
"output_tokens": 21393,
|
| 1008 |
+
"cached_input_tokens": 7916288,
|
| 1009 |
+
"cache_write_input_tokens": 0,
|
| 1010 |
+
"reasoning_output_tokens": 8065,
|
| 1011 |
+
"cache_reported_input_tokens": 8091847,
|
| 1012 |
+
"cache_write_reported_input_tokens": 8091847,
|
| 1013 |
+
"reasoning_reported_output_tokens": 21393,
|
| 1014 |
+
"schema": "rlebench/token-usage/1",
|
| 1015 |
+
"source": "Provider terminal response usage",
|
| 1016 |
+
"response_count": 160,
|
| 1017 |
+
"reported_responses": {
|
| 1018 |
+
"input_tokens": 160,
|
| 1019 |
+
"output_tokens": 160,
|
| 1020 |
+
"cached_input_tokens": 160,
|
| 1021 |
+
"cache_write_input_tokens": 160,
|
| 1022 |
+
"reasoning_output_tokens": 160,
|
| 1023 |
+
"cache_reported_input_tokens": 160,
|
| 1024 |
+
"cache_write_reported_input_tokens": 160,
|
| 1025 |
+
"reasoning_reported_output_tokens": 160
|
| 1026 |
+
},
|
| 1027 |
+
"known_cached_input_tokens": 7916288,
|
| 1028 |
+
"known_input_tokens": 8091847,
|
| 1029 |
+
"known_output_tokens": 21393,
|
| 1030 |
+
"known_cache_write_input_tokens": 0,
|
| 1031 |
+
"known_reasoning_output_tokens": 8065,
|
| 1032 |
+
"cache_hit_rate": 0.9783042116342536,
|
| 1033 |
+
"uncached_input_tokens": 175559,
|
| 1034 |
+
"cost_usd": null,
|
| 1035 |
+
"response_ids_complete": true,
|
| 1036 |
+
"usage_source": "provider-raw",
|
| 1037 |
+
"request_attempts": 160,
|
| 1038 |
+
"rejected_attempts": 0,
|
| 1039 |
+
"unresolved_attempts": [],
|
| 1040 |
+
"audit_complete": true,
|
| 1041 |
+
"accounting": "reported-responses"
|
| 1042 |
+
},
|
| 1043 |
+
"codex_wall_time_s": 800.707309,
|
| 1044 |
+
"kinex_wall_time_s": 1708.264043,
|
| 1045 |
+
"kinex_version": "0.10.3",
|
| 1046 |
+
"kinex_revision": "a1eaece0bd01a25177a43c2091c10b213b416aa5",
|
| 1047 |
+
"same_roboenv_source": true
|
| 1048 |
},
|
| 1049 |
{
|
| 1050 |
+
"task_key": "task02/03",
|
| 1051 |
+
"instruction": "turn on the stove and put the moka pot on it",
|
| 1052 |
+
"instruction_policy": "original_native",
|
| 1053 |
"codex_success": true,
|
| 1054 |
"kinex_success": true,
|
| 1055 |
+
"codex_steps": 2013,
|
| 1056 |
+
"kinex_steps": 706,
|
| 1057 |
+
"codex_usage": {
|
| 1058 |
+
"accounting": "reported-responses",
|
| 1059 |
+
"audit_complete": true,
|
| 1060 |
+
"cache_hit_rate": 0.967178914688548,
|
| 1061 |
+
"cache_reported_input_tokens": 2225094,
|
| 1062 |
+
"cache_write_input_tokens": 0,
|
| 1063 |
+
"cache_write_reported_input_tokens": 2225094,
|
| 1064 |
+
"cached_input_tokens": 2152064,
|
| 1065 |
+
"completed_turns": 1,
|
| 1066 |
+
"cost_usd": null,
|
| 1067 |
+
"failed_turns": 0,
|
| 1068 |
+
"input_tokens": 2225094,
|
| 1069 |
+
"known_cache_write_input_tokens": 0,
|
| 1070 |
+
"known_cached_input_tokens": 2152064,
|
| 1071 |
+
"known_input_tokens": 2225094,
|
| 1072 |
+
"known_output_tokens": 13673,
|
| 1073 |
+
"known_reasoning_output_tokens": 5025,
|
| 1074 |
+
"output_tokens": 13673,
|
| 1075 |
+
"reasoning_output_tokens": 5025,
|
| 1076 |
+
"reasoning_reported_output_tokens": 13673,
|
| 1077 |
+
"reported_responses": {
|
| 1078 |
+
"cache_reported_input_tokens": 50,
|
| 1079 |
+
"cache_write_input_tokens": 50,
|
| 1080 |
+
"cache_write_reported_input_tokens": 50,
|
| 1081 |
+
"cached_input_tokens": 50,
|
| 1082 |
+
"input_tokens": 50,
|
| 1083 |
+
"output_tokens": 50,
|
| 1084 |
+
"reasoning_output_tokens": 50,
|
| 1085 |
+
"reasoning_reported_output_tokens": 50
|
| 1086 |
+
},
|
| 1087 |
+
"response_count": 50,
|
| 1088 |
+
"response_ids_complete": true,
|
| 1089 |
+
"schema": "rlebench/token-usage/1",
|
| 1090 |
+
"source": "Codex token_usage_record per response",
|
| 1091 |
+
"uncached_input_tokens": 73030,
|
| 1092 |
+
"unidentified_usage_records": 0
|
| 1093 |
+
},
|
| 1094 |
+
"kinex_usage": {
|
| 1095 |
+
"input_tokens": 1221677,
|
| 1096 |
+
"output_tokens": 6604,
|
| 1097 |
+
"cached_input_tokens": 1146496,
|
| 1098 |
+
"cache_write_input_tokens": 0,
|
| 1099 |
+
"reasoning_output_tokens": 1955,
|
| 1100 |
+
"cache_reported_input_tokens": 1221677,
|
| 1101 |
+
"cache_write_reported_input_tokens": 1221677,
|
| 1102 |
+
"reasoning_reported_output_tokens": 6604,
|
| 1103 |
+
"schema": "rlebench/token-usage/1",
|
| 1104 |
+
"source": "Provider terminal response usage",
|
| 1105 |
+
"response_count": 47,
|
| 1106 |
+
"reported_responses": {
|
| 1107 |
+
"input_tokens": 47,
|
| 1108 |
+
"output_tokens": 47,
|
| 1109 |
+
"cached_input_tokens": 47,
|
| 1110 |
+
"cache_write_input_tokens": 47,
|
| 1111 |
+
"reasoning_output_tokens": 47,
|
| 1112 |
+
"cache_reported_input_tokens": 47,
|
| 1113 |
+
"cache_write_reported_input_tokens": 47,
|
| 1114 |
+
"reasoning_reported_output_tokens": 47
|
| 1115 |
+
},
|
| 1116 |
+
"known_cached_input_tokens": 1146496,
|
| 1117 |
+
"known_input_tokens": 1221677,
|
| 1118 |
+
"known_output_tokens": 6604,
|
| 1119 |
+
"known_cache_write_input_tokens": 0,
|
| 1120 |
+
"known_reasoning_output_tokens": 1955,
|
| 1121 |
+
"cache_hit_rate": 0.9384608206588158,
|
| 1122 |
+
"uncached_input_tokens": 75181,
|
| 1123 |
+
"cost_usd": null,
|
| 1124 |
+
"response_ids_complete": true,
|
| 1125 |
+
"accounting": "reported-responses",
|
| 1126 |
+
"usage_source": "provider-raw",
|
| 1127 |
+
"request_attempts": 47,
|
| 1128 |
+
"rejected_attempts": 0,
|
| 1129 |
+
"unresolved_attempts": [],
|
| 1130 |
+
"audit_complete": true
|
| 1131 |
+
},
|
| 1132 |
+
"codex_wall_time_s": 531.974225,
|
| 1133 |
+
"kinex_wall_time_s": 371.824558,
|
| 1134 |
+
"kinex_version": "0.10.3",
|
| 1135 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1136 |
+
"same_roboenv_source": true
|
| 1137 |
},
|
| 1138 |
{
|
| 1139 |
+
"task_key": "task02/01",
|
| 1140 |
+
"instruction": "put both the alphabet soup and the tomato sauce in the basket",
|
| 1141 |
+
"instruction_policy": "original_native",
|
| 1142 |
"codex_success": true,
|
| 1143 |
"kinex_success": true,
|
| 1144 |
+
"codex_steps": 969,
|
| 1145 |
+
"kinex_steps": 1094,
|
| 1146 |
+
"codex_usage": {
|
| 1147 |
+
"accounting": "reported-responses",
|
| 1148 |
+
"audit_complete": true,
|
| 1149 |
+
"cache_hit_rate": 0.9509386255123954,
|
| 1150 |
+
"cache_reported_input_tokens": 832121,
|
| 1151 |
+
"cache_write_input_tokens": 0,
|
| 1152 |
+
"cache_write_reported_input_tokens": 832121,
|
| 1153 |
+
"cached_input_tokens": 791296,
|
| 1154 |
+
"completed_turns": 1,
|
| 1155 |
+
"cost_usd": null,
|
| 1156 |
+
"failed_turns": 0,
|
| 1157 |
+
"input_tokens": 832121,
|
| 1158 |
+
"known_cache_write_input_tokens": 0,
|
| 1159 |
+
"known_cached_input_tokens": 791296,
|
| 1160 |
+
"known_input_tokens": 832121,
|
| 1161 |
+
"known_output_tokens": 6802,
|
| 1162 |
+
"known_reasoning_output_tokens": 1635,
|
| 1163 |
+
"output_tokens": 6802,
|
| 1164 |
+
"reasoning_output_tokens": 1635,
|
| 1165 |
+
"reasoning_reported_output_tokens": 6802,
|
| 1166 |
+
"reported_responses": {
|
| 1167 |
+
"cache_reported_input_tokens": 27,
|
| 1168 |
+
"cache_write_input_tokens": 27,
|
| 1169 |
+
"cache_write_reported_input_tokens": 27,
|
| 1170 |
+
"cached_input_tokens": 27,
|
| 1171 |
+
"input_tokens": 27,
|
| 1172 |
+
"output_tokens": 27,
|
| 1173 |
+
"reasoning_output_tokens": 27,
|
| 1174 |
+
"reasoning_reported_output_tokens": 27
|
| 1175 |
+
},
|
| 1176 |
+
"response_count": 27,
|
| 1177 |
+
"response_ids_complete": true,
|
| 1178 |
+
"schema": "rlebench/token-usage/1",
|
| 1179 |
+
"source": "Codex token_usage_record per response",
|
| 1180 |
+
"uncached_input_tokens": 40825,
|
| 1181 |
+
"unidentified_usage_records": 0
|
| 1182 |
+
},
|
| 1183 |
+
"kinex_usage": {
|
| 1184 |
+
"input_tokens": 1546168,
|
| 1185 |
+
"output_tokens": 7604,
|
| 1186 |
+
"cached_input_tokens": 1461888,
|
| 1187 |
+
"cache_write_input_tokens": 0,
|
| 1188 |
+
"reasoning_output_tokens": 2069,
|
| 1189 |
+
"cache_reported_input_tokens": 1546168,
|
| 1190 |
+
"cache_write_reported_input_tokens": 1546168,
|
| 1191 |
+
"reasoning_reported_output_tokens": 7604,
|
| 1192 |
+
"schema": "rlebench/token-usage/1",
|
| 1193 |
+
"source": "Provider terminal response usage",
|
| 1194 |
+
"response_count": 55,
|
| 1195 |
+
"reported_responses": {
|
| 1196 |
+
"input_tokens": 55,
|
| 1197 |
+
"output_tokens": 55,
|
| 1198 |
+
"cached_input_tokens": 55,
|
| 1199 |
+
"cache_write_input_tokens": 55,
|
| 1200 |
+
"reasoning_output_tokens": 55,
|
| 1201 |
+
"cache_reported_input_tokens": 55,
|
| 1202 |
+
"cache_write_reported_input_tokens": 55,
|
| 1203 |
+
"reasoning_reported_output_tokens": 55
|
| 1204 |
+
},
|
| 1205 |
+
"known_cached_input_tokens": 1461888,
|
| 1206 |
+
"known_input_tokens": 1546168,
|
| 1207 |
+
"known_output_tokens": 7604,
|
| 1208 |
+
"known_cache_write_input_tokens": 0,
|
| 1209 |
+
"known_reasoning_output_tokens": 2069,
|
| 1210 |
+
"cache_hit_rate": 0.9454910462511189,
|
| 1211 |
+
"uncached_input_tokens": 84280,
|
| 1212 |
+
"cost_usd": null,
|
| 1213 |
+
"response_ids_complete": true,
|
| 1214 |
+
"accounting": "reported-responses",
|
| 1215 |
+
"usage_source": "provider-raw",
|
| 1216 |
+
"request_attempts": 55,
|
| 1217 |
+
"rejected_attempts": 0,
|
| 1218 |
+
"unresolved_attempts": [],
|
| 1219 |
+
"audit_complete": true
|
| 1220 |
+
},
|
| 1221 |
+
"codex_wall_time_s": 272.779105,
|
| 1222 |
+
"kinex_wall_time_s": 406.249632,
|
| 1223 |
+
"kinex_version": "0.10.3",
|
| 1224 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1225 |
+
"same_roboenv_source": true
|
| 1226 |
},
|
| 1227 |
{
|
| 1228 |
+
"task_key": "task02/02",
|
| 1229 |
+
"instruction": "put both the cream cheese box and the butter in the basket",
|
| 1230 |
+
"instruction_policy": "original_native",
|
| 1231 |
"codex_success": true,
|
| 1232 |
"kinex_success": true,
|
| 1233 |
+
"codex_steps": 722,
|
| 1234 |
+
"kinex_steps": 1011,
|
| 1235 |
+
"codex_usage": {
|
| 1236 |
+
"accounting": "reported-responses",
|
| 1237 |
+
"audit_complete": true,
|
| 1238 |
+
"cache_hit_rate": 0.9515345728205089,
|
| 1239 |
+
"cache_reported_input_tokens": 784518,
|
| 1240 |
+
"cache_write_input_tokens": 0,
|
| 1241 |
+
"cache_write_reported_input_tokens": 784518,
|
| 1242 |
+
"cached_input_tokens": 746496,
|
| 1243 |
+
"completed_turns": 1,
|
| 1244 |
+
"cost_usd": null,
|
| 1245 |
+
"failed_turns": 0,
|
| 1246 |
+
"input_tokens": 784518,
|
| 1247 |
+
"known_cache_write_input_tokens": 0,
|
| 1248 |
+
"known_cached_input_tokens": 746496,
|
| 1249 |
+
"known_input_tokens": 784518,
|
| 1250 |
+
"known_output_tokens": 6225,
|
| 1251 |
+
"known_reasoning_output_tokens": 1371,
|
| 1252 |
+
"output_tokens": 6225,
|
| 1253 |
+
"reasoning_output_tokens": 1371,
|
| 1254 |
+
"reasoning_reported_output_tokens": 6225,
|
| 1255 |
+
"reported_responses": {
|
| 1256 |
+
"cache_reported_input_tokens": 28,
|
| 1257 |
+
"cache_write_input_tokens": 28,
|
| 1258 |
+
"cache_write_reported_input_tokens": 28,
|
| 1259 |
+
"cached_input_tokens": 28,
|
| 1260 |
+
"input_tokens": 28,
|
| 1261 |
+
"output_tokens": 28,
|
| 1262 |
+
"reasoning_output_tokens": 28,
|
| 1263 |
+
"reasoning_reported_output_tokens": 28
|
| 1264 |
+
},
|
| 1265 |
+
"response_count": 28,
|
| 1266 |
+
"response_ids_complete": true,
|
| 1267 |
+
"schema": "rlebench/token-usage/1",
|
| 1268 |
+
"source": "Codex token_usage_record per response",
|
| 1269 |
+
"uncached_input_tokens": 38022,
|
| 1270 |
+
"unidentified_usage_records": 0
|
| 1271 |
+
},
|
| 1272 |
+
"kinex_usage": {
|
| 1273 |
+
"input_tokens": 2453051,
|
| 1274 |
+
"output_tokens": 9069,
|
| 1275 |
+
"cached_input_tokens": 2363776,
|
| 1276 |
+
"cache_write_input_tokens": 0,
|
| 1277 |
+
"reasoning_output_tokens": 2858,
|
| 1278 |
+
"cache_reported_input_tokens": 2453051,
|
| 1279 |
+
"cache_write_reported_input_tokens": 2453051,
|
| 1280 |
+
"reasoning_reported_output_tokens": 9069,
|
| 1281 |
+
"schema": "rlebench/token-usage/1",
|
| 1282 |
+
"source": "Provider terminal response usage",
|
| 1283 |
+
"response_count": 84,
|
| 1284 |
+
"reported_responses": {
|
| 1285 |
+
"input_tokens": 84,
|
| 1286 |
+
"output_tokens": 84,
|
| 1287 |
+
"cached_input_tokens": 84,
|
| 1288 |
+
"cache_write_input_tokens": 84,
|
| 1289 |
+
"reasoning_output_tokens": 84,
|
| 1290 |
+
"cache_reported_input_tokens": 84,
|
| 1291 |
+
"cache_write_reported_input_tokens": 84,
|
| 1292 |
+
"reasoning_reported_output_tokens": 84
|
| 1293 |
+
},
|
| 1294 |
+
"known_cached_input_tokens": 2363776,
|
| 1295 |
+
"known_input_tokens": 2453051,
|
| 1296 |
+
"known_output_tokens": 9069,
|
| 1297 |
+
"known_cache_write_input_tokens": 0,
|
| 1298 |
+
"known_reasoning_output_tokens": 2858,
|
| 1299 |
+
"cache_hit_rate": 0.9636065454815248,
|
| 1300 |
+
"uncached_input_tokens": 89275,
|
| 1301 |
+
"cost_usd": null,
|
| 1302 |
+
"response_ids_complete": true,
|
| 1303 |
+
"accounting": "reported-responses",
|
| 1304 |
+
"usage_source": "provider-raw",
|
| 1305 |
+
"request_attempts": 84,
|
| 1306 |
+
"rejected_attempts": 0,
|
| 1307 |
+
"unresolved_attempts": [],
|
| 1308 |
+
"audit_complete": true
|
| 1309 |
+
},
|
| 1310 |
+
"codex_wall_time_s": 231.024585,
|
| 1311 |
+
"kinex_wall_time_s": 611.087641,
|
| 1312 |
+
"kinex_version": "0.10.3",
|
| 1313 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1314 |
+
"same_roboenv_source": true
|
| 1315 |
},
|
| 1316 |
{
|
| 1317 |
+
"task_key": "task02/04",
|
| 1318 |
+
"instruction": "put the black bowl in the bottom drawer of the cabinet and close it",
|
| 1319 |
+
"instruction_policy": "original_native",
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1320 |
"codex_success": false,
|
| 1321 |
"kinex_success": true,
|
| 1322 |
+
"codex_steps": 3687,
|
| 1323 |
+
"kinex_steps": 1483,
|
| 1324 |
+
"codex_usage": {
|
| 1325 |
+
"accounting": "reported-responses",
|
| 1326 |
+
"audit_complete": true,
|
| 1327 |
+
"cache_hit_rate": 0.9745746678635921,
|
| 1328 |
+
"cache_reported_input_tokens": 3356377,
|
| 1329 |
+
"cache_write_input_tokens": 0,
|
| 1330 |
+
"cache_write_reported_input_tokens": 3356377,
|
| 1331 |
+
"cached_input_tokens": 3271040,
|
| 1332 |
+
"completed_turns": 1,
|
| 1333 |
+
"cost_usd": null,
|
| 1334 |
+
"failed_turns": 0,
|
| 1335 |
+
"input_tokens": 3356377,
|
| 1336 |
+
"known_cache_write_input_tokens": 0,
|
| 1337 |
+
"known_cached_input_tokens": 3271040,
|
| 1338 |
+
"known_input_tokens": 3356377,
|
| 1339 |
+
"known_output_tokens": 23446,
|
| 1340 |
+
"known_reasoning_output_tokens": 11600,
|
| 1341 |
+
"output_tokens": 23446,
|
| 1342 |
+
"reasoning_output_tokens": 11600,
|
| 1343 |
+
"reasoning_reported_output_tokens": 23446,
|
| 1344 |
+
"reported_responses": {
|
| 1345 |
+
"cache_reported_input_tokens": 67,
|
| 1346 |
+
"cache_write_input_tokens": 67,
|
| 1347 |
+
"cache_write_reported_input_tokens": 67,
|
| 1348 |
+
"cached_input_tokens": 67,
|
| 1349 |
+
"input_tokens": 67,
|
| 1350 |
+
"output_tokens": 67,
|
| 1351 |
+
"reasoning_output_tokens": 67,
|
| 1352 |
+
"reasoning_reported_output_tokens": 67
|
| 1353 |
+
},
|
| 1354 |
+
"response_count": 67,
|
| 1355 |
+
"response_ids_complete": true,
|
| 1356 |
+
"schema": "rlebench/token-usage/1",
|
| 1357 |
+
"source": "Codex token_usage_record per response",
|
| 1358 |
+
"uncached_input_tokens": 85337,
|
| 1359 |
+
"unidentified_usage_records": 0
|
| 1360 |
+
},
|
| 1361 |
+
"kinex_usage": {
|
| 1362 |
+
"input_tokens": 1990960,
|
| 1363 |
+
"output_tokens": 12374,
|
| 1364 |
+
"cached_input_tokens": 1883904,
|
| 1365 |
+
"cache_write_input_tokens": 0,
|
| 1366 |
+
"reasoning_output_tokens": 5081,
|
| 1367 |
+
"cache_reported_input_tokens": 1990960,
|
| 1368 |
+
"cache_write_reported_input_tokens": 1990960,
|
| 1369 |
+
"reasoning_reported_output_tokens": 12374,
|
| 1370 |
+
"schema": "rlebench/token-usage/1",
|
| 1371 |
+
"source": "Provider terminal response usage",
|
| 1372 |
+
"response_count": 59,
|
| 1373 |
+
"reported_responses": {
|
| 1374 |
+
"input_tokens": 59,
|
| 1375 |
+
"output_tokens": 59,
|
| 1376 |
+
"cached_input_tokens": 59,
|
| 1377 |
+
"cache_write_input_tokens": 59,
|
| 1378 |
+
"reasoning_output_tokens": 59,
|
| 1379 |
+
"cache_reported_input_tokens": 59,
|
| 1380 |
+
"cache_write_reported_input_tokens": 59,
|
| 1381 |
+
"reasoning_reported_output_tokens": 59
|
| 1382 |
+
},
|
| 1383 |
+
"known_cached_input_tokens": 1883904,
|
| 1384 |
+
"known_input_tokens": 1990960,
|
| 1385 |
+
"known_output_tokens": 12374,
|
| 1386 |
+
"known_cache_write_input_tokens": 0,
|
| 1387 |
+
"known_reasoning_output_tokens": 5081,
|
| 1388 |
+
"cache_hit_rate": 0.9462289548760398,
|
| 1389 |
+
"uncached_input_tokens": 107056,
|
| 1390 |
+
"cost_usd": null,
|
| 1391 |
+
"response_ids_complete": true,
|
| 1392 |
+
"accounting": "reported-responses",
|
| 1393 |
+
"usage_source": "provider-raw",
|
| 1394 |
+
"request_attempts": 59,
|
| 1395 |
+
"rejected_attempts": 0,
|
| 1396 |
+
"unresolved_attempts": [],
|
| 1397 |
+
"audit_complete": true
|
| 1398 |
+
},
|
| 1399 |
+
"codex_wall_time_s": 919.204412,
|
| 1400 |
+
"kinex_wall_time_s": 624.888736,
|
| 1401 |
+
"kinex_version": "0.10.3",
|
| 1402 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1403 |
+
"same_roboenv_source": true
|
| 1404 |
},
|
| 1405 |
{
|
| 1406 |
+
"task_key": "task02/05",
|
| 1407 |
+
"instruction": "put the white mug in the center of the left plate and put the yellow and white mug in the center of the right plate, with each mug resting on its plate",
|
| 1408 |
+
"instruction_policy": "modified",
|
| 1409 |
"codex_success": true,
|
| 1410 |
"kinex_success": true,
|
| 1411 |
+
"codex_steps": 634,
|
| 1412 |
+
"kinex_steps": 765,
|
| 1413 |
+
"codex_usage": {
|
| 1414 |
+
"accounting": "reported-responses",
|
| 1415 |
+
"audit_complete": true,
|
| 1416 |
+
"cache_hit_rate": 0.9493350495033964,
|
| 1417 |
+
"cache_reported_input_tokens": 509662,
|
| 1418 |
+
"cache_write_input_tokens": 0,
|
| 1419 |
+
"cache_write_reported_input_tokens": 509662,
|
| 1420 |
+
"cached_input_tokens": 483840,
|
| 1421 |
+
"completed_turns": 1,
|
| 1422 |
+
"cost_usd": null,
|
| 1423 |
+
"failed_turns": 0,
|
| 1424 |
+
"input_tokens": 509662,
|
| 1425 |
+
"known_cache_write_input_tokens": 0,
|
| 1426 |
+
"known_cached_input_tokens": 483840,
|
| 1427 |
+
"known_input_tokens": 509662,
|
| 1428 |
+
"known_output_tokens": 5212,
|
| 1429 |
+
"known_reasoning_output_tokens": 1183,
|
| 1430 |
+
"output_tokens": 5212,
|
| 1431 |
+
"reasoning_output_tokens": 1183,
|
| 1432 |
+
"reasoning_reported_output_tokens": 5212,
|
| 1433 |
+
"reported_responses": {
|
| 1434 |
+
"cache_reported_input_tokens": 20,
|
| 1435 |
+
"cache_write_input_tokens": 20,
|
| 1436 |
+
"cache_write_reported_input_tokens": 20,
|
| 1437 |
+
"cached_input_tokens": 20,
|
| 1438 |
+
"input_tokens": 20,
|
| 1439 |
+
"output_tokens": 20,
|
| 1440 |
+
"reasoning_output_tokens": 20,
|
| 1441 |
+
"reasoning_reported_output_tokens": 20
|
| 1442 |
+
},
|
| 1443 |
+
"response_count": 20,
|
| 1444 |
+
"response_ids_complete": true,
|
| 1445 |
+
"schema": "rlebench/token-usage/1",
|
| 1446 |
+
"source": "Codex token_usage_record per response",
|
| 1447 |
+
"uncached_input_tokens": 25822,
|
| 1448 |
+
"unidentified_usage_records": 0
|
| 1449 |
+
},
|
| 1450 |
+
"kinex_usage": {
|
| 1451 |
+
"input_tokens": 2532606,
|
| 1452 |
+
"output_tokens": 7281,
|
| 1453 |
+
"cached_input_tokens": 2444032,
|
| 1454 |
+
"cache_write_input_tokens": 0,
|
| 1455 |
+
"reasoning_output_tokens": 1949,
|
| 1456 |
+
"cache_reported_input_tokens": 2532606,
|
| 1457 |
+
"cache_write_reported_input_tokens": 2532606,
|
| 1458 |
+
"reasoning_reported_output_tokens": 7281,
|
| 1459 |
+
"schema": "rlebench/token-usage/1",
|
| 1460 |
+
"source": "Provider terminal response usage",
|
| 1461 |
+
"response_count": 88,
|
| 1462 |
+
"reported_responses": {
|
| 1463 |
+
"input_tokens": 88,
|
| 1464 |
+
"output_tokens": 88,
|
| 1465 |
+
"cached_input_tokens": 88,
|
| 1466 |
+
"cache_write_input_tokens": 88,
|
| 1467 |
+
"reasoning_output_tokens": 88,
|
| 1468 |
+
"cache_reported_input_tokens": 88,
|
| 1469 |
+
"cache_write_reported_input_tokens": 88,
|
| 1470 |
+
"reasoning_reported_output_tokens": 88
|
| 1471 |
+
},
|
| 1472 |
+
"known_cached_input_tokens": 2444032,
|
| 1473 |
+
"known_input_tokens": 2532606,
|
| 1474 |
+
"known_output_tokens": 7281,
|
| 1475 |
+
"known_cache_write_input_tokens": 0,
|
| 1476 |
+
"known_reasoning_output_tokens": 1949,
|
| 1477 |
+
"cache_hit_rate": 0.9650265378823236,
|
| 1478 |
+
"uncached_input_tokens": 88574,
|
| 1479 |
+
"cost_usd": null,
|
| 1480 |
+
"response_ids_complete": true,
|
| 1481 |
+
"accounting": "reported-responses",
|
| 1482 |
+
"usage_source": "provider-raw",
|
| 1483 |
+
"request_attempts": 88,
|
| 1484 |
+
"rejected_attempts": 0,
|
| 1485 |
+
"unresolved_attempts": [],
|
| 1486 |
+
"audit_complete": true
|
| 1487 |
+
},
|
| 1488 |
+
"codex_wall_time_s": 195.147246,
|
| 1489 |
+
"kinex_wall_time_s": 569.224064,
|
| 1490 |
+
"kinex_version": "0.10.3",
|
| 1491 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1492 |
+
"same_roboenv_source": true
|
| 1493 |
},
|
| 1494 |
{
|
| 1495 |
+
"task_key": "task02/06",
|
| 1496 |
+
"instruction": "pick up the book and place it in the back compartment of the caddy, between the two large side compartments",
|
| 1497 |
+
"instruction_policy": "modified",
|
| 1498 |
"codex_success": true,
|
| 1499 |
"kinex_success": true,
|
| 1500 |
+
"codex_steps": 325,
|
| 1501 |
+
"kinex_steps": 441,
|
| 1502 |
+
"codex_usage": {
|
| 1503 |
+
"accounting": "reported-responses",
|
| 1504 |
+
"audit_complete": true,
|
| 1505 |
+
"cache_hit_rate": 0.956185574299006,
|
| 1506 |
+
"cache_reported_input_tokens": 707210,
|
| 1507 |
+
"cache_write_input_tokens": 0,
|
| 1508 |
+
"cache_write_reported_input_tokens": 707210,
|
| 1509 |
+
"cached_input_tokens": 676224,
|
| 1510 |
+
"completed_turns": 1,
|
| 1511 |
+
"cost_usd": null,
|
| 1512 |
+
"failed_turns": 0,
|
| 1513 |
+
"input_tokens": 707210,
|
| 1514 |
+
"known_cache_write_input_tokens": 0,
|
| 1515 |
+
"known_cached_input_tokens": 676224,
|
| 1516 |
+
"known_input_tokens": 707210,
|
| 1517 |
+
"known_output_tokens": 6550,
|
| 1518 |
+
"known_reasoning_output_tokens": 2424,
|
| 1519 |
+
"output_tokens": 6550,
|
| 1520 |
+
"reasoning_output_tokens": 2424,
|
| 1521 |
+
"reasoning_reported_output_tokens": 6550,
|
| 1522 |
+
"reported_responses": {
|
| 1523 |
+
"cache_reported_input_tokens": 25,
|
| 1524 |
+
"cache_write_input_tokens": 25,
|
| 1525 |
+
"cache_write_reported_input_tokens": 25,
|
| 1526 |
+
"cached_input_tokens": 25,
|
| 1527 |
+
"input_tokens": 25,
|
| 1528 |
+
"output_tokens": 25,
|
| 1529 |
+
"reasoning_output_tokens": 25,
|
| 1530 |
+
"reasoning_reported_output_tokens": 25
|
| 1531 |
+
},
|
| 1532 |
+
"response_count": 25,
|
| 1533 |
+
"response_ids_complete": true,
|
| 1534 |
+
"schema": "rlebench/token-usage/1",
|
| 1535 |
+
"source": "Codex token_usage_record per response",
|
| 1536 |
+
"uncached_input_tokens": 30986,
|
| 1537 |
+
"unidentified_usage_records": 0
|
| 1538 |
+
},
|
| 1539 |
+
"kinex_usage": {
|
| 1540 |
+
"input_tokens": 1345973,
|
| 1541 |
+
"output_tokens": 6571,
|
| 1542 |
+
"cached_input_tokens": 1272960,
|
| 1543 |
+
"cache_write_input_tokens": 0,
|
| 1544 |
+
"reasoning_output_tokens": 2281,
|
| 1545 |
+
"cache_reported_input_tokens": 1345973,
|
| 1546 |
+
"cache_write_reported_input_tokens": 1345973,
|
| 1547 |
+
"reasoning_reported_output_tokens": 6571,
|
| 1548 |
+
"schema": "rlebench/token-usage/1",
|
| 1549 |
+
"source": "Provider terminal response usage",
|
| 1550 |
+
"response_count": 54,
|
| 1551 |
+
"reported_responses": {
|
| 1552 |
+
"input_tokens": 54,
|
| 1553 |
+
"output_tokens": 54,
|
| 1554 |
+
"cached_input_tokens": 54,
|
| 1555 |
+
"cache_write_input_tokens": 54,
|
| 1556 |
+
"reasoning_output_tokens": 54,
|
| 1557 |
+
"cache_reported_input_tokens": 54,
|
| 1558 |
+
"cache_write_reported_input_tokens": 54,
|
| 1559 |
+
"reasoning_reported_output_tokens": 54
|
| 1560 |
+
},
|
| 1561 |
+
"known_cached_input_tokens": 1272960,
|
| 1562 |
+
"known_input_tokens": 1345973,
|
| 1563 |
+
"known_output_tokens": 6571,
|
| 1564 |
+
"known_cache_write_input_tokens": 0,
|
| 1565 |
+
"known_reasoning_output_tokens": 2281,
|
| 1566 |
+
"cache_hit_rate": 0.9457544839309555,
|
| 1567 |
+
"uncached_input_tokens": 73013,
|
| 1568 |
+
"cost_usd": null,
|
| 1569 |
+
"response_ids_complete": true,
|
| 1570 |
+
"accounting": "reported-responses",
|
| 1571 |
+
"usage_source": "provider-raw",
|
| 1572 |
+
"request_attempts": 54,
|
| 1573 |
+
"rejected_attempts": 0,
|
| 1574 |
+
"unresolved_attempts": [],
|
| 1575 |
+
"audit_complete": true
|
| 1576 |
+
},
|
| 1577 |
+
"codex_wall_time_s": 241.745547,
|
| 1578 |
+
"kinex_wall_time_s": 358.977488,
|
| 1579 |
+
"kinex_version": "0.10.3",
|
| 1580 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1581 |
+
"same_roboenv_source": true
|
| 1582 |
},
|
| 1583 |
{
|
| 1584 |
+
"task_key": "task02/07",
|
| 1585 |
+
"instruction": "put the white mug in the center of the plate and put the chocolate pudding immediately to the right of the plate",
|
| 1586 |
+
"instruction_policy": "modified",
|
| 1587 |
"codex_success": true,
|
| 1588 |
"kinex_success": true,
|
| 1589 |
+
"codex_steps": 1016,
|
| 1590 |
+
"kinex_steps": 1153,
|
| 1591 |
+
"codex_usage": {
|
| 1592 |
+
"accounting": "reported-responses",
|
| 1593 |
+
"audit_complete": true,
|
| 1594 |
+
"cache_hit_rate": 0.9487277032832223,
|
| 1595 |
+
"cache_reported_input_tokens": 701841,
|
| 1596 |
+
"cache_write_input_tokens": 0,
|
| 1597 |
+
"cache_write_reported_input_tokens": 701841,
|
| 1598 |
+
"cached_input_tokens": 665856,
|
| 1599 |
+
"completed_turns": 1,
|
| 1600 |
+
"cost_usd": null,
|
| 1601 |
+
"failed_turns": 0,
|
| 1602 |
+
"input_tokens": 701841,
|
| 1603 |
+
"known_cache_write_input_tokens": 0,
|
| 1604 |
+
"known_cached_input_tokens": 665856,
|
| 1605 |
+
"known_input_tokens": 701841,
|
| 1606 |
+
"known_output_tokens": 6754,
|
| 1607 |
+
"known_reasoning_output_tokens": 1904,
|
| 1608 |
+
"output_tokens": 6754,
|
| 1609 |
+
"reasoning_output_tokens": 1904,
|
| 1610 |
+
"reasoning_reported_output_tokens": 6754,
|
| 1611 |
+
"reported_responses": {
|
| 1612 |
+
"cache_reported_input_tokens": 24,
|
| 1613 |
+
"cache_write_input_tokens": 24,
|
| 1614 |
+
"cache_write_reported_input_tokens": 24,
|
| 1615 |
+
"cached_input_tokens": 24,
|
| 1616 |
+
"input_tokens": 24,
|
| 1617 |
+
"output_tokens": 24,
|
| 1618 |
+
"reasoning_output_tokens": 24,
|
| 1619 |
+
"reasoning_reported_output_tokens": 24
|
| 1620 |
+
},
|
| 1621 |
+
"response_count": 24,
|
| 1622 |
+
"response_ids_complete": true,
|
| 1623 |
+
"schema": "rlebench/token-usage/1",
|
| 1624 |
+
"source": "Codex token_usage_record per response",
|
| 1625 |
+
"uncached_input_tokens": 35985,
|
| 1626 |
+
"unidentified_usage_records": 0
|
| 1627 |
+
},
|
| 1628 |
+
"kinex_usage": {
|
| 1629 |
+
"input_tokens": 1086944,
|
| 1630 |
+
"output_tokens": 7094,
|
| 1631 |
+
"cached_input_tokens": 980736,
|
| 1632 |
+
"cache_write_input_tokens": 0,
|
| 1633 |
+
"reasoning_output_tokens": 2071,
|
| 1634 |
+
"cache_reported_input_tokens": 1086944,
|
| 1635 |
+
"cache_write_reported_input_tokens": 1086944,
|
| 1636 |
+
"reasoning_reported_output_tokens": 7094,
|
| 1637 |
+
"schema": "rlebench/token-usage/1",
|
| 1638 |
+
"source": "Provider terminal response usage",
|
| 1639 |
+
"response_count": 40,
|
| 1640 |
+
"reported_responses": {
|
| 1641 |
+
"input_tokens": 40,
|
| 1642 |
+
"output_tokens": 40,
|
| 1643 |
+
"cached_input_tokens": 40,
|
| 1644 |
+
"cache_write_input_tokens": 40,
|
| 1645 |
+
"reasoning_output_tokens": 40,
|
| 1646 |
+
"cache_reported_input_tokens": 40,
|
| 1647 |
+
"cache_write_reported_input_tokens": 40,
|
| 1648 |
+
"reasoning_reported_output_tokens": 40
|
| 1649 |
+
},
|
| 1650 |
+
"known_cached_input_tokens": 980736,
|
| 1651 |
+
"known_input_tokens": 1086944,
|
| 1652 |
+
"known_output_tokens": 7094,
|
| 1653 |
+
"known_cache_write_input_tokens": 0,
|
| 1654 |
+
"known_reasoning_output_tokens": 2071,
|
| 1655 |
+
"cache_hit_rate": 0.9022875143521654,
|
| 1656 |
+
"uncached_input_tokens": 106208,
|
| 1657 |
+
"cost_usd": null,
|
| 1658 |
+
"response_ids_complete": true,
|
| 1659 |
+
"accounting": "reported-responses",
|
| 1660 |
+
"usage_source": "provider-raw",
|
| 1661 |
+
"request_attempts": 40,
|
| 1662 |
+
"rejected_attempts": 0,
|
| 1663 |
+
"unresolved_attempts": [],
|
| 1664 |
+
"audit_complete": true
|
| 1665 |
+
},
|
| 1666 |
+
"codex_wall_time_s": 260.996992,
|
| 1667 |
+
"kinex_wall_time_s": 349.052826,
|
| 1668 |
+
"kinex_version": "0.10.3",
|
| 1669 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1670 |
+
"same_roboenv_source": true
|
| 1671 |
},
|
| 1672 |
{
|
| 1673 |
+
"task_key": "task02/08",
|
| 1674 |
+
"instruction": "put both the alphabet soup and the cream cheese box fully inside the basket",
|
| 1675 |
+
"instruction_policy": "modified",
|
| 1676 |
"codex_success": true,
|
| 1677 |
"kinex_success": true,
|
| 1678 |
+
"codex_steps": 1115,
|
| 1679 |
+
"kinex_steps": 926,
|
| 1680 |
+
"codex_usage": {
|
| 1681 |
+
"accounting": "reported-responses",
|
| 1682 |
+
"audit_complete": true,
|
| 1683 |
+
"cache_hit_rate": 0.9604073765886478,
|
| 1684 |
+
"cache_reported_input_tokens": 1406070,
|
| 1685 |
+
"cache_write_input_tokens": 0,
|
| 1686 |
+
"cache_write_reported_input_tokens": 1406070,
|
| 1687 |
+
"cached_input_tokens": 1350400,
|
| 1688 |
+
"completed_turns": 1,
|
| 1689 |
+
"cost_usd": null,
|
| 1690 |
+
"failed_turns": 0,
|
| 1691 |
+
"input_tokens": 1406070,
|
| 1692 |
+
"known_cache_write_input_tokens": 0,
|
| 1693 |
+
"known_cached_input_tokens": 1350400,
|
| 1694 |
+
"known_input_tokens": 1406070,
|
| 1695 |
+
"known_output_tokens": 9396,
|
| 1696 |
+
"known_reasoning_output_tokens": 4471,
|
| 1697 |
+
"output_tokens": 9396,
|
| 1698 |
+
"reasoning_output_tokens": 4471,
|
| 1699 |
+
"reasoning_reported_output_tokens": 9396,
|
| 1700 |
+
"reported_responses": {
|
| 1701 |
+
"cache_reported_input_tokens": 40,
|
| 1702 |
+
"cache_write_input_tokens": 40,
|
| 1703 |
+
"cache_write_reported_input_tokens": 40,
|
| 1704 |
+
"cached_input_tokens": 40,
|
| 1705 |
+
"input_tokens": 40,
|
| 1706 |
+
"output_tokens": 40,
|
| 1707 |
+
"reasoning_output_tokens": 40,
|
| 1708 |
+
"reasoning_reported_output_tokens": 40
|
| 1709 |
+
},
|
| 1710 |
+
"response_count": 40,
|
| 1711 |
+
"response_ids_complete": true,
|
| 1712 |
+
"schema": "rlebench/token-usage/1",
|
| 1713 |
+
"source": "Codex token_usage_record per response",
|
| 1714 |
+
"uncached_input_tokens": 55670,
|
| 1715 |
+
"unidentified_usage_records": 0
|
| 1716 |
+
},
|
| 1717 |
+
"kinex_usage": {
|
| 1718 |
+
"input_tokens": 1201086,
|
| 1719 |
+
"output_tokens": 6987,
|
| 1720 |
+
"cached_input_tokens": 1119616,
|
| 1721 |
+
"cache_write_input_tokens": 0,
|
| 1722 |
+
"reasoning_output_tokens": 2114,
|
| 1723 |
+
"cache_reported_input_tokens": 1201086,
|
| 1724 |
+
"cache_write_reported_input_tokens": 1201086,
|
| 1725 |
+
"reasoning_reported_output_tokens": 6987,
|
| 1726 |
+
"schema": "rlebench/token-usage/1",
|
| 1727 |
+
"source": "Provider terminal response usage",
|
| 1728 |
+
"response_count": 43,
|
| 1729 |
+
"reported_responses": {
|
| 1730 |
+
"input_tokens": 43,
|
| 1731 |
+
"output_tokens": 43,
|
| 1732 |
+
"cached_input_tokens": 43,
|
| 1733 |
+
"cache_write_input_tokens": 43,
|
| 1734 |
+
"reasoning_output_tokens": 43,
|
| 1735 |
+
"cache_reported_input_tokens": 43,
|
| 1736 |
+
"cache_write_reported_input_tokens": 43,
|
| 1737 |
+
"reasoning_reported_output_tokens": 43
|
| 1738 |
+
},
|
| 1739 |
+
"known_cached_input_tokens": 1119616,
|
| 1740 |
+
"known_input_tokens": 1201086,
|
| 1741 |
+
"known_output_tokens": 6987,
|
| 1742 |
+
"known_cache_write_input_tokens": 0,
|
| 1743 |
+
"known_reasoning_output_tokens": 2114,
|
| 1744 |
+
"cache_hit_rate": 0.9321697197369714,
|
| 1745 |
+
"uncached_input_tokens": 81470,
|
| 1746 |
+
"cost_usd": null,
|
| 1747 |
+
"response_ids_complete": true,
|
| 1748 |
+
"accounting": "reported-responses",
|
| 1749 |
+
"usage_source": "provider-raw",
|
| 1750 |
+
"request_attempts": 44,
|
| 1751 |
+
"rejected_attempts": 0,
|
| 1752 |
+
"unresolved_attempts": [],
|
| 1753 |
+
"audit_complete": true
|
| 1754 |
+
},
|
| 1755 |
+
"codex_wall_time_s": 412.342365,
|
| 1756 |
+
"kinex_wall_time_s": 389.816733,
|
| 1757 |
+
"kinex_version": "0.10.3",
|
| 1758 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1759 |
+
"same_roboenv_source": true
|
| 1760 |
},
|
| 1761 |
{
|
| 1762 |
+
"task_key": "task02/09",
|
| 1763 |
+
"instruction": "put both moka pots on the stove and turn the stove on",
|
| 1764 |
+
"instruction_policy": "modified",
|
| 1765 |
"codex_success": true,
|
| 1766 |
"kinex_success": true,
|
| 1767 |
+
"codex_steps": 560,
|
| 1768 |
+
"kinex_steps": 890,
|
| 1769 |
+
"codex_usage": {
|
| 1770 |
+
"accounting": "reported-responses",
|
| 1771 |
+
"audit_complete": true,
|
| 1772 |
+
"cache_hit_rate": 0.9551717429852196,
|
| 1773 |
+
"cache_reported_input_tokens": 714527,
|
| 1774 |
+
"cache_write_input_tokens": 0,
|
| 1775 |
+
"cache_write_reported_input_tokens": 714527,
|
| 1776 |
+
"cached_input_tokens": 682496,
|
| 1777 |
+
"completed_turns": 1,
|
| 1778 |
+
"cost_usd": null,
|
| 1779 |
+
"failed_turns": 0,
|
| 1780 |
+
"input_tokens": 714527,
|
| 1781 |
+
"known_cache_write_input_tokens": 0,
|
| 1782 |
+
"known_cached_input_tokens": 682496,
|
| 1783 |
+
"known_input_tokens": 714527,
|
| 1784 |
+
"known_output_tokens": 5355,
|
| 1785 |
+
"known_reasoning_output_tokens": 1306,
|
| 1786 |
+
"output_tokens": 5355,
|
| 1787 |
+
"reasoning_output_tokens": 1306,
|
| 1788 |
+
"reasoning_reported_output_tokens": 5355,
|
| 1789 |
+
"reported_responses": {
|
| 1790 |
+
"cache_reported_input_tokens": 24,
|
| 1791 |
+
"cache_write_input_tokens": 24,
|
| 1792 |
+
"cache_write_reported_input_tokens": 24,
|
| 1793 |
+
"cached_input_tokens": 24,
|
| 1794 |
+
"input_tokens": 24,
|
| 1795 |
+
"output_tokens": 24,
|
| 1796 |
+
"reasoning_output_tokens": 24,
|
| 1797 |
+
"reasoning_reported_output_tokens": 24
|
| 1798 |
+
},
|
| 1799 |
+
"response_count": 24,
|
| 1800 |
+
"response_ids_complete": true,
|
| 1801 |
+
"schema": "rlebench/token-usage/1",
|
| 1802 |
+
"source": "Codex token_usage_record per response",
|
| 1803 |
+
"uncached_input_tokens": 32031,
|
| 1804 |
+
"unidentified_usage_records": 0
|
| 1805 |
+
},
|
| 1806 |
+
"kinex_usage": {
|
| 1807 |
+
"input_tokens": 1251289,
|
| 1808 |
+
"output_tokens": 7233,
|
| 1809 |
+
"cached_input_tokens": 1150848,
|
| 1810 |
+
"cache_write_input_tokens": 0,
|
| 1811 |
+
"reasoning_output_tokens": 2182,
|
| 1812 |
+
"cache_reported_input_tokens": 1251289,
|
| 1813 |
+
"cache_write_reported_input_tokens": 1251289,
|
| 1814 |
+
"reasoning_reported_output_tokens": 7233,
|
| 1815 |
+
"schema": "rlebench/token-usage/1",
|
| 1816 |
+
"source": "Provider terminal response usage",
|
| 1817 |
+
"response_count": 47,
|
| 1818 |
+
"reported_responses": {
|
| 1819 |
+
"input_tokens": 47,
|
| 1820 |
+
"output_tokens": 47,
|
| 1821 |
+
"cached_input_tokens": 47,
|
| 1822 |
+
"cache_write_input_tokens": 47,
|
| 1823 |
+
"reasoning_output_tokens": 47,
|
| 1824 |
+
"cache_reported_input_tokens": 47,
|
| 1825 |
+
"cache_write_reported_input_tokens": 47,
|
| 1826 |
+
"reasoning_reported_output_tokens": 47
|
| 1827 |
+
},
|
| 1828 |
+
"known_cached_input_tokens": 1150848,
|
| 1829 |
+
"known_input_tokens": 1251289,
|
| 1830 |
+
"known_output_tokens": 7233,
|
| 1831 |
+
"known_cache_write_input_tokens": 0,
|
| 1832 |
+
"known_reasoning_output_tokens": 2182,
|
| 1833 |
+
"cache_hit_rate": 0.9197299744503468,
|
| 1834 |
+
"uncached_input_tokens": 100441,
|
| 1835 |
+
"cost_usd": null,
|
| 1836 |
+
"response_ids_complete": true,
|
| 1837 |
+
"accounting": "reported-responses",
|
| 1838 |
+
"usage_source": "provider-raw",
|
| 1839 |
+
"request_attempts": 47,
|
| 1840 |
+
"rejected_attempts": 0,
|
| 1841 |
+
"unresolved_attempts": [],
|
| 1842 |
+
"audit_complete": true
|
| 1843 |
+
},
|
| 1844 |
+
"codex_wall_time_s": 216.244925,
|
| 1845 |
+
"kinex_wall_time_s": 412.870858,
|
| 1846 |
+
"kinex_version": "0.10.3",
|
| 1847 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1848 |
+
"same_roboenv_source": true
|
| 1849 |
},
|
| 1850 |
{
|
| 1851 |
+
"task_key": "task02/10",
|
| 1852 |
+
"instruction": "put the yellow and white mug in the microwave and close it",
|
| 1853 |
+
"instruction_policy": "original_native",
|
| 1854 |
"codex_success": true,
|
| 1855 |
"kinex_success": true,
|
| 1856 |
+
"codex_steps": 1210,
|
| 1857 |
+
"kinex_steps": 1950,
|
| 1858 |
+
"codex_usage": {
|
| 1859 |
+
"accounting": "reported-responses",
|
| 1860 |
+
"audit_complete": true,
|
| 1861 |
+
"cache_hit_rate": 0.9540138123860797,
|
| 1862 |
+
"cache_reported_input_tokens": 1349079,
|
| 1863 |
+
"cache_write_input_tokens": 0,
|
| 1864 |
+
"cache_write_reported_input_tokens": 1349079,
|
| 1865 |
+
"cached_input_tokens": 1287040,
|
| 1866 |
+
"completed_turns": 1,
|
| 1867 |
+
"cost_usd": null,
|
| 1868 |
+
"failed_turns": 0,
|
| 1869 |
+
"input_tokens": 1349079,
|
| 1870 |
+
"known_cache_write_input_tokens": 0,
|
| 1871 |
+
"known_cached_input_tokens": 1287040,
|
| 1872 |
+
"known_input_tokens": 1349079,
|
| 1873 |
+
"known_output_tokens": 10556,
|
| 1874 |
+
"known_reasoning_output_tokens": 4068,
|
| 1875 |
+
"output_tokens": 10556,
|
| 1876 |
+
"reasoning_output_tokens": 4068,
|
| 1877 |
+
"reasoning_reported_output_tokens": 10556,
|
| 1878 |
+
"reported_responses": {
|
| 1879 |
+
"cache_reported_input_tokens": 39,
|
| 1880 |
+
"cache_write_input_tokens": 39,
|
| 1881 |
+
"cache_write_reported_input_tokens": 39,
|
| 1882 |
+
"cached_input_tokens": 39,
|
| 1883 |
+
"input_tokens": 39,
|
| 1884 |
+
"output_tokens": 39,
|
| 1885 |
+
"reasoning_output_tokens": 39,
|
| 1886 |
+
"reasoning_reported_output_tokens": 39
|
| 1887 |
+
},
|
| 1888 |
+
"response_count": 39,
|
| 1889 |
+
"response_ids_complete": true,
|
| 1890 |
+
"schema": "rlebench/token-usage/1",
|
| 1891 |
+
"source": "Codex token_usage_record per response",
|
| 1892 |
+
"uncached_input_tokens": 62039,
|
| 1893 |
+
"unidentified_usage_records": 0
|
| 1894 |
+
},
|
| 1895 |
+
"kinex_usage": {
|
| 1896 |
+
"input_tokens": 3336633,
|
| 1897 |
+
"output_tokens": 12622,
|
| 1898 |
+
"cached_input_tokens": 3199872,
|
| 1899 |
+
"cache_write_input_tokens": 0,
|
| 1900 |
+
"reasoning_output_tokens": 3738,
|
| 1901 |
+
"cache_reported_input_tokens": 3336633,
|
| 1902 |
+
"cache_write_reported_input_tokens": 3336633,
|
| 1903 |
+
"reasoning_reported_output_tokens": 12622,
|
| 1904 |
+
"schema": "rlebench/token-usage/1",
|
| 1905 |
+
"source": "Provider terminal response usage",
|
| 1906 |
+
"response_count": 82,
|
| 1907 |
+
"reported_responses": {
|
| 1908 |
+
"input_tokens": 82,
|
| 1909 |
+
"output_tokens": 82,
|
| 1910 |
+
"cached_input_tokens": 82,
|
| 1911 |
+
"cache_write_input_tokens": 82,
|
| 1912 |
+
"reasoning_output_tokens": 82,
|
| 1913 |
+
"cache_reported_input_tokens": 82,
|
| 1914 |
+
"cache_write_reported_input_tokens": 82,
|
| 1915 |
+
"reasoning_reported_output_tokens": 82
|
| 1916 |
+
},
|
| 1917 |
+
"known_cached_input_tokens": 3199872,
|
| 1918 |
+
"known_input_tokens": 3336633,
|
| 1919 |
+
"known_output_tokens": 12622,
|
| 1920 |
+
"known_cache_write_input_tokens": 0,
|
| 1921 |
+
"known_reasoning_output_tokens": 3738,
|
| 1922 |
+
"cache_hit_rate": 0.9590122737502147,
|
| 1923 |
+
"uncached_input_tokens": 136761,
|
| 1924 |
+
"cost_usd": null,
|
| 1925 |
+
"response_ids_complete": true,
|
| 1926 |
+
"accounting": "reported-responses",
|
| 1927 |
+
"usage_source": "provider-raw",
|
| 1928 |
+
"request_attempts": 82,
|
| 1929 |
+
"rejected_attempts": 0,
|
| 1930 |
+
"unresolved_attempts": [],
|
| 1931 |
+
"audit_complete": true
|
| 1932 |
+
},
|
| 1933 |
+
"codex_wall_time_s": 399.959252,
|
| 1934 |
+
"kinex_wall_time_s": 717.324727,
|
| 1935 |
+
"kinex_version": "0.10.3",
|
| 1936 |
+
"kinex_revision": "caac19a8a36272972f762e0f74cfe381e8500048",
|
| 1937 |
+
"same_roboenv_source": true
|
| 1938 |
}
|
| 1939 |
],
|
| 1940 |
+
"preserved_robodojo_comparison": {
|
| 1941 |
+
"baseline_revision": "1a7f1c9f178ba7aeb659e5b527a77978e9136072",
|
| 1942 |
+
"summary": {
|
| 1943 |
+
"completed": 42,
|
| 1944 |
+
"codex_successes": 31,
|
| 1945 |
+
"codex_success_rate_completed_subset": 0.7380952380952381,
|
| 1946 |
+
"kinex_successes_same_subset": 35,
|
| 1947 |
+
"kinex_successes_all42": 35,
|
| 1948 |
+
"all42_comparison_complete": true
|
| 1949 |
+
},
|
| 1950 |
+
"tasks": [
|
| 1951 |
+
{
|
| 1952 |
+
"task_key": "task04/01",
|
| 1953 |
+
"native_id": "robodojo/make-toast",
|
| 1954 |
+
"status": "completed",
|
| 1955 |
+
"codex_success": false,
|
| 1956 |
+
"kinex_success": false,
|
| 1957 |
+
"kinex_version": "0.10.3"
|
| 1958 |
+
},
|
| 1959 |
+
{
|
| 1960 |
+
"task_key": "task04/02",
|
| 1961 |
+
"native_id": "robodojo/classify-objects-by-language",
|
| 1962 |
+
"status": "completed",
|
| 1963 |
+
"codex_success": true,
|
| 1964 |
+
"kinex_success": true,
|
| 1965 |
+
"kinex_version": "0.10.0"
|
| 1966 |
+
},
|
| 1967 |
+
{
|
| 1968 |
+
"task_key": "task04/03",
|
| 1969 |
+
"native_id": "robodojo/store-laptop-and-headphones",
|
| 1970 |
+
"status": "completed",
|
| 1971 |
+
"codex_success": false,
|
| 1972 |
+
"kinex_success": false,
|
| 1973 |
+
"kinex_version": "0.10.3"
|
| 1974 |
+
},
|
| 1975 |
+
{
|
| 1976 |
+
"task_key": "task04/04",
|
| 1977 |
+
"native_id": "robodojo/cover-blocks",
|
| 1978 |
+
"status": "completed",
|
| 1979 |
+
"codex_success": true,
|
| 1980 |
+
"kinex_success": true,
|
| 1981 |
+
"kinex_version": "0.10.2"
|
| 1982 |
+
},
|
| 1983 |
+
{
|
| 1984 |
+
"task_key": "task04/05",
|
| 1985 |
+
"native_id": "robodojo/play-xylophone",
|
| 1986 |
+
"status": "completed",
|
| 1987 |
+
"codex_success": true,
|
| 1988 |
+
"kinex_success": true,
|
| 1989 |
+
"kinex_version": "0.10.0"
|
| 1990 |
+
},
|
| 1991 |
+
{
|
| 1992 |
+
"task_key": "task04/06",
|
| 1993 |
+
"native_id": "robodojo/store-tools-in-toolbox",
|
| 1994 |
+
"status": "completed",
|
| 1995 |
+
"codex_success": false,
|
| 1996 |
+
"kinex_success": false,
|
| 1997 |
+
"kinex_version": "0.10.3"
|
| 1998 |
+
},
|
| 1999 |
+
{
|
| 2000 |
+
"task_key": "task04/07",
|
| 2001 |
+
"native_id": "robodojo/insert-tubes",
|
| 2002 |
+
"status": "completed",
|
| 2003 |
+
"codex_success": true,
|
| 2004 |
+
"kinex_success": true,
|
| 2005 |
+
"kinex_version": "0.10.3"
|
| 2006 |
+
},
|
| 2007 |
+
{
|
| 2008 |
+
"task_key": "task04/08",
|
| 2009 |
+
"native_id": "robodojo/deposit-coin",
|
| 2010 |
+
"status": "completed",
|
| 2011 |
+
"codex_success": true,
|
| 2012 |
+
"kinex_success": true,
|
| 2013 |
+
"kinex_version": "0.10.2"
|
| 2014 |
+
},
|
| 2015 |
+
{
|
| 2016 |
+
"task_key": "task04/09",
|
| 2017 |
+
"native_id": "robodojo/fasten-screws",
|
| 2018 |
+
"status": "completed",
|
| 2019 |
+
"codex_success": true,
|
| 2020 |
+
"kinex_success": true,
|
| 2021 |
+
"kinex_version": "0.10.3"
|
| 2022 |
+
},
|
| 2023 |
+
{
|
| 2024 |
+
"task_key": "task04/10",
|
| 2025 |
+
"native_id": "robodojo/play-stacking-toy",
|
| 2026 |
+
"status": "completed",
|
| 2027 |
+
"codex_success": true,
|
| 2028 |
+
"kinex_success": true,
|
| 2029 |
+
"kinex_version": "0.10.3"
|
| 2030 |
+
},
|
| 2031 |
+
{
|
| 2032 |
+
"task_key": "task04/11",
|
| 2033 |
+
"native_id": "robodojo/align-blocks",
|
| 2034 |
+
"status": "completed",
|
| 2035 |
+
"codex_success": true,
|
| 2036 |
+
"kinex_success": true,
|
| 2037 |
+
"kinex_version": "0.10.0"
|
| 2038 |
+
},
|
| 2039 |
+
{
|
| 2040 |
+
"task_key": "task04/12",
|
| 2041 |
+
"native_id": "robodojo/arrange-largest-number",
|
| 2042 |
+
"status": "completed",
|
| 2043 |
+
"codex_success": true,
|
| 2044 |
+
"kinex_success": true,
|
| 2045 |
+
"kinex_version": "0.10.3"
|
| 2046 |
+
},
|
| 2047 |
+
{
|
| 2048 |
+
"task_key": "task04/14",
|
| 2049 |
+
"native_id": "robodojo/build-tower",
|
| 2050 |
+
"status": "completed",
|
| 2051 |
+
"codex_success": true,
|
| 2052 |
+
"kinex_success": true,
|
| 2053 |
+
"kinex_version": "0.10.3"
|
| 2054 |
+
},
|
| 2055 |
+
{
|
| 2056 |
+
"task_key": "task04/15",
|
| 2057 |
+
"native_id": "robodojo/classify-objects",
|
| 2058 |
+
"status": "completed",
|
| 2059 |
+
"codex_success": true,
|
| 2060 |
+
"kinex_success": true,
|
| 2061 |
+
"kinex_version": "0.10.2"
|
| 2062 |
+
},
|
| 2063 |
+
{
|
| 2064 |
+
"task_key": "task04/16",
|
| 2065 |
+
"native_id": "robodojo/fill-egg-holder",
|
| 2066 |
+
"status": "completed",
|
| 2067 |
+
"codex_success": true,
|
| 2068 |
+
"kinex_success": false,
|
| 2069 |
+
"kinex_version": "0.10.3"
|
| 2070 |
+
},
|
| 2071 |
+
{
|
| 2072 |
+
"task_key": "task04/17",
|
| 2073 |
+
"native_id": "robodojo/fill-pen-holder",
|
| 2074 |
+
"status": "completed",
|
| 2075 |
+
"codex_success": false,
|
| 2076 |
+
"kinex_success": false,
|
| 2077 |
+
"kinex_version": "0.10.3"
|
| 2078 |
+
},
|
| 2079 |
+
{
|
| 2080 |
+
"task_key": "task04/18",
|
| 2081 |
+
"native_id": "robodojo/fold-clothes",
|
| 2082 |
+
"status": "completed",
|
| 2083 |
+
"codex_success": true,
|
| 2084 |
+
"kinex_success": true,
|
| 2085 |
+
"kinex_version": "0.10.3"
|
| 2086 |
+
},
|
| 2087 |
+
{
|
| 2088 |
+
"task_key": "task04/20",
|
| 2089 |
+
"native_id": "robodojo/general-pickup",
|
| 2090 |
+
"status": "completed",
|
| 2091 |
+
"codex_success": true,
|
| 2092 |
+
"kinex_success": true,
|
| 2093 |
+
"kinex_version": "0.10.0"
|
| 2094 |
+
},
|
| 2095 |
+
{
|
| 2096 |
+
"task_key": "task04/21",
|
| 2097 |
+
"native_id": "robodojo/hang-mugs",
|
| 2098 |
+
"status": "completed",
|
| 2099 |
+
"codex_success": true,
|
| 2100 |
+
"kinex_success": true,
|
| 2101 |
+
"kinex_version": "0.10.3"
|
| 2102 |
+
},
|
| 2103 |
+
{
|
| 2104 |
+
"task_key": "task04/23",
|
| 2105 |
+
"native_id": "robodojo/imitate-sorting-sequence",
|
| 2106 |
+
"status": "completed",
|
| 2107 |
+
"codex_success": true,
|
| 2108 |
+
"kinex_success": true,
|
| 2109 |
+
"kinex_version": "0.10.2"
|
| 2110 |
+
},
|
| 2111 |
+
{
|
| 2112 |
+
"task_key": "task04/24",
|
| 2113 |
+
"native_id": "robodojo/insert-key",
|
| 2114 |
+
"status": "completed",
|
| 2115 |
+
"codex_success": true,
|
| 2116 |
+
"kinex_success": true,
|
| 2117 |
+
"kinex_version": "0.10.2"
|
| 2118 |
+
},
|
| 2119 |
+
{
|
| 2120 |
+
"task_key": "task04/25",
|
| 2121 |
+
"native_id": "robodojo/make-kong",
|
| 2122 |
+
"status": "completed",
|
| 2123 |
+
"codex_success": false,
|
| 2124 |
+
"kinex_success": true,
|
| 2125 |
+
"kinex_version": "0.10.3"
|
| 2126 |
+
},
|
| 2127 |
+
{
|
| 2128 |
+
"task_key": "task04/27",
|
| 2129 |
+
"native_id": "robodojo/match-and-pick-from-conveyor",
|
| 2130 |
+
"status": "completed",
|
| 2131 |
+
"codex_success": true,
|
| 2132 |
+
"kinex_success": true,
|
| 2133 |
+
"kinex_version": "0.10.3"
|
| 2134 |
+
},
|
| 2135 |
+
{
|
| 2136 |
+
"task_key": "task04/28",
|
| 2137 |
+
"native_id": "robodojo/organize-table",
|
| 2138 |
+
"status": "completed",
|
| 2139 |
+
"codex_success": false,
|
| 2140 |
+
"kinex_success": true,
|
| 2141 |
+
"kinex_version": "0.10.3"
|
| 2142 |
+
},
|
| 2143 |
+
{
|
| 2144 |
+
"task_key": "task04/29",
|
| 2145 |
+
"native_id": "robodojo/pack-objects-into-box",
|
| 2146 |
+
"status": "completed",
|
| 2147 |
+
"codex_success": true,
|
| 2148 |
+
"kinex_success": true,
|
| 2149 |
+
"kinex_version": "0.10.2"
|
| 2150 |
+
},
|
| 2151 |
+
{
|
| 2152 |
+
"task_key": "task04/31",
|
| 2153 |
+
"native_id": "robodojo/pick-from-conveyor-by-image",
|
| 2154 |
+
"status": "completed",
|
| 2155 |
+
"codex_success": false,
|
| 2156 |
+
"kinex_success": false,
|
| 2157 |
+
"kinex_version": "0.10.3"
|
| 2158 |
+
},
|
| 2159 |
+
{
|
| 2160 |
+
"task_key": "task04/32",
|
| 2161 |
+
"native_id": "robodojo/play-tic-tac-toe",
|
| 2162 |
+
"status": "completed",
|
| 2163 |
+
"codex_success": true,
|
| 2164 |
+
"kinex_success": true,
|
| 2165 |
+
"kinex_version": "0.10.3"
|
| 2166 |
+
},
|
| 2167 |
+
{
|
| 2168 |
+
"task_key": "task04/33",
|
| 2169 |
+
"native_id": "robodojo/plug-in-charger",
|
| 2170 |
+
"status": "completed",
|
| 2171 |
+
"codex_success": true,
|
| 2172 |
+
"kinex_success": true,
|
| 2173 |
+
"kinex_version": "0.10.3"
|
| 2174 |
+
},
|
| 2175 |
+
{
|
| 2176 |
+
"task_key": "task04/34",
|
| 2177 |
+
"native_id": "robodojo/pour-balls-into-vase",
|
| 2178 |
+
"status": "completed",
|
| 2179 |
+
"codex_success": true,
|
| 2180 |
+
"kinex_success": true,
|
| 2181 |
+
"kinex_version": "0.10.3"
|
| 2182 |
+
},
|
| 2183 |
+
{
|
| 2184 |
+
"task_key": "task04/35",
|
| 2185 |
+
"native_id": "robodojo/pour-by-language",
|
| 2186 |
+
"status": "completed",
|
| 2187 |
+
"codex_success": false,
|
| 2188 |
+
"kinex_success": true,
|
| 2189 |
+
"kinex_version": "0.10.0"
|
| 2190 |
+
},
|
| 2191 |
+
{
|
| 2192 |
+
"task_key": "task04/36",
|
| 2193 |
+
"native_id": "robodojo/pour-liquid-into-cup",
|
| 2194 |
+
"status": "completed",
|
| 2195 |
+
"codex_success": true,
|
| 2196 |
+
"kinex_success": true,
|
| 2197 |
+
"kinex_version": "0.10.3"
|
| 2198 |
+
},
|
| 2199 |
+
{
|
| 2200 |
+
"task_key": "task04/38",
|
| 2201 |
+
"native_id": "robodojo/press-by-number",
|
| 2202 |
+
"status": "completed",
|
| 2203 |
+
"codex_success": true,
|
| 2204 |
+
"kinex_success": true,
|
| 2205 |
+
"kinex_version": "0.10.3"
|
| 2206 |
+
},
|
| 2207 |
+
{
|
| 2208 |
+
"task_key": "task04/39",
|
| 2209 |
+
"native_id": "robodojo/push-t",
|
| 2210 |
+
"status": "completed",
|
| 2211 |
+
"codex_success": false,
|
| 2212 |
+
"kinex_success": true,
|
| 2213 |
+
"kinex_version": "0.10.3"
|
| 2214 |
+
},
|
| 2215 |
+
{
|
| 2216 |
+
"task_key": "task04/41",
|
| 2217 |
+
"native_id": "robodojo/put-bottles-into-dustbin",
|
| 2218 |
+
"status": "completed",
|
| 2219 |
+
"codex_success": true,
|
| 2220 |
+
"kinex_success": true,
|
| 2221 |
+
"kinex_version": "0.10.2"
|
| 2222 |
+
},
|
| 2223 |
+
{
|
| 2224 |
+
"task_key": "task04/42",
|
| 2225 |
+
"native_id": "robodojo/solve-equation",
|
| 2226 |
+
"status": "completed",
|
| 2227 |
+
"codex_success": true,
|
| 2228 |
+
"kinex_success": true,
|
| 2229 |
+
"kinex_version": "0.10.0"
|
| 2230 |
+
},
|
| 2231 |
+
{
|
| 2232 |
+
"task_key": "task04/43",
|
| 2233 |
+
"native_id": "robodojo/sort-nesting-dolls-by-size",
|
| 2234 |
+
"status": "completed",
|
| 2235 |
+
"codex_success": true,
|
| 2236 |
+
"kinex_success": true,
|
| 2237 |
+
"kinex_version": "0.10.3"
|
| 2238 |
+
},
|
| 2239 |
+
{
|
| 2240 |
+
"task_key": "task04/45",
|
| 2241 |
+
"native_id": "robodojo/stack-blocks",
|
| 2242 |
+
"status": "completed",
|
| 2243 |
+
"codex_success": true,
|
| 2244 |
+
"kinex_success": true,
|
| 2245 |
+
"kinex_version": "0.10.2"
|
| 2246 |
+
},
|
| 2247 |
+
{
|
| 2248 |
+
"task_key": "task04/46",
|
| 2249 |
+
"native_id": "robodojo/stack-blocks-by-language",
|
| 2250 |
+
"status": "completed",
|
| 2251 |
+
"codex_success": true,
|
| 2252 |
+
"kinex_success": true,
|
| 2253 |
+
"kinex_version": "0.10.0"
|
| 2254 |
+
},
|
| 2255 |
+
{
|
| 2256 |
+
"task_key": "task04/48",
|
| 2257 |
+
"native_id": "robodojo/stack-bowls",
|
| 2258 |
+
"status": "completed",
|
| 2259 |
+
"codex_success": true,
|
| 2260 |
+
"kinex_success": true,
|
| 2261 |
+
"kinex_version": "0.10.2"
|
| 2262 |
+
},
|
| 2263 |
+
{
|
| 2264 |
+
"task_key": "task04/51",
|
| 2265 |
+
"native_id": "robodojo/swap-t",
|
| 2266 |
+
"status": "completed",
|
| 2267 |
+
"codex_success": true,
|
| 2268 |
+
"kinex_success": true,
|
| 2269 |
+
"kinex_version": "0.10.2"
|
| 2270 |
+
},
|
| 2271 |
+
{
|
| 2272 |
+
"task_key": "task04/52",
|
| 2273 |
+
"native_id": "robodojo/swap-blocks",
|
| 2274 |
+
"status": "completed",
|
| 2275 |
+
"codex_success": false,
|
| 2276 |
+
"kinex_success": false,
|
| 2277 |
+
"kinex_version": "0.10.3"
|
| 2278 |
+
},
|
| 2279 |
+
{
|
| 2280 |
+
"task_key": "task04/53",
|
| 2281 |
+
"native_id": "robodojo/sweep-blocks",
|
| 2282 |
+
"status": "completed",
|
| 2283 |
+
"codex_success": false,
|
| 2284 |
+
"kinex_success": true,
|
| 2285 |
+
"kinex_version": "0.10.2"
|
| 2286 |
+
}
|
| 2287 |
+
],
|
| 2288 |
+
"limitations": [
|
| 2289 |
+
"Historical Kinex baseline mixes three runtime versions.",
|
| 2290 |
+
"24 tasks share RoboEnv source/build; 18 use an older build.",
|
| 2291 |
+
"The first six completed Codex tasks used original request defaults; later attempts use 50 retries."
|
| 2292 |
+
]
|
| 2293 |
+
},
|
| 2294 |
"limitations": [
|
| 2295 |
+
"Single seed and one native episode per task; these are observed success rates, not population estimates.",
|
| 2296 |
+
"Wall time is operational latency under different concurrency, not a controlled speed benchmark; Kinex used up to five concurrent tasks and the new Codex dispatcher admits up to twenty.",
|
| 2297 |
+
"LIBERO Kinex uses caac19a. The retained RoboCasa Kinex reference uses a1eaece; the two intervening Kinex commits change session timing and footer display.",
|
| 2298 |
+
"The ordinary Codex baseline uses its native exec tools and the public RoboEnv SDK/CLI. Kinex uses its robotics tools.",
|
| 2299 |
+
"RoboDojo retains its previously published protocol and historical source differences."
|
| 2300 |
]
|
| 2301 |
}
|
data.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
episodes.csv
CHANGED
|
@@ -1,4 +1,24 @@
|
|
| 1 |
task_key,native_id,status,run_status,success,steps,input_tokens,cached_input_tokens,output_tokens
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
task04/01,robodojo/make-toast,completed,finished,False,6884,13262282,13103360,42295
|
| 3 |
task04/02,robodojo/classify-objects-by-language,completed,finished,True,1642,1667619,1598848,9371
|
| 4 |
task04/03,robodojo/store-laptop-and-headphones,completed,finished,False,5696,18189657,17666048,45612
|
|
|
|
| 1 |
task_key,native_id,status,run_status,success,steps,input_tokens,cached_input_tokens,output_tokens
|
| 2 |
+
task01/01,robocasa/load-condiments-in-fridge,completed,finished,False,5496,7077229,6989312,20496
|
| 3 |
+
task01/02,robocasa/filter-microwavable-item,completed,finished,False,5903,4438106,4354176,22048
|
| 4 |
+
task01/03,robocasa/store-dumplings,completed,finished,False,5053,11115392,11003904,29598
|
| 5 |
+
task01/04,robocasa/divide-buffet-trays,completed,finished,False,5435,11714063,11596032,32111
|
| 6 |
+
task01/05,robocasa/make-cheesecake-filling,completed,finished,True,5224,12355153,12213888,35229
|
| 7 |
+
task01/06,robocasa/multistep-steaming,completed,finished,False,5416,9630097,9514624,24897
|
| 8 |
+
task01/07,robocasa/scale-portioning,completed,finished,False,1426,1217225,1176320,8600
|
| 9 |
+
task01/08,robocasa/scrub-cutting-board,completed,finished,False,1212,775611,745088,6607
|
| 10 |
+
task01/09,robocasa/prepare-veggie-dip,completed,finished,False,2216,1973478,1920896,11695
|
| 11 |
+
task01/10,robocasa/prepare-vegetable-roasting,completed,finished,True,3969,2989000,2922496,13452
|
| 12 |
+
task02/01,libero/10-0,completed,finished,True,969,832121,791296,6802
|
| 13 |
+
task02/02,libero/10-1,completed,finished,True,722,784518,746496,6225
|
| 14 |
+
task02/03,libero/10-2,completed,finished,True,2013,2225094,2152064,13673
|
| 15 |
+
task02/04,libero/bowl-into-bottom-drawer,completed,finished,False,3687,3356377,3271040,23446
|
| 16 |
+
task02/05,libero/10-4,completed,finished,True,634,509662,483840,5212
|
| 17 |
+
task02/06,libero/10-5,completed,finished,True,325,707210,676224,6550
|
| 18 |
+
task02/07,libero/10-6,completed,finished,True,1016,701841,665856,6754
|
| 19 |
+
task02/08,libero/10-7,completed,finished,True,1115,1406070,1350400,9396
|
| 20 |
+
task02/09,libero/10-8,completed,finished,True,560,714527,682496,5355
|
| 21 |
+
task02/10,libero/mug-into-microwave,completed,finished,True,1210,1349079,1287040,10556
|
| 22 |
task04/01,robodojo/make-toast,completed,finished,False,6884,13262282,13103360,42295
|
| 23 |
task04/02,robodojo/classify-objects-by-language,completed,finished,True,1642,1667619,1598848,9371
|
| 24 |
task04/03,robodojo/store-laptop-and-headphones,completed,finished,False,5696,18189657,17666048,45612
|
episodes.json
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|
index.html
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
<!doctype html>
|
| 2 |
-
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><meta name="description" content="Codex robotics benchmark: original and revised instructions, seed 0, GPT-6 Astra high. Explore task videos, sessions and evidence
|
| 3 |
-
<body><header class="hero"><div class="wrap"><div class="topbar"><a class="brand" href="#"><span class="brand-icon" aria-hidden="true">◇</span> RLE-Bench</a><nav aria-label="Primary"><a href="#tasks">Tasks</a><a href="https://huggingface.co/datasets/RLE-Bench/codex-benchmark" target="_blank" rel="noopener">Dataset ↗</a><a href="REPORT.md" target="_blank" rel="noopener">Report ↗</a></nav></div><div class="hero-content"><div class="eyebrow">ROBOTICS EVALUATION / STOCK CODEX</div><h1>Codex Benchmark</h1><p>GPT-6 Astra · high · seed 0 · original & modified instructions</p><div class="hero-bottom"><span>
|
| 4 |
-
<main class="wrap"><div id="loading" role="status">Loading benchmark…</div><div id="browse" hidden><div class="edition"><span id="edition-count">Loading results…</span><span class="chip" id="edition-label">Benchmark results</span></div><section aria-labelledby="overview-heading"><div class="section-top"><h2 id="overview-heading">Task overview</h2><a href="summary.json" target="_blank" rel="noopener">Summary JSON ↗</a></div><div class="table-wrap"><table><thead><tr><th>Environment</th><th>Coverage</th><th>Native success</th><th>Input tokens</th><th>Output tokens</th><th>Usage coverage</th></tr></thead><tbody id="overview"></tbody></table></div><p class="footnote">Current recorded result per task · Seed 0 · Native success criteria. Each task shows its latest result and the instruction used for that recording. Rates include valid completed results only; pending tasks are excluded. No pooled score across environments. Input includes cached tokens; output includes reasoning tokens.</p></section><section id="failure-overview" lang="zh-CN" aria-label="失败原因"></section><p id="readiness-note" class="readiness-note" hidden></p><div id="tasks" class="task-controls"><nav id="family-nav" aria-label="Task families"></nav><div class="filters"><label class="search-label"><span class="sr-only">Search tasks</span><input id="search" type="search" placeholder="Search
|
| 5 |
<footer class="wrap"><span>RLE-Bench · Codex Benchmark</span><nav aria-label="Downloads"><a href="data.json">Data JSON</a><a href="episodes.csv">CSV</a><a href="DATA_FORMAT.md">Data format</a><a href="THIRD_PARTY_NOTICES.md">Attribution</a><a href="https://huggingface.co/datasets/RLE-Bench/codex-benchmark">Dataset ↗</a></nav></footer></body></html>
|
|
|
|
| 1 |
<!doctype html>
|
| 2 |
+
<html lang="en"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><meta name="description" content="Codex robotics benchmark: original and revised instructions, seed 0, GPT-6 Astra high. Explore task videos, sessions and evidence across RoboCasa, LIBERO and RoboDojo."><title>Codex Benchmark · RLE-Bench</title><link rel="icon" href="favicon.svg"><link rel="stylesheet" href="style.css"><script defer src="app.js"></script></head>
|
| 3 |
+
<body><header class="hero"><div class="wrap"><div class="topbar"><a class="brand" href="#"><span class="brand-icon" aria-hidden="true">◇</span> RLE-Bench</a><nav aria-label="Primary"><a href="#tasks">Tasks</a><a href="https://huggingface.co/datasets/RLE-Bench/codex-benchmark" target="_blank" rel="noopener">Dataset ↗</a><a href="REPORT.md" target="_blank" rel="noopener">Report ↗</a></nav></div><div class="hero-content"><div class="eyebrow">ROBOTICS EVALUATION / STOCK CODEX</div><h1>Codex Benchmark</h1><p>GPT-6 Astra · high · seed 0 · original & modified instructions</p><div class="hero-bottom"><span>62 tasks across three environments. Native success criteria.</span><a class="hero-button" href="#task04">Explore the results <span>↗</span></a></div></div></div></header>
|
| 4 |
+
<main class="wrap"><div id="loading" role="status">Loading benchmark…</div><div id="browse" hidden><div class="edition"><span id="edition-count">Loading results…</span><span class="chip" id="edition-label">Benchmark results</span></div><section aria-labelledby="overview-heading"><div class="section-top"><h2 id="overview-heading">Task overview</h2><a href="summary.json" target="_blank" rel="noopener">Summary JSON ↗</a></div><div class="table-wrap"><table><thead><tr><th>Environment</th><th>Coverage</th><th>Native success</th><th>Input tokens</th><th>Output tokens</th><th>Usage coverage</th></tr></thead><tbody id="overview"></tbody></table></div><p class="footnote">Current recorded result per task · Seed 0 · Native success criteria. Each task shows its latest result and the instruction used for that recording. Rates include valid completed results only; pending tasks are excluded. No pooled score across environments. Input includes cached tokens; output includes reasoning tokens.</p></section><section id="failure-overview" lang="zh-CN" aria-label="失败原因"></section><p id="readiness-note" class="readiness-note" hidden></p><div id="tasks" class="task-controls"><nav id="family-nav" aria-label="Task families"></nav><div class="filters"><label class="search-label"><span class="sr-only">Search tasks</span><input id="search" type="search" placeholder="Search 62 tasks…"></label><label><span class="sr-only">Filter by status</span><select id="status-filter"><option value="all">All tasks</option><option value="completed">Results available</option><option value="pending">Pending</option></select></label><label><span class="sr-only">Filter by failure category</span><select id="failure-filter" aria-label="Filter by failure category"><option value="all">All failure categories</option></select></label><span id="filter-count" aria-live="polite"></span></div></div><div id="families"></div><p id="empty" hidden>No tasks match this search.</p><section class="protocol-box"><h2>The evaluation protocol</h2><div class="protocol-grid"><div><b>Instruction versions</b><p>Modified instructions are labeled, with changes highlighted. Each recording retains the instruction it used; native success predicates are unchanged.</p></div><div><b>Single seed</b><p>One independent episode, seed 0. GPT-6 Astra with high reasoning effort. Stock Codex CLI 0.160.0; each task starts with a fresh workspace and no inherited trajectory.</p></div><div><b>Native control budgets</b><p>RoboCasa and LIBERO: 6,000 steps at 20 Hz. RoboDojo: 7,500 frames at 25 Hz. Stepped simulation, with an eight-hour wall-clock limit.</p></div></div></section></div><article id="detail" hidden></article></main>
|
| 5 |
<footer class="wrap"><span>RLE-Bench · Codex Benchmark</span><nav aria-label="Downloads"><a href="data.json">Data JSON</a><a href="episodes.csv">CSV</a><a href="DATA_FORMAT.md">Data format</a><a href="THIRD_PARTY_NOTICES.md">Attribution</a><a href="https://huggingface.co/datasets/RLE-Bench/codex-benchmark">Dataset ↗</a></nav></footer></body></html>
|
licenses/LIBERO-LICENSE.txt
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2023 Lifelong Robot Learning
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
licenses/robocasa-LICENSE.txt
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2026 the RoboCasa Team
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
| 22 |
+
|
| 23 |
+
This software includes the partial implementation of Deepmind Mujoco https://github.com/deepmind/mujoco.
|
| 24 |
+
Deepmind Mujoco is licensed under the Apache License, Version 2.0 (the "License");
|
| 25 |
+
you may not use the files except in compliance with the License.
|
| 26 |
+
|
| 27 |
+
You may obtain a copy of the License at
|
| 28 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
licenses/robosuite-LICENSE.txt
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2022 Stanford Vision and Learning Lab and UT Robot Perception and Learning Lab
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
| 22 |
+
|
| 23 |
+
This software includes the partial implementation of Deepmind Mujoco https://github.com/deepmind/mujoco.
|
| 24 |
+
Deepmind Mujoco is licensed under the Apache License, Version 2.0 (the "License");
|
| 25 |
+
you may not use the files except in compliance with the License.
|
| 26 |
+
|
| 27 |
+
You may obtain a copy of the License at
|
| 28 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
manifest.json
CHANGED
|
@@ -1,21 +1,24 @@
|
|
| 1 |
{
|
| 2 |
".gitattributes": "51dec59a208c8656ce7d6614b7b547022f7a4720b1cdad6a20aeeda717ec090b",
|
| 3 |
-
"DATA_FORMAT.md": "
|
| 4 |
-
"README.md": "
|
| 5 |
-
"REPORT.md": "
|
| 6 |
-
"THIRD_PARTY_NOTICES.md": "
|
| 7 |
-
"app.js": "
|
| 8 |
"attempt-history.json": "8555f458c1dd362f5a07447342994fddf265ff23bbc40e9bd6f9984a892ddb4e",
|
| 9 |
-
"comparison.json": "
|
| 10 |
-
"data.json": "
|
| 11 |
-
"episodes.csv": "
|
| 12 |
-
"episodes.json": "
|
| 13 |
"failure-reviews.json": "c66a997c9ecd6f7f43f130a8a33cd64c7a5b2049033e9ccde653ee30e4796f23",
|
| 14 |
"favicon.svg": "46063791218655f70c828ef71a1232a34fd38181da938338badec929ad226819",
|
| 15 |
-
"index.html": "
|
|
|
|
| 16 |
"licenses/RoboDojo-LICENSE.txt": "7794bb06af8fe5485ca912454ad1665ccd7846c45f9c2b1258658e480948cdbe",
|
| 17 |
-
"
|
|
|
|
|
|
|
| 18 |
"style.css": "dd31b30951c220ebb82bc96224b5e640913945679e42b2e0a42e2c373a329e6c",
|
| 19 |
-
"summary.json": "
|
| 20 |
-
"task-index.json": "
|
| 21 |
}
|
|
|
|
| 1 |
{
|
| 2 |
".gitattributes": "51dec59a208c8656ce7d6614b7b547022f7a4720b1cdad6a20aeeda717ec090b",
|
| 3 |
+
"DATA_FORMAT.md": "e2fd4f6b5c8be9d9e008db7d0c1f5d52b419d68328a1c86625c9f67c49b03403",
|
| 4 |
+
"README.md": "43d89b6c59fdee0a1f5bddf3bd21cf0d809fee832f115771da95389f40f421b8",
|
| 5 |
+
"REPORT.md": "696454f04ac34467d2313e28b1e58b7095b02dc09e2b33b4b597594e10e3b59d",
|
| 6 |
+
"THIRD_PARTY_NOTICES.md": "8f38aa1ab2e03203b438e095a6325b3914541efe434290a9a2f786087d8f657d",
|
| 7 |
+
"app.js": "eb013cc9d24382410953a1a8c81e446a6592a81b32887749d17bfc863936556b",
|
| 8 |
"attempt-history.json": "8555f458c1dd362f5a07447342994fddf265ff23bbc40e9bd6f9984a892ddb4e",
|
| 9 |
+
"comparison.json": "012c7a64ecfde3352d258e332eae3b52de3d0c253663841e1a7dc9af3b3dfab7",
|
| 10 |
+
"data.json": "ca7388d97a0ed696d3b173eef5f8264bc76451b864baf2678df3144943bfe53e",
|
| 11 |
+
"episodes.csv": "ef6dc2beac16fbf02f603a770de4f3045eede1696d04c25cb0de789745fd4c69",
|
| 12 |
+
"episodes.json": "4dfa2559eee240954b0c2bf1dc8e15085f03d34dde749a21497c805cb7f57129",
|
| 13 |
"failure-reviews.json": "c66a997c9ecd6f7f43f130a8a33cd64c7a5b2049033e9ccde653ee30e4796f23",
|
| 14 |
"favicon.svg": "46063791218655f70c828ef71a1232a34fd38181da938338badec929ad226819",
|
| 15 |
+
"index.html": "80694c5b55656fd3392b354a0d122586bb86a2a1466ca23d2031092299d72202",
|
| 16 |
+
"licenses/LIBERO-LICENSE.txt": "e2885fd30a08381b799c4a33385522b23d637b4051b8f9a7f9f2519944b68ff6",
|
| 17 |
"licenses/RoboDojo-LICENSE.txt": "7794bb06af8fe5485ca912454ad1665ccd7846c45f9c2b1258658e480948cdbe",
|
| 18 |
+
"licenses/robocasa-LICENSE.txt": "5da18670b3f00c59847b1ded9c28dee59940d963b1e03b528b0108d9c5a09885",
|
| 19 |
+
"licenses/robosuite-LICENSE.txt": "177978cbece0a4c454c2aaec5b3f145b39270814874c43109da9e829c39d9cba",
|
| 20 |
+
"protocol.json": "8368960a967825a8d721cf982e61d028c5d72091530b6d54316e7626955e6dbd",
|
| 21 |
"style.css": "dd31b30951c220ebb82bc96224b5e640913945679e42b2e0a42e2c373a329e6c",
|
| 22 |
+
"summary.json": "752f4bccc2f906bc94f05113a9fd906d8bc1ad512d3a4057ae885479b1b94611",
|
| 23 |
+
"task-index.json": "c5bee5272347efefba25529144c9e752bf3ca2d336942c6b57a34036d1774517"
|
| 24 |
}
|
protocol.json
CHANGED
|
@@ -4,14 +4,52 @@
|
|
| 4 |
"effort": "high",
|
| 5 |
"seed": 0,
|
| 6 |
"episodes_per_task": 1,
|
| 7 |
-
"max_steps": 7500,
|
| 8 |
-
"control_frequency_hz": 25,
|
| 9 |
"mode": "stepped",
|
| 10 |
"safety_timeout_sec": 28800,
|
| 11 |
"fresh_workspace": true,
|
| 12 |
"resume_trajectory": false,
|
| 13 |
"source_jobs": [],
|
| 14 |
"families": [
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 15 |
{
|
| 16 |
"id": "task04",
|
| 17 |
"name": "RoboDojo",
|
|
@@ -34,5 +72,26 @@
|
|
| 34 |
}
|
| 35 |
],
|
| 36 |
"usage_accounting": "reported-responses",
|
| 37 |
-
"retry_policy_note": "The first six completed results use original Codex request defaults. Remaining tasks and two authorized fresh replacements use a 50-retry request policy."
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 38 |
}
|
|
|
|
| 4 |
"effort": "high",
|
| 5 |
"seed": 0,
|
| 6 |
"episodes_per_task": 1,
|
|
|
|
|
|
|
| 7 |
"mode": "stepped",
|
| 8 |
"safety_timeout_sec": 28800,
|
| 9 |
"fresh_workspace": true,
|
| 10 |
"resume_trajectory": false,
|
| 11 |
"source_jobs": [],
|
| 12 |
"families": [
|
| 13 |
+
{
|
| 14 |
+
"id": "task01",
|
| 15 |
+
"name": "RoboCasa",
|
| 16 |
+
"total": 10,
|
| 17 |
+
"completed": 10,
|
| 18 |
+
"pending": 0,
|
| 19 |
+
"successes": 2,
|
| 20 |
+
"valid_results": 10,
|
| 21 |
+
"success_rate": 0.2,
|
| 22 |
+
"input_tokens": 63285354,
|
| 23 |
+
"cached_input_tokens": 62436736,
|
| 24 |
+
"output_tokens": 204733,
|
| 25 |
+
"usage_complete": 10,
|
| 26 |
+
"control_frequency_hz": 20,
|
| 27 |
+
"max_control_steps": 6000,
|
| 28 |
+
"preflight_results": 0,
|
| 29 |
+
"formal_results": 10,
|
| 30 |
+
"modified_results": 10,
|
| 31 |
+
"original_results": 0
|
| 32 |
+
},
|
| 33 |
+
{
|
| 34 |
+
"id": "task02",
|
| 35 |
+
"name": "LIBERO Long",
|
| 36 |
+
"total": 10,
|
| 37 |
+
"completed": 10,
|
| 38 |
+
"pending": 0,
|
| 39 |
+
"successes": 9,
|
| 40 |
+
"valid_results": 10,
|
| 41 |
+
"success_rate": 0.9,
|
| 42 |
+
"input_tokens": 12586499,
|
| 43 |
+
"cached_input_tokens": 12106752,
|
| 44 |
+
"output_tokens": 93969,
|
| 45 |
+
"usage_complete": 10,
|
| 46 |
+
"control_frequency_hz": 20,
|
| 47 |
+
"max_control_steps": 6000,
|
| 48 |
+
"preflight_results": 0,
|
| 49 |
+
"formal_results": 10,
|
| 50 |
+
"modified_results": 5,
|
| 51 |
+
"original_results": 5
|
| 52 |
+
},
|
| 53 |
{
|
| 54 |
"id": "task04",
|
| 55 |
"name": "RoboDojo",
|
|
|
|
| 72 |
}
|
| 73 |
],
|
| 74 |
"usage_accounting": "reported-responses",
|
| 75 |
+
"retry_policy_note": "The first six completed results use original Codex request defaults. Remaining tasks and two authorized fresh replacements use a 50-retry request policy. All new RoboCasa and LIBERO runs use explicit 50-retry SSE request policy; no automatic episode retry.",
|
| 76 |
+
"family_protocols": {
|
| 77 |
+
"task01": {
|
| 78 |
+
"episodes": 1,
|
| 79 |
+
"seed": 0,
|
| 80 |
+
"max_steps": 6000,
|
| 81 |
+
"control_frequency_hz": 20
|
| 82 |
+
},
|
| 83 |
+
"task02": {
|
| 84 |
+
"episodes": 1,
|
| 85 |
+
"seed": 0,
|
| 86 |
+
"init_state_index": 0,
|
| 87 |
+
"max_steps": 6000,
|
| 88 |
+
"control_frequency_hz": 20
|
| 89 |
+
},
|
| 90 |
+
"task04": {
|
| 91 |
+
"episodes": 1,
|
| 92 |
+
"seed": 0,
|
| 93 |
+
"max_steps": 7500,
|
| 94 |
+
"control_frequency_hz": 25
|
| 95 |
+
}
|
| 96 |
+
}
|
| 97 |
}
|
publication-manifest.json
CHANGED
|
@@ -6,44 +6,44 @@
|
|
| 6 |
"bytes": 215
|
| 7 |
},
|
| 8 |
"DATA_FORMAT.md": {
|
| 9 |
-
"sha256": "
|
| 10 |
"bytes": 1129
|
| 11 |
},
|
| 12 |
"README.md": {
|
| 13 |
-
"sha256": "
|
| 14 |
-
"bytes":
|
| 15 |
},
|
| 16 |
"REPORT.md": {
|
| 17 |
-
"sha256": "
|
| 18 |
-
"bytes":
|
| 19 |
},
|
| 20 |
"THIRD_PARTY_NOTICES.md": {
|
| 21 |
-
"sha256": "
|
| 22 |
-
"bytes":
|
| 23 |
},
|
| 24 |
"app.js": {
|
| 25 |
-
"sha256": "
|
| 26 |
-
"bytes":
|
| 27 |
},
|
| 28 |
"attempt-history.json": {
|
| 29 |
"sha256": "8555f458c1dd362f5a07447342994fddf265ff23bbc40e9bd6f9984a892ddb4e",
|
| 30 |
"bytes": 36661
|
| 31 |
},
|
| 32 |
"comparison.json": {
|
| 33 |
-
"sha256": "
|
| 34 |
-
"bytes":
|
| 35 |
},
|
| 36 |
"data.json": {
|
| 37 |
-
"sha256": "
|
| 38 |
-
"bytes":
|
| 39 |
},
|
| 40 |
"episodes.csv": {
|
| 41 |
-
"sha256": "
|
| 42 |
-
"bytes":
|
| 43 |
},
|
| 44 |
"episodes.json": {
|
| 45 |
-
"sha256": "
|
| 46 |
-
"bytes":
|
| 47 |
},
|
| 48 |
"failure-reviews.json": {
|
| 49 |
"sha256": "c66a997c9ecd6f7f43f130a8a33cd64c7a5b2049033e9ccde653ee30e4796f23",
|
|
@@ -54,32 +54,44 @@
|
|
| 54 |
"bytes": 227
|
| 55 |
},
|
| 56 |
"index.html": {
|
| 57 |
-
"sha256": "
|
| 58 |
-
"bytes":
|
|
|
|
|
|
|
|
|
|
|
|
|
| 59 |
},
|
| 60 |
"licenses/RoboDojo-LICENSE.txt": {
|
| 61 |
"sha256": "7794bb06af8fe5485ca912454ad1665ccd7846c45f9c2b1258658e480948cdbe",
|
| 62 |
"bytes": 1091
|
| 63 |
},
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 64 |
"protocol.json": {
|
| 65 |
-
"sha256": "
|
| 66 |
-
"bytes":
|
| 67 |
},
|
| 68 |
"style.css": {
|
| 69 |
"sha256": "dd31b30951c220ebb82bc96224b5e640913945679e42b2e0a42e2c373a329e6c",
|
| 70 |
"bytes": 16560
|
| 71 |
},
|
| 72 |
"summary.json": {
|
| 73 |
-
"sha256": "
|
| 74 |
-
"bytes":
|
| 75 |
},
|
| 76 |
"task-index.json": {
|
| 77 |
-
"sha256": "
|
| 78 |
-
"bytes":
|
| 79 |
},
|
| 80 |
"manifest.json": {
|
| 81 |
-
"sha256": "
|
| 82 |
-
"bytes":
|
| 83 |
}
|
| 84 |
}
|
| 85 |
}
|
|
|
|
| 6 |
"bytes": 215
|
| 7 |
},
|
| 8 |
"DATA_FORMAT.md": {
|
| 9 |
+
"sha256": "e2fd4f6b5c8be9d9e008db7d0c1f5d52b419d68328a1c86625c9f67c49b03403",
|
| 10 |
"bytes": 1129
|
| 11 |
},
|
| 12 |
"README.md": {
|
| 13 |
+
"sha256": "43d89b6c59fdee0a1f5bddf3bd21cf0d809fee832f115771da95389f40f421b8",
|
| 14 |
+
"bytes": 1688
|
| 15 |
},
|
| 16 |
"REPORT.md": {
|
| 17 |
+
"sha256": "696454f04ac34467d2313e28b1e58b7095b02dc09e2b33b4b597594e10e3b59d",
|
| 18 |
+
"bytes": 1565
|
| 19 |
},
|
| 20 |
"THIRD_PARTY_NOTICES.md": {
|
| 21 |
+
"sha256": "8f38aa1ab2e03203b438e095a6325b3914541efe434290a9a2f786087d8f657d",
|
| 22 |
+
"bytes": 2319
|
| 23 |
},
|
| 24 |
"app.js": {
|
| 25 |
+
"sha256": "eb013cc9d24382410953a1a8c81e446a6592a81b32887749d17bfc863936556b",
|
| 26 |
+
"bytes": 24218
|
| 27 |
},
|
| 28 |
"attempt-history.json": {
|
| 29 |
"sha256": "8555f458c1dd362f5a07447342994fddf265ff23bbc40e9bd6f9984a892ddb4e",
|
| 30 |
"bytes": 36661
|
| 31 |
},
|
| 32 |
"comparison.json": {
|
| 33 |
+
"sha256": "012c7a64ecfde3352d258e332eae3b52de3d0c253663841e1a7dc9af3b3dfab7",
|
| 34 |
+
"bytes": 84718
|
| 35 |
},
|
| 36 |
"data.json": {
|
| 37 |
+
"sha256": "ca7388d97a0ed696d3b173eef5f8264bc76451b864baf2678df3144943bfe53e",
|
| 38 |
+
"bytes": 907017
|
| 39 |
},
|
| 40 |
"episodes.csv": {
|
| 41 |
+
"sha256": "ef6dc2beac16fbf02f603a770de4f3045eede1696d04c25cb0de789745fd4c69",
|
| 42 |
+
"bytes": 5479
|
| 43 |
},
|
| 44 |
"episodes.json": {
|
| 45 |
+
"sha256": "4dfa2559eee240954b0c2bf1dc8e15085f03d34dde749a21497c805cb7f57129",
|
| 46 |
+
"bytes": 667599
|
| 47 |
},
|
| 48 |
"failure-reviews.json": {
|
| 49 |
"sha256": "c66a997c9ecd6f7f43f130a8a33cd64c7a5b2049033e9ccde653ee30e4796f23",
|
|
|
|
| 54 |
"bytes": 227
|
| 55 |
},
|
| 56 |
"index.html": {
|
| 57 |
+
"sha256": "80694c5b55656fd3392b354a0d122586bb86a2a1466ca23d2031092299d72202",
|
| 58 |
+
"bytes": 4340
|
| 59 |
+
},
|
| 60 |
+
"licenses/LIBERO-LICENSE.txt": {
|
| 61 |
+
"sha256": "e2885fd30a08381b799c4a33385522b23d637b4051b8f9a7f9f2519944b68ff6",
|
| 62 |
+
"bytes": 1080
|
| 63 |
},
|
| 64 |
"licenses/RoboDojo-LICENSE.txt": {
|
| 65 |
"sha256": "7794bb06af8fe5485ca912454ad1665ccd7846c45f9c2b1258658e480948cdbe",
|
| 66 |
"bytes": 1091
|
| 67 |
},
|
| 68 |
+
"licenses/robocasa-LICENSE.txt": {
|
| 69 |
+
"sha256": "5da18670b3f00c59847b1ded9c28dee59940d963b1e03b528b0108d9c5a09885",
|
| 70 |
+
"bytes": 1418
|
| 71 |
+
},
|
| 72 |
+
"licenses/robosuite-LICENSE.txt": {
|
| 73 |
+
"sha256": "177978cbece0a4c454c2aaec5b3f145b39270814874c43109da9e829c39d9cba",
|
| 74 |
+
"bytes": 1474
|
| 75 |
+
},
|
| 76 |
"protocol.json": {
|
| 77 |
+
"sha256": "8368960a967825a8d721cf982e61d028c5d72091530b6d54316e7626955e6dbd",
|
| 78 |
+
"bytes": 2574
|
| 79 |
},
|
| 80 |
"style.css": {
|
| 81 |
"sha256": "dd31b30951c220ebb82bc96224b5e640913945679e42b2e0a42e2c373a329e6c",
|
| 82 |
"bytes": 16560
|
| 83 |
},
|
| 84 |
"summary.json": {
|
| 85 |
+
"sha256": "752f4bccc2f906bc94f05113a9fd906d8bc1ad512d3a4057ae885479b1b94611",
|
| 86 |
+
"bytes": 1922
|
| 87 |
},
|
| 88 |
"task-index.json": {
|
| 89 |
+
"sha256": "c5bee5272347efefba25529144c9e752bf3ca2d336942c6b57a34036d1774517",
|
| 90 |
+
"bytes": 192726
|
| 91 |
},
|
| 92 |
"manifest.json": {
|
| 93 |
+
"sha256": "65df6bd8f0503fc521d5101cd2d41f2c4d1bd408d3f8ea672fd8363211d68e14",
|
| 94 |
+
"bytes": 1979
|
| 95 |
}
|
| 96 |
}
|
| 97 |
}
|
summary.json
CHANGED
|
@@ -1,8 +1,48 @@
|
|
| 1 |
{
|
| 2 |
-
"planned_tasks":
|
| 3 |
-
"published_results":
|
| 4 |
"pending_tasks": 0,
|
| 5 |
"families": [
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
{
|
| 7 |
"id": "task04",
|
| 8 |
"name": "RoboDojo",
|
|
@@ -25,13 +65,13 @@
|
|
| 25 |
}
|
| 26 |
],
|
| 27 |
"progress": {
|
| 28 |
-
"finished":
|
| 29 |
-
"
|
| 30 |
-
"native_failures": 11,
|
| 31 |
-
"native_successes": 31,
|
| 32 |
-
"needs_review": 0,
|
| 33 |
"queued": 0,
|
| 34 |
-
"
|
|
|
|
|
|
|
|
|
|
| 35 |
},
|
| 36 |
"interrupted_attempts": 4,
|
| 37 |
"usage_incomplete_tasks": [],
|
|
|
|
| 1 |
{
|
| 2 |
+
"planned_tasks": 62,
|
| 3 |
+
"published_results": 62,
|
| 4 |
"pending_tasks": 0,
|
| 5 |
"families": [
|
| 6 |
+
{
|
| 7 |
+
"id": "task01",
|
| 8 |
+
"name": "RoboCasa",
|
| 9 |
+
"total": 10,
|
| 10 |
+
"completed": 10,
|
| 11 |
+
"pending": 0,
|
| 12 |
+
"successes": 2,
|
| 13 |
+
"valid_results": 10,
|
| 14 |
+
"success_rate": 0.2,
|
| 15 |
+
"input_tokens": 63285354,
|
| 16 |
+
"cached_input_tokens": 62436736,
|
| 17 |
+
"output_tokens": 204733,
|
| 18 |
+
"usage_complete": 10,
|
| 19 |
+
"control_frequency_hz": 20,
|
| 20 |
+
"max_control_steps": 6000,
|
| 21 |
+
"preflight_results": 0,
|
| 22 |
+
"formal_results": 10,
|
| 23 |
+
"modified_results": 10,
|
| 24 |
+
"original_results": 0
|
| 25 |
+
},
|
| 26 |
+
{
|
| 27 |
+
"id": "task02",
|
| 28 |
+
"name": "LIBERO Long",
|
| 29 |
+
"total": 10,
|
| 30 |
+
"completed": 10,
|
| 31 |
+
"pending": 0,
|
| 32 |
+
"successes": 9,
|
| 33 |
+
"valid_results": 10,
|
| 34 |
+
"success_rate": 0.9,
|
| 35 |
+
"input_tokens": 12586499,
|
| 36 |
+
"cached_input_tokens": 12106752,
|
| 37 |
+
"output_tokens": 93969,
|
| 38 |
+
"usage_complete": 10,
|
| 39 |
+
"control_frequency_hz": 20,
|
| 40 |
+
"max_control_steps": 6000,
|
| 41 |
+
"preflight_results": 0,
|
| 42 |
+
"formal_results": 10,
|
| 43 |
+
"modified_results": 5,
|
| 44 |
+
"original_results": 5
|
| 45 |
+
},
|
| 46 |
{
|
| 47 |
"id": "task04",
|
| 48 |
"name": "RoboDojo",
|
|
|
|
| 65 |
}
|
| 66 |
],
|
| 67 |
"progress": {
|
| 68 |
+
"finished": 62,
|
| 69 |
+
"running": 0,
|
|
|
|
|
|
|
|
|
|
| 70 |
"queued": 0,
|
| 71 |
+
"needs_review": 0,
|
| 72 |
+
"interrupted": 0,
|
| 73 |
+
"native_successes": 42,
|
| 74 |
+
"native_failures": 20
|
| 75 |
},
|
| 76 |
"interrupted_attempts": 4,
|
| 77 |
"usage_incomplete_tasks": [],
|
task-index.json
CHANGED
|
@@ -1,4 +1,1467 @@
|
|
| 1 |
[
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 2 |
{
|
| 3 |
"key": "task04/01",
|
| 4 |
"family": "task04",
|
|
|
|
| 1 |
[
|
| 2 |
+
{
|
| 3 |
+
"key": "task01/01",
|
| 4 |
+
"family": "task01",
|
| 5 |
+
"slot": "01",
|
| 6 |
+
"native_id": "robocasa/load-condiments-in-fridge",
|
| 7 |
+
"title": "Load condiments in fridge",
|
| 8 |
+
"catalog_instruction": "Place the specified condiments from the counter on the top shelf of the fridge. Move any existing top-shelf items to other shelves. Release the items and move the gripper clear of the stored items.",
|
| 9 |
+
"native_instruction": "Place the {condiment1} and {condiment2} from the counter to the top shelf of the fridge. If the existing items in the fridge are on the top shelf, move them to other shelves.",
|
| 10 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 11 |
+
"status": "completed",
|
| 12 |
+
"episode_id": "task01-01-seed0-formal",
|
| 13 |
+
"planned_protocol": {
|
| 14 |
+
"episodes": 1,
|
| 15 |
+
"seed": 0,
|
| 16 |
+
"control_frequency_hz": 20,
|
| 17 |
+
"max_control_steps": 6000,
|
| 18 |
+
"timeout_s": 28800,
|
| 19 |
+
"mode": "stepped",
|
| 20 |
+
"instruction_policy": "modified"
|
| 21 |
+
},
|
| 22 |
+
"catalog_id": "robocasa/load-condiments-in-fridge",
|
| 23 |
+
"native_identity": {
|
| 24 |
+
"task_name": "LoadCondimentsInFridge",
|
| 25 |
+
"max_steps": 6000
|
| 26 |
+
},
|
| 27 |
+
"instruction_revision": {
|
| 28 |
+
"native_instruction": "Place the {condiment1} and {condiment2} from the counter to the top shelf of the fridge. If the existing items in the fridge are on the top shelf, move them to other shelves.",
|
| 29 |
+
"instruction": "Place the specified condiments from the counter on the top shelf of the fridge. Move any existing top-shelf items to other shelves. Release the items and move the gripper clear of the stored items.",
|
| 30 |
+
"native_instruction_kind": "source",
|
| 31 |
+
"diff": [
|
| 32 |
+
{
|
| 33 |
+
"op": "equal",
|
| 34 |
+
"original": "Place the ",
|
| 35 |
+
"modified": "Place the "
|
| 36 |
+
},
|
| 37 |
+
{
|
| 38 |
+
"op": "replace",
|
| 39 |
+
"original": "{condiment1}",
|
| 40 |
+
"modified": "specified"
|
| 41 |
+
},
|
| 42 |
+
{
|
| 43 |
+
"op": "equal",
|
| 44 |
+
"original": " ",
|
| 45 |
+
"modified": " "
|
| 46 |
+
},
|
| 47 |
+
{
|
| 48 |
+
"op": "replace",
|
| 49 |
+
"original": "and {condiment2}",
|
| 50 |
+
"modified": "condiments"
|
| 51 |
+
},
|
| 52 |
+
{
|
| 53 |
+
"op": "equal",
|
| 54 |
+
"original": " from the counter ",
|
| 55 |
+
"modified": " from the counter "
|
| 56 |
+
},
|
| 57 |
+
{
|
| 58 |
+
"op": "replace",
|
| 59 |
+
"original": "to",
|
| 60 |
+
"modified": "on"
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"op": "equal",
|
| 64 |
+
"original": " the top shelf of the fridge. ",
|
| 65 |
+
"modified": " the top shelf of the fridge. "
|
| 66 |
+
},
|
| 67 |
+
{
|
| 68 |
+
"op": "replace",
|
| 69 |
+
"original": "If",
|
| 70 |
+
"modified": "Move"
|
| 71 |
+
},
|
| 72 |
+
{
|
| 73 |
+
"op": "equal",
|
| 74 |
+
"original": " ",
|
| 75 |
+
"modified": " "
|
| 76 |
+
},
|
| 77 |
+
{
|
| 78 |
+
"op": "replace",
|
| 79 |
+
"original": "the",
|
| 80 |
+
"modified": "any"
|
| 81 |
+
},
|
| 82 |
+
{
|
| 83 |
+
"op": "equal",
|
| 84 |
+
"original": " existing ",
|
| 85 |
+
"modified": " existing "
|
| 86 |
+
},
|
| 87 |
+
{
|
| 88 |
+
"op": "insert",
|
| 89 |
+
"original": "",
|
| 90 |
+
"modified": "top-shelf "
|
| 91 |
+
},
|
| 92 |
+
{
|
| 93 |
+
"op": "equal",
|
| 94 |
+
"original": "items",
|
| 95 |
+
"modified": "items"
|
| 96 |
+
},
|
| 97 |
+
{
|
| 98 |
+
"op": "delete",
|
| 99 |
+
"original": " in the fridge are on the top shelf, move them",
|
| 100 |
+
"modified": ""
|
| 101 |
+
},
|
| 102 |
+
{
|
| 103 |
+
"op": "equal",
|
| 104 |
+
"original": " to other shelves.",
|
| 105 |
+
"modified": " to other shelves."
|
| 106 |
+
},
|
| 107 |
+
{
|
| 108 |
+
"op": "insert",
|
| 109 |
+
"original": "",
|
| 110 |
+
"modified": " Release the items and move the gripper clear of the stored items."
|
| 111 |
+
}
|
| 112 |
+
],
|
| 113 |
+
"changed": true,
|
| 114 |
+
"preserve_whitespace": true,
|
| 115 |
+
"label": "Modified instruction",
|
| 116 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/01-load-condiments-in-fridge/task.yaml"
|
| 117 |
+
},
|
| 118 |
+
"run_status": "finished",
|
| 119 |
+
"attempt_history": [],
|
| 120 |
+
"status_note": "Evaluation and native evidence complete."
|
| 121 |
+
},
|
| 122 |
+
{
|
| 123 |
+
"key": "task01/02",
|
| 124 |
+
"family": "task01",
|
| 125 |
+
"slot": "02",
|
| 126 |
+
"native_id": "robocasa/filter-microwavable-item",
|
| 127 |
+
"title": "Filter microwavable item",
|
| 128 |
+
"catalog_instruction": "Remove the specified fruit from the bowl and place it on the small plate. Then place the bowl with only the specified meat in the microwave, close the door, and press the start button to microwave the meat. Release the bowl and move the gripper clear of it.",
|
| 129 |
+
"native_instruction": "Remove the {fruit} from the bowl and place it on the small plate. Then place the bowl with only the {meat} in the microwave, close the door, and press the start button to microwave the {meat}.",
|
| 130 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 131 |
+
"status": "completed",
|
| 132 |
+
"episode_id": "task01-02-seed0-formal",
|
| 133 |
+
"planned_protocol": {
|
| 134 |
+
"episodes": 1,
|
| 135 |
+
"seed": 0,
|
| 136 |
+
"control_frequency_hz": 20,
|
| 137 |
+
"max_control_steps": 6000,
|
| 138 |
+
"timeout_s": 28800,
|
| 139 |
+
"mode": "stepped",
|
| 140 |
+
"instruction_policy": "modified"
|
| 141 |
+
},
|
| 142 |
+
"catalog_id": "robocasa/filter-microwavable-item",
|
| 143 |
+
"native_identity": {
|
| 144 |
+
"task_name": "FilterMicrowavableItem",
|
| 145 |
+
"max_steps": 6000
|
| 146 |
+
},
|
| 147 |
+
"instruction_revision": {
|
| 148 |
+
"native_instruction": "Remove the {fruit} from the bowl and place it on the small plate. Then place the bowl with only the {meat} in the microwave, close the door, and press the start button to microwave the {meat}.",
|
| 149 |
+
"instruction": "Remove the specified fruit from the bowl and place it on the small plate. Then place the bowl with only the specified meat in the microwave, close the door, and press the start button to microwave the meat. Release the bowl and move the gripper clear of it.",
|
| 150 |
+
"native_instruction_kind": "source",
|
| 151 |
+
"diff": [
|
| 152 |
+
{
|
| 153 |
+
"op": "equal",
|
| 154 |
+
"original": "Remove the ",
|
| 155 |
+
"modified": "Remove the "
|
| 156 |
+
},
|
| 157 |
+
{
|
| 158 |
+
"op": "replace",
|
| 159 |
+
"original": "{fruit}",
|
| 160 |
+
"modified": "specified fruit"
|
| 161 |
+
},
|
| 162 |
+
{
|
| 163 |
+
"op": "equal",
|
| 164 |
+
"original": " from the bowl and place it on the small plate. Then place the bowl with only the ",
|
| 165 |
+
"modified": " from the bowl and place it on the small plate. Then place the bowl with only the "
|
| 166 |
+
},
|
| 167 |
+
{
|
| 168 |
+
"op": "replace",
|
| 169 |
+
"original": "{meat}",
|
| 170 |
+
"modified": "specified meat"
|
| 171 |
+
},
|
| 172 |
+
{
|
| 173 |
+
"op": "equal",
|
| 174 |
+
"original": " in the microwave, close the door, and press the start button to microwave the ",
|
| 175 |
+
"modified": " in the microwave, close the door, and press the start button to microwave the "
|
| 176 |
+
},
|
| 177 |
+
{
|
| 178 |
+
"op": "replace",
|
| 179 |
+
"original": "{meat}.",
|
| 180 |
+
"modified": "meat. Release the bowl and move the gripper clear of it."
|
| 181 |
+
}
|
| 182 |
+
],
|
| 183 |
+
"changed": true,
|
| 184 |
+
"preserve_whitespace": true,
|
| 185 |
+
"label": "Modified instruction",
|
| 186 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/02-filter-microwavable-item/task.yaml"
|
| 187 |
+
},
|
| 188 |
+
"run_status": "finished",
|
| 189 |
+
"attempt_history": [],
|
| 190 |
+
"status_note": "Evaluation and native evidence complete."
|
| 191 |
+
},
|
| 192 |
+
{
|
| 193 |
+
"key": "task01/03",
|
| 194 |
+
"family": "task01",
|
| 195 |
+
"slot": "03",
|
| 196 |
+
"native_id": "robocasa/store-dumplings",
|
| 197 |
+
"title": "Store dumplings",
|
| 198 |
+
"catalog_instruction": "Place two dumplings into each of the tupperware containers and then place both containers on a shelf in the fridge. Release the dumplings and containers and move the gripper clear of them.",
|
| 199 |
+
"native_instruction": "Place two dumplings into each of the tupperware containers and then place the containers in the fridge.",
|
| 200 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 201 |
+
"status": "completed",
|
| 202 |
+
"episode_id": "task01-03-seed0-formal",
|
| 203 |
+
"planned_protocol": {
|
| 204 |
+
"episodes": 1,
|
| 205 |
+
"seed": 0,
|
| 206 |
+
"control_frequency_hz": 20,
|
| 207 |
+
"max_control_steps": 6000,
|
| 208 |
+
"timeout_s": 28800,
|
| 209 |
+
"mode": "stepped",
|
| 210 |
+
"instruction_policy": "modified"
|
| 211 |
+
},
|
| 212 |
+
"catalog_id": "robocasa/store-dumplings",
|
| 213 |
+
"native_identity": {
|
| 214 |
+
"task_name": "StoreDumplings",
|
| 215 |
+
"max_steps": 6000
|
| 216 |
+
},
|
| 217 |
+
"instruction_revision": {
|
| 218 |
+
"native_instruction": "Place two dumplings into each of the tupperware containers and then place the containers in the fridge.",
|
| 219 |
+
"instruction": "Place two dumplings into each of the tupperware containers and then place both containers on a shelf in the fridge. Release the dumplings and containers and move the gripper clear of them.",
|
| 220 |
+
"native_instruction_kind": "source",
|
| 221 |
+
"diff": [
|
| 222 |
+
{
|
| 223 |
+
"op": "equal",
|
| 224 |
+
"original": "Place two dumplings into each of the tupperware containers and then place ",
|
| 225 |
+
"modified": "Place two dumplings into each of the tupperware containers and then place "
|
| 226 |
+
},
|
| 227 |
+
{
|
| 228 |
+
"op": "replace",
|
| 229 |
+
"original": "the",
|
| 230 |
+
"modified": "both"
|
| 231 |
+
},
|
| 232 |
+
{
|
| 233 |
+
"op": "equal",
|
| 234 |
+
"original": " containers",
|
| 235 |
+
"modified": " containers"
|
| 236 |
+
},
|
| 237 |
+
{
|
| 238 |
+
"op": "insert",
|
| 239 |
+
"original": "",
|
| 240 |
+
"modified": " on a shelf"
|
| 241 |
+
},
|
| 242 |
+
{
|
| 243 |
+
"op": "equal",
|
| 244 |
+
"original": " in the fridge.",
|
| 245 |
+
"modified": " in the fridge."
|
| 246 |
+
},
|
| 247 |
+
{
|
| 248 |
+
"op": "insert",
|
| 249 |
+
"original": "",
|
| 250 |
+
"modified": " Release the dumplings and containers and move the gripper clear of them."
|
| 251 |
+
}
|
| 252 |
+
],
|
| 253 |
+
"changed": true,
|
| 254 |
+
"preserve_whitespace": true,
|
| 255 |
+
"label": "Modified instruction",
|
| 256 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/03-store-dumplings/task.yaml"
|
| 257 |
+
},
|
| 258 |
+
"run_status": "finished",
|
| 259 |
+
"attempt_history": [],
|
| 260 |
+
"status_note": "Evaluation and native evidence complete."
|
| 261 |
+
},
|
| 262 |
+
{
|
| 263 |
+
"key": "task01/04",
|
| 264 |
+
"family": "task01",
|
| 265 |
+
"slot": "04",
|
| 266 |
+
"native_id": "robocasa/divide-buffet-trays",
|
| 267 |
+
"title": "Divide buffet trays",
|
| 268 |
+
"catalog_instruction": "Gather the specified vegetables from the fridge and place them on a tray on the dining counter. Then gather the specified meats from the fridge and place them on the other tray. Release the food and move the gripper clear of the food and both trays.",
|
| 269 |
+
"native_instruction": "Gather the {vegetables} from the fridge and place them on a tray on the dining counter. Then gather the {meats} from the fridge and place them on the other tray.",
|
| 270 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 271 |
+
"status": "completed",
|
| 272 |
+
"episode_id": "task01-04-seed0-formal",
|
| 273 |
+
"planned_protocol": {
|
| 274 |
+
"episodes": 1,
|
| 275 |
+
"seed": 0,
|
| 276 |
+
"control_frequency_hz": 20,
|
| 277 |
+
"max_control_steps": 6000,
|
| 278 |
+
"timeout_s": 28800,
|
| 279 |
+
"mode": "stepped",
|
| 280 |
+
"instruction_policy": "modified"
|
| 281 |
+
},
|
| 282 |
+
"catalog_id": "robocasa/divide-buffet-trays",
|
| 283 |
+
"native_identity": {
|
| 284 |
+
"task_name": "DivideBuffetTrays",
|
| 285 |
+
"max_steps": 6000
|
| 286 |
+
},
|
| 287 |
+
"instruction_revision": {
|
| 288 |
+
"native_instruction": "Gather the {vegetables} from the fridge and place them on a tray on the dining counter. Then gather the {meats} from the fridge and place them on the other tray.",
|
| 289 |
+
"instruction": "Gather the specified vegetables from the fridge and place them on a tray on the dining counter. Then gather the specified meats from the fridge and place them on the other tray. Release the food and move the gripper clear of the food and both trays.",
|
| 290 |
+
"native_instruction_kind": "source",
|
| 291 |
+
"diff": [
|
| 292 |
+
{
|
| 293 |
+
"op": "equal",
|
| 294 |
+
"original": "Gather the ",
|
| 295 |
+
"modified": "Gather the "
|
| 296 |
+
},
|
| 297 |
+
{
|
| 298 |
+
"op": "replace",
|
| 299 |
+
"original": "{vegetables}",
|
| 300 |
+
"modified": "specified vegetables"
|
| 301 |
+
},
|
| 302 |
+
{
|
| 303 |
+
"op": "equal",
|
| 304 |
+
"original": " from the fridge and place them on a tray on the dining counter. Then gather the ",
|
| 305 |
+
"modified": " from the fridge and place them on a tray on the dining counter. Then gather the "
|
| 306 |
+
},
|
| 307 |
+
{
|
| 308 |
+
"op": "replace",
|
| 309 |
+
"original": "{meats}",
|
| 310 |
+
"modified": "specified meats"
|
| 311 |
+
},
|
| 312 |
+
{
|
| 313 |
+
"op": "equal",
|
| 314 |
+
"original": " from the fridge and place them on the other tray.",
|
| 315 |
+
"modified": " from the fridge and place them on the other tray."
|
| 316 |
+
},
|
| 317 |
+
{
|
| 318 |
+
"op": "insert",
|
| 319 |
+
"original": "",
|
| 320 |
+
"modified": " Release the food and move the gripper clear of the food and both trays."
|
| 321 |
+
}
|
| 322 |
+
],
|
| 323 |
+
"changed": true,
|
| 324 |
+
"preserve_whitespace": true,
|
| 325 |
+
"label": "Modified instruction",
|
| 326 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/04-divide-buffet-trays/task.yaml"
|
| 327 |
+
},
|
| 328 |
+
"run_status": "finished",
|
| 329 |
+
"attempt_history": [],
|
| 330 |
+
"status_note": "Evaluation and native evidence complete."
|
| 331 |
+
},
|
| 332 |
+
{
|
| 333 |
+
"key": "task01/05",
|
| 334 |
+
"family": "task01",
|
| 335 |
+
"slot": "05",
|
| 336 |
+
"native_id": "robocasa/make-cheesecake-filling",
|
| 337 |
+
"title": "Make cheesecake filling",
|
| 338 |
+
"catalog_instruction": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer bowl, lower the mixer head fully, and turn the speed knob to begin making cheesecake filling.",
|
| 339 |
+
"native_instruction": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer bowl and then turn the speed knob to begin making cheesecake filling.",
|
| 340 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 341 |
+
"status": "completed",
|
| 342 |
+
"episode_id": "task01-05-seed0-formal",
|
| 343 |
+
"planned_protocol": {
|
| 344 |
+
"episodes": 1,
|
| 345 |
+
"seed": 0,
|
| 346 |
+
"control_frequency_hz": 20,
|
| 347 |
+
"max_control_steps": 6000,
|
| 348 |
+
"timeout_s": 28800,
|
| 349 |
+
"mode": "stepped",
|
| 350 |
+
"instruction_policy": "modified"
|
| 351 |
+
},
|
| 352 |
+
"catalog_id": "robocasa/make-cheesecake-filling",
|
| 353 |
+
"native_identity": {
|
| 354 |
+
"task_name": "MakeCheesecakeFilling",
|
| 355 |
+
"max_steps": 6000
|
| 356 |
+
},
|
| 357 |
+
"instruction_revision": {
|
| 358 |
+
"native_instruction": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer bowl and then turn the speed knob to begin making cheesecake filling.",
|
| 359 |
+
"instruction": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer bowl, lower the mixer head fully, and turn the speed knob to begin making cheesecake filling.",
|
| 360 |
+
"native_instruction_kind": "source",
|
| 361 |
+
"diff": [
|
| 362 |
+
{
|
| 363 |
+
"op": "equal",
|
| 364 |
+
"original": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer ",
|
| 365 |
+
"modified": "Add the butter stick, sugar cube, and cream cheese stick to the stand mixer "
|
| 366 |
+
},
|
| 367 |
+
{
|
| 368 |
+
"op": "replace",
|
| 369 |
+
"original": "bowl",
|
| 370 |
+
"modified": "bowl, lower the mixer head fully,"
|
| 371 |
+
},
|
| 372 |
+
{
|
| 373 |
+
"op": "equal",
|
| 374 |
+
"original": " and",
|
| 375 |
+
"modified": " and"
|
| 376 |
+
},
|
| 377 |
+
{
|
| 378 |
+
"op": "delete",
|
| 379 |
+
"original": " then",
|
| 380 |
+
"modified": ""
|
| 381 |
+
},
|
| 382 |
+
{
|
| 383 |
+
"op": "equal",
|
| 384 |
+
"original": " turn the speed knob to begin making cheesecake filling.",
|
| 385 |
+
"modified": " turn the speed knob to begin making cheesecake filling."
|
| 386 |
+
}
|
| 387 |
+
],
|
| 388 |
+
"changed": true,
|
| 389 |
+
"preserve_whitespace": true,
|
| 390 |
+
"label": "Modified instruction",
|
| 391 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/05-make-cheesecake-filling/task.yaml"
|
| 392 |
+
},
|
| 393 |
+
"run_status": "finished",
|
| 394 |
+
"attempt_history": [],
|
| 395 |
+
"status_note": "Evaluation and native evidence complete."
|
| 396 |
+
},
|
| 397 |
+
{
|
| 398 |
+
"key": "task01/06",
|
| 399 |
+
"family": "task01",
|
| 400 |
+
"slot": "06",
|
| 401 |
+
"native_id": "robocasa/multistep-steaming",
|
| 402 |
+
"title": "Multistep steaming",
|
| 403 |
+
"catalog_instruction": "Turn on the sink faucet. Move the specified vegetable from the counter into the sink while the water is running. Turn off the faucet, move the vegetable into the pot next to the stove, and move the pot to the specified burner.",
|
| 404 |
+
"native_instruction": "Turn on the sink faucet. Then move the {vegetable} from the counter to the sink. Turn off the sink. Move the vegetable from the sink to the pot next to the stove. Finally move the pot to the {burner} burner.",
|
| 405 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 406 |
+
"status": "completed",
|
| 407 |
+
"episode_id": "task01-06-seed0-formal",
|
| 408 |
+
"planned_protocol": {
|
| 409 |
+
"episodes": 1,
|
| 410 |
+
"seed": 0,
|
| 411 |
+
"control_frequency_hz": 20,
|
| 412 |
+
"max_control_steps": 6000,
|
| 413 |
+
"timeout_s": 28800,
|
| 414 |
+
"mode": "stepped",
|
| 415 |
+
"instruction_policy": "modified"
|
| 416 |
+
},
|
| 417 |
+
"catalog_id": "robocasa/multistep-steaming",
|
| 418 |
+
"native_identity": {
|
| 419 |
+
"task_name": "MultistepSteaming",
|
| 420 |
+
"max_steps": 6000
|
| 421 |
+
},
|
| 422 |
+
"instruction_revision": {
|
| 423 |
+
"native_instruction": "Turn on the sink faucet. Then move the {vegetable} from the counter to the sink. Turn off the sink. Move the vegetable from the sink to the pot next to the stove. Finally move the pot to the {burner} burner.",
|
| 424 |
+
"instruction": "Turn on the sink faucet. Move the specified vegetable from the counter into the sink while the water is running. Turn off the faucet, move the vegetable into the pot next to the stove, and move the pot to the specified burner.",
|
| 425 |
+
"native_instruction_kind": "source",
|
| 426 |
+
"diff": [
|
| 427 |
+
{
|
| 428 |
+
"op": "equal",
|
| 429 |
+
"original": "Turn on the sink faucet. ",
|
| 430 |
+
"modified": "Turn on the sink faucet. "
|
| 431 |
+
},
|
| 432 |
+
{
|
| 433 |
+
"op": "replace",
|
| 434 |
+
"original": "Then move",
|
| 435 |
+
"modified": "Move"
|
| 436 |
+
},
|
| 437 |
+
{
|
| 438 |
+
"op": "equal",
|
| 439 |
+
"original": " the ",
|
| 440 |
+
"modified": " the "
|
| 441 |
+
},
|
| 442 |
+
{
|
| 443 |
+
"op": "replace",
|
| 444 |
+
"original": "{vegetable}",
|
| 445 |
+
"modified": "specified vegetable"
|
| 446 |
+
},
|
| 447 |
+
{
|
| 448 |
+
"op": "equal",
|
| 449 |
+
"original": " from the counter ",
|
| 450 |
+
"modified": " from the counter "
|
| 451 |
+
},
|
| 452 |
+
{
|
| 453 |
+
"op": "replace",
|
| 454 |
+
"original": "to",
|
| 455 |
+
"modified": "into"
|
| 456 |
+
},
|
| 457 |
+
{
|
| 458 |
+
"op": "equal",
|
| 459 |
+
"original": " the ",
|
| 460 |
+
"modified": " the "
|
| 461 |
+
},
|
| 462 |
+
{
|
| 463 |
+
"op": "replace",
|
| 464 |
+
"original": "sink.",
|
| 465 |
+
"modified": "sink while the water is running."
|
| 466 |
+
},
|
| 467 |
+
{
|
| 468 |
+
"op": "equal",
|
| 469 |
+
"original": " Turn off the ",
|
| 470 |
+
"modified": " Turn off the "
|
| 471 |
+
},
|
| 472 |
+
{
|
| 473 |
+
"op": "replace",
|
| 474 |
+
"original": "sink.",
|
| 475 |
+
"modified": "faucet,"
|
| 476 |
+
},
|
| 477 |
+
{
|
| 478 |
+
"op": "equal",
|
| 479 |
+
"original": " ",
|
| 480 |
+
"modified": " "
|
| 481 |
+
},
|
| 482 |
+
{
|
| 483 |
+
"op": "replace",
|
| 484 |
+
"original": "Move",
|
| 485 |
+
"modified": "move"
|
| 486 |
+
},
|
| 487 |
+
{
|
| 488 |
+
"op": "equal",
|
| 489 |
+
"original": " the vegetable ",
|
| 490 |
+
"modified": " the vegetable "
|
| 491 |
+
},
|
| 492 |
+
{
|
| 493 |
+
"op": "replace",
|
| 494 |
+
"original": "from the sink to",
|
| 495 |
+
"modified": "into"
|
| 496 |
+
},
|
| 497 |
+
{
|
| 498 |
+
"op": "equal",
|
| 499 |
+
"original": " the pot next to the ",
|
| 500 |
+
"modified": " the pot next to the "
|
| 501 |
+
},
|
| 502 |
+
{
|
| 503 |
+
"op": "replace",
|
| 504 |
+
"original": "stove.",
|
| 505 |
+
"modified": "stove,"
|
| 506 |
+
},
|
| 507 |
+
{
|
| 508 |
+
"op": "equal",
|
| 509 |
+
"original": " ",
|
| 510 |
+
"modified": " "
|
| 511 |
+
},
|
| 512 |
+
{
|
| 513 |
+
"op": "replace",
|
| 514 |
+
"original": "Finally",
|
| 515 |
+
"modified": "and"
|
| 516 |
+
},
|
| 517 |
+
{
|
| 518 |
+
"op": "equal",
|
| 519 |
+
"original": " move the pot to the ",
|
| 520 |
+
"modified": " move the pot to the "
|
| 521 |
+
},
|
| 522 |
+
{
|
| 523 |
+
"op": "replace",
|
| 524 |
+
"original": "{burner}",
|
| 525 |
+
"modified": "specified"
|
| 526 |
+
},
|
| 527 |
+
{
|
| 528 |
+
"op": "equal",
|
| 529 |
+
"original": " burner.",
|
| 530 |
+
"modified": " burner."
|
| 531 |
+
}
|
| 532 |
+
],
|
| 533 |
+
"changed": true,
|
| 534 |
+
"preserve_whitespace": true,
|
| 535 |
+
"label": "Modified instruction",
|
| 536 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/06-multistep-steaming/task.yaml"
|
| 537 |
+
},
|
| 538 |
+
"run_status": "finished",
|
| 539 |
+
"attempt_history": [],
|
| 540 |
+
"status_note": "Evaluation and native evidence complete."
|
| 541 |
+
},
|
| 542 |
+
{
|
| 543 |
+
"key": "task01/07",
|
| 544 |
+
"family": "task01",
|
| 545 |
+
"slot": "07",
|
| 546 |
+
"native_id": "robocasa/scale-portioning",
|
| 547 |
+
"title": "Scale portioning",
|
| 548 |
+
"catalog_instruction": "Take the specified meat from the fridge and place it on the digital scale on the counter by the fridge. Release it and move the gripper clear while waiting a few seconds for a reading. Then move it to the plate on the dining counter, release it, and move the gripper clear.",
|
| 549 |
+
"native_instruction": "Take the {meat} from the fridge and place it on the digital scale on the counter by the fridge. Wait a few seconds for a reading, then move it to the plate on the dining counter.",
|
| 550 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 551 |
+
"status": "completed",
|
| 552 |
+
"episode_id": "task01-07-seed0-formal",
|
| 553 |
+
"planned_protocol": {
|
| 554 |
+
"episodes": 1,
|
| 555 |
+
"seed": 0,
|
| 556 |
+
"control_frequency_hz": 20,
|
| 557 |
+
"max_control_steps": 6000,
|
| 558 |
+
"timeout_s": 28800,
|
| 559 |
+
"mode": "stepped",
|
| 560 |
+
"instruction_policy": "modified"
|
| 561 |
+
},
|
| 562 |
+
"catalog_id": "robocasa/scale-portioning",
|
| 563 |
+
"native_identity": {
|
| 564 |
+
"task_name": "ScalePortioning",
|
| 565 |
+
"max_steps": 6000
|
| 566 |
+
},
|
| 567 |
+
"instruction_revision": {
|
| 568 |
+
"native_instruction": "Take the {meat} from the fridge and place it on the digital scale on the counter by the fridge. Wait a few seconds for a reading, then move it to the plate on the dining counter.",
|
| 569 |
+
"instruction": "Take the specified meat from the fridge and place it on the digital scale on the counter by the fridge. Release it and move the gripper clear while waiting a few seconds for a reading. Then move it to the plate on the dining counter, release it, and move the gripper clear.",
|
| 570 |
+
"native_instruction_kind": "source",
|
| 571 |
+
"diff": [
|
| 572 |
+
{
|
| 573 |
+
"op": "equal",
|
| 574 |
+
"original": "Take the ",
|
| 575 |
+
"modified": "Take the "
|
| 576 |
+
},
|
| 577 |
+
{
|
| 578 |
+
"op": "replace",
|
| 579 |
+
"original": "{meat}",
|
| 580 |
+
"modified": "specified meat"
|
| 581 |
+
},
|
| 582 |
+
{
|
| 583 |
+
"op": "equal",
|
| 584 |
+
"original": " from the fridge and place it on the digital scale on the counter by the fridge. ",
|
| 585 |
+
"modified": " from the fridge and place it on the digital scale on the counter by the fridge. "
|
| 586 |
+
},
|
| 587 |
+
{
|
| 588 |
+
"op": "replace",
|
| 589 |
+
"original": "Wait",
|
| 590 |
+
"modified": "Release it and move the gripper clear while waiting"
|
| 591 |
+
},
|
| 592 |
+
{
|
| 593 |
+
"op": "equal",
|
| 594 |
+
"original": " a few seconds for a ",
|
| 595 |
+
"modified": " a few seconds for a "
|
| 596 |
+
},
|
| 597 |
+
{
|
| 598 |
+
"op": "replace",
|
| 599 |
+
"original": "reading,",
|
| 600 |
+
"modified": "reading."
|
| 601 |
+
},
|
| 602 |
+
{
|
| 603 |
+
"op": "equal",
|
| 604 |
+
"original": " ",
|
| 605 |
+
"modified": " "
|
| 606 |
+
},
|
| 607 |
+
{
|
| 608 |
+
"op": "replace",
|
| 609 |
+
"original": "then",
|
| 610 |
+
"modified": "Then"
|
| 611 |
+
},
|
| 612 |
+
{
|
| 613 |
+
"op": "equal",
|
| 614 |
+
"original": " move it to the plate on the dining ",
|
| 615 |
+
"modified": " move it to the plate on the dining "
|
| 616 |
+
},
|
| 617 |
+
{
|
| 618 |
+
"op": "replace",
|
| 619 |
+
"original": "counter.",
|
| 620 |
+
"modified": "counter, release it, and move the gripper clear."
|
| 621 |
+
}
|
| 622 |
+
],
|
| 623 |
+
"changed": true,
|
| 624 |
+
"preserve_whitespace": true,
|
| 625 |
+
"label": "Modified instruction",
|
| 626 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/07-scale-portioning/task.yaml"
|
| 627 |
+
},
|
| 628 |
+
"run_status": "finished",
|
| 629 |
+
"attempt_history": [],
|
| 630 |
+
"status_note": "Evaluation and native evidence complete."
|
| 631 |
+
},
|
| 632 |
+
{
|
| 633 |
+
"key": "task01/08",
|
| 634 |
+
"family": "task01",
|
| 635 |
+
"slot": "08",
|
| 636 |
+
"native_id": "robocasa/scrub-cutting-board",
|
| 637 |
+
"title": "Scrub cutting board",
|
| 638 |
+
"catalog_instruction": "Pick up the sponge from the counter and scrub across a broad area of the cutting board, keeping the sponge grasped and in contact with the board throughout the scrubbing motion. Once finished, release the sponge and retract the gripper well away from it.",
|
| 639 |
+
"native_instruction": "Pick up the sponge from the counter and clean the cutting board by briefly scrubbing or pressing down on the cutting board. Once finished, release the sponge.",
|
| 640 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 641 |
+
"status": "completed",
|
| 642 |
+
"episode_id": "task01-08-seed0-formal",
|
| 643 |
+
"planned_protocol": {
|
| 644 |
+
"episodes": 1,
|
| 645 |
+
"seed": 0,
|
| 646 |
+
"control_frequency_hz": 20,
|
| 647 |
+
"max_control_steps": 6000,
|
| 648 |
+
"timeout_s": 28800,
|
| 649 |
+
"mode": "stepped",
|
| 650 |
+
"instruction_policy": "modified"
|
| 651 |
+
},
|
| 652 |
+
"catalog_id": "robocasa/scrub-cutting-board",
|
| 653 |
+
"native_identity": {
|
| 654 |
+
"task_name": "ScrubCuttingBoard",
|
| 655 |
+
"max_steps": 6000
|
| 656 |
+
},
|
| 657 |
+
"instruction_revision": {
|
| 658 |
+
"native_instruction": "Pick up the sponge from the counter and clean the cutting board by briefly scrubbing or pressing down on the cutting board. Once finished, release the sponge.",
|
| 659 |
+
"instruction": "Pick up the sponge from the counter and scrub across a broad area of the cutting board, keeping the sponge grasped and in contact with the board throughout the scrubbing motion. Once finished, release the sponge and retract the gripper well away from it.",
|
| 660 |
+
"native_instruction_kind": "source",
|
| 661 |
+
"diff": [
|
| 662 |
+
{
|
| 663 |
+
"op": "equal",
|
| 664 |
+
"original": "Pick up the sponge from the counter and ",
|
| 665 |
+
"modified": "Pick up the sponge from the counter and "
|
| 666 |
+
},
|
| 667 |
+
{
|
| 668 |
+
"op": "insert",
|
| 669 |
+
"original": "",
|
| 670 |
+
"modified": "s"
|
| 671 |
+
},
|
| 672 |
+
{
|
| 673 |
+
"op": "equal",
|
| 674 |
+
"original": "c",
|
| 675 |
+
"modified": "c"
|
| 676 |
+
},
|
| 677 |
+
{
|
| 678 |
+
"op": "replace",
|
| 679 |
+
"original": "l",
|
| 680 |
+
"modified": "rub across a broad ar"
|
| 681 |
+
},
|
| 682 |
+
{
|
| 683 |
+
"op": "equal",
|
| 684 |
+
"original": "ea",
|
| 685 |
+
"modified": "ea"
|
| 686 |
+
},
|
| 687 |
+
{
|
| 688 |
+
"op": "replace",
|
| 689 |
+
"original": "n",
|
| 690 |
+
"modified": " of"
|
| 691 |
+
},
|
| 692 |
+
{
|
| 693 |
+
"op": "equal",
|
| 694 |
+
"original": " the cutting board",
|
| 695 |
+
"modified": " the cutting board"
|
| 696 |
+
},
|
| 697 |
+
{
|
| 698 |
+
"op": "insert",
|
| 699 |
+
"original": "",
|
| 700 |
+
"modified": ", keeping the sponge grasped and in contact with the"
|
| 701 |
+
},
|
| 702 |
+
{
|
| 703 |
+
"op": "equal",
|
| 704 |
+
"original": " b",
|
| 705 |
+
"modified": " b"
|
| 706 |
+
},
|
| 707 |
+
{
|
| 708 |
+
"op": "replace",
|
| 709 |
+
"original": "y",
|
| 710 |
+
"modified": "oard"
|
| 711 |
+
},
|
| 712 |
+
{
|
| 713 |
+
"op": "equal",
|
| 714 |
+
"original": " ",
|
| 715 |
+
"modified": " "
|
| 716 |
+
},
|
| 717 |
+
{
|
| 718 |
+
"op": "replace",
|
| 719 |
+
"original": "b",
|
| 720 |
+
"modified": "th"
|
| 721 |
+
},
|
| 722 |
+
{
|
| 723 |
+
"op": "equal",
|
| 724 |
+
"original": "r",
|
| 725 |
+
"modified": "r"
|
| 726 |
+
},
|
| 727 |
+
{
|
| 728 |
+
"op": "replace",
|
| 729 |
+
"original": "i",
|
| 730 |
+
"modified": "oughout th"
|
| 731 |
+
},
|
| 732 |
+
{
|
| 733 |
+
"op": "equal",
|
| 734 |
+
"original": "e",
|
| 735 |
+
"modified": "e"
|
| 736 |
+
},
|
| 737 |
+
{
|
| 738 |
+
"op": "delete",
|
| 739 |
+
"original": "fly",
|
| 740 |
+
"modified": ""
|
| 741 |
+
},
|
| 742 |
+
{
|
| 743 |
+
"op": "equal",
|
| 744 |
+
"original": " scrubbing ",
|
| 745 |
+
"modified": " scrubbing "
|
| 746 |
+
},
|
| 747 |
+
{
|
| 748 |
+
"op": "insert",
|
| 749 |
+
"original": "",
|
| 750 |
+
"modified": "m"
|
| 751 |
+
},
|
| 752 |
+
{
|
| 753 |
+
"op": "equal",
|
| 754 |
+
"original": "o",
|
| 755 |
+
"modified": "o"
|
| 756 |
+
},
|
| 757 |
+
{
|
| 758 |
+
"op": "replace",
|
| 759 |
+
"original": "r press",
|
| 760 |
+
"modified": "t"
|
| 761 |
+
},
|
| 762 |
+
{
|
| 763 |
+
"op": "equal",
|
| 764 |
+
"original": "i",
|
| 765 |
+
"modified": "i"
|
| 766 |
+
},
|
| 767 |
+
{
|
| 768 |
+
"op": "delete",
|
| 769 |
+
"original": "ng down ",
|
| 770 |
+
"modified": ""
|
| 771 |
+
},
|
| 772 |
+
{
|
| 773 |
+
"op": "equal",
|
| 774 |
+
"original": "on",
|
| 775 |
+
"modified": "on"
|
| 776 |
+
},
|
| 777 |
+
{
|
| 778 |
+
"op": "delete",
|
| 779 |
+
"original": " the cutting board",
|
| 780 |
+
"modified": ""
|
| 781 |
+
},
|
| 782 |
+
{
|
| 783 |
+
"op": "equal",
|
| 784 |
+
"original": ". Once finished, release the sponge",
|
| 785 |
+
"modified": ". Once finished, release the sponge"
|
| 786 |
+
},
|
| 787 |
+
{
|
| 788 |
+
"op": "insert",
|
| 789 |
+
"original": "",
|
| 790 |
+
"modified": " and retract the gripper well away from it"
|
| 791 |
+
},
|
| 792 |
+
{
|
| 793 |
+
"op": "equal",
|
| 794 |
+
"original": ".",
|
| 795 |
+
"modified": "."
|
| 796 |
+
}
|
| 797 |
+
],
|
| 798 |
+
"changed": true,
|
| 799 |
+
"preserve_whitespace": true,
|
| 800 |
+
"label": "Modified instruction",
|
| 801 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/08-scrub-cutting-board/task.yaml",
|
| 802 |
+
"native_predicates_unchanged": true,
|
| 803 |
+
"evaluation_status": "Evaluated with this modified instruction.",
|
| 804 |
+
"review_reason": "Clarify broad board coverage, maintained grasp/contact, and final retraction without numeric thresholds."
|
| 805 |
+
},
|
| 806 |
+
"run_status": "finished",
|
| 807 |
+
"attempt_history": [],
|
| 808 |
+
"status_note": "Evaluation and native evidence complete."
|
| 809 |
+
},
|
| 810 |
+
{
|
| 811 |
+
"key": "task01/09",
|
| 812 |
+
"family": "task01",
|
| 813 |
+
"slot": "09",
|
| 814 |
+
"native_id": "robocasa/prepare-veggie-dip",
|
| 815 |
+
"title": "Prepare veggie dip",
|
| 816 |
+
"catalog_instruction": "Pick the specified vegetable and the cream cheese from the fridge, place both fully inside the blender, and turn it on.",
|
| 817 |
+
"native_instruction": "Pick the {vegetable} and the cream cheese from the fridge, place them in the blender, and turn it on.",
|
| 818 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 819 |
+
"status": "completed",
|
| 820 |
+
"episode_id": "task01-09-seed0-formal",
|
| 821 |
+
"planned_protocol": {
|
| 822 |
+
"episodes": 1,
|
| 823 |
+
"seed": 0,
|
| 824 |
+
"control_frequency_hz": 20,
|
| 825 |
+
"max_control_steps": 6000,
|
| 826 |
+
"timeout_s": 28800,
|
| 827 |
+
"mode": "stepped",
|
| 828 |
+
"instruction_policy": "modified"
|
| 829 |
+
},
|
| 830 |
+
"catalog_id": "robocasa/prepare-veggie-dip",
|
| 831 |
+
"native_identity": {
|
| 832 |
+
"task_name": "PrepareVeggieDip",
|
| 833 |
+
"max_steps": 6000
|
| 834 |
+
},
|
| 835 |
+
"instruction_revision": {
|
| 836 |
+
"native_instruction": "Pick the {vegetable} and the cream cheese from the fridge, place them in the blender, and turn it on.",
|
| 837 |
+
"instruction": "Pick the specified vegetable and the cream cheese from the fridge, place both fully inside the blender, and turn it on.",
|
| 838 |
+
"native_instruction_kind": "source",
|
| 839 |
+
"diff": [
|
| 840 |
+
{
|
| 841 |
+
"op": "equal",
|
| 842 |
+
"original": "Pick the ",
|
| 843 |
+
"modified": "Pick the "
|
| 844 |
+
},
|
| 845 |
+
{
|
| 846 |
+
"op": "replace",
|
| 847 |
+
"original": "{vegetable}",
|
| 848 |
+
"modified": "specified vegetable"
|
| 849 |
+
},
|
| 850 |
+
{
|
| 851 |
+
"op": "equal",
|
| 852 |
+
"original": " and the cream cheese from the fridge, place ",
|
| 853 |
+
"modified": " and the cream cheese from the fridge, place "
|
| 854 |
+
},
|
| 855 |
+
{
|
| 856 |
+
"op": "replace",
|
| 857 |
+
"original": "them",
|
| 858 |
+
"modified": "both"
|
| 859 |
+
},
|
| 860 |
+
{
|
| 861 |
+
"op": "equal",
|
| 862 |
+
"original": " ",
|
| 863 |
+
"modified": " "
|
| 864 |
+
},
|
| 865 |
+
{
|
| 866 |
+
"op": "replace",
|
| 867 |
+
"original": "in",
|
| 868 |
+
"modified": "fully inside"
|
| 869 |
+
},
|
| 870 |
+
{
|
| 871 |
+
"op": "equal",
|
| 872 |
+
"original": " the blender, and turn it on.",
|
| 873 |
+
"modified": " the blender, and turn it on."
|
| 874 |
+
}
|
| 875 |
+
],
|
| 876 |
+
"changed": true,
|
| 877 |
+
"preserve_whitespace": true,
|
| 878 |
+
"label": "Modified instruction",
|
| 879 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/09-prepare-veggie-dip/task.yaml"
|
| 880 |
+
},
|
| 881 |
+
"run_status": "finished",
|
| 882 |
+
"attempt_history": [],
|
| 883 |
+
"status_note": "Evaluation and native evidence complete."
|
| 884 |
+
},
|
| 885 |
+
{
|
| 886 |
+
"key": "task01/10",
|
| 887 |
+
"family": "task01",
|
| 888 |
+
"slot": "10",
|
| 889 |
+
"native_id": "robocasa/prepare-vegetable-roasting",
|
| 890 |
+
"title": "Prepare vegetable roasting",
|
| 891 |
+
"catalog_instruction": "Pick the specified vegetable from the fridge and hold it under running water from the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting. Release it and move the gripper clear.",
|
| 892 |
+
"native_instruction": "Pick the {vegetable} from the fridge and hold it under the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting.",
|
| 893 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 894 |
+
"status": "completed",
|
| 895 |
+
"episode_id": "task01-10-seed0-formal",
|
| 896 |
+
"planned_protocol": {
|
| 897 |
+
"episodes": 1,
|
| 898 |
+
"seed": 0,
|
| 899 |
+
"control_frequency_hz": 20,
|
| 900 |
+
"max_control_steps": 6000,
|
| 901 |
+
"timeout_s": 28800,
|
| 902 |
+
"mode": "stepped",
|
| 903 |
+
"instruction_policy": "modified"
|
| 904 |
+
},
|
| 905 |
+
"catalog_id": "robocasa/prepare-vegetable-roasting",
|
| 906 |
+
"native_identity": {
|
| 907 |
+
"task_name": "PrepareVegetableRoasting",
|
| 908 |
+
"max_steps": 6000
|
| 909 |
+
},
|
| 910 |
+
"instruction_revision": {
|
| 911 |
+
"native_instruction": "Pick the {vegetable} from the fridge and hold it under the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting.",
|
| 912 |
+
"instruction": "Pick the specified vegetable from the fridge and hold it under running water from the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting. Release it and move the gripper clear.",
|
| 913 |
+
"native_instruction_kind": "source",
|
| 914 |
+
"diff": [
|
| 915 |
+
{
|
| 916 |
+
"op": "equal",
|
| 917 |
+
"original": "Pick the ",
|
| 918 |
+
"modified": "Pick the "
|
| 919 |
+
},
|
| 920 |
+
{
|
| 921 |
+
"op": "replace",
|
| 922 |
+
"original": "{vegetable}",
|
| 923 |
+
"modified": "specified vegetable"
|
| 924 |
+
},
|
| 925 |
+
{
|
| 926 |
+
"op": "equal",
|
| 927 |
+
"original": " from the fridge and hold it under",
|
| 928 |
+
"modified": " from the fridge and hold it under"
|
| 929 |
+
},
|
| 930 |
+
{
|
| 931 |
+
"op": "insert",
|
| 932 |
+
"original": "",
|
| 933 |
+
"modified": " running water from"
|
| 934 |
+
},
|
| 935 |
+
{
|
| 936 |
+
"op": "equal",
|
| 937 |
+
"original": " the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting.",
|
| 938 |
+
"modified": " the sink faucet to wash it. Then place it on the tray next to the sink to prepare for roasting."
|
| 939 |
+
},
|
| 940 |
+
{
|
| 941 |
+
"op": "insert",
|
| 942 |
+
"original": "",
|
| 943 |
+
"modified": " Release it and move the gripper clear."
|
| 944 |
+
}
|
| 945 |
+
],
|
| 946 |
+
"changed": true,
|
| 947 |
+
"preserve_whitespace": true,
|
| 948 |
+
"label": "Modified instruction",
|
| 949 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task01/10-prepare-vegetable-roasting/task.yaml"
|
| 950 |
+
},
|
| 951 |
+
"run_status": "finished",
|
| 952 |
+
"attempt_history": [],
|
| 953 |
+
"status_note": "Evaluation and native evidence complete."
|
| 954 |
+
},
|
| 955 |
+
{
|
| 956 |
+
"key": "task02/01",
|
| 957 |
+
"family": "task02",
|
| 958 |
+
"slot": "01",
|
| 959 |
+
"native_id": "libero/10-0",
|
| 960 |
+
"title": "LIBERO-10-01",
|
| 961 |
+
"catalog_instruction": "put both the alphabet soup and the tomato sauce in the basket",
|
| 962 |
+
"native_instruction": "put both the alphabet soup and the tomato sauce in the basket",
|
| 963 |
+
"instruction_source": "runtime native task.instruction",
|
| 964 |
+
"status": "completed",
|
| 965 |
+
"episode_id": "task02-01-seed0-formal",
|
| 966 |
+
"planned_protocol": {
|
| 967 |
+
"episodes": 1,
|
| 968 |
+
"seed": 0,
|
| 969 |
+
"control_frequency_hz": 20,
|
| 970 |
+
"max_control_steps": 6000,
|
| 971 |
+
"timeout_s": 28800,
|
| 972 |
+
"mode": "stepped",
|
| 973 |
+
"instruction_policy": "original_native"
|
| 974 |
+
},
|
| 975 |
+
"catalog_id": "libero/libero-10-01",
|
| 976 |
+
"native_identity": {
|
| 977 |
+
"suite_name": "libero_10",
|
| 978 |
+
"task_id": 0,
|
| 979 |
+
"init_state_index": 0,
|
| 980 |
+
"max_steps": 6000
|
| 981 |
+
},
|
| 982 |
+
"run_status": "finished",
|
| 983 |
+
"attempt_history": [],
|
| 984 |
+
"status_note": "Evaluation and native evidence complete."
|
| 985 |
+
},
|
| 986 |
+
{
|
| 987 |
+
"key": "task02/02",
|
| 988 |
+
"family": "task02",
|
| 989 |
+
"slot": "02",
|
| 990 |
+
"native_id": "libero/10-1",
|
| 991 |
+
"title": "LIBERO-10-02",
|
| 992 |
+
"catalog_instruction": "put both the cream cheese box and the butter in the basket",
|
| 993 |
+
"native_instruction": "put both the cream cheese box and the butter in the basket",
|
| 994 |
+
"instruction_source": "runtime native task.instruction",
|
| 995 |
+
"status": "completed",
|
| 996 |
+
"episode_id": "task02-02-seed0-formal",
|
| 997 |
+
"planned_protocol": {
|
| 998 |
+
"episodes": 1,
|
| 999 |
+
"seed": 0,
|
| 1000 |
+
"control_frequency_hz": 20,
|
| 1001 |
+
"max_control_steps": 6000,
|
| 1002 |
+
"timeout_s": 28800,
|
| 1003 |
+
"mode": "stepped",
|
| 1004 |
+
"instruction_policy": "original_native"
|
| 1005 |
+
},
|
| 1006 |
+
"catalog_id": "libero/libero-10-02",
|
| 1007 |
+
"native_identity": {
|
| 1008 |
+
"suite_name": "libero_10",
|
| 1009 |
+
"task_id": 1,
|
| 1010 |
+
"init_state_index": 0,
|
| 1011 |
+
"max_steps": 6000
|
| 1012 |
+
},
|
| 1013 |
+
"run_status": "finished",
|
| 1014 |
+
"attempt_history": [],
|
| 1015 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1016 |
+
},
|
| 1017 |
+
{
|
| 1018 |
+
"key": "task02/03",
|
| 1019 |
+
"family": "task02",
|
| 1020 |
+
"slot": "03",
|
| 1021 |
+
"native_id": "libero/10-2",
|
| 1022 |
+
"title": "LIBERO-10-03",
|
| 1023 |
+
"catalog_instruction": "turn on the stove and put the moka pot on it",
|
| 1024 |
+
"native_instruction": "turn on the stove and put the moka pot on it",
|
| 1025 |
+
"instruction_source": "runtime native task.instruction",
|
| 1026 |
+
"status": "completed",
|
| 1027 |
+
"episode_id": "task02-03-seed0-formal",
|
| 1028 |
+
"planned_protocol": {
|
| 1029 |
+
"episodes": 1,
|
| 1030 |
+
"seed": 0,
|
| 1031 |
+
"control_frequency_hz": 20,
|
| 1032 |
+
"max_control_steps": 6000,
|
| 1033 |
+
"timeout_s": 28800,
|
| 1034 |
+
"mode": "stepped",
|
| 1035 |
+
"instruction_policy": "original_native"
|
| 1036 |
+
},
|
| 1037 |
+
"catalog_id": "libero/libero-10-03",
|
| 1038 |
+
"native_identity": {
|
| 1039 |
+
"suite_name": "libero_10",
|
| 1040 |
+
"task_id": 2,
|
| 1041 |
+
"init_state_index": 0,
|
| 1042 |
+
"max_steps": 6000
|
| 1043 |
+
},
|
| 1044 |
+
"run_status": "finished",
|
| 1045 |
+
"attempt_history": [],
|
| 1046 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1047 |
+
},
|
| 1048 |
+
{
|
| 1049 |
+
"key": "task02/04",
|
| 1050 |
+
"family": "task02",
|
| 1051 |
+
"slot": "04",
|
| 1052 |
+
"native_id": "libero/bowl-into-bottom-drawer",
|
| 1053 |
+
"title": "LIBERO-10-04",
|
| 1054 |
+
"catalog_instruction": "put the black bowl in the bottom drawer of the cabinet and close it",
|
| 1055 |
+
"native_instruction": "put the black bowl in the bottom drawer of the cabinet and close it",
|
| 1056 |
+
"instruction_source": "runtime native task.instruction",
|
| 1057 |
+
"status": "completed",
|
| 1058 |
+
"episode_id": "task02-04-seed0-formal",
|
| 1059 |
+
"planned_protocol": {
|
| 1060 |
+
"episodes": 1,
|
| 1061 |
+
"seed": 0,
|
| 1062 |
+
"control_frequency_hz": 20,
|
| 1063 |
+
"max_control_steps": 6000,
|
| 1064 |
+
"timeout_s": 28800,
|
| 1065 |
+
"mode": "stepped",
|
| 1066 |
+
"instruction_policy": "original_native"
|
| 1067 |
+
},
|
| 1068 |
+
"catalog_id": "libero/libero-10-04",
|
| 1069 |
+
"native_identity": {
|
| 1070 |
+
"suite_name": "libero_10",
|
| 1071 |
+
"task_id": 3,
|
| 1072 |
+
"init_state_index": 0,
|
| 1073 |
+
"max_steps": 6000
|
| 1074 |
+
},
|
| 1075 |
+
"run_status": "finished",
|
| 1076 |
+
"attempt_history": [],
|
| 1077 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1078 |
+
},
|
| 1079 |
+
{
|
| 1080 |
+
"key": "task02/05",
|
| 1081 |
+
"family": "task02",
|
| 1082 |
+
"slot": "05",
|
| 1083 |
+
"native_id": "libero/10-4",
|
| 1084 |
+
"title": "LIBERO-10-05",
|
| 1085 |
+
"catalog_instruction": "put the white mug in the center of the left plate and put the yellow and white mug in the center of the right plate, with each mug resting on its plate",
|
| 1086 |
+
"native_instruction": "put the white mug on the left plate and put the yellow and white mug on the right plate",
|
| 1087 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 1088 |
+
"status": "completed",
|
| 1089 |
+
"episode_id": "task02-05-seed0-formal",
|
| 1090 |
+
"planned_protocol": {
|
| 1091 |
+
"episodes": 1,
|
| 1092 |
+
"seed": 0,
|
| 1093 |
+
"control_frequency_hz": 20,
|
| 1094 |
+
"max_control_steps": 6000,
|
| 1095 |
+
"timeout_s": 28800,
|
| 1096 |
+
"mode": "stepped",
|
| 1097 |
+
"instruction_policy": "modified"
|
| 1098 |
+
},
|
| 1099 |
+
"catalog_id": "libero/libero-10-05",
|
| 1100 |
+
"native_identity": {
|
| 1101 |
+
"suite_name": "libero_10",
|
| 1102 |
+
"task_id": 4,
|
| 1103 |
+
"init_state_index": 0,
|
| 1104 |
+
"max_steps": 6000
|
| 1105 |
+
},
|
| 1106 |
+
"instruction_revision": {
|
| 1107 |
+
"native_instruction": "put the white mug on the left plate and put the yellow and white mug on the right plate",
|
| 1108 |
+
"instruction": "put the white mug in the center of the left plate and put the yellow and white mug in the center of the right plate, with each mug resting on its plate",
|
| 1109 |
+
"native_instruction_kind": "source",
|
| 1110 |
+
"diff": [
|
| 1111 |
+
{
|
| 1112 |
+
"op": "equal",
|
| 1113 |
+
"original": "put the white mug ",
|
| 1114 |
+
"modified": "put the white mug "
|
| 1115 |
+
},
|
| 1116 |
+
{
|
| 1117 |
+
"op": "insert",
|
| 1118 |
+
"original": "",
|
| 1119 |
+
"modified": "in the center "
|
| 1120 |
+
},
|
| 1121 |
+
{
|
| 1122 |
+
"op": "equal",
|
| 1123 |
+
"original": "o",
|
| 1124 |
+
"modified": "o"
|
| 1125 |
+
},
|
| 1126 |
+
{
|
| 1127 |
+
"op": "replace",
|
| 1128 |
+
"original": "n",
|
| 1129 |
+
"modified": "f"
|
| 1130 |
+
},
|
| 1131 |
+
{
|
| 1132 |
+
"op": "equal",
|
| 1133 |
+
"original": " the left plate and put the yellow and white mug ",
|
| 1134 |
+
"modified": " the left plate and put the yellow and white mug "
|
| 1135 |
+
},
|
| 1136 |
+
{
|
| 1137 |
+
"op": "insert",
|
| 1138 |
+
"original": "",
|
| 1139 |
+
"modified": "in the center "
|
| 1140 |
+
},
|
| 1141 |
+
{
|
| 1142 |
+
"op": "equal",
|
| 1143 |
+
"original": "o",
|
| 1144 |
+
"modified": "o"
|
| 1145 |
+
},
|
| 1146 |
+
{
|
| 1147 |
+
"op": "replace",
|
| 1148 |
+
"original": "n",
|
| 1149 |
+
"modified": "f"
|
| 1150 |
+
},
|
| 1151 |
+
{
|
| 1152 |
+
"op": "equal",
|
| 1153 |
+
"original": " the right plate",
|
| 1154 |
+
"modified": " the right plate"
|
| 1155 |
+
},
|
| 1156 |
+
{
|
| 1157 |
+
"op": "insert",
|
| 1158 |
+
"original": "",
|
| 1159 |
+
"modified": ", with each mug resting on its plate"
|
| 1160 |
+
}
|
| 1161 |
+
],
|
| 1162 |
+
"changed": true,
|
| 1163 |
+
"preserve_whitespace": true,
|
| 1164 |
+
"label": "Modified instruction",
|
| 1165 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task02/05-libero-10-05/task.yaml",
|
| 1166 |
+
"native_predicates_unchanged": true,
|
| 1167 |
+
"evaluation_status": "Evaluated with this modified instruction.",
|
| 1168 |
+
"review_reason": "Authorized full current-source LIBERO comparison refresh."
|
| 1169 |
+
},
|
| 1170 |
+
"run_status": "finished",
|
| 1171 |
+
"attempt_history": [],
|
| 1172 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1173 |
+
},
|
| 1174 |
+
{
|
| 1175 |
+
"key": "task02/06",
|
| 1176 |
+
"family": "task02",
|
| 1177 |
+
"slot": "06",
|
| 1178 |
+
"native_id": "libero/10-5",
|
| 1179 |
+
"title": "LIBERO-10-06",
|
| 1180 |
+
"catalog_instruction": "pick up the book and place it in the back compartment of the caddy, between the two large side compartments",
|
| 1181 |
+
"native_instruction": "pick up the book and place it in the back compartment of the caddy",
|
| 1182 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 1183 |
+
"status": "completed",
|
| 1184 |
+
"episode_id": "task02-06-seed0-formal",
|
| 1185 |
+
"planned_protocol": {
|
| 1186 |
+
"episodes": 1,
|
| 1187 |
+
"seed": 0,
|
| 1188 |
+
"control_frequency_hz": 20,
|
| 1189 |
+
"max_control_steps": 6000,
|
| 1190 |
+
"timeout_s": 28800,
|
| 1191 |
+
"mode": "stepped",
|
| 1192 |
+
"instruction_policy": "modified"
|
| 1193 |
+
},
|
| 1194 |
+
"catalog_id": "libero/libero-10-06",
|
| 1195 |
+
"native_identity": {
|
| 1196 |
+
"suite_name": "libero_10",
|
| 1197 |
+
"task_id": 5,
|
| 1198 |
+
"init_state_index": 0,
|
| 1199 |
+
"max_steps": 6000
|
| 1200 |
+
},
|
| 1201 |
+
"instruction_revision": {
|
| 1202 |
+
"native_instruction": "pick up the book and place it in the back compartment of the caddy",
|
| 1203 |
+
"instruction": "pick up the book and place it in the back compartment of the caddy, between the two large side compartments",
|
| 1204 |
+
"native_instruction_kind": "source",
|
| 1205 |
+
"diff": [
|
| 1206 |
+
{
|
| 1207 |
+
"op": "equal",
|
| 1208 |
+
"original": "pick up the book and place it in the back compartment of the caddy",
|
| 1209 |
+
"modified": "pick up the book and place it in the back compartment of the caddy"
|
| 1210 |
+
},
|
| 1211 |
+
{
|
| 1212 |
+
"op": "insert",
|
| 1213 |
+
"original": "",
|
| 1214 |
+
"modified": ", between the two large side compartments"
|
| 1215 |
+
}
|
| 1216 |
+
],
|
| 1217 |
+
"changed": true,
|
| 1218 |
+
"preserve_whitespace": true,
|
| 1219 |
+
"label": "Modified instruction",
|
| 1220 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task02/06-libero-10-06/task.yaml",
|
| 1221 |
+
"native_predicates_unchanged": true,
|
| 1222 |
+
"evaluation_status": "Evaluated with this modified instruction.",
|
| 1223 |
+
"review_reason": "Authorized full current-source LIBERO comparison refresh."
|
| 1224 |
+
},
|
| 1225 |
+
"run_status": "finished",
|
| 1226 |
+
"attempt_history": [],
|
| 1227 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1228 |
+
},
|
| 1229 |
+
{
|
| 1230 |
+
"key": "task02/07",
|
| 1231 |
+
"family": "task02",
|
| 1232 |
+
"slot": "07",
|
| 1233 |
+
"native_id": "libero/10-6",
|
| 1234 |
+
"title": "LIBERO-10-07",
|
| 1235 |
+
"catalog_instruction": "put the white mug in the center of the plate and put the chocolate pudding immediately to the right of the plate",
|
| 1236 |
+
"native_instruction": "put the white mug on the plate and put the chocolate pudding to the right of the plate",
|
| 1237 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 1238 |
+
"status": "completed",
|
| 1239 |
+
"episode_id": "task02-07-seed0-formal",
|
| 1240 |
+
"planned_protocol": {
|
| 1241 |
+
"episodes": 1,
|
| 1242 |
+
"seed": 0,
|
| 1243 |
+
"control_frequency_hz": 20,
|
| 1244 |
+
"max_control_steps": 6000,
|
| 1245 |
+
"timeout_s": 28800,
|
| 1246 |
+
"mode": "stepped",
|
| 1247 |
+
"instruction_policy": "modified"
|
| 1248 |
+
},
|
| 1249 |
+
"catalog_id": "libero/libero-10-07",
|
| 1250 |
+
"native_identity": {
|
| 1251 |
+
"suite_name": "libero_10",
|
| 1252 |
+
"task_id": 6,
|
| 1253 |
+
"init_state_index": 0,
|
| 1254 |
+
"max_steps": 6000
|
| 1255 |
+
},
|
| 1256 |
+
"instruction_revision": {
|
| 1257 |
+
"native_instruction": "put the white mug on the plate and put the chocolate pudding to the right of the plate",
|
| 1258 |
+
"instruction": "put the white mug in the center of the plate and put the chocolate pudding immediately to the right of the plate",
|
| 1259 |
+
"native_instruction_kind": "source",
|
| 1260 |
+
"diff": [
|
| 1261 |
+
{
|
| 1262 |
+
"op": "equal",
|
| 1263 |
+
"original": "put the white mug ",
|
| 1264 |
+
"modified": "put the white mug "
|
| 1265 |
+
},
|
| 1266 |
+
{
|
| 1267 |
+
"op": "insert",
|
| 1268 |
+
"original": "",
|
| 1269 |
+
"modified": "in the center "
|
| 1270 |
+
},
|
| 1271 |
+
{
|
| 1272 |
+
"op": "equal",
|
| 1273 |
+
"original": "o",
|
| 1274 |
+
"modified": "o"
|
| 1275 |
+
},
|
| 1276 |
+
{
|
| 1277 |
+
"op": "replace",
|
| 1278 |
+
"original": "n",
|
| 1279 |
+
"modified": "f"
|
| 1280 |
+
},
|
| 1281 |
+
{
|
| 1282 |
+
"op": "equal",
|
| 1283 |
+
"original": " the plate and put the chocolate pudding ",
|
| 1284 |
+
"modified": " the plate and put the chocolate pudding "
|
| 1285 |
+
},
|
| 1286 |
+
{
|
| 1287 |
+
"op": "insert",
|
| 1288 |
+
"original": "",
|
| 1289 |
+
"modified": "immediately "
|
| 1290 |
+
},
|
| 1291 |
+
{
|
| 1292 |
+
"op": "equal",
|
| 1293 |
+
"original": "to the right of the plate",
|
| 1294 |
+
"modified": "to the right of the plate"
|
| 1295 |
+
}
|
| 1296 |
+
],
|
| 1297 |
+
"changed": true,
|
| 1298 |
+
"preserve_whitespace": true,
|
| 1299 |
+
"label": "Modified instruction",
|
| 1300 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task02/07-libero-10-07/task.yaml",
|
| 1301 |
+
"native_predicates_unchanged": true,
|
| 1302 |
+
"evaluation_status": "Evaluated with this modified instruction.",
|
| 1303 |
+
"review_reason": "Authorized full current-source LIBERO comparison refresh."
|
| 1304 |
+
},
|
| 1305 |
+
"run_status": "finished",
|
| 1306 |
+
"attempt_history": [],
|
| 1307 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1308 |
+
},
|
| 1309 |
+
{
|
| 1310 |
+
"key": "task02/08",
|
| 1311 |
+
"family": "task02",
|
| 1312 |
+
"slot": "08",
|
| 1313 |
+
"native_id": "libero/10-7",
|
| 1314 |
+
"title": "LIBERO-10-08",
|
| 1315 |
+
"catalog_instruction": "put both the alphabet soup and the cream cheese box fully inside the basket",
|
| 1316 |
+
"native_instruction": "put both the alphabet soup and the cream cheese box in the basket",
|
| 1317 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 1318 |
+
"status": "completed",
|
| 1319 |
+
"episode_id": "task02-08-seed0-formal",
|
| 1320 |
+
"planned_protocol": {
|
| 1321 |
+
"episodes": 1,
|
| 1322 |
+
"seed": 0,
|
| 1323 |
+
"control_frequency_hz": 20,
|
| 1324 |
+
"max_control_steps": 6000,
|
| 1325 |
+
"timeout_s": 28800,
|
| 1326 |
+
"mode": "stepped",
|
| 1327 |
+
"instruction_policy": "modified"
|
| 1328 |
+
},
|
| 1329 |
+
"catalog_id": "libero/libero-10-08",
|
| 1330 |
+
"native_identity": {
|
| 1331 |
+
"suite_name": "libero_10",
|
| 1332 |
+
"task_id": 7,
|
| 1333 |
+
"init_state_index": 0,
|
| 1334 |
+
"max_steps": 6000
|
| 1335 |
+
},
|
| 1336 |
+
"instruction_revision": {
|
| 1337 |
+
"native_instruction": "put both the alphabet soup and the cream cheese box in the basket",
|
| 1338 |
+
"instruction": "put both the alphabet soup and the cream cheese box fully inside the basket",
|
| 1339 |
+
"native_instruction_kind": "source",
|
| 1340 |
+
"diff": [
|
| 1341 |
+
{
|
| 1342 |
+
"op": "equal",
|
| 1343 |
+
"original": "put both the alphabet soup and the cream cheese box ",
|
| 1344 |
+
"modified": "put both the alphabet soup and the cream cheese box "
|
| 1345 |
+
},
|
| 1346 |
+
{
|
| 1347 |
+
"op": "insert",
|
| 1348 |
+
"original": "",
|
| 1349 |
+
"modified": "fully "
|
| 1350 |
+
},
|
| 1351 |
+
{
|
| 1352 |
+
"op": "equal",
|
| 1353 |
+
"original": "in",
|
| 1354 |
+
"modified": "in"
|
| 1355 |
+
},
|
| 1356 |
+
{
|
| 1357 |
+
"op": "insert",
|
| 1358 |
+
"original": "",
|
| 1359 |
+
"modified": "side"
|
| 1360 |
+
},
|
| 1361 |
+
{
|
| 1362 |
+
"op": "equal",
|
| 1363 |
+
"original": " the basket",
|
| 1364 |
+
"modified": " the basket"
|
| 1365 |
+
}
|
| 1366 |
+
],
|
| 1367 |
+
"changed": true,
|
| 1368 |
+
"preserve_whitespace": true,
|
| 1369 |
+
"label": "Modified instruction",
|
| 1370 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task02/08-libero-10-08/task.yaml",
|
| 1371 |
+
"native_predicates_unchanged": true,
|
| 1372 |
+
"evaluation_status": "Evaluated with this modified instruction.",
|
| 1373 |
+
"review_reason": "Authorized full current-source LIBERO comparison refresh."
|
| 1374 |
+
},
|
| 1375 |
+
"run_status": "finished",
|
| 1376 |
+
"attempt_history": [],
|
| 1377 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1378 |
+
},
|
| 1379 |
+
{
|
| 1380 |
+
"key": "task02/09",
|
| 1381 |
+
"family": "task02",
|
| 1382 |
+
"slot": "09",
|
| 1383 |
+
"native_id": "libero/10-8",
|
| 1384 |
+
"title": "LIBERO-10-09",
|
| 1385 |
+
"catalog_instruction": "put both moka pots on the stove and turn the stove on",
|
| 1386 |
+
"native_instruction": "put both moka pots on the stove",
|
| 1387 |
+
"instruction_source": "Current reviewed benchmark instruction; each recording retains its measured instruction.",
|
| 1388 |
+
"status": "completed",
|
| 1389 |
+
"episode_id": "task02-09-seed0-formal",
|
| 1390 |
+
"planned_protocol": {
|
| 1391 |
+
"episodes": 1,
|
| 1392 |
+
"seed": 0,
|
| 1393 |
+
"control_frequency_hz": 20,
|
| 1394 |
+
"max_control_steps": 6000,
|
| 1395 |
+
"timeout_s": 28800,
|
| 1396 |
+
"mode": "stepped",
|
| 1397 |
+
"instruction_policy": "modified"
|
| 1398 |
+
},
|
| 1399 |
+
"catalog_id": "libero/libero-10-09",
|
| 1400 |
+
"native_identity": {
|
| 1401 |
+
"suite_name": "libero_10",
|
| 1402 |
+
"task_id": 8,
|
| 1403 |
+
"init_state_index": 0,
|
| 1404 |
+
"max_steps": 6000
|
| 1405 |
+
},
|
| 1406 |
+
"instruction_revision": {
|
| 1407 |
+
"native_instruction": "put both moka pots on the stove",
|
| 1408 |
+
"instruction": "put both moka pots on the stove and turn the stove on",
|
| 1409 |
+
"native_instruction_kind": "source",
|
| 1410 |
+
"diff": [
|
| 1411 |
+
{
|
| 1412 |
+
"op": "equal",
|
| 1413 |
+
"original": "put both moka pots on the stove",
|
| 1414 |
+
"modified": "put both moka pots on the stove"
|
| 1415 |
+
},
|
| 1416 |
+
{
|
| 1417 |
+
"op": "insert",
|
| 1418 |
+
"original": "",
|
| 1419 |
+
"modified": " and turn the stove on"
|
| 1420 |
+
}
|
| 1421 |
+
],
|
| 1422 |
+
"changed": true,
|
| 1423 |
+
"preserve_whitespace": true,
|
| 1424 |
+
"label": "Modified instruction",
|
| 1425 |
+
"source_url": "https://huggingface.co/datasets/RLE-Bench/codex-benchmark/resolve/main/catalog/task02/09-libero-10-09/task.yaml",
|
| 1426 |
+
"native_predicates_unchanged": true,
|
| 1427 |
+
"evaluation_status": "Evaluated with this modified instruction.",
|
| 1428 |
+
"review_reason": "Authorized full current-source LIBERO comparison refresh."
|
| 1429 |
+
},
|
| 1430 |
+
"run_status": "finished",
|
| 1431 |
+
"attempt_history": [],
|
| 1432 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1433 |
+
},
|
| 1434 |
+
{
|
| 1435 |
+
"key": "task02/10",
|
| 1436 |
+
"family": "task02",
|
| 1437 |
+
"slot": "10",
|
| 1438 |
+
"native_id": "libero/mug-into-microwave",
|
| 1439 |
+
"title": "LIBERO-10-10",
|
| 1440 |
+
"catalog_instruction": "put the yellow and white mug in the microwave and close it",
|
| 1441 |
+
"native_instruction": "put the yellow and white mug in the microwave and close it",
|
| 1442 |
+
"instruction_source": "runtime native task.instruction",
|
| 1443 |
+
"status": "completed",
|
| 1444 |
+
"episode_id": "task02-10-seed0-formal",
|
| 1445 |
+
"planned_protocol": {
|
| 1446 |
+
"episodes": 1,
|
| 1447 |
+
"seed": 0,
|
| 1448 |
+
"control_frequency_hz": 20,
|
| 1449 |
+
"max_control_steps": 6000,
|
| 1450 |
+
"timeout_s": 28800,
|
| 1451 |
+
"mode": "stepped",
|
| 1452 |
+
"instruction_policy": "original_native"
|
| 1453 |
+
},
|
| 1454 |
+
"catalog_id": "libero/libero-10-10",
|
| 1455 |
+
"native_identity": {
|
| 1456 |
+
"suite_name": "libero_10",
|
| 1457 |
+
"task_id": 9,
|
| 1458 |
+
"init_state_index": 0,
|
| 1459 |
+
"max_steps": 6000
|
| 1460 |
+
},
|
| 1461 |
+
"run_status": "finished",
|
| 1462 |
+
"attempt_history": [],
|
| 1463 |
+
"status_note": "Evaluation and native evidence complete."
|
| 1464 |
+
},
|
| 1465 |
{
|
| 1466 |
"key": "task04/01",
|
| 1467 |
"family": "task04",
|