hbfreed's picture
Add files using upload-large-folder tool
c7b14e3 verified
Raw
History Blame Contribute Delete
3.74 kB
2026-07-23:22:19:14 WARNING [config.evaluate_config:287] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2026-07-23:22:19:20 INFO [_cli.run:388] Selected Tasks: ['humaneval', 'mbpp']
2026-07-23:22:19:22 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
2026-07-23:22:19:22 INFO [evaluator:239] Initializing local-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8399/v1/completions', 'tokenizer': 'outputs/qwen35_pruned_keep25', 'num_concurrent': 32, 'tokenized_requests': False, 'max_retries': 3}
2026-07-23:22:19:22 INFO [models.openai_completions:42] Remote tokenizer not supported. Using huggingface tokenizer backend.
2026-07-23:22:19:22 INFO [models.api_models:179] Using max length 2048 - 1
2026-07-23:22:19:22 INFO [models.api_models:200] Using tokenizer huggingface
2026-07-23:22:19:30 INFO [evaluator_utils:446] Selected tasks:
2026-07-23:22:19:30 INFO [evaluator_utils:480] Task: humaneval (humaneval/humaneval.yaml)
2026-07-23:22:19:30 INFO [evaluator_utils:480] Task: mbpp (mbpp/mbpp.yaml)
2026-07-23:22:19:30 INFO [evaluator:314] humaneval: Using gen_kwargs: {'until': ['\nclass', '\ndef', '\n#', '\nif', '\nprint'], 'max_gen_toks': 1024, 'do_sample': False}
2026-07-23:22:19:30 INFO [evaluator:314] mbpp: Using gen_kwargs: {'until': ['[DONE]'], 'do_sample': False}
2026-07-23:22:19:30 INFO [api.task:312] Building contexts for humaneval on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10/10 [00:00<00:00, 669.16it/s]
2026-07-23:22:19:30 INFO [api.task:312] Building contexts for mbpp on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10/10 [00:00<00:00, 135.22it/s]
2026-07-23:22:19:30 INFO [evaluator:585] Running generate_until requests
2026-07-23:22:19:30 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked.
Requesting API: 0%| | 0/10 [00:00<?, ?it/s] Requesting API: 10%|β–ˆ | 1/10 [01:31<13:45, 91.77s/it] Requesting API: 20%|β–ˆβ–ˆ | 2/10 [01:31<05:03, 37.90s/it] Requesting API: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10/10 [01:31<00:00, 9.20s/it]
Requesting API: 0%| | 0/10 [00:00<?, ?it/s] Requesting API: 10%|β–ˆ | 1/10 [00:03<00:28, 3.16s/it] Requesting API: 20%|β–ˆβ–ˆ | 2/10 [00:03<00:11, 1.41s/it] Requesting API: 50%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 5/10 [00:05<00:04, 1.04it/s] Requesting API: 60%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 6/10 [00:05<00:03, 1.31it/s] Requesting API: 90%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ | 9/10 [00:06<00:00, 2.43it/s] Requesting API: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10/10 [00:23<00:00, 3.79s/it] Requesting API: 100%|β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ| 10/10 [00:23<00:00, 2.35s/it]
2026-07-23:22:21:30 INFO [loggers.evaluation_tracker:247] Saving results aggregated
2026-07-23:22:21:30 INFO [loggers.evaluation_tracker:119] Saving per-task samples to outputs/evals/protocolF_gate3/student/*.jsonl
local-completions ({'model': 'student', 'base_url': 'http://127.0.0.1:8399/v1/completions', 'tokenizer': 'outputs/qwen35_pruned_keep25', 'num_concurrent': 32, 'tokenized_requests': False, 'max_retries': 3}), gen_kwargs: ({}), limit: 10.0, num_fewshot: None, batch_size: 1
| Tasks |Version| Filter |n-shot| Metric | |Value| |Stderr|
|---------|------:|-----------|-----:|---------|---|----:|---|-----:|
|humaneval| 1|create_test| 0|pass@1 |↑ | 0|Β± | 0|
|mbpp | 1|none | 3|pass_at_1|↑ | 0|Β± | 0|