| 2026-07-23:21:58:26 WARNING [config.evaluate_config:287] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT. |
| 2026-07-23:21:58:33 INFO [_cli.run:388] Selected Tasks: ['mmlu_pro', 'gpqa_diamond_zeroshot', 'minerva_math500', 'ifeval', 'gsm8k_cot_zeroshot'] |
| 2026-07-23:21:58:35 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234 |
| 2026-07-23:21:58:35 WARNING [evaluator:226] generation_kwargs: {'max_gen_toks': 1280} specified through cli, these settings will update set parameters in yaml tasks. Ensure 'do_sample=True' for non-greedy decoding! |
| 2026-07-23:21:58:35 INFO [evaluator:239] Initializing local-chat-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8399/v1/chat/completions', 'num_concurrent': 32, 'tokenized_requests': False, 'max_retries': 3} |
| 2026-07-23:21:58:35 INFO [models.api_models:179] Using max length 2048 - 1 |
| 2026-07-23:21:58:35 INFO [models.api_models:200] Using tokenizer None |
|
Generating test split: 0%| | 0/12032 [00:00<?, ? examples/s]
Generating test split: 100%|██████████| 12032/12032 [00:00<00:00, 130527.55 examples/s] |
|
Generating validation split: 0%| | 0/70 [00:00<?, ? examples/s]
Generating validation split: 100%|██████████| 70/70 [00:00<00:00, 16487.04 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 5882.03 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 57671.90 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 60575.64 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10353.75 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55956.35 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59606.95 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10440.27 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56297.33 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59894.52 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10421.74 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56485.03 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 60117.06 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10421.00 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55947.18 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59463.25 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10308.67 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56257.74 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59577.96 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10358.86 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55622.74 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59308.25 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10358.86 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55755.30 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59239.33 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10346.09 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55855.75 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59560.52 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10676.02 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56148.11 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59580.91 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10722.81 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55920.01 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59653.31 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10736.14 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56088.26 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59599.63 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10846.40 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56265.18 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59696.35 examples/s] |
|
Filter: 0%| | 0/70 [00:00<?, ? examples/s]
Filter: 100%|██████████| 70/70 [00:00<00:00, 10649.69 examples/s] |
|
Filter: 0%| | 0/12032 [00:00<?, ? examples/s]
Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55632.12 examples/s]
Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59335.94 examples/s] |
|
Generating train split: 0%| | 0/198 [00:00<?, ? examples/s]
Generating train split: 100%|██████████| 198/198 [00:00<00:00, 2754.71 examples/s] |
|
Map: 0%| | 0/198 [00:00<?, ? examples/s]
Map: 100%|██████████| 198/198 [00:00<00:00, 1745.18 examples/s] |
| 2026-07-23:21:58:57 INFO [evaluator_utils:446] Selected tasks: |
| 2026-07-23:21:58:57 INFO [evaluator_utils:462] Group: mmlu_pro |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_biology (mmlu_pro/mmlu_pro_biology.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_business (mmlu_pro/mmlu_pro_business.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_chemistry (mmlu_pro/mmlu_pro_chemistry.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_computer_science (mmlu_pro/mmlu_pro_computer_science.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_economics (mmlu_pro/mmlu_pro_economics.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_engineering (mmlu_pro/mmlu_pro_engineering.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_health (mmlu_pro/mmlu_pro_health.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_history (mmlu_pro/mmlu_pro_history.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_law (mmlu_pro/mmlu_pro_law.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_math (mmlu_pro/mmlu_pro_math.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_other (mmlu_pro/mmlu_pro_other.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_philosophy (mmlu_pro/mmlu_pro_philosophy.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_physics (mmlu_pro/mmlu_pro_physics.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_psychology (mmlu_pro/mmlu_pro_psychology.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: gpqa_diamond_zeroshot (gpqa/zeroshot/gpqa_diamond_zeroshot.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: gsm8k_cot_zeroshot (gsm8k/gsm8k-cot-zeroshot.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: ifeval (ifeval/ifeval.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: minerva_math500 (minerva_math/minerva_math500.yaml) |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_biology: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_business: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_chemistry: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_computer_science: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_economics: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_engineering: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_health: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_history: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_law: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_math: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_other: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_philosophy: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_physics: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_psychology: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0} |
| 2026-07-23:21:58:57 INFO [evaluator:314] minerva_math500: Using gen_kwargs: {'until': ['Problem:'], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280} |
| 2026-07-23:21:58:57 INFO [evaluator:314] ifeval: Using gen_kwargs: {'until': [], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280} |
| 2026-07-23:21:58:57 INFO [evaluator:314] gsm8k_cot_zeroshot: Using gen_kwargs: {'until': ['Q:', '</s>', '<|im_end|>'], 'do_sample': False, 'max_gen_toks': 1280} |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_biology on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 2828.64it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_business on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3239.60it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_chemistry on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3167.66it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_computer_science on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3093.60it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_economics on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3127.74it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_engineering on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3112.66it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_health on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3215.75it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_history on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 2917.78it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_law on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3058.63it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_math on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3091.32it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_other on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3156.70it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_philosophy on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3223.66it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_physics on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3233.60it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_psychology on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 3131.01it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for gpqa_diamond_zeroshot on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 1608.49it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for minerva_math500 on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 398.62it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for ifeval on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 52103.16it/s] |
| 2026-07-23:21:58:57 INFO [api.task:312] Building contexts for gsm8k_cot_zeroshot on rank 0... |
|
0%| | 0/10 [00:00<?, ?it/s]
100%|██████████| 10/10 [00:00<00:00, 1775.59it/s] |
| 2026-07-23:21:58:57 INFO [evaluator:585] Running generate_until requests |
| 2026-07-23:21:58:57 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked. |
|
Requesting API: 0%| | 0/140 [00:00<?, ?it/s]2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6509', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5885', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6359', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6022', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6034', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6371', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5764', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6256', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5926', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5268', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5319', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5445', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5367', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5370', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5001', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5031', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5037', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4707', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4804', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4927', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5300', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4861', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5568', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5151', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4932', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4766', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5400', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5687', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5578', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5815', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5553', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5835', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5411', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5722', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4988', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5125', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4964', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5185', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5111', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4950', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5085', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5181', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4960', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3637', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3659', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3789', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3779', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5566', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4088', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4005', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3847', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4810', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4911', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5039', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4918', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4908', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4813', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4963', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5242', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4904', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9508', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9440', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9581', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9622', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9459', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '10496', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9415', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9681', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9441', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '11029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7108', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7138', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8636', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7438', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7189', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8474', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7222', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5362', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5451', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5201', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5480', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5398', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5641', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5638', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5610', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4412', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4453', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4354', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3933', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4535', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4594', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4854', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4470', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4204', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4145', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4756', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4276', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4146', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4172', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4577', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4521', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5055', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4660', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5063', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5027', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4735', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6095', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4772', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5197', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4808', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6049', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5628', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5244', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5560', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5251', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5596', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5977', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5647', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6509', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5885', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6359', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6022', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6034', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6371', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5764', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6256', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5926', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5268', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5319', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5445', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5367', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5370', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5001', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5031', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5037', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4707', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4804', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4927', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5300', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4861', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5568', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5151', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4932', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4766', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5400', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5687', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5578', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5815', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5553', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5835', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5411', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5722', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4988', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5125', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4964', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5185', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5111', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4950', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5085', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5181', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4960', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3637', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3659', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3789', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3779', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5566', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4088', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4005', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3847', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4810', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4911', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5039', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4918', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4908', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4813', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4963', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5242', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4904', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9508', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9440', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9581', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9622', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9459', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '10496', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9415', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9681', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9441', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '11029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7108', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7138', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8636', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7438', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7189', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8474', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7222', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5362', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5451', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5201', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5480', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5398', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5641', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5638', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5610', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4412', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4453', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4354', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3933', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4535', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4594', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4854', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4470', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4204', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4145', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4756', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4276', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4146', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4172', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4577', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4521', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5055', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4660', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5063', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5027', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4735', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6095', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4772', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5197', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4808', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6049', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5628', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5244', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5560', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5251', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5596', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5977', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5647', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
| 2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2 |
| 2026-07-23:21:58:59 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying... |
| 2026-07-23:21:58:59 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying. |
|
Requesting API: 0%| | 0/140 [00:02<?, ?it/s] |
| 2026-07-23:21:58:59 ERROR [models.api_models:559] Exception:ServerDisconnectedError('Server disconnected'), (no outputs), retrying. |
| Traceback (most recent call last): |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/bin/lm_eval", line 10, in <module> |
| sys.exit(cli_evaluate()) |
| ^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/__main__.py", line 10, in cli_evaluate |
| parser.execute(args) |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/_cli/harness.py", line 60, in execute |
| args.func(args) |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/_cli/run.py", line 391, in _execute |
| results = simple_evaluate( |
| ^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/utils.py", line 575, in _wrapper |
| return fn(*args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/evaluator.py", line 358, in simple_evaluate |
| results = evaluate( |
| ^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/utils.py", line 575, in _wrapper |
| return fn(*args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/evaluator.py", line 596, in evaluate |
| resps = getattr(lm, reqtype)(cloned_reqs) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 825, in generate_until |
| asyncio.run( |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/runners.py", line 195, in run |
| return runner.run(main) |
| ^^^^^^^^^^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/runners.py", line 118, in run |
| return self._loop.run_until_complete(task) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/base_events.py", line 691, in run_until_complete |
| return future.result() |
| ^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 631, in get_batched_requests |
| return await tqdm_asyncio.gather(*tasks, desc="Requesting API") |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py", line 79, in gather |
| res = [await f for f in cls.as_completed(ifs, loop=loop, timeout=timeout, |
| ^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/tasks.py", line 631, in _wait_for_one |
| return f.result() # May raise f.exception(). |
| ^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py", line 76, in wrap_awaitable |
| return i, await f |
| ^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 193, in async_wrapped |
| return await copy(fn, *args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 112, in __call__ |
| do = await self.iter(retry_state=retry_state) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 157, in iter |
| result = await action(retry_state) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/_utils.py", line 111, in inner |
| return call(*args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 413, in exc_check |
| raise retry_exc.reraise() |
| ^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 184, in reraise |
| raise self.last_attempt.result() |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 449, in result |
| return self.__get_result() |
| ^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result |
| raise self._exception |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 116, in __call__ |
| result = await fn(*args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 560, in amodel_call |
| raise e |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 528, in amodel_call |
| response.raise_for_status() |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client_reqrep.py", line 655, in raise_for_status |
| raise ClientResponseError( |
| aiohttp.client_exceptions.ClientResponseError: 400, message='Bad Request', url='http://127.0.0.1:8399/v1/chat/completions' |
| Task exception was never retrieved |
| future: <Task finished name='Task-159' coro=<tqdm_asyncio.gather.<locals>.wrap_awaitable() done, defined at /home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py:75> exception=ServerDisconnectedError('Server disconnected')> |
| Traceback (most recent call last): |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py", line 76, in wrap_awaitable |
| return i, await f |
| ^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 193, in async_wrapped |
| return await copy(fn, *args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 112, in __call__ |
| do = await self.iter(retry_state=retry_state) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 157, in iter |
| result = await action(retry_state) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/_utils.py", line 111, in inner |
| return call(*args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 413, in exc_check |
| raise retry_exc.reraise() |
| ^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 184, in reraise |
| raise self.last_attempt.result() |
| ^^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 449, in result |
| return self.__get_result() |
| ^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result |
| raise self._exception |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 116, in __call__ |
| result = await fn(*args, **kwargs) |
| ^^^^^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 560, in amodel_call |
| raise e |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 516, in amodel_call |
| async with session.post( |
| ^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client.py", line 1683, in __aenter__ |
| self._resp: _RetType_co = await self._coro |
| ^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client.py", line 856, in _request |
| resp = await handler(req) |
| ^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client.py", line 834, in _connect_and_send_request |
| await resp.start(conn) |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client_reqrep.py", line 558, in start |
| message, payload = await protocol.read() # type: ignore[union-attr] |
| ^^^^^^^^^^^^^^^^^^^^^ |
| File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/streams.py", line 705, in read |
| await self._waiter |
| aiohttp.client_exceptions.ServerDisconnectedError: Server disconnected |
|
|