hbfreed's picture
Add files using upload-large-folder tool
8717f59 verified
Raw
History Blame Contribute Delete
379 kB
2026-07-23:21:58:26 WARNING [config.evaluate_config:287] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2026-07-23:21:58:33 INFO [_cli.run:388] Selected Tasks: ['mmlu_pro', 'gpqa_diamond_zeroshot', 'minerva_math500', 'ifeval', 'gsm8k_cot_zeroshot']
2026-07-23:21:58:35 INFO [evaluator:214] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
2026-07-23:21:58:35 WARNING [evaluator:226] generation_kwargs: {'max_gen_toks': 1280} specified through cli, these settings will update set parameters in yaml tasks. Ensure 'do_sample=True' for non-greedy decoding!
2026-07-23:21:58:35 INFO [evaluator:239] Initializing local-chat-completions model, with arguments: {'model': 'student', 'base_url': 'http://127.0.0.1:8399/v1/chat/completions', 'num_concurrent': 32, 'tokenized_requests': False, 'max_retries': 3}
2026-07-23:21:58:35 INFO [models.api_models:179] Using max length 2048 - 1
2026-07-23:21:58:35 INFO [models.api_models:200] Using tokenizer None
Generating test split: 0%| | 0/12032 [00:00<?, ? examples/s] Generating test split: 100%|██████████| 12032/12032 [00:00<00:00, 130527.55 examples/s]
Generating validation split: 0%| | 0/70 [00:00<?, ? examples/s] Generating validation split: 100%|██████████| 70/70 [00:00<00:00, 16487.04 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 5882.03 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 57671.90 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 60575.64 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10353.75 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55956.35 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59606.95 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10440.27 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56297.33 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59894.52 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10421.74 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56485.03 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 60117.06 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10421.00 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55947.18 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59463.25 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10308.67 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56257.74 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59577.96 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10358.86 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55622.74 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59308.25 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10358.86 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55755.30 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59239.33 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10346.09 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55855.75 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59560.52 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10676.02 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56148.11 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59580.91 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10722.81 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55920.01 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59653.31 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10736.14 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56088.26 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59599.63 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10846.40 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 56265.18 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59696.35 examples/s]
Filter: 0%| | 0/70 [00:00<?, ? examples/s] Filter: 100%|██████████| 70/70 [00:00<00:00, 10649.69 examples/s]
Filter: 0%| | 0/12032 [00:00<?, ? examples/s] Filter: 58%|█████▊ | 7000/12032 [00:00<00:00, 55632.12 examples/s] Filter: 100%|██████████| 12032/12032 [00:00<00:00, 59335.94 examples/s]
Generating train split: 0%| | 0/198 [00:00<?, ? examples/s] Generating train split: 100%|██████████| 198/198 [00:00<00:00, 2754.71 examples/s]
Map: 0%| | 0/198 [00:00<?, ? examples/s] Map: 100%|██████████| 198/198 [00:00<00:00, 1745.18 examples/s]
2026-07-23:21:58:57 INFO [evaluator_utils:446] Selected tasks:
2026-07-23:21:58:57 INFO [evaluator_utils:462] Group: mmlu_pro
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_biology (mmlu_pro/mmlu_pro_biology.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_business (mmlu_pro/mmlu_pro_business.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_chemistry (mmlu_pro/mmlu_pro_chemistry.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_computer_science (mmlu_pro/mmlu_pro_computer_science.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_economics (mmlu_pro/mmlu_pro_economics.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_engineering (mmlu_pro/mmlu_pro_engineering.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_health (mmlu_pro/mmlu_pro_health.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_history (mmlu_pro/mmlu_pro_history.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_law (mmlu_pro/mmlu_pro_law.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_math (mmlu_pro/mmlu_pro_math.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_other (mmlu_pro/mmlu_pro_other.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_philosophy (mmlu_pro/mmlu_pro_philosophy.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_physics (mmlu_pro/mmlu_pro_physics.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:470] Task: mmlu_pro_psychology (mmlu_pro/mmlu_pro_psychology.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: gpqa_diamond_zeroshot (gpqa/zeroshot/gpqa_diamond_zeroshot.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: gsm8k_cot_zeroshot (gsm8k/gsm8k-cot-zeroshot.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: ifeval (ifeval/ifeval.yaml)
2026-07-23:21:58:57 INFO [evaluator_utils:480] Task: minerva_math500 (minerva_math/minerva_math500.yaml)
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_biology: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_business: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_chemistry: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_computer_science: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_economics: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_engineering: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_health: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_history: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_law: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_math: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_other: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_philosophy: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_physics: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] mmlu_pro_psychology: Using gen_kwargs: {'until': ['Question:'], 'max_gen_toks': 1280, 'do_sample': False, 'temperature': 0.0}
2026-07-23:21:58:57 INFO [evaluator:314] minerva_math500: Using gen_kwargs: {'until': ['Problem:'], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280}
2026-07-23:21:58:57 INFO [evaluator:314] ifeval: Using gen_kwargs: {'until': [], 'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 1280}
2026-07-23:21:58:57 INFO [evaluator:314] gsm8k_cot_zeroshot: Using gen_kwargs: {'until': ['Q:', '</s>', '<|im_end|>'], 'do_sample': False, 'max_gen_toks': 1280}
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_biology on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 2828.64it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_business on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3239.60it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_chemistry on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3167.66it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_computer_science on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3093.60it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_economics on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3127.74it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_engineering on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3112.66it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_health on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3215.75it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_history on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 2917.78it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_law on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3058.63it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_math on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3091.32it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_other on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3156.70it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_philosophy on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3223.66it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_physics on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3233.60it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for mmlu_pro_psychology on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 3131.01it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for gpqa_diamond_zeroshot on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 1608.49it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for minerva_math500 on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 398.62it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for ifeval on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 52103.16it/s]
2026-07-23:21:58:57 INFO [api.task:312] Building contexts for gsm8k_cot_zeroshot on rank 0...
0%| | 0/10 [00:00<?, ?it/s] 100%|██████████| 10/10 [00:00<00:00, 1775.59it/s]
2026-07-23:21:58:57 INFO [evaluator:585] Running generate_until requests
2026-07-23:21:58:57 INFO [models.api_models:747] Tokenized requests are disabled. Context + generation length is not checked.
Requesting API: 0%| | 0/140 [00:00<?, ?it/s]2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6509', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5885', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6359', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6022', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6034', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6371', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5764', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6256', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5926', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5268', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5319', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5445', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5367', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5370', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5001', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5031', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5037', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4707', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4804', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4927', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5300', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4861', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5568', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5151', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4932', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4766', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5400', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5687', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5578', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5815', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5553', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5835', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5411', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5722', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4988', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5125', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4964', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5185', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5111', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4950', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5085', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5181', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4960', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3637', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3659', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3789', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3779', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5566', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4088', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4005', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3847', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4810', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4911', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5039', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4918', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4908', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4813', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4963', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5242', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4904', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9508', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9440', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9581', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9622', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9459', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '10496', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9415', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9681', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9441', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '11029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7108', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7138', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8636', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7438', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7189', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8474', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7222', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5362', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5451', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5201', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5480', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5398', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5641', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5638', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5610', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4412', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4453', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4354', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3933', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4535', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4594', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4854', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4470', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4204', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4145', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4756', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4276', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4146', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4172', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4577', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4521', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5055', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4660', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5063', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5027', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4735', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6095', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4772', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5197', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:56 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4808', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6049', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5628', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5244', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5560', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5251', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5596', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5977', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:57 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:57 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5647', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:57 INFO [models.api_models:607] Retry attempt 1
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6509', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5885', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6359', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6022', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6034', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6371', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5764', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6256', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5926', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5268', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5319', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5445', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5367', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5370', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5001', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5031', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5037', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4707', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4804', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4927', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5300', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4861', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5568', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5151', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4932', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4766', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5400', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5687', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5578', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5815', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5553', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5835', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5411', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5722', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4988', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5125', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4964', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5185', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5111', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4950', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5085', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5181', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4960', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3637', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3659', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3789', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3779', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5566', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4088', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4005', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3847', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4810', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4911', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5039', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4918', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4908', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4813', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4963', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5242', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4904', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9508', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9440', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9581', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9622', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9459', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '10496', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9415', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9681', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '9441', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '11029', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7108', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7138', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8636', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8186', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7438', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7189', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '8474', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7222', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '7212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5362', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5540', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5451', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5201', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5480', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5398', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5641', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:57 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5638', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5610', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4412', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4453', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4354', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '3933', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4535', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4010', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4594', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4854', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4470', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4204', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4145', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4756', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4276', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4146', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4172', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4577', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4521', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5055', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4502', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4660', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5063', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5027', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5047', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4735', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6095', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5658', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4751', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4772', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5197', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '4808', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6049', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '6212', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5628', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5321', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5244', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5560', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5251', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5596', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5977', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:58 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:58 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5647', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
2026-07-23:21:58:58 INFO [models.api_models:607] Retry attempt 2
2026-07-23:21:58:59 WARNING [models.api_models:523] API request failed! Status code: 400, Response text: {"error":{"message":"This model's maximum context length is 2048 tokens. However, you requested 1280 output tokens and your prompt contains at least 769 input tokens, for a total of at least 2049 tokens. Please reduce the length of the input prompt or the number of requested output tokens. (parameter=input_tokens, value=769)","type":"BadRequestError","param":"input_tokens","code":400}}. Retrying...
2026-07-23:21:58:59 ERROR [models.api_models:559] Exception:ClientResponseError(RequestInfo(url=URL('http://127.0.0.1:8399/v1/chat/completions'), method='POST', headers=<CIMultiDictProxy('Host': '127.0.0.1:8399', 'Authorization': 'Bearer ', 'Accept': '*/*', 'Accept-Encoding': 'gzip, deflate', 'User-Agent': 'Python/3.12 aiohttp/3.14.1', 'Content-Length': '5982', 'Content-Type': 'application/json')>, real_url=URL('http://127.0.0.1:8399/v1/chat/completions')), (), status=400, message='Bad Request', headers=<CIMultiDictProxy('Date': 'Fri, 24 Jul 2026 04:58:58 GMT', 'Server': 'uvicorn', 'Content-Length': '388', 'Content-Type': 'application/json')>), (no outputs), retrying.
Requesting API: 0%| | 0/140 [00:02<?, ?it/s]
2026-07-23:21:58:59 ERROR [models.api_models:559] Exception:ServerDisconnectedError('Server disconnected'), (no outputs), retrying.
Traceback (most recent call last):
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/bin/lm_eval", line 10, in <module>
sys.exit(cli_evaluate())
^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/__main__.py", line 10, in cli_evaluate
parser.execute(args)
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/_cli/harness.py", line 60, in execute
args.func(args)
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/_cli/run.py", line 391, in _execute
results = simple_evaluate(
^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/utils.py", line 575, in _wrapper
return fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/evaluator.py", line 358, in simple_evaluate
results = evaluate(
^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/utils.py", line 575, in _wrapper
return fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/evaluator.py", line 596, in evaluate
resps = getattr(lm, reqtype)(cloned_reqs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 825, in generate_until
asyncio.run(
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/runners.py", line 195, in run
return runner.run(main)
^^^^^^^^^^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/runners.py", line 118, in run
return self._loop.run_until_complete(task)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/base_events.py", line 691, in run_until_complete
return future.result()
^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 631, in get_batched_requests
return await tqdm_asyncio.gather(*tasks, desc="Requesting API")
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py", line 79, in gather
res = [await f for f in cls.as_completed(ifs, loop=loop, timeout=timeout,
^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/asyncio/tasks.py", line 631, in _wait_for_one
return f.result() # May raise f.exception().
^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py", line 76, in wrap_awaitable
return i, await f
^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 193, in async_wrapped
return await copy(fn, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 112, in __call__
do = await self.iter(retry_state=retry_state)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 157, in iter
result = await action(retry_state)
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/_utils.py", line 111, in inner
return call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 413, in exc_check
raise retry_exc.reraise()
^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 184, in reraise
raise self.last_attempt.result()
^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 449, in result
return self.__get_result()
^^^^^^^^^^^^^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
raise self._exception
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 116, in __call__
result = await fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 560, in amodel_call
raise e
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 528, in amodel_call
response.raise_for_status()
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client_reqrep.py", line 655, in raise_for_status
raise ClientResponseError(
aiohttp.client_exceptions.ClientResponseError: 400, message='Bad Request', url='http://127.0.0.1:8399/v1/chat/completions'
Task exception was never retrieved
future: <Task finished name='Task-159' coro=<tqdm_asyncio.gather.<locals>.wrap_awaitable() done, defined at /home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py:75> exception=ServerDisconnectedError('Server disconnected')>
Traceback (most recent call last):
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tqdm/asyncio.py", line 76, in wrap_awaitable
return i, await f
^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 193, in async_wrapped
return await copy(fn, *args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 112, in __call__
do = await self.iter(retry_state=retry_state)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 157, in iter
result = await action(retry_state)
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/_utils.py", line 111, in inner
return call(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 413, in exc_check
raise retry_exc.reraise()
^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/__init__.py", line 184, in reraise
raise self.last_attempt.result()
^^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 449, in result
return self.__get_result()
^^^^^^^^^^^^^^^^^^^
File "/home/henry/.local/share/uv/python/cpython-3.12.12-linux-x86_64-gnu/lib/python3.12/concurrent/futures/_base.py", line 401, in __get_result
raise self._exception
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/tenacity/asyncio/__init__.py", line 116, in __call__
result = await fn(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 560, in amodel_call
raise e
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/lm_eval/models/api_models.py", line 516, in amodel_call
async with session.post(
^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client.py", line 1683, in __aenter__
self._resp: _RetType_co = await self._coro
^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client.py", line 856, in _request
resp = await handler(req)
^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client.py", line 834, in _connect_and_send_request
await resp.start(conn)
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/client_reqrep.py", line 558, in start
message, payload = await protocol.read() # type: ignore[union-attr]
^^^^^^^^^^^^^^^^^^^^^
File "/home/henry/Documents/PythonProjects/variable-reap/.venv/lib/python3.12/site-packages/aiohttp/streams.py", line 705, in read
await self._waiter
aiohttp.client_exceptions.ServerDisconnectedError: Server disconnected