fix: reduce max_tokens and add retries to handle Groq TPM rate limits gracefully e3fee95 hamba-ho commited on 17 days ago
fix: switch default model to llama-3.1-8b-instant to bypass 70b daily quota limit 1fed431 hamba-ho commited on 17 days ago
fix: reduce stop sequences to 4 to comply with Groq API limits ac8779d hamba-ho commited on 17 days ago
fix: bind stop sequences to llm to prevent hallucination loop and rate limit 429 errors dacd14f hamba-ho commited on 17 days ago