Fury / query_llm /README.md
rasa2's picture
Update Fury broker Space
661b502 verified
|
Raw
History Blame Contribute Delete
1.73 kB

minimal-RITS-query

A tiny dependency-free Python client for calling a RITS-hosted chat model.

The script reads .env from the repository root and uses:

  • LLM_API_KEY
  • LLM_BASE_URL, for example https://inference-3scale-apicast-production.apps.rits.fmaas.res.ibm.com/
  • LLM_MODEL

Usage

Call the default model from .env:

python3 rits_query.py "Explain retrieval augmented generation in one sentence."

Call a specific model. The script uses LLM_BASE_URL from .env, derives the RITS route from the model name, and appends /v1:

python3 rits_query.py --model ibm/granite-4.0-micro "Write a short haiku about logs."

Read the prompt from a file. --gpus is accepted for job compatibility and defaults to 1 if you do not pass it:

python3 rits_query.py \
  --model meta-llama/llama-3-3-70b-instruct \
  --gpus 1 \
  --input-file /tmp/hf_fury_job_50f6mo8l/prompt.txt

You can also override the endpoint and key:

python3 rits_query.py \
  --base-url "https://your-rits-host/" \
  --api-key "$LLM_API_KEY" \
  --model ibm/granite-4.0-micro \
  "Hello"

If you do not pass --model, the script uses LLM_MODEL from .env. If LLM_BASE_URL is the RITS host root, the script builds a route from the final model name by replacing dots with hyphens. For example, model ibm/granite-4.0-micro becomes: https://inference-3scale-apicast-production.apps.rits.fmaas.res.ibm.com/granite-4-0-micro/v1. If LLM_BASE_URL already ends in /some-model/v1, the script replaces that model route with the requested model. If that model route is not deployed on the configured RITS host, RITS will return 404; in that case pass the exact model base URL with --base-url.