| --- |
| title: OpenResearcher |
| emoji: π |
| colorFrom: purple |
| colorTo: green |
| sdk: docker |
| app_port: 7860 |
| pinned: false |
| --- |
| |
| OpenResearcher DeepSearch agent. Inference runs locally in this Space via |
| ZeroGPU by default; a cloud OpenAI-compatible endpoint is used only as a |
| backup if local inference fails. An "Inference Mode" switch in the sidebar |
| (Auto / Local only / Cloud only) lets you override this per request. The |
| agent also plans an explicit research goal (with a 1-5 vagueness rating) up |
| front and stops as soon as that goal is met, instead of always running to |
| the round limit. |
|
|
| ## Configuration (Space secrets, all optional) |
|
|
| | Secret | Description | |
| |---|---| |
| | `MODEL_NAME` | Local model to run via ZeroGPU. Defaults to `OpenResearcher/Nemotron-3-Nano-30B-A3B`. | |
| | `OPENAI_BASE_URL` | Cloud backup, used only if local inference errors out. Base URL of an OpenAI-compatible API; `/chat/completions` is appended automatically. | |
| | `OPENAI_MODEL` | Model name for the cloud backup. Defaults to `MODEL_NAME`. | |
| | `OPENAI_API_KEY` | Bearer token for the cloud backup. | |
| | `JINA_MCP_BASE_URL` | Self-hosted Jina MCP server. Used exclusively for web search (`search_web` tool) and as the primary scraping backend (`read_url` tool). Defaults to `https://leon4gr45-jina.hf.space`. Its own API key is configured on that server, so none is needed here. | |
| | `MAX_CONTEXT_TOKENS` | Fallback input budget (approx. tokens), only used if the context length can't be auto-detected. Defaults to `100000`. | |
|
|
| The ReAct loop's conversation history is kept under a per-model token |
| budget automatically. That budget is auto-detected once per run: for the |
| local model, straight from its loaded config (`max_position_embeddings` |
| etc.); for the cloud backup, by querying `{OPENAI_BASE_URL}/models` and |
| reading off whichever context-length field the server reports (e.g. vLLM's |
| `max_model_len`, OpenRouter's `context_length` - the plain OpenAI API |
| itself doesn't expose this, so falls back to `MAX_CONTEXT_TOKENS` there). |
| In Auto mode, the tighter of the two detected limits is used, since either |
| backend might end up serving the request. Older rounds get truncated |
| (recent rounds stay verbatim) before each generation call, and again as an |
| emergency retry if a "context length exceeded" error still slips through. |
|
|
| Web search uses only Jina's `search_web` tool. (A self-hosted SearXNG |
| instance was tried, but Hugging Face's abuse-detector auto-paused that |
| Space over a process-name/cmdline match on "searxng" β not viable here.) |
| Page fetching tries Jina's `read_url` first, then falls back to a direct |
| HTTP fetch, then to Serper's scrape API if a key is provided in the UI. |
|
|
| This Space requires ZeroGPU (or another GPU) hardware to run local model |
| inference. |
|
|
| Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference |
|
|