openweb / README.md
Leon4gr45's picture
Force redeploy with binary path fix
2a0d3d8 verified
|
Raw
History Blame Contribute Delete
2.84 kB
metadata
title: OpenResearcher
emoji: 🏃
colorFrom: purple
colorTo: green
sdk: docker
app_port: 7860
pinned: false

OpenResearcher DeepSearch agent. Inference runs locally in this Space via ZeroGPU by default; a cloud OpenAI-compatible endpoint is used only as a backup if local inference fails. An "Inference Mode" switch in the sidebar (Auto / Local only / Cloud only) lets you override this per request. The agent also plans an explicit research goal (with a 1-5 vagueness rating) up front and stops as soon as that goal is met, instead of always running to the round limit.

Configuration (Space secrets, all optional)

Secret Description
MODEL_NAME Local model to run via ZeroGPU. Defaults to OpenResearcher/Nemotron-3-Nano-30B-A3B.
OPENAI_BASE_URL Cloud backup, used only if local inference errors out. Base URL of an OpenAI-compatible API; /chat/completions is appended automatically.
OPENAI_MODEL Model name for the cloud backup. Defaults to MODEL_NAME.
OPENAI_API_KEY Bearer token for the cloud backup.
JINA_MCP_BASE_URL Self-hosted Jina MCP server. Used exclusively for web search (search_web tool) and as the primary scraping backend (read_url tool). Defaults to https://leon4gr45-jina.hf.space. Its own API key is configured on that server, so none is needed here.
MAX_CONTEXT_TOKENS Fallback input budget (approx. tokens), only used if the context length can't be auto-detected. Defaults to 100000.

The ReAct loop's conversation history is kept under a per-model token budget automatically. That budget is auto-detected once per run: for the local model, straight from its loaded config (max_position_embeddings etc.); for the cloud backup, by querying {OPENAI_BASE_URL}/models and reading off whichever context-length field the server reports (e.g. vLLM's max_model_len, OpenRouter's context_length - the plain OpenAI API itself doesn't expose this, so falls back to MAX_CONTEXT_TOKENS there). In Auto mode, the tighter of the two detected limits is used, since either backend might end up serving the request. Older rounds get truncated (recent rounds stay verbatim) before each generation call, and again as an emergency retry if a "context length exceeded" error still slips through.

Web search uses only Jina's search_web tool. (A self-hosted SearXNG instance was tried, but Hugging Face's abuse-detector auto-paused that Space over a process-name/cmdline match on "searxng" — not viable here.) Page fetching tries Jina's read_url first, then falls back to a direct HTTP fetch, then to Serper's scrape API if a key is provided in the UI.

This Space requires ZeroGPU (or another GPU) hardware to run local model inference.

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference