Biomanus
Biomanus is a graph-guided extension of the Biomni A1 biomedical agent. It adds MCP server generation, MCP ToolGraph construction, GraphRAG-style tool routing, dynamic MCP registration, and benchmark runners for BioAgent-Bench and LAB-Bench.
This README explains how to install the environment, prepare MCP servers, build the MCP graph, and run the experiments.
1. Repository Layout
biomni/ Core agent, tools, graph router, retriever, and utilities
biomni_env/ Conda environment files and setup scripts
demo_web/backend/tool2mcp.py Utility for converting bioinformatics tools into MCP servers
build_generated_mcp_graph.py Build MCP server catalogs and graph artifacts
run_bioagent_bench.py BioAgent-Bench runner
evaluate_bioagent_bench.py BioAgent-Bench evaluator
experiments/lab_bench/scripts/ LAB-Bench runners and graph utilities
graph_outputs/ Example prebuilt graph artifacts
Large runtime artifacts, downloaded data lakes, benchmark outputs, node modules, and Git history are intentionally not included in this upload package.
2. Install the Environment
Create and activate the Biomni environment:
cd /path/to/Biomanus
conda env create -f biomni_env/environment.yml
conda activate biomni_e1
For a smaller environment, use:
conda env create -f biomni_env/fixed_env.yml
conda activate biomni_e1
Install extra packages used by MCP generation and benchmark execution:
pip install fastmcp mcp openai anthropic python-dotenv pymupdf pipreqs
pip install langchain langchain-openai langchain-anthropic langgraph
Because this upload package is source-first and does not require a wheel install, run commands from the repository root or export:
export PYTHONPATH="$PWD:$PYTHONPATH"
export BIOMNI_EXECUTION_ENV_PREFIX="$CONDA_PREFIX"
export BIOMNI_EXECUTION_PYTHON="$CONDA_PREFIX/bin/python"
If your experiments need command-line bioinformatics tools, install them into the same conda environment. The full upstream Biomni setup script is available at:
bash biomni_env/setup.sh
This full setup can take many hours and requires substantial disk space.
3. Configure the LLM Backend
Biomanus uses Biomni's Custom OpenAI-compatible backend in the main experiments. For DeepSeek:
export DEEPSEEK_API_KEY="your_key"
export DEEPSEEK_BASE_URL="https://api.deepseek.com/v1"
export DEEPSEEK_MODEL_NAME="deepseek-chat"
export BIOMNI_SOURCE="Custom"
export BIOMNI_LLM="$DEEPSEEK_MODEL_NAME"
export BIOMNI_CUSTOM_BASE_URL="$DEEPSEEK_BASE_URL"
export BIOMNI_CUSTOM_API_KEY="$DEEPSEEK_API_KEY"
export BIOMNI_LLM_PROVIDER="deepseek"
For another OpenAI-compatible endpoint, keep BIOMNI_SOURCE=Custom and replace the base URL, model name, and API key.
4. Prepare MCP Servers
Graph-based execution requires local MCP server directories. Each generated server should follow this structure:
demo_web/backend/data/mcp_generated/
mcp_<server_name>/
app/
<server_name>_server.py
<server_name>_shim_server.py # optional but recommended
Dockerfile
docker-compose.yml
requirements.txt
environment.yaml
Create the MCP directories:
mkdir -p demo_web/backend/data/mcp_generated
mkdir -p demo_web/backend/data/merged_prefer_help_txt
Option A: Copy Existing MCP Servers
If you already generated MCP servers elsewhere, copy the mcp_* directories into:
demo_web/backend/data/mcp_generated/
Then patch generated servers so they expose stdio-compatible MCP entrypoints:
python patch_raw_mcp_servers.py --mcp-root demo_web/backend/data/mcp_generated
python patch_generated_mcp_shims.py --mcp-root demo_web/backend/data/mcp_generated
Option B: Generate One MCP Server from a Tool Manual
Use tool2mcp.py to convert a bioinformatics command-line tool and its help/manual text into an MCP server.
Example using a local help text:
python demo_web/backend/tool2mcp.py \
--name kallisto \
--manual docs/kallisto_help.txt \
--output_location demo_web/backend/data/mcp_generated \
--llm_provider deepseek
Example using the installed tool's help command:
python demo_web/backend/tool2mcp.py \
--name fastqc \
--manual=--help \
--run_help_command True \
--output_location demo_web/backend/data/mcp_generated \
--llm_provider deepseek
The converter calls an LLM, so the corresponding API key must be configured. Supported providers include deepseek, openai, claude, azure, and gemini.
5. Build the MCP Graph
Once MCP servers exist under demo_web/backend/data/mcp_generated, build a graph over all servers:
python build_generated_mcp_graph.py \
--mcp-root demo_web/backend/data/mcp_generated \
--help-root demo_web/backend/data/merged_prefer_help_txt \
--output-root graph_outputs \
--preset all \
--python-cmd "$BIOMNI_EXECUTION_PYTHON"
The output directory will look like:
graph_outputs/mcp_generated_graph_all_<timestamp>/
server_catalog.json
graph_nodes.json
graph_edges.json
server_semantics.json
graph_summary.json
For BioAgent-Bench, you can build a smaller graph containing high-signal benchmark servers:
python build_generated_mcp_graph.py \
--mcp-root demo_web/backend/data/mcp_generated \
--help-root demo_web/backend/data/merged_prefer_help_txt \
--output-root graph_outputs \
--preset bioagent-bench \
--python-cmd "$BIOMNI_EXECUTION_PYTHON"
This upload package includes an example graph under graph_outputs/. Use it for inspection or as a template. For executable experiments on a new machine, rebuild the graph after placing MCP servers locally so server_catalog.json contains valid local commands.
6. Run BioAgent-Bench
First download or clone BioAgent-Bench and set the metadata/data paths:
export BIOAGENT_BENCH_DATASET_ROOT="/path/to/bioagent-bench/dataset"
export BIOAGENT_BENCH_METADATA="/path/to/bioagent-bench/src/task_metadata.json"
Run a single task:
python run_bioagent_bench.py \
--task transcript-quant \
--dataset-root "$BIOAGENT_BENCH_DATASET_ROOT" \
--metadata "$BIOAGENT_BENCH_METADATA" \
--mcp-graph graph_outputs/mcp_generated_graph_benchmark_<timestamp> \
--executable-mcp-only \
--output-root experiments/bioagent-bench-runs \
--execution-env-prefix "$CONDA_PREFIX"
Run all supported tasks:
python run_bioagent_bench.py \
--all \
--dataset-root "$BIOAGENT_BENCH_DATASET_ROOT" \
--metadata "$BIOAGENT_BENCH_METADATA" \
--mcp-graph graph_outputs/mcp_generated_graph_benchmark_<timestamp> \
--executable-mcp-only \
--output-root experiments/bioagent-bench-runs \
--execution-env-prefix "$CONDA_PREFIX"
Useful ablation switches:
--no-graph-retriever # disable graph-guided routing
--no-mcp # disable MCP registration/execution
--no-use-tool-retriever # disable internal Biomni tool retrieval
Evaluate BioAgent-Bench outputs:
python evaluate_bioagent_bench.py \
--all \
--runs-root experiments/bioagent-bench-runs \
--dataset-root "$BIOAGENT_BENCH_DATASET_ROOT" \
--judge-mode rule \
--output experiments/bioagent-bench-runs/latest_evaluation.json
For LLM-based judging, set the judge provider/key and use --judge-mode llm or --judge-mode both.
7. Run LAB-Bench
First install or clone LAB-Bench and set:
export LAB_BENCH_ROOT="/path/to/LAB-Bench"
Run a small debug pass:
python experiments/lab_bench/scripts/run_labbench_with_hypobioos.py \
--evals DbQA SeqQA \
--splits test \
--debug \
--graph-dir graph_outputs/mcp_generated_graph_all_<timestamp> \
--output-root experiments/lab_bench/results/debug \
--agent-root experiments/lab_bench/agent_runtime/debug \
--compact-results-path experiments/lab_bench/results/debug/results.jsonl \
--reasoning-log-path experiments/lab_bench/results/debug/reasoning.log \
--compact-output-only
Run the main DbQA/SeqQA protocol:
python experiments/lab_bench/scripts/run_labbench_with_hypobioos.py \
--evals DbQA SeqQA \
--splits test \
--dev-size 45 \
--test-size 315 \
--seed 20260514 \
--graph-dir graph_outputs/mcp_generated_graph_all_<timestamp> \
--mcp-server-top-k 8 \
--mcp-tool-top-k 12 \
--mcp-server-candidate-pool 24 \
--output-root experiments/lab_bench/results/main \
--agent-root experiments/lab_bench/agent_runtime/main \
--compact-results-path experiments/lab_bench/results/main/results.jsonl \
--reasoning-log-path experiments/lab_bench/results/main/reasoning.log \
--compact-output-only
For parallel sharded runs, call the same runner repeatedly with --shard-index and --shard-count, writing all shards to the same compact JSONL with --skip-existing-results.
8. Recommended Reproduction Order
- Create and activate the conda environment.
- Configure the LLM backend.
- Generate or copy MCP servers into
demo_web/backend/data/mcp_generated. - Patch MCP shims.
- Build the MCP graph.
- Run a single BioAgent-Bench task as a smoke test.
- Run LAB-Bench in
--debugmode. - Launch full benchmark runs and evaluate.
9. Common Issues
MCP root does not exist
Create demo_web/backend/data/mcp_generated and copy or generate mcp_* server directories.
Help root does not exist
Create an empty help directory:
mkdir -p demo_web/backend/data/merged_prefer_help_txt
server_catalog.json points to old paths
Rebuild the graph on the current machine after placing MCP servers locally. The graph stores executable command paths.
Missing API key
Export DEEPSEEK_API_KEY, OPENAI_API_KEY, ANTHROPIC_API_KEY, or the matching provider key before running agents or MCP generation.
LAB-Bench import errors
Set LAB_BENCH_ROOT to the local LAB-Bench checkout and install LAB-Bench dependencies in the active environment.
BioAgent-Bench cannot find metadata
Set BIOAGENT_BENCH_METADATA and BIOAGENT_BENCH_DATASET_ROOT, or pass --metadata and --dataset-root explicitly.