Safetensors

EviBack: Reinforcement Learning Search Agents via Evidence-Constrained Teacher Backoff

Xiao Ma*  Zhiquan Hu*  Yi Wei  Chenchen Zhao  Yijun Chen  Jicheng Zhao  Yuming Lii†  Chuang Dai
NEXTAI Research Institute, Chery Group.

*Equal Contribution, i†Corresponding Author

Case 1
Case 2
Case 3

πŸ“£ Updates

  • [2026.07.30] πŸ”₯ We have released the code and models for EviBack.
  • [2026.07.28] πŸ”₯ Our paper is in public on arxiv.

πŸ’‘ Introduction

EviBack is a search-agent training framework with an Evidence-Constrained Teacher. It keeps inference Actor-only and uses the Teacher as a fallback when the Actor group receives no useful reward. The package provides the training, retrieval, evaluation, and environment components needed to reproduce the method. E2E-APE supplies the development-time prompt-engineering instructions used to freeze the Teacher prompt. it is not used during Actor inference.

πŸ—ΊοΈ Code map

.
β”œβ”€β”€ src/
β”‚   └── eviback/
β”‚       β”œβ”€β”€ teacher/
β”‚       β”‚   └── Stage A, conditional Stage B, client, merge, provenance
β”‚       β”œβ”€β”€ rewards/
β”‚       β”‚   └── Actor-first all-zero gate and paper reward
β”‚       β”œβ”€β”€ training/
β”‚       β”‚   └── VERL adapter and post-normalization reference logic
β”‚       β”œβ”€β”€ retrieval/
β”‚       β”‚   └── retriever client, optional server, public schema
β”‚       β”œβ”€β”€ inference.py
β”‚       β”‚   └── iterative Actor-only search
β”‚       β”œβ”€β”€ evaluation.py
β”‚       β”‚   └── aligned metrics and macro summaries
β”‚       └── aggregate.py
β”‚           └── paired bootstrap comparison
└── release/
    └── source freeze, environment, licenses, data decision, blockers

πŸ”§ Install

Choose one accelerator profile before installing:

Profile Runtime Environment and requirements Device argument
GPU NVIDIA CUDA 12.4 environment.yml, requirements-gpu.txt cuda
NPU Ascend 910B + CANN environment-npu.yml, requirements-npu.txt npu

The full installation procedures are GPU installation and Ascend NPU installation. Do not install the CUDA and NPU runtime profiles into the same environment.

The pure reward, parser, mock Teacher, evaluation, and data JSONL paths need no GPU:

python -m venv .venv
. .venv/bin/activate
pip install -e '.[test,data]'
python -m pytest -q tests
python scripts/run_reward_smoke.py

For GPU formal training, use Python 3.11, a CUDA-compatible PyTorch build, Transformers 4.57.6, vLLM 0.11.0, Ray, and a compatible VERL checkout. Apply third_party/verl/group_postnorm.patch before training. For Ascend NPU formal training, use the matched torch-npu and vllm-ascend stack described in docs/install_npu.md. The complete historical environment is not yet recovered; consult release/environment_snapshot.txt.

🌱 Prepare data

Obtain each upstream QA dataset under its own terms and arrange it as described in datasets-examples/README.md. Build the frozen 5,100/3,500 splits:

python scripts/prepare_data.py \
  --raw-root /path/to/raw-datasets \
  --output-root artifacts/data \
  --format parquet

python scripts/prepare_data.py \
  --check-only \
  --manifest datasets-examples/manifests/train_5100.json

Sampling is a stable SHA-256 order after normalized-question deduplication. The manifests publish quotas, seeds, and the original derived-artifact hashes without redistributing the data.

πŸ“ Services

Start a Search-R1-compatible dense retriever. Index, corpus, model, host, and port are explicit parameters:

bash scripts/start_retriever.sh \
  --index /path/to/e5_Flat.index \
  --corpus /path/to/docs.jsonl \
  --model intfloat/e5-base-v2 \
  --device cuda \
  --port 8000

Use --device cuda for NVIDIA GPUs or --device npu for Ascend NPUs.

Serve the training Teacher through an OpenAI-compatible chat-completions API, then export its endpoint and model. API credentials, when needed, are read from EVIBACK_TEACHER_API_KEY.

export EVIBACK_RETRIEVER_ENDPOINT=http://127.0.0.1:8000/retrieve
export EVIBACK_TEACHER_ENDPOINT=http://127.0.0.1:8001/v1/chat/completions
export EVIBACK_TEACHER_MODEL=your-glm-4.7-flash-revision

πŸ“ˆ Train

Start the formal VERL training job after installing VERL, applying the patch, and preparing the datasets and services:

bash scripts/train_actor.sh \
  --config configs/training/qwen3_1.7b.yaml

The entry point supports --model, --train-data, --eval-data, --retriever-endpoint, --teacher-endpoint, --teacher-model, --output-dir, --seed, --group-size, and --teacher-fallback-scale. Unknown arguments are forwarded as VERL/Hydra overrides.

The formal 1.7B configuration freezes 5,100 examples, 79 steps, eight rollouts, and seed 42. Model-scale and lambda matrices are documented in configs/training/qwen3_scale_matrix.yaml. Hardware-specific batch sizes remain explicit user overrides because they depend on the available accelerator.

πŸ“Š Evaluate

References are joined only after inference. The evaluator rejects duplicate or misaligned (data_source, index) keys and mismatched question text.

Models Download Link HuggingFace Download Link ModelScope
QWen3-0.6B πŸ€— HuggingFace πŸ”· ModelScope
QWen3-1.7B πŸ€— HuggingFace πŸ”· ModelScope
QWen3-4B πŸ€— HuggingFace πŸ”· ModelScope
bash scripts/run_evaluation.sh \
  --predictions artifacts/inference/evi_actor.jsonl \
  --references artifacts/data/eval_3500.jsonl \
  --output artifacts/evaluation/evi_actor.json \
  --model-checkpoint /path/to/actor-checkpoint \
  --retriever-revision your-index-revision

Reported metrics are legacy EM, token F1, valid-answer rate, mean search calls, duplicate-query rate, maximum-turn rate, single-hop macro, multi-hop macro, and paired bootstrap confidence intervals.

πŸ“Ž E2E-APE

E2E-APE is the project's development-time prompt-engineering process, not a runtime component. The public artifacts are the English prompt in ape/controller_instructions.md, the Chinese prompt in ape/controller_instructions_cn.md, and the English contract in ape/contract_en.md. They describe sample construction, labeling, prompt ablation, evaluation, and strategy selection for an agent/model to carry out. The generated runner, scorer, ablation, and selection code is intentionally not included. ape/prompts/frozen_policy.json records the selected prompt identifiers, while ape/benchmark/ contains metadata and schema only; the raw benchmark is not redistributed. See ape/README.md.

πŸ“’ Citation

If you find our work useful for your research, please consider citing the paper :

@article{ma2026eviback,
title={Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff},
  author={Xiao Ma, Zhiquan Hu, Yi Wei, Chenchen Zhao, Yijun Chen, Jicheng Zhao, Yuming Li, Chuang Dai},
  year={2026},
  eprint={2607.23955},
  archivePrefix={arXiv},
  primaryClass={cs.AI}
}

πŸ“œ License

The models in this repository are licensed under the Apache 2.0 License. We claim no rights over the your generated contents, granting you the freedom to use them while ensuring that your usage complies with the provisions of this license. You are fully accountable for your use of the models, which must not involve sharing any content that violates applicable laws, causes harm to individuals or groups, disseminates personal information intended for harm, spreads misinformation, or targets vulnerable populations.


Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for chery-nextai/eviback