---
license: mit
base_model: BAAI/bge-m3
library_name: sentence-transformers
pipeline_tag: feature-extraction
tags: [retrieval, nuclear-physics, NSR, EXFOR, bge-m3, dense-retrieval, sentence-transformers, NSR-CPT]
language: [en]
datasets: [NYSgpt/nsr-eval]
---
# NSR Encoder
*Dense retrieval over nuclear-physics literature (277,068 Nuclear Science References).*
[**NSR Collection**](https://huggingface.co/collections/NYSgpt/nsr-6a83f426fd0a4b01d54af471) | [**NSR Reranker**](https://huggingface.co/NYSgpt/nsr-reranker)
- ✅ **Nearly triples its base on expert queries:** R@10 **0.487 vs 0.171** for stock
`BAAI/bge-m3` (+186%), on 4,998 held-out expert keyword queries against all 277,068
papers.
- ✅ **Largest gains where retrieval is hardest:** title-only documents +262% R@10,
pre-1970 papers +307%.
- ✅ **Trained on expert-written queries:** NSR indexers hand-write a structured keyword
abstract for every paper — no click logs, no synthetic questions.
- ✅ **Drop-in:** standard `sentence-transformers` bi-encoder, 1024-d cosine vectors,
TEI-servable, MIT-licensed.
`ncbi/MedCPT` learned biomedical search from PubMed click logs. Nuclear physics has
something rarer: for ~200,000 papers, an NSR indexer wrote a canonical description of
what was measured and deduced. This model learned retrieval from **39,568** of those
expert query→paper pairs, plus EXFOR experiment→paper links, in 33 minutes on one GPU.
First stage of a two-stage stack — pair it with the
[NSR Reranker](https://huggingface.co/NYSgpt/nsr-reranker) for best results.
---
## Details
| Property | nsr-encoder |
|---|---|
| **Type** | Dense bi-encoder (single vector) |
| **Total parameters** | ~568M |
| **Backbone** | [BAAI/bge-m3](https://huggingface.co/BAAI/bge-m3) |
| **Output** | 1024-d normalized vector |
| **Similarity** | Cosine |
| **Sequence length** | 128 query / 256 passage |
| **Training signal** | 39.6k expert-written query→paper pairs (NSR keyword abstracts, EXFOR entries) |
| **Built for** | Search (query→document) over nuclear-physics literature |
| **Expert-keyword R@10** | 0.487 |
| **License** | MIT |
---
## Performance
**Expert keyword queries** (`KW`, n = 4,998) — an NSR indexer's structured keyword
abstract as the query, the paper it describes as the gold. Every arm is scored on the
same held-out queries, retrieved against the same 277,068 documents, with the same metric
code — on a benchmark ([NSR Eval](https://huggingface.co/datasets/NYSgpt/nsr-eval))
frozen before any training and split by paper.
| Rank | Arm | R@1 | R@10 | nDCG@10 |
|---:|---|---:|---:|---:|
| — | RRF(FTS + this model) *(the production arm)* | **0.344** | **0.542** | **0.437** |
| **1** | **nsr-encoder (ours)** | **0.252** | **0.487** | **0.363** |
| 2 | RRF(FTS + stock bge-m3) | 0.212 | 0.282 | 0.244 |
| 3 | Postgres FTS | 0.161 | 0.165 | 0.163 |
| 4 | BAAI/bge-m3 (stock base) | 0.080 | 0.171 | 0.121 |
**Where the gain lands** — R@10 by segment, stock base vs this model:
| Segment | n | stock bge-m3 | **nsr-encoder** | Δ |
|---|---:|---:|---:|---:|
| title-only documents | 4,141 | 0.127 | **0.458** | **+262%** |
| has-abstract documents | 857 | 0.383 | **0.631** | +65% |
| pre-1970 | 602 | 0.076 | **0.311** | **+307%** |
| 1970–1999 | 2,686 | 0.152 | **0.507** | +233% |
| 2000+ | 1,710 | 0.232 | **0.519** | +123% |
| journal articles | 4,250 | 0.179 | **0.495** | +176% |
| other reference types | 748 | 0.120 | **0.447** | +271% |
Pre-registered gate — ≥ stock base on R@10 and nDCG@10 in every segment, ≥ 10%
relative on blended KW — **passed** in all 8 segments.
**EXFOR queries** (`EX`, n = 4,997) — an experiment's title and reaction codes as the
query. Near-saturated for every dense arm; this model still leads: R@1 **0.911** vs
0.871 for stock bge-m3, R@10 0.966 vs 0.958.
---
## Training
| | |
|---|---|
| Objective | in-batch contrastive, dense only (`--unified_finetuning False`) |
| Trainer | FlagEmbedding `finetune.embedder.encoder_only.m3` |
| Pairs | **39,568** — all EXFOR links + keyword abstracts stratified across era / richness / reference-type cells |
| Negatives / group | 7 hard negatives mined from the base model's own embedding space + a lexical arm · group size 8 |
| Batch / epochs / lr / temperature | 16 / 1 / 1e-5 / 0.02 |
| Precision | bf16 |
| Hardware / wall-clock | 1× NVIDIA L40S (`g6e.xlarge`) · 33.5 min |
Split by paper; every benchmark paper is excluded as a query source and positive.
---
## How to run
```python
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("NYSgpt/nsr-encoder")
query = model.encode(["92Zr(n,γ) cross section, stellar nucleosynthesis"], normalize_embeddings=True)
docs = model.encode(
["Neutron capture cross sections of 92Zr and their astrophysical implications ..."],
normalize_embeddings=True,
)
print(query @ docs.T)
```
Serves cleanly on Hugging Face TEI: `--model-id NYSgpt/nsr-encoder`.
**For best quality, add the second stage:** re-score this model's top-50 with the
[NSR Reranker](https://huggingface.co/NYSgpt/nsr-reranker).
---
## 📬 Contact
Questions, results, or a use case to share? Open a discussion in the
[Community tab](https://huggingface.co/NYSgpt/nsr-encoder/discussions).
## Citation
```bibtex
@misc{nsrencoder2026,
title = {NSR Encoder: dense retrieval over the Nuclear Science References corpus},
author = {NYSgpt},
year = {2026},
url = {https://huggingface.co/NYSgpt/nsr-encoder}
}
```