relweave-1.7b-base

Typed entity and relation extraction from English text into a knowledge graph. relweave-1.7b-base reads a passage and returns its entities (with all their mentions, including pronouns and descriptions such as "the company") and the typed relations between them, following a fixed business schema of 6 entity types and 21 relation types (companies, people, places, ownership, subsidiaries, executives, board members, acquisitions, and more).

It is the smaller variant of relweave-4b-base, trained with the same data and recipe: two LoRA adapters on Qwen3-1.7B (4-bit) that work together:

  • generator/: writes the entities and relations as text, under constrained decoding.
  • head/: a pair-classification head that scores every entity pair; its scores remove relations the generator invented and add ones it missed.

Use it through the relweave library, which chunks long documents, runs both parts, merges entities across chunks and returns a graph in JSON Graph Format.

Usage

pip install git+https://github.com/memlocator/relweave   # PyPI release to follow
relweave run report.txt --weights chrullis/relweave-1.7b-base --out graph.json
from relweave import Extractor

ex = Extractor(weights="chrullis/relweave-1.7b-base")
graph = ex.run(open("report.txt").read())
print(graph.to_json())                              # JSON Graph Format v2: nodes, edges, mention spans, evidence

A CUDA GPU is needed (the base model is 4-bit; the 1.7B needs less memory than the 4B).

Deployment

relweave runs in two steps per chunk: the generator writes entities and relations, then the pair head scores every entity pair. The generator is most of the run time, and it can be served by vLLM:

vllm serve chrullis/relweave-1.7b-base --max-model-len 4096
# or the official container (same arguments):
docker run --gpus all -p 8000:8000 --ipc=host vllm/vllm-openai --model chrullis/relweave-1.7b-base --max-model-len 4096
from relweave import Extractor
ex = Extractor(weights="chrullis/relweave-1.7b-base", generator_url="http://localhost:8000")
graph = ex.run(open("report.txt").read())
  • The model at the repository root is the generator with its LoRA already merged (see Files), so vLLM needs no --enable-lora and no adapter arguments. Tested with vLLM 0.31.
  • The endpoint is a component, not a chat model: it answers in relweave's line format only when prompted the way it was trained and decoded under the per-chunk grammar relweave sends with each request. Call it through relweave (generator_url), which also runs the pair head and merges entities across chunks.
  • The pair head always runs locally in relweave (it needs the model's hidden states, which vLLM does not return), on a GPU with about 3 GB free.
  • On an 8 GB card the 1.7B server and the pair head fit together with --gpu-memory-utilization 0.52 --max-model-len 3072 --max-num-batched-tokens 1024 --enforce-eager.
  • Without a CUDA toolkit installed, start vLLM with VLLM_USE_FLASHINFER_SAMPLER=0.

Choosing a model

model system F1 (validation / test) generator on vLLM, 8 GB GPU generator, transformers backend vLLM weights in memory
relweave-4b-base 0.781 / 0.783 about 42 chunks/min (fp8) about 4.5 chunks/min 4.2 GB (fp8)
relweave-1.7b-base 0.737 / 0.722 about 96 chunks/min (bf16) about 6 chunks/min 3.3 GB (bf16)

Speeds on one RTX 3070 Ti (8 GB), 64 test chunks of about 200 words. Over 377 held-out chunks the 1.7B scores 0.048 below the 4B (95% interval 0.030 to 0.067).

Results

Strict typed relation F1 (a relation counts only with the right type, direction and both entities), on held-out English Wikipedia passages about Nordic and European companies. Evaluation sets share no documents, entities or facts with the training data; the test set was reported, never tuned on.

system validation test
relweave-1.7b-base (generator + pair head, union) 0.737 0.722
generator alone 0.622 0.626
relweave-4b-base, for comparison 0.781 0.783

Test precision 0.798, recall 0.659. Over validation, test and a fresh test set (377 chunks) the 1.7B scores 0.732 against the 4B's 0.780: 0.048 lower, 95% interval 0.030 to 0.067. Use the 4B unless memory or speed rules it out.

On the fresh test set (64 chunks), the merged generator served by vLLM scores 0.749 for the whole system against 0.736 with the transformers backend.

Schema

The schema is written once, as Python classes in bench/ontology/business.py. The docstring of each class is its definition. Prompts, decoding grammars, label validation and scoring are derived from it. The tables below are generated from those classes.

Entity types (6)

type definition
PERSON A human being.
ORG A company, institution, family acting as owner, or other organised body; publications, imprints and brands are Orgs.
OBJECT A physical thing: vehicle, vessel, building, artwork.
PLACE A country, region, city or other location.
COORDINATE A numeric map coordinate.
EVENT Something that happened at a time: a sale, meeting, election, ceremony.

Relation types (21)

Edge attributes such as dates, shares and roles exist in the classes but are not part of the scored output. Symmetric relations have no direction.

relation endpoints (source -> target) definition
EMPLOYED_BY Person -> Org Person works for Org in a non-executive role
EXECUTIVE_OF Person -> Org Person holds an executive title at Org (CEO, CFO, managing director) or chairs its board or supervisory board (title chairman, vice chairman, honorary chairman)
BOARD_MEMBER_OF Person -> Org Person sits on the board of Org as a member; a board chair is EXECUTIVE_OF instead
MEMBER_OF Person -> Org Person is a member of Org (club, party, association)
FAMILY_OF Person -> Person (symmetric) Family tie; kind is what the target is to the source: spouse, parent, child or sibling
ASSOCIATE_OF Person -> Person (symmetric) Persons described as associates, partners or close contacts
MET_WITH Person -> Person (symmetric) Two persons met in person
COMMUNICATED_WITH Person -> Person (symmetric) Two persons communicated (call, email, message, letter)
LOCATED_IN Object or Org or Person or Place -> Place Source is or was resident or situated in the Place (dated if the text says when), or a Place lies within another; not a birthplace, not a market; an Org that is based in a Place is HEADQUARTERED_IN
BORN_IN Person -> Place Person was born in the Place
OPERATES_IN Org -> Place Org has operations, offices, plants, stores or sales in the Place
HEADQUARTERED_IN Org -> Place Org has its headquarters in, or is based in, the Place
FOUNDED Org or Person -> Org Person or Org founded or co-founded the Org; when Orgs merge to form a new Org, each merging Org FOUNDED it
OWNS_STAKE_IN Org or Person -> Org Person or Org owns shares in, a stake in, or controls through ownership the Org; share is the stated fraction
SUBSIDIARY_OF Org -> Org Source Org is a subsidiary, division or brand of the target Org; publications, magazines, newspapers, imprints and product brands are Orgs linked this way, never Objects; in a chain of owners, link only to the nearest parent the text states
ACQUIRED Org or Person -> Org Person or Org bought or took over the target Org; only when the text names the buyer
HAS_COORDINATE Place -> Coordinate Place has the stated numeric coordinate
OWNS_OBJECT Org or Person -> Object Person or Org owns the Object (a physical thing: vehicle, vessel, building, artwork); brands and publications are Orgs
TRANSFERRED_OBJECT Event -> Object The Event transferred the Object (sale, delivery, seizure)
PARTICIPATED_IN Org or Person -> Event Person or Org took part in the Event; role is buyer, seller, host, attendee
HELD_AT Event -> Place The Event took place at the Place

Training

The training, validation and test data are published as chrullis/relweave-business-data (CC BY-SA 4.0, with the source article of every passage).

The training data are passages from English Wikipedia articles about Nordic and European companies, annotated by a large language model following written annotation conventions and spot-checked; later rounds corrected model drafts instead of labelling from scratch. The generator is a QLoRA fine-tune (rank 32) of Qwen3-1.7B; the pair head is trained together with its own LoRA on the generator's own entity lists. The full methodology, including the evaluation protocol and the experiments that did not help, is in the library repository (docs/methodology.md).

Limitations

  • At most 40 entities per chunk. The output grammar and schema cap a chunk at 40 entities. Long text must be chunked. relweave chunks documents into whole paragraphs of up to 200 words with one paragraph of overlap, and merges entities and relations across chunks.
  • Repetition loops in decoding. The generator can fill relation blocks until a cap. A stop-on-repeat check reduced capped blocks from 109 to 7 on one 4B run, but loops are not eliminated.
  • Business-only schema. The 6 entity types and 21 relation types are fixed. New entity or relation types need labelled data and training (see the recipe in the relweave repository); a version that reads a user-defined schema at run time is in development and is not part of this model.
  • Coreference is weaker than entity finding. Proper names are found reliably. Error analysis of the 4B (not repeated for the 1.7B): on validation 150 of 1,981 gold entities (7.6%) are missed: 60% never written (divisions and brands with ordinary-word names), 25% written with another type, 15% grouped differently (a company merged with its predecessor).
  • Lists are the biggest error source. For the 4B, of 250 missed test relations, 159 sit in long sentences, mostly coordinated lists where later items are dropped (375 of 522 list items found). Invented relations mostly link co-occurring entities (49 of 161) or entities that share a neighbour (34).
  • Labels and test size. Labels come from one annotator model under one convention. The test set has 105 chunks. All text is English business Wikipedia, cleaner than news or filings.

Licences and attribution

component licence what it requires
Wikipedia article text (training, validation and test chunks) CC BY-SA 4.0 attribution; share-alike for any redistributed text or data derived from it
Re-DocRED (experiments only) MIT keep the copyright notice
Qwen3-1.7B (base model) Apache 2.0 keep the licence and notices

The model weights (both adapters and the pair head) are released under the Apache License 2.0, like the base model. They were trained on passages from English Wikipedia (CC BY-SA 4.0; attribution: Wikipedia contributors, https://en.wikipedia.org). Redistributing the training text or labels derived from it falls under CC BY-SA 4.0. Other data tried in experiments was not used to train relweave-1.7b-base.

Files

path contents used by
config.json, model.safetensors, tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json (root) the generator as one plain bf16 model: the 4-bit Qwen3-1.7B base the adapter was trained on, dequantized, with the generator LoRA merged in. The pair head is not in it. vLLM and other servers
generator/ the generator LoRA adapter (PEFT) and tokenizer, unmerged; relweave_config.json holds the base model id and prompt settings relweave, transformers backend (adapter on the 4-bit base)
head/ the pair head: its own LoRA adapter (PEFT) on the 4-bit base, head.safetensors (the classifier over the probe hidden states) and head_config.json (label set, layers read, union settings) relweave, always local
schema.json the business schema as JSON (types, definitions, endpoints) relweave

relweave downloads only generator/, head/ and schema.json. vLLM loads only the root model; its download also fetches the adapters' .safetensors files (a few hundred MB), which it does not use.

Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train chrullis/relweave-1.7b-base