Instructions to use chrullis/relweave-1.7b-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use chrullis/relweave-1.7b-base with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
relweave-1.7b-base
Typed entity and relation extraction from English text into a knowledge graph. relweave-1.7b-base reads a passage and returns its entities (with all their mentions, including pronouns and descriptions such as "the company") and the typed relations between them, following a fixed business schema of 6 entity types and 21 relation types (companies, people, places, ownership, subsidiaries, executives, board members, acquisitions, and more).
It is the smaller variant of relweave-4b-base, trained with the same data and recipe: two LoRA adapters on Qwen3-1.7B (4-bit) that work together:
generator/: writes the entities and relations as text, under constrained decoding.head/: a pair-classification head that scores every entity pair; its scores remove relations the generator invented and add ones it missed.
Use it through the relweave library, which chunks long documents, runs both parts, merges entities across chunks and returns a graph in JSON Graph Format.
Usage
pip install git+https://github.com/memlocator/relweave # PyPI release to follow
relweave run report.txt --weights chrullis/relweave-1.7b-base --out graph.json
from relweave import Extractor
ex = Extractor(weights="chrullis/relweave-1.7b-base")
graph = ex.run(open("report.txt").read())
print(graph.to_json()) # JSON Graph Format v2: nodes, edges, mention spans, evidence
A CUDA GPU is needed (the base model is 4-bit; the 1.7B needs less memory than the 4B).
Deployment
relweave runs in two steps per chunk: the generator writes entities and relations, then the pair head scores every entity pair. The generator is most of the run time, and it can be served by vLLM:
vllm serve chrullis/relweave-1.7b-base --max-model-len 4096
# or the official container (same arguments):
docker run --gpus all -p 8000:8000 --ipc=host vllm/vllm-openai --model chrullis/relweave-1.7b-base --max-model-len 4096
from relweave import Extractor
ex = Extractor(weights="chrullis/relweave-1.7b-base", generator_url="http://localhost:8000")
graph = ex.run(open("report.txt").read())
- The model at the repository root is the generator with its LoRA already merged (see Files), so vLLM needs no
--enable-loraand no adapter arguments. Tested with vLLM 0.31. - The endpoint is a component, not a chat model: it answers in relweave's line format only when prompted the way it
was trained and decoded under the per-chunk grammar relweave sends with each request. Call it through relweave
(
generator_url), which also runs the pair head and merges entities across chunks. - The pair head always runs locally in relweave (it needs the model's hidden states, which vLLM does not return), on a GPU with about 3 GB free.
- On an 8 GB card the 1.7B server and the pair head fit together with
--gpu-memory-utilization 0.52 --max-model-len 3072 --max-num-batched-tokens 1024 --enforce-eager. - Without a CUDA toolkit installed, start vLLM with
VLLM_USE_FLASHINFER_SAMPLER=0.
Choosing a model
| model | system F1 (validation / test) | generator on vLLM, 8 GB GPU | generator, transformers backend | vLLM weights in memory |
|---|---|---|---|---|
| relweave-4b-base | 0.781 / 0.783 | about 42 chunks/min (fp8) | about 4.5 chunks/min | 4.2 GB (fp8) |
| relweave-1.7b-base | 0.737 / 0.722 | about 96 chunks/min (bf16) | about 6 chunks/min | 3.3 GB (bf16) |
Speeds on one RTX 3070 Ti (8 GB), 64 test chunks of about 200 words. Over 377 held-out chunks the 1.7B scores 0.048 below the 4B (95% interval 0.030 to 0.067).
Results
Strict typed relation F1 (a relation counts only with the right type, direction and both entities), on held-out English Wikipedia passages about Nordic and European companies. Evaluation sets share no documents, entities or facts with the training data; the test set was reported, never tuned on.
| system | validation | test |
|---|---|---|
| relweave-1.7b-base (generator + pair head, union) | 0.737 | 0.722 |
| generator alone | 0.622 | 0.626 |
| relweave-4b-base, for comparison | 0.781 | 0.783 |
Test precision 0.798, recall 0.659. Over validation, test and a fresh test set (377 chunks) the 1.7B scores 0.732 against the 4B's 0.780: 0.048 lower, 95% interval 0.030 to 0.067. Use the 4B unless memory or speed rules it out.
On the fresh test set (64 chunks), the merged generator served by vLLM scores 0.749 for the whole system against 0.736 with the transformers backend.
Schema
The schema is written once, as Python classes in bench/ontology/business.py. The docstring of each class is its
definition. Prompts, decoding grammars, label validation and scoring are derived from it. The tables below are
generated from those classes.
Entity types (6)
| type | definition |
|---|---|
| PERSON | A human being. |
| ORG | A company, institution, family acting as owner, or other organised body; publications, imprints and brands are Orgs. |
| OBJECT | A physical thing: vehicle, vessel, building, artwork. |
| PLACE | A country, region, city or other location. |
| COORDINATE | A numeric map coordinate. |
| EVENT | Something that happened at a time: a sale, meeting, election, ceremony. |
Relation types (21)
Edge attributes such as dates, shares and roles exist in the classes but are not part of the scored output. Symmetric relations have no direction.
| relation | endpoints (source -> target) | definition |
|---|---|---|
| EMPLOYED_BY | Person -> Org | Person works for Org in a non-executive role |
| EXECUTIVE_OF | Person -> Org | Person holds an executive title at Org (CEO, CFO, managing director) or chairs its board or supervisory board (title chairman, vice chairman, honorary chairman) |
| BOARD_MEMBER_OF | Person -> Org | Person sits on the board of Org as a member; a board chair is EXECUTIVE_OF instead |
| MEMBER_OF | Person -> Org | Person is a member of Org (club, party, association) |
| FAMILY_OF | Person -> Person (symmetric) | Family tie; kind is what the target is to the source: spouse, parent, child or sibling |
| ASSOCIATE_OF | Person -> Person (symmetric) | Persons described as associates, partners or close contacts |
| MET_WITH | Person -> Person (symmetric) | Two persons met in person |
| COMMUNICATED_WITH | Person -> Person (symmetric) | Two persons communicated (call, email, message, letter) |
| LOCATED_IN | Object or Org or Person or Place -> Place | Source is or was resident or situated in the Place (dated if the text says when), or a Place lies within another; not a birthplace, not a market; an Org that is based in a Place is HEADQUARTERED_IN |
| BORN_IN | Person -> Place | Person was born in the Place |
| OPERATES_IN | Org -> Place | Org has operations, offices, plants, stores or sales in the Place |
| HEADQUARTERED_IN | Org -> Place | Org has its headquarters in, or is based in, the Place |
| FOUNDED | Org or Person -> Org | Person or Org founded or co-founded the Org; when Orgs merge to form a new Org, each merging Org FOUNDED it |
| OWNS_STAKE_IN | Org or Person -> Org | Person or Org owns shares in, a stake in, or controls through ownership the Org; share is the stated fraction |
| SUBSIDIARY_OF | Org -> Org | Source Org is a subsidiary, division or brand of the target Org; publications, magazines, newspapers, imprints and product brands are Orgs linked this way, never Objects; in a chain of owners, link only to the nearest parent the text states |
| ACQUIRED | Org or Person -> Org | Person or Org bought or took over the target Org; only when the text names the buyer |
| HAS_COORDINATE | Place -> Coordinate | Place has the stated numeric coordinate |
| OWNS_OBJECT | Org or Person -> Object | Person or Org owns the Object (a physical thing: vehicle, vessel, building, artwork); brands and publications are Orgs |
| TRANSFERRED_OBJECT | Event -> Object | The Event transferred the Object (sale, delivery, seizure) |
| PARTICIPATED_IN | Org or Person -> Event | Person or Org took part in the Event; role is buyer, seller, host, attendee |
| HELD_AT | Event -> Place | The Event took place at the Place |
Training
The training, validation and test data are published as chrullis/relweave-business-data (CC BY-SA 4.0, with the source article of every passage).
The training data are passages from English Wikipedia articles about Nordic and European companies, annotated by a
large language model following written annotation conventions and spot-checked; later rounds corrected model drafts
instead of labelling from scratch. The generator is a QLoRA fine-tune (rank 32) of Qwen3-1.7B; the pair head is trained
together with its own LoRA on the generator's own entity lists. The full methodology, including the evaluation
protocol and the experiments that did not help, is in the library repository (docs/methodology.md).
Limitations
- At most 40 entities per chunk. The output grammar and schema cap a chunk at 40 entities. Long text must be
chunked.
relweavechunks documents into whole paragraphs of up to 200 words with one paragraph of overlap, and merges entities and relations across chunks. - Repetition loops in decoding. The generator can fill relation blocks until a cap. A stop-on-repeat check reduced capped blocks from 109 to 7 on one 4B run, but loops are not eliminated.
- Business-only schema. The 6 entity types and 21 relation types are fixed. New entity or relation types need labelled data and training (see the recipe in the relweave repository); a version that reads a user-defined schema at run time is in development and is not part of this model.
- Coreference is weaker than entity finding. Proper names are found reliably. Error analysis of the 4B (not repeated for the 1.7B): on validation 150 of 1,981 gold entities (7.6%) are missed: 60% never written (divisions and brands with ordinary-word names), 25% written with another type, 15% grouped differently (a company merged with its predecessor).
- Lists are the biggest error source. For the 4B, of 250 missed test relations, 159 sit in long sentences, mostly coordinated lists where later items are dropped (375 of 522 list items found). Invented relations mostly link co-occurring entities (49 of 161) or entities that share a neighbour (34).
- Labels and test size. Labels come from one annotator model under one convention. The test set has 105 chunks. All text is English business Wikipedia, cleaner than news or filings.
Licences and attribution
| component | licence | what it requires |
|---|---|---|
| Wikipedia article text (training, validation and test chunks) | CC BY-SA 4.0 | attribution; share-alike for any redistributed text or data derived from it |
| Re-DocRED (experiments only) | MIT | keep the copyright notice |
| Qwen3-1.7B (base model) | Apache 2.0 | keep the licence and notices |
The model weights (both adapters and the pair head) are released under the Apache License 2.0, like the base model.
They were trained on passages from English Wikipedia (CC BY-SA 4.0; attribution: Wikipedia contributors,
https://en.wikipedia.org). Redistributing the training text or labels derived from it falls under CC BY-SA 4.0. Other
data tried in experiments was not used to train relweave-1.7b-base.
Files
| path | contents | used by |
|---|---|---|
config.json, model.safetensors, tokenizer.json, tokenizer_config.json, chat_template.jinja, generation_config.json (root) |
the generator as one plain bf16 model: the 4-bit Qwen3-1.7B base the adapter was trained on, dequantized, with the generator LoRA merged in. The pair head is not in it. | vLLM and other servers |
generator/ |
the generator LoRA adapter (PEFT) and tokenizer, unmerged; relweave_config.json holds the base model id and prompt settings |
relweave, transformers backend (adapter on the 4-bit base) |
head/ |
the pair head: its own LoRA adapter (PEFT) on the 4-bit base, head.safetensors (the classifier over the probe hidden states) and head_config.json (label set, layers read, union settings) |
relweave, always local |
schema.json |
the business schema as JSON (types, definitions, endpoints) | relweave |
relweave downloads only generator/, head/ and schema.json. vLLM loads only the root model; its download also
fetches the adapters' .safetensors files (a few hundred MB), which it does not use.
- Downloads last month
- -