relweave-4b-base

Typed entity and relation extraction from English text into a knowledge graph. relweave-4b-base reads a passage and returns its entities (with all their mentions, including pronouns and descriptions such as "the company") and the typed relations between them, following a fixed business schema of 6 entity types and 21 relation types (companies, people, places, ownership, subsidiaries, executives, board members, acquisitions, and more).

It is two LoRA adapters on Qwen3-4B (4-bit) that work together:

  • generator/: writes the entities and relations as text, under constrained decoding.
  • head/: a pair-classification head that scores every entity pair; its scores remove relations the generator invented and add ones it missed.

Use it through the relweave library, which chunks long documents, runs both parts, merges entities across chunks and returns a graph in JSON Graph Format.

Usage

pip install git+https://github.com/memlocator/relweave   # PyPI release to follow
relweave run report.txt --out graph.json
from relweave import Extractor

graph = Extractor().run(open("report.txt").read())  # downloads chrullis/relweave-4b-base
print(graph.to_json())                              # JSON Graph Format v2: nodes, edges, mention spans, evidence

A CUDA GPU with about 5 GB of free memory is needed (the base model is 4-bit).

Results

Strict typed relation F1 (a relation counts only with the right type, direction and both entities), on held-out English Wikipedia passages about Nordic and European companies. Evaluation sets share no documents, entities or facts with the training data; the test set was reported, never tuned on.

evaluation validation test
end to end (the model finds the entities itself) 0.781 0.783
relation step with gold entities 0.839 0.822

Test precision 0.818, recall 0.751. The pair head adds about +0.10 F1 over the generator alone.

Schema

The schema is written once, as Python classes in bench/ontology/business.py. The docstring of each class is its definition. Prompts, decoding grammars, label validation and scoring are derived from it. The tables below are generated from those classes.

Entity types (6)

type definition
PERSON A human being.
ORG A company, institution, family acting as owner, or other organised body; publications, imprints and brands are Orgs.
OBJECT A physical thing: vehicle, vessel, building, artwork.
PLACE A country, region, city or other location.
COORDINATE A numeric map coordinate.
EVENT Something that happened at a time: a sale, meeting, election, ceremony.

Relation types (21)

Edge attributes such as dates, shares and roles exist in the classes but are not part of the scored output. Symmetric relations have no direction.

relation endpoints (source -> target) definition
EMPLOYED_BY Person -> Org Person works for Org in a non-executive role
EXECUTIVE_OF Person -> Org Person holds an executive title at Org (CEO, CFO, managing director) or chairs its board or supervisory board (title chairman, vice chairman, honorary chairman)
BOARD_MEMBER_OF Person -> Org Person sits on the board of Org as a member; a board chair is EXECUTIVE_OF instead
MEMBER_OF Person -> Org Person is a member of Org (club, party, association)
FAMILY_OF Person -> Person (symmetric) Family tie; kind is what the target is to the source: spouse, parent, child or sibling
ASSOCIATE_OF Person -> Person (symmetric) Persons described as associates, partners or close contacts
MET_WITH Person -> Person (symmetric) Two persons met in person
COMMUNICATED_WITH Person -> Person (symmetric) Two persons communicated (call, email, message, letter)
LOCATED_IN Object or Org or Person or Place -> Place Source is or was resident or situated in the Place (dated if the text says when), or a Place lies within another; not a birthplace, not a market; an Org that is based in a Place is HEADQUARTERED_IN
BORN_IN Person -> Place Person was born in the Place
OPERATES_IN Org -> Place Org has operations, offices, plants, stores or sales in the Place
HEADQUARTERED_IN Org -> Place Org has its headquarters in, or is based in, the Place
FOUNDED Org or Person -> Org Person or Org founded or co-founded the Org; when Orgs merge to form a new Org, each merging Org FOUNDED it
OWNS_STAKE_IN Org or Person -> Org Person or Org owns shares in, a stake in, or controls through ownership the Org; share is the stated fraction
SUBSIDIARY_OF Org -> Org Source Org is a subsidiary, division or brand of the target Org; publications, magazines, newspapers, imprints and product brands are Orgs linked this way, never Objects; in a chain of owners, link only to the nearest parent the text states
ACQUIRED Org or Person -> Org Person or Org bought or took over the target Org; only when the text names the buyer
HAS_COORDINATE Place -> Coordinate Place has the stated numeric coordinate
OWNS_OBJECT Org or Person -> Object Person or Org owns the Object (a physical thing: vehicle, vessel, building, artwork); brands and publications are Orgs
TRANSFERRED_OBJECT Event -> Object The Event transferred the Object (sale, delivery, seizure)
PARTICIPATED_IN Org or Person -> Event Person or Org took part in the Event; role is buyer, seller, host, attendee
HELD_AT Event -> Place The Event took place at the Place

Training

The training data are passages from English Wikipedia articles about Nordic and European companies, annotated by a large language model following written annotation conventions and spot-checked; later rounds corrected model drafts instead of labelling from scratch. The generator is a QLoRA fine-tune (rank 32) of Qwen3-4B; the pair head is trained together with its own LoRA on the generator's own entity lists. The full methodology, including the evaluation protocol and the experiments that did not help, is in the library repository (docs/methodology.md).

Limitations

  • At most 40 entities per chunk. The output grammar and schema cap a chunk at 40 entities. Long text must be chunked. relweave chunks documents into whole paragraphs of up to 200 words with one paragraph of overlap, and merges entities and relations across chunks.
  • Repetition loops in decoding. The generator can fill relation blocks until a cap. A stop-on-repeat check reduced capped blocks from 109 to 7 on one run, but loops are not eliminated.
  • Business-only schema. The 6 entity types and 21 relation types are fixed. New entity or relation types need labelled data and training (see the recipe in the relweave repository); a version that reads a user-defined schema at run time is in development and is not part of this model.
  • Coreference is weaker than entity finding. Proper names are found reliably. On validation 150 of 1,981 gold entities (7.6%) are missed: 60% never written (divisions and brands with ordinary-word names), 25% written with another type, 15% grouped differently (a company merged with its predecessor).
  • Lists are the biggest error source. Of 250 missed test relations, 159 sit in long sentences, mostly coordinated lists where later items are dropped (375 of 522 list items found). Invented relations mostly link co-occurring entities (49 of 161) or entities that share a neighbour (34).
  • Labels and test size. Labels come from one annotator model under one convention. The test set has 105 chunks. All text is English business Wikipedia, cleaner than news or filings.

Licences and attribution

component licence what it requires
Wikipedia article text (training, validation and test chunks) CC BY-SA 4.0 attribution; share-alike for any redistributed text or data derived from it
Re-DocRED (experiments only) MIT keep the copyright notice
Qwen3-4B (base model) Apache 2.0 keep the licence and notices

The model weights (both adapters and the pair head) are released under the Apache License 2.0, like the base model. They were trained on passages from English Wikipedia (CC BY-SA 4.0; attribution: Wikipedia contributors, https://en.wikipedia.org). Redistributing the training text or labels derived from it falls under CC BY-SA 4.0. Other data tried in experiments was not used to train relweave-4b-base.

Files

path contents
generator/ LoRA adapter (PEFT format) and tokenizer; relweave_config.json holds the prompt settings
head/ LoRA adapter (PEFT format) and tokenizer; head.safetensors the pair head; head_config.json its label set, layers read and union settings
schema.json the business schema as JSON (types, definitions, endpoints)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support