Spaces:
Paused
Dataset β Database Field Mapping
Maps every field from the HuggingFace dataset
jamescalam/ai-arxiv2-chunks
to its destination in the PostgreSQL schema.
Quick Reference
| HF Dataset Field | Type | β | DB Table | DB Column | Transform |
|---|---|---|---|---|---|
doi |
string |
β | documents |
arxiv_id |
direct copy |
id |
string |
β | documents |
arxiv_id |
fallback if doi absent |
title |
string |
β | documents |
title |
direct copy |
doi + title |
string |
β | documents |
full_citation |
computed: "{title}. arXiv:{doi}" |
doi |
string |
β | documents |
source_url |
computed: "https://arxiv.org/abs/{doi}" |
source |
string |
β | documents |
metadata["source"] |
stored in JSONB |
source |
string |
β | chunks |
metadata["source"] |
stored in JSONB |
chunk |
string |
β | chunks |
content |
direct copy |
chunk-id |
int64 |
β οΈ | chunks |
chunk_index |
field name mismatch β always 0 |
| (computed) | float[384] |
β | chunks |
embedding |
encoded locally via all-MiniLM-L6-v2 |
summary |
string |
β | β | β | not stored |
authors |
string |
β | β | β | not stored |
categories |
string |
β | β | β | not stored |
primary_category |
string |
β | β | β | not stored |
published |
string |
β | β | β | not stored |
updated |
string |
β | β | β | not stored |
comment |
string? |
β | β | β | not stored |
journal_ref |
string? |
β | β | β | not stored |
references |
list |
β | β | β | not stored |
Detailed Mapping
documents table
| DB Column | HF Field(s) | Ingest Code | Example |
|---|---|---|---|
id |
β | gen_random_uuid() auto-generated |
235692e6-67ae-... |
arxiv_id |
doi (primary), id (fallback) |
row.get("doi", row.get("arxiv_id", uuid4())) |
2401.09350 |
title |
title |
row.get("title") |
Foundations of Vector Retrieval |
source_url |
derived from doi |
f"https://arxiv.org/abs/{doi}" |
https://arxiv.org/abs/2401.09350 |
full_citation |
title + doi |
f"{title}. arXiv:{doi}" |
Foundations of Vector Retrieval. arXiv:2401.09350 |
metadata |
source |
{"source": row.get("source", "")} |
{"source": "http://arxiv.org/pdf/2401.09350"} |
created_at |
β | NOW() (DB default) |
2026-03-20T15:04:19Z |
Note on
source_url: The dataset'ssourcefield contains the PDF URL (http://arxiv.org/pdf/...). The ingest script ignores this and instead constructs an abstract URL (https://arxiv.org/abs/...) fromdoi. The raw PDF URL is preserved only insidemetadata["source"].
chunks table
| DB Column | HF Field(s) | Ingest Code | Example |
|---|---|---|---|
id |
β | gen_random_uuid() auto-generated |
a1b2c3d4-... |
document_id |
β | FK from documents.id upsert |
235692e6-... |
content |
chunk (primary), text (fallback) |
row.get("chunk", row.get("text", "")) |
"Sebastian Bruch # Foundations of Vector Retrieval..." |
embedding |
β | SentenceTransformer.encode(content) |
[0.012, -0.034, ..., 0.091] (384 floats) |
chunk_index |
β οΈ chunk-id |
row.get("chunk_index", 0) |
Always 0 β see bug below |
metadata |
source |
{"source": row.get("source", "")} |
{"source": "http://arxiv.org/pdf/2401.09350"} |
Known Issues
β οΈ Bug 1 β chunk_index always 0 (field name mismatch)
Root cause: The HuggingFace dataset names the field chunk-id (with a hyphen),
but the ingest script looks up chunk_index (with an underscore):
# ingest_data.py line 178 β CURRENT (broken)
chunk_index = row.get("chunk_index", 0) # 'chunk_index' never exists β always 0
# CORRECT
chunk_index = row.get("chunk-id", row.get("chunk_index", 0))
Impact:
- All chunks have
chunk_index = 0 - The unique constraint
UNIQUE(document_id, chunk_index)means only one chunk per document is stored β all subsequent chunks of the same document are silently skipped as duplicates - 55 documents Γ 1 chunk each = 55 rows instead of the expected ~5 000 unique chunks
Fix: See ingest_data.py β the lookup key must be chunk-id.
β Fields Not Stored (9 dataset fields dropped)
Nine dataset fields are read during ingestion but discarded without being written to any DB column:
| HF Field | Content | Why Useful | Suggested Destination |
|---|---|---|---|
summary |
Full paper abstract (228β1 920 chars) | Richer context for retrieval; could be searchable | documents.metadata["summary"] |
authors |
Comma-separated author names | Citation display; author-based filtering | documents.metadata["authors"] |
categories |
All arXiv categories (cs.LG, cs.AI, ...) |
Category-based RBAC or filtering | documents.metadata["categories"] |
primary_category |
Primary arXiv category | UI display; query routing | documents.metadata["primary_category"] |
published |
Publication date (YYYYMMDD) |
Temporal filtering; freshness ranking | documents.metadata["published"] |
updated |
Last update date (YYYYMMDD) |
Detect revised papers | documents.metadata["updated"] |
comment |
Author comments (nullable) | Version notes, e.g. "Published in NeurIPS 2023" | documents.metadata["comment"] |
journal_ref |
Journal reference (nullable) | Peer-review status indicator | documents.metadata["journal_ref"] |
references |
List of cited paper IDs | Citation graph; could link documents | documents.metadata["references"] |
All nine fields can be preserved by expanding the _get_or_create_document call in
ingest_data.py to write them into documents.metadata:
# Expanded metadata β stores all dropped fields
psycopg2.extras.Json({
"source": row.get("source", ""),
"summary": row.get("summary"),
"authors": row.get("authors"),
"categories": row.get("categories"),
"primary_category": row.get("primary_category"),
"published": row.get("published"),
"updated": row.get("updated"),
"comment": row.get("comment"),
"journal_ref": row.get("journal_ref"),
"references": row.get("references", []),
})
Full Dataset Schema (for reference)
All 15 fields available in jamescalam/ai-arxiv2-chunks (train split, 241k rows):
| Field | Type | Nullable | Range / Notes |
|---|---|---|---|
doi |
string | no | 10 chars β arXiv ID (e.g. 2401.09350) |
id |
string | no | 12β14 chars β same arXiv ID, alternate field |
chunk-id |
int64 | no | 0β936 β position of this chunk within the paper |
chunk |
string | no | 401β2 020 chars β the actual text passage |
title |
string | no | 8β162 chars |
summary |
string | no | 228β1 920 chars β full abstract |
source |
string | no | 31 chars β PDF URL (http://arxiv.org/pdf/{id}) |
authors |
string | no | 7β6 970 chars β comma-separated |
categories |
string | no | 5β107 chars β space-separated arXiv tags |
primary_category |
string | no | 5β17 chars β e.g. cs.LG |
published |
string | no | 8 chars β YYYYMMDD |
updated |
string | no | 8 chars β YYYYMMDD |
comment |
string | yes | 4β398 chars β nullable |
journal_ref |
string | yes | 8β194 chars β nullable |
references |
list[{id}] | no | Nested list of {"id": "..."} objects |
Re-ingestion After Fixes
After applying the chunk-id field fix and expanding metadata,
clear existing data and re-ingest:
# Remove only the ingested content (preserves users/roles/logs)
docker exec agentic-rag-db-1 psql -U postgres -d agentic_rag -c "
TRUNCATE chunks, role_document_access, documents RESTART IDENTITY CASCADE;"
# Re-ingest with correct field mapping
docker-compose run --rm backend python -m scripts.ingest_data --limit 5000
After fixing chunk-id, a --limit 5000 run will produce ~5 000 unique chunks
(instead of 55) because each document now contributes multiple chunk positions.