Spaces:
Sleeping
Download docs/DATA_SOURCES.md from shield137/shockmap-api: direct link, hf CLI and curl.
- Browser
- Download file 8.26 kB
-
https://huggingface.co/spaces/shield137/shockmap-api/resolve/main/docs/DATA_SOURCES.md
- Command line
-
hf download hf://spaces/shield137/shockmap-api/docs/DATA_SOURCES.md
-
curl -L -o DATA_SOURCES.md https://huggingface.co/spaces/shield137/shockmap-api/resolve/main/docs/DATA_SOURCES.md
Data Sources
Everything in this doc is what you need to collect manually. The code is built to absorb it without modification — drop the file in the right place, restart the backend, done.
1. Chinese Provincial Environmental Notices (HIGHEST PRIORITY)
This is the differentiator. Everything else is supporting evidence.
Hebei Province EPB
- Source: https://hbsthjt.hebei.gov.cn/
- Section: 公示公告 → 行政处罚公告 (Administrative Penalty Notices)
- What to look for: Notices mentioning 制药 (pharma), 化工 (chemical), 停产 (halt production), 整改 (rectification)
- Drop into:
data/seed/epb_notices.json - Schema:
{
"id": "hebei_2026_04_15",
"source_url": "https://hbsthjt.hebei.gov.cn/notice/...",
"scraped_at": "2026-04-15T10:00:00Z",
"factory_name_zh": "石家庄某制药企业",
"factory_name_en": "Shijiazhuang Pharmaceutical Enterprise",
"industry": "API manufacturing",
"violation_type": "环保不达标",
"violation_type_en": "Environmental compliance failure",
"severity": "HIGH",
"duration_days_estimate": 30,
"linked_apis": ["para_aminophenol"],
"raw_text_zh": "...full notice text in Chinese...",
"gemini_translation": "...full English translation..."
}
To run the scraper that does this automatically:
cd ingestion
playwright install chromium
python scrape_hebei_epb.py
# Output appended to data/seed/epb_notices.json
Other Chinese provinces to consider
- Jiangsu: http://hbj.jiangsu.gov.cn/
- Zhejiang: https://sthjt.zj.gov.cn/
- Shandong: http://sthjt.shandong.gov.cn/
- Hubei: http://sthjt.hubei.gov.cn/
The scraper is parameterizable — just change BASE_URL and LISTING_PATH at the top of scrape_hebei_epb.py.
2. FDA Import Alerts on Chinese Pharma Facilities
- Source: https://www.accessdata.fda.gov/cms_ia/ialist.html
- Filter: Country = China, Industry = 66 (Pharmaceuticals) or 56 (Cosmetics) or 53 (Drugs)
- What to look for: OAI status (Official Action Indicated), Refused For Import, Detention Without Physical Examination
- Drop into:
data/seed/fda_alerts.json - Schema:
{
"id": "fda_66-40_2026-03-12",
"alert_number": "66-40",
"publish_date": "2026-03-12",
"firm_name": "Hebei Welcome Pharmaceutical Co.",
"city": "Shijiazhuang",
"products": ["Penicillin G Potassium API"],
"linked_apis": ["penicillin_g_potassium"],
"reason": "Data integrity violations during pre-approval inspection",
"source_url": "https://www.accessdata.fda.gov/cms_ia/importalert_..."
}
To run the scraper:
cd ingestion
python scrape_fda_alerts.py
# Output: data/seed/fda_alerts.json
This requires no auth — runs in 60 seconds.
3. DGCI&S Trade Data (India's official import statistics)
- Source: https://commerce.gov.in/eidb/
- What to download: Monthly import data by HS code
- HS 29 — Organic chemicals (covers most APIs)
- HS 30 — Pharmaceutical products (finished formulations)
- HS 2941 — Antibiotics specifically
- Format: CSV download (sometimes Excel — convert to CSV)
- Drop into:
data/seed/trade_data.csv - Required columns:
month,api_id,hs_code,country_origin,import_value_usd,import_quantity_kg
2024-01,para_aminophenol,29222910,China,42500000,1180000
2024-02,para_aminophenol,29222910,China,38000000,1050000
Mapping HS codes to API IDs: Many APIs share HS codes (commodity-level). The mapping is in data/seed/hs_code_mapping.json (placeholder). You'll need to research which 8-digit HS codes correspond to which API. Alternative: just track at the HS code level and don't try to map every drug.
Easier alternative: Pharmexcil publishes monthly digest PDFs with import volumes already broken down by chemical. Source: https://pharmexcil.com/
4. NLEM 2022 Drug List (to expand from 20 → all 800+ drugs)
- Source: https://cdsco.gov.in/opencms/opencms/en/NLEM-2022/
- Format: PDF
- Action: For each drug not in
data/seed/drugs.json, add an entry:
{
"id": "drug_id_lowercase",
"name": "Display Name",
"generic_name": "Generic Name",
"nlem_tier": "TIER_1",
"patient_population_estimate": 50000000,
"primary_apis": ["api_id_1"],
"has_substitute": false,
"therapeutic_class": "antibiotic"
}
Tier mapping: NLEM 2022 doesn't have explicit tiers. Approximate by category:
- TIER_1: critical / life-saving (insulin, antibiotics, paracetamol, anti-TB)
- TIER_2: chronic disease management (statins, antihypertensives)
- TIER_3: specialty / less common
Each entry takes ~3 minutes. Top 50 drugs would be ~2.5 hours of work and get you to demo-quality data density.
5. Historical Disruption Events (the GNN training labels)
- Source: News archives + WHO drug shortage database + FDA shortage database
- Drop into:
data/seed/historical_disruptions.json(already has 5 events; add more) - Schema:
{
"date": "2024-01-15",
"source_event": "Hebei pharma plant environmental inspection wave",
"province": "Hebei",
"severity": 0.7,
"duration_days": 21,
"affected_drugs": ["paracetamol", "ibuprofen"],
"lead_time_days": 23,
"indian_consumer_price_impact_pct": 100.0,
"citation_url": "https://news-source-url"
}
This is what makes the GNN training real instead of circular. With 20+ events with measured indian_consumer_price_impact_pct as labels, the GNN learns actual market response patterns rather than just regurgitating edge weights.
Where to find them:
- LiveMint, Economic Times, Business Standard archives — search "API shortage India"
- WHO drug shortage database: https://list.essentialmeds.org/
- FDA Drug Shortage database: https://www.accessdata.fda.gov/scripts/drugshortages/
Aim for 20 events spanning 2018-2026.
6. Policy Snippets (RAG quality)
- Drop into:
data/seed/policy_snippets.json(already has 10; needs ~20 more) - What to add:
- ORF report — full text from "Securing India's Pharmaceutical Supply Chain" (Nov 2025)
- NLEM 2022 preamble — first 3 paragraphs
- NITI Aayog PLI scheme document — sections on bulk drugs and KSMs
- Department of Pharmaceuticals annual reports — sections on import dependency
- WHO essential medicines criteria — sections on supply security
Format: Each entry is a 2-4 sentence chunk (Qdrant indexes these for semantic search).
{
"id": "orf_2025_05",
"source": "ORF Research Brief, Nov 2025, p.12",
"source_url": "https://orfonline.org/...",
"text": "The fragility of the global paracetamol supply chain was exposed during the 2024 Hebei environmental inspections...",
"keywords": ["paracetamol", "Hebei", "fragility"]
}
Quality > quantity. 20 well-chosen snippets give better RAG answers than 200 random paragraphs.
7. Live source URLs (verifier badge)
For every alert in alerts.json, the source_url field MUST resolve to a real page. Currently many are fake (e.g. reuters.com/business/pharma/jiangsu-industrial-accident-impacts-pharma — 404).
Action: Replace fake URLs with real ones. If the news source has paywalled, link to archive.org snapshot:
https://web.archive.org/web/2024*/<original-url>
A judge clicking through and seeing real Chinese government text or real news is worth 30 minutes of slide content.
Refresh schedule (Phase 2)
For production, schedule the scrapers:
# Hebei EPB — every 6 hours
0 */6 * * * cd /app/ingestion && python scrape_hebei_epb.py
# FDA — daily at 3am UTC
0 3 * * * cd /app/ingestion && python scrape_fda_alerts.py
# DGCI&S — manual monthly (data only published monthly)
After scraper runs, hit POST /api/v1/ingest/refresh to make the backend re-read the JSON files without restart.
TL;DR — what to do if you have 1 hour
- Run
python scrape_fda_alerts.py— populates real FDA alerts (10 min, no setup). - Manually fetch 3-5 Hebei EPB notices, paste into
epb_notices.json(20 min). - Add 10 more drugs to
drugs.jsonfrom NLEM PDF (20 min). - Add 5 more historical disruptions with real news URLs (10 min).
That's enough to make the demo feel real.