Instructions to use chenhaodev/patient-edu-qa with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use chenhaodev/patient-edu-qa with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf chenhaodev/patient-edu-qa:Q8_0 # Run inference directly in the terminal: llama cli -hf chenhaodev/patient-edu-qa:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf chenhaodev/patient-edu-qa:Q8_0 # Run inference directly in the terminal: llama cli -hf chenhaodev/patient-edu-qa:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf chenhaodev/patient-edu-qa:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf chenhaodev/patient-edu-qa:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf chenhaodev/patient-edu-qa:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf chenhaodev/patient-edu-qa:Q8_0
Use Docker
docker model run hf.co/chenhaodev/patient-edu-qa:Q8_0
- LM Studio
- Jan
- vLLM
How to use chenhaodev/patient-edu-qa with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "chenhaodev/patient-edu-qa" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "chenhaodev/patient-edu-qa", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/chenhaodev/patient-edu-qa:Q8_0
- Ollama
How to use chenhaodev/patient-edu-qa with Ollama:
ollama run hf.co/chenhaodev/patient-edu-qa:Q8_0
- Unsloth Desktop
- Pi
How to use chenhaodev/patient-edu-qa with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chenhaodev/patient-edu-qa:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "chenhaodev/patient-edu-qa:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use chenhaodev/patient-edu-qa with Docker Model Runner:
docker model run hf.co/chenhaodev/patient-edu-qa:Q8_0
- Lemonade
How to use chenhaodev/patient-edu-qa with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull chenhaodev/patient-edu-qa:Q8_0
Run and chat with the model
lemonade run user.patient-edu-qa-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use chenhaodev/patient-edu-qa with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chenhaodev/patient-edu-qa:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default chenhaodev/patient-edu-qa:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use chenhaodev/patient-edu-qa with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf chenhaodev/patient-edu-qa:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "chenhaodev/patient-edu-qa:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
patient-edu-qa · 患者教育问答 Harness
面向患者(非医生)的疾病科普问答系统。患者用自然语言(中文,口语化/碎片化均可)提问, 系统由 1-3B 小模型路由器 做结构化意图拆分与红旗判定,从 多病种 FAISS 索引 检索证据, 最终由 大模型(唯一生成出口) 输出连贯、统一风格的回答。
设计核心:先路由、再检索、最后由唯一的大模型生成。小模型不直接写用户可见的文字, 它只输出一个紧凑的结构化 JSON 计划。
用户输入 ──▶ 1-3B 路由器(纯分类)
· 并列拆分 sub_intents[] (每个 = 一个 RAG 寻址单元)
· mode = grounded | alert
· red_flags[](危险信号)
│
▼
多RAG:每个子意图选一个病种索引 → 检索证据(中文 query / 英文 UpToDate 语料,bge-m3 跨语言)
│
▼
大模型(唯一出口):grounded path 引用证据 / alert path 建议就医
1. 组件与架构
| 组件 | 职责 |
|---|---|
| 1-3B 路由器 | 纯分类:意图拆分 + mode + 红旗,不做生成 |
| 多 RAG | 每病种一个 FAISS 索引,跨语言检索(BAAI/bge-m3) |
| 大模型(唯一出口) | grounded 引用证据 / alert 建议就医 |
部署后端(DEPLOY)
DEPLOY=gpu(默认):unslothLoRA 路由器 +sentence-transformers(BAAI/bge-m3)嵌入,走 torch/GPU。DEPLOY=cpu:llama.cppGGUF 路由器 + GGUF 嵌入,纯 CPU 推理,无需 torch/GPU。
路由器后端(ROUTER_BACKEND)
rules:纯规则引擎,零权重,确定/可解释,开箱即用。model:纯微调后的 1-3B LoRA 路由器。hybrid(默认):mode/red_flags走规则(安全、权威);category/intent走微调后的小模型(路由更准),小模型失败时回退规则。
路由器输出 schema
{
"mode": "alert", // grounded | alert
"red_flags": [ // mode=alert 时逐条列出
{"type": "chronicity", "trigger": "duration:2周", "severity": "high"}
],
"sub_intents": [ // 并列子意图,每个 = 一个 RAG 寻址单元
{
"intent": "self_care", // 归属意图 taxonomy(决定 prompt 模版)
"entity": "headache",
"category": "brain-and-nerves", // 选定 FAISS 索引(与 data/rag/<category> 强一致)
"level": "basics",
"tmpl": "self_care",
"risk": "high"
}
]
}
两个生成路径(都在大模型内)
grounded:RAG 命中 → 大模型走证据引用路径(「据 UpToDate 患者教育」逐条引用)。alert:命中红旗(如 2 周慢性头痛)→ 大模型首句前置「建议就医」,自疗建议降为辅助,不强求 RAG 有证据。
2. 目录结构
patient-edu-qa/
├── harness/ # 可运行的系统骨架
│ ├── config.py # 路径 + 后端/API/超参配置(含 DEPLOY/GGUF 路径,均可用环境变量覆盖)
│ ├── router.py # 路由器:rules / model / hybrid / GGUF 后端(llama_cpp 惰性导入)
│ ├── rag.py # 多RAG检索器(每病种一个 FAISS 索引,bge-m3 / GGUF 嵌入)
│ ├── prompts.py # 意图→prompt模版 + grounded/alert 双路径
│ ├── llm.py # 大模型客户端(OpenAI 兼容,流式)
│ ├── pipeline.py # 编排:输入→route→检索→组装→生成
│ └── __main__.py # 交互式 CLI(REPL / 单句查询)
├── scripts/ # 构建管线
│ ├── extract_patient_education.py # 解析 UpToDate Patient Education → Q&A 语料
│ ├── red_flag_rules.py # 红旗规则引擎(病程/体征/人群)
│ ├── gen_router_data.py # curated 弱标注(规则+taxonomy)
│ ├── gen_router_augment.py # 合成告警/复合样本(补规则盲区)
│ ├── merge_router_data.py # 合并 → router SFT train/val
│ ├── build_multi_rag.py # 构建 30 病种 FAISS 索引
│ ├── train_router.py # 训练 1-3B LoRA 路由器
│ └── merge_lora_gguf.py # LoRA 合并回基座 → FP16 safetensors(GGUF 前置)
├── router-q8_0.gguf # CPU 部署路由器(GGUF,~1.83GB,仓库根目录)
├── data/
│ ├── patient_education_qa.jsonl # 问答块(原始解析)
│ ├── patient_questions_clean.jsonl # 清洗患者问题
│ ├── rag/<category>/{index.faiss,docs.json,meta.json} # 30 病种索引
│ └── router/router_sft_{train,val}.jsonl # 路由器 SFT 数据
└── README.md
嵌入用
BAAI/bge-m3(multilingual,公开模型)。其 GGUF 量化版(bge-m3-q8_0.gguf)未随本仓库上传, 可另行从 HF 下载或自行量化(见 §5 导出流程)。CPU 部署时把EMBED_GGUF_PATH指向本地bge-m3-q8_0.gguf。
3. 环境安装
# 底层依赖
uv pip install openai faiss-cpu sentence-transformers beautifulsoup4
# 训练/推理小模型路由器时才需要
uv pip install unsloth trl datasets transformers
# CPU/GGUF 部署(DEPLOY=cpu)时才需要
uv pip install llama-cpp-python
镜像:官方
huggingface.co在部分网络(如本沙箱)不可达,用export HF_ENDPOINT=https://hf-mirror.com。 GPU:uv run python -c "import torch; print(torch.cuda.is_available())"应输出True。
大模型 API(OpenAI 兼容,默认对接 vLLM):
export LLM_BASE_URL=http://...:8000/v1
export LLM_API_KEY=...
export LLM_MODEL=minimax
4. 运行 Harness
python -m harness # 交互式 REPL
python -m harness "什么是糖尿病,平时饮食要注意什么" # 单句查询
# 部署后端:默认 gpu;CPU 部署
DEPLOY=cpu \
ROUTER_GGUF_PATH=router-q8_0.gguf \
EMBED_GGUF_PATH=/path/to/bge-m3-q8_0.gguf \
python -m harness "我最近2周一直头疼,有什么好的免吃药的方法舒缓"
# 路由器后端切换(rules / model / hybrid)
ROUTER_BACKEND=model python -m harness "..."
降级兜底:若
llama_cpp未安装或 GGUF 缺失,路由自动回退到规则引擎,保证任何机器上都能跑通。
5. 构建管线(从零复现)
# 1) 解析 UpToDate Patient Education → Q&A 语料
python3 scripts/extract_patient_education.py
# 2) 红旗规则引擎(规则层)
python3 -c "from scripts.red_flag_rules import scan; print(scan('我最近2周一直头疼'))"
# 3) 生成路由器 SFT 数据
python3 scripts/gen_router_data.py
python3 scripts/gen_router_augment.py
python3 scripts/merge_router_data.py # → data/router/router_sft_{train,val}.jsonl
# 4) 训练 1-3B LoRA 路由器(基座 Qwen/Qwen3-1.7B,可用 ROUTER_BASE 覆盖)
python3 scripts/train_router.py # → checkpoints/.../final
# 5) 构建 30 病种多 RAG 索引
python3 scripts/build_multi_rag.py # → data/rag/<30病种>/
# 6) 导出 + 量化 CPU/GGUF 资产(可选,供无 GPU 机器部署)
# a. LoRA 合并回基座 → FP16 safetensors(脚本硬编码本地 LoRA/缓存路径,MERGE_OUT 指定输出)
MERGE_OUT=/tmp/router-merged python3 scripts/merge_lora_gguf.py
# b. safetensors → GGUF f16(llama.cpp 官方转换器)
PYTHONPATH=/path/to/llama.cpp/gguf-py \
python3 /path/to/llama.cpp/convert_hf_to_gguf.py /tmp/router-merged \
--outfile router-f16.gguf --outtype f16
# c. f16 → Q8_0(~1.83GB,近无损)
/path/to/llama.cpp/build/bin/llama-quantize router-f16.gguf router-q8_0.gguf Q8_0
# d. bge-m3 嵌入同理:bge-m3-f32.gguf → bge-m3-q8_0.gguf
6. 验证结果
Alert 路径(慢性头痛红旗,不套 RAG,直接建议就医)
PLAN mode=alert, red_flags=[chronicity/2周], sub_intents=[self_care/brain-and-nerves]
→ 大模型:首句提示「持续两周需重视、建议就医」,再给居家的非药物缓解建议
Grounded 路径(糖尿病定义+饮食,走 RAG 证据)
PLAN mode=grounded, intent=diet, category=diabetes
→ 检索 4 条证据(What is diabetes?/type 2 等),大模型分点作答并标来源
复合问题(多子意图)架构已支持
PLAN sub_intents=[self_care/brain-and-nerves, self_care/sleep]
→ 从两个索引各检索证据,合并后交给大模型
案例:Harness vs 裸大模型
- 过敏性鼻炎用药(grounded):裸大模型直接给具体药名、无来源、有个别非循证说法;Harness 检索到 UpToDate「Medications that may help symptoms」等证据,逐条引用;检索为空时正确拒答并建议就医而非编造。
- 持续 2 周头痛(alert):裸大模型把「2周」当普通场景,红旗只作文末附加提醒;Harness 首句前置「建议就医」,自疗建议降为辅助,再次强调面诊。
工程经验
- LoRA 的核心价值 = schema 对齐:无 LoRA 时 1.7B 输出自由 JSON(
{"intents":[...]})并长篇 thinking;微调后每个 query 严格输出约定的{mode, red_flags, sub_intents},可直接被 pipeline 解析。 - 为何默认 hybrid:1.7B 在少数类 alert 样本(如 180/2313)上对红旗判定不如确定性规则可靠(曾把「2周头疼」漏判为 grounded)。红旗是安全攸关且规则引擎已可解释地做对,故
mode/red_flags走规则、category/intent走小模型。 - 类别名与 RAG 目录名强一致:
asthma-allergyvsallergies-and-asthma曾导致检索永久为空,务必对齐并加缺索引兜底。
7. 已构建资产
| 资产 | 路径 | 规模 |
|---|---|---|
| 原始 Q&A 语料 | data/patient_education_qa.jsonl |
34,077 块 |
| 清洗患者问题 | data/patient_questions_clean.jsonl |
12,093 条 |
| 红旗规则引擎 | scripts/red_flag_rules.py |
5 类红旗,26 组用例 |
| 路由器 SFT 数据 | data/router/router_sft_{train,val}.jsonl |
train 2,313 / val 350 |
| 多病种 FAISS 索引 | data/rag/<30病种>/ |
30 索引 / 21,847 文档 |
| 路由器 LoRA | checkpoints/Qwen-Qwen3-1.7B-router-lora/final |
1-3B 路由模型 |
| CPU 路由器 GGUF | router-q8_0.gguf(仓库根目录) |
~1.83GB |
| 可运行骨架 | harness/ |
双路径 + hybrid 后端跑通 |
免责声明
本项目用于疾病科普与就医建议参考,不构成医疗诊断。若出现胸痛、呼吸困难、意识改变等急重症, 请立即就医或拨打急救电话。红旗规则为启发式,非完备医学词典;最终措辞由大模型负责任地建议就医。
- Downloads last month
- -
8-bit