Chekhov0919 commited on
Commit
3283bf4
·
verified ·
1 Parent(s): 42cdc42

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +79 -0
README.md CHANGED
@@ -1,3 +1,82 @@
1
  ---
2
  license: mit
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: mit
3
+ language:
4
+ - en
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - mental-health
8
+ - counseling
9
+ - lora
10
+ - peft
11
+ - diffusion-language-model
12
+ - LLaDA
13
  ---
14
+
15
+ # BiGraph-Diffuse
16
+
17
+ LoRA adapter for **BiGraph-Diffuse**, a retrieval-augmented diffusion language model for empathetic mental health counseling.
18
+
19
+ ## Model Overview
20
+
21
+ This is a [LoRA](https://arxiv.org/abs/2106.09685) adapter fine-tuned on **LLaDA-8B-Instruct**, a discrete diffusion language model. The adapter is trained on counseling dialogues to generate empathetic, psychologically grounded counselor responses.
22
+
23
+ ### Architecture
24
+
25
+ - **Base Model**: LLaDA-8B-Instruct (discrete diffusion LM)
26
+ - **Adapter**: LoRA (rank=32, alpha=64, dropout=0.1)
27
+ - **Target Modules**: `q_proj`, `k_proj`, `v_proj`, `o_proj`
28
+ - **Task**: Causal language modeling with masked diffusion loss
29
+
30
+ Full architecture includes **BiGraph-RAG**, a bipartite graph retrieval system that augments generation with relevant psychological knowledge. Code available at the [GitHub repo](https://github.com/Chekhov0919/BiGraph-Diffuse).
31
+
32
+ ## Usage
33
+
34
+ ```python
35
+ import torch
36
+ from transformers import AutoModelForCausalLM, AutoTokenizer
37
+ from peft import PeftModel
38
+
39
+ # Load base model
40
+ base_model_path = "path/to/LLaDA-8B-Instruct"
41
+ tokenizer = AutoTokenizer.from_pretrained(base_model_path, trust_remote_code=True)
42
+ tokenizer.padding_side = "left"
43
+
44
+ base_model = AutoModelForCausalLM.from_pretrained(
45
+ base_model_path,
46
+ torch_dtype=torch.bfloat16,
47
+ device_map="auto",
48
+ trust_remote_code=True
49
+ )
50
+
51
+ # Load LoRA adapter
52
+ model = PeftModel.from_pretrained(base_model, "Chekhov0919/BiGraph-Diffuse")
53
+ model.eval()
54
+
55
+ # Generate with diffusion
56
+ # See GitHub repo for full inference code with BiGraph-RAG integration
57
+ ```
58
+
59
+ For the complete inference pipeline with BiGraph-RAG retrieval, refer to the [GitHub repository](https://github.com/Chekhov0919/BiGraph-Diffuse).
60
+
61
+ ## Training
62
+
63
+ | Setting | Value |
64
+ |---------|-------|
65
+ | Base Model | LLaDA-8B-Instruct |
66
+ | Dataset | CPsyCounD (counseling dialogues) |
67
+ | LoRA rank | 32 |
68
+ | LoRA alpha | 64 |
69
+ | LoRA dropout | 0.1 |
70
+ | Batch size | 2 × 32 (gradient accumulation) |
71
+ | Learning rate | 3e-5 |
72
+ | Epochs | 5 |
73
+ | LR scheduler | Cosine |
74
+ | Mask token ID | 126336 |
75
+
76
+ ## Citation
77
+
78
+ Please stay tuned — citation information will be added upon publication.
79
+
80
+ ## License
81
+
82
+ MIT