ADOHAHA123
/

test-sft

Safetensors

qwen3

Model card Files Files and versions

xet

Community

ADOHAHA123 commited on Jan 16

Commit

42ecfc6

verified ·

1 Parent(s): c5a870c

Update README.md

Browse files

Files changed (1) hide show

README.md +1 -295

README.md CHANGED Viewed

@@ -1,291 +1,3 @@
----
-language:
-  - zh
-  - en
-license: apache-2.0
-library_name: transformers
-pipeline_tag: text-generation
-tags:
-  - roleplay
-  - dialogue
-  - multi-turn
-  - qwen
-  - sft
-  - chat
-base_model: Qwen/Qwen3-32B
----
-# HER-SFT-Qwen-32B
-<p align="center">
-  <a href="https://arxiv.org/abs/xxxx.xxxxx"><img src="https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv" alt="Paper"></a>
-  <a href="https://huggingface.co/datasets/ADOHAHA123/HER-Dataset"><img src="https://img.shields.io/badge/🤗%20Dataset-HER--Dataset-yellow" alt="Dataset"></a>
-  <a href="https://huggingface.co/ADOHAHA123/HER-Qwen3-32B"><img src="https://img.shields.io/badge/🤗%20Model-HER--RL-blue" alt="HER-RL"></a>
-  <a href="https://huggingface.co/ADOHAHA123/HER-SFT-Qwen3-32B"><img src="https://img.shields.io/badge/🤗%20Model-HER--SFT-green" alt="HER-SFT"></a>
-  <a href="https://github.com/your-username/HER"><img src="https://img.shields.io/badge/GitHub-Code-black?logo=github" alt="GitHub"></a>
-</p>
-HER-SFT (Human Emulation Reasoning - Supervised Fine-Tuning) is a role-playing language model built upon Qwen-32B base model. This is the supervised fine-tuned version of HER, trained on reasoning-augmented role-playing data constructed through reverse engineering synthesis.
-HER-SFT serves as the foundation for HER-RL (the reinforcement learning enhanced version) and demonstrates strong role-playing capabilities through **Dual-layer Thinking**:
-- **System Thinking** (third-person): LLM's meta-level planning on how to portray the character
-- **Role Thinking** (first-person): Character's inner thoughts and cognitive processes
-This model achieves competitive performance on role-playing benchmarks, with HER-RL further improving upon it through preference-aligned reinforcement learning.
-## Model Information
-- **Base Model**: Qwen-32B
-- **Training Method**: Supervised Fine-Tuning (SFT)
-- **Training Data**: Reasoning-augmented role-playing dialogues with dual-layer thinking
-- **Model Size**: 32B parameters
-- **Enhanced Version**: [HER-RL](../her-rl-qwen-32b) (RL-enhanced version available)
-## Key Features
-Our model generates responses with rich, interleaved structure:
-- `<system_thinking>`: Third-person analysis of how to portray the role
-- `<role_thinking>`: Character's inner thoughts (invisible to others)
-- `<role_action>`: Character's physical actions and expressions
-- Speech: Natural dialogue text
-This hierarchical approach enables more nuanced and authentic character portrayal.
-## How to Use
-### Quick Start: Interactive Chat Demo
-The easiest way to try the model is using our interactive chat demo:
-```bash
-cd chat_demo
-python chat_demo.py
-```
-This will start an interactive session where you can:
-1. Choose a scenario from classic literature (Pride and Prejudice, The Great Gatsby, etc.)
-2. Select which character the AI should play
-3. Select which character you want to play
-4. Start chatting with the AI character!
-**Demo Options:**
-```bash
-# Show the model's reasoning process (system thinking)
-python chat_demo.py --show-think
-# Show character's inner thoughts (role thinking)
-python chat_demo.py --show-rolethink
-# Directly specify scenario and character
-python chat_demo.py --scenario 0 --character 1
-# Use simple built-in scenarios
-python chat_demo.py --simple
-```
-**Chat Commands:**
-- `quit` / `exit` / `q` - Exit the chat
-- `clear` - Clear conversation history
-- `history` - View conversation history
-- `prompt` - View the full prompt
-See `chat_demo/README.md` for detailed instructions.
-### Programmatic Usage
-```python
-from transformers import AutoModelForCausalLM, AutoTokenizer
-model_name = "your-username/her-sft-qwen-32b"
-tokenizer = AutoTokenizer.from_pretrained(model_name)
-model = AutoModelForCausalLM.from_pretrained(
-    model_name,
-    torch_dtype="auto",
-    device_map="auto"
-)
-# Example: Role-playing as Elizabeth Bennet from Pride and Prejudice
-system_prompt = """You are Elizabeth Bennet from Pride and Prejudice.
-===Elizabeth Bennet's Profile===
-The protagonist, intelligent and strong-willed. Quick-witted with a playful sense of humor. Values honesty and integrity. Maintains composure under pressure while harboring deep emotions beneath the surface.
-Background: Second of five daughters in the Bennet family. Known for her intelligence, independence, and refusal to conform to societal expectations.
-Personality: Quick-witted with a playful sense of humor. Values honesty and integrity. Maintains composure under pressure.
-===Current Scenario===
-The scene is set in Mr. Bennet's private study. Elizabeth has been summoned unexpectedly after Lady Catherine's confrontational visit, where she refused to promise not to marry Mr. Darcy. The tension is palpable as Mr. Bennet holds a mysterious letter.
-===Output Format===
-Your output should follow this structure:
-1. System Thinking: Wrapped in <system_thinking></system_thinking> tags - third-person analysis of how to portray the role
-2. Role-play Response: Including <role_thinking> for inner thoughts, <role_action> for actions, and plain text for speech"""
-user_input = "Well, my dear Lizzy, I trust you are not too greatly troubled by recent events?"
-messages = [
-    {"role": "system", "content": system_prompt},
-    {"role": "user", "content": user_input}
-]
-text = tokenizer.apply_chat_template(
-    messages,
-    tokenize=False,
-    add_generation_prompt=True
-)
-model_inputs = tokenizer([text], return_tensors="pt").to(model.device)
-generated_ids = model.generate(
-    **model_inputs,
-    max_new_tokens=512,
-    temperature=0.8,
-    top_p=0.9
-)
-generated_ids = [
-    output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
-]
-response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
-print(response)
-```
-## Framework Overview
-<p align="center">
-  <img src="figure2github.png" alt="HER Framework" width="100%">
-</p>
-<p align="center">
-  <em>HER Framework: Dual-layer Thinking for Cognitive-Level Persona Simulation</em>
-</p>
-## Training Methodology
-HER-SFT employs a comprehensive training pipeline:
-1. **Dual-layer Thinking**: Separates hidden third-person system thinking (how the LLM plans to portray the character) from first-person role thinking (the character's actual inner thoughts). This dual-layer structure enables more authentic and cognitively grounded character simulation.
-2. **Reverse Engineering Data Synthesis**: We curate reasoning-augmented role-playing data through a three-stage reverse synthesis pipeline:
-   - Stage 1: Role thinking enhancement - Enriching characters' psychological activities
-   - Stage 2: Pattern diversification - Creating diverse response patterns
-   - Stage 3: System thinking generation - Adding meta-level planning analysis
-3. **Multi-layered Response Generation**: Models learn to generate system thinking (planning), role thinking (inner thoughts), role actions, and speech in an interleaved manner, avoiding monotonous patterns.
-## Performance
-### Main Leaderboard Results
-| Rank | Model | CoSER Avg | CoSER SC | CoSER AN | CoSER CF | CoSER SQ | MiniMax Avg | MiniMax Worlds (50%) | MiniMax Stories (25%) | MiniMax Pref (25%) | 95% CI |
-|------|-------|-----------|----------|----------|----------|----------|-------------|----------------------|----------------------|--------------------|---------|
-| 1 | Claude-4.5-Opus | **62.43** | 63.74 | **64.28** | 58.45 | 63.24 | 76.62 | 67.23 | 82.10 | 89.90 | [75.5, 77.7] |
-| 2 | Gemini-3-Pro | 61.80 | **65.95** | 60.42 | **58.34** | 62.49 | 75.60 | 62.72 | 83.87 | 93.08 | [74.5, 76.7] |
-| 3 | GPT-5.1 | 61.10 | 64.95 | 53.99 | 60.13 | 65.35 | 80.63 | 76.62 | 72.21 | 97.05 | [79.6, 81.6] |
-| 4 | Gemini-2.5-Pro | 60.68 | 61.05 | 60.80 | 57.48 | 63.40 | 68.23 | 52.36 | 82.11 | 86.08 | [67.1, 69.3] |
-| 5 | DeepSeek-v3.2 | 58.68 | 55.85 | 57.07 | 57.44 | 64.35 | 60.27 | 45.81 | 66.64 | 82.83 | [59.2, 61.4] |
-| 6 | MiniMax-M2-RP | 57.30 | 60.03 | 50.11 | 49.30 | **69.77** | **84.65** | **80.55** | 79.97 | **97.51** | [83.6, 85.7] |
-| 7 | DeepSeek-v3.1 | 53.50 | 50.15 | 53.18 | 53.93 | 56.72 | 64.22 | 51.11 | 66.45 | 88.21 | [62.9, 65.5] |
-| 8 | HER-RL | 53.12 | 54.33 | 47.26 | 52.78 | 58.12 | 65.73 | 59.13 | 57.74 | 86.90 | [63.0, 68.4] |
-| **9** | **HER-SFT (this model)** | **50.92** | **50.52** | **45.99** | **49.78** | **57.37** | **58.44** | **47.29** | **52.78** | **86.40** | **[56.5, 60.4]** |
-| 10 | Grok-4.1-Fast | 47.40 | 49.21 | 47.57 | 42.64 | 50.17 | 48.47 | 29.87 | 47.51 | 86.64 | [47.4, 49.5] |
-| 11 | Claude-4.5-Sonnet | 45.21 | 47.18 | 36.02 | 47.55 | 50.09 | 69.35 | 55.72 | 75.66 | 90.28 | [68.2, 70.5] |
-| 12 | Claude-3.7-Think | 39.73 | 44.84 | 31.00 | 42.45 | 40.65 | 61.25 | 50.66 | 59.53 | 84.15 | [58.5, 64.0] |
-| 13 | CoSER-70B | 35.95 | 35.05 | 31.16 | 32.28 | 45.33 | 45.38 | 34.32 | 30.32 | 82.58 | [43.5, 47.2] |
-| 14 | GPT-5-Mini | 32.97 | 38.10 | 24.60 | 27.20 | 42.00 | 57.63 | 43.32 | 50.11 | 93.78 | [55.9, 59.3] |
-| 15 | GPT-4o-240806 | 27.69 | 34.00 | 14.90 | 22.90 | 38.90 | 66.39 | 64.96 | 46.23 | 89.40 | [64.1, 68.7] |
-| 16 | GPT-OSS-120B | 26.12 | 32.80 | 14.80 | 21.50 | 35.40 | 60.72 | 47.27 | 56.65 | 91.71 | [58.0, 63.4] |
-| 17 | Qwen3-32B | 22.86 | 30.56 | 19.61 | 15.52 | 30.56 | 50.76 | 40.38 | 32.82 | 89.48 | [48.4, 53.2] |
-**CoSER Benchmark**: Evaluates role-playing quality on 0-100 scale across four dimensions:
-- **SC** (Story Consistency): Narrative coherence and plot continuity
-- **AN** (Anthropomorphism): Human-like behavior and natural expression
-- **CF** (Character Fidelity): Consistency with character profile and traits
-- **SQ** (Storyline Quality): Overall narrative quality and engagement
-**MiniMax Role-Play Bench**: Comprehensive evaluation on 0-100 scale:
-- **Worlds** (50%): Basic text quality and world-building
-- **Stories** (25%): Diversity and content logic
-- **Preferences** (25%): Interaction quality and user satisfaction
-**HER-SFT** achieves **rank 9** overall, demonstrating strong baseline performance:
-- **Significant improvement over base model**: +28.06 CoSER average vs Qwen3-32B baseline (+122.8% relative improvement)
-- **Strong preference scores (86.40)**: Competitive interaction quality, nearly matching HER-RL (86.90)
-- **Solid foundation for RL**: HER-RL builds upon this SFT model to achieve further gains (+2.20 CoSER average)
-The supervised fine-tuning establishes a strong baseline, with reinforcement learning providing incremental improvements in character consistency and narrative coherence.
-## Example Use Case
-### Elizabeth Bennet from Pride and Prejudice
-**System Prompt** (abbreviated):
-```
-You are Elizabeth Bennet from Pride and Prejudice. The protagonist, intelligent and strong-willed...
-===Current Scenario===
-The scene is set in Mr. Bennet's private study. Elizabeth has been summoned after Lady Catherine's visit...
-```
-**User Input**:
-```
-Well, my dear Lizzy, I trust you are not too greatly troubled by recent events?
-```
-**Model Output**:
-```
-<system_thinking>
-I need to generate a response for Elizabeth Bennet at the beginning of this scene.
-Context Analysis: Mr. Bennet has just spoken, picking up a letter and asking if I am "greatly troubled" by recent events. His tone is sardonic, yet there is underlying tension. The scenario implies the aftermath of Lady Catherine's visit.
-Character State (Elizabeth):
-- Internal: Feeling a mix of relief (that I stood my ground) and anxiety (about the consequences)
-- External: Need to appear composed, cannot let him see me crumble
-Goal: Answer honestly but carefully. Convey that I am holding up well.
-Plan:
-- Role Thinking: Reflect on the complexity of feelings
-- Role Action: Use physical action to show composure
-- Speech: Direct answer showing resilience
-</system_thinking>
-Elizabeth Bennet: <role_thinking>His tone is light, but the air feels heavy. I cannot let him see how much Lady Catherine's intrusion still stings, nor how anxious I am about what he knows</role_thinking><role_action>takes a steadying breath, smoothing the folds of her dress to mask a tremor</role_action>I believe I can manage, Father. Though I must admit, I am curious about what this letter contains.
-```
-## Training Data
-The model is trained on a comprehensive role-playing dataset that includes:
-- Authentic multi-turn, multi-character dialogues from classic literature
-- Character profiles with detailed background information
-- Scenario descriptions and conversation contexts
-- Rich annotations including system thinking, role thinking, and role actions
-- Diverse response patterns avoiding monotonous structures
-## Ethical Considerations
-We have conducted safety checks on the training dataset and implemented safeguards. However, users should be aware that:
-- The models may generate content that reflects biases present in the training data
-- Role-playing as certain characters might involve generating content with specific personality traits or behaviors
-- Users should implement appropriate content filtering when deploying these models in production applications
-- The models include safety evaluation dimensions to minimize harmful outputs
-## Model Comparison
-| Feature | HER-SFT (this model) | HER-RL |
-|---------|---------------------|---------|
-| Training Method | Supervised Fine-Tuning | SFT + Reinforcement Learning |
-| CoSER Average | 50.92 | 53.12 (+2.20) |
-| MiniMax Average | 58.44 | 65.73 (+7.29) |
-| Character Consistency | Good | Better (+2.00 CF) |
-| Interaction Quality | 86.40 Pref | 86.90 Pref (+0.50) |
-| Best Use Case | General role-playing | Preference-aligned interactions |
 ## Citation
 If you use HER-SFT in your research, please cite our paper:
@@ -303,10 +15,4 @@ If you use HER-SFT in your research, please cite our paper:
 Apache-2.0
-## Acknowledgments
-This model is based on Qwen-32B developed by Alibaba Cloud. We thank the Qwen team for their excellent base model.
-## Related Models
-- **[HER-RL](../her-rl-qwen-32b)**: Enhanced version with reinforcement learning for improved character consistency and preference alignment

 ## Citation
 If you use HER-SFT in your research, please cite our paper:
 Apache-2.0
+## Acknowledgments