YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Model Card for moelora_qkvow2
Model Description
moelora_qkvow2 is a task-aware Mixture-of-Experts LoRA (MoELoRA) adapter designed for domain-specialized large language models, with a particular focus on medical multi-department reasoning.
This adapter introduces a two-level expert hierarchy:
- Task-level parameter adaptation via a mixture of LoRA experts (MoELoRA).
- Token-level computation routing via the original Mixture-of-Experts (MoE) mechanism in the base model.
MoELoRA-QKVOW2 injects task-conditioned LoRA experts into both:
- Attention projections (
q_proj,k_proj,v_proj,o_proj) - Feed-forward network (FFN) layers (
w1,w2,w3, corresponding to SwiGLU)
This configuration provides higher expressive capacity than attention-only LoRA variants while remaining parameter-efficient.
Architecture Overview
- Adapter type: MoELoRA (Mixture-of-Experts LoRA)
- Base model:
SNOWTEAM/sft_medico-mistral
(derived frommistralai/Mixtral-8x7B-Instruct-v0.1) - Backbone architecture: Mixtral (MoE Transformer)
LoRA Expert Configuration
| Component | Value | Description |
|---|---|---|
| Number of LoRA experts | 8 | Size of the shared LoRA expert pool. Each expert represents a reusable low-rank adaptation component that can be combined across different tasks. |
| Number of tasks (departments) | 16 | Number of distinct task / department conditions supported by the model. Each task is represented by a dedicated task embedding that controls expert selection. |
| Task embedding dimension | 64 | Dimensionality of the task embedding used to compute gating weights over LoRA experts. |
| LoRA rank (r) | 16 | Rank of the low-rank matrices used in each LoRA expert, controlling adaptation capacity. |
| LoRA alpha | 32 | Scaling factor applied to LoRA updates, balancing the contribution of adaptation parameters relative to the frozen base model. |
| LoRA dropout | 0.1 | Dropout rate applied to LoRA modules during training to improve robustness and prevent overfitting. |
Target Modules
LoRA experts are injected into the following linear layers:
- Attention:
q_proj,k_proj,v_proj,o_proj - FFN (SwiGLU):
w1,w2,w3
How MoELoRA Works
MoELoRA introduces task-conditioned parameter modulation on top of a MoE backbone.
- For each task (e.g., medical department), a task embedding is used to compute gating weights over multiple LoRA experts.
- The resulting task-aware LoRA parameter update is computed once per task and remains fixed for all tokens in the input.
- During inference, each token still follows the original Mixtral MoE routing to select FFN experts.
- The final computation combines:
- Token-selected MoE expert weights, and
- Task-selected LoRA parameter updates.
This design cleanly separates:
- Task-level specialization (MoELoRA), and
- Token-level computation routing (Mixtral MoE).
Intended Use
MoELoRA-QKVOW2 is intended for:
- Medical or domain-specific LLMs requiring multi-department specialization
- Parameter-efficient adaptation of large MoE models
- Research on conditional parameter modulation and expert-based adaptation
This adapter is not a standalone model
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = AutoModelForCausalLM.from_pretrained(
"SNOWTEAM/sft_medico-mistral",
torch_dtype="auto",
device_map="auto"
)
model = PeftModel.from_pretrained(
base_model,
"SNOWTEAM/moelora_qkvow2"
)
tokenizer = AutoTokenizer.from_pretrained("SNOWTEAM/sft_medico-mistral")