You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Model Card for moelora_qkvow2

Model Description

moelora_qkvow2 is a task-aware Mixture-of-Experts LoRA (MoELoRA) adapter designed for domain-specialized large language models, with a particular focus on medical multi-department reasoning.

This adapter introduces a two-level expert hierarchy:

  1. Task-level parameter adaptation via a mixture of LoRA experts (MoELoRA).
  2. Token-level computation routing via the original Mixture-of-Experts (MoE) mechanism in the base model.

MoELoRA-QKVOW2 injects task-conditioned LoRA experts into both:

  • Attention projections (q_proj, k_proj, v_proj, o_proj)
  • Feed-forward network (FFN) layers (w1, w2, w3, corresponding to SwiGLU)

This configuration provides higher expressive capacity than attention-only LoRA variants while remaining parameter-efficient.


Architecture Overview

  • Adapter type: MoELoRA (Mixture-of-Experts LoRA)
  • Base model: SNOWTEAM/sft_medico-mistral
    (derived from mistralai/Mixtral-8x7B-Instruct-v0.1)
  • Backbone architecture: Mixtral (MoE Transformer)

LoRA Expert Configuration

Component Value Description
Number of LoRA experts 8 Size of the shared LoRA expert pool. Each expert represents a reusable low-rank adaptation component that can be combined across different tasks.
Number of tasks (departments) 16 Number of distinct task / department conditions supported by the model. Each task is represented by a dedicated task embedding that controls expert selection.
Task embedding dimension 64 Dimensionality of the task embedding used to compute gating weights over LoRA experts.
LoRA rank (r) 16 Rank of the low-rank matrices used in each LoRA expert, controlling adaptation capacity.
LoRA alpha 32 Scaling factor applied to LoRA updates, balancing the contribution of adaptation parameters relative to the frozen base model.
LoRA dropout 0.1 Dropout rate applied to LoRA modules during training to improve robustness and prevent overfitting.

Target Modules

LoRA experts are injected into the following linear layers:

  • Attention: q_proj, k_proj, v_proj, o_proj
  • FFN (SwiGLU): w1, w2, w3

How MoELoRA Works

MoELoRA introduces task-conditioned parameter modulation on top of a MoE backbone.

  • For each task (e.g., medical department), a task embedding is used to compute gating weights over multiple LoRA experts.
  • The resulting task-aware LoRA parameter update is computed once per task and remains fixed for all tokens in the input.
  • During inference, each token still follows the original Mixtral MoE routing to select FFN experts.
  • The final computation combines:
    • Token-selected MoE expert weights, and
    • Task-selected LoRA parameter updates.

This design cleanly separates:

  • Task-level specialization (MoELoRA), and
  • Token-level computation routing (Mixtral MoE).

Intended Use

MoELoRA-QKVOW2 is intended for:

  • Medical or domain-specific LLMs requiring multi-department specialization
  • Parameter-efficient adaptation of large MoE models
  • Research on conditional parameter modulation and expert-based adaptation

This adapter is not a standalone model

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained(
    "SNOWTEAM/sft_medico-mistral",
    torch_dtype="auto",
    device_map="auto"
)

model = PeftModel.from_pretrained(
    base_model,
    "SNOWTEAM/moelora_qkvow2"
)

tokenizer = AutoTokenizer.from_pretrained("SNOWTEAM/sft_medico-mistral")


Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support