Text Classification
Transformers
Safetensors
trained-from-scratch
deberta
sequence-classification
multiple-choice
question-answering
Instructions to use dahaludba/QSolver_Scratch with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use dahaludba/QSolver_Scratch with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="dahaludba/QSolver_Scratch")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("dahaludba/QSolver_Scratch", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 5,383 Bytes
41c6cfb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 | ---
license: mit
base_model: microsoft/deberta-v3-xsmall
library_name: transformers
pipeline_tag: text-classification
tags:
- trained-from-scratch
- deberta
- sequence-classification
- multiple-choice
- question-answering
metrics:
- map@3
- accuracy
- f1
---
# QSolver_Scratch
QSolver_Scratch is a lightweight, custom transformer model trained **completely from scratch** (random weight initialization via `.from_config()`) using a modified DeBERTa-v3 architecture.
It uses a question-option explosion technique that reformulates 5-choice multiple-choice questions into binary sequence classification tasks (`Question: {q}\nOption: {opt_val}`). Predictions across option pairs are grouped into 5-way logit arrays to evaluate choices using Mean Average Precision at 3 (MAP@3), Accuracy, Precision, Recall, and F1-score across 5-Fold Cross-Validation.
---
## Model Details
- **Model Name:** QSolver_Scratch
- **Repository ID:** `dahaludba/QSolver_Scratch`
- **Tokenizer:** `microsoft/deberta-v3-xsmall`
- **Training Strategy:** Trained from scratch (random initial weights, no pre-trained weights used)
- **Architecture Base:** DeBERTa-v3 (Custom Reduced Variant)
- **Task Type:** Binary Sequence Classification for Multiple-Choice Ranking
- **Maximum Sequence Length:** 256 tokens
- **Total Training Duration:** 1 hour 15 minutes
- **License:** MIT License
---
## Custom Architecture Specifications
To build a lightweight, fast-executing encoder model, the base `microsoft/deberta-v3-xsmall` architecture configuration was customized with reduced layer count and embedding dimensions:
- **Number of Hidden Layers (`num_hidden_layers`):** 4 (reduced from standard 12)
- **Hidden Layer Dimension (`hidden_size`):** 256
- **Attention Heads (`num_attention_heads`):** 4
- **Feed-Forward Intermediate Dimension (`intermediate_size`):** 1024
- **Pooler Hidden Dimension (`pooler_hidden_size`):** 256
- **Number of Output Labels (`num_labels`):** 1 (binary classification logit per prompt-option pair)
---
## Data Pipeline & Exploded Formatter
Each multiple-choice sample containing a prompt and 5 candidate choices (A, B, C, D, E) is exploded into 5 independent text instances:
```text
Text: "Question: {question_text}\nOption: {option_value}"
Label: 1.0 (if option is correct) | 0.0 (if option is incorrect)
```
During evaluation, outputs are reshaped into batches of size $(N, 5)$ where $N$ is the number of question groups. Logits are sorted in descending order to derive the top-3 predicted option ranks.
---
## Hyperparameters
| Hyperparameter | Value |
| :--- | :--- |
| Initialization Method | Random Weights (`AutoModelForSequenceClassification.from_config`) |
| Learning Rate | 1e-4 |
| LR Scheduler | Cosine Decay |
| Warmup Steps | 100 |
| Weight Decay | 0.01 |
| Per Device Train Batch Size | 64 |
| Per Device Eval Batch Size | 64 |
| Gradient Accumulation Steps | 1 |
| Epochs per Fold | 5 |
| Precision | Mixed Precision FP16 (`fp16=True`) |
| Primary Metric for Best Model | `eval_map@3` (`greater_is_better=True`) |
| Seed | 42 |
---
## Experiment Tracking & Cross-Validation Results
Training across all 5 folds completed in **1 hour and 15 minutes**. Individual run telemetry and evaluation metrics are recorded on Weights & Biases:
- **Fold 1:** [W&B Run zb6ldwgg](https://wandb.ai/24f2002963-dl-genai-project/24f2002963-t22026/runs/zb6ldwgg?nw=nwuser24f2002963)
- **Fold 2:** [W&B Run i6sht758](https://wandb.ai/24f2002963-dl-genai-project/24f2002963-t22026/runs/i6sht758?nw=nwuser24f2002963)
- **Fold 3:** [W&B Run d6zqwe0x](https://wandb.ai/24f2002963-dl-genai-project/24f2002963-t22026/runs/d6zqwe0x?nw=nwuser24f2002963)
- **Fold 4:** [W&B Run az2xrt2a](https://wandb.ai/24f2002963-dl-genai-project/24f2002963-t22026/runs/az2xrt2a?nw=nwuser24f2002963)
- **Fold 5:** [W&B Run hkzxclzi](https://wandb.ai/24f2002963-dl-genai-project/24f2002963-t22026/runs/hkzxclzi?nw=nwuser24f2002963)
---
## How to Load and Run Inference
Below is a Python snippet showing how to load a fold checkpoint and perform multi-choice scoring:
```python
import torch
import numpy as np
from transformers import AutoTokenizer, AutoModelForSequenceClassification
REPO_ID = "dahaludba/QSolver_Scratch"
SUBFOLDER = "fold_1"
TOKENIZER_NAME = "microsoft/deberta-v3-xsmall"
# Load Tokenizer and Fold Model
tokenizer = AutoTokenizer.from_pretrained(TOKENIZER_NAME)
model = AutoModelForSequenceClassification.from_pretrained(REPO_ID, subfolder=SUBFOLDER)
model.eval()
question = "What organelle is known as the powerhouse of the cell?"
options = [
"Nucleus",
"Mitochondria",
"Ribosome",
"Golgi Apparatus",
"Endoplasmic Reticulum"
]
# Format exploded inputs
input_texts = [f"Question: {question}\nOption: {opt}" for opt in options]
# Tokenize inputs
inputs = tokenizer(
input_texts,
padding=True,
truncation=True,
max_length=256,
return_tensors="pt"
)
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits.squeeze(-1).cpu().numpy()
# Rank option choices by predicted score
option_letters = ["A", "B", "C", "D", "E"]
top3_indices = np.argsort(-logits)[:3]
top3_predictions = [f"{option_letters[idx]} ({options[idx]})" for idx in top3_indices]
print("Top-3 Predicted Options:", top3_predictions)
```
---
## License
This project and all associated model weights are released under the **MIT License**. |