Cross-prompt essay scoring checkpoints: ASAP 2.0

Seven-fold leave-one-prompt-out on ASAP 2.0: fold k holds out prompt P(k+1) and is trained only on the other six.

Subfolder Contents
scorer-k5/fold_k Reference-conditioned scorer, K=5 in-context references
scorer-k{7,9,11}/fold_k Scorers trained with K=7, 9, 11 references
noref/fold_k Scorer trained without references
simulator/fold_k Score-conditioned essay simulator that writes the references
4b-scorer-k5/fold_k, 4b-noref/fold_k The same scorers on Qwen/Qwen3.5-4B
judge Full-parameter scorer without references trained on all seven prompts; for reporting simulator read-back only

Loading

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "sjin4861/sail-asap2", subfolder="scorer-k5/fold_0")
tok = AutoTokenizer.from_pretrained("sjin4861/sail-asap2", subfolder="scorer-k5/fold_0")

Scorers decode greedily (temperature 0, 64 new tokens, thinking disabled); simulators sample with temperature 0.8, top-p 0.95, top-k 50 and up to 2,048 new tokens. Prompts are rendered by the accompanying code.

Each subfolder holds the checkpoint selected on source-prompt validation for that fold (SOURCE.txt names it; trainer_state.json gives its step and epoch). Adapters are LoRA rank 32, alpha 64, on all linear layers, seed 42.

Trained on essays from public essay-scoring corpora; use under those datasets' terms.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sjin4861/sail-asap2

Finetuned
Qwen/Qwen3.5-9B
Adapter
(759)
this model