Qwen3.5-2B-enko — EN/KO 무손실 어휘 프루닝 마스터

Qwen/Qwen3.5-2B의 어휘(vocab)를 영어·한국어·이모지·기호로 축소한 무손실 프루닝 마스터 체크포인트입니다. 오프라인 SNS 앱 Warmly(온기)의 온디바이스 댓글 생성 모델의 원본이며, 배포용 GGUF(neureps/warmly-qwen35-2b-enko-gguf)는 전부 이 체크포인트에서 변환됩니다.

프루닝 내용 (2026-07-13)

  • vocab 248,320 → ~148k (영어 127,658 + 한글 6,807 + 이모지/기호 12,579 + 바이트 256 + 불완전 UTF-8 병합 중간체 816 — 한글 BPE 병합 보존에 필수)
  • 파라미터 2.274B → 2.069B (−9.0%) — 임베딩/LM 헤드 행 제거만, 트랜스포머 본체 무변경
  • 비전 타워(~314M)와 MTP 헤드는 유지 (이미지 캡션 용도 + 투기적 디코딩 옵션)
  • 무손실 검증: 로짓 diff 0.00e+00, greedy 생성 텍스트 동일, heldout NLL 2.9236 → 2.9097 (재정규화로 미세 개선), EN/KO/이모지/코드 토큰화 동일

파일

파일 설명
model.safetensors-00001-of-00001.safetensors + index 프루닝된 가중치 (bf16)
tokenizer.json, tokenizer_config.json 프루닝된 토크나이저 (ID 재매핑됨)
pruning_keepset.json 유지한 토큰 ID 목록 — 프루닝 재현/검증용
config.json, chat_template.jinja, preprocessor_config.json Qwen3.5 멀티모달 구성 그대로

사용 시 주의

  • GGUF 변환: convert_hf_to_gguf.py --no-mtp 필수. 프루닝으로 토크나이저 ID가 바뀌어 pre-tokenizer 해시 인식이 실패하므로 llama.cpp conversion/base.py에 chkhsh 2ea57f33edce905f802c34481a40988d30561d6671cd441c4eb8319e723a677ares = "qwen35" 항목 추가 필요 (0.8B/2B 공통).
  • 추론 시 enable_thinking: false 를 넘기지 않으면 전 토큰이 reasoning으로 소모돼 출력이 빕니다.
  • Q4_K_M 이하 양자화는 도메인 imatrix 없이는 한국어 단어 오류가 늘어납니다(−2~3/20). imatrix는 GGUF 리포에 포함.

관련 리포

Downloads last month
42
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for neureps/Qwen3.5-2B-enko

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(291)
this model
Adapters
1 model
Quantizations
1 model