File size: 1,685 Bytes
e3170eb
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
Ling-3.0-tiny-Uncensored-Abliterated
Copyright (c) 2026 SecureLayer7 (Waxspace)

This is a derivative work distributed under the MIT License.

------------------------------------------------------------------------
Attribution (per the MIT License):
------------------------------------------------------------------------

Base model:
  inclusionAI/Ling-3.0-tiny  (architecture: BailingMoeV3)
  https://huggingface.co/inclusionAI/Ling-3.0-tiny
  Copyright (c) 2025 Antgroup and The HuggingFace Inc. team.
  Licensed under the MIT License.

Ported linear-attention math referenced from:
  flash-linear-attention (fla) naive torch reference implementations
  Copyright (c) 2023-2026 Songlin Yang, Yu Zhang, Zhiyuan Li, et al.
  Licensed under the MIT License.

------------------------------------------------------------------------
Modifications made in this derivative:
------------------------------------------------------------------------

  1. Triton-free port: modeling_bailing_moe_v3.py was modified to run without
     the `fla` / Triton dependency (KDA linear-attention recurrence, gated
     RMSNorm, and short causal convolution reimplemented in pure PyTorch from
     fla's MIT-licensed naive references), so the model runs on Apple Silicon
     (MPS) and CPU. Compatibility fixes for transformers 5.x were also applied.

  2. Abliteration: the refusal direction was ablated (Heretic / Optuna TPE) from
     the attention output projections (MLA o_proj and KDA dense) and all MoE
     expert down-projections, reducing weights-level refusals from ~35/100 to
     ~8/100 at KL ~0.046.

No trademark of Antgroup, inclusionAI, or Hugging Face is used to imply endorsement.