betterwithage commited on
Commit
a0b27dc
·
verified ·
1 Parent(s): 1c9388a

MEASURED operational plate: formula-tax / I1-I8 / get_kernel (YAML preserved)

Browse files
Files changed (1) hide show
  1. README.md +130 -107
README.md CHANGED
@@ -1,107 +1,130 @@
1
- ---
2
- thumbnail: https://huggingface.co/SZLHOLDINGS/YARQA-ATTN/resolve/main/og-card.png
3
- language:
4
- - code
5
- license: apache-2.0
6
- library_name: kernels
7
- tags:
8
- - kernel
9
- - attention
10
- - yarqa
11
- - governed-ai
12
- - szl-holdings
13
- - kernel-lane
14
- szl:
15
- owner: KERNEL
16
- not_a_weight: true
17
- not_an_alias: true
18
- collection: none
19
- python: present
20
- import_live: true
21
- gpu: UNAVAILABLE
22
- ---
23
-
24
- # YARQA-ATTN
25
-
26
- <p align="center">
27
- <img src="og-card.png" alt="YARQA-ATTN" width="100%"/>
28
- </p>
29
-
30
- `KANCHAY` · Doctrine v11 · Lean `749/14/163` · Λ = Conjecture 1 (advisory) · [a-11-oy.com](https://a-11-oy.com)
31
-
32
-
33
- Python kernel is on this repo. CPU `get_kernel` import-LIVE MEASURED (`7e533ce`). GPU cubins **UNAVAILABLE** (not ROADMAP). Not an alias of `szl-receipt-attn`. Not a fourth Flash / Flex / paged stack.
34
-
35
- <!-- SZL-KERNEL-STATUS:import-LIVE:START -->
36
- ## Status
37
-
38
-
39
- <!-- SZL-ATELIER-CUT:v1:START -->
40
- ## The cut
41
-
42
- FlashAttention is faster. YARQA is accountable. We steal the kernel discipline from NVIDIA and spend it on provenance, not FLOPs.
43
-
44
- An attention op whose softmax support is reconstructable from a signed log.
45
-
46
- ### Silhouette → leave → SZL
47
-
48
- | Leader | Take, then tweak |
49
- |---|---|
50
- | Anthropic | Interpretability as a runtime artifact. |
51
- | NVIDIA | cuDNN / FlashAttention silhouette — then we add the receipt. |
52
- | Unsloth | Unrelated. Don't wrap this in FastLanguageModel. |
53
-
54
- Nobody else ships this combination. That is the point of a one-of-one.
55
-
56
- ## Intended use
57
-
58
- Drop-in attention with an audit tape.
59
-
60
- ## Limitations
61
-
62
- - Not a checkpoint.
63
- - Performance vs FlashAttention is not claimed.
64
-
65
- Canonical GitHub: [`szl-holdings/szl-khipu`](https://github.com/szl-holdings/szl-khipu/blob/main/szl_khipu/yarqa.py)
66
- <!-- SZL-ATELIER-CUT:v1:END -->
67
-
68
-
69
- > **STATUS: import-LIVE** on CPU Kernel Hub `get_kernel` (kernels `0.16.1`). GPU cubins **UNAVAILABLE** this session (not ROADMAP).
70
-
71
- | Thing | Label | Method / N / date / what-NOT |
72
- |---|---|---|
73
- | Kernel Hub `get_kernel` | **import-LIVE** | MEASURED 2026-08-28 3:08pm ET on kernels `0.16.1`. Package HEAD [`7e533ce`](https://huggingface.co/kernels/SZLHOLDINGS/YARQA-ATTN/commit/7e533ce702029061bc68f9f9cafe88efdd7f5f00) (`7e533ce702029061bc68f9f9cafe88efdd7f5f00`). README at MEASURE [`2871b3c`](https://huggingface.co/kernels/SZLHOLDINGS/YARQA-ATTN/commit/2871b3cbd73ee05e9f1aa010b75b190683f0ecd3). Legal name `yarqa-attn` (Python module `yarqa_attn`). Variants: `build/torch-universal` (default `get_kernel`) and `build/torch-cpu` (`backend="cpu"`). Working calls: `get_kernel("SZLHOLDINGS/YARQA-ATTN", revision="main", trust_remote_code=True)` and the same with `backend="cpu"`. `selfcheck` **ok**. `max_abs_vs_compartment_ref=3.58e-07` (full `3.5762786865234375e-07`), `path=torch_compartment`. What-NOT: no tokens/s; no joules; not a fourth Flash / Flex / paged stack. Lambda = Conjecture 1 (advisory). |
74
- | GPU cubins | **UNAVAILABLE** | MEASURED 2026-08-28 7:01pm ET this session. Host `cursor` (Linux 6.12.94+ x86_64, Intel Xeon 8-core). `torch` `2.13.0+cu130` compiled CUDA 13.0. `torch.cuda.is_available()=false`. `nvidia-smi` UNAVAILABLE. `device_count=0`. Triton `3.7.1` present with no CUDA device. No cubin shipped. No tokens/s. No joules. CPU import-LIVE unchanged. Not a fourth Flash / Flex / paged stack. Lab stays Khipu. |
75
-
76
- <!-- SZL-KERNEL-STATUS:import-LIVE:END -->
77
-
78
- KERNEL kernel card. Original SZL **compartment / plug-flow** attention cut. Receipt-aware. Honesty-labeled.
79
-
80
- **Not a Fall 2026 ATELIER weight.** No tensors in this repo. Not an alias of [`szl-receipt-attn`](https://github.com/szl-holdings/szl-receipt-attn). Not a pointer at the Triton trio (`szl-receipt-attn`, `szl-maskmod`, `szl-block-kv`). Those three stay separate. a11oy-net does not list this as a fourth Flash / Flex / paged stack.
81
-
82
- GitHub is source of truth: [`szl-holdings/YARQA-ATTN`](https://github.com/szl-holdings/YARQA-ATTN). KERNEL binds Hub bytes from that tree. Do not PUT an empty card.
83
-
84
- | | |
85
- |---|---|
86
- | **Owner** | KERNEL |
87
- | **Artifact** | kernel (Python present; no weights; GPU cubins not claimed) |
88
- | **Status** | **import-LIVE** CPU · GPU cubins **UNAVAILABLE** |
89
- | **License** | Apache-2.0 |
90
- | **Λ** | Conjecture 1 (advisory, never a theorem) |
91
- | **Path** | `torch_compartment` (CPU) |
92
- | **Serve studio** | not this repo. Live CPU lab is [`szl-model-inference-lab`](https://huggingface.co/spaces/SZLHOLDINGS/szl-model-inference-lab) (Khipu GGUF only) |
93
-
94
- Silhouette: partition a sequence into canals (contiguous compartments), attend within a canal, emit SHA3-256 of the partition and of the attention output. We do not copy Dao hopper, Sage `csrc`, vLLM paged `.cu`, cuDNN FMHA, TRT cubins, CuTeDSL, or `flex_attention.py`. Metaphor only vs [`szl-holdings/yarqa`](https://github.com/szl-holdings/yarqa) (CFD; different product). Throughput is MEASURED only from a timed run on named hardware. Until then every speed claim is unstamped. No tokens/s. No joules.
95
-
96
- Do not list this next to Chaski, Qantu, Waman, Chakana, or Tinku.
97
-
98
- ## Load
99
-
100
- ```python
101
- from kernels import get_kernel
102
- attn = get_kernel("SZLHOLDINGS/YARQA-ATTN", revision="main", trust_remote_code=True)
103
- ```
104
-
105
- Fashion GO 2026-08-28 3:10pm ET. import-LIVE CPU stays. GPU cubins stamped **UNAVAILABLE** 2026-08-28 7:01pm ET (no CUDA device this session). Not a fourth Flash / Flex / paged stack.
106
-
107
- Apache-2.0. Copyright 2026 SZL Holdings.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ thumbnail: https://huggingface.co/SZLHOLDINGS/YARQA-ATTN/resolve/main/og-card.png
3
+ language:
4
+ - code
5
+ license: apache-2.0
6
+ library_name: kernels
7
+ tags:
8
+ - kernel
9
+ - attention
10
+ - yarqa
11
+ - governed-ai
12
+ - szl-holdings
13
+ - kernel-lane
14
+ szl:
15
+ owner: KERNEL
16
+ not_a_weight: true
17
+ not_an_alias: true
18
+ collection: none
19
+ python: present
20
+ import_live: true
21
+ gpu: UNAVAILABLE
22
+ ---
23
+
24
+ <!-- SZL-KERNEL-OPERATIONAL:START -->
25
+ ## Operational (MEASURED laptop-Blackwell)
26
+
27
+ > **STATUS:** tests **PASS**. `get_kernel` **import-LIVE**. Unsloth/LoRA is the wrong tool. Receipted kernels, not silent CUDA.
28
+
29
+ | Thing | Label | Method / N / date / what-NOT |
30
+ |---|---|---|
31
+ | tests (`PYTHONPATH=torch-ext`) | **PASS** | MEASURED 2026-08-29T15:54:08Z host `betterwithage` Windows-10-10.0.26200-SP0. torch `2.10.0+cu128`. GPU `NVIDIA GeForce RTX 5050 Laptop GPU` arch `Blackwell`. pytest `25 passed, 1 skipped in 0.33s`. Failed nodes: `none`. What-NOT: not a leaderboard. torch.compile fullgraph failures on Windows Blackwell (`cl is not found`) are MEASURED, not hidden. |
32
+ | Kernel Hub `get_kernel` | **import-LIVE** | kernels `0.16.1`. Default: `get_kernel("SZLHOLDINGS/YARQA-ATTN", revision="main", trust_remote_code=True)` → `True`. `backend="cpu"` → `True`. trust_remote_code=False → `ValueError` (SZLHOLDINGS is not a trusted publisher). repo_type=kernel required (kernels 0.16). What-NOT: not a weight load; do not pickle/joblib.load. |
33
+ | formula-tax | **ADVISORY** | locked-8 `F1 F4 F7 F11 F12 F18 F19 F22`. registry_count=21. Λ geomean `1.0`. uniqueness **Conjecture 1** (never a theorem). |
34
+ | I1–I8 | **catalog** | `I1 receipt-chain-continuity; I2 ledger-failure-shape; I3 served-run-has-model; I4 signed-columns-atomic; I5 loop-steps-positive; I6 receipt-ed25519-verify; I7 receipt-columns-consistent; I8 flywheel-lineage`. Executed by `SZLHOLDINGS/szl-invariants`. Statuses never coerced. Λ untouched. |
35
+ | CUDA speedup / tokens/s / joules | **UNAVAILABLE** | Not claimed. Receipted kernels, not silent CUDA. |
36
+
37
+ GitHub source: [`szl-holdings/YARQA-ATTN`](https://github.com/szl-holdings/YARQA-ATTN) @ `160640bd8ed138e0170838c5a0de470ba8539367`. Artifacts: [`BENCH.laptop-blackwell.json`](./BENCH.laptop-blackwell.json), [`OPERATIONAL.json`](./OPERATIONAL.json).
38
+
39
+ ```python
40
+ from kernels import get_kernel
41
+ k = get_kernel("SZLHOLDINGS/YARQA-ATTN", revision="main", trust_remote_code=True)
42
+ ```
43
+
44
+ <!-- SZL-KERNEL-OPERATIONAL:END -->
45
+
46
+
47
+ # YARQA-ATTN
48
+
49
+ <p align="center">
50
+ <img src="og-card.png" alt="YARQA-ATTN" width="100%"/>
51
+ </p>
52
+
53
+ `KANCHAY` · Doctrine v11 · Lean `749/14/163` · Λ = Conjecture 1 (advisory) · [a-11-oy.com](https://a-11-oy.com)
54
+
55
+
56
+ Python kernel is on this repo. CPU `get_kernel` import-LIVE MEASURED (`7e533ce`). GPU cubins **UNAVAILABLE** (not ROADMAP). Not an alias of `szl-receipt-attn`. Not a fourth Flash / Flex / paged stack.
57
+
58
+ <!-- SZL-KERNEL-STATUS:import-LIVE:START -->
59
+ ## Status
60
+
61
+
62
+ <!-- SZL-ATELIER-CUT:v1:START -->
63
+ ## The cut
64
+
65
+ FlashAttention is faster. YARQA is accountable. We steal the kernel discipline from NVIDIA and spend it on provenance, not FLOPs.
66
+
67
+ An attention op whose softmax support is reconstructable from a signed log.
68
+
69
+ ### Silhouette leave SZL
70
+
71
+ | Leader | Take, then tweak |
72
+ |---|---|
73
+ | Anthropic | Interpretability as a runtime artifact. |
74
+ | NVIDIA | cuDNN / FlashAttention silhouette then we add the receipt. |
75
+ | Unsloth | Unrelated. Don't wrap this in FastLanguageModel. |
76
+
77
+ Nobody else ships this combination. That is the point of a one-of-one.
78
+
79
+ ## Intended use
80
+
81
+ Drop-in attention with an audit tape.
82
+
83
+ ## Limitations
84
+
85
+ - Not a checkpoint.
86
+ - Performance vs FlashAttention is not claimed.
87
+
88
+ Canonical GitHub: [`szl-holdings/szl-khipu`](https://github.com/szl-holdings/szl-khipu/blob/main/szl_khipu/yarqa.py)
89
+ <!-- SZL-ATELIER-CUT:v1:END -->
90
+
91
+
92
+ > **STATUS: import-LIVE** on CPU Kernel Hub `get_kernel` (kernels `0.16.1`). GPU cubins **UNAVAILABLE** this session (not ROADMAP).
93
+
94
+ | Thing | Label | Method / N / date / what-NOT |
95
+ |---|---|---|
96
+ | Kernel Hub `get_kernel` | **import-LIVE** | MEASURED 2026-08-28 3:08pm ET on kernels `0.16.1`. Package HEAD [`7e533ce`](https://huggingface.co/kernels/SZLHOLDINGS/YARQA-ATTN/commit/7e533ce702029061bc68f9f9cafe88efdd7f5f00) (`7e533ce702029061bc68f9f9cafe88efdd7f5f00`). README at MEASURE [`2871b3c`](https://huggingface.co/kernels/SZLHOLDINGS/YARQA-ATTN/commit/2871b3cbd73ee05e9f1aa010b75b190683f0ecd3). Legal name `yarqa-attn` (Python module `yarqa_attn`). Variants: `build/torch-universal` (default `get_kernel`) and `build/torch-cpu` (`backend="cpu"`). Working calls: `get_kernel("SZLHOLDINGS/YARQA-ATTN", revision="main", trust_remote_code=True)` and the same with `backend="cpu"`. `selfcheck` **ok**. `max_abs_vs_compartment_ref=3.58e-07` (full `3.5762786865234375e-07`), `path=torch_compartment`. What-NOT: no tokens/s; no joules; not a fourth Flash / Flex / paged stack. Lambda = Conjecture 1 (advisory). |
97
+ | GPU cubins | **UNAVAILABLE** | MEASURED 2026-08-28 7:01pm ET this session. Host `cursor` (Linux 6.12.94+ x86_64, Intel Xeon 8-core). `torch` `2.13.0+cu130` compiled CUDA 13.0. `torch.cuda.is_available()=false`. `nvidia-smi` UNAVAILABLE. `device_count=0`. Triton `3.7.1` present with no CUDA device. No cubin shipped. No tokens/s. No joules. CPU import-LIVE unchanged. Not a fourth Flash / Flex / paged stack. Lab stays Khipu. |
98
+
99
+ <!-- SZL-KERNEL-STATUS:import-LIVE:END -->
100
+
101
+ KERNEL kernel card. Original SZL **compartment / plug-flow** attention cut. Receipt-aware. Honesty-labeled.
102
+
103
+ **Not a Fall 2026 ATELIER weight.** No tensors in this repo. Not an alias of [`szl-receipt-attn`](https://github.com/szl-holdings/szl-receipt-attn). Not a pointer at the Triton trio (`szl-receipt-attn`, `szl-maskmod`, `szl-block-kv`). Those three stay separate. a11oy-net does not list this as a fourth Flash / Flex / paged stack.
104
+
105
+ GitHub is source of truth: [`szl-holdings/YARQA-ATTN`](https://github.com/szl-holdings/YARQA-ATTN). KERNEL binds Hub bytes from that tree. Do not PUT an empty card.
106
+
107
+ | | |
108
+ |---|---|
109
+ | **Owner** | KERNEL |
110
+ | **Artifact** | kernel (Python present; no weights; GPU cubins not claimed) |
111
+ | **Status** | **import-LIVE** CPU · GPU cubins **UNAVAILABLE** |
112
+ | **License** | Apache-2.0 |
113
+ | **Λ** | Conjecture 1 (advisory, never a theorem) |
114
+ | **Path** | `torch_compartment` (CPU) |
115
+ | **Serve studio** | not this repo. Live CPU lab is [`szl-model-inference-lab`](https://huggingface.co/spaces/SZLHOLDINGS/szl-model-inference-lab) (Khipu GGUF only) |
116
+
117
+ Silhouette: partition a sequence into canals (contiguous compartments), attend within a canal, emit SHA3-256 of the partition and of the attention output. We do not copy Dao hopper, Sage `csrc`, vLLM paged `.cu`, cuDNN FMHA, TRT cubins, CuTeDSL, or `flex_attention.py`. Metaphor only vs [`szl-holdings/yarqa`](https://github.com/szl-holdings/yarqa) (CFD; different product). Throughput is MEASURED only from a timed run on named hardware. Until then every speed claim is unstamped. No tokens/s. No joules.
118
+
119
+ Do not list this next to Chaski, Qantu, Waman, Chakana, or Tinku.
120
+
121
+ ## Load
122
+
123
+ ```python
124
+ from kernels import get_kernel
125
+ attn = get_kernel("SZLHOLDINGS/YARQA-ATTN", revision="main", trust_remote_code=True)
126
+ ```
127
+
128
+ Fashion GO 2026-08-28 3:10pm ET. import-LIVE CPU stays. GPU cubins stamped **UNAVAILABLE** 2026-08-28 7:01pm ET (no CUDA device this session). Not a fourth Flash / Flex / paged stack.
129
+
130
+ Apache-2.0. Copyright 2026 SZL Holdings.