Commit History

Card: parity table at 1954de4 and real-document results at the model defaults
5265f03
verified

handwoven8588 commited on

Card: table generated from idle-bracketed gate runs at 1954de4; state what encode peak measures and how much wall time varies
d1021c8
verified

handwoven8588 commited on

Pin weights to the requested revision; harden the varlen probe; card: all tiers measured on one GPU with stated batch sizes
1954de4
verified

handwoven8588 commited on

Auto mode probes the varlen kernel and falls back; eager accepts a 2-D mask
4c04d49
verified

handwoven8588 commited on

Three-tier attention: torch varlen, flash_attn, eager (inline mask)
8187d6a
verified

handwoven8588 commited on

README: drop pipeline_tag (feature-extraction)
e361c6f
verified

handwoven8588 commited on

README: reframe bf16 weights — reason is flash_attn half-precision + runtime-bf16 (not download size)
aa2548a
verified

handwoven8588 commited on

README pass: base_model_relation=quantized, fix CPU-eager wording (bf16, not bit-identical), drop internal refs + first-person plural
7961156
verified

handwoven8588 commited on

README: fill 3090 Ti perf table (bf16 flash 2.1GB/162k tok/s vs fp32 eager 6.7GB/52k tok/s; cosine 0.9986)
b283422
verified

handwoven8588 commited on

v2: bf16 weights (547->274MB) + from_pretrained torch_dtype fix (loads bf16 natively) + corrected model tree (base_model=nomic-ai/CodeRankEmbed only) + bf16-derivative README
581207a
verified

handwoven8588 commited on

CodeRankEmbed with native flash-attn varlen forward (derivative of nomic-ai/CodeRankEmbed; identical weights; flash-vs-fp32 parity cosine 0.99999, eager fallback bit-identical)
22d2b3c
verified

handwoven8588 commited on