Noema Overfit

Paged MoE model bundles for Noema's Overfit expert-paging runtime. Each subfolder is one .noema-paged package: a resident.gguf (all non-expert weights, always loaded) plus experts-*.bin page files that are streamed and evicted on demand, described by manifest.json.

These are runtime-specific packages, not standalone GGUF files. Each manifest.json declares its required Noema native contract: the Qwen and Gemma packages below require contract v3, while DeepSeek V4 requires contract v4.

Download the complete .noema-paged folder and preserve its layout. In Noema, add the package to Stored, select Automatic under Overfit (Paged Experts), and run the canary test before relying on it. A fast local SSD is strongly recommended. Expert paging increases runnable capacity; it does not turn storage into RAM or guarantee interactive speed.

Models

`DeepSeek-V4-Flash-0731-UD-IQ4_NL-00001-of-00004.noema-paged/`

Base model deepseek-ai/DeepSeek-V4-Flash-0731
Source GGUF unsloth/DeepSeek-V4-Flash-0731-GGUF/UD-IQ4_NL — 4 shards, 136.66 GB
Upstream revisions Base 7872f01b1d1fe23eabc4c98b48bffcef5a386062 · GGUF fbbb5b93fb787c21338159b0af3318bb3f4d9768
Architecture deepseek4 (256 routed experts, 6 active, 43 MoE layers)
Quantization UD-IQ4_NL
Resident weights resident.gguf — 7.81 GB
Expert pages experts-000.binexperts-007.bin — 128.85 GB total
Manifest Format v1 · native contract v4 · 33,024 records · 16,384-byte alignment
Source SHA-256 e53b8a27242a271af2ebee6171763e07913fb54dbacb3745afa76e0c50c062ff
Package fingerprint dcd6721b6c851be88ab192280dc9defc855e3f7f31e69b8db3bdcd8024727b12
Manifest SHA-256 d9a22388dca29a6abb4c296283f60f7a8b94b88b91d38e1107de05df5d46a32c
License MIT (upstream model)

`gemma-4-26B-A4B-it-qat-UD-Q4_K_XL.noema-paged/`

Base model google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
Source GGUF unsloth/gemma-4-26B-A4B-it-qat-GGUFUD-Q4_K_XL (14.25 GB)
Architecture gemma4 (128 experts, 8 active, 30 MoE layers, fused gate/up)
Resident weights resident.gguf — 1.40 GB
Expert pages experts-000.bin — 12.96 GB
Alignment 16384 bytes

`Qwen3.6-35B-A3B-UD-Q4_K_M.noema-paged/`

Base model Qwen/Qwen3.6-35B-A3B
Architecture qwen35moe (256 experts, 8 active, 40 MoE layers)
Source GGUF Qwen3.6-35B-A3B-UD-Q4_K_M.gguf (22.13 GB)
Resident weights resident.gguf — 2.57 GB
Expert pages experts-000.bin, experts-001.bin — 19.57 GB total
Alignment 16384 bytes

`Qwen3.5-122B-A10B-Q4_K_M.noema-paged/`

Architecture qwen35moe (256 experts, 8 active, 48 MoE layers)
Source GGUF Qwen3.5-122B-A10B-Q4_K_M (2 shards, 74.2 GB)
Resident weights resident.gguf — 4.0 GB
Expert pages experts-000.binexperts-004.bin — 65 GB
Alignment 16384 bytes

Packages are generated with Noema's paged-model conversion tooling. File sizes, source fingerprints, and SHA-256 integrity hashes are recorded in each package's manifest.json.

Licenses

Noema's packaging does not replace the upstream model licenses. Use each bundle according to the terms attached to its base model:

Downloads last month
316
GGUF
Model size
7B params
Architecture
deepseek4
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for NoemaAI-labs/Noema-Overfit

Quantized
(152)
this model