infosave commited on
Commit
b05a511
·
verified ·
1 Parent(s): 5590d8e

Refresh GLM Q2TP and Q4TP performance evidence

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -107,7 +107,7 @@ Neither the 108.27 GiB nor the 155.68 GiB file needs to fit in VRAM. CMF keeps w
107
  in host memory and detects the available adapter budget. When dynamic pooling
108
  is enabled, routed experts use a bounded global GPU pool: resident experts run
109
  on the GPU, while cache misses are completed exactly on CPU and accumulated
110
- into the same MoE result. Q4TP automatic mode and systems without a supported
111
  adapter use the same file on CPU.
112
 
113
  Useful controls:
 
107
  in host memory and detects the available adapter budget. When dynamic pooling
108
  is enabled, routed experts use a bounded global GPU pool: resident experts run
109
  on the GPU, while cache misses are completed exactly on CPU and accumulated
110
+ into the same MoE result. Automatic mode and systems without a supported
111
  adapter use the same file on CPU.
112
 
113
  Useful controls: