The model is very good.

#7
by deniiiiiij - opened

In my personal tests, this is undoubtedly the champion among all the llm I can run on my 5060ti at an adequate speed. BigBang delivers 70t/s with llamacpp; here are the launch parameters:
./build/bin/llama-server
-m /endless-frontier_BigBang-v1-IQ4_XS.gguf
--host 0.0.0.0 --port 8080
--flash-attn on
--ctx-size 128000
-ctk q4_0 -ctv q4_0
-ngl 999
--n-cpu-moe 13
--reasoning on
--reasoning-preserve
--no-mmap
--mlock
--fit off
-t 7
-tb 4
--batch-size 512
--ubatch-size 256
--spec-type draft-mtp
--spec-draft-n-max 2
--spec-draft-p-min 0.0
--no-mmproj
--temp 0.6 --top-p 0.95 --top-k 20
--min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.1

repeat-penalty 1.1 otherwise the model starts to loop, takes a very long time to think through complex tasks, but follows the instructions perfectly. Thank you, developers.

Sign up or log in to comment