POSH prebuilt llama-server (aarch64-musl)

Static aarch64-linux-musl build of llama.cpp's llama-server, published for the POSH Android app. POSH's on-device GGUF engine downloads this binary into its Alpine/proot sandbox instead of compiling llama.cpp on the phone (which takes 10-30 minutes and can fail on low-RAM devices).

Built CPU-only, GGML_NATIVE=OFF (portable armv8-a baseline), static musl link. llama.cpp is MIT-licensed.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support