Ignore my previous comment as I was testing the configuration on llama.cpp, which isn't as optimized as MLX based frameworks. However after trying this on mlx-vlm, I'm actually curious how you got this running long enough for your sanity testing. I tried to test it myself, and mlx-vlm is currently suffering from intense RAM spikes (4x the usage) which then very quickly causes OOMs. I opened an issue and I am working with collaborators who confirmed this is an issue.
Anthony A
anthonyya
AI & ML interests
None yet
Recent Activity
commentedon an article 5 days ago
Block Diffusion on Apple Silicon with 3.7× Speedup for Qwopus 3.6 27B liked a model 5 days ago
lemuralabs/Qwen3.6-27B-V2-abliterated-uncensored-6-bit-mlx updated a model 5 days ago
anthonyya/Qwen3.6-27B-DFlash-8bitOrganizations
None yet