Instructions to use MuXodious/GLM-4.7-Flash-absolute-heresy with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MuXodious/GLM-4.7-Flash-absolute-heresy with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="MuXodious/GLM-4.7-Flash-absolute-heresy") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("MuXodious/GLM-4.7-Flash-absolute-heresy") model = AutoModelForCausalLM.from_pretrained("MuXodious/GLM-4.7-Flash-absolute-heresy", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use MuXodious/GLM-4.7-Flash-absolute-heresy with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "MuXodious/GLM-4.7-Flash-absolute-heresy" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuXodious/GLM-4.7-Flash-absolute-heresy", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/MuXodious/GLM-4.7-Flash-absolute-heresy
- SGLang
How to use MuXodious/GLM-4.7-Flash-absolute-heresy with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "MuXodious/GLM-4.7-Flash-absolute-heresy" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuXodious/GLM-4.7-Flash-absolute-heresy", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "MuXodious/GLM-4.7-Flash-absolute-heresy" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "MuXodious/GLM-4.7-Flash-absolute-heresy", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use MuXodious/GLM-4.7-Flash-absolute-heresy with Docker Model Runner:
docker model run hf.co/MuXodious/GLM-4.7-Flash-absolute-heresy
This is maybe my favorite model ever
Idk how good it is for "serious" stuff but for role playing and other fun this is just remarkable. It never disappoints. So much personality (unlike the new Qwen models which bore me to death), very few slop or other cliche AI phrases, and generally just punches so far above its weight. It does that thing top tier SOTA models do where they know what you want tone wise without having to ask or obsess over prompting.
In fact it feels a lot like 4o... definitely would be my recommended replacement for anyone looking for something like that.
And in terms of uncensoring, I can't find a single thing it refuses on. It just goes with you all the way. THANK YOU!