Text Generation
Safetensors
GGUF
Chinese
English
qwen3
fortune-telling
qwen
qwen2.5
ollama
conversational
Instructions to use Tbata7/FortuneQwen3_4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Tbata7/FortuneQwen3_4b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Tbata7/FortuneQwen3_4b:Q8_0 # Run inference directly in the terminal: llama cli -hf Tbata7/FortuneQwen3_4b:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Tbata7/FortuneQwen3_4b:Q8_0 # Run inference directly in the terminal: llama cli -hf Tbata7/FortuneQwen3_4b:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Tbata7/FortuneQwen3_4b:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf Tbata7/FortuneQwen3_4b:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Tbata7/FortuneQwen3_4b:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Tbata7/FortuneQwen3_4b:Q8_0
Use Docker
docker model run hf.co/Tbata7/FortuneQwen3_4b:Q8_0
- LM Studio
- Jan
- vLLM
How to use Tbata7/FortuneQwen3_4b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Tbata7/FortuneQwen3_4b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Tbata7/FortuneQwen3_4b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Tbata7/FortuneQwen3_4b:Q8_0
- Ollama
How to use Tbata7/FortuneQwen3_4b with Ollama:
ollama run hf.co/Tbata7/FortuneQwen3_4b:Q8_0
- Unsloth Studio
How to use Tbata7/FortuneQwen3_4b with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Tbata7/FortuneQwen3_4b to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Tbata7/FortuneQwen3_4b to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Tbata7/FortuneQwen3_4b to start chatting
- Docker Model Runner
How to use Tbata7/FortuneQwen3_4b with Docker Model Runner:
docker model run hf.co/Tbata7/FortuneQwen3_4b:Q8_0
- Lemonade
How to use Tbata7/FortuneQwen3_4b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Tbata7/FortuneQwen3_4b:Q8_0
Run and chat with the model
lemonade run user.FortuneQwen3_4b-Q8_0
List all available models
lemonade list
- Atomic Chat
| license: other | |
| language: | |
| - zh | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - fortune-telling | |
| - qwen | |
| - qwen2.5 | |
| - qwen3 | |
| - gguf | |
| - ollama | |
| <div align="center"> | |
| # FortuneQwen3_4b | |
| [**中文**](./README.md) | [**English**](./README_EN.md) | |
| </div> | |
| This is a 4B parameter model fine-tuned based on the Qwen3 architecture, specifically designed for Fortune Telling tasks. This repository provides merged Safetensors weights, GGUF quantized files, and a Modelfile for Ollama. | |
| ## About This Repository | |
| This repository contains the model in three formats: | |
| 1. **GGUF Quantized Model (Recommended)**: | |
| * Filename: `FortuneQwen3_4b_q8_0.gguf` (or other versions) | |
| * Description: Pre-converted GGUF format (Int8 quantized), ready for use with `llama.cpp` or `Ollama`. | |
| 2. **Modelfile**: | |
| * Filename: `Modelfile` | |
| * Description: Configuration file for Ollama import, defining system prompts and parameters. | |
| 3. **Hugging Face Safetensors**: | |
| * Filename: `model.safetensors`, etc. | |
| * Description: Full model parameters with LoRA weights already merged, suitable for Transformers-based inference, further fine-tuning, or exporting custom GGUF files. | |
| ## Quick Start | |
| ### Option 1: Using Ollama (Recommended) | |
| You can quickly create an Ollama model using the pre-converted GGUF file found in this repository. | |
| 1. **Clone this repository**: | |
| ```bash | |
| git clone https://huggingface.co/Tbata7/FortuneQwen3_4b | |
| cd FortuneQwen3_4b | |
| ``` | |
| 2. **Create the model**: | |
| ```bash | |
| # This uses the local Modelfile and GGUF file | |
| ollama create FortuneQwen3_q8:4b -f Modelfile | |
| ``` | |
| 3. **Run the model**: | |
| ```bash | |
| ollama run FortuneQwen3_q8:4b | |
| ``` | |
| ### Option 2: Using llama.cpp | |
| If you prefer to use the GGUF file directly with `llama.cpp`: | |
| ```bash | |
| ./llama-cli -m FortuneQwen3_4b_q8_0.gguf -p "Your question here..." -n 512 | |
| ``` | |
| ## Advanced Usage: Exporting Custom GGUF | |
| If you wish to use a different quantization level (e.g., q4_k, q6_k, fp16), you can export a custom GGUF from the Safetensors weights using `llama.cpp`. | |
| 1. **Prepare Environment**: | |
| Ensure you have `llama.cpp` python dependencies installed. | |
| 2. **Convert Model**: | |
| Use the `convert_hf_to_gguf.py` script. You must specify the `--outtype` parameter to control the output type. | |
| * **Export as FP16 (No quantization)**: | |
| ```bash | |
| python llama.cpp/convert_hf_to_gguf.py ./FortuneQwen3_4b --outfile FortuneQwen3_4b_fp16.gguf --outtype f16 | |
| ``` | |
| * **Export as Int8 (q8_0)**: | |
| ```bash | |
| python llama.cpp/convert_hf_to_gguf.py ./FortuneQwen3_4b --outfile FortuneQwen3_4b_q8_0.gguf --outtype q8_0 | |
| ``` | |
| * **Other Quantizations**: | |
| First export as f16, then use the `llama-quantize` tool: | |
| ```bash | |
| ./llama-quantize FortuneQwen3_4b_fp16.gguf FortuneQwen3_4b_q4_k_m.gguf q4_k_m | |
| ``` | |
| ## Model Information | |
| - **Base Architecture**: Qwen3:4B | |
| - **Task**: Fortune Telling / I-Ching Interpretation | |
| - **Context Window**: 32768 | |
| - **Fine-tuning Framework**: LLaMA-Factory | |
| ## Disclaimer | |
| This model is for entertainment and research purposes only. Please believe in science. | |