Text Generation
Transformers
Safetensors
English
qwen2
aethersearch
agentic-search
search-augmented-generation
supervised-fine-tuning
conversational
text-generation-inference
Instructions to use muradil211/AetherSearch_SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use muradil211/AetherSearch_SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="muradil211/AetherSearch_SFT") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("muradil211/AetherSearch_SFT") model = AutoModelForCausalLM.from_pretrained("muradil211/AetherSearch_SFT", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use muradil211/AetherSearch_SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "muradil211/AetherSearch_SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "muradil211/AetherSearch_SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/muradil211/AetherSearch_SFT
- SGLang
How to use muradil211/AetherSearch_SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "muradil211/AetherSearch_SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "muradil211/AetherSearch_SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "muradil211/AetherSearch_SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "muradil211/AetherSearch_SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use muradil211/AetherSearch_SFT with Docker Model Runner:
docker model run hf.co/muradil211/AetherSearch_SFT
| library_name: transformers | |
| pipeline_tag: text-generation | |
| base_model: Qwen/Qwen2.5-3B-Instruct | |
| datasets: | |
| - muradil211/AetherSearch_SFT | |
| tags: | |
| - aethersearch | |
| - agentic-search | |
| - search-augmented-generation | |
| - supervised-fine-tuning | |
| - qwen2 | |
| language: | |
| - en | |
| <div align="center"> | |
| <img src="assets/aethersearch-mark.svg" alt="AetherSearch monogram" width="144"> | |
| # 🔭 AetherSearch SFT | |
| ### A compact search agent that learns to reason, retrieve, and answer | |
| Fine-tuned from **Qwen2.5-3B-Instruct** on **2,000 complete search trajectories**. | |
| <p> | |
| <a href="https://huggingface.co/Qwen/Qwen2.5-3B-Instruct"><img src="https://img.shields.io/badge/Base-Qwen2.5--3B--Instruct-7C3AED?style=flat-square" alt="Base model: Qwen2.5-3B-Instruct"></a> | |
| <img src="https://img.shields.io/badge/Weights-BF16-0F766E?style=flat-square" alt="Weights: BF16"> | |
| <a href="https://huggingface.co/datasets/muradil211/AetherSearch_SFT"><img src="https://img.shields.io/badge/Trajectories-2%2C000-F59E0B?style=flat-square" alt="Training trajectories: 2,000"></a> | |
| <img src="https://img.shields.io/badge/Context-32K-2563EB?style=flat-square" alt="Context window: 32K"> | |
| </p> | |
| [🏠 Project](https://github.com/Muradil-mamat-211/AetherSearch) · | |
| [🧪 Training code](https://github.com/Muradil-mamat-211/AetherSearch/tree/main/sft) · | |
| [📚 Dataset](https://huggingface.co/datasets/muradil211/AetherSearch_SFT) | |
| </div> | |
| > 🔌 **Bring your own retriever.** AetherSearch SFT is a search-agent policy, | |
| > not a self-contained QA service. The host runtime must execute each | |
| > `<search>...</search>` request and return evidence inside | |
| > `<information>...</information>`. | |
| ## ✨ Highlights | |
| - 🔎 **Search-native behavior** — learns when and what to search before answering. | |
| - 🔁 **Single- and multi-search trajectories** — trained on 1,025 single-search | |
| and 975 multi-search examples. | |
| - 🧾 **Evidence-in-the-loop reasoning** — retrieved passages stay visible as | |
| context while being excluded from the training loss. | |
| - ⚡ **Compact 3B backbone** — built on Qwen2.5-3B-Instruct for accessible | |
| experimentation and deployment. | |
| - 🧪 **Reproducible release** — public trainer, launcher, data checksum, schema | |
| tests, and artifact manifest are included or linked. | |
| ## 🧠 How it works | |
| ```text | |
| Question | |
| │ | |
| ▼ | |
| <think>reason about what is missing</think> | |
| │ | |
| ▼ | |
| <search>focused retrieval query</search> ─────► Search / RAG backend | |
| ▲ │ | |
| └──── <information>retrieved evidence</information> ◄────┘ | |
| │ | |
| ├── repeat the search loop when more evidence is needed | |
| ▼ | |
| <answer>evidence-grounded final answer</answer> | |
| ``` | |
| The model produces the reasoning, search, and answer spans. Your runtime owns | |
| retrieval: parse a completed `<search>` span, run the query, append the result | |
| as `<information>`, and resume generation until the model emits `<answer>`. | |
| ## 📊 Model at a glance | |
| | Field | Value | | |
| |---|---| | |
| | 🧱 Base model | [`Qwen/Qwen2.5-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct) | | |
| | 🧬 Base revision | `aa8e72537993ba99e69dfaafa59ed015b17504d1` | | |
| | 🏗️ Architecture | Qwen2 causal language model | | |
| | 🔢 Parameters | 3,085,938,688 | | |
| | 🎛️ Weight dtype | BF16 | | |
| | 📏 Context | 32,768 positions; training sequences capped at 4,096 | | |
| | 📚 Training data | 2,000 complete trajectories | | |
| | 🔍 Search mix | 1,025 single-search + 975 multi-search trajectories | | |
| | 🎓 Training stage | One full-trajectory SFT stage | | |
| ## 🚀 Quick start | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "muradil211/AetherSearch_SFT" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto", | |
| ) | |
| model.config.use_cache = True | |
| model.eval() | |
| ``` | |
| > 💡 Loading the checkpoint is only the first step. For end-to-end use, wrap | |
| > generation in the retrieval loop shown above and preserve the XML protocol | |
| > exactly. | |
| ## 🧬 Checkpoint identity | |
| This model was trained once on the 2,000 records in the canonical | |
| `final_sft_2000.jsonl` dataset, using the same configuration as the public | |
| AetherSearch SFT-2000 training code. The release contains the final model | |
| artifacts and reproducible code, not server-local logs or optimizer state. | |
| **Dataset SHA-256** | |
| ```text | |
| fec609652d3832c7a6c0ee2861c6f946b6cf7c3d3d40fc5d9be9b75df6325dcb | |
| ``` | |
| ## 🧪 Training recipe | |
| | Setting | Value | Setting | Value | | |
| |---|---:|---|---:| | |
| | Epochs | 1 | Learning rate | `2e-6` | | |
| | Scheduler | Cosine | Global batch size | 24 | | |
| | Precision | BF16 + TF32 | Max sequence length | 4,096 | | |
| | Padding | Dynamic | Distributed training | DeepSpeed ZeRO-3 | | |
| The training configuration matches the public SFT-2000 recipe: one epoch, | |
| learning rate `2e-6`, cosine scheduling, BF16, TF32, gradient checkpointing, | |
| dynamic padding, effective global batch size 24, and DeepSpeed ZeRO-3. On the | |
| three-worker training topology, per-device batch size 1 and gradient | |
| accumulation 8 resolve to that global batch. The completed checkpoint is | |
| exported as `final_model/`. | |
| The public launcher is hardware-topology independent: it uses the devices made | |
| visible by the surrounding runtime and derives gradient accumulation to keep | |
| global batch 24 unchanged. It does not embed physical GPU IDs, node addresses, | |
| NCCL fabric settings, allocator tuning, or server-local paths. | |
| ### 🎯 Supervision contract | |
| - ⬛ System, user, and question tokens are **masked**. | |
| - ⬛ Complete `<information>...</information>` spans are **masked**. | |
| - ✅ Assistant `<think>`, `<search>`, and `<answer>` spans are **supervised**. | |
| - ✅ The final assistant `<|im_end|>` token is **supervised**. | |
| The trainer, launcher, configuration, checksum, and schema tests are published | |
| in the [AetherSearch SFT directory](https://github.com/Muradil-mamat-211/AetherSearch/tree/main/sft). | |
| ## 📦 Files and integrity | |
| The release contains two BF16 SafeTensors shards, the shard index, model and | |
| generation configuration, tokenizer assets, this model card, the project logo, | |
| and `MODEL_MANIFEST.sha256`. It intentionally excludes optimizer states, | |
| intermediate checkpoints, `training_args.bin`, evaluation bundles, and all log | |
| files. | |
| After download, verify the release from its repository directory: | |
| ```bash | |
| sha256sum -c MODEL_MANIFEST.sha256 | |
| ``` | |
| ## ⚠️ Limitations | |
| Generated searches and answers can be incorrect, unsupported, or unsafe; | |
| retrieval and answer verification remain the caller's responsibility. No | |
| evaluation result is claimed by this model card. | |
| ## 📜 Terms | |
| No additional blanket license is asserted here. Review the | |
| [Qwen2.5-3B-Instruct license](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE) | |
| and the [AetherSearch SFT data attribution and rights status](https://github.com/Muradil-mamat-211/AetherSearch/blob/main/sft/ATTRIBUTION.md) | |
| before redistribution or downstream use. | |
| --- | |
| <div align="center"> | |
| **Built for experiments in agentic search and retrieval-augmented reasoning.** 🔎✨ | |
| </div> | |