Text Generation
Transformers
Safetensors
English
ivme_conversate_s_v2_instruct
from-scratch
experimental
causal-lm
small-language-model
instruct-pretrained
custom_code
Instructions to use IvmeLabs/Ivme-Conversate-S-v2-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use IvmeLabs/Ivme-Conversate-S-v2-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="IvmeLabs/Ivme-Conversate-S-v2-Instruct", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("IvmeLabs/Ivme-Conversate-S-v2-Instruct", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use IvmeLabs/Ivme-Conversate-S-v2-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "IvmeLabs/Ivme-Conversate-S-v2-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IvmeLabs/Ivme-Conversate-S-v2-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/IvmeLabs/Ivme-Conversate-S-v2-Instruct
- SGLang
How to use IvmeLabs/Ivme-Conversate-S-v2-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "IvmeLabs/Ivme-Conversate-S-v2-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IvmeLabs/Ivme-Conversate-S-v2-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "IvmeLabs/Ivme-Conversate-S-v2-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "IvmeLabs/Ivme-Conversate-S-v2-Instruct", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use IvmeLabs/Ivme-Conversate-S-v2-Instruct with Docker Model Runner:
docker model run hf.co/IvmeLabs/Ivme-Conversate-S-v2-Instruct
| license: apache-2.0 | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| language: | |
| - en | |
| tags: | |
| - from-scratch | |
| - experimental | |
| - causal-lm | |
| - small-language-model | |
| - instruct-pretrained | |
| datasets: | |
| - HuggingFaceH4/ultrachat_200k | |
| - allenai/soda | |
| - openbmb/UltraInteract_sft | |
| - microsoft/orca-math-word-problems-200k | |
| - databricks/databricks-dolly-15k | |
| - b-mc2/sql-create-context | |
| # Ivme-Conversate-S-v2-Instruct | |
| 9,021,600 parameters. Standard decoder-only Transformer (tied | |
| embeddings, multi-head attention, RoPE, SwiGLU, RMSNorm) -- matching | |
| Ivme-Conversate-v2-Base's proven recipe exactly, deliberately with zero | |
| architectural novelty. | |
| Trained single-epoch on ~900M tokens, instruct-heavy from the start rather | |
| than base-pretrain-then-finetune: UltraChat-200k (real multi-turn dialogue) | |
| as the dominant 45% share, plus SODA, UltraInteract reasoning traces, | |
| orca-math, dolly-15k instructions, and sql-create-context. All sources | |
| permissively licensed (MIT/CC-BY/CC-BY-SA). | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model = AutoModelForCausalLM.from_pretrained( | |
| "ivmelabs/Ivme-Conversate-S-v2-Instruct", trust_remote_code=True | |
| ) | |
| tok = AutoTokenizer.from_pretrained("ivmelabs/Ivme-Conversate-S-v2-Instruct") | |
| ids = tok("Hello!", return_tensors="pt").input_ids | |
| out = model.generate(ids, max_new_tokens=80, do_sample=True, temperature=0.8, top_k=40) | |
| print(tok.decode(out[0])) | |
| ``` | |
| Note: no KV-cache in this architecture -- `.generate()` works but is O(n^2) | |
| rather than O(n), fine for short samples, not tuned for long-form serving. | |