Instructions to use bharatgenai/AyurParam with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bharatgenai/AyurParam with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="bharatgenai/AyurParam", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("bharatgenai/AyurParam", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use bharatgenai/AyurParam with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "bharatgenai/AyurParam" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bharatgenai/AyurParam", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/bharatgenai/AyurParam
- SGLang
How to use bharatgenai/AyurParam with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "bharatgenai/AyurParam" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bharatgenai/AyurParam", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "bharatgenai/AyurParam" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "bharatgenai/AyurParam", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use bharatgenai/AyurParam with Docker Model Runner:
docker model run hf.co/bharatgenai/AyurParam
vLLM/SGLang Compatibility & DynamicCache Bug Report
- vLLM and SGLang Serving Failure (ValueError: Model architectures not supported)
Location: Repository setup / config.json
Issue: The Hugging Face repo automatically suggests serving commands such as vllm serve "bharatgenai/AyurParam" and python3 -m sglang.launch_server --model-path "bharatgenai/AyurParam".
Root Cause: Custom architectures relying on trust_remote_code=True / auto_map (ParamBharatGenForCausalLM) are not supported natively by vLLM or SGLang out of the box without defining custom model implementations or registering custom C++/CUDA kernels inside their respective codebases.
Error Triggered:
ValueError: Model architectures ['ParamBharatGenForCausalLM'] are not supported yet.
Suggested Fix: Provide instructions for overriding the model architecture locally (if compatible with Llama/Gemma), or clarify that the model requires a standard Hugging Face Transformers pipeline/FastAPI server for deployment.
- Transformers 4.40+ DynamicCache Breaking Change
Location: modeling_parambharatgen.py
Issue: Generating text using recent versions of the transformers library (4.40+) fails during model inference.
Root Cause: Modern transformers generation passes DynamicCache objects for past_key_values rather than legacy tuples. The indexing and concatenation logic inside ParamBharatGenAttention in modeling_parambharatgen.py assumes past_key_values is a tuple of tensors, raising AttributeError / TypeError.
Suggested Fix: Update modeling_parambharatgen.py to support transformers.cache_utils.Cache / DynamicCache interface methods (cache_kwargs, .update()).
- Unused / Backup Configuration Files
Location: tokenizer_config.json.bak, chat_template_bck.jinja
Issue: The repository contains residual .bak and _bck configuration files from development.
Suggested Fix: Clean up unused files to avoid confusion when downloading or cloning the repository programmatically.
Environment Information:
Model: bharatgenai/AyurParam
Library: transformers >= 4.40.0, vLLM >= 0.4.0