vLLM/SGLang Compatibility & DynamicCache Bug Report

#6
by AA-C - opened
  1. vLLM and SGLang Serving Failure (ValueError: Model architectures not supported)
    Location: Repository setup / config.json

Issue: The Hugging Face repo automatically suggests serving commands such as vllm serve "bharatgenai/AyurParam" and python3 -m sglang.launch_server --model-path "bharatgenai/AyurParam".

Root Cause: Custom architectures relying on trust_remote_code=True / auto_map (ParamBharatGenForCausalLM) are not supported natively by vLLM or SGLang out of the box without defining custom model implementations or registering custom C++/CUDA kernels inside their respective codebases.

Error Triggered:

ValueError: Model architectures ['ParamBharatGenForCausalLM'] are not supported yet.
Suggested Fix: Provide instructions for overriding the model architecture locally (if compatible with Llama/Gemma), or clarify that the model requires a standard Hugging Face Transformers pipeline/FastAPI server for deployment.

  1. Transformers 4.40+ DynamicCache Breaking Change
    Location: modeling_parambharatgen.py

Issue: Generating text using recent versions of the transformers library (4.40+) fails during model inference.

Root Cause: Modern transformers generation passes DynamicCache objects for past_key_values rather than legacy tuples. The indexing and concatenation logic inside ParamBharatGenAttention in modeling_parambharatgen.py assumes past_key_values is a tuple of tensors, raising AttributeError / TypeError.

Suggested Fix: Update modeling_parambharatgen.py to support transformers.cache_utils.Cache / DynamicCache interface methods (cache_kwargs, .update()).

  1. Unused / Backup Configuration Files
    Location: tokenizer_config.json.bak, chat_template_bck.jinja

Issue: The repository contains residual .bak and _bck configuration files from development.

Suggested Fix: Clean up unused files to avoid confusion when downloading or cloning the repository programmatically.

Environment Information:
Model: bharatgenai/AyurParam

Library: transformers >= 4.40.0, vLLM >= 0.4.0

Sign up or log in to comment