Instructions to use afifaimran/meno_llm_3b_v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use afifaimran/meno_llm_3b_v2 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("hishab/titulm-llama-3.2-3b-v2.0") model = PeftModel.from_pretrained(base_model, "afifaimran/meno_llm_3b_v2") - Transformers
How to use afifaimran/meno_llm_3b_v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="afifaimran/meno_llm_3b_v2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("afifaimran/meno_llm_3b_v2", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use afifaimran/meno_llm_3b_v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "afifaimran/meno_llm_3b_v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "afifaimran/meno_llm_3b_v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/afifaimran/meno_llm_3b_v2
- SGLang
How to use afifaimran/meno_llm_3b_v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "afifaimran/meno_llm_3b_v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "afifaimran/meno_llm_3b_v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "afifaimran/meno_llm_3b_v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "afifaimran/meno_llm_3b_v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Studio
How to use afifaimran/meno_llm_3b_v2 with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for afifaimran/meno_llm_3b_v2 to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for afifaimran/meno_llm_3b_v2 to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for afifaimran/meno_llm_3b_v2 to start chatting
Load model with FastModel
pip install unsloth from unsloth import FastModel model, tokenizer = FastModel.from_pretrained( model_name="afifaimran/meno_llm_3b_v2", max_seq_length=2048, ) - Docker Model Runner
How to use afifaimran/meno_llm_3b_v2 with Docker Model Runner:
docker model run hf.co/afifaimran/meno_llm_3b_v2
meno_llm_3b_v2
This repository contains a LoRA adapter fine-tuned on top of:
Base model: hishab/titulm-llama-3.2-3b-v2.0
Description
This model is being developed as a Bangla empathetic conversational assistant for health-related question answering.
It is intended to be used in a retrieval-augmented generation (RAG) pipeline, where relevant context is retrieved first and then passed to the model for response generation.
Important Note
This repository contains only the LoRA adapter, not the full standalone base model.
To use it, load it on top of:
hishab/titulm-llama-3.2-3b-v2.0
Current Status
- Fine-tuned adapter uploaded successfully
- Produces empathetic Bangla answers
- Works better when used with retrieved supporting context
- Still under improvement for grounding, medical faithfulness, and weak-context handling
Intended Use
- Bangla conversational response generation
- Empathetic health-related QA
- RAG-based assistant workflows
- Research and development use
Limitations
- This model is still under active development
- It may produce incorrect or weakly grounded answers if retrieval quality is poor
- It should not be treated as a substitute for professional medical advice
Loading
Use this adapter together with the base TiTuLM model.
Base model:
hishab/titulm-llama-3.2-3b-v2.0
Adapter repo:
afifaimran/meno_llm_3b_v2
Author
Afifa Imran
- Downloads last month
- 1
Model tree for afifaimran/meno_llm_3b_v2
Base model
hishab/titulm-llama-3.2-3b-v2.0