Instructions to use stabilityai/japanese-stablelm-instruct-beta-7b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use stabilityai/japanese-stablelm-instruct-beta-7b with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="stabilityai/japanese-stablelm-instruct-beta-7b")

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("stabilityai/japanese-stablelm-instruct-beta-7b")
model = AutoModelForCausalLM.from_pretrained("stabilityai/japanese-stablelm-instruct-beta-7b")

Inference
Notebooks
Google Colab
Kaggle
Local Apps Settings

vLLM

How to use stabilityai/japanese-stablelm-instruct-beta-7b with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "stabilityai/japanese-stablelm-instruct-beta-7b"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "stabilityai/japanese-stablelm-instruct-beta-7b",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker

docker model run hf.co/stabilityai/japanese-stablelm-instruct-beta-7b

SGLang

How to use stabilityai/japanese-stablelm-instruct-beta-7b with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "stabilityai/japanese-stablelm-instruct-beta-7b" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "stabilityai/japanese-stablelm-instruct-beta-7b",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "stabilityai/japanese-stablelm-instruct-beta-7b" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "stabilityai/japanese-stablelm-instruct-beta-7b",
		"prompt": "Once upon a time,",
		"max_tokens": 512,
		"temperature": 0.5
	}'

Docker Model Runner
How to use stabilityai/japanese-stablelm-instruct-beta-7b with Docker Model Runner:
```
docker model run hf.co/stabilityai/japanese-stablelm-instruct-beta-7b
```
Browse Quantizations to use this model in llama.cpp, Ollama, LM Studio, or any compatible app.

Japanese-StableLM-Instruct-Beta-7B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

Model Description

japanese-stablelm-instruct-beta-7b is a 7B-parameter decoder-only language model based on japanese-stablelm-base-beta-7b and further fine tuned on Databricks Dolly-15k, Anthropic HH, and other public data.

This model is also available in a larger 70b version, or a faster version with a specialized tokenizer.

Usage

First install additional dependencies in requirements.txt:

pip install -r requirements.txt

Then start generating text with japanese-stablelm-instruct-beta-7b by using the following code snippet:

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_name = "stabilityai/japanese-stablelm-instruct-beta-7b"
tokenizer = AutoTokenizer.from_pretrained(model_name)

# The next line may need to be modified depending on the environment
model = AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16, low_cpu_mem_usage=True, device_map="auto")

def build_prompt(user_query, inputs):
    sys_msg = "<s>[INST] <<SYS>>\nあなたは役立つアシスタントです。\n<<SYS>>\n\n"
    p = sys_msg + user_query + "\n\n" + inputs + " [/INST] "
    return p

user_inputs = {
    "user_query": "与えられたことわざの意味を小学生でも分かるように教えてください。",
    "inputs": "情けは人のためならず"
}
prompt = build_prompt(**user_inputs)

input_ids = tokenizer.encode(
    prompt,
    add_special_tokens=True,
    return_tensors="pt"
)

# this is for reproducibility.
# feel free to change to get different result
seed = 23  
torch.manual_seed(seed)

tokens = model.generate(
    input_ids.to(device=model.device),
    max_new_tokens=128,
    temperature=0.99,
    top_p=0.95,
    do_sample=True,
)

out = tokenizer.decode(tokens[0], skip_special_tokens=True)
print(out)

We suggest playing with different generation config (top_p, repetition_penalty etc) to find the best setup for your tasks. For example, use higher temperature for roleplay task, lower temperature for reasoning.

Model Details

Model type: japanese-stablelm-instruct-beta-7b model is an auto-regressive language model based on the Llama2 transformer architecture.
Language(s): Japanese
License: Llama2 Community License.
Contact: For questions and comments about the model, please join Stable Community Japan. For future announcements / information about Stability AI models, research, and events, please follow https://twitter.com/StabilityAI_JP.

Training Dataset

The following datasets were used for the instruction training. Note these are Japanese translated versions of the original datasets, shared by kunishou.

Use and Limitations

Intended Use

The model is intended to be used by all individuals as a foundation for application-specific fine-tuning without strict limitations on commercial use.

Limitations and bias

The pre-training dataset may have contained offensive or inappropriate content even after applying data cleansing filters which can be reflected in the model generated text. We recommend users exercise reasonable caution when using these models in production systems. Do not use the model for any applications that may cause harm or distress to individuals or groups.

Authors

This model was developed by the Research & Development team at Stability AI Japan, and the development was co-led by Takuya Akiba and Meng Lee. The members of the team are as follows:

Acknowledgements

We thank Meta Research for releasing Llama 2 under an open license for others to build on.

We are grateful for the contributions of the EleutherAI Polyglot-JA team in helping us to collect a large amount of pre-training data in Japanese. Polyglot-JA members includes Hyunwoong Ko (Project Lead), Fujiki Nakamura (originally started this project when he commited to the Polyglot team), Yunho Mo, Minji Jung, KeunSeok Im, and Su-Kyeong Jang.

We are also appreciative of AI Novelist/Sta (Bit192, Inc.) and the numerous contributors from Stable Community Japan for assisting us in gathering a large amount of high-quality Japanese textual data for model training.

Downloads last month: 62

Safetensors

Model size

7B params

Tensor type

F32

F16

Model tree for stabilityai/japanese-stablelm-instruct-beta-7b

Quantizations

3 models

Datasets used to train stabilityai/japanese-stablelm-instruct-beta-7b

Collection including stabilityai/japanese-stablelm-instruct-beta-7b

Japanese Stable LM

Collection

Suite of LLMs focusing on Japanese usage • 15 items • Updated Jan 9, 2025 • 21