Instructions to use microsoft/Phi-3-small-128k-instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use microsoft/Phi-3-small-128k-instruct with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="microsoft/Phi-3-small-128k-instruct", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("microsoft/Phi-3-small-128k-instruct", trust_remote_code=True, dtype="auto")

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use microsoft/Phi-3-small-128k-instruct with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "microsoft/Phi-3-small-128k-instruct"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "microsoft/Phi-3-small-128k-instruct",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/microsoft/Phi-3-small-128k-instruct

SGLang

How to use microsoft/Phi-3-small-128k-instruct with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "microsoft/Phi-3-small-128k-instruct" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "microsoft/Phi-3-small-128k-instruct",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "microsoft/Phi-3-small-128k-instruct" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "microsoft/Phi-3-small-128k-instruct",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use microsoft/Phi-3-small-128k-instruct with Docker Model Runner:
```
docker model run hf.co/microsoft/Phi-3-small-128k-instruct
```

mwirth-epo commited on Jul 23, 2024

Commit

8993bb8

verified ·

1 Parent(s): f80aaa3

update positional_embedding.py

Browse files

Relates to bf16 for query_states and key_states issue.
Apply fix in https://huggingface.co/microsoft/Phi-3-small-8k-instruct/commit/f196467b67c13127747a03c142e09aa6841447b8 also for this model

Files changed (1) hide show

positional_embedding.py +3 -3

positional_embedding.py CHANGED Viewed

@@ -269,10 +269,10 @@ class RotaryEmbedding(torch.nn.Module):
         return (
             apply_rotary_pos_emb(
                 q, cos_cached[seqlen_offset:seq_len], sin_cached[seqlen_offset:seq_len], seq_dimension=seq_dimension
-            ),
             apply_rotary_pos_emb(
                 k, cos_cached[seqlen_offset:seq_len], sin_cached[seqlen_offset:seq_len], seq_dimension=seq_dimension
-            ),
         )
     @classmethod
@@ -285,4 +285,4 @@ class RotaryEmbedding(torch.nn.Module):
         )
         if config.rope_scaling is not None:
             kwargs["longrope_config"] = LongRopeConfig.from_dict(config.rope_scaling)
-        return cls(**kwargs)

         return (
             apply_rotary_pos_emb(
                 q, cos_cached[seqlen_offset:seq_len], sin_cached[seqlen_offset:seq_len], seq_dimension=seq_dimension
+            ).to(q.dtype),
             apply_rotary_pos_emb(
                 k, cos_cached[seqlen_offset:seq_len], sin_cached[seqlen_offset:seq_len], seq_dimension=seq_dimension
+            ).to(k.dtype),
         )
     @classmethod
         )
         if config.rope_scaling is not None:
             kwargs["longrope_config"] = LongRopeConfig.from_dict(config.rope_scaling)
+        return cls(**kwargs)