Buckets:
hf-doc-build/doc-dev / sagemaker /pr_2709 /en /tutorials /sagemaker-sdk /sagemaker-sdk-quickstart.md
Amazon SageMaker SDK Quickstart
Deploy a model from the Hugging Face Hub to a live SageMaker endpoint in a few minutes with the SageMaker Python SDK.
1 · Deploy
Point ModelBuilder at a Hub model ID and create the endpoint.
2 · Invoke
Send a JSON request to the live endpoint and get a prediction.
3 · Delete
One call deletes the endpoint and stops all charges.
Prerequisites
- An AWS account. If you do not have one, follow the AWS setup guide.
- The SageMaker Python SDK v3, which provides
ModelBuilderfor inference andModelTrainerfor training:
pip install "sagemaker>=3.0.0"
- An IAM execution role. In SageMaker Studio or on a SageMaker notebook instance,
get_execution_role()returns it automatically. In a local environment you must pass the role ARN yourself — both setups are shown in Set up the SageMaker SDK.
Deploy a model from the Hub
Create a session, then point ModelBuilder at a model ID from the Hub:
from sagemaker.core.helper.session_helper import Session, get_execution_role
from sagemaker.serve import ModelBuilder, ModelServer
from sagemaker.serve.builder.schema_builder import SchemaBuilder
from sagemaker.core import image_uris
sess = Session()
role = get_execution_role()
# any Hub model works
model_id = "cardiffnlp/twitter-roberta-base-sentiment-latest"
instance_type = "ml.m5.xlarge"
# Retrieve the Hugging Face PyTorch inference DLC image URI
inference_image = image_uris.retrieve(
framework="huggingface",
region=sess.boto_region_name,
# Transformers version
version="4.51.3",
base_framework_version="pytorch2.6.0", # PyTorch version
# Python version
py_version="py312",
image_scope="inference",
instance_type=instance_type,
)
# Sample request/response used by ModelBuilder to set up serialization
sample_input = {"inputs": "I love how simple this was!"}
sample_output = [{"label": "positive", "score": 0.99}]
model_builder = ModelBuilder(
# Hub model ID, loaded at deploy time
model=model_id,
model_server=ModelServer.MMS,
image_uri=inference_image,
# tells the Inference Toolkit which pipeline to serve
env_vars={"HF_TASK": "text-classification"},
role_arn=role,
sagemaker_session=sess,
instance_type=instance_type,
schema_builder=SchemaBuilder(sample_input=sample_input, sample_output=sample_output),
)
model_builder.build()
predictor = model_builder.deploy(initial_instance_count=1, instance_type=instance_type)
Invoke the endpoint
The request and response bodies are JSON, and every request needs an inputs key:
import json
res = predictor.invoke(
body=json.dumps({"inputs": "I love how simple this was!"}),
content_type="application/json",
)
print(json.loads(res.body.read()))
Clean up
Delete the endpoint when you are done:
predictor.delete()
What's next
- Deploy models covers deploying models you trained in SageMaker or stored in S3, batch transform jobs, and custom inference code.
- Train models covers
ModelTrainer: training scripts, distributed training, spot instances, and metrics. - For a full LLM recipe, see the example Fine-Tuning LLMs with TRL CLI on SageMaker.
Xet Storage Details
- Size:
- 3.45 kB
- Xet hash:
- 1e5e09f44d372193820f2c29dc64db8e6300308fdd43918f92357a2a71681cf7
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.