Buckets:
hf-doc-build/doc-dev / sagemaker /pr_2709 /en /tutorials /sagemaker-sdk /sagemaker-sdk-quickstart.md
| # Amazon SageMaker SDK Quickstart | |
| Deploy a model from the Hugging Face Hub to a live SageMaker endpoint in a few minutes with the SageMaker Python SDK. | |
| 1 · Deploy | |
| Point ModelBuilder at a Hub model ID and create the endpoint. | |
| 2 · Invoke | |
| Send a JSON request to the live endpoint and get a prediction. | |
| 3 · Delete | |
| One call deletes the endpoint and stops all charges. | |
| ## Prerequisites | |
| - An AWS account. If you do not have one, follow the [AWS setup guide](https://docs.aws.amazon.com/sagemaker/latest/dg/gs-set-up.html). | |
| - The SageMaker Python SDK v3, which provides `ModelBuilder` for inference and `ModelTrainer` for training: | |
| ```bash | |
| pip install "sagemaker>=3.0.0" | |
| ``` | |
| - An IAM execution role. In SageMaker Studio or on a SageMaker notebook instance, `get_execution_role()` returns it automatically. In a local environment you must pass the role ARN yourself — both setups are shown in [Set up the SageMaker SDK](./setup-sagemaker-sdk). | |
| ## Deploy a model from the Hub | |
| Create a session, then point `ModelBuilder` at a model ID from the Hub: | |
| ```python | |
| from sagemaker.core.helper.session_helper import Session, get_execution_role | |
| from sagemaker.serve import ModelBuilder, ModelServer | |
| from sagemaker.serve.builder.schema_builder import SchemaBuilder | |
| from sagemaker.core import image_uris | |
| sess = Session() | |
| role = get_execution_role() | |
| # any Hub model works | |
| model_id = "cardiffnlp/twitter-roberta-base-sentiment-latest" | |
| instance_type = "ml.m5.xlarge" | |
| # Retrieve the Hugging Face PyTorch inference DLC image URI | |
| inference_image = image_uris.retrieve( | |
| framework="huggingface", | |
| region=sess.boto_region_name, | |
| # Transformers version | |
| version="4.51.3", | |
| base_framework_version="pytorch2.6.0", # PyTorch version | |
| # Python version | |
| py_version="py312", | |
| image_scope="inference", | |
| instance_type=instance_type, | |
| ) | |
| # Sample request/response used by ModelBuilder to set up serialization | |
| sample_input = {"inputs": "I love how simple this was!"} | |
| sample_output = [{"label": "positive", "score": 0.99}] | |
| model_builder = ModelBuilder( | |
| # Hub model ID, loaded at deploy time | |
| model=model_id, | |
| model_server=ModelServer.MMS, | |
| image_uri=inference_image, | |
| # tells the Inference Toolkit which pipeline to serve | |
| env_vars={"HF_TASK": "text-classification"}, | |
| role_arn=role, | |
| sagemaker_session=sess, | |
| instance_type=instance_type, | |
| schema_builder=SchemaBuilder(sample_input=sample_input, sample_output=sample_output), | |
| ) | |
| model_builder.build() | |
| predictor = model_builder.deploy(initial_instance_count=1, instance_type=instance_type) | |
| ``` | |
| ## Invoke the endpoint | |
| The request and response bodies are JSON, and every request needs an `inputs` key: | |
| ```python | |
| import json | |
| res = predictor.invoke( | |
| body=json.dumps({"inputs": "I love how simple this was!"}), | |
| content_type="application/json", | |
| ) | |
| print(json.loads(res.body.read())) | |
| ``` | |
| ## Clean up | |
| Delete the endpoint when you are done: | |
| ```python | |
| predictor.delete() | |
| ``` | |
| ## What's next | |
| - [Deploy models](./deploy-sagemaker-sdk) covers deploying models you trained in SageMaker or stored in S3, batch transform jobs, and custom inference code. | |
| - [Train models](./training-sagemaker-sdk) covers `ModelTrainer`: training scripts, distributed training, spot instances, and metrics. | |
| - For a full LLM recipe, see the example [Fine-Tuning LLMs with TRL CLI on SageMaker](../../examples/sagemaker-sdk-fine-tune-trl-cli). | |
Xet Storage Details
- Size:
- 3.45 kB
- Xet hash:
- 1e5e09f44d372193820f2c29dc64db8e6300308fdd43918f92357a2a71681cf7
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.