Deploying on AWS documentation

Amazon SageMaker SDK Quickstart

Hugging Face's logo
Join the Hugging Face community

and get access to the augmented documentation experience

to get started

Amazon SageMaker SDK Quickstart

Deploy a model from the Hugging Face Hub to a live SageMaker endpoint in a few minutes with the SageMaker Python SDK.

1 · Deploy

Point ModelBuilder at a Hub model ID and create the endpoint.

2 · Invoke

Send a JSON request to the live endpoint and get a prediction.

3 · Delete

One call deletes the endpoint and stops all charges.

Prerequisites

  • An AWS account. If you do not have one, follow the AWS setup guide.
  • The SageMaker Python SDK v3, which provides ModelBuilder for inference and ModelTrainer for training:
pip install "sagemaker>=3.0.0"
  • An IAM execution role. In SageMaker Studio or on a SageMaker notebook instance, get_execution_role() returns it automatically. In a local environment you must pass the role ARN yourself — both setups are shown in Set up the SageMaker SDK.

Deploy a model from the Hub

Create a session, then point ModelBuilder at a model ID from the Hub:

from sagemaker.core.helper.session_helper import Session, get_execution_role
from sagemaker.serve import ModelBuilder, ModelServer
from sagemaker.serve.builder.schema_builder import SchemaBuilder
from sagemaker.core import image_uris

sess = Session()
role = get_execution_role()

# any Hub model works
model_id = "cardiffnlp/twitter-roberta-base-sentiment-latest"
instance_type = "ml.m5.xlarge"

# Retrieve the Hugging Face PyTorch inference DLC image URI
inference_image = image_uris.retrieve(
    framework="huggingface",
    region=sess.boto_region_name,
    # Transformers version
    version="4.51.3",
    base_framework_version="pytorch2.6.0", # PyTorch version
    # Python version
    py_version="py312",
    image_scope="inference",
    instance_type=instance_type,
)

# Sample request/response used by ModelBuilder to set up serialization
sample_input = {"inputs": "I love how simple this was!"}
sample_output = [{"label": "positive", "score": 0.99}]

model_builder = ModelBuilder(
    # Hub model ID, loaded at deploy time
    model=model_id,
    model_server=ModelServer.MMS,
    image_uri=inference_image,
    # tells the Inference Toolkit which pipeline to serve
    env_vars={"HF_TASK": "text-classification"},
    role_arn=role,
    sagemaker_session=sess,
    instance_type=instance_type,
    schema_builder=SchemaBuilder(sample_input=sample_input, sample_output=sample_output),
)
model_builder.build()

predictor = model_builder.deploy(initial_instance_count=1, instance_type=instance_type)

Invoke the endpoint

The request and response bodies are JSON, and every request needs an inputs key:

import json

res = predictor.invoke(
    body=json.dumps({"inputs": "I love how simple this was!"}),
    content_type="application/json",
)
print(json.loads(res.body.read()))

Clean up

Delete the endpoint when you are done:

predictor.delete()

What’s next

Update on GitHub