Reference¶
Serving behavior of the AutoGluon container. These settings apply when the image runs as a SageMaker endpoint; training jobs follow the standard SageMaker training toolkit conventions.
Inference handler¶
At startup the server imports the handler file named by SAGEMAKER_PROGRAM from /opt/ml/model/code/ (the code/ directory of your model artifact). The handler must define two functions:
Function |
Signature |
Purpose |
|---|---|---|
|
|
Called once per worker at startup with |
|
|
Called for every |
Request and response handling:
request_bodyis astrfortext/*,application/json, andapplication/jsonlrequests, andbytesfor every other content type.input_content_typecomes from theContent-Typeheader (defaultapplication/json).output_content_typecomes from theAcceptheader, falling back toSAGEMAKER_DEFAULT_INVOCATIONS_ACCEPT.transform_fnmust return a(body, content_type)tuple wherebodyisstrorbytes.Raising
ValueErrorreturns HTTP 400 with the error message. Any other exception returns HTTP 500 and is logged to CloudWatch.input_fn,predict_fn, andoutput_fnare not used. Put that logic intransform_fn.
Note
Migrating from the AutoGluon 1.5 autogluon-inference image? Handlers that split their logic across input_fn / predict_fn / output_fn must be merged into a single transform_fn.
Extra Python packages¶
If the model artifact contains code/requirements.txt, the container installs it with uv pip install once at startup, before the server starts.
To install from a private AWS CodeArtifact repository, set CA_REPOSITORY_ARN (see below). The execution role then needs codeartifact:GetAuthorizationToken, codeartifact:GetRepositoryEndpoint, codeartifact:ReadFromRepository, and sts:GetServiceBearerToken.
Environment variables¶
Set these in the Environment of the SageMaker model.
Variable |
Default |
Description |
|---|---|---|
|
|
Handler file, relative to |
|
|
|
|
|
Number of Gunicorn worker processes. Each worker loads its own copy of the model. |
|
|
Per-request worker timeout, in seconds |
|
|
Server log level ( |
|
unset |
CodeArtifact repository ARN ( |
SageMaker sets SAGEMAKER_BIND_TO_PORT automatically; otherwise the server listens on port 8080.
Known limitations¶
x86 only. There are no ARM64 images.
One model per container. Multi-model endpoints are not supported.
Model loading happens before
/pingsucceeds. RaiseContainerStartupHealthCheckTimeoutInSecondsin theProductionVariantfor large models.