Run Pretrained Tabular Foundation Models on Amazon SageMaker¶
Tabular foundation models are large pretrained models that make predictions on a new dataset in context: you pass a set of labeled rows together with the rows to predict, and the model predicts them in a single forward pass. Because they’re pretrained on large collections of diverse tables, they generalize to unseen tables out of the box — no hyperparameter tuning or lengthy training required.
That makes the workflow much simpler than training your own tabular predictor, which requires you to first fit a predictor on your data and then manage the trained artifact. With foundation models you skip the fit step entirely and go straight to running batch predictions or deploying an endpoint.
AutoGluon-Cloud exposes this workflow through TabularFoundationModel, with models like Mitra, TabICLv2, TabDPT-Turbo and Nori available out of the box. For time series forecasting, see Time Series Foundation Models.
Create the model¶
Important
Before running any code below, follow the Setup tutorial to register the IAM role and S3 bucket that SageMaker will use. The examples assume those resources are saved in ~/.autogluon/cloud.yaml.
from autogluon.cloud import TabularFoundationModel
model = TabularFoundationModel(model_id="mitra-classifier")
The rest of the tutorial reuses this model object.
Available models¶
The following model_id values are currently supported (also available via TabularFoundationModel.list_models()). Each model_id targets a single task — pick a *-classifier model for classification (binary or multiclass) and a *-regressor model for regression.
Model ID |
Task |
Documentation |
Weights |
|---|---|---|---|
|
Classification |
||
|
Regression |
||
|
Classification |
||
|
Classification |
||
|
Regression |
||
|
Classification |
||
|
Regression |
||
|
Regression |
||
|
Regression |
Mitra runs on CPU by default; the other models default to GPU instances. For background on tabular foundation models in AutoGluon, see the Tabular Foundational Models tutorial.
Data¶
The examples use the Adult Income dataset, where the task is to predict whether a person earns more than $50K a year. Tabular foundation models are designed for small to medium context sizes, so we take a random sample of the training set as the labeled context:
import pandas as pd
train_data = pd.read_csv("https://autogluon.s3.amazonaws.com/datasets/Inc/train.csv").sample(2000, random_state=0)
train_data.head()
| age | workclass | fnlwgt | education | education-num | marital-status | occupation | relationship | race | sex | capital-gain | capital-loss | hours-per-week | native-country | class | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 6118 | 51 | Private | 39264 | Some-college | 10 | Married-civ-spouse | Exec-managerial | Wife | White | Female | 0 | 0 | 40 | United-States | >50K |
| 23204 | 58 | Private | 51662 | 10th | 6 | Married-civ-spouse | Other-service | Wife | White | Female | 0 | 0 | 8 | United-States | <=50K |
| 29590 | 40 | Private | 326310 | Some-college | 10 | Married-civ-spouse | Craft-repair | Husband | White | Male | 0 | 0 | 44 | United-States | <=50K |
| 18116 | 37 | Private | 222450 | HS-grad | 9 | Never-married | Sales | Not-in-family | White | Male | 0 | 2339 | 40 | El-Salvador | <=50K |
| 33964 | 62 | Private | 109190 | Bachelors | 13 | Married-civ-spouse | Exec-managerial | Husband | White | Male | 15024 | 0 | 40 | United-States | >50K |
The input is a regular tabular DataFrame: one row per example, one column per feature, plus the label column — here, class. The rows to predict must contain the same feature columns. Drop the label from the test split since it’s what we want to predict:
test_data = (
pd.read_csv("https://autogluon.s3.amazonaws.com/datasets/Inc/test.csv")
.sample(500, random_state=0)
.drop(columns=["class"])
)
test_data.head()
| age | workclass | fnlwgt | education | education-num | marital-status | occupation | relationship | race | sex | capital-gain | capital-loss | hours-per-week | native-country | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 5454 | 41 | Self-emp-not-inc | 408498 | HS-grad | 9 | Married-civ-spouse | Exec-managerial | Husband | White | Male | 0 | 0 | 50 | United-States |
| 6111 | 39 | Private | 746786 | Bachelors | 13 | Married-civ-spouse | Prof-specialty | Husband | White | Male | 0 | 0 | 55 | United-States |
| 5282 | 50 | Private | 62593 | Assoc-voc | 11 | Married-civ-spouse | Farming-fishing | Husband | Asian-Pac-Islander | Male | 0 | 0 | 40 | United-States |
| 3046 | 31 | Private | 248178 | Some-college | 10 | Married-civ-spouse | Other-service | Husband | Black | Male | 0 | 0 | 35 | United-States |
| 2162 | 43 | State-gov | 52849 | Bachelors | 13 | Married-civ-spouse | Prof-specialty | Husband | White | Male | 0 | 0 | 40 | United-States |
Inference modes¶
TabularFoundationModel supports two inference modes on SageMaker. In both, every call sends the labeled train_data as context along with the rows to predict, and AutoGluon fits a TabularPredictor wrapping the foundation model on that context before predicting.
Batch prediction — launch a one-off SageMaker job that scores a dataset and writes the results to S3. Compute spins up, runs, and shuts down automatically. Best for offline scoring of larger datasets where minutes of startup latency are fine.
Real-time inference — deploy the model to a long-running SageMaker endpoint and send requests over HTTPS. Lowest per-request latency, supports GPU instances. You pay for the endpoint as long as it’s up, so best when you need predictions on demand and have steady traffic.
Note
Serverless inference is not supported for tabular foundation models, since SageMaker Serverless Inference does not provide sufficient resources to run them.
The examples below all reuse the train_data and test_data DataFrames loaded above.
Batch prediction¶
Use predict() to score a dataset as a one-off job. It returns a Series of predictions:
predictions = model.predict(
test_data=test_data,
train_data=train_data,
label="class",
)
For classification, use predict_proba() to get class probabilities. By default it returns both the predictions and the probabilities, computed in the same job:
predictions, probabilities = model.predict_proba(
test_data=test_data,
train_data=train_data,
label="class",
)
The job also writes the predictions to S3 as a CSV. By default they land at {cloud_output_path}/{job_name}/predictions.csv; pass predictions_path to choose an explicit destination:
predictions = model.predict(
test_data=test_data,
train_data=train_data,
label="class",
predictions_path="s3://my-bucket/predictions/2026-06-02.csv",
)
For long-running jobs you can return immediately with wait=False. predict() then returns a JobPredictionFuture you can poll with .status() and resolve with .result():
future = model.predict(
test_data=test_data,
train_data=train_data,
label="class",
wait=False,
)
print(future.job_name, future.status()) # 'ag-...', 'InProgress'
predictions = future.result() # blocks until the job finishes, returns a Series
Real-time inference¶
Deploy the model to a SageMaker endpoint with deploy(), then send requests through the returned TabularEndpoint. Pick an instance_type based on cost and latency requirements (defaults to ml.m5.4xlarge for Mitra and ml.g5.xlarge for the other models):
endpoint = model.deploy(instance_type="ml.m5.4xlarge") # takes a few minutes
predictions = endpoint.predict(
data=test_data,
train_data=train_data,
label="class",
)
predictions, probabilities = endpoint.predict_proba(
data=test_data,
train_data=train_data,
label="class",
)
The endpoint holds no data between requests, so each request carries its own labeled context. This means you can use a single endpoint to serve predictions for different datasets and tasks of the same type — for example, a mitra-classifier endpoint can classify rows from any table, as long as each request includes the matching train_data and label.
Each request payload — train_data and data combined — is limited to 6 MB. For larger inputs, use batch prediction instead.
The endpoint stays active — and billed — until you delete it:
endpoint.delete_endpoint()
Invoke the endpoint without AutoGluon-Cloud¶
The deployed endpoint is a normal SageMaker endpoint, so you can invoke it from any AWS SDK. Unlike the trained-predictor case, foundation model endpoints need the labeled context and the label column name bundled into every request — so plain CSV is not supported. Use AutoGluon-Cloud’s native application/x-autogluon envelope instead.
Reattaching to an existing endpoint¶
To send requests to an endpoint that’s already running (e.g. from a previous session, or one a
teammate deployed), build a TabularEndpoint directly from the
endpoint name:
from autogluon.cloud import TabularEndpoint
endpoint = TabularEndpoint(endpoint_name="my-existing-endpoint")
Pass a configured boto3.Session to use a non-default AWS profile or region. The endpoint must
have been deployed via AutoGluon-Cloud, since the request payload format is AutoGluon-specific.
Choosing an instance type¶
Defaults work for most users — read on if you want to optimize for cost or throughput, or the default instance type isn’t available in your account or region.
Mitra defaults to the CPU instance ml.m5.4xlarge for both batch prediction and real-time endpoints. TabICLv2, TabDPT-Turbo and Nori default to the GPU instance ml.g5.xlarge.
A few rules of thumb:
Context size drives cost. Every request is processed together with the full
train_data, so latency grows with the number of context rows and columns. Sample down the context if latency matters more than accuracy.For CPU, scaling vCPUs helps. Larger
m5/c6iinstances reduce latency at higher cost.Slow deploy? If a deploy takes much longer than ~6 min for GPU or ~4 min for CPU, the region is likely out of capacity for that instance type — try a different instance type or region.