Run Pretrained Tabular Foundation Models on Amazon SageMaker

Tabular foundation models are large pretrained models that make predictions on a new dataset in context: you pass a set of labeled rows together with the rows to predict, and the model predicts them in a single forward pass. Because they’re pretrained on large collections of diverse tables, they generalize to unseen tables out of the box — no hyperparameter tuning or lengthy training required.

That makes the workflow much simpler than training your own tabular predictor, which requires you to first fit a predictor on your data and then manage the trained artifact. With foundation models you skip the fit step entirely and go straight to running batch predictions or deploying an endpoint.

AutoGluon-Cloud exposes this workflow through TabularFoundationModel, with models like Mitra, TabICLv2, TabDPT-Turbo and Nori available out of the box. For time series forecasting, see Time Series Foundation Models.

Create the model

Important

Before running any code below, follow the Setup tutorial to register the IAM role and S3 bucket that SageMaker will use. The examples assume those resources are saved in ~/.autogluon/cloud.yaml.

from autogluon.cloud import TabularFoundationModel

model = TabularFoundationModel(model_id="mitra-classifier")

The rest of the tutorial reuses this model object.

Available models

The following model_id values are currently supported (also available via TabularFoundationModel.list_models()). Each model_id targets a single task — pick a *-classifier model for classification (binary or multiclass) and a *-regressor model for regression.

Model ID

Task

Documentation

Weights

mitra-classifier

Classification

MitraModel

autogluon/mitra-classifier

mitra-regressor

Regression

MitraModel

autogluon/mitra-regressor

mitra-v2-classifier

Classification

MitraModel

autogluon/mitra-classifier-2

tabicl-v2-classifier

Classification

TabICLModel

jingang/TabICL

tabicl-v2-regressor

Regression

TabICLModel

jingang/TabICL

tabdpt-turbo-classifier

Classification

TabDPTTurboModel

Layer6/TabDPT

tabdpt-turbo-regressor

Regression

TabDPTTurboModel

Layer6/TabDPT

nori-regressor

Regression

NoriModel

Synthefy/Nori

nori-30m-regressor

Regression

NoriModel

Synthefy/Nori-30M

Mitra runs on CPU by default; the other models default to GPU instances. For background on tabular foundation models in AutoGluon, see the Tabular Foundational Models tutorial.

Data

The examples use the Adult Income dataset, where the task is to predict whether a person earns more than $50K a year. Tabular foundation models are designed for small to medium context sizes, so we take a random sample of the training set as the labeled context:

import pandas as pd

train_data = pd.read_csv("https://autogluon.s3.amazonaws.com/datasets/Inc/train.csv").sample(2000, random_state=0)
train_data.head()
age workclass fnlwgt education education-num marital-status occupation relationship race sex capital-gain capital-loss hours-per-week native-country class
6118 51 Private 39264 Some-college 10 Married-civ-spouse Exec-managerial Wife White Female 0 0 40 United-States >50K
23204 58 Private 51662 10th 6 Married-civ-spouse Other-service Wife White Female 0 0 8 United-States <=50K
29590 40 Private 326310 Some-college 10 Married-civ-spouse Craft-repair Husband White Male 0 0 44 United-States <=50K
18116 37 Private 222450 HS-grad 9 Never-married Sales Not-in-family White Male 0 2339 40 El-Salvador <=50K
33964 62 Private 109190 Bachelors 13 Married-civ-spouse Exec-managerial Husband White Male 15024 0 40 United-States >50K

The input is a regular tabular DataFrame: one row per example, one column per feature, plus the label column — here, class. The rows to predict must contain the same feature columns. Drop the label from the test split since it’s what we want to predict:

test_data = (
    pd.read_csv("https://autogluon.s3.amazonaws.com/datasets/Inc/test.csv")
    .sample(500, random_state=0)
    .drop(columns=["class"])
)
test_data.head()
age workclass fnlwgt education education-num marital-status occupation relationship race sex capital-gain capital-loss hours-per-week native-country
5454 41 Self-emp-not-inc 408498 HS-grad 9 Married-civ-spouse Exec-managerial Husband White Male 0 0 50 United-States
6111 39 Private 746786 Bachelors 13 Married-civ-spouse Prof-specialty Husband White Male 0 0 55 United-States
5282 50 Private 62593 Assoc-voc 11 Married-civ-spouse Farming-fishing Husband Asian-Pac-Islander Male 0 0 40 United-States
3046 31 Private 248178 Some-college 10 Married-civ-spouse Other-service Husband Black Male 0 0 35 United-States
2162 43 State-gov 52849 Bachelors 13 Married-civ-spouse Prof-specialty Husband White Male 0 0 40 United-States

Inference modes

TabularFoundationModel supports two inference modes on SageMaker. In both, every call sends the labeled train_data as context along with the rows to predict, and AutoGluon fits a TabularPredictor wrapping the foundation model on that context before predicting.

  • Batch prediction — launch a one-off SageMaker job that scores a dataset and writes the results to S3. Compute spins up, runs, and shuts down automatically. Best for offline scoring of larger datasets where minutes of startup latency are fine.

  • Real-time inference — deploy the model to a long-running SageMaker endpoint and send requests over HTTPS. Lowest per-request latency, supports GPU instances. You pay for the endpoint as long as it’s up, so best when you need predictions on demand and have steady traffic.

Note

Serverless inference is not supported for tabular foundation models, since SageMaker Serverless Inference does not provide sufficient resources to run them.

The examples below all reuse the train_data and test_data DataFrames loaded above.

Batch prediction

Use predict() to score a dataset as a one-off job. It returns a Series of predictions:

predictions = model.predict(
    test_data=test_data,
    train_data=train_data,
    label="class",
)

For classification, use predict_proba() to get class probabilities. By default it returns both the predictions and the probabilities, computed in the same job:

predictions, probabilities = model.predict_proba(
    test_data=test_data,
    train_data=train_data,
    label="class",
)

The job also writes the predictions to S3 as a CSV. By default they land at {cloud_output_path}/{job_name}/predictions.csv; pass predictions_path to choose an explicit destination:

predictions = model.predict(
    test_data=test_data,
    train_data=train_data,
    label="class",
    predictions_path="s3://my-bucket/predictions/2026-06-02.csv",
)

For long-running jobs you can return immediately with wait=False. predict() then returns a JobPredictionFuture you can poll with .status() and resolve with .result():

future = model.predict(
    test_data=test_data,
    train_data=train_data,
    label="class",
    wait=False,
)

print(future.job_name, future.status())  # 'ag-...', 'InProgress'

predictions = future.result()  # blocks until the job finishes, returns a Series

Real-time inference

Deploy the model to a SageMaker endpoint with deploy(), then send requests through the returned TabularEndpoint. Pick an instance_type based on cost and latency requirements (defaults to ml.m5.4xlarge for Mitra and ml.g5.xlarge for the other models):

endpoint = model.deploy(instance_type="ml.m5.4xlarge")  # takes a few minutes

predictions = endpoint.predict(
    data=test_data,
    train_data=train_data,
    label="class",
)
predictions, probabilities = endpoint.predict_proba(
    data=test_data,
    train_data=train_data,
    label="class",
)

The endpoint holds no data between requests, so each request carries its own labeled context. This means you can use a single endpoint to serve predictions for different datasets and tasks of the same type — for example, a mitra-classifier endpoint can classify rows from any table, as long as each request includes the matching train_data and label.

Each request payload — train_data and data combined — is limited to 6 MB. For larger inputs, use batch prediction instead.

The endpoint stays active — and billed — until you delete it:

endpoint.delete_endpoint()

Invoke the endpoint without AutoGluon-Cloud

The deployed endpoint is a normal SageMaker endpoint, so you can invoke it from any AWS SDK. Unlike the trained-predictor case, foundation model endpoints need the labeled context and the label column name bundled into every request — so plain CSV is not supported. Use AutoGluon-Cloud’s native application/x-autogluon envelope instead.

Payload format — boto3 example

Each DataFrame is serialized as base64-encoded parquet, with the label column name carried in inference_kwargs. This is what autogluon.cloud.TabularEndpoint.predict() sends under the hood:

import base64
import io
import json
import boto3
import pandas as pd

def df_to_b64(df: pd.DataFrame) -> str:
    return base64.b64encode(df.to_parquet()).decode("ascii")

train_data = pd.read_csv("https://autogluon.s3.amazonaws.com/datasets/Inc/train.csv").sample(2000, random_state=0)
test_data = pd.read_csv("https://autogluon.s3.amazonaws.com/datasets/Inc/test.csv").drop(columns=["class"])

payload = {
    "version": 1,
    "data": df_to_b64(test_data),
    "train_data": df_to_b64(train_data),
    "inference_kwargs": {
        "label": "class",
    },
}

client = boto3.client("sagemaker-runtime")
response = client.invoke_endpoint(
    EndpointName=ENDPOINT_NAME,
    ContentType="application/x-autogluon",
    Accept="application/x-parquet",  # or "application/json", "text/csv"
    Body=json.dumps(payload).encode("utf-8"),
)
result = pd.read_parquet(io.BytesIO(response["Body"].read()))

For classification models, result contains the predicted label in the class column plus one <class>_proba column per class. For regression models, it contains a single column with the predictions.

Reattaching to an existing endpoint

To send requests to an endpoint that’s already running (e.g. from a previous session, or one a teammate deployed), build a TabularEndpoint directly from the endpoint name:

from autogluon.cloud import TabularEndpoint

endpoint = TabularEndpoint(endpoint_name="my-existing-endpoint")

Pass a configured boto3.Session to use a non-default AWS profile or region. The endpoint must have been deployed via AutoGluon-Cloud, since the request payload format is AutoGluon-specific.

Choosing an instance type

Defaults work for most users — read on if you want to optimize for cost or throughput, or the default instance type isn’t available in your account or region.

Mitra defaults to the CPU instance ml.m5.4xlarge for both batch prediction and real-time endpoints. TabICLv2, TabDPT-Turbo and Nori default to the GPU instance ml.g5.xlarge.

A few rules of thumb:

  • Context size drives cost. Every request is processed together with the full train_data, so latency grows with the number of context rows and columns. Sample down the context if latency matters more than accuracy.

  • For CPU, scaling vCPUs helps. Larger m5 / c6i instances reduce latency at higher cost.

  • Slow deploy? If a deploy takes much longer than ~6 min for GPU or ~4 min for CPU, the region is likely out of capacity for that instance type — try a different instance type or region.