Forecasting Time Series - Evaluation Metrics¶
Picking the right evaluation metric is one of the most important choices when using an AutoML framework. This page lists the forecast evaluation metrics available in AutoGluon, explains when different metrics should be used, and describes how to define custom evaluation metrics.
When using AutoGluon, you can specify the metric using the eval_metric argument to TimeSeriesPredictor, for example:
from autogluon.timeseries import TimeSeriesPredictor
predictor = TimeSeriesPredictor(eval_metric="MASE")
AutoGluon will use the provided metric to tune model hyperparameters, rank models, and construct the final ensemble for prediction.
Note
AutoGluon always reports all metrics in a higher-is-better format.
For this purpose, some metrics are multiplied by -1.
For example, if we set eval_metric="MASE", the predictor will actually report -MASE (i.e., MASE score multiplied by -1). This means the test_score will be between 0 (most accurate forecast) and \(-\infty\) (least accurate forecast).
Currently, AutoGluon supports following evaluation metrics:
Mean quantile loss. |
|
Scaled quantile loss. |
|
Weighted quantile loss. |
|
Mean absolute error. |
|
Mean absolute error with a bias penalty. |
|
Mean absolute percentage error. |
|
Mean absolute scaled error. |
|
Mean squared error. |
|
Root mean squared error. |
|
Root mean squared logarithmic error. |
|
Root mean squared scaled error. |
|
Symmetric mean absolute percentage error. |
|
Weighted absolute percentage error. |
|
Weighted absolute percentage error with a bias penalty. |
Alternatively, you can define a custom forecast evaluation metric.
Which evaluation metric to choose?¶
If you are not sure which evaluation metric to pick, here are three questions that can help you make the right choice for your use case.
1. Are you interested in a point forecast or a probabilistic forecast?
If your goal is to generate an accurate probabilistic forecast, you should use WQL, MQL or SQL metrics.
These metrics are based on the quantile loss and measure the accuracy of the quantile forecasts.
By default, AutoGluon predicts quantile levels [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9].
To predict a different set of quantiles, you can use quantile_levels argument:
predictor = TimeSeriesPredictor(eval_metric="WQL", quantile_levels=[0.1, 0.5, 0.75, 0.9])
All remaining forecast metrics described on this page are point forecast metrics.
Note that if you select the eval_metric to a point forecast metric when creating the TimeSeriesPredictor, then the forecast minimizing this metric will always be provided in the "mean" column of the predictions dataframe.
2. Do you care more about accurately predicting time series with large values?
If the answer is “yes” (for example, if it’s important to more accurately predict sales of popular products), you should use scale-dependent metrics like WQL, MQL, MAE, RMSE, or WAPE.
These metrics are also well-suited for dealing with sparse (intermittent) time series that have lots of zeros.
If you additionally want to penalize forecasts that are systematically biased (consistently over- or under-forecasting demand), consider the bias-penalized metrics MAEB or WAPEB.
To measure (rather than penalize) the direction and magnitude of the forecast bias, you can evaluate a trained predictor with the BIAS metric via predictor.evaluate(..., metrics="BIAS"). Since a low BIAS value can be achieved by a poor forecast (both over- and under-forecasting are penalized only in absolute terms), BIAS cannot be used as an eval_metric for model selection.
If the answer is “no” (you care equally about all time series in the dataset), consider scaled metrics like SQL, MASE and RMSSE. Alternatively, percentage-based metrics MAPE and SMAPE can also be used to equalize the scale across time series. However, these percentage-based metrics have some well-documented limitations, so we don’t recommend using them in practice.
Note that both scaled and percentage-based metrics are poorly suited for sparse (intermittent) data.
3. (Point forecast only) Do you want to estimate the mean or the median?
To estimate the median, you need to use metrics such as MAE, MASE or WAPE.
If your goal is to predict the mean (expected value), you should use MSE, RMSE or RMSSE metrics.
When dealing with intermittent or sparse time series, you can use the bias-penalized metrics MAEB and WAPEB. These balance between a median-seeking accuracy term and a mean-seeking bias penalty, which discourages the chronic under-forecasting that minimizing MAE or WAPE alone tends to reward.
Point forecast metrics¶
We use the following notation in mathematical definitions of point forecast metrics:
\(y_{i,t}\) - observed value of time series \(i\) at time \(t\)
\(f_{i,t}\) - predicted value of time series \(i\) at time \(t\)
\(N\) - number of time series (number of items) in the dataset
\(T\) - length of the observed time series
\(H\) - length of the forecast horizon (
prediction_length)
- class autogluon.timeseries.metrics.MAE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean absolute error.
(1)¶\[\operatorname{MAE} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N}\sum_{t=T+1}^{T+H} |y_{i,t} - f_{i,t}|\]Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
not sensitive to outliers
prefers models that accurately estimate the median
References
- class autogluon.timeseries.metrics.MAEB(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean absolute error with a bias penalty.
Adds a penalty for systematically biased (over- or under-) forecasts to
MAE. Defined as the mean absolute error plus the absolute mean forecast bias.(2)¶\[\operatorname{MAEB} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} |y_{i,t} - f_{i,t}| + \left| \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} (f_{i,t} - y_{i,t}) \right|\]The first term measures forecast accuracy (as in MAE), while the second term penalizes forecast bias (the mean net over- or under-forecast). This discourages models that achieve low error by systematically under-forecasting demand, which is a common failure mode when optimizing MAE on intermittent (sparse) time series.
Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
equivalent to
MAEwhen the forecast is unbiased (mean net error is zero)
See
WAPEBfor a scale-free version of this metric that normalizes the error across time series.
- class autogluon.timeseries.metrics.MAPE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean absolute percentage error.
(3)¶\[\operatorname{MAPE} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} \frac{ |y_{i,t} - f_{i,t}|}{|y_{i,t}|}\]Properties:
should only be used if all time series have positive values
undefined for time series that contain zero values
penalizes overprediction more heavily than underprediction
References
- class autogluon.timeseries.metrics.MASE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean absolute scaled error.
Normalizes the absolute error for each time series by the historical seasonal error of this time series.
(4)¶\[\operatorname{MASE} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \frac{1}{a_i} \sum_{t=T+1}^{T+H} |y_{i,t} - f_{i,t}|\]where \(a_i\) is the historical absolute seasonal error defined as
(5)¶\[a_i = \frac{1}{T-m} \sum_{t=m+1}^T |y_{i,t} - y_{i,t-m}|\]and \(m\) is the seasonal period of the time series (
eval_metric_seasonal_period).Properties:
scaled metric (normalizes the error for each time series by the scale of that time series)
undefined for constant time series
not sensitive to outliers
prefers models that accurately estimate the median
References
- class autogluon.timeseries.metrics.MSE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean squared error.
Using this metric will lead to forecast of the mean.
(6)¶\[\operatorname{MSE} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N}\sum_{t=T+1}^{T+H} (y_{i,t} - f_{i,t})^2\]Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
heavily penalizes models that cannot quickly adapt to abrupt changes in the time series
sensitive to outliers
prefers models that accurately estimate the mean (expected value)
References
- class autogluon.timeseries.metrics.RMSE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Root mean squared error.
(7)¶\[\operatorname{RMSE} = \sqrt{\frac{1}{N} \frac{1}{H} \sum_{i=1}^{N}\sum_{t=T+1}^{T+H} (y_{i,t} - f_{i,t})^2}\]Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
heavily penalizes models that cannot quickly adapt to abrupt changes in the time series
sensitive to outliers
prefers models that accurately estimate the mean (expected value)
References
- class autogluon.timeseries.metrics.RMSLE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Root mean squared logarithmic error.
Applies a logarithmic transformation to the predictions before computing the root mean squared error. Assumes both the ground truth and predictions are positive. If negative predictions are given, they will be clipped to zero.
(8)¶\[\operatorname{RMSLE} = \sqrt{\frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} (\ln(1 + y_{i,t}) - \ln(1 + f_{i,t}))^2}\]Properties:
undefined for time series with negative values
penalizes models that underpredict more than models that overpredict
insensitive to effects of outliers and scale, best when targets can vary or trend exponentially
References
- class autogluon.timeseries.metrics.RMSSE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Root mean squared scaled error.
Normalizes the absolute error for each time series by the historical seasonal error of this time series.
(9)¶\[\operatorname{RMSSE} = \sqrt{\frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \frac{1}{s_i} \sum_{t=T+1}^{T+H} (y_{i,t} - f_{i,t})^2}\]where \(s_i\) is the historical squared seasonal error defined as
(10)¶\[s_i = \frac{1}{T-m} \sum_{t=m+1}^T (y_{i,t} - y_{i,t-m})^2\]and \(m\) is the seasonal period of the time series (
eval_metric_seasonal_period).Properties:
scaled metric (normalizes the error for each time series by the scale of that time series)
undefined for constant time series
heavily penalizes models that cannot quickly adapt to abrupt changes in the time series
sensitive to outliers
prefers models that accurately estimate the mean (expected value)
References
- class autogluon.timeseries.metrics.SMAPE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Symmetric mean absolute percentage error.
(11)¶\[\operatorname{SMAPE} = 2 \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} \frac{ |y_{i,t} - f_{i,t}|}{|y_{i,t}| + |f_{i,t}|}\]Properties:
should only be used if all time series have positive values
poorly suited for sparse & intermittent time series that contain zero values
penalizes overprediction more heavily than underprediction
References
- class autogluon.timeseries.metrics.WAPE(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Weighted absolute percentage error.
Defined as sum of absolute errors divided by the sum of absolute time series values in the forecast horizon.
(12)¶\[\operatorname{WAPE} = \frac{1}{\sum_{i=1}^{N} \sum_{t=T+1}^{T+H} |y_{i, t}|} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} |y_{i,t} - f_{i,t}|\]Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
not sensitive to outliers
prefers models that accurately estimate the median
If
self.horizon_weightis provided, both the errors and the target time series in the denominator will be re-weighted.References
- class autogluon.timeseries.metrics.WAPEB(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Weighted absolute percentage error with a bias penalty.
Adds a penalty for systematically biased (over- or under-) forecasts to
WAPE. Defined as the sum of absolute errors plus the absolute total forecast bias, divided by the sum of absolute time series values in the forecast horizon.(13)¶\[\operatorname{WAPEB} = \frac{\sum_{i=1}^{N} \sum_{t=T+1}^{T+H} |y_{i,t} - f_{i,t}| + \left| \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} (f_{i,t} - y_{i,t}) \right|}{\sum_{i=1}^{N} \sum_{t=T+1}^{T+H} |y_{i, t}|}\]The numerator combines forecast accuracy (absolute error, as in WAPE) with forecast bias (the absolute net over- or under-forecast). This discourages models that achieve low error by systematically under-forecasting demand, which is a common failure mode for intermittent (sparse) time series. WAPEB is the scale-free counterpart of
MAEB, analogous to how WAPE relates to MAE.This metric was used to rank submissions in the VN1 Forecasting Accuracy Challenge (for non-negative target values, \(\sum |y_{i,t}|\) equals the total actual sales used to normalize the score in the competition).
Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
equivalent to
WAPEwhen the forecast is unbiased (net error is zero)well-suited for sparse (intermittent) time series that contain many zeros
If
self.horizon_weightis provided, the errors, the biases, and the target time series in the denominator will all be re-weighted.References
- class autogluon.timeseries.metrics.BIAS(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean forecast bias (signed).
Measures the average net over- or under-forecast. Positive values indicate that the forecast systematically over-predicts the target, negative values indicate systematic under-prediction.
(14)¶\[\operatorname{BIAS} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} (f_{i,t} - y_{i,t})\]Forecast bias is a standard diagnostic in demand and retail forecasting, where the direction of the error matters: over-forecasting ties up inventory and capital, while under-forecasting causes stockouts.
Warning
Unlike other metrics, BIAS is not a loss function: its optimum is 0, and both positive and negative values indicate a worse forecast. For this reason it can only be used to evaluate a trained predictor via
evaluate()and cannot be passed aseval_metricfor model selection (evaluate_only=True). To penalize forecast bias during model selection, useMAEBorWAPEBinstead, which combine forecast accuracy with an absolute bias penalty.Properties:
scale-dependent (time series with large absolute value contribute more to the metric)
Probabilistic forecast metrics¶
In addition to the notation listed above, we use following notation to define probabilistic forecast metrics:
\(f_{i,t}^q\) - predicted quantile \(q\) of time series \(i\) at time \(t\)
\(\rho_q(y, f) \) - quantile loss at level \(q\) defined as
- class autogluon.timeseries.metrics.MQL(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Mean quantile loss.
Also known as mean pinball loss.
Defined as the quantile loss averaged over all time series, time steps in the forecast horizon, and quantile levels.
(15)¶\[\operatorname{MQL} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} \frac{1}{|\mathcal{Q}|} \sum_{q \in \mathcal{Q}} \rho_q(y_{i,t}, f^q_{i,t})\]where \(\mathcal{Q}\) is the set of quantile levels.
Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
equivalent to MAE if
quantile_levels = [0.5]
References
- class autogluon.timeseries.metrics.SQL(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Scaled quantile loss.
Also known as scaled pinball loss.
Normalizes the quantile loss for each time series by the historical seasonal error of this time series.
(16)¶\[\operatorname{SQL} = \frac{1}{N} \frac{1}{H} \sum_{i=1}^{N} \frac{1}{a_i} \sum_{t=T+1}^{T+H} \sum_{q} \rho_q(y_{i,t}, f^q_{i,t})\]where \(a_i\) is the historical absolute seasonal error defined as
(17)¶\[a_i = \frac{1}{T-m} \sum_{t=m+1}^T |y_{i,t} - y_{i,t-m}|\]and \(m\) is the seasonal period of the time series (
eval_metric_seasonal_period).Properties:
scaled metric (normalizes the error for each time series by the scale of that time series)
undefined for constant time series
equivalent to MASE if
quantile_levels = [0.5]
References
- class autogluon.timeseries.metrics.WQL(prediction_length: int = 1, seasonal_period: int | None = None, horizon_weight: Sequence[float] | None = None)[source]¶
Weighted quantile loss.
Also known as weighted pinball loss.
Defined as total quantile loss divided by the sum of absolute time series values in the forecast horizon.
(18)¶\[\operatorname{WQL} = \frac{1}{\sum_{i=1}^{N} \sum_{t=T+1}^{T+H} |y_{i, t}|} \sum_{i=1}^{N} \sum_{t=T+1}^{T+H} \sum_{q} \rho_q(y_{i,t}, f^q_{i,t})\]Properties:
scale-dependent (time series with large absolute value contribute more to the loss)
equivalent to WAPE if
quantile_levels = [0.5]
If
horizon_weightis provided, both the errors and the target time series in the denominator will be re-weighted.References
Custom forecast metrics¶
If none of the built-in metrics meet your requirements, you can provide a custom evaluation metric to AutoGluon.
To define a custom metric, you need to create a class that inherits from TimeSeriesScorer and implements the compute_metric method according to the following API specification:
- TimeSeriesScorer.compute_metric(data_future: TimeSeriesDataFrame, predictions: TimeSeriesDataFrame, target: str = 'target', **kwargs) float[source]¶
Internal method that computes the metric for given forecast & actual data.
This method should be implemented by all custom metrics.
- Parameters:
data_future (TimeSeriesDataFrame) – Actual values of the time series during the forecast horizon (
prediction_lengthvalues for each time series in the dataset). Must have the same index aspredictions.predictions (TimeSeriesDataFrame) – Data frame with predictions for the forecast horizon. Contain columns “mean” (point forecast) and the columns corresponding to each of the quantile levels. Must have the same index as
data_future.target (str, default = "target") – Name of the column in
data_futurethat contains the target time series.
- Returns:
score – Value of the metric for given forecast and data. If self.greater_is_better_internal is True, returns score in greater-is-better format, otherwise in lower-is-better format.
- Return type:
float
Custom mean squared error metric¶
Here is an example of how you can define a custom mean squared error (MSE) metric using TimeSeriesScorer.
import sklearn.metrics
from autogluon.timeseries.metrics import TimeSeriesScorer
class MeanSquaredError(TimeSeriesScorer):
greater_is_better_internal = False
optimum = 0.0
def compute_metric(self, data_future, predictions, target, **kwargs):
return sklearn.metrics.mean_squared_error(y_true=data_future[target], y_pred=predictions["mean"])
The internal method compute_metric returns metric in lower-is-better format, so we need to set greater_is_better_internal=False.
This will tell AutoGluon that the metric value must be multiplied by -1 to convert it to greater-is-better format.
Note
Custom metrics must be defined in a separate Python file and imported so that they can be pickled (Python’s serialization protocol). If a custom metric is not picklable, AutoGluon may crash during fit if you enable hyperparameter tuning. In the above example, you would want to create a new python file such as my_metrics.py with class MeanSquaredError defined in it, and then use it via from my_metrics import MeanSquaredError.
We can use the custom metric to measure accuracy of a forecast generated by the predictor.
import pandas as pd
from autogluon.timeseries import TimeSeriesPredictor, TimeSeriesDataFrame
# Create dummy dataset
data = TimeSeriesDataFrame.from_iterable_dataset(
[
{"start": pd.Period("2023-01-01", freq="D"), "target": list(range(15))},
{"start": pd.Period("2023-01-01", freq="D"), "target": list(range(30, 45))},
]
)
prediction_length = 3
train_data, test_data = data.train_test_split(prediction_length=prediction_length)
predictor = TimeSeriesPredictor(prediction_length=prediction_length, verbosity=0).fit(train_data, hyperparameters={"Naive": {}})
predictions = predictor.predict(train_data)
mse = MeanSquaredError(prediction_length=predictor.prediction_length)
mse_score = mse(
data=test_data,
predictions=predictions,
target=predictor.target,
)
print(f"{mse.name_with_sign} = {mse_score}")
-MeanSquaredError = -4.666666666666667
Note that the metric value has been multiplied by -1 because we set greater_is_better_internal=False when defining the metric.
When we call the metric, TimeSeriesScorer takes care of splitting test_data into past & future parts, validating that predictions have correct timestamps, and ensuring that the score is reported in greater-is-better format.
During the metric call, the method compute_metric that we implemented receives as input the following arguments:
Test data corresponding to the forecast horizon
data_future = test_data.slice_by_timestep(-prediction_length, None)
data_future
| target | ||
|---|---|---|
| item_id | timestamp | |
| 0 | 2023-01-13 | 12 |
| 2023-01-14 | 13 | |
| 2023-01-15 | 14 | |
| 1 | 2023-01-13 | 42 |
| 2023-01-14 | 43 | |
| 2023-01-15 | 44 |
Predictions for the forecast horizon
predictions.round(2)
| mean | 0.1 | 0.2 | 0.3 | 0.4 | 0.5 | 0.6 | 0.7 | 0.8 | 0.9 | ||
|---|---|---|---|---|---|---|---|---|---|---|---|
| item_id | timestamp | ||||||||||
| 0 | 2023-01-13 | 11.0 | 9.72 | 10.16 | 10.48 | 10.75 | 11.0 | 11.25 | 11.52 | 11.84 | 12.28 |
| 2023-01-14 | 11.0 | 9.19 | 9.81 | 10.26 | 10.64 | 11.0 | 11.36 | 11.74 | 12.19 | 12.81 | |
| 2023-01-15 | 11.0 | 8.78 | 9.54 | 10.09 | 10.56 | 11.0 | 11.44 | 11.91 | 12.46 | 13.22 | |
| 1 | 2023-01-13 | 41.0 | 39.72 | 40.16 | 40.48 | 40.75 | 41.0 | 41.25 | 41.52 | 41.84 | 42.28 |
| 2023-01-14 | 41.0 | 39.19 | 39.81 | 40.26 | 40.64 | 41.0 | 41.36 | 41.74 | 42.19 | 42.81 | |
| 2023-01-15 | 41.0 | 38.78 | 39.54 | 40.09 | 40.56 | 41.0 | 41.44 | 41.91 | 42.46 | 43.22 |
Note that both data_future and predictions cover the same time range.
Custom quantile loss metric¶
The metric can be computed on any columns of the predictions DataFrame.
For example, here is how we can define the mean quantile loss metric that measures the accuracy of the quantile forecast.
class MeanQuantileLoss(TimeSeriesScorer):
needs_quantile = True
greater_is_better_internal = False
optimum = 0.0
def compute_metric(self, data_future, predictions, target, **kwargs):
quantile_columns = [col for col in predictions if col != "mean"]
total_quantile_loss = 0.0
for q in quantile_columns:
total_quantile_loss += sklearn.metrics.mean_pinball_loss(y_true=data_future[target], y_pred=predictions[q], alpha=float(q))
return total_quantile_loss / len(quantile_columns)
Here we set needs_quantile=True to tell AutoGluon that this metric is evaluated on the quantile forecasts.
In this case, models such as DirectTabularModel will train a regression model from autogluon.tabular with problem_type="quantile" under the hood.
If needs_quantile=False, these models will use problem_type="regression" instead.
Custom mean absolute scaled error metric¶
Finally, here is how we can define the mean absolute scaled error (MASE) metric. Unlike previously discussed metrics, MASE is computed using both past and future time series values. The past values are used to compute the scale by which we normalize the error during the forecast horizon.
class MeanAbsoluteScaledError(TimeSeriesScorer):
greater_is_better_internal = False
optimum = 0.0
optimized_by_median = True
equivalent_tabular_regression_metric = "mean_absolute_error"
def save_past_metrics(
self, data_past: TimeSeriesDataFrame, target: str = "target", seasonal_period: int = 1, **kwargs
) -> None:
seasonal_diffs = data_past[target].groupby(level="item_id").diff(seasonal_period).abs()
self._abs_seasonal_error_per_item = seasonal_diffs.groupby(level="item_id").mean().fillna(1.0)
def clear_past_metrics(self):
self._abs_seasonal_error_per_item = None
def compute_metric(
self, data_future: TimeSeriesDataFrame, predictions: TimeSeriesDataFrame, target: str = "target", **kwargs
) -> float:
mae_per_item = (data_future[target] - predictions["mean"]).abs().groupby(level="item_id").mean()
return (mae_per_item / self._abs_seasonal_error_per_item).mean()
We compute the metrics on past data using save_past_metrics method.
Doing this in a separate method allows AutoGluon to avoid redundant computations when fitting the weighted ensemble, which requires thousands of metric evaluations.
Because we set optimized_by_median=True, AutoGluon will automatically paste the median forecast into the "mean" column of predictions.
This is done for consistency: if TimeSeriesPredictor is trained with a point forecast metric, the optimal point forecast will always be stored in the "mean" column.
Finally, the equivalent_tabular_regression_metric is used by forecasting models that fit tabular regression models from autogluon.tabular under the hood.
Using custom metrics in TimeSeriesPredictor¶
Now that we have created several custom metrics, let’s use them for training and evaluating models.
predictor = TimeSeriesPredictor(eval_metric=MeanQuantileLoss()).fit(train_data, hyperparameters={"Naive": {}, "SeasonalNaive": {}, "Theta": {}})
Beginning AutoGluon training...
AutoGluon will save models to '/home/ci/autogluon/docs/tutorials/timeseries/AutogluonModels/ag-20260811_120820'
=================== System Info ===================
AutoGluon Version: 1.6.2.dev0
Python Version: 3.13.11
Operating System: Linux
Platform Machine: x86_64
Platform Version: #1 SMP Thu Jun 25 14:43:50 UTC 2026
CPU Count: 8
Pytorch Version: 2.13.0+cu130
CUDA Version: 13.0
GPU Memory: GPU 0: 14.57/14.57 GB
Total GPU Memory: Free: 14.57 GB, Allocated: 0.00 GB, Total: 14.57 GB
GPU Count: 1
Memory Avail: 27.81 GB / 30.94 GB (89.9%)
Disk Space Avail: 213.69 GB / 255.99 GB (83.5%)
===================================================
Fitting with arguments:
{'enable_ensemble': True,
'eval_metric': MeanQuantileLoss,
'hyperparameters': {'Naive': {}, 'SeasonalNaive': {}, 'Theta': {}},
'known_covariates_names': [],
'num_val_windows': 1,
'prediction_length': 1,
'quantile_levels': [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9],
'random_seed': 123,
'refit_every_n_windows': 1,
'refit_full': False,
'skip_model_selection': False,
'target': 'target',
'verbosity': 2}
Inferred time series frequency: 'D'
Provided train_data has 24 rows, 2 time series. Median time series length is 12 (min=12, max=12).
Provided data contains following columns:
target: 'target'
AutoGluon will gauge predictive performance using evaluation metric: 'MeanQuantileLoss'
This metric's sign has been flipped to adhere to being higher_is_better. The metric score can be multiplied by -1 to get the metric value.
===================================================
Starting training. Start time is 2026-08-11 12:08:20
Models that will be trained: ['Naive', 'SeasonalNaive', 'Theta']
Training timeseries model Naive.
-0.3323 = Validation score (-MeanQuantileLoss)
0.02 s = Training runtime
0.01 s = Validation (prediction) runtime
Training timeseries model SeasonalNaive.
-2.3263 = Validation score (-MeanQuantileLoss)
0.02 s = Training runtime
0.01 s = Validation (prediction) runtime
Training timeseries model Theta.
-0.2525 = Validation score (-MeanQuantileLoss)
0.02 s = Training runtime
0.40 s = Validation (prediction) runtime
Fitting 1 ensemble(s), in 1 layers.
Training ensemble model WeightedEnsemble.
Ensemble weights: {'Theta': 1.0}
-0.2525 = Validation score (-MeanQuantileLoss)
2.40 s = Training runtime
0.40 s = Validation (prediction) runtime
Training complete. Models trained: ['Naive', 'SeasonalNaive', 'Theta', 'WeightedEnsemble']
Total runtime: 2.91 s
Best model: Theta
Best model score: -0.2525
We can also evaluate a trained predictor using these custom metrics
predictor.evaluate(test_data, metrics=[MeanAbsoluteScaledError(), MeanQuantileLoss(), MeanSquaredError()])
Model not specified in predict, will default to the model with the best validation score: Theta
{'MeanAbsoluteScaledError': np.float64(-0.07215007215007217),
'MeanQuantileLoss': -0.2525248845418294,
'MeanSquaredError': -0.2550760126517704}
That’s all it takes to create and use custom forecasting metrics in AutoGluon!
You can have a look at the AutoGluon source code for example implementations of point and quantile forecasting metrics.
If you create a custom metric, consider submitting a PR so that we can officially add it to AutoGluon.
For more tutorials, refer to Forecasting Time Series - Quick Start and Forecasting Time Series - In Depth.
Customizing the training loss for individual models¶
While eval_metric is used for model selection and weighted ensemble construction, it usually has no effect on the training loss of the individual forecasting models.
In some models such as AutoETS or AutoARIMA, the training loss is fixed and cannot be changed.
In contrast, for GluonTS-based deep learning models the training loss can be changed by modifying the distr_output hyperparameter.
By default, most GluonTS models set the distr_output to the heavy‑tailed StudentTOutput distribution for increased robustness to outliers.
You can replace the default StudentTOutput with any built‑in Output from the gluonts.torch.distributions module.
For example, here we train two versions of PatchTST with different outputs and losses:
NormalOutput- the model outputs parameters of a Gaussian distribution and trains with the negative log-likelihood loss.QuantileOutput- the model outputs a quantile forecast and trains with the quantile loss.
from autogluon.timeseries import TimeSeriesPredictor
from gluonts.torch.distributions import NormalOutput, QuantileOutput
predictor = TimeSeriesPredictor(...)
predictor.fit(
train_data,
hyperparameters={
"PatchTST": [
{"distr_output": NormalOutput()},
{"distr_output": QuantileOutput(quantiles=predictor.quantile_levels)},
]
}
)
You can define a custom loss function for the GluonTS models by defining a subclass of gluonts.torch.distributions.Output
and providing it as a distr_output to the model.
from gluonts.torch.distributions import Output
class MyCustomOutput(Output):
# implement methods of gluonts.torch.distributions.Output
...
predictor.fit(train_data, hyperparameters={"PatchTST": {"distr_output": MyCustomOutput()}})
You can find examples of Output implementations in the GluonTS code base (e.g., QuantileOutput or NormalOutput).