What you will be able to do
- Train a single-series forecasting model with CREATE SNOWFLAKE.ML.FORECAST and generate predictions with its FORECAST method
- Read the forecast output columns and set the prediction interval
- Forecast many series at once with SERIES_COLNAME, and forecast with exogenous features by supplying their future values
- Choose CONFIG_OBJECT training options such as method, frequency, aggregation and on_error
Key concept
Forecast model object — In Snowflake, forecasting is a class you instantiate rather than a function you call row by row. Creating an instance of SNOWFLAKE.ML.FORECAST trains a model on your historical time series. You then call that instance's methods to predict future values of a numeric metric.
1.Forecasting is a model you train, then call
Snowflake's built-in forecasting is not a scalar function you drop into a SELECT list. It is a class, SNOWFLAKE.ML.FORECAST. You create an instance of it, which trains a model on your history, and then you call methods on that instance to get predictions. The documentation's standard example is forecasting sales by item for the next two weeks.
Before you write any SQL, three things must be in place. First, select a database, schema and virtual warehouse. Second, you must own the schema or hold the CREATE SNOWFLAKE.ML.FORECAST privilege in it. Third, you need a table or view with at least one timestamp column and one numeric column. The timestamps should be at a fixed interval with few gaps. Training can still handle real-world data with missing, duplicate or misaligned time steps. You can run the whole process from the AI & ML Studio in Snowsight or from SQL.
CREATE SNOWFLAKE.ML.FORECAST model1(
INPUT_DATA => TABLE(v1),
TIMESTAMP_COLNAME => 'date',
TARGET_COLNAME => 'sales'
);Only three arguments are required. INPUT_DATA is a table or view reference wrapped in the TABLE keyword. TIMESTAMP_COLNAME names the time column, and TARGET_COLNAME names the numeric metric to predict. When training succeeds, Snowflake replies "Instance MODEL1 successfully created." From then on, model1 is an object you can call methods on.
Checkpoint 1 of 6· Check yourself
An analyst's role can read the sales view but cannot create a forecasting model in a schema it does not own. What is missing?
To create a forecasting model, you must either own the schema or hold the forecast-specific creation privilege in it.
“Confirm that you own your schema or have CREATE SNOWFLAKE.ML.FORECAST privileges in the schema you’ve chosen.”Source: docs.snowflake.com
Sources1
2.Calling FORECAST and reading the prediction interval
You get predictions from the instance's FORECAST method, using the model!METHOD syntax. For a model trained without extra features, FORECASTING_PERIODS says how many steps ahead to predict:
call model1!FORECAST(FORECASTING_PERIODS => 3);Every forecast returns the same five columns. Dashboards and downstream queries are built on these columns:
| Column | Type | Meaning |
|---|---|---|
| SERIES | VARIANT | Series value; NULL if the model was trained on a single time series |
| TS | TIMESTAMP_NTZ | Timestamp of the predicted point |
| FORECAST | FLOAT | Forecast target value |
| LOWER_BOUND | FLOAT | Lower boundary of the prediction interval |
| UPPER_BOUND | FLOAT | Upper boundary of the prediction interval |
LOWER_BOUND and UPPER_BOUND express the model's uncertainty. The default prediction_interval is 0.95, which means 95% of future points are expected to fall between the two bounds. To use a different level, pass prediction_interval in a CONFIG_OBJECT. In the documentation's toy example, the series is perfectly linear and the model makes zero errors, so both bounds equal the forecast.
CALL is convenient for interactive use. To filter, join or save the output, put the method call inside TABLE( ) in a FROM clause and leave out the CALL keyword:
CREATE TABLE my_forecasts AS
SELECT * FROM TABLE(model1!FORECAST(FORECASTING_PERIODS => 3));Checkpoint 2 of 6· Fill the gap
Which configuration key widens or narrows the band between LOWER_BOUND and UPPER_BOUND?
CALL model1!FORECAST(FORECASTING_PERIODS => 3, CONFIG_OBJECT => {' ? ': 0.8});prediction_interval sets the share of future points expected to fall inside the bounds. Its default is 0.95, and here it is set to 0.8.
Source: docs.snowflake.comCheckpoint 3 of 6· Exam question
An analyst owns the table `DAILY_SALES(sale_date DATE, revenue NUMBER)` with three years of one continuous daily series. The business wants a 30-day revenue forecast from Snowflake's built-in ML function. Which approach produces it?
Correct answer: D — Run CREATE SNOWFLAKE.ML.FORECAST sales_model(INPUT_DATA => TABLE(daily_sales), TIMESTAMP_COLNAME => 'sale_date', TARGET_COLNAME => 'revenue'), then CALL sales_model!FORECAST(FORECASTING_PERIODS => 30).
- A. Incorrect. No such one-shot procedure exists; training and prediction are separate steps, and the trained model persists as an object in the schema.
- B. Incorrect. Methods on a model object use the bang syntax with a CALL statement and named arguments such as FORECASTING_PERIODS, not a dotted TABLE() reference.
- C. Incorrect. Forecasting is not a window function; it is a model class that must be created as a schema object and then invoked through its methods.
- D. Correct. Forecasting is a two-step flow: train a model object from the table with timestamp and target columns, then call its FORECAST method with the number of periods to predict.
3.Many series at once, and forecasting with features
Real forecasts rarely involve a single series. To forecast sales separately for every store/item combination, build one column that identifies each series and pass it as SERIES_COLNAME. The column can hold a value of any type, or an array built from several columns. The documentation's example uses [store_id, item] AS store_item. A single CREATE statement then trains a model for every series:
CREATE SNOWFLAKE.ML.FORECAST model2(
INPUT_DATA => TABLE(v3),
SERIES_COLNAME => 'store_item',
TIMESTAMP_COLNAME => 'date',
TARGET_COLNAME => 'sales'
);Calling model2!FORECAST with only FORECASTING_PERIODS predicts every series, and the SERIES column shows which rows belong to which series. If you need just one series, name it with SERIES_VALUE. Only that series' forecast is generated, which is cheaper than forecasting everything and filtering afterwards:
CALL model2!FORECAST(SERIES_VALUE => [2,'umbrella'], FORECASTING_PERIODS => 2);Checkpoint 4 of 6· Check yourself
A dashboard needs forecasts only for store 2's umbrellas, using model2, which was trained on every store/item series. What is the most efficient call?
SERIES_VALUE generates the forecast for only the named series, so no work is spent on the series you would filter out.
“Specifying one series with the FORECAST method is more efficient than filtering the results of a multi-series forecast”Source: docs.snowflake.com
Features are additional factors you believe influence the metric, such as temperature, humidity or holidays. These are also called exogenous variables. No special argument is needed to declare them: any column in the training data other than the timestamp and the target is treated as a feature. The trade-off comes at prediction time. Because the model learned how sales respond to temperature, it needs future temperatures to make a prediction. You therefore supply a view of future feature values as INPUT_DATA, and FORECASTING_PERIODS is no longer used. The forecast timestamps come from that view:
CALL model3!FORECAST(
INPUT_DATA => TABLE(v2_forecast),
TIMESTAMP_COLNAME =>'date'
);Checkpoint 5 of 6· Put it in order
Put the steps for forecasting sales with weather and holiday features in order
- 1.Call FORECAST with INPUT_DATA pointing at the future-features view
- 2.Build a view holding future dates and future values for each feature
- 3.Train the model with CREATE SNOWFLAKE.ML.FORECAST, naming only the timestamp and target columns
- 4.Build a training view with date, sales and the feature columns
The features must be in the training data before the model can learn from them. The future values are needed only when you generate the forecast, and they replace FORECASTING_PERIODS.
“To generate forecasts with this model, you must provide future values for the features to the model”Source: docs.snowflake.com
Sources1
4.Tuning training with CONFIG_OBJECT
CREATE SNOWFLAKE.ML.FORECAST also accepts a CONFIG_OBJECT. Its keys control how the model is trained. These are the keys that most often decide whether a forecasting job fits the business need:
| Key | Default | What it controls |
|---|---|---|
| method | 'best' | 'best' uses an ensemble (Prophet, ARIMA, Exponential Smoothing and a GBM-based algorithm). 'fast' uses a single GBM-based algorithm and is recommended for 10,000 or more series |
| frequency | n/a (inferred) | The time series frequency as a string such as '1 day'; units range from seconds to years |
| aggregation_target | Same as aggregation_numeric, or 'MEAN' | How target values are combined, for example 'MEAN', 'MEDIAN', 'SUM', 'MIN', 'MAX', 'FIRST', 'LAST' |
| evaluate | TRUE | Whether additional cross-validation models are trained to produce evaluation metrics |
| on_error | 'ABORT' | 'skip' lets training continue for the other series when one series fails |
| lower_bound / upper_bound | NULL | Thresholds the forecast will never go below or above |
Several of these keys trade accuracy against cost. With 'best', the default, several algorithms compete. 'fast' trains a single GBM, which is quicker but may be less accurate. Setting evaluate to FALSE skips the extra cross-validation training. frequency and the aggregation_* keys work together: you declare the grain you want, for example '1 day', and choose how values are combined at that grain. With on_error set to 'skip', you can find the series that failed with the model's SHOW_TRAINING_LOGS method.
For warehouse sizing, the documentation names two key factors: the number of rows and columns, and the number of distinct series. Finally, a trained model never changes. When new data arrives, you train a new model. Unused models still incur storage cost until you drop them.
Checkpoint 6 of 6· Match them up
Match each requirement to the training option that addresses it
Tap a term, then the definition that fits it.
Each option targets a different constraint: series count, time grain, failure isolation or evaluation cost.
“We recommend using ‘fast’ when your training data has 10,000 or more individual series.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.To query forecast output, write SELECT * FROM TABLE(CALL model!FORECAST(...)).Why is that wrong?
Inside a FROM clause you leave out CALL and wrap the method call in TABLE( ): SELECT * FROM TABLE(model1!FORECAST(FORECASTING_PERIODS => 3)).
Covered in Calling FORECAST and reading the prediction interval
2.A model trained with exogenous features can be forecast with FORECASTING_PERIODS alone, like a plain model.Why is that wrong?
A feature model needs future values of each feature. You pass them as INPUT_DATA, and the forecast timestamps come from that data rather than from a period count.
Covered in Many series at once, and forecasting with features
3.By default, the forecast model trains a single fast GBM algorithm.Why is that wrong?
The default method is 'best', an ensemble that picks the best algorithm for the data. 'fast' is the single-GBM option you have to request.
Covered in Tuning training with CONFIG_OBJECT
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Have a table or view with at least two columns: one timestamp column and one numeric column.”
↩︎ Forecasting is a model you train, then call“Instead, put the call in parentheses, preceded by the TABLE keyword.”
↩︎ Calling FORECAST and reading the prediction interval“To create a forecasting model for multiple series at once, use the series_colname parameter.”
↩︎ Many series at once, and forecasting with features“additional columns in the input data are assumed to be features for use in training.”
↩︎ Many series at once, and forecasting with features“Models are immutable and cannot be updated in place. Train a new model instead.”
↩︎ Tuning training with CONFIG_OBJECT“Forecasting uses a machine learning algorithm that predicts future numeric data from historical time series data.”
↩︎ Key concept“As shown in the example above, when calling the method, omit the CALL command.”
↩︎ Exam trap 1“In this variation of the FORECAST method, you do not specify the number of timestamps to predict.”
↩︎ Exam trap 2“Note that the model has inferred the interval between timestamps from the training data.”
↩︎ Prediction“Confirm that you own your schema or have CREATE SNOWFLAKE.ML.FORECAST privileges in the schema you’ve chosen.”
↩︎ Checkpoint“Specifying one series with the FORECAST method is more efficient than filtering the results of a multi-series forecast”
↩︎ Checkpoint“To generate forecasts with this model, you must provide future values for the features to the model”
↩︎ Checkpoint - 2.
“The default value of 0.95 means 95% of future points are expected to fall within the interval [lower_bound, upper_bound] from the forecast result.”
↩︎ Calling FORECAST and reading the prediction interval - 3.
“'fast': Uses a single algorithm - a GBM based algorithm - to train the model.”
↩︎ Tuning training with CONFIG_OBJECT“'best': Uses an ensemble of models to determine the best algorithm for the data.”
↩︎ Exam trap 3“We recommend using ‘fast’ when your training data has 10,000 or more individual series.”
↩︎ Checkpoint