7.3 Running Inference and Organizing Models in Model Registry

Key Takeaways

  • ML.PREDICT scores rows in batch with SQL; ML.FORECAST, ML.DETECT_ANOMALIES, ML.RECOMMEND, and ML.EXPLAIN_PREDICT cover forecasting, anomalies, recommendations, and explanations.

  • Wrapping ML.PREDICT in a scheduled query or Dataform action refreshes a predictions table automatically.

  • Low-latency online predictions need the BigQuery ML model registered in Model Registry and deployed to an endpoint, not ML.PREDICT.

  • Registering at training time uses MODEL_REGISTRY = 'VERTEX_AI' with VERTEX_AI_MODEL_ID; ALTER MODEL ... SET OPTIONS (vertex_ai_model_id = ...) registers an existing model.

  • Version aliases such as production let applications follow the approved model version without code changes.

Last updated: October 2026

7.3 Running Inference and Organizing Models in Model Registry

Core Focus: A trained model is useful only when it produces predictions where decisions happen. The exam guide asks you to "perform inference using BigQuery ML models" and "organize models in Model Registry." This section covers batch inference in SQL, the specialized inference functions, online prediction through endpoints, and versioning models in Model Registry.


Batch Inference in BigQuery

Most BigQuery ML predictions are batch: score a table of rows with SQL and write the results to another table that dashboards or applications read.

FunctionUse withReturns
ML.PREDICTClassification, regression, k-means, imported and remote modelspredicted_<label> (and predicted_<label>_probs for classifiers); for k-means, the nearest centroid
ML.FORECASTARIMA_PLUS, ARIMA_PLUS_XREGFuture values with prediction intervals
ML.DETECT_ANOMALIESARIMA_PLUS, KMEANS, autoencoder, PCAAn is_anomaly flag and anomaly probability per row
ML.RECOMMENDMATRIX_FACTORIZATIONPredicted ratings or confidence for user-item pairs
ML.EXPLAIN_PREDICTMany supervised modelsPredictions plus the features that contributed most
-- Score active customers nightly and keep the results in a table
CREATE OR REPLACE TABLE `customer_analytics.churn_scores` AS
SELECT
  customer_id,
  predicted_has_churned,
  (SELECT prob FROM UNNEST(predicted_has_churned_probs) WHERE label = TRUE) AS churn_probability
FROM ML.PREDICT(
  MODEL `customer_analytics.churn_model`,
  (SELECT * FROM `customer_analytics.active_customers`)
);

-- Forecast daily units for the next 30 days with 90% prediction intervals
SELECT *
FROM ML.FORECAST(
  MODEL `inventory.daily_units_arima`,
  STRUCT(30 AS horizon, 0.9 AS confidence_level)
);

-- Flag unusual sensor readings from an ARIMA_PLUS model
SELECT *
FROM ML.DETECT_ANOMALIES(
  MODEL `iot.temperature_arima`,
  STRUCT(0.95 AS anomaly_prob_threshold)
);

Wrap these statements in a scheduled query or a Dataform action to refresh predictions automatically. For classifiers, ML.PREDICT also accepts a threshold, for example STRUCT(0.7 AS threshold) as a third argument, when the default 0.5 does not match the business trade-off between precision and recall.

Models Trained Elsewhere

BigQuery ML can import models trained outside BigQuery (TensorFlow, TensorFlow Lite, ONNX, XGBoost) from Cloud Storage with CREATE MODEL ... OPTIONS (model_type = 'ONNX', model_path = 'gs://...') and then run ML.PREDICT on them in SQL. It can also export many BigQuery ML models to Cloud Storage with EXPORT MODEL for serving elsewhere.


Online Inference: When SQL Is Not Enough

ML.PREDICT runs as a query job, so it is not designed for an app that needs one prediction in milliseconds for each user request. For online predictions:

  1. Register the BigQuery ML model in Model Registry.
  2. Deploy it from Model Registry to an endpoint (Vertex AI endpoints, now listed under Gemini Enterprise Agent Platform). No custom serving container is needed for registered BigQuery ML models, although some model types have deployment limits.
  3. Call the endpoint's prediction API from the application.
NeedChoose
Score millions of rows nightly for a dashboardBatch: ML.PREDICT in a scheduled query
Return a churn score while a support agent is on the phoneOnline: Model Registry and an endpoint
Forecast next quarter's demand once a weekBatch: ML.FORECAST

Model Registry

Model Registry is the central place to manage models from BigQuery ML, AutoML, and custom training: versions, aliases, evaluation results, and deployment. BigQuery ML models can be registered without exporting them.

Registering a BigQuery ML Model

-- Register at training time
CREATE OR REPLACE MODEL `customer_analytics.churn_model`
OPTIONS (
  model_type = 'logistic_reg',
  input_label_cols = ['has_churned'],
  model_registry = 'vertex_ai',
  vertex_ai_model_id = 'customer_churn',
  vertex_ai_model_version_aliases = ['staging']
) AS
SELECT * EXCEPT (customer_id) FROM `customer_analytics.customer_features`;

-- Register an existing model
ALTER MODEL `customer_analytics.segment_model`
SET OPTIONS (vertex_ai_model_id = 'customer_segments');
  • MODEL_REGISTRY = 'VERTEX_AI' registers the model once training completes.
  • VERTEX_AI_MODEL_ID names the registry entry; each BigQuery ML model can be registered to only one model ID, and training a new BigQuery ML model with the same ID adds a new version under it.
  • VERTEX_AI_MODEL_VERSION_ALIASES attaches aliases such as staging or production; moving an alias to a newer version is a clean way to promote a model.
  • You can also register from the console (the model's Registry tab) or with bq update --model --vertex_ai_model_id.
  • Registering requires Agent Platform (Vertex AI) permissions, such as the roles/aiplatform.admin role.

What You Get from the Registry

CapabilityWhy it matters
VersionsKeep every retrained model and compare them
AliasesPoint applications at production instead of a version number
Evaluation viewCompare metrics across versions before promoting
DeploymentDeploy a version to an endpoint for online prediction
MonitoringWatch deployed models for skew and drift

Common Exam Traps

  • Using ML.PREDICT for low-latency app requests: batch SQL is the wrong tool; register and deploy the model to an endpoint.
  • Exporting a BigQuery ML model just to manage versions: Model Registry manages BigQuery ML models in place.
  • Hard-coding version numbers in applications: use aliases so promotion doesn't require application changes.
  • Forecasting with ML.PREDICT: time-series models use ML.FORECAST with a horizon.
Test Your Knowledge

Which query forecasts the next 30 days from an ARIMA_PLUS model with 90% prediction intervals?

A

SELECT * FROM ML.DETECT_ANOMALIES(MODEL m, STRUCT(30 AS horizon))

B

SELECT * FROM ML.EVALUATE(MODEL m, STRUCT(30 AS days, 0.9 AS confidence_level))

C

SELECT * FROM ML.PREDICT(MODEL m, STRUCT(30 AS horizon, 0.9 AS confidence_level))

D

SELECT * FROM ML.FORECAST(MODEL m, STRUCT(30 AS horizon, 0.9 AS confidence_level))

Test Your Knowledge

A team retrains a BigQuery ML model every month and wants applications to always call the approved version without code changes. What should they use?

A

Model Registry version aliases such as production

B

A table snapshot of the model, refreshed after each monthly retraining run

C

A new BigQuery dataset name for each monthly model so versions never collide

D

EXPORT MODEL to a new Cloud Storage path each month and update the path

Test Your Knowledge

A customer-support application must display a churn score within milliseconds while an agent is on a call. The model was trained with BigQuery ML. What should the team do?

A

Register the model in Model Registry and deploy it to an endpoint

B

Run ML.PREDICT from the application as a query for every incoming call

C

Export all predictions to a CSV file in Cloud Storage every night

D

Retrain the model with AutoML so it can return results faster

Sections you finish are checked off in the contents.