7.3 Running Inference and Organizing Models in Model Registry
Key Takeaways
ML.PREDICT scores rows in batch with SQL; ML.FORECAST, ML.DETECT_ANOMALIES, ML.RECOMMEND, and ML.EXPLAIN_PREDICT cover forecasting, anomalies, recommendations, and explanations.
Wrapping ML.PREDICT in a scheduled query or Dataform action refreshes a predictions table automatically.
Low-latency online predictions need the BigQuery ML model registered in Model Registry and deployed to an endpoint, not ML.PREDICT.
Registering at training time uses MODEL_REGISTRY = 'VERTEX_AI' with VERTEX_AI_MODEL_ID; ALTER MODEL ... SET OPTIONS (vertex_ai_model_id = ...) registers an existing model.
Version aliases such as production let applications follow the approved model version without code changes.
7.3 Running Inference and Organizing Models in Model Registry
Core Focus: A trained model is useful only when it produces predictions where decisions happen. The exam guide asks you to "perform inference using BigQuery ML models" and "organize models in Model Registry." This section covers batch inference in SQL, the specialized inference functions, online prediction through endpoints, and versioning models in Model Registry.
Batch Inference in BigQuery
Most BigQuery ML predictions are batch: score a table of rows with SQL and write the results to another table that dashboards or applications read.
| Function | Use with | Returns |
|---|---|---|
ML.PREDICT | Classification, regression, k-means, imported and remote models | predicted_<label> (and predicted_<label>_probs for classifiers); for k-means, the nearest centroid |
ML.FORECAST | ARIMA_PLUS, ARIMA_PLUS_XREG | Future values with prediction intervals |
ML.DETECT_ANOMALIES | ARIMA_PLUS, KMEANS, autoencoder, PCA | An is_anomaly flag and anomaly probability per row |
ML.RECOMMEND | MATRIX_FACTORIZATION | Predicted ratings or confidence for user-item pairs |
ML.EXPLAIN_PREDICT | Many supervised models | Predictions plus the features that contributed most |
-- Score active customers nightly and keep the results in a table
CREATE OR REPLACE TABLE `customer_analytics.churn_scores` AS
SELECT
customer_id,
predicted_has_churned,
(SELECT prob FROM UNNEST(predicted_has_churned_probs) WHERE label = TRUE) AS churn_probability
FROM ML.PREDICT(
MODEL `customer_analytics.churn_model`,
(SELECT * FROM `customer_analytics.active_customers`)
);
-- Forecast daily units for the next 30 days with 90% prediction intervals
SELECT *
FROM ML.FORECAST(
MODEL `inventory.daily_units_arima`,
STRUCT(30 AS horizon, 0.9 AS confidence_level)
);
-- Flag unusual sensor readings from an ARIMA_PLUS model
SELECT *
FROM ML.DETECT_ANOMALIES(
MODEL `iot.temperature_arima`,
STRUCT(0.95 AS anomaly_prob_threshold)
);
Wrap these statements in a scheduled query or a Dataform action to refresh predictions automatically. For classifiers, ML.PREDICT also accepts a threshold, for example STRUCT(0.7 AS threshold) as a third argument, when the default 0.5 does not match the business trade-off between precision and recall.
Models Trained Elsewhere
BigQuery ML can import models trained outside BigQuery (TensorFlow, TensorFlow Lite, ONNX, XGBoost) from Cloud Storage with CREATE MODEL ... OPTIONS (model_type = 'ONNX', model_path = 'gs://...') and then run ML.PREDICT on them in SQL. It can also export many BigQuery ML models to Cloud Storage with EXPORT MODEL for serving elsewhere.
Online Inference: When SQL Is Not Enough
ML.PREDICT runs as a query job, so it is not designed for an app that needs one prediction in milliseconds for each user request. For online predictions:
- Register the BigQuery ML model in Model Registry.
- Deploy it from Model Registry to an endpoint (Vertex AI endpoints, now listed under Gemini Enterprise Agent Platform). No custom serving container is needed for registered BigQuery ML models, although some model types have deployment limits.
- Call the endpoint's prediction API from the application.
| Need | Choose |
|---|---|
| Score millions of rows nightly for a dashboard | Batch: ML.PREDICT in a scheduled query |
| Return a churn score while a support agent is on the phone | Online: Model Registry and an endpoint |
| Forecast next quarter's demand once a week | Batch: ML.FORECAST |
Model Registry
Model Registry is the central place to manage models from BigQuery ML, AutoML, and custom training: versions, aliases, evaluation results, and deployment. BigQuery ML models can be registered without exporting them.
Registering a BigQuery ML Model
-- Register at training time
CREATE OR REPLACE MODEL `customer_analytics.churn_model`
OPTIONS (
model_type = 'logistic_reg',
input_label_cols = ['has_churned'],
model_registry = 'vertex_ai',
vertex_ai_model_id = 'customer_churn',
vertex_ai_model_version_aliases = ['staging']
) AS
SELECT * EXCEPT (customer_id) FROM `customer_analytics.customer_features`;
-- Register an existing model
ALTER MODEL `customer_analytics.segment_model`
SET OPTIONS (vertex_ai_model_id = 'customer_segments');
MODEL_REGISTRY = 'VERTEX_AI'registers the model once training completes.VERTEX_AI_MODEL_IDnames the registry entry; each BigQuery ML model can be registered to only one model ID, and training a new BigQuery ML model with the same ID adds a new version under it.VERTEX_AI_MODEL_VERSION_ALIASESattaches aliases such asstagingorproduction; moving an alias to a newer version is a clean way to promote a model.- You can also register from the console (the model's Registry tab) or with
bq update --model --vertex_ai_model_id. - Registering requires Agent Platform (Vertex AI) permissions, such as the
roles/aiplatform.adminrole.
What You Get from the Registry
| Capability | Why it matters |
|---|---|
| Versions | Keep every retrained model and compare them |
| Aliases | Point applications at production instead of a version number |
| Evaluation view | Compare metrics across versions before promoting |
| Deployment | Deploy a version to an endpoint for online prediction |
| Monitoring | Watch deployed models for skew and drift |
Common Exam Traps
- Using
ML.PREDICTfor low-latency app requests: batch SQL is the wrong tool; register and deploy the model to an endpoint. - Exporting a BigQuery ML model just to manage versions: Model Registry manages BigQuery ML models in place.
- Hard-coding version numbers in applications: use aliases so promotion doesn't require application changes.
- Forecasting with
ML.PREDICT: time-series models useML.FORECASTwith ahorizon.
Which query forecasts the next 30 days from an ARIMA_PLUS model with 90% prediction intervals?
SELECT * FROM ML.DETECT_ANOMALIES(MODEL m, STRUCT(30 AS horizon))
SELECT * FROM ML.EVALUATE(MODEL m, STRUCT(30 AS days, 0.9 AS confidence_level))
SELECT * FROM ML.PREDICT(MODEL m, STRUCT(30 AS horizon, 0.9 AS confidence_level))
SELECT * FROM ML.FORECAST(MODEL m, STRUCT(30 AS horizon, 0.9 AS confidence_level))
A team retrains a BigQuery ML model every month and wants applications to always call the approved version without code changes. What should they use?
Model Registry version aliases such as production
A table snapshot of the model, refreshed after each monthly retraining run
A new BigQuery dataset name for each monthly model so versions never collide
EXPORT MODEL to a new Cloud Storage path each month and update the path
A customer-support application must display a churn score within milliseconds while an agent is on a call. The model was trained with BigQuery ML. What should the team do?
Register the model in Model Registry and deploy it to an endpoint
Run ML.PREDICT from the application as a query for every incoming call
Export all predictions to a CSV file in Cloud Storage every night
Retrain the model with AutoML so it can return results faster
Sections you finish are checked off in the contents.