Vertex AI Cost Metrics
This document defines the shared cost-observability contract for all Hummingbird services that call Vertex AI. It covers two obligations:
- Prometheus metrics — estimated cost counters scraped into Mimir.
- Vertex request labels — billing metadata attached to every API call.
All services that make Vertex AI calls MUST implement both.
hummingbird-ai emits shared model request, token, and cost metrics for its
consumers. The dashboard continues to expose independent database-backed
gauges with the same names and labels.
1. Prometheus metrics (estimated cost)
Every Vertex AI caller MUST emit the following counters on its scraped
/metrics endpoint. Using identical names and label sets across services
enables a single Grafana dashboard without per-app query logic.
Counters
| Name | Unit | Labels | Description |
|---|---|---|---|
hummingbird_vertex_cost_dollars_total |
USD | model, workflow |
Estimated cost based on MODEL_PRICING x tokens |
hummingbird_vertex_tokens_total |
tokens | model, workflow, direction |
Token counts per API call |
hummingbird_vertex_requests_total |
requests | model, workflow, status |
API call attempts |
Labels
Applications MUST set these labels. Additional labels are permitted if cardinality remains bounded.
| Label | Values |
|---|---|
model |
Model name as sent to Vertex (e.g. gemini-2.5-flash, claude-sonnet-4-6) |
workflow |
Workflow or operation name (e.g. code-review, renovate-babysit) |
direction |
input, output, cache_read, cache_creation |
status |
success, error, retry, parse_error |
Request statuses are mutually exclusive per model attempt: success means the
response was accepted, error is a terminal model-call failure, retry marks
an attempt that will be retried, and parse_error marks a terminal
response-validation failure. Token and cost metrics still include every model
response, including responses rejected by validation.
Applications MUST NOT set cluster, namespace, app, service, or
pod — Alloy adds these automatically at scrape time via relabel rules.
Cardinality
model: bounded by deployed model set (currently 2-4)workflow: bounded by workflow config entries (~10 for agent, 1 for dashboard)direction: fixed set of 4status: fixed set of 4
Total series per app: model x workflow x max(direction, status) = 4 x 10 x 4 = 160. Well within Mimir per-metric limits.
2. Vertex request labels (billed cost)
Every Vertex API call MUST include billing labels so that GCP Cloud Billing rows can be attributed to a specific application and workflow once BigQuery export is enabled.
Required labels
| Key | Value | Example |
|---|---|---|
app |
Application name (matches Kubernetes app) | agent, dashboard |
workflow |
Workflow or operation name within the app | code-review, analyze-failures |
Callers MUST set both keys on every generateContent,
streamGenerateContent, or rawPredict request.
How to attach
| Endpoint | Mechanism |
|---|---|
generateContent (Gemini) |
Top-level "labels" key in JSON body |
rawPredict / streamRawPredict |
X-Vertex-AI-Labels HTTP header |
For generateContent, add the labels object to the request body:
{
"contents": [...],
"labels": {"app": "agent", "workflow": "code-review"}
}
For rawPredict, base64-encode the labels JSON and pass as a header
(shown decoded for clarity):
import base64
import json
labels = {"app": "dashboard", "workflow": "analyze-failures"}
headers["X-Vertex-AI-Labels"] = base64.b64encode(
json.dumps(labels).encode(),
).decode()
Constraints (GCP requirements)
- Keys and values: lowercase letters, numbers, underscores, dashes only
- Max 63 characters each
- Keys MUST start with a letter
- Labels are only forwarded for PayGo consumption (Provisioned Throughput silently ignores them)
Reference
Example PromQL
Total estimated spend yesterday (all apps):
sum(increase(hummingbird_vertex_cost_dollars_total[1d]))
Per-app daily cost:
sum by (app)(increase(hummingbird_vertex_cost_dollars_total[1d]))
Per-workflow breakdown for the agent:
sum by (workflow, model)(
increase(hummingbird_vertex_cost_dollars_total{app="hummingbird-agent"}[1d])
)
Token consumption by direction:
sum by (app, direction)(increase(hummingbird_vertex_tokens_total[1d]))
Implementing apps
| Application | Port | Workflow values |
|---|---|---|
| hummingbird-agent | 9090 | from wf_cfg.name per workflow |
| hummingbird-dashboard | 8080 | analyze-failures |