Vertex AI Cost Metrics

This document defines the shared cost-observability contract for all Hummingbird services that call Vertex AI. It covers two obligations:

  1. Prometheus metrics — estimated cost counters scraped into Mimir.
  2. Vertex request labels — billing metadata attached to every API call.

All services that make Vertex AI calls MUST implement both.

hummingbird-ai emits shared model request, token, and cost metrics for its consumers. The dashboard continues to expose independent database-backed gauges with the same names and labels.

1. Prometheus metrics (estimated cost)

Every Vertex AI caller MUST emit the following counters on its scraped /metrics endpoint. Using identical names and label sets across services enables a single Grafana dashboard without per-app query logic.

Counters

Name Unit Labels Description
hummingbird_vertex_cost_dollars_total USD model, workflow Estimated cost based on MODEL_PRICING x tokens
hummingbird_vertex_tokens_total tokens model, workflow, direction Token counts per API call
hummingbird_vertex_requests_total requests model, workflow, status API call attempts

Labels

Applications MUST set these labels. Additional labels are permitted if cardinality remains bounded.

Label Values
model Model name as sent to Vertex (e.g. gemini-2.5-flash, claude-sonnet-4-6)
workflow Workflow or operation name (e.g. code-review, renovate-babysit)
direction input, output, cache_read, cache_creation
status success, error, retry, parse_error

Request statuses are mutually exclusive per model attempt: success means the response was accepted, error is a terminal model-call failure, retry marks an attempt that will be retried, and parse_error marks a terminal response-validation failure. Token and cost metrics still include every model response, including responses rejected by validation.

Applications MUST NOT set cluster, namespace, app, service, or pod — Alloy adds these automatically at scrape time via relabel rules.

Cardinality

  • model: bounded by deployed model set (currently 2-4)
  • workflow: bounded by workflow config entries (~10 for agent, 1 for dashboard)
  • direction: fixed set of 4
  • status: fixed set of 4

Total series per app: model x workflow x max(direction, status) = 4 x 10 x 4 = 160. Well within Mimir per-metric limits.

2. Vertex request labels (billed cost)

Every Vertex API call MUST include billing labels so that GCP Cloud Billing rows can be attributed to a specific application and workflow once BigQuery export is enabled.

Required labels

Key Value Example
app Application name (matches Kubernetes app) agent, dashboard
workflow Workflow or operation name within the app code-review, analyze-failures

Callers MUST set both keys on every generateContent, streamGenerateContent, or rawPredict request.

How to attach

Endpoint Mechanism
generateContent (Gemini) Top-level "labels" key in JSON body
rawPredict / streamRawPredict X-Vertex-AI-Labels HTTP header

For generateContent, add the labels object to the request body:

{
  "contents": [...],
  "labels": {"app": "agent", "workflow": "code-review"}
}

For rawPredict, base64-encode the labels JSON and pass as a header (shown decoded for clarity):

import base64
import json

labels = {"app": "dashboard", "workflow": "analyze-failures"}
headers["X-Vertex-AI-Labels"] = base64.b64encode(
    json.dumps(labels).encode(),
).decode()

Constraints (GCP requirements)

  • Keys and values: lowercase letters, numbers, underscores, dashes only
  • Max 63 characters each
  • Keys MUST start with a letter
  • Labels are only forwarded for PayGo consumption (Provisioned Throughput silently ignores them)

Reference

Example PromQL

Total estimated spend yesterday (all apps):

sum(increase(hummingbird_vertex_cost_dollars_total[1d]))

Per-app daily cost:

sum by (app)(increase(hummingbird_vertex_cost_dollars_total[1d]))

Per-workflow breakdown for the agent:

sum by (workflow, model)(
  increase(hummingbird_vertex_cost_dollars_total{app="hummingbird-agent"}[1d])
)

Token consumption by direction:

sum by (app, direction)(increase(hummingbird_vertex_tokens_total[1d]))

Implementing apps

Application Port Workflow values
hummingbird-agent 9090 from wf_cfg.name per workflow
hummingbird-dashboard 8080 analyze-failures