Hummingbird Dashboard
Web dashboard and CLI for monitoring Konflux build pipeline status.
Features
- Web Dashboard - Real-time view of build, test, and release status
- Commits View - Detailed pipeline status per commit with expandable releases
- Components View - Latest status per component grouped by state
- Merge Requests View - Track build status across MR versions
- CLI Tool - Command-line access to pipeline data in YAML/JSON/table formats
- Failed Releases - View and manage failed releases with LLM-powered analysis
- Failed Push Builds - Detect and auto-retry failed on-push Konflux builds
- Auto-Rerun - Automatically retry transient release and push build failures
- Blocked Snapshots - Block/unblock snapshots from auto-rerun with error pattern matching
- Smart Triggering - Application-aware rules for determining affected components
- Error Budget Alerts - Slack notifications when CVE R-Time error budget hits 20%, 5%, or 0% remaining
- Blame Page -
/cve/blameview listing late unblamed tickets with per-category breakdown and per-ticket stage detail - ProdSec CVE-HUM SLA Auto-Blame - Auto-locks ProdSec blame when CVE-HUM time exceeds 6h
- Image Pull Availability - Estimated SLI-UE1-B availability from manifest and sampled checks
- RPM Tracker API - Read-only snapshot of tracker-owned tables in the shared database
Prerequisites
- Python 3.11+
- PostgreSQL database (via hummingbird-status or port-forward)
Installation
cd hummingbird-dashboard
pip install -e .
Usage
Local Development
cd hummingbird-dashboard
# Option 1: Port-forward to production database
./dev.sh port-forward # In terminal 1
DATABASE_URL=postgresql://postgres@localhost:15432/events ./dev.sh start # In terminal 2
# Option 2: Use local database (via hummingbird-status)
cd ../hummingbird-status && ./dev.sh db-start && ./dev.sh db-init
cd ../hummingbird-dashboard
DATABASE_URL=postgresql://postgres:dev@localhost:5432/events ./dev.sh start
The dashboard runs at http://localhost:8080 with live reload.
Web UI
| Endpoint | Description |
|---|---|
/ |
Dashboard with all applications |
/apps/{app}/commits |
Commits view for an application |
/apps/{app}/components |
Components view for an application |
/mrs |
Merge requests overview (filterable) |
/mr/{project}/{iid} |
MR detail with build status per version |
/releases/failed |
Failed releases with analysis |
/releases/blocked |
Blocked snapshots management |
/releases/analysis-log |
LLM analysis run history |
/cve/status |
CVE ticket status dashboard |
/cve/run-log |
CVE analysis run history |
/cve/open |
Open CVE trackers with SLO status |
/cve/closed |
Closed tickets with VEX reconciliation |
/cve/r-time |
CVE R-Time (notified to done) metrics |
/cve/blame |
Late unblamed tickets and per-category blame breakdown |
/availability |
Estimated image pull availability (SLI-UE1-B) |
/health |
Health check endpoint |
/metrics |
Prometheus metrics |
Each CVE tracker page (/cve/status, /cve/open, /cve/closed,
/cve/r-time) has a “Download CSV” button at the bottom of the page that downloads
the current view (same filters as the page) as a CSV file, via a matching
/cve/{tab}/csv route (for example /cve/open/csv?slo=24). A “Show CVE ID”
toggle in the header reveals the CVE ID next to each HUM ticket key
(HUM-1234/CVE-2026-56789); the preference persists via localStorage.
CVE Search API
GET /api/cve-runs/search — search CVE analysis run history.
| Parameter | Type | Default | Description |
|---|---|---|---|
package |
string | — | Filter by package name (substring match) |
ticket |
string | — | Filter by ticket key (substring match) |
q |
string | — | Free-text search across ticket details |
log |
string | — | Search CronJob log output |
days |
int | 30 | Time window in days (1–365) |
limit |
int | 500 | Max results (1–2000) |
At least one of package, ticket, q, or log is required.
The response includes two result sets:
results— per-ticket matches from the JSONB details columnlog_matches— per-run matches from CronJob log output, each with alog_excerptshowing the matching lines in context
Example:
# Find runs whose logs mention "merge train" — show excerpts
curl -s '/api/cve-runs/search?log=merge+train' | jq '.log_matches[] | {run_at, log_excerpt}'
# Find tickets for a package, also searching logs
curl -s '/api/cve-runs/search?package=openssl&log=advisory+failed&days=7' \
| jq '{tickets: [.results[].detail.key], log_hits: .log_matches | length}'
# List all ticket keys matching a free-text query
curl -s '/api/cve-runs/search?q=needs-attention&days=14' | jq '[.results[].detail.key] | unique'
CLI
# View latest commit status
hummingbird-dashboard --application myapp --format table commit
# View specific commit
hummingbird-dashboard --application myapp commit abc1234
# View multiple commits
hummingbird-dashboard --application myapp --format table commit --limit 10
# Component status overview
hummingbird-dashboard --application myapp component
# JSON output
hummingbird-dashboard --format json commit | jq .
# Auto-rerun failed releases
hummingbird-dashboard auto-rerun
# Analyze failed releases via LLM
hummingbird-dashboard analyze-failures
# Probe catalog image pull availability (HUM-721)
hummingbird-dashboard check-image-pulls
hummingbird-dashboard check-image-pulls --dry-run
hummingbird-dashboard check-image-pulls --force
REST API Reference
All endpoints return JSON. Read-only GET endpoints are unauthenticated.
Interactive OpenAPI/Swagger documentation is available at /docs.
RPM Tracker
GET /api/rpm-upstream-state returns the collection generation and checkout
SHA with upstream, package status, and rollout rows from the independently
migrated rpm_tracker schema. The dashboard does not create or write these
tables. Run the tracker initializer before deploying this API; until a compatible
tracker schema is present, this endpoint returns HTTP 503. See
RPM Tracker for migration and storage details.
Tier 1 — Core Pipeline Status
GET /api/apps/{app_name}/commits
Pipeline status for commits in an application.
| Parameter | Type | Default | Description |
|---|---|---|---|
sha |
string | — | Filter by commit SHA (prefix) |
component |
string | — | Filter by component name |
limit |
int | 10 | Max commits to return (max 100) |
curl -s 'http://localhost:8080/api/apps/rpms/commits?limit=5' | jq .
curl -s 'http://localhost:8080/api/apps/rpms/commits?sha=abc123&component=openssl' | jq .
GET /api/apps/{app_name}/components
Latest build status per component, grouped by state.
| Parameter | Type | Default | Description |
|---|---|---|---|
search |
string | — | Filter components by name (substring) |
curl -s 'http://localhost:8080/api/apps/rpms/components' | jq .
curl -s 'http://localhost:8080/api/apps/rpms/components?search=openssl' | jq .
Tier 2 — Operational Visibility
GET /api/dashboard
Full overview across all applications: aggregate build counts, failed releases, component staleness, and recent activity.
curl -s 'http://localhost:8080/api/dashboard' | jq .
GET /api/releases/failed
List of currently failed releases with analysis results.
| Parameter | Type | Default | Description |
|---|---|---|---|
application |
string | — | Filter by application name |
curl -s 'http://localhost:8080/api/releases/failed' | jq .
curl -s 'http://localhost:8080/api/releases/failed?application=rpms' | jq .
GET /api/releases/{name}/analysis
LLM failure analysis for a specific release (root cause, classification, recommendation).
curl -s 'http://localhost:8080/api/releases/my-release-abc/analysis' | jq .
GET /api/releases/auto-rerun-log
History of automatic rerun attempts.
| Parameter | Type | Default | Description |
|---|---|---|---|
limit |
int | 50 | Max entries |
curl -s 'http://localhost:8080/api/releases/auto-rerun-log?limit=10' | jq .
GET /api/releases/analysis-log
History of LLM analysis runs.
| Parameter | Type | Default | Description |
|---|---|---|---|
limit |
int | 50 | Max entries |
curl -s 'http://localhost:8080/api/releases/analysis-log?limit=10' | jq .
GET /api/releases/analysis-costs
Cumulative token usage and estimated cost for LLM analysis runs.
curl -s 'http://localhost:8080/api/releases/analysis-costs' | jq .
GET /api/mrs
Merge requests across monitored repositories.
| Parameter | Type | Default | Description |
|---|---|---|---|
state |
string | opened |
MR state: opened, merged, closed, all |
project |
string | — | Filter by project (repository) name |
curl -s 'http://localhost:8080/api/mrs' | jq .
curl -s 'http://localhost:8080/api/mrs?state=merged&project=rpms' | jq .
GET /api/mrs/projects
List of all projects (repositories) with tracked merge requests.
curl -s 'http://localhost:8080/api/mrs/projects' | jq .
GET /api/mr/{project}/{iid}
Detail for a single merge request, including all versions and per-version build status.
curl -s 'http://localhost:8080/api/mr/rpms/42' | jq .
GET /api/cve/status
Current CVE ticket status across all tracked packages.
curl -s 'http://localhost:8080/api/cve/status' | jq .
GET /api/cve/open
Open HUM CVE tracker tickets with SLO status. Notified is the later of HUM ticket created and fix available (upstream or Fedora). SLO is green PASS unless a fix is available and elapsed time is within 8h of the selected SLO (yellow AT RISK) or past it (red FAIL).
| Parameter | Type | Default | Description |
|---|---|---|---|
slo |
float | 24 | SLO threshold hours (24, 72, or 168) |
include_next_release |
bool | true | Include cve-next-release tickets |
curl -s 'http://localhost:8080/api/cve/open?slo=24' | jq .
GET /api/cve/closed
Closed HUM CVE tracker tickets with Red Hat CSAF VEX reconciliation
(HUM-5843). Analysis stores package-scoped vex_status and Jira
resolution; the dashboard computes MATCH from those facts
(Done-Errata → fixed, Not a Bug → known_not_affected,
package_not_listed, or fixed – a downstream package rebuilt by an
unrelated dependency’s fix is not a contradiction of a Not a Bug verdict).
Shows Hummingbird vex_status, match state
(matched / pending / mismatch), and vex_resolved (first
scan time where VEX agreed with the Jira resolution). Event metadata is
merged across a ticket’s rows in occurred_at order (HUM-6091): later
non-empty fields overlay earlier ones, so Close/VEX facts are not stuck
on the first cve_published row. The Resolution column uses Jira
Closed / {resolution} when stored analysis text is not already Closed.
The Closed tab shows the VEX timestamp as a link to the CVE’s Red Hat
CSAF VEX document when the package is named in CSAF. package_not_listed
(Not a Bug, package absent from the document) shows N/A with no link:
this ticket did not produce a VEX change.
| Parameter | Type | Default | Description |
|---|---|---|---|
days |
int | 30 | Time window (0 = all time) |
curl -s 'http://localhost:8080/api/cve/closed?days=30' | jq .
GET /api/cve/r-time
CVE R-Time (notified to done) for Done-Errata tickets. Include CVE-HUM starts at
CVE publication (NVD datePublished / cve_published), falling back to HUM
ticket creation only when CVE publication is missing – CVE publication wins
even when the HUM ticket predates it, so an embargoed ticket’s R-Time starts
at public disclosure rather than at ticket creation (which can predate
disclosure by days while the CVE is still privately embargoed). Exclude
CVE-HUM starts at the later of HUM created and CVE publication – same as HUM
created for a normal ticket (Exclude’s purpose is to ignore filing lag after
disclosure), but an embargoed ticket’s clock must never start before public
disclosure in either mode, so it also starts at CVE publication there.
Delivery is the catalog
image publish when the package is a catalog image source, otherwise the RPM
publish. Include VEX requires delivery, VEX updated, and Closed / Done-Errata;
Exclude VEX requires delivery and close only. The done timestamp is the latest
of the required gates, and only when all of them exist. Incomplete tickets stay
on the table with elapsed time until now; the days window filters by
Done-Errata close (fallback: notified), same as completed tickets use done.
CVE-HUM is filing lag after NVD (max(HUM created − NVD, 0));
HUM-first tickets show 0.0h. ADV-VEX is Done-Errata close to VEX feed update
(max(VEX − close, 0)); close is recorded when the advisory MR merges.
Upstream/Fedora are not start or end. Pre-built (delivery before notification)
is informational only. Rows without catalog_image_source in metadata still
require an image until the collector restamps them. Each entry includes
package onboarding timestamps and delay analysis (rpm_first_published_at,
pkg_lag_hours, and our_delay_hours). Default excludes
cve-next-release, CVE-HUM time, and VEX time.
| Parameter | Type | Default | Description |
|---|---|---|---|
days |
int | 30 | Time window (0 = all time) |
slo |
float | 168 | SLO threshold hours (24, 72, or 168) |
include_next_release |
bool | false | Include cve-next-release tickets |
include_hum_cve_time |
bool | false | Include NVD-to-HUM filing lag in R-Time |
include_vex_time |
bool | false | Require VEX feed update before R-Time ends |
curl -s 'http://localhost:8080/api/cve/r-time?days=10&slo=168' | jq .
GET /api/cve/run-log
History of CVE analysis CronJob runs.
| Parameter | Type | Default | Description |
|---|---|---|---|
limit |
int | 100 | Max entries |
curl -s 'http://localhost:8080/api/cve/run-log?limit=10' | jq .
GET /api/cve-export
Export CVE sync data as JSON for prod-to-preprod replication. Valid section
values are cve_analysis_log, cve_ticket_events, package_lifecycle, and
cve_ticket_blame.
Optional pagination parameters:
| Parameter | Type | Default | Description |
|---|---|---|---|
section |
string | — | Section to export |
limit |
int | 2000 | Rows per page |
offset |
int | 0 | Pagination offset |
curl -s -H "Authorization: Bearer $CVE_REPORT_TOKEN" \
'http://localhost:8080/api/cve-export' | jq '.cve_ticket_events | length'
curl -s -H "Authorization: Bearer $CVE_REPORT_TOKEN" \
'http://localhost:8080/api/cve-export?section=cve_ticket_events&limit=1000&offset=0' \
| jq '.count,.has_more'
# Package-scoped lifecycle milestones (rpm_first_published, etc.)
curl -s -H "Authorization: Bearer $CVE_REPORT_TOKEN" \
'http://localhost:8080/api/cve-export?section=package_lifecycle&limit=10000&offset=0' \
| jq '.rows[] | select(.event_type == "rpm_first_published")'
Event type validation. The import path (POST /api/cve-import) validates
every event_type value against a canonical allowlist defined in
hummingbird_dashboard/sources.py:
CANONICAL_TICKET_EVENT_TYPES— values accepted forcve_ticket_events:cve_published,osidb_flaw_created,osidb_affect_created,cve_embargo_lifted,hum_ticket_created,hum_ticket_closed,upstream_fix_merged,fedora_update_available,rpm_fix_published_to_pulp,image_rebuilt_on_quay,vex_resolved.CANONICAL_PACKAGE_EVENT_TYPES— values accepted forpackage_lifecycle:rpm_first_published.
Legacy names (e.g. jira_created, delivered_in_rpm, …) are
mapped to their canonical equivalents via _LEGACY_EVENT_KEY_MAP.
Unknown names are rejected with HTTP 400.
POST /api/cve-import
Replace local CVE sync data from a payload previously returned by
/api/cve-export.
curl -s -X POST \
-H "Authorization: Bearer $CVE_REPORT_TOKEN" \
-H "Content-Type: application/json" \
--data @cve-export.json \
'http://localhost:8080/api/cve-import' | jq .
curl -s -X POST \
-H "Authorization: Bearer $CVE_REPORT_TOKEN" \
-H "Content-Type: application/json" \
--data '{"section":"cve_ticket_events","rows":[],"reset":true}' \
'http://localhost:8080/api/cve-import' | jq .
To copy prod CVE sync rows onto preprod (both sections; first page of each
section uses reset=true):
CVE_REPORT_TOKEN=... ./hummingbird-cve-analysis/scripts/copy_prod_to_preprod.sh
CVE_REPORT_TOKEN=... ./hummingbird-cve-analysis/scripts/copy_prod_to_preprod.sh --dry-run
The script prefers paginated /api/cve-export. If that endpoint returns HTTP
500, it retries unpaged /api/cve-export. Analysis logs may fall back to
/api/cve/run-log (latest 100 runs). Ticket events are not reconstructed
from Open/Closed/R-Time: those views omit ticket-event metadata. Before
import, ticket events collapse pre-HUM-5918 names (jira_closed,
delivered_in_image, …) onto canonical types so a null legacy row cannot
wipe a timestamp or VEX label. Prod ticket-event rows often omit
catalog_image_source; the copy fills that flag from the catalog source
map (same rule as the collector) and keeps prod’s True/False when present.
The copy fails if the catalog map cannot be built. Ticket-event types such as
osidb_flaw_created, osidb_affect_created, and cve_embargo_lifted copy
with the rest of cve_ticket_events because they are on
CANONICAL_TICKET_EVENT_TYPES.
occurred_at conflict resolution on import (both the paginated and
full-replace code paths) matches the live collector path (HUM-6860):
rpm_fix_published_to_pulp and image_rebuilt_on_quay are mutable delivery
timestamps where a non-null incoming value wins, so a sync can update them,
without letting a null incoming value erase an already-recorded delivery
timestamp; every other event_type is write-once, keeping the existing
timestamp once set so a sync or backfill re-run cannot overwrite an
established immutable milestone.
cve_analysis_log is paged 5 rows at a time by default. Each row includes
full log_output (up to 512 KiB). Ticket events default to 1000 rows per page.
Tier 3 — Analytics & SRE Metrics
GET /api/stats/build-throughput
Build counts over time (successful, failed, total).
| Parameter | Type | Default | Description |
|---|---|---|---|
days |
int | 30 | Lookback window in days |
curl -s 'http://localhost:8080/api/stats/build-throughput?days=7' | jq .
GET /api/stats/release-success-rate
Release success/failure ratios over time.
| Parameter | Type | Default | Description |
|---|---|---|---|
days |
int | 30 | Lookback window in days |
application |
string | — | Filter by application |
curl -s 'http://localhost:8080/api/stats/release-success-rate?days=14&application=rpms' | jq .
GET /api/stats/auto-rerun-effectiveness
Success rate and time-to-resolution for automatic reruns.
| Parameter | Type | Default | Description |
|---|---|---|---|
days |
int | 30 | Lookback window in days |
curl -s 'http://localhost:8080/api/stats/auto-rerun-effectiveness?days=7' | jq .
GET /api/components/staleness
Components ranked by time since last successful build.
curl -s 'http://localhost:8080/api/components/staleness' | jq .
GET /api/releases/similar-failures
Find releases with error messages similar to a given string.
| Parameter | Type | Default | Description |
|---|---|---|---|
error |
string | — | Error text to match against |
curl -s 'http://localhost:8080/api/releases/similar-failures?error=timeout+connecting' | jq .
GET /api/apps/{app_name}/last-successful
Timestamp and SHA of the last fully successful pipeline per component.
curl -s 'http://localhost:8080/api/apps/rpms/last-successful' | jq .
Tier 4 — System
GET /api/service-status/json
Health of upstream services the dashboard depends on (database, Konflux API, KubeArchive, etc.).
curl -s 'http://localhost:8080/api/service-status/json' | jq .
GET /api/settings
Current runtime settings (auto-rerun enabled, analysis enabled, blocked patterns, etc.).
curl -s 'http://localhost:8080/api/settings' | jq .
GET /health
Enriched health check returning service version, uptime, database connectivity, and dependency status.
curl -s 'http://localhost:8080/health' | jq .
Configuration
Environment Variables
| Variable | Default | Description |
|---|---|---|
DATABASE_URL |
postgresql://...localhost:5432/ |
PostgreSQL connection URL |
PORT |
8080 |
Web server port |
Authentication Variables
These control the retrigger functionality (requires OAuth proxy in production):
| Variable | Default | Description |
|---|---|---|
TRIGGER_AUTH_MODE |
oauth |
oauth (production) or local |
TRIGGER_AUTH_GROUP |
konflux-hummingbird-admin-access |
OpenShift group required to retrigger |
TRIGGER_LOCAL_USER |
local-dev |
Username when TRIGGER_AUTH_MODE=local |
CVE_REPORT_TOKEN |
Bearer token for CVE API auth |
For local development with retrigger enabled:
TRIGGER_AUTH_MODE=local DATABASE_URL=... ./dev.sh start
Auto-Rerun Variables
These control the auto-rerun cronjob and failure analysis:
| Variable | Default | Description |
|---|---|---|
AUTO_RERUN_MIN_AGE_MINUTES |
30 |
Minimum failure age before retrying |
AUTO_RERUN_MAX_RETRIES |
3 |
Max rerun attempts per snapshot+plan |
KONFLUX_KUBECONFIG_PATH |
Path to Konflux kubeconfig file | |
RELEASE_NAMESPACE |
Namespace where releases are created | |
MANAGED_NAMESPACE |
Namespace for fetching release PLRs | |
KUBEARCHIVE_URL |
KubeArchive API URL for archived resources | |
GOOGLE_APPLICATION_CREDENTIALS |
Path to GCP SA key for Vertex AI | |
GOOGLE_CLOUD_PROJECT |
GCP project ID for Vertex AI | |
ANALYSIS_MODEL_API_KEY |
Gemini API key (fallback if no GCP creds) | |
ANALYSIS_MODEL |
gemini-2.5-flash |
LLM model for failure analysis |
ANALYSIS_MODEL_REGION |
global |
Vertex AI region |
MAX_ANALYSES_PER_CYCLE |
5 |
Max releases to analyze per cycle |
SLACK_WEBHOOK_URL |
Slack webhook for rerun/analysis notifications | |
DASHBOARD_URL |
Dashboard base URL for Slack links | |
GITLAB_SLACK_MAP_PATH |
/etc/hummingbird/gitlab-slack-map.yaml |
GitLab→Slack map for author @-mentions (infra-mounted; missing omits mention) |
PUSH_RERUN_MIN_AGE_MINUTES |
10 |
Min push build failure age before retrying |
PUSH_RERUN_MAX_RETRIES |
3 |
Max retry attempts per component+sha |
Error Budget Alert Variables
The dashboard pod periodically evaluates CVE R-Time error budget (168h SLO, 30-day window) and posts to Slack on worsening transitions only:
| Remaining budget | Slack alert |
|---|---|
| ≤20% | Warning |
| ≤5% | Critical |
| 0% | Exhausted (sprint-stop) |
Incoming webhooks always post to the channel they were created for; the
payload cannot select #team-hummingbird. Create the webhook for that
channel and set ERROR_BUDGET_SLACK_WEBHOOK_URL (or point the shared
SLACK_WEBHOOK_URL at that webhook).
State is stored in dashboard_settings (error_budget_rtime_alert_state) so
the same threshold is not re-alerted. The stored state is claimed with a
compare-and-set before Slack is called, and rolled back if the post fails, so
a transient Slack error is retried on the next interval and concurrent
evaluators cannot double-post the same transition. Recovery updates the
stored state silently so a later re-worsening can alert again. No message is
sent when the webhook URL is unset, or when alerts are disabled via the
R-Time page toggle (error_budget_alerts_enabled in dashboard_settings;
default enabled).
| Variable | Default | Description |
|---|---|---|
ERROR_BUDGET_SLACK_WEBHOOK_URL |
SLACK_WEBHOOK_URL |
Incoming webhook created for the alert channel (typically #team-hummingbird) |
SLACK_WEBHOOK_URL |
Fallback webhook if the error-budget URL is unset | |
DASHBOARD_URL |
Base URL for the R-Time deep link in alerts | |
ERROR_BUDGET_ALERT_INTERVAL_SECONDS |
900 |
Seconds between background checks |
Failures
The failures section is accessible via the “Failures” nav item and provides two views selectable by tab: Releases and Push Builds.
Failed Releases
The /failures/releases page shows all failed releases with:
- LLM Analysis — Each failure is analyzed by Gemini with root cause, classification, and recommendation
- Failure Classification — Transient, Configuration, Code, External Service, or Unknown
- Rerun History — Past rerun attempts and outcomes per snapshot
- Error Pattern Matching — Auto-block snapshots matching known error patterns
- Auto-Rerun — Automatically retry transient failures (configurable via dashboard toggle)
- Analysis Toggle — Enable/disable LLM analysis from the dashboard
CLI Subcommands
The auto-rerun subcommand retries eligible failed releases:
hummingbird-dashboard auto-rerun --application myapp
The auto-rerun-push subcommand retries eligible failed push builds:
hummingbird-dashboard auto-rerun-push --application myapp
The analyze-failures subcommand runs LLM analysis on unanalyzed failures:
hummingbird-dashboard analyze-failures --managed-namespace rhtap-releng-tenant
All three are designed to run as Kubernetes CronJobs. Slack messages for auto-rerun, push-retry, analysis, and manual UI reruns include a GitLab MR link (and author @-mention when mapped) when the failure SHA resolves to an MR.
Image Pull Availability
The /availability page tracks HUM-721 (SLI-UE1-B): can a customer actually
pull the images we publish. Continuously pulling every catalog image on a
tight schedule is not feasible (terabytes of registry traffic), so the check
is split into three tiers, run by check-image-pulls:
- Manifest HEAD (exhaustive) — a cheap HEAD request, no image bytes
transferred, against every
latest*image:variant tag in the catalog on every run. This is the headline “Estimated Availability” number. A digest cache skips the live HEAD when the tag’s digest is unchanged, the previous check succeeded, and it was confirmed within the last 6 hours. - Blob sample (sampled) — a real HEAD against every blob (config + layers) of a small random sample of currently-available images each run, confirming referenced layers still exist in storage.
- Full-pull sample (sampled) — a real GET of one image’s config blob (a few KB, following the registry’s CDN redirect) each run — genuine end-to-end retrieval, not just a HEAD.
Both sample tiers prioritize tags published in the last hour over older ones, since “did the thing we just shipped actually land” is the highest-value failure mode. The resulting availability % is estimated: the manifest tier is exhaustive, but the blob/full-pull tiers are sampled, not exhaustive — the page shows all three separately rather than blending them into one number.
- SLO: 95% target, 28-day rolling window, based on the manifest tier
- Error Budget: manifest-check misses allowed within the window
hummingbird-dashboard check-image-pulls
hummingbird-dashboard check-image-pulls --dry-run
hummingbird-dashboard check-image-pulls --force
| Option | Description |
|---|---|
--dry-run |
Probe without writing results to the database |
--force |
Force live manifest HEAD checks (ignore digest cache) |
--catalog-api-base |
Catalog API base URL (default: CATALOG_API_BASE or built-in) |
Designed to run as a Kubernetes CronJob (every 15 minutes). Checks can be
disabled from the dashboard via the image_pull_checks_enabled setting.
Merge Requests
The /mrs page shows merge requests across all monitored repositories with:
- State filter - Open, merged, closed, or all MRs
- Project filter - Filter by specific repository
- Build status - Aggregate status across all components
The /mr/{project}/{iid} detail page shows:
- All MR versions - Each head commit SHA that was pushed to the MR
- Build status per version - Full pipeline status (build, snapshot, test, release)
- Links to Konflux UI - Direct links to PipelineRuns and Snapshots
Versions are ordered by latest event timestamp, so force-pushed commits appear in the correct position even if they reuse an earlier SHA.
Component Status
The dashboard tracks component build status with these states:
| Status | Icon | Description |
|---|---|---|
| Success | ✅ | Build passed for expected commit |
| Superseded | 🔄 | Expected commit not built, but newer succeeded |
| Failed | ❌ | Build failed for expected commit |
| Stale | ⚠️ | Build not triggered for expected commit |
| Running | ⏳ | Build in progress |
| Missing | ❓ | No build found |
Components are grouped by status: Failed/Stale → Running → OK → Missing.
Trigger Rules
The dashboard uses application-specific rules to determine which components are affected by a push:
- containers: Changes in
images/{component}trigger builds - rpms: Changes in
rpms/{component}ormock/mock.cfgtrigger builds - tools: Changes in
{component}directory trigger builds
Certain files are excluded from triggering (README, templates, etc.).
Development
See the main README for development workflows.
Running Tests
cd hummingbird-dashboard
pip install -e ".[dev]"
pytest
The PostgreSQL migration regression test uses the CI PostgreSQL sidecar via
DATABASE_URL. Locally, set DATABASE_URL to a disposable PostgreSQL database
to run it; without one, the test is skipped.
Project Structure
hummingbird-dashboard/
├── Containerfile
├── alembic.ini
├── alembic/
│ ├── env.py # Alembic environment (DB URL, include_object guard)
│ ├── script.py.mako # Migration template
│ └── versions/ # Migration files
├── dev.sh
├── hummingbird_dashboard/
│ ├── analysis.py # LLM failure analysis (Gemini)
│ ├── cli.py # CLI entry point
│ ├── db/ # SQLAlchemy engine/session + ORM models
│ │ ├── __init__.py # get_db_url(), get_engine(), get_sessionmaker(), upgrade_head()
│ │ └── models.py
│ ├── konflux.py # Konflux API client
│ ├── models.py # Data models (Component, PushEvent)
│ ├── sources.py # PostgreSQL queries
│ ├── table.py # CLI table formatting
│ ├── triggers.py # Application-specific trigger rules
│ ├── views.py # Status computation and aggregation
│ └── web/
│ ├── app.py # FastAPI application
│ └── templates/ # Jinja2 templates
└── tests/
Building Container Image
cd hummingbird-dashboard
podman build -f Containerfile -t hummingbird-dashboard .
Database Migrations
The dashboard uses Alembic to manage its schema.
Migrations run automatically on every startup via alembic upgrade head.
On first startup against an existing environment, the dashboard detects that the
schema is already present (via the blocked_snapshots sentinel table) and
automatically stamps the baseline migration as applied before upgrading. No
manual intervention is required.
Local development — migrations run automatically via ./dev.sh db-seed:
./dev.sh db-start
./dev.sh db-seed # runs alembic upgrade head then seeds test data
./dev.sh start
License
This project is licensed under the GNU General Public License v3.0 or later - see the LICENSE file for details.