Message Bus Architecture

The Hummingbird message bus is an event-driven architecture built on AWS SNS. Events from multiple sources flow through a central topic, enabling subscribers to filter and process only the events they need.

Architecture

flowchart TD
    subgraph Publishers
        GL[GitLab Webhooks]
        K8S[Kubernetes Clusters]
    end

    subgraph MessageBus [Message Bus]
        SNS[(SNS Topic)]
    end

    subgraph Archiver
        ARCH[SNS S3 Archiver]
        S3[(S3 Bucket)]
    end

    subgraph StatusDB [Status Database]
        SQS1[SQS Queue]
        WORKER[hummingbird-status]
        PG[(PostgreSQL)]
    end

    subgraph ConsumerQueue [Consumer Queue]
        SQS2[SQS Queue]
    end

    subgraph Catalog [Container Catalog]
        SYNC[SyncFunction]
        DDB[(DynamoDB)]
    end

    subgraph Consumers
        CONSUMER[Event Consumer]
    end

    GL -->|gitlab-event-forwarder| SNS
    K8S -->|kubernetes-event-forwarder| SNS
    SNS --> ARCH
    ARCH --> S3
    SNS --> SQS1
    SQS1 --> WORKER
    WORKER --> PG
    SNS -->|"FilterPolicy\nkind=Release"| SYNC
    SYNC -->|fetch manifests| Registry[(OCI Registry)]
    SYNC -->|read/write| DDB
    SNS --> SQS2
    SQS2 -->|new events| CONSUMER
    S3 -.->|historic events| CONSUMER
    PG -.->|aggregated status| CONSUMER

Components

Component Role Description Staging
hummingbird-events-topic Infrastructure Central SNS topic for all events Yes
gitlab-event-forwarder Publisher Receives GitLab webhooks, publishes to SNS Yes
kubernetes-event-forwarder Publisher Watches K8s resources, publishes changes to SNS Yes
sns-s3-archiver Subscriber Archives all events to S3 for querying/replay Yes
hummingbird-status Subscriber Ingests events to PostgreSQL for structured queries Yes
pipeline-timing Subscriber Stores pipeline timing in its own schema and queue No
hummingbird-agent Subscriber Processes events to drive automated workflows Yes
workqueue-service Subscriber Queues and dispatches work items from events Yes
container-catalog Subscriber Incrementally syncs image metadata to DynamoDB on Release events No

For subscriber rows, the Staging column describes running consumers. Pipeline Timing queues and subscriptions can be provisioned in both environments before the workers run; their presence does not mark the service as deployed.

Message Format

All messages include standard attributes for filtering:

Attribute Description Examples
source Event origin gitlab, kubernetes
event_type Type of event push, merge_request, pipeline, build, MODIFIED

Additional attributes vary by source - see individual publisher docs for details.

GitLab Events

Published by gitlab-event-forwarder:

Attribute Description Example
project_path Full project path redhat/hummingbird/containers
group_path Full group path redhat/hummingbird

When configured, Pipeline Hooks use event_type: pipeline; Job Hooks use event_type: build.

Kubernetes Events

Published by kubernetes-event-forwarder:

Attribute Description Example
cluster Logical kubeconfig cluster entry name konflux-rh03
namespace Object namespace production
kind Resource kind PipelineRun, TaskRun, Snapshot, Release
api_version API version tekton.dev/v1, appstudio.redhat.com/v1alpha1
object_name Resource name nginx-7d8c4c9d6f

Published kinds depend on the deployed forwarder configuration. TaskRun source activation is separate rollout work. The forwarder obtains cluster from the selected kubeconfig context’s context.cluster value, not the API server URL.

Subscription Filtering

SNS filter policies enable subscribers to receive only relevant events:

{
  "source": ["gitlab"],
  "event_type": ["push", "merge_request"]
}
{
  "source": ["kubernetes"],
  "kind": ["Deployment"],
  "event_type": ["MODIFIED"]
}

Kubernetes Release events (used by container-catalog sync Lambda):

{
  "kind": ["Release"]
}

See hummingbird-events-topic for subscription setup instructions.

Event Flow Example

  1. Developer pushes to GitLab repository
  2. GitLab sends webhook to gitlab-event-forwarder Lambda
  3. Lambda validates token, extracts metadata, publishes to SNS
  4. SNS delivers to all matching subscribers:
    • sns-s3-archiver stores event in S3
    • Event consumers process new events in real-time

Archive keys

The S3 archiver writes one deterministic key for each SNS message. Duplicate delivery rewrites the same object. GitLab pipeline and job identities include the project and their numeric IDs; all Kubernetes identities include logical cluster, namespace, and name. URL-valued cluster attributes from older forwarders retain their name-only archive identities.

See SNS S3 Archiver for the exact key format and the machine-readable archive-key contract used by the archiver and GitLab forwarder tests.

Consuming Events

Consumers can receive events from multiple sources:

  • New events: Subscribe to SNS topic for real-time processing
  • Historic events: Download from S3 archive for replay or catch-up
  • Structured queries: Query PostgreSQL via hummingbird-status for pipeline status, component state, and cross-referenced data

The S3 archive enables consumers to bootstrap state, then switch to live SNS events. Hummingbird Status provides a queryable view of its current pipeline status data. Timeline ingestion for GitLab pipelines/jobs and TaskRuns is separate follow-up work.

Replaying Events

The sns-s3-archiver stores complete SNS records with decoded payloads, enabling easy replay:

# Download archived events
aws s3 sync s3://bucket-name/2025/12/13/ ./local-events/

# Browse events
zcat ./local-events/2025/12/13/12/34/*.json.gz | jq
import gzip
import json
from pathlib import Path

# Replay archived events to handler
for path in sorted(Path("./local-events").rglob("*.json.gz")):
    with gzip.open(path, "rt") as f:
        sns_record = json.load(f)
    event = {"Records": [{"Sns": sns_record}]}
    my_handler.lambda_handler(event, None)

See sns-s3-archiver documentation for details on storage format and the _decode_message pattern for handlers.

Staging Environment

Production and staging use independent SNS topics. Each environment has its own publishers, archiver, and consumer queues so that changes can be tested in staging before production deployment.

flowchart TD
    GL["GitLab Projects\n(dual webhook delivery)"]
    K8S["Konflux Cluster\n(K8s resources)"]

    subgraph production ["Production"]
        GEF_P["gitlab-event-forwarder\n(Lambda)"]
        KEF_P["kubernetes-event-forwarder\n(K8s pod)"]
        SNS_P["SNS: arr-hummingbird-prod-events"]
        prodConsumers["consumers\n(status, agent, workqueue,\narchiver, etc.)"]

        GEF_P --> SNS_P
        KEF_P --> SNS_P
        SNS_P --> prodConsumers
    end

    subgraph staging ["Staging"]
        GEF_S["gitlab-event-forwarder\n(Lambda)"]
        KEF_S["kubernetes-event-forwarder\n(K8s pod)"]
        SNS_S["SNS: arr-hummingbird-staging-events"]
        stagConsumers["consumers\n(status, agent, workqueue,\narchiver)"]

        GEF_S --> SNS_S
        KEF_S --> SNS_S
        SNS_S --> stagConsumers
    end

    GL --> GEF_P
    GL --> GEF_S
    K8S --> KEF_P
    K8S --> KEF_S

How it works

  • Duplicate webhook delivery: GitLab sends every event to both gitlab-event-forwarder.hummingbird-project.io (prod) and gitlab-event-forwarder.staging.hummingbird-project.io (staging). Each forwarder publishes to its own SNS topic.
  • Dual Kubernetes watches: Both prod and staging kubernetes-event-forwarder pods watch the same Konflux cluster resources and publish to their respective topics.
  • Consumer isolation: Staging consumers (hummingbird-status, hummingbird-agent, workqueue-service) have separate SQS queues subscribed to the staging topic, with dedicated Postgres databases.
  • Staging archiver: A separate sns-s3-archiver subscribes to the staging topic and archives to its own S3 bucket with a shorter retention period than production.