Hummingbird AI

hummingbird-ai is a synchronous Python library used in-process by Hummingbird agents to call Gemini and Anthropic models through normalized interfaces.

Features

  • Shared Gemini API/Vertex AI and Anthropic Vertex AI adapters.
  • Model construction, regional selection, token pricing, and usage accounting.
  • Retryable model-call handling and Prometheus request/token/cost metrics.
  • HTTP sessions and Google Application Default Credentials authentication.

The library is not a service and has no deployment of its own. Model adapter instances are scoped to individual conversations. Callers provide bounded concurrency for independent work.

Model interface

hummingbird_ai.models defines the ModelAdapter protocol and normalized ToolDef, ToolCall, Usage, ModelResponse, and ModelError types. Adapters accept provider-neutral tool definitions and return normalized responses. Their make_user_content() and make_tool_responses() methods create provider-specific conversation content; callers keep contents opaque.

Adapters also implement to_canonical() and from_canonical() for persisted conversation history. The canonical message format lets an agent resume a session with a different provider without storing either provider’s native wire format.

Provider adapters

Gemini

GeminiModel supports the direct Gemini API with an API key and Vertex AI with Google credentials. It translates the normalized tool and response types to Gemini REST JSON. Gemini’s implicit prompt cache is reflected in Usage.cache_read_tokens from cachedContentTokenCount.

Anthropic on Vertex AI

AnthropicVertexModel uses Vertex AI rawPredict. It handles Anthropic’s assistant/user turn alternation and associates tool results with the conversation-local tool-use IDs. For prompt caching it adds sliding cache_control breakpoints to shallow copies of messages; the original conversation history is not mutated. Transient messages can use ephemeral so warning text is excluded from the cache write prefix.

When Anthropic web search is enabled, the adapter adds the server-side search tool. Search blocks remain opaque to the agent tool loop and survive canonical session serialization. The adapter also follows pause_turn continuations and merges content and usage from those provider-side calls into the returned response.

Provider wire formats

The agent uses normalized types; adapters translate them to provider payloads:

Normalized data Gemini Anthropic on Vertex AI
System instruction systemInstruction top-level system string
Tool definitions functionDeclarations with JSON Schema tools with input_schema
Tool calls functionCall parts tool_use content blocks
Tool results functionResponse parts tool_result content blocks in a user turn

Gemini API-key requests use the direct Gemini endpoint; Vertex requests use Google authentication and a configured region. Anthropic requests use Vertex AI rawPredict. The adapters keep provider-specific message normalization inside generate(); the agent appends the returned raw_content without interpreting it.

Canonical sessions

Adapters convert their native conversation content to and from the shared canonical message format at persistence boundaries. Anthropic thinking and server-tool blocks are preserved as opaque pass-through fields and restored in their original order. This allows sessions saved by one provider to resume on another while keeping the agent loop independent of provider payload shapes. Extended thinking blocks in assistant responses are part of the cached prefix. They do not break cache hits when the following user message contains only tool_result blocks.

Prompt caching

Gemini Vertex AI prompt caching is implicit. The adapter reads cachedContentTokenCount from usage metadata and reports it as cache_read_tokens; it adds no cache annotations to requests.

Anthropic Vertex caching uses explicit cache_control: {"type": "ephemeral"} annotations. The adapter uses two of the four available breakpoints on messages: B2 writes the latest prefix and B1 reads the previous prefix. The prefix hash includes prior tools and system instructions, so extra breakpoints on those components are redundant. Breakpoints are added to shallow copies so the agent-owned history is unchanged. Ephemeral turns keep warning text out of the cache: B1 still reads the last stable prefix, while B2 and the tracked write position advance only to the last stable message before the transient tail. If there is no stable message to annotate, no breakpoint is moved.

Vertex AI Anthropic prompt caches have a five-minute TTL refreshed on a hit. Cache writes use 1.25 times the base input-token rate and reads use 0.10 times that rate. Minimum cacheable prefixes are 1,024 tokens for Sonnet/Opus and 4,096 for Haiku. A newly constructed adapter starts without a prior breakpoint position, so the first resumed call writes a fresh prefix before later calls can read it.

T1:     U1·B2                                       prev=0
T2:     U1·B1  M1  U2·B2                            prev=2
T3(E):  U1  M1  U2·B1  M2·B2  U3+E                  prev=3
T4(E):  U1  M1  U2  M2·B1  U3  M3·B2  U4+E          prev=5
T5:     U1  M1  U2  M2  U3  M3·B1  U4  M4  U5·B2    prev=8

Calls, retries, and accounting

hummingbird_ai.calls.generate_with_retry() owns model-call retry and usage accounting. A caller selects either an elapsed-time retry budget or a maximum attempt count. It can also pass response validation into the same attempt loop, so validation failures do not multiply a caller’s transport retry limit. Provider-specific HTTP status handling remains in each adapter.

Each returned ModelResponse contributes its input, output, cache-read, and cache-creation tokens and estimated cost exactly once, even when caller-side validation rejects that response. The library emits shared Vertex request, token, and cost counters while callers retain workflow-specific error handling and summaries.

Pricing is kept in models.MODEL_PRICING as per-million-token rates for input, output, cache reads, and cache creation. models.estimate_cost() returns an estimated USD cost or None for an unpriced model.

HTTP and authentication

hummingbird_ai.http.new_session() creates requests sessions with configurable transport retries. Model adapters disable transport-level retries for 429 so that rate-limit failures use the model-call retry policy. VertexAuth uses Google Application Default Credentials and resolves the project from an explicit argument, GOOGLE_CLOUD_PROJECT, then ADC.

requests.Timeout is caught before requests.ConnectionError in adapters because ConnectTimeout is a subclass of both. The urllib3 retry adapter does not retry POST requests (non-idempotent), so all transport errors from model calls propagate to the generate_with_retry retry layer.

Prerequisites

  • Python 3.11 or later.
  • Google credentials and a project for Vertex AI models; Gemini API key mode is also supported.

Installation

The library is installed as an in-repository dependency of both agents:

make hummingbird-agent/setup
make hummingbird-cve-agent/setup

Usage

from hummingbird_ai import models

model = models.build_model(
    "gemini-2.5-flash",
    google_api_key="",
    google_cloud_project="my-project",
    model_regions={"gemini-2.5-": "us-central1"},
)

Shared calls and usage accounting are provided by hummingbird_ai.calls. Agent-specific prompts, tools, response parsing, and decisions remain in the consuming agent.

Configuration

Setting Use
GOOGLE_API_KEY Passed by the caller for direct Gemini API access
GOOGLE_CLOUD_PROJECT Vertex AI project fallback when credentials do not specify a project

Model regions and Vertex labels are passed by the consuming agent when it constructs an adapter.

Development

See the main README for development workflows.

make check                         # lint
make test                          # run tests

License

This project is licensed under the GNU General Public License v3.0 or later - see the LICENSE file for details.