Hummingbird AI
hummingbird-ai is a synchronous Python library used in-process by Hummingbird
agents to call Gemini and Anthropic models through normalized interfaces.
Features
- Shared Gemini API/Vertex AI and Anthropic Vertex AI adapters.
- Model construction, regional selection, token pricing, and usage accounting.
- Retryable model-call handling and Prometheus request/token/cost metrics.
- HTTP sessions and Google Application Default Credentials authentication.
The library is not a service and has no deployment of its own. Model adapter instances are scoped to individual conversations. Callers provide bounded concurrency for independent work.
Model interface
hummingbird_ai.models defines the ModelAdapter protocol and normalized
ToolDef, ToolCall, Usage, ModelResponse, and ModelError types. Adapters
accept provider-neutral tool definitions and return normalized responses.
Their make_user_content() and make_tool_responses() methods create
provider-specific conversation content; callers keep contents opaque.
Adapters also implement to_canonical() and from_canonical() for persisted
conversation history. The canonical message format lets an agent resume a
session with a different provider without storing either provider’s native
wire format.
Provider adapters
Gemini
GeminiModel supports the direct Gemini API with an API key and Vertex AI with
Google credentials. It translates the normalized tool and response types to
Gemini REST JSON. Gemini’s implicit prompt cache is reflected in
Usage.cache_read_tokens from cachedContentTokenCount.
Anthropic on Vertex AI
AnthropicVertexModel uses Vertex AI rawPredict. It handles Anthropic’s
assistant/user turn alternation and associates tool results with the
conversation-local tool-use IDs. For prompt caching it adds sliding
cache_control breakpoints to shallow copies of messages; the original
conversation history is not mutated. Transient messages can use ephemeral
so warning text is excluded from the cache write prefix.
When Anthropic web search is enabled, the adapter adds the server-side search
tool. Search blocks remain opaque to the agent tool loop and survive canonical
session serialization. The adapter also follows pause_turn continuations and
merges content and usage from those provider-side calls into the returned
response.
Provider wire formats
The agent uses normalized types; adapters translate them to provider payloads:
| Normalized data | Gemini | Anthropic on Vertex AI |
|---|---|---|
| System instruction | systemInstruction |
top-level system string |
| Tool definitions | functionDeclarations with JSON Schema |
tools with input_schema |
| Tool calls | functionCall parts |
tool_use content blocks |
| Tool results | functionResponse parts |
tool_result content blocks in a user turn |
Gemini API-key requests use the direct Gemini endpoint; Vertex requests use
Google authentication and a configured region. Anthropic requests use Vertex
AI rawPredict. The adapters keep provider-specific message normalization
inside generate(); the agent appends the returned raw_content without
interpreting it.
Canonical sessions
Adapters convert their native conversation content to and from the shared canonical message format at persistence boundaries. Anthropic thinking and server-tool blocks are preserved as opaque pass-through fields and restored in their original order. This allows sessions saved by one provider to resume on another while keeping the agent loop independent of provider payload shapes. Extended thinking blocks in assistant responses are part of the cached prefix. They do not break cache hits when the following user message contains only tool_result blocks.
Prompt caching
Gemini Vertex AI prompt caching is implicit. The adapter reads
cachedContentTokenCount from usage metadata and reports it as
cache_read_tokens; it adds no cache annotations to requests.
Anthropic Vertex caching uses explicit cache_control: {"type": "ephemeral"}
annotations. The adapter uses two of the four available breakpoints on
messages: B2 writes the latest prefix and B1 reads the previous prefix. The
prefix hash includes prior tools and system instructions, so extra breakpoints
on those components are redundant. Breakpoints are added to shallow copies so
the agent-owned history is unchanged. Ephemeral turns keep warning text out of
the cache: B1 still reads the last stable prefix, while B2 and the tracked write
position advance only to the last stable message before the transient tail. If
there is no stable message to annotate, no breakpoint is moved.
Vertex AI Anthropic prompt caches have a five-minute TTL refreshed on a hit. Cache writes use 1.25 times the base input-token rate and reads use 0.10 times that rate. Minimum cacheable prefixes are 1,024 tokens for Sonnet/Opus and 4,096 for Haiku. A newly constructed adapter starts without a prior breakpoint position, so the first resumed call writes a fresh prefix before later calls can read it.
T1: U1·B2 prev=0
T2: U1·B1 M1 U2·B2 prev=2
T3(E): U1 M1 U2·B1 M2·B2 U3+E prev=3
T4(E): U1 M1 U2 M2·B1 U3 M3·B2 U4+E prev=5
T5: U1 M1 U2 M2 U3 M3·B1 U4 M4 U5·B2 prev=8
Calls, retries, and accounting
hummingbird_ai.calls.generate_with_retry() owns model-call retry and usage
accounting. A caller selects either an elapsed-time retry budget or a maximum
attempt count. It can also pass response validation into the same attempt loop,
so validation failures do not multiply a caller’s transport retry limit.
Provider-specific HTTP status handling remains in each adapter.
Each returned ModelResponse contributes its input, output, cache-read, and
cache-creation tokens and estimated cost exactly once, even when caller-side
validation rejects that response. The library emits shared Vertex request,
token, and cost counters while callers retain workflow-specific error handling
and summaries.
Pricing is kept in models.MODEL_PRICING as per-million-token rates for input,
output, cache reads, and cache creation. models.estimate_cost() returns an
estimated USD cost or None for an unpriced model.
HTTP and authentication
hummingbird_ai.http.new_session() creates requests sessions with configurable
transport retries. Model adapters disable transport-level retries for 429 so
that rate-limit failures use the model-call retry policy. VertexAuth uses
Google Application Default Credentials and resolves the project from an
explicit argument, GOOGLE_CLOUD_PROJECT, then ADC.
requests.Timeout is caught before requests.ConnectionError in adapters
because ConnectTimeout is a subclass of both. The urllib3 retry adapter
does not retry POST requests (non-idempotent), so all transport errors
from model calls propagate to the generate_with_retry retry layer.
Prerequisites
- Python 3.11 or later.
- Google credentials and a project for Vertex AI models; Gemini API key mode is also supported.
Installation
The library is installed as an in-repository dependency of both agents:
make hummingbird-agent/setup
make hummingbird-cve-agent/setup
Usage
from hummingbird_ai import models
model = models.build_model(
"gemini-2.5-flash",
google_api_key="",
google_cloud_project="my-project",
model_regions={"gemini-2.5-": "us-central1"},
)
Shared calls and usage accounting are provided by hummingbird_ai.calls.
Agent-specific prompts, tools, response parsing, and decisions remain in the
consuming agent.
Configuration
| Setting | Use |
|---|---|
GOOGLE_API_KEY |
Passed by the caller for direct Gemini API access |
GOOGLE_CLOUD_PROJECT |
Vertex AI project fallback when credentials do not specify a project |
Model regions and Vertex labels are passed by the consuming agent when it constructs an adapter.
Development
See the main README for development workflows.
make check # lint
make test # run tests
License
This project is licensed under the GNU General Public License v3.0 or later - see the LICENSE file for details.