# Hummingbird AI

LLMS index: [llms.txt](https://hummingbird-project.io/llms.txt) | Full content: [llms-full.txt](https://hummingbird-project.io/llms-full.txt)

---

`hummingbird-ai` is a synchronous Python library used in-process by Hummingbird
agents to call Gemini and Anthropic models through normalized interfaces.

## Features

- Shared Gemini API/Vertex AI and Anthropic Vertex AI adapters.
- Model construction, regional selection, token pricing, and usage accounting.
- Retryable model-call handling and Prometheus request/token/cost metrics.
- HTTP sessions and Google Application Default Credentials authentication.

The library is not a service and has no deployment of its own. Model adapter
instances are scoped to individual conversations. Callers provide bounded
concurrency for independent work.

## Model interface

`hummingbird_ai.models` defines the `ModelAdapter` protocol and normalized
`ToolDef`, `ToolCall`, `Usage`, `ModelResponse`, and `ModelError` types. Adapters
accept provider-neutral tool definitions and return normalized responses.
Their `make_user_content()` and `make_tool_responses()` methods create
provider-specific conversation content; callers keep `contents` opaque.

Adapters also implement `to_canonical()` and `from_canonical()` for persisted
conversation history. The canonical message format lets an agent resume a
session with a different provider without storing either provider's native
wire format.

## Provider adapters

### Gemini

`GeminiModel` supports the direct Gemini API with an API key and Vertex AI with
Google credentials. It translates the normalized tool and response types to
Gemini REST JSON. Gemini's implicit prompt cache is reflected in
`Usage.cache_read_tokens` from `cachedContentTokenCount`.

### Anthropic on Vertex AI

`AnthropicVertexModel` uses Vertex AI `rawPredict`. It handles Anthropic's
assistant/user turn alternation and associates tool results with the
conversation-local tool-use IDs. For prompt caching it adds sliding
`cache_control` breakpoints to shallow copies of messages; the original
conversation history is not mutated. Transient messages can use `ephemeral`
so warning text is excluded from the cache write prefix.

When Anthropic web search is enabled, the adapter adds the server-side search
tool. Search blocks remain opaque to the agent tool loop and survive canonical
session serialization. The adapter also follows `pause_turn` continuations and
merges content and usage from those provider-side calls into the returned
response.

## Provider wire formats

The agent uses normalized types; adapters translate them to provider payloads:

| Normalized data    | Gemini                                  | Anthropic on Vertex AI                      |
| ------------------ | --------------------------------------- | ------------------------------------------- |
| System instruction | `systemInstruction`                     | top-level `system` string                   |
| Tool definitions   | `functionDeclarations` with JSON Schema | `tools` with `input_schema`                 |
| Tool calls         | `functionCall` parts                    | `tool_use` content blocks                   |
| Tool results       | `functionResponse` parts                | `tool_result` content blocks in a user turn |

Gemini API-key requests use the direct Gemini endpoint; Vertex requests use
Google authentication and a configured region. Anthropic requests use Vertex
AI `rawPredict`. The adapters keep provider-specific message normalization
inside `generate()`; the agent appends the returned `raw_content` without
interpreting it.

## Canonical sessions

Adapters convert their native conversation content to and from the shared
canonical message format at persistence boundaries. Anthropic thinking and
server-tool blocks are preserved as opaque pass-through fields and restored in
their original order. This allows sessions saved by one provider to resume on
another while keeping the agent loop independent of provider payload shapes.
Extended thinking blocks in assistant responses are part of the cached prefix.
They do not break cache hits when the following user message contains only tool_result blocks.

## Prompt caching

Gemini Vertex AI prompt caching is implicit. The adapter reads
`cachedContentTokenCount` from usage metadata and reports it as
`cache_read_tokens`; it adds no cache annotations to requests.

Anthropic Vertex caching uses explicit `cache_control: {"type": "ephemeral"}`
annotations. The adapter uses two of the four available breakpoints on
messages: B2 writes the latest prefix and B1 reads the previous prefix. The
prefix hash includes prior tools and system instructions, so extra breakpoints
on those components are redundant. Breakpoints are added to shallow copies so
the agent-owned history is unchanged. Ephemeral turns keep warning text out of
the cache: B1 still reads the last stable prefix, while B2 and the tracked write
position advance only to the last stable message before the transient tail. If
there is no stable message to annotate, no breakpoint is moved.

Vertex AI Anthropic prompt caches have a five-minute TTL refreshed on a hit.
Cache writes use 1.25 times the base input-token rate and reads use 0.10 times
that rate. Minimum cacheable prefixes are 1,024 tokens for Sonnet/Opus and
4,096 for Haiku. A newly constructed adapter starts without a prior breakpoint
position, so the first resumed call writes a fresh prefix before later calls
can read it.

```text
T1:     U1·B2                                       prev=0
T2:     U1·B1  M1  U2·B2                            prev=2
T3(E):  U1  M1  U2·B1  M2·B2  U3+E                  prev=3
T4(E):  U1  M1  U2  M2·B1  U3  M3·B2  U4+E          prev=5
T5:     U1  M1  U2  M2  U3  M3·B1  U4  M4  U5·B2    prev=8
```

## Calls, retries, and accounting

`hummingbird_ai.calls.generate_with_retry()` owns model-call retry and usage
accounting. A caller selects either an elapsed-time retry budget or a maximum
attempt count. It can also pass response validation into the same attempt loop,
so validation failures do not multiply a caller's transport retry limit.
Provider-specific HTTP status handling remains in each adapter.

Each returned `ModelResponse` contributes its input, output, cache-read, and
cache-creation tokens and estimated cost exactly once, even when caller-side
validation rejects that response. The library emits shared Vertex request,
token, and cost counters while callers retain workflow-specific error handling
and summaries.

Pricing is kept in `models.MODEL_PRICING` as per-million-token rates for input,
output, cache reads, and cache creation. `models.estimate_cost()` returns an
estimated USD cost or `None` for an unpriced model.

## HTTP and authentication

`hummingbird_ai.http.new_session()` creates requests sessions with configurable
transport retries. Model adapters disable transport-level retries for 429 so
that rate-limit failures use the model-call retry policy. `VertexAuth` uses
Google Application Default Credentials and resolves the project from an
explicit argument, `GOOGLE_CLOUD_PROJECT`, then ADC.

`requests.Timeout` is caught before `requests.ConnectionError` in adapters
because `ConnectTimeout` is a subclass of both. The urllib3 retry adapter
does not retry `POST` requests (non-idempotent), so all transport errors
from model calls propagate to the `generate_with_retry` retry layer.

## Prerequisites

- Python 3.11 or later.
- Google credentials and a project for Vertex AI models; Gemini API key mode is
  also supported.

## Installation

The library is installed as an in-repository dependency of both agents:

```bash
make hummingbird-agent/setup
make hummingbird-cve-agent/setup
```

## Usage

```python
from hummingbird_ai import models

model = models.build_model(
    "gemini-2.5-flash",
    google_api_key="",
    google_cloud_project="my-project",
    model_regions={"gemini-2.5-": "us-central1"},
)
```

Shared calls and usage accounting are provided by `hummingbird_ai.calls`.
Agent-specific prompts, tools, response parsing, and decisions remain in the
consuming agent.

## Configuration

| Setting                | Use                                                                  |
| ---------------------- | -------------------------------------------------------------------- |
| `GOOGLE_API_KEY`       | Passed by the caller for direct Gemini API access                    |
| `GOOGLE_CLOUD_PROJECT` | Vertex AI project fallback when credentials do not specify a project |

Model regions and Vertex labels are passed by the consuming agent when it
constructs an adapter.

## Development

See the main [README][readme] for development workflows.

```bash
make check                         # lint
make test                          # run tests
```

## License

This project is licensed under the GNU General Public License v3.0 or later -
see the [LICENSE][license] file for details.

[readme]: https://gitlab.com/redhat/hummingbird/tools/-/blob/main/README.md
[license]: https://gitlab.com/redhat/hummingbird/tools/-/blob/main/LICENSE
