AI
Robin, Heron and Raven: three models built to execute, made to operate with Birdhouse OS.
The flock
Our smallest model
Quick, light, everywhere.
Built for the high volume, low latency decisions that make up most of an agent's day. Small enough to run close to your data, cheap enough to call on every event.
Best for
Our general execution model
Patient and precise.
The everyday workhorse. It writes and reviews code, drafts and summarises, and follows a structured task through several steps, without the cost of reaching for the largest model.
Best for
Our smartest model
The one that works it out.
For the hardest work: writing complex code, finding complex security vulnerabilities, and multi-step problems where getting it right matters more than getting it fast.
Best for
Benchmarks
How each model scores across the work it does.
| Benchmark (SWE-bench) | Heron v0.6 |
|---|---|
| Coding | 87.0 |
| Reasoning | 68.2 |
| Tool use | 79.6 |
| Long context | 84.1 |
| Task Execution | 89.1 |
How they work together
Birdhouse OS sends each task to the lightest model able to handle it, and moves up the family only when the work calls for it. Most of the load lands on Robin and Heron, Raven takes the hard problems, and a third party model is the last resort rather than the default.