Announcement
New Models

Claude Sonnet 5 + Haiku 4.5

July 27, 2026
Thesean

The same product, two new grades: ship-like/claude-sonnet-5 and ship-like/claude-haiku-4-5 match their reference models' quality and behavior at a fraction of the cost — guaranteed at least half, backed by our quality SLA.

CONTENTS

Ship gives you the highest intelligence per dollar of any frontier model by finding the cheapest execution that still meets a model's quality grade. We first introduced it for Claude Opus 4.8 and GPT 5.6 Sol. Today we're adding Claude Sonnet 5 and Claude Haiku 4.5 with the same quality SLA, the same cost reduction, and the same ease of use.

Equivalent Performance

Base model

A fraction of the cost

Just name a reference model as a quality grade, and we fulfill each request in the cheapest way that still meets that grade. We measure both Capability Equivalence and Behavioral Equivalence to the reference model, so that the model performs as well as expected, and it works the same way within your harness when you switch to ship.

TL;DR. Change model="claude-sonnet-5" to model="ship-like/claude-sonnet-5". Same answers, same behavior, for at least 50% less (in our evals 30% for Sonnet 5 and ~40% for Haiku 4.5) with a quality SLA.

New to Ship? Start with Introducing Ship.

Capability equivalence

Across every public benchmark we tested, Ship lands inside its reference model's margin of error — at a fraction of the cost per task. Use the selector to switch between Sonnet 5 and Haiku 4.5.

Behavioral equivalence

A drop-in replacement has to act like the model it replaces, or you'd have to change your harness (prompts, parsers, tools, etc.). The selected model's Ship grade has a high distributional overlap with it across nine behavioral axes — 98% for Sonnet 5 (tool calls 20.0 → 19.3, steps 20.4 → 20.2) and 97% for Haiku 4.5 (18.6 → 18.5, 20.0 → 19.9).

Behavioral fingerprint

Base model

normalized axes · Ship dashed against the selected base

Zoom into tool use and the two grades reach for the same tools in the same proportions — so a prompt tuned to the selected model's tool habits behaves the same on Ship. The time budget splits the same way, too — most of it in model generation, the rest in tool execution — so latency profiles match, not just totals.

Tool-call mix (average)

Base model

Where the time goes (average)

model generation vs tool execution

Step back to the whole trace and the equivalence is holistic: embed each run and Ship lands inside the selected model's cloud, while a model with similar benchmark performance sits clearly apart.

Ship traces sit inside the base cloud

Base model

task-centered embeddings · t-SNE

And it's task-for-task, not just on average. Bucket every task by how often the selected model solves it, and Ship tracks it bucket-for-bucket far more tightly than any other model — the same hard tasks stay hard, the same easy ones stay easy.

Other models within each Claude Sonnet 5 solve-rate bucket

Base model

tasks sorted by base solve rate · color = solve rate

Spend the savings on quality

The savings don't have to go into your pocket. Because Ship delivers the same grade for less, the budget you free up can be reinvested into more reasoning — pushing quality past where your budget used to cap it.

Max-reasoning quality, at a lower-reasoning price — ARC-AGI-2

Claude Sonnet 5 · low → medium → high → max · lower-left is cheaper

Get started

Both new ship models are available now. Swap one string:

model="ship-like/claude-sonnet-5"  ·  model="ship-like/claude-haiku-4-5"

Bring your private evals or production traffic and we'll compare Ship against your current model directly. Get an API key · Talk to us about an SLA.

Thesean