Changelog

What shipped.

The short version of what's new. For the detail behind any release — or what's coming — talk to us.

July 2026

Continuous quality monitoring

ReasonRank now stays on watch between decisions: live quality drift alerts, judge health checks, and dataset freshness — so a verified switch keeps earning its keep.

July 2026

Framework adapters

Connect your existing agent stack with less glue code. Bring your own harness stays the default — adapters just shorten the path to first evidence.

July 2026

Richer scoring for agent work

An expanded scorer library built for agents and skills, beyond plain correctness — so quality means what your team means by quality.

July 2026

Golden datasets

Every verdict now pins the exact cases it was measured on. Re-run the same evidence later and compare apples to apples.

July 2026

Record & replay

Capture real traffic and replay it against candidate models — locally, on your keys — before anything changes in production.

July 2026

Judge calibration

LLM judges are measured against your human labels before their scores count, and re-checked on a schedule afterward.

Spring 2026

Verified savings

The core loop went live for private-beta teams: recommend the cheapest model that holds quality, gate it in CI, then verify the dollars on production traffic.