Early access

The eval layer
for your skills.

A skill that works in one AI coding harness doesn't always work in another — and model updates change behavior without warning. Write your evals once, run them across the harnesses you use, and see exactly what broke.

Launch updates only. No spam.

Your skill works. Until it doesn't.

Harness drift

A skill written for one agent runtime behaves differently in the next. Tool names, output formats, system prompts — everything shifts.

Model drift

The model behind your harness updates, and yesterday's reliable skill starts failing. Silently, with no alert and no changelog you can act on.

No signal

Without evals you find out from your users. With evals tied to a single harness, you only ever see half the picture.

Author. Run. Score. Improve.

  1. 1

    Author

    Write evals in code next to your skills: tasks, assertions, graders. Plain files, versioned in your repo.

  2. 2

    Run

    Execute the same suite across each harness and model you target — on your machine or in CI.

  3. 3

    Scorecard

    Compare results side by side. See which harness, which model, and which skill regressed — down to the failing case.

  4. 4

    Improve

    Fix the skill, re-run, watch the delta. Ship with proof it works everywhere you care about.

Built for people shipping agent skills

Skill authors

Keep prompts and skills reliable as harnesses and models move under your feet. Know the moment a change breaks something.

Agent teams

Gate skill changes in CI the way you gate code. Review a scorecard before anything reaches production users.

Be first in line.

We're building the eval layer for agent skills. Early access goes to the waitlist.

Request access