For DevRel and DX teams
milo turns your spec, CLI, MCP server and docs into real developer tasks, runs coding agents on them against your sandbox, checks what actually happened, and tells you which file to change.
| Agent gets | Task success | Lift vs spec |
|---|---|---|
| Spec only | 33% | – |
| Your current kit | 20% | −13 pts |
| With milo's fixes | 62% | +29 pts |
Illustrative example. Real reports show 95% intervals and every run behind each number.
What you bring
No rewrite and no new platform. milo reads the artifacts your developers already use and gives you a lint report in minutes, before any agent runs.
However complete they are. Thin, generated specs are where agents struggle most.
Crawled from their help. milo flags prompts, missing JSON output and commands that hang without a terminal.
milo reads your tool list and checks descriptions, schemas and the token cost agents pay on every turn.
Checked for broken links and used to write tasks that match how developers really use your API.
Already running Claude Code plugin evals? milo adds outcome checks to the suite you have.
Bring a test account. Each run gets its own isolated workspace, and every request is logged and checked.
How it works
milo reads your spec, CLI, MCP server and docs and lists what trips agents up, each with a fix. No model calls.
milo drafts real developer tasks from your API and you approve them. A fifth are held out, so later fixes are judged on tasks they weren't tuned for.
Agents run the tasks through Claude Code's plugin evals, each with its own sandbox account. milo then checks the account: the records, the webhooks, the requests.
Failures are traced to a file with a patch, and each patch is re-run before you see it. In CI, milo fails the PR if agent success drops.
Claude Code
Verify Acme webhooks and record paid invoices
Graded on outcomes
The same tasks run with only your spec, with your current kit and with milo's fixes, over repeated trials, so a better number means a better kit.
State, webhooks and requests checked
Outcomes
Lift, with confidence intervals
Lift
Reliability across repeated trials
pass^k
Whether agents use your kit at all
Activation
Fixes, not just scores
milo reads every failed run and traces it to your spec, CLI, MCP server, skill or docs. Each patch is a diff against your files, re-run on the affected tasks before you see it. Nothing is applied for you.
1$ milo diagnose2FAIL csv-import 2 of 3 runs left duplicate customers34 cause skill skills/acme/SKILL.md never loaded:5 its description doesn't mention imports6 patch skill + Use before writing code that imports,7 syncs or bulk-creates Acme records8 check re-ran the 3 affected tasks: 1/3 → 3/39101 patch helped, 0 dropped. See milo/patches/.About milo
Lorem ipsum dolor sit amet, consectetur adipiscing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua. Ut enim ad minim veniam, quis nostrud exercitation ullamco laboris nisi ut aliquip ex ea commodo consequat.
Bring your spec and a test account.