01 / The generatorMethod 4.4

Ratchet
Loop

Build. Compare. Keep the better version.

Give Grok Build or Claude Code a goal, a real reference, and a check that must hold. Get a complete run kit. Keep the strongest accepted version, not simply the latest change.

Fig. 01 / The championAccept · preserve
Progress is what you keep.A new attempt must earn its place against the accepted version.
  1. Build ×2
  2. Blind duel vs the champion
  3. Win → keep
  4. Lose → preserve champion

Claude Code · Grok Build
No account. Complete prompts.

Brief No 01, 02 and 06 are required
0 chars

Say what the finished thing has to do.

Why, and what works

Don’t name the tools or how to build it — that limits the result to what you already thought of. Write as much as you like: nothing is ever cut, and the box scrolls instead of pushing the page down.

02 · The bar — what does good look like? no references

Give it something real to aim at — a site, a file, a screenshot, or just a sentence. It has to be something the agent can actually open and hold up next to its own work. Empty, it cannot be filed — that is what the second button below is for.

Drop a file here, or pick any of these — they all work from the keyboard too:

Which references work
Typed path
The agent opens it directly. Prefer one relative to your reporefs/ or ./refs/hud.png — because that resolves wherever the run works. An absolute path like C:\refs only opens if the agent is on this machine, and a sandboxed one is not.
Dropped file
Browsers only give us the file name, never the folder. Tell us the folder once and we remember it.
Uploaded
Becomes a web link the run can open from anywhere. Use this if you are sharing with a team.
A web link
We screenshot it for you, desktop and phone, plus the text of the page. Real pictures, or an honest blank if it would not load.
  • Capture · Kind · Reference · StateThe bar is empty. Everything you add shows up here — exactly what the agent will be given, anything wrong with it, and the screenshots where we have them.
03 · Where it runs — the window you paste into, not the model Claude Code
04 · Loop strength — what differs is attempts a round, and what you ship Duel

We pick one for you from your brief. This line always says why, and clicking a plate overrides it.

05 · Caps — optional, and you should no caps

USD, from Claude Code’s estimated session cost. This is a prompt-managed boundary limit, not a hard billing ceiling.

Leave one blank and the prompt says so in words, so the run knows you meant it and cannot invent a limit of its own.

What a cap really does

That has bitten someone: they left this empty expecting no limit, the run read “set a maximum round count”, picked 8, and stopped there. That 8 came out of a cautionary sentence in the prompt that happened to name it, so the number this field carries has to be yours.

Cost These runs can reach hundreds of dollars. Nothing here stops that but you — set a cap you can afford to lose, and keep an eye on it.

What a token limit actually does. It does not squeeze the work into the budget. The run goes at full speed and stops when the money is gone — so you get fewer improvements, not a half-built thing. The prompt tells it to stop before starting a round it cannot afford to finish, and to hand you the best version it reached with a note saying it ran out rather than finished.

06 · What must always be true — required, and worth more than the rest nothing of your own

Everything else the agent checks, it decided on by reading your brief.

How to write one

This one it cannot touch or reword. It runs every round exactly as you wrote it, which makes it the only part of the run that is entirely yours. One sentence is enough.

Plain words. The agent works out how to test it — your sentence stays exactly as written, and the test has to fit it, never the other way round.

Every kit also carries checks for errors, repeatability, keyboard access, and unexpected third-party calls. The binding checks can reject a candidate in Duel and Ratchet. Speed and transfer size are measured against round zero but are advisory: the judges decide whether an increase earns its cost. Gauntlet carries findings to its critic rather than using a gate. A check that cannot run is reported as unrunnable, never as a pass; an inapplicable check needs a written reason.

Selecting a file uploads it to a public URL. Anyone with that URL can read it. Keep credentials and private data out of the file; “sealed” means protected from the builder, not private online. Stored copies become eligible for deletion after 12 hours; downloaded kits keep their own copy.

Nothing uploaded. A Node.js check runs exactly as written. It must print a final JSON line with satisfied: true or false and exit 0 for a valid verdict.

Captures Nothing pending

Web links on the bar are screenshotted before you file — desktop, phone, and the text of the page. Nothing is waiting on one.

Cannot file yet — 01 · The goal, 02 · The bar, 06 · What must always be true still empty. “I don’t know the bar yet” needs only the goal.

Two ways out of this sheet, and the second is not a lesser one. Filing needs 01, 02 and 06. The other writes a different, much shorter prompt: go and find three to five real references, propose the criteria with the measurement that decides each, propose the sealed suite, sketch what it would and would not build — then stop, so you can disagree before the money is spent. Cut what is wrong and bring it back here.

The filed brief

Unfiled


        
        

Nothing filed yet.

Your prompt appears here in full — nothing is ever cut short. It is the record of what was filed, and the kit carries it to disk as run-kit/PROMPT.md, which is what the run reads. It includes your goal, every reference and how to get it, how the loop works, your limits, and when to stop.

02 / The method

Why keep a champion?

Understand the loop, the evidence, and the limits.

03 / For agents

A better brief starts here.

Give your assistant the field guide, not guesswork.

04 / Relay beta

A run beyond one session.

Explore supervised rounds for Grok Build.