slipway · a project framework for Claude Code
The agent said “done.” The tests were failing.
Slipway is the project template I built after too many days like that. It takes you from a rough idea to a release, with checks that fail when you or the agent skip a step. It assumes an engineer is reading the diffs.
npx use-slipway acme --dry-runPrints every step it would take. Changes nothing.
npx use-slipway acmeCreates ./acme and a private GitHub repo. Then open Claude Code and run /bootstrap.
- The command installs session hooks into the new project’s
.claude/settings.json;--no-harnessskips it.
The package is use-slipway; slipway on npm is unrelated.
A slipway is the ramp a ship is built on. It stays there until it’s ready, then goes down the ramp once.
What the command creates
One command sets up a repo with planning docs, eleven Claude Code skills and CI checks. The skills take you through the work one step at a time.
- docs/Templates for the product frame, the PRD and the milestones, with the questions to answer already written in.
- .claude/skills/Eleven skills, run as slash commands:
/kickoff,/log-feature,/work-ticket,/close-milestoneand seven more. Each is a written procedure the agent follows. - ci/checks/Plain Node scripts that GitHub Actions runs on every pull request.
- .claude/settings.jsonHooks for Claude Code sessions. One runs
pnpm verify:fastwhen the agent tries to finish, and sends it back once if that fails. - pnpm statusOne command that reads the repo and prints which step you’re on and what to do next.
A rule only counts if something fails when it’s broken.
What problem does it solve?
Coding agents write code quickly. They also tell you it works when it doesn’t, and build each feature as if the others didn’t exist.
I’ve been building products with coding agents since mid-2024. It went in four stages.
- No structure. One session at a time, no plan, no docs. The agent said “Done and pushed.” I found the bug; it said “You’re absolutely right.” Whole days went to circling the same bug. I couldn’t trust anything on the screen.
- Tickets and CI. I started planning in issues, with a defined scope and PRs small enough to review, and let CI do the checking. Quality improved right away. None of this was new. It’s how teams already work.
- Skills. I wrote skills for the repeated work (logging a bug, shaping a feature, working a ticket) and the results got consistent enough to run several tickets at once. Then the product stopped agreeing with itself: the same number calculated a different way in each feature, a drawer on one screen and a full page on the next. Every PR passed review on its own. Nothing checked that they added up to one product.
- Slipway. I put what worked into one template. Most of what I learned the hard way is now either a check that fails or a written rule with a review date. A rule that nothing checks gets skipped, by me as much as by the agent. A new project starts here instead of at stage one.
Writing the code was never the slow part. Deciding what to build, proving it works and keeping features consistent with each other were. Slipway is my attempt at those three. The third is the least finished. It won’t get you an MVP from one prompt.
Three questions I ask before shipping
- Does it work? CI runs the checks on every pull request. A hook runs
pnpm verify:fastwhen the agent tries to finish, and sends it back once if that fails. - Is it useful?
/kickoffwrites down the one question the product answers./log-featureasks whether a new feature serves it. One that doesn’t waits. - Does it fit? It reuses the patterns the product already has, and gets its numbers from the same place. Slipway doesn’t check this yet. It’s planned.
What goes wrong, and what catches it
Each row is a script in CI or a hook in the Claude Code session. The CI scripts run on plain Node with nothing to install, so a broken dependency can’t silently turn one off.
| What goes wrong | What happens instead |
|---|---|
| The agent says it’s finished while tests fail | When the agent tries to end its turn, a hook runs pnpm verify:fast. If it fails, the agent gets the output and is sent back once. If it stops again, it’s told to say what is still failing. |
| Tests exist but CI never runs them | Every script named test, lint, check, build or typecheck has to be called by a CI workflow, or the build fails. |
| A check that can’t fail | Each check ships with deliberately broken examples. If it doesn’t fail on them, for exactly the expected reasons, the build fails. |
| A PR that doesn’t say how it was tested | The PR description has to say what was run, and link its issue or say why there isn’t one. |
| “Make it faster” as an acceptance criterion | An issue whose acceptance is a bare adjective gets labelled needs-shape, and /work-ticket won’t start on it. |
| A plan nobody pushed back on | The PRD needs a written review that names the exact version it read. |
| Building before checking anyone wants it | Nothing past the first thin version starts until each “will anyone want this?” risk has a result written down: a test result against a pass mark set first, a note that you’re relying on experience with what would prove you wrong, or a decision to build ahead. |
What will my workflow look like?
Eight steps, the same ones on every project. The early ones are mostly you deciding what to build. After that the agent writes most of the code and you review it.
You don’t have to remember where you are. pnpm status reads the repo and prints the step you’re on and what to do next. Agent sessions get the same line when they start. Before /bootstrap has run, the line reads:
**Next:** Step 0 (you + agent) — Bootstrap: run /bootstrap …
- 0Bootstrapyou + agent · 1–2 h
One command creates the repo.
/bootstrapadds the app, then runs each check against something broken to show it fails. It lists the ones you have to try yourself. - 1Frameyou · an afternoon
/kickoffasks one question at a time: who uses this, when, and what question it answers for them. - 2Test the riskyou · days to weeks
Pick the assumption that sinks the product if it’s wrong. Write down what result counts as a pass. Then run the cheapest test: talk to people, or do the job by hand for a few of them.
- 3Shapeyou · 1–2 days
A PRD, the decisions that are expensive to change later, and 3–5 milestones. Then a fresh session reviews the plan and looks for holes.
- 4Skeletonagent · days
The thinnest version that works end to end, deployed through CI, with analytics and error tracking on. It can run alongside step 2.
- 5Buildagent · the rest of the milestone
One milestone at a time, in small PRs. Each PR says what was run to test it.
- 6Closeyou · about 7 minutes to a pull request, timed once, on the strongest model
/close-milestonewrites a retro from the issues, PRs and git log. Out of time? Cut scope and close anyway. - 7Learnyou · weekly
Usage numbers and conversations with users decide the next milestone.
→ back to 5
What most of my days look like
- Pick the next piece of the active milestone. Anything outside it goes on the milestone’s “not doing” list or into a later one, not into today’s change.
- Size it. A change you can describe in one sentence, with no new behaviour, is just a PR. One session of work gets an issue with acceptance criteria you can check. Anything bigger, or anything that adds a new concept, gets a short feature doc first.
- Give it to a fresh agent session. The agent looks for an existing helper before writing a new one.
- If
pnpm verify:fastfails when the agent tries to finish, it’s sent back once with the output. If it stops anyway, it’s told to say what is failing. - Review the PR. It says what was run. CI runs the same
pnpm verifyyou run locally, so passing on your machine and passing in CI mean the same thing.
What does a successful slipway project look like?
You can leave it for two weeks, come back, run one command and know what’s next.
I switch between projects a lot, so this is the part I care about most. It works because the state of the project lives in the repo, not in my head or a chat history. Open one that’s going well and this is what you find:
- docs/product/FRAME.mdThe one question the product answers. Every feature either helps answer it or is listed as out of scope.
- docs/product/evidence/Notes from interviews and tests, each with the pass mark written before the test ran.
- docs/PRD.mdThe plan. Every requirement has an ID that issues and milestones refer to.
- docs/reviews/The review that challenged the plan, and which version it read.
- decisions.mdWhy the stack, the data model and the money types are what they are, written on the day they were decided.
- docs/milestones/One active milestone, with a time budget and the conditions for abandoning it, both written before it started. Each finished one has a retro.
- docs/product/metrics.mdWeekly numbers and what users said. The next milestone comes from here.
- process/lessons/Every rule the project has learned, each saying where it’s enforced or, if that isn’t built yet, what would reopen it. Rules that need judgment carry a review date.
And outside the repo
- A live URL from step 4 on, with analytics checked to fire.
- Pull requests that say what was tested. The template also asks what wasn’t.
What does it ask of me?
It makes you answer “why” before it builds, and that takes days, sometimes weeks.
Before any code, /kickoff asks what question the product answers and who is asking it, one question at a time, and writes your answers into the repo. Before a new feature, /log-feature asks whether it serves that question.
This can feel like paperwork, and sometimes I’ve felt that too. Not every doubt about demand needs a test. Sometimes you’re the user, or every product in the category has the feature, or you know the domain. Slipway accepts that: you can mark that risk as settled by experience, as long as you write down what would prove you wrong.
The time it takes: an afternoon to frame the product, days to weeks to test the riskiest assumption, a day or two to plan.
Who it’s for
- Engineers who own the build, alone or as the only engineer working with people who own the problem.
- Engineers who’d rather be in the code and want something beside them asking the product questions. It doesn’t replace your technical judgment.
Who it isn’t for
- A weekend project. If the goal is a demo by Sunday, this is the wrong tool.
- Building without an engineer. It assumes someone who can read the diff and argue with the plan.
- Anyone after a big library of agents or personas. Slipway is eleven skills, one path and a set of checks.
Can I use it on a project I already have?
Not yet. Today slipway starts new projects and keeps them updated. It can’t be added to an existing repo in one command.
Updates. A project started from slipway takes new versions with /sync-slipway. It shows a plan first, works on its own branch, merges rather than overwrites files you changed, and never edits your own documents: the PRD, the frame, your decisions.
Existing repos. Adopting a repo that didn’t start from slipway is planned and not started. The issue is written, with its acceptance criteria, and nothing is built. Help sorting an old backlog is also planned, and comes before it.
Until then, you can copy pieces. The checks are Node scripts that run with no install step; take one along with the helpers it imports from ci/checks/lib/. The lessons are Markdown files. The agent rules fit on one page. I have no record of anyone doing this outside a project that started from slipway.
How releases are published. CI stages each release from a version tag. It goes public only after I approve that exact package at npm with a second factor.
What you need
- Node 24, pnpm 10, git and the GitHub CLI, logged in to a GitHub account.
- Claude Code, for the skills, the hooks and the session status.
- GitHub, for issue forms, required checks and Actions. On another host you’d rewire the checks that depend on it.
- A paid GitHub plan, if you want branch protection on a private repository.
- A default stack you can change: TypeScript, React with Vite, Cloudflare. You confirm or replace it in week one, and write down why.
What the first run asks
- Claude Code asks whether you trust the new folder, and lists the two permissions the project pre-approves:
git stash listandgit stash apply. /bootstrapasks three things about your product, one at a time: which GitHub project board new issues go to (“none” skips the board), the timezone your deadlines are read in, and whether the product has money or other math that must be exact.- The first CI run on
mainis red on purpose.
What’s proven so far?
One engineer builds it in the open and uses it on real work. Here’s what that has shown, and what it hasn’t.
Shown working
- Every check has been seen failing. Each has deliberately broken examples it must fail on, for exactly the expected reasons.
- Slipway’s own changes go through the same three sizes of change it gives a project. Its feature docs, decisions, lessons and reviews are in the repo. It has no product frame, PRD or milestone of its own.
- One real product has used it: a private web app with one engineer.
/kickoffshaped the product, the checks ran on its real pull requests and issues, and the setup script protected its repository. The hook that runs when the agent stops caught a problem in a live session that became a fix in slipway (#58). - Seven updates, one that met the target. The target is one command, three questions at most, and the owner able to say what changed. One sync met it, two missed, and four weren’t recorded fully enough to score. Each is dated in SLIPWAY.md.
- One milestone closed. On 2026-10-06,
/close-milestonetook about seven minutes to open a pull request and about half an hour to merge, and asked one question. Every item planned for the milestone had been built, and the retro was written from issues, pull requests andgit log. Timed once, on the strongest model at high effort. It also turned up a bug: a feature doc is marked shipped whole even when one section still owes a check (#284, open).
Not proven yet
Each line is dated 2026-10-07, in SLIPWAY.md, under “Not verified here — read as unknown”.
- Hooks in a live session. The hook that runs when the agent stops has only been seen blocking for a missing install, not for a failing test. The approval prompts haven’t been seen in a live session. The other hooks and the session-start status have only been seen on sample input.
- Any other kind of close. A milestone that was abandoned, extended or cut short; a project that uses GitHub milestones; a close that changes the PRD.
- The intake and ticket skills on a real project. This repository has no record of how those runs went.
- The issue check has a limit. It catches adjectives. It doesn’t catch a criterion that can’t fail: “returns HTTP 200” passes.
In progress
The work under way now, in order, is epic #312.
Planned, not started
- A check that a feature doc past draft says how it fits the product (the patterns it reuses, where its numbers come from), and a pause at each milestone to list what should be shared and decide.
Why I built this
I’ve been building software for twenty years: mining, web hosting, e-commerce, beauty, identity and fine-grained authorization, restaurants, and now healthcare. I’m a staff engineer, and most of what I know I learned the hard way.
Agents made code cheap to write. They didn’t make it correct, and the people using what I build still expect what they always did: that it works, that it’s useful, and that it fits together.
Quality, transparency and integrity are the values I lead engineering by. In slipway they mean checks that can fail, pull requests that say what was tested, a page that says what isn’t proven, and never weakening a check to get past it.
Agents make rewriting cheap, which makes it tempting to skip the thinking. Slipway is how I keep doing the thinking. I start every project from it now.
— Mat Dupont