slipway

slipway · a project framework for Claude Code

AI changed how we build software. It didn’t change what users expect from it.

Slipway is the framework I start every project from: a path from a rough idea to a release that holds up, with Claude Code doing the building and checks that fail when a step gets skipped. By an engineer, for engineers.

npx use-slipway acme --dry-runPrints every step it would take. Changes nothing.

npx use-slipway acmeCreates ./acme and a private GitHub repo. Then open Claude Code and run /bootstrap.

  • Branch protection on a private repository needs a paid GitHub plan.
  • The first CI run on main is red on purpose.
  • The command installs session hooks into the new project’s .claude/settings.json; --no-harness skips it.

The package is use-slipway; slipway on npm is unrelated.

A slipway is the ramp a ship is built on. It stays there until it’s ready, then goes down the ramp once.

8
steps from a rough idea to a product people use
1
command tells you what to do next
1
milestone active at a time
0
packages to install for the checks to run
The one idea

A rule only counts if something fails when it’s broken.

What problem does it solve?

Agents made writing code cheap. They didn’t make it right, useful or coherent. That part is still engineering.

I’ve been building products with coding agents since mid-2024. It went in four stages.

  1. No structure. One session at a time, no plan, no docs. The agent said “Done and pushed.” I found the bug; it said “You’re absolutely right.” Whole days went to circling the same bug. I couldn’t trust anything I saw on the screen.
  2. Tickets and CI. I started planning in issues: a defined scope, PRs small enough to review, and CI doing the checking. The quality I could expect changed overnight. None of it was new. It was ordinary engineering discipline, applied to agents.
  3. Skills. I wrote skills for the repeated work, like logging a bug, shaping a feature and working a ticket, and the results became consistent. I could run several tickets at once. Then I had a director’s problems: the same number calculated a different way in each feature, a drawer in one screen and a full page in the next. It felt fast. It wasn’t.
  4. Slipway. Everything that worked, with most of what I learned the hard way turned into a check that fails where one can catch it, and a dated rule where it can’t, so the next project starts at stage four.

The agent’s coding was never the bottleneck. Everything around it was: knowing what to build, proving it works, and keeping it coherent as it grows. Slipway is built to get you to a release that holds up sooner. It won’t get you to a one-shot MVP.

Three questions before anything ships

  • Does it work? Tests the agent is sent back to when they fail, and CI that runs the same checks on every pull request.
  • Is it useful? It answers the question the product exists to answer, for the person asking it.
  • Does it feel right? It fits the product: the patterns it reuses, and where its numbers come from. Slipway doesn’t check this one yet; it’s next on the list below.

What goes wrong, and what catches it

A test suite nobody runs has a 100% pass rate. Each row below is a check in CI or a hook in the agent’s session. The checks run on plain Node with nothing to install, so a broken dependency can’t quietly switch one off.

What goes wrongWhat fails instead
The agent finishes its turn with tests failingA hook runs pnpm verify:fast when the agent tries to stop. It’s sent back once with the failure, and told to say what is failing if it stops red anyway.
A check that can’t failEvery check has a known-bad example it must go red on, for exactly the expected reasons.PC1
Tests or scripts that CI never runsEvery gated script has to be invoked by a CI workflow.W1
A pull request with no record of what was runThe PR description has to say what was verified, and link its issue or say why there is none.P1
“Make it faster” as the acceptance criterionAn issue that uses a bare adjective as a criterion is labelled needs-shape, and /work-ticket won’t start it.I1
A plan nobody argued withThe PRD needs a written review that names the exact version it read.R1
Building before anyone tested demandNo milestone past the first skeleton starts until each value risk has a result against a bar written first, or is settled by experience with what would prove it wrong written down.K1
A milestone that runs overOne milestone is active at a time. Past its time budget, it needs a written decision.MS1
A lesson nothing enforcesEvery lesson names where it’s enforced or, if it isn’t built yet, the event that reopens it. The judgment ones get a review date, and fail when it passes.L1

What will my workflow look like?

Eight steps from a rough idea to something people use. You do the thinking; the agent does the building.

You never have to remember where you are. pnpm status reads the repo and prints which step you’re on and what to do next. Agent sessions get the same line when they start.

  1. 0Bootstrapyou + agent · 1–2 h

    One command makes the repo. /bootstrap adds the app and runs the probes that show each gate can go red, listing the ones you run yourself.

  2. 1Frameyou · an afternoon

    /kickoff asks one question at a time: who reaches for this, when, and what it answers for them.

  3. 2Test the riskyou · days to weeks

    The cheapest test that could fail. Talk to people, do the job by hand. Write the bar first.

  4. 3Shapeyou · 1–2 days

    A PRD, the decisions that are costly to change later, 3–5 milestones. Then a fresh session argues with it.

  5. 4Skeletonagent · days

    The thinnest real path, deployed through CI with analytics and errors wired. Runs alongside step 2.

  6. 5Buildagent · the rest of the milestone

    One milestone at a time. Small PRs, each carrying its evidence.

  7. 6Closeyou · about 7 minutes to a pull request, once, on the strongest model

    A retro written from the record, on the one run so far. Out of time? Cut scope and close.

  8. 7Learnyou · weekly

    Usage numbers and real conversations pick the next milestone.

    → back to 5

What most of my days look like

  1. Take the next slice from the active milestone. Anything that isn’t in it goes to its “not doing” list or a later milestone, not into today’s change.
  2. Size it. One sentence and no new behaviour is a plain PR. One session’s work gets an issue with checkable acceptance. Bigger, or a new concept, gets a short feature doc with a contract first.
  3. Hand it to a fresh agent session. It looks for an existing helper before writing a new one. Once a project wires a duplication check, a ratchet never lets its number go up.
  4. When pnpm verify:fast is red, the agent is sent back once to fix it, and told to say what’s failing if it stops anyway.
  5. Review a PR that says what was verified. CI runs the same pnpm verify I do. There’s no second definition of green.

What does a successful slipway project look like?

You can leave it for two weeks, come back, run one command and know what’s next.

I switch between projects a lot, so this is the part I care about most. It works because the state of the project lives in the repo, not in my head or a chat history. Open one that’s going well and this is what you find:

And outside the repo

  • A live URL from the walking skeleton on, with analytics checked to fire.
  • Pull requests that say what was verified. The template also asks what wasn’t.

The skeleton ships in days, and each milestone is a bet you’re allowed to lose.

What does it ask of me?

It will ask you why. That’s the cost, and most of the point.

Before it builds anything, slipway asks what question the product answers and who’s asking. Before a feature, it asks which part of that question the feature serves. /kickoff asks one question at a time, so the reasoning ends up on paper instead of in your head.

That can feel like a barrier, and I’ve felt it too. Plenty of good features start as experience: you’re the user, or it’s table stakes in the domain. Nobody trusts a finance app with half its features. Some products can’t be tested by hand in a spreadsheet first. Slipway lets you settle a value risk by experience (it’s table stakes, you’re the user, or you know the domain), when you write down what would prove you wrong.

The time it takes: an afternoon to frame, days to weeks to test the riskiest assumption, a day or two to shape.

Who it’s for

  • Engineers who own the build, alone or as the only engineer beside people who own the problem.
  • Heads-down engineers who want a product mind next to them. It doesn’t replace technical judgement. It asks the questions a good product partner would.

Who it isn’t for

  • A weekend one-shot. If the goal is a demo by Sunday, this is the wrong tool.
  • Building without an engineer. It assumes someone who can read the diff and argue with the plan.
  • Collectors of skills. It isn’t a pack of hundreds of agents or a cast of role-play personas. It’s a short path, a dozen or so skills, and checks that fail.

Can I use it on a project I already have?

Honest answer: not in one command yet. Today slipway starts new projects, and keeps them up to date.

Updates are built in. A project started from slipway takes newer versions with /sync-slipway. It prints a plan before it touches anything, works on its own branch, merges instead of overwriting a file you changed, and never edits the documents that are yours: the PRD, the frame, your decisions.

Adopting an existing repo is a later bet, not under way. Adopting a project that didn’t start from slipway is filed, with its acceptance written, and not built. Help triaging an old backlog, which issues still hold against the code and which can go, is another later bet, ahead of it.

Until then, pieces travel on their own. The checks are Node scripts that run with no install step. The lessons are plain Markdown files. The working rules for agents fit on one page. A check brings the helpers in ci/checks/lib/ it imports, and slipway records none used outside a project that started from it.

What you’re running. Each release is staged by CI from a version tag, and goes public only when I approve that exact package at npm with a second factor.

What it expects

  • Node 24, pnpm 10, git and the GitHub CLI, logged in to a GitHub account.
  • Claude Code for the agent side: the hooks, the skills, the session status.
  • Questions the first time. Claude Code asks whether you trust the new folder, and lists the two permissions the project pre-approves: git stash list and git stash apply. /bootstrap then asks three things about your product, one at a time: which GitHub project board new issues go to (“none” skips the board), the timezone your deadlines are read in, and whether the product has money or other math that must be exact.
  • GitHub for issue forms, required checks and Actions. Moving elsewhere means rewiring the checks tied to it.
  • A default stack you can change: TypeScript, React with Vite, Cloudflare. You confirm or replace it in week one, and write down why.

What’s proven so far?

Slipway is being built in the open, by one engineer, and used on real work. Here’s where it stands.

Shown working

  • Every check is seen failing before it’s trusted. Each has known-bad examples it must go red on, for exactly the expected reasons.
  • Slipway’s own work goes through its lanes. Its feature docs, decisions, lessons and reviews are in the repo.
  • A private product with one engineer, a web app, has used it. It ran the gates on real pull requests and issues, and its repository was protected by the setup script. /kickoff shaped its product, and the Stop hook caught a problem in a live session that became a fix in slipway (#58).
  • Seven real syncs, scored together. The bar is one command, three questions at most, and the owner able to say what changed. One sync met it. Two missed it. Four were not fully scored: for two, the record says nothing on whether the owner could say what changed, and for the other two it has no command count. Each is dated in SLIPWAY.md.
  • The first milestone close has run. On 2026-10-06, /close-milestone took about seven minutes from the start of the session to an open pull request, and about half an hour to merge, on the strongest model at high effort. It asked one question: which of the open issues belonged to the milestone. One run; nothing else has been timed.
  • That project had built every item of its first milestone. The close ended it before its appetite ran out. Every gate line was shown by a CI run or a recorded result, and the retro was written from issues, pull requests and git log: the work that crept in, the rework and its cause, the code-health numbers, one lesson for the project.
  • The first real close used a fix made that day. With no GitHub milestone, slipway’s close now lists the project’s open issues and asks which belong, instead of reporting nothing left (#258, fixed in #264 on 2026-10-06, before the close ran). It also found a bug: it marks a feature doc shipped whole even when one section still owes a check (#284, open).

Not proven yet

Each line is dated 2026-10-07, in SLIPWAY.md, under “Not verified here — read as unknown”.

  • The agent hooks and approval prompts in a live session. The Stop hook has only been seen blocking on a missing install, not refusing on a failing test. The approval prompts haven’t been seen in a live session, and the other hooks and the session-start status have been seen only on sample input.
  • A close that isn’t the first one’s path. A milestone ended as killed, extended or cut; a project with a GitHub milestone; a close that changes the PRD.
  • The skills’ settings on a real project. This repository holds no record of the intake and ticket skills’ runs there.
  • A limit of the issue check. It catches adjectives, not criteria that can’t fail: “returns HTTP 200” passes.

Next

  • A “fit” check for features (the patterns a feature reuses, where its numbers come from), and a pause at each milestone to reorganise what should be shared.
  • The whole plan, in order, is epic #43.

Why I built this

I’ve been building software for twenty years: mining, web hosting, e-commerce, beauty, identity and fine-grained authorization, restaurants, and now healthcare. I’m a staff engineer, and most of what I know I learned the hard way.

Agents have changed how I work. They didn’t change what the people using my software expect: that it works, that it’s useful, and that it feels right.

Three values run through how I lead engineering, and through slipway. Quality: gates that can actually fail. Transparency: every PR says what was verified, and this page says what isn’t proven. Integrity: you never weaken a gate to get past it.

Refactoring has never been cheaper, and debt still compounds. Slipway is how I stay deliberate at agent speed. It’s what I start every project from now.

— Mat Dupont