Skip to content

Design, build, grade, ship

A skill is not done when it looks finished. It is done when it clears each stage of a defined lifecycle and passes a single publish gate.

What the code audit doesn’t catch yet

The audit reads your code without running it. It cannot tell who can see whose data, whether someone can change a price at checkout, or whether your backups work.

It also does not look for these yet:

  • Passwords and keys left in your own code. People report AI bills of $4,368 and $55,444 after one leaked key.
  • Database tables anyone can open. That is how a 2025 Lovable flaw exposed data.

The pre-ship checklist covers these by hand. It is free and there is nothing to install.

For developers

The checks give the same findings for the same input every time, use only the Python standard library, and make no model call. Every rule has a stable id in a numbered family, and ship_check chains them into one GO or NO-GO verdict. One real finding, as it comes back raw:

CP009  error  app.py:26
os.system() runs a shell and is a command-injection sink — use subprocess.run([...]) with an argument list

In plain English: Your code hands a whole command line to the shell. Whoever controls that text can end your command and start their own.

Every skill takes the same path

  1. 01

    Architect

    skill-design-plan-architect

    Turn a one-paragraph brief into a typed design plan: an anatomy claim, a file list, four worker parcels, and the sister skills to bridge to.

  2. 02

    Orchestrator dispatch

    skill-build-pipeline

    Fan out four worker subagents in parallel — SKILL.md body, references, scripts, and evals — then integrate their output on disk.

  3. 03

    Dogfood

    skill-ralph-loop-advisor

    Score the skill's loop, tools, memory, and context profile. Optional: purely mechanical skills with deterministic output can skip it.

  4. 04

    Tier-fix sweep

    skill-dogfood-triage

    Classify each finding as Tier 1 mechanical, Tier 2 calibration, or Tier 3 architectural, then fix cheapest-first.

  5. 05

    Bridge wiring

    skill-bridge-patcher

    Wire the new skill into its sister skills' hand-offs, using one of four canonical bridge patterns.

  6. 06

    Preflight and grade

    skill-preflight-check · skill-quality-grader

    Validate the skill for upload and grade its content against the quality rubric. This is the terminal gate.

Each check has a number

Findings cite a specific rule, not a preference. A sample of the families the toolkit enforces:

FamilyChecksOwner
P001–P017Claude.ai upload parser compliance plus Anthropic's Skills best-practices specskill-preflight-check
Q001–Q010Content-quality rubricskill-quality-grader
DD001–DD007Stale documentation referencesskill-doc-drift-sweep
CE001–CE007Context-engineering on workflowsskill-context-audit
LS001–LS010LLM/agent security posturellm-security-best-practices
MS001–MS020MCP server project correctnessskill-mcp-server-builder
SA001–SA014Sub-agent definition hygieneskill-subagent-audit
ET001–ET019Whether the eval discipline is written down and practisedeval-and-testing-best-practices
TC001–TC004A run against a declared required-tool-set (advisory)eval-and-testing-best-practices
V001–V010Hazeley voice guidelinkedin-content-strategy

A rule is a claim until something measures it

Numbered rules say what good looks like. They do not say whether the skills pass, whether the score can be trusted, or whether a passing score could have been reached without doing the work. That is a separate discipline, and the toolkit audits itself against it.

The discipline is auditable

ET001–ET019 checks whether a repository has written down and practised an eval discipline: the vocabulary, the split between capability and regression suites, the graduation bar, a recorded baseline. It runs against this toolkit and against any other project. Findings are advisory and each one names the reference that explains it.

A score you cannot fake

Berkeley’s 2026 audit reached near-perfect scores on eight well-known agent benchmarks without solving a single task, by exploiting how each score was computed. So the grader, its ground truth, and the baseline it compares against must sit outside the write scope of the thing being graded. A validity check runs first, because a metric that rewards a low number scores a missing artifact as its best result.

Reliability is not capability

Passing once in several attempts says the behaviour exists. Passing every time says it can be relied on. The two are different numbers, and reporting only the first hides which one you have. Path is never graded: the outcome is the gate, and trajectory checks stay advisory.

Where this stands today. Every shipped skill carries a graded eval suite, and the deterministic scorecard runs in CI on every change. That is a capability baseline: it runs each case once. The repeated-trial sweep that would produce a reliability figure is wired and scheduled, and has not yet been run, so no reliability number is published here. The ledger in the repository records that gap instead of rounding past it.

19 workflows chain the tools end-to-end

audit-and-bridge

W4 - library health -> wiring: ecosystem-audit a collection, bridge each overlap/gap edge, then re-audit to confirm.

Type thisRun onSonnet

Audit the skill collection in skills/ for overlap, then propose bridges that let each skill delegate to the right downstream skill.

author-prompt

W12 - author a prompt from an intent when none exists yet: interview the brief into a prompt-spec, render it through the profile for its kind, lint it against PD001-PD014, read it adversarially against its own spec, and stop at the command that would measure it.

best-practices-pass

W9 - audit a repo for software-development best-practice readiness, triage the gaps, then advise a ranked adoption plan (advisory, no verdict).

Type thisRun onSonnet

Audit the monorepo in /repos/backend-services for engineering practices and prioritize adoption by implementation cost.

build-hermes-agent

Build a Hermes agent end-to-end: scaffold a Track-A or Track-B project, validate HM001-HM008, then fix to clean. Counterpart of the build-hermes-agent.js workflow (skill-hermes-agent).

Type thisRun onSonnet

Scaffold a Track A Hermes agent at /projects/new-agent. Validate the project structure to ensure all dependencies resolve cleanly.

build-skill

W1 - author a new skill end-to-end: architect -> 4-worker fan-out -> dogfood-triage -> bridge -> ship gate. Formalizes skill-build-pipeline.

Type thisRun onOpus

Create the skill from .claude/plans/data-sync.md into skills/connectors/data-sync/. Implement all stages, error recovery, and retry logic outlined in the locked design document.

cc-dir-tuneup

W18 - drive a project's Claude Code directory to clean: audit CD plus the sibling families, bootstrap missing rungs with the scaffolder, fix cheapest-first by family, re-audit; secrets and the user directory always go to a person.

change-history-catchup

W16 - clear the decision-log backlog: window the pending evidence, write each window in order, then ingest once. Serial by necessity - the writer reads the ledger for entries its window supersedes.

design-workflow

W15 - brief -> validated workflow.json: decide single call vs workflow vs agent, draft the spec, then validate WF001-WF012 and fix until clean.

eval-sweep

W13 - join the eval-discipline audit (ET001-ET019), the deterministic k=1 scorecard, and an opt-in judged pass@k/pass^k sweep into one ledger-ready record. Reports; never gates.

extend-rule

W7 - add or tighten a numbered rule (P/M/A/Q/R) across the five touchpoints, then audit completeness.

Type thisRun onSonnet

Rule-303 in api-consistency-validator is too loose on trailing slashes. Add stricter logic and propagate through linter, formatter, CI, git hooks, and the web interface.

friction-postmortem

W6 - tune a skill after real use: quantify trailing friction from a transcript, classify it, fold a bounded patch into the right artifact, re-gate.

Type thisRun onOpus

Can you trace through session.jsonl to identify why skills/api-refactor didn't meet expectations, then propose one specific, bounded fix?

harvest-repo

W10 - first-contact harvest evaluation of an external repo: Gate 0 recon -> dedup sweep -> analysis record + registry rows -> HR audit. Formalizes skill-repo-harvest.

Type thisRun onSonnet

Evaluate https://github.com/torvalds/linux and create a durable record of the practices most worth emulating.

prospect-to-design

W5 - demand -> design: mine session transcripts for recurring workflows, rank by reach, design the top candidate, then build it.

Type thisRun onOpus

My past sessions are in ~/.claude/transcripts. Find recurring workflows, score by time-saving potential, design the top three as skills.

release-new-skill

W8 - toolkit propagation gate: audit the cross-cutting registries for a new skill and emit a release checklist.

Type thisRun onSonnet

New skill: doc-coauthoring at ~/.claude/skills/doc-coauthoring/. Is it registered in all required registries? Give me the release checklist.

sdlc-loop

W14 - drive a repository's SDLC posture to clean: collect the facts, audit the intent/spec/plan chain and control backing (SD001-SD014), triage, fix cheapest-first, re-audit; escalate what survives two rounds.

select-skill-prompts

W11 - the best published example prompt for each skill, by measurement: scaffold a suite from a skills tree, review it, then generate -> rubric -> routing -> pairwise as client-executed contracts, and gate the result.

Type thisRun onSonnet

Pick the example prompt for every skill in skills/ by testing, not vibes. Generate options, check which reach the intended skill, keep the best. I want evidence behind each published line.

ship-skill

W2 - pre-publish GO/NO-GO gate over one skill: preflight -> grade -> ecosystem-audit, fastest-first with early-exit, then a verdict.

Type thisRun onSonnet

Run the GO/NO-GO check on skills/skill-mcp-server-builder before I upload it. Tell me pass or fail, nothing else.

triage-and-fix

W3 - finding -> tiered fix sweep until clean: run a gate, classify findings into Tier 1/2/3, fix cheapest-first, re-gate.

Type thisRun onSonnet

Run the gate against skills/skill-bridge-patcher, sort findings cheapest-first, fix what's fixable, then re-gate until it's clean.

vibe-coder-reposition

W17 - repositioning brief -> verified copy, harvested repos, two new skills, and an external-comparator report. Three stages in one workflow: content (fact-check, draft, lint, opus-panel verify, harvest, design two plans), build (nested build-skill per clean plan), comparator (landing_page_audit + a final relint, reports only). Run each stage as a separate Workflow() call.

Each prompt above was selected out of 75–175 candidates by a scored funnel. The full set, including the single-verifier asks, is on Prompts.

One verdict, in the output

Hand-rolling this discipline means re-deciding what “good” means on every review. The gate settles it once: a skill is GO or NO-GO, and the reason is a named rule that fires the same way for every author, every time.

ship_checkGO/ NO-GO