Skip to content

The whole discipline behind one MCP connection

Jackdaws runs as a hosted Model Context Protocol server: 130 tools that verify, grade, design, bridge, and eval Claude Skills, callable from claude.ai, Claude Code, or any MCP client you run. Subscription-gated, metered per user, and built so the generative work runs on your model, not a marked-up one.

Deterministic verifiers, generative agents

Layer 1 · verify

The numbered rules, as tools

Preflight, quality grading, hygiene, security posture, doc drift, subagent lint, eval discipline, and the rest of the read-only rule families. Each verifier accepts your skill’s files inline and returns findings that cite a rule id, never a mood. The same engine the toolkit runs against itself, with nothing to install locally. That includes the checks on measurement itself — whether your eval discipline is written down, and whether a passing score could have been reached without doing the work. See how the toolkit measures itself.

Layer 2 · generate

Agents that do the authoring

design_skill_from_brief turns a paragraph into a typed design plan. bridge_skills wires hand-offs, generate_skill_evals drafts the eval suite, and optimize_description tightens trigger phrasing. Every output passes back through a deterministic validation gate before it counts as done.

A tool response is not paid once

It enters the caller’s transcript and is re-sent as input on every turn that follows. A listing returned once is billed again on every later turn it survives. A tool’s own description is worse: it sits in the cached prefix for the whole session, whether or not the model ever reaches for that tool.

MT001 – MT005

tool_token_audit

It reads an MCP server and reports the shapes that make that cost unbounded. A collection tool with no limit or cursor, so the caller cannot ask for less than everything. No output cap anywhere in the server. A tool returning a whole file where an identifier or a range would do. A description heavy enough to weigh on every turn. A surface large enough that the definitions alone deserve deferred loading.

The boundary

Two questions, cleanly split

mcp_security_audit asks whether a hostile caller can hurt you. This one asks what answering costs. We pointed it at our own two servers before publishing it. It found five list tools with no bound and no output cap on either. Those findings are open in the repository, not in a slide.

The same question runs on the other side of the connection. harness_lint now checks an agent harness for the five ways a prompt cache is lost. A tool inventory shipped with no breakpoint. A prefix edited mid-session. History rewritten instead of appended to. A cache nobody measures. Stable values serialized differently on every run. Each failure is silent. It arrives as a bill, not a stack trace.

Bring your own model

Generative tools resolve a model in three tiers, in order. The server never requires you to pay for someone else’s model markup to use the discipline.

  1. 01

    Your client's model, via MCP sampling

    Clients that advertise sampling (Claude Code, Claude Desktop, custom harnesses) run the generative agents on their own model. The server contributes the prompts and the rules; the model cost stays yours.

  2. 02

    A server-held key, when provisioned

    Operators can configure a server-side Anthropic key as a fallback for clients that cannot sample. Optional, and off by default.

  3. 03

    A generation contract, everywhere else

    With neither, the call returns the agent's system prompt, rendered user prompt, and output schema for your own model to execute in-conversation, closed by the deterministic validate_layer2_output gate. Layer 2 is never unavailable to a subscriber.

The boring parts, done properly

OAuth, not API keys in prompts

Authentication is standard OAuth through Scalekit. The claude.ai connector registers via dynamic client registration; automation uses per-user machine-to-machine credentials.

Fail-closed entitlements

Tool calls and workflow prompts are gated on an active subscription with a strict allowlist. Listing stays open, so you can see the surface before subscribing.

Metered and rate-limited per user

Every call lands in a usage ledger keyed to the calling identity, with per-user rate limits. Usage is attributable, not pooled guesswork.

Sandboxed custom rules

Custom verification rules run as hash-pinned WebAssembly modules with no filesystem, network, or environment access, under fuel and memory budgets.

A pinned tool surface

The server's tool definitions are locked in CI. What the catalog says is deployed is what is deployed.

Content-in verifiers

The hosted verifiers accept your skill's files inline, materialized into a bounded sandbox per call. No shared mounts, no uploads that persist.

Getting connected

One subscription, one identity

Sign in, pay, and land in your account. claude.ai, Claude Code and Google Antigravity sign in with the same email; your own harness uses an API key you create there. Either way the verifiers, agents, and workflow prompts are live.

The connect page carries the per-client steps: configuration shape, which clients lend a model, and what each one meters.

A wiki your agents can read

The toolkit’s best-practices deep modules are published as a generated, drift-guarded knowledge bundle in Open Knowledge Format. Remote agents read it as MCP resources, so your assistants cite the same discipline the verifiers enforce.