Skip to content

Cairn — Getting Started

Mission control for humans and AI agents. Cairn organizes work around outcomes and proves progress with evidence — not a status someone typed. This guide gets a whole team started, whether or not you write code.

The human-first work loop described here (Capture → Shape → Commit → Flow → Review → Learn) works with no AI at all. On top of it, Cairn serves context to external AI agents, lets them propose work for you to decide on, and lets them do some of it — claim authorized work, attach evidence, link a pull request, submit for review — inside a boundary you set per agent. Everything in that last group is reversible, and an agent cannot mark its own work verified.


Part 1 · For everyone (no code required)

What Cairn is

Most trackers organize activity — tickets, columns, statuses you drag around. Cairn organizes outcomes: the change you're trying to produce, the reasoning behind it, and the evidence that it happened.

Four things make it different from a ticket tracker:

  • Outcome-first. The center of gravity is a Mission's outcome and success measures — never a to-do list. Tasks live underneath the outcome.
  • Progress is evidenced, not typed. You don't mark something "done." You attach proof — a merged PR, a green test run, a moved metric — and the status advances on its own.
  • Decisions are remembered. Why you chose something, the trade-offs, and when to revisit are recorded and kept even when a decision is later changed.
  • Everything is auditable. Every change writes to a tamper-evident timeline with who did it and when.

The building blocks

A dozen plain-language objects make up the whole system. You meet them in roughly this order as you use the app.

Object What it is
Signal Anything that might deserve attention — a bug, a customer ask, an idea, an alert. Capturing is lightweight and open to everyone; not yet a commitment.
Mission A bounded piece of work defined by its outcome, with success measures, constraints, assumptions, and a review date. Never a task list.
Brief The living, versioned shared understanding of a Mission. A version can be approved.
Proposal & Commitment Propose a plan or scope change, then commit the Mission with a rationale, what's out of scope, and a review date.
Work Item A concrete unit of execution under a Mission, with acceptance conditions (what "done" verifiably means), an owner, a claimant, and dependencies.
Deliverable A reviewable output — PR, design, report, release — that enters a shared review queue.
Evidence An objective, source-linked fact (merged PR, passing CI, a moved metric). Verified evidence against an acceptance condition is what advances a Mission.
Decision A recorded choice with rationale and trade-offs. Revisiting creates a new record and supersedes the old — history preserved.
Risk Something that could knock a Mission off course, with likelihood, impact, indicators. Escalating pulls a human in.
Attention Inbox What needs a person, in four lanes (Review, Decide, Clarify, Intervene), generated from the record — not noisy notifications.
Timeline The tamper-evident, chronological record of what changed, with the actor and an integrity check.
Pulse A plain summary — what changed, what's blocked, what decisions are needed — assembled from work history.
Actor Anyone who acts: a person, an AI agent, an integration. Each has an identity, a permission level, and an audit trail. Agents are never anonymous keys.
Handoff How an actor ends its work: what is done, what is not, what was decided, what a human must pick up. A claim, never proof.
Agent Run One bounded session by an agent — what it read, how long it took, what it reported spending, and how it ended. Every proposal points back to one.

A Mission moves along a fixed status chain. You never set these by hand past commitment — evidence advances them:

Captured → Shaped → Proposed → Committed → Authorized →
Executing → Reviewable → Verified → Shipped → Adopted

Start your workspace

You need nothing installed to use a running Cairn — just a browser.

  1. Sign in. Enter your email; Cairn sends a one-time magic link. No password. (Teams can also enable "Sign in with GitHub".)
  2. Create a workspace. A workspace is your organization; the creator becomes the owner. Give it a name and a short URL slug (e.g. acme).
  3. Invite your team. Open Team in the sidebar: admins and owners invite teammates by email, pick a role, and can revoke an invitation that hasn't been accepted. The invitee gets an emailed link, lasting seven days, that shows them who invited them and to what before asking them to sign in. Belong to several workspaces? Switch between them with the picker at the top.

The work loop

How a regular task travels through Cairn, from a stray thought to a verified outcome:

  1. Capture. Someone drops a Signal in Intake — "customers keep hitting checkout timeouts." No process, no assignee.
  2. Shape into a Mission. At triage you classify or dismiss signals (with a reason), then shape the worthwhile ones into a Mission. Cairn makes you state the outcome — "Checkout never times out again" — not a task. This opens the Mission Canvas, which always leads with the outcome.
  3. Build shared understanding & commit. Write the Brief and approve a version. Then Shape → Commit: submit a Proposal, and Commit with a rationale, what's out of scope, and a review date. → Committed.
  4. Do the work. Break the outcome into Work Items, each with acceptance conditions ("p95 < 300ms"). Claim them, move them along, link dependencies. Attach Deliverables for review. → Executing, then Reviewable.
  5. Verify with evidence. Attach Evidence against each acceptance condition — pick the condition, link the proof, mark it verified. Each flips ○ → ✓; when all are met the Mission reaches Verified. Nobody typed "done."
  6. Learn & steer. Record Decisions (revisit later without losing history). Raise and escalate Risks. The Attention Inbox shows what needs a person; the Timeline proves what happened; the Pulse summarizes for a stakeholder.

The mental model: a Mission is an outcome, and progress is evidenced. Everything else hangs off those two ideas.

Running a team

Roles separate "who can see" from "who can change" from "who can administer":

Role Can do
viewer Read everything and capture Signals. For stakeholders who watch outcomes.
member Do the work: shape Missions, commit, add work items, review deliverables, verify evidence, record decisions, raise risks.
admin Everything a member can, plus manage people (invite, set roles) and integrations (register agents/automations, issue tokens).
owner Full control; the person who created the workspace.
  • Workstreams group ongoing work that never really ends (reliability, onboarding). Attach Missions to a Workstream so a bounded outcome has a home.
  • Attention Inbox turns live conditions (a review waiting, a decision pending, a blocked item, an escalated risk) into a short, lane-sorted list; items clear themselves when you act on the underlying object.
  • Pulse gives a manager the "what changed / what's blocked / what needs a decision" digest without anyone writing a status report. It can be written by a model per audience (team, executive, stakeholder) — every claim cites the record it came from, and anything the model cannot ground is dropped before you see it.

Working with AI agents

Cairn is designed so an outside AI tool — Claude Code, Cursor, or anything speaking MCP — can understand your work without being handed your database, and can help without being able to change anything on its own.

Connecting one. Open Connect an agent. You register the agent as an Actor (it gets a name and an identity, not an anonymous key), choose what it may do, and copy a ready-made config for your tool. The credential is shown once.

What it may do is a deliberate choice:

Scope What the agent can do
Read only Read the context it is authorized to see — missions, briefs, acceptance conditions, decisions, dependencies, pulses. It cannot write anything.
Read + propose Also submit plans, work items, decisions, scope changes, estimates, and report blockers or ask questions — all of which wait for a person.

Nothing an agent submits takes effect by itself. Proposals land in the Decide lane of your Attention Inbox with what they actually say — the plan's steps, a decision's options, an estimate's basis — and the badge tells you an agent wrote it. Accepting is what makes anything real. Questions land in Clarify and blockers in Intervene; you answer them there. An agent reporting a blocker does not mark the work blocked — that stays your call.

Questions come back answered. When an agent asks something it cannot work out, your answer does not just close a notification: it becomes part of that Mission's context. The next run reads it automatically, so nobody answers the same question twice — and the list of what was ambiguous is usually a hint that the Brief needs a sentence it does not have yet.

Work is handed over, not abandoned. When an agent finishes a session it leaves a Handoff: what it completed, what it did not, what it decided along the way, what risks it is leaving, and what you need to do. If anything was left unfinished it must say what you should do about it — and that lands in your inbox until someone picks it up.

One place shows you every actor. The Actor Console lists everyone who acts here — people and agents together, because both are governed the same way — with what each may do, what they are working on, how their sessions went, and what they cost. Rates like "proposals accepted" only appear once there is enough history to mean something; below that it says so rather than printing a confident number.

Turning an agent off. Admins can Disable an agent from the console. It stops being able to act and its live credentials are revoked, but it stays on the roster with its runs, proposals and history intact — Cairn withdraws rather than deletes, so the record of what a thing did survives the decision to stop using it. Enabling it again does not resurrect the old credentials; issue a new one. Individual credentials can also be revoked from an agent's page. People are not disabled this way: they act through a session, not a credential, so removing someone is a membership question answered on the Team page.

Some actions stop and wait for you. Every action has a permission level, and so does every person and agent. Committing a Mission or recording a Decision needs Level 4 — an admin, or someone individually trusted with it. When someone below that tries, the action does not happen: it becomes a checkpoint in your inbox showing exactly what was attempted, and you choose how far to lend your authority — just this once, for a day, or from now on. The last two are recorded as a capability on that person, so you can see and revoke them later.

Agents are told how to work here. Cairn serves a set of workflows to every connected tool — how to catch up on a Mission, how to check whether work is ready to start, how to propose a plan or surface a decision — so two different AI tools follow the same process instead of improvising. In Claude Code and Cursor they show up as slash commands. You can see the list on the Connect page.

You can always see what an agent did. Every call it makes, allowed or refused, is on the Timeline, and every session is recorded as an Agent Run with its tool calls and duration. Revoking its credential stops the next call. An agent also cannot bury you: repeat submissions are folded into the original, and one agent may only have a handful of undecided items waiting on you at a time.


Part 2 · For builders (self-hosting & integrating)

Architecture at a glance

A TypeScript monorepo. Nothing here needs an AI provider to run.

  • Web — Next.js (App Router, RSC). Port 3100.
  • API — NestJS. Port 3400. The web app proxies /api/* to it, so external clients should use the app origin.
  • Database — PostgreSQL (pgvector image). Every meaningful change is written to an append-only, per-organization hash-chained event store in the same transaction — this is what makes the audit timeline tamper-evident.
  • Queue — Redis + BullMQ for background work.
  • Multi-tenant — every row is scoped to an organization; isolation is enforced at the data layer, not by discipline.
  • Observability — OpenTelemetry traces on every request; the trace id rides along on each event.

Run it locally

No provider keys required for the default stack.

# 1 · start Postgres, Redis, and friends
docker compose up -d

# 2 · install, migrate the database, and seed
pnpm install
pnpm --filter @cairn/api db:migrate
pnpm --filter @cairn/api db:seed

# 3 · run the web + api dev servers
pnpm dev
# web  → http://localhost:3100
# api  → http://localhost:3400

Signing in during local dev: no SMTP is configured out of the box, so the magic link is printed to the API server console. Enter your email in the app, then copy the link from the log line starting with [auth] magic link for …. Configure EMAIL_SERVER to send real email later.

Inbound intake API

To feed Signals in automatically — from a webhook, an email-forwarding gateway, or a monitoring alert — use the token-authenticated intake endpoint.

1 · Register an integration and issue a token (admin):

# register an integration actor
curl -X POST http://localhost:3400/api/orgs/<slug>/actors \
  -H "Content-Type: application/json" --cookie "authjs.session-token=…" \
  -d '{"type":"integration","displayName":"Email gateway","permissionLevel":1}'

# issue a scoped credential — the cairn_… token is shown ONCE
curl -X POST http://localhost:3400/api/orgs/<slug>/actors/<actorId>/credentials \
  -H "Content-Type: application/json" --cookie "authjs.session-token=…" \
  -d '{"scope":"intake:write"}'

2 · Post signals with the token:

curl -X POST http://localhost:3400/api/intake/signals \
  -H "Authorization: Bearer cairn_…" \
  -H "Content-Type: application/json" \
  -d '{
    "source": "email",
    "rawContent": "Customer reports checkout timeouts at peak",
    "summary": "Checkout timeouts",
    "urgency": "high"
  }'

The Signal lands in the workspace's Intake, attributed to the integration, ready to triage and shape. Only a hash of the token is stored, and it carries only the intake:write scope.

Key endpoints

App endpoints are under /api/orgs/:slug/… and use your session; the intake endpoint uses a bearer token. A selection:

Method & path What it does
POST /api/orgs Create a workspace (you become owner)
GET · POST /api/orgs/:slug/invitations List outstanding / invite by email + role (admin)
DELETE /api/orgs/:slug/invitations/:id Revoke an invitation that hasn't been accepted (admin)
GET /api/invitations/:token What an invitation link points at (no session — the token is the credential)
POST /api/invitations/accept Redeem an invitation as the invited address
GET · POST /api/orgs/:slug/signals List / capture Signals
POST /api/intake/signals Inbound capture with a bearer token
GET · POST /api/orgs/:slug/missions List / create Missions
GET /api/orgs/:slug/missions/:id/canvas The assembled Mission Canvas
POST /api/orgs/:slug/missions/:id/work-items Add Work Items; claim & transition
POST /api/orgs/:slug/missions/:id/deliverables Attach Deliverables for review
POST /api/orgs/:slug/missions/:id/evidence Attach Evidence (advances status)
GET /api/orgs/:slug/attention The Attention Inbox, by lane
GET /api/orgs/:slug/timeline The tamper-evident audit timeline
GET /api/orgs/:slug/pulse The templated status digest
GET /api/orgs/:slug/search?q= Search objects → { hits, mode, degraded }
GET /api/orgs/:slug/ai/status Whether AI is configured and working
GET /api/orgs/:slug/mcp/connection Setup recipes + what each scope grants
GET /api/orgs/:slug/agent-runs Agent sessions: tool calls, cost, outcome
GET /api/orgs/:slug/agent-requests Blockers and questions raised by agents
POST /api/mcp The MCP endpoint (bearer credential)
GET · POST /api/orgs/:slug/connectors List / link connectors (link: admin, L5)
GET /api/orgs/:slug/connectors/:id/events Recent deliveries — metadata, not payloads
POST /api/connectors/:id/webhook Inbound webhook (signed, unauthenticated)
GET · POST /api/orgs/:slug/missions/:id/resources Where a Mission ships

When an agent can change the shared record

mcp:write adds four more tools — capture a signal, create work items, record a decision, close a run. It buys less than the name suggests, deliberately, and the asymmetry is the whole shape of it:

tool stops for a human? why
record_decision always A decision binds the team. A Level-4 action by a non-human actor routes to approval however the agent is provisioned — there is no grant that makes its decision final.
create_work_item no Reversible, so the ladder does not stop it. Use propose_work_items if you want it reviewed.
create_signal no Capture is universal — a signal is raw and untriaged and carries no authority.
close_agent_run no Ending your own session binds nobody.

So the rule is: an agent may change what the organization knows, and may not change what it has decided without a person saying so.

When a decision waits, the Attention Inbox shows the decision itself — the question, what was chosen, the rejected alternatives, the rationale — not "Record decision needs Level 4 approval". Approving something you cannot read is not approval. Approving does not execute it either: the agent retries, and the grant is spent by that retry, so one approval authorizes one decision.

Two things an agent cannot do here: close a run over work it still holds (create_handoff exists for exactly that, and says what is left), and record a decision with no rejected alternative — which is an announcement, not a decision, and is refused before it reaches anyone.

When an agent can act

mcp:execute adds the seven execute tools — claim work, declare what a session is for, attach evidence, attach a deliverable, link a pull request, submit test results, submit for review. Granting it is three separate facts, and an action passes only if all three hold:

  1. the credential carries mcp:execute,
  2. the actor is trusted to Level 3, and
  3. a Context Contract admits that action on that Mission.

So a leaked token, an over-trusted actor, and a mis-scoped contract each fail alone. Without a contract every execute call is refused — which is the gate Cairn's policy engine has demanded for non-human actors since Phase 0 and which nothing reached until now.

Everything an agent does here is reversible. A claim can be released, evidence superseded, a deliverable rejected, a review sent back. Nothing merges, deploys, deletes, or reaches a customer.

An agent cannot verify its own work. Evidence it attaches is recorded unverified, whatever it puts in the call. It may name the acceptance condition it believes the evidence satisfies — and that condition must be one actually declared on the Work Item, or the call is refused, because otherwise an agent invents a condition, satisfies it, and the record reads as though something had been proven. A person confirms before a Mission moves.

That is the same line Cairn draws everywhere: a CI run ingested from GitHub does verify, because nobody in the organization authored it; an agent reporting that its own tests passed does not. On the Mission Canvas the two sit side by side, labelled observed and agent-reported.

start_agent_run does not start a run — Cairn opens one on an agent's first call, which is what keeps the record of what it did observed rather than self-reported. The tool declares what the run is for, and the Actor Console shows that beside how the run ended, so reviewing a session is a comparison.

Connecting an external system (webhooks)

A connector is how Cairn is reached from outside — the first place it stops only reading and starts touching systems other people can see. Two things are worth knowing before you link one.

Linking is a Level-5 action. It grants Cairn reach into a shared system, so it always waits for a human, even when an owner asks. The first request returns 202 with a checkpoint id; approve it in the Attention Inbox and submit again.

Reach is enumerated, never implied. A connector may touch exactly the resources you list and nothing else — there is deliberately no "all repositories" option, and an empty list reaches nothing. Install the provider-side app on the same set: what GitHub enforces is worth more than what Cairn does, because a broken Cairn still cannot reach a repository the App was never installed on.

# 1 · link (returns 202 with a checkpointId the first time)
curl -X POST http://localhost:3400/api/orgs/<slug>/connectors \
  -H "Content-Type: application/json" --cookie "authjs.session-token=…" \
  -d '{"provider":"github","displayName":"Acme GitHub",
       "resources":[{"ref":"acme/payments-api"},{"ref":"acme/docs"}]}'

# 2 · approve it, then repeat step 1 — now it returns the connector
curl -X POST http://localhost:3400/api/orgs/<slug>/checkpoints/<id>/decide \
  -H "Content-Type: application/json" --cookie "authjs.session-token=…" \
  -d '{"grant":true,"scope":"once"}'

The successful response carries the webhook path and a signing secret shown exactly once. It is stored encrypted (an HMAC signature has to be recomputed, so unlike an actor credential it cannot be stored as a hash) and no screen will show it again — rotate to issue a new one.

Point the provider at POST /api/connectors/<connectorId>/webhook, signed with X-Hub-Signature-256: sha256=<hmac> for GitHub, or X-Cairn-Signature-256 for the generic provider. The HMAC is over the exact request bytes.

What happens to a delivery, in order: oversized bodies are refused unread; an unknown or disabled connector answers a bare 404 (identical, so the endpoint cannot be used to discover ids); a bad or missing signature is 401, counted on the connector and audited at a rate limit; a redelivery returns { duplicate: true }; and an authentic delivery about a resource you never granted is accepted but shelved as ignored — so a misconfigured hook is visible rather than silent. Everything else is stored, audited, and processed on a queue with exponential backoff, with an exhausted delivery recorded as connector.event_failed.

Only a github and a generic adapter ship today. The other providers in the roadmap are declared but not linkable: declaring a provider is a plan, shipping an adapter is a capability, and a connector that could never verify a delivery would sit in the UI looking configured.

Where a Mission ships

A Mission can declare the external resource it ships in, on its Canvas under Ships in. This does two jobs. As a control, the first declaration confines work on that Mission to what it names (before it, any connected resource is allowed). As identity, it is what lets an inbound pull-request event find its Mission by a recorded fact rather than by matching on a title.

curl -X POST http://localhost:3400/api/orgs/<slug>/missions/<missionId>/resources \
  -H "Content-Type: application/json" --cookie "authjs.session-token=…" \
  -d '{"resourceId":"<connectorResourceId>"}'

What GitHub events become

Once a connector is linked and a Mission declares it ships in that repository, inbound deliveries stop being raw records and become facts about the work.

GitHub Cairn Verifying?
PR opened / reopened a Deliverable in the Review Queue — (it is the thing to review)
PR merged Evidence (pr) yes
PR closed unmerged Evidence (pr) no
Review approved Evidence (approval) yes
Changes requested Evidence (approval) no
Check suite / run / workflow completed Evidence (ci) only when green
Deployment succeeded / failed Evidence (deploy) only on success
Pushes, issues, stars, … nothing, with the reason recorded —

A pull request is the thing; everything that happens to it is a fact about it. That is why a merge does not update the PR record — it is new evidence attached to it. Attaching evidence re-derives the Mission's state through the same ProgressService a human's evidence goes through, so a Mission moves committed → executing → reviewable without anyone typing a status.

It stops at reviewable, deliberately. Reaching verified requires every declared acceptance condition to be met by verifying evidence naming that condition (P1-12), and no webhook knows which condition a green build met. A Mission that said what "done" means must not be auto-verified by CI passing; saying what was proven stays a human's or a contracted agent's act.

A Mission is found by its binding, never by matching. Cairn does not read a Mission id out of a branch name or a PR title. If no Mission ships in the repository, nothing is derived and the delivery says so. If two Missions do, it is recorded as ambiguous rather than guessed — a pull request is the deliverable of one piece of work, and creating it against both would put one PR in two review queues.

Everything derived this way is marked origin: connector and attributed to no actor, and the UI says "from GitHub" / "observed". This matters beyond bookkeeping: a pull request's title is written by whoever opened it, which on a public repository is a stranger. A reviewer reading it should know whether a colleague wrote those words.

Connecting an AI agent (MCP)

POST /api/mcp speaks MCP over Streamable HTTP, on both protocol revisions from the same URL: 2026-07-28 (no initialize handshake — each request carries its own version envelope, and clients discover the server with server/discover) and the 2025-era handshake that every older client uses. Point a client at the endpoint and it connects on whichever it speaks; nothing needs configuring either way. GET /api/mcp/info reports which revisions are served, without a credential.

An agent authenticates with a short-lived Actor credential and gets a tool surface fixed by that credential's scope. There are 29 tools in four families, and each scope is a superset of the one above it:

Scope Adds Tools
mcp:read Read authorized context. Changes nothing. 10
mcp:propose Plans, work items, decisions, scope changes, estimates, blockers, questions — each waits for a human. 9
mcp:execute Claim work, attach evidence and deliverables, link a PR, submit for review. All reversible. 7
mcp:write Capture a signal, create work items, record a decision. 3

GET /api/mcp/info describes the whole surface without authentication.

Three independent checks gate every call — the credential's scope, the actor's permission level, and (for execute and above) a Context Contract naming the Missions and actions that agent may touch. Without a contract an agent reaches every Mission in the workspace; grant its first one and it reaches only what it holds contracts for, and revoking them all leaves it with nothing rather than with everything. Both allowed and refused calls are written to the timeline, each refusal naming why.

The propose tools never change the record: they create a Proposal or an agent request that waits for a person. The execute tools do change it, but only in ways a person can undo — and record_decision always routes to a human no matter how the agent is provisioned, so the wider scopes let an agent raise a question, never answer it. In-product setup lives at /connect.

The server also implements the MCP prompts capability: prompts/list and prompts/get serve the workflows an agent should follow (plan_a_mission, check_before_you_start, catch_up_on_a_mission, …). Their requirement lines are generated from the same contract the validator enforces, so guidance cannot drift from what the endpoint accepts, and they are scope-filtered like the tools.

# what an agent is doing right now, and what it is asking for
curl -s localhost:3100/api/orgs/<slug>/agent-runs --cookie "authjs.session-token=…"
curl -s localhost:3100/api/orgs/<slug>/agent-requests?state=open --cookie "authjs.session-token=…"

Permission levels

Separate from team roles, every actor — a person, an AI agent, an integration — has a permission level from 1 to 5. It's a ceiling on the class of action they may take, and it's the backbone of the governance the AI layers rest on. Nobody deploys silently.

Level Class of action Today
1 · Observe Search, read, summarize, explain. Agents read context over MCP
2 · Propose Draft plans, decisions, scope changes, risks. Agents propose; humans decide
3 · Execute (reversible) Branches, drafts, evidence, deliverables, work items. Agents execute, inside a Context Contract
4 · Execute (governed) Merge approved changes, invoke integrations. Stops at a human checkpoint
5 · Sensitive Production deploys, access/data changes — always require human approval. Always stops for a human, even for an owner

For agents the level is only one of three gates: the credential's scope and the Context Contract are the others, and a call has to satisfy all three.

Levels 4 and 5 do not mean "an agent may do this." They mean the attempt is recognised rather than refused outright: it stops, lands in someone's inbox with the action it was trying to take, and proceeds only if a person grants it — once, for a day, or from now on. Linking a connector is Level 5, so it waits for a human even when an owner is the one asking.


The human-first work loop is complete and tested. Agents read context, propose work, and execute the reversible parts of it under the levels above. What is not built: agents merging, deploying, or acting outside a contract — those stop for a person by design, not by omission.