Open source · Apache-2.0 · v0.2.0-beta.3

Coding agents that ask before they push.

docket runs a small team of AI agents on your codebase. One plans, one writes, one reviews. Every file edit, shell command and API call passes one gate first, and the risky ones wait for you.

  • Roles
  • Policy gate
  • Human approval
  • Audit trail
docket — the tool-call gate
$ docket policies test pre_tool_call implementer 'git push origin production'
  Result: require_approval
$ docket pod myapp dispatch
→ Dispatching 1 pending task(s) through: lead → implementer → reviewer
  ⋯ tool.ask       tool=bash policy_id=high-risk-deploy
  ⋯ approval.deny  channel=timeout
$ docket audit verify
✓ 6 chained line(s) verified clean.

Most agent tools ask the model to behave. docket doesn’t ask.

The usual way

  • A prompt says “never push to production”, and the model decides whether to listen.
  • One agent plans, writes and grades its own work.
  • Edits land on the branch you have checked out.
  • When something goes wrong, you scroll a chat log.

With docket

  • git push origin production is held by code until a person says yes.
  • A Lead plans, an Implementer writes, a Reviewer and a Tester gate the result.
  • The Implementer works in its own git worktree. Your checkout stays clean.
  • Every decision lands in a hash-chained log you can verify.
01

A team, not a lone agent.

Each role gets only the tools its job needs. A Reviewer has no write tool at all, so it cannot be talked into editing.

02

One gate. No side door.

Built-in tools and MCP tools alike pass the same policy check before they run. There is no setting that turns it off.

03

Evidence, not promises.

Runs, traces, approvals and token counts stay queryable after the turn. docket audit verify tells you if the log was touched.

1
gate for every tool call
0
code paths around it
4
ways to approve: CLI, HTTP, MCP, Telegram
120s
of silence, and the answer is no

How it works

You write the task. The pod does the rest, in order.

docket init sets up a pod: a small team scoped to one project. docket pod <id> dispatch runs a task through it.

  1. 01LeadPlans and delegates. Never edits code.
  2. 02ImplementerWrites the change in its own git worktree.
  3. 03VerifyYour command. A nonzero exit fails the task.
  4. 04ReviewerRead-only. Can send it back.
  5. 05TesterPASS or FAIL. Nothing in between.
  6. 06EvidenceRun, trace and audit record.

Reviewer and Tester are optional. When they are on, their verdict blocks the task: it is not advice a model can argue past. Blueprints shape the same pipeline for research, content, ops or product work.

The gate

Pick a command. See what happens.

Every action takes the same path: policy, risk check, a person if needed, budget, then execution. Here is what that looks like.

role: implementergit push origin production
Held

Risky commands wait for a person.

Deploys, money and secrets are high-risk. The call pauses until someone approves it from the CLI, HTTP, MCP or Telegram. No answer in 120 seconds means denied.

Proof

Real output. No mockups.

Captured from real runs against a local model on 2026-09-18. Every line is what the CLI printed.

Terminal output: an Implementer’s git push origin production is held for approval by the high-risk-deploy policy, nobody answers, the call is denied on timeout without executing, and the audit chain verifies clean.
A push to production is held, nobody answers, it is denied without running, and the audit chain checks out.
Terminal output: the Implementer’s own workspace and a separate git worktree on its own branch, while the main checkout stays clean.
The Implementer’s edit lives only in its own worktree. Your branch is untouched.

Three ways in

Use it from your terminal, from your tools, or inside your app.

docket

CLI

Set up a pod and dispatch tasks against your own repo. The fastest way to a governed run.

docket harness run

Harness

One agent, one turn, driven by another program. It never waits for a person: anything that needs approval ends the run as blocked.

docket-runtime

Engine

Send your app’s own tools through the same policy, approval and audit gate. Two dependencies. Built from source.

Get started

Your first governed run.

Install, point it at a model, and dispatch a task. Local models work.

install
brew tap yielab/docket-cli https://github.com/yielab/docket
brew install docket-cli
first run
docket models provider add local http://127.0.0.1:8081/v1 \
  --model local-model --ctx 32768 --max-tokens 4096
docket models preset local

cd ~/code/myapp
docket init
docket pod myapp delegate "Create FIRST_TURN.md containing: governed first turn"
docket pod myapp dispatch

docket runs list
docket trace tail myapp
docket audit verify

Needs Python 3.11+, Git, Bash, and an OpenAI-compatible endpoint with tool calling: hosted (OpenRouter, Vercel AI Gateway) or local (llama.cpp, vLLM, LM Studio).

Read the full quick start →

Straight answers

What docket is not.

All known limits →
  • Not a dashboard. It exposes a read API so you can build or plug in your own.
  • Not a general orchestration framework. It runs one supervised team per project.
  • Not finished. It is beta, built for a single operator, and breaking changes between releases are expected.
  • Not a network jail. The fetch tool is allowlisted, but an allowed shell can still reach the network. Run untrusted work inside a container.

Put a gate in front of your agents.

Open source under Apache-2.0. Read the code, run it on your own machine, and check every decision it made.