About
The dojo skill is a self-improvement training loop that uses surplus quota time to automatically analyze the ax graph, run experiments, and generate reports. It triggers via specific commands like "/dojo" and requires axctl on PATH, utilizing embedded DuckDB without a database daemon. Developers should use it for automated backtesting, proposal generation, and issue reporting during unused quota windows.
Quick Install
Claude Code
Recommendednpx skills add Necmttn/ax -a claude-code/plugin add https://github.com/Necmttn/axgit clone https://github.com/Necmttn/ax.git ~/.claude/skills/dojoCopy and paste this command in Claude Code to install this skill
Documentation
ax:dojo - overnight training loop
You are entering a budget-bounded self-improvement loop. The brain is
ax dojo agenda --json; you are the thin driver. Spec:
docs/superpowers/specs/2026-06-13-ax-dojo-design.md (in the Necmttn/ax repo).
Entry
- Run
ax dojo agenda --json. If it fails with a connection error, tell the user to runax doctorand STOP. - If
budget.has_surplusis false: report the envelope and STOP unless the user re-invokes with--force(then pass--forceon every lap). - On Claude Code: enter loop mode now - invoke the
/loopskill with/dojoas the recurring prompt (dynamic mode, self-paced). Each wakeup re-runs this skill from the top; that is expected and correct. On Codex (no /loop): run as ONE long turn - do not end the turn until a stop condition below is met.
The lap
ax dojo agenda --json-> agenda.- STOP conditions (write the report, then stop):
budget.has_surplusis false- now >=
budget.deadline itemsis empty
- Otherwise: take
items[0], follow its playbook below, then go to 1. Completed work self-clears: the item vanishes from the next agenda because the underlying system recorded it (verdict locked, brief consumed, proposal created). If the same item survives 2 laps untouched, skip it and note why in the report.
Playbooks by kind
- verdict_pending -
ax improve verdict <id>to see the suggested verdict + checkpoint evidence; confirm with--set <verdict>only when the evidence supports it. Distinguish "pattern resolved" from "artifact never fired" before locking no_longer_needed. - brief_unfilled - open the
.ax/tasks/*.mdbrief, do what it says in the target files, then run the reconciler it names (ax skills lint/ax improve lint). - routing_backtest - judgment-flagged routing classes: backtest the
pattern against dispatch history (
ax dispatches --candidates), check false-positive risk, thenax routing tune --apply=<ids> --days=<window>or reject with a written rationale in the report. - proposal_mint -
ax improve recommend; accept the grounded ones (ax improve accept <id>) so briefs exist for the next lap. - experiment - heavy item. Work ONLY in a fresh worktree
(
git worktree add .claude/worktrees/dojo-<slug> -b dojo/<slug>). Reproduce the churn pattern, attempt the fix/hook/skill, capture evidence. If it will not finish inside this budget: package it as a goal file (objective + checkpoint index + gates) under docs/superpowers/goals/ so the NEXT dojo session resumes it. Output = an improve proposal; merging the proposal is what activates anything. NEVER merge, never touch main. - New hooks specifically - author via @ax/hooks-sdk, then run BOTH
validators and embed their output in the proposal:
ax hooks backtest <file> --json→ cases caught (benefit side): would-block/ would-warn rates, false-positive count, cases with evidence.ax hooks bench <file> --json→ per-fire p50/p95 from real bun spawns, est fires/day from tool_call history, installed-chain budget vs --budget-ms default 250 (cost side). Reject the hook when daily cost (fires/day × p95) or an installed-chain budget overrun outweighs the benefit shown by backtest. Both ledgers must appear in the proposal; neither alone is sufficient.
- spar - only present when invoked with --spar and spendable >= 30%.
One task, one delta, scored. Concrete flow:
- Pick a landed task:
ax sessions here --days=30- note its commit sha fromax sessions near <sha>orgit log. ax dojo spar-plan <sha>- captures the baseline (prompt + cost/turns/churn) and writes~/.ax/dojo/spar/<id>.md; the command prints the exactgit worktree addcommand to run next.- Read the brief at
~/.ax/dojo/spar/<id>.md; run the printedgit worktree add .claude/worktrees/dojo-spar-<id> -b dojo/spar-<id> <parentSha>command to pin the worktree at the parent SHA. - Apply exactly ONE delta in the delta section (skill on/off, hook on/off, prompt change, thinking level, or model override) - no compound changes.
- Do the task in that worktree; let it finish naturally.
ax dojo spar-score <id>- auto-discovers the variant session from the worktree cwd; or pass--variant-session=<id>if there are multiple sessions. Writes the receipt to~/.ax/dojo/spar/<id>-report.md.- Append the receipt to the dojo report. Track multi-run campaigns as goal files under docs/superpowers/goals/ so the next session can resume.
- Pick a landed task:
- explore - free investigation, retro-meta style: follow a hunch
through
ax recall/ax sessions churn, and convert anything real into a proposal or outbox draft. - Upstream findings (any lap) - an ax bug or improvement found while
training (items of kind
upstream_draftare handled by this same rule): runax dojo draft --title=<title> --kind=bug|improvementto stage it to~/.ax/dojo/outbox/<slug>.md(complete issue draft: title, body, repro, session refs written by the command). NEVER publish from the dojo - the user reviews and publishes in the morning (ax-repo skill / gh).
Exit - the morning report
Run ax dojo report --since=<loop-start-iso> --notes-file=<lap-notes-path> to
write ~/.ax/dojo/reports/<YYYY-MM-DD>.md. The command collects the budget
envelope, per-lap item log (from the lap notes file), proposals created,
verdicts locked, and outbox drafts awaiting review - pass it the ISO timestamp
you recorded when the loop started and the scratch file you appended notes to.
Then tell the user the report path and the top 3 things awaiting their review.
For upstream findings (ax bugs or improvements discovered during training), stage
them with ax dojo draft --title=<title> --kind=bug|improvement before the
report step - never publish directly. The draft lands in
~/.ax/dojo/outbox/<slug>.md; the user reviews and publishes via ax-repo skill /
gh in the morning.
Hard rails
- worktrees only; never write on main; never merge anything
- proposals are the only activation path
- outbox only; nothing leaves the machine
- respect the deadline even mid-item: checkpoint, report, stop
GitHub Repository
Frequently asked questions
What is the dojo skill?
dojo is a Claude Skill by Necmttn. Skills package instructions and resources that Claude loads on demand, so Claude can perform dojo-related tasks without extra prompting.
How do I install dojo?
Use the install commands on this page: add dojo to Claude Code as a plugin, or clone its repository into your skills directory, then restart Claude so it picks up the skill.
What category does dojo belong to?
dojo is in the Testing category, tagged ai, testing, design, and data.
Is dojo free to use?
Yes. dojo is listed on AIMCP and free to install.
Related Skills
This Claude Skill runs the lm-evaluation-harness to benchmark LLMs across 60+ standardized academic tasks like MMLU and GSM8K. It's designed for developers to compare model quality, track training progress, or report academic results. The tool supports various backends including HuggingFace and vLLM models.
This skill provides comprehensive knowledge for implementing Cloudflare Cron Triggers to schedule Workers using cron expressions. It covers setting up periodic tasks, maintenance jobs, and automated workflows while handling common issues like invalid cron expressions and timezone problems. Developers can use it for configuring scheduled handlers, testing cron triggers, and integrating with Workflows and Green Compute.
This Claude Skill provides a Playwright-based toolkit for testing local web applications through Python scripts. It enables frontend verification, UI debugging, screenshot capture, and log viewing while managing server lifecycles. Use it for browser automation tasks but run scripts directly rather than reading their source code to avoid context pollution.
This skill helps developers complete finished work by verifying tests pass and then presenting structured integration options. It guides the workflow for merging, creating PRs, or cleaning up branches after implementation is done. Use it when your code is ready and tested to systematically finalize the development process.
