TarkOS

An operating layer that makes a Claude Code agent run on files, not prompts.

Most agent setups keep the agent's behaviour in the prompt, so every session starts by re-teaching it. TarkOS moves that layer into version-controlled files that load in every session: a short posture, a decision tray, a correction-to-rule pipeline, verification gates and an engine that works a backlog while I'm away.

01 · Scale
852
agent sessions since Feb 2026, 32,000+ commits
02 · Evals
6 of 8
instruction rules that helped in a blind A/B, same split on two models
03 · Rebuild
82 KB to 44 KB
always-on context after a measured rebuild, no rules lost
01 · Problem

Every session started from zero. I wanted it to start from files.

I run Claude Code agents every day. Out of the box, each session forgets how I want it to work, and a lot of the steering lives in the prompt: check your work, don’t delete things, ask before the risky calls. So every session began with re-teaching, and the agent’s “done” was a claim I still had to check.

A fresh-context grader I built later made the second part concrete: it re-checked the agent’s own “done” claims and found 7 of 12 over-claimed.

GoalA fresh session, running on its instructions and mechanisms alone, behaves the way I’d want without in-session steering, and proves its work instead of asserting it.

02 · What it is

A generic core, a thin agent config, and files as the bus

TarkOS is two layers. The core holds the mechanisms and is generic: it ships empty of content and knows nothing about my domain. The host agent keeps its identity, mission and its own rules. The host’s instruction file joins them with two @-import lines, so the core text is never pasted in.

An autonomous engine loads only in away mode. Everything talks through git-backed files (tasks, decisions, project state, a capture ledger) rather than function calls, which keeps every step auditable and revertible. TarkOS Cockpit is a desktop surface over those same files.

fig.1 · tarkos · architecture DWG TK-03 · REV C
FIG. 1 · Solid = loads or data flow. Amber dashed = the capture loop: a correction in a session becomes a written rule in the core. Cyan = the install doctor's check. Dashed boxes load only when needed.

Most of the code in TarkOS, and in the tools around it, is written by AI agents under my specification, test contracts and review of the result. My part is the architecture, the decisions and the taste. That the agents can do the rest reliably is the thing TarkOS exists to prove.

The install manifest lists 9 mechanisms. The boundary between core and host is enforced by instruments, not by judgement: a leak suite for identity tokens, a reference checker that allows shipped docs to point only at shipped components, and a frontmatter ratchet.

03 · How it works

Five mechanisms, each with an artifact you can open

  1. M1: Posture

    Five always-on points: be proactive, decide what’s yours to decide, run autonomously, evolve from corrections, and report simply. When I’m away, the agent logs each estimate as an assumption and opens the next session with an assumptions round.

  2. M2: Decision tray

    Reversible and confident: do it, then leave a record that one command can undo. Irreversible or unsure: a decision card with options, consequences and a recommendation, which I ratify when I choose. The agent never ratifies its own card.

  3. M3: Capture

    When I correct the agent, the correction is written into the file that owns it, with a ledger entry. The next session loads the rule instead of repeating the mistake.

  4. M4: Verification gates

    Verdicts come from output, never from exit codes. A fix carries evidence that the test failed before and passes after. A done-claim whose evidence is older than the branch tip fails.

  5. M5: Autonomous engine

    An orchestrator dispatches task cards to at most two workers, logs every dispatch and acceptance, and ends the run with one morning card: what changed, what waits for me, what was measured, what stays unverified.

Installing on a new agent starts with a read-only doctor. It reports what would collide or is missing, and it never deletes anything:

fig.2 · tarkos-doctor + provision throwaway host
FIG. 2 · The doctor on a throwaway demo agent: two RED findings (missing import lines), then the provisioner seeds the config. Real output, verbatim.

Decisions that wait for me show up in TarkOS Cockpit with the agent’s recommendation and its stated confidence:

fig.3 · cockpit · decisions synthetic demo data
FIG. 3 · Cockpit’s decision view. The projects and decisions in the capture are synthetic demo data. Try the live demo (opens in a new tab) (synthetic data).
04 · Key decisions & trade-offs

Five calls that shaped it

D1 · Autonomy

Route every decision by confidence and reversibility

Alternatives
  • Ask about everything
  • Maximum aggression
  • Let the agent tune its own thresholds
Why

Autonomy grows by shrinking irreversibility. High confidence and reversible: act, then keep or revert. Medium: act and flag it. Low confidence or irreversible: always ask. Autonomy levels are per category; promotion is only offered after 90% prediction accuracy over at least 10 decisions, demotion below 75% is automatic, and money or anything irreversible never moves.

Trade-off

Some decisions wait for me that the agent would have got right. That’s cheaper than undoing an irreversible one.

Outcome

The first live run: 14 do-then-confirm records, all kept, none reverted. Later the agent caught its own noise filter hiding 468 of 468 records from me.

D2 · Loading

Two import lines instead of a hook

Alternatives
  • SessionStart hook injection
  • A regex prompt manifest
  • Paste the core into the instruction file
Why

The core first loaded through a SessionStart hook, so subagents got it too. It cost latency, and the prompt manifest only ever saw my prompt, not what the task needed later: all 4 of its measured injections went unused. Two @-import lines are visible in /context and need no hooks.

Trade-off

Hosts that strip the instruction file need the optional hook loader.

Outcome

Hooks stay where they’re the right tool: a rule that keeps being violated gets a mechanism that enforces it, never louder prose.

D3 · Rebuild

A measured zero-base rebuild: delete beats reorganise

Alternatives
  • Keep adding rules
  • Reorganise the files
  • Start a new harness
Why

A measurement tool showed a lot of always-on context, routing split across two systems, and 33 of 47 skills that never fired. Every cut went to a staging folder with a manifest, a seven-day watch and a one-commit revert, and a floor of invariants could not be simplified.

Trade-off

Slower than deleting outright. One earlier session misread six live mechanisms as dead and deleted one; since then the rule is to report, not cut, wherever no regression gate exists.

Outcome

Always-on context 82 KB → 44.3 KB. The ten largest files 726 KB → 253 KB, with no section headings lost.

D4 · Portability

A generic core and a thin per-agent config

Alternatives
  • Publish v1 without the identity-laden scripts
  • Publish as it is
Why

Mechanisms go to TarkOS; identity stays with the agent. The test for every rule: would a stranger’s agent in another domain need it? The package is generated from the live repo, not maintained by hand.

Trade-off

The agent recommended publishing a smaller v1 first. I overrode it and chose to generalise first, which delayed the release.

Outcome

Identity leaks in shipped files went from 153 to 0. The export is being rebuilt for release and is still unpublished.

D5 · Engine

Rebuild the engine as an orchestrator with workers

Alternatives
  • Rebuild but keep the worker relay
  • Not now
Why

The engine had grown to 15 files and 62 mechanisms, and its reference docs had zero reads across 2,405 telemetry rows. I noticed I was babysitting runs that were supposed to run without me.

Trade-off

It dropped session chaining across the context window, which had taken weeks to build.

Outcome

Worker-slot use went from 30% to 94% on the next night run.

05 · Evals

Measure the instructions, and the instruments

Blind A/B of rules
6 of 8

Arm A gets only the rules under test, arm B gets none. Same model, questions and criteria written before the run, a blind pairwise judge plus keyword checks. The same split held on a second model.

Cheaper judge
20× cheaper

One batched judge instead of ten per-question judges, with 8 of 10 verdicts agreeing. Both disagreements were cautious ties.

Done claims
7 of 12

The agent’s “done” claims that a fresh-context grader found over-claimed. It’s why verdicts come from output.

Instruments check themselves
1 scorer bug

The agent found that its own autonomy scorer counted the harness’s injected messages as my prompts. Fixing it was a measurement correction, not an improvement, and a regression test now guards it.

06 · A session, replayed

What an overnight run actually looks like

Recorded from a real session while I was asleep. The agent scored my launcher’s publishability gates on a fresh export and found two of them red: 7 of 14. Instead of just reporting it, it widened the fix, wrote a failing test first, fixed the bug and the identity hits, and handed the result to a worker with zero context for a re-score: 10 of 14. One verification run went red with no failing test. It was recorded as a flake and handed on as the next task, not hidden.

fig.4 · session.replay ~/launcher · overnight run
Step 18/18Elapsed 1:36:55Model claude
result 7/14 → 10/14 · a flaky run recorded, not hiddenRecorded from a real TarkOS session (overnight run, 2026-09-23). Paths shortened, private details removed, timings real.

Other sessions are documented only in their verification reports. One example is a launcher performance fix that took frame-time p95 at four panes from 764.7 ms to 31.0 ms with the same harness, after the agent first caught its own measurement error.

07 · What isn't done

What isn’t done yet

  1. G1: It isn't public yet.

    The open-source export exists, but it is being rebuilt for release: the harness is frozen at a stable version while the export is reworked.

  2. G2: "Portable" is proven on one install.

    My own agent is the only production install so far. A second agent is the real test.

  3. G3: The main goal isn't moving yet.

    The measure I care about is fewer corrections per session. It’s going the wrong way: my corrections per 50 sessions sat between 60 and 77 in earlier stretches and reached 153 in the latest one, while the ratio of problems the agent reports itself to problems I catch fell to 0.85. I track the two together on purpose, so the agent can’t look better by going quiet. The first suspect is below: a rule that exists but doesn’t load can’t prevent a correction.

  4. G4: Loading the right rule at the right moment is still weak.

    Measured retrieval recall is about 46%. That measurement is what removed the old loader hook, and it’s the next thing to improve.

  5. G5: Some things were retired when measured.

    Session chaining across the context window and the wrapper that relaunched sessions were built, measured and removed.

08 · What's next

Next

An open-source release under Apache-2.0, an install on a second agent in a different domain, and an English-only core. The measure for the release: a senior engineer who has never seen it understands it in three minutes and installs it in fifteen.