Every session started from zero. I wanted it to start from files.
I run Claude Code agents every day. Out of the box, each session forgets how I want it to work, and a lot of the steering lives in the prompt: check your work, don’t delete things, ask before the risky calls. So every session began with re-teaching, and the agent’s “done” was a claim I still had to check.
A fresh-context grader I built later made the second part concrete: it re-checked the agent’s own “done” claims and found 7 of 12 over-claimed.
GoalA fresh session, running on its instructions and mechanisms alone, behaves the way I’d want without in-session steering, and proves its work instead of asserting it.
A generic core, a thin agent config, and files as the bus
TarkOS is two layers. The core holds the mechanisms and is generic: it ships empty of content and knows nothing about my domain. The host agent keeps its identity, mission and its own rules. The host’s instruction file joins them with two @-import lines, so the core text is never pasted in.
An autonomous engine loads only in away mode. Everything talks through git-backed files (tasks, decisions, project state, a capture ledger) rather than function calls, which keeps every step auditable and revertible. TarkOS Cockpit is a desktop surface over those same files.
Most of the code in TarkOS, and in the tools around it, is written by AI agents under my specification, test contracts and review of the result. My part is the architecture, the decisions and the taste. That the agents can do the rest reliably is the thing TarkOS exists to prove.
The install manifest lists 9 mechanisms. The boundary between core and host is enforced by instruments, not by judgement: a leak suite for identity tokens, a reference checker that allows shipped docs to point only at shipped components, and a frontmatter ratchet.
Five mechanisms, each with an artifact you can open
-
M1: Posture
Five always-on points: be proactive, decide what’s yours to decide, run autonomously, evolve from corrections, and report simply. When I’m away, the agent logs each estimate as an assumption and opens the next session with an assumptions round.
-
M2: Decision tray
Reversible and confident: do it, then leave a record that one command can undo. Irreversible or unsure: a decision card with options, consequences and a recommendation, which I ratify when I choose. The agent never ratifies its own card.
-
M3: Capture
When I correct the agent, the correction is written into the file that owns it, with a ledger entry. The next session loads the rule instead of repeating the mistake.
-
M4: Verification gates
Verdicts come from output, never from exit codes. A fix carries evidence that the test failed before and passes after. A done-claim whose evidence is older than the branch tip fails.
-
M5: Autonomous engine
An orchestrator dispatches task cards to at most two workers, logs every dispatch and acceptance, and ends the run with one morning card: what changed, what waits for me, what was measured, what stays unverified.
Installing on a new agent starts with a read-only doctor. It reports what would collide or is missing, and it never deletes anything:
Decisions that wait for me show up in TarkOS Cockpit with the agent’s recommendation and its stated confidence:
Five calls that shaped it
Route every decision by confidence and reversibility
- Alternatives
- Ask about everything
- Maximum aggression
- Let the agent tune its own thresholds
- Why
Autonomy grows by shrinking irreversibility. High confidence and reversible: act, then keep or revert. Medium: act and flag it. Low confidence or irreversible: always ask. Autonomy levels are per category; promotion is only offered after 90% prediction accuracy over at least 10 decisions, demotion below 75% is automatic, and money or anything irreversible never moves.
- Trade-off
Some decisions wait for me that the agent would have got right. That’s cheaper than undoing an irreversible one.
- Outcome
The first live run: 14 do-then-confirm records, all kept, none reverted. Later the agent caught its own noise filter hiding 468 of 468 records from me.
Two import lines instead of a hook
- Alternatives
- SessionStart hook injection
- A regex prompt manifest
- Paste the core into the instruction file
- Why
The core first loaded through a SessionStart hook, so subagents got it too. It cost latency, and the prompt manifest only ever saw my prompt, not what the task needed later: all 4 of its measured injections went unused. Two
@-import lines are visible in/contextand need no hooks.- Trade-off
Hosts that strip the instruction file need the optional hook loader.
- Outcome
Hooks stay where they’re the right tool: a rule that keeps being violated gets a mechanism that enforces it, never louder prose.
A measured zero-base rebuild: delete beats reorganise
- Alternatives
- Keep adding rules
- Reorganise the files
- Start a new harness
- Why
A measurement tool showed a lot of always-on context, routing split across two systems, and 33 of 47 skills that never fired. Every cut went to a staging folder with a manifest, a seven-day watch and a one-commit revert, and a floor of invariants could not be simplified.
- Trade-off
Slower than deleting outright. One earlier session misread six live mechanisms as dead and deleted one; since then the rule is to report, not cut, wherever no regression gate exists.
- Outcome
Always-on context 82 KB → 44.3 KB. The ten largest files 726 KB → 253 KB, with no section headings lost.
A generic core and a thin per-agent config
- Alternatives
- Publish v1 without the identity-laden scripts
- Publish as it is
- Why
Mechanisms go to TarkOS; identity stays with the agent. The test for every rule: would a stranger’s agent in another domain need it? The package is generated from the live repo, not maintained by hand.
- Trade-off
The agent recommended publishing a smaller v1 first. I overrode it and chose to generalise first, which delayed the release.
- Outcome
Identity leaks in shipped files went from 153 to 0. The export is being rebuilt for release and is still unpublished.
Rebuild the engine as an orchestrator with workers
- Alternatives
- Rebuild but keep the worker relay
- Not now
- Why
The engine had grown to 15 files and 62 mechanisms, and its reference docs had zero reads across 2,405 telemetry rows. I noticed I was babysitting runs that were supposed to run without me.
- Trade-off
It dropped session chaining across the context window, which had taken weeks to build.
- Outcome
Worker-slot use went from 30% to 94% on the next night run.
Measure the instructions, and the instruments
Arm A gets only the rules under test, arm B gets none. Same model, questions and criteria written before the run, a blind pairwise judge plus keyword checks. The same split held on a second model.
One batched judge instead of ten per-question judges, with 8 of 10 verdicts agreeing. Both disagreements were cautious ties.
The agent’s “done” claims that a fresh-context grader found over-claimed. It’s why verdicts come from output.
The agent found that its own autonomy scorer counted the harness’s injected messages as my prompts. Fixing it was a measurement correction, not an improvement, and a regression test now guards it.
What an overnight run actually looks like
Recorded from a real session while I was asleep. The agent scored my launcher’s publishability gates on a fresh export and found two of them red: 7 of 14. Instead of just reporting it, it widened the fix, wrote a failing test first, fixed the bug and the identity hits, and handed the result to a worker with zero context for a re-score: 10 of 14. One verification run went red with no failing test. It was recorded as a flake and handed on as the next task, not hidden.
Other sessions are documented only in their verification reports. One example is a launcher performance fix that took frame-time p95 at four panes from 764.7 ms to 31.0 ms with the same harness, after the agent first caught its own measurement error.
What isn’t done yet
-
G1: It isn't public yet.
The open-source export exists, but it is being rebuilt for release: the harness is frozen at a stable version while the export is reworked.
-
G2: "Portable" is proven on one install.
My own agent is the only production install so far. A second agent is the real test.
-
G3: The main goal isn't moving yet.
The measure I care about is fewer corrections per session. It’s going the wrong way: my corrections per 50 sessions sat between 60 and 77 in earlier stretches and reached 153 in the latest one, while the ratio of problems the agent reports itself to problems I catch fell to 0.85. I track the two together on purpose, so the agent can’t look better by going quiet. The first suspect is below: a rule that exists but doesn’t load can’t prevent a correction.
-
G4: Loading the right rule at the right moment is still weak.
Measured retrieval recall is about 46%. That measurement is what removed the old loader hook, and it’s the next thing to improve.
-
G5: Some things were retired when measured.
Session chaining across the context window and the wrapper that relaunched sessions were built, measured and removed.
Next
An open-source release under Apache-2.0, an install on a second agent in a different domain, and an English-only core. The measure for the release: a senior engineer who has never seen it understands it in three minutes and installs it in fifteen.