.pi/extensions/_shared/ is installed by deploy-scenario.sh from shared/extensions/
and is not in the scenario's workspace/, so pi-diff listed it as merely
"UNTRACKED in live workspace" -- no comparison at all. The file most likely to
drift, because it is the one shared between scenarios, was the one the diff had
nothing to say about.
Now compared against shared/extensions/ and reported like any other file, with a
missing source called out separately.
Also fixes the counter: the new branch incremented a variable named `changed`
while the summary reads `DIFFS`, so a detected difference printed a diff and then
claimed the installation agreed, and exited 0. Verified by planting drift: exit 1
and "1 difference(s) found", then 0 after redeploying.
Two findings that changed the plan rather than confirming it:
23. Skills require a tool literally named `read`. Curator's tools are all
domain-specific, so every --skill argument was discarded in silence. The
planned split into curator-core / video-arr / books-ingest was inert before
it was written; the policy stays in APPEND_SYSTEM.md. memo-inbox is
unaffected because it registers a restricted `read` override, which is why
the earlier note generalised wrongly from it.
24. A long-lived session is worth far more than the startup it saves: 99.97% of
input read from cache on a continuing conversation against 0% on a new one.
That is what makes the generated tool list necessary rather than merely
tidy -- anything varying at the front of the prompt destroys it -- and it
makes rotation a cost to be bounded rather than applied eagerly.
profile.toml now describes the phase-3 configuration that is actually deployed,
including that the empty `skills` list is a finding and not an oversight.
pi_rpc gains --system-prompt support and no longer guesses whether a `read` tool
will exist; extension_registers_read has to be stated.
harness-layering.md records what transfers from a widely-shared account of
building a personal coding harness on pi, and what does not. The layering frame
holds and the cache-hit figure was the useful part. Its central recommendation --
installing third-party packages -- is disqualifying for an unattended agent
holding tracker credentials, and its discipline layer (AGENTS.md) is precisely
what we block, because it is discovered from every parent directory.
curator-tools.ts registers no schema of its own: it fetches the specs from the
backend bridge, so contracts.py stays the single owner and there is no TypeScript
copy to drift. It refuses to activate without a bridge URL and token, because an
agent that silently loses its tools still answers -- from the model's memory of
what a media library might contain.
deploy-scenario.sh now vendors listed shared/extensions modules into
.pi/extensions/_shared/. A tracked extension importing from shared/ cannot
resolve that path once installed outside the repository, so the deployed tree has
to be self-contained; this overwrites rather than merges, keeping the repository
authoritative. common.sh gains toml_list, using tomllib rather than more awk
because an array can span lines or carry comments.
verify-generated.sh checks that generated regions in tracked prompts match the
backend that generates them, and is wired into the pre-commit hook. This is
needed because of finding 22 below: the tool list has to be copied into the
prompt, and a copy drifts silently.
Two findings recorded in docs/pi-runtime-notes.md, both measured:
21. An extension that fails to import is silent -- exit 0, empty stderr, no
tools. A missing --extension path exits 1 with a clear message, but a
module that throws while loading reports nothing. The agent then invented a
complete library listing with plausible episode counts, quality and size.
A later identical run said it had no data instead, so the failure is both
silent and inconsistent.
22. --system-prompt suppresses the tool list. The customPrompt branch returns
before toolsList and guidelines are built, so promptSnippet and
promptGuidelines are inert. The tools stay callable over the provider API,
so tool use becomes a coin flip: one run in four looked at the library and
three said they had not been given any results.
SYSTEM.phase3.md is staged alongside the deployed SYSTEM.md rather than replacing
it: profile.toml still describes the phase-0 configuration that is actually
running, and the live service is untouched.
Records the risk policy as decided -- low_write only, high_write and destructive
refused outright with no confirmation flow, unclassified actions defaulting to
destructive so a missing classification fails closed.
Also records the measured fact-pack leak (filesystem path, quality profile id,
internal row id and a raw byte count all reaching the model), and the three
problems found while building it: the sqlite3 context manager not closing
connections, the five call sites writing a status the new CHECK constraint
rejects, and fallback_answer maintaining a diverged second copy of the receipt.
Production database migrated 0 -> 3 with row counts unchanged and both integrity
checks clean. Records the two problems found while building it: the newer-schema
guard was unreachable as first written, and create_control_plan -- the idempotency
gate for every write -- was check-then-insert, so a duplicate request surfaced as
an IntegrityError rather than as "already planned".
Phase 0 is done: commits 6df2f20, 1580504, 0722f1a, 643cdd5 in the
server-management repository, tests 37 -> 58.
Two items were not implemented as the plan specified, and the reasons are
recorded rather than left as silent gaps.
The plan called for .pi/SYSTEM.md to list seven tools. Phase 0's agent genuinely
has none, and naming tools that do not exist invites the model to call them, so
the deployed file states the absence instead. The tool-bearing specification is
kept in §4b.
The plan called for CURATOR_HOST to be narrowed to a specific LAN address and for
IPAddressDeny=any. Neither is safe here: the journal shows real traffic from both
the LAN and 127.0.0.1, so narrowing the bind address breaks one of them, and the
main service needs egress to Telegram and the model provider, which
IPAddressDeny would cut. The real fix is authentication, already tracked as P2-6.
The workspace now holds .pi/SYSTEM.md and .pi/APPEND_SYSTEM.md and nothing else.
SYSTEM.md is rewritten for what is actually deployed. The version committed in
07dd648 described five tools that will not exist until phase 3; shipping it would
have invited the model to call tools it does not have. The capability section now
states plainly that the agent has no tools and that every fact arrives in the
request. The phase-3 target, including the full tool-bearing launch contract, is
recorded in the plan as §4b together with why each part cannot be enabled sooner.
profile.toml likewise describes the deployed configuration rather than the target,
so that deploy-scenario.sh validates against reality and the path check means
something.
Recovered from the retired SKILL.md and folded into SYSTEM.md: the rule that the
current request's schema and length limits override everything else, and that a
JSON task returns exactly one JSON value with no fences. Phase 0's four prompt
types all depend on it, and it was the one part of that file not already covered.
Measured before and after on the real workspace, with flags read from the code
rather than transcribed (docs/evidence/2026-08-27-curator-phase0-prompt.md):
- expert coding assistant framing: present -> gone
- pointer to pi's own documentation: present -> gone
- the 64-line media policy: absent -> present
- workspace AGENTS.md: loaded -> blocked
- parent-directory AGENTS.md: LEAKED -> blocked
- <available_skills>: absent both times
Two things this confirms on the production configuration rather than a synthetic
probe. --skill was genuinely a no-op: it pointed at a real 64-line SKILL.md and
the skills block was still absent, because pi emits it only when a tool named
read is active and --no-tools deactivates everything. And --no-context-files is
the only switch that stops parent-directory pollution: a marker planted in
/home/claw/pi-workspaces/AGENTS.md reached the prompt before and not after.
Deleting the now-dead AGENTS.md and SKILL.md from the workspace changed the
prompt length by zero bytes, which is the proof that they were dead.
README corrected. Its table cited 960 characters as Curator's system prompt after
the change; that figure came from a few-line stub SYSTEM.md in the isolation
probe, and the real prompt is 3539 -- larger, not smaller. Presenting the stub
measurement as Curator's was misleading, and "72% smaller" was wrong. The prompt
grew because roughly 1.9 KB of pi scaffolding was replaced by domain policy that
had never loaded at all. The mechanism claim is unaffected.
README leads with the finding that motivated the repository -- pi emits the
skills section only when a tool named 'read' is active, so Curator's
--no-tools --skill combination made its policy unreachable -- with the measured
before/after table and instructions to reproduce it at zero token cost.
AGENTS.md sets seven rules for anyone changing this repository. The third is the
one that matters most: verify pi's behaviour with a probe rather than inferring it
from the docs. Three claims in the first draft of these documents were wrong and
were only corrected by running one.
The _template scenario carries the isolation defaults and inline warnings at the
places where mistakes have already cost time: --no-tools disabling the skills
mechanism, cwd anchoring .pi discovery, the read override being mandatory rather
than optional, and allowed-tools frontmatter not being enforced in 0.84.3.
Scenarios
- memo-inbox: mirrored by copying; the live directory was not moved or modified
and the service was not restarted. All four tracked files match byte for byte
(pi-diff.sh reports SAME). Marked deploy = "mirror" so deploy-scenario.sh
refuses --apply: applying a mirror would invert the direction of truth and
could change a service in daily use.
- curator: target configuration, not yet deployed. .pi/SYSTEM.md replaces pi's
coding-assistant prompt; durable role text is in .pi/APPEND_SYSTEM.md;
profile.toml is the single source of truth for the launch contract.
- pi-grok: registered only. It is genuinely a coding agent, so the isolation
baseline does not apply in full.
Corrections to the documentation, found by testing rather than by reading
- AGENTS.override.md does NOT block parent-directory context files; it only
shadows its own directory. Verified: with an override file in the workspace, a
marker in /tmp/AGENTS.md still reached the system prompt. The only effective
switch is --no-context-files, so durable role text must live in
.pi/APPEND_SYSTEM.md, which is a system-prompt file and unaffected by -nc.
Verified end state: no coding-assistant framing, no pi-docs block, own
identity and role text present, no parent pollution, only own skills/tools.
- PI_CODING_AGENT_DIR isolates settings/models/auth/trust/extensions/skills/
prompts/themes under the agent directory -- stronger than the --no-* flags
because it also repoints credentials -- but does NOT cover ~/.agents/skills.
Measured: find-skills, modsearch and summarize still leak. So it complements
--no-skills rather than replacing it.
- --append-system-prompt accepts a file path, which pi-grok relies on.
- cwd is what anchors .pi discovery: a probe that forgot cwd silently lost
.pi/SYSTEM.md and kept the coding-assistant persona.
Tooling (all dry-run by default; none of them restarts a service)
- pi-diff.sh: compares tracked config against the live install in both
directions, with a key-redacted comparison for models.json
- deploy-scenario.sh: installs a workspace and renders profile.toml into
.pi/launch.json, then checks that every referenced path exists
- deploy-runtime.sh: renders models.json from its template, refusing placeholder
or missing keys. Verified byte-identical to the live file
- pi-backup.sh / pi-restore.sh: archives outside the repo, sha256 manifest
verified before any restore, live paths preserved rather than overwritten
Fixed while testing: pi-backup.sh compared the destination against the repo root
literally, so a relative --dest ./backups wrote credential archives into the work
tree. Now canonicalised with realpath; ./backups, an absolute in-repo path and
./docs/../backups are all refused.
Extracted from pi-workspaces/memo-inbox/.pi/extensions/memo-guard.ts, which has
enforced these patterns in production since 2026-07.
Exports:
- inside() / safeRealPath() / makePathResolver(): path containment that resolves
symlinks before checking, so a link inside an allowed root cannot escape it
- registerRestrictedRead(): a path-restricted 'read' that shadows the built-in.
Required rather than optional: pi emits the skills block only when a tool named
'read' is active and skill bodies load through it, while the built-in 'read'
accepts absolute paths and could reach the service's credential files
- installGuard(): the two capability layers, setActiveTools plus a tool_call
block, re-asserted on resources_discover as well as session_start
- truncate(): byte-aware truncation ahead of pi's 50 KB / 2000 line caps
- registerBridgeTools() / fetchBridgeSpecs(): loopback HTTP bridge, with the
baseUrl asserted to be loopback. Since registerTool accepts a plain JSON
Schema object, the backend can own the schema instead of a drifting copy
Verified against a real pi process with zero model tokens
(shared/extensions/tests/run-guard-checks.sh, 14 assertions):
active tools are exactly the declared set, the read override wins with
source=cli, the skills section is present and contains only the scenario's own
skill, and reads of an outside file, a ../ traversal and an absolute path to
~/.config/curator/curator.env are all denied.
The deny-path fixture is named .txt and renamed to .env only inside the temp
work directory, because verify-no-secrets.sh correctly refused to track a file
called *.env.sample.
Generalises PiRPC from pi-workspaces/memo-inbox/telegram-gateway/gateway.py,
which has run this pattern in production since 2026-07, and closes the four gaps
both existing scenarios shared:
- explicit minimal env, so provider and backend API keys never reach the node
process (verified: 6 variables, an injected secret is withheld)
- start_new_session plus killpg on stop, so a stuck node tree cannot outlive the
turn (verified: no orphan after stop)
- the loading-isolation flags are part of the launch contract instead of
something each caller has to remember
- a per-turn deadline enforced with RPC abort rather than by killing the process
Retains the original's proven mechanics: strict newline-only JSONL framing,
correlation by id, agent_settled as terminal event, and receipts harvested from
tool_execution_end rather than from model prose.
PiLaunchConfig warns when skills are configured but no 'read' tool can be
active, which is exactly the condition that silently disabled Curator's SKILL.md.
Includes a zero-token smoke test: it drives a real pi process with get_state
only, so no model call is billed.
Mirrors ~/.pi/agent/ as the authoritative copy. models.json becomes
models.json.template with ${ZENMUX_API_KEY} substituted; the real value stays
in secrets/zenmux.env, which is untracked and enforced by the pre-commit guard.
Excluded with rationale: auth.json, trust.json, models-store.json, sessions/,
herdr-agent-state.ts (installer-managed, overwritten on reinstall) and the
third-party skills under ~/.agents/skills.
Recorded during migration: the configured fallback model zenmux/x-ai/grok-4.6 is
absent from models.json, so pi falls back to an undeclared custom model id with
no context window, cost table or thinkingLevelMap. Fixing that is a behaviour
change and is deferred rather than folded into this zero-change migration.
Two defects found by testing the guard against itself:
1. When invoked through the .git/hooks/pre-commit symlink, deriving the repo
root from dirname(BASH_SOURCE)/.. resolved to .git/ instead of the work
tree, so the hook scanned nothing and never blocked. Use
'git rev-parse --show-toplevel' instead.
2. The blanket secrets/ rule rejected secrets/.gitkeep. Replaced with an
explicit allowlist: .gitkeep, README.md, *.example, *.template.
Verified: a staged file containing a Telegram bot token now aborts the commit
and leaves HEAD unchanged.
Establishes this repository as the authoritative source for Pi agent
configuration across scenarios, starting with the documentation layer.
Key verified findings (probe harness included, zero model tokens):
- The skills section of the system prompt is emitted only when an active tool
named 'read' exists (system-prompt.js:59,113). Therefore --no-tools silently
makes every SKILL.md unreachable and --skill a no-op.
- registerTool accepts a plain JSON Schema object, so tool definitions can be
served from a backend instead of duplicated in TypeScript.
- An extension can shadow a built-in tool by name, which is how a dedicated
agent gets a path-restricted 'read' while still satisfying the rule above.
- .pi/SYSTEM.md replaces pi's coding-assistant prompt, but the replacement
branch contributes neither the tool list nor the guidelines.
- Without --no-skills/--no-extensions, user-global resources leak into every
scenario; probed leak was find-skills, modsearch, summarize.
Measured effect of the full baseline: system prompt 2619 -> 960 characters,
coding-assistant framing and pi-docs paths removed, skill finally reachable.
Secrets are guarded by scripts/verify-no-secrets.sh, installed as a pre-commit
hook. Backups deliberately live outside the repository.