Files
pi-agent-config/docs/gateway-patterns.md
T
Kai cbba8faabc docs: pi 0.84.3 runtime mechanics, isolation baseline, personality layering, gateway patterns
Establishes this repository as the authoritative source for Pi agent
configuration across scenarios, starting with the documentation layer.

Key verified findings (probe harness included, zero model tokens):

- The skills section of the system prompt is emitted only when an active tool
  named 'read' exists (system-prompt.js:59,113). Therefore --no-tools silently
  makes every SKILL.md unreachable and --skill a no-op.
- registerTool accepts a plain JSON Schema object, so tool definitions can be
  served from a backend instead of duplicated in TypeScript.
- An extension can shadow a built-in tool by name, which is how a dedicated
  agent gets a path-restricted 'read' while still satisfying the rule above.
- .pi/SYSTEM.md replaces pi's coding-assistant prompt, but the replacement
  branch contributes neither the tool list nor the guidelines.
- Without --no-skills/--no-extensions, user-global resources leak into every
  scenario; probed leak was find-skills, modsearch, summarize.

Measured effect of the full baseline: system prompt 2619 -> 960 characters,
coding-assistant framing and pi-docs paths removed, skill finally reachable.

Secrets are guarded by scripts/verify-no-secrets.sh, installed as a pre-commit
hook. Backups deliberately live outside the repository.
2026-08-26 22:47:53 -07:00

9.2 KiB
Raw Blame History

Gateway Patterns

Patterns for hosting a Pi agent behind a long-running service (Telegram, HTTP, queue). Distilled from the memo-inbox scenario, which has run these in production since 2026-07, and from the defects found in the curator scenario, which reinvented the same problems worse.

Mechanism references: pi-runtime-notes.md §§1318.


Pattern 1 — One long-lived RPC process, not one process per message

Do: start pi --mode rpc once, drive it turn by turn over stdin.

Don't: spawn pi --print per message.

Per-message spawning pays node startup, resource discovery, skill loading and session deserialization on every single turn. The curator scenario spawned 24 processes per Telegram message; with a 120 s per-invocation timeout its worst case was ~480 s of model time for one message, with no overall budget.

An RPC process additionally gives you abort, set_model, set_thinking_level, compact and get_session_stats at runtime — none of which are reachable from --print.

Framing is strict JSONL

Split on \n only, strip a trailing \r. Do not use a generic line reader: Node's readline also splits on U+2028/U+2029, which are legal inside JSON strings. The reference implementation:

buffer = b""
while chunk := await proc.stdout.read(65536):
    buffer += chunk
    while b"\n" in buffer:
        line, buffer = buffer.split(b"\n", 1)
        line = line.rstrip(b"\r")
        if not line:
            continue
        await events.put(json.loads(line))

Correlate by id, and understand what success means

Every command may carry an id; the matching response echoes it. But success: true only means accepted. Failures after acceptance never produce a second response for that id — they arrive as events. So a request is complete when you have seen both the response and the terminal event you care about (agent_settled, or message_end for a reply).


Pattern 2 — Rotate sessions on a counter

Pi has no session TTL, rotation or size cap, and auto-compaction only fires near the context window — with a 1.05 M-token model that is effectively never.

memo-inbox rotates after a fixed number of prompts:

PI_SESSION_ROTATE_AFTER_PROMPTS=24

The curator scenario did not. Its per-chat session reached 174 KB / 74 messages, growing ~4.7 KB per turn because every turn embedded the full backend fact payload of that turn. Latency, cost and timeout risk degrade monotonically.

Rotate on whichever bound you can measure: prompt count, session bytes, or get_session_stats().contextUsage.percent. Prefer starting a fresh --session-id suffix over relying on compact, so that history is auditable.


Pattern 3 — Derive write receipts from tool_execution_end, never from prose

This is the single most valuable pattern in memo-inbox. The gateway subscribes to tool completion events and extracts the receipt from the tool result:

elif (event.get("type") == "tool_execution_end"
      and event.get("toolName") in MUTATION_TOOLS):
    result = event.get("result") or {}
    receipt = "\n".join(b.get("text", "") for b in result.get("content") or [])

The model's prose is then a summary, not the record of what happened.

Compare the curator scenario, which asked the model to phrase the outcome and relied on prompt rules to keep the distinction. Observed production output when has_file was false:

《My Brilliant Career》4K 版已成功加入库中,但目前还没有 4K 文件。

"加入库中" reads as in the library while the item had only been added to the tracker. Prose guardrails did not hold. Deterministic receipts do.

Rule: anything a user might act on — created, added, deleted, scheduled — must be rendered by code from a tool result, not generated by the model.


Pattern 4 — Structured extraction via a terminating tool

For classification and extraction there is no need for session state, skills, or prose parsing. Define a tool whose parameters is the output schema:

{
  name: "emit_extraction",
  parameters: <generated JSON Schema>,
  constrainedSampling: { type: "json_schema", strict: "prefer" },
  async execute(_id, params) {
    return { content: [{ type: "text", text: "recorded" }],
             details: params, terminate: true };
  }
}

Run these calls with --no-session so they cannot pollute the conversational history, and at a lower thinking level than the conversational turn.

This removes an entire failure class. The curator scenario extracted JSON with re.compile(r"\{.*\}", re.DOTALL) — a greedy match from the first { to the last } — and reused one post-processor across three different schemas, which injected an items: [] key into every intent plan.


Pattern 5 — Separate the stateful and stateless roles

Role Session Skills Thinking Process
Conversation / answering persistent, rotated yes high long-lived RPC
Classification / extraction / synthesis --no-session no lowmedium short-lived --print --mode json

Sharing one session across both roles causes mode confusion: the curator session alternated "只输出 JSON,不回答问题" and "直接回答用户,不要输出 JSON" for 37 consecutive turns, and each turn carried the previous turn's full fact payload.


Pattern 6 — Project a fact pack; do not hand over raw backend responses

Everything you put in a prompt is context you pay for and surface you expose. Project backend results through an explicit allowlist before they reach the model.

Observed in a single production curator prompt: filesystem paths (/mnt/truenas/multimedia/tv/…, /mnt/unRaid/tv4k/…), size_on_disk, quality_profile_id, internal database ids, a static capabilities block resent every turn, and an action_result.regular_matches object that duplicated verbatim an object already present under library.matches.

None of it was needed to answer "is there a 4K version".

Budget by estimated tokens, not characters — for CJK text the ratio is close to 1:1, so an 80 000-character cap is an 80 000-token cap.


Pattern 7 — Close the timeout budget, and kill the process group

Per-invocation timeouts do not compose. Define one deadline per user message, allocate it across stages, and enforce it with RPC abort rather than by killing the process.

When you do have to kill: pi is a node CLI that may spawn children. Launch with start_new_session=True and terminate with killpg, otherwise orphaned node processes keep consuming provider quota.


Pattern 8 — Keep provider and backend credentials out of the node process

subprocess inherits the parent environment by default. A gateway unit that carries *_API_KEY for Radarr, Sonarr, Plex, Telegram and a search provider hands all of them to pi and to every extension it loads.

Pass an explicit minimal env= (PATH, HOME, LANG, plus the few PI_*/bridge variables the extension genuinely needs). The provider key belongs in ~/.pi/agent/models.json, which pi reads itself.

Both scenarios in this repository currently inherit the full environment. This is the highest-value shared fix.


Pattern 9 — Bridge host capabilities over loopback HTTP

There is no RPC command to inject a tool result from the host, so host-side capabilities must be exposed as extension tools that call back out. memo-inbox does this for document parsing (127.0.0.1:8090) and OCR.

Recommended shape:

  • backend listens on 127.0.0.1 only, with a per-start shared secret passed to the extension through env;
  • backend serves both the tool definitions (JSON Schema) and the tool invocations, so the schema has one source of truth;
  • the extension is a generic proxy — registerTool accepts a plain JSON Schema object (verified), so tool definitions need not be duplicated in TypeScript;
  • every tool truncates its own output before returning.

Pattern 10 — Observability is already in the stream

Do not build separate instrumentation. message_update carries usage (input / output / cacheRead / cacheWrite / cost) and get_session_stats returns totals plus contextUsage. Persist per-turn model, thinking level, latency, tokens, cost and fallback into the service's own event ledger.

The curator scenario discarded stderr on success and kept only the last line on failure, so "how often does the primary model time out" was answerable only by grepping journald for a Chinese notice string.


Conformance

Pattern memo-inbox curator (as found)
1 long-lived RPC yes no — 24 spawns/message
2 session rotation yes — 24 prompts no — unbounded
3 deterministic receipts yes no — model prose, observed failure
4 terminating-tool schema n/a (tools are the output) no — greedy regex
5 role separation n/a (single role) no — shared session
6 fact-pack projection n/a no — raw responses
7 closed timeout budget partial (600 s event wait) no
8 minimal env no no
9 loopback bridge yes n/a (no tools at all)
10 usage observability not persisted not persisted

Shared implementations live in shared/lib/py/ and shared/extensions/.