“Fully autonomous agents that decide everything on their own are an illusion. What actually drives agents in production systems is meticulously engineered context hygiene, thinking-budget model routing, and deterministic parallel slicing.”
When building multi-agent architectures today, common failure patterns are strikingly consistent: giving agents overly broad autonomy leads to infinite token-burning death spirals; accumulating tens of thousands of lines of logs and raw files causes cognitive context degradation; and uncontrolled parallelization creates chaotic merge conflicts.
To overcome these pitfalls and operate multi-agent systems on solid engineering foundations, I share our account/project-level orchestration architecture and production-proven practices.
1. Protecting the Main Context: “The Main Agent Only Conducts and Judges”
In an interactive session, the main agent is the conductor steering the entire project’s trajectory. Once the conductor’s context becomes polluted, judgment degrades across the board. Therefore, main context protection is the first principle of our system.
[Main Agent] (Senior: Judgment, Architecture, Synthesis)
│
├─► Delegation Check: "Do I personally need to inspect intermediate outputs to judge?"
│ ├─ No ──► Delegate to sub-agent (Summary/path contract)
│ └─ Yes ──► Execute directly in main (Targeted reading)
│
└─► Output Contract Enforcement
└─ Return only [Path / Diff summary / Verification results] (No raw file dumps)
Production Operating Norms
- Output Contract: When delegating to sub-agents, dumping raw file contents back into the main context is strictly prohibited. Large outputs exceeding 100 lines must be written directly to disk by the child, returning only
[File path / Change summary / Validation results]to the main session. - Proactive Bulk Read Interception: If the main agent anticipates reading more than three files or an aggregate of 400+ lines, it intercepts itself and dispatches an exploration sub-agent, receiving only the distilled conclusion.
- Strict Read/Write Turn Separation: Exploration (reading/diagnosing) and modification (writing) are never intermingled within a single execution turn; they are separated into distinct phases.
- 40% Handoff Hard Gate: When remaining session context drops below 40%, new tasks halt immediately. The agent synthesizes an indexed handoff prompt containing key progress and decisions, transitioning cleanly to a fresh session.
- L4 Historical Document Isolation: AI agents are barred from arbitrary searches across historical archives (
docs/reports/, legacy specs) unless explicitly requested by the user, avoiding context bloat and hallucination.
2. Model Tiering: “Allocated by ‘Thinking Budget’, Not by Raw Scale or Risk”
Assigning top-tier models simply because a task feels “important” is the primary driver of latency and cost blowups. Model tiers must be classified strictly by the quantity of reasoning required.
| Tier | Primary Role | Allocation Principle & Representative Models |
|---|---|---|
| Pinnacle | Complex trade-offs, hard architecture, high-stakes arbitration | Explicit invocation only (Never auto-assigned). Requires user prompt or escalation via [Return] |
| Advanced | Orchestration, directional decisions, debugging | Default tier for interactive main. Handles open-ended planning and synthesis (Claude Opus, GPT Sol) |
| Standard | Scoped execution, deterministic queries, repetitive implementation | Dedicated to sliced execution lanes. Code implementation, refactoring, batch updates (Claude Sonnet, DeepSeek Flash) |
Session Caps and Official Cross-Provider Pairs
- Session Cap: The main agent’s tier (Advanced) serves as the ceiling for that session. If an Advanced session encounters a Pinnacle dispatch, rather than throwing an escalation error, it automatically falls back to an Advanced model from the same provider to preserve session flow.
- Conductor ↔ Worker Cross-Provider Pairing:
While provider isolation (e.g., Claude main invoking Claude sub-agents) is typical, we codified an official cross-provider pair for optimal speed and cost:
- Conductor (Main):
Gemini 3.8 Flash(Rapid orchestration, massive context window) - Worker (Standard Routine):
DeepSeek 4.1 Flash(Cost-effective, lightning-fast code generation) - Registered in the
model-tiers.jsonrouting table with gateway filters, routine execution slices dispatched by a Gemini main are seamlessly routed via the Paseo daemon to DeepSeek workers.
- Conductor (Main):
3. The Sub-Agent Paradox: “The Problem Is Context Flooding, Not Missing Context”
A common anxiety when spawning sub-agents is: “What if it lacks the background knowledge to do the job?” However, in modern CLI and runtime environments like Claude Code or OpenCode, sub-agents already start by carrying a massive backpack containing tens of thousands of tokens (system prompts, global rules, project configurations, and dozens of MCP tool schemas).
[Sub-Agent Entry Context]
┌──────────────────────────────────────────────┐
│ Runtime Prompts + Global Rules (Tens of K tokens) │ ──► Attention Dilution
├──────────────────────────────────────────────┤
│ 🔍 Briefing Role: Selective Focusing Lens │
│ - Exact target files (write_set) │
│ - One-line quality anchor │ ──► Cuts through noise to focus on the task
│ - Strict output contract │
└──────────────────────────────────────────────┘
The essence of a sub-agent briefing is not injecting more data, but serving as a focusing lens that cuts away peripheral noise so the model can concentrate on its specific target.
The Three [Return] Protocols Against Runaway Agents
To prevent stuck sub-agents from hallucinating forced workarounds, they are mandated to halt immediately and signal a [Return] under three conditions:
- ① Unspecified policy or architectural decisions emerge
- ② The same failure repeats twice without new evidence
- ③ Execution time exceeds 1.5x of the estimate (excluding external I/O waits)
A
[Return]is not a failure; it is a structured signal requesting re-classification or specification from the main agent. This single mechanism eliminates 99% of wasted tokens and erratic code edits.
4. Parallel Delegation and Slicing: same tree + write_set disjoint
Spamming git worktree instances across multiple parallel agents frequently halts pipelines due to branch merge conflicts, duplicate dependency installs, and local port collisions. Production requires the simplest, lightest isolation model.
[Task Slicing & Parallel Dispatch]
Main: Global planning & slicing (Computed exactly once)
│
├─► [Wave 1] Lock upstream dependencies (Shared interface/type contracts)
│
└─► [Wave 2] Parallel dispatch of independent execution (Simultaneous invocation)
├─ Lane A: auth/login.ts (write_set A)
└─ Lane B: profile/user.ts (write_set B)
* Shared bottlenecks (package.json, etc.) reserved for a single lane
Core Principles of Parallelization
- Disjoint Write Sets in a Single Tree (
write_set disjoint): Within the same working tree, target file lists (write_set) across agents must never overlap. Worktrees are reserved strictly for explicit user requirements. - Upfront Slicing Is the Highest-Yield Investment: Computing how to divide tasks happens in the main context exactly once per instruction. Locking down contracts upfront ensures sub-agents finish without collision, bringing total rollback cost to zero.
- Single-Lane Ownership of Shared Bottlenecks:
Congested global files like
package.json, global router configurations, and core stylesheets must be assigned exclusively to a single lane. Shared types are finalized before parallel waves begin. - Time Budget Controls: Total sequential duration is bounded upfront. Long chains exceeding half a day (4 hours) are broken into smaller phases. Any single slice running past 1.5x of its estimate is immediately halted for user realignment.
5. Conclusion: From Prompt Engineering to “Systems Engineering”
The success of a multi-agent system hinges less on the raw ‘intelligence’ of any individual model, and far more on how effectively the system isolates context and failure.
- Keep the main context immaculate via output contracts and handoff gates,
- Route models based on thinking budget rather than perceived prestige,
- Lay down rails for sub-agents with disjoint write sets and strict
[Return]protocols.
When these engineering norms are codified into system configurations (~/.agents/) and repository rules (AGENTS.md), AI agents cease to be capricious toys and begin operating like a dependable junior/mid-level engineering team.