Sovereign Life Hub
The system, demonstrated in geometric order.
A governance substrate that ships as a life-management product, described here the way it was actually built: from a handful of definitions and axioms, through propositions the system enforces on itself, to the open questions it has not yet closed. Every count below was read live from the Sovereignhub MCP at grounding time. None was recalled from memory.
“An emotion, which is a passion, ceases to be a passion, as soon as we form a clear and distinct idea thereof.”
On the last day of February 2026, mid-match, the founder caught a familiar spike of fear before it finished firing, named it, and watched it lose its grip on his attention. Spinoza wrote the proposition in 1677. The insight that night was narrower and more useful: the proposition compiles. Noticing is not preparation for the intervention. Noticing is the intervention, and a thing that works once under pressure can be made to work on a schedule.
Sovereign is that observation given infrastructure. It captures the raw material of a life (thoughts, spending, signals, skills, incidents), forces the resulting beliefs through a hygiene discipline before they harden, and watches its own operation with the same instruments it points at everything else. What follows is a portrait of that machine as it stood on July 1: its vocabulary, its invariants, the propositions it enforces on itself, and, kept deliberately at the end rather than swept under the prose, what it does not yet know about itself.
Definitions
The system speaks at three altitudes, and most confusion about what Sovereign is comes from mixing them. So the vocabulary comes first, each term pinned to the register where it is true. Substrate language is wrong in product copy; marketing language is wrong when cited as architecture. The altitude distinction is load-bearing, not stylistic.
agent_traces, chained by parent_trace_id so any dispatch tree can be reconstructed after the fact. Token counts are mandatory fields, not optional metadata, including for the expensive lead sessions.The governing thesis is that execution-only platforms drift, because the gap between performing an action and governing that action is where real work fails. Competing tools solve execution and bolt governance on afterward, and the bolt-on rarely holds, because governance has to be designed in, not retrofitted. Sovereign was built governance-first: the orchestration layer, the multi-tier agent architecture, and the trace-and-review discipline are the primary differentiator, not the surface chrome.
The founder-is-the-proof pillar follows from the same thesis. The founder operates Sovereign at a real cadence, running engineering and personal-life governance in parallel, and when that cadence surfaces a structural failure, the substrate adapts. The documented example is a worktree-workflow drift incident: the failure surfaced a gap, the founder articulated a reframe, and the substrate updated to make the underlying condition visible as ambient signal rather than settling for a one-off fix. Cadence surfaces the issue, the issue generates a governance update, the update closes the gap. That loop is the substrate working as designed.∎
v3, the boardroom pivot (2026-04-26). Under v2, Sovereign was a metacognitive life OS reached primarily through chat, supplemented by dashboards and ambient capture. Under v3 it is a boardroom in the sense of Def. 2, and a first slice shipped 2026-06-10 as S9.5-BOARDROOM-V1: the existing /today route became the Command Center by accretion rather than by a new route.
v3.1, the SaaS focus (2026-06-10). The founder resolved a standing personal-versus-SaaS fork in favor of SaaS: the system is good enough for his own use as it stands, and his use case has converged with the customer's, so effort now goes toward serving other people. The product articulation, captured verbatim in governance: Sovereign's best current use case is long-term planning and a visual that keeps the models and the operator in cohesion. The dual customer is the operator and their agents together.
Axioms
Six invariants constrain everything else in this document. None of them arrived as theory. Each was ratified after a specific failure made its absence expensive, which is why every plate below carries a citation back to the precedent that holds it. They are axioms in the working sense: the propositions of Parts III through VII assume them without re-arguing them.
No durable claim about live state is made from memory. The claim is read at the state layer at the moment of authoring, and if the state layer is unreachable, the author halts loud rather than substituting recollection.
A role and the system it makes claims about are two halves of one operation: the role performs synthesis and disposition, the state layer holds the truth. Claim-level and task-level falsification checks are downstream safety nets; this is the architecture they sit on.
Settled Precedent #35 · Grounded-Reality InvariantThe mind that produced the work is never admitted as its judge. A different trace reviews, and for anything a human will see, the reviewer sees what the human sees: rendered output, not diffs.
Backend and frontend critics sit at the boundary between work and record. One screen per dispatch, browser-verified before handoff.
Precedents #2 · #24 · different-trace review · rendered verificationGovernance-touching work moves in two phases: an analysis returns to the operator for disposition before anything is built.
The rule exists because a single-phase dispatch once skipped its own review loop. Two-phase dispatch has since run successfully in parallel four-agent waves.
Precedent #29 · two-phase dispatchWhen a capture has a defined schema and more than one required field, the interface is a structured form. Freeform capture stays message-as-action, because there the text is itself the payload.
Reached after three spec iterations on the Discord money channel, each adding parser complexity without closing the root category error.
Precedent #33 · forms over parsingExternal surfaces never touch MCP. Every one of them routes through a scoped HTTP API enforcing surface identity, user identity, scope authorization, and source tagging.
MCP is the agent-trust surface only. Handler reuse between the scoped API and MCP is permitted under four explicit conditions, and any new external surface requires a published API spec before build.
Precedent #34 · external surfaces auth boundaryEvery dispatch resolves to a single register of decision: ops, design, infrastructure, product, data, or security. An agent that discovers a second axis mid-work stops and escalates.
Multi-register dispatches split agent attention, fragment the critic's register-specific checklist, and let cross-register violations slip through review undetected.
Precedent #5 · one decision axis per dispatchAn axiom is easiest to trust when you watch it catch a gap in its own corpus. Two files independently claim the #35 slot: the Grounded-Reality Invariant cited throughout this document, and a 2026-04-28 draft on the CEO session disposition cycle. Neither has ceded the number; the active file says so itself. Separately, governance/all-opus-mandate.md still carries active status while the 2026-06-10 version log already narrates its generalization into an apex-model-class mandate. That is a live instance of exactly the drift Axiom I exists to catch, not yet reconciled at the file-status layer. Both are carried into the Open Questions, unresolved on purpose, rather than silently tidied here.
On Method
Method here means one thing: the machine is arranged so that its own weaknesses of mind, and its agents' weaknesses, are structural problems with structural answers. Four tiers of authority, a critic gate at the boundary between work and record, and a trace chain that makes every dispatch auditable after the fact.
No mind, human or model, is admitted as the judge of its own work.
Axiom II, made mechanical. Every wave that produces work also produces a review by a different trace: a backend critic for structure and a frontend critic that must verify rendered output in a browser, not a code diff, before a SHIP verdict counts. The author and reviewer separation is enforced at dispatch, not requested in prose.∎
Authority descends by tier; evidence returns by trace, or it did not happen.
Tier 0 is the operator: strategic authority, outcomes, blast-radius constraints. Tier 1 is the CEO advisory layer, reached through chat and web, bound by Axiom I for every durable claim. Tier 2 is the Lead agent, which owns the plan, execute, gate cycle and is boot-gated by the middleware (Prop. III). Tier 3 is seven domain agents: shell, editor, ops, viz, mcp, data, auth, each held to a 40 percent context ceiling because sessions that cross it degrade and require rework, which compounds cost rather than saving it. Four auxiliary roles (Devil's Advocate pre- and post-sprint, disposable fix, research) are dispatched ad hoc by Lead outside the standard chain.
Every dispatch writes an immutable row to agent_traces, chained by parent_trace_id, with token counts as mandatory fields. The full dispatch tree of any piece of work can be reconstructed after the fact, which is the property everything else in this Part depends on.∎
A discipline that lives in prose will eventually be skipped. A discipline that lives in the middleware cannot be.
Dispatch and status-update tools are blocked at the middleware layer until a session calls get_active_context and get_roadmap_tasks; skipping the sequence returns a governance violation with a pointer to the source file. The gate is architectural, not aspirational, because the aspirational version already failed: an audit found eleven shipped deliverables with no roadmap task behind them and more than fifteen traces where the boot sequence had been silently skipped. The founding insight of the whole system applies to the system: what is not caught in the act simply continues.∎
A conclusion is entitled to exactly as much force as the evidence that would survive an attempt to kill it.
Every tier defaults to the apex model available. The rule hardened after a QA walkthrough found lower-tier frontend work failing in the browser despite green tests and a clean critic SHIP verdict: the canonical telling has 2,272 tests passing while 58 percent of components failed rendered QA, broken design-token fidelity, CSS specificity conflicts, missing responsive breakpoints. Tests and code review check structure, not aesthetics, so the failure was invisible until a human opened a browser. The resulting All-Opus mandate (Precedent #26) became the default for every role, with narrow exceptions: pure backend work with zero user-facing surface and full test coverage, a cost-controlled Sonnet path for briefing emails, and a dedicated Haiku 4.5 exception for the Insight Engine's structured synthesis. On 2026-06-10 the mandate generalized from a brand, Opus, to a capability class, apex-model, currently the Fable family, to be re-validated at every model transition.
Then the proposition was applied to the mandate itself. A 2026-07-02 review corrected the incident's attribution (the figures trace to an April walkthrough recorded in lesson 122, not to the Sprint 6 wave the corpus long blamed), found the lesson's internal counts inconsistent, and, when the true counts proved unrecoverable, marked the figures unverifiable at source. The largest cost line in the system was resting on a number nobody could reproduce. It now faces a designed falsification battery (GOV-MANDATE-FALSIFY) that must measure its own apex-class baseline rather than cite the folklore figure, with outcome semantics written before the run: a pass widens the exception list, a fail strengthens the mandate, and either way the challenge goes on the record.∎
Note what did not happen. The mandate was not revoked when its evidence degraded; it holds while the experiment is designed, because prudence and epistemics are different questions. What changed is its standing: from a fact everyone repeats to a hypothesis with a scheduled attempt on its life. That is the same treatment the lessons store applies to any belief (Part V), pointed at the system's most expensive one.
One state layer, reached one way, or there is no state at all: only competing recollections.
111 tools mediate all state, organized in 24 domain groups, up from 97 at the last dated registry snapshot (2026-04-25). Two entire groups shipped since that snapshot and did not exist in it: KB Topics (Scholar V2) and Workspace File I/O, both called out in the registry itself as not yet present when it was written. Agents reach state through MCP and nothing else; external surfaces reach it only through the scoped API of Axiom V. The taxonomy below is the system's entire vocabulary of action.∎
On Law
Sovereign's law is jurisprudence, not legislation. Practice rules first, and the statute is the record of what survived. That ordering is deliberate, and it produces two ordinary conditions a naive reading would call defects: rules that bind before they are filed, and files that stay active after reality has moved on. The corpus is candid about both.
The written law is a description of behavior that already proved itself; its lifecycle exists so the description can be corrected without being erased.
62 files across five categories, each carrying a status that moves through a fixed lifecycle. The lifecycle's proof that it works is a pair of clean supersessions: Precedent #33 (forms over parsing) and Precedent #34 (the auth boundary) each shipped a v1, ran it through review, then promoted a sharper v2 to active while the v1 flipped to retired. Same number, same principle, better text.∎
| Category | Count | What it holds |
|---|---|---|
| governance | 22 | Orientation, protocol, and policy documents: the substrate narrative, pivot declarations, session-start protocols, cost strategy. |
| specs | 25 | Feature and build specifications, from the Insight Engine design to individual TSGs. |
| precedents | 9 | Standalone precedent files. The numbered Settled Precedent series runs past #35; many earlier ones live directly in the root CLAUDE.md tiers rather than as their own governance-file row. |
| lessons | 5 | Durable incident postmortems (INC-001, INC-002) and audit-visibility lessons. Distinct from the 216-row lessons_learned store of Part V. |
| agent-rules | 1 | The MCP tool registry itself. |
The paperwork lags in both directions, and the lag is informative rather than embarrassing. A majority of the corpus, 31 of 62 files, sits in draft, including at least one the corpus itself calls an enforced rule in its own header (Precedent #5): draft-but-binding. In the other direction, all-opus-mandate.md carries active status while a later version-log entry already narrates its generalization: active-but-superseded. And the file that governs grounded claims about live state, Precedent #35, is itself an ungrounded claim about which file it is, since two files hold the number. A system that runs faster than its own record-keeping will always show this pattern. The failure mode would be pretending otherwise; the discipline is naming it, which is what the Open Questions footer is for.
On Knowledge
Two knowledge stores exist side by side and are easy to conflate. The lessons governance category of Part IV holds five durable incident postmortems. A separate lessons_learned table, reached through get_lessons, holds 216 maturity-scored entries across seven life domains. This Part is about the second store, and about the older idea it implements: that knowledge comes in kinds, ordered not by how strongly a belief is felt but by how it was earned.
The mind orders its ideas by provenance: what arrived through raw experience is held more loosely than what survived dialectic, and both more loosely than what checked out against the world.
The first kind of knowledge is capture. Thoughts, signals, journal fragments, lived experience: the richest input stream and the least trustworthy, taken raw. The scoring system weights lived experience below dialectic and outside research on purpose, in its own words, because lived experience is where cognitive distortions are most convincing, and outside research checks against a dataset larger than one life.
The second kind is the lesson. A captured belief becomes a lesson when it is versioned, domain-tagged, and scored, and the score is explicit about what it measures: epistemic hygiene, not truth. Age, challenge status, source diversity, evolution, and language, none of which require empirical verifiability. That design choice makes the store usable in domains where beliefs are unfalsifiable, by flagging untested beliefs rather than unscientific ones.∎
Eighty-one and a half percent of what this system knows is about itself. The old term for reflexive knowledge is idea ideae, the idea of an idea, and a young substrate mostly thinking about its own operation is exactly what a bootstrap phase should look like. But the ratio is a diagnostic, not a trophy: a life OS whose knowledge stays overwhelmingly self-referential as it matures would be a mirror studying its own glass. The share is therefore treated as a trend to watch, not a fact to admire.
A memory that cannot be challenged is not knowledge; it is exposure.
Of the 216 lessons, 11 currently carry a logged devil's-advocate challenge, the mechanism that raises a lesson's maturity by forcing a counter-argument onto the record. Staleness flags and configurable challenge intensity work the same seam. Prop. IV showed the mechanism running at full authority: the lesson under the system's largest cost line was challenged, found unverifiable, and re-marked, with the mandate's future now assigned to an experiment instead of a memory.∎
To notice is already to intervene: name the pattern, and the pattern loses bandwidth.
The epigraph of this document, made into a product surface. The Insight Engine, approved 2026-04-16, is the observation layer's primary primitive and the reason Sovereign is described as a metacognitive life OS rather than a productivity suite with governance bolted on. Ten trigger primitives (threshold_crossed, state_changed, milestone_hit, streak_started, streak_broken, first_in_window, record_broken, anomaly_detected, correlation_surfaced, inactivity_detected) exist as data rows rather than code, so roughly 500 entities can generate on the order of 5,000 potential insights without a new code path per insight.
The economics keep the noticing quiet: deterministic templates at $0 per user per month, a daily-briefing Haiku pass at roughly $0.36 under the Precedent #26 exception, moment-triggered Haiku at roughly $0.01, an aggregate-only proposer at roughly $0.06, totaling near $0.43 per user per month, about 2.9 percent of a $15 subscription. The render target is dual: a boardroom panel as the ambient recall surface (everything noticed, one place) and per-page bands as the precision surface (relevant to where you are standing). The register is fixed by design: the system notices, lowercase; it does not nag.∎
Build status, stated plainly. The 40-trigger content sprint shipped. The core engine (schema, ten handlers, evaluation cron) is currently blocked. UI surfacing, Haiku-mode integration, and five mockups are pending. Approved and partially built, currently stalled at the core-engine phase, is the accurate present-tense description.
One layer above the lessons sits an empty shelf. Scholar V2's topic layer is built, six dedicated MCP tools, none of which existed in the April registry, and list_kb_topics returned zero rows at grounding time. That is consistent with the roadmap rather than contradictory to it: S10-SCHOLAR, the task that would populate it under a Skills through Scholar through Notes through KB through Skills loop, is still status planned. The shelf was carpentered before the books arrived.
On Time and Persistence
“Everything, in so far as it is in itself, endeavours to persist in its own being.”
The sprint record is the system's conatus made visible: not persistence as stasis, but persistence as a steady increase in the power of acting, purchased by continuous self-revision. Fifty-four version events since genesis on 2026-03-02, and a numbering history with visible scar tissue, because a system that never renumbers is a system that never learned anything worth reorganizing around.
The system persists by revising itself; supersession is not its failure mode but its metabolism.
S1 through S8 shipped in sequence, from multi-tenant core through billing to AgentOps polish, each closing with its full task set done or done-minus-one-cut. S9, Growth and Launch, is the largest sprint tracked at 128 tasks: 94 done, 16 pending, 11 planned, 4 cut, 2 in review, 1 in progress, and the marker task S9-00 itself still reads in_progress. Most of the sprint has shipped; the tail, design-system enforcement, visual regression infrastructure, and the API wave, has not. S9.5, Boardroom v1, is the first physical embodiment of the v3 pivot and is effectively closed. S10, the Insight Engine, is approved and partially built with its core engine blocked, per Prop. IX. S11 through S13 are specced rather than built: the deploy manifest phase 1 analysis plus Insight Engine phases 4 and 5, a promoted API and Discord bridge spec with early build motion (2 of 8 tasks), and a finer-grained permission model for agent actions, currently in design.∎
Two housekeeping notes from reconciling live status against the corpus. The sprint numbering was reshuffled at least once: three retired placeholder tasks (OLD-S4, OLD-S5, OLD-S7) mark earlier definitions of billing, polish, and growth-and-launch sprints that were cut and renumbered. And an early S11-00 placeholder, Security Audit Team and MAESTRO Threat Model, still shows status planned alongside the active S11 deploy-manifest thread; this overview follows the active thread, consistent with the build brief's own framing. A record with no renumbering would mean nothing was learned mid-course. This one renumbered, kept the retired markers visible, and carried the ambiguity to the footer instead of deleting it.
On Others
“There is no individual thing in nature, which is more useful to man, than a man who lives in obedience to reason.”
Logged 2026-05-23 and locked as foundational vision rather than a near-term feature, this precedent exists so that future drift toward growth-optimized sharing gets recognized as drift, not as natural product evolution. It is the destination the definitions of Part I were always pointing at.
Two people who cannot read each other's internal weightings will each experience the other as unreasonable, and neither will be wrong about the experience.
Most friction between people is not disagreement about facts. It is two people operating from different internal weightings who cannot see each other's models, so behavior that is reasonable given one person's weights reads as difficult, selfish, or oblivious to the other. Sovereign's purpose beyond personal life management is to make internal weightings legible: first to oneself, which is the solo product today, eventually to trusted others, which is the unbuilt future state.
Why raw sharing fails, twice. A naive version of this looks like deeper social media, sharing thought patterns directly. Most people cannot articulate their own weights even to themselves, let alone hand them over raw, so the individual layer has to come first. And raw interiority is too vulnerable to share even with a trusted partner: privacy is the precondition for honest self-documentation, not an obstacle to it. Remove the privacy and the honesty goes with it.
Synthesis, not raw thought. The resolution the precedent proposes: two Sovereigns negotiate on behalf of their humans and surface patterns at a synthesis layer, output in the shape of “your weighting on one thing and theirs on another are in tension here, and here is where you align.” The humans see the synthesis. They do not see each other's interiors. Discretion becomes the system's job specifically because humans handle it unreliably under emotional load.∎
- Solo Sovereign, the current state. Build an honest model of yourself because no one else is watching.
- Opt-in pairwise synthesis, the future state. Two Sovereigns negotiate what is safe to surface between trusted partners.
- Small-group synthesis, further out. Families, teams, communities, on the same opt-in basis.
Explicitly not social media, which optimizes for attention; this optimizes for understanding, a different objective function entirely. Explicitly not a near-term build target: the solo layer has to be deep and trustworthy before pairwise synthesis is even safe to consider. And the precedent sets a test for its own future: if a feature proposal cannot trace back to serving cross-Sovereign synthesis as eventually understood, it is suspect.