C2 — A Single-Human Company OS
An experimental operating system for a single-human company: a deterministic operational spine, governed agent teams, scoped memory, and a structured model of organizational reality in PostgreSQL.
The problem
Chat interfaces make AI look easy to operate and hide how hard it is to run anything with it for more than an afternoon. The failure modes are familiar to anyone who has tried: the agent forgets what mattered last week; every agent sees everything because scoping is tedious; the conversation history quietly becomes the system of record; an automation writes to a production table with nobody's name on the row; three tools hold three copies of the same fact; nobody can say which agent owned a decision or where a piece of "knowledge" came from.
I run one business and one life with a great deal of software. The question I am working on is not "how do I make an assistant answer more questions." It is closer to: what architecture lets one person operate a sophisticated organization with software and AI while keeping structure, continuity, authority, and control? That is an organizational design problem and a systems engineering problem at the same time, and C2 is where I am working through both.
C2 is not a chatbot, an agent wrapper, a pile of automations, a vector database, or a Slack bot, although it contains a Slack front door, workflows, a vector store, and agents. It is an attempt to make those parts one governed operating system: agents, deterministic software, memory, authority, operational state, tools, and human approval working together with explicit boundaries between them.
System model
The architecture has four subsystems, and the separation between them is the design.
C2 Spine is the deterministic operational substrate: a PostgreSQL database of record, a workflow engine that is the only automated writer, a governed write API, a read-only console, and the network edges that make it reachable. Projects and work, CRM, finance, knowledge, content, IT assets, health, and the agent-operations ledger all live here. The premise is simple: agents should not have to reconstruct organizational reality from conversation history. Important state should exist deterministically, in a governed system, with provenance.
C2 Brain is the semantic memory layer, deliberately separate from the Spine. Memory is derived and reconstructible; the record is not. They should never share a failure domain, and retrieval load should never land on the system of record.
The agent harness is the governed execution layer: one orchestrator, a container per active domain team, managers and specialists inside each team, independent validators, a small set of orchestration strategies, and a proposal-and-approval path that ends with a human click in Slack.
Cognition, local model serving, is a scaffold with nothing in it yet. Every model call today goes to a hosted provider or to a first-party agent runtime.
A restrained biological reading helps keep the roles straight. The Spine is the skeleton: structure and durable state. The Brain is long-term memory. Teams are specialized regions, each responsible for a domain. The harness is the nervous system, carrying signals between regions and turning intent into coordinated work. Tools are the muscles that actually do things. Governance is the immune system: it does not think, it refuses. I use the metaphor only where it explains a boundary, because the technical meaning matters more than the picture. Postgres does not reason. Models do not hold authority. The intelligence of the system comes from specialized agents, scoped memory, controlled tools, and a structured model of the organization working together.
Specialized teams, not one omniscient agent
C2 is built around logically separated teams rather than one universal agent. A team is a directory of data the engine mounts: agent definitions, a manifest declaring what the team owns and may read, a Compose file, and a database identity conferred by a reviewed migration. The engine is one image; the team is the data.
Four teams are active today and routable by the orchestrator: operations (projects, tasks, delivery, cost control), marketing (research, writing, and every write to this website), engineering (changes to C2 itself), and health (personal health interpretation, with its own stricter runtime rules). Ten more exist as seeded charters with a placeholder manager and no container, identity, or route: command, administration, intel, logistics, strategy, technology, learning, finance, revenue, and governance. Activating one is a reviewed pull request that flips its status and applies the migration that creates its role. The orchestrator cannot route to a seeded team, and the test suite pins that.
Inside a team the shape is consistent. A manager is the only agent that can spawn others; it returns a plan naming a strategy, the workers, and a validator, and the harness checks every name against the team's manifest. Specialists have narrow lanes and explicit source allowlists. An independent validator the manager cannot bypass reviews the result with a fresh context that deliberately excludes the creator's reasoning. Any plan that would change state must use the creator-verifier strategy and must end in a human approval.
The separation is worth the overhead for reasons that compound: smaller authority per identity, smaller blast radius when a prompt is compromised, clearer ownership, narrower retrieval scope, more predictable behavior, and audits that can name a team rather than "the AI." A team does not receive every tool, credential, memory namespace, or table simply because it belongs to C2. The operating principle is that capability is scoped to organizational responsibility. Operations can read projects and propose task changes; it cannot read the finance ledger, and its cost controller reasons from operational cost records instead. That is a grant, not an instruction.
Memory and world state
C2 distinguishes three things that are easy to conflate: what a team is doing, what the organization knows, and what is actually true.
Execution state is mission-local: the objective, the runs, the validation result, the proposal identifiers, and the events. It lives in a small per-team store that is disposable by design. If losing it would cost a durable fact, that fact was stored in the wrong place.
Semantic memory is what agents carry between sessions: documents, research, summaries, things learned. This is the Brain's job. Its infrastructure runs today as a memory server backed by its own Postgres with pgvector, with embedding indexes for similarity search and full-text indexes beside them, and per-client credentials so that C2 agents and my own coding tools are distinct consumers. The memory design (recorded as a proposed decision, not a finished one) defines four layers: execution state, operational truth, canonical reviewed knowledge, and semantic recall. It also defines memory scopes as logical namespaces over one shared substrate, not a brain per team. A team reads the shared namespace, its own namespace, and the project scopes it is working in; it writes only its own namespace and its session records. Nothing writes the shared namespace directly. A cross-domain fact first lands in the learner's namespace with provenance and reaches the shared layer only by promotion.
This matters because a marketing team should keep continuity about campaigns, audience, and publishing decisions without every other team absorbing that working context, while all teams still share one picture of organizations, people, projects, decisions, research, assets, and relationships. The rule that keeps recall honest is that semantic memory never overrides the record: when recall contradicts operational truth, that is a reconciliation signal, not a correction.
Operational truth is the Spine. Which brings me to the part of C2 I think is most important.
The life ontology
The C2 Postgres database is a structured representation of the organization and the world it operates in. I call it the life ontology because it is not a task database or a CRM with extra tables; it is one relational model of people, organizations, clients, projects, milestones, work items, decisions and lessons, money, assets, subscriptions, infrastructure, knowledge, content, research, operational events, and the relationships among them. The schema is organized into 26 domain modules, 16 active and 10 planned as scaffolds, registered in the database itself. That registry is not documentation: workflows and Slack commands are generated from it, and a view that lists unregistered tables must return zero rows.
A few conventions carry most of the weight. Every entity has a universal identifier; identities in external systems are recorded in one mapping table so no table grows a per-system column. Nothing is hard-deleted. Who did something is decided by the database, not the caller: provenance triggers stamp the acting identity from transaction state and overwrite anything a client sent. A single actionable-record table holds every task, reminder, and follow-up, so there is one task list. A change log records what happened, distinct from what is to be done. An outcomes table holds wins, losses, and the lesson from each, which is the reflection layer and, honestly, the table most likely to sit empty.
The distinction between memory and state is fundamental. Semantic memory helps an agent remember or retrieve context. The ontology tells the system what is actually true according to the governed system of record. The language model is not the brain on its own; the intelligence emerges from specialized agents, memory, tools, and this structured model of organizational reality. Postgres supplies durable world state that models and deterministic services reason over. It does not reason.
Much of the ontology is still sparsely populated. The tables, constraints, vocabularies, and data dictionary exist; the seeding worksheets are written; loading them through the governed path is current work. I would rather say that plainly than imply a full organization already lives in it.
Governed agency
C2 is not an attempt to maximize agent freedom. It is an attempt to make agent capability governable, and the design principle is that an agent can be intelligent without being authoritative. Agents analyze, retrieve, synthesize, plan, and prepare actions. Anything that modifies a system, spends money, contacts a person, publishes content, or creates an external consequence needs stronger authority than a model holds.
The enforcement sits as close to the data as possible, in layers that hold independently. Each team connects to Postgres under its own role, and the roles are narrow by construction: broad read, narrow write, and for a runtime-facing team, no business writes at all. An agent role can insert an approval request but cannot update one, which is the privilege that would let it approve its own proposal. It cannot edit its own budget limit, prompt version, or evaluation record. The tool server that a team's runtime uses enumerates its role's write privileges at startup and refuses to serve if they exceed the allowed set. Prompts describe behavior; grants confer authority.
DECISION: Agents propose; a separate identity executes after a human click. No model-facing process holds a credential that can change business state or publish.
Every write to the record goes through a small governed API rather than raw SQL. Those functions validate the target against the module registry, reject unknown columns by name, refuse provenance columns, and for updates carry the expected prior values so a stale approval fails instead of overwriting. There is no delete function, on purpose.
Above the database, the harness adds the human boundary. A mission is the durable envelope for a request: objective, team, strategy, runs, validation, proposals, approvals, and events, with the run identifiers carried into every tool call so that a proposal can be traced to the mission that filed it. Anything that would change state becomes a proposal row that a trigger pins to pend [EDITORIAL: this sentence arrived truncated in the source text; restore it before publishing]
Tools define capability. A skill is a reusable procedure an agent may load, and a skill can never add a tool, credential, network path, or grant. The team's tool surface is a per-team MCP server spawned per run with the team's identity, exposing four primitives and a set of semantic wrappers, and every call is audited as a row in the record. Third-party MCP servers may attach beside it only as non-authoritative. Website publishing uses the same shape with a different target: the marketing team drafts through the CMS's own MCP endpoint under a drafts-only role, the proposal carries a content fingerprint of the draft as I read it, and the approvals service publishes exactly that version or refuses.
Infrastructure
Everything runs as containers on one host today, organized as several Compose projects that could be split across machines: the Spine stack (Postgres, a graph database, the workflow engine, a reverse proxy with an internal certificate authority, an outbound tunnel), the Brain stack (the memory server and its own Postgres), the agent stack (the orchestrator, one container per active team, and the Slack entry as its own project), and a private media archive with its own database and machine-learning service. The multi-machine layout is designed and supported by the repository structure, and not yet executed. The graph database runs and is monitored but nothing reads or writes it yet.
Ingress is deliberately narrow. The workflow engine is bound to loopback and reached only through the proxy on the local network. Inbound webhooks arrive through an outbound-dialing tunnel that routes only webhook paths and refuses everything else, with request signatures verified at the edge. The conversational front door uses a socket connection that needs no inbound port at all. Configuration lives in git; runtime state, the encryption keys, and the certificate authority root live outside the working tree; secrets are kept out of the repository structurally rather than by ignore rules. The whole system can be rebuilt from the repository plus committed schema snapshots, which is the success criterion I guard most carefully.
The workflow engine is version-controlled in a way I have come to value: workflows are identified by id rather than name, reconciled in one direction with guards that refuse to sync from anything but a clean, current main branch, and a scheduled job that opens a pull request when the editor drifts from the repository and deploys only when a human has merged. No job merges a pull request.
Model access is hosted today: a hosted model gateway for most provider calls, and a first-party coding runtime that runs three of the four teams on their own subscription sessions with configuration passed as code and an allowlisted environment. A local inference tier is planned and not started.
Development evolution
C2 did not start as a multi-agent diagram. It started as deterministic automation solving operational problems, and the agent architecture grew out of what that automation could not do.
The first layer was the database and the workflows: a PostgreSQL schema with governed write functions and provenance triggers, a workflow engine writing through them, and Slack slash commands for projects, tasks, reminders, daily briefs, and captures. That surface still runs and it is still the right tool for anything that should happen the same way every time. A read-only console followed, and the privilege model for it was settled by decision record before the first page was written.
The team model came next: one shared runtime, one team per domain, one identity chain per team, and the rule that a manifest requests authority while only a reviewed migration confers it. Then orchestration through a single entry point, with missions as the durable envelope and human approval as a state transition rather than a suggestion. Then first-party agent runtimes reaching the record only through a per-team tool server that reads and proposes. Then shared skills, CMS publishing as a governed write boundary, one Slack front door, and a private media archive that agents can search but never publish from.
The decision records are honest about their own status. Eight of them are still marked proposed while the design they describe is running in production, and the repository tracks decision status and delivery status as two axes that are allowed to disagree. A continuous-integration check refuses any pull request that changes the system without changing the documents that describe it.
What I am learning
[EDITORIAL: the opening of this paragraph arrived truncated in the source text; restore it before publishing] like as data, and where a human sits in the state machine. Once those were right, the agents could be given real work.
Memory is an organizational-design problem. Deciding what a team may remember, what is shared, and who can promote a learned fact into common knowledge looks a lot like designing an organization's information flow. Treating memory as one substrate with scoped views, rather than a brain per team, was the decision that made the rest coherent.
Specialization looks like an org chart because it is one. The teams, their charters, and the value chain across them read like a company's staff sections. That was not decoration; ownership of a domain is what makes scoping, validation, and audit meaningful.
Deterministic state is non-negotiable once probabilistic systems touch real operations. Recall that contradicts the record has to lose. A cron job and an agent draw the same line: derived is written, judged is proposed.
LESSON: Orchestration complexity grows quickly. A follow-up question became a whole mission, a read of a proposal was credited as filing it, and the front door lost its socket for three days without a log line saying so. Each became a design change, and each was cheaper to find because the system writes down what it did.
Simpler is usually better. Most of what C2 does every day is deterministic workflow, and I keep the slash commands on a separate app from the agent precisely so the difference is visible. Not every automation should become an agent.
Current state
Implemented and running: the PostgreSQL system of record with 26 modules (groupings of domain-specific tables. For example, CRM consists of customtables for a personal CRM), the governed write API and provenance model, per-team database roles, the workflow engine with version-controlled workflows and one-direction reconciliation, the read-only console, the orchestrator and four active teams, missions and orchestration strategies, per-team MCP tool surfaces with call auditing, the approvals service and Slack front door, fingerprint-verified website publishing, the media archive with sanitized staging, an offline test suite that pins the governance invariants, and an infrastructure report that posts on a schedule whether or not anything is wrong.
Actively evolving: seeding the ontology from the data dictionary and worksheets, activating the intelligence collection chain that has been wired but has not run, ratifying the proposed decision records, and closing the small deploy gaps recorded in the status file.
Proposed: the layered memory substrate and its memory manager. The Brain's infrastructure runs; the scoped namespaces, retrieval contract, consolidation, and promotion path are designed and not built. The manifests declare memory namespaces that nothing enforces yet.
Planned: ten seeded teams awaiting activation, ten planned schema modules, the multi-machine split, and the local inference tier.
Experimental: a model bake-off harness that records evaluations to a disposable table, and the first-party runtime path itself, which is proven end to end and still young.
OPEN QUESTION: As autonomous systems gain authority, agent identity, configuration integrity, provenance, and authorization stop being software-design problems and become security problems. Today's immune system is roles, grants, triggers, scoped credentials, fingerprints, and approvals. The direction I am designing toward is cryptographic agent identity, signed and validated configuration, runtime integrity, and eventually identity and signing that remain sound through the post-quantum transition. None of that is implemented in C2 today.
From C2 to Ally
C2 is where I am working through the engineering problem. Ally is the venture hypothesis that asks what this architecture could become if it were turned into a coherent product for other founder-operators. The two are kept separate on purpose: this page is the system as it actually exists, with its proposed parts labeled; the Ally page is the direction.