Reconciling realiy: a real world framework for financial engineering (draft)

Jul 19, 2026

Epistemology before ontology

Reconciliation is applied epistemology. Companion to Reflections on financial advisory and portfolio management.

Reconciliation lives in the basement of high finance. No quant enters the profession dreaming of matching custodial files at seven in the morning; the glamour is presumed to reside several floors above, amid the alpha signals, the investment theses, the models, and the backtests with their obliging Sharpe ratios. Financial-engineering programs indulge this preference. They teach stochastic calculus, market microstructure, and machine learning as though the inputs to these magnificent engines arrived clean, complete, punctual, and true. I’m not your back-office b**tch, a senior quant engineer once protested when I asked him to reconcile an account’s performance. He was wrong. It is, technically, middle office.

The title financial engineering is an honorific more than a description. Much of the field contains little engineering; academic finance, less still. The predictable result is a glut of armchair philosophers of markets and a scarcity of people who can force an idea through the recalcitrant machinery of commercial reality: vendors, custodians, settlement, taxes, fees, short locates, transaction costs, and the client who wires out cash halfway through a rebalance. The engineering lies precisely in that passage. An idea that has never survived contact with a custodial file is not yet an idea about markets. It is an idea about mathematics.

Engineering, moreover, is not synonymous with coding. Code can increasingly be delegated, and as machines grow more adept at drafting what humans specify, its manual production occupies a diminishing share of the craft. Code is becoming abundant; engineering judgment is not. The most valuable practitioner is increasingly an orchestrator: someone able to design an end-to-end process, decompose it into tractable parts, direct agents and systems, inspect their outputs, and remain answerable for the behavior of the whole rather than merely execute one narrow technical task. What cannot so readily be delegated is the cast of mind that programming has traditionally imposed: the explicit management of state rather than its tacit assumption; the preference for pure functions and declared side effects; the taste for a simple mechanism that fails predictably over a clever one that fails inventively; and the presumption, fundamental to fault-tolerant design, that every input will eventually arrive late, duplicated, corrupted, or wrong. Such habits are seldom acquired except through injury. One learns to respect state from the bug that corrupted it, and fault tolerance from the pager that sounded because none had been built. The subtitle of Reflections stated the conclusion: all portfolio managers should be engineers. This essay concerns what, exactly, that engineering entails.

The first fact of practice is severe but simple: every representation of the world may be wrong; one must nevertheless act upon it. Fiduciary duty is often a duty of speed. The deposit must be invested, the withdrawal funded, the alpha refreshed, the loss harvested, and the portfolio rebalanced today—at today’s prices, on today’s data. The duty of care does not adjourn while the data is repaired. Nor does responsibility for the resulting trade error ordinarily remain with the defective file. You selected the vendor, designed the ingestion, and decided what to trust, to what degree, and at what moment. The file was wrong may explain the failure. It does not excuse it. Selecting, supervising, and constraining the file was itself part of the work.

Between the world and its representations lies a gap that finance can neither abolish nor ignore. It prices across that gap, settles across it, and occasionally falls into it. Reconciliation is the discipline of measuring the distance before the distance measures you.

I recently compared notes with a fintech founder and the crew he had assembled at his Hacker House in Atherton. Their project was lifetime financial planning by simulation: every consequential choice would be rehearsed in a digital twin before being made in life itself. What to own and when to sell it; which bill to pay first; how much risk to insure against; whether to refinance; which college to attend; when to retire; when to buy the house—all would be run forward, cheaply and repeatedly, through a model intended to contain the whole financial environment. The necessity of that discipline becomes clearest when the representation aspires to reproduce not merely an account but an entire financial life.

Take the ambition seriously, and its scope is not merely the tax code plus an assortment of optimizers. It is everything: every price of every instrument; every interest rate, exchange rate, yield curve, credit spread, and volatility surface; every statute, regulation, benefit formula, insurance term, debt covenant, fee, penalty, vesting date, tuition schedule, and household cash flow; every relevant fact about the person; and every plausible path by which any of these might change. The replica would have to contain not simply the financial world as it stood, but the branching futures through which a person and that world might travel together.

Among those recruited to construct it were philosophers working on ontology—an intellectually respectable and, in the age of large language models, distinctly fashionable concern. The instinct was sound. A system that proposes to reason about an entire financial life must first decide what kinds of things inhabit that life, which distinctions matter, and how the entities relate: a person to an account, an account to an asset, an asset to a price, a marriage to a tax status, a mortgage to a house, an obligation to a date. Before the machine can reason about the world, it must possess some account of what the world contains.

The harder problem begins where naming ends. It is not merely ontological but epistemological. The system must know not only what an account, a price, a marriage, a tax liability, or a future obligation is, but which claims about each are true; how those claims came to be known; when they were last verified; which authority supports them; how much confidence they deserve; and what should happen when credible sources disagree. A custodian may report one position while an adviser’s ledger reports another. A market-data vendor may revise a price after the trade has already been placed. A tax rule may be enacted, interpreted, challenged, and applied on different dates. Ontology supplies the replica with its nouns and relations. Epistemology decides whether any sentence composed from them deserves belief. Reconciliation is the machinery by which that belief is tested, revised, and earned.

The ambition is magnificent because simulation is cheap and life is not: better a thousand mistakes in the copy than one in the original. One imagines the Federal Reserve taking notes. Yet such a replica would constitute derived state on an imperial scale, and derived state remains subject to the elementary law at issue here: unless it is continually anchored to its source, it drifts. Prices move, statutes change, custodians restate, households surprise themselves, and the world declines the courtesy of remaining still. A twin that is not continuously reconciled against its original will therefore cease to be a twin and become, by degrees too small to notice until they are too large to repair, a fiction with a dashboard.

What follows is therefore a prolegomenon to any future Laplacean ledger. Before a simulated world can be trusted, someone must build the machinery that keeps it faithful to the world it simulates. That machinery is reconciliation. Its reach marks the frontier of safe application, for the twin and for every lesser system alike: one may automate exactly as far as one can verify, and no farther.

In Reflections, I approached the problem through three aphorisms. The first was that information travels at the speed of light, while financial information moves at T+1—at least in settled form. The second was that concurrent write access to the same custodial account is the original sin: it creates race conditions and yields undefined behavior. The third completes the argument: the custodian’s firmware is broken—not necessarily as an engineering artifact, but as a social contract—and the only dependable remedies are exclusive access and explicit sequencing.

Together, these propositions describe a distributed system that predates the term itself: replicas without consensus, messages without exactly-once semantics, clocks without synchronization, and participants without a shared definition of done. We do not choose this system; we inherit it, fax machines included. A later essay may consider the architecture one would design from first principles—single-writer custody, event-native records, and finality that moves at the speed of information rather than settlement. For the present, however, we must work inside the system history has supplied. Thoughtful practice consists in supplying, through design and vigilance, the guarantees that history omitted.

A minimal formalism

The discussion so far has been conducted in prose, and prose eventually cheats. Before the engineering artifacts arrive, it is worth fixing notation precise enough to hold them — not rigor for its own sake, but addressability: every claim in this essay should be a statement about a named object.

Let \(x^{*}(t)\) denote the true state of the world at time t — the account as it is, the price as it printed, the liability as enacted. No one holds \(x^{*}\); it is the thing fiduciary duty is owed toward and the thing no file contains. What one holds are sources. A source \(V\) — a custodian, an accountant, a market-data vendor, one's own middle office — supplies assertions

\[x^V\!\left( \underset{\substack{\Big\uparrow\\[-0.2ex] \scriptstyle \mathrm{knowledge\ time}}}{s}, \; \underset{\substack{\Big\uparrow\\[-0.2ex] \scriptstyle \mathrm{observation\ time}}}{t} \right), \qquad t \le s,\]

what is known at knowledge time s about the state at observation time t, with the shorthand \(x^V_t := x^V(t, t)\) for the in-time reading. The constraint \(t \le s\) holds for realized observations because sources report the past, and the object is the knowledge triangle of the sections that follow: a row \(x^V(s, \cdot)\) is a vintage — the world as known at one instant; a column \(x^V(\cdot, t)\) a revision history — the biography of one observation date; the diagonal the in-time series — what each date knew about itself; and \(x^V(\text{now}, \cdot)\) the best knowledge — the current retrospect. Write \(x^V(\infty, t)\) for the settled terminal value, once the revisions stop. A source's error then splits in two:

\[x^V(s, t) - x^{*}(t) \;=\; \underbrace{x^V(s, t) - x^V(\infty, t)}_{\text{pending revision}} \;+\; \underbrace{x^V(\infty, t) - x^{*}(t)}_{\text{bias}}\]

The first term the source will repair by itself, given time; the second it will never repair, because it does not know it is wrong. A good source is one whose pending revisions decay quickly in \(s - t\) and whose bias is small — quality is a property of the whole triangle, not of any single print. One clarification prevents a category error: the uncertainty here is epistemic, not ontic. Unlike a quantum state, the past is not undecided; \(x^{*}(t)\) is perfectly definite and merely unobserved.

Ground truth never appears in a computable expression. What can be computed are cross-source residuals,

\[\delta^{VW}(s, t) \;=\; x^V(s, t) - x^W(s, t) \;=\; e^V(s, t) - e^W(s, t),\]

and this is the first structural fact about reconciliation: it is inference about the unobservable errors \(e^V = x^V - x^{*}\) from the observable disagreements \(\delta^{VW}\). The subtraction cancels whatever the sources share — two witnesses with a common failure mode agree with each other and with nothing else — which is the formal content of the essay's recurring demand for independent witnesses: the value of a second source is not its accuracy but the orthogonality of its bias. Triangulate until no plausible error direction lies in the common blind spot.

Now the dynamics. The state moves under two hands,

\[x^{*}(t+1) \;=\; x^{*}(t) + a_t + w_t, \qquad a_t = \pi\!\left(x^V_t\right),\]

where \(a_t\) is the side effect you induce — chosen by a policy \(\pi\), a fixed rule from belief to action; from belief, not truth, because belief is all \(\pi\) can be given — and \(w_t\) is everything the world's other writers do: fees, dividends, journals, client flows, reorganizations. The expected law of motion is

\[x^{E}_{t+1} \;=\; x^V_t + \pi\!\left(x^V_t\right) + \hat w_t,\]

with \(\hat w_t\) the anticipated portion of the exogenous term — declared dividends, scheduled fees; zero in the simplest case. Tomorrow the source reports \(x^V_{t+1}\), and it will not equal the expectation. Define the innovation and decompose it:

\[\nu_{t+1} \;:=\; x^V_{t+1} - x^{E}_{t+1} \;=\; \underbrace{x^V(t{+}1,\, t) - x^V(t,\, t)}_{\text{revision of the base}} \;+\; \underbrace{\tilde a_t - a_t}_{\text{execution}} \;+\; \underbrace{w_t - \hat w_t}_{\text{exogenous surprise}} \;+\; \underbrace{\eta_{t+1}}_{\text{fresh observation error}}\]

where \(\tilde a_t\) is the action as actually executed — fills, partials, rejects — and \(\eta_{t+1}\) is the new print's own error, the revisions it has not yet received. This identity is, in a strict sense, the subject of the essay. Every morning's surprise is the sum of exactly four things: the base you acted on was revised; the action executed was not the action sent; the world did something unanticipated; and the newest reading is itself provisional. The contract-based reconciliation of the later section is this identity with each term bound to evidence — the execution term to fills and pending settlements, the exogenous term to declared entitlements, the revision term to the source's own restatements — and a clock on every binding. What survives after each term has claimed its share of \(\nu\) is the break.

The frame this places us in is not exotic; it is filtering. \(x^{*}\) is a latent state, the sources are sensors with lag, noise, and the unusual habit of revising their past readings, the internal book is the posterior, and \(\nu\) is the innovation. Fifty years of state-space discipline transfer intact, including the central habit: a healthy filter is monitored through its innovations, not its states. When the pipeline is sound, \(\nu\) is small, patternless, and fully attributed. A break is a structured residual — and a recurring structured residual is a misspecified model of the world, which the later taxonomy will call a convention error.

The counterfactual is where the money is. With hindsight \(s > t\), the action one should have taken is \(a^{*}_t = \pi(x^V(s, t))\) — the same policy, evaluated on revised knowledge: a counterfactual over information sets, not over worlds. The action error and its bound:

\[\Delta a_t \;=\; \pi\!\left(x^V(t, t)\right) - \pi\!\left(x^V(t{+}1, t)\right), \qquad \lVert \Delta a_t \rVert \;\le\; L_\pi \,\lVert r_t \rVert,\]

where \(r_t = x^V(t{+}1, t) - x^V(t, t)\) is the overnight revision and \(L_\pi\) the policy's sensitivity to its inputs — its Lipschitz constant: the worst-case action change per unit change in input. Its impact on the world is \(\Delta a_t\) itself, carried forward by the dynamics: a mis-trade does not decay of its own accord; it persists until compensated, at whatever the price has become. The bound names the only two levers that exist. Shrink \(\lVert r \rVert\): better sources, second witnesses, arbitration before use — the ingestion disciplines. Or shrink \(L_\pi\): make the policy less sensitive to precisely the inputs most likely to revise — which is what regularization means operationally, and why a good system is regularized is an engineering theorem rather than a taste. Tolerance bands acquire an exact meaning here: a no-trade band is a region where \(\pi\) is locally constant, \(L_\pi = 0\), so revisions within tolerance induce identically zero action error. One does not ignore small breaks out of laziness; one designs the policy so that small breaks are provably harmless.

Finally, drift. Let \(b_t\) be the internal book — the firm's stored belief about \(x^{*}(t)\) — and \(\varepsilon_t := b_t - x^{*}(t)\) its error. A book rolled forward on its own beliefs, \(b_{t+1} = b_t + \tilde a_t + \hat w_t\), accumulates error like a random walk:

\[\varepsilon_{t+1} \;=\; \varepsilon_t + (\hat w_t - w_t) + \cdots \;\;\Longrightarrow\;\; \operatorname{Var}(\varepsilon_T) = O(T).\]

The self-rolled book does not merely err; it wanders. A book re-anchored on the record, \(b_{t+1} = x^C(t{+}1, \cdot) + \text{in-flight}\) — the newest record plus only what has not yet reached it — carries error bounded by the record's own lag window: \(O(1)\) against \(O(\sqrt T)\). The anchor rule of the coming sections is that choice, and nothing more.

With the notation fixed, the essay's artifacts become statements. The stored triangle is \(x^V(s, t)\) made durable; best_knowledge is its bottom row and in_time its diagonal; \(\sigma\) is a knowledge cutoff, the row at which a query freezes the triangle; an epoch — cutoff, methodology version, engine version, the Restatement section's triple — makes every published number a pure function of a vintage. The custodian and the accountant are the sensors \(x^C\) and \(x^A\); IBOR is the posterior estimate of \(x^{*}(\text{now})\); OBOR is the pinned argument actually fed to \(\pi\), kept so that \(\Delta a\) can later be audited rather than argued. And reconciliation, throughout, is the discipline of the innovation: measure \(\nu\), explain it term by term, put a clock on every explanation, and treat what survives as the truth trying to reach you.

Three disciplines

Before reconciliation is a process, it is a set of capabilities, and they come in three tiers of ascending difficulty. They are worth specifying before any implementation, because the tiers have an awkward property: they are built from the bottom and justified from the top.

The first is lossless record-keeping: everything received — every file, every message, every acknowledgment — kept verbatim, forever. Storage is cheap, and the vintage you failed to keep is unrecoverable at any price. But the tier is easier to satisfy than to satisfy well. A dump of everything into flat files honors the letter — nothing is lost — and betrays the purpose, because the archive does not exist for its own sake; it exists to feed the tiers above. Governance needs to replay it: re-ingest a corrected feed, reprocess a quarantined batch, prove what arrived when. Attribution needs to address it: retrieve the exact assertion, from the exact source, as of the exact moment. An archive that can be neither replayed nor addressed is not a library but a landfill — the vintage technically present and practically unrecoverable, which is the expensive way of having failed to keep it. Lossless is the floor. Addressable is the requirement.

The second is governance: the ingested record is presumed defective — late, duplicated, restated, occasionally wrong — and the handling of defects is policy rather than heroics at the point of failure. But policy undersells the requirement. The policies must be encoded: a platform whose API makes them first-class — declarable, versioned, testable, changeable — and whose abstractions match the policy patterns that actually occur: quarantine, precedence, quorum, tolerance, escalation. The governing constraint has a name in the software literature — Martin's Dependency Inversion Principle: high-level modules must not depend on low-level modules; both must depend on abstractions. Concretely, the attribution tier must not depend on how any particular vendor's defects are handled, and the vendor handlers must not leak upward; both depend on the abstraction between them — a stream of clean assertions carrying provenance, behind which the arbitration rules can change without the tiers above noticing anything but a version number. Invert the dependency and a policy change is a new implementation of a stable interface. Fail to invert it and every policy change makes the higher tiers a little more opaque, until nobody can say which rule produced which number. This is not something most engineers can do. It is something most engineers believe they can do, which is worse, and the signature failure is not underbuilding but overbuilding: a rules engine general enough for policies nobody has, configurable enough that the configuration becomes the new opacity. The right platform encodes the patterns actually observed and nothing speculative. Restraint is the hard part.

The third is attribution: the ability to say, of any derived number, which inputs produced it, under which assumptions, as known when. Performance data is where this tier earns its reputation. GIPS composites want one view, the portfolio manager presents another, the advisor a third, the end investor a fourth — and these are not four formats of one number but four defensible definitions of it. Each must be derivable from the same fact base, attributable to its inputs, revisable coherently when the inputs move, and reconcilable against the custodian's arithmetic, the accountant's, and whatever performance software the other side runs. The first tier is an archive; the second, a system of law; the third, an epistemology.

The tiers have a treacherous build order. Logistics runs bottom-up: the archive exists on day one, policies accrete as defects arrive, attribution comes last, if ever. The dependencies run top-down: what the archive must preserve is dictated by what attribution must address; what the policy layer must expose is dictated by what attribution must see through. Build in delivery order while designing in dependency order, or the early tiers harden into shapes the later ones cannot use — the flat-file landfill, the opaque rules engine — and the third tier arrives to find its foundations poured wrong. Built from the bottom; designed from the top. The only defense is to hold the end state in view from the first commit.

None of this is separable from the people. The engineers who build such platforms are rarely the ones who keep them: the builder ships the platform he imagines is wanted and moves on; the keeper inherits the toil the imagination missed. Toil is not noise — it is the ledger on which that misalignment is written, the transaction cost of improvement paid daily by those without the authority to remove its cause. Reading that ledger, and realigning authority with cost, takes a capable engineering lead and an honest engineering culture; no framework substitutes for either. I do not believe the engineering can be separated from the engineers. A first-class data platform cannot be built without a first-class engineering team: the team is the only way to the product, and the product is what drives the team. They compound together, or not at all.

Two tools serve the three tiers, and the next section takes them up in turn. Separation of sources is the simpler instrument: keep the record as received and the truth as derived, side by side, with a policy bridging the two. Bitemporality remembers everything: each assertion is kept with the time it was learned, so any past view can be reconstructed and any derived number attributed — the third tier's natural substrate. The tools attack the same problem — sources that arrive late, wrong, and revised — and they are substitutes as often as complements. Bitemporality is the more powerful and the more expensive: it confuses analysts, taxes every query, and earns its keep only where revision itself is the subject — performance, backtests, audit. Separation of sources costs almost nothing and covers a surprising share of the ground. A shop can get away with one of the two for a long time. The mistake is implementing neither, and discovering that during a restatement.

Sources of truth

The purpose of financial engineering is to systematically induce side effects in the world. Trade, journal, sweep, elect, withhold: each operation attempts to move an account from one state to another — an increment \(a_t\) to the true state \(x^{*}\) — and whether to act, and how, depends upon the state of the world now.

Here the difficulty begins, because no one has ever seen a portfolio. One sees files about portfolios The operations desk is an involuntary Yogācārin: it never touches the world itself, only representations of it — vijñapti-mātra, delivered nightly over SFTP. — never \(x^{*}(t)\), only assertions \(x^V(s, t)\), each stamped with a source and two times. The policy is therefore fed belief, \(a_t = \pi(x^V_t)\), because belief is all it can be given; the formalism's law of motion is not a modeling choice but a description of the predicament. The first engineering question, then, is not simply what is true? It is: which source shall we treat as true, for which attribute, in which frame, at what knowledge time, and on whose authority?

The phrase source of truth obscures three distinct ideas, and the notation separates them. The source of record is the party whose books are authoritative: whatever your database may report, the account is what the custodian's ledger says it is. In the formalism's terms, record status zeroes the bias by contract — \(x^{*}(t) = x^C(\infty, t)\) for settled units and cash, not as an empirical finding but as a definition the parties have signed — so that all of a record's error is pending revision, repairable in principle by waiting. The source of truth is a narrower engineering designation: the input your system elects to treat as authoritative for a specified attribute at a specified time — the particular \(x^V(\sigma, \cdot)\) actually handed downstream. One vendor may govern closing prices, another corporate actions, while the custodian governs settled positions. Derived state is everything computed from those inputs: the internal book \(b_t\), the model, the firm's present belief about the account. The internal book is not the account. It is a belief about the account, carrying an error \(\varepsilon_t\) that no query can display — and beliefs are precisely the kind of thing that must remain revisable.

Every assertion also arrives inside a frame, and the frame is wider than the coordinates: \((V, s, t)\) name the source and the two times; the frame adds a cut time, a timezone, an accounting basis, an adjustment policy, a revision policy. A custodial file is not the account in the abstract; it is one vintage of the ledger — the account as represented at 21:47 Eastern, under settlement-date accounting, subject to later restatement. Prices are plural — the consolidated close, the exchange's official close, an evaluated price are different sources, not competing guesses at one number. Corporate actions are announced, amended, postponed, and sometimes reversed: a column with a habit of rewriting itself. A number without a frame is not data. It is a rumor with digits.

Separation of sources

At bootstrap, taking a single source as gospel — one \(x^V\), no witness — is often rational. It may be the only practicable way to begin, and the expected cost of error — defect frequency multiplied by book size — may be small enough to absorb.

An enterprise under load faces a different arithmetic. Given enough accounts, instruments, and mornings, a one-in-ten-thousand defect becomes a daily occurrence. Nor do defects distribute themselves politely. They gather in the difficult corners: foreign dividends with withholding, midstream symbol changes, odd-lot corporate actions, reorganizations interpreted with clerical originality. At scale, single-sourcing an attribute is a decision to be wrong coherently.

The remedies are dull because foundations usually are: redundancy, cross-validation, and explicit arbitration. Any fact capable of moving money should, where practicable, have a second independent witness — a \(W\) that makes the residual \(\delta^{VW}\) computable at all, since a lone source's error has no observable shadow. Independence is the load-bearing word: the subtraction cancels whatever the sources share, so the second witness is valued for the orthogonality of its bias, not for its accuracy. Sources should be compared before use — agreement checks and control totals on \(\delta\) at ingestion, not merely examined after damage. And disagreement must be resolved by policy — a precedence rule, quorum, or escalation path that determines mechanically which source prevails. At scale, truth is not an input. It is a procedure.

Beneath every ingestion pipeline is one pattern. Ingest the source of record from the vendor, as published. Resolve it — validate, arbitrate, fill — into a source of truth. Act on the truth, never on the raw record. The resolution step exists because the two series have different masters: the record answers to the vendor's schedule, the truth answers to yours. Act on the record directly and the business breaks every time the vendor is late — and vendors are late. Let the truth float free of the record and it drifts, quietly and compoundingly. The whole design problem is that tension — absorb the record's delays without inheriting them; track the record without being hostage to it — and the policy that resolves it should be written down, versioned, and boring.

One pipeline can carry the whole apparatus. Consider benchmark weights: a provider publishes them on a lag — sometimes a long one — yet the weights must exist, every morning, for optimization and performance to run. The design persists two series, separately. The provider's prints, stored as published, are the source of record for what the provider said; late arrivals backfill it, and revision is its job. The operational series is the designated source of truth for everything downstream; it is written once per day, at decision time, and never revised — only superseded. Between the two sits a derivation. Note, too, what the design does not require: no bitemporal machinery — two plain tables and a policy carry the whole weight, which is the point of the separation.

 SOURCE OF RECORD  the prints as published; persisted,
 append-only; late arrivals backfill; revision is its job

 ══ p(t-2) ═══════════ p(t-1) missing ═══════ p(t-1) arrived ══▶
                                                  
                     p(t-1) absent at 8:30:   p(t) absent too, but
        present:     derive                   p(t-1) backfilled
        use the      b(t-1)=drift(p(t-2))     overnight: derive
        print         one hop from the       b(t)=drift(p(t-1)),
                     record                   re-anchored on the
                                             not drift(b(t-1))
                                                  
 ══ b(t-2) ═══════════ b(t-1) ══════════════════ b(t) ════════▶
    pinned            pinned                pinned

 SOURCE OF TRUTH  the operational series; persisted separately;
 never revised, only superseded  pinned at decision time

Each morning the deriver reads the record as it stands and applies the arbitration policy: if the print is there, promote it; if it is not, drift the newest print forward by returns and promote that. The separation is what buys resilience — a delayed record does not block the truth, because the truth is derived, and the business runs. The anchor rule is what keeps resilience from curdling into error: tomorrow's derivation reads the updated record, never yesterday's truth. A drifted value is derived state, one hop from a print; it is never itself drifted. The letter in the diagram is deliberate: the operational series is an internal book for one attribute, and it inherits the formalism's drift bound — re-anchored on the record, its error is confined to the record's lag window, \(O(1)\); rolled forward on itself, it walks, \(\operatorname{Var}(\varepsilon_T) = O(T)\). A late print therefore costs one day of estimation error, extinguished at the next anchoring; drifting the drift is the random walk, discovered months later by whoever reconciles the performance.

The pipeline instantiates the section's trichotomy exactly: the prints are the source of record, the operational series is the engineered source of truth, and the drift is derived state, promoted only by policy. The record keeps revising its estimate of \(t{-}1\) as better information arrives — delayed updates are its design property, not a defect. The truth is pinned at decision time, never revised, only superseded. Performance measured against the pinned value answers what did we do, given what we knew; performance measured against the revised column answers what would we have done, knowing better. Both are legitimate questions. Woe to the shop that cannot tell them apart.

Bitemporality

The practical consequence is bitemporality. A sound system preserves both the time a fact concerns and the time the system came to believe it — the formalism's observation time and knowledge time, which the temporal-database literature has carried since Snodgrass as valid time and transaction time. Restatement then becomes an ordinary write rather than a special calamity, and a historical decision can be audited against the information available when the decision was made — not against the cleaner, corrected knowledge available now.

The object underneath is the knowledge triangle, discretized to days — and it has been rediscovered by every field that takes revision seriously; only the names change. Vintage is the macroeconomists' term: the real-time datasets of Croushore and Stark at the Philadelphia Fed, and ALFRED at the St. Louis Fed, exist to serve GDP as it stood on a given day, not as it later became. The revision history is the actuaries' object, developed one lag at a time in run-off triangles. For realized observations the constraint \(t \le s\) empties the upper half; facts asserted with future effect — a declared dividend, an announced split — are the licensed exception, and the reason the schema of the next section will not enforce the triangle.

The benchmark pipeline of the previous subsection lives in this picture: the source of record owns a column, revising its estimate of \(t{-}1\) as information arrives; the operational truth owns the diagonal — pinned, one cell per day.

                       observation date t ─────────▶

                   1        2        3        4        5
             ┌────────────────────────────────────────────────┐
knowledge 1   x(1,1)                                        
date      2   x(2,1)  x(2,2)           (s < t: not yet     
  s       3   x(3,1)  x(3,2)  x(3,3)    observed)           
         4   x(4,1)  x(4,2)  x(4,3)  x(4,4)                
         5   x(5,1)* x(5,2)* x(5,3)* x(5,4)* x(5,5)*       
             └────────────────────────────────────────────────┘

     row s, read across ─▶   a vintage: the world as known on s
     column t, read down    a revision history: the biography of t
     diagonal               the in-time series x(t,t)
     bottom row *            the current retrospect

The diagonal deserves emphasis. An honest backtest walks the triangle row by row: on simulated date s the policy is fed its own vintage and nothing below it — \(a_s = \pi\!\left(x(s, \cdot)\right)\), the law of motion taken literally. Where observations are available contemporaneously, the freshest frontier lies on the diagonal. Where publication lags intervene, that frontier recedes into the triangle and becomes ragged: the strategy must use the newest observation actually available in row s, not the observation that later history says belonged there.

A strategy simulated against the bottom row instead — today's cleaned, adjusted, and restated history — evaluates \(\pi\) on the retrospect where the frontier was owed, and the edge it reports includes the formalism's counterfactual \(\Delta a\), harvested as alpha. Vendors sell the cure as point-in-time data; the disease is lookahead bias wearing a respectable face. The halfway house also has a name — pseudo real-time, the nowcasters' term for an honest calendar over bottom-row values — which removes the availability lookahead and leaves the revision lookahead intact; a buyer should ask which of the two is actually in the box. Decisions are made from the knowledge available in their own row and judged, years later, from the bottom row. A system that stores only the latter has quietly destroyed the evidence of what deciding was actually like.

Naively materialized, the triangle is quadratic in time and almost entirely redundant. Most facts are printed once and never touched again. Each column is therefore piecewise constant, changing only when a revision arrives. The efficient representation is sparse and event-shaped: store the first assertion and each subsequent revision, append-only, then reconstruct x(s, t) by selecting the latest version of observation t recorded no later than knowledge date s.

One more column

In Postgres the two dimensions cost exactly two columns: date carries the observation date \(t\), updated_at the knowledge time \(s\). The second must be assigned by the database, never by the application — a client that supplies its own knowledge time is a witness writing its own timestamps.

create table observation (
    series_id   bigint       not null,
    date        date         not null,               -- observation date
    updated_at  timestamptz  not null default now(), -- knowledge time
    value       numeric,                             -- null is a tombstone
    primary key (series_id, date, updated_at)
);
revoke update, delete on observation from app_rw;

The table is a ledger: insert-only, enforced by grants rather than by good intentions. A revision is a new row for the same (series_id, date) with a later updated_at; a deletion is a tombstone. Nothing is destroyed, so every vintage stays reconstructible. The canonical read rebuilds one — the row \(x(\sigma, \cdot)\): for each key, the newest assertion at or before the cutoff:

select distinct on (series_id, date) *
from observation
where updated_at <= :sigma
order by series_id, date, updated_at desc;

One column added; every query changed. That is the honest price of bitemporality, and the tax is not the column — it is that every query must now declare which cross-section of the triangle it wants, and SQL will not force the declaration. A naive where date = ... returns every revision, and a sum quietly double-counts them. A join between two bitemporal tables read at different knowledge times manufactures a chimera — a pairing \(\big(x(\sigma_1, \cdot),\, y(\sigma_2, \cdot)\big)\) that no single row of any triangle ever held. The retrospect used where the frontier was owed is lookahead bias, compiled and cached. Each failure returns a number; none returns an error. The bug type-checks.

The remedy is ergonomic before it is technical: make the correct reading the default reading. The ledger lives in its own schema and is never queried casually; every bitemporal table then exposes the same small set of doors, with the same names, so that the suffix does the thinking. Two doors carry nearly all the traffic — the triangle's bottom row and its diagonal, made addressable.

-- the ledger keeps its own schema; api is what analysts see
create schema ledger; create schema api;

-- best knowledge (the bottom row): what we now believe happened.
-- the bare name resolves here, so the naive query is the correct one
create view api.observation as
select distinct on (series_id, date) *
from ledger.observation
order by series_id, date, updated_at desc;

-- in time (the diagonal): what was known by the end of each
-- observation date. Move the cutoff to taste (the 8:30 run, say)
create view api.observation_in_time as
select distinct on (series_id, date) *
from ledger.observation
where updated_at < date + interval '1 day'
order by series_id, date, updated_at desc;

best_knowledge answers questions about the world: research, current reporting, reconciliation against the newest custodial file. in_time answers questions about decisions: backtests, performance as it was published, the audit of what the model saw. The rule of thumb fits on an index card — best knowledge for what happened; in time for what was decided — and the join discipline follows from it: like joins like, best knowledge to best knowledge, in time to in time. The one legitimate mixed join is the revision study — in_time against best_knowledge, per observation date — and it deserves its own named view, because how wrong were the first prints is a question a quant shop should ask on purpose, not produce by accident.

Two more doors serve the specialists. observation_asof(σ) — a function, since views take no parameters — rebuilds an arbitrary vintage; it is the epoch builder's tool, and when several bitemporal tables meet in one query they must share a \(\sigma\), declared once in a leading CTE and threaded through: the query's epoch header. observation_revisions, the ledger filtered to one (series_id, date), reads out a column of the triangle — the biography of a number, which is where every break investigation starts: what did it say, when did it change, to what. And the frontier a backtest must walk — the newest observation per series available on each simulated day — is asof evaluated day by day under a lateral join, the closest thing Postgres has to the as-of join kdb+ built a career on.

The doors are boilerplate, so they are manufactured, never handwritten: one template applied to every bitemporal table, so that _in_time means the same thing on prices as on weights as on flows. The suffix becomes the type system SQL doesn't have. Writes get one door as well — assert_observation(series_id, date, value), which stamps the knowledge time itself; the application never touches updated_at. And when best_knowledge runs hot — it will; it is the default — materialize it as a real table maintained by trigger on ledger insert: the old current-plus-history pattern, where reads become a plain primary-key lookup and correctness survives because the current table is derived. If the two ever disagree, the ledger wins, and the current table is truncated and rebuilt from it — a cache, not a book. The physics cooperate: the ledger's primary key (series_id, date, updated_at) is exactly the index distinct on wants, an insert-only table never bloats, and updated_at is monotone, so a BRIN index keeps time-slicing cheap as history grows.

The border crossing everyone actually faces is the join: a unitemporal table — flows, trades, an instrument master, anything with a date and no knowledge axis — meeting a bitemporal one. The join is underdetermined, because the unitemporal schema cannot say which cross-section it deserves; the question to settle before writing the on clause is where \(\sigma\) comes from, and there are exactly three answers. If it comes from nowhere — the question is about the world, and the flows just want the best current estimate of the prices — join the bare name on date, and the default door does the rest. If the whole query is an epoch, \(\sigma\) comes from the header: join observation_asof((select t from sigma)) like any other table. The interesting case is when \(\sigma\) comes from each row — the unitemporal table records decisions, and every row carries its own knowledge cutoff. That is a per-row as-of, which in Postgres is a lateral:

select t.trade_id, o.value as price_as_known
from trade t
left join lateral (
    select value
    from ledger.observation o
    where o.series_id  = t.series_id
      and o.date       = t.price_date       -- the observation wanted
      and o.updated_at <= t.executed_at     -- knowledge available then
    order by o.updated_at desc
    limit 1
) o on true;

When the unitemporal side carries only a date and no timestamp, the row's true cutoff is unrecoverable and a convention must stand in for it — which is exactly what in_time is: join it on date, and the convention lives in one view definition rather than scattered across queries. And when the observation may not yet exist at the cutoff — publication lag — the decision join becomes a double as-of: the newest observation date at or before the one wanted, and within it the newest assertion known in time.

left join lateral (
    select *
    from ledger.observation o
    where o.series_id  = t.series_id
      and o.date       <= t.date            -- newest available observation
      and o.updated_at <= t.decided_at      -- known at decision time
    order by o.date desc, o.updated_at desc
    limit 1
) o on true;

The order by does the work — observation date first, knowledge time second — and the result is the ragged frontier of the earlier subsection, rendered in six lines: the join a backtest actually runs. The rule condenses to a border policy. A unitemporal table crossing into a bitemporal one is a tourist, and it must declare something: the bare name if its question is about the world, its own timestamp or in_time if its question is about a decision, the \(\sigma\) header if the query is an epoch.

One honest caveat, in the spirit of the first aphorism. now() is assigned at transaction start, but transactions commit out of order, so a vintage read near the leading edge is not repeatable: a straggler can commit later carrying an earlier updated_at, and the same \(\sigma\) returns different answers before and after. The fix is a watermark — publish only from a \(\sigma\) safely behind the oldest in-flight transaction (pg_current_snapshot() gives the horizon), or simply lag by a grace interval. Only behind the watermark does the vintage \(x(\sigma, \cdot)\) become what the epoch will require — a pure function of \(\sigma\). Even your own database settles on a delay. Information travels at the speed of light; knowledge, it turns out, moves at T+ε.

Restatement

Nothing tests the design like performance. Published returns are the most derivative data a firm produces — sleeve returns feed blended returns, pre-tax feeds after-tax, dailies compound into inception-to-date — and the temptation is to store each series as a fact and maintain it. The trap is the word maintain. One bad price in an equity sleeve, discovered months late, must restate the sleeve's pre-tax return, its after-tax return, the blended pre-tax and after-tax returns of every portfolio containing it, and every cumulative number downstream since the error. If those series are mutable rows, coherence is a discipline problem across thousands of tables — someone must remember every dependent — and discipline loses. The principle is the opposite: coherence is never maintained; it is derived. Store only observations bitemporally — prices, transactions, flows, lots, corporate actions, custodial positions — and make every return, at every level of aggregation, a pure function of an addressed state of the fact store. Then internal consistency is a theorem, not a chore.

The address needs three coordinates. A published number is \(R = f(\sigma, m, c)\): \(\sigma\) a knowledge-time cutoff over the facts — a vintage; \(m\) a methodology version — the assumption set, itself data; \(c\) a calculation-engine version. Call the triple an epoch. The coherence rule then fits in a sentence: numbers published together come from one epoch. The blend is never computed from stored sleeve returns; it is computed from the same \(\sigma\) the sleeve returns came from, so sleeve, blend, pre-tax, after-tax, and inception-to-date cannot disagree — they are projections of a single snapshot. The bad price stops being a coordination problem: the correction is a new fact — old observation date, new knowledge time, a cell appended to day \(d\)'s column — which produces a new vintage, which produces a new value everywhere at once. No one restates the equity sleeve and forgets the blend, because no one restates anything; everything re-derives.

     corrected fact: price(x, day d)
     (old observation date, new knowledge time)
                      
                      
        valuation(equity sleeve, d) 
                      
                      
        pre-tax return(equity, d) 
                              
                              
   after-tax return(eq, d)    blend pre-tax(d) 
                              
                              
   blend after-tax(d)         cum blend pre-tax
                              (d..today) 
               
   cum after-tax  sleeve and
   blend (d..today) 

   = the dirty cone: recomputed. Outside the cone 
  other sleeves, other accounts, all days before d 
  content-addressed cache hits.

Re-deriving everything is affordable for two structural reasons. The dependency graph is explicit — price to valuation to sleeve return to after-tax to blend to cumulative — so a correction dirties exactly its downstream cone, not the universe. And returns chain-link: cumulatives are folds over daily atoms, so a fix on day \(d\) recomputes day \(d\) and re-multiplies forward, trivial even months back. Content-address the intermediate artifacts and unchanged subtrees become cache hits — the Nix and Bazel move, applied to returns. The literature is squarely on point: Build Systems à la Carte argues that Excel is a build system; incremental view maintenance and log-derived views are the database renderings of the same idea. A performance database is a build system whose artifacts happen to be returns.

The methodology change is the harder event, and it is harder because it is different in kind. A data correction is a new cell in the same triangle; a change in how a transaction type is handled is a different triangle. Hence \(m\): the assumption set is versioned data, never an edit to code. To rebake, build the epoch \((\sigma, m_2)\) in parallel while \((\sigma, m_1)\) keeps serving — blue-green deployment for analytics, so no partially rebaked state is ever visible. A hermetic \(f\) makes the rebake embarrassingly parallel by account, and the graph scopes it: a change to option-assignment handling invalidates only the accounts that ever had an assignment, and the graph knows which. Best of all, the difference between two epochs is a query, not a project — and that diff is the impact study compliance wants before the flip, and the substance of the client letter after it. Reputational exposure is worst when the firm cannot say who is affected, by how much, since when; epochs reduce the question to a select statement.

Two conventions with regulatory teeth complete the design. Error corrections apply retroactively by nature; methodology changes apply prospectively by convention — a seam, itself a recorded and footnoted fact — and the two must never masquerade as each other. The statistical agencies draw the same line — regular against benchmark revisions — with the methodology convention reversed: a comprehensive revision rebakes the whole triangle under the new \(m\) and restates decades, because the agency's audience wants one consistent history; the firm's audience was shown the old numbers, so the firm moves a seam instead. Same machinery, opposite promotion. And published is not a table but a pointer: which epoch clients see is promoted by policy and moved only as an auditable event — the OBOR move, one level up. What was actually sent to clients is a book the firm itself keeps the record of: the statement archive is immutable, and a restatement is not an update but a superseding event carrying the old number, the new one, the cause, and both epoch identifiers, generated mechanically. Performance, in the end, is not bitemporal but tri-temporal — what happened, what was known, and how it was counted: \(t\), \(\sigma\), \(m\) — and the whole apparatus leaves exactly one mutable thing: a pointer.

BOR bores

The value identity closed with the alphabet as four evaluation policies; the industry keeps the four as institutions, and they map onto the section's trichotomy — though not the way the acronyms suggest. CBOR, the custodial book of record, names the source of record — with one care the notation makes exact: the custodian's ledger is the sensor \(x^C\), while the CBOR file ingested at T+1 is one vintage of it, \(x^C(s, \cdot)\) — cut at their time, on their basis, restatable. One never holds the record; one holds its vintages. The record is kept at lot level, for one account, and it admits surprises: fees, journals, transfers, reorgs, and the custodian's own corrections arrive there first, on nobody's schedule but theirs — exogenous terms of \(w_t\), surfacing in the file before your book has heard of them.

ABOR, the accounting book, is the second sensor, \(x^A\) — not a rival truth but a record-keeper with a different jurisdiction: the accountant is authoritative for NAV, accruals, fees, and income, as the custodian is for settled units and cash. Its error profile is the custodian's mirror. It is unashamedly revisionist because its contract is the terminal value — being right at \(s = \infty\) is what it sells — and it buys the small bias with large, slow pending revisions: accruals true up, expenses adjust, valuations correct. And it doubles as the independent witness for reconciliation: \(\delta^{CA}\) is computable, and valuable precisely because the accountant shares neither the custodian's failure modes nor your incentives — the orthogonality the formalism demanded of a second source.

IBOR, the investment book, is derived state, full stop: the posterior estimate of \(x^{*}(\text{now})\) — the newest custodial vintage, plus everything in flight (today's fills, pending settlements, estimated corporate actions), projected to the present. It is, in the macroeconomists' word, a nowcast, and it lives past the leading edge of every triangle, where no column yet has a cell. News crosses this boundary in both directions, and the census already gave the law: your own actions you witness at the fill, so IBOR leads CBOR on everything endogenous; the world's actions surface in the custodial file first, so the record leads the belief on everything exogenous — \(\tilde a\) before the record, \(w\) after it. Each book is the other's news service, with opposite lags. IBOR exists because of the T+1 aphorism: the gap between decision-time and settlement-time truth grew painful enough that the industry gave the patch a four-letter name.

OBOR, the operating book — whatever one actually transacts against on a given day — is the formalism's closing promise made institutional: the pinned argument actually fed to \(\pi\), kept so that the action error \(\Delta a\) can later be audited rather than argued. It is assembled each morning from three inputs — the newest CBOR vintage, the in-flight ledger that is IBOR, and the corrections surfaced by last night's reconciliation against the accountant — and then frozen, so that every trade answers to what was believed rather than to what became true. The discipline that keeps the belief honest is re-derivation, and here the drift bound comes home to the object it was proved for: rebuilt each morning as \(x^C(t, \cdot)\) plus in-flight, the book's error is confined to the record's lag window; rolled forward from yesterday's book alone, it walks away from the custodian's at \(O(\sqrt T)\). The two canonical failures are the two ways to skip a step: transacting off the raw custodial vintage (stale — selling shares this morning's fill already sold), and never re-anchoring (the rolling position-keeper whose divergence is discovered by whoever reconciles the performance).

Measured against the trichotomy, the four books refuse to line up one-to-one, and the misalignment is itself the lesson. The three labels are not kinds of book but orthogonal properties, each with a formal referent: record is a fact about authority — whose settled column defines the truth, the bias zeroed by contract; truth is a fact about designation — which \(x^V(\sigma, \cdot)\) your system elects to hand to \(\pi\); derived is a fact about lineage — whether the object lies in the image of some \(f\). CBOR holds record authority, yet what you designate as truth each morning is a vintage of it, the ledger itself being unreachable: you hand downstream \(x^C(s, \cdot)\), never \(x^C\). ABOR is both at once: derived in lineage — the accountant computes NAV from custodial data and prices — yet a record in authority, because the struck NAV is official by agreement, however it was computed; its bias is zeroed by signature, not by measurement. Record status is conferred by social contract, not by computational primitivity. IBOR is the opposite case: never a record, derived through and through, and nonetheless the designated truth for one attribute — position, now — because the present has no record-keeper. No column has a cell at now; past the leading edge, the posterior is the only candidate for truth there is. And OBOR, derived for market state, is a genuine record for exactly one attribute: the argument of \(\pi\) — what the firm believed when it acted. No party but the firm is authoritative over its own past beliefs, which is why revising the operating book is not correction but falsification.

What separates the four books, in the end, is less content than revision contract — the shape of the column each is willing to write. OBOR's column has one cell; never revising is its entire value. IBOR never revises either: it is superseded, morning by morning, each edition a frontier cell that stands as issued. CBOR's columns are nearly flat — it restates rarely, though with authority when it does. ABOR's columns converge by design; revision is the mechanism of its contract. This is why the operating book cannot double as the accounting book, however tempting one golden book sounds. Stability at decision time is a flat column; correctness in retrospect is a column that converges to the terminal value; a column both flat and convergent would have had to be right at first print, which is the one promise no sensor can make. The contracts are incompatible; the books must be plural.

 THE RECORDS — kept by others; authoritative,
 each in its own jurisdiction; restatable

 ┌──────────────────────────┐    ┌──────────────────────────┐
 │ CBOR · the custodian     │    │ ABOR · the accountant    │
 │ the sensor x^C           │    │ the sensor x^A           │
 │ settled units and cash,  │    │ NAV, accruals, income,   │
 │ lot level, one account — │    │ fees; trade-date basis   │
 │ sees no sleeves          │    │ revised often: right     │
 │ mostly final; restates   │    │ in retrospect is its job │
 │ rarely, surprises freely │    │ asks: what is it worth,  │
 │ asks: what has settled?  │    │ officially?              │
 └────────────┬─────────────┘    └────────────┬─────────────┘
              │ T+1 file — a vintage of       │
              │ the ledger, not the ledger    │ independent
              │ exogenous news arrives        │ witness: its
              │ here first                    │ recon feeds
              ▼                               │ the next
                                              │ derivation
 THE BELIEFS — yours; derived state;          │
 may be partitioned into sleeves              │
 (the partition must sum to CBOR)             │
                                              │
 ┌────────────────────────────────────────┐   │
 │ IBOR · newest record + the in-flight   │   │
 │ ledger: fills, pending settlements,    │   │
 │ estimated corporate actions            │   │
 │ the posterior of x*(now)               │   │
 │ leads CBOR on your own actions;        │   │
 │ trails it on the world's               │   │
 │ in time, never revised — rebuilt each  │   │
 │ morning, never rolled from yesterday   │   │
 │ asks: what do I hold, right now?       │   │
 └───────────────────┬────────────────────┘   │
                     │ decision-time snapshot,│
                     │ promoted to truth by   │
                     │ policy, then frozen    │
                     ▼                        │
 ┌────────────────────────────────────────┐   │
 │ OBOR · the operating book of the day   │   │
 │ the pinned argument of π               │   │
 │ pinned: never revised, only superseded │   │
 │ asks: what did we believe when we      │   │
 │ acted?                                 │   │
 └───────────────────┬────────────────────┘   │
                     │                        │
                     ▼                        ▼
      reconciliation — the beliefs against the records;
          its findings anchor tomorrow's derivation

Sleeves complicate the picture in an instructive way. The custodian keeps one account, at lot level; a sleeved account partitions the state into sub-portfolios by designation — \(x = \sum_k x^{(k)}\), this lot to the equity model, that one to the ladder, the overlay owning the short options. The custodian neither knows nor enforces the partition. Every sleeve-level book — ABOR, IBOR, OBOR alike — is therefore derived state twice over: derived from the records, and derived again through a partition map that exists only in your metadata. The map must also migrate: corporate actions arrive at account level, and a split that multiplies the custodian's lots multiplies your assignments with it, or silently orphans them.

The partition has one law — it must sum to the account, lot by lot for units and in aggregate for cash — and that law is the sleeve ledger's only witness. Here the value identity's kernel warning stops being hypothetical. An account-level book can be checked against the custodian and the accountant; a sleeve ledger has no external counterparty at all — no second source, hence no computable \(\delta\), hence errors with no observable shadow. Reconciliation acquires an internal layer, \(\sum_k x^{(k)} = x\), and with it a class of break invisible from outside: a reassignment across sleeves slides the ledger along a level set of the sum, so the account ties perfectly to CBOR while the sleeves misallocate — quietly poisoning per-sleeve performance, attribution, and tax lots without producing a single external symptom. The most dangerous kind of book is derived state that nothing outside the firm can contradict.

Cash is where the fiction strains hardest — the aphorism again: sleeving worsens the concurrency problem because cash is shared state. There is one custodial cash balance and \(n\) sleeve cash ledgers. A dividend arrives at the account and must be attributed to the sleeve that held the payer; a fee, a sweep, a client withdrawal arrives at the account and must be allocated by policy — every exogenous term of \(w_t\) without an allocation rule becomes a break that someone resolves by hand, differently each time. A sleeve's cash can go negative while the account is flush, and the reverse; only the sum is real. Meanwhile each sleeve's strategy is a would-be writer to the shared account — \(n\) policies proposing actions on one state — so the partition converts one concurrency problem into \(n\) of them, which is why the working answer, in practice, is a single overlay writer that serializes every sleeve's intentions. But that belongs to the next section.

The value identity

The formalism kept the state abstract, and abstraction was the point: one set of coordinates for prices, weights, and positions alike. But the essay owes the reader one attribute worked in full, and the candidate selects itself — the value of the account, the number every book must eventually answer for and the only one the client is ever shown. Specialize the custodian sensor \(x^C\) to it, writing \(V^C(s,t)\) for the custodian's assertion at knowledge time \(s\) of the account's value on date \(t\), and the innovation decomposition stops being a schema and becomes a morning.

Start with the simplest account worth having: long-only equities, no trading, no flows. The custodian's value is already a sum of ledgers,

\[V^C(s,t) \;=\; \sum_i q_i(s,t)\,p_i(s,t) \;+\; K(s,t) \;+\; D(s,t),\]

settled units at their marks, settled cash, and pending dividends — the receivable, which is where accrual lives. Each factor is bitemporal in its own right, and the sum inherits every clock. Day over day, with holdings still, the evolution is a base and three flows:

\[V^C_t \;=\; \underbrace{V^C(t,\,t{-}1)}_{\text{yesterday, as restated}} \;+\; \underbrace{\textstyle\sum_i q_i\,\Delta p_{i,t}}_{U_t:\ \text{unrealized PnL}} \;+\; \underbrace{I_t}_{\text{income}} \;-\; \underbrace{F_t}_{\text{fees}} \;+\; \underbrace{T_t}_{\text{net transfers}}.\]

The base deserves the stare. It is \(V^C(t, t{-}1)\) — yesterday as this morning's file restates it — not \(V^C(t{-}1, t{-}1)\), yesterday as yesterday told it. The gap between the two is the formalism's revision term, and it arrives not as an announcement but as a silently different opening balance; a reconciler who compares today's file only against today's expectations has left the first term of \(\nu\) unwatched.

Income has a geography worth fixing. A dividend enters value exactly once, at the ex-date, as units held through record times the declared rate, posted to the receivable; the pay date merely migrates it, \(D\) down, \(K\) up, value-neutral. A declared-but-unpaid dividend is also the licensed occupant of the formalism's empty triangle — a fact asserted now with future effect — which is why \(\hat I_t\) is computable in advance and belongs in \(\hat w_t\). And not every custodian carries the receivable: a cash-basis file dips at ex — the price drops with nothing posted against it — and recovers at pay. Against any accrual-basis book this opens a disagreement that is born at ex, lives as exactly the receivable, and dies at pay, on schedule. Hold that shape; it returns below.

The return is the identity solved for growth, with what was given or taken stripped out — the end-of-day value purged of transfers, over the base:

\[r_t \;=\; \frac{V^C(t,t) \,-\, T_t \,-\, V^C(t,\,t{-}1)}{V^C(t,\,t{-}1)},\]

where \(T_t\) nets contributions and withdrawals of cash and securities alike, the in-kind leg valued by convention — and the convention is policy, because a mispriced transfer in is manufactured performance. Flow timing is a convention too: end-of-day here, Modified Dietz weighting in general, valuation at every large external flow under GIPS. The denominator of a return is a policy document. One more choice hides in the base: computed against the restated \(V^C(t,t{-}1)\), the return answers what happened; against the original print, what was published. Either is defensible; mixing them is not — a chain of returns whose bases are drawn from different rows compounds a vintage that never existed, which is the Restatement section's epoch rule wearing arithmetic.

On the quiet day the innovation is a null experiment. Nothing was traded, so the execution term vanishes identically and \(\nu_t\) collapses to the other three: the restated base, the surprises in income, fees, and flows, and the new print's own \(\eta_t\). A desk that watches the quiet days learns its sources' noise floor and revision habits before the loud days arrive — the filter tuned on silence.

Now trade. Fills arrive intraday at prices \(\tilde p_f\), in signed quantities \(\Delta q_f\), with explicit costs \(c_t\) — commissions, exchange and regulatory fees — and the evolution gains a term:

\[V^C_t \;=\; V^C(t,\,t{-}1) \;+\; \underbrace{\textstyle\sum_i q_{i,t-1}\,\Delta p_{i,t}}_{\text{holding}} \;+\; \underbrace{\textstyle\sum_{f}\, \Delta q_f\,\big(p_{i(f),t} - \tilde p_f\big)}_{G_t:\ \text{execution to close}} \;-\; c_t \;+\; I_t \;-\; F_t \;+\; T_t,\]

where \(i(f)\) names the instrument of fill \(f\). Two details in the execution term carry more than their weight. It sums over fills, not over net changes: an intraday round trip nets to zero shares and not to zero dollars, and a book that computes trading PnL from \(\Delta q_i\) has defined the round trip out of existence. And it is the formalism's \(\tilde a_t - a_t\) given a ledger: between the order as sent and the position as settled stand the fill, the allocation, and the affirmation — five assertions with five knowledge times, of which the custodian ever sees only the last. Your average price against the broker's per-fill prices, embedded against explicit commissions: the execution clause of tomorrow's reconciliation is a stack of small cross-source residuals, each with a name. Settlement adds the wedge: under T+1 the fills move your units today and the custodian's settled book tomorrow, and between the two sits a payable-or-receivable that exists for one day by design — trade-date and settlement-date accounting are two bases telling the truth on different schedules, the ex-to-pay shape again with a shorter fuse.

Note also what the identity refuses to see. Sells crystallize PnL against particular lots, but value is indifferent to the partition of PnL into realized and unrealized — the total ties whether the lots are right or scrambled. An account can reconcile to the penny with its tax lots corrupted, its wash sales miscounted, its sleeves misassigned: the identity holds while the lots lie. File the observation; it becomes a theorem shortly.

Sell short and the identity grows a financing wing. Let \(q\) be signed, so the holding term already marks the shorts — price up, value down — and write

\[V^C_t \;=\; V^C(t,\,t{-}1) \;+\; \sum_i q_{i,t-1}\,\Delta p_{i,t} \;+\; G_t \;-\; c_t \;+\; I_t \;-\; \Pi_t \;+\; \Phi_t \;-\; F_t \;+\; T_t.\]

\(I_t\) is now entitlement income on the longs only. \(\Pi_t\) is its mirror: payments in lieu owed on shorts held over the record date (with a twin on the receiving side, if the longs are lent) — cash-identical to a dividend and tax-distinct from one, a distinction the client's 1099 will insist on even though the value identity cannot. \(\Phi_t\) collects the rate-times-balance terms: credit interest on free cash; rebate on the short proceeds held as collateral — the reference rate minus the borrow fee, negative for hard-to-borrow names, where the rebate is a fee wearing a credit's clothing; margin interest against the debit. Every one of these is a product of two bitemporal factors on different clocks: the balance revises with settlement, the rate with the prime broker's rerate, and the accrual booked daily meets the monthly statement's true-up like a first print meeting its revision. The algebra of that meeting is the product rule,

\[\Delta(\rho K) \;=\; K\,\Delta\rho \;+\; \rho\,\Delta K \;+\; \Delta\rho\,\Delta K,\]

a restated balance under a rerated rate — so even attributing the revision of one accrual is itself a small decomposition problem. And the terms couple through the balances: one failed settlement moves the debit, which moves the margin interest, which moves the collateral, which moves the rebate. The identity is linear; the evidence graph is not. One event, three clauses.

Now read the identity as a census, asking of each term: who witnesses it first, who is bound by it, and how does it revise. Your own \(\Delta q_f\) you witness at the fill; the custodian meets it at settlement — belief leads record on everything endogenous. The world's terms run the other way: fees, journals, client flows surface in the custodian's file before your book has heard of them. Entitlements answer to the issuer through its agent — announced, amended, occasionally reversed. Marks answer to a pricing hierarchy that restates its bad closes. The financing terms answer to a prime broker who behaves like an accountant: accrue now, true up monthly. Fees answer to your own schedule, which is to say to a computation. Hardly any two terms share an authority, a lag, and a revision contract — and that, not vendor history, is why the books are plural. A book of record is an evaluation policy: an assignment, to every term of the identity, of a source and a knowledge cutoff. The alphabet is four extreme policies. IBOR evaluates each term at its freshest witness and fills the still-empty cells with estimates — the identity, nowcast. CBOR restricts to what the record has settled — authoritative and late. ABOR evaluates on accrual basis and keeps rolling toward the bottom row — willing to be wrong now in order to be right later. OBOR evaluates once and pins — the row against which a decision can be audited. Fresh, settled, right, still: the identity offers four virtues and forbids their conjunction, because its terms revise on clocks no single cutoff can flatter. A golden book is not an economy; it is a request that the custodian, the accountant, and your own blotter share a clock. The books are plural because the clocks are.

Why, then, does the difficulty compound rather than settle? First, because the headline reconciliation is underdetermined: the observable residual \(\delta\) against any counterparty is one equation, and the identity has just supplied ten unknowns. A total that ties is weak evidence — term errors offset, and an offsetting pair is two breaks wearing a zero. Identification requires a witness per term, which is what the contract-based reconciliation promised by the formalism exists to supply: the identity is the contract schema, each term a clause, each clause bound to its own evidence on its own clock. Second, because the null hypothesis is not zero. The bases disagree by design — the receivable between ex and pay, the settlement wedge between T and T+1, accrual against cash everywhere — so the expected residual is not \(\delta = 0\) but a bridge with a term structure, born on schedule and dying on schedule. A monitor that alarms on the bridge trains its operators to ignore alarms, and an operator trained to ignore alarms is the last control, disarmed. Breaks are departures from the curve, not from zero. Third — the filed observation — because value is a linear functional on a richer state than it measures. Let \(\ell\) be the full ledger: lots with basis, tax character, sleeve assignment, balances by type. Reconciling \(V(\ell)\) controls error only up to the kernel of \(V\): any corruption that slides the ledger along a level set — realized swapped for unrealized, a dividend relabeled a payment in lieu, a lot reassigned across sleeves — is invisible to the tie. The dangerous breaks live in the kernel of the headline. The remedy is the formalism's triangulation made literal linear algebra: reconcile a family of functionals — units by security, cash by balance type, lot counts, character totals, sleeve sums — chosen so that their kernels intersect trivially on the error directions that matter. A second witness is valuable for the orthogonality of its bias; a second functional, for the transversality of its kernel. Fourth, because revisions do not always arrive as new cells. A late-discovered corporate action rewrites a block of a column at once; an elective action makes even the sign of the pending term a choice someone must make and record; a rerate re-prices an accrual whose balance was itself restated. The triangle's cells are not independent, and some revisions are operators on its history, not entries in it.

Multiply it out — terms by sources by bases by clocks — and the product is not an accident of bad vendors; the product is the business. The engineering that follows does not shrink it. It names every factor: a source of truth per term, a revision contract per source, a bridge per basis, a witness per kernel direction. A value, at scale, is not a fact. It is a quorum.

An implementation

Two columns, four doors

The triangle costs exactly two columns, and the second is never the application's to write. Here is the emitted DDL for one fact family:

create table ledger.price (
    source      text not null,
    security_id text not null,
    date        date not null,                                   -- observation date t
    updated_at  timestamptz not null default clock_timestamp(),  -- knowledge time s
    close       numeric,
    retracted   boolean not null default false,   -- explicit, not a null pun
    primary key (source, security_id, date, updated_at)
);
revoke update, delete on ledger.price from public;

Three refusals in a dozen lines. revoke update, delete makes the ledger append-only by grant rather than by good intentions. Retraction is an explicit flag rather than a null tombstone — with several value columns the pun turns ambiguous — so a dead fact is a fact about death, with its own knowledge time. And updated_at defaults to clock_timestamp(), now() is pinned at transaction start; clock_timestamp() at the call. Primary keys should not collide merely because two facts shared a transaction. while the write door, api.assert_price(...), takes every column except that one. A witness never writes its own timestamps.

The doors make the correct reading the default reading:

create view api.price as                    -- best knowledge: the bottom row
select * from (
    select distinct on (source, security_id, date) *
    from ledger.price
    order by source, security_id, date, updated_at desc
) t where not t.retracted;

create function api.price_asof(sigma timestamptz)   -- x(σ, ·): the epoch's read
returns setof ledger.price language sql stable as $$
    select * from (
        select distinct on (source, security_id, date) *
        from ledger.price
        where updated_at <= sigma
        order by source, security_id, date, updated_at desc
    ) t where not t.retracted
$$;

Both doors read latest-then-not-retracted, so a retracted fact vanishes from every vintage after the retraction and from none before it — history stays honest about what was believed. And the doors are manufactured, never handwritten: the whole migration is for spec in RECORDS + TRUTHS: conn.execute(spec.ddl()), which is why _in_time means the same thing on prices as on fills as on weights. The suffix is the type system SQL doesn't have, because it is the same code everywhere.

The earlier caveat — knowledge moves at T+ε — now has a function body:

create function meta.safe_sigma(grace interval default '5 seconds')
returns timestamptz language sql stable as $$
    select least(
        now() - grace,
        coalesce((select min(xact_start)
                  from pg_stat_activity
                  where backend_xid is not null
                    and pid <> pg_backend_pid()), now())
    )
$$;

A straggler can commit late carrying an early timestamp, so the watermark sits behind the oldest in-flight writer. Only behind it is \(x(\sigma, \cdot)\) a pure function of \(\sigma\) — which is the purity every epoch below will spend.

The pin is a primary key

OBOR's contract — never revised, only superseded — is not a policy in this schema. It is a key:

create table truth.obor_term (
    account_id text not null,
    term       text not null,
    date       date not null,
    pinned_at  timestamptz not null default clock_timestamp(),
    amount     numeric, source text, sigma timestamptz,
    primary key (account_id, term, date)   -- revision is falsification; the key enforces it
);

No knowledge axis, because the pinned plane refuses to have one: date advances and columns do not. A second write to the same key is an integrity error, and the error is the contract — revising the operating book is not merely forbidden but unwritable. The matching evaluator is the one place \(\sigma\) is deliberately dead:

@evaluator("value", "OBOR")
def obor_value(conn, account_id, date, sigma):
    """The pinned row. sigma is accepted and ignored: that is the point."""
    return val(conn, """
        select amount from truth.obor_term
        where account_id = %(a)s and date = %(d)s and term = 'value'
    """, a=account_id, d=date)

A book of record, in this code, is data: evaluators register per (term, source), and the four books are four tuples of clauses over the registry — fresh, settled, right, still, as the value identity ordered them. Assembly runs under the promise the sleeves subsection deferred — one writer to the shared account:

class overlay_writer:
    """One writer to the shared account: every sleeve's intentions pass
    through here, serialized by an advisory transaction lock."""
    def __enter__(self):
        execute(self.conn,
                "select pg_advisory_xact_lock(hashtext(%(a)s))",
                a=self.account_id)
        

\(n\) sleeves, \(n\) policies, one state, one lock. The concurrency section, in a line of Postgres.

Reuse is a replayed trace

Coherence is derived, so derivation must be honest about its inputs. A rule receives a context and cannot read outside it; every read leaves a fingerprint: The verifying traces of Build Systems à la Carte — the Restatement section's citation, now cashed: the machinery of cloud build caches, Shake to Bazel, in about fifty lines against two tables.

class _Trace:                    # the context handed to a rule
    def facts(self, sql, **params):
        got = rows(self._b.conn, sql, sigma=self._b.epoch.sigma, **params)
        self.entries.append({"t": "facts", "sql": sql,
                             "params": params, "h": _hrows(got)})
        return got

    def dep(self, kind, **key):
        payload, cache_key = self._b.get(kind, **key)
        self.entries.append({"t": "dep", "kind": kind, "key": key, "h": cache_key})
        return payload

    def param(self, *path):
        v = self._b.params
        for p in path:
            v = v[p]
        self.entries.append({"t": "param", "path": list(path), "h": _h(v)})
        return v

Fact reads hash their result sets; dependencies record the child's cache key; parameter reads hash a slice of \(m\). Reuse under a new epoch is then a replay, not a guess:

def _verifies(self, trace) -> bool:
    for e in trace:
        if e["t"] == "facts":
            got = rows(self.conn, e["sql"], sigma=self.epoch.sigma, **e["params"])
            if _hrows(got) != e["h"]: return False
        elif e["t"] == "dep":
            _, child_key = self.get(e["kind"], **e["key"])
            if child_key != e["h"]: return False
        elif e["t"] == "param":
            
    return True

The recorded queries rerun at the new \(\sigma\) and must hash identical. Two consequences fall out rather than being computed. The dirty cone is exact: a corrected mark breaks precisely the fact-read hashes downstream of it, and yesterday's valuation replays clean. And the cache crosses epochs: two cutoffs that disagree as timestamps but agree on every fact a rule touched share an artifact — content beats coordinates.

A rule, for shape — note where the denominator comes from:

@rule("return.daily")
def return_daily(ctx, account_id, date, prev):
    v1 = num(ctx.dep("valuation", account_id=account_id, date=date)["value"])
    v0 = num(ctx.dep("valuation", account_id=account_id, date=prev)["value"])
    t  = num(ctx.dep("flows.net",  account_id=account_id, date=date)["value"])
    timing = ctx.param("returns", "flow_timing")
    denom = v0 + t if timing == "start_of_day" else v0
    return {"value": str((v1 - t - v0) / denom)}

flow_timing is read from \(m\), so changing it is a new methodology — a different triangle — never an edit to this function. The denominator of a return is a policy document, and here the document is enforceable.

The Restatement section promised that the impact study is a select statement. Kept, literally:

select coalesce(a.kind, b.kind) as kind,
       coalesce(a.key,  b.key)  as key,
       a.cache_key as in_a, b.cache_key as in_b
from      (select * from build.manifest where epoch_id = %(a)s) a
full join (select * from build.manifest where epoch_id = %(b)s) b
       on a.kind = b.kind and a.key = b.key
where a.cache_key is distinct from b.cache_key;

A curve, not zero

The reconciliation runner's inner lines, trimmed to the live branch — a missing witness short-circuits to watch above them, because lateness is data, not an exception:

v = EVALUATORS[(term, sv)](conn, account_id, date, epoch.sigma)
w = EVALUATORS[(term, sw)](conn, account_id, date, epoch.sigma)
delta    = Decimal(v) - Decimal(w)
br       = sum((BRIDGES[b](conn, account_id, date, epoch.sigma) for b in bnames), ZERO)
residual = delta - br                       # the alarm variable
status   = "ok" if abs(residual) <= tolerance else "break"

The bridge is computed from the ledger, never stored as a number someone once believed:

@bridge("settlement_wedge")
def settlement_wedge(conn, account_id, date, sigma):
    cd = newest_cbor_date(conn, account_id, sigma)
    
    for sec, qty, px_fill, costs in rows(conn, """
        select security_id, qty, price, costs from api.fill_asof(%(sigma)s)
        where account_id = %(a)s and date > %(cd)s and date <= %(d)s
    """, a=account_id, cd=cd, d=date, sigma=sigma):
        total += Decimal(qty) * (price_at(conn, sec, date, sigma)
                                 - Decimal(px_fill)) - Decimal(costs)
    return total

Born at execution, dead at settlement, on schedule — the null hypothesis with a term structure. The kernel argument of the value identity also compiles. The sleeve ledger has no external counterparty, so its one law is its only witness, and this functional sees exactly the reassignments the headline cannot:

def sleeve_sum_law(conn, account_id, date, sigma):
    account = units_by_security(conn, account_id, date, sigma)     # the record
    sleeved = {sec: Decimal(u) for sec, u in rows(conn, """
        select security_id, sum(units) from api.sleeve_units_asof(%(sigma)s)
        where account_id = %(a)s and date = %(d)s group by security_id
    """, a=account_id, d=date, sigma=sigma)}                        # your designation
    return [{"security_id": s, "orphaned": account.get(s, ZERO) - sleeved.get(s, ZERO)}
            for s in sorted(set(account) | set(sleeved))
            if account.get(s, ZERO) != sleeved.get(s, ZERO)]

Last, the innovation, with \(\eta\) defined the only honest way — as what the evidence fails to explain:

def innovation(observed_close, actual, expected):
    base_revision = actual.base - expected.base_as_issued            # the silent opener
    execution     = (actual.execution - actual.costs) - expected.execution
    exogenous     = actual.world() - expected.w_hat                  # the world's surprise
    nu            = observed_close - expected.close()
    observation   = nu - base_revision - execution - exogenous      # eta, by construction
    return Innovation(base_revision, execution, exogenous, observation)

Four addends, four different conversations: a data conversation, an execution conversation, a market-and-operations conversation, and a data-quality conversation. The morning meeting, typed.

Notice, finally, what the platform declines to make writable. The operating book has no update path. A client's own timestamp has no column. A drift of a drift has no code path, because the arbiter's input is always the record's frontier. Lookahead has no door, because every read below the epoch band carries one \(\sigma\). The earlier sections could only recommend these disciplines; the schema refuses them. That is what natively means — the prose demoted to comments, and the aphorisms promoted to constraints.

The atomic transaction framework

An enterprise seldom acts upon the world directly. It instructs a custodian, routes to a broker, delegates to a middle office. Often it should: these functions benefit from specialization, scale, and regulatory moats.

But the moment an instruction crosses a firm boundary, one is executing a distributed transaction without a coordinator, a shared clock, a common log, or a common runtime. There are only messages moving through channels that may delay, duplicate, drop, or reorder them.

Lamport showed that in such a system before and after are constructions, not observations. The Two Generals Problem showed that no finite exchange over an unreliable channel creates common knowledge. Finance answered these theorems in its customary fashion: it lowered the standard from certainty to evidence and attached a deadline.

This answer is not foolish. It is sufficient—but only when engineered deliberately.

Whenever a unit of work is handed to a counterparty, five questions require explicit answers.

← Back to all posts