Agents — getting your data AI-ready without giving it the keys.

Agents

AI-readiness is a data-readiness problem.

Everyone is buying models. Almost nobody is asking what the model is reading. A model amplifies whatever you hand it — so every extra copy in your stack is a place where it reads garbage, fluently and at your per-token rate. The xerocopy discipline is how you get the most out of AI: fewer copies, one source of truth, contracts at the boundary.

01

AI-readiness is data-readiness.

Everyone is asking which model to buy. The honest answer is that the model is the least of it. An LLM does not clean up what you hand it — it amplifies it, including the copies. If the retrieval layer serves a stale export, a shadow table with different null semantics, or two sources that quietly disagree, the agent reads all of it with equal confidence and answers fluently from the wrong one.

This is why xerocopy starts on the data side, not the model side. You don't make data AI-ready by embedding harder; you make it AI-ready by removing everything the model shouldn't have to read. One source of truth, one contract, one owner answering for it — the same discipline that keeps a warehouse honest is what keeps an agent honest. Get the data right and almost any model will do. Get it wrong and no model will save you.

The model amplifies your data. If your data is a pile of copies, the model amplifies the pile.
02

Without the discipline, you pay tokens for garbage.

Every extra copy of your data is a bill you pay three times. Once in storage. Once in the sync jobs that keep the copies alive. And once per token — because whatever exists gets embedded, whatever gets embedded gets retrieved, and whatever gets retrieved gets shipped into the prompt. The vector index nobody reconciled, the mart with yesterday's numbers, the extract with different null semantics: the model doesn't know they're garbage. It reads them like gospel, at your per-token rate.

Contradictions are worse than gaps. When two copies disagree, the agent doesn't flag it — it picks one, confidently, and narrates its reasoning. So you pay for the tokens, then pay again for a human to check the answer, which is the very cost the agent was supposed to remove. Fewer copies isn't a storage optimization. It is the difference between an agent that answers and an agent you have to audit.

Garbage in, garbage out — except with agents, garbage in is billed per token.
03

Least copy is least blast radius.

The same argument runs in the security direction. Every place a row exists is a place it can be exfiltrated; every tool an agent holds is a capability it can be tricked into using. Bhattarai & Vu's "Lethal Trifecta" — untrusted input, sensitive data, and external side-effects colocated in one context — is exactly what a stack full of copies hands you. AgentGuardian's answer is the same instinct at the capability level: grant the smallest set, deny the rest by default.

Fewer copies, fewer capabilities, narrower contracts. The stack that is cheapest to operate is also the cheapest to defend — and the smallest retrievable corpus is the smallest prompt to poison.

See: [1] [2]

What the data never touched, it cannot leak.
04

When agents act, the contract acts first.

Reading is only half of it. The 2026 governance literature — deterministic pre-action authorization, natural-language policy compiled into checks a machine evaluates, attestation of actions rather than agents, cryptographic proof of what actually ran — converges on one claim: the model proposes, the policy commits. Every tool call copies authority from your governance into a runtime, and that copy has to sit under a contract you wrote on purpose, checked outside the model, before the side effect happens.

And the evidence has to outlive the agent's own transcript. If the runtime and the guardrail don't emit verification a third party can check, you don't have governance — you have logs. For a data platform this is a familiar request: emit the alarm with the row; record which policy version approved which write, at the moment of the write. Same discipline, one boundary up the stack.

See: [3] [4] [5] [6] [7]

The model proposes. The policy commits. Never the other way around.
05

The zero-copy discipline is the AI-readiness spec.

Strip away the vocabulary and xerocopy was never only about saving storage. It is about refusing intermediate states — copies that drift, schemas that diverge, truths that disagree. An AI stack built without that discipline is a pile of intermediate states with a language model standing on top, reading all of them with equal confidence, and billing you for the privilege.

Concretely, here is how our principles translate:

No intermediate buffers
The agent queries the system of record in place. No nightly export to a vector store that silently disagrees with the warehouse by Tuesday — and no tokens spent reconciling the two.
One source of truth
Every retrieval resolves to a row with a lineage, an owner, and a freshness contract. If two sources disagree, the agent doesn't get to pick — neither does it get asked.
Contracts at the boundary
Schemas, freshness, and null semantics are declared, versioned, and checked — because an LLM will not notice an ambiguous column; it will confidently narrate it.
Least copy, least blast radius
What the data never touched, it cannot leak. Smaller surfaces mean smaller retrievable corpora, which means smaller prompts to poison — and smaller prompts to pay for.
Publish deltas, not copies
Change feeds and in-place reads keep the model's view within one policy boundary of the truth — the same mechanism that keeps analytics honest keeps agents honest.
AI-readiness isn't a layer you add on top of your data. It's the property your data has when there is nothing in between to distrust.

References

The 2026 agent-governance shelf.

All entries are arXiv preprints from 2026 and are not yet peer-reviewed. Included because the ideas travel, not because the venue confers authority.

  1. [1]

    Uchibeke, U.

    Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents

    arXiv:2603.20953 (2026) — Open Agent Passport (OAP)

  2. [2]
  3. [3]
  4. [4]

    Palumbo, N. et al.

    Policy Compiler for Secure Agentic Systems (PCAS)

    arXiv:2602.16708 (2026) — Datalog-derived enforcement

  5. [5]
  6. [6]

    Bhattarai, M. & Vu, M.

    Trustworthy Agentic AI Requires Deterministic Architectural Boundaries

    arXiv:2602.09947 (2026) — the Lethal Trifecta

  7. [7]

    Jin, X. et al.

    Proof-of-Guardrail in AI Agents and What (Not) to Trust from It

    arXiv:2603.05786 (2026) — TEE-backed proofs

Coda

Before you point an agent at your warehouse, delete the copies it shouldn't read.

We help teams get to AI-ready from the data side — one source of truth, contracts at the boundary, and no tokens spent on garbage. Audit the data and controls first, or build a bounded workflow when the source is ready.