scriptkittyos & labs requisition in production

sanction

Proof that the work happened.

A consequential act, by a person, a device, or an AI agent, becomes a signed record of who did it, under what rule, with which witnesses. Anyone can check it, without trusting the institution that produced it.

requisition receipt device DL-2291 · seq 00418
actionfiremain.valve.repair
boundat proposal · re-verified at execution

rule in force work complete, then CDI inspects, then QAR inspects three distinct people · inspectors present at the equipment

witnesses
humanMM2qual MMpresent14:02Z
humanAD1qual CDIpresent14:38Z
humanAZ1qual QARpresent15:11Z
machinediag-70.83n/aconstrains only · never independent
SATISFIED verifiable offline
against the published key

the problem

One failure, in every institution that keeps its own records.

The party who benefits from the record being wrong is the party who controls the record. A sailor signs his own maintenance check. A guard logs his own rounds. An agency decides what its own release contains. An agent reports what it did. The rules requiring a second person already exist and are written down. They are attested by the institution being reviewed, in a form nobody outside can verify.

87% of the USS Bonhomme Richard's fire stations were in inactive maintenance status the morning of the fire that destroyed her. PACFLT command investigation, 2021
$285B estimated DoD deferred maintenance backlog for FY2025, up from $137B in FY2020. GAO-26-107255
264% increase in Coast Guard cutter maintenance deferred since FY2018. $179M in FY2024. GAO-25-107222

The Navy has a word for the result. Gundecking: signing off checks that were never performed. It is not laziness. It is what happens when the workload exceeds the time available and the record is the cheapest thing to fake.

how it works

One mechanism, configured per deployment.

The record is a byproduct of doing the work, not a form filled out afterward. Where paperwork costs more than the task, the paperwork gets skipped or falsified. So the proof is produced by the act itself.

Witness

Every approving act is a signed record carrying attributes someone can vouch for: identity, qualification, billet, presence at the equipment, time. A person, a device, a decoder, or an agent can each be a witness. They are not interchangeable.

Predicate

Each deployment declares, per class of action, what an independent second act means. The gate asks one question: does this witness set satisfy the rule in force? A rule nobody on the current roster could satisfy is refused when it is configured, with a reason, instead of at three in the morning.

Receipt

Every outcome, executed or refused, seals a receipt recording the rule, the witnesses, and the evaluation trace. A stranger can re-run the decision offline and reach the same verdict without asking permission.

Three rules the language enforces, not convention

  • The proposer is never a witness. Naval aviation's "no one inspects their own work", made structural.
  • A machine witness never counts toward independence. A diagnosis can constrain what is reachable. It cannot stand in for a person.
  • A neural act is never sole for an irreversible action. Configurable, but the shipped default refuses, and the receipt records the choice.

What never loosens

  • Approval binds to the exact bytes and is re-verified at execution. The same primitive as PSD2 dynamic linking and FIDO transaction confirmation, generalized beyond payments and extended to multi-party approval.
  • Policy only narrows. An advisory hold can escalate but never approve. If the gate is unavailable, nothing runs.
  • Silence is a finding. The schedule is declared, so a check that never opens is as visible as one that fails.

four systems

Each stands alone. Composed, they are Sanction.

A customer running one part gets a real property, not a fragment. Which components a deployment takes is configuration, never a fork.

Requisition

accountability plane

Witnesses, independence predicates, byte-bound approval re-verified at execution, and signed, chain-bound receipts a stranger verifies offline against a published, append-only key registry. An agent proposal is one input among several; strip the agent out and the mechanism is unchanged.

In productionLicensed from Script Kitty OS & Labs with possible donation

Ultraviolet

purple-team SOC

A security operations center on the BEAM. Sensors, detection, replay, investigation, human approval as a plain workflow, and execution behind dry-run-default adapters. Complete on its own, with no authority plane in the picture.

Open source · Apache-2.0Intended for possible donation

Trinity

governed agent

A personal AI agent that runs on your machine, remembers you, learns procedures, and acts through tools. Its authority is pluggable: under Requisition it can only propose, and every proposal is a governed act with a receipt.

Soon to be open source · Apache-2.0Intended for possible donation

beam_mcp

transport

Model Context Protocol for the BEAM. It carries calls and holds no authority. No approvals, no receipts, no masking, by design. Whatever wraps it decides those.

Open source · Apache-2.0Intended for possible donation
the platform

Sanction OS

Where Sanction runs. Not a member of it. The platform is what makes it something an organization can operate: rules configured without writing code, a tenant boundary, the evaluation harness run against the deployed build, registry operators chosen per deployment, and a path to accreditation.

The default is maximal and configuration goes downward, because a system that starts permissive and adds controls has a moment in its history when it was permissive. Self-hosted Sanction and platform-hosted Sanction OS are two offers, not one offer at two prices.

three layers

Each layer holds a guarantee the layer below cannot.

Requisition alone has limits Requisition alone cannot fix. The four together have different limits. The platform adds what neither can without one. Each claim is tested one configuration at a time.

Requisitionalone
The record is honest. Who did it, under what rule, with which witnesses, and that nothing was altered afterward. A lie about the past has to cohere with records the liar does not control, and is caught the first time physics disagrees.
what it cannot do alone Be its own independent observer. A device that reports work it never did, or a witness set that is all in on it, produces a record that is intact and wrong. Requisition proves the record was not altered; it cannot see what was never entered.
Sanctioncomposed
An independent observer exists. Ultraviolet is a second operator with a different threat model. Trinity makes the agent a governed actor. beam_mcp carries the calls and holds no authority.
what composition closes A detection can stand as a witness. An observer's own queries are consequential acts with receipts, so an observer that goes bad is attributable. What was claimed is compared against what was observed, by a party that did not make the claim.
Sanction OSplatform
An enterprise can run it. Predicates configured without code, and a rule change is itself a governed act with its own witnesses and its own receipt. A tenant boundary. Continuous authority to operate.
what the platform closes Registry operators chosen per deployment across independent institutions. Never a single command in a defense deployment, never a carrier in a trade one. A configuration with no evaluation row is not supported.

What it does not do.

It does not make judgments correct. Discretion, whether an emergency waiver, a departure from specification, or "as the commanding officer sees fit", is recorded as a decision by someone with the authority to make it, bound to the bytes, with a reason. The system never evaluates the judgment.

Two signatures do not stop rubber-stamping. The clinical literature is blunt that an independent double check is a weak safeguard when it is the only one. What deters falsification is comparing the claimed outcome against the observed state afterward.

Presence is graded, not absolute. A static code on a piece of equipment can be photographed and replayed. That is evidence, not proof, and the receipt says which level it holds. An action requiring more than an asset can attest is refused with the gap named.

where it runs

Five deployments. One rule registry.

These are not five products. They are five configurations of the same mechanism, and the shared failure is the one above: the party who benefits from the record being wrong controls the record.

Shipboard and unit maintenanceNavy 3-M, NAMP, Coast Guard SFLC

the receipt provesWho performed, who inspected, in what order, with what qualification, and that each was at the equipment. The repair is declared before it starts and compared against the equipment's observed state afterward. A check that never opens on schedule is flagged, not merely absent.

does not fixWhether the repair chosen was the right one. It proves the declared outcome was reached; it does not supply engineering judgment about what should have been declared.

Custody of released recordsFOIA, congressional production, court-ordered disclosure

the receipt provesCompleteness and integrity of a set without disclosing its contents. Each withholding emits its own signed record: item, exemption, deciding authority, date. The count of withheld items is checkable even when the items are not.

does not fixWhether the withholding was justified, or whether the right material was collected in the first place. It makes a decision attributable, not correct.

Detention and custodial careuse of force, medication, welfare checks, grievances

the receipt provesThat a record was not altered after the fact, that the person who signed held the qualification and was where the record says, and that the rounds due on the schedule were walked. Verification does not run through the facility.

does not fixWhether force was justified or care was adequate. It establishes what happened and who is answerable, for people who now have a record they can trust.

Payment authorizationwhere the primitive is already regulated and proven

the receipt provesApproval bound to the exact bytes and re-verified at execution. An approval that merely describes a transaction can be satisfied while the executed payload differs from the one approved; a bound one cannot.

does not fixFraud where every party is genuine and the transaction is intended. Binding proves what was approved, not that approving it was wise.

Clinical action with a brain-computer interfacethe case that forced the general mechanism

the receipt provesThat a person who acts by thinking has an independently checkable record that the thing they remember doing actually happened, with a clinician's approval alongside it where the action is irreversible. Confidence decides which classes of action are reachable, never whether a human keeps control.

does not fixWhether a decoded intent was the person's intent. Two decoders agreeing does not certify intent under a shared perturbation. Confidence bounds what is reachable; it never establishes truth.

evidence

Measured, not declared.

The benchmark, its frozen methodology and its per-trial results are public, so an assessor can examine the evidence before the system is installed.

0 unauthorized effects published deposit
61 attack trials across nine attack families frozen protocol
5.9% upper bound of the 95% confidence interval, stated rather than omitted [0.0%, 5.9%]
2 standalone verifiers, Elixir and dependency-free Python, that must never disagree offline · exit codes defined

Independent examiners hold a frozen specification of the authority boundary and test against it. The goalposts do not move.

The published benchmark carries its own errata. Corrections append and are dated; nothing above the correction line is edited. A defect found in a shipped verifier is a published row, not a quiet fix.

No change lands without a failing test first. A criterion that cannot fail is not a criterion, and a check that cannot fail is not a check.

Every claim on this page that a build cannot yet satisfy is a standing obligation, tracked by name. Fielded scope and accreditation status for a given environment are stated separately, on request.

Script Kitty Labs

Founded by Ayla Croft. The system came out of years spent breaking the model.

8 system cards The ART agent red-teaming benchmark, from the paper Ayla co-authored (arXiv:2507.20526), is cited in eight frontier model system cards.
DEF CON 33 She designed and ran the in-person AI red-teaming competition. Then she quit to start this company, because the result was always the same: the model breaks. If the model is always breakable, the thing that holds has to be the system around it.
Built on the BEAM Elixir/OTP supervision and isolation, chosen because accountability infrastructure has to stay up while everything around it fails. Ayla's published research.

Requisition was developed under the working name HolyTrinity; the published benchmark and the position paper The Model Proposes, the System Authorizes keep that name. ONE-Bench, an open benchmark for brain-computer-interface control pipelines, is published alongside.

deployment posture

The questions a program office asks second, answered first.

Device
Government-furnished, managed, with derived credentials on the device. No personal devices, no unmanaged storage. Media stays in the managed store; the record holds only its hash, so anything later determined to be controlled can be purged without breaking the chain.
Signatures
Two signatures, two jobs. The witness's identity signature answers who, and chains to department PKI. The chain signature answers whether anything was altered, and verifies offline against a published registry. The algorithm is a recorded field, so a national security system runs a suite-approved cipher without a redesign.
Disconnected
Every device keeps its own signed chain and never rewrites it. Chains merge on reconnection with inclusion and consistency proofs. Ordering is logical rather than clock-based, so a ship under emission control produces a record that reconciles without a server and verifies without trusting one.
Borders
A verdict can cross a jurisdiction while the data stays home. Relay receipts carry the result and its depth, never the payload.
Configuration
A unit configures its own rules without writing code. A proposed rule is checked before it takes effect: satisfiable, sourced, reachable given the actual roster, and no looser than the rule it replaces. Emergency authority is a declared path with a signed record of who invoked it and a required reconstitution afterward.
Accreditation
Containerized for a software factory with continuous authority to operate, targeting the impact level that controlled unclassified maintenance data requires. Simulation establishes what the plane does on a class of action; hosting, registries and attestation are stated per environment.

timing

The deadline is regulatory.

OMB M-25-21

Federal agencies must identify and manage high-impact AI, and use cases that fail the minimum practices must be discontinued. The reporting deadline is September 2026.

whitehouse.gov, April 2025

CISA and Five Eyes agentic AI guidance

Six national cyber agencies jointly name accountability gaps, meaning the inability to trace decisions, audit actions, or assign responsibility when agents act on their own, as one of five risk categories for agentic AI. They treat governance and human oversight as prerequisites, not options.

Careful Adoption of Agentic AI Services, cisa.gov, May 2026

the foundation

Yes, the cats are real.

Script Kitty runs alongside rescue cats and open-source pet tech. The Script Kitty Foundation is the part of this that does that, and one day the platform gets a physical home where research and fostering share a room. That page is not up yet.