The Flagship Product
Charlie Owl, the Nightingale OS accountability companion

Charlie Owl

The evidence layer for governed hospital AI. Charlie Owl keeps the receipts for every algorithm that touches a patient — what the model said, what the clinician decided, who attested — so a surveyor, a board, or a court gets a record instead of a reconstruction. The same spine holds the designations: the verified promise that a hospital's stroke, trauma, cancer, or maternity care works — Charlie Owl keeps it continuously provable, protecting the patients it covers, the community's access to that care, and the revenue it carries. It keeps the receipts for every algorithm that touches a patient, too — running on documents, attestations, public data, and exports you already own. No EHR integration required.

Screenshot of the Owl Designation Ledger — a readiness board of criteria marked green, amber, or red, each with its evidence or its gap attached
The Designation Ledger — the flagship readiness board
AI Governance — the Evidence Layer

When the algorithm is wrong,
someone kept the receipts.

Boards, counsel, and Joint Commission responsible-AI expectations are asking one question — and so is every discovery request now landing on a health system: what did the model say, who decided, and can you prove it. Charlie Owl is the answer, produced continuously. Deploy the AI, because you can produce the record.

Algorithm Inventory

Every model touching patients gets an entry: what it is, what it claims, who owns it, and where it runs. The EHR's sepsis score, the vendor's deterioration index, the scheduling optimizer — inventoried, not assumed.

Human-in-the-Loop Evidence

For each recommendation: what the model said, what the clinician decided, and who attested. When the question comes — from a surveyor, a lawyer, or a family — the answer is a receipt, not a reconstruction.

Calibration & Drift Monitoring

Vendor models get graded against their own outcomes. Subgroup calibration, drift over time, and a scorecard the quality committee can read — because "the vendor says it works" is not governance.

Tamper-Evident by Construction

Every decision lands in a hash-chained, append-only ledger — each row referencing the last, completeness-checked on a schedule. Nobody can quietly rewrite what the model said last Tuesday.

Governance in the moment, not only in the record

Grading a vendor’s model after the fact is one half of AI governance. The other half is what happens while software is being written — increasingly, written by an AI coding agent inside a clinical codebase. Nuthatch evaluates each agent write at the tool call, before the bytes land, and receipts every evaluation to a hash-chained ledger: what was attempted, which control matched, who the actor was, what the verdict was. Non-compliance becomes immediate, attributed, and provable. Every control in the starter pack starts by asking a human; blocking controls are the next phase, and each earns the right to refuse on its own only on demonstrated evidence. Meet Charlie Nuthatch ↓

Screenshot of the Owl algorithm-accountability scorecard — vendor models graded against their own outcomes, with subgroup calibration and drift over time
Algorithm accountability — the vendor-model scorecard and calibration wall
Why Now — The Record Is the Defense

Every theory in court right now
ends in the same demand: produce the record.

Clinical AI stopped being a governance exercise and became a discovery exercise. Two waves are running at once, and neither is answered by a vendor's assurance — only by what the hospital itself can show. Cited as filed, not as decided.

Wave one — the payer algorithms

Suits over automated coverage denial are past the pleadings. Lokken v. UnitedHealth drew a March 2026 discovery order on the nH Predict model; in Kisting-Leung v. Cigna the ERISA claims survived against a review averaging 1.2 seconds per file. The $556M Kaiser settlement is a chart-review False Claims Act matter, not an AI case — we don't count it as one.

Wave two — the scribes and the agents

Saucedo v. Sharp (Nov 2025) is the first AI-scribe class action against a health system, over 100,000+ encounters. Washington v. Sutter Health / MemorialCare (Apr 2026) plead CIPA at $5,000 per interception. Winters v. OpenAI (Jul 2026) is the first missed-diagnosis claim against an AI agent. The exposure moved from the payer to the bedside.

Regulators ask the same question

The Texas Attorney General's Sept 2024 settlement with Pieces Technologies turned on hallucination-rate claims the vendor could not substantiate. Enforcement and litigation converge on one question, and it is a records question.

What has actually worked

In Lisota v. Heartland (dismissed Jan 2026), the documented operating posture was the defense — governance in practice, not a disclaimer in a contract. Charlie Owl is that posture, kept continuously: inventory, human-in-the-loop evidence, calibration, and a hash-chained receipt for each.

The Trust Gap

The frontline doesn’t trust ungoverned AI.
It’s right not to.

Nurse AI use nearly tripled in a year (15% to 44%, Incredible Health 2026) while 83% still say the output isn’t reliable enough to act on unchecked, 60% of RNs don’t trust their employer to put patient safety first in AI deployment (NNU 2024), and 40% of nurses under outcome-prediction scores cannot override them. The same surveys name what builds trust — and each one is an Owl mechanism, not a promise.

Consulted, not surprised

Nurses consulted when AI tools are chosen report 74% trust versus 38% when they never are — and only 8% of employers always include the frontline (Incredible Health 2026). Owl is built and governed by clinicians; deployment starts with the people whose license is on the line.

Cited, not asserted

Over 60% of clinicians say transparent citations to peer-reviewed evidence would raise their trust in AI (Elsevier 2026). Every Owl answer carries its citation to the active standard or published evidence — never generated prose. When it declines, the refusal names the policy.

Validated over time, not at launch

88% of physicians rate “validation of safety and efficacy by a trusted entity, monitored over time” the top adoption factor (AMA 2026). Owl’s algorithm scorecards are exactly that: calibration and drift, re-checked continuously, kept as evidence a committee can read.

Overridable, and the override is honored

40% of nurses under AI outcome-prediction scores say they cannot override them (NNU 2024). In Owl the model only proposes; a named clinician decides and attests, and the hash chain records both — the override is evidence, not insubordination.

Not just Joint Commission.
The spine 148 programs share.

A hospital answers to far more than one accreditor, and no one holds the portfolio in a single place. Owl does. The Joint Commission gets the deepest treatment — a tracer engine and mock-survey mode — but the same continuous-readiness spine carries every body below. This is the “no one holds the portfolio” wedge, made concrete: each of these is covered in the product today, in honest tiers — 50 programs modeled end-to-end with 583 individually-cited survey criteria behind a live readiness score, 58 pre-cited certifier packs — 13 curated plus 45 derived from the modeled programs’ own cited criteria, 611 criteria with every one carrying a standards citation, and a 148-program designation roster — 43 regulatory, 39 certification, 35 accreditation, and 31 voluntary-excellence, each entry with a validated designating authority and honest tier — that a facility can register, attest, and revenue-model against. Breadth is never dressed up as depth.

Deemed-status accreditors

The Joint Commission (tracer engine + mock-survey), DNV / NIAHO, HFAP, CIHQ, and ACHC as alternative accreditor tracks, and CARF for behavioral-health and rehab programs.

CMS & federal

Conditions of Participation (42 CFR 482), the CMS-2567 Statement of Deficiencies and Plan of Correction, EMTALA, NHSN / CDC infection surveillance, MQSA mammography, and Section 1557 language access.

Pharmacy, lab & safety

DEA controlled-substance and diversion, SAMHSA OTP/MAT, 340B, USP 797/800 compounding, CLIA / CAP lab, and OSHA workplace safety.

Nursing, privacy & state scope

ANCC Magnet and shared governance, HIPAA / OCR privacy and Safe-Harbor de-identification, and an all-50-state scope-of-practice matrix.

Screenshot of the Charlie Owl compliance cockpit — the continuous-readiness home spanning multiple accreditation domains
The Owl cockpit — one home across the whole certifier portfolio

Coverage depth varies by body: the Joint Commission carries the full tracer and mock-survey engine, while several accreditors are catalog-level reference tracks today. A handful of compliance surfaces are not yet wired to a citation chain and are not claimed as citation-grade until they are.

The Designation Ledger

A designation is a promise,
not a billing code.

Trauma verification, stroke certification, cancer accreditation — each is an outside body's promise that a high-acuity service line actually works, and the license to bill for it. Losing one is a patient-safety event, a loss of community access to specialty care, and a multi-million-dollar, press-release event — in that order. Charlie Owl tracks every criterion, its evidence status, and what a lapse puts at risk, cited to the actual standards.

Continuous Readiness

A living readiness score per designation — every criterion green, amber, or red, with the evidence (or the gap) attached. Survey day becomes a walk-through instead of a scramble, because readiness was never allowed to decay in the dark.

Survey-Day Tooling

Live tracer management, Immediate Jeopardy assessment, document requests captured as they happen, and plan-of-correction workflows that gate a finding until the correction is evidenced. Drill mode runs the same walk before the real one.

Survey Mode, in Your Pocket

Walk the halls with the surveyor: tracer prompts, two-tap evidence pull, findings logged in real time — offline-capable, because surveys don't wait for hospital wifi.

Regulatory Change-Watch

When a standard changes, the criteria that depend on it flag themselves for review. Charlie Owl never silently fabricates a readiness score or a regulatory update — when it doesn't know, it says so, with a citation.

How Continuous Readiness Actually Works

A readiness score with four inputs,
on a spine you can verify.

“Living green/amber/red” is the outcome. Here is the machinery under it. Each accreditation domain carries a running 0–100 readiness score that moves on four measurable inputs, with the weakest domains flagged. Bands: ready ≥ 90 · watch 75–89 · at-risk < 75. The score is a prioritization aid — the product labels it exactly that, never as compliance evidence or a guarantee of survey outcome.

Open findings

Unresolved findings from tracers, drills, and prior surveys weigh the domain down until each is closed with evidence attached.

Criterion gaps

Every criterion a designation requires that has no evidence behind it yet — the standard's own checklist, scored against what you can show.

Stale & expiring evidence

Evidence carries a freshness clock. Competencies, policies, and attestations decay toward their review date — and the score decays with them, before they lapse.

Mock-survey pass rate

Your team's scored performance on the sandboxed drill walk feeds back in — practice that never touches the real evidence chain, but that the score can see.

The spine underneath the number

Hash-chained Decision log

Every mutation, attestation, and refusal is a row in an append-only, hash-chained ledger — each row referencing the last, completeness-checked on a schedule. A surveyor can verify the chain end-to-end.

Tracer engine

A Joint Commission patient tracer walks one patient backward through every standard touchpoint; a system tracer walks one process end-to-end. Each touchpoint maps to the standard chapter it tests, with a finding slot attached.

Sandboxed mock-survey mode

Simulates an accreditor's questions and emits a readiness scorecard with per-element gaps and actions. By construction it refuses to turn practice into real evidence, to pool across facilities, or to auto-fire remediation.

Citation & deficiency tooling

Structured CMS-2567 Plan-of-Correction composition and SAFER-matrix placement — the deficiency response drafted against the actual standard, not a blank template.

Evidence-loop refusals

A surface cannot publish a readiness claim without a verified citation chain. When the audit trail would auto-modify a surface's own citation, the evidence loop refuses — a machine-readable code, naming the policy.

Honest empty states

Boards render live evidence only where a facility has loaded it. Where nothing is seeded, Owl shows the empty state — it never fabricates a screenshot, a score, or a survey finding to fill the gap.

Screenshot of the Owl tracer engine — a patient traced backward through every standard touchpoint, each mapped to the standard chapter it tests
Tracer engine — one patient, every standard touchpoint
Screenshot of an Owl evidence-loop refusal — a machine-readable code naming the policy and the citation instead of a fabricated result
Evidence-loop refusal — cited, machine-readable, never a dead end
Screenshot of Owl Survey Mode — the in-pocket tracer walk with two-tap evidence pull and live findings capture
Survey Mode — the tracer walk, in your pocket
Screenshot of the Owl CMS-2567 deficiency and Plan-of-Correction composer, drafted against the cited standard
CMS-2567 Statement of Deficiencies & Plan-of-Correction composer
The Second Name — Charlie Nuthatch

Owl proves what happened.
Nuthatch is there while it happens.

We asked a hard question about our own platform: how does our AI governance stop a model from making an error in real time? We audited ourselves against four rungs — prevent, detect, prove, improve — and the honest answer was that our detection was world-class, our evidence spine was world-class, and our prevent rung was one permission deny. Everything else fired after the code already existed. Nuthatch is what we built because we did not like our own answer.

Nuthatch does not make bad code impossible. Nuthatch makes non-compliance immediate, attributed, and provable, and it makes bypass impossible to do silently.

For your engineers — the harness

Nuthatch hangs off an AI coding agent’s own tool-call boundary and evaluates each attempted write with deterministic rules before the bytes exist on disk. Twenty controls come in the starter pack — identifiers in a logger, a credential written to a file, a store read that swallows its failure into an empty list, a permission check that fails open, a raw outbound call off the reviewed transport, an unwitnessed bypass — and every one of them starts at “ask.” Not one can deny today. A control earns “deny” by demonstrating a low false-positive rate against real evidence, and the tool that reports which controls have earned it deliberately refuses to make the change: flipping a control from asking a human to blocking one has to have a human’s name on it.

For your survey binder — the packet

What the harness produces is read by someone who will never open the codebase: an attestation packet naming the controls that were in force and the signature state of the pack they came from, the receipt totals, whether the chain and its signed tree heads verify, and every break-glass event rendered individually with its named actor and its stated reason. It closes with a methodology section stating its own limits — including that observations are not blocks, and that this is evidence for conformance, never a compliance certification. The ledger is hash-chained and tamper-evident: an edited, deleted, or missing record is detectable rather than invisible. It is designed so that an independent auditor holding only a public verification key can confirm a record is genuine, without being shown anything else in the log — and a separate, dependency-free verifier does exactly that, offline. The trail stores no file contents; file payloads are recorded only as cryptographic fingerprints, and credential-shaped material is scrubbed from recorded command text before anything lands. Both the packet and the verifier are built; publication is roadmap, not shipped.

Three tiers of question, told apart on purpose

T0 is decidable from the tool call alone — a secret being written, an identifier key in a logger, a bypass flag. A refusal is legitimate here. T1 needs cached repository context on a strict time budget — does that data-model delegate exist, did this change raise a baseline, is the cited commit really an ancestor — and on timeout it degrades to asking, never to passing. T2 is not decidable in real time by anything, and Nuthatch does not pretend otherwise: it does not judge the content, it forces provenance, and it starts an obligation clock routed to a named human. dose-gate’s epistemic claim is “this constant came from the registry,” not “this constant is correct.”

What Nuthatch does not do

Bypass is possible and we say so: a bypass flag, a subshell, a script the agent then runs, an unproxied tool. Break-glass demands a written reason of at least forty characters, refuses to arm without a chained receipt of its own, and burns the grant before the bypass applies — under hard caps on how long a grant lives and how many times it is used. So a bypass is witnessed, not prevented. Files referenced without a tool call are invisible to it. Writes that arrive shaped as shell commands rather than edits are measured, not caught — and a control exists to measure that gap rather than hide it. The starter pack is unsigned today, and every packet says so. It supports one coding agent — Claude Code — today; other agent transports are roadmap. It certifies nothing and makes no one compliant — it produces evidence a human being still has to read.

Nuthatch receipts are designed to register as an Owl evidence source, so the governance file a hospital hands a surveyor covers both the algorithms that touch patients and the agents that touch the code. That join is built into the band grammar and the receipt format; wiring it into the Owl board is gated on a design partner. The developer page →

Why the Standard Matters

The stakes are measured in outcomes,
not only dollars.

A service line runs on its designation, and a community's access to specialty care runs on both. When one lapses, the public evidence on what a community loses is peer-reviewed and — kept honest — associational, not causal. We cite it the way we cite everything else.

1.9M
Neurons lost each minute a large-vessel ischemic stroke goes untreated
A modeled estimate for a typical large-vessel supratentorial stroke — illustrates treatment urgency, not a direct measurement of delay against disability (Saver, “Time Is Brain—Quantified,” Stroke 2006, PMID 16339467)
21%
Higher adjusted odds of in-hospital death for injured patients whose drive time rose after a nearby trauma center closed
Associated with, not caused by; adjusted OR 1.21 (95% CI 1.04–1.40), observational, California 1999–2009; relative odds, not absolute risk — the exposed group was small (5,122 of ~271,000) (Hsia et al., J Trauma Acute Care Surg 2014, PMID 24625549)
+0.67 pp
More preterm births the year after a remote rural county lost hospital-based obstetric services
Non-urban-adjacent rural counties; associated with, not caused by; 95% CI 0.02–1.33 — an interval that nears zero, stated honestly. Births in hospitals with no obstetric unit rose +3.06 pp (CI 2.66–3.46) (Kozhimannil et al., JAMA 2018, PMID 29522161)

How to read these

These are the honest stakes, honestly labeled. The stroke figure is a modeled estimate for a typical large-vessel stroke, not a measurement. The trauma and obstetric figures are observational associations, adjusted for the confounders their authors could measure, and drawn from datasets a decade old. They do not prove that any one closure caused any one death, and Charlie Owl never claims they do. They establish the stakes of the promise a designation makes — which is exactly why keeping it provable, continuously, is worth doing.

Registrar Amplification

Abstract once. Submit everywhere.

Registry abstraction still means a human hunting through charts — about 39 minutes per case before acuity, by peer-reviewed stopwatch study. The same hospitals abstract again for cardiology, surgery, stroke, and cancer registries. Charlie Owl proposes; a named registrar attests; the hash chain keeps the receipt.

Amplification, Never Replacement

The registry standards require human registrars — so amplification is the only compliant architecture, and it's ours by design. The registrar's judgment is the product; Charlie Owl removes the chart-hunting around it.

Your File Will Not Bounce

National registry submissions are rejected whole-file on a single critical flag. Charlie Owl validates locally, before submission, against the same rules — so the deadline week isn't spent discovering what the registry would have told you.

The Substrate

Governance is the architecture,
not a feature bolted on.

Hash-Chained Audit Trail

Every decision is recorded in an append-only ledger — each row hashed, chained to the previous, and completeness-checked by a scheduled job. A surveyor can verify the entire chain. No trust required.

Refusal-Grade Honesty

When Charlie Owl won't do something, it returns a machine-readable code naming the policy and the citation. It never guesses, never fabricates, never silently succeeds. Honesty isn't a feature — it's the foundation.

Governance Enforced at Boot

Before the server accepts a single request, it scans itself: hardcoded clinical thresholds, missing patient-data gates, silently swallowed errors. If a check fails, the server refuses to start. This is not a compliance PDF — it is the compiler.

Evidence-Cited Recommendations

Every automated recommendation carries its source: the published guideline, the review date, and an investigational flag when it isn't clinically validated. If the evidence base changes, the recommendation flags itself.

The Standard

Continuous survey readiness is the bar
this category should be judged by.

The readiness market judges itself by two things: surviving survey day, and a mock-survey quote every cycle. Charlie Owl's position is that always-ready is the only honest bar — because the patients, the community, and the revenue do not wait for survey week. This is a path to continuous survey readiness. That is the goal, and that is the standard.

The accreditor already agrees

The Joint Commission itself sells a “Continuous Survey Readiness Program” — one federal purchase order obligated $9.71M over roughly six years (USAspending.gov, VA, 2010–2016). Continuous readiness is not our marketing idea; it is the thing the incumbent already charges seven figures for. We think it should be the product, not the premium tier.

Don't rent your evidence file from the accreditor

The evidence a hospital uses to prove readiness — and to contest a finding, defend an Immediate Jeopardy, or switch accreditors — should not live inside the accreditor's own tool. The conflict of interest is the same one that keeps an EHR from auditing itself. Your receipts should be yours, portable across every certifier your portfolio spans.

Regulatory Tailwinds

The rules are changing in your favor.

DriverStatusWhat Charlie Owl Does
Joint Commission / CHAI RUAIH (responsible AI)June 2026AI governance evidence packs, inventory, human-in-the-loop receipts — and, for AI-authored code, tool-call-level receipts
ONC HTI-1 (algorithm transparency)2026Every algorithm shows its work; every refusal names the policy
Payer-algorithm litigation (Lokken, Kisting-Leung)In discovery, 2026Hash-chained receipts for what the model said and who decided
Scribe & AI-agent suits (Saucedo, Washington, Winters)2025–2026Algorithm inventory, intended-use and consent evidence, per-encounter trail
Trauma / stroke / cancer designation cyclesOngoingDesignation ledger, tracers, Immediate Jeopardy, plan-of-correction
Legacy trauma-registry end-of-life12–18 mo windowThe no-integration migration path: abstract once, submit everywhere
CMS CoP (hospital conditions)OngoingAccreditation readiness, evidence pack generation

Why No Incumbent Can Follow

The EHR vendor cannot audit itself. The AI vendor cannot grade its own homework. And the accreditor that sells continuous readiness cannot also be the neutral keeper of the file a hospital would use to contest its own finding. None of them is built on an enforced-governance substrate. The conflict of interest is structural — that's the moat. And when the record is subpoenaed rather than surveyed, a file kept by the party being questioned is worth less than one kept on a neutral, tamper-evident spine.

See Charlie Owl on your own compliance calendar.

We are seeking design partners — any hospital standing up AI governance, and community hospitals defending trauma, stroke, or specialty designations. The first receipts come from data you already have.

info@nightingaleos.com · kristen@nightingaleos.com