SOURCE OF TRUTH · OPEN PILOT

YOUR DOCUMENTS
vs. WHAT ACTUALLY
GOVERNS THEM

Somewhere in your manuals a deadline is wrong, a threshold is stale, or an obligation was never written down at all. SOT reads the document and the rule together and tells you exactly where they disagree — quoting both sides, every time.

Reads both sides — the rules that govern you and the documents that run your operation
Reports the disagreements — with the clause on each side quoted, not summarized
Names what's uncovered — obligations your procedures never mention
Writes the correction — replacement language you can paste, cited to its source
[01] How it works

Three inputs. One report
you can hand to an auditor.

[01]

Bring your document

A manual, SOP, policy set, contract, or requirements spreadsheet. PDF, Word, Excel, CSV, HTML, Markdown, or plain text — parsed into sections by its own headings or worksheets.

[02]

Pick the governing source

The corpus that has authority over it: a statute set, a standard, a master agreement, an internal policy of record. Adding one is a data change, not an engineering project.

[03]

Read the findings

Contradictions with both quotes and a suggested fix, plus the uncovered-obligation list. Verified findings are marked. Share it, print it, or hand it to the auditor.

What a finding looks like
sot · findings
$ sot check --document claims-manual.pdf --against ca-workers-comp 3 findings · 1 gap · 27 sections read CONFLICT §2.3 Utilization Review Timing document "20 business days, extensions at reviewer discretion" L412 source Cal. Lab. Code § 4610(i) — 5 working days, 14 max L18 fix "…decisions issued within 5 working days of receipt, extended only where information is reasonably required, never beyond 14 calendar days." review second model — upheld OK §2.1 Claim Form Delivery § 5401(a) L88 OK §3.1 Temporary Disability Rate § 4650(a) L233 GAP §§ 4060–4062.2 — medical dispute panel process nothing in this document addresses it
[01]

Contradiction finding

Section by section, SOT compares what your document says against what the governing source requires. A mismatched deadline, threshold, rate, or weakened standard comes back as a finding — with both quotes side by side.

[02]

Coverage gaps

The inverse question, which most reviews never ask: which obligations does nothing in your document address? Silent omissions surface as their own list, cited to the requirement they miss.

[03]

Cited, never vibes

Every finding carries the document line and the source clause it turns on. If a claim can't be pinned to text on both sides, it doesn't ship as a finding.

[04]

Second-model review

Findings are re-examined by a stronger model whose job is to knock them down. Only what survives the challenge reaches you — and anything left unverified is labeled, never quietly upgraded.

[05]

Measured, not asserted

Each domain ships with a labeled test set — real requirements, planted errors, known answers. Precision and recall are numbers we publish per deployment, including the misses.

[06]

Nothing retained

Files are parsed in memory and deleted when the report renders. No account, no storage, no training on your documents. Run it in your own environment when policy requires.

[02] Try it

Upload something and see.

Any policy document or spreadsheet. It runs against a live workers' compensation corpus — the vertical we've benchmarked — so you can judge the output on rules you can look up yourself. Your file is deleted when the report renders.

Drop a file here — or click to browse PDF · Word (.docx) · Excel (.xlsx) · CSV / TSV · HTML · Markdown · text · JSON
Up to 15 MB. Sections come from headings, numbered clauses, or worksheets.
Uploading…

Findings

[03] Proof

Numbers, including
the ones that aren't perfect.

A domain is a corpus plus a labeled test set. Workers' compensation is the one that's live: 51 sections with known answers, seeded with contradictions — and with traps that look wrong but are correct, which must not be flagged.

[CALIFORNIA]

100% / 100%

Precision / recall across 27 labeled sections. 13 contradictions found, every one upheld under second-model review; both coverage gaps caught.

[TEXAS]

100% / 91%

Nothing wrong was reported, but one contradiction was missed — a benefit-eligibility threshold. Published as measured, not rounded up.

[TEST SET]

51 sections

Includes paraphrased contradictions with no numbers to match on, and stricter-than-required text that must come back clean.

Open the worked example Read a full report
Same engine, other domains

Insurance & claims [LIVE]

Claims-handling manuals against state statute — deadlines, caps, benefit calculations.

Banking & lending

Credit, AML, and servicing policy against supervisory guidance and internal policy of record.

Healthcare operations

Clinical SOPs and compliance plans against federal conditions and accreditation standards.

Contracts & obligations

Statements of work and vendor terms against the master agreement they inherit from.

Safety & quality

Plant procedures against OSHA, ISO, and corporate standards — across sites that drifted apart.

Public programs

Program manuals against enabling statute, funding conditions, and eligibility rules.

[04] FAQ

Fair questions.

[01]What does SOT actually compare?

Two things you already have: a document that governs (statute, standard, master agreement, policy of record) and a document that operates (manual, SOP, procedure, spreadsheet of controls). Most review tools read one and summarize it. SOT reads both and reports the delta — contradictions, omissions, and the corrected wording.

[02]Is this a search or retrieval product?

No. Retrieval shortlists passages by similarity and hopes the right one made the cut; when it doesn't, the miss is invisible. SOT evaluates each section of your document against the entire requirement set, so a missed contradiction is a defect the test set catches rather than a silent hole. Retrieval only enters when a corpus is too large to hold at once, and never decides what's true.

[03]How do you keep false positives down?

Two ways. Findings are re-argued by a stronger model instructed to refute them, and anything that fails that challenge is dropped. And the whole pipeline is scored against labeled test sets built to punish over-flagging — including sections written to look wrong but that are actually stricter than required, which must not be flagged.

[04]What can I upload?

PDF, Word (.docx), Excel (.xlsx), CSV or TSV, HTML, Markdown, plain text, and JSON exports. Spreadsheets are read worksheet by worksheet, so a controls matrix or requirements register works as well as prose. Documents need headings, worksheets, or numbered sections so there are boundaries to reason about.

[05]What happens to my file?

It's parsed in memory, analyzed, and deleted when the report renders — there is no database and no storage bucket. Nothing is used for training. For work that can't leave your perimeter, the same system runs inside your environment.

[06]How would I know it's right?

Don't take the claim; read the receipts. Every finding quotes both sides so you can check it in seconds, and each domain publishes precision and recall against a labeled corpus — including where it fell short. The workers' compensation numbers, and the misses, are on the demo.

[07]How do we start?

Upload a document below and read what comes back. For a real engagement, a pilot takes one policy set and one authority corpus and returns a scored report in the first session — no integration, no schema work, no data warehouse.

Start with the document
you'd least like to defend.

Upload it above, or bring us a policy set and the source that governs it — you'll have a scored report in the first session.