Bring your document
A manual, SOP, policy set, contract, or requirements spreadsheet. PDF, Word, Excel, CSV, HTML, Markdown, or plain text — parsed into sections by its own headings or worksheets.
Somewhere in your manuals a deadline is wrong, a threshold is stale, or an obligation was never written down at all. SOT reads the document and the rule together and tells you exactly where they disagree — quoting both sides, every time.
A manual, SOP, policy set, contract, or requirements spreadsheet. PDF, Word, Excel, CSV, HTML, Markdown, or plain text — parsed into sections by its own headings or worksheets.
The corpus that has authority over it: a statute set, a standard, a master agreement, an internal policy of record. Adding one is a data change, not an engineering project.
Contradictions with both quotes and a suggested fix, plus the uncovered-obligation list. Verified findings are marked. Share it, print it, or hand it to the auditor.
Section by section, SOT compares what your document says against what the governing source requires. A mismatched deadline, threshold, rate, or weakened standard comes back as a finding — with both quotes side by side.
The inverse question, which most reviews never ask: which obligations does nothing in your document address? Silent omissions surface as their own list, cited to the requirement they miss.
Every finding carries the document line and the source clause it turns on. If a claim can't be pinned to text on both sides, it doesn't ship as a finding.
Findings are re-examined by a stronger model whose job is to knock them down. Only what survives the challenge reaches you — and anything left unverified is labeled, never quietly upgraded.
Each domain ships with a labeled test set — real requirements, planted errors, known answers. Precision and recall are numbers we publish per deployment, including the misses.
Files are parsed in memory and deleted when the report renders. No account, no storage, no training on your documents. Run it in your own environment when policy requires.
Any policy document or spreadsheet. It runs against a live workers' compensation corpus — the vertical we've benchmarked — so you can judge the output on rules you can look up yourself. Your file is deleted when the report renders.
A domain is a corpus plus a labeled test set. Workers' compensation is the one that's live: 51 sections with known answers, seeded with contradictions — and with traps that look wrong but are correct, which must not be flagged.
Precision / recall across 27 labeled sections. 13 contradictions found, every one upheld under second-model review; both coverage gaps caught.
Nothing wrong was reported, but one contradiction was missed — a benefit-eligibility threshold. Published as measured, not rounded up.
Includes paraphrased contradictions with no numbers to match on, and stricter-than-required text that must come back clean.
Claims-handling manuals against state statute — deadlines, caps, benefit calculations.
Credit, AML, and servicing policy against supervisory guidance and internal policy of record.
Clinical SOPs and compliance plans against federal conditions and accreditation standards.
Statements of work and vendor terms against the master agreement they inherit from.
Plant procedures against OSHA, ISO, and corporate standards — across sites that drifted apart.
Program manuals against enabling statute, funding conditions, and eligibility rules.
Two things you already have: a document that governs (statute, standard, master agreement, policy of record) and a document that operates (manual, SOP, procedure, spreadsheet of controls). Most review tools read one and summarize it. SOT reads both and reports the delta — contradictions, omissions, and the corrected wording.
No. Retrieval shortlists passages by similarity and hopes the right one made the cut; when it doesn't, the miss is invisible. SOT evaluates each section of your document against the entire requirement set, so a missed contradiction is a defect the test set catches rather than a silent hole. Retrieval only enters when a corpus is too large to hold at once, and never decides what's true.
Two ways. Findings are re-argued by a stronger model instructed to refute them, and anything that fails that challenge is dropped. And the whole pipeline is scored against labeled test sets built to punish over-flagging — including sections written to look wrong but that are actually stricter than required, which must not be flagged.
PDF, Word (.docx), Excel (.xlsx), CSV or TSV, HTML, Markdown, plain text, and JSON exports. Spreadsheets are read worksheet by worksheet, so a controls matrix or requirements register works as well as prose. Documents need headings, worksheets, or numbered sections so there are boundaries to reason about.
It's parsed in memory, analyzed, and deleted when the report renders — there is no database and no storage bucket. Nothing is used for training. For work that can't leave your perimeter, the same system runs inside your environment.
Don't take the claim; read the receipts. Every finding quotes both sides so you can check it in seconds, and each domain publishes precision and recall against a labeled corpus — including where it fell short. The workers' compensation numbers, and the misses, are on the demo.
Upload a document below and read what comes back. For a real engagement, a pilot takes one policy set and one authority corpus and returns a scored report in the first session — no integration, no schema work, no data warehouse.
Upload it above, or bring us a policy set and the source that governs it — you'll have a scored report in the first session.