The exam paper · compiled from live diligence

Ten questions, in writing. How many can you answer today?

Health-system diligence for AI documentation vendors has converged on ten questions, asked in writing, scored by committees that have heard “we spot-check” from everyone. Most can be answered tonight. The four highlighted below demand evidence that, at most vendors, does not exist yet: a finding that shows its work, an edge-case battery with objective ground truth, traceable validation records, a grounded coding gate. In the cycles we have seen, these four are where prose answers die, because “trust us” is not an artifact.

The exam paper · verbatim, de-identified from live diligence

  1. 1How do you ensure AI output never contradicts approved clinical guidelines, and how are guideline updates governed?
  2. 2How is medico-legal responsibility separated between physician and AI?
  3. 3When your system flags missing documentation or diagnoses, can we see why? Source logic, evidence, examples?
  4. 4Will the system monitor or score physician documentation behavior, and who has access to those metrics?
  5. 5Can that data be used for credentialing, review, or discipline?
  6. 6How does this comply with our EHR constraints and security policies?
  7. 7How are your AI agents tested for accuracy, safety, bias, and explainability? What scenarios validate edge cases and failure conditions? Is validation evidence captured and traceable?
  8. 8Confirm full compliance with our QA certification SOP, every gate, before go-live.
  9. 9Detail your AI-specific testing: scope, methodology, validation, governance.
  10. 10How do you ensure coders review AI-generated codes before any downstream use?

The plain ones belong to programs you already run: clinical governance, medico-legal, monitoring access, EHR security, QA certification. Lithrim supplies the evidence artifact where it plugs in. The four highlighted are where a paragraph is not enough.

The four answers, as artifacts.

Question 3

A finding that shows its evidence

From transcript to verdict, every link is cited: the source line, the note span, the deterministic check, the reviewer votes, the verdict. A reviewer clicks from any verdict back to the exact lines that produced it, or reruns the grade and watches the same chain re-derive. Omissions get the reverse: the chain cites the transcript line that should have appeared and didn't.

Question 7

An edge-case battery with objective ground truth

Twelve synthetic visits, each carrying its own provable ground truth, so pass and fail are objective. Clean notes clear; the failure classes a validation reviewer asks about (fabricated exam, source contradiction, misattributed speaker, high-severity omissions) are each caught with the quote that proves it. Every run is signed and traceable.

Question 9

Traceable validation records

Criteria, judge configurations, and every run are hashed, versioned, and logged. Your validation evidence pins the exact criteria version it was graded under, and quarterly updates ship as regrade diffs, so improving the checks never invalidates yesterday's audit trail.

Question 10

A grounded coding gate

AI-generated codes are checked against the source and a live terminology standard before anything moves downstream. The check knows direction: generalizing a diagnosis is safe, sharpening it beyond what the record supports is an upcode, and the gate names it with the terminology as evidence.

Walk into your next diligence cycle with the four answers written.

A one-week pilot leaves you with artifact-linked responses to the four: the evidence chain, the signed edge-case battery, the traceable validation records, the grounded coding gate, plus a map of the other six to the programs that own them. Your counsel and security team see exactly which artifact backs each answer.

Start the pilot conversation