How to evaluate ambient AI vendors
What to ask, what a good answer sounds like, and what evidence is worth the weight you put on it.
Every published guide to buying ambient AI was written by someone who sells ambient AI, or sells access to the people who do.
Orinyx does not sell an ambient scribe. We verify clinical AI output against authoritative sources, which means we have no scribe product to defend in this guide and no marketplace placement to sell you. That is the only reason this page can be written the way it is. Take the reasoning on its own terms; nothing here requires you to become an Orinyx customer, and the questions below work just as well without us.
This guide is for the people who actually do the evaluating: CMIOs, chief quality and chief physician officers, clinical informatics teams, and the compliance, privacy, and health information management colleagues who get pulled in late and should be pulled in early.
What ambient AI does, and where the risk actually sits
You know the category. A microphone captures the encounter, a model produces a transcript, and a second stage generates a structured clinical note from that transcript. The documentation relief is real and it is the main reason adoption has moved faster than shared evaluation standards, a point the Coalition for Health AI makes explicitly in its Ambient AI work group materials.
The risk does not sit where most procurement checklists look for it. Transcription accuracy is the easiest part of the stack to measure and the part vendors have largely solved to a reasonable standard. The risk sits in the second stage, where a model decides what the conversation meant, what belongs in the assessment, what to carry forward, and what to leave out. That is a judgment step, and judgment steps are what a purchase decision has to be able to interrogate.
One structural note worth having in front of you: in the United States, ambient documentation tools are not currently regulated as medical devices by the FDA, a point CHAI states plainly in its Ambient AI Testing and Evaluation Framework. Other jurisdictions draw the line differently. NHS England guidance, developed with the MHRA, treats ambient scribing products that go beyond plain transcription into generative summarisation as medical devices, and its AVT Supplier Registry requires suppliers to hold at least MHRA Class 1 registration. Whatever you conclude about that gap, the practical consequence for a US health system is direct: there is no regulatory floor doing this work for you. The evaluation is yours to run.
Transcript-faithful is not the same as clinically correct
This is the distinction that most buyer's guides skip, and it is the one that changes what you ask for.
When a vendor describes a guardrail that checks the note for hallucinations or confabulations, ask what the check compares the note against. In most ambient systems, the answer is the transcript. Every claim in the note is scored on whether the conversation supports it. That is a legitimate and useful measurement. It catches misheard drug names, dropped medications, misattributed statements, and invented findings, which together are a large share of the errors these systems make.
But notice the boundary. A transcript-fidelity check cannot catch an error that originates in the transcript. If a clinician states a dose incorrectly mid-shift, the microphone hears it correctly, the model records it faithfully, and the guardrail scores it as fully supported. The note is transcript-faithful. It is also clinically wrong. The same holds for a reversed guideline recited from memory, a drug interaction that goes unmentioned, or a monitoring parameter that belongs to a different drug class. The reference material has no opinion on any of it, because the reference material is a recording of a room.
Catching that second class of error requires a different reference entirely: external clinical authority. Drug labeling. Current guidelines. The patient's own chart. Nothing about the conversation can supply it.
We wrote this argument out at length, with a worked example, in transcript-faithful is not the same as clinically correct. The short version for your purposes: a vendor's safety layer answers the question the vendor scoped. Your job in evaluation is to find out precisely what that question was, so you can decide, deliberately, who is accountable for everything outside it.
This is not an argument against ambient AI. It is an argument for knowing where the guarantee stops.
Acceptance rate is not accuracy rate
You will be shown an acceptance rate, or an edit rate, or some variant: the percentage of notes clinicians sign with minimal or no changes. It is usually high, and it is usually presented as evidence of quality.
It is evidence of something. It is not evidence of correctness, and the reason is worth stating carefully because it is easy to overstate in the other direction.
Acceptance measures whether a clinician found the note acceptable enough to sign, under real conditions: end of a clinic day, a queue of notes, cognitive load already spent on the patient. A note that reads fluently, uses the right structure, and contains no obvious defect will be accepted. A subtle clinical error, by definition, does not present as an obvious defect. Fluency and correctness are produced by the same generation process and are not correlated in the way a reader intuitively assumes them to be.
There is a second problem, which is that acceptance is measured on the notes the system produced, judged by the person the system was built to please. It is a satisfaction signal wearing a quality signal's clothes. High acceptance tells you the workflow fits, clinicians are not fighting the tool, and adoption is likely to hold. Those are real and valuable findings. They are answers to a different question than "is this note correct."
The practical test: ask the vendor whether any note ever accepted by a clinician was later found, on structured review, to contain a clinical error. If the answer is no, the vendor is not looking. If the answer is yes, and they can characterise the error types, you are talking to someone with a real quality program.
CHAI's Ambient AI Implementation Playbook makes a related point from the governance side, recommending that organisations measure success through clinician wellness and patient experience rather than adoption alone.
The questions to ask, and what a good answer sounds like
Listicles give you the questions. The questions are the easy part. What separates a useful evaluation from a performative one is knowing what a strong answer sounds like, so you can tell the difference in the room.
1. What does your safety check compare the note against?
A good answer names the reference explicitly and without hedging. "The transcript" is a good answer. It is honest, it is accurate for most systems on the market, and a vendor who says it plainly is a vendor you can trust on harder questions. A better answer names a second reference for a defined subset of claims: medication assertions checked against labeling, for instance, with a clear statement of what is and is not in scope.
A weak answer describes the mechanism without naming the reference. "Our proprietary guardrail model scores every claim for support" tells you nothing until you know what the support is measured against. Push once. If the answer stays at the mechanism level after a direct question, that is itself information.
2. Who defined the benchmark, curated the cases, and scored the results?
Three separate roles. Ask about them separately, because a vendor can be honest about one and quiet about the others.
A good answer discloses which of the three were internal, publishes or shares a methods description, and volunteers what the benchmark was not designed to test. The strongest version offers to be evaluated on cases the vendor did not select.
A weak answer is a single headline number with no methods behind it, or the phrase "independently validated" without naming who validated it, on what cases, under what protocol, and whether that party had a commercial relationship with the vendor at the time.
3. What happens when the clinician says something wrong?
A good answer is direct: the system is designed to be faithful to the encounter, so a spoken error will be documented as spoken. That is the correct answer and a well-run vendor gives it without flinching. If they then describe anything that sits outside the transcript, ask which claim types it covers and what it checks against.
A weak answer claims the product catches clinical errors without naming the external source it checks against. That claim needs a reference behind it, and if there is one, the vendor will be able to name it in a sentence.
4. How will we know when the model changes?
Silent model updates are a named problem, not a hypothetical. CHAI's Implementation Playbook lists systems that drift as vendors quietly update them among the recurring issues its work group surfaced across five lifecycle stages.
A good answer includes version notification with lead time, a changelog you can actually read, a contractual right to re-test before a material update reaches production, and retained evaluation data so you can compare before and after on the same cases.
A weak answer is "we continuously improve the model." That is a description of a process with no observability attached.
5. What happens to our audio, our transcripts, and our notes?
A good answer states a default retention period, in the contract, in days. It states deletion behaviour on request and on termination. It states, unambiguously, whether your data is used to train or improve models, with an opt-out that is contractual rather than a setting. It names subprocessors and confirms BAA coverage extends to each of them.
A weak answer is "we are HIPAA compliant." The Joint Commission and CHAI guidance on the Responsible Use of AI in Healthcare recommends going considerably further: updating business associate and data use agreements to define how data may be used, minimise what is shared, prohibit re-identification, and reserve audit rights. That last item is the one most often missing and the cheapest to add before signature.
6. What was this evaluated on, and does it look like our patients?
A good answer describes the specialties, encounter types, patient populations, languages, and accents in the evaluation set, and says candidly where coverage is thin. It supports you running your own evaluation on your own population rather than treating that request as a lack of trust. The Joint Commission and CHAI guidance recommends exactly this: ask how tools were trained, validated, and tested, and evaluate them against your own patient populations.
A weak answer treats performance as a single portable number. It is not. CHAI's own framework flags that its metrics were derived primarily from narrative clinical note generation and that some will apply differently, or not at all, to flowsheets, orders, and structured data capture.
What evidence to demand, and what evidence is weak
Evidence worth weighting:
- A methods description you could hand to a biostatistician without embarrassment: denominator, case selection, adjudication procedure, inter-rater agreement.
- Results broken out by error type and clinical severity, not one aggregate accuracy figure. An aggregate number hides exactly the errors you care most about, because the severe ones are rare.
- Performance on a held-out set the vendor did not curate.
- Clinician adjudication against a defined error taxonomy, rather than a binary accurate/inaccurate judgment.
- Results generated on your population, or a documented commitment to produce them during the pilot.
- Public transparency artifacts. CHAI maintains a public registry where solutions can be listed with model card detail; presence there is not a guarantee of quality, but it is a disclosure posture worth noting.
Evidence that carries less weight than it appears to:
- An accuracy percentage with no denominator and no methods section.
- Acceptance rates, adoption rates, and edit rates, for the reasons above.
- Time-saved surveys and burnout scores. These measure something real and important, and they measure workflow fit, not note correctness. Treat them as evidence of the benefit, not evidence of the safety.
- Testimonials, NPS, and reference calls arranged by the vendor.
- A comparison against a single hand-picked baseline. The comparator was chosen by the party being measured.
- "Validated by [named institution]" where the institution was a design partner and the protocol is not published.
None of this means a vendor producing weak evidence is a weak vendor. Most of the market is in the same position because the standards are new. It means you should not treat the evidence as stronger than it is when you write the risk section of your governance memo.
Who verifies the output, and why the vendor grading itself is not verification
There are three distinct roles in any clinical AI deployment: the party that generates the output, the party that grades it, and the party that is accountable when it is wrong. In most ambient AI arrangements today, the first two are the same organisation and the third is you.
That is the structural problem, and it is not a claim about anyone's integrity. A self-administered benchmark measures what its author chose to measure, under conditions its author chose to set. It can be entirely accurate and still tell you nothing about performance on the claims the author did not think to test, or against the standard the author was not trying to meet. Nothing can credibly audit itself, and the failure is architectural rather than moral.
Clinician review is the layer most systems rely on to close this, and it is genuine. It is also attestation under time pressure rather than measurement, it is not sampled, and it produces no record you can hand to a board. Documentation error long predates AI: patient-reported error rates in ambulatory notes and systematic reviews of documentation deficiency both point to a baseline problem that human review has never fully caught.
There is a group inside your organisation that already sees note-level error patterns at volume, and is usually not in the selection meeting: clinical documentation integrity, health information management, and coding. They read thousands of notes after the fact, they detect systematic drift earlier than almost anyone, and they will be the first to notice when an ambient system starts producing a consistent structural defect. Putting them on the evaluation committee costs nothing and materially improves what you catch. The same is true of clinical informatics, who will own the thing operationally long after the procurement is closed.
One thing a vendor can do to close the independence gap is submit to a benchmark run by a party that does not sell the tool being measured. That is what the Orinyx benchmark is: verification of clinical AI output against authoritative sources, in a pre-deployment sandbox, producing a report a governance committee can read. Phase 1 runs on synthetic data, so there is no PHI and no EHR connection involved. To be precise about what it is not: it is not a regulatory clearance, not a safety certification, and not a substitute for your own evaluation. It is a second reading by someone with no product in the result.
What to write into the contract and the pilot
Most of what protects you is decided before signature and cannot be retrofitted.
In the contract:
- A defined note-quality standard, not just an uptime SLA. Uptime is the easy commitment and the one least connected to your risk.
- Notification of material model changes with lead time, and the right to re-test before those changes reach production.
- Data retention defaults stated in days, deletion behaviour on request and on termination, and an explicit contractual position on training use.
- A named subprocessor list, with BAA coverage confirmed for each, and audit rights reserved.
- Participation in your safety event reporting program as a contractual obligation rather than a courtesy.
In the pilot design:
- Pre-specify what you will measure and what result would make you decline, before the pilot begins. A pilot without a defined failure condition is a rollout with extra steps.
- Include a held-out set the vendor did not curate, and include the specialties and encounter types where you expect performance to be worst rather than best.
- Score by error type and severity, with clinician adjudication.
- Treat ambient documentation as recording governed by state law, and build a workflow for a patient withdrawing consent mid-visit.
- Fund ambient AI as shared infrastructure rather than a per-provider productivity charge, which matters most for part-time clinicians and trainees who are otherwise excluded by seat math.
- Plan monitoring past go-live. Pre-deployment controls do not reliably predict real-world behaviour, so build in shadow-run and post-deployment monitoring.
The evaluation in one page
- Establish what the vendor's safety check compares the note against. Usually the transcript. Get it said out loud.
- Separate transcript fidelity from clinical correctness in your own risk documentation, and decide explicitly who owns the gap.
- Discount acceptance rate as a quality signal. Keep it as a workflow signal.
- Ask who defined, curated, and scored every benchmark you are shown.
- Demand results broken out by error type and severity, on cases the vendor did not pick.
- Put clinical documentation integrity, HIM, and clinical informatics on the committee, not just in the rollout.
- Write model-change notification, retention defaults, subprocessor BAA coverage, and audit rights into the contract.
- Define the pilot's failure condition before the pilot starts.
- Plan for monitoring after go-live, because pre-deployment performance does not predict production performance.
If you want the second reading on your own evaluation, that is the work we do. If you would rather just take the checklist and run it yourself, that was the point of writing it down.
Further reading
- CHAI Ambient AI Work Group — vendor-agnostic consensus guidance across procurement, pre-deployment, piloting, deployment, and monitoring.
- Joint Commission and CHAI, Guidance on the Responsible Use of AI in Healthcare — governance structure, BAA and data-use agreement terms, and evaluating vendor tools against your own patient populations.
- NHS England, guidance on AI-enabled ambient scribing products — the clearest published statement of where a regulator draws the line between transcription and summarisation.
- Transcript-faithful is not the same as clinically correct — the long-form version of the core distinction on this page.
- Bell SK, et al. Frequency and Types of Patient-Reported Errors in Electronic Health Record Ambulatory Care Notes. JAMA Network Open (2020)
- Shahbodaghi A, et al. Documentation Errors and Deficiencies in Medical Records: A Systematic Review. Journal of Health Management (2024)