# Orinyx > Orinyx is an independent verification platform for clinical AI. Hospitals use it to test AI products before deployment and to keep verifying them once they are live. Orinyx has no clinical AI product of its own, because nothing can credibly audit itself. Orinyx is one platform with two phases. Phase 1, Deployment Readiness: hospitals test clinical AI products in the Orinyx sandbox before deployment. Every assertion the AI makes is verified against authoritative sources, and AI-generated clinical summaries are scored for accuracy against the source encounter. The output is a citable verification report a governance committee can act on. Phase 1 runs on synthetic data, with no PHI, no EHR connection, and no security review required to begin. Phase 2, Runtime Verification: once a product goes live, the same verification runs continuously at runtime with a full audit trail, firing upstream of the clinician. The same verification that grades an AI product before deployment keeps running once it is live. The test is the monitor. Orinyx is tool-agnostic and clinician-facing only in that it surfaces flagged recommendations for review. It is a software platform, not a diagnostic tool. ## Pages - [Home](/): Overview of the Orinyx independent verification platform, the two phases, security, and FAQs. - [Benchmark](/benchmark): Phase 0 entry point. Independent benchmark of a clinical AI tool's outputs against authoritative sources. - [Evaluating ambient AI vendors](/evaluating-ambient-ai): A buyer's guide for hospital leadership: what to ask an ambient AI scribe vendor, what a good answer sounds like, and what evidence to demand. - [Blog](/blog): Articles on clinical AI safety, verification, governance, and oversight. - [AI Scribe Patient Consent: What a Defensible Consent Flow Requires](/blog/ambient-scribe-patient-consent): What consent for an ambient AI scribe actually requires: recording, transmission, retention, and a consent record the documentation tool did not write. - [Automation Bias in Clinical AI: What Happens When Humans Stop Checking](/blog/clinical-ai-automation-bias): What automation bias is, the evidence that human review atrophies when AI checks AI, and how hospitals can keep verification real. Sourced for CMIOs. - [Benchmarking Clinical AI: What We Measure, What We Withhold, and What Changes for Agents](/blog/benchmarking-agentic-clinical-ai): How Orinyx benchmarks clinical AI: catch, false-flag, and abstain rates, hash-pinned case sets, rotation policy, and the measures planned for agentic systems. - [When Medicare starts paying for outcomes, the evidence becomes the product](/blog/medicare-sams-outcome-evidence): CMS has proposed a named Medicare payment category for clinical software, and signaled that future payment rates will track impact on patient outcomes. That turns a vendor's evidence package into a financial instrument, and raises a question procurement has not had to ask before: who measures the outcome that determines the payment? - [When an AI Has a Stake in Its Own Recommendation](/blog/ai-incentive-bias): An AI can be correct and conflicted at once. Why Illinois now requires conflict-free AI auditors, and what clinical AI procurement should ask for. - [Why an AI System Can't Audit Its Own Output](/blog/ai-self-audit-problem): An AI system can't verify its own output: generator and checker share the same blind spots. Why clinical AI needs independent verification. - [Why you can't make clinical AI accurate enough to skip governance](/blog/clinical-ai-accuracy-governance): A higher accuracy score doesn't close the clinical AI safety gap. Here's what governance does that model training cannot, and why the FDA's 2026 CDS deregulation makes this urgent. - [Transcript-faithful is not the same as clinically correct](/blog/transcript-faithful-not-clinically-correct): A transcript-faithful clinical note can still be clinically wrong. Why that gap matters for AI governance, from Orinyx, an independent safety layer for clinical AI. ## Optional - [Request a founding benchmark](/#demo): Founding design-partner cohort with a complimentary diagnostic benchmark for the founding cohort.