Posted on
May 11, 2026
Clinical Reasoning vs. Standard Scribing: Owning the Accuracy Gap in AI Documentation
Clinical Reasoning vs. Standard Scribing: Owning the Accuracy Gap
The Clinical Library Playbook for CMIOs Who Refuse to Lose Revenue to Narrative-Only Documentation
TL;DR — Why This Matters to Your Organization
Most AI scribes function as sophisticated text summarizers—they listen, compress, and paste prose into notes. But insurance-billed care doesn't run on prose. It runs on discrete, coded data: LVEF stored as a FHIR R4 Observation (LOINC 33878-0), NYHA class as a coded value, PHQ-9 totals mapped to LOINC 44249-1. When these values live only in narrative, 96127 billing fails, echo authorizations get denied for "insufficient medical necessity," and eCQM reporting collapses. Scribing.io's clinical reasoning engine closes this accuracy gap by extracting, scoring, and writing structured clinical data in real time—turning ambient conversation into defensible, billable, auditable documentation. This playbook details exactly how, why the AMA's transparency framework validates the approach, and what CMIOs should demand from any AI documentation vendor in 2026.
Table of Contents
Beyond Summarization — Why "Clinical Reasoning" Is the New Standard
Scribing.io Clinical Logic — A Before-and-After Case in Cardiology-Leaning IM
The Information Gain Gap — What Transparency Frameworks Miss About Discrete Data
Technical Reference: ICD-10 Documentation Standards
CMIO Vendor Evaluation Matrix: 11 Questions That Separate Reasoning Engines from Summarizers
Free Discrete Data Gap Report — Your Next Step
Beyond Summarization — Why "Clinical Reasoning" Is the New Standard for AI Documentation
The AMA's June 2026 Annual Meeting resolution correctly identifies that AI tools in healthcare must be transparent, evidence-based, and physician-led. What the resolution does not address—and what represents the most consequential gap in current industry guidance—is the distinction between text summarization and clinical reasoning in AI-generated documentation.
This distinction is not semantic. It is the difference between a note that reads well and a note that works: one that supports billing, survives audit, feeds quality measures, and automates prior authorization. Scribing.io was built on the thesis that this gap—between narrative plausibility and computational utility—is the single largest source of preventable revenue loss in ambulatory documentation today.
Text summarization takes a physician-patient conversation and produces compressed narrative. The output is prose. It may mention that "the patient's ejection fraction was noted to be 35%" or that "a PHQ-9 was administered with a score of 14." These are linguistically accurate statements. They are also computationally useless—invisible to billing engines, unreachable by clinical decision support, and indefensible in a payer audit that requires discrete data elements. A JAMA study on clinical documentation burden found that the downstream cost of rework from inadequate structured data capture far exceeds the time "saved" by narrative-only tools.
Clinical reasoning, as implemented in Scribing.io's engine, performs a fundamentally different operation:
Identifies clinical markers in real-time audio—LVEF values, NYHA functional classifications, PHQ-9 item-level responses—using specialty-trained NLP models. These are not keyword matchers. They parse clinical semantics: "his squeeze is down to about 30" maps to LVEF 30%, not a grip strength measurement.
Validates extracted data against clinical logic. A PHQ-9 total of 14 from only 7 captured items triggers a completeness prompt, not silent acceptance. An LVEF of 35% is classified as HFrEF, not HFpEF, based on the 2022 AHA/ACC/HFSA Heart Failure Guidelines threshold of ≤40%.
Structures the data as coded, discrete elements—FHIR R4 Observations, QuestionnaireResponses, and Conditions with appropriate LOINC, ICD-10, and SNOMED CT codes.
Writes those elements into the EHR using FHIR APIs where available and HL7 v2 ORU messages or vendor-specific adapters (e.g., Epic's Flowsheet Write API, Oracle Health's discrete data bridge) where FHIR write access is limited.
Surfaces billing opportunities—96127 for standardized instrument scoring, correct ICD-10 linkage, medical necessity language—before the encounter closes.
Every extracted value includes provenance metadata: the timestamp in the conversation where the value was spoken, the clinical logic applied, and the coding standard used. This directly satisfies the AMA's call for "explainability" in clinical AI—not as a whitepaper abstraction, but as an audit trail attached to every discrete data point.
This architecture matters across specialties. The same discrete-data gap that causes echo denials in cardiology causes missed developmental screening codes in Pediatrics and lost psychotherapy add-on billing in Psychiatry. The pathology is identical: narrative that humans can read but systems cannot consume.
For CMIOs evaluating AI scribe vendors: if a system cannot tell you the LOINC code it would assign to a captured value, it is a summarizer. If it can—and can write that code to a flowsheet—it has clinical reasoning. The revenue, compliance, and quality implications of this distinction are detailed throughout this playbook.
Scribing.io Clinical Logic — A Before-and-After Case in Cardiology-Leaning Internal Medicine
The following scenario is constructed from operational patterns observed across cardiology-leaning internal medicine groups. It illustrates the measurable impact of replacing narrative-only documentation with Scribing.io's clinical reasoning pipeline.
Before: Standard Scribe Summarizer
A six-provider IM group with a significant heart failure panel uses a market-leading ambient AI scribe. The tool listens to encounters and generates well-formatted SOAP notes. During a typical HF follow-up, the physician discusses:
A recent echocardiogram showing LVEF of 30%
The patient's current NYHA Class III functional status
A PHQ-9 screen with a total score of 14 (moderate depression)
The AI summarizer dutifully includes all three data points in the Assessment & Plan narrative. But:
LVEF is not written to a discrete field. It exists only as text in the note body. When the practice submits a prior authorization for a repeat echocardiogram (CPT 93306), the payer's automated review cannot locate a structured LVEF value to validate medical necessity. Result: denial.
NYHA Class appears as "patient is NYHA Class III" in prose. The eCQM engine looking for heart failure quality measures (e.g., CMS 144v12 — Heart Failure: Beta-Blocker Therapy for LVSD) cannot find a discrete, coded value. Result: quality measure gap.
PHQ-9 total score is embedded in a sentence. Without a discrete LOINC-coded entry (44249-1) with instrument name, administration date, and item-level responses, the practice cannot bill 96127 (Brief emotional/behavioral assessment) per CMS fee schedule requirements. Result: missed revenue.
Quantified weekly impact for a single provider:
Problem | Frequency (5-Day Week) | Financial Impact | Administrative Burden |
|---|---|---|---|
Echo (93306) prior auth denials — missing discrete LVEF | ~3 denials | ~$1,080 ($360 × 3) | Resubmission: ~1.5 hrs each (4.5 hrs total) |
Missed 96127 billing — PHQ-9 not stored discretely | ~22 eligible units unbilled | ~$220 ($10 × 22) | Retrospective chart correction: not feasible at scale |
PA resubmissions requiring manual LVEF/NYHA documentation | ~2 resubmissions | Staff cost: ~$75 (1.5 hrs × $25/hr × 2) | 3 hours staff time |
Weekly Total (Single Provider) | ~$1,375 lost | ~7.5 hours staff/provider time |
Extrapolated across a six-provider group: approximately $8,250 per week in lost revenue and avoidable administrative cost—or $33,000+ per month.
After: Scribing.io Clinical Reasoning Engine — Step-by-Step Logic Breakdown
Same encounter. Same conversation. Same physician workflow. Zero additional clicks beyond confirmation. Here is exactly what happens under the hood:
Step 1: PHQ-9 Completeness Check and Auto-Scoring. The engine detects that the physician has discussed 7 of 9 PHQ-9 items conversationally ("sleep has been terrible—waking up at 3 AM," "appetite is down," "trouble concentrating at work," etc.). Scribing.io's instrument-aware parser maps each conversational fragment to its corresponding PHQ-9 item code. Items 8 (psychomotor changes) and 9 (suicidal ideation) have not been mentioned. The system surfaces a real-time prompt: "Items 8 and 9 (psychomotor changes, suicidal ideation) not yet captured. Confirm or skip?" The physician confirms both are negative (score 0 each). The system auto-scores the total (14), stores it as a FHIR R4 QuestionnaireResponse with all 9 individual item LOINC codes, and writes the total as an Observation resource (LOINC 44249-1) with instrument name ("PHQ-9"), administration date, scoring method, and provider attestation. This discrete entry is the evidentiary foundation for 96127 billing.
Step 2: LVEF as Coded Observation. The spoken value "EF is 30%" is captured by the NLP layer. Clinical reasoning applies here in three ways: (a) the system disambiguates "EF" from other possible meanings based on the cardiology encounter context; (b) it validates the value against the referenced echo report date, flagging if the stated value diverges by >5 points from the most recent discrete echo result in the chart; (c) it writes the value as a FHIR R4 Observation (LOINC 10230-1 for LVEF by 2D echo, or LOINC 33878-0 depending on documented method) with the numeric value (30), unit (%), date, performing provider, and reference range. For EHRs where FHIR write access to flowsheets remains limited—still common in 2026 despite ONC mandates—Scribing.io's adapter layer uses HL7 v2 ORU^R01 messages or vendor-specific APIs (Epic Flowsheet API, Oracle Health discrete write bridge) to place the value in the correct flowsheet row.
Step 3: NYHA Classification as Discrete Coded Value. "NYHA Class III" is extracted and stored as a coded Observation with SNOMED CT mapping (420913000 for NYHA Class III). This discrete value feeds CMS 144v12 and CMS 145v12 measure engines automatically and supports HCC risk adjustment documentation when paired with the appropriate HF diagnosis code.
Step 4: Payer-Preferred Medical Necessity Language Injection. The system inserts contextually populated justification into the Assessment section: "Repeat echocardiography indicated to assess interval change in LVEF (currently 30%, NYHA III) following 90 days of optimized guideline-directed medical therapy, consistent with ACC/AHA 2022 HF guideline recommendations and payer medical policy criteria." This language is template-driven but dynamically populated with the patient's actual discrete values—the LVEF and NYHA class already captured in Steps 2 and 3. This is not boilerplate pasted blindly; it references the specific clinical data that the payer's utilization review algorithm will query.
Step 5: Billing Surface with ICD-10 Linkage. Before the encounter closes, Scribing.io presents actionable billing recommendations: 96127 ×2 (PHQ-9 screening and scoring), linked to ICD-10 codes F33.1 (Major depressive disorder, recurrent, moderate) and Z13.31 (Encounter for screening for depression). The physician taps once to confirm. No manual code lookup. No after-hours chart correction.
Quantified 30-day outcome:
Metric | Before (Summarizer) | After (Scribing.io) | Delta |
|---|---|---|---|
Echo (93306) denials per provider/month | ~12 | 0 | −12 denials |
96127 units captured per provider/month | 0 | ~88 | +$880/month/provider |
PA resubmission hours per provider/month | ~12 hours | ~0.5 hours | −11.5 hours reclaimed |
Documentation close time | Next-day or later | Same-day (encounter close) | ~40 min/day reclaimed per provider |
eCQM measure capture rate (CMS 144/145) | <40% (manual abstraction) | >95% (automated discrete capture) | +55 percentage points |
Net revenue recovered per provider/month | Baseline | +$900 net (after Scribing.io cost) | Positive ROI from month 1 |
Transparency without structured data output is an aspiration. Structured data output is transparency—every value traceable, every code auditable, every billing decision defensible.
The Information Gain Gap — What Transparency Frameworks Miss About Discrete Data
The AMA's 2026 framework advances important principles: graded evidence hierarchies, explainability mandates, annual audits of AI tools used by payers. These are necessary governance rails. But they operate at the policy layer, and the accuracy crisis in AI documentation lives at the data layer.
Here is the gap the current industry conversation has not addressed: Most AI scribes only summarize text. For insurance-billed care, accuracy equals discrete data.
This is not a nuance. It is the central architectural failure of the first generation of ambient AI scribes, and it explains why health systems that deployed these tools in 2023–2025 are now facing a secondary documentation crisis: notes that look complete but cannot function computationally.
Why Narrative Accuracy ≠ Clinical Accuracy
A summarizer that correctly transcribes "ejection fraction is 35%" into a note has achieved narrative accuracy—the words are right. But the EHR's billing engine, quality measure calculator, CDS alert system, prior authorization module, and risk adjustment workflow cannot read prose. They read discrete, coded data elements:
LVEF must exist as a coded Observation. Depending on measurement method, this maps to LOINC 33878-0 (LVEF by echocardiography), LOINC 10230-1 (LVEF by 2D echo), or LOINC 32324-6 (LVEF by ventriculography). The value, unit, date, and method must be machine-readable.
NYHA Class must be stored as a discrete, coded value—not "patient is NYHA III" in text. SNOMED CT provides the ontology (420913000 for Class III). Without this, heart failure quality measures (CMS 144, CMS 145) cannot be automatically calculated, and MIPS reporting requires manual abstraction that introduces both delay and error.
PHQ-9 must be captured as a FHIR QuestionnaireResponse with item-level codes and a total score Observation (LOINC 44249-1). The CMS National Correct Coding Initiative edits for 96127 require that the instrument be identified, the score be documented discretely, and the screening be linked to an appropriate ICD-10 code. Prose alone fails every one of these requirements.
The EHR Write Problem Nobody Talks About
Even if a vendor claims "structured data output," the question CMIOs must ask is: where does it land? Many EHRs' FHIR APIs, despite ONC certification requirements, do not expose full write access to flowsheets, SmartData elements, or instrument-specific fields. A system that produces a FHIR Observation resource but can only deposit it in a note attachment has accomplished nothing—the data is still trapped in an unqueryable container.
Scribing.io addresses this through a three-tier write strategy:
FHIR R4 native write where the EHR supports it (increasingly available for Observation and QuestionnaireResponse resources in Epic February 2026+ and Oracle Health Millennium 2025.02+).
HL7 v2 ORU^R01 messages routed through the EHR's integration engine (Rhapsody, Mirth, Cloverleaf) to populate flowsheet rows directly—the same pathway used by lab instruments and monitoring devices.
Vendor-specific adapters for environments where neither FHIR write nor HL7 v2 routing is configured, using Epic Interconnect Flowsheet Write APIs, Oracle Health PowerChart custom objects, or MEDITECH Expanse discrete data services.
The result: discrete data that is queryable by the billing engine, visible to CDS, consumable by eCQM calculators, and defensible in payer audit—regardless of the EHR platform or its FHIR maturity level.
How This Connects to the AMA Framework
The AMA's augmented intelligence principles call for AI tools that are "transparent in their design and function" and that "support the physician-patient relationship." Scribing.io's provenance metadata—linking every discrete value to a specific moment in the clinical conversation, with the reasoning chain that produced the code—is the most literal implementation of these principles currently available in production clinical AI. A 2024 NIH-funded review of ambient AI scribes concluded that "the gap between narrative output and structured EHR data remains the primary barrier to realizing the efficiency promise of these tools." Clinical reasoning engines that close this gap are not incremental improvements; they are a category correction.
Technical Reference: ICD-10 Documentation Standards
Denial rates for cardiology and behavioral health services correlate directly with ICD-10 specificity. Generic or truncated codes trigger automated payer edits; maximum-specificity codes with supporting discrete data pass. Scribing.io's clinical reasoning engine enforces specificity at the point of capture—not as a retrospective coding correction.
Heart Failure Documentation
When a provider states "the patient has systolic heart failure," a summarizer might document the phrase verbatim. Scribing.io's clinical logic maps the spoken description to the maximum-specificity ICD-10 code. "Systolic" heart failure that is explicitly chronic requires I50.22—not the unspecified I50.9, which triggers payer edits and risks downcoding. If the patient also carries a hypertension diagnosis and the physician indicates hypertensive etiology, the engine applies I11.0 (Hypertensive heart disease with heart failure), which is critical for accurate HCC risk adjustment (HCC 85, coefficient 0.368 in CMS-HCC V28). The system will not assign I11.0 without explicit documentation of the causal relationship—preventing upcoding while ensuring that legitimate specificity is never lost to narrative ambiguity.
Depression Screening and Diagnosis Documentation
For behavioral health comorbidities in the cardiology-IM population, two coding paths must be managed simultaneously:
Screening code: Z13.31 – Encounter for screening for depression. This supports the 96127 charge when a PHQ-9 is administered as a screening instrument—common in HF follow-ups where depression prevalence exceeds 20% per published meta-analyses.
Diagnosis code: When the PHQ-9 score supports a clinical diagnosis, the engine maps to F33.1 – Major depressive disorder, recurrent, moderate. The specificity here matters enormously: "recurrent" (not "single episode") and "moderate" (not "unspecified") are required for clean claim adjudication. Scribing.io infers "recurrent" only when the chart history contains a prior depressive episode; "moderate" is assigned when the PHQ-9 total falls between 10 and 19, consistent with the APA scoring interpretation. If the score is 14 but no prior episode exists in the chart, the system defaults to F32.1 (single episode, moderate) and flags the discrepancy for physician review.
This is clinical reasoning applied to coding—not a lookup table, not a keyword match, but a logic chain that respects diagnostic criteria, chart history, and payer-specific specificity requirements.
Specificity Enforcement Workflow
Clinical Input | Summarizer Output (Typical) | Scribing.io Clinical Reasoning Output | Revenue/Compliance Impact |
|---|---|---|---|
"Chronic systolic heart failure" | I50.9 (Heart failure, unspecified) or prose only | I50.22 (Chronic systolic CHF) + discrete LVEF + NYHA | Clean PA, correct HCC, eCQM capture |
"Depression screen was 14" | "PHQ-9 score of 14" in note text | LOINC 44249-1 discrete Observation + Z13.31 + F33.1 | 96127 billable, screening measure met |
"NYHA Class III" | "Patient is NYHA Class III" in prose | SNOMED 420913000 coded Observation | CMS 144/145 auto-calculated |
"Hypertension causing the heart failure" | "HTN-related HF" in Assessment | I11.0 with causal link documented discretely | HCC 85 captured, RAF accuracy |
CMIO Vendor Evaluation Matrix: 11 Questions That Separate Reasoning Engines from Summarizers
Every CMIO evaluating ambient AI documentation tools in 2026 should require written answers to these questions before signing a contract. A vendor that cannot answer questions 1–5 is a summarizer, regardless of marketing claims.
# | Question | What a Summarizer Says | What Scribing.io Demonstrates |
|---|---|---|---|
1 | What LOINC code do you assign to a captured LVEF value? | "We include LVEF in the note text." | LOINC 33878-0 or 10230-1, written as FHIR Observation with value, unit, date, method. |
2 | How do you store PHQ-9 results? | "The score appears in the Assessment." | FHIR QuestionnaireResponse (9 item codes) + Observation (LOINC 44249-1) with instrument name and date. |
3 | Can your output trigger 96127 billing automatically? | "Coders can identify it from the note." | 96127 ×1–2 surfaced pre-encounter-close with ICD-10 linkage (F33.1, Z13.31) for one-tap confirmation. |
4 | Where does NYHA class land in the EHR? | "It's documented in the HPI or Assessment." | Discrete coded Observation (SNOMED 420913000) in the problem-linked flowsheet, queryable by eCQM engines. |
5 | What is your EHR write strategy when FHIR write access is unavailable? | "We generate a PDF/CDA document." | HL7 v2 ORU^R01 via integration engine, or vendor-specific adapter (Epic Flowsheet API, Oracle Health bridge). |
6 | Do you include provenance metadata for extracted values? | "The note is the documentation." | Every value links to conversation timestamp, extraction logic, and coding standard applied. |
7 | How do you handle incomplete screening instruments? | "We document what was discussed." | Real-time prompt for missing items with confirm/skip workflow; prevents partial scores from being finalized. |
8 | Can you insert payer-specific medical necessity language? | "Providers can use templates." | Context-populated necessity statements referencing actual discrete LVEF, NYHA, GDMT duration, and guideline citations. |
9 | How do you prevent ICD-10 truncation? | "We suggest codes based on note content." | Logic chain enforces maximum specificity: I50.22 not I50.9, F33.1 not F33.9, validated against chart history. |
10 | What is your eCQM pass-through rate for HF measures? | "We support quality reporting." | >95% automated capture rate for CMS 144/145 via discrete LVEF + NYHA + medication class data. |
11 | Can you demonstrate positive ROI in a 30-day pilot with my data? | "Our customers report satisfaction improvements." | Free Discrete Data Gap Report on your last 20 HF/depression encounters in 48 hours. Measured in dollars and hours. |
Free Discrete Data Gap Report — Your Next Step
Bring your last 20 HF and depression encounters. We'll deliver a free Discrete Data Gap Report in 48 hours: which notes are missing LVEF/NYHA/PHQ-9, the 96127 units you could have billed, and which charts are denial-prone—with a live plan to light up discrete writes in your EHR in under 2 weeks.
What you'll receive:
Chart-by-chart audit: discrete field presence vs. narrative-only mention for LVEF, NYHA, PHQ-9
Estimated 96127 units missed in the sample period, with projected annualized revenue recovery
Denial risk score for pending/recent echo PAs based on medical necessity language completeness
EHR-specific write pathway recommendation (FHIR, HL7 v2, or vendor adapter) with implementation timeline
Request your Discrete Data Gap Report at Scribing.io →
The gap between narrative-only AI scribes and clinical reasoning engines is not closing on its own. Summarizers are getting better at prose. They are not getting better at discrete data. Every week a cardiology-leaning IM practice operates without structured LVEF, NYHA, and PHQ-9 capture, it loses roughly $1,300 per provider in preventable denials and missed charges—and burns staff hours on rework that a reasoning engine eliminates at the point of care.
The question for CMIOs is no longer whether AI should document clinical encounters. It is whether your AI can think about what it documents—or merely repeat it in fewer words.
Scribing.io is the clinical reasoning engine that closes the accuracy gap. The data proves it. The workflow proves it. Your 20 charts will prove it.



