Psychiatry

Everyday medical support built on trust, quality checkups, and personal attention to your overall wellness.

Psychiatrist's workspace with digital tablet showing a structured mental status exam documentation interface for AI-assisted behavioral health notes

Mental Status Exam AI Custom Instructions: The Clinical Library Playbook for Behavioral Health Documentation

  • The Note Bloat Crisis: Why Generic AI Scribes Fail Behavioral Health

  • Scribing.io Clinical Logic: The Before-and-After of MSE Custom Instructions

  • The Clinical Relevance Filter: Architecture of a Focused MSE

  • The Mood–Affect Granularity Grid: Dimensional Precision That Survives Audits

  • Automated Risk Signal Surfacing: SI/HI, Psychosis, and Agitation Tagging

  • Interactive Complexity (90785): Auto-Detection and Flagging Logic

  • Telehealth Documentation: Conditional Capture Without Bloat

  • Technical Reference: ICD-10 Documentation Standards

  • EHR Export: Clean, Minimal Fields That Map to Payer Audit Logic

  • Implementation Pathway and Workflow Audit

The Note Bloat Crisis: Why Generic AI Scribes Fail Behavioral Health

Every behavioral health medical director has seen it: a 45-minute medication management visit produces a three-page AI-generated note where twelve lines describe the patient's recent vacation, the MSE reads "mood: good, affect: appropriate, thought process: linear," and nowhere does the note document the specific severity, functional impact, or risk qualifiers that payers require to justify the billed service.

This is note bloat—and it is the central documentation failure in behavioral health AI scribing today. It is also the exact problem that Scribing.io was architected to eliminate through Mental Status Exam AI custom instructions built on a Clinical Relevance Filter that no template-only approach can replicate.

Note bloat occurs because most ambient AI scribes were designed for procedural medicine. They transcribe everything said in the encounter, organize it by SOAP section, and present a draft that faithfully reproduces the conversation. In cardiology or orthopedics, this approach works reasonably well because the clinical signal-to-noise ratio is high: a patient describing chest pain on exertion is producing documentable data with nearly every sentence. (For a comparison of how AI scribe accuracy plays out in procedural specialties, see our analysis of ambient AI accuracy rates in Cardiology.)

Behavioral health is structurally different. The therapeutic encounter is nonlinear, narrative-rich, and deliberately conversational. A psychiatrist discussing a patient's weekend isn't gathering social history—they're observing affect reactivity, testing cognitive flexibility, and assessing interpersonal relatedness. The clinical data isn't in what the patient says about the weekend; it's in how they say it: the flattened prosody, the incongruent smile, the tangential drift when the topic approaches a trauma trigger.

When an AI scribe indiscriminately transcribes the content of these conversations, three failures cascade:

  1. Clinical signal drowns in social noise. The note captures "Patient discussed attending her daughter's soccer game and reported enjoying the weather" but fails to document that affect was constricted with restricted range and mood was dysphoric with moderate severity—the observations the clinician actually made during that exchange.

  2. MSE defaults to meaningless boilerplate. Without explicit custom instructions requiring granular descriptors, AI models default to the language patterns most common in their training data. In psychiatry notes, that means "appropriate," "normal," "intact," and "good"—words that communicate nothing to a payer auditor and provide no baseline for tracking clinical change. The AMA's CPT documentation guidelines explicitly tie E/M level selection to the complexity of medical decision-making documented in the note; boilerplate MSE language undermines every MDM argument.

  3. Risk documentation becomes implicit rather than explicit. A patient's fleeting mention of "not wanting to be here anymore" gets buried in a paragraph of narrative rather than surfaced as a structured risk assessment with ideation characterization, intent, plan, means, and protective factors—the elements the APA Practice Guidelines consider standard of care for suicide risk documentation.

The competitor landscape reflects this gap. Leading platforms emphasize template flexibility, nonlinear conversation handling, and the ability to train the AI on a clinician's writing style. These are valuable features. But they address the formatting of the note without addressing the clinical logic of what belongs in the note versus what must be excluded. Offering "full template control" still requires the clinician to manually construct the intelligence layer—the rules about what constitutes a clinically relevant MSE observation versus documentable social chatter. That cognitive burden is what burns out behavioral health providers and leads to the 9–12 minutes of post-visit note editing that erases the time savings ambient AI was supposed to deliver.

The solution is not a better template. It is a Clinical Relevance Filter embedded in the AI's custom instructions—a decisional architecture that triages every captured utterance and observation against explicit clinical relevance criteria before it enters the note.

Scribing.io Clinical Logic: The Before-and-After of MSE Custom Instructions

This section presents the concrete clinical and financial transformation when a behavioral health practice deploys Mental Status Exam AI custom instructions built on Scribing.io's Clinical Relevance Filter architecture.

Before: The Documentation Tax

A 6-clinician outpatient psychiatry group—three psychiatrists, two PMHNPs, one psychologist—runs a mixed caseload of medication management (99213–99215) and psychotherapy add-ons (90833/90836). Each clinician sees 18–22 patients per day. Their existing AI scribe produces draft notes that require 9–12 minutes of editing per encounter.

Editing Burden Breakdown: Pre-Implementation

Editing Task

Avg. Time per Note

Root Cause

Stripping non-clinical social content from HPI/Subjective

2–3 min

AI transcribes all patient speech indiscriminately

Rewriting MSE with clinically specific language

3–4 min

AI defaults to "appropriate/normal/intact" boilerplate

Adding SI/HI risk qualifiers and safety plan documentation

1–2 min

Risk mentions captured in narrative but not structured

Adjusting E/M or psychotherapy code selection

1–2 min

Note content doesn't support billed complexity level

Removing or adding telehealth-specific elements

0.5–1 min

Telehealth fields appear on in-person visits or are missing on virtual ones

Total per note

9–12 min


Financial impact: The group experiences a 14% rate of claim delays or downgrades driven by payer queries on medical necessity—specifically, auditors citing insufficient MSE specificity and missing risk documentation. Clinicians, aware of audit risk, defensively downcode: billing 99214 when documentation supports 99215, or omitting the 90833 psychotherapy add-on because the note doesn't clearly delineate the psychotherapy component from the E/M service. Conservative estimates place the monthly revenue loss at ~$7,800 across the group, combining denied claims, downcoded visits, and the administrative cost of responding to payer queries.

Provider burnout compounds the financial loss. At 20 patients/day × 10 minutes of editing, each clinician loses 3.3 hours daily to note remediation. A 2019 study published in the Annals of Internal Medicine found that physicians spend nearly two hours on EHR work for every hour of direct patient care—a ratio that behavioral health notes, with their narrative complexity, push even further.

After: Scribing.io MSE Custom Instructions with Clinical Relevance Filter

Scribing.io deploys its MSE Custom Instructions package configured with the Clinical Relevance Filter, the Mood–Affect Granularity Grid, automated risk signal tagging, Interactive Complexity detection, and conditional telehealth documentation.

Performance Comparison: Before vs. After Scribing.io Deployment

Metric

Before Scribing.io

After Scribing.io

Change

Average note editing time

9–12 min

~2 min

↓ 78–83%

Payer denial-related queries

14% of claims

~5.6% of claims

↓ 60%

Defensive downcoding frequency

Routine (est. 30% of eligible visits)

Rare (<5% of eligible visits)

↓ ~85%

Monthly net collections impact

−$7,800 (lost revenue)

+$8,600 (recovered + new)

+$16,400 swing

Clinician hours recovered per week (group)

0 (baseline)

~5 hrs/clinician/week

+30 hrs/week (group)

90785 Interactive Complexity capture

Inconsistent (clinician-dependent)

Auto-flagged when criteria detected

Est. +12–18 billable add-ons/month

The Three-Tier Triage Mechanism

When a Scribing.io session begins, the Clinical Relevance Filter applies a classification layer to every segment of captured audio. Content is triaged into three categories:

  • Clinically Documentable: Observations, symptoms, functional status, risk indicators, treatment response, medication effects, and MSE-relevant behavioral data. → Structured into the appropriate note section.

  • Clinically Informative but Non-Documentable: Contextual social content that informed the clinician's assessment (e.g., a story about a family conflict that revealed irritability and paranoid ideation). → The clinical observation derived from the content is documented; the narrative content itself is auto-redacted.

  • Non-Clinical Social Chatter: Greetings, weather discussion, scheduling logistics, casual banter used for rapport-building. → Auto-redacted from the note entirely.

This three-tier triage is what no "template flexibility" or "learn your style" feature can replicate—because it requires clinical decisional logic, not pattern matching on format preferences.

The Clinical Relevance Filter: Architecture of a Focused MSE

The Clinical Relevance Filter is not a post-processing cleanup tool. It is an inference-time decisional layer that governs what enters the note and how it is structured. This distinction matters operationally: post-processing tools let bloated content into the draft and then try to trim it, which still presents the clinician with noise to evaluate. Inference-time filtering means the first draft the clinician sees is already focused.

Hard-Limiting MSE Output to Abnormalities and Clinically Meaningful Change

The default behavior of most AI scribes is additive: capture everything, let the clinician subtract. Scribing.io inverts this. The MSE section is hard-limited to two categories of output:

  1. Abnormalities — Any MSE domain where the patient's presentation deviates from normative expectations for their demographic and clinical context.

  2. Clinically Meaningful Change — Any MSE domain where the patient's current presentation differs from their documented baseline in prior visits, regardless of whether the current presentation is technically "abnormal."

A stable patient with chronic flat affect will have affect documented (because it is abnormal). A patient with fully normal MSE across all domains receives a streamlined "MSE within normal limits with the following specifications:" followed by only the domains the clinician flags for specificity. The note never defaults to twelve lines of "normal, normal, normal, intact, intact, appropriate."

This approach aligns with the documentation standards outlined in CMS Evaluation and Management documentation guidance, which ties reimbursement to the complexity of problems addressed and data reviewed—not to the volume of normal findings listed.

The Mood–Affect Granularity Grid: Dimensional Precision That Survives Audits

Mood and affect are the two MSE domains most frequently cited in payer audits and most commonly reduced to meaningless descriptors by AI scribes. Scribing.io's custom instructions enforce a structured grid that replaces single-word descriptors with dimensional observations:

Mood–Affect Granularity Grid

Domain

Dimension

Documentation Output Examples

Affect

Range

Full / Restricted / Constricted / Flat / Labile

Intensity

Heightened / Appropriate to context / Blunted / Absent

Stability

Stable throughout / Fluctuating (specify triggers) / Rapidly shifting

Congruence

Congruent with stated mood / Incongruent (specify: e.g., smiling while describing hopelessness)

Mood

Valence

Euthymic / Dysphoric / Euphoric / Anxious / Irritable / Anhedonic (patient's own words quoted when clinically relevant)

Severity

Mild / Moderate / Severe (anchored to functional impact: e.g., "moderate—patient reports missing 3 days of work this week due to low motivation")

Why This Grid Matters to Payers

Consider the difference between these two MSE excerpts for the same patient:

Generic AI output: "Mood: depressed. Affect: appropriate."

Scribing.io output: "Mood: dysphoric, moderate severity—patient states 'I can't make myself care about anything,' reports missing 3 days of work this week. Affect: constricted range, blunted intensity, stable throughout session, incongruent at one point (smiled briefly when describing argument with spouse, then returned to tearful presentation)."

The first tells an auditor nothing. The second demonstrates medical necessity for ongoing treatment, supports a 99215 MDM complexity argument, provides a measurable baseline for next visit comparison, and contains the functional impact language that the JAMA Psychiatry literature on measurement-based care identifies as critical for treatment planning documentation.

Each dimension in the grid maps to a specific custom instruction within Scribing.io's MSE module. The AI is instructed to observe and classify along each dimension independently, which prevents the common failure mode where a single adjective ("appropriate") collapses four distinct clinical observations into one undifferentiated word.

Automated Risk Signal Surfacing: SI/HI, Psychosis, and Agitation Tagging

Risk documentation failures drive a disproportionate share of behavioral health claim denials and, more critically, represent medicolegal exposure. Scribing.io's custom instructions include an automated risk signal detection layer that operates across three tiers:

Tier 1: Suicidal and Homicidal Ideation Structuring

When any utterance or clinical observation triggers SI/HI detection, the system auto-generates a structured risk assessment framework requiring the clinician to confirm or edit:

  • Ideation characterization: Passive vs. active; frequency; duration; onset

  • Intent: Present / Absent / Ambivalent (with supporting quotes)

  • Plan: Specific vs. vague vs. denied

  • Means: Access assessment; lethal means counseling documented Y/N

  • Protective factors: Explicitly enumerated (children, religious beliefs, therapeutic alliance, stated reasons for living)

  • Risk level assignment: Low / Moderate / High with clinical rationale

  • Intervention: Safety plan reviewed/updated, crisis resources provided, collateral contact made, hospitalization considered/recommended

This structure aligns with the Columbia Suicide Severity Rating Scale (C-SSRS) framework and the SAMHSA clinical best practices for suicide risk assessment. Payer auditors reviewing medical necessity for high-frequency visits or intensive outpatient referrals look for exactly this level of structured risk documentation.

Tier 2: Psychosis Indicators

Hallucinations (auditory, visual, tactile), delusions (persecutory, grandiose, referential), disorganized thought process, and paranoid ideation are flagged and structured when detected—whether the patient reports them directly or the clinician observes behavioral correlates (e.g., the patient appears to respond to internal stimuli during the session).

Tier 3: Agitation and Behavioral Escalation

Psychomotor agitation, pressured speech, hostility, poor impulse control, and behavioral dysregulation are tagged with severity descriptors and linked to the MSE motor activity and behavior domains. These signals support medical necessity for medication adjustments, frequency of visits, and level-of-care decisions.

All three tiers feed into the note's Assessment and Plan section, where risk level directly informs treatment rationale—the linkage that payer audit algorithms check when evaluating whether a billed service level matches documented clinical complexity.

Interactive Complexity (90785): Auto-Detection and Flagging Logic

CPT 90785 (Interactive Complexity) is one of the most under-billed add-on codes in behavioral health. The AMA CPT codebook defines Interactive Complexity as applicable when specific communication factors complicate the delivery of a psychiatric procedure. Scribing.io's custom instructions detect and flag the following qualifying criteria in real time:

Interactive Complexity (90785) Auto-Detection Criteria

Criterion

Detection Trigger

Documentation Output

Third-party involvement requiring management

Caregiver/family member speaks during session; interpreter present

"Session included [mother/interpreter/case manager] whose emotional distress required direct clinical management for [X minutes]."

Patient communication difficulties

Language barrier; cognitive impairment affecting communication; use of play/art/other non-verbal modalities

"Communication complexity: patient required [interpreter services / modified language / non-verbal assessment tools]."

Emotional/behavioral dysregulation during session

Trauma trigger activation; dissociative episode; behavioral escalation requiring de-escalation intervention

"Interactive complexity: patient experienced [dissociative episode / acute trauma response] at [time point], requiring [X minutes] of stabilization before psychiatric evaluation could continue."

Conflicting guardian/patient treatment goals

Detected disagreement between patient and accompanying party regarding treatment plan

"Conflicting treatment perspectives between patient and [guardian/spouse] regarding [medication adherence / hospitalization / treatment modality], requiring mediation and separate clinical assessment of each party's concerns."

When any of these criteria are detected, Scribing.io generates a 90785 flag in the coding suggestion panel with pre-populated documentation supporting the add-on. The clinician reviews, confirms, and the billable add-on is captured. Across the 6-clinician group in our reference scenario, this auto-detection recovered an estimated 12–18 billable 90785 add-ons per month that were previously missed—representing approximately $960–$1,440 in monthly revenue at average reimbursement rates.

Telehealth Documentation: Conditional Capture Without Bloat

Telehealth documentation requirements add another layer of note bloat when AI scribes apply them indiscriminately. Scribing.io's custom instructions handle telehealth metadata conditionally:

  • Session type detection: The system identifies whether the encounter is in-person or telehealth based on session initiation metadata (platform integration) or clinician input at session start.

  • Telehealth = TRUE: The note auto-populates patient location (originating site), provider location (distant site), consent for telehealth (documented or referenced to standing consent), and Modifier 95 (synchronous audio-video) applied to the claim. These fields are structured as a single-line metadata block at the note header—not woven into the clinical narrative.

  • Telehealth = FALSE: All telehealth fields are suppressed entirely. No "N/A" placeholders, no empty telehealth sections, no vestigial Modifier 95 references.

This conditional logic prevents two common errors: telehealth metadata appearing on in-person visit notes (which triggers audit flags) and telehealth notes missing required location/consent documentation (which triggers claim denials). CMS telehealth billing requirements specify that originating site and consent must be documented—but they do not require that documentation to consume multiple note paragraphs.

Technical Reference: ICD-10 Documentation Standards

Behavioral health ICD-10 coding is where vague MSE documentation creates the most direct revenue impact. Payers deny or downgrade claims when the documented clinical picture does not support the specificity level of the reported diagnosis code. Scribing.io's MSE custom instructions are engineered to generate documentation that maps directly to ICD-10 specificity requirements.

Depressive Disorders

F33.1 Major depressive disorder, recurrent, moderate; F41.1 Generalized anxiety disorder; F31.9 Bipolar disorder — these codes require documentation that distinguishes between single episode and recurrent, and between mild, moderate, and severe. The Mood–Affect Granularity Grid directly supports this: when the custom instructions enforce severity documentation anchored to functional impact ("moderate—missing 3 days of work/week"), the note provides the clinical evidence that justifies "moderate" over "mild" or the nonspecific "unspecified." Without that granularity, coders are forced to select unspecified codes, which carry higher denial rates.

Anxiety, Bipolar, and Behavioral Codes

For unspecified; R45.851 Suicidal ideations; R46.89 Other symptoms and signs involving appearance and behavior; R41.840 Attention and concentration deficit — the R-codes in particular are critical secondary diagnosis codes that many practices fail to capture. R45.851 (Suicidal ideations) should be reported as a secondary code whenever SI is documented, regardless of whether the primary diagnosis is MDD, PTSD, or borderline personality disorder. Scribing.io's risk signal surfacing automatically flags when SI documentation meets the threshold for R45.851 reporting and includes it in the coding suggestion panel.

R46.89 (Other symptoms and signs involving appearance and behavior) captures MSE-documented behavioral abnormalities—psychomotor agitation, bizarre behavior, poor grooming—that don't map to a specific psychiatric diagnosis code but support medical necessity for the visit. R41.840 (Attention and concentration deficit) is flagged when cognitive MSE domains reveal impairment, supporting referral for neuropsychological testing or justifying medication adjustments targeting cognitive symptoms.

Maximum Specificity Logic

Scribing.io's coding suggestion engine applies a specificity hierarchy: the system will never suggest a nonspecific code when the documented MSE and clinical narrative contain sufficient information to support a more specific code. This is operationalized through a mapping layer that cross-references MSE output (severity, episode type, symptom cluster) against the CMS ICD-10-CM code set to select the highest-specificity match. The clinician retains final code selection authority, but the default suggestion is always maximum specificity—reversing the typical pattern where defaults trend toward unspecified codes that invite denials.

EHR Export: Clean, Minimal Fields That Map to Payer Audit Logic

A focused MSE has limited value if the export to the EHR re-introduces bloat through field mapping problems. Scribing.io's export architecture is designed around minimal-field structured output:

  • Discrete MSE fields: Each dimension of the Mood–Affect Granularity Grid exports as a discrete, queryable field in the EHR—not embedded in free-text narrative. This enables measurement-based care tracking (comparing affect range across visits) and supports population health reporting.

  • Risk assessment as structured data: SI/HI assessment components export as discrete fields (ideation: Y/N, intent: Y/N, plan: Y/N, means access: Y/N, risk level: low/moderate/high) alongside the narrative risk assessment paragraph. The structured data feeds patient safety dashboards; the narrative provides the clinical context for auditors.

  • Coding suggestions as metadata: Suggested CPT and ICD-10 codes export as metadata tags, not as embedded note text. This prevents the common failure where AI-generated coding suggestions appear in the clinical note itself—a compliance risk that auditors flag as potential upcoding evidence.

  • Telehealth metadata as header fields: Location, consent, and modifier data export to the EHR's encounter-level metadata fields, not to the clinical narrative body.

The result is a note that, when printed or viewed in the EHR, reads as a focused clinical document—not a data dump decorated with template artifacts.

Implementation Pathway and Workflow Audit

Deploying MSE custom instructions is not a software toggle. It requires calibration to each practice's documentation conventions, payer mix, and clinical workflow. Scribing.io's implementation follows a structured pathway:

  1. Baseline Audit (Day 1–3): Scribing.io's clinical documentation team reviews 10–15 sample notes from the practice, identifying specific note bloat patterns, MSE granularity gaps, missed billing opportunities, and compliance risks.

  2. Custom Instruction Configuration (Day 4–7): The Clinical Relevance Filter is calibrated to the practice's specialty mix (psychiatry vs. psychology vs. therapy-only), payer requirements (commercial vs. Medicare vs. Medicaid behavioral health carve-outs), and EHR platform.

  3. Parallel Run (Week 2): Clinicians use Scribing.io alongside their existing workflow for one week. Each note is generated in both systems; the clinical team compares output quality, editing time, and coding accuracy.

  4. Go-Live and Optimization (Week 3+): Full deployment with weekly optimization reviews for the first month, then monthly thereafter.

Book a 15-minute Workflow Audit to receive a plug-and-play MSE Custom Instruction set: a Mood–Affect granularity grid, Interactive Complexity (90785) trigger prompts, telehealth Modifier 95 capture without bloat, and a redlined sample of your current note showing exactly what to keep, cut, or restructure to pass payer audits and shorten editing time by 5–8 minutes per note. Schedule your audit at Scribing.io.

Behavioral health documentation has tolerated vague, bloated notes for decades because the alternative—manually writing precise, audit-proof MSE language for every encounter—was unsustainable at clinical volume. Mental Status Exam AI custom instructions built on a Clinical Relevance Filter change that calculus permanently. The question for medical directors is no longer whether AI scribing works for behavioral health. It's whether your current AI scribe is producing notes that would survive a payer audit—or notes that are costing you $7,800 a month in revenue you've already earned.

Still not sure? Book a free discovery call now.

Frequently

asked question

Answers to your asked queries

Can we get started today?

Can I edit or review notes before they go into my EHR?

Does Scribing.io work with telehealth and video visits?

Is Scribing.io HIPAA compliant?

Is patient data used to train your AI models?

Still not sure? Book a free discovery call now.

Frequently

asked question

Answers to your asked queries

Can we get started today?

Can I edit or review notes before they go into my EHR?

Does Scribing.io work with telehealth and video visits?

Is Scribing.io HIPAA compliant?

Is patient data used to train your AI models?

Still not sure? Book a free discovery call now.

Frequently

asked question

Answers to your asked queries

Can we get started today?

Can I edit or review notes before they go into my EHR?

Does Scribing.io work with telehealth and video visits?

Is Scribing.io HIPAA compliant?

Is patient data used to train your AI models?

Image

Clinical Precision.
Zero Documentation Debt

Finish Your Charts - Go Home on Time.

Clinical Precision.
Zero Documentation Debt

Finish Your Charts - Go Home on Time.