Why Benchmarking Matters

RAF score benchmarking is the process of comparing a Medicare Advantage plan's average Risk Adjustment Factor scores against peer plans, regional averages, historical baselines, and clinically expected norms — transforming raw scores into actionable intelligence that identifies documentation gaps, coding underperformance, or over-coding risk before it affects revenue or triggers a RADV audit.

A RAF score in isolation tells you very little. Knowing that your plan's average RAF is 1.05 is meaningless unless you understand how that compares to what it should be. Benchmarking transforms raw risk scores into actionable intelligence by placing them in context against peers, historical performance, and clinical expectations.

The financial stakes are significant. For a 50,000-member Medicare Advantage plan, every 0.01 gap between actual and optimal average RAF represents roughly $5.2 million in annual capitation payments. Plans that do not systematically benchmark their RAF scores operate with a blindfold over their most critical revenue driver.

  • Revenue Validation: Benchmarking confirms whether CMS payments reflect the true clinical burden of your population, or whether documentation and coding gaps are suppressing reimbursement
  • Compliance Calibration: Scores that run significantly above benchmarks can signal over-coding risk and potential RADV audit exposure, while scores running below may indicate missed conditions
  • Operational Insight: Comparing RAF performance across provider networks, regions, and disease categories reveals where operational processes are working and where they are failing

With CMS-HCC V28 now the sole model for 2026 payment calculations, historical benchmarks from V24 are no longer valid comparisons. Every plan needs to re-establish baselines using V28-specific data, making benchmarking more critical than it has been in the past decade.

The Underperformance Gap

Industry analysis shows that the average MA plan underperforms its optimal RAF by 0.03 to 0.08 points. For mid-size plans, this translates to $15M to $42M in annual revenue left on the table due to documentation gaps, missed recapture, and coding inconsistencies.

V28 Baseline Reset

CMS projected a 3.12% average RAF decline under V28. Plans experiencing declines greater than 5% relative to clinically equivalent populations should investigate whether the drop reflects model changes or operational underperformance.

Benchmarking Methodologies

Effective RAF benchmarking requires a structured methodology that accounts for population differences, model changes, and data timing. There is no single correct benchmark; the most accurate picture comes from triangulating multiple approaches.

  • Actuarial Expected RAF: Starting with your population's demographic profile and known clinical history, actuaries calculate an expected RAF score that represents what the plan should achieve assuming complete documentation and coding accuracy. The gap between expected and actual RAF is the most precise measure of underperformance.
  • Year-Over-Year Trend Analysis: Comparing your plan's RAF trajectory across payment years reveals whether performance is improving, stable, or declining. This method must account for V28 transition effects, membership churn, and any provider network changes that alter the clinical mix.
  • Peer Cohort Comparison: Grouping your plan against competitors with similar geographic footprints, enrollment sizes, and dual-eligible ratios enables meaningful cross-plan benchmarking. CMS publishes contract-level RAF data that supports this analysis, though it lags by 12 to 18 months.
  • Clinical Encounter-Based Benchmarking: Analyzing encounter data to identify conditions documented in the medical record but not captured as HCCs provides a ground-truth benchmark. This approach requires integration between clinical and coding data sources.
  • HCC Category-Specific Analysis: Rather than benchmarking a single average RAF, break performance down by disease family. A plan may perform well on diabetes capture but underperform significantly on cardiovascular or renal categories.

The most mature organizations combine all five methodologies into a composite benchmarking framework that provides both a top-line performance view and granular diagnostic detail.

Internal vs External Benchmarks

Benchmarking sources divide into two categories, each offering distinct advantages and limitations.

Internal Benchmarks

  • Historical Plan Performance: Your own year-over-year RAF trends, adjusted for V28 model effects, provide the most controlled comparison since the population and operational context are consistent
  • Provider-Level Variation: Comparing RAF outcomes across your contracted provider groups reveals which networks capture risk effectively and which leave value on the table
  • Regional Sub-Plan Analysis: For plans operating across multiple counties or states, regional benchmarks highlight geographic variation in documentation practices and coding quality
  • Prospective vs Retrospective Gap: Comparing batch-scored prospective RAF estimates against final CMS-assigned retrospective scores measures the accuracy of your internal risk prediction

External Benchmarks

  • CMS National Averages: Published in the annual Rate Announcement and Advance Notice, these provide a top-level reference point for community and institutional populations
  • County-Level Benchmarks: County base rates and demographic profiles allow localized comparisons that account for regional health status differences
  • Industry Reports: Organizations like Milliman, Wakely, and Oliver Wyman publish MA-specific benchmarking data that enables peer comparison without direct data sharing
  • CMS Bid Data: Contract-level bid summaries reveal competitor RAF assumptions, providing insight into whether your scores align with market expectations

The most effective risk adjustment analytics programs maintain a balanced scorecard that integrates both internal and external benchmarks, updated at least quarterly.

Free Resources

Three downloads risk adjustment teams actually use

Checklists, playbooks, and frameworks — built for analysts, auditors, and VPs working RAF, RADV, and HCC.

Checklist

2026 RADV Audit Readiness Checklist

12-point compliance checklist for documentation, diagnosis code validation, extrapolation defense, and pre-audit scrub workflows.

Playbook

RAF Score Optimization Playbook

Tactical guide for analysts: HCC recapture workflows, V28 transition impacts, prospective gap-closure plays, and KPIs that move RAF lift.

Playbook

Risk Adjustment Analytics Playbook

How payer leaders sequence prospective and retrospective risk adjustment for compounding RAF lift. Deployment patterns, KPIs, and a VP-level operating rhythm.

Benchmark Your RAF Scores: Our RAF Score tools provide population-level benchmarking against industry averages, helping you identify where your plan over- or underperforms. Explore RAF Score tools →

Key Metrics to Compare

RAF benchmarking extends beyond a single average score. A comprehensive benchmarking framework tracks multiple interrelated metrics that together reveal the full picture of risk capture performance.

RAF score benchmarks by Medicare Advantage plan segment under CMS-HCC V28 (2026). Revenue impact estimates assume ~$9,500 annual capitation per member.

Plan Segment Avg RAF Score High Performer (90th %ile) Low Performer (10th %ile) Revenue Impact per 0.01 Gap
Large MA Plan (>100K members) 1.08–1.12 1.18+ <1.02 ~$950K/yr
Mid-Size MA Plan (25K–100K) 1.05–1.09 1.15+ <0.99 ~$285K/yr
Small MA Plan (<25K members) 1.02–1.07 1.13+ <0.97 ~$95K/yr
D-SNP (Dual-Eligible) 1.48–1.65 1.80+ <1.35 ~$1.9M/yr
PACE Program 2.10–2.40 2.60+ <1.90 ~$2.4M/yr
National MA Average 1.09 1.20+ <0.98 Varies by size
  • Average Plan RAF Score: The headline metric, compared against actuarial expected, national average (approximately 1.10 for community), and peer cohort. Track separately for community, institutional, ESRD, and new enrollee segments.
  • HCC Recapture Rate: The percentage of prior-year HCCs successfully re-documented in the current payment year. Top-performing plans achieve 85% or higher; the industry average hovers around 75%. Every unrecaptured HCC represents lost revenue.
  • HCC Prevalence by Category: The number of unique HCCs per member, broken down by disease family. Compare against expected prevalence based on population demographics and known chronic disease rates.
  • Suspect Condition Conversion Rate: Of conditions identified as likely present through predictive analytics, what percentage are ultimately confirmed and coded? Rates below 40% signal either poor targeting or provider engagement failures.
  • Encounter Submission Completeness: The percentage of eligible encounters successfully submitted to CMS within required timelines. Submission gaps directly suppress RAF scores regardless of documentation quality.
  • Coding Specificity Index: Measures whether coders consistently select the most specific ICD-10 code, which under V28 determines whether a diagnosis maps to an HCC at all
  • RAF Variance Distribution: Analyzing the distribution of individual member RAF changes year-over-year identifies outlier patterns that warrant investigation for both under-coding and over-coding

Identifying Underperformance

Underperformance in RAF scoring rarely presents as a single obvious failure. It typically manifests as a collection of moderate gaps across multiple dimensions that compound into significant revenue impact.

  • Systematic RAF Drift: A gradual decline in average RAF over three or more quarters that exceeds expected V28 model effects. If your population demographics and clinical mix are stable but RAF is declining, the problem is operational, not clinical.
  • Low-RAF High-Cost Members: Members with RAF scores below 1.0 but medical costs exceeding $20,000 annually represent the clearest signal of missed risk capture. These members are generating losses because their documented conditions do not reflect their actual resource consumption.
  • HCC Desert Providers: Provider groups within your network that consistently generate fewer HCCs per encounter than peers treating similar populations. This pattern typically indicates documentation deficiencies rather than healthier patients.
  • Disease-Specific Gaps: Comparing your plan's HCC prevalence rates against epidemiological expectations reveals disease categories where capture systematically falls short. If national data suggests 18% diabetes prevalence among your demographic but your HCC data shows 12%, there is a capture gap.
  • New Enrollee Underperformance: New members typically show lower RAF scores in their first year due to incomplete data. Plans that fail to close this gap within 12 months are leaving first-year revenue on the table.

The key is distinguishing between underperformance that reflects true operational gaps and variation that results from legitimate differences in population health. Provider-level analytics are essential for this distinction, allowing plans to control for case mix when evaluating coding performance.

Action Plans for Improvement

Once benchmarking identifies specific areas of underperformance, the next step is translating findings into targeted action plans with measurable outcomes.

  • Prioritize by Revenue Impact: Rank identified gaps by estimated revenue at risk. An HCC recapture gap in diabetes affecting 5,000 members generates more recoverable revenue than a rare condition gap affecting 50 members, even if the per-member RAF impact is larger for the rare condition.
  • Deploy Targeted Chart Reviews: For members identified as likely under-coded, initiate retrospective chart reviews focused on the specific disease categories where benchmarking shows the largest gaps. Use automated RAF scoring tools to pre-identify high-probability candidates.
  • Provider Education Programs: When underperformance concentrates in specific provider groups, develop documentation improvement programs tailored to the specific HCC categories those providers are missing. Generic training produces generic results.
  • Encounter Submission Remediation: If benchmarking reveals submission completeness below 95%, treat encounter operations as a top priority. No amount of documentation improvement matters if encounters do not reach CMS.
  • Prospective Interventions: Shift from retrospective correction to prospective capture by embedding risk adjustment intelligence into point-of-care workflows, enabling providers to address documentation gaps during the patient visit rather than after the fact.
  • Quarterly Benchmarking Cadence: Establish a formal quarterly review cycle where benchmarking results are presented to executive leadership, improvement actions are tracked, and targets are updated based on progress

Organizations that approach benchmarking as a continuous operational discipline rather than an annual exercise consistently outperform peers. The goal is not a single point-in-time correction but a systematic capability that prevents underperformance from accumulating in the first place.

Key Insight: RAF score benchmarking under CMS-HCC V28 requires re-establishing every baseline from scratch. Plans that carried V24-era benchmarks into 2026 are measuring against an obsolete standard. The organizations gaining competitive advantage are those that built V28-native benchmarking frameworks during the transition period and now have 18 months of V28 trend data to inform their performance targets.

Ready to Benchmark Your RAF Scores?

See how our Medicare Risk Revenue Intelligence platform helps plans identify underperformance, track benchmarks, and close RAF gaps with real-time analytics.

Schedule a Demo