Predictive Risk Scoring for U.S. Medicare Advantage: Mitigating Risk Adjustment Fraud via ML-Driven Encounter Data Analytics Aligned with CMS Version 28
Main Article Content
Abstract
Background: Medicare Advantage (MA) reimburses private plans through the CMS Hierarchical Condition Category (CMS-HCC) risk-adjustment model, and the transition to CMS-HCC Version 28 has reshaped diagnostic mappings and payment calculations. Risk-adjustment fraud — diagnosis upcoding, unsupported HCC submissions, and risk-score inflation through health risk assessments — drives billions of dollars in overpayments annually and threatens payment integrity.
Objective: This study develops and empirically evaluates a machine learning-driven predictive risk scoring framework, aligned with CMS-HCC Version 28 structures, for early detection of abnormal risk-adjustment patterns in MA encounter data.
Methods: A reproducible synthetic dataset of 50,000 beneficiary-year records was constructed with a simplified V28-style HCC structure, utilization generated from true clinical burden, and 5,075 fraudulent records (10.1%) injected across three documented fraud typologies: severity upcoding, unsupported HCC submission, and HRA-only diagnosis inflation. Four gradient-based and ensemble models — XGBoost, LightGBM, Random Forest, and CatBoost — were comparatively evaluated on fraud classification (Accuracy, Precision, Recall, F1, ROC-AUC, PR-AUC) and risk-score regression (MAE, RMSE), with SHAP-based explainability analysis.
Results: All four models performed strongly, with no single algorithm dominating: LightGBM achieved the highest F1-score (0.956), Random Forest the highest precision (0.977) and lowest false-positive rate (0.25%), and CatBoost the highest recall (0.967) and best regression accuracy (MAE 0.064). Detection recall exceeded 97% for unsupported HCC submissions and HRA-only inflation, with severity upcoding hardest at 92.2%. SHAP analysis identified the year-over-year risk-score jump, submitted risk score, provider coding-intensity percentile, and share of HCCs lacking service-record support as dominant predictors.
Conclusions: Ensemble machine learning over encounter-derived integrity features enables accurate, interpretable pre-payment screening of MA risk-adjustment submissions. Validation on operational encounter data is the essential next step.