Healthcare Policy · Risk Adjustment

Simplifying Risk Adjustment

By Syed Muzayan Mehmud · Healthcare Actuary · LinkedIn ↗
Read the Article ↓
1. Abstract / Overview

Tens of billions of dollars in risk adjustment payments flow each year based on systems that reward sophisticated claim coding over actual patient health. Efforts to curb this gaming have consistently fallen short, not because regulators haven't tried, but because every fix leaves the underlying incentive intact. As long as payments depend on claim data, organizations will invest in optimizing that data.

Healthcare risk adjustment methodologies have become increasingly complex, largely due to the belief that complexity enhances accuracy. However, a disconnect exists between how healthcare risk is assessed and how payments are adjusted.

This paper argues that the only durable solution is to sever the connection between risk adjustment and claim data entirely, and demonstrates that doing so is feasible. A simplified approach using self-reported health status, with no diagnosis codes or drug records, achieves comparable payment accuracy to the current system when measured the way it actually matters: at the health plan level, where payments are made. Accuracy is not a barrier to simplification.

2. The Disconnect

The Disconnect

Accurate risk adjustment is crucial for ensuring health plans receive appropriate compensation for the risks they manage, discouraging risk selection and fostering market stability. Risk adjusted payments are critical for health plan financials and a significant portion of the premium they collect.

While model accuracy is not the only criterion in the development of a risk adjustment methodology, it is an important one that has guided the evolution of these models towards more complexity. We have added more condition categories, hierarchies within conditions, Diagnosis Codes (Dx), National Drug Codes (NDC), different sets of model coefficients for population subgroups, condition interactions, and so on.

However, more complexity comes at a significant cost. There is the cost to administer the program and a cost borne by health plans and providers with limited capacity to analyze and keep up with an elaborate method. There is also a large cost to society in the form of large potential overpayments due to differential diagnosis-coding, and large health plan investments into finding more codes that get passed to the consumer in the form of higher premiums. In the Medicare Advantage (MA) program alone, overpayments due to over-coding are estimated to cost taxpayers billions of dollars. [1] [2] [3]

Efforts to address over-coding have taken various forms: removing or adjusting specific diagnosis codes, adding safeguards against duplicative submissions, or developing proxies designed to be harder to game. These approaches share a common limitation. None of them sever the underlying incentive. As long as risk adjustment payments depend on the quantity and quality of claim data, organizations that invest heavily in documentation will continue to outperform those that do not. Regulatory adjustments change the rules; sophisticated participants retool and adapt. The game continues under a different name.

Has a perceived increase in the accuracy of risk adjusted payments been worth the complexity and added costs to society?

Healthcare risk adjustment accuracy is typically measured using member-level measures such as r-square, or predictive ratios for subgroups such as demographic cohorts or disease conditions.

However, there is a glaring disconnect in how we measure accuracy and how risk adjustment really works. We don't transfer payments between members of health plans or cut checks to demographic cohorts or those for specific diseases. Risk adjustment takes place at the level of health plans. These plans typically have thousands of members, with varying mix of demographics, disease conditions and so on. The measurement of payment accuracy should not rely on member-level metrics or predictive ratios for selected subgroups, but rather on calculating the difference between predicted risk for a health plan and the actual risk incurred.

3. A Better Way

A Better Way

Once we look at accuracy in terms of how entities get risk adjusted instead of how members get risk scored, we find that far simpler models approach the accuracy of more complex ones.

Before I show with real data how this can work, let's spend a little time thinking through how this can ever work. It may seem counterintuitive that a less accurate model at the member level can perform just as well at the health plan level, where risk adjustment payments are actually made.

The key to understanding this is that error dynamics are different when we are thinking of total variability within a population vs. the variability within non-random groups of members. If risk adjustment models were perfect (i.e., r-square of 1.0), then their predictions would be perfect regardless of whether we measure them at the member or the group level. But these models are not perfect. The best ones have large magnitudes of errors, with r-squares well below 0.5. This leaves the door open to unexpected results when we roll up predictions and associated errors from a member to a group level.

A simple example is worth a few paragraphs. So, consider the following illustration, that also foreshadows how we will analyze real data.

Exhibit 1 shows data on five hypothetical individuals. Their annual expenditure is in the Actual column. These expenditures are not in dollar terms but are scaled so that they average to 1.0 over the population. There are predictions for the risk of these members from two different models, Prediction 1 (P1) and Prediction 2 (P2). How do we tell which model predicts expenditures better?

Exhibit 1: Member-level data
Person Actual P1 (Prediction 1) P2 (Prediction 2) Plan
11.271.071.18Plan 1
21.571.241.61Plan 1
30.781.050.92Plan 2
41.081.100.46Plan 2
50.290.540.83Plan 2
Avg1.001.001.00
69.9% ★26.4%
MAPE21.6% ★28.5%
Source: author's analysis. · ★ Winner at member level

We can calculate measures such as the coefficient of determination (R²) and Mean Absolute Error (MAE). In this example, P1 has an R² of 70% and P2 has an R² of a distant 26% by comparison. Further, the MAE for P1 is significantly better (22%) relative to P2 (29%). MAE is a better measure for risk adjustment payment accuracy as it is the difference between what was expected and what happened.

By either measure, clearly, prediction 1 is the better model.

Or is it? Two of our hypothetical individuals belong to plan 1, and three to plan 2. This is indicated in the Plan column. Recall that we don't make risk adjustment payments to individuals; we make these to health plans. We need to roll up this data at the plan level¹ and re-calculate our performance metrics.

Once we do this, we get the table in Exhibit 2, and the tables have turned! P1 has an R² of 61% on a group basis, whereas P2 has a staggering R² of over 99%! Further, the MAE for P1 is now higher at 21% relative to P2 at just 2%.

Exhibit 2: Plan-level data
Plan Actual P1 (Prediction 1) P2 (Prediction 2)
Plan 11.331.121.31
Plan 20.670.880.69
Avg1.001.001.00
61.0%99.6% ★
MAPE20.5%2.2% ★
Source: author's analysis. · ★ Winner at plan level — the reverse of Exhibit 1

In this illustration, it would be a mistake to pick P1 over P2 as the better model. On the metric that ultimately matters, payment accuracy to risk adjusted entities — this is a no contest. We risk making a similar mistake within large scale healthcare risk adjustment programs if we don't align the measurement of risk adjustment accuracy with how these payments get made.

3.1 Testing with Data

To keep with the earlier illustration, I have actual healthcare costs that are scaled to 1.0 over more than ninety-thousand real-world observations from Agency for Healthcare Research & Quality's (AHRQ) Medical Expenditure Panel Survey (MEPS). The data is for years 2017–2019, including households with Medicare, Medicaid, and Commercial coverage.

Representing the status-quo method, I have risk scores from Centers for Medicaid & Medicare (CMS) Hierarchical Condition Category (HCC) model — the same model that is used to adjust payments to plans in the Affordable Care Act (ACA) program. This model uses demographic information such as age, gender, diagnosis codes as well as prescription drug codes. Thousands of diagnosis codes and NDCs are organized into more than 150 HCC indicators, including hierarchies, condition interactions, pharmacy flags and so on.

For my alternative approach, I use a simple method based on age, gender, and just twelve health status survey questions. The questions are along the lines of what most of us experience when filling out paperwork in a doctor's office. These include "have you been diagnosed with cancer?", "do you have diabetes?", and "do you have difficulty walking up a flight of stairs?" etc.

An approach that utilizes detailed diagnosis-codes will be most accurate in a concurrent application. I chose this comparison to give my alternative approach its stiffest test.

The member-level performance from the two models is presented in Exhibit 3.

Exhibit 3: Member-level comparison of HCC (Dx) and Simple (No Dx) predictions
Metric HCC Model (Dx codes) Simple Model (12 questions)
21.1%
13.8%
MAPE 95.8%
102.3%
At the member level, the HCC model outperforms on both metrics. High MAPE values reflect the large individual-level variance in healthcare costs.
Source: author's analysis. Based on 90,853 MEPS observations, 2017–2019.

As expected, the highly complex diagnosis-based HCC risk adjustment model performs well, with an R² of over 20%. The R² of my simple model isn't too far behind at 14%. I note that the performance of the diagnosis-based HCC model used in the Medicare Advantage (MA) program is around 11%.

Next, I assemble hypothetical health plans from the data, that look just like real health plan entities, with typical group sizes as well as a typical distribution of expenditures that is observed for real-world health plans. Then, I recalculate our metrics just like we saw in the illustration above.

Exhibit 4 shows the R² and MAE results for the two very different approaches by plan size. The diagnosis (and Rx)-code based HCC model, and the ultra-simple model that relies only on demographics and a few questions about a member's health status.

Exhibit 4: Plan-level comparison of HCC (Dx) and Simple (No Dx) predictions
R² by Plan Size
Higher is better
MAPE by Plan Size
Lower is better · 1-Member bar reflects member-level variance
Source: author's analysis.

We see the relatively lower performance of the simple model replicated for plan size of 1 (which is equivalent to the prior table for individuals). As soon as we started measuring payment accuracy between groups of people, the performance gap closes between the diagnosis-based and the simple model. In fact, the simple model tends to perform better.

This is a striking result. A simplistic model relying on twelve questions about health status outperforms a model relying on over 10,000 diagnosis codes and over 10,000 National Drug Codes (NDCs). There are two reasons for this, both of which accrue from a central point of this paper. One, we need to measure accuracy of models in a way that represents groups of members that are risk adjusted, and crucially, we need to build our models in that way. We get significantly better payment accuracy when we fit our models to groups of individuals representative of health plans rather than individual members.

I note that for group sizes around 500 members, the more complex diagnosis-code based model performs slightly better. More than 99% of the members in MA or ACA are in plans with higher membership, however, we still need to think about how to address payment accuracy for the tiny health plans, which are relatively few. One relatively simple way to accomplish this is through a two-sided risk corridor. In the following set of results, I implemented a 25% corridor, where a health plan is credited 25% of the difference between actual aggregate risk and that estimated by the model. Such an approach can be implemented below a very low threshold of plan size and is likely not necessary to do for larger plans that comprise almost the entire membership in the MA or ACA programs.

Exhibit 5: Same as Exhibit 4, with an added set of results combining the Simple method with a risk corridor
R² by Plan Size (with Corridor)
Higher is better
MAPE by Plan Size (with Corridor)
Lower is better
Source: author's analysis.
4. Further Exploration

Further Exploration

The way we measure accuracy of risk adjustment models informs policy and the direction of the development of risk adjustment models. We can do better at aligning the way we measure model performance to how these models work in practice.

This paper demonstrates that we need not chase complexity on the grounds of risk adjusted payment accuracy. There is an enormous potential for desirable outcomes if we can simplify risk adjustment methodologies. The specific option discussed in this paper is self-reported health status, with no reliance on healthcare claim data.

Towards the objectives of simplicity, mitigating gaming, and higher payment accuracy, there are several areas for further research and development:

  • Research into candidate localities for a pilot MA program. In this way we can run a parallel calculation and test the effectiveness of a health status survey compared to the current healthcare claim data and diagnosis-code based approach.
  • To address concerns around gaming of survey answers, such a survey should be centrally administered and controlled. This will go a long way towards leveling the playing field. We need research into how technology can be used to accomplish this efficiently. Further research should focus on identifying the most predictive and least manipulable survey questions, as well as developing robust mechanisms for centralized, unbiased survey administration.
  • We already have marketing regulations designed to prevent discriminatory practices including member selection based on health conditions. There can be further research into appropriate regulations that would prevent gaining an unfair advantage for health status survey completions (e.g., disallowing coaching, incentives, etc.).
  • Research into additional resources that could be deployed by the central administering authority to mitigate uneven survey responses where necessary (e.g. multi-lingual support, access to technology, educational programs, etc.).
  • Members have concerns about privacy. In the MA program detailed data is already centrally collected on members in an identifiable way. In the ACA program, detailed member-level data is collected centrally in the EDGE program but is somewhat de-identified. In this paper I am advocating for a group level approach, so central data collection could be completely de-identified, lessening privacy concerns in this approach vs. the reality today. We don't care about linking health status survey responses to an individual or even a member identifier; we just care about aggregating those responses up to the health plan level.
  • Along with measuring accuracy, the way we build and calibrate a risk adjustment model also depends greatly on whether we take a member-level or a group-level view. More research is needed on analyzing group-level risks with real world characteristics and for selecting variables and calibrating models on that basis.
  • Of course, model accuracy is not the only consideration in evaluating the success of a risk adjustment program. One consideration is the resources required to execute a given methodology. A far simpler risk adjustment approach will also be easier to administer. The efficiencies gained could unlock resources for researching other desirable features of risk adjustment policy, for example, using social determinants of health towards improving healthcare equities in vulnerable populations.
  • An important limitation of the data I used in this analysis is that all the information is self-reported. Participants in the MEPS survey were asked about each hospitalization, outpatient or doctors' visit. Professional coders then translated detailed survey responses from a medical visit into diagnosis codes and prescription drug codes. The general health status questions that I used in my alternate approach were mostly yes/no questions from a different part of the survey, and were representative of how this process might work in practice. As I noted above, there is evidence here that self-reported data can explain a large amount of variation in healthcare spending, perhaps even greater than what we currently observe in a program like MA. Still, we need to do a pilot study to assess the applicability and generalizability of the results presented in this paper.
5. Conclusion

Conclusion

The two main points in this paper are:

  1. We need to align how we measure healthcare risk adjustment accuracy with how these adjustments are actually administered.
  2. Once this alignment is achieved, a far simpler approach becomes feasible — one that could not only improve payment accuracy but also eliminate the need for healthcare claim data.

The second point is important. Tens of billions of dollars move to or between health plans in each of the MA and ACA programs each year. As long as healthcare claim data drives risk adjustment, incentives will favor measuring differential data quality rather than the actual health status of members. This undermines the purpose of risk adjustment. The only durable fix is to sever the connection between risk adjustment and claim data entirely.

A methodology grounded in self-reported health status eliminates the incentive to code more aggressively, because there are no codes to chase. The resources currently spent on identifying, documenting, auditing, and adjudicating diagnoses could be redirected toward actual healthcare — toward improving the health of members rather than measuring it for payment purposes.

To assess the feasibility and impact of this simplified methodology, a pilot program should be conducted alongside the existing approach. This would allow for real-world testing, refinement of the model, and provide insights into potential implementation challenges and solutions.

1 This includes rescaling to 1.0 over the plan-level.

[1] Committee for a Responsible Federal Budget / Health Systems Innovation Network. Medicare Advantage Overpayments ↗

[2] HHS Office of Inspector General, OEI-03-17-00474. Some Medicare Advantage Organizations Submitted More Diagnoses Than Allowed, Resulting in Substantial Overpayments ↗

[3] SSRN Working Paper. papers.ssrn.com ↗

FAQ

Frequently Asked Questions

People already fill out health status questionnaires at the doctor's office. They do it routinely, sometimes at every visit, sometimes repeatedly. Chronic conditions, functional limitations, general health: these are standard intake questions across the country. Shifting the survey to plan enrollment and annual renewal simply changes when it is administered, not what it asks.

Delivery is not a technical problem. The survey can go out through a mobile app, a web form, or an automated phone line. AI voice tools can handle it in any language without a human operator. Multiple channels and follow-up reminders can support participation in the same way outreach already works for annual wellness visits.

The backend is equally tractable. Collecting, storing, and analyzing survey responses at this scale is routine cloud infrastructure. Storage costs are negligible. Analysis pipelines can return risk scores quickly. This is a substantially simpler and cheaper data operation than the current system, which requires ingesting, coding, auditing, and adjudicating millions of diagnosis and drug records every year.

Efforts to curb over-coding tend to follow the same playbook: remove an HCC, adjust a coefficient, add a new safeguard. The problem is that the organizations gaming the system are sophisticated. They can retool their analytics in days and their clinical interventions in weeks. Every rule change triggers adaptation. It is whack-a-mole — and the mole always wins eventually, because the underlying incentive has not changed. Risk adjustment payments still depend on claim data, so the financial reward for coding more aggressively remains intact regardless of which specific codes or rules are in play.

The only way to stop the game is to change what the game is played on.

The honest obstacle is not technical. The data in this paper show a viable path. The challenge is that any significant change to a system moving tens of billions of dollars requires evidence it works before anyone will accept the risk of transition.

A parallel pilot is the most credible path forward: run the simplified approach alongside the current one, generate real-world data, and let the results make the case. No existing payments are disrupted. Plans, regulators, and policymakers can evaluate performance before committing to anything. If the results hold up in practice as they do in this analysis, the policy case becomes difficult to argue against.

Health plans are an easy target. A few probably deserve real scrutiny. But the honest answer is no.

Outright fabrication of diagnoses without clinical support is fraud, and it should be treated that way. But the vast majority of what gets labeled upcoding is legal. Health plans have built teams, software, and clinical programs to find the codes the current rules reward. That is not a moral failure. It is a rational response to incentives that regulators created and have maintained for years.

Health plans have real obligations: to their members, to financial stability, to the people whose coverage depends on the organization staying solvent. Expecting them to voluntarily leave legitimate revenue on the table is wishful thinking. You cannot design a system that rewards a behavior and then blame the participants for doing it. Fix the rules, and the behavior follows.

Risk adjustment payments are large enough that every plan has reason to invest heavily in optimizing them. That is rational behavior under the current rules.

But everyone else is doing the same thing. When every plan in a market builds sophisticated coding operations, whatever competitive edge one plan got from doing it first vanishes. You spend more to stand still.

That spending is not trivial. Actuarial and data science teams that could be designing better benefits or pricing products more creatively are instead running code gap analyses. Clinical teams are optimizing documentation rather than care delivery. Compliance staff are fighting audits. These are some of a plan's most talented people, and the current system funnels them into a regulatory accounting exercise disconnected from member experience.

What happens if the system gets simpler? The coding advantage disappears, but so does the cost of chasing it. Plans compete on what they do for members, not on how well they document diagnoses. The resources that go into risk adjustment compliance go back into the things that actually differentiate a plan in the market.

That is a better deal for plans and for members. Not because anyone is gaming the system today, but because the system itself imposes a cost on everyone that no one benefits from.

Any system can be gamed if designed carelessly. The question is whether it has to be.

Self-reported data creates a gaming risk only if plans control how the survey reaches members. A centrally administered survey, where the regulator owns the process end to end, removes that lever. Plans cannot coach responses, target selectively, or incentivize members to answer in a particular way. Marketing rules in Medicare Advantage already prevent plans from communicating with members in ways that select on health status. The same regulatory framework extends naturally to survey administration. The playing field is flat by design.

Some members will misreport. Response rates will vary across plans. Those are real concerns, but they do not compare to what the current system already contends with: widespread coding errors alongside deliberate upcoding, and a costly audit process that samples a small fraction of members and extrapolates findings to entire populations. That is a blunt instrument generating its own overhead and disputes, and it still does not fully solve the problem. The survey approach has a smaller, more contained version of the same challenge, and avoids that entire class of administrative burden.

Other Research and Journalism

Links are not necessarily endorsements of views.