Consumer AI as Health Plan Decision Support
A benchmark of LLM plan recommendations against a claims-based pricing engine
We benchmarked the July 2026 consumer AI lineup from OpenAI and Anthropic on a realistic health plan selection task: five ACA-compliant plans, twelve claims-anchored member profiles, and four progressively richer levels of health information. Across 1,564 responses, frontier models recommended the benchmark-optimal plan in roughly half of cases and performed much better when additional health details and market prices were provided.
AI optimizes on expected annual cost
Across current frontier models, 80–86% of top recommendations are justified as minimizing expected annual cost. Risk protection is the runner-up’s job: switch the view below — GPT-5.5 cites downside protection for its second-choice plan in 48% of responses, but for its top pick in only 9%.
| Model | Expected annual cost | Usability / cost certainty | Risk protection | Other / mixed |
|---|---|---|---|---|
| GPT-5.5 | 86% | 4% | 9% | 1% |
| GPT-5.4-mini | 81% | 9% | 10% | 0% |
| Claude Sonnet 5 | 80% | 6% | 10% | 3% |
| Claude Haiku 4.5 | 65% | 22% | 11% | 2% |
Rows may not total 100% due to rounding.
Frontier models are better — and reasoning does the work
Correct recommendations improve from 36% for the previous generation to 57% for today’s frontier model. Reasoning effort matters as much as model family: the same GPT-5.4-mini drops from 56% to 33% correct when reasoning is disabled — and its self-consistency, whether the pick matches its own cost math, falls from 90% to 60%.
| Model | % correct | Mean regret $ | Self-consistency |
|---|---|---|---|
| GPT-5.5 | 57% | $836 | 87% |
| GPT-5.4-mini (high effort) | 56% | $940 | 90% |
| Claude Sonnet 5 | 47% | $1,326 | 80% |
| Claude Haiku 4.5 | 40% | $1,408 | 64% |
| GPT-5.4-mini (no reasoning) (same model as above, reasoning disabled) | 33% | $1,780 | 60% |
| GPT-5.4 (previous generation, March 2026) | 36% | $1,890 | 36% |
Correct = recommended plan within $250/yr of the benchmark-optimal plan.
AI gets better when consumers share their health information
Telling the AI about conditions and medications lifts correct recommendations to 61% across four models, with GPT-5.5 reaching 73%. Adding more detailed clinical information produces no further gains and, in some cases, reduces accuracy, while supplying market-level prices raises GPT-5.5's accuracy to 90%.
| Series | Demographics only (T-0) | Expected utilization category (T-hcgov) | Conditions and medications (T-1) | Detailed clinical information (T-1.5) | + Market prices |
|---|---|---|---|---|---|
| GPT-5.5 | 48% | 58% | 73% | 48% | 90% |
| GPT-5.4-mini | 38% | 60% | 71% | 54% | not tested |
| Claude Sonnet 5 | 29% | 56% | 54% | 48% | 73% |
| Claude Haiku 4.5 | 46% | 29% | 46% | 40% | not tested |
| All four models | 40% | 51% | 61% | 47% | not tested |
More detail can backfire for the regular-care cohort
The change in accuracy isn’t one-directional. Healthy members improve, from 75% to 83%, and intensive-needs members dip only modestly — but regular-care members collapse from 72% to 28% when detailed clinical information is added. For this cohort, modest out-of-pocket differences — not premiums — decide the optimal plan, and models increasingly underestimate those costs.
| Cohort | Correct, conditions and medications (T-1) | Correct, detailed clinical information (T-1.5) | Mean error (T-1) | Mean error (T-1.5) |
|---|---|---|---|---|
| Healthy | 75% | 83% | +$687 | +$196 |
| Regular care | 72% | 28% | +$586 | −$325 |
| Intensive needs | 44% | 41% | +$47 | +$383 |
Open circle: conditions and medications (T-1). Arrow tip: detailed clinical information (T-1.5).
Stable answers are not the same as correct ones
Ask the same question twice and you may not get the same answer. GPT-5.5’s four responses disagree in 29% of scenarios, rising to 83% for Claude Haiku 4.5. Stability is not accuracy, though, and GPT-5.5 shows it most sharply: when it recommends the wrong plan, all four of its responses agree on that same wrong plan 62% of the time.
| Model | Not unanimous | Unanimous and correct | Split, correct | Split, wrong | Unanimous and wrong | Share of wrong-majority scenarios that were unanimous |
|---|---|---|---|---|---|---|
| GPT-5.5 | 29% | 21 | 6 | 8 | 13 | 62% |
| GPT-5.4-mini (high effort) | 58% | 17 | 9 | 19 | 3 | 14% |
| Claude Sonnet 5 | 56% | 15 | 8 | 19 | 6 | 24% |
| Claude Haiku 4.5 | 83% | 4 | 17 | 23 | 4 | 15% |
Unanimous means all four responses named the same plan. Correct means the plan named by a majority of the four came within $250/yr of the benchmark-optimal plan for that member. Errors repeated is the share of a model’s wrong-majority scenarios that were unanimous. Rows may not total 100% due to rounding.
AI could change how consumers shop for health insurance
Unlike traditional enrollment tools, AI can collect rich member context through natural conversation — conditions, medications, and expected care. Both sides of the equation matter: consumer information improves recommendations, while better insurance information allows the model to make better use of it.
AI gives consumers a new way to compare plans. A consumer can describe conditions, medications, and expected care in conversation and get recommendations tailored to their circumstances.
AI advice creates new risks for consumers. The models can be inconsistent or confidently wrong, and their recommendations can be affected by inaccurate price information or ambiguous plan documents.
Better-informed plan selection could affect who chooses which plans. If AI materially reduces the information barrier in plan choice, it could change selection patterns across plans, an issue for insurers and actuaries as well as policymakers.
The technology still has limitations. But it is evolving quickly from a tool that answers insurance questions into one that can help consumers make insurance decisions. And this year's open enrollment may mark a new phase in health insurance — one in which AI becomes a practical part of how consumers choose their health plans.
Suggested citation: Dai, K. (2026). Consumer AI as Health Plan Decision Support: A benchmark of LLM plan recommendations against a claims-based pricing engine. https://doi.org/10.2139/ssrn.7235019
Also available on Zenodo.
Licensed under CC BY 4.0.
← Back to all researchEverything behind these findings
Complete benchmark methodology, model-by-model results, intervention experiments, robustness checks, and limitations.