Consumer AI as Health Plan Decision Support
A benchmark of LLM plan recommendations against a claims-based pricing engine
We benchmarked the July 2026 consumer AI lineup from OpenAI and Anthropic on a realistic health plan selection task: five ACA-compliant plans, twelve claims-anchored member profiles, and four progressively richer levels of health information, 1,680 benchmarked AI responses in all.
AI optimizes on expected annual cost
Across current frontier models, 80–86% of top recommendations are justified as minimizing expected annual cost. Risk protection is the runner-up’s job: switch the view below — GPT-5.5 cites downside protection for its second-choice plan in 48% of responses, but for its top pick in only 9%.
| Model | Expected annual cost | Usability / cost certainty | Risk protection | Other / mixed |
|---|---|---|---|---|
| GPT-5.5 | 86% | 4% | 9% | 1% |
| GPT-5.4-mini | 81% | 9% | 10% | 0% |
| Claude Sonnet 5 | 80% | 6% | 10% | 3% |
| Claude Haiku 4.5 | 65% | 22% | 11% | 2% |
Rows may not total 100% due to rounding.
Frontier models are better — and reasoning does the work
Correct recommendations improve from 35% for the previous generation to roughly 53% for today’s frontier model. Reasoning effort matters as much as model family: the same GPT-5.4-mini drops from 52% to 33% correct when reasoning is disabled — and its self-consistency, whether the pick matches its own cost math, falls from 90% to 60%.
| Model | % correct | Mean regret $ | Self-consistency |
|---|---|---|---|
| GPT-5.5 | 53% | $1,261 | 87% |
| GPT-5.4-mini (high effort) | 52% | $1,302 | 90% |
| Claude Sonnet 5 | 44% | $1,822 | 80% |
| Claude Haiku 4.5 | 43% | $1,664 | 64% |
| GPT-5.4-mini (no reasoning) (same model as above, reasoning disabled) | 33% | $2,005 | 60% |
| GPT-5.4 (previous generation, March 2026) | 35% | $1,842 | 37% |
Correct = recommended plan within $250/yr of the benchmark-optimal plan.
More information helps — until the data runs out
Telling the AI about conditions and medications lifts correct recommendations to 60% across four models — the peak of the ladder. Adding more detailed clinical information produces no further gains and, in some cases, reduces accuracy, while supplying market-level prices narrows most of the accuracy gap: 85% correct for Claude Sonnet 5 and 95% for GPT-5.5.
| Series | Demographics only (T-0) | Expected utilization category (T-hcgov) | Conditions and medications (T-1) | Detailed clinical information (T-1.5) | + Market prices (5 intervention profiles) |
|---|---|---|---|---|---|
| GPT-5.5 | 56% | 50% | 65% | 40% | 95% |
| GPT-5.4-mini | 29% | 52% | 75% | 50% | not tested |
| Claude Sonnet 5 | 35% | 48% | 54% | 40% | 85% |
| Claude Haiku 4.5 | 44% | 33% | 46% | 48% | not tested |
| All four models | 41% | 46% | 60% | 44% | not tested |
*Illustrative: the market-price intervention on the five hardest (medium- and high-utilization) profiles. On those profiles, correct recommendations rise from 25% to 85% (Claude Sonnet 5) and from 50% to 95% (GPT-5.5) with market-level prices.
More detail can backfire for the regular-care cohort
The change in accuracy isn’t one-directional. Healthy members hold steady at 71% and intensive-needs members dip only modestly — but regular-care members collapse from 72% to 28% when detailed clinical information is added. For this cohort, modest out-of-pocket differences — not premiums — decide the optimal plan, and models increasingly underestimate those costs.
| Cohort | Correct, conditions and medications (T-1) | Correct, detailed clinical information (T-1.5) | Mean error (T-1) | Mean error (T-1.5) |
|---|---|---|---|---|
| Healthy | 71% | 71% | +$34 | −$457 |
| Regular care | 72% | 28% | −$503 | −$1,414 |
| Intensive needs | 44% | 41% | +$146 | +$481 |
○ Conditions and medications (T-1) → ● Detailed clinical information (T-1.5)
The gap is a data problem and it can be narrowed
Unlike traditional enrollment tools, AI collects rich member context through natural conversation — conditions, medications, expected care. This study shows conversational intake works, and that the largest remaining gains come from pairing AI with structured insurance data.
A price book for common healthcare services raises correct recommendations to 95%. Structured allowed-amount data is the highest-leverage input.
Drug-level formulary detail and unambiguous cost-sharing rules further reduce estimation errors.
Consistent, machine-readable plan documents make AI recommendations more accurate and transparent.
Human judgment still matters when plan costs depend on information that is missing, ambiguous, or hard to verify. But the question is no longer whether AI becomes part of plan selection — it’s how to integrate it so recommendations are accurate, transparent, and trustworthy.
Suggested citation: Dai, K. (2026). Consumer AI as Health Plan Decision Support: A benchmark of LLM plan recommendations against a claims-based pricing engine. https://doi.org/10.5281/zenodo.21813452
Also distributed by SSRN.
Licensed under CC BY 4.0.
← Back to all researchEverything behind these findings
Complete benchmark methodology, model-by-model results, intervention experiments, robustness checks, and limitations.