Research/

Consumer AI as Health Plan Decision Support

A benchmark of LLM plan recommendations against a claims-based pricing engine

We benchmarked the July 2026 consumer AI lineup from OpenAI and Anthropic on a realistic health plan selection task: five ACA-compliant plans, twelve claims-anchored member profiles, and four progressively richer levels of health information. Across 1,564 responses, frontier models recommended the benchmark-optimal plan in roughly half of cases and performed much better when additional health details and market prices were provided.

KD
Khris Dai, FSA
Founder & CEO, Visuary AI
Read Full Report
Finding 01 — What AI optimizes for

AI optimizes on expected annual cost

Across current frontier models, 80–86% of top recommendations are justified as minimizing expected annual cost. Risk protection is the runner-up’s job: switch the view below — GPT-5.5 cites downside protection for its second-choice plan in 48% of responses, but for its top pick in only 9%.

Stated rationale for the recommended plan
Share of each model's responses citing that rationale for the given rank.
Stated rationale for the recommended plan. Share of each model's responses citing each rationale.
ModelExpected annual costUsability / cost certaintyRisk protectionOther / mixed
GPT-5.586%4%9%1%
GPT-5.4-mini81%9%10%0%
Claude Sonnet 580%6%10%3%
Claude Haiku 4.565%22%11%2%

Rows may not total 100% due to rounding.

Finding 02 — Model generations

Frontier models are better — and reasoning does the work

Correct recommendations improve from 36% for the previous generation to 57% for today’s frontier model. Reasoning effort matters as much as model family: the same GPT-5.4-mini drops from 56% to 33% correct when reasoning is disabled — and its self-consistency, whether the pick matches its own cost math, falls from 90% to 60%.

Correct recommendations, by model
Higher is better · share of responses recommending the benchmark-optimal plan
Model leaderboard: correct recommendations, mean regret, and self-consistency, for each model serving.
Model% correctMean regret $Self-consistency
GPT-5.557%$83687%
GPT-5.4-mini (high effort)56%$94090%
Claude Sonnet 547%$1,32680%
Claude Haiku 4.540%$1,40864%
GPT-5.4-mini (no reasoning) (same model as above, reasoning disabled)33%$1,78060%
GPT-5.4 (previous generation, March 2026)36%$1,89036%

Correct = recommended plan within $250/yr of the benchmark-optimal plan.

Finding 03 — The information ladder

AI gets better when consumers share their health information

Telling the AI about conditions and medications lifts correct recommendations to 61% across four models, with GPT-5.5 reaching 73%. Adding more detailed clinical information produces no further gains and, in some cases, reduces accuracy, while supplying market-level prices raises GPT-5.5's accuracy to 90%.

Correct recommendations, by what the consumer disclosed
Share of responses recommending the benchmark-optimal plan. Toggle models below; hover any point for detail.
Correct recommendations by information level. Share of responses recommending the benchmark-optimal plan.
SeriesDemographics only (T-0)Expected utilization category (T-hcgov)Conditions and medications (T-1)Detailed clinical information (T-1.5)+ Market prices
GPT-5.548%58%73%48%90%
GPT-5.4-mini38%60%71%54%not tested
Claude Sonnet 529%56%54%48%73%
Claude Haiku 4.546%29%46%40%not tested
All four models40%51%61%47%not tested
0%25%50%75%100%Demographics onlyT-0Expected utilization categoryT-hcgovConditions and medicationsT-1Detailed clinical informationT-1.5+ Market pricesintervention+25–42 pts*73%90%

Finding 04 — Utilization cohorts

More detail can backfire for the regular-care cohort

The change in accuracy isn’t one-directional. Healthy members improve, from 75% to 83%, and intensive-needs members dip only modestly — but regular-care members collapse from 72% to 28% when detailed clinical information is added. For this cohort, modest out-of-pocket differences — not premiums — decide the optimal plan, and models increasingly underestimate those costs.

Where the estimates go wrong
Mean signed out-of-pocket estimation error by cohort — negative means the model estimated below the member's actual cost.
Correct-recommendation rate and mean signed out-of-pocket estimation error by member utilization cohort, at the two highest information levels. Negative error means the model estimated below the member's actual cost.
CohortCorrect, conditions and medications (T-1)Correct, detailed clinical information (T-1.5)Mean error (T-1)Mean error (T-1.5)
Healthy75%83%+$687+$196
Regular care72%28%+$586−$325
Intensive needs44%41%+$47+$383

Open circle: conditions and medications (T-1). Arrow tip: detailed clinical information (T-1.5).

Finding 05 — Stability

Stable answers are not the same as correct ones

Ask the same question twice and you may not get the same answer. GPT-5.5’s four responses disagree in 29% of scenarios, rising to 83% for Claude Haiku 4.5. Stability is not accuracy, though, and GPT-5.5 shows it most sharply: when it recommends the wrong plan, all four of its responses agree on that same wrong plan 62% of the time.

Agreement across four repeated responses
Share of scenarios by whether the four responses agree and whether the pick is correct.
Agreement across four repeated responses, by model. Share of scenarios by whether the four responses agree and whether the pick is correct. Counts are scenarios.
ModelNot unanimousUnanimous and correctSplit, correctSplit, wrongUnanimous and wrongShare of wrong-majority scenarios that were unanimous
GPT-5.529%21681362%
GPT-5.4-mini (high effort)58%17919314%
Claude Sonnet 556%15819624%
Claude Haiku 4.583%41723415%

Unanimous means all four responses named the same plan. Correct means the plan named by a majority of the four came within $250/yr of the benchmark-optimal plan for that member. Errors repeated is the share of a model’s wrong-majority scenarios that were unanimous. Rows may not total 100% due to rounding.

What this means

AI could change how consumers shop for health insurance

Unlike traditional enrollment tools, AI can collect rich member context through natural conversation — conditions, medications, and expected care. Both sides of the equation matter: consumer information improves recommendations, while better insurance information allows the model to make better use of it.

Consumer choice

AI gives consumers a new way to compare plans. A consumer can describe conditions, medications, and expected care in conversation and get recommendations tailored to their circumstances.

Consumer protection

AI advice creates new risks for consumers. The models can be inconsistent or confidently wrong, and their recommendations can be affected by inaccurate price information or ambiguous plan documents.

Risk pools

Better-informed plan selection could affect who chooses which plans. If AI materially reduces the information barrier in plan choice, it could change selection patterns across plans, an issue for insurers and actuaries as well as policymakers.

The technology still has limitations. But it is evolving quickly from a tool that answers insurance questions into one that can help consumers make insurance decisions. And this year's open enrollment may mark a new phase in health insurance — one in which AI becomes a practical part of how consumers choose their health plans.

Suggested citation: Dai, K. (2026). Consumer AI as Health Plan Decision Support: A benchmark of LLM plan recommendations against a claims-based pricing engine. https://doi.org/10.2139/ssrn.7235019

Also available on Zenodo.

Licensed under CC BY 4.0.

← Back to all research

Everything behind these findings

Complete benchmark methodology, model-by-model results, intervention experiments, robustness checks, and limitations.