Research/

Consumer AI as Health Plan Decision Support

A benchmark of LLM plan recommendations against a claims-based pricing engine

We benchmarked the July 2026 consumer AI lineup from OpenAI and Anthropic on a realistic health plan selection task: five ACA-compliant plans, twelve claims-anchored member profiles, and four progressively richer levels of health information, 1,680 benchmarked AI responses in all.

KD
Khris Dai, FSA
Founder & CEO, Visuary AI
Read Full Report
Finding 01 — What AI optimizes for

AI optimizes on expected annual cost

Across current frontier models, 80–86% of top recommendations are justified as minimizing expected annual cost. Risk protection is the runner-up’s job: switch the view below — GPT-5.5 cites downside protection for its second-choice plan in 48% of responses, but for its top pick in only 9%.

Stated rationale for the recommended plan
Share of each model's responses citing that rationale for the given rank.
Stated rationale for the recommended plan. Share of each model's responses citing each rationale.
ModelExpected annual costUsability / cost certaintyRisk protectionOther / mixed
GPT-5.586%4%9%1%
GPT-5.4-mini81%9%10%0%
Claude Sonnet 580%6%10%3%
Claude Haiku 4.565%22%11%2%

Rows may not total 100% due to rounding.

Finding 02 — Model generations

Frontier models are better — and reasoning does the work

Correct recommendations improve from 35% for the previous generation to roughly 53% for today’s frontier model. Reasoning effort matters as much as model family: the same GPT-5.4-mini drops from 52% to 33% correct when reasoning is disabled — and its self-consistency, whether the pick matches its own cost math, falls from 90% to 60%.

Correct recommendations, by model
Higher is better · share of responses recommending the benchmark-optimal plan
Model leaderboard: correct recommendations, mean regret, and self-consistency, for each model serving.
Model% correctMean regret $Self-consistency
GPT-5.553%$1,26187%
GPT-5.4-mini (high effort)52%$1,30290%
Claude Sonnet 544%$1,82280%
Claude Haiku 4.543%$1,66464%
GPT-5.4-mini (no reasoning) (same model as above, reasoning disabled)33%$2,00560%
GPT-5.4 (previous generation, March 2026)35%$1,84237%

Correct = recommended plan within $250/yr of the benchmark-optimal plan.

Finding 03 — The information ladder

More information helps — until the data runs out

Telling the AI about conditions and medications lifts correct recommendations to 60% across four models — the peak of the ladder. Adding more detailed clinical information produces no further gains and, in some cases, reduces accuracy, while supplying market-level prices narrows most of the accuracy gap: 85% correct for Claude Sonnet 5 and 95% for GPT-5.5.

Correct recommendations, by what the consumer disclosed
Share of responses recommending the benchmark-optimal plan. Toggle models below; hover any point for detail.
Correct recommendations by information level. Share of responses recommending the benchmark-optimal plan.
SeriesDemographics only (T-0)Expected utilization category (T-hcgov)Conditions and medications (T-1)Detailed clinical information (T-1.5)+ Market prices (5 intervention profiles)
GPT-5.556%50%65%40%95%
GPT-5.4-mini29%52%75%50%not tested
Claude Sonnet 535%48%54%40%85%
Claude Haiku 4.544%33%46%48%not tested
All four models41%46%60%44%not tested
0%25%50%75%100%Demographics onlyT-0Expected utilization categoryT-hcgovConditions and medicationsT-1Detailed clinical informationT-1.5+ Market pricesintervention · 5 profiles+45–55 ptson the 5 intervention profiles*85%95%

*Illustrative: the market-price intervention on the five hardest (medium- and high-utilization) profiles. On those profiles, correct recommendations rise from 25% to 85% (Claude Sonnet 5) and from 50% to 95% (GPT-5.5) with market-level prices.

Finding 04 — Utilization cohorts

More detail can backfire for the regular-care cohort

The change in accuracy isn’t one-directional. Healthy members hold steady at 71% and intensive-needs members dip only modestly — but regular-care members collapse from 72% to 28% when detailed clinical information is added. For this cohort, modest out-of-pocket differences — not premiums — decide the optimal plan, and models increasingly underestimate those costs.

Where the estimates go wrong
Mean signed out-of-pocket estimation error by cohort — negative means the model estimated below the member's actual cost.
Correct-recommendation rate and mean signed out-of-pocket estimation error by member utilization cohort, at the two highest information levels. Negative error means the model estimated below the member's actual cost.
CohortCorrect, conditions and medications (T-1)Correct, detailed clinical information (T-1.5)Mean error (T-1)Mean error (T-1.5)
Healthy71%71%+$34−$457
Regular care72%28%−$503−$1,414
Intensive needs44%41%+$146+$481

○ Conditions and medications (T-1) → ● Detailed clinical information (T-1.5)

What this means

The gap is a data problem and it can be narrowed

Unlike traditional enrollment tools, AI collects rich member context through natural conversation — conditions, medications, expected care. This study shows conversational intake works, and that the largest remaining gains come from pairing AI with structured insurance data.

Market-level prices

A price book for common healthcare services raises correct recommendations to 95%. Structured allowed-amount data is the highest-leverage input.

Formularies & benefit rules

Drug-level formulary detail and unambiguous cost-sharing rules further reduce estimation errors.

Clear plan documents

Consistent, machine-readable plan documents make AI recommendations more accurate and transparent.

Human judgment still matters when plan costs depend on information that is missing, ambiguous, or hard to verify. But the question is no longer whether AI becomes part of plan selection — it’s how to integrate it so recommendations are accurate, transparent, and trustworthy.

Suggested citation: Dai, K. (2026). Consumer AI as Health Plan Decision Support: A benchmark of LLM plan recommendations against a claims-based pricing engine. https://doi.org/10.5281/zenodo.21813452

Also distributed by SSRN.

Licensed under CC BY 4.0.

← Back to all research

Everything behind these findings

Complete benchmark methodology, model-by-model results, intervention experiments, robustness checks, and limitations.