Ramie Fathy, MD· Mass General BrighamBoston
$open dermoscopy-llm-eval
/research/dermoscopy-llm-dashboard

100 images. 17 models.
10,200 repeated evaluations.

Aggregate evaluation,
multimodal LLMs on
dermoscopy.
JAAD 2025 (in press)
Question · Method · Limit

How do image-classification results differ across 17 models and 6 prompting arms on the same 100 dermoscopic images? This fixed benchmark does not establish equivalence to clinicians, patient-level accuracy, or improvement over time.

Dataset spans 8 dermatologic diagnoses tested across 17 multimodal LLMs from OpenAI and Google. 6 prompting strategies cover zero-shot label list, few-shot exemplars, primer board, anti-anchoring, and free-form variants.

Back to research →Browse apps →
MethodAggregate accuracy, sensitivity,
and specificity across all
model × prompt × diagnosis
combinations.CitationTadros AR, Zhuo W, Fathy RA,
et al. JAAD 2025 (in press).DisclaimerResearch dashboard. Not medical
advice. Not for clinical use.

Dermoscopy LLM Evaluation Dashboard

Filter models, compare prompting arms, and explore cost vs accuracy tradeoffs.

OpenAIGemini
Loading dashboard…
Dataset: Dermoscopy LLM evaluation summary.Total trials: –