How do image-classification results differ across 17 models and 6 prompting arms on the same 100 dermoscopic images? This fixed benchmark does not establish equivalence to clinicians, patient-level accuracy, or improvement over time.
Dataset spans 8 dermatologic diagnoses tested across 17 multimodal LLMs from OpenAI and Google. 6 prompting strategies cover zero-shot label list, few-shot exemplars, primer board, anti-anchoring, and free-form variants.