Artificial intelligence is becoming a routine feature of health decision-making. A 2026 Gallup poll found that AI now plays some role in health decisions for nearly half of Americans, with over a third using it to research symptoms before or after a clinical visit. Platforms have taken note. Health apps, telemedicine services and online pharmacies are increasingly fielding the same question: which AI health tool should we integrate, and how much should we trust it?
A new study presented at the Association for Computing Machinery’s 2026 FAccT conference offers a useful anchor for that question. Led by researchers at Penn State University, it evaluated how accurately large language model chatbots, including ChatGPT, responded to health questions assessed by nine board-certified physicians. The overall accuracy rate was around 76%. That figure invites cautious optimism. It also conceals a significant problem.
Why the average chatbot accuracy figure misleads
Averaged across all health domains, the chatbots’ 76% accuracy sounds reasonable for a supplementary tool. But the Penn State researchers identified clear performance gaps across specialties. Dermatology, along with mental health and internal medicine, fell toward the bottom of the accuracy range.
For platforms, this matters more than the headline number. An overall average can create a false impression of reliability across every use case while hiding real weaknesses in specific clinical areas. Specialty-level evidence is the figure to ask for.
Why general health questions suit text-based AI
Large language models perform well where the diagnostic process maps cleanly onto language. Common illnesses, general medical questions and routine health inquiries follow patterns that LLMs can match against enormous volumes of text-based training data. A question about whether certain symptoms suggest a respiratory infection, or what a given medication does, fits squarely within what a well-trained language model can handle. The reasoning is text-in, text-out.
Why dermatology breaks the text-based pattern
Dermatology does not follow that pattern. The Penn State researchers attributed the lower performance in skin conditions to a specific technical reason: dermatology questions frequently require image analysis, and that is an area where text-based AI systems remain considerably less capable. The point is not that AI is bad at dermatology in general. It is that the wrong category of AI is bad at it.
When a patient describes a lesion in words, something is already lost. Colour, texture, border irregularity, location relative to surrounding tissue, the way light catches the surface: none of these translate cleanly into symptom descriptions. A clinician examining a skin condition is not primarily processing language. They are processing a visual field with trained pattern recognition built across thousands of cases. Asking a language model to replicate that through text alone is the wrong tool for the problem.
Image-based AI maps colour, texture and contour directly from the photograph
Real-world evidence confirms the text-based gap
The Penn State finding is not isolated. A 2025 post-market clinical follow-up study of a CE-marked clinical decision support system, integrated into a major private healthcare network in Portugal, reached a closely related conclusion through a different route.
What the clinical data showed
The study, known as ESSENCE, evaluated how well the system’s condition suggestions matched the physician’s most likely diagnosis across multiple specialties. Across most areas, the correspondence was strong, with physicians’ most probable diagnosis included among the AI’s suggestions in the majority of cases. Dermatology was the clear exception. The authors noted this finding directly: dermatological diagnosis depends primarily on visual pattern recognition rather than symptom description, and a text-based system is inherently limited in that context. The gap was not a quality problem with that particular system. It reflected the structural ceiling of the approach.
The consistency across two independent sources
Two distinct studies, one a controlled evaluation of general-purpose LLMs and one a real-world PMCF study of a regulated medical device, arrived at the same conclusion by different methods. Text-based AI for dermatology has a ceiling that cannot be engineered away through better prompting or larger training sets. The ceiling is architectural.
Chatbot or image-based AI: the choice platforms face
For a product team evaluating how to add AI dermatology capability to a telemedicine service, pharmacy app or consumer health platform, this distinction has direct practical implications. The choice is not between a good AI dermatology tool and a bad one. It is between two fundamentally different categories of system.
The architecture question
A general-purpose health chatbot processes text. A purpose-built dermatology AI processes images. These are different technical problems with different data requirements, different validation approaches and, under EU regulations, different classification requirements. Under the EU Medical Devices Regulation, image-processing tools for dermatological analysis are classified at a minimum as Class IIa medical devices, reflecting the higher risk profile of image-based clinical decision support. A text-based health chatbot and an image-analysis dermatology API are not interchangeable, even if both carry AI branding.
The integration question
Platforms evaluating AI dermatology features should be asking a specific set of questions. Has the system been validated on image input, not text descriptions? What does the clinical evidence look like in real-world dermatology use cases, not averaged across all health domains? What is the regulatory pathway, and does it reflect the actual risk class of image-based analysis? How does the system handle the handoff from AI output to clinical review? The last question matters more than it might appear. AI outputs in this context are informational condition suggestions. Clinical responsibility remains with the professional reviewing the case. The architecture of that review gate, and how it is documented and maintained, is part of what regulators are evaluating.
What the evidence says about image-based dermatology AI
There is a growing body of clinical evidence for AI systems built specifically around image analysis in dermatology. This evidence looks different from general health AI benchmarks.
Suggestion accuracy as the right measure
The appropriate metric is suggestion accuracy: whether the correct condition appears within the AI’s ranked output, typically evaluated across a defined number of suggestions. This is meaningfully different from asking whether a chatbot correctly answers a health question. In a real-world comparative evaluation across 91 confirmed cases, Autoderm’s image-based API recorded a top-5 suggestion accuracy of 93% and a treatment pathway accuracy of 95%. A separate peer-reviewed primary care study found that 92% of GPs using the system in real-world consultations rated it useful, with a 34% reduction in specialist referrals. These figures come from validated image-analysis infrastructure, not from text-based symptom assessment.
What real-world deployment adds
Real-world deployment data matters alongside controlled evaluations. A post-market clinical follow-up study involving over a thousand patients in an active telemedicine deployment recorded a top-5 recall rate of approximately 75%, reflecting performance under everyday conditions rather than curated test sets. The gap between controlled and real-world figures is expected and instructive. It reflects the variability of user-submitted images, the diversity of presentation, and the conditions under which platforms actually operate. Because an image-based system reads the photograph itself rather than a description of it, how image quality shapes suggestion accuracy becomes a variable platforms can influence directly, through the capture experience they design. Both numbers matter. Together they give a more complete picture than either alone.
Questions worth asking before you build
If your platform is evaluating a skin AI integration, four questions will clarify the decision faster than any feature comparison.
First: does the system process images or text? If it processes text descriptions of symptoms, the Penn State finding applies directly. Dermatology accuracy will be capped by the underlying architecture regardless of the vendor’s marketing language.
Second: what is the CE classification, and does it match the actual use case? Image-based AI for dermatological triage carries a different regulatory profile from a general health chatbot. Platforms integrating AI features inherit responsibility for ensuring the tool is appropriately classified for its intended use.
Third: what does the clinical evidence cover? Published peer-reviewed studies and real-world deployment data are meaningfully different from benchmark tests. Both matter, and the clinical domains covered by the evidence should match the use case.
Fourth: how does the AI output reach a clinician? AI in this context provides informational condition suggestions. That output should support a professional review step, not replace it. The architecture of that handoff is both a clinical and a regulatory question.
The Penn State finding is a useful corrective to overclaiming. It shows that text-based AI has a real and measurable role in health information. It also shows, clearly, where that role ends. For platforms serious about adding dermatology capability, the practical implication is straightforward: the tool needs to work the way diagnosis actually works in this specialty. That means image first.
Autoderm provides informational condition suggestions based on image analysis. All outputs are reviewed by a qualified clinician before any clinical decision is made. Autoderm does not provide diagnoses, medical advice, or treatment recommendations.
References
- Mingole B, Majumdar A, Choudhury FA, et al. Dr. GPT will see you now, but should it? Exploring the benefits and harms of large language models in medical diagnosis using crowdsourced clinical cases [preprint]. arXiv. 2026. DOI: 10.48550/arXiv.2506.13805. Presented at ACM FAccT 2026, Montreal.
- Pimenta A, Kini N, Cotte F, et al. Appropriateness and utility of a clinical decision support system at the digital front door [preprint]. Research Square. January 8, 2026. DOI: 10.21203/rs.3.rs-8157860/v1
- Escalé-Besa A, Yélamos O, Vidal-Alaball J, et al. Exploring the potential of artificial intelligence in improving skin lesion diagnosis in primary care. Sci Rep. 2023;13:4293. DOI: 10.1038/s41598-023-31340-1