A dermatology AI is only as reliable as the image it receives. The most common source of integration disappointment is not the algorithm. It is the photograph, and platform builders who understand that early design better capture journeys and get stronger output quality at scale.
This article looks at how Autoderm’s model reads a skin image, what degrades its output, and what that means when you are designing the capture experience your users will follow. Image-based analysis is the right tool for dermatology precisely because the specialty resists text-based AI, a limitation we cover separately in why AI chatbots struggle with skin conditions. Once the input is an image, the image itself becomes the variable.
What the model is actually analysing
How convolutional neural networks read a skin image
Autoderm’s algorithm is built on a convolutional neural network (CNN), a class of model trained to recognise patterns in pixel data. During training, the network works through hundreds of thousands of annotated smartphone images and learns which visual features point to which conditions: texture, colour distribution, lesion borders, surrounding skin tone, and how those elements sit in relation to one another.
The V2.3.0 model has been trained on over 150,000 dermatologist-annotated images drawn from a proprietary dataset of more than three million real-world smartphone photographs, with new images added at a rate of approximately 100,000 per month. Each image in the training set required agreement from at least two of three independent board-certified dermatologists before annotation was accepted.
In practice, the model recognises skin conditions in images that resemble the ones it learned from. A submission that departs from that pattern, whether out of focus, cluttered, or oddly framed, leaves the network working with less relevant signal. The output shows it.
Why the model ranks, rather than selects
Autoderm returns a ranked list of possible skin conditions rather than a single output. This is by design. The ranking reflects the distribution of visual similarity scores across all conditions in the model’s scope: the condition ranked first is the one whose learned visual pattern most closely matches the submitted image, and the rest follow in order of similarity.
This is where image quality does its work. A well-captured, focused image of a single lesion concentrates the model’s pattern-matching on the relevant features. A cluttered or poorly framed image distributes that signal across multiple competing visual patterns, which can push the correct condition lower in the ranking or weight an anatomically irrelevant structure more heavily than the lesion itself.
The most common causes of poor output
Competing features in the frame
The model does not know which element in a photograph is the one the user is concerned about. It analyses the entire image and ranks conditions based on the totality of what it sees. A photograph of a suspicious mole that also includes a large expanse of hair, a belly button, previous scarring, and several unrelated skin markings gives the model a large number of competing visual features to weigh.
The result is that the model may correctly identify several of the conditions visible in the frame, while the condition of actual clinical interest ranks lower than it should. The output is not wrong, precisely, but it is not targeted at what the user intended to submit.
Anatomical ambiguity
CNNs identify visual similarity, not anatomical intent. Certain body parts, viewed from specific angles and at certain distances, share enough visual characteristics with other body parts that the model will classify them similarly. A single finger photographed end-on, for example, shares enough visual structure with certain other appendages that the model’s ranking may reflect that ambiguity. That is not a flaw in the algorithm. It is simply how visual similarity models behave, and better framing resolves it.
The same principle applies at condition level. Conditions that look alike, in colour distribution, border shape, or surface texture, will compete in the ranking. When the submitted image is well-framed and clearly focused, the distinguishing features between similar conditions become more visible to the model and rankings become more reliable.
Lighting and focus
Colour carries real weight in skin condition classification. The distinction between erythema and pigmentation, or between a seborrhoeic keratosis and an atypical naevus, involves colour information that can be distorted by artificial lighting, shadow, overexposure, or motion blur. Poor natural lighting flattens the tonal variation the model relies on to distinguish conditions that look similar in shape but differ in colour presentation.
Strong directional light blows out the raised red lesion and washes the colour out of the mole beside it.
Focus matters for the same reason. Many conditions are told apart by surface texture, and a blurred image removes exactly that information. Where texture is the distinguishing feature, blur degrades the ranking directly.
Close and reasonably lit, but soft focus and an angled camera soften both lesion borders.
What a well-captured image looks like
Framing and crop
The single change that helps most is cropping tightly to the area of concern before submission, or photographing only that area. The image should contain the lesion or affected area, a small margin of surrounding skin for context, and as little else as possible.
If a user has multiple concerns, each concern should be submitted as a separate image. A photograph intended to capture a suspicious mole on the trunk that also shows hair, navel, and adjacent skin markings will produce less targeted output than a photograph cropped or repositioned to show only the mole and the skin immediately around it.
Lighting conditions
Natural daylight, indirect rather than direct, produces the most accurate colour reproduction. Direct sunlight creates harsh shadows and overexposed highlights that distort surface colour. Artificial indoor lighting, particularly tungsten or warm-spectrum sources, shifts colour temperature in ways that affect how the model reads pigmentation. If natural light is not available, a ring light or diffuse flash at a slight angle reduces shadow without overexposing the lesion surface.
The camera should be held perpendicular to the skin surface. Angled shots introduce perspective distortion and can create shadows across the lesion that partially obscure the features the model uses for classification.
What to exclude from the frame
Jewellery, clothing edges, pointing fingers, drawn markings, adhesive dressings, and any other non-skin element in the frame should be removed or repositioned before capture. Each of these introduces visual features that have no representation in the training data and that the model will attempt to classify anyway, adding noise to the output.
This applies to scale markers as well. Users sometimes place a coin or ruler next to a lesion to indicate size. From the model’s perspective, this is an unrecognised foreign object sharing the frame with the lesion. It does not assist classification and may interfere with it.
The same two lesions, framed three ways
Everything above is easier to see in a single worked example. The subject is one limb carrying two adjacent lesions: a brown pigmented mole and a raised red lesion beside it. Nothing about the skin changes across the photographs that follow. Only the camera does.
The photographs below were submitted to Autoderm and the rankings shown are what the model actually returned for each one.
✗ Too far from the skin
Both lesions are present, but each occupies a small fraction of the frame.
The camera is well back from the limb. The two lesions sit small in the middle of the frame while most of the image is unaffected skin, body hair, the edge of the limb and the room behind it. This is what the model returned:
| Rank | Condition suggestion | Confidence distribution |
| 1 | Nevus (benign mole) | 28.92% |
| 2 | Atypical melanocytic lesion | 26.70% |
| 3 | Dermal nevus | 21.95% |
| 4 | Seborrhoeic keratosis | 14.83% |
| 5 | Malignant melanoma | 7.61% |
The lesions are simply too small to contribute much signal. The model analyses everything it is given, and here most of what it is given is background. The output spreads thinly across several pigmented-lesion categories, and the raised red lesion does not appear anywhere in the top five. In effect, the model never saw it.
✗ Sharp, close, well lit, and still two conditions in one frame
Technically the best photograph in the set, yet both lesions sit inside a single frame.
Nothing is wrong with this photograph as a photograph. It is close, in focus, evenly lit and steady. The problem is what is in the frame, not how it was captured, and that alone produced the least reliable output of the set:
| Rank | Condition suggestion | Confidence distribution |
| 1 | Malignant melanoma | 48.29% |
| 2 | Atypical melanocytic lesion | 23.08% |
| 3 | Nevus (benign mole) | 10.84% |
| 4 | Basal cell carcinoma | 9.84% |
| 5 | Dermatofibroma | 7.94% |
The failure here is framing, not photography. A brown pigmented area sitting immediately beside a red-pink area is, to a visual similarity model, a single lesion with asymmetry, colour variegation and an irregular border. That is the classic melanoma pattern, and the model matched it. The pattern only existed because two separate conditions were framed as one.
This is the most important point in the example. Sharpness did not help. Lighting did not help. Getting closer actively made things worse, because it resolved two lesions clearly enough for the model to read them as one.
✓ What a correct submission looks like
Each lesion belongs in its own image, cropped so that it and a small margin of surrounding skin are all that is in the frame:
The mole on its own
| Rank | Condition suggestion | Confidence distribution |
| 1 | Atypical Melanocytic Lesion (Atypical Nevus) | 31.1% |
| 2 | Nevus (Benign Mole) | 24.4% |
| 3 | Seborrhoeic Keratosis | 20.9% |
| 4 | Lentigo | 16.0% |
| 5 | Malignant Melanoma | 7.6% |
| Rank | Condition suggestion | Confidence distribution |
| 1 | Furuncle (Deep Folliculitis) | 56.7% |
| 2 | Folliculitis | 16.6% |
| 3 | Dermatofibroma | 11.1% |
| 4 | Pseudofolliculitis Barbae | 8.8% |
| 5 | Dermal Nevus | 6.8% |
Submitted this way, each image contains a single condition. The model is no longer weighing two lesions against one another, and no longer has the opportunity to read them as a single irregular lesion. Each concern gets its own assessment, which is what the user intended in the first place.
The lesson is not simply to get closer, and it is not simply to shoot in better light. The distant photograph produced a safer output than the sharp one, precisely because the two lesions were never resolved clearly enough to merge. The lesson is one concern per image.
What this means for platform integration
Building image guidance into the user journey
For platform builders, the point is simple: image quality is a user experience problem before it is an algorithm problem. Users who submit poor images are rarely careless. They just do not know what a good image looks like in this context. Guidance built into the capture flow, before the image is submitted, does more for output quality than anything that happens after.
That guidance does not need to be extensive. A short instruction set at the point of capture, covering crop, lighting, and what to keep out of the frame, deals with most quality issues. Some integrations surface a preview step that prompts the user to confirm the lesion is clearly visible and centred before submission. Others use image quality pre-screening to flag likely blur or low-light conditions before the image reaches the API.
A post-market clinical follow-up study covering 1,092 real-world clinical observations across Sweden, Norway, Finland and the United Kingdom found a real-world top-5 recall of approximately 75%. That figure represents open-set clinical conditions where the correct diagnosis may fall outside the conditions evaluated. The gap between that figure and the 93% top-5 suggestion accuracy recorded in the Coachella Study (white paper, V2.2, n=91, board-certified dermatologist reference standard) reflects, in part, the difference between controlled image capture and unconstrained user submissions. Platform integrations that invest in capture guidance close part of that gap.
How suggestion accuracy scales with image quality at volume
For a single image, poor capture means a ranking that is less targeted than it could be. Across thousands of submissions a month, the same problem becomes a measurable drop in output quality, and that drop reaches user satisfaction, clinician confidence in the tool, and eventually the internal ROI case for the integration itself.
At volume, even a modest lift in average image quality moves a meaningful share of outputs from loosely targeted to clinically useful. Clearer capture guidance is the cheapest way to get that lift.
This matters even more under a Gate-After deployment model, where users see the ranked condition suggestions directly, before any clinician review. That first output is the product as far as the user is concerned. Platforms that treat capture guidance as a UX priority rather than a technical footnote see the difference in how their users engage with it.
Autoderm provides informational condition suggestions with confidence levels. All outputs are intended as decision support and require professional clinical review before any clinical decision is made. Autoderm does not provide medical diagnoses, prescriptions, or treatment recommendations. Clinical responsibility remains with the healthcare professional and the platform’s defined care pathways.
References
- iDoc24 AB, Visiba Group AB. Visiba Care post-market clinical follow-up report [PMCF report]. September 2021. Autoderm V2.0; n=1,092 real-world observations across Sweden, Norway, Finland and the United Kingdom; EU MDR Annex XIV. https://autoderm.ai/clinical-evidence/
- Autoderm clinical evidence portfolio: V2.3.0 training data and deployment architecture. Autoderm; 2026. https://autoderm.ai/clinical-evidence/