AI Dermatology Tools for Melanoma Detection in Primary Care
Early detection tools could narrow racial disparities in melanoma survival rates.

Melanoma detection lives or dies on timing. Catch it early and the five-year survival rate clears 99%. Miss it until stage IV and that number collapses to 30%. The overall five-year survival rate across all stages sits at 95%, but that average is doing a lot of hiding: it flattens the difference between a lesion caught at a routine physical and one that gets discovered after it has already spread. Stage at diagnosis is the one variable a clinician can actually move, and it happens to be the exact variable that a new wave of AI dermatology tools is aimed at.
The disparity gets sharper, and uglier, once you split it by race. Five-year survival for Black patients with melanoma sits at 70%, against 95% for white patients. That gap isn't biology doing the work; it's mostly delayed or missed diagnosis. Darker-skinned patients face melanoma mortality rates 3.75 times higher than lighter-skinned patients, even though their incidence rate is dramatically lower (0.9 per 100,000 versus 22 per 100,000 in lighter-skinned populations). Lower incidence, higher death rate: that asymmetry is not a footnote you skim past. It's a structural problem, and it's one that any new detection tool has to actually solve rather than just avoid making worse.
How large and fast-growing the melanoma burden has become
The raw numbers are not small, and they are not slowing down. Projections put new U.S. melanoma diagnoses at 234,680 for 2026, split between roughly 122,680 noninvasive cases and 112,000 invasive ones, with invasive melanoma expected to rank as the fourth most commonly diagnosed cancer in the country. Between 2015 and 2025, new invasive melanoma cases rose 42%. That's not a slow drift upward; it's a sustained, compounding climb, the kind of growth curve that makes health systems nervous for a reason.
Incidence sits at 22.3 per 100,000 (based on 2019 to 2023 cases), with a death rate of 2.0 per 100,000 (2020 to 2024). About 2.2% of Americans will be diagnosed with melanoma at some point in their life. As of 2023, an estimated 1,573,288 people in the U.S. were living with a melanoma diagnosis, past or present. Zoom out globally and the pattern holds: over 331,000 new cases and 58,000 deaths worldwide in 2022 alone. The U.S. burden isn't some strange outlier; it's one piece of a global trend.
Here's the genuinely odd part, the part worth sitting with for a second: death rates are falling while incidence keeps climbing. From 2011 to 2020, the death rate dropped 5% per year in adults under 50 and 3% per year in those over 50, even as more people kept getting diagnosed. Better treatment is saving more of the people who get caught. But more people are getting caught in the first place, which means the system needs to find and evaluate a growing number of lesions with a dermatology workforce that isn't growing anywhere near as fast. That mismatch, more patients chasing a flat or shrinking pool of specialist time, is the demand problem sitting underneath everything else in this piece.
Why primary care is where most suspicious lesions are — and where detection most often falls short
Ask a patient with a weird-looking mole who they call first, and it's almost never a dermatologist. It's their primary care doctor. PCPs are the actual front door for skin cancer concerns in this country, whether or not the system was built with that in mind.
The tool most PCPs lean on is the ABCDE mnemonic (asymmetry, border, color, diameter, evolving), and it's a fine teaching device with a real weakness: its performance swings a lot from clinician to clinician, and it's particularly bad at catching melanomas that don't look like the textbook picture. Most PCPs also haven't been trained in dermoscopy, the magnified, polarized-light exam dermatologists use to look past the surface of a lesion. That's not a knock on PCPs; it's a training and equipment gap built into the structure of how primary care works.
The evidence for that gap is not subtle. A UK primary care audit found that only 13% of skin cancer referrals ended up as a confirmed skin cancer diagnosis, well short of the local audit standard of 50%. Among the urgent, suspected-cancer referral track specifically, only 16% were confirmed as high-risk skin cancers, and malignant melanoma made up just 5.7% of those referrals. Read that carefully and you will notice it points in two directions at once: (i) some cancers are being missed and monitored too long, and (ii) a much larger number of low-risk lesions are being referred unnecessarily, clogging dermatology schedules that already have too much on them. A tool that could address both problems at once, (i) catching more real melanomas and (ii) stopping the flood of benign lesions into specialist schedules, would be solving both ends of the same rope. That's not a failure of effort on the part of primary care doctors. It's a failure of the tools they've had to work with.
How the dermatology shortage amplifies the primary care detection problem
Even a perfect referral would still run into a waiting room. The average wait for a dermatology appointment nationally is 34.5 days, up 7% since 2017. In major cities, that average has climbed to 32.3 days, a 46% jump since 2009. Medicaid-insured patients often wait six to twelve months, which means the patients with the least access to a primary care doctor in the first place are also the ones stuck longest once they finally get referred.
In rural communities and urban safety-net systems, that wait isn't just inconvenient. It's an oncologic event. Lesions don't pause while they wait for an appointment; they evolve, and stages advance. The dermatology workforce shortfall isn't projected to fix itself anytime soon, which means the referral bottleneck isn't a temporary strain the system will grow out of. It's a structural feature of American healthcare, baked in for the foreseeable future.
That's the actual problem AI tools in this space are trying to solve, and it's worth stating precisely because the framing gets muddled so often. The question isn't "can an algorithm read a lesion as well as a board-certified dermatologist?" The more useful question, the one with real clinical stakes, is: can AI help a primary care doctor make a better call at the point of first contact, before that lesion ever joins a six-month queue?
What the clinical evidence says about AI accuracy in melanoma detection
Systematic reviews published between 2023 and 2025 point in a consistent direction: AI performs roughly on par with experienced dermatologists, and noticeably better than generalist clinicians working alone. Pooled data show AI sensitivity of 86.3% and specificity of 78.4%, against generalist clinician sensitivity of 64.6% and specificity of 72.8%. That sensitivity gap, over 20 points, is where most of the clinical value sits. Sensitivity is the measure that tells you how many real cancers a test actually catches, so a 20-point gap means a meaningful number of melanomas that a generalist might miss.
A 2025 systematic review and meta-analysis covering many tens of thousands of test images found pooled AI sensitivity of 0.91 and specificity of 0.64, with an area under the ROC curve of 0.88. Worth naming honestly: there's a real tradeoff between sensitivity and specificity here, and it shows up across nearly every study in this space. Push sensitivity up (catch more real cancers) and specificity tends to fall (flag more benign lesions as suspicious). Which side of that tradeoff matters more depends entirely on the clinical setting; a screening tool in a low-access clinic might reasonably favor sensitivity, while a confirmatory test before a biopsy might want more specificity.
The 2024 multicenter prospective trial by Heinlein and colleagues put real numbers on that tradeoff. AI hit sensitivity of 92.1%, against 73.4% for expert dermatologists. But specificity flipped the other way: dermatologists came in at 82.8%, AI at 67.3%. Neither figure alone tells the full story. The finding that actually matters, the one that should shape how anyone thinks about deploying these tools, is what happened when human judgment and AI were combined: specificity jumped to 90.7%, the best diagnostic balance of anything tested in the study. That's not a hopeful marketing line. It's the empirical result, and it points toward these tools working as decision support bolted onto a clinician's judgment, not a replacement for it.
A separate randomized controlled trial, the first of its kind in this space, published in Nature Digital Medicine in 2024 by Han and colleagues, backed this up: AI assistance improved diagnostic accuracy overall, and the improvement was largest among generalists. Generalists are, of course, exactly who works in primary care.
How the 2024 BJD prospective trial and Scandinavian feasibility study tested AI in actual primary care settings
Lab accuracy numbers are one thing. Real clinic conditions, with real patients and real time pressure, are another. A 2024 prospective trial published in the British Journal of Dermatology tested an AI-based clinical decision support tool on actual primary care patients rather than a curated set of pretty, well-lit images. The tool held up: high diagnostic accuracy in prospective, real-world use, adding genuine clinical value for PCPs sizing up a skin lesion for melanoma risk.
A companion feasibility study out of the Scandinavian Journal of Primary Health Care, also from 2024, tested the tool two ways: 15 PCPs working through a near-live simulation, and 25 PCPs reviewing dermoscopic images with and without AI support in a separate reader study. In both arms, AI support raised physician diagnostic accuracy. Just as telling, the tool scored 84.8 on the System Usability Scale, landing in the "good" range. That matters more than it sounds like it should: a tool that's statistically excellent but annoying to use in the middle of a 15-minute appointment is a tool that quietly stops getting used. Adoption generally depends on (i) diagnostic accuracy and (ii) the day-to-day experience of using the tool, and these two studies are among the first to test both at once in a primary care setting rather than a dermatology clinic.
Worth saying plainly: both studies worked with small-to-moderate sample sizes in specific clinical contexts, mostly outside the U.S. Whether these results generalize to U.S. primary care at scale is still an open question, not a settled one.
DermaSensor: what FDA clearance means and what the pivotal trial actually showed
On January 17, 2024, the FDA cleared DermaSensor as the first AI-enabled medical device meant specifically for use in primary care to help detect skin cancer. That's a regulatory first, and it signals two things at once: the underlying technology has matured enough to clear FDA review, and the unmet need in primary care is real enough that the agency saw a reason to create this pathway.
The device itself uses elastic scattering spectroscopy, a mouthful of a term for something fairly simple in concept: a handheld, non-invasive tool that emits pulses of light at a lesion and reads the pattern of light that scatters back, which reflects the lesion's underlying tissue architecture. An AI algorithm processes that signal and flags melanoma, squamous cell carcinoma, and basal cell carcinoma. No dermoscope, no specialized imaging setup. Just a handheld device built to sit in a primary care exam room, not a dermatology suite.
The pivotal trial behind the clearance, DERM-SUCCESS, ran across hundreds of lesions at 22 clinics in the U.S. and Australia. Histopathology confirmed 224 of those lesions as cancer, a minority of everything evaluated. The device hit sensitivity of 95.5%, against a PCP sensitivity of 83%, and a a very high negative predictive value. That NPV figure is arguably the most useful number in the whole trial for a primary care setting: it tells you, as a PCP, how confident you can be that a negative result genuinely means "not cancer." The device also demonstrated non-inferiority to a high dermatologist sensitivity benchmark.
Specificity, though, came in at just 20.7%, and that limitation deserves to be stated as plainly as the good news: roughly one in six lesions the device flagged as positive turned out to be cancerous. That's low. It means the device is built to work as a safety net, not a rule-in test. It tells a PCP when a lesion is probably fine, more than it tells them when a lesion is definitely cancer.
A utility sub-study inside the same trial found that using the device raised physician diagnostic sensitivity by a meaningful margin, and referral sensitivity substantially. Physician confidence in their own management decisions rose considerably, which is a big jump for a single handheld device to produce. An independent validation study run at UPMC found an AUC of 0.79, identical to the AUC from the FDA pivotal trial. That kind of replication in an outside, investigator-initiated study is generally the sort of confirmation you would want to see for a device this new.
A skin tone sub-analysis within DERM-SUCCESS found very high sensitivity across skin cancer types, with performance holding up similarly across Fitzpatrick skin type subgroups. Given the mortality gap laid out at the start of this piece, that finding matters more than a routine subgroup analysis normally would.
Other AI approaches moving through the evidence pipeline
DermaSensor isn't the only approach in motion, and the variety is worth a quick tour because each one is solving a slightly different piece of the puzzle. Multimodal AI systems that combine dermoscopic imaging with molecular and genomic data, built on datasets like MRA-MIDAS, are still research-stage but pointed toward individualized risk stratification rather than a single binary answer.
The C4C Risk Score, developed at Anglia Ruskin University in 2024, takes a completely different route: no imaging at all. It's a machine-learning model that distilled 22 clinical features down to 7, things like whether a lesion changed size, color, or shape, whether it looked pink or inflamed, and even the patient's hair color at age 15. It hit notably higher accuracy than the older 7-Point Checklist and the Williams score. No hardware, no imaging, just a smarter checklist. That makes it a candidate for very low-resource settings where even a handheld device might be one piece of equipment too many.
Then there's teledermatology paired with AI triage: images captured and stored, then screened by an algorithm before a dermatologist ever looks at them, letting specialists weigh in asynchronously instead of needing to be available in real time. Dermoscopy-enabled eConsult platforms work along similar lines, letting PCPs send images out for AI-assisted or specialist review, chipping away at the wait-time problem through a different channel than a point-of-care device does. What all of these approaches share is a boundary they don't cross: none of them replace a dermatologist for a confirmed or complicated case. They're all aimed at making the first encounter, wherever it happens, sharper than it would otherwise be.
Where these tools fit inside a primary care workflow — and where they do not
The practical question for you, if your clinic is considering one of these tools, is not whether it works in a study. It is whether it adds a step to a busy visit or replaces one. Most of what's cleared and in the pipeline right now adds a layer of decision support on top of an existing encounter rather than rebuilding the visit from scratch.
A point-of-care device like DermaSensor is meant to be used during the visit itself: your patient comes in, you examine the lesion, and the device gets used before you make a referral decision. The high NPV supports a specific, narrow use case: confidently telling a patient they don't need a dermatology referral for a low-risk lesion, which chips away at the unnecessary-referral half of the problem described earlier. The high sensitivity supports the flip side: catching a lesion a PCP might have otherwise just told the patient to keep an eye on.
Teledermatology and eConsult tools sit in a different spot in the timeline, between the visit and the specialist, working asynchronously with their own latency and workflow quirks. None of this happens automatically just because a device shows up in a supply closet. The Scandinavian feasibility study used near-live simulation training, and your clinic will likely need something similar: brief but structured onboarding, not just a device sitting on a shelf with a manual you never open. There's also a documentation and liability layer that's still being worked out. FDA clearance sets a regulatory floor, but each institution's policy on how an AI-assisted call gets documented and who's accountable for it is still very much a work in progress.
None of these tools replace clinical judgment on an atypical or genuinely confusing lesion, a full skin exam, a biopsy, or a dermatologist managing a confirmed high-risk case. The value proposition is strongest exactly where the shortage math from earlier in this piece is worst: rural practices, safety-net clinics, and any setting where the referral queue stretches for months instead of weeks.
Equity considerations that should shape how health systems deploy these tools
Go back to the number that opened this piece: 3.75 times higher melanoma mortality in darker-skinned patients despite lower incidence. That is the baseline you should measure any new detection tool against, because a tool that performs unevenly across skin tones will not fix that gap; it can just as easily make it worse while appearing, on paper, like an improvement.
DermaSensor's sub-analysis showing consistent sensitivity across Fitzpatrick skin type subgroups is a genuinely encouraging sign, and it's the right kind of evidence to be asking for from any device in this category. But a sub-analysis inside a larger trial isn't the same as a dedicated, adequately powered prospective study built specifically around diverse populations. That distinction matters, because a lot of the AI dermatology research that came before this generation of tools was trained mostly on images from lighter-skinned patients, a known and well-documented limitation in the field. Whether the newer wave of tools has actually closed that gap, or has simply narrowed it enough to pass a sub-analysis, is a question you and your health system should keep asking as you decide where and how to deploy these devices.


