Melanoma-Li
Benign or malignant from a dermoscopic image, tuned by nested cross-validation and then scored on 1,512 images from a separate collection.
- Year
- 2025
- Kind
- Research project
- Stack
- PyTorch, Optuna
- Data
- HAM10000 for training; the ISIC Archive’s MSK collection as the outside set
Background
Telling benign from malignant lesions in dermoscopic images is a crowded problem, and a high score on images from the training collection says little on its own. The question worth answering was whether the result holds on images gathered elsewhere.
Method
The model was trained on HAM10000, 10,015 dermoscopic images, with its diagnoses collapsed to benign and malignant. Five outer folds each held out a fifth of the images. Inside each, 50 Optuna trials tuned the hyperparameters on four inner folds, so a fold’s score never saw the data that tuned it.
Each fold trains three ImageNet-pretrained networks, ResNet-101, EfficientNet-B4 and DenseNet-121, each with squeeze-and-excitation attention, and combines them with learned softmax weights. Augmentation mixes each image with another of the same class. For the outside images the five fold models’ probabilities are averaged, and the desktop app shows a Grad-CAM map with each prediction.
- From the project’s configuration (5 outer folds, 4 inner, 50 trials) and its cross-validation code; AUCs from slide 50 of its deck.
- Folds are numbered from 0, as the code and the slide’s legend number them. A fold’s held-out share is drawn as one block; the folds are drawn at random, stratified by class.
- A fold model is ResNet-101, EfficientNet-B4 and DenseNet-121, combined by learned softmax weights; the best of its trials sets how it is trained on its four fifths.
- The outside images, the ISIC Archive’s MSK collection, took no part in training or tuning. The average is the mean of the five fold models’ probabilities.
- The mean AUC within the folds, 0.952, is on other images and is not on this scale.
Results
Across the five outer folds the mean AUC was 0.952, with accuracy 0.922, sensitivity 0.786 and specificity 0.955. On 1,512 images from the ISIC Archive’s MSK collection, a public set collected separately from HAM10000, the five fold models scored AUC 0.906 to 0.929 on their own and 0.946 averaged. At a threshold of 0.34 the average reached accuracy 0.896, sensitivity 0.82 and specificity 0.92.
Limitations
The threshold of 0.34 was chosen on the outside set’s own labels, as the one that maximises F1, so the accuracy, sensitivity and specificity quoted at it are the best that set allows rather than a forecast; the AUC of 0.946 does not depend on a threshold and is unaffected.
A sensitivity of 0.82 means about one malignant lesion in six is missed at that threshold; moving the threshold toward sensitivity would catch more of them at the cost of specificity.
The outside set is a public archive, not a clinic’s consecutive patients, and there is no prospective or calibration study, so these numbers say how well the model ranks archive images, not how it would serve in practice. Collapsing the diagnoses to two classes also hides which lesions are confused with which, so the figures cannot say whether the lesions missed are melanomas or other malignant lesions.
Public repository, MIT, 2025.