Melanoma-Li

Benign or malignant from a dermoscopic image, tuned by nested cross-validation and then scored on 1,512 images from a separate collection.

Year
2025
Kind
Research project
Stack
PyTorch, Optuna
Data
HAM10000 for training; the ISIC Archive’s MSK collection as the outside set

Background

Telling benign from malignant lesions in dermoscopic images is a crowded problem, and a high score on images from the training collection says little on its own. The question worth answering was whether the result holds on images gathered elsewhere.

Method

The model was trained on HAM10000, 10,015 dermoscopic images, with its diagnoses collapsed to benign and malignant. Five outer folds each held out a fifth of the images. Inside each, 50 Optuna trials tuned the hyperparameters on four inner folds, so a fold’s score never saw the data that tuned it.

Each fold trains three ImageNet-pretrained networks, ResNet-101, EfficientNet-B4 and DenseNet-121, each with squeeze-and-excitation attention, and combines them with learned softmax weights. Augmentation mixes each image with another of the same class. For the outside images the five fold models’ probabilities are averaged, and the desktop app shows a Grad-CAM map with each prediction.

HAM10000, 10,015 imagesfold 0fold 1fold 2fold 3fold 4fold modelsaveragedheld out, a fifthBfold 0’s four training fifths50 trials,each on all four0.900.910.920.930.940.95fold20143averagedANested cross-validation5 outer foldsBInside one outer fold4 inner foldsCAUC on 1,512 outside images0.90 to 0.95
  1. From the project’s configuration (5 outer folds, 4 inner, 50 trials) and its cross-validation code; AUCs from slide 50 of its deck.
  2. Folds are numbered from 0, as the code and the slide’s legend number them. A fold’s held-out share is drawn as one block; the folds are drawn at random, stratified by class.
  3. A fold model is ResNet-101, EfficientNet-B4 and DenseNet-121, combined by learned softmax weights; the best of its trials sets how it is trained on its four fifths.
  4. The outside images, the ISIC Archive’s MSK collection, took no part in training or tuning. The average is the mean of the five fold models’ probabilities.
  5. The mean AUC within the folds, 0.952, is on other images and is not on this scale.

Results

Across the five outer folds the mean AUC was 0.952, with accuracy 0.922, sensitivity 0.786 and specificity 0.955. On 1,512 images from the ISIC Archive’s MSK collection, a public set collected separately from HAM10000, the five fold models scored AUC 0.906 to 0.929 on their own and 0.946 averaged. At a threshold of 0.34 the average reached accuracy 0.896, sensitivity 0.82 and specificity 0.92.

ROC curves, true positive rate against false positive rate, on the outside images. The averaged model’s curve, AUC 0.946, lies above those of the five fold models, whose AUCs run from 0.906 to 0.929.
ROC curves on the 1,512 outside images: the five fold models, and their average in black, labelled Dynamic Ensemble as on the project’s slides.

Limitations

The threshold of 0.34 was chosen on the outside set’s own labels, as the one that maximises F1, so the accuracy, sensitivity and specificity quoted at it are the best that set allows rather than a forecast; the AUC of 0.946 does not depend on a threshold and is unaffected.

A sensitivity of 0.82 means about one malignant lesion in six is missed at that threshold; moving the threshold toward sensitivity would catch more of them at the cost of specificity.

The outside set is a public archive, not a clinic’s consecutive patients, and there is no prospective or calibration study, so these numbers say how well the model ranks archive images, not how it would serve in practice. Collapsing the diagnoses to two classes also hides which lesions are confused with which, so the figures cannot say whether the lesions missed are melanomas or other malignant lesions.

Public repository, MIT, 2025.