What a 50%-accurate model taught me


Photo by Kay Nauwelaerts on Unsplash
We built a skin-cancer triage app in a week and won a presentation award for it. The model was about 50% accurate. Both of those facts belong in the write-up.
Two true things
In one week, a team I was part of shipped a working skin-cancer triage application: a ResNet50 retrained by transfer learning on 2,500 ISIC-archive images across five classes, an interface, and a database behind it. We presented it and it won the Biomedical Engineering Presentation Award at 86%.
The model reached about 50% accuracy across five classes, on a dataset with a well-documented bias toward fair skin.
Both of those sentences are true, and the interesting question is what you do with the second one.
The easy version
The easy version of this project on a CV reads: built an AI skin-cancer triage app with explainability, awarded 86%. Every word is defensible. It is also, as a whole, misleading — because a reader will assume the thing works, and it does not work well enough to be trusted near a patient.
Five-class accuracy near 50% on a biased dataset is not a clinical tool. It is a demonstration that a pipeline runs end to end.
What the project actually demonstrated
Something worth claiming, once you stop overclaiming:
- A team can take a model, an interface and a data layer from nothing to working in one week.
- The integration problem — getting a trained network into something a person can actually use — is real work, and it was my part of it.
- A biased training set produces a biased model, and you find that out by measuring, not by hoping.
That last point is the one I actually carry around. The dataset’s fair-skin bias is documented and known. It was still surprising how directly it showed up.
Knowing what your system is not good enough to do is a separate result, and it needs stating separately.
Splitting the claims
The habit I took from this: state the delivery claim and the performance claim apart from each other, and never let the first one imply the second.
“We built this in a week” and “this is accurate enough to rely on” are different assertions with different evidence. Collapsing them is how a demo gets described as a product — and in medical applications that gap is not a presentation problem, it is a safety one.