Cell type annotation: manual markers, automated references, or both?
How to annotate single-cell clusters so the labels survive peer review — the failure modes of each approach and a workflow that combines them.
Annotation is where a single-cell analysis becomes a biological claim, and where most of them become indefensible. Two approaches dominate, each with a characteristic way of failing.
Manual marker-based annotation
You look at ranked marker genes per cluster and assign labels from knowledge or a marker database. It is transparent and it works well for well-characterised tissues like blood.
It fails through confirmation bias. If you expect five cell types, you will find five, because clusters are labelled by whichever expected marker appears highest. Rare or unexpected populations get absorbed into their nearest famous neighbour. It also fails silently when a marker is not specific in your tissue — CD68 is a macrophage marker in many contexts and expressed elsewhere in others.
Automated reference-based annotation
Tools such as SingleR, Azimuth, CellTypist and scmap transfer labels from an annotated reference. Fast, reproducible, and free of the expectation bias above.
Its failure mode is the reference. A label can only be as good as what exists in the reference: annotate tumour tissue against a healthy atlas and every malignant cell will be confidently called something normal. Automated tools also tend to be overconfident at boundaries — transitional states get assigned a crisp label with a high score, and the transitional biology disappears.
The workflow that survives review
- Cluster at two resolutions. One coarse, one fine. Real cell types are stable across both; artefacts appear and vanish.
- Run automated annotation first, blind. Before you look, so it cannot be tuned to your expectations.
- Then look at markers. Ranked genes per cluster, with the automated call alongside.
- Score gene sets independently. A module score for each expected lineage gives a third, continuous view that does not force a discrete choice.
- Reconcile explicitly. Where the three agree, label confidently. Where they disagree, label broadly ("myeloid") rather than precisely, and say why in the text.
- Check the leftovers. A cluster with no convincing identity is either doublets, stressed cells, ambient contamination — or the interesting thing in your dataset. Check the first three before you claim the fourth.
Reporting that pre-empts objections
Show a figure with the automated labels and the final labels side by side, and a table of the top markers per cluster. State the reference used and its version. Give the percentage of cells left unassigned — an honest small number is more credible than a tidy 100%, which usually means ambiguity was forced into a label.
One habit worth adopting
Write the annotation rationale as you go, cluster by cluster, one line each: what it was called, by what evidence, and what was ruled out. It takes minutes during the analysis and it is the document you will want a year later when a reviewer asks why cluster 7 is what you say it is.