Single-cell RNA sequencing has changed how we study biology.
For the first time, researchers can examine gene expression one cell at a time rather than relying on an average signal from an entire tissue or sample. To make these enormous, high-dimensional datasets manageable, researchers commonly use clustering: grouping cells with similar transcriptional profiles into recognizable populations.
Clustering is a practical and valuable tool. It can reveal major cell populations and provide a starting point for interpreting complex tissues.
A cluster is an analytical grouping, not a direct observation of a fixed biological truth. The familiar UMAP plot, with its distinct colored islands and annotated cell types, is useful for orientation. Still, it can make biology appear more discrete and uniform than it really is.
When “similar” is not enough.
Cells assigned to the same cluster can differ in important ways. They might be at different stages of maturity, have different activation states, or differ in how closely they resemble the human cell type researchers are trying to model. Within a population that appears uniform, some cells may be on target, while others might still be in an intermediate state or developing toward a different identity. Smaller off-target populations can also blend into larger groups or fall near a cluster boundary, making them difficult to identify through clustering and annotation alone.
This is not a flaw in scRNA-seq or clustering. It reflects the underlying biology: cell identity is often a spectrum, and transitions between states can be gradual. How the data is analyzed can also influence where these boundaries appear and how they shift between identities.
Beyond the cluster label.
At CapyBio, we use standard scRNA-seq data to ask a more specific question:
CapyBio’s Capybara technology compares individual cells with human reference data to quantify how closely they resemble specific cell identities. Rather than depending solely on marker-gene expression or cluster annotation, researchers can evaluate the distribution of cellular identities throughout a sample. This helps reveal differences between cells that may look similar on a UMAP, but vary in how closely they match the intended cell type.
From descriptive data to decisions.
This deeper view can help researchers:
- Compare differentiation outcomes across batches, protocols, and cell lines.
- Identify cells or subpopulations that do not sufficiently align with the intended target identity.
- Detect variation that may influence assay performance or data interpretation.
- Determine whether a cell model is ready for the next experimental stage.
- Identify where protocol refinement could improve model quality.
Clustering will remain essential for exploring single-cell data. But for a cell model to support rigorous discovery, the key question is not only, “What cluster is this cell in?”
CapyBio helps turn that question into a quantitative, actionable measure so teams can move beyond descriptive maps and make better-informed decisions about the models underlying their research.
