Finding the Right Cancer Culprits Using Mutational Heterogeneity
Imagine you’re a police officer on patrol. You get a call: several 30-year-old Caucasian men were seen breaking into homes nearby and stealing heirlooms. Witnesses say the suspects ran into a convention center. When you arrive, the whole center is hosting an antique show — full of 30-year-old Caucasian men carrying heirlooms. How do you catch the real culprits?
You could arrest everyone who matches the description and question them. Or you could look for clues that most of these people don’t belong, or call the station for more context to narrow the crowd down. Without that extra context, you risk arresting the wrong men while the real thieves get away.
Why Mutation Rate Variability Matters
Cancer genomics faces a similar problem today. Better sequencing technology now lets researchers detect every genomic mutation in a cancer cell.
But mutation rates aren’t the same across all genes. That sounds like a small detail — but it can trick scientists into thinking a mutated gene causes cancer, when really that gene just “matches the description.” It’s simply more prone to mutation, with no real role in oncogenesis.
This is exactly what happened with olfactory genes in lung cancer. Researchers found high mutation rates in these genes1. But olfactory genes have no obvious reason to drive lung tumors, so the scientists ultimately concluded these mutations played no role in transforming lung epithelial cells1.
In Nature, Lawrence et al. dug deeper into this problem. They showed that failing to correct for mutation rate differences across the genome can produce false positives for cancer-associated genes1. To show why this correction matters, the authors compared datasets with similar mutation frequencies against datasets with different average mutation frequencies. Ignoring mutation variability, they found, increases false identification of cancer-associated genes.
The authors also warned about a related risk. As sample sizes grow in “big data” sets — like the American Society of Clinical Oncology’s “CancerLinQ™”2 and the Cancer Genome Atlas3 — failing to correct for mutation rates could make the false-positive problem worse, by lowering the bar for statistical significance.
Lawrence identified three contextual factors researchers often fail to correct for:
- Patient-specific context — mutation rates vary between samples of the same cancer type.
- Sequence-specific context — mutation rates vary depending on the nucleotides surrounding a sequence.
- Replication/transcription-specific context — mutation rates vary depending on when a gene is replicated or transcribed.
Using the olfactory gene example above, plus 3,083 tumor-control pairs across 27 cancer types, the authors showed why these context factors matter. They built a new algorithm for context-based analysis, called MutSigCV.
What the Data Showed
Lawrence et al. studied 3,083 tumor-normal pairs across 27 tumor types, all with variable average mutation rates. Across all of them, they found a 1,000-fold difference in median mutation frequency. The pattern was clear:
- Hematological and pediatric cancers had the lowest variance.
- Cancers linked to environmental causes, like smoking and radiation, had the highest variance.
This variability shows why treatment needs to be context-specific — not just across cancer types, but even between patients with the same cancer.
The authors corrected for mutation rates tied to tissue type, known carcinogens, and cancer type. Even then, they still found high mutational variability within samples of the same cancer type. Since carcinogens couldn’t fully explain this, Lawrence et al. proposed that a gene’s nucleotide makeup might also drive variability in mutation rate.
To test this, they checked mutational patterns across multiple tumors using 96 possible mutation types (accounting for flanking bases). They then plotted the results on a radial chart1. Certain tumor types clustered around specific mutated sequences with matching flanking nucleotides — for example, lung cancers showed mostly C-to-A mutations — though even this varied within a single cancer type.
Both the mutation rate variance and the sequence patterns mattered. But the biggest driver of mutational heterogeneity turned out to be regional differences across the genome — accounting for more than a fivefold difference in median mutation rates1. Lawrence et al. traced this to two factors:
- How often a gene is transcribed.
- When its DNA section gets replicated.
Mutation rates were highest in genes with low transcription and late DNA replication. When the authors compared the falsely-implicated olfactory genes to real cancer-associated genes, they found a clear difference: olfactory genes have lower transcription rates and replicate later, while cancer-associated genes transcribe more and replicate earlier.
In other words, both types of genes build up mutations — but for different underlying reasons. Without correcting for replication and transcription timing, researchers can easily mistake a gene with high raw mutation counts for a true cancer-associated gene.
Why This Matters for Drug Discovery
As the authors put it, “the rich variation in mutational spectrum across tumours underscores the problems with using an overly simplistic model of the average mutational process for a tumour type and failing to account for heterogeneity within a tumor type.”
Their new algorithm, MutSigCV, corrects for these context-dependent quirks, letting researchers analyze cancer genomes without generating these false positives.
Using MutSigCV, Lawrence et al. took a list of 450 suspected cancer-associated genes in lung carcinoma and narrowed it to just 11 genes truly linked to cancer1. That’s a striking demonstration of why context-specific analysis matters so much in cancer genomics. Without it, whole genome sequencing aimed at finding new drug targets could lead pharmaceutical and biotech companies toward targets that merely “fit the description” — innocent bystanders, not the real cancer culprits.
References:
1 Lawrence, M. S. et al. Mutational heterogeneity in cancer and the search for new cancer-associated genes. Nature, doi:10.1038/nature12213 (2013).
2 DeMartino, J. K. & Larsen, J. K. Data Needs in Oncology: “Making Sense of The Big Data Soup“. Journal of the National Comprehensive Cancer Network 11, S-1-S-12 (2013).
3 Network, C. G. A. R. Comprehensive genomic characterization of squamous cell lung cancers. Nature 489, 519-525, doi:10.1038/nature11404 (2012).