Criteria for Selecting "Key Mutations" to Define Clades in Augur/Nextstrain

Hi, everyone

I’m trying to generate a clade.tsv file to annotate specific clades in my tree.json. I know I can get mutation info by clicking/hovering on the tree as per the documentation, but I’m unsure about the criteria for selecting which mutations are critical.

What standards or thresholds (e.g., mutation frequency, nucleotide vs. amino acid, uniqueness to the clade) should one use to pick the “key” defining mutations to ensure accurate and stable clade labeling?

Thanks

Hi @hanguojun,

I don’t think there are any concrete recommendations because the standards/thresholds will vary per pathogen. The only clade criteria that I know of is for seasonal influenza, which you can read about in https://doi.org/10.1111/irv.70230.

Best,
Jover

Hi @joverlee,

Thanks for pointing out the influenza paper! I’ve read through it, but it seems to focus mainly on the criteria for when to assign a new clade name, rather than how to select the specific defining mutations from the many that accumulate in a lineage.

My question is more about the practical filtering step: when a clade has dozens of mutations, what standards do people typically use to pick the 1–3 “key” ones for the clade.tsv?

I’m trying to make sure the mutations I put in clade.tsv are both biologically meaningful and stable over time, rather than just the first few that happen to appear in the tree. Any guidance or references on this specific filtering step would be super helpful!

Best,
Guojun