Advice on organizing wildcards phylogenetic workflow

I am hoping to eventually create and maintain a Nextclade dataset of a currently unsupported pathogen and I’m looking for a little advice on how to organize the phylogenetic workflow. My pathogen has multiple subtypes and and several genes people traditionally use to generate phylogenetic trees. I would like to be able to be able to select the “change dataset” option in Auspice to view data for each combination of subtype and gene, but I’m unsure about the best way to define the subtype and gene wildcards within the phylogenetic workflow. I was looking at the Nextstrain repository for measles where they define each combination of gene-geography as a seprate build via the defaults/config.yaml file and the Snakefile. Would this be the ideal way to set things up? I’m relatively inexperienced with Nextstrain, Nextclade, and Snakemake so I want to make sure I start “correctly” from the very beginning! Thanks so much in advance.

Hi @eam,

Yes, I would recommend following the pattern in the measles repo. It’s relatively new so we don’t have any documentation for it, but basically the full dataset name is captured by a build wildcard which can be broken down into different parts separated by /. Each part gets its own dropdown in Auspice.

As you mentioned, the measles datasets are named with two parts: {gene}/{geography}. gene is used as a wildcard in the workflow because there are gene-specific Snakemake rules. On the other hand, there are no geography-specific Snakemake rules so it doesn’t need to be defined as a wildcard in the workflow.

Thank you so much! I assume it makes the most sense to start with the basic pathogen repository and go from there?

Yes, you can use the pathogen-repo-guide for initial file structure. From there, you can refer to the Creating a phylogenetic workflow tutorial to understand the individual steps of the workflow, and the measles repo to understand how to implement in Snakemake.