# All samples dropped during augur filter

**URL:** <https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390>\
**Category:** Uncategorized\
**Created:** [March 4, 2021, 8:49pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390 "2021-03-04T20:49:21Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 4, 2021, 8:49pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/1 "2021-03-04T20:49:21Z")

</div>

Hi - I see a few questions related to this same issue but none that have been answered so bumping again - when running a basic global build, I keep getting an error in the augur filter step saying ‘all samples have been dropped! Check filter rules and metadata file format.’. It then deletes the filtered.fasta file and exits. I’ve looked at all the log files and they’re almost all empty so the actual issue is very difficult to diagnose. I’ve used many different subsets of sequences, both my own and those downloaded directly from GISAID, as well as the metadata format downloaded directly from GISAID as well. I’ve been able to build trees successfully in the past but would get this error seemingly randomly, and now have been getting it every time. Happy to post other results/logs if that would be helpful but most everything is deleted when the program exits, so there isn’t much to show. Any help would be much appreciated! Thank you!

---

<div class="post-metadata">

**Author:** ![james](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/james/32/11_2.png) [@james](https://nextstrain.discourse.group/u/james)\
**Post date:** [March 8, 2021, 1:41am UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/2 "2021-03-08T01:41:02Z")

</div>

Hi @aeroder – could you post the output that Snakemake prints at the filtering step which may give us some clues.

P.S. the output from `augur filter` should soon be more descriptive thanks to [soon-to-be released work](https://github.com/nextstrain/augur/pull/679) by @jlhudd.

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 8, 2021, 8:21pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/3 "2021-03-08T20:21:12Z")

</div>

![Screen Shot 2021-03-04 at 3.46.12 PM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/d7908138cccf95a3c945baf0594f8bf602796e33.png)

Here is a screenshot of the error message. It doesn’t say much but maybe will be helpful!

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 8, 2021, 11:20pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/4 "2021-03-08T23:20:59Z")

</div>

Thank you for sharing the filter log, @aeroder. My best guess is that there is a mismatch between strain names in the sequence and metadata inputs, but without access to those input data, I can’t be sure.

We just released a new version of Augur (11.2.0) that improves the filter report and always prints the full report to the logs even if all samples have been dropped. Are you able to upgrade your Augur installation so you can re-run this filter step and post the improved report here?

If you’ve installed Augur with pip, you can upgrade with:

```bash
python3 -m pip install --upgrade nextstrain-augur

```

If you are running your analysis with the Nextstrain CLI and Docker, you can get the latest Augur by running:

```bash
nextstrain update

```

If you are running your ncov workflow with `snakemake --use-conda`, you update [the conda environment file to reference `nextstrain-augur==11.2.0`](https://github.com/nextstrain/ncov/blob/bbabb4624042846939b850285b13c5d02a523519/workflow/envs/nextstrain.yaml#L22) (instead of 11.1.2).

If you have installed Augur with Bioconda, the latest version should be available by tonight or tomorrow morning (new Bioconda packages require manual human approval while the other approaches above do not). You’ll be able to upgrade with:

```bash
conda activate nextstrain
conda update --all

```

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 10, 2021, 5:20pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/5 "2021-03-10T17:20:06Z")

</div>

The update seemed to fix the issue! If I get it again, I’ll post the filter message here. Thank you!

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 12, 2021, 3:14pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/6 "2021-03-12T15:14:38Z")

</div>

Okay after one successful run, I’m getting the same error again. I’m attaching a screenshot of the error message which says that all samples were dropped because there is no sequence data. However when I grep for the sequence name in the fasta file and the metadata file, I’ve confirmed that a number of them are exactly the same. I’m sure there is something small I’m missing but any help would be appreciated! Thank you!!

 ![Screen Shot 2021-03-12 at 10.12.50 AM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/263c9a4ada6803385c829f520feec44ba0ec347a.png) ![Screen Shot 2021-03-12 at 10.09.59 AM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/3ca6897fa9f41e32734e8d9d2c075511785d4ee8.png)

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 15, 2021, 5:12pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/7 "2021-03-15T17:12:32Z")

</div>

@aeroder, would you mind sharing the output of the following command (you may need to use `cat -A` if you’re on a Linux system)?

```bash
cat -e 031121-usethis.fasta | grep 'USA/DC-HP00054/2020'

```

I’m wondering if the issue is related to whitespace characters that Augur isn’t handling properly, since your strain names look fine in the metadata and sequence data.

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 15, 2021, 7:37pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/8 "2021-03-15T19:37:32Z")

</div>

![Screen Shot 2021-03-15 at 3.36.12 PM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/74c73863235867cceaf89aa1f0394f9bcdbbca61.png)

This is the result that I got. I downloaded these sequences directly from GISAID so I’m not sure if the issue arises when I append my sequences to the GISAID file or if it is in the GISAID sequences when I download them

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 15, 2021, 8:36pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/9 "2021-03-15T20:36:52Z")

</div>

Ok, that looks like we’d expect. I should have asked this at the same time, but what do you see for the metadata with a similar command?

```bash
cat -e metadata-4.tsv | grep 'USA/DC-HP00054/2020'

```

As another test, could you generate a sequence index for the sequences and search for the same sample there? The index should take ~10 minutes to build.

```bash
# Build the sequence index.
augur index --sequences 031121-usethis.fasta --output sequence_index.tsv

# Search for a specific sample, showing whitespace characters.
cat -e sequence_index.tsv | grep 'USA/DC-HP00054/2020'

```

Behind the scenes, Augur creates [a set of strains from the metadata](https://github.com/nextstrain/augur/blob/8df4b4d3a500102f7aa3ffb40f78b107c01ae937/augur/filter.py#L177) and [a set of strains from sequence data](https://github.com/nextstrain/augur/blob/8df4b4d3a500102f7aa3ffb40f78b107c01ae937/augur/filter.py#L239). It calculates [the set of available strains as the intersection of the metadata and sequence strain sets](https://github.com/nextstrain/augur/blob/8df4b4d3a500102f7aa3ffb40f78b107c01ae937/augur/filter.py#L249). Since these sets consist only of strain names, the bug must be related to how we’re parsing those names from the current data. Now, my best guess is that there’s an additional whitespace character in the metadata, since the sequence id looks fine.

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 16, 2021, 1:48pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/10 "2021-03-16T13:48:51Z")

</div>

Here is the results of the cat command on the metadata file:

 ![Screen Shot 2021-03-16 at 9.43.56 AM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/fe1f0ece9efc3547b4a91821e1bc973bc22518da.png)

And here is the result on the sequence index (as a note, the sequence index only took about 15 seconds to build)

 ![Screen Shot 2021-03-16 at 9.46.16 AM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/3c4662d3bb2ea5a5a70ace4cdca3b68df1a0c8a7.png)

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 16, 2021, 5:04pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/11 "2021-03-16T17:04:06Z")

</div>

Ok, I have one final idea, since `cat -e` doesn’t do everything I thought it did on OS X. To show non-printing characters (including tabs), we need to use `cat -vet`. Can you share the output of the following commands?

```bash
# Search metadata.
cat -vet metadata-4.tsv | grep "USA/DC-HP00054/2020"

# Search sequences.
cat -vet 031121-usethis.fasta | grep "USA/DC-HP00054/2020"

# Search sequence index.
cat -vet sequence_index.tsv | grep "USA/DC-HP00054/2020"

```

I’m sorry this is so complicated! Whitespace is the bane of bioinformatics. Here is an example output from my computer for the metadata (tabs are shown as `^I`):

```bash
$ cat -vet data/example_metadata.tsv | grep Wuhan/WH01/2019
Wuhan/WH01/2019^Incov^IEPI_ISL_406798^ILR757998^I2019-12-26^IAsia^IChina^IHubei^IWuhan^IAsia^IChina^IHubei^Igenome^I29866^IHuman^I44^IMale^IGeneral Hospital of Central Theater Command of People's Liberation Army of China^IBGI & Institute of Microbiology, Chinese Academy of Sciences & Shandong First Medical University & Shandong Academy of Medical Sciences & General Hospital of Central Theater Command of People's Liberation Army of China^IWeijun Chen et al^Ihttps://www.gisaid.org^I?^I2020-01-30$

```

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 16, 2021, 5:47pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/12 "2021-03-16T17:47:31Z")

</div>

No problem - I really appreciate you helping me diagnose this!

Here is the output of those three commands.

I don’t see any obvious whitespace aside from the tabs. I’ve also tried looking in both files for blank lines but don’t see any. I seem to get the ‘all samples dropped’ error seemingly randomly (I’m sure it isn’t random, but I can never identify what exactly triggers it).

 ![Screen Shot 2021-03-16 at 1.46.23 PM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/cf70063dc3328b13a170922f633303982ae841f6.png)

 ![Screen Shot 2021-03-16 at 1.46.33 PM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/11a4c2bc45f307398a70adf046a34d45d80719d0.png)

 ![Screen Shot 2021-03-16 at 1.46.39 PM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/df6919f418d408e9f44383638c6ca4ff04fceec6.png)

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 16, 2021, 6:34pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/13 "2021-03-16T18:34:22Z")

</div>

Interesting…everything looks good unless somehow we have an issue processing Windows-style line endings.

One more thing to check is the header row of the metadata file. What do you get when you run this command?

```bash
cat -vet metadata-4.tsv | head -n 1

```

---

<div class="post-metadata">

**Author:** ![aeroder](https://avatars.discourse-cdn.com/v4/letter/a/977dab/32.png) [@aeroder](https://nextstrain.discourse.group/u/aeroder)\
**Post date:** [March 16, 2021, 6:39pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/14 "2021-03-16T18:39:19Z")

</div>

Here is the result of that:

 ![Screen Shot 2021-03-16 at 2.37.44 PM](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/1b0b4a43110a49f75f624b71f9eb8ba8fa65d16b.png)

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 17, 2021, 9:32pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/15 "2021-03-17T21:32:53Z")

</div>

Thanks, @aeroder! I’m a little stumped now, since none of the most logical explanations I can think of seem to explain the bug. I’m going to experiment with metadata that have Windows-style line endings just to confirm that isn’t the issue, but I’ll follow up with you about maybe getting a copy of your data to see if I can reproduce the problem locally.

---

<div class="post-metadata">

**Author:** ![underscore](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/underscore/32/118_2.png) [@underscore](https://nextstrain.discourse.group/u/underscore)\
**Post date:** [March 21, 2021, 2:11pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/16 "2021-03-21T14:11:33Z")

</div>

Basically the same issue here (macOS Catalina) where all sequences are filtered. Filter rules & log files are not informative. Can provide `FASTA` & `TSV` if you wish.

 ![s](https://canada1.discourse-cdn.com/flex031/uploads/nextstrain/original/1X/92ab13c39c081a65fab28031d0b454a085e7f88d.jpeg)

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 23, 2021, 5:12pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/17 "2021-03-23T17:12:11Z")

</div>

Hi @underscore, it would be great if you could provide FASTA and metadata that are producing this issue. You can message me directly to work out the transfer, if you cannot share these data publicly.

---

<div class="post-metadata">

**Author:** ![underscore](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/underscore/32/118_2.png) [@underscore](https://nextstrain.discourse.group/u/underscore)\
**Post date:** [March 23, 2021, 5:59pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/18 "2021-03-23T17:59:44Z")

</div>

million thanks.

AFAIK the problem was never with the FASTA files but with the metadata where I found various issues (not explained at [Preparing your data — Nextstrain documentation](https://docs.nextstrain.org/en/latest/tutorials/SARS-CoV-2/steps/data-prep.html)) For example

- “XX” in date field is not allowed
- the minimum number of columns is 12
- the 2 Wuhan sequences need to be always included
- added dummy dates
- …

After correcting these errors, the filter issue suddenly disappeared 😀

---

<div class="post-metadata">

**Author:** ![jlhudd](https://yyz1.discourse-cdn.com/flex031/user_avatar/nextstrain.discourse.group/jlhudd/32/583_2.png) [@jlhudd](https://nextstrain.discourse.group/u/jlhudd)\
**Post date:** [March 23, 2021, 8:03pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/19 "2021-03-23T20:03:43Z")

</div>

So strange. None of those issues _should_ affect the filtering step, but this is helpful to know. It sounds like we need a way to validate metadata and sequence index files independently from `augur filter`.

For our own team’s notes, we could consider adding a subcommand for `metadata` to the existing `augur validate` command. A simpler immediate step would be to implement [a validation schema for metadata](https://snakemake.readthedocs.io/en/stable/snakefiles/configuration.html#validation) and apply the Snakemake’s validate function to each metadata input at the beginning of the workflow.

---

<div class="post-metadata">

**Author:** ![pratibha](https://avatars.discourse-cdn.com/v4/letter/p/eada6e/32.png) [@pratibha](https://nextstrain.discourse.group/u/pratibha)\
**Post date:** [August 4, 2021, 5:49pm UTC](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390/20 "2021-08-04T17:49:06Z")

</div>

Hey. I am still getting the same error.I have tried all the things which mentioned here.can some one please let me know the best solution asap.

The command which i had used:

augur filter --sequences data/sequences2021.03.fasta --metadata data/mumbai2021.03\_metadata.tsv --include defaults/include.txt --max-date 2021.03 --min-date 2021.08 --min-length 20000 --output results/filtered.fasta 2\>&1 | tee logs/filtered.txt

[Next page](https://nextstrain.discourse.group/t/all-samples-dropped-during-augur-filter/390.md?page=2)
