PDMB Sequencing report, potential outliers

Hi again @mattk,

I’m back with a couple more PDMB & sequencing-related questions I was hoping you could help me with.

First, I’m looking to include some sort of sequencing report in my methods. The information in the 2024 data paper was a good start, but other information like what fraction of the reads were human, what was in the controls, and if any samples were removed for reasons other than insufficient read depth (etc) would be appreciated.

Second, in analysing the PDMB stool/gut data I found a group of what look like outliers - 42 samples with much lower sequencing depth than the rest of the samples (median 124k reads vs 1.45m) . This group also seems to have more taxa associated with the oral microbiome than the gut microbiome, as in lots of Streptococcus, Rothia, Neisseria, Actinomyces, etc. Is it possible that these samples were misidentified as stool samples? Moving forward, would you suggest excluding these from analysis? The sample names are below, thanks for your help! - Adam

FOX_928949
FOX_468215
FOX_563676
FOX_337847
FOX_498006
FOX_563567
FOX_489847
181
088
192
FOX_739702
130
201
FOX_808609
242
FOX_963562
FOX_600771
FOX_644064
FOX_712622
FOX_112824
FOX_317663
FOX_677694
FOX_912118
063
FOX_283957
FOX_284498
FOX_589353
FOX_096521
FOX_253293
FOX_499510
FOX_380673
FOX_876367
FOX_917197
FOX_519587
FOX_428825
FOX_741697
FOX_354214
FOX_562209
FOX_489040
243
168
274

Hi Adam (@alemkow),

Thanks for reaching out. I’m taking a look at these questions and will get back to you.

Best,

Matt

Hi @alemkow,

I wanted to provide a brief update on this thread. We are working on providing the sequencing report data on Fox DEN for researchers. These things can take some time, so we appreciate your patience!

Best,

Matt

Thanks very much for the update Matt, it is appreciated.

Hi @alemkow (Cc @MYSchmidt ),

We’ve uploaded the PDMB metagenomic sequencing QC data to FoxDEN. You can find the files under Fox DEN > Resources > PD Microbiome QC and Read Mapping Metrics:

  • qc_metrics_consented.csv — per-sample sequencing quality control metrics (read depths at each processing stage, Q-scores, GC content, deduplication rates)
  • sample_overview_consented.csv — per-sample read counts at the genus, species, and strain levels, and mapping to KEGG Orthology groups

A full data dictionary describing all column names, definitions, and units is available at Fox Insight Parkinson's Disease Microbiome Study Sequencing Reports: Quality Control and Mapping Metrics .

We hope that this additional information will be helpful in determining exclusions. Please let me know if you have any feedback at all.

Best,

Matt

Amazing! Thank you @mattk !