Hi everyone!
I have come to another unexpected finding with the SomaScan CSF data (under project 151), and I cannot find any information online for the SomaScan version used with PPMI (for PPMI we have 4,785 aptamers, whereas online I can see information mostly for versions with 7k and 11k aptamers). So, I’m bringing my problem here to see whether anyone else had the same issue and could share some thoughts ![]()
Basically, some of the gene IDs/symbols I get from SomaScan’s sequence ID are different depending on whether I do the mapping based on the CSV files provided in PPMI or based on newer versions of SomaScan (for example this link is the one I’m using but I also get the same from this repository). Just to give 3 examples (there are many more!):
- Sequence ID 2728-62 has 3 gene IDs from PPMI’s CSVs (LAMA1; LAMB1; LAMC1), whereas online I get only 1 gene ID (TNC)
- Sequence ID 11089-7 has 6 gene IDs from PPMI’s CSVs (IGHA1; IGHA2; IGK; IGL; JCHAIN; PIGR), whereas online I get only 2 gene IDs (IGHA2; IGHA1)
- Sequence ID 2783-18 is mapped to gene CCL3L3 from PPMI’s CSVs, whereas online it’s mapped to CCL3L1
(Notice that I’m giving the gene symbol to make it easier to explain but it’s the same with the gene (Entrez) IDs)
Do you have any idea on why this might be the case, and which cases I should “choose” for example for enrichment analysis, because as you can see I can get very different results depending on where I look.
I will really appreciate any comments!