Skip to content

Request clarification of AMBER scoring for Reverse Illumina reads #67

Description

@RhysCAllen

Greetings CAMI team,

My question is about how AMBER treats the scoring of reverse reads.

For genome binning and taxon binning challenge, the @@SEQUENCEID column can contain CAMI read IDs.

So for example, for the plant rhizosphere dataset from CAMI II, HiSeq Illumina reads, our upload to AMBER for binning of user-created assemblies (not gold-standard assemblies) might look like this:

@Version:0.9.0
@SampleID:rhimgCAMI2_short_read_sample_0
#
@@SEQUENCEID	BINID	TAXID	_CONTIG_	
S0R10862/1	bin_11	384	c_000000003390	
S0R10862/1	bin_11	384	c_000000003390	
S0R136448/1	bin_11	384	c_000000003390	
S0R136448/1	bin_11	384	c_000000003390	
S0R171187/1	bin_11	384	c_000000003390	
S0R171187/1	bin_11	384	c_000000003390	
S0R214209/1	bin_11	384	c_000000003390	
S0R214209/1	bin_11	384	c_000000003390	

...... millions of lines .....

S0R12628509/1	bin_5	75612	c_000000000862	
S0R13262523/1	bin_5	75612	c_000000000862	
S0R13672719/1	bin_5	75612	c_000000000862	
S0R14374433/1	bin_5	75612	c_000000000862	
S0R14512923/1	bin_5	75612	c_000000000862	
S0R14502081/2	bin_5	75612	c_000000000862	
S0R14817008/1	bin_5	75612	c_000000000862	
S0R15182902/2	bin_5	75612	c_000000000862	
S0R15537387/1	bin_5	75612	c_000000000862	

I'm scraping the reads per contig from my sample_0.sam file, which was used to determine differential abundance when creating these bins.

I learned that a requirement of the SAM record standard is that both read pairs have the same name. So, most mapping software ignores the /1 and /2, storing read direction in bitwise flags instead. I can see that happening in my data.

The raw reads for example include unique IDs S0R9909/1 (forward) and S0R9909/2 (reverse).

However, in my .sam file, there are now two entries for S0R9909/1, and the bit flags indicate that one of these is reverse, and one is forward.

grep "S0R9909/1" sample_0.sam
S0R9909/1 BH:changed:3    83    c_000000005423    2706    45    150=    =    2622    -234    CTTGGGCACCACTACCACTTGGCCGAACAGTCCGTTGGCCACGTTGCCCGCATCCCCTTCTCCGCCGAAGGTCGCGCCACGGCTGCTGGCGGCAAACGCCCCTTCGCGCTCGGCGTACAAGGTGTAGGAACGGGTGGCGCCGGGGGCAAT    02212$$121(1222121$$112#132231211113123121112231/1012111022021111114201211111121131231411113324211111/241111031111122332132122/23211222121111111111222    NM:i:0    AM:i:45
S0R9909/1 BH:changed:3    163    c_000000005423    2622    45    150=    =    2706    234    AATCGGTTGACCGGCCGGGGTGCGGCCGACGCTGGCCAGGCGCATCTCTTCTTCGGTCACGGTGTTGCGATAGGTGCGGCCGCCCTTGGGCACCACTACCACTTGGCCGAACAGTCCGTTGGCCACGTTGCCCGCATCCCCTTCTCCGCC    22(2*2232322222$2222$221,22231113222241222223131331332%23232142233221323#22121221222135222224231#34232232221242231321202222211)23222112$3221%22132$211    NM:i:0    AM:i:45

A consequence of this behavior is that all of the paired reads that are mapped to my assemblies are missing the reverse sequence ID in my AMBER upload file, and instead, the forward read seqID is listed twice. The only reads with a "/2" tag are unpaired singletons.

Does the AMBER scoring method penalize for these "missing" reverse reads? Or does it take the behavior of the mapping software into account already, maybe by detecting "duplicate" forward reads, and scoring them as forward and reverse?

Thanks for any additional info you might have.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions