Greetings CAMI team,
My question is about how AMBER treats the scoring of reverse reads.
For genome binning and taxon binning challenge, the @@SEQUENCEID column can contain CAMI read IDs.
So for example, for the plant rhizosphere dataset from CAMI II, HiSeq Illumina reads, our upload to AMBER for binning of user-created assemblies (not gold-standard assemblies) might look like this:
@Version:0.9.0
@SampleID:rhimgCAMI2_short_read_sample_0
#
@@SEQUENCEID BINID TAXID _CONTIG_
S0R10862/1 bin_11 384 c_000000003390
S0R10862/1 bin_11 384 c_000000003390
S0R136448/1 bin_11 384 c_000000003390
S0R136448/1 bin_11 384 c_000000003390
S0R171187/1 bin_11 384 c_000000003390
S0R171187/1 bin_11 384 c_000000003390
S0R214209/1 bin_11 384 c_000000003390
S0R214209/1 bin_11 384 c_000000003390
...... millions of lines .....
S0R12628509/1 bin_5 75612 c_000000000862
S0R13262523/1 bin_5 75612 c_000000000862
S0R13672719/1 bin_5 75612 c_000000000862
S0R14374433/1 bin_5 75612 c_000000000862
S0R14512923/1 bin_5 75612 c_000000000862
S0R14502081/2 bin_5 75612 c_000000000862
S0R14817008/1 bin_5 75612 c_000000000862
S0R15182902/2 bin_5 75612 c_000000000862
S0R15537387/1 bin_5 75612 c_000000000862
I'm scraping the reads per contig from my sample_0.sam file, which was used to determine differential abundance when creating these bins.
I learned that a requirement of the SAM record standard is that both read pairs have the same name. So, most mapping software ignores the /1 and /2, storing read direction in bitwise flags instead. I can see that happening in my data.
The raw reads for example include unique IDs S0R9909/1 (forward) and S0R9909/2 (reverse).
However, in my .sam file, there are now two entries for S0R9909/1, and the bit flags indicate that one of these is reverse, and one is forward.
grep "S0R9909/1" sample_0.sam
S0R9909/1 BH:changed:3 83 c_000000005423 2706 45 150= = 2622 -234 CTTGGGCACCACTACCACTTGGCCGAACAGTCCGTTGGCCACGTTGCCCGCATCCCCTTCTCCGCCGAAGGTCGCGCCACGGCTGCTGGCGGCAAACGCCCCTTCGCGCTCGGCGTACAAGGTGTAGGAACGGGTGGCGCCGGGGGCAAT 02212$$121(1222121$$112#132231211113123121112231/1012111022021111114201211111121131231411113324211111/241111031111122332132122/23211222121111111111222 NM:i:0 AM:i:45
S0R9909/1 BH:changed:3 163 c_000000005423 2622 45 150= = 2706 234 AATCGGTTGACCGGCCGGGGTGCGGCCGACGCTGGCCAGGCGCATCTCTTCTTCGGTCACGGTGTTGCGATAGGTGCGGCCGCCCTTGGGCACCACTACCACTTGGCCGAACAGTCCGTTGGCCACGTTGCCCGCATCCCCTTCTCCGCC 22(2*2232322222$2222$221,22231113222241222223131331332%23232142233221323#22121221222135222224231#34232232221242231321202222211)23222112$3221%22132$211 NM:i:0 AM:i:45
A consequence of this behavior is that all of the paired reads that are mapped to my assemblies are missing the reverse sequence ID in my AMBER upload file, and instead, the forward read seqID is listed twice. The only reads with a "/2" tag are unpaired singletons.
Does the AMBER scoring method penalize for these "missing" reverse reads? Or does it take the behavior of the mapping software into account already, maybe by detecting "duplicate" forward reads, and scoring them as forward and reverse?
Thanks for any additional info you might have.
Greetings CAMI team,
My question is about how AMBER treats the scoring of reverse reads.
For genome binning and taxon binning challenge, the @@SEQUENCEID column can contain CAMI read IDs.
So for example, for the plant rhizosphere dataset from CAMI II, HiSeq Illumina reads, our upload to AMBER for binning of user-created assemblies (not gold-standard assemblies) might look like this:
I'm scraping the reads per contig from my sample_0.sam file, which was used to determine differential abundance when creating these bins.
I learned that a requirement of the SAM record standard is that both read pairs have the same name. So, most mapping software ignores the /1 and /2, storing read direction in bitwise flags instead. I can see that happening in my data.
The raw reads for example include unique IDs S0R9909/1 (forward) and S0R9909/2 (reverse).
However, in my .sam file, there are now two entries for S0R9909/1, and the bit flags indicate that one of these is reverse, and one is forward.
A consequence of this behavior is that all of the paired reads that are mapped to my assemblies are missing the reverse sequence ID in my AMBER upload file, and instead, the forward read seqID is listed twice. The only reads with a "/2" tag are unpaired singletons.
Does the AMBER scoring method penalize for these "missing" reverse reads? Or does it take the behavior of the mapping software into account already, maybe by detecting "duplicate" forward reads, and scoring them as forward and reverse?
Thanks for any additional info you might have.