LoRDEC hybrid error corrected read usage

wangchuang2017

于 2019-12-04 16:53:21 发布

阅读量118

点赞数

本文链接：https://blog.csdn.net/u010608296/article/details/103389793

版权

第三代测序技术同时被 3 个专栏收录

257 篇文章 24 订阅

订阅专栏

生信工具Bioinformatics Tools

77 篇文章 40 订阅

订阅专栏

基因组组装assembly

53 篇文章 17 订阅

订阅专栏

Hi,

I am trying a donovo assembly of a reptilian genome (size comparable to humans) with ALLPATHS-LG. I have two illumina libraries paired-end and mate-pairs. In addition to it, I have a pacbio library.

I used LoRDEC to correct the errors in the pacbio data. For this I utilized the short reads from illumina (to get the deBruijn graph). I also carried out the trim-split step given in LoRDEC. My question is do I use the corrected pacbio reads (as is) or do I use the corrected-trimmed-split pacbio reads as long reads in ALLPATHS-LG deno assembly. pipeline

I am asking this because according to the LoRDEC manual "The output is the set of corrected reads also in FASTA format. In these corrected sequences: uppercase symbol denote correct nucleotides, while lowercase denote nucleotides left un-corrected."

Also, I plan to improve upon the correction process by using the corrected pacbio reads (either as corrected or as corrected-trim-split fasta files) as the input for the succeeding step of error correction with an increment in k-mer value and repeat the same. Could anyone tell me if the above steps are meaningful or if they are wrong, suggest an alternative iterated correction protocol.

Thanks

The untrimmed un-split reads, contains uncorrected regions either because of lack of coverage or this regions of high errors and it could not be corrected. so using it could lead to miss-assemble.

regarding correction with different k-mer I think there is a suggested value by LoRDEC depends on the genome, as you mentioned it is a big genome, as I remember you should use 21.

Also I a recommend using HALC, based on this Efficiency of PacBio long read correction by 2nd generation Illumina sequencing

Regarding assembly: If you have high PacBio coverage (>20X) you can use canu for assembly without short reads.

also there is other ways to use PacBio

filling gaps (case of low coverage)
or hybrid assembly using tools like dbg2olc (Just an example, follow this link for more)

follow this post
C: Why we need 100X coverage to get a high-quality assembly?

wangchuang2017

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
打赏
0
评论
LoRDEC hybrid error corrected read usage

Hi,I am trying a donovo assembly of a reptilian genome (size comparable to humans) with ALLPATHS-LG. I have two illumina libraries paired-end and mate-pairs. In addition to it, I have a pacbio libra...
复制链接

扫一扫