Academic
Publications
RAIphy: Phylogenetic classification of metagenomics samples using iterative refinement of relative abundance index profiles

RAIphy: Phylogenetic classification of metagenomics samples using iterative refinement of relative abundance index profiles,10.1186/1471-2105-12-41,BM

RAIphy: Phylogenetic classification of metagenomics samples using iterative refinement of relative abundance index profiles   (Citations: 1)
BibTex | RIS | RefWorks Download
BACKGROUND: Computational analysis of metagenomes requires the taxonomical assignment of the genome contigs assembled from DNA reads of environmental samples. Because of the diverse nature of microbiomes, the length of the assemblies obtained can vary between a few hundred bp to a few hundred Kbp. Current taxonomic classification algorithms provide accurate classification for long contigs or for short fragments from organisms that have close relatives with annotated genomes. These are significant limitations for metagenome analysis because of the complexity of microbiomes and the paucity of existing annotated genomes. RESULTS: We propose a robust taxonomic classification method, RAIphy, that uses a novel sequence similarity metric with iterative refinement of taxonomic models and functions effectively without these limitations. We have tested RAIphy with synthetic metagenomics data ranging between 100 bp to 50 Kbp. Within a sequence read range of 100 bp-1000 bp, the sensitivity of RAIphy ranges between 38%-81% outperforming the currently popular composition-based methods for reads in this range. Comparison with computationally more intensive sequence similarity methods shows that RAIphy performs competitively while being significantly faster. The sensitivity-specificity characteristics for relatively longer contigs were compared with the PhyloPythia and TACOA algorithms. RAIphy performs better than these algorithms at varying clade-levels. For an acid mine drainage (AMD) metagenome, RAIphy was able to taxonomically bin the sequence read set more accurately than the currently available methods, Phymm and MEGAN, and more accurately in two out of three tests than the much more computationally intensive method, PhymmBL. CONCLUSIONS: With the introduction of the relative abundance index metric and an iterative classification method, we propose a taxonomic classification algorithm that performs competitively for a large range of DNA contig lengths assembled from metagenome data. Because of its speed, simplicity, and accuracy RAIphy can be successfully used in the binning process for a broad range of metagenomic data obtained from environmental samples.
Journal: BMC Bioinformatics , vol. 12, no. 1, pp. 41-14, 2011
Cumulative Annual
View Publication
The following links allow you to view full publications. These links are maintained by other sources not affiliated with Microsoft Academic Search.
    • ...Typically, oligonucleotide frequencies signatures are used: among these, tetranucleotide frequencies signature is the most used [13], [14], [17], [20]; sometimes frequencies of short oligonucleotides (length up to 6) are used together [16], [19]; a few tools adopted oligonucleotides longer than 6 nucleotides [8], [15], [21]...

    Fabio Goriet al. Genomic signatures for metagenomic data analysis: Exploiting the rever...

Sort by: