A simulation study comparing supertree and combined analysis methods using SMIDGen

Access full-text files

Date

2010-01-04

Authors

Swenson, M. Shel
Barbancon, Francois
Warnow, Tandy
Linder, C. Randal

Journal Title

Journal ISSN

Volume Title

Publisher

Algorithms for Molecular Biology

Abstract

Background:Supertree methods comprise one approach to reconstructing large molecular phylogenies given multi-marker datasets: trees are estimated on each marker and then combined into a tree (the "supertree") on the entire set of taxa. Supertrees can be constructed using various algorithmic techniques, with the most common being matrix representation with parsimony (MRP). When the data allow, the competing approach is a combined analysis (also known as a "supermatrix" or "total evidence" approach) whereby the different sequence data matrices for each of the different subsets of taxa are concatenated into a single supermatrix, and a tree is estimated on that supermatrix. Results: In this paper, we describe an extensive simulation study we performed comparing two supertree methods, MRP and weighted MRP, to combined analysis methods on large model trees. A key contribution of this study is our novel simulation methodology (Super-Method Input Data Generator, or SMIDGen) that better reflects biological processes and the practices of systematists than earlier simulations. We show that combined analysis based upon maximum likelihood outperforms MRP and weighted MRP, giving especially big improvements when the largest subtree does not contain most of the taxa. Conclusions: This study demonstrates that MRP and weighted MRP produce distinctly less accurate trees than combined analyses for a given base method (maximum parsimony or maximum likelihood). Since there are situations in which combined analyses are not feasible, there is a clear need for better supertree methods. The source tree and combined datasets used in this study can be used to test other supertree and combined analysis methods.

Description

M. Shel Swenson and Tandy Warnow are with the Department of Computer Sciences, The University of Texas at Austin, Austin TX, USA -- Francois Barbancon is with Microsoft, Redmond WA, USA -- C. Randal Linder is with the Section of Integrative Biology, The University of Texas at Austin, Austin TX, USA

LCSH Subject Headings

Citation

Swenson, M. Shel, François Barbançon, Tandy Warnow, and C. Randal Linder. “A Simulation Study Comparing Supertree and Combined Analysis Methods Using SMIDGen.” Algorithms for Molecular Biology 5, no. 1 (January 4, 2010): 8. doi:10.1186/1748-7188-5-8.