This year, the Hennig meeting was held at Riverside, California, and I'm fortunate enough to be able to be there (in great part thanks to the Kurt Pickett award, given by the Willi Hennig Society). As in previous occasions, I will try to give a short review of many of talks!
23 jun
I want to remark the excellent reception given by the organizers, and the opportunity to meet some wonderful heteropteran researchers [which is huge for an heteroptera aficionado! :)]
24 Jun
The meeting start up with a plenary talk given by Quentin Wheeler, with the title of Cladistic cosmology. Basically it is about of building a “Hubble like” project for taxonomy, specifically the description of new species, and facilities to make access to museum specimens across the web, both as metadata, but using some form of virtual specimens (e.g. scans of herbaria types, 3d scans of specimens, etc.). I like his point, but I think he miss the point of sampling: we not only need better understanding of current museum data, but also, we need to go to the field and look for never seen before specimens!
The first symposium was organized by Ward Wheeler about linguistics. The first talk, was a video-conf given by Peter Whitely about the phonetic character definition used for an analysis of Uto-Aztec languages. The actual analysis was presented by Ward, that shows some of the particularities of using phonetic characters in comparison with DNA data, such as having 121 letters (against 4) and short strings (against long sequence fragments). Peter Forster talks about the using of phylogenetic networks for language data, and how it can be associated with what we know about the history inferred from mitochondrial molecular markers. You can found his program here. Joanna Nichols also uses phylogenetics networks with emphasis of using different sources of cognate data. To finish the season, John Wenzel shows a reanalysis of some previous indo-european analysis, showing some problems with the character coding, and how much of the current resolution was a by product of the errors in the codings, and the method used. I do not know much about the state of the field of phylogenetic linguistics (although I like the use of these techniques in linguistics!), I found a nice blog that usually talk about phylogenetics in linguistics here.
The next session was on phylogenomics, and was organized by Gonzalo Giribet, so are heavily biased towards invertebrates, but for a frustrated entomologist like me, it was wonderful :D! First Gonzalo talks about how new advances in sequencing technology has increased the amount of available data for phylogenetics, and how the problem was shifted from phylogenetic analysis to orthology recognition among this huge amount of new sequences. Then Torsten Struck shows how the selection of outgroups affects the position of one most mysterious group of annelids, the Myzostomida. Sebastian Kvist, explores the genomics of the blood feeding proteins used by Hirudinid leeches, and the realtionships of their host bacteria, a group related with nitrogen-fixing bacteria found in plants! Kevin Kocot was interested on orthology detection using phylogenetic trees, so he implement a program to detect the best possible set of ortholog sequences. Björn von Reumont talks about the pancrustacean phylogenomics, and how the acquisition of newer data from Remipedia is key in the problem and that results shows than remipedians are more closely related to hexapodans than to more traditional crustaceas, he also hints the existence of some apomorphies for such group based on nervous system. Vanessa González talks about the relationships among Bivalbia, that as many molluscan taxa, seems to be well known, but their phylogenetic relationships are just barely explored.
Then the poster season starts, most of the posters present a lot of entomological work on heteropterans (specially reduviids), so that makes me so happy, I like most the poster from Michael Forthman as he works with Ectrichodiins the group that I see for an small [very small!] amount of time xD, so hopefully he will construct a nice phylogeny of the group! :) The winning poster was given by Guayang Zhang and it is about the elongation of the front legs in harpactorines.
25 Jun
Students day! Most of students talks where given this day, so it will allow the judges to give the student awards with a complete overview of all the material presented by students.
Ansel Payne gives a wonderful talk about the general relationships of Hymenopteran superfamilies. Rebecca Dikow explores a huge data set (in term of characters) using bacterial genomes, so her trees have brutal lengths over 37 millions of steps! Ross Mounce talks about measuring support using supertrees, he also remind us about the importance of using machine readable data for phylogenetics, and that he thinks that using svg files will be great for publishing data, as you can include a lot of non-published metadata inside the files. Julieta Gallego presents a talk about the stability of molecular clocks (as I am a co-author of the talk, maybe I will talk about this in other time!). John Denton gives a talk about the usage of iterative alignments to improve the searches under likelihood approaches. Greame Oatley speaks about a cool group of birds [give me a passeriform, and I will love it ;)], and their species limits, in a work in part phylogeographic, in part taxonomic. Juanita Rodriguez (from Colombia! But she is working in Utah), talks about the phylogeny and biogeography of a group of Pompilid wasps. Fernando Gelin talks about Polybia, a wasp genus. Federico López (another colombian! And old friend of my from undergrad days! He is now at Vermont) gives a talk about the phylogenetics of a group of Vespids. To finish the first student season, I give my own talk about biogeography of amphibians.
In between student talks, there is an interesting symposium about the species problem. I'm personally no see so much in that problem, but is nice to see a lot of talk about that. Fortunatelly, I guess, although the problem is recognized, most taxonomists and systematist can continue to work, and work well, even if there is no consensus in what an species is. Brent Mischler spokes about the phylogenetic (topologic-monophyletic) species concept. Rudolf Meier talks whats about the problem of discussing the topic, and how most people skip the problem at all, and why he thinks it is important to discuss the subject. For me, this is the best talk of the symposium. Finally Kipling Will talks about the influence of the species definition in alpha-taxonomy, specifically how it is affected by some recent approaches, like bar-coding. Then a nice discussion (which includes Quentin Wheeler) was given by the four panelists of the symposium.
The second student seasson begin with a talk given by Christine Hayes about the the phylogenetic relationships of a group of phorid flies. Anna Dal Molin explore several techniques to visualize large groups of trees. Pei-Luen Li talks about a worldwide distributed (but concentrated in Hawai'i) group of asparagales, and Cecilia Waichert explore some adaptation hypothesis about nesting behavior using a phylogeny of Pomipilid wasps.
26 Jun
Pablo Goloboff, gives his talk about the implementation of Iter-PCR, focusing on the effects of changing the ingenuous algorithms, with the use of algorithms that take previous information into account, and makes feasible the use of iter-PCR in TNT. Jean DeLaet, gives an insight on his future program for phylogenetic analysis, Anagallis, that will include his algorithms for dealing with non-applicable data! I'm very curious about it [hopefully it will be open source!]. Nico Franz speaks about using taxonomy ontologies, although I think that kind of work is nice to discover some particular things, I guess that the missing thing is that they at the moment are no using character data, and I think that the key for this kind of databases will be both explicit references, and explicit data supporting each taxonomic affirmation. Lenka Drabkova talks about the plant cytokinins. Claudia Szumik shows the results of a the search of areas of endemism, using also phylogenetic information, with a huge data set of mammals. Then Santiago Catalano presents GB-to-TNT a user friendly program to build huge matrices, and also some nice taxonomy mapping tools found in the newer version of TNT. Daniel Janies, shows that super-trees are not a real solution to the uses of super-matrices, as they fails to produce approximate results of a super-matrix. In one of the most controversial talks Donald Buth argues that mitochondrial data, as non intrinsic source of data, and susceptible of the host-parasite interactions, must not be used for phylogenetic analysis. Personally I do not agree with him, but a friend of me, tells me that the same concerns have been raised in population genetics (a.k.a. Phylogeography) literature. So maybe, mitochondrial data is not good for small taxonomic levels, by the same reason of gene vs species trees, but at large levels, it will provide a nice source of evidence (as the difference between gene and species trees at this levels are no so important). This is an intriguing line of though, as it is precisely the opposite to the most received view, from at least five years ago. John Wiens gives a review of all of his papers about the use of phylogenies for study ancestral ecology. I was too tired after Jon talk, so I miss the next ones. Later Santiago gives a nice talk about the usage of morphometric data on several and different kinds of studies. Tim Crowe talks about the results of a whole life dedicated to study guinea fowls at different taxonomic levels. Dalton Amorim, shows a new analysis of phorid flies based on morphology, one of the few talks using explicit morphological data! Continuing with morphological data, and dipterans, Torsten Dikow shows its phylogenetic analysis Mydidae, he also shows interest on detecting unstable terminals. Efraín de Luna, talks about the use of morphometric data in phylogenetics. To finish the day, Nico Franz gives a talk about how his analysis of some particular group of curculionids evolve in time from the initial data matrix to the definitive one. I think this is a nice exercise, but ultimately, it is difficult to extract some information about it, except from the anecdotal experience.
Later this day was the banquet, that as is a tradition in the hennig meetings was free for students! The banquet speech was given by Dennis Stevenson, the current editor of the journal of the society (Cladistics), and it was funny, although it was somewhat small!
The winners of the awards? Ok, there is a lot of Marie Stoppes and Kurt Pickett awards for traveling students. The Don Rosen award for the best poster was given to Guayang Zhang, as I tell you before. Cecilia Waichert receives the Lars Brundin award, for the second best talk, and John Denton receives the Willi Hennig award for the best student talk! Congratulation to the winners, all are well deserved, and I think the judge have a difficult time selecting these three, because there are a lot of great talks given by students :)!
27 Jun
I think this is the most difficult day to give a talk: everyone is tired of the amount of talks in previous days, as it is the day after the banquet, a lot of people are dehydrated and overtighted, and as it is the last day, you have to check-out from the hotel...
Here this is the last symposium, this one, about molecular clocks organized by Sean Brady, who discuss several sources of error. Tracy Heath talks about how to model dating in bayesian analysis, and Elizabeth Murray has a more concrete talk on the dating of some group of Hymenoptera.
Robert Sansom presents some differences between results using hard and soft anatomy in vertebrates, and the implications for the fossil record, that is only based on hard parts. Mark Simmons shows some problems from using multiple matrices, with small overlap. He present a lot of interesting results for some particular examples, so it is important to known how general they are. Mari Källersjö gives one of the most beautiful talks of the meeting, about a working project of a particular group of plants. Everyone loves his talk :D. Veronica Pereyra talks about thysanopterans, but unfortunately I have to make the check out, so I lost it :-(. Ronald Clouse talks about the phylogeography of an opilion that lives in southeastern US, that is a remaining of a gondwanian clade (as Florida was long time ago, part of African plate!). Ulf Jondelius, talks about the acoelan worms, a poorly known taxon, that he was sampled across the world. To finish the meeting Steve Farris talks about the recent surge of the three taxon statements authors (Ebach, Williams, etc.).
General overview
The meeting was very well organized by John Hartley and Christiane Weirauch, that move around all the time to make sure everything is right. They have a full group of students that are helping with everything, so it was huge :).
On the meeting itself, it has a lot of students, with several very nice talks, and a lot of discussion, that is always welcomed :D. One of the things that I note, is in the same line of previous meetings in which few morphology is used, here I see that although most studies only use molecules for phylogenetic analysis, people is more interested in link several of their findings with well detailed morphological data, hopefully, in the next few years, fully integrative, genomic and morphological data will be used in conjunction in almost every work!
The next year, the meeting will be held in Rostock, Germany, a city on the coast of the baltic sea. Hopefully I will be able to make it, and several of my cladistists ends will be there :D Also, if any one reading this blog interested in phylogenetics, and living somewhere in Europe, I encourage you to go, even if you don't work with parsimony! [In fact, a lot of people that does not use parsimony analysis shows their results at the meeting! And nobody complains about it, at much, someone as why do you not use parsimony, but just that] So! See you there :)!
Acknowledgments
I receive a lot of funding support from several institutions that make this trip possible: CONICET, FONCyT, GBIF (well, this one in the future) provides monetary funding, whereas INSUE allows me to work at their facilities. I receive a Kurt Pickett travel award from the Willi Hennig Society. At personal level, my advisors (Pablo and Claudia) provides me a lot of support, Julí convince me to go, Santí helps me with the trip logistics, and they, with Ross, they are excellent room-mates :)!
Mostrando las entradas con la etiqueta trees. Mostrar todas las entradas
Mostrando las entradas con la etiqueta trees. Mostrar todas las entradas
jueves, julio 05, 2012
lunes, octubre 19, 2009
The behemot! for free! / El behemot gratuito!
Santí noted that the "behemot"'s paper [1] is now available for free from Wiley/Blackwell web-page, this is the link:
Or you can jump directly to the paper abstract and download the pdf:
If you want to look the results, Pablo post a quick guide to manage the trees and extract info from them. If you want to check you favorite group, or just play for a moment, there are no excuses ;)!
And the best of it, as they are macros, you can use for your own trees! and don't forget to check out the TNT wiki!
Santí me hizo notar que el paper donde el "behemot" vio la luz [1] se encuentra disponible de manera gratuita en el sitio de blackwell. El enlace es este:
O pueden ir directamente a la pagina del artículo y descargar el pdf:
Para quienes quieran jugar con los resultados, Pablo publicó una guía rápida para manejar los árboles y poder extraer info de ellos. Así que si tienen curiosidad por ver como quedo el grupo que ustedes trabajan, o simplemente quieren pasar un rato, ya no hay excusas ;)
Y lo mejor de todo, como son macros, se pueden usar en sus propios resultados :D! No se olviden de consultar la wiki de TNT ;)
[1] Goloboff, P.A. et al. 2009. Phylogenetic analysis of 73 060 taxa corroborates major eukaryotic groups. Cladistics 25: 211-230. DOI: 10.1111/j.1096-0031.2009.00255.x
sábado, mayo 16, 2009
Algorithms for phylogenetics 0b: Trees
Another basic component of phylogenetic analysis' programs are the trees. Trees can be the result of an analysis (as in cladistic analysis' programs) or part of the data used (as in biogeography analysis' programs).
Trees are one of the most well known data structures of computer science. In computers science a tree is a collection of elements (nodes), nodes are related by a "kindship" relation that imposes a hierarchy on the nodes. Kinship is a pairwise relationship: one node is the ancestor and the other the descendant. A node can have an indefinite number of descendants, but, at maximum, only one ancestor.
In general, in phylogenetics, the term node is used only for elements with descendants, and terminal of leafs to elements without descendants (in computer science the used terms are internal node and terminal node, or leaf node. Here I will use the terminology from computer science, and when using node I refer to any element of the tree).
A trees has at maximum a single root node. A root node is an internal node without ancestor. Different from most abstract trees of computer science, the only elements with have the user data are the terminals (from the data matrix). Data stored by internal nodes is inferred through algoritmic means.
There are many ways to represent a tree. For me, the most natural one is using structures and pointers. Here is the basic node
typedef struct tagPHYLONODE PHYLONODE;This structure is a node from a binary tree, that is, each internal node have two descendants, the pointers left and right, note that they are pointers to PHYLONODE elements. The pointer to ancestor is anc. In this structure, the root node has anc as NULL (i.e. points to no elements), and terminal nodes has left and right as NULL. The array recons store the data (as bitfields).
struct tagPHYLONODE {
PHYLONODE* anc;
PHYLONODE* left;
PHYLONODE* right;
STATES* recons;
};
I usually store the information from a terminal (extracted from data matrix) into an independent structure, for example
typedef struct tagTERMINAL TERMINAL;And include a pointer to TERMINAL as part of PHYLONODE
struct tagTERMINAL {
char name [32];
STATES* chars;
};
struct tagPHYLONODE {
... //Otros campos aquí
TERMINAL* term;
};
Internal nodes has term as NULL.A nice property of trees, is that subtrees have the same properties of a tree, then recursive algorithms a very natural way to work with trees. Nevertheless, I always prefer to have an independent structure to store the whole tree
typedef struct tagPHYLOTREE PHYLOTREE;This structure store a pointer to the root node (root), an array to the nodes of the tree (nodeArray), a pointer to a data matrix (matix) and the number of nodes.
struct tagPHYLOTREE {
unsigned numNodes;
PHYLONODE* root;
PHYLONODE* nodeArray;
DATAMATRIX* matrix;
};
This approach has the advantage of having several independent sets of trees, each tree with his on characterisitcs (id, name, data matrix), that are no properties of any node, but of the whole set of nodes (A typical case: in biogeography, several cladograms form different organism are used).
Different from structures it is possible to store trees as coordinate arrays. This option is used in [1] for the example code, and it is also the way in which TNT macro language uses trees (the language does not support structures). The declarations of arrays is simple
int* anc;In this case, all arrays are dynamic (they are declared as pointers, but as soon as they are assigned, they can be treated as arrays), and instead of have a pointer, they store the number that is the index of the array. As it is a coordinate array, every element with the same index is the same node. For example anc [17] = 20, assigns the node 20 as the ancestor of node 17. In this case, left [20] or right [20] must be equal to 17. chars [20] has the sates assigned to node 20. If you look at the tree entry in wikipedia (specially the entry on heaps) you will find that is the way used to keep trees.
int* left;
int* right;
STATES** chars;
For me, the bad side of that approximation is that requires an strict bookkeeping (take note that in TNT macro language, TNT automatically updates the information, so that is not a problem!). This is my own experience, I worked with tree reconciliation, and work with someone that does not know working with pointers. So I try a coordinate array approach. But we not know how many nodes would be on the tree after reconciliation, then the code turns messy and extremely difficult to update.
That experience does not show if some alternative is best than the other, just that sometimes the nature of the problem might constraint the way to resolve it, and that some styles of programming are better than others for some programmers ;). In this series I always will use structures.
Just to have some code, and to mark the recursivity of trees, here are a function to store a tree in a parenthetical notation (as in Hennig86, NONA and TNT)
#include < stdio.h >This functions used standard IO operations. The function PrintTree puts the tree header for Hennig86/NONA/TNT trees. And call PrintNode on the root node. PrintNode checks if the node is a terminal, if this is true, it writes the terminal name, if not, the node is an internal one, so it prints the parenthesis and call recursively PrintNode on each one of its descendants.
void PrintTree (FILE* out, PHYLOTREE* tree) {
fprintf (out, "tread \'Simple tree output\'\n");
PrintNode (out, tree - > root)
fprintf (out, ";\np/;");
}
void PrintNode (FILE* out, PHYLONODE* node) {
if (node - > term != NULL) {
fprinf ("%s ", node - > term - > name);
return;
}
fprintf (out, "(");
PrintNode (out, node - > left);
PrintNode (out, node- > right);
fprintf (out, ") ");
}
I think that some time ago Mike Keesey write a post about trees under a object oriented paradigm. I'm unable to find the post, this one is the most closely alternative :P.
References
[1] Goloboff, P. A. 1997. Self weighted optimization: tree searches and character state reconstruction under implied transformation costs. Cladistics 13: 225-245. doi: 10.1111/j.1096-0031.1997.tb00317.x
Previous post on Algorithms for phylogenetics
0a. Bitfields
Etiquetas:
algorithms,
program design,
programming,
software,
trees
viernes, julio 18, 2008
“Weighting” tress with TreeFitter
TreeFitter [1] is a program for matching phylogenies with associations used in biogeography and co-evolutionary studies. It has some problems, as seems that Treefitter move to open-source, maybe this problems can be solved! Here I address some problems produced with the 'weighting tree' procedure. The analogous problems is found in some phylogenetic methods and programs, and I address it laterally (they consequences are fully discussed in [2]).
As many biogeographic programs [3, 4, 5] TreeFitter only dealt with perfectly dicotomic trees. To overcome this fault, it implements a weighting of trees. Then you can put all the dicotomic trees found in the phylogenetic analysis and put a fractional weight to each tree. For example, if you found four trees, each one would be weighted by 1/4. This is in fact a majority rule consensus. Ronquist, who is a defender of bayesian methods, see the weighting of trees as a positive characteristic as it covers the 'uncertainty' of the analysis [6]. Then it express the 'confidence' (support in a most relaxed version) of each clade. But contrary to intuitive expectations, majority rule consensus have nothing to do with the support of a determinate clade. Instead, they favoring ambiguous topologies! [2, 7].
Take this example (after [2] and [7]):
At first look it seems that this case can be solved weighting the whole islands instead of each tree, but within each island it is possible to have the same problems of the first example, and in more complex cases, identifying 'topology' islands became difficult and maybe impossible if there are several combinations in independent clades!
Solutions?
Of course the best solution for the problem of multiple trees in TreeFitter without using weights is a new version that dealt with polytomic trees. I guess that the resistance against polytomic trees is because they 'imply' simultaneous speciation. I do not hold that kind of idea ;), and I have no reason to think that polytomic trees supports such interpretation. Even if that is the interpretation is better than weighting (by the way, if you think that polytomic trees implies multiple speciation, the weighting implies fractional speciation! I think that it is a more problematic idea than 'multiple' speciation!).
But before a new version of TreeFitter--or a similar program--arrives it is necessary to found a solution to the problem of multiple trees. I am not happy with the solution that I propose here, but I have no other idea, so here I go...
Use an Adams consensus to detect the terminal/clades that produce the multiple trees, and remove it form the analysis, so you keep the stable part of the topology. I do not like removal of evidence, but it seems safer than relying in the biased solution of tree weights. Maybe some want to re-run the analysis, but I think that is preferable to use the reduced tree as it is based directly on the whole evidence, then, the effect of data removal, I hope, is minor, as the stable part of the original topology is conserved. Also, TNT [11, 12] has several tools to identify moving taxons, so an analysis of the trees with TNT, provides several ways to found stable topologies.
A second problem is that in TreeFitter, the extinctions had a cost greater than 0 [6], so removing terminals could increase the extinction value, but I think that this is a minor problem compared with given weight to ambiguous data. Actually we can think that the cost increase by the new 'extinctions' is the penalty for the ambiguity of the data.
Weighting trees in other methods
As far as I know there are other phylogenetic methods/programs that weight trees, and suffer from the same flaws as the tree weighting under TreeFitter.
First, is the majority rule consensus. It seems that the main reason to prefer majority rule consensus is because they provide results with more resolution. But as examples shows [2, 7] the resolution created by a majority rule tree is coupled ore with the ambiguity generated by a particular topology skeleton. Maybe a more interesting solution could be the use of a minority rule consensus (note that supported clades, always appear as they minority rule is 100%), an advantage is that at least they knowledge that the number of clade instances are not related with the 'support' of the clade. A comparison between the majority and minority rule consensus can show the parts of the tree that are somewhat unstable. But I think that this kind of analysis is better performed using a combination of an strict consensus and an Adams consensus [8], or several of the tools from TNT [11, 12].
Under Bayesian analysis, the frequency of clades is recognized as a 'posterior probability' (a more catholic interpretation might be clade support). Bayesian analysis differs from typical consensus because they made a consensus from trees with different optimality value. But all explored topologies are taken as they are found, and branches are never collapsed, so they are subject of the problems of majority rule consensus [2, 7]. Then in cases where a particular terminal/clade is producing ambiguity, the final topology favors the ambiguous topology. This can produce some illogical results when the data sets even with a single optimal tree, are not very decisive [9], with several near-optimal fit trees. In that case, is is possible that the method prefer the ambiguous topologies from the sub-optimal trees, over the topology in the optimal one, even with high probability values! I recommend you to read [2] for a full criticism of bayesian analysis.
There is also other form of majority rule consensus is used: for support measuring using resampling (like jacknife or bootstrap). In this case the use of majority rule reflects the amount of support for each clade. First, from each resampled matrix analyzed a strict consensus tree is build, then the final majority rule consensus shows the amount of times in which a supported clade appears. In this case there is no bias against or for a particular clade, because the trees from resampled matrix are collapsed. This is not the case for PAUP [2, 10], because in PAUP the trees are weighted in each resampled search, then, the majority rule consensus is a majority rule consensus of several individual, and un-collapsed trees, then it is fully prone to the ambiguity problems of majority rule consensus.
[1] Ronquist, F. 2001. TreeFitter, program and documentation. Available at: http://www.ebc.uu.se/systzoo/research/treefitter/treefitter.html
[2] Goloboff, P.A., Pol, D. 2005. Parsimony and Bayesian phylogenetics. In: Albert, V.A. Ed. Parsimony, phylogeny, and genomics. Oxford univ., Oxford, pp. 148-159.
[3] Page, R. D. M. 1993. Component 2.0, program and documentation. Available at: http://taxonomy.zoology.gla.ac.uk/rod/cpw.html
[4] Page, R. D. M. 1994. TreeMap 1.0, program and documentation. Available at: http://taxonomy.zoology.gla.ac.uk/rod/treemap.html
[5] Ronquist, F. 1996. DiVa, program and documentation. Available at: http://www.ebc.uu.se/systzoo/research/diva/diva.html
[6] Ronquist, F. 2003. Parsimony analysis of coevolving species associations. In: Page, R.D.M. Ed. Tangled trees. Chicago Univ., Chicago, pp. 22-64.
[7] Sharkey, M.J., Leathers, J.W. 2001. Majority does not rule: the trouble with majority-rule consensus trees. Cladistics 17: 282-284. doi: 10.1006/clad.2001.0174
[8] Kearney, M. 2002. Fragmentary taxa, missing data, and ambiguity: mistaken assumptions and conclusions. Systematic biology 51: 369-381. doi: 10.1080/10635150252899824
[9] Goloboff, P.A. 1991. Homoplasy and the choice among cladograms. Cladistics 7: 215-232. doi: 10.1111/j.1096-0031.1991.tb00035.x
[10] Swoford, D. 1998. PAUP, program and documentation, Sinauer, Sunderland (USA).
[11] Goloboff, P.A., Farris, J.S., Nixon, K.C. 2008. TNT, program and documentation. Available at: http://www.zmuc.dk/public/phylogeny/TNT/
[12] Goloboff, P.A., Farris, J.S., Nixon, K.C. 2008. TNT, a free program for phylogenetic analysis. Cladistics, in press. doi: 10.1111/j.1096-0031.2008.00217.x
As many biogeographic programs [3, 4, 5] TreeFitter only dealt with perfectly dicotomic trees. To overcome this fault, it implements a weighting of trees. Then you can put all the dicotomic trees found in the phylogenetic analysis and put a fractional weight to each tree. For example, if you found four trees, each one would be weighted by 1/4. This is in fact a majority rule consensus. Ronquist, who is a defender of bayesian methods, see the weighting of trees as a positive characteristic as it covers the 'uncertainty' of the analysis [6]. Then it express the 'confidence' (support in a most relaxed version) of each clade. But contrary to intuitive expectations, majority rule consensus have nothing to do with the support of a determinate clade. Instead, they favoring ambiguous topologies! [2, 7].
Take this example (after [2] and [7]):
#nexusThe ambiguity is caused by the taxon '10' of data set 'new' that jumps to several positions among the tree. '10' has not influence on the selected tree because it is a product of a dispersal in all topologies. '10' is inestable in all topologies that include (4,5), then the majority rule gives more weight to that topology than to alternative topology, in which '10' position is not ambiguous. In this case both 'tree islands' have the same evidential weight (by the way, that is the reason to prefer strict consensus over other consensus!). But when weights are applied the topology showing (4,5) are preferred as there is more topologies with that clade, then the first tree is preferred because they lack of resolution!
ptree new1 weight=0.143 (1,((6,(7,(8,9))),(10,(4,(3,(2,5))))));
range new1 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree new2 weight=0.143 (1,((6,(7,(8,9))),(2,(3,(4,(5,10))))));
range new2 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree new3 weight=0.143 (1,((6,(7,(8,9))),(2,(3,(5,(4,10))))));
range new3 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree new4 weight=0.143 (1,((6,(7,(8,9))),(2,(3,(10,(4,5))))));
range new4 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree new5 weight=0.143 (1,((6,(7,(8,9))),(2,((3,10),(4,5)))));
range new5 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree new6 weight=0.143 (1,((6,(7,(8,9))),(2,(10,(3,(4,5))))));
range new6 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree new7 weight=0.143 (1,((6,(7,(8,9))),((2,10),(3,(4,5)))));
range new7 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i, 10:g;
ptree indie1 (1,((6,(7,(8,9))),(4,(3,(2,5)))));
range indie1 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i;
ptree indie2 (1,((6,(7,(8,9))),(2,(3,(4,5)))));
range indie2 1:a, 2:b, 3:c, 4:d, 5:e, 6:f, 7:g, 8:h, 9:i;
htree host1 (a,((f,(g,(h,i))),(d,(c,(b,e)))));
htree host2 (a,((f,(g,(h,i))),(b,(c,(d,e)))));
At first look it seems that this case can be solved weighting the whole islands instead of each tree, but within each island it is possible to have the same problems of the first example, and in more complex cases, identifying 'topology' islands became difficult and maybe impossible if there are several combinations in independent clades!
Solutions?
Of course the best solution for the problem of multiple trees in TreeFitter without using weights is a new version that dealt with polytomic trees. I guess that the resistance against polytomic trees is because they 'imply' simultaneous speciation. I do not hold that kind of idea ;), and I have no reason to think that polytomic trees supports such interpretation. Even if that is the interpretation is better than weighting (by the way, if you think that polytomic trees implies multiple speciation, the weighting implies fractional speciation! I think that it is a more problematic idea than 'multiple' speciation!).
But before a new version of TreeFitter--or a similar program--arrives it is necessary to found a solution to the problem of multiple trees. I am not happy with the solution that I propose here, but I have no other idea, so here I go...
Use an Adams consensus to detect the terminal/clades that produce the multiple trees, and remove it form the analysis, so you keep the stable part of the topology. I do not like removal of evidence, but it seems safer than relying in the biased solution of tree weights. Maybe some want to re-run the analysis, but I think that is preferable to use the reduced tree as it is based directly on the whole evidence, then, the effect of data removal, I hope, is minor, as the stable part of the original topology is conserved. Also, TNT [11, 12] has several tools to identify moving taxons, so an analysis of the trees with TNT, provides several ways to found stable topologies.
A second problem is that in TreeFitter, the extinctions had a cost greater than 0 [6], so removing terminals could increase the extinction value, but I think that this is a minor problem compared with given weight to ambiguous data. Actually we can think that the cost increase by the new 'extinctions' is the penalty for the ambiguity of the data.
Weighting trees in other methods
As far as I know there are other phylogenetic methods/programs that weight trees, and suffer from the same flaws as the tree weighting under TreeFitter.
First, is the majority rule consensus. It seems that the main reason to prefer majority rule consensus is because they provide results with more resolution. But as examples shows [2, 7] the resolution created by a majority rule tree is coupled ore with the ambiguity generated by a particular topology skeleton. Maybe a more interesting solution could be the use of a minority rule consensus (note that supported clades, always appear as they minority rule is 100%), an advantage is that at least they knowledge that the number of clade instances are not related with the 'support' of the clade. A comparison between the majority and minority rule consensus can show the parts of the tree that are somewhat unstable. But I think that this kind of analysis is better performed using a combination of an strict consensus and an Adams consensus [8], or several of the tools from TNT [11, 12].
Under Bayesian analysis, the frequency of clades is recognized as a 'posterior probability' (a more catholic interpretation might be clade support). Bayesian analysis differs from typical consensus because they made a consensus from trees with different optimality value. But all explored topologies are taken as they are found, and branches are never collapsed, so they are subject of the problems of majority rule consensus [2, 7]. Then in cases where a particular terminal/clade is producing ambiguity, the final topology favors the ambiguous topology. This can produce some illogical results when the data sets even with a single optimal tree, are not very decisive [9], with several near-optimal fit trees. In that case, is is possible that the method prefer the ambiguous topologies from the sub-optimal trees, over the topology in the optimal one, even with high probability values! I recommend you to read [2] for a full criticism of bayesian analysis.
There is also other form of majority rule consensus is used: for support measuring using resampling (like jacknife or bootstrap). In this case the use of majority rule reflects the amount of support for each clade. First, from each resampled matrix analyzed a strict consensus tree is build, then the final majority rule consensus shows the amount of times in which a supported clade appears. In this case there is no bias against or for a particular clade, because the trees from resampled matrix are collapsed. This is not the case for PAUP [2, 10], because in PAUP the trees are weighted in each resampled search, then, the majority rule consensus is a majority rule consensus of several individual, and un-collapsed trees, then it is fully prone to the ambiguity problems of majority rule consensus.
[1] Ronquist, F. 2001. TreeFitter, program and documentation. Available at: http://www.ebc.uu.se/systzoo/research/treefitter/treefitter.html
[2] Goloboff, P.A., Pol, D. 2005. Parsimony and Bayesian phylogenetics. In: Albert, V.A. Ed. Parsimony, phylogeny, and genomics. Oxford univ., Oxford, pp. 148-159.
[3] Page, R. D. M. 1993. Component 2.0, program and documentation. Available at: http://taxonomy.zoology.gla.ac.uk/rod/cpw.html
[4] Page, R. D. M. 1994. TreeMap 1.0, program and documentation. Available at: http://taxonomy.zoology.gla.ac.uk/rod/treemap.html
[5] Ronquist, F. 1996. DiVa, program and documentation. Available at: http://www.ebc.uu.se/systzoo/research/diva/diva.html
[6] Ronquist, F. 2003. Parsimony analysis of coevolving species associations. In: Page, R.D.M. Ed. Tangled trees. Chicago Univ., Chicago, pp. 22-64.
[7] Sharkey, M.J., Leathers, J.W. 2001. Majority does not rule: the trouble with majority-rule consensus trees. Cladistics 17: 282-284. doi: 10.1006/clad.2001.0174
[8] Kearney, M. 2002. Fragmentary taxa, missing data, and ambiguity: mistaken assumptions and conclusions. Systematic biology 51: 369-381. doi: 10.1080/10635150252899824
[9] Goloboff, P.A. 1991. Homoplasy and the choice among cladograms. Cladistics 7: 215-232. doi: 10.1111/j.1096-0031.1991.tb00035.x
[10] Swoford, D. 1998. PAUP, program and documentation, Sinauer, Sunderland (USA).
[11] Goloboff, P.A., Farris, J.S., Nixon, K.C. 2008. TNT, program and documentation. Available at: http://www.zmuc.dk/public/phylogeny/TNT/
[12] Goloboff, P.A., Farris, J.S., Nixon, K.C. 2008. TNT, a free program for phylogenetic analysis. Cladistics, in press. doi: 10.1111/j.1096-0031.2008.00217.x
Etiquetas:
bayesian fireworks,
biogeography,
dispersal,
models,
trees
lunes, noviembre 26, 2007
A brief sketch of a program for phylogenetics
The last week is a nice week for me, as I received a grant to begin a graduate scholarship :) As my main objective is coupled with the development of a program, and my own interest in phylogenetics, and in a lesser extent, in software engineering, i will try to show here how I think a phyligenetic-related software would be implemented.
Do not take it for granted! The software projects in which i am involved to date, are smaller and short-term, and then I usually make some (ugly) shortcuts for immediate solutions but with low flexibility to modifications. Here my sketch take into account the possibility of further refinements of each part, so it is far more modular than my small projects.
Every program is composed with two parts. The first, the more messier and the most bored part to write, is the user interface (UI). The second, is the program itself (the core). In small projects you could mix both parts without much pain, but as the program growths leaving both parts mixed makes the program difficult to maintain: you would rewrite several code parts to introduce new input formats.
Ideally you would try to make the core to be totally independent from the UI, but usually it is not possible (unless your program do noting). An effort to construct a well designed communication between the UI and the core, is always well payed in the future.
I start my design with the core, and in every moment thinking how the core would communicate with a black-box UI. In phylogenetics you have two major parts: the matrix, and the tree. Both are composed of more smaller elements. The matrix is a list of labels, usually with some coupled data (taxon-character matrix, a taxon-distribution matrix). The tree is a topology in which each vertex is associated with the matrix labels.
You could think that both the tree and the matrix are the same, after all, the tree is only an specific way to sort the terminals in the matrix (it seems that TreeFitter[1] works in that way). But there is a difference, the terminals in the matrix are unique, but you could have many tress associated with the same matrix, terminal 'a' are independent in two trees, but they could point to the same element in a matrix.
Apart from terminals, that have a fixed content, internal nodes of a tree are coupled with the same sort of data as terminals, but this is a variable content, dependent on the tree topology. So we could implement a single fixed matrix, and one (or more) variable matrix, when the tree is unused, the variable matrix can be discarded.
To communicate the core with the UI, at this moment, I prefer string streams. They could be find in standard library (although with some little spelling differences), in that way the UI always translate between formats to send to the core specific strings, for example the tree in parenthesis format, so the tree construction would not interact with files, or console input, for example.
In that way, the UI is a central of translation from console, file, command dialogs, etc., that can be implemented for different systems (e.g. Win, Linux, command line), but the core remains the same.
[1] Ronquist, F. TreeFitter, program and documentation, available on line at: http://www.ebc.uu.se/systzoo/research/treefitter/treefitter.html
Do not take it for granted! The software projects in which i am involved to date, are smaller and short-term, and then I usually make some (ugly) shortcuts for immediate solutions but with low flexibility to modifications. Here my sketch take into account the possibility of further refinements of each part, so it is far more modular than my small projects.
Every program is composed with two parts. The first, the more messier and the most bored part to write, is the user interface (UI). The second, is the program itself (the core). In small projects you could mix both parts without much pain, but as the program growths leaving both parts mixed makes the program difficult to maintain: you would rewrite several code parts to introduce new input formats.Ideally you would try to make the core to be totally independent from the UI, but usually it is not possible (unless your program do noting). An effort to construct a well designed communication between the UI and the core, is always well payed in the future.
I start my design with the core, and in every moment thinking how the core would communicate with a black-box UI. In phylogenetics you have two major parts: the matrix, and the tree. Both are composed of more smaller elements. The matrix is a list of labels, usually with some coupled data (taxon-character matrix, a taxon-distribution matrix). The tree is a topology in which each vertex is associated with the matrix labels.
You could think that both the tree and the matrix are the same, after all, the tree is only an specific way to sort the terminals in the matrix (it seems that TreeFitter[1] works in that way). But there is a difference, the terminals in the matrix are unique, but you could have many tress associated with the same matrix, terminal 'a' are independent in two trees, but they could point to the same element in a matrix.
Apart from terminals, that have a fixed content, internal nodes of a tree are coupled with the same sort of data as terminals, but this is a variable content, dependent on the tree topology. So we could implement a single fixed matrix, and one (or more) variable matrix, when the tree is unused, the variable matrix can be discarded.To communicate the core with the UI, at this moment, I prefer string streams. They could be find in standard library (although with some little spelling differences), in that way the UI always translate between formats to send to the core specific strings, for example the tree in parenthesis format, so the tree construction would not interact with files, or console input, for example.
In that way, the UI is a central of translation from console, file, command dialogs, etc., that can be implemented for different systems (e.g. Win, Linux, command line), but the core remains the same.
[1] Ronquist, F. TreeFitter, program and documentation, available on line at: http://www.ebc.uu.se/systzoo/research/treefitter/treefitter.html
Suscribirse a:
Entradas (Atom)
