This year, the Hennig meeting was held at Riverside, California, and I'm fortunate enough to be able to be there (in great part thanks to the Kurt Pickett award, given by the Willi Hennig Society). As in previous occasions, I will try to give a short review of many of talks!
23 jun
I want to remark the excellent reception given by the organizers, and the opportunity to meet some wonderful heteropteran researchers [which is huge for an heteroptera aficionado! :)]
24 Jun
The meeting start up with a plenary talk given by Quentin Wheeler, with the title of Cladistic cosmology. Basically it is about of building a “Hubble like” project for taxonomy, specifically the description of new species, and facilities to make access to museum specimens across the web, both as metadata, but using some form of virtual specimens (e.g. scans of herbaria types, 3d scans of specimens, etc.). I like his point, but I think he miss the point of sampling: we not only need better understanding of current museum data, but also, we need to go to the field and look for never seen before specimens!
The first symposium was organized by Ward Wheeler about linguistics. The first talk, was a video-conf given by Peter Whitely about the phonetic character definition used for an analysis of Uto-Aztec languages. The actual analysis was presented by Ward, that shows some of the particularities of using phonetic characters in comparison with DNA data, such as having 121 letters (against 4) and short strings (against long sequence fragments). Peter Forster talks about the using of phylogenetic networks for language data, and how it can be associated with what we know about the history inferred from mitochondrial molecular markers. You can found his program here. Joanna Nichols also uses phylogenetics networks with emphasis of using different sources of cognate data. To finish the season, John Wenzel shows a reanalysis of some previous indo-european analysis, showing some problems with the character coding, and how much of the current resolution was a by product of the errors in the codings, and the method used. I do not know much about the state of the field of phylogenetic linguistics (although I like the use of these techniques in linguistics!), I found a nice blog that usually talk about phylogenetics in linguistics here.
The next session was on phylogenomics, and was organized by Gonzalo Giribet, so are heavily biased towards invertebrates, but for a frustrated entomologist like me, it was wonderful :D! First Gonzalo talks about how new advances in sequencing technology has increased the amount of available data for phylogenetics, and how the problem was shifted from phylogenetic analysis to orthology recognition among this huge amount of new sequences. Then Torsten Struck shows how the selection of outgroups affects the position of one most mysterious group of annelids, the Myzostomida. Sebastian Kvist, explores the genomics of the blood feeding proteins used by Hirudinid leeches, and the realtionships of their host bacteria, a group related with nitrogen-fixing bacteria found in plants! Kevin Kocot was interested on orthology detection using phylogenetic trees, so he implement a program to detect the best possible set of ortholog sequences. Björn von Reumont talks about the pancrustacean phylogenomics, and how the acquisition of newer data from Remipedia is key in the problem and that results shows than remipedians are more closely related to hexapodans than to more traditional crustaceas, he also hints the existence of some apomorphies for such group based on nervous system. Vanessa González talks about the relationships among Bivalbia, that as many molluscan taxa, seems to be well known, but their phylogenetic relationships are just barely explored.
Then the poster season starts, most of the posters present a lot of entomological work on heteropterans (specially reduviids), so that makes me so happy, I like most the poster from Michael Forthman as he works with Ectrichodiins the group that I see for an small [very small!] amount of time xD, so hopefully he will construct a nice phylogeny of the group! :) The winning poster was given by Guayang Zhang and it is about the elongation of the front legs in harpactorines.
25 Jun
Students day! Most of students talks where given this day, so it will allow the judges to give the student awards with a complete overview of all the material presented by students.
Ansel Payne gives a wonderful talk about the general relationships of Hymenopteran superfamilies. Rebecca Dikow explores a huge data set (in term of characters) using bacterial genomes, so her trees have brutal lengths over 37 millions of steps! Ross Mounce talks about measuring support using supertrees, he also remind us about the importance of using machine readable data for phylogenetics, and that he thinks that using svg files will be great for publishing data, as you can include a lot of non-published metadata inside the files. Julieta Gallego presents a talk about the stability of molecular clocks (as I am a co-author of the talk, maybe I will talk about this in other time!). John Denton gives a talk about the usage of iterative alignments to improve the searches under likelihood approaches. Greame Oatley speaks about a cool group of birds [give me a passeriform, and I will love it ;)], and their species limits, in a work in part phylogeographic, in part taxonomic. Juanita Rodriguez (from Colombia! But she is working in Utah), talks about the phylogeny and biogeography of a group of Pompilid wasps. Fernando Gelin talks about Polybia, a wasp genus. Federico López (another colombian! And old friend of my from undergrad days! He is now at Vermont) gives a talk about the phylogenetics of a group of Vespids. To finish the first student season, I give my own talk about biogeography of amphibians.
In between student talks, there is an interesting symposium about the species problem. I'm personally no see so much in that problem, but is nice to see a lot of talk about that. Fortunatelly, I guess, although the problem is recognized, most taxonomists and systematist can continue to work, and work well, even if there is no consensus in what an species is. Brent Mischler spokes about the phylogenetic (topologic-monophyletic) species concept. Rudolf Meier talks whats about the problem of discussing the topic, and how most people skip the problem at all, and why he thinks it is important to discuss the subject. For me, this is the best talk of the symposium. Finally Kipling Will talks about the influence of the species definition in alpha-taxonomy, specifically how it is affected by some recent approaches, like bar-coding. Then a nice discussion (which includes Quentin Wheeler) was given by the four panelists of the symposium.
The second student seasson begin with a talk given by Christine Hayes about the the phylogenetic relationships of a group of phorid flies. Anna Dal Molin explore several techniques to visualize large groups of trees. Pei-Luen Li talks about a worldwide distributed (but concentrated in Hawai'i) group of asparagales, and Cecilia Waichert explore some adaptation hypothesis about nesting behavior using a phylogeny of Pomipilid wasps.
26 Jun
Pablo Goloboff, gives his talk about the implementation of Iter-PCR, focusing on the effects of changing the ingenuous algorithms, with the use of algorithms that take previous information into account, and makes feasible the use of iter-PCR in TNT. Jean DeLaet, gives an insight on his future program for phylogenetic analysis, Anagallis, that will include his algorithms for dealing with non-applicable data! I'm very curious about it [hopefully it will be open source!]. Nico Franz speaks about using taxonomy ontologies, although I think that kind of work is nice to discover some particular things, I guess that the missing thing is that they at the moment are no using character data, and I think that the key for this kind of databases will be both explicit references, and explicit data supporting each taxonomic affirmation. Lenka Drabkova talks about the plant cytokinins. Claudia Szumik shows the results of a the search of areas of endemism, using also phylogenetic information, with a huge data set of mammals. Then Santiago Catalano presents GB-to-TNT a user friendly program to build huge matrices, and also some nice taxonomy mapping tools found in the newer version of TNT. Daniel Janies, shows that super-trees are not a real solution to the uses of super-matrices, as they fails to produce approximate results of a super-matrix. In one of the most controversial talks Donald Buth argues that mitochondrial data, as non intrinsic source of data, and susceptible of the host-parasite interactions, must not be used for phylogenetic analysis. Personally I do not agree with him, but a friend of me, tells me that the same concerns have been raised in population genetics (a.k.a. Phylogeography) literature. So maybe, mitochondrial data is not good for small taxonomic levels, by the same reason of gene vs species trees, but at large levels, it will provide a nice source of evidence (as the difference between gene and species trees at this levels are no so important). This is an intriguing line of though, as it is precisely the opposite to the most received view, from at least five years ago. John Wiens gives a review of all of his papers about the use of phylogenies for study ancestral ecology. I was too tired after Jon talk, so I miss the next ones. Later Santiago gives a nice talk about the usage of morphometric data on several and different kinds of studies. Tim Crowe talks about the results of a whole life dedicated to study guinea fowls at different taxonomic levels. Dalton Amorim, shows a new analysis of phorid flies based on morphology, one of the few talks using explicit morphological data! Continuing with morphological data, and dipterans, Torsten Dikow shows its phylogenetic analysis Mydidae, he also shows interest on detecting unstable terminals. Efraín de Luna, talks about the use of morphometric data in phylogenetics. To finish the day, Nico Franz gives a talk about how his analysis of some particular group of curculionids evolve in time from the initial data matrix to the definitive one. I think this is a nice exercise, but ultimately, it is difficult to extract some information about it, except from the anecdotal experience.
Later this day was the banquet, that as is a tradition in the hennig meetings was free for students! The banquet speech was given by Dennis Stevenson, the current editor of the journal of the society (Cladistics), and it was funny, although it was somewhat small!
The winners of the awards? Ok, there is a lot of Marie Stoppes and Kurt Pickett awards for traveling students. The Don Rosen award for the best poster was given to Guayang Zhang, as I tell you before. Cecilia Waichert receives the Lars Brundin award, for the second best talk, and John Denton receives the Willi Hennig award for the best student talk! Congratulation to the winners, all are well deserved, and I think the judge have a difficult time selecting these three, because there are a lot of great talks given by students :)!
27 Jun
I think this is the most difficult day to give a talk: everyone is tired of the amount of talks in previous days, as it is the day after the banquet, a lot of people are dehydrated and overtighted, and as it is the last day, you have to check-out from the hotel...
Here this is the last symposium, this one, about molecular clocks organized by Sean Brady, who discuss several sources of error. Tracy Heath talks about how to model dating in bayesian analysis, and Elizabeth Murray has a more concrete talk on the dating of some group of Hymenoptera.
Robert Sansom presents some differences between results using hard and soft anatomy in vertebrates, and the implications for the fossil record, that is only based on hard parts. Mark Simmons shows some problems from using multiple matrices, with small overlap. He present a lot of interesting results for some particular examples, so it is important to known how general they are. Mari Källersjö gives one of the most beautiful talks of the meeting, about a working project of a particular group of plants. Everyone loves his talk :D. Veronica Pereyra talks about thysanopterans, but unfortunately I have to make the check out, so I lost it :-(. Ronald Clouse talks about the phylogeography of an opilion that lives in southeastern US, that is a remaining of a gondwanian clade (as Florida was long time ago, part of African plate!). Ulf Jondelius, talks about the acoelan worms, a poorly known taxon, that he was sampled across the world. To finish the meeting Steve Farris talks about the recent surge of the three taxon statements authors (Ebach, Williams, etc.).
General overview
The meeting was very well organized by John Hartley and Christiane Weirauch, that move around all the time to make sure everything is right. They have a full group of students that are helping with everything, so it was huge :).
On the meeting itself, it has a lot of students, with several very nice talks, and a lot of discussion, that is always welcomed :D. One of the things that I note, is in the same line of previous meetings in which few morphology is used, here I see that although most studies only use molecules for phylogenetic analysis, people is more interested in link several of their findings with well detailed morphological data, hopefully, in the next few years, fully integrative, genomic and morphological data will be used in conjunction in almost every work!
The next year, the meeting will be held in Rostock, Germany, a city on the coast of the baltic sea. Hopefully I will be able to make it, and several of my cladistists ends will be there :D Also, if any one reading this blog interested in phylogenetics, and living somewhere in Europe, I encourage you to go, even if you don't work with parsimony! [In fact, a lot of people that does not use parsimony analysis shows their results at the meeting! And nobody complains about it, at much, someone as why do you not use parsimony, but just that] So! See you there :)!
Acknowledgments
I receive a lot of funding support from several institutions that make this trip possible: CONICET, FONCyT, GBIF (well, this one in the future) provides monetary funding, whereas INSUE allows me to work at their facilities. I receive a Kurt Pickett travel award from the Willi Hennig Society. At personal level, my advisors (Pablo and Claudia) provides me a lot of support, Julí convince me to go, Santí helps me with the trip logistics, and they, with Ross, they are excellent room-mates :)!
Mostrando las entradas con la etiqueta parsimony. Mostrar todas las entradas
Mostrando las entradas con la etiqueta parsimony. Mostrar todas las entradas
jueves, julio 05, 2012
domingo, agosto 07, 2011
Hennig 30, July 29-August 2, 2011, Sao José do Rio Preto, Sao Paulo, Brazil
Just a week ago I was at the 30th meeting of the Willi Hennig Society (my third in four years! :D). It was in Sao José do Rio Preto, an small city in the state of Sao Paulo, Brazil. There is a lot of people! (More than 200! most of them are students!). Here I will try to give an small review of the whole meeting :) [In the same way as I do in previous ones ;)]
July 29
This was the day of the welcome party, I meet with some good old friends and new people. Also, I'm very tired after a long trip from Campinas to Sao Jose xD.
July 30
The real meeting starts here, with an small presentation of the presidency of the Willi Hennig (by Rudolf Meier) and the organization committee (Dalton Amorim and Fernando Noll). Then a talk by Mario Viaro about the evolution of the prepositión “ate.” It is nice to see some strong similarities between cladistics and linguistics! But I think that most of the talk is based on more or less informed speculation, I hope that new quantitative approaches (like the one show by Ward Wheeler) will invigorate the field!
The first symposium was about bioinformatic tools for the analysis of diseases. Dan Janies, comment of a new way to dealt with they web services of SupraMap, and how it evolves from a monolithic application, to a more piecemeal approach (as unix/linux users know it), it will be more flexible and then more useful (And as a developer, a lot easier to maintain!). If you are curious, its new web site is GisBank.
Then Julián Villabona, a long time friend, give a review of the different tools to understand the geographic spreading of dengue, based on his own previous and current work. I think that the nice lesson is the importance of using new quantitative approaches from historical biogeography, in this clearly phylogeographic framework.
There are two more talks on this symposium, but they are more orientated in bioinformatic questions (data mining, and so), a subject that I do not really feel, so I can take anything about it :P.
The second symposium is on Phylogenetics and conservation. The first talk, of course, was by Dan Faith, that review his PD measure, that I think is the best way to take phylogeny into account in a diversity framework. He shows some way to expand the measure to complementarity, I personally do not like that approach, as it miss one of the most robust things of the PD measure: that is independent of the sampling.
I skip Roseli Pellens talk, but I see the talk of Ronald Clouse about Schizomida of Micronesia, both in terms of phylogeny of the group and at population levels. He found some particular things about its distribution, as it seems that the explanation of its distribution is based on long dispersal, but they are very well separated even inside the small islands. I will be looking forward for his publication on this particular data set :D!
There is a poster season here, but there were a lot of posters, and most of them on very technical issues of particular groups, which is wonderful :D (I love to see biodiversity research) but I'm very ignorant on most of that groups :P
July 31
The third symposium was on biogeography, specially, neotropics. Dalton Amorim shows how the “relationships” among areas shows that there are different historical components on the neotropics, each with different times. Although I don't like the methodology of his work (I'm anti “cladistic biogeography”), I guess that their own study shows that looking for hierarchical patterns in biogeography is not the right route!
Eduardo Almeida talks about the breakup of Gondwana, and the alternatives. I feel he gives much weight to the “expanding earth” hypothesis. It is important to note that this is a fringe hypothesis in geology, in fact, it is almost never mentioned in books on tectonics :P. Actual plate tectonics is one of the our most powerful theories, so the expanding earth require a more compressive mechanism set. Sadly all the discussion on the expanding earth eclipsed the most interesting part of his talk, that is about the biogeographic patterns of Colletidae bees.
To end the symposium, Juan Morrone gives a review of some of the american transition zones, in Mexico and in Patagonia. This prompt an interesting discussion, and I feel, like in the talk of Amorim, that their own research shows that the hierarchic classification in biogeography is not feasible.
Then contributing papers season start. I was very afraid (there is a lot of people xD), but fortunatelly all the discussion on Juan's talk gives me a nice time to keep my mind and provide a nice starting point for my talk ;).
Then Cyrille D'Haese talk about the relationships of a beauty group of colorful and “giant” collembolans of New Caledonia and South west Asia. Fernando Dagosta discuss the relationships of a characid group of fishes (I guess that you can ask Marcos about this talk :P jejeje). Jerome Murienne talk was about niche modelling, and how to integrate them into a phylogenetic framework. I think that this is a very difficult topic, and I was gland that Jerome was very cautions on its presentation ;) as more approaches to integrete phylogeny and distribution modelling are brutally ad-hoc. Denis Machado talk was about the phylogeny of freshwater stingray's tapeworms.
Wei Song Hwang gives a talk about the phylogeny of Reduviidae (specifically Reduviinae subfamily, a non-monophyletic assemblage). It was a terrific talk, on a group that requires an extensive revision although I will prefer more morphological data to be coupled with the molecular data. Anyway, I love this talk. [Side note, I guess it helps that bugs are theo only group that I was on the verge to work with it xD]
Marcos Mirande shows its update of its phylogeny of Characidae--there is a lot of characid papers in this meeting! :)--with new characters, new taxa, and molecular sequences. Owen Davies, who I meet in South Africa, study a group of small african birds of the genus Cisticola, and how old taxonomic revision was good and also wrong, I like this appreciations for old works :D. To close the season, Guanyang Zhang talk about the phylogeny of Harpactorini, a tribe of neotropical Harpactorinae, another grouop of assassin bugs, that, if you live in south America, you can see waiting on leaves with their front legs raised into the air: they have some stiky glands to capture prey, and that was Guanyong main subject.
Another poster season, I just have few time to check it out the other posters :-/, as I have my own, and I was very busy there: I have a pair of very good discussions ;).
August 1
There is another symposium, this one about techniques for molecular phylogenetics. Lone Aagesen explore some different weighting schemes to dealt with full genomic data and explore some groupings detected under different gene duplications among angiosperms. Torsten Dikow about the effects on combining morphology and molecules in Asilidae, as in many cases morphological data set produce an excellent retrieval of actual classification, whereas combined data does not.
Then Ross Mounce give a great talk about how ILD was abused in the literature, and why the way in it was presented in publications made most of the results difficult to repeat--If you read Ross's blog, you will know that he is, rightfully ;), obsessed with result reproduction!--He also shows how there are several ways to extract information on the results of this test, beyond the probability of the test.
As I mentioned before, Ward uses the dynamic homology framework to explore the phylogenetic relationships among uto-aztec languages, I think that this approach provides a real improvement on the analysis of linguistic data, and I am waiting for its publication :D!
Then Mario DePinna gives a talk about ontogeny and rooting. I guess that this talk is most on the same line of the works Nelson and Platnick, that I despise, and I thing are hot topic about 20 years ago, but not now.
Pedro Romano was interested on the use of ordered and unordered multistate characters, and how that decision is not usually well explored in most phylogenetic works. Then Apurva Narechania shows a methodology called “Radical” to explore the data concatenation. I think that his approach can be expanded in several ways, for example, looking for stability after adding taxa (instead of characters), that I feel, is a unexplored field.
Tim Crowe compare the evidence value of nuclear vs. mitochondrial molecular data, and show how their behavior is just not as usually supposed (mitochondria work well on tips, nuclear data work well on deep nodes) in his extensive data set of Galliformes.
Herbert Ferrarezzi explore a way to code polymorphic taxa, instead of using single specimen molecular sequences. Jaime Rodriguez present gives a talk about the use of spectral firms to identify insects, he tries to link it with phylogenies, but I really do not understand how this was done.
The talk of Claudia Szumik was about new character systems used in the classification of Embioptera. I like her work as it is a clear example that in many groups, the “lack” of informative morphological characters, rest on traditional grounds (i.e. taxonomic research centered on some particular character systems), rather than in a really uninformative morphology.
To end the day, Gustavo Hermes speak about a difficult wasp group, the Eumeninae.
The banquet!
This day was the banquet, in a place with a lot of food of several sources: fish, shrimps, sushi, meat (a lot!). It was great ;). And of course there are the Banquet speech (about “addiction on cladistics,” an excellent talk given by John Wenzel), and the Society awards for students, both the travel grants (Marie Stoppes) and the Don Rosen award for the best poster (to Rafaela Lopes), the Lars Brundin award for the second student talk (to Gustavo Hermes), and the Willi Hennig award for the best student talk (to me! I was pretty surprised :D, and really happy! :D). Congrats to the winners :), I'm proud of be in such nice company :D!
August 2
I must admit that I was pretty tired after two days of brazilian joy ;) so, I just barely play attention to several of this talks :P, and loss the talks from Phillippe Grandcolas and Nobuhiro Minaka.
Hilton Japyssu talk about how to codify different behavior patterns, in particular grooming on rodents. They produce a very well supported phylogeny with a lot of this rutines, congruent with several other data sources. Then Carlos Alberts, on the same line work with behavior patterns of Arini parrots.
Later, Santiago Catalano present his method about landmark alignment, and how it produces better results than alignments that not take into account the phylogeny. Another hit for the dynamic homology approach.
One of the talks I want to hear, is the one by Jan DeLaet on character weighting. He was very interested on implied weighting approaches, but it seems that he prefers linear functions instead of concave ones, so he present some properties for a particular function schema that he call, “self calibrated.”
To finish the meeting there is a symposium on theoretical and phylosophical aspects on cladistics. The first talk was by Pablo Goloboff about some improvements to implied weights, that is, to use different weighting functions among different characters, and the possibility to weight blocks of characters instead individual characters (that is mostly for molecular data).
James Carpenter gives a talk about the meaning of homology vs. synapomorphy. Of course, synapomorphy is not homology, as you can also have homology in symplesiomorphy! In the same line Kevin Nixon (who was not there, but send a video of his talk) was on the same line. This might be an old fashion discussion, but given that there is a lot of recent publications by people like Malte Ebach, David Williams and company (they publish a book, and have several papers in journals like zootaxa) it is important to recall this arguments again.
Andy Brower, tries to show how really most of objections on parsimony are not so strong, and that actually, parsimony is not as bad as their criticizers argue: the results are nearly identical with alternative methods. There are several topics I agree with Andy, but there are others that not, mostly he uses some sociological explanations in some cases, but ignore them in others (and of course, I'm very realist :P ejeje, but that is my own problem xD).
The second talk of Jan was about the topics touched on its 2005 paper, that is, about the use of gaps weighting and how they are related to dynamic homology, and of course, why really using everything in 1, is not really a straight forward consequence of using parsimony.
To finish, Steve Farris criticize the most recent attempts of users of 2 taxon analysis approach (already mentioned: Ebach, Williams, Nelson, Platnick), and how they change their own words when new criticisms where made.
Overview
A very long piece here :P! It was a very nice meeting. With a lot of students, that as great, and I think the most important objective of the WHS: to bring students into phylogenetics. But there is disappointment, apart from brazilian students there are few from other countries. I guess that it is a consequence of recent economic downside, as there are just few US students (3 or 4), and just one European students (just Ross). It is shocking when you see such small representation from first world countries. Representation from latinamerica is also scarce: just one from Argentina (were was everybody?), and nobody from other latin American countries. Of course, there are some colombians working on Brazil, but they count as Brazilian students (as its fundings are given by brazilian institutions). I don't know what happens, it is not lack of publicity: there is the same amount as in any other Hennig meeting. It is not the price, after all, this is one of the cheapest international meetings out there!
Nevertheless, that is not a problem of the organization, so I just can tell that the meeting was great, a lot of students, a nice people, and several interesting talks and posters :) Now, it is time to think on the next one, which will be somewhere in the U.S., hopefully I can made it :D!
Acknowledgments
Of course, a lot of fundings is required to do this trip, and I want to acknowledge the support given by the FONCyT, CONICET and the INSUE, and the Willi Hennig Society that gives me a Marie Stoppes award, and the Organization committee that give us a discount rate at the hotel, and everyday lunch ;)! Obrigado! :D
sábado, febrero 12, 2011
Who is a cladist?
Some time ago, someone ask me about which makes cladistics different from other phylogenetic approaches. For most people the answer is straightforward: a cladist is the one that uses parsimony (e.g. [1][2]).
I think parsimony is very important for cladists, but I do not think that you are only a cladists when use parsimony. There is a lot of papers from author that use parsimony, but which I'm not think that are cladistic papers. Also there are a lot of works form author that do not use parsimony, but they are clearly under the cladistic framework (I think on the old one authors: Hennig, Brundin, Wydodzinsky, etc., all of them are cladists, but no one uses parsimony, at least in an explicit [=numerical] way).
So, if it is not an algorithm, where is the difference? When are you a cladists?
A key part of cladistic thinking is monophyly, but monophyly actually is usually understood as a topological term (i.e. just the relationships depicted on a tree), but I think this passage from Hennig [3] is clear:
“The supposition that two or more species are more closely related to one another than to any other species, and that, together they form a monophyletic group, can only be confirmed by demonstrating their common possession of derivative characters (“synapomorphy”). When such character have been demonstrated, then the supposition has been confirmed that they have been inherited from an ancestral species common only to the species showing these characters.”
Then, for a cladist, monophyly is not just about the relationship, but the characters that support such relationship. When a cladist answer the question “Is this group monophyletic?” he/she not only shows a tree, he/she also shows the characters that support such grouping.
Then cladistics is not about the particular algorithm used to infer the relationships, but about the characters that support such groups. It is not estrange then that most cladists used morphology in their analyzes. But even, using just molecular data, you can be a cladist, when referring to each clade you are also discussing the synapomorphies of that clade (either molecular and/or morphological). Then also, you might prefer some particular algorithm over another.
Here is an example, there are two papers, one by Wiens et al. [4] about the phylogeny of Squamates, and other from Beutel et al. [5] on the phylogeny of Holometabola. Both papers use parsimony and bayesian analysis on their data, and both seems to prefer the tree resulting from the bayesian analysis. Both are about the same size (in pages), and discuss highl level relationships of a largely diverse group using both molecular and morphological data sets. One of this works is clearly a cladistic work, and the other not.
If you look to Beutel et al. paper, you will find a lot of discussion about the characters that support each clade. For them, support is not just the Bremer' support value, jackknife frequency or Bayesian probability, they are important of course, but they are meaningless without an explicit reference to the characters on the node. So, even if they prefer Bayesian analysis, they continue to be cladistis.
On the other hand, you have de Wiens et al. paper. For that authors, just the relationships are important, they discuss groupings and bayesian probabilies, but no discussion on any explicit character for any group is made. Then, even in the case of Wiens et al. only use the results of parsimony, they are not cladists: characters are not important for them.
A cladist is not the one which write a paper with the more up-to-date search strategies, or support measures, or larges amounts of new molecular data, or parsimony analysis. All of that things are important. But in a cladistic paper, the important think are the characters that support the phylogeny. A good cladistic paper is a paper about the characters.
References
[1] Felsenstein, J. 2001. The troubled growth of statistical phylogenetics. Syst. Biol. 50: 465-467. Doi: 10.1080/10635150119297 [available free at Syst. Biol. site]
[2] Williams, D.M., Ebach, M.C. 2007. Foundations of systematics and biogeography. Springer, New York (USA).
[3] Hennig, W. 1965. Phylogenetic systematics. Ann. Rev. Entomol. 10: 97-116. Doi: 10.1146/annurev.en.10.010165.000525 [available free here]
[4] Wiens, J.J. et al. 2010. Combining phylogenomics and fossils in higher-level squamate reptile phylogeny: molecular data change the placement of fossil taxa. Syst. Biol. 59: 674-658. doi: 10.1093/sysbio/syq048 [available free at Syst. Biol. site]
[5] Beutel, R.G. et al. In press. Morphological and molecular evidence converge upon a robust phylogeny of the megadiverse Holometabola. Cladistics. Doi:10.1111/j.1096-0031.2010.00338.x [available free at Cladistics site]
lunes, octubre 19, 2009
The behemot! for free! / El behemot gratuito!
Santí noted that the "behemot"'s paper [1] is now available for free from Wiley/Blackwell web-page, this is the link:
Or you can jump directly to the paper abstract and download the pdf:
If you want to look the results, Pablo post a quick guide to manage the trees and extract info from them. If you want to check you favorite group, or just play for a moment, there are no excuses ;)!
And the best of it, as they are macros, you can use for your own trees! and don't forget to check out the TNT wiki!
Santí me hizo notar que el paper donde el "behemot" vio la luz [1] se encuentra disponible de manera gratuita en el sitio de blackwell. El enlace es este:
O pueden ir directamente a la pagina del artículo y descargar el pdf:
Para quienes quieran jugar con los resultados, Pablo publicó una guía rápida para manejar los árboles y poder extraer info de ellos. Así que si tienen curiosidad por ver como quedo el grupo que ustedes trabajan, o simplemente quieren pasar un rato, ya no hay excusas ;)
Y lo mejor de todo, como son macros, se pueden usar en sus propios resultados :D! No se olviden de consultar la wiki de TNT ;)
[1] Goloboff, P.A. et al. 2009. Phylogenetic analysis of 73 060 taxa corroborates major eukaryotic groups. Cladistics 25: 211-230. DOI: 10.1111/j.1096-0031.2009.00255.x
lunes, abril 27, 2009
A phylogeny of 73060 eukaryotes
Finally, the behemoth has seen the light :). Our paper with a parsimony analysis of 73060 eukariotic species (and 7800 mol+morf characters) was just published (as “online early”) in Cladistics [doi:10.1111/j.1096-0031.2009.00255.x].

Pablo does a wonderful work optimizing every aspect of the tree-searches in TNT. And all of the guys worked really hard to manage that amount of data!
At first I was surprised with the high accuracy of the trees founded, because the data set is full of missing entries. Also, I fill happy because the inclusion of morphological data, even at this huge scale, produce better results than molecules alone!
Just few months ago, this was posted in dechronization:
He [Cassey Dunn] makes a convincing case for the idea that a revolution in analytical techniques will be needed as we enter an era during which computational capabilities will be more limiting than data availability.I think our study shows exactly the inverse: that our actual search capabilities are good enough, but we do not have sufficient data (the largest gene set is SSU with 20000 species, and a handful of genes has more than 10000 species).
The second lesson?... We do not need super-trees!
Etiquetas:
molecular phylogenetics,
morphological phylogenetics,
parsimony,
software,
tnt
viernes, abril 24, 2009
Algorithms for phylogenetics 0a: Bitfields
Intro
I want to start a series of post about the algorithms used in phylogenetic analyses. I feel a bit disappointed with the entry on computational phylogenetics (and related subjects) in wikipedia, that are somewhat biased towards model-based methods and "bioinformatics", with a poor representation of algorithms for phylogenetics.
There are two outstanding examples of algorithms for phylogenetics on the web. The book of Wheeler et al. [1], and, in a similar vein, a manuscript by DeLaet [2]. But their presentation of the algorithms is somewhat general, that is be nice for educational purpouse, but in some cases, far away from the actual implementation in computer software.
As the presentation of the algorithms in [1] and [2] is excellent, I want to focus on the implementation. The idea is to go from simpler to more complex algorithms (as far as I can comprehend them xD).
I want to acknowledge Pablo Goloboff. He always have the time to discuss with me about algorithms (and any topic on methodological phylogenetics) :D. Of course, any error in my implementations are my own errors!
As this series progress, I will post the source code on google source or sourceforge, if someone wants to check it, requires an knowledge of C programming language. Thanks to Rubén who showme some tips with the HTML :).
Bitfields
Characters are a basic component of phylogenetic analysis. For simplicity assume that we only have a single character in each terminal. Each character can be stored into a variable, assigning to them 0, 1, 2... as indicated by the character state. This idea is implicit in the original description of Wagner's "groundpland divergence" [3], and in Farris' formalization [4, see 5]. Then operations between characters would be operations of ranges. But it is more simpler to think about characters as a set of states [4, 6]. Then operations between characters would be set operations. This set has a peculiar property: they are perfectly boolean, that is, a taxon has or has not a particular state. This can be translated in two positive consequences: (i) character states can be stored as a bitfield: (ii) character operations are bifield operations.
Computers store numbers as binary numbers. A bitfields use this characteristic to represent a set of positioned bits. This example shows it better:
State binary representation "Number" (human)As can be seen, each state has his on bit position, and the states combination is just the union of its states. Many programming languages include operations to work directly on bits (not, and, or and xor), then the translation from set operations to bit operations is direct.
0 0001 1
1 0010 2
2 0100 4
3 1000 8
0, 1 0011 3
1, 3 0101 5
In this series I will use char to store character states. In C, char has 8 bits, then only 8 states can be stored. I define the bitfield with a typedef, so changing the number of bits on the bitfield (and then the number of states) is easy.
#define BITS_ON_STATE 8This small example reads the caracters of a taxon, from a file (in), and store it into an array of characters (chars). The number of characters is nc, and the functions getc (in stdio.h) reads a single character from the file, SkipSpaces ignores the blank characters, and isdigit (ctype.h) detects if the character is a number.
#define ALL_BITS_ON 255
typedef char STATES;
for (j = 0; j < nc; ++ j) {
error = SkipSpaces (in);
if (error != NO_ERROR) return error;
c = getc (in);
if (isdigit (c))
chars [j] = 1 < < (c - '0');
else if ((c == '?') || (c == '-'))
chars [j] = ALL_BITS_ON;
else return BAD_FORMAT;
}
The operation chars [j] 1 < < (c – '0') substract the ASCII value of '0' (30) from the ASCII value stored in c (a number, so it is form 30 to 39). The resulting value is used to shift the bits of 1. For example if a '0' is readed the operations are:chars [j] = 1 < < (30 – 30)If a 5 is readed
chars [j] = 1 < < 0
[0000 0001 < < 0 = 0000 0001]
chars [j] = 1
chars [j] = 1 < < (35 – 30)The 1 is shifted five bits to the left.
chars [j] = 1 < < 5
[0000 0001 < < 5 = 0010 0000]
chars [j] = 32
Note that inapplicables and unknowns are coded with all bits on.
This work nicely with non additive characters. For additive characters, it is possible to “recode it” as several binary characters [7][8], then it is possible to use the character as independent binary characters.
There are more sophisticated usages of bitfields, packing several characters in a single variable (for example, eight 4-bit characters can be stored into a singel 32-bits variable) [8][9]. This packing allows an speed up during searches, because several characters can be examined simultaneously.
References
[1] Wheeler, W. et al. 2006. Dynamic homology and phylogenetic systematics: a unified approach using POY. American Mus. of Natural History, published in cooperation with NASA. online: http://research.amnh.org/scicomp/pdfs/wheeler/Wheeler_etal2006b.pdf
[2] DeLaet, J. 2005. Pseudocode for some tree search algorithms in phylogenetics. Manuscript online: http://www.plantsystematics.org/publications/jdelaet./algora.pdf
[3] Wagner, W.H. 1961. Problems in the classification of ferns. In: Recent advances in botany. Toronto Univ. Press. pp. 841-844.
[4] Farris, J.S. 1970. Methods for computing Wagner trees. Syst. Zool. 19: 83-92.
[5] Goloboff, P.A. et al. 2006. Continuos characters analyzed as such. Cladistics 22: 589-601. doi: 10.1111/j.1096-0031.2006.00122.x
[6] Fitch, W.M. 1971. Toward defining the course of evolution: minimum change for a specific tree topology. Syst. Zool. 20: 406-416.
[7] Farris, J.S. et al. 1970. A numerical approach to phylogenetic systematics. Syst. Zool. 19: 172-189.
[8] Moilanen, A. 1999. Searching for most parsimonious trees with simulated evolutionary optimization. Cladistics 15: 39-50. doi: 10.1111/j.1096-0031.1999.tb00393.x
[9] Goloboff, P.A. 2002. Optimization of polytomies: state and parallel set operations. Mol Phyl. Evol. 22: 269-275. doi: 10.1006/mpev.2001.1049
Etiquetas:
algorithms,
parsimony,
program design,
programming,
software
sábado, diciembre 22, 2007
Homology and parsimony
Today is a transaltion day ;). This post is a translation of my previous post in spanish about the subject.
This post is somewhat inspired in a discussion by Ebach and Williams in their blog systematics & biogeography (they had many post about the subject!). Here I want to show the relationship between parsimony (or cladistic analysis) and homology.
Homology
It is a long tradition of discussion about the definition of homology. I use a definition similar to the traditional one, then two structures are homologs when both are considered the same structure in different organism. There are countless arguments to acept two structures as being the same: God's plan, natural order, morphotype, or the one I endorse, by common ancestry.
If two structures are homologs, it implies that they are the same structure, and they are inherited from a common ancestor of the two examined organims. Inheritance implies several things by definition: specially from genetics and development biology. Moreover, the structures could not be highly similar, even if they are the same, because they can be strongly modified. Nevertheless both are the same, so it implies that structures change across the time, in our assumptions include several process that promote the differentiation, such as genetic interaction and population dynamics.
Cladistic analysis do not have a direct interest in these phenomena associated with homology. The process are of interest in other fields like evolutionary biology, development biology, genetics, molecular biology, population genetics, ecology, etc. But lack of interest do not imply that they could be ignored! In the character definition for a cladistic analysis (i.e selection of homolog characters) several of those factors would be taken into account to propose character limits and character codification.
Unfortunately, as process is not of direct interest for cladistics allows the erroneous idea that the process is irrelevant in character definition. The the so called pattern cladists (as Nelson and Platnick [1] and recently Pleijel [2], Brower [3] and Ebach and Williams) to think that cladistics is free of evolutionary thinking, but it is a totally wrong position: the use of homolog characters implies a framework based on origin by common ancestry, the every character--and not only its apomorphic state--are the same structure [4, 5].
But it is more, implications of evolution are big for several characters, specially if there is a well knowledge of the character, the the idea of Kluge [6] who claims that only assumption of cladistic analysis is 'descent with modification' is also wrong. When you include inheritance and evolution, many things are included in the definition of character. Of course, for some characters we only got a morphological knowledge of the character, then assuming a simple 'descent with modification' seems to be correct, but for other characters (as in many vertebrates) we got knowledge from development and genetics of the structure, in such cases the assumption are far more complex than descent with modification.
Parsimony: the algorithm
The basic principle of parsimony algorithm is fairly simple. If you want to know the character state in the node 'x', which descendants are 'y' and 'z', and we know their character states, then the character state of 'x' is the intersection between states of 'y' and 'z', if the intersection is void then the state of 'x' is the union of 'y' and 'z' states. Using an union implies a character change (a step). Of course, optimization of states is more complex, but for the present discussion we only need the basic steps.
State assignation using the parsimony algorithm: (A) and (B) the state of ancestor 'x' is equal to the intersection of descendants 'y' and 'z', i this case, white; (c) 'y' and 'z' do not share any state, then the state assigned to 'x' is the state union of his descendants.As you see, the algorithm is 'independent' of the data used, you can use any kind of character (in computer science this problem is know as the coloring problem, and the 'characters' are colors of a map) and any kind of terminal. It is a common character of every method formalized in an algorithmic form. Then it is necessary to provide a proper justification to use the algorithm in a particular problem.
Parsimony: cladistic analysis
With the evolutive concept of homology, its union with the algorithmic parsimony is direct. If a character, no matter its state, is the same between two organisms that share a common ancestor, it implies that the character is inherited from the common ancestor, then the common ancestor would have the character.
Moreover if both organisms share the same form of the character, that is the same state, then that state would be in its common ancestor (the character is the same!), but if both organisms had different forms of the characters, we do not know which form would be present in the common ancestor, so we assume that it could be any of the both states, in that case, if both organisms really share the same character then a transformation would be happen.
This is exactly the same description of the parsimony algorithm. In cladistics the use of parsimony algorithm is justified because used characters are homologs. In this context the algorithm maximize our homology propositions when minimize transformations: this allows that most terminals with the same states would be contiguous. Then the hypothesis ad hoc of homoplasy are minimized [7].
Parsimony and homology tests
It is a common idea between cladists to say that parsimony is a test of homology: congruence. O disagree, because as I argument here the basis of parsimony is assuming from the very beginning that character are homologs! Any homology test would be previous to a rigorous cladistic analysis.
Homology test could have several forms, they could be morphology arguments (usally put under 'similarity' label), anatomical position, structural organization, ontogeny, genetics, and in most cases the 'test' is a conjunction of these procedures--for these reasons defining a character implies an strong theoretical background--. After examining all of those alternatives you got a good character. Is for these reasons that homoplasy is an ad hoc hypothesis: homoplasy is only justified in realtion with the cladogram.
Of course, character revision is always welcomed, and homoplasic ones maybe demand a close examination, but it is equally valid to examine every character. It is possible that codification form some characters is dubious, in such cases it is possible to use, as an exploratory devise different codifications (similar to the proposition of Ramirez [8] for morphology, and Wheeler [9] for molecules). But beware: this codification is not supported by the cladogram, because the argument used to defend that codification is the same used for homoplasy: it is justified only in relation with the cladoogram, the the codification is an ad hoc codification. In ambiguous cases I prefer a weighting schema as proposed by Neff [10]: because we know little about the character, and we have some doubts about its coding, it is better that it has a lower weight than characters that we know better.
**
Bonus: an historical speculation
He I show homology and parsimony ideas in a separated fashion and then I fuse them. I do it in that way to clarify the argument. But historically the development is intertwined from the beginning. If you read Wagner [11] in the algorithmic pathway, and Hennig [12] from the logical point it is clear that both positions are very close. Both visions were fused in an excelent fashion by Farris and its collaborators [7, 13, 14], which ideas (especially from [14]) agree in many points exposed here. Then from the beginning cladistic analysis and the parsimony algorithm walk together.
Wagner, Hennig and Farris development their ideas from a morphology context. At the same time Dayoff [15] experiment with several algorithms for molecular sequences, which at least today, homology ideas for molecular biologist are different to the morphological concept, I do not know what homology ideas used molecular biologist form 60s, but it seems that she did not believe that two bases (in case of Dayoff, two aminoacids) equal in two organisms imply common origin, the idea of point mutations precludes the idea. It is worth to note that Dayoff never could find a way to assign states in ancestors.
References
[1] Nelson, G., Platnick, N. 1981. Systematics and biogeography. Columbia Univ., New York.
[2] Pleijel, F. 1995. On character coding for phylogeny reconstruction. Cladistics 11: 309-315.
[3] Brower, A.V.Z. 2000. Evolution is not a necessary assumption of cladistics. Cladistics 16: 143-154.
[4] Fitzhugh, K. 2006. The philosophical basis of character coding for inference of phylogenetic hypothesis. Zoologica scripta 35: 261-286.
[5] Grant, T., kluge, A.G. 2004. Transformation series as an ideographic character concept. Cladistics 20: 23-31.
[6] Kluge, A.G. 2003. On the deduction of species relationships: a précis. Cladistics 19: 233-239.
[7] Farris, J.S. 1983. The logical basis of phylogenetic systematics. In: Platnick, M., Funk, V.A. (Eds.), Advances in cladistics, vol. 2. Columbia Univ., New York, Pp. 7-36.
[8] Ramirez, M.J. 2007. Homology as a parsimony problem: a dynamic homology approach for morphological data. Cladistics 23: 588-612.
[9] Wheeler, W. 1996. Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics 12: 1-9.
[10] Neff, N.A. 1986. A rational basis for a priori character weighting. Systematic zoology 35: 110-123.
[11] Wagner, W.H. 1961. Problems in the classification of ferns. In: Recent advances in botany, vol. 1, Univ. Toronto, Toronto, Pp. 841-844.
[12] Hennigh, W. 1966. Phylogenetic systematics. Univ. Illinois, Urbana.
[13] Kluge, A.G., Farris, J.S. 1969. Quantitative phyletic and the evolution of anurans. Systematic zoology 18: 1-32.
[14] Farris, J.S., Kluge, A.G., Eckardt, M.J. 1970. A numerical approach to phylogenetic systematics. Systematic zoology 19: 172-189.
[15] Dayoff, M.O. 1969. Computer analysis of protein evolution. Scientific american 221: 87-95.
viernes, diciembre 14, 2007
Really huge news about TNT
I'm out of the city, so no computer for long posts :'(... but I'm happy to give a flash news about TNT: The Willy Hennig Society takes the sponsorship of TNT (a software by Pablo Goloboff, Steve Farris and Kevin Nixon), so from now, the program is free! The are some simple conditions: personal use, and a citation of the program--and the sponsor, that is the WHS ;)--in published results!
In case that you don't know, TNT is the most faster program for phylogenetic analysis under parsimony, implements several new and efficient heuristic algoritms [1,2], and a powerful script/macro language. If you are doing cladistics/phylogenetics, you should surely dream with this program!
I love the matrix editor! Is easy to use, and more straightforward than WinClada, NDE or Mesquite (yep!... far better than mesquite!).
You could download it at: http://www.zmuc.dk/public/phylogeny/TNT/
read the license agreement and enjoy :)
I give the proper citations when I return to Bogotá :P
[1] Nixon, K.C. 1999. The parsimony ratchet a new method for rapid parsimony analysis. Cladistics 15: 407-414.
[2] Goloboff, P. A. 1999. Analyzing large data sets in reasonable times: solutions for composite optima. Cladistics 15: 415-428.
In case that you don't know, TNT is the most faster program for phylogenetic analysis under parsimony, implements several new and efficient heuristic algoritms [1,2], and a powerful script/macro language. If you are doing cladistics/phylogenetics, you should surely dream with this program!
I love the matrix editor! Is easy to use, and more straightforward than WinClada, NDE or Mesquite (yep!... far better than mesquite!).
You could download it at: http://www.zmuc.dk/public/phylogeny/TNT/
read the license agreement and enjoy :)
[1] Nixon, K.C. 1999. The parsimony ratchet a new method for rapid parsimony analysis. Cladistics 15: 407-414.
[2] Goloboff, P. A. 1999. Analyzing large data sets in reasonable times: solutions for composite optima. Cladistics 15: 415-428.
Suscribirse a:
Entradas (Atom)
