Just two weeks before, a seminar about biogeography and phylogenetics starts here, in Tucumán. We were about 8 people taking about a biogeographic paper (the next week, it will be a phylogenetic one). So there are my own impressions about the papar (I posted it late, because I'm out of the town because of the of the holly week break).
Sanmartín, I., van der Mark, P., Ronquist, F. 2008. Inferring dispersal: a Bayesian approach to phylogenetic-based island biogeography, with special reference to the Canary islands. J. biogeografy 35: 428-449. DOI: 10.1111/j.1365-2699.2008.01885.x
Personally I do not like Bayesian methods in phylogenetics, they are based on several flawed principles [1], but I will try to focus in the other aspects of the proposed framework of Sanmartín et al., instead of attacking the Bayesian principles.
As many others, Sanmartín et al., attack 'vicariance' biogeography on the grounds that they only counts vicariance and ignore dispersal, I think that the tale is somewhat different, vicariance biogeographers are fully aware of dispersal, but they can't find grounds to found a general dispersal pattern, in the other hand vicariance provides a fully explicable pattern that is shared by several and unrelated organisms, then dispersal, can only be addressed group-wise. But in the case of the islands without any connection to land, vicariance can not be used as a general explanation. But can dispersal provide one?
Sanmartín et al., argue that the answer is 'yes', but instead of show why such conclusion arises they jump to their own model of dispersal, the causal mechanism that put several different organisms in the same model is never given, then their own asseveration that dispersal can produce 'concerted' (i.e. general) patterns, which is the most interesting question raised by they are never answered. What is striking is that they show several different models of dispersion proposed for canarian taxa.
Moreover the model developed, is a symmetrical one, then if in fact it is a common dispersal route, by an oceanic stream for example, it can be masked by their model that presupposes that both dispersal directions are equally probable.
Then they jump on the model frenzy, they are as embedded on the power of model approaches, that they raise 'new questions', that were irrelevant before their model. The most striking one is the 'carrying capacity' of their model. Carrying capacity is a term borrowed from ecology, and in this case is somewhat equated with island raw richness. This parameter, that seems to be originated from the 'island biogeography' of Wilson, can be of interest to ecologist, but they relevance in the search of common biogeographical patterns is never clearly showed (and I do not think that it has some empirical meaning), but now, this is an important question: most of Sanmartín et al. discussion, are about carrying capacities, not about dispersal models, which taxa are best fitted to the model, dispersal rates for different groups, or common routes of dispersal. Sanmartín et al. transform the parameter calibration into the research answer, losing the important questions in the middle.
When given the results, Sanmartín et al. contradicts many of its initial reasons to choose a likelihood model, the first one is that the likelihood of each model allows a selection of the models, but as the model (and its parameters) with the best likelihood seems to be illogical, they move on a sub-optimal models, they claim that it is not enough data, but when is enough data? How can I selecet a model only guided by its unexpected results? Why no limit the model to certain values if we are afraid of illogical answers? I think that Sanmartín et al. Results actually undermined their claims.
Sadly the paper is more like a discourse from a politician in campaign: all about the promises of wonderful results that might be open for likelihood models, but the actual use of the model with the empirical case is disappointing.
If a dispersal biogeography based on 'concerted' patterns of dispersal is to be used, their user would answer several questions: (1) why expect a 'common ' dispersal pattern? (2) is their model adequate to the desired answer? (3) Is the pattern only spatial (i.e. a common dipsersal route) or spatial and temporal (i.e. a common dispersal route, roughly at the same time)? (4) What about the taxa that departs from the selected model? Questions (1) is never answered by Sanmartín et al., question (2) is a clear no (the symmetrical model contradicts directly their aim), maybe they can give an answer to question (3) but they are more interested in parameter estimation than in biological questions, the question (4) is never answered, even when they accept that there are several different models proposed for different taxa in the Canarias.
[1] Goloboff, P.A., Pol, D. 2005. Parsimony and Bayesian phylogenetics. In: Albert, V.A. Ed. Parsimony, phylogeny, and genomics. Oxford univ., Oxford, pp. 148-159.
lunes, marzo 31, 2008
viernes, febrero 29, 2008
A new look!
It is really nice! The page of the Willi Hennig Society has a new design, it is far better than the old one :)... I hope, some new features (like the data matrices, additional data, comments for papers, an open manuscrips storage, blogs from leader cladists...) sooner, but the first step is cool!
Another page was launched, but by now I do not see it :P, is the Encyclopedia Of Life (EOF, this is a horrible acronym, I see it as 'End Of File'), I read the review of Rod Page (see also the review of Chris Taylor), then I hope that many of problems showed by Rod will be solved sooner (they have 50 millions of US$ for 5 years!!)
And the site for the XXVII WHS meeting (at Tucumán, my new home town :D!) would be launched sooner, stay tuned! ;)
Another page was launched, but by now I do not see it :P, is the Encyclopedia Of Life (EOF, this is a horrible acronym, I see it as 'End Of File'), I read the review of Rod Page (see also the review of Chris Taylor), then I hope that many of problems showed by Rod will be solved sooner (they have 50 millions of US$ for 5 years!!)
And the site for the XXVII WHS meeting (at Tucumán, my new home town :D!) would be launched sooner, stay tuned! ;)
miércoles, enero 30, 2008
Taxonomy: its time to change!
I just read a post by Christopher Taylor about a 'taxonomic problem' with Drosophila melanogaster. The curiosity of the case prompt me to made this post, that I have in my mind for a long time.
One of the main advantages of a classification is the information retrieval. Biological classification allows to find relationships and characters of a determinate specie. There is a set laws to manage the classification, know as the codes (they are different codes for animals, plants, bacteria, etc.).
In this age of intent and massive databases, one could thing that taxonomy are ready to made a direct jump, but unfortunately that seems not to be the case. The codes are really old ones, and many taxonomist are afraid of changing the rules would collapse the system.
I do not think that taxonomist would leave the data designers to propose a new system, but I think that a deep collaboration, coupled with some changes in the code structures will speed-up taxonomic research. The main objective of taxonomist is to describe and classify species, not to be lawyers!
I identify some problems that I think that be are the most important.
Linnean names
Linnean names are nice, they provide a some form of order in a time of great explorations around the world (XVIII-XIX century), when europeans became surprised with the diversity around the world. Many things are change from that time. By now we have several phylogenetic methods, which show how arbitrary a clade name would be (I like this post about the subject), and of course, the tree of life have far more divisions than Linnean ranks. Now we have databases to store names and search algorithms and find it quickly.
A rank-free taxonomy seems to be more adequate to store our information about the phylogeny of species. That is not an embrace of 'phylocode', as I prefer a taxonomy based on specific characters, or better a combination of characters and topology, instead of a topology based taxonomy (note that apomorphy based definition are also based on topology). Also, there is an special utility from ranked names: they can serve as Landmarks when browsing the tree of life.
The nested nature of biological classification allows a rank free taxonomy without the pains imagined by [1], and using “landmarks” helps in search--and writing--of abstracts, titles of papers, and of course, the search of a particular clade, as you can see using GenBank. Of course, genus and species can be (and I think would be) of obligatory nature. That allows a continuity with Linnean taxonomy.
If linnean names are more landmarks, then many of the laws for synonyms, and so one, need to be changed to a more practical usage.
Synonyms
Perhaps, the most annoying characteristic of taxonomy is the synonymy, rank changes, and several related problems. This problems are actually a burden of actual codes, and can be changed with changing the laws, without harming the actual classification.
Synonyms are really bad for databasesas the same entity is labeled with two different names, or (worst!) a same name is applied to different entities, in other cases, the range of the synonymy seems to be overlapping.
Of course, no matter what phylocoders say, the same problem applies to their 'phylogenetic' names: as new revision is published, the names continue but with completely different meaning. At least in traditional taxonomy, it is possible to reject some names.
It seems better to relax some of the naming rules, so a new classification would be clearly different from the original one. If a huge family is discovered to be massively paraphyletic, I think that it is no point allowing the original name to survive, it is simply an brutally wrong name.
For example, “reptiles” for a long time include many steam ammniotes, therapsids “mammal-like reptiles”, anapsids (as turtles), lepidosaurs (lizards and snakes) and crocodyles. They are amniotes but exclude birds and mammals. I can see a reason to retain that ugly name. It is simple, synonymize it with their monophyletic equivalent (Amniota), and never more use it. There is no way that a modern paper allows some confusion with the old ones. The only valid use of the name is when someone shows that the original reptiles, are monophyletic.
Names can be used instead to show different possible classifications. For example, the different arrangements of Arthropoda receive different names: Atelocerata (Myriapods, like centipede, and insects) vs. Pancrustacea (Crustaceans and insects), Mandibulata (Myriposd, crustaceans and insects) vs. schizoramia (crustaceans and arachnids). These names identify different entities, each one attached to a different phylogenetic proposal (note that form phylocoders, atelocerata, pancrustacea and mandibulata can be the same entity!).
Types
The reptile and arthropod example, are possible because there are no types fixing the names of that groups. At family levels, family names, or genus names are ruled by typification of names. Which allows more confusion, solutions to actual research.
There is an example: Lygaelidae was a large family of bugs (Insecta: Heteroptera), for long time it was believed that it was paraphyletic [2], but this was only demostrated by Henry [3]. He propose a new whole classification of Lygelids, elevating to family range no less than 7 subfamilies, and restructuring the meaning of Lygaelidae to only 3 subfamilies. Is the new Lygaelidae the same of, say 20 or 30 years ago? Of course not. Then typification was creating name stability (as is created by topology by phylocoders) but at cost of the loss of name utility.
As in the case of reptiles, there are no Lygaelidae any more. Any new reference i the litarature to 'Lygalelidae' only confuses with the initial meaning of the group (which is a synonym of the superfamly Lyageoidea). Of course you can use 'Lygaeidae sensu Henry' but it is only a clumsy (and error prone) way to give a new name.
Another nice example was provided by the previously mentioned post of Chris. It is about Drosophila. In this case, the usage can be against the taxonomic practice. For m, the solution is simple: no more Drosophila. But as there is a huge number of users of the name, that surely don't care about taxonomy, that outnumbered the number of Drosphilic taxonomic publications, there is when a commission can rule. The practical solution is to maintain Drosophila to the molecular people. Then solutions in cases of conflict would be guided by practical options, rather than some old described type.
This, of course, allows to made classification changes without 'using the types' (as far as the original characters were examined!). And free taxonomist to depend on some poorly known species (or even, specimen) to nominate genera and families.
Speaking about types...
Type specimens have a particular property: they are in the first world, but they are collected in the third world. The reason for this is historical, but their consequences are seen more acute today. Museums are measured by their amount of 'type specimens', and there are particular politics to borrow that specimens (only borrowing one on time, certifications, curators permission...). Also, some taxonomist, specially from the old past, simply nominate a wonderful amount of new species, only giving a superficial description (some color, some illustrations of genital parts) and based the whole 'description' on a type species designation. This old practice continued in several obscure papers in the third world.
Then typing, although seems to be a reasonable way to be objective, is more harming I guess that typing will never gone, but at the moment there are several nice perspectives to free from borrowing politics. Approaches like [4] with a great emphasis on characters and images, can change the situation. Researches far away from type specimenes can see high quality pics of several specimen parts. A side consequence is that the concept of type becomes lost, it is impossible to pic every part from a single specimen, and is possible that it ends destroyed, then the new typing would be more responsible, as it would be based in several different specimens.
Moreover a destypification increase general collections value, that is more the quantity and quality (e.g. fresh specimenes) of material available, than a particular specimen collected in 1816, saved from a fire in 1874, harmfully damaged by bad curation in 1903...
Actually there are many phylogenetic work without using type specimens, that is, the major bulk of molecular phylogenies, and I guess several morphological ones. I think that they do it in a very objective and testable way. If they can live without types, why classical taxonomist do not?
The data matrix
A second question from the previous section, is how a non-typified research can be objective? The answer is that instead on focusing on a particular specimen, phylogeneticists use a data matrix of taxon an characters.
A non type taxonomy enforce the use of well delimited characters, it is the only way to show the reality of the new designation. Look at some recent revisions with a phylogenetic analysis, and compare it with a revision without it (for example some of both see the pubs of AMNH). The character matrix allows to a quick examination of several characters, it is possible to see which state each character has in each taxon. New technologies (see [4]) couple specimen, characters and images for each cell entry. By default a matrix provide a multi-entry key, the identification tools are better.
There are some nice things of using a character matrix. The first, is an increasing interest in provide well defined characters [4]. As character are used for phylogenetic analysis, they would be stricter. Other characteristics like color patterns, length measurements would be restricted to a simpler description. Another advantage is that it provides a quick classification of a new species.
A nice real example was provided with the dinosaur paleontologist researches. The y publish some quick and small reports in high profile journals (like Nature or Science) with small descriptions, but as they have a great database of characters, several points of the anatomy of the new described fossil are immediately 'published', long before the detailed description in a more specialized journal.
Thinking on databasing
Of course, using a character matrix is direct consequence for storage: well defined characters and images enter smoothly in a database [4].
I think that the new challenges of the 'biodiversity crisis' as well as the 'taxonomic crisis' can be solved with a thinking of data storage. How can we store the data more efficiently? How can we link taxonomic and publication data? How changes in our knowledge about phylogeny could change the previous publication data, how the harm can be minimized?
It is important to a new taxonomy to keep the great advances made from Linneaus times, a start from the scratch is clearly a wrong solution. But also taxonomist would be able to made some concessions in their practice, and update it to new data architecture of the world.
It is time that taxonomy became a useful discussion about actual data, facing the massive extinction that the man is producing around the world, it seems weird that a taxonomic study would began searching for old papers from XVIII century, which only utility is that they provide a name, descriptions, characters and other things from that papers are of low value (by the way, taxonomy is the only field of science that continue using such old data. For historians old text are the source of investigation, for taxonomist is a more lawyer-like activity of searching for an 'old case'. Catalogs are some nice curiosities, and surely valuable for historians, but what is their actual value for taxonomist? They are important only because points to papers that establish a name).
If taxonomy is the main objective, then useful data storage is the main objective. Book keeping and law courts are not part of knowing biodiversity.
References
[1] Dominguez, E., Wheeler, Q. 1997. Taxonomic stability is ignorance. Cladistics 13: 367-372. doi: 10.1111/j.1096-0031.1997.tb00325.x
[2] Schuh, R. T., Slater, J. A. 1995. True Bugs of the World. Cornell Univ. New York
[3] Henry, T. J. 1997. Phylogenetic analysis of family groups within the infraorder Pentatomomorpha (Hemiptera: Heteroptera), with emphasis on the Lygaeoidea. Annals of the Entomological Society of America 90: 275-301
[4] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic Biology 56: 283-294. doi: 10.1080/10635150701313848
One of the main advantages of a classification is the information retrieval. Biological classification allows to find relationships and characters of a determinate specie. There is a set laws to manage the classification, know as the codes (they are different codes for animals, plants, bacteria, etc.).
In this age of intent and massive databases, one could thing that taxonomy are ready to made a direct jump, but unfortunately that seems not to be the case. The codes are really old ones, and many taxonomist are afraid of changing the rules would collapse the system.
I do not think that taxonomist would leave the data designers to propose a new system, but I think that a deep collaboration, coupled with some changes in the code structures will speed-up taxonomic research. The main objective of taxonomist is to describe and classify species, not to be lawyers!
I identify some problems that I think that be are the most important.
Linnean names
Linnean names are nice, they provide a some form of order in a time of great explorations around the world (XVIII-XIX century), when europeans became surprised with the diversity around the world. Many things are change from that time. By now we have several phylogenetic methods, which show how arbitrary a clade name would be (I like this post about the subject), and of course, the tree of life have far more divisions than Linnean ranks. Now we have databases to store names and search algorithms and find it quickly.
A rank-free taxonomy seems to be more adequate to store our information about the phylogeny of species. That is not an embrace of 'phylocode', as I prefer a taxonomy based on specific characters, or better a combination of characters and topology, instead of a topology based taxonomy (note that apomorphy based definition are also based on topology). Also, there is an special utility from ranked names: they can serve as Landmarks when browsing the tree of life.
The nested nature of biological classification allows a rank free taxonomy without the pains imagined by [1], and using “landmarks” helps in search--and writing--of abstracts, titles of papers, and of course, the search of a particular clade, as you can see using GenBank. Of course, genus and species can be (and I think would be) of obligatory nature. That allows a continuity with Linnean taxonomy.
If linnean names are more landmarks, then many of the laws for synonyms, and so one, need to be changed to a more practical usage.
Synonyms
Perhaps, the most annoying characteristic of taxonomy is the synonymy, rank changes, and several related problems. This problems are actually a burden of actual codes, and can be changed with changing the laws, without harming the actual classification.
Synonyms are really bad for databasesas the same entity is labeled with two different names, or (worst!) a same name is applied to different entities, in other cases, the range of the synonymy seems to be overlapping.
Of course, no matter what phylocoders say, the same problem applies to their 'phylogenetic' names: as new revision is published, the names continue but with completely different meaning. At least in traditional taxonomy, it is possible to reject some names.
It seems better to relax some of the naming rules, so a new classification would be clearly different from the original one. If a huge family is discovered to be massively paraphyletic, I think that it is no point allowing the original name to survive, it is simply an brutally wrong name.
For example, “reptiles” for a long time include many steam ammniotes, therapsids “mammal-like reptiles”, anapsids (as turtles), lepidosaurs (lizards and snakes) and crocodyles. They are amniotes but exclude birds and mammals. I can see a reason to retain that ugly name. It is simple, synonymize it with their monophyletic equivalent (Amniota), and never more use it. There is no way that a modern paper allows some confusion with the old ones. The only valid use of the name is when someone shows that the original reptiles, are monophyletic.
Names can be used instead to show different possible classifications. For example, the different arrangements of Arthropoda receive different names: Atelocerata (Myriapods, like centipede, and insects) vs. Pancrustacea (Crustaceans and insects), Mandibulata (Myriposd, crustaceans and insects) vs. schizoramia (crustaceans and arachnids). These names identify different entities, each one attached to a different phylogenetic proposal (note that form phylocoders, atelocerata, pancrustacea and mandibulata can be the same entity!).
Types
The reptile and arthropod example, are possible because there are no types fixing the names of that groups. At family levels, family names, or genus names are ruled by typification of names. Which allows more confusion, solutions to actual research.
There is an example: Lygaelidae was a large family of bugs (Insecta: Heteroptera), for long time it was believed that it was paraphyletic [2], but this was only demostrated by Henry [3]. He propose a new whole classification of Lygelids, elevating to family range no less than 7 subfamilies, and restructuring the meaning of Lygaelidae to only 3 subfamilies. Is the new Lygaelidae the same of, say 20 or 30 years ago? Of course not. Then typification was creating name stability (as is created by topology by phylocoders) but at cost of the loss of name utility.
As in the case of reptiles, there are no Lygaelidae any more. Any new reference i the litarature to 'Lygalelidae' only confuses with the initial meaning of the group (which is a synonym of the superfamly Lyageoidea). Of course you can use 'Lygaeidae sensu Henry' but it is only a clumsy (and error prone) way to give a new name.
Another nice example was provided by the previously mentioned post of Chris. It is about Drosophila. In this case, the usage can be against the taxonomic practice. For m, the solution is simple: no more Drosophila. But as there is a huge number of users of the name, that surely don't care about taxonomy, that outnumbered the number of Drosphilic taxonomic publications, there is when a commission can rule. The practical solution is to maintain Drosophila to the molecular people. Then solutions in cases of conflict would be guided by practical options, rather than some old described type.
This, of course, allows to made classification changes without 'using the types' (as far as the original characters were examined!). And free taxonomist to depend on some poorly known species (or even, specimen) to nominate genera and families.
Speaking about types...
Type specimens have a particular property: they are in the first world, but they are collected in the third world. The reason for this is historical, but their consequences are seen more acute today. Museums are measured by their amount of 'type specimens', and there are particular politics to borrow that specimens (only borrowing one on time, certifications, curators permission...). Also, some taxonomist, specially from the old past, simply nominate a wonderful amount of new species, only giving a superficial description (some color, some illustrations of genital parts) and based the whole 'description' on a type species designation. This old practice continued in several obscure papers in the third world.
Then typing, although seems to be a reasonable way to be objective, is more harming I guess that typing will never gone, but at the moment there are several nice perspectives to free from borrowing politics. Approaches like [4] with a great emphasis on characters and images, can change the situation. Researches far away from type specimenes can see high quality pics of several specimen parts. A side consequence is that the concept of type becomes lost, it is impossible to pic every part from a single specimen, and is possible that it ends destroyed, then the new typing would be more responsible, as it would be based in several different specimens.
Moreover a destypification increase general collections value, that is more the quantity and quality (e.g. fresh specimenes) of material available, than a particular specimen collected in 1816, saved from a fire in 1874, harmfully damaged by bad curation in 1903...
Actually there are many phylogenetic work without using type specimens, that is, the major bulk of molecular phylogenies, and I guess several morphological ones. I think that they do it in a very objective and testable way. If they can live without types, why classical taxonomist do not?
The data matrix
A second question from the previous section, is how a non-typified research can be objective? The answer is that instead on focusing on a particular specimen, phylogeneticists use a data matrix of taxon an characters.
A non type taxonomy enforce the use of well delimited characters, it is the only way to show the reality of the new designation. Look at some recent revisions with a phylogenetic analysis, and compare it with a revision without it (for example some of both see the pubs of AMNH). The character matrix allows to a quick examination of several characters, it is possible to see which state each character has in each taxon. New technologies (see [4]) couple specimen, characters and images for each cell entry. By default a matrix provide a multi-entry key, the identification tools are better.
There are some nice things of using a character matrix. The first, is an increasing interest in provide well defined characters [4]. As character are used for phylogenetic analysis, they would be stricter. Other characteristics like color patterns, length measurements would be restricted to a simpler description. Another advantage is that it provides a quick classification of a new species.
A nice real example was provided with the dinosaur paleontologist researches. The y publish some quick and small reports in high profile journals (like Nature or Science) with small descriptions, but as they have a great database of characters, several points of the anatomy of the new described fossil are immediately 'published', long before the detailed description in a more specialized journal.
Thinking on databasing
Of course, using a character matrix is direct consequence for storage: well defined characters and images enter smoothly in a database [4].
I think that the new challenges of the 'biodiversity crisis' as well as the 'taxonomic crisis' can be solved with a thinking of data storage. How can we store the data more efficiently? How can we link taxonomic and publication data? How changes in our knowledge about phylogeny could change the previous publication data, how the harm can be minimized?
It is important to a new taxonomy to keep the great advances made from Linneaus times, a start from the scratch is clearly a wrong solution. But also taxonomist would be able to made some concessions in their practice, and update it to new data architecture of the world.
It is time that taxonomy became a useful discussion about actual data, facing the massive extinction that the man is producing around the world, it seems weird that a taxonomic study would began searching for old papers from XVIII century, which only utility is that they provide a name, descriptions, characters and other things from that papers are of low value (by the way, taxonomy is the only field of science that continue using such old data. For historians old text are the source of investigation, for taxonomist is a more lawyer-like activity of searching for an 'old case'. Catalogs are some nice curiosities, and surely valuable for historians, but what is their actual value for taxonomist? They are important only because points to papers that establish a name).
If taxonomy is the main objective, then useful data storage is the main objective. Book keeping and law courts are not part of knowing biodiversity.
References
[1] Dominguez, E., Wheeler, Q. 1997. Taxonomic stability is ignorance. Cladistics 13: 367-372. doi: 10.1111/j.1096-0031.1997.tb00325.x
[2] Schuh, R. T., Slater, J. A. 1995. True Bugs of the World. Cornell Univ. New York
[3] Henry, T. J. 1997. Phylogenetic analysis of family groups within the infraorder Pentatomomorpha (Hemiptera: Heteroptera), with emphasis on the Lygaeoidea. Annals of the Entomological Society of America 90: 275-301
[4] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic Biology 56: 283-294. doi: 10.1080/10635150701313848
viernes, enero 18, 2008
DataTube? I can't wait!!
I just read a wonderful news at WiredScience, Google will be hosting open scientific data on the web [http://research.google.com]! WiredScience says that the interface will be similar to the one from YouTube, with annotations and comments.
I can't wait to see many morphological matrices, and morphological pics! I think that the excellent proposal of Ramírez et al. [1] can be coupled with that project :).
[1] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic biology 56: 283-294. doi: 10.1080/10635150701313848
I can't wait to see many morphological matrices, and morphological pics! I think that the excellent proposal of Ramírez et al. [1] can be coupled with that project :).
[1] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic biology 56: 283-294. doi: 10.1080/10635150701313848
sábado, diciembre 22, 2007
Homology and parsimony
Today is a transaltion day ;). This post is a translation of my previous post in spanish about the subject.
This post is somewhat inspired in a discussion by Ebach and Williams in their blog systematics & biogeography (they had many post about the subject!). Here I want to show the relationship between parsimony (or cladistic analysis) and homology.
Homology
It is a long tradition of discussion about the definition of homology. I use a definition similar to the traditional one, then two structures are homologs when both are considered the same structure in different organism. There are countless arguments to acept two structures as being the same: God's plan, natural order, morphotype, or the one I endorse, by common ancestry.
If two structures are homologs, it implies that they are the same structure, and they are inherited from a common ancestor of the two examined organims. Inheritance implies several things by definition: specially from genetics and development biology. Moreover, the structures could not be highly similar, even if they are the same, because they can be strongly modified. Nevertheless both are the same, so it implies that structures change across the time, in our assumptions include several process that promote the differentiation, such as genetic interaction and population dynamics.
Cladistic analysis do not have a direct interest in these phenomena associated with homology. The process are of interest in other fields like evolutionary biology, development biology, genetics, molecular biology, population genetics, ecology, etc. But lack of interest do not imply that they could be ignored! In the character definition for a cladistic analysis (i.e selection of homolog characters) several of those factors would be taken into account to propose character limits and character codification.
Unfortunately, as process is not of direct interest for cladistics allows the erroneous idea that the process is irrelevant in character definition. The the so called pattern cladists (as Nelson and Platnick [1] and recently Pleijel [2], Brower [3] and Ebach and Williams) to think that cladistics is free of evolutionary thinking, but it is a totally wrong position: the use of homolog characters implies a framework based on origin by common ancestry, the every character--and not only its apomorphic state--are the same structure [4, 5].
But it is more, implications of evolution are big for several characters, specially if there is a well knowledge of the character, the the idea of Kluge [6] who claims that only assumption of cladistic analysis is 'descent with modification' is also wrong. When you include inheritance and evolution, many things are included in the definition of character. Of course, for some characters we only got a morphological knowledge of the character, then assuming a simple 'descent with modification' seems to be correct, but for other characters (as in many vertebrates) we got knowledge from development and genetics of the structure, in such cases the assumption are far more complex than descent with modification.
Parsimony: the algorithm
The basic principle of parsimony algorithm is fairly simple. If you want to know the character state in the node 'x', which descendants are 'y' and 'z', and we know their character states, then the character state of 'x' is the intersection between states of 'y' and 'z', if the intersection is void then the state of 'x' is the union of 'y' and 'z' states. Using an union implies a character change (a step). Of course, optimization of states is more complex, but for the present discussion we only need the basic steps.
State assignation using the parsimony algorithm: (A) and (B) the state of ancestor 'x' is equal to the intersection of descendants 'y' and 'z', i this case, white; (c) 'y' and 'z' do not share any state, then the state assigned to 'x' is the state union of his descendants.As you see, the algorithm is 'independent' of the data used, you can use any kind of character (in computer science this problem is know as the coloring problem, and the 'characters' are colors of a map) and any kind of terminal. It is a common character of every method formalized in an algorithmic form. Then it is necessary to provide a proper justification to use the algorithm in a particular problem.
Parsimony: cladistic analysis
With the evolutive concept of homology, its union with the algorithmic parsimony is direct. If a character, no matter its state, is the same between two organisms that share a common ancestor, it implies that the character is inherited from the common ancestor, then the common ancestor would have the character.
Moreover if both organisms share the same form of the character, that is the same state, then that state would be in its common ancestor (the character is the same!), but if both organisms had different forms of the characters, we do not know which form would be present in the common ancestor, so we assume that it could be any of the both states, in that case, if both organisms really share the same character then a transformation would be happen.
This is exactly the same description of the parsimony algorithm. In cladistics the use of parsimony algorithm is justified because used characters are homologs. In this context the algorithm maximize our homology propositions when minimize transformations: this allows that most terminals with the same states would be contiguous. Then the hypothesis ad hoc of homoplasy are minimized [7].
Parsimony and homology tests
It is a common idea between cladists to say that parsimony is a test of homology: congruence. O disagree, because as I argument here the basis of parsimony is assuming from the very beginning that character are homologs! Any homology test would be previous to a rigorous cladistic analysis.
Homology test could have several forms, they could be morphology arguments (usally put under 'similarity' label), anatomical position, structural organization, ontogeny, genetics, and in most cases the 'test' is a conjunction of these procedures--for these reasons defining a character implies an strong theoretical background--. After examining all of those alternatives you got a good character. Is for these reasons that homoplasy is an ad hoc hypothesis: homoplasy is only justified in realtion with the cladogram.
Of course, character revision is always welcomed, and homoplasic ones maybe demand a close examination, but it is equally valid to examine every character. It is possible that codification form some characters is dubious, in such cases it is possible to use, as an exploratory devise different codifications (similar to the proposition of Ramirez [8] for morphology, and Wheeler [9] for molecules). But beware: this codification is not supported by the cladogram, because the argument used to defend that codification is the same used for homoplasy: it is justified only in relation with the cladoogram, the the codification is an ad hoc codification. In ambiguous cases I prefer a weighting schema as proposed by Neff [10]: because we know little about the character, and we have some doubts about its coding, it is better that it has a lower weight than characters that we know better.
**
Bonus: an historical speculation
He I show homology and parsimony ideas in a separated fashion and then I fuse them. I do it in that way to clarify the argument. But historically the development is intertwined from the beginning. If you read Wagner [11] in the algorithmic pathway, and Hennig [12] from the logical point it is clear that both positions are very close. Both visions were fused in an excelent fashion by Farris and its collaborators [7, 13, 14], which ideas (especially from [14]) agree in many points exposed here. Then from the beginning cladistic analysis and the parsimony algorithm walk together.
Wagner, Hennig and Farris development their ideas from a morphology context. At the same time Dayoff [15] experiment with several algorithms for molecular sequences, which at least today, homology ideas for molecular biologist are different to the morphological concept, I do not know what homology ideas used molecular biologist form 60s, but it seems that she did not believe that two bases (in case of Dayoff, two aminoacids) equal in two organisms imply common origin, the idea of point mutations precludes the idea. It is worth to note that Dayoff never could find a way to assign states in ancestors.
References
[1] Nelson, G., Platnick, N. 1981. Systematics and biogeography. Columbia Univ., New York.
[2] Pleijel, F. 1995. On character coding for phylogeny reconstruction. Cladistics 11: 309-315.
[3] Brower, A.V.Z. 2000. Evolution is not a necessary assumption of cladistics. Cladistics 16: 143-154.
[4] Fitzhugh, K. 2006. The philosophical basis of character coding for inference of phylogenetic hypothesis. Zoologica scripta 35: 261-286.
[5] Grant, T., kluge, A.G. 2004. Transformation series as an ideographic character concept. Cladistics 20: 23-31.
[6] Kluge, A.G. 2003. On the deduction of species relationships: a précis. Cladistics 19: 233-239.
[7] Farris, J.S. 1983. The logical basis of phylogenetic systematics. In: Platnick, M., Funk, V.A. (Eds.), Advances in cladistics, vol. 2. Columbia Univ., New York, Pp. 7-36.
[8] Ramirez, M.J. 2007. Homology as a parsimony problem: a dynamic homology approach for morphological data. Cladistics 23: 588-612.
[9] Wheeler, W. 1996. Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics 12: 1-9.
[10] Neff, N.A. 1986. A rational basis for a priori character weighting. Systematic zoology 35: 110-123.
[11] Wagner, W.H. 1961. Problems in the classification of ferns. In: Recent advances in botany, vol. 1, Univ. Toronto, Toronto, Pp. 841-844.
[12] Hennigh, W. 1966. Phylogenetic systematics. Univ. Illinois, Urbana.
[13] Kluge, A.G., Farris, J.S. 1969. Quantitative phyletic and the evolution of anurans. Systematic zoology 18: 1-32.
[14] Farris, J.S., Kluge, A.G., Eckardt, M.J. 1970. A numerical approach to phylogenetic systematics. Systematic zoology 19: 172-189.
[15] Dayoff, M.O. 1969. Computer analysis of protein evolution. Scientific american 221: 87-95.
Bases de datos relacionales para datos filogenéticos
Hoy es un día de traducciones ;). Este post es la versión en español de mi post previo en 'ingles' sobre el tema.
Hace un buen tiempo cuando estudiaba ingeniería de sistemas (=ciencias de los computadores), el paradigma reinante para las bases de datos eran las bases de datos relacionales, pero con el rápido crecimiento de internet, computadores más poderosas, y motores de búsqueda eficaces y populares--como Google--pusieron las bases de datos basadas en texto en la cima. Las bases de datos basadas en texto son el paradigma en el cual se han desarrollado las bases de datos para datos filogenéticos, como GenBank y TreeBASE.
En las bases de datos basadas en texto (BDTex) el énfasis se coloca en estandarizar los archivos de entrada, que deben contener todos los datos necesarios para hacer una búsqueda. Seguramente el lector conoce los archivos que se pueden descargar de GenBank, o las matrices/árboles descargados desde TreeBASE. Ellos tienen muchos campos, como una clasificación, identificación de las secuencias o caracteres, un algoritmo de string matching (coincidencia de letras/palabras? No se como se diga esto en español :P) es usado para encontrar coincidencias exactas o similares a la solicitud del usuario.
Entonces, la principal fuente de investigaciones para las BDTex esta basada en los algoritmos de string matching. El blog iPhylo de Rod Page, tiene varios posts, y enlaces a artículos y manuscritos (aquí o aquí) sobre el tema. Pero yo creo que las bases de BDTex son una aproximación erronea.
La coincidencia de letras (string matching) es una buena herramienta para buscar palabras claves a lo largo de la red, o varios grupos de letras, como la búsqueda de secuencias en GenBank (yo supongo que esa es la principal razón para que sea una BDTex!), pero parece problemática para una base de datos de taxonomía/filogenia. Quiero resaltar algunos puntos:
- Ambigüedad: aunque la entrada para cargar archivos en los servidores estén basados en plantillas de 'cortar y pager' (o usen asistentes paso a paso), es responsabilidad de quien carga el archivo llenarlo adecuadamente. Como Page llama la atención, en muchas de las matrices de TreeBASE los nombres de los terminales no son nombres científicos.
- Información irrelevante: cuando uno descarga un archivo, este suele estar lleno de información que el usuario no desea, y otra información no esta disponible. Además como las búsquedas se basan en coincidencias de texto, cosas como notas, comentarios y otros campos pueden confundir a los algoritmos (recuerden, se trata de algoritmos de string matching!).
- Formateado: Al ser una base de datos basada en texto, se tiene que hacer un compromiso en hacer los archivos legibles (=estandarizados) por varios programas lo que hace que los campos sean rigidos y existe una gran dificultad para incluir nuevos campos e información. Por ejemplo, la información geográfica en GenBank no es obligatoria--hasta donde se--incluso para los conjuntos de datos usados en análisis filogeográficos! La información geográfica y de material examinado para TreeBASE parece imposible si su implementación continua unida al formato nexus.
- Taxonomía: Como consecuencia del formateado rígido, clasificaciones alternativas y búsquedas usando sinónimos no pueden ser implementadas, o requieren búsquedas adicionales en bases de datos alternativas.
- Filogenia: ¿Alguna vez han intentado buscar una 'estructura de árbol' en TreeBASE?
Mi propia idea de como debe ser la estructura de la base de datos, en forma simplificada, y basada en mis propios intereses y búsquedas que yo he realizado es esta:

Por supuesto, un diseño adecuado requiere muchos años de desarrollo, con montones de entrevistas a varios taxónomos para permitir un producto que pueda ser usado adecuadamente alrededor del mundo! (A pesar que en sus post el apoya las BDTex, este post de Rod Page da varias ideas muy interesantes sobre el trabajo multidisciplinario de una base de datos para filogenias, por supuesto hay muchas cosas en la que yo no estoy de acuerdo xD).
Esta es la descripción de las diferentes tablas de mi bosquejo. La tabla 'taxon' guarda la nomeclatura del nombre actual de un taxon, sea una especie o una entidad supra-especifica, puede incluir la diagnosis (enlazada con la tabla de entrada de caracteres!), el 'concepto filogenético', enlace a dibujos o fotografías, el espécimen tipo (con enlace a la tabla de especímenes) y cosas de ese estilo. La tabla de 'synonyms' y 'classification' son tablas secundarias que guardan solo la relación entre dos taxones: los sinónimos (se puede incluir el motivo de la sinonimia), y el siguiente taxon más inclusivo de un taxon particular (y puede incluir el autor de esa propuesta de inclusión). Como cada taxon es tratado de forma independiente uno puede incluir tantos sinónimos como desee, o múltiples clasificaciones propuestas.
La tabla 'specimen' puede guardar información especifica de el material examinado, enlace a fotografías del espécimen, la localidad donde fue colectado, y cosas así.
Además hay una sección de caracteres! La tabla 'character' guarda el nombre, descripción, y quizá una bibliografía y dibujos o fotos del carácter, con 'character equivalences' es posible guardar caracteres equivalentes usados en otros estudios, permitiendo referencias cruzadas entre estudios que tengan diferentes grupos de terminales, la tabla 'character entry' guarda la información especifica de un carácter y un taxon, puede ser una celda de una matriz de caracteres morfológicos, o un fragmento específico de una secuencia molecular.
Este diseño puede ayudar en las cosas en donde las BDTex fallan. La ambigüedad es reducida, porque cada entrada es única y especifica: al cargar un archivo en la base de datos se debe identificar la naturaleza especifica de la entrada. BDTex pueden desarrollarse con un estándar como ese, pero en este caso es posible incluir muchos sinónimos, y nombres específicos de caracteres. Un proceso curatorial para la taxonomía puede ser posible sin destruir la integridad del conjunto de datos, y como los resultados filogenéticos deben ser introducidos en forma de una clasificación, el distanciamiento entre la filogenia y la clasificación [1] se vera reducido.
Cuando se realiza una búsqueda el usuario puede recuperar solo la información específica que desea: por ejemplo todos los caracteres para la cabeza de Hymenoptera, usando las equivalencias los caracteres pueden ser organizados en una forma relativamente adecuada, y usando la tabla de clasificaciones es posible encontrar caracteres de la cabeza usados en diferentes estudios, caracteres que pueden estar incluidos si se utiliza una clasificación alternativa, o caracteres que pueden estar presentes al ser incluidos en clasificaciones más inclusivas (por ejemplo, caracteres usados para definir Hexapoda o Arthropoda), de esta forma se consigue un verdadero sistema de recuperación de la información basado en la clasificación [2].
Tal y como muestran Nixon et al. [3] un único formato de archivo no es algo bueno para una base de datos, más bien, es preferible la estructura de tablas, y mecanismos de reporte que puedan producir entradas en diferentes formatos, por ejemplo recuperar secuencias en el formato de GenBank, en el de TNT, y en POY (fasta), o un archivo de distribuciones listo para usar con NDM.
La tabla de clasificaciones puede ayudar para localizar estudios que apoyen o rechacen una clasificación particular, el usuario puede encontrar la evidencia para un agrupamiento y para la clasificación alternativa. Pueden desarrollarse y usarse algoritmos que realicen la tarea de convertir la estructura de un árbol en una solicitud de búsqueda--Page a posteado sobre el tema con relativa frecuencia ;)--. Pero es más poderoso que las busquedas basadas en texto pues la misma base de datos posee la información (la clasificación jerárquica) para realizar la búsqueda.
Espero algún día hacerme millonario (lo dudo :P) o recibir financiación--ojala ;)--para elaborar esta enorme tarea, o al menos que alguien en la red, tenga ideas similares. Hasta entonces el único camino parece ser el sufrimiento continuo con las bases de datos 'taxonómicas' y 'filogenéticas' que esta por la red...
Pd. Como se puede ver, quizá esta base de datos pondria más cosas sobre el investigador que cargue los datos... pero despues de haber estado en campo por meses, examinar material por horas, día tras día, escribir reportes y manuscritos, cargar la información es solo una parte de toda la investigación!
[1] Franz, N.M. 2005. On the lack of good scientific reasons for the growing phylogeny/classification gap. Cladistics 21: 495-500.
[2] Farris, J.S. 1979. The information content of the phylogenetic system. Systematic zoology 28: 483-519.
[3] Nixon, K.C., Carpenter, J.M., Borgardt, S.J. 2001. Beyond NEXUS: universal cladistic data objects. Cladistics 17: S53-S59.
Etiquetas:
bases de datos,
taxonomía digital
jueves, diciembre 20, 2007
Relational databases for phylogenetic data
A long time ago, when I'm studying computer science, the only paradigm for databases were realtional databases, then under internet fast growth, more powerful computers, and popular searching engines--like Google--put text based databases on the top. Text based databases are the paradigm used to construct some databases for phylogenetic data, like GenBank, and TreeBASE.
Under text based databases the emphasis is put on somewhat standardized input files, that keep all the possible data necessary to make the search, you surely know the files retrieved from GenBank, or the matrices/trees from treeBASE, they have several fields, like classification, identification of sequence or character, a string matching algorithm is used to found exact and similar matches from the user query.
Then, the principal source of research in a text based database is to implement matching algorithms, the blog iPhylo from Rod Page, have several posts, paper and manuscripts links (here, here) about that subject. But I think that a text based databases are a wrong approach.
I remark some points:
Relational databases are a different concept from text-based databases. They are based on several independent tables connected to key-fields and in several cases using a secondary tables to match keys form different tables. The searching engine is based on specific keys (for complex searches) and specific fields form each table.
My own idea of the structure of database, in a sketch fashion, and based on my usual queries is like this:
Of course, a proper database design will need several years of development, with tons of interviews of several taxonomists to allow a product that could be used in a right way for several people around the world! (Although he had several post endorsing string matching databases, this post of Rod Page provide several nice ideas about the integrative work of a phylogenetic databases, of course there are several things that i don-t like xD).
Here I explain some part of the different tables, the table 'taxon' store the actual nomeclature of a taxon name, a species or a supra-specific entity, its name, author, it may include the diagnosis (linked to character entry!) a 'phylogenetic concept', a link to pictures, the type specimen (with a link to specimen!) and things like that. The table 'synonyms' and 'classification' are secondary tables to store only a relationship between two taxons: synonyms (it could include the motive of the synonymy), and the next inclusive taxon to a particular taxon (it might include the author of this inclusive relationship), as each taxon is independent you could include as many synonyms as you know, or as many classifications proposed.
The 'specimen' table could store specific information about the examined material, links to pictures of the specimen, the locality where the specimen was collected, and so on.
It is also a character section, the table 'character' store the name, description, and maybe bibliography and pictures of the character, with 'character equivalences' it is possible to store equivalent characters used in other studies, allowing cross reference between studies that used different taxon scopes, the 'character entry' table store the specific information for the character and the taxon, it could be a single cell in from a morphology matrix, or an specific sequence fragment.
This design could help in queries/actions in which text-based databases fail. The ambiguity is reduced, as each entry is a single one: the uploader would identify the particular nature of their entries in an specific way, text based databases could be developed with that standard in mind, but here it is possible to keep an strict species naming, as many synonyms as you want, and specific character names. A curator process for the taxonomy could be possible without harming the whole database data, and as phylogenetic results would be introduced in a form of classification, the gap between phylogeny and classification [1] would be reduced.
When you perform a search you could retrieve only the specific information that you want: for example all head characters from Hymenoptera, using the equivalences the characters could be more or less organized, and using the classification table you could retrieve head characters used for many different studies, alternative characters that could match under alternative classifications, or possible characters that could be present as they are scored for more inclusive classifications (for example, a character used to define Hexapoda and Arthropoda), a truly information retrieval system based on classification [2].
As is remarked by Nixon et al. [3] a single file format is not a good thing to a database, instead, it is preferable a table structure, and report tools that could produce entries in different formats, for example retrieving sequences in a GenBank format, in a TNT format, and a POY (fasta) format, or a distribution file ready to use in NDM.
The classification table could help to retrieve studies that support or reject a particular classification, you could found the evidence for one grouping and for the alternative classification. Particular algorithms need to be developed to translate the tree structure to a query--Page posted about the subject frequently ;)--. But it is more powerful that a text based search because the same database have the information needed (the hierarchic classification) to perform the search.
I hope sometime I became rich (I doubt it) or receive a grant--I hope ;)--to perform this huge task, or at least that someone around the net, had similar ideas. Until then the only path is continuous suffering with some 'taxonomic' and 'phylogenetic' databases around the net...
Pd. As you note, maybe a database like that put some burden on researcher that upload the data... but if you go to field trips for months, examine material for hours, day after day, write reports and manuscripts, uploading information it is part of the whole research!
[1] Franz, N.M. 2005. On the lack of good scientific reasons for the growing phylogeny/classification gap. Cladistics 21: 495-500.
[2] Farris, J.S. 1979. The information content of the phylogenetic system. Systematic zoology 28: 483-519.
[3] Nixon, K.C., Carpenter, J.M., Borgardt, S.J. 2001. Beyond NEXUS: universal cladistic data objects. Cladistics 17: S53-S59.
Under text based databases the emphasis is put on somewhat standardized input files, that keep all the possible data necessary to make the search, you surely know the files retrieved from GenBank, or the matrices/trees from treeBASE, they have several fields, like classification, identification of sequence or character, a string matching algorithm is used to found exact and similar matches from the user query.
Then, the principal source of research in a text based database is to implement matching algorithms, the blog iPhylo from Rod Page, have several posts, paper and manuscripts links (here, here) about that subject. But I think that a text based databases are a wrong approach.
I remark some points:
- Ambiguity: even if the entry for uploads is based on a 'cut and paste' template (or a wizard) the uploader is left with the responsibility to fill it adequately. As Page remarks, in many TreeBASE matrices terminal names are not properly scientific names.
- Irrelevant information: when you retrieve a file, it is usually full of information that you don't want, other information it is not provided. As searches are based on text matches, several notes, comments and other fields contain information, its function seems to be to confuse the searching algorithms (remember, they are string matching algorithms!).
- Formatting: As is text based, a compromise to made the files available to several programs made the entry fields to be rigid and difficult to modify and include new fields/information. For example in GenBank geographic information is not mandatory--as far as I know--even for phylogeographic datasets! Geographic and examined material for TreeBASE seems to be impossible to implement using the nexus format.
- Taxonomy: As a consequence of the rigid format, alternative taxonomies and synonyms searches can not be implemented, or require searches on alternative databases.
- Phylogeny: Why not try a 'tree structure' search on TreeBASE?
Relational databases are a different concept from text-based databases. They are based on several independent tables connected to key-fields and in several cases using a secondary tables to match keys form different tables. The searching engine is based on specific keys (for complex searches) and specific fields form each table.
My own idea of the structure of database, in a sketch fashion, and based on my usual queries is like this:

Of course, a proper database design will need several years of development, with tons of interviews of several taxonomists to allow a product that could be used in a right way for several people around the world! (Although he had several post endorsing string matching databases, this post of Rod Page provide several nice ideas about the integrative work of a phylogenetic databases, of course there are several things that i don-t like xD).
Here I explain some part of the different tables, the table 'taxon' store the actual nomeclature of a taxon name, a species or a supra-specific entity, its name, author, it may include the diagnosis (linked to character entry!) a 'phylogenetic concept', a link to pictures, the type specimen (with a link to specimen!) and things like that. The table 'synonyms' and 'classification' are secondary tables to store only a relationship between two taxons: synonyms (it could include the motive of the synonymy), and the next inclusive taxon to a particular taxon (it might include the author of this inclusive relationship), as each taxon is independent you could include as many synonyms as you know, or as many classifications proposed.
The 'specimen' table could store specific information about the examined material, links to pictures of the specimen, the locality where the specimen was collected, and so on.
It is also a character section, the table 'character' store the name, description, and maybe bibliography and pictures of the character, with 'character equivalences' it is possible to store equivalent characters used in other studies, allowing cross reference between studies that used different taxon scopes, the 'character entry' table store the specific information for the character and the taxon, it could be a single cell in from a morphology matrix, or an specific sequence fragment.
This design could help in queries/actions in which text-based databases fail. The ambiguity is reduced, as each entry is a single one: the uploader would identify the particular nature of their entries in an specific way, text based databases could be developed with that standard in mind, but here it is possible to keep an strict species naming, as many synonyms as you want, and specific character names. A curator process for the taxonomy could be possible without harming the whole database data, and as phylogenetic results would be introduced in a form of classification, the gap between phylogeny and classification [1] would be reduced.
When you perform a search you could retrieve only the specific information that you want: for example all head characters from Hymenoptera, using the equivalences the characters could be more or less organized, and using the classification table you could retrieve head characters used for many different studies, alternative characters that could match under alternative classifications, or possible characters that could be present as they are scored for more inclusive classifications (for example, a character used to define Hexapoda and Arthropoda), a truly information retrieval system based on classification [2].
As is remarked by Nixon et al. [3] a single file format is not a good thing to a database, instead, it is preferable a table structure, and report tools that could produce entries in different formats, for example retrieving sequences in a GenBank format, in a TNT format, and a POY (fasta) format, or a distribution file ready to use in NDM.
The classification table could help to retrieve studies that support or reject a particular classification, you could found the evidence for one grouping and for the alternative classification. Particular algorithms need to be developed to translate the tree structure to a query--Page posted about the subject frequently ;)--. But it is more powerful that a text based search because the same database have the information needed (the hierarchic classification) to perform the search.
I hope sometime I became rich (I doubt it) or receive a grant--I hope ;)--to perform this huge task, or at least that someone around the net, had similar ideas. Until then the only path is continuous suffering with some 'taxonomic' and 'phylogenetic' databases around the net...
Pd. As you note, maybe a database like that put some burden on researcher that upload the data... but if you go to field trips for months, examine material for hours, day after day, write reports and manuscripts, uploading information it is part of the whole research!
[1] Franz, N.M. 2005. On the lack of good scientific reasons for the growing phylogeny/classification gap. Cladistics 21: 495-500.
[2] Farris, J.S. 1979. The information content of the phylogenetic system. Systematic zoology 28: 483-519.
[3] Nixon, K.C., Carpenter, J.M., Borgardt, S.J. 2001. Beyond NEXUS: universal cladistic data objects. Cladistics 17: S53-S59.
viernes, diciembre 14, 2007
Muy buena noticia sobre TNT
En estos momentos no estoy en la ciudad, así que no puedo escribir posts muy largos :'(... pero no podia dejar pasar esta gran noticia sobre TNT. La Willy Hennig Society a tomado la financiación y apoyo a TNT (un programa de Pablo Goloboff, Steve Farris y Kevin Nixon), así que a partir de ahora el programa es gratuito! Solo hay que cumplir algunas condiciones muy simples: uso personal, y citar el programa--y como no la financiación por partde de la WHS--al publicar resultados.
En caso de que no lo sepan, TNT es el programa más rápido para análisis filogenético de parsimonia, que implementa muchos de los más recientes y eficientes algoritmos de búsqueda [1, 2], y un excelente y poderoso lenguaje de macros/scripts. Si ustedes realizan análisis cladísticos/filogenéticos, este es el programa ideal.
Yo adoro el editor de caracteres! Es muy fácil de usar, y mucho más intuitivo que WinClada, NDE o Mesquite (sip! Es mucho mejor que el de Mesquite!).
Pueden descargarlo aquí: http://www.zmuc.dk/public/phylogeny/TNT/
lean la licencia y a disfrútenlo :)
Cuando regrese a Bogotá completare las citas ;)
[1] Nixon, K.C. 1999. The parsimony ratchet a new method for rapid parsimony analysis. Cladistics 15: 407-414.
[2] Goloboff, P. A. 1999. Analyzing large data sets in reasonable times: solutions for composite optima. Cladistics 15: 415-428.
En caso de que no lo sepan, TNT es el programa más rápido para análisis filogenético de parsimonia, que implementa muchos de los más recientes y eficientes algoritmos de búsqueda [1, 2], y un excelente y poderoso lenguaje de macros/scripts. Si ustedes realizan análisis cladísticos/filogenéticos, este es el programa ideal.
Yo adoro el editor de caracteres! Es muy fácil de usar, y mucho más intuitivo que WinClada, NDE o Mesquite (sip! Es mucho mejor que el de Mesquite!).
Pueden descargarlo aquí: http://www.zmuc.dk/public/phylogeny/TNT/
lean la licencia y a disfrútenlo :)
[1] Nixon, K.C. 1999. The parsimony ratchet a new method for rapid parsimony analysis. Cladistics 15: 407-414.
[2] Goloboff, P. A. 1999. Analyzing large data sets in reasonable times: solutions for composite optima. Cladistics 15: 415-428.
Really huge news about TNT
I'm out of the city, so no computer for long posts :'(... but I'm happy to give a flash news about TNT: The Willy Hennig Society takes the sponsorship of TNT (a software by Pablo Goloboff, Steve Farris and Kevin Nixon), so from now, the program is free! The are some simple conditions: personal use, and a citation of the program--and the sponsor, that is the WHS ;)--in published results!
In case that you don't know, TNT is the most faster program for phylogenetic analysis under parsimony, implements several new and efficient heuristic algoritms [1,2], and a powerful script/macro language. If you are doing cladistics/phylogenetics, you should surely dream with this program!
I love the matrix editor! Is easy to use, and more straightforward than WinClada, NDE or Mesquite (yep!... far better than mesquite!).
You could download it at: http://www.zmuc.dk/public/phylogeny/TNT/
read the license agreement and enjoy :)
I give the proper citations when I return to Bogotá :P
[1] Nixon, K.C. 1999. The parsimony ratchet a new method for rapid parsimony analysis. Cladistics 15: 407-414.
[2] Goloboff, P. A. 1999. Analyzing large data sets in reasonable times: solutions for composite optima. Cladistics 15: 415-428.
In case that you don't know, TNT is the most faster program for phylogenetic analysis under parsimony, implements several new and efficient heuristic algoritms [1,2], and a powerful script/macro language. If you are doing cladistics/phylogenetics, you should surely dream with this program!
I love the matrix editor! Is easy to use, and more straightforward than WinClada, NDE or Mesquite (yep!... far better than mesquite!).
You could download it at: http://www.zmuc.dk/public/phylogeny/TNT/
read the license agreement and enjoy :)
[1] Nixon, K.C. 1999. The parsimony ratchet a new method for rapid parsimony analysis. Cladistics 15: 407-414.
[2] Goloboff, P. A. 1999. Analyzing large data sets in reasonable times: solutions for composite optima. Cladistics 15: 415-428.
lunes, noviembre 26, 2007
A brief sketch of a program for phylogenetics
The last week is a nice week for me, as I received a grant to begin a graduate scholarship :) As my main objective is coupled with the development of a program, and my own interest in phylogenetics, and in a lesser extent, in software engineering, i will try to show here how I think a phyligenetic-related software would be implemented.
Do not take it for granted! The software projects in which i am involved to date, are smaller and short-term, and then I usually make some (ugly) shortcuts for immediate solutions but with low flexibility to modifications. Here my sketch take into account the possibility of further refinements of each part, so it is far more modular than my small projects.
Every program is composed with two parts. The first, the more messier and the most bored part to write, is the user interface (UI). The second, is the program itself (the core). In small projects you could mix both parts without much pain, but as the program growths leaving both parts mixed makes the program difficult to maintain: you would rewrite several code parts to introduce new input formats.
Ideally you would try to make the core to be totally independent from the UI, but usually it is not possible (unless your program do noting). An effort to construct a well designed communication between the UI and the core, is always well payed in the future.
I start my design with the core, and in every moment thinking how the core would communicate with a black-box UI. In phylogenetics you have two major parts: the matrix, and the tree. Both are composed of more smaller elements. The matrix is a list of labels, usually with some coupled data (taxon-character matrix, a taxon-distribution matrix). The tree is a topology in which each vertex is associated with the matrix labels.
You could think that both the tree and the matrix are the same, after all, the tree is only an specific way to sort the terminals in the matrix (it seems that TreeFitter[1] works in that way). But there is a difference, the terminals in the matrix are unique, but you could have many tress associated with the same matrix, terminal 'a' are independent in two trees, but they could point to the same element in a matrix.
Apart from terminals, that have a fixed content, internal nodes of a tree are coupled with the same sort of data as terminals, but this is a variable content, dependent on the tree topology. So we could implement a single fixed matrix, and one (or more) variable matrix, when the tree is unused, the variable matrix can be discarded.
To communicate the core with the UI, at this moment, I prefer string streams. They could be find in standard library (although with some little spelling differences), in that way the UI always translate between formats to send to the core specific strings, for example the tree in parenthesis format, so the tree construction would not interact with files, or console input, for example.
In that way, the UI is a central of translation from console, file, command dialogs, etc., that can be implemented for different systems (e.g. Win, Linux, command line), but the core remains the same.
[1] Ronquist, F. TreeFitter, program and documentation, available on line at: http://www.ebc.uu.se/systzoo/research/treefitter/treefitter.html
Do not take it for granted! The software projects in which i am involved to date, are smaller and short-term, and then I usually make some (ugly) shortcuts for immediate solutions but with low flexibility to modifications. Here my sketch take into account the possibility of further refinements of each part, so it is far more modular than my small projects.
Every program is composed with two parts. The first, the more messier and the most bored part to write, is the user interface (UI). The second, is the program itself (the core). In small projects you could mix both parts without much pain, but as the program growths leaving both parts mixed makes the program difficult to maintain: you would rewrite several code parts to introduce new input formats.Ideally you would try to make the core to be totally independent from the UI, but usually it is not possible (unless your program do noting). An effort to construct a well designed communication between the UI and the core, is always well payed in the future.
I start my design with the core, and in every moment thinking how the core would communicate with a black-box UI. In phylogenetics you have two major parts: the matrix, and the tree. Both are composed of more smaller elements. The matrix is a list of labels, usually with some coupled data (taxon-character matrix, a taxon-distribution matrix). The tree is a topology in which each vertex is associated with the matrix labels.
You could think that both the tree and the matrix are the same, after all, the tree is only an specific way to sort the terminals in the matrix (it seems that TreeFitter[1] works in that way). But there is a difference, the terminals in the matrix are unique, but you could have many tress associated with the same matrix, terminal 'a' are independent in two trees, but they could point to the same element in a matrix.
Apart from terminals, that have a fixed content, internal nodes of a tree are coupled with the same sort of data as terminals, but this is a variable content, dependent on the tree topology. So we could implement a single fixed matrix, and one (or more) variable matrix, when the tree is unused, the variable matrix can be discarded.To communicate the core with the UI, at this moment, I prefer string streams. They could be find in standard library (although with some little spelling differences), in that way the UI always translate between formats to send to the core specific strings, for example the tree in parenthesis format, so the tree construction would not interact with files, or console input, for example.
In that way, the UI is a central of translation from console, file, command dialogs, etc., that can be implemented for different systems (e.g. Win, Linux, command line), but the core remains the same.
[1] Ronquist, F. TreeFitter, program and documentation, available on line at: http://www.ebc.uu.se/systzoo/research/treefitter/treefitter.html
jueves, noviembre 15, 2007
Now in english
All the few previous post I made for this blog were in spanish, my mother language, for now I will try to post in both english and spanish. In english because everybody know that it is the science language, I want that other people interested in systematics, biogeography and specifically cladistics could read my blog! In spanish because I never found another cladistics blog in spanish. The content of posts would not necessarily overlap, I would try to provide translations of each post, but for my own convenience, I want to post technical posts--dealing with computer programming, databasing, and 'phyloinformatics', as Page call it--in english, and philosophical and methodological in spanish. I also want to post about true bug's (Insecta: Heteroptera) phylogenetics, I think I would use english.
I also want to increase the rate of postings ;) I hope, I could post at least one time per week ;).
As you could see, I am not a good english writer, so I hope that this blog could help me to write somewhat better :).
I also want to increase the rate of postings ;) I hope, I could post at least one time per week ;).
As you could see, I am not a good english writer, so I hope that this blog could help me to write somewhat better :).
miércoles, noviembre 07, 2007
Homología y parsimonia
Este post esta inspirado ligeramente en la discusión ke han colocado Ebach y Williams en su blog systematics & biogeography. Mi intención es mostrar como se relaciona el análisis de parsimonia (o análisis cladístico) con la homología.
Homología
El termino homología tiene una larga historia de discusiones. Yo uso un poko la definición tradicional y tomo como homologas dos estructuras ke son consideradas la misma estructura en dos organismos diferentes. Uno puede usar múltiples argumentos para decir que significa ke una estructura sea la misma: el plan de Dios, el orden del universo, el tipo primordial, o la ke yo prefiero que es por ancestría común.
Cuando digo ke dos estructuras son homologas, implica ke son la misma estructura, y que esta a sido heredada de el ancestro común de ambos organismos. La herencia implica un montón de procesos, que se hayan implícitos en la definición: especialmente genética y biología del desarrollo. Además, las estructuras, a pesar de ser las mismas no suelen ser iguales, pueden estar fuertemente modificadas, pero aún así son la misma estructura, esto implica ke hay una transformación de las homologías, así entran en juego otros factores a nuestra definición como son los procesos ke desencadenan esas diferencias, como la interacción genética y la dinámica de las poblaciones.
En un análisis cladístico no hay una preocupación directa por todos esos fenómenos asociados a la homología. Esos procesos son más de interés de otras ramas de la biología como la evolución, le biología del desarrollo, la genética, la biología molecular, la genética de poblaciones, la ecología, etc. pero ke no sean de interés no implica ke sean ignorables: en la definición de caracteres (homologos) para el análisis cladístico muchos de esos factores entran en consideración para dar limite a los caracteres y como codificarlos.
Desafortunadamente, el ke el 'proceso' no sea preocupación de la cladística a llevado a la errónea idea ke es irrelevante en la definición de los caracteres. Así los cladístas de patrón (como Nelson, Platnick, o más recientemente Pleijel y Brower) consideran ke la cladística esta totalmente libre de todo pensamiento evolutivo, pero todo lo contrario: el usar caracteres homologos implica ke estamos trabajando en un marco ke esta basado en el origen por ancestría común, por lo tanto todos los estados de un carácter--y no solo los apomorficos--son la misma estructura.
Pero no solo eso, la implicación evolutiva es enorme para muchos caracteres, en especial si tenemos un buen conocimiento para el carácter, por lo ke la idea de Kluge (e.g. 2003) de ke solo es necesario 'descendencia con modificación' es también errada. Al introducir herencia y evolución, muchas más cosas se encuentran embebidas en la definición de cada carácter. Claro esta, eso depende de cada carácter, en muchas situaciones solamente tenemos un conocimiento morfológico del carácter, con lo ke lo único ke podemos asumir es la simple 'descendencia con modificación', pero en muchos grupos (como los vertebrados), el conocimiento ke se tiene sobre los caracteres es mucho más amplio, incluyendo por ejemplo un buen conocimiento del desarrollo de la estructura.
Parsimonia: algoritmo
El algoritmo básico de parsimonia es bastante sencillo. Si keremos establecer el estado de carácter presente en un nodo A, cuyos descendientes son B y C, y de los cuales ya sabemos su estado de carácter, el estado de A es la intersección de los estados de B y C, si esa intersección es un conjunto vacío entonces el estado es la unión de los estados B y C. Esta unión implica ke ha habido un cambio (lo ke es llamado un paso). Optimizar los estados es un poko más complicado, pero esas operaciones son lo básico del algoritmo.
Asignación de estados con el algoritmo de parsimonia: (A) y (B) el estado del nodo ancestro 'x' es igual a la intersección del estado de los descendientes 'y', 'z', en este caso, blanco; (C) 'y' y 'z' no comparten ningún estado, por lo tanto el estado asignado a 'x' es la unión de los estados de los descendientes.
Es fácil notar ke el algoritmo es 'independiente' de los datos tomados, uno puede tomar cualkier clase de 'carácter' (en ciencias de computadores este problema es conocido como el problema del coloreado, y los 'caracteres' son colores en un mapa) y cualkier clase de terminales de análisis. Esa es una característica propia de los métodos formalizados como algoritmos. Dado esas circunstancias es necesario proveer una justificación del uso de un algoritmo particular para cada problema.
En cladística la justificación para el uso de parsimonia viene dada por el concepto de homología.
Parsimonia: análisis cladístico
Dado el concepto evolutivo de homología, su unión con el concepto algoritmico de parsimonia es directo. Si decimos ke un carácter, cualquiera ke sea su estado, es el mismo en dos organismos ke comparten un ancestro común, eso implica ke ese carácter es heredado del ancestro común, por lo tanto el ancestro común debe tener ese carácter.
Si además ambos organismos comparten la misma expresión del carácter, es decir su mismo estado, entonces ese mismo estado se encuentra en su ancestro común, si por el contrario ambos organismos tienen diferentes expresiones del carácter, el ancestro común tiene el carácter, pero no sabemos ke expresión pueda tener, por lo ke asumimos ke tiene alguno de los dos estados, y damos por hecho, ke si en realidad ambos organismos comparten ese carácter, debió haber ocurrido una transformación.
Como se puede ver, esa es exactamente la misma descripción del algoritmo de parsimonia. En cladística el algoritmo esta justificado pues los caracteres usados son tomados como homologos. En ese contexto el algoritmo maximiza nuestras proposiciones de homología al minimizar las transformaciones: de esa forma se asegura ke la mayor cantidad de organismos con el mismo estados de carácter kedan contiguos. Esto es minimización de hipótesis ad hoc de homoplasia (Farris, 1983).
Parsimonia y los tests de homología
Es una idea común entre los cladístas el argumentar ke parsimonia es un test de homología: la congruencia. Yo estoy en desacuerdo, pues como he argumentado akí la base de la parsimonia es asumir de plano ke las entradas para un carácter son homologas. Para mi los tests de homología son una condición previa del análisis cladístico.
Esos test pueden venir de muchas formas, bien sea basado en argumentos de morfología (usualmente llamada 'similitud'), posición anatómica, organización estructural, ontogenía, genética, en la mayor parte de los casos muchas de esas pruebas son conjuntas--por lo ke definir un carácter marca una fuerte carga teórica--es despues de haber examinado todas esas alternativas ke podemos tener un buen carácter. Es por eso ke una homoplasía es una hipótesis ad hoc: la homoplasia en el carácter solo esta justificada en relación con el cladograma.
Por supuesto, es necesario recurrir a la revisión de caracteres, y ciertamente es bueno revisar los caracteres homoplasicos, pero es igualmente valioso examinar todos los caracteres por igual. Es posible ke tengamos alguna duda acerca de como codificar caracteres, debido a su complejidad o porke son muy poko conocidos, en estos casos puede ser posible utilizar como apoyo el resultado ke producen diferentes tipos de codificación (esa es la alternativa ke usa Ramírez, en prensa, y en molecular Wheeler, 1996). Pero debo llamar la atención sobre eso: la codificación no esta 'avalada' por el cladograma, pues el argumento ke se utiliza para defender esa codificación es el mismo usado para la homoplasia: solo se justifica en relación con el cladograma, por lo tanto la codificación tiene un fuerte elemento ad hoc. Para esos casos yo preferiría ke se utilizara un eskema de pesado como el de Neff (1986): dado ke conocemos poko del carácter, tanto ke tenemos nuestras dudas para codificarlo, es mejor ke tenga un peso más bajo ke el de caracteres ke conocemos mejor.
**
Adicional: una especulación histórica
En este ensayo yo presente las ideas de homología y parsimonia por separado y luego las fusione. Lo hice así porke considero ke da claridad a la discusión. Sin embargo, históricamente el desarrollo fue muy entrelazado desde el inicio. Leyendo a Wagner (1961, desde el lado algoritmico) y a Hennig (1968, desde el lado lógico) es claro como ambas ideas son intercambiables desde el inicio. Ambas visiones kedaron integradas muy bien en Farris et al. 1970, ke a pesar de ser un intento de axiomatización, contiene ideas (para mi muy estimulantes) ke coinciden con muchos de los puntos expuestos akí. Así desde el inicio las ideas del análisis cladístico y de parsimonia estuvieron juntas.
Wagner, Hennig y Farris desarrollaron sus ideas desde un contexto de la morfología. Dayoff (1969) también experimento con varios algoritmos para secuencias moleculares, donde al menos por ahora, las ideas de homología son en muchos casos diferentes a el concepto de morfología, en su época (mediados de los 60s), no se exactamente ke pensaran de homología los moleculares, pero era claro ke ellos no creían ke dos bases (en el caso de Dayoff, dos aminoacidos) iguales en dos organismos implicaba un origen común, la idea de mutaciones constantes precluyo la idea. Es interesante ver ke Dayoff no pudo encontrar una forma adecuada de asignar estados en los ancestros.
Referencias
Dayoff, M.O. (1969) Computer analysis of protein evolution. Sci. Amer. 221: 87-95.
Farris, J.S. (1983) The logical basis of phylogenetic systematics. En: Platnick, N., Funk, V.A. (Eds.), Advances in cladistics, vol. 2. Columbia, New York. Pp. 7-36.
Farris, J.S., Kluge, A.G., Eckardt, M.J. (1970) A numerical approach to phylogenetic systematics. Syst. Zool. 19: 172-189.
Hennig, W. (1968) Elementos de una sistemática filogenética. Eudeba, Buenos Aires.
Kluge, A.G. (2003) On deduction of species relationships: a précis. Cladistics 19: 233-239.
Neff, N.A. (1986) A rational basis for a priori character weighting. Syst. Zool. 35: 110-123.
Ramírez, M.J. (2007) Homology as a parsimony problem: a dynamic homology approach for morgological data. Cladistics 23:xx-xx.
Wagner, W.H. (1961) Problems in the classification of ferns. En: Recent advances in botany, vol. 1 Univ. Toronto, Toronto. Pp. 841-844.
Wheeler, W.C. (1996) Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics 12: 1-9.
Homología
El termino homología tiene una larga historia de discusiones. Yo uso un poko la definición tradicional y tomo como homologas dos estructuras ke son consideradas la misma estructura en dos organismos diferentes. Uno puede usar múltiples argumentos para decir que significa ke una estructura sea la misma: el plan de Dios, el orden del universo, el tipo primordial, o la ke yo prefiero que es por ancestría común.
Cuando digo ke dos estructuras son homologas, implica ke son la misma estructura, y que esta a sido heredada de el ancestro común de ambos organismos. La herencia implica un montón de procesos, que se hayan implícitos en la definición: especialmente genética y biología del desarrollo. Además, las estructuras, a pesar de ser las mismas no suelen ser iguales, pueden estar fuertemente modificadas, pero aún así son la misma estructura, esto implica ke hay una transformación de las homologías, así entran en juego otros factores a nuestra definición como son los procesos ke desencadenan esas diferencias, como la interacción genética y la dinámica de las poblaciones.
En un análisis cladístico no hay una preocupación directa por todos esos fenómenos asociados a la homología. Esos procesos son más de interés de otras ramas de la biología como la evolución, le biología del desarrollo, la genética, la biología molecular, la genética de poblaciones, la ecología, etc. pero ke no sean de interés no implica ke sean ignorables: en la definición de caracteres (homologos) para el análisis cladístico muchos de esos factores entran en consideración para dar limite a los caracteres y como codificarlos.
Desafortunadamente, el ke el 'proceso' no sea preocupación de la cladística a llevado a la errónea idea ke es irrelevante en la definición de los caracteres. Así los cladístas de patrón (como Nelson, Platnick, o más recientemente Pleijel y Brower) consideran ke la cladística esta totalmente libre de todo pensamiento evolutivo, pero todo lo contrario: el usar caracteres homologos implica ke estamos trabajando en un marco ke esta basado en el origen por ancestría común, por lo tanto todos los estados de un carácter--y no solo los apomorficos--son la misma estructura.
Pero no solo eso, la implicación evolutiva es enorme para muchos caracteres, en especial si tenemos un buen conocimiento para el carácter, por lo ke la idea de Kluge (e.g. 2003) de ke solo es necesario 'descendencia con modificación' es también errada. Al introducir herencia y evolución, muchas más cosas se encuentran embebidas en la definición de cada carácter. Claro esta, eso depende de cada carácter, en muchas situaciones solamente tenemos un conocimiento morfológico del carácter, con lo ke lo único ke podemos asumir es la simple 'descendencia con modificación', pero en muchos grupos (como los vertebrados), el conocimiento ke se tiene sobre los caracteres es mucho más amplio, incluyendo por ejemplo un buen conocimiento del desarrollo de la estructura.
Parsimonia: algoritmo
El algoritmo básico de parsimonia es bastante sencillo. Si keremos establecer el estado de carácter presente en un nodo A, cuyos descendientes son B y C, y de los cuales ya sabemos su estado de carácter, el estado de A es la intersección de los estados de B y C, si esa intersección es un conjunto vacío entonces el estado es la unión de los estados B y C. Esta unión implica ke ha habido un cambio (lo ke es llamado un paso). Optimizar los estados es un poko más complicado, pero esas operaciones son lo básico del algoritmo.
Asignación de estados con el algoritmo de parsimonia: (A) y (B) el estado del nodo ancestro 'x' es igual a la intersección del estado de los descendientes 'y', 'z', en este caso, blanco; (C) 'y' y 'z' no comparten ningún estado, por lo tanto el estado asignado a 'x' es la unión de los estados de los descendientes.Es fácil notar ke el algoritmo es 'independiente' de los datos tomados, uno puede tomar cualkier clase de 'carácter' (en ciencias de computadores este problema es conocido como el problema del coloreado, y los 'caracteres' son colores en un mapa) y cualkier clase de terminales de análisis. Esa es una característica propia de los métodos formalizados como algoritmos. Dado esas circunstancias es necesario proveer una justificación del uso de un algoritmo particular para cada problema.
En cladística la justificación para el uso de parsimonia viene dada por el concepto de homología.
Parsimonia: análisis cladístico
Dado el concepto evolutivo de homología, su unión con el concepto algoritmico de parsimonia es directo. Si decimos ke un carácter, cualquiera ke sea su estado, es el mismo en dos organismos ke comparten un ancestro común, eso implica ke ese carácter es heredado del ancestro común, por lo tanto el ancestro común debe tener ese carácter.
Si además ambos organismos comparten la misma expresión del carácter, es decir su mismo estado, entonces ese mismo estado se encuentra en su ancestro común, si por el contrario ambos organismos tienen diferentes expresiones del carácter, el ancestro común tiene el carácter, pero no sabemos ke expresión pueda tener, por lo ke asumimos ke tiene alguno de los dos estados, y damos por hecho, ke si en realidad ambos organismos comparten ese carácter, debió haber ocurrido una transformación.
Como se puede ver, esa es exactamente la misma descripción del algoritmo de parsimonia. En cladística el algoritmo esta justificado pues los caracteres usados son tomados como homologos. En ese contexto el algoritmo maximiza nuestras proposiciones de homología al minimizar las transformaciones: de esa forma se asegura ke la mayor cantidad de organismos con el mismo estados de carácter kedan contiguos. Esto es minimización de hipótesis ad hoc de homoplasia (Farris, 1983).
Parsimonia y los tests de homología
Es una idea común entre los cladístas el argumentar ke parsimonia es un test de homología: la congruencia. Yo estoy en desacuerdo, pues como he argumentado akí la base de la parsimonia es asumir de plano ke las entradas para un carácter son homologas. Para mi los tests de homología son una condición previa del análisis cladístico.
Esos test pueden venir de muchas formas, bien sea basado en argumentos de morfología (usualmente llamada 'similitud'), posición anatómica, organización estructural, ontogenía, genética, en la mayor parte de los casos muchas de esas pruebas son conjuntas--por lo ke definir un carácter marca una fuerte carga teórica--es despues de haber examinado todas esas alternativas ke podemos tener un buen carácter. Es por eso ke una homoplasía es una hipótesis ad hoc: la homoplasia en el carácter solo esta justificada en relación con el cladograma.
Por supuesto, es necesario recurrir a la revisión de caracteres, y ciertamente es bueno revisar los caracteres homoplasicos, pero es igualmente valioso examinar todos los caracteres por igual. Es posible ke tengamos alguna duda acerca de como codificar caracteres, debido a su complejidad o porke son muy poko conocidos, en estos casos puede ser posible utilizar como apoyo el resultado ke producen diferentes tipos de codificación (esa es la alternativa ke usa Ramírez, en prensa, y en molecular Wheeler, 1996). Pero debo llamar la atención sobre eso: la codificación no esta 'avalada' por el cladograma, pues el argumento ke se utiliza para defender esa codificación es el mismo usado para la homoplasia: solo se justifica en relación con el cladograma, por lo tanto la codificación tiene un fuerte elemento ad hoc. Para esos casos yo preferiría ke se utilizara un eskema de pesado como el de Neff (1986): dado ke conocemos poko del carácter, tanto ke tenemos nuestras dudas para codificarlo, es mejor ke tenga un peso más bajo ke el de caracteres ke conocemos mejor.
**
Adicional: una especulación histórica
En este ensayo yo presente las ideas de homología y parsimonia por separado y luego las fusione. Lo hice así porke considero ke da claridad a la discusión. Sin embargo, históricamente el desarrollo fue muy entrelazado desde el inicio. Leyendo a Wagner (1961, desde el lado algoritmico) y a Hennig (1968, desde el lado lógico) es claro como ambas ideas son intercambiables desde el inicio. Ambas visiones kedaron integradas muy bien en Farris et al. 1970, ke a pesar de ser un intento de axiomatización, contiene ideas (para mi muy estimulantes) ke coinciden con muchos de los puntos expuestos akí. Así desde el inicio las ideas del análisis cladístico y de parsimonia estuvieron juntas.
Wagner, Hennig y Farris desarrollaron sus ideas desde un contexto de la morfología. Dayoff (1969) también experimento con varios algoritmos para secuencias moleculares, donde al menos por ahora, las ideas de homología son en muchos casos diferentes a el concepto de morfología, en su época (mediados de los 60s), no se exactamente ke pensaran de homología los moleculares, pero era claro ke ellos no creían ke dos bases (en el caso de Dayoff, dos aminoacidos) iguales en dos organismos implicaba un origen común, la idea de mutaciones constantes precluyo la idea. Es interesante ver ke Dayoff no pudo encontrar una forma adecuada de asignar estados en los ancestros.
Referencias
Dayoff, M.O. (1969) Computer analysis of protein evolution. Sci. Amer. 221: 87-95.
Farris, J.S. (1983) The logical basis of phylogenetic systematics. En: Platnick, N., Funk, V.A. (Eds.), Advances in cladistics, vol. 2. Columbia, New York. Pp. 7-36.
Farris, J.S., Kluge, A.G., Eckardt, M.J. (1970) A numerical approach to phylogenetic systematics. Syst. Zool. 19: 172-189.
Hennig, W. (1968) Elementos de una sistemática filogenética. Eudeba, Buenos Aires.
Kluge, A.G. (2003) On deduction of species relationships: a précis. Cladistics 19: 233-239.
Neff, N.A. (1986) A rational basis for a priori character weighting. Syst. Zool. 35: 110-123.
Ramírez, M.J. (2007) Homology as a parsimony problem: a dynamic homology approach for morgological data. Cladistics 23:xx-xx.
Wagner, W.H. (1961) Problems in the classification of ferns. En: Recent advances in botany, vol. 1 Univ. Toronto, Toronto. Pp. 841-844.
Wheeler, W.C. (1996) Optimization alignment: the end of multiple sequence alignment in phylogenetics? Cladistics 12: 1-9.
jueves, noviembre 01, 2007
Pasando a BPR3
Ayer me entere ke existe una iniciativa, BPR3, blogs sobre investigación revisada por pares, y muy contento con esa idea me voy a poner a modificar mis post previos :), y realizar los futuros dentro de esa iniciativa :D

Pueden ver lo ke dice la gente de BPR3 en su blog oficial: bpr3.org
Y aquí pueden leer un poco sobre lo ke pienso de esa idea (en mi blog personal).

Pueden ver lo ke dice la gente de BPR3 en su blog oficial: bpr3.org
Y aquí pueden leer un poco sobre lo ke pienso de esa idea (en mi blog personal).
martes, septiembre 11, 2007
Tasa de 'poder explicativo'
Recientemente, Grant y Kluge (2007) propusieron la tasa de poder explicativo, una medida basada en el soporte de Bremer, para medir el soporte de las ramas. A diferencia del soporte de Bremer tradicional, esta medida es escalada, con lo cual es posible comparar el valor de soporte de varias ramas. En este pequeño ensayo trato de mostrar ke su medida no logra su cometido, y ke las soluciones ke ellos critican son mucho mejores.
La tasa de poder explicativo (REP por su nombre en ingles, Ratio of explanatory power) consiste en escalar el soporte de Bremer con la diferencia entre el mejor árbol posible y el peor. Es decir, el denominador es el numerador del indice de retención para todo el árbol. Así si un clado solo desaparece en el peor árbol, tanto el denominador como el numerador son iguales, y por lo tanto el REP=1, si un grupo es contradicho (o es ambiguo) y soportado en los árboles más parsimoniosos (i.e. esta en algunos árboles, pero no aparece en el consenso estricto), el REP=0. Aunque no es una idea explorada por Grant y Kluge, uno puede extender la medida para grupos no soportados entre los árboles más parsimoniosos, dando valores negativos, cuyo minimo seria REP=-1, en donde un clado solo aparezca en el peor árbol (en este caso, nunca aparece!).
Como es un valor escalado del soporte de Bremer, su utilidad se limita a comparar entre clados de diferentes conjuntos de datos, pero no para comprar clados dentro de un mismo conjunto de datos. Grant y Kluge, intentan dar una enumeración de los posibles usos de su medida, como por ejemplo para discriminar hipótesis biogeograficas, o para hacer superaboles. Es interesante hacer notar, que en un paper anterior, Grant y Kluge (2003, p. 384) estaban en contra de utilizar los valores de soporte altos, entre otras cosas para hacer análisis biogeograficos:
“[E]pistemologically, all clades in the least refuted cladogram(s) provide an equally valid basis for testing evolutionary scenarios (e.g., biogeographic hypotheses)...although it may be tempting to base such studies exclusively on the more strongly supported parts of a given cladogram, that procedure is a slippery slope to a verificationist interpretation of better-supported clades as more reliable or certain or closer to truth.”
Así si según Grant y Kluge uno no debe comparar entre clados de un mismo conjunto de datos, porke si podemos hacerlo con diferentes conjuntos? Es más, tiene más lógica ke uno compare el soporte de un clado de Dipteros, con por ejemplo, un clado de Asteraceas o Dinoflagelados?
Una utilidad ke parece más real, aunque de plano rechazada por Granty y Kluge, es komo una medida similar al soporte particionado de Bremer (Grant y Kluge estan en contra, continuamente recalcan ke se debe usar para comparar diferentes conjuntos de terminales) o en experimentos como los de Rydin y Källersjö (2002; que son de diferentes terminales, pero del mismo conjunto de datos). En ambos casos, el escalamiento es relativo al conjunto usado y permitiría ver en ke conjunto de datos, el soporte es relativamente más fuerte para cada grupo. Quizá sea la única utilidad posible para este método, ke de otro lado es solo una forma más difícil de expresar el soporte de Bremer.
Parecería mucho mejor una medida ke mostrar cuanto de la evidencia apoya al grupo en el árbol actual y cuan otra lo contradice! Grant y Kluge parecen no distinguir eso, ellos si un grupo es soportado en todos los árboles más parsimoniosos implica, como en la democracia norteamericana, ke toda la evidencia posible favorece al grupo:
“For example, two groups on the same tree may have identical [Bremer support] (and, therefore, identical relative explanatory power)” (p. 484).
Pero dado un conjunto de datos, uno puede tener caracteres ke apoyan la rama sin ser contradichos, mientras en otra rama, los caracteres ke la soportan son contradichos con homoplasia dentro del clado (i.e reversiones) o en ramas independientes (i.e. 'paralelismos'). En casos reales, lo más probable es ke tengamos el segundo caso, pero en diferentes grados, y una medida de soporte escalada debería ser capaz de distinguir esos grados. Esa medida es la Diferencia relativa de ajuste (RFD, relative fit difference, Goloboff y Farris, 2001).
La ventaja de esa medida escalada, discriminar grupos de acuerdo con la proporción de evidencia favorable ke tienen. Así, en contra de Grant y Kluge, uno puede encontrar ke hay grupos ke a pesar de tener un soporte de Bremer igual (es decir, ke tienen la misma cantidad de evidencia favorable), son contradichos de forma diferente, y por lo tanto no podemos decir ke tienen el mismo poder explicativo! Usando las palabras de Grant y Kluge, uno de los grupos esta más refutado (con más evidencia en contra) que el otro! Esa razón esta claramente expuesta en Goloboff y Farris, 2001 (p. S30-S31), contra Grant y Kluge que argumentan ke el RFD fue presentado sin mostrar porke dicha medida es útil, cosa ke al REP le keda más difícil demostrar!
El RDF no es, ni pretende ser una alternativa al soporte de Bremer, más bien, es un complemento a esa medida, tanto mejor para medir el 'poder explicativo'. Cabe anotar ke el RFD esta implementado en las versiones más recientes de NONA/Piwe (Goloboff, 1998) y en TNT (Goloboff et al., 2007).
El RDF no es perfecto, tal y como lo muestran Goloboff y Farris. Para mi el principal, es ke depende mucho de la pareja de árboles en las ke se basa la comparación. Así ke el tema rekiere seguir siendo investigado, pero más en el sentido de RFD (discriminar grupos dentro de un mismo conjunto de datos) que en el REP (escalar el soporte de Bremer con una constante).
Nota lateral: En el paper del 2003, Grant y Kluge cuando discuten el soporte de Bremer no lo relacionan con su definición de soporte tal como parece kedar implícito en el paper del 2007. Hasta donde se, ese enlace fue hecho por primera vez, al menos de forma impresa, por Ramírez (2005) en una replica al paper de Grant y Kluge.
Goloboff, P.A. (1998) NONA/Piwe, programa de computo disponible gratis en: http://www.zmuc.dk/public/phylogeny/Nona-PeeWee/
Goloboff, P.A., Farris, J.S. (2001) Cladistics 17: S26-S34.
Goloboff, P.A. (2007) TNT, programa de computo, demo disponible en: http://www.zmuc.dk/public/phylogeny/TNT/
Grant, T., Kluge, A. G. (2007) Mol. Phyl. Evol. 44: 483-487.
Grant, T., Kluge, A. G. (2003) 19: Cladistics, 379-418.
Ramírez, M. J. (2005) Cladistics 21: 83-89.
Rydin, C., Källersjö, M. (2002) Cladistics 18: 475-513.
La tasa de poder explicativo (REP por su nombre en ingles, Ratio of explanatory power) consiste en escalar el soporte de Bremer con la diferencia entre el mejor árbol posible y el peor. Es decir, el denominador es el numerador del indice de retención para todo el árbol. Así si un clado solo desaparece en el peor árbol, tanto el denominador como el numerador son iguales, y por lo tanto el REP=1, si un grupo es contradicho (o es ambiguo) y soportado en los árboles más parsimoniosos (i.e. esta en algunos árboles, pero no aparece en el consenso estricto), el REP=0. Aunque no es una idea explorada por Grant y Kluge, uno puede extender la medida para grupos no soportados entre los árboles más parsimoniosos, dando valores negativos, cuyo minimo seria REP=-1, en donde un clado solo aparezca en el peor árbol (en este caso, nunca aparece!).
Como es un valor escalado del soporte de Bremer, su utilidad se limita a comparar entre clados de diferentes conjuntos de datos, pero no para comprar clados dentro de un mismo conjunto de datos. Grant y Kluge, intentan dar una enumeración de los posibles usos de su medida, como por ejemplo para discriminar hipótesis biogeograficas, o para hacer superaboles. Es interesante hacer notar, que en un paper anterior, Grant y Kluge (2003, p. 384) estaban en contra de utilizar los valores de soporte altos, entre otras cosas para hacer análisis biogeograficos:
“[E]pistemologically, all clades in the least refuted cladogram(s) provide an equally valid basis for testing evolutionary scenarios (e.g., biogeographic hypotheses)...although it may be tempting to base such studies exclusively on the more strongly supported parts of a given cladogram, that procedure is a slippery slope to a verificationist interpretation of better-supported clades as more reliable or certain or closer to truth.”
Así si según Grant y Kluge uno no debe comparar entre clados de un mismo conjunto de datos, porke si podemos hacerlo con diferentes conjuntos? Es más, tiene más lógica ke uno compare el soporte de un clado de Dipteros, con por ejemplo, un clado de Asteraceas o Dinoflagelados?
Una utilidad ke parece más real, aunque de plano rechazada por Granty y Kluge, es komo una medida similar al soporte particionado de Bremer (Grant y Kluge estan en contra, continuamente recalcan ke se debe usar para comparar diferentes conjuntos de terminales) o en experimentos como los de Rydin y Källersjö (2002; que son de diferentes terminales, pero del mismo conjunto de datos). En ambos casos, el escalamiento es relativo al conjunto usado y permitiría ver en ke conjunto de datos, el soporte es relativamente más fuerte para cada grupo. Quizá sea la única utilidad posible para este método, ke de otro lado es solo una forma más difícil de expresar el soporte de Bremer.
Parecería mucho mejor una medida ke mostrar cuanto de la evidencia apoya al grupo en el árbol actual y cuan otra lo contradice! Grant y Kluge parecen no distinguir eso, ellos si un grupo es soportado en todos los árboles más parsimoniosos implica, como en la democracia norteamericana, ke toda la evidencia posible favorece al grupo:
“For example, two groups on the same tree may have identical [Bremer support] (and, therefore, identical relative explanatory power)” (p. 484).
Pero dado un conjunto de datos, uno puede tener caracteres ke apoyan la rama sin ser contradichos, mientras en otra rama, los caracteres ke la soportan son contradichos con homoplasia dentro del clado (i.e reversiones) o en ramas independientes (i.e. 'paralelismos'). En casos reales, lo más probable es ke tengamos el segundo caso, pero en diferentes grados, y una medida de soporte escalada debería ser capaz de distinguir esos grados. Esa medida es la Diferencia relativa de ajuste (RFD, relative fit difference, Goloboff y Farris, 2001).
La ventaja de esa medida escalada, discriminar grupos de acuerdo con la proporción de evidencia favorable ke tienen. Así, en contra de Grant y Kluge, uno puede encontrar ke hay grupos ke a pesar de tener un soporte de Bremer igual (es decir, ke tienen la misma cantidad de evidencia favorable), son contradichos de forma diferente, y por lo tanto no podemos decir ke tienen el mismo poder explicativo! Usando las palabras de Grant y Kluge, uno de los grupos esta más refutado (con más evidencia en contra) que el otro! Esa razón esta claramente expuesta en Goloboff y Farris, 2001 (p. S30-S31), contra Grant y Kluge que argumentan ke el RFD fue presentado sin mostrar porke dicha medida es útil, cosa ke al REP le keda más difícil demostrar!
El RDF no es, ni pretende ser una alternativa al soporte de Bremer, más bien, es un complemento a esa medida, tanto mejor para medir el 'poder explicativo'. Cabe anotar ke el RFD esta implementado en las versiones más recientes de NONA/Piwe (Goloboff, 1998) y en TNT (Goloboff et al., 2007).
El RDF no es perfecto, tal y como lo muestran Goloboff y Farris. Para mi el principal, es ke depende mucho de la pareja de árboles en las ke se basa la comparación. Así ke el tema rekiere seguir siendo investigado, pero más en el sentido de RFD (discriminar grupos dentro de un mismo conjunto de datos) que en el REP (escalar el soporte de Bremer con una constante).
Nota lateral: En el paper del 2003, Grant y Kluge cuando discuten el soporte de Bremer no lo relacionan con su definición de soporte tal como parece kedar implícito en el paper del 2007. Hasta donde se, ese enlace fue hecho por primera vez, al menos de forma impresa, por Ramírez (2005) en una replica al paper de Grant y Kluge.
Goloboff, P.A. (1998) NONA/Piwe, programa de computo disponible gratis en: http://www.zmuc.dk/public/phylogeny/Nona-PeeWee/
Goloboff, P.A., Farris, J.S. (2001) Cladistics 17: S26-S34.
Goloboff, P.A. (2007) TNT, programa de computo, demo disponible en: http://www.zmuc.dk/public/phylogeny/TNT/
Grant, T., Kluge, A. G. (2007) Mol. Phyl. Evol. 44: 483-487.
Grant, T., Kluge, A. G. (2003) 19: Cladistics, 379-418.
Ramírez, M. J. (2005) Cladistics 21: 83-89.
Rydin, C., Källersjö, M. (2002) Cladistics 18: 475-513.
martes, noviembre 14, 2006
Secuencias ambientales y monstruos marinos
Hace poco alguien que estaba 'asesorando' a una estudiante en su trabajo de grado, me pregunto si para trabajar secuencias ambientares era necesario hacer un análisis filogenético. El y la estudiante solo querían hacer alineamientos simples usando FASTA.
Sidenote: Las secuencia ambiental es el resultado de amplificar material genético (generalmente ARN ribosomal) con el fin de identificar población bacteriana en sustratos donde es prácticamente imposible aislar bacterias domesticables (i.e. cultivables), como el fondo de los océanos, termales, pozos de desperdicios, pero también en sustratos más tradicionales, como suelos. Estos estudios han detectado una gran variedad de bacterias y archeas nunca vistas (de hecho en muchos casos solo se conocen sus secuencias!) (DeLong & Pace, 2001).
Los alineamientos simples de FASTA son muy poderosos, permiten una identificación rápida y acertada de secuencias conocidas. Un bonito ejemplo es el de Carr (2002), quienes secuenciaron unos misteriosos-y putrefactos-restos encontrados en la costa de Bermudas, su análisis mostró que los restos eran de un cachalote, un monstruo marino-que lo diga Herman Melville ;)-pero nada que excite a un 'cripto-zoólogo'.
Los alineamientos simples de FASTA no son más que claves, automatizadas, muy similares a las hechas por los taxónomos desde los tiempos de Linneo (de eso hablare en otro post), y son muy buenos si las secuencias ya se encuentran en GenBank.
Pero la filogenética es una herramienta mucho más fuerte que FASTA. En primer lugar no se requiere que la secuencia del individuo a identificar no este en GenBank, eso si, se necesita, como en todo análisis filogenético, que existan secuencias de parientes de los individuos en el nivel de resolución deseado. Es ese el aspecto que hace este análisis superior a FASTA: la(s) muestra(s) prueba(s) es(son) colocada(s) en un clado especifico. Los niveles de similitud de FASTA no pueden conseguirlo, pues, como bien saben los cladistas desde tiempos de Hennig, esta similitud no distingue entre similitud por apomorfía o por plesiomorfía. Es posible hacer una identificación partiendo desde niveles filogenéticos amplios e ir mejorando la resolución hacia niveles más restringidos. En el mundo de los cetaceos, ya algunos autores propusieron un protocolo de este estilo para identificar especies de ballenas (Ross & Murugan, 2006), aunque su método se ve afectado porque usan Neighbuor joining (Farris et al., 1996), el espíritu de su procedimiento es como el que defiendo aquí.
La filogenética ya se a usado para hacer identificaciones. Por ejemplo, Allard et al. (1995) demostro que un supuesto ADN de Dinosaurio del mesozoico-a la parque jurasico-era en realidad una contaminación humana. O en un estudio más clásico, (Ou et al., 1992) que había evidencia que un odontólogo contagio a sus pacientes!
Todos estos ejemplos (y si uno se pone a buscar va a encontrar muchos más!) muestran uno de
los puntos más fuertes que tiene el análisis filogenético: y es su asombroso poder predictivo, un punto que, misteriosamente para mi al menos, a sido objetado por alguno que otro de los eminentes 'popperianos' y/o 'realistas escepticos' que escriben sobre la filosofía del análisis filogenético (no los cito, porque de ellos hablare después en otro post). FASTA, aunque muy eficaz, no tiene ese poder predictivo al no poner nunca las secuencias en un contexto filogenético.
Sidenote: Dado que en las secuencias ambientales la expectativa es encontrar cosas desconocidas y nuevas, yo recomendé a mi colega y la estudiante que realizaran el análisis filogenético. Ellos no se veían muy convencidos. Tristemente, aun en estos dias, la gente cree que el análisis filogenético es algo críptico y difícil de comprender :(.
Allard et al. (1995) Science 268: 1192.
Carr, S.M. et al. (2002) Biol. Bull 202: 1-5.
DeLong, E.F., Pace, N.R. (2001) Syst. Biol. 50: 470-478.
Farris, J.S. et al. (1996) Cladistics 12: 99-124.
Ou, C-Y. et al. (1992) Science 256: 1165-1171.
Ross, H.A., Murugan, S. (2006) Mol. Phyl. Evol. 40: 866-871.
Sidenote: Las secuencia ambiental es el resultado de amplificar material genético (generalmente ARN ribosomal) con el fin de identificar población bacteriana en sustratos donde es prácticamente imposible aislar bacterias domesticables (i.e. cultivables), como el fondo de los océanos, termales, pozos de desperdicios, pero también en sustratos más tradicionales, como suelos. Estos estudios han detectado una gran variedad de bacterias y archeas nunca vistas (de hecho en muchos casos solo se conocen sus secuencias!) (DeLong & Pace, 2001).
Los alineamientos simples de FASTA son muy poderosos, permiten una identificación rápida y acertada de secuencias conocidas. Un bonito ejemplo es el de Carr (2002), quienes secuenciaron unos misteriosos-y putrefactos-restos encontrados en la costa de Bermudas, su análisis mostró que los restos eran de un cachalote, un monstruo marino-que lo diga Herman Melville ;)-pero nada que excite a un 'cripto-zoólogo'.
Los alineamientos simples de FASTA no son más que claves, automatizadas, muy similares a las hechas por los taxónomos desde los tiempos de Linneo (de eso hablare en otro post), y son muy buenos si las secuencias ya se encuentran en GenBank.
Pero la filogenética es una herramienta mucho más fuerte que FASTA. En primer lugar no se requiere que la secuencia del individuo a identificar no este en GenBank, eso si, se necesita, como en todo análisis filogenético, que existan secuencias de parientes de los individuos en el nivel de resolución deseado. Es ese el aspecto que hace este análisis superior a FASTA: la(s) muestra(s) prueba(s) es(son) colocada(s) en un clado especifico. Los niveles de similitud de FASTA no pueden conseguirlo, pues, como bien saben los cladistas desde tiempos de Hennig, esta similitud no distingue entre similitud por apomorfía o por plesiomorfía. Es posible hacer una identificación partiendo desde niveles filogenéticos amplios e ir mejorando la resolución hacia niveles más restringidos. En el mundo de los cetaceos, ya algunos autores propusieron un protocolo de este estilo para identificar especies de ballenas (Ross & Murugan, 2006), aunque su método se ve afectado porque usan Neighbuor joining (Farris et al., 1996), el espíritu de su procedimiento es como el que defiendo aquí.
La filogenética ya se a usado para hacer identificaciones. Por ejemplo, Allard et al. (1995) demostro que un supuesto ADN de Dinosaurio del mesozoico-a la parque jurasico-era en realidad una contaminación humana. O en un estudio más clásico, (Ou et al., 1992) que había evidencia que un odontólogo contagio a sus pacientes!
Todos estos ejemplos (y si uno se pone a buscar va a encontrar muchos más!) muestran uno de
los puntos más fuertes que tiene el análisis filogenético: y es su asombroso poder predictivo, un punto que, misteriosamente para mi al menos, a sido objetado por alguno que otro de los eminentes 'popperianos' y/o 'realistas escepticos' que escriben sobre la filosofía del análisis filogenético (no los cito, porque de ellos hablare después en otro post). FASTA, aunque muy eficaz, no tiene ese poder predictivo al no poner nunca las secuencias en un contexto filogenético.
Sidenote: Dado que en las secuencias ambientales la expectativa es encontrar cosas desconocidas y nuevas, yo recomendé a mi colega y la estudiante que realizaran el análisis filogenético. Ellos no se veían muy convencidos. Tristemente, aun en estos dias, la gente cree que el análisis filogenético es algo críptico y difícil de comprender :(.
Allard et al. (1995) Science 268: 1192.
Carr, S.M. et al. (2002) Biol. Bull 202: 1-5.
DeLong, E.F., Pace, N.R. (2001) Syst. Biol. 50: 470-478.
Farris, J.S. et al. (1996) Cladistics 12: 99-124.
Ou, C-Y. et al. (1992) Science 256: 1165-1171.
Ross, H.A., Murugan, S. (2006) Mol. Phyl. Evol. 40: 866-871.
lunes, enero 30, 2006
Buenas y malas filogenias
¿Podemos confiar en un análisis filogenético si las relaciones de un grupo se alejan de la expectativa (dada por otros análisis anteriores)?
Seguramente la pregunta no tendría sentido para autores como Kluge. Según sus razonamientos (Kluge 1997, 2003) cada vez conocemos más acerca de la filogenia de un grupo gracias a los 'ciclos de investigación'. Nuestra pregunta implicaría que no hemos usado toda la evidencia disponible para dilucidar las relaciones. Como ejemplo empírico usaría, por ejemplo, Eernisse y Kluge (1993), más recientemente. Pero, el dibujo de Kluge, de adicionar cada vez más evidencia no es tan sencillo. A la hora de planear un análisis filogenético no solo es crucial usar la mayor parte de la evidencia ya usada, sino también hacer una selección cuidadosa de los terminales. Por ejemplo, Giribet et al. (2001) al hacer el análisis filogenético de los artrópodos no utilizó toda la información que el y sus colaboradores han compilado en otros problemas (por ejemplo las relaciones entre los insectos y los quelicerados). Un análisis de esa magnitud, aunque fuera manejable, tendría el problema en que los crustaceos quedarían fuertemente submuestreados.
Pero uno de los resultados más sorprendentes de Giribet et al. es que ¡Drosophila y un dipluro forman parte de los crustaceos! Esto contradice toda expectativa, puesto que la monofilia de los hexápodos es muy bien fundamentada a nivel morfológico y molecular (Wheeler et al. 2001), y las moscas forman claramente un grupo muy anidado dentro de hexápoda. Giribet et al., reconocen las limitaciones de su análisis y no hacen afirmaciones dentro de la filogenia de pancrustacea (aunque aseguran que el grupo esta bine soportado). Es una lastima que Giribet et al. no revisarán que pasaba si se excluía estos taxones problemáticos, aunque en su análisis Drosophila se encuentra dentro de los insectos en el 91% de los costos explorados. De mi experiencia con los datos morfológicos usados por Giribet et al., eliminar Drosophila, Diplura y Balanus (los taxones problemáticos), ambas topologías comparten 11 sinapomorfias (de 12 en los datos completos) para Mandibulata, cuatro (5) para Myriapoda, cinco (7) para Pancrustacea, 6 (8) para Crustacea todas y (7) para Hexapoda, además de que 6 de esos caracteres dejan de ser homoplasicos. Así que para el objetivo de Giribet et al. que es establecer las relaciones entre los grandes grupos de artrópodos ('clases') los resultados son muy acertados. Seria un error que ellos pensaran que su conjunto de datos provee evidencia para contradecir la monofilia de los hexápodos, puesto que el conjunto de la evidencia que ellos examinan se solapa ligeramente ¡con intentos más exhaustivos para esta materia!
De otro lado, otros estudios pueden no ser tan concluyentes. Por ejemplo Silva de Paula et al. (2005) no tienen la precaución de Giribet et al. En un conjunto de datos fuertemente sesgado hacia los triatomineaos (chinches transmisoras de la enfermedad de Chagas) y algunas muestras de sus familiares (chinches reduviidos), llegan a la conclusión que los triatomineos no son monofileticos. La filogenia de los reduviidos no es bien conocida, pero si existe algo esperado en este grupo es la monofilia de los harpactorinos (otros reduviidos) (e.g. Davis 1969, Clayton 1990). Algunas sinapomorfias de este grupo son la membranización del abdomen, una celda cuadrada en el hemielitro, espina tibial y ausencia de glándula vermiforme. Pero en el análisis usado por Silva de Paula, los nueve harpactorinos escogidos parecen divididos en dos grupos de tres terminales cada uno y los otros tres dispersos en otras ramas del análisis. Los únicos grupos reconocidble son Rhondius y Triatomini que constituyen la mayor parte de los terminales del estudio y cuya monofilia queda supuestamente refutada. Silva de Paula et al. simplemente recogieron los taxones disponibles en genbank sin preocuparse por hacer una selección o determinar si la excesiva asimetría del estudio podía sesgar sus resultados. En una exploración preliminar que yo realice usando algunos de esos terminales los resultados dependen según los terminales escogidos.
Ridyn y Källersjö (2002) propusieron un marco de trabajo que consiste en hacer matrices con un contenido diferente de taxones (aunque con los mismos datos). Lo deseable es que los resultados a nivel de los clados encontrados (y uno esperaría también de las sianpomorfias) fueran los mismos a pesar que los terminales son distintos. Un caso de la historia de la ciencia es más o menos claro: gracias a diferentes estudios que usan diferentes conjuntos de taxones, siempre hemos encontrado las relaciones entre los tres 'dominios' de la vida y de los grandes grupos ('reinos') de cada dominio. Casos similares se han visto con los 'filums' de animales y en el ejemplo de Rydin y Källersjö para grandes grupos de plantas vasculares. Esta clase de test podría ser de gran utilidad precisamente para detectar estas instancias de incongruencia entre resultados esperados y los encontrados (en especial cuando la expectativa es fuertemente destruida como en el caso de Giribet et al.). Esta idea es una extensión del trabajo de Siddall y Whiting (1999), si los resultados son productos de ramas largas, ¡pues deben diferir cuando se remueven las ramas largas!
La lección es que la selección de las relaciones filogenéticas no es tan sencilla como Kluge quiere verla. Una investigación simplemente no refuta una anterior, pues en muchos casos la comparación es prácticamente imposible. Existen factores externos al cladograma como tal, como lo son la selección de la evidencia a usar (por ejemplo, si un 'sistema' de caracteres esta solo explorado en un grupo eso puede sesgar el resultado solo hacia ese grupo, y aun sabiendo que esos caracteres están disponibles lo mejor puede ser excluirlos por el momento mientras se recopila información para completar las observaciones en otros taxa) y los terminales que se incluyen en el análisis. Una segunda fuente de decisiones externas constituye el limites en donde aceptamos lo que cladograma expresa, este limite seria preferible delimitarlo antes del análisis (e.g. Rydin y Källersjö 2002), en otras ocasiones es una exploración de los datos (e.g. Giribet et al. 2001) la que nos permite identificar ese limite.
Clayton, R.A. 1990. A phylogenetic analysis of the Reduviidae (Hemiptera: Heteroptera) with redescripotions of the subfamilies and tribes. Ph. D. Dissertation, George Washington University, D.C.
Davis, N.T. 1969. Contributions to the morphology and phylogeny of the Reduvioidea. Part IV. The harpactoroid complex. Ann. Entomol. Soc. Am. 62, 74-94.
Eernisse, D.J., Kluge, A.G. 1993. Taxonomic congruence versus total evidence, and amniote phylogeny inferred from fossils, molecules, and morphology. Mol. Biol. Evol. 19, 1170-1195.
Giribet, G., Edgecombe, G.D., Wheeler, W.C. 2001. Arthropod phylogeny based on eight molecular loci and morphology. Nature 413, 157-161.
Kluge, A.G. 1997. Sophisticated falsification and research cycles: consequences for differential character weighting in phylogenetic systematics. Zool. Scripta 26, 349-360.
Kluge, A.G. 2003. On deduction of species relationships: a précis. Cladistics 19, 233-239.
Rydin, C., Källersjö, M. 2002. Taxon sampling and seed plant phylogeny. Cladistics 18, 485-513.
Siddall, M.E., Whiting, M.F. 1999. Long-branch abstractions. Cladistics 15, 9-24.
Silva de Paula, A., Diotaiuti, L., Schofield, C.J. 2005. Testing the sister-group relationship of the Rhodniini and Triatomini (Insecta: Hemiptera: Reduviidae: Traitominae). Mol. Phyl. Evol. 35, 712-718.
Wheeler, W.C., Whiting, M., Wheeler, Q.D., Carpenter, J.M. 2001. The phylogeny of the extant hexapod orders. Cladistics 17, 113-169.
Seguramente la pregunta no tendría sentido para autores como Kluge. Según sus razonamientos (Kluge 1997, 2003) cada vez conocemos más acerca de la filogenia de un grupo gracias a los 'ciclos de investigación'. Nuestra pregunta implicaría que no hemos usado toda la evidencia disponible para dilucidar las relaciones. Como ejemplo empírico usaría, por ejemplo, Eernisse y Kluge (1993), más recientemente. Pero, el dibujo de Kluge, de adicionar cada vez más evidencia no es tan sencillo. A la hora de planear un análisis filogenético no solo es crucial usar la mayor parte de la evidencia ya usada, sino también hacer una selección cuidadosa de los terminales. Por ejemplo, Giribet et al. (2001) al hacer el análisis filogenético de los artrópodos no utilizó toda la información que el y sus colaboradores han compilado en otros problemas (por ejemplo las relaciones entre los insectos y los quelicerados). Un análisis de esa magnitud, aunque fuera manejable, tendría el problema en que los crustaceos quedarían fuertemente submuestreados.
Pero uno de los resultados más sorprendentes de Giribet et al. es que ¡Drosophila y un dipluro forman parte de los crustaceos! Esto contradice toda expectativa, puesto que la monofilia de los hexápodos es muy bien fundamentada a nivel morfológico y molecular (Wheeler et al. 2001), y las moscas forman claramente un grupo muy anidado dentro de hexápoda. Giribet et al., reconocen las limitaciones de su análisis y no hacen afirmaciones dentro de la filogenia de pancrustacea (aunque aseguran que el grupo esta bine soportado). Es una lastima que Giribet et al. no revisarán que pasaba si se excluía estos taxones problemáticos, aunque en su análisis Drosophila se encuentra dentro de los insectos en el 91% de los costos explorados. De mi experiencia con los datos morfológicos usados por Giribet et al., eliminar Drosophila, Diplura y Balanus (los taxones problemáticos), ambas topologías comparten 11 sinapomorfias (de 12 en los datos completos) para Mandibulata, cuatro (5) para Myriapoda, cinco (7) para Pancrustacea, 6 (8) para Crustacea todas y (7) para Hexapoda, además de que 6 de esos caracteres dejan de ser homoplasicos. Así que para el objetivo de Giribet et al. que es establecer las relaciones entre los grandes grupos de artrópodos ('clases') los resultados son muy acertados. Seria un error que ellos pensaran que su conjunto de datos provee evidencia para contradecir la monofilia de los hexápodos, puesto que el conjunto de la evidencia que ellos examinan se solapa ligeramente ¡con intentos más exhaustivos para esta materia!
De otro lado, otros estudios pueden no ser tan concluyentes. Por ejemplo Silva de Paula et al. (2005) no tienen la precaución de Giribet et al. En un conjunto de datos fuertemente sesgado hacia los triatomineaos (chinches transmisoras de la enfermedad de Chagas) y algunas muestras de sus familiares (chinches reduviidos), llegan a la conclusión que los triatomineos no son monofileticos. La filogenia de los reduviidos no es bien conocida, pero si existe algo esperado en este grupo es la monofilia de los harpactorinos (otros reduviidos) (e.g. Davis 1969, Clayton 1990). Algunas sinapomorfias de este grupo son la membranización del abdomen, una celda cuadrada en el hemielitro, espina tibial y ausencia de glándula vermiforme. Pero en el análisis usado por Silva de Paula, los nueve harpactorinos escogidos parecen divididos en dos grupos de tres terminales cada uno y los otros tres dispersos en otras ramas del análisis. Los únicos grupos reconocidble son Rhondius y Triatomini que constituyen la mayor parte de los terminales del estudio y cuya monofilia queda supuestamente refutada. Silva de Paula et al. simplemente recogieron los taxones disponibles en genbank sin preocuparse por hacer una selección o determinar si la excesiva asimetría del estudio podía sesgar sus resultados. En una exploración preliminar que yo realice usando algunos de esos terminales los resultados dependen según los terminales escogidos.
Ridyn y Källersjö (2002) propusieron un marco de trabajo que consiste en hacer matrices con un contenido diferente de taxones (aunque con los mismos datos). Lo deseable es que los resultados a nivel de los clados encontrados (y uno esperaría también de las sianpomorfias) fueran los mismos a pesar que los terminales son distintos. Un caso de la historia de la ciencia es más o menos claro: gracias a diferentes estudios que usan diferentes conjuntos de taxones, siempre hemos encontrado las relaciones entre los tres 'dominios' de la vida y de los grandes grupos ('reinos') de cada dominio. Casos similares se han visto con los 'filums' de animales y en el ejemplo de Rydin y Källersjö para grandes grupos de plantas vasculares. Esta clase de test podría ser de gran utilidad precisamente para detectar estas instancias de incongruencia entre resultados esperados y los encontrados (en especial cuando la expectativa es fuertemente destruida como en el caso de Giribet et al.). Esta idea es una extensión del trabajo de Siddall y Whiting (1999), si los resultados son productos de ramas largas, ¡pues deben diferir cuando se remueven las ramas largas!
La lección es que la selección de las relaciones filogenéticas no es tan sencilla como Kluge quiere verla. Una investigación simplemente no refuta una anterior, pues en muchos casos la comparación es prácticamente imposible. Existen factores externos al cladograma como tal, como lo son la selección de la evidencia a usar (por ejemplo, si un 'sistema' de caracteres esta solo explorado en un grupo eso puede sesgar el resultado solo hacia ese grupo, y aun sabiendo que esos caracteres están disponibles lo mejor puede ser excluirlos por el momento mientras se recopila información para completar las observaciones en otros taxa) y los terminales que se incluyen en el análisis. Una segunda fuente de decisiones externas constituye el limites en donde aceptamos lo que cladograma expresa, este limite seria preferible delimitarlo antes del análisis (e.g. Rydin y Källersjö 2002), en otras ocasiones es una exploración de los datos (e.g. Giribet et al. 2001) la que nos permite identificar ese limite.
Clayton, R.A. 1990. A phylogenetic analysis of the Reduviidae (Hemiptera: Heteroptera) with redescripotions of the subfamilies and tribes. Ph. D. Dissertation, George Washington University, D.C.
Davis, N.T. 1969. Contributions to the morphology and phylogeny of the Reduvioidea. Part IV. The harpactoroid complex. Ann. Entomol. Soc. Am. 62, 74-94.
Eernisse, D.J., Kluge, A.G. 1993. Taxonomic congruence versus total evidence, and amniote phylogeny inferred from fossils, molecules, and morphology. Mol. Biol. Evol. 19, 1170-1195.
Giribet, G., Edgecombe, G.D., Wheeler, W.C. 2001. Arthropod phylogeny based on eight molecular loci and morphology. Nature 413, 157-161.
Kluge, A.G. 1997. Sophisticated falsification and research cycles: consequences for differential character weighting in phylogenetic systematics. Zool. Scripta 26, 349-360.
Kluge, A.G. 2003. On deduction of species relationships: a précis. Cladistics 19, 233-239.
Rydin, C., Källersjö, M. 2002. Taxon sampling and seed plant phylogeny. Cladistics 18, 485-513.
Siddall, M.E., Whiting, M.F. 1999. Long-branch abstractions. Cladistics 15, 9-24.
Silva de Paula, A., Diotaiuti, L., Schofield, C.J. 2005. Testing the sister-group relationship of the Rhodniini and Triatomini (Insecta: Hemiptera: Reduviidae: Traitominae). Mol. Phyl. Evol. 35, 712-718.
Wheeler, W.C., Whiting, M., Wheeler, Q.D., Carpenter, J.M. 2001. The phylogeny of the extant hexapod orders. Cladistics 17, 113-169.
miércoles, enero 25, 2006
Selección de funciones en cladística
Este post esta basado en mi proyecto de grado, defendido públicamente el 17de enero de 2006. Pueden descargar la presentación y el manuscrito en mi página académica http://ciencias.uis.edu.co/labsist/salvador/relfun.htm. Este proyecto fue dirigido por Daniel Miranda.
Goloboff (1993) desarrollo una alternativa a la parsimonia tradicional (minimizar los pasos homoplasicos), que tiene en cuanta cuan homoplasico es el carácter. Entre más homoplasicos el aumento la distorsión (alejamiento del optimo ideal: no homoplasia) es mucho más pequeño. Es decir que si vamos a escoger entre dos soluciones, se prefiere la que minimize la homoplasia, en los caracteres menos homoplasicos (en vez de minimizarla en todos los caracteres sin importar la homoplasia).
Sin embargo, existen muchas funciones para conseguir esto (en piwe (Goloboff, 1998), tnt (Goloboff et al., 2005) y paup (Swofford, 2001) esta implementada h / (h + k), o su inverso, donde h es la homoplasia del carácter y k una constante de concavidad, entre menor sea el valor de k, mayor sera la fuerza a favor de los caracteres poco homoplasicos. En tnt es posible implementar otras funciones, por ejemplo ln(h), raiz de h, o una con los valores definidos por el usuario). Así que ¿como escogemos las funciones?
La solución usada aquí esta basada en el trabajo de Goloboff (1997) y Ramírez (2003). Consiste en escoger la función que maximize la estabilidad de la topología ante la perturbación de los datos. Las perturbaciones fueron basadas en jackknife: de caracteres, terminales o combinando las dos. También es posible observar la estabilidad revisando solamente el número de nodos muy bien soportados usando búsquedas rápidas (Goloboff & Farris, 2001).
En concordancia con el trabajo de Goloboff (1997), la parsimonia tradicional produjo los resultados menos resueltos y con la menor cantidad de nodos estables. El resultado más interesante para mi, fue que con pocas replicas (20 replicas de eliminación) y sin importar el árbol de referencia usado (uno con búsquedas exhaustivas o usando consensos de Goloboff-Farris), la cantidad nodos estables es la misma que con muchas replicas. Dado ese resultado en el manuscrito se sugirió una búsqueda rápida inicial y luego una búsqueda un poco más exhaustiva para la selección de una sola función.
Goloboff en sus observaciones del manuscrito argumenta que seria mejor explorar la estabilidad de los nodos sin importar el tipo de función usada (similar a la aproximación de Wheeler, 1995 y Giribet, 2003 para datos moleculares).
Creo que es posible tener una solución que tome lo mejor de las dos propuestas. Hacer una búsqueda rápida a lo largo de muchas funciones, y seleccionar el vecindario de funciones optimas, y limitando los resultados de la estabilidad a lo largo de las funciones, solo para los resultados óptimos en el (o los) vecindario(s) de funciones escogido.
Agradezco mucho a P. Goloboff y M. Ramírez que leyeron el manuscrito y proporcionaron muchas ideas y discusión, tanto en su comunicación conmigo, como en sus escritos sobre el tema. D. Miranda dirigió mi proyecto y su ayuda en mi estancia en la universidad es invaluable. El proyecto fue financiado por Colciencias (1102-05-13563).
Giribet, G. 2003. Stability in phylogenetic formulations and its relationships to nodal support. Syst. Biol. 52, 554-564.
Goloboff, P.A. 1993. Estimating character weights during tree search. Cladistics 9, 83-91.
Goloboff, P.A. 1997. Self-weighted optimization: tree searches and character state reconstructions under implied transformation costs. Cladistics 13, 225-245.
Goloboff, P.A. 1998. PiWe/NONA, manual y programa distribuido por el autor. Disponible en internet en http://www.zmuc.dk/public/phylogeny/Nona-PeeWee/
Goloboff, P.A., Farris, J.S. 2001. Methods for quick consensus estimation. Cladistics 17, s26-s34.
Goloboff, P.A., Farris, J.S., Nixon, K.C. 2005. TNT, manual y programa distribuido por los autores. Disponible en internet en http://www.zmuc.dk/public/phylogeny/TNT/
Ramírez, M.J. 2003. The spider subfamily Amaurobioidinae (Aranea, Anyphaenidae): a phylogenetic revision at the generic level. Bull. AMNH 277.
Swofford, D.L. 2001. Paup*, manual y programa distribuido por Sinauer, Sunderland (USA).
Wheeler, W.C. 1995. Sequence alignment, parameter sensitivity, and the phylogenetic analysis of molecular data. Syst Biol. 44, 321-331.
Goloboff (1993) desarrollo una alternativa a la parsimonia tradicional (minimizar los pasos homoplasicos), que tiene en cuanta cuan homoplasico es el carácter. Entre más homoplasicos el aumento la distorsión (alejamiento del optimo ideal: no homoplasia) es mucho más pequeño. Es decir que si vamos a escoger entre dos soluciones, se prefiere la que minimize la homoplasia, en los caracteres menos homoplasicos (en vez de minimizarla en todos los caracteres sin importar la homoplasia).
Sin embargo, existen muchas funciones para conseguir esto (en piwe (Goloboff, 1998), tnt (Goloboff et al., 2005) y paup (Swofford, 2001) esta implementada h / (h + k), o su inverso, donde h es la homoplasia del carácter y k una constante de concavidad, entre menor sea el valor de k, mayor sera la fuerza a favor de los caracteres poco homoplasicos. En tnt es posible implementar otras funciones, por ejemplo ln(h), raiz de h, o una con los valores definidos por el usuario). Así que ¿como escogemos las funciones?
La solución usada aquí esta basada en el trabajo de Goloboff (1997) y Ramírez (2003). Consiste en escoger la función que maximize la estabilidad de la topología ante la perturbación de los datos. Las perturbaciones fueron basadas en jackknife: de caracteres, terminales o combinando las dos. También es posible observar la estabilidad revisando solamente el número de nodos muy bien soportados usando búsquedas rápidas (Goloboff & Farris, 2001).
En concordancia con el trabajo de Goloboff (1997), la parsimonia tradicional produjo los resultados menos resueltos y con la menor cantidad de nodos estables. El resultado más interesante para mi, fue que con pocas replicas (20 replicas de eliminación) y sin importar el árbol de referencia usado (uno con búsquedas exhaustivas o usando consensos de Goloboff-Farris), la cantidad nodos estables es la misma que con muchas replicas. Dado ese resultado en el manuscrito se sugirió una búsqueda rápida inicial y luego una búsqueda un poco más exhaustiva para la selección de una sola función.
Goloboff en sus observaciones del manuscrito argumenta que seria mejor explorar la estabilidad de los nodos sin importar el tipo de función usada (similar a la aproximación de Wheeler, 1995 y Giribet, 2003 para datos moleculares).
Creo que es posible tener una solución que tome lo mejor de las dos propuestas. Hacer una búsqueda rápida a lo largo de muchas funciones, y seleccionar el vecindario de funciones optimas, y limitando los resultados de la estabilidad a lo largo de las funciones, solo para los resultados óptimos en el (o los) vecindario(s) de funciones escogido.
Agradezco mucho a P. Goloboff y M. Ramírez que leyeron el manuscrito y proporcionaron muchas ideas y discusión, tanto en su comunicación conmigo, como en sus escritos sobre el tema. D. Miranda dirigió mi proyecto y su ayuda en mi estancia en la universidad es invaluable. El proyecto fue financiado por Colciencias (1102-05-13563).
Giribet, G. 2003. Stability in phylogenetic formulations and its relationships to nodal support. Syst. Biol. 52, 554-564.
Goloboff, P.A. 1993. Estimating character weights during tree search. Cladistics 9, 83-91.
Goloboff, P.A. 1997. Self-weighted optimization: tree searches and character state reconstructions under implied transformation costs. Cladistics 13, 225-245.
Goloboff, P.A. 1998. PiWe/NONA, manual y programa distribuido por el autor. Disponible en internet en http://www.zmuc.dk/public/phylogeny/Nona-PeeWee/
Goloboff, P.A., Farris, J.S. 2001. Methods for quick consensus estimation. Cladistics 17, s26-s34.
Goloboff, P.A., Farris, J.S., Nixon, K.C. 2005. TNT, manual y programa distribuido por los autores. Disponible en internet en http://www.zmuc.dk/public/phylogeny/TNT/
Ramírez, M.J. 2003. The spider subfamily Amaurobioidinae (Aranea, Anyphaenidae): a phylogenetic revision at the generic level. Bull. AMNH 277.
Swofford, D.L. 2001. Paup*, manual y programa distribuido por Sinauer, Sunderland (USA).
Wheeler, W.C. 1995. Sequence alignment, parameter sensitivity, and the phylogenetic analysis of molecular data. Syst Biol. 44, 321-331.
Suscribirse a:
Entradas (Atom)
