There are some nice things on line lately :).
First, Rod Page talk at NHM of London, he puts his slide-show on line, there is also a video and a pod-cast of the talk (I do not see them :P). The slide-show can give a nice idea about the talk. There are two important things on the talk: (a) the importance of the availability of data! I just agree with him: data must be freely available, and previously published data must be easy accessible; and (b) Scientific data must be readily usable, for example the “Encyclopedia of life” is just a fun site just like wikipedia or the web tree of life, they are full of nice pics and info, but cientifically irrelevant: no hard data are attached to it (i.e. morfological info, character matrices), this contrast with the most simpler ispaces of Page!
Although I do not like some of the comments, Malte Ebach made a good point against the revival of a “pragmatic” classification (for horticulture, for example). I think that the only rule of classification is the phylogeny. If an immutable classification is the objective, the only way I can thing is an “alphabetic” or “numeric” system, but the preservation of “traditional” names is not the solution. Maybe, it is better a change of the rules...
And speaking of change of the rules, I recommend the continuos checking of several of Mike Keesey posts. I do not think that phylocode is the solution (in fact I think that it is even worse than traditional Linnean classification), but he writes interesting things about databasing!
Also, it is seems to change some rules about the publication of names for taxonomy (via Evolving thoughts). I thing that only peer reviewed publication must count, at least from the last fifty years or so. Also I thing that new names, at least for species must be published on journals instead of books...
Mostrando las entradas con la etiqueta databasing. Mostrar todas las entradas
Mostrando las entradas con la etiqueta databasing. Mostrar todas las entradas
sábado, marzo 21, 2009
jueves, octubre 30, 2008
Hennig XVII: “Live” blogging, day 3
Today was a highly theoretical day ;)... Most of the talks were highly methodological, with some scattered practical works.
Ward Wheeler, in a line similar to Grant and Kluge, argues that “objective support”, like Bremer support of Likelihood ratios are different to “average support”. His main argument, is that objective support is a better measure than average support, because it is based on a direct comparison of optimality criterion... I'm not agree xD.
Next, Pablo Goloboff showed that the common argument against weighting, that a character which is poor in clade is underweighted in a clade in which character has low homoplasy, is a problematic one, and that parsimony, as we know it, imply homogeneous weighting across the whole cladogram.
John Wenzel seems to be unsuccessful to show a different way to attack consensus trees. I think that agreement subtrees, the method that he defends is not as good as reduced consensus that can be found with TNT.
In an interesting talk, from philosophical, and statistical point of view, Chris Randle, showed that as actually implemented, Bayesian analysis in not bayesian, because the impossibility to implement a real definition of clade priors.
After coffee break, Steve Farris give an entertaining and clever talk about some misrepresentation of ideas of support by Grant & Kluge, and of course, re-affirms his masterful conclusion from his 1983 classic: parsimony is minimization of ad hoc hypotheses of homoplasy. I'm very happy to see the one that gives shape to actual numerical cladistics (and, I thinks, the major contributor of the theoretical development of phylogenetics in general!).
Then, a bunch of papers based on Pablo's implementation of continuous characters, using Farris' optimization, using Opiliones. But the most interesting contribution was from Santiago Catalano, who shows that landmark data can be viewed as a generalization of Sankoff's parsimony!
Afternoon starts with a presentation of the possibilities of EOL.org (Encyclopedia of Life) by Torstein Dikow, actually, apart of being as wonderful as Wikipedia, I do not see any application for EOL... (see Page's blog!)
Matthew Yoder, shows some wonderful ways to work using open source, in the development of his software for multi-author phylogenetic studies, with his sever-based Mx.
Then Rasmus Hovmoller, shows some interesting work to understand the spreading of avian influenza A, unfortunately, external problems was an obstacle to enjoy their results.
Federico López gives a talk about using conservation indexes to conservation in amazonia. I'm quite suspicious of that kind of indexes (although they are also bad, I think that Faith's PD is far better than Vane-Wright indexes!).
Norberto Giannini shows a new way to treat correlation of characters (“comparative method”) into a truly phylogenetic way. The method is excellent, and I think a real improvement in that field!
Fernando Noll, showed a beautiful work of behavioral data for Meliponini bees, that include oviposition and nest architecture.
Jeffrey Skevington use dragonflies from Fidji, and he tries to explain the origin of sexual bias on this beautiful insects. Juan Larrain shows his molecular analysis of a group of mosses, and compare his results with a preliminar set of morphological characters.
Martín Ramírez gives an excellent talk about the usefulness of ontologies for phylogenetic analysis! I feel that ontologies are an important step in the maintainability of morphological data (and their subsequent usage), but I think that although wonderful, the re-using, specially from authors extern to the original work, seems to be difficult (at least, as actually doing).
To finish the day, Johnatan Liria talks about k selection using some of my old TNT scripts xD...
Etiquetas:
databasing,
digital taxonomy,
hennig meeting,
molecular phylogenetics
miércoles, enero 30, 2008
Taxonomy: its time to change!
I just read a post by Christopher Taylor about a 'taxonomic problem' with Drosophila melanogaster. The curiosity of the case prompt me to made this post, that I have in my mind for a long time.
One of the main advantages of a classification is the information retrieval. Biological classification allows to find relationships and characters of a determinate specie. There is a set laws to manage the classification, know as the codes (they are different codes for animals, plants, bacteria, etc.).
In this age of intent and massive databases, one could thing that taxonomy are ready to made a direct jump, but unfortunately that seems not to be the case. The codes are really old ones, and many taxonomist are afraid of changing the rules would collapse the system.
I do not think that taxonomist would leave the data designers to propose a new system, but I think that a deep collaboration, coupled with some changes in the code structures will speed-up taxonomic research. The main objective of taxonomist is to describe and classify species, not to be lawyers!
I identify some problems that I think that be are the most important.
Linnean names
Linnean names are nice, they provide a some form of order in a time of great explorations around the world (XVIII-XIX century), when europeans became surprised with the diversity around the world. Many things are change from that time. By now we have several phylogenetic methods, which show how arbitrary a clade name would be (I like this post about the subject), and of course, the tree of life have far more divisions than Linnean ranks. Now we have databases to store names and search algorithms and find it quickly.
A rank-free taxonomy seems to be more adequate to store our information about the phylogeny of species. That is not an embrace of 'phylocode', as I prefer a taxonomy based on specific characters, or better a combination of characters and topology, instead of a topology based taxonomy (note that apomorphy based definition are also based on topology). Also, there is an special utility from ranked names: they can serve as Landmarks when browsing the tree of life.
The nested nature of biological classification allows a rank free taxonomy without the pains imagined by [1], and using “landmarks” helps in search--and writing--of abstracts, titles of papers, and of course, the search of a particular clade, as you can see using GenBank. Of course, genus and species can be (and I think would be) of obligatory nature. That allows a continuity with Linnean taxonomy.
If linnean names are more landmarks, then many of the laws for synonyms, and so one, need to be changed to a more practical usage.
Synonyms
Perhaps, the most annoying characteristic of taxonomy is the synonymy, rank changes, and several related problems. This problems are actually a burden of actual codes, and can be changed with changing the laws, without harming the actual classification.
Synonyms are really bad for databasesas the same entity is labeled with two different names, or (worst!) a same name is applied to different entities, in other cases, the range of the synonymy seems to be overlapping.
Of course, no matter what phylocoders say, the same problem applies to their 'phylogenetic' names: as new revision is published, the names continue but with completely different meaning. At least in traditional taxonomy, it is possible to reject some names.
It seems better to relax some of the naming rules, so a new classification would be clearly different from the original one. If a huge family is discovered to be massively paraphyletic, I think that it is no point allowing the original name to survive, it is simply an brutally wrong name.
For example, “reptiles” for a long time include many steam ammniotes, therapsids “mammal-like reptiles”, anapsids (as turtles), lepidosaurs (lizards and snakes) and crocodyles. They are amniotes but exclude birds and mammals. I can see a reason to retain that ugly name. It is simple, synonymize it with their monophyletic equivalent (Amniota), and never more use it. There is no way that a modern paper allows some confusion with the old ones. The only valid use of the name is when someone shows that the original reptiles, are monophyletic.
Names can be used instead to show different possible classifications. For example, the different arrangements of Arthropoda receive different names: Atelocerata (Myriapods, like centipede, and insects) vs. Pancrustacea (Crustaceans and insects), Mandibulata (Myriposd, crustaceans and insects) vs. schizoramia (crustaceans and arachnids). These names identify different entities, each one attached to a different phylogenetic proposal (note that form phylocoders, atelocerata, pancrustacea and mandibulata can be the same entity!).
Types
The reptile and arthropod example, are possible because there are no types fixing the names of that groups. At family levels, family names, or genus names are ruled by typification of names. Which allows more confusion, solutions to actual research.
There is an example: Lygaelidae was a large family of bugs (Insecta: Heteroptera), for long time it was believed that it was paraphyletic [2], but this was only demostrated by Henry [3]. He propose a new whole classification of Lygelids, elevating to family range no less than 7 subfamilies, and restructuring the meaning of Lygaelidae to only 3 subfamilies. Is the new Lygaelidae the same of, say 20 or 30 years ago? Of course not. Then typification was creating name stability (as is created by topology by phylocoders) but at cost of the loss of name utility.
As in the case of reptiles, there are no Lygaelidae any more. Any new reference i the litarature to 'Lygalelidae' only confuses with the initial meaning of the group (which is a synonym of the superfamly Lyageoidea). Of course you can use 'Lygaeidae sensu Henry' but it is only a clumsy (and error prone) way to give a new name.
Another nice example was provided by the previously mentioned post of Chris. It is about Drosophila. In this case, the usage can be against the taxonomic practice. For m, the solution is simple: no more Drosophila. But as there is a huge number of users of the name, that surely don't care about taxonomy, that outnumbered the number of Drosphilic taxonomic publications, there is when a commission can rule. The practical solution is to maintain Drosophila to the molecular people. Then solutions in cases of conflict would be guided by practical options, rather than some old described type.
This, of course, allows to made classification changes without 'using the types' (as far as the original characters were examined!). And free taxonomist to depend on some poorly known species (or even, specimen) to nominate genera and families.
Speaking about types...
Type specimens have a particular property: they are in the first world, but they are collected in the third world. The reason for this is historical, but their consequences are seen more acute today. Museums are measured by their amount of 'type specimens', and there are particular politics to borrow that specimens (only borrowing one on time, certifications, curators permission...). Also, some taxonomist, specially from the old past, simply nominate a wonderful amount of new species, only giving a superficial description (some color, some illustrations of genital parts) and based the whole 'description' on a type species designation. This old practice continued in several obscure papers in the third world.
Then typing, although seems to be a reasonable way to be objective, is more harming I guess that typing will never gone, but at the moment there are several nice perspectives to free from borrowing politics. Approaches like [4] with a great emphasis on characters and images, can change the situation. Researches far away from type specimenes can see high quality pics of several specimen parts. A side consequence is that the concept of type becomes lost, it is impossible to pic every part from a single specimen, and is possible that it ends destroyed, then the new typing would be more responsible, as it would be based in several different specimens.
Moreover a destypification increase general collections value, that is more the quantity and quality (e.g. fresh specimenes) of material available, than a particular specimen collected in 1816, saved from a fire in 1874, harmfully damaged by bad curation in 1903...
Actually there are many phylogenetic work without using type specimens, that is, the major bulk of molecular phylogenies, and I guess several morphological ones. I think that they do it in a very objective and testable way. If they can live without types, why classical taxonomist do not?
The data matrix
A second question from the previous section, is how a non-typified research can be objective? The answer is that instead on focusing on a particular specimen, phylogeneticists use a data matrix of taxon an characters.
A non type taxonomy enforce the use of well delimited characters, it is the only way to show the reality of the new designation. Look at some recent revisions with a phylogenetic analysis, and compare it with a revision without it (for example some of both see the pubs of AMNH). The character matrix allows to a quick examination of several characters, it is possible to see which state each character has in each taxon. New technologies (see [4]) couple specimen, characters and images for each cell entry. By default a matrix provide a multi-entry key, the identification tools are better.
There are some nice things of using a character matrix. The first, is an increasing interest in provide well defined characters [4]. As character are used for phylogenetic analysis, they would be stricter. Other characteristics like color patterns, length measurements would be restricted to a simpler description. Another advantage is that it provides a quick classification of a new species.
A nice real example was provided with the dinosaur paleontologist researches. The y publish some quick and small reports in high profile journals (like Nature or Science) with small descriptions, but as they have a great database of characters, several points of the anatomy of the new described fossil are immediately 'published', long before the detailed description in a more specialized journal.
Thinking on databasing
Of course, using a character matrix is direct consequence for storage: well defined characters and images enter smoothly in a database [4].
I think that the new challenges of the 'biodiversity crisis' as well as the 'taxonomic crisis' can be solved with a thinking of data storage. How can we store the data more efficiently? How can we link taxonomic and publication data? How changes in our knowledge about phylogeny could change the previous publication data, how the harm can be minimized?
It is important to a new taxonomy to keep the great advances made from Linneaus times, a start from the scratch is clearly a wrong solution. But also taxonomist would be able to made some concessions in their practice, and update it to new data architecture of the world.
It is time that taxonomy became a useful discussion about actual data, facing the massive extinction that the man is producing around the world, it seems weird that a taxonomic study would began searching for old papers from XVIII century, which only utility is that they provide a name, descriptions, characters and other things from that papers are of low value (by the way, taxonomy is the only field of science that continue using such old data. For historians old text are the source of investigation, for taxonomist is a more lawyer-like activity of searching for an 'old case'. Catalogs are some nice curiosities, and surely valuable for historians, but what is their actual value for taxonomist? They are important only because points to papers that establish a name).
If taxonomy is the main objective, then useful data storage is the main objective. Book keeping and law courts are not part of knowing biodiversity.
References
[1] Dominguez, E., Wheeler, Q. 1997. Taxonomic stability is ignorance. Cladistics 13: 367-372. doi: 10.1111/j.1096-0031.1997.tb00325.x
[2] Schuh, R. T., Slater, J. A. 1995. True Bugs of the World. Cornell Univ. New York
[3] Henry, T. J. 1997. Phylogenetic analysis of family groups within the infraorder Pentatomomorpha (Hemiptera: Heteroptera), with emphasis on the Lygaeoidea. Annals of the Entomological Society of America 90: 275-301
[4] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic Biology 56: 283-294. doi: 10.1080/10635150701313848
One of the main advantages of a classification is the information retrieval. Biological classification allows to find relationships and characters of a determinate specie. There is a set laws to manage the classification, know as the codes (they are different codes for animals, plants, bacteria, etc.).
In this age of intent and massive databases, one could thing that taxonomy are ready to made a direct jump, but unfortunately that seems not to be the case. The codes are really old ones, and many taxonomist are afraid of changing the rules would collapse the system.
I do not think that taxonomist would leave the data designers to propose a new system, but I think that a deep collaboration, coupled with some changes in the code structures will speed-up taxonomic research. The main objective of taxonomist is to describe and classify species, not to be lawyers!
I identify some problems that I think that be are the most important.
Linnean names
Linnean names are nice, they provide a some form of order in a time of great explorations around the world (XVIII-XIX century), when europeans became surprised with the diversity around the world. Many things are change from that time. By now we have several phylogenetic methods, which show how arbitrary a clade name would be (I like this post about the subject), and of course, the tree of life have far more divisions than Linnean ranks. Now we have databases to store names and search algorithms and find it quickly.
A rank-free taxonomy seems to be more adequate to store our information about the phylogeny of species. That is not an embrace of 'phylocode', as I prefer a taxonomy based on specific characters, or better a combination of characters and topology, instead of a topology based taxonomy (note that apomorphy based definition are also based on topology). Also, there is an special utility from ranked names: they can serve as Landmarks when browsing the tree of life.
The nested nature of biological classification allows a rank free taxonomy without the pains imagined by [1], and using “landmarks” helps in search--and writing--of abstracts, titles of papers, and of course, the search of a particular clade, as you can see using GenBank. Of course, genus and species can be (and I think would be) of obligatory nature. That allows a continuity with Linnean taxonomy.
If linnean names are more landmarks, then many of the laws for synonyms, and so one, need to be changed to a more practical usage.
Synonyms
Perhaps, the most annoying characteristic of taxonomy is the synonymy, rank changes, and several related problems. This problems are actually a burden of actual codes, and can be changed with changing the laws, without harming the actual classification.
Synonyms are really bad for databasesas the same entity is labeled with two different names, or (worst!) a same name is applied to different entities, in other cases, the range of the synonymy seems to be overlapping.
Of course, no matter what phylocoders say, the same problem applies to their 'phylogenetic' names: as new revision is published, the names continue but with completely different meaning. At least in traditional taxonomy, it is possible to reject some names.
It seems better to relax some of the naming rules, so a new classification would be clearly different from the original one. If a huge family is discovered to be massively paraphyletic, I think that it is no point allowing the original name to survive, it is simply an brutally wrong name.
For example, “reptiles” for a long time include many steam ammniotes, therapsids “mammal-like reptiles”, anapsids (as turtles), lepidosaurs (lizards and snakes) and crocodyles. They are amniotes but exclude birds and mammals. I can see a reason to retain that ugly name. It is simple, synonymize it with their monophyletic equivalent (Amniota), and never more use it. There is no way that a modern paper allows some confusion with the old ones. The only valid use of the name is when someone shows that the original reptiles, are monophyletic.
Names can be used instead to show different possible classifications. For example, the different arrangements of Arthropoda receive different names: Atelocerata (Myriapods, like centipede, and insects) vs. Pancrustacea (Crustaceans and insects), Mandibulata (Myriposd, crustaceans and insects) vs. schizoramia (crustaceans and arachnids). These names identify different entities, each one attached to a different phylogenetic proposal (note that form phylocoders, atelocerata, pancrustacea and mandibulata can be the same entity!).
Types
The reptile and arthropod example, are possible because there are no types fixing the names of that groups. At family levels, family names, or genus names are ruled by typification of names. Which allows more confusion, solutions to actual research.
There is an example: Lygaelidae was a large family of bugs (Insecta: Heteroptera), for long time it was believed that it was paraphyletic [2], but this was only demostrated by Henry [3]. He propose a new whole classification of Lygelids, elevating to family range no less than 7 subfamilies, and restructuring the meaning of Lygaelidae to only 3 subfamilies. Is the new Lygaelidae the same of, say 20 or 30 years ago? Of course not. Then typification was creating name stability (as is created by topology by phylocoders) but at cost of the loss of name utility.
As in the case of reptiles, there are no Lygaelidae any more. Any new reference i the litarature to 'Lygalelidae' only confuses with the initial meaning of the group (which is a synonym of the superfamly Lyageoidea). Of course you can use 'Lygaeidae sensu Henry' but it is only a clumsy (and error prone) way to give a new name.
Another nice example was provided by the previously mentioned post of Chris. It is about Drosophila. In this case, the usage can be against the taxonomic practice. For m, the solution is simple: no more Drosophila. But as there is a huge number of users of the name, that surely don't care about taxonomy, that outnumbered the number of Drosphilic taxonomic publications, there is when a commission can rule. The practical solution is to maintain Drosophila to the molecular people. Then solutions in cases of conflict would be guided by practical options, rather than some old described type.
This, of course, allows to made classification changes without 'using the types' (as far as the original characters were examined!). And free taxonomist to depend on some poorly known species (or even, specimen) to nominate genera and families.
Speaking about types...
Type specimens have a particular property: they are in the first world, but they are collected in the third world. The reason for this is historical, but their consequences are seen more acute today. Museums are measured by their amount of 'type specimens', and there are particular politics to borrow that specimens (only borrowing one on time, certifications, curators permission...). Also, some taxonomist, specially from the old past, simply nominate a wonderful amount of new species, only giving a superficial description (some color, some illustrations of genital parts) and based the whole 'description' on a type species designation. This old practice continued in several obscure papers in the third world.
Then typing, although seems to be a reasonable way to be objective, is more harming I guess that typing will never gone, but at the moment there are several nice perspectives to free from borrowing politics. Approaches like [4] with a great emphasis on characters and images, can change the situation. Researches far away from type specimenes can see high quality pics of several specimen parts. A side consequence is that the concept of type becomes lost, it is impossible to pic every part from a single specimen, and is possible that it ends destroyed, then the new typing would be more responsible, as it would be based in several different specimens.
Moreover a destypification increase general collections value, that is more the quantity and quality (e.g. fresh specimenes) of material available, than a particular specimen collected in 1816, saved from a fire in 1874, harmfully damaged by bad curation in 1903...
Actually there are many phylogenetic work without using type specimens, that is, the major bulk of molecular phylogenies, and I guess several morphological ones. I think that they do it in a very objective and testable way. If they can live without types, why classical taxonomist do not?
The data matrix
A second question from the previous section, is how a non-typified research can be objective? The answer is that instead on focusing on a particular specimen, phylogeneticists use a data matrix of taxon an characters.
A non type taxonomy enforce the use of well delimited characters, it is the only way to show the reality of the new designation. Look at some recent revisions with a phylogenetic analysis, and compare it with a revision without it (for example some of both see the pubs of AMNH). The character matrix allows to a quick examination of several characters, it is possible to see which state each character has in each taxon. New technologies (see [4]) couple specimen, characters and images for each cell entry. By default a matrix provide a multi-entry key, the identification tools are better.
There are some nice things of using a character matrix. The first, is an increasing interest in provide well defined characters [4]. As character are used for phylogenetic analysis, they would be stricter. Other characteristics like color patterns, length measurements would be restricted to a simpler description. Another advantage is that it provides a quick classification of a new species.
A nice real example was provided with the dinosaur paleontologist researches. The y publish some quick and small reports in high profile journals (like Nature or Science) with small descriptions, but as they have a great database of characters, several points of the anatomy of the new described fossil are immediately 'published', long before the detailed description in a more specialized journal.
Thinking on databasing
Of course, using a character matrix is direct consequence for storage: well defined characters and images enter smoothly in a database [4].
I think that the new challenges of the 'biodiversity crisis' as well as the 'taxonomic crisis' can be solved with a thinking of data storage. How can we store the data more efficiently? How can we link taxonomic and publication data? How changes in our knowledge about phylogeny could change the previous publication data, how the harm can be minimized?
It is important to a new taxonomy to keep the great advances made from Linneaus times, a start from the scratch is clearly a wrong solution. But also taxonomist would be able to made some concessions in their practice, and update it to new data architecture of the world.
It is time that taxonomy became a useful discussion about actual data, facing the massive extinction that the man is producing around the world, it seems weird that a taxonomic study would began searching for old papers from XVIII century, which only utility is that they provide a name, descriptions, characters and other things from that papers are of low value (by the way, taxonomy is the only field of science that continue using such old data. For historians old text are the source of investigation, for taxonomist is a more lawyer-like activity of searching for an 'old case'. Catalogs are some nice curiosities, and surely valuable for historians, but what is their actual value for taxonomist? They are important only because points to papers that establish a name).
If taxonomy is the main objective, then useful data storage is the main objective. Book keeping and law courts are not part of knowing biodiversity.
References
[1] Dominguez, E., Wheeler, Q. 1997. Taxonomic stability is ignorance. Cladistics 13: 367-372. doi: 10.1111/j.1096-0031.1997.tb00325.x
[2] Schuh, R. T., Slater, J. A. 1995. True Bugs of the World. Cornell Univ. New York
[3] Henry, T. J. 1997. Phylogenetic analysis of family groups within the infraorder Pentatomomorpha (Hemiptera: Heteroptera), with emphasis on the Lygaeoidea. Annals of the Entomological Society of America 90: 275-301
[4] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic Biology 56: 283-294. doi: 10.1080/10635150701313848
viernes, enero 18, 2008
DataTube? I can't wait!!
I just read a wonderful news at WiredScience, Google will be hosting open scientific data on the web [http://research.google.com]! WiredScience says that the interface will be similar to the one from YouTube, with annotations and comments.
I can't wait to see many morphological matrices, and morphological pics! I think that the excellent proposal of Ramírez et al. [1] can be coupled with that project :).
[1] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic biology 56: 283-294. doi: 10.1080/10635150701313848
I can't wait to see many morphological matrices, and morphological pics! I think that the excellent proposal of Ramírez et al. [1] can be coupled with that project :).
[1] Ramírez, M. J. et al. 2007. Linking of digital images to phylogenetic data matrices using a morphological ontology. Systematic biology 56: 283-294. doi: 10.1080/10635150701313848
jueves, diciembre 20, 2007
Relational databases for phylogenetic data
A long time ago, when I'm studying computer science, the only paradigm for databases were realtional databases, then under internet fast growth, more powerful computers, and popular searching engines--like Google--put text based databases on the top. Text based databases are the paradigm used to construct some databases for phylogenetic data, like GenBank, and TreeBASE.
Under text based databases the emphasis is put on somewhat standardized input files, that keep all the possible data necessary to make the search, you surely know the files retrieved from GenBank, or the matrices/trees from treeBASE, they have several fields, like classification, identification of sequence or character, a string matching algorithm is used to found exact and similar matches from the user query.
Then, the principal source of research in a text based database is to implement matching algorithms, the blog iPhylo from Rod Page, have several posts, paper and manuscripts links (here, here) about that subject. But I think that a text based databases are a wrong approach.
I remark some points:
Relational databases are a different concept from text-based databases. They are based on several independent tables connected to key-fields and in several cases using a secondary tables to match keys form different tables. The searching engine is based on specific keys (for complex searches) and specific fields form each table.
My own idea of the structure of database, in a sketch fashion, and based on my usual queries is like this:
Of course, a proper database design will need several years of development, with tons of interviews of several taxonomists to allow a product that could be used in a right way for several people around the world! (Although he had several post endorsing string matching databases, this post of Rod Page provide several nice ideas about the integrative work of a phylogenetic databases, of course there are several things that i don-t like xD).
Here I explain some part of the different tables, the table 'taxon' store the actual nomeclature of a taxon name, a species or a supra-specific entity, its name, author, it may include the diagnosis (linked to character entry!) a 'phylogenetic concept', a link to pictures, the type specimen (with a link to specimen!) and things like that. The table 'synonyms' and 'classification' are secondary tables to store only a relationship between two taxons: synonyms (it could include the motive of the synonymy), and the next inclusive taxon to a particular taxon (it might include the author of this inclusive relationship), as each taxon is independent you could include as many synonyms as you know, or as many classifications proposed.
The 'specimen' table could store specific information about the examined material, links to pictures of the specimen, the locality where the specimen was collected, and so on.
It is also a character section, the table 'character' store the name, description, and maybe bibliography and pictures of the character, with 'character equivalences' it is possible to store equivalent characters used in other studies, allowing cross reference between studies that used different taxon scopes, the 'character entry' table store the specific information for the character and the taxon, it could be a single cell in from a morphology matrix, or an specific sequence fragment.
This design could help in queries/actions in which text-based databases fail. The ambiguity is reduced, as each entry is a single one: the uploader would identify the particular nature of their entries in an specific way, text based databases could be developed with that standard in mind, but here it is possible to keep an strict species naming, as many synonyms as you want, and specific character names. A curator process for the taxonomy could be possible without harming the whole database data, and as phylogenetic results would be introduced in a form of classification, the gap between phylogeny and classification [1] would be reduced.
When you perform a search you could retrieve only the specific information that you want: for example all head characters from Hymenoptera, using the equivalences the characters could be more or less organized, and using the classification table you could retrieve head characters used for many different studies, alternative characters that could match under alternative classifications, or possible characters that could be present as they are scored for more inclusive classifications (for example, a character used to define Hexapoda and Arthropoda), a truly information retrieval system based on classification [2].
As is remarked by Nixon et al. [3] a single file format is not a good thing to a database, instead, it is preferable a table structure, and report tools that could produce entries in different formats, for example retrieving sequences in a GenBank format, in a TNT format, and a POY (fasta) format, or a distribution file ready to use in NDM.
The classification table could help to retrieve studies that support or reject a particular classification, you could found the evidence for one grouping and for the alternative classification. Particular algorithms need to be developed to translate the tree structure to a query--Page posted about the subject frequently ;)--. But it is more powerful that a text based search because the same database have the information needed (the hierarchic classification) to perform the search.
I hope sometime I became rich (I doubt it) or receive a grant--I hope ;)--to perform this huge task, or at least that someone around the net, had similar ideas. Until then the only path is continuous suffering with some 'taxonomic' and 'phylogenetic' databases around the net...
Pd. As you note, maybe a database like that put some burden on researcher that upload the data... but if you go to field trips for months, examine material for hours, day after day, write reports and manuscripts, uploading information it is part of the whole research!
[1] Franz, N.M. 2005. On the lack of good scientific reasons for the growing phylogeny/classification gap. Cladistics 21: 495-500.
[2] Farris, J.S. 1979. The information content of the phylogenetic system. Systematic zoology 28: 483-519.
[3] Nixon, K.C., Carpenter, J.M., Borgardt, S.J. 2001. Beyond NEXUS: universal cladistic data objects. Cladistics 17: S53-S59.
Under text based databases the emphasis is put on somewhat standardized input files, that keep all the possible data necessary to make the search, you surely know the files retrieved from GenBank, or the matrices/trees from treeBASE, they have several fields, like classification, identification of sequence or character, a string matching algorithm is used to found exact and similar matches from the user query.
Then, the principal source of research in a text based database is to implement matching algorithms, the blog iPhylo from Rod Page, have several posts, paper and manuscripts links (here, here) about that subject. But I think that a text based databases are a wrong approach.
I remark some points:
- Ambiguity: even if the entry for uploads is based on a 'cut and paste' template (or a wizard) the uploader is left with the responsibility to fill it adequately. As Page remarks, in many TreeBASE matrices terminal names are not properly scientific names.
- Irrelevant information: when you retrieve a file, it is usually full of information that you don't want, other information it is not provided. As searches are based on text matches, several notes, comments and other fields contain information, its function seems to be to confuse the searching algorithms (remember, they are string matching algorithms!).
- Formatting: As is text based, a compromise to made the files available to several programs made the entry fields to be rigid and difficult to modify and include new fields/information. For example in GenBank geographic information is not mandatory--as far as I know--even for phylogeographic datasets! Geographic and examined material for TreeBASE seems to be impossible to implement using the nexus format.
- Taxonomy: As a consequence of the rigid format, alternative taxonomies and synonyms searches can not be implemented, or require searches on alternative databases.
- Phylogeny: Why not try a 'tree structure' search on TreeBASE?
Relational databases are a different concept from text-based databases. They are based on several independent tables connected to key-fields and in several cases using a secondary tables to match keys form different tables. The searching engine is based on specific keys (for complex searches) and specific fields form each table.
My own idea of the structure of database, in a sketch fashion, and based on my usual queries is like this:

Of course, a proper database design will need several years of development, with tons of interviews of several taxonomists to allow a product that could be used in a right way for several people around the world! (Although he had several post endorsing string matching databases, this post of Rod Page provide several nice ideas about the integrative work of a phylogenetic databases, of course there are several things that i don-t like xD).
Here I explain some part of the different tables, the table 'taxon' store the actual nomeclature of a taxon name, a species or a supra-specific entity, its name, author, it may include the diagnosis (linked to character entry!) a 'phylogenetic concept', a link to pictures, the type specimen (with a link to specimen!) and things like that. The table 'synonyms' and 'classification' are secondary tables to store only a relationship between two taxons: synonyms (it could include the motive of the synonymy), and the next inclusive taxon to a particular taxon (it might include the author of this inclusive relationship), as each taxon is independent you could include as many synonyms as you know, or as many classifications proposed.
The 'specimen' table could store specific information about the examined material, links to pictures of the specimen, the locality where the specimen was collected, and so on.
It is also a character section, the table 'character' store the name, description, and maybe bibliography and pictures of the character, with 'character equivalences' it is possible to store equivalent characters used in other studies, allowing cross reference between studies that used different taxon scopes, the 'character entry' table store the specific information for the character and the taxon, it could be a single cell in from a morphology matrix, or an specific sequence fragment.
This design could help in queries/actions in which text-based databases fail. The ambiguity is reduced, as each entry is a single one: the uploader would identify the particular nature of their entries in an specific way, text based databases could be developed with that standard in mind, but here it is possible to keep an strict species naming, as many synonyms as you want, and specific character names. A curator process for the taxonomy could be possible without harming the whole database data, and as phylogenetic results would be introduced in a form of classification, the gap between phylogeny and classification [1] would be reduced.
When you perform a search you could retrieve only the specific information that you want: for example all head characters from Hymenoptera, using the equivalences the characters could be more or less organized, and using the classification table you could retrieve head characters used for many different studies, alternative characters that could match under alternative classifications, or possible characters that could be present as they are scored for more inclusive classifications (for example, a character used to define Hexapoda and Arthropoda), a truly information retrieval system based on classification [2].
As is remarked by Nixon et al. [3] a single file format is not a good thing to a database, instead, it is preferable a table structure, and report tools that could produce entries in different formats, for example retrieving sequences in a GenBank format, in a TNT format, and a POY (fasta) format, or a distribution file ready to use in NDM.
The classification table could help to retrieve studies that support or reject a particular classification, you could found the evidence for one grouping and for the alternative classification. Particular algorithms need to be developed to translate the tree structure to a query--Page posted about the subject frequently ;)--. But it is more powerful that a text based search because the same database have the information needed (the hierarchic classification) to perform the search.
I hope sometime I became rich (I doubt it) or receive a grant--I hope ;)--to perform this huge task, or at least that someone around the net, had similar ideas. Until then the only path is continuous suffering with some 'taxonomic' and 'phylogenetic' databases around the net...
Pd. As you note, maybe a database like that put some burden on researcher that upload the data... but if you go to field trips for months, examine material for hours, day after day, write reports and manuscripts, uploading information it is part of the whole research!
[1] Franz, N.M. 2005. On the lack of good scientific reasons for the growing phylogeny/classification gap. Cladistics 21: 495-500.
[2] Farris, J.S. 1979. The information content of the phylogenetic system. Systematic zoology 28: 483-519.
[3] Nixon, K.C., Carpenter, J.M., Borgardt, S.J. 2001. Beyond NEXUS: universal cladistic data objects. Cladistics 17: S53-S59.
Suscribirse a:
Entradas (Atom)
