ClassyFire: automated chemical classification with a comprehensive, computable taxonomy

There is no author summary for this article yet. Authors can add summaries to their articles on ScienceOpen to make them more accessible to a non-specialist audience.

Abstract

Background Scientists have long been driven by the desire to describe, organize, classify, and compare objects using taxonomies and/or ontologies. In contrast to biology, geology, and many other scientific disciplines, the world of chemistry still lacks a standardized chemical ontology or taxonomy. Several attempts at chemical classification have been made; but they have mostly been limited to either manual, or semi-automated proof-of-principle applications. This is regrettable as comprehensive chemical classification and description tools could not only improve our understanding of chemistry but also improve the linkage between chemistry and many other fields. For instance, the chemical classification of a compound could help predict its metabolic fate in humans, its druggability or potential hazards associated with it, among others. However, the sheer number (tens of millions of compounds) and complexity of chemical structures is such that any manual classification effort would prove to be near impossible. Results We have developed a comprehensive, flexible, and computable, purely structure-based chemical taxonomy (ChemOnt), along with a computer program (ClassyFire) that uses only chemical structures and structural features to automatically assign all known chemical compounds to a taxonomy consisting of >4800 different categories. This new chemical taxonomy consists of up to 11 different levels (Kingdom, SuperClass, Class, SubClass, etc.) with each of the categories defined by unambiguous, computable structural rules. Furthermore each category is named using a consensus-based nomenclature and described (in English) based on the characteristic common structural properties of the compounds it contains. The ClassyFire webserver is freely accessible at http://classyfire.wishartlab.com/. Moreover, a Ruby API version is available at https://bitbucket.org/wishartlab/classyfire_api, which provides programmatic access to the ClassyFire server and database. ClassyFire has been used to annotate over 77 million compounds and has already been integrated into other software packages to automatically generate textual descriptions for, and/or infer biological properties of over 100,000 compounds. Additional examples and applications are provided in this paper. Conclusion ClassyFire, in combination with ChemOnt (ClassyFire’s comprehensive chemical taxonomy), now allows chemists and cheminformaticians to perform large-scale, rapid and automated chemical classification. Moreover, a freely accessible API allows easy access to more than 77 million “ClassyFire” classified compounds. The results can be used to help annotate well studied, as well as lesser-known compounds. In addition, these chemical classifications can be used as input for data integration, and many other cheminformatics-related tasks. Electronic supplementary material The online version of this article (doi:10.1186/s13321-016-0174-y) contains supplementary material, which is available to authorized users.

Related collections

Most cited references 22

Record: found
Abstract: found
Article: not found

Gene Ontology: tool for the unification of biology

Michael Ashburner, Catherine A. Ball, Judith Blake … (2002)

Genomic sequencing has made it clear that a large fraction of the genes specifying the core biological functions are shared by all eukaryotes. Knowledge of the biological role of such shared proteins in one organism can often be transferred to other organisms. The goal of the Gene Ontology Consortium is to produce a dynamic, controlled vocabulary that can be applied to all eukaryotes even as knowledge of gene and protein roles in cells is accumulating and changing. To this end, three independent ontologies accessible on the World-Wide Web (http://www.geneontology.org) are being constructed: biological process, molecular function and cellular component.

0 comments Cited 15534 times – based on 0 reviews      Review now

Bookmark

Record: found
Abstract: not found
Article: not found

Description of several chemical structure file formats used by computer programs developed at Molecular Design Limited

Arthur Dalby, James G Nourse, W. Hounshell … (1992)

0 comments Cited 128 times – based on 0 reviews      Review now

Bookmark

Record: found
Abstract: found
Article: found

Is Open Access

T3DB: the toxic exposome database

David Wishart, David L. Arndt, Allison Pon … (2014)

The exposome is defined as the totality of all human environmental exposures from conception to death. It is often regarded as the complement to the genome, with the interaction between the exposome and the genome ultimately determining one's phenotype. The ‘toxic exposome’ is the complete collection of chronically or acutely toxic compounds to which humans can be exposed. Considerable interest in defining the toxic exposome has been spurred on by the realization that most human injuries, deaths and diseases are directly or indirectly caused by toxic substances found in the air, water, food, home or workplace. The Toxin-Toxin-Target Database (T3DB - www.t3db.ca) is a resource that was specifically designed to capture information about the toxic exposome. Originally released in 2010, the first version of T3DB contained data on nearly 2900 common toxic substances along with detailed information on their chemical properties, descriptions, targets, toxic effects, toxicity thresholds, sequences (for both targets and toxins), mechanisms and references. To more closely align itself with the needs of epidemiologists, toxicologists and exposome scientists, the latest release of T3DB has been substantially upgraded to include many more compounds (>3600), targets (>2000) and gene expression datasets (>15 000 genes). It now includes extensive data on ‘normal’ toxic compound concentrations in human biofluids as well as detailed chemical taxonomies, informative chemical ontologies and a large number of referential NMR, MS/MS and GC-MS spectra. This manuscript describes the most recent update to the T3DB, which was previously featured in the 2010 NAR Database Issue.

0 comments Cited 119 times – based on 0 reviews      Review now

Bookmark

All references

Author and article information

Journal

Title: Journal of Cheminformatics

Abbreviated Title: J Cheminform

Publisher: Springer Nature

ISSN (Electronic): 1758-2946

Publication date Created: December 2016

Publication date (Print): November 2016

Volume: 8

Issue: 1

Article

DOI: 10.1186/s13321-016-0174-y

SO-VID: 6c95432b-4640-4d0b-be01-dfbe90896bb5

History

Data availability:

Comments

Comment on this article

scite_

Cited by 349

See all cited by

Most referenced authors 1,558

See all reference authors

ClassyFire: automated chemical classification with a comprehensive, computable taxonomy

Read this article at

Abstract

Related collections

ScienceOpen Research

Most cited references 22

Gene Ontology: tool for the unification of biology

Description of several chemical structure file formats used by computer programs developed at Molecular Design Limited

T3DB: the toxic exposome database

Author and article information

Journal

Article

History

Comments

Comment on this article

Similar content 1,321

Cited by 349

Most referenced authors 1,558