2015
Ambiguity of non-systematic chemical identifiers within and between small-molecule databases
Abstract: BackgroundA wide range of chemical compound databases are currently available for pharmaceutical research. To retrieve compound information, including structures, researchers can query these chemical databases using non-systematic identifiers. These are source-dependent identifiers (e.g., brand names, generic names), which are usually assigned to the compound at the point of registration. The correctness of non-systematic identifiers (i.e., whether an identifier matches the associated structure) can only be as…
Search citation statements
Paper Sections
Select...
18
5
2
0
Citation Types
6
15
0
0
Year Published
2016
2026
Publication Types
Select...
15
4
2
1
Relationship
2
20
Authors
Journals
Cited by 22 publications
(21 citation statements)
References 43 publications
6
15
0
0
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Within the same database, the percentage of identifiers with multiplicity larger than one varies from 0% to 100%, whereas the ambiguity of names ranges from 0.07% to 29.4%. When mapping between databases, these ambiguities and multiplicities lead to larger inconsistencies, and this agree with previous observations regarding small molecules databases [ 31 , 35 ]. The inconsistencies when mapping using metabolite names range from 0% to 81.2%.…”
Section: Discussion
supporting
confidence: 90%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Within the same database, the percentage of identifiers with multiplicity larger than one varies from 0% to 100%, whereas the ambiguity of names ranges from 0.07% to 29.4%. When mapping between databases, these ambiguities and multiplicities lead to larger inconsistencies, and this agree with previous observations regarding small molecules databases [ 31 , 35 ]. The inconsistencies when mapping using metabolite names range from 0% to 81.2%.…”
Section: Discussion
supporting
confidence: 90%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Finally, when desalted formulas were compared, which further eliminated discrepancies that could be attributed to salt-parent, stereo, and geometric isomer differences, 3% of the mismatches remained, indicating a significant remaining level of gross structural errors. These results are consistent with a recently published report quantifying the ambiguity of chemical identifiers in a selection of public databases, 35 but to our knowledge, the high degree of inconsistency of chemical supplier-provided information relative to manual curation results had not previously been reported. (Note: These results were based on an analysis of 90% of the current ToxCast sample inventory, as of March, 2014, using canonicalized SMILES strings for both sets of parent and desalted structures generated from the original SD files in ACD/ChemFolder, v2012, Toronto, Canada.…”
Section: Building the Toxcast Chemical Library
supporting
confidence: 92%
Abstract
Smart CitationsHow this paper cites the one you are viewing
“…Within the same database, the percentage of identifiers with multiplicity larger than one varies from 0 % to 100 %, whereas the ambiguity of names ranges from 0.07 % to 29.4 %. When mapping between databases, these ambiguities and multiplicities lead to larger inconsistencies, and this agree with previous observations regarding small molecules databases [262,266]. The inconsistencies when mapping using metabolite names range from 0 % to 81.2 %.…”
Section: Discussion
supporting
confidence: 88%
“…In agreement with earlier studies [358,357], our observations in Chapter 2 also suggested that mapping without manual curation implies high risk of mismatch. Yet this is not the most efficient approach [357]. Detailed study has shown that manual curation has resolved only a small part of the inconsistency in ChemSpider [357].…”
Section: The Lack Of Standards In Gems
supporting
confidence: 92%
