close
2015
DOI: 10.1186/s13321-015-0102-6
|Get access via publisher |Summarize |Cite
|
Sign up to set email alerts

Ambiguity of non-systematic chemical identifiers within and between small-molecule databases

Abstract: BackgroundA wide range of chemical compound databases are currently available for pharmaceutical research. To retrieve compound information, including structures, researchers can query these chemical databases using non-systematic identifiers. These are source-dependent identifiers (e.g., brand names, generic names), which are usually assigned to the compound at the point of registration. The correctness of non-systematic identifiers (i.e., whether an identifier matches the associated structure) can only be as… Show more

Search citation statements

Order By: Relevance

Paper Sections

Select...
18
5
2
0

Citation Types

6
15
0
0

Year Published

2016
2016
2026
2026

Publication Types

Select...
15
4
2
1

Relationship

2
20

Authors

Journals

citations

Cited by 22 publications

(21 citation statements)
references

References 43 publications

6
15
0
0
Order By: Relevance
How this paper cites the one you are viewing
“…Within the same database, the percentage of identifiers with multiplicity larger than one varies from 0% to 100%, whereas the ambiguity of names ranges from 0.07% to 29.4%. When mapping between databases, these ambiguities and multiplicities lead to larger inconsistencies, and this agree with previous observations regarding small molecules databases [ 31 , 35 ]. The inconsistencies when mapping using metabolite names range from 0% to 81.2%.…”
Section: Discussion
supporting
confidence: 90%
How this paper cites the one you are viewing
“…Within the same database, the percentage of identifiers with multiplicity larger than one varies from 0% to 100%, whereas the ambiguity of names ranges from 0.07% to 29.4%. When mapping between databases, these ambiguities and multiplicities lead to larger inconsistencies, and this agree with previous observations regarding small molecules databases [ 31 , 35 ]. The inconsistencies when mapping using metabolite names range from 0% to 81.2%.…”
Section: Discussion
supporting
confidence: 90%
How this paper cites the one you are viewing
“…Finally, when desalted formulas were compared, which further eliminated discrepancies that could be attributed to salt-parent, stereo, and geometric isomer differences, 3% of the mismatches remained, indicating a significant remaining level of gross structural errors. These results are consistent with a recently published report quantifying the ambiguity of chemical identifiers in a selection of public databases, 35 but to our knowledge, the high degree of inconsistency of chemical supplier-provided information relative to manual curation results had not previously been reported. (Note: These results were based on an analysis of 90% of the current ToxCast sample inventory, as of March, 2014, using canonicalized SMILES strings for both sets of parent and desalted structures generated from the original SD files in ACD/ChemFolder, v2012, Toronto, Canada.…”
Section: Building the Toxcast Chemical Library
supporting
confidence: 92%
How this paper cites the one you are viewing
“…Within the same database, the percentage of identifiers with multiplicity larger than one varies from 0 % to 100 %, whereas the ambiguity of names ranges from 0.07 % to 29.4 %. When mapping between databases, these ambiguities and multiplicities lead to larger inconsistencies, and this agree with previous observations regarding small molecules databases [262,266]. The inconsistencies when mapping using metabolite names range from 0 % to 81.2 %.…”
Section: Discussion
supporting
confidence: 88%
“…In agreement with earlier studies [358,357], our observations in Chapter 2 also suggested that mapping without manual curation implies high risk of mismatch. Yet this is not the most efficient approach [357]. Detailed study has shown that manual curation has resolved only a small part of the inconsistency in ChemSpider [357].…”
Section: The Lack Of Standards In Gems
supporting
confidence: 92%
“…In order to improve the correctness of mapping, we need to get back to the root of the problem and answer the question about the reasons for the inconsistent mapping. Although previous studies have suggested that internal ambiguity is low and solving this will only solve part of the inconsistent problem [357], our findings in Chapter 4 agree with earlier analysis on public databases [358] that the problem is 144 rooted in the internal inconsistency of each database. As shown in Chapter 4, the majority of identifiers link to multiple names in the same database.…”
Section: The Lack Of Standards In Gems
supporting
confidence: 90%
See 1 more Smart Citation