EDBT 2026 Demo / reviewers in the wild / expert
Karl T. Mueller
dblp:99/7389
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2011
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 80% Query processing and optimization · 20% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
indexing |
0.1 | 1 | 2011 | Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011 |
Information retrieval › indexing › index compression
index pruning |
0.1 | 1 | 2011 | Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011 |
Query processing and optimization › selection queries
partial match query |
0.1 | 1 | 2011 | Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011 |
Information retrieval
query processing |
0.1 | 1 | 2011 | Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011 |
Information retrieval
search engines |
0.1 | 1 | 2011 | Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital Documents · ACM Trans. Inf. Syst. 2011 |
Methods — techniques the papers use, named apart from their topics
support vector machine · 0.1conditional random field · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2011 | Identifying, Indexing, and Ranking Chemical Formulae and Chemical Names in Digital DocumentsabstractEnd-users utilize chemical search engines to search for chemical formulae and chemical names. Chemical search engines identify and index chemical formulae and chemical names appearing in text documents to support efficient search and retrieval in the future. Identifying chemical formulae and chemical names in text automatically has been a hard problem that has met with varying degrees of success in the past. We propose algorithms for chemical formula and chemical name tagging using Conditional Random Fields (CRFs) and Support Vector Machines (SVMs) that achieve higher accuracy than existing (published) methods. After chemical entities have been identified in text documents, they must be indexed. In order to support user-provided search queries that require a partial match between the chemical name segment used as a keyword or a partial chemical formula, all possible (or a significant number of) subformulae of formulae that appear in any document and all possible subterms (e.g., “methyl”) of chemical names (e.g., “methylethyl ketone”) must be indexed. Indexing all possible subformulae and subterms results in an exponential increase in the storage and memory requirements as well as the time taken to process the indices. We propose techniques to prune the indices significantly without reducing the quality of the returned results significantly. Finally, we propose multiple query semantics to allow users to pose different types of partial search queries for chemical entities. We demonstrate empirically that our search engines improve the relevance of the returned results for search queries involving chemical entities. Bingjun Sun, Prasenjit Mitra 0001, C. Lee Giles, Karl T. Mueller |
ACM Trans. Inf. Syst. | 4 |