EDBT 2026 Demo / reviewers in the wild / expert
Dezhao Song
dblp:59/8700
· DBLP profile ↗
17ranked-venue papers
10as first author
1since 2021 · last 2022
0000-0002-2553-3108ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 7 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
4 papers |
Knowledge graphs · 50% Information retrieval · 36% Data integration and cleaning · 15% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge graphs › knowledge graph construction
enterprise knowledge graph |
0.4 | 1 | 2019 | Building and Querying an Enterprise Knowledge Graph · IEEE Trans. Serv. Comput. 2019 |
Knowledge graphs
knowledge graph construction |
0.4 | 1 | 2019 | Building and Querying an Enterprise Knowledge Graph · IEEE Trans. Serv. Comput. 2019 |
Knowledge graphs
knowledge graph querying |
0.4 | 1 | 2019 | Building and Querying an Enterprise Knowledge Graph · IEEE Trans. Serv. Comput. 2019 |
Information retrieval
candidate selection |
0.3 | 1 | 2017 | Linking Heterogeneous Data in the Semantic Web Using Scalable and Domain-Independent Candidate Selection · IEEE Trans. Knowl. Data Eng. 2017 |
Data integration and cleaning
entity matching |
0.3 | 1 | 2017 | Linking Heterogeneous Data in the Semantic Web Using Scalable and Domain-Independent Candidate Selection · IEEE Trans. Knowl. Data Eng. 2017 |
Knowledge graphs › semantic web
semantic web data |
0.3 | 1 | 2017 | Linking Heterogeneous Data in the Semantic Web Using Scalable and Domain-Independent Candidate Selection · IEEE Trans. Knowl. Data Eng. 2017 |
Information retrieval › document retrieval › domain-specific retrieval
financial information retrieval |
0.2 | 1 | 2016 | Interacting with Financial Data using Natural Language · SIGIR 2016 |
Data integration and cleaning › entity resolution
instance coreference resolution |
0.2 | 1 | 2016 | Ontology Instance Linking: Towards Interlinked Knowledge Graphs · AAAI 2016 |
Information retrieval › query formulation
natural language querying |
0.2 | 1 | 2016 | Interacting with Financial Data using Natural Language · SIGIR 2016 |
Information retrieval
question answering |
0.2 | 1 | 2016 | Interacting with Financial Data using Natural Language · SIGIR 2016 |
Information retrieval
search interfaces |
0.2 | 1 | 2016 | Interacting with Financial Data using Natural Language · SIGIR 2016 |
Knowledge graphs
entity linking |
0.1 | 1 | 2019 | Building and Querying an Enterprise Knowledge Graph · IEEE Trans. Serv. Comput. 2019 |
Methods — techniques the papers use, named apart from their topics
natural language translation to queries · 0.8RDF modeling · 0.8unsupervised learning · 0.3histsim · 0.3DisNGram · 0.3natural language generation · 0.2named entity recognition · 0.2entity coreference resolution · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Multi-label legal document classification: A deep learning-based approach with label-attention and domain-specific pre-training
Dezhao Song, Andrew Vold, Kanika Madan, Frank Schilder |
Inf. Syst. | 1 |
| 2019 | Building and Querying an Enterprise Knowledge GraphabstractInformation providers are faced with a critical challenge to process, retrieve and present information to their users in order to satisfy their complex information needs, because data has been increasing in an unprecedented manner, coming from diverse sources, and covering a variety of domains in heterogeneous formats. In this paper, we present Thomson Reuters' effort in developing a family of services for building and querying an enterprise knowledge graph in order to address this challenge. We first acquire data from various sources via different approaches. Furthermore, we mine useful information from the data by adopting a variety of techniques, including Named Entity Recognition and Relation Extraction; such mined information is further integrated with existing structured data (e.g., via Entity Linking techniques) in order to obtain relatively comprehensive descriptions of the entities. By modeling the data as an RDF graph model, we enable easy data management and the embedding of rich semantics in our data. Finally, in order to facilitate the querying of this mined and integrated data, i.e., the knowledge graph, we propose TR Discover, a natural language interface that allows users to ask questions of our knowledge graph in their own words; such natural language questions are then translated into executable queries for answer retrieval. We evaluate our services, i.e., named entity recognition, relation extraction, entity linking and natural language interface, on real-world datasets, and demonstrate and discuss their practicability and limitations. Dezhao Song, Frank Schilder, Shai Hertz, Giuseppe Saltini, Charese Smiley, Phani Nivarthi, Oren Hazai, Dudi Landau, Mike Zaharkin, Tom Zielund, Hugo Molina-Salgado, Chris Brew, Dan Bennett |
IEEE Trans. Serv. Comput. | 1 |
| 2018 | The E2E NLG Challenge: A Tale of Two SystemsabstractThis paper presents the two systems we entered into the 2017 E2E NLG Challenge: TemplGen, a templated-based system and SeqGen, a neural network-based system.Through the automatic evaluation, SeqGen achieved competitive results compared to the template-based approach and to other participating systems as well.In addition to the automatic evaluation, in this paper we present and discuss the human evaluation results of our two systems. Charese Smiley, Elnaz Davoodi, Dezhao Song, Frank Schilder |
INLG | 3 |
| 2017 | Linking Heterogeneous Data in the Semantic Web Using Scalable and Domain-Independent Candidate SelectionabstractDue to the decentralized nature of the Semantic Web, the same real-world entity may be described in various data sources with different ontologies and assigned syntactically distinct identifiers. In order to facilitate data utilization and consumption in the Semantic Web, without compromising the freedom of people to publish their data, one critical problem is to appropriately interlink such heterogeneous data. This interlinking process is sometimes referred to as Entity Matching, i.e., finding which identifiers refer to the same real-world entity. In this paper, we propose two candidate selection algorithms to improve the scalability of entity matching systems. First of all, we propose HistSim that utilizes the matching histories of the instances to prune instance pairs that are not sufficiently similar to the same pool of other instances. A sigmoid function based thresholding method is proposed to automatically adjust the threshold for such commonality on-the-fly. Furthermore, we propose DisNGram that selects candidate instance pairs by computing a character-level similarity metric on discriminating literal values that are chosen using domain-independent unsupervised learning. Instances are indexed on the chosen predicates' literal values to enable efficient look-up for similar instances. Finally, in order to be able to handle heterogeneous datasets with a large number of predicates, a mechanism for automatically determining predicate comparability is proposed. We evaluate our two candidate selection algorithms against six state-of-the-art systems on three Semantic Web datasets, and demonstrate that our proposed algorithms frequently outperform state-of-the-art systems on F1-score and runtime. Dezhao Song, Jeff Heflin |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Ontology Instance Linking: Towards Interlinked Knowledge GraphsabstractDue to the decentralized nature of the Semantic Web, the same real-world entity may be described in various data sources with different ontologies and assigned syntactically distinct identifiers. In order to facilitate data utilization and consumption in the Semantic Web, without compromising the freedom of people to publish their data, one critical problem is to appropriately interlink such heterogeneous data. This interlinking process is sometimes referred to as Entity Coreference, i.e., finding which identifiers refer to the same real-world entity. In this paper, we first summarize state-of-the-art algorithms in detecting such coreference relationships between ontology instances. We then discuss various techniques in scaling entity coreference to large-scale datasets. Finally, we present well-adopted evaluation datasets and metrics, and compare the performance of the state-of-the-art algorithms on such datasets. Jeff Heflin, Dezhao Song |
AAAI | 2 |
| 2016 | When to Plummet and When to Soar: Corpus Based Verb Selection for Natural Language GenerationabstractFor data-to-text tasks in Natural Language Generation (NLG), researchers are often faced with choices about the right words to express phenomena seen in the data.One common phenomenon centers around the description of trends between two data points and selecting the appropriate verb to express both the direction and intensity of movement.Our research shows that rather than simply selecting the same verbs again and again, variation and naturalness can be achieved by quantifying writers' patterns of usage around verbs. Charese Smiley, Vassilis Plachouras, Frank Schilder, Hiroko Bretz, Jochen L. Leidner, Dezhao Song |
INLG | 6 |
| 2016 | Interacting with Financial Data using Natural LanguageabstractFinancial and economic data are typically available in the form of tables and comprise mostly of monetary amounts, numeric and other domain-specific fields. They can be very hard to search and they are often made available out of context, or in forms which cannot be integrated with systems where text is required, such as voice-enabled devices. This work presents a novel system that enables both experts in the finance domain and non-expert users to search financial data with both keyword and natural language queries. Our system answers the queries with an automatically generated textual description using Natural Language Generation (NLG). The answers are further enriched with derived information, not explicitly asked in the user query, to provide the context of the answer. The system is designed to be flexible in order to accommodate new use cases without significant development effort, thus allowing fast integration of new datasets. Vassilis Plachouras, Charese Smiley, Hiroko Bretz, Ola Taylor, Jochen L. Leidner, Dezhao Song, Frank Schilder |
SIGIR | 6 |
| 2015 | Natural Language Question Answering and Analytics for Diverse and Interlinked DatasetsabstractPrevious systems for natural language questions over complex linked datasets require the user to enter a complete and well-formed question, and present the answers as raw lists of entities. Using a feature-based grammar with a full formal semantics, we have developed a system that is able to support rich autosuggest, and to deliver dynamically generated analytics for each result that it returns. Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew |
HLT-NAACL | 1 |
| 2015 | TR Discover: A Natural Language Interface for Querying and Analyzing Interlinked Datasets
Dezhao Song, Frank Schilder, Charese Smiley, Chris Brew, Tom Zielund, Hiroko Bretz, Robert Martin, Chris Dale, John Duprey, Johanna Harrison |
ISWC (2) | 1 |
| 2015 | Multimodal Entity Coreference for Cervical Dysplasia DiagnosisabstractCervical cancer is the second most common type of cancer for women. Existing screening programs for cervical cancer, such as Pap Smear, suffer from low sensitivity. Thus, many patients who are ill are not detected in the screening process. Using images of the cervix as an aid in cervical cancer screening has the potential to greatly improve sensitivity, and can be especially useful in resource-poor regions of the world. In this paper, we develop a data-driven computer algorithm for interpreting cervical images based on color and texture. We are able to obtain 74% sensitivity and 90% specificity when differentiating high-grade cervical lesions from low-grade lesions and normal tissue. On the same dataset, using Pap tests alone yields a sensitivity of 37% and specificity of 96%, and using HPV test alone gives a 57% sensitivity and 93% specificity. Furthermore, we develop a comprehensive algorithmic framework based on Multimodal Entity Coreference for combining various tests to perform disease classification and diagnosis. When integrating multiple tests, we adopt information gain and gradient-based approaches for learning the relative weights of different tests. In our evaluation, we present a novel algorithm that integrates cervical images, Pap, HPV, and patient age, which yields 83.21% sensitivity and 94.79% specificity, a statistically significant improvement over using any single source of information alone. Dezhao Song, Sharon X. Huang, Joseph Patruno, Hector Muñoz-Avila, Jeff Heflin, L. Rodney Long, Sameer K. Antani |
IEEE Trans. Medical Imaging | 1 |
| 2014 | Exploring Linked Data with contextual tag clouds
Xingjian Zhang 0006, Dezhao Song, Sambhawa Priya, Zachary A. Daniels, Kelly Reynolds, Jeff Heflin |
J. Web Semant. | 2 |
| 2013 | Infrastructure for Efficient Exploration of Large Scale Linked Data via Contextual Tag Clouds
Xingjian Zhang 0006, Dezhao Song, Sambhawa Priya, Jeff Heflin |
ISWC (1) | 2 |
| 2013 | Semantator: Semantic annotator for converting biomedical text to linked data
Cui Tao, Dezhao Song, Deepak K. Sharma, Christopher G. Chute |
J. Biomed. Informatics | 2 |
| 2012 | Scalable and Domain-Independent Entity Coreference: Establishing High Quality Data Linkages across Heterogeneous Data Sources
Dezhao Song |
ISWC (2) | 1 |
| 2012 | Accuracy vs. Speed: Scalable Entity Coreference on the Semantic Web with On-the-Fly PruningabstractOne challenge for the Semantic Web is to scalably establish high quality owl: same As links between co referent ontology instances in different data sources, traditional approaches that exhaustively compare every pair of instances do not scale well to large datasets. In this paper, we propose a pruning-based algorithm for reducing the complexity of entity co reference. First, we discard candidate pairs of instances that are not sufficiently similar to the same pool of other instances. A sigmoid function based thresholding method is proposed to automatically adjust the threshold for such commonality on-the-fly. In our prior work, each instance is associated with a context graph consisting of neighboring RDF nodes. In this paper, we speed up the comparison for a single pair of instances by pruning insignificant context in the graph, this is accomplished by evaluating its potential contribution to the final similarity measure. We evaluate our system on three Semantic Web instance categories. We verify the effectiveness of our thresholding and context pruning methods by comparing to nine state-of-the-art systems. We show that our algorithm frequently outperforms those systems with a runtime speedup factor of 18 to 24 while maintaining competitive F1-scores. For datasets of up to 1 million instances, this translates to as much as 370 hours improvement in runtime. Dezhao Song, Jeff Heflin |
Web Intelligence | 1 |
| 2011 | Automatically Generating Data Linkages Using a Domain-Independent Candidate Selection Approach
Dezhao Song, Jeff Heflin |
ISWC (1) | 1 |
| 2010 | Domain-independent entity coreference in RDF graphsabstractIn this paper, we present a novel entity coreference algorithm for Semantic Web instances. The key issues include how to locate context information and how to utilize the context appropriately. To collect context information, we select a neighborhood (consisting of triples) of each instance from the RDF graph. To determine the similarity between two instances, our algorithm computes the similarity between comparable property values in the neighborhood graphs. The similarity of distinct URIs and blank nodes is computed by comparing their outgoing links. To provide the best possible domain-independent matches, we examine an appropriate way to compute the discriminability of triples. To reduce the impact of distant nodes, we explore a distance-based discounting approach. We evaluated our algorithm using different instance categories in two datasets. Our experiments show that the best results are achieved by including both our triple discrimination and discounting approaches. Dezhao Song, Jeff Heflin |
CIKM | 1 |