Denghao Ma

dblp:166/0023 · DBLP profile ↗
← Back
9ranked-venue papers in the field
6as first author
7since 2021 · last 2025
0000-0001-7597-9354ORCID · corroborated

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 7 (4 first)Information Retrieval & Web Search · 2 (2 first)
YearPublicationVenuePosition
2025 Sentence Extraction Framework with High Relevance and Divergence for Document Summarization
Huiwen Xue, Baoan Li, Denghao Ma, Xueqiang Lv, Xiaoxi Wang
DASFAA (1)3
2024 Reusing Keywords for Fine-grained Representations and Matchings
Li Chong, Denghao Ma, Yueguo Chen, Xueqiang Lv
DASFAA (2)2
2024 Domain-specific Answer Sentence Selection with Terminology Augmentation and Cascade Attention
Xiaotong Lyu, Denghao Ma, Yueguo Chen
DASFAA (5)2
2023 A Principled Decomposition of Pointwise Mutual Information for Intention Template Discovery
abstract
With the rise of Artificial Intelligence (AI), question answering systems have become common for users to interact with computers, e.g., ChatGPT and Siri. These systems require a substantial amount of labeled data to train their models. However, the labeled data is scarce and challenging to be constructed. The construction process typically involves two stages: discovering potential sample candidates and manually labeling these candidates. To discover high-quality candidate samples, we study the intention paraphrase template discovery task: Given some seed questions or templates of an intention, discover new paraphrase templates that describe the intention and are diverse to the seeds enough in text. As the first exploration of the task, we identify the new quality requirements, i.e., relevance, divergence and popularity, and identify the new challenges, i.e., the paradox of divergent yet relevant paraphrases, and the conflict of popular yet relevant paraphrases. To untangle the paradox of divergent yet relevant paraphrases, in which the traditional bag of words falls short, we develop usage-centric modeling, which represents a question/template/answer as a bag of usages that users engaged (e.g., up-votes), and uses a usage-flow graph to interrelate templates, questions and answers. To balance the conflict of popular yet relevant paraphrases, we propose a new and principled decomposition for the well-known Pointwise Mutual Information from the usage perspective (usage-PMI), and then develop a Bayesian inference framework over the usage-flow graph to estimate the usage-PMI. Extensive experiments over three large CQA corpora show strong performance advantage over the baselines adopted from paraphrase identification task. We release 885,000 paraphrase templates of high quality discovered by our proposed PMI decomposition model, and the data is available in site https://github.com/Para-Questions/Intention\_template\_discovery.
Denghao Ma, Kevin Chen-Chuan Chang, Yueguo Chen, Xueqiang Lv
CIKM1
2023 Category-Highlighting Transformer Network for Question Retrieval
Denghao Ma, Li Chong, Yueguo Chen
DASFAA (3)1
2023 Two-stage Interest Calibration Network for Reranking Hotels
Denghao Ma, Jiajia Sun, Yueguo Chen, Genliang Yi
DASFAA (4)1
2022 Definition-Augmented Jointly Training Framework for Intention Phrase Mining
Denghao Ma, Yueguo Chen, Changyu Wang, Hongbin Pei, Yitao Zhai, Gang Zheng 0006
DASFAA (3)1
2018 Interpreting Fine-Grained Categories from Natural Language Queries of Entity Search
Denghao Ma, Yueguo Chen, Xiaoyong Du 0001, Yuanzhe Hao
DASFAA (1)1
2018 Leveraging Fine-Grained Wikipedia Categories for Entity Search
abstract
Ad-hoc entity search, which is to retrieve a ranked list of relevant entities in response to a query of natural language question, has been widely studied. It has been shown that category matching of entities, especially when matching to fine-grained entity types/categories, is critical to the performance of entity search. However, the potentials of the fine-grained Wikipedia entity categories, has not been well exploited by existing studies. Based on the observation of how people describe entities of a specific type, we propose a headword-and-modifier model to deeply interpret both queries and fine-grained entity types/categories. Probabilistic generative models are designed to effectively estimate the relevance of headwords and modifiers as a pattern-based matching problem, taking the Wikipedia type taxonomy as an important input to address the ad-hoc representations of concepts/entities in queries. Extensive experimental results on three widely-used test sets: INEX-XER 2009, SemSearch-LS and TREC-Entity, show that our method achieves a significant improvement of the entity search performance over the state-of-the-art methods.
Denghao Ma, Yueguo Chen, Kevin Chen-Chuan Chang, Xiaoyong Du 0001, Chuanfei Xu, Yi Chang 0001
WWW1