Shi Wang 0002

dblp:55/2449-2 · DBLP profile ↗
← Back
14ranked-venue papers in the field
1as first author
5since 2021 · last 2025
0000-0002-1329-2415ORCID · conflict

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 10 (1 first)Data Mining & Knowledge Discovery · 2Information Retrieval & Web Search · 2
YearPublicationVenuePosition
2025 Bridging the Gap: Aligning Language Model Generation with Structured Information Extraction via Controllable State Transition
abstract
Large language models (LLMs) achieve superior performance in generative tasks. However, due to the natural gap between language model generation and structured information extraction in three dimensions: task type, output format, and modeling granularity, they often fall short in structured information extraction, a crucial capability for effective data utilization on the web. In this paper, we define the generation process of the language model as the controllable state transition, aligning the generation and extraction processes to ensure the integrity of the output structure and adapt to the goals of the information extraction task. Furthermore, we propose the Structure2Text decider to help the language model understand the fine-grained extraction information, which converts the structured output into natural language and makes state decisions, thereby focusing on the task-specific information kernels, and alleviating language model hallucinations and incorrect content generation. We conduct extensive experiments and detailed analyses on myriad information extraction tasks, including named entity recognition, relation extraction, and event argument extraction. Our method not only achieves significant performance improvements but also considerably enhances the model's capability to generate precise and relevant content, making the extracted content easy to parse.
Hao Li 0156, Yubing Ren, Yanan Cao 0001, Fang Fang 0009, Zheng Lin 0001, Shi Wang 0002
WWW7
2024 Generative Models for Complex Logical Reasoning over Knowledge Graphs
abstract
Answering complex logical queries over knowledge graphs (KGs) is a fundamental yet challenging task. Recently, query representation has been a mainstream approach to complex logical reasoning, making the target answer and query closer in the embedding space. However, there are still two limitations. First, prior methods model the query as a fixed vector, but ignore the uncertainty of relations on KGs. In fact, different relations may contain different semantic distributions. Second, traditional representation frameworks fail to capture the joint distribution of queries and answers, which can be learned by generative models that have the potential to produce more coherent answers. To alleviate these limitations, we propose a novel generative model, named DiffCLR, which exploits the diffusion model for complex logical reasoning to approximate query distributions. Specifically, we first devise a query transformation to convert logical queries into input sequences by dynamically constructing contextual subgraphs. Then, we integrate them into the diffusion model to execute a multi-step generative process, and a structure-enhanced self-attention is further designed for incorporating the structural features embodied in KGs. Experimental results on two benchmark datasets show our model effectively outperforms state-of-the-art methods, particularly in multi-hop chain queries with significant improvement.
Yu Liu 0118, Yanan Cao 0001, Shi Wang 0002, Qingyue Wang, Guanqun Bi
WSDM3
2022 ECCKG: An Eventuality-Centric Commonsense Knowledge Graph
Cun-gen Cao 0001, Zhiwen Chen 0007, Shi Wang 0002
KSEM (1)4
2022 CKGAC: A Commonsense Knowledge Graph About Attributes of Concepts
Cun-gen Cao 0001, Zhiwen Chen 0007, Shi Wang 0002
KSEM (1)4
2021 A Property-Based Method for Acquiring Commonsense Knowledge
Cun-gen Cao 0001, Yuting Cao, Shi Wang 0002
KSEM4
2020 HIN: Hierarchical Inference Network for Document-Level Relation Extraction
Hengzhu Tang, Yanan Cao 0001, Zhenyu Zhang 0006, Jiangxia Cao, Fang Fang 0009, Shi Wang 0002, Pengfei Yin
PAKDD (1)6
2020 High Quality Candidate Generation and Sequential Graph Attention Network for Entity Linking
abstract
Entity Linking (EL) is a task for mapping mentions in text to corresponding entities in knowledge base (KB). This task usually includes candidate generation (CG) and entity disambiguation (ED) stages. Recent EL systems based on neural network models have achieved good performance, but they still face two challenges: (i) Previous studies evaluate their models without considering the differences between candidate entities. In fact, the quality (gold recall in particular) of candidate sets has an effect on the EL results. So, how to promote the quality of candidates needs more attention. (ii) In order to utilize the topical coherence among the referred entities, many graph and sequence models are proposed for collective ED. However, graph-based models treat all candidate entities equally which may introduce much noise information. On the contrary, sequence models can only observe previous referred entities, ignoring the relevance between the current mention and its subsequent entities. To address the first problem, we propose a multi-strategy based CG method to generate high recall candidate sets. For the second problem, we design a Sequential Graph Attention Network (SeqGAT) which combines the advantages of graph and sequence methods. In our model, mentions are dealt with in a sequence manner. Given the current mention, SeqGAT dynamically encodes both its previous referred entities and subsequent ones, and assign different importance to these entities. In this way, it not only makes full use of the topical consistency, but also reduce noise interference. We conduct experiments on different types of datasets and compare our method with previous EL system on the open evaluation platform. The comparison results show that our model achieves significant improvements over the state-of-the-art methods.
Zheng Fang 0002, Yanan Cao 0001, Zhenyu Zhang 0006, Yanbing Liu 0007, Shi Wang 0002
WWW6
2019 Answer-Focused and Position-Aware Neural Network for Transfer Learning in Question Generation
Kangli Zi, Xingwu Sun, Yanan Cao 0001, Shi Wang 0002, Xiaoming Feng, Zhaobo Ma, Cun-gen Cao 0001
KSEM (2)4
2016 A Practical Method of Identifying Chinese Metaphor Phrases from Corpus
Jianhui Fu, Shi Wang 0002, Cun-gen Cao 0001
KSEM2
2016 Extracting Knowledge from Web Tables Based on DOM Tree Similarity
Cun-gen Cao 0001, Jianhui Fu, Shi Wang 0002
KSEM5
2015 Tree Based Shape Similarity Measurement for Chinese Characters
abstract
In Chinese, there are many characters which are similar in shape, and this phenomenon usually induces writing errors. As one important issue in spelling automatic correction, shape similarity measurement is still a challenging problem. To address this issue, we propose a component-tree based method in this paper, which is based on the hypothesis “characters are similar if their construction and components are both similar”. Firstly, we decompose each character to a tree recursively, in which the root node is the character and the leaf nodes are atomic parts, called strokes. Then, we align any pair of trees using their minimal super-tree and calculate their similarity from bottom to up based on weighted edit distance. Finally, the cognitive prominence is used to adjust the similarity scores. In text proofreading experiments, our method achieved 97% precision and 95.6% recall, which can be applied in practical systems.
Yanan Cao 0001, Shi Wang 0002, Cun-gen Cao 0001
KSEM2
2014 A Practical Approach to Extracting Names of Geographical Entities and Their Relations from the Web
Cun-gen Cao 0001, Shi Wang 0002
KSEM2
2007 Learning Concepts from Text Based on the Inner-Constructive Model
Shi Wang 0002, Yanan Cao 0001, Cun-gen Cao 0001
KSEM1
2007 A Google-Based Statistical Acquisition Model of Chinese Lexical Concepts
Shi Wang 0002, Cun-gen Cao 0001
KSEM2