VLDB 2026 Research / reviewers in the wild / expert
Cun-gen Cao 0001
dblp:06/584 · also Cungen Cao 0001
· DBLP profile ↗
38ranked-venue papers
6as first author
5since 2021 · last 2022
0009-0007-3250-1001ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 18 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 first-authorTheory of computation · 2Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | ECCKG: An Eventuality-Centric Commonsense Knowledge Graph
Cun-gen Cao 0001, Zhiwen Chen 0007, Shi Wang 0002 |
KSEM (1) | 2 |
| 2022 | CKGAC: A Commonsense Knowledge Graph About Attributes of Concepts
Cun-gen Cao 0001, Zhiwen Chen 0007, Shi Wang 0002 |
KSEM (1) | 2 |
| 2021 | SOM-NCSCM : An Efficient Neural Chinese Sentence Compression Model Enhanced with Self-Organizing MapabstractSentence Compression (SC), which aims to shorten sentences while retaining important words that express the essential meanings, has been studied for many years in many languages, especially in English.However, improvements on Chinese SC task are still quite few due to several difficulties: scarce of parallel corpora, different segmentation granularity of Chinese sentences, and imperfect performance of syntactic analyses.Furthermore, entire neural Chinese SC models have been under-investigated so far.In this work, we construct an SC dataset of Chinese colloquial sentences from a real-life question answering system in the telecommunication domain, and then, we propose a neural Chinese SC model enhanced with a Self-Organizing Map (SOM-NCSCM), to gain a valuable insight from the data and improve the performance of the whole neural Chinese SC model in a valid manner. 1 Experimental results show that our SOM-NCSCM can significantly benefit from the deep investigation of similarity among data, and achieve a promising F1 score of 89.655 and BLEU4 score of 70.116, which also provides a baseline for further research on the Chinese SC task. Kangli Zi, Shi Wang 0002, Yu Liu 0118, Jicun Li, Yanan Cao 0001, Cun-gen Cao 0001 |
EMNLP (1) | 6 |
| 2021 | Knowledge Enhanced Sequential Entity LinkingabstractEntity Linking (EL) is the task of mapping mentions in texts to the corresponding entities in knowledge bases. Existing studies mostly focus on joint disambiguation based on the topical coherence, including graph and sequence models. Sequence models alleviate the complexity caused by graph models, but exist the error propagation that incorrectly disambiguated entities are likely to induce further errors when predicting future mentions. Moreover, it is a huge expense to construct the relationship between entities to explore structured knowledge. To address these problems, we propose a novel method, Knowledge Enhanced Sequential Entity Linking (KESEL), which converts global EL into a sequence decision problem and applies a pre-trained language model to better fuse entity knowledge. Specifically, we firstly utilize multiple features to learn local contextual representations of mentions and candidates respectively. Next, a sequential ERNIE model is introduced to generate knowledgeable representations by dynamically integrating the knowledge of previously referred entities into subsequent mentions disambiguation. Finally, by concatenating the above learned contextual and knowledgeable representations, we make full use of multi-semantic information to improve the performance of EL. Extensive experiments show that our method can achieve competitive or state-of-the-art results. Yu Liu 0118, Shi Wang 0002, Kangli Zi, Jicun Li, Cun-gen Cao 0001 |
IJCNN | 5 |
| 2021 | A Property-Based Method for Acquiring Commonsense Knowledge
Cun-gen Cao 0001, Yuting Cao, Shi Wang 0002 |
KSEM | 2 |
| 2020 | HAPE: A programmable big knowledge graph platform
Ruqian Lu, Chaoqun Fei, Chuanqing Wang, Shunfeng Gao, Han Qiu 0001, Songmao Zhang, Cun-gen Cao 0001 |
Inf. Sci. | 7 |
| 2019 | Answer-Focused and Position-Aware Neural Network for Transfer Learning in Question Generation
Kangli Zi, Xingwu Sun, Yanan Cao 0001, Shi Wang 0002, Xiaoming Feng, Zhaobo Ma, Cun-gen Cao 0001 |
KSEM (2) | 7 |
| 2019 | Reasoning and querying web-scale open data based on DL-LiteA in a divide-and-conquer way
Zhenzhen Gu, Songmao Zhang, Cun-gen Cao 0001 |
J. Web Semant. | 3 |
| 2016 | Knowledge Extraction from Chinese Records of Cyber Attacks Based on a Semantic Grammar
Fang Fang 0009, Luchen Zhang, Cun-gen Cao 0001 |
KSEM | 4 |
| 2016 | A Practical Method of Identifying Chinese Metaphor Phrases from Corpus
Jianhui Fu, Shi Wang 0002, Cun-gen Cao 0001 |
KSEM | 4 |
| 2016 | Extracting Knowledge from Web Tables Based on DOM Tree Similarity
Cun-gen Cao 0001, Jianhui Fu, Shi Wang 0002 |
KSEM | 2 |
| 2016 | The M-computations induced by accessibility relations in nonstandard models M of Hoare logic
Cun-gen Cao 0001, Yuefei Sui, Zaiyue Zhang |
Frontiers Comput. Sci. | 1 |
| 2016 | A Seed-Based Method for Generating Chinese Confusion SetsabstractIn natural language, people often misuse a word (called a “confused word”) in place of other words (called “confusing words”). In misspelling corrections, many approaches to finding and correcting misspelling errors are based on a simple notion called a “confusion set.” The confusion set of a confused word consists of confusing words. In this article, we propose a new method of building Chinese character confusion sets. Our method is composed of two major phases. In the first phase, we build a list of seed confusion sets for each Chinese character, which is based on measuring similarity in character pinyin or similarity in character shape. In this phase, all confusion sets are constructed manually, and the confusion sets are organized into a graph, called a “seed confusion graph” (SCG), in which vertices denote characters and edges are pairs of characters in the form (confused character, confusing character). In the second phase, we extend the SCG by acquiring more pairs of (confused character, confusing character) from a large Chinese corpus. For this, we use several word patterns (or patterns) to generate new confusion pairs and then verify the pairs before adding them into a SCG. Comprehensive experiments show that our method of extending confusion sets is effective. Also, we shall use the confusion sets in Chinese misspelling corrections to show the utility of our method. Cun-gen Cao 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2015 | Two-Phased Event Causality Acquisition: Coupling the Boundary Identification and Argument Identification ApproachesabstractEvent causality is indispensable for knowledge-driven intelligent systems. In this paper, we propose a supervised method of extracting event causalities such as forest is cut down $$\rightarrow $$ forest is destroyed from web text. While relation identification using lexico-syntactic patterns (LSPs) is not novel, it is still challenging to extract the event expressions with necessary arguments from identified causality mentions. To address this issue, our method divides event-pair extraction into two phases: event boundary identification and missing argument identification. In the first phase, we propose a Naive Baysian probability method to identify the boundary of causal events, and extract the corresponding text fragments as event expressions. Secondly, we learn a multi-class decision tree (LADTree) to identify the missing argument for each incomplete event. Experimental results showed the good effectiveness of our approach on a large-scale open corpus. Yanan Cao 0001, Cun-gen Cao 0001, Jingzun Zhang, Wenjia Niu |
KSEM | 2 |
| 2015 | Tree Based Shape Similarity Measurement for Chinese CharactersabstractIn Chinese, there are many characters which are similar in shape, and this phenomenon usually induces writing errors. As one important issue in spelling automatic correction, shape similarity measurement is still a challenging problem. To address this issue, we propose a component-tree based method in this paper, which is based on the hypothesis “characters are similar if their construction and components are both similar”. Firstly, we decompose each character to a tree recursively, in which the root node is the character and the leaf nodes are atomic parts, called strokes. Then, we align any pair of trees using their minimal super-tree and calculate their similarity from bottom to up based on weighted edit distance. Finally, the cognitive prominence is used to adjust the similarity scores. In text proofreading experiments, our method achieved 97% precision and 95.6% recall, which can be applied in practical systems. Yanan Cao 0001, Shi Wang 0002, Cun-gen Cao 0001 |
KSEM | 3 |
| 2015 | The Double-Level Default Description Logic D 3 LabstractWe propose the default description logic $$\mathcal {D}2\mathcal {L}$$ and the double-level default description logic $$\mathcal {D}3\mathcal {L}$$ . $$\mathcal {D}2\mathcal {L}$$ embeds normal defaults inside the basic description logic $$\mathcal {ALC},$$ and $$\mathcal {D}3\mathcal {L}$$ augments $$\mathcal {D}2\mathcal {L}$$ with normal double-level defaults. Double-level defaults are defaults of defaults and can be used to represent default inheritance of default properties of concepts in ontologies. A $$\mathcal {D}3\mathcal {L}$$ knowledge base ( $$\mathcal {D}3\mathcal {L}$$ -KB) can be divided into two levels of knowledge bases, and correspondingly its extensions can be computed in two steps. $$\mathcal {D}3\mathcal {L}$$ is more expressive than $$\mathcal {D}2\mathcal {L}$$ since there is a $$\mathcal {D}3\mathcal {L}$$ -KB that cannot reduce to any $$\mathcal {D}2\mathcal {L}$$ -KB. Specifically, there is a $$\mathcal {D}3\mathcal {L}$$ -KB such that the set of all its extensions cannot be exactly generated by any $$\mathcal {D}2\mathcal {L}$$ -KB. Liangjun Zang, Weimin Wang 0002, Cun-gen Cao 0001 |
KSEM | 4 |
| 2015 | A Chinese Framework of Semantic Taxonomy and Description: Preliminary Experimental Evaluation Using Web Information ExtractionabstractThe Chinese Framework of Semantic Taxonomy and Description (FSTD) is a linguistic resource that stores lexical and predicate-argument semantics about events or states in Chinese text, developed with the application of knowledge acquisition from Chinese text in mind. In this paper we build a web information extraction system, called NkiExtractor, to evaluate FSTD experimentally. We use two metrics: grammar coverage measures whether there is a semantic category of FSTD that corresponds to an event description in text, and extraction precision measures whether the correct predicate-argument structure can be extracted from text. Experimental results show that FSTD is a fairly comprehensive and effective resource for knowledge acquisition. We also discuss future work for expanding FSTD and improving extraction precision of NkiExtractor. Liangjun Zang, Weimin Wang 0002, Fang Fang 0009, Cong Cao 0001, Cun-gen Cao 0001 |
KSEM | 10 |
| 2014 | A Practical Approach to Extracting Names of Geographical Entities and Their Relations from the Web
Cun-gen Cao 0001, Shi Wang 0002 |
KSEM | 1 |
| 2013 | Relational Operations and Uncertainty Measure in Rough Relational DatabaseabstractThe traditional relational database model (RDM) is not effective for dealing with imprecise and uncertain data as it deals with precise and unambiguous data. Hence, Beaubouef et al. proposed the rough relational database model (RRDM) for the management of uncertainty in relational databases. Beaubouef et al. defined the corresponding rough relational operators in rough relational databases as in ordinary relational databases. And to give an effective measure of uncertainty in rough relational databases, they defined the rough relation entropy. In this paper, we further discuss the issues of relational operations and uncertainty measure in rough relational databases. We give some new definitions for rough relational operators and rough relation entropy in rough relational databases. Furthermore, we discuss the basic properties of rough relational operators and rough relation entropy, as well as the connections between rough relational operators and rough relation entropy. Feng Jiang 0019, Xiaoyan Wan, Yuefei Sui, Cun-gen Cao 0001, Junwei Du |
Fundam. Informaticae | 4 |
| 2013 | A Survey of Commonsense Knowledge Acquisition
Liangjun Zang, Cong Cao 0001, Yanan Cao 0001, Yuming Wu, Cun-gen Cao 0001 |
J. Comput. Sci. Technol. | 5 |
| 2011 | A Chinese time ontology for the Semantic Web
Chunxia Zhang 0001, Cun-gen Cao 0001, Yuefei Sui, Xindong Wu 0001 |
Knowl. Based Syst. | 2 |
| 2011 | A hybrid approach to outlier detection based on boundary region
Feng Jiang 0019, Yuefei Sui, Cun-gen Cao 0001 |
Pattern Recognit. Lett. | 3 |
| 2008 | Using Sense Recognition to Resolve the Problem of Polysemy in Building a Taxonomic Hierarchy
Lei Liu 0039, Lu Hong Diao, Shuying Yan, Cun-gen Cao 0001 |
KES (1) | 5 |
| 2007 | Learning Concepts from Text Based on the Inner-Constructive Model
Shi Wang 0002, Yanan Cao 0001, Cun-gen Cao 0001 |
KSEM | 4 |
| 2007 | A Chinese Time Ontology
Chunxia Zhang 0001, Cun-gen Cao 0001, Yuefei Sui, Zhendong Niu |
KSEM | 2 |
| 2007 | A Google-Based Statistical Acquisition Model of Chinese Lexical Concepts
Shi Wang 0002, Cun-gen Cao 0001 |
KSEM | 3 |
| 2006 | NKIMathE - A Multi-purpose Knowledge Management Environment for Mathematical Concepts
Qingtian Zeng, Cun-gen Cao 0001, Hua Duan, Yongquan Liang 0001 |
KSEM | 2 |
| 2006 | A Tree Construction of the Preferable Answer Sets for Prioritized Basic Disjunctive Logic Programs
Zaiyue Zhang, Yuefei Sui, Cun-gen Cao 0001 |
TAMC | 3 |
| 2005 | A Survey of Computational Emotion Research
Donglei Zhang, Cun-gen Cao 0001, Xi Yong, Haitao Wang 0009, Yu Pan 0004 |
IVA | 2 |
| 2005 | Context-Restricted, Role-Oriented Emotion Knowledge Acquisition and Representation
Xi Yong, Cun-gen Cao 0001, Haitao Wang 0009 |
KES (2) | 2 |
| 2004 | Ontology-Based Web Agents Using Concept Description Flow
Nengfu Xie, Cun-gen Cao 0001, Bingxian Ma, Chunxia Zhang 0001, Jinxin Si |
KES | 2 |
| 2004 | Knowledge modeling and acquisition of traditional Chinese herbal drugs and formulae from text
Cun-gen Cao 0001, Haitao Wang 0009, Yuefei Sui |
Artif. Intell. Medicine | 1 |
| 2004 | Domain-Specific Ontology of Botany
Fang Gu, Cun-gen Cao 0001, Yuefei Sui |
J. Comput. Sci. Technol. | 2 |
| 2004 | Domain-Specific Formal Ontology of Archaeology and Its Application in Knowledge Acquisition and Analysis
Chunxia Zhang 0001, Cun-gen Cao 0001, Fang Gu, Jinxin Si |
J. Comput. Sci. Technol. | 2 |
| 2002 | Extracting and Sharing Knowledge from Medical Texts
Cun-gen Cao 0001 |
J. Comput. Sci. Technol. | 1 |
| 2002 | Progress in the Development of National Knowledge Infrastructure
Cun-gen Cao 0001, Qiangze Feng, Fang Gu, Jinxin Si, Yuefei Sui, Haitao Wang 0009, Qingtian Zeng, Chunxia Zhang 0001, Yufei Zheng, Xiaobin Zhou |
J. Comput. Sci. Technol. | 1 |
| 2002 | Designing a Top-Level Ontology of Human Beings: A Multi-Perspective Approach
Fang Gu, Cun-gen Cao 0001 |
J. Comput. Sci. Technol. | 3 |
| 2001 | Medical Knowledge Acquisition from the Electronic Encyclopedia of China
Cun-gen Cao 0001 |
AIME | 1 |