VLDB 2026 Research / reviewers in the wild / expert
Liangjun Zang
dblp:65/1380
· DBLP profile ↗
30ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0003-3483-521XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-View Graph Learning with Dynamic Evidential Fusion for Response ForecastingabstractThe dissemination of misinformation on social media is likely to lead to severe social conflicts, therefore, the prediction of individual responses to news events becomes a significant task of social importance. The existing belief-centered method overlooks the fact that people tend to respond similarly to certain types of news. To tackle this issue, we propose a Multi-View framework with Dynamic Evidential Fusion(MVDEF) for Response Forecasting. To carry out the task, we first utilize Large Language Models to extract news topics and user beliefs from the news content and user profiles. Then, we construct three single-view graphs: a user-news interaction graph, a belief-aware graph, and a topic-aware graph. We dynamically evaluate the reliability of each view for different samples and then integrate the results based on their corresponding uncertainty mass. Additionally, we introduce a pseudo-view to enhance the interaction between these three views. Extensive experiments demonstrate that our model achieves excellent performance on real-world Twitter data. Further analysis reveals the model’s capability in unseen user scenarios, underscoring its practical applicability in real-world problems. Liangjun Zang, Zhaoxing Li, Songlin Hu 0001 |
IJCNN | 4 |
| 2024 | MT-MNA: Multiple Network Alignment with Absent Priori AnnotationabstractNetwork alignment serves as a key methodology in data analysis, enabling a understanding of intricate structures within multiple networks. The discovery of potential aligned pairs between networks is a prerequisite for a wide spectrum of applications, e.g. cross-domain recommendation or re-identification attack. Presently, most existing network alignment methods involve only two networks, remaining challenges in scenarios with multiple networks. In addition, acquiring cross-network links often poses a significant challenge, resulting in numerous instances where priori annotation information is absent. Our research confronts these problems by proposing Multi-hop Transformation Multiple Network Alignment with Absent Priori Annotation (MT-MNA). The core idea is to simultaneously align multiple networks and address the scenarios of absent priori annotation by fully utilizing the rich-annotated network pairs. Through rigorous experiments on real-world datasets, our method demonstrates superior precision and adaptability to the challenging condition of absent priori annotation in multiple network alignment, overperforming existing state-of-the-art methods. Haichao Fu, Liangjun Zang, Baojie Tian, Songlin Hu 0001 |
CSCWD | 2 |
| 2024 | Adaptive Mixture of Domain-aware Experts for Detecting Social BotsabstractSocial bot detection has received widely attention from academic and industrial communities. However, existing bot detection methods are far from perfect. To mimic genuine users on social networks, advanced social bots are often active in multiple domains and have mixed characteristics of multiple domains. It is unreasonable to classify a bot with only one domain. To effectively extract and fuse features from multiple domains, we propose a novel method for Domain-aware Social Bot Detection (DSBD). Specifically, we first use a prompt-based method for zero-shot domain classification to obtain accurate domain distribution for any user. We then aggregate multiple domain expert representations through a domain gate, and finally use the fused representation to classify. Experimental results show that our approach consistently outperforms all baselines and that our fusion strategy perform well in various settings especially zero-shot situation. Qianqian Lu, Shilong Li 0003, Wei Zhou 0019, Liangjun Zang |
CSCWD | 5 |
| 2024 | Explainable Deepfake Detection with Human PromptsabstractFacial manipulation techniques pose a significant threat to society due to the prevalence of deepfake content on the internet. While previous efforts have focused on developing accurate deepfake detection models, these models may be limited in real-world scenarios due to the lack of confidence that human analysts have in their results. Therefore, this study presents a novel approach to improve the practicality of deepfake detection models by incorporating human understanding. We propose a human prompt based deepfake detection framework that overlays Human-enhanced artifacts attention onto image artifact attention, which utilizes vision prompts to improve the model’s responsiveness and feedback ability while preserving its precision and generalizability. The deepfake detection model achieves an AUC score of 0.99 on the FaceForensics++ dataset and exhibits graceful generalization when evaluated on the Celeb-DF dataset. Furthermore, the model generates "possible area of manipulation" that provides an intuitive signal to facilitate interpretation of the detection process, bridging the gap between machine and human perception of "fake". Our proposed approach can potentially mitigate the harm caused by deepfakes and provide a more reliable solution for real-world applications. Xiaorong Ma, Zhaoxing Li, Yesheng Chai, Liangjun Zang, Jizhong Han |
CSCWD | 5 |
| 2024 | Towards More Effective and Transferable Poisoning Attacks against Link Prediction on GraphsabstractWith the impressive performances achieved by graph representation learning models on tasks such as link prediction, their vulnerability to imperceptible adversarial perturbations has also come to light. However, adversarial attacks struggle to balance effectiveness and transferability, particularly when attempting to deliver effective attacks on various target models in one shot and achieving effective outcomes in both availability and integrity attack settings. This study explores a novel way to mitigate attack performance across various models under availability and integrity attack settings. To fulfill these objectives, we develop a Scoring & Update (SU) framework for performing adversarial attacks against link prediction on graphs. Specifically, we iteratively score perturbations through a parameter-frozen and transferable surrogate model and then update the perturbation with scoring feedback to learn a more effective and transferable adversarial perturbation. Extensive experiments on two real-world datasets show that our attack model is more effective than six attack baselines against six popular target models for graph representation under both availability and integrity attack settings. Code is available at https://github.com/anonymousaccept/STAA. Baojie Tian, Liangjun Zang, Haichao Fu, Jizhong Han, Songlin Hu 0001 |
CSCWD | 2 |
| 2024 | Capture Long-Range Dependency with Meta-Path Transformer for De-Anonymization of Q&A SitesabstractThe expeditious advancement of social question-and-answer (Q&A) platforms has led to the valuable yet challenging practice of anonymous knowledge sharing. Despite implementing various anonymity techniques, the persistent threat of potential privacy breaches remains a paramount concern. To tackle this issue, we introduce the task of de-anonymization within Q&A communities and provide a bilingual dataset (Chinese and English) for research. In this paper, we propose a novel de-anonymization framework called MPT, effectively improving the model’s ability to capture long-range dependencies between nodes by integrating graph neural networks(GNNs) and language models(LMs). Specifically, we use GNN to extract structural features, and then we encode and fuse node representations from multiple meta-paths using Transformer and attention mechanisms. Extensive experiments on Zhihu and Quora data sets show that our model significantly outperforms the baseline model. In addition, our model possesses a degree of interpretability, enabling a comprehensive comprehension of the underlying factors contributing to user privacy breaches and facilitating the implementation of appropriate safeguards. The dataset1and code2utilized in this study have been made publicly accessible. Baojie Tian, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
CSCWD | 2 |
| 2024 | Triple-Based Data Augmentation for Event Temporal Extraction via Reinforcement Learning from Human FeedbackabstractEvent temporal relation (TempRel) extraction is a vital branch of information extraction. However, it is expensive to construct human-labeled training datasets, which leads to the lack of high-quality training data and inadequate training of the model. To mitigate this problem, we introduce a novel RLHF-based TempRel Data Augmentation method (RTDA) which aims at generating training data with high quality in terms of diversity and semantics fluency. Experimental results on three popular datasets TBD, TDD-Man, and TDD-Auto show that the model equipped with our method achieves state-of-the-art performance. Liangjun Zang, Shuchong Wei, Songlin Hu 0001 |
CSCWD | 2 |
| 2024 | Leveraging Evolution Patterns to Enhance Script Event Prediction by Large Language Models
Shuchong Wei, Liangjun Zang, Songlin Hu 0001 |
DASFAA (5) | 2 |
| 2024 | DBPrompt: A Database Anomaly Operation Detection and Analysis via Prompt Learning
Huazhen Zhong, Xuejian Wang, Wenjie Xiao, Xuehai Tang, Liangjun Zang |
ICIC (8) | 7 |
| 2024 | HIDD: Human-perception-centric Incremental Deepfake DetectionabstractFacial manipulation techniques pose a significant societal threat due to the widespread dissemination of deepfake content on the internet. Existing efforts for deepfake detection exhibit inadequate generalization performance when encountering unseen or degraded samples. We attribute this limitation to the overfitting of minor forgery patterns and variations in data distribution among disparate datasets. To tackle this issue, we introduce an innovative human-perception-centric incremental deepfake detection framework to enhance the generalization capabilities of deepfake detection models through continuous learning from a limited set of new samples. Firstly, the model leverages human perceptual salience to discern and comprehend significant artifacts, thereby mitigating overfitting to minor features. Subsequently, in the incremental learning process, we utilize multi-perspective knowledge distillation and a replay strategy to maintain the performance of the old model and minimize the feature distance between old and new samples. This comprehensive approach mitigates feature-level overfitting and addresses distribution differences among various datasets in the incremental phase. We conducted thorough experiments on four benchmark datasets (FF++, DFDC-P, CDF2, and DFD), and the experimental results demonstrate the superior performance of our method. Xiaorong Ma, Yesheng Chai, Zhaoxing Li, Jiao Dai, Liangjun Zang, Jizhong Han |
ICME | 7 |
| 2024 | HDDA: Human-perception-centric Deepfake Detection AdapterabstractFacial manipulation techniques pose a significant societal threat due to the prevalent presence of deepfake content online. Current deepfake detection methods demonstrate subpar generalization performance when applied to unseen samples. The cause of this limitation lies in the overfitting of minor forgery patterns and variations in data distribution across different datasets. To tackle this issue, we introduce an innovative Human-perception-centric Deepfake Detection Adapter, namely HDDA, to enhance the generalization ability of deepfake detection models. This adaptation primarily involves two stages. During the pre-training stage, the model utilizes human perception salience to spot significant artifacts, thus reducing overfitting to minor features. In the subsequent fine-tuning stage, we introduce an efficient parameter tuning module named Deepfake Detection Adapter. The Adapter introduces two types of lightweight yet specialized adapter modules to the pre-trained model while keeping the backbone network frozen. It fine-tunes the pre-trained model through the adapter to adapt new and unseen datasets, thereby enhancing generalization. We conducted comprehensive experiments on various standard deepfake detection benchmarks to validate the effectiveness of our approach, particularly in showcasing a compelling advantage under cross-dataset and cross-manipulation settings. Xiaorong Ma, Yesheng Chai, Jiao Dai, Zhaoxing Li, Liangjun Zang, Jizhong Han |
IJCNN | 6 |
| 2024 | Event Temporal Relation Extraction based on Retrieval-Augmented on LLMsabstractEvent temporal relation (TempRel) is a primary subject of the event relation extraction task. However, the inherent ambiguity of TempRel increases the difficulty of the task. With the rise of prompt engineering, it is important to design effective prompt templates and verbalizers to extract relevant knowledge. The traditional manually designed templates struggle to extract precise temporal knowledge. This paper introduces a novel retrieval-augmented TempRel extraction approach, leveraging knowledge retrieved from large language models (LLMs) to enhance prompt templates and verbalizers. Our method capitalizes on the diverse capabilities of various LLMs to generate a wide array of ideas for template and verbalizer design. Our proposed method fully exploits the potential of LLMs for generation tasks and contributes more knowledge to our design. Empirical evaluations across three widely recognized datasets demonstrate the efficacy of our method in improving the performance of event temporal relation extraction tasks. Liangjun Zang, Shuchong Wei, Songlin Hu 0001 |
IJCNN | 2 |
| 2023 | Improving Event Representation with Supervision from Available Semantic Resources
Shuchong Wei, Liangjun Zang, Songlin Hu 0001 |
DASFAA (3) | 2 |
| 2023 | UDAD: An Accurate Unsupervised Database Anomaly Detection MethodabstractDatabase systems are widely employed to store crucial data across domains. However, an increasing emergence of stealthy abnormal database access behaviors, such as re-identification and differential attacks, has been observed. These behaviors exhibit short durations and similarities to normal actions, challenging existing detection methods. Moreover, current approaches lack granularity in pinpointing anomalies at the operational level. They treat entire sequences of operations as anomalies, though the majority likely represent normal behavior, with only a few as anomalies. This paper presents UDAD, a novel method for precisely detecting stealthy abnormal database access behaviors. By transforming SQL statements into semantic vectors, we enhance the learning of embedded semantic information. Through the integration of an attention-based BiLSTM model and an autoencoder, UDAD achieves accurate detection and precise localization of abnormal operations. We evaluate UDAD on publicly available datasets, demonstrating its superiority over state-of-the-art methods. Huazhen Zhong, Weifang Zhang, Wenjie Xiao, Xuehai Tang, Liangjun Zang |
IPCCC | 7 |
| 2022 | ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence EmbeddingabstractContrastive learning has been attracting much attention for learning unsupervised sentence embeddings. The current state-of-the-art unsupervised method is the unsupervised SimCSE (unsup-SimCSE). Unsup-SimCSE takes dropout as a minimal data augmentation method, and passes the same input sentence to a pre-trained Transformer encoder (with dropout turned on) twice to obtain the two corresponding embeddings to build a positive pair. As the length information of a sentence will generally be encoded into the sentence embeddings due to the usage of position embedding in Transformer, each positive pair in unsup-SimCSE actually contains the same length information. And thus unsup-SimCSE trained with these positive pairs is probably biased, which would tend to consider that sentences of the same or similar length are more similar in semantics. Through statistical observations, we find that unsup-SimCSE does have such a problem. To alleviate it, we apply a simple repetition operation to modify the input sentence, and then pass the input sentence and its modified counterpart to the pre-trained Transformer encoder, respectively, to get the positive pair. Additionally, we draw inspiration from the community of computer vision and introduce a momentum contrast, enlarging the number of negative pairs without additional calculations. The proposed two modifications are applied on positive and negative pairs separately, and build a new sentence embedding method, termed Enhanced Unsup-SimCSE (ESimCSE). We evaluate the proposed ESimCSE on several benchmark datasets w.r.t the semantic text similarity (STS) task. Experimental results show that ESimCSE outperforms the state-of-the-art unsup-SimCSE by an average Spearman correlation of 2.02% on BERT-base. Xing Wu 0002, Chaochen Gao, Liangjun Zang, Jizhong Han, Zhongyuan Wang 0006, Songlin Hu 0001 |
COLING | 3 |
| 2022 | Emotionflow: Capture the Dialogue Level Emotion TransitionsabstractEmotion recognition in conversations (ERC) has attracted increasing interests in recent years, due to its wide range of applications, such as customer service analysis, health-care consultation, etc. One key challenge of ERC is that users' emotions would change due to the impact of others' emotions. That is, the emotions within the conversation can spread among the communication participants. However, the spread impact of emotions in a conversation is rarely addressed in existing researches. To this end, we propose EmotionFlow for ERC with the consideration of the spread of participants' emotions during a conversation. EmotionFlow first encodes users' utterance by concatenating the context with an auxiliary question, which helps to learn user-specific features. Then, conditional random field is applied to capture the sequential information at emotional level. We conduct extensive experiments on a public dataset Multimodal EmotionLines Dataset (MELD), and the results demonstrate the effectiveness of our proposed model. Liangjun Zang, Rong Zhang 0006, Songlin Hu 0001, Longtao Huang |
ICASSP | 2 |
| 2022 | A Knowledge/Data Enhanced Method for Joint Event and Temporal Relation ExtractionabstractUnderstanding temporal relations (TempRels) between events is an important task that could benefit many downstream NLP applications. This task inevitably faces the challenges of both a limited amount of high-quality training data and a very biased distribution of TempRels. These problems will substantially hurt the performance of extraction systems because they are inclined to predict dominant TempRels when training with a limited amount of data. To alleviate those issues, we propose a Knowledge/Data Enhanced method for Event and TempRel Extraction, which integrates the temporal commonsense knowledge, data augmentation and Focal Loss function into one single extraction system. Altogether, these components improve the performance of the system on two public benchmark datasets TB-Dense and MATRES1. Liangjun Zang, Songlin Hu 0001 |
ICASSP | 2 |
| 2021 | Fed-Tra: Improving Accuracy of Deep Learning Model on Non-iid in Federated Learning
Wenjie Xiao, Xuehai Tang, Biyu Zhou, Wang Wang, Yangchen Dong, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
ICA3PP (1) | 6 |
| 2020 | Symmetric Metric Learning with Adaptive Margin for RecommendationabstractMetric learning based methods have attracted extensive interests in recommender systems. Current methods take the user-centric way in metric space to ensure the distance between user and negative item to be larger than that between the current user and positive item by a fixed margin. While they ignore the relations among positive item and negative item. As a result, these two items might be positioned closely, leading to incorrect results. Meanwhile, different users usually have different preferences, the fixed margin used in those methods can not be adaptive to various user biases, and thus decreases the performance as well. To address these two problems, a novel Symmetic Metric Learning with adaptive margin (SML) is proposed. In addition to the current user-centric metric, it symmetically introduces a positive item-centric metric which maintains closer distance from positive items to user, and push the negative items away from the positive items at the same time. Moreover, the dynamically adaptive margins are well trained to mitigate the impact of bias. Experimental results on three public recommendation datasets demonstrate that SML produces a competitive performance compared with several state-of-the-art methods. Fuqing Zhu, Wanhui Qian, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
AAAI | 5 |
| 2020 | An Event-Oriented Neural Ranking Model for News RetrievalabstractEvent-oriented news retrieval (ENR) is the task of retrieving news articles related to the specific event in response to the event-oriented query. Previous approaches usually focus on optimizing traditional retrieval models through hand-crafted features from the perspective of new articles. However, these approaches often fail to work well in reality, as they do not consider the essential natures of the event, i.e., dynamics, coupling. In this paper, we propose a novel and effective event-oriented neural ranking model for news retrieval (ENRMNR). Our model exploits a deep attention mechanism to tackle the dynamics and coupling derived from event evolution. Specifically, the word-level bidirectional attention allows the model to identify which query words about the subevent are related to the news article words, and vice-versa, in order to tackle the dynamics. Moreover, the hierarchical attention at passage-level and document-level allows it to capture fine-grained event representations for the coupling between different events within a news article. Experimental results on real-world datasets demonstrate that ENRMNR model significantly outperforms competitive models. Wanhui Qian, Liangjun Zang, Fuqing Zhu, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001 |
CIKM | 3 |
| 2020 | A Rating Bias Formulation based on Fuzzy Set for RecommendationabstractIn recommender systems, the user uncertain preference results in unexpected ratings. Previous approaches (e.g., BiasMF) only adjust the rating value based on the bias vector, ignoring the uncertainty of rating. This paper makes an initial attempt in integrating the influence of user uncertain degree and user rating bias into the matrix factorization framework, simultaneously. An approach based on fuzzy set, called fuZzy Matrix Factorization (ZMF), is proposed. Specifically, a fuzzy set of like is defined for each user, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Then, the user uncertain preference matrix is obtained, which could explain and represent the user bias and uncertainty effectively. Furthermore, to enhance the computational impact on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Besides, the proposed approach could be extended to others due to independency on additional data sources. Experimental results on three datasets show that ZMF produces an effective improvement. Fuqing Zhu, Jiao Dai, Liangjun Zang, Yipeng Su, Jizhong Han, Songlin Hu 0001 |
IJCNN | 4 |
| 2020 | AutoSUM: Automating Feature Extraction and Multi-user Preference Simulation for Entity Summarization
Dongjun Wei, Fuqing Zhu, Liangjun Zang, Wei Zhou 0019, Songlin Hu 0001 |
PAKDD (2) | 4 |
| 2020 | Yet another approach to understanding news event evolution
Shangwen Lv, Longtao Huang, Liangjun Zang, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001 |
World Wide Web | 3 |
| 2019 | A Fuzzy Set Based Approach for Rating BiasabstractIn recommender systems, the user uncertain preference results in unexpected ratings. This paper makes an initial attempt in integrating the influence of user uncertain degree into the matrix factorization framework. Specifically, a fuzzy set of like for each user is defined, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Furthermore, to enhance the computational effect on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Experimental results on three real-world datasets show that the proposed approach produces stable improvements compared with others. Jiao Dai, Fuqing Zhu, Liangjun Zang, Songlin Hu 0001, Jizhong Han |
AAAI | 4 |
| 2019 | Mask and Infill: Applying Masked Language Model for Sentiment TransferabstractThis paper focuses on the task of sentiment transfer on non-parallel text, which modifies sentiment attributes (e.g., positive or negative) of sentences while preserving their attribute-independent contents. Existing methods adopt RNN encoder-decoder structure to generate a new sentence of a target sentiment word by word, which is trained on a particular dataset from scratch and have limited ability to produce satisfactory sentences. When people convert the sentiment attribute of a given sentence, a simple but effective approach is to only replace the sentiment tokens of the sentence with other expressions indicative of the target sentiment, instead of building a new sentence from scratch. Such a process is very similar to the task of Text Infilling or Cloze. With this intuition, we propose a two steps approach: Mask and Infill. In the \emph{mask} step, we identify and mask the sentiment tokens of a given sentence. In the \emph{infill} step, we utilize a pre-trained Masked Language Model (MLM) to infill the masked positions by predicting words or phrases conditioned on the context\footnote{In this paper, \emph{content} and \emph{context} are equivalent, \emph{style}, \emph{attribute} and \emph{label} are equivalent.}and target sentiment. We evaluate our model on two review datasets \emph{Yelp} and \emph{Amazon} by quantitative, qualitative, and human evaluations. Experimental results demonstrate that our model achieve state-of-the-art performance on both accuracy and BLEU scores. Xing Wu 0002, Tao Zhang 0101, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
IJCAI | 3 |
| 2019 | LMLSTM: Extract Event-Oriented Keyphrase From News Stream
Longtao Huang, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
IJCNN | 3 |
| 2018 | An Interactivity-Based Personalized Mutual Reinforcement Model for Microblog Topic Summarization
Lu Zhang 0038, Liangjun Zang, Longtao Huang, Jizhong Han, Songlin Hu 0001 |
PRICAI (1) | 2 |
| 2015 | The Double-Level Default Description Logic D 3 LabstractWe propose the default description logic $$\mathcal {D}2\mathcal {L}$$ and the double-level default description logic $$\mathcal {D}3\mathcal {L}$$ . $$\mathcal {D}2\mathcal {L}$$ embeds normal defaults inside the basic description logic $$\mathcal {ALC},$$ and $$\mathcal {D}3\mathcal {L}$$ augments $$\mathcal {D}2\mathcal {L}$$ with normal double-level defaults. Double-level defaults are defaults of defaults and can be used to represent default inheritance of default properties of concepts in ontologies. A $$\mathcal {D}3\mathcal {L}$$ knowledge base ( $$\mathcal {D}3\mathcal {L}$$ -KB) can be divided into two levels of knowledge bases, and correspondingly its extensions can be computed in two steps. $$\mathcal {D}3\mathcal {L}$$ is more expressive than $$\mathcal {D}2\mathcal {L}$$ since there is a $$\mathcal {D}3\mathcal {L}$$ -KB that cannot reduce to any $$\mathcal {D}2\mathcal {L}$$ -KB. Specifically, there is a $$\mathcal {D}3\mathcal {L}$$ -KB such that the set of all its extensions cannot be exactly generated by any $$\mathcal {D}2\mathcal {L}$$ -KB. Liangjun Zang, Weimin Wang 0002, Cun-gen Cao 0001 |
KSEM | 1 |
| 2015 | A Chinese Framework of Semantic Taxonomy and Description: Preliminary Experimental Evaluation Using Web Information ExtractionabstractThe Chinese Framework of Semantic Taxonomy and Description (FSTD) is a linguistic resource that stores lexical and predicate-argument semantics about events or states in Chinese text, developed with the application of knowledge acquisition from Chinese text in mind. In this paper we build a web information extraction system, called NkiExtractor, to evaluate FSTD experimentally. We use two metrics: grammar coverage measures whether there is a semantic category of FSTD that corresponds to an event description in text, and extraction precision measures whether the correct predicate-argument structure can be extracted from text. Experimental results show that FSTD is a fairly comprehensive and effective resource for knowledge acquisition. We also discuss future work for expanding FSTD and improving extraction precision of NkiExtractor. Liangjun Zang, Weimin Wang 0002, Fang Fang 0009, Cong Cao 0001, Cun-gen Cao 0001 |
KSEM | 1 |
| 2013 | A Survey of Commonsense Knowledge Acquisition
Liangjun Zang, Cong Cao 0001, Yanan Cao 0001, Yuming Wu, Cun-gen Cao 0001 |
J. Comput. Sci. Technol. | 1 |