Zongcheng Ji

dblp:117/4334 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0002-4963-8422ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 BiMarker: Enhancing text watermark detection for large language models with bipolar watermarks
Qiuping Yi, Zongcheng Ji, Yijian Lu, Shun Zou, Yanqi Li, Keyang Xiao, Hongliang Liang
Neurocomputing3
2026 An end-to-end approach for fixing concurrency bugs via SHB-based context extractor
Qiuping Yi, Keyang Xiao, Zongcheng Ji, Hongliang Liang
J. Syst. Softw.4
2026 CosFormer: A Code Semantic-Aware Transformer for Vulnerability Detection
abstract
Deep learning-based vulnerability detection has made significant strides, surpassing traditional static and dynamic analysis methods. However, existing approaches, including Graph Neural Networks (GNNs) and Transformer-based models, still struggle to fully capture complex code semantics. In this paper, we proposeCosFormer, a novel Code Semantic-aware Transformer tailored for vulnerability detection.CosFormerintroduces two key components: Code Semantic-aware Embedding, which enhances semantic representation at both the token and line levels, and Spatial Dependency-aware Encoding, which integrates structural dependencies from Control Flow Graphs (CFGs) and Program Dependency Graphs (PDGs) to guide attention toward vulnerability-relevant code. We evaluateCosFormeron four benchmark datasets, including a real-world dataset, and demonstrate its superior performance.CosFormerachieves the highest F1 scores across all tasks, outperforming state-of-the-art GNN-based and Transformer-based models, as well as large language models (LLMs). Notably,CosFormerachieves a Cross F1 of 69.37 and a Mixed F1 of 58.42 in generalization evaluation, surpassing all baselines. These results highlightCosFormer’s effectiveness and robustness in detecting vulnerabilities across diverse and previously unseen codebases.
Qianyue Wei, Qiuping Yi, Zongcheng Ji, Hongliang Liang
IEEE Trans. Software Eng.4
2025 DuST: Chinese NER using dual-grained syntax-aware transformer network
Yinlong Xiao, Zongcheng Ji, Jianqiang Li 0002
Inf. Process. Manag.2
2025 DualFLAT: Dual Flat-Lattice Transformer for domain-specific Chinese named entity recognition
Yinlong Xiao, Zongcheng Ji, Jianqiang Li 0002, Qing Zhu 0004
Inf. Process. Manag.2
2024 LLET: Lightweight Lexicon-Enhanced Transformer for Chinese NER
abstract
The Flat-LAttice Transformer (FLAT) has achieved notable success in Chinese named entity recognition (NER) by integrating lexical information into the widely-used Transformer encoder. FLAT enhances each sentence by constructing a flat lattice, a token sequence with characters and matched lexicon words, and calculating self-attention among tokens. However, FLAT faces a quadruple complexity challenge, especially with lengthy sentences containing numerous matched words, significantly increasing memory and computational costs. To alleviate this issue, we propose a novel lightweight lexicon-enhanced Transformer (LLET) for Chinese NER. Specifically, we introduce two distinct variants that focus on character attention to characters and words, both jointly and separately. Experimental results conducted on four public Chinese NER datasets show that both variants achieve significant memory savings while maintaining comparable performance when compared to FLAT.
Zongcheng Ji, Yinlong Xiao
ICASSP1
2024 Dust: Dual-Grained Syntax-Aware Transformer Network for Chinese Named Entity Recognition
abstract
Named Entity Recognition (NER) is a fundamental task in natural language processing. Syntax plays a significant role in helping to recognize the boundaries and types of entities. In comparison to English, Chinese NER, due to the absence of explicit delimiters, often faces challenges in determining entity boundaries. Similarly, syntactic parsing results can also lead to errors caused by wrong segmentation. In this paper, we propose the dual-grained syntax-aware Transformer network to mitigate the noise from single-grained syntactic parsing results by incorporating dual-grained syntactic information. Specifically, we first introduce syntax-aware Transformers to model dual-grained syntax-aware features and a contextual Transformer to model contextual features. We then design a triple feature aggregation module to dynamically fuse these features. We validate the effectiveness of our approach on three public datasets.
Yinlong Xiao, Zongcheng Ji
ICASSP2
2024 MVT: Chinese NER Using Multi-View Transformer
abstract
Integrating lexical knowledge in Chinese named entity recognition (NER) has been proven effective. Among the existing methods, Flat-LAttice Transformer (FLAT) has achieved great success in both performance and efficiency. FLAT performs lexical enhancement for each sentence by constructing a flat lattice (i.e., a sequence of tokens including the characters in a sentence and the matched words in a lexicon) and calculating self-attention with a fully-connected structure. However, the different interactions between tokens, which can bring different aspects of semantic information for Chinese NER, cannot be well captured by self-attention with a fully-connected structure. In this paper, we propose a novel Multi-View Transformer (MVT) to effectively capture the different interactions between tokens. We first define four views to capture four different token interaction structures. We then construct a view-aware visible matrix for each view according to the corresponding structure and introduce a view-aware dot-product attention for each view to limit the attention scope by incorporating the corresponding visible matrix. Finally, we design three different MVT variants to fuse the multi-view features at different levels of the Transformer architecture. Experimental results conducted on four public Chinese NER datasets show the effectiveness of the proposed method. Specifically, on the most challenging dataset Weibo, which is in an informal text style, MVT outperforms FLAT in F1 score by 2.56%, and when combined with BERT, MVT outperforms FLAT in F1 score by 3.03%.
Yinlong Xiao, Zongcheng Ji, Jianqiang Li 0002
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 A Neural Transition-based Joint Model for Disease Named Entity Recognition and Normalization
abstract
Zongcheng Ji, Tian Xia, Mei Han, Jing Xiao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Zongcheng Ji, Tian Xia 0004, Jing Xiao 0006
ACL/IJCNLP (1)1
2021 Leveraging Large-Scale Weakly Labeled Data for Semi-Supervised Mass Detection in Mammograms
abstract
Mammographic mass detection is an integral part of a computer-aided diagnosis system. Annotating a large number of mammograms at pixel-level in order to train a mass detection model in a fully supervised fashion is costly and time-consuming. This paper presents a novel self-training framework for semi-supervised mass detection with soft image-level labels generated from diagnosis reports by Mammo-RoBERTa, a RoBERTa-based natural language processing model fine-tuned on the fully labeled data and associated mammography reports. Starting with a fully supervised model trained on the data with pixel-level masks, the proposed framework iteratively refines the model itself using the entire weakly labeled data (image-level soft label) in a self-training fashion. A novel sample selection strategy is proposed to identify those most informative samples for each iteration, based on the current model output and the soft labels of the weakly labeled data. A soft cross-entropy loss and a soft focal loss are also designed to serve as the image-level and pixel-level classification loss respectively. Our experiment results show that the proposed semi-supervised framework can improve the mass detection accuracy on top of the supervised baseline, and outperforms the previous state-of-the-art semi-supervised approaches with weakly labeled data, in some cases by a large margin.
Yuxing Tang, Zhenjie Cao, Zongcheng Ji, Jing Xiao 0006, Peng Chang 0002
CVPR5
2020 A study of deep learning approaches for medication and adverse drug event extraction from clinical text
abstract
OBJECTIVE: This article presents our approaches to extraction of medications and associated adverse drug events (ADEs) from clinical documents, which is the second track of the 2018 National NLP Clinical Challenges (n2c2) shared task. MATERIALS AND METHODS: The clinical corpus used in this study was from the MIMIC-III database and the organizers annotated 303 documents for training and 202 for testing. Our system consists of 2 components: a named entity recognition (NER) and a relation classification (RC) component. For each component, we implemented deep learning-based approaches (eg, BI-LSTM-CRF) and compared them with traditional machine learning approaches, namely, conditional random fields for NER and support vector machines for RC, respectively. In addition, we developed a deep learning-based joint model that recognizes ADEs and their relations to medications in 1 step using a sequence labeling approach. To further improve the performance, we also investigated different ensemble approaches to generating optimal performance by combining outputs from multiple approaches. RESULTS: Our best-performing systems achieved F1 scores of 93.45% for NER, 96.30% for RC, and 89.05% for end-to-end evaluation, which ranked #2, #1, and #1 among all participants, respectively. Additional evaluations show that the deep learning-based approaches did outperform traditional machine learning algorithms in both NER and RC. The joint model that simultaneously recognizes ADEs and their relations to medications also achieved the best performance on RC, indicating its promise for relation extraction. CONCLUSION: In this study, we developed deep learning approaches for extracting medications and their attributes such as ADEs, and demonstrated its superior performance compared with traditional machine learning algorithms, indicating its uses in broader NER and RC tasks in the medical domain.
Qiang Wei 0002, Zongcheng Ji, Zhiheng Li 0004, Jingcheng Du, Jun Xu 0007, Yang Xiang 0003, Firat Tiryaki, Stephen Wu 0004, Yaoyun Zhang, Cui Tao, Hua Xu 0001
J. Am. Medical Informatics Assoc.2
2020 Deep learning in clinical natural language processing: a methodical review
abstract
OBJECTIVE: This article methodically reviews the literature on deep learning (DL) for natural language processing (NLP) in the clinical domain, providing quantitative analysis to answer 3 research questions concerning methods, scope, and context of current research. MATERIALS AND METHODS: We searched MEDLINE, EMBASE, Scopus, the Association for Computing Machinery Digital Library, and the Association for Computational Linguistics Anthology for articles using DL-based approaches to NLP problems in electronic health records. After screening 1,737 articles, we collected data on 25 variables across 212 papers. RESULTS: DL in clinical NLP publications more than doubled each year, through 2018. Recurrent neural networks (60.8%) and word2vec embeddings (74.1%) were the most popular methods; the information extraction tasks of text classification, named entity recognition, and relation extraction were dominant (89.2%). However, there was a "long tail" of other methods and specific tasks. Most contributions were methodological variants or applications, but 20.8% were new methods of some kind. The earliest adopters were in the NLP community, but the medical informatics community was the most prolific. DISCUSSION: Our analysis shows growing acceptance of deep learning as a baseline for NLP research, and of DL-based NLP in the medical community. A number of common associations were substantiated (eg, the preference of recurrent neural networks for sequence-labeling named entity recognition), while others were surprisingly nuanced (eg, the scarcity of French language clinical NLP with deep learning). CONCLUSION: Deep learning has not yet fully penetrated clinical NLP and is growing rapidly. This review highlighted both the popular and unique trends in this active field.
Stephen Wu 0004, Kirk Roberts, Surabhi Datta, Jingcheng Du, Zongcheng Ji, Yuqi Si, Sarvesh Soni, Qiang Wei 0002, Yang Xiang 0003, Bo Zhao 0001, Hua Xu 0001
J. Am. Medical Informatics Assoc.5
2020 A study of entity-linking methods for normalizing Chinese diagnosis and procedure terms to ICD codes
Zongcheng Ji, Stephen Wu 0004, Weiyan Lin, Wenzhen Li, Guohong Xiao, Hua Xu 0001, Yi Zhou 0005
J. Biomed. Informatics2
2019 Relation Extraction from Clinical Narratives Using Pre-trained Language Models
Qiang Wei 0002, Zongcheng Ji, Yuqi Si, Jingcheng Du, Firat Tiryaki, Stephen Wu 0004, Cui Tao, Kirk Roberts, Hua Xu 0001
AMIA2
2019 Extracting entities with attributes in clinical text via joint deep learning
abstract
OBJECTIVE: Extracting clinical entities and their attributes is a fundamental task of natural language processing (NLP) in the medical domain. This task is typically recognized as 2 sequential subtasks in a pipeline, clinical entity or attribute recognition followed by entity-attribute relation extraction. One problem of pipeline methods is that errors from entity recognition are unavoidably passed to relation extraction. We propose a novel joint deep learning method to recognize clinical entities or attributes and extract entity-attribute relations simultaneously. MATERIALS AND METHODS: The proposed method integrates 2 state-of-the-art methods for named entity recognition and relation extraction, namely bidirectional long short-term memory with conditional random field and bidirectional long short-term memory, into a unified framework. In this method, relation constraints between clinical entities and attributes and weights of the 2 subtasks are also considered simultaneously. We compare the method with other related methods (ie, pipeline methods and other joint deep learning methods) on an existing English corpus from SemEval-2015 and a newly developed Chinese corpus. RESULTS: Our proposed method achieves the best F1 of 74.46% on entity recognition and the best F1 of 50.21% on relation extraction on the English corpus, and 89.32% and 88.13% on the Chinese corpora, respectively, which outperform the other methods on both tasks. CONCLUSIONS: The joint deep learning-based method could improve both entity recognition and relation extraction from clinical text in both English and Chinese, indicating that the approach is promising.
Xue Shi, Yingping Yi, Buzhou Tang, Qingcai Chen, Xiaolong Wang 0001, Zongcheng Ji, Yaoyun Zhang, Hua Xu 0001
J. Am. Medical Informatics Assoc.7
2018 Linking Fine-Grained Locations in User Comments (Extended Abstract)
abstract
Many domain-specific websites host a profile page for each entity (e.g., locations on Foursquare, movies on IMDb, and products on Amazon), and users can post comments on it. When commenting on an entity, users often mention other entities for reference or comparison. Compared with web pages and tweets, disambiguating the mentioned entities in user comments has not received much attention. This paper investigates linking fine-grained locations in Foursquare comments. We demonstrate that the focal location, i.e., the location that a comment is posted on, provides rich contexts for linking. To exploit such information, we represent the Foursquare data in a graph, which includes locations, comments, and their relations. A probabilistic model named FocalLink is proposed to estimate the probability that a user mentions a location when commenting on a focal location, by following different kinds of relations. Experimental results show that FocalLink is consistently superior to different baselines.
Jialong Han, Aixin Sun, Gao Cong, Wayne Xin Zhao, Zongcheng Ji, Minh C. Phan
ICDE5
2018 Linking Fine-Grained Locations in User Comments
abstract
Many domain-specific websites host a profile page for each entity (e.g., locations on Foursquare, movies on IMDb, and products on Amazon) for users to post comments on. When commenting on an entity, users often mention other entities for reference or comparison. Compared with web pages and tweets, the problem of disambiguating the mentioned entities in user comments has not received much attention. This paper investigates linking fine-grained locations in Foursquare comments. We demonstrate that the focal location, i.e., the location that a comment is posted on, provides rich contexts for the linking task. To exploit such information, we represent the Foursquare data in a graph, which includes locations, comments, and their relations. A probabilistic model named FocalLink is proposed to estimate the probability that a user mentions a location when commenting on a focal location, by following different kinds of relations. Experimental results show that FocalLink is consistently superior under different collective linking settings.
Jialong Han, Aixin Sun, Gao Cong, Wayne Xin Zhao, Zongcheng Ji, Minh C. Phan
IEEE Trans. Knowl. Data Eng.5
2016 Joint Recognition and Linking of Fine-Grained Locations from Tweets
abstract
Many users casually reveal their locations such as restaurants, landmarks, and shops in their tweets. Recognizing such fine-grained locations from tweets and then linking the location mentions to well-defined location profiles (e.g., with formal name, detailed address, and geo-coordinates etc.) offer a tremendous opportunity for many applications. Different from existing solutions which perform location recognition and linking as two sub-tasks sequentially in a pipeline setting, in this paper, we propose a novel joint framework to perform location recognition and location linking simultaneously in a joint search space. We formulate this end-to-end location linking problem as a structured prediction problem and propose a beam-search based algorithm. Based on the concept of multi-view learning, we further enable the algorithm to learn from unlabeled data to alleviate the dearth of labeled data. Extensive experiments are conducted to recognize locations mentioned in tweets and link them to location profiles in Foursquare. Experimental results show that the proposed joint learning algorithm outperforms the state-of-the-art solutions, and learning from unlabeled data improves both the recognition and linking accuracy.
Zongcheng Ji, Aixin Sun, Gao Cong, Jialong Han
WWW1
2013 Learning to rank for question routing in community question answering
abstract
This paper focuses on the problem of Question Routing (QR) in Community Question Answering (CQA), which aims to route newly posted questions to the potential answerers who are most likely to answer them. Traditional methods to solve this problem only consider the text similarity features between the newly posted question and the user profile, while ignoring the important statistical features, including the question-specific statistical feature and the user-specific statistical features. Moreover, traditional methods are based on unsupervised learning, which is not easy to introduce the rich features into them. This paper proposes a general framework based on the learning to rank concepts for QR. Training sets consist of triples (q, asker, answerers) are first collected. Then, by introducing the intrinsic relationships between the asker and the answerers in each CQA session to capture the intrinsic labels/orders of the users about their expertise degree of the question q, two different methods, including the SVM-based and RankingSVM-based methods, are presented to learn the models with different example creation processes from the training set. Finally, the potential answerers are ranked using the trained models. Extensive experiments conducted on a real world CQA dataset from Stack Overflow show that our proposed two methods can both outperform the traditional query likelihood language model (QLLM) as well as the state-of-the-art Latent Dirichlet Allocation based model (LDA). Specifically, the RankingSVM-based method achieves statistical significant improvements over the SVM-based method and has gained the best performance.
Zongcheng Ji, Bin Wang 0004
CIKM1
2013 Improving Alignment of System Combination by Using Multi-objective Optimization
abstract
This paper proposes a multi-objective optimization framework which supports heterogeneous information sources to improve alignment in machine translation system combination techniques.In this area, most of techniques usually utilize confusion networks (CN) as their central data structure to compact an exponential number of an potential hypotheses, and because better hypothesis alignment may benefit constructing better quality confusion networks, it is natural to add more useful information to improve alignment results.However, these information may be heterogeneous, so the widely-used Viterbi algorithm for searching the best alignment may not apply here.In the multi-objective optimization framework, each information source is viewed as an independent objective, and a new goal of improving all objectives can be searched by mature algorithms.The solutions from this framework, termed Pareto optimal solutions, are then combined to construct confusion networks.Experiments on two Chinese-to-English translation datasets show significant improvements, 0.97 and 1.06 BLEU points over a strong Indirected Hidden Markov Model-based (IHMM) system, and 4.75 and 3.53 points over the best single machine translation systems.
Tian Xia 0004, Zongcheng Ji, Shaodan Zhai, Yidong Chen 0001, Qun Liu 0001
EMNLP2
2012 Question-answer topic model for question retrieval in community question answering
abstract
The major challenge for Question Retrieval (QR) in Community Question Answering (CQA) is the lexical gap between the queried question and the historical questions. This paper proposes a novel Question-Answer Topic Model (QATM) to learn the latent topics aligned across the question-answer pairs to alleviate the lexical gap problem, with the assumption that a question and its paired answer share the same topic distribution. Experiments conducted on a real world CQA dataset from Yahoo! Answers show that combining both parts properly can get more knowledge than each part or both parts in a simple mixing way and combining our QATM with the state-of-the-art translation-based language model, where the topic and translation information is learned from the question-answer pairs at two different grained semantic levels respectively, can significantly improve the QR performance.
Zongcheng Ji, Bin Wang 0004, Ben He 0001
CIKM1
2012 Dual role model for question recommendation in community question answering
abstract
Question recommendation that automatically recommends a new question to suitable users to answer is an appealing and challenging problem in the research area of Community Question Answering (CQA). Unlike in general recommender systems where a user has only a single role, each user in CQA can play two different roles (dual roles) simultaneously: as an asker and as an answerer. To the best of our knowledge, this paper is the first to systematically investigate the distinctions between the two roles and their different influences on the performance of question recommendation in CQA. Moreover, we propose a Dual Role Model (DRM) to model the dual roles of users effectively. With different indepen-dence assumptions, two variants of DRM are achieved. Finally, we present the DRM based approach to question recommendation which provides a mechanism for naturally integrating the user relation between the answerer and the asker with the content re-levance between the answerer and the question into a uni-fied probabilistic framework. Experiments using a real-world data crawled from Yahoo! Answers show that: (1) there are evident distinctions between the two roles of users in CQA. Additionally, the answerer role is more effective than the asker role for modeling candidate users in question recommendation; (2) compared with baselines utilizing a single role or blended roles based methods, our DRM based approach consistently and significantly improves the performance of question recommendation, demonstrating that our approach can model the user in CQA more reasonably and precisely.
Zongcheng Ji, Bin Wang 0004
SIGIR2