EDBT 2026 Demo / reviewers in the wild / expert
Nankai Lin
dblp:235/2956
· DBLP profile ↗
50ranked-venue papers
13as first author
49since 2021 · last 2026
0000-0003-2838-8273ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 9 first-author · 31 since 2021Human-computer interaction and ubiquitous computing · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chameleon: Benchmarking Detection and Backtracking on Commercial-Grade AI-Generated VideosabstractThe proliferation of AI-Generated Content (AIGC), especially deepfake videos, poses a severe threat to social trust by enabling fraud, privacy violations and disinformation. Existing AI-generated video detection (AGVD) benchmarks focus on open-source model generated videos, yet commercial closed-source models produce more realistic, temporally coherent videos that are underexplored in detection research. To fill this gap, we present Chameleon, a commercial-grade dataset with 1,700 AI-generated videos from 600 real-world sources across three key domains (News, Speech, Recommendation), featuring high resolution, rich annotations and 3D consistency metrics for dynamic scene spatial coherence, shifting detection from face-centric forgery to holistic scene forensics. This benchmark assesses models on two core tasks: accurate AI video detection in real-world conditions and forensic backtracking of original sources. Experimental results reveal critical limitations of existing methods in detecting and backtracking high-fidelity, spatiotemporally consistent videos from commercial closed-source models, highlighting current methods’ flawed forensic reasoning and establishing Chameleon as a vital challenge for AIGC security research. The code and data are available at https://github.com/lxixim/Chameleon. Xingming Liao, Meiyu Zeng, Canyu Chen, Nankai Lin, Zhuowei Wang 0001, Aimin Yang 0002 |
ICMR | 4 |
| 2026 | Multi-scenario CTR prediction via enhanced scene-aware transformer frameworkabstractAbstract In the context of multi-scenario recommendation, multi-scenario click-through-rate (MS-CTR) prediction plays a crucial role in effectively personalizing recommendations on commercial platforms. However, when confronted with multi-scenario data, models often face challenges such as overfitting, inadequate feature representation, and unstable optimization, which hinder the performance and reliability of CTR prediction. To tackle these challenges, this study introduces the enhanced scene-aware transformer (ESAT) framework for MS-CTR prediction. This framework is divided into five modules. First, the structural position-aware scene encoding module converts scene attributes into fixed-dimensional embedding vectors and the scene adaptive transformation module uses a nonlinear transformation to dynamically adjust scene features. The cross-scene regularization module uses multi-sample dropout technology to enhance generalization ability and prevent overfitting. The scene-aware discriminative learning module applies contrastive learning to optimize the similarity between samples. Finally, the hierarchical stability control module introduces ClippyGrad optimizer and $L_\infty $ regularization to accurately control gradient updates, avoid excessive steps, and improve training stability. Experiments conducted on large-scale multi-scenario datasets confirm the effectiveness of the proposed module. The results indicate that the ESAT model substantially elevates performance in MS-CTR prediction tasks, especially in terms of generalization to new scenes, precision in feature representation, and stability in parameter updates. Weizhong Liu, Qifeng Bai, Feiyan Pang, Nankai Lin, Aimin Yang 0002 |
Comput. J. | 5 |
| 2026 | Advancing LLMs for Chinese semantic error correction: Example selection and re-scoring
Nankai Lin, Shengyi Jiang, Lianxi Wang 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Jailbreaking? One Step Is Enough!abstractWeixiong Zheng, Peijian Zeng, YiWei Li, Hongyan Wu, Nankai Lin, Junhao Chen, Aimin Yang, Yongmei Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weixiong Zheng, Peijian Zeng, Nankai Lin, Aimin Yang 0002, Yongmei Zhou |
ACL (1) | 5 |
| 2025 | JUDICIOUS: Evaluating Robustness of Large Language Models in the Legal Realm
Ziling Dai, Nankai Lin |
CogSci | 2 |
| 2025 | Rethinking Vocabulary Augmentation: Addressing the Challenges of Low-Resource Languages in Multilingual ModelsabstractThe performance of multilingual language models (MLLMs) is notably inferior for low-resource languages (LRL) compared to high-resource ones, primarily due to the limited available corpus during the pre-training phase. This inadequacy stems from the under-representation of low-resource language words in the subword vocabularies of MLLMs, leading to their misidentification as unknown or incorrectly concatenated subwords. Previous approaches are based on frequency sorting to select words for augmenting vocabularies. However, these methods overlook the fundamental disparities between model representation distributions and frequency distributions. To address this gap, we introduce a novel Entropy-Consistency Word Selection (ECWS) method, which integrates semantic and frequency metrics for vocabulary augmentation. Our results indicate an improvement in performance, supporting our approach as a viable means to enrich vocabularies inadequately represented in current MLLMs. Nankai Lin, Peijian Zeng, Weixiong Zheng, Shengyi Jiang, Dong Zhou 0001, Aimin Yang 0002 |
COLING | 1 |
| 2025 | Pseudo-label Data Construction Method and Syntax-enhanced Model for Chinese Semantic Error RecognitionabstractChinese Semantic Error Recognition (CSER) has always been a weak link in Chinese language processing due to the complexity and obscureness of Chinese semantics. Existing research has gradually focused on leveraging pre-trained models to perform CSER. Although some researchers have attempted to integrate syntax information into the pre-trained language model, it requires training the models from scratch, which is time-consuming and laborious. Furthermore, despite the existence of datasets for CSER, the constrained size of these datasets impairs the performance of the models. Thus, in order to address the difficulty posed by a limited sample set and the need of annotating samples with semantic-level errors, we propose a Pseudo-label Data Construction method for CSER (PDC-CSER), generating pseudo-labels for augmented samples based on perplexity and model respectively, which overcomes the difficulty of constructing pseudo-label data containing semantic-level errors and ensures the quality of pseudo-labels. Moreover, we propose a CSER method with the Dependency Syntactic Attention mechanism (CSER-DSA) to explicitly infuse dependency syntactic information only in the fine-tuning stage, achieving robust performance, and simultaneously reducing substantial computing power and time cost. Results demonstrate that the pseudo-label technology PDC-CSER and the semantic error recognition method CSER-DSA surpass the existing models Nankai Lin, Shengyi Jiang, Lianxi Wang 0001, Aimin Yang 0002 |
COLING | 2 |
| 2025 | WATER: A Two-Stage in-Context Learning Debiasing Framework for Multilingual Text ClassificationabstractRecently, Large Language Models (LLMs) have shown remarkable success across a variety of tasks, with rapid advancements in supporting multilingual capabilities. However, these models exhibit varying degrees of demographic biases in text classification tasks. Most existing research focuses on debiasing pre-trained models or addressing biases in monolingual text classification, resulting in limited exploration in multilingual contexts. To solve the above problems, this paper introduces a tWo-stAge in-conText learning dEbiasing fRamework (WATER). Our approach does not require updating the model's parameters and is adaptable to any language. It includes three key modules: sample selection, sample filtering, and template filling and prediction. In the first stage, we leverage a sample selection module to identify text that closely matches the model embeddings. In the second stage, we introduce an innovative Contextual Disparity Measure (CDM) in the sample filtering module to filter out samples that effectively address the bias associated with specific attributes. Finally, the template filling and prediction module is used to fill the selected samples into the template and input them into the model to complete the multilingual text classification task. Our experimental results verify the effectiveness of our method in mitigating biases related to four sensitive attributes of gender, age, race, and country, demonstrating its potential to improve the fairness and accuracy of LLMs in multilingual classification tasks. Zeyong Long, Dong Zhou 0001, Zhijin Chen, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 6 |
| 2025 | Unraveling the Efficacy of In-Context Learning in Indonesian Grammatical Error CorrectionabstractGrammatical error correction (GEC) is of great importance in natural language processing (NLP). However, due to limited language resources, research on the Indonesian GEC remains scarce. In this paper, we propose an InDonesian In-cOntext-guided grammaticaL Error CorrecTion (IDIOLECT) method, aimed at enhancing the performance of large language models (LLMs) on Indonesian GEC task. Specifically, we calculate sentence similarity to select suitable in-context learning (ICL) demonstrations for each sample in the training set and test set, thereby aiding the model in more effectively identifying and correcting grammatical errors. This study further investigates the effects of ICL configurations, demonstration ordering, and demonstration quantity on model performance. The results indicate that the proposed method effectively improves the performance of LLMs in Indonesian GEC task. Shengyi Jiang, Xuming Li, Nankai Lin, Lixian Xiao, Lianxi Wang 0001 |
CSCWD | 4 |
| 2025 | FairTriplet: Balancing Fairness and Accuracy in Contextual Pre-Trained Models Through Prefix TuningabstractNatural language processing models learn powerful language representation abilities from vast amounts of data, but they also inherit societal biases embedded in that data. Current research on debiasing often struggles to balance the removal of model bias with the preservation of model performance. Most existing approaches depend on fine-tuning model parameters, which can introduce uncertainties in model performance due to the modifications made to these parameters. In this paper, we propose a novel debiasing framework called FairTriplet. First, this framework employs prefix tuning to freeze the parameters of the original pre-trained model. Then, it optimizes the prefix parameters through two debiasing terms. These two debiasing terms function by reducing the semantic distance between social groups (e.g., male and female) and increasing the semantic distance between social groups and neutral attributes (e.g., family and occupation) in the semantic space. This approach not only removes bias from the model but also preserves its performance. Experimental results demonstrate that FairTriplet achieves state-of-the-art (SOTA) levels in debiasing while maintaining model performance on GLUE downstream tasks. Zeyong Long, Weixiong Zheng, Dong Zhou 0001, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 5 |
| 2025 | Enhancing Cross-Lingual Aspect-Based Sentiment Analysis with Code-Mixed In-Context Demonstrations and Language-Specific TagsabstractCross-lingual Aspect-based Sentiment Analysis (XABSA) aims to extract aspect-level sentiments across multiple languages. This task typically relies on source language data to train models and transfer them to target languages, so it faces significant challenges such as data scarcity and language disparities. To this end, this study proposes a code-Mixed In-conteXt lEaRning (MIXER). We design four kinds of demonstration retrieval libraries to introduce Code-mixed In-Context Demonstrations (CICD), which use the code-mixed mechanism to integrate the features of the target language and enrich the target language's knowledge while retaining the source language's knowledge. Language-Specific tags (LST) are introduced to enhance the model's understanding of multilingual demonstrations. To validate the effectiveness of MIXER, we conduct extensive experiments on the SemEval-2016 dataset, comparing its performance against existing XABSA methods. The experimental results show that MIXER performs better than existing XABSA methods with average F1 scores on the Mistral and Llama3 improved by 1.59% and 1.44%, respectively, highlighting its potential for broader multilingual applications. Meiyu Zeng, Xingming Liao, Yongmei Zhou, Nankai Lin, Aimin Yang 0002 |
CSCWD | 5 |
| 2025 | Conditional Independent Test in the Presence of Measurement Error with Causal Structure LearningabstractTesting conditional independence is a critical task, particularly in causal discovery and learning in Bayesian networks. However, in many real-world scenarios, variables are often measured with errors, such as those introduced by insufficient measurement accuracy, complicating the testing process. This paper focuses on testing conditional independence in the linear non-Gaussian measurement error model, under the condition that measurement error noise follows a Gaussian distribution. By leveraging high-order cumulants, we derive rank constraints on the cumulant matrix and establish their role in effectively assessing conditional independence, even in the presence of measurement errors. Based on these theoretical results, we leverage the rank constraints of the cumulant matrix as a tool for conditional independence testing and incorporate it into the PC algorithm, resulting in the PC-ME algorithm — a method designed to learn causal structures from observed data while accounting for measurement errors. Experimental results demonstrate that the proposed method outperforms existing approaches, particularly in cases other methods encounter difficulties. Hongbin Zhang 0008, Kezhou Chen, Nankai Lin, Aimin Yang 0002, Zhifeng Hao 0004, Zhengming Chen 0002 |
IJCAI | 3 |
| 2025 | Filter-enhanced Contrast Variational AutoEncoders for sequential recommendationabstractAbstract Data augmentation-based contrastive learning has been successfully employed in Variational AutoEncoders sequence recommendation systems to tackle the issue of data sparsity. Nevertheless, this strategy is generally less advantageous for tail users. The prospective transmission of information from head-to-tail users to alleviate long-tail impact is encouraging. However, data augmentation distorts the original sequence and embeds stochastic noise into latent variables, impeding the decoder’s capacity to accurately identify the user’s true preferences. In addition, contrastive learning seeks to achieve consistency in the latent variables of both the original and augmented data. However, the presence of noise in the augmented data might hamper the encoding of latent variables from the original data, especially impacting head users. In order to address these challenges, this work introduces a new sequence recommendation model called the Filter-enhanced Contrastive Variational Autoencoder (FeCVAE). It employs Fourier filters and adversarial attack training to minimize the impact of stochastic noise, thereby improving the quality of latent variables and facilitating more accurate decoder outputs. Moreover, a user enhancer is introduced to leverage knowledge from head users to empower tail users, thereby alleviating the long-tail effect. The efficacy of FeCVAE is demonstrated through comprehensive experiments across four benchmark datasets. Zhijin Chen, Nankai Lin, Aimin Yang 0002, Dong Zhou 0001 |
Comput. J. | 2 |
| 2025 | A novel curriculum learning framework for multi-label emotion classificationabstractAbstract Curriculum learning (CL) is a training strategy that imitates how humans learn, by gradually introducing more complex samples and information to the model. However, in multi-label emotion classification (MEC) tasks, using a traditional CL approach can result in overfitting on easy samples and lead to biased training. Additionally, the sample difficulty varies as the model trains. To address these challenges, we propose a novel CL framework for MEC tasks called CLF-MEC. Unlike traditional approaches that assess difficulty at the sample level, we utilize category-level assessment to determine the difficulty level of samples. As the model identifies a category well, the score for that category’s samples is reduced, ensuring dynamic changes in the sample difficulty are accounted for. Our CL framework employs two training modes, namely “learning” and “tackling.” These two processes are trained alternatively to imitate the “learning-tackling” process in human learning. This ensures that samples from hard-to-learn categories receive more attention. During the “tackling” process, our method transforms the task of dealing with hard samples into an “easy” learning task by utilizing contrastive learning to enhance the semantic representation of those hard samples. Experimental results demonstrate that our CLF-MEC framework has achieved significant improvements in MEC. Nankai Lin, Peijian Zeng, Qifeng Bai, Dong Zhou 0001, Aimin Yang 0002 |
Comput. J. | 1 |
| 2025 | Corpus and unsupervised benchmark: Towards Tagalog grammatical error correction
Nankai Lin, Hongbin Zhang 0008, Menglan Shen, Shengyi Jiang, Aimin Yang 0002 |
Comput. Speech Lang. | 1 |
| 2025 | A Chinese Spelling Check Method Based on Reverse Contrastive Learning
Nankai Lin, Sihui Fu, Shengyi Jiang, Aimin Yang 0002 |
J. Comput. Sci. Technol. | 1 |
| 2025 | A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error CorrectionabstractCurrently, the majority of research in grammatical error correction (GEC) is concentrated on universal languages, such as English and Chinese. Many low-resource languages lack accessible evaluation corpora. How to efficiently construct high-quality evaluation corpora for GEC in low-resource languages has become a significant challenge. To fill these gaps, in this article, we present a framework for constructing GEC corpora. Specifically, we focus on Indonesian as our research language and construct an evaluation corpus for Indonesian GEC using the proposed framework, addressing the limitations of existing evaluation corpora in Indonesian. Furthermore, we investigate the feasibility of utilizing existing large language models (LLMs), such as GPT-3.5-Turbo and GPT-4, to streamline corpus annotation efforts in GEC tasks. The results demonstrate significant potential for enhancing the performance of LLMs in low-resource language settings. Our code and corpus can be obtained from https://github.com/GKLMIP/GEC-Construction-Framework . Nankai Lin, Meiyu Zeng, Shengyi Jiang, Lixian Xiao, Aimin Yang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2025 | GS2F: Multimodal Fake News Detection Utilizing Graph Structure and Guided Semantic FusionabstractThe prevalence of fake news online has become a significant societal concern. To combat this, multimodal detection techniques based on images and text have shown promise. Yet, these methods struggle to analyze complex relationships within and between modalities due to the diverse discriminative elements in the news content. In addition, research on multimodal and multi-class fake news detection remains insufficient. To address the above challenges, in this article, we propose a novel detection model, GS 2 F, leveraging g raph s tructure and g uided s emantic f usion. Specifically, we construct a multimodal graph structure to align two modalities and employ graph contrastive learning for refined fusion representations. Furthermore, a guided semantic fusion module is introduced to maximize the utilization of single-modal information and a dynamic contribution assignment layer is designed to weigh the importance of image, text, and multimodal features. Experimental results on Fakeddit demonstrate that our model outperforms existing methods, marking a step forward in the multimodal and multi-class fake news detection. Dong Zhou 0001, Qiang Ouyang, Nankai Lin, Yongmei Zhou, Aimin Yang 0002 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2025 | Ensuring accuracy and fairness: a de-biasing framework for sequential recommendation
Qifeng Bai, Nankai Lin, Meiyu Zeng, Guanqiu Qin, Dong Zhou 0001, Aimin Yang 0002 |
User Model. User Adapt. Interact. | 2 |
| 2025 | A contrastive news recommendation framework based on curriculum learning
Xingran Zhou, Nankai Lin, Weixiong Zheng, Dong Zhou 0001, Aimin Yang 0002 |
User Model. User Adapt. Interact. | 2 |
| 2024 | A Retrieval-Augmented Contrastive Framework for Legal Case Retrieval Based on Event Information
Changyong Fan, Nankai Lin, Dong Zhou 0001, Yongmei Zhou, Aimin Yang 0002 |
ACML | 2 |
| 2024 | MLCL: A Framework for Reducing Language Imbalance in Sino-Tibetan Languages through Adapter Structures
Jiajun Fang, Aimin Yang 0002, Dong Zhou 0001, Nankai Lin |
ACML | 5 |
| 2024 | Enhancing Aspect Sentiment Quad Prediction through Dual-Sequence Data Augmentation and Contrastive Learning
Nankai Lin, Pinmo Wu, Dong Zhou 0001, Aimin Yang 0002 |
ACML | 2 |
| 2024 | HiRAG: A Historical Information-Driven Retrieval-Augmented Generation Framework for Background Summarization
Dong Zhou 0001, Binli Zeng, Nankai Lin, Yongmei Zhou, Aimin Yang 0002 |
ACML | 3 |
| 2024 | GPF: Generative Prediction Fusion for Multi-Label Emotion ClassificationabstractMulti-label Emotion Classification (MEC) is a fundamental and challenging task in natural language processing. The MEC task aims to recognize at least an emotion from a sentence. Previous seq2seq based models required to transform the set of labels into a sequence. However, the labels in the sentence are unordered. In this paper, we propose GPF as a framework of generative prediction fusion for MEC. Specifically, we utilize the non-autoregressive decoder to simultaneously generate the set of labels and propose an output fusion strategy for MEC. Meanwhile, we develop a multi-label contrastive learning to enhance the representation of our model. Experiment results on three distinct language datasets demonstrate the effectiveness of our model. Shiqiao Huang, Nankai Lin, Mianshen Xu |
CSCWD | 3 |
| 2024 | Composited-Nested-Learning with Data Augmentation for Nested Named Entity RecognitionabstractNested Named Entity Recognition (NNER) focuses on addressing overlapped entity recognition. Compared to Flat Named Entity Recognition (FNER), annotated resources are scarce in the corpus for NNER. Data augmentation is an effective approach to address the insufficient annotated corpus. However, there is a significant lack of exploration in data augmentation methods for NNER. Due to the presence of nested entities in NNER, existing data augmentation methods cannot be directly applied to NNER tasks. Therefore, in this work, we focus on data augmentation for NNER and resort to more expressive structures, Composited-Nested-Label Classification (CNLC) in which constituents are combined by nested-word and nested-label, to model nested entities. The dataset is augmented using the Composited-Nested-Learning (CNL). In addition, we propose the Confidence Filtering Mechanism (CFM) for a more efficient selection of generated data. Experimental results demonstrate that this approach results in improvements in ACE2004 and ACE2005 and alleviates the impact of sample imbalance. Xingming Liao, Nankai Lin, Lianglun Cheng, Zhuowei Wang 0001, Chong Chen 0010 |
CSCWD | 2 |
| 2024 | A Vision Enhanced Framework for Indonesian Multimodal Abstractive Text-Image SummarizationabstractMultimodal abstractive summarization (MAS) is a technique that generates a brief summary by processing input text and images. While preceding investigations on MAS have prioritized the utilization of visual features to amplify the quality of summaries, such advancements have predominantly been realized within high-resource languages, most notably English and Chinese. However, in the case of resource-scarce languages like Indonesian, the research and available resources pertaining to multimodal abstractive summarization remains constrained. In addition, the heterogeneity between visual and textual features may impact the quality of summary generation. Therefore, it is crucial to investigate vision-enhanced generative models to improve summary quality. To address the problem of insufficient resources of Indonesian MAS, we constructed the E-Liputan dataset, which is a summary-guided multimodal generative summarization dataset in Indonesian. We employed a two-stage methodology: first, we utilized a mask strategy to effectively address text denoising, thereby facilitating the pre-training of the visual encoder. Second, we fine-tuned an end-to-end multimodal summarization model and proposed a summary-guided multimodal interactive co-attention learning fusion, which facilitated the seamless integration and fusion of modalities within the model. Through the employment of these methodologies, the proposed multimodal model effectively acquired a richer repertoire of visual feature information oriented towards summarization, ultimately leading to enhanced precision and accuracy in generating summaries. We conducted extensive experiments on the E-Liputan dataset and found that our model outperformed the baseline models. Our findings suggest that investigating vision-enhanced generative models for MAS can significantly improve summary quality, particularly in resource-scarce languages. Yutao Song, Nankai Lin, Lingbao Li, Shengyi Jiang |
CSCWD | 2 |
| 2024 | ADSE: Adversarial Debiasing Framework Based on Sinusoidal Embedding for Sequential RecommendationabstractSequential recommendation plays a key role in recommender systems, where the goal is to predict a user’s future points of interest by analyzing his or her historical interactions. This process not only requires the system to be able to accurately identify and recommend items that are likely to be of interest to the user but also ensures that all items receive equal exposure to prevent over-concentration or marginalization of items due to algorithmic bias. To address these challenges, in this paper, we propose a novel Adversarial Debiasing framework based on Sinusoidal Embedding for sequential recommendation, ADSE. This framework employs sinusoidal position embeddings to extract positional information between sequences more precisely and utilizes a dropout strategy to optimize the handling of cold-start sequences, aiming to resolve the cold-start issue while maintaining the semantics of the original sequences. Additionally, adversarial training was incorporated to reduce implicit bias due to assuming interactions in the calculation of exposure. Qifeng Bai, Nankai Lin, Junheng He, Zhijin Chen, Dong Zhou 0001, Aimin Yang 0002 |
ICWS | 2 |
| 2024 | ACTOR: Advancing Argument Components Identification Through In-Context Learning and Proximity Information Awareness
Peijian Zeng, Weixiong Zheng, Nankai Lin, Aimin Yang 0002, Shengyi Jiang |
NLPCC (5) | 4 |
| 2024 | A Chinese Grammatical Error Correction Model Based On Grammatical Generalization And Parameter SharingabstractAbstract Chinese grammatical error correction (CGEC) is a significant challenge in Chinese natural language processing. Deep-learning-based models tend to have tens of millions or even hundreds of millions of parameters since they model the target task as a sequence-to-sequence problem. This may require a vast quantity of annotated corpora for training and parameter tuning. However, there are currently few open-source annotated corpora for the CGEC task; the existing researches mainly concentrate on using data augmentation technology to alleviate the data-hungry problem. In this paper, rather than expanding training data, we propose a competitive CGEC model from a new insight for reducing model parameters. The model contains three main components: a sequence learning module, a grammatical generalization module and a parameter sharing module. Experimental results on two Chinese benchmarks demonstrate that the proposed model could achieve competitive performance over several baselines. Even if the parameter number of our model is reduced by 1/3, it could reach a comparable $F_{0.5}$ value of 30.75%. Furthermore, we utilize English datasets to evaluate the generalization and scalability of the proposed model. This could provide a new feasible research direction for CGEC research. Nankai Lin, Xiaotian Lin, Yingwen Fu, Shengyi Jiang, Lianxi Wang 0001 |
Comput. J. | 1 |
| 2024 | Towards fair decision: A novel representation method for debiasing pre-trained models
Junheng He, Nankai Lin, Qifeng Bai, Dong Zhou 0001, Aimin Yang 0002 |
Decis. Support Syst. | 2 |
| 2024 | Addressing class-imbalance challenges in cross-lingual aspect-based sentiment analysis: Dynamic weighted loss and anti-decoupling
Nankai Lin, Meiyu Zeng, Xingming Liao, Weizhong Liu, Aimin Yang 0002, Dong Zhou 0001 |
Expert Syst. Appl. | 1 |
| 2024 | Global information enhancement and subgraph-level weakly contrastive learning for lightweight weakly supervised document-level event extraction
Guanqiu Qin, Nankai Lin, Menglan Shen, Qifeng Bai, Dong Zhou 0001, Aimin Yang 0002 |
Expert Syst. Appl. | 2 |
| 2023 | Simplifying Aspect-Sentiment Quadruple Prediction with Cartesian Product Operation
Jigang Wang, Aimin Yang 0002, Dong Zhou 0001, Nankai Lin, Weifeng Huang |
ICIC (4) | 4 |
| 2023 | Towards Malay Abbreviation Disambiguation: Corpus and Unsupervised Model
Haoyuan Bu, Nankai Lin, Lianxi Wang 0001, Shengyi Jiang |
NLPCC (2) | 2 |
| 2023 | Towards Malay named entity recognition: an open-source dataset and a multi-task frameworkabstractNamed entity recognition (NER) is a key component of many natural language processing (NLP) applications. The majority of advanced research, however, has not been widely applied to low-resource languages represented by Malay due to the data-hungry problem. In this paper, we present a system for building a Malay NER dataset (MS-NER) of 20,146 sentences through labelled datasets of homologous languages and iterative optimisation. Additionally, we propose a Multi-Task framework, namely MTBR, to integrate boundary information more effectively for NER. Specifically, boundary detection is treated as an auxiliary task and an enhanced Bidirectional Revision module with a gated ignoring mechanism is proposed to undertake conditional label transfer. This can reduce error propagation by the auxiliary task. We conduct extensive experiments on Malay, Indonesian, and English. Experimental results show that MTBR could achieve competitive performance and tends to outperform multiple baselines. The constructed dataset and model would be made available to the public as a new, reliable benchmark for Malay NER. Yingwen Fu, Nankai Lin, Zhihe Yang, Shengyi Jiang |
Connect. Sci. | 2 |
| 2023 | Cross-Lingual Named Entity Recognition for Heterogenous LanguagesabstractPrevious works on cross-lingual Named Entity Recognition (NER) have achieved great success. However, few of them consider the effect of language families between the source and target languages. In this study, we find that the cross-lingual NER performance of a target language would decrease when its source language is changed from the same (homogenous) into a different (heterogenous) language family with that target language. To improve the NER performance in this situation, we propose a novel cross-lingual NER framework based on self-distillation mechanism and Bilateral-Branch Network (SD-BBN). SD-BBN learns source-language NER knowledge from supervised datasets and obtains target-language knowledge from weakly supervised datasets. These two kinds of knowledge are then fused based on self-distillation mechanism for better identifying entities in the target language. We evaluate SD-BBN on 9 language datasets from 4 different language families. Results show that SD-BBN tends to outperform baseline methods. Remarkably, when the target and source languages are heterogenous, SD-BBN can achieve a greater boost. Our results might suggest that obtaining language-specific knowledge from the target language is essential for improving cross-lingual NER when the source and target languages are heterogenous. This finding could provide a novel insight into further research. Yingwen Fu, Nankai Lin, Shengyi Jiang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Self-Training With Double Selectors for Low-Resource Named Entity RecognitionabstractNamed Entity Recognition (NER) is fundamental to multiple downstream natural language processing (NLP) tasks, but most advanced NER methods heavily rely on massive labeled data with high cost. In this paper, we explore the effectiveness of self-training for low-resource NER. It is one of the semi-supervised approaches to reduce the reliance on manual annotation. However, random pseudo sample selection in standard self-training framework may cause serious error propagation, especially for token-level tasks. To that end, this paper focuses on pseudo sample selection and proposes a new self-training framework with double selectors, namely auxiliary judge task and entropy-based confidence measurement. Specifically, the auxiliary judge task is proposed to filter out the pseudo samples with wrong predictions. The entropy-based confidence measurement is introduced to select pseudo samples with high quality. In addition, to make full use of all pseudo samples, we propose a cumulative function based on the idea of curriculum learning to prompt the model to learn from easy samples to hard ones. Samples with low quality are filtered out through the double selectors, which is more conducive to the training of student models. Experimental results on five NER benchmark datasets from different languages indicate the effectiveness of the proposed framework over several advanced baselines. Yingwen Fu, Nankai Lin, Xiaohui Yu 0009, Shengyi Jiang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | CL-XABSA: Contrastive Learning for Cross-Lingual Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA), an extensively researched area in the field of natural language processing (NLP), predicts the sentiment expressed in a text relative to the corresponding aspect. Unfortunately, most languages lack sufficient annotation resources; thus, an increasing number of recent researchers have focused on cross-lingual aspect-based sentiment analysis (XABSA). However, most recent studies focus only on cross-lingual data alignment instead of model alignment. Therefore, we propose a novel framework, CL-XABSA: contrastive learning for cross-lingual aspect-based sentiment analysis. Based on contrastive learning, we close the distance between samples with the same label in different semantic spaces, achieving convergence of semantic spaces of different languages. Specifically, we design two contrastive objectives, token-level contrastive learning of token embeddings (TL-CTE) and sentiment-level contrastive learning of token embeddings (SL-CTE), to unify the semantic space of source and target languages. Since CL-XABSA can receive datasets in multiple languages during training, it can be further extended to multilingual aspect-based sentiment analysis (MABSA). To further improve the model performance, we perform knowledge distillation with target-language unlabeled data. In the distillation XABSA task, we further explore the effectiveness of different data (source dataset, translated dataset, and code-switched dataset). The results demonstrate that the proposed method has a certain improvement in the three XABSA tasks, distillation XABSA and MABSA. The source code of this paper is publicly available athttps://github.com/GKLMIP/CL-XABSA. Nankai Lin, Yingwen Fu, Xiaotian Lin, Dong Zhou 0001, Aimin Yang 0002, Shengyi Jiang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Deps-SAN: Neural Machine Translation with Dependency-Scaled Self-Attention Network
Ru Peng, Nankai Lin, Shengyi Jiang, Tianyong Hao, Junbo Zhao 0002 |
ICONIP (3) | 2 |
| 2022 | LaoPLM: Pre-trained Language Models for LaoabstractTrained on the large corpus, pre-trained language models (PLMs) can capture different levels of concepts in context and hence generate universal language representations. They can benefit from multiple downstream natural language processing (NLP) tasks. Although PTMs have been widely used in most NLP applications, especially for high-resource languages such as English, it is under-represented in Lao NLP research. Previous work on Lao has been hampered by the lack of annotated datasets and the sparsity of language resources. In this work, we construct a text classification dataset to alleviate the resource-scarce situation of the Lao language. In addition, we present the first transformer-based PTMs for Lao with four versions: BERT-Small , BERT-Base , ELECTRA-Small , and ELECTRA-Base . Furthermore, we evaluate them on two downstream tasks: part-of-speech (POS) tagging and text classification. Experiments demonstrate the effectiveness of our Lao models. We release our models and datasets to the community, hoping to facilitate the future development of Lao NLP applications. Nankai Lin, Yingwen Fu, Chuwei Chen, Shengyi Jiang |
LREC | 1 |
| 2022 | A Fine-Grained Social Bias Measurement Framework for Open-Domain Dialogue Systems
Aimin Yang 0002, Qifeng Bai, Jigang Wang, Nankai Lin, Xiaotian Lin, Guanqiu Qin, Junheng He |
NLPCC (2) | 4 |
| 2022 | A simple but effective method for Indonesian automatic text summarisationabstractAutomatic text summarisation (ATS) (therein two main approaches–abstractive summarisation and extractive summarisation are involved) is an automatic procedure for extracting critical information from the text using a specific algorithm or method. Due to the scarcity of corpus, abstractive summarisation achieves poor performance for low-resource language ATS tasks. That’s why it is common for researchers to apply extractive summarisation to low-resource language instead of using abstractive summarisation. As an emerging branch of extraction-based summarisation, methods based on feature analysis quantitate the significance of information by calculating utility scores of each sentence in the article. In this study, we propose a simple but effective extractive method based on the Light Gradient Boosting Machine regression model for Indonesian documents. Four features are extracted, namely PositionScore, TitleScore, the semantic representation similarity between the sentence and the title of document, the semantic representation similarity between the sentence and sentence’s cluster center. We define a formula for calculating the sentence score as the objective function of the linear regression. Considering the characteristics of Indonesian, we use Indonesian lemmatisation technology to improve the calculation of sentence score. The results show that our method is more applicable. Nankai Lin, Shengyi Jiang |
Connect. Sci. | 1 |
| 2022 | Multi-label emotion classification based on adversarial multi-task learning
Nankai Lin, Sihui Fu, Xiaotian Lin, Lianxi Wang 0001 |
Inf. Process. Manag. | 1 |
| 2022 | Unsupervised Character Embedding Correction and Candidate Word DenoisingabstractInthis paper, we take Indonesian as the research object, and propose a multiple filter correction framework (MFCF). The main idea of MFCF is to remove noise from candidate words to increase the probability of correct words being selected. In MFCF, we use window search algorithm (WSA) to filter the candidate words in the dictionary. When searching for candidate words whose Levenshtein distance is 1, WSA reduces the candidate word search space by an average of 71%. When searching for candidate words whose Levenshtein distance is 2, the search space is reduced by an average of 55%. The reduction in search space has brought about an increase in search speed. When WSA searches for candidate words with Levenshtein distance equal to 1 and 2, the speed exceeds the current advanced search algorithm. A character vector-based candidate word scoring model (CWSM-CV) is also introduced in this paper. CWSM-CV is a simple but unsupervised method. In MFCF, we use CWSM-CV to filter the correct word in the candidate word list. Through exploring the feasibility of using word vector-based candidate word scoring model to score candidate words (CWSM-WV), we find the necessity of denoising the candidate word list and verified it with experiments. In order to apply this finding to the text correction, a new set of evaluation indicators are proposed to replace accuracy. Finally, we recommend that researchers who correct text in low-resource languages make the model an open system and publish it for users to use. The system receives user feedback as new data to gradually reduce the negative impact of data volume. Kengtao Zheng, Nankai Lin, Shengyi Jiang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Research on Pseudo-label Technology for Multi-label News Classification
Lianxi Wang 0001, Xiaotian Lin, Nankai Lin |
ICDAR (2) | 3 |
| 2021 | Pre-trained Models and Evaluation Data for the Myanmar Language
Shengyi Jiang, Xiuwen Huang, Xiaonan Cai, Nankai Lin |
ICONIP (6) | 4 |
| 2021 | Pre-trained Language Models for Tagalog with Multi-source Data
Shengyi Jiang, Yingwen Fu, Xiaotian Lin, Nankai Lin |
NLPCC (1) | 4 |
| 2021 | A Framework for Indonesian Grammar Error CorrectionabstractGrammatical Error Correction (GEC) is a challenge in Natural Language Processing research. Although many researchers have been focusing on GEC in universal languages such as English or Chinese, few studies focus on Indonesian, which is a low-resource language. In this article, we proposed a GEC framework that has the potential to be a baseline method for Indonesian GEC tasks. This framework treats GEC as a multi-classification task. It integrates different language embedding models and deep learning models to correct 10 types of Part of Speech (POS) error in Indonesian text. In addition, we constructed an Indonesian corpus that can be utilized as an evaluation dataset for Indonesian GEC research. Our framework was evaluated on this dataset. Results showed that the Long Short-Term Memory model based on word-embedding achieved the best performance. Its overall macro-average F 0.5 in correcting 10 POS error types reached 0.551. Results also showed that the framework can be trained on a low-resource dataset. Nankai Lin, Xiaotian Lin, Kanoksak Wattanachote, Shengyi Jiang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2020 | Multi-domain Sentiment Classification on Self-constructed Indonesian Dataset
Nankai Lin, Sihui Fu, Xiaotian Lin, Shengyi Jiang |
NLPCC (1) | 1 |