EDBT 2026 Demo / reviewers in the wild / expert
Lianxi Wang 0001
dblp:23/9624-1 · also Lian-xi Wang 0001
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0002-6746-6476ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Advancing LLMs for Chinese semantic error correction: Example selection and re-scoring
Nankai Lin, Shengyi Jiang, Lianxi Wang 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Feature-level enhanced syntactic-semantic graph networks via optimal transport for aspect-based sentiment analysis
Xinfeng Liao, Xuanqi Chen, Lianxi Wang 0001, Ziying Rong, Jiahuan Yang, Zhuowei Chen |
Knowl. Inf. Syst. | 3 |
| 2025 | CLCO: Enhancing Chinese-Centric Low-Resource Machine Translation Through Domain-Specific Corpus Optimization
Songxi Xu, Lianxi Wang 0001 |
IEEE Big Data | 2 |
| 2025 | Pseudo-label Data Construction Method and Syntax-enhanced Model for Chinese Semantic Error RecognitionabstractChinese Semantic Error Recognition (CSER) has always been a weak link in Chinese language processing due to the complexity and obscureness of Chinese semantics. Existing research has gradually focused on leveraging pre-trained models to perform CSER. Although some researchers have attempted to integrate syntax information into the pre-trained language model, it requires training the models from scratch, which is time-consuming and laborious. Furthermore, despite the existence of datasets for CSER, the constrained size of these datasets impairs the performance of the models. Thus, in order to address the difficulty posed by a limited sample set and the need of annotating samples with semantic-level errors, we propose a Pseudo-label Data Construction method for CSER (PDC-CSER), generating pseudo-labels for augmented samples based on perplexity and model respectively, which overcomes the difficulty of constructing pseudo-label data containing semantic-level errors and ensures the quality of pseudo-labels. Moreover, we propose a CSER method with the Dependency Syntactic Attention mechanism (CSER-DSA) to explicitly infuse dependency syntactic information only in the fine-tuning stage, achieving robust performance, and simultaneously reducing substantial computing power and time cost. Results demonstrate that the pseudo-label technology PDC-CSER and the semantic error recognition method CSER-DSA surpass the existing models Nankai Lin, Shengyi Jiang, Lianxi Wang 0001, Aimin Yang 0002 |
COLING | 4 |
| 2025 | Unraveling the Efficacy of In-Context Learning in Indonesian Grammatical Error CorrectionabstractGrammatical error correction (GEC) is of great importance in natural language processing (NLP). However, due to limited language resources, research on the Indonesian GEC remains scarce. In this paper, we propose an InDonesian In-cOntext-guided grammaticaL Error CorrecTion (IDIOLECT) method, aimed at enhancing the performance of large language models (LLMs) on Indonesian GEC task. Specifically, we calculate sentence similarity to select suitable in-context learning (ICL) demonstrations for each sample in the training set and test set, thereby aiding the model in more effectively identifying and correcting grammatical errors. This study further investigates the effects of ICL configurations, demonstration ordering, and demonstration quantity on model performance. The results indicate that the proposed method effectively improves the performance of LLMs in Indonesian GEC task. Shengyi Jiang, Xuming Li, Nankai Lin, Lixian Xiao, Lianxi Wang 0001 |
CSCWD | 6 |
| 2025 | OTESGN: Optimal Transport-Enhanced Syntactic-Semantic Graph Networks for Aspect-Based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) aims to identify aspect terms and determine their sentiment polarity. While dependency trees combined with contextual semantics provide structural cues, existing approaches often rely on dotproduct similarity and fixed graphs, which limit their ability to capture nonlinear associations and adapt to noisy contexts. To address these limitations, we propose the Optimal TransportEnhanced Syntactic-Semantic Graph Network (OTESGN), a model that jointly integrates structural and distributional signals. Specifically, a Syntactic Graph-Aware Attention module models global dependencies with syntax-guided masking, while a Semantic Optimal Transport Attention module formulates aspect-opinion association as a distribution matching problem solved via the Sinkhorn algorithm. An Adaptive Attention Fusion mechanism balances heterogeneous features, and contrastive regularization enhances robustness. Extensive experiments on three benchmark datasets (Rest14, Laptop14, and Twitter) demonstrate that OTESGN delivers state-of-the-art performance. Notably, it surpasses competitive baselines by up to +1.30 Macro-F1 on Laptop14 and +1.01 on Twitter. Ablation studies and visualization analyses further highlight OTESGN's ability to capture finegrained sentiment associations and suppress noise from irrelevant context. Xinfeng Liao, Xuanqi Chen, Lianxi Wang 0001, Jiahuan Yang, Zhuowei Chen, Ziying Rong |
ICDM | 3 |
| 2024 | Enhancing Hindi Feature Representation through Fusion of Dual-Script Word EmbeddingsabstractPretrained language models excel in various natural language processing tasks but often neglect the integration of different scripts within a language, constraining their ability to capture richer semantic information, such as in Hindi. In this work, we present a dual-script enhanced feature representation method for Hindi. We combine single-script features from Devanagari and Romanized Hindi Roberta using concatenation, addition, cross-attention, and convolutional networks. The experiment results show that using a dual-script approach significantly improves model performance across various tasks. The addition fusion technique excels in sequence generation tasks, while for text classification, the CNN-based dual-script enhanced representation performs best with longer sentences, and the addition fusion technique is more effective for shorter sequences. Our approach shows significant advantages in multiple natural language processing tasks, providing a new perspective on feature representation for Hindi. Our code has been released on https://github.com/JohnnyChanV/Hindi-Fusion. Lianxi Wang 0001, Yujia Tian, Zhuowei Chen |
LREC/COLING | 1 |
| 2024 | An Effective Deployment of Diffusion LM for Data Augmentation in Low-Resource Sentiment ClassificationabstractSentiment classification (SC) often suffers from low-resource challenges such as domainspecific contexts, imbalanced label distributions, and few-shot scenarios.The potential of the diffusion language model (LM) for textual data augmentation (DA) remains unexplored, moreover, textual DA methods struggle to balance the diversity and consistency of new samples.Most DA methods either perform logical modifications or rephrase less important tokens in the original sequence with the language model.In the context of SC, strong emotional tokens could act critically on the sentiment of the whole sequence.Therefore, contrary to rephrasing less important context, we propose DiffusionCLS to leverage a diffusion LM to capture in-domain knowledge and generate pseudo samples by reconstructing strong label-related tokens.This approach ensures a balance between consistency and diversity, avoiding the introduction of noise and augmenting crucial features of datasets.Dif-fusionCLS also comprises a Noise-Resistant Training objective to help the model generalize.Experiments demonstrate the effectiveness of our method in various low-resource scenarios including domain-specific and domain-general problems.Ablation studies confirm the effectiveness of our framework's modules, and visualization studies highlight optimal deployment conditions, reinforcing our conclusions. Zhuowei Chen, Lianxi Wang 0001, Yuben Wu, Xinfeng Liao, Yujia Tian, Junyang Zhong |
EMNLP | 2 |
| 2024 | A Chinese Grammatical Error Correction Model Based On Grammatical Generalization And Parameter SharingabstractAbstract Chinese grammatical error correction (CGEC) is a significant challenge in Chinese natural language processing. Deep-learning-based models tend to have tens of millions or even hundreds of millions of parameters since they model the target task as a sequence-to-sequence problem. This may require a vast quantity of annotated corpora for training and parameter tuning. However, there are currently few open-source annotated corpora for the CGEC task; the existing researches mainly concentrate on using data augmentation technology to alleviate the data-hungry problem. In this paper, rather than expanding training data, we propose a competitive CGEC model from a new insight for reducing model parameters. The model contains three main components: a sequence learning module, a grammatical generalization module and a parameter sharing module. Experimental results on two Chinese benchmarks demonstrate that the proposed model could achieve competitive performance over several baselines. Even if the parameter number of our model is reduced by 1/3, it could reach a comparable $F_{0.5}$ value of 30.75%. Furthermore, we utilize English datasets to evaluate the generalization and scalability of the proposed model. This could provide a new feasible research direction for CGEC research. Nankai Lin, Xiaotian Lin, Yingwen Fu, Shengyi Jiang, Lianxi Wang 0001 |
Comput. J. | 5 |
| 2023 | Towards Malay Abbreviation Disambiguation: Corpus and Unsupervised Model
Haoyuan Bu, Nankai Lin, Lianxi Wang 0001, Shengyi Jiang |
NLPCC (2) | 3 |
| 2023 | Feature selection considering interaction, redundancy and complementarity for outlier detection in categorical data
Lianxi Wang 0001, Yubing Ke |
Knowl. Based Syst. | 1 |
| 2022 | Multi-label emotion classification based on adversarial multi-task learning
Nankai Lin, Sihui Fu, Xiaotian Lin, Lianxi Wang 0001 |
Inf. Process. Manag. | 4 |
| 2021 | Research on Pseudo-label Technology for Multi-label News Classification
Lianxi Wang 0001, Xiaotian Lin, Nankai Lin |
ICDAR (2) | 1 |
| 2021 | A feature selection method via analysis of relevance, redundancy, and interaction
Lianxi Wang 0001, Shengyi Jiang, Siyu Jiang |
Expert Syst. Appl. | 1 |
| 2016 | Efficient feature selection based on correlation measure between continuous and discrete features
Sheng-Yi Jiang, Lianxi Wang 0001 |
Inf. Process. Lett. | 2 |