VLDB 2026 Research / reviewers in the wild / expert
Xiangping Wu 0001
dblp:52/10131
· DBLP profile ↗
16ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-5267-2250ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FSE: Continual learning for named entity recognition by fast-slow experts
Yunan Zhang 0003, Xiangping Wu 0001, Qingcai Chen |
Pattern Recognit. Lett. | 4 |
| 2025 | Local and Global Aware Document Image Enhancement with Residual Denoising Diffusion ModelabstractIn document image enhancement scenarios, due to the limitations of high computational complexity caused by high-resolution input images, current methods often process these original degraded images by cropping them into patches of specified sizes. However, previous approaches that solely rely on cropped patches or merely use the document enhancement result of the global image as a reference are difficult to fully utilize global image information. This limitation often results in inconsistent enhancement effects across different regions of the same image. In this paper, we introduce LGA-Doc, a novel two-stage local-global information aware generative framework for document image enhancement. Our approach employs a context-aware image feature fusion module that facilitates feature interaction between local document patches and the global image, enabling deep integration of multi-granularity information. The experimental results demonstrate that our method achieves state-of-the-art performance on both the deblurring dataset and the binarization evaluation dataset. Ablation studies further validate the effectiveness of our local-global information aware module. Hongrui Tie, Heng Li 0014, Xiangping Wu 0001, Qingcai Chen |
ICMR | 3 |
| 2025 | CLSurCoder: LLMs Based Cross-Lingual Transfer Learning for Low-Resource Language Surgical Records Coding
Dawen Chu, Fen Yang, Bingchen Zhong, Xiangping Wu 0001, Shuoran Jiang, Qingcai Chen |
NLPCC (2) | 6 |
| 2024 | TDeLTA: A Light-Weight and Robust Table Detection Method Based on Learning Text ArrangementabstractThe diversity of tables makes table detection a great challenge, leading to existing models becoming more tedious and complex. Despite achieving high performance, they often overfit to the table style in training set, and suffer from significant performance degradation when encountering out-of-distribution tables in other domains. To tackle this problem, we start from the essence of the table, which is a set of text arranged in rows and columns. Based on this, we propose a novel, light-weighted and robust Table Detection method based on Learning Text Arrangement, namely TDeLTA. TDeLTA takes the text blocks as input, and then models the arrangement of them with a sequential encoder and an attention module. To locate the tables precisely, we design a text-classification task, classifying the text blocks into 4 categories according to their semantic roles in the tables. Experiments are conducted on both the text blocks parsed from PDF and extracted by open-source OCR tools, respectively. Compared to several state-of-the-art methods, TDeLTA achieves competitive results with only 3.1M model parameters on the large-scale public datasets. Moreover, when faced with the cross-domain data under the 0-shot setting, TDeLTA outperforms baselines by a large margin of nearly 7%, which shows the strong robustness and transferability of the proposed model. Xiangping Wu 0001, Qingcai Chen, Heng Li 0014, Zhixiang Cai, Qitian Wu |
AAAI | 2 |
| 2024 | ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-Order OptimizationabstractLowering the memory requirement in full-parameter training on large models has become a hot research area. MeZO fine-tunes the large language models (LLMs) by just forward passes in a zeroth-order SGD optimizer (ZO-SGD), demonstrating excellent performance with the same GPU memory usage as inference. However, the simulated perturbation stochastic approximation for gradient estimate in MeZO leads to severe oscillations and incurs a substantial time overhead. Moreover, without momentum regularization, MeZO shows severe over-fitting problems. Lastly, the perturbation-irrelevant momentum on ZO-SGD does not improve the convergence rate. This study proposes ZO-AdaMU to resolve the above problems by adapting the simulated perturbation with momentum in its stochastic approximation. Unlike existing adaptive momentum methods, we relocate momentum on simulated perturbation in stochastic gradient approximation. Our convergence analysis and experiments prove this is a better way to improve convergence stability and rate in ZO-SGD. Extensive experiments demonstrate that ZO-AdaMU yields better generalization for LLMs fine-tuning across various NLP tasks than MeZO and its momentum variants. Shuoran Jiang, Qingcai Chen, Youcheng Pan, Yang Xiang 0003, Yukang Lin, Xiangping Wu 0001, Chuanyi Liu, Xiaobao Song |
AAAI | 6 |
| 2024 | Confounder balancing in adversarial domain adaptation for pre-trained large models fine-tuningabstractThe excellent generalization, contextual learning, and emergence abilities in the pre-trained large models (PLMs) handle specific tasks without direct training data, making them the better foundation models in the adversarial domain adaptation (ADA) methods to transfer knowledge learned from the source domain to target domains. However, existing ADA methods fail to account for the confounder properly, which is the root cause of the source data distribution that differs from the target domains. This study proposes a confounder balancing method in adversarial domain adaptation for PLMs fine-tuning (CadaFT), which includes a PLM as the foundation model for a feature extractor, a domain classifier and a confounder classifier, and they are jointly trained with an adversarial loss. This loss is designed to improve the domain-invariant representation learning by diluting the discrimination in the domain classifier. At the same time, the adversarial loss also balances the confounder distribution among source and unmeasured domains in training. Compared to newest ADA methods, CadaFT can correctly identify confounders in domain-invariant features, thereby eliminating the confounder biases in the extracted features from PLMs. The confounder classifier in CadaFT is designed as a plug-and-play and can be applied in the confounder measurable, unmeasurable, or partially measurable environments. Empirical results on natural language processing and computer vision downstream tasks show that CadaFT outperforms the newest GPT-4, LLaMA2, ViT and ADA methods. Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001, Yukang Lin |
Neural Networks | 5 |
| 2024 | BaSFormer: A Balanced Sparsity Regularized Attention Network for TransformerabstractAttention networks often make decisions relying solely on a few pieces of tokens, even if those reliances are not truly indicative of the underlying meaning or intention of the full context. This can lead to over-fitting in transformers and hinder their ability to generalize. Attention regularization and sparsity-based methods have been used to overcome this issue. However, these methods cannot guarantee that all tokens have sufficient receptive fields for global information inference. Thus, the impact of individual biases cannot be effectively reduced. As a result, the generalization of these approaches improved slightly from the training data to new data. To address these limitations, we propose a balanced sparsity (BaS) regularized attention network on top of the transformers, called BaSFormer. BaS regularization introduces the K-regular graph constraint on self-attention connections, which replaces SoftMax with SparseMax in the attention transformation. In BaS-regularized self-attention, SparseMax assigns zero attention scores to low-scoring connections, highlighting influential and meaningful contexts. The K-regular graph constraint ensures that all tokens have an equal-sized receptive field to aggregate information, which facilitates the involvement of global tokens in the feature update of each layer and reduces the impact of individual biases. Given that there is no continuous loss can be used for the K-regular graph regularization, we propose an exponential extremum loss with an augmented Lagrangian function. The experimental results showed that BaSFormer improved the effectiveness of debiasing compared to that of the newest LLMs, such as the GPT-3.5, GPT-4 and LLaMA. In addition, BaSFormer achieves new state-of-the-art (SOTA) results in text generation tasks. Interestingly, this work also shows that BaSFormer can learn hierarchical linguistic dependencies in gradient attributions, which improves interpretability and adversarial robustness. Shuoran Jiang, Qingcai Chen, Yang Xiang 0003, Youcheng Pan, Xiangping Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Learning to Improve Out-of-Distribution Generalization via Self-Adaptive Language MaskingabstractAlthough the pre-trained Transformers learned general linguistic knowledge from large-scale corpus, they still over-fit on the lexical biases when fine-tuning on specific datasets. This problem limits the generalizability of pre-trained models, particularly when learning over out-of-distribution (OOD) data. To address this issue, this paper proposes a self-adaptive language masking (AdaLMask) paradigm to fine-tune the pre-trained Transformers. AdaLMask obviates lexical biases by eliminating the dependence on semantically inessential words. Specifically, AdaLMask learns a Gumbel-Softmax distribution to determine the desired masking positions, and the distribution parameters are optimized via a representation-invariant (RInv) objective to ensure the masked positions are semantically lossless. Four natural language processing tasks are chosen to evaluate the effectiveness of the proposed method on the robustness of lexical biases and OOD generalization. All empirical results demonstrate that the AdaLMask paradigm substantially improves the OOD generalization of pre-trained Transformers. Shuoran Jiang, Youcheng Pan, Qingcai Chen, Yang Xiang 0003, Xiangping Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Foreground and Text-lines Aware Document Image RectificationabstractThis paper aims at the distorted document image rectification problem, the objective to eliminate the geometric distortion in the document images and realize document intelligence. Improving the readability of distorted documents is crucial to effectively extract information from deformed images. According to our observations, the foreground and text-line of the original warped image can represent the deformation tendency. However, previous distorted image rectification methods pay little attention to the readability of the warped paper. In this paper, we focus on the foreground and text-line regions of distorted paper and proposes a global and local fusion method to improve the rectification effect of distorted images and enhance the readability of document images. We introduce cross attention to capture the features of the foreground and text-lines in the warped document and effectively fuse them. The proposed method is evaluated quantitatively and qualitatively on the public DocUNet benchmark and DIR300 Dataset, which achieve state-of-the-art performances. Experimental analysis shows the proposed method can well perform overall geometric rectification of distorted images and effectively improve document readability (using the metrics of Character Error Rate and Edit Distance). The code is available at https://github.com/xiaomore/Document-Image-Dewarping. Heng Li 0014, Xiangping Wu 0001, Qingcai Chen, Qianjin Xiang |
ICCV | 2 |
| 2021 | Decomposing word embedding with the capsule network
Xin Liu 0054, Qingcai Chen, Yan Liu 0004, Joanna Siebert, Baotian Hu, Xiangping Wu 0001, Buzhou Tang |
Knowl. Based Syst. | 6 |
| 2021 | LCSegNet: An Efficient Semantic Segmentation Network for Large-Scale Complex Chinese Character RecognitionabstractComplex scene character recognition is a challenging yet important task in machine learning, especially for languages with large character sets, such as Chinese, which is composed of hieroglyphics with large-scale categories and similar glyphs. Recently, state-of-the-art methods based on semantic segmentation have achieved great success in scene parsing and have been applied in scene text recognition. However, because of limitations in terms of memory and computation, they are only applied in the small category recognition tasks, such as tasks involving English alphabets and digits. In this paper, we propose an efficient semantic segmentation model based on label coding (LC), called LCSegNet, to recognize large-scale Chinese characters. First, to reduce the number of labels, we design a new label coding method based on the Wubi Chinese characters code, called Wubi-CRF. In this method, glyphs and structure information of Chinese characters are encoded into 140-bit labels. Second, we employ an efficient semantic segmentation model for pixel-wise prediction and utilize a conditional random field (CRF) module to learn the constraint rules of Wubi-like coding. Finally, experiments are conducted on three benchmarks: a large Chinese text dataset in the wild (CTW), ICDAR2019-ReCTS, and HIT-OR3C dataset. Results show that the proposed method achieves state-of-the-art performances in both complex scene and handwritten character recognition tasks. Xiangping Wu 0001, Qingcai Chen, Yulun Xiao, Xin Liu 0054, Baotian Hu |
IEEE Trans. Multim. | 1 |
| 2020 | AdaHGNN: Adaptive Hypergraph Neural Networks for Multi-Label Image ClassificationabstractMulti-label image classification is an important and challenging task in computer vision and multimedia fields. Most of the recent works only capture the pair-wise dependencies among multiple labels through statistical co-occurrence information, which cannot model the high-order semantic relations automatically. In this paper, we propose a high-order semantic learning model based on adaptive hypergraph neural networks (AdaHGNN) to boost multi-label classification performance. Firstly, an adaptive hypergraph is constructed by using label embeddings automatically. Secondly, image features are decoupled into feature vectors corresponding to each label, and hypergraph neural networks (HGNN) are employed to correlate these vectors and explore the high-order semantic interactions. In addition, multi-scale learning is used to reduce sensitivity to object size inconsistencies. Experiments are conducted on four benchmarks: MS-COCO, NUS-WIDE, Visual Genome, and Pascal VOC 2007, which cover large, medium, and small-scale categories. State-of-the-art performances are achieved on three of them. Results and analysis demonstrate that the proposed method has the ability to capture high-order semantic dependencies. Xiangping Wu 0001, Qingcai Chen, Yulun Xiao, Baotian Hu |
ACM Multimedia | 1 |
| 2020 | Gated Semantic Difference Based Sentence Semantic Equivalence IdentificationabstractThis article proposes a novel sentence semantic equivalence identification (SSEI) method by using the semantic difference features between sentences. The lexical differences of a sentence pair are first extracted, and the bidirectional long short term memory (BiLSTM) network is then applied on them to generate the semantic difference representations. Finally, an efficient gate mechanism is proposed to integrate the semantic differences with existing models (called base model) to enhance their encoding capability in the SSEI task. Exhaustive experiments conducted on the standard Quora corpus, and the Large-scale Chinese Question Matching Corpus (LCQMC) show that the proposed gated semantic difference (GSD) method brings significant improvement for different existing state-of-the-art models. When the bidirectional encoder representations from transformers model (BERT) is used as the base model, the accuracy for SSEI on Quora is improved from 90.63% to 91.98%, and the F1 score on the LCQMC is improved from 87.0% to 87.7%, which outperforms the best-published results. Xin Liu 0054, Qingcai Chen, Xiangping Wu 0001, Yang Hua 0004, Dongfang Li 0002, Buzhou Tang, Xiaolong Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Stroke Sequence-Dependent Deep Convolutional Neural Network for Online Handwritten Chinese Character RecognitionabstractWe propose a novel model, called stroke sequence-dependent deep convolutional neural network (SSDCNN), which uses the stroke sequence information and eight-directional features of Chinese characters for online handwritten Chinese character recognition (OLHCCR). SSDCNN learns the representation of OLHCCs by incorporating the natural sequence information of the strokes. Furthermore, it naturally incorporates the eight-directional features. First, SSDCNN inputs the stroke sequence and transforms it into stacks of feature maps following the writing order of the strokes. Second, the fixed-length, stroke sequence-dependent representations of OLHCC are derived through convolutional, residual, and max-pooling operations. Third, the stroke sequence-dependent representation is combined with the eight-directional features via a number of fully connected neural network layers. Finally, the Chinese characters are recognized using a softmax classifier. The SSDCNN is trained in two stages: 1) the whole architecture is pretrained using the training data until the performance converges to an acceptable degree. 2) The stroke sequence-dependent representation is combined with the eight-directional features by a fully connected neural network and a softmax layer for further training. The model was experimentally evaluated on the OLHCCR competition tasks of International Conference on Document Analysis and Recognition (ICDAR) 2013. The recognition error was a maximum 58.28% lower in SSDCNN than in a model using the eight-directional features alone (5.13% versus 2.14%). Owing to its high accuracy (97.86%), the proposed SSDCNN reduced the recognition error by approximately 18.0% as compared with that of the winning system in the ICDAR 2013 competition. SSDCNN integrated with an adaptation mechanism, called the SSDCNN+Adapt model, and reached a new state-of-the-art (SOTA) standard with an accuracy of 97.94%. The SSDCNN exploits the stroke sequence information to learn high-quality OLHCC representations. Moreover, the learned representation and the classical eight-directional features complement each other within the SSDCNN architecture. Xin Liu 0054, Baotian Hu, Qingcai Chen, Xiangping Wu 0001, Jinghan You |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2019 | Unconstrained Offline Handwritten Word Recognition by Position Embedding Integrated ResNets ModelabstractThe state-of-the-art methods usually integrate with linguistic knowledge in the recognizer, which makes models more complicated and hard for resource-lacking languages. This letter proposes a new method for unconstrained offline handwritten word recognition by combining position embeddings with residual networks (ResNets) and bidirectional long short-term memory (BiLSTM) networks. At first, ResNets are used to extract abundant features from the input image. Then, position embeddings are used as indices of the character sequence corresponding to a word. By combining the ResNets features with each position embedding, the model generates different inputs for the BiLSTM networks. Finally, the state sequence of the BiLSTM is used to recognize corresponding characters. Without additional language resource, the proposed model achieved the best result on two public corpora, i.e., the 2017 ICDAR word-level information extraction in historical handwritten records competition and the RIMES public dataset on character error rate. Xiangping Wu 0001, Qingcai Chen, Jinghan You, Yulun Xiao |
IEEE Signal Process. Lett. | 1 |
| 2014 | A Short Texts Matching Method Using Shallow Features and Deep Features
Longbiao Kang, Baotian Hu, Xiangping Wu 0001, Qingcai Chen |
NLPCC | 3 |