VLDB 2026 Research / reviewers in the wild / expert
Xiangcheng Du
dblp:252/5303
· DBLP profile ↗
23ranked-venue papers
5as first author
22since 2021 · last 2025
0000-0002-4268-6114ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Exemplar-based Framework for Chinese Text RecognitionabstractThis paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize them by retrieving analogous exemplars from a database. With text-line level annotations, we design the exemplar discovery network to simultaneously recognize texts and capture individual character positions in a weak-supervision manner. The exemplar retrieval module is then crafted to identify the most similar exemplar and propagate the corresponding character label. This enables us to effectively rectify the misrecognized characters and boost the performance of scene text recognition. Experiments on four scenarios of Chinese texts demonstrate the effectiveness of our proposed framework. Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001 |
AAAI | 2 |
| 2025 | Achieving Ensemble-Like Performance in a Single Model: A Feature Diversification Framework for Image-Text MatchingabstractModel ensembling is a widely used technique that enhances performance in image-text matching tasks by combining multiple models, each trained with different initializations. However, the inefficiencies associated with training several models and generating outputs from them constrain their practical applicability. In this paper, we argue that while the parameters of two randomly initialized models can differ significantly, their feature distributions can be similar at certain stages. By employing a proposed technique called cross-modal realignment, we demonstrate that features derived from differently initialized models maintain similarity at the feature extraction stage and can be effectively transformed by fine-tuning a small number of parameters. These findings provide an efficient way to achieve ensemble-like performance within a single model. Specifically, we propose a Feature Diversification Framework (FDF) that emulates the outputs of multiple model initializations to generate diverse features from a common shared feature. Firstly, we introduce feature conversion methods to transform shared features into a set of distinct features. Next, a realignment training strategy is presented to optimize negative pairs for realigning these transformed features, thereby enhancing their diversification to resemble the outputs of different models. Additionally, we propose a reweighting module that assigns weights to these features, enabling a weighted fusion approach for robust feature representation. Extensive experiments on the Flickr30K and MS-COCO datasets demonstrate the effectiveness and generalizability of our framework. Zhao Zhou, Yingbin Zheng, Xiangcheng Du, Cheng Jin 0001 |
AAAI | 5 |
| 2025 | Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided LearningabstractImage-text matching is a crucial task that bridges visual and linguistic modalities. Recent research typically formulates it into the problem of maximizing the margin with the truly hardest negatives to enhance the learning efficiency and avoid the poor local optima. We argue that such formulation can lead to a serious limitation, i.e., under this formulation, conventional trainers would confine their horizon within the hardest negative examples, while other negative examples offer a range of semantic differences not present in the hardest negatives. In this paper, we propose an efficient negative distribution guided training framework for image-text matching to unlock the substantial promotion space left by the above limitation. Rather than simply incorporating additional negative examples into the training objective, which could diminish both the leading role of the hardest negatives in training and the effect of a large margin learning in producing a robust matching model, our central idea is to supply the objective with distributional information on the entire set of negative examples. To be precise, we first construct the sample similarity matrix based on several pretrained models to extract the distributional information of the entire negative sample dataset. Then we encode it into a margin regularization module to smooth the similarities differences of all negatives. This enhancement facilitates the capture of fine-grained semantic differences and guides the main learning process by maximizing the margin with hard negative examples. Furthermore, we propose a hardest negative rectification module to address the instability in hardest negative selection based on predicted similarity and to correct erroneous hardest negatives. We evaluate our method in combination with several state-of-the-art image-text matching methods, and our quantitative and qualitative experiments demonstrate its significant generalizability and effectiveness. Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Cheng Jin 0001 |
AAAI | 3 |
| 2025 | GoLoColor: Towards Global-Local Semantic Aware Image ColorizationabstractOwing to powerful generative priors, Text-to-Image (T2I) diffusion models have achieved promising results in image colorization task. However, recent advanced methods primarily integrate global semantics. Such practice neglects local semantics, yielding suboptimal colorization performance. In this paper, we present a novel global-local semantic aware colorization method named GoLoColor, which performs semantic awareness at both global and local levels. The GoLoColor includes Global Aware (GoA) module, Local Aware module (LoA) and Semantic Aggregation (SA) module for semantic understanding. Specifically, the GoA produces global semantic embedding to represent whole image, while the LoA provides semantic support for local objects, particularly in scenes containing multiple entities. The SA module facilitates semantic interaction between local and global semantic embedding to produce richer semantic information. Finally, a controlled T2I diffusion model is utilized to produce color image guided by the aggregated semantic embedding. Comprehensive experiments demonstrate that our method achieves superior performance and can produce realistic colorization. Tianai Yue, Xiangcheng Du, Jing Liu 0080, Zhongli Fang |
ICASSP | 2 |
| 2025 | Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image ColorizationabstractRecent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enhances the prior while considering the semantic discrepancy between color and grayscale image pairs. The SeAda employs a semantic adapter to produce refined semantic embeddings and a controlled T2I diffusion model to create reasonably colored images. Specifically, the semantic adapter transfers the embedding from grayscale to color domain, while the diffusion model utilizes the refined embedding and prior knowledge to achieve realistic and diverse results. We also design a three-staged training strategy to improve semantic comprehension and prior integration for further performance improvement. Extensive experiments on public datasets demonstrate that our method outperforms existing state-of-the-art techniques, yielding superior performance in image colorization. Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Peizhu Gong, Cheng Jin 0001 |
IJCAI | 1 |
| 2024 | Efficient Scene Text Image Super-Resolution with Semantic GuidanceabstractScene text image super-resolution has significantly improved the accuracy of scene text recognition. However, many existing methods emphasize performance over efficiency and ignore the practical need for lightweight solutions in deployment scenarios. Faced with the issues, our work proposes an efficient framework called SGENet to facilitate deployment on resource-limited platforms. SGENet contains two branches: super-resolution branch and semantic guidance branch. We apply a lightweight pre-trained recognizer as a semantic extractor to enhance the understanding of text information. Meanwhile, we design the visual-semantic alignment module to achieve bidirectional alignment between image features and semantics, resulting in the generation of high-quality prior guidance. We conduct extensive experiments on benchmark dataset, and the proposed SGENet achieves excellent performance with fewer computational costs. LeoWu TomyEnrique, Xiangcheng Du, Kangliang Liu, Zhao Zhou, Cheng Jin 0001 |
ICASSP | 2 |
| 2024 | Fine-Grained Scene Image Classification with Modality-Agnostic AdapterabstractWhen dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In other words, models are deliberately designed based on prior intuitions about the importance of different modalities. In this paper, we present a new multi-modal feature fusion approach named MAA (Modality-Agnostic Adapter), trying to make the model learn the importance of different modalities in different cases adaptively, without giving a prior setting in the model architecture. More specifically, we eliminate the modal differences in distribution and then use a modality-agnostic Transformer encoder for a semantic-level feature fusion. Our experiments demonstrate that MAA achieves state-of-the-art results on benchmarks by applying the same modalities with previous methods. Besides, it is worth mentioning that new modalities can be easily added when using MAA and further boost the performance. Zhao Zhou, Xiangcheng Du, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001 |
ICME | 3 |
| 2024 | Minutes to Seconds: Speeded-up DDPM-based Image Inpainting with Coarse-to-Fine SamplingabstractFor image inpainting, the existing Denoising Diffusion Probabilistic Model (DDPM) based method i.e. RePaint can produce high-quality images for any inpainting form. It utilizes a pre-trained DDPM as a prior and generates inpainting results by conditioning on the reverse diffusion process, namely denoising process. However, this process is significantly time-consuming. In this paper, we propose an efficient DDPM-based image inpainting method which includes three speed-up strategies. First, we utilize a pre-trained Light-Weight Diffusion Model (LWDM) to reduce the number of parameters. Second, we introduce a skip-step sampling scheme of Denoising Diffusion Implicit Models (DDIM) for the denoising process. Finally, we propose Coarse-to-Fine Sampling (CFS), which speeds up inference by reducing image resolution in the coarse stage and decreasing denoising timesteps in the refinement stage. We conduct extensive experiments on both faces and general-purpose image inpainting tasks, and our method achieves competitive performance with approximately 60 times speedup. Xiangcheng Du, LeoWu TomyEnrique, Yingbin Zheng, Cheng Jin 0001 |
ICME | 2 |
| 2024 | MultiColor: Image Colorization by Learning from Multiple Color Spaces
Xiangcheng Du, Zhao Zhou, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001 |
ACM Multimedia | 1 |
| 2024 | Refined and Locality-Enhanced Feature for Handwritten Mathematical Expression Recognition
Xiangcheng Du, Ziang Liu 0019, Daoguo Dong, Liang He 0001 |
PRCV (7) | 2 |
| 2024 | Cross-domain document layout analysis using document style guide
Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Yingbin Zheng, Xin Li 0110, Tianlong Ma, Cheng Jin 0001, Liang He 0001 |
Expert Syst. Appl. | 3 |
| 2023 | DDT: Dual-branch Deformable Transformer for Image DenoisingabstractTransformer is beneficial for image denoising tasks since it can model long-range dependencies to overcome the limitations presented by inductive convolutional biases. However, directly applying the transformer structure to remove noise is challenging because its complexity grows quadratically with the spatial resolution. In this paper, we propose an efficient Dual-branch Deformable Transformer (DDT) denoising network which captures both local and global interactions in parallel. We divide features with a fixed patch size and a fixed number of patches in local and global branches, respectively. In addition, we apply deformable attention operation in both branches, which helps the network focus on more important regions and further reduces computational complexity. We conduct extensive experiments on real-world and synthetic denoising tasks, and the proposed DDT achieves state-of-the-art performance with significantly fewer computational costs. Kangliang Liu, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001 |
ICME | 2 |
| 2023 | Image Layer Modeling for Complex Document Layout GenerationabstractDocument layout analysis (DLA) plays an essential role in information extraction and document understanding. At present, DLA has reached the milestone achievement; however, DLA of non-Manhattan is still challenging because of annotation data limitations. In this paper, we propose an image layer modeling method to mitigate this issue. The image layer modeling method generates document images of non-Manhattan layouts by superimposing images under pre-defined aesthetic rules. Due to the lack of evaluation benchmark for non-Manhattan layout, we have constructed a manually-labeled non-Manhattan layout fine-grained segmentation dataset. To the best of our knowledge, this is the first manually-labeled non-Manhattan layout fine-grained segmentation dataset. Extensive experimental results verify that our proposed image layer modeling method can better deal with the fine-grained segmented document of the non-Manhattan layout. Tianlong Ma, Xingjiao Wu, Xiangcheng Du, Cheng Jin 0001 |
ICME | 3 |
| 2023 | Modeling Stroke Mask for End-to-End Text ErasingabstractScene text erasing aims to wipe text regions in scene images with reasonable background. Most previous approaches employ scene text detectors to assist localization of the text regions. However, detected text boxes contain both text strokes and background clutters, and directly in-painting on the whole boxes may remain text artifacts and make regions unnatural. In this paper, we present an end-to-end network that focuses on modeling text stroke masks that provide more accurate locations to compute erased images. The network consists of two stages, i.e., a basic network with stroke generation and a refinement network with stroke awareness. The basic network predicts the text stroke masks and initial erasing results simultaneously. The refinement network receives the masks as supervision to generate natural erased results. Experiments on both synthetic and real-world scene images demonstrate the effectiveness of our framework in producing high quality erasing results. Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Tianlong Ma, Xingjiao Wu, Cheng Jin 0001 |
WACV | 1 |
| 2023 | Progressive scene text erasing with self-supervision
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Tianlong Ma, Cheng Jin 0001 |
Comput. Vis. Image Underst. | 1 |
| 2023 | DRFN: A unified framework for complex document layout analysis
Xingjiao Wu, Tianlong Ma, Xiangcheng Du, Ziling Hu, Jing Yang 0023, Liang He 0001 |
Inf. Process. Manag. | 3 |
| 2023 | Reading Scene Text with Aggregated Temporal Convolutional EncoderabstractReading scene text in the natural image is of fundamental importance in many real-world problems. Text recognition has a profound effect on information processing by enabling automated extraction and interpretation. Recent scene text recognition methods employ the encoder-decoder framework, which constructs the encoder by obtaining the visual representations based on the last layer of the backbone network and then feeding them into a sequence model. In this article, we propose a novel encoder structure that performs the feature extractor and the sequence modeling within a unified framework. The introduced Aggregated Temporal Convolutional Encoder (ATCE) first incorporates the temporal convolutional layers to consider the long-term temporal relationship in the encoder stage. The aggregation of these temporal convolution modules is designed to utilize visual features from different levels, by augmenting the standard architecture with deeper aggregation to better fuse information across modules. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on several scene text recognition benchmarks for both Chinese and English; the experiments demonstrate the complementary ability with different decoder variants and the effectiveness of our proposed approach. Tianlong Ma, Xiangcheng Du, Xingjiao Wu, Zhao Zhou, Yingbin Zheng, Cheng Jin 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Scene Text Recognition with Heuristic Local AttentionabstractScene text recognition is considered as a sequence labeling problem. For the text recognition task, the alignment between the scene text image and the output text is coincident, which means the latter characters corresponding to the image region will also be behind. However, the existing global attention-based method focuses too much irrelevant information which leads to alignment drift. Contrary, local attention selects the subset of feature representation most relevant to the current character. In this paper, we explore the local attention mechanism and attempt to replace the global attention to implement decoding. Therefore, we revise several variants of local attention methods and provide a comprehensive comparison, which is missing in the scene text recognition literature so far. Specially, we introduce two Heuristic approaches for Local Attention (HLA) and prove that monotonic alignment improves performance significantly. Evaluations on the benchmarks show that the local attention method outperforms the existing global attention methods. Tianlong Ma, Xiangcheng Du, Xiutao Cui |
IEEE Big Data | 2 |
| 2022 | 3D Clues Guided Convolution for Depth CompletionabstractDepth completion is a task that recovers a dense depth map from a sparse depth map with the corresponding color image. Recently, the intensive depth generation guided by image clues in the color map has achieved good results. Color images can provide structural and semantic information as guidance information, but cannot provide the more important information about geometric relationships. In this paper, we propose a novel network to learn latent 3D cues from RGB images and depth images. More specifically, the network contains a 3D clues extractor and a dense depth generator. The extractor is designed to fusion and extract the 3D joint clues from the color image and sparse depth. The generator is trained with the sparse depth map and 3D clues to producing a more accurate dense depth map. Extensive experiments show that our proposed method has a significant improvement over existing image-guided methods. Zhichao Fu, Xingjiao Wu, Xiangcheng Du, Tianlong Ma, Liang He 0001 |
ICIP | 4 |
| 2022 | Document Layout Analysis Via Positional EncodingabstractDocument layout analysis plays a vital role in computer vision research. Current document layout analysis methods mostly use pixel-based classification for document layout analysis. However, the method based on pixel classification is insufficient for maintaining the continuity of the classification area. In this paper, we propose a document layout analysis method based on positional encoding and bounding box specification. We maintain the continuity of the analysis area by constructing a document layout analysis framework based on the bounding box. In addition, we also integrate a positional encoding module in the framework to maintain the detailed information in the document layout analysis and modeling process. Experimental results prove that our proposed method has achieved state-of-the-art results. Ejian Zhou, Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Tianlong Ma, Liang He 0001 |
ICIP | 4 |
| 2021 | Unsupervised Learning Boost Person Re-identification and Real World ApplicationabstractPerson re-identification (Re-ID) is a retrieval problem based on computer vision, playing an important role in surveillance applications where we tried to identify the same person among surveillance photographs. At present, most person re-identification technologies and methods are based on convolutional neural networks (CNNs). Vision Transformers are merged recently and tend to displace pure CNNs in various computer vision tasks. In this paper, we first try ResNet to accomplish the task, with MGN and other tricks to improve the precision, then we explore the Swin Transformer, a pure transformer-based model. We make a large scale unlabeled dataset in which people acting various activities by cutting pictures in videos from YouTube, DINO and MoBY are employed for ResNet and Swin Transformer separately to perform unsupervised pre-training for improving the generalization ability of the learned person re-identification feature representation. Finally, We also make a dataset based on border inspection BoderCheck by using DBSCAN to cluster unlabeled pictures with several adaptations, a multi-datasets training method and a data augmentation are proposed to tackle the half-length vs full-body matching problem. Extensive experiments indicate that our method achieves state-of-the-art results on several mainstream benchmarks and behave well on BoderCheck dataset. Yangsheng Lin, Xiangcheng Du, Yining Lin |
IEEE BigData | 3 |
| 2021 | Document Layout Analysis via Dynamic Residual Feature FusionabstractThe document layout analysis (DLA) aims to split the document image into different interest regions and understand the role of each region, which has wide application such as optical character recognition (OCR) systems and document retrieval. However, it is a challenge to build a DLA system because the training data is very limited and lacks an efficient model. In this paper, we propose an end-to-end united network named Dynamic Residual Fusion Network (DRFN) for the DLA task. Specifically, we design a dynamic residual feature fusion module which can fully utilize low-dimensional information and maintain high-dimensional category information. Besides, to deal with the model overfitting problem that is caused by lacking enough data, we propose the dynamic select mechanism for efficient fine-tuning in limited train data. We experiment with two challenging datasets and demonstrate the effectiveness of the proposed module. Xingjiao Wu, Ziling Hu, Xiangcheng Du, Jing Yang 0023, Liang He 0001 |
ICME | 3 |
| 2020 | Scene Text Recognition with Temporal Convolutional EncoderabstractTexts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label sequence. In this paper, we study text recognition framework by considering the long-term temporal dependencies in the encoder stage. We demonstrate that the proposed Temporal Convolutional Encoder with increased sequential extents improves the accuracy of text recognition. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on seven datasets and the experiments demonstrate the effectiveness of our proposed approach. Xiangcheng Du, Tianlong Ma, Yingbin Zheng, Hao Ye 0005, Xingjiao Wu, Liang He 0001 |
ICASSP | 1 |