EDBT 2026 Demo / reviewers in the wild / expert
Xiaojun Bi 0002
dblp:19/3851-2 · also Xiao-Jun Bi 0002
· DBLP profile ↗
43ranked-venue papers
16as first author
33since 2021 · last 2026
0000-0002-5382-1000ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 14 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCC3: a novel structure-connected cognition cube network for Manchu word recognition
Xiaojun Bi 0002, Wenhao Tao, Haipeng Sun |
Expert Syst. Appl. | 1 |
| 2026 | Lightweight Dongba character recognition: A novel and simple baseline
Zheng Chen 0017, Jianyu Yue, Xiaojun Bi 0002 |
Expert Syst. Appl. | 3 |
| 2026 | A weakly supervised preference alignment framework for robust ancient Chinese translation
Xiaojun Bi 0002, Junyao Xing |
Pattern Recognit. | 1 |
| 2026 | RGCA: Recognition-Guided Context-Aware Network for Manchu word inpainting
Xiaojun Bi 0002, Wenhao Tao |
Pattern Recognit. | 1 |
| 2026 | TIMTQE: Benchmarking Machine Translation Quality Estimation for Text ImagesabstractAccurate quality estimation is critical for the reliable machine translation of text from historical document images. However, existing quality estimation methods typically presuppose clean source text, failing to account for the visual noise and optical character recognition errors prevalent in authentic documents. This challenge is compounded by the lack of standardized benchmarks. To address these issues, we introduce TIMTQE, a novel multimodal and multilingual benchmark for quality estimation on text images. TIMTQE comprises a human-annotated corpus of historical documents and a large-scale synthetic dataset, designed for reference-free quality prediction directly from noisy images. We evaluate modern multimodal large language models and demonstrate that end-to-end models achieve substantially greater robustness on high-noise images compared to conventional cascaded systems that rely on separate optical character recognition and quality estimation modules. The dataset and resources are publicly available athttps://github.com/thinklis/TIMTQE. Xiaojun Bi 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveabstractVideo text-based visual question answering (Video TextVQA) aims to answer questions by explicitly reading and reasoning about the text involved in a video. Most works in this field follow a frame-level framework which suffers from redundant text entities and implicit relation modeling, resulting in limitations in both accuracy and efficiency. In this paper, we rethink the Video TextVQA task from an instance-oriented perspective and propose a novel model termed GAT (Gather and Trace). First, to obtain accurate reading result for each video text instance, a context-aggregated instance gathering module is designed to integrate the visual appearance, layout characteristics, and textual contents of the related entities into a unified textual representation. Then, to capture dynamic evolution of text in the video flow, an instance-focused trajectory tracing module is utilized to establish spatio-temporal relationships between instances and infer the final answer. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. GAT outperforms existing Video TextVQA methods, video-language pretraining methods, and video large language models in both accuracy and inference speed. Notably, GAT surpasses the previous state-of-the-art Video TextVQA methods by 3.86% in accuracy and achieves ten times of faster inference speed than video large language models. The source code is available at https://github.com/zhangyan-ucas/GAT. Gangyan Zeng, Daiqing Wu, Huawen Shen, Binbin Li 0003, Yu Zhou 0015, Can Ma, Xiaojun Bi 0002 |
ACM Multimedia | 8 |
| 2025 | Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionabstractRemoving various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a cumbersome and highly complex document processing system. Although recent studies attempt to unify multiple tasks, they often suffer from limited scalability due to handcrafted prompts and heavy preprocessing, and fail to fully exploit inter-task synergy within a shared architecture. To address the aforementioned challenges, we propose Uni-DocDiff, a Unified and highly scalable Doc ument restoration model based on Dif fusion. Uni-DocDiff develops a learnable task prompt design, ensuring exceptional scalability across diverse tasks. To further enhance its multi-task capabilities and address potential task interference, we devise a novel Prior Pool, a simple yet comprehensive mechanism that combines both local high-frequency features and global low-frequency features. Additionally, we design the Prior Fusion Module (PFM), which enables the model to adaptively select the most relevant prior information for each specific task. Extensive experiments show that the versatile Uni-DocDiff achieves performance comparable or even superior performance compared with task-specific expert models, and simultaneously holds the task scalability for seamless adaptation to new tasks. Fangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang, Binbin Li 0003, Xiaojun Bi 0002, Yu Zhou 0015 |
ACM Multimedia | 6 |
| 2025 | ShuiAttNet: Fully convolutional attention network for Shuishu character recognition
Xiaojun Bi 0002, Weizheng Qiao |
Expert Syst. Appl. | 1 |
| 2025 | Lightweight Text-VQA: Superior trade-off between performance and capacity
Xiaojun Bi 0002, Jianyu Yue, Zheng Chen 0017 |
Expert Syst. Appl. | 1 |
| 2025 | MFNet: Multiscale fusion network for Dongba character image inpainting
Yanlong Luo, Xiaojun Bi 0002 |
Expert Syst. Appl. | 2 |
| 2025 | LCFFNet: A Lightweight Cross-scale Feature Fusion Network for human pose estimation
Xuelian Zou, Xiaojun Bi 0002 |
Neural Networks | 2 |
| 2025 | HyperSL: A Spectral Foundation Model for Hyperspectral Image InterpretationabstractThis work has delivered a novelty foundational model for hyperspectral remote sensing images. Current approaches for hyperspectral data interpretation often require specialized models that are specifically tailored to individual datasets or tasks. In some specific tasks, the availability of hyperspectral data is often limited, posing significant challenges to training due to data insufficiency. Furthermore, the diverse structure of hyperspectral data often complicates the transfer of knowledge from other available datasets, severely limiting cross-scenario capabilities. To bridge this gap, we introduce a highly adaptable foundational model with strong transferability, capable of processing all forms of hyperspectral data across diverse spectral bands and ranges. Compared to previous methods, our approach: 1) standardizes all types spectral vectors into a common token format, enabling a single model for multi-source hyperspectral data; 2) aligns spectral features across different ranges by embedding wavelength information into position encoding; 3) has been pre-trained on over 300 million spectral instances worldwide, ensuring broad generalization; 4) transfers learned knowledge effortlessly to downstream tasks and new datasets without architectural modifications or training from scratch. Experimental results demonstrate that, compared to other mainstream methods, our approach achieves state-of-the-art classification performance across various datasets with different spectral characteristics in both supervised and unsupervised learning settings, while also delivering impressive results in change detection tasks. The source code and the pretrained weights are available at https://github.com/kkweil/HyperSL. Baisen Liu, Xiaojun Bi 0002, Changdong Yu, Yushi Chen 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Leveraging Multiple Source Cities in Selective Transfer Learning for Traffic Prediction With Limited DataabstractTraffic prediction with limited data becomes increasingly momentous and attracts a lot of attention because the urban data scarcity problem is common and often leads to low prediction precision in the practical application. Cross-city transfer learning based on deep learning can effectively alleviate the above problem by transferring knowledge data-rich source cities to data-poor target cities. Recently, a selectively cross-city method is proposed and achieves state-of-the-art precision. Nonetheless, its knowledge transfer is only suitable for utilizing one source city. The diversity of knowledge from just one source city is usually inadequate for effective transferring. To address this problem, this paper proposesCross-Multiple-Source-cities selective transfer learning viaVirtualCity (CMSVC) for traffic prediction with limited data that can effectively exploit knowledge of multiple source cities. We propose a novel virtual city mechanism to integrate beneficial regions for the target city from multiple source cities. To implement this mechanism, we compare the time-series and geographic information of the source and target cities by adopting appropriate similarity metrics. Additionally, we adopt the depth-first search algorithm to extract areas so as to maintain the geographical adjacency relationship. After the virtual city mechanism, these beneficial regions are fused and can be the input for the graph neural network and meta-learning mechanism to obtain weights for selective learning. We evaluate our method on four traffic real-world datasets. The extensive experimental results demonstrate that CMSVC outperforms the state-of-the-art method. The source code of CMSVC is available athttps://github.com/pku-smart-city/source_code/tree/main/CMSVC Xiaojun Bi 0002, Wudi Li, Siyuan Qi |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Omni-scale feature learning for lightweight image dehazing
Zheng Chen 0017, Xiaojun Bi 0002, Jianyu Yue |
Appl. Intell. | 2 |
| 2024 | Appearance-Pose Joint Coordinates Information Collaboration Model for clothes-changing person re-identification
Xiaojun Bi 0002 |
Expert Syst. Appl. | 1 |
| 2024 | Oscar: Omni-scale robust contrastive learning for Text-VQA
Jianyu Yue, Xiaojun Bi 0002, Zheng Chen 0017 |
Expert Syst. Appl. | 2 |
| 2024 | Heterogeneous-branch integration framework: Introducing first-order predicate logic in Logical Reasoning Question Answering
Jianyu Yue, Xiaojun Bi 0002, Zheng Chen 0017 |
Neurocomputing | 2 |
| 2024 | A Universal Representation Mechanism for Multisource Remote Sensing ImageabstractDeep learning-based processing of hyperspectral remote sensing images (HSIs) has emerged as a research hotspot. In recent years, the delivery of numerous novel HSI acquisition platforms has resulted in an exponential growth in the number of available HSI datasets from various sources. However, due to variations among hyperspectral sensors, the acquired data frequently consists of various spectral dimensions. This results in a challenge since standard deep learning approaches often require a different model for each HSI source, which impedes the construction of fundamental models for HSIs. To address this issue, we propose a unified representation mechanism for multisource HSIs that can transform spectra from numerous dimensions to a shared representation space, yielding a scalable pretraining model. Compared to existing methods, our strategy has the following advantages: 1) compatibility with HSIs of arbitrary spectral dimensions, ranges, and resolutions; 2) fully leveraging existing multisource HSIs; 3) spontaneously capturing spectral features via self-supervised learning; and 4) pretrained on large-scale multisource HSI datasets and a considerable enhancement in classification accuracy. In three reconstruction test sets, the PSNRs are 27.07, 22.27, and 29.35 dB, and the SSIMs are 0.93, 0.82, and 0.88, respectively. Compared with the randomly initialized model, there are 8.28% and 1.4% improvements on the Indian Pines dataset and Pavia University dataset respectively. Baisen Liu, Xiaojun Bi 0002, Yihan He |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Stronger Heterogeneous Feature Learning for Visible-Infrared Person Re-IdentificationabstractAbstract Visible-Infrared person re-identification (VI-ReID) is of great importance in the field of intelligent surveillance. It enables re-identification of pedestrians between daytime and dark scenarios, which can help police find escaped criminals at night. Currently, existing methods suffer from inadequate utilisation of cross-modality information, missing modality-specific discriminative information and weaknesses in perceiving differences between different modalities. To solve the above problems, we innovatively propose a stronger heterogeneous feature learning (SHFL) method for VI-ReID. First, we innovatively propose a Cross-Modality Group-wise constraint to solve the problem of inadequate utilization of cross-modality information. Secondly, we innovatively propose a Second-Order Homogeneous Invariant Regularizer to address the problem that missing modality-specific discriminative information. Finally, we innovatively propose a Modality-Aware Batch Normalization to address the problem of weaknesses in perceiving differences between different modalities. Extensive experimental results on two generic VI-ReID datasets demonstrate that the proposed final method outperforms the state-of-the-art algorithms. Xiaojun Bi 0002, Changdong Yu |
Neural Process. Lett. | 2 |
| 2024 | Hybrid ViT-CNN Network for Fine-Grained Image ClassificationabstractIn recent years, vision transformer (ViT) has achieved remarkable breakthroughs in fine-grained visual classification (FGVC) because of its self-attention mechanism that excels in extracting distinctive features from different pixels. However, pure ViT falls short in capturing the crucial multi-scale, local, and low-layer features that hold significance for FGVC. To compensate for these shortcomings, a new hybrid network called HVCNet is designed, which fuses the advantages of ViT and convolutional neural networks (CNN). The three modifications in the original ViT are: 1) using a multi-scale image-to-tokens (MIT) module instead of directly tokenizing the raw input image, thus enabling the network to capture the features at different scales; 2) substituting feed-forward network in ViT's encoder with mixed convolution feed-forward (MCF) module, which enhances the capability of the network in capturing the local and multi-scale features; 3) designing multi-layer feature selection (MFS) module to address the issue of deep-layer tokens in ViT to avoid ignoring the local and low-layer features. The experiment results indicate that the proposed method surpasses state-of-the-art methods on publicly datasets. Ran Shao, Xiaojun Bi 0002, Zheng Chen 0017 |
IEEE Signal Process. Lett. | 2 |
| 2024 | Information Dropping Data Augmentation for Machine Translation Quality EstimationabstractMachine translation quality estimation (QE) refers to the quality assessment of machine translations without a given reference translation. Supervised QE models based on neural networks have achieved state-of-the-art results. But this method requires large-scale training data, which requires bilingual experts to create high-quality labels. This is often very costly. Therefore, we propose a sentence-level machine translation QE data augmentation method based on information dropping. Firstly, we calculate the subwords information of the target translation based on the conditional language model. Subsequently, some subwords in the target translation are randomly deleted or replaced. We obtain the pseudo quality score by calculating the remaining information. Finally, the original and augmented data are combined to train the final model. This pseudo-data generation method based on information dropping strategy enables us to obtain more faithful and diverse training samples without requiring additional corpus resources. Experimental results show that we improve the correlation with human judgment by an average of 5.96% in the seven translation directions of the MLQE-PE dataset, while improving the model's robustness to low adequacy samples. In addition, the method does not require any modifications to the model architecture. Xiaojun Bi 0002, Tao Liu 0038, Zheng Chen 0017 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | Model-Guided Generative Adversarial Networks for Unsupervised Fine-Grained Image GenerationabstractUnsupervised fine-grained image generation is a challenging issue in computer vision. Although many recent significant advances have improved performance, the ability to synthesize photo-realistic images in an unsupervised manner remains extremely difficult. The existing methods compose an image via complex three-stage generative adversarial networks and impose constraints between the latent codes. This pipeline focuses on the disentanglement and ignores the quality of generated images. In this article, we propose a novel two-stage approach for unsupervised fine-grained image generation, termed Model-Guided Generative Adversarial Networks (MG-GAN). We introduce an attention module for exploring the correlation between fine-grained latent codes and image features in the foreground generation stage. The attention module enables the network to automatically focus on the color details and semantic concepts of objects related to different fine-grained classes. Furthermore, we incorporate knowledge distillation strategy and design a simple but effective inverse background image generator as a teacher to guide the background image generation. With the help of knowledge learned in the pre-trained inverse background image generator, a comfortable canvas is synthesized and combined with foreground object more reasonably. Extensive experiments on three popularly fine-grained datasets demonstrate that our approach achieves state-of-the-art performance and is even competitive with semi-supervised method. Jian Xiao 0009, Xiaojun Bi 0002 |
IEEE Trans. Multim. | 2 |
| 2023 | Multiple attentional aggregation network for handwritten Dongba character recognition
Yanlong Luo, Xiaojun Bi 0002 |
Expert Syst. Appl. | 3 |
| 2023 | Structural redundancy reduction based efficient training for lightweight person re-identification
Xiaojun Bi 0002 |
Inf. Sci. | 3 |
| 2023 | Lightweight image de-snowing: A better trade-off between network capacity and performance
Zheng Chen 0017, Xiaojun Bi 0002, Jianyu Yue |
Neural Networks | 3 |
| 2023 | Word self-update contrastive adversarial networks for text-to-image synthesis
Jian Xiao 0009, Xiaojun Bi 0002 |
Neural Networks | 3 |
| 2023 | Improving Human Pose Estimation Based on Stacked Hourglass Network
Xuelian Zou, Xiaojun Bi 0002, Changdong Yu |
Neural Process. Lett. | 2 |
| 2023 | Retrospective Multi-granularity Fusion Network for Chinese Idiom Cloze-style Reading ComprehensionabstractChinese idiom cloze-style reading comprehension task is of great significance for improving the machine’s ability to understand Chinese idioms, which is one of the essential application requirements in advanced artificial intelligence. Existing methods suffer from an insufficient deep semantic understanding of the text. To solve this problem, this paper proposes a novelRetrospective Multi-granularity Fusion Network (RMFNet)for Chinese idiom cloze-style reading comprehension. Our RMFNet is equipped with two novel modules to model deeper contextual information of passage and Chinese idioms, respectively. First, we propose a novelMulti-granularity Passage Fusion (MgPF)module, which enhances the passage representation by integrating different semantic perspectives. Second, we propose aRetrospective Reading (Re \(^2\) )module that implements a back-and-forth reading mechanism to concentrate on critical Chinese idioms, thereby generating an ultimate memory for the whole text. Notably, the intuition of the MgPF module and the Re \(^2\) module is based on human reading strategies in the real world. The strategies in these modules are similar to how humans perceive the text. Extensive experiments are conducted on Chinese benchmark datasets to evaluate the effectiveness and superiority of the proposed method. Our RMFNet achieves state-of-the-art performance and in-depth analysis verifies its capability for understanding the deep semantics of the text. Jianyu Yue, Xiaojun Bi 0002, Zheng Chen 0017, Yu Zhang 0038 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | Multi-task Learning with Auxiliary Cross-attention Transformer for Low-Resource Multi-dialect Speech Recognition
Zhengjia Dan, Yue Zhao 0013, Xiaojun Bi 0002, Licheng Wu |
NLPCC (1) | 3 |
| 2022 | LRP-net: A lightweight recursive pyramid network for single image deraining
Xiaojun Bi 0002, Zheng Chen 0017, Jianyu Yue |
Neurocomputing | 1 |
| 2022 | LightweightDeRain: learning a lightweight multi-scale high-order feedback network for single image de-raining
Zheng Chen 0017, Xiaojun Bi 0002, Yu Zhang 0038, Jianyu Yue |
Neural Comput. Appl. | 2 |
| 2022 | Domain generalization and adaptation based on second-order style information
Xiaojun Bi 0002 |
Pattern Recognit. | 2 |
| 2021 | Person Re-Identification Based on Graph Relation Learning
Xiaojun Bi 0002 |
Neural Process. Lett. | 2 |
| 2019 | An enhanced high-order Boltzmann machine for feature engineering
Xiaojun Bi 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2019 | Early Alzheimer's disease diagnosis based on EEG spectral images using deep learning
Xiaojun Bi 0002 |
Neural Networks | 1 |
| 2019 | Contractive Slab and Spike Convolutional Deep Belief Network
Xiaojun Bi 0002 |
Neural Process. Lett. | 2 |
| 2018 | A niche-elimination operation based NSGA-III algorithm for many-objective optimization
Xiaojun Bi 0002 |
Appl. Intell. | 1 |
| 2018 | Contractive Slab and Spike Convolutional Deep Boltzmann Machine
Xiaojun Bi 0002 |
Neurocomputing | 1 |
| 2018 | A many-objective evolutionary algorithm based on hyperplane projection and penalty distance selection
Xiaojun Bi 0002 |
Nat. Comput. | 1 |
| 2018 | A Light Dual-Task Neural Network for Haze RemovalabstractSingle-image dehazing is a challenging problem due to its ill-posed nature. Existing methods rely on a suboptimal two-step approach, where an intermediate product like a depth map is estimated, based on which the haze-free image is subsequently generated using an artificial prior formula. In this paper, we propose a light dual-task Neural Network called LDTNet that restores the haze-free image in one shot. We use transmission map estimation as an auxiliary task to assist the main task, haze removal, in feature extraction and to enhance the generalization of the network. In LDTNet, the haze-free image and the transmission map are produced simultaneously. As a result, the artificial prior is reduced to the smallest extent. Extensive experiments demonstrate that our algorithm achieves superior performance against the state-of-the-art methods on both synthetic and real-world images. Yu Zhang 0038, Xinchao Wang, Xiaojun Bi 0002, Dacheng Tao |
IEEE Signal Process. Lett. | 3 |
| 2017 | An improved NSGA-III algorithm based on objective space decomposition for many-objective optimization
Xiaojun Bi 0002 |
Soft Comput. | 1 |
| 2012 | Classification-based self-adaptive differential evolution and its application in multi-lateral multi-issue negotiation
Xiaojun Bi 0002 |
Frontiers Comput. Sci. | 1 |
| 2011 | Classification-based self-adaptive differential evolution with fast and reliable convergence performance
Xiaojun Bi 0002 |
Soft Comput. | 1 |