VLDB 2026 Research / reviewers in the wild / expert
Xue Xia 0005
dblp:168/6125-5
· DBLP profile ↗
28ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0002-2872-7151ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-CancerabstractSingle-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%). Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002 |
AAAI | 4 |
| 2026 | Multi-scale interleaved transformer network for image deraining
Yue Que 0001, Hanqing Xiong, Xue Xia 0005 |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Cross-Structure and Semantic Enhancement for Diabetic Retinopathy GradingabstractChallenges such as highly variable lesion appearances and complex structural distributions hinder model performance in diabetic retinopathy (DR) grading tasks. To address these issues, we focus on guiding the network toward discriminative feature representation by prioritizing diagnostically relevant information and modeling intricate dependencies between features and DR grades. Specifically, we propose a DR grading network that integrates the Kolmogorov-Arnold Network (KAN) into a Convolution-Vision Transformer (CNN-ViT) cooperative framework, where convolutions and Transformer encoders capture spatial patterns, hierarchical structures, and context, while KAN enhances non-linear semantic dependency modeling. Additionally, we introduce a Cross-Structure (CS) attention module to emphasize relevant features. The proposed modules form the Convolution-Cross-Structure-KAN (CCSK) block, which serves as the backbone of our network, CCSKFormer, enabling more accurate DR grading. The proposed model achieves outstanding performance on two public datasets, with comparisons and ablation studies further validating the effectiveness of the individual modules (https://github.com/xia-xx-cv/CCSKformer). Xue Xia 0005, Zipeng Lin, Jingying Zhu, Jiebin Yan, Yuming Fang 0001 |
ICME | 1 |
| 2025 | Hybrid Mamba-Transformer with Frequency Enhancement for Single Image Deraining
Yue Que 0001, Wenjun Xia, Xue Xia 0005 |
PRCV (9) | 3 |
| 2025 | Texture-Aware Network for Enhancing Inner Smoke Representation in Visual Smoke Density EstimationabstractABSTRACT Smoke often appears before visible flames in the early stages of fire disasters, making accurate pixel‐wise detection essential for fire alarms. Although existing segmentation models effectively identify smoke pixels, they generally treat all pixels within a smoke region as having the same prior probability. This assumption of rigidity, common in natural object segmentation, fails to account for the inherent variability within smoke. We argue that pixels within smoke exhibit a probabilistic relationship with both smoke and background, necessitating density estimation to enhance the representation of internal structures within the smoke. To this end, we propose enhancements across the entire network. First, we improve the backbone by adaptively integrating scene information into texture features through separate paths, enabling smoke‐tailored feature representation for further exploit. Second, we introduce a texture‐aware head with long convolutional kernels to integrate both global and orientation‐specific information, enhancing representation for intricate smoke structure. Third, we develop a dual‐task decoder for simultaneous density and location recovery, with the frequency‐domain alignment in the final stage to preserve internal smoke details. Extensive experiments on synthetic and real smoke datasets demonstrate the effectiveness of our approach. Specifically, comparisons with 17 models show the superiority of our method, with mean IoU improvements of 4.88%, 2.63%, and 3.17% on three test sets. (The code will be available on https://github.com/xia‐xx‐cv/TANet_smoke ). Xue Xia 0005, Yajing Peng, Zichen Li, Jinting Shi, Yuming Fang 0001 |
IET Comput. Vis. | 1 |
| 2024 | Benchmarking deep models on retinal fundus disease diagnosis and a large-scale datasetabstractRetinal fundus imaging contributes to monitoring the vision of patients by providing views of the interior surface of the eyes. Machine learning models greatly aided ophthalmologists in detecting retinal disorders from color fundus images. Hence, the quality of the data is pivotal for enhancing diagnosis algorithms, which ultimately benefits vision care and maintenance. To facilitate further research in this domain, we introduce the Eye Disease Diagnosis and Fundus Synthesis (EDDFS) dataset, comprising 28,877 fundus images. These include 15,000 healthy samples and a diverse range of images depicting various disorders such as diabetic retinopathy, age-related macular degeneration, glaucoma, pathological myopia, hypertension retinopathy, retinal vein occlusion, and Laser photocoagulation. In addition to providing the dataset, we propose a Transformer-joint convolution network for automated eye disease screening. Firstly, a co-attention structure is integrated to capture long-range attention information along with local features. Secondly, a cross-stage feature fusion module is designed to extract multi-level and disease-related information. By leveraging the dataset and our proposed network, we establish benchmarks for disease screening and grading tasks. Our experimental results underscore the network’s proficiency in both multi-label and single-label disease diagnosis, while also showcasing the dataset’s capability in supporting fundus synthesis. (The dataset and code will be available on https://github.com/xia-xx-cv/EDDFS_dataset). Xue Xia 0005, Guobei Xiao, Kun Zhan, Jinhua Yan, Yuming Fang 0001, Guofu Huang |
Signal Process. Image Commun. | 1 |
| 2024 | Learning content-aware feature fusion for guided depth map super-resolution
Yifan Zuo 0001, Xiaoshui Huang, Xue Xia 0005, Yuming Fang 0001 |
Signal Process. Image Commun. | 6 |
| 2024 | Denoising Diffusion Probabilistic Model for Face Sketch-to-Photo SynthesisabstractThe field of face sketch-to-photo synthesis involves generating photographic facial images with enhanced details and a heightened sense of style realism. In recent years, the advancement of deep learning techniques has significantly contributed to the development of methods for synthesizing photographic face images from sketches. Nevertheless, challenges remain in synthesizing facial photographs with richer details and more accurate structural representation. This paper introduces a novel architecture for face sketch-to-photo synthesis, using denoising diffusion probabilistic models (DDPM). Our approach simplifies the complex transformation process into sequential forward and backward denoising steps. We incorporate a pretrained coarse generator to effectively encode sketch information, integrating it into each backward step to guide the generative process toward accurate photo space representation. Furthermore, we design a detail diffusion branch to refine the coarse photo face generated from the coarse generator. By deeply fusing multiscale detail features from this branch with a sophisticated conditional noise predictor, our model effectively captures the correlation between detail and stylistic elements both in sketches and in photographic faces. Extensive experimental evaluations on three datasets show the effectiveness of our model, emphasizing its ability to synthesize facial photographs with remarkable realism and rich detail. The synthesized facial images consistently demonstrate superior face recognition accuracy, surpassing that of state-of-the-art methods. Yue Que 0001, Li Xiong 0018, Weiguo Wan, Xue Xia 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Video Quality Assessment for Online Processing: From Spatial to Temporal SamplingabstractWith the rapid development of multimedia processing and deep learning technologies, especially in the field of video understanding, video quality assessment (VQA) has achieved significant progress. Although researchers have moved from designing efficient video quality mapping models to various research directions, in-depth exploration of the effectiveness-efficiency trade-offs of spatio-temporal modeling in VQA models is still less sufficient. Considering the fact that videos have highly redundant information, this paper investigates this problem from the perspective of joint spatial and temporal sampling, aiming to seek the answer to how little information we should keep at least when feeding videos into the VQA models while with acceptable performance sacrifice. To this end, we drastically sample the video’s information from both spatial and temporal dimensions, and the heavily squeezed video is then fed into a stable VQA model. Comprehensive experiments regarding joint spatial and temporal sampling are conducted on six public video quality databases, and the results demonstrate the acceptable performance of the VQA model when throwing away most of the video information. Furthermore, with the proposed joint spatial and temporal sampling strategy, we make an initial attempt to design an online VQA model, which is instantiated by as simple as possible a spatial feature extractor, a temporal feature fusion module, and a global quality regression module. Through quantitative and qualitative experiments, we verify the feasibility of online VQA model by simplifying itself and reducing input. Jiebin Yan, Yuming Fang 0001, Xuelin Liu, Xue Xia 0005, Weide Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Bayesian Uncertainty Calibration for Federated Time Series AnalysisabstractDeep learning models for time series analysis often require large-scale labeled datasets for training. However, acquiring such datasets is cost-intensive and challenging, particularly for individual institutions. To overcome this challenge and concern about data confidentiality among different institutions, federated learning (FL) servers as a viable solution to this dilemma by offering a decentralized learning framework. However, the datasets collected by each institution often suffer from imbalance and may not adhere to uniform protocols, leading to diverse data distributions. To address this problem, we design a global model to approximate the global data distribution of all participant clients, then transfer it to local clients as an induction in the training phase. While discrepancies between the approximate distribution and the actual distribution result in uncertainty in the predicted results. Moreover, the diverse data distributions among various clients within the FL framework, combined with the inherent lack of reliability and interpretability in deep learning models, further amplify the uncertainty of the prediction results. To address these issues, we propose an uncertainty calibration method based on Bayesian deep learning techniques, which captures uncertainty by learning a fidelity transformation to reconstruct the output of time series regression and classification tasks, utilizing deterministic pre-trained models. Extensive experiments on the regression dataset (C-MAPSS) and classification datasets (ESR, Sleep-EDF, HAR, and FD) in the Independent and Identically Distributed (IID) and non-IID settings show that our approach effectively calibrates uncertainty within the FL framework and facilitates better generalization performance in both the regression and classification tasks, achieving state-of-the-art performance. Weide Liu, Xue Xia 0005, Zhenghua Chen, Yuming Fang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Edge-reinforced attention network for smoke semantic segmentation
Lin Zhang 0061, Feiniu Yuan, Xue Xia 0005 |
Multim. Tools Appl. | 3 |
| 2022 | Eye Disease Diagnosis and Fundus Synthesis: A Large-Scale Dataset and BenchmarkabstractAs one of the most common imaging modalities, retinal fundus imaging offers images of interior surface of eyes for initial examination of disorders. Data-driven machine learning methods, especially deep learning models in recent years, provide automatic ophthalmological disease diagnosis techniques from color fundus images. Data with high quality, diversity and balanced distribution supports deep model-based eye disease diagnosis. However, many existing datasets focus on a specific kind of eye disease, and some suffer from label noise or quality degeneration, which hinders automatic screening algorithms from dealing with multiple eye diseases. To solve this, we propose a high-quality dataset containing 28877 color fundus images for deep learning-based diagnosis. Except for 15000 healthy samples, the dataset consists of 8 eye disorders including diabetic retinopathy, agerelated macular degeneration, glaucoma, pathological myopia, hypertension, retinal vein occlusion, LASIK spot and others. Based on this, we propose a co-attention network for disease diagnosis, establish benchmark on screening and grading tasks, and demonstrate that the proposed dataset supports generative adversarial network-based image synthesis. The dataset will be made publicly available. Xue Xia 0005, Kun Zhan, Guobei Xiao, Jinhua Yan, Zhuxiang Huang, Guofu Huang, Yuming Fang 0001 |
MMSP | 1 |
| 2022 | FundusGAN: A One-Stage Single Input GAN for Fundus Synthesis
Xue Xia 0005, Yuming Fang 0001 |
PRCV (2) | 2 |
| 2022 | Texture-aware Network for Smoke Density EstimationabstractSmoke density estimation, also termed as soft segmentation, was developed from pixel-wise smoke (hard) segmen-tation and it aims at providing transparency and segmentation confidence for each pixel. The key difference between them lies in that segmentation focuses on classifying pixels into smoke and non-smoke ones, while density estimation obtains inner transparency of smoke component rather than treat all smoke pixels as an equal value. Based on this, we propose a texture-aware network being able to capture inner transparency of smoke components rather than merely focus on general smoke distribution for pixel-wise smoke density estimation. Besides, we adapt the Squeeze-and-Excitation (SE) layer for smoke feature extraction by involving max values for robustness. In order to represent inhomogeneous smoke pixels, we proposed a simple yet efficient attention-based texture-aware module that involves both gradient and semantic information. Experimental results show that our method outperforms others in both single image density estimation or segmentation and video smoke detection. Xue Xia 0005, Kun Zhan, Yajing Peng, Yuming Fang 0001 |
VCIP | 1 |
| 2022 | Cubic-cross convolutional attention and count prior embedding for smoke segmentation
Feiniu Yuan, Zeshu Dong, Lin Zhang 0061, Xue Xia 0005, Jinting Shi |
Pattern Recognit. | 4 |
| 2022 | Subjective and Objective Quality of Experience of Free Viewpoint VideosabstractFree viewpoint videos (FVVs) provide immersive experiences for end-users, and they have been applied in many applications, such as movies, sports, and TV shows. However, the development of quantifying the quality of experience (QoE) of FVVs is still relatively slow due to the high costs of data collection and limited public databases. In this paper, we conduct a comprehensive study on FVV QoE. First, we construct the largest, to the best of our knowledge, FVV QoE database called Youku-FVV from two complex real scenarios, i. e., entertainment and sports. Specifically, Youku-FVV originates from the videos captured by dozens of real cameras arranged annularly. We use these videos to generate virtual viewpoints, which make up FVVs together with real views. In constructing the FVV QoE database, we consider both internal and external influencing factors of QoE, which correspond to FVV generation and playback, respectively. Besides, we make an initial attempt to train an efficient no reference FVV QoE prediction model using this database, where several sparse frame sampling strategies are validated. And we demonstrate the feasibility of striving for the balance between effectiveness and efficiency of FVV QoE prediction. The proposed FVV QoE database and source codes are publicly available at https://github.com/QTJiebin/FVV_QoE. Jiebin Yan, Jing Li 0026, Yuming Fang 0001, Zhaohui Che, Xue Xia 0005, Yang Liu 0293 |
IEEE Trans. Image Process. | 5 |
| 2021 | A confidence prior for image dehazing
Feiniu Yuan, Yu Zhou 0009, Xue Xia 0005, Xueming Qian |
Pattern Recognit. | 3 |
| 2021 | A Gated Recurrent Network With Dual Classification Assistance for Smoke Semantic SegmentationabstractSmoke has semi-transparency property leading to highly complicated mixture of background and smoke. Sparse or small smoke is visually inconspicuous, and its boundary is often ambiguous. These reasons result in a very challenging task of separating smoke from a single image. To solve these problems, we propose a Classification-assisted Gated Recurrent Network (CGRNet) for smoke semantic segmentation. To discriminate smoke and smoke-like objects, we present a smoke segmentation strategy with dual classification assistance. Our classification module outputs two prediction probabilities for smoke. The first assistance is to use one probability to explicitly regulate the segmentation module for accuracy improvement by supervising a cross-entropy classification loss. The second one is to multiply the segmentation result by another probability for further refinement. This dual classification assistance greatly improves performance at image level. In the segmentation module, we design an Attention Convolutional GRU module (Att-ConvGRU) to learn the long-range context dependence of features. To perceive small or inconspicuous smoke, we design a Multi-scale Context Contrasted Local Feature structure (MCCL) and a Dense Pyramid Pooling Module (DPPM) for improving the representation ability of our network. Extensive experiments validate that our method significantly outperforms existing state-of-art algorithms on smoke datasets, and also obtain satisfactory results on challenging images with inconspicuous smoke and smoke-like objects. Feiniu Yuan, Lin Zhang 0027, Xue Xia 0005, Qinghua Huang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Image dehazing based on a transmission fusion strategy by automatic image matting
Feiniu Yuan, Yu Zhou 0009, Xue Xia 0005, Jinting Shi, Yuming Fang 0001, Xueming Qian |
Comput. Vis. Image Underst. | 3 |
| 2020 | A Wave-Shaped Deep Neural Network for Smoke Density EstimationabstractSmoke density estimation from a single image is a totally new but highly ill-posed problem. To solve the problem, we stack several convolutional encoder-decoder structures together to propose a wave-shaped neural network, termed W-Net. Stacking encoder-decoders directly increases the network depth, leading to the enlargement of receptive fields for encoding more semantic information. To maximize the degrees of feature re-usage, we copy and resize the outputs of encoding layers to corresponding decoding layers, and then concatenate them to implement short-cut connections for improving spatial accuracy. The crests and troughs of W-Net are special structures containing abundant localization and semantic information, so we also use short-cut connections between these structures and decoding layers. Estimated smoke density is useful in many applications, such as smoke segmentation, smoke detection, disaster simulation. Experimental results show that our method outperforms existing methods on both smoke density estimation and segmentation. It also achieves satisfying results in visual detection of auto exhausts. Feiniu Yuan, Lin Zhang 0061, Xue Xia 0005, Qinghua Huang, Xuelong Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Encoding pairwise Hamming distances of Local Binary Patterns for visual smoke recognition
Feiniu Yuan, Jinting Shi, Xue Xia 0005, Lin Zhang 0061 |
Comput. Vis. Image Underst. | 3 |
| 2019 | Co-occurrence matching of local binary patterns for improving visual adaption and its application to smoke recognitionabstractIt is challenging to recognize smoke from visual scenes due to large variations of smoke colors, textures and shapes. To improve robustness, we propose a novel feature extraction method based on similarity and dissimilarity matching measures of Local Binary Patterns (LBP). Given two bit‐sequences of an LBP code pair, the similarity and dissimilarity matching measures are defined as the ratios of the 1–1 bitwise matching number to the 0–0 bitwise matching number and the 1–0 number to the 0–1 number, respectively. To capture local code variations, we calculate the measures between LBP codes of a center pixel and its neighbors. Then we compare each measure with its global mean to propose Similarity Matching based Local Binary Patterns (SMLBP) and Dissimilarity Matching based Local Binary Patterns (DMLBP). Since SMLBP and DMLBP extract spatial variations of the 1st order LBP codes, they actually represent the 2nd order variations of pixel values. Furthermore, we adopt different mapping modes and multi‐scale neighborhoods to obtain rotation and scale invariances. Finally, we concatenate the histograms of LBP, SMLBP and DMLBP to generate a feature vector containing 1st and 2nd order information. Experiments show that our method obviously outperforms existing methods. Feiniu Yuan, Jinting Shi, Xue Xia 0005, Qinghua Huang, Xuelong Li 0001 |
IET Comput. Vis. | 3 |
| 2019 | Fusing texture, edge and line features for smoke recognitionabstractTo improve recognition accuracy, the authors fuse texture, edge and line information to propose a feature extraction method for smoke recognition. The Canny operator is proposed to generate an edge image from an original image, and then adopt the Hough transform to extract straight lines from the edge image. The lines are rasterised to generate a discrete line image and two local patterns are proposed for the edge and line images. The first one is local boundary summation pattern (LBSP) that computes the sum of binary pixel values along the boundary of a local region around a centre pixel. The second one is called local region summation pattern (LRSP) that sums up the binary values of pixels in a local region around the centre pixel. Besides LBSP and LRSP, LBPs with three mapping modes (LBP_M3) to achieve traditional texture information are also extracted. Finally, the authors concatenate the histograms of LBP_M3, LBSP and LRSP to generate a feature vector, and use support vector machine for classifying and testing. Experiments show that authors’ method outperforms most of existing traditional methods for smoke recognition. Although this method has low dimensional features, it also obtains good performance for multi‐class texture classification. Feiniu Yuan, Xue Xia 0005, Bang Jun Lei, Jinting Shi |
IET Image Process. | 3 |
| 2019 | Deep smoke segmentation
Feiniu Yuan, Lin Zhang 0061, Xue Xia 0005, Boyang Wan, Qinghua Huang, Xuelong Li 0001 |
Neurocomputing | 3 |
| 2019 | Convolutional neural networks based on multi-scale additive merging layers for visual smoke recognition
Feiniu Yuan, Lin Zhang 0061, Boyang Wan, Xue Xia 0005, Jinting Shi |
Mach. Vis. Appl. | 4 |
| 2018 | Mixed co-occurrence of local binary patterns and Hamming-distance-based local binary patterns
Feiniu Yuan, Xue Xia 0005, Jinting Shi |
Inf. Sci. | 2 |
| 2018 | Learning multi-scale and multi-order features from 3D local differences for visual smoke recognition
Feiniu Yuan, Xue Xia 0005, Jinting Shi, Lin Zhang 0061, Jifeng Huang |
Inf. Sci. | 2 |
| 2016 | High-order local ternary patterns with locality preserving projection for smoke detection and image classification
Feiniu Yuan, Jinting Shi, Xue Xia 0005, Yuming Fang 0001, Zhijun Fang 0001, Tao Mei 0001 |
Inf. Sci. | 3 |