VLDB 2026 Research / reviewers in the wild / expert
Xinwei Xue
dblp:47/4548
· DBLP profile ↗
27ranked-venue papers
14as first author
19since 2021 · last 2026
0000-0003-0082-3034ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiEn: Hierarchical ensemble learning for semi-supervised medical image segmentation
Long Ma 0002, Xinwei Xue, Chengpei Xu, Weimin Wang 0007, Yi Wang 0037 |
Neurocomputing | 4 |
| 2026 | Model-aware ellipse detection via parametric correlation learningabstractEllipse detection presents a significant challenge in computer vision and pattern recognition, often hindered by traditional parameter regression methods that fail to account for the unique geometric characteristics and complex parameter interactions of ellipses. These limitations frequently result in imprecise detections, notably with small or partially occluded ellipses. To overcome these challenges, we propose EDNet, a novel ellipse detection network that exploits the geometric properties of ellipses, thus moving beyond the reliance on internal textures. EDNet improves ellipse detection by refining the loss function to better capture the relationship between the error and each parameters during training. It features a LoG-like Edge Detection Module (LEDM) and an Edge Guided Module (EGM) for precise boundary extraction and multi-scale feature enhancement. Additionally, an auxiliary component estimates ellipse vertices, boosting accuracy for occluded ellipses. Experimental results on two wildly-used benchmark datasets demonstrate that EDNet achieves significant improvements, with an average detection accuracy increase of 6% and 10% over leading state-of-the-art models. • We concentrate on the geometric characteristics of ellipse detection via Edge Detection Module and Edge Guided Module. • We design an auxiliary head for the estimation of four ellipse vertices, invoking additional feature attention on these pivotal points. • We establish the relations between the error and geometric characteristics of the ellipse by a model-aware loss function. Qi Jia 0001, Zezheng Liu, Yu Liu 0012, Yi Wang 0037, Xinwei Xue, Weimin Wang 0007 |
Signal Process. | 5 |
| 2026 | Enhancing Underwater Images via Resonant FusionabstractRecent advances in learning-based underwater image enhancement have achieved remarkable progress. However, the inherent diversity and complexity of underwater scenes still limit the ability of existing approaches to simultaneously restore fine structural details and global image layouts. To address this challenge, we propose a Resonant Fusion (ReFu) framework that explicitly leverages complementary information in both spatial and frequency domains. Specifically, we design a frequency decomposer and a spatial decomposer to capture high- and low-frequency cues from different perspectives. A resonant fuser is then introduced to adaptively integrate high-frequency resonances for detail refinement and low-frequency resonances for structural consistency. This fine-grained cross-domain fusion significantly improves structural preservation and detail enhancement, thereby generating visually more natural and perceptually friendly underwater images. Extensive quantitative and qualitative evaluations across diverse underwater benchmarks show that ReFu consistently surpasses state-of-the-art methods by a clear margin. Comprehensive ablation studies further validate the effectiveness of each module and prove the necessity of the proposed ReFu mechanism. Our code is available at https://github.com/CircleQa/ReFu-main. Xinwei Xue, Zimeng Xu, Jincheng Yuan, Jingchun Zhou, Chengpei Xu, Xiaoke Shang, Long Ma 0002, Weimin Wang 0007 |
IEEE Trans. Image Process. | 1 |
| 2025 | Real-Time Detection of Injection Attacks in Industrial Multi-agent Systems
Kai Di, Chengge Duan, Xinwei Xue |
PDCAT | 4 |
| 2025 | Crossing the Chasm: A practical architecture augmentation for low-quality object detection
Xinwei Xue, Haoze Zheng, Yuechao Gao, Tengyu Ma 0004, Long Ma 0002, Qi Jia 0001 |
Neurocomputing | 1 |
| 2025 | ASF-Net: Robust video deraining via temporal alignment and online adaptive learning
Xinwei Xue, Long Ma 0002, Risheng Liu |
Pattern Recognit. | 1 |
| 2025 | Rectangling for Stitched Image via Pixel-Wise Deformation LearningabstractImage rectangling involves filling in the blanks created during image stitching through deformation techniques. However, existing methods still struggle with incomplete filling and distortion of content, ultimately affecting the overall visual impression and potentially hindering subsequent tasks such as recognition. In this work, we design a pixel-wise deformation framework that utilizes explicit edge guidance to maintain consistency of texture and structure, yielding rectangular images with natural structure. Specifically, we decouple motion into region-level and pixel-level components through uniform mesh warping and pixel-wise deformation to precisely rearrange the spatial distribution of all pixels. Uniform deformation preserves local structure within divided patches, while pixel-wise motion coordinates the consistency between patches. Their combination provides robust and accurate pixel-wise offsets for structure-preserved rectangling. To further bolster the consistency of structure and texture, we leverage edge information to establish structural constraints and design an edge-guided enhancement module to aid in restoring fine texture details. Additionally, stitched images encompass both meaningful content and blank spaces, we innovatively incorporate a mask predictor, which acts as a guiding beacon, directing the network's attention solely towards content-rich regions to facilitate precise pixel-wise motion estimation. Experimental results demonstrate that our approach achieves state-of-the-art performance in rectifying irregular boundaries while contributing to downstream visual perception tasks. Xiaomei Feng, Qi Jia 0001, Yu Liu 0012, Weimin Wang 0007, Yuqing Liu 0001, Xinwei Xue |
IEEE Trans. Multim. | 6 |
| 2024 | Joint edge detection learning for recurrent homography estimationabstractHomography estimation plays a pivotal role in aligning image pairs across multiple viewpoints. Existing methods focus mainly on texture alignment, whereas overlooking the influence of geometric structures, thereby resulting in inaccurate homography estimation. In this paper, we propose a novel recurrent homography estimation framework with joint edge detection learning. We find that edge detection explores extra anchors for homography estimation, and meanwhile homography provides complementary information of cross views for edge detection refinement. Unlike traditional edge detection applied to individual images, our approach establishes structural consistency constraints to reinforce mutual edges while suppressing unreliable structures. Specifically, the detected edges guide and enhance the texture features through a specifically designed edge-aware fusion module. Ultimately, we recurrently compute the correlation of fusion features from small to large scales for homography regression. Our experimental results demonstrate that the proposed method reduces the matching error by 41.7% than state-of-the-art methods. Furthermore, our network excels in detecting edges with extensive details even under dramatic perspective changes. Code is available at https://github.com/edmandzhao/edge-detection-for-RHE. Qi Jia 0001, Zikun Zhao, Xiaomei Feng, Jinyuan Liu 0001, Yu Liu 0012, Xinwei Xue |
ICME | 6 |
| 2024 | CSUNet: Contour-Sensitive Underwater Salient Object Detection
Yi Wang 0037, Shijun Yan, Tianzhu Wang, Zhihan Wang, Weirong Sun, Yu Zhao 0054, Xinwei Xue |
MMAsia | 8 |
| 2024 | Learning Deep Scene Curve for Fast and Robust Underwater Image EnhancementabstractLearning-based approaches inspired by the scattering model for enhancing underwater imagery have gained prominence. Nevertheless, these methods often suffer from time-consuming attributable to their sizable model dimensions. Moreover, they face challenges in adapting unknown scenes, primarily because the scattering model's original design was intended for atmospheric rather than marine condition. To address these obstacles, we begin by investigating the inherent differences in imaging characteristics between atmospheric and marine conditions based on statistical distributions. Building on these observations, we introduce an efficient and effective algorithm called Deep Scene Curve, abbreviated as DSC. This method comprises two essential steps: scene-irrelevant zero-mean adjustment and scene-oriented hyperparameter estimation. The first step transforms scene features into a unified zero-mean space, thereby reducing interference from scene-specific attributes. In the second step, we employ a lightweight neural network to estimate scene-oriented hyperparameters for a defined pixel-level curve based on underwater observations. This approach enables us to generate a deep curve that excels in both adaptability and efficiency, as substantiated by extensive experiments. Notably, our method achieves a significant 56% improvement in average inference time while reducing FLOPs by 92% compared to existing techniques. Furthermore, our extensive experiments in low-light image enhancement tasks highlight the potential advantages of DSC. Xinwei Xue, Yidong Han, Long Ma 0002, Risheng Liu |
IEEE Signal Process. Lett. | 1 |
| 2024 | Edge-Aware Correlation Learning for Unsupervised Progressive Homography EstimationabstractHomography estimation aligns image pairs in cross-views, which is a crucial and fundamental computer vision problem. Existing methods only consider correspondences of texture features for homography estimation, leading to unpleasant artifacts and misalignments introduced by mismatches, especially for low-texture image pairs. In contrast to others, we introduce intuitive structural information as an additional clue that is more sensitive to human vision and low-texture scenarios. In this paper, we propose an edge-aware unsupervised progressive network that couples texture and edge correlation to comprehensively explore potential matching features for homography estimation. To explore robust edge and texture features, we employ a multiscale network to capture feature pyramids with different receptive fields. Then, we design an edge-aware correlation module tailored for homography regression, which plugs in multiscale features to capture accurate correlation maps. Specifically, the edge-aware correlation module leverages the feature-selecting strategy for edge features to capture discriminative matching edges and further guides the texture correlation unit to focus on correctly matched textures. Finally, we leverage multiscale edge-aware correlation maps to predict homography progressively from coarse to fine. Experimental results demonstrate that our proposed method improves PSNR by 11.09% on the real large parallax dataset and reduces matching error by 32.04% on the synthetic COCO dataset, yielding more accurate alignment results than previous state-of-the-art methods. Xiaomei Feng, Qi Jia 0001, Zikun Zhao, Yu Liu 0012, Xinwei Xue, Xin Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Investigating intrinsic degradation factors by multi-branch aggregation for real-world underwater image enhancement
Xinwei Xue, Long Ma 0002, Qi Jia 0001, Risheng Liu, Xin Fan 0001 |
Pattern Recognit. | 1 |
| 2022 | Best of Both Worlds: See and Understand Clearly in the DarkabstractRecently, with the development of intelligent technology, the perception of low-light scenes has been gaining widespread attention. However, existing techniques usually focus on only one task (e.g., enhancement) and lose sight of the others (e.g., detection), making it difficult to perform all of them well at the same time. To overcome this limitation, we propose a new method that can handle visual quality enhancement and semantic-related tasks (e.g., detection, segmentation) simultaneously in a unified framework. Specifically, we build a cascaded architecture to meet the task requirements. To better enhance the entanglement in both tasks and achieve mutual guidance, we develop a new contrastive-alternative learning strategy for learning the model parameters, to largely improve the representational capacity of the cascaded architecture. Notably, the contrastive learning mechanism establishes the communication between two objective tasks in essence, which actually extends the capability of contrastive learning to some extent. Finally, extensive experiments are performed to fully validate the advantages of our method over other state-of-the-art works in enhancement, detection, and segmentation. A series of analytical evaluations are also conducted to reveal our effectiveness. The code is available at https://github.com/k914/contrastive-alternative-learning. Xinwei Xue, Long Ma 0002, Yi Wang 0037, Xin Fan 0001, Risheng Liu |
ACM Multimedia | 1 |
| 2021 | Physics-inspired Learning for Structure-Aware Texture-Sensitive Underwater Image EnhancementabstractRecently, improving the visual quality of underwater images using deep learning-based methods has drawn considerable attention. Unfortunately, diverse environmental factors (e.g., blue/green color distortion) severely limit their performance in real-world environments. Therefore, strengthening the superiority of the underwater image enhancement method is critical. In this paper, we devote ourselves to develop a new architecture with strong superiority and adaptability. Inspired by the underwater imaging principle, we establish a novel physics-inspired learning model that is easy to realize. A Structure-Aware Texture-Sensitive Network (SATS-Net) is further developed to portray the model. The structure-aware module is responsible for structural information, and the texture-sensitive module is responsible for textural information. Thus, SATS-Net successfully incorporates robust characterization absorbed from the physical principle to achieve strong robustness and adaptability. We conduct extensive experiments to demonstrate that SATS-Net outperforms existing advanced techniques in various real-world underwater environments. Xinwei Xue, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
ACML | 1 |
| 2021 | Temporal Rain Decomposition with Spatial Structure Guidance for Video DerainingabstractRecently, removing rain streaks from videos has drawn wide concerns in vision and multimedia communities. But existing works ignore the depicts of image inherent structure and rain location to cause details loss, and their adopted manners of exploiting temporal information are still insufficient. In this work, we propose a multi-frame deraining network with temporal rain decomposition and spatial structure guidance to more effectively accomplish video deraining. A learnable decomposition method is defined to learn the distribution of rain, where the location map acts on a single-frame deraining block. We construct a multi-frame fusion module with a detailed guidance map to integrate temporal and spatial information. Many evaluated experiments demonstrate that our algorithm performs favorably on video deraining tasks compared with other methods. The elaborate ablation study in terms of network architecture fully indicates the effectiveness of our network. Xinwei Xue, Ying Ding 0006, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001 |
ICASSP | 1 |
| 2021 | GTA-Net: Gradual Temporal Aggregation Network for Fast Video DerainingabstractRecently, the development of intelligent technology arouses the requirements of high-quality videos. Rain streak is a frequent and inevitable factor to degrade the video. Many researchers have put their energies into eliminating the adverse effects of rainy video. Unfortunately, how to fully utilize the temporal information from rainy video is still in suspense. In this work, to effectively exploit temporal information, we develop a simple but effective network, Gradual Temporal Aggregation Network (GTA-Net for short). To be specific, according to the temporal distance between rainy frames and the reference frame, we divide the rainy frames into different groups. A multi-stream coarse temporal aggregation module is first performed to aggregate different temporal information with equal status and importance. Then we design a single-stream fine temporal aggregation module to further fuse the integrated frames that maintain the different distances with the target frame. In this way of coarse-to-fine, we not only achieve superior performance, but also gain the surprising execution speed owing to abandon the time-consuming alignment operation. Plenty of experimental results demonstrate that our GTA-Net performs favorably compared to other state-of-the-art approaches. The meticulous ablation study further indicates the effectiveness of our designed GTA-Net. Xinwei Xue, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
ICASSP | 1 |
| 2021 | Searching Frame-Recurrent Attentive Deformable Network for Real-Time Video DerainingabstractVideo deraining has become an issue of great interest since rain streaks inevitably affect video quality. Most of the existing works focus on heuristically designing the network architecture to integrate available information derived from the temporal dimension. However, their inferences take a long time, so that the practicability is somewhat ignored. To solve this problem, we develop a real-time video deraining network in a frame-recurrent manner. It includes a fast attentive deformable alignment module and an automatically-discovered spatial-temporal reconstruction module. In which, the alignment is composed of a single newly-built deformable convolution under the channel attention mechanism to keep the accurate motion consistency and reduce time-consuming by a wide margin. The reconstruction part for the first time introduces the architecture search technique for video deraining to automatically discover a high-effective architecture by designing an effective and compact search space. Experimental results demonstrate remarkable superiority both in computational efficiency and actual performance compared to other state-of-the-art approaches. Xinwei Xue, Long Ma 0002, Yi Wang 0037, Risheng Liu, Xin Fan 0001 |
ICME | 1 |
| 2021 | Underwater Species Detection using Channel Sharpening AttentionabstractWith the continuous exploration of marine resources, underwater artificial intelligent robots play an increasingly important role in the fish industry. However, the detection of underwater objects is a very challenging problem due to the irregular movement of underwater objects, the occlusion of sand and rocks, the diversity of water illumination, and the poor visibility and low color contrast in the underwater environment. In this article, we first propose a real-world underwater object detection dataset (UODD), which covers more than 3K images of the most common aquatic products. Then we propose Channel Sharpening Attention Module (CSAM) as a plug-and-play module to further fuse high-level image information, providing the network with the privilege of selecting feature maps. Fusion of original images through CSAM can improve the accuracy of detecting small and medium objects, thereby improving the overall detection accuracy. We also use Water-Net as a preprocessing method to remove the haze and color cast in complex underwater scenes, which shows a satisfactory detection result on small-sized objects. In addition, we use the class weighted loss as the training loss, which can accurately describe the relationship between classification and precision of bounding boxes of targets, and the loss function converges faster during the training process. Experimental results show that the proposed method reaches a maximum AP of 50.1%, outperforming other traditional and state-of-the-art detectors. In addition, our model only needs an average inference time of 25.4 ms per image, which is quite fast and might suit the real-time scenario. Lihao Jiang, Yi Wang 0037, Qi Jia 0001, Shengwei Xu, Yu Liu 0012, Xin Fan 0001, Risheng Liu, Xinwei Xue, Ruili Wang 0001 |
ACM Multimedia | 9 |
| 2021 | Joint Luminance and Chrominance Learning for Underwater Image EnhancementabstractRecently, learning-based works have been widely-investigated to enhance underwater images. However, interactions between various degradation factors (e.g., color distortion and haze effects) inevitably cause negative interference during the inference phase. Thus, these works cannot fully remove degraded factors. To address this problem, we propose a novel Joint Luminance and Chrominance Learning Network (JLCL-Net). Concretely, we reformulate the task as luminance reconstruction (for haze removal), and chrominance correction (for color correction) sub-tasks by separating the luminance and chrominance (i.e., color appearance) of the underwater images. In this way, we successfully realize the disentanglement in degraded factors to avoid introducing interference. We specify the reconstruction by integrating the atmospheric scattering model, which endows the adaptive dehazing ability over different scenarios. The correction learns to compensate for color by a simple network to reverse the color attenuation process. To this end, we obtain our JLCL-Net. To better train it, we design a new multi-stage cross-space training strategy, which progressively updates the network parameters to enlarge the network potentiality. Extensive evaluations are presented to fully verify our superiority against other methods. Xinwei Xue, Zhenhua Hao, Long Ma 0002, Yi Wang 0037, Risheng Liu |
IEEE Signal Process. Lett. | 1 |
| 2020 | Fine-Grained Action Recognition on a Novel Basketball DatasetabstractCurrently most works on action recognition focus on the coarsely-grained actions, while the fine-grained action recognition is seldom addressed which is of vital importance in many applications such as video retrieval. To tackle this issue, in this paper, we release a challenging dataset by annotating the fine-grained actions in basketball game videos. A benchmark evaluation of the state-of-the-art approaches for action recognition is also provided on our dataset. Furthermore, we propose an approach by integrating the NTS-Net into two-stream network so as to locate the most informative regions and extract more discriminative features for fine-grained action recognition. Our experiments show that the proposed approach significantly outperforms the existing approaches. Xiaofan Gu, Xinwei Xue |
ICASSP | 2 |
| 2020 | Sequential Deep Unrolling With Flow Priors For Robust Video DerainingabstractVideo deraining has attracted wide attention since the urgent demand of high-quality video in recent years. The indistinct details and nonideal deraining effects are the most common defects in existing techniques, whose cause lies in the insufficient usage of single-frame image and temporal information. To effectively settle video deraining, we establish a new deraining model with flow priors to simultaneously introduce spatial and temporal information for accurately depicting the enhancement model of the current frame. A sequential deep unrolling framework is substantially presented by solving this model based on optimization techniques. The ablation study indicates our effectiveness as far as the design of architecture. Plenty of subjective and objective evaluations fully demonstrate our superiority in detail recovery and deraining effects against other state-of-the-are video deraining approaches. Xinwei Xue, Ying Ding 0006, Pan Mu, Long Ma 0002, Risheng Liu, Xin Fan 0001 |
ICASSP | 1 |
| 2020 | Multi-Scale Features Joint Rain Removal For Single ImageabstractThe presence of rain and haze often cause degradation of images. Therefore, it is important to remove rain or haze and recover the background in outdoor vision systems. Due to the limited size of the network acceptance domain, the pixel value of each spatial position can only be inferred from the surrounding small local area; thus, it is often difficult to remove long rain streaks using existing methods. Therefore, we propose a feature joint dense network (FJDN) to extract multi-scale aggregation features. First, we design a multiscale feature extraction module that uses four dilated convolutional layers to extract multi-scale features. These multi-scale features are then combined into one feature map. We also aggregate three multi-scale features in feature joint dense block (FJDB). By using multi-scale features, we can effectively detect rain streaks of different lengths. Finally, we perform multiple experiments to visually and quantitatively compare our method with several existing methods, demonstrating its superiority. The proposed method is also applied to image dehazing. Xinwei Xue, Zhenhua Hao, Ying Ding 0006, Qi Jia 0001, Risheng Liu |
ICIP | 1 |
| 2013 | Multi-scale bidirectional local template patterns for real-time human detectionabstractIn this paper, a feature named multi-scale bidirectional local template patterns (MBLTP) is proposed for human detection. As an extension of bidirectional local template patterns (BLTP), MBLTP not only integrates the textural and gradient information according to the four predefined templates but also calculates information for additional feature vectors by adjusting the scale of the training samples. These additional feature vectors contain multi-scale information on the samples, which can make the feature more discriminative than its original form. Experimental results for an INRIA dataset show that the detection rate of our proposed MBLTP feature outperforms those of other features such as the multi-level histogram of orientated gradient (multi-level HOG), multi scale block histogram of template (MB-HOT), and HOG-LBP. Moreover, in order to make our feature meet real-time requirements, an implementation based on a graphic process unit (GPU) is adopted to accelerate the calculation. Jiu Xu, Ning Jiang 0002, Xinwei Xue, Heming Sun, Wenxin Yu 0001, Satoshi Goto |
MMSP | 3 |
| 2012 | Motion robust rain detection and removal from videosabstractWeather such as rain and snow cause difficulties in processing the videos captured. Since the appearance of rain drops can affect the performance of human tracking and reduce the efficiency of video compression, detection and removal of rain is a challenging problem in outdoor surveillance systems. In this paper, we propose a new algorithm for rain detection, which is based on joint spatial and wavelet domain features. This approach is robust to the videos with moving objects in the rain. Experimental results demonstrated its better performance in comparison with the existing approaches in the subjective quality. Xinwei Xue, Xin Jin 0002, Satoshi Goto |
MMSP | 1 |
| 2009 | HDR image compression using optimized tone mapping modelabstractIn this paper, we propose a coding algorithm for High Dynamic Range Images (HDRI). Our encoder applies a tone mapping model based on scaled μ-Lawencoding, followed by a conventional Low Dynamic Range Image (LDRI) encoder. The tone mapping model is designed to minimize the difference between the tone mappedHDRI and its LDR version. By virtue of the nature of the model, not only the quality of the HDRI but also the one of LDRI are improved, compared with a state of the art in conventional HDRI compression. Furthermore the error caused by our tone mapping model encoding is theoretically analyzed. Nagisa Sugiyama, Hironori Kaida, Xinwei Xue, Takao Jinno, Nicola Adami, Masahiro Okuda |
ICASSP | 3 |
| 2005 | Constructing Comprehensive Behaviors: A Simulation Study
Thomas C. Henderson, Xinwei Xue |
CAINE | 2 |
| 2001 | Variable-conductance, level-set curvature for image denoisingabstractThis paper describes a partial differential equation for denoising images. The proposed method is demonstrably superior to anisotropic diffusion (and its many variations) for denoising images that are approximately piecewise constant. The method relies on an equation that is the level-set equivalent of the anisotropic diffusion equation proposed by Perona and Malik (1990). This proposed equation has come up in the literature, but has failed to be fully utilized due to a lack of analysis and the need for a stable, accurate numerical implementation. Our analysis shows that the proposed method is more aggressive than anisotropic diffusion at enhancing and preserving edges, and is less sensitive to the edge contrast parameter. Empirical results confirm these advantages, and show that for certain classes of images, one should always prefer the proposed method over anisotropic diffusion. Ross T. Whitaker, Xinwei Xue |
ICIP (3) | 2 |