VLDB 2026 Research / reviewers in the wild / expert
Xinyi Ying
dblp:262/3992
· DBLP profile ↗
19ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-3683-1477ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A mutual information-based framework for generalized image fusion via common-unique decoupling
Liyuan Pan, Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002 |
Knowl. Based Syst. | 6 |
| 2026 | Probing Deep Into Temporal Profile Makes the Infrared Small Target Detector Much BetterabstractInfrared small target (IRST) detection is challenging in simultaneously achieving precise, robust, and efficient performance due to extremely dim targets and strong interference. Current learning-based methods attempt to leverage "more" information from both the spatial and the short-term temporal domains, but suffer from unreliable performance under complex conditions while incurring computational redundancy. In this paper, we explore the "more essential" information from a more crucial domain for the detection. Through theoretical analysis, we reveal that the global temporal saliency and correlation information in the temporal profile demonstrate significant superiority in distinguishing target signals from other signals. To investigate whether such superiority is preferentially leveraged by well-trained networks, we built the first prediction attribution tool in this field and verified the importance of the temporal profile information. Inspired by the above conclusions, we remodel the IRST detection task as a one-dimensional signal anomaly detection task, and propose an efficient deep temporal probe network (DeepPro) that only performs calculations in the time dimension for IRST detection. We conducted extensive experiments to fully validate the effectiveness of our method. The experimental results are exciting, as our DeepPro outperforms existing state-of-the-art IRST detection methods on widely-used benchmarks with extremely high efficiency, and achieves a significant improvement on dim targets and in complex scenarios. We provide a new modeling domain, a new insight, a new method, and a new performance, which can promote the development of IRST detection. Ruojing Li, Wei An 0003, Yingqian Wang 0002, Xinyi Ying, Yimian Dai, Longguang Wang, Yulan Guo, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Dynamic High-Frequency Convolution for Infrared Small Target DetectionabstractInfrared small targets are typically tiny and locally salient, which belong to high-frequency components (HFCs) in images. Single-frame infrared small target (SIRST) detection is challenging, since there are many HFCs along with targets, such as bright corners, broken clouds, and other clutters. Current learning-based methods rely on the powerful capabilities of deep networks, but neglect explicit modeling and discriminative representation learning of various HFCs, which is important to distinguish targets from other HFCs. To address the aforementioned issues, we propose a dynamic high-frequency convolution (DHiF) to translate the discriminative modeling process into the generation of a dynamic local filter bank. Especially, DHiF is sensitive to HFCs, owing to the dynamic parameters of its generated filters being symmetrically adjusted within a zero-centered range according to Fourier transformation properties. Combining with standard convolution operations, DHiF can adaptively and dynamically process different HFC regions and capture their distinctive grayscale variation characteristics for discriminative representation learning. DHiF functions as a drop-in replacement for standard convolution and can be used in arbitrary SIRST detection networks without significant decrease in computational efficiency. To validate the effectiveness of our DHiF, we conducted extensive experiments across different SIRST detection networks on real-scene datasets. Compared to other state-of-the-art convolution operations, DHiF exhibits superior detection performance with promising improvement. Codes are available at https://github.com/TinaLRJ/DHiF. Ruojing Li, Wei An 0003, Xinyi Ying, Yingqian Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Redundancy-aware masked graph autoencoder for overlapping community detection in attributed networks
Hongkai Xie, Xinyi Ying, Xiaofeng Wang 0004, Xiaofeng Huang, Junzheng Jiang, Daying Quan |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Multimodal image generation and fusion through content-style hybrid disentanglementabstract• Research highlight 1: We propose a novel cross-task hybrid training methodology for multimodal images, offering a simple yet unified solution that simultaneously addresses both image generation and fusion tasks. • Research highlight 2: Building upon mutual-supervised multimodal image pairs, we innovatively integrate single-modality self-supervision to develop a hybrid-supervised decoupling framework with a dedicated loss function, achieving robust separation of content-style representations. • Research highlight 3: Extensive experiments spanning on four modalities and seven popular datasets demonstrate our method’s consistent superiority and impressive cross-task capability. Ablation studies further reveal that our framework learns generalized representations transferable across different image processing tasks. Multimodal image fusion and cross-modal translation are fundamental yet challenging tasks in computer vision, with their performance directly impacting downstream applications. Existing approaches typically treat these tasks independently, developing specialized models that fail to exploit the intrinsic relationships between different modalities. This limitation not only restricts model generalizability but also hinders further performance improvements. In this paper, we propose a joint optimization framework for image generation and fusion. Specifically, we generalize multimodal image tasks as the fusion and transformation of cross-modal features, and design a hybrid task training strategy. At the data level, we introduce a self-supervised and mutual-supervised hybrid mechanism for content-style feature decoupling, which achieves superior feature separation through stepwise training on intra-modal and cross-modal data. At the model level, we construct a triple-branch decoupling head along with fusion and transformation modules to ensure synchronous and efficient execution of dual tasks. Our method not only breaks through the single task limitation of the model, but also innovatively introduces mixed supervision into multimodal processing. We conduct comprehensive experiments covering four modalities fusion tasks on seven popular datasets. Extensive experimental results demonstrate that our method achieves superior performance on two tasks as compared of the respective state-of-the-art methods, and show impressive cross-task generalization capability. Huanxin Zou, Jun Li 0020, Hao Chen 0046, Xinyi Ying, Shitian He, Yingqian Wang 0002, Liyuan Pan |
Knowl. Based Syst. | 5 |
| 2025 | Visible-Thermal Tiny Object Detection: A Benchmark Dataset and BaselinesabstractVisible-thermal small object detection (RGBT SOD) is a significant yet challenging task with a wide range of applications, including video surveillance, traffic monitoring, search and rescue. However, existing studies mainly focus on either visible or thermal modality, while RGBT SOD is rarely explored. Although some RGBT datasets have been developed, the insufficient quantity, limited diversity, unitary application, misaligned images and large target size cannot provide an impartial benchmark to evaluate RGBT SOD algorithms. In this paper, we build the first large-scale benchmark with high diversity for RGBT SOD (namely RGBT-Tiny), including 115 paired sequences, 93 K frames and 1.2 M manual annotations. RGBT-Tiny contains abundant objects (7 categories) and high-diversity scenes (8 types that cover different illumination and density variations). Note that, over 81% of objects are smaller than 16×16, and we provide paired bounding box annotations with tracking ID to offer an extremely challenging benchmark with wide-range applications, such as RGBT image fusion, object detection and tracking. In addition, we propose a scale adaptive fitness (SAFit) measure that exhibits high robustness on both small and large objects. The proposed SAFit can provide reasonable performance evaluation and promote detection performance. Based on the proposed RGBT-Tiny dataset, extensive evaluations have been conducted with IoU and SAFit metrics, including 30 recent state-of-the-art algorithms that cover four different types (i.e., visible generic object detection, visible SOD, thermal SOD and RGBT object detection). Xinyi Ying, Wei An 0003, Ruojing Li, Boyang Li 0007, Zhaoxu Li, Yingqian Wang 0002, Mingyuan Hu, Zaiping Lin, Shilin Zhou 0001, Li Liu 0002, Weidong Sheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Infrared Small Target Detection in Satellite Videos: A New Dataset and a Novel Recurrent Feature Refinement FrameworkabstractMultiframe infrared small target (MIRST) detection in satellite videos has been a long-standing, fundamental yet challenging task for decades, and the challenges can be summarized as follows. First, the extremely small target size, highly complex clutter & noise and various satellite motions result in limited feature representation, high false alarms and difficult motion analyses. In addition, existing methods are primarily designed for static or slightly adjusted perspectives captured by short-distance platforms, which cannot generalize well to complex background motion in satellite videos. Second, the lack of a large-scale publicly available MIRST dataset in satellite videos greatly hinders the algorithm development. To address the aforementioned challenges, in this article, we first build a large-scale dataset for MIRST detection in satellite videos (namely IRSatVideo-LEO), and then develop a recurrent feature refinement (RFR) framework as the baseline method for satellite motion estimation and compensation. Specifically, IRSatVideo-LEO is a semi-simulated dataset with synthesized satellite motion, target appearance, trajectory, and intensity, which can provide a standard toolbox for satellite video generation and a reliable evaluation platform to facilitate algorithm development. For the baseline method, RFR is proposed to be equipped with existing powerful CNN-based methods for long-term temporal dependency exploitation and integrated motion compensation and MIRST detection. Specifically, a pyramid deformable alignment (PDA) module is proposed to achieve effective feature alignment, and a temporal-spatial–frequent modulation (TSFM) module is proposed to achieve efficient feature aggregation and enhancement. Extensive experiments have been conducted to demonstrate the effectiveness and superiority of our scheme. The comparative results show that ResUNet equipped with RFR outperforms the state-of-the-art MIRST detection methods. The dataset and code are available athttps://github.com/XinyiYing/RFR. Xinyi Ying, Li Liu 0002, Zaiping Lin, Yangsi Shi, Yingqian Wang 0002, Ruojing Li, Boyang Li 0007, Shilin Zhou 0001, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Multi-Scale Direction-Aware Network for Infrared Small Target DetectionabstractInfrared small target detection faces the problem that it is difficult to effectively separate the background and the target. Existing deep learning-based methods focus on edge and shape features, but ignore the richer structural differences and detailed information embedded in high-frequency components from different directions, thereby failing to fully exploit the value of high-frequency directional features in target perception. To address this limitation, we propose a multi-scale direction-aware network (MSDA-Net), which is the first attempt to integrate the high-frequency directional features of infrared small targets as domain prior knowledge into neural networks. Specifically, to fully mine the high-frequency directional features, on the one hand, a high-frequency direction injection (HFDI) module without trainable parameters is constructed to inject the high-frequency directional information of the original image into the network. On the other hand, a multi-scale direction-aware (MSDA) module is constructed, which promotes the full extraction of local relations at different scales and the full perception of key features in different directions. In addition, considering the characteristics of infrared small targets, we construct a feature aggregation (FA) structure to address target disappearance in high-level feature maps, and a feature calibration fusion (FCF) module to alleviate feature bias during cross-layer feature fusion. Extensive experimental results show that our MSDA-Net achieves state-of-the-art (SOTA) results on multiple public datasets. The code can be available at https://github.com/YuChuang1205/MSDA-Net. Jinmiao Zhao, Zelin Shi, Chuang Yu 0003, Yunpeng Liu 0001, Xinyi Ying, Yimian Dai |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Motion and Appearance Decoupling Representation for Event CamerasabstractEvent cameras, with high temporal resolution and high dynamic range, have shown great potential under extreme scenarios such as high-speed movement and low illumination. However, previous event representation methods typically aggregate event data into a single dense tensor, often overlooking the dynamic changes of events within a given time unit. This limitation can introduce historical artifacts and semantic inconsistencies, ultimately degrading model performance. Inspired by human visual prior, we propose a motion and appearance decoupling (MAD) event representation to disentangle the mixed spatial-temporal event tensor into two independent branches. This bio-inspired design helps the network extract discriminative temporal (i.e., motion) and spatial (i.e., appearance) information, thus reducing the network's learning burden toward complex high-level interpretation tasks. In our method, the event motion guided attention module (EMGA) is designed to achieve temporal and spatial feature interaction and fusion sequentially. Based on EMGA, three specially designed decoder heads are proposed for several representative event-based tasks (i.e., object detection, semantic segmentation, and human pose estimation). Experimental results demonstrate that our method achieves state-of-the-art performance on the above three tasks, which reveals that our method is an easy-to-implement replacement for currently event-based methods. Our code is available at: https://github.com/ChenYichen9527/MAD-representation. Boyang Li 0007, Yingqian Wang 0002, Xinyi Ying, Longguang Wang, Chushu Zhang, Yulan Guo, Wei An 0003 |
IEEE Trans. Image Process. | 4 |
| 2024 | ICPR 2024 Competition on Resource-Limited Infrared Small Target Detection Challenge: Methods and Results
Boyang Li 0007, Xinyi Ying, Ruojing Li, Yongxian Liu, Yangsi Shi, Xin Zhang 0170, Mingyuan Hu, Yukai Zhang, Dongli Tang, Qiang Ling 0002, Zaiping Lin, Weidong Sheng, Chenxu Peng, Huoren Yang, Lingjie Liu, Zelin Shi, Yunpeng Liu 0001, Chuang Yu 0003, Jinmiao Zhao, Heng Xiang, Tianyu Li 0005, Minghang Zhou, Chenxi Lan, Dongyu Xi, Chaofan Qiao, Yupeng Gao, Yongxu Liu 0006, Deping Chen, Xiaopeng Song, Jiuping Yang, Zhaobing Qiu, Rixiang Ni, Changhai Luo, Shuyuan Zheng, Baojin Huang, Xiaoqi Zhou, Qingshan Guo, Dangxuan Wu, Haodong Zeng, Qiang Fu 0017, Yimian Dai, Renke Kou, Jian Song 0007, Changfeng Feng, Zihao Xiong, Mengxuan Xiao, Yingxu Liu, Quanyi Zhao |
ICPR (34) | 2 |
| 2023 | Mapping Degeneration Meets Label Evolution: Learning Infrared Small Target Detection with Single Point SupervisionabstractTraining a convolutional neural network (CNN) to detect infrared small targets in a fully supervised manner has gained remarkable research interests in recent years, but is highly labor expensive since a large number of per-pixel annotations are required. To handle this problem, in this paper, we make the first attempt to achieve infrared small target detection with point-level supervision. Interestingly, during the training phase supervised by point labels, we discover that CNNs first learn to segment a cluster of pixels near the targets, and then gradually converge to predict groundtruth point labels. Motivated by this “mapping degeneration” phenomenon, we propose a label evolution framework named label evolution with single point supervision (LESPS) to progressively expand the point label by leveraging the intermediate predictions of CNNs. In this way, the network predictions can finally approximate the updated pseudo labels, and a pixel-level target mask can be obtained to train CNNs in an end-to-end manner. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Experimental results show that CNNs equipped with LESPS can well recover the target masks from corresponding point labels, and can achieve over 70% and 95% of their fully supervised performance in terms of pixel-level intersection over union (IoU) and object-level probability of detection (Pd), respectively. Code is available at https://github.com/XinyiYing/LESPS. Xinyi Ying, Li Liu 0002, Yingqian Wang 0002, Ruojing Li, Zaiping Lin, Weidong Sheng, Shilin Zhou 0001 |
CVPR | 1 |
| 2023 | Exploring Fine-Grained Sparsity in Convolutional Neural Networks for Efficient InferenceabstractNeural networks contain considerable redundant computation, which drags down the inference efficiency and hinders the deployment on resource-limited devices. In this paper, we study the sparsity in convolutional neural networks and propose a generic sparse mask mechanism to improve the inference efficiency of networks. Specifically, sparse masks are learned in both data and channel dimensions to dynamically localize and skip redundant computation at a fine-grained level. Based on our sparse mask mechanism, we develop SMPointSeg, SMSR, and SMStereo for point cloud semantic segmentation, single image super-resolution, and stereo matching tasks, respectively. It is demonstrated that our sparse masks are well compatible to different model components and network architectures to accurately localize redundant computation, with computational cost being significantly reduced for practical speedup. Extensive experiments show that our SMPointSeg, SMSR, and SMStereo achieve state-of-the-art performance on benchmark datasets in terms of both accuracy and efficiency. Longguang Wang, Yulan Guo, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Learning scalable dynamic filter in convolutional networks
Shuanglin Wu, Xinyi Ying, Longguang Wang, Jun-Gang Yang, Wei An 0003 |
Pattern Recognit. Lett. | 3 |
| 2023 | Incorporating Deep Background Prior Into Model-Based Method for Unsupervised Moving Vehicle Detection in Satellite VideosabstractBackground reconstruction is a key step of moving object detection in satellite videos. Most existing model-based methods exploit low-rank prior to recover background, which have achieved good performance but suffered degradation under complex and dynamic scenes. In this paper, we introduce a deep background prior into model-based methods for moving vehicle detection in satellite videos. Our deep background prior is obtained by a background reconstruction network, which can learn to reconstruct background from consecutive frames. By applying our deep background prior into model-based methods, a closed-form solution can be obtained via alternating direction method of multipliers (ADMM) and then detection results can be acquired through iterative optimization. More importantly, our background reconstruction network can be trained in an unsupervised way by introducing specifically designed loss, thus relieving the dependence on large-scale labeled dataset. Extensive experimental results demonstrate the efficiency and effectiveness of the proposed method. Ting Liu 0017, Xinyi Ying, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | DSFNet: Dynamic and Static Fusion Network for Moving Object Detection in Satellite VideosabstractMoving object detection (MOD) in satellite videos remains challenging due to the extremely small size of the interested targets and the highly complex background. Both the intra-frame (static) and inter-frame (dynamic) information are of great importance to MOD. In this letter, we propose a two-stream detection network named dynamic and static fusion network (DSFNet) to tackle the MOD problem in satellite videos. Specifically, the DSFNet is composed of a 2-D backbone to extract static context information from a single frame and a lightweight 3-D backbone to extract dynamic motion cues from consecutive frames. Then the extracted static and dynamic features are fused and fed into the detection head to detect the moving targets in satellite videos. We conduct extensive experiments on videos collected from Jilin-1 satellite and the results have demonstrated the effectiveness and robustness of the proposed DSFNet. Experimental results show that our DSFNet achieves the-state-of-the-art performance. Xinyi Ying, Ruojing Li, Shuanglin Wu, Li Liu 0002, Wei An 0003 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Exploring Sparsity in Image Super-Resolution for Efficient InferenceabstractCurrent CNN-based super-resolution (SR) methods process all locations equally with computational resources being uniformly assigned in space. However, since missing details in low-resolution (LR) images mainly exist in regions of edges and textures, less computational resources are required for those flat regions. Therefore, existing CNN-based methods involve redundant computation in flat regions, which increases their computational cost and limits their applications on mobile devices. In this paper, we explore the sparsity in image SR to improve inference efficiency of SR networks. Specifically, we develop a Sparse Mask SR (SMSR) network to learn sparse masks to prune redundant computation. Within our SMSR, spatial masks learn to identify "important" regions while channel masks learn to mark redundant channels in those "unimportant" regions. Consequently, redundant computation can be accurately localized and skipped while maintaining comparable performance. It is demonstrated that our SMSR achieves state-of-the-art performance with 41%/33%/27% FLOPs being reduced for ×2/3/4 SR. Code is available at: https://github.com/LongguangWang/SMSR. Longguang Wang, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003, Yulan Guo |
CVPR | 4 |
| 2021 | Light Field Image Super-Resolution Using Deformable ConvolutionabstractLight field (LF) cameras can record scenes from multiple perspectives, and thus introduce beneficial angular information for image super-resolution (SR). However, it is challenging to incorporate angular information due to disparities among LF images. In this paper, we propose a deformable convolution network (i.e., LF-DFnet) to handle the disparity problem for LF image SR. Specifically, we design an angular deformable alignment module (ADAM) for feature-level alignment. Based on ADAM, we further propose a collect-and-distribute approach to perform bidirectional alignment between the center-view feature and each side-view feature. Using our approach, angular information can be well incorporated and encoded into features of each view, which benefits the SR reconstruction of all LF images. Moreover, we develop a baseline-adjustable LF dataset to evaluate SR performance under different disparity variations. Experiments on both public and our self-developed datasets have demonstrated the superiority of our method. Our LF-DFnet can generate high-resolution images with more faithful details and achieve state-of-the-art reconstruction accuracy. Besides, our LF-DFnet is more robust to disparity variations, which has not been well addressed in literature. Yingqian Wang 0002, Jun-Gang Yang, Longguang Wang, Xinyi Ying, Tianhao Wu 0014, Wei An 0003, Yulan Guo |
IEEE Trans. Image Process. | 4 |
| 2020 | A Stereo Attention Module for Stereo Image Super-ResolutionabstractIn stereo image super-resolution (SR), exploiting both intra-view and cross-view information is significant but challenging. As existing single image SR (SISR) methods are powerful in intra-view information exploitation, in this letter, we propose a generic stereo attention module (SAM) to extend arbitrary SISR networks for stereo image SR. Specifically, we apply two identical pretrained SISR networks to stereo images. The extracted stereo features at different stages are fed to SAMs to interact cross-view information. Finally, the intra-view and cross-view information is incorporated by SISR networks for stereo image SR. Experiments on the KITTI2012, KITTI2015 and Middlebury datasets have demonstrated the effectiveness of our scheme. Using SAM, we can exploit cross-view information while maintaining the superiority of intra-view information exploitation, resulting in notable performance gain to SISR networks. Moreover, SRResNet equipped with our SAM outperforms the state-of-the-art stereo SR methods. Source code is available at https://github.com/XinyiYing/SAM. Xinyi Ying, Yingqian Wang 0002, Longguang Wang, Weidong Sheng, Wei An 0003, Yulan Guo |
IEEE Signal Process. Lett. | 1 |
| 2020 | Deformable 3D Convolution for Video Super-ResolutionabstractThe spatio-temporal information among video sequences is significant for video super-resolution (SR). However, the spatio-temporal information cannot be fully used by existing video SR methods since spatial feature extraction and temporal motion compensation are usually performed sequentially. In this paper, we propose a deformable 3D convolution network (D3Dnet) to incorporate spatio-temporal information from both spatial and temporal dimensions for video SR. Specifically, we introduce deformable 3D convolution (D3D) to integrate deformable convolution with 3D convolution, obtaining both superior spatio-temporal modeling capability and motion-aware modeling flexibility. Extensive experiments have demonstrated the effectiveness of D3D in exploiting spatio-temporal information. Comparative results show that our network achieves state-of-the-art SR performance. Code is available at: https://github.com/XinyiYing/D3Dnet. Xinyi Ying, Longguang Wang, Yingqian Wang 0002, Weidong Sheng, Wei An 0003, Yulan Guo |
IEEE Signal Process. Lett. | 1 |