EDBT 2026 Demo / reviewers in the wild / expert
Yaoqi Sun
dblp:226/6091
· DBLP profile ↗
52ranked-venue papers
5as first author
49since 2021 · last 2026
0000-0001-8874-241XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 1 first-author · 27 since 2021Artificial intelligence and machine learning · 24 · 4 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | 2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information PropagationabstractDespite recent progress in adapting State Space Models such as Mamba to vision tasks, their intrinsic 1D scanning mechanism imposes limitations when applied to inherently 2D-structured data like images. Existing adaptations, including VMamba and 2DMamba, either suffer from inconsistency between scanning order and spatial locality or restrict inter-patch communication to singular paths, hindering effective information propagation. In this paper, we propose 2D-CrossScan, a novel 2D-compatible scan framework that enables spatially consistent, multi-path hidden state propagation by integrating modified state equations over two-dimensional neighborhoods. Furthermore, we mitigate redundant information accumulation due to overlapping paths via cross-directional subtraction. To fully align with the 2D spatial structure, we introduce a multi-directional scanning strategy that starts simultaneously from all four corners of the image, enabling diverse propagation paths and better feature integration. Our approach maintains efficiency, requiring only minimal architectural changes to existing Mamba variants. Experimental results demonstrate substantial improvements in multiple visual tasks, including object detection and semantic segmentation on PANDA and COCO datasets. Compared to baseline SSM-based methods, 2D-CrossScan consistently yields better spatial representations, as confirmed by extensive effective receptive field visualizations and attention analyses. These results highlight the importance of geometry-aware state propagation and validate 2D-CrossScan as a simple yet powerful extension to SSMs for vision. Longlong Yu 0001, Wenxi Li, Yaoqi Sun, Chenggang Yan 0001 |
AAAI | 3 |
| 2026 | Temporal Calibrating and Distilling for Scene-Text Aware Text-Video RetrievalabstractExisting text-video retrieval methods mainly focus on singlemodal video content (i.e., visual entities), often overlooking heterogeneous scene text that is ubiquitous in human environments. Although scene text in videos provides finegrained semantics for cross-modal retrieval, effectively utilizing it presents two key challenges: (1) Temporally dense scene text disrupts sync with sparse video frames, obstructing video understanding;(2) Redundant scene text and irrelevant video frames hinder the learning of discriminative temporal clues for retrieval. To address them, we propose a temporal scene-text calibrating and distilling (TCD) network for textvideo retrieval. Specifically, we first design a window-OCR captioner that aggregates dense scene text into OCR captions to facilitate feature interaction. Next, we devise a heterogeneous semantics calibration module that leverages scene text as a self-supervised signal to temporally align window-level OCR captions and frame-level video features. Further, we introduce a context-guided temporal clue distillation module to learn the complementary and relevant details between scene text and video modalities, thereby obtaining discriminative temporal clues for retrieval. Extensive experiments show that our TCD achieves state-of-the-art performance on three scene-text related benchmarks. Zhiqian Zhao, Liang Li 0003, Xichun Sheng, Yaoqi Sun, Fang Kang, Chenggang Yan 0001 |
AAAI | 5 |
| 2026 | Multi-annotation agreement and prediction consistency networks: Improving semi-supervised segmentation of medical images with ambiguous boundaries
Shuai Wang 0003, Tengjin Weng, Yang Shen 0011, Zhidong Zhao, Yixiu Liu, Pengfei Jiao, Zhiming Cheng, Yaoqi Sun, Yaqi Wang 0002 |
Artif. Intell. Medicine | 10 |
| 2026 | Feature boosting and scale-aware network with multi-modal information for underwater salient object detection
Tingyu Wang 0002, Junzhe Lu 0002, Bin Wan, Rongfeng Lu, Yaoqi Sun, Duanpo Wu, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Lightweight multi-scale weight pruning network for salient object detectionabstractSalient object detection (SOD) is fundamental to computer vision, yet deep learning approaches often suffer from high computational costs, limiting deployment on resource-constrained devices. We propose a Lightweight Multi-scale Weight Pruning Network (LMWP-Net) to balance high performance with low complexity. LMWP-Net employs an encoder–decoder architecture featuring two key components: a Multi-scale Weight Pruning Module (MWPM) for efficient feature extraction and redundancy reduction, and a Multi-scale Attention Fusion Module (MAFM) for effective integration via attention mechanisms. Extensive experiments on public datasets demonstrate that LMWP-Net consistently outperforms existing lightweight methods and achieves competitive accuracy against state-of-the-art models. Remarkably, compared to the prominent BANet, LMWP-Net achieves a 94.6% reduction in parameters and a 99.5% reduction in FLOPs, validating its superior efficiency and effectiveness for real-time applications. The implemented code is publicly available at https://github.com/IMOP-lab/LMWP-Net . Xichun Sheng, Yaoqi Sun, Gaopeng Huang, Ya-Hong Chen, Jin Liu 0025, Xiaoshuai Zhang, Xingru Huang |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | TriFTM-Net: Tri-Path Fourier-Temporal Modulation Network for macular edema pathology segmentation and reconstruction in high-precision intraoperative navigationabstractOphthalmic diseases such significantly impair the vision of numerous individuals globally. Accurate and real-time 3D reconstruction of macular edema and retinal tears is crucial for improving surgical efficiency and success rates. However, lesion areas often exhibit considerable noise and high heterogeneity, and the imaging devices employed may introduce electronic noise and artifacts. Current 2D medical image segmentation techniques fail to achieve optimal outcomes. To overcome these challenges, we propose the Tri-Path Fourier-Temporal Modulation Network (TriFTM-Net). TriFTM-Net synergistically integrates spatial, frequency, and spatiotemporal features. This design effectively augments both feature representation and extraction. TriFTM-Net comprises three critical modules: the Tri-Path Spectral Hierarchical Encoder (TPSHE), which amplifies feature representation by integrating tri-path features; the Feature Re-Modulation (FRM), which reduces noise interference and enhances feature extraction; and the Hierarchical Feature Reconstruction Module (HFRM), which improves detail preservation in upsampled images. Comparative analysis with thirteen baseline methods demonstrates that our approach achieves the highest Dice scores, IoU, and Kappa coefficient on the OIMHS dataset.Our code is publicly available at https://github.com/IMOP-lab/TriFTM-Net. Xingru Huang, Shuaibin Chen, Gaopeng Huang, Zhaoyang Xu, Wenbin Zhang 0002, Jian Huang 0015, Jin Liu 0025, Xiaoshuai Zhang, Shaowei Jiang, Huiyu Zhou 0001, Yaoqi Sun |
Neural Networks | 15 |
| 2026 | DMDNet: Dual-branch multi-modal deep fusion network for V-D-T salient object detection
Yaoqi Sun, Bin Wan, Haibing Yin, Yahong Chen |
Neural Networks | 1 |
| 2026 | Spatial-Temporal Clue Reasoning Chain for Long Video Question Answering
Haibo Gong, Chenggang Yan 0001, Yaoqi Sun, Liang Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Empirical Study on Fusion Strategy in RGB-T Salient Object DetectionabstractIn the research field of RGB-Thermal saliency object detection (RGB-T SOD), the effective exploitation of the complementary characteristics of the two modalities represents a major challenge for enhancing detection performance. Current fusion methodologies can be roughly classified into early fusion and middle fusion strategies, with prevalent techniques primarily encompassing concatenation, summation, and multiplication of the two modalities. To in depth assess the efficacy of these fusion strategies, we took an empirical investigation on them. Our findings demonstrate that the concatenation of middle features constitutes a more advantageous fusion strategy, yielding superior performance and demonstrating enhanced stability. Furthermore, observing the unique properties of thermal (T) images, we introduced gamma correction as a novel data augmentation methodology to RGB-T SOD. We subsequently evaluated the responses across varying correction parameter ranges, revealing that while the response to this data augmentation technique differs across various models, data augmentation is found to be effective in general. Building upon these findings, we proposed the Gamma Correction Network (GaCNet). Specifically, we also integrated image pyramid mechanism in a lightweight manner, which facilitates a more effective recovery of fine-grained image details. Significant improvement was achieved on commonly used RGB-T testing datasets, especially in VT821 dataset, manifesting the effectiveness of our method. Shuai Wang 0003, Qiang Zhao 0005, Junbo Ma, Xichun Sheng, Yaoqi Sun, Hongfa Wen, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | LLFeat: Noise-Aware Feature Matching Under Various Low-Light Conditions
Longjian Zeng, Zunjie Zhu, Ming Lu 0002, Bolun Zheng, Rongfeng Lu, Tingyu Wang 0002, Zhongtian Zheng, Yaoqi Sun, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Prompt Learning With Knowledge Regularization for Pre-Trained Vision-Language ModelsabstractPrompt learning is an effective way to adapt pre-trained models to downstream tasks by training a small number of additional learnable prompts. Recent studies address several early challenges by combining generalized knowledge from frozen pre-trained VL models with task-specific knowledge from training data as guidance for prompt learning. However, existing methods still struggle with the generalization-adaptation (GA) trade-off dilemma: excessive reliance on generalized knowledge hinders adaptation to downstream tasks, while overemphasis on task-specific knowledge undermines the inherent generalization capabilities of pre-trained models. To address this issue, we propose a novel prompt learning method called Prompt Learning with Knowledge Regularization (PLKR). PLKR effectively mitigates the GA trade-off dilemma by offering greater flexibility in adapting to task-specific knowledge while minimizing the disruption of pre-trained knowledge. Specifically, we propose category-invariant and topology-invariant knowledge regularization to preserve generalized knowledge: the former enhances category-level discriminative capabilities while allowing flexible task-specific learning, and the latter maintains global topological stability during adaptation to new tasks. Through the proposed regularization, PLKR improves the performance on both base and new tasks. We evaluate the effectiveness of our approach on four representative tasks over 11 datasets. Experimental results show our method outperforms existing SOTA methods by a large margin. Boyang Guo, Liang Li 0003, Yaoqi Sun, Chenggang Yan 0001, Xichun Sheng |
IEEE Trans. Multim. | 4 |
| 2026 | Hybrid Debiasing Transformer With Adaptive Regularization for Video Moment LocalizationabstractVideo Moment Localization (VML) is a task that seeks to pinpoint the most pertinent segment within an untrimmed video using a linguistic query. Previous works expose the severe data bias issues in VML and note that models avoid understanding visual-textual content by adapting the timestamp distribution. The work investigates data biases from both intrinsic and extrinsic perspectives: The former arises primarily from moment boundary ambiguity and the inputoutput information imbalance. The latter is attributed to the longtail distribution and the semantic bias resulting from the limited tail samples. To reduce the issues, we develop a hybrid multimodal debiasing network with a temporal consistency constraint for VML. Firstly, we propose a multi-temporal Transformer to alleviate boundary ambiguity by merging frame-wise features into segment-wise representations and dynamically aligning with moment boundaries. Subsequently, we implement a temporal consistency constraint to accentuate action information from complex moment context and mitigate the intrinsic bias caused by information imbalance. Moreover, we develop a hybrid linguistic activation module to mitigate the long-tail bias, which offers prior guidance to emphasize distinguishing clues from tail samples. Additionally, we introduce the prior-guided Transformer to alleviate the semantic bias by learning the global semantics of sentences, thereby circumventing the tail-sample overfitting issue. Comprehensive experiments demonstrate the efficacy of our proposed method across three datasets. Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Hongkui Wang, Yaoqi Sun, Zunjie Zhu |
IEEE Trans. Multim. | 6 |
| 2025 | SdalsNet: Self-Distilled Attention Localization and Shift Network for Unsupervised Camouflaged Object DetectionabstractUnsupervised camouflaged object detection (UCOD) poses significant challenges, primarily attributed to the absence of human labels. Existing UCOD methodologies, leveraging attention mechanisms, often struggle to achieve precise localization of camouflaged objects. To overcome this limitation, we introduce a groundbreaking fully unsupervised algorithm for attention-guided camouflaged object localization, shift, and inference, termed the self-distilled attention localization and shift network (SdalsNet). In this study, we formulate an attention localization methodology aimed at accurately identifying the central coordinate of the camouflaged object. Furthermore, we propose four distinct loss functions tailored to refine the precision of attentional positioning. These loss functions effectively constrain the distances between three types of class tokens, facilitating seamless attentional shifting across the input sample. Additionally, we design a sophisticated prediction inference technique to reconstruct the binary output of an attention map, thereby providing a comprehensive understanding of the detected camouflaged objects. Experimental results on four challenging COD benchmark datasets corroborate the effectiveness of our proposed approach, demonstrating notable superiority over state-of-the-art methods. Peiyao Shou, Yixiu Liu, Wei Wang 0335, Yaoqi Sun, Zhigao Zheng 0001, Shangdong Zhu, Chenggang Yan 0001 |
AAAI | 4 |
| 2025 | Heterogeneous Prompt-Guided Entity Inferring and Distilling for Scene-Text Aware Cross-Modal RetrievalabstractIn cross-modal retrieval, comprehensive image understanding is vital while the scene text in images can provide fine-grained information to understand visual semantics. Current methods fail to make full use of scene text. They suffer from the semantic ambiguity of independent scene text and overlook the heterogeneous concepts in image-caption pairs. In this paper, we propose a heterogeneous prompt-guided entity inferring and distilling (HOPID) network to explore the nature connection of scene text in images and captions and learn a property-centric scene text representation. Specifically, we propose to align scene text in images and captions via heterogeneous prompt, which consists of visual and text prompt. For text prompt, we introduce the discriminative entity inferring module to reason key scene text words from captions, while visual prompt highlights the corresponding scene text in images. Furthermore, to secure a robust scene text representation, we design a perceptive entity distilling module that distills the beneficial information of scene text at a fine-grained level. Extensive experiments show that the proposed method significantly outperforms existing approaches on two public cross-modal retrieval benchmarks. Zhiqian Zhao, Liang Li 0003, Yaoqi Sun, Xichun Sheng, Haibing Yin, Shaowei Jiang |
AAAI | 4 |
| 2025 | Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain AdaptationabstractThis paper explores the Class-Incremental Source-Free Unsupervised Domain Adaptation (CI-SFUDA) problem, where the unlabeled target data come incrementally without access to labeled source instances. This problem poses two challenges, the interference of similar source-class knowledge in target-class representation learning and the shocks of new target knowledge to old ones. To address them, we propose the Multi-Granularity Class Prototype Topology Distillation (GROTO) algorithm, which effectively transfers the source knowledge to the class-incremental target domain. Concretely, we design the multi-granularity class prototype self-organization module and the prototype topology distillation module. First, we mine the positive classes by modeling accumulation distributions. Next, we introduce multi-granularity class prototypes to generate reliable pseudo-labels, and exploit them to promote the positive-class target feature self-organization. Second, the positive-class prototypes are leveraged to construct the topological structures of source and target feature spaces. Then, we perform the topology distillation to continually mitigate the shocks of new target knowledge to old ones. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on three public datasets. Peihua Deng, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Ying Fu 0001, Liang Li 0003 |
CVPR | 5 |
| 2025 | Debiased Teacher for Day-to-Night Domain Adaptive Object Detection
Liang Li 0003, Haibing Yin, Yaoqi Sun, Chenggang Yan 0001 |
ICCV | 5 |
| 2025 | Hypergraph-Guided Federated Distillation Learning for Efficient and Robust Multi-center fMRI Data Analysis
Yidan Xu, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Xiangmin Han, Yue Gao 0002 |
MICCAI (11) | 6 |
| 2025 | VGNC: Reducing the Overfitting of Sparse-view 3DGS via Validation-guided Gaussian Number ControlabstractSparse-view 3D reconstruction is a fundamental yet challenging task in practical 3D reconstruction applications. Recently, many methods based on 3D Gaussian Splatting (3DGS) have been proposed to address sparse-view 3D reconstruction. Although these methods have made considerable advancements, they still show significant issues with overfitting. To reduce the overfitting, we introduce VGNC, a novel Validation-guided Gaussian Number Control approach based on generative novel view synthesis (NVS) models. To the best of our knowledge, this is the first attempt to alleviate the overfitting issue of sparse-view 3DGS with generative validation images. Specifically, we first introduce a validation image generation method based on a generative NVS model. We then propose a Gaussian number control strategy that utilizes generated validation images to determine optimal Gaussian numbers, thereby reducing the issue of overfitting. We conducted detailed experiments on various sparse-view 3DGS baselines and datasets to evaluate the effectiveness of VGNC. Extensive experiments show that our approach not only reduces overfitting but also improves rendering quality on the test set while decreasing the number of Gaussians. This reduction lowers storage demands and accelerates both training and rendering. Our code is available at: https://github.com/LinLif1869/VGNC. Rongfeng Lu, Haofan Ren, Ming Lu 0002, Yaoqi Sun, Chenggang Yan 0001, Anke Xue |
ACM Multimedia | 6 |
| 2025 | Improving Integrated Satellite-Terrestrial Cell-Free Massive MIMO Systems by Rate-Splitting Multiple AccessabstractWe investigate the spectral and energy efficiencies of the uplink in an integrated satellite-terrestrial cell-free massive multiple-input multiple-output (IST-CF-mMIMO) system assisted by rate-splitting multiple access (RSMA). In the IST-CF-mMIMO system, the terrestrial users employ RSMA to transmit a message as a superposition of two parts with different power to the terrestrial access points and low-Earth-orbit satellite. Taking realistic conditions such as the spatially correlated Ricean fading channels, imperfect channel knowledge, and successive interference cancellation into account, we derive rigorous closed-form expressions for uplink achievable spectral and energy efficiencies and evaluate these performance metrics across a range of system configurations. Additionally, to enhance the system energy efficiency, we formulate the design of users’ power control coefficients as an energy efficiency optimization problem and design an efficient algorithm based on Lagrangian dual transformation and quadratic transformation techniques to solve it. Comprehensive simulations validate our theoretical propositions and evaluate the efficacy of the proposed energy efficiency maximization algorithm. Yao Zhang 0016, Jintao Shen, Yaoqi Sun, Xichun Sheng, Haitao Zhao 0004, Hongbo Zhu 0002 |
IEEE Internet Things J. | 5 |
| 2025 | Lightweight three-stream encoder-decoder network for multi-modal salient object detection
Junzhe Lu 0002, Tingyu Wang 0002, Bin Wan, Qiang Zhao 0005, Shuai Wang 0003, Yaoqi Sun, Yang Zhou 0052, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2025 | Multidimensional Directionality-Enhanced Segmentation via large vision model
Xingru Huang, Changpeng Yue, Jian Huang 0015, Zhengyao Jiang, Mingkuan Wang, Zhaoyang Xu, Guangyuan Zhang, Jin Liu 0025, Tianyun Zhang, Xiaoshuai Zhang, Shaowei Jiang, Yaoqi Sun |
Medical Image Anal. | 15 |
| 2025 | Progressive Decision Boundary Shifting for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) is attracting more attention from researchers for boosting the task-specific generalization on target domain. It focuses on addressing the domain shift between the labeled source domain and the unlabeled target domain. Recent biclassifier-based UDA models perform category-level alignment to reduce domain shift, and meanwhile, self-training is used for improving the discriminability of target instances. However, the error accumulation problem of instances with high semantic uncertainty may cause discriminability degradation and category-level misalignment. To solve this issue, we design the progressive decision boundary shifting algorithm, where stable category information of target instances is explored for learning a discriminability structure on target domain. Specifically, we first model the semantic uncertainty of instances by progressively shifting decision boundaries of category. Then, we introduce the uncertainty decoupling in a contrastive manner, where the discriminative information is learned from the source domain for instance with low semantic uncertainty. Furthermore, we minimize the predictive entropy of instances with high semantic uncertainty to reduce their prediction confidence. Extensive experiments on three popular datasets show that our model outperforms the current state-of-the-art (SOTA) UDA methods. Liang Li 0003, Tongyu Lu, Yaoqi Sun, Chenggang Yan 0001, Qingming Huang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Domain Shared and Specific Prompt Learning for Incremental Monocular Depth Estimation
Zhiwen Yang 0003, Liang Li 0003, Tingyu Wang 0002, Yaoqi Sun, Chenggang Yan 0001 |
ACM Multimedia | 5 |
| 2024 | Upping the Game: How 2D U-Net Skip Connections Flip 3D SegmentationabstractIn the present study, we introduce an innovative structure for 3D medical image segmentation that effectively integrates 2D U-Net-derived skip connections into the architecture of 3D convolutional neural networks (3D CNNs). Conventional 3D segmentation techniques predominantly depend on isotropic 3D convolutions for the extraction of volumetric features, which frequently engenders inefficiencies due to the varying information density across the three orthogonal axes in medical imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI). This disparity leads to a decline in axial-slice plane feature extraction efficiency, with slice plane features being comparatively underutilized relative to features in the time-axial. To address this issue, we introduce the U-shaped Connection (uC), utilizing simplified 2D U-Net in place of standard skip connections to augment the extraction of the axial-slice plane features while concurrently preserving the volumetric context afforded by 3D convolutions. Based on uC, we further present uC 3DU-Net, an enhanced 3D U-Net backbone that integrates the uC approach to facilitate optimal axial-slice plane feature utilization. Through rigorous experimental validation on five publicly accessible datasets—FLARE2021, OIMHS, FeTA2021, AbdomenCT-1K, and BTCV, the proposed method surpasses contemporary state-of-the-art models. Notably, this performance is achieved while reducing the number of parameters and computational complexity. This investigation underscores the efficacy of incorporating 2D convolutions within the framework of 3D CNNs to overcome the intrinsic limitations of volumetric segmentation, thereby potentially expanding the frontiers of medical image analysis. Our implementation is available at https://github.com/IMOP-lab/U-Shaped-Connection. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Shaowei Jiang, Yaoqi Sun |
NeurIPS | 7 |
| 2024 | Enhanced local distribution learning for real image super-resolution
Yaoqi Sun, Aiai Huang, Chenggang Yan 0001, Bolun Zheng |
Comput. Vis. Image Underst. | 1 |
| 2024 | ADNet: Anti-noise dual-branch network for road defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | On the Performance of Cell-Free IoT Systems With RSMA and Downlink TrainingabstractThis letter establishes a novel transmission framework that amalgamates downlink (DL) training with rate-splitting multiple access, thereby being expected to enhance the spectral efficiency (SE) of a cell-free massive multiple-input multiple-output enabled Internet of Things (IoT) system. Considering a correlated Ricean fading environment coupled with imperfect channel knowledge, we derive a closed-form expression for the achievable SE and evaluate the DL SE under a variety of system configurations. Our comprehensive simulations corroborate the theoretical findings and yield critical insights pertinent to the system’s architectural design. Yao Zhang 0016, Wenchao Xia, Haitao Zhao 0004, Yaoqi Sun, Hongkui Wang, Hongbo Zhu 0002 |
IEEE Internet Things J. | 5 |
| 2024 | Dynamic interactive refinement network for camouflaged object detection
Yaoqi Sun, Lidong Ma, Peiyao Shou, Hongfa Wen, Yixiu Liu, Chenggang Yan 0001, Haibing Yin |
Neural Comput. Appl. | 1 |
| 2024 | TMNet: Triple-modal interaction encoder and multi-scale fusion decoder network for V-D-T salient object detection
Bin Wan, Chengtao Lv, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Hongkui Wang, Chenggang Yan 0001 |
Pattern Recognit. | 4 |
| 2024 | Multiple-environment Self-adaptive Network for aerial-view geo-localization
Tingyu Wang 0002, Zhedong Zheng, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001, Tat-Seng Chua |
Pattern Recognit. | 3 |
| 2024 | Rethinking Pooling for Multi-Granularity Features in Aerial-View Geo-LocalizationabstractVision-based aerial-view geo-localization aims to match drone- and satellite-views of the same geographical location. Several feature partition strategies divide spatial features to mine contextual information. However, the compression from fine-grained features to visual descriptors is ill-considered, that is, classical pooling destroys discriminative features while increasing the sensitivity of networks to contextual information. In order to clarify this, we first review existing pooling layer and analyze their pros and cons when applied in feature compression. Inspired by the appearance of aerial views, we then summarize an ideal feature compression operation, i.e., precisely highlighting the central target while maximizing the use of environmental information in a feature-smoothing manner. To achieve the above process, we propose a distance-dependent parameter initialization strategy and form a novel pooling called$D^{2}$-GeM pooling, which can explicitly guide the network to compress fine-grained features in multiple patterns. Extensive experiments on public benchmark University-1652 substantiate that our strategy attains more appealing results without additional costs. Tingyu Wang 0002, Yaoqi Sun, Chenggang Yan 0001 |
IEEE Signal Process. Lett. | 4 |
| 2024 | SDPL: Shifting-Dense Partition Learning for UAV-View Geo-LocalizationabstractCross-view geo-localization aims to match images of the same target from different platforms, e.g., drone and satellite. It is a challenging task due to the changing appearance of targets and environmental content from different views. Most methods focus on obtaining more comprehensive information through feature map segmentation, while inevitably destroying the image structure, and are sensitive to the shifting and scale of the target in the query. To address the above issues, we introduce simple yet effective part-based representation learning, shifting-dense partition learning (SDPL). We propose a dense partition strategy (DPS), dividing the image into multiple parts to explore contextual information while explicitly maintaining the global structure. To handle scenarios with non-centered targets, we further propose the shifting-fusion strategy, which generates multiple sets of parts in parallel based on various segmentation centers, and then adaptively fuses all features to integrate their anti-offset ability. Extensive experiments show that SDPL is robust to position shifting, and performs competitively on two prevailing benchmarks, University-1652 and SUES-200. In addition, SDPL shows satisfactory compatibility with a variety of backbone networks (e.g., ResNet and Swin).https://github.com/C-water/SDPL_release. Tingyu Wang 0002, Haoran Li 0025, Rongfeng Lu, Yaoqi Sun, Bolun Zheng, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Learning Cross-View Geo-Localization Embeddings via Dynamic Weighted Decorrelation RegularizationabstractIn the domain of cross-view geo-localization, the challenge lies in accurately matching images captured from distinct perspectives, such as aerial drone imagery and satellite imagery of the same geographical location. Existing methods predominantly concentrate on minimizing distances between feature embeddings in the representational space, inadvertently overlooking the significance of reducing embedding redundancy. This oversight potentially hampers the extraction of diverse and distinctive visual patterns critical for precise localization. This work argues that minimizing embedding redundancy is a pivotal factor in enhancing a model’s ability to discriminate diverse scene characteristics. To support this claim, we introduce a straightforward yet effective regularization technique, termed dynamic weighted decorrelation regularization (DWDR). DWDR serves to actively promote the learning of orthogonal feature channels within neural networks. By dynamically adjusting weights, DWDR targets the minimization of interchannel correlations, guiding the correlation matrix toward diagonality, indicative of independence among channels. The dynamic weighting mechanism adaptively prioritizes the decorrelation of channels that remain highly correlated throughout training. Additionally, we devise a symmetrical sampling strategy for cross-view scenarios to ensure that the training examples are balanced across different imaging platforms in a batch. Despite its simplicity, the integration of DWDR and the proposed sampling scheme yields remarkable performance across four extensive benchmark datasets: University-1652, CVUSA, CVACT, and VIGOR. Notably, in stringent conditions, such as when constrained to exceedingly compact feature dimensions of 64, our methodology significantly outperforms conventional baselines, thereby affirming its efficacy and robustness under challenging constraints. Tingyu Wang 0002, Zhedong Zheng, Zunjie Zhu, Yaoqi Sun, Chenggang Yan 0001, Yi Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Quality-Aware Selective Fusion Network for V-D-T Salient Object DetectionabstractDepth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often unreliable in some challenging scenarios, which will result in the performance degradation of the two-modal based salient object detection (SOD). Meanwhile, some researchers pay attention to the triple-modal SOD task, namely the visible-depth-thermal (VDT) SOD, where they attempt to explore the complementarity of the RGB image, the depth image, and the thermal image. However, existing triple-modal SOD methods fail to perceive the quality of depth maps and thermal images, which leads to performance degradation when dealing with scenes with low-quality depth and thermal images. Therefore, in this paper, we propose a quality-aware selective fusion network (QSF-Net) to conduct VDT salient object detection, which contains three subnets including the initial feature extraction subnet, the quality-aware region selection subnet, and the region-guided selective fusion subnet. Firstly, except for extracting features, the initial feature extraction subnet can generate a preliminary prediction map from each modality via a shrinkage pyramid architecture, which is equipped with the multi-scale fusion (MSF) module. Then, we design the weakly-supervised quality-aware region selection subnet to generate the quality-aware maps. Concretely, we first find the high-quality and low-quality regions by using the preliminary predictions, which further constitute the pseudo label that can be used to train this subnet. Finally, the region-guided selective fusion subnet purifies the initial features under the guidance of the quality-aware maps, and then fuses the triple-modal features and refines the edge details of prediction maps through the intra-modality and inter-modality attention (IIA) module and the edge refinement (ER) module, respectively. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art methods with a large margin. Our code and results are available at https://github.com/Lx-Bao/QSFNet. Liuxin Bao, Xiaofei Zhou 0003, Xiankai Lu, Yaoqi Sun, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | MFFNet: Multi-Modal Feature Fusion Network for V-D-T Salient Object DetectionabstractThis article discusses the limitations of single- and two-modal salient object detection (SOD) methods and the emergence of multi-modal SOD techniques that integrate Visible, Depth, or Thermal information. However, current multi-modal methods often rely on simple fusion techniques such as addition, multiplication and concatenation, to combine the different modalities, which is ineffective for challenging scenes, such as low illumination and background messy. To address this issue, we propose a novel multi-modal feature fusion network (MFFNet) for V-D-T salient object detection, where the two key points are the triple-modal deep fusion encoder and the progressive feature enhancement decoder. The MFFNet's triple-modal deep fusion (TDF) module is designed to integrate the features of the three modalities and explore their complementarity by utilizing mutual optimization during the encoding phase. In addition, the progressive feature enhancement decoder consists of the weighted context-enhanced feature (WCF) module, region optimization (RO) module and boundary perception (BP) module to produce region-aware and contour-aware features. After that, a multi-scale fusion (MF) module is proposed to integrate these features and generate high-quality saliency maps. We conduct extensive experiments on the VDT-2048 dataset, and our results show that the proposed MFFNet outperforms 12 state-of-the-art multi-modal methods. Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | GFNet: gated fusion network for video saliency prediction
Songhe Wu, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Appl. Intell. | 3 |
| 2023 | AutoDeconJ: a GPU-accelerated ImageJ plugin for 3D light-field deconvolution with optimal iteration numbers predictingabstractMOTIVATION: Light-field microscopy (LFM) is a compact solution to high-speed 3D fluorescence imaging. Usually, we need to do 3D deconvolution to the captured raw data. Although there are deep neural network methods that can accelerate the reconstruction process, the model is not universally applicable for all system parameters. Here, we develop AutoDeconJ, a GPU-accelerated ImageJ plugin for 4.4× faster and more accurate deconvolution of LFM data. We further propose an image quality metric for the deconvolution process, aiding in automatically determining the optimal number of iterations with higher reconstruction accuracy and fewer artifacts. RESULTS: Our proposed method outperforms state-of-the-art light-field deconvolution methods in reconstruction time and optimal iteration numbers prediction capability. It shows better universality of different light-field point spread function (PSF) parameters than the deep learning method. The fast, accurate and general reconstruction performance for different PSF parameters suggests its potential for mass 3D reconstruction of LFM data. AVAILABILITY AND IMPLEMENTATION: The codes, the documentation and example data are available on an open source at: https://github.com/Onetism/AutoDeconJ.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Changqing Su, Yaoqi Sun, Chenggang Yan 0001, Haibing Yin |
Bioinform. | 4 |
| 2023 | SMINet: Semantics-aware multi-level feature interaction network for surface defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Haibing Yin, Ji Hu 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | CANet: Context-aware Aggregation Network for Salient Object Detection of Surface Defects
Bin Wan, Xiaofei Zhou 0003, Mang Xiao, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | Depth-guided deep filtering network for efficient single image bokeh rendering
Bolun Zheng, Xiaofei Zhou 0003, Aiai Huang, Yaoqi Sun, Chuqiao Chen, Chenggang Yan 0001, Shanxin Yuan |
Neural Comput. Appl. | 5 |
| 2022 | Gait Recognition in the Wild with Multi-hop Temporal SwitchabstractExisting studies for gait recognition are dominated by in-the-lab scenarios. Since people live in real-world senses, gait recognition in the wild is a more practical problem that has recently attracted the attention of the community of multimedia and computer vision. Current methods that obtain state-of-the-art performance on in-the-lab benchmarks achieve much worse accuracy on the recently proposed in-the-wild datasets because these methods can hardly model the varied temporal dynamics of gait sequences in unconstrained scenes. Therefore, this paper presents a novel multi-hop temporal switch method to achieve effective temporal modeling of gait patterns in real-world scenes. Concretely, we design a novel gait recognition network, named Multi-hop Temporal Switch Network (MTSGait), to learn spatial features and multi-scale temporal features simultaneously. Different from existing methods that use 3D convolutions for temporal modeling, our MTSGait models the temporal dynamics of gait sequences by 2D convolutions. By this means, it achieves high efficiency with fewer model parameters and reduces the difficulty in optimization compared with 3D convolution-based models. Based on the specific design of the 2D convolution kernels, our method can eliminate the misalignment of features among adjacent frames. In addition, a new sampling strategy, i.e., non-cyclic continuous sampling, is proposed to make the model learn more robust temporal features. Finally, the proposed method achieves superior performance on two public gait in-the-wild datasets, i.e., GREW and Gait3D, compared with state-of-the-art methods. Jinkai Zheng, Xinchen Liu, Xiaoyan Gu 0001, Yaoqi Sun, Chuang Gan 0001, Jiyong Zhang 0001, Wu Liu 0005, Chenggang Yan 0001 |
ACM Multimedia | 4 |
| 2022 | Bidirectional difference locating and semantic consistency reasoning for change captioningabstractChange captioning is an emerging task to describe the changes between a pair of images. The difficulty in this task is to discover the differences between the two images. Recently, some methods have been proposed to address this problem. However, they all employ unidirectional difference localization to identify the changes. This can lead to ambiguity about the nature of the changes. Instead, we propose a framework with bidirectional difference localization and semantic consistency reasoning to describe the image changes. First, we locate the changes in the two images by capturing bidirectional differences. Then we design a decoder with spatial-channel attention to generate the change caption. Finally, we introduce semantic consistency reasoning to constrain our bidirectional difference localization module and spatial-channel attention module. Extensive experiments on three public data sets show that the performance of our proposed model outperforms the state-of-the-art change captioning models by a large margin. Yaoqi Sun, Liang Li 0003, Tongyv Lu, Bolun Zheng, Chenggang Yan 0001, Yongjun Bao, Guiguang Ding, Gregory Slabaugh |
Int. J. Intell. Syst. | 1 |
| 2022 | Learning Frequency Domain Priors for Image DemoireingabstractImage demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin. Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Each Part Matters: Local Patterns Facilitate Cross-View Geo-LocalizationabstractCross-view geo-localization is to spot images of the same geographic target from different platforms,e.g., drone-view cameras and satellites. It is challenging in the large visual appearance changes caused by extreme viewpoint variations. Existing methods usually concentrate on mining the fine-grained feature of the geographic target in the image center, but underestimate the contextual information in neighbor areas. In this work, we argue that neighbor areas can be leveraged as auxiliary information, enriching discriminative clues for geo-localization. Specifically, we introduce a simple and effective deep neural network, called Local Pattern Network (LPN), to take advantage of contextual information in an end-to-end manner. Without using extra part estimators, LPN adopts a square-ring feature partition strategy, which provides the attention according to the distance to the image center. It eases the part matching and enables the part-wise representation learning. Owing to the square-ring partition design, the proposed LPN has good scalability to rotation variations and achieves competitive results on three prevailing benchmarks,i.e., University-1652, CVUSA and CVACT. Besides, we also show the proposed LPN can be easily embedded into other frameworks to further boost performance. Tingyu Wang 0002, Zhedong Zheng, Chenggang Yan 0001, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng, Yi Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Age-Invariant Face Recognition by Multi-Feature Fusionand Decomposition with Self-attentionabstractDifferent from general face recognition, age-invariant face recognition (AIFR) aims at matching faces with a big age gap. Previous discriminative methods usually focus on decomposing facial feature into age-related and age-invariant components, which suffer from the loss of facial identity information. In this article, we propose a novel Multi-feature Fusion and Decomposition (MFD) framework for age-invariant face recognition, which learns more discriminative and robust features and reduces the intra-class variants. Specifically, we first sample multiple face images of different ages with the same identity as a face time sequence. Then, the multi-head attention is employed to capture contextual information from facial feature series, extracted by the backbone network. Next, we combine feature decomposition with fusion based on the face time sequence to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against the aging process. Besides, we also mitigate imbalanced age distribution in the training data by a re-weighted age loss. We experimented with the proposed MFD over the popular CACD and CACD-VS datasets, where we show that our approach improves the AIFR performance than previous state-of-the-art methods. We simultaneously show the performance of MFD on LFW dataset. Chenggang Yan 0001, Lixuan Meng, Liang Li 0003, Jian Yin 0003, Jiyong Zhang 0001, Yaoqi Sun, Bolun Zheng |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2021 | Heuristic Depth Estimation with Progressive Depth Reconstruction and Confidence-Aware LossabstractRecently deep learning-based depth estimation has shown the promising result, especially with the help of sparse depth reference samples. Existing works focus on directly inferring the depth information from sparse samples with high confidence. In this paper, we propose a Heuristic Depth Estimation Network (HDEN) with progressive depth reconstruction and confidence-aware loss. The HDEN leverages the reference samples with low confidence to distill the spatial geometric and local semantic information for dense depth prediction. Specifically, we first train a U-NET network to generate a coarse-level dense reference map. Second, the progressive depth reconstruction module successively reconstructs the fine-level dense depth map from different scales, where a multi-level upsampling block is designed to recover the local structure of object. Finally, the confidence-aware loss is proposed to trigger the reference samples with low confidence, which enforces the model focusing on estimating the depth of the tiny structure. Extensive experiments on the NYU-Depth-v2 and KITTI-Odometry dataset show the effectiveness of our method. Visualization results demonstrate that the dense depth maps generated by HDEN have better consistency at the entity edge with RGB image. Liang Li 0003, Chenggang Yan 0001, Yaoqi Sun, Tao Shen 0004, Jiyong Zhang 0001 |
ACM Multimedia | 4 |
| 2021 | Cross-modal semantic correlation learning by Bi-CNN networkabstractAbstract Cross modal retrieval can retrieve images through a text query and vice versa. In recent years, cross modal retrieval has attracted extensive attention. The purpose of most now available cross modal retrieval methods is to find a common subspace and maximize the different modal correlation. To generate specific representations consistent with cross modal tasks, this paper proposes a novel cross modal retrieval framework, which integrates feature learning and latent space embedding. In detail, we proposed a deep CNN and a shallow CNN to extract the feature of the samples. The deep CNN is used to extract the representation of images, and the shallow CNN uses a multi‐dimensional kernel to extract multi‐level semantic representation of text. Meanwhile, we enhance the semantic manifold by constructing cross modal ranking and within‐modal discriminant loss to improve the division of semantic representation. Moreover, the most representative samples are selected by using online sampling strategy, so that the approach can be implemented on a large‐scale data. This approach not only increases the discriminative ability among different categories, but also maximizes the relativity between different modalities. Experiments on three real word datasets show that the proposed method is superior to the popular methods. Liang Li 0003, Chenggang Yan 0001, Yaoqi Sun, Jiyong Zhang 0001 |
IET Image Process. | 5 |
| 2021 | Evolution of ICTs-empowered-identification: A general re-ranking method for person re-identification
Tongkun Xu, Bolun Zheng, Yaoqi Sun, Anan Liu, Zhendong Mao 0001, Chenggang Yan 0001 |
Pattern Recognit. Lett. | 5 |
| 2021 | Dynamic Selective Network for RGB-D Salient Object DetectionabstractRGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models. Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding |
IEEE Trans. Image Process. | 5 |
| 2020 | Real-World Automatic Makeup via Identity Preservation Makeup NetabstractThis paper focuses on the real-world automatic makeup problem. Given one non-makeup target image and one reference image, the automatic makeup is to generate one face image, which maintains the original identity with the makeup style in the reference image. In the real-world scenario, face makeup task demands a robust system against the environmental variants. The two main challenges in real-world face makeup could be summarized as follow: first, the background in real-world images is complicated. The previous methods are prone to change the style of background as well; second, the foreground faces are also easy to be affected. For instance, the ``heavy'' makeup may lose the discriminative information of the original identity. To address these two challenges, we introduce a new makeup model, called Identity Preservation Makeup Net (IPM-Net), which preserves not only the background but the critical patterns of the original identity. Specifically, we disentangle the face images to two different information codes, i.e., identity content code and makeup style code. When inference, we only need to change the makeup style code to generate various makeup images of the target person. In the experiment, we show the proposed method achieves not only better accuracy in both realism (FID) and diversity (LPIPS) in the test set, but also works well on the real-world images collected from the Internet. Zhikun Huang, Zhedong Zheng, Chenggang Yan 0001, Hongtao Xie 0001, Yaoqi Sun, Jiyong Zhang 0001 |
IJCAI | 5 |
| 2019 | Image classification base on PCA of multi-view deep representation
Yaoqi Sun, Liang Li 0003, Liang Zheng 0007, Ji Hu 0002, Wenchao Li 0004, Yatong Jiang, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2019 | Deep fusion based video saliency detection
Hongfa Wen, Xiaofei Zhou 0003, Yaoqi Sun, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 3 |