VLDB 2026 Research / reviewers in the wild / expert
Guoqing Wang 0001
dblp:17/356-1
· DBLP profile ↗
77ranked-venue papers
7as first author
71since 2021 · last 2026
0000-0002-3938-5994ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 4 first-author · 50 since 2021Artificial intelligence and machine learning · 24 · 2 first-author · 23 since 2021Security and privacy · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Beyond Illusion: Generalized and Efficient Mirror DetectionabstractReflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby identifying the mirror regions. However, the exploration of extended scenes with dynamic content changes is rarely investigated. Therefore, we propose the MirrorSAM designed for MD based on the Segment Anything Model (SAM). Specifically, due to the varying reflections produced by mirrors in different positions and the complex visual space that interferes with localization, we design the hierarchical mixture of direction experts (HMDE) in the low-rank space to reduce biases towards entities in SAM and dynamically adjust experts based on the input scene. We observe differences in depth between mirrors and adjacent areas, and propose the depth token calibration (DTC), which introduces a learnable depth token to generate the depth map and serve as an error correction factor. We further formulate the selective pixel-prototype contrastive (SPPC) loss, selecting partially confusable samples to promote the decoupling of mirror and non-mirror representations. Extensive experiments conducted on four mirror benchmarks and two settings demonstrate that our approach surpasses state-of-the-art methods with few trainable parameters and FLOPs. We further extend to four transparent surface benchmarks to validate generalization. Mingfeng Zha, Guoqing Wang 0001, Tianyu Li 0003, Wei Dong 0010, Peng Wang 0023, Yang Yang 0002 |
AAAI | 2 |
| 2026 | RIR-Agent: An interactive framework for effective and adaptive restoration of remote sensing imagery
Junyu Liu, Tianyu Li 0003, Lanyue Liang, Gang Fu 0003, Guoqing Wang 0001, Quan Rui, Xiongxin Tang, Shuyuan Zhu, Yang Yang 0002 |
Expert Syst. Appl. | 5 |
| 2026 | Think Twice Before Determining: Toward Scene-Aware Visual Reasoning for Mirror DetectionabstractMirror detection (MD) aims to overcome interference caused by reflections and locate mirror regions. Existing methods focus on designing components to explicitly establish the associations between physical entities and corresponding imagings, or utilizing rotation to construct symmetric consistency. We observe that: a) incomplete and incorrect correspondence between entities and imagings; b) other physical materials (e.g., glass) exhibit characteristics partially similar to mirrors, causing confusion when they co-occur; c) complex interfering factors (e.g., occlusion) and reflection mechanisms may expand vector space several times over. To address these issues in a unified manner, we formulate the scene-aware visual reasoning network (SVRNet) based on visual prompts. Specifically, we construct the prototype-guided prompt chain reasoning (PPCR) that generates a mixed chain of thought reasoning based on maximal difference heterogeneous prototypes to construct comprehensive spatial location and semantic perception. Noise may accumulate gradually through the chain, and crucial clues may also disappear. Therefore, we design the prompt evolution (PE) to filter out noise and enhance the coupling between prompts. We further develop the mixture of prompt injection expert (MPIE) to dynamically select the optimal injection strategy in the low-rank space based on specific scene. Due to reflection interference and random parameter space introducing potential ambiguity, we formulate the three-way evidence-aware (TEA) loss to quantify the uncertainty, thereby providing reliable predictions. To leverage historical knowledge and further disentangle representations, we propose the frequency prototype contrastive (FPC) loss for learning more generalizable features across images. Finally, we relabel 25,828 images and formulate the first point-supervised MD framework. Extensive experiments conducted on four mirror benchmarks under three settings demonstrate that our method surpasses state-of-the-art approaches. Promising results are also achieved on six related benchmarks, showing its generality. Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Jiayi Ma 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | DSPFusion: Image Fusion via Degradation and Semantic Dual-Prior GuidanceabstractExisting infrared-visible image fusion methods are mainly tailored for high-quality source images. Although recent studies have begun to explore degradation-aware fusion, most existing methods still focus on specific degradation types, while unified frameworks that aim to handle diverse degradations often depend on auxiliary textual prompts, which limits their practicality in automatic fusion scenarios. This work presents a Degradation and Semantic Prior dual-guided framework for degraded image Fusion (DSPFusion), which jointly performs degradation-aware restoration and complementary information aggregation in a unified architecture without relying on auxiliary prompts. Specifically, it first extracts modality-specific degradation priors from degraded infrared and visible images, while capturing compact semantic embeddings from paired source images as low-quality semantic priors to encode global scene context. Then, a semantic prior diffusion model is devised to restore high-quality scene semantic priors in a compact latent space, providing global scene guidance with low computational overhead and enabling over $30\times $ inference speedup compared with mainstream diffusion model-based image fusion schemes, such as DDFM. Guided by the restored semantic priors and degradation priors, the enhancement and fusion network adaptively suppresses degradations and aggregates complementary information. Extensive experiments under both degraded and normal scenarios demonstrate that DSPFusion effectively handles representative degradations, preserves complementary information, and achieves competitive performance with low computational cost, thereby broadening the practical application scope of image fusion. The source code is publicly available at https://github.com/Linfeng-Tang/DSPFusion. Linfeng Tang, Yeda Wang, Guoqing Wang 0001, Yixuan Yuan, Jiayi Ma 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Hierarchical Consistency Learning for Test-Time Adaptation in Camouflage PerceptionabstractCamouflaged object detection (COD) aims to localize targets that exhibit minimal perceptual differences from backgrounds through physical attributes. Existing methods, constrained by the static train-then-freeze paradigm, suffer from domain rigidity and annotation dependency, limiting their adaptability to scene variations and unseen camouflage patterns. To overcome these, we propose the hierarchical consistency learning (HCL) framework, which integrates test-time adaptation for dynamic representation recalibration. Specifically, we design the hierarchical representation reconstruction (HRR) to alleviate feature entanglement by synergizing spatial reconstruction with dual-stream frequency-domain decomposition, enhancing robustness against appearance homogenization. The pixel and spectrum inference provide structural and contextual priors. We further introduce task affinity guidance (TAG) to propagate knowledge across branches via channel-wise affinity, aligning local discriminative cues and mitigating semantic drift. To ensure semantic invariance, we formulate the prototype consistency calibration (PCC), which aggregates region features into compact prototypes and establishes prototype-feature similarity. This imposes implicit and hierarchical constraints that bridge task and representation gaps. Extensive experiments across four camouflaged and four underwater object benchmarks, under three degradation settings, demonstrate that our method consistently outperforms state-of-the-art approaches, highlighting its robustness and generalization under distribution shifts. Mingfeng Zha, Tianyu Li 0003, Guoqing Wang 0001, Yunqiang Pei, Chaofan Qiao, Jiening Zhang, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 3 |
| 2026 | SNN-FT: Temporal-Coded Spiking Neural Networks for Fourier TransformabstractThe Fourier transform (FT) stands as a fundamental tool in modern signal processing with widespread applications across various scientific and engineering fields. Therefore, there remains a need for continued research efforts to devise energy-efficient implementations of the FT. Due to their inherent energy efficiency, biologically plausible spiking neural networks (SNNs) emerge as a promising alternative solution. However, current SNN implementations of the FT suffer from two key shortcomings, namely, high latency and reduced accuracy. In this article, we analyze the underlying causes of these limitations and highlight deficiencies in the existing spike-based encoding mechanisms and spiking neuron models. We then propose a new SNN-based FT (SNN-FT) based on a logarithmically polarized time-to-first-spike (TTFS) encoding method (called LP-TTFS) along with a novel piecewise spiking neuron (PTSN) model based on ternary spikes (referred to as PTSN). The resulting SNN-FT is mathematically equivalent to the conventional FT and demonstrates superior performance in accuracy as well as reduced latency. We assess the performance of the proposed SNN-FT alternative through extensive experiments on FT-based applications, such as radar and audio signal processing, and the obtained results demonstrate the efficacy of SNN-FT and its superiority over the existing approaches. This study unveils a novel energy-efficient neuromorphic computing technique with great potential for FT applications across diverse scientific and engineering domains. Shuai Wang 0058, Haorui Zheng, Ammar Belatreche, Guoqing Wang 0001, Yeying Jin, Jibin Wu, Malu Zhang, Yang Yang 0002, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Leveraging Asynchronous Spiking Neural Networks for Ultra Efficient Event-Based Visual ProcessingabstractEvent cameras encode visual information by generating asynchronous and sparse event streams, which hold great potential for low latency and low power consumption. Despite many successful implementations of event camera-based applications, most of them accumulate the events into frames and then utilize conventional frame-based computer vision algorithms. These frame-based methods, though typically effective, diminish the inherent advantages of the event camera's low latency and low power consumption. To solve the above problems, we propose ASGCN, which efficiently processes data on an event-by-event basis and dynamically evolves into a corresponding dynamic representation, enabling low latency and high sparsity of data representation. The sparsity computation is further improved by introducing brain-inspired spiking neural networks, resulting in low power consumption for ASGCN. Extensive and diverse experiments demonstrate the energy efficiency and low latency advantages of our processing pipeline. Especially on real-world event camera datasets, our pipeline consumes more than 10,000 times less energy and achieves similar performance compared to current frame-based methods. Dingyi Zeng, Honglin Cao, Wanlong Liu, Yichen Xiao, Chengzhuo Lu, Wenyu Chen 0001, Malu Zhang, Guoqing Wang 0001, Yang Yang 0002 |
AAAI | 9 |
| 2025 | Iterative Predictor-Critic Code Decoding for Real-World Image DehazingabstractWe propose a novel Iterative Predictor-Critic Code Decoding framework for real-world image dehazing, abbreviated as IPC-Dehaze, which leverages the high-quality codebook prior encapsulated in a pre-trained VQGAN. Apart from previous codebook-based methods that rely on oneshot decoding, our method utilizes high-quality codes obtained in the previous iteration to guide the prediction of the Code-Predictor in the subsequent iteration, improving code prediction accuracy and ensuring stable dehazing performance. Our idea stems from the observations that 1) the degradation of hazy images varies with haze density and scene depth, and 2) clear regions play crucial cues in restoring dense haze regions. However, it is nontrivial to progressively refine the obtained codes in subsequent iterations, owing to the difficulty in determining which codes should be retained or replaced at each iteration. Another key insight of our study is to propose CodeCritic to capture interrelations among codes. The CodeCritic is used to evaluate code correlations and then resample a set of codes with the highest mask scores, i.e., a higher score indicates that the code is more likely to be rejected, which helps retain more accurate codes and predict difficult ones. Extensive experiments demonstrate the superiority of our method over state-of-the-art methods in real-world dehazing. Our project page can be found at https://github.com/Jiayi-Fu/IPC-Dehaze. Jiayi Fu, Zikun Liu 0001, Chunle Guo, Hyunhee Park, Guoqing Wang 0001, Chongyi Li |
CVPR | 7 |
| 2025 | Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyabstractA prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived through the multiplication structure of down-projection and up-projection matrices, exemplified by methods such as LoRA and Adapter. In this work, we observe an approximate orthogonality among any two row or column vectors within any weight matrix of the backbone parameters; however, this property is absent in the vectors of the down/up-projection matrices. Approximate orthogonality implies a reduction in the upper bound of the model's generalization error, signifying that the model possesses enhanced generalization capability. If the fine-tuned down/up-projection matrices were to exhibit this same property as the pre-trained backbone matrices, could the generalization capability of fine-tuned ViTs be further augmented? To address this question, we propose an Approximately Orthogonal Fine-Tuning (AOFT) strategy for representing the low-rank weight matrices. This strategy employs a single learnable vector to generate a set of approximately orthogonal vectors, which form the down/up-projection matrices, thereby aligning the properties of these matrices with those of the backbone. Extensive experimental results demonstrate that our method achieves competitive performance across a range of downstream image classification tasks, confirming the efficacy of the enhanced generalization capability embedded in the down/up-projection matrices. Yiting Yang, Qingsen Yan, Haokui Zhang, Wei Dong 0010, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ICCV | 7 |
| 2025 | InteractGuide: LLM-Enhanced Multimodal Reasoning for User-Centric Interaction Recommendations in AR-HRI AuthoringabstractAugmented Reality (AR) enhances Human-Robot Interaction (HRI) by offering diverse interaction methods. However, existing systems often fail to resolve the conflict between a user's implicit preferences and physical ergonomics, leading to suboptimal experiences. We introduce InteractGuide, a novel framework that, for the first time, uses a Large Language Model (LLM) as a central reasoning engine to dynamically balance these competing factors. Our system translates physiological signals into a symbolic ''Preference Memory'' that the LLM reasons over, alongside real-time ergonomic and contextual data, to provide personalized interaction recommendations. A 29-participant study confirms our architecture improves efficiency and experience compared to single-factor approaches, showing the potential of LLMs as reasoning engines for complex AR-HRI. This work presents a validated end-to-end architecture for user-centric interaction adaptation, demonstrating the potential of LLMs as reasoning engines in complex AR-HRI systems. Yunqiang Pei, Hongrong Yang, Guoqing Wang 0001, Peng Wang 0023, Chaoning Zhang, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 4 |
| 2025 | S2NN: Sub-bit Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To further explore the compression and acceleration potential of SNNs, we propose Sub-bit Spiking Neural Networks (S$^2$NNs) that represent weights with less than one bit. Specifically, we first establish an S$^2$NN baseline by leveraging the clustering patterns of kernels in well-trained binary SNNs. This baseline is highly efficient but suffers from \textit{outlier-induced codeword selection bias} during training. To mitigate this issue, we propose an \textit{outlier-aware sub-bit weight quantization} (OS-Quant) method, which optimizes codeword selection by identifying and adaptively scaling outliers. Furthermore, we propose a \textit{membrane potential-based feature distillation} (MPFD) method, improving the performance of highly compressed S$^2$NN via more precise guidance from a teacher model. Extensive results on vision reveal that S$^2$NN outperforms existing quantized SNNs in both performance and efficiency, making it promising for edge computing applications. Wenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche, Shuai Wang 0058, Yimeng Shan, Honglin Cao, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 9 |
| 2025 | Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence ModelingabstractThe explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiotemporal spike trains, making them well-suited for long sequence modeling. However, RF neurons exhibit limited effective memory capacity and a trade-off between energy efficiency and training speed on complex temporal tasks. Inspired by the dendritic structure of biological neurons, we propose a Dendritic Resonate-and-Fire (D-RF) model, which explicitly incorporates a multi-dendritic and soma architecture. Each dendritic branch encodes specific frequency bands by utilizing the intrinsic oscillatory dynamics of RF neurons, thereby collectively achieving comprehensive frequency representation. Furthermore, we introduce an adaptive threshold mechanism into the soma structure. This mechanism adjusts the firing threshold according to historical spiking activity, thereby reducing redundant spikes while maintaining training efficiency in long-sequence tasks. Extensive experiments demonstrate that our method maintains competitive accuracy while substantially ensuring sparse spikes without compromising computational efficiency during training. These results underscore its potential as an effective and efficient solution for long sequence modeling on edge platforms. Dehao Zhang, Malu Zhang, Shuai Wang 0058, Wenjie Wei, Zeyu Ma 0002, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 7 |
| 2025 | MSCRS: Multi-modal Semantic Graph Prompt Learning Framework for Conversational Recommender SystemsabstractConversational Recommender Systems (CRSs) aim to provide personalized recommendations by interacting with users through conversations. Most existing studies of CRS focus on extracting user preferences from conversational contexts. However, due to the short and sparse nature of conversational contexts, it is difficult to fully capture user preferences by conversational contexts only. We argue that multi-modal semantic information can enrich user preference expressions from diverse dimensions (e.g., a user preference for a certain movie may stem from its magnificent visual effects and compelling storyline). In this paper, we propose a multi-modal semantic graph prompt learning framework for CRS, named MSCRS. First, we extract textual and image features of items mentioned in the conversational contexts. Second, we capture higher-order semantic associations within different semantic modalities (collaborative, textual, and image) by constructing modality-specific graph structures. Finally, we propose an innovative integration of multi-modal semantic graphs with prompt learning, harnessing the power of large language models to comprehensively explore high-dimensional semantic relationships. Experimental results demonstrate that our proposed method significantly improves accuracy in item recommendation, as well as generates more natural and contextually relevant content in response generation. Code and extended multi-modal CRS datasets are available at https://github.com/BIAOBIAO12138/MSCRS-main. Yibiao Wei, Jie Zou 0001, Weikang Guo, Guoqing Wang 0001, Xing Xu 0001, Yang Yang 0002 |
SIGIR | 4 |
| 2025 | AttentionAR: AR Adaptation and Warning for Real-World Safety via Attention Modeling and MLLM Reasoning
Yunqiang Pei, Renming Huang, Mingfeng Zha, Guoqing Wang 0001, Peng Wang 0023, Qiao Kang, Yang Yang 0002, Heng Tao Shen |
UIST | 4 |
| 2025 | Unlocking spatial textures: Gradient-guided pansharpening for enhancing multispectral imagery
Lanyue Liang, Tianyu Li 0003, Guoqing Wang 0001, Lin Mei 0001, Xiongxin Tang, Chaofan Qiao, Dongyu Xie |
Neurocomputing | 3 |
| 2025 | Evidence-Based Multi-Feature Fusion for Adversarial RobustnessabstractThe accumulation of adversarial perturbations in the feature space makes it impossible for Deep Neural Networks (DNNs) to know what features are robust and reliable, and thus DNNs can be fooled by relying on a single contaminated feature. Numerous defense strategies attempt to improve their robustness by denoising, deactivating, or recalibrating non-robust features. Despite their effectiveness, we still argue that these methods are under-explored in terms of determining how trustworthy the features are. To address this issue, we propose a novel Evidence-based Multi-Feature Fusion (termed EMFF) for adversarial robustness. Specifically, our EMFF approach introduces evidential deep learning to help DNNs quantify the belief mass and uncertainty of the contaminated features. Subsequently, a novel multi-feature evidential fusion mechanism based on Dempster's rule is proposed to fuse the trusted features of multiple blocks within an architecture, which further helps DNNs avoid the induction of a single manipulated feature and thus improve their robustness. Comprehensive experiments confirm that compared with existing defense techniques, our novel EMFF method has obvious advantages and effectiveness in both scenarios of white-box and black-box attacks, and also prove that by integrating into several adversarial training strategies, we can improve the robustness of across distinct architectures, including traditional CNNs and recent vision Transformers with a few extra parameters and almost the same cost. Zheng Wang 0044, Xing Xu 0001, Lei Zhu 0002, Yi Bin, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM CollaborationabstractReferring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. It serves as an essential testing ground for Multimodal Large Language Models (MLLMs). To advance this field, we introduced a new REC dataset in our previous conference paper, characterized by two key features. First, it is designed with controllable difficulty levels, requiring multi-level fine-grained reasoning across object categories, attributes, and multi-hop relationships. Second, it incorporates negative text and images generated through fine-grained editing and augmentation, explicitly testing a model's ability to reject scenarios where the target object is absent-an often-overlooked yet critical challenge in existing datasets. In this extended work, we propose two new methods to tackle the challenges of fine-grained REC by combining the strengths of Specialist Models and MLLMs. The first method adaptively assigns simple cases to faster, lightweight models and reserves complex ones for powerful MLLMs, balancing accuracy and efficiency. The second method lets a specialist generate a set of possible object regions, and the MLLM selects the most plausible one using its reasoning ability. These collaborative strategies lead to significant improvements on our dataset and other challenging benchmarks. Our results show that combining specialized and general-purpose models offers a practical path toward solving complex real-world vision-language tasks. Xuzheng Yang, Junzhuo Liu 0002, Peng Wang 0023, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Toward Generalized and Realistic Unpaired Image Dehazing via Region-Aware Physical ConstraintsabstractSupervised dehazing models, trained on synthetic hazy-clean image pairs, often face a notable decline in performance when applied to real-world scenes. Consequently, CycleGAN-based unpaired dehazing methods are proposed to improve the model’s generalization. One successful approach among these methods involves decomposing the physical properties of the atmospheric scattering model (ASM). However, estimating physical properties individually from input images is difficult without supervised labels, which ignores the semantic consistency between different physical regions. We claim semantic region information can offer additional geometric spatial constraints for estimating physical properties, as natural images can be divided into regions with similar scene depths. Motivated by this, we propose a novel generalized and realistic unpaired image dehazing framework via region-aware physical constraints (RPC-Dehaze). Our approach utilizes fine-grained semantic region maps from the Segment Anything Model (SAM) in a specially designed region prompt enhancement module. This enables the dehazing and hazing cyclic networks to learn region-aware physical constraints, leading to accurate estimation of haze imaging physical properties. In contrast to existing unpaired methods that treat dehazing and hazing networks equally, we incorporate Retinex theory into the hazing network, allowing it to learn diverse illumination effects in different regions. We adaptively refine the Retinex-based illumination component, resulting in more realistic hazy images. To further facilitate unsupervised learning in our framework, we propose a physical consensual contrastive regularization to ensure compact representation constraints in the latent feature space. Extensive experiments on synthetic and real image datasets show our method surpasses state-of-the-art unpaired dehazing methods in both effectiveness and generalization capability. Kaihao Lin, Guoqing Wang 0001, Tianyu Li 0003, Yuhui Wu 0001, Chongyi Li, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DMM: Disparity-Guided Multispectral Mamba for Oriented Object Detection in Remote SensingabstractMultispectral oriented object detection faces challenges due to both inter-modal and intra-modal discrepancies. Recent studies often rely on transformer-based models to address these issues and achieve cross-modal fusion detection. However, the quadratic computational complexity of transformers limits their performance in remote sensing imagery. Inspired by the efficiency and lower complexity of Mamba in long sequence tasks, we propose Disparity-guided Multispectral Mamba (DMM), a multispectral oriented object detection framework comprised of a Disparity-guided Cross-modal Fusion Mamba (DCFM) module, a Multi-scale Target-aware Attention (MTA) module, and a Target-Prior Aware (TPA) auxiliary task. The DCFM module leverages disparity information between modalities to adaptively merge features from RGB and IR images, mitigating inter-modal conflicts. The MTA module aims to enhance feature representation by focusing on relevant target regions within the RGB modality, addressing intra-modal variations. The TPA auxiliary task utilizes single-modal labels to guide the optimization of the MTA module, ensuring it focuses on targets and their local context. Extensive experiments on the DroneVehicle and VEDAI datasets demonstrate the effectiveness of our method, which outperforms state-of-the-art methods while maintaining computational efficiency. Code will be available at https://github.com/Another-0/DMM. Minghang Zhou, Tianyu Li 0003, Chaofan Qiao, Dongyu Xie, Guoqing Wang 0001, Ningjuan Ruan, Lin Mei 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Heterogeneous Experts and Hierarchical Perception for Underwater Salient Object DetectionabstractExisting underwater salient object detection (USOD) methods design fusion strategies to integrate multimodal information, but lack exploration of modal characteristics. To address this, we separately leverage the RGB and depth branches to learn disentangled representations, formulating the heterogeneous experts and hierarchical perception network (HEHP). Specifically, to reduce modal discrepancies, we propose the hierarchical prototype guided interaction (HPI), which achieves fine-grained alignment guided by the semantic prototypes, and then refines with complementary modalities. We further design the mixture of frequency experts (MoFE), where experts focus on modeling high- and low-frequency respectively, collaborating to explicitly obtain hierarchical representations. To efficiently integrate diverse spatial and frequency information, we formulate the four-way fusion experts (FFE), which dynamically selects optimal experts for fusion while being sensitive to scale and orientation. Since depth maps with poor quality inevitably introduce noises, we design the uncertainty injection (UI) to explore high uncertainty regions by establishing pixel-level probability distributions. We further formulate the holistic prototype contrastive (HPC) loss based on semantics and patches to learn compact and general representations across modalities and images. Finally, we employ varying supervision based on branch distinctions to implicitly construct difference modeling. Extensive experiments on two USOD datasets and four relevant underwater scene benchmarks validate the effect of the proposed method, surpassing state-of-the-art binary detection models. Impressive results on seven natural scene benchmarks further demonstrate the scalability. Mingfeng Zha, Guoqing Wang 0001, Yunqiang Pei, Tianyu Li 0003, Xiongxin Tang, Chongyi Li, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Image Process. | 2 |
| 2025 | Geometric Matching for Cross-Modal RetrievalabstractDespite its significant progress, cross-modal retrieval still suffers from one-to-many matching cases, where the multiplicity of semantic instances in another modality could be acquired by a given query. However, existing approaches usually map heterogeneous data into the learned space as deterministic point vectors. In spite of their remarkable performance in matching the most similar instance, such deterministic point embedding suffers from the insufficient representation of rich semantics in one-to-many correspondence. To address the limitations, we intuitively extend a deterministic point into a closed geometry and develop geometric representation learning methods for cross-modal retrieval. Thus, a set of points inside such a geometry could be semantically related to many candidates, and we could effectively capture the semantic uncertainty. We then introduce two types of geometric matching for one-to-many correspondence, i.e., point-to-rectangle matching (dubbed P2RM) and rectangle-to-rectangle matching (termed R2RM). The former treats all retrieved candidates as rectangles with zero volume (equivalent to points) and the query as a box, while the latter encodes all heterogeneous data into rectangles. Therefore, we could evaluate semantic similarity among heterogeneous data by the Euclidean distance from a point to a rectangle or the volume of intersection between two rectangles. Additionally, both strategies could be easily employed for off-the-self approaches and further improve the retrieval performance of baselines. Under various evaluation metrics, extensive experiments and ablation studies on several commonly used datasets, two for image-text matching and two for video-text retrieval, demonstrate our effectiveness and superiority. Zheng Wang 0044, Zhenwei Gao, Yang Yang 0002, Guoqing Wang 0001, Chengbo Jiao, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | ScanERU: Interactive 3D Visual Grounding Based on Embodied Reference UnderstandingabstractAiming to link natural language descriptions to specific regions in a 3D scene represented as 3D point clouds, 3D visual grounding is a very fundamental task for human-robot interaction. The recognition errors can significantly impact the overall accuracy and then degrade the operation of AI systems. Despite their effectiveness, existing methods suffer from the difficulty of low recognition accuracy in cases of multiple adjacent objects with similar appearance. To address this issue, this work intuitively introduces the human-robot interaction as a cue to facilitate the development of 3D visual grounding. Specifically, a new task termed Embodied Reference Understanding (ERU) is first designed for this concern. Then a new dataset called ScanERU is constructed to evaluate the effectiveness of this idea. Different from existing datasets, our ScanERU dataset is the first to cover semi-synthetic scene integration with textual, real-world visual, and synthetic gestural information. Additionally, this paper formulates a heuristic framework based on attention mechanisms and human body movements to enlighten the research of ERU. Experimental results demonstrate the superiority of the proposed method, especially in the recognition of multiple identical objects. Our codes and dataset are available in the ScanERU repository. Yunqiang Pei, Guoqing Wang 0001, Peiwei Li, Yang Yang 0002, Yinjie Lei, Heng Tao Shen |
AAAI | 3 |
| 2024 | Weakly-Supervised Mirror Detection via Scribble AnnotationsabstractMirror detection is of great significance for avoiding false recognition of reflected objects in computer vision tasks. Existing mirror detection frameworks usually follow a supervised setting, which relies heavily on high quality labels and suffers from poor generalization. To resolve this, we instead propose the first weakly-supervised mirror detection framework and also provide the first scribble-based mirror dataset. Specifically, we relabel 10,158 images, most of which have a labeled pixel ratio of less than 0.01 and take only about 8 seconds to label. Considering that the mirror regions usually show great scale variation, and also irregular and occluded, thus leading to issues of incomplete or over detection, we propose a local-global feature enhancement (LGFE) module to fully capture the context and details. Moreover, it is difficult to obtain basic mirror structure using scribble annotation, and the distinction between foreground (mirror) and background (non-mirror) features is not emphasized caused by mirror reflections. Therefore, we propose a foreground-aware mask attention (FAMA), integrating mirror edges and semantic features to complete mirror regions and suppressing the influence of backgrounds. Finally, to improve the robustness of the network, we propose a prototype contrast loss (PCL) to learn more general foreground features across images. Extensive experiments show that our network outperforms relevant state-of-the-art weakly supervised methods, and even some fully supervised methods. The dataset and codes are available at https://github.com/winter-flow/WSMD. Mingfeng Zha, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Yang Yang 0002, Wenbin Qian, Heng Tao Shen |
AAAI | 3 |
| 2024 | Diffusion Models as Optimizers for Efficient Planning in Offline RL
Renming Huang, Yunqiang Pei, Guoqing Wang 0001, Yang Yang 0002, Peng Wang 0023, Heng Tao Shen |
ECCV (51) | 3 |
| 2024 | Region-Aware Distribution Contrast: A Novel Approach to Multi-task Partially Supervised Learning
Tianyu Li 0003, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Jie Zou 0001 |
ECCV (51) | 3 |
| 2024 | Domain Prompt Learning Framework for Real Image DehazingabstractSupervised dehazing models trained on synthetic datasets exhibit severe performance degradation in real-world scenarios due to the domain gap. Unsupervised methods are proposed to process real hazy images, while they suffer from the complex training procedure. In this paper, we present a universal domain prompt learning framework (DPLF) for boosting the performance of supervised dehazing models in real scenarios by introducing prompt learning and well-designed domain adapters. We train the learnable text prompts by CLIP feature alignment, which can discriminate between real hazy and clean images, and use these prompts as unsupervised text constraints. Notably, the distribution gap between synthetic and real haze can be regarded as the difference of the haze-relevant style domain. Motivated by this, we design the style domain prompt adapter to align features from synthetic and real haze domains. Extensive experiments on real-world datasets demonstrate significant performance improvement of the baseline dehazing models with our DPLF. Kaihao Lin, Guoqing Wang 0001, Yuhui Wu 0001, Shuhang Gu, Xing Xu 0001, Yang Yang 0002 |
ICME | 2 |
| 2024 | Shapley Ensemble Adversarial AttackabstractAn intuitive strategy for generating more transferable adversarial examples is to absorb the advantages of different surrogates in an ensemble. However, existing ensemble adversarial attacks simply average the outputs of various models while ignoring their different contributions. To quantify the importance of various surrogates, we propose a novel Shapley Ensemble Adversarial Attack (dubbed SEAA) – an effective algorithm that allocates weights based on Shapley values of the surrogates. Specifically, the synthesis of adversarial examples in an ensemble adversarial attack is firstly regarded as a cooperative game process of multiple surrogates from the perspective of contribution. As the Shapley value is known as an effective decision metric in cooperative game theory, it is thus intuitively introduced into this work to accurately evaluate the contribution of a surrogate by reweighing its importance at each iteration, thus avoiding local optimality and steering the generation of adversarial examples. Comprehensive experimental results demonstrate our effectiveness. Zheng Wang 0044, Yi Bin, Lei Zhu 0002, Guoqing Wang 0001, Yang Yang 0002 |
ICME | 5 |
| 2024 | Temporal Self-Paced Proposal Learning for Weakly-Supervised Video Moment Retrieval and Highlight DetectionabstractThe Weakly-Supervised Moment Retrieval and Highlight Detection (WS-MRHD) task aims at retrieving target moments and highlights in an untrimmed video with a semantic relevant text query. One of the most challenging problems in this task is the absence of reliable temporal supervision signals. In this paper, we propose a Temporal Self-paced Proposal Learning (TSPL) method to perform a progressive temporal proposal selection mechanism. It productively improves the effectiveness of contrastive learning even when the frame-level annotations are inaccessible. Specifically, our proposed TSPL method consists of three key components: (1) The Variance-Based Instance Selection (VBIS) module leverages self-paced learning for dynamic temporal proposal selection. (2) A Highlight Broadcasting (HB) module to combine reliable time spans and assign frame-level pseudo labels. (3) A Negative Sample Learning (NSL) module to align the text query with relevant video segments. By dynamically selecting the appropriate temporal proposals for training, our TSPL method conducts more reliable cross-modal alignment thus remarkably boosting retrieval performance. The extensive experiments on two WS-MRHD public benchmarks verify our proposed TSPL method substantially outperforms current state-of-the-art methods. Liqing Zhu, Xun Jiang 0001, Fumin Shen, Guoqing Wang 0001, Yang Yang 0002, Xing Xu 0001 |
ICME | 4 |
| 2024 | Towards Real-time Video Compressive Sensing on Mobile Devices
Lishun Wang, Huan Wang 0014, Guoqing Wang 0001, Xin Yuan 0002 |
ACM Multimedia | 4 |
| 2024 | Emotion Recognition in HMDs: A Multi-task Approach Using Physiological Signals and Occluded FacesabstractPrior research on emotion recognition in extended reality (XR) has faced challenges due to the occlusion of facial expressions by Head-Mounted Displays (HMDs). This limitation hinders accurate Facial Expression Recognition (FER), which is crucial for immersive user experiences. This study aims to overcome the occlusion challenge by integrating physiological signals with partially visible facial expressions to enhance emotion recognition in XR environments. We employed a multi-task approach, utilizing a feature-level fusion to fuse Electroencephalography (EEG) and Galvanic Skin Response (GSR) signals with occluded facial expressions. The model predicts valence and arousal simultaneously from both macro-and micro-expression. Our method demonstrated improved accuracy in emotion recognition under partial occlusion conditions. The integration of temporal physiological signals with other modalities significantly enhanced performance, particularly for half-face emotion recognition. The study presents a novel approach to emotion recognition in XR, addressing the limitations of facial occlusion by HMDs. The findings suggest that physiological signals are vital for interpreting emotions in occluded scenarios, offering potential for real-time applications and advancing social XR applications. Yunqiang Pei, Jialei Tang, Qihang Tang, Mingfeng Zha, Dongyu Xie, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 6 |
| 2024 | Improving Interaction Comfort in Authoring Task in AR-HRI through Dynamic Dual-Layer Interaction AdjustmentabstractPrevious research has demonstrated the potential of Augmented Reality in enhancing psychological comfort in Human-Robot Interaction (AR-HRI) through shared robot intent, enhanced visual feedback, and increased expressiveness and creativity in interaction methods. However, the challenge of selecting interaction methods that enhance physical comfort in varying scenarios remains. This study purposes a dynamic dual-layer interaction adjustment mechanism to improve user comfort and interaction efficiency. The mechanism comprises two models: an general layer model, grounded in ergonomics principles, identifies appropriate areas for various interaction methods; a individual layer model predicts user discomfort levels using physiological signals. Interaction methods are dynamically adjusted based on discomfort level changes, enabling the system to adapt to individual differences and dynamic changes, thereby reducing misjudgments and enhancing comfort management. The mechanism's success in authoring tasks validates its effectiveness, significantly advancing AR-HRI and fostering more comfortable and enhancing efficient human-centered interactions. Yunqiang Pei, Hongrong Yang, Qihang Tang, Jialei Tang, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 7 |
| 2024 | Cascaded Adversarial Attack: Simultaneously Fooling Rain Removal and Semantic Segmentation NetworksabstractWhen applying high-level visual algorithms to rainy scenes, it is customary to preprocess the rainy images using low-level rain removal networks, followed by visual networks to achieve the desired objectives. Such a setting has never been explored by adversarial attack methods, which are only limited to attacking one kind of them. Considering the deficiency of multi-functional attacking strategies and the significance for open-world perception scenarios, we are the first to propose a Cascaded Adversarial Attack (CAA) setting, where the adversarial example can simultaneously attack different-level tasks, such as rain removal and semantic segmentation in an integrated system. Specifically, our attack on the rain removal network aims to preserve rain streaks in the output image, while for the semantic segmentation network, we employ powerful existing adversarial attack methods to induce misclassification of the image content. Importantly, CAA innovatively utilizes binary masks to effectively concentrate the aforementioned two significantly disparate perturbation distributions on the input image, enabling attacks on both networks. Additionally, we propose two variants of CAA, which minimize the differences between the two generated perturbations by introducing a carefully designed perturbation interaction mechanism, resulting in enhanced attack performance. Extensive experiments validate the effectiveness of our methods, demonstrating their superior ability to significantly degrade the performance of the downstream task compared to methods that solely attack a single network. Zhiwen Wang 0004, Yuhui Wu 0001, Zheng Wang 0044, Jiwei Wei, Tianyu Li 0003, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 6 |
| 2024 | JoReS-Diff: Joint Retinex and Semantic Priors in Diffusion Model for Low-light Image EnhancementabstractLow-light image enhancement (LLIE) has achieved promising performance by employing conditional diffusion models. Despite the success of some conditional methods, previous methods may neglect the importance of a sufficient formulation of task-specific condition strategy, resulting in suboptimal visual outcomes. In this study, we propose JoReS-Diff, a novel approach that incorporates Retinex- and semantic-based priors as the additional pre-processing condition to regulate the generating capabilities of the diffusion model. We first leverage pre-trained decomposition network to generate the Retinex prior, which is updated with better quality by an adjustment network and integrated into a refinement network to implement Retinex-based conditional generation at both feature- and image-levels. Moreover, the semantic prior is extracted from the input image with an off-the-shelf semantic segmentation model and incorporated through semantic attention layers. By treating Retinex- and semantic-based priors as the condition, JoReS-Diff presents a unique perspective for establishing an diffusion model for LLIE and similar image enhancement tasks. Extensive experiments validate the rationality and superiority of our approach. Yuhui Wu 0001, Guoqing Wang 0001, Zhiwen Wang 0004, Yang Yang 0002, Tianyu Li 0003, Malu Zhang, Chongyi Li, Heng Tao Shen |
ACM Multimedia | 2 |
| 2024 | Generalizing ISP Model by Unsupervised Raw-to-raw MappingabstractISP (Image Signal Processor) serves as a pipeline converting unprocessed raw images to sRGB images, positioned before nearly all visual tasks. Due to the varying spectral sensitivities of cameras, raw images captured by different cameras exist in different color spaces, making it challenging to deploy ISP across cameras with consistent performance. To address this challenge, it is intuitively to incorporate a raw-to-raw mapping (mapping raw images across camera color spaces) module into the ISP. However, the lack of paired data (i.e., images of the same scene captured by different cameras) makes it difficult to train a raw-to-raw model using supervised learning methods. In this paper, we aim to achieve ISP generalization by proposing the first unsupervised raw-to-raw model. To be specific, we propose a CSTPP (Color Space Transformation Parameters Predictor) module to predict the space transformation parameters in a patch-wise manner, which can accurately perform color space transformation and flexibly manage complex lighting conditions. Additionally, we design a CycleGAN-style training framework to realize unsupervised learning, overcoming the deficiency of paired data. Our proposed unsupervised model achieved performance comparable to that of the state-of-the-art semi-supervised method in raw-to-raw task. Furthermore, to assess its ability to generalize the ISP model across different cameras, we for the first formulated cross-camera ISP task and demonstrated the performance of our method through extensive experiments. The codes are released at https://github.com/ydxxxx/Unsupervised-Raw-to-raw-Mapping. Dongyu Xie, Chaofan Qiao, Lanyue Liang, Zhiwen Wang 0004, Tianyu Li 0003, Qiao Liu 0003, Chongyi Li, Guoqing Wang 0001, Yang Yang 0002 |
ACM Multimedia | 8 |
| 2024 | Physics-Constrained Comprehensive Optical Neural NetworksabstractWith the advantages of low latency, low power consumption, and high parallelism, optical neural networks (ONN) offer a promising solution for time-sensitive and resource-limited artificial intelligence applications. However, the performance of the ONN model is often diminished by the gap between the ideal simulated system and the actual physical system. To bridge the gap, this work conducts extensive experiments to investigate systematic errors in the optical physical system within the context of image classification tasks. Through our investigation, two quantifiable errors—light source instability and exposure time mismatches—significantly impact the prediction performance of ONN. To address these systematic errors, a physics-constrained ONN learning framework is constructed, including a well designed loss function to mitigate the effect of light fluctuations, a CCD adjustment strategy to alleviate the effects of exposure time mismatches and a ’physics-prior based’ error compensation network to manage other systematic errors, ensuring consistent light intensity across experimental results and simulations. In our experiments, the proposed method achieved a test classification accuracy of 96.5% on the MNIST dataset, a substantial improvement over the 61.6% achieved with the original ONN. For the more challenging QuickDraw16 and Fashion MNIST datasets, experimental accuracy improved from 63.0% to 85.7% and from 56.2% to 77.5%, respectively. Moreover, the comparison results further demonstrate the effectiveness of the proposed physics-constrained ONN learning framework over state-of-the-art ONN approaches. This lays the groundwork for more robust and precise optical computing applications. Yanbing Liu 0006, Jianwei Qin, Xi Yue, Guoqing Wang 0001, Tianyu Li 0003, Fangwei Ye |
NeurIPS | 6 |
| 2024 | Towards a Flexible Semantic Guided Model for Single Image Enhancement and RestorationabstractLow-light image enhancement (LLIE) investigates how to improve the brightness of an image captured in illumination-insufficient environments. The majority of existing methods enhance low-light images in a global and uniform manner, without taking into account the semantic information of different regions. Consequently, a network may easily deviate from the original color of local regions. To address this issue, we propose a semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that adaptively integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in the LLIE task. We further present a refined framework SKF++ with two new techniques: (a) Extra convolutional branch for intra-class illumination and color recovery through extracting local information and (b) Equalization-based histogram transformation for contrast enhancement and high dynamic range adjustment. Extensive experiments on various benchmarks of LLIE task and other image processing tasks show that models equipped with the SKF/SKF++ significantly outperform the baselines and our SKF/SKF++ generalizes to different models and scenes well. Besides, the potential benefits of our method in face detection and semantic segmentation in low-light conditions are discussed. Yuhui Wu 0001, Guoqing Wang 0001, Shaochong Liu, Yang Yang 0002, Wei Liu 0005, Xiongxin Tang, Shuhang Gu, Chongyi Li, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Reconstruction flow recurrent network for compressed video quality enhancement
Zhengning Wang, Xuhang Liu, Chuan Wang 0001, Ting Jiang 0005, Tianjiao Zeng, Zhenni Zeng, Guoqing Wang 0001, Shuaicheng Liu |
Pattern Recognit. | 7 |
| 2024 | Towards constructing a DOE-based practical optical neural system for ship recognition in remote sensing images
Yanbing Liu 0006, Shaochong Liu, Tianyu Li 0003, Guoqing Wang 0001 |
Signal Process. | 6 |
| 2024 | Dual Domain Perception and Progressive Refinement for Mirror DetectionabstractMirror detection aims to discover mirror regions in images to avoid misidentifying reflected objects. Existing methods mainly mine clues from spatial domain. We observe that the frequencies inside and outside the mirror region are distinctive. Besides, the low-frequency representing the feature semantics can help to locate the mirror region, and the high-frequency representing the details can refine it. Motivated by this, we introduce frequency guidance and propose the dual domain perception progressive refinement network (DPRNet) to mine dual-domain information. Specifically, we first decouple the images into high-frequency and low-frequency components by Laplace pyramid and vision Transformer, respectively, and design the frequency interaction alignment (FIA) module to integrate frequency features to initially localize the mirror region. To handle scale variations, we propose the multi-order feature perception (MOFP) module to adaptively aggregate adjacent features with progressive and gating mechanisms. We further propose the separation-based difference fusion (SDF) module to establish associations between entities and imagings and discover the correct boundary to mine the complete mirror region. Extensive experiments show that DPRNet outperforms the state-of-the-art method by an average of 3% with only about one-fifth of the parameters and FLOPs on four datasets. Our DPRNet also achieves promising performance on remote sensing and camouflage scenarios, validating its generalization. The code is available athttps://github.com/winter-flow/DPRNet. Mingfeng Zha, Feiyang Fu, Yunqiang Pei, Guoqing Wang 0001, Tianyu Li 0003, Xiongxin Tang, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Density-Aware Cloud Removal of Remote Sensing Imagery Using a Global-Local Fusion TransformerabstractCloud cover poses a significant challenge in remote sensing image processing, affecting the extraction and analysis of terrestrial features. Despite advancements in multitemporal cloud removal methods, single-image declouding remains crucial for emergency response and disaster management, where rapid acquisition of cloud-free imagery is essential. Traditional approaches often rely on synthetic aperture radar (SAR) or cloud masks as guidance for cloud removal, introducing additional complexities and dependencies on extensive data. To address these limitations, we propose a density-aware cloud removal using a global-local fusion Transformer (DCR-GLFT), which leverages density information as guidance and does not rely on extensive data. Specifically, our method employs density labels to guide the cloud removal process through two primary stages: cloud density estimation and density-guided cloud removal. A cloud density classifier is proposed in the first stage, trained with roughly estimated ground truth, to generate density labels for guiding subsequent removal processes. The second stage integrates cloud density information with cloud-ground image features using a Transformer-based network, enabling precise and nuanced cloud removal while preserving underlying surface details through the integration of both global and local features. The proposed method achieved the state-of-the-art results (peak signal-to-noise ratio (PSNR) of 28.93 and structural similarity index measure (SSIM) of 0.84) on the renowned cloud-removal dataset SEN12MS-CR, even without utilizing SAR data for guidance. This accomplishment highlights its significant advancement in the single-image cloud removal task. Our code will be made available athttps://github.com/ruiquan1214/DCR-GLFT.git. Quan Rui, Shiyuan He, Tianyu Li 0003, Guoqing Wang 0001, Ningjuan Ruan, Lin Mei 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Pixel Bleach Network for Detecting Face Forgery Under CompressionabstractThe existing face forgery algorithms have achieved remarkable progress in how to generate reasonable facial images and can even successfully deceive human beings. Considering public security, face forgery detection is of vital importance, making it essential to design face forgery detection algorithms to detect forgery images over the Internet. Despite the great success achieved by the existing Deepfake detection algorithms, they usually failed to achieve satisfactory Deepfake detection performance when deployed to handle the forgery videos in practice. One significant reason is compression. The videos over the Internet are inevitably compressed considering the transmission efficiency. To address this issue, in this paper, we propose a generic, simple yet effective “bleaching” pre-processing module based on the generative model and the high-level feature representations to produce ableached image, which shares a similar appearance with the compressed images. The bleached images with recovered information can be identified accurately by the optimized Deepfake detection models without retraining. The proposed method has utilized a redesigned feature representation, which serves as a navigator to effectively and sufficiently alter the feature distribution in the high-dimensional space to remedy the difference between real facial images and forgery counterparts. Thus, the proposed method can successfully avoid misclassification. Comprehensive and extensive experiments are carried out on four low-quality Faceforensics++ datasets, demonstrating the effectiveness of our method in recovering the information loss caused by the compression artifacts across various backbones and compression. Congrui Li, Ziqiang Zheng, Yi Bin, Guoqing Wang 0001, Yang Yang 0002, Xuesheng Li, Heng Tao Shen |
IEEE Trans. Multim. | 4 |
| 2024 | Adaptive Multi-scale Degradation-Based Attack for Boosting the Adversarial TransferabilityabstractThe vulnerability of deep neural networks to adversarial examples has raised huge concerns about the security of these algorithms. Black-box adversarial attacks have received a lot of attention as an influential method for evaluating model robustness. While various sophisticated adversarial attack methods have been proposed, the success rate in the black-box scenario still needs to be improved. To address these issues, we develop an Adaptive Multi-scale Degradation-based Attack method calledAMDA. The intuitive motivation behind our approach is that different models tend to have similar attention regions for low-scale images. Specifically, AMDA uses degraded images to generate perturbations at different scales and fuses these perturbations to generate adversarial examples that are insensitive to model changes. Furthermore, we design an adaptive multi-scale perturbation fusion that evaluates the transferability of perturbations at different scales based on noise and adaptively allocates fusion weights to prioritize strong transferability attacks and avoid being compromised by local optima. Extensive experimental results on the ImageNet, CIFAR-100, and CIFAR-10 datasets demonstrate that the proposed AMDA algorithm exhibits competitive performance for both normally trained models and defense models. Ran Ran 0001, Jiwei Wei, Chaoning Zhang, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 4 |
| 2024 | Towards Robust Person Re-Identification by Adversarial Training With Dynamic Attack StrategyabstractRecently, person re-identification has gained significant attention from both academic and industry fields due to its potential applications in surveillance and security. However, the security of re-identification systems has not been widely investigated, and they are vulnerable to adversarial attacks, which can significantly degrade their performance. Although numerous sophisticated adversarial training methods have been proposed for image classification, metric analysis systems such as person re-identification have not been fully explored. In this paper, we develop a novel adversarial training framework with a dynamic attack strategy for person re-identification, to further enhance the robustness of the model. Specifically, we gradually increase the perturbation budget during the generation until the generated adversarial examples reach a certain level of attack strength. As the iterations progress, the model becomes more robust, and our framework can generate stronger adversarial examples to continuously explore the robustness bounds of the model. Moreover, to alleviate the conflict between the adversarial robustness and natural generalization of the model, we design a novel performance alignment loss to further constrain the adversarial example generation process, which can make the generated adversarial examples as close as possible to the clean samples in terms of performance. Experiments on two widely used person re-ID benchmark datasets demonstrate the effectiveness and superiority of our proposed method. Jiwei Wei, Shiyuan He, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 4 |
| 2024 | Runge-Kutta Guided Feature Augmentation for Few-Sample LearningabstractDeep Neural Networks (DNNs) have primarily been demonstrated to be successful when large-scale labeled data are available. However, DNNs usually fail when tasked in few-sample learning scenarios, and the results will be much worse when the limited data show large intra-class variation and inter-class similarity (a.k.a fine-grained classification). To solve this challenging task, the idea of carrying out feature augmentation is visited and better achieved by exploring the merit of the forward Euler method in solving ordinary differential equations (ODEs), and a novel high-order feature augmentation (HFA) model with ResNet is proposed. Specifically, the proposed method leverages the stacked residual structure to model the direction of feature change over the initial state, and uses the triplet loss as constraint to model the step size of change in an adaptive manner. As a result, the initial features can then be augmented by a residual structure with a forward Eulerian form to generate features of the same subcategory with a similar representation as the input image. Furthermore, the proposed augmentation mechanism enjoys two additional benefits: a) it can help avoid the over-fitting issue when learned with insufficient training data; b) it can be used seamlessly with any residual structure-based classification network, and the ResNet used in this paper remains unchanged during testing. Extensive experiments are carried out on fine-grained visual categorization benchmarks, and the results demonstrate that our approach can significantly improve the categorization performance when the training data is highly insufficient. Jiwei Wei, Yang Yang 0002, Xiang Guan, Xing Xu 0001, Guoqing Wang 0001, Heng Tao Shen |
IEEE Trans. Multim. | 5 |
| 2024 | Align and Retrieve: Composition and Decomposition Learning in Image Retrieval With Text FeedbackabstractWe study the task of image retrieval with text feedback, where a reference image and modification text are composed to retrieve the desired target image. To accomplish this goal, existing methods always get the multimodal representations through different feature encoders and then adopt different strategies to model the correlation between the composed inputs and the target image. However, the multimodal query brings more challenges as it requires not only the synergistic understanding of the semantics from the heterogeneous multimodal inputs but also the ability to accurately build the underlying semantic correlation existing in each inputs-target triplet, i.e., reference image, modification text, and target image. In this paper, we tackle these issues with a novel Align and Retrieve (AlRet) framework. First, our proposed methods employ the contrastive loss in the feature encoders to learn meaningful multimodal representation while making the subsequent correlation modeling process in a more harmonious space. Then we propose to learn the accurate correlation between the composed inputs and target image in a novel composition-and-decomposition paradigm. Specifically, the composition network couples the reference image and modification text into a joint representation to learn the correlation between the joint representation and target image. The decomposition network conversely decouples the target image into visual and text subspaces to exploit the underlying correlation between the target image with each query element. The composition-and-decomposition paradigm forms a closed loop, which can be optimized simultaneously to promote each other in the performance. Massive comparison experiments on three real-world datasets confirm the effectiveness of the proposed method. Yahui Xu, Yi Bin, Jiwei Wei, Yang Yang 0002, Guoqing Wang 0001, Heng Tao Shen |
IEEE Trans. Multim. | 5 |
| 2023 | Learning Semantic-Aware Knowledge Guidance for Low-Light Image EnhancementabstractLow-light image enhancement (LLIE) investigates how to improve illumination and produce normal-light images. The majority of existing methods improve low-light images via a global and uniform manner, without taking into account the semantic information of different regions. Without semantic priors, a network may easily deviate from a region's original color. To address this issue, we propose a novel semantic-aware knowledge-guided framework (SKF) that can assist a low-light enhancement model in learning rich and diverse priors encapsulated in a semantic segmentation model. We concentrate on incorporating semantic knowledge from three key aspects: a semantic-aware embedding module that wisely integrates semantic priors in feature representation space, a semantic-guided color histogram loss that preserves color consistency of various instances, and a semantic-guided adversarial loss that produces more natural textures by semantic priors. Our SKF is appealing in acting as a general framework in LLIE task. Extensive experiments show that models equipped with the SKF significantly outperform the baselines on multiple datasets and our SKF generalizes to different models and scenes well. The code is available at Semantic-Aware-Low-Light-Image-Enhancement. Yuhui Wu 0001, Guoqing Wang 0001, Yang Yang 0002, Jiwei Wei, Chongyi Li, Heng Tao Shen |
CVPR | 3 |
| 2023 | Cross-Subject Mental Fatigue Detection based on Separable Spatio-Temporal Feature AggregationabstractCross-subject mental fatigue detection via Electroencephalography (EEG) is challenging because EEG from different individuals varies greatly. Existing works have exploited domain adaption to alleviate the individual discrepancy due to personality, gender and so on. However, the distributions of data from new subjects and old ones are aligned by deceiving the domain discriminator. An inevitable issue of such a paradigm is that the samples near the decision boundary are easy to be misclassified. To address this issue, we propose a Separable Spatio-temporal Feature Aggregation (SSFA) that consists of a Spatio-temporal Feature Extractor (SFE) and a Separable Feature Aggregation mechanism (SFA). Specifically, SFE utilizes the spatio-temporal information in EEG and automatically tune the weights of temporal and spatial features, so as to update the model along the optimal direction and obtain more discriminative features. In addition, SFA employs two classifiers combined with sliced Wasserstein Discrepancy to aggregate each separate class together, facilitating the mapping of the new subjects to the support region of the old subjects. Leave-one-subject-out experiments conducted on a public fatigue dataset show that the proposed method performs better than state-of-the-art on many evaluation metrics especially with an accuracy of 85.91%. Yalan Ye, Yutuo He, Wanjing Huang, Qiaosen Dong, Guoqing Wang 0001 |
ICASSP | 6 |
| 2023 | Region-Aware Semantic Consistency for Unsupervised Domain-Adaptive Semantic SegmentationabstractAs acquiring pixel-wise labels for semantic segmentation is labor-intensive, unsupervised domain adaptation (UDA) techniques aim to transfer knowledge from synthetic data to real-scene data. To overcome the distribution misalignment between the source domain and the target domain, Teacher-Student (TS) methods are widely-used and promising. In TS methods, the student resorts to the one-hot pseudo labels generated by the teacher. However, the generated one-hot pseudo labels are dubious and ignore the semantic correlation among classes. Besides, in the same position of the same image, the output distributions between the student and the teacher should be consistent. Such prediction consistency is defined as Region-Aware Semantic Consistency (RASC). Correspondingly, we propose an RASC module to assimilate the output distributions of the teacher and the student. Our RASC module is flexible and easily plugged into TS state-of-the-arts (SOTAs) based on either CNNs or Transformers. Yixuan Zhou 0001, Xing Xu 0001, Guoqing Wang 0001, Fumin Shen, Yang Yang 0002 |
ICME | 4 |
| 2023 | Faster Video Moment Retrieval with Point-Level SupervisionabstractVideo Moment Retrieval (VMR) aims at retrieving the most relevant events from an untrimmed video with natural language queries. Existing VMR methods suffer from two defects: (1) massive expensive temporal annotations are required to obtain satisfying performance; (2) complicated cross-modal interaction modules are deployed, which lead to high computational cost and low efficiency for the retrieval process. To address these issues, we propose a novel method termed Cheaper and Faster Moment Retrieval (CFMR), which balances the retrieval accuracy, efficiency, and annotation cost for VMR. Specifically, our proposed CFMR method learns from point-level supervision where each annotation is a single frame randomly located within the target moment. Such a labeling strategy achieves 6 times cheaper than the conventional annotations of event boundaries. Furthermore, we also design a concept-based multimodal alignment mechanism to bypass the usage of cross-modal interaction modules during the inference process, remarkably improving retrieval efficiency. The experimental results on three widely used VMR benchmarks demonstrate our proposed CFMR method achieves superior comprehensive performance to current state-of-the-art methods. Moreover, it significantly accelerates the retrieval speed with more than 100 times FLOPs compared to existing approaches with point-level supervision. Our open-source implementation is available at https://github.com/CFM-MSG/Code_CFMR. Xun Jiang 0001, Zailei Zhou, Xing Xu 0001, Yang Yang 0002, Guoqing Wang 0001, Heng Tao Shen |
ACM Multimedia | 5 |
| 2023 | Multimodal Physiological Signals Fusion for Online Emotion RecognitionabstractMultimodal physiological-based emotion recognition is one of the most available but challenging studies due to complexity of emotions and individual differences in physiological signals. However, existing studies mainly combine multimodal data to fuse multimodal information in offline scenarios, ignoring data/modalities correlation among multimodal data and individual differences of non-stationary physiological signals in online scenarios. In this paper, we propose a novel Online Multimodal HyperGraph Learning (OMHGL) method to fuse multimodal information for emotion recognition based on time-series physiological signals. Our method consists of multimodal hypergraph fusion and online hypergraph learning. Specifically, the multimodal hypergraph fusion can fuse multimodal physiological signals to effectively obtain emotionally dependent information via leveraging multimodal information and higher-order correlations among multimodal data/modalities. The online hypergraph learning is designed to learn new information from online data by updating hypergraph projection. As a result, the proposed online emotion recognition model can be more effective for emotion recognition of target subjects when target data arrive in an online manner. Experimental results have demonstrated that the proposed method significantly outperforms the baselines and compared state-of-the-art methods in online emotion recognition tasks. Tongjie Pan, Yalan Ye, Hecheng Cai, Shudong Huang, Yang Yang 0002, Guoqing Wang 0001 |
ACM Multimedia | 6 |
| 2023 | Cross-modal Consistency Learning with Fine-grained Fusion Network for Multimodal Fake News DetectionabstractPrevious studies on multimodal fake news detection have observed the mismatch between text and images in the fake news and attempted to explore the consistency of multimodal news based on global features of different modalities. However, they fail to investigate this relationship between fine-grained fragments in multimodal content. To gain public trust, fake news often includes relevant parts in the text and the image, making such multimodal content appear consistent. Using global features may suppress potential inconsistencies in irrelevant parts. Therefore, in this paper, we propose a novel Consistency-learning Fine-grained Fusion Network (CFFN) that separately explores the consistency and inconsistency from high-relevant and low-relevant word-region pairs. Specifically, for a multimodal post, we divide word-region pairs into high-relevant and low-relevant parts based on their relevance scores. For the high-relevant part, we follow the cross-modal attention mechanism to explore the consistency. For low-relevant part, we calculate inconsistency scores to capture inconsistent points. Finally, a selection module is used to choose the primary clue (consistency or inconsistency) for identifying the credibility of multimodal news. Extensive experiments on two public datasets demonstrate that our CFFN substantially outperforms all the baselines. Our code can be found at: https://github.com/uestc-lj/CFFN/. Jun Li 0112, Yi Bin, Jie Zou 0001, Jiwei Wei, Guoqing Wang 0001, Yang Yang 0002 |
MMAsia | 5 |
| 2023 | Hypercomplex context guided interaction modeling for scene graph generation
Zheng Wang 0044, Xing Xu 0001, Yadan Luo, Guoqing Wang 0001, Yang Yang 0002 |
Pattern Recognit. | 4 |
| 2023 | Visual Embedding Augmentation in Fourier Domain for Deep Metric LearningabstractDeep Metric Learning (DML) is very effective for many computer vision applications such as image retrieval or cross-modal matching. The common paradigm for DML is to seek metric spaces that can encode semantically similar objects close while locating the dissimilar ones far away from each other. To make features more discriminative, the mainstream methods usually design various specific loss functions to seek the help of hard negatives through complex hard mining strategies or hard synthesizing with additional networks. In spite of their fruitfulness, these approaches ignore the impact of low-level information in images on the performance, which may degrade the discerning ability of learned embedding. To alleviate these problems, we introduce a simple yet effective augmentation method to generate more hard negatives by swapping the low-frequency spectra of negative instances with anchors in the Fourier domain. Specifically, unlike previous methods, our proposed approach does not involve any complex design strategies but enriches hard negatives by manipulating the low-level variability of images only with simple Fourier transforms. In addition, our method is treated as a universal plug-in, which can be incorporated into different models for performance improvement. In the end, we conduct extensive experiments to evaluate our method on the widely-used datasets including CUB-200–2011, CARS-196, and Stanford Online Products. Our quantitative results demonstrate that the proposed plug-in outperforms previous approaches consistently and significantly across different datasets and evaluation metrics. Zheng Wang 0044, Zhenwei Gao, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Less is Better: Exponential Loss for Cross-Modal MatchingabstractDeep metric learning has become a key component of cross-modal retrieval. By learning to pull the features of matched instances closer while pushing the features of mismatched instances farther away, one can learn highly robust multi-modal representations. Most existing cross-modal retrieval methods leverage vanilla triplet loss to train the network, which cannot adaptively penalize pairs with different hardness. Although various weighting strategies have been designed for unimodal matching tasks, few weighting strategies have been applied to cross-modal tasks due to the specificity of cross-modal tasks. While few weighting strategies are designed for cross-modal scenarios, they usually involve a lot of hyper-parameters, which require a lot of computational resources to fine-tune. In this paper, we introduce a new exponential loss, which can assign appropriate weights to individual positive and negative pairs according to their similarity so that it can adaptively penalize pairs with different hardness. Furthermore, the exponential loss has only two hyper-parameters, making it easier to find the optimal parameters to suit various data distributions in practice. Exponential loss can be universally applied to well-established cross-modal models and further boost their retrieval performance. We exhaustively ablate our method on Image-Text matching, Video-Text matching, as well as unimodal Image matching. Experimental results show that a standard model trained with exponential loss can achieve noticeable performance gains. Jiwei Wei, Yang Yang 0002, Xing Xu 0001, Jingkuan Song, Guoqing Wang 0001, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Physics Guided Remote Sensing Image Synthesis Network for Ship DetectionabstractAutomatic detection and localization of objects in remote sensing images are of great significance for remote sensing systems. Existing frameworks usually train an object detection network using collected remote sensing images. However, these models usually perform poorly due to the lack of large-scale training datasets, which is often the case for special remote sensing scenarios, e.g., the detection of ships in the open sea. Although image synthesis is a common strategy to alleviate the issue of data insufficiency, the trained model still performs poorly when being tested on real-world scenes. Aimed at this, a novel sensor-related image synthesis framework, dubbed as remote sensing-image synthesis pipeline (RS-ISP), is developed to address the lack of on-orbit remote sensing images. Specifically, our RS-ISP introduces two novel designs to ensure the distribution consistency between the generated images and the real images: 1) the first is a novel pipeline for modeling the physical process of noise production during image capture using specific sensors and 2) the second is the design of a detection-oriented image harmonization model. Similar to the existing design, our model first produces coarse synthetic images by copy–paste operation, on which the proposed harmonization process is used to reduce the variation in the pasted foreground and background. By incorporating these two designs into a unified framework, our RS-ISP is designed and used to produce large-scale synthetic images used to train the object detection model for detecting ships in remote sensing images. Comparative experiments demonstrated that RS-ISP increased the [email protected] from 0.148 to 0.498 for the ship detection task. Code will be publicly available. Weichang Zhang, Rui Zhang 0138, Guoqing Wang 0001, Yang Yang 0002, Die Hu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Modeling Lateral Control Behaviors of Distracted Drivers for Haptic-Shared Steering SystemabstractHaptic-shared steering (HSS) systems have been reported to enhance vehicle safety and reduce the workload for drivers. However, few studies have focused on modeling the lateral vehicle control behaviors of distracted drivers for HSS. The current study models this type of behavior in a series of high-fidelity driving simulator experiments with 18 participants. Two experimental conditions for a double lane change task are tested: HGT-Constant (haptic guidance torque with a constant gain) and HGT-Adaptive (haptic guidance torque with an adaptive gain). A gated recurrent unit (GRU) network is used to model lateral control behavior during driving. The effectiveness of the GRU network-based lateral control model of distracted drivers is benchmarked using a state-of-the-art long short-term memory network, a back propagation network, an extreme learning machine, and a traditional two-point visual model with neuromuscular dynamics. Experimental results indicate that the GRU network has the highest accuracy in terms of root mean square error, mean absolute error, mean absolute percentage error, and determination coefficient. In addition, a simulation model of the driver-vehicle-road closed-loop system demonstrates that the proposed model predicts driver behavior with acceptable lateral position error. Feixiang Xu, Shi-Yong Feng, Edric John Cruz Nacpil, Zheng Wang 0039, Guoqing Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2023 | Quaternion Relation Embedding for Scene Graph GenerationabstractAs an important visual understanding task, scene graph generation has been drawing widespread attention and could boost a broad range of downstream vision applications. Traditional scene graph generation methods based on different context refinements are trained with probabilistic chain rule, which treats objects and relationships as independent entities. Despite their surprisingly great progress, such a plain formulation unconsciously ignores the latent geometric structure of entities and relationships. To address this issue, we move beyond the traditional real-valued representations and useQuaternionRelationEmbedding (QuatRE) to generate scene graphs with more expressive hypercomplex representations. More specifically, we introduce the concept of quaternion representations, hyper-complex valued with three imaginary components for objects entities, then formulate the relation triplets with Hamilton product. Benefiting from explicitly modeling the latent inter-dependencies among all imaginary components and strong expressive capacity, our proposed QuatRE method could better capture the interactions between entities. More importantly, our novel QuatRE method can be treated as a plug-in and well generalized into other methods for performance improvement as it involves no additional layers. Finally, extensive comparisons of our proposed method against the state-of-the-art methods on two large-scale and widely-used datasets, i.e. Visual Genome and Open Images, demonstrated our superiority and generalization capability on various metrics for biased or unbiased inference. Zheng Wang 0044, Xing Xu 0001, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2023 | Multi-Modal Transformer With Global-Local Alignment for Composed Query Image RetrievalabstractIn this paper, we study the composed query image retrieval, which aims at retrieving the target image similar to the composed query, i.e., a reference image and the desired modification text. Compared with conventional image retrieval, this task is more challenging as it not only requires precisely aligning the composed query and target image in a common embedding space, but also simultaneously extracting related information from the reference image and modification text. In order to properly extract related information from the composed query, existing methods usually embed vision-language inputs using different feature encoders, e.g., CNN for images and LSTM/BERT for text, and then employ a complicated manually-designed composition module for learning the joint image-text representation. However, the architecture discrepancy in feature encoders would restrict the vision-language plenitudinous interaction. Meanwhile, certain complicated composition designs might significantly hamper the generalization ability of the model. To tackle these problems, we propose a new framework termed ComqueryFormer, which effectively processes the composed query with the Transformer for this task. Specifically, to eliminate the architecture discrepancy, we leverage a unified transformer-based architecture to homogeneously encode the vision-language inputs. Meanwhile, instead of the complicated composition module, the neat yet effective cross-modal transformer is adopted to hierarchically fuse the composed query at various vision scales. On the other hand, we introduce an efficient global-local alignment module to narrow the distance between the composed query and the target image. It not only considers the divergence in the global joint embedding space but also forces the model to focus on the local detail differences. Extensive experiments on three real-world datasets demonstrate the superiority of our ComqueryFormer. Our code can be found at:https://github.com/uestc-xyh/ComqueryFormer. Yahui Xu, Yi Bin, Jiwei Wei, Yang Yang 0002, Guoqing Wang 0001, Heng Tao Shen |
IEEE Trans. Multim. | 5 |
| 2022 | NAS-StegNet: Lightweight Image Steganography Networks via Neural Architecture Search
Zhixian Wang, Guoqing Wang 0001, Yang Yang 0002 |
ICONIP (3) | 2 |
| 2022 | ARRA: Absolute-Relative Ranking Attack against Image RetrievalabstractWith the extensive application of deep learning, adversarial attacks especially query-based attacks receive more concern than ever before. However, the scenarios assumed by existing query-based attacks against image retrieval are usually too simple to satisfy the attack demand. In this paper, we propose a novel method termed Absolute-Relative Ranking Attack (ARRA) that considers a more practical attack scenario. Specifically, we propose two compatible goals for the query-based attack, i.e., absolute ranking attack and relative ranking attack, which aim to change the relative order of chosen candidates and assign the specific ranks to chosen candidates in retrieval list respectively. We further devise the Absolute Ranking Loss (ARL) and Relative Ranking Loss (RRL) for the above goals and implement our ARRA by minimizing their combination with black-box optimizers and evaluate the attack performance by attack success rate and normalized ranking correlation. Extensive experiments conducted on widely-used SOP and CUB-200 datasets demonstrate the superiority of the proposed approach over the baselines. Moreover, the attack result on a real-world image retrieval system, i.e., Huawei Cloud Image Search, also proves the practicability of our ARRA approach. Xing Xu 0001, Zailei Zhou, Yang Yang 0002, Guoqing Wang 0001, Heng Tao Shen |
ACM Multimedia | 5 |
| 2022 | Rethinking Open-World Object Detection in Autonomous Driving ScenariosabstractExisting object detection models have been demonstrated to successfully discriminate and localize the predefined object categories under the seen or similar situations. However, the open-world object detection as required by autonomous driving perception systems refers to recognizing unseen objects under various scenarios. On the one hand, the knowledge gap between seen and unseen object categories poses extreme challenges for models trained with supervision only from the seen object categories. On the other hand, the domain differences across different scenarios also cause an additional urge to take the domain gap into consideration by aligning the sample or label distribution. Aimed at resolving these two challenges simultaneously, we firstly design a pre-training model to formulate the mappings between visual images and semantic embeddings from the extra annotations as guidance to link the seen and unseen object categories through a self-supervised manner. Within this formulation, the domain adaptation is then utilized for extracting the domain-agnostic feature representations and alleviating the misdetection of unseen objects caused by the domain appearance changes. As a result, the more realistic and practical open-world object detection problem is visited and resolved by our novel formulation, which could detect the unseen categories from unseen domains without any bounding box annotations while there is no obvious performance drop in detecting the seen categories. We are the first to formulate a unified model for open-world task and establish a new state-of-the-art performance for this challenge. Zeyu Ma 0002, Yang Yang 0002, Guoqing Wang 0001, Xing Xu 0001, Heng Tao Shen |
ACM Multimedia | 3 |
| 2021 | Learning Hierarchal Channel Attention for Fine-grained Visual ClassificationabstractLearning delicate feature representation of object parts plays a critical role in fine-grained visual classification tasks. However, advanced deep convolutional neural networks trained for general visual classification tasks usually tend to focus on the coarse-grained information while ignoring the fine-grained one, which is of great significance for learning discriminative representation. In this work, we explore the great merit of multi-modal data in introducing semantic knowledge and sequential analysis techniques in learning hierarchical feature representation for generating discriminative fine-grained features. To this end, we propose a novel approach, termed Channel Cusum Attention ResNet (CCA-ResNet ), for multi-modal joint learning of fine-grained representation. Specifically, we use feature-level multi-modal alignment to connect image and text classification models for joint multi-modal training. Through joint training, image classification models trained with semantic level labels tend to focus on the most discriminative parts, which enhances the cognitive ability of the model. Then, we propose a Channel Cusum Attention (CCA ) mechanism to equip feature maps with hierarchical properties through unsupervised reconstruction of local and global features. The benefits brought by the CCA are in two folds: a) allowing fine-grained features from early layers to be preserved in the forward propagation of deep networks; b) leveraging the hierarchical properties to facilitate multi-modal feature alignment. We conduct extensive experiments to verify that our proposed model can achieve state-of-the-art performance on a series of fine-grained visual classification benchmarks. Xiang Guan, Guoqing Wang 0001, Xing Xu 0001, Yi Bin |
ACM Multimedia | 2 |
| 2021 | Imbalanced Source-free Domain AdaptationabstractConventional Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a well-labeled source domain to an unlabeled target domain only when data from both domains is simultaneously accessible, which is challenged by the recent Source-free Domain Adaptation (SFDA). However, we notice that the performance of existing SFDA methods would be dramatically degraded by intra-domain class imbalance and inter-domain label shift. Unfortunately, class-imbalance is a common phenomenon in real-world domain adaptation applications. To address this issue, we present Imbalanced Source-free Domain Adaptation (ISFDA) in this paper. Specifically, we first train a uniformed model from the source domain, and then propose secondary label correction, curriculum sampling, plus intra-class tightening and inter-class separation to overcome the joint presence of covariate shift and label shift. Extensive experiments on three imbalanced benchmarks verify that ISFDA could perform favorably against existing UDA and SFDA methods under various conditions of class-imbalance, and outperform existing SFDA methods by over 15% in terms of per-class average accuracy on a large-scale long-tailed imbalanced dataset. Xinhao Li 0002, Jingjing Li 0001, Lei Zhu 0002, Guoqing Wang 0001, Zi Huang |
ACM Multimedia | 4 |
| 2021 | Progressive Graph Attention Network for Video Question AnsweringabstractVideo question answering~(Video-QA) is a task of answering a natural language question related to the content of a video. Existing methods generally explore the single interactions between objects or between frames, which are insufficient to deal with the sophisticated scenes in videos. To tackle this problem, we propose a novel model, termed Progressive Graph Attention Network (PGAT), which can jointly explore the multiple visual relations on object-level, frame-level and clip-level. Specifically, in the object-level relation encoding, we design two kinds of complementary graphs, one for learning the spatial and semantic relations between objects from the same frame, the other for modeling the temporal relations between the same object from different frames. The frame-level graph explores the interactions between diverse frames to record the fine-grained appearance change, while the clip-level graph models the temporal and semantic relations between various actions from clips. These different-level graphs are concatenated in a progressive manner to learn the visual relations from low-level to high-level. Furthermore, we for the first time identified that there are serious answer biases with TGIF-QA, a very large Video-QA dataset, and reconstructed a new dataset based on it to overcome the biases, called TGIF-QA-R. We evaluate the proposed model on three benchmark datasets and the new TGIF-QA-R, and the experimental results demonstrate that our model significantly outperforms other state-of-the-art models. Our codes and dataset are available at https://github.com/PengLiang-cn/PGAT. Shuangji Yang, Yi Bin, Guoqing Wang 0001 |
ACM Multimedia | 4 |
| 2021 | Disentangled Representation Learning and Enhancement Network for Single Image De-RainingabstractIn this paper, we present a disentangled representation learning and enhancement network (DRLE-Net) to address the challenging single image de-raining problems, i.e., raindrop and rain streak removal. Specifically, the DRLE-Net is formulated as a multi-task learning framework, and an elegant knowledge transfer strategy is designed to train the encoder of DRLE-Net to embed a rainy image into two separated latent spaces representing the task (clean image reconstruction in this paper) relevant and irrelevant variations respectively, such that only the essential task-relevant factors will be used by the decoder of DRLE-Net to generate high-quality de-raining results. Furthermore, visual attention information is modeled and fed into the disentangled representation learning network to enhance the task-relevant factor learning. To facilitate the optimization of the hierarchical network, a new adversarial loss formulation is proposed and used together with the reconstruction loss to train the proposed DRLE-Net. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and DRLE-Net is demonstrated to produce significantly better results than state-of-the-art models. Guoqing Wang 0001, Changming Sun, Xing Xu 0001, Jingjing Li 0001, Zheng Wang 0044, Zeyu Ma 0002 |
ACM Multimedia | 1 |
| 2021 | Meta Self-Paced Learning for Cross-Modal MatchingabstractCross-modal matching has attracted growing attention due to the rapid emergence of the multimedia data on the web and social applications. Recently, many re-weighting methods have been proposed for accelerating model training by designing a mapping function from similarity scores to weights. However, these re-weighting methods are difficult to be universally applied in practice since manually pre-set weighting functions inevitably involve hyper-parameters. In this paper, we propose a Meta Self-Paced Network (Meta-SPN) that automatically learns a weighting scheme from data for cross-modal matching. Specifically, a meta self-paced network composed of a fully connected neural network is designed to fit the weight function, which takes the similarity score of the sample pairs as input and outputs the corresponding weight value. Our meta self-paced network considers not only the self-similarity scores, but also their potential interactions (e.g., relative-similarity) when learning the weights. Motivated by the success of meta-learning, we use the validation set to update the meta self-paced network during the training of the matching network. Experiments on two image-text matching benchmarks and two video-text matching benchmarks demonstrate the generalization and effectiveness of our method. Jiwei Wei, Xing Xu 0001, Zheng Wang 0044, Guoqing Wang 0001 |
ACM Multimedia | 4 |
| 2021 | PFFN: Progressive Feature Fusion Network for Lightweight Image Super-ResolutionabstractRecently, convolutional neural network (CNN) has been the core ingredient of modern models, triggering the surge of deep learning in super-resolution (SR). Despite the great success of these CNN-based methods which are prone to be deeper and heavier, it is impracticable to directly apply these methods for some low-budget devices due to the superfluous computational overhead. To alleviate this problem, a novel lightweight SR network named progressive feature fusion network (PFFN) is developed to seek for better balance between performance and running efficiency. Specifically, to fully exploit the feature maps, a novel progressive attention block (PAB) is proposed as the main building block of PFFN. The proposed PAB adopts several parallel but connected paths with pixel attention, which could significantly increase the receptive field of each layer, distill useful information and finally learn more discriminative feature representations. In PAB, a powerful dual attention module (DAM) is further incorporated to provide the channel and spatial attention mechanism in fairly lightweight manner. Besides, we construct a pretty concise and effective upsampling module with the help of multi-scale pixel attention, named MPAU. All of the above modules ensure the network can benefit from attention mechanism while still being lightweight enough. Furthermore, a novel training strategy following the cosine annealing learning scheme is proposed to maximize the representation ability of the model. Comprehensive experiments show that our PFFN achieves the best performance against all existing lightweight state-of-the-art SR methods with less number of parameters and even performs comparably to computationally expensive networks. Dongyang Zhang 0001, Changyu Li, Ning Xie 0003, Guoqing Wang 0001, Jie Shao 0001 |
ACM Multimedia | 4 |
| 2021 | Hierarchical Composition Learning for Composed Query Image RetrievalabstractComposed query image retrieval is a growing research topic. The object is to retrieve images not only generally resemble the reference image, but differ according to the desired modification text. Existing methods mainly explore composing modification text with global feature or local entity descriptor of reference image. However, they ignore the fact that modification text is indeed diverse and arbitrary. It not only relates to abstractive global feature or concrete local entity transformation, but also often associates with the fine-grained structured visual adjustment. Thus, it is insufficient to emphasize the global or local entity visual for the query composition. In this work, we tackle this task by hierarchical composition learning. Specifically, the proposed method first encodes images into three representations consisting of global, entity and structure level representations. Structure level representation is richly explicable, which explicitly describes entities as well as attributes and relationships in the image with a directed graph. Based on these, we naturally perform hierarchical composition learning by fusing modification text and reference image in the global-entity-structure manner. It can transform the visual feature conditioned on modification text to target image in a coarse-to-fine manner, which takes advantage of the complementary information among three levels. Moreover, we introduce a hybrid space matching to explore global, entity and structure alignments which can get high performance and good interpretability. Yahui Xu, Yi Bin, Guoqing Wang 0001, Yang Yang 0002 |
MMAsia | 3 |
| 2021 | Context-Enhanced Representation Learning for Single Image Deraining
Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
Int. J. Comput. Vis. | 1 |
| 2021 | Multi-Scale Deep Representation Aggregation for Vein RecognitionabstractThe recent success of Deep Convolutional Neural Network (DCNN) for various computer vision tasks such as image recognition has already demonstrated its robust feature representation ability. However, the limitation of training database on small scale vein recognition tasks restricts its performance because the recognition result of DCNN depends heavily on the number of trainsets. This motivates the design of a Multi-Scale Deep Representation Aggregation (MSDRA) model based on a pre-trained DCNN for vein recognition. First, the multi-scale feature maps are extracted by a pre-trained DCNN model. Second, a local mean threshold approach is designed to preliminarily remove the noisy information of multi-scale feature maps and generate the selected feature maps. Third, we propose an Unsupervised Vein Information Mining (UVIM) method to localize vein information of selected feature maps for generating a binary vein information mask, and then the vein information mask is utilized to keep useful deep representation and discard the background information. Finally, the discriminative multi-scale deep representations, which are generated by using the vein information mask to aggregate multi-scale feature maps, are concatenated into the final compact feature vectors, and then a Support Vector Machine (SVM) is introduced for final recognition. Our proposed model outperforms the state-of-the-art methods on two benchmark vein databases. Moreover, an additional experiment using the subset of PolyU Palmprint database illustrates the system's generalization ability and robustness. Zaiyu Pan, Jun Wang 0071, Guoqing Wang 0001, Jihong Zhu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Attentive Feature Refinement Network for Single Rainy Image RestorationabstractDespite the fact that great progress has been made on single image deraining tasks, it is still challenging for existing models to produce satisfactory results directly, and it often requires a single or multiple refinement stages to gradually improve the quality. However, in this paper, we demonstrate that existing image-level refinement with a stage-independent learning design is problematic with the side effect of over/under-deraining. To resolve this issue, we for the first time propose the mechanism of learning to carry out refinement on the unsatisfactory features, and propose a novel attentive feature refinement (AFR) module. Specifically, AFR is designed as a two-branched network for simultaneous rain-distribution-aware attention map learning and attention guided hierarchy-preserving feature refinement. Guided by task-specific attention, coarse features are progressively refined to better model the diversified rainy effects. By using a separable convolution as the basic component, our AFR module introduces little computation overhead and can be readily integrated into most rainy-to-clean image translation networks for achieving better deraining results. By incorporating a series of AFR modules into a general encoder-decoder network, AFR-Net is constructed for deraining and it achieves new state-of-the-art results on both synthetic and real images. Furthermore, by using AFR-Net as a teacher model, we explore the use of knowledge distillation to successfully learn a student model that is also able to achieve state-of-the-art results but with a much faster inference speed (i.e., it only takes 0.08 second to process a 512×512 rainy image). Code and pre-trained models are available at 〈 https://github.com/RobinCSIRO/AFR-Net 〉 . Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-Weighted Co-Occurrence Descriptor Encoding for Vein RecognitionabstractDespite being highly secure, vein recognition suffers from the high inter-class similarity and intra-class variation resulting from the uncontrolled image capture, making the design of discriminative and robust representation very important. The recent success of convolutional neural network (CNN) for various image understanding tasks makes it a promising method for feature extraction. However, limited variability in small-scale datasets leads to systems derived from the direct training or fine-tuning not transferable and unreliable for practical biometric applications. This motivates the design of a multi-weighted co-occurrence descriptor encoding (MWCDE) model for vein recognition. Instead of directly conducting a feed-forward operation with a pre-trained CNN for obtaining the semantic features from the fully connected layers, co-occurrence features among convolutional filters are modeled first in MWCDE by a simple convolution between an indicator filter in a higher layer with a to-be-reweighted filter in a lower layer, and a redundancy-driven indicator filter selection algorithm is designed for filtering out some ambiguous representations. Second, another hard feature weighting strategy with a binary masking scheme is proposed for discarding noisy background and feature redundancy. The selected high-order descriptors are then embedded and aggregated into the compact feature vectors with a saliency driven spatial weighted Fisher vector algorithm, followed by the introduction of a generalized support vector machine for recognition. Extensive experiments with three benchmark vein datasets demonstrate that the proposed framework can achieve state-of-the-art results, and an additional experiment with the PolyU multispectral palmprint database illustrates its generalization ability. Code is available at (https://github.com/RobinCSIRO/MWCDE-for-Vein-Recognition). Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Learning a Compact Vein Discrimination Model With GANerated SamplesabstractDespite the great success achieved by convolutional neural networks (CNNs) in various image understanding tasks, it is still difficult for CNNs to be applied to vein recognition tasks due to the problems of insufficient training datasets, intra-class variations, and inter-class similarities. Besides, due to the essential requirement on the storage of millions of parameters for CNN, it is challenging to use a CNN for designing a vein-based embedded person identification system. In this paper, these two problems are addressed by learning a discriminative and compact vein recognition model. For the first problem, a hierarchical generative adversarial network (HGAN) consisting of a constrained CNN and a CycleGAN is proposed for data augmentation. Two similarity losses are defined for estimating the self-similarity and inter-class dissimilarity, and a CycleGAN model is properly trained with these two losses for better task-specific training sample generation. After obtaining a baseline vein recognition model fine-tuned on the augmented datasets, the existence of parameter redundancy in the over-parameterized network motivates the proposal of model compression by way of filter pruning and low rank approximation, thus making the compressed model more suitable for deployment on embedded systems. Through the vein recognition experiments with two different datasets and an additional palmprint recognition experiment, the proposed algorithms are shown to yield a highly compact model while keeping the accuracy acceptable for application. Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | Cascaded Attention Guidance Network for Single Rainy Image RestorationabstractRestoring a rainy image with raindrops or rainstreaks of varying scales, directions, and densities is an extremely challenging task. Recent approaches attempt to leverage the rain distribution (e.g., location) as prior to generate satisfactory results. However, concatenation of a single distribution map with the rainy image or with intermediate feature maps is too simplistic to fully exploit the advantages of such priors. To further explore this valuable information, an advanced cascaded attention guidance network, dubbed as CAG-Net, is formulated and designed as a three-stage model. In the first stage, a multitask learning network is constructed for producing the attention map and coarse de-raining results simultaneously. Subsequently, the coarse results and the rain distribution map are concatenated and fed to the second stage for results refinement. In this stage, the attention map generation network from the first stage is used to formulate a novel semantic consistency loss for better detail recovery. In the third stage, a novel pyramidal "whereand- how" learning mechanism is formulated. At each pyramid level, a two-branch network is designed to take the features from previous stages as inputs to generate better attention-guidance features and de-raining features, which are then combined via a gating scheme to produce the final de-raining results. Moreover, the uncertainty maps are also generated in this stage for more accurate pixel-wise loss calculation. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and CAG-Net is demonstrated to produce significantly better results than state-of-the-art models. Code will be publicly available after paper acceptance. Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Image Process. | 1 |
| 2019 | ERL-Net: Entangled Representation Learning for Single Image De-RainingabstractDespite the significant progress achieved in image de-raining by training an encoder-decoder network within the image-to-image translation formulation, blurry results with missing details indicate the deficiency of the existing models. By interpreting the de-raining encoder-decoder network as a conditional generator, within which the decoder acts as a generator conditioned on the embedding learned by the encoder, the unsatisfactory output can be attributed to the low-quality embedding learned by the encoder. In this paper, we hypothesize that there exists an inherent mapping between the low-quality embedding to a latent optimal one, with which the generator (decoder) can produce much better results. To improve the de-raining results significantly over existing models, we propose to learn this mapping by formulating a residual learning branch, that is capable of adaptively adding residuals to the original low-quality embedding in a representation entanglement manner. Using an embedding learned this way, the decoder is able to generate much more satisfactory de-raining results with better detail recovery and rain artefacts removal, providing new state-of-the-art results on four benchmark datasets with considerable improvement (i.e., on the challenging Rain100H data, an improvement of 4.19dB on PSNR and 5% on SSIM is obtained). The entanglement can be easily adopted into any encoder-decoder based image restoration networks. Besides, we propose a series of evaluation metrics to investigate the specific contribution of the proposed entangled representation learning mechanism. Codes are available at 〈https://github.com/RobinCSIRO/ERL-Net-for-Single-Image-Deraining〉. Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
ICCV | 1 |
| 2018 | Bimodal Vein Data Mining via Cross-Selected-Domain Knowledge TransferabstractRecent success in large-scale image recognition challenge (i.e., ImageNet) fully demonstrates the capability of deep neural network (DNN) in learning complex and semantic representation, and this also motivates the generation of transfer learning model, which fine-tunes state-of-the-art DNN models with other small-scale databases for better performance. Driven by such an idea, a task-specific DNN model fine-tuned from VGG-face is constructed for both gender and identity recognition with hand vein information. Unlike the traditional transfer learning models, which fine-tune directly from source to target, we leverage the coarse-to-fine scheme to train the task-specific models in a step-aware way, such that the inherent correlation between the neighboring databases could serve as initialization base to relieve the problem of over-fitting, which is inevitable with the small-scaled hand vein database, and also speed up the convergence. Besides, the task-driven network training idea, which involves joint optimization of linear regression classifier and network parameters, is also adopted during training of each model to obtain more discriminative representation for specified tasks. Instead of adopting the trained linear regression classifier for gender and identity classification, the large margin distribution machine (LDM) is introduced to ensure the discriminative and generalization performance of the model simultaneously, and it should be noted that before feeding the gender feature vector into the LDM, a supervised feature selection step is incorporated to improve the classification performance by discarding the redundant feature and highlighting the important ones for gender classification. Rigorous experiments using the lab-made database are conducted to demonstrate the effectiveness and feasibility of the proposed model. What is more, additional experiment with a subset of the PolyU database illustrates its generalization ability and robustness. Jun Wang 0071, Guoqing Wang 0001, Mei Zhou |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Quality-Specific Hand Vein Recognition SystemabstractVein images generally appear darker with low contrast, which require contrast enhancement during preprocessing to design satisfactory hand vein recognition system. However, the modification introduced by contrast enhancement (CE) is reported to bring side effects through pixel intensity distribution adjustments. Furthermore, the inevitable results of fake vein generation or information loss occur and make nearly all vein recognition systems unconvinced. In this paper, a “CE-free” quality-specific vein recognition system is proposed, and three improvements are involved. First, a high-quality lab-vein capturing device is designed to solve the problem of low contrast from the view of hardware improvement. Then, a high quality lab-made database is established. Second, CFISH score, a fast and effective measurement for vein image quality evaluation, is proposed to obtain quality index of lab-made vein images. Then, unsupervised $K$ -means with optimized initialization and convergence condition is designed with the quality index to obtain the grouping results of the database, namely, low quality (LQ) and high quality (HQ). Finally, discriminative local binary pattern (DLBP) is adopted as the basis for feature extraction. For the HQ image, DLBP is adopted directly for feature extraction, and for the LQ one. CE_DLBP could be utilized for discriminative feature extraction for LQ images. Based on the lab-made database, rigorous experiments are conducted to demonstrate the effectiveness and feasibility of the proposed system. What is more, an additional experiment with PolyU database illustrates its generalization ability and robustness. Jun Wang 0071, Guoqing Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |