EDBT 2026 Demo / reviewers in the wild / expert
Qing Guo 0005
dblp:25/3038-5
· DBLP profile ↗
143ranked-venue papers
19as first author
120since 2021 · last 2026
0000-0003-0974-9299ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 13 first-author · 56 since 2021Artificial intelligence and machine learning · 71 · 8 first-author · 64 since 2021Security and privacy · 13 · 13 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving SystemsabstractMultimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scenarios. Existing patch-based attack methods are primarily designed for object detection models. Due to the more complex architectures and strong reasoning capabilities of MLLMs, these approaches perform poorly when transferred to MLLM-based systems. To address these limitations, we propose PhysPatch, a physically realizable and transferable adversarial patch framework tailored for MLLM-based AD systems. PhysPatch jointly optimizes patch location, shape, and content to enhance attack effectiveness and real-world applicability. It introduces a semantic-based mask initialization strategy for realistic placement, an SVD-based local alignment loss with patch-guided crop-resize to improve transferability, and a potential field-based mask refinement method. Extensive experiments across open-source, commercial, and reasoning-capable MLLMs demonstrate that PhysPatch significantly outperforms state-of-the-art (SOTA) methods in steering MLLM-based AD systems toward target-aligned perception and planning outputs. Moreover, PhysPatch consistently places adversarial patches in physically feasible regions of AD scenes, ensuring strong real-world applicability and deployability. Qi Guo 0008, Xiaojun Jia, Shanmin Pang, Simeng Qin, Lin Wang 0026, Ju Jia, Yang Liu 0003, Qing Guo 0005 |
AAAI | 8 |
| 2026 | FreeMem: Enhancing Consistency in Long Video Generation via Tuning-Free MemoryabstractText-to-Video (T2V) generation has advanced greatly, yet maintaining consistency remains challenging, especially for tuning-free long video generation. We attribute the consistency problem to cumulative deviations for long video generation at three levels: the random noise lacking correlation results initial deviation between frames; discrepancy in semantic feature tokens between denoising network blocks gradually accumulates as the frame count grows, leading to greater deviations; attention mechanisms struggle to capture global relationships across distant frames in long videos. To address these, we propose FreeMem, a tuning-free framework leveraging hierarchical memory update and injection: the noise memory stabilizes consistency by manipulating low and high frequency components in the initial noise space; the token memory combats inconsistency through adaptive fusion of historical and current semantic feature tokens between denoising network blocks; and the attention memory establishes persistent cache to model long-range relationships within self attention layers. Evaluated on VBench, FreeMem improves subject and background consistency matrics across various methods, offering a practical solution for low-cost, high-consistency long video generation. Jibin Peng, Di Lin 0002, Zhecheng Xu, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Qing Guo 0005 |
AAAI | 10 |
| 2026 | Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New ThinkingabstractIn this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential relationships among agent decisions. Since each decision influences the subsequent choices, this forms a hierarchical and nested tree-like structure of interdependencies. While modeling tree-like data in Euclidean spaces could cause distortion, which results in a loss of agent decision structure information. Motivated by this, we reconsider model agent behaviors in hyperbolic space and propose the Hyperbolic Multi-Agent Representations (HMAR) method, which projects the agent behaviors into a Poincaré ball and leverages hyperbolic neural networks to learn agent policy representations. Additionally, we designed a contrastive loss function to train this network, minimizing the distance in feature space between different representations of the same agent while maximizing the distance between representations of distinct agents. Experimental results provide empirical evidence for the effectiveness of the HMAR method in cooperative and competitive environments, demonstrating the potential of hyperbolic agent representations for effective decision-making in multi-agent environments. Bohao Qu, Xiaofeng Cao 0002, Bing Li 0001, Menglin Zhang, Tuan-Anh Vu, Di Lin 0002, Qing Guo 0005 |
AAAI | 7 |
| 2026 | MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM AgentsabstractPhysical adversarial attacks in driving scenarios can expose critical vulnerabilities in visual perception models. However, developing such attacks remains non-trivial due to diverse real-world environmental influences. Existing approaches either struggle to generalize to dynamic environments or fail to achieve consistent physical attack performance. To address these challenges, we propose MAGIC (Mastering Physical Adversarial Generation In Context), a novel framework powered by multi-modal LLM agents to automatically understand the scene context during testing time and generate adversarial patches through synergistic interaction of language and vision understanding. Specifically, MAGIC orchestrates three specialized LLM agents: the adv-patch generation agent masters the creation of deceptive patches via strategic prompt manipulation for text-to-image models; the adv-patch deployment agent ensures contextual coherence by determining optimal deployment strategies based on scene understanding; and the self-examination agent completes this trilogy by providing critical oversight and iterative refinement of both processes. We validate our approach with both digital and physical scenarios, i.e., nuImage and real-world scenes, where both statistical and visual results demonstrate that our MAGIC is powerful and effective for attacking widely applied object detection systems, such as YOLO and DETR series. Yun Xing 0001, Nhat Chung, Jie Zhang 0002, Ivor W. Tsang, Yang Liu 0003, Lei Ma 0003, Qing Guo 0005 |
AAAI | 8 |
| 2026 | Learn Like Humans: Use Meta-cognitive Reflection for Efficient Self-ImprovementabstractWhile Large Language Models (LLMs) enable complex autonomous behavior, current agents remain constrained by static, humandesigned prompts that limit adaptability.Existing self-improving frameworks attempt to bridge this gap but typically rely on inefficient, multi-turn recursive loops that incur high computational costs.To address this, we propose Metacognitive Agent Reflective Self-improvement (MARS), a framework that achieves efficient self-evolution within a single recurrence cycle.Inspired by educational psychology, MARS mimics human learning by integrating principle-based reflection (abstracting normative rules to avoid errors) and procedural reflection (deriving step-by-step strategies for success).By synthesizing these insights into optimized instructions, MARS allows agents to systematically refine their reasoning logic without continuous online feedback.Extensive experiments on six benchmarks demonstrate that MARS outperforms state-of-the-art self-evolving systems while significantly reducing computational overhead. Xinmeng Hou, Bohao Qu, Wuqi Wang, Peiliang Gong, Qing Guo 0005 |
ACL (1) | 5 |
| 2026 | From Language to Driving: A Dual-Loop SLM-Enhanced Framework for Multi-Planner Scheduling via a Domain-Specific LanguageabstractJiawei Liu, Xun Gong, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor Tsang, Yunfeng hu, Hong Chen, Qing Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xun Gong 0007, Muli Yang, Xingrui Yu, Fen Fang, Xulei Yang, Ivor W. Tsang, Yunfeng Hu 0003, Hong Chen 0003, Qing Guo 0005 |
ACL (1) | 10 |
| 2026 | Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent CuesabstractGlass is a prevalent material among solid objects in every-day life, yet segmentation methods struggle to distinguish it from opaque materials due to its transparency and reflection. While it is known that human perception relies on boundary and reflective-object features to distinguish glass objects, the existing literature has not yet sufficiently captured both properties when handling transparent objects. Hence, we propose incorporating both of these powerful visual cues via the Boundary Feature Enhancement and Reflection Feature Enhancement modules in a mutually beneficial way. Our proposed framework, TransCues, is a pyramidal transformer encoder-decoder architecture to segment transparent objects. We empirically show that these two modules can be used together effectively, improving overall performance across various benchmark datasets, including glass object semantic segmentation, mirror object semantic segmentation, and generic segmentation datasets. Our method outperforms the state-of-the-art by a large margin, achieving +4.2% mIoU on Trans10K-v2, +5.6% mIoU on MSD, +10.1% mIoU on RGBD-Mirror, +13.1% mIoU on TROSD, and +8.3% mIoU on Stanford2D3D, showing the effectiveness of our method against glass objects. Tuan-Anh Vu, Hai Nguyen-Truong, Ziqiang Zheng, Binh-Son Hua, Qing Guo 0005, Ivor W. Tsang, Sai-Kit Yeung |
WACV | 5 |
| 2026 | Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with DiffusionabstractAbstract Text-to-image diffusion techniques have shown exceptional capabilities in producing high-quality, dense visual predictions from open-vocabulary text. This indicates a strong correlation between visual and textual domains in open concepts and that diffusion-based text-to-image models can capture rich and diverse information for computer vision tasks. However, we found that those advantages do not hold for learning of features of camouflaged individuals because of the significant blending between their visual boundaries and their surroundings. In this paper, while leveraging the benefits of diffusion-based techniques and text-image models in open-vocabulary settings, we aim to address a challenging problem in computer vision: open-vocabulary camouflaged instance segmentation (OVCIS). Specifically, we propose a method built upon state-of-the-art diffusion empowered by open-vocabulary to learn multi-scale textual-visual features for camouflaged object representation learning. Such cross-domain representations are desirable in segmenting camouflaged objects where visual cues subtly distinguish the objects from the background, and in segmenting novel object classes which are not seen in training. To enable such powerful representations, we devise complementary modules to effectively fuse cross-domain features, and to engage relevant features towards respective foreground objects. We validate and compare our method with existing ones on several benchmark datasets of camouflaged and generic open-vocabulary instance segmentation. The experimental results confirm the advances of our method over existing ones. We believe that our proposed method would open a new avenue for handling camouflages such as computer vision-based surveillance systems, wildlife monitoring, and military reconnaissance. Tuan-Anh Vu, Duc Thanh Nguyen, Qing Guo 0005, Nhat Chung, Binh-Son Hua, Ivor W. Tsang, Sai-Kit Yeung |
Int. J. Comput. Vis. | 3 |
| 2026 | CamoVid60K: A Large-Scale Video Dataset for Moving Camouflaged Animals UnderstandingabstractAbstract We have been witnessing remarkable success led by the power of neural networks driven by a significant scale of training data in handling various computer vision tasks. However, less attention has been paid to monitoring the camouflaged animals, the masters of hiding themselves in the background. Robust and precise segmentation of camouflaged animals is challenging even for domain experts due to their similarity to the environment. Although several efforts have been made in camouflaged animal image segmentation, to the best of our knowledge, limited work exists on camouflaged animal video understanding (CAVU). Biologists often prefer videos for monitoring and understanding animal behaviors, as videos provide redundant information and temporal consistency. However, the scarcity of labeled video data significantly hinders progress in this area. To address these challenges, we present CamoVid60K , a diverse, large-scale, and accurately annotated video dataset of camouflaged animals. This dataset comprises 218 videos with 62,774 finely annotated frames, covering 70 animal categories, which surpasses all previous datasets in terms of the number of videos/frames and species included. CamoVid60K also offers more diverse downstream tasks in computer vision, such as camouflaged animal classification, detection, and task-specific segmentation (semantic, referring, motion), etc. We have benchmarked several state-of-the-art algorithms on the proposed CamoVid60K dataset, and the experimental results provide valuable insights for future research directions. Our dataset serves as a novel and challenging benchmark to stimulate the development of more powerful camouflaged animal video segmentation algorithms, with substantial room for further improvement. Tuan-Anh Vu, Ziqiang Zheng, Chengyang Song, Qing Guo 0005, Ivor W. Tsang, Sai-Kit Yeung |
Int. J. Comput. Vis. | 4 |
| 2026 | Knowledge-Guided Adversarial Training for Infrared Object Detection via Thermal Radiation Modeling
Shukun Xiong, Maoxun Yuan, Ranjie Duan, Qing Guo 0005, Haibin Duan, Xingxing Wei 0001 |
Int. J. Comput. Vis. | 6 |
| 2026 | Ensuring consistency with benign predictions: Differential privacy-guided certified defense against poisoning-based backdoor attacks
Yukun Yan, Jie Zhang 0073, Peng Tang 0002, Rui Chen 0012, Qilong Han, Haibo Hu 0001, Qing Guo 0005 |
Inf. Sci. | 7 |
| 2026 | FOCUS: Frequency-Optimized Conditioning of diffUSion models for mitigating catastrophic forgetting during test-time adaptation
Gabriel Tjio, Jie Zhang 0002, Xulei Yang, Nhat Chung, Xiaofeng Cao 0002, Ivor W. Tsang, Chee Keong Kwoh 0001, Qing Guo 0005 |
Mach. Vis. Appl. | 9 |
| 2026 | Exploring Security Vulnerabilities in Multilingual Speech Translation Systems via Deceptive InputsabstractAs speech translation (ST) systems become increasingly prevalent, understanding their vulnerabilities is crucial for ensuring robust and reliable communication. However, limited work has explored this issue in depth. This paper explores methods of compromising these systems through imperceptible audio manipulations. Specifically, we present two approaches: (1) adapting perturbation-based techniques used for automatic speech recognition (ASR) attacks to the ST context, making our work the first to apply this approach to ST, and (2) proposing a novel music generation-based method to guide targeted translation, while also conducting more practical over-the-air attacks in the physical world. Our experiments reveal that carefully crafted audio perturbations can mislead translation models to produce targeted, harmful outputs, while adversarial music achieve this goal more covertly, exploiting the natural imperceptibility of music. These attacks have proven effective across multiple languages and translation models, highlighting a systemic vulnerability in current ST architectures. Beyond immediate security concerns, our findings highlight broader challenges in the robustness and interpretability of neural speech systems. Chang Liu 0089, Haolin Wu 0001, Cong Wu 0003, Weiming Zhang 0001, Nenghai Yu, Tianwei Zhang 0004, Qing Guo 0005, Jie Zhang 0073 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | FOOLSDEDIT: Deceptively Steering Your Edits Towards Targeted Attribute-Aware DistributionabstractGuided image synthesis methods, like SDEdit based on the diffusion model, excel at creating realistic images from user inputs such as stroke paintings. However, existing efforts mainly focus on image quality, often overlooking a key point: the diffusion model represents a data distribution, not individual images. This introduces a low but critical chance of generating images that contradict user intentions, raising ethical concerns. For example, a user inputting a stroke painting with female characteristics might, with some probability, get male faces from SDEdit. To expose this potential vulnerability, we propose the Targeted Attribute Generative Attack (TAGA), whose objective is to force SDEdit to generate data distributions aligned with a specified attribute (i.e.,targeted attribute like male), without changing the attribute of the input image. Empirical studies reveal that traditional adversarial noise struggles to achieve TAGA, while natural perturbations such as exposure and motion blur can easily influence attributes of the generated images. Inspired by the observation, we design attack methodFOOLSDEDITto achieve effective TAGA against SDEdit. It aims to search for an optimized strategy to execute attacks within a weighted graph-based attack architecture, which is formulated to model diverse strategies derived from both exposure and motion blur perturbations. Comprehensive experiments on two commonly used datasets and three social attributes present thatFOOLSDEDITforces SDEdit to generate targeted attribute-aware distributions, achieving significantly more effective TAGA than the baselines. We also validated empirically thatFOOLSDEDITcould induce bias in downstream tasks of SDEdit which rely on the generated data under attack. Our work reveals critical vulnerabilities in diffusion-based image generation models and paves the way for future research on model auditing and bias mitigation. Qi Zhou 0012, Dongxia Wang 0002, Tianlin Li, Yang Liu 0003, Kui Ren 0001, Wenhai Wang, Qing Guo 0005 |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2026 | Toward Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-Critical ScenariosabstractAutonomous driving has made significant progress in both academia and industry, including performance improvements in perception tasks and the development of end-to-end autonomous driving systems. However, the safety and robustness assessment of autonomous driving has not received sufficient attention. Current evaluations of autonomous driving are typically conducted in natural driving scenarios. However, accidents often occur in edge cases, also known as safety-critical scenarios. These safety-critical scenarios are difficult to collect, and there is currently no clear definition of what constitutes a safety-critical scenario. In this work, we explore the safety and robustness of autonomous driving in safety-critical scenarios. First, we provide a definition of safety-critical scenarios, including static traffic scenarios such as adversarial attack scenarios and natural distribution shifts, as well as dynamic traffic scenarios such as accident scenarios. Then, we develop an autonomous driving test framework to comprehensively evaluate autonomous driving systems, encompassing not only the assessment of perception modules but also system-level evaluations. Our work systematically constructs a safety verification process for autonomous driving, providing technical support for the industry to establish standardized test framework. Jingzheng Li, Xianglong Liu 0001, Shikui Wei, Yufei Ge, Bing Li 0001, Qing Guo 0005, Xianqi Yang, Yanjun Pu, Qianren Mao, Jiakai Wang |
IEEE Trans. Image Process. | 7 |
| 2026 | Time-Variant Image Inpainting via Interactive Distribution Transition EstimationabstractIn this work, we focus on a novel and practical task, i.e., Time-vAriant iMage inPainting (TAMP). The aim of TAMP is to restore a damaged target image by leveraging the complementary information from a reference image, where both images capture the same scene but with a significant time gap in between, i.e., time-variant images. Different from conventional reference-guided image inpainting, the reference image under TAMP setup presents significant content distinction to the target image and potentially also suffers from damages. Such an application frequently happens in our daily life to restore a damaged image by referring to another reference image, where there is no guarantee of the reference image's source and quality. In particular, our study finds that even SOTA reference-guided image inpainting methods fail to achieve plausible results due to the chaotic image complementation. To address such an ill-posed problem, we propose a novel Interactive Distribution Transition Estimation (InDiTE) module which interactively complements the time-variant images with appropriate semantics thus facilitate the restoration of damaged regions. To further boost the performance, we propose our TAMP solution, namely Interactive Distribution Transition Estimation-driven Diffusion (InDiTE-Diff), which integrates InDiTE with SOTA diffusion model and conducts latent cross-reference during sampling. Moreover, considering the lack of benchmarks for TAMP task, we newly assembled a dataset, i.e., TAMP-Street, based on existing image and mask datasets. We conduct experiments on the TAMP-Street datasets under two different time-variant image inpainting settings, which show our method consistently outperform SOTA reference-guided image inpainting methods for solving TAMP. Yun Xing 0001, Qing Guo 0005, Yihao Huang 0001, Xiaofeng Cao 0002, Luqi Gong, Di Lin 0002, Ivor W. Tsang, Lei Ma 0003 |
IEEE Trans. Image Process. | 2 |
| 2025 | Concept Matching with Agent for Out-of-Distribution DetectionabstractThe remarkable achievements of Large Language Models (LLMs) have captivated the attention of both academia and industry, transcending their initial role in dialogue generation. To expand the usage scenarios of LLM, some works enhance the effectiveness and capabilities of the model by introducing more external information, which is called the agent paradigm. Based on this idea, we propose a new method that integrates the agent paradigm into out-of-distribution (OOD) detection task, aiming to improve its robustness and adaptability. Our proposed method, Concept Matching with Agent (CMA), employs neutral prompts as agents to augment the CLIP-based OOD detection process. These agents function as dynamic observers and communication hubs, interacting with both In-distribution (ID) labels and data inputs to form vector triangle relationships. This triangular framework offers a more nuanced approach than the traditional binary relationship, allowing for better separation and identification of ID and OOD inputs. Our extensive experimental results showcase the superior performance of CMA over both zero-shot and training-required methods in a diverse array of real-world scenarios. Yuxiao Lee, Xiaofeng Cao 0002, Jingcai Guo, Wei Ye 0001, Qing Guo 0005, Yi Chang 0001 |
AAAI | 5 |
| 2025 | Efficient Universal Goal Hijacking with Semantics-guided Prompt OrganizationabstractYihao Huang, Chong Wang, Xiaojun Jia, Qing Guo, Felix Juefei-Xu, Jian Zhang, Yang Liu, Geguang Pu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yihao Huang 0001, Chong Wang 0013, Xiaojun Jia, Qing Guo 0005, Felix Juefei-Xu, Jian Zhang 0087, Yang Liu 0003, Geguang Pu |
ACL (1) | 4 |
| 2025 | SafeGuider: Robust and Practical Content Safety Control for Text-to-Image ModelsabstractText-to-image models have shown remarkable capabilities in generating high-quality images from natural language descriptions. However, these models are highly vulnerable to adversarial prompts, which can bypass safety measures and produce harmful content. Despite various defensive strategies, achieving robustness against attacks while maintaining practical utility in real-world applications remains a significant challenge. To address this issue, we first conduct an empirical study of the text encoder in the Stable Diffusion (SD) model, which is a widely used and representative text-to-image model. Our findings reveal that the [EOS] token acts as a semantic aggregator, exhibiting distinct distributional patterns between benign and adversarial prompts in its embedding space. Building on this insight, we introduce SafeGuider, a two-step framework designed for robust safety control without compromising generation quality. SafeGuider combines an embedding-level recognition model with a safety-aware feature erasure beam search algorithm. This integration enables the framework to maintain high-quality image generation for benign prompts while ensuring robust defense against both in-domain and out-of-domain attacks. SafeGuider demonstrates exceptional effectiveness in minimizing attack success rates, achieving a maximum rate of only 5.48% across various attack scenarios. Moreover, instead of refusing to generate or producing black images for unsafe prompts, SafeGuider generates safe and meaningful images, enhancing its practical utility. In addition, SafeGuider is not limited to the SD model and can be effectively applied to other text-to-image models, such as the Flux model, demonstrating its versatility and adaptability across different architectures. We hope that SafeGuider can shed some light on the practical deployment of secure text-to-image systems. Peigui Qi, Kunsheng Tang, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu, Tianwei Zhang 0004, Qing Guo 0005, Jie Zhang 0073 |
CCS | 7 |
| 2025 | SceneTAP: Scene-Coherent Typographic Adversarial Planner against Vision-Language Models in Real-World EnvironmentsabstractLarge vision-language models (LVLMs) have shown remarkable capabilities in interpreting visual content. While existing works demonstrate these models’ vulnerability to deliberately placed adversarial texts, such texts are often easily identifiable as anomalous. In this paper, we present the first approach to generate scene-coherent typographic adversarial attacks that mislead advanced LVLMs while maintaining visual naturalness through the capability of the LLM-based agent. Our approach addresses three critical questions: what adversarial text to generate, where to place it within the scene, and how to integrate it seamlessly. We propose a training-free, multi-modal LLM-driven scene-coherent typographic adversarial planning (SceneTAP) that employs a three-stage process: scene understanding, adversarial planning, and seamless integration. The SceneTAP utilizes chain-of-thought reasoning to comprehend the scene, formulate effective adversarial text, strategically plan its placement, and provide detailed instructions for natural integration within the image. This is followed by a scene-coherent TextDiffuser that executes the attack using a local diffusion mechanism. We extend our method to real-world scenarios by printing and placing generated patches in physical environments, demonstrating its practical implications. Extensive experiments show that our scene-coherent adversarial text successfully misleads state-of-the-art LVLMs, including ChatGPT-4o, even after capturing new images of physical setups. Our evaluations demonstrate a significant increase in attack success rates while maintaining visual naturalness and contextual appropriateness. This work highlights vulnerabilities in current vision-language models to sophisticated, scene-coherent adversarial attacks and provides insights into potential defense mechanisms. We release our code at https://github.com/tsingqguo/scenetap. Jie Zhang 0002, Di Lin 0002, Tianwei Zhang 0004, Ivor W. Tsang, Yang Liu 0003, Qing Guo 0005 |
CVPR | 8 |
| 2025 | Engage for All: Making Ordinary Image Descriptions Appealing Again!
Yuyan Chen, Jinghan Cao, Yu Guan 0001, Ming-Hsuan Yang 0001, Qing Guo 0005 |
ICCV | 7 |
| 2025 | Social Debiasing for Fair Multi-Modal LLMsabstractMulti-modal Large Language Models (MLLMs) have dramatically advanced the research field and delivered powerful vision-language understanding capabilities. However, these models often inherit deep-rooted social biases from their training data, leading to uncomfortable responses with respect to attributes such as race and gender. This paper addresses the issue of social biases in MLLMs by i) introducing a comprehensive counterfactual dataset with multiple social concepts (CMSC), which complements existing datasets by providing 18 diverse and balanced social concepts; and ii) proposing a counter-stereotype debiasing (CSD) strategy that mitigates social biases in MLLMs by leveraging the opposites of prevalent stereotypes. CSD incorporates both a novel bias-aware data sampling method and a loss rescaling method, enabling the model to effectively reduce biases. We conduct extensive experiments with four prevalent MLLM architectures. The results demonstrate the advantage of the CMSC dataset and the edge of CSD strategy in reducing social biases compared to existing competing methods, without compromising the overall performance on general multi-modal reasoning benchmarks. Harry Cheng 0002, Qing Guo 0005, Ming-Hsuan Yang 0001, Tian Gan 0002, Weili Guan, Liqiang Nie |
ICCV | 3 |
| 2025 | NoiseController: Towards Consistent Multi-View Video Generation via Noise Decomposition and Collaboration
Haotian Dong, Xin Wang 0118, Di Lin 0002, Yipeng Wu, Kairui Yang, Ping Li 0016, Qing Guo 0005 |
ICCV | 9 |
| 2025 | VideoShield: Regulating Diffusion-based Video Generation Models via WatermarkingabstractArtificial Intelligence Generated Content (AIGC) has advanced significantly, particularly with the development of video generation models such as text-to-video (T2V) models and image-to-video (I2V) models. However, like other AIGC types, video generation requires robust content control. A common approach is to embed watermarks, but most research has focused on images, with limited attention given to videos. Traditional methods, which embed watermarks frame-by-frame in a post-processing manner, often degrade video quality. In this paper, we propose VideoShield, a novel watermarking framework specifically designed for popular diffusion-based video generation models. Unlike post-processing methods, VideoShield embeds watermarks directly during video generation, eliminating the need for additional training. To ensure video integrity, we introduce a tamper localization feature that can detect changes both temporally (across frames) and spatially (within individual frames). Our method maps watermark bits to template bits, which are then used to generate watermarked noise during the denoising process. Using DDIM Inversion, we can reverse the video to its original watermarked noise, enabling straightforward watermark extraction. Additionally, template bits allow precise detection for potential spatial and temporal modification. Extensive experiments across various video models (both T2V and I2V models) demonstrate that our method effectively extracts watermarks and detects tamper without compromising video quality. Furthermore, we show that this approach is applicable to image generation models, enabling tamper detection in generated images as well. Codes and models are available at https://github.com/hurunyi/VideoShield. Runyi Hu, Jie Zhang 0073, Yiming Li 0004, Jiwei Li 0001, Qing Guo 0005, Han Qiu 0001, Tianwei Zhang 0004 |
ICLR | 5 |
| 2025 | Trajectory-LLM: A Language-based Data Generator for Trajectory Prediction in Autonomous DrivingabstractVehicle trajectory prediction is a crucial aspect of autonomous driving, which requires extensive trajectory data to train prediction models to understand the complex, varied, and unpredictable patterns of vehicular interactions. However, acquiring real-world data is expensive, so we advocate using Large Language Models (LLMs) to generate abundant and realistic trajectories of interacting vehicles efficiently. These models rely on textual descriptions of vehicle-to-vehicle interactions on a map to produce the trajectories. We introduce Trajectory-LLM (Traj-LLM), a new approach that takes brief descriptions of vehicular interactions as input and generates corresponding trajectories. Unlike language-based approaches that translate text directly to trajectories, Traj-LLM uses reasonable driving behaviors to align the vehicle trajectories with the text. This results in an "interaction-behavior-trajectory" translation process. We have also created a new dataset, Language-to-Trajectory (L2T), which includes 240K textual descriptions of vehicle interactions and behaviors, each paired with corresponding map topologies and vehicle trajectory segments. By leveraging the L2T dataset, Traj-LLM can adapt interactive trajectories to diverse map topologies. Furthermore, Traj-LLM generates additional data that enhances downstream prediction models, leading to consistent performance improvements across public benchmarks. The source code is released at https://github.com/TJU-IDVLab/Traj-LLM. Kairui Yang, Gengjie Lin, Haotian Dong, Yipeng Wu, Die Zuo, Jibin Peng, Ziyuan Zhong, Xin Wang 0118, Qing Guo 0005, Xiaosong Jia, Junchi Yan, Di Lin 0002 |
ICLR | 11 |
| 2025 | TGSR: Template-Guided Semantic Resampling against Adversarial Tracking AttacksabstractDeep object tracking has made significant strides, demonstrating impressive accuracy across diverse visual scenarios. However, recent studies revealed that visual object trackers remain vulnerable to adversarial attacks specifically designed to disrupt tracking tasks. While image resampling—reconstructing images using resampled coordinates and bilinear interpolation—has shown promise in enhancing tracker robustness by disrupting adversarial patterns, this naive approach overlooks crucial semantic information and scale variations between frames relative to the initial object template. These variations typically arise from changes in the target object or background. To address this limitation, we propose a template-guided semantic resampling (TGSR) method to counter adversarial tracking attacks. Our approach comprises two key components: template-aware joint semantic and appearance implicit representation (T-SAIR) and template-aware predictive resampling (T-PRES). T-SAIR estimates semantic embeddings and corresponding pixel colors at arbitrary coordinates based on incoming frames and the object template, while T-PRES predicts pixel-wise coordinate shifts in response to scale changes relative to the template. The integration of these modules enables our method to effectively reconstruct frames, neutralizing adversarial perturbations while preserving semantic information relative to the object template. Extensive experimental evaluation against three attacks across typical tracking methods demonstrates the effectiveness of our approach. Xuhong Ren, Jianlang Chen, Wanli Xue, Lei Ma 0003, Qing Guo 0005, Jianjun Zhao 0001, Shengyong Chen |
ICME | 5 |
| 2025 | PhysLight: Accurate rPPG Heart Rate Measurement with Adaptive Video RelightingabstractFacial video-based remote physiological measurement (rPPG) can non-invasively estimate vital signs, such as heart rate (HR), which often faces challenges under varying lighting conditions. We propose the PhysLight framework to enhance the accuracy of rPPG heart rate measurement through adaptive video relighting. Our approach subtly modifies illumination in video frames to improve detection accuracy while maintaining visual quality. The framework includes a GenLightNet to extract ideal lighting priors and a WipeLightNet module to refine poorly lit videos. Extensive evaluations on benchmark datasets show that our method significantly improves HR estimation reliability, outperforming existing baselines and enhancing non-contact physiological monitoring in diverse environments. Menglin Zhang, Xiaoxin Guo, Bohao Qu, Xiaofeng Cao 0002, Shuifa Sun, Qing Guo 0005 |
ICME | 6 |
| 2025 | Cowpox: Towards the Immunity of VLM-based Multi-Agent SystemsabstractVision Language Model (VLM) Agents are stateful, autonomous entities capable of perceiving and interacting with their environments through vision and language. Multi-agent systems comprise specialized agents who collaborate to solve a (complex) task. A core security property is robustness, stating that the system maintains its integrity during adversarial attacks. Multi-agent systems lack robustness, as a successful exploit against one agent can spread and infect other agents to undermine the entire system’s integrity. We propose a defense Cowpox to provably enhance the robustness of a multi-agent system by a distributed mechanism that improves the recovery rate of agents by limiting the expected number of infections to other agents. The core idea is to generate and distribute a special cure sample that immunizes an agent against the attack before exposure. We demonstrate the effectiveness of Cowpox empirically and provide theoretical robustness guarantees. Yutong Wu 0009, Jie Zhang 0073, Yiming Li 0004, Chao Zhang 0008, Qing Guo 0005, Han Qiu 0001, Nils Lukas, Tianwei Zhang 0004 |
ICML | 5 |
| 2025 | Defending LVLMs Against Vision Attacks Through Partial-Perception SupervisionabstractRecent studies have raised significant concerns regarding the vulnerability of Large Vision Language Models (LVLMs) to maliciously injected or perturbed input images, which can mislead their responses. Existing defense methods show that such vision attacks are sensitive to image modifications especially cropping, using majority voting across responses of modified images as corrected responses. However, these modifications often result in partial images and distort the semantics, which reduces response quality on clean images after voting. Instead of directly using responses from partial images for voting, we investigate using them to supervise (guide) the LVLM’s responses to the original images at inference time. We propose a black-box, training-free method called DPS (Defense through Partial-Perception Supervision). In this approach, the model is prompted using the responses generated by a model that perceives only a partial image. With DPS, the model can adjust its response based on partial image understanding when under attack, while confidently maintaining its original response for clean input. Empirical experiments show our method outperforms the baseline, cutting the average attack success rate by 76.3% across six datasets on three popular models. Qi Zhou 0012, Dongxia Wang 0002, Tianlin Li, Yun Lin 0001, Yang Liu 0003, Jin Song Dong 0001, Qing Guo 0005 |
ICML | 7 |
| 2025 | Diversifying Policy Behaviors with Extrinsic Behavioral CuriosityabstractImitation learning (IL) has shown promise in various applications (e.g. robot locomotion) but is often limited to learning a single expert policy, constraining behavior diversity and robustness in unpredictable real-world scenarios. To address this, we introduce Quality Diversity Inverse Reinforcement Learning (QD-IRL), a novel framework that integrates quality-diversity optimization with IRL methods, enabling agents to learn diverse behaviors from limited demonstrations. This work introduces Extrinsic Behavioral Curiosity (EBC), which allows agents to receive additional curiosity rewards from an external critic based on how novel the behaviors are with respect to a large behavioral archive. To validate the effectiveness of EBC in exploring diverse locomotion behaviors, we evaluate our method on multiple robot locomotion tasks. EBC improves the performance of QD-IRL instances with GAIL, VAIL, and DiffAIL across all included environments by up to 185%, 42%, and 150%, even surpassing expert performance by 20% in Humanoid. Furthermore, we demonstrate that EBC is applicable to Gradient-Arborescence-based Quality Diversity Reinforcement Learning (QD-RL) algorithms, where it substantially improves performance and provides a generic technique for learning behavioral diverse policies. The source code of this work is provided at https://github.com/vanzll/EBC. Zhenglin Wan, Xingrui Yu, David Mark Bossens, Yueming Lyu, Qing Guo 0005, Flint Xiaofeng Fan, Yew-Soon Ong, Ivor W. Tsang |
ICML | 5 |
| 2025 | A Dual Large Language Models Architecture with Herald Guided Prompts for Parallel Fine Grained Traffic Signal ControlabstractLeveraging large language models (LLMs) in traffic signal control (TSC) improves optimization efficiency and interpretability compared to traditional reinforcement learning (RL) methods. However, existing LLM-based approaches are limited by fixed time signal durations and are prone to hallucination errors, while RL methods lack robustness in signal timing decisions and suffer from poor generalization. To address these challenges, this paper proposes HeraldLight, a dual LLMs architecture enhanced by Herald guided prompts. The Herald Module extracts contextual information and forecasts queue lengths for each traffic phase based on real-time conditions. The first LLM, LLM-Agent, uses these forecasts to make fine grained traffic signal control, while the second LLM, LLM-Critic, refines LLM-Agent's outputs, correcting errors and hallucinations. These refined outputs are used for score-based fine-tuning to improve accuracy and robustness. Simulation experiments using CityFlow on real world datasets covering 224 intersections in Jinan (12), Hangzhou (16), and New York (196) demonstrate that HeraldLight outperforms state of the art baselines, achieving a 20.03% reduction in average travel time across all scenarios and a 10.74% reduction in average queue length on the Jinan and Hangzhou scenarios. The source code is available on GitHub: https://github.com/BUPT-ANTlab/HeraldLight. Qing Guo 0005, Xiaocong Li |
ICPADS | 1 |
| 2025 | Retrieval Augmented Generation-Enhanced Distributed LLM Agents for Generalizable Traffic Signal Control with Emergency VehiclesabstractWith increasing urban traffic complexity, Traffic Signal Control (TSC) is essential for optimizing traffic flow and improving road safety. Large Language Models (LLMs) emerge as promising approaches for TSC. However, they are prone to hallucinations in emergencies, leading to unreliable decisions that may cause substantial delays for emergency vehicles. Moreover, diverse intersection types present substantial challenges for traffic state encoding and cross-intersection training, limiting generalization across heterogeneous intersections. Therefore, this paper proposes Retrieval Augmented Generation (RAG)-enhanced distributed LLM agents with Emergency response for Generalizable TSC (REG-TSC). Firstly, this paper presents an emergencyaware reasoning framework, which dynamically adjusts reasoning depth based on the emergency scenario and is equipped with a novel Reviewer-based Emergency RAG (RERAG) to distill specific knowledge and guidance from historical cases, enhancing the reliability and rationality of agents' emergency decisions. Secondly, this paper designs a type-agnostic traffic representation and proposes a Reward-guided Reinforced Refinement ($\mathbf{R}^{3}$) for heterogeneous intersections.$\mathbf{R}^{3}$adaptively samples training experience from diverse intersections with environment feedbackbased priority and fine-tunes LLM agents with a designed reward-weighted likelihood loss, guiding REG-TSC toward highreward policies across heterogeneous intersections. On three realworld road networks with 17 to 177 heterogeneous intersections, extensive experiments show that REG-TSC reduces travel time by 42.00%, queue length by 62.31%, and emergency vehicle waiting time by 83.16%, outperforming other state-of-the-art methods. Qing Guo 0005, Shengzhe Xu |
ICPADS | 2 |
| 2025 | Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
Xingrui Yu, Zhenglin Wan, David Mark Bossens, Yueming Lyu, Qing Guo 0005, Ivor W. Tsang |
AAMAS | 5 |
| 2025 | Mask Image WatermarkingabstractWe present MaskWM, a simple, efficient, and flexible framework for image watermarking. MaskWM has two variants: (1) MaskWM-D, which supports global watermark embedding, watermark localization, and local watermark extraction for applications such as tamper detection; (2) MaskWM-ED, which focuses on local watermark embedding and extraction, offering enhanced robustness in small regions to support fine-grined image protection. MaskWM-D builds on the classical encoder-distortion layer-decoder training paradigm. In MaskWM-D, we introduce a simple masking mechanism during the decoding stage that enables both global and local watermark extraction. During training, the decoder is guided by various types of masks applied to watermarked images before extraction, helping it learn to localize watermarks and extract them from the corresponding local areas. MaskWM-ED extends this design by incorporating the mask into the encoding stage as well, guiding the encoder to embed the watermark in designated local regions, which improves robustness under regional attacks. Extensive experiments show that MaskWM achieves state-of-the-art performance in global and local watermark extraction, watermark localization, and multi-watermark embedding. It outperforms all existing baselines, including the recent leading model WAM for local watermarking, while preserving high visual quality of the watermarked images. In addition, MaskWM is highly efficient and adaptable. It requires only 20 hours of training on a single A6000 GPU, achieving 15× computational efficiency compared to WAM. By simply adjusting the distortion layer, MaskWM can be quickly fine-tuned to meet varying robustness requirements. Runyi Hu, Jie Zhang 0073, Shiqian Zhao, Nils Lukas, Jiwei Li 0001, Qing Guo 0005, Han Qiu 0001, Tianwei Zhang 0004 |
NeurIPS | 6 |
| 2025 | AngleRoCL: Angle-Robust Concept Learning for Physically View-Invariant Adversarial PatchesabstractCutting-edge works have demonstrated that text-to-image (T2I) diffusion models can generate adversarial patches that mislead state-of-the-art object detectors in the physical world, revealing detectors' vulnerabilities and risks. However, these methods neglect the T2I patches' attack effectiveness when observed from different views in the physical world (i.e., angle robustness of the T2I adversarial patches). In this paper, we study the angle robustness of T2I adversarial patches comprehensively, revealing their angle-robust issues, demonstrating that texts affect the angle robustness of generated patches significantly, and task-specific linguistic instructions fail to enhance the angle robustness. Motivated by the studies, we introduce Angle-Robust Concept Learning (AngleRoCL), a simple and flexible approach that learns a generalizable concept (i.e., text embeddings in implementation) representing the capability of generating angle-robust patches. The learned concept can be incorporated into textual prompts and guides T2I models to generate patches with their attack effectiveness inherently resistant to viewpoint variations. Through extensive simulation and physical-world experiments on five SOTA detectors across multiple views, we demonstrate that AngleRoCL significantly enhances the angle robustness of T2I adversarial patches compared to baseline methods. Our patches maintain high attack success rates even under challenging viewing conditions, with over 50% average relative improvement in attack effectiveness across multiple angles. This research advances the understanding of physically angle-robust patches and provides insights into the relationship between textual concepts and physical properties in T2I-generated contents. We released our code at https://github.com/tsingqguo/anglerocl. Wenjun Ji, Luyang Ying, Deng-Ping Fan, Yuyi Wang 0001, Ming-Ming Cheng, Ivor W. Tsang, Qing Guo 0005 |
NeurIPS | 8 |
| 2025 | Open-Vocabulary Part Segmentation via Progressive and Boundary-Aware StrategyabstractOpen-vocabulary part segmentation (OVPS) struggles with structurally connected boundaries due to the inherent conflict between continuous image features and discrete classification mechanism. To address this, we propose PBAPS, a novel training-free framework specifically designed for OVPS. PBAPS leverages structural knowledge of object-part relationships to guide a progressive segmentation from objects to fine-grained parts. To further improve accuracy at challenging boundaries, we introduce a Boundary-Aware Refinement (BAR) module that identifies ambiguous boundary regions by quantifying classification uncertainty, enhances the discriminative features of these ambiguous regions using high-confidence context, and adaptively refines part prototypes to better align with the specific image. Experiments on Pascal-Part-116, ADE20K-Part-234, PartImageNet demonstrate that PBAPS significantly outperforms state-of-the-art methods, achieving 46.35\% mIoU and 34.46\% bIoU on Pascal-Part-116. Our code is available at https://github.com/TJU-IDVLab/PBAPS. Xinlong Li, Di Lin 0002, Shaoyiyi Gao, Qing Guo 0005 |
NeurIPS | 6 |
| 2025 | DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible PatchesabstractStereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous works have shown that repeating optimized textures can effectively mislead stereo depth estimation in digital settings. However, our research reveals that these naively repeated textures perform poorly in physical implementations, $\textit{i.e.}$, when deployed as patches, limiting their practical utility for stress-testing stereo depth estimation systems. In this work, for the first time, we discover that introducing regular intervals among the repeated textures, creating a grid structure, significantly enhances the patch attack performance. Through extensive experimentation, we analyze how variations of this novel structure influence the adversarial effectiveness. Based on these insights, we develop a novel stereo depth attack that jointly optimizes both the interval structure and texture elements. Our generated adversarial patches can be inserted into any scenes and successfully attack advanced stereo depth estimation methods of different paradigms, $\textit{i.e.}$, RAFT-Stereo and STTR. Most critically, our patch can also attack commercial RGB-D cameras (Intel RealSense) in real-world conditions, demonstrating their practical relevance for security assessment of stereo systems. The code is officially released at: https://github.com/WiWiN42/DepthVanish Yun Xing 0001, Nhat Chung, Jie Zhang 0050, Ivor W. Tsang, Ming-Ming Cheng, Yang Liu 0003, Lei Ma 0003, Qing Guo 0005 |
NeurIPS | 9 |
| 2025 | CamLopa: A Hidden Wireless Camera Localization Framework via Signal Propagation Path AnalysisabstractHidden wireless cameras pose significant privacy threats, necessitating effective detection and localization methods. However, existing localization solutions often require impractical activity spaces, expensive specialized devices, or pre-collected training data, limiting their practical deployment. To address these limitations, we introduce CamLopa, a training-free wireless camera localization framework that operates with minimal activity space constraints using low-cost, commercial-off-the-shelf (COTS) devices. CamLopa can achieve detection and localization in just 45 seconds of user activities with a Raspberry Pi board. During this short period, it analyzes the causal relationship between wireless traffic and user movement to detect the presence of a hidden camera. Upon detection, CamLopa utilizes a novel azimuth localization model based on wireless signal propagation path analysis for localization. This model leverages the time ratio of user paths crossing the First Fresnel Zone (FFZ) to determine the camera's azimuth angle. Subsequently, CamLopa refines the localization by identifying the camera's quadrant. We evaluate CamLopa across various devices and environments, demonstrating its effectiveness with a 95.37% detection accuracy for snooping cameras and an average localization error of 17.23°, under the significantly reduced activity space requirements and without the need for training. Our code and demo are available at https://github.com/CamLoPA/CamLoPA-Code. Xiang Zhang 0011, Jie Zhang 0073, Zehua Ma, Jinyang Huang, Meng Li 0006, Huan Yan 0004, Peng Zhao 0024, Zijian Zhang 0001, Bin Liu 0016, Qing Guo 0005, Tianwei Zhang 0004, Nenghai Yu |
SP | 10 |
| 2025 | DiffLoc: WiFi Hidden Camera Localization Based on Electromagnetic Diffraction
Xiang Zhang 0011, Jie Zhang 0073, Huan Yan 0004, Jinyang Huang, Zehua Ma, Bin Liu 0016, Meng Li 0006, Kejiang Chen, Qing Guo 0005, Tianwei Zhang 0004, Zhi Liu 0002 |
USENIX Security Symposium | 9 |
| 2025 | When Translators Refuse to Translate: A Novel Attack to Speech Translation Systems
Haolin Wu 0001, Chang Liu 0089, Jing Chen 0003, Ruiying Du, Kun He 0008, Yu Zhang 0036, Cong Wu 0003, Tianwei Zhang 0004, Qing Guo 0005, Jie Zhang 0073 |
USENIX Security Symposium | 9 |
| 2025 | Security-reliability analysis in uplink cognitive satellite-terrestrial networks with LEO relayingabstractThis paper investigates the security and reliability performance of hybrid cognitive satellite-terrestrial networks employing a Low Earth Orbit (LEO) satellite as a decode-and-forward (DF) relay. The terrestrial user (TU) operates within an underlay cognitive radio (CR) network, where the primary user (PU) shares its spectrum with the TU while imposing interference power constraints to protect its quality-of-service. To counteract eavesdropping from a terrestrial adversary, the TU incorporates artificial noise (AN) into its transmission, creating a tradeoff between security and reliability. The TU-to-LEO and TU-to-PU links are modeled using Shadowed Rician and Nakagami- m fading, respectively. Key performance metrics, including the outage probability (OP) and intercept probability (IP), are analyzed under varying system parameters such as power-splitting factor, channel conditions, and interference thresholds. Analytical results are validated through Monte Carlo simulations , and simplified approximations are presented for practical implementation. Results demonstrate the efficacy of the proposed approach in balancing security and reliability. Qing Guo 0005 |
Comput. Networks | 2 |
| 2025 | EfficientDeRain+: Learning Uncertainty-Aware Filtering via RainMix Augmentation for High-Efficiency Deraining
Qing Guo 0005, Hua Qi, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Di Lin 0002, Wei Feng 0005, Song Wang 0002 |
Int. J. Comput. Vis. | 1 |
| 2025 | Towards robust DeepFake distortion attack via adversarial autoaugment
Qi Guo 0008, Shanmin Pang, Qing Guo 0005 |
Neurocomputing | 4 |
| 2025 | Efficient Place Recognition With Complex Number Framework for Robust Environmental Description and Real-Time PerformanceabstractThe rapid development of autonomous driving and robotics has led to increasingly higher demands for high-precision localization, making it an indispensable key component of modern autonomous systems. High-definition maps have become widely used to enhance localization accuracy, particularly in complex and dynamic environments. However, persistent challenges such as sensor measurement errors, environmental drift, and inaccuracies remain, making place recognition an increasingly critical research focus. This paper introduces a novel place recognition method based on a complex number framework, wherein environmental features are stored in both the real and imaginary parts of a complex number. This innovative approach significantly improves the richness and accuracy of environmental descriptions, thereby enhancing the system’s ability to distinguish between highly similar environments. Furthermore, a descriptor similarity evaluation strategy based on the angle between complex number vectors is proposed, which enhances the distinction of descriptors, especially in scenarios where the environments share similar characteristics, thus reducing the likelihood of mismatches. To ensure real-time performance, a hierarchical retrieval filtering process is introduced that effectively combines both coarse and fine-grained search strategies, optimizing the KD-Tree search and filtering out irrelevant matches. Experimental results on public datasets and real-world vehicle data across diverse scenarios validate the proposed method’s effectiveness and robustness, showing notable gains in recognition accuracy and real-time performance in complex environments. Xiangmo Zhao, Wuqi Wang, Xia Wu 0004, Chunyun Zheng, Qing Guo 0005 |
IEEE Internet Things J. | 7 |
| 2025 | Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language AttackabstractVision-language pre-training (VLP) models excel at interpreting both images and text but remain vulnerable to multimodal adversarial examples (AEs). Advancing the generation of transferable AEs, which succeed across unseen models, is key to developing more robust and practical VLP models. Previous approaches augment image-text pairs to enhance diversity within the adversarial example generation process, aiming to improve transferability by expanding the contrast space of image-text features. However, these methods focus solely on diversity around the current AEs, yielding limited gains in transferability. To address this issue, we propose to increase the diversity of AEs by leveraging the intersection regions along the adversarial trajectory during optimization. Specifically, we propose sampling from adversarial evolution triangles composed of clean, historical, and current adversarial examples to enhance adversarial diversity. We provide a theoretical analysis to demonstrate the effectiveness of the proposed adversarial evolution triangle. Moreover, we find that redundant inactive dimensions can dominate similarity calculations, distorting feature matching and making AEs model-dependent with reduced transferability. Hence, we propose to generate AEs in the semantic image-text feature contrast space, which can project the original feature space into a semantic corpus subspace. The proposed semantic-aligned subspace can reduce the image feature redundancy, thereby improving adversarial transferability. Extensive experiments across different datasets and models demonstrate that the proposed method can effectively improve adversarial transferability and outperform state-of-the-art adversarial attack methods. Xiaojun Jia, Sensen Gao, Qing Guo 0005, Simeng Qin, Ke Ma 0001, Yihao Huang 0001, Yang Liu 0003, Ivor W. Tsang, Xiaochun Cao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | The NeRF Signature: Codebook-Aided Watermarking for Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have been gaining attention as a significant form of 3D content representation. With the proliferation of NeRF-based creations, the need for copyright protection has emerged as a critical issue. Although some approaches have been proposed to embed digital watermarks into NeRF, they often neglect essential model-level considerations and incur substantial time overheads, resulting in reduced imperceptibility and robustness, along with user inconvenience. In this paper, we extend the previous criteria for image watermarking to the model level and propose NeRF Signature, a novel watermarking method for NeRF. We employ a Codebook-aided Signature Embedding (CSE) that does not alter the model structure, thereby maintaining imperceptibility and enhancing robustness at the model level. Furthermore, after optimization, any desired signatures can be embedded through the CSE, and no fine-tuning is required when NeRF owners want to use new binary signatures. Then, we introduce a joint pose-patch encryption watermarking strategy to hide signatures into patches rendered from a specific viewpoint for higher robustness. In addition, we explore a Complexity-Aware Key Selection (CAKS) scheme to embed signatures in high visual complexity patches to enhance imperceptibility. The experimental results demonstrate that our method outperforms other baseline methods in terms of imperceptibility and robustness. Ziyuan Luo, Anderson Rocha 0001, Boxin Shi, Qing Guo 0005, Haoliang Li, Renjie Wan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | FaceTracer: Unveiling Source Identities From Swapped Face Images and Videos for Fraud PreventionabstractFace-swapping techniques have advanced rapidly with the evolution of deep learning, leading to widespread use and growing concerns about potential misuse, especially in cases of fraud. While many efforts have focused on detecting swapped face images or videos, these methods are insufficient for tracing the malicious users behind fraudulent activities. Intrusive watermark-based approaches also fail to trace unmarked identities, limiting their practical utility. To address these challenges, we introduce FaceTracer, the first non-intrusive framework specifically designed to trace the identity of the source person from swapped face images or videos. Specifically, FaceTracer leverages a disentanglement module that effectively suppresses identity information related to the target person while isolating the identity features of the source person. This allows us to extract robust identity information that can directly link the swapped face back to the original individual, aiding in uncovering the actors behind fraudulent activities. Extensive experiments demonstrate FaceTracer's effectiveness across various face-swapping techniques, successfully identifying the source person in swapped content and enabling the tracing of malicious actors involved in fraudulent activities. Additionally, FaceTracer shows strong transferability to unseen face-swapping methods including commercial applications and robustness against transmission distortions and adaptive attacks. Zhongyi Zhang 0001, Jie Zhang 0073, Wenbo Zhou 0004, Xinghui Zhou, Qing Guo 0005, Weiming Zhang 0001, Tianwei Zhang 0004, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Efficient Generation of Targeted and Transferable Adversarial Examples for Vision-Language Models via Diffusion ModelsabstractAdversarial attacks, particularly targeted transfer-based attacks, can be used to assess the adversarial robustness of large visual-language models (VLMs), allowing for a more thorough examination of potential security flaws before deployment. However, previous transfer-based adversarial attacks incur high costs due to high iteration counts and complex method structure. Furthermore, due to the unnaturalness of adversarial semantics, the generated adversarial examples have low transferability. These issues limit the utility of existing methods for assessing robustness. To address these issues, we propose AdvDiffVLM, which uses diffusion models to generate natural, unrestricted and targeted adversarial examples via score matching. Specifically, AdvDiffVLM uses Adaptive Ensemble Gradient Estimation (AEGE) to modify the score during the diffusion model’s reverse generation process, ensuring that the produced adversarial examples have natural adversarial targeted semantics, which improves their transferability. Simultaneously, to improve the quality of adversarial examples, we use the GradCAM-guided Mask Generation (GCMG) to disperse adversarial semantics throughout the image rather than concentrating them in a single area. Finally, AdvDiffVLM embeds more target semantics into adversarial examples after multiple iterations. Experimental results show that our method generates adversarial examples 5x to 10x faster than state-of-the-art (SOTA) transfer-based adversarial attacks while maintaining higher quality adversarial examples. Furthermore, compared to previous transfer-based adversarial attacks, the adversarial examples generated by our method have better transferability. Notably, AdvDiffVLM can successfully attack a variety of commercial VLMs in a black-box environment, including GPT-4V. The code is available athttps://github.com/gq-max/AdvDiffVLM Qi Guo 0008, Shanmin Pang, Xiaojun Jia, Yang Liu 0003, Qing Guo 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Scale-Invariant Adversarial Attack Against Arbitrary-Scale Super-ResolutionabstractThe advent of local continuous image function (LIIF) has garnered significant attention for arbitrary-scale super-resolution (SR) techniques. However, while the vulnerabilities of fixed-scale SR have been assessed, the robustness of continuous representation-based arbitrary-scale SR against adversarial attacks remains an area warranting further exploration. The elaborately designed adversarial attacks for fixed-scale SR are scale-dependent, which will cause time-consuming and memory-consuming problems when applied to arbitrary-scale SR. To address this concern, we propose a simple yet effective “scale-invariant” SR adversarial attack method with good transferability, termed SIAGT. Specifically, we propose to construct resource-saving attacks by exploiting finite discrete points of continuous representation. In addition, we formulate a coordinate-dependent loss to enhance the cross-model transferability of the attack. The attack can significantly deteriorate the SR images while introducing imperceptible distortion to the targeted low-resolution (LR) images. Experiments carried out on three popular LIIF-based SR approaches and four classical SR datasets show remarkable attack performance and transferability of SIAGT. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Xiaojun Jia, Weikai Miao, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Compromising LLM Driven Embodied Agents With Contextual Backdoor Attacks
Aishan Liu, Yuguang Zhou, Xianglong Liu 0001, Tianyuan Zhang 0004, Siyuan Liang 0004, Jiakai Wang, Yanjun Pu, Tianlin Li, Wenbo Zhou 0004, Qing Guo 0005, Dacheng Tao |
IEEE Trans. Inf. Forensics Secur. | 11 |
| 2025 | Lighting is Unreliable: Adversarial Video Relighting Against rPPG Heart Rate Measurement
Menglin Zhang, Xiaoxin Guo, Xiaofeng Cao 0002, Shuifa Sun, Huazhu Fu, Qing Guo 0005 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | C-NeRF: Representing Scene Changes as Directional Consistency Difference-Based NeRFabstractIn this work, we aim to detect the changes caused by object variations in a scene represented by the neural radiance fields (NeRFs). Given an arbitrary view and two sets of scene images captured at different timestamps, we can predict scene changes in that view, which has significant potential applications in scene monitoring and measuring. We conducted preliminary studies and found that such an exciting task cannot be easily achieved by utilizing existing NeRFs and 2D change detection (CD) methods with many false or missing detections. The main reason is that the 2D CD is based on the pixel appearance difference between spatial-aligned image pairs and neglects the stereo information in the NeRF. To address the limitations, we propose the C-NeRF to represent scene changes as directional consistency difference-based NeRF, which mainly contains three modules. We first build two aligned NeRFs from pre-change and post-change scenes. Then, we identify the change points based on the direction-consistent constraint; that is, real change points have similar change representations across view directions, but fake change points do not. Finally, we design the change map rendering process based on the built NeRFs and can generate the change map of an arbitrarily specified view direction. To validate the effectiveness, we build a new dataset containing ten scenes covering diverse scenarios with different changing objects. Our approach surpasses state-of-the-art CD methods and NeRF-based methods by a significant margin. Rui Huang 0006, Haojie Tao, Binbin Jiang, Qingyi Zhao, Liang Wang 0001, Qing Guo 0005 |
IEEE Trans. Image Process. | 6 |
| 2025 | Adversarial Exposure Attack on Diabetic Retinopathy Imagery GradingabstractDiabetic Retinopathy (DR) is a leading cause of vision loss around the world. To help diagnose it, numerous cutting-edge works have built powerful deep neural networks (DNNs) to automatically grade DR via retinal fundus images (RFIs). However, RFIs are commonly affected by camera exposure issues that may lead to incorrect grades. The mis-graded results can potentially pose high risks to an aggravation of the condition. In this paper, we study this problem from the viewpoint of adversarial attacks. We identify and introduce a novel solution to an entirely new task, termed as adversarial exposure attack, which is able to produce natural exposure images and mislead the state-of-the-art DNNs. We validate our proposed method on a real-world public DR dataset with three DNNs, e.g., ResNet50, MobileNet, and EfficientNet, demonstrating that our method achieves high image quality and success rate in transferring the attacks. Our method reveals the potential threats to DNN-based automatic DR grading and would benefit the development of exposure-robust DR grading methods in the future. Yupeng Cheng, Qing Guo 0005, Felix Juefei-Xu, Huazhu Fu, Shangwei Lin 0001, Weisi Lin |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | CarveNet: Carving Point-Block for Complex 3D Shape Completionabstract3D point cloud completion is very challenging because it relies on accurately understanding the complex 3D shapes (e.g., high-curvature, concave/convex, and hollowed-out 3D shapes) and the unknown & diverse patterns of the partially available point clouds. In this paper, we propose a novel solution, i.e.,Point-block Carving(PC), for completing the complex 3D point cloud completion. Given the partial point cloud as the guidance, we carve a 3D block that contains the uniformly distributed 3D points, yielding the entire point cloud. We propose a new network architecture to achieve PC, i.e.,CarveNet. This network conducts the exclusive convolution on each block point, where the convolutional kernels are trained on the 3D shape data. CarveNet determines which point should be carved to recover the complete shapes' details effectively. Furthermore, we propose a sensor-aware method for data augmentation, i.e.,SensorAug, for training CarveNet on richer patterns of partial point clouds, thus enhancing the completion power of the network. The extensive evaluations on the ShapeNet, ShapNet-55/34 and KITTI datasets demonstrate the generality of our approach on the partial point clouds with diverse patterns. On these datasets, CarveNet successfully outperforms the state-of-the-art methods. Qing Guo 0005, Zhijie Wang 0014, Lubo Wang, Haotian Dong, Felix Juefei-Xu, Di Lin 0002, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003 |
IEEE Trans. Multim. | 1 |
| 2025 | Common Corruption Robustness of Point Cloud Detectors: Benchmark and EnhancementabstractObject detection through LiDAR-based point cloud has recently been important in autonomous driving. Although achieving high accuracy on public benchmarks, the state-of-the-art detectors may still go wrong and cause a heavy loss due to the widespread corruptions in the real world like rain, snow, sensor noise,etc. Nevertheless, there is a lack of a large-scale dataset covering diverse scenes and realistic corruption types with different severities to develop practical and robust point cloud detectors, which is challenging due to the heavy collection costs. To alleviate the challenge and start the first step for robust point cloud detection, we propose the physical-aware simulation methods to generate degraded point clouds under different real-world common corruptions. Then, for the first attempt, we construct a benchmark based on the physical-aware common corruptions for point cloud detectors, which contains a total of 1,122,150 examples covering 7,481 scenes, 25 common corruption types, and 6 severities. With such a novel benchmark, we conduct extensive empirical studies on 12 state-of-the-art detectors that contain 6 different detection frameworks. Thus we get several insight observations revealing the vulnerabilities of the detectors and indicating the enhancement directions. Moreover, we further study the effectiveness of existing robustness enhancement methods based on data augmentation, data denoising, test-time adaptation. The benchmark can potentially be a new platform for evaluating point cloud detectors, opening a door for developing novel robustness enhancement methods. Shuangzhi Li 0002, Zhijie Wang 0014, Felix Juefei-Xu, Qing Guo 0005, Lei Ma 0003 |
IEEE Trans. Multim. | 4 |
| 2025 | HRC-Net: Learning Visual Hypothesis, Representative, and Collaboration for Multi-Domain Image InpaintingabstractMulti-domain image inpainting utilizes complementary contextual information from auxiliary domain images to restore corrupted regions. While existing methods reconstruct auxiliary images to provide additional guidance, they face fundamental limitations: recovered pixels with complex patterns often lack representative details, while oversimplified patterns offer insufficient contextual information. To address these challenges, we propose HRC-Net, a novel framework incorporating three generative sub-networks for the comprehensive image inpainting task. Our architecture consists of: (1) A Hypothesis Sub-network that enables robust samplings of pixel-wise hypotheses from multi-domain inputs; (2) A Representative Sub-network that learns to score hypothesis quality based on contextual relevance; and (3) a Collaboration Sub-network that optimizes adaptive fusion kernels to integrate the most pertinent details. Together, these components model the joint distribution of representative scores and convolutional kernels, fostering a precise interaction between auxiliary hypotheses and target image corruption to meticulously repair the target image. Extensive evaluations across multiple benchmark datasets demonstrate HRC-Net's superior performance, significantly outperforming state-of-the-art methods in both quantitative metrics and visual quality. Xin Wang 0118, Di Lin 0002, Wanchao Su, Ji Du, Jie Zhang 0090, Haotian Dong, Ke Xu 0010, Qing Guo 0005, Ping Li 0016 |
ACM Trans. Graph. | 9 |
| 2024 | Personalization as a Shortcut for Few-Shot Backdoor Attack against Text-to-Image Diffusion ModelsabstractAlthough recent personalization methods have democratized high-resolution image synthesis by enabling swift concept acquisition with minimal examples and lightweight computation, they also present an exploitable avenue for highly accessible backdoor attacks. This paper investigates a critical and unexplored aspect of text-to-image (T2I) diffusion models - their potential vulnerability to backdoor attacks via personalization. By studying the prompt processing of popular personalization methods (epitomized by Textual Inversion and DreamBooth), we have devised dedicated personalization-based backdoor attacks according to the different ways of dealing with unseen tokens and divide them into two families: nouveau-token and legacy-token backdoor attacks. In comparison to conventional backdoor attacks involving the fine-tuning of the entire text-to-image diffusion model, our proposed personalization-based backdoor attack method can facilitate more tailored, efficient, and few-shot attacks. Through comprehensive empirical study, we endorse the utilization of the nouveau-token backdoor attack due to its impressive effectiveness, stealthiness, and integrity, markedly outperforming the legacy-token backdoor attack. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Jie Zhang 0002, Yutong Wu 0009, Ming Hu 0003, Tianlin Li, Geguang Pu, Yang Liu 0003 |
AAAI | 3 |
| 2024 | Cosalpure: Learning Concept from Group Images for Robust Co-Saliency DetectionabstractCo-salient object detection (CoSOD) aims to identify the common and salient (usually in the foreground) regions across a given group of images. Although achieving sig-nificant progress, state-of-the-art CoSODs could be easily affected by some adversarial perturbations, leading to sub-stantial accuracy reduction. The adversarial perturbations can mislead CoSODs but do not change the high-level se-mantic information (e.g., concept) of the co-salient objects. In this paper, we propose a novel robustness enhancement framework by first learning the concept of the co-salient ob-jects based on the input group images and then leveraging this concept to purify adversarial perturbations, which are subsequently fed to CoSODs for robustness enhancement. Specifically, we propose Cosalpure containing two modules, i.e., group-image concept learning and concept-guided diffusion purification. For the first module, we adopt a pre-trained text-to-image diffusion model to learn the con-cept of co-salient objects within group images where the learned concept is robust to adversarial examples. For the second module, we map the adversarial image to the latent space and then perform diffusion generation by embedding the learned concept into the noise prediction function as an extra condition. Our method can effectively alleviate the in-fluence of the SOTA adversarial attack containing different adversarial patterns, including exposure and noise. The ex-tensive results demonstrate that our method could enhance the robustness of Cos ODs significantly. The project is avail-able at https://vllen.github.io/CosalPure/. Jiayi Zhu 0002, Qing Guo 0005, Felix Juefei-Xu, Yihao Huang 0001, Yang Liu 0003, Geguang Pu |
CVPR | 2 |
| 2024 | Light the Night: A Multi-Condition Diffusion Framework for Unpaired Low-Light Enhancement in Autonomous DrivingabstractVision-centric perception systems for autonomous driving have gained considerable attention recently due to their cost-effectiveness and scalability, especially compared to LiDAR-based systems. However, these systems often struggle in low-light conditions, potentially compromising their performance and safety. To address this, our paper introduces LightDiff, a domain-tailored framework designed to enhance the low-light image quality for autonomous driving applications. Specifically, we employ a multi-condition controlled diffusion model. LightDiff works without any human-collected paired data, leveraging a dynamic data degradation process instead. It incorporates a novel multi-condition adapter that adaptively controls the input weights from different modalities, including depth maps, RGB images, and text captions, to effectively illuminate dark scenes while maintaining context consistency. Furthermore, to align the enhanced images with the detection model's knowledge, LightDiff employs perception-specific scores as rewards to guide the diffusion training process through reinforcement learning. Extensive experiments on the nuScenes datasets demonstrate that LightDiff can significantly improve the performance of several state-of-the-art 3D detectors in night-time conditions while achieving high visual quality scores, highlighting its potential to safeguard autonomous driving. Zhengzhong Tu, Xinyu Liu 0009, Qing Guo 0005, Felix Juefei-Xu, Runsheng Xu, Hongkai Yu |
CVPR | 5 |
| 2024 | Boosting Transferability in Vision-Language Attacks via Diversification Along the Intersection Region of Adversarial Trajectory
Sensen Gao, Xiaojun Jia, Xuhong Ren, Ivor W. Tsang, Qing Guo 0005 |
ECCV (57) | 5 |
| 2024 | Event Trojan: Asynchronous Event-Based Backdoor Attacks
Ruofei Wang, Qing Guo 0005, Haoliang Li, Renjie Wan |
ECCV (7) | 2 |
| 2024 | SAIR: Learning Semantic-Aware Implicit Representation
Canyu Zhang 0002, Qing Guo 0005, Song Wang 0002 |
ECCV (4) | 3 |
| 2024 | SPY-Watermark: Robust Invisible Watermarking for Backdoor AttackabstractBackdoor attack aims to deceive a victim model when facing backdoor instances while maintaining its performance on benign data. Current methods use manual patterns or special perturbations as triggers, while they often overlook the robustness against data corruption, making backdoor attacks easy to defend in practice. To address this issue, we propose a novel backdoor attack method named Spy-Watermark, which remains effective when facing data collapse and backdoor defense. Therein, we introduce a learnable watermark embedded in the latent domain of images, serving as the trigger. Then, we search for a watermark that can withstand collapse during image decoding, cooperating with several anti-collapse operations to further enhance the resilience of our trigger against data corruption. Extensive experiments are conducted on CIFAR10, GTSRB, and ImageNet datasets, demonstrating that Spy-Watermark overtakes ten state-of-the-art methods in terms of robustness and stealthiness. Ruofei Wang, Renjie Wan, Zongyu Guo, Qing Guo 0005, Rui Huang 0006 |
ICASSP | 4 |
| 2024 | Architecture-Agnostic Iterative Black-Box Certified Defense Against Adversarial PatchesabstractThe adversarial patch attack aims to fool image classifiers within a bounded, contiguous region of arbitrary changes. To address this problem in a trustworthy way, the certified patch defense methods are proposed. However, the state-of-the-art certified defenses inevitably needed to access the size of the adversarial patch, which is unreasonable and impractical in real-world attack scenarios. To improve the feasibility of the architecture-agnostic certified defense in a black-box setting, we propose a novel two-stage Iterative Black-box Certified Defense method, termed IBCD. In the first stage, it estimates the patch size in a search-based manner by evaluating the size relationship between the patch and mask with pixel masking. In the second stage, the accuracy results are calculated by the existing white-box certified defense methods with the estimated patch size. The experiments conducted on two popular model architectures and two datasets verify the effectiveness and efficiency of IBCD. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Ming Hu 0003, Yang Liu 0003, Geguang Pu |
ICASSP | 3 |
| 2024 | IRAD: Implicit Representation-driven Image Resampling against Adversarial AttacksabstractWe introduce a novel approach to counter adversarial attacks, namely, image resampling. Image resampling transforms a discrete image into a new one, simulating the process of scene recapturing or rerendering as specified by a geometrical transformation. The underlying rationale behind our idea is that image resampling can alleviate the influence of adversarial perturbations while preserving essential semantic information, thereby conferring an inherent advantage in defending against adversarial attacks. To validate this concept, we present a comprehensive study on leveraging image resampling to defend against adversarial attacks. We have developed basic resampling methods that employ interpolation strategies and coordinate shifting magnitudes. Our analysis reveals that these basic methods can partially mitigate adversarial attacks. However, they come with apparent limitations: the accuracy of clean images noticeably decreases, while the improvement in accuracy on adversarial examples is not substantial.We propose implicit representation-driven image resampling (IRAD) to overcome these limitations. First, we construct an implicit continuous representation that enables us to represent any input image within a continuous coordinate space. Second, we introduce SampleNet, which automatically generates pixel-wise shifts for resampling in response to different inputs. Furthermore, we can extend our approach to the state-of-the-art diffusion-based method, accelerating it with fewer time steps while preserving its defense capability. Extensive experiments demonstrate that our method significantly enhances the adversarial robustness of diverse deep models against various attacks while maintaining high accuracy on clean images. Tianlin Li, Xiaofeng Cao 0002, Ivor W. Tsang, Yang Liu 0003, Qing Guo 0005 |
ICLR | 6 |
| 2024 | LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking AttacksabstractVisual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial perturbations in the incoming frames. This can lead to significant robustness and security issues when these trackers are deployed in the real world. To achieve high accuracy on both clean and adversarial data, we propose building a spatial-temporal continuous representation using the semantic text guidance of the object of interest. This novel continuous representation enables us to reconstruct incoming frames to maintain semantic and appearance consistency with the object of interest and its clean counterparts. As a result, our proposed method successfully defends against different SOTA adversarial tracking attacks while maintaining high accuracy on clean data. In particular, our method significantly increases tracking accuracy under adversarial attacks with around 90% relative improvement on UAV123, which is even higher than the accuracy on clean data. Jianlang Chen, Xuhong Ren, Qing Guo 0005, Felix Juefei-Xu, Di Lin 0002, Wei Feng 0005, Lei Ma 0003, Jianjun Zhao 0001 |
ICLR | 3 |
| 2024 | AdvGPS: Adversarial GPS for Multi-Agent Perception AttackabstractThe multi-agent perception system collects visual data from sensors located on various agents and leverages their relative poses determined by GPS signals to effectively fuse information, mitigating the limitations of single-agent sensing, such as occlusion. However, the precision of GPS signals can be influenced by a range of factors, including wireless transmission and obstructions like buildings. Given the pivotal role of GPS signals in perception fusion and the potential for various interference, it becomes imperative to investigate whether specific GPS signals can easily mislead the multi-agent perception system. To address this concern, we frame the task as an adversarial attack challenge and introduce ADVGPS, a method capable of generating adversarial GPS signals which are also stealthy for individual agents within the system, significantly reducing object detection accuracy. To enhance the success rates of these attacks in a black-box scenario, we introduce three types of statistically sensitive natural discrepancies: appearance-based discrepancy, distribution-based discrepancy, and task-aware discrepancy. Our extensive experiments on the OPV2V dataset demonstrate that these attacks substantially undermine the performance of state-of-the-art methods, showcasing remarkable transferability across different point cloud based 3D detection systems. This alarming revelation underscores the pressing need to address security implications within multi-agent perception systems, thereby underscoring a critical area of research. The code is available at https://github.com/jinlong17/AdvGPS. Xinyu Liu 0009, Jianwu Fang, Felix Juefei-Xu, Qing Guo 0005, Hongkai Yu |
ICRA | 6 |
| 2024 | RUNNER: Responsible UNfair NEuron Repair for Enhancing Deep Neural Network FairnessabstractDeep Neural Networks (DNNs), an emerging software technology, have achieved impressive results in a variety of fields. However, the discriminatory behaviors towards certain groups (a.k.a. unfairness) of DNN models increasingly become a social concern, especially in high-stake applications such as loan approval and criminal risk assessment. Although there has been a number of works to improve model fairness, most of them adopt an adversary to either expand the model architecture or augment training data, which introduces excessive computational overhead. Recent work diagnoses responsible unfair neurons first and fixes them with selective retraining. Unfortunately, existing diagnosis process is time-consuming due to multi-step training sample analysis, and selective retraining may cause a performance bottleneck due to indirectly adjusting unfair neurons on biased samples. In this paper, we propose Responsible UNfair NEuron Repair (RUNNER) that improves existing works in three key aspects: (1) efficiency: we design the Importance-based Neuron Diagnosis that identifies responsible unfair neurons in one step with a novel importance criterion of neurons; (2) effectiveness: we design the Neuron Stabilizing Retraining by adding a loss term that measures the activation distance of responsible unfair neurons from different subgroups in all sources; (3) generalization: we investigate the effectiveness on both structured tabular data and large-scale unstructured image data, which is often ignored in prior studies. Our extensive experiments across 5 datasets show that RUUNER can effectively and efficiently diagnose and repair the DNNs regarding unfairness. On average, our approach significantly reduces computing overhead from 341.7s to 29.65s, and achieves improved fairness up to 79.3%. Besides, RUNNER also keeps state-of-the-art results on the unstructured dataset. Tianlin Li, Jian Zhang 0087, Shiqian Zhao, Yihao Huang 0001, Aishan Liu, Qing Guo 0005, Yang Liu 0003 |
ICSE | 7 |
| 2024 | MetaRepair: Learning to Repair Deep Neural Networks from Repairing Experiences
Yun Xing 0001, Qing Guo 0005, Xiaofeng Cao 0002, Ivor W. Tsang, Lei Ma 0003 |
ACM Multimedia | 2 |
| 2024 | Sim2Real-Fire: A Multi-modal Simulation Dataset for Forecast and Backtracking of Real-world Forest FireabstractThe latest research on wildfire forecast and backtracking has adopted AI models, which require a large amount of data from wildfire scenarios to capture fire spread patterns. This paper explores using cost-effective simulated wildfire scenarios to train AI models and apply them to the analysis of real-world wildfire. This solution requires AI models to minimize the Sim2Real gap, a brand-new topic in the fire spread analysis research community. To investigate the possibility of minimizing the Sim2Real gap, we collect the Sim2Real-Fire dataset that contains 1M simulated scenarios with multi-modal environmental information for training AI models. We prepare 1K real-world wildfire scenarios for testing the AI models. We also propose a deep transformer, S2R-FireTr, which excels in considering the multi-modal environmental information for forecasting and backtracking the wildfire. S2R-FireTr surpasses state-of-the-art methods in real-world wildfire scenarios. Keqiu Li, Li Guohui, Changqing Ji, Lubo Wang, Die Zuo, Qing Guo 0005, Manyu Wang 0001, Di Lin 0002 |
NeurIPS | 8 |
| 2024 | ColJailBreak: Collaborative Generation and Editing for Jailbreaking Text-to-Image Deep GenerationabstractThe commercial text-to-image deep generation models (e.g. DALL·E) can produce high-quality images based on input language descriptions. These models incorporate a black-box safety filter to prevent the generation of unsafe or unethical content, such as violent, criminal, or hateful imagery. Recent jailbreaking methods generate adversarial prompts capable of bypassing safety filters and producing unsafe content, exposing vulnerabilities in influential commercial models. However, once these adversarial prompts are identified, the safety filter can be updated to prevent the generation of unsafe images. In this work, we propose an effective, simple, and difficult-to-detect jailbreaking solution: generating safe content initially with normal text prompts and then editing the generations to embed unsafe content. The intuition behind this idea is that the deep generation model cannot reject safe generation with normal text prompts, while the editing models focus on modifying the local regions of images and do not involve a safety strategy. However, implementing such a solution is non-trivial, and we need to overcome several challenges: how to automatically confirm the normal prompt to replace the unsafe prompts, and how to effectively perform editable replacement and naturally generate unsafe content. In this work, we propose the collaborative generation and editing for jailbreaking text-to-image deep generation (ColJailBreak), which comprises three key components: adaptive normal safe substitution, inpainting-driven injection of unsafe content, and contrastive language-image-guided collaborative optimization. We validate our method on three datasets and compare it to two baseline methods. Our method could generate unsafe content through two commercial deep generation models including GPT-4 and DALL·E 2. Yizhuo Ma, Shanmin Pang, Qi Guo 0008, Qing Guo 0005 |
NeurIPS | 5 |
| 2024 | Geometry Awakening: Cross-Geometry Learning Exhibits Superiority over Individual StructuresabstractRecent research has underscored the efficacy of Graph Neural Networks (GNNs) in modeling diverse geometric structures within graph data. However, real-world graphs typically exhibit geometrically heterogeneous characteristics, rendering the confinement to a single geometric paradigm insufficient for capturing their intricate structural complexities. To address this limitation, we examine the performance of GNNs across various geometries through the lens of knowledge distillation (KD) and introduce a novel cross-geometric framework. This framework encodes graphs by integrating both Euclidean and hyperbolic geometries in a space-mixing fashion. Our approach employs multiple teacher models, each generating hint embeddings that encapsulate distinct geometric properties. We then implement a structure-wise knowledge transfer module that optimally leverages these embeddings within their respective geometric contexts, thereby enhancing the training efficacy of the student model. Additionally, our framework incorporates a geometric optimization network designed to bridge the distributional disparities among these embeddings. Experimental results demonstrate that our model-agnostic framework more effectively captures topological graph knowledge, resulting in superior performance of the student models when compared to traditional KD methodologies. Yadong Sun, Xiaofeng Cao 0002, Yu Wang 0152, Wei Ye 0001, Jingcai Guo, Qing Guo 0005 |
NeurIPS | 6 |
| 2024 | Voxel Proposal Network via Multi-Frame Knowledge Distillation for Semantic Scene CompletionabstractSemantic scene completion is a difficult task that involves completing the geometry and semantics of a scene from point clouds in a large-scale environment. Many current methods use 3D/2D convolutions or attention mechanisms, but these have limitations in directly constructing geometry and accurately propagating features from related voxels, the completion likely fails while propagating features in a single pass without considering multiple potential pathways. And they are generally only suitable for static scenes and struggle to handle dynamic aspects. This paper introduces Voxel Proposal Network (VPNet) that completes scenes from 3D and Bird's-Eye-View (BEV) perspectives. It includes Confident Voxel Proposal based on voxel-wise coordinates to propose confident voxels with high reliability for completion. This method reconstructs the scene geometry and implicitly models the uncertainty of voxel-wise semantic labels by presenting multiple possibilities for voxels. VPNet employs Multi-Frame Knowledge Distillation based on the point clouds of multiple adjacent frames to accurately predict the voxel-wise labels by condensing various possibilities of voxel relationships. VPNet has shown superior performance and achieved state-of-the-art results on the SemanticKITTI and SemanticPOSS datasets. Lubo Wang, Di Lin 0002, Kairui Yang, Qing Guo 0005, Wuyuan Xie, Miaohui Wang, Lingyu Liang, Ping Li 0016 |
NeurIPS | 5 |
| 2024 | Defense against Adversarial Cloud Attack on Remote Sensing Salient Object DetectionabstractDetecting the salient objects in a remote sensing image has wide applications. Many existing deep learning methods have been proposed for Salient Object Detection (SOD) in remote sensing images with remarkable results. However, the recent adversarial attack examples, generated by changing a few pixel values on the original image, could result in a collapse for the well-trained deep learning model. Different with existing methods adding perturbation to original images, we propose to jointly tune adversarial exposure and additive perturbation for attack and constrain image close to cloudy image as Adversarial Cloud. Cloud is natural and common in remote sensing images, however, camouflaging cloud based adversarial attack and defense for remote sensing images are not well studied before. Furthermore, we design DefenseNet as a learnable pre-processing to the adversarial cloudy images to preserve the performance of the deep learning based remote sensing SOD model, without tuning the already deployed deep SOD model. By considering both regular and generalized adversarial examples, the proposed DefenseNet can defend the proposed Adversarial Cloud in white-box setting and other attack methods in black-box setting. Experimental results on a synthesized benchmark from the public remote sensing dataset (EORSSD) show the promising defense against adversarial cloud attacks. Huiming Sun, Lan Fu, Qing Guo 0005, Zibo Meng, Tianyun Zhang, Yuewei Lin, Hongkai Yu |
WACV | 4 |
| 2024 | Dodging DeepFake Detection via Implicit Spatial-Domain Notch FilteringabstractThe current high-fidelity generation and high-precision detection of DeepFake images are at an arms race. We believe that producing DeepFakes that are highly realistic and “detection evasive” can serve the ultimate goal of improving future generation DeepFake detection capabilities. In this paper, we propose a simple yet powerful pipeline to reduce the artifact patterns of fake images without hurting image quality by performing implicit spatial-domain notch filtering. We first demonstrate that frequency-domain notch filtering, although famously shown to be effective in removing periodic noise in the spatial domain, is infeasible for our task at hand due to the manual designs required for the notch filters. We, therefore, resort to a learning-based approach to reproduce the notch filtering effects, but solely in the spatial domain. We adopt a combination of adding overwhelming spatial noise for breaking the periodic noise pattern and deep image filtering to reconstruct the noise-free fake images, and we name our method DeepNotch. Deep image filtering provides a specialized filter for each pixel in the noisy image, producing filtered images with high fidelity compared to their DeepFake counterparts. Moreover, we also use the semantic information of the image to generate an adversarial guidance map to add noise intelligently. Our large-scale evaluation on 3 representative DeepFake detection methods (tested on 16 types of DeepFakes) has demonstrated that our technique significantly reduces the accuracy of these 3 fake image detection methods, 36.79% on average and up to 97.02% in the best case. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Yang Liu 0003, Geguang Pu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | AdaAug+: A Reinforcement Learning-Based Adaptive Data Augmentation for Change DetectionabstractData augmentation (DA) increases the diversity of training data to improve the model generalization ability. Most of the DA methods focus on image classification or object detection. Directly using the existing DA methods on change detection tasks not only falls short of fully exploring the specificity of the change image pairs but also leads to longer training times. In this article, we first propose a mask-guided mixing (MGM) DA for change detection, which mixes the change regions of the current training sample based on prediction results and labels to generate high-quality samples with more positive samples. We then propose a new reinforcement learning (RL)-based Adaptive DA method, AdaAug+, to adaptively select the optimal DA policy for the training samples. An actor selects the best augmentation operation from the operation set according to the image pair. The augmented image pairs make it easier for the change detector to learn the optimal parameters and improve the final detection performance. To reduce the training time, we identify and remove the redundant training samples during the training process by our redundancy searching policy. We have conducted various experiments on four remote sensing change detection datasets with different change detectors. The experimental results demonstrate that AdaAug+ achieves promising performance compared to the state-of-the-art DA methods and requires less training time. Rui Huang 0006, Jieda Wei, Qing Guo 0005 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Texture Re-Scalable Universal Adversarial PerturbationabstractUniversal adversarial perturbation (UAP), also known as image-agnostic perturbation, is a fixed perturbation map that can fool the classifier with high probabilities on arbitrary images, making it more practical for attacking deep models in the real world. Previous UAP methods generate a scale-fixed and texture-fixed perturbation map for all images, which ignores the multi-scale objects in images and usually results in a low fooling ratio. Since the widely used convolution neural networks tend to classify objects according to semantic information stored in local textures, it seems a reasonable and intuitive way to improve the UAP from the perspective of utilizing local contents effectively. In this work, we find that the fooling ratios significantly increase when we add a constraint to encourage a small-scale UAP map and repeat it vertically and horizontally to fill the whole image domain. To this end, we propose texture scale-constrained UAP (TSC-UAP), a simple yet effective UAP enhancement method that automatically generates UAPs with category-specific local textures that can fool deep models more easily. Through a low-cost operation that restricts the texture scale, TSC-UAP achieves a considerable improvement in the fooling ratio and attack transferability for both data-dependent and data-free UAP methods. Experiments conducted on two state-of-the-art UAP methods, eight popular CNN models and four classical datasets show the remarkable performance of TSC-UAP. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Ming Hu 0003, Xiaojun Jia, Xiaochun Cao, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Minimalism is King! High-Frequency Energy-Based Screening for Data-Efficient Backdoor AttacksabstractGiven the effectiveness of deep neural networks in various fields, the security of neural networks has received great attention. The backdoor attack, which induces malicious behaviors of models by poisoning part of the training set, still remains a challenging problem. Many recent efforts have proposed different ways of embedding backdoors to improve the stealthiness of backdoor attacks. Yet, lowering the percentage of poisoned samples is one of the most direct ways to increase stealthiness. A recent study (Filtering-and-Updating strategy, FUS) has revealed that the sample selection for poisoning is also crucial, as different samples contribute differently to the final decision boundary of the network. Concretely, they utilize each sample’s forgetting events during the training stage to identify which samples will contribute more to the network’s prediction. The training phase of their search method, however, is computationally expensive and slow. To overcome this, in this paper, we propose an efficient sample selection strategy based on the high-frequency energy (HFE) of training samples with a global screening and updating strategy, which can not only achieve a higher backdoor-attack success rate but also reduce the searching time by a factor of 4320 compared to FUS (12 hours vs 10 seconds). The extensive experiment results on CIFAR-10, CIFAR-100, and ImageNet-10 have shown that our proposed method is much simpler, faster, and more efficient. Yuan Xun, Xiaojun Jia, Jindong Gu, Qing Guo 0005, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Adversarial Relighting Against Face RecognitionabstractDeep face recognition (FR) has achieved significantly high accuracy on several challenging datasets and fosters successful real-world applications, even showing high robustness to the illumination variation that is usually regarded as a main threat to the FR system. However, in the real world, illumination variation caused by diverse lighting conditions cannot be fully covered by the limited face dataset. In this paper, we study the threat of lighting against FR from a new angle,i.e.,adversarial attack, and identify a new task,i.e.,adversarial relighting. Given a face image, adversarial relighting aims to produce a naturally relighted counterpart while fooling the state-of-the-art deep FR methods. To this end, we first propose the physical model-based adversarial relighting attack (ARA) denoted asalbedo-quotient-based adversarial relighting attack (AQ-ARA). It generates natural adversarial lighting under the guidance of FR systems and synthesizes adversarially relighted face images. Moreover, we propose theauto-predictive adversarial relighting attack (AP-ARA)by training an adversarial relighting network (ARNet) to automatically predict the adversarial lighting in a one-step manner according to different input faces, allowing efficiency-sensitive applications. More importantly, we propose to transfer the above digital attacks tophysical ARA (Phy-ARA)through a precise relighting device, making the estimated adversarial lighting condition reproducible in the real world. We validate our methods on several state-of-the-art deep FR methods on two public datasets. The extensive and insightful results demonstrate our work can generate realistic adversarial relighted face images fooling face recognition tasks easily, revealing the threat of specific light directions and strengths. Qian Zhang 0051, Qing Guo 0005, Ruijun Gao, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Transductive Reward Inference on GraphabstractIn this study, we present a transductive inference approach on that reward information propagation graph, which enables the effective estimation of rewards for unlabelled data in offline reinforcement learning. Reward inference is the key to learning effective policies in practical scenarios, while direct environmental interactions are either too costly or unethical and the reward functions are rarely accessible, such as in healthcare and robotics. Our research focuses on developing a reward inference method based on the contextual properties of information propagation on graphs that capitalizes on a constrained number of human reward annotations to infer rewards for unlabelled data. We leverage both the available data and limited reward annotations to construct a reward propagation graph, wherein the edge weights incorporate various influential factors pertaining to the rewards. Subsequently, we employ the constructed graph for transductive reward inference, thereby estimating rewards for unlabelled data. Furthermore, we establish the existence of a fixed point during several iterations of the transductive inference process and demonstrate its at least convergence to a local optimum. Empirical evaluations on locomotion and robotic manipulation tasks validate the effectiveness of our approach. The application of our inferred rewards improves the performance in offline reinforcement learning tasks. Bohao Qu, Xiaofeng Cao 0002, Qing Guo 0005, Yi Chang 0001, Ivor W. Tsang, Chengqi Zhang |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | Natural & Adversarial Bokeh Rendering via Circle-of-Confusion Predictive NetworkabstractBokeh effect is a natural shallow depth-of-field phenomenon that blurs the out-of-focus part in photography. In recent years, a series of works have proposed automatic and realistic bokeh rendering methods for artistic and aesthetic purposes. They usually employ cutting-edge data-driven deep generative networks with complex training strategies and network architectures. However, these works neglect that the bokeh effect, as a real phenomenon, can inevitably affect the subsequent visual intelligent tasks like recognition, and their data-driven nature prevents them from studying the influence of bokeh-related physical parameters (i.e., depth-of-the-field) on the intelligent tasks. To fill this gap, we study a totally new problem, i.e.,natural & adversarial bokeh rendering, which consists of two objectives: rendering realistic and natural bokeh and fooling the visual perception models (i.e., bokeh-based adversarial attack). To this end, beyond the pure data-driven solution, we propose a hybrid alternative by taking the respective advantages of data-driven and physical-aware methods. Specifically, we propose thecircle-of-confusion predictive network (CoCNet)by taking the all-in-focus image and depth image as inputs to estimate circle-of-confusion parameters for each pixel, which are employed to render the final image through a well-known physical model of bokeh. With the hybrid solution, our method could achieve more realistic rendering results with the naive training strategy and a much lighter network. Moreover, we propose the adversarial bokeh attack by fixing the CoCNet while optimizing the depth map w.r.t. the visual perception tasks. Then, we are able to study the vulnerability of deep neural networks according to the depth variations in the real world. The extensive experiments show that our method produces more realistic bokeh than the state-of-the-art methods while fooling the powerful deep neural networks with a high accuracy drop. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Geguang Pu, Yang Liu 0003 |
IEEE Trans. Multim. | 3 |
| 2024 | Faire: Repairing Fairness of Neural Networks via Neuron Condition SynthesisabstractDeep Neural Networks (DNNs) have achieved tremendous success in many applications, while it has been demonstrated that DNNs can exhibit some undesirable behaviors on concerns such as robustness, privacy, and other trustworthiness issues. Among them, fairness (i.e., non-discrimination) is one important property, especially when they are applied to some sensitive applications (e.g., finance and employment). However, DNNs easily learn spurious correlations between protected attributes (e.g., age, gender, race) and the classification task and develop discriminatory behaviors if the training data is imbalanced. Such discriminatory decisions in sensitive applications would introduce severe social impacts. To expose potential discrimination problems in DNNs before putting them in use, some testing techniques have been proposed to identify the discriminatory instances (i.e., instances that show defined discrimination 1 ). However, how to repair DNNs after detecting such discrimination is still challenging. Existing techniques mainly rely on retraining on a large number of discriminatory instances generated by testing methods, which requires huge time overhead and makes the repairing inefficient. In this work, we propose the method Faire to effectively and efficiently repair the fairness issues of DNNs, without using additional data (e.g., discriminatory instances). Our basic idea is inspired by the traditional program repair method that synthesizes proper condition checking. To repair traditional programs, a typical method is to localize the program defects and repair the program logic by adding condition checking. Similarly, for DNNs, we try to understand the unfair logic and reformulate it with well-designed condition checking. In this article, we synthesize the condition that can reduce the effect of features relevant to the protected attributes in the DNN. Specifically, we first perform the neuron-based analysis and check the functionalities of neurons to identify neurons whose outputs could be regarded as features relevant to protected attributes and original tasks. Then a new condition layer is added after each hidden layer to penalize neurons that are accountable for the protected features (i.e., intermediate features relevant to protected attributes) and promote neurons that are accountable for the non-protected features (i.e., intermediate features relevant to original tasks). In sum, the repair rate 2 of Faire reaches up to more than 99%, which outperforms other methods, and the whole repairing process only takes no more than 340 s. The evaluation results demonstrate that our approach can effectively and efficiently repair the individual discriminatory instances of the target model. Tianlin Li, Xiaofei Xie, Jian Wang 0067, Qing Guo 0005, Aishan Liu, Lei Ma 0003, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Background-Mixed Augmentation for Weakly Supervised Change DetectionabstractChange detection (CD) is to decouple object changes (i.e., object missing or appearing) from background changes (i.e., environment variations) like light and season variations in two images captured in the same scene over a long time span, presenting critical applications in disaster management, urban development, etc. In particular, the endless patterns of background changes require detectors to have a high generalization against unseen environment variations, making this task significantly challenging. Recent deep learning-based methods develop novel network architectures or optimization strategies with paired-training examples, which do not handle the generalization issue explicitly and require huge manual pixel-level annotation efforts. In this work, for the first attempt in the CD community, we study the generalization issue of CD from the perspective of data augmentation and develop a novel weakly supervised training algorithm that only needs image-level labels. Different from general augmentation techniques for classification, we propose the background-mixed augmentation that is specifically designed for change detection by augmenting examples under the guidance of a set of background changing images and letting deep CD models see diverse environment variations. Moreover, we propose the augmented & real data consistency loss that encourages the generalization increase significantly. Our method as a general framework can enhance a wide range of existing deep learning-based detectors. We conduct extensive experiments in two public datasets and enhance four state-of-the-art methods, demonstrating the advantages of our method. We release the code at https://github.com/tsingqguo/bgmix. Rui Huang 0006, Ruofei Wang, Qing Guo 0005, Jieda Wei, Yuxiang Zhang 0003, Wei Fan 0001, Yang Liu 0003 |
AAAI | 3 |
| 2023 | Distilling Cross-Temporal Contexts for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to recognize glosses in a sign language video. State-of-the-art methods typically have two modules, a spatial perception module and a temporal aggregation module, which are jointly learned end-to-end. Existing results in [9, 20, 25, 36] have indicated that, as the frontal component of the over-all model, the spatial perception module used for spatial feature extraction tends to be insufficiently trained. In this paper, we first conduct empirical studies and show that a shallow temporal aggregation module allows more thor-ough training of the spatial perception module. However, a shallow temporal aggregation module cannot well capture both local and global temporal context information in sign language. To address this dilemma, we propose a cross-temporal context aggregation (CTCA) model. Specifically, we build a dual-path network that contains two branches for perceptions of local temporal context and global temporal context. We further design a cross-context knowledge distil-lation learning objective to aggregate the two types of con-text and the linguistic prior. The knowledge distillation en-ables the resultant one-branch temporal aggregation mod-ule to perceive local-global temporal and semantic context. This shallow temporal perception module structure facili-tates spatial perception module learning. Extensive exper-iments on challenging CSLR benchmarks demonstrate that our method outperforms all state-of-the-art methods. Leming Guo, Wanli Xue, Qing Guo 0005, Bo Liu 0005, Kaihua Zhang 0001, Tiantian Yuan, Shengyong Chen |
CVPR | 3 |
| 2023 | Evading DeepFake Detectors via Adversarial Statistical ConsistencyabstractIn recent years, as various realistic face forgery techniques known as DeepFake improves by leaps and bounds, more and more DeepFake detection techniques have been proposed. These methods typically rely on detecting statistical differences between natural (i.e., real) and DeepFake-generated images in both spatial and frequency domains. In this work, we propose to explicitly minimize the statistical differences to evade state-of-the-art DeepFake detectors. To this end, we propose a statistical consistency attack (StatAttack) against DeepFake detectors, which contains two main parts. First, we select several statistical-sensitive natural degradations (i.e., exposure, blur, and noise) and add them to the fake images in an adversarial way. Second, we find that the statistical differences between natural and DeepFake images are positively associated with the distribution shifting between the two kinds of images, and we propose to use a distribution-aware loss to guide the optimization of different degradations. As a result, the feature distributions of generated adversarial examples is close to the natural images. Furthermore, we extend the StatAttack to a more powerful version, MStatAttack, where we extend the single-layer degradation to multi-layer degradations sequentially and use the loss to tune the combination weights jointly. Comprehensive experimental results on four spatial-based detectors and two frequency-based detectors with four datasets demonstrate the effectiveness of our proposed attack method in both white-box and black-box settings. Qing Guo 0005, Yihao Huang 0001, Xiaofei Xie, Lei Ma 0003, Jianjun Zhao 0001 |
CVPR | 2 |
| 2023 | CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image ClassificationabstractThis paper presents a CLIP-based unsupervised learning method for annotation-free multi-label image classification, including three stages: initialization, training, and inference. At the initialization stage, we take full advantage of the powerful CLIP model and propose a novel approach to extend CLIP for multi-label predictions based on globallocal image-text similarity aggregation. To be more specific, we split each image into snippets and leverage CLIP to generate the similarity vector for the whole image (global) as well as each snippet (local). Then a similarity aggregator is introduced to leverage the global and local similarity vectors. Using the aggregated similarity scores as the initial pseudo labels at the training stage, we propose an optimization framework to train the parameters of the classification network and refine pseudo labels for unobserved labels. During inference, only the classification network is used to predict the labels of the input image. Extensive experiments show that our method outperforms state-of-the-art unsupervised methods on MS-COCO, PASCAL VOC 2007, PASCAL VOC 2012, and NUS datasets and even achieves comparable results to weakly supervised classification methods. Rabab Abdelfattah, Qing Guo 0005, Xiaofeng Wang 0007, Song Wang 0002 |
ICCV | 2 |
| 2023 | CVSformer: Cross-View Synthesis Transformer for Semantic Scene CompletionabstractSemantic scene completion (SSC) requires an accurate understanding of the geometric and semantic relationships between the objects in the 3D scene for reasoning the occluded objects. The popular SSC methods voxelize the 3D objects, allowing the deep 3D convolutional network (3D CNN) to learn the object relationships from the complex scenes. However, the current networks lack the controllable kernels to model the object relationship across multiple views, where appropriate views provide the relevant information for suggesting the existence of the occluded objects. In this paper, we propose Cross-View Synthesis Transformer (CVSformer), which consists of Multi-View Feature Synthesis and Cross-View Transformer for learning cross-view object relationships. In the multi-view feature synthesis, we use a set of 3D convolutional kernels rotated differently to compute the multi-view features for each voxel. In the cross-view transformer, we employ the cross-view fusion to comprehensively learn the cross-view relationships, which form useful information for enhancing the features of individual views. We use the enhanced features to predict the geometric occupancies and semantic labels of all voxels. We evaluate CVSformer on public datasets, where CVS-former yields state-of-the-art results. Our code is available at https://github.com/donghaotian123/CVSformer. Haotian Dong, Enhui Ma, Lubo Wang, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016, Lingyu Liang, Kairui Yang, Di Lin 0002 |
ICCV | 6 |
| 2023 | Leveraging Inpainting for Single-Image Shadow RemovalabstractFully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are much lower than those of fully-supervised methods. In this work, we find that pretraining shadow removal networks on the image inpainting dataset can reduce the shadow remnants significantly: a naive encoder-decoder network gets competitive restoration quality w.r.t. the state-of-the-art methods via only 10% shadow & shadow-free image pairs. After analyzing networks with/without inpainting pretraining via the information stored in the weight (IIW), we find that inpainting pretraining improves restoration quality in non-shadow regions and enhances the generalization ability of networks significantly. Additionally, shadow removal fine-tuning enables networks to fill in the details of shadow regions. Inspired by these observations we formulate shadow removal as an adaptive fusion task that takes advantage of both shadow removal and image inpainting. Specifically, we develop an adaptive fusion network consisting of two encoders, an adaptive fusion block, and a decoder. The two encoders are responsible for extracting the features from the shadow image and the shadow-masked image respectively. The adaptive fusion block is responsible for combining these features in an adaptive manner. Finally, the decoder converts the adaptive fused features to the desired shadow-free result. The extensive experiments show that our method empowered with inpainting outperforms all state-of-the-art methods. We have realized codes and models in https://github.com/tsingqguo/inpaint4shadow Qing Guo 0005, Rabab Abdelfattah, Di Lin 0002, Wei Feng 0005, Ivor W. Tsang, Song Wang 0002 |
ICCV | 2 |
| 2023 | CopyRNeRF: Protecting the CopyRight of Neural Radiance FieldsabstractNeural Radiance Fields (NeRF) have the potential to be a major representation of media. Since training a NeRF has never been an easy task, the protection of its model copyright should be a priority. In this paper, by analyzing the pros and cons of possible copyright protection solutions, we propose to protect the copyright of NeRF models by replacing the original color representation in NeRF with a watermarked color representation. Then, a distortion-resistant rendering scheme is designed to guarantee robust message extraction in 2D renderings of NeRF. Our proposed method can directly protect the copyright of NeRF models while maintaining high rendering quality and bit accuracy when compared among optional solutions. Project page: https://luo-ziyuan.github.io/copyrnerf. Ziyuan Luo, Qing Guo 0005, Ka Chun Cheung, Simon See, Renjie Wan |
ICCV | 2 |
| 2023 | FAIRER: Fairness as Decision Rationale AlignmentabstractDeep neural networks (DNNs) have made significant progress, but often suffer from fairness issues, as deep models typically show distinct accuracy differences among certain subgroups (e.g., males and females). Existing research addresses this critical issue by employing fairness-aware loss functions to constrain the last-layer outputs and directly regularize DNNs. Although the fairness of DNNs is improved, it is unclear how the trained network makes a fair prediction, which limits future fairness improvements. In this paper, we investigate fairness from the perspective of decision rationale and define the parameter parity score to characterize the fair decision process of networks by analyzing neuron influence in various subgroups. Extensive empirical studies show that the unfair issue could arise from the unaligned decision rationales of subgroups. Existing fairness regularization terms fail to achieve decision rationale alignment because they only constrain last-layer outputs while ignoring intermediate neuron alignment. To address the issue, we formulate the fairness as a new task, i.e., decision rationale alignment that requires DNNs’ neurons to have consistent responses on subgroups at both intermediate processes and the final prediction. To make this idea practical during optimization, we relax the naive objective function and propose gradient-guided parity alignment, which encourages gradient-weighted consistency of neurons across subgroups. Extensive experiments on a variety of datasets show that our method can significantly enhance fairness while sustaining a high level of accuracy and outperforming other approaches by a wide margin. Tianlin Li, Qing Guo 0005, Aishan Liu, Mengnan Du, Yang Liu 0003 |
ICML | 2 |
| 2023 | Fairness via Group Contribution MatchingabstractFairness issues in Deep Learning models have recently received increasing attention due to their significant societal impact. Although methods for mitigating unfairness are constantly proposed, little research has been conducted to understand how discrimination and bias develop during the standard training process. In this study, we propose analyzing the contribution of each subgroup (i.e., a group of data with the same sensitive attribute) in the training process to understand the cause of such bias development process. We propose a gradient-based metric to assess training subgroup contribution disparity, showing that unequal contributions from different subgroups are one source of such unfairness. One way to balance the contribution of each subgroup is through oversampling, which ensures that an equal number of samples are drawn from each subgroup during each training iteration. However, we have found that even with a balanced number of samples, the contribution of each group remains unequal, resulting in unfairness under the oversampling strategy. To address the above issues, we propose an easy but effective group contribution matching (GCM) method to match the contribution of each subgroup. Our experiments show that our GCM effectively improves fairness and outperforms other methods significantly. Tianlin Li, Anran Li 0001, Mengnan Du, Aishan Liu, Qing Guo 0005, Guozhu Meng, Yang Liu 0003 |
IJCAI | 6 |
| 2023 | FedSDG-FS: Efficient and Secure Feature Selection for Vertical Federated LearningabstractVertical Federated Learning (VFL) enables multiple data owners, each holding a different subset of features about largely overlapping sets of data sample(s), to jointly train a useful global model. Feature selection (FS) is important to VFL. It is still an open research problem as existing FS works designed for VFL either assumes prior knowledge on the number of noisy features or prior knowledge on the post-training threshold of useful features to be selected, making them unsuitable for practical applications. To bridge this gap, we propose the Federated Stochastic Dual-Gate based Feature Selection (FedSDG-FS) approach. It consists of a Gaussian stochastic dual-gate to efficiently approximate the probability of a feature being selected, with privacy protection through Partially Homomorphic Encryption without a trusted third-party. To reduce overhead, we propose a feature importance initialization method based on Gini impurity, which can accomplish its goals with only two parameter transmissions between the server and the clients. Extensive experiments on both synthetic and real-world datasets show that FedSDG-FS significantly outperforms existing approaches in terms of achieving accurate selection of high-quality features as well as building global models with improved performance. Anran Li 0001, Hongyi Peng, Lan Zhang 0002, Qing Guo 0005, Han Yu 0001, Yang Liu 0003 |
INFOCOM | 5 |
| 2023 | ALA: Naturalness-aware Adversarial Lightness AttackabstractMost researchers have tried to enhance the robustness of deep neural networks (DNNs) by revealing and repairing the vulnerability of DNNs with specialized adversarial examples. Parts of the attack examples have imperceptible perturbations restricted by Lp norm. However, due to their high-frequency property, the adversarial examples can be defended by denoising methods and are hard to realize in the physical world. To avoid the defects, some works have proposed unrestricted attacks to gain better robustness and practicality. It is disappointing that these examples usually look unnatural and can alert the guards. In this paper, we propose Adversarial Lightness Attack (ALA), a white-box unrestricted adversarial attack that focuses on modifying the lightness of the images. The shape and color of the samples, which are crucial to human perception, are barely influenced. To obtain adversarial examples with a high attack success rate, we propose unconstrained enhancement in terms of the light and shade relationship in images. To enhance the naturalness of images, we craft the naturalness-aware regularization according to the range and distribution of light. The effectiveness of ALA is verified on two popular datasets for different tasks (i.e., ImageNet for image classification and Places-365 for scene recognition). Yihao Huang 0001, Liangru Sun, Qing Guo 0005, Felix Juefei-Xu, Jiayi Zhu 0002, Jincao Feng, Yang Liu 0003, Geguang Pu |
ACM Multimedia | 3 |
| 2023 | DistXplore: Distribution-Guided Testing for Evaluating and Enhancing Deep Learning SystemsabstractDeep learning (DL) models are trained on sampled data, where the distribution of training data differs from that of real-world data (i.e., the distribution shift), which reduces the model's robustness. Various testing techniques have been proposed, including distribution-unaware and distribution-aware methods. However, distribution-unaware testing lacks effectiveness by not explicitly considering the distribution of test cases and may generate redundant errors (within same distribution). Distribution-aware testing techniques primarily focus on generating test cases that follow the training distribution, missing out-of-distribution data that may also be valid and should be considered in the testing process. In this paper, we propose a novel distribution-guided approach for generating valid test cases with diverse distributions, which can better evaluate the model's robustness (i.e., generating hard-to-detect errors) and enhance the model's robustness (i.e., enriching training data). Unlike existing testing techniques that optimize individual test cases, DistXplore optimizes test suites that represent specific distributions. To evaluate and enhance the model's robustness, we design two metrics: distribution difference, which maximizes the similarity in distribution between two different classes of data to generate hard-to-detect errors, and distribution diversity, which increase the distribution diversity of generated test cases for enhancing the model's robustness. To evaluate the effectiveness of DistXplore in model evaluation and enhancement, we compare DistXplore with 14 state-of-the-art baselines on 10 models across 4 datasets. The evaluation results show that DisXplore not only detects a larger number of errors (e.g., 2×+ on average). Furthermore, DistXplore achieves a higher improvement in empirical robustness (e.g., 5.2% more accuracy improvement than the baselines on average). Longtian Wang, Xiaofei Xie, Xiaoning Du 0001, Qing Guo 0005, Chao Shen 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2023 | FedCSS: Joint Client-and-Sample Selection for Hard Sample-Aware Noise-Robust Federated LearningabstractFederated Learning (FL) enables a large number of data owners (a.k.a. FL clients) to jointly train a machine learning model without disclosing private local data. The importance of local data samples to the FL model vary widely. This is exacerbated by the presence of noisy data, which exhibit large losses similar to important (hard) samples. Currently, there lacks an FL approach that can effectively distinguish hard samples (which are beneficial) from noisy samples (which are harmful). To bridge this gap, we propose the Federated Client and Sample Selection (FedCSS) approach. It is a bilevel optimization approach for FL client-and-sample selection to achieve hard sample-aware noise-robust learning in a privacy preserving manner. It performs meta-learning based online approximation to iteratively update global FL models, select the most positively influential samples and deal with training data noise. Theoretical analysis shows that it is guaranteed to converge in an efficient manner. Experimental comparison against six state-of-the-art baselines on five real-world datasets in the presence of data noise and heterogeneity shows that it achieves up to 26.4% higher test accuracy, while saving communication and computation costs by at least 41.5% and 1.2%, respectively. Anran Li 0001, Jiabao Guo, Hongyi Peng, Qing Guo 0005, Han Yu 0001 |
Proc. ACM Manag. Data | 5 |
| 2023 | Coarse-to-Fine Task-Driven Inpainting for Geoscience ImagesabstractThe processing and recognition of geoscience images have wide applications. Most of existing researches focus on understanding the high-quality geoscience images by assuming that all the images are clear. However, in many real-world cases, the geoscience images might contain occlusions during the image acquisition. This problem actually implies the image inpainting problem in computer vision and multimedia. As far as we know, all the existing image inpainting algorithms learn to repair the occluded regions for a better visualization quality, they are excellent for natural images but not good enough for geoscience images, and they never consider the following gescience task when developing inpainting methods. This paper aims to repair the occluded regions for a better geoscience task performance and advanced visualization quality simultaneously, without changing the current deployed deep learning based geoscience models. Because of the complex context of geoscience images, we propose a coarse-to-fine encoder-decoder network with the help of designed coarse-to-fine adversarial context discriminators to reconstruct the occluded image regions. Due to the limited data of geoscience images, we propose a MaskMix based data augmentation method, which augments inpainting masks instead of augmenting original images, to exploit the limited geoscience image data. The experimental results on three public geoscience datasets for remote sensing scene recognition, cross-view geolocation and semantic segmentation tasks respectively show the effectiveness and accuracy of the proposed method. The code is available at:https://github.com/HMS97/Task-driven-Inpainting. Huiming Sun, Jin Ma 0005, Qing Guo 0005, Qin Zou 0001, Shaoyue Song, Yuewei Lin, Hongkai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Structure-Informed Shadow Removal NetworksabstractExisting deep learning-based shadow removal methods still produce images with shadow remnants. These shadow remnants typically exist in homogeneous regions with low-intensity values, making them untraceable in the existing image-to-image mapping paradigm. We observe that shadows mainly degrade images at the image-structure level (in which humans perceive object shapes and continuous colors). Hence, in this paper, we propose to remove shadows at the image structure level. Based on this idea, we propose a novel structure-informed shadow removal network (StructNet) to leverage the image-structure information to address the shadow remnant problem. Specifically, StructNet first reconstructs the structure information of the input image without shadows and then uses the restored shadow-free structure prior to guiding the image-level shadow removal. StructNet contains two main novel modules: 1) a mask-guided shadow-free extraction (MSFE) module to extract image structural features in a non-shadow-to-shadow directional manner; and 2) a multi-scale feature & residual aggregation (MFRA) module to leverage the shadow-free structure information to regularize feature consistency. In addition, we also propose to extend StructNet to exploit multi-level structure information (MStructNet), to further boost the shadow removal performance with minimum computational overheads. Extensive experiments on three shadow removal benchmarks demonstrate that our method outperforms existing shadow removal methods, and our StructNet can be integrated with existing methods to improve them further. Yuhao Liu 0001, Qing Guo 0005, Lan Fu, Zhanghan Ke, Ke Xu 0010, Wei Feng 0005, Ivor W. Tsang, Rynson W. H. Lau |
IEEE Trans. Image Process. | 2 |
| 2023 | ArchRepair: Block-Level Architecture-Oriented Repairing for Deep Neural NetworksabstractOver the past few years, deep neural networks (DNNs) have achieved tremendous success and have been continuously applied in many application domains. However, during the practical deployment in industrial tasks, DNNs are found to be erroneous-prone due to various reasons such as overfitting and lacking of robustness to real-world corruptions during practical usage. To address these challenges, many recent attempts have been made to repair DNNs for version updates under practical operational contexts by updating weights (i.e., network parameters) through retraining, fine-tuning, or direct weight fixing at a neural level. Nevertheless, existing solutions often neglect the effects of neural network architecture and weight relationships across neurons and layers. In this work, as the first attempt, we initiate to repair DNNs by jointly optimizing the architecture and weights at a higher (i.e., block level). We first perform empirical studies to investigate the limitation of whole network-level and layer-level repairing, which motivates us to explore a novel repairing direction for DNN repair at the block level. To this end, we need to further consider techniques to address two key technical challenges, i.e., block localization , where we should localize the targeted block that we need to fix; and how to perform joint architecture and weight repairing . Specifically, we first propose adversarial-aware spectrum analysis for vulnerable block localization that considers the neurons’ status and weights’ gradients in blocks during the forward and backward processes, which enables more accurate candidate block localization for repairing even under a few examples. Then, we further propose the architecture-oriented search-based repairing that relaxes the targeted block to a continuous repairing search space at higher deep feature levels. By jointly optimizing the architecture and weights in that space, we can identify a much better block architecture. We implement our proposed repairing techniques as a tool, named ArchRepair , and conduct extensive experiments to validate the proposed method. The results show that our method can not only repair but also enhance accuracy and robustness, outperforming the state-of-the-art DNN repair techniques. Hua Qi, Zhijie Wang 0014, Qing Guo 0005, Jianlang Chen, Felix Juefei-Xu, Fuyuan Zhang, Lei Ma 0003, Jianjun Zhao 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Can You Spot the Chameleon? Adversarially Camouflaging Images from Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods. In this paper, we address this problem from the perspective of adversarial attacks and identify a novel task: adversarial co-saliency attack. Specially, given an image selected from a group of images containing some common and salient objects, we aim to generate an adversarial version that can mislead CoSOD methods to predict incorrect co-salient regions. Note that, compared with general white-box adversarial attacks for classification, this new task faces two additional challenges: (1) low success rate due to the diverse appearance of images in the group; (2) low transferability across CoSOD methods due to the considerable difference between CoSOD pipelines. To address these challenges, we propose the very first blackbox joint adversarial exposure and noise attack (Jadena), where we jointly and locally tune the exposure and additive perturbations of the image according to a newly designed high-feature-level contrast-sensitive loss function. Our method, without any information on the state-of-the-art CoSOD methods, leads to significant performance degradation on various co-saliency detection datasets and makes the co-salient objects undetectable. This can have strong practical benefits in properly securing the large number of personal photos currently shared on the Internet. Moreover, our method is potential to be utilized as a metric for evaluating the robustness of CoSOD methods. Ruijun Gao, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 2 |
| 2022 | MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image InpaintingabstractAlthough achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applications. Image-level predictive filtering is a widely used restoration technique by predicting suitable kernels adaptively according to different input scenes. Inspired by this inherent advantage, we explore the possibility of addressing image inpainting as a filtering task. To this end, we first study the advantages and challenges of the image-level predictive filtering for inpainting: the method can preserve local structures and avoid artifacts but fails to fill large missing areas. Then, we propose the semantic filtering by conducting filtering on deep feature level, which fills the missing semantic information but fails to recover the details. To address the issues while adopting the respective advantages, we propose a novel filtering technique, i.e., Multi-level Interactive Siamese Filtering (MISF) containing two branches: kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). These two branches are interactively linked: SIFB provides multi-level features for KPB while KPB predicts dynamic kernels for SIFB. As a result, the final method takes the advantage of effective semantic & image-level filling for high-fidelity inpainting. Moreover, we discuss the relationship between MISF and the naive encoder-decoder-based inpainting, inferring that MISF provides novel dynamic convolutional operations to enhance the high generalization capability across scenes. We validate our method on three challenging datasets, i.e., Dunhuang, Places2, and CelebA. Our method outperforms state-of-the-art baselines on four metrics, i.e.,$L_{1}$, PSNR, SSIM, and LPIPS. Qing Guo 0005, Di Lin 0002, Ping Li 0016, Wei Feng 0005, Song Wang 0002 |
CVPR | 2 |
| 2022 | Optimistic Exploration Based on Categorical-DQN for Cooperative Markov Games
Chengwei Zhang 0001, Qing Guo 0005, Kangjie Zheng, Wanqing Fang, Xintian Zhao |
DAI | 3 |
| 2022 | A3GAN: Attribute-Aware Anonymization Networks for Face De-identificationabstractFace de-identification (De-ID) removes face identity information in face images to avoid personal privacy leakage. Existing face De-ID breaks the raw identity by cutting out the face regions and recovering the corrupted regions via deep generators, which inevitably affect the generation quality and cannot control generation results according to subsequent intelligent tasks (eg., facial expression recognition). In this work, for the first attempt, we think the face De-ID from the perspective of attribute editing and propose an attribute-aware anonymization network (A3GAN) by formulating face De-ID as a joint task of semantic suppression and controllable attribute injection. Intuitively, the semantic suppression removes the identity-sensitive information in embeddings while the controllable attribute injection automatically edits the raw face along the attributes that benefit De-ID. To this end, we first design a multi-scale semantic suppression network with a novel suppressive convolution unit (SCU), which can remove the face identity along multi-level deep features progressively. Then, we propose an attribute-aware injective network (AINet) that can generate De-ID-sensitive attributes in a controllable way (i.e., specifying which attributes can be changed and which cannot) and inject them into the latent code of the raw face. Moreover, to enable effective training, we design a new anonymization loss to let the injected attributes shift far away from the original ones. We perform comprehensive experiments on four datasets covering four different intelligent tasks including face verification, face detection, facial expression recognition, and fatigue detection, all of which demonstrate the superiority of our face De-ID over state-of-the-art methods. Liming Zhai, Qing Guo 0005, Xiaofei Xie, Lei Ma 0003, Yi Estelle Wang, Yang Liu 0003 |
ACM Multimedia | 2 |
| 2022 | Generative Status Estimation and Information Decoupling for Image Rain RemovalabstractImage rain removal requires the accurate separation between the pixels of the rain streaks and object textures. But the confusing appearances of rains and objects lead to the misunderstanding of pixels, thus remaining the rain streaks or missing the object details in the result. In this paper, we propose SEIDNet equipped with the generative Status Estimation and Information Decoupling for rain removal. In the status estimation, we embed the pixel-wise statuses into the status space, where each status indicates a pixel of the rain or object. The status space allows sampling multiple statuses for a pixel, thus capturing the confusing rain or object. In the information decoupling, we respect the pixel-wise statuses, decoupling the appearance information of rain and object from the pixel. Based on the decoupled information, we construct the kernel space, where multiple kernels are sampled for the pixel to remove the rain and recover the object appearance. We evaluate SEIDNet on the public datasets, achieving state-of-the-art performances of image rain removal. The experimental results also demonstrate the generalization of SEIDNet, which can be easily extended to achieve state-of-the-art performances on other image restoration tasks (e.g., snow, haze, and shadow removal). Di Lin 0002, Xin Wang 0118, Miaohui Wang, Wuyuan Xie, Qing Guo 0005, Ping Li 0016 |
NeurIPS | 8 |
| 2022 | Countering Malicious DeepFakes: Survey, Battleground, and Horizon
Felix Juefei-Xu, Run Wang 0001, Yihao Huang 0001, Qing Guo 0005, Lei Ma 0003, Yang Liu 0003 |
Int. J. Comput. Vis. | 4 |
| 2022 | DARTSRepair: Core-failure-set guided DARTS for network robustness to common corruptions
Xuhong Ren, Jianlang Chen, Felix Juefei-Xu, Wanli Xue, Qing Guo 0005, Lei Ma 0003, Jianjun Zhao 0001, Shengyong Chen |
Pattern Recognit. | 5 |
| 2022 | Let There Be Light: Improved Traffic Surveillance via Detail Preserving Night-to-Day TransferabstractIn recent years, image and video surveillance have made considerable progresses to the Intelligent Transportation Systems (ITS) with the help of deep Convolutional Neural Networks (CNNs). As one of the state-of-the-art perception approaches, detecting the interested objects in each frame of video surveillance is widely desired by ITS. Currently, object detection shows remarkable efficiency and reliability in standard scenarios such as daytime scenes with favorable illumination conditions. However, in face of adverse conditions such as the nighttime, object detection loses its accuracy significantly. One of the main causes of the problem is the lack of sufficient annotated detection datasets of nighttime scenes. In this paper, we propose a framework to alleviate the accuracy decline when object detection is taken to adverse conditions by using image translation method. We propose to utilize style translation based StyleMix method to acquire pairs of day time image and nighttime image as training data for following nighttime to daytime image translation. To alleviate the detail corruptions caused by Generative Adversarial Networks (GANs), we propose to utilize Kernel Prediction Network (KPN) based method to refine the nighttime to daytime image translation. The KPN network is trained with object detection task together to adapt the trained daytime model to nighttime vehicle detection directly. Experiments on vehicle detection verified the accuracy and effectiveness of the proposed approach. Lan Fu, Hongkai Yu, Felix Juefei-Xu, Qing Guo 0005, Song Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Object-Aware Ghost Identification and Elimination for Dynamic Scene MosaicabstractComposite ghost is a common phenomenon that widely exists in dynamic scene image mosaic and significantly affects the naturalness of mosaic. To remove the ghost effectively and produce visually natural mosaic, we propose a novel image mosaic method by jointly identifying composite ghost and eliminating ghost regions without distorting, splitting, and duplicating objects. Specifically, our main contributions are three-fold:First, we propose themotion-awarecomposite ghost identification to localize the potential composite ghosts in the mosaic region (i.e., overlapping area between two images to be stitched) by detecting the salient-moving objects in two stitched images.Second, we design theobject-awarealternative region selection strategy to produce ghostless regions that can replace the localized composite ghosts while avoiding object distortion, object separation, and object repetition.Third, we realize theimage interpolation-basedcomposite ghost elimination that can generate natural stitched image by eliminating the composite ghost of the initial blending result with the selected image source. We validate the proposed method on challenging datasets and show that our method outperform the state-of-the-art methods. Zhe Zhang 0039, Xuhong Ren, Wanli Xue, Chengwei Zhang 0001, Qing Guo 0005, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | FakeLocator: Robust Localization of GAN-Based Face ManipulationsabstractFull face synthesis and partial face manipulation by virtue of the generative adversarial networks (GANs) and its variants have raised wide public concerns. In the multi-media forensics area, detecting and ultimately locating the image forgery has become an imperative task. In this work, we investigate the architecture of existing GAN-based face manipulation methods and observe that the imperfection of upsampling methods therewithin could be served as an important asset for GAN-synthesized fake image detection and forgery localization. Based on this basic observation, we have proposed a novel approach, termedFakeLocator, to obtain high localization accuracy, at full resolution, on manipulated facial images. To the best of our knowledge, this is the very first attempt to solve the GAN-based fake localization problem with a gray-scale fakeness map that preserves more information of fake regions. To improve the universality ofFakeLocatoracross multifarious facial attributes, we introduce an attention mechanism to guide the training of the model. To improve the universality ofFakeLocatoracross different DeepFake methods, we propose partial data augmentation and single sample clustering on the training images. Experimental results on popular FaceForensics++, DFFD datasets and seven different state-of-the-art GAN-based face generation methods have shown the effectiveness of our method. Compared with the baselines, our method performs better on various metrics. Moreover, the proposed method is robust against various real-world facial image degradations such as JPEG compression, low-resolution, noise, and blur. Yihao Huang 0001, Felix Juefei-Xu, Qing Guo 0005, Yang Liu 0003, Geguang Pu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Pasadena: Perceptually Aware and Stealthy Adversarial Denoise AttackabstractImage denoising can remove natural noise that widely exists in images captured by multimedia devices due to low-quality imaging sensors, unstable image transmission processes, or low light conditions. Recent works also find that image denoising benefits the high-level vision tasks,e.g., image classification. In this work, we try to challenge this common sense and explore a totally new problem,i.e., whether the image denoising can be given the capability of fooling the state-of-the-art deep neural networks (DNNs) while enhancing the image quality. To this end, we initiate the very first attempt to study this problem from the perspective of adversarial attack and propose theadversarial denoise attack. More specifically, our main contributions are three-fold:First, we identify a new task that stealthily embeds attacks inside the image denoising module widely deployed in multimedia devices as an image post-processing operation to simultaneously enhance the visual image quality and fool DNNs.Second, we formulate this new task as a kernel prediction problem for image filtering and propose theadversarial-denoising kernel predictionthat can produce adversarial-noiseless kernels for effective denoising and adversarial attacking simultaneously.Third, we implement an adaptiveperceptual region localizationto identify semantic-related vulnerability regions with which the attack can be more effective while not doing too much harm to the denoising. We name the proposed method asPasadena(Perceptually Aware and Stealthy Adversarial DENoise Attack) and validate our method on the NeurIPS’17 adversarial competition dataset, CVPR2021-AIC-VI: unrestricted adversarial attacks on ImageNet, and Tiny-ImageNet-C dataset. The comprehensive evaluation and analysis demonstrate that our method not only realizes denoising but also achieves a significantly higher success rate and transferability over state-of-the-art attacks. Yupeng Cheng, Qing Guo 0005, Felix Juefei-Xu, Shangwei Lin 0001, Wei Feng 0005, Weisi Lin, Yang Liu 0003 |
IEEE Trans. Multim. | 2 |
| 2022 | NPC: Neuron Path Coverage via Characterizing Decision Logic of Deep Neural NetworksabstractDeep learning has recently been widely applied to many applications across different domains, e.g., image classification and audio recognition. However, the quality of Deep Neural Networks (DNNs) still raises concerns in the practical operational environment, which calls for systematic testing, especially in safety-critical scenarios. Inspired by software testing, a number of structural coverage criteria are designed and proposed to measure the test adequacy of DNNs. However, due to the blackbox nature of DNN, the existing structural coverage criteria are difficult to interpret, making it hard to understand the underlying principles of these criteria. The relationship between the structural coverage and the decision logic of DNNs is unknown. Moreover, recent studies have further revealed the non-existence of correlation between the structural coverage and DNN defect detection, which further posts concerns on what a suitable DNN testing criterion should be. In this article, we propose the interpretable coverage criteria through constructing the decision structure of a DNN. Mirroring the control flow graph of the traditional program, we first extract a decision graph from a DNN based on its interpretation, where a path of the decision graph represents a decision logic of the DNN. Based on the control flow and data flow of the decision graph, we propose two variants of path coverage to measure the adequacy of the test cases in exercising the decision logic. The higher the path coverage, the more diverse decision logic the DNN is expected to be explored. Our large-scale evaluation results demonstrate that: The path in the decision graph is effective in characterizing the decision of the DNN, and the proposed coverage criteria are also sensitive with errors, including natural errors and adversarial examples, and strongly correlate with the output impartiality. Xiaofei Xie, Tianlin Li, Jian Wang 0067, Lei Ma 0003, Qing Guo 0005, Felix Juefei-Xu, Yang Liu 0003 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2022 | DeepRepair: Style-Guided Repairing for Deep Neural Networks in the Real-World Operational EnvironmentabstractDeep neural networks (DNNs) are continuously expanding their application to various domains due to their high performance. Nevertheless, a well-trained DNN after deployment could oftentimes raise errors during practical use in the operational environment due to the mismatching between distributions of the training dataset and the potential unknown noise factors in the operational environment, e.g., weather, blur, noise, etc. Hence, it poses a rather important problem for the DNNs’ real-world applications: how to repair the deployed DNNs for correcting the failure samples under the deployed operational environment while not harming their capability of handling normal or clean data with limited failure samples we can collect. In this article, we propose astyle-guided data augmentation for repairing DNN in the operational environment, which learns and introduces the unknown failure patterns within the failure samples into the training data via the style transfer. Moreover, we further propose theclustering-based failure data generationfor much more effective style-guided data augmentation. We conduct a large-scale evaluation with 15 degradation factors that may happen in the real world and compare with four state-of-the-art data augmentation methods and two DNN repairing methods. Our technique successfully repairs three convolutional neural networks and two recurrent neural networks with averaging 62.88% and 39.02% accuracy enhancements on the 15 failure patterns, respectively, achieving higher repairing performance than state-of-the-art repairing methods on the most failure patterns with even better accuracy on clean datasets. Hua Qi, Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Jianjun Zhao 0001 |
IEEE Trans. Reliab. | 3 |
| 2021 | EfficientDeRain: Learning Pixel-wise Dilation Filtering for High-Efficiency Single-Image DerainingabstractSingle-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however, significantly affects these methods' efficiency and effectiveness for many efficiency-critical applications. To fill this gap, in this paper, we regard the single-image deraining as a general image-enhancing problem and originally propose a model-free deraining method, i.e., EfficientDeRain, which is able to process a rainy image within 10 ms (i.e., around 6 ms on average), over 80 times faster than the state-of-the-art method (i.e., RCDNet), while achieving similar de-rain effects. We first propose novel pixel-wise dilation filtering. In particular, a rainy image is filtered with the pixel-wise kernels estimated from a kernel prediction network, by which suitable multi-scale kernels for each pixel can be efficiently predicted. Then, to eliminate the gap between synthetic and real data, we further propose an effective data augmentation method (i.e., RainMix) that helps to train the network for handling real rainy images. We perform a comprehensive evaluation on both synthetic and real-world rainy datasets to demonstrate the effectiveness and efficiency of our method. We release the model and code in https://github.com/tsingqguo/efficientderain.git. Qing Guo 0005, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Xiaofei Xie, Wei Feng 0005, Yang Liu 0003, Jianjun Zhao 0001 |
AAAI | 1 |
| 2021 | Auto-Exposure Fusion for Single-Image Shadow RemovalabstractShadow removal is still a challenging task due to its inherent background-dependent1and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this task by formulating it as an exposure fusion problem to address the challenges. Intuitively, we first estimate multiple over-exposure images w.r.t. the input image to let the shadow regions in these images have the same color with shadow-free areas in the input image. Then, we fuse the original input with the over-exposure images to generate the final shadow-free counterpart. Nevertheless, the spatial-variant property of the shadow requires the fusion to be sufficiently ‘smart’, that is, it should automatically select proper over-exposure pixels from different images to make the final output natural. To address this challenge, we propose the shadow-aware FusionNet that takes the shadow image as input to generate fusion weight maps across all the over-exposure images. Moreover, we propose the boundary-aware RefineNet to eliminate the remaining shadow trace further. We conduct extensive experiments on the ISTD, ISTD+, and SRD datasets to validate our method’s effectiveness and show better performance in shadow regions and comparable performance in non-shadow regions over the state-of-the-art methods. We release the code in https://github.com/tsingqguo/exposure-fusion-shadow-removal. Lan Fu, Changqing Zhou, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 3 |
| 2021 | Learning to Adversarially Blur Visual Object TrackingabstractMotion blur caused by the moving of the object or camera during the exposure can be a key challenge for visual object tracking, affecting tracking accuracy significantly. In this work, we explore the robustness of visual object trackers against motion blur from a new angle, i.e., adversarial blur attack (ABA). Our main objective is to online transfer input frames to their natural motion-blurred counterparts while misleading the state-of-the-art trackers during the tracking process. To this end, we first design the motion blur synthesizing method for visual tracking based on the generation principle of motion blur, considering the motion information and the light accumulation process. With this synthetic method, we propose optimization-based ABA (OP-ABA) by iteratively optimizing an adversarial objective function against the tracking w.r.t. the motion and light accumulation parameters. The OP-ABA is able to produce natural adversarial examples but the iteration can cause heavy time cost, making it unsuitable for attacking real-time trackers. To alleviate this issue, we further propose one-step ABA (OS-ABA) where we design and train a joint adversarial motion and accumulation predictive network (JAMANet) with the guidance of OP-ABA, which is able to efficiently estimate the adversarial motion and accumulation parameters in a one-step way. The experiments on four popular datasets (e.g., OTB100, VOT2018, UAV123, and LaSOT) demonstrate that our methods are able to cause significant accuracy drops on four state-of-the-art trackers with high transferability. Please find the source code at https://github.com/tsingqguo/ABA Qing Guo 0005, Ziyi Cheng, Felix Juefei-Xu, Lei Ma 0003, Xiaofei Xie, Yang Liu 0003, Jianjun Zhao 0001 |
ICCV | 1 |
| 2021 | Deepmix: Online Auto Data Augmentation for Robust Visual Object TrackingabstractOnline updating of the object model via samples from historical frames is of great importance for accurate visual object tracking. Recent works mainly focus on constructing effective and efficient updating methods while neglecting the training samples for learning discriminative object models, which is also a key part of a learning problem. In this paper, we propose the DeepMix that takes historical samples’ embeddings as input and generates augmented embeddings online, enhancing the state-of-the-art online learning methods for visual object tracking. More specifically, we first propose the online data augmentation for tracking that online augments the historical samples through object-aware filtering. Then, we propose MixNet which is an offline trained network for performing online data augmentation within one-step, enhancing the tracking accuracy while preserving high speeds of the state-of-the-art online learning methods. The extensive experiments on three different tracking frameworks, i.e., DiMP, DSiam, and SiamRPN++, and three large-scale and challenging datasets, i.e., OTB-2015, LaSOT, and VOT, demonstrate the effectiveness and advantages of the proposed method. Ziyi Cheng, Xuhong Ren, Felix Juefei-Xu, Wanli Xue, Qing Guo 0005, Lei Ma 0003, Jianjun Zhao 0001 |
ICME | 5 |
| 2021 | Bias Field Poses a Threat to DNN-Based X-Ray RecognitionabstractChest X-ray plays a key role in screening and diagnosis of many lung diseases including the COVID-19. Many works construct deep neural networks (DNNs) for chest X-ray images to realize automated and efficient diagnosis of lung diseases. However, bias field caused by the improper medical image acquisition process widely exists in the chest X-ray images while the robustness of DNNs to the bias field is rarely explored, posing a threat to the X-ray-based automated diagnosis system. In this paper, we study this problem based on the adversarial attack and propose a brand new attack, i.e., adversarial bias field attack where the bias field instead of the additive noise works as the adversarial perturbations for fooling DNNs. This novel attack poses a key problem: how to locally tune the bias field to realize high attack success rate while maintaining its spatial smoothness to guarantee high realisticity. These two goals contradict each other and thus has made the attack significantly challenging. To overcome this challenge, we propose the adversarial-smooth bias field attack that can locally tune the bias field with joint smooth & adversarial constraints. As a result, the adversarial X-ray images can not only fool the DNNs effectively but also retain very high level of realisticity. We validate our method on real chest X-ray datasets with powerful DNNs, e.g., ResNet50, DenseNet121, and MobileNet, and show different properties to the state-of-the-art attacks in both image realisticity and attack transferability. Our method reveals the potential threat to the DNN-based X-ray automated diagnosis and can definitely benefit the development of bias-field-robust automated diagnosis system. Binyu Tian, Qing Guo 0005, Felix Juefei-Xu, Wen Le Chan, Yupeng Cheng, Xiaohong Li 0001, Xiaofei Xie, Shengchao Qin |
ICME | 2 |
| 2021 | AVA: Adversarial Vignetting Attack against Visual RecognitionabstractVignetting is an inherent imaging phenomenon within almost all optical systems, showing as a radial intensity darkening toward the corners of an image. Since it is a common effect for photography and usually appears as a slight intensity variation, people usually regard it as a part of a photo and would not even want to post-process it. Due to this natural advantage, in this work, we study the vignetting from a new viewpoint, i.e., adversarial vignetting attack (AVA), which aims to embed intentionally misleading information into the vignetting and produce a natural adversarial example without noise patterns. This example can fool the state-of-the-art deep convolutional neural networks (CNNs) but is imperceptible to human. To this end, we first propose the radial-isotropic adversarial vignetting attack (RI-AVA) based on the physical model of vignetting, where the physical parameters (e.g., illumination factor and focal length) are tuned through the guidance of target CNN models. To achieve higher transferability across different CNNs, we further propose radial-anisotropic adversarial vignetting attack (RA-AVA) by allowing the effective regions of vignetting to be radial-anisotropic and shape-free. Moreover, we propose the geometry-aware level-set optimization method to solve the adversarial vignetting regions and physical parameters jointly. We validate the proposed methods on three popular datasets, i.e., DEV, CIFAR10, and Tiny ImageNet, by attacking four CNNs, e.g., ResNet50, EfficientNet-B0, DenseNet121, and MobileNet-V2, demonstrating the advantages of our methods over baseline methods on both transferability and image quality. Binyu Tian, Felix Juefei-Xu, Qing Guo 0005, Xiaofei Xie, Xiaohong Li 0001, Yang Liu 0003 |
IJCAI | 3 |
| 2021 | AdvFilter: Predictive Perturbation-aware Filtering against Adversarial Attack via Multi-domain LearningabstractHigh-level representation-guided pixel denoising and adversarial training are independent solutions to enhance the robustness of CNNs against adversarial attacks by pre-processing input data and re-training models, respectively. Most recently, adversarial training techniques have been widely studied and improved while the pixel denoising-based method is getting less attractive. However, it is still questionable whether there exists a more advanced pixel denoising-based method and whether the combination of the two solutions benefits each other. To this end, we first comprehensively investigate two kinds of pixel denoising methods for adversarial robustness enhancement (i.e., existing additive-based and unexplored filtering-based methods) under the loss functions of image-level and semantic-level, respectively, showing that pixel-wise filtering can obtain much higher image quality (e.g., higher PSNR) as well as higher robustness (e.g., higher accuracy on adversarial examples) than existing pixel-wise additive-based method. However, we also observe that the robustness results of the filtering-based method rely on the perturbation amplitude of adversarial examples used for training. To address this problem, we propose predictive perturbation-aware & pixel-wise filtering, where dual-perturbation filtering and an uncertainty-aware fusion module are designed and employed to automatically perceive the perturbation amplitude during the training and testing process. The method is termed as AdvFilter. Moreover, we combine adversarial pixel denoising methods with three adversarial training-based methods, hinting that considering data and models jointly is able to achieve more robust CNNs. The experiments conduct on NeurIPS-2017DEV, SVHN and CIFAR10 datasets and show advantages over enhancing CNNs' robustness, high generalization to different models and noise levels. Yihao Huang 0001, Qing Guo 0005, Felix Juefei-Xu, Lei Ma 0003, Weikai Miao, Yang Liu 0003, Geguang Pu |
ACM Multimedia | 2 |
| 2021 | JPGNet: Joint Predictive Filtering and Generative Network for Image InpaintingabstractImage inpainting aims to restore the missing regions of corrupted images and make the recovery result identical to the originally complete image, which is different from the common generative task emphasizing the naturalness or realism of generated images. Nevertheless, existing works usually regard it as a pure generation problem and employ cutting-edge deep generative techniques to address it. The generative networks can fill the main missing parts with realistic contents but usually distort the local structures or introduce obvious artifacts. In this paper, for the first time, we formulate image inpainting as a mix of two problems, i.e., predictive filtering and deep generation. Predictive filtering is good at preserving local structures and removing artifacts but falls short to complete the large missing regions. The deep generative network can fill the numerous missing pixels based on the understanding of the whole scene but hardly restores the details identical to the original ones. To make use of their respective advantages, we propose the joint predictive filtering and generative network (JPGNet) that contains three branches: predictive filtering & uncertainty network (PFUNet), deep generative network, and uncertainty-aware fusion network (UAFNet). The PFUNet can adaptively predict pixel-wise kernels for filtering-based inpainting according to the input image and output an uncertainty map. This map indicates the pixels should be processed by filtering or generative networks, which is further fed to the UAFNet for a smart combination between filtering and generative results. Note that, our method as a novel framework for the image inpainting problem can benefit any existing generation-based methods. We validate our method on three public datasets, i.e., Dunhuang, Places2, and CelebA, and demonstrate that our method can enhance three state-of-the-art generative methods (i.e., StructFlow, EdgeConnect, and RFRNet) significantly with slightly extra time costs. We have released the code at https://github.com/tsingqguo/jpgnet. Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Yang Liu 0003, Song Wang 0002 |
ACM Multimedia | 1 |
| 2021 | Exploring the Effects of Blur and Deblurring to Visual Object TrackingabstractThe existence of motion blur can inevitably influence the performance of visual object tracking. However, in contrast to the rapid development of visual trackers, the quantitative effects of increasing levels of motion blur on the performance of visual trackers still remain unstudied. Meanwhile, although image-deblurring can produce visually sharp videos for pleasant visual perception, it is also unknown whether visual object tracking can benefit from image deblurring or not. In this paper, we present a Blurred Video Tracking (BVT) benchmark to address these two problems, which contains a large variety of videos with different levels of motion blurs, as well as ground-truth tracking results. To explore the effects of blur and deblurring to visual object tracking, we extensively evaluate 25 trackers on the proposed BVT benchmark and obtain several new interesting findings. Specifically, we find that light motion blur may improve the accuracy of many trackers, but heavy blur usually hurts the tracking performance. We also observe that image deblurring is helpful to improve tracking accuracy on heavily-blurred videos but hurts the performance of lightly-blurred videos. According to these observations, we then propose a new general GAN-based scheme to improve a tracker's robustness to motion blur. In this scheme, a fine-tuned discriminator can effectively serve as an adaptive blur assessor to enable selective frames deblurring during the tracking process. We use this scheme to successfully improve the accuracy of 6 state-of-the-art trackers on motion-blurred videos. Qing Guo 0005, Wei Feng 0005, Ruijun Gao, Yang Liu 0003, Song Wang 0002 |
IEEE Trans. Image Process. | 1 |
| 2020 | An Attentional Recurrent Neural Network for Personalized Next Location RecommendationabstractMost existing studies on next location recommendation propose to model the sequential regularity of check-in sequences, but suffer from the severe data sparsity issue where most locations have fewer than five following locations. To this end, we propose an Attentional Recurrent Neural Network (ARNN) to jointly model both the sequential regularity and transition regularities of similar locations (neighbors). In particular, we first design a meta-path based random walk over a novel knowledge graph to discover location neighbors based on heterogeneous factors. A recurrent neural network is then adopted to model the sequential regularity by capturing various contexts that govern user mobility. Meanwhile, the transition regularities of the discovered neighbors are integrated via the attention mechanism, which seamlessly cooperates with the sequential regularity as a unified recurrent framework. Experimental results on multiple real-world datasets demonstrate that ARNN outperforms state-of-the-art methods. Qing Guo 0005, Zhu Sun 0001, Jie Zhang 0002, Yin Leng Theng |
AAAI | 1 |
| 2020 | SPARK: Spatial-Aware Online Incremental Attack Against Visual Tracking
Qing Guo 0005, Xiaofei Xie, Felix Juefei-Xu, Lei Ma 0003, Zhongguo Li, Wanli Xue, Wei Feng 0005, Yang Liu 0003 |
ECCV (25) | 1 |
| 2020 | FakePolisher: Making DeepFakes More Detection-Evasive by Shallow ReconstructionabstractAt this moment, GAN-based image generation methods are still imperfect, whose upsampling design has limitations in leaving some certain artifact patterns in the synthesized image. Such artifact patterns can be easily exploited (by recent methods) for difference detection of real and GAN-synthesized images. However, the existing detection methods put much emphasis on the artifact patterns, which can become futile if such artifact patterns were reduced. Yihao Huang 0001, Felix Juefei-Xu, Run Wang 0001, Qing Guo 0005, Lei Ma 0003, Xiaofei Xie, Weikai Miao, Yang Liu 0003, Geguang Pu |
ACM Multimedia | 4 |
| 2020 | DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat RhythmsabstractAs the GAN-based face image and video generation techniques, widely known as DeepFakes, have become more and more matured and realistic, there comes a pressing and urgent demand for effective DeepFakes detectors. Motivated by the fact that remote visual photoplethysmography (PPG) is made possible by monitoring the minuscule periodic changes of skin color due to blood pumping through the face, we conjecture that normal heartbeat rhythms found in the real face videos will be disrupted or even entirely broken in a DeepFake video, making it a potentially powerful indicator for DeepFake detection. In this work, we propose DeepRhythm, a DeepFake detection technique that exposes DeepFakes by monitoring the heartbeat rhythms. DeepRhythm utilizes dual-spatial-temporal attention to adapt to dynamically changing face and fake types. Extensive experiments on FaceForensics++ and DFDC-preview datasets have confirmed our conjecture and demonstrated not only the effectiveness, but also the generalization capability of DeepRhythm over different datasets by various DeepFakes generation techniques and multifarious challenging degradations. Hua Qi, Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003, Jianjun Zhao 0001 |
ACM Multimedia | 2 |
| 2020 | Amora: Black-box Adversarial Morphing AttackabstractNowadays, digital facial content manipulation has become ubiquitous and realistic with the success of generative adversarial networks (GANs), making face recognition (FR) systems suffer from unprecedented security concerns. In this paper, we investigate and introduce a new type of adversarial attack to evade FR systems by manipulating facial content, called adversarial morphing attack (a.k.a. Amora). In contrast to adversarial noise attack that perturbs pixel intensity values by adding human-imperceptible noise, our proposed adversarial morphing attack works at the semantic level that perturbs pixels spatially in a coherent manner. To tackle the black-box attack problem, we devise a simple yet effective joint dictionary learning pipeline to obtain a proprietary optical flow field for each attack. Our extensive evaluation on two popular FR systems demonstrates the effectiveness of our adversarial morphing attack at various levels of morphing intensity with smiling facial expression manipulations. Both open-set and closed-set experimental results indicate that a novel black-box adversarial attack based on local deformation is possible, and is vastly different from additive noise attacks. The findings of this work potentially pave a new research direction towards a more thorough understanding and investigation of image-based adversarial attacks and defenses. Run Wang 0001, Felix Juefei-Xu, Qing Guo 0005, Yihao Huang 0001, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003 |
ACM Multimedia | 3 |
| 2020 | DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake VoicesabstractWith the recent advances in voice synthesis, AI-synthesized fake voices are indistinguishable to human ears and widely are applied to produce realistic and natural DeepFakes, exhibiting real threats to our society. However, effective and robust detectors for synthesized fake voices are still in their infancy and are not ready to fully tackle this emerging threat. In this paper, we devise a novel approach, named DeepSonar, based on monitoring neuron behaviors of speaker recognition (SR) system, i.e., a deep neural network (DNN), to discern AI-synthesized fake voices. Layer-wise neuron behaviors provide an important insight to meticulously catch the differences among inputs, which are widely employed for building safety, robust, and interpretable DNNs. In this work, we leverage the power of layer-wise neuron activation patterns with a conjecture that they can capture the subtle differences between real and AI-synthesized fake voices, in providing a cleaner signal to classifiers than raw inputs. Experiments are conducted on three datasets (including commercial products from Google, Baidu, etc) containing both English and Chinese languages to corroborate the high detection rates (98.1% average accuracy) and low false alarm rates (about 2% error rate) of DeepSonar in discerning fake voices. Furthermore, extensive experimental results also demonstrate its robustness against manipulation attacks (e.g., voice conversion and additive real-world noises). Our work further poses a new insight into adopting neuron behaviors for effective and robust AI aided multimedia fakes forensics as an inside-out approach instead of being motivated and swayed by various artifacts introduced in synthesizing fakes. Run Wang 0001, Felix Juefei-Xu, Yihao Huang 0001, Qing Guo 0005, Xiaofei Xie, Lei Ma 0003, Yang Liu 0003 |
ACM Multimedia | 4 |
| 2020 | Watch out! Motion is Blurring the Vision of Your Deep Neural NetworksabstractThe state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, making the study of which greatly important especially for the widely adopted real-time image processing tasks (e.g., object detection, tracking). In this paper, we initiate the first step to comprehensively investigate the potential hazards of blur effect for DNN, caused by object motion. We propose a novel adversarial attack method that can generate visually natural motion-blurred adversarial examples, named motion-based adversarial blur attack (ABBA). To this end, we first formulate the kernel-prediction-based attack where an input image is convolved with kernels in a pixel-wise way, and the misclassification capability is achieved by tuning the kernel weights. To generate visually more natural and plausible examples, we further propose the saliency-regularized adversarial kernel prediction, where the salient region serves as a moving object, and the predicted kernel is regularized to achieve naturally visual effects. Besides, the attack is further enhanced by adaptively tuning the translations of object and background. A comprehensive evaluation on the NeurIPS'17 adversarial competition dataset demonstrates the effectiveness of ABBA by considering various kernel sizes, translations, and regions. The in-depth study further confirms that our method shows a more effective penetrating capability to the state-of-the-art GAN-based deblurring mechanisms compared with other blurring methods. We release the code to \url{https://github.com/tsingqguo/ABBA}. Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Jian Wang 0067, Wei Feng 0005, Yang Liu 0003 |
NeurIPS | 1 |
| 2020 | Selective Spatial Regularization by Reinforcement Learned Decision Making for Object TrackingabstractSpatial regularization (SR) is known as an effective tool to alleviate the boundary effect of correlation filter (CF), a successful visual object tracking scheme, from which a number of state-of-the-art visual object trackers can be stemmed. Nevertheless, SR highly increases the optimization complexity of CF and its target-driven nature makes spatially-regularized CF trackers may easily lose the occluded targets or the targets surrounded by other similar objects. In this paper, we propose selective spatial regularization (SSR) for CF-tracking scheme. It can achieve not only higher accuracy and robustness, but also higher speed compared with spatially-regularized CF trackers. Specifically, rather than simply relying on foreground information, we extend the objective function of CF tracking scheme to learn the target-context-regularized filters using target-context-driven weight maps. We then formulate the online selection of these weight maps as a decision making problem by a Markov Decision Process (MDP), where the learning of weight map selection is equivalent to policy learning of the MDP that is solved by a reinforcement learning strategy. Moreover, by adding a special state, representing not-updating filters, in the MDP, we can learn when to skip unnecessary or erroneous filter updating, thus accelerating the online tracking. Finally, the proposed SSR is used to equip three popular spatially-regularized CF trackers to significantly boost their tracking accuracy, while achieving much faster online tracking speed. Besides, extensive experiments on five benchmarks validate the effectiveness of SSR. Qing Guo 0005, Rui-Ze Han, Wei Feng 0005, Zhihao Chen 0004 |
IEEE Trans. Image Process. | 1 |
| 2019 | Exploiting Side Information for Recommendation
Qing Guo 0005, Zhu Sun 0001, Yin Leng Theng |
ICWE | 1 |
| 2019 | Modeling Heterogeneous Influences for Point-of-Interest Recommendation in Location-Based Social Networks
Qing Guo 0005, Zhu Sun 0001, Jie Zhang 0002, Yin Leng Theng |
ICWE | 1 |
| 2019 | Fast and object-adaptive spatial regularization for correlation filters based tracking
Qing Guo 0005, Wei Feng 0005 |
Neurocomputing | 2 |
| 2019 | Dynamic Saliency-Aware Regularization for Correlation Filter-Based Object TrackingabstractWith a good balance between tracking accuracy and speed, correlation filter (CF) has become one of the best object tracking frameworks, based on which many successful trackers have been developed. Recently, spatially regularized CF tracking (SRDCF) has been developed to remedy the annoying boundary effects of CF tracking, thus further boosting the tracking performance. However, SRDCF uses a fixed spatial regularization map constructed from a loose bounding box and its performance inevitably degrades when the target or background show significant variations, such as object deformation or occlusion. To address this problem, we propose a new dynamic saliency-aware regularized CF tracking (DSAR-CF) scheme. In DSAR-CF, a simple yet effective energy function, which reflects the object saliency and tracking reliability in the spatial-temporal domain, is defined to guide the online updating of the regularization weight map using an efficient level-set algorithm. Extensive experiments validate that the proposed DSAR-CF leads to better performance in terms of accuracy and speed than the original SRDCF. Wei Feng 0005, Rui-Ze Han, Qing Guo 0005, Jianke Zhu, Song Wang 0002 |
IEEE Trans. Image Process. | 3 |
| 2018 | Background-Suppressed Correlation Filters for Visual TrackingabstractCorrelation filters (CF) visual object tracking is a powerful framework, with excellent tracking accuracy and beyond real-time frame rate. Its performance, however, can be severely degraded in cluttered background images. In this paper, we propose background-suppressed correlation filters (BSCF), a better CF tracking scheme, which can significantly improve the reliability and accuracy of CF trackers, without harming their beyond real-time speed. Specifically, we present a unified BSCF object function. We show that both the correlation filters and BS weight map can be efficiently and jointly solved in frequency domain. Extensive experiments on OTB-100 benchmark validate the effectiveness and generality of BS in improving multiple CF trackers with higher accuracy and robustness while maintaining their fast tracking speed. We also show BS boosted CF tracker can achieve comparable accuracy of the state-of-the-art spatially-regularized CF tracker but is 14 times faster. Zhihao Chen 0004, Qing Guo 0005, Wei Feng 0005 |
ICME | 2 |
| 2018 | Content-Related Spatial Regularization for Visual Object TrackingabstractSpatial regularization (SR), being an effective tool to alleviate the boundary effects, can significantly improve the accuracy and robustness of correlation filters (CF) based visual object tracking. The core of SR is a spatially variant weight map that is used to regularize the online learned correlation filters by selecting more meaningful samples. However, most existing trackers apply a data-independent SR weight map. In this paper, we show that a content-related spatial regularization (CRSR) can help to further boost both the tracking accuracy and robustness. Specifically, we present to consider both frame saliency and spatial preference to online generate the CRSR weight map and propose a simple yet effective saliency-embedded CF objective function to simultaneously optimize the filters and CRSR weight map in spatial-temporal domain. Extensive experiments validate that our content-related SR outperforms the classical SR, with higher tracking accuracy and almost two times faster speed. Rui-Ze Han, Qing Guo 0005, Wei Feng 0005 |
ICME | 2 |
| 2018 | Fast Spatially-Regularized Correlation Filters for Visual Object Tracking
Qing Guo 0005, Wei Feng 0005 |
PRICAI (1) | 2 |
| 2018 | Frequency-tuned active contour model
Qing Guo 0005, Shuifa Sun, Xuhong Ren, Fangmin Dong, Bruce Zhi Gao, Wei Feng 0005 |
Neurocomputing | 1 |
| 2017 | Frequency-tuned ACM for biomedical image segmentationabstractBiomedical images are usually corrupted by strong noise and intensity inhomogeneity simultaneously. Existing region-based active contour models (RACMs) easily fail when segmenting such images. In the frequency domain, we propose a generalized RACM that presents a new way to understand the essence of classical RACMs whose segmentation results are determined by a frequency filter to extract the proposed frequency boundary energy. Then, we introduce the difference of Gaussians as the optimal filter to exclude strong noise and intensity inhomogeneity effectively. We show superior performance of the model by comparing with six state-of-the-art methods on challenge biomedical images and segmenting an optical coherence tomography image sequence. Qing Guo 0005, Shuifa Sun, Fangmin Dong, Wei Feng 0005, Bruce Zhi Gao, Siyu Ma |
ICASSP | 1 |
| 2017 | Selective object and context trackingabstractRobust appearance model is significantly important to state-of-the-art trackers. However, such trackers highly rely on the reliability of foreground appearance model. When the foreground is seriously occluded or the scene contains multiple objects with similar appearance, such foundation is destroyed. To extend the ability of trackers to handle these difficulties, we propose selective object and context tracking to locate the target according to the reliability of the foreground appearance model which is determined by two measures about whether the target is occluded or surrounded by similar objects. Extensive experiments show that our method achieves better performance than state-of-the-art trackers on VOT TIR-2015 dataset and is able to track the target even when the foreground appearance is completely unreliable. Ce Zhou, Qing Guo 0005, Wei Feng 0005 |
ICASSP | 2 |
| 2017 | Learning Dynamic Siamese Network for Visual Object TrackingabstractHow to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achieving balanced accuracy and beyond realtime speed. However, they still have a big gap to classification & updating based trackers in tolerating the temporal changes of objects and imaging conditions. In this paper, we propose dynamic Siamese network, via a fast transformation learning model that enables effective online learning of target appearance variation and background suppression from previous frames. We then present elementwise multi-layer fusion to adaptively integrate the network outputs using multi-level deep features. Unlike state-of-theart trackers, our approach allows the usage of any feasible generally- or particularly-trained features, such as SiamFC and VGG. More importantly, the proposed dynamic Siamese network can be jointly trained as a whole directly on the labeled video sequences, thus can take full advantage of the rich spatial temporal information of moving objects. As a result, our approach achieves state-of-the-art performance on OTB-2013 and VOT-2015 benchmarks, while exhibits superiorly balanced accuracy and real-time response over state-of-the-art competitors. Qing Guo 0005, Wei Feng 0005, Ce Zhou, Rui Huang 0006, Song Wang 0002 |
ICCV | 1 |
| 2017 | Structure-Regularized Compressive Tracking With Online Data-Driven SamplingabstractBeing a powerful appearance model, compressive random projection derives effective Haar-like features from non-rotated 4-D-parameterized rectangles, thus supporting fast and reliable object tracking. In this paper, we show that such successful fast compressive tracking scheme can be further significantly improved by structural regularization and online data-driven sampling. Our major contribution is threefold. First, we find that superpixel-guided compressive projection can generate more discriminative features by sufficiently capturing rich local structural information of images. Second, we propose fast directional integration that enables low-cost extraction of feasible Haar-like features from arbitrarily rotated 5-D-parameterized rectangles to realize more accurate object localization. Third, beyond naive dense uniform sampling, we present two practical online data-driven sampling strategies to produce less yet more effective candidate and training samples for object detection and classifier updating, respectively. Extensive experiments on real-world benchmark data sets validate the superior performance, i.e., much better object localization ability and robustness, of the proposed approach over state-of-the-art trackers. Qing Guo 0005, Wei Feng 0005, Ce Zhou, Chi-Man Pun |
IEEE Trans. Image Process. | 1 |
| 2016 | Structure-regularized compressive trackingabstractCompressive random projection is a powerful appearance model to derive effective Haar-like features from non-rotated 4D rectangles, which can support fast and reliable object tracking. In this paper, we show that such successful compressive tracking scheme can be further significantly improved by structural regularization. Specifically, we propose two effective structural regularizations. First, we find that, guided by superpixels, compressive random projection can always generate more discriminative features by sufficiently capturing the rich local structure information of images. Second, we present fast directional integration to enable low-cost extraction of feasible Haar-like features from arbitrarily rotated 5D rectangles to realize more accurate object localization. We compare the proposed structure-regularized compressive tracker with a number of state-of-the-art methods. Extensive experiments on challenging benchmark dataset validate the superior performance and comparable real-time speed of the proposed approach. Qing Guo 0005, Wei Feng 0005, Ce Zhou |
ICME | 1 |
| 2013 | On-line boosting based real-time tracking with efficient HOGabstractIn this paper, a real-time visual tracking system that delivers superior performance under difficult situations is proposed. The system is based on Histogram of Oriented Gradient (HOG) within the on-line boosting framework. For environmental adaptation, the HOG feature is calculated with blocks of random scale, position and aspect ratio which form a feature pool. The on-line boosting can then select the best distinguishable features from this pool for the robust tracking. The randomness of the blocks guarantees the existence of those features. Three experiments are conducted to highlight different characteristics of this new system. The first experiment proves the validity for the system to be able to pick out the best possible HOG features. The second experiment shows its robustness against bad illuminations and small foreground background difference. The third experiment demonstrates its advancement compared with the Haar-based state-of-the-art system. All those are offered without sacrificing the computation load. Shuifa Sun, Qing Guo 0005, Fangmin Dong, Bang Jun Lei |
ICASSP | 2 |
| 2013 | Image denoising algorithm based on contourlet transform for optical coherence tomography heart tube imageabstractOptical coherence tomography (OCT) is becoming an increasingly important imaging technology in the Biomedical field. However, the application of OCT is limited by the ubiquitous noise. In this study, the noise of OCT heart tube image is first verified as being multiplicative based on the local statistics (i.e. the linear relationship between the mean and the standard deviation of certain flat area). The variance of the noise is evaluated in log-domain. Based on these, a joint probability density function is constructed to take the inter-direction dependency in the contourlet domain from the logarithmic transformed image into account. Then, a bivariate shrinkage function is derived to denoise the image by the maximum a posteriori estimation. Systemic comparative experiments are made to synthesis images, OCT heart tube images and other OCT tissue images by subjective assessment and objective metrics. The experiment results are analysed based on the denoising results and the predominance degree of the proposed algorithm with respect to the wavelet-based algorithm. The results show that the proposed algorithm improves the signal-to-noise ratio, whereas preserving the edges and has more advantages on the images containing multi-direction information like OCT heart tube image. Qing Guo 0005, Fangmin Dong, Shuifa Sun, Bang Jun Lei, Bruce Zhi Gao |
IET Image Process. | 1 |