EDBT 2026 Demo / reviewers in the wild / expert
Hao Cheng 0015
dblp:09/5158-15
· DBLP profile ↗
20ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-3246-6636ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | QoMEX 2026 Grand Challenge on Video Quality Assessment for Asymmetric Encoded Videos: Methods and Results
Yixu Chen, Hai Wei, Pierre R. Lebreton, Patrick Le Callet, Alexander Kopte, Amritha Premkumar, Anna Meyer, Baojun Li, Changsheng Gao, Christian Herglotz, Christian Timmerer, Dandan Zhu 0001, Diwakara Reddy, Dong Liu 0002, Dounia Hammou, Guangtao Zhai, Hadi Amirpour, Hao Cheng 0015, Hichem Faraoun, Jonas Janzen, Krishna Srikar Durbha, Li Li 0040, Marc Windsheimer, MohammadAli Hamidi, Mykyta Skipenko, Paul Wawerek-Lopez, Pragyadipta Adhya, Prajit T. Rajendran, Rafal Mantiuk, Shien Ke, Sid Ahmed Fezza, Simon Deniffel, Wei Sun 0029, Weixia Zhang, Xiangguang Chen, Zuowei Cao, Minhao Tang, Xiaoyan Sun 0001, Xingwei Liu, Yeganeh Chatri, Yenan Xu |
QoMEX | 19 |
| 2025 | Not Just Text: Uncovering Vision Modality Typographic Threats in Image Generation ModelsabstractCurrent image generation models can effortlessly produce high-quality, highly realistic images, but this also increases the risk of misuse. In various Text-to-Image or Image-to-Image tasks, attackers can generate a series of images containing inappropriate content by simply editing the language modality input. To mitigate this security concern, numerous guarding or defensive strategies have been proposed, with a particular emphasis on safeguarding language modality. However, in practical applications, threats in the vision modality, particularly in tasks involving the editing of real-world images, present heightened security risks as they can easily infringe upon the rights of the image owner. Therefore, this paper employs a method named typographic attack to reveal that various image generation models are also susceptible to threats within the vision modality. Furthermore, we also evaluate the defense performance of various existing methods when facing threats in the vision modality and uncover their ineffectiveness. Finally, we propose the Vision Modal Threats in Image Generation Models (VMT-IGMs) dataset, which would serve as a baseline for evaluating the vision modality vulnerability of various image generation models.Warning: This paper includes content that may cause discomfort or distress. Potentially disturbing content has been blocked and blurred. Hao Cheng 0015, Erjia Xiao, Jiayan Yang, Jiahang Cao, Qiang Zhang 0029, Jize Zhang, Kaidi Xu, Jindong Gu, Renjing Xu |
CVPR | 1 |
| 2025 | Event Masked Autoencoder: Point-wise Action Recognition with Event-Based CamerasabstractDynamic vision sensors (DVS) are bio-inspired devices that capture visual information in the form of asynchronous events, which encode changes in pixel intensity with high temporal resolution and low latency. These events provide rich motion cues that can be exploited for various computer vision tasks, such as action recognition. However, most existing DVS-based action recognition methods lose temporal information during data transformation or suffer from noise and outliers caused by sensor imperfections or environmental factors. To address these challenges, we propose a novel framework that preserves and exploits the spatiotemporal structure of event data for action recognition. Our framework consists of two main components: 1) a point-wise event masked autoencoder (MAE) that learns a compact and discriminative representation of event patches by reconstructing them from masked raw event camera points data; 2) an improved event points patch generation algorithm that leverages an event data inlier model and point-wise data augmentation techniques to enhance the quality and diversity of event points patches. To the best of our knowledge, our approach introduces the pre-train method into event camera raw points data for the first time, and we propose a novel event points patch embedding to utilize transformer-based models on event cameras. Jingkai Sun, Qiang Zhang 0029, Jiahang Cao, Hao Cheng 0015, Renjing Xu |
ICASSP | 5 |
| 2025 | TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination via Latent Truthful-Guided Pre-InterventionabstractObject Hallucination (OH) has been acknowledged as one of the major trustworthy challenges in Large Vision-Language Models (LVLMs). Recent advancements in Large Language Models (LLMs) indicate that internal states, such as hidden states, encode the "overall truthfulness" of generated responses. However, it remains under-explored how internal states in LVLMs function and whether they could serve as "per-token" hallucination indicators, which is essential for mitigating OH. In this paper, we first conduct an in-depth exploration of LVLM internal states with OH issues and discover that (1) LVLM internal states are high-specificity per-token indicators of hallucination behaviors. Moreover, (2) different LVLMs encode universal patterns of hallucinations in common latent subspaces, indicating that there exist "generic truthful directions" shared by various LVLMs. Based on these discoveries, we propose Truthful-Guided Pre-Intervention (TruthPrInt) that first learns the truthful direction of LVLM decoding and then applies truthful-guided inference-time intervention during LVLM decoding. We further propose TruthPrInt to enhance both cross-LVLM and cross-data hallucination detection transferability by constructing and aligning hallucination latent subspaces. We evaluate TruthPrInt in extensive experimental settings, including in-domain and out-of-domain scenarios, over popular LVLMs and OH benchmarks. Experimental results indicate that TruthPrInt significantly outperforms state-of-the-art methods. Codes will be available at https://github.com/jinhaoduan/TruthPrInt. Jinhao Duan, Fei Kong, Hao Cheng 0015, James Diffenderfer, Bhavya Kailkhura, Lichao Sun 0001, Xiaofeng Zhu 0001, Xiaoshuang Shi, Kaidi Xu |
ICCV | 3 |
| 2025 | Mamba Policy: Towards Efficient 3D Diffusion Policy with Hybrid Selective State ModelsabstractDiffusion models have been widely employed in the field of 3D manipulation due to their efficient capability to learn distributions, allowing for precise prediction of action trajectories. However, diffusion models typically rely on large parameter UNet backbones as policy networks, which can be challenging to deploy on resource-constrained devices. Recently, the Mamba model has emerged as a promising solution for efficient modeling, offering low computational complexity and strong performance in sequence modeling. In this work, we propose the Mamba Policy, a lighter but stronger policy that reduces the parameter count by over 80% compared to the original policy network while achieving superior performance. Specifically, we introduce the XMamba Block, which effectively integrates input information with conditional features and leverages a combination of Mamba and Attention mechanisms for deep feature extraction. Extensive experiments demonstrate that the Mamba Policy excels on the Adroit, Dexart, and MetaWorld datasets, requiring significantly fewer computational resources. Additionally, we highlight the Mamba Policy’s enhanced robustness in long-horizon scenarios compared to baseline methods and explore the performance of various Mamba variants within the Mamba Policy framework. Real-world experiments are also conducted to further validate its effectiveness. Our open-source project page can be found at https://andycao1125.github.io/mamba_policy/. Jiahang Cao, Qiang Zhang 0029, Jingkai Sun, Hao Cheng 0015, Yulin Li 0001, Jun Ma 0008, Kun Wu 0001, Yecheng Shao, Yijie Guo, Renjing Xu |
IROS | 5 |
| 2025 | Transfer Attack for Bad and Good: Explain and Boost Adversarial Transferability across Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) demonstrate exceptional performance in cross-modality interaction, yet they also suffer adversarial vulnerabilities. In particular, the transferability of adversarial examples remains an ongoing challenge. In this paper, we specifically analyze the manifestation of adversarial transferability among MLLMs and identify the key factors that influence this characteristic. We discover that the transferability of MLLMs exists in cross-LLM scenarios with the same vision encoder and indicate two key Factors that may influence transferability. We provide two semantic-level data augmentation methods, Adding Image Patch (AIP) and Typography Augment Transferability Method (TATM), which boost the transferability of adversarial examples across MLLMs. To explore the potential impact in the real world, we utilize two tasks that can have both negative and positive societal impacts: 1. Harmful Content Insertion and 2. Information Protection. Hao Cheng 0015, Erjia Xiao, Jiayan Yang, Jinhao Duan, Yichi Wang 0002, Jiahang Cao, Qiang Zhang 0029, Le Yang 0007, Kaidi Xu, Jindong Gu, Renjing Xu |
ACM Multimedia | 1 |
| 2025 | Jailbreak-AudioBench: In-Depth Evaluation and Analysis of Jailbreak Threats for Large Audio Language ModelsabstractLarge Language Models (LLMs) demonstrate impressive zero-shot performance across a wide range of natural language processing tasks. Integrating various modality encoders further expands their capabilities, giving rise to Multimodal Large Language Models (MLLMs) that process not only text but also visual and auditory modality inputs. However, these advanced capabilities may also pose significant safety problems, as models can be exploited to generate harmful or inappropriate content through jailbreak attack. While prior work has extensively explored how manipulating textual or visual modality inputs can circumvent safeguards in LLMs and MLLMs, the vulnerability of audio-specific Jailbreak on Large Audio-Language Models (LALMs) remains largely underexplored. To address this gap, we introduce \textbf{Jailbreak-AudioBench}, which consists of the Toolbox, curated Dataset, and comprehensive Benchmark. The Toolbox supports not only text-to-audio conversion but also various editing techniques for injecting audio hidden semantics. The curated Dataset provides diverse explicit and implicit jailbreak audio examples in both original and edited forms. Utilizing this dataset, we evaluate multiple state-of-the-art LALMs and establish the most comprehensive Jailbreak benchmark to date for audio modality. Finally, Jailbreak-AudioBench establishes a foundation for advancing future research on LALMs safety alignment by enabling the in-depth exposure of more powerful jailbreak threats, such as query-based audio editing, and by facilitating the development of effective defense mechanisms. Hao Cheng 0015, Erjia Xiao, Yichi Wang 0002, Le Yang 0007, Chao Shen 0001, Philip Torr 0001, Jindong Gu, Renjing Xu |
NeurIPS | 1 |
| 2025 | VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment: Methods and ResultsabstractThis paper reviews the VCIP 2025 Grand Challenge on Live Broadcasting Video Quality Assessment. The competition aims to foster innovation in both subjective and objective VQA techniques tailored to live broadcasting videos, addressing the unique challenges posed by live streaming impairments while emphasizing the evaluation of QoE. The grand challenge used live broadcasting database LBVD which consists of 1013 videos focusing on distortion in live broadcasting videos. The competition had 14 participants and 5 teams submitted valid solutions for the final testing phase. The proposed solutions have shown significant progress in areas such as combining traditional feature engineering with deep learning models, achieved state-of-the-art performances for LBVD. Team ATHENA-Live-QoE and Team HZX Force tied for the first position. The dataset can be found at https://github.com/cpf0079/LBVD. Wenqi Fei, Yuhua Zhang, MohammadAli Hamidi, Hadi Amirpour, Erjia Xiao, Zhenjie Su, Hao Cheng 0015, Yu Liu 0023, Wei Zhou 0021, Yanbiao Ma, Renjing Xu, Long Chen 0015, Xiaoshuai Hao, Yipo Huang, Tushar Shinde |
VCIP | 10 |
| 2024 | Shifting Attention to Relevance: Towards the Predictive Uncertainty Quantification of Free-Form Large Language ModelsabstractJinhao Duan, Hao Cheng, Shiqi Wang, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, Kaidi Xu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jinhao Duan, Hao Cheng 0015, Shiqi Wang 0002, Alex Zavalny, Chenan Wang, Renjing Xu, Bhavya Kailkhura, Kaidi Xu |
ACL (1) | 2 |
| 2024 | ACT-Diffusion: Efficient Adversarial Consistency Training for One-Step Diffusion ModelsabstractThough diffusion models excel in image generation, their step-by-step denoising leads to slow generation speeds. Consistency training addresses this issue with single-step sampling but often produces lower-quality generations and requires high training costs. In this paper, we show that optimizing consistency training loss minimizes the Wasserstein distance between target and generated distributions. As timestep increases, the upper bound accumulates previous consistency training losses. Therefore, larger batch sizes are needed to reduce both current and accumulated losses. We propose Adversarial Consistency Training (ACT), which directly minimizes the Jensen-Shannon (JS) divergence between distributions at each timestep using a discriminator. Theoretically, ACT enhances generation quality, and convergence. By incorporating a discriminator into the consistency training framework, our method achieves improved FID scores on CIFAR10 and ImageNet 64×64 and LSUN Cat 256 ×256 datasets, retains zero-shot image inpainting capabilities, and uses less than 1/6 of the original batch size and fewer than 1/2 of the model parameters and training steps compared to the baseline method, this leads to a substantial reduction in resource consumption. Our code is available: https://github.com/kong13661/ACT Fei Kong, Jinhao Duan, Lichao Sun 0001, Hao Cheng 0015, Renjing Xu, Heng Tao Shen, Xiaofeng Zhu 0001, Xiaoshuang Shi, Kaidi Xu |
CVPR | 4 |
| 2024 | Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Models
Hao Cheng 0015, Erjia Xiao, Jindong Gu, Le Yang 0007, Jinhao Duan, Jize Zhang, Jiahang Cao, Kaidi Xu, Renjing Xu |
ECCV (59) | 1 |
| 2024 | DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
Le Yang 0007, Ziwei Zheng, Yizeng Han, Hao Cheng 0015, Shiji Song, Gao Huang 0001, Fan Li 0003 |
ECCV (46) | 4 |
| 2024 | DONE: Dynamic Neural Representation Via Hyperplane Neural ODEabstractMuch progress has been made in dynamic scene reconstruction by neural rendering techniques. Even if existing methods based on Neural Radiance Field (NeRF) achieve marvelous fidelity, they suffer from either complex camera settings, dense training timestamps, or limited ability to inter- and extrapolate. In this paper, a novel paradigm is presented to address these issues. Our method, called Dynamic Ode NErf (DONE), comprises two main modules. The first one trains a deformable neural representation in a static scene to store neural features. The second module establishes a deterministic neural ordinary differential equation to model the dynamics of the scene object. We also present a new 360-degree dynamic dataset for the research of entire 360-degree dynamic scene reconstruction. Experiments visually and quantitatively illustrate the effectiveness of the proposed model. Hao Cheng 0015, Renjing Xu |
ICASSP | 3 |
| 2024 | Gaining the Sparse Rewards by Exploring Lottery Tickets in Spiking Neural NetworksabstractDeploying energy-efficient deep learning algorithms on computational-limited devices, such as robots, is still a pressing issue for real-world applications. Spiking Neural Networks (SNNs), a novel brain-inspired algorithm, offer a promising solution due to their low-latency and low-energy properties over traditional Artificial Neural Networks (ANNs). Despite their advantages, the dense structure of deep SNNs can still result in extra energy consumption. The Lottery Ticket Hypothesis (LTH) posits that within dense neural networks, there exist winning Lottery Tickets (LTs), namely sub-networks, that can be obtained without compromising performance. Inspired by this, this paper delves into the spiking-based LTs (SLTs), examining their unique properties and potential for extreme efficiency. Then, two significant sparse Rewards are gained through comprehensive explorations and meticulous experiments on SLTs across various dense structures. Moreover, a sparse algorithm tailored for spiking transformer structure, which incorporates convolution operations into the Patch Embedding Projection (ConvPEP) module, has been proposed to achieve Multi-level Sparsity (MultiSp). MultiSp refers to (1) Patch number sparsity; (2) ConvPEP weights sparsity and binarization; and (3) ConvPEP activation layer binarization. Extensive experiments demonstrate that our method achieves extreme sparsity with only a slight performance decrease, paving the way for deploying energy-efficient neural networks in robotics and beyond. Hao Cheng 0015, Jiahang Cao, Erjia Xiao, Mengshu Sun, Renjing Xu |
IROS | 1 |
| 2024 | Energy-based Active Learning for Bringing Beam-induced Domain Gap for 3D Object DetectionabstractIn many real-world applications, 16-beam LiDAR-based 3D object detection (3DOD) is indispensable in scene understanding. However, the absence of well-labeled large-scale 16-beam LiDAR datasets impedes the development of these 3DOD methods. To avoid annotation costs in developing datasets, we proposed an energy-based active learning method for cross-beam domain adaptation, which effectively transfers the knowledge from the existing well-labeled 64-beam counterpart. Specifically, the cross-beam domain gap between the source (64-beam) and the target (16-beam) domain is reduced by aligning the deep features based on an energy-based feature-matching loss term during training. Moreover, the proposed energy-based active learning method enables the sampling strategy to shed light on selecting the most valuable 16-beam target samples to be manually labeled, which are then added to the training set. Experimental results show that our method can effectively transfer the knowledge from the 64-beam domain to the 16-beam one, and successfully learns a high-performance 16-beam 3DOD model with only a small portion of unlabeled data to annotate. Le Yang 0007, Yixuan Yan, Hao Cheng 0015, Fan Li 0003 |
MobiCom | 4 |
| 2024 | Spiking Neural Network as Adaptive Event Stream SlicerabstractEvent-based cameras are attracting significant interest as they provide rich edge information, high dynamic range, and high temporal resolution. Many state-of-the-art event-based algorithms rely on splitting the events into fixed groups, resulting in the omission of crucial temporal information, particularly when dealing with diverse motion scenarios (e.g., high/low speed). In this work, we propose SpikeSlicer, a novel-designed event processing framework capable of splitting events stream adaptively. SpikeSlicer utilizes a low-energy spiking neural network (SNN) to trigger event slicing. To guide the SNN to fire spikes at optimal time steps, we propose the Spiking Position-aware Loss (SPA-Loss) to modulate the neuron's state. Additionally, we develop a Feedback-Update training strategy that refines the slicing decisions using feedback from the downstream artificial neural network (ANN). Extensive experiments demonstrate that our method yields significant performance improvements in event-based object tracking and recognition. Notably, SpikeSlicer provides a brand-new SNN-ANN cooperation paradigm, where the SNN acts as an efficient, low-energy data processor to assist the ANN in improving downstream performance, injecting new perspectives and potential avenues of exploration. Jiahang Cao, Hao Cheng 0015, Qiang Zhang 0029, Shibo Zhou, Renjing Xu |
NeurIPS | 4 |
| 2024 | Spiking Denoising Diffusion Probabilistic ModelsabstractSpiking neural networks (SNNs) have ultra-low energy consumption and high biological plausibility due to their binary and bio-driven nature compared with artificial neural networks (ANNs). While previous research has primarily focused on enhancing the performance of SNNs in classification tasks, the generative potential of SNNs remains relatively unexplored. In our paper, we put forward Spiking Denoising Diffusion Probabilistic Models (SDDPM), a new class of SNN-based generative models that achieve high sample quality. To fully exploit the energy efficiency of SNNs, we propose a purely Spiking U-Net architecture, which achieves comparable performance to its ANN counterpart using only 4 time steps, resulting in significantly reduced energy consumption. Extensive experimental results reveal that our approach achieves state-of-the-art on the generative tasks and substantially outperforms other SNN-based generative models, achieving up to 12× and 6× improvement on the CIFAR-10 and the CelebA datasets, respectively. Moreover, we propose a threshold-guided strategy that can further improve the performances by 2.69% in a training-free manner. The SDDPM symbolizes a significant advancement in the field of SNN generation, injecting new perspectives and potential avenues of exploration. Our code is available at https://github.com/AndyCao1125/SDDPM. Jiahang Cao, Hanzhong Guo, Hao Cheng 0015, Qiang Zhang 0029, Renjing Xu |
WACV | 4 |
| 2023 | RBFormer: Improve Adversarial Robustness of Transformers by Robust Bias
Hao Cheng 0015, Jinhao Duan, Jiahang Cao, Lyutianyang Zhang, Jize Zhang, Kaidi Xu, Renjing Xu |
BMVC | 1 |
| 2023 | Improve Video Representation with Temporal Adversarial AugmentationabstractRecent works reveal that adversarial augmentation benefits the generalization of neural networks (NNs) if used in an appropriate manner. In this paper, we introduce Temporal Adversarial Augmentation (TA), a novel video augmentation technique that utilizes temporal attention. Unlike conventional adversarial augmentation, TA is specifically designed to shift the attention distributions of neural networks with respect to video clips by maximizing a temporal-related loss function. We demonstrate that TA will obtain diverse temporal views, which significantly affect the focus of neural networks. Training with these examples remedies the flaw of unbalanced temporal information perception and enhances the ability to defend against temporal shifts, ultimately leading to better generalization. To leverage TA, we propose Temporal Video Adversarial Fine-tuning (TAF) framework for improving video representations. TAF is a model-agnostic, generic, and interpretability-friendly training strategy. We evaluate TAF with four powerful models (TSM, GST, TAM, and TPN) over three challenging temporal-related benchmarks (Something-something V1&V2 and diving48). Experimental results demonstrate that TAF effectively improves the test accuracy of these models with notable margins without introducing additional parameters or computational costs. As a byproduct, TAF also improves the robustness under out-of-distribution (OOD) settings. Code is available at https://github.com/jinhaoduan/TAF. Jinhao Duan, Quanfu Fan, Hao Cheng 0015, Xiaoshuang Shi, Kaidi Xu |
IJCAI | 3 |
| 2019 | Adversarial Robustness vs. Model Compression, or Both?abstractIt is well known that deep neural networks (DNNs) are vulnerable to adversarial attacks, which are implemented by adding crafted perturbations onto benign examples. Min-max robust optimization based adversarial training can provide a notion of security against adversarial attacks. However, adversarial robustness requires a significantly larger capacity of the network than that for the natural training with only benign examples. This paper proposes a framework of concurrent adversarial training and weight pruning that enables model compression while still preserving the adversarial robustness and essentially tackles the dilemma of adversarial training. Furthermore, this work studies two hypotheses about weight pruning in the conventional setting and finds that weight pruning is essential for reducing the network model size in the adversarial setting; training a small model from scratch even with inherited initialization from the large model cannot achieve neither adversarial robustness nor high standard accuracy. Code is available at https://github.com/yeshaokai/Robustness-Aware-Pruning-ADMM. Shaokai Ye, Xue Lin 0001, Kaidi Xu, Sijia Liu 0001, Hao Cheng 0015, Jan-Henrik Lambrechts, Huan Zhang 0001, Aojun Zhou, Kaisheng Ma, Yanzhi Wang 0001 |
ICCV | 5 |