EDBT 2026 Demo / reviewers in the wild / expert
Qinhong Jiang
dblp:262/3567
· DBLP profile ↗
23ranked-venue papers
2as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 since 2021Security and privacy · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Electromagnetic interference (EMI) backdoor: An EMI-based backdoor attack against computer vision systemsabstractRecently, computer vision systems, for example, smart traffic surveillance systems, facial recognition systems, etc., have significantly changed our daily life. Even though the neural networks in such systems are known to suffer from backdoor attacks, causing the backdoored models to behave well on benign samples but maliciously on controlled samples (with triggers applied to activate the backdoor), it is generally believed that most of the triggers, when used in physical attacks, are noticeable to victim users and not robust in various settings, such as different angles, distances, lighting conditions, etc. In this paper, we leverage electromagnetic interference (EMI) to produce a specific pattern distortion in images captured by the camera system and utilize the pattern distortion as the backdoor trigger. To avoid the overhead of manually collecting poisoned images, we introduce a simulation sample generation approach, converting clean images to poisoned ones by simulating the distortion caused by EMI against the camera system. Additionally, we propose a contrast loss function to enhance the generalization of backdoor features, improving triggers’ capability to activate the embedded backdoors. We conduct extensive physical experiments using diverse deep neural networks across various camera systems in different practical environments, achieving a 92.54% average backdoor success rate. Mengjie Sun, Peizhuo Lv, Shengzhi Zhang, Jianshuo Liu, Kai Chen 0012, Hong Li 0004, Zhi Li 0018, Qinhong Jiang, Limin Sun 0001 |
J. Comput. Secur. | 8 |
| 2025 | PhantomLiDAR: Cross-modality Signal Injection Attacks against LiDAR
Zizhi Jin, Qinhong Jiang, Xuancun Lu, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
NDSS | 2 |
| 2025 | GhostShot: Manipulating the Image of CCD Cameras with Electromagnetic Interference
Yanze Ren, Qinhong Jiang, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
NDSS | 2 |
| 2024 | Understanding Impacts of Electromagnetic Signal Injection Attacks on Object DetectionabstractObject detection can localize and identify objects in images, and it is extensively employed in critical multimedia applications such as security surveillance and autonomous driving. Despite the success of existing object detection models, they are often evaluated in ideal scenarios where captured images guarantee the accurate and complete representation of the detecting scenes. However, images captured by image sensors may be affected by different factors in real applications, including cyber-physical attacks. In particular, attackers can exploit hardware properties within the systems to inject electromagnetic interference so as to manipulate the images. Such attacks can cause noisy or incomplete information about the captured scene, leading to incorrect detection results, potentially granting attackers malicious control over critical functions of the systems. This paper presents a research work that comprehensively quantifies and analyzes the impacts of such attacks on state-of-the-art object detection models in practice. It also sheds light on the underlying reasons for the incorrect detection outcomes. Youqian Zhang, Eugene Yujun Fu, Qinhong Jiang, Chen Yan 0001, Sze-Yiu Chau, Grace Ngai, Hong Va Leong, Xiapu Luo, Wenyuan Xu 0001 |
ICME | 4 |
| 2024 | GhostType: The Limits of Using Contactless Electromagnetic Interference to Inject Phantom Keys into Analog Circuits of Keyboards
Qinhong Jiang, Yanze Ren, Yan Long 0002, Chen Yan 0001, Yumai Sun, Xiaoyu Ji 0001, Kevin Fu, Wenyuan Xu 0001 |
NDSS | 1 |
| 2024 | EM Eye: Characterizing Electromagnetic Side-channel Eavesdropping on Embedded Cameras
Yan Long 0002, Qinhong Jiang, Chen Yan 0001, Tobias Alam, Xiaoyu Ji 0001, Wenyuan Xu 0001, Kevin Fu |
NDSS | 2 |
| 2024 | Watch Your Speed: Injecting Malicious Voice Commands via Time-Scale ModificationabstractExisting adversarial example (AE) attacks against automatic speech recognition (ASR) systems focus on adding deliberate noises to input audio. In this paper, we propose a new attack that purely speeds up or slows down original audio instead of adding perturbations, and we call it Time-Scale Modification Adversarial Example (TSMAE). By investigating the impact of speed variation on 100, 000 pieces of audio clips, we found that misrecognition manifests in three categories: delete, substitution, and insertion. These are the accumulated results caused by the misrecognition of both the acoustic and language models inside an ASR system. Despite the challenges, i.e., ASR systems are typically black-box and reveal no gradient information, we managed to launch one-segment untargeted and targetedTSMAEattacks based on particle swarm optimization algorithms. Our untargeted attacks only require modifying the speed of one segment (e.g., 20 ms), and our targeted attacks can generate meaningful yet benign audio to cause an ASR system to output a malicious output, e.g., “open the door”. We validate the feasibility ofTSMAEon two open-source ASR models (e.g., DeepSpeech and Sphinx) and four commercial ones (e.g., IBM, Google, Baidu, and iFLYTEK). Results show that our untargeted attack can successfully attack all 6 ASR models with one segment modification, and our targeted attack is robust to various factors, such as model versions and speech sources. Finally, both attacks can bypass existing open-source defense methods, and our insights call attention to the defense’s focus from coping with perturbation to emerging adversarial example attacks. Xiaoyu Ji 0001, Qinhong Jiang, Chaohao Li, Zhuoyang Shi, Wenyuan Xu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Graph-DETR4D: Spatio-Temporal Graph Modeling for Multi-View 3D Object DetectionabstractMulti-View 3D object detection (MV3D) has made tremendous progress by leveraging multiple perspective features through surrounding cameras. Despite demonstrating promising prospects in various applications, accurately detecting objects through camera view in the 3D space is extremely difficult due to the ill-posed issue in monocular depth estimation. Recently, Graph-DETR3D presents a novel graph-based 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves competitive performance. Although it enriches the query representations with 2D image features through a learnable 3D graph, it still suffers from limited depth and velocity estimation abilities due to the adoption of a single-frame input setting. To solve this problem, we introduce a unified spatial-temporal graph modeling framework to fully leverage the multi-view imagery cues under the multi-frame inputs setting. Thanks to the flexibility and sparsity of the dynamic graph architecture, we lift the original 3D graph into the 4D space with an effective attention mechanism to automatically perceive imagery information at both spatial and temporal levels. Moreover, considering the main latency bottleneck lies in the image backbone, we propose a novel dense-sparse distillation framework for multi-view 3D object detection, to reduce the computational budget while sacrificing no detection accuracy, making it more suitable for real-world deployment. To this end, we propose Graph-DETR4D, a faster and stronger multi-view 3D object detection framework, built on top of Graph-DETR3D. Extensive experiments on nuScenes and Waymo benchmarks demonstrate the effectiveness and efficiency of Graph-DETR4D. Notably, our best model achieves 62.0% NDS on nuScenes test leaderboard. Code is available at https://github.com/zehuichen123/Graph-DETR4D. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Wu 0005, Feng Zhao 0004 |
IEEE Trans. Image Process. | 6 |
| 2023 | BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004 |
ICLR | 5 |
| 2023 | GlitchHiker: Uncovering Vulnerabilities of Image Signal Transmission with IEMI
Qinhong Jiang, Xiaoyu Ji 0001, Chen Yan 0001, Zhixin Xie, Haina Lou, Wenyuan Xu 0001 |
USENIX Security Symposium | 1 |
| 2022 | SimIPU: Simple 2D Image and 3D Point Cloud Unsupervised Pre-training for Spatial-Aware Visual RepresentationsabstractPre-training has become a standard paradigm in many computer vision tasks. However, most of the methods are generally designed on the RGB image domain. Due to the discrepancy between the two-dimensional image plane and the three-dimensional space, such pre-trained models fail to perceive spatial information and serve as sub-optimal solutions for 3D-related tasks. To bridge this gap, we aim to learn a spatial-aware visual representation that can describe the three-dimensional space and is more suitable and effective for these tasks. To leverage point clouds, which are much more superior in providing spatial information compared to images, we propose a simple yet effective 2D Image and 3D Point cloud Unsupervised pre-training strategy, called SimIPU. Specifically, we develop a multi-modal contrastive learning framework that consists of an intra-modal spatial perception module to learn a spatial-aware representation from point clouds and an inter-modal feature interaction module to transfer the capability of perceiving spatial information from the point cloud encoder to the image encoder, respectively. Positive pairs for contrastive losses are established by the matching algorithm and the projection matrix. The whole framework is trained in an unsupervised end-to-end fashion. To the best of our knowledge, this is the first study to explore contrastive learning pre-training strategies for outdoor multi-modal datasets, containing paired camera images and LIDAR point clouds. Zhenyu Li 0007, Liangji Fang, Qinhong Jiang, Xianming Liu 0005, Junjun Jiang, Bolei Zhou, Hang Zhao 0021 |
AAAI | 5 |
| 2022 | Deformable Feature Aggregation for Dynamic Multi-modal 3D Object Detection
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004 |
ECCV (8) | 5 |
| 2022 | Unsupervised Domain Adaptation for Monocular 3D Object Detection via Self-training
Zhenyu Li 0007, Liangji Fang, Qinhong Jiang, Xianming Liu 0005, Junjun Jiang |
ECCV (9) | 5 |
| 2022 | AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object DetectionabstractObject detection through either RGB images or the LiDAR point clouds has been extensively explored in autonomous driving. However, it remains challenging to make these two data sources complementary and beneficial to each other. In this paper, we propose AutoAlign, an automatic feature fusion strategy for 3D object detection. Instead of establishing deterministic correspondence with camera projection matrix, we model the mapping relationship between the image and point clouds with a learnable alignment map. This map enables our model to automate the alignment of non-homogenous features in a dynamic and data-driven manner. Specifically, a cross-attention feature alignment module is devised to adaptively aggregate pixel-level image features for each voxel. To enhance the semantic consistency during feature alignment, we also design a self-supervised cross-modal feature interaction module, through which the model can learn feature aggregation with instance-level feature guidance. Extensive experimental results show that our approach can lead to 2.3 mAP and 7.0 mAP improvements on the KITTI and nuScenes datasets respectively. Notably, our best model reaches 70.9 NDS on the nuScenes testing leaderboard, achieving competitive performance among various state-of-the-arts. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004, Bolei Zhou, Hang Zhao 0021 |
IJCAI | 5 |
| 2022 | Graph-DETR3D: Rethinking Overlapping Regions for Multi-View 3D Object Detectionabstract3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. However, accurately detecting objects through perspective views in the 3D space is extremely difficult due to the lack of depth information. Recently, DETR3D introduces a novel 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves state-of-the-art performance. In this paper, with intensive pilot experiments, we quantify the objects located at different regions and find that the "truncated instances'' (i.e., at the border regions of each image) are the main bottleneck hindering the performance of DETR3D. Although it merges multiple features from two adjacent views in the overlapping regions, DETR3D still suffers from insufficient feature aggregation, thus missing the chance to fully boost the detection performance. In an effort to tackle the problem, we propose Graph-DETR3D to automatically aggregate multi-view imagery information through graph structure learning. It constructs a dynamic 3D graph between each object query and 2D feature maps to enhance the object representations, especially at the border regions. Besides, Graph-DETR3D benefits from a novel depth-invariant multi-scale training strategy, which maintains the visual depth consistency by simultaneously scaling the image size and the object depth. Extensive experiments on the nuScenes dataset demonstrate the effectiveness and efficiency of our Graph-DETR3D. Notably, our best model achieves 49.5 NDS on the nuScenes test leaderboard, achieving new state-of-the-art in comparison with various published image-view 3D object detectors. Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004 |
ACM Multimedia | 5 |
| 2022 | Shape Prior Guided Instance Disparity Estimation for 3D Object DetectionabstractIn this paper, we propose a novel system named Disp R-CNN for 3D object detection from stereo images. Many recent works solve this problem by first recovering point clouds with disparity estimation and then apply a 3D detector. The disparity map is computed for the entire image, which is costly and fails to leverage category-specific prior. In contrast, we design an instance disparity estimation network (iDispNet) that predicts disparity only for pixels on objects of interest and learns a category-specific shape prior for more accurate disparity estimation. To address the challenge from scarcity of disparity annotation in training, we propose to use a statistical shape model to generate dense disparity pseudo-ground-truth without the need of LiDAR point clouds, which makes our system more widely applicable. Experiments on the KITTI dataset show that, when LiDAR ground-truth is not used at training time, Disp R-CNN outperforms previous state-of-the-art methods based on stereo input by 20 percent in terms of average precision for all categories. The code and pseudo-ground-truth data are available at the project page: https://github.com/zju3dv/disprcnn. Jiaming Sun 0002, Qing Shuai, Qinhong Jiang, Guofeng Zhang 0001, Hujun Bao, Xiaowei Zhou 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | MonoEF: Extrinsic Parameter Free Monocular 3D Object DetectionabstractMonocular 3D object detection is an important task in autonomous driving. It can be easily intractable where there exists ego-car pose change w.r.t. ground plane. This is common due to the slight fluctuation of road smoothness and slope. Due to the lack of insight in industrial application, existing methods on open datasets neglect the camera pose information, which inevitably results in the detector being susceptible to camera extrinsic parameters. The perturbation of objects is very popular in most autonomous driving cases for industrial products. To this end, we propose a novel method to capture camera pose to formulate the detector free from extrinsic perturbation. Specifically, the proposed framework predicts camera extrinsic parameters by detecting vanishing point and horizon change. A converter is designed to rectify perturbative features in the latent space. By doing so, our 3D detector works independent of the extrinsic parameter variations and produces accurate results in realistic cases, e.g., potholed and uneven roads, where almost all existing monocular detectors fail to handle. Experiments demonstrate our method yields the best performance compared with the other state-of-the-arts by a large margin on both KITTI 3D and nuScenes datasets. Yunsong Zhou, Hongzi Zhu, Cheng Wang 0043, Qinhong Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Multimodal Motion Prediction With Stacked TransformersabstractPredicting multiple plausible future trajectories of the nearby vehicles is crucial for the safety of autonomous driving. Recent motion prediction approaches attempt to achieve such multimodal motion prediction by implicitly regularizing the feature or explicitly generating multiple candidate proposals. However, it remains challenging since the latent features may concentrate on the most frequent mode of the data while the proposal-based methods depend largely on the prior knowledge to generate and select the proposals. In this work, we propose a novel transformer framework for multimodal motion prediction, termed as mmTransformer. A novel network architecture based on stacked transformers is designed to model the multimodality at feature level with a set of fixed independent proposals. A region-based training strategy is then developed to induce the multimodality of the generated proposals. Experiments on Argoverse dataset show that the proposed model achieves the state-of-the-art performance on motion prediction, substantially improving the diversity and the accuracy of the predicted trajectories. Demo video and code are available at https://decisionforce.github.io/mmTransformer. Jinghuai Zhang, Liangji Fang, Qinhong Jiang, Bolei Zhou |
CVPR | 4 |
| 2021 | Monocular 3D Object Detection: An Extrinsic Parameter Free ApproachabstractMonocular 3D object detection is an important task in autonomous driving. It can be easily intractable where there exists ego-car pose change w.r.t. ground plane. This is common due to the slight fluctuation of road smoothness and slope. Due to the lack of insight in industrial application, existing methods on open datasets neglect the cam-era pose information, which inevitably results in the detector being susceptible to camera extrinsic parameters. The perturbation of objects is very popular in most autonomous driving cases for industrial products. To this end, we propose a novel method to capture camera pose to formulate the detector free from extrinsic perturbation. Specifically, the proposed framework predicts camera extrinsic parameters by detecting vanishing point and horizon change. A converter is designed to rectify perturbative features in the latent space. By doing so, our 3D detector works independent of the extrinsic parameter variations and produces accurate results in realistic cases, e.g., potholed and uneven roads, where almost all existing monocular detectors fail to handle. Experiments demonstrate our method yields the best performance compared with the other state-of-the-arts by a large margin on both KITTI 3D and nuScenes datasets. Yunsong Zhou, Hongzi Zhu, Cheng Wang 0043, Qinhong Jiang |
CVPR | 6 |
| 2020 | TPNet: Trajectory Proposal Network for Motion PredictionabstractMaking accurate motion prediction of the surrounding traffic agents such as pedestrians, vehicles, and cyclists is crucial for autonomous driving. Recent data-driven motion prediction methods have attempted to learn to directly regress the exact future position or its distribution from massive amount of trajectory data. However, it remains difficult for these methods to provide multimodal predictions as well as integrate physical constraints such as traffic rules and movable areas. In this work we propose a novel two-stage motion prediction framework, Trajectory Proposal Network (TPNet). TPNet first generates a candidate set of future trajectories as hypothesis proposals, then makes the final predictions by classifying and refining the proposals which meets the physical constraints. By steering the proposal generation process, safe and multimodal predictions are realized. Thus this framework effectively mitigates the complexity of motion prediction problem while ensuring the multimodal output. Experiments on four large-scale trajectory prediction datasets, i.e. the ETH, UCY, Apollo and Argoverse datasets, show that TPNet achieves the state-of-the-art results both quantitatively and qualitatively. Liangji Fang, Qinhong Jiang, Jianping Shi, Bolei Zhou |
CVPR | 2 |
| 2020 | Disp R-CNN: Stereo 3D Object Detection via Shape Prior Guided Instance Disparity EstimationabstractIn this paper, we propose a novel system named Disp R-CNN for 3D object detection from stereo images. Many recent works solve this problem by first recovering a point cloud with disparity estimation and then apply a 3D detector. The disparity map is computed for the entire image, which is costly and fails to leverage category-specific prior. In contrast, we design an instance disparity estimation network (iDispNet) that predicts disparity only for pixels on objects of interest and learns a category-specific shape prior for more accurate disparity estimation. To address the challenge from scarcity of disparity annotation in training, we propose to use a statistical shape model to generate dense disparity pseudo-ground-truth without the need of LiDAR point clouds, which makes our system more widely applicable. Experiments on the KITTI dataset show that, even when LiDAR ground-truth is not available at training time, Disp R-CNN achieves competitive performance and outperforms previous state-of-the-art methods by 20% in terms of average precision. The code will be available at https://github.com/zju3dv/disprcnn. Jiaming Sun 0002, Qinhong Jiang, Xiaowei Zhou 0001, Hujun Bao |
CVPR | 5 |
| 2020 | Recursive Social Behavior Graph for Trajectory PredictionabstractSocial interaction is an important topic in human trajectory prediction to generate plausible paths. In this paper, we present a novel insight of group-based social interaction model to explore relationships among pedestrians. We recursively extract social representations supervised by group-based annotations and formulate them into a social behavior graph, called Recursive Social Behavior Graph. Our recursive mechanism explores the representation power largely. Graph Convolutional Neural Network then is used to propagate social interaction information in such a graph. With the guidance of Recursive Social Behavior Graph, we surpass state-of-the-art methods on ETH and UCY dataset for 11.1% in ADE and 10.8% in FDE in average, and successfully predict complex social behaviors. Jianhua Sun 0003, Qinhong Jiang, Cewu Lu |
CVPR | 2 |
| 2020 | Dynamic and Static Context-Aware LSTM for Multi-agent Motion Prediction
Chaofan Tao, Qinhong Jiang, Lixin Duan, Ping Luo 0002 |
ECCV (21) | 2 |