VLDB 2026 Research / reviewers in the wild / expert
Wenxuan Zhu
dblp:223/7913
· DBLP profile ↗
14ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FreqMamba: Frequency-aware multi-graph fusion with Mamba for traffic flow prediction
Xiaohang Zhao, Haiyang Chi, Qingwang Wang, Sunyan Hong, Wenxuan Zhu, Lixue Liu, Bidong Chen, Yirong Zhu |
Knowl. Based Syst. | 8 |
| 2025 | 4D-Bench: Benchmarking Multi-Modal Large Language Models for 4D Object Understanding
Wenxuan Zhu, Bing Li 0024, Cheng Zheng 0002, Jinjie Mai, Jun Chen 0021, Letian Jiang, Abdullah Hamdi, Sara Rojas Martinez, Chia-Wen Lin, Mohamed Elhoseiny 0001, Bernard Ghanem |
ICCV | 1 |
| 2025 | Physical Adversarial Patch Attack for Optical Fine-Grained Aircraft RecognitionabstractDeep neural networks (DNNs) have been widely used in remote sensing but demonstrated to be sensitive with adversarial examples. By introducing carefully designed perturbations to clean images, DNNs can be led to incorrect predictions. Adversarial patch is commonly used to conduct adversarial attack, where traditional methods optimize its content and position separately, neglecting the coupling relation of two factors. In this paper, we propose a black-box attack framework targeting fine-grained aircraft recognition, named PatchGen, simultaneously optimizing both content and position of physical adversarial patches. For the requirements of physical attack, we further constrain the patch in object region and utilize elaborate criteria to evaluate its naturalness to alleviate the distortion when applying the patch in real world. We comprehensively validate our method in fine-grained aircraft classification, extending to object detection subsequently. Extensive experiments demonstrate that the proposed method achieves superior attack performance efficiently for classification and detection tasks in digital domain. Moreover, we validate the effectiveness of the adversarial patch under diverse circumstances in the physical world and prove that our method can be applied to different models as well as various domains. Ke Li 0024, Di Wang 0011, Wenxuan Zhu, Quan Wang 0006, Xinbo Gao 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Unleashing Channel Potential: Space-Frequency Selection Convolution for SAR Object DetectionabstractDeep Convolutional Neural Networks (DCNNs) have achieved remarkable performance in synthetic aperture radar (SAR) object detection, but this comes at the cost of tremendous computational resources, partly due to extracting redundant features within a single convolutional layer. Recent works either delve into model compression methods or focus on the carefully-designed lightweight models, both of which result in performance degradation. In this paper, we propose an efficient convolution module for SAR object detection, called SFS-Conv, which increases feature diversity within each convolutional layer through a shunt-perceive-select strategy. Specifically, we shunt input feature maps into space and frequency aspects. The former perceives the context of various objects by dynamically adjusting receptive field, while the latter captures abundant frequency variations and textural features via fractional Gabor transformer. To adaptively fuse features from space and frequency aspects, a parameter-free feature selection module is proposed to ensure that the most representative and distinctive information are preserved. With SFS-Conv, we build a lightweight SAR object detection network, called SFS-CNet. Experimental results show that SFS-CNet outperforms state-of-the-art (SoTA) models on a series of SAR object detection benchmarks, while simultaneously reducing both the model size and computational cost. Ke Li 0024, Di Wang 0011, Zhangyuan Hu, Wenxuan Zhu, Quan Wang 0006 |
CVPR | 4 |
| 2024 | TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks
Jinjie Mai, Wenxuan Zhu, Sara Rojas 0001, Jesus Zarzar, Abdullah Hamdi, Guocheng Qian, Bing Li 0024, Silvio Giancola, Bernard Ghanem |
ECCV (12) | 2 |
| 2024 | Transferable Physical Adversarial Patch Attack for Remote Sensing Object DetectionabstractDeep neural networks (DNNs) have been widely used in remote sensing but demonstrated to be vulnerable with adversarial examples. By adding elaborately designed perturbations on the clean images, DNNs may output wrong prediction. Research on adversarial attack contributes to the study of model robustness. However, previous methods mainly focus on white-box scenario or digital domain for classification tasks, while the vulnerability of remote sensing detectors has not been fully explored. Aiming at attacking black-box remote sensing detectors in physical domain, we propose to generate a transferable physical adversarial patch (TPAP) as the perturbations. Specifically, the initial patch is optimized by a U-Net and modified by the plane mask and position mask before applied to the clean image. By attacking a surrogate model, TPAP can be transferred to the target model. Abundant experimental results validate the attack ability of TPAP and evaluate the robustness of current one-stage detectors. Di Wang 0011, Wenxuan Zhu, Ke Li 0024, Pengfei Yang 0001 |
IGARSS | 2 |
| 2024 | Exploring Learngene via Stage-wise Weight Sharing for Initializing Variable-sized Models
Shiyu Xia, Wenxuan Zhu, Xu Yang 0021, Xin Geng 0001 |
IJCAI | 2 |
| 2024 | Vivid-ZOO: Multi-View Video Generation with Diffusion ModelabstractWhile diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of massive captioned multi-view videos and the complexity of modeling such multi-dimensional distribution. To this end, we propose a novel diffusion-based pipeline that generates high-quality multi-view videos centered around a dynamic 3D object from text. Specifically, we factor the T2MVid problem into viewpoint-space and time components. Such factorization allows us to combine and reuse layers of advanced pre-trained multi-view image and 2D video diffusion models to ensure multi-view consistency as well as temporal coherence for the generated multi-view videos, largely reducing the training cost. We further introduce alignment modules to align the latent spaces of layers from the pre-trained multi-view and the 2D video diffusion models, addressing the reused layers' incompatibility that arises from the domain gap between 2D and multi-view data. In support of this and future research, we further contribute a captioned multi-view video dataset. Experimental results demonstrate that our method generates high-quality multi-view videos, exhibiting vivid motions, temporal coherence, and multi-view consistency, given a variety of text prompts. Bing Li 0024, Cheng Zheng 0002, Wenxuan Zhu, Jinjie Mai, Biao Zhang 0005, Peter Wonka, Bernard Ghanem |
NeurIPS | 3 |
| 2024 | A fragmentation-aware redundancy elimination scheme for inline backup systems
Wenxuan Zhu, Dan Feng 0001, Wei Huang 0013, Nan Jiang 0013, Meng Chen 0024, Renxin Xia |
Future Gener. Comput. Syst. | 2 |
| 2024 | DiagSWin: A multi-scale vision transformer with diagonal-shaped windows for object detection and segmentation
Ke Li 0024, Di Wang 0011, Gang Liu 0006, Wenxuan Zhu, Haodi Zhong, Quan Wang 0006 |
Neural Networks | 4 |
| 2024 | GR-GAN: A unified adversarial framework for single image glare removal and denoising
Cong Niu, Ke Li 0024, Di Wang 0011, Wenxuan Zhu, Jinhui Dong |
Pattern Recognit. | 4 |
| 2023 | Causal Deep Reinforcement Learning Using Observational DataabstractDeep reinforcement learning (DRL) requires the collection of interventional data, which is sometimes expensive and even unethical in the real world, such as in the autonomous driving and the medical field. Offline reinforcement learning promises to alleviate this issue by exploiting the vast amount of observational data available in the real world. However, observational data may mislead the learning agent to undesirable outcomes if the behavior policy that generates the data depends on unobserved random variables (i.e., confounders). In this paper, we propose two deconfounding methods in DRL to address this problem. The methods first calculate the importance degree of different samples based on the causal inference technique, and then adjust the impact of different samples on the loss function by reweighting or resampling the offline dataset to ensure its unbiasedness. These deconfounding methods can be flexibly combined with existing model-free DRL algorithms such as soft actor-critic and deep Q-learning, provided that a weak condition can be satisfied by the loss functions of these algorithms. We prove the effectiveness of our deconfounding methods and validate them experimentally. Wenxuan Zhu, Chao Yu 0004, Qiang Zhang 0008 |
IJCAI | 1 |
| 2020 | A Preliminary Study of Fusion ARTs with Adaptively Information Intensity Attenuation ControllingabstractFusion ART is an enhanced version of Adaptive Resonance Theory (ART) which is derived from a biologically-plausible theory of human cognitive information processing. Due to its well-established ability of learning associative mappings across multimodal pattern channels in an online and incremental manner, fusion ART has been widely applied in many real world learning problems. In this paper, we take a Fusion Architecture for Learning, Cognition, and Navigation (FALCON) as the specification and essential backbone of fusion ART and introduce an intensity attenuation controller δ for adaptively adjusting the intensity of information captured from the environment, by taking inspiration from Broadbent-Treisman Filter-Attenuation's perceptual model of environmental attention. Particularly, we propose both an adaptive δ detection algorithm as well as a δ-based pruning algorithm to enhance the learning performance of FALCON while reduce the redundant memory storage incurred by the "detrimental δ". To verify the effectiveness and efficiency of our proposed method, comprehensive experimental studies are carried out on a classical minefield navigation task. Wenxuan Zhu, Yaqing Hou, Qiang Zhang 0008, Hong-Wei Ge, Xin Yang 0011, Liang Feng 0001, Xinghua Qu |
IJCNN | 1 |
| 2018 | Adaptively Shaping Reinforcement Learning Agents via Human Reward
Chao Yu 0004, Tianpei Yang, Wenxuan Zhu, Yuchen Li 0006, Hong-Wei Ge, Jiankang Ren |
PRICAI (1) | 4 |