VLDB 2026 Research / reviewers in the wild / expert
Chenxi Du
dblp:259/7223
· DBLP profile ↗
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Redundancy in Diffusion Transformers: A Temporal-Spatial Joint Caching Strategy for Efficient SamplingabstractDiffusion Transformers (DiTs) achieve impressive generative performance but suffer from significant inference latency. Feature caching–based acceleration methods reduce total computation by reusing results from earlier timesteps, but they largely ignore that temporal redundancy is dynamic and inconsistent across timesteps. Our analysis reveals this variability. More crucially, we identify a previously underexplored form of efficiency, namely spatial redundancy, characterized by high similarity between adjacent transformer blocks within the same timestep. Motivated by this dual-dimensional redundancy, we propose Temporal-Spatial Joint Cache, a training-free inference acceleration strategy that dynamically determines optimal reuse operations across temporal and spatial dimensions. Our approach features a redundancy-guided operation selector that estimates local feature stability using second-order divided differences, enabling fine-grained decisions between full computation, temporal cache, and spatial cache. Furthermore, we use interpolation-based feature prediction to capture local feature evolution for more accurate reuse. In addition, we propose a bounded cache distance control mechanism to mitigate error accumulation from excessive reuse. Together, these components allow our method to deliver substantial inference speedups without retraining or compromising generation fidelity, offering a new perspective on efficiency in diffusion transformer inference. Chenxi Du, Yongheng Deng, Ju Ren 0001, Yaoxue Zhang |
KDD (1) | 1 |
| 2026 | FB-4D: Spatial-Temporal Coherent Dynamic 3D Content Generation with Feature BanksabstractWith the rapid advancements in diffusion models and 3D generation techniques, dynamic 3D content generation has become a crucial research area. However, achieving high-fidelity 4D (dynamic 3D) generation with strong spatial-temporal consistency remains a challenging task. Inspired by recent findings that pretrained diffusion features capture rich correspondences, we propose FB-4D, a novel 4D generation framework that integrates a Feature Bank mechanism to enhance both spatial and temporal consistency in generated frames. In FB-4D, we store features extracted from previous frames and fuse them into the process of generating subsequent frames, ensuring consistent characteristics across both time and multiple views. To ensure a compact representation, the Feature Bank is updated by a proposed dynamic merging mechanism. Leveraging this Feature Bank, we demonstrate for the first time that generating additional reference sequences through multiple autoregressive iterations can continuously improve generation performance. Experimental results show that FB-4D significantly outperforms existing methods in terms of rendering quality, spatial-temporal consistency, and robustness. It surpasses all multi-view generation tuning-free approaches by a large margin and achieves performance on par with training-based methods. Our code and data will be publicly available to support future research. Huan-ang Gao, Wenyi Li 0001, Haohan Chi, Chenxi Du, Yiqian Liu, Mingju Gao, Guiyu Zhang, Zongzheng Zhang, Li Yi 0001, Hongyang Li 0001, Hao Zhao 0002 |
WACV | 6 |
| 2026 | Symmetrical Semantic-Visual Refinement for Cross-Domain Iris Presentation Attack Detection
Botong Li, Kaiyue Shi, Chenxi Du, Hang Zou 0002, Hui Zhang 0061 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Toward Generalized Iris Presentation Attack Detection: A Mask-and-Distill Mixture of Experts ApproachabstractIris Presentation Attack Detection (PAD) is critical for securing recognition systems, yet its practical deployment is severely hindered by the poor generalization of models across different acquisition devices and diverse datasets. To address this persistent cross-domain challenge, we first introduce a comprehensive evaluation framework, the Iris Presentation Attack Detection Cross-Domain-Testing (IPAD-CDT) Protocol, designed to evaluate the model robustness in these scenarios. Our core contribution is a novel Masked Mixture-of-Experts (MMoE) method, which enhances the generalization of Transformer-based architectures. MMoE introduces a structured information asymmetry, where "student" Experts learn robust features from masked inputs by distilling knowledge from an unmasked "teacher" Expert via a cosine distance loss. This mask-and-distill mechanism effectively mitigates overfitting and guides the model to learn domain-invariant cues. By integrating MMoE into a CLIP-based model, we conduct extensive experiments on our IPAD-CDT protocol. The results demonstrate that our method sets a new state-of-the-art, significantly outperforming existing models, especially in the challenging cross-dataset and cross-device settings. Hang Zou 0002, Chenxi Du, Ajian Liu 0001, Yuan Zhang 0023, Jing Liu 0062, Jun Wan 0001, Hui Zhang 0061, Zhenan Sun |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | La-SoftMoE CLIP for Unified Physical-Digital Face Attack DetectionabstractFacial recognition systems are susceptible to both physical and digital attacks, posing significant security risks. Traditional approaches often treat these two attack types separately due to their distinct characteristics. Thus, when being combined attacked, almost all methods could not deal. Some studies attempt to combine the sparse data from both types of attacks into a single dataset and try to find a common feature space, which is often impractical due to the space is difficult to be found or even non-existent. To overcome these challenges, we propose a novel approach that uses the sparse model to handle sparse data, utilizing different parameter groups to process distinct regions of the sparse feature space. Specifically, we employ the Mixture of Experts (MoE) framework in our model, expert parameters are matched to tokens with varying weights during training and adaptively activated during testing. However, the traditional MoE struggles with the complex and irregular classification boundaries of this problem. Thus, we introduce a flexible self-adapting weighting mechanism, enabling the model to better fit and adapt. In this paper, we proposed La-SoftMoE CLIP, which allows for more flexible adaptation to the Unified Attack Detection (UAD) task, significantly enhancing the model’s capability to handle diversity attacks. Experiment results demonstrate that our proposed method has SOTA performance. Hang Zou 0002, Chenxi Du, Hui Zhang 0061, Yuan Zhang 0023, Ajian Liu 0001, Jun Wan 0001, Zhen Lei 0001 |
IJCB | 2 |
| 2024 | Unsupervised Domain Adaptation for Cross-Device Iris Liveness Detection Model Transfer
Xiuying Wu, Chenxi Du, Hui Zhang 0061, Jing Liu 0062, Hang Zou 0002 |
ICPR (28) | 2 |