VLDB 2026 Research / reviewers in the wild / expert
Wenshuo Chen
dblp:193/3404
· DBLP profile ↗
11ranked-venue papers
3as first author
9since 2021 · last 2025
0009-0002-1966-6059ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DCTdiff: Intriguing Properties of Image Generative Modeling in the DCT SpaceabstractThis paper explores image modeling from the frequency space and introduces DCTdiff, an end-to-end diffusion generative paradigm that efficiently models images in the discrete cosine transform (DCT) space. We investigate the design space of DCTdiff and reveal the key design factors. Experiments on different frameworks (UViT, DiT), generation tasks, and various diffusion samplers demonstrate that DCTdiff outperforms pixel-based diffusion models regarding generative quality and training efficiency. Remarkably, DCTdiff can seamlessly scale up to 512$\times$512 resolution without using the latent diffusion paradigm and beats latent diffusion (using SD-VAE) with only 1/4 training cost. Finally, we illustrate several intriguing properties of DCT image modeling. For example, we provide a theoretical proof of why `image diffusion can be seen as spectral autoregression', bridging the gap between diffusion and autoregressive models. The effectiveness of DCTdiff and the introduced properties suggest a promising direction for image modeling in the frequency space. The code is at https://github.com/forever208/DCTdiff. Mang Ning, Mingxiao Li 0002, Jianlin Su, Haozhe Jia, Lanmiao Liu, Martin Benes 0001, Wenshuo Chen, Albert Ali Salah, Itir Önal |
ICML | 7 |
| 2025 | ANT: Adaptive Neural Temporal-Aware Text-to-Motion ModelabstractWhile diffusion models advance text-to-motion generation, their static semantic conditioning ignores temporal-frequency demands: early denoising requires structural semantics for motion foundations while later stages need localized details for text alignment. This mismatch mirrors biological morphogenesis where developmental phases demand distinct genetic programs. Inspired by epigenetic regulation governing morphological specialization, we propose **(ANT)**, an **A**daptive **N**eural **T**emporal-Aware architecture. ANT orchestrates semantic granularity through: **(i) Semantic Temporally Adaptive (STA) Module:** Automatically partitions denoising into low-frequency structural planning and high-frequency refinement via spectral analysis. **(ii) Dynamic Classifier-Free Guidance scheduling (DCFG):** Adaptively adjusts conditional to unconditional ratio enhancing efficiency while maintaining fidelity. Extensive experiments show that ANT can be applied to various baselines, significantly improving model performance, and achieving state-of-the-art semantic alignment on StableMoFusion. Wenshuo Chen, Kuimou Yu, Haozhe Jia, Kaishen Yuan, Zexu Huang, Songning Lai, Hongru Xiao, Erhang Zhang, Lei Wang 0108, Yutao Yue |
ACM Multimedia | 1 |
| 2025 | Physics-Informed Representation Alignment for Sparse Radio-Map ReconstructionabstractWith the rapid development of wireless communication technology, the efficient utilization of spectrum resources, optimization of communication quality, and intelligent communication have become critical. Radio map reconstruction is essential for enabling advanced applications, yet challenges such as complex signal propagation and sparse observational data hinder accurate reconstruction in practical scenarios. Existing methods often fail to align physical constraints with data-driven features, particularly under sparse measurement conditions. To address these issues, we propose Physics-Aligned Radio Map Diffusion Model (PhyRMDM), a novel framework that establishes cross-domain representation alignment between physical principles and neural network features through dual learning pathways. The proposed model integrates Physics-Informed Neural Networks (PINNs) with a representation alignment mechanism that explicitly enforces consistency between Helmholtz equation constraints and environmental propagation patterns. Our architecture employs two synergistic U-Nets: the first ensures physical consistency by minimizing PDE residuals and boundary conditions through latent space alignment, while the second refines predictions via diffusion-based denoising with attention-guided feature fusion. This dual alignment strategy enables simultaneous satisfaction of wave propagation laws and data distribution characteristics. Experimental results demonstrate significant improvements over state-of-the-art methods, achieving NMSE of 0.0031 and RMSE of 0.0125 under Static Radio Map (SRM) conditions, and NMSE of 0.0047 with RMSE of 0.0146 in Dynamic Radio Map (DRM) scenarios. The proposed representation alignment paradigm provides 37.2% accuracy enhancement in ultra-sparse cases (1% sampling rate), confirming its effectiveness in bridging physics-based modeling and deep learning for radio map reconstruction. These advancements establish a new framework for sparse signal environment characterization, with direct applications in 5G/6G network optimization and intelligent spectrum management. The code can be found on the website: https://github.com/Hxxxz0/RMDM Haozhe Jia, Wenshuo Chen, Lei Wang 0108, Hongru Xiao, Nanqian Jia, Keming Wu, Songning Lai, Yutao Yue |
ACM Multimedia | 2 |
| 2025 | From Guesswork to Guarantee: Towards Faithful Multimedia Web Forecasting with TimeSieveabstractThe domain of time series forecasting has gained significant attention due to its critical applications in multimedia-rich web traffic (including video streaming workloads and dynamic content delivery) and cross-platform advertisement click predictions, which are essential for web operations planning. While models like TimeSieve have demonstrated strong capabilities in predicting web visitation metrics, they suffer from critical unfaithfulness issues, including sensitivity to random seeds, input noise, layer noise, and parametric perturbations. To address these limitations, we propose Faithful TimeSieve (FTS), an enhanced framework designed to improve prediction reliability and robustness. Our approach systematically detects and mitigates unfaithfulness in TimeSieve, significantly enhancing its stability and consistency. Experimental results demonstrate that FTS substantially improves the model's faithfulness, setting a new standard for temporal forecasting methods. This advancement not only increases TimeSieve's reliability but also contributes to more robust temporal modeling, particularly crucial for web traffic forecasting where prediction accuracy directly impacts operational decisions. Our work thus represents a significant step toward more dependable time series predictions in web-related applications. Songning Lai, Ninghui Feng, Jiechao Gao, Hao Wang 0220, Haochen Sui, Xin Zou 0001, Wenshuo Chen, Lijie Hu, Hang Zhao 0010, Xuming Hu, Yutao Yue |
ACM Multimedia | 8 |
| 2025 | Learning New Concepts, Remembering the Old: Continual Learning for Multimodal Concept Bottleneck ModelsabstractConcept Bottleneck Models (CBMs) enhance the interpretability of AI systems, particularly by bridging visual input with human-understandable concepts, effectively acting as a form of multimodal interpretability model. However, existing CBMs typically assume static datasets, which fundamentally limits their adaptability to real-world, continuously evolving multimodal data streams. To address this, we define a novel continual learning task for CBMs: simultaneously handling concept-incremental and class-incremental learning. This task requires models to continuously acquire new concepts (often representing cross-modal attributes) and classes while robustly preserving previously learned knowledge. To tackle this challenging problem, we propose CONceptual Continual Incremental Learning (CONCIL), a novel framework that fundamentally re-imagines concept and decision layer updates as linear regression problems. This reformulation eliminates the need for gradient-based optimization, thereby effectively preventing catastrophic forgetting. Crucially, CONCIL relies solely on recursive matrix operations, rendering it highly computationally efficient and well-suited for real-time and large-scale multimodal data applications. Experimental results compellingly demonstrate that CONCIL achieves ''absolute knowledge memory'' and significantly surpasses the performance of traditional CBM methods in both concept- and class-incremental settings, thus establishing a new paradigm for continual learning in CBMs, particularly valuable for dynamic multimodal understanding. Songning Lai, Mingqian Liao, Zhangyi Hu, Wenshuo Chen, Hongru Xiao, Jianheng Tang 0001, Haicheng Liao, Yutao Yue |
ACM Multimedia | 5 |
| 2025 | Text2Weight: Bridging Natural Language and Neural Network Weight SpacesabstractHow far are we really from automatically generating neural networks? While neural network weight generation shows promise, current approaches struggle with generalization to unseen tasks and practical application exploration. To address this, we propose T2W, a diffusion transformer framework that generates task-specific weights conditioned on natural language descriptions. T2W hierarchically processes network parameters into uniform blocks, integrates text embeddings from CLIP via a prior attention mechanism, and employs adversarial training with weight-space augmentation to enhance generalization. Experiments on Cifar100, Caltech256, and TinyImageNet demonstrate T2W's ability to produce high-quality weights for unseen tasks, outperforming optimization-based initialization and enabling novel applications such as weight enhancement and text-guided model fusion. Our work bridges textual semantics with weight-space dynamics, supported by an open-source dataset of text-weight pairs, advancing the practicality of generative models in neural network parameter synthesis. Our code is available on https://github.com/TianSuya/T2W. Wenshuo Chen, Zexi Li 0001, Songning Lai, Jiemin Wu, Yutao Yue |
ACM Multimedia | 2 |
| 2025 | Can Audio Language Models Listen Between the Lines? A Study on Metaphorical Reasoning via UnspokenabstractRecent advancements in Audio Language Models (ALMs) have led to significant improvements in speech-related tasks. However, their capacity for profound metaphorical reasoning, especially when derived from audio-specific cues, has yet to be thoroughly investigated. To address this gap, we introduce Unspoken, a bilingual (Chinese-English) question answering benchmark designed to assess ALMs' comprehension of non-literal, metaphor-rich audio. Unlike prior text-centric evaluations, Unspoken emphasizes prosody, phonetic ambiguity, emotional inflection, and other nuanced acoustic features critical to metaphor understanding but often lost in transcription. We construct a high-quality dataset of 2,764 manually curated and validated QA pairs, spanning three reasoning dimensions: semantic, acoustic, and contextual, and covering six common types of metaphors. Evaluation across 23 mainstream ALMs reveals a substantial performance gap: the best model achieves only 69.5% accuracy, significantly below the human average of 81.1%. By analyzing the error patterns, we identify five key failure modes that reveal fundamental limitations in current models' reasoning capabilities. Unspoken not only sets a new standard for evaluating metaphorical reasoning in audio but also pioneers a novel research direction that moves beyond transcription-based assessments. Grounding metaphor understanding in authentic human communication scenarios offers deep insight for developing more cognitively capable ALMs. The data and codes are available at https://github.com/Hongru0306/UNSPOKEN. Hongru Xiao, Xiang Li 0064, Duyi Pan, ZhixueSong ZhixueSong, Jiale Han 0001, Songning Lai, Wenshuo Chen, Benyou Wang |
ACM Multimedia | 8 |
| 2024 | SATO: Stable Text-to-Motion FrameworkabstractIs the Text to Motion model robust? Recent advancements in Text to Motion models primarily stem from more accurate predictions of specific actions. However, the text modality typically relies solely on pre-trained Contrastive Language-Image Pretraining (CLIP) models. Our research has uncovered a significant issue with the text-tomotion model: its predictions often exhibit inconsistent outputs, resulting in vastly different or even incorrect poses when presented with semantically similar or identical text inputs. In this paper, we undertake an analysis to elucidate the underlying causes of this instability, establishing a clear link between the unpredictability of model outputs and the erratic attention patterns of the text encoder module. Consequently, we introduce a formal framework aimed at addressing this issue, which we term the Stable Text-to-Motion Framework (SATO). SATO consists of three modules, each dedicated to stable attention, stable prediction, and maintaining a balance between accuracy and robustness trade-off. We present a methodology for constructing an SATO that satisfies the stability of attention and prediction. To verify the stability of the model, we introduced a new textual synonym perturbation dataset based on HumanML3D and KIT-ML. Results show that SATO is significantly more stable against synonyms and other slight perturbations while keeping its high accuracy performance. Codes and models are released at Wenshuo Chen, Hongru Xiao, Erhang Zhang, Lijie Hu, Lei Wang 0108, Mengyuan Liu 0004, Chen Chen 0001 |
ACM Multimedia | 1 |
| 2024 | Towards Multi-dimensional Explanation Alignment for Medical ClassificationabstractThe lack of interpretability in the field of medical image analysis has significant ethical and legal implications. Existing interpretable methods in this domain encounter several challenges, including dependency on specific models, difficulties in understanding and visualization, and issues related to efficiency. To address these limitations, we propose a novel framework called Med-MICN (Medical Multi-dimensional Interpretable Concept Network). Med-MICN provides interpretability alignment for various angles, including neural symbolic reasoning, concept semantics, and saliency maps, which are superior to current interpretable methods. Its advantages include high prediction accuracy, interpretability across multiple dimensions, and automation through an end-to-end concept labeling process that reduces the need for extensive human training effort when working with new datasets. To demonstrate the effectiveness and interpretability of Med-MICN, we apply it to four benchmark datasets and compare it with baselines. The results clearly demonstrate the superior performance and interpretability of our Med-MICN. Lijie Hu, Songning Lai, Wenshuo Chen, Hongru Xiao, Jingfeng Zhang, Di Wang 0015 |
NeurIPS | 3 |
| 2018 | Real-Time Scalable Visual Tracking via Quadrangle Kernelized Correlation FiltersabstractCorrelation filter (CF) has been widely used in tracking tasks due to its simplicity and high efficiency. However, conventional CF-based trackers fail to handle the scale variation that occurs when the targeted object is moving, which is one of the most notable unsolved problems of visual object tracking. In this paper, we propose a scalable visual tracking algorithm based on kernelized correlation filters, referred to as quadrangle kernelized correlation filters (QKCF). Unlike existing complicated scalable trackers that either perform the correlation filtering operation multiple times or extract many candidate windows at various scales, our tracker intends to estimate the scale of the object based on the positions of its four corners, which can be detected using a new Gaussian training output matrix within one filtering process. After obtaining four peak values corresponding to the four corners, we measure the detection confidence of each part response by evaluating its spatial and temporal smoothness. On top of it, a weighted Bayesian inference framework is employed to estimate the final location and size of the bounding box from the response matrix, where the weights are synchronized with the calculated detection likelihoods. Experiments are performed on the OTB-100 data set and 16 benchmark sequences with significant scale variations. The results demonstrate the superiority of the proposed method in terms of both effectiveness and robustness, compared with the state-of-the-art methods. Guiguang Ding, Wenshuo Chen, Sicheng Zhao, Jungong Han, Qiaoyan Liu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | Accelerated Manhattan hashing via bit-remapping with location information
Wenshuo Chen, Guiguang Ding, Zijia Lin, Iyad Jafar, Jisheng Pei |
Multim. Tools Appl. | 1 |