VLDB 2026 Research / reviewers in the wild / expert
Wanyue Zhang
dblp:260/9519
· DBLP profile ↗
13ranked-venue papers
6as first author
12since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated ObjectsabstractWe present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-based contact maps conditioned on the object trajectory with an articulation-aware feature representation, revealing rich bimanual patterns for manipulation. The learned contact prior is then used to guide our hand motion generator, producing diverse and realistic bimanual motions for object movement and articulation. Our work offers key insights into feature representation and contact prior for articulated objects, demonstrating their effectiveness in taming the complex, high-dimensional space of bimanual hand-object interactions. Through comprehensive quantitative experiments, we demonstrate a clear step towards simplified and high-quality hand-object animations that surpass the state of the art in motion quality and diversity. Project page: https://vcai.mpi-inf.mpg.de/projects/bimart/. Wanyue Zhang, Rishabh Dabral, Vladislav Golyanik, Vasileios Choutas, Eduardo Alvarado, Thabo Beeler, Marc Habermann, Christian Theobalt |
CVPR | 1 |
| 2025 | Beyond One-Size-Fits-All: Adaptive Fine-Tuning for LLMs Based on Data Inherent Heterogeneity
Wanyue Zhang, Yangyifan Xu, Shuo Ren 0002, Jiajun Zhang 0001 |
NLPCC (1) | 1 |
| 2025 | Channel Estimation Based on Block Sparsity in Wavenumber-Domain for Holographic MIMO SystemsabstractThis paper investigates channel estimation for holographic MIMO (HMIMO) systems that achieve continuous electromagnetic aperture synthesis through ultra-dense configurations of antenna elements within constrained spatial dimensions. We aim to exploit the inherent sparsity characteristics of the HMIMO generated by ultra-dense antenna configurations, thereby reducing the error and complexity caused by the large number of antennas in channel estimation process. Unlike previous algorithms that rely solely on sparsity, the proposed algorithm considers the block sparse structure of the channel, thereby yielding more accurate channel estimation performances. Specifically, we first establish the block sparse representation in the wavenumber domain of the channel for HMIMO systems. Subsequently, the channel estimation problem is formulated as the block sparse matrix recovery problem, incorporating an l2,1norm term in the objective function to promote block sparse structure. The block sparse-based gradient descent (BSBGD) algorithm is proposed to solve the formulated problem. Simulation results show that channel estimation results obtained by the proposed algorithm are more accurate. Fengxi Gu, Chen Liu 0005, Yunchao Song, Wanyue Zhang |
VTC2025-Fall | 4 |
| 2024 | ROAM: Robust and Object-Aware Motion Generation Using Neural Pose DescriptorsabstractExisting automatic approaches for 3D virtual character motion synthesis supporting scene interactions do not gen-eralise well to new objects outside training distributions, even when trained on extensive motion capture datasets with diverse objects and annotated interactions. This paper addresses this limitation and shows that robustness and generalisation to novel scene objects in 3D object-aware character synthesis can be achieved by training a motion model with as few as one reference object. We leverage an implicit feature representation trained on object-only datasets, which encodes an SE(3)-equivariant descriptor field around the object. Given an unseen object and a reference pose-object pair, we optimise for the object-aware pose that is closest in the feature space to the reference pose. Finally, we use l-NSM, i.e., our motion generation model that is trained to seamlessly transition from locomotion to object interaction with the proposed bidirectional pose blending scheme. Through comprehensive numerical comparisons to state-of-the-art methods and in a user study, we demonstrate substantial improvements in 3D virtual character motion and interaction quality and robustness to scenarios with unseen objects. Our project page is available at https://vcai.mpi-inf.mpg.de/projects/ROAM/. Wanyue Zhang, Rishabh Dabral, Thomas Leimkühler, Vladislav Golyanik, Marc Habermann, Christian Theobalt |
3DV | 1 |
| 2024 | Hybrid-Field Channel Estimation for XL-MIMO: A Proximal Gradient Algorithm on the Fixed-Rank Matrix ManifoldabstractThis paper investigates hybrid-field channel estimation for extremely large-scale MIMO (XL-MIMO) systems. Different from the previous algorithms based on sparsity, the proposed algorithm considers both the low-rank structure and the sparsity of the channel, leading to a more accurate estimation performance. Particularly, we formulate the estimation problem of the hybrid-field channel as a sparse matrix recovery problem with a fixed-rank constraint. However, this problem encounters two challenges: the non-smooth$\ell_{1}-\mathbf{norm}$term and the non-convex fixed-rank matrix constraint. To address these challenges, we employ the proximal operator to handle the non-smooth$\ell_{1}-\mathbf{norm}$term and consider the set of matrices with a fixed rank as the fixed-rank matrix manifold. Then a fixed-rank matrix manifold-based proximal gradient (FRM-PG) algorithm is proposed to smoothly search the solution on the fixed-rank matrix manifold, where the descent direction is derived in each iteration. Simulation results demonstrate that the proposed algorithm outperforms the classical channel estimation algorithms. Wanyue Zhang, Yunchao Song, Chen Liu 0005, Mujun Qian |
ICC | 1 |
| 2024 | Improved Design of Resource Hopping Based Multiple Access for Grant-Free Random Access in 6G mMTC SystemabstractIn order to satisfy the increasingly massive connection in mMTC system, multiple access technology is a key enabler in the future 6G mMTC. Recently, an emerging multiple access scheme named resource hopping based multiple access (RHMA) has been proposed to achieve reliable user identification and data detection in grant-free random access for mMTC. However, the collision resolution of RHMA is still limited for the future 6G mMTC requirements. Therefore, an improved design is proposed in this paper to enhance the collision resolution capability of RHMA. Specifically, successive interference cancellation (SIC) is combined with user identification and segment decoding at the receiver of RHMA. Also, the user identification of RHMA is improved to eliminate the false alarm user caused by collision and blind channel estimation is considered to recover the signal on the colliding segments. The simulation results show that the improved design is able to achieve a higher collision resolution capability of RHMA. Wanyue Zhang, Guangkai Li, Yiyan Ma, Wanqiao Wang, Botao Feng, Bo Ai 0001 |
VTC Spring | 1 |
| 2024 | Revisiting pretraining for semi-supervised learning in the low-label regime
Xun Xu 0002, Jingyi Liao, Lile Cai, Kangkang Lu 0001, Wanyue Zhang, Yasin Yazici, Chuan-Sheng Foo |
Neurocomputing | 6 |
| 2024 | Unsupervised Pansharpening Based on Double-Cycle ConsistencyabstractMultispectral (MS) pansharpening can improve the spatial resolution of MS images by fusing panchromatic (PAN) images, which have important applications in the fields of smart agriculture and environmental monitoring. However, existing supervised algorithms treat the original MS images as ground truth and generate training data under Wald’s protocol, resulting in a gap between the learned degradation process of the model and reality. This leads to the model having poor generalization and impractical. Unsupervised pansharpening methods often struggle to fully explore the rich information contained in images, leading to suboptimal pansharpening outcomes. In this work, we propose an unsupervised pansharpening algorithm based on double-cycle consistency that can learn directly from the original MS images without relying on artificially simulated degradation processes. Specifically, the network with cross-domain correlation information interaction is developed to achieve a deep fusion of spatial and spectral features. To address the inaccurate degradation mechanism representation of MS images, a spatial information extraction module based on scale invariance is developed to achieve an accurate representation. Meanwhile, double-cycle consistency loss is proposed to reduce the information loss caused by simulated degradation during the cycle process. Experimental results show that this method outperforms existing unsupervised pansharpening methods in both quantitative and qualitative evaluation of full-resolution images. Lijun He 0001, Zhihan Ren 0002, Wanyue Zhang, Fan Li 0003, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation with Reliable Voted Pseudo LabelsabstractIn this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unla-beled data by scaling up or down point clouds and then predicting the scales. Second, to capture the local structure in a self-supervised manner, we propose to project a 3D local area onto a 2D plane and then learn to reconstruct the squeezed region. Moreover, to effectively transfer the knowledge from source domain, we propose to vote pseudo labels for target samples based on the labels of their nearest source neighbors in the shared feature space. To avoid the noise caused by incorrect pseudo labels, we only select re-liable target samples, whose voting consistencies are high enough, for enhancing adaptation. The voting method is able to adaptively select more and more target samples during training, which in return facilitates adaptation because the amount of labeled target data increases. Experiments on PointDA (ModelNet-10, ShapeNet-10 and ScanNet-10) and Sim-to-Real (ModelNet-11, ScanObjectNN-11, ShapeNet-9 and ScanObjectNN-9) demonstrate the effectiveness of our method. Hehe Fan, Xiaojun Chang, Wanyue Zhang, Ying Sun 0001, Mohan Kankanhalli |
CVPR | 3 |
| 2022 | Open-Set Semi-Supervised Learning for 3D Point Cloud UnderstandingabstractSemantic understanding of 3D point cloud relies on learning models with massively annotated data, which, in many cases, are expensive or difficult to collect. This has led to an emerging research interest in semi-supervised learning (SSL) for 3D point cloud. It is commonly assumed in SSL that the unlabeled data are drawn from the same distribution as that of the labeled ones; This assumption, however, rarely holds true in realistic environments. Blindly using out-of-distribution (OOD) unlabeled data could harm SSL performance. In this work, we propose to selectively utilize unlabeled data through sample weighting, so that only conducive unlabeled data would be prioritized. To estimate the weights, we adopt a bi-level optimization framework which iteratively optimizes a meta-objective on a held-out validation set and a task-objective on a training set. Faced with the instability of efficient bi-level optimizers, we further propose three regularization techniques to enhance the training stability. Extensive experiments on 3D point cloud classification and segmentation tasks verify the effectiveness of our proposed method. We also demonstrate the feasibility of a more efficient training strategy. Our code is released on Github1. Xian Shi, Xun Xu 0002, Wanyue Zhang, Xiatian Zhu, Chuan-Sheng Foo, Kui Jia |
ICPR | 3 |
| 2022 | Few-Shot Adaptation of Pre-Trained Networks for Domain ShiftabstractDeep networks are prone to performance degradation when there is a domain shift between the source (training) data and target (test) data. Recent test-time adaptation methods update batch normalization layers of pre-trained source models deployed in new target environments with streaming data. Although these methods can adapt on-the-fly without first collecting a large target domain dataset, their performance is dependent on streaming conditions such as mini-batch size and class-distribution which can be unpredictable in practice. In this work, we propose a framework for few-shot domain adaptation to address the practical challenges of data-efficient adaptation. Specifically, we propose a constrained optimization of feature normalization statistics in pre-trained source models supervised by a small target domain support set. Our method is easy to implement and improves source model performance with as little as one sample per class for classification tasks. Extensive experiments on 5 cross-domain classification and 4 semantic segmentation datasets show that our proposed method achieves more accurate and reliable performance than test-time adaptation, while not being constrained by streaming conditions. Wenyu Zhang 0003, Wanyue Zhang, Chuan-Sheng Foo |
IJCAI | 3 |
| 2021 | On Automatic Data Augmentation for 3D Point Cloud Classification
Wanyue Zhang, Xun Xu 0002, Fayao Liu, Le Zhang 0001, Chuan-Sheng Foo |
BMVC | 1 |
| 2020 | An LSTM Approach to Temporal 3D Object Detection in LiDAR Point Clouds
Wanyue Zhang, Abhijit Kundu, Caroline Pantofaru, David A. Ross, Thomas A. Funkhouser, Alireza Fathi |
ECCV (18) | 2 |