VLDB 2026 Research / reviewers in the wild / expert
Bo Wang 0019
dblp:72/6811-19
· DBLP profile ↗
18ranked-venue papers
2as first author
13since 2021 · last 2025
0000-0002-0127-2281ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature AdaptationabstractZero-shot human-object interaction (HOI) detection remains a challenging task, particularly in generalizing to unseen actions. Existing methods address this challenge by tapping Vision-Language Models (VLMs) to access knowledge beyond the training data. However, they either struggle to distinguish actions involving the same object or demonstrate limited generalization to unseen classes. In this paper, we introduce HOLa (Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation), a novel approach that both enhances generalization to unseen classes and improves action distinction. In training, HOLa decomposes VLM text features for given HOI classes via low-rank factorization, producing class-shared basis features and adaptable weights. These features and weights form a compact HOI representation that preserves shared information across classes, enhancing generalization to unseen classes. Subsequently, we refine action distinction by adapting weights for each HOI class and introducing human-object tokens to enrich visual interaction representations. To further distinguish unseen actions, we guide the weight adaptation with LLM-derived action regularization. Experimental results show that our method sets a new state-of-the-art across zero-shot HOI settings on HICO-DET, achieving an unseen-class mAP of 27.91 in the unseen-verb setting. Our code is available at https://github.com/ChelsieLei/HOLa. Qinqian Lei, Bo Wang 0019, Robby T. Tan |
ICCV | 2 |
| 2025 | 3DOT: Texture Transfer for 3DGS Objects from a Single Reference ImageabstractImage-based 3D texture transfer from a single 2D reference image enables practical customization of 3D object appearances with minimal manual effort.
Adapted 2D editing and text-driven 3D editing approaches can serve this purpose. However, 2D editing typically involves frame-by-frame manipulation, often resulting in inconsistencies across views, while text-driven 3D editing struggles to preserve texture characteristics from reference images.
To tackle these challenges, we introduce \textbf{3DOT}, a \textbf{3D} Gaussian Splatting \textbf{O}bject \textbf{T}exture Transfer method based on a single reference image, integrating: 1) progressive generation, 2) view-consistency gradient guidance, and 3) prompt-tuned gradient guidance.
To ensure view consistency, progressive generation starts by transferring texture from the reference image and gradually propagates it to adjacent views.
View-consistency gradient guidance further reinforces coherence by conditioning the generation model on feature differences between consistent and inconsistent outputs.
To preserve texture characteristics, prompt-tuning-based gradient guidance learns a token that describes differences between original and reference textures, guiding the transfer for faithful texture preservation across views.
Overall, 3DOT combines these strategies to achieve effective texture transfer while maintaining structural coherence across viewpoints.
Extensive qualitative and quantitative evaluations confirm that our three components enable convincing and effective 2D-to-3D texture transfer. Our project page is available here: https://massyzs.github.io/3DOT_web/. Xiao Cao, Beibei Lin, Bo Wang 0019, Robby T. Tan |
NeurIPS | 3 |
| 2024 | Few-Shot Learning from Augmented Label-Uncertain Queries in Bongard-HOIabstractDetecting human-object interactions (HOI) in a few-shot setting remains a challenge. Existing meta-learning methods struggle to extract representative features for classification due to the limited data, while existing few-shot HOI models rely on HOI text labels for classification. Moreover, some query images may display visual similarity to those outside their class, such as similar backgrounds between different HOI classes. This makes learning more challenging, especially with limited samples. Bongard-HOI epitomizes this HOI few-shot problem, making it the benchmark we focus on in this paper. In our proposed method, we introduce novel label-uncertain query augmentation techniques to enhance the diversity of the query inputs, aiming to distinguish the positive HOI class from the negative ones. As these augmented inputs may or may not have the same class label as the original inputs, their class label is unknown. Those belonging to a different class become hard samples due to their visual similarity to the original ones. Additionally, we introduce a novel pseudo-label generation technique that enables a mean teacher model to learn from the augmented label-uncertain inputs. We propose to augment the negative support set for the student model to enrich the semantic information, fostering diversity that challenges and enhances the student’s learning. Experimental results demonstrate that our method sets a new state-of-the-art (SOTA) performance by achieving 68.74% accuracy on the Bongard-HOI benchmark, a significant improvement over the existing SOTA of 66.59%. In our evaluation on HICO-FS, a more general few-shot recognition dataset, our method achieves 73.27% accuracy, outperforming the previous SOTA of 71.20% in the 5- way 5-shot task. Qinqian Lei, Bo Wang 0019, Robby T. Tan |
AAAI | 2 |
| 2024 | Domain-Adaptive 2D Human Pose Estimation via Dual Teachers in Extremely Low-Light Conditions
Yihao Ai, Bo Wang 0019, Yu Cheng 0009, Xinchao Wang, Robby T. Tan |
ECCV (47) | 3 |
| 2024 | Remote Sensing Domain Adaptive Alignment via Student-Teacher LearningabstractImage Alignment between Synthetic Aperture Radar (SAR) and Electro-Optical (EO) imagery is a task that has comprehensive remote sensing capabilities. Traditional deep learning-based SAR-optical image matching models heavily rely on supervised learning with expensive annotated datasets, leading to reduced accuracy and overfitting when encountering insufficient data during training. To tackle this issue, this paper proposes a student-teacher framework for Domain Adaptation (DA) approach, transferring deep learning models from well-annotated source domains like normal outside optical images to non-annotated SAR-EO target domains. Additionally, in contrast to previous methods which usually use CNN or ordinary Transformer structure to extract features from image pairs, we use self and cross attention mechanisms in Transformer to obtain feature descriptors that are conditioned on both multimodal images. The larger global receptive field and better feature extraction capability provided by this Transformer shows its ability to accommodate large disparities of multimodal data and manage fewer textures such as rural areas with forest or desert in satellite images, where previous backbones usually struggle to produce repeatable and correct interest points. Qiuhang Liu, Weilong Yan, Bo Wang 0019, Bharadwaj Veeravalli, Robby T. Tan |
IGARSS | 3 |
| 2024 | EZ-HOI: VLM Adaptation via Guided Prompt Learning for Zero-Shot HOI DetectionabstractDetecting Human-Object Interactions (HOI) in zero-shot settings, where models must handle unseen classes, poses significant challenges. Existing methods that rely on aligning visual encoders with large Vision-Language Models (VLMs) to tap into the extensive knowledge of VLMs, require large, computationally expensive models and encounter training difficulties. Adapting VLMs with prompt learning offers an alternative to direct alignment. However, fine-tuning on task-specific datasets often leads to overfitting to seen classes and suboptimal performance on unseen classes, due to the absence of unseen class labels. To address these challenges, we introduce a novel prompt learning-based framework for Efficient Zero-Shot HOI detection (EZ-HOI). First, we introduce Large Language Model (LLM) and VLM guidance for learnable prompts, integrating detailed HOI descriptions and visual semantics to adapt VLMs to HOI tasks. However, because training datasets contain seen-class labels alone, fine-tuning VLMs on such datasets tends to optimize learnable prompts for seen classes instead of unseen ones. Therefore, we design prompt learning for unseen classes using information from related seen classes, with LLMs utilized to highlight the differences between unseen and related seen classes. Quantitative evaluations on benchmark datasets demonstrate that our EZ-HOI achieves state-of-the-art performance across various zero-shot settings with only 10.35\% to 33.95\% of the trainable parameters compared to existing methods. Code is available at https://github.com/ChelsieLei/EZ-HOI. Qinqian Lei, Bo Wang 0019, Robby T. Tan |
NeurIPS | 2 |
| 2023 | DSFNet: Dual Space Fusion Network for Occlusion-Robust 3D Dense Face AlignmentabstractSensitivity to severe occlusion and large view angles limits the usage scenarios of the existing monocular 3D dense face alignment methods. The state-of-the-art 3DMM-based method, directly regresses the model's coefficients, underutilizing the low-level 2D spatial and semantic information, which can actually offer cues for face shape and orientation. In this work, we demonstrate how modeling 3D facial geometry in image and model space jointly can solve the occlusion and view angle problems. Instead of predicting the whole face directly, we regress image space features in the visible facial region by dense prediction first. Subsequently, we predict our model's coefficients based on the regressed feature of the visible regions, leveraging the prior knowledge of whole face geometry from the morphable models to complete the invisible regions. We further propose a fusion network that combines the advantages of both the image and model space predictions to achieve high robustness and accuracy in unconstrained scenarios. Thanks to the proposed fusion module, our method is robust not only to occlusion and large pitch and roll view angles, which is the bene- fit of our image space approach, but also to noise and large yaw angles, which is the benefit of our model space method. Comprehensive evaluations demonstrate the superior performance of our method compared with the state-of-the-art methods. On the 3D dense face alignment task, we achieve 3.80% NME on the AFLW2000-3D dataset, which outperforms the state-of-the-art method by 5.5%. Code is available at https://github.com/1hyfst/DSFNet. Heyuan Li, Bo Wang 0019, Yu Cheng 0009, Mohan Kankanhalli, Robby T. Tan |
CVPR | 2 |
| 2023 | High-fidelity point cloud completion with low-resolution recovery and noise-aware upsamplingabstractCompleting an unordered partial point cloud is a challenging task. Existing approaches that rely on decoding a latent feature to recover the complete shape, often lead to the completed point cloud being over-smoothing, losing details, and noisy. Instead of decoding a whole shape, we propose to decode and refine a low-resolution (low-res) point cloud first, and then perform a patch-wise noise-aware upsampling rather than interpolating the whole sparse point cloud at once, which tends to lose details. Regarding the possibility of lacking details of the initially decoded low-res point cloud, we propose an iterative refinement to recover the geometric details and a symmetrization process to preserve the trustworthy information from the input partial point cloud. After obtaining a sparse and complete point cloud, we propose a patch-wise upsampling strategy. Patch-based upsampling allows to recover fine details better rather than decoding a whole shape. The patch extraction approach is to generate training patch pairs between the sparse and ground-truth point clouds with an outlier removal step to suppress the noisy points from the sparse point cloud. Together with the low-res recovery, our whole pipeline can achieve high-fidelity point cloud completion. Comprehensive evaluations are provided to demonstrate the effectiveness of the proposed method and its components. Ren-Wu Li, Bo Wang 0019, Lin Gao 0004, Ling-Xiao Zhang, Chun-Peng Li |
Graph. Model. | 2 |
| 2023 | Dual Networks Based 3D Multi-Person Pose Estimation From Monocular VideoabstractMonocular 3D human pose estimation has made progress in recent years. Most of the methods focus on single persons, which estimate the poses in the person-centric coordinates, i.e., the coordinates based on the center of the target person. Hence, these methods are inapplicable for multi-person 3D pose estimation, where the absolute coordinates (e.g., the camera coordinates) are required. Moreover, multi-person pose estimation is more challenging than single pose estimation, due to inter-person occlusion and close human interactions. Existing top-down multi-person methods rely on human detection (i.e., top-down approach), and thus suffer from the detection errors and cannot produce reliable pose estimation in multi-person scenes. Meanwhile, existing bottom-up methods that do not use human detection are not affected by detection errors, but since they process all persons in a scene at once, they are prone to errors, particularly for persons in small scales. To address all these challenges, we propose the integration of top-down and bottom-up approaches to exploit their strengths. Our top-down network estimates human joints from all persons instead of one in an image patch, making it robust to possible erroneous bounding boxes. Our bottom-up network incorporates human-detection based normalized heatmaps, allowing the network to be more robust in handling scale variations. Finally, the estimated 3D poses from the top-down and bottom-up networks are fed into our integration network for final 3D poses. To address the common gaps between training and testing data, we do optimization during the test time, by refining the estimated 3D human poses using high-order temporal constraint, re-projection loss, and bone length regularizations. We also introduce a two-person pose discriminator that enforces natural two-person interactions. Finally, we apply a semi-supervised method to overcome the 3D ground-truth data scarcity. Our evaluations demonstrate the effectiveness of the proposed method and its individual components. Our code and pretrained models are available publicly: https://github.com/3dpose/3D-Multi-Person-Pose. Yu Cheng 0009, Bo Wang 0019, Robby T. Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Bottom-up 2D pose estimation via dual anatomical centers for small-scale persons
Yu Cheng 0009, Yihao Ai, Bo Wang 0019, Xinchao Wang, Robby T. Tan |
Pattern Recognit. | 3 |
| 2021 | Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular VideosabstractDespite the recent progress, 3D multi-person pose estimation from monocular videos is still challenging due to the commonly encountered problem of missing information caused by occlusion, partially out-of-frame target persons, and inaccurate person detection. To tackle this problem, we propose a novel framework integrating graph convolutional networks (GCNs) and temporal convolutional networks (TCNs) to robustly estimate camera-centric multi-person 3D poses that does not require camera parameters. In particular, we introduce a human-joint GCN, which unlike the existing GCN, is based on a directed graph that employs the 2D pose estimator's confidence scores to improve the pose estimation results. We also introduce a human-bone GCN, which models the bone connections and provides more information beyond human joints. The two GCNs work together to estimate the spatial frame-wise 3D poses and can make use of both visible joint and bone information in the target frame to estimate the occluded or missing human-part information. To further refine the 3D pose estimation, we use our temporal convolutional networks (TCNs) to enforce the temporal and human-dynamics constraints. We use a joint-TCN to estimate person-centric 3D poses across frames, and propose a velocity-TCN to estimate the speed of 3D joints to ensure the consistency of the 3D pose estimation in consecutive frames. Finally, to estimate the 3D human poses for multiple persons, we propose a root-TCN that estimates camera-centric 3D poses without requiring camera parameters. Quantitative and qualitative evaluations demonstrate the effectiveness of the proposed method. Yu Cheng 0009, Bo Wang 0019, Bo Yang 0070, Robby T. Tan |
AAAI | 2 |
| 2021 | Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up NetworksabstractIn monocular video 3D multi-person pose estimation, inter-person occlusion and close interactions can cause human detection to be erroneous and human-joints grouping to be unreliable. Existing top-down methods rely on human detection and thus suffer from these problems. Existing bottom-up methods do not use human detection, but they process all persons at once at the same scale, causing them to be sensitive to multiple-persons scale variations. To address these challenges, we propose the integration of top-down and bottom-up approaches to exploit their strengths. Our top-down network estimates human joints from all persons instead of one in an image patch, making it robust to possible erroneous bounding boxes. Our bottom-up network incorporates human-detection based normalized heatmaps, allowing the network to be more robust in handling scale variations. Finally, the estimated 3D poses from the top-down and bottom-up networks are fed into our integration network for final 3D poses. Besides the integration of top-down and bottom-up networks, unlike existing pose discriminators that are designed solely for a single person, and consequently cannot assess natural inter-person interactions, we propose a two-person pose discriminator that enforces natural two-person interactions. Lastly, we also apply a semi-supervised method to overcome the 3D ground-truth data scarcity. Quantitative and qualitative evaluations show the effectiveness of the proposed method. Our code is available publicly.1 Yu Cheng 0009, Bo Wang 0019, Bo Yang 0070, Robby T. Tan |
CVPR | 2 |
| 2021 | OctField: Hierarchical Implicit Functions for 3D ModelingabstractRecent advances in localized implicit functions have enabled neural implicit representation to be scalable to large scenes.However, the regular subdivision of 3D space employed by these approaches fails to take into account the sparsity of the surface occupancy and the varying granularities of geometric details. As a result, its memory footprint grows cubically with the input volume, leading to a prohibitive computational cost even at a moderately dense decomposition. In this work, we present a learnable hierarchical implicit representation for 3D surfaces, coded OctField, that allows high-precision encoding of intricate surfaces with low memory and computational budget. The key to our approach is an adaptive decomposition of 3D scenes that only distributes local implicit functions around the surface of interest. We achieve this goal by introducing a hierarchical octree structure to adaptively subdivide the 3D space according to the surface occupancy and the richness of part geometry. As octree is discrete and non-differentiable, we further propose a novel hierarchical network that models the subdivision of octree cells as a probabilistic process and recursively encodes and decodes both octree structure and surface geometry in a differentiable manner. We demonstrate the value of OctField for a range of shape modeling and reconstruction tasks, showing superiority over alternative approaches. Jia-Heng Tang, Weikai Chen 0001, Jie Yang 0038, Bo Wang 0019, Songrun Liu, Bo Yang 0070, Lin Gao 0004 |
NeurIPS | 4 |
| 2020 | 3D Human Pose Estimation Using Spatio-Temporal Networks with Explicit Occlusion TrainingabstractEstimating 3D poses from a monocular video is still a challenging task, despite the significant progress that has been made in the recent years. Generally, the performance of existing methods drops when the target person is too small/large, or the motion is too fast/slow relative to the scale and speed of the training data. Moreover, to our knowledge, many of these methods are not designed or trained under severe occlusion explicitly, making their performance on handling occlusion compromised. Addressing these problems, we introduce a spatio-temporal network for robust 3D human pose estimation. As humans in videos may appear in different scales and have various motion speeds, we apply multi-scale spatial features for 2D joints or keypoints prediction in each individual frame, and multi-stride temporal convolutional networks (TCNs) to estimate 3D joints or keypoints. Furthermore, we design a spatio-temporal discriminator based on body structures as well as limb motions to assess whether the predicted pose forms a valid pose and a valid movement. During training, we explicitly mask out some keypoints to simulate various occlusion cases, from minor to severe occlusion, so that our network can learn better and becomes robust to various degrees of occlusion. As there are limited 3D ground truth data, we further utilize 2D video data to inject a semi-supervised learning capability to our network. Experiments on public data sets validate the effectiveness of our method, and our ablation studies show the strengths of our network's individual submodules. Yu Cheng 0009, Bo Yang 0070, Bo Wang 0019, Robby T. Tan |
AAAI | 3 |
| 2019 | Occlusion-Aware Networks for 3D Human Pose Estimation in VideoabstractOcclusion is a key problem in 3D human pose estimation from a monocular video. To address this problem, we introduce an occlusion-aware deep-learning framework. By employing estimated 2D confidence heatmaps of keypoints and an optical-flow consistency constraint, we filter out the unreliable estimations of occluded keypoints. When occlusion occurs, we have incomplete 2D keypoints and feed them to our 2D and 3D temporal convolutional networks (2D and 3D TCNs) that enforce temporal smoothness to produce a complete 3D pose. By using incomplete 2D keypoints, instead of complete but incorrect ones, our networks are less affected by the error-prone estimations of occluded keypoints. Training the occlusion-aware 3D TCN requires pairs of a 3D pose and a 2D pose with occlusion labels. As no such a dataset is available, we introduce a ``Cylinder Man Model'' to approximate the occupation of body parts in 3D space. By projecting the model onto a 2D plane in different viewing angles, we obtain and label the occluded keypoints, providing us plenty of training data. In addition, we use this model to create a pose regularization constraint, preferring the 2D estimations of unreliable keypoints to be occluded. Our method outperforms state-of-the-art methods on Human 3.6M and HumanEva-I datasets. Yu Cheng 0009, Bo Yang 0070, Bo Wang 0019, Wending Yan, Robby T. Tan |
ICCV | 3 |
| 2017 | Data-Driven Rank Aggregation with Application to Grand Challenges
James Fishbaugh, Marcel Prastawa, Bo Wang 0019, Patrick Reynolds, Stephen R. Aylward, Guido Gerig |
MICCAI (2) | 3 |
| 2016 | Modeling 4D pathological changes by leveraging normative models
Bo Wang 0019, Marcel Prastawa, Andrei Irimia, Avishek Saha, Wei Liu 0036, S. Y. Matt Goh, Paul M. Vespa, John D. Van Horn, Guido Gerig |
Comput. Vis. Image Underst. | 1 |
| 2009 | Ultrasound Speckle Reduction via Super Resolution and Nonlinear Diffusion
Bo Wang 0019, Tian Cao 0001, Yuguo Dai, Dong C. Liu |
ACCV (3) | 1 |