Jiahao Wang 0001

dblp:34/5354-1 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-8768-4913ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Mamba-Reg: Vision Mamba Also Needs Registers
abstract
Similar to Vision Transformers, this paper identifies artifacts also present within the feature maps of Vision Mamba. These artifacts, corresponding to high-norm tokens emerging in low-information background areas of images, appear much more severe in Vision Mamba—they exist prevalently even with the tiny-sized model and activate extensively across background regions. To mitigate this issue, we follow the prior solution of introducing register tokens into Vision Mamba. To better cope with Mamba blocks’ uni-directional inference paradigm, two key modifications are introduced: 1) evenly inserting registers throughout the input token sequence, and 2) recycling registers for final decision predictions. We term this new architecture Mamba®. Qualitative observations suggest, compared to vanilla Vision Mamba, Mamba®’s feature maps appear cleaner and more focused on semantically meaningful regions. Quantitatively, Mamba®attains stronger performance and scales better. For example, on the ImageNet benchmark, our Mamba®-B attains 83.0% accuracy, significantly outperforming Vim-B’s 81.8%; furthermore, we provide the first successful scaling to the large model size with 341M parameters, attaining competitive accuracies of 83.6% and 84.5% for 224×224 and 384×384 inputs, respectively. Additional validation on the downstream semantic segmentation task also supports Mamba®’s efficacy. Code is available at https://github.com/wangf3014/Mamba-Reg.
Feng Wang 0047, Jiahao Wang 0001, Sucheng Ren, Guoyizhe Wei, Jieru Mei, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie
CVPR2
2025 EasyRet3D: Uncalibrated Multi-View Multi-Human 3D Reconstruction and Tracking
abstract
Current methods performing 3D human pose estimation from multi-view still bear several key limitations. First, most methods require manual intrinsic and extrinsic camera calibration, which is laborious and difficult in many settings. Second, more accurate models rely on further training on the same datasets they evaluate, severely limiting their generalizability in real-world settings. We address these limitations with EasyRet3D (Easy REconstruction and Tracking in 3D), which simultaneously reconstructs and tracks 3D humans in a global coordinate frame across all views with uncalibrated cameras and videos in the wild. EasyRet3D is a compositional framework that composes our proposed modules (Automatic Calibration module, Adaptive Stitching Module, and Optimization Module) and off-the-shelf, large pre-trained models at intermediate steps to avoid manual intrinsic and extrinsic calibration and task-specific training. EasyRet3D outperforms all existing multi-view 3D tracking or pose estimation methods in Panoptic, EgoHumans, Shelf, and Human3.6M datasets. Code and demos will be released on the project website.
Junjie Oscar Yin, Jiahao Wang 0001, Yi Zhang 0099, Alan L. Yuille
WACV3
2024 Structure-Aware Sparse-View X-Ray 3D Reconstruction
abstract
X-ray, known for its ability to reveal internal structures of objects, is expected to provide richer information for 3D reconstruction than visible light. Yet, existing NeRF algorithms overlook this nature of X-ray, leading to their limitations in capturing structural contents of imaged objects. In this paper, we propose a framework, Structure-Aware X-ray Neural Radiodensity Fields (SAX-NeRF), for sparse-view X-ray 3D reconstruction. Firstly, we design a Line Segment-based Transformer (Lineformer) as the backbone of SAX-NeRF. Linefomer captures internal structures of objects in 3D space by modeling the dependencies within each line segment of an X-ray. Secondly, we present a Masked Local-Global (MLG) ray sampling strategy to extract contextual and geometric information in 2D projection. Plus, we collect a larger-scale dataset X3D covering wider X-ray applications. Experiments on X3D show that SAX-NeRF surpasses previous NeRF-based methods by 12.56 and 2.49 dB on novel view synthesis and CT reconstruction. https://github.com/caiyuanhao1998/SAX-NeRF
Yuanhao Cai, Jiahao Wang 0001, Alan L. Yuille, Zongwei Zhou, Angtian Wang
CVPR2
2024 Radiative Gaussian Splatting for Efficient X-Ray Novel View Synthesis
Yuanhao Cai, Yixun Liang, Jiahao Wang 0001, Angtian Wang, Yulun Zhang 0001, Xiaokang Yang 0001, Zongwei Zhou, Alan L. Yuille
ECCV (1)3
2024 Generating Images with 3D Annotations Using Diffusion Models
abstract
Diffusion models have emerged as a powerful generative method, capable of producing stunning photo-realistic images from natural language descriptions. However, these models lack explicit control over the 3D structure in the generated images. Consequently, this hinders our ability to obtain detailed 3D annotations for the generated images or to craft instances with specific poses and distances. In this paper, we propose 3D Diffusion Style Transfer (3D-DST), which incorporates 3D geometry control into diffusion models. Our method exploits ControlNet, which extends diffusion models by using visual prompts in addition to text prompts. We generate images of the 3D objects taken from 3D shape repositories~(e.g., ShapeNet and Objaverse), render them from a variety of poses and viewing directions, compute the edge maps of the rendered images, and use these edge maps as visual prompts to generate realistic images. With explicit 3D geometry control, we can easily change the 3D structures of the objects in the generated images and obtain ground-truth 3D annotations automatically. This allows us to improve a wide range of vision tasks, e.g., classification and 3D pose estimation, in both in-distribution (ID) and out-of-distribution (OOD) settings. We demonstrate the effectiveness of our method through extensive experiments on ImageNet-100/200, ImageNet-R, PASCAL3D+, ObjectNet3D, and OOD-CV. The results show that our method significantly outperforms existing methods, e.g., 3.8 percentage points on ImageNet-100 using DeiT-B. Our code is available at <https://ccvl.jhu.edu/3D-DST/>
Wufei Ma, Qihao Liu, Jiahao Wang 0001, Angtian Wang, Xiaoding Yuan, Yi Zhang 0099, Zihao Xiao 0001, Guofeng Zhang 0020, Beijia Lu, Ruxiao Duan, Yongrui Qi, Adam Kortylewski, Yaoyao Liu 0001, Alan L. Yuille
ICLR3
2024 OOD-CV-v2 : An Extended Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images
abstract
Enhancing the robustness of vision algorithms in real-world scenarios is challenging. One reason is that existing robustness benchmarks are limited, as they either rely on synthetic data or ignore the effects of individual nuisance factors. We introduce OOD-CV-v2, a benchmark dataset that includes out-of-distribution examples of 10 object categories in terms of pose, shape, texture, context and the weather conditions, and enables benchmarking of models for image classification, object detection, and 3D pose estimation. In addition to this novel dataset, we contribute extensive experiments using popular baseline methods, which reveal that: 1) Some nuisance factors have a much stronger negative effect on the performance compared to others, also depending on the vision task. 2) Current approaches to enhance robustness have only marginal effects, and can even reduce robustness. 3) We do not observe significant differences between convolutional and transformer architectures. We believe our dataset provides a rich test bed to study robustness and will help push forward research in this area.
Bingchen Zhao, Jiahao Wang 0001, Wufei Ma, Artur Jesslen, Siwei Yang, Shaozuo Yu, Oliver Zendel 0001, Christian Theobalt, Alan L. Yuille, Adam Kortylewski
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Animal3D: A Comprehensive Dataset of 3D Animal Pose and Shape
abstract
Accurately estimating the 3D pose and shape is an essential step towards understanding animal behavior, and can potentially benefit many downstream applications, such as wildlife conservation. However, research in this area is held back by the lack of a comprehensive and diverse dataset with high-quality 3D pose and shape annotations. In this paper, we propose Animal3D, the first comprehensive dataset for mammal animal 3D pose and shape estimation. Animal3D consists of 3379 images collected from 40 mammal species, high-quality annotations of 26 key-points, and importantly the pose and shape parameters of the SMAL [50] model. All annotations were labeled and checked manually in a multi-stage process to ensure highest quality results. Based on the Animal3D dataset, we benchmark representative shape and pose estimation models at: (1) supervised learning from only the Animal3D data, (2) synthetic to real transfer from synthetically generated images, and (3) fine-tuning human pose and shape estimation models. Our experimental results demonstrate that predicting the 3D shape and pose of animals across species remains a very challenging task, despite significant advances in human pose estimation. Our results further demonstrate that synthetic pre-training is a viable strategy to boost the model performance. Overall, Animal3D opens new directions for facilitating future research in animal 3D pose and shape estimation, and is publicly available.
Jiacong Xu, Yi Zhang 0099, Wufei Ma, Artur Jesslen, Pengliang Ji, Qixin Hu, Qihao Liu, Jiahao Wang 0001, Wei Ji 0011, Chen Wang 0049, Xiaoding Yuan, Prakhar Kaushik, Guofeng Zhang 0020, Jie Liu 0044, Yushan Xie, Yawen Cui, Alan L. Yuille, Adam Kortylewski
ICCV10
2022 SAGA: Stochastic Whole-Body Grasping with Contact
Jiahao Wang 0001, Yan Zhang 0054, Otmar Hilliges, Fisher Yu 0001, Siyu Tang 0001
ECCV (6)2
2019 Gaussian field estimator with manifold regularization for retinal image registration
Jiahao Wang 0001, Jun Chen 0019, Shuaibin Zhang, Xiaoguang Mei, Jun Huang 0008, Jiayi Ma 0001
Signal Process.1