VLDB 2026 Research / reviewers in the wild / expert
Angtian Wang
dblp:227/4739
· DBLP profile ↗
24ranked-venue papers
6as first author
21since 2021 · last 2025
0009-0006-9189-5277ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adventurer: Optimizing Vision Mamba Architecture Designs for Efficiencyabstractseries models where we treat images as sequences of patch tokens and employ uni-directional language models to learn visual representations. This modeling paradigm allows us to process images in a recurrent formulation with linear complexity relative to the sequence length, which can effectively address the memory and computation explosion issues posed by high-resolution and fine-grained images. In detail, we introduce two simple designs that seamlessly integrate image inputs into the causal inference framework: a global pooling token placed at the beginning of the sequence and a flipping operation between every two layers. Extensive empirical studies highlight that compared with the existing plain architectures such as DeiT [46] and Vim [57], Adventurer offers an optimal efficiency-accuracy trade-off. For example, our Adventurer-Base attains a competitive test accuracy of 84.3% on the standard ImageNet-1k benchmark with 216 images/s training throughput, which is 3.8× and 6.2× faster than Vim and DeiT to achieve the same result. As Adventurer offers great computation and memory efficiency and allows scaling with linear complexity, we hope this architecture can benefit future explorations in modeling long sequences for high-resolution or fine-grained images. Code is available at https://github.com/wangf3014/Adventurer. Feng Wang 0047, Timing Yang, Yaodong Yu, Sucheng Ren, Guoyizhe Wei, Angtian Wang, Wei Shao 0008, Yuyin Zhou, Alan L. Yuille, Cihang Xie |
CVPR | 6 |
| 2025 | Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question AnsweringabstractFor vision-language models (VLMs), understanding the dynamic properties of objects and their interactions in 3D scenes from videos is crucial for effective reasoning about high-level temporal and action semantics. Although humans are adept at understanding these properties by constructing 3D and temporal (4D) representations of the world, current video understanding models struggle to extract these dynamic semantics, arguably because these models use cross-frame reasoning without underlying knowledge of the 3D/4D scenes.
In this work, we introduce **DynSuperCLEVR**, the first video question answering dataset that focuses on language understanding of the dynamic properties of 3D objects. We concentrate on three physical concepts—*velocity*, *acceleration*, and *collisions*—within 4D scenes. We further generate three types of questions, including factual queries, future predictions, and counterfactual reasoning that involve different aspects of reasoning on these 4D dynamic properties.
To further demonstrate the importance of explicit scene representations in answering these 4D dynamics questions, we propose **NS-4DPhysics**, a **N**eural-**S**ymbolic VideoQA model integrating **Physics** prior for **4D** dynamic properties with explicit scene representation of videos.
Instead of answering the questions directly from the video text input, our method first estimates the 4D world states with a 3D generative model powered by a physical prior, and then uses neural symbolic reasoning to answer the questions based on the 4D world states.
Our evaluation on all three types of questions in DynSuperCLEVR shows that previous video question answering models and large multimodal models struggle with questions about 4D dynamics, while our NS-4DPhysics significantly outperforms previous state-of-the-art models. Xingrui Wang, Wufei Ma, Angtian Wang, Adam Kortylewski, Alan L. Yuille |
ICLR | 3 |
| 2025 | WorldWeaver: Generating Long-Horizon Video Worlds via Rich PerceptionabstractGenerative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these issues, we introduce WorldWeaver, a robust framework for long video generation that jointly models RGB frames and perceptual conditions within a unified long-horizon modeling scheme. Our training framework offers three key advantages. First, by jointly predicting perceptual conditions and color information from a unified representation, it significantly enhances temporal consistency and motion dynamics. Second, by leveraging depth cues, which we observe to be more resistant to drift than RGB, we construct a memory bank that preserves clearer contextual information, improving quality in long-horizon video generation. Third, we employ segmented noise scheduling for training prediction groups, which further mitigates drift and reduces computational cost. Extensive experiments on both diffusion and rectified flow-based models demonstrate the effectiveness of WorldWeaver in reducing temporal drift and improving the fidelity of generated videos. Xueqing Deng, Shoufa Chen, Angtian Wang, Qiushan Guo, Mingfei Han 0003, Zeyue Xue, Mengzhao Chen, Ping Luo 0002 |
NeurIPS | 4 |
| 2024 | HISR: Hybrid Implicit Surface Representation for Photorealistic 3D Human ReconstructionabstractNeural reconstruction and rendering strategies have demonstrated state-of-the-art performances due, in part, to their ability to preserve high level shape details. Existing approaches, however, either represent objects as implicit surface functions or neural volumes and still struggle to recover shapes with heterogeneous materials, in particular human skin, hair or clothes. To this aim, we present a new hybrid implicit surface representation to model human shapes. This representation is composed of two surface layers that represent opaque and translucent regions on the clothed human body. We segment different regions automatically using visual cues and learn to reconstruct two signed distance functions (SDFs). We perform surface-based rendering on opaque regions (e.g., body, face, clothes) to preserve high-fidelity surface normals and volume rendering on translucent regions (e.g., hair). Experiments demonstrate that our approach obtains state-of-the-art results on 3D human reconstructions, and also shows competitive performances on other objects. Angtian Wang, Yuanlu Xu, Nikolaos Sarafianos, Robert Maier 0001, Edmond Boyer, Alan L. Yuille, Tony Tung |
AAAI | 1 |
| 2024 | Structure-Aware Sparse-View X-Ray 3D ReconstructionabstractX-ray, known for its ability to reveal internal structures of objects, is expected to provide richer information for 3D reconstruction than visible light. Yet, existing NeRF algorithms overlook this nature of X-ray, leading to their limitations in capturing structural contents of imaged objects. In this paper, we propose a framework, Structure-Aware X-ray Neural Radiodensity Fields (SAX-NeRF), for sparse-view X-ray 3D reconstruction. Firstly, we design a Line Segment-based Transformer (Lineformer) as the backbone of SAX-NeRF. Linefomer captures internal structures of objects in 3D space by modeling the dependencies within each line segment of an X-ray. Secondly, we present a Masked Local-Global (MLG) ray sampling strategy to extract contextual and geometric information in 2D projection. Plus, we collect a larger-scale dataset X3D covering wider X-ray applications. Experiments on X3D show that SAX-NeRF surpasses previous NeRF-based methods by 12.56 and 2.49 dB on novel view synthesis and CT reconstruction. https://github.com/caiyuanhao1998/SAX-NeRF Yuanhao Cai, Jiahao Wang 0001, Alan L. Yuille, Zongwei Zhou, Angtian Wang |
CVPR | 5 |
| 2024 | Radiative Gaussian Splatting for Efficient X-Ray Novel View Synthesis
Yuanhao Cai, Yixun Liang, Jiahao Wang 0001, Angtian Wang, Yulun Zhang 0001, Xiaokang Yang 0001, Zongwei Zhou, Alan L. Yuille |
ECCV (1) | 4 |
| 2024 | iNeMo: Incremental Neural Mesh Models for Robust Class-Incremental Learning
Tom Fischer, Yaoyao Liu 0001, Artur Jesslen, Prakhar Kaushik, Angtian Wang, Alan L. Yuille, Adam Kortylewski, Eddy Ilg |
ECCV (77) | 6 |
| 2024 | NOVUM: Neural Object Volumes for Robust Object Classification
Artur Jesslen, Guofeng Zhang 0025, Angtian Wang, Wufei Ma, Alan L. Yuille, Adam Kortylewski |
ECCV (4) | 3 |
| 2024 | Generating Images with 3D Annotations Using Diffusion ModelsabstractDiffusion models have emerged as a powerful generative method, capable of producing stunning photo-realistic images from natural language descriptions. However, these models lack explicit control over the 3D structure in the generated images. Consequently, this hinders our ability to obtain detailed 3D annotations for the generated images or to craft instances with specific poses and distances. In this paper, we propose 3D Diffusion Style Transfer (3D-DST), which incorporates 3D geometry control into diffusion models. Our method exploits ControlNet, which extends diffusion models by using visual prompts in addition to text prompts. We generate images of the 3D objects taken from 3D shape repositories~(e.g., ShapeNet and Objaverse), render them from a variety of poses and viewing directions, compute the edge maps of the rendered images, and use these edge maps as visual prompts to generate realistic images. With explicit 3D geometry control, we can easily change the 3D structures of the objects in the generated images and obtain ground-truth 3D annotations automatically. This allows us to improve a wide range of vision tasks, e.g., classification and 3D pose estimation, in both in-distribution (ID) and out-of-distribution (OOD) settings. We demonstrate the effectiveness of our method through extensive experiments on ImageNet-100/200, ImageNet-R, PASCAL3D+, ObjectNet3D, and OOD-CV. The results show that our method significantly outperforms existing methods, e.g., 3.8 percentage points on ImageNet-100 using DeiT-B. Our code is available at <https://ccvl.jhu.edu/3D-DST/> Wufei Ma, Qihao Liu, Jiahao Wang 0001, Angtian Wang, Xiaoding Yuan, Yi Zhang 0099, Zihao Xiao 0001, Guofeng Zhang 0020, Beijia Lu, Ruxiao Duan, Yongrui Qi, Adam Kortylewski, Yaoyao Liu 0001, Alan L. Yuille |
ICLR | 4 |
| 2024 | Semantic Flow: Learning Semantic Fields of Dynamic Scenes from Monocular VideosabstractIn this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual points, Semantic Flow learns semantics from continuous flows that contain rich 3D motion information. As there is 2D-to-3D ambiguity problem in the viewing direction when extracting 3D flow features from 2D video frames, we consider the volume densities as opacity priors that describe the contributions of flow features to the semantics on the frames. More specifically, we first learn a flow network to predict flows in the dynamic scene, and propose a flow feature aggregation module to extract flow features from video frames. Then, we propose a flow attention module to extract motion information from flow features, which is followed by a semantic network to output semantic logits of flows. We integrate the logits with
volume densities in the viewing direction to supervise the flow features with semantic labels on video frames. Experimental results show that our model is able to learn from multiple dynamic scenes and supports a series of new tasks such as instance-level scene editing, semantic completions, dynamic scene tracking and semantic adaption on novel scenes. Fengrui Tian, Yueqi Duan, Angtian Wang, Jianfei Guo, Shaoyi Du |
ICLR | 3 |
| 2024 | From Pixel to Cancer: Cellular Automata in Computed Tomography
Yuxiang Lai, Xiaoxi Chen, Angtian Wang, Alan L. Yuille, Zongwei Zhou |
MICCAI (1) | 3 |
| 2024 | Neural Textured Deformable Meshes for Robust Analysis-by-SynthesisabstractHuman vision demonstrates higher robustness than current AI algorithms under out-of-distribution scenarios. It has been conjectured such robustness benefits from performing analysis-by-synthesis. Our paper formulates triple vision tasks in a consistent manner using approximate analysis-by-synthesis by render-and-compare algorithms on neural features. In this work, we introduce Neural Textured Deformable Meshes (NTDM), which involve the object model with deformable geometry that allows optimization on both camera parameters and object geometries. The deformable mesh is parameterized as a neural field, and covered by whole-surface neural texture maps, which are trained to have spatial discriminability. During inference, we extract the feature map of the test image and subsequently optimize the 3D pose and shape parameters of our model using differentiable rendering to best reconstruct the target feature map. We show that our analysis-by-synthesis is much more robust than conventional neural networks when evaluated on real-world images and even in challenging out-of-distribution scenarios, such as occlusion and domain shift. Our algorithms are competitive with standard algorithms when tested on conventional performance measures. Angtian Wang, Wufei Ma, Alan L. Yuille, Adam Kortylewski |
WACV | 1 |
| 2024 | Robust Category-Level 3D Pose Estimation from Diffusion-Enhanced Synthetic DataabstractObtaining accurate 3D object poses is vital for numerous computer vision applications, such as 3D reconstruction and scene understanding. However, annotating real-world objects is time-consuming and challenging. While synthetically generated training data is a viable alternative, the domain shift between real and synthetic data is a significant challenge. In this work, we aim to narrow the performance gap between models trained on synthetic data and fully supervised models trained on a large amount of real data. We achieve this by approaching the problem from two perspectives: 1) We introduce P3D-Diffusion, a new synthetic dataset with accurate 3D annotations generated with a graphics-guided diffusion model. 2) We propose Cross-domain 3D Consistency, CC3D, for unsupervised domain adaptation of neural mesh models. In particular, we exploit the spatial relationships between features on the mesh surface and a contrastive learning scheme to guide the domain adaptation process. Combined, these two approaches enable our models to perform competitively with state-of-the-art models using only 10% of the respective real training images, while outperforming the SOTA model by a wide margin using only 50% of the real training data. By encouraging the diversity of synthetic data and generating the images with an OOD-aware manner, our model further demonstrates robust generalization to out-of-distribution scenarios despite being trained with minimal real data. The code is available at https://github.com/YangYY06/synthetic_3d. Wufei Ma, Angtian Wang, Xiaoding Yuan, Alan L. Yuille, Adam Kortylewski |
WACV | 3 |
| 2023 | 3D-Aware Neural Body Fitting for Occlusion Robust 3D Human Pose EstimationabstractRegression-based methods for 3D human pose estimation directly predict the 3D pose parameters from a 2D image using deep networks. While achieving state-of-the-art performance on standard benchmarks, their performance degrades under occlusion. In contrast, optimization-based methods fit a parametric body model to 2D features in an iterative manner. The localized reconstruction loss can potentially make them robust to occlusion, but they suffer from the 2D-3D ambiguity. Motivated by the recent success of generative models in rigid object pose estimation, we propose 3D-aware Neural Body Fitting (3DNBF) - an approximate analysis-by-synthesis approach to 3D human pose estimation with SOTA performance and occlusion robustness. In particular, we propose a generative model of deep features based on a volumetric human representation with Gaussian ellipsoidal kernels emitting 3D pose-dependent feature vectors. The neural features are trained with contrastive learning to become 3D-aware and hence to overcome the 2D-3D ambiguity. Experiments show that 3DNBF outperforms other approaches on both occluded and standard benchmarks. Code is available at https://github.com/edz-o/3DNBF Yi Zhang 0099, Pengliang Ji, Angtian Wang, Jieru Mei, Adam Kortylewski, Alan L. Yuille |
ICCV | 3 |
| 2023 | VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis
Angtian Wang, Peng Wang 0001, Adam Kortylewski, Alan L. Yuille |
ICLR | 1 |
| 2023 | CoKe: Contrastive Learning for Robust Keypoint DetectionabstractIn this paper, we introduce a contrastive learning framework for keypoint detection (CoKe). Keypoint detection differs from other visual tasks where contrastive learning has been applied because the input is a set of images in which multiple keypoints are annotated. This requires the contrastive learning to be extended such that the keypoints are represented and detected independently, which enables the contrastive loss to make the keypoint features different from each other and from the background. Our approach has two benefits: It enables us to exploit contrastive learning for keypoint detection, and by detecting each key-point independently the detection becomes more robust to occlusion compared to holistic methods, such as stacked hourglass networks, which attempt to detect all keypoints jointly. Our CoKe framework introduces several technical innovations. In particular, we introduce: (i) A clutter bank to represent non-keypoint features; (ii) a keypoint bank that stores prototypical representations of keypoints to approximate the contrastive loss between keypoints; and (iii) a cumulative moving average update to learn the key-point prototypes while training the feature extractor. Our experiments on a range of diverse datasets (PASCAL3D+, MPII, ObjectNet3D) show that our approach works as well, or better than, alternative methods for keypoint detection, even for human keypoints, for which the literature is vast. Moreover, we observe that CoKe is exceptionally robust to partial occlusion and previously unseen object poses. Yutong Bai, Angtian Wang, Adam Kortylewski, Alan L. Yuille |
WACV | 2 |
| 2022 | Robust Category-Level 6D Pose Estimation with Coarse-to-Fine Rendering of Neural Features
Wufei Ma, Angtian Wang, Alan L. Yuille, Adam Kortylewski |
ECCV (9) | 2 |
| 2022 | OOD-CV: A Benchmark for Robustness to Out-of-Distribution Shifts of Individual Nuisances in Natural Images
Bingchen Zhao, Shaozuo Yu, Wufei Ma, Mingxin Yu, Shenxiao Mei, Angtian Wang, Ju He, Alan L. Yuille, Adam Kortylewski |
ECCV (8) | 6 |
| 2021 | NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose Estimation
Angtian Wang, Adam Kortylewski, Alan L. Yuille |
ICLR | 1 |
| 2021 | Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D PoseabstractWe study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D pose annotation from the labelled to unlabelled images reliably, despite unseen 3D views and nuisance variations such as the object shape, texture, illumination or scene context. In our approach, objects are represented as 3D cuboid meshes composed of feature vectors at each mesh vertex. The model is initialized from a few labelled images and is subsequently used to synthesize feature representations of unseen 3D views. The synthesized views are matched with the feature representations of unlabelled images to generate pseudo-labels of the 3D pose. The pseudo-labelled data is, in turn, used to train the feature extractor such that the features at each mesh vertex are more invariant across varying 3D views of the object. Our model is trained in an EM-type manner alternating between increasing the 3D pose invariance of the feature extractor and annotating unlabelled data through neural view synthesis and matching. We demonstrate the effectiveness of the proposed semi-supervised learning framework for 3D pose estimation on the PASCAL3D+ and KITTI datasets. We find that our approach outperforms all baselines by a wide margin, particularly in an extreme few-shot setting where only 7 annotated images are given. Remarkably, we observe that our model also achieves an exceptional robustness in out-of-distribution scenarios that involve partial occlusion. Angtian Wang, Shenxiao Mei, Alan L. Yuille, Adam Kortylewski |
NeurIPS | 1 |
| 2021 | Compositional Convolutional Neural Networks: A Robust and Interpretable Model for Object Recognition Under Occlusion
Adam Kortylewski, Qing Liu 0017, Angtian Wang, Yihong Sun, Alan L. Yuille |
Int. J. Comput. Vis. | 3 |
| 2020 | Robust Object Detection Under Occlusion With Context-Aware CompositionalNetsabstractDetecting partially occluded objects is a difficult task.Our experimental results show that deep learning approaches, such as Faster R-CNN, are not robust at object detection under occlusion.Compositional convolutional neural networks (CompositionalNets) have been shown to be robust at classifying occluded objects by explicitly representing the object as a composition of parts.In this work, we propose to overcome two limitations of Compositional-Nets which will enable them to detect partially occluded objects: 1) CompositionalNets, as well as other DCNN architectures, do not explicitly separate the representation of the context from the object itself.Under strong object occlusion, the influence of the context is amplified which can have severe negative effects for detection at test time.In order to overcome this, we propose to segment the context during training via bounding box annotations.We then use the segmentation to learn a context-aware CompositionalNet that disentangles the representation of the context and the object.2) We extend the part-based voting scheme in Compo-sitionalNets to vote for the corners of the object's bounding box, which enables the model to reliably estimate bounding boxes for partially occluded objects.Our extensive experiments show that our proposed model can detect objects robustly, increasing the detection performance of strongly occluded vehicles from PASCAL3D+ and MS-COCO by 41% and 35% respectively in absolute performance relative to Faster R-CNN. Angtian Wang, Yihong Sun, Adam Kortylewski, Alan L. Yuille |
CVPR | 1 |
| 2019 | Hyper-Pairing Network for Multi-phase Pancreatic Ductal Adenocarcinoma Segmentation
Yuyin Zhou, Yingwei Li 0002, Zhishuai Zhang, Yan Wang 0033, Angtian Wang, Elliot K. Fishman, Alan L. Yuille, Seyoun Park |
MICCAI (2) | 5 |
| 2018 | Weakly Supervised Region Proposal Network and Object Detection
Peng Tang 0005, Xinggang Wang, Angtian Wang, Yongluan Yan, Wenyu Liu 0001, Junzhou Huang, Alan L. Yuille |
ECCV (11) | 3 |