Tianyu Wang 0035

dblp:35/8397-35 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0001-9032-8488ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 58% Vision and language · 15% Segmentation and scene understanding · 10%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d scene understanding
1.422025
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer · CVPR 2025
Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition · ECCV (23) 2022
Computer vision › Vision and language › vision-language model › multimodal large language model
3d multimodal large language model
0.912025
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer · CVPR 2025
Computer vision › 3D vision
3d reconstruction
0.912025
DCHM: Depth-Consistent Human Modeling for Multiview Detection · ICCV 2025
Computer vision › Vision and language
multimodal dialogue
0.912025
3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer · CVPR 2025
Computer vision › 3D vision › 3d object detection
multi-view pedestrian detection
0.912025
DCHM: Depth-Consistent Human Modeling for Multiview Detection · ICCV 2025
Computer vision › Image recognition and object detection
pedestrian detection
0.912025
DCHM: Depth-Consistent Human Modeling for Multiview Detection · ICCV 2025
Computer vision › 3D vision
depth estimation
0.812024
Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth Estimation · CVPR 2024
Computer vision › 3D vision
motion estimation
0.812024
Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth Estimation · CVPR 2024
Computer vision › 3D vision › motion estimation
rigid motion estimation
0.812024
Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth Estimation · CVPR 2024
Computer vision › 3D vision › depth estimation › self-supervised depth estimation
self-supervised monocular depth estimation
0.812024
Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth Estimation · CVPR 2024
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.612022
Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition · ECCV (23) 2022
Computer vision › 3D vision › 3d scene understanding › scene decomposition
unsupervised 3d scene decomposition
0.612022
Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition · ECCV (23) 2022
Machine learning › Learning paradigms
unsupervised learning
0.612022
Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition · ECCV (23) 2022
Computer vision › Segmentation and scene understanding › image segmentation
multi-view segmentation
0.312025
DCHM: Depth-Consistent Human Modeling for Multiview Detection · ICCV 2025

Methods — techniques the papers use, named apart from their topics

unified instruction tuning · 0.9superpixel segmentation · 0.9omni superpoint transformer · 0.9hybrid pre-training · 0.9gaussian splatting · 0.9self-supervised learning · 0.8scale alignment · 0.8pseudo-labeling · 0.8spatially invariant object-centric learning · 0.6
YearPublicationVenuePosition
2025 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer
abstract
Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we introduce 3D-LLaVA, a simple yet highly powerful 3D LMM designed to act as an intelligent assistant in comprehending, reasoning, and interacting with the 3D world. Unlike existing top-performing methods that rely on complicated pipelines—such as offline multi-view feature extraction or additional task-specific heads—3D-LLaVA adopts a minimalist design with integrated architecture and only takes point clouds as input. At the core of 3D-LLaVA is a new Omni Superpoint Transformer (OST), which integrates three functionalities: (1) a visual feature selector that converts and selects visual tokens, (2) a visual prompt encoder that embeds interactive visual prompts into the visual token space, and (3) a referring mask decoder that produces 3D masks based on text description. This versatile OST is empowered by the hybrid pretraining to obtain perception priors and leveraged as the visual connector that bridges the 3D data to the LLM. After performing unified instruction tuning, our 3D-LLaVA reports impressive results on various benchmarks. The code and model will be released at https://github.com/djiajunustc/3D-LLaVA.
Jiajun Deng, Tianyu He, Tianyu Wang 0035, Feras Dayoub, Ian D. Reid 0001
CVPR4
2025 DCHM: Depth-Consistent Human Modeling for Multiview Detection
abstract
Multiview pedestrian detection typically involves two stages: human modeling and pedestrian localization. Human modeling represents pedestrians in 3D space by fusing multiview information, making its quality crucial for detection accuracy. However, existing methods often introduce noise and have low precision. While some approaches reduce noise by fitting on costly multiview 3D annotations, they often struggle to generalize across diverse scenes. To eliminate reliance on human-labeled annotations and accurately model humans, we propose Depth-Consistent Human Modeling (DCHM), a framework designed for consistent depth estimation and multiview fusion in global coordinates. Specifically, our proposed pipeline with superpixel-wise Gaussian Splatting achieves multiview depth consistency in sparse-view, large-scaled, and crowded scenarios, producing precise point clouds for pedestrian localization. Extensive validations demonstrate that our method significantly reduces noise during human modeling, outperforming previous state-of-the-art baselines. Additionally, to our knowledge, DCHM is the first to reconstruct pedestrians and perform multiview segmentation in such a challenging setting. Code is available on the \href{https://jiahao-ma.github.io/DCHM/}{project page}.
Tianyu Wang 0035, Miaomiao Liu 0001, David Ahmedt-Aristizabal
ICCV2
2024 Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth Estimation
abstract
This paper focuses on self-supervised monocular depth estimation in dynamic scenes trained on monocular videos. Existing methods jointly estimate pixel-wise depth and motion, relying mainly on an image reconstruction loss. Dynamic regions11Dynamic regions indicate regions covered by moving objects. remain a critical challenge for these methods due to the inherent ambiguity in depth and motion estimation, resulting in inaccurate depth estimation. This paper proposes a self-supervised training framework exploiting pseudo depth labels for dynamic regions from training data. The key contribution of our framework is to decouple depth estimation for static and dynamic regions of images in the training data. We start with an unsupervised depth estimation approach, which provides reliable depth estimates for static regions and motion cues for dynamic regions and allows us to extract moving object information at the instance level. In the next stage, we use an object network to estimate the depth of those moving objects assuming rigid motions. Then, we propose a new scale alignment module to address the scale ambiguity between estimated depths for static and dynamic regions. We can then use the depth labels generated to train an end-to-end depth estimation network and improve its performance. Extensive experiments on the Cityscapes and KITTI datasets show that our self-training strategy consistently outperforms existing self-/unsupervised depth estimation methods. Our code is available at https://github.com/HoangChuongNguyen/mono-consistent-depth.git
Hoang Chuong Nguyen, Tianyu Wang 0035, José M. Álvarez 0004, Miaomiao Liu 0001
CVPR2
2022 Spatially Invariant Unsupervised 3D Object-Centric Learning and Scene Decomposition
Tianyu Wang 0035, Miaomiao Liu 0001, Kee Siong Ng
ECCV (23)1