Junyi Zhang 0004

dblp:00/1627-4 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-9291-3098ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 58% Video understanding and tracking · 14% Generative modeling · 11%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d reconstruction
1.722025
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion · ICLR 2025
ST4RTrack: Simultaneous 4D Reconstruction and Tracking in the World · ICCV 2025
Computer vision › 3D vision › correspondence estimation
semantic correspondence
1.422024
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence · CVPR 2024
A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence · NeurIPS 2023
Computer vision › 3D vision › motion estimation › 3d motion tracking
3d point tracking
0.912025
ST4RTrack: Simultaneous 4D Reconstruction and Tracking in the World · ICCV 2025
Computer vision › 3D vision
camera pose estimation
0.912025
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion · ICLR 2025
Computer vision › 3D vision
depth estimation
0.912025
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion · ICLR 2025
Computer vision › 3D vision › 3d reconstruction
dynamic 3d reconstruction
0.912025
ST4RTrack: Simultaneous 4D Reconstruction and Tracking in the World · ICCV 2025
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.912025
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion · ICLR 2025
Computer vision › Video understanding and tracking
feature tracking
0.912025
ST4RTrack: Simultaneous 4D Reconstruction and Tracking in the World · ICCV 2025
Robotics › Motion planning and robot control
robot learning
0.912025
Pre-training Auto-regressive Robotic Models with 4D Representations · ICML 2025
Computer vision › 3D vision › depth estimation
video depth estimation
0.912025
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion · ICLR 2025
Computer vision › Video understanding and tracking › human action analysis
action understanding
0.812024
From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding · CVPR 2024
Machine learning › Generative modeling
diffusion model
0.712023
LayoutDiffusion: Improving Graphic Layout Generation by Discrete Diffusion Probabilistic Models · ICCV 2023
Machine learning › Generative modeling › diffusion model › diffusion-based representation learning
diffusion model features
0.712023
A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence · NeurIPS 2023
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.712023
LayoutDiffusion: Improving Graphic Layout Generation by Discrete Diffusion Probabilistic Models · ICCV 2023
Machine learning › Representation and self-supervised learning
visual representation
0.712023
A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence · NeurIPS 2023
Visual content generation and editing › layout generation
graphic layout generation
0.712023
LayoutDiffusion: Improving Graphic Layout Generation by Discrete Diffusion Probabilistic Models · ICCV 2023
Computer vision › Image recognition and object detection
human-object interaction detection
0.612022
Mining Cross-Person Cues for Body-Part Interactiveness Learning in HOI Detection · ECCV (4) 2022

Methods — techniques the papers use, named apart from their topics

transfer learning · 1.6reprojection loss · 0.9pointmap representation · 0.9pointmap prediction · 0.9monocular depth estimation · 0.9fine-tuning · 0.9feed-forward reconstruction · 0.9feed-forward framework · 0.9autoregressive modeling · 0.9verb taxonomy hierarchy · 0.8discrete denoising diffusion · 0.7block-wise transition matrix · 0.7
YearPublicationVenuePosition
2025 ST4RTrack: Simultaneous 4D Reconstruction and Tracking in the World
abstract
Dynamic 3D reconstruction and point tracking in videos are typically treated as separate tasks, despite their deep connection. We propose St4RTrack, a feed-forward framework that simultaneously reconstructs and tracks dynamic video content in a world coordinate frame from RGB inputs. This is achieved by predicting two appropriately defined pointmaps for a pair of frames captured at different moments. Specifically, we predict both pointmaps at the same moment, in the same world, capturing both static and dynamic scene geometry while maintaining 3D correspondences. Chaining these predictions through the video sequence with respect to a reference frame naturally computes long-range correspondences, effectively combining 3D reconstruction with 3D tracking. Unlike prior methods that rely heavily on 4D ground truth supervision, we employ a novel adaptation scheme based on a reprojection loss. We establish a new extensive benchmark for world-frame reconstruction and tracking, demonstrating the effectiveness and efficiency of our unified, data-driven framework. Our code, model, and benchmark will be released.
Haiwen Feng, Junyi Zhang 0004, Qianqian Wang 0002, Yufei Ye 0001, Pengcheng Yu, Michael J. Black, Trevor Darrell, Angjoo Kanazawa
ICCV2
2025 MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
abstract
Estimating geometry from dynamic scenes, where objects move and deform over time, remains a core challenge in computer vision. Current approaches often rely on multi-stage pipelines or global optimizations that decompose the problem into subtasks, like depth and flow, leading to complex systems prone to errors. In this paper, we present Motion DUSt3R (MonST3R), a novel geometry-first approach that directly estimates per-timestep geometry from dynamic scenes. Our key insight is that by simply estimating a pointmap for each timestep, we can effectively adapt DUSt3R’s representation, previously only used for static scenes, to dynamic scenes. However, this approach presents a significant challenge: the scarcity of suitable training data, namely dynamic, posed videos with depth labels. Despite this, we show that by posing the problem as a fine-tuning task, identifying several suitable datasets, and strategically training the model on this limited data, we can surprisingly enable the model to handle dynamics, even without an explicit motion representation. Based on this, we introduce new optimizations for several downstream video-specific tasks and demonstrate strong performance on video depth and camera pose estimation, outperforming prior work in terms of robustness and efficiency. Moreover, MonST3R shows promising results for primarily feed-forward 4D reconstruction. Interactive 4D results, source code, and trained models are available at: https://monst3r-project.github.io/.
Junyi Zhang 0004, Charles Herrmann, Junhwa Hur, Varun Jampani, Trevor Darrell, Forrester Cole, Deqing Sun, Ming-Hsuan Yang 0001
ICLR1
2025 Pre-training Auto-regressive Robotic Models with 4D Representations
abstract
Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in robotics have struggled to achieve similar success, limited by either the need for costly robotic annotations or the lack of representations that effectively model the physical world. In this paper, we introduce ARM4R, an **A**uto-regressive **R**obotic **M**odel that leverages low-level **4**D **R**epresentations learned from human video data to yield a better pre-trained robotic model. Specifically, we focus on utilizing 3D point tracking representations from videos derived by lifting 2D representations into 3D space via monocular depth estimation across time. These 4D representations maintain a shared geometric structure between the points and robot state representations up to a linear transformation, enabling efficient transfer learning from human video data to low-level robotic control. Our experiments show that ARM4R can transfer efficiently from human video data to robotics and consistently improves performance on tasks across various robot environments and configurations.
Dantong Niu, Yuvan Sharma, Haoru Xue, Giscard Biamby, Junyi Zhang 0004, Ziteng Ji, Trevor Darrell, Roei Herzig
ICML5
2024 From Isolated Islands to Pangea: Unifying Semantic Space for Human Action Understanding
abstract
Action understanding has attracted long-term attention. It can be formed as the mapping from the physical space to the semantic space. Typically, researchers built datasets according to idiosyncratic choices to define classes and push the envelope of benchmarks respectively. Datasets are incompatible with each other like “Isolated Islands” due to semantic gaps and various class granularities, e.g., do housework in dataset A and wash plate in dataset B. We argue that we need a more principled semantic space to concentrate the community efforts and use all datasets together to pursue generalizable action learning. To this end, we design a structured action semantic space in view of verb taxonomy hierarchy and covering massive actions. By aligning the classes of previous datasets to our semantic space, we gather (image/video/skeleton/McCap] datasets into a unified database in a unified label system, i.e., bridging “isolated islands” into a “Pangea”. Accordingly, we propose a novel model mapping from the physical space to semantic space to fully use Pangea. In extensive experiments, our new system shows significant superiority, especially in transfer learning. Our code and data will be made public at https://mvig-rhos.com/pangea.
Yong-Lu Li 0001, Xinpeng Liu 0002, Yiming Dou, Yikun Ji, Junyi Zhang 0004, Yixing Li, Jingru Tan, Cewu Lu
CVPR7
2024 Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
abstract
While pre-trained large-scale vision models have shown significant promise for semantic correspondence, their features often struggle to grasp the geometry and orientation of instances. This paper identifies the importance of being geometry-aware for semantic correspondence and reveals a limitation of the features of current foundation models under simple post-processing. We show that incorporating this information can markedly enhance semantic correspondence performance with simple but effective solutions in both zero-shot and supervised settings. We also construct a new challenging benchmark for semantic correspondence built from an existing animal pose estimation dataset, for both pre-training validating models. Our method achieves a [email protected] score of 65.4 (zero-shot) and 85.6 (supervised) on the challenging SPair-71k dataset, surpassing the state of the art by 5.5p and 11.0p absolute gains, respectively. Our code and datasets are publicly available at: https://telling-left-from-right.github.io.
Junyi Zhang 0004, Charles Herrmann, Junhwa Hur, Varun Jampani, Deqing Sun, Ming-Hsuan Yang 0001
CVPR1
2023 LayoutDiffusion: Improving Graphic Layout Generation by Discrete Diffusion Probabilistic Models
abstract
Creating graphic layouts is a fundamental step in graphic designs. In this work, we present a novel generative model named LayoutDiffusion for automatic layout generation. As layout is typically represented as a sequence of discrete tokens, LayoutDiffusion models layout generation as a discrete denoising diffusion process. It learns to reverse a mild forward process, in which layouts become increasingly chaotic with the growth of forward steps and layouts in the neighboring steps do not differ too much. Designing such a mild forward process is however very challenging as layout has both categorical attributes and ordinal attributes. To tackle the challenge, we summarize three critical factors for achieving a mild forward process for the layout, i.e., legality, coordinate proximity and type disruption. Based on the factors, we propose a block-wise transition matrix coupled with a piece-wise linear noise schedule. Experiments on RICO and PubLayNet datasets show that LayoutDiffusion outperforms state-of-the-art approaches significantly. Moreover, it enables two conditional layout generation tasks in a plug-and-play manner without re-training and achieves better performance than existing methods. Project page: https://layoutdiffusion.github.io.
Junyi Zhang 0004, Shizhao Sun, Jian-Guang Lou, Dongmei Zhang 0001
ICCV1
2023 A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic Correspondence
abstract
Text-to-image diffusion models have made significant advances in generating and editing high-quality images. As a result, numerous approaches have explored the ability of diffusion model features to understand and process single images for downstream tasks, e.g., classification, semantic segmentation, and stylization. However, significantly less is known about what these features reveal across multiple, different images and objects. In this work, we exploit Stable Diffusion (SD) features for semantic and dense correspondence and discover that with simple post-processing, SD features can perform quantitatively similar to SOTA representations. Interestingly, the qualitative analysis reveals that SD features have very different properties compared to existing representation learning features, such as the recently released DINOv2: while DINOv2 provides sparse but accurate matches, SD features provide high-quality spatial information but sometimes inaccurate semantic matches. We demonstrate that a simple fusion of these two features works surprisingly well, and a zero-shot evaluation using nearest neighbors on these fused features provides a significant performance gain over state-of-the-art methods on benchmark datasets, e.g., SPair-71k, PF-Pascal, and TSS. We also show that these correspondences can enable interesting applications such as instance swapping in two images. Project page: https://sd-complements-dino.github.io/.
Junyi Zhang 0004, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera, Varun Jampani, Deqing Sun, Ming-Hsuan Yang 0001
NeurIPS1
2022 Mining Cross-Person Cues for Body-Part Interactiveness Learning in HOI Detection
Yong-Lu Li 0001, Xinpeng Liu 0002, Junyi Zhang 0004, Cewu Lu
ECCV (4)4