Dadong Jiang

dblp:359/4764 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Visual content generation and editing · 34% Virtual and augmented reality · 20% Geometric modeling and processing · 14%
Artificial intelligence
2 papers
3D vision · 38% Video understanding and tracking · 38% Generative modeling · 23%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.912025
Timeformer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction · ICCV 2025
Computer vision › Video understanding and tracking › temporal modeling
temporal relation modeling
0.912025
Timeformer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction · ICCV 2025
Visual content generation and editing
3d content generation
0.912025
FlexiTex: Enhancing Texture Generation via Visual Guidance · AAAI 2025
Geometric modeling and processing › 3d reconstruction › 3d scene reconstruction
dynamic scene reconstruction
0.912025
Timeformer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction · ICCV 2025
Visual content generation and editing › texture synthesis
text-to-texture generation
0.912025
FlexiTex: Enhancing Texture Generation via Visual Guidance · AAAI 2025
Visual content generation and editing
texture synthesis
0.912025
FlexiTex: Enhancing Texture Generation via Visual Guidance · AAAI 2025
Virtual and augmented reality
visual guidance
0.912025
FlexiTex: Enhancing Texture Generation via Visual Guidance · AAAI 2025
Image and video coding
neural video representation
0.812024
DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes · CVPR 2024
Computational photography and imaging › snapshot compressive imaging
video reconstruction
0.812024
DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes · CVPR 2024
Rendering › image-based rendering
depth-image-based rendering
0.712023
CollabVr: Reprojection-Based Edge-Client Collaborative Rendering for Real-Time High-Quality Mobile Virtual Reality · RTSS 2023
Virtual and augmented reality › virtual reality
mobile virtual reality
0.712023
CollabVr: Reprojection-Based Edge-Client Collaborative Rendering for Real-Time High-Quality Mobile Virtual Reality · RTSS 2023
Machine learning › Generative modeling
diffusion model
0.312025
FlexiTex: Enhancing Texture Generation via Visual Guidance · AAAI 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
0.312025
FlexiTex: Enhancing Texture Generation via Visual Guidance · AAAI 2025
Geometric modeling and processing
implicit neural representation
0.212024
DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes · CVPR 2024

Methods — techniques the papers use, named apart from their topics

visual guidance enhancement · 1.7two-stream optimization · 1.7transformer · 1.7direction-aware adaptation · 1.7diffusion model · 1.73d gaussian splatting · 1.7foveated patch subdivision · 1.3interpolation sampling · 0.8implicit neural representation · 0.8cross-channel attention · 0.8parallel streaming · 0.7
YearPublicationVenuePosition
2025 4D-Editor: Interactive Object-Level Editing in Dynamic Neural Radiance Fields via Semantic Distillation
abstract
This paper targets interactive object-level editing (e.g., deletion, recoloring, transformation, composition) in dynamic scenes. Recently, some methods aiming for flexible editing static scenes represented by neural radiance field (NeRF) have shown impressive synthesis quality, while similar capabilities in time-variant dynamic scenes remain limited. To solve this problem, we propose 4D-Editor, an interactive semantic-driven editing framework, allowing editing multiple objects in a dynamic NeRF with user strokes on a single frame. Specifically, we extend the original dynamic NeRF by incorporating Hybrid Semantic Feature Distillation to maintain spatial-temporal consistency after editing. In addition, a Recursive Selection Refinement module is presented to significantly boost object segmentation accuracy within a dynamic NeRF to aid the editing process. Moreover, we develop Multi-view Reprojection Inpainting to fill holes caused by incomplete scene capture after editing. Extensive quantitative and qualitative experiments on real application scenarios demonstrate that 4D-Editor achieves photo-realistic editing on dynamic NeRFs. Project page: https://patrickddj.github.ioI/4D-Editor
Dadong Jiang, Zhihui Ke, Xiaobo Zhou 0003, Tie Qiu 0001, Xidong Shi
3DV1
2025 FlexiTex: Enhancing Texture Generation via Visual Guidance
abstract
Recent texture generation methods achieve impressive results due to the powerful generative prior they leverage from large-scale text-to-image diffusion models. However, abstract textual prompts are limited in providing global textural or shape information, which results in the texture generation methods producing blurry or inconsistent patterns. To tackle this, we present FlexiTex, embedding rich information via visual guidance to generate a high-quality texture. The core of FlexiTex is the Visual Guidance Enhancement module, which incorporates more specific information from visual guidance to reduce ambiguity in the text prompt and preserve high-frequency details. To further enhance the visual guidance, we introduce a Direction-Aware Adaptation module that automatically designs direction prompts based on different camera poses, avoiding the Janus problem and maintaining semantically global consistency. Benefiting from the visual guidance, FlexiTex produces quantitatively and qualitatively sound results, demonstrating its potential to advance texture generation for real-world applications.
Dadong Jiang, Xianghui Yang, Zibo Zhao 0001, Zeqiang Lai, Shaoxiong Yang, Chunchao Guo, Xiaobo Zhou 0003, Zhihui Ke
AAAI1
2025 Timeformer: Capturing Temporal Relationships of Deformable 3D Gaussians for Robust Reconstruction
abstract
Dynamic scene reconstruction is a long-term challenge in 3D vision. Recent methods extend 3D Gaussian Splatting to dynamic scenes via additional deformation fields and apply explicit constraints like motion flow to guide the deformation. However, they learn motion changes from individual timestamps independently, making it challenging to reconstruct complex scenes, particularly when dealing with violent movement, extreme-shaped geometries, or reflective surfaces. To address the above issue, we design a plug-and-play module called TimeFormer to enable existing deformable 3D Gaussians reconstruction methods with the ability to implicitly model motion patterns from a learning perspective. Specifically, TimeFormer includes a Cross-Temporal Transformer Encoder, which adaptively learns the temporal relationships of deformable 3D Gaussians. Furthermore, we propose a two-stream optimization strategy that transfers the motion knowledge learned from TimeFormer to the base stream during the training phase. This allows us to remove TimeFormer during inference, thereby preserving the original rendering speed. Extensive experiments in the multi-view and monocular dynamic scenes validate qualitative and quantitative improvement brought by TimeFormer. Project Page: https://patrickddj.github.io/TimeFormer/
Dadong Jiang, Zhi Hou, Zhihui Ke, Xianghui Yang, Xiaobo Zhou 0003, Tie Qiu 0001
ICCV1
2024 DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes
abstract
Implicit neural representations for video (NeRV) have recently become a novel way for high-quality video representation. However, existing works employ a single network to represent the entire video, which implicitly con-fuse static and dynamic information. This leads to an inability to effectively compress the redundant static information and lack the explicitly modeling of global temporal-coherent dynamic details. To solve above problems, we propose DS-NeRV, which decomposes videos into sparse learnable static codes and dynamic codes without the need for explicit optical flow or residual supervision. By setting different sampling rates for two codes and applying weighted sum and interpolation sampling methods, DS-NeRV efficiently utilizes redundant static information while maintaining high-frequency details. Additionally, we design a cross-channel attention-based (CCA) fusion module to efficiently fuse these two codes for frame decoding. Our approach achieves a high quality reconstruction of 31.2 PSNR with only 0.35M parameters thanks to separate static and dynamic codes representation and outperforms existing NeRV methods in many downstream tasks. Our project website is at https://haoyan14.github.io/DS-NeRV/.
Zhihui Ke, Xiaobo Zhou 0003, Tie Qiu 0001, Xidong Shi, Dadong Jiang
CVPR6
2023 CollabVr: Reprojection-Based Edge-Client Collaborative Rendering for Real-Time High-Quality Mobile Virtual Reality
abstract
Collaborative mobile virtual reality (VR) has recently emerged as a promising solution to provide an immersive user experience with low motion-to-photo (MTP) latency. The rendering tasks are usually divided into background and foreground ones, which are executed in the edge server and head-mounted display (HMD), respectively. Assuming that the background images are static, they can be reused for temporal redundancy reduction in transmission. However, in high dynamic high-quality scenes, background images are continuously changing, making the temporal reuse strategy ineffective, leading to high MTP latency and hence motion sickness. In this paper, we propose CollabVR, a reprojection-based edge-client collaborative rendering approach for real-time high-quality mobile VR in high dynamic scenes. The key idea is to reduce the spatial redundancy in transmission by exploiting the high similarity between the left view and right view. With CollabVR, only one view is rendered and encoded in the edge server and then be transmitted to, and decoded at the HMD. The another view is reprojected by utilizing depth image-based rendering (DIBR) in the HMD, thereby greatly reducing the MTP latency even in high dynamic scenes. Furthermore, we propose a foveated-based multi-level patch subdivision strategy to achieve real-time reprojection in the resourced-limited HMD. A parallel streaming strategy is also proposed to fill holes that exist in the reprojected image. Experiments we conducted using Commercial Off-The-Shelf (COTS) devices indicate that CollabVR can reduce the average MTP latency by up to 36% compared to the baseline methods.
Zhihui Ke, Xiaobo Zhou 0003, Dadong Jiang, Tie Qiu 0001
RTSS3