Yingji Zhong

dblp:42/4270 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-0946-1079ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 4 since 2021Computer networks · 3 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
3D vision · 50% Face, body and person analysis · 18% Generative modeling · 18%
Computer graphics and multimedia
4 papers
Rendering · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
novel view synthesis
1.932025
Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting · NeurIPS 2025
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs · CVPR 2024
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs · CVPR 2025
Rendering
neural radiance fields
1.822026
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views · AAAI 2026
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs · CVPR 2024
Rendering › neural radiance fields
sparse-input neural radiance fields
1.822026
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views · AAAI 2026
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs · CVPR 2024
Rendering
novel view synthesis
1.012026
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views · AAAI 2026
Computer vision › Face, body and person analysis
person re-identification
0.922021
Progressive Feature Enhancement for Person Re-Identification · IEEE Trans. Image Process. 2021
Robust Partial Matching for Person Search in the Wild · CVPR 2020
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting · NeurIPS 2025
Machine learning › Generative modeling
diffusion model
0.912025
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs · CVPR 2025
Computer vision › 3D vision › 3d reconstruction › multi-view reconstruction
sparse-view reconstruction
0.912025
Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting · NeurIPS 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs · CVPR 2025
Rendering › gaussian splatting
3d gaussian splatting
0.912025
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs · CVPR 2025
Rendering
neural rendering
0.912025
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs · CVPR 2025
Computer vision › 3D vision
3d scene reconstruction
0.812024
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs · CVPR 2024
Computer vision › Image recognition and object detection › visual attention modeling
attention-based feature enhancement
0.512021
Progressive Feature Enhancement for Person Re-Identification · IEEE Trans. Image Process. 2021
Machine learning › Deep learning architectures and training › multi-scale learning
multi-scale feature learning
0.512021
Progressive Feature Enhancement for Person Re-Identification · IEEE Trans. Image Process. 2021
Computer vision › Face, body and person analysis › person re-identification
partial matching
0.412020
Robust Partial Matching for Person Search in the Wild · CVPR 2020
Computer vision › Face, body and person analysis
person search
0.412020
Robust Partial Matching for Person Search in the Wild · CVPR 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.312026
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views · AAAI 2026
Computer vision › 3D vision › novel view synthesis
sparse view synthesis
0.312025
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs · CVPR 2025
Rendering › neural rendering
gaussian splatting rendering
0.312025
Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting · NeurIPS 2025
Computer vision › Image recognition and object detection › object detection
bounding box refinement
0.112020
Robust Partial Matching for Person Search in the Wild · CVPR 2020

Methods — techniques the papers use, named apart from their topics

learnable codebook · 2.0dual-level semantic guidance · 2.0bi-directional verification · 2.0attention mechanism · 2.0trajectory initialization · 1.7scene-grounding guidance · 1.7noise injection · 1.7gaussian dropout · 1.7transformer · 1.5contrastive loss · 1.5
YearPublicationVenuePosition
2026 Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views
abstract
Neural Radiance Fields (NeRF) have shown remarkable capabilities for photorealistic novel view synthesis. One major deficiency of NeRF is that dense inputs are typically required, and the rendering quality will drop drastically given sparse inputs. In this paper, we highlight the effectiveness of rendered semantics from dense novel views, and show that rendered semantics can be treated as a more robust form of augmented data than rendered RGB. Our method enhances NeRF’s performance by incorporating guidance derived from the rendered semantics. The rendered semantic guidance encompasses two levels: the supervision level and the feature level. The supervision-level guidance incorporates a bi-directional verification module that decides the validity of each rendered semantic label, while the feature-level guidance integrates a learnable codebook that encodes semantic-aware information, which is queried by each point via the attention mechanism to obtain semanticrelevant predictions. The overall semantic guidance is embedded into a self-improved pipeline.We also introduce a more challenging sparse-input indoor benchmark, where the number of inputs is limited to as few as 6. Experiments demonstrate the effectiveness of our method and it exhibits superior performance compared to existing approaches.
Yingji Zhong, Kaichen Zhou, Zhihao Li 0002, Lanqing Hong, Zhenguo Li, Dan Xu 0002
AAAI1
2025 Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs
abstract
Despite recent successes in novel view synthesis using 3D Gaussian Splatting (3DGS), modeling scenes with sparse inputs remains a challenge. In this work, we address two critical yet overlooked issues in real-world sparse-input modeling: extrapolation and occlusion. To tackle these issues, we propose to use a reconstruction by generation pipeline that leverages learned priors from video diffusion models to provide plausible interpretations for regions outside the field of view or occluded. However, the generated sequences exhibit inconsistencies that do not fully benefit subsequent 3DGS modeling. To address the challenge of inconsistencies, we introduce a novel scene-grounding guidance based on rendered sequences from an optimized 3DGS, which tames the diffusion model to generate consistent sequences. This guidance is training-free and does not require any fine-tuning of the diffusion model. To facilitate holistic scene modeling, we also propose a trajectory initialization method. It effectively identifies regions that are outside the field of view and occluded. We further design a scheme tailored for 3DGS optimization with generated sequences. Experiments demonstrate that our method significantly improves upon the baseline and achieves state-of-the-art performance on challenging benchmarks.
Yingji Zhong, Zhihao Li 0002, Dave Zhenyu Chen, Lanqing Hong, Dan Xu 0002
CVPR1
2025 Quantifying and Alleviating Co-Adaptation in Sparse-View 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has demonstrated impressive performance in novel view synthesis under dense-view settings. However, in sparse-view scenarios, despite the realistic renderings in training views, 3DGS occasionally manifests appearance artifacts in novel views. This paper investigates the appearance artifacts in sparse-view 3DGS and uncovers a core limitation of current approaches: the optimized Gaussians are overly-entangled with one another to aggressively fit the training views, which leads to a neglect of the real appearance distribution of the underlying scene and results in appearance artifacts in novel views. The analysis is based on a proposed metric, termed Co-Adaptation Score (CA), which quantifies the entanglement among Gaussians, i.e., co-adaptation, by computing the pixel-wise variance across multiple renderings of the same viewpoint, with different random subsets of Gaussians. The analysis reveals that the degree of co-adaptation is naturally alleviated as the number of training views increases. Based on the analysis, we propose two lightweight strategies to explicitly mitigate the co-adaptation in sparse-view 3DGS: (1) random gaussian dropout; (2) multiplicative noise injection to the opacity. Both strategies are designed to be plug-and-play, and their effectiveness is validated across various methods and benchmarks. We hope that our insights into the co-adaptation effect will inspire the community to achieve a more comprehensive understanding of sparse-view 3DGS.
Kangjie Chen, Yingji Zhong, Youyu Chen, Minghan Qin, Haoqian Wang
NeurIPS2
2024 CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
abstract
Neural Radiance Fields (NeRF) have shown impressive capabilities for photorealistic novel view synthesis when trained on dense inputs. However, when trained on sparse inputs, NeRF typically encounters issues of incorrect density or color predictions, mainly due to insufficient coverage of the scene causing partial and sparse supervision, thus leading to significant performance degradation. While existing works mainly consider ray-level consistency to construct 2D learning regularization based on rendered color, depth, or semantics on image planes, in this paper we propose a novel approach that models 3D spatial field consistency to improve NeRF's performance with sparse inputs. Specifically, we first adopt a voxel-based ray sampling strategy to ensure that the sampled rays intersect with a certain voxel in 3D space. We then randomly sample additional points within the voxel and apply a Transformer to infer the properties of other points on each ray, which are then incorporated into the volume rendering. By backpropagating through the rendering loss, we enhance the consistency among neighboring points. Additionally, we propose to use a contrastive loss on the encoder output of the Transformer to further improve consistency within each voxel. Exper-iments demonstrate that our method yields significant improvement over different radiance fields in the sparse inputs setting, and achieves comparable performance with current works. The project page for this paper is available at https://zhongyingji.github.io/CVT-xRF.
Yingji Zhong, Lanqing Hong, Zhenguo Li, Dan Xu 0002
CVPR1
2021 Progressive Feature Enhancement for Person Re-Identification
abstract
Most of person Re-Identification (ReID) works extract features from the top CNN layer for person image matching. The top CNN layer commonly corresponds to large receptive fields, thus is not effective in depicting visual cues at multiple scales, e.g., both global appearance and local details. This work proposes a Progressive Feature Enhancement (PFE) algorithm to spot and fuse multi-scale discriminative cues from different CNN layers into a single feature vector. The basic idea is to progressively learn complementary features with a layer-specific supervision from deep to shallow layers. The layer-specific supervision is inferred by the proposed Masked Feature Augmentation (MFA) module. For each CNN layer, MFA indicates cues that have been captured in its deeper layers. MFA hence supervises each layer to depict additional visual cues missed by its deeper layers. This framework effectively learns multi-scale features without requiring extra part annotations or dividing body parts. To further facilitate the layer-specific feature generation, a Two-Stage Attention Module (TSAM) is proposed to filter pixel-wise and channel-wise noises on intermediate feature maps. Extensive experiments on four ReID datasets show that our approach achieves competitive performance, e.g., with ResNet50 backbone, it achieves rank1 accuracy of 95.1%, 88.2%, 79.1% and 71.6% on Market-1501, DukeMTMC-ReID, MSMT17 and CUHK03 Detected, respectively, outperforming many state-of-the-art works.
Yingji Zhong, Yaowei Wang 0001, Shiliang Zhang
IEEE Trans. Image Process.1
2020 Robust Partial Matching for Person Search in the Wild
abstract
Various factors like occlusions, backgrounds, etc., would lead to misaligned detected bounding boxes , e.g., ones covering only portions of human body. This issue is common but overlooked by previous person search works. To alleviate this issue, this paper proposes an Align-to-Part Network (APNet) for person detection and re-Identification (reID). APNet refines detected bounding boxes to cover the estimated holistic body regions, from which discriminative part features can be extracted and aligned. Aligned part features naturally formulate reID as a partial feature matching procedure, where valid part features are selected for similarity computation, while part features on occluded or noisy regions are discarded. This design enhances the robustness of person search to real-world challenges with marginal computation overhead. This paper also contributes a Large-Scale dataset for Person Search in the wild (LSPS), which is by far the largest and the most challenging dataset for person search. Experiments show that APNet brings considerable performance improvement on LSPS. Meanwhile, it achieves competitive performance on existing person search benchmarks like CUHK-SYSU and PRW.
Yingji Zhong, Xiaoyu Wang 0002, Shiliang Zhang
CVPR1
2009 Cross layer multicarrier MIMO cognitive cooperation scheme for wireless hybrid ad hoc networks
Yingji Zhong, Kyung Sup Kwak, Dongfeng Yuan
Comput. Commun.1
2008 A novel cross layer game knowledge sharing algorithm based on neural fuzzy connection admission controller for cellular Ad Hoc networking
Yingji Zhong, Kyung Sup Kwak, Dongfeng Yuan
Comput. Commun.1
2007 A Novel Cross Layer Power Control Game Algorithm Based on Neural Fuzzy Connection Admission Controller in Cellular Ad Hoc Networks
Dongfeng Yuan, Yingji Zhong
ISNN (1)3
2006 Nonlinear optimization for energy efficiency in IEEE 802.11a wireless LANs
Dongfeng Yuan, Song Ci, Yingji Zhong
Comput. Commun.4
2005 A New QoS Routing Optimal Algorithm in Mobile Ad Hoc Networks Based on Hopfield Neural Network
Dongfeng Yuan, Song Ci, Yingji Zhong
ISNN (3)4