VLDB 2026 Research / reviewers in the wild / expert
Yan Zhang 0057
dblp:04/3348-57
· DBLP profile ↗
17ranked-venue papers
0as first author
7since 2021 · last 2025
0000-0002-9621-7321ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ANIR: Adaptive Neural Implicit Representation for 3D shape reconstruction and generation
Kun Liu 0021, Yan Zhang 0057, Yanwen Guo 0001, Jie Guo 0001 |
Comput. Aided Des. | 2 |
| 2024 | LiDAR-Net: A Real-Scanned 3D Point Cloud Dataset for Indoor ScenesabstractIn this paper, we present LiDAR-Net, a new real-scanned indoor point cloud dataset, containing nearly 3.6 billion precisely point-level annotated points, covering an expansive area of 30,000m2. It encompasses three prevalent daily environments, including learning scenes, working scenes, and living scenes. LiDAR-Net is characterized by its non-uniform point distribution, e.g., scanning holes and scanning lines. Additionally, it meticulously records and an-notates scanning anomalies, including reflection noise and ghost. These anomalies stem from specular reflections on glass or metal, as well as distortions due to moving persons. LiDAR-Net's realistic representation of non-uniform distribution and anomalies significantly enhances the training of deep learning models, leading to improved generalization in practical applications. We thoroughly evaluate the performance of state-of-the-art algorithms on LiDAR-Net and provide a detailed analysis of the results. Crucially, our research identifies several fundamental challenges in understanding indoor point clouds, contributing essential insights to future explorations in this field. Our dataset can be found online: http://lidar-net.njumeta.com. Yanwen Guo 0001, Yuanqi Li, Dayong Ren, Xiaohong Zhang 0009, Liang Pu, Changfeng Ma, Xiaoyu Zhan, Jie Guo 0001, Mingqiang Wei, Yan Zhang 0057, Piaopiao Yu, Shuangyu Yang, Donghao Ji, Huisheng Ye |
CVPR | 11 |
| 2024 | GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive LocalizationabstractAbstract With the emergence of large‐scale Text‐to‐Image(T2I) models and implicit 3D representations like Neural Radiance Fields (NeRF), many text‐driven generative editing methods based on NeRF have appeared. However, the implicit encoding of geometric and textural information poses challenges in accurately locating and controlling objects during editing. Recently, significant advancements have been made in the editing methods of 3D Gaussian Splatting, a real‐time rendering technology that relies on explicit representation. However, these methods still suffer from issues including inaccurate localization and limited manipulation over editing. To tackle these challenges, we propose GSEditPro, a novel 3D scene editing framework which allows users to perform various creative and precise editing using text prompts only. Leveraging the explicit nature of the 3D Gaussian distribution, we introduce an attention‐based progressive localization module to add semantic labels to each Gaussian during rendering. This enables precise localization on editing areas by classifying Gaussians based on their relevance to the editing prompts derived from cross‐attention layers of the T2I model. Furthermore, we present an innovative editing optimization method based on 3D Gaussian Splatting, obtaining stable and refined editing results through the guidance of Score Distillation Sampling and pseudo ground truth. We prove the efficacy of our method through extensive experiments. Yanhao Sun, Runze Tian, Xinyao Liu, Yan Zhang 0057, Kai Xu 0004 |
Comput. Graph. Forum | 5 |
| 2023 | Deep graph learning for spatially-varying indoor lighting prediction
Jiayang Bai, Jie Guo 0001, Zhenyu Chen 0001, Piaopiao Yu, Yan Zhang 0057, Yanwen Guo 0001 |
Sci. China Inf. Sci. | 8 |
| 2023 | Local-to-Global Panorama Inpainting for Locale-Aware Indoor Lighting PredictionabstractPredicting panoramic indoor lighting from a single perspective image is a fundamental but highly ill-posed problem in computer vision and graphics. To achieve locale-aware and robust prediction, this problem can be decomposed into three sub-tasks: depth-based image warping, panorama inpainting and high-dynamic-range (HDR) reconstruction, among which the success of panorama inpainting plays a key role. Recent methods mostly rely on convolutional neural networks (CNNs) to fill the missing contents in the warped panorama. However, they usually achieve suboptimal performance since the missing contents occupy a very large portion in the panoramic space while CNNs are plagued by limited receptive fields. The spatially-varying distortion in the spherical signals further increases the difficulty for conventional CNNs. To address these issues, we propose a local-to-global strategy for large-scale panorama inpainting. In our method, a depth-guided local inpainting is first applied on the warped panorama to fill small but dense holes. Then, a transformer-based network, dubbed PanoTransformer, is designed to hallucinate reasonable global structures in the large holes. To avoid distortion, we further employ cubemap projection in our design of PanoTransformer. The high-quality panorama recovered at any locale helps us to capture spatially-varying indoor illumination with physically-plausible global structures and fine details. Jiayang Bai, Jie Guo 0001, Zhenyu Chen 0001, Yan Zhang 0057, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | ShadowMover: Automatically Projecting Real Shadows onto Virtual ObjectabstractInserting 3D virtual objects into real-world images has many applications in photo editing and augmented reality. One key issue to ensure the reality of the composite whole scene is to generate consistent shadows between virtual and real objects. However, it is challenging to synthesize visually realistic shadows for virtual and real objects without any explicit geometric information of the real scene or manual intervention, especially for the shadows on the virtual objects projected by real objects. In view of this challenge, we present, to our knowledge, the first end-to-end solution to fully automatically project real shadows onto virtual objects for outdoor scenes. In our method, we introduce the Shifted Shadow Map, a new shadow representation that encodes the binary mask of shifted real shadows after inserting virtual objects in an image. Based on the shifted shadow map, we propose a CNN-based shadow generation model named ShadowMover which first predicts the shifted shadow map for an input image and then automatically generates plausible shadows on any inserted virtual object. A large-scale dataset is constructed to train the model. Our ShadowMover is robust to various scene configurations without relying on any geometric information of the real scene and is free of manual intervention. Extensive experiments validate the effectiveness of our method. Piaopiao Yu, Jie Guo 0001, Zhenyu Chen 0001, Chen Wang 0149, Yan Zhang 0057, Yanwen Guo 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | FAME: 3D Shape Generation via Functionality-Aware Model EvolutionabstractWe introduce a modeling tool which can evolve a set of 3D objects in a functionality-aware manner. Our goal is for the evolution to generate large and diverse sets of plausible 3D objects for data augmentation, constrained modeling, as well as open-ended exploration to possibly inspire new designs. Starting with an initial population of 3D objects belonging to one or more functional categories, we evolve the shapes through part recombination to produce generations of hybrids or crossbreeds between parents from the heterogeneous shape collection. Evolutionary selection of offsprings is guided both by a functional plausibility score derived from functionality analysis of shapes in the initial population and user preference, as in a design gallery. Since cross-category hybridization may result in offsprings not belonging to any of the known functional categories, we develop a means for functionality partial matching to evaluate functional plausibility on partial shapes. We show a variety of plausible hybrid shapes generated by our functionality-aware model evolution, which can complement existing datasets as training data and boost the performance of contemporary data-driven segmentation schemes, especially in challenging cases. Our tool supports constrained modeling, allowing users to restrict or steer the model evolution with functionality labels. At the same time, unexpected yet functional object prototypes can emerge during open-ended exploration owing to structure breaking when evolving a heterogeneous collection. Yanran Guan, Han Liu 0003, Kun Liu 0021, Kangxue Yin, Ruizhen Hu, Oliver van Kaick, Yan Zhang 0057, Ersin Yumer, Nathan Carr 0001, Radomír Mech, Hao (Richard) Zhang |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2020 | Progressive Scene Segmentation Based on Self-Attention MechanismabstractSemantic scene segmentation is vital for a large variety of applications as it enables understanding of 3D data. Nowadays, various approaches based upon point clouds ignore the mathematical distribution of points and treat the points equally. The methods following this direction neglect the imbalance problem of samples that naturally exists in scenes. To avoid these issues, we propose a two-stage semantic scene segmentation framework based on self-attention mechanism and achieved state-of-the-art performance on 3D scene understanding tasks. We split the whole task into two small ones which efficiently relief the sample imbalance issue. In addition, we have designed a new self-attention block which could be inserted into submanifold convolution networks to model the long-range dependencies that exists among points. The proposed network consists of an encoder and a decoder, with the spatial-wise and channel-wise attention modules inserted. The two-stage network shares a U-Net architecture and is an end-to-end trainable framework which could predict the semantic label for the scene point clouds fed into it. Experiments on standard benchmarks of 3D scenes implies that our network could perform at par or better than the existing state-of-the-art methods. Yunyi Pan, Yuan Gan, Kun Liu 0021, Yan Zhang 0057 |
ICPR | 4 |
| 2020 | Qualitative photo collage by quartet analysis and active learning
Yuan Gan, Yan Zhang 0057, Zhengxing Sun, Hao (Richard) Zhang |
Comput. Graph. | 2 |
| 2019 | PartNet: A Recursive Part Decomposition Network for Fine-Grained and Hierarchical Shape SegmentationabstractDeep learning approaches to 3D shape segmentation are typically formulated as a multi-class labeling problem. These models are trained for a fixed set of labels, which greatly limits their flexibility and adaptivity. We opt for top-down recursive decomposition and develop the first deep learning model for hierarchical segmentation of 3D shapes, based on recursive neural networks. Starting from a full shape represented as a point cloud, our model performs recursive binary decomposition, where the decomposition network at all nodes in the hierarchy share weights. At each node, a node classifier is trained to determine the type (adjacency or symmetry) and stopping criteria of its decomposition. The features extracted in higher level nodes are recursively propagated to lower level ones. Thus, the meaningful decompositions in higher levels provide strong contextual cues constraining the segmentations in lower levels. Meanwhile, to increase the segmentation accuracy at each node, we enhance the recursive contextual feature with the shape feature extracted for the corresponding part. Our method segments a 3D shape in point cloud into an arbitrary number of parts, depending on the shape complexity, showing strong generality and flexibility. It achieves the state-of-the-art performance, both for fine-grained and semantic segmentation, on the public benchmark and a new benchmark of fine-grained segmentation proposed in this work. We also demonstrate its application for fine-grained part refinements in image-to-shape reconstruction. Fenggen Yu, Kun Liu 0021, Yan Zhang 0057, Chenyang Zhu 0002, Kai Xu 0004 |
CVPR | 3 |
| 2019 | Qualitative Organization of Photo Collections via Quartet Analysis and Active Learning
Yuan Gan, Yan Zhang 0057, Zhengxing Sun, Hao (Richard) Zhang |
Graphics Interface | 2 |
| 2019 | VERAM: View-Enhanced Recurrent Attention Model for 3D Shape ClassificationabstractMulti-view deep neural network is perhaps the most successful approach in 3D shape classification. However, the fusion of multi-view features based on max or average pooling lacks a view selection mechanism, limiting its application in, e.g., multi-view active object recognition by a robot. This paper presents VERAM, a view-enhanced recurrent attention model capable of actively selecting a sequence of views for highly accurate 3D shape classification. VERAM addresses an important issue commonly found in existing attention-based models, i.e., the unbalanced training of the subnetworks corresponding to next view estimation and shape classification. The classification subnetwork is easily overfitted while the view estimation one is usually poorly trained, leading to a suboptimal classification performance. This is surmounted by three essential view-enhancement strategies: 1) enhancing the information flow of gradient backpropagation for the view estimation subnetwork, 2) devising a highly informative reward function for the reinforcement training of view estimation and 3) formulating a novel loss function that explicitly circumvents view duplication. Taking grayscale image as input and AlexNet as CNN architecture, VERAM with 9 views achieves instance-level and class-level accuracy of 95.5 and 95.3 percent on ModelNet10, 93.7 and 92.1 percent on ModelNet40, both are the state-of-the-art performance under the same number of views. Song-Le Chen, Yan Zhang 0057, Zhixin Sun, Kai Xu 0004 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | 3D Shape Segmentation Based on Viewpoint Entropy and Projective Fully Convolutional Networks Fusing Multi-view FeaturesabstractThis paper introduces an architecture for segmenting 3D shapes into labeled semantic parts. Our architecture combines viewpoint selection method based on viewpoint entropy, multi-view image-based Fully Convolutional Networks (FCNs) and graph cuts optimization method to yield coherent segmentation of 3D shapes. First, we select iteratively a fixed number of perspectives with the maximum viewpoint entropy from existing viewpoints that can cover the shape's triangles, to maximally and automatically adjust the distance between the viewpoint and the center point of the shape to make sure the shape projected to fill the render window as wide as possible. Second, the image-based FCN is used for efficient view-based reasoning about 3D shape parts. In this process, global features generated by max view pooling are concatenated with every single view's feature in the fully connected layer before upsampling. Then, the multi-view FCN outputs confidence maps per part, which are then input into the projection layer that contains the mapping relationship of every shape's triangles and their projective pixels' positions in the rendered images from selected perspectives. And then, the FCN outputs are projected back onto 3D shape surfaces. and max view pooling is applied to the output of the projection layer so that every triangle of each shape has a unique probability for each label. Finally, graph cuts algorithm is implemented for the final segmentation result. Panpan Shui, Pengyu Wang 0004, Fenggen Yu, Bingyang Hu, Yuan Gan, Kun Liu 0021, Yan Zhang 0057 |
ICPR | 7 |
| 2018 | 3D shape segmentation via shape fully convolutional networks
Pengyu Wang 0004, Yuan Gan, Panpan Shui, Fenggen Yu, Yan Zhang 0057, Song-Le Chen, Zhengxing Sun |
Comput. Graph. | 5 |
| 2018 | Corrigendum to "3D shape segmentation via shape fully convolutional networks" [Computers & Graphics 70 (2018) 128-139]
Pengyu Wang 0004, Yuan Gan, Panpan Shui, Fenggen Yu, Yan Zhang 0057, Song-Le Chen, Zhengxing Sun |
Comput. Graph. | 5 |
| 2018 | 3D shape segmentation via shape fully convolutional networks
Pengyu Wang 0004, Yuan Gan, Panpan Shui, Fenggen Yu, Yan Zhang 0057, Song-Le Chen, Zhengxing Sun |
Comput. Graph. | 5 |
| 2018 | Semi-Supervised Co-Analysis of 3D Shape Styles from Projected LinesabstractWe present a semi-supervised co-analysis method for learning 3D shape styles from projected feature lines , achieving style patch localization with only weak supervision. Given a collection of 3D shapes spanning multiple object categories and styles, we perform style co-analysis over projected feature lines of each 3D shape and then back-project the learned style features onto the 3D shapes. Our core analysis pipeline starts with mid-level patch sampling and pre-selection of candidate style patches. Projective features are then encoded via patch convolution. Multi-view feature integration and style clustering are carried out under the framework of partially shared latent factor (PSLF) learning, a multi-view feature learning scheme. PSLF achieves effective multi-view feature fusion by distilling and exploiting consistent and complementary feature information from multiple views, while also selecting style patches from the candidates. Our style analysis approach supports both unsupervised and semi-supervised analysis. For the latter, our method accepts both user-specified shape labels and style-ranked triplets as clustering constraints. We demonstrate results from 3D shape style analysis and patch localization as well as improvements over state-of-the-art methods. We also present several applications enabled by our style analysis. Fenggen Yu, Yan Zhang 0057, Kai Xu 0004, Ali Mahdavi-Amiri, Hao (Richard) Zhang |
ACM Trans. Graph. | 2 |