VLDB 2026 Research / reviewers in the wild / expert
Zhiyong Su
dblp:56/3816
· DBLP profile ↗
32ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0001-9483-5268ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OpenSGG-VL: Open-Vocabulary 3DSGG with Orthogonal Residual Fusion and Iterative Relation RefinementabstractOpen-vocabulary 3D scene graph generation (3DSGG) aims to predict object categories and relation triplets from 3D scans while generalizing beyond a fixed label set. Prior open-vocabulary methods commonly rely on 2D features as an intermediate bridge to connect 3D representations with language, which can introduce misalignment and unstable multimodal fusion. In this work, we propose OpenSGG-VL, a framework that learns text-aligned 3D instance embeddings via large-scale 3D-Text contrastive learning, further enhanced by lightweight pose features to retain spatial cues. At inference, we extract 2D embeddings from RGB views and fuse them with 3D embeddings using Orthogonal Residual Fusion (ORF), which preserves the dominant semantic direction while injecting complementary geometric residuals. For open-set relation prediction, we employ an LLM and formulate relation inference as an iterative, scene-consistent refinement process with grouped relation decoding and multi-stage optimization. Experiments on 3DSSG demonstrate that our method achieves strong open-vocabulary performance and produces more coherent scene graphs than prior baselines. Liang Xiao 0001, Zhiyong Su |
ICMR | 4 |
| 2026 | A general framework for interactive semantic segmentation refinement of point clouds
Jinsheng Sun, Zhiyong Su |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | CASA-SDF: Curriculum-aware spatial adaptation with curvature-guided density for neural implicit surface reconstruction
Zhiyong Su, Liang Xiao 0001 |
Neurocomputing | 3 |
| 2026 | Open-World Point Cloud Semantic Segmentation: A Human-in-the-Loop FrameworkabstractOpen-world point cloud semantic segmentation (OW-Seg) aims to predict point labels for both base and novel classes in real-world scenarios. However, existing methods depend on additional samples to support the prediction on query sample, which is labor-intensive to collect and annotate. Some methods further rely on additional learning stages to adapt to novel classes, limiting their practicality in dynamic scenarios. In addition, the intra-class distribution shifts across samples introduce biased class representations (prototypes), resulting in sub-optimal predictions. To address these limitations, we propose HOW-Seg, the first human-in-the-loop framework for OW-Seg. Instead of relying on additional samples, HOW-Seg constructs class prototypes in the query sample feature space based on sparse point-level annotations, thereby avoiding cross-sample distribution shifts. Considering the lack of granularity of initial prototypes, we introduce an interactive prototype disambiguation mechanism to refine ambiguous prototypes. To further enrich contextual awareness, we propose a prototype label assignment module, which employs a dense conditional random field (CRF) upon the prototypes to optimize their label assignments. Through iterative human feedback, HOW-Seg dynamically improves its predictions, achieving high-quality segmentation for both base and novel classes. Experiments demonstrate that with sparse annotations (e.g., one-class-one-click), HOW-Seg surpasses the state-of-the-art generalized few-shot segmentation (GFS-Seg) method under the 5-shot setting. When using advanced backbones (e.g., Stratified Transformer) and denser annotations (e.g., 10 clicks), HOW-Seg achieves 85.27% mIoU on S3DIS and 66.37% mIoU on ScanNetv2, significantly outperforming other alternatives. The source code will be publicly available at https://github.com/Pengz98/HOW-Seg. Songru Yang, Jinsheng Sun, Zhiyong Su |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | UGD: An Unsupervised Geometric Distance for Evaluating Real-World Noisy Point Cloud DenoisingabstractExisting quantitative evaluation metrics for point cloud denoising methods require both the denoised point cloud and the corresponding ground-truth clean point cloud to compute a representative geometric distance. This requirement is highly problematic in real-world scenarios, where ground-truth clean point clouds are often unavailable. In this paper, we propose a simple yet effective unsupervised geometric distance (UGD) for real-world noisy point cloud denoising, calculated solely from noisy point clouds. The core idea is to learn a patch-wise prior model from a set of clean point clouds and then employ this prior model as the ground-truth to quantify the degradation by measuring the geometric variations of the denoised point cloud. To this end, we first learn a pristine Gaussian Mixture Model (GMM) with extracted patch-wise quality-aware features from a set of pristine clean point clouds by a patch-wise feature extraction network, which serves as the ground-truth for the quantitative evaluation. Then, the UGD is defined as the weighted sum of distances between each patch of the denoised point cloud and the learned pristine GMM model in the patch space. To train the employed patch-wise feature extraction network, we propose a self-supervised training framework through multi-task learning, which includes pair-wise quality ranking, distortion classification, and distortion distribution prediction. Quantitative experiments with synthetic noise confirm that the proposed UGD achieves comparable performance to supervised full-reference metrics. Moreover, experimental results on real-world data demonstrate that the proposed UGD enables unsupervised evaluation of point cloud denoising methods based exclusively on noisy point clouds. Zhiyong Su, Jincan Wu, Yanke Li |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Simultaneous Vision-Language Knowledge Transfer for Zero-Shot Human-Object Interaction Detection
Zhiyong Su, Liang Xiao 0001 |
PRCV (12) | 3 |
| 2025 | SITF: A Self-Supervised Iterative Training Framework for Point Cloud Denoising
Zhiyong Su, Changchang Wang |
Comput. Aided Des. | 1 |
| 2025 | No-reference geometry quality assessment for colorless point clouds via list-wise rank learning
Bingxu Xie, Chao Chu, Zhiyong Su |
Comput. Graph. | 5 |
| 2025 | Hypergraph convolutional network based weakly supervised point cloud semantic segmentation with scene-level annotations
Zhuheng Lu, Yuewei Dai, Zhiyong Su |
Neurocomputing | 5 |
| 2025 | The Worse the Better: Content-Aware Viewpoint Generation Network for Projection-Related Point Cloud Quality AssessmentabstractExisting projection-related point cloud quality assessment (PCQA) methods commonly adopt a straightforward but content-independent projection strategy, which selects a certain number of viewpoints to obtain projected images of degraded point clouds for further assessment. Through experimental studies, however, we observed the instability of final predicted quality scores, which change significantly over different viewpoint settings. Inspired by the “wooden barrel theory”, given the default content-independent viewpoints of existing projection-related PCQA approaches, this paper presents a novel content-aware viewpoint generation network (CAVGN) to learn better viewpoints by taking the distribution of geometric and attribute features of degraded point clouds into consideration. Firstly, the proposed CAVGN extracts multi-scale geometric and texture features of the entire input point cloud, respectively. Then, for each default content-independent viewpoint, the extracted geometric and texture features are refined to focus on its corresponding visible part of the input point cloud. Finally, the refined geometric and texture features are concatenated to generate an optimized viewpoint. To train the proposed CAVGN, we present a self-supervised viewpoint ranking network (SSVRN) to select the viewpoint with the worst quality projected image to construct a default-optimized viewpoint dataset, which consists of thousands of paired default viewpoints and corresponding optimized viewpoints. Experimental results show that the projection-related PCQA methods can achieve higher performance using the viewpoints generated by the proposed CAVGN. The source code can be found athttps://github.com/yokeno1/CAVGN1. Zhiyong Su, Bingxu Xie, Jincan Wu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Fine-Grained Metrics for Point Cloud Semantic Segmentation
Zhuheng Lu, Yuewei Dai, Zhiyong Su |
PRCV (11) | 5 |
| 2024 | Mesh deformation-based single-view 3D reconstruction of thin eyeglasses frames with differentiable renderingabstractWith the support of Virtual Reality (VR) and Augmented Reality (AR) technologies, the 3D virtual eyeglasses try-on application is well on its way to becoming a new trending solution that offers a "try on" option to select the perfect pair of eyeglasses at the comfort of your own home. Reconstructing eyeglasses frames from a single image with traditional depth and image-based methods is extremely difficult due to their unique characteristics such as lack of sufficient texture features, thin elements, and severe self-occlusions. In this paper, we propose the first mesh deformation-based reconstruction framework for recovering high-precision 3D full-frame eyeglasses models from a single RGB image, leveraging prior and domain-specific knowledge. Specifically, based on the construction of a synthetic eyeglasses frame dataset, we first define a class-specific eyeglasses frame template with pre-defined keypoints. Then, given an input eyeglasses frame image with thin structure and few texture features, we design a keypoint detector and refiner to detect predefined keypoints in a coarse-to-fine manner to estimate the camera pose accurately. After that, using differentiable rendering, we propose a novel optimization approach for producing correct geometry by progressively performing free-form deformation (FFD) on the template mesh. We define a series of loss functions to enforce consistency between the rendered result and the corresponding RGB input, utilizing constraints from inherent structure, silhouettes, keypoints, per-pixel shading information, and so on. Experimental results on both the synthetic dataset and real images demonstrate the effectiveness of the proposed algorithm. Ziyue Ji, Weiguang Kang, Zhiyong Su |
Graph. Model. | 5 |
| 2023 | Joint data and feature augmentation for self-supervised representation learning on point cloudsabstractTo deal with the exhausting annotations, self-supervised representation learning from unlabeled point clouds has drawn much attention, especially centered on augmentation-based contrastive methods. However, specific augmentations hardly produce sufficient transferability to high-level tasks on different datasets. Besides, augmentations on point clouds may also change underlying semantics. To address the issues, we propose a simple but efficient augmentation fusion contrastive learning framework to combine data augmentations in Euclidean space and feature augmentations in feature space. In particular, we propose a data augmentation method based on sampling and graph generation. Meanwhile, we design a data augmentation network to enable a correspondence of representations by maximizing consistency between augmented graph pairs. We further design a feature augmentation network that encourages the model to learn representations invariant to the perturbations using an encoder perturbation. We comprehensively conduct extensive object classification experiments and object part segmentation experiments to validate the transferability of the proposed framework. Experimental results demonstrate that the proposed framework is effective to learn the point cloud representation in a self-supervised manner, and yields state-of-the-art results in the community. The source code is publicly available at: https://github.com/VCG-NJUST/AFSRL. Zhuheng Lu, Yuewei Dai, Zhiyong Su |
Graph. Model. | 4 |
| 2023 | Video driven adaptive grasp planning of virtual hand using deep reinforcement learning
Yihe Wu, Zhenning Zhang, Dong Qiu, Zhiyong Su |
Multim. Tools Appl. | 5 |
| 2022 | Structure-Aware Denoising for Real-world Noisy Point Clouds with Complex Structures
Guoxing Sun 0002, Chao Chu, Jialin Mei, Zhiyong Su |
Comput. Aided Des. | 5 |
| 2022 | A weakly supervised framework for real-world point cloud classification
An Deng, Yunchao Wu, Zhuheng Lu, Zhiyong Su |
Comput. Graph. | 6 |
| 2022 | Point cloud denoising review: from classical to deep learning-based approaches
Guoxing Sun 0002, Zhiyong Su |
Graph. Model. | 5 |
| 2022 | Imitative Collaboration: A mirror-neuron inspired mixed reality collaboration method with remote hands and local replicas
Zhenning Zhang, Zhiyong Su |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | Slicing-Tracking-Detection: Simultaneous Multi-Cylinder Detection From Large-Scale and Complex Point CloudsabstractMultiple cylinders detection from large-scale and complex point clouds is a historical but challenging problem, considering the efficiency and accuracy. We propose a novel framework, named slicing-tracking-detection (STD), that detects multiple cylinders accurately and simultaneously from point clouds of large-scale and complex process plants. In this framework, the 3D cylinder detection problem is reformulated as a cylinder ingredients tracking task based on multi-object tracking (MOT). First, we generate slices from the input point cloud, and render them to slice sequence. Then, the cycle of a cylinder is modeled with a Markov Decision Process (MDP), where the ingredient is tracked with a template and the miss tracking is associated with ingredient proposals through reinforcement learning. Finally, by applying MDP for each cylinder, multiple cylinders can be detected simultaneously and accurately. Extensive experiments show that the proposed STD framework can significantly outperform the state-of-the-art approaches in efficiency, accuracy, and robustness. The source code is available at http://zhiyongsu.github.io. Zhuheng Lu, Weiwei Mao, Yuewei Dai, Zhiyong Su |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | An augmented reality-based multimedia environment for experimental education
Zhenning Zhang, Zichen Li, Zhiyong Su |
Multim. Tools Appl. | 4 |
| 2021 | Learning to Hash for Personalized Image AuthenticationabstractThis paper takes a fresh look at the image authentication problem and proposes an alternative framework for personalized authentication based on the hash learning technology. Conventional image authentication methods tend to provide a general authentication framework for all images with the fixed quantization strategy and fixed control parameters determined based on given limited images and attacks. However, they may suffer from more or less misjudgments in practice, not to mention performance degradation when encountering out-of-sample images. Instead of proposing a new feature extraction algorithm, a novel personalized authentication framework which incorporates the distance metric learning technology and supervised quantization strategy to the process of image authentication is proposed in this paper. The tamper detection task is reformulated as a new supervised manipulation classification problem. For each input image, various content-preserving and content-changing samples are generated automatically firstly. Then, feature representations of all samples can be obtained by existing feature extraction methods. After that, a weighted large margin for manipulation classification (WLMMC) scheme is proposed to learn an effective feature mapping space to improve the classification performance between content-changing samples and content-preserving samples. During the quantization stage, a novel supervised personalized quantization strategy (SPQ), which is motivated by the observation that different attacks have different degrees of influence on feature components, is proposed to learn more compact yet discriminative binary codes for each input image. Effectiveness of the proposed framework is qualitatively and quantitatively demonstrated on a variety of images. Extensive experiments show that the proposed framework can significantly improve the authentication performance over the state-of-the-art techniques while achieve more compact hash codes flexibly as required. Zhiyong Su, Jialin Mei |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Robust 2D Engineering CAD Graphics Hashing for Joint Topology and Geometry Authentication via Covariance-Based DescriptorsabstractThis paper investigates the joint authentication of topology and geometry information of 2D engineering computer-aided design graphics, which focus more on topological modeling than geometric modeling of objects. A robust hashing scheme is proposed for joint topology and geometry authentication. The covariance matrices of descriptors are explored to fuse and encode both topology and geometry features of different types into a compact representation. First, a normalized binary shape texture is rendered for each geometric object through the render-to-texture technique. Then, for each geometric object, geometry features are computed based on statistical features that are extracted from image rings. Additionally, topology features are generated according to the topological relations among joint objects. To generate hash codes of the graphic, all geometric objects are first grouped according to their geometry features. Then, for each group, the covariance matrices of descriptors are applied to fuse both the topology and geometry features of all objects, and the intermediate hash codes of each group are computed based on the covariance matrices. The final hash sequence is formed by concatenating the intermediate hash codes that correspond to each group. Secret keys are introduced into both feature extraction and hash construction. The hashes are robust against topology-preserving graphic manipulations and sensitive to malicious attacks. By decomposing the hashes, the locations of tampered objects can be determined. Experimental results are presented to evaluate the performance and show the effectiveness of the method. Zhiyong Su, Yuewei Dai |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | Topology based 2D engineering drawing and 3D model matching for process plant
Rui Wen 0005, Weiqing Tang, Zhiyong Su |
Graph. Model. | 3 |
| 2017 | A unified framework for authenticating topology integrity of 2D heterogeneous engineering CAD drawings
Zhiyong Su, Yaobin Mao, Yuewei Dai, Weiqing Tang |
Multim. Tools Appl. | 1 |
| 2015 | Topology authentication for piping isometric drawings
Zhiyong Su, Guangjie Liu 0001, Weiqing Tang |
Comput. Aided Des. | 1 |
| 2014 | Authenticating topological integrity of process plant models through digital watermarking
Zhiyong Su, Guangjie Liu 0001, Jianshou Kong, Yuewei Dai |
Multim. Tools Appl. | 1 |
| 2013 | Watermarking 3D CAPD models for topology verification
Zhiyong Su, Jianshou Kong, Yuewei Dai, Weiqing Tang |
Comput. Aided Des. | 1 |
| 2013 | Topology authentication for CAPD models based on Laplacian coordinates
Zhiyong Su, Yuewei Dai, Weiqing Tang |
Comput. Graph. | 1 |
| 2010 | Evaluating the Reliability of Web Services Based on BPEL Code Structure Analysis and Run-Time Information CaptureabstractIn this article, an approach is proposed to evaluate the reliability of Web services, where three kinds of Web services are discussed, they are atomic service without the structural information, structural activity based composite service which is composed from other services by using a structural activity mechanism, and BPEL flow process based composite service which composes all kinds of atomic services and activity based composite services using BPEL language in an orchestration style. Firstly, the reliability of atomic service is evaluated based on an extended UDDI model, then, the reliability of activity based composite services is evaluated using BPEL code structure analysis, the reliability of atomic service, and function transition probability, finally, the reliability of BPEL flow process based composite web service was evaluated by a recursive algorithm. Case study and experimental results show the significance of the approach. Bixin Li, Xiaocong Fan, Zhiyong Su |
APSEC | 4 |
| 2010 | WS-PSC Monitor: A Tool Chain for Monitoring Temporal and Timing Properties in Composite Service Based on Property Sequence Chart
Pengcheng Zhang 0001, Zhiyong Su, Yuelong Zhu, Wenrui Li 0002, Bixin Li |
RV | 2 |
| 2008 | Extending PSC for Monitoring the Timed Properties in Composite ServicesabstractDue to the dynamically evolving attribute, validation of composite services must be extended from design time to run-time. Dynamical verification techniques, such as runtime monitoring, have been first class activities to be performed during the execution of composite services. For a kind of composite services, nonfunctional properties, such as timed properties, are as important as functional properties and need to be monitored at run-time. However, using traditional logic and formalism, these timed properties are not easily represented for general software engineers. In order to deal with this problem, we first extend a novel notation (Property Sequence Chart) with time constructs. Then, we give its semantics in terms of timed Buchi automata and measure its expressiveness based on recently proposed real-time specification patterns. Finally, we propose a novel framework to monitor two kinds of timed properties in composite services: the accomplished time of basic service operations and some additional timed assumptions of the composition process. Our framework provides a completely graphical front-end which can friendly help general software engineers to monitor the timed properties in composite services. Pengcheng Zhang 0001, Bixin Li, Zhiyong Su, Mingjie Sun |
APSEC | 3 |
| 2008 | A user-oriented Web service reliability modelabstractIn this article, a user-oriented software reliability model was proposed to evaluate the reliability of Web services Two kinds of Web services were discussed: atomic services without the structural information and the composite services consisting of atomic services. Firstly, the reliability of atomic service was evaluated based on an extended UDDI model. Then, the overall reliability of composite service was evaluated using BPEL structure chart-based model and the reliabilities of all atomic services being used. A case study is designed, implemented, and analyzed to support our model. The experimental results show the significances of the model. Bixin Li, Zhiyong Su, Xufang Gong |
SMC | 2 |