EDBT 2026 Demo / reviewers in the wild / expert
Zhenzhong Kuang
dblp:144/1438 · also Zhengzhong Kuang
· DBLP profile ↗
47ranked-venue papers
16as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 12 first-author · 18 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 10 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IE-SRGS: An Internal-External Knowledge Fusion Framework for High-Fidelity 3D Gaussian Splatting Super-ResolutionabstractReconstructing high-resolution (HR) 3D Gaussian Splatting (3DGS) models from low-resolution (LR) inputs remains challenging due to the lack of fine-grained textures and geometry. Existing methods typically rely on pre-trained 2D super-resolution (2DSR) models to enhance textures, but suffer from 3D Gaussian ambiguity arising from cross-view inconsistencies and domain gaps inherent in 2DSR models. We propose IE-SRGS, a novel 3DGS SR paradigm that addresses this issue by jointly leveraging the complementary strengths of external 2DSR priors and internal 3DGS features. Specifically, we use 2DSR and depth estimation models to generate HR images and depth maps as external knowledge, and employ multi-scale 3DGS models to produce cross-view consistent, domain-adaptive counterparts as internal knowledge. A mask-guided fusion strategy is introduced to integrate these two sources and synergistically exploit their complementary strengths, effectively guiding the 3D Gaussian optimization toward high-fidelity reconstruction. Extensive experiments on both synthetic and real-world benchmarks show that IE-SRGS consistently outperforms state-of-the-art methods in both quantitative accuracy and visual fidelity. Tieshi Zhong, Shuo Chang, Weiliu Wang, Chengkai Wang, Yifei Chen 0019, Tongyu Hu, Zhenzhong Kuang, Xuefei Yin, Yanming Zhu 0001 |
AAAI | 9 |
| 2026 | KF-GS: Kalman filter-guided Gaussian splatting for real-time high-quality dynamic scene reconstruction
Qingyuan Tang, Yufei Yin, Yanming Zhu 0001, Zhou Yu 0001, Zhenzhong Kuang, Jiajun Ding, Jifa He |
J. Vis. Commun. Image Represent. | 5 |
| 2026 | Compositional Text-to-Image Synthesis With Training-Free Layout-Guided DiffusionabstractRecent text-to-image (T2I) diffusion models have made significant strides in generating high-quality images from diverse textual prompts. Despite this progress, these models often face challenges in accurately understanding and synthesizing complex prompts, primarily due to their limited compositional capabilities. In this study, we propose a novel approach for compositional T2I synthesis using layout-guided diffusion models, which do not require additional training. Specifically, we leverage the chain-of-code prompting technique of large language models to interpret textual prompts and generate object layouts with spatial coherence. To enhance the alignment between generated images and textual descriptions, we introduce two innovative layoutguided loss functions: Patch-oriented Cross-Attention (PCA) loss and Region-oriented Cross-Attention (RCA) loss. The PCA loss emphasizes high activation values for image patches that attend to all tokens in the prompt across the layout. The RCA loss enhances the average attention within the layout, thereby increasing the accuracy of generating objects and their associated attributes within specified regions. These proposed loss functions reassign cross-attention in diffusion models during the denoising process. Our comprehensive experiments consistently demonstrate the effectiveness of our approach in improving semantic alignment between generated images and a diverse range of textual prompts, while ensuring high usability as a ready-to-use plugin. Our code is available athttps://github.com/gxl-groups/Compositional-T2I. Xiaoling Gu, Lingwei Luo, Shengqi Wu, Zizhao Wu, Zhenzhong Kuang, Zhou Yu 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | What we Need is Explicit Controllability: Training 3D Gaze Estimator Using Only Facial Images
Tingwei Li, Jun Bao, Zhenzhong Kuang, Buyu Liu |
ICCV | 3 |
| 2025 | Gaussian Splatting-based Scene Reconstruction with Adaptive Sampling and Region-based RenderingabstractRecently, 3D Gaussian Splatting (3DGS) based methods, such as Scaffold-GS, have exhibited state-of-the-art performance for high-fidelity scene reconstruction. However, existing methods may suffer from degraded details because they usually pay the same attention to the whole image set and their contents. In fact, some parts of the image data may contain more information than others. Thus, it is reasonable to treat them differently. Motivated by this, in this paper, we propose a new 3DGS-based method to adaptively choose the data that are more valuable for training. Our method consists of two key parts: Adaptive Sampling (AS) and Region-based Rendering (RR). AS focuses on collecting difficult data for 3DGS training. RR focuses on enabling training with partial image data. Our method can produce more fine-grained details for scene reconstruction and has good scalability. On multiple public datasets, we verify the effectiveness of the proposed method by conducting comparative and ablation experiments. The quantitative and qualitative experimental results show that our method has achieved state-of-the-art results compared with the baselines. Takahiko Furuya, Zhenzhong Kuang, Xiaoling Gu, Jiajun Ding |
IJCNN | 3 |
| 2025 | FutureGS: Structured Gaussian Fields for Future-Aware Dynamic Scene Modeling
Mingyang Ding, Tingting Han 0003, Jiajun Ding, Min Tan 0005, Zhenzhong Kuang |
ACM Multimedia | 8 |
| 2025 | FedGA: Federated Learning via Gradient Adaptive AggregationabstractIn modern lives, the rapid proliferation of Internet of Things (IoT) devices has made them indispensable tools for data collection and analysis across various domains. However, growing concerns over data ownership and privacy have hindered effective data sharing among IoT devices, leading to the persistent challenge of data silos. Federated Learning (FL) has emerged as a promising solution to this problem by enabling collaborative model training without direct data exchange. Despite its potential, FL faces two critical limitations: severe catastrophic forgetting for historical knowledge and inefficient average aggregation. To address these challenges, this paper proposes FedGA, an innovative FL framework that leverages cosine similarity-based weighted aggregation to enhance model convergence speed. Furthermore, FedGA incorporates a mechanism to memorize historical models, thereby significantly alleviating catastrophic forgetting. Extensive experiments on three public datasets validate the effectiveness of FedGA, demonstrating its superior performance in both accuracy and training efficiency compared to state-of-the-art methods. The results highlight FedGA’s capability to overcome the key shortcomings of existing FL approaches, making it a robust solution for practical IoT applications. Changfeng Hu, Min Tan 0005, Tingting Han 0003, Zhenzhong Kuang |
SMC | 5 |
| 2025 | ViewCloud: A lightweight multi-view point cloud representation for efficient 3D recognition and cross-domain retrieval
Zhihe Wu, Yaomin Wang, Zhenzhong Kuang, Jiajun Ding, Min Tan 0005, Xuefei Yin, Yanming Zhu 0001 |
Comput. Aided Des. | 3 |
| 2025 | Imp: Highly Capable Large Multimodal Models for Mobile DevicesabstractBy harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp—a family of highly capable LMMs at the 2B$\sim$4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s. Zhenwei Shao, Zhou Yu 0001, Jun Yu 0002, Xuecheng Ouyang, Lihao Zheng 0001, Zhenbiao Gai, Zhenzhong Kuang, Jiajun Ding |
IEEE Trans. Multim. | 8 |
| 2025 | VC-GS: view-consistent deblurring Gaussian splatting via alternating branch optimization
Qida Cao, Jiajun Ding, Zhenyang Liu, Zhenzhong Kuang, Yijie Shao, Yilan Shen |
Vis. Comput. | 4 |
| 2025 | Innovative AI techniques for photorealistic 3D clothed human reconstruction from monocular images or videos: a survey
Xiaoling Gu, Zhenzhong Kuang, Fei-wei Qin, Zizhao Wu |
Vis. Comput. | 3 |
| 2025 | Collaborative neural radiance fields for novel view synthesis
Junqing Yuan, Mengting Fan, Zhenyang Liu, Tongxuan Han, Zhenzhong Kuang, Chihao Pan, Jiajun Ding |
Vis. Comput. | 5 |
| 2024 | Facial Identity Anonymization via Intrinsic and Extrinsic Attention DistractionabstractThe unprecedented capture and application of face images raise increasing concerns on anonymization to fight against privacy disclosure. Most existing methods may suffer from the problem of excessive change of the identity-independent information or insufficient identity protection. In this paper, we present a new face anonymization approach by distracting the intrinsic and extrinsic identity attentions. On the one hand, we anonymize the identity information in the feature space by distracting the intrinsic identity attention. On the other, we anonymize the visual clues (i.e. appearance and geometry structure) by distracting the extrinsic identity attention. Our approach allows for flexible and intuitive manipulation of face appearance and geometry structure to produce diverse results, and it can also be used to instruct users to perform personalized anonymization. We conduct extensive experiments on multiple datasets and demonstrate that our approach outperforms state-of-the-art methods. Zhenzhong Kuang, Yingjie Shen, Jun Yu 0002 |
CVPR | 1 |
| 2024 | Latent Representation Reorganization for Face Privacy ProtectionabstractThe issue of face privacy protection has aroused wide social concern along with the increasing applications of face images. The latest methods focus on achieving a good privacy-utility tradeoff so that the protected results can still be used to support the downstream computer vision tasks. However, they may suffer from limited flexibility in manipulating this tradeoff because the practical requirements may vary under different scenarios. In this paper, we present a novel recurrent latent representation reorganization (LReOrg) framework to deal with the problem. LReOrg relies on two key modules to deal with the privacy-utility tradeoff, where the first one is responsible for anonymizing the privacy sensitive information and the other is responsible for recovering the destroyed useful insensitive information according to user requirements. LReOrg is advantageous in: (a) enabling users to recurrently process fine-grained attributes; (b) providing flexible control over privacy-utility tradeoff by manipulating which attributes to anonymize or preserve using cross-modal keywords; and (c) eliminating the need of data annotations for network training. The experimental results on benchmark datasets have reported the superior ability of our approach for providing flexible protection on facial information. Zhenzhong Kuang, Jianan Lu, Chenhui Hong, Haobin Huang, Suguo Zhu, Jun Yu 0002, Jianping Fan 0007 |
ACM Multimedia | 1 |
| 2024 | Self-supervised learning of rotation-invariant 3D point set features using transformer and its self-distillationabstractInvariance against rotations of 3D objects is an important property in analyzing 3D point set data. Conventional 3D point set DNNs having rotation invariance typically obtain accurate 3D shape features via supervised learning by using labeled 3D point sets as training samples. However, due to the rapid increase in 3D point set data and the high cost of labeling, a framework to learn rotation-invariant 3D shape features from numerous unlabeled 3D point sets is required. This paper proposes a novel self-supervised learning framework for acquiring accurate and rotation-invariant 3D point set features at object-level. Our proposed lightweight DNN architecture decomposes an input 3D point set into multiple global-scale regions, called tokens, that preserve the spatial layout of partial shapes composing the 3D object. We employ a self-attention mechanism to refine the tokens and aggregate them into an expressive rotation-invariant feature per 3D point set. Our DNN is effectively trained by using pseudo-labels generated by a self-distillation framework. To facilitate the learning of accurate features, we propose to combine multi-crop and cut-mix data augmentation techniques to diversify 3D point sets for training. Through a comprehensive evaluation, we empirically demonstrate that, (1) existing rotation-invariant DNN architectures designed for supervised learning do not necessarily learn accurate 3D shape features under a self-supervised learning scenario, and (2) our proposed algorithm learns rotation-invariant 3D point set features that are more accurate than those learned by existing algorithms. Takahiko Furuya, Zhoujie Chen, Ryutarou Ohbuchi, Zhenzhong Kuang |
Comput. Vis. Image Underst. | 4 |
| 2024 | ZS-SRT: An efficient zero-shot super-resolution training method for Neural Radiance Fields
Yongbo He, Chengkai Wang, Zhenzhong Kuang, Jiajun Ding, Fei-wei Qin, Jun Yu 0002, Jianping Fan 0001 |
Neurocomputing | 5 |
| 2024 | iDesigner: making intelligent fashion designs
Xiaoling Gu, Qiming Yao, Xiaojun Gong, Zhenzhong Kuang |
Multim. Tools Appl. | 4 |
| 2024 | Confidence correction for trained graph convolutional networksabstractAdopting Graph Convolutional Networks (GCNs) for transductive node classification is a hot research direction in artificial intelligence . Vanilla GCNs are primarily under-confident and struggle to clarify the final classification results explicitly due to the lack of supervision. Existing works mainly alleviated this issue by improving annotation deficiency and introducing addition regularization terms. However, these methods need to re-train the model from the beginning, which is computationally expensive for large dataset and model. To deal with this problem, a novel confidence correction mechanism (CCM) for trained GCNs is proposed in this work. Such mechanism aims at calibrating the confidence output of each node in the inference stage by jointly inferring the feature and predicted pseudo label. Specifically, in the inference stage, it uses the predicted pseudo label to select target-related features over all network to obtain a more confident and better result. Such selectivity is formulated as an optimization problem to maximize the category score of each node. In addition, the greedy optimization strategy is utilized to solve this problem and we have mathematically proven that the proposed mechanism can reach the local optimum by mathematical induction . Note that such mechanism is flexible and can be introduced to most GCN-based model. Extensive experimental results on benchmark datasets show that the proposed method can promote the confidence of the final target category and improve the performance of GCNs in the inference stage. Junqing Yuan, Huanlei Guo, Chenyi Zhou, Jiajun Ding, Zhenzhong Kuang, Zhou Yu 0001 |
Pattern Recognit. | 5 |
| 2024 | Semantic-aware hyper-space deformable neural radiance fields for facial avatar reconstruction
Kaixin Jin, Xiaoling Gu, Zhenzhong Kuang, Zizhao Wu, Min Tan 0005, Jun Yu 0002 |
Pattern Recognit. Lett. | 4 |
| 2023 | Hyperplane patch mixing-and-folding decoder and weighted chamfer distance loss for 3D point set reconstructionabstractAbstract 3D point set reconstruction is an important and challenging 3D shape analysis task. Current state-of-the-art algorithms for 3D point set reconstruction employ a deep neural network (DNN) having an encoder–decoder architecture. Recently, the decoder DNNs that transform multiple 2D planar patches to reconstruct a 3D shape have seen some success. These “patch-folding” decoders are adept at approximating smooth surfaces in 3D objects. However, 3D point sets generated by these decoders often lack local geometrical details, as 2D planar patches tend to overly constrain the patch folding process. In this paper, we propose a novel decoder DNN for 3D point sets called Hyperplane Mixing and Folding Net (HMF-Net). HMF-Net uses less constrained hyperplane, not 2D plane, patches as its input to the folding process. HMF-Net has, as its core building block, a stack of token-mixing layers to effectively learn global consistency among the hyperplane patches. In addition to HMF-Net, we also propose a novel loss for 3D point set reconstruction called Weighted Chamfer Distance (WCD). WCD tries to weight, or amplify, loss from parts of shape that are highly variable across training samples by emphasizing higher point-pair distance values between a generated point set and a groundtruth point set. This helps the decoder DNN learn shape details better. We comprehensively evaluate our algorithm under three 3D point set reconstruction scenarios, that are, shape completion, shape upsampling, and shape reconstruction from 2D images. Experimental results demonstrate that our algorithm yields accuracies higher than the existing algorithms for 3D point set reconstruction. Takahiko Furuya, Wujie Liu, Ryutarou Ohbuchi, Zhenzhong Kuang |
Vis. Comput. | 4 |
| 2022 | Delegate-based Utility Preserving Synthesis for Pedestrian Image AnonymizationabstractThe rapidly growing application of pedestrian images has aroused wide concern on visual privacy protection because personal information is under the risk of privacy disclosure. Anonymization is regarded as an effective solution by identity obfuscation. Most recent methods focus on face, but it is not enough when the presence of human body carries lots of identifiable information. This paper presents a new delegate-based utility preserving synthesis (DUPS) approach for pedestrian image anonymization. This is challenging because one may expect that the anonymized image can still be useful in various computer vision tasks. We model DUPS as an adaptive translation process from source to target. To provide a comprehensive identity protection, we first perform anonymous delegate sampling based on image-level differential privacy. To synthesize anonymous images, we then introduce an adaptive translation network and optimize it with a multi-task loss function. Our approach is theoretically sound and can generate diverse results by preserving data utility. The experiments on multiple datasets show that DUPS can not only achieve superior anonymization performance against deep pedestrian recognizers, but also can obtain a better tradeoff between privacy protection and utility preservation compared with state-of-the-art methods. Zhenzhong Kuang, Longbin Teng, Zhou Yu 0001, Jun Yu 0002, Jianping Fan 0001, Mingliang Xu 0001 |
ACM Multimedia | 1 |
| 2021 | Effective De-identification Generative Adversarial Network for Face AnonymizationabstractThe growing application of face images and modern AI technology has raised another important concern in privacy protection. In many real scenarios like scientific research, social sharing and commercial application, lots of images are released without privacy processing to protect people's identity. In this paper, we develop a novel effective de-identification generative adversarial network (DeIdGAN) for face anonymization by seamlessly replacing a given face image with a different synthesized yet realistic one. Our approach consists of two steps. First, we anonymize the input face to obfuscate its original identity. Then, we use our designed de-identification generator to synthesize an anonymized face. During the training process, we leverage a pair of identity-adversarial discriminators to explicitly constrain identity protection by pushing the synthesized face away from the predefined sensitive faces to resist re-identification and identity invasion. Finally, we validate the effectiveness of our approach on public datasets. Compared with existing methods, our approach can not only achieve better identity protection rates but also preserve superior image quality and data reusability, which suggests the state-of-the-art performance. Zhenzhong Kuang, Huigui Liu, Jun Yu 0002, Aikui Tian, Jianping Fan 0001, Noboru Babaguchi |
ACM Multimedia | 1 |
| 2021 | Reproducibility Companion Paper: Campus3D: A Photogrammetry Point Cloud Benchmark for Outdoor Scene Hierarchical UnderstandingabstractThis companion paper is to support the replication of paper "Campus3D: A Photogrammetry Point Cloud Benchmark for Outdoor Scene Hierarchical Understanding", which was presented at ACM Multimedia 2020. The supported paper's main purpose was to provide a photogrammetry point cloud-based dataset with hierarchical multilabels to facilitate the area of 3D deep learning. Based on this provided dataset and source code, in this work, we build a complete package to reimplement the proposed methods and experiments (i.e., the hierarchical learning framework and the benchmarks of the hierarchical semantic segmentation task). Specifically, this paper contains the technical details of the package, including file structure, dataset preparation, installation package, and the conduction of the experiment. We also present the replicated experiment results and indicate our contributions to the original implementation. Yuqing Liao, Zekun Tong, Yabang Zhao, Andrew Lim 0001, Zhenzhong Kuang, Cise Midoglu |
ACM Multimedia | 6 |
| 2021 | Reproducibility Companion Paper: Visual Relation of Interest DetectionabstractIn this companion paper, we provide the details of the reproducibility artifacts of the paper "Visual Relation of Interest Detection" presented at MM'20. Visual Relation of Interest Detection (VROID) aims to detect visual relations that are important for conveying the main content of an image. In this paper, we explain the file structure of the source code and publish the details of our ViROI dataset, which can be used to retrain the model with custom parameters. We also detail the scripts for component analysis and comparison with other methods and list the parameters that can be modified for custom training and inference. Fan Yu 0003, Tongwei Ren, Jinhui Tang 0001, Gangshan Wu, Jingjing Chen 0001, Zhenzhong Kuang |
ACM Multimedia | 7 |
| 2021 | Unnoticeable synthetic face replacement for image privacy protection
Zhenzhong Kuang, Zhiqiang Guo, Jinglong Fang, Jun Yu 0002, Noboru Babaguchi, Jianping Fan 0001 |
Neurocomputing | 1 |
| 2021 | Deep embedding of concept ontology for hierarchical fashion recognition
Zhenzhong Kuang, Xin Zhang 0063, Jun Yu 0002, Zongmin Li, Jianping Fan 0001 |
Neurocomputing | 1 |
| 2020 | Aggregating diverse deep attention networks for large-scale plant species identification
Haixi Zhang, Zhenzhong Kuang, Xianlin Peng, Guiqing He, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 2 |
| 2020 | Partial person re-identification with two-stream network and reconstruction
Suguo Zhu, Xiaowei Gong, Zhenzhong Kuang, Junping Du 0001 |
Neurocomputing | 3 |
| 2019 | Deep Mixture of Diverse Experts for Large-Scale Visual RecognitionabstractIn this paper, a deep mixture of diverse experts algorithm is developed to achieve more efficient learning of a huge (mixture) network for large-scale visual recognition application. First, a two-layer ontology is constructed to assign large numbers of atomic object classes into a set of task groups according to the similarities of their learning complexities, where certain degrees of inter-group task overlapping are allowed to enable sufficient inter-group message passing. Second, one particular base deep CNNs with M+1 outputs is learned for each task group to recognize its M atomic object classes and identify one special class of "not-in-group", where the network structure (numbers of layers and units in each layer) of the well-designed deep CNNs (such as AlexNet, VGG, GoogleNet, ResNet) is directly used to configure such base deep CNNs. For enhancing the separability of the atomic object classes in the same task group, two approaches are developed to learn more discriminative base deep CNNs: (a) our deep multi-task learning algorithm that can effectively exploit the inter-class visual similarities; (b) our two-layer network cascade approach that can improve the accuracy rates for the hard object classes at certain degrees while effectively maintaining the high accuracy rates for the easy ones. Finally, all these complementary base deep CNNs with diverse but overlapped outputs are seamlessly combined to generate a mixture network with larger outputs for recognizing tens of thousands of atomic object classes. Our experimental results have demonstrated that our deep mixture of diverse experts algorithm can achieve very competitive results on large-scale visual recognition. Qiuyu Chen, Zhenzhong Kuang, Jun Yu 0002, Wei Zhang 0016, Jianping Fan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Effective 3-D Shape Retrieval by Integrating Traditional Descriptors and Pointwise ConvolutionabstractThe applications of isometric 3-D objects have recently received sufficient attention and, thus, it is very attractive to retrieve such isometric 3-D objects from large-scale collections. Although existing approaches have presented some interesting ideas, their performance is limited to their ability on feature representation. To improve the performance of 3-D object (shape) recognition, some recent algorithms prefer using complicated deep neural networks to learn discriminative features, but they consume huge amounts of computing resources. Instead, this paper presents a more effective solution by seamlessly integrating the traditional local descriptor with a deep pointwise convolutional network to extract 1-D features for shape recognition and retrieval. To reduce the costs of designing a complicated deep network, the first step of our algorithm is to describe the shape deformation by sampling a set of intrinsic point descriptors. Then, we introduce a simple yet effective pointwise convolutional network to integrate these descriptors as a global feature and the learning process can be significantly accelerated with the help of downsampling. Furthermore, a knowledge transfer strategy is used to upgrade our feature by compensating for information loss. Finally, we carry out experimental evaluations over popular shape benchmarks, and the results suggest that our approach exhibits superior accuracy rates and robustness on shape recognition and retrieval. Zhenzhong Kuang, Jun Yu 0002, Suguo Zhu, Zongmin Li, Jianping Fan 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Deep Point Convolutional Approach for 3D Model RetrievalabstractWith the increasing popularity of 3D models, retrieving deformable 3D objects is becoming a crucial task. The state-of-the-art methods use complex deep neural networks to address this problem, which require lots of computational resources. In this paper, we develop a more effective solution by using point convolution. Our algorithm takes local point descriptors as the input and produces a global vector for shape retrieval. To save the efforts of designing complex deep convolutional neural network (CNN), we first use intrinsic point descriptors to describe the shape deformations. Then, a simple but effective point CNN network is developed to integrate the local shape information by performing subspace compression and fusion, which depends on an end-to-end learning process to link the local and global information for discriminative shape representation. The experimental results on popular benchmarks have verified that our algorithm is able to outperform the state-of-the-art methods. Zhenzhong Kuang, Jun Yu 0002, Jianping Fan 0001, Min Tan 0005 |
ICME | 1 |
| 2018 | Integrating multi-level deep learning and concept ontology for large-scale visual recognition
Zhenzhong Kuang, Jun Yu 0002, Zongmin Li, Baopeng Zhang, Jianping Fan 0001 |
Pattern Recognit. | 1 |
| 2018 | Leveraging Content Sensitiveness and User Trustworthiness to Recommend Fine-Grained Privacy Settings for Social Image SharingabstractTo configure successful privacy settings for social image sharing, two issues are inseparable: 1) content sensitiveness of the images being shared; and 2) trustworthiness of the users being granted to see the images. This paper aims to consider these two inseparable issues simultaneously to recommend fine-grained privacy settings for social image sharing. For achieving more compact representation of image content sensitiveness (privacy), two approaches are developed: 1) a deep network is adapted to extract 1024-D discriminative deep features; and 2) a deep multiple instance learning algorithm is adopted to identify 280 privacy-sensitive object classes and events. Second, users on the social network are clustered into a set of representative social groups to generate a discriminative dictionary for user trustworthiness characterization. Finally, both the image content sensitiveness and the user trustworthiness are integrated to train a tree classifier to recommend fine-grained privacy settings for social image sharing. Our experimental studies have demonstrated both the efficiency and the effectiveness of our proposed algorithms. Jun Yu 0002, Zhenzhong Kuang, Baopeng Zhang, Wei Zhang 0016, Dan Lin 0001, Jianping Fan 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2017 | Deep Mixture of Experts with Diverse Task SpacesabstractIn this paper, a deep mixture algorithm is developed to support large-scale visual recognition (e.g., recognizing tens of thousands of object classes) by seamlessly combining a set of base deep CNNs (AlexNet) with diverse task spaces, e.g., such base deep CNNs (i.e., diverse experts) are trained to recognize different subsets of tens of thousands of object classes rather than the same set of object classes. Our experimental results have demonstrated that our deep mixture algorithm can achieve very competitive results on large-scale visual recognition. Jianping Fan 0001, Zhenzhong Kuang, Zhou Yu 0001, Jun Yu 0002 |
ICMLA | 3 |
| 2017 | Privacy Setting Recommendation for Image SharingabstractThis paper aims to simultaneously consider two inseparable issues for privacy setting recommendation: (1) sensitiveness of visual content of the images being shared; and (2) trustworthiness of users being granted. First, an object-based approach is developed for image content sensitiveness (privacy) representation. Secondly, the users on a social network are clustered into a set of representative social groups to generate a discriminative dictionary for user trustworthiness characterization. Finally, a tree classifier is trained hierarchically to recommend appropriate privacy settings for image sharing. Jun Yu 0002, Zhenzhong Kuang, Zhou Yu 0001, Dan Lin 0001, Jianping Fan 0001 |
ICMLA | 2 |
| 2017 | iPrivacy: Image Privacy Protection by Identifying Sensitive Objects via Deep Multi-Task LearningabstractTo achieve automatic recommendation of privacy settings for image sharing, a new tool called iPrivacy (image privacy) is developed for releasing the burden from users on setting the privacy preferences when they share their images for special moments. Specifically, this paper consists of the following contributions: 1) massive social images and their privacy settings are leveraged to learn the object-privacy relatedness effectively and identify a set of privacy-sensitive object classes automatically; 2) a deep multi-task learning algorithm is developed to jointly learn more representative deep convolutional neural networks and more discriminative tree classifier, so that we can achieve fast and accurate detection of large numbers of privacy-sensitive object classes; 3) automatic recommendation of privacy settings for image sharing can be achieved by detecting the underlying privacy-sensitive objects from the images being shared, recognizing their classes, and identifying their privacy settings according to the object-privacy relatedness; and 4) one simple solution for image privacy protection is provided by blurring the privacy-sensitive objects automatically. We have conducted extensive experimental studies on real-world images and the results have demonstrated both the efficiency and effectiveness of our proposed approach. Jun Yu 0002, Baopeng Zhang, Zhenzhong Kuang, Dan Lin 0001, Jianping Fan 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2017 | HD-MTL: Hierarchical Deep Multi-Task Learning for Large-Scale Visual RecognitionabstractIn this paper, a hierarchical deep multi-task learning (HD-MTL) algorithm is developed to support large-scale visual recognition (e.g., recognizing thousands or even tens of thousands of atomic object classes automatically). First, multiple sets of multi-level deep features are extracted from different layers of deep convolutional neural networks (deep CNNs), and they are used to achieve more effective accomplishment of the coarseto- fine tasks for hierarchical visual recognition. A visual tree is then learned by assigning the visually-similar atomic object classes with similar learning complexities into the same group, which can provide a good environment for determining the interrelated learning tasks automatically. By leveraging the inter-task relatedness (inter-class similarities) to learn more discriminative group-specific deep representations, our deep multi-task learning algorithm can train more discriminative node classifiers for distinguishing the visually-similar atomic object classes effectively. Our hierarchical deep multi-task learning (HD-MTL) algorithm can integrate two discriminative regularization terms to control the inter-level error propagation effectively, and it can provide an end-to-end approach for jointly learning more representative deep CNNs (for image representation) and more discriminative tree classifier (for large-scale visual recognition) and updating them simultaneously. Our incremental deep learning algorithms can effectively adapt both the deep CNNs and the tree classifier to the new training images and the new object classes. Our experimental results have demonstrated that our HD-MTL algorithm can achieve very competitive results on improving the accuracy rates for large-scale visual recognition. Jianping Fan 0001, Zhenzhong Kuang, Yu Zheng 0006, Ji Zhang 0005, Jun Yu 0002, Jinye Peng 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Multiscale shape context and re-ranking for deformable shape retrieval
Zongmin Li, Zhenzhong Kuang, Yujie Liu 0002, Jiayan Wang |
Comput. Graph. | 2 |
| 2016 | Shape similarity assessment based on partial feature aggregation and ranking lists
Zhenzhong Kuang, Zongmin Li, Yujie Liu 0002, Changqing Zou |
Pattern Recognit. Lett. | 1 |
| 2015 | A More Effective Method for Image Representation: Topic Model Based on Latent Dirichlet AllocationabstractNowadays, the Bag-of-words (BoW) representation is well applied to recent state-of-the-art image retrieval works. However, with the rapid growth in the number of images, the dimension of the dictionary increases substantially which leads to great storage and CPU cost. Besides, the local features do not convey any semantic information which is very important in image retrieval. In this paper, we propose to use "topics" instead of "visual words" as the image representation by topic model to reduce the feature dimension and mine more high level semantic information. We call this as Bag-of-Topics (BoT) which is a type of statistical model for discovering the abstract "topics" from the words. We extract the topics by Latent Dirichlet Allocation (LDA) and calculate the similarity between the images using BoT model instead of BoW directly. The results show that the dimension of the image representation has been reduced significantly, while the retrieval performance is improved. Zongmin Li, Yante Li, Zhenzhong Kuang, Yujie Liu 0002 |
CAD/Graphics | 4 |
| 2015 | Discovering the Latent Similarities of the KNN Graph by Metric TransformationabstractThe manifold of the dataset turns out to be quite useful in refining the retrieval results, and the diffusion process provides an efficient solution by careful selection of the similarity neighborhood which is usually modeled as the K-nearest neighborhood (KNN) graph. However, existing works are sensitive to the topology noises induced by the first K neighbors. In this paper, we tackle the problem by studying metric transformation which aims at finding new functional relationship to dig the latent similarity. The advantage of the approach lies in its robustness towards the varying K values; that is to say, it could preserve high similarity performances even if K is very large. Except for discussing only the global KNN (i.e. the same K for all neighborhoods) graph, we also investigate to specify a different K for each neighborhood by incorporating the new penalized consensus information (PCI). We show that PCI works superior compared with the original consensus information for denoising. Experiments on multiple affinity matrices have corroborated the superiority of our method with surprising good results. Zhenzhong Kuang, Zongmin Li, Jianping Fan 0001 |
ICMR | 1 |
| 2015 | Retrieval of non-rigid 3D shapes from multiple aspects
Zhenzhong Kuang, Zongmin Li, Xiaxia Jiang, Yujie Liu 0002, Hua Li 0009 |
Comput. Aided Des. | 1 |
| 2015 | Modal function transformation for isometric 3D shape representation
Zhenzhong Kuang, Zongmin Li, Yujie Liu 0002 |
Comput. Graph. | 1 |
| 2015 | Exploration in improving retrieval quality and robustness for deformable non-rigid 3D shapes
Zhenzhong Kuang, Zongmin Li, Xiaxia Jiang, Yujie Liu 0002 |
Multim. Tools Appl. | 1 |
| 2014 | Graph Contexts for Retrieving Deformable Non-rigid 3D ShapesabstractDeformable non-rigid 3D shape retrieval plays an important role in various applications. Although there are many related works, their precision and robustness are not ideal. In this paper, we develop a novel retrieval method by using graph contexts, which consists of three steps. Initially, we evaluate the performance of spectral distances for deformable shape representation, which has not been studied in detail before. Then, we create a weighted L2distance for similarity measurement based on the spectra of Laplace-Beltrami operator. Finally, a new local graph diffusion method is introduced to reduce the mismatch error in feature space and the time cost of diffusion has reduced a lot simultaneously. Our experiment results on SHREC'11 Non-rigid dataset have reached the best reported retrieval performance (MAP: 99.9%). Zhenzhong Kuang, Zongmin Li, Xiaxia Jiang, Yujie Liu 0002 |
ICPR | 1 |
| 2014 | Evidence-based SVM fusion for 3D model retrieval
Zongmin Li, Zhenzhong Kuang, Yongzhou Gan, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2013 | GBI-SA: GBI Feature with Subtle Adjustment for Robust Non-rigid 3D Shape RetrievalabstractDeformable shape retrieval has posed a challenge to researchers, which is becoming more and more important. This paper addresses non-rigid 3D shape retrieval in terms of a subtle adjustment strategy. For this purpose, we first construct a new global shape descriptor based on biharmonic distance, which uses the connectivity between vertices for natural shape representation. Then, in order to enhance retrieval ability, we propose a subtle adjustment method to provide discriminative information. Finally, we depend on experiment to evaluate the performance of the proposed method. Comparison results have proved that the proposed GBI-SA retrieval method is superior to state-of-the-art methods, including SD-GDM, mesh SIFT, MDS-CM-BOF, Shape DNA, and WESD. Zhenzhong Kuang, Zongmin Li, Yujie Liu 0002 |
CAD/Graphics | 1 |