Zhendong Yang

dblp:14/1820 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 GeoVoronoi: A Voronoi Diagram Generation System for Large-Scale Geographical Point Data via Spatial Attribute Association
abstract
Voronoi diagram is commonly used to visualize geographical point dataset with a collection of plane-partitioned facets. As the size of the geographical point dataset increases, facets are densely distributed, and present different sizes and irregular shapes, leading to overdrawing and confusion problems, and hampering the visual perception of Voronoi diagram and insightful exploration of geographical point data. In this paper, we propose a novel Voronoi diagram generation framework to visualize and explore large-scale geographical point datasets. Firstly, an attribute-based blue noise sampling model is designed to select a subset of points to generate the simplified Voronoi diagram, retaining both the spatial distribution and attribute relationship of the original large-scale geographical points. Then a couple of optimization schemes are integrated into the sampling model to replace the representative points, aiming to enhance the visual perception of Voronoi diagram, such as shape balance and color characterization. Furthermore, we implement an interactive online Voronoi diagram generation tool, GeoVoronoi, enabling users to generate meaningful facets according to their requirements. Quantitative comparisons, case studies and user studies based on real-world datasets have demonstrated the effectiveness of our proposed method in the generation of credible Voronoi diagram and in-depth exploration of geographical point datasets.
Zhiguang Zhou, Haoxuan Wang 0001, Zhendong Yang, Yuanyuan Chen 0013, Ying Lai, Wei Chen 0001, Yuwei Meng
IEEE Trans. Big Data3
2024 CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
abstract
Recent vision-language models have achieved tremendous advances. However, their computational costs are also escalating dramatically, making model acceleration exceedingly critical. To pursue more efficient vision-language Transformers, this paper introduces Cross-Guided Ensemble of Tokens (CrossGET), a general acceleration framework for vision-language Transformers. This framework adaptively combines tokens in real-time during inference, significantly reducing computational costs while maintaining high performance. CrossGET features two primary innovations: 1) Cross-Guided Matching and Ensemble. CrossGET leverages cross-modal guided token matching and ensemble to effectively utilize cross-modal information, achieving wider applicability across both modality-independent models, e.g., CLIP, and modality-dependent ones, e.g., BLIP2. 2) Complete-Graph Soft Matching. CrossGET introduces an algorithm for the token-matching mechanism, ensuring reliable matching results while facilitating parallelizability and high efficiency. Extensive experiments have been conducted on various vision-language tasks, such as image-text retrieval, visual reasoning, image captioning, and visual question answering. The performance on both classic multimodal architectures and emerging multimodal LLMs demonstrates the framework’s effectiveness and versatility. The code is available at https://github.com/sdc17/CrossGET.
Dachuan Shi, Chaofan Tao, Anyi Rao, Zhendong Yang, Chun Yuan 0003, Jiaqi Wang 0003
ICML4
2024 Spatio-temporal adaptive convolution and bidirectional motion difference fusion for video action recognition
Mingwei Tang, Zhendong Yang, Jie Hu 0007, Mingfeng Zhao
Expert Syst. Appl.3
2023 Transparent Shape from a Single View Polarization Image
abstract
This paper presents a learning-based method for transparent surface estimation from a single view polarization image. Existing shape from polarization(SfP) methods have the difficulty in estimating transparent shape since the inherent transmission interference heavily reduces the reliability of physics-based prior. To address this challenge, we propose the concept of physics-based prior confidence, which is inspired by the characteristic that the transmission component in the polarization image has more noise than reflection. The confidence is used to determine the contribution of the interfered physics-based prior. Then, we build a network(TransSfP) with multi-branch architecture to avoid the destruction of relationships between different hierarchical inputs. To train and test our method, we construct a dataset for transparent shape from polarization with paired polarization images and ground-truth normal maps. Extensive experiments and comparisons demonstrate the superior accuracy of our method. Our cdataset and code are publicly available at https://github.com/shaomq2187/TransSfP
Mingqi Shao, Chongkun Xia, Zhendong Yang, Junnan Huang, Xueqian Wang 0001
ICCV3
2023 From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft Labels
abstract
Knowledge Distillation (KD) uses the teacher’s logits as soft labels to guide the student, while self-KD does not need a real teacher to require the soft labels. This work unifies the formulations of the two tasks by decomposing and reorganizing the generic KD loss into a Normalized KD (NKD) loss and customized soft labels for both target class (image’s category) and non-target classes named Universal Self-KD (USKD). We decompose the KD loss and find the non-target loss from it forces the student’s non-target logits to match the teacher’s, but the sum of the two non-target logits is different, preventing them from being identical. NKD normalizes the non-target logits to equalize their sum. It can be generally used for KD and self-KD to better use the soft labels for distillation. USKD generates customized soft labels for both target and non-target classes without a teacher. It smooths the target logit of the student as the soft target label and uses the rank of the intermediate feature to generate the soft non-target labels with Zipf’s law. For KD with teachers, NKD achieves state-of-the-art performance on CIFAR-100 and ImageNet, boosting the ImageNet Top-1 accuracy of Res-18 from 69.90% to 71.96% with a Res-34 teacher. For self-KD without teachers, USKD is the first method that can be effectively applied to both CNN and ViT models with negligible additional time and memory cost, resulting in new state-of-the-art results, such as 1.17% and 0.55% accuracy gains on ImageNet for MobileNet and DeiT-Tiny, respectively. Code is available at https://github.com/yzd-v/cls_KD.
Zhendong Yang, Ailing Zeng, Tianke Zhang, Chun Yuan 0003, Yu Li 0003
ICCV1
2023 Accurate 3D Face Reconstruction with Facial Component Tokens
abstract
Accurately reconstructing 3D faces from monocular images and videos is crucial for various applications, such as digital avatar creation. However, the current deep learning-based methods face significant challenges in achieving accurate reconstruction with disentangled facial parameters and ensuring temporal stability in single-frame methods for 3D face tracking on video data. In this paper, we propose TokenFace, a transformer-based monocular 3D face reconstruction model. TokenFace uses separate tokens for different facial components to capture information about different facial parameters and employs temporal transformers to capture temporal information from video data. This design can naturally disentangle different facial components and is flexible to both 2D and 3D training data. Trained on hybrid 2D and 3D data, our model shows its power in accurately reconstructing faces from images and producing stable results for video data. Experimental results on popular benchmarks NoWand Stirling demonstrate that TokenFace achieves state-of-the-art performance, outperforming existing methods on all metrics by a large margin.
Tianke Zhang, Xuangeng Chu, Yunfei Liu 0001, Lijian Lin, Zhendong Yang, Zhengzhuo Xu, Chengkun Cao, F. Richard Yu, Changyin Zhou, Chun Yuan 0003, Yu Li 0003
ICCV5
2023 UPop: Unified and Progressive Pruning for Compressing Vision-Language Transformers
abstract
Real-world data contains a vast amount of multimodal information, among which vision and language are the two most representative modalities. Moreover, increasingly heavier models, e.g., Transformers, have attracted the attention of researchers to model compression. However, how to compress multimodal models, especially vison-language Transformers, is still under-explored. This paper proposes the Unified and Progressive Pruning (UPop) as a universal vison-language Transformer compression framework, which incorporates 1) unifiedly searching multimodal subnets in a continuous optimization space from the original model, which enables automatic assignment of pruning ratios among compressible modalities and structures; 2) progressively searching and retraining the subnet, which maintains convergence between the search and retrain to attain higher compression ratios. Experiments on various tasks, datasets, and model architectures demonstrate the effectiveness and versatility of the proposed UPop framework. The code is available at https://github.com/sdc17/UPop.
Dachuan Shi, Chaofan Tao, Zhendong Yang, Chun Yuan 0003, Jiaqi Wang 0003
ICML4
2022 Focal and Global Knowledge Distillation for Detectors
abstract
Knowledge distillation has been applied to image classification successfully. However, object detection is much more sophisticated and most knowledge distillation methods have failed on it. In this paper, we point out that in object detection, the features of the teacher and student vary greatly in different areas, especially in the foreground and background. If we distill them equally, the uneven differences between feature maps will negatively affect the distillation. Thus, we propose Focal and Global Distillation (FGD). Focal distillation separates the foreground and background, forcing the student to focus on the teacher's critical pixels and channels. Global distillation rebuilds the relation between different pixels and transfers it from teachers to students, compensating for missing global information in focal distillation. As our method only needs to calculate the loss on the feature map, FGD can be applied to various detectors. We experiment on various detectors with different backbones and the results show that the student detector achieves excellent mAP improvement. For example, ResNet-50 based RetinaNet, Faster RCNN, RepPoints and Mask RCNN with our distillation method achieve 40.7%, 42.0%, 42.0% and 42.1% mAP on COCO2017, which are 3.3, 3.6, 3.4 and 2.9 higher than the baseline, respectively. Our codes are available at https://github.com/yzd-v/FGD.
Zhendong Yang, Xiaohu Jiang, Yuan Gong 0002, Zehuan Yuan, Danpei Zhao, Chun Yuan 0003
CVPR1
2022 Masked Generative Distillation
Zhendong Yang, Mingqi Shao, Dachuan Shi, Zehuan Yuan, Chun Yuan 0003
ECCV (11)1
2020 SHAMC: A Secure and highly available database system in multi-cloud environment
Liangmin Wang 0001, Zhendong Yang, Xiangmei Song
Future Gener. Comput. Syst.2
2008 Robust Subspace Clustering by Logarithmic Hyperbolic Cosine Function
abstract
As an important category of clustering methods, subspace clustering algorithms have arisen particular attention during the last decade. Most subspace clustering algorithms are designed by first constructing a similarity matrix and then using spectral clustering algorithms to perform clustering. How to learn a suitable representation matrix to construct the similarity matrix is essential to the clustering performance. In most existing algorithms, the representation matrix is solved by norm-minimization, which commonly enforces the error matrix with nuclear norm or sparsity norm. However, these methods may fail to achieve satisfactory performance for real data contaminated by complex noise. To this end, we propose a novel robust subspace clustering method based on the Logarithmic Hyperbolic Cosine Function (LHCF). We theoretically analyze the grouping effect, as well as the convergence behavior, which illustrates that highly correlated samples can be grouped into the same cluster. Experimental results conducted on the Extended Yale B dataset show that the newly proposed algorithm yields better clustering performance compared with some advanced methods.
Long Shi 0002, Jun Wang 0089, Zhendong Yang, Badong Chen
IEEE Signal Process. Lett.4