Zhihao Peng 0002

dblp:231/8964-2 · DBLP profile ↗
← Back
12ranked-venue papers
8as first author
10since 2021 · last 2025
0000-0001-8273-9527ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
abstract
Endoscopic procedures are essential for diagnosing and treating internal diseases, and multi-modal large language models (MLLMs) are increasingly applied to assist in endoscopy analysis. However, current benchmarks are limited, as they typically cover specific endoscopic scenarios and a small set of clinical tasks, failing to capture the real-world diversity of endoscopic scenarios and the full range of skills needed in clinical workflows. To address these issues, we introduce EndoBench, the first comprehensive benchmark specifically designed to assess MLLMs across the full spectrum of endoscopic practice with multi-dimensional capacities. EndoBench encompasses 4 distinct endoscopic scenarios, 12 specialized clinical tasks with 12 secondary subtasks, and 5 levels of visual prompting granularities, resulting in 6,832 rigorously validated VQA pairs from 21 diverse datasets. Our multi-dimensional evaluation framework mirrors the clinical workflow—spanning anatomical recognition, lesion analysis, spatial localization, and surgical operations—to holistically gauge the perceptual and diagnostic abilities of MLLMs in realistic scenarios. We benchmark 23 state-of-the-art models, including general-purpose, medical-specialized, and proprietary MLLMs, and establish human clinician performance as a reference standard. Our extensive experiments reveal: (1) proprietary MLLMs outperform open-source and medical-specialized models overall, but still trail human experts; (2) medical-domain supervised fine-tuning substantially boosts task-specific accuracy; and (3) model performance remains sensitive to prompt format and clinical task complexity. EndoBench establishes a new standard for evaluating and advancing MLLMs in endoscopy, highlighting both progress and persistent gaps between current models and expert clinical reasoning. We publicly release our benchmark and code.
Boyun Zheng, Wenting Chen, Zhihao Peng 0002, Zhenfei Yin, Jiancong Hu, Yixuan Yuan
NeurIPS4
2024 F2TNet: FMRI to T1w MRI Knowledge Transfer Network for Brain Multi-phenotype Prediction
Wuyang Li, Yu Jiang 0013, Zhihao Peng 0002, Pengyu Wang 0005, Xiang Li 0001, Tianming Liu 0001, Junwei Han 0001, Yixuan Yuan
MICCAI (11)4
2024 Hierarchical Graph Learning with Small-World Brain Connectomes for Cognitive Prediction
Yu Jiang 0013, Zhihao Peng 0002, Yixuan Yuan
MICCAI (5)3
2024 GBT: Geometric-Oriented Brain Transformer for Autism Diagnosis
Zhihao Peng 0002, Yu Jiang 0013, Pengyu Wang 0005, Yixuan Yuan
MICCAI (12)1
2024 fTSPL: Enhancing Brain Analysis with FMRI-Text Synergistic Prompt Learning
Pengyu Wang 0005, Huaqi Zhang, Zhihao Peng 0002, Yixuan Yuan
MICCAI (12)4
2023 Deep Attention-Guided Graph Clustering With Dual Self-Supervision
abstract
Existing deep embedding clustering methods fail to sufficiently utilize the available off-the-shelf information from feature embeddings and cluster assignments, limiting their performance. To this end, we propose a novel method, namely deep attention-guided graph clustering with dual self-supervision (DAGC). Specifically, DAGC first utilizes a heterogeneity-wise fusion module to adaptively integrate the features of the auto-encoder and the graph convolutional network in each layer and then uses a scale-wise fusion module to dynamically concatenate the multi-scale features in different layers. Such modules are capable of learning an informative feature embedding via an attention-based mechanism. In addition, we design a distribution-wise fusion module that leverages cluster assignments to acquire clustering results directly. To better explore the off-the-shelf information from the cluster assignments, we develop a dual self-supervision solution consisting of a soft self-supervision strategy with a Kullback-Leibler divergence loss and a hard self-supervision strategy with a pseudo supervision loss. Extensive experiments on nine benchmark datasets validate that our method consistently outperforms state-of-the-art methods. Especially, our method improves the ARI by more than 10.29% over the best baseline. The code will be publicly available athttps://github.com/ZhihaoPENG-CityU/DAGC.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
IEEE Trans. Circuits Syst. Video Technol.1
2023 EGRC-Net: Embedding-Induced Graph Refinement Clustering Network
abstract
Existing graph clustering networks heavily rely on a predefined yet fixed graph, which can lead to failures when the initial graph fails to accurately capture the data topology structure of the embedding space. In order to address this issue, we propose a novel clustering network called Embedding-Induced Graph Refinement Clustering Network (EGRC-Net), which effectively utilizes the learned embedding to adaptively refine the initial graph and enhance the clustering performance. To begin, we leverage both semantic and topological information by employing a vanilla auto-encoder and a graph convolution network, respectively, to learn a latent feature representation. Subsequently, we utilize the local geometric structure within the feature embedding space to construct an adjacency matrix for the graph. This adjacency matrix is dynamically fused with the initial one using our proposed fusion architecture. To train the network in an unsupervised manner, we minimize the Jeffreys divergence between multiple derived distributions. Additionally, we introduce an improved approximate personalized propagation of neural predictions to replace the standard graph convolution network, enabling EGRC-Net to scale effectively. Through extensive experiments conducted on nine widely-used benchmark datasets, we demonstrate that our proposed methods consistently outperform several state-of-the-art approaches. Notably, EGRC-Net achieves an improvement of more than 11.99% in Adjusted Rand Index (ARI) over the best baseline on the DBLP dataset. Furthermore, our scalable approach exhibits a 10.73% gain in ARI while reducing memory usage by 33.73% and decreasing running time by 19.71%. The code for EGRC-Net will be made publicly available at https://github.com/ZhihaoPENG-CityU/EGRC-Net.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
IEEE Trans. Image Process.1
2022 Maximum Entropy Subspace Clustering Network
abstract
Deep subspace clustering networks have attracted much attention in subspace clustering, in which an auto-encoder non-linearly maps the input data into a latent space, and a fully connected layer named self-expressiveness module is introduced to learn the affinity matrix via a typical regularization term (e.g., sparse or low-rank). However, the adopted regularization terms ignore the connectivity within each subspace, limiting their clustering performance. In addition, the adopted framework suffers from the coupling issue between the auto-encoder module and the self-expressiveness module, making the network training non-trivial. To tackle these two issues, we propose a novel deep subspace clustering method named Maximum Entropy Subspace Clustering Network (MESC-Net). Specifically, MESC-Net maximizes the entropy of the affinity matrix to promote the connectivity within each subspace, in which its elements corresponding to the same subspace are uniformly and densely distributed. Meanwhile, we design a novel framework to explicitly decouple the auto-encoder module and the self-expressiveness module. Besides, we also theoretically prove that the learned affinity matrix satisfies the block-diagonal property under the assumption of independent subspaces. Extensive quantitative and qualitative results on commonly used benchmark datasets validate MESC-Net significantly outperforms state-of-the-art methods. The code is publicly available athttps://github.com/ZhihaoPENG-CityU/MESC.
Zhihao Peng 0002, Yuheng Jia, Hui Liu 0032, Junhui Hou, Qingfu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Adaptive Attribute and Structure Subspace Clustering Network
abstract
Deep self-expressiveness-based subspace clustering methods have demonstrated effectiveness. However, existing works only consider the attribute information to conduct the self-expressiveness, limiting the clustering performance. In this paper, we propose a novel adaptive attribute and structure subspace clustering network (AASSC-Net) to simultaneously consider the attribute and structure information in an adaptive graph fusion manner. Specifically, we first exploit an auto-encoder to represent input data samples with latent features for the construction of an attribute matrix. We also construct a mixed signed and symmetric structure matrix to capture the local geometric structure underlying data samples. Then, we perform self-expressiveness on the constructed attribute and structure matrices to learn their affinity graphs separately. Finally, we design a novel attention-based fusion module to adaptively leverage these two affinity graphs to construct a more discriminative affinity graph. Extensive experimental results on commonly used benchmark datasets demonstrate that our AASSC-Net significantly outperforms state-of-the-art methods. In addition, we conduct comprehensive ablation studies to discuss the effectiveness of the designed modules. The code is publicly available at https://github.com/ZhihaoPENG-CityU/AASSC-Net.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
IEEE Trans. Image Process.1
2021 Attention-driven Graph Clustering Network
abstract
The combination of the traditional convolutional network (i.e., an auto-encoder) and the graph convolutional network has attracted much attention in clustering, in which the auto-encoder extracts the node attribute feature and the graph convolutional network captures the topological graph feature. However, the existing works (i) lack a flexible combination mechanism to adaptively fuse those two kinds of features for learning the discriminative representation and (ii) overlook the multi-scale information embedded at different layers for subsequent cluster assignment, leading to inferior clustering results. To this end, we propose a novel deep clustering method named Attention-driven Graph Clustering Network (AGCN). Specifically, AGCN exploits a heterogeneity-wise fusion module to dynamically fuse the node attribute feature and the topological graph feature. Moreover, AGCN develops a scale-wise fusion module to adaptively aggregate the multi-scale features embedded at different layers. Based on a unified optimization framework, AGCN can jointly perform feature learning and cluster assignment in an unsupervised fashion. Compared with the existing deep clustering methods, our method is more flexible and effective since it comprehensively considers the numerous and discriminative information embedded in the network and directly produces the clustering results. Extensive quantitative and qualitative results on commonly used benchmark datasets validate that our AGCN consistently outperforms state-of-the-art methods.
Zhihao Peng 0002, Hui Liu 0032, Yuheng Jia, Junhui Hou
ACM Multimedia1
2020 Non-Negative Transfer Learning With Consistent Inter-Domain Distribution
abstract
In this letter, we propose a novel transfer learning approach, which simultaneously exploits the intra-domain differentiation and inter-domain correlation to comprehensively solve the drawbacks many existing transfer learning methods suffer from, i.e., they either are unable to handle the negative samples or have strict assumptions on the distribution. Specifically, the sample selection strategy is introduced to handle negative samples by using the local geometry structure and the label information of source samples. Furthermore, the pseudo target label is imposed to slack the assumption on the inter-domain distribution for considering the inter-domain correlation. Then, an efficient alternating iterative algorithm is proposed to solve the formulated optimization problem with multiple constraints. The extensive experiments conducted on eleven real-world datasets show the superiority of our method over state-of-the-art approaches, i.e., our method achieves 11.23% improvement on the MNIST dataset.
Zhihao Peng 0002, Yuheng Jia, Junhui Hou
IEEE Signal Process. Lett.1
2020 Active Transfer Learning
abstract
A major assumption in data mining and machine learning is that the training set and test set come from the same domain. They share the same feature space and have the same distribution. However, in many real-world applications, the training set and test set usually come from different domains. Thus, there might be negative similarities between different domains so that the negative transfer problem caused by negative similarity may happen. In this paper, we propose a novel method named active transfer learning (ATL) to solve the above problem. Specifically, the orthogonal projection matrix and the weight coefficient vector are introduced to extend maximum mean discrepancy (MMD) so that it can minimize MMD and simultaneously eliminate the negative transfer. To find the informative and discriminative subsets from the source domain, we then propose an information diversity term by using the local geometric structure information of the source samples. Besides, by using the label information of source samples, our method can guarantee the selected subsets as discriminative as possible. Finally, to efficiently implement the proposed method, an alternating optimization approach, which is based on the alternating direction method of multipliers (ADMM), is designed to solve the optimization problem. To demonstrate the effectiveness of the proposed ATL model, experiments are conducted on five real-world data sets. The experimental results show the superiority of our method over the state-of-the-art methods.
Zhihao Peng 0002, Wei Zhang 0005, Na Han, Xiaozhao Fang, Peipei Kang, Luyao Teng
IEEE Trans. Circuits Syst. Video Technol.1