VLDB 2026 Research / reviewers in the wild / expert
Guo Zhong
dblp:258/9048
· DBLP profile ↗
54ranked-venue papers
17as first author
49since 2021 · last 2026
0000-0002-6428-5645ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 9 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 14 since 2021Databases, data management, data science and information retrieval · 10 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-view subspace clustering with adaptive weighted reconstruction loss
Guo Zhong, Yalin Wang 0012, Mingdong Tang, Xueming Yan, Jianghao Lin |
Int. J. Approx. Reason. | 1 |
| 2026 | MADAT: Missing-aware dynamic adaptive transformer model for medical prognosis prediction with incomplete multimodal data
Jianbin He, Guoheng Huang, Xiaochen Yuan, Chi-Man Pun, Guo Zhong, Bai Ying Lei, Haojiang Li |
Medical Image Anal. | 5 |
| 2026 | Error-Resilient incomplete multi-View clustering: Mitigating imputation-induced error accumulation
Xuanlong Ma, Fenfang Xie, Guo Zhong |
Pattern Recognit. | 4 |
| 2025 | Lorentz Transformation Neural NetworkabstractWe propose a novel neural network architecture, the Lorentz Transformation Neural Network (LTNN), which utilizes Lorentz transformations to generate a complex computation matrix that enhances the network’s expressive power. Furthermore, LTNN is lightweight due to the shared weight matrices in the computation matrix. LTNN treats the input and output as coordinates in high-dimensional spacetime, with the weight matrices in each layer representing the velocity components of a spacetime reference frame. During training, these weight matrices are transformed into a computation matrix via Lorentz transformations, describing the coordinate transformations between different reference frames. We evaluate LTNN on four datasets: California Housing Prices, Iris, MNIST, and Fashion-MNIST. Experimental results demonstrate that LTNN outperforms conventional neural networks and quaternion neural networks in terms of both accuracy and parameter efficiency. Wenyuan Li 0007, Jingchao Wang 0002, Guoheng Huang, Tongxu Lin, Guo Zhong, Xiaochen Yuan, Chi-Man Pun, An Zeng |
ICIP | 5 |
| 2025 | Superpixel-Enhanced Quaternion Feature Fusion and Contextualization Graph Contrastive Learning for Cervical Cancer Diagnosis
Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Lianglun Cheng, Chi-Man Pun, Guo Zhong, Qingjian Ye |
ICONIP (2) | 8 |
| 2025 | Adaptive Query Prompting for Multi-Domain Landmark DetectionabstractMedical landmark detection is crucial in various medical imaging modalities and procedures. Although deep learning-based methods have achieve promising performance, they are mostly designed for specific anatomical regions or tasks. In this work, we propose a universal model for multi-domain landmark detection by leveraging transformer architecture and developing a prompting component, named as Adaptive Query Prompting (AQP). Transformer backbone architecture is suitable for our study due to its ability to capture long-range dependencies crucial in multi-domain landmark detection. Specifically, transformers excel at understanding global anatomical relationships, which span across entire images. Instead of embedding additional modules in the backbone network, we design a separate module to generate prompts that can be effectively extended to any other transformer network. In our proposed AQP, prompts are learnable parameters maintained in a memory space called prompt pool. The central idea is to keep the backbone frozen and then optimize prompts to instruct the model inference process. Furthermore, we employ a lightweight decoder to decode landmarks from the extracted features, namely Light-MLD. Thanks to the lightweight nature of the decoder and AQP, we can handle multiple datasets by sharing the backbone encoder and then only perform partial parameter tuning without incurring much additional cost. It has the potential to be extended to more landmark detection tasks. We conduct experiments on three widely used X-ray datasets for different medical landmark detection tasks. Our proposed Light-MLD coupled with AQP achieves SOTA performance on many metrics even without the use of elaborate structural designs or complex frameworks. Qiusen Wei, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Jianwen Huang |
IJCNN | 5 |
| 2025 | Cross-View Geo-Localization via Learning Correspondence Semantic Similarity Knowledge
Guanli Chen, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
MMM (1) | 5 |
| 2025 | Locality adaptive incomplete multi-view subspace clustering
Guo Zhong, Shengqi Wu, Yuzhi Liang, Xiuyun Zhu |
Data Min. Knowl. Discov. | 1 |
| 2025 | A plug-and-play data-driven approach for anti-money laundering in bitcoin
Yuzhi Liang, Weijing Wu, Ruiju Liang, Kai Lei, Guo Zhong, Qingqing Gan, Jinsheng Huang |
Expert Syst. Appl. | 6 |
| 2025 | The Structure-sharing Hypergraph Reasoning Attention Module for CNNs
Jingchao Wang 0002, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Tongxu Lin, Chi-Man Pun, Fenfang Xie |
Expert Syst. Appl. | 4 |
| 2025 | Visual-linguistic Diagnostic Semantic Enhancement for medical report generation
Jiahong Chen, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zhe Tan, Chi-Man Pun |
J. Biomed. Informatics | 4 |
| 2025 | Implicit Supervision-Assisted Graph Collaborative Filtering for Third-Party Library RecommendationabstractThird-party libraries (TPLs) play a crucial role in software development. Utilizing TPL recommender systems can aid software developers in promptly finding useful TPLs. A number of TPL recommendation approaches have been proposed and among them graph neural network (GNN)-based recommendation is attracting the most attention. However, GNN-based approaches generate node representations through multiple convolutional aggregations, which is prone to introducing noise, resulting in the over-smoothing issue. In addition, due to the high sparsity of labelled data, node representations may be biased in real-world scenarios. To address these issues, this paper presents a TPL recommendation method named Implicit Supervision-assisted Graph Collaborative Filtering (ISGCF). Specifically, it takes the App-TPL interaction relationships as input and employs a popularity-debiased method to generate denoised App and TPL graphs. This reduces the noise introduced during graph convolution and alleviates the over-smoothing issue. It also employs a novel implicitly-supervised loss function to exploit the labelled data to learn enhanced node representations. Extensive experiments on a large-scale real-world dataset demonstrate that ISGCF achieves a significant performance advantage over other state-of-the-art TPL recommendation methods in Recall, NDCG and MAP. The experiments also validate the superiority of ISGCF in mitigating the over-smoothing problem. Lianrong Chen, Mingdong Tang, Naidan Mei, Fenfang Xie, Guo Zhong, Qiang He 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2025 | Psanet: prototype-guided salient attention for few-shot segmentation
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
Vis. Comput. | 6 |
| 2025 | Weakly supervised semantic segmentation via saliency perception with uncertainty-guided noise suppression
Guoheng Huang, Xiaochen Yuan, Zewen Zheng, Guo Zhong, Xuhang Chen 0002, Chi-Man Pun |
Vis. Comput. | 5 |
| 2024 | IMAN: An Adaptive Network for Robust NPC Mortality Prediction with Missing ModalitiesabstractAccurate prediction of mortality in nasopharyngeal carcinoma (NPC), a complex malignancy particularly challenging in advanced stages, is crucial for optimizing treatment strategies and improving patient outcomes. However, this predictive process is often compromised by the high-dimensional and heterogeneous nature of NPC-related data, coupled with the pervasive issue of incomplete multi-modal data, manifesting as missing radiological images or incomplete diagnostic reports. Traditional machine learning approaches suffer significant performance degradation when faced with such incomplete data, as they fail to effectively handle the high-dimensionality and intricate correlations across modalities. Even advanced multi-modal learning techniques like Transformers struggle to maintain robust performance in the presence of missing modalities, as they lack specialized mechanisms to adaptively integrate and align the diverse data types, while also capturing nuanced patterns and contextual relationships within the complex NPC data. To address these problem, we introduce IMAN: an adaptive network for robust NPC mortality prediction with missing modalities. IMAN features three integrated modules: the Dynamic Cross-Modal Calibration (DCMC) module employs adaptive, learnable parameters to scale and align medical images and field data; the Spatial-Contextual Attention Integration (SCAI) module enhances traditional Transformers by incorporating positional information within the self-attention mechanism, improving multi-modal feature integration; and the Context-Aware Feature Acquisition (CAFA) module adjusts convolution kernel positions through learnable offsets, allowing for adaptive feature capture across various scales and orientations in medical image modalities. Extensive experiments on our proprietary NPC dataset demonstrate IMAN’s robustness and high predictive accuracy, even with missing data. Compared to existing methods, IMAN consistently outperforms in scenarios with incomplete data, representing a significant advancement in mortality prediction for medical diagnostics and treatment planning. Our code is available at https://github.com/king-huoye/BIBM-2024/tree/master. Yejing Huo, Guoheng Huang, Lianglun Cheng, Jianbin He, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun |
BIBM | 7 |
| 2024 | FAQNet: Frequency-Aware Quaternion Network for Endoscopic Highlight RemovalabstractDue to the built-in light source within the endoscope, the illumination of bodily mucous can cause the formation of highlight regions due to reflection. This not only interferes with the diagnosis conducted by doctors but also poses a challenge to subsequent computer vision tasks. To tackle this issue, we introduce FAQNet, a network specifically designed for endoscopic image highlight removal. FAQNet seamlessly integrates multi-channel information leveraging quaternion convolution and spatial channel attention within our Quaternion Multi-Channel Fusion (QMCF) Module. This allows it to capture intricate details of color, texture, spatial information, and highlight characteristics within the imaged organ. Additionally, by employing frequency domain transformation and dilated convolution, the Contextual Information Integration (CII) Module effectively enlarges the receptive field, organizing contextual information between highlight regions and their surrounding areas. Lastly, the PixelShuffle Upsampling (PSU) Module generates the restored image. We validate our model’s performance on two benchmark datasets, demonstrating its superiority over existing highlight removal methodologies. Dingzhou Zhu, Guoheng Huang, Xiaochen Yuan, Xuhang Chen 0002, Guo Zhong, Chi-Man Pun |
BIBM | 5 |
| 2024 | FOPS-V: Feature-Aware Optimization and Parallel Scale Fusion for 3D Human Reconstruction in Video
Guoheng Huang, Lianglun Cheng, Yejing Huo, Xuhang Chen 0002, Xiaochen Yuan, Guo Zhong, Chi-Man Pun |
ICONIP (8) | 7 |
| 2024 | ROSAL: Semi-supervised Active Learning with Representation Aggregation and Outlier for Endoscopy Image Classification
Xiaocong Huang, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Xuhang Chen 0002, Chi-Man Pun, Jianwu Chen |
ICONIP (11) | 3 |
| 2024 | Medical Visual Prompting (MVP): A Unified Framework for Versatile and High-Quality Medical Image SegmentationabstractAccurate segmentation of lesion regions is crucial for clinical diagnosis and treatment across various diseases. While deep convolutional networks have achieved satisfactory results in medical image segmentation, they face challenges such as the loss of lesion shape information due to continuous convolution and downsampling, as well as the high cost of manually labeling lesions with varying shapes and sizes. To address these issues, we propose a novel Medical Visual Prompting (MVP) framework that leverages pre-training and prompting concepts from Natural Language Processing (NLP). The framework utilizes three key components: Super-Pixel Guided Prompting (SPGP) for superpixelating the input image, Image Embedding Guided Prompting (IEGP) for freezing patch embedding and merging with superpixels to provide visual prompts, and Adaptive Attention Mechanism Guided Prompting (AAGP) for pinpointing prompt content and efficiently adapting all layers. By integrating SPGP, IEGP, and AAGP, the MVP framework enables the segmentation network to better learn shape prompting information and facilitates mutual learning across different tasks. Extensive experiments conducted on five datasets demonstrate the superior performance of the proposed method in various challenging medical image tasks while simplifying single-task medical segmentation models. This novel framework offers improved performance with fewer parameters and holds significant potential for accurate segmentation of lesion regions in various medical tasks, making it clinically valuable. Guoheng Huang, Zijin Lin, Guo Zhong, Shenghong Luo |
SMC | 5 |
| 2024 | Cross-Modality Disentangled Information Bottleneck Strategy for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) has been a pivotal domain in current research area which utilizes diverse information carriers such as videos containing multiple modal-ities to understand the user's sentiment. With the success of multimodal fusion techniques, lots of fusion strategies have been proposed to obtain a favorable multimodal joint representation for MSA. However, existing studies hardly consider the problem of redundant information in unimodal, resulting in the joint representation may contain much redundant information from different modalities, thus limiting the accuracy of sentiment prediction. In this work, we propose a Cross-Modality Disentangled Information Bottleneck Strategy (CMDIBS), which consists of a Cross-Modality Knowledge Awareness (CMKA) module and a Multimodal Disentangled Information Bottleneck (MDIB) mechanism. Specifically, the CMKA module encourages in-teractions among different modalities to learn the sentiment embedding relevant to the predicted goals. In particular, MDIB mechanism aims to maximize the mutual information (MI) between the multimodal joint representation and the predicted label, and maximize the MI between the style embedding with the label and the input data while constraining the MI between the multimodal joint representation and the style embedding to obtain a succinct and efficient multimodal joint representation. Experimental results on the benchmark datasets, namely CMU-MOSI and CMU-MOSEI, indicated that the proposed method surpasses existing approaches and attains SOTA performance. Zhengnan Deng, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Lian Huang, Chi-Man Pun |
SMC | 3 |
| 2024 | IAMS-Net: An Illumination-Adaptive Multi-Scale Lesion Segmentation NetworkabstractIn recent years, many Lesion segmentation (LS) models based on UNet have been proposed. However, existing researches rarely consider the influence of illumination change leads to the weak boundary area. Such as melanomas and polyps, the demarcation of the boundary between the diseased area and the surrounding tissue remains particularly challenging. To overcome these challenges, we propose an IlluminationAdaptive Multi-scale Lesion Segmentation Network (IAMS-Net). In IAMS-Net, we integrate Illumination-Adaptive MultiStream Attention (IAMA) and Contour Perception Module (CPM). In the decoding stage, the IAMA is used as a bridge between the encoder and the decoder to solve the adverse effects of illumination changes on the segmentation of weak boundary lesions. In order to further enhance the boundary features lost due to illumination change in the low-contrast lesion area, we introduce the CPM to improve the perception of the integrity of the lesion area. Subsequently, we performed comparison and ablation experiments using the publicly available ISIC2018 dataset and the individually collected data set BoreIllumination(BI). Yisen Zheng, Guoheng Huang, Lianglun Cheng, Xiaochen Yuan, Guo Zhong, Shenghong Luo |
SMC | 6 |
| 2024 | Binary spectral clustering for multi-view data
Xueming Yan, Guo Zhong, Yaochu Jin, Xiaohua Ke, Fenfang Xie, Guoheng Huang |
Inf. Sci. | 2 |
| 2024 | Multi-task subspace clustering
Guo Zhong, Chi-Man Pun |
Inf. Sci. | 1 |
| 2024 | Black-box reversible adversarial examples with invertible neural network
Jielun Huang, Guoheng Huang, Xiaochen Yuan, Fenfang Xie, Chi-Man Pun, Guo Zhong |
Image Vis. Comput. | 7 |
| 2024 | Cross-domain visual prompting with spatial proximity knowledge distillation for histological image classification
Guoheng Huang, Lianglun Cheng, Guo Zhong, Weihuang Liu, Xuhang Chen 0002, Muyan Cai |
J. Biomed. Informatics | 4 |
| 2024 | Progressive normalizing flow with learnable spectrum transform for style transfer
Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Yiwen Zeng |
Knowl. Based Syst. | 4 |
| 2024 | Nonnegative Tensor Representation With Cross-View Consensus for Incomplete Multi-View ClusteringabstractTensors capture the multi-dimensional structure of multi-view data naturally, resulting in richer and more meaningful data representations. This produces more accurate clustering results for challenging incomplete multi-view clustering (IMVC) tasks. However, previous tensor learning-based IMVC (TLIMVC) methods often build a tensor representation by simply stacking view-specific representations. Consequently, the learned tensor representation lacks good interpretability since each entry of it could not directly reveals the similarity relationship of the corresponding two samples. In addition, most of them only focus on exploring the high-order correlations among views, while the underlying consensus information is not fully exploited. To this end, we propose a novel TLIMVC method named Nonnegative Tensor Representation with Cross-view Consensus (NTRC$^{2}$) in this paper. Specifically, a nonnegative constraint and view-specific consensus are jointly integrated into the framework of the tensor based self-representation learning, which enables the method to simultaneously explore the consensus and complementary information of multi-view data more fully. An Augmented Lagrangian Multiplier based optimization algorithm is derived to optimize the objective function. Experiments on several challenging benchmark datasets verify our NTRC$^{2}$method's effectiveness and competitiveness against state-of-the-art methods. Guo Zhong, Juanchun Wu, Xueming Yan, Xuanlong Ma |
IEEE Signal Process. Lett. | 1 |
| 2024 | GDN-CMCF: A Gated Disentangled Network With Cross-Modality Consensus Fusion for Multimodal Named Entity RecognitionabstractMultimodal named entity recognition (MNER) is a crucial task in social systems of artificial intelligence that requires precise identification of named entities in sentences using both visual and textual information. Previous methods have focused on capturing fine-grained visual features and developing complex fusion procedures. However, these approaches overlook the heterogeneity gap and loss of original modality uniqueness that may occur during fusion, leading to incorrect entity identification. This article proposes a novel approach for MNER called a gated disentangled network with cross-modality consensus fusion (GDN-CMCF) to address the above challenges. Specifically, to eliminate cross-modality variation, we propose a cross-modality consensus fusion module that generates a consensus representation by learning inter-and intramodality interactions with a designed commonality constraint. We then introduce a gated disentanglement module to separate modality-relevant features from support and auxiliary modalities, which further filters out extraneous information while retaining the uniqueness of unimodal features. Experimental results on two real public datasets are provided to verify the effectiveness of our proposed GDN-CMCF. The source code of this article can be found at https://github.com/HaoDavis/ GDN-CMCF. Guoheng Huang, Zihao Dai, Guo Zhong, Xiaochen Yuan, Chi-Man Pun |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | Quaternion Cross-Modality Spatial Learning for Multi-Modal Medical Image SegmentationabstractRecently, the Deep Neural Networks (DNNs) have had a large impact on imaging process including medical image segmentation, and the real-valued convolution of DNN has been extensively utilized in multi-modal medical image segmentation to accurately segment lesions via learning data information. However, the weighted summation operation in such convolution limits the ability to maintain spatial dependence that is crucial for identifying different lesion distributions. In this paper, we propose a novel Quaternion Cross-modality Spatial Learning (Q-CSL) which explores the spatial information while considering the linkage between multi-modal images. Specifically, we introduce to quaternion to represent data and coordinates that contain spatial information. Additionally, we propose Quaternion Spatial-association Convolution to learn the spatial information. Subsequently, the proposed De-level Quaternion Cross-modality Fusion (De-QCF) module excavates inner space features and fuses cross-modality spatial dependency. Our experimental results demonstrate that our approach compared to the competitive methods perform well with only 0.01061 M parameters and 9.95G FLOPs. Junyang Chen 0001, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zewen Zheng, Chi-Man Pun, Jian Zhu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Learning From Incorrectness: Active Learning With Negative Pre-Training and Curriculum Querying for Histological Tissue ClassificationabstractPatch-level histological tissue classification is an effective pre-processing method for histological slide analysis. However, the classification of tissue with deep learning requires expensive annotation costs. To alleviate the limitations of annotation budgets, the application of active learning (AL) to histological tissue classification is a promising solution. Nevertheless, there is a large imbalance in performance between categories during application, and the tissue corresponding to the categories with relatively insufficient performance are equally important for cancer diagnosis. In this paper, we propose an active learning framework called ICAL, which contains Incorrectness Negative Pre-training (INP) and Category-wise Curriculum Querying (CCQ) to address the above problem from the perspective of category-to-category and from the perspective of categories themselves, respectively. In particular, INP incorporates the unique mechanism of active learning to treat the incorrect prediction results that obtained from CCQ as complementary labels for negative pre-training, in order to better distinguish similar categories during the training process. CCQ adjusts the query weights based on the learning status on each category by the model trained by INP, and utilizes uncertainty to evaluate and compensate for query bias caused by inadequate category performance. Experimental results on two histological tissue classification datasets demonstrate that ICAL achieves performance approaching that of fully supervised learning with less than 16% of the labeled data. In comparison to the state-of-the-art active learning algorithms, ICAL achieved better and more balanced performance in all categories and maintained robustness with extremely low annotation budgets. The source code will be released at https://github.com/LactorHwt/ICAL. Lianglun Cheng, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Chi-Man Pun, Muyan Cai |
IEEE Trans. Medical Imaging | 5 |
| 2024 | SCDet: decoupling discriminative representation for dark object detection via supervised contrastive learning
Tongxu Lin, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Xiaocong Huang, Chi-Man Pun |
Vis. Comput. | 4 |
| 2023 | Locality Preserving Multiview Graph Hashing For Large Scale Remote Sensing Image SearchabstractHashing is very popular for remote sensing image search. This article proposes a multiview hashing with learnable parameters to retrieve the queried images for a large-scale remote sensing dataset. Existing methods always neglect that real-world remote sensing data lies on a low- dimensional manifold embedded in high-dimensional ambient space. Unlike previous methods, this article proposes to learn the consensus compact codes in a view-specific low-dimensional subspace. Furthermore, we have added a hyperparameter learnable module to avoid complex parameter tuning. In order to prove the effectiveness of our method, we carried out experiments on three widely used remote sensing data sets and compared them with seven state-of-the-art methods. Extensive experiments show that the proposed method can achieve competitive results compared to the other method. Wenyun Li 0001, Guo Zhong, Chi-Man Pun |
ICASSP | 2 |
| 2023 | Data Representation by Joint Hypergraph Embedding and Sparse Coding (Extended Abstract)abstractMatrix factorization (MF), a popular unsupervised learning technique for data representation, has been widely applied in data mining and machine learning. According to different application scenarios, one can impose different constraints on the factorization to find the desired basis, which captures high-level semantics for the given data, and learns the compact representation corresponding to the basis. We note that almost all previous work on MF in data mining has ignored to find such a basis, which can carry high-order semantics in the data. In this work, we propose a novel MF framework called Joint Hypergraph Embedding and Sparse Coding, in which the obtained basis captures high-order semantic information in data. Experimental results on data clustering demonstrate that the proposed method consistently outperforms the other state-of-the-art matrix factorization methods. Guo Zhong, Chi-Man Pun |
ICDE | 1 |
| 2023 | High-Order Collaborative Filtering for Third-Party Library RecommendationabstractDevelopers of mobile applications (apps) can enhance their work efficiency by reusing suitable third-party libraries (TPLs). TPLs recommendation methods have been proposed to assist app developers in quickly finding useful TPLs, but the existing methods, such as those based on graph neural networks (GNN), have limitations in extracting high-order neighborhood information that can enhance the representation ability of nodes. To address this issue, we propose a novel hypergraph neural network method based on collaborative filtering, called High-order Collaborative Filtering (HCF). We first build two hypergraphs by fully exploiting the TPLs usage records in apps and then extracting the high-order neighborhood information from the hypergraphs. The neighborhood information extracted contains less noise compared to classic GNNs, thereby alleviating the problem of over-smoothing in GNN node representations. This advantage enables HCF to recommend more accurate and diverse TPLs for apps development. Extensive experiments on a real-world dataset demonstrate that HCF significantly outperforms the state-of-the-art methods in terms of recommendation accuracy and diversity. Lianrong Chen, Naidan Mei, Yingying He, Wanping Liu, Guo Zhong, Mingdong Tang |
ICWS | 5 |
| 2023 | A Neural Inference of User Social Interest for Item RecommendationabstractAbstract User-generated content is daily produced in social media, as such user interest summarization is critical to distill salient information from massive information for recommendation tasks. While the interested messages (e.g., tags or posts) from a single user are usually sparse becoming a bottleneck for existing methods, we propose a neural inference method (NIGraphNet) by mining user social interest for item recommendation. It can unearth user latent topics combined with user relation learning. Specifically, we exploit a neural variational inference approach to learn the distributions between user interests and hidden topics. (We denote it as interest-topic distributions in the following.) Then, we adopt a unified graph-based training loss that jointly learns the hidden topics and user relations for item recommendation. Experiments on two datasets collected from well-known social media platforms demonstrate the superior performance of our model in the tasks of user interest summarization and item recommendation. Further discussions also show that exploiting the latent topic representations and user relations is conducive to the user’s automatic language understanding. Junyang Chen 0001, Mengzhu Wang, Ge Fan, Guo Zhong, Ou Liu, Wenfeng Du, Zhenghua Xu 0001, Zhiguo Gong |
Data Sci. Eng. | 5 |
| 2023 | A new focused crawler using an improved tabu search algorithm incorporating ontology and host informationabstractTo solve the problems of incomplete topic description and repetitive crawling of visited hyperlinks in traditional focused crawling methods, in this paper, we propose a novel focused crawler using an improved tabu search algorithm with domain ontology and host information (FCITS_OH), where a domain ontology is constructed by formal concept analysis to describe topics at the semantic and knowledge levels. To avoid crawling visited hyperlinks and expand the search range, we present an improved tabu search (ITS) algorithm and the strategy of host information memory. In addition, a comprehensive priority evaluation method based on Web text and link structure is designed to improve the assessment of topic relevance for unvisited hyperlinks. Experimental results on both tourism and rainstorm disaster domains show that the proposed focused crawlers overmatch the traditional focused crawlers for different performance metrics. Jingfa Liu, Guo Zhong, Zhihe Yang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2023 | Simultaneous Laplacian embedding and subspace clustering for incomplete multi-view data
Guo Zhong, Chi-Man Pun |
Knowl. Based Syst. | 1 |
| 2023 | Self-taught Multi-view Spectral Clustering
Guo Zhong, Chi-Man Pun |
Pattern Recognit. | 1 |
| 2023 | RBA-GCN: Relational Bilevel Aggregation Graph Convolutional Network for Emotion RecognitionabstractEmotion recognition in conversation (ERC) has received increasing attention from researchers due to its wide range of applications. As conversation has a natural graph structure, numerous approaches used to model ERC based on graph convolutional networks (GCNs) have yielded significant results. However, the aggregation approach of traditional GCNs suffers from the node information redundancy problem, leading to node discriminant information loss. Additionally, single-layer GCNs lack the capacity to capture long-range contextual information from the graph. Furthermore, the majority of approaches are based on textual modality or stitching together different modalities, resulting in a weak ability to capture interactions between modalities. To address these problems, we present the relational bilevel aggregation graph convolutional network (RBA-GCN), which consists of three modules: the graph generation module (GGM), similarity-based cluster building module (SCBM) and bilevel aggregation module (BiAM). First, GGM constructs a novel graph to reduce the redundancy of target node information. Then, SCBM calculates the node similarity in the target node and its structural neighborhood, where noisy information with low similarity is filtered out to preserve the discriminant information of the node. Meanwhile, BiAM is a novel aggregation method that can preserve the information of nodes during the aggregation process. This module can construct the interaction between different modalities and capture long-range contextual information based on similarity clusters. On both the IEMOCAP and MELD datasets, the weighted average F1 score of RBA-GCN has a 2.17$\sim$5.21% improvement over that of the most advanced method. Guoheng Huang, Fenghuan Li, Xiaochen Yuan, Chi-Man Pun, Guo Zhong |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | TempEE: Temporal-Spatial Parallel Transformer for Radar Echo Extrapolation Beyond AutoregressionabstractMeteorological radar reflectivity data (i.e. radar echo) significantly influences precipitation prediction. It can facilitate accurate and expeditious forecasting of short-term heavy rainfall bypassing the need for complex Numerical Weather Prediction (NWP) models. In comparison to conventional models, Deep Learning (DL)-based radar echo extrapolation algorithms exhibit higher effectiveness and efficiency. Nevertheless, the development of reliable and generalized echo extrapolation algorithm is impeded by three primary challenges: cumulative error spreading, imprecise representation of sparsely distributed echoes, and inaccurate description of non-stationary motion processes. To tackle these challenges, this paper proposes a novel radar echo extrapolation algorithm called Temporal-Spatial Parallel Transformer, referred to asTempEE.TempEEavoids using auto-regression and instead employs a one-step forward strategy to prevent cumulative error spreading during the extrapolation process. Additionally, we propose the incorporation of a Multi-level Temporal-Spatial Attention mechanism to improve the algorithm’s capability of capturing both global and local information while emphasizing task-related regions, including sparse echo representations, in an efficient manner. Furthermore, the algorithm extracts spatio-temporal representations from continuous echo images using a parallel encoder to model the non-stationary motion process for echo extrapolation. The superiority of ourTempEEhas been demonstrated in the context of the classic radar echo extrapolation task, utilizing a real-world dataset. Extensive experiments have further validated the efficacy and indispensability of various components withinTempEE. Shengchao Chen, Ting Shu 0001, Huan Zhao 0004, Guo Zhong, Xunlai Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | QGD-Net: A Lightweight Model Utilizing Pixels of Affinity in Feature Layer for Dermoscopic Lesion SegmentationabstractRESPONSE: Pixels with location affinity, which can be also called "pixels of affinity," have similar semantic information. Group convolution and dilated convolution can utilize them to improve the capability of the model. However, for group convolution, it does not utilize pixels of affinity between layers. For dilated convolution, after multiple convolutions with the same dilated rate, the pixels utilized within each layer do not possess location affinity with each other. To solve the problem of group convolution, our proposed quaternion group convolution uses the quaternion convolution, which promotes the communication between to promote utilizing pixels of affinity between channels. In quaternion group convolution, the feature layers are divided into 4 layers per group, ensuring the quaternion convolution can be performed. To solve the problem of dilated convolution, we propose the quaternion sawtooth wave-like dilated convolutions module (QS module). QS module utilizes quaternion convolution with sawtooth wave-like dilated rates to effectively leverage the pixels that share the location affinity both between and within layers. This allows for an expanded receptive field, ultimately enhancing the performance of the model. In particular, we perform our quaternion group convolution in QS module to design the quaternion group dilated neutral network (QGD-Net). Extensive experiments on Dermoscopic Lesion Segmentation based on ISIC 2016 and ISIC 2017 indicate that our method has significantly reduced the model parameters and highly promoted the precision of the model in Dermoscopic Lesion Segmentation. And our method also shows generalizability in retinal vessel segmentation. Jingchao Wang 0002, Guoheng Huang, Guo Zhong, Xiaochen Yuan, Chi-Man Pun |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Simultaneous multi-graph learning and clustering for multiview data
Xuanlong Ma, Xueming Yan, Jingfa Liu, Guo Zhong |
Inf. Sci. | 4 |
| 2022 | A novel focused crawler combining Web space evolution and domain ontology
Jingfa Liu, Xin Li 0159, Qiansheng Zhang, Guo Zhong |
Knowl. Based Syst. | 4 |
| 2022 | Local Learning-based Multi-task Clustering
Guo Zhong, Chi-Man Pun |
Knowl. Based Syst. | 1 |
| 2022 | Multi-view spectral clustering by simultaneous consensus graph learning and discretization
Guo Zhong, Ting Shu 0001, Guoheng Huang, Xueming Yan |
Knowl. Based Syst. | 1 |
| 2022 | Improved Normalized Cut for Multi-View ClusteringabstractSpectral clustering (SC) algorithms have been successful in discovering meaningful patterns since they can group arbitrarily shaped data structures. Traditional SC approaches typically consist of two sequential stages, i.e., performing spectral decomposition of an affinity matrix and then rounding the relaxed continuous clustering result into a binary indicator matrix. However, such a two-stage process could make the obtained binary indicator matrix severely deviate from the ground true one. This is because the former step is not devoted to achieving an optimal clustering result. To alleviate this issue, this paper presents a general joint framework to simultaneously learn the optimal continuous and binary indicator matrices for multi-view clustering, which also has the ability to tackle the conventional single-view case. Specially, we provide theoretical proof for the proposed method. Furthermore, an effective alternate updating algorithm is developed to optimize the corresponding complex objective. A number of empirical results on different benchmark datasets demonstrate that the proposed method outperforms several state-of-the-arts in terms of six clustering metrics. Guo Zhong, Chi-Man Pun |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Data Representation by Joint Hypergraph Embedding and Sparse CodingabstractMatrix factorization (MF), a popular unsupervised learning technique for data representation, has been widely applied in data mining and machine learning. According to different application scenarios, one can impose different constraints on the factorization to find the desired basis, which captures high-level semantics for the given data, and learns the compact representation corresponding to the basis. We note that almost all previous work on MF in data mining has ignored to find such a basis, which can carry high-order semantics in the data. In this article, we propose a novel MF framework called Joint Hypergraph Embedding and Sparse Coding (JHESC), in which the obtained basis captures high-order semantic information in data. Specifically, we first propose a new hypergraph learning model to obtain a more discriminative basis by hypergraph-based Laplacian Eigenmap, then sparse coding is conducted on the learned basis such that the new representation has stronger identification capability. In addition, we extend the proposed method to the reproducing kernel Hilbert space for dealing with nonlinear data more effectively. Extensive experimental results on data clustering demonstrate that the proposed method consistently outperforms the other state-of-the-art matrix factorization methods. Guo Zhong, Chi-Man Pun |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Latent Low-rank Graph Learning for Multimodal ClusteringabstractMultimodal clustering has become a fundamental and important problem in the data mining community since the development of multimedia technology over the last two decades has led to a tremendous increase in unlabeled multimodal data. Although a panoply of multimodal subspace clustering methods shows promising performance via fusing information from different views of multimodal data, most of them consist of two sequential steps, i.e., learning a consensus affinity matrix from the original data and then feeding the resulting affinity matrix into the framework of spectral clustering. However, this leads to the suboptimal clustering performance due to the following limitations: 1) the two steps of learning the affinity matrix and clustering are carried out independently; 2) the affinity matrix may be unreliable; 3) the post-processing requirement, such as K-means. To address these issues, we propose a novel multimodal subspace clustering method via adaptively learning a similarity graph on a latent low-rank representation space. In particular, the number of connected components of the learned graph is precisely equal to the number of clusters, i.e., the optimal solution of the associated problem directly reveals the clustering structure of data. Extensive evaluations on several benchmark multimodal datasets demonstrate that the proposed approach outperforms state-of-the-art methods. Guo Zhong, Chi-Man Pun |
ICDE | 1 |
| 2021 | RPCA-induced self-representation for subspace clustering
Guo Zhong, Chi-Man Pun |
Neurocomputing | 1 |
| 2020 | A Unified Framework for Multi-view Spectral ClusteringabstractIn the era of big data, multi-view clustering has drawn considerable attention in machine learning and data mining communities due to the existence of a large number of unlabeled multi-view data in reality. Traditional spectral graph theoretic methods have recently been extended to multi-view clustering and shown outstanding performance. However, most of them still consist of two separate stages: learning a fixed common real matrix (i.e., continuous labels) of all the views from original data, and then applying K-means to the resulting common label matrix to obtain the final clustering results. To address these, we design a unified multi-view spectral clustering scheme to learn the discrete cluster indicator matrix in one stage. Specifically, the proposed framework directly obtain clustering results without performing K-means clustering. Experimental results on several famous benchmark datasets verify the effectiveness and superiority of the proposed method compared to the state-of-the-arts. Guo Zhong, Chi-Man Pun |
ICDE | 1 |
| 2020 | Revisiting Nyström extension for hypergraph clustering
Guo Zhong, Chi-Man Pun |
Neurocomputing | 1 |
| 2020 | Nonnegative self-representation with a fixed rank constraint for subspace clustering
Guo Zhong, Chi-Man Pun |
Inf. Sci. | 1 |
| 2020 | Simultaneously learning feature-wise weights and local structures for multi-view subspace clustering
Guo Zhong, Ting Shu 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Subspace clustering by simultaneously feature selection and similarity learning
Guo Zhong, Chi-Man Pun |
Knowl. Based Syst. | 1 |