VLDB 2026 Research / reviewers in the wild / expert
Guangzhi Ma
dblp:08/3298
· DBLP profile ↗
24ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0001-5726-1672ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Out-of-Distribution Detection in Transfer Learning Through Intuitionistic Fuzzy Set-Based PredictionabstractOut-of-distribution(OOD) detection involves training a model on in-distribution (ID) samples to determine whether a given input belongs to classes that were unknown during training (i.e., OOD data). While recent advances in vision-language models like CLIP have improved OOD detection, effectively handling the uncertainty inherent in OOD scenarios remains a major challenge, primarily due to the unknown nature, diversity, and absence of OOD samples during training. To address this, we propose a novel method that leverages intuitionistic fuzzy sets (IFS) to explicitly model uncertainty through membership, non-membership, and hesitation degrees. Specifically, we use a pre-trained CLIP-based model to learn positive and negative prompts, which are then used to construct IFS. We further introduce a new loss function to guide the model in building appropriate IFS from ID data, and propose a hesitation-based scoring function for OOD detection during inference. Extensive experiments across standard benchmarks and evaluation metrics show that our approach outperforms state-of-the-art (SOTA) methods, demonstrating the superiority of applying fuzzy logic in addressing the inherent uncertainties of OOD detection. Guangzhi Ma, Jie Lu 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | MO2Tracker: Polyp Tracking by Multi-objective Optimization ApproachabstractMulti-Object Tracking (MOT) of polyps in colonoscopy videos serves as a foundational component for intelligent medical workflows. In tracking-by-detection algorithms, managing the equilibrium among competing optimization objectives during data association presents a complex issue. Current approaches, including threshold-based weighted linear aggregation methodologies and hierarchical cascade matching schemes, demonstrate persistent limitations when deployed in endoscopic polyp monitoring scenarios characterized by vigorous polyp movements and complex morphological deformations. Thus, this study proposes a novel polyp tracking framework named MO2Tracker, which can simultaneously optimize multiple objectives during the data association phase. Specifically, MO2Tracker integrates three discriminative similarity cues in parallel: polyp coordinates, detection confidence scores, and appearance consistency, to establish associations between new detections and existing trajectories. To resolve this optimization challenge, we implement Pareto Front analysis via a multi-objective optimization algorithm, deriving Pareto Front solutions. Subsequently, we design two novel selection paradigms, knee-oriented selection and human-inspired selection, to filter optimal detector-trajectory pairings from the Pareto Front. Extensive experiments conducted on two public datasets and one private dataset validate the superior performance of MO2Tracker. On the public dataset SUN-SEG, MO2Tracker achieves improvements of +4.1% in Multi-object Tracking Accuracy (MOTA), +7.2% in Higher Order Tracking Accuracy(HOTA), and +5.8% in IDF1 Score. Guangzhi Ma, Dongming Yu, Wanyu Qiu, Xianyuan Wang, Enmin Song |
SMC | 2 |
| 2025 | Image Grayscale Enhancement through Frame Accumulation with Shaped-Function Signal
Enmin Song, Guangzhi Ma, Wanyu Qiu |
SMC | 3 |
| 2025 | Multiview Classification Through Learning From Interval-Valued DataabstractThe classification problem concerning crisp-valued data has been well resolved. However, interval-valued data, where all of the observations' features are described by intervals, are also a common data type in real-world scenarios. For example, the data extracted by many measuring devices are not exact numbers but intervals. In this article, we focus on a highly challenging problem called learning from interval-valued data (LIND), where we aim to learn a classifier with high performance on interval-valued observations. First, we obtain the estimation error bound of the LIND problem based on the Rademacher complexity. Then, we give the theoretical analysis to show the strengths of multiview learning on classification problems, which inspires us to construct a new algorithm called multiview interval information extraction (Mv-IIE) approach for improving classification accuracy on interval-valued data. The experiment comparisons with several baselines on both synthetic and real-world datasets illustrate the superiority of the proposed framework in handling interval-valued data. Moreover, we describe an application of Mv-IIE that we can prevent data privacy leakage by transforming crisp-valued (raw) data into interval-valued data. Guangzhi Ma, Jie Lu 0001, Zhen Fang 0001, Feng Liu 0003, Guangquan Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Spatio-temporal scale information fusion of Functional Near-Infrared Spectroscopy signal for depression detection
Jitao Zhong, Guangzhi Ma, Lu Zhang 0071, Quanhong Wang, Shi Qiao 0006, Hong Peng 0003, Bin Hu 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Quaternion Deformable Local Binary Pattern and Pose-Correction Facial Decomposition for Color Facial Expression Recognition in the WildabstractFacial expression recognition (FER) in the wild is a more challenging topic than that under laboratory-controlled conditions. The major obstacles of FER in the wild are head pose variations, illumination changes, and different skin colors. To address these problems, we propose a framework named quaternion deformable local binary pattern (QDLBP)-Net for color FER in the wild. First, to eliminate the interferences of head pose variations, a pose-correction facial decomposition (PCFD) strategy is proposed to correct the head pose and decompose the facial image into five emotion-related regions. Then, to handle the problems of illumination changes and different skin colors, an effective feature descriptor named “QDLBP” is developed. QDLBP extracts color quaternion features from each emotional region, which not only computes the strength of emotional features, but also maintains the spectral correlation between color channels. Finally, a quaternion classification network (QC-Net) is proposed to classify the quaternion features from five emotional regions into seven basic expressions. The experimental results on three in-the-wild FER datasets and two nonfrontal pose variation datasets exhibit the effectiveness and superiority of QDLBP-Net by showing clear performance improvements over other state-of-the-art (SOTA) FER methods. Yu Zhou 0049, Guangzhi Ma, Enmin Song |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Multiclass Classification With Fuzzy-Feature Observations: Theory and AlgorithmsabstractThe theoretical analysis of multiclass classification has proved that the existing multiclass classification methods can train a classifier with high classification accuracy on the test set, when the instances are precise in the training and test sets with same distribution and enough instances can be collected in the training set. However, one limitation with multiclass classification has not been solved: how to improve the classification accuracy of multiclass classification problems when only imprecise observations are available. Hence, in this article, we propose a novel framework to address a new realistic problem called multiclass classification with imprecise observations (MCIMO), where we need to train a classifier with fuzzy-feature observations. First, we give the theoretical analysis of the MCIMO problem based on fuzzy Rademacher complexity. Then, two practical algorithms based on support vector machine and neural networks are constructed to solve the proposed new problem. The experiments on both synthetic and real-world datasets verify the rationality of our theoretical analysis and the efficacy of the proposed algorithms. Guangzhi Ma, Jie Lu 0001, Feng Liu 0003, Zhen Fang 0001, Guangquan Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2024 | Fuzzy Machine Learning: A Comprehensive Framework and Systematic ReviewabstractMachine learning draws its power from various disciplines, including computer science, cognitive science, and statistics. Although machine learning has achieved great advancements in both theory and practice, its methods have some limitations when dealing with complex situations and highly uncertain environments. Insufficient data, imprecise observations, and ambiguous information/relationships can all confound traditional machine learning systems. To address these problems, researchers have integrate machine leaning from different aspects, and fuzzy techniques including fuzzy sets, fuzzy systems, fuzzy logic, fuzzy measures, fuzzy relations, and so on. This paper presents a systematic review of fuzzy machine learning, from theory, approach to application, with the overall objective of providing an overview of recent achievements in the field of fuzzy machine learning. To this end, the concepts and frameworks discussed are divided into five categories: (a) fuzzy classical machine learning; (b) fuzzy transfer learning; (c) fuzzy data stream learning; (d) fuzzy reinforcement learning; and (e) fuzzy recommender systems. The literature presented should provide researchers with a solid understanding of the current progress in fuzzy machine learning research and its applications. Jie Lu 0001, Guangzhi Ma, Guangquan Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | Domain Adaptation With Interval-Valued Observations: Theory and AlgorithmsabstractUnsupervised Domain Adaptation (UDA) focuses on enhancing the model performance on an unlabeled target domain by leveraging knowledge from a source domain. The source and target domains usually share different distributions. Existing UDA research primarily concentrates on image data characterized by crisp-valued features. However, interval-valued data, where all the observations’ features are described by intervals, is also a common type of data in real-world scenarios. For instance, measurement instruments are unable to provide exact numerical outcomes, instead employing intervals to describe their results. Hence, this paper focuses on the highly challenging context known as domain adaptation with interval-valued observations. In this environment, the objective is to improve classification accuracy within an unlabeled target domain by capitalizing on knowledge gleaned from a labeled source domain, where both domains exclusively feature interval-valued observations. To address this, we first establish an upper bound on the risk in the interval-valued target domain, underpinning our analysis with rigorous theoretical insights. Subsequently, guided by our theoretical analysis, a new model based on Takagi-Sugeno Fuzzy rules and a Self-supervised Pseudo-labeling strategy (SP-TSF) is developed to address the proposed problem. Takagi-Sugeno fuzzy rules are harnessed to handle the inherent uncertainty intrinsic to interval-valued data, while a pseudo-labeling strategy is developed to augment distribution alignment between the source and target domains, each characterized by interval-valued observations. Extensive experiments on both synthetic and realworld datasets verify the rationality of our theoretical analysis and the efficacy of the proposed model. Guangzhi Ma, Jie Lu 0001, Feng Liu 0003, Zhen Fang 0001, Guangquan Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Multisource Domain Adaptation With Interval-Valued Target Data via Fuzzy Neural NetworksabstractMulti-source domain adaptation (MSDA) refers to the task of adapting a model from multiple source domains to a target domain that shares a different distribution with all source domains. However, most existing MSDA works focus on crispvalued data, while such data may not be available in some realworld scenarios. For example, data extracted by many measuring devices are not exact numbers but rather intervals. In this paper, a highly challenging problem called MSDA with interval-valued target data is presented. The objective is to learn a new model for interval-valued target data by leveraging knowledge from source models trained on multiple crisp-valued source data. First, a theoretical analysis is given to inform the appropriate combination of multi-source models. Then, we propose a new neural network model based on a fuzzy transformation function and fuzzy distances to address the proposed problem. The fuzzy transformation function is applied to extract valuable crisp-valued information from interval-valued target data, while fuzzy distances are designed to guide the fusion of multiple source models. Experiments on both synthetic and real-world datasets verify the superiority of our proposed MSDA method for classification task. Furthermore, the results of the ablation study and parameter sensitivity analysis illustrate the rationality of the proposed fuzzy distance-based model. Guangzhi Ma, Jie Lu 0001, Guangquan Zhang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2023 | Compact Selective Transformer Based on Information Entropy for Facial Expression Recognition in the WildabstractFacial expression recognition (FER) in the wild is a challenging task due to pose variations, occlusions, etc. Many studies employ region-based methods to relieve the influence of occlusions and pose variations. However, these methods often neglect the global relationship between local regions. To address these problems, we introduce a compact selective transformer into ResNet-50 (R-CST) for in-the-wild FER. First, we develop a compact transformer to capture the global relationship between local regions outputted by the intermediate of ResNet-50. Then, an information entropy-based selective module is added to the compact transformer to select discriminative information and drop the background and occlusions. Finally, we combine the intermediate features and the last convolutional features of R-CST for emotion classification. Experimental results on three in-the-wild FER datasets demonstrate that the proposed R-CST outperforms several state-of-the-art FER models.Codes are available at https://github.com/Gabrella/R-CST. Liyuan Guo, Guangzhi Ma |
ICIP | 3 |
| 2022 | APRNet: A 3D Anisotropic Pyramidal Reversible Network With Multi-Modal Cross-Dimension Attention for Brain Tissue Segmentation in MR ImagesabstractBrain tissue segmentation in multi-modal magnetic resonance (MR) images is significant for the clinical diagnosis of brain diseases. Due to blurred boundaries, low contrast, and intricate anatomical relationships between brain tissue regions, automatic brain tissue segmentation without prior knowledge is still challenging. This paper presents a novel 3D fully convolutional network (FCN) for brain tissue segmentation, called APRNet. In this network, we first propose a 3D anisotropic pyramidal convolutional reversible residual sequence (3DAPC-RRS) module to integrate the intra-slice information with the inter-slice information without significant memory consumption; secondly, we design a multi-modal cross-dimension attention (MCDA) module to automatically capture the effective information in each dimension of multi-modal images; then, we apply 3DAPC-RRS modules and MCDA modules to a 3D FCN with multiple encoded streams and one decoded stream for constituting the overall architecture of APRNet. We evaluated APRNet on two benchmark challenges, namely MRBrainS13 and iSeg-2017. The experimental results show that APRNet yields state-of-the-art segmentation results on both benchmark challenge datasets and achieves the best segmentation performance on the cerebrospinal fluid region. Compared with other methods, our proposed approach exploits the complementary information of different modalities to segment brain tissue regions in both adult and infant MR images, and it achieves the average Dice coefficient of 87.22% and 93.03% on the MRBrainS13 and iSeg-2017 testing data, respectively. The proposed method is beneficial for quantitative brain analysis in the clinical study, and our code is made publicly available. Yuzhou Zhuang, Hong Liu 0005, Enmin Song, Guangzhi Ma, Chih-Cheng Hung |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Learning from Imprecise Observations: An Estimation Error Bound based on Fuzzy Random VariablesabstractIn the problem of multi-class classification, researchers have proved that we can train a classifier that has good performance on the test set, as long as the training and test sets are precisely drawn from the same distribution and the size of the training set approaches infinity. However, in a realworld situation, such precise observations are often unavailable in some cases. For example, readings on analogue measurement equipment are not precise numbers but intervals since there is only a finite number of decimals available. Hence, in this paper, we propose a more realistic problem called learning from imprecise observations (LIMO), where we train a classifier with fuzzy observations (i.e., fuzzy vectors). We prove the estimation error bound of this novel problem based on the distribution of fuzzy random variables. This bound demonstrates that we can always learn the best classifier when we have infinite fuzzy observations. We also develop a practical algorithm to train a classifier using fuzzy observations. The experiment results verify the efficacy of our theory and algorithm. Guangzhi Ma, Feng Liu 0003, Guangquan Zhang 0001, Jie Lu 0001 |
FUZZ-IEEE | 1 |
| 2021 | Dual link distributed source coding scheme for the transmission of satellite hyperspectral imagery
Ahmed Hagag, Ibrahim Omara, Souleyman Chaib, Guangzhi Ma, Fathi E. Abd El-Samie |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | A novel approach for ear recognition: learning Mahalanobis distance features from deep CNNs
Ibrahim Omara, Ahmed Hagag, Guangzhi Ma, Fathi E. Abd El-Samie, Enmin Song |
Mach. Vis. Appl. | 3 |
| 2020 | LDM-DAGSVM: Learning Distance Metric via DAG Support Vector Machine for Ear Recognition ProblemabstractRecently, the ear recognition system takes more increasingly interesting for many applications, especially, in immigration system, forensic, and surveillance applications. For face re-identification and image classification, metric learning has significantly improved machine learning accuracies by using K-Nearest Neighbor (KNN) and Support Vector Machine (SVM) classifiers. However, metric learning via SVM has not yet been investigated for the ear recognition problem. To achieve better generalization ability than the traditional previous classifiers, a novel framework for ear recognition is proposed based on learning distance metric (LDM) via SVM since the LDM and the directed acyclic graph SVM (DAGSVM) are two emerging techniques which perform outstanding in dealing with classification problems. This work considers metric learning for SVM by proposing a hybrid learning distance metric and directed acyclic graph SVM (LDM-DAGSVM) model for ear recognition system. Different from existing ear biometric methods, the proposed approach aims to learn a Mahalanobis distance metric via SVM to maximize the inter-class variations and minimize the intra-class variations, simultaneously. The experiments are conducted on complicated ear datasets and the results can achieve better performance compared with the state-of-the-art ear recognition methods. The proposed approach can get classification accuracy up to 98.79%, 98.70%, and 84.30% for AWE, AME and WPUT ear datasets, respectively. Ibrahim Omara, Guangzhi Ma, Enmin Song |
IJCB | 2 |
| 2020 | Cascaded hybrid residual U-Net for glioma segmentation
Jiaosong Long, Guangzhi Ma, Hong Liu 0005, Enmin Song, Chih-Cheng Hung, Renchao Jin, Yuzhou Zhuang, DaiYang Liu |
Multim. Tools Appl. | 2 |
| 2020 | A Two-Stage Convolutional Neural Networks for Lung Nodule DetectionabstractEarly detection of lung cancer is an effective way to improve the survival rate of patients. It is a critical step to have accurate detection of lung nodules in computed tomography (CT) images for the diagnosis of lung cancer. However, due to the heterogeneity of the lung nodules and the complexity of the surrounding environment, it is a challenge to develop a robust nodule detection method. In this study, we propose a two-stage convolutional neural networks (TSCNN) for lung nodule detection. The first stage based on the improved U-Net segmentation network is to establish an initial detection of lung nodules. During this stage, in order to obtain a high recall rate without introducing excessive false positive nodules, we propose a new sampling strategy for training. Simultaneously, a two-phase prediction method is also proposed in this stage. The second stage in the TSCNN architecture based on the proposed dual pooling structure is built into three 3D-CNN classification networks for false positive reduction. Since the network training requires a significant amount of training data, we designed a random mask as the data augmentation method in this study. Furthermore, we have improved the generalization ability of the false positive reduction model by means of ensemble learning. We verified the proposed architecture on the LUNA dataset in our experiments, which showed that the proposed TSCNN architecture did obtain competitive detection performance. Haichao Cao, Hong Liu 0005, Enmin Song, Guangzhi Ma, Renchao Jin, Tengying Liu, Chih-Cheng Hung |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Learning deep CNNs for impulse noise removal in images
Guangzhi Ma, Enmin Song |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Tunable discounting and visual exploration for language models
Junfei Guo, Qi Han 0006, Guangzhi Ma, Hong Liu 0005, Seth van Hooland |
Neurocomputing | 3 |
| 2015 | A novel method for fusion of differently exposed images based on spatial distribution of intensity for ubiquitous multimedia
Mali Yu, Enmin Song, Renchao Jin, Hong Liu 0005, Guangzhi Ma |
Multim. Tools Appl. | 6 |
| 2013 | Pattern classification of dermoscopy images: A perceptually uniform model
Qaisar Abbas, M. Emre Celebi 0001, Carmen Serrano, Irene Fondón, Guangzhi Ma |
Pattern Recognit. | 5 |
| 2012 | Multiple costs based decision making with back-propagation neural networks
Guangzhi Ma, Enmin Song, Chih-Cheng Hung, Dongshan Huang |
Decis. Support Syst. | 1 |
| 2011 | Semi-supervised multi-class Adaboost by exploiting unlabeled data
Enmin Song, Dongshan Huang, Guangzhi Ma, Chih-Cheng Hung |
Expert Syst. Appl. | 3 |