VLDB 2026 Research / reviewers in the wild / expert
Chao Zhang 0072
dblp:94/3019-72
· DBLP profile ↗
16ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-7705-8505ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entity-relation enhanced bidirectional information fusion for relational triples extraction
Haoran Liao, Yuan Rong, Mingming Kong, Chao Zhang 0072, Xianjun Tian, Vladimir Simic 0001, Dragan Pamucar |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Refining data granularity and feature fusion for boundary refinement in instance segmentation
Yumeng Yan, Mingming Kong, Maochao Zhang, Shunnan Zhao, Chao Zhang 0072 |
Signal Process. Image Commun. | 5 |
| 2025 | IdCo: Joint Identification and Contrastive Learning for Masked Face RecognitionabstractDeep face recognition models suffer significant accuracy declines when encountering masked faces, limiting their real-world applicability in the wild. This decline stems from misalignment between masked and regular faces in the feature space. However, improper adjustments to correct this misalignment would degrade regular face feature extracting, which is crucial for practical use. To obtain better performance in both masked face recognition (MFR) and regular face recognition (FR), we propose a joint Identification and Contrastive learning framework (IdCo), which combines identification learning with instance discrimination-based contrastive learning. In this framework, a modified loss function—Supervised Contrastive loss with Hard negative samples (SupConH)—is introduced in the contrastive learning module. This loss incorporates a hard negative sampling technique into the Supervised Contrastive loss (SupCon) to enhance learning effectiveness. Our framework promotes alignment between masked and regular faces without compromising discriminative power of learned features. To reduce the high training costs of contrastive learning, we adopt a stage-wise training strategy. In the first stage, the model is trained using identification learning alone, and IdCo is incorporated in the second stage. Extensive experiments on popular masked and regular face datasets show that IdCo outperforms state-of-the-art methods in both MFR and FR tasks. Qingtong Xu, Chao Zhang 0072, Ce Zhu |
MMSP | 2 |
| 2025 | FITE-GAT: Enhancing aspect-level sentiment classification with FT-RoBERTa induced trees and graph attention network
Mengmeng Fan, Mingming Kong, Xi Wang 0048, Fei Hao 0001, Chao Zhang 0072 |
Expert Syst. Appl. | 5 |
| 2025 | Fc-gcn: A formal concept-enhanced graph convolution network model
Chao Zhang 0072, Fei Hao 0001, Jinhai Li 0001, Qing Wan, Kyuwon Park, Xueyang Qin, Vincenzo Loia |
Soft Comput. | 2 |
| 2024 | Distance-reconstructed dependency enhanced aspect-based sentiment analysis with sentiment strength
Mingming Kong, Le Feng, Chao Zhang 0072, Fei Hao 0001, Yumeng Yan |
Neurocomputing | 3 |
| 2024 | Text-Vision Relationship Alignment for Referring Image SegmentationabstractAbstract Referring image segmentation aims to segment object in an image based on a referring expression. Its difficulty lies in aligning expression semantics with visual instances. The existing methods based on semantic reasoning are limited by the performance of external syntax parser and do not explicitly explore the relationships between visual instances. This article proposes an end-to-end method for referring image segmentation by aligning ’linguistic relationship’ with ’visual relationships’. This method does not rely on external syntax parser for expression parsing. In this paper, the expression is adaptively and structurally parsed into three components: ’subject’, ’object’, and ’linguistic relationship’ by the Semantic Component Parser (SCP) in a learnable manner. Instances Activation Map Module (IAM) locates multiple visual instances based on the subject and object. In addition, the Relationship Based Visual Localization Module (RBVL) firstly enables each instance of the image to learn global knowledge, then decodes the visual relationships between these visual instances, and finally aligns the visual relationships with the linguistic relationships to further accurately locate the target object. The experimental results show that the proposed method improves performance by 4– 9% compared with baseline method on multiple referring image segmentation datasets. Mingxing Pu, Bing Luo 0003, Chao Zhang 0072, Fayou Xu, Mingming Kong |
Neural Process. Lett. | 3 |
| 2024 | Blind Image Deblurring via Minimizing Similarity Between Fuzzy Sets on Image PixelsabstractMost existing image deblurring methods construct statistical prior to describe the difference between blur and clear image. They discard the position information and ignore pixel feature changing in deblurring, which results in inferior restoration performance for images unsatisfying corresponding assumptions. Intuitively, fuzziness of pixel belonging to different image regions will reduce along with image deblurring. This phenomenon could intrinsically describe the pixel characteristic. To this end, we analyze fuzziness of pixels and objects in a blurry image, and utilize the similarity between two fuzzy objects on image pixels to depict the blur degree of an image, which is inspired by overlap functions and overlap indices. To minimize the similarity between fuzzy objects, we introduce the non-parameters model to construct an integer programming problem. Energy minimization could significantly reduce the similarity between two fuzzy objects. Experimental results show that the proposed method can achieve better performance than the state-of-the-art blind deblurring methods on benchmark datasets and natural images. Junge Peng, Bing Luo 0003, Chao Zhang 0072, Zheng Pei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Blind image deblurring via content adaptive method
Zhongzhe Cheng, Bing Luo 0003, Bo Li 0054, Zheng Pei 0001, Chao Zhang 0072 |
Signal Process. Image Commun. | 6 |
| 2022 | Local-Global Interaction and Progressive Aggregation for Video Salient Object Detection
Dingyao Min, Chao Zhang 0072, Yukang Lu, Keren Fu, Qijun Zhao |
ICONIP (6) | 2 |
| 2022 | Mutual-Guidance Transformer-Embedding Network for Video Salient Object DetectionabstractVideo salient object detection (VSOD) aims at locating the most attractive objects presented in video sequences by exploiting spatial and temporal cues. Previous methods mainly utilize convolutional neural networks (CNNs) to fuse or complement across RGB and optical flow cues via simple strategies. To take full advantage of CNNs and recently emerged Transformers, this letter proposes a novel mutual-guidance Transformer-embedding network, called MGT-Net, where a mutual-guidance multi-head attention mechanism (MGMA) explores more sophisticated long-range cross-modal interactions. Such a mechanism is designed into a new mutual-guidance Transformer (MGTrans) module that can propagate long-range contextual dependencies based on information of the other modality. To the best of our knowledge, MGT-Net is the first VSOD model that embeds Transformers as modules into CNNs for improved performance. Prior to MGTrans, we also propose and deploy a feature purification module (FPM) to purify noisy backbone features. Experimental results on five benchmark datasets demonstrate the state-of-the-art performance of MGT-Net. Dingyao Min, Chao Zhang 0072, Yukang Lu, Keren Fu, Qijun Zhao |
IEEE Signal Process. Lett. | 2 |
| 2022 | MA-GANet: A Multi-Attention Generative Adversarial Network for Defocus Blur DetectionabstractBackground clutters pose challenges to defocus blur detection. Existing approaches often produce artifact predictions in background areas with clutter and relatively low confident predictions in boundary areas. In this work, we tackle the above issues from two perspectives. Firstly, inspired by the recent success of self-attention mechanism, we introduce channel-wise and spatial-wise attention modules to attentively aggregate features at different channels and spatial locations to obtain more discriminative features. Secondly, we propose a generative adversarial training strategy to suppress spurious and low reliable predictions. This is achieved by utilizing a discriminator to identify predicted defocus map from ground-truth ones. As such, the defocus network (generator) needs to produce ‘realistic’ defocus map to minimize discriminator loss. We further demonstrate that the generative adversarial training allows exploiting additional unlabeled data to improve performance, a.k.a. semi-supervised learning, and we provide the first benchmark on semi-supervised defocus detection. Finally, we demonstrate that the existing evaluation metrics for defocus detection generally fail to quantify the robustness with respect to thresholding. For a fair and practical evaluation, we introduce an effective yet efficient$AUF_\beta $metric. Extensive experiments on three public datasets verify the superiority of the proposed methods compared against state-of-the-art approaches. Xun Xu 0002, Le Zhang 0001, Chao Zhang 0072, Chuan-Sheng Foo, Ce Zhu |
IEEE Trans. Image Process. | 4 |
| 2020 | MultiANet: a Multi-Attention Network for Defocus Blur DetectionabstractDefocus blur detection is a challenging task because of obscure homogenous regions and interferences of background clutter. Most existing deep learning-based methods mainly focus on building wider or deeper network to capture multi-level features, neglecting to extract the feature relationships of intermediate layers, thus hindering the discriminative ability of network. Moreover, fusing features at different levels have been demonstrated to be effective. However, direct integrating without distinction is not optimal because low-level features focus on fine details only and could be distracted by background clutters. To address these issues, we propose the Multi-Attention Network for stronger discriminative learning and spatial guided low-level feature learning. Specifically, a channel-wise attention module is applied to both high-level and low-level feature maps to capture channel-wise global dependencies. In addition, a spatial attention module is employed to low-level features maps to emphasize effective detailed information. Experimental results show the performance of our network is superior to the state-of-the-art algorithms. Xun Xu 0002, Chao Zhang 0072, Ce Zhu |
MMSP | 3 |
| 2019 | C3AE: Exploring the Limits of Compact Model for Age EstimationabstractAge estimation is a classic learning problem in computer vision. Many larger and deeper CNNs have been proposed with promising performance, such as AlexNet, VggNet, GoogLeNet and ResNet. However, these models are not practical for the embedded/mobile devices. Recently, MobileNets and ShuffleNets have been proposed to reduce the number of parameters, yielding lightweight models. However, their representation has been weakened because of the adoption of depth-wise separable convolution. In this work, we investigate the limits of compact model for small-scale image and propose an extremely Compact yet efficient Cascade Context-based Age Estimation model(C3AE). This model possesses only 1/9 and 1/2000 parameters compared with MobileNets/ShuffleNets and VggNet, while achieves competitive performance. In particular, we re-define age estimation problem by two-points representation, which is implemented by a cascade model. Moreover, to fully utilize the facial context information, multi-branch CNN network is proposed to aggregate multi-scale context. Experiments are carried out on three age estimation datasets. The state-of-the-art performance on compact model has been achieved with a relatively large margin. Chao Zhang 0072, Shuaicheng Liu, Xun Xu 0002, Ce Zhu |
CVPR | 1 |
| 2018 | Image Ordinal Classification and Understanding: Grid Dropout with Masking LabelabstractImage ordinal classification refers to predicting a discrete target value which carries ordering correlation among image categories. The limited size of labeled ordinal data renders modern deep learning approaches easy to overfit. To tackle this issue, neuron dropout and data augmentation were proposed which, however, still suffer from over-parameterization and breaking spatial structure, respectively. To address the issues, we first propose a grid dropout method that randomly dropout/blackout some areas of the training image. Then we combine the objective of predicting the blackout patches with classification to take advantage of the spatial information. Finally we demonstrate the effectiveness of both approaches by visualizing the Class Activation Map (CAM) and discover that grid dropout is more aware of the whole facial areas and more robust than neuron dropout for small training dataset. Experiments are conducted on a challenging age estimation dataset-Adience dataset with very competitive results compared with state-of-the-art methods. Chao Zhang 0072, Ce Zhu, Jimin Xiao, Xun Xu 0002, Yipeng Liu 0001 |
ICME | 1 |
| 2018 | Visual aesthetic understanding: Sample-specific aesthetic classification and deep activation map visualization
Chao Zhang 0072, Ce Zhu, Xun Xu 0002, Yipeng Liu 0001, Jimin Xiao, Tammam Tillo |
Signal Process. Image Commun. | 1 |