EDBT 2026 Demo / reviewers in the wild / expert
Fangming Zhong
dblp:172/2719
· DBLP profile ↗
27ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0002-2775-8175ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HaNa: Hardness and Noise-Aware Robust Cross-modal RetrievalabstractNoisy correspondence in cross-modal retrieval introduces significant challenges due to its inherent difficulty in identification and correction. Although existing methods attempt to minimize the influence of noisy samples by the weighting mechanism, these methods still struggle with performance degradation under increasing noise levels. Specifically, the clean samples are assigned the same weight of 1, which ignores the sample hardness. In addition, the weights for noisy samples are approaching 0, leading to the overlook of sample diversity. To address these issues, we propose a Hardness and Noise-aware (HaNa) robust cross-modal retrieval method. HaNa introduces a momentum-based reweighting mechanism to adaptively balance learning difficulty across clean samples, avoiding overfitting risk and accumulative partitioning bias. Moreover, HaNa addresses the limitation that weights for noisy data are approaching 0 from a new perspective to fully employ the diversity of samples to further improve its generalization. It employs an Asymmetric Noise-aware Regularization Loss (ANRL) to treat identified noisy data as negative samples for optimization. Extensive experiments demonstrate that HaNa achieves superior matching accuracy and stability, especially in high-noise scenarios, outperforming state-of-the-art methods. Fangming Zhong, Haiquan Yu, Cun Zhu, Suhua Zhang |
AAAI | 1 |
| 2026 | Cross-domain few-shot compact multi-modal feature fusion for hyperspectral images classification with supervised contrastive learning
Suhua Zhang, Zhikui Chen, Huicen Guo, Fangming Zhong |
Expert Syst. Appl. | 4 |
| 2026 | GARE-Net: Geometric contextual aggregation and regional contextual enhancement network for image-text matching
Fangming Zhong, Zhikui Chen, Suhua Zhang |
Expert Syst. Appl. | 1 |
| 2026 | Semi-supervised time series classification via sequence neural process
Zhikui Chen, Fangming Zhong |
Knowl. Inf. Syst. | 3 |
| 2026 | An infrared and visible image fusion architecture based on foreground-aware salient object detection network
Zhikui Chen, Fangming Zhong |
Multim. Syst. | 3 |
| 2026 | HGACH: hypergraph attention convolutional hashing for semi-supervised cross-modal retrieval
Fangming Zhong, Cun Zhu, Haiquan Yu, Chenglong Chu, Suhua Zhang |
Multim. Syst. | 1 |
| 2025 | Retention Enhanced Cross-modal Attention for Multi-Hop VQAabstractExploring multimodal information from external knowledge bases in Visual Question Answering (VQA) reasoning tasks presents a significant challenge. Current methods, such as attention-based and graph-based approaches, have limitations in effectively capturing contextual key information. It is essential to enhance the model’s cognitive understanding of the associations between the question and the corresponding external knowledge to improve accuracy. To overcome these challenges, this paper introduces RECA, a Retention Enhanced Cross-modal Attention method for multi-hop VQA inspired by the hypergraph. Specifically, the contextual information after cross-modal attention fusion are further enhanced by the proposed retention module based on Retentive Network. Finally, we introduced a mimicry loss through model distillation. This enabled our model to enhance collaborative learning and improve generalisation performance. Our method is evaluated on KVQA, FVQA, PQ, and PQL datasets, demonstrating state-of-the-art performance. Chenglong Chu, Fangming Zhong |
ICASSP | 4 |
| 2025 | Cross-domain multimodal feature enhancement hypergraph neural network for few-shot hyperspectral images classification
Suhua Zhang, Zhikui Chen, Fangming Zhong |
Expert Syst. Appl. | 3 |
| 2025 | GPO++: A dynamic pooling routing network based on pooling operator for cross-modal retrieval
Fangming Zhong, Chuanyu Bing, Suhua Zhang |
Expert Syst. Appl. | 1 |
| 2024 | Dual-Mix for Cross-Modal Retrieval with Noisy LabelsabstractCross-modal retrieval with deep neural networks heavily relies on accurate annotation. However, existing methods may easily suffer from the scarcity and validity of annotations due to the expensive cost of manual labeling. In addition, it is inevitable that noisy labels are imposed during labeling. To this end, it is worthwhile to explore the potential of noisy labels in cross-modal retrieval. In this work, we propose a novel framework entitled Dual-Mix for Cross-Modal Retrieval with noisy labels (DMCM). It consists of two components, which are mixing the robust loss functions and mixing augmentation for noisy samples. In the first mixing stage, the normalized generalized cross entropy and mean absolute error are combined to boost each other. Then, after separating clean and noisy samples by Beta Mixture Model, we mix these samples via augmentation to further address the scarcity of labeled samples. Extensive experiments demonstrate the significant superiority of our DMCM. Fangming Zhong |
ICASSP | 4 |
| 2024 | Context-Aware and Contrastiveness-Driven Feature Learning for Cross-Domain Few-Shot Hyperspectral Image ClassificationabstractFew-shot learning has attracted considerable attention in the field of hyperspectral image (HSI) classification due to its suitability in addressing the challenges encountered in numerous real-world scenarios. However, the scarcity of labeled samples poses a significant challenge in learning informative and discriminative features, limiting the potential for achieving higher accuracy. In this paper, we propose a contextual information aggregation module (CIAM) as part of the feature extraction network for few-shot hyperspectral image classification which can aggregate more spatial-spectral information for each pixel from the neighbored pixels. Meanwhile, supervised contrastive learning is introduced to learn more discriminative representations for addressing specific challenges of high inter-class similarity and large intra-class variance in hyperspectral images. Extensive experiments on two benchmark datasets show that our proposed method achieves the state-of-the-art results. Suhua Zhang, Fangming Zhong, Zhikui Chen |
ICASSP | 2 |
| 2024 | APDF: An active preference-based deep forest expert system for overall survival prediction in gastric cancer
Qiucen Li, Zedong Du, Weihan Zhang, Fangming Zhong, Z. Jane Wang 0001, Zhikui Chen |
Expert Syst. Appl. | 6 |
| 2023 | Hypergraph-Enhanced Hashing for Unsupervised Cross-Modal Retrieval via Robust Similarity GuidanceabstractUnsupervised cross-modal hashing retrieval across image and text modality is a challenging task because of the suboptimality of similarity guidance, i.e., the joint similarity matrix constructed by existing methods does not possess clear enough guiding significance. How to construct more robust similarity matrix is the key to solve this problem. The unsupervised cross-modal retrieval methods based on graph have a good performance in mining semantic information of input samples, but the graph hashing based on traditional affinity graph cannot capture the high-order semantic information of input samples effectively. In order to overcome the aforementioned limitations, this paper presents a novel hypergraph-based approach for unsupervised cross-modal retrieval that differs from previous works in two significant ways. Firstly, to address the ubiquitous redundant information present in current methods, this paper introduces a robust similarity matrix constructing method. Secondly, we propose a novel hypergraph enhanced module that produces embedding vectors by hypergraph convolution and attention mechanism for input data, capturing important high-order semantics. Our approach is evaluated on the NUS-WIDE and MIRFlickr datasets, and yields state-of-the-art performance for unsupervised cross-modal retrieval. Fangming Zhong, Chenglong Chu, Zhikui Chen |
ACM Multimedia | 1 |
| 2022 | Noise Suppression for Improved Few-Shot LearningabstractFew-shot learning (FSL) aims to generalize from few labeled samples. Recently, metric-based methods have achieved surprising classification performance on many FSL benchmarks. However, those methods ignore the impact of noise, making the few-shot learning still tricky. In this work, we identify that noise suppression is important to improve the performance of FSL algorithms. Hence, we proposed a novel attention-based contrastive learning model with discrete cosine transform input (ACL-DCT), which can suppress the noise in input images, image labels, and learned features, respectively. ACL-DCT takes the transformed frequency domain representations by DCT as input and removes the high-frequency part to suppress the input noise. Besides, an attention-based alignment of the feature maps and a supervised contrastive loss are used to mitigate the feature and label noise. We evaluate our ACL-DCT by comparing previous methods on two widely used datasets for few-shot classification (i.e., miniImageNet and CUB). The results indicate that our proposed method outperforms the state-of-the-art methods. Zhikui Chen, Tiandong Ji, Suhua Zhang, Fangming Zhong |
ICASSP | 4 |
| 2022 | Dual-Attention Network for Few-Shot SegmentationabstractFew-shot segmentation aims at segmenting target object areas with only a few labeled samples. Previous methods extract class-specific prototypes to guide segmentation. How-ever, using one or more prototypes to represent the whole object inevitably drops vital spatial information, ignoring many details in original images. To address the issue, we propose a Dual-Attention Network (DANet) for few-shot segmentation. Firstly, a light-dense attention module is proposed to set up pixel-wise relations between feature pairs at different levels to activate object regions, which can leverage semantic information in a coarse-to-fine manner. Secondly, in contrast to the previous prototype-based methods that offer a holistic representation for each object class, we propose a prototypical channel attention module which incorporates channel interdependencies to enhance the discriminative capacity of features. The extensive experiments on two benchmarks show that our approach outperforms the state-of-the-arts in most cases. Zhikui Chen, Suhua Zhang, Fangming Zhong |
ICASSP | 4 |
| 2021 | Multiple-Input Multiple-Output Fusion Network for Generalized Zero-Shot LearningabstractGeneralized zero-shot learning (GZSL) has attracted considerable attention recently, which trains models with data from seen classes and tests on data from both seen and unseen classes. Most of the existing methods attempt to find a mapping from visual space to semantic space, such mapping can easily result in the domain shift problem. To address this issue, we propose a Multiple-Input Multiple-Output Fusion Network to GZSL. It can generate similar common semantic representation to paired inputs even with only the class semantic embeddings. This makes it possible to synthesize pseudo samples from attributes of unseen classes. Extensive experiments carried out on three benchmark datasets show the effectiveness of the proposed model. Fangming Zhong, Guangze Wang, Zhikui Chen, Xu Yuan 0002, Feng Xia 0001 |
ICASSP | 1 |
| 2021 | CHOP: An orthogonal hashing method for zero-shot cross-modal retrieval
Xu Yuan 0002, Guangze Wang, Zhikui Chen, Fangming Zhong |
Pattern Recognit. Lett. | 4 |
| 2021 | A Sparse Deep Transfer Learning Model and Its Application for Smart AgricultureabstractThe introduction of deep transfer learning (DTL) further reduces the requirement of data and expert knowledge in various uses of applications, helping DNN‐based models effectively reuse information. However, it often transfers all parameters from the source network that might be useful to the task. The redundant trainable parameters restrict DTL in low‐computing‐power devices and edge computing, while small effective networks with fewer parameters have difficulty transferring knowledge due to structural differences in design. For the challenge of how to transfer a simplified model from a complex network, in this paper, an algorithm is proposed to realize a sparse DTL, which only transfers and retains the most necessary structure to reduce the parameters of the final model. Sparse transfer hypothesis is introduced, in which a compressing strategy is designed to construct deep sparse networks that distill useful information in the auxiliary domain, improving the transfer efficiency. The proposed method is evaluated on representative datasets and applied for smart agriculture to train deep identification models that can effectively detect new pests using few data samples. Zhikui Chen, Fangming Zhong |
Wirel. Commun. Mob. Comput. | 4 |
| 2020 | HDMFH: Hypergraph Based Discrete Matrix Factorization Hashing for Multimodal RetrievalabstractIn recent years, hashing based cross-modal retrieval methods have attracted considerable attention for the high retrieval efficiency and low storage cost. However, most of the existing methods neglect the high-order relationship among data samples. In addition, most of them can only deal with two modalities, e.g., image and text, without discussing the scenario of multiple modalities. To address these issues, in this paper, we propose a novel cross-modal hashing method, named Hypergraph Based Discrete Matrix Factorization Hashing (HDMFH), for multimodal retrieval. Different from most previous approaches, our method based on hypergraph regularization and matrix factorization can handle the cross-modal retrieval of more than two modalities, which is known as multimodal retrieval. Extensive experiments demonstrate that HDMFH outperforms the state-of-the-art cross-modal hashing methods. Jing Gao 0007, Zhikui Chen, Fangming Zhong |
ICASSP | 4 |
| 2020 | Semantic Augmentation Hashing for Zero-Shot Image RetrievalabstractHashing technique has been widely applied to large-scale image retrieval due to its efficacy in storage and retrieval. However, due to the explosive growth of multimedia data on the web, existing hashing approaches can hardly achieve satisfactory performance on the newly-emerging images of new classes. In this paper, we propose a novel Semantic Augmentation Hashing (SAH) for zero-shot image retrieval. The class semantic embeddings are used as an intermediate space between visual features and binary codes to align visual features to corresponding class semantics and to transfer knowledge from seen classes to unseen classes simultaneously. Extensive experiments conducted on two datasets with different scales demonstrate the superiority of our method as compared against the state-of-the-arts. Fangming Zhong, Zhikui Chen, Geyong Min, Feng Xia 0001 |
ICASSP | 1 |
| 2020 | UCMH: Unpaired cross-modal hashing with matrix factorization
Jing Gao 0007, Fangming Zhong, Zhikui Chen |
Neurocomputing | 3 |
| 2020 | A novel strategy to balance the results of cross-modal hashing
Fangming Zhong, Zhikui Chen, Geyong Min, Feng Xia 0001 |
Pattern Recognit. | 1 |
| 2019 | An Exploration of Cross-Modal Retrieval for Unseen Concepts
Fangming Zhong, Zhikui Chen, Geyong Min |
DASFAA (2) | 1 |
| 2019 | STCMH with minimal semantic lossabstractCross‐modal hashing (CMH) has received widespread attention due to high retrieval efficiency, which plays an extremely important role in cross‐modal retrieval. Recently, many CMH methods have been proposed to establish the semantic connection of different modalities. However, most of these methods only use a simple quantisation strategy, resulting in large quantisation error, and inferior hash codes. To address this issue, in this study, the authors propose a novel self‐taught CMH (STCMH) to minimise the semantic encoding loss. In particular, the common semantic representations across different modalities are first learnt based on collective matrix factorisation. Then, the quantisation procedure based on orthogonal transformation is integrated to encode the semantic representations into discriminative binary codes. Moreover, similarity preservation is imposed to further boost the discriminative power. Finally, hashing functions learning is formulated as a binary classification problem by self‐taught scheme. Experimental results on three public datasets demonstrate that STCMH significantly outperforms most state‐of‐the‐art CMH methods. Jianing Du, Zhikui Chen, Fangming Zhong, Xiru Qiu |
IET Image Process. | 3 |
| 2018 | Combinative hypergraph learning in subspace for cross-modal ranking
Fangming Zhong, Zhikui Chen, Geyong Min, Zhaolong Ning, Hua Zhong 0006, Yueming Hu 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Deep Discrete Cross-Modal Hashing for Cross-Media Retrieval
Fangming Zhong, Zhikui Chen, Geyong Min |
Pattern Recognit. | 1 |
| 2017 | BRGP: a balanced RDF graph partitioning algorithm for cloud storageabstractSummary The continuous growth of resource description framework (RDF) data poses an important challenge on RDF data partitioning that is a vital technique for effective cloud storage. Recently, many partitioning algorithms for large RDF data have been developed, and most of them are based on graph partitioning. However, existing graph partitioning methods could not partition asymmetric RDF data effectively, resulting in a lower performance for cloud storage. This paper proposes a balanced RDF graph partitioning algorithm for storing massive RDF data on cloud. We first devise a modularity‐based multi‐level label propagation algorithm (MMLP) to partition RDF graph roughly and then use a balanced K‐mediods clustering algorithm for finalk‐way partitioning. Balanced RDF graph partitioning algorithm designs an effective label update rule and a balanced modification strategy to achieve a high quality coarsening result and make the partition as equilibrium as possible. Experiments are carried on two representative RDF benchmarks and one real RDF dataset by comparison with two representative graph partitioning methods, that is, METIS and MLP+METIS. Results demonstrate that our proposed scheme can produce a high‐quality partition for massive RDF data storage on cloud. Copyright © 2016 John Wiley & Sons, Ltd. Yonglin Leng, Zhikui Chen, Fangming Zhong, Xiongjiu Li, Yueming Hu 0001 |
Concurr. Comput. Pract. Exp. | 3 |