Zhangmin Huang

dblp:289/7654 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0001-9294-4196ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Trusted Mamba Contrastive Network for Multi-View Clustering
abstract
Multi-view clustering can partition data samples into their categories by learning a consensus representation in an unsupervised way and has received more and more attention in recent years. However, there is an untrusted fusion problem. The reasons for this problem are as follows: 1) The current methods ignore the presence of noise or redundant information in the view; 2) The similarity of contrastive learning comes from the same sample rather than the same cluster in deep multi-view clustering. It causes multi-view fusion in the wrong direction. This paper proposes a novel multi-view clustering network to address this problem, termed as Trusted Mamba Contrastive Network (TMCN). Specifically, we present a new Trusted Mamba Fusion Network (TMFN), which achieves a trusted fusion of multi-view data through a selective mechanism. Moreover, we align the fused representation and the view-specific representation using the Average-similarity Contrastive Learning (AsCL) module. AsCL increases the similarity of view presentation from the same cluster, not merely from the same sample. Extensive experiments show that the proposed method achieves state-of-the-art results in deep multi-view clustering tasks. The source code is available at https://github.com/HackerHyper/TMCN.
Xin Zou 0001, Lei Liu 0029, Zhangmin Huang, Chang Tang, Li-Rong Dai 0001
ICASSP4
2025 Dynamic SRM Curriculum for Trustworthy Multi-modal Classification
abstract
Trustworthy multi-modal learning integrates multiple sources of data reliably. However, the current methods still focus on performance improvement by developing deep multi-modal networks. These approaches frequently encounter challenges due to the inherent non-convex nature of deep neural networks and their vulnerability to local minima, ultimately leading to a diminished ability for generalization. To address this problem, we present a novel curriculum termed the Dynamic SRM Curriculum (DSRMC). Within DSRMC, the deep trustworthy multi-modal networks undergo training with data provided sequentially, progressing from simple to complex samples. This training strategy mimics the human learning process, commencing with fundamental concepts and gradually advancing to tackle more complex and abstract ideas. Building upon DSRMC, we propose an innovative Curriculum Trustworthy Multi-modal Learning (CTML) method. CTML makes it easier to place the learned model in a flatter area, which improves its overall ability for generalization. Comprehensive experiments on three public datasets demonstrate that the proposed CTML performs better than state-of-the-art methods, achieving a maximum improvement of 6.7% on macroF1.
Cui Yu, Xin Zou 0001, Zhangmin Huang, Chenshu Hu, Jun Sun 0014, Bo Lyu, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
ICASSP4
2025 CLIP Multi-modal Hashing for Multimedia Retrieval
Mingkai Sheng, Zhangmin Huang, Jingfei Chang, Jinling Jiang, Lei Liu 0029
MMM (1)3
2025 Decoupling Neural Networks to Leverage Uniform Representation and Balance Personalization and Collaboration in Federated Learning
abstract
Federated learning (FL), a distributed learning paradigm focused on preserving data privacy, faces challenges due to varying data distributions among clients, impacting global model performance. To mitigate data heterogeneity, we propose FedUB-a personalized FL framework leveraging uniform feature representation and balancing personalization and collaboration in the classifier. Specifically, the uniform representation (UR) in FedUB provides all clients with a shared feature extractor and a common representation centroid (RC). Achieving this uniformity involves incorporating a regularization term to reduce the gap between global and local RCs. Additionally, an importance estimation of the parameters in the classifier is provided to partition the parameters into two parts: the personalized component and the collaborated component. Specifically, the personalized component adapts to local data, while the collaborated component prevents the classifier from overfitting local data. Theoretically, we establish the existence of the UR, demonstrating its effectiveness in reducing the average generalization bound. Experiments on benchmark datasets consistently demonstrate the performance gains and improved generalization behavior of FedUB.
Zhangmin Huang, Shaojie Tang 0001, Bo Lyu, Lingfang Zeng
IEEE Trans. Neural Networks Learn. Syst.1
2024 Adaptive Confidence Multi-View Hashing for Multimedia Retrieval
abstract
The multi-view hash method converts heterogeneous data from multiple views into binary hash codes, which is one of the critical technologies in multimedia retrieval. However, the current methods mainly explore the complementarity among multiple views while lacking confidence in learning and fusion. Moreover, in practical application scenarios, the single-view data contains redundant noise. To conduct confidence learning and eliminate unnecessary noise, we propose a novel Adaptive Confidence Multi-View Hashing (ACMVH) method. First, a confidence network is developed to extract useful information from various single-view features and remove noise information. Furthermore, an adaptive confidence multi-view network is employed to measure the confidence of each view and then fuse multi-view features through a weighted summation. Lastly, a dilation network is designed to further enhance the feature representation of the fused features. To the best of our knowledge, we pioneer the application of confidence learning into the field of multimedia retrieval. Extensive experiments on two public datasets show that the proposed ACMVH performs better than state-of-the-art methods (maximum increase of 3.24%). The source code is available at https://github.com/HackerHyper/ACMVH.
Zhangmin Huang, Lei Liu 0029, Lingfang Zeng
ICASSP3
2024 Boosted Curriculum Multi-View Hashing for Multimedia Retrieval
abstract
The multi-view hash method plays a pivotal role in multimedia retrieval, transforming diverse data from multiple perspectives into binary hash codes. While existing methods primarily emphasize complementarity across multiple views, they often face challenges associated with the non-convex nature of deep neural networks, ultimately causing a decrease in generalization ability. To overcome this limitation, we propose a novel curriculum calledAutomatic Multiple Loss Curriculum(AMLC). In AMLC, the deep multi-view hashing network undergoes training with data presented sequentially, progressing from simple to complex samples. This training strategy mirrors the human learning process, commencing with fundamental concepts and progressively advancing to tackle more intricate and abstract ideas. Building upon AMLC, we propose theBoosted Curriculum Multi-View Hashing(BCMVH) method. BCMVH facilitates the positioning of the learned model in a more flat region, enhancing its overall generalization capability. Extensive experiments conducted on three public datasets demonstrate that the proposed BCMVH outperforms state-of-the-art methods, achieving a maximum improvement of 3.17% in terms of mean Average Precision.
Zhangmin Huang, Lei Liu 0029, Chang Tang, Li-Rong Dai 0001
IEEE Signal Process. Lett.2
2023 Deep Metric Multi-View Hashing for Multimedia Retrieval
abstract
Learning the hash representation of multi-view heterogeneous data is an important task in multimedia retrieval. However, existing methods fail to effectively fuse the multi-view features and utilize the metric information provided by the dissimilar samples, leading to limited retrieval precision. Current methods utilize weighted sum or concatenation to fuse the multi-view features. We argue that these fusion methods cannot capture the interaction among different views. Furthermore, these methods ignored the information provided by the dissimilar samples. We propose a novel deep metric multi-view hashing (DMMVH) method to address the mentioned problems. Extensive empirical evidence is presented to show that gate-based fusion is better than typical methods. We introduce deep metric learning to the multi-view hashing problems, which can utilize metric information of dissimilar samples. On the MIR-Flickr25K, MS COCO, and NUS-WIDE, our method outperforms the current state-of-the-art methods by a large margin (up to 15.28 mean Average Precision (mAP) improvement).
Xiaohu Ruan, Yongli Cheng, Zhangmin Huang, Lingfang Zeng
ICME4
2022 pFedGF: Enabling Personalized Federated Learning via Gradient Fusion
abstract
Data heterogeneity is one of the main challenges faced by federated learning (FL). Unlike traditional FL methods (e.g. FedAvg) which train a global model for all clients, personalized federated learning (PFL) can address the above problem by training a personalized model for each client. Current mainstream PFL researches first obtain a global model through collaborative training among all clients and then fine-tune the global model on each client's local data to obtain personalized models. However, this two-staged approach has a drawback: when the heterogeneity of different clients is large, the obtained final global model can deviate from the distributions of all clients, and therefore is not a good starting point for updating personalized models. In this paper, we propose pFedGF, a new PFL method based on gradient fusion. Different from traditional two-staged PFL, in each round of pFedGF, each client maintains two gradients simultaneously, a global gradient to capture information from all clients, and a local gradient that reflects the specific distribution of each client. The two gradients are fused to obtain the updated direction of the personalized model for each client. We carried out experiments on MNIST, FMNIST, and CIFAR-10 datasets. The results demonstrate that in the presence of data heterogeneity, pFedGF outperforms other PFL methods.
Xinghao Wu, Jianwei Niu 0002, Xuefeng Liu 0001, Tao Ren 0001, Zhangmin Huang, Zhetao Li
IPDPS5
2021 CIC-FL: Enabling Class Imbalance-Aware Clustered Federated Learning over Shifted Distributions
Yanan Fu, Xuefeng Liu 0001, Shaojie Tang 0001, Jianwei Niu 0002, Zhangmin Huang
DASFAA (1)5