Mingfu Xiong

dblp:167/9610 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0003-0487-0356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 4SNet: Spatial and Spectrum Self-adaptive Synergy Network for Visible-Infrared Person Re-identification
Mingfu Xiong, Feiyang Luo, Yifei Guo, Aziz Alotaibi, Sambit Bakshi, Javier Del Ser, Khan Muhammad 0001
Pattern Recognit.1
2026 HPRNet: Human Parsing Reconstruction With Non-Local Multi-Scale Perception Network for Cloth-Changing Person Re-Identification
abstract
Cloth-changing Person Re-Identification (CC-ReID) is a challenging data modeling task that involves identifying specific pedestrians wearing different outfits. Existing methods primarily focus on altering clothing color and directly reconstructing appearance to extract features independent of the clothes. Real pedestrians differ in height, body shape, etc. Such methods are prone to losing the intrinsic information of the original sample (i.e., the person identity) owing to the absence of contextual phenomena (e.g., texture structure and local correlation), which decreases the recognition performance. To address this problem, we propose a framework called HPRNet, or ”Human Parsing Reconstruction with Non-Local Multi-Scale Perception Network,” which includes a non-local weighted multi-scale perception (NWMP) module and a parsing reconstruction exploration (PRE) module. In particular, the proposed NWMP module effectively captures the global receptive field of a sample and obtains a contextual correlation between non-neighboring pixels within the sample image. The PRE module was used to achieve a more accurate reconstruction of human body components with a clothing parsing model to better distinguish features related to or unrelated to clothes. Extensive experiments were conducted on CC-ReID public datasets (LTCC, PRCC, and CCVID) to demonstrate the effectiveness and competitiveness of the proposed method with state-of-the-art (SOTA) baselines for this complex modeling task.
Mingfu Xiong, Longlong Ge, Ruimin Hu, Khan Muhammad 0001, Sambit Bakshi, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Online visual multi-object tracking via blind super-resolution and clustering learning
Mingfu Xiong, Guoheng Wei, Jiuxing Zhou, Qingrui Meng
Vis. Comput.3
2025 Towards Hazardous Activity Recognition for A Novel Real-World Dataset
abstract
Detecting hazardous activities is essential for ensuring safety. However, existing datasets often lack coverage of the nuanced and diverse hazards present in indoor environments, which hinders the development of a specialized model. To address this, we introduce the Real-World Hazardous Activities Dataset (RHAD), a novel and diverse video dataset specifically curated for recognizing hazardous activities in real-world indoor settings. Leveraging RHAD, we introduce HazardNet, a hybrid deep-learning architecture designed for hazardous activity recognition. HazardNet integrates local and global spatial-temporal representation modules to effectively capture complex patterns, enabling a robust understanding of the activity. We perform comprehensive evaluations by benchmarking against a range of state-of-the-art activity recognition models. Experimental results show that our proposed model performs significantly better, surpassing the latest model, VideoMamba, with a 9.2% accuracy gain. Moreover, by providing the dataset and an effective recognition model, our work lays the foundation for further research, paving the way for enhanced safety measures and preventive interventions. The dataset and code are available at https://github.com/ShehzadCS18/RHAD.
Shehzad Ali, Md Tanvir Islam, Ikhyun Lee, Mingfu Xiong, Minh-Son Dao, Saeed Anwar, Sambit Bakshi, Khan Muhammad 0001
ACM Multimedia4
2025 STCGen: Sketch-based Text-to-Clothing Image Generation with Contour and Style Consistency
abstract
In modern fashion design field, it is a mainstream practice to generate clothing images by combining sketch and text. However, the image quality generated by existing multimodal methods combining sketches and text descriptions is suboptimal, as the clothing in the generated image often lacks contour accuracy and stylistic coherence. In this paper, we present STCGen, an advanced multimodal framework that uses both sketches and text to generate clothing images with improved contours and more consistent style. First, we introduce the sketch prior embedding module, which processes sketches to extract key structural features and ensure the consistency of contours, thereby enhancing image details. Second, we propose a cross space attention mechanism to address the issue of text information loss and ensure stylistic consistency, thereby enhancing overall image coherence. Finally, we propose a network simplification scheme to reduce complexity without compromising the quality of resulting images. Experimental results demonstrate that our method excels in generating high-fidelity clothing images.
Chunxia Xiao, Ruhan He, Jia Chen 0012, Mingfu Xiong, Tao Peng 0006, Xinrong Hu
MMAsia7
2025 Adaptive Clustering and Weighted Regularization Contrastive Learning Framework for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) has recently gained significant attention from researchers. ReID matches images of the same person from different camera views in various scenes without any labels. Existing clustering methods primarily rely on a fixed threshold (the maximum distance between sample points and clustering centroids) and overlook the importance of adjusting this threshold during continuous model optimization. This mismatch between clustering thresholds and inter- or intra-class spacing reduces clustering accuracy. To address this issue, this study proposes an Adaptive Clustering and Weighted Regularization Contrastive Learning (ACWRCL) framework for unsupervised person ReID. The ACWRCL framework comprises two main components: (1) the Clustering Threshold Adaptive Adjustment (CTAA) module, and (2) the Weighted Regularization Contrastive Learning (WRCL) module. The CTAA module dynamically adjusts the clustering threshold to align with model optimization, ensuring that the threshold remains within an appropriate range to prevent under- or over-robustness in the clustering model. The WRCL module uses the similarity ratio between the query sample and the clustering centroid relative to the overall similarity of all samples with the same labels as the query sample. This ratio is used as the weight in the loss function to penalize incorrect clustering and improve pseudo-label generation accuracy. Extensive experiments on public ReID datasets—Market-1501, MSMT17, Veri776, CUHK03, and PersonX—demonstrate the effectiveness of the proposed method.
Mingfu Xiong, Kaikang Hu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001, Javier Del Ser, Xiaokang Yang 0001, Bin Sheng 0001
IEEE Trans. Multim.1
2025 Face photo-sketch portraits transformation via generation pipeline
Mengsi Guo, Mingfu Xiong, Xinrong Hu, Tao Peng 0006
Vis. Comput.2
2024 PDET: Progressive Diversity Expansion Transformer for Cross-Modality Visible-Infrared Person Re-identification
Mingfu Xiong, Jingbang Liang, Yifei Guo, Ikhyun Lee, Sambit Bakshi, Khan Muhammad 0001
ICPR (14)1
2024 Domain generalized person reidentification based on skewness regularity of higher-order statistics
Mingfu Xiong, Ruimin Hu, Zhongyuan Wang 0001, Javier Del Ser, Khan Muhammad 0001, Zixiang Xiong
Knowl. Based Syst.1
2024 Inter-camera Identity Discrimination for Unsupervised Person Re-identification
abstract
Unsupervised person re-identification (Re-ID) has garnered significant attention because of its data-friendly nature, as it does not require labeled data. Existing approaches primarily address this challenge by employing feature-clustering techniques to generate pseudo-labels. In addition, camera-proxy-based methods have emerged because of their impressive ability to cluster sample identities. However, these methods often blur the distinctions between individuals within inter-camera views, which is crucial for effective person re-ID. To address this issue, this study introduces an inter-camera-identity-difference-based contrastive learning framework for unsupervised person Re-ID. The proposed framework comprises two key components: (1) a different sample cross-view close-range penalty module and (2) the same sample cross-view long-range constraint module. The former aims at penalizing excessive similarity among different subjects across inter-camera views, whereas the latter mitigates the challenge of excessive dissimilarity among the same subject across camera views. To validate the performance of our method, we conducted extensive experiments on three existing person Re-ID datasets (Market-1501, MSMT17, and PersonX). The results demonstrate the effectiveness of the proposed method, which shows a promising performance. The code is available at https://github.com/hooldylan/IIDCL .
Mingfu Xiong, Kaikang Hu, Zhihan Lyu, Zhongyuan Wang 0001, Ruimin Hu, Khan Muhammad 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Cross-cycle Transformer-based Stitching Method for Low-resolution Borehole Images
abstract
The stitching of borehole images has an important predictive role in safety analysis in the field of geotechnical engineering and intelligent geological exploration. Applying traditional image stitching methods that designed specifically for high-resolution images to low-resolution images will lead to blurred stitching results, stitching seams, fewer matched feature points and difficulties in massive image stitching. To address these problems, we propose an autoencoder-based coarse-to-fine feature extraction network, which can extract image features with high semantic and improves the accuracy of the feature point matching. Besides, we design a cross-cycle Transformer-based image stitching framework, which increase the number of matching feature points by Cross-QuadTree attention and stitch image by affine transformation. Experimental results show that the proposed method can effectively stitch low-resolution geotechnical borehole images with satisfactory visual quality.
Jia Chen 0012, Zhenpeng Fu, Mingfu Xiong, Xinrong Hu, Tao Peng 0006
ICME4
2022 A Speech Enhancement Method Combining Two-Branch Communication and Spectral Subtraction
Ruhan He, Yajun Tian, Zhenghao Chang, Mingfu Xiong
ICONIP (5)5
2022 Facial expressions recognition with multi-region divided attention networks for smart education cloud applications
Yifei Guo, Mingfu Xiong, Zhongyuan Wang 0001, Xinrong Hu, Mohammad Hijji
Neurocomputing3
2022 Multi-scale Interest Dynamic Hierarchical Transformer for sequential recommendation
Nana Huang, Ruimin Hu, Mingfu Xiong, Xiaoran Peng, Xiaodong Jia 0005, Lingkun Zhang
Neural Comput. Appl.3
2021 Cascaded Cross-Domain Fusion of Virtual Try-On
abstract
Image-based virtual try-on, aiming to fit new in-shop clothes into a person image, has gained extensive attention in the fields of computer vision and image process community. However, the existing methods are difficult to generate photo-realistic try-on images when large-scale deformations or large occlusions occur. To address this issue, we propose a novel two stage visual try-on network. Specifically, in the first stage, we used a shape matching model to learn the geometric transformation of in-shop clothes. For the second stage, an U-net with cascaded attention mechanism is presented to learn the composition mask which adjust the clothes and rendered persons. The adjusted clothes and the rendered person are combined by the composition mask to get the final try-on result. Experimental results have shown that our method can generate photo-realistic images with no occlusion.
Xinrong Hu, Tao Peng 0006, Mingfu Xiong, Feng Yu 0017, Li Li 0094
BIBM4
2021 A Triplet Appearance Parsing Network for Person Re-Identification
abstract
As one of the specific vision tasks, person re-identification has become a prevalent research topic in the field of multimedia and computer vision. However, existing feature extraction methods, originating from the quality of the bounding boxes which could cause the inhomogeneity and incoherence of person representation for cluttered backgrounds, are difficult to adapt the challenges of the harsh real-world scenarios. This study develops a Triplet person Appearances Parsing Framework (TAPF) which eliminates the surrounding interference factors of bounding boxes for person re-identification. The framework consists of a triplet person parsing network and an integration mechanism for person local and global appearance information. Concretely, the triplet parsing network includes a channel parsing module, a position parsing module and a color parsing module, which are used to extract the person channel parsing descriptor, regional descriptor and color perception descriptor, respectively. Then, a local and global flatten gaussian operations are performed to integrate the person appearance parsing descriptors to obtain more discriminative features for the person representation. The experimental results have been conducted to validate our proposed algorithm can achieve a better performance for person re-identification on several public datasets, i.e., VIPeR and Market-1501, respectively.
Mingfu Xiong, Zhongyuan Wang 0001, Ruhan He, Xinrong Hu, Xiao Qin 0001, Jia Chen 0012
ICASSP1
2020 Triple Attention Network for Clothing Parsing
Ruhan He, Mingfu Xiong, Xiao Qin 0001, Junping Liu, Xinrong Hu
ICONIP (1)3
2020 Mobile person re-identification with a lightweight trident CNN
Mingfu Xiong, Dan Chen 0001, Xiaoqiang Lu
Sci. China Inf. Sci.1
2020 HybridGAN: hybrid generative adversarial networks for MR image synthesis
Jia Chen 0012, Mingfu Xiong, Tao Peng 0006, Minghua Jiang, Xiao Qin 0001
Multim. Tools Appl.3
2019 Rain Streak Removal via Multi-scale Mixture Exponential Power Model
abstract
Rain streaks severely hamper the visible performance of the outdoor surveillance videos, which becomes an attractive issue in recent computer vision research. Existing methods usually encode rain streaks into Gaussian Mixture Model (GM-M). However, the limited number of Gaussian components in the GMM compromises the ability of the model in fitting real noise, such as sparse noise, which is exactly the characteristic of the rain streaks. In this paper, a novel model named Mixture Exponential Power Model (MEPM) is exploited. It sets multiple Laplace noise components and expands the representation capability for the sparse noise. Moreover, considering that the rain streaks in a video occur in different distances from the camera, we encode rain streaks into Multi-scale Mixture Exponential Power Model. The model is opti-mized by expectation-maximization (EM) algorithm and La-grange multiplier strategy. Experiments are implemented on synthetic and real rain videos and verify the superiority of the proposed method, compared with state-of-the-art methods.
Jun Chen 0001, Zhen Han 0002, Mingfu Xiong, Chao Liang 0001, Zhongyuan Wang 0001
ICASSP4
2019 Person re-identification with multiple similarity probabilities using deep metric learning for efficient smart security applications
Mingfu Xiong, Dan Chen 0001, Jun Chen 0001, Jingying Chen 0001, Benyun Shi, Chao Liang 0001, Ruimin Hu
J. Parallel Distributed Comput.1
2016 Person Re-Identification via Multiple Coarse-to-Fine Deep Metrics
abstract
Person re-identification, aiming to identify images of the same person from various cameras views in different places, has attracted a lot of research interests in the field of artificial intelligence and multimedia. As one of its popular research directions, the metric learning method plays an important role for seeking a proper metric space to generate accurate feature comparison. However, the existing metric learning methods mainly aim to learn an optimal distance metric function through a single metric, making them difficult to consider multiple similar relationships between the samples. To solve this problem, this paper proposes a coarse-to-fine deep metric learning method equipped with multiple different Stacked Auto-Encoder (SAE) networks and classification networks. In the perspective of the human's visual mechanism, the multiple different levels of deep neural networks simulate the information processing of the brain's visual system, which employs different patterns to recognize the character of objects. In addition, a weighted assignment mechanism is presented to handle the different measure manners for final recognition accuracy. The experimental results conducted on two public datasets, i.e., VIPeR and CUHK have shown the prospective performance of the proposed method.
Mingfu Xiong, Jun Chen 0001, Zheng Wang 0007, Zhongyuan Wang 0001, Ruimin Hu, Chao Liang 0001, Daming Shi 0001
ECAI1