EDBT 2026 Demo / reviewers in the wild / expert
Fuqing Zhu
dblp:200/8105
· DBLP profile ↗
28ranked-venue papers
4as first author
15since 2021 · last 2025
0000-0001-7061-3329ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 12 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evidence-Claim Relevance Decoupling for Multimodal Fact-CheckingabstractWith the proliferation of misinformation on social media, the need for fact-checking has increased, garnering the attention of researchers. The main task of fact-checking is to analyze the semantic relationship and logic between the evidence and the claim, so as to determine the veracity of the claim. The semantic relationship between the claim and the evidence can be strong or weak. However, previous studies ignores the disparities in the strength of semantic relationship between evidence and claims in the process of evidence modeling for fact-checking, neglects the role of weak semantically related evidence in verifying claims. Weakly semantically related evidence logically determines the veracity of the claim sometimes. This paper proposes an evidence-claim relevance decoupling (ECRD) framework for multimodal fact-checking, where the claim and the evidence could be separated by decoupling the strong and weak relevant semantic parts in multimodal dual-stream models. Specifically, we first project text and image modalities into two distinct spaces. One space focuses on tightly coupled semantics between the claim and evidence, reducing the gap. Another space explores the loosely coupled parts, analyzing the various meaning between claim and evidence to reveal potential underlying connections. Then, the fully connected visual and textual graphs are constructed for above two spaces. Finally, the responses obtained by interacting in above graphs are fed into the fusion layer. On FACTIFY dataset, which features diverse text-image relationships, and MultiFact dataset, experimental results show that the proposed ECRD framework could improve the performance of classification significantly compared to baseline models and produces a competitive performance. Jianliang Zeng, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
CSCWD | 2 |
| 2025 | Mutual Information-Driven Rationale Distillation for Hateful Memes DetectionabstractHateful memes typically disseminate discriminatory statements directly or indirectly based on race, religion, or other characteristics. Previous research has made enlightening exploration in detecting explicit hateful memes. However, these methods overlook the analysis of implicit hate, which is particularly challenging as it requires comprehensive background knowledge to be accurately identified. To this end, this paper proposes a mutual information-driven rationale distillation model (MIRD) that leverages multimodal reasoning knowledge extracted from MLLM to reveal hateful memes. The model is divided into two stages. The first stage is abductive reasoning, which provides contextual background information to guide the revelation of the implicit meaning of memes. The second stage is rationale distillation, which transfers the background knowledge provided by MLLM to small models, thereby reducing computational requirements while generating labels and explanations for memes. Besides, we design a mutual information module aimed at maximizing information sharing between label prediction and explanation generation tasks, enhancing the consistency between tasks. Extensive experimental results demonstrate that MIRD achieves state-of-the-art performance on three widely-used datasets. Chuanpeng Yang, Fuqing Zhu |
ECAI | 3 |
| 2025 | Visual Perception Uncertainty Learning for Hallucination Detection in Large Vision-Language ModelsabstractHallucination remains a significant challenge which constrains the development of large vision-language models (LVLMs). Therefore, reliable hallucination detection has become a critical step in LVLMs evaluation and real-world deployment. Many previous studies have explored hallucination detection in LVLMs, with uncertainty-based approaches being widely adopted due to the independence from external tools and relatively low resource consumption. However, we observe that uncertainty does not always completely correlate with hallucination. Therefore, uncertainty-based methods may fail in certain cases, such as instances exhibiting high uncertainty but non-hallucination. To address this issue, we propose a framework called Visual Perception Uncertainty Learning (VisPUL) for hallucination detection in LVLMs. Specifically, VisPUL integrates visual information into uncertainty learning directly, allowing to capture uncertainty and visual-text consistency simultaneously. VisPUL improves the insufficiency of uncertainty methods that rely only on text output, providing enhanced generalizability and reliability. Extensive experiments conducted on the M-HalDetect and POPE datasets, covering both open-ended and yes-or-no tasks. Experimental results demonstrate that VisPUL significantly outperforms several strong baseline methods across different LVLMs. Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 2 |
| 2024 | Uncertainty-Guided Modal Rebalance for Hateful Memes DetectionabstractHateful memes detection is a challenging multimodal understanding task that requires comprehensive learning of vision, language, and cross-modal interactions.Previous research has focused on developing effective fusion strategies for integrating hate information from different modalities.However, these methods excessively rely on cross-modal fusion features, ignoring the modality uncertainty caused by the contribution degree of each modality to hate sentiment and the modality imbalance caused by the dominant modality suppressing the optimization of another modality.To this end, this paper proposes an Uncertainty-guided Modal Rebalance (UMR) framework for hateful memes detection.The uncertainty of each meme is explicitly formulated by designing stochastic representation drawn from a Gaussian distribution for aggregating cross-modal features with unimodal features adaptively.The modality imbalance is alleviated by improving cosine loss from the perspectives of intermodal feature and weight vectors constraints.In this way, the suppressed unimodal representation ability in multimodal models would be unleashed, while the learning of modality contribution would be further promoted.Extensive experimental results demonstrate that the proposed UMR produces the state-of-the-art performance on four widely-used datasets. Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ACL (1) | 3 |
| 2024 | Few-Shot Learning for Cold-Start RecommendationabstractCold-start is a significant problem in recommender systems. Recently, with the development of few-shot learning and meta-learning techniques, many researchers have devoted themselves to adopting meta-learning into recommendation as the natural scenario of few-shots. Nevertheless, we argue that recent work has a huge gap between few-shot learning and recommendations. In particular, users are locally dependent, not globally independent in recommendation. Therefore, it is necessary to formulate the local relationships between users. To accomplish this, we present a novel Few-shot learning method for Cold-Start (FCS) recommendation that consists of three hierarchical structures. More concretely, this first hierarchy is the global-meta parameters for learning the global information of all users; the second hierarchy is the local-meta parameters whose goal is to learn the adaptive cluster of local users; the third hierarchy is the specific parameters of the target user. Both the global and local information are formulated, addressing the new user’s problem in accordance with the few-shot records rapidly. Experimental results on two public real-world datasets show that the FCS method could produce stable improvements compared with the state-of-the-art. Songlin Hu 0001, Fuqing Zhu, Qiannan Zhu |
LREC/COLING | 3 |
| 2024 | Uncertainty-Aware Cross-Modal Alignment for Hate Speech DetectionabstractHate speech detection has become an urgent task with the emergence of huge multimodal harmful content (, memes) on social media platforms. Previous studies mainly focus on complex feature extraction and fusion to learn discriminative information from memes. However, these methods ignore two key points: 1) the misalignment of image and text in memes caused by the modality gap, and 2) the uncertainty between modalities caused by the contribution degree of each modality to hate sentiment. To this end, this paper proposes an uncertainty-aware cross-modal alignment (UCA) framework for modeling the misalignment and uncertainty in multimodal hate speech detection. Specifically, we first utilize the cross-modal feature encoder to capture image and text feature representations in memes. Then, a cross-modal alignment module is applied to reduce semantic gaps between modalities by aligning the feature representations. Next, a cross-modal fusion module is designed to learn semantic interactions between modalities to capture cross-modal correlations, providing complementary features for memes. Finally, a cross-modal uncertainty learning module is proposed, which evaluates the divergence between unimodal feature distributions to to balance unimodal and cross-modal fusion features. Extensive experiments on five publicly available datasets show that the proposed UCA produces a competitive performance compared with the existing multimodal hate speech detection methods. Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
LREC/COLING | 2 |
| 2024 | Integrating Open-domain Knowledge via Large Language Model for Multimodal Fake News DetectionabstractMultimodal fake news propagation on social media has emerged as a primary concern for both the public and the government. Distinguishing fake news from genuine content is challenging due to their close resemblance. Moreover, fake news often presents information that contradicts real-world facts, underscoring the necessity of incorporating comprehensive open-domain knowledge. However, current knowledge-based detection methods struggle to improve two key areas in integrating open-domain knowledge for fake news detection: i) query acquisition for retrieval and ii) utilization of retrieved knowledge. The challenges stem from limitations in understanding the semantics and reasoning about the relationship between open-domain knowledge and the content of the news. With the emergence of the Large Language Model (LLM), Natural Language Processing (NLP) has witnessed a revolution where impressive semantic understanding and reasoning abilities are significantly improved. This paper proposes a novel Open-domain Knowledge Integrated (OKI) framework for multimodal fake news detection, featuring two LLM-based agents that collaboratively leverage open-domain knowledge. One agent generates appropriate queries to retrieve knowledge, while the other filters out irrelevant retrieved knowledge. Experimental results demonstrate a significant performance improvement of OKI over established baselines on the Weibo and Twitter datasets. Anbin Xie, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
CSCWD | 2 |
| 2024 | Pyramidal Cross-Modal Transformer with Sustained Visual Guidance for Multi-Label Image ClassificationabstractMulti-label image classification poses a formidable challenge due to the presence of multiple objects in each image, rendering it notably complex to decipher the visual content comprehensively. Discriminating between multiple objects necessitates the establishment of robust visual label dependencies. Previous methods attempt to formulate cross-modal interaction or one-shot co-occurrence relationship guidance. However, it not only exhibits limitations when handling occluded or blurry objects but also fails to fully leverage the diverse hierarchical properties for sustainably guiding the learning process of label dependencies. To sustainably establish hierarchical visual label dependencies, this paper introduces a Pyramidal Cross-modal Transformer framework for MLIC tasks. Specifically, the pyramidal visual guidance layer parses the visual features into a multi-resolution pyramid structure, allowing the updated visual-related information to provide sustained guidance for label semantics. This surpasses the conventional pre-processing of co-occurrence relationships. Besides, the hybrid modal interaction layer is proposed to effectively mitigate the semantic disparities between visual and label information with modal-blended indiscriminate attention, replacing vanilla self-attention. Several combination blocks consisting of these two layers are integrated and embedded within the encoder-decoder structure to facilitate the exploration of meticulous visual label dependencies. Extensive experiments on two widely-used benchmarks, including MS-COCO and PASCAL VOC 2007, consistently demonstrate that PCMT could provide state-of-the-art results. Ruyun Wang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ICMR | 3 |
| 2023 | Invariant Meets Specific: A Scalable Harmful Memes Detection FrameworkabstractHarmful memes detection is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. Current research on this task mainly focuses on multimodal dual-stream models. However, the existing works ignore the misalignment of the memes caused by the modality gap. Moreover, the cross-modal interaction in the dual-stream models is insufficient to identify harmful memes. To this end, this paper proposes a scalable invariant and specific modality (ISM) representations framework via graph neural networks. The proposed ISM framework provides a comprehensive and disentangled view for memes and promotes inter-modal interaction. Specifically, ISM projects each modality to two distinct spaces. The first space is modality-invariant, learning the corresponding commonalities and reducing the modality gap. The second space is modality-specific, holding the distinctive characteristics of each modality and complementing the common latent features captured in invariant spaces. Then, we construct fully connected visual and textual graphs for each space. The unimodal graphs are fused to dynamically balance inter-modal and intra-modal relationships, which are complementary to the dual-stream models. Finally, an adaptive module is designed to weigh the proportion of each fusion graph for memes. Moreover, the mainstream multimodal dual-stream models could be employed as the backbone flexibly. Extensive experiments on five publicly available datasets show that the proposed ISM provides a stable improvement over baselines and produces a competitive performance compared with the existing harmful memes detection methods. Chuanpeng Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 2 |
| 2023 | ContE: contextualized knowledge graph embedding for circular relations
Shangwen Lv, Fuqing Zhu, Longtao Huang, Songlin Hu 0001 |
Data Min. Knowl. Discov. | 4 |
| 2022 | Cross-Layer Aggregation with Transformers for Multi-Label Image ClassificationabstractMulti-label image classification task aims to predict multiple object labels in a given image and faces the challenge of variable-sized objects. Limited by the size of CNN convolution kernels, existing CNN-based methods have difficulty capturing global dependencies and effectively fusing multiple layers features, which is critical for this task. Recently, transformers have utilized multi-head attention to extract feature with long range dependencies. Inspired by this, this paper proposes a Cross-layer Aggregation with Transformers (CAT) framework, which leverages transformers to capture the long range dependencies of CNN-based features with Long Range Dependencies module and aggregate the features layer by layer with Cross-Layer Fusion module. To make the framework efficient, a multi-head pre-max attention is designed to reduce the computation cost when fusing the high-resolution features of lower-layers. On two widely-used benchmarks (i.e., VOC2007 and MS-COCO), CAT provides a stable improvement over the baseline and produces a competitive performance. Weibo Zhang, Fuqing Zhu, Jizhong Han, Tao Guo 0006, Songlin Hu 0001 |
ICASSP | 2 |
| 2022 | UFI: A Unified Feature Interaction Framework for Multi-Label Image ClassificationabstractMulti-label image classification (MLIC) is a more challenging task compared with single-label image classification due to multiple concepts targets, and complex visual relationships should be formulated. Convolutional Neural Network (CNN) and Visual Transformer (ViT) have shown superior performance in local and global feature representations, respectively. However, the interactions between local and global features are neglected in current works. To further formulate the critical interactions, this paper designs a Unified Feature Interaction (UFI) framework, aiming to integrate the selected local features with global features based on CNN and ViT, simultaneously. The proposed UFI includes two key modules: Class-Related Feature Selection (CRFS) and Feature Interaction Attention (FIA) modules. Specifically, according to the activation map, CRFS selects target regions by the preliminary calculation of predicted scores. FIA enables the significant local-global feature interaction based on the selected target regions and whole image. We initially attempted to interact with local and global features for multi-label image classification. UFI provides a stable improvement over the baseline and produces a new state-of-the-art result on MS-COCO and VOC2007. Weibo Zhang, Ziang Yang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ICME | 4 |
| 2022 | Multimodal Hate Speech Detection via Cross-Domain Knowledge TransferabstractNowadays, the hate speech diffusion of texts and images in social network has become the mainstream compared with the diffusion of texts-only, raising the pressing needs of multimodal hate speech detection task. Current research on this task mainly focuses on the construction of multimodal models without considering the influence of the unbalanced and widely distributed samples for various attacks in hate speech. In this situation, introducing enhanced knowledge is necessary for understanding the attack category of hate speech comprehensively. Due to the high correlation between hate speech detection and sarcasm detection tasks, this paper makes an initial attempt of common knowledge transfer based on the above two tasks, where hate speech detection and sarcasm detection are defined as primary and auxiliary tasks, respectively. A scalable cross-domain knowledge transfer (CDKT) framework is proposed, where the mainstream vision-language transformer could be employed as backbone flexibly. Three modules are included, bridging the semantic, definition and domain gaps simultaneously between primary and auxiliary tasks. Specifically, semantic adaptation module formulates the irrelevant parts between image and text in primary and auxiliary tasks, and disentangles with the text representation to align the visual and word tokens. Definition adaptation module assigns different weights to the training samples of auxiliary task by measuring the correlation between samples of the auxiliary and primary task. Domain adaptation module minimizes the feature distribution gap of samples in two tasks. Extensive experiments show that the proposed CDKT provides a stable improvement compared with baselines and produces a competitive performance compared with some existing multimodal hate speech detection methods. Chuanpeng Yang, Fuqing Zhu, Guihua Liu, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 2 |
| 2021 | An Adaptive Hybrid Framework for Cross-domain Aspect-based Sentiment AnalysisabstractCross-domain aspect-based sentiment analysis aims to utilize the useful knowledge in a source domain to extract aspect terms and predict their sentiment polarities in a target domain. Recently, methods based on adversarial training have been applied to this task and achieved promising results. In such methods, both the source and target data are utilized to learn domain-invariant features through deceiving a domain discriminator. However, the task classifier is only trained on the source data, which causes the aspect and sentiment information lying in the target data can not be exploited by the task classifier. In this paper, we propose an Adaptive Hybrid Framework (AHF) for cross-domain aspect-based sentiment analysis. We integrate pseudo-label based semi-supervised learning and adversarial training in a unified network. Thus the target data can be used not only to align the features via the training of domain discriminator, but also to refine the task classifier. Furthermore, we design an adaptive mean teacher as the semi-supervised part of our network, which can mitigate the effects of noisy pseudo labels generated on the target data. We conduct experiments on four public datasets and the experimental results show that our framework significantly outperforms the state-of-the-art methods. Fuqing Zhu, Pu Song, Jizhong Han, Tao Guo 0006, Songlin Hu 0001 |
AAAI | 2 |
| 2021 | Aligning the training and evaluation of unsupervised text style TransferabstractIn the text style transfer task, models modify the attribute style of given texts while keeping the style-irrelevant content unchanged. Previous work has proposed many approaches on the non-parallel corpus (without style-to-style training pairs). These approaches are mostly motivated by heuristic intuition and fail to precisely control texts’ attributes, such as the amount of preserved semantics, which leaves discrepancies between training and evaluation. This paper proposes a novel training method based on the evaluation metrics to address the discrepancy issue. Specifically, the model first evaluates different aspects of the transferred texts and provides the differentiable quality approximations by employing extra supervising modules. Then the model is optimized by bridging the gap between approximations and expectations. Extensive experiments conducted on two sentiment style datasets demonstrate the effectiveness of our proposal compared with some competitive baselines. Wanhui Qian, Fuqing Zhu, Jinzhu Yang, Jizhong Han, Songlin Hu 0001 |
ICASSP | 2 |
| 2020 | Symmetric Metric Learning with Adaptive Margin for RecommendationabstractMetric learning based methods have attracted extensive interests in recommender systems. Current methods take the user-centric way in metric space to ensure the distance between user and negative item to be larger than that between the current user and positive item by a fixed margin. While they ignore the relations among positive item and negative item. As a result, these two items might be positioned closely, leading to incorrect results. Meanwhile, different users usually have different preferences, the fixed margin used in those methods can not be adaptive to various user biases, and thus decreases the performance as well. To address these two problems, a novel Symmetic Metric Learning with adaptive margin (SML) is proposed. In addition to the current user-centric metric, it symmetically introduces a positive item-centric metric which maintains closer distance from positive items to user, and push the negative items away from the positive items at the same time. Moreover, the dynamically adaptive margins are well trained to mitigate the impact of bias. Experimental results on three public recommendation datasets demonstrate that SML produces a competitive performance compared with several state-of-the-art methods. Fuqing Zhu, Wanhui Qian, Liangjun Zang, Jizhong Han, Songlin Hu 0001 |
AAAI | 3 |
| 2020 | An Event-Oriented Neural Ranking Model for News RetrievalabstractEvent-oriented news retrieval (ENR) is the task of retrieving news articles related to the specific event in response to the event-oriented query. Previous approaches usually focus on optimizing traditional retrieval models through hand-crafted features from the perspective of new articles. However, these approaches often fail to work well in reality, as they do not consider the essential natures of the event, i.e., dynamics, coupling. In this paper, we propose a novel and effective event-oriented neural ranking model for news retrieval (ENRMNR). Our model exploits a deep attention mechanism to tackle the dynamics and coupling derived from event evolution. Specifically, the word-level bidirectional attention allows the model to identify which query words about the subevent are related to the news article words, and vice-versa, in order to tackle the dynamics. Moreover, the hierarchical attention at passage-level and document-level allows it to capture fine-grained event representations for the coupling between different events within a news article. Experimental results on real-world datasets demonstrate that ENRMNR model significantly outperforms competitive models. Wanhui Qian, Liangjun Zang, Fuqing Zhu, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001 |
CIKM | 4 |
| 2020 | Integrating External Event Knowledge for Script LearningabstractScript learning aims to predict the subsequent event according to the existing event chain.Recent studies focus on event co-occurrence to solve this problem.However, few studies integrate external event knowledge to solve this problem.With our observations, external event knowledge can provide additional knowledge like temporal or causal knowledge for understanding event chain better and predicting the right subsequent event.In this work, we integrate event knowledge from ASER (Activities, States, Events and their Relations) knowledge base to help predict the next event.We propose a new approach consisting of knowledge retrieval stage and knowledge integration stage.In the knowledge retrieval stage, we select relevant external event knowledge from ASER.In the knowledge integration stage, we propose three methods to integrate external knowledge into our model and infer final answers.Experiments on the widely-used Multi-Choice Narrative Cloze (MCNC) task show our approach achieves state-of-the-art performance compared to other methods. Shangwen Lv, Fuqing Zhu, Songlin Hu 0001 |
COLING | 2 |
| 2020 | Structural Position Network for Aspect-Based Sentiment Classification
Pu Song, Wei Jiang 0028, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
ICANN (2) | 3 |
| 2020 | A Rating Bias Formulation based on Fuzzy Set for RecommendationabstractIn recommender systems, the user uncertain preference results in unexpected ratings. Previous approaches (e.g., BiasMF) only adjust the rating value based on the bias vector, ignoring the uncertainty of rating. This paper makes an initial attempt in integrating the influence of user uncertain degree and user rating bias into the matrix factorization framework, simultaneously. An approach based on fuzzy set, called fuZzy Matrix Factorization (ZMF), is proposed. Specifically, a fuzzy set of like is defined for each user, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Then, the user uncertain preference matrix is obtained, which could explain and represent the user bias and uncertainty effectively. Furthermore, to enhance the computational impact on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Besides, the proposed approach could be extended to others due to independency on additional data sources. Experimental results on three datasets show that ZMF produces an effective improvement. Fuqing Zhu, Jiao Dai, Liangjun Zang, Yipeng Su, Jizhong Han, Songlin Hu 0001 |
IJCNN | 2 |
| 2020 | AutoSUM: Automating Feature Extraction and Multi-user Preference Simulation for Entity Summarization
Dongjun Wei, Fuqing Zhu, Liangjun Zang, Wei Zhou 0019, Songlin Hu 0001 |
PAKDD (2) | 3 |
| 2019 | A Fuzzy Set Based Approach for Rating BiasabstractIn recommender systems, the user uncertain preference results in unexpected ratings. This paper makes an initial attempt in integrating the influence of user uncertain degree into the matrix factorization framework. Specifically, a fuzzy set of like for each user is defined, and the membership function is utilized to measure the degree of an item belonging to the fuzzy set. Furthermore, to enhance the computational effect on sparse matrix, the uncertain preference is formulated as a side-information for fusion. Experimental results on three real-world datasets show that the proposed approach produces stable improvements compared with others. Jiao Dai, Fuqing Zhu, Liangjun Zang, Songlin Hu 0001, Jizhong Han |
AAAI | 3 |
| 2019 | Multi-hop Selector Network for Multi-turn Response Selection in Retrieval-based ChatbotsabstractChunyuan Yuan, Wei Zhou, Mingming Li, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunyuan Yuan, Wei Zhou 0019, Shangwen Lv, Fuqing Zhu, Jizhong Han, Songlin Hu 0001 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | SPL: Exploiting Unlabeled Data for Multi-label Image ClassificationabstractThe utilization of the unlabeled data provides a beneficial attempt for improving the generalization ability of the convolutional neural network (CNN) model, just as what is applied in person re-identification task. Different from that, multi-label image classification aims to predict multiple labels for each given image. The unlabeled data should be properly assigned multiple labels for regularizing the training process of CNN model. To make full use of the unlabeled data, this paper proposes a soft pseudo labeling (SPL) method for multi-label image classification. Specifically, the unlabeled samples are first generated by DCGAN and WGAN-GP. Then, the virtual multiple labels of the generated unlabeled samples are assigned based on an initial confidence value by SoftMax function. Finally, both the generated samples and original training samples are fed into the network as input, in order to learn a CNN model with stronger generalization ability. On three public multi-label image classification datasets (i.e., WIDER-Attribute, NUS-WIDE and MS-COCO), SPL provides a stable improvement over the baseline and produces a competitive performance compared with some existing multi-label image classification methods. Weibo Zhang, Fuqing Zhu, Jiao Dai, Songlin Hu 0001, Jizhong Han, Tao Guo 0006 |
ICME | 2 |
| 2018 | Pseudo-positive regularization for deep person re-identification
Fuqing Zhu, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001 |
Multim. Syst. | 1 |
| 2018 | A novel two-stream saliency image fusion CNN architecture for person re-identification
Fuqing Zhu, Xiangwei Kong 0001, Haiyan Fu, Qi Tian 0001 |
Multim. Syst. | 1 |
| 2018 | A loss combination based deep model for person re-identification
Fuqing Zhu, Xiangwei Kong 0001, Haiyan Fu, Ming Li 0011 |
Multim. Tools Appl. | 1 |
| 2017 | Part-Based Deep Hashing for Large-Scale Person Re-IdentificationabstractLarge-scale is a trend in person re-identi- fication (re-id). It is important that real-time search be performed in a large gallery. While previous methods mostly focus on discriminative learning, this paper makes the attempt in integrating deep learning and hashing into one framework to evaluate the efficiency and accuracy for large-scale person re-id. We integrate spatial information for discriminative visual representation by partitioning the pedestrian image into horizontal parts. Specifically, Part-based Deep Hashing (PDH) is proposed, in which batches of triplet samples are employed as the input of the deep hashing architecture. Each triplet sample contains two pedestrian images (or parts) with the same identity and one pedestrian image (or part) of the different identity. A triplet loss function is employed with a constraint that the Hamming distance of pedestrian images (or parts) with the same identity is smaller than ones with the different identity. In the experiment, we show that the proposed PDH method yields very competitive re-id accuracy on the large-scale Market-1501 and Market-1501+500K datasets. Fuqing Zhu, Xiangwei Kong 0001, Liang Zheng 0001, Haiyan Fu, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |