Guannan Dong

dblp:182/7334 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Agent Consultation and Uncertainty-guided Voting for text-to-image person retrieval
Mingcheng Ni, Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Guannan Dong, Yong Cheng 0001
Eng. Appl. Artif. Intell.5
2025 Unveiling Local Well-posedness Influence for Cross-modal Person Re-Identification
abstract
The existing cross-modal retrieval methods trend toward the conventional multi-modal alignment while ignoring the localization bias caused by visual hallucination, including color pollution and appearance-like occlusion due to uncontrollable factors such as weather, illumination, and occlusion. This feature blinding misleads the model to lock in the pseudo-real position and further leads to local unmatched. To this end, we discuss cross-modal local alignment well-posedness by making a phased local modal-masking to calibrate the undisturbed actual local alignment from entity, attribute, and appearance. Specifically, we introduce a mask-based local well-posedness modeling (MLWM) strategy, including text-based entity masking (TEM), text-based attribute-specific masking (TAM), and image-based appearance masking (IAM) to phased collaboratively consider image prompting-based text entities, image prompting-based text attributes, and text prompting-based appearance inference contrast, respectively. Finally, we dynamically optimize the weights of positively correlated image-text pairs by comparing the similarity between original and reconstructed features. Experimental results demonstrate that our method is effective on three public datasets.
Guannan Dong, Aichun Zhu, Mingcheng Ni, Yifeng Li 0002
ICASSP2
2025 Style-Texture Collaborative Learning for Face Forgery Detection
Pengcheng Jia, Guannan Dong, Shubo Wang, Aichun Zhu
PRCV (15)3
2024 TVPR: Text-to-Video Person Retrieval and a New Benchmark
abstract
Most existing methods for text-based person retrieval focus on text-to-image person retrieval. Nevertheless, due to the lack of dynamic information provided by isolated frames, the performance is hampered when the person is obscured or variable motion details are missed in isolated frames. To overcome this, we propose a novel Text-to-Video Person Retrieval (TVPR) task. Since there is no dataset or benchmark that describes person videos with natural language, we construct a large-scale cross-modal person video dataset containing detailed natural language annotations, termed as Text-to-Video Person Reidentification (TVPReid) dataset. In this paper, we introduce a Multielement Feature Guided Fragments Learning (MFGF) strategy, which leverages the cross-modal text-video representations to provide strong text-visual and text-motion matching information to tackle uncertain occlusion conflicting and variable motion details. Specifically, we establish two potential cross-modal spaces for text and video feature collaborative learning to progressively reduce the semantic difference between text and video. To evaluate the effectiveness of the proposed MFGF, extensive experiments have been conducted on TVPReid dataset. To the best of our knowledge, MFGF is the first successful attempt to use video for text-based person retrieval task and has achieved state-of-the-art performance on TVPReid dataset. The TVPReid dataset will be publicly available to benefit future research.
Xu Zhang 0075, Fan Ni, Guannan Dong, Aichun Zhu, Mingcheng Ni, Hui Liu 0026
ACM Multimedia3
2024 Hierarchical Discrepancy-Aware Interaction Network for Face Forgery Detection
Pengcheng Jia, Guannan Dong, Aichun Zhu
PRCV (15)2
2024 EESSO: Exploiting Extreme and Smooth Signals via Omni-frequency learning for Text-based Person Retrieval
Jingyi Xue, Zijie Wang 0003, Guannan Dong, Aichun Zhu
Image Vis. Comput.3
2024 Potential source-information dominated learning for composed cross-modal person re-identification
Xiangyun Zhang, Guannan Dong, Kunyu Wu, Yong Cheng 0001, Aichun Zhu
Multim. Tools Appl.3
2022 Temporal Relation Inference Network for Multimodal Speech Emotion Recognition
abstract
Speech emotion recognition (SER) is a non-trivial task for humans, while it remains challenging for automatic SER due to the linguistic complexity and contextual distortion. Notably, previous automatic SER systems always regarded multi-modal information and temporal relations of speech as two independent tasks, ignoring their association. We argue that the valid semantic features and temporal relations of speech are both meaningful event relationships. This paper proposes a novel temporal relation inference network (TRIN) to help tackle multi-modal SER, which fully considers the underlying hierarchy of phonetic structure and its associations between various modalities under the sequential temporal guidance. Mainly, we design a temporal reasoning calibration module to imitate real and abundant contextual conditions. Unlike the previous works, which assume all multiple modalities are related, it infers the dependency relationship between the semantic information from the temporal level and learns to handle the multi-modal interaction sequence with a flexible order. To enhance the feature representation, an innovative temporal attentive fusion unit is developed to magnify the details embedded in a single modality from semantic level. Meanwhile, it aggregates the feature representation from both the temporal and semantic levels to maximize the integrity of feature representation by an adaptive feature fusion mechanism to selectively collect the implicit complementary information to strengthen the dependencies between different information subspaces. Extensive experiments conducted on two benchmark datasets demonstrate the superiority of our TRIN method against some state-of-the-art SER methods.
Guannan Dong, Chi-Man Pun, Zheng Zhang 0006
IEEE Trans. Circuits Syst. Video Technol.1
2021 Deep Collaborative Multi-Modal Learning for Unsupervised Kinship Estimation
abstract
Kinship verification is a long-standing research challenge in computer vision. The visual differences presented to the face have a significant effect on the recognition capabilities of the kinship systems. We argue that aggregating multiple visual knowledge can better describe the characteristics of the subject for precise kinship identification. Typically, the age-invariant features can represent more natural facial details. Such age-related transformations are essential for face recognition due to the biological effects of aging. However, the existing methods mainly focus on employing the single-view image features for kinship identification, while more meaningful visual properties such as race and age are directly ignored in the feature learning step. To this end, we propose a novel deep collaborative multi-modal learning (DCML) to integrate the underlying information presented in facial properties in an adaptive manner to strengthen the facial details for effective unsupervised kinship verification. Specifically, we construct a well-designed adaptive feature fusion mechanism, which can jointly leverage the complementary properties from different visual perspectives to produce composite features and draw greater attention to the most informative components of spatial feature maps. Particularly, an adaptive weighting strategy is developed based on a novel attention mechanism, which can enhance the dependencies between different properties by decreasing the information redundancy in channels in a self-adaptive manner. Moreover, we propose to use self-supervised learning to further explore the intrinsic semantics embedded in raw data and enrich the diversity of samples. As such, we could further improve the representation capabilities of kinship feature learning and mitigate the multiple variations from original visual images. To validate the effectiveness of the proposed method, extensive experimental evaluations conducted on four widely-used datasets show that our DCML method is always superior to some state-of-the-art kinship verification methods.
Guannan Dong, Chi-Man Pun, Zheng Zhang 0006
IEEE Trans. Inf. Forensics Secur.1
2021 Kinship Verification Based on Cross-Generation Feature Interaction Learning
abstract
Kinship verification from facial images has been recognized as an emerging yet challenging technique in many potential computer vision applications. In this paper, we propose a novel cross-generation feature interaction learning (CFIL) framework for robust kinship verification. Particularly, an effective collaborative weighting strategy is constructed to explore the characteristics of cross-generation relations by corporately extracting features of both parents and children image pairs. Specifically, we take parents and children as a whole to extract the expressive local and non-local features. Different from the traditional works measuring similarity by distance, we interpolate the similarity calculations as the interior auxiliary weights into the deep CNN architecture to learn the whole and natural features. These similarity weights not only involve corresponding single points but also excavate the multiple relationships cross points, where local and non-local features are calculated by using these two kinds of distance measurements. Importantly, instead of separately conducting similarity computation and feature extraction, we integrate similarity learning and feature extraction into one unified learning process. The integrated representations deduced from local and non-local features can comprehensively express the informative semantics embedded in images and preserve abundant correlation knowledge from image pairs. Extensive experiments demonstrate the efficiency and superiority of the proposed model compared to some state-of-the-art kinship verification methods.
Guannan Dong, Chi-Man Pun, Zheng Zhang 0006
IEEE Trans. Image Process.1
2018 Optimal Downlink Transmission in Massive MIMO Enabled SWIPT Systems with Zero-Forcing Precoding
abstract
This paper investigates the downlink transmission of massive multiple-input-multiple-output (MIMO) simultaneous wireless information and power transfer (SWIPT) systems. The base station (BS) is equipped with large scale antenna array to provide users with concurrent information and energy supplies. Considering the short communication range between users and the BS in SWIPT systems, the transmission channels are modeled as Rician fading channels to capture both the line-of-sight (LOS) and non-LOS propagations. The approximate and asymptotic expressions of the achievable rate are first derived, and a sum achievable rate optimization problem is formulated based on the asymptotic expression subject to the quality-of-service (QoS) and transmit power constraints. An iterative optimization framework is proposed to solve the original non-linear non-convex problem, and iterative successive convex approximation (SCA) method is introduced in the framework to transmit the non-convex subproblem into convex form. The convergence and effectiveness of the proposed framework are analyzed and proved through analysis and intensive simulations. Results show that the proposed framework can achieve the optimal system performance as the exhaustive search method does.
Guannan Dong, Haixia Zhang 0001, Dongfeng Yuan
GLOBECOM1
2016 Linear Programming Based Pilot Allocation in TDD Massive Multiple-Input Multiple-Output Systems
abstract
By providing substantial gains in terms of both spectral and energy-efficiency, Massive MIMO is expected to be the promising enabler for the fifth generation (5G) communications. However the performance of massive MIMO is greatly affected by pilot contamination due to the insufficiency of pilot sequences. To overcome this, we propose a linear programming based pilot allocation with the purpose of alleviating the effect of pilot contamination and maximizing the system throughput. We first formulate the pilot allocation as a user clustering problem, which can be converted to a linear programming one by introducing the integer factor and constraint relaxation. An efficient linear programming algorithm is proposed to solve the problem. Simulation results demonstrate that the proposed scheme outperforms the other candidates in the presence of pilot contamination.
Guannan Dong, Haixia Zhang 0001, Dongfeng Yuan
VTC Spring1