Jie Guo 0008

dblp:77/2751-8 · DBLP profile ↗
← Back
28ranked-venue papers
14as first author
20since 2021 · last 2026
0000-0003-4975-0315ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 IFDMA With Low-Complexity Bayesian-Optimal Receiver for High-Mobility Massive Connectivity
Yuhao Chi, Lingfei Zhao, Lei Liu 0005, Yao Ge 0001, Shunqi Huang, Jie Guo 0008, Min Sheng
ICC6
2026 Oversampled IFDM: Low-Complexity Detection with Bayes-Optimal Performance
Zheng Shen, Yuhao Chi, Lei Liu 0005, Yao Ge 0001, Jie Guo 0008, Min Sheng
ISIT5
2026 Information entropy-guided knowledge graph recommendation
Jie Guo 0008, Yunfei Zhao 0004, Bin Song 0001
Knowl. Based Syst.1
2026 Adaptive Graph Convolution With Diffusion Models for Multimodal Recommendation
Jie Guo 0008, Bin Song 0001
IEEE Trans. Knowl. Data Eng.1
2026 TextBridge: A Text-Centered Framework for Enhanced Multimodal Integration and Retrieval
abstract
Despite significant advancements in multimodal pre-training, effectively integrating and using latent semantic information across multiple modalities remains a challenge. In this paper, we introduce TextBridge, a text-centered framework that uses the text modality as a semantic anchor to guide cross-modal integration and alignment. TextBridge employs frozen encoders from state-of-the-art pre-trained models and introduces an innovative modality bridge module that enhances semantic alignment and reduces redundancy among different modal features. The framework also incorporates a multi-projection text feature fusion method, enhancing the alignment and integration of text features from diverse modalities into a cohesive semantic representation. To optimize the integration of multimodal information, we make the text encoder trainable and use a text-centered contrastive loss function to enhance the model's ability to capture complementary information across modalities. Extensive experiments on the M5Product dataset demonstrate that TextBridge significantly outperforms the SCALE model in mean average precision (mAP) and precision (Prec), underscoring its effectiveness in multimodal retrieval tasks.
Jie Guo 0008, Haiyang Jing, Bin Song 0001
IEEE Trans. Multim.1
2026 LETTER: Self-Harmonized Representation Learning for Multimodal Recommendation
abstract
Multimodal recommender systems try to integrate multimedia data (images, texts, etc.) with user-item historical records to better model user preference. However, most previous methods largely ignored the underlying fine-grained attribute features of items, which makes it difficult to fully explore users' nuanced attention across individual and combined attributes, resulting in low recommendation performance. To address these issues, this paper proposes a novel and effective self-harmonized representation learning network for multimodal recommendation, named LETTER. LETTER has the ability to effectively optimize the user and item representations for multimodal recommendation. Specifically, we design a factorized attribute interaction module that captures diverse combinations of item latent attributes using a bilinear pooling strategy. Then a dual graph convolution module is established to learn the modality-specific representations from user-item interactive and item semantic relations. Finally, we design a preference self-harmonization module that adaptively identifies the salient influencing factors of user preference, thus refining user and item representations to improve recommendation accuracy. We conduct extensive experiments on three real-world datasets, demonstrating that LETTER outperforms state-of-the-art multimodal recommendation methods.
Jie Guo 0008, Longyu Wen, Yunfei Zhao 0004, Bin Song 0001, Yuhao Chi
IEEE Trans. Multim.1
2025 Video-text retrieval based on multi-grained hierarchical aggregation and semantic similarity optimization
Jie Guo 0008, Shujie Lan, Bin Song 0001, Mengying Wang 0003
Neurocomputing1
2025 Large Language Models and Artificial Intelligence Generated Content Technologies Meet Communication Networks
abstract
Artificial intelligence generated content (AIGC) technologies, with a predominance of large language models (LLMs), have demonstrated remarkable performance improvements in various applications, which have attracted great interests from both academia and industry. Although some noteworthy advancements have been made in this area, a comprehensive exploration of the intricate relationship between AIGC and communication networks remains relatively limited. To address this issue, this article conducts an exhaustive survey from dual standpoints: first, it scrutinizes the integration of LLMs and AIGC technologies within the domain of communication networks and second, it investigates how the communication networks can further bolster the capabilities of LLMs and AIGC. Additionally, this research explores the promising applications along with the challenges encountered during the incorporation of these AI technologies into communication networks. Through these detailed analyses, our work aims to deepen the understanding of how LLMs and AIGC can synergize with and enhance the development of advanced intelligent communication networks, contributing to a more profound comprehension of next-generation intelligent communication networks.
Jie Guo 0008, Meiting Wang, Hang Yin 0007, Bin Song 0001, Yuhao Chi, F. Richard Yu, Chau Yuen
IEEE Internet Things J.1
2025 Multi-Scale Semantic Communication for Object Detection: Single and Cross-Domain Scenarios
abstract
With the rapid popularity of vision-driven communication applications, object detection has become one of the fundamental techniques for performing practical vision tasks. In traditional communication systems, images are compressed for transmission, reconstructed at the receiver, and then processed by existing object detection algorithms. However, transmitting large amounts of images consumes significant storage and communication resources. To address this challenge, a semantic communication-based image reconstruction scheme has been proposed for object detection, which transmits only the semantic information relevant to image reconstruction. However, this method is prone to losing key information, such as object position and texture details, leading to degraded object detection performance. Additionally, it is sensitive to environmental factors such as weather and lighting, resulting in poor adaptability across multiple scenarios. To address these issues, we propose a multi-scale semantic communication framework for object detection that transmits only multi-scale semantic features relevant to the task and employs decoupling at the receiver to separate positional and classification information of target objects without requiring image reconstruction. To improve adaptability across multiple scenarios, we introduce a cross-domain object detection technique that ensures reliable object detection in new scenarios by optimizing the framework’s multi-scale semantic encoder through domain adversarial learning. Numerical results demonstrate that the proposed framework achieves mean average precision improvements of$15.4\% \sim 38.5\%$over the traditional communication framework within low to medium signal-to-noise ratio regions in additive white Gaussian noise and Rayleigh fading channels.
Jie Guo 0008, Hang Yin 0007, Bin Song 0001, Yuhao Chi, Zhaoyang Zhang 0001, Chau Yuen, Dusit Niyato
IEEE Trans. Wirel. Commun.1
2024 Multitask Fine-Grained Feature Mining for Multilabel Remote Sensing Image Classification
abstract
Multilabel remote sensing image classification can provide comprehensive object-level semantic descriptions of remote sensing images. However, most existing methods cannot fully mine the fine-grained features of images and labels, resulting in low classification accuracy. To address this issue, we propose a novel multitask framework for multilabel remote sensing image classification. The framework establishes the class-specific feature extraction as a binary classification auxiliary task to assist the main multilabel classification task, which can improve the model’s local and global feature extraction ability. Meanwhile, the framework updates the label correlation graph using the graph transformer layer to accurately identify label node pairs with potential correlation, which effectively mines the correlation of multiple labels to generate more accurate label co-occurrence embedding for image label prediction. Experimental results on UCM, AID, and DFC15 multilabel datasets show that the proposed method outperforms existing state-of-the-art methods.
Jie Guo 0008, Hao Sun 0033, Jinheng Han, Bin Song 0001, Yuhao Chi, Bingxi Song
IEEE Trans. Geosci. Remote. Sens.1
2024 SPACE: Self-Supervised Dual Preference Enhancing Network for Multimodal Recommendation
abstract
Multimodal recommendation is an emerging task with the goal of improving the effectiveness of the recommendation system by utilizing multimodal data (images, texts, etc.). Most previous methods have struggled with the ability to mine item semantic relationships while guaranteeing accurate modeling of user modality preferences, resulting in low recommendation accuracy. To address this issue, this paper proposes a novel and effective Self-suPervised duAl preference enhanCing nEtwork for multimodal recommendation, named SPACE, which further mines user preferences towards historical interactions and multimodal features of items to obtain more precise user and item representation. Specifically, we design an interaction preference enhancing module to learn both interactive and latent semantic relationships between users and items. Then, a modality preference enhancing module is established by introducing self-supervised learning (SSL), which aims to strengthen the role of dominant modality-specific representation of items. Finally, the enhanced interaction and modality representations are fused, and the recommendation performance is largely improved by utilizing dual joint prediction. Extensive experiments are conducted on three real-world datasets, and the simulation results demonstrate that the proposed SPACE model outperforms the state-of-the-art multimodal recommendation methods.
Jie Guo 0008, Longyu Wen, Bin Song 0001, Yuhao Chi, F. Richard Yu
IEEE Trans. Multim.1
2023 Attention-guided Multi-step Fusion: A Hierarchical Fusion Network for Multimodal Recommendation
abstract
The main idea of multimodal recommendation is the rational utilization of the item's multimodal information to improve the recommendation performance. Previous works directly integrate item multimodal features with item ID embeddings, ignoring the inherent semantic relations contained in the multimodal features. In this paper, we propose a novel and effective aTtention-guided Multi-step FUsion Network for multimodal recommendation, named TMFUN. Specifically, our model first constructs modality feature graph and item feature graph to model the latent item-item semantic structures. Then, we use the attention module to identify inherent connections between user-item interaction data and multimodal data, evaluate the impact of multimodal data on different interactions, and achieve early-step fusion of item features. Furthermore, our model optimizes item representation through the attention-guided multi-step fusion strategy and contrastive learning to improve recommendation performance. The extensive experiments on three real-world datasets show that our model has superior performance compared to the state-of-the-art models.
Jie Guo 0008, Hao Sun 0033, Bin Song 0001, F. Richard Yu
SIGIR2
2023 Interpretability for reliable, efficient, and self-cognitive DNNs: From theories to applications
Xu Kang 0002, Jie Guo 0008, Bin Song 0001, Binghuang Cai, Hongyu Sun 0001, Zhebin Zhang
Neurocomputing2
2023 Black-box attacks on image classification model with advantage actor-critic algorithm in latent space
Xu Kang 0002, Bin Song 0001, Jie Guo 0008, Hao Qin 0001, Xiaojiang Du, Mohsen Guizani
Inf. Sci.3
2023 Inter-Intra Modal Representation Augmentation With Trimodal Collaborative Disentanglement Network for Multimodal Sentiment Analysis
abstract
Recently, Multimodal Sentiment Analysis (MSA) is a challenging research area given its complex nature, and humans express emotional cues across various modalities such as language, facial expressions, and speech. Representation and fusion of features are the most crucial tasks in multimodal sentiment analysis research. However, in the current research, most methods ignore the importance of eliminating potential irrelevant features in the original features of each modality and cross-modal common feature. Moreover, the features extracted from all the modalities contain cluttered background noise and different occlusions noise, which negatively affects feature alignment. Different from these methods, we propose a novel Trimodal Collaborative Disentanglement Network (TCDN) to solve these problems in this paper. This work can obtain effective sentiment results on two aspects: i) Trimodal collaborative uses L1-norm to eliminate irrelevant features and unify the characteristics of the three modals (inter-modal). ii) Disentanglement network introduces an adversary noise by combining the original features of various single modalities and the common representation, alleviating the background noises within each modality (intramodal). This inter-intra modal feature augmentation method is the first work to obtain the common representation by implementing data augmentation as far as we know. Extensive experiments are completed on two benchmark datasets, including MOSI and MOSEI, demonstrating the superiority of the TCDN model over the state-of-the-art methods.
Chen Chen 0128, Hansheng Hong, Jie Guo 0008, Bin Song 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 A VAE-Based User Preference Learning and Transfer Framework for Cross-Domain Recommendation
abstract
The core idea of cross-domain recommendation is to alleviate the problem of data scarcity. Previous methods have made brilliant successes. However, many of them mainly focus on learning an ideal mapping function across-domains, ignoring the user preferences within a specific domain, which leads to suboptimal results. In this paper, we propose a Cross-Domain Recommendation Variational AutoEncoder framework (CDRVAE), a novel extension of a variational autoencoder on cross-domain recommendations for user behaviour distribution modeling. It applies a new hybrid architecture of VAE as the backbone and simultaneously constructs two information flows, within-domain and cross-domain modeling. For the former, an asymmetric codec structure is designed to reconstruct preference distribution from domain-specific latent factors. To relieve the posterior collapse dilemma, a combined prior is employed to increase the distribution complexity. The equivalent transition by a transformation matrix and the unobserved interaction generation by cross-domain reconstruction contribute to the latter. We combine all the above components for the more accurate and reliable user features. Extensive experiments are conducted on three public benchmark datasets to validate the effectiveness of the proposed CDRVAE. Experimental results demonstrate that CDRVAE is consistently superior to other state-of-the-art alternative baseline models.
Tong Zhang 0015, Chen Chen 0128, Dan Wang 0002, Jie Guo 0008, Bin Song 0001
IEEE Trans. Knowl. Data Eng.4
2023 Trust-Aware Multi-Task Knowledge Graph for Recommendation
abstract
Data sparsity and cold start problems are common in recommender systems. Adding some side information, such as knowledge graph and users' trust relationship, is an effective method to alleviate these problems. However, few work jointly explore the fine-grained implicit relationships between the external heterogeneous graphs to enhance the recommendation accuracy. To address this issue, in this paper, we propose a new method named Trust-aware Multi-task Knowledge Graph (TMKG), which uses multi-task learning to integrate two kinds of side information of trust graph and knowledge graph in an end-to-end manner. Firstly, we mine the intra-graph and inter-graph high-order connections through the node propagation and aggregation, and optimize the embedding of nodes through the implicit relationships obtained. Furthermore, through the shared cross unit, the connection relationships between each layer is mined, and the high-order interaction of nodes of different layers is obtained. We conduct extensive experiments on real-world datasets and prove that our model has the superior performance compared with the state-of-the-art models.
Jie Guo 0008, Bin Song 0001, Chen Chen 0128, Jianglong Chang, F. Richard Yu
IEEE Trans. Knowl. Data Eng.2
2023 HGAN: Hierarchical Graph Alignment Network for Image-Text Retrieval
abstract
Image-text retrieval (ITR) is a challenging task in the field of multimodal information processing due to the semantic gap between different modalities. In recent years, researchers have made great progress in exploring the accurate alignment between image and text. However, existing works mainly focus on the fine-grained alignment between image regions and sentence fragments, which ignores the guiding significance of context background information. Actually, integrating the local fine-grained information and global context background information can provide more semantic clues for retrieval. In this paper, we propose a novel Hierarchical Graph Alignment Network (HGAN) for image-text retrieval. First, to capture the comprehensive multimodal features, we construct the feature graphs for the image and text modality respectively. Then, a multi-granularity shared space is established with a designed Multi-granularity Feature Aggregation and Rearrangement (MFAR) module, which enhances the semantic corresponding relations between the local and global information, and obtains more accurate feature representations for the image and text modalities. Finally, the ultimate image and text features are further refined through three-level similarity functions to achieve the hierarchical alignment. To justify the proposed model, we perform extensive experiments on MS-COCO and Flickr30 K datasets. Experimental results show that the proposed HGAN outperforms the state-of-the-art methods on both datasets, which demonstrates the effectiveness and superiority of our model.
Jie Guo 0008, Meiting Wang, Bin Song 0001, Yuhao Chi, Jianglong Chang
IEEE Trans. Multim.1
2022 Task-Oriented Image Transmission for Scene Classification in Unmanned Aerial Systems
abstract
The vigorous developments of the Internet of Things make it possible to extend its computing and storage capabilities to computing tasks in the aerial system with the collaboration of cloud and edge, especially for artificial intelligence (AI) tasks based on deep learning (DL). Collecting a large amount of image/video data, unmanned aerial vehicles (UAVs) can only hand over intelligent analysis tasks to the back-end mobile edge computing (MEC) server due to their limited storage and computing capabilities. How to efficiently transmit the most correlated information for the AI model is a challenging topic. Inspired by task-oriented communication in recent years, we propose a new aerial image transmission paradigm for the scene classification task. A lightweight model is developed on the front-end UAV for semantic block transmission with the perception of images and channel states. To achieve the tradeoff between transmission latency and classification accuracy, deep reinforcement learning (DRL) is applied to explore the semantic blocks which have the greatest contribution to the back-end classifier under various channel states. Experimental results show that the proposed method can significantly improve classification accuracy by more than 4% under the same conditions, compared to other semantic saliency learning methods.
Xu Kang 0002, Bin Song 0001, Jie Guo 0008, Zhijin Qin, F. Richard Yu
IEEE Trans. Commun.3
2021 Dual Attention Transfer in Session-based Recommendation with Multi-dimensional Integration
abstract
Session-based recommendation (SBR) is widely used in e-commerce to predict the anonymous user's next click action according to a short sequence. Many previous studies have shown the potential advantages of applying Graph Neural Networks (GNN) to SBR tasks. However, the existing SBR models using GNN to solve user preference problems are only based on one single dataset to obtain one recommendation model during training. While the single dataset has the problems including the excessive sparse data source and the long-distance relationship of items. Therefore, introducing the dual transfer, which can enrich the data source, to SBR is absolutely necessary. To this end, a new method is proposed in this paper, which is called dual attention transfer based on multi-dimensional integration (DAT-MDI): (i) DAT uses a potential mapping method based on a slot attention mechanism to extract the user's representation information in different sessions between multiple domains. (ii) MDI combines the graph neural network for the graphs (session graph and global graph) and the gate recurrent unit (GRU) for the sequence to learn the item representation in each session. Then the multi-level session representation are combined by a soft-attention mechanism. We do a variety of experiments on four benchmark datasets which have shown that the superiority of the DAT-MDI model over the state-of-the-art methods.
Chen Chen 0128, Jie Guo 0008, Bin Song 0001
SIGIR2
2020 Context-Aware Object Detection for Vehicular Networks Based on Edge-Cloud Cooperation
abstract
Due to high mobility and high dynamic environments, object detection for vehicular networks is one of the most challenging tasks. However, the development of integration techniques, such as software-defined networking (SDN) and network function visualization (NFV), in networking, caching, and computing provides us with new approaches. In this article, we propose a novel context-aware object detection method based on edge-cloud cooperation. Specifically, an object detection model based on deep learning is established in the cloud server. Different from other methods, to further explore the underlying inner spatial features of collected images, the visual objects of images are regarded as nodes and the spatial relations between objects as edges, then a type of message-passing method is employed to update the nodes' features. In the mobile edge computing (MEC) servers, the context information and captured images of the vehicular environments are extracted and then are used to adjust the object detection model from the cloud server. In this way, the cloud server cooperates with the MEC servers to realize context-aware object detection, which improves the adaptation and performance of the detection model under different scenarios. The simulation results also demonstrate that the proposed method is more accurate and faster than the previous methods.
Jie Guo 0008, Bin Song 0001, F. Richard Yu, Xiaojiang Du, Mohsen Guizani
IEEE Internet Things J.1
2020 Super-Sparse On-Off Division Multiple Access: Replacing Repetition With Idling
abstract
A very low-complexity on-off division multiple access (ODMA) scheme is proposed for K-user non-orthogonal multiple access (NOMA) systems. At the transmission side, each user employs the same length-m channel code whose coded bits, after modulation, are sent in a random time-hopping manner. Specifically, m coded bits are randomly scheduled and sent using n time slots with n≫m, i.e., only m slots are used for signal transmission and the other n- m slots are idle. The slot selection, referred to as an on-off pattern, is unique to each user, and it is the only means of user separation. Consequently, at each time slot only a very few users (i.e., 2 or 3) may simultaneously access the channel, leading to a super-sparse access system. Due to the sparse access property, a very low-complexity iterative multi-user decoding method can be implemented on an almost tree-like factor graph. Compared with existing iteratively decodable code division multiple access (CDMA) schemes, such as sparse-CDMA and interleave division multiple access (IDMA), ODMA does not rely on repetition (spreading) or user interleaving. In fact, we show that in using extrinsic information transfer (EXIT) analysis and simulation, idling is more effective than repetition in terms of enhancing the multi-user iterative decoding performance. By replacing repetition with idling, a remarkable multi-user decoding performance gain is achieved and, at the same time, the decoding complexity is significantly reduced.
Guanghui Song, Kui Cai 0001, Yuhao Chi, Jie Guo 0008, Jun Cheng 0001
IEEE Trans. Commun.4
2019 Deep neural network-aided Gaussian message passing detection for ultra-reliable low-latency communications
Jie Guo 0008, Bin Song 0001, Yuhao Chi, Lahiru Jayasinghe, Chau Yuen, Yong Liang Guan 0001, Xiaojiang Du, Mohsen Guizani
Future Gener. Comput. Syst.1
2018 FPAN: Fine-grained and progressive attention localization network for data retrieval
Bin Song 0001, Jie Guo 0008, Yanling Zhang, Xiaojiang Du, Mohsen Guizani
Comput. Networks3
2017 An Unbalanced Data Hybrid-Sampling Algorithm Based on Multi-Information Fusion
abstract
The emergence of big data bringsnewissues and challenges for the data imbalance problem.Therefore, unbalanced data sampling technology has been a hot research topic in the field of big data.However, the existing sampling methods cannot accurately define the harmful and useless samplescontained in the originaldataset. That is, based on the single information of the dataset, a large number of actuallyharmful samples are being used for sampling, which results in a sharp decline in the identifiable performance of the sampled data. In order to overcome the problems caused by only using one kind of information, an unbalanced data hybrid-sampling algorithm based on multi-information fusion(MIFS)is presented in this paper. The MIFS combines the feature information learned by the boostingmodel with the position information of the data to define the sample, and then divides the samples into different subsets by the information contained. According to the definition of samples, the algorithm performs corresponding under-sampling and over-sampling on these subsets. Experiments show that the MIFS method can improve the performance of sampling operations and produce a high F-score and AUC against bothminority and majority classes in the classification of balanced data.
Bin Song 0001, Jie Guo 0008, Xiaojiang Du
GLOBECOM3
2017 Object detection among multimedia big data in the compressive measurement domain under mobile distributed architecture
Jie Guo 0008, Bin Song 0001, F. Richard Yu, Zheng Yan 0002, Laurence T. Yang
Future Gener. Comput. Syst.1
2016 Significance Evaluation of Video Data Over Media Cloud Based on Compressed Sensing
abstract
Given the varying communication environment between the media cloud and users, there is a need to ensure the most significant part of a video will be successfully transmitted. Although there exist some techniques to evaluate the significance of video data in traditional video coding methods, such as H.264, the evaluation algorithms are often simple and inaccurate. This paper presents a novel significance evaluation method for video data based on compressed sensing. Specifically, we propose a method to obtain a trained dictionary directly by using the measurements of the video data, and then keep the sparse components and generate a saliency map. Since the sparse components can reflect the essential parts of videos, we discuss how to analyze the area and distribution of salient regions. At last, we present a computing method that gives the degree of significance of a frame. Experimental results show that the proposed saliency map reflects the focus points of humans. The method can be used in the distribution of video data over “wireless” transmissions and provide good video quality to mobile users.
Jie Guo 0008, Bin Song 0001, Xiaojiang Du
IEEE Trans. Multim.1
2014 Estimation of measurements for block-based compressed video sensing: study of correlation noise in measurement domain
abstract
Compressed video sensing (CVS) is an application of compressed sensing theory which samples a signal below the Shannon–Nyquist rate. However, previous research about CVS has largely ignored the inter‐frame correlation analysis in the measurement domain, and then is not able to remove the time redundancy. In this study, the authors consider the estimation of the measurements of a block in any possible position in a frame by introducing a correlation noise (CN) between the actual and the estimated measurements. In this work, they first establish a correlation model (CM) in the pixel domain between a block which is in a random unknown position in a frame and the adjacent non‐overlapping blocks that they already have. Then, a novel measurement domain CM is presented to approximate the measurements for the random block. Lastly, they employ the CN to characterise the accuracy of the CM in the measurement domain. The simulation results show that the proposed model can make an accurate estimation to the actual measurements of an arbitrary block in a frame and that by using the proposed CN to perform motion estimation, they can improve the peak signal‐to‐noise ratio of the video sequences by 0.1–1.7 dB compared with the existing methods.
Bin Song 0001, Jie Guo 0008, Lingquan Li, Haixiao Liu
IET Image Process.2