Yuhui Guo

dblp:159/3872 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0002-4833-7003ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 11 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Semantics to Spectrum: A New Lens on Graph Augmentation Strategy
abstract
Graph augmentation is a cornerstone of effective graph contrastive learning, yet existing methods often rely on random designed perturbations, which may distort latent semantics and impair representation quality. In this work, we argue that semantic consistency can be effectively approximated by low-frequency components in the spectral domain, offering a principled proxy for guiding augmentation. Based on this insight, we propose Frequency-Aware Graph Contrastive Learning (FA-GCL), a novel framework that explicitly preserves low-frequency signals while selectively perturbing high-frequency components. By aligning augmentation with frequency-aware decomposition, FA-GCL generates diverse yet semantically coherent views, mitigating semantic drift and enhancing representational discrimination. Extensive experiments across multiple benchmarks demonstrate that FA-GCL consistently outperforms state-of-the-art baselines with statistically significant gains, validating its exclusive merits.
Xiangping Zheng 0002, Xiuxin Hao, Bo Wu 0026, Wei Li 0109, Yuhui Guo, Xun Liang 0001, Zhiwen Yu 0001
AAAI7
2026 Scalable Multimodal Localization for Underground Parking Lots Using Distributed Antenna Systems
abstract
Achieving accurate and flexible localization in global navigation satellite system (GNSS)-denied underground spaces is critical for Internet of Things (IoT)-enabled logistics and personnel operations. To this end, this paper proposes a scalable multimodal localization framework requiring only a single ultra-wideband (UWB) transmitter connected to a distributed antenna system, avoiding the deployment of multiple additional anchors. An adjacency-masked, change-point-aware hidden Markov model (AC-HMM) is developed for region identification using only two-path UWB measurements, which avoids full channel impulse response (CIR) processing and reduces computational complexity. A multi-scale factor-graph maximum a posteriori inference method (MS-FGM) is then proposed for dynamic localization by fusing UWB and magnetic-field residuals with region and motion constraint factors. Multi-scale temporal aggregation is further introduced to mitigate motion-induced fluctuations and improve localization accuracy and stability. Experiments conducted along the roadways of an underground parking lot demonstrate an average region classification accuracy of 96.81% and a mean positioning error of 0.88 m, outperforming existing methods by up to 42.19%.
Yihong Zheng, Zhaoming Lu, Xinghe Chu, Yinzhe Zhou, Ziwen Luo, Zhiqun Hu, Yuhui Guo
IEEE Internet Things J.7
2026 Bridging the Gap: Seamless Indoor-Outdoor GNSS Positioning for Smartphones in Tunnel Environments via Communication Leaky Coaxial Cables
abstract
To achieve seamless and continuous positioning for smartphones in tunnels, where the Global Navigation Satellite System (GNSS) signals are unavailable, this paper proposes a new method that leverages the existing leaky coaxial cable (LCX) infrastructure from public land mobile networks to introduce the GNSS signals directly from outside the tunnels. This approach introduces three key innovations. First, a continuous hybrid GNSS-LCX channel is proposed which uses a waveguide-to-wireless model to provide continuous GNSS signal coverage via existing 5G LCX without requiring dedicated hardware. Based on this model, a bidirectional clock bias cancellation mechanism is designed for GNSS signals inside the tunnel. This method establishes a quantitative mapping between the position in the tunnel and the pseudorange observations incorporating clock bias, enabling the residual latency from smartphones and wired transmission in the tunnel to be modeled as a function of signal propagation distance and user tracks within the tunnel environment. Furthermore, a 5G-enhanced factor graph optimization (FGO) method is proposed, which integrates 5G measurements and tunnel topology to suppress positioning fluctuations caused by multipath effects in tunnels, while mitigating GNSS positioning latency and ambiguity induced by 5G cell handovers. Field tests in a 150-m tunnel show 1.21 m median accuracy, achieving 46% higher accuracy than fingerprinting and 3.2×faster convergence than conventional methods. This solution reuses 5G-LCX infrastructure for cost-effective and consistent tunnel positioning that bridges the gap between open-sky and underground tunnel environments.
Yinzhe Zhou, Zhaoming Lu, Yihong Zheng, Ziwen Luo, Shuya Zhou, Yuhui Guo, Xinghe Chu
IEEE Internet Things J.6
2023 MVRACE: Multi-view Graph Contrastive Encoding for Graph Neural Network Pre-training
Bo Wu 0026, Xun Liang 0001, Xiangping Zheng 0002, Yuhui Guo, Xuan Zhang 0009
CogSci4
2023 LogLG: Weakly Supervised Log Anomaly Detection via Log-Event Graph Construction
Hongcheng Guo, Yuhui Guo, Jian Yang 0030, Zhoujun Li 0001, Tieqiao Zheng, Liangfan Zheng, Weichao Hou, Bo Zhang 0096
DASFAA (4)2
2023 Modeling High-Order Relation to Explore User Intent with Parallel Collaboration Views
Xiangping Zheng 0002, Xun Liang 0001, Bo Wu 0026, Yuhui Guo, Sensen Zhang, Yuefeng Ma
DASFAA (2)5
2023 Cross-Modal Matching and Adaptive Graph Attention Network for RGB-D Scene Recognition
abstract
Despite the significant advances in RGB-D scene recognition, there are several major limitations that need further investigation. For example, simply extracting modal-specific features neglects the complex relationships among multiple modalities of features. Moreover, cross-modal features have not been considered in most existing methods. To address these concerns, we propose to integrate the tasks of cross-modal matching and modal-specific recognition, termed as Matching-to-Recognition Network (MRNet). Specifically, the cross-modal matching network enhances the descriptive power of the recognition network via a layer-wise semantic loss. The recognition network obtains multi-modal features from a two-stream CNN: global features are obtained by a higher-layer of a CNN to preserve the semantic content, and local layout features are learned by the graph attention network, thus better capturing the key object regions and modelling their relationships. Extensive experiments results demonstrate the MRNet achieves superior performance to state-of-the-art methods, especially for recognition solely based on single modality.
Yuhui Guo, Xun Liang 0001, James T. Kwok, Xiangping Zheng 0002, Bo Wu 0026, Yuefeng Ma
ICASSP1
2023 Intent Does Matter! Propagating High-Order Relations for Exploring Interest Preferences
abstract
Session-based recommendation (SBR) aims to predict the user’s action at the next timestamp according to an anonymous yet short interaction sequence (i.e., session). Almost all the existing SBR solutions for user preference are only based on the current session without exploiting the high-order relations among other sessions, which may restrict the SBR representation ability and even deteriorate the performance. To this end, we propose a Hyper-relation alignment hyperGraph Convolutional Network, called Hyra-GCN, for better inferring the user preference of the current session. Specifically, we first model session-based data as a hyper-graph capable of representing high-order relationships to exploit item transitions over sessions in a more subtle manner. Subsequently, we explore self-supervised learning on item-session hypergraphs, so as to alleviate the problem of data sparsity. Experimental results on real-world datasets demonstrate the effectiveness of our proposed Hyra-GCN against state-of-the-art baselines.
Xiangping Zheng 0002, Xun Liang 0001, Bo Wu 0026, Junlan Feng, Yuhui Guo, Sensen Zhang
ICASSP5
2023 A Multi-scale Interaction Motion Network for Action Recognition Based on Capsule Network
abstract
Recently, action recognition has achieved impressive performance, mainly due to the aid of deep convolutional neural networks and large datasets. Traditionally, most efforts in action recognition have focused on capturing motion information by dense optical flow, but optical flow extraction is very time-consuming. Moreover, prior arts seek to improve accuracy but neglect the part-whole relationship between objects in videos, which may be self-defeating and even deteriorate the performance of methods. To circumvent the above challenges, we present a novel collaborative multipath capsule network (CMCN) for action recognition. In particular, we propose a plug-and-play collaborative multipath block containing spatiotemporal, channel, and motion units, which are complementary and crucial information for action recognition. We exploit the interaction of these three units and selectively emphasize informative spatial-temporal motion to reduce the expensive computational costs. Subsequently, we explore a new capsule voting procedure to reduce the computation used in the capsule dynamic routing mechanism. The critical insight is that the same type of capsules simulates the same entity in different positions, and their voting results should be consistent. This strategy lessens the number of learning parameters that backward pass in the training process, and thus strengthens part-whole relationships in a video. Extensive experiments on multiple real-world datasets for action recognition demonstrate that our model significantly outperforms state-of-the-art models.
Xiangping Zheng 0002, Xun Liang 0001, Bo Wu 0026, Yuhui Guo, Xuan Zhang 0009, Yuefeng Ma
SDM5
2023 Dual-aware Domain Mining and Cross-aware Supervision for Weakly-supervised Semantic Segmentation
abstract
Weakly Supervised Semantic Segmentation with image-level annotation uses localization maps from the classifier to generate pseudo labels. However, such localization maps focus only on sparse salient object regions, it is difficult to generate high-quality segmentation labels, which deviates from the requirement of semantic segmentation. To address this issue, we propose a dual-aware domain mining and cross-aware supervision (DDMCAS) method for weakly-supervised semantic segmentation. Specifically, we propose a dual-aware domain mining (DDM) module consisting of graph-based global reasoning unit and salient-region extension controller, which produces dense localization maps by exploring object features in salient regions and adjacent non-salient regions simultaneously. In order to further bridge the gap between salient regions and adjacent non-salient regions to generate more refined localization maps, we propose a cross-aware supervision (CAS) strategy to recover missing parts of the target objects and enhance weak attention in adjacent non-salient regions, leading to pseudo labels of higher quality for training the segmentation network. Based on the generated pseudo-labels, extensive experiments on PASCAL VOC 2012 dataset demonstrate that our method outperforms state-of-the-art methods using image-level labels for weakly supervised semantic segmentation.
Yuhui Guo, Xun Liang 0001, Bo Wu 0026, Xiangping Zheng 0002, Xuan Zhang 0009
ACM Trans. Knowl. Discov. Data1
2023 Diffuse and Smooth: Beyond Truncated Receptive Field for Scalable and Adaptive Graph Representation Learning
abstract
As the scope of receptive field and the depth of Graph Neural Networks (GNNs) are two completely orthogonal aspects for graph learning, existing GNNs often have shallow layers with truncated-receptive field and far from achieving satisfactory performance. In this article, we follow the idea of decoupling graph convolution into propagation and transformation processes, which generates representations over a sequence of increasingly larger neighborhoods. Though this manner can enlarge the receptive field, it has two critical problems unsolved: how to find the suitable receptive field to avoid under-smoothing or over-smoothing? and how to balance different diffusion operators for better capturing the local and global dependencies? We tackle these challenges and propose a S calable, A daptive G raph C onvolutional N etworks ( SAGCN ) with Transformer architecture. Concretely, we propose a novel non-heuristic metric method that quickly finds the suitable number of diffusing iterations and produces smoothed local embeddings that enable the truncated receptive field to become scalable and independent of prior experience. Furthermore, we devise smooth2seq and diffusion-based position schemes introduced into Transformer architecture for better capturing local and global information among embeddings. Experimental results show that SAGCN enjoys high accuracy, scalability and efficiency on various open benchmarks and is competitive with other state-of-the-art competitors.
Xun Liang 0001, Yuhui Guo, Xiangping Zheng 0002, Bo Wu 0026, Sensen Zhang, Zhiying Li 0004
ACM Trans. Knowl. Discov. Data3
2022 Eureka: Neural Insight Learning for Knowledge Graph Reasoning
abstract
The human recognition system has presented the remarkable ability to effortlessly learn novel knowledge from only a few trigger events based on prior knowledge, which is called insight learning. Mimicking such behavior on Knowledge Graph Reasoning (KGR) is an interesting and challenging research problem with many practical applications. Simultaneously, existing works, such as knowledge embedding and few-shot learning models, have been limited to conducting KGR in either “seen-to-seen” or “unseen-to-unseen” scenarios. To this end, we propose a neural insight learning framework named Eureka to bridge the “seen” to “unseen” gap. Eureka is empowered to learn the seen relations with sufficient training triples while providing the flexibility of learning unseen relations given only one trigger without sacrificing its performance on seen relations. Eureka meets our expectation of the model to acquire seen and unseen relations at no extra cost, and eliminate the need to retrain when encountering emerging unseen relations. Experimental results on two real-world datasets demonstrate that the proposed framework also outperforms various state-of-the-art baselines on datasets of both seen and unseen relations.
Xuan Zhang 0009, Xun Liang 0001, Bo Wu 0026, Xiangping Zheng 0002, Sensen Zhang, Yuhui Guo, Xinyao Liu
COLING6
2022 Graph Fine-Grained Contrastive Representation Learning
abstract
Existing graph contrastive methods have benefited from ingenious data augmantations and mutual information estimation operations that are carefully designated to augment graph views and maximize the agreement between representations produced at the aftermost layer of two view networks. However, the design of graph CL schemes is coarse-grained and difficult to capture the universal and intrinsic properties across intermediate layers. To address this problem, we propose a novel fine-grained graph contrastive learning model (FGCL), which decomposes graph CL into global-to-local levels and disentangles the two graph views into hierarchical graphs by pooling operation to capture both global and local dependencies across views and across layers. To prevent layers mismatch and automatically assign proper hierarchical representations of the augmented graph (Key view) for each pooling layer of the original graph (Query view), we propose a sematic-aware layer allocation strategy to integrate positive guidance from diverse representations rather than a fixed layer manually. Experimental results demonstrate the advantages of our model on graph classification task. This suggests that the proposed fine-grained graph CL presents great potential for graph representation learning.
Xun Liang 0001, Yuhui Guo, Xiangping Zheng 0002, Bo Wu 0026
ICASSP3
2022 Improving Dynamic Graph Convolutional Network with Fine-Grained Attention Mechanism
abstract
Graph convolutional network (GCN) is a novel framework that utilizes a pre-defined Laplacian matrix to learn graph data effectively. With its powerful nonlinear fitting ability, GCN can produce high-quality node embedding. However, generalized GCN can only handle static graphs, whereas a large number of graphs are dynamic and evolve over time, which limits the application field of GCN. Facing the challenge, GCN with recurrent neural network (e.g., RNN) is naturally combined to acquire dynamic graph changes through joint training. However, these methods must use the node information during the entire timeline and ignore two subtle factors: the influence of nodes change with time and are related to the frequency of events. Therefore, we propose a stable and scalable dynamic GCN method using a fine-grained attention mechanism named FADGC. We use GCN to obtain static node vectors at each timestep and integrate node influence factors with multi-head attention for graph time-series learning. Experiments on multiple datasets show that our approach can better capture the inherent special characteristics of different dynamic graphs and achieve higher performance compared with related approaches.
Bo Wu 0026, Xun Liang 0001, Xiangping Zheng 0002, Yuhui Guo
ICASSP4
2022 Adaptive Attention Graph Capsule Network
abstract
From the perspective of the spatial domain, Graph Convolutional Network (GCN) is essentially a process of iteratively aggregating neighbor nodes. However, the existing GCNs using simple average or sum aggregation may neglect the characteristics of each node and the topology between nodes, resulting in a large amount of early-stage information lost during the graph convolution step. To tackle the above challenge, we innovatively propose an adaptive attention graph capsule network, named AA-GCN, for graph classification. We explore various propagation mechanisms of graphs and present an attention mechanism combined with graph propagation and capsules to generate capsule nodes, preserving the spatial topology between nodes. We also propose a graph adaptive attention mechanism to investigate the context information in different global GCN layers, so as to effectively improve the next dynamic routing connection and the final graph classification. Experiments show that our proposed algorithm achieves either state-of-the-art or competitive results across all the datasets.
Xiangping Zheng 0002, Xun Liang 0001, Bo Wu 0026, Yuhui Guo
ICASSP4
2022 CoNet: Co-Embedding by Reinforcing Graph Feature and Topology Information
abstract
Sparsity and smoothness are two main factors that affect the performance of Graph Convolutional Networks (GCNs). Sparsity ensures that models have the first-class generalization ability, while smoothness benefits to reduce noise and make edges reliable. As real-world graphs are often incom-plete and noisy, most GCNs learn node embeddings only acting them as ground-truth information, which unavoidably lead to suboptimal solutions. This paper proposes a co-embedding network (CoNet), jointly learns embeddings by fusing the global and local dependencies to capture the uni-versal and intrinsic properties. We proposed NodeNet and EdgeNet modules, which aggregate global node information and refine local topology structure respectively. Moreover, we further introduce two piplines of variational auto-encoders to fuse the intermediate latent variables of each module to ob-tain co-embeddings via Knowledge Distillation strategy. Ex-tensive experiments on multiple benchmarks show that our proposed approach achieves better performance than existing methods on the graph node classification task.
Xun Liang 0001, Yuhui Guo, Bo Wu 0026, Xiangping Zheng 0002
ICME3
2022 Cross-Pixel Dependency with Boundary-Feature Transformation for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation with image-level labels is a challenging problem that typically relies on the initial responses generated by the classification network to locate object regions. However, such initial responses only cover the most discriminative parts of the object and may incorrectly activate in the background regions. To address this problem, we propose a Cross-pixel Dependency with Boundary-feature Transformation (CDBT) method for weakly supervised semantic segmentation. Specifically, we develop a boundary-feature transformation mechanism, to build strong connections among pixels belonging to the same object but weak connections among different objects. Moreover, we design a cross-pixel dependency module to enhance the initial responses, which exploits context appearance information and refines the prediction of current pixels by the relations of global channel pixels, thus generating pseudo labels of higher quality for training the semantic segmentation network. Extensive experiments on the PASCAL VOC 2012 segmentation benchmark demonstrate that our method outperforms state-of-the-art methods using image-level labels as weak supervision.
Yuhui Guo, Xun Liang 0001, Bo Wu 0026, Xiangping Zheng 0002
ICMR1
2022 When True Becomes False: Few-Shot Link Prediction beyond Binary Relations through Mining False Positive Entities
abstract
Recently, the link prediction task on Hyper-relational Knowledge Graphs (HKGs) has been a hot spot, which aims to predict new facts beyond binary relations. Although previous models have accomplished considerable achievements, there remain three challenges: i) the previous models neglect the existence of False Positive Entities (FPEs), which are true entities in the binary triples, yet becomes false when encountering the query statements of HKGs; ii) Due to the sparse interactions, the models are not capable of coping with long-tail hyper-relations, which are ubiquitous in the real-world; iii) The models are generally transductive learning processes, and have difficulty in adapting new hyper-relations. To tackle the above issues, we firstly propose the task of few-shot link prediction on HKGs and devise hyper-relation-aware attention networks with a contrastive loss, which are empowered to encode all entities including FPEs effectively and increase the distance between the true entities and FPEs through contrastive learning. With few-shot references available, the proposed model then learns the representations of their long-tail hyper-relations and predicts new links by calculating the likelihood between queries and references. Furthermore, our model is inductive and can be scalable to any new hyper-relation effortlessly. Since it is the first trial on few-shot link prediction for HKGs, we also modify the existing few-shot learning approaches on binary relational data to work with HKGs as baselines. Experimental results on three real-world datasets show the superiority of our model over various state-of-the-art baselines.
Xuan Zhang 0009, Xun Liang 0001, Xiangping Zheng 0002, Bo Wu 0026, Yuhui Guo
ACM Multimedia5
2022 Charge Own Job: Saliency Map and Visual Word Encoder for Image-Level Semantic Segmentation
Yuhui Guo, Xun Liang 0001, Xiangping Zheng 0002, Bo Wu 0026, Xuan Zhang 0009
ECML/PKDD (3)1
2022 MULTIFORM: Few-Shot Knowledge Graph Completion via Multi-modal Contexts
Xuan Zhang 0009, Xun Liang 0001, Xiangping Zheng 0002, Bo Wu 0026, Yuhui Guo
ECML/PKDD (2)5
2022 Graph Capsule Network with a Dual Adaptive Mechanism
abstract
While Graph Convolutional Networks (GCNs) have been extended to various fields of artificial intelligence with their powerful representation capabilities, recent studies have revealed that their ability to capture the part-whole structure of the graph is limited. Furthermore, though many GCNs variants have been proposed and obtained state-of-the-art results, they face the situation that much early information may be lost during the graph convolution step. To this end, we innovatively present an Graph Capsule Network with a Dual Adaptive Mechanism (DA-GCN) to tackle the above challenges. Specifically, this powerful mechanism is a dual-adaptive mechanism to capture the part-whole structure of the graph. One is an adaptive node interaction module to explore the potential relationship between interactive nodes. The other is an adaptive attention-based graph dynamic routing to select appropriate graph capsules, so that only favorable graph capsules are gathered and redundant graph capsules are restrained for better capturing the whole structure between graphs. Experiments demonstrate that our proposed algorithm has achieved the most advanced or competitive results on all datasets.
Xiangping Zheng 0002, Xun Liang 0001, Bo Wu 0026, Yuhui Guo, Xuan Zhang 0009
SIGIR4
2021 RGB-D Scene Recognition based on Object-Scene Relation (Student Abstract)
abstract
We develop a RGB-D scene recognition model based on object-scene relation(RSBR). First learning a Semantic Network in the semantic domain that classifies the label of a scene on the basis of the labels of all object types. Then, we design an Appearance Network in the appearance domain that recognizes the scene according to local captions. We enforce the Semantic Network to guide the Appearance Network in the learning procedure. Based on the proposed RSBR model, we obtain the state-of-the-art results of RGB-D scene recognition on SUN RGB-D and NYUD2 datasets.
Yuhui Guo, Xun Liang 0001
AAAI1
2021 Graph Ensemble Networks for Semi-supervised Embedding Learning
Xun Liang 0001, Bo Wu 0026, Zhenyu Guan 0003, Yuhui Guo, Xiangping Zheng 0002
KSEM5
2021 RGB-D Scene Recognition based on Object-Scene Relation and Semantics-Preserving Attention
abstract
Scene recognition is challenging due to intra-class diversity and inter-class similarity. Previous works recognize scenes either with global representations or with intermediate representations of objects. By contrast, we investigate more discriminative sequential representation of object-to-scene relations (SOSRs) for scene recognition. Particularly, we develop an Attention-Preserving Memory-Learning (APML) model, which enforces the Memory Network of the semantic domain to guide the Learning Network of the appearance domain in the learning procedure. Accordingly, we allocate semantics-preserving attention to different objects, which is more effective to seek the key encoded SOSR and discard the misleading encoded SOSR between objects and scene without requiring extra labeled data. Based on the proposed APML networks, we obtain the state-of-the-art results of RGB-D scene recognition on SUN RGB-D and NYUD2 datasets.
Yuhui Guo, Xun Liang 0001
ICMR1
2020 Software trustworthiness evaluation model based on a behaviour trajectory matrix
Yuhui Guo
Inf. Softw. Technol.2
2014 Fast intra partition algorithm for HEVC screen content coding
abstract
Since the publication of the High Efficiency Video Coding standard as the newest video coding standard, several extensions have been made. Among these, the use of the screen content coding in many fields is one of the important extensions. In terms of coding tree unit (CTU) partitioning, rate distortion optimization is still used in screen content coding. The complexity of the process has resulted in problems in relation to real-time application. Thus, this paper proposes a fast-deciding CTU partition mode algorithm based on entropy and coding bits. Experimental results show that the proposed algorithm can save 32% of encoding time on average compared with the default algorithm in HM-12.1+RExt-5.1 with only 0.8% bit rate increment in coding performance.
Mengmeng Zhang 0008, Yuhui Guo, Huihui Bai 0001
VCIP2