Weiyu Guo

dblp:93/1549 · DBLP profile ↗
← Back
29ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 9 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
abstract
Source-Free Object Detection (SFOD) aims to adapt a source-pretrained object detector to a target domain without access to source data. However, existing SFOD methods predominantly rely on internal knowledge from the source model, which limits their capacity to generalize across domains and often results in biased pseudo-labels, thereby hindering both transferability and discriminability. In contrast, Vision Foundation Models (VFMs), pretrained on massive and diverse data, exhibit strong perception capabilities and broad generalization, yet their potential remains largely untapped in the SFOD setting. In this paper, we propose a novel SFOD framework that leverages VFMs as external knowledge sources to jointly enhance feature alignment and label quality. Specifically, we design three VFM-based modules: (1) Patch-weighted Global Feature Alignment (PGFA) distills global features from VFMs using patch-similarity–based weighting to enhance global feature transferability; (2) Prototype-based Instance Feature Alignment (PIFA) performs instance-level contrastive learning guided by momentum-updated VFM prototypes; and (3) Dual-source Enhanced Pseudo-label Fusion (DEPF) fuses predictions from detection VFMs and teacher models via an entropy-aware strategy to yield more reliable supervision. Extensive experiments on six benchmarks demonstrate that our method achieves state-of-the-art SFOD performance, validating the effectiveness of integrating VFMs to simultaneously improve transferability and discriminability.
Huizai Yao, Sicheng Zhao, Pengteng Li, Shuo Lu, Weiyu Guo, Yunfan Lu, Yijie Xu, Hui Xiong 0001
AAAI6
2026 Adapting Graph Models via Target Integrity Assessment and Source Distribution Hypothesis
abstract
This article addresses the challenge of domain adaptation on graphs, a specialized form of graph transfer learning (GTL), which involves adapting a graph model trained on source graphs to unlabeled target graphs that significantly differ in distribution. Traditional methods often rely heavily on the source graph to transfer learned task knowledge, but certain situations may render the source graph unavailable or restricted due to privacy or security concerns, thus impeding the usability and flexibility of graph model adaptation. Therefore, this article studies the problem of source-free domain adaptation (SFDA) in graph transfer learning (GTL). Our objective is to adapt a pretrained model to effectively operate on the target graph without the need to access the source graph. To achieve this, we first incorporate a weighted information maximization loss to enhance the model's discriminative ability on the target graph, where we introduce the concept of posterior integrities of target nodes to assess their optimization confidence. Then, we estimate the distributions of the source graph and generate synthesized source nodes. We propose a reconstruction decoder to enhance the authenticity of the synthesized nodes and use adversarial learning to align the distributions between graphs, leading to improved adaptation of the model. Finally, extensive experimental results on a range of publicly accessible datasets demonstrate the superior performance of our method over the state of the art.
Ziyue Qiao, Xiaomin Yu, Weiyu Guo, Xiao Luo 0001, Meng Xiao 0001, Hui Xiong 0001
IEEE Trans. Neural Networks Learn. Syst.3
2025 BrainCognizer: Brain Decoding with Human Visual Cognition Simulation for fMRI-to-Image Reconstruction
abstract
Brain decoding is a key neuroscience field that reconstructs the visual stimuli from brain activity with fMRI, which helps illuminate how the brain represents the world. fMRI-to-image reconstruction has achieved impressive progress by leveraging diffusion models. However, brain signals infused with prior knowledge and associations exhibit a significant information asymmetry when compared to raw visual features, still posing challenges for decoding fMRI representations under the supervision of images. Consequently, the reconstructed images often lack fine-grained visual fidelity, such as missing attributes and distorted spatial relationships. To tackle this challenge, we propose BrainCognizer, a novel brain decoding model inspired by human visual cognition, which explores multilevel semantics and correlations without fine-tuning of generative models. Specifically, BrainCognizer introduces two modules: the Cognitive Integration Module which incorporates prior human knowledge to extract hierarchical region semantics; and the Cognitive Correlation Module which captures contextual semantic relationships across regions. Incorporating these two modules enhances intra-region semantic consistency and maintains interregion contextual associations, thereby facilitating fine-grained brain decoding. Moreover, we quantitatively interpret our components from a neuroscience perspective and analyze the associations between different visual patterns and brain functions. Extensive quantitative and qualitative experiments demonstrate that BrainCognizer outperforms state-of-the-art approaches on multiple evaluation metrics. Our code is released publicly at https://github.com/Grace160/BrainCognizer.
Guoying Sun, Weiyu Guo, Tong Shao, Yang Yang 0002, Haijin Zeng, Jingyong Su
BIBM2
2025 Revisiting Noise Resilience Strategies in Gesture Recognition: Short-Term Enhancement in sEMG Analysis
abstract
Gesture recognition based on surface electromyography (sEMG) has been gaining importance in many 3D Interactive Scenes. However, sEMG is easily influenced by various forms of noise in real-world environments, leading to challenges in providing long-term stable interactions through sEMG. Existing methods often struggle to enhance model noise resilience through various predefined data augmentation techniques. In this work, we revisit the problem from a short-term enhancement perspective to improve precision and robustness against various common noisy scenarios with learnable denoise using sEMG intrinsic pattern information and sliding-window attention. We propose a Short Term Enhancement Module(STEM), which can be easily integrated with various models. STEM offers several benefits: 1) Noise-resistant, enhanced robustness against noise without manual data augmentation; 2) Adaptability, adaptable to various models; and 3) Inference efficiency, achieving short-term enhancement through minimal weight-sharing in an efficient attention mechanism. In particular, we incorporate STEM into a transformer, creating the Short-Term Enhanced Transformer (STET). Compared with best-competing approaches, the impact of noise on STET is reduced by more than 20\%. We report promising results on classification and regression tasks and demonstrate that STEM generalizes across different gesture recognition tasks. The code is available at https://anonymous.4open.science/r/short_term_semg.
Weiyu Guo, Ziyue Qiao, Ying Sun 0006, Yijie Xu, Hui Xiong 0001
ICML1
2025 Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
abstract
Understanding long video content is a complex endeavor that often relies on densely sampled frame captions or end-to-end feature selectors, yet these techniques commonly overlook the logical relationships between textual queries and visual elements. In practice, computational constraints necessitate coarse frame subsampling, a challenge analogous to “finding a needle in a haystack.” To address this issue, we introduce a semantics-driven search framework that reformulates keyframe selection under the paradigm of Visual Semantic-Logical Search (VSLS). Specifically, we systematically define four fundamental logical dependencies: 1) spatial co-occurrence, 2) temporal proximity, 3) attribute dependency, and 4) causal order. These relations dynamically update frame sampling distributions through an iterative refinement process, enabling context-aware identification of semantically critical frames tailored to specific query requirements. Our method establishes new state-of-the-art performance on the manually annotated benchmark in keyframe selection metrics. Furthermore, when applied to downstream video question-answering tasks, the proposed approach demonstrates the best performance gains over existing methods on LongVideoBench and Video-MME, validating its effectiveness in bridging the logical gap between textual queries and visual-temporal reasoning. The code will be publicly available.
Weiyu Guo, Shaoguang Wang, JianXiang He, Yijie Xu, Jinhui Ye, Ying Sun 0006, Hui Xiong 0001
NeurIPS1
2025 See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
abstract
We introduce See&Trek, the first training-free prompting framework tailored to enhance the spatial understanding of Multimodal Large Language Models (MLLMs) under vision-only constraints. While prior efforts have incorporated modalities like depth or point clouds to improve spatial reasoning, purely visual-spatial understanding remains underexplored. See&Trek addresses this gap by focusing on two core principles: increasing visual diversity and motion reconstruction. For visual diversity, we conduct Maximum Semantic Richness Sampling, which employs an off-the-shell perception model to extract semantically rich keyframes that capture scene structure. For motion reconstruction, we simulate visual trajectories and encode relative spatial positions into keyframes to preserve both spatial relations and temporal coherence. Our method is training&GPU-free, requiring only a single forward pass, and can be seamlessly integrated into existing MLLMs. Extensive experiments on the VSI-Bench and STI-Bench show that See&Trek consistently boosts various MLLMs performance across diverse spatial reasoning tasks with the most +3.5% improvement, offering a promising path toward stronger spatial intelligence.
Pengteng Li, Pinhao Song, Wuyang Li, Huizai Yao, Weiyu Guo, Yijie Xu, Dugang Liu, Hui Xiong 0001
NeurIPS5
2025 Corporate Carbon Emission Prediction: Combining Structured and Unstructured Data
Jiaguan Shen, Weiyu Guo
PAKDD (7)2
2024 SpGesture: Source-Free Domain-adaptive sEMG-based Gesture Recognition with Jaccard Attentive Spiking Neural Network
abstract
Surface electromyography (sEMG) based gesture recognition offers a natural and intuitive interaction modality for wearable devices. Despite significant advancements in sEMG-based gesture recognition models, existing methods often suffer from high computational latency and increased energy consumption. Additionally, the inherent instability of sEMG signals, combined with their sensitivity to distribution shifts in real-world settings, compromises model robustness. To tackle these challenges, we propose a novel SpGesture framework based on Spiking Neural Networks, which possesses several unique merits compared with existing methods: (1) Robustness: By utilizing membrane potential as a memory list, we pioneer the introduction of Source-Free Domain Adaptation into SNN for the first time. This enables SpGesture to mitigate the accuracy degradation caused by distribution shifts. (2) High Accuracy: With a novel Spiking Jaccard Attention, SpGesture enhances the SNNs' ability to represent sEMG features, leading to a notable rise in system accuracy. To validate SpGesture's performance, we collected a new sEMG gesture dataset which has different forearm postures, where SpGesture achieved the highest accuracy among the baselines ($89.26\%$). Moreover, the actual deployment on the CPU demonstrated a latency below 100ms, well within real-time requirements. This impressive performance showcases SpGesture's potential to enhance the applicability of sEMG in real-world scenarios. The code is available at https://github.com/guoweiyu/SpGesture/.
Weiyu Guo, Ying Sun 0006, Yijie Xu, Ziyue Qiao, Yongkui Yang, Hui Xiong 0001
NeurIPS1
2024 A hybrid approach combining data envelopment analysis and recurrent neural network for predicting the efficiency of research institutions
Gaomin Zhang, Weiyu Guo, Zhongcheng Guan
Expert Syst. Appl.2
2024 CrowdTrans: Learning top-down visual perception for crowd counting by transformer
Weiyu Guo, Shaopeng Yang, Yuheng Ren, Yongzhen Huang
Neurocomputing1
2024 Information filtering and interpolating for semi-supervised graph domain adaptation
abstract
Graph domain adaptation, which falls under the umbrella of graph transfer learning, involves transferring knowledge from a labeled source graph to improve prediction accuracy on an unlabeled target graph, where both graphs have identical label spaces but exhibit distribution discrepancies due to temporal data shifts or distinct data collection methods. This adaptation is complicated by the challenges of graph-specific domain discrepancies and cross-graph label scarcity. This paper proposes a semi-supervised G raph domain adaptation method via I nformation F iltering and I nterpolating (GIFI). Specifically, GIFI utilizes a parameterized graph reduction module and a variational information bottleneck to adequately filter out irrelevant information from the source and target graphs to eliminate distribution discrepancy. GIFI also introduces an interpolation-enhanced pseudo-labeling strategy for cross-graph semi-supervised learning, which can mitigate model over-fitting on domain-specific features and limited labeled nodes, thus improving the model’s adaptation and discriminative capability. Experimental results on various graph domain adaptation benchmarks demonstrate GIFI’s superior performance over state-of-the-art methods. Our code is available at https://github.com/joe817/GIFI .
Ziyue Qiao, Meng Xiao 0001, Weiyu Guo, Xiao Luo 0001, Hui Xiong 0001
Pattern Recognit.3
2023 Wavelet-SVDD: Anomaly Detection and Segmentation with Frequency Domain Attention
Linhui Zhou, Weiyu Guo
ADMA (2)2
2023 Interactive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement Learning
abstract
Personalized interior decoration design often incurs high labor costs. Recent efforts in developing intelligent interior design systems have focused on generating textual requirement-based decoration designs while neglecting the problem of how to mine homeowner's hidden preferences and choose the proper initial design. To fill this gap, we propose an Interactive Interior Design Recommendation System (IIDRS) based on reinforcement learning (RL). IIDRS aims to find an ideal plan by interacting with the user, who provides feedback on the gap between the recommended plan and their ideal one. To improve decision-making efficiency and effectiveness in large decoration spaces, we propose a Decoration Recommendation Coarse-to-Fine Policy Network (DecorRCFN). Additionally, to enhance generalization in online scenarios, we propose an object-aware feedback generation method that augments model training with diversified and dynamic textual feedback. Extensive experiments on a real-world dataset demonstrate our method outperforms traditional methods by a large margin in terms of recommendation accuracy. Further user studies demonstrate that our method reaches higher real-world user satisfaction than baseline methods.
He Zhang 0030, Ying Sun 0006, Weiyu Guo, Haonan Lu, Xiaodong Lin 0004, Hui Xiong 0001
ACM Multimedia3
2023 Explainable enterprise credit rating using deep feature crossing
Weiyu Guo, Zhijiang Yang
Expert Syst. Appl.1
2023 Multi-Attention Feature Fusion Network for Accurate Estimation of Finger Kinematics From Surface Electromyographic Signals
abstract
Simultaneous and proportional control (SPC) based on surface electromyographic (sEMG) signals has led to a broad range of applications. However, due to the limitation in the generalization and stability of current machine learning algorithms, these methods can only estimate less than 15 simultaneuous and proportional (SP) categories of finger movement. In this article, a novel deep learning algorithm, named multiattention feature fusion network (MAFN), is proposed to estimate comprehensive finger movement (up to 28 categories SP movements) from sEMG signals. MAFN is based on the multihead attention mechanism, which adaptively extracts essential features for analyzing the joint angles from the extracted sEMG features. Furthermore, a real-time exponential smoothing algorithm is designed for further improvement of the prediction stability. MAFN was evaluated on 28 finger movements of 38 subjects in the Ninapro_db2 dataset, and benchmarked with the state-of-the-art methods, such as temporal convolutional network (TCN) and long short term memory network (LSTM). The results demonstrated that the average Pearson correlation coefficient, root mean squared error of MAFN (0.84 ± 0.03,0.09 ± 0.01) were significantly higher than those of TCN (0.52 ± 0.06,pppp< 0.001). These improvements led to more stable and accurate movement predictions. Additionally, the time delay and power consumption of MAFN when applied to sEMG signals on a portable device are only 83.4 ms and 3 W, which implies prospective commercial applications.
Weiyu Guo, Ning Jiang 0001, Dario Farina, Jingyong Su, Zheng Wang 0027, Chuang Lin 0001, Hui Xiong 0001
IEEE Trans. Hum. Mach. Syst.1
2022 Context-Enhanced Stereo Transformer
Weiyu Guo, Zhaoshuo Li, Yongkui Yang, Zheng Wang 0027, Russell H. Taylor, Mathias Unberath, Alan L. Yuille, Yingwei Li 0002
ECCV (32)1
2022 CrowdFormer: An Overlap Patching Vision Transformer for Top-Down Crowd Counting
abstract
Crowd counting methods typically predict a density map as an intermediate representation of counting, and achieve good performance. However, due to the perspective phenomenon, there is a scale variation in real scenes, which causes the density map-based methods suffer from a severe scene generalization problem because only a limited number of scales are fitted in density map prediction and generation. To address this issue, we propose a novel vision transformer network, i.e., CrowdFormer, and a density kernels fusion framework for more accurate density map estimation and generation, respectively. Thereafter, we incorporate these two innovations into an adaptive learning system, which can take both the annotation dot map and original image as input, and jointly learns the density map estimator and generator within an end-to-end framework. The experimental results demonstrate that the proposed model achieves the state-of-the-art in the terms of MAE and MSE (e.g., it achieved a MAE of 67.1 and MSE of 301.6 on NWPU-Crowd dataset.), and confirm the effectiveness of the proposed two designs. The code is https://github.com/special-yang/Top_Down-CrowdCounting.
Shaopeng Yang, Weiyu Guo, Yuheng Ren
IJCAI2
2022 Efficient convolutional networks learning through irregular convolutional kernels
Weiyu Guo, Jiabin Ma, Yidong Ouyang, Liang Wang 0001, Yongzhen Huang
Neurocomputing1
2021 CNN-DMA: A Predictable and Scalable Direct Memory Access Engine for Convolutional Neural Network with Sliding-window Filtering
abstract
Memory bandwidth utilization has become the key performance bottleneck for state-of-the-art variants of neural network kernels. Current structures such as depth-wise, point-wise and atrous convolutions have already introduced diverse and discontinuous memory access patterns, which impact efficient activation supply due to more frequent cache misses and consequently high-penalty DRAM pre-charging. To handle this, GPU achieves efficient parallelization with sophisticated optimization of CUDA program to reduce memory footprints, which demands high engineering efforts. In this work, we in contrast propose a programmable direct memory access engine for convolutional neural networks (CNN-DMA) supporting a fast supply of activation for independent and scalable computing units. The CNN-DMA favours a predictable activation streaming approach which completely avoids penalties by bus contention, cache misses and less carefully designed low-level programs. Furthermore, we enhance the baseline DMA with the capability of out-of-order data supply to filter out unique sliding-windows to boost the performance of the computing infrastructure. Experiments on state-of-the-art neural networks show that CNN-DMA achieves optimal DRAM access efficiency for point-wise convolution layers, while reduces 30% to 70% rounds of computation with sliding-window filtering.
Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Weiguang Chen, Wenxuan Chen, Weiyu Guo, Zhibin Yu 0001
ACM Great Lakes Symposium on VLSI10
2021 Towards CSI-based diversity activity recognition via LSTM-CNN encoder-decoder neural network
Linlin Guo, Hang Zhang 0011, Weiyu Guo, Guangqiang Diao, Bingxian Lu, Chuang Lin 0001, Lei Wang 0005
Neurocomputing4
2020 Explainable Enterprise Rating Using Attention Based Convolutional Neural Network
Weiyu Guo
WISA1
2018 RotateConv: Making Asymmetric Convolutional Kernels Rotatable
abstract
In deep Convolutional Neural Networks(CNN), the design of kernel shapes influences a lot on the model size and performance. In this work, our proposed method, RotateConv, applies a novel kernel shape to massively reduce the number of parameters while maintaining considerable performance. The new shape is extremely simple as a line segment one, and we equip it with the rotatable ability which aims to learn diverse features with respect to different angles. The kernel weights and angles are learned simultaneously during end-to-end training via the standard back-propagation algorithm. There are two variants of RotateConv that only have 2 and 4 parameters respectively depending on whether using weight sharing, which are much compressed than the normal 3×3 kernel with 9 parameters. In experiments, we validate our RotateConv with two classical models, ResNet and DenseNet, on four image classification benchmark datasets, namely MNIST, CIFAR10, CIFAR100 and SVHN.
Jiabin Ma, Weiyu Guo, Wei Wang 0115, Liang Wang 0001
ICPR2
2016 Personalized ranking with pairwise Factorization Machines
Weiyu Guo, Liang Wang 0001, Tieniu Tan
Neurocomputing1
2016 Coupled Topic Model for Collaborative Filtering With User-Generated Content
abstract
The user-generated content (UGC) is a type of dyadic information that provides description of the interaction between users and items (such as rating, purchasing, etc.). Most conventional methods incorporate either a user profile or the item description, which cannot well utilize this kind of content information. Some other works jointly consider user ratings and reviews, but they are based on the factorization technique and have difficulty in providing explanations on generated recommendations. In this study, a coupled topic model (CoTM) for recommendation with UGC is developed. By combining UGC and ratings, the method discussed in this study captures both the content-based preferences and collaborative preferences and, thus, can explain both the user and item latent spaces using the topics discovered from the UGC. The learned topics in CoTM can also serve as proper explanations for the generated recommendations. Experimental results show that the proposed CoTM model yields significant improvements over the compared competitive methods on two typical datasets, that is, MovieLens-10M and Citation-network V1. The topics discovered by CoTM can be used not only to illustrate the topic distributions of users and items, but also to explain the generated user-item recommendations.
Weiyu Guo, Song Xu 0002, Yongzhen Huang, Liang Wang 0001, Tieniu Tan
IEEE Trans. Hum. Mach. Syst.2
2015 Multiple Attribute Aware Personalized Ranking
Weiyu Guo, Liang Wang 0001, Tieniu Tan
APWeb1
2015 Social-Relational Topic Model for Social Networks
abstract
Social networking services, such as Twitter and Sina Weibo, have tremendous popularity in recent years. Mass of short texts and social links are aggregated into these service platforms. To realize personalized services on social network, topic inference from both short texts and social links plays more and more important role. Most conventional topic modeling methods focus on analyzing formal texts, e.g., papers, news and blogs, and usually assume that the links are only generated by topical factors. As a result, on social network, the learned topics of these methods are usually affected by topic-irrelevant links. Recently, a few approaches use artificial priors to recognize the links generated by the popularity factor in topic modeling. However, employing global priors, these methods can not well capture the distinct properties of each link and still suffer from the effect of topic-irrelevant links. To address the above limitations, we propose a novel Social-Relational Topic Model (SRTM), which can alleviate the effect of topic-irrelevant links by analyzing relational users' topics of each link. SRTM jointly models texts and social links for learning the topic distribution and topical influence of each user. The experimental results show that, our model outperforms the state-of-the-arts in topic modeling and social link prediction.
Weiyu Guo, Liang Wang 0001, Tieniu Tan
CIKM1
2007 The potential energy of knowledge flow
abstract
Abstract A knowledge flow is invisible but it plays an important role in ordering knowledge exchange when working in a team. It can help achieve effective team knowledge management by modeling, optimizing, monitoring, and controlling the operation of knowledge flow processes. This paper proposes the notion of knowledge energy as the driving force behind the formation of an autonomous knowledge flow network, and explores the underlying principles. Knowing these principles helps teams and the support systems improve cooperation by monitoring the knowledge energy of nodes, by evaluating and adjusting knowledge flows, and by adopting appropriate strategies. A knowledge flow network management mechanism can help improve the efficiency of knowledge‐intensive distributed teamwork. Copyright © 2006 John Wiley & Sons, Ltd.
Hai Zhuge, Weiyu Guo
Concurr. Comput. Pract. Exp.2
2007 Virtual knowledge service market - For effective knowledge flow within knowledge grid
Hai Zhuge, Weiyu Guo
J. Syst. Softw.2
2003 Theory and Algorithm for Rule Base Refinement
Hai Zhuge, Yunchuan Sun, Weiyu Guo
IEA/AIE3