Guanlin Wu

dblp:203/0815 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-9968-2977ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toward Optimal Mixture of Experts System for 3D Object Detection: A Game of Accuracy, Efficiency and Adaptivity
abstract
Autonomous vehicles, open-world robots, and other automated systems rely on accurate, efficient perception modules for real-time object detection. Although high-precision models improve reliability, their processing time and computational overhead can hinder real-time performance and raise safety concerns. This paper introduces an Edge-based Mixture-of-Experts Optimal Sensing (EMOS) System that addresses the challenge of co-achieving accuracy, latency and scene adaptivity, further demonstrated in the open-world autonomous driving scenarios. Algorithmically, EMOS fuses multimodal sensor streams via an Adaptive Multimodal Data Bridge and uses a scenario-aware MoE switch to activate only a complementary set of specialized experts as needed. The proposed hierarchical backpropagation and a multiscale pooling layer let model capacity scale with real-world demand complexity. System-wise, an edge-optimized runtime with accelerator-aware scheduling (e.g., ONNX/TensorRT), zero-copy buffering, and overlapped I/O-compute enforces explicit latency/accuracy budgets across diverse driving conditions. Experimental results establish EMOS as the new state of the art: on KITTI, it increases average AP by 3.17% while running $2.6\times$2.6× faster on Nvidia Jetson. On nuScenes, it improves accuracy by 0.2% mAP and 0.5% NDS, with 34% fewer parameters and a $15.35\times$15.35× Nvidia Jetson speedup. Leveraging multimodal data and intelligent experts cooperation, EMOS delivers accurate, efficient and edge-adaptive perception system for autonomous vehicles, thereby ensuring robust, timely responses in real-world scenarios.
Linshen Liu, Guanlin Wu, Junyue Jiang, Hao (Frank) Yang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Climber: Toward Efficient Scaling Laws for Large Recommendation Models
abstract
Transformer-based generative models have achieved remarkable success across domains with various scaling law manifestations. However, our extensive experiments reveal persistent challenges when applying Transformer to recommendation systems: (1) Transformer scaling is not ideal with increased computational resources, due to structural incompatibilities with recommendation-specific features such as multi-source data heterogeneity; (2) critical online inference latency constraints (tens of milliseconds) that intensify with longer user behavior sequences and growing computational demands. We propose Climber, an efficient recommendation framework comprising two synergistic components: the model architecture for efficient scaling and the co-designed acceleration techniques. Our proposed model adopts two core innovations: (1) multi-scale sequence extraction that achieves a time complexity reduction by a constant factor, enabling more efficient scaling with sequence length; (2) dynamic temperature modulation adapting attention distributions to the multi-scenario and multi-behavior patterns. Complemented by acceleration techniques, Climber achieves a 5.15× throughput gain without performance degradation by adopting a ''single user, multiple item'' batched processing and memory-efficient Key-Value caching.
Songpei Xu, Da Guo, Xianwen Guo, Bin Huang 0012, Guanlin Wu, Chuanjiang Luo
CIKM7
2025 HumanMM: Global Human Motion Recovery from Multi-shot Videos
abstract
In this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understanding, but are of great challenge to be recovered due to abrupt shot transitions, partial occlusions, and dynamic backgrounds presented in such videos. Existing methods primarily focus on single-shot videos, where continuity is maintained within a single camera view, or simplify multi-shot alignment in camera space only. In this work, we tackle the challenges by integrating an enhanced camera pose estimation with Human Motion Recovery (HMR) by incorporating a shot transition detector and a robust alignment module for accurate pose and orientation continuity across shots. By leveraging a custom motion integrator, we effectively mitigate the problem of foot sliding and ensure temporal consistency in human pose. Extensive evaluations on our created multi-shot dataset from public 3D human datasets demonstrate the robustness of our method in reconstructing realistic human motion in world coordinates.
Guanlin Wu, Zhuokai Zhao, Xiaoke Jiang, Zhuoheng Li, Hao (Frank) Yang, Haoqian Wang, Lei Zhang 0001
CVPR2
2025 Emotion Transfer with Enhanced Prototype for Unseen Emotion Recognition in Conversation
abstract
Current Emotion Recognition in Conversation (ERC) research follows a closed-domain assumption.However, there is no clear consensus on emotion classification in psychology, which presents a challenge for models when it comes to recognizing previously unseen emotions in real-world applications.To bridge this gap, we introduce the Unseen Emotion Recognition in Conversation (UERC) task for the first time and propose ProEmoTrans, a solid prototype-based emotion transfer framework.This prototype-based approach shows promise but still faces key challenges: First, implicit expressions complicate emotion definition, which we address by proposing an LLM-enhanced description approach.Second, utterance encoding in long conversations is difficult, which we tackle with a proposed parameter-free mechanism for efficient encoding and overfitting prevention.Finally, the Markovian flow nature of emotions is hard to transfer, which we address with an improved Attention Viterbi Decoding (AVD) method to transfer seen emotion transitions to unseen emotions.Extensive experiments on three datasets show that our method serves as a strong baseline for preliminary exploration in this new area.
Cong Cao 0001, Hao Peng 0001, Guanlin Wu, Zhifeng Hao 0004, Lei Jiang 0003, Yanbing Liu 0007, Philip S. Yu
EMNLP4
2025 Towards Accurate and Efficient 3D Object Detection for Autonomous Driving: A Mixture of Experts Computing System on Edge
abstract
This paper presents Edge-based Mixture of Experts (MoE) Collaborative Computing (EMC2), an optimal computing system designed for autonomous vehicles (AVs) that simultaneously achieves low-latency and high-accuracy 3D object detection. Unlike conventional approaches, EMC2 incorporates a scenario-aware MoE architecture specifically optimized for edge platforms. By effectively fusing LiDAR and camera data, the system leverages the complementary strengths of sparse 3D point clouds and dense 2D images to generate robust multimodal representations. To enable this, EMC2 employs an adaptive multimodal data bridge that performs multi-scale preprocessing on sensor inputs, followed by a scenario-aware routing mechanism that dynamically dispatches features to dedicated expert models based on object visibility and distance. In addition, EMC2 integrates joint hardware-software optimizations, including hardware resource utilization optimization and computational graph simplification, to ensure efficient and real-time inference on resource-constrained edge devices. Experiments on open-source benchmarks clearly show the EMC2 advancements as an end-to-end system. On the KITTI dataset, it achieves an average accuracy improvement of 3.58% and a 159.06% inference speedup compared to 15 baseline methods on Jetson platforms, with similar performance gains on the nuScenes dataset, highlighting its capability to advance reliable, real-time 3D object detection tasks for AVs. The official implementation is available at https://github.com/LinshenLiu622/EMC2.
Linshen Liu, Boyan Su, Junyue Jiang, Guanlin Wu, Cong Guo 0003, Ceyu Xu, Hao (Frank) Yang
ICCV4
2025 T-T: Table Transformer for Tagging-based Aspect Sentiment Triplet Extraction
abstract
Aspect sentiment triplet extraction (ASTE) aims to extract triplets composed of aspect terms, opinion terms, and sentiment polarities from given sentences. The table tagging method is a popular approach to addressing this task, which encodes a sentence into a 2-dimensional table, allowing for the tagging of relations between any two words. Previous efforts have focused on designing various downstream relation learning modules to better capture interactions between tokens in the table, revealing that a stronger capability in relation capture can lead to greater improvements in the model. Motivated by this, we attempt to directly utilize transformer layers as downstream relation learning modules. Due to the powerful semantic modeling capability of transformers, it is foreseeable that this will lead to excellent improvement. However, owing to the quadratic relation between the length of the table and the length of the input sentence sequence, using transformers directly faces two challenges: overly long table sequences and unfair local attention interaction. To address these challenges, we propose a novel Table-Transformer (T-T) for the tagging-based ASTE method. Specifically, we introduce a stripe attention mechanism with a loop-shift strategy to tackle these challenges. The former modifies the global attention mechanism to only attend to a 2-dimensional local attention window, while the latter facilitates interaction between different attention windows. Extensive and comprehensive experiments demonstrate that the T-T, as a downstream relation learning module, achieves state-of-the-art performance with lower computational costs.
Chaodong Tong, Cong Cao 0001, Hao Peng 0001, Qian Li 0033, Guanlin Wu, Lei Jiang 0003, Yanbing Liu 0007, Philip S. Yu
IJCAI6
2025 STAMImputer: Spatio-Temporal Attention MoE for Traffic Data Imputation
abstract
Traffic data imputation is fundamentally important to support various applications in intelligent transportation systems such as traffic flow prediction. However, existing time-to-space sequential methods often fail to effectively extract features in block-wise missing data scenarios. Meanwhile, the static graph structure for spatial feature propagation significantly constrains the model's flexibility in handling the distribution shift issue for the nonstationary traffic data. To address these issues, this paper proposes a Spatio-Temporal Attention Mixture of experts network named STAMImputer for traffic data imputation. Specifically, we introduce a Mixture of Experts (MoE) framework to capture latent spatio-temporal features and their influence weights, effectively imputing block missing. A novel Low-rank guided Sampling Graph ATtention (LrSGAT) mechanism is designed to dynamically balance the local and global correlations across road networks. The sampled attention vectors are utilized to generate dynamic graphs that capture real-time spatial correlations. Extensive experiments are conducted on four traffic datasets for evaluation. The result shows STAMImputer achieves significantly performance improvement compared with existing SOTA approaches. Our codes are available at https://github.com/RingBDStack/STAMImupter.
Yiming Wang 0010, Hao Peng 0001, Senzhang Wang, Haohua Du, Jia Wu 0001, Guanlin Wu
IJCAI7
2025 Unsupervised Graph Clustering with Deep Structural Entropy
abstract
Research on Graph Structure Learning (GSL) provides key insights for graph-based clustering, yet current methods like Graph Neural Networks (GNNs), Graph Attention Networks (GATs), and contrastive learning often rely heavily on the original graph structure. Their performance deteriorates when the original graph's adjacency matrix is too sparse or contains noisy edges unrelated to clustering. Moreover, these methods depend on learning node embeddings and using traditional techniques like k-means to form clusters, which may not fully capture the underlying graph structure between nodes. To address these limitations, this paper introduces DeSE, a novel unsupervised graph clustering framework incorporating Deep Structural Entropy. It enhances the original graph with quantified structural information and deep neural networks to form clusters. Specifically, we first propose a method for calculating structural entropy with soft assignment, which quantifies structure in a differentiable form. Next, we design a Structural Learning layer (SLL) to generate an attributed graph from the original feature data, serving as a target to enhance and optimize the original structural graph, thereby mitigating the issue of sparse connections between graph nodes. Finally, our clustering assignment method (ASS), based on GNNs, learns node embeddings and a soft assignment matrix to cluster on the enhanced graph. The ASS layer can be stacked to meet downstream task requirements, minimizing structural entropy for stable clustering and maximizing node consistency with edge-based cross-entropy loss. Extensive comparative experiments are conducted on four benchmark datasets against eight representative unsupervised graph clustering baselines, demonstrating the superiority of the DeSE in both effectiveness and interpretability.
Jingyun Zhang 0001, Hao Peng 0001, Li Sun 0008, Guanlin Wu, Zhengtao Yu 0001
KDD (2)4
2025 Towards Physics-informed Spatial Intelligence with Human Priors: An Autonomous Driving Pilot Study
abstract
How to integrate and verify spatial intelligence in foundation models remains an open challenge. Current practice often proxies Visual-Spatial Intelligence (VSI) with purely textual prompts and VQA-style scoring, which obscures geometry, invites linguistic shortcuts, and weakens attribution to genuinely spatial skills. We introduce Spatial Intelligence Grid (SIG): a structured, grid-based schema that explicitly encodes object layouts, inter-object relations, and physically grounded priors. As a complementary channel to text, SIG provides a faithful, compositional representation of scene structure for foundation-model reasoning. Building on SIG, we derive SIG-informed evaluation metrics that quantify a model’s intrinsic VSI, which separates spatial capability from language priors. In few-shot in-context learning with state-of-the-art multimodal LLMs (e.g. GPT- and Gemini-family models), SIG yields consistently larger, more stable, and more comprehensive gains across all VSI metrics compared to VQA-only representations, indicating its promise as a data-labeling and training schema for learning VSI. We also release SIGBench, a benchmark of 1.4K driving frames annotated with ground-truth SIG labels and human gaze traces, supporting both grid-based machine VSI tasks and attention-driven, human-like VSI tasks in autonomous-driving scenarios.
Guanlin Wu, Boyan Su, Yang Zhao 0013, Yichen Lin, Hao (Frank) Yang
NeurIPS1
2025 Structural Information-based Hierarchical Diffusion for Offline Reinforcement Learning
abstract
Diffusion-based generative methods have shown promising potential for modeling trajectories from offline reinforcement learning (RL) datasets, and hierarchical diffusion has been introduced to mitigate variance accumulation and computational challenges in long-horizon planning tasks. However, existing approaches typically assume a fixed two-layer diffusion hierarchy with a single predefined temporal scale, which limits adaptability to diverse downstream tasks and reduces flexibility in decision making. In this work, we propose SIHD, a novel Structural Information-based Hierarchical Diffusion framework for effective and stable offline policy learning in long-horizon environments with sparse rewards. Specifically, we analyze structural information embedded in offline trajectories to construct the diffusion hierarchy adaptively, enabling flexible trajectory modeling across multiple temporal scales. Rather than relying on reward predictions from localized sub-trajectories, we quantify the structural information gain of each state community and use it as a conditioning signal within the corresponding diffusion layer. To reduce overreliance on offline datasets, we introduce a structural entropy regularizer that encourages exploration of underrepresented states while avoiding extrapolation errors from distributional shifts. Extensive evaluations show that SIHD significantly outperforms state-of-the-art baselines in decision-making performance and demonstrates superior generalization across diverse scenarios.
Xianghua Zeng, Hao Peng 0001, Yicheng Pan 0001, Angsheng Li, Guanlin Wu
NeurIPS5
2024 Energy-Efficient MIMO Integrated Sensing and Communications With On-Off Nontransmission Power
abstract
This paper investigates the energy efficiency of a multiple-input multiple-output (MIMO) integrated sensing and communications (ISAC) system for Internet of things (IoT), in which one multi-antenna IoT transceiver transmits unified ISAC signals to a multi-antenna communication user (CU) and at the same time use the echo signals to estimate an extended target. We focus on one particular ISAC transmission block and take into account the practical on-off non-transmission power at the IoT transceiver. Under this setup, we minimize the energy consumption at the transceiver while ensuring a minimum average data rate requirement for communication and a maximum Cramér-Rao bound (CRB) requirement for target estimation, by jointly optimizing the transmit covariance matrix and the “on” duration for active transmission. We obtain the optimal solution to the rate-and-CRB-constrained energy minimization problem in a semi-closed form. Interestingly, the obtained optimal solution is shown to unify the spectrum-efficient and energy-efficient communications and sensing designs. In particular, for the special MIMO sensing case with rate constraint inactive, the optimal solution follows the isotropic transmission with shortest “on” duration, in which the IoT transceiver radiates the required sensing energy by using sufficiently high power over the shortest duration. For the general ISAC case, the optimal transmit covariance solution is of full rank and follows the eigenmode transmission based on the communication channel, while the optimal “on” duration is determined based on both the rate and CRB constraints. Numerical results show that the proposed ISAC design achieves significantly reduced energy consumption as compared to the benchmark schemes based on isotropic transmission, always-on transmission, and sensing or communications only designs, especially when the rate and CRB constraints become stringent.
Guanlin Wu, Yuan Fang 0002, Jie Xu 0002, Zhiyong Feng 0001, Shuguang Cui
IEEE Internet Things J.1
2023 Dyna-PPO reinforcement learning with Gaussian process for the continuous action decision-making in autonomous driving
Guanlin Wu, Wenqi Fang, Ji Wang 0002, Pin Ge, Jiang Cao, Yang Ping, Peng Gou
Appl. Intell.1
2023 Approximation algorithms for a virtual machine allocation problem with finite types
Lifeng Guo, Changhong Lu, Guanlin Wu
Inf. Process. Lett.3
2022 Qauxi: Cooperative multi-agent reinforcement learning with knowledge transferred from auxiliary task
Wenqian Liang, Ji Wang 0002, Weidong Bao 0001, Xiaomin Zhu 0001, Guanlin Wu, Dayu Zhang, Liyuan Niu
Neurocomputing5
2022 Contrastive autoencoder for anomaly detection in multivariate time series
Hao Zhou 0032, Ke Yu 0001, Xuan Zhang 0007, Guanlin Wu, Anis Yazidi
Inf. Sci.4
2022 An Edge Storage Acceleration Service for Collaborative Mobile Devices
abstract
Fueled by the advances in the Internet of Things, and the growing capacity of smart mobile devices at the edge of the Internet, we have witnessed a growing trend in research and development for edge computing and edge storage, which extends the abilities of single mobile device on the edge through on-demand collaboration among multiple geographically distributed mobile devices. In this article, we address several technical challenges that are unique for collaborative storage at the edge due to the unique characteristics of mobile devices. First, we formalize the collaborative storage problem as an optimization problem. Second, we design an Acceleration Algorithm for Collaborative Storage, called A2CS, based on the architecture of Alternating Direction Method of Multipliers (ADMM). Specifically, we use the Nesterov’s Acceleration strategy and the step size rules in the process of updating variables and determining the optimal speed of convergence. We develop a novel collaborative storage policy in order to guide the whole lifecycle of collaborative storage. Finally, we conduct a series of experiments for acceleration performance analysis and validation. We show that A2CS delivers a better convergence performance with different step size rules, compared with two existing approaches: the ADMM baseline and the ADMM-OR (ADMM with Over-Relaxation), achieving the acceleration percentage by at least 25.33 percent and at most 64.01 percent. In addition, by conducting the utility performance comparison analysis with the existing Average Distribution Strategy (ADS) and the existing Distance Preferred Distribution Strategy (DPDS), we show the advantage of A2CS over both ADS and DPDS with respect to the total utility and energy consumption.
Xiong Gao, Weidong Bao 0001, Xiaomin Zhu 0001, Guanlin Wu, Ling Liu 0001
IEEE Trans. Serv. Comput.4
2021 Identifying the module structure of swarms using a new framework of network-based time series clustering
Kongjing Gu, Ziyang Mao, Xiaojun Duan, Guanlin Wu, Liang Yan 0003
Eng. Appl. Artif. Intell.4
2017 A Lightweight Recommendation Framework for Mobile User's Link Selection in Dense Network
abstract
With the proliferation of mobile devices and the development of communication technology, mobile devices have permeated every aspect of our daily lives. However, in dense network where large crowd of mobile devices try to access to the network simultaneously, the severe interference between mobile devices may incur a remarkable deterioration of the wireless communication quality. How to improve individual's experience in such scenario is a critical yet open problem. Inspired by the mobile device users' usage pattern as well as the characteristic of most wireless communication systems, we propose a framework offering uplink/downlink selection recommendation to different mobile device users to enhance their utility in this paper. The design of the framework starts with formulating the problem as a link selection game. Analysis shows that the game can be categorized as a generalized ordinal potential game whose Nash Equilibrium is guaranteed. We then devise a distributed link selection algorithm to generate a Nash Equilibrium of the game. To accommodate to the characteristic of dense network and the capacity limitation of mobile device, the design of the algorithm shows a light-weight property and does not require each mobile device user to know others' current selection. The probability of incomplete information gathering is also considered. Extensive experiments are conducted to demonstrate the effectiveness and superiority of the proposed framework. Experimental results show that the global average utility increase rate reaches above 20%, and about 70% mobile device users can benefit from using our framework.
Ji Wang 0002, Xiaomin Zhu 0001, Weidong Bao 0001, Guanlin Wu
ICDCS4
2017 Towards collaborative storage scheduling using alternating direction method of multipliers for mobile edge cloud
Guanlin Wu, Junjie Chen 0007, Weidong Bao 0001, Xiaomin Zhu 0001, Wenhua Xiao, Ji Wang 0002
J. Syst. Softw.1