VLDB 2026 Research / reviewers in the wild / expert
Guanbo Wang
dblp:294/8693
· DBLP profile ↗
19ranked-venue papers
7as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On-policy Reinforcement Fine-tuning with Offline reward for Multi-step Embodied PlanningabstractEmbodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and verbal goals.While recent vision-language models (VLMs) excel at static perception tasks, they struggle in interactive environments.Reinforcement learning (RL) offers a natural way to address this limitation, yet online RL approaches suffer from costly interaction and sparse rewards in embodied settings.This paper introduces ORBIT, an On-policy Reinforcement finetuning (RFT) framework with offline rewards for EmBodIed Task Planning, that preserves the generalization benefits of RFT while addressing the challenges of costly interaction and sparse rewards, supported by solid theoretical guarantees.Our approach is evaluated on EmbodiedBench, a recent benchmark for interactive embodied tasks, covering both in-domain and out-of-domain scenarios.Experimental results show that ORBIT achieves SOTA performance on EB-ALFRED, outperforming all closed-source and online-RL-based methods, while being substantially more efficient in training speed and computational cost, remaining robust to sub-optimal expert trajectories, and exhibiting strong generalization to unseen environments.We released all code and data at https://github.com/mail-taii/Reinforced- Reasoning-for Chloe Gu, Guanbo Wang, Wenhao Li 0001, Bo Jin 0003 |
ACL (1) | 4 |
| 2026 | EarAuth: Towards Practical Cardiac Vibration Authentication on COTS Wireless Earbuds
Yongjian Fu 0004, Wenpeng Zhu, Yingjun Wu, Hao Pan 0003, Guanbo Wang, Yongheng Deng, Yaoxue Zhang, Ju Ren 0001 |
INFOCOM | 6 |
| 2026 | Scale-Aware Attention and Multi-Modal Prompt Learning With Fusion Adapter for RGBT TrackingabstractFusing visible (RGB) and thermal (T) images for RGBT tracking has received growing interest in the field of computer vision. However, how to improve the robustness of the tracker to target scale variety, effectively apply visual prompts to multimodal tracking tasks, and enhance the multimodal fusion effectiveness are still urgent challenges in the field of RGBT tracking. To this purpose, this work proposes an RGBT tracking framework integrating scale-aware dilation attention, multimodal prompt interaction learning, and cross- fusion adapter, named MPANet. Firstly, a scale-aware dilation attention (SADA) module is put forward to enhance the flexibility of the tracker in the presence of target scale variations by embedding convolutions with different dilation rates into the self-attention. Subsequently, a multimodal prompt interaction learning (MPIL) module is constructed, which combines global token adaptive attention and spatial attention to efficiently learn visual prompts from different modalities and achieve intermodal prompt interactions. Finally, a cross-fusion adapter (CFA) is developed to facilitate the adaptability of the network to different modalities in the process of multimodal information fusion through the adapter mechanism. Extensive experiments on public RGBT benchmark tracking datasets such as GTOT, RGBT234, LasHeR and VTUAV demonstrate that the proposed method outperforms existing advanced trackers and achieves state-of-the-art performance. Victor S. Sheng, Yujun Ma, Xiaoguo Liang, Guanbo Wang |
IEEE Trans. Multim. | 6 |
| 2025 | ConCISE: Confidence-guided Compression in Step-by-step Efficient ReasoningabstractZiqing Qiao, Yongheng Deng, Jiali Zeng, Dong Wang, Lai Wei, Guanbo Wang, Fandong Meng, Jie Zhou, Ju Ren, Yaoxue Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Ziqing Qiao, Yongheng Deng, Jiali Zeng, Guanbo Wang, Fandong Meng, Jie Zhou 0016, Ju Ren 0001, Yaoxue Zhang |
EMNLP | 6 |
| 2025 | FedAF: Alignment-Augmented Fusion for Federated Multimodal Learning with Small LabelsabstractFederated multimodal learning is an emerging advancement in artificial intelligence, enabling the integration of data from diverse modalities while preserving data privacy. However, limited labeled data and modality heterogeneity on the clients pose significant challenges for effective federated multimodal model training. To address these challenges, this paper introduces FedAF, a novel alignment-augmented fusion framework tailored for federated multimodal learning. FedAF extracts unbiased and complementary information from multiple modalities with small data, enabling effective modality fusion and feature alignment for improving system performance. The framework introduces a three-stage strategy. First, FedAF utilizes labeled data to create unbiased anchor points, addressing disparities in client feature distributions. Second, FedAF employs a weighted enhancement contrast fusion scheme to improve feature clustering and reduce feature overlap. Finally, a multimodal semisupervised algorithm mitigates data heterogeneity and overfitting. Extensive experiments demonstrate that FedAF significantly outperforms baseline methods, showcasing its effectiveness in federated multimodal learning scenarios. Guanbo Wang, Yongheng Deng, Yingjun Wu, Xinyi Li 0005, Tuowei Wang, Yaoxue Zhang, Ju Ren 0001 |
IWQoS | 1 |
| 2025 | Efficient Randomized Experiments Using Foundation ModelsabstractRandomized experiments are the preferred approach for evaluating the effects of interventions, but they are costly and often yield estimates with substantial uncertainty. On the other hand, in silico experiments leveraging foundation models offer a cost-effective alternative that can potentially attain higher statistical precision. However, the benefits of in silico experiments come with a significant risk: statistical inferences are not valid if the models fail to accurately predict experimental responses to interventions.
In this paper, we propose a novel approach that integrates the predictions from multiple foundation models with experimental data while preserving valid statistical inference. Our estimator is consistent and asymptotically normal, with asymptotic variance no larger than the standard estimator based on experimental data alone. Importantly, these statistical properties hold even when model predictions are arbitrarily biased. Empirical results across several randomized experiments show that our estimator offers substantial precision gains, equivalent to a reduction of up to 20\% in the sample size needed to match the same precision as the standard estimator based on experimental data alone. Piersilvio De Bartolomeis, Javier Abad, Guanbo Wang, Konstantin Donhauser, Raymond M. Duch, Fanny Yang, Issa J. Dahabreh |
NeurIPS | 3 |
| 2025 | Gains: Fine-grained Federated Domain Adaptation in Open SetabstractConventional federated learning (FL) assumes a closed world with a fixed total number of clients. In contrast, new clients continuously join the FL process in real-world scenarios, introducing new knowledge. This raises two critical demands: detecting new knowledge, i.e., knowledge discovery, and integrating it into the global model, i.e., knowledge adaptation. Existing research focuses on coarse-grained knowledge discovery, and often sacrifices source domain performance and adaptation efficiency. To this end, we propose a fine-grained federated domain adaptation approach in open set (Gains). Gains splits the model into an encoder and a classifier, empirically revealing features extracted by the encoder are sensitive to domain shifts while classifier parameters are sensitive to class increments. Based on this, we develop fine-grained knowledge discovery and contribution-driven aggregation techniques to identify and incorporate new knowledge. Additionally, an anti-forgetting mechanism is designed to preserve source domain performance, ensuring balanced adaptation. Experimental results on multi-domain datasets across three typical data-shift scenarios demonstrate that Gains significantly outperforms other baselines in performance for both source-domain and target-domain clients. Code is available at: https://github.com/Zhong-Zhengyi/Gains. Zhengyi Zhong, Wenzheng Jiang, Weidong Bao 0001, Ji Wang 0002, Cheems Wang, Guanbo Wang, Yongheng Deng, Ju Ren 0001 |
NeurIPS | 6 |
| 2025 | Fighting against forest fire: A lightweight real-time detection approach for forest fire based on synthetic images
Guanbo Wang, Pengfei Yu 0003, Zhaisheng Ding, Zongshan Wang, Shidong Xie |
Expert Syst. Appl. | 1 |
| 2025 | SCDFuse: A semantic complementary distillation framework for joint infrared and visible image fusion and denoising
Shidong Xie, Yongsheng Zang, Jinde Cao, Dongming Zhou 0001, Mingchuan Tan, Zhaisheng Ding, Guanbo Wang |
Knowl. Based Syst. | 8 |
| 2025 | DPMNet: A Remote Sensing Forest Fire Real-Time Detection Network Driven by Dual Pathways and Multidimensional Interactions of FeaturesabstractA fundamental challenge in remote sensing-based forest fire detection lies in accurately discerning fire characteristics on various scales against the backdrop of intricate and heterogeneous forest landscapes. In response to this challenge, we propose a dual-path network (DPMNet) with multidimensional feature interaction for real time remote sensing forest fire detection. Initially, a dual-path backbone network is designed, integrating coarse-grained and fine-grained parallel pathways, working in tandem to capture both global visual features and nuanced local texture details. Subsequently, we develop the Multidimensional Interactive Feature Pyramid Network (MiFPN), a novel structure that amalgamates information streams from varied levels through a three-branch structure and engenders profound fusion and dynamic interaction of features across multiple scales. Thereafter, the Context-Enriched Adaptive Fusion Module (CEAFM) is proposed, which emerges to meticulously blend macroscopic visual elements harvested via coarse-grained conduits, employing a multi-faceted pathway strategy to bolster the model’s overarching comprehension and precision in forest fire detection. Finally, the Enhanced Contextual Pooling Bottleneck (ECPB) is put forward, an integration that augments the model’s spatial perception and contextual acumen through the incorporation of dilated convolution and global pooling techniques. Extensive experiments are conducted on the remote sensing forest fire dataset in order to confirm the efficacy of DPMNet. The experimental results demonstrate that our DPMNet achieves satisfactory performance in terms of real-time performance as well as accuracy and provides an effective solution for real-time detection of remote sensing forest fires based on UAVs. Guanbo Wang, Victor S. Sheng, Yujun Ma, Hongwei Ding 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | RelayRec: Empowering Privacy-Preserving CTR Prediction via Cloud-Device Relay LearningabstractClick-through rate (CTR) prediction holds paramount importance across numerous applications, profoundly impacting user experience and business profitability. The freshness of a CTR prediction model significantly influences its performance, since users’ needs and interests may be changing over time, thereby requiring the model to be updated frequently. However, stringent data protection regulations have constrained the collection of users’ personal data, posing challenges to traditional model refreshing strategies that rely on centralized data collection. On-device learning techniques, such as federated learning (FL), offer a viable solution by enabling model training on devices without compromising user privacy. Nevertheless, the scarcity of training data with diverse distributions among devices presents considerable obstacles to on-device learning effectiveness. To address these challenges, we introduce RelayRec, a cloud-device relay learning framework designed for privacy-preserving CTR prediction. To establish competent initial models for devices, RelayRec categorizes pre-regulation cloud data into user preference groups, training preference-specific models for devices. Furthermore, a cloud-based automated model selector is developed to identify suitable initial models for devices. To elevate the relay learning performance of these initial models, we incorporate a personalized collaborative learning mechanism that aggregates device models based on user preferences. Extensive experimental evaluations underscore RelayRec’s superior performance compared to state-of-the-art benchmarks, affirming its efficacy in privacy-preserving CTR prediction. Yongheng Deng, Guanbo Wang, Sheng Yue 0001, Wei Rao 0003, Qin Zu, Ju Ren 0001, Yaoxue Zhang |
IPSN | 2 |
| 2024 | M4SFWD: A Multi-Faceted synthetic dataset for remote sensing forest wildfires detection
Guanbo Wang, Xun Lang, Yanling Feng, Zhaisehng Ding, Shidong Xie |
Expert Syst. Appl. | 1 |
| 2024 | Real-time detection algorithm for non-motorized vehicles based on D-YOLO model
Hongwei Ding 0006, Guanbo Wang |
Multim. Tools Appl. | 5 |
| 2024 | RFWNet: A Multiscale Remote Sensing Forest Wildfire Detection Network With Digital Twinning, Adaptive Spatial Aggregation, and Dynamic Sparse FeaturesabstractReal-time detection of forest fires through remote sensing is a challenging task, especially in the context of limited data availability. In response to this challenge, this article leverages the digital twin (DT) concept to create a comprehensive and high-fidelity synthetic forest wildfire dataset. Alongside this, we have made available a high-resolution forest fire remote sensing dataset from real scenarios, meticulously collected and annotated by our research team. Aiming for precision in detecting forest fires via remote sensing, we present the remote sensing forest wildfire detection network (RFWNet) and its lightweight version, RFWNet-nano. More specifically, our network’s backbone, grounded on deformable convolution network v3 (DCNv3), develops a multigroup mechanism, amplifying its ability to perceive the correlations overextended distances. Utilizing our dual-path dynamic sparse attention (DDSA), we meld coarse-grained regional selection with granular token-to-token attention, adeptly capturing the evolving contours of fires and smoke. To address diverse scenarios, our Vanilla Head design, backed by a profound training approach and simultaneous stacked activations, accurately identifies flames and smoke across multiple scales. Furthermore, we advocate for a 24/7 real-time monitoring system, synergizing drones, edge computing devices, and NVIDIA GPUs. Our experimental outcomes indicate that relative to numerous prevailing object detection algorithms, RFWNet and RFWNet-nano both manifest considerable superiority in terms of quantitative precision and visual results, substantiating the robustness and preeminence of our methodologies. Guanbo Wang, Shuhua Ye, Hongwei Ding 0001, Shidong Xie |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | CLARE: Conservative Model-Based Reward Learning for Offline Inverse Reinforcement Learning
Sheng Yue 0001, Guanbo Wang, Wei Shao 0006, Sen Lin 0001, Ju Ren 0001, Junshan Zhang |
ICLR | 2 |
| 2023 | FedINC: An Exemplar-Free Continual Federated Learning Framework with Small Labeled DataabstractFederated learning (FL) has shown great promise for privacy-preserving learning by enabling collaborative training on decentralized clients. However, in realistic FL scenarios, clients often collect new data continuously, join or exit learning dynamically. As a result, the global model tends to forget old knowledge while learning new knowledge. Meanwhile, labeling the continuously arriving data in real-time is usually challenging. Therefore, the catastrophic forgetting problem intertwined with the label deficiency issue poses significant challenges for both learning new knowledge and consolidating old knowledge. To address these challenges, we develop a novel exemplar-free continual federated learning framework named FedINC, to learn a global incremental model with limited labeled data. We begin by excavating the cause of catastrophic forgetting via in-depth empirical studies. Based on that, we introduce targeted mechanisms for FedINC, including a hybrid contrastive learning mechanism to efficiently learn new knowledge with limited labeled data, a plastic feature regularization mechanism to preserve old task's representation space, a prototype-guided regularization mechanism to mitigate feature overlap between old and new classes while aligning the features of non-iid clients, and a prototype evolution mechanism for flexible and efficient incremental classification. Extensive experiments demonstrate the superior performance of FedINC in terms of both convergence speed and accuracy of the global model. Yongheng Deng, Sheng Yue 0001, Tuowei Wang, Guanbo Wang, Ju Ren 0001, Yaoxue Zhang |
SenSys | 4 |
| 2022 | The THUEE System Description for the IARPA OpenASR21 ChallengeabstractThis paper describes the THUEE team's speech recognition system for the IARPA Open Automatic Speech Recognition Challenge (OpenASR21), with further experiment explorations. We achieve outstanding results under both the Constrained and Constrained-plus training conditions. For the Constrained training condition, we construct our basic ASR system based on the standard hybrid architecture. To alleviate the Out-Of-Vocabulary (OOV) problem, we extend the pronunciation lexicon using Grapheme-to-Phoneme (G2P) techniques for both OOV and potential new words. Standard acoustic model structures such as CNN-TDNN-F and CNN-TDNN-F-A are adopted. In addition, multiple data augmentation techniques are applied. For the Constrained-plus training condition, we use the self-supervised learning framework wav2vec2.0. We experiment with various fine-tuning techniques with the Connectionist Temporal Classification (CTC) criterion on top of the publicly available pre-trained model XLSR-53. We find that the frontend feature extractor plays an important role when applying the wav2vec2.0 pre-trained model to the encoder-decoder based CTC/Attention ASR architecture. Extra improvements can be achieved by using the CTC model finetuned in the target language as the frontend feature extractor. Haoyu Wang 0014, Shuzhou Chai, Guanbo Wang, Guoguo Chen, Weiqiang Zhang 0001 |
INTERSPEECH | 5 |
| 2022 | TRC-YOLO: A real-time detection method for lightweight targets based on mobile devicesabstractAbstract Object detection is one of the main tasks of computer vision. Object detection algorithms usually rely on deep convolutional neural networks, which require the host device to have high computing capabilities, greatly limiting the application of object detection methods for mobile devices with limited computing capabilities, such as embedded devices. Among the current object detection algorithms, the you only look once (YOLO) series takes both speed and accuracy into consideration and is one of the most commonly used methods for object detection. In this article, TRC‐YOLO is proposed, which improves the mean average precision (mAP) and real‐time detection speed of the model while reducing the size of the model. In TRC‐YOLO, the convolution kernel of YOLO v4‐tiny is pruned and an expansive convolution layer is introduced into the residual module of the network to produce an hourglass Cross Stage Partial ResNet (CSPResNet) structure. A receptive field block (RFB) that simulates human vision is also added, increasing the receptive field of the model and strengthening the feature extraction ability of the network. In addition, the convolutional block attention module is applied, which combines spatial attention and channel attention, to enhance the effective features of the model and reduce the negative impact of noise on the model. The size of the TRC‐YOLO model is 17.8 MB, which is 5.9 MB smaller than YOLO v4‐tiny, and the model parameter is 2.983 billion floating point operations per second (BFLOP/s) (3.834 BFLOP/s less than YOLO v4‐tiny). In addition, TRC‐YOLO achieves a real‐time performance of 36.9 frames per second on a Jetson Xavier NX, and its mAP on the PASCAL VOC dataset is 66.4 (3.83 higher than YOLO v4‐tiny). In addition, the mAP of TRC‐YOLO on the MS COCO dataset is 37.7, which is 1.9 higher than that of the baseline model. Guanbo Wang, Hongwei Ding 0001, Bo Li 0025, Liyong Bao |
IET Comput. Vis. | 1 |
| 2022 | Trident-YOLO: Improving the precision and speed of mobile device object detectionabstractAbstract This paper introduce an efficient object detection network named Trident‐You Only Look Once (YOLO), which is designed for mobile devices with limited computing power. The new architecture is improved based on YOLO v4‐tiny. The authors redesign the network structure and propose a trident feature pyramid network (Trident‐FPN), which can improve the precision and recall of lightweight object detection. Specifically, Trident‐FPN increases the computational complexity by only a small amount of floating point operations per second (FLOPs) and obtains a multi‐scale feature map of the model, which significantly lightweight object detection performance. To enlarge the receptive field of the network with the fewest FLOPs, this paper redesign the receptive field block (RFB) and spatial pyramid pooling (SPP) layer and propose tinier cross‐stage partial RFBs and smaller cross‐stage partial SPPs. This paper present extensive experiments, and Trident‐YOLO shows strong performance compared to that of other popular models on the PASCAL VOC and MS COCO. On the MS COCO and PASCAL VOC 2007 test sets, the mean average precision (mAP) of Trident‐YOLO improved by 4.5% and 5.0%, respectively. Trident‐YOLO also reduce the network size by more than 54.4% compared to YOLO v4‐tiny. With a 23.7% FLOP reduction, the FPS is improved by 1.9 on an Nvidia Jetson Xavier NX. Guanbo Wang, Hongwei Ding 0001, Bo Li 0025, Rencan Nie |
IET Image Process. | 1 |