VLDB 2026 Research / reviewers in the wild / expert
Penghong Wang
dblp:240/6176
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-5401-5676ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Neural Skill-Based Meta-Reinforcement Learning for Efficient Adaptability in Robotic Manipulation TasksabstractRobotic manipulation tasks frequently share foundational structures. Meta-reinforcement learning aims to develop generalizable policies that leverage these shared structures. However, existing methods usually struggle to efficiently encode this task-specific knowledge into their policy networks: hierarchical policies depend on task-agnostic action-level skills or intrinsic rewards, while context-based paradigms suffer from Markov Decision Process ambiguity. To address these challenges, we propose a hierarchical neural skill-based meta-reinforcement learning framework. This framework includes neural skill generation, neural skill decoding, and policy network construction. A Transformer-based neural skill generation unit sequentially generates hierarchical neural skills conditioned on a task description. The decoding mechanism translates these neural skills into network parameters for a policy. Using these decoded parameters, the policy network is constructed layer by layer to process the state information. Unlike previous work, our method treats skills as abstractions of layer-wise network parameters, allowing task-specific knowledge embedded in neural skills to directly configure the policy network. Experimental results demonstrate that our method has enhanced flexibility and efficiency. Hao Wang 0212, Wenrui Li 0001, Penghong Wang, Xianqi Zhang, Xiaopeng Fan 0001 |
IEEE Signal Process. Lett. | 3 |
| 2026 | DV-Hop Localization Based on Probability Distance Estimation and Expected Hop Distance CorrectionabstractDistance estimation and theoretical derivation in 3D space form the foundation basis for improving localization performance in wireless sensor networks (WSNs). Localization is a pivotal challenge in wireless sensor network (WSN) applications. To address this issue, we propose a probability-based distance estimation (PDE) model and a distance correction strategy based on expected hops (DCSEH). First, the PDE model is constructed from the multi-hop probability distribution of nodes, from which the upper bound and average distance for anchor nodes to detect target nodes under different hop counts are derived. Second, the DCSEH strategy effectively mitigates transmission-path detours in wireless node communication. Finally, the constructed loss function is embedded into a multi-objective genetic algorithm to predict the position of each unknown node in three-dimensional space. Extensive experiments demonstrate that the proposed method achieves state-of-the-art 3D localization performance on both random and multimodal datasets. Penghong Wang, Hao Wang 0212, Wenrui Li 0001, Hengyu Man, Xin Yue, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | TS-Net: Assembling Task-specific Features from Multiple Feature Levels for Multi-task LearningabstractMulti-task learning (MTL) has become an attractive topic that leverages shared knowledge to improve performance and enhance generalization. However, most existing works neglect the varying contribution of multi-level features to sub-task representations. In this paper, we explore the impact of multilevel features on different tasks and propose a novel level-assembling MTL architecture named TS-Net. TS-Net integrates multi-level features into multi-task representations by combining task-specific and task-generic features. We first introduce a Task-Specific Feature Capturing Block (TSFCB) to aggregate task-specific features by dynamically assembling features for input samples and prioritizing more relevant feature levels. In addition, we present a Multi-Task Mixture-of-Experts (MTMoE) module to facilitate cross-task interaction. In MTMoE, task-generic features are captured and integrated with task-specific features through a gating mechanism, allowing TS-Net to effectively share knowledge across tasks. Extensive experiments demonstrate that TS-Net exhibits superior performance across a range of tasks, including detection, segmentation, and image reconstruction. Zhaolin Wan, Penghong Wang, Xiaopeng Fan 0001 |
ICASSP | 3 |
| 2025 | Multi-Timescale Motion-Decoupled Spiking Transformer for Audio-Visual Zero-Shot LearningabstractAudio-visual zero-shot learning (ZSL) has been extensively researched for its capability to classify video data from unseen classes during training. Nevertheless, current methodologies often struggle with background scene biases and inadequate motion detail. This paper proposes a novel dual-stream Multi-Timescale Motion-Decoupled Spiking Transformer (MDST++), which decouples contextual semantic information and sparse dynamic motion information. The recurrent joint learning unit is proposed to extract contextual semantic information and capture joint knowledge across various modalities to understand the environment of actions. By converting RGB images to events, our method captures motion information more accurately and mitigates background scene biases. Moreover, we introduce a discrepancy analysis block to model audio motion information. To enhance the robustness of SNNs in extracting temporal and motion cues, we dynamically adjust the threshold of Leaky Integrate-and-Fire neurons based on global motion and contextual semantic information. Our experiments validate the effectiveness of MDST++, demonstrating their consistent superiority over state-of-the-art methods on mainstream benchmarks. Additionally, incorporating motion and multi-timescale information significantly improves HM and ZSL accuracy by 26.2% and 39.9%. Wenrui Li 0001, Penghong Wang, Wangmeng Zuo, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | DV-Hop Localization Based on Distance Estimation Using Multinode and Hop Loss in IoTabstractSensor location awareness is a critical issue in internet of things applications. For more accurate location estimation, the two issues should be considered extensively: 1) how to sufficiently utilize the connection information between multiple nodes and 2) how to select a suitable solution from multiple solutions obtained by the Euclidean distance loss. In this paper, a DV-Hop localization based on the distance estimation using multi-node (DEMN) and the hop loss in WSNs is proposed to address the two issues. In DEMN, when multiple anchor nodes can detect an unknown node, the distance expectation between the unknown node and an anchor node is calculated using the cross domain information and is considered as the expected distance between them, which narrows the search space. When minimizing the traditional Euclidean distance loss, multiple solutions may exist. To select a suitable solution, the hop loss is proposed, which minimizes the difference between the real and its predicted hops. Finally, the Euclidean distance loss calculated by the DEMN and the hop loss are embedded into the multi-objective optimization algorithm. The experimental results show that the proposed method gains 86.11% location accuracy in the randomly distributed network, which is 6.05% better than the DEM-DV-Hop, while DEMN and the hop loss can contribute 2.46% and 3.41%, respectively. Penghong Wang, Wenrui Li 0001, Xiaopeng Fan 0001, Debin Zhao |
IEEE Internet Things J. | 1 |
| 2024 | 3D many-objective DV-hop localization model with NSGA3
Penghong Wang, Hangjuan Li, Xingjuan Cai |
Soft Comput. | 1 |
| 2024 | Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot LearningabstractThe spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling SNNs (binary spike sequences) with transformers (float-point sequences) to jointly explore the temporal-semantic information still facing challenges. In this paper, we introduce a novel Spiking Tucker Fusion Transformer (STFT) for audio-visual zero-shot learning (ZSL). The STFT leverage the temporal and semantic information from different time steps to generate robust representations. The time-step factor (TSF) is introduced to dynamically synthesis the subsequent inference information. To guide the formation of input membrane potentials and reduce the spike noise, we propose a global-local pooling (GLP) which combines the max and average pooling operations. Furthermore, the thresholds of the spiking neurons are dynamically adjusted based on semantic and temporal cues. Integrating the temporal and semantic information extracted by SNNs and Transformers are difficult due to the increased number of parameters in a straightforward bilinear model. To address this, we introduce a temporal-semantic Tucker fusion module, which achieves multi-scale fusion of SNN and Transformer outputs while maintaining full second-order interactions. Our experimental results demonstrate the effectiveness of the proposed approach in achieving state-of-the-art performance in three benchmark datasets. The harmonic mean (HM) improvement of VGGSound, UCF101 and ActivityNet are around 15.4%, 3.9%, and 14.9%, respectively. Wenrui Li 0001, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Reservoir Computing Transformer for Image-Text RetrievalabstractAlthough the attention mechanism in transformers has proven successful in image-text retrieval tasks, most transformer models suffer from a large number of parameters. Inspired by brain circuits that process information with recurrent connected neurons, we propose a novel Reservoir Computing Transformer Reasoning Network (RCTRN) for image-text retrieval. The proposed RCTRN employs a two-step strategy to focus on feature representation and data distribution of different modalities respectively. Specifically, we send visual and textual features through a unified meshed reasoning module, which encodes multi-level feature relationships with prior knowledge and aggregates the complementary outputs in a more effective way. The reservoir reasoning network is proposed to optimize memory connections between features at different stages and address the data distribution mismatch problem introduced by the unified scheme. To investigate the significance of the low power dissipation and low bandwidth characteristics of RRN in practical scenarios, we deployed the model in the wireless transmission system, demonstrating that RRN's optimization of data structures also has a certain robustness against channel noise. Extensive experiments on two benchmark datasets, Flickr30K and MS-COCO, demonstrate the superiority of RCTRN in terms of performance and low-power dissipation compared to state-of-the-art baselines. Wenrui Li 0001, Zhengyu Ma, Liang-Jian Deng, Penghong Wang, Jinqiao Shi, Xiaopeng Fan 0001 |
ACM Multimedia | 4 |
| 2022 | Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel CodingabstractAudio-visual parsing (AVP) is a newly emerged multimodal perception task, which detects and classifies audio-visual events in video. However, most existing AVP networks only use a simple attention mechanism to guide audio-visual multimodal events, and are implemented in a single end. This makes it unable to effectively capture the relationship between audio-visual events, and is not suitable for implementation in the network transmission scenario. In this paper, we focus on these problems and propose a distributed audio-visual parsing network (DAVPNet) based on multimodal transformer and deep joint source channel coding (DJSCC). Multimodal transformers are used to enhance the attention calculation between audio-visual events, and DJSCC is used to apply DAVP tasks to network transmission scenarios. Finally, the Look, Listen, and Parse (LLP) dataset is used to test the algorithm performance, and the experimental results show that the DAVPNet has superior parsing performance. Penghong Wang, Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001 |
ICASSP | 1 |
| 2020 | A Gaussian error correction multi-objective positioning model with NSGA-IIabstractSummary Distance vector‐hop (DVHop), as a range‐independent positioning algorithm, is a significant positioning method in wireless sensor networks (WSNs). It is composed of three parts, including connectivity detection, distance estimation, and position estimation. However, this simple positioning method results in a larger positioning error. Therefore, to enhance the positioning precision, this paper investigates the characteristic of error distribution between the estimated and real distance in the DVHop algorithm and reveals that the error is subjecting to the Gaussian distribution, N∼(0,1/3CR). Furthermore, to improve positioning accuracy, we propose a Gaussian error correction multi‐objective positioning model with non‐dominated sorting (NSGA‐II), which named GGAII‐DVHop. Finally, this model is tested on three complex network topologies, and the results demonstrate that it is significantly superior to other four algorithms in both positioning precision and robustness. Penghong Wang, Jianrou Huang, Zhihua Cui, Jinjun Chen |
Concurr. Comput. Pract. Exp. | 1 |
| 2020 | Weight convergence analysis of DV-hop localization algorithm with GA
Xingjuan Cai, Penghong Wang, Zhihua Cui, Wensheng Zhang 0002, Jinjun Chen |
Soft Comput. | 2 |
| 2019 | Malicious code detection based on CNNs and multi-objective algorithm
Zhihua Cui, Penghong Wang, Xingjuan Cai, Wensheng Zhang 0002 |
J. Parallel Distributed Comput. | 3 |