EDBT 2026 Demo / reviewers in the wild / expert
Wenlong Dong
dblp:85/4198
· DBLP profile ↗
18ranked-venue papers
7as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semantic-Guided Visual Byte-Pair Encoding for Unified Autoregressive Multimodal Modeling
Wenlong Dong, Lijian Gao, Qirong Mao |
ICIC (12) | 2 |
| 2026 | Keyword Mamba: Spoken keyword spotting with state space models
Hanyu Ding, Wenlong Dong, Qirong Mao |
Comput. Speech Lang. | 2 |
| 2026 | HCATRE-AVAD: Hierarchical cross-alignment and temporal relational encoding for weakly supervised audio-visual anomaly detection
Nuku Atta Kordzo Abiew, Lijian Gao, Godbless Mensah, Wenlong Dong, Qirong Mao |
Image Vis. Comput. | 4 |
| 2026 | A blockchain-assisted lightweight authentication scheme for smart home environments
Xiujun Wang, Wenlong Dong, Juyan Li |
J. Netw. Comput. Appl. | 2 |
| 2026 | Adaptive Key Role Guided Hierarchical Relation Inference for Enhanced Group-Level Emotion RecognitionabstractIn this paper, we propose a novel hierarchical relational network, termed Key Role Guided Hierarchical Relation Inference (KR-HRI), for enhanced group-level emotion recognition (GER). Unlike existing methods that adopt a coarse-grained approach to model interactions among all individuals, our approach adaptively identifies and emphasizes key individuals who play a crucial role in conveying group-level emotions. By integrating coarse-grained relationship modeling with fine-grained key individual enhancement and leveraging global scene information, our method effectively refines discriminative feature generation while minimizing irrelevant interference. We introduce a Multi-branch Interaction Module (MIM) to dynamically fuse features from both the global scene and local individual branches using a localized mask integration strategy. This comprehensive approach enhances the interaction between global and local features, resulting in robust group-level emotion representations. Extensive experiments on three widely adopted GER datasets demonstrate that our framework consistently outperforms state-of-the-art methods, validating the effectiveness and robustness of our proposed approach. Qing Zhu 0002, Qirong Mao, Wenlong Dong, Xiuyan Shao, Xiaohua Huang 0003, Wenming Zheng |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Key Clues Guided Video Character Social Relationship Recognition Enhanced by LLMabstractVideo Character Social Relationship Recognition (VCSRR) requires a comprehensive consideration about spatio-temporal and multi-modal clues in videos. Most existing methods mainly focus on integrating multi-modal clues and modeling interactions among characters. However, they fail to discover key clues in the complex video data or fully understand the clues related to social relationships. In this article, we propose a novel Large Language Model Enhanced Key Clues Selection (LE-KCS) framework to address the aforementioned issues. The core of LE-KCS is to mine multi-scale key clues from the perspectives of time, space and multi-modality, then transfer the knowledge about social relationships of the Large Language Model to VCSRR for understanding the selected clues. We evaluated LE-KCS on the MovieGraphs dataset and the experimental results indicate that our proposed LE-KCS achieves state-of-the-art performance. Wenlong Dong, Qing Zhu 0002, Qirong Mao |
ICASSP | 1 |
| 2025 | RTAGrasp: Learning Task-Oriented Grasping from Human Videos via Retrieval, Transfer, and AlignmentabstractTask-oriented grasping (TOG) is crucial for robots to accomplish manipulation tasks, requiring the determination of TOG positions and directions. Existing methods either rely on costly manual TOG annotations or only extract coarse grasping positions or regions from human demonstrations, limiting their practicality in real-world applications. To address these limitations, we introduce RTAGrasp, a Retrieval, Transfer, and Alignment framework inspired by human grasping strategies. Specifically, our approach first effortlessly constructs a robot memory from human grasping demonstration videos, extracting both TOG position and direction constraints. Then, given a task instruction and a visual observation of the target object, RTAGrasp retrieves the most similar human grasping experience from its memory and leverages semantic matching capabilities of vision foundation models to transfer the TOG constraints to the target object in a training-free manner. Finally, RTAGrasp aligns the transferred TOG constraints with the robot's action for execution. Evaluations on the public TOG benchmark, TaskGrasp dataset, show the competitive performance of RTAGrasp on both seen and unseen object categories compared to existing baseline methods. Real-world experiments further validate its effectiveness on a robotic arm. Our code, appendix, and video are available at https://sites.google.com/view/rtagrasp/home. Wenlong Dong, Dehao Huang, Jiangshan Liu, Chao Tang 0001, Hong Zhang 0013 |
ICRA | 1 |
| 2025 | Leveraging Semantic and Geometric Information for Zero-Shot Robot-to-Human HandoverabstractHuman-robot interaction (HRI) encompasses a wide range of collaborative tasks, with handover being one of the most fundamental. As robots become more integrated into human environments, the potential for service robots to assist in handing objects to humans is increasingly promising. In robot-to-human (R2H) handover, selecting the optimal grasp is crucial for success, as it requires avoiding interference with the human's preferred grasp region and minimizing intrusion into their workspace. Existing methods either inadequately consider geometric information or rely on data-driven approaches, which often struggle to generalize across diverse objects. To address these limitations, we propose a novel zero-shot system that combines semantic and geometric information to generate optimal handover grasps. Our method first identifies grasp regions using semantic knowledge from vision-language models (VLMs) and, by incorporating customized visual prompts, achieves finer granularity in region grounding. A grasp is then selected based on grasp distance and approach angle to maximize human ease and avoid interference. We validate our approach through ablation studies and real-world comparison experiments. Results demonstrate that our system improves handover success rates and provides a more user-preferred interaction experience. Videos, appendixes and more are available at https://sites.google.com/view/vlm-handover. Jiangshan Liu, Wenlong Dong, Jiankun Wang 0001, Max Q.-H. Meng |
ICRA | 2 |
| 2025 | HGDiffuser: Efficient Task-Oriented Grasp Generation via Human-Guided Grasp Diffusion ModelsabstractTask-oriented grasping (TOG) is essential for robots to perform manipulation tasks, requiring grasps that are both stable and compliant with task-specific constraints. Humans naturally grasp objects in a task-oriented manner to facilitate subsequent manipulation tasks. By leveraging human grasp demonstrations, current methods can generate high-quality robotic parallel-jaw task-oriented grasps for diverse objects and tasks. However, they still encounter challenges in maintaining grasp stability and sampling efficiency. These methods typically rely on a two-stage process: first performing exhaustive task-agnostic grasp sampling in the 6-DoF space, then applying demonstration-induced constraints (e.g., contact regions and wrist orientations) to filter candidates. This leads to inefficiency and potential failure due to the vast sampling space. To address this, we propose the Human-guided Grasp Diffuser (HGDiffuser), a diffusion-based framework that integrates these constraints into a guided sampling process. Through this approach, HGDiffuser directly generates 6-DoF task-oriented grasps in a single stage, eliminating exhaustive task-agnostic sampling. Furthermore, by incorporating Diffusion Transformer (DiT) blocks as the feature backbone, HGDiffuser improves grasp generation quality compared to MLP-based methods. Experimental results demonstrate that our approach significantly improves the efficiency of task-oriented grasp generation, enabling more effective transfer of human grasping strategies to robotic systems. To access the source code and supplementary videos, visit https://sites.google.com/ view/hgdiffuser. Dehao Huang, Wenlong Dong, Chao Tang 0001, Hong Zhang 0013 |
IROS | 2 |
| 2025 | StyU-STD: Style-Diverse Sample Generation from Unlabeled Data for Query-by-Example Spoken Term DetectionabstractIn recent years, query-by-example spoken term detection (QbE-STD) techniques have made significant progress in detection accuracy and speed. However, this task also encounters situations where labeled data is scarce or even nonexistent, with only unlabeled data available. Although some solutions exist, they still struggle to effectively handle highly variable speech, especially when it comes to differing styles. To address this issue, we propose a self-supervised learning method named Style-diverse sample generation from Unlabeled data for query-by-example Spoken Term Detection (StyU-STD). The core idea is to generate samples with the same content but different styles for learning. Specifically, we randomly extract segments from the speech to be tested as positive samples, while segments randomly extracted from other speech data are labeled as negative samples of the speech to be tested. In addition, various transformations are applied to alter the style of both positive and negative samples while preserving their original content. Then, the generated sample pairs are used to train the Style Suppressed Convolutional Network, which focuses more on content-related information in speech and effectively reduces the interference caused by style differences. The experimental results show that, across multiple datasets, our method outperforms existing methods, achieving higher accuracy and robustness. Hanyu Ding, Lijian Gao, Wenlong Dong, Xiangrui Li, Qirong Mao |
SMC | 3 |
| 2025 | A secure lightweight identity authentication and key agreement scheme for internet of drones
Wenlong Dong, Xiujun Wang, Juyan Li |
Comput. Networks | 1 |
| 2025 | EMP3D: an emergency medical procedures 3D dataset with pose and shape
Hehao Bao, Keying Du, Xinyi Su, Jianqi Fan, Wenlong Dong |
Frontiers Comput. Sci. | 6 |
| 2025 | FoundationGrasp: Generalizable Task-Oriented Grasping With Foundation ModelsabstractTask-oriented grasping (TOG), which refers to synthesizing grasps on an object that are configurationally compatible with the downstream manipulation task, is the first milestone towards tool manipulation. Analogous to the activation of two brain regions responsible for semantic and geometric reasoning during cognitive processes, modeling the intricate relationship between objects, tasks, and grasps necessitates rich semantic and geometric prior knowledge about these elements. Existing methods typically restrict the prior knowledge to a closed-set scope, limiting their generalization to novel objects and tasks out of the training set. To address such a limitation, we propose FoundationGrasp, a foundation model-based TOG framework that leverages the open-ended knowledge from foundation models to learn generalizable TOG skills. Extensive experiments are conducted on the contributed Language and Vision Augmented TaskGrasp (LaViA-TaskGrasp) dataset, demonstrating the superiority of FoundationGrasp over existing methods when generalizing to novel object instances, object classes, and tasks out of the training set. Furthermore, the effectiveness of FoundationGrasp is validated in real-robot grasping and manipulation experiments on a 7-DoF robotic arm. Our code, data, appendix, and video are publicly available athttps://sites.google.com/view/foundationgrasp. Note to Practitioners—This research is motivated by the challenge of generalizable task-oriented grasping skill learning. Solving such a challenge could significantly improve the robot’s level of automation and intelligence in tool manipulation for household and industrial tasks. Existing methods struggle with handling unseen objects and tasks in dynamic, open-world environments. To overcome this limitation, we propose to leverage the open-ended knowledge from foundation models to improve the generalization capabilities of existing TOG methods. This way, the robot can perform TOG w.r.t. unseen objects and tasks, facilitating downstream tool manipulation. Overall, this research has broad applicability to various scenarios involving tool manipulation, such as cleaning kitchenware and assembling parts in industrial contexts. Chao Tang 0001, Dehao Huang, Wenlong Dong, Ruinian Xu, Hong Zhang 0013 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | MLETune: Streamlining Database Knob Tuning via Multi-LLMs Experts Guided Deep Reinforcement LearningabstractAutomatic knob tuning has emerged as a critical field of study within database optimization, focusing on simplifying the configuration of database parameters to boost performance, particularly in the realm of advanced modern database management systems with myriad adjustable knobs. The primary challenge revolves around identifying the ideal knob configurations that can markedly enhance system efficiency. Various machine learning techniques have been devised to automate this tuning process. Nevertheless, these methods frequently entail running extensive workloads, resulting in significant time and resource consumption. This inefficiency arises from their reliance on runtime feedback or the limited exploitation of domain knowledge.To overcome these limitations, we propose MLETune, a novel deep reinforcement learning-based approach guided by multi-Large Language Models (LLMs) experts. Our method leverages a remix retrieval-augmented generation algorithm to harness knowledge and distill expert guidance effectively. Additionally, we utilize a genetic algorithm for coarse-grained exploration based on system and query-level knob knowledge to expedite the cold start process in deep reinforcement learning. By classifying and compressing metrics and optimizing tuning knobs based on workload and knob-level insights, we aim to reduce the search space efficiently. In addition, adopting a delayed update strategy helps mitigate the training time required for the deep reinforcement learning model. Our extensive experiments demonstrate that MLETune outperforms existing methods by identifying superior configurations in significantly less time, showing an average improvement of ${6x}$, along with achieving up to a $26 \%$ performance improvement. Wenlong Dong, Wei Liu 0279, Mengshu Hou, Shuhuan Fan |
ICPADS | 1 |
| 2023 | A Specific Emitter Identification Method Based on Time-Frequency Feature ExtractionabstractWith the rapid growth of the Internet of Things (IoT), fundamental security measures of wireless networks have become a basic requirement. Aiming at the identification of wireless transmitters with the same parameters, this paper proposes a specific emitter identification (SEI) method based on time-frequency feature extraction. Received signals go through preprocessing, i.e., multipath effect estimation and Doppler frequency compensation, to mitigate the channel effect. Then the time-frequency spectrum is generated and a time-frequency feature extraction network is constructed to achieve feature extraction and identification task. Real-world data are used to verify the effectiveness of the proposed method. The overall identification accuracy for stationary emitters reaches 92.9%. Besides, the proposed preprocessing method improves moving emitter identification accuracy by 12%. Wenlong Dong, Yuqi Wang 0002, Guangcai Sun, Mengdao Xing |
IGARSS | 1 |
| 2022 | Synthetic Aperture Passive Localization for Frequency Hopping SignalabstractFrequency hopping (FH) signal is one of the research hotspots of passive positioning. Aiming at the problem of FH signal localization, this paper proposes a synthetic aperture passive positioning method. The method estimates and compensates for the baseband modulation of the received signal. Then the received signal vectors are arranged into a two-dimensional matrix. The Doppler frequency of each pulse is compensated by the Doppler frequency compensate matrix. The cost function is constructed by a two-dimensional focus of the received signal, and the emitter position is directly obtained through a gird search. Simulation and experimental data verify the effectiveness of the proposed method. Wenlong Dong, Yuqi Wang 0002, Guangcai Sun, Mengdao Xing, Xiaoniu Yang |
IGARSS | 1 |
| 2000 | Postprocessing of Compressed 3D Graphic Data
Ka Man Cheang, Wenlong Dong, Jiankun Li, C.-C. Jay Kuo |
J. Vis. Commun. Image Represent. | 2 |
| 1999 | Refinement of 3D Meshes by Selective SubdivisionabstractAn adaptive subdivision method is proposed in this work for automatic post-processing of a 3D graphic model of coarse resolution. The method is an improved version of the Modified Butterfly Scheme (MBS) developed by Zorin et al. The main contribution of this work is to exploit the local smoothness information of a surface for adaptive refinement of a coarse 3D graphic model. With this approach, we can avoid unnecessary subdivision in relative smooth legions. The new algorithm not only reduces the computational complexity but also reduces the storage space. It is demonstrated via experiments that a more visual-pleasing representation can be obtained. Wenlong Dong, Jiankun Li, C.-C. Jay Kuo |
ICIP (4) | 1 |