Hongyu Deng

dblp:185/3905 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Dual-Space Optimization for data-free model merging
Hongyu Deng, Zhengde Tian
Eng. Appl. Artif. Intell.1
2025 Poster: Radar-Enhanced Robotic Material Perception with Vision Language Models
abstract
Robotic perception has been significantly advanced by integrating vision with additional sensing modalities such as acoustic and tactile sensors. However, existing methods largely emphasize external object properties, including appearance and geometry, while neglecting internal material attributes (e.g., composition) that are crucial for robust and reliable robotic manipulation. In this work, we propose augmenting robotic systems with radar sensing and introduce CRMaterial, a new camera-radar fusion framework powered by vision language models (VLMs) for accurate object material identification. Preliminary experiments demonstrate that our system improves material identification accuracy by 2.5× compared to a camera-only baseline.
Hongyu Deng, Jiangyou Zhu, He Henry Chen
MobiCom1
2025 FuseGrasp: Radar-Camera Fusion for Robotic Grasping of Transparent Objects
abstract
Transparent objects are prevalent in everyday environments, but their distinct physical properties pose significant challenges for camera-guided robotic arms. Current research is mainly dependent on camera-only approaches, which often falter in suboptimal conditions, such as low-light environments. In response to this challenge, we present FuseGrasp, the first radar-camera fusion system tailored to enhance the transparent objects manipulation. FuseGrasp exploits the weak penetrating property of millimeter-wave (mmWave) signals, which causes transparent materials to appear opaque, and combines it with the precise motion control of a robotic arm to acquire high-quality mmWave radar images of transparent objects. The system employs a carefully designed deep neural network to fuse radar and camera imagery, thereby improving depth completion and elevating the success rate of object grasping. Nevertheless, training FuseGrasp effectively is non-trivial, due to limited radar image datasets for transparent objects. We address this issue utilizing large RGB-D dataset, and propose an effective two-stage training approach: we first pre-train FuseGrasp on a large public RGB-D dataset of transparent objects, then fine-tune it on a self-built small RGB-D-Radar dataset. Furthermore, as a byproduct, FuseGrasp can determine the composition of transparent objects, such as glass or plastic, leveraging the material identification capability of mmWave radar. This identification result facilitates the robotic arm in modulating its grip force appropriately. Extensive testing reveals that FuseGrasp significantly improves the accuracy of depth reconstruction and material identification for transparent objects. Moreover, real-world robotic trials have confirmed that FuseGrasp markedly enhances the handling of transparent items.
Hongyu Deng, Tianfan Xue, He Henry Chen
IEEE Trans. Mob. Comput.1
2025 SRAL-IRS: Swift, Robust, and Accurate IRS-Aided Localization With COTS WiFi
abstract
Intelligent Reflecting Surface (IRS) has emerged as a crucial technology for indoor localization in future wireless networks. However, existing IRS-aided systems that utilize commercial off-the-shelf (COTS) WiFi devices face challenges due to two key factors. First, the passive reflective nature of IRS generates relatively weak reflected signals, which can be easily drowned out by multipath interference and noise. Secondly, existing works require a large number of IRS codebooks for fine-grained scanning of the entire area, resulting in significant time requirements. This paper introduces SRAL-IRS, which enables swift, robust, and accurate IRS-aided indoor localization using commercial WiFi devices. Through intelligent codebook design and theoretical derivation, we reveal the theoretical relationship between the IRS codebook modifications, target position, and variations in the received signal. Subsequently, SRAL-IRS mitigates the impact of environmental and hardware noise by carefully selecting and combining WiFi subcarriers. Finally, by utilizing the received data from multiple codebooks and subcarriers and leveraging the orthogonality between the noise subspace and the IRS-reflected signal subspace, we can achieve precise localization even with a limited number of IRS codebooks. Real-world implementation of SRAL-IRS, utilizing our designed IRS prototype and COTS WiFi devices, validates its feasibility and effectiveness.
Dongheng Zhang, Hongyu Deng, Xuecheng Xie, Fengquan Zhan, Yan Chen 0007
IEEE Trans. Wirel. Commun.3
2024 Practical Challenge and Solution for IRS-Aided Indoor Localization System
abstract
Intelligent reflecting surfaces (IRS) is a novel integrated sensing and communication technology that can manipulate the propagation of wireless signals. However, existing IRS-based sensing systems require directional antennas for signal transmission, incompatible with commercial WiFi devices. This paper proposes an IRS-aided localization system using omnidirectional antennas and reveals two critical challenges in practical deployment. First, the accurate distance between the IRS and the transmitter is needed for IRS codebook design, but practice measurements invariably introduce centimeter-level bias, which seriously affects localization accuracy. We derive a linear relationship between measurement bias and localization error for calibration. Second, only relative IRS phase change under different bias voltages can be measurable, not the absolute phase offset, introducing an unknown fixed phase offset in reflections. We solve this challenge by eliminating the signals that are not related to the IRS. Experiments validate the proposed calibration techniques, proving that our system achieves high-precision passive localization.
Dongheng Zhang, Hongyu Deng, Fengquan Zhan, Yan Chen 0007
ICASSP3
2024 LLM for Complex Signal Processing in FPGA-based Software Defined Radios: A Case Study on FFT
abstract
This paper investigates the potential of large language models (LLMs) in accelerating the development of complex signal-processing algorithms on field-programmable gate arrays (FPGAs) for software-defined radio (SDR) systems. Using the Fast Fourier Transform (FFT) algorithm as a case study, we identify two common challenges in applying LLMs to realize intricate wireless communication algorithms on FPGA: 1) handling convoluted mathematical problems and 2) scheduling the execution of sub-modules within the hardware structure. To overcome the first problem, we adapt the chain-of-thought (CoT) prompting technique with a length-limit strategy to enhance the LLM’s Verilog writing performance. To handle the second problem, we develop a novel iterative in-context learning (IICL) prompting scheme that utilizes the iterative structure within the FFT module to perform in-context learning (ICL). These efforts significantly reduce the LLM’s error rate in completing the FFT implementation task and make possible the successful generation of a 64-point FFT module in Verilog, marking a significant milestone as the first LLM-written complex signal-processing algorithm for wireless communication on FPGA.
Yuyang Du 0001, Hongyu Deng, Soung Chang Liew, Yulin Shao, Kexin Chen 0003, He Henry Chen
VTC Fall2
2024 Practical Passive Indoor Localization With Intelligent Reflecting Surface
abstract
Intelligent reflecting surface has gained significant attention for supporting integrated sensing and communication (ISAC) by manipulating wireless signals. However, existing IRS-based sensing systems commonly utilize directional horn antennas for signal transmission, which is incompatible with commercial WiFi devices and contradicts the concept of ISAC. In this paper, we propose an IRS-aided localization system with omnidirectional antennas to overcome these limitations. We reveal two critical challenges for implementing such a system. First, precise distance between the IRS and transmitter is requisite for IRS codebook design, yet practical measurement introduces centimeter-scale biases, causing severe localization errors. We derive a linear relationship between measurement bias and localization error, which serves as the basis for a calibration method. Second, in practice, we can only obtain the relative phase change of the IRS under different bias voltages, but cannot measure the phase offset introduced by IRS at zero voltage. Consequently, an unknown fixed offset appears when calculating the channel response of the signal reflected by the IRS, which destroys subsequent IRS codebook design. We resolve this problem by eliminating the signals that are not related to the IRS. Extensive experiments validate the proposed calibration techniques and demonstrate that our system achieves accurate passive localization.
Dongheng Zhang, Hongyu Deng, Fengquan Zhan, Yan Chen 0007
IEEE Trans. Mob. Comput.3
2024 CDKM: Common and Distinct Knowledge Mining Network With Content Interaction for Dense Captioning
abstract
The dense captioning task aims at detecting multiple salient regions of an image and describing them separately in natural language. Although significant advancements in the field of dense captioning have been made, there are still some limitations to existing methods in recent years. On the one hand, most dense captioning methods lack strong target detection capabilities and struggle to cover all relevant content when dealing with target-intensive images. On the other hand, current transformer-based methods are powerful but neglect the acquisition and utilization of contextual information, hindering the visual understanding of local areas. To address these issues, we propose a common and distinct knowledge-mining network with content interaction for the task of dense captioning. Our network has a knowledge mining mechanism that improves the detection of salient targets by capturing common and distinct knowledge from multi-scale features. We further propose a content interaction module that combines region features into a unique context based on their correlation. Our experiments on various benchmarks have shown that the proposed method outperforms the current state-of-the-art methods.
Hongyu Deng, Yushan Xie, Qi Wang 0079, Weijian Ruan, Wu Liu 0005, Yong-Jin Liu 0001
IEEE Trans. Multim.1
2023 LCM-Captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text
Qi Wang 0079, Hongyu Deng, Zhenguo Yang, Yazhou Wang 0006, Gefei Hao
Neural Networks2
2023 AA-trans: Core attention aggregating transformer with information entropy selector for fine-grained visual classification
Qi Wang 0079, JianJun Wang, Hongyu Deng, Yazhou Wang 0006, Gefei Hao
Pattern Recognit.3
2021 SpeedNet: Indoor Speed Estimation With Radio Signals
abstract
Indoor human speed estimation is critical to in-home health monitoring of elderly people since it can provide the moving status of the human. Contactless indoor speed estimation with radio signals is challenging due to the complicated relationship between the speed of moving human and radio signals. In this article, we propose an indoor speed estimation framework, SpeedNet, to estimate the speed from the radio signals. Specifically, SpeedNet first extracts the dominant path signal reflected from the human through the beamforming technique. Then, SpeedNet obtains the doppler frequency shift (DFS) corresponding to the moving human by analyzing the short-time Fourier transform (STFT) spectrogram of the dominant path signal. Finally, SpeedNet trains a deep neural network composed of convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) to utilize the spatial and temporal features of DFS to estimate the speed of moving human. The experimental results show that SpeedNet can estimate the human moving speed with an average accuracy of 96.33% in a typical indoor environment, which is better than the state-of-the-art approaches.
Yan Chen 0007, Hongyu Deng, Dongheng Zhang, Yang Hu 0006
IEEE Internet Things J.2
2020 Estimating Indoor Human Speed via Radio Signals
abstract
Indoor human speed estimation, which can provide the moving status of human, is attracting considerable critical attention, especially in the field of in-home health monitoring of elderly people. Since the relationship between the moving of human and radio signal is very complicated, indoor speed estimation via radio signals is non-trivial and challenging. To address the challenge, in this paper, we propose a SpeedNet framework to estimate the speed of moving human from the radio signals. Specifically, SpeedNet first utilizes the beamforming technique to extract the dominant path signal reflected from individuals. Then, with short time Fourier transform (STFT), SpeedNet analyzes the spectrogram of the dominant path signal and obtains the doppler frequency shift (DFS) that corresponds to the moving human. Finally, SpeedNet exploits the spatial and temporal features of the DFS through a deep neural network, which consists of convolutional neural networks (CNN) and long short-term memory networks (LSTM), to estimate the speed of moving human. Extensive experiments show that compared with the state-of-the-art approaches, SpeedNet can achieve much better speed estimation performance with a mean absolute percentage error (MAPE) of 3.67% in a typical indoor environment.
Hongyu Deng, Dongheng Zhang, Yang Hu 0006, Yan Chen 0007
GLOBECOM1
2018 Cognitive radio : A method to achieve spectrum sharing in LTE-R system
abstract
In order to solve the problem of spectrum waste in the LTE-Railway (LTE-R) system, the paper uses Cognitive Radio (CR) to improve the ability of spectrum sharing on Vehicle-to- Ground communication. By constructing a novel Cognitive Radio Network (CRN) in LTE-R system, the Cognitive LTE-R eNodeB (C-eNodeB) can work with Vehicle Gateway (VG) and allocate idle and wasted spectrum resources to the passengers communicating devices to improve spectrum utilization of LTE-R, without impacting train-ground communication. Aiming at the novel CRN architecture, a C-eNodeB Queue Management Strategy (QMS) based on Type of Service (ToS) value priority is proposed to reduce the Real-Time (RT) service delay of Secondary Users (SU) caused by FIFO QMS. The simulation results show that the proposed CRN effectively improves the spectrum utilization of LTE-R system and the C-eNodeB QMS based on the ToS value priority significantly reduce the delay of RT business of passengers.
Hongyu Deng, Yiming Wang 0003, Cheng Wu 0001
NOMS1