Jiahui Li 0006

dblp:153/2952-6 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021
YearPublicationVenuePosition
2025 Spatial Frequency Interleaving Residual Autoencoder for Indoor Radio Map Reconstruction
abstract
Indoor radio maps with frequency domain data are difficult to reconstruct when only limited measurements at a few locations are available. Naive convolutional neural networks suffer from flawed structures in the frequency domain when predicting these radio maps, resulting in overly smoothed predictions. We propose a Spatial Frequency Interleaving Residual Autoencoder (SFIRA) architecture to tackle this problem, along with a Procedural Radio Map Generation (PRMG) method to address the data deficiency of indoor radio maps. Experimental results indicate that the proposed architecture achieves lower Normalized Root Mean Square Error (NRMSE). Visualization of the reconstruction further suggests that the proposed methods are effective.
Shitong Chai, Jiahui Li 0006, Mengyao Ma, Junwen Xie, Xiaopeng Fan 0001, Xianqi Zhang
ICASSP2
2022 Distributed Audio-Visual Parsing Based On Multimodal Transformer and Deep Joint Source Channel Coding
abstract
Audio-visual parsing (AVP) is a newly emerged multimodal perception task, which detects and classifies audio-visual events in video. However, most existing AVP networks only use a simple attention mechanism to guide audio-visual multimodal events, and are implemented in a single end. This makes it unable to effectively capture the relationship between audio-visual events, and is not suitable for implementation in the network transmission scenario. In this paper, we focus on these problems and propose a distributed audio-visual parsing network (DAVPNet) based on multimodal transformer and deep joint source channel coding (DJSCC). Multimodal transformers are used to enhance the attention calculation between audio-visual events, and DJSCC is used to apply DAVP tasks to network transmission scenarios. Finally, the Look, Listen, and Parse (LLP) dataset is used to test the algorithm performance, and the experimental results show that the DAVPNet has superior parsing performance.
Penghong Wang, Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001
ICASSP2
2022 Constellation Design for Deep Joint Source-Channel Coding
abstract
Deep learning-based joint source-channel coding (JSCC) has shown excellent performance in image and feature transmission. However, the output values of the JSCC encoder are continuous, which makes the constellation of modulation complex and dense. It is hard and expensive to design radio frequency chains for transmitting such full-resolution constellation points. In this paper, two methods of mapping the full-resolution constellation to finite constellation are proposed for real system implementation. The constellation mapping results of the pro- posed methods correspond to regular constellation and irregular constellation, respectively. We apply the methods to existing deep JSCC models and evaluate them on AWGN channels with different signal-to-noise ratios (SNRs). Experimental results show that the proposed methods outperform the traditional uniform quadrature amplitude modulation (QAM) constellation mapping method by only adding a few additional parameters.
Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001
IEEE Signal Process. Lett.2
2021 SNR-Adaptive Deep Joint Source-Channel Coding for Wireless Image Transmission
abstract
Considering the problem of joint source-channel coding (JSCC) for multi-user transmission of images over noisy channels, an autoencoder-based novel deep joint source-channel coding scheme is proposed in this paper. In the proposed JSCC scheme, the decoder can estimate the signal-to-noise ratio (SNR) and use it to adaptively decode the transmitted image. Experiments demonstrate that the proposed scheme achieves impressive results in adaptability for different SNRs and is robust to the noise in the SNR estimation of the decoder. To the best of our knowledge, this is the first deep JSCC scheme that focuses on the adaptability for different SNRs and can be applied to multi-user scenarios.
Mingze Ding, Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001
ICASSP2
2021 MSFC: Deep Feature Compression in Multi-Task Network
abstract
With the remarkable success of deep learning, a novel AI-deployment strategy on mobile devices called collaborative intelligence (CI) is proposed recently, which can greatly improve the efficiency of neural network by distributing work-loads between mobile devices and the cloud. In order to reduce transmission overhead, feature maps obtained from mobile devices need to be compressed before being transmitted to the cloud. In this paper, we propose a multi-scale feature compression (MSFC) framework for applying complex multi-task learning network in CI deployment scenarios, which consists of a multi-scale feature fusion (MSFF) module, a single-stream feature codec (SSFC) and a multi-scale feature reconstruction (MSFR) module. When applied to the popular multi-task network Mask R-CNN, experimental results show that with less than 2% accuracy degradation, the proposed MSFC can compress the 32-bit floating point feature to 0.012 bits on average.
Mengyao Ma, Jiahui Li 0006, Xiaopeng Fan 0001
ICME4
2021 Deep Joint Source-Channel Coding for Multi-Task Network
abstract
Multi-task learning (MTL) is an efficient way to improve the performance of related tasks by sharing knowledge. However, most existing MTL networks run on a single end and are not suitable for collaborative intelligence (CI) scenarios. In this work, we propose an MTL network with a deep joint source-channel coding (JSCC) framework, which allows operating under CI scenarios. We first propose a feature fusion based MTL network (FFMNet) for joint object detection and semantic segmentation. Compared with other MTL networks, FFMNet gets higher performance with fewer parameters. Then FFMNet is split into two parts, which run on a mobile device and an edge server respectively. The feature generated by the mobile device is transmitted through the wireless channel to the edge server. To reduce the transmission overhead of the intermediate feature, a deep JSCC network is designed. By combining two networks together, the whole model achieves 512 compression for the intermediate feature and a performance loss within 2% on both tasks. At last, by training with noise, the FFMNet with JSCC is robust to various channel conditions and outperforms the separate source and channel coding scheme.
Jiahui Li 0006, Mengyao Ma, Xiaopeng Fan 0001
IEEE Signal Process. Lett.3