Guanchong Niu

dblp:223/8319 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-0571-2571ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SAASN: A Spatially-Aware Adaptive Scheduling Network for Large-Scale Multi-UAV Tasking
Jiaqing Xiong, Xiaohui Li 0001, Feixiang Liu, Guanchong Niu
WCNC4
2026 DAM-SC: Text-Driven Semantic Communication for Ultra-Low Bitrate Video Transmission Using Dynamic Attribute Models
Longfei Zhou, Xiaohui Li 0001, Feixiang Liu, Guanchong Niu
WCNC4
2026 Mamba-HTEPA: A multi-branch structure framework for multimodal grading of meningiomas using Mamba
Zhuo Zhang 0020, Yihan Wen, Wendi Liang, Guanchong Niu, Quanfeng Ma
Expert Syst. Appl.5
2026 MSET: Multimodal Semantic-Enhanced Real-World Beam Prediction via Temporal Modeling With Visual Foundation Models
abstract
While machine learning (ML) has been explored for beam prediction, many methods remain constrained by single-modality inputs, shallow temporal modeling, and limited robustness to interference and domain shift. We introduce Multimodal Semantic-Enhanced Real-World Beam Prediction via Temporal Modeling with Visual Foundation Models (MSET), a multimodal framework that couples visual semantics with positional priors in a causal, lightweight design. A visual foundation model (VFM), instantiated as a Swin-Transformer, learns rich spatial and semantic cues from RGB images and Segment Anything Model (SAM)-derived region masks, and a lightweight ResNet-18 student distills this knowledge for efficient inference. Short-horizon dynamics are captured by a causal Temporal Convolutional Network (TCN) with an adaptive receptive field, whose volatility-driven depth gate expands context under motion spikes and contracts it in calm periods. On top of semantic-aware frame embeddings, a Temporally Aware Cross-Attention (TACA) aligns original and semantic-enhanced tokens, while a Mambaconditioned GPS prior, implemented via a selective state-space model (SSM), issues a location-conditioned single query with positional bias to attend the fused tokens. We further extend inference to dense deployments with multiple proximate candidates and address target selection under ambiguity. Experiments on DeepSense 6G dataset show consistent Top-k improvements across single/multi-target and day/night scenarios, indicating reduced sweep reliance and strong generalization under realistic V2I dynamics.
Feixiang Liu, Xiaohui Li 0001, Wenhui Gao, Jiaqing Xiong, Guanchong Niu, Chung Shue Chen
IEEE Internet Things J.5
2026 Rule-Semantic Generative Calibration Blur Detection for UAV Imagery
Yihan Wen, Zhuo Zhang 0020, Xianping Ma, Peipei Zhu, Jinglei Li, Guanchong Niu, Qiguang Miao
IEEE Trans. Circuits Syst. Video Technol.6
2025 SAW: Semantic-Aware WebRTC Transmission Using Diffusion-Based Scalable Video Coding
abstract
As video transmission systems expand into various complex scenarios, real-time video coding methods are essential for maintaining low latency and high perceptual quality across varying network conditions. In this work, we propose service-aware Web real-time communication (WebRTC), a semantic-assisted WebRTC system built on scalable video coding (SVC). Specifically, this system is structured with three layers: 1)$\mathcal {L}_{1}$extracts and down-samples semantic information at the encoder, employing a novel super-resolution (SR) method named BUS-DDIM at the decoder to enhance the transmission efficiency and machine vision recognition rate; 2)$\mathcal {L}_{2}$adaptively compresses high-quality video by discarding frames with little motion at the encoder to address latency issues under poor network conditions, and utilize the adjacent frame-guided denoised interpolation model called the adjacent frame-guided denoised diffusion implicit model for restoring the video; and 3)$\mathcal {L}_{3}$transmits high-quality video tailored for users with high-definition video requirements and favorable network conditions. These layers dynamically enhance the visual experience and ensure low latency across various network environments. Experiments are conducted on diverse videos to validate the effectiveness of the proposed framework. The performance evaluation under real-time scenarios indicates significant enhancements in video quality and transmission efficiency, showcasing compatibility and versatility across various applications.
Yihan Wen, Jinglei Li, Chung Shue Chen, Guanchong Niu
IEEE Internet Things J.6
2024 Semantic-Based Motion Detection Method for Unmanned Aerial Vehicle Data Transmission
abstract
Unmanned Aerial Vehicles (UAVs) are crucial for wireless network transmissions, particularly in challenging en-vironments of regular inspection. However, transmitting high-resolution video data from UAV s poses challenges due to limited resources and significant data volumes. Traditional video compression methods, removing redundant information with a single frame, suffer from quality loss as compression rates increase. To address these issues, we propose a novel framework, namely Semantic-based Motion Detection Compression (SMDC) to perform the video compression with high-quality resolution. The proposed framework incorporates Generative Diffusion Change Detection (TransC-GD-CD), a robust semantic-based change detection method, to accurately detect motion between adjacent video frames. Specifically, frames with slight motion are eliminated, thereby reducing network bandwidth requirements for the UAV inspection. Furthermore, a neural network-based interpolation technique is integrated to restore the information loss and ensure smooth playback. Experimental results show that SMDC outperforms traditional compression methods based on H.264, achieving higher video quality at matched bitrates. The exceptional performance of SMDC promises its potential as an effective solution for high-resolution video transmission in scenarios with limited bandwidth.
Yihan Wen, Feixiang Liu, Qi Cao 0001, Guanchong Niu
ICC4
2024 Enhanced Facial Restoration with Misinformation-Filtered Guide-Denoising Diffusion Probabilistic Models
abstract
Most of the existing generation models encounter notable challenges in complex scenes, particularly with inaccuracies in facial organs and textures that do not align with actual conditions. Traditional face restoration methods, heavily dependent on facial geometry and reference priors, often generate incorrect facial images that contribute misleading prior information. In this study, a Misinformation-Flitered GuideDenoising Diffusion Probabilistic Models (MF-GDDPM) is proposed to address these issue. Specifically, MF-GDDPM employs low-pass filtering to remove high-frequency details that contain misleading prior information. This process results in filtered low-dimensional facial contours that guide the diffusion model in generating high-quality facial images. To further enhance the fidelity of the generated results, a dualstream encoder within the Denoising Unet is constructed to process facial contours and high-dimensional details separately, while the Attention Feature Fusion (AFF) attention mechanism ensures the fidelity of image restoration. We have also incorporated the Natural Image Quality Evaluator (NIQE), a deep learning-based image quality assessment tool, into our framework as a novel loss function to crucially ensure the naturalness of restored images. Overall, the proposed method marks a significant improvement in generating accurate and clear facial images using diffusion models.
Wendi Liang, Yihan Wen, Jianuo Jiang, Tat-Ming Lok, Guanchong Niu
ICIP6
2024 Zero-Error Capacity of Broadcast Channels with Two Binary Outputs
abstract
This paper begins a systematic study of the zeroerror capacity problem of broadcast channels, where the message sent can be decoded by each receiver with zero error. A graph set is used to represent the broadcast channel. We particularly consider a set of two graphs, where each graph in the set contains only one edge. The corresponding zero-error capacity is determined in this paper.
Guanchong Niu, Yanlin Geng, Baoming Bai
ITW3
2023 Vision-Based Target Localization with Cooperative UAVs Towards Indoor Surveillance
abstract
Unmanned aerial vehicles (UAVs) have been widely adopted for a variety of civilian and military applications. Despite their many advantages, most UAVs are not suitable for indoor missions due to the lack of global position system (GPS) and the large size of UAV with diverse sensors. In this work, a lightweight and cost-effective indoor multi-UAV surveillance system is presented for accurate target localization. The proposed system employs a vision-based architecture, leveraging ORB-SLAM for self-localization and YOLOv3 for object detection. The triangulation method is employed for target positioning, offering reliable performance with negligible time delays and improved detection compared to depth camera approaches. After that, the Riccati observer is introduced to address the stochastic nature of the system. Through a series of experiments involving real UAVs, the system demonstrates its ability to accurately localize and track targets in indoor environments, updating the UAVs’ positions in real-time with optimal performance.
Guanchong Niu, Qi Cao 0001, Chung Shue Chen
VTC Fall1
2022 3D map reconstruction using a monocular camera for smart cities
abstract
Abstract Large-scale high-resolution three-dimensional (3D) maps play a vital role in the development of smart cities. In this work, a novel deep learning-based multi-view-stereo method is proposed for reconstructing the 3D maps in large-scale urban environments by exploiting a monocular camera. Compared with other existing works, the proposed method can perform 3D depth estimation more efficiently in terms of computational complexity and graphics processing unit memory usage. As a result, the proposed method can practically perform depth estimation for each pixel before generating 3D maps for even large-scale scenes. Extensive experiments on the well-known DTU dataset and real-life data collected on our campus confirm the good performance of the proposed method.
Taimeng Fu, Guanchong Niu, Zixiao Liu, Man-On Pun
J. Supercomput.3
2021 UAV-Enabled 3D Indoor Positioning and Navigation Based on VLC
abstract
The 3D indoor positioning and indoor navigation (IPIN) system is of great significance for promoting and expanding indoor intelligent services and applications. The rapid development of unmanned aerial vehicles (UAVs) has provided new opportunities in this field. However, in contrast to their outdoor applications, IPIN for UAVs is more challenging since the Global Positioning System (GPS) is in general inaccessible in indoor environments. In this work, we propose a UAV-enabled 3D IPIN system based on visible light communication (VLC). Firstly, a novel VLC-based indoor positioning scheme is developed using a fusion algorithm based on the dynamic time warping (DTW) method with visible light intensity sequence (VLIS) and inertial measurement unit (IMU) data. To reduce the workload of fingerprint measurements, we propose to modularize a floor site using a standard symmetric structure for VLC positioning. In this manner, the navigation can be achieved by recognizing the edge of each module. Furthermore, since the sampling frequency of IMU is much higher than that of VLIS, discrete Kalman filter (KF) is introduced to correct the location measured by IMU when VLIS is unavailable. A proof-of-concept IPIN prototype is constructed. Field experiments confirm the effectiveness of our proposed IPIN system.
Guanchong Niu, Man-On Pun, Chung Shue Chen
ICC1
2020 Magnetic Field Strength Sequence-based Indoor Localization Using Multi-level Link-node Models
abstract
This work investigates geomagnetism-based indoor localization by exploiting the magnetometers built-in smartphones. The main challenge arises from the fact that the localization accuracy is handicapped by the limited dimensionality of the magnetometer data. To cope with this problem, the magnetic field strength (MFS) sequence has been proposed to improve the localization accuracy. However, it remains an open challenge to derive the exact location through the MFS. In this work, a novel multi-level link-node model containing geometry and topology information is first proposed to construct the MFS sequence fingerprint database which can be easily constructed by crowdsourced technology. By exploiting this database, a hybrid approach combining dynamic time warping (DTW), pedestrian dead reckoning (PDR) and k-nearest neighbor (kNN) is developed to achieve efficient MFS sequence matching and accurate indoor localization. The proposed method provides a convenient and pervasive indoor localization solution that uses only the built-in sensors of smartphones. Without knowing the initial position, the user's location can be quickly determined. Extensive experimental results confirm that the proposed approach can achieve accurate and efficient MFS sequence matching results while providing accurate initial and end position estimation in the indoor environments.
Guanchong Niu, Man-On Pun
ICC2
2020 Joint Hybrid Precoding and Power Allocation via Rank-Constrained D.C. Programming
abstract
Hybrid analog and digital precoding have recently been proposed for massive multiple-input multiple-output (MIMO) systems. However, it is challenging to jointly optimize the design of precoding and power allocation as the optimization problem is highly non-convex. In this work, the non-convex problem is cast into the D.C. (difference of two convex functions) programming framework. To cope with the high dimensionality of the problem, an iterative rank-constrained D.C. programming technique is developed. As a result, the proposed algorithm is capable of directly maximizing the weighted sum-rate (WSR) of all users while taking into account the quality of service (QoS) requirement of each user. Furthermore, the monotonic convergence behavior of the proposed iterative algorithm is proved. Finally, simulation results confirm the effectiveness of the proposed iterative algorithm.
Guanchong Niu, Qi Cao 0001, Man-On Pun
ICC1
2018 Large-Area Super-Resolution 3D Digital Maps for Indoor and Outdoor Wireless Channel Modeling
abstract
This paper reports our recent work on creating the world's first super-resolution 3D digital maps for indoor and outdoor wireless channel modeling. By exploiting the recent technological breakthroughs in unmanned aerial vehicle (UAV), Light Detection and Ranging (LiDAR) and the Simultaneous Localization and Mapping (SLAM) technology, this work develops a surveying system prototype to produce super-resolution 3D digital maps of centimeter-level resolution for large areas covering both indoor and outdoor environments. In addition, the resulting maps are designed to precisely capture building shapes even for skyscrapers with wall material information. It is believed that these new maps can potentially revolutionize the network planning practice commonly performed in the telecommunication industry by providing highly accurate indoor and outdoor channel models. By exploiting these new channel models, wireless service operators will be able to better optimize their wireless networks with reduced necessity of labor-intensive drive tests.
Guanchong Niu, Man-On Pun
VTC Spring2