Gang Li 0020

dblp:62/2655-20 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0003-3141-4428ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Computer networks · 7 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Low-light image segmentation using sine-curve enhancement and cross-resolution attention mechanism
Linkang Xu, Gang Li 0020, Xiangxin Ji, Yue Song 0005
Eng. Appl. Artif. Intell.2
2026 RPE-IDet: Toward robust low-light object Detection via RGB-Pseudo-event Interactive feature fusion
Gang Li 0020, Jinping Zhang, Zhongpan Zhu, Mingke Gao
Neurocomputing1
2026 On the Stability of Spatially Distributed Cavity Laser and Boundary of Resonant Beam SLIPT
abstract
Spatially distributed cavity (SDC) lasers are a promising technology for simultaneous light information and power transfer (SLIPT), offering benefits such as increased mobility and intrinsic safety, which are advantageous for various Internet of Things (IoT) devices. However, achieving beam transmission over meter-level long working distances presents significant challenges from cavity stability constraints, manufacturing/ assembly tolerances, and diffraction losses. This paper conducts a theoretical investigation of the fundamental restrictions limiting long-range resonant beam generation. We investigate cavity stability and beam characteristics, and propose a binary-search-based Monte Carlo simulation algorithm as well as a linear approximation algorithm to quantify the maximum acceptable tolerances for stable operation. Numerical results indicate that the stable region contracts sharply as distance increases. For fixed-component systems, an acceptable tolerance of 0.01 mm restricts the achievable transmission distance to less than 2 m. To address this limitation, we also prove the feasibility of long-range beam formation using precision adjustable elements, paving the way for advanced engineering applications. Experimental results verified this assumption, demonstrating that by tuning the stable region during assembly, the transmission distance could be extended to 2.8 m. This work provides essential theoretical insights and practical design guidelines for realizing stable, long-range SDC systems.
Mingliang Xiong, Zeqian Guo, Qingwen Liu 0001, Gang Wang 0014, Gang Li 0020, Bin He 0003
IEEE Internet Things J.6
2026 Self-Aligning Resonant Beam for Simultaneous Wireless Power Transfer and Duplex Communication
abstract
Sustainable energy supply and high-speed communications are two significant needs for the upcoming 6G applications. This paper introduces a self-aligning resonant beam system for simultaneous light information and power transfer (SLIPT), employing a novel coupled spatially distributed resonator (CSDR). The system utilizes a resonant beam for efficient power delivery and a second-harmonic beam for concurrent data transmission, inherently minimizing echo interference and enabling bidirectional communication. Through comprehensive analyses, we investigate the CSDR’s stable region, beam evolution, and power characteristics in relation to working distance and device parameters. Numerical simulations validate the CSDR-SLIPT system’s feasibility by identifying a stable beam waist location for achieving accurate mode-match coupling between two spatially distributed resonant cavities and demonstrating its operational range and efficient power delivery across varying distances. The research reveals the system’s benefits in terms of both safety and energy transmission efficiency. We also demonstrate the trade-off among the reflectivities of the cavity mirrors in the CSDR. Besides, an experiment was conducted to verified the feasibility of self-aligning beam generation and safety under the designed structure. These findings offer valuable design insights for resonant beam systems, advancing SLIPT with significant potential for remote device connectivity.
Mingliang Xiong, Qingwen Liu 0001, Hao Deng 0002, Gang Wang 0014, Jianchen Zhu, Gang Li 0020, Bin He 0003
IEEE J. Sel. Areas Commun.6
2026 Beyond textual rationales: Anatomy-grounded chain-of-thought for traceable radiology reasoning
Jun Yang 0056, Mengyuan Xu, Mingliang Xiong, Wen Fang 0001, Mingqing Liu 0002, Hao Deng 0002, Bin He 0003, Gang Li 0020, Qingwen Liu 0001
Knowl. Based Syst.10
2025 Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing
abstract
Due to control limitations in the denoising process and the lack of training, zero-shot video editing methods often struggle to meet user instructions, resulting in generated videos that are visually unappealing and fail to fully satisfy expectations. To address this problem, we propose Align-A-Video, a video editing pipeline that incorporates human feedback through reward fine-tuning. Our approach consists of two key steps: 1) Deterministic Reward Fine-tuning. To reduce optimization costs for expected noise distributions, we propose a deterministic reward tuning strategy. This method improves tuning stability by increasing sample determinism, allowing the tuning process to be completed in minutes; 2) Feature Propagation Across Frames. We optimize a selected anchor frame and propagate its features to the remaining frames, improving both visual quality and semantic fidelity. This approach avoids temporal consistency degradation from reward optimization. Extensive qualitative and quantitative experiments confirm the effectiveness of using reward fine-tuning in Align-A-Video, significantly improving the overall quality of generated videos.
Yingkang Zhong, Jiangchuan Mu, Mingliang Xiong, Wen Fang 0001, Mingqing Liu 0002, Hao Deng 0001, Bin He 0003, Gang Li 0020, Qingwen Liu 0001
CVPR10
2025 Toward Cognitive Digital Twin System of Human-Robot Collaboration Manipulation
abstract
Multielement decision-making is crucial for the robust deployment of human-robot collaboration (HRC) systems in flexible manufacturing environments with personalized tasks and dynamic scenes. Large Language Models (LLMs) have recently demonstrated remarkable reasoning capabilities in various robotic tasks, potentially offering this capability. However, the application of LLMs to actual HRC systems requires the timely and comprehensive capturing of real-scene information. In this study, we suggest incorporating real scene data into LLMs using digital twin (DT) technology and present a cognitive digital twin prototype system of HRC manipulation, known as HRC-CogiDT. Specifically, we initially construct a scene semantic graph encoding the geometric information of entities, spatial relations between entities, actions of humans and robots, and collaborative activities. Subsequently, we devise a prompt that merges scene semantics with prior knowledge of activities, linking the real scene with LLMs. To evaluate performance, we compile an HRC scene understanding dataset and set up a laboratory-level experimental platform. Empirical results indicate that HRC-CogiDT can swiftly perceive scene changes and make high-level decisions based on varying task requirements, such as task planning, anomaly detection, and schedule reasoning. This study provides promising insights for the future applications of LLMs in robotics.Note to Practitioners—Recently, LLMs have demonstrated significant success in various robotic tasks, suggesting their potential as a powerful tool for robotic decision-making. Motivated by this, to improve the production efficiency of HRC in flexible manufacturing, we innovatively combine LLMs with DT technology, and propose a cognitive DT system for HRC, aiming to integrate LLMs into the decision-making loop of HRC system. Experiments conducted in a laboratory-scale platform indicate that the proposed system can handle different decision-making needs in different HRC activities. This system can provide professional guidance to operators in a comprehensible form and serve as a medium for monitoring the safety and standardization of the manipulation process. Future work will explore the use of virtual space provided by the proposed system to optimize the decision outputs of LLMs to make the proposed system more broadly applicable.
Xin Li 0093, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou, Gang Li 0020, Xiang Li 0010
IEEE Trans Autom. Sci. Eng.5
2025 A Biological Structure-Inspired Infrastructure Health Assessment Method
abstract
Structural health monitoring with wireless sensor networks(WSNs) plays an increasingly critical role in modern municipal. While there still remains a gap in getting a comprehensive understanding of complex structural data. Skin diseases can be well diagnosed with modern medical technology, from which similar methods can be learned to improve the intuitiveness and accuracy of structural health monitoring. This work proposes a multi-layered skin-like architecture based method(MSHA), which each layer has its own functions and on the whole presents the disaster situation. This biological structure-inspired architecture has three layers: 1) data substrate layer; 2) connectivity structure layer; 3) pathological manifestation layer, to simulate the three-layer structure of skin. First, a temporal feature extraction method is proposed, which can provide temporal correlations of each node. Second, a dimensional independent spacial feature extraction method is proposed, these first two methods form the connectivity structure layer, which is mainly composed of a spatio-temporal correlation model. Third, a structural health evaluation method for the pathological manifestation layer is proposed to fuse heterogeneous data and calculate multi-granularity structural risk with the features extracted from the connectivity structure layer. The experimental results show that MSHA can achieve not only accurate prediction and intuitive disaster situation, but also low energy consumption and longer lifetime. Note to Practitioners—This paper was motivated by the problem of effectively monitoring the structural health of infrastructure with wireless sensor networks. Existing methods for infrastructure structural health monitoring have space for improvement in maximizing the utilization of correlations between sensor nodes, temporal sequences, and between different sensor types, and fail to produce intuitive results to characterize structural health. In this paper, inspired by the multilayer structure of skin, we propose a new method for analyzing and presenting the structural changes of infrastructure under wireless sensor network data, which fully analyzes the relationship between sensor data in time, space, and type, and is able to predict the structural data in the future period and visually express the current safety situation with size and color information. In this paper, we mathematically describe the model construction method and feature transfer of the constructed intelligent monitoring method. We validate the proposed method using actual tunnel data. The experimental results show that the method is feasible and can accurately indicate the current safety situation in each area while achieving high accuracy in predicting future data, as well as good performance in terms of energy consumption and life cycle. In the future, we will explore optimal strategies for the placement of sensor nodes and improve the convenience and adaptability of the models across multiple scenarios.
Gang Li 0020, Bin He 0003, Bin Cheng 0008
IEEE Trans Autom. Sci. Eng.2
2025 Deep Learning for Low-Light Vision: A Comprehensive Survey
abstract
Visual recognition in low-light environments is a challenging problem since degraded images are the stacking of multiple degradations (noise, low light and blur, etc.). It has received extensive attention from academia and industry in the era of deep learning. Existing surveys focus on low-light image enhancement (LLIE) methods and normal-light visual recognition methods, while few comprehensive surveys of low-light-related vision tasks. This article provides a comprehensive survey of the latest advancements in low-light vision, including methods, datasets, and evaluation metrics, in two aspects: visual quality-driven and recognition quality-driven. On the visual quality-driven aspect, we survey a large number of very recent LLIE methods. On the recognition quality-driven aspect, we survey low-light object detection techniques in the deep learning era using more intuitive categorization method. Furthermore, a quantitative benchmarking of different methods is conducted on several widely adopted low-light vision-related datasets. Finally, we discuss the challenges that exist in low-light vision and future directions worth exploring. We provide a public website that will continue to track developments in this promising field.
Qian Zhao 0014, Gang Li 0020, Bin He 0003, Runjie Shen
IEEE Trans. Neural Networks Learn. Syst.2
2025 Resonant Beam Enabled Multi-Target Localization
abstract
In the era of the Internet of everything (IoE) and the metaverse, there is a growing demand for high-accuracy indoor positioning for applications such as autonomous robots, virtual reality, and smartphones. This paper proposed a resonant beam phase-based passive localization (RBPPL) system optimized for high-precision indoor positioning in multi-access scenarios. By leveraging the self-alignment characteristic and integrating the analysis of resonant beam phase, angle of arrival (AoA) matching and binocular disparity method for 3D point coordinate acquisition, the RBPPL system achieves binocular passive multi-access 3D positioning with an error within 4 cm at a distance of 8 m. We present a novel multi-access AoA estimation method that overcomes the challenges of spot overlap in traditional CMOS-based angle analysis. We propose a telescope system to correct the phase and focus the propagation direction of optical resonant beam systems. Simulations demonstrate the system’s robustness and high accuracy. The proposed RBPPL system, optimized for multi-access scenarios, offers a promising solution for high-accuracy indoor positioning, supporting various IoE and metaverse applications. Future work will focus on real-world deployment and its potential in complex multi-access scenarios.
Guangkun Zhang, Mengyuan Xu, Wen Fang 0001, Mingliang Xiong, Mingqing Liu 0002, Gang Li 0020, Bin He 0003, Qingwen Liu 0001
IEEE Trans. Wirel. Commun.8
2024 A digital twin system for Task-Replanning and Human-Robot control of robot manipulation
Xin Li 0093, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou, Gang Li 0020, Zhongpan Zhu
Adv. Eng. Informatics5
2024 Landslide displacement prediction with step-like curve based on convolutional neural network coupled with bi-directional gated recurrent unit optimized by attention mechanism
Shaoqiang Meng, Zhenming Shi, Gang Li 0020, Hongchao Zheng
Eng. Appl. Artif. Intell.4
2024 Graph-based geometric structure line parsing
Feng Li 0046, Gang Li 0020, Bin He 0003, Ping Lu 0013, Bin Cheng 0008
Neurocomputing2
2024 Laser Ranger-Based Baseline Measurement for Collaborative Localization
abstract
To address challenges in outdoor multi-robot collaborative localization (MRCL) due to low GPS accuracy, we propose a system using three UGVs, each equipped with a shared camera and a laser rangefinder. Our trilateral localization algorithm combines least-squares matrix and gradient descent methods, resulting in an 80.4% improvement in accuracy compared to traditional methods. The system mitigates the limitations of GPS accuracy by utilizing accurate baseline measurements and optimizing the localization process. These advancements have potential applications in transportation, production, and logistics, enhancing MRCL performance in outdoor environments.
Mingqing Liu 0002, Yihan Zhu, Qingwen Liu 0001, Qunhui Yang, Gang Li 0020, Bin He 0003
IEEE Internet Things J.7
2023 A Survey of Crowdsensing and Privacy Protection in Digital City
abstract
The key pillar of developing digital city is the ubiquitous sensing of people and the environment. Crowdsensing requires a large number of users to participate in the collection of sensing data, and these data may carry sensitive information, such as identity and location related to the users or sensing object. If this information is eavesdropped, intercepted, and leaked, this may seriously harm the interests of individuals, organizations, and even countries. Therefore, from a privacy perspective, users may be reluctant to open data. While relying on mobile devices used by a large number of ordinary users as the basic sensing unit, it is necessary to include a variety of communication methods to realize the distribution of sensing tasks and to collect the sensing data. Then, to complete the complex crowdsensing tasks, it is important to ensure privacy security in the context of crowdsensing because it is a key problem. In this article, we comb through the development status of crowdsensing in the digital city, emphatically analyze the privacy protection in crowdsensing under the background of digital city, and qualitatively evaluate the existing privacy protection technologies for crowdsensing. Finally, this article presents research challenges and future directions that should be addressed to improve the performance of privacy protection technologies for crowdsensing systems.
Bin He 0003, Gang Li 0020, Bin Cheng 0008
IEEE Trans. Comput. Soc. Syst.3
2022 A Data Agent Inspired by Interpersonal Interaction Behaviors for Wireless Sensor Networks
abstract
Event monitoring is the main purpose of monitoring applications based on wireless sensor networks (WSNs). To realize the autonomous representation of tunnel disasters in WSNs, this work proposes a data agent (DA) inspired by interpersonal interaction behaviors, which bridges the gap between data and humans. The DA divides event monitoring into four states: 1) active perception; 2) understanding; 3) thinking; and 4) representation, to realize a kind of human-like event monitoring. First, a DA behavior model based on a finite-state machine is proposed, which can adaptively switch among these four states. Second, four methods, including active perception based on a dual neuron perception structure, understanding for forming the information granularity, thinking for forming the disaster knowledge graph, and representation using grayscale and knowledge graphs are proposed to realize the four states of the DA. The experimental results show that the DA inspired by interpersonal interaction can make WSNs not only achieve efficient and accurate disaster self-monitoring but also represent the disaster situation intuitively and quickly through grayscale and knowledge graphs.
Gang Li 0020, Bin He 0003, Zhipeng Wang 0006, Yanmin Zhou
IEEE Internet Things J.1
2022 Blockchain-Enhanced Spatiotemporal Data Aggregation for UAV-Assisted Wireless Sensor Networks
abstract
Wireless sensor networks (WSNs) are widely used in the field of monitoring. For data collection of sensor nodes in large-scale monitoring scenarios, unmanned aerial vehicle (UAV)-assisted WSNs have emerged. For the security and validity of data collection, a blockchain-enhanced data collection framework for UAV-assisted WSNs is presented in this article. To reduce data redundancy in WSNs, a sparsity-optimized and compressed sensing-based spatiotemporal data aggregation model is built. A UAV identity authentication mechanism based on a Merkle tree is also designed to ensure the security of data transmission. By combining blockchain building and data aggregation, a disaster semantic blockchain (DSB) based on a data reconstruction-directed consensus mechanism is presented. Disaster semantics are extracted by analyzing the semantic association relationship of disaster, background, event, and sensor data. The experimental results show that the blockchain-enhanced spatiotemporal data aggregation effectively increases the network life cycle and data reconstruction accuracy. The disaster situation can be described accurately through the DSB.
Gang Li 0020, Bin He 0003, Zhipeng Wang 0006, Jie Chen 0003
IEEE Trans. Ind. Informatics1
2021 A data-efficient goal-directed deep reinforcement learning method for robot visuomotor skill
Rong Jiang 0003, Zhipeng Wang 0006, Bin He 0003, Yanmin Zhou, Gang Li 0020, Zhongpan Zhu
Neurocomputing5
2018 A traffic congestion aware vehicle-to-vehicle communication framework based on Voronoi diagram and information granularity
Gang Li 0020, Bin He 0003, Aimin Du
Peer-to-Peer Netw. Appl.1
2017 Intelligent Self-Adaptation Data Behavior Control Inspired by Speech Acts
abstract
Wireless sensor networks (WSNs) are a promising technology for collecting information by utilizing various types of small-sized sensors. Implementing efficient data acquisition and transmission is important for real-time monitoring and data analysis, owing to massive and heterogeneous data with low precision. In this work, the data behavior model inspired by speech acts is built, including the data correlation model and the data comment model. Intelligent self-adaptation data behavior control is then proposed, of which the main idea is to make sensor nodes (SNs) process data intelligently. In the proposed data behavior control, the temporal compression behavior is achieved through variable-cycle transmission and the multivariate spatial compression behavior is motivated by compressed sensing (CS) theory. The multivariate spatial-temporal compression composed of the above two compression behaviors can adjust itself through data comment behaviors including the large-cycle error self-correction behavior based on the node credibility and the large-cycle event self-adaptation behavior based on the information granularity. The data behavior control presented has been validated by the experiments in the Shanghai metro tunnel. The experiment results show that these intelligent data behaviors inspired by speech acts can make WSNs more effective and intelligent with no change in the existing network structure.
Bin He 0003, Gang Li 0020
ACM Trans. Sens. Networks2
2015 Salient object detection based on meanshift filtering and fusion of colour information
abstract
Colour and its spatial distribution are the main information currently used to detect salient objects in an image, but this cannot always guarantee satisfying performance. To deal with this problem, a salient object detection algorithm has been presented based on meanshift filtering and fusion of colour information. Superpixel segmentation is used to analyse the images by sets of pixels instead of single pixel, which improves the robustness of the algorithm to noises, as well as the efficiency. Meanshift filtering is used to detect the modes of every superpixel in spatial domain and range domain, respectively, which is the basis of the subsequent calculation. Each target therefore offers almost the same saliency and the spatial distribution of which will be easier to analyse. The fusion of colour contrast and colour concentration as well as centre prior is used as criterion to evaluate the saliency of every single superpixel. According to the tests of the algorithm on the open popular dataset, it has been proved that the algorithm presented in this work shows better results in both the aspect of effectiveness and efficiency, compared with its traditional equivalents.
Gang Li 0020, Bin He 0003, Xiaojiao Tao
IET Image Process.3