EDBT 2026 Demo / reviewers in the wild / expert
Xiangwei Zhu
dblp:60/8945
· DBLP profile ↗
29ranked-venue papers
0as first author
27since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 12 · 12 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Systems, architecture and hardware · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Real-time event-based visual rotational odometry with time-aware mean minimization
Hexiong Yao, Mingjun Ouyang, Zhiqiang Dai, Xiangwei Zhu |
Expert Syst. Appl. | 5 |
| 2026 | A time synchronization attack detection framework for communication base stations featuring ultra-low false alarm rate and balanced detection speed
Chuhan Huang, Xiangwei Zhu, Chonggui Yi |
Expert Syst. Appl. | 4 |
| 2026 | Learning Spatiotemporal Variations Using BiLSTM for Smartphone Precise Point Positioning Correction in Urban AreasabstractGlobal navigation satellite system (GNSS) provides precise, real-time, all-weather location services for Internet of Things (IoT) applications. However, smartphone positioning accuracy is easily affected by signal obstructions such as multipath and non-line-of-sight effects in challenging urban environments. Therefore, we present a bidirectional long short-term memory (BiLSTM) neural network capable of learning the spatiotemporal mechanisms in positioning to mitigate precise point positioning (PPP) errors. We derive the relationships between time, velocity, displacement, and position in GNSS positioning and introduce the concepts of instantaneous displacement error, cumulative displacement error, and estimated positioning error as pseudo-error variables. Based on the spatiotemporal variation mechanisms, we construct a BiLSTM network to learn error patterns in smartphone PPP technology to improve positioning accuracy. Validated on the Google Smartphone Decimeter Challenge (GSDC) dataset, our method reduces positioning errors by 50–75%, thereby achieving a positioning accuracy of 2–3 meters in three directions in challenging urban environments. Compared to other advanced learning-based positioning methods, our approach demonstrates strong versatility under data-constrained conditions, on unseen devices, and across untrained trajectories, while maintaining comparable positioning accuracy. Thus, the proposed mechanisms and methods have the potential to be applied to various ubiquitous global positioning applications. Wanqing Li 0005, Jiangbo Song, Bo Xu 0029, Xiangwei Zhu |
IEEE Internet Things J. | 4 |
| 2026 | Policy distillation-based multiagent actor-critic for cooperative UAV path planning in complex environments
Huidong Liu, Jiangshan Ai, Xianlei Long, Yong Li 0023, Xiangwei Zhu, Fuqiang Gu |
Knowl. Based Syst. | 6 |
| 2026 | Fontify : One-shot font generation via in-context learning
Xiangwei Zhu |
Knowl. Based Syst. | 2 |
| 2026 | CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic SegmentationabstractDue to the substantial domain gaps in Remote Sensing (RS) images that are characterized by variabilities such as location, wavelength, and sensor type, Remote Sensing Domain Generalization (RSDG) has emerged as a critical and valuable research frontier, focusing on developing models that generalize effectively across diverse scenarios. However, research in this area remains underexplored: (1) Current cross-domain methods primarily focus on Domain Adaptation (DA), which adapts models to predefined domains rather than to unseen ones; (2) Few studies target the RSDG issue, especially for semantic segmentation tasks. Existing related models are developed for specific unknown domains, struggling with issues of underfitting on other unseen scenarios; (3) Existing RS foundation models tend to prioritize in-domain performance over cross-domain generalization. To this end, we introduce the first vision foundation model for RSDG semantic segmentation, CrossEarth. CrossEarth demonstrates strong cross-domain generalization through a specially designed data-level Earth-Style Injection pipeline and a model-level Multi-Task Training pipeline. In addition, for the semantic segmentation task, we have curated an RSDG benchmark comprising 32 semantic segmentation scenarios across various regions, spectral bands, platforms, and climates, providing comprehensive evaluations of the generalizability of future RSDG models. Extensive experiments on this collection demonstrate the superiority of CrossEarth over existing state-of-the-art methods. Ziyang Gong, Zhixiang Wei, Di Wang 0023, Xiaoxing Hu, Xianzheng Ma, Hongruixuan Chen, Yuru Jia, Yupeng Deng 0002, Zhenming Ji, Xiangwei Zhu, Xue Yang 0005, Naoto Yokoya, Jing Zhang 0037, Bo Du 0001, Junchi Yan, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | Factor Graph-Based Tightly Coupled PPP-B2b/INS for Real-Time Precise Positioning
Ruite Yi, Xiangwei Zhu, Mingjun Ouyang, Chengchao Bai, Guangteng Fan |
IEEE Signal Process. Lett. | 2 |
| 2025 | Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity RecognitionabstractGrounded Multimodal Named Entity Recognition (GMNER) is an emerging information extraction (IE) task, aiming to simultaneously extract entity spans, types, and corresponding visual regions of entities from given sentence-image pairs data. Recent unified methods employing machine reading comprehension or sequence generation-based frameworks show limitations in this difficult task. The former, utilizing human-designed type queries, struggles to differentiate ambiguous entities, such as Jordan (Person) and off-White x Jordan (Shoes). The latter, following the one-by-one decoding order, suffers from exposure bias issues. We maintain that these works misunderstand the relationships of multimodal entities. To tackle these, we propose a novel unified framework named Multi-grained Query-guided Set Prediction Network (MQSPN) to learn appropriate relationships at intra-entity and inter-entity levels. Specifically, MQSPN explicitly aligns textual entities with visual regions by employing a set of learnable queries to strengthen intra-entity connections. Based on distinct intra-entity modeling, MQSPN reformulates GMNER as a set prediction, guiding models to establish appropriate inter-entity relationships from a optimal global matching perspective. Additionally, we incorporate a query-guided Fusion Net (QFNet) as a glue network to boost better alignment of two-level relationships. Extensive experiments demonstrate that our approach achieves state-of-the-art performances in widely used benchmarks. Jielong Tang, Ziyang Gong, Jianxing Yu, Xiangwei Zhu, Jian Yin 0001 |
AAAI | 5 |
| 2025 | Towards Efficient Image-goal Navigation: A Self-supervised Transformer-based Reinforcement Learning ApproachabstractImage-goal navigation is a crucial yet challenging task that requires an agent to navigate to a goal location specified by an image. Modular methods decompose the problem into distinct subtasks and often involve explicit map construction, which can struggle in complex, unstructured environments. In contrast, end-to-end deep reinforcement learning (DRL)-based methods directly output actions from visual input, with recent improvements focusing on enhancing embedding fusion between the current observation and the goal image. However, both approaches fail to fully leverage the rich temporal relationships present in the agent’s visual-action history. In this paper, we address this limitation by employing a self-supervised transformer to predict masked portions of the agent’s visual-action embeddings. To promote spatio-temporal reasoning, a dual-attention shared transformer is utilized for both masked representation learning and policy generation. Our method demonstrates superior performance and generalization ability compared to 12 existing baselines across the Gibson, MP3D, and HM3D datasets. Code and trained models are available at https://github.com/hujch23/DaMVA. Jiaocheng Hu, Xiangwei Zhu |
IROS | 4 |
| 2025 | A GNSS Time Synchronization Attack Detection Method for Commercial Off-the-Shelf Receivers: Cumulative Second-Order Difference of PseudorangesabstractThe security of time synchronization is crucial in the Internet of Things (IoT), as it enables sensors and actuators to collect and transmit various types of data. However, root and critical nodes in the IoT are typically connected to the global navigation satellite system (GNSS) receivers, increasing the susceptibility of the system to GNSS-based time synchronization attacks (TSAs). When subjected to TSAs, the scheduling and messages of nodes will conflict, especially in distributed IoT systems. To address this issue, a TSA detection method named the cumulative second-order difference of pseudoranges (CD2-P) is proposed in this article. Furthermore, we propose an alarm strategy called CD2-P-alarm based on TSA characteristics. When an attacker tries to manipulate the GNSS receiver’s time, this method can detect the time discrepancy and provide an early warning. This article verifies that the proposed method outperforms the other methods by using software-defined receivers based on the Texas spoofing test battery data set. Compared with the state-of-the-art methods, the proposed method can provide more than a 10% improvement in detection probability. Moreover, the practical applicability of the proposed method on a U-blox receiver and an Android mobile phone is validated through real-world experiments. This indicates that the proposed method can be seamlessly integrated into current commercial receivers without the need for modification. By focusing on processing the output data produced by commercial off-the-shelf receivers, the method effectively increases the real-time security of existing IoT systems. Wenqiang Wei, Yibo Si, Xiangwei Zhu |
IEEE Internet Things J. | 5 |
| 2025 | Image Compression and Transmission System Based on Narrowband Satellite Internet of ThingsabstractThe Satellite Internet of Things (SatIoT) enables integrated space–air–ground connectivity and information exchange through satellite communication networks and distributed sensors. Its primary advantage lies in achieving global coverage, effectively addressing communication challenges in signal-deprived regions. However, image transmission in SatIoT systems is severely constrained by limited satellite bandwidth, placing high demands on the efficiency and adaptability of compression algorithms. To address these challenges, this work adopts a co-design approach at both the algorithmic and system levels. At the algorithmic level, we implement a deep learning-based image compression framework. To improve the Rate-Distortion performance of the algorithm, 1) we propose a plug-and-play feature extraction module that integrates the strengths of Transformer and CNN to capture both global and local features, preserving critical image details while eliminating redundancy. 2) A Multi Channel Scaling Enhancement module is proposed, which can efficiently fuse features without performance loss and reduce the computational overhead through channel scaling. The experimental results show that they significantly improve the Rate-Distortion performance of mainstream compression algorithms with relatively low computational consumption. At the system level, we design a complete SatIoT communication chain comprising an edge terminal, the Tiantong satellite, and a ground receiving terminal. The optimized compression algorithm is deployed on the edge terminal to enable edge computing, supporting image acquisition, compression, and satellite-based transmission. The system achieves reliable, end-to-end image transmission from signal-deprived regions to ground networks, greatly improving the efficiency and robustness of satellite-based image delivery. Guanzhong Liao, Yiheng Fan, Qianyao Xu, Mingjun Ouyang, Xiangwei Zhu |
IEEE Internet Things J. | 5 |
| 2025 | A Novel GNSS Decentralized Cooperative Positioning Algorithm for Internet of VehiclesabstractWith cooperative positioning (CP) in Internet of Vehicles (IoV), the positioning performance can be improved by utilizing the positioning information provided by neighboring vehicles. However, in urban canyons, the CP performance based on global navigation satellite system (GNSS) is severely degraded by multipath and non-line-of-sight (NLOS) effects, and a robust CP algorithm is required. This article proposes a GNSS decentralized CP (DCP) algorithm based on the generalized extreme studentized deviate test (GESD) filter. Further, a new two-stage CP framework is established, consisting of two modules, the independent and connected modules. The independent module is the first filter layer detection for GNSS pseudorange residual error, operating independently within each vehicle to mitigate the impact of abnormal measurements, such as multipath bias and satellite faults. When receiving data from neighboring vehicles, the connected module is activated to detect shared GNSS pseudorange errors and relative pseudorange measurements, further reducing the impact of anomalous measurements. The outdoor experiments validate the superiority of the DCP algorithm over the single-point positioning (SPP) method based on receiver autonomous integrity monitoring fault detection and exclusion (RAIM-FDE) in both GPS-only and GPS/BDS combination strategies, especially in blocked situations. The positioning root-mean-square error (RMSE) of the DCP algorithm for horizontal positioning is about 1.00 m in static scenes and about 8.00 m in dynamic scenes. Compared with the SPP-RAIM, the DCP algorithm can improve about 25.22%–43.04% and 16.55%–40.17% on average under GPS/BDS combination strategies in the horizontal and 3-D directions, respectively. Hexiong Yao, Mingjun Ouyang, Zhiqiang Dai, Xiangwei Zhu, Qianqiang Lin |
IEEE Internet Things J. | 7 |
| 2025 | Multimodal speech emotion recognition via modality constraint with hierarchical bottleneck feature fusion
Ying Wang 0078, Jianjun Lei 0002, Xiangwei Zhu |
Speech Commun. | 3 |
| 2025 | Spike-BRGNet: Efficient and Accurate Event-Based Semantic Segmentation With Boundary Region-Guided Spiking Neural NetworksabstractEvent-based semantic segmentation in traffic scenes has attracted considerable attention in autonomous driving systems due to the advantages of event cameras such as high dynamic range, low latency, and low energy consumption. However, existing Artificial Neural Network (ANN)-based methods rely on conventional image frames, often neglecting the spatial-temporal dynamics inherent in event streams and consuming higher energy costs, significantly limiting their applicability in energy-constrained environments. In this study, we introduce Spike-BRGNet, a Spike-driven Boundary Region-Guided Network that efficiently extracts boundary information utilizing only events to guide the segmentation encoder, while preserving the energy efficiency of Spiking Neural Networks (SNNs). Specifically, to explore the implicit information from events, we design a three-branch spiking encoder that consists of semantic detail (SD), context aggregation (CA), and boundary aware (BA) branches to capture specific features. Then, a spiking multi-scale context aggregation (SMSCA) module is proposed to enhance the semantics of the CA branch. Finally, a novel boundary region-guided loss function and a dynamic surrogate gradient function, EvAF, are designed to optimize the model. Extensive experiments show that our model outperforms state-of-the-art (SOTA) SNN-based methods on DDD17 (+1.57%) and DSEC dataset (+1.91%). Furthermore, Spike-BRGNet consumes$17.76\times $less energy than ANN-based models, showing superior energy-saving performance. Xianlei Long, Xiaxin Zhu, Fangming Guo, Chao Chen 0004, Xiangwei Zhu, Fuqiang Gu, Songyu Yuan, Chunlong Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Tightly Coupled RTK-Visual-Inertial Integration With a Novel Sliding Ambiguity Window Optimization FrameworkabstractAccurate and reliable navigation is fundamental for intelligent transportation applications such as autonomous driving. Current research community has gained advances in global navigation satellite system (GNSS) and its fusion with visual-inertial navigation system (VINS). However, existing optimization-based integrations fail to fully utilize the constant characteristic of carrier phase ambiguity when facing frequent cycle slips and degraded GNSS signals, leading to erratic navigation output in complex environments. To address these limitations, this work proposes a tightly coupled GNSS real-time kinematic (RTK)-visual-inertial integration with a novel sliding ambiguity window optimization framework to achieve high-precision and robust positioning. Specifically, a sliding ambiguity window framework is built to associate float single-differenced ambiguities estimated by factor graph optimization and integer double-differenced ambiguities obtained by LAMBDA ambiguity resolution (AR). The framework can maintain continuous AR constraints across epochs and extend ambiguity-fixed solutions even during AR failures. Additionally, a VINS-aided single&dual-frequency hybrid cycle slip detection method is proposed, combining available dual-frequency GNSS observations and VINS information to perform reliable cycle slip detection for all single- and dual-frequency carrier phases. The superiority of the proposed method is verified by real-world vehicle traveling experiments in campus and urban scenarios. Results show that our proposed method can achieve robust centimeter-level positioning accuracy in both trajectories, outperforming the state-of-the-art optimization-based integration method by 72.2% and 66.0%, respectively. Chufeng Duan, Shengquan Li 0002, Quan Tang 0001, Zhiqiang Dai, Xiangwei Zhu |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Parsing All Adverse Scenes: Severity-Aware Semantic Segmentation with Mask-Enhanced Cross-Domain ConsistencyabstractAlthough recent methods in Unsupervised Domain Adaptation (UDA) have achieved success in segmenting rainy or snowy scenes by improving consistency, they face limitations when dealing with more challenging scenarios like foggy and night scenes. We argue that these prior methods excessively focus on weather-specific features in adverse scenes, which exacerbates the existing domain gaps. To address this issue, we propose a new metric to evaluate the severity of all adverse scenes and offer a novel perspective that enables task unification across all adverse scenarios. Our method focuses on Severity, allowing our model to learn more consistent features and facilitate domain distribution alignment, thereby alleviating domain gaps. Unlike the vague descriptions of consistency in previous methods, we introduce Cross-domain Consistency, which is quantified using the Structure Similarity Index Measure (SSIM) to measure the distance between the source and target domains. Specifically, our unified model consists of two key modules: the Merging Style Augmentation Module (MSA) and the Severity Perception Mask Module (SPM). The MSA module transforms all adverse scenes into augmented scenes, effectively eliminating weather-specific features and enhancing Cross-domain Consistency. The SPM module incorporates a Severity Perception mechanism, guiding a Mask operation that enables our model to learn highly consistent features from the augmented scenes. Our unified framework, named PASS (Parsing All adverSe Scenes), achieves significant performance improvements over state-of-the-art methods on widely-used benchmarks for all adverse scenes. Notably, the performance of PASS is superior to Semi-Unified models and even surpasses weather-specific models. Fuhao Li, Ziyang Gong, Yupeng Deng 0002, Xianzheng Ma, Renrui Zhang, Zhenming Ji, Xiangwei Zhu |
AAAI | 7 |
| 2024 | CoDA: Instructive Chain-of-Domain Adaptation with Severity-Aware Visual Prompt Tuning
Ziyang Gong, Fuhao Li, Yupeng Deng 0002, Deblina Bhattacharjee, Xianzheng Ma, Xiangwei Zhu, Zhenming Ji |
ECCV (77) | 6 |
| 2024 | Human State Recognition Using Ultra Wideband Radar Based on CvTabstractEmotions have a significant impact on an individual’s life, and positive emotions can enhance their quality of life. Moreover, emotions play a crucial role in the field of medicine. However, the complexity of human emotions has hindered their rapid development in the medical field. Therefore, it is crucial to rapidly and accurately identify human emotions. In the field of emotion recognition, radar is widely used for monitoring human physiological features due to its strong penetrability and non-contact advantages. In deep learning, the Transformer model has been introduced into radar signal processing due to its excellent global perception ability, offering new possibilities for fast and accurate detection of human emotions. In this paper, we propose a human emotion recognition system based on ultra-wideband radar, utilizing the Convolutional Vision Transformer (CvT) model as the deep learning model for emotion classification. The CvT model incorporates convolution into the Vision Transformer architecture, combining the advantages of both in image recognition tasks. Unlike previous works with convolutional networks, CvT fully leverages the benefits of convolution while retaining the characteristics of the Transformer. We pre-trained the model on a publicly available radar dataset and then conducted experimental validation using our collected dataset. The experimental results demonstrate that our network outperforms traditional convolutional approaches, achieving a test accuracy of 86.25%, which is of significant importance for radar-based emotion recognition. Du Li, Yiyu Xu, Xuelin Yuan, Xiangwei Zhu |
IEEE Internet Things J. | 5 |
| 2024 | GLRT-Based Spacetime Detection Algorithms Via Joint DoA and Doppler Shift Method for GNSS Spoofing InterferenceabstractThe critical applications of global navigation satellite system in the field of national defense and security are seriously limited by the susceptibility of its system to spoofing interference. At present, The spoof interference detection algorithms commonly utilize temporal consistency and other time-frequency information, as well as angle resolution and other spatial information, which the types of referenced information are limited and singular. In this article, a spacetime spoofing interference detection algorithm is proposed, which simultaneously combines both temporal and spatial information. A joint generalized likelihood ratio test framework involving relative Direction of Arrival (DoA) and Doppler shift difference is established, and the characteristics of inconsistent incoming angular angle and similar Doppler shift of spoofing signals are fully utilized. Experimental findings indicate that the proposed algorithm can effectively detect spoofing in scenarios involving a limited number of spoofed satellites. Moreover, the detection probability has been reached 99% compared with the traditional DoA estimation technique and metric combination technique, which verifies the efficacy and dependability of the proposed approach. Xinzhi Peng, Chuhan Huang, Xiangwei Zhu, Zhengkun Chen, Xuelin Yuan |
IEEE Internet Things J. | 3 |
| 2024 | ORB-NeuroSLAM: A Brain-Inspired 3-D SLAM System Based on ORB FeaturesabstractIntelligent navigation is a fundamental technology that enables unmanned systems to achieve autonomy in the intelligent era. However, existing navigation schemes suffer from high computational complexity and power consumption, as well as low robustness in complex or unknown environments. To address these challenges, this paper proposes a novel 3D brain-inspired simultaneous localization and mapping (SLAM) method, called ORB-NeuroSLAM, based on the oriented FAST and rotated BRIEF (ORB) features. The proposed method takes inspiration from the robust and low-power navigation capabilities of humans and animals. The ORB-NeuroSLAM leverages the ORB features of camera images to compute robot self-motion and visual cues. Then, continuous attractor neural networks (CANNs) model multilayered head direction cells and three-dimensional grid cells that exist in animal brains. These cells are utilized jointly to represent the robot poses. Efficient and robust methods for loop closure detection and experience map construction were also developed. The proposed method was verified on 10 KITTI datasets, and experimental results demonstrate that it outperforms state-of-the-art brain-inspired SLAM methods in terms of accuracy and efficiency. Additionally, it is comparable to state-of-the-art visual SLAM method ORB-SLAM3. Dan Shen 0003, Gelu Liu, Fangwen Yu, Fuqiang Gu, Xiangwei Zhu |
IEEE Internet Things J. | 7 |
| 2024 | R²-GVIO: A Robust, Real-Time GNSS-Visual-Inertial State Estimator in Urban Challenging EnvironmentsabstractVisual-Inertial Odometry (VIO) often suffers from drifting, particularly in large-scale environments. Concurrently, the Global Navigation Satellite System (GNSS) signals will be intermittent or even inaccessible in obstructed environments. In this work, we present a robust, real-time GNSS-visual-inertial state estimator, abbreviated as R-GVIO, that achieves drift-free six-degree-of-freedom accurate global positioning in urban challenging environments. Specifically, we integrate GNSS into the optimization-based monocular and stereo VIO, enabling the fusion of GNSS Real-time Kinematic (RTK), reprojection error, and IMU pre-integration within a factor graph framework. In addition, to enhance the R-GVIO’s robustness in GNSS-unfriendly environments, we propose an online GNSS measurement outlier detection and culling algorithm based on velocity constraints and a time synchronization mechanism. Furthermore, we propose a GNSS-aided initialization method that provides accurate gravity direction, IMU zero bias, and local-global external parameters. Moreover, we adopt the marginalization strategy of covisibility graph keyframes and add the GNSS factor constraint on the covisibility keyframe pose, ensuring the real-time performance of the system. Experiments on public datasets and real-world experiments verify that in challenging environments, including urban canyons, complex indoor-outdoor, and large-scale environments, the positioning accuracy and robustness of the proposed R-GVIO algorithm are superior to the state-of-the-art algorithms VINS-Fusion, ORB-SLAM3, IC-GVINS, and GICI-LIB. To make contribute to the community, we open sourced the complex environment datasets “SYSU-Campus-GVI” on GitHub. Jiangbo Song, Wanqing Li 0005, Chufeng Duan, Xiangwei Zhu |
IEEE Internet Things J. | 4 |
| 2024 | Corrections to "R₂-GVIO: A Robust, Real-Time GNSS-Visual-Inertial State Estimator in Urban Challenging Environments"abstractPresents corrections to the paper, (Corrections to “R₂-GVIO: A Robust, Real-Time GNSS-Visual-Inertial State Estimator in Urban Challenging Environments”). Jiangbo Song, Wanqing Li 0005, Chufeng Duan, Xiangwei Zhu |
IEEE Internet Things J. | 4 |
| 2024 | SG-VIO: Monocular Visual-Inertial Odometry With Tightly Coupled Structural Lines and Gravity to Avoid DegeneracyabstractVisual-inertial odometry (VIO) has played an important role in the field of the Internet of Things, providing a variety of devices and systems with high-precision and reliable positioning and navigation capabilities. In particular, indoor environments have become an important scenario for its application. However, the lack of robustness of point-based VIO systems in low-textured man-made environments often leads to failure. Based on this issue, this article proposes an innovative monocular VIO approach to fully utilize the available information in man-made environments. In the front end, the inertial measurement unit measurement model is defined by the preintegration method. In image data, first, the line features undergo a uniformization process, which reduces the redundant features and improves the accuracy of line feature matching. Then, the vanishing points in the image are detected using the Manhattan world assumption, and the structural line features are classified as either parallel or perpendicular to gravity based on vanishing points. In the back end, a novel residual term is defined for structural line features and gravity, deriving the corresponding Jacobian. This approach effectively addresses the issue of structural line degeneracy and continuously optimizes gravity, while also increasing the utilization of structural lines. A sliding window nonlinear optimization method is employed to minimize the sum of residuals. We tested the proposed system and the state-of-the-art VIO systems on both the public data sets and our collected data set to validate the effectiveness of the proposed system. Hexiong Yao, Yuexin Ma, Peijing Li, Chunlei Zhai, Jiangbo Song, Mingjun Ouyang, Zhiqiang Dai, Xiangwei Zhu |
IEEE Internet Things J. | 8 |
| 2024 | CASIT: Collective Intelligent Agent System for Internet of ThingsabstractIn the last few years, the bottleneck of bandwidth in Internet of Thing (IoT) has driven expectations to figure out new ways to preprocess the information needed to be transmitted. The ways which were used before are not smart enough and they cannot align to the users’ need. Large language model (LLM)-based intelligent agent is a very hot concept in AI community, which aims to save various problems via adapting LLM to different industries. In this article, we present a collective intelligent agent system for the IoT (CASIT) that is a pioneering LLM-agent-based IoT system. We put forward a IoT framework that can be used to lots of scenarios. CASIT refers to a system based on multiple intelligent LLM agents, which realizes complex tasks through cooperation and makes full use of collective intelligence. In order to solve the problems, we designed the Memory Mechanism and Summary Mechanism that enable LLMs to efficiently process the data by comparing historical data with Local Knowledge and Chat History in the prompt. After experimental verification, we have found that our framework could accurately conclude the abnormal information, and it outperforms the single LLM system when we input 200 sets of temperature and humidity data from five different places. The system provides a new solution and method for information processing in all IoT systems. Our framework may also provide refreshing ideas for edge computing and semantic communication. Ningze Zhong, Yi Wang 0095, Yingyue Zheng, Mingjun Ouyang, Dan Shen 0003, Xiangwei Zhu |
IEEE Internet Things J. | 8 |
| 2024 | Interplay of Request Number and Cache Size in Coded CachingabstractCoded caching is an effective method to reduce the traffic load on network bottleneck. While the heterogeneities on the number of requests and cache sizes in coded caching have been studied independently, their joint impact is still unclear. This paper investigates coded caching in scenarios with heterogeneous number of requests and cache sizes. We propose two achievable schemes. The first scheme, based on file grouping and multi-round decentralized coded caching, is demonstrated to be order optimal under the worst setting, i.e., when user with the i-th smallest cache has the i-th largest number of requests. Moreover, we obtain an important insight that the lower bound of the achievable rate is predominantly influenced by users with high$\frac {X_{i}}{M_{i}}$ratios, where$X_{i}$and$M_{i}$represent the number of requests and cache size of user i, respectively. Since the achievable rate of our first scheme is difficult to analyze in the general setting, we further propose the second scheme to derive a tractable upper bound. Based on the insight, the second scheme rearranges the users according to their$\frac {X_{i}}{M_{i}}$ratios and employs a threshold to divide them into the head users who may have a large impact on the lower bound, and the tail users who may have a small impact on the lower bound. The server transmits the demands of the head users directly while applying the first scheme in groups among the tail users. Under the general setting, the gap between the rate of our second scheme and the lower bound is proved to be within a logarithmic factor. Simulations are conducted to verify the superior performance of our proposed schemes. Kai Huang 0012, Xiaoxia Wang 0001, Jinbei Zhang, Kechao Cai, Xiangwei Zhu |
IEEE Trans. Commun. | 5 |
| 2023 | Dual-attention assisted deep reinforcement learning algorithm for energy-efficient resource allocation in Industrial Internet of Things
Ying Wang 0078, Fengjun Shang, Jianjun Lei 0002, Xiangwei Zhu, Haoming Qin, Jiayu Wen |
Future Gener. Comput. Syst. | 4 |
| 2022 | BAT: Block and token self-attention for speech emotion recognition
Jianjun Lei 0002, Xiangwei Zhu, Ying Wang 0078 |
Neural Networks | 2 |
| 2016 | GPU-accelerated main road extraction in Polarimetric SAR images based on MRFabstractRoad extraction in SAR (Synthetic Aperture Radar) images is a difficult task because of speckle noise and various interferences. PolSAR (Polarimetric SAR) measures road's reflectivity in four polarizations and provides more information of road, which potentially indicates higher extraction performance than that in single-polarization SAR cases. However, the implementation of the traditional MRF (Markov Random Field) based road extraction method needs excessive computation operations, which is relatively time consuming and difficult for application. This letter presents a GPU-accelerated road extraction method in PolSAR images with five stages. Firstly, the total power is calculated to form the span image. Then, pixels of the similar scattering property are included in the refined Lee filtering process. Thirdly, the edges are detected by the bi-windows edge detection algorithm. Fourthly, line pre-processing and graph construction are done to decrease the false line elements and form the line graph. Finally, the MRF is used to extract road line elements. The experiments have proved that such a GPU-accelerated approach can improve the computing efficiency, and the effectiveness is demonstrated by using the airborne PolSAR image acquired over the Oberpfaffenhofen area in Germany. Jianghua Cheng, Wenxia Ding, Xiangwei Zhu, Gui Gao |
IECON | 3 |
| 2016 | A robust real-time indoor navigation technique based on GPU-accelerated feature matchingabstractRobust feature tracking is a basic requisite for indoor navigation. The Scale Invariant Features Transform (SIFT) features are invariant to image translation, scaling, rotation, and partially invariant to illumination changes. Therefore, it is widely used for image matching based indoor navigation. However, the implementation of the traditional SIFT algorithm needs excessive computation operations, which is relatively time consuming and difficult for indoor navigation like real-time application. This article proposes a SIFT acceleration strategy based on graphics processing unit (GPU). It is proposed as the following four steps. Firstly, Gaussian pyramid is divided into different blocks. And in each GPU block, DoG (Difference of Gaussian) scale-space is computed. Secondly, local keypoints detection is done in GPU, and each keypoint is processed in one block to calculate gradient orientation and magnitude. Thirdly, the keypoint descriptor is formulated as vectors, and each vector is computed in one GPU block. Finally, GPU-accelerated SIFT is introduced into the indoor navigation system to address viewpoint changes. Our main contribution of this article is using an optimized GPU accelerated scheme for real-time indoor navigation. The experiments have proved that such approach can improve the computing efficiency and reduce the chances of mismatches. Jianghua Cheng, Xiangwei Zhu, Wenxia Ding, Gui Gao |
IPIN | 2 |