VLDB 2026 Research / reviewers in the wild / expert
Kaiwei Wang
dblp:45/11142
· DBLP profile ↗
54ranked-venue papers
7as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 6 since 2021Computer networks · 8 · 7 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Snorkel-Guided GAN-Based Framework for Robust Traffic Data Imputation in Sparse IoT Traffic Environments
Kaiwei Wang, Mona Jaber |
ICC | 1 |
| 2026 | An Augmented GNSS-DAS Architecture for Continuous and Robust Positioning
Kaiwei Wang, Ruikang Zhong, Mona Jaber, Moussa Ayyash |
ICC | 1 |
| 2026 | A Real-Time Demonstration Platform for DAS-Based Urban Traffic Monitoring
Kaiwei Wang, Chia-Yen Chiang, Mona Jaber, Ruikang Zhong, Peter Hayward |
INFOCOM | 1 |
| 2026 | Addressing inconsistency confusion: Temporal causal inference via fuzzy knowledge-data link consistency mining for root cause tracing
Pengyu Song, Kaiwei Wang |
Expert Syst. Appl. | 3 |
| 2026 | P2U-SLAM: A Monocular Wide-FoV SLAM System Based on Point Uncertainty and Pose Uncertainty
Kailun Yang 0001, Ze Wang 0009, Kaiwei Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | One-Step Event-Driven High-Speed AutofocusabstractHigh-speed autofocus in extreme scenes remains a significant challenge. Traditional methods rely on repeated sampling around the focus position, resulting in "focus hunting". Event-driven methods have advanced focusing speed and improved performance in low-light conditions; however, current approaches still require at least one lengthy round of "focus hunting", involving the collection of a complete focus stack. We introduce the Event Laplacian Product (ELP) focus detection function, which combines event data with grayscale Laplacian information, redefining focus search as a detection task. This innovation enables the first one-step event-driven autofocus, cutting focusing time by up to two-thirds and reducing focusing error by 24 times on the DAVIS346 dataset and 22 times on the EVK4 dataset. Additionally, we present an auto-focus pipeline tailored for event-only cameras, achieving accurate results across a range of challenging motion and lighting conditions. All datasets and code are available in https://github.com/YuHanBaozju/ELP. Yuhan Bao, Shaohua Gao, Wenyong Li, Kaiwei Wang |
CVPR | 4 |
| 2025 | Omnidirectional Multi-Object TrackingabstractPanoramic imagery, with its 360° field of view, offers comprehensive information to support Multi-Object Tracking (MOT) in capturing spatial and temporal relationships of surrounding objects. However, most MOT algorithms are tailored for pinhole images with limited views, impairing their effectiveness in panoramic settings. Additionally, panoramic image distortions, such as resolution loss, geometric deformation, and uneven lighting, hinder direct adaptation of existing MOT methods, leading to significant performance degradation. To address these challenges, we propose OmniTrack, an omnidirectional MOT framework that incorporates Tracklet Management to introduce temporal cues, FlexiTrack Instances for object localization and association, and the CircularStatE Module to alleviate image and geometric distortions. This integration enables tracking in panoramic field-of-view scenarios, even under rapid sensor motion. To mitigate the lack of panoramic MOT datasets, we introduce the QuadTrack dataset—a comprehensive panoramic dataset collected by a quadruped robot, featuring diverse challenges such as panoramic fields of view, intense motion, and complex environments. Extensive experiments on the public JRDB dataset and the newly introduced QuadTrack benchmark demonstrate the state-of-the-art performance of the proposed framework. OmniTrack achieves a HOTA score of 26.92% on JRDB, representing an improvement of 3.43%, and further achieves 23.45% on QuadTrack, surpassing the baseline by 6.81%. The established dataset and source code are available at https://github.com/xifen523/OmniTrack. Hao Shi 0004, Mengfei Duan, Chang Huang, Kaiwei Wang, Kailun Yang 0001 |
CVPR | 8 |
| 2025 | Low-Light Image Enhancement Using Event-Based Illumination EstimationabstractLow-light image enhancement (LLIE) aims to improve the visibility of images captured in poorly lit environments. Prevalent event-based solutions primarily utilize events triggered by motion, i.e., ''motion events'' to strengthen only the edge texture, while leaving the high dynamic range and excellent low-light responsiveness of event cameras largely unexplored. This paper instead opens a new avenue from the perspective of estimating the illumination using ''temporal-mapping'' events, i.e., by converting the timestamps of events triggered by a transmittance modulation into brightness values. The resulting fine-grained illumination cues facilitate a more effective decomposition and enhancement of the reflectance component in low-light images through the proposed Illumination-aided Reflectance Enhancement module. Furthermore, the degradation model of temporal-mapping events under low-light condition is investigated for realistic training data synthesizing. To address the lack of datasets under this regime, we construct a beam-splitter setup and collect EvLowLight dataset that includes images, temporal-mapping events, and motion events. Extensive experiments across 5 synthetic datasets and our real-world EvLowLight dataset substantiate that the devised pipeline, dubbed RetinEV, excels in producing well-illuminated, high dynamic range images, outperforming previous state-of-the-art event-based methods by up to 6.62 dB, while maintaining an efficient inference speed of 35.6 frame-per-second on a 640X480 image. Lei Sun 0009, Yuhan Bao, Jiajun Zhai, Jingyun Liang, Yulun Zhang 0001, Kaiwei Wang, Danda Pani Paudel, Luc Van Gool |
ICCV | 6 |
| 2025 | SF-TIM: A Simple Framework for Enhancing Quadrupedal Robot Jumping Agility by Combining Terrain Imagination and MeasurementabstractDynamic jumping on high platforms and over gaps differentiates legged robots from wheeled counterparts. Compared to walking on rough terrains, dynamic locomotion on abrupt surfaces requires fusing proprioceptive and exteroceptive perception for explosive movements. In this paper, we propose SF-TIM (Simple Framework combining Terrain Imagination and Measurement), a single-policy method that enhances quadrupedal robot jumping agility, while preserving their fundamental blind walking capabilities. In addition, we introduce a terrain-guided reward design specifically to assist quadrupedal robots in high jumping, improving their performance in this task. To narrow the simulation-to-reality gap in quadrupedal robot learning, we introduce a stable and high-speed elevation map generation framework, enabling zero-shot simulation-to-reality transfer of locomotion ability. Our algorithm has been deployed and validated on both the small-/large-size quadrupedal robots, demonstrating its effectiveness in real-world applications: the robot has successfully traversed various high platforms and gaps, showing the robustness of our proposed approach. A demo video has been made available at https://flysoaryun.github.io/SF-TIM. Ze Wang 0009, Long Xu 0002, Hao Shi 0004, Zunwang Ma, Zhen Chu, Fei Gao 0011, Kailun Yang 0001, Kaiwei Wang |
IROS | 10 |
| 2025 | EgoEvGesture: Gesture Recognition Based on Egocentric Event CameraabstractEgocentric gesture recognition is a pivotal technology for enhancing natural human-computer interaction, yet traditional RGB-based solutions suffer from motion blur and illumination variations in dynamic scenarios. While event cameras show distinct advantages in handling high dynamic range with ultra-low power consumption, existing RGB-based architectures face inherent limitations in processing asynchronous event streams due to their synchronous frame-based nature. Moreover, from an egocentric perspective, event cameras record data that includes events generated by both head movements and hand gestures, thereby increasing the complexity of gesture recognition. To address this, we propose a novel network architecture specifically designed for event data processing, incorporating (1) a lightweight CNN with asymmetric depthwise convolutions to reduce parameters while preserving spatiotemporal features, (2) a plug-and-play state-space model as context block that decouples head movement noise from gesture dynamics, and (3) a parameter-free Bins-Temporal Shift Module (BTSM) that shifts features along bins and temporal dimensions to fuse sparse events efficiently. We further establish the EgoEvGesture dataset, the first large-scale dataset for egocentric gesture recognition using event cameras. Experimental results demonstrate that our method achieves 62.7% accuracy tested on unseen subjects with only 7M parameters, 3.1% higher than state-of-the-art approaches. Notable misclassifications in freestyle motions stem from high interpersonal variability and unseen test patterns differing from training data. Moreover, our approach achieved a remarkable accuracy of 97.0% on the DVS128 Gesture, demonstrating the effectiveness and generalization capability of our method on public datasets. The dataset and models are made available at https://github.com/3190105222/EgoEv_Gesture. Luming Wang, Hao Shi 0004, Xiaoting Yin, Kailun Yang 0001, Kaiwei Wang |
SMC | 5 |
| 2025 | EI-Nexus: Towards Unmediated and Flexible Inter-Modality Local Feature Extraction and Matching for Event-Image DataabstractEvent cameras, with high temporal resolution and high dynamic range, have limited research on the inter-modality local feature extraction and matching of event-image data. We propose EI-Nexus, an unmediated and flexible framework that integrates two modality-specific keypoint extractors and a feature matcher. To achieve keypoint extraction across viewpoint and modality changes, we bring Local Feature Distillation (LFD), which transfers the viewpoint consistency from a well-learned image extractor to the event extractor, ensuring robust feature correspondence. Furthermore, with the help of Context Aggregation (CA), a remarkable enhancement is observed in feature matching. We further establish the first two inter-modality feature matching benchmarks, MVSEC-RPE and EC-RPE, to assess relative pose estimation on event-image data. Our approach outperforms traditional methods that rely on explicit modal transformation, offering more unmediated and adaptable feature extraction and matching, achieving better keypoint similarity and state-of-the-art results on the MVSEC-RPE and ECRPE benchmarks. The source code and benchmarks will be made publicly available at EI-Nexus. Zhonghua Yi, Hao Shi 0004, Kailun Yang 0001, Ze Wang 0009, Diyang Gu, Kaiwei Wang |
WACV | 8 |
| 2025 | A Unified Framework for Event-Based Frame Interpolation With Ad-Hoc Deblurring in the WildabstractEffective video frame interpolation hinges on the adept handling of motion in the input scene. Prior work acknowledges asynchronous event information for this, but often overlooks whether motion induces blur in the video, limiting its scope to sharp frame interpolation. We instead propose a unified framework for event-based frame interpolation that performs deblurring ad-hoc and thus works both on sharp and blurry input videos. Our model consists in a bidirectional recurrent network that incorporates the temporal dimension of interpolation and fuses information from the input frames and the events adaptively based on their temporal proximity. To enhance the generalization from synthetic data to real event cameras, we integrate self-supervised framework with the proposed model to enhance the generalization on real-world datasets in the wild. At the dataset level, we introduce a novel real-world high-resolution dataset with events and color videos named HighREV, which provides a challenging evaluation setting for the examined task. Extensive experiments show that our network consistently outperforms previous state-of-the-art methods on frame interpolation, single image deblurring, and the joint task of both. Experiments on domain transfer reveal that self-supervised training effectively mitigates the performance degradation observed when transitioning from synthetic data to real-world data. Code and datasets are available at https://github.com/AHupuJR/REFID. Lei Sun 0009, Daniel Gehrig, Christos Sakaridis, Mathias Gehrig, Jingyun Liang, Zhijie Xu, Kaiwei Wang, Luc Van Gool, Davide Scaramuzza 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2025 | Offboard Occupancy Refinement With Hybrid Propagation for Autonomous DrivingabstractVision-based occupancy prediction, also known as 3D Semantic Scene Completion (SSC), presents a significant challenge in computer vision. Previous methods, confined to onboard processing, struggle with simultaneous geometric and semantic estimation, continuity across varying viewpoints, and single-view occlusion. Our paper introduces OccFiner, a novel offboard framework designed to enhance the accuracy of vision-based occupancy predictions. OccFiner operates in two hybrid phases: 1) a multi-to-multi local propagation network that implicitly aligns and processes multiple local frames for correcting onboard model errors and consistently enhancing occupancy accuracy across all distances. 2) the region-centric global propagation, focuses on refining labels using explicit multi-view geometry and integrating sensor bias, particularly for increasing the accuracy of distant occupied voxels. Extensive experiments demonstrate that OccFiner improves both geometric and semantic accuracy across various types of coarse occupancy, setting a new state-of-the-art performance on the SemanticKITTI dataset. Notably, OccFiner significantly boosts the performance of vision-based SSC models, achieving accuracy levels competitive with established LiDAR-based onboard SSC methods. Furthermore, OccFiner is the first to achieve automatic annotation of SSC in a purely vision-based approach. Quantitative experiments prove that OccFiner successfully facilitates occupancy data loop-closure in autonomous driving. Additionally, we quantitatively and qualitatively validate the superiority of the offboard approach on city-level SSC static maps. The source code will be made publicly available at https://github.com/MasterHow/OccFiner Hao Shi 0004, Song Wang 0019, Jiaming Zhang 0001, Xiaoting Yin, Guangming Wang 0001, Jianke Zhu, Kailun Yang 0001, Kaiwei Wang |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2024 | Temporal-Mapping Photography for Event Cameras
Yuhan Bao, Lei Sun 0009, Yuqin Ma, Kaiwei Wang |
ECCV (55) | 4 |
| 2024 | Simultaneous Time Synchronization and Mutual Localization for Multi-robot SystemabstractMutual localization stands as a foundational component within various domains of multi-robot systems. Nevertheless, in relative pose estimation, time synchronization is usually underappreciated and rarely addressed, although it significantly influences estimation accuracy. In this paper, we introduce time synchronization into mutual localization to recover the time offset and relative poses between robots simultaneously. Under a constant velocity assumption in a short time, we fuse time offset estimation with our previous bearing-based mutual localization by a novel error representation. Based on the error model, we formulate a joint optimization problem and utilize semi-definite relaxation (SDR) to furnish a lossless relaxation. By solving the relaxed problem, time synchronization and relative pose estimation can be achieved when time drift between robots is limited. To enhance the application range of time offset estimation, we further propose an iterative method to recover the time offset from coarse to fine. Comparisons between the proposed method and the existing ones through extensive simulation tests present prominent benefits of time synchronization on mutual localization. Moreover, real-world experiments are conducted to show the practicality and robustness. Xiangyong Wen, Yingjian Wang 0001, Kaiwei Wang, Chao Xu 0001, Fei Gao 0011 |
ICRA | 4 |
| 2024 | MambaMOS: LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space ModelabstractLiDAR-based Moving Object Segmentation (MOS) aims to locate and segment moving objects in point clouds of the current scan using motion information from previous scans. Despite the promising results achieved by previous MOS methods, several key issues, such as the weak coupling of temporal and spatial information, still need further study. In this paper, we propose a novel LiDAR-based 3D Moving Object Segmentation with Motion-aware State Space Model, termed MambaMOS. Firstly, we develop a novel embedding module, the Time Clue Bootstrapping Embedding (TCBE), to enhance the coupling of temporal and spatial information in point clouds and alleviate the issue of overlooked temporal clues. Secondly, we introduce the Motion-aware State Space Model (MSSM) to endow the model with the capacity to understand the temporal correlations of the same object across different time steps. Specifically, MSSM emphasizes the motion states of the same object at different time steps through two distinct temporal modeling and correlation steps. We utilize an improved state space model to represent these motion differences, significantly modeling the motion states. Finally, extensive experiments on the SemanticKITTI-MOS and KITTI-Road benchmarks demonstrate that the proposed MambaMOS achieves state-of-the-art performance. The source code is publicly available at https://github.com/Terminal-K/MambaMOS Kang Zeng, Hao Shi 0004, Jiacheng Lin, Siyu Li 0002, Jintao Cheng, Kaiwei Wang, Zhiyong Li 0001, Kailun Yang 0001 |
ACM Multimedia | 6 |
| 2024 | Exploring event-based human pose estimation with 3D event representations
Xiaoting Yin, Hao Shi 0004, Jiaan Chen, Ze Wang 0009, Yaozu Ye, Kailun Yang 0001, Kaiwei Wang |
Comput. Vis. Image Underst. | 7 |
| 2024 | Human Gait Recognition Based on Frontal-View Walking Sequences Using Multi-modal Feature Representations and LearningabstractAbstract Despite that much progress has been reported in gait recognition, most of these existing works adopt lateral-view parameters as gait features, which requires large area of data collection environment and limits the applications of gait recognition in real-world practice. In this paper, we adopt frontal-view walking sequences rather than lateral-view sequences and propose a new gait recognition method based on multi-modal feature representations and learning. Specifically, we characterize walking sequences with two different kinds of frontal-view gait features representations, including holistic silhouette and dense optical flow. Pedestrian regions extraction is achieved by an improved YOLOv7 algorithm called Gait-YOLO algorithm to eliminate the effects of background interference. Multi-modal fusion module (MFM) is proposed to explore the intrinsic connections between silhouette and dense optical flow features by using squeeze and excitation operations at the channel and spatial levels. Gait feature encoder is further used to extract global walking characteristics, enabling efficient multi-modal information fusion. To validate the efficacy of the proposed method, we conduct experiments on CASIA-B and OUMVLP gait databases and compare performance of our proposed method with other existing state-of-the-art gait recognition methods. Muqing Deng, Zebang Zhong, Kaiwei Wang, Junrong Liao |
Neural Process. Lett. | 5 |
| 2024 | Behind Every Domain There is a Shift: Adapting Distortion-Aware Vision Transformers for Panoramic Semantic SegmentationabstractIn this paper, we address panoramic semantic segmentation which is under-explored due to two critical challenges: (1) image distortions and object deformations on panoramas; (2) lack of semantic annotations in the$360^\circ$imagery. To tackle these problems, first, we propose the upgraded Transformer for Panoramic Semantic Segmentation, ie, Trans4PASS+, equipped withDeformable Patch Embedding (DPE)andDeformable MLP (DMLPv2)modules for handling object deformations and image distortions whenever (before or after adaptation) and wherever (shallow or deep levels). Second, we enhance theMutual Prototypical Adaptation (MPA)strategy via pseudo-label rectification for unsupervised domain adaptive panoramic segmentation. Third, aside from Pinhole-to-Panoramic (Pin2Pan) adaptation, we create a new dataset (SynPASS) with 9,080 panoramic images, facilitating Synthetic-to-Real (Syn2Real) adaptation scheme in$360^\circ$imagery. Extensive experiments are conducted, which cover indoor and outdoor scenarios, and each of them is investigated withPin2PanandSyn2Realregimens. Trans4PASS+ achieves state-of-the-art performances on four domain adaptive panoramic semantic segmentation benchmarks. Code is available athttps://github.com/jamycheung/Trans4PASS. Jiaming Zhang 0001, Kailun Yang 0001, Hao Shi 0004, Simon Reiß, Kunyu Peng, Chaoxiang Ma, Haodong Fu, Philip Torr 0001, Kaiwei Wang, Rainer Stiefelhagen |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2024 | LF-VISLAM: A SLAM Framework for Large Field-of-View Cameras With Negative Imaging Plane on Mobile AgentsabstractSimultaneous Localization And Mapping (SLAM) has become a crucial aspect in the fields of autonomous driving and robotics. One crucial component of visual SLAM is the Field-of-View (FoV) of the camera, as a larger FoV allows for a wider range of surrounding elements and features to be perceived. However, when the FoV of the camera reaches the negative half-plane, traditional methods for representing image feature points using$[u,v,1]^T$become ineffective. While the panoramic FoV is advantageous for loop closure, its benefits are not easily realized under large-attitude-angle differences where loop-closure frames cannot be easily matched by existing methods. As loop closure on wide-FoV panoramic data further comes with a large number of outliers, traditional outlier rejection methods are not directly applicable. To address these issues, we propose LF-VISLAM, aVisualInertialSLAMframework for cameras with extremelyLargeFoV with loop closure. A three-dimensional vector with unit length is introduced to effectively represent feature points even on the negative half-plane. The attitude information of the SLAM system is leveraged to guide the feature point detection of the loop closure. Additionally, a new outlier rejection method based on the unit length representation is integrated into the loop closure module. We collect the PALVIO dataset using aPanoramicAnnularLens (PAL) system with an entire FoV of$360^\circ{\times}(40^\circ{\sim}120^\circ)$and an Inertial Measurement Unit (IMU) forVisualInertialOdometry (VIO) to address the lack of panoramic SLAM datasets. Experiments on the established PALVIO and public datasets show that the proposed LF-VISLAM outperforms state-of-the-art SLAM methods. Our code will be open-sourced at https://github.com/flysoaryun/LF-VISLAM.Note to Practitioners—Motivated by the challenges of handling large-FoV cameras in SLAM applications, this paper proposes LF-VISLAM, a novel SLAM framework that uses a large-FoV camera and IMU sensors. Our framework is equipped with a loop closure thread that can use attitude information to eliminate accumulated errors. We have made algorithmic adjustments and optimizations to the negative half-plane features to better adapt to cameras with large FoV. Experimental evaluations demonstrate that LF-VISLAM significantly outperforms traditional SLAM methods. Additionally, the code will be open-sourced, providing easy access for research and implementation. Overall, LF-VISLAM is a promising solution to improve the performance of SLAM in challenging environments of autonomous driving and robotics. Ze Wang 0009, Kailun Yang 0001, Hao Shi 0004, Peng Li 0034, Fei Gao 0011, Kaiwei Wang |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2024 | Minimalist and High-Quality Panoramic Imaging With PSF-Aware TransformersabstractHigh-quality panoramic images with a Field of View (FoV) of 360° are essential for contemporary panoramic computer vision tasks. However, conventional imaging systems come with sophisticated lens designs and heavy optical components. This disqualifies their usage in many mobile and wearable applications where thin and portable, minimalist imaging systems are desired. In this paper, we propose a Panoramic Computational Imaging Engine (PCIE) to achieve minimalist and high-quality panoramic imaging. With less than three spherical lenses, a Minimalist Panoramic Imaging Prototype (MPIP) is constructed based on the design of the Panoramic Annular Lens (PAL), but with low-quality imaging results due to aberrations and small image plane size. We propose two pipelines, i.e. Aberration Correction (AC) and Super-Resolution and Aberration Correction (SR&AC), to solve the image quality problems of MPIP, with imaging sensors of small and large pixel size, respectively. To leverage the prior information of the optical system, we propose a Point Spread Function (PSF) representation method to produce a PSF map as an additional modality. A PSF-aware Aberration-image Recovery Transformer (PART) is designed as a universal network for the two pipelines, in which the self-attention calculation and feature extraction are guided by the PSF map. We train PART on synthetic image pairs from simulation and put forward the PALHQ dataset to fill the gap of real-world high-quality PAL images for low-level vision. A comprehensive variety of experiments on synthetic and real-world benchmarks demonstrates the impressive imaging results of PCIE and the effectiveness of the PSF representation. We further deliver heuristic experimental findings for minimalist and high-quality panoramic imaging, in terms of the choices of prototype and pipeline, network architecture, training strategies, and dataset construction. Our dataset and code will be available at https://github.com/zju-jiangqi/PCIE-PART. Shaohua Gao, Kailun Yang 0001, Zhonghua Yi, Hao Shi 0004, Lei Sun 0009, Kaiwei Wang |
IEEE Trans. Image Process. | 8 |
| 2024 | CoBEV: Elevating Roadside 3D Object Detection With Depth and Height ComplementarityabstractRoadside camera-driven 3D object detection is a crucial task in intelligent transportation systems, which extends the perception range beyond the limitations of vision-centric vehicles and enhances road safety. While previous studies have limitations in using only depth or height information, we find both depth and height matter and they are in fact complementary. The depth feature encompasses precise geometric cues, whereas the height feature is primarily focused on distinguishing between various categories of height intervals, essentially providing semantic context. This insight motivates the development of Complementary-BEV (CoBEV), a novel end-to-end monocular 3D object detection framework that integrates depth and height to construct robust BEV representations. In essence, CoBEV estimates each pixel's depth and height distribution and lifts the camera features into 3D space for lateral fusion using the newly proposed two-stage complementary feature selection (CFS) module. A BEV feature distillation framework is also seamlessly integrated to further enhance the detection accuracy from the prior knowledge of the fusion-modal CoBEV teacher. We conduct extensive experiments on the public 3D detection benchmarks of roadside camera-based DAIR-V2X-I and Rope3D, as well as the private Supremind-Road dataset, demonstrating that CoBEV not only achieves the accuracy of the new state-of-the-art, but also significantly advances the robustness of previous methods in challenging long-distance scenarios and noisy camera disturbance, and enhances generalization by a large margin in heterologous settings with drastic changes in scene and camera parameters. For the first time, the vehicle AP score of a camera model reaches 80% on DAIR-V2X-I in terms of easy mode. The source code will be made publicly available at CoBEV. Hao Shi 0004, Chengshan Pang, Jiaming Zhang 0001, Kailun Yang 0001, Huajian Ni, Yining Lin, Rainer Stiefelhagen, Kaiwei Wang |
IEEE Trans. Image Process. | 9 |
| 2023 | Delivering Arbitrary-Modal Semantic SegmentationabstractMultimodal fusion can make semantic segmentation more robust. However, fusing an arbitrary number of modalities remains underexplored. To delve into this problem, we create the Deliver arbitrary-modal segmentation benchmark, covering Depth, LiDAR, multiple Views, Events, and RGB. Aside from this, we provide this dataset in four severe weather conditions as well as five sensor failure cases to exploit modal complementarity and resolve partial outages. To make this possible, we present the arbitrary cross-modal segmentation model CMNEXT. It encompasses a Self-Query Hub (SQ-Hub) designed to extract effective information from any modality for subsequent fusion with the RGB representation and adds only negligible amounts of parameters ~0.1M) per additional modality. On top, to efficiently and flexibly harvest discriminative cues from the auxiliary modalities, we introduce the simple Parallel Pooling Mixer (PPX). With extensive experiments on a total of six benchmarks, our CMNEXT achieves state-of-the-art performance on the Deliver, Kitti-360, MFNet, NYU Depth V2, UrbanLF, and MCubeS datasets, allowing to scale from 1 to 81 modalities. On the freshly collected Deliver, the quad-modal CMNEXT reaches up to 66.30% in mIoU with a +9.10% gain as compared to the mono-modal baseline.11The Deliver dataset and our code will be made publicly available at https://jamycheung.github.io/DELIVER.html. Jiaming Zhang 0001, Ruiping Liu 0001, Hao Shi 0004, Kailun Yang 0001, Simon Reiß, Kunyu Peng, Haodong Fu, Kaiwei Wang, Rainer Stiefelhagen |
CVPR | 8 |
| 2023 | Event-Based Frame Interpolation with Ad-hoc DeblurringabstractThe performance of video frame interpolation is inherently correlated with the ability to handle motion in the input scene. Even though previous works recognize the utility of asynchronous event information for this task, they ignore the fact that motion may or may not result in blur in the input video to be interpolated, depending on the length of the exposure time of the frames and the speed of the motion, and assume either that the input video is sharp, restricting themselves to frame interpolation, or that it is blurry, including an explicit, separate deblurring stage before interpolation in their pipeline. We instead propose a general method for event-based frame interpolation that performs deblurring ad-hoc and thus works both on sharp and blurry input videos. Our model consists in a bidirectional recurrent network that naturally incorporates the temporal dimension of interpolation and fuses information from the input frames and the events adaptively based on their temporal proximity. In addition, we introduce a novel real-world high-resolution dataset with events and color videos named HighREV, which provides a challenging evaluation setting for the examined task. Extensive experiments on the standard CoPro benchmark and on our dataset show that our network consistently outperforms previous state-of-the-art methods on frame interpolation, single image deblurring and the joint task of interpolation and deblurring. Our code and dataset are available at https://github.com/AHupuJR/REFID. Lei Sun 0009, Christos Sakaridis, Jingyun Liang, Kai Zhang 0008, Jiezhang Cao, Kaiwei Wang, Luc Van Gool |
CVPR | 8 |
| 2023 | LEO Mega-Constellations Routing Algorithm Based on Area SegmentationabstractLow earth orbit (LEO) mega-constellations in future 6G have attracted the attention from both academia and industry. However, due to the high dynamic characteristic of the LEO satellite network topology and the limited on-board resources, existing approaches relying on high on-board processing capabilities and monitoring the global network state, result in intolerable packet loss rate and excessive signalling overhead. In this paper, a routing algorithm based on area segmentation for LEO mega-constellations is proposed according to the topological characteristics of the network. Specifically, we divide the LEO mega-constellations into multiple areas with four quadrant parts. In addition, the transmission cluster is defined consisted of two adjacent parts based on transmission direction. With the relative geographical location and transmission clusters, we joint intra-area and inter-area routing to realize multi-path routing and forwarding by periodically updating the link state, instead of globally calculating. Simulation results demonstrate that the proposed algorithm can achieve higher throughput, decrease the packet loss rate by 22% and reduce the signalling overhead significantly. Jiaxin Zhang 0001, Shuang Zheng 0005, Kaiwei Wang, Peng Wang 0062, Xing Zhang 0001 |
WCNC | 4 |
| 2023 | PanoFlow: Learning 360° Optical Flow for Surrounding Temporal UnderstandingabstractOptical flow estimation is a basic task in self-driving and robotics systems, which enables to temporally interpret traffic scenes. Autonomous vehicles clearly benefit from the ultra-wide Field of View (FoV) offered by 360° panoramic sensors. However, due to the unique imaging process of panoramic cameras, models designed for pinhole images do not directly generalize satisfactorily to 360° panoramic images. In this paper, we put forward a novel network framework——PANO FLOW, to learn optical flow for panoramic images. To overcome the distortions introduced by equirectangular projection in panoramic transformation, we design a Flow Distortion Augmentation (FDA) method, which contains radial flow distortion (FDA-R) or equirectangular flow distortion (FDA-E). We further look into the definition and properties of cyclic optical flow for panoramic videos, and hereby propose a Cyclic Flow Estimation (CFE) method by leveraging the cyclicity of spherical images to infer 360° optical flow and converting large displacement to relatively small displacement. PanoFlow is applicable to any existing flow estimation method and benefits from the progress of narrow-FoV flow estimation. In addition, we create and release a synthetic panoramic dataset FlowScape based on CARLA to facilitate training and quantitative analysis. PanoFlow achieves state-of-the-art performance on the public OmniFlowNet and the fresh established FlowScape benchmarks. Our proposed approach reduces the End-Point-Error (EPE) on FlowScape by 27.3%. On OmniFlowNet, PanoFlow achieves an EPE of 3.17 pixels, a 55.5% error reduction from the best published result (7.12 pixels). We also qualitatively validate our method via an outdoor collection vehicle and a public real-world OmniPhotos dataset, indicating strong potential and robustness for real-world navigation applications. Code and dataset are publicly available at PanoFlow. Hao Shi 0004, Yifan Zhou 0001, Kailun Yang 0001, Xiaoting Yin, Ze Wang 0009, Yaozu Ye, Shi Meng, Peng Li 0034, Kaiwei Wang |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2022 | Efficient Human Pose Estimation via 3D Event Point CloudabstractHuman Pose Estimation (HPE) based on RGB images has experienced a rapid development benefiting from deep learning. However, event-based HPE has not been fully studied, which remains great potential for applications in extreme scenes and efficiency-critical conditions. In this paper, we are the first to estimate 2D human pose directly from 3D event point cloud. We propose a novel representation of events, the rasterized event point cloud, aggregating events on the same position of a small time slice. It maintains the 3D features from multiple statistical cues and significantly reduces memory consumption and computation complexity, proved to be efficient in our work. We then leverage the rasterized event point cloud as input to three different back-bones, PointNet, DGCNN, and Point Transformer, with two linear layer decoders to predict the location of human key-points. We find that based on our method, PointNet achieves promising results with much faster speed, whereas Point Transfomer reaches much higher accuracy, even close to previous event-frame-based methods. A comprehensive set of results demonstrates that our proposed method is consistently effective for these 3D backbone models in event-driven human pose estimation. Our method based on PointNet with 2048 points input achieves 82.46mm in MPJPE3Don the DHP19 dataset, while only has a latency of 12.29ms on an NVIDIA Jetson Xavier NX edge computing platform, which is ideally suitable for real-time detection with event cameras. Code is available at https: //github.com/ MasterHow/EventPointPose. Jiaan Chen, Hao Shi 0004, Yaozu Ye, Kailun Yang 0001, Lei Sun 0009, Kaiwei Wang |
3DV | 6 |
| 2022 | Event-Based Fusion for Motion Deblurring with Cross-modal Attention
Lei Sun 0009, Christos Sakaridis, Jingyun Liang, Kailun Yang 0001, Yaozu Ye, Kaiwei Wang, Luc Van Gool |
ECCV (18) | 8 |
| 2022 | LF-VIO: A Visual-Inertial-Odometry Framework for Large Field-of-View Cameras with Negative PlaneabstractVisual-inertial-odometry has attracted extensive attention in the field of autonomous driving and robotics. The size of Field of View (FoV) plays an important role in Visual-Odometry (VO) and Visual-Inertial-Odometry (VIO), as a large FoV enables to perceive a wide range of surrounding scene elements and features. However, when the field of the camera reaches the negative half plane, one cannot simply use$[u, v, 1]^{T}$to represent the image feature points anymore. To tackle this issue, we propose LF-VIO, a real-time VIO framework for cameras with extremely large FoV.We leverage a threedimensional vector with unit length to represent feature points, and design a series of algorithms to overcome this challenge. To address the scarcity of panoramic visual odometry datasets with ground-truth location and pose, we present the PALVIO dataset, collected with a Panoramic Annular Lens (PAL) system with an entire FoV of 36$0^{\circ}\times(40^{\circ}\sim 120^{\circ})$and an IMU sensor. With a comprehensive variety of experiments, the proposed LF-VIO is verified on both the established PALVIO benchmark and a public fisheye camera dataset with a FoV of$360^{\circ}\times(0^{\circ}\sim 93.5^{\circ})$. LF-VIO outperforms state-of-the-art visual-inertial-odometry methods. Our dataset and code are made publicly available at https://github.com/flysoaryun/LF-VIO Ze Wang 0009, Kailun Yang 0001, Hao Shi 0004, Peng Li 0034, Fei Gao 0011, Kaiwei Wang |
IROS | 6 |
| 2022 | CSFlow: Learning Optical Flow via Cross Strip Correlation for Autonomous DrivingabstractOptical flow estimation is an essential task in self-driving systems, which helps autonomous vehicles perceive temporal continuity information of surrounding scenes. The calculation of all-pair correlation plays an important role in many existing state-of-the-art optical flow estimation methods. However, the reliance on local knowledge often limits the model’s accuracy under complex street scenes. In this paper, we propose a new deep network architecture for optical flow estimation in autonomous driving——CSFlow, which consists of two novel modules: Cross Strip Correlation module (CSC) and Correlation Regression Initialization module (CRI). CSC utilizes a striping operation across the target image and the attended image to encode global context into correlation volumes, while maintaining high efficiency. CRI is used to maximally exploit the global context for optical flow initialization. Our method has achieved state-of-the-art accuracy on the public autonomous driving dataset KITTI-2015. Code is publicly available at https://github.com/MasterHow/CSFlow. Hao Shi 0004, Yifan Zhou 0001, Kailun Yang 0001, Xiaoting Yin, Kaiwei Wang |
IV | 5 |
| 2022 | Omnisupervised Omnidirectional Semantic SegmentationabstractModern efficient Convolutional Neural Networks (CNNs) are able to perform semantic segmentation both swiftly and accurately, which covers typically separate detection tasks desired by Intelligent Vehicles (IV) in a unified way. Most of the current semantic perception frameworks are designed to work with pinhole cameras and benchmarked against public datasets with narrow Field-of-View (FoV) images. However, there is a large accuracy downgrade when a pinhole-yielded CNN is taken to omnidirectional imagery, causing it unreliable for surrounding perception. In this paper, we propose an omnisupervised learning framework for efficient CNNs, which bridges multiple heterogeneous data sources that are already available in the community, bypassing the labor-intensive process to have manually annotated panoramas, while improving their reliability in unseen omnidirectional domains. Being omnisupervised, the efficient CNN exploits both labeled pinhole images and unlabeled panoramas. The framework is based on our specialized ensemble method that considers the wide-angle and wrap-around features of omnidirectional images, to automatically generate panoramic labels for data distillation. A comprehensive variety of experiments demonstrates that the proposed solution helps to attain significant generalizability gains in panoramic imagery domains. Our approach outperforms state-of-the-art efficient segmenters on highly unconstrained IDD20K and PASS datasets. Kailun Yang 0001, Xinxin Hu, Yicheng Fang, Kaiwei Wang, Rainer Stiefelhagen |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Hierarchical visual localization for visually impaired people using multimodal images
Ruiqi Cheng, Weijian Hu, Yicheng Fang, Kaiwei Wang, Zhijie Xu |
Expert Syst. Appl. | 5 |
| 2020 | DS-PASS: Detail-Sensitive Panoramic Annular Semantic Segmentation through SwaftNet for Surrounding SensingabstractSemantically interpreting the traffic scene is crucial for autonomous transportation and robotics systems. However, state-of-the-art semantic segmentation pipelines are dominantly designed to work with pinhole cameras and train with narrow Field-of-View (FoV) images. In this sense, the perception capacity is severely limited to offer higher-level confidence for upstream navigation tasks. In this paper, we propose a network adaptation framework to achieve Panoramic Annular Semantic Segmentation (PASS), which allows to re-use conventional pinhole-view image datasets, enabling modern segmentation networks to comfortably adapt to panoramic images. Specifically, we adapt our proposed SwaftNet to enhance the sensitivity to details by implementing attention-based lateral connections between the detail-critical encoder layers and the context-critical decoder layers. We benchmark the performance of efficient segmenters on panoramic segmentation with our extended PASS dataset, demonstrating that the proposed real-time SwaftNet outperforms state-of-the-art efficient networks. Furthermore, we assess real-world performance when deploying the Detail-Sensitive PASS (DS-PASS) system on a mobile robot and an instrumented vehicle, as well as the benefit of panoramic semantics for visual odometry, showing the robustness and potential to support diverse navigational applications. Kailun Yang 0001, Xinxin Hu, Kaite Xiang, Kaiwei Wang, Rainer Stiefelhagen |
IV | 5 |
| 2020 | In Defense of Multi-Source Omni-Supervised Efficient ConvNet for Robust Semantic Segmentation in Heterogeneous Unseen DomainsabstractSemantic segmentation renders a unified way of surrounding perception, where most of driving scene detection tasks can be covered by running a single efficient ConvNet through a forward pass. However, current frameworks posit the closed-world paradigm expressed as a single source of distribution over a predetermined set of visual classes, forgetting that a deep model must be deployed in the wild facing unseen domains and unforeseen hazards. In spite of being accurate in its comfort zone, the segmentation model may not generalize well to a new domain. In addition, a model trained with single dataset is heavily limited in terms of recognizable classes. In this paper, we propose an omni-supervised learning framework for semantic segmentation which is able to leverage heterogeneous data sources. Our omni-supervised training framework incorporates all available labeled and unlabeled data, meanwhile bridges multiple training sets to be capable of recognizing more classes that are needed for autonomous navigation application at hand in the new domain. A comprehensive variety of experiments shows that with the proposed multi-source omni-supervised learning solution, an efficient ConvNet like our ERF-PSPNet attains significant robustness gains in open domains that are of critical relevance to real deployment of vision algorithms. Our approach surpasses the state of the art on the highly unconstrained PASS and IDD20K datasets. Kailun Yang 0001, Xinxin Hu, Kaiwei Wang, Rainer Stiefelhagen |
IV | 3 |
| 2020 | CFVL: A Coarse-to-Fine Vehicle Localizer with Omnidirectional Perception across Severe Appearance VariationsabstractVisual localization in vehicle navigation remains a crucial image retrieval task to determine the best matched image. Developing an efficient algorithm to address the localization issues of vehicle is highly difficult, for severe appearance variations with vehicles moving around can bring about significant challenges and big obstacles. In this paper, we propose the CFVL framework which takes panoramas into use in the localizer and the system processes from coarse to fine, in order to attain more robust and stable descriptors. NetVALD descriptors based on explicit panorama construction, which are regarded robust to appearance changes, are extracted in the coarse stage, while Geodesc keypoint descriptors, which are believed to detect more detailed information, are utilized in the fine stage, so as to perceive the accurate localization. A comprehensive set of experiments is carried on several datasets with different appearances across seasonal cycling, illumination variations, diverse traversals, and so on, to verify the effectiveness of the coarse stage and fine stage in our system. Brute Force (BF) matching and Fundamental Matrix mapping are utilized to match and locate correct locations after coarse stage and after fine stage. The accuracy of the coarse matching and fine matching are verified separately. Our system is demonstrated to be with high location recall, generalization capacity across different environments. Yicheng Fang, Kaiwei Wang, Ruiqi Cheng, Kailun Yang 0001 |
IV | 2 |
| 2020 | AttenNet: Deep Attention Based Retinal Disease Classification in OCT Images
Jun Wu 0022, Jianchun Zhao, Dayong Ding, Ningjiang Chen, Chunhui Jiang, Xuan Zou, Yuan Tian 0017, Zongjiang Shang, Kaiwei Wang, Xirong Li 0001, Gang Yang 0001, Jianping Fan 0001 |
MMM (2) | 15 |
| 2020 | Universal Semantic Segmentation for Fisheye Urban Driving ImagesabstractSemantic segmentation is a critical method in the field of autonomous driving. When performing semantic image segmentation, a wider field of view (FoV) helps to obtain more information about the surrounding environment, making automatic driving safer and more reliable, which could be offered by fisheye cameras. However, large public fisheye datasets are not available, and the fisheye images captured by the fisheye camera with large FoV comes with large distortion, so commonly-used semantic segmentation model cannot be directly utilized. In this paper, a seven degrees of freedom (DoF) augmentation method is proposed to transform rectilinear image to fisheye image in a more comprehensive way. In the training process, rectilinear images are transformed into fisheye images in seven DoF, which simulates the fisheye images taken by cameras of different positions, orientations and focal lengths. The result shows that training with the seven-DoF augmentation can improve the model's accuracy and robustness against different distorted fisheye data. This seven-DoF augmentation provides a universal semantic segmentation solution for fisheye cameras in different autonomous driving applications. Also, we provide specific parameter settings of the augmentation for autonomous driving. At last, we tested our universal semantic segmentation model on real fisheye images and obtained satisfactory results. The code and configurations are released at https://github.com/Yaozhuwa/FisheyeSeg. Yaozu Ye, Kailun Yang 0001, Kaite Xiang, Kaiwei Wang |
SMC | 5 |
| 2020 | PASS: Panoramic Annular Semantic SegmentationabstractPixel-wise semantic segmentation is capable of unifying most of driving scene perception tasks, and has enabled striking progress in the context of navigation assistance, where an entire surrounding sensing is vital. However, current mainstream semantic segmenters are predominantly benchmarked against datasets featuring narrow Field of View (FoV), and a large part of vision-based intelligent vehicles use only a forward-facing camera. In this paper, we propose a Panoramic Annular Semantic Segmentation (PASS) framework to perceive the whole surrounding based on a compact panoramic annular lens system and an online panorama unfolding process. To facilitate the training of PASS models, we leverage conventional FoV imaging datasets, bypassing the efforts entailed to create fully dense panoramic annotations. To consistently exploit the rich contextual cues in the unfolded panorama, we adapt our real-time ERF-PSPNet to predict semantically meaningful feature maps in different segments, and fuse them to fulfill panoramic scene parsing. The innovation lies in the network adaptation to enable smooth and seamless segmentation, combined with an extended set of heterogeneous data augmentations to attain robustness in panoramic imagery. A comprehensive variety of experiments demonstrates the effectiveness for real-world surrounding perception in a single PASS, while the adaptation proposal is exceptionally positive for state-of-the-art efficient networks. Kailun Yang 0001, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, Kaiwei Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2019 | Oval Shape Constraint based Optic Disc and Cup Segmentation in Fundus Photographs
Jun Wu 0022, Kaiwei Wang, Zongjiang Shang, Jie Xu 0010, Dayong Ding, Xirong Li 0001, Gang Yang 0001 |
BMVC | 2 |
| 2019 | ACNET: Attention Based Network to Exploit Complementary Features for RGBD Semantic SegmentationabstractCompared to RGB semantic segmentation, RGBD semantic segmentation can achieve better performance by taking depth information into consideration. However, it is still problematic for contemporary segmenters to effectively exploit RGBD information since the feature distributions of RGB and depth (D) images vary significantly in different scenes. In this paper, we propose an Attention Complementary Network (ACNet) that selectively gathers features from RGB and depth branches. The main contributions lie in the Attention Complementary Module (ACM) and the architecture with three parallel branches. More precisely, ACM is a channel attention-based module that extracts weighted features from RGB and depth branches. The architecture preserves the inference of the original RGB and depth branches, and enables the fusion branch at the same time. Based on the above structures, ACNet is capable of exploiting more high-quality features from different channels. We evaluate our model on SUN-RGBD and NYUDv2 datasets, and prove that our model outperforms state-of-the-art methods. In particular, a mIoU score of 48.3% on NYUDv2 test set is achieved with ResNet50. We will release our source code based on PyTorch and the trained segmentation model at https://github.com/anheidelonghu/ACNet. Xinxin Hu, Kailun Yang 0001, Lei Fei, Kaiwei Wang |
ICIP | 4 |
| 2019 | Can we PASS beyond the Field of View? Panoramic Annular Semantic Segmentation for Real-World Surrounding PerceptionabstractPixel-wise semantic segmentation unifies distinct scene perception tasks in a coherent way, and has catalyzed notable progress in autonomous and assisted navigation, where a whole surrounding perception is vital. However, current mainstream semantic segmenters are normally benchmarked against datasets with narrow Field of View (FoV), and most vision-based navigation systems use only a forward-view camera. In this paper, we propose a Panoramic Annular Semantic Segmentation (PASS) framework to perceive the entire surrounding based on a compact panoramic annular lens system and an online panorama unfolding process. To facilitate the training of PASS models, we leverage conventional FoV imaging datasets, bypassing the effort entailed to create dense panoramic annotations. To consistently exploit the rich contextual cues in the unfolded panorama, we adapt our real-time ERF-PSPNet to predict semantically meaningful feature maps in different segments and fuse them to fulfill smooth and seamless panoramic scene parsing. Beyond the enlarged FoV,we extend focal length-related and style transfer-based data augmentations, to robustify the semantic segmenter against distortions and blurs in panoramic imagery. A comprehensive variety of experiments demonstrates the qualified robustness of our proposal for realworld surrounding understanding. Kailun Yang 0001, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, Dongming Sun, Kaiwei Wang |
IV | 7 |
| 2019 | A Coarse-to-fine Cascading Model for Cataract Nuclear Segmentation in Slit-lamp PhotographsabstractA nuclear cataract is an age-related chronic and priority ophthalmic disease in which a clouding of the lens in the human eye affects vision. Automatic segmentation of nuclear region based on slit-lamp photographs is a basic step for computer-aided diagnosis such as nuclear cataract grading. However, slit-lamp photographs collected from a clinic scenario often have complex background containing the eyelids, sclera and cornea with spectral highlights. The existing efforts using traditional image processing that have unsatisfactory results, and the deep learning method using standard Faster R-CNN tends to obtain a bigger nuclear contour. In this paper, we propose a coarse-to-fine deep learning solution to localize nuclear regions by cascading the Faster R-CNN in a two-stage framework. First, a nuclear ROI (region of interest) predictor is pre-trained to localize a rough position and remove complex backgrounds. Then, a fine nuclear locator is applied to predict a more compact nuclear bounding box. Finally, an ellipse-like nuclear contour is fitted based on its bounding box. Evaluated on a clinical dataset of 884 slit-lamp photographs, the proposed method outperforms the state-of-the-art, improving the overlapping rate (IoU) by 0.33% from 67.98% to 68.31%, and increasing the success rate by 2.55% from 85.71% to 88.26%. Jun Wu 0022, Xianfang Rong, Zhennan Zhao, Dayong Ding, Xirong Li 0001, Zongjiang Shang, Kaiwei Wang, Xixi He, Xiangjia Zhu, Wenwen He, Yinglei Zhang |
VCIP | 8 |
| 2019 | A generic parallel computational framework of lifting wavelet transform for online engineering surface filtration
Yuanping Xu, Chaolong Zhang 0002, Zhijie Xu, Jiliu Zhou, Kaiwei Wang, Jian Huang 0017 |
Signal Process. | 5 |
| 2018 | Intersection Navigation for People with Visual Impairment
Ruiqi Cheng, Kaiwei Wang, Shufei Lin |
ICCHP (2) | 2 |
| 2018 | KrNet: A Kinetic Real-Time Convolutional Neural Network for Navigational Assistance
Shufei Lin, Kaiwei Wang, Kailun Yang 0001, Ruiqi Cheng |
ICCHP (2) | 2 |
| 2018 | Visual Localization of Key Positions for Visually Impaired PeopleabstractOn the off-the-shelf navigational assistance devices, the localization precision is limited to the signal error of global navigation satellite system (GNSS). During travelling outdoors, the inaccurately localization perplexes visually impaired people, especially at key positions, such as gates, bus stations or intersections. The visual localization is a feasible approach to improving the positioning precision of assistive devices. Using multiple image descriptors, the paper proposes a robust and efficient visual localization algorithm, which takes advantage of priori GNSS signals and multi-modal images to achieve the accurate localization of key positions. In the experiments, we implement the approach on the wearable system and test the performance of visual localization under practical scenarios. Ruiqi Cheng, Kaiwei Wang, Longqing Lin, Kailun Yang 0001 |
ICPR | 2 |
| 2018 | Unifying terrain awareness through real-time semantic segmentationabstractActive research on computer vision accelerates the progress in autonomous driving. Following this trend, we aim to leverage the recently emerged methods for Intelligent Vehicles (IV), and transfer them to develop navigation assistive technologies for the Visually Impaired (VI). This topic grows notoriously challenging as it requires to detect a variety of scenes towards higher level of assistance. Computer vision based techniques with monocular detectors or depth sensors sprung up within years of research. These separate approaches achieved remarkable results with relatively low processing time, and improved the mobility of visually impaired people to a large extent. However, running all detectors jointly increases the latency and burdens the computational resources. In this paper, we put forward to seize pixel-wise semantic segmentation to cover the perception needs of navigational assistance in a unified way. This is critical not only for the terrain awareness regarding traversable areas, sidewalks, stairs and water hazards, but also for the avoidance of short-range obstacles, fast-approaching pedestrians and vehicles. At the heart of our proposal is a combination of efficient residual factorized network (ERFNet), pyramid scene parsing network (PSPNet) and 3D point cloud based segmentation. This approach proves to be with qualified accuracy and speed for real-world applications by a comprehensive set of experiments on a wearable navigation system. Kailun Yang 0001, Luis Miguel Bergasa, Eduardo Romera, Ruiqi Cheng, Tianxue Chen, Kaiwei Wang |
Intelligent Vehicles Symposium | 6 |
| 2018 | An Environmental Perception and Navigational Assistance System for Visually Impaired Persons Based on Semantic Stixels and Sound InteractionabstractAssistive technologies aim at enhancing personal mobility of individuals with disabilities to improve their independence and access to social life. For the visually impaired, perception during navigation comprises a major ingredient of independent living. With the development of computer vision, it is possible to meet the richer needs of visually impaired people. However, research on navigation assistance for the visually impaired is still relatively unexplored when compared with the active progress in autonomous driving which is already in full swing. In respond to this issue, we aim to leverage the study of the Stixel-World for automotive systems and transfer it to develop assistive technology for visually impaired people. The impressive research results of deep learning also suppose benefits for vision-based technology. Precisely, semantic segmentation is a task that enables identification of different objects uniformly. Inspired by these observations, we design a set of wearable visual aids, while the core algorithm is based on the stixel representations for three-dimensional world combined with pixel-wise semantic segmentation. Predetermined conditions for stixels in automotive research, such as camera angles, position fixes, and unsuitable assumptions made about the real world are optimized to fit the needs of navigation assistance in our algorithm, along with the incorporation of traversability-related semantic information. We also propose a sound mapping scheme, so that the environmental awareness about geographic and semantic information are conveyed to the visually impaired through acoustic feedback. Kailun Yang 0001, Weijian Hu, Kaiwei Wang |
SMC | 4 |
| 2018 | Real-time pedestrian crossing lights detection algorithm for the visually impaired
Ruiqi Cheng, Kaiwei Wang, Kailun Yang 0001, Ningbo Long |
Multim. Tools Appl. | 2 |
| 2017 | On Joint BBU/RRH Resource Allocation in Heterogeneous Cloud-RANsabstractCloud radio access network (Cloud-RAN) is a promising wireless network architecture that can satisfy the fast growing mobile data traffic and improve the performance of Internet of Things. In this paper, we propose an energy-efficient resource allocation scheme based on heterogeneous Cloud-RAN jointly considering the remote radio head (RRH) antenna resource with baseband unit (BBU) computation resource. We formulate our joint resource allocation problem and decompose it into two subproblems. The first subproblem is a network-wide beamforming vectors optimization problem, and it is solved by weighted minimum mean square error approach. Based on the optimized beamforming vector, we propose an algorithm to get the RRH-user equipment clusters. The second subproblem is a BBU scheduling problem, and we reformulate it as a bin packing problem which aims to minimize the number of BBUs in working model to save more energy. Compared to some existed works which form the BBU scheduling problem as a bin packing problem, we propose a bin packing algorithm based on the best-fit-decreasing method, which has better performance. With simulation results and detailed analysis, the system performance of our proposed joint resource allocation scheme is verified, which is more energy-efficient than other existing schemes. Kaiwei Wang, Wuyang Zhou, Shiwen Mao |
IEEE Internet Things J. | 1 |
| 2016 | Energy Efficient Joint Resource Scheduling for Delay-Aware Traffic in Cloud-RANabstractIn this paper, we focus on the energy efficient joint resource scheduling scheme in time varying Cloud-RAN with delay sensitive traffic. We jointly consider the computation resources provided by baseband units (BBUs), which are modeled as the data processing rate of each virtual machine (VM) provided by the BBUs, and the antenna resources provided by the remote radio heads (RRHs), which are modeled as the beamforming vectors for each user equipment (UE) considering the limited fronthaul capacity and per-UE QoS requirement. Based on the Lyapunov optimization method, we divide the original problem into two subproblems, i.e., the BBU processing rate scheduling problem and the network-wide beamforming strategy problem. The first subproblem can be formulated as a convex programming problem, and we can solve the second subproblem with a weighted minimum mean square error (WMMSE) approach. With the detailed theoretical analysis and simulation results, it is clear that we can achieve a trade- off between energy efficiency and traffic delay, which can be controlled by the control parameter V. Kaiwei Wang, Wuyang Zhou, Shiwen Mao |
GLOBECOM | 1 |
| 2014 | Traffic-aware graph-based dynamic frequency reuse for heterogeneous Cloud-RANabstractIn order to allocate frequency resource among users in heterogeneous Cloud-RAN (Het-Cloud-RAN) as well as avoid the co-channel interference, we propose a traffic-aware graph-based dynamic frequency reuse scheme. Cloud-RAN is a new network structure comprised of baseband units (BBU) and remote radio heads (RRH), which has the ability to reduce the energy costs of network. Two kinds of RRHs, including macro RRHs and pico RRHs, are considered to form the two tiers of the Het-Cloud-RAN. We construct the interference graph where the vertices represent different RRHs in the network, and then apply our proposed scheme in it. We use a graph coloring method to allocate different length of bandwidth to cells in different tiers based on their traffic demands. Our scheme also considers the time varying traffic and can adjust the frequency allocation based on the change of traffic demands in each cell without re-performing the whole scheme. With the application of our scheme in Het-Cloud-RAN, we can not only improve the spectrum utilization and efficiency, but also reduce the energy consumption through reducing the number of working BBUs. The system level simulation results show clearly that our new scheme in Het-Cloud-RAN outperforms the existing frequency allocation scheme based on traditional cellular system in many ways including throughput and energy efficiency. Kaiwei Wang, Ming Zhao 0008, Wuyang Zhou |
GLOBECOM | 1 |
| 2014 | Graph-based dynamic frequency reuse in Cloud-RANabstractThis work aims to apply a new dynamic frequency reuse scheme based on fractional frequency reuse (FFR) in Cloud-RAN structure. FFR is one of the most efficient schemes to mitigate the inter-cell interference in OFDM system, while strict FFR in traditional cellular system is still facing some problems such as the inefficient use of spectrum. Cloud-RAN is a new network structure comprised of baseband units (BBU) and remote radio heads (RRH), which has the ability to reduce the energy costs of network. We use a graph coloring method to allocate spectrum resource among different cell zones depending on different traffic demands as well as avoiding the inter-cell interference. With the application of our proposed scheme based on Cloud-RAN structure, we can not only improve the spectrum utilization, but also reduce the energy consumption through reducing the working number of BBUs. The simulation results show clearly that our new scheme in Cloud-RAN structure can achieve higher throughput and energy efficiency compared with the strict FFR in traditional cellular system. Kaiwei Wang, Ming Zhao 0008, Wuyang Zhou |
WCNC | 1 |
| 2012 | Queuing method in combined channel aggregation and fragmentation strategy for dynamic spectrum accessabstractTo achieve highly efficient and flexible spectrum sharing in cognitive radio networks, we proposed a queuing method in combined channel aggregation and fragmentation (CAF) strategy for dynamic spectrum access (DSA). We derived the balance equation and evaluated various system performance metrics including blocking probability, dropping probability, spectrum utilization, throughput and mean queuing time by a continuous time Markov chain (CTMC) model. Numerical results show that this CAF with queuing strategy greatly lowers the dropping probability without obviously impacting the blocking probability, and slightly increases the spectrum utilization and throughput of the secondary network compared with the non-queuing case. Moreover, by tuning the bandwidth requirement of each secondary user, different system performance can be achieved. Sihai Zhang, Kaiwei Wang, Wuyang Zhou |
PIMRC | 3 |