Siqi Li 0001

dblp:34/180-1 · also Si-Qi Li 0001 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
25since 2021 · last 2026
0000-0001-9720-826XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 16 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 13 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HIRNet: Hypergraph-Induced Iterative Reasoning Network for Crowd Counting
abstract
In recent years, point regression-based crowd counting has demonstrated remarkable advantages for accurate counting in highly congested scenes. However, existing single-pass regression frameworks lack an explicit self-correction mechanism, making it difficult to rectify localization shifts or even false positives induced by local visual ambiguity in heavily occluded regions. Moreover, straightforward regression predictions struggle to capture complex cluster-level higher-order relations in dense crowds, limiting the model’s ability to leverage global context to resolve semantic confusion. To address these issues, we reformulate point regression as an iterative optimization process and propose a Hypergraph-induced Iterative Reasoning Network, HIRNet. Specifically, we first design a 2D bidirectional hypergraph iterative reasoning module (HIR-2D), which treats intermediate-round regression points as an explicit structural prior and dynamically constructs bidirectional hypergraphs along the spatial and channel dimensions. By performing higher-order feature aggregation in both dimensions, HIR-2D enforces long-range spatial consistency and recalibrates channel features, thereby effectively suppressing background distractions. Second, to mitigate the matching jitter of regressed point sets across iterations, we introduce an Anchor-Matching Iterative Deep Supervision strategy that locks the bipartite matching from the final round to ensure consistent gradient directions across rounds. Furthermore, we devise a Monotonic Improvement Loss that explicitly constrains the model to progressively reduce geometric errors throughout iterations, stabilizing the reasoning dynamics. Extensive experiments on three mainstream datasets demonstrate that our method substantially outperforms current advanced approaches.
Mengqi Lei, Siqi Li 0001
ICMR3
2026 SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
Mengqi Lei, Siqi Li 0001, Xinhu Zheng, Shaoyi Du, Yue Gao 0002
Int. J. Comput. Vis.3
2026 Event-based facial expression recognition via large vision-language models
Siqi Li 0001, Yongji Zhang, Yue Gao 0002
Pattern Recognit.1
2026 H3Former: Hypergraph-Based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
abstract
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on feature-selection mechanisms or region-proposal strategies to localize discriminative regions for semantic analysis. However, these methods often fail to capture discriminative cues comprehensively while introducing substantial category-agnostic redundancy. To address these limitations, we propose $\text {H}^{3}$ Former, a novel token-to-region framework that leverages high-order semantic relations to aggregate local fine-grained representations with structured region-level modeling. Specifically, we propose the Semantic-Aware Aggregation Module (SAAM), which exploits multi-scale contextual cues to dynamically construct a weighted hypergraph among tokens. By applying hypergraph convolution, SAAM captures high-order semantic dependencies and progressively aggregates token features into compact region-level representations. Furthermore, we introduce the Hyperbolic Hierarchical Contrastive Loss (HHCL), which enforces hierarchical semantic constraints in a non-Euclidean embedding space. The HHCL enhances inter-class separability and intra-class consistency while preserving the intrinsic hierarchical relationships among fine-grained categories. Comprehensive experiments conducted on four standard FGVC benchmarks validate the superiority of our $\text {H}^{3}$ Former framework. Code is available at https://github.com/xiaozhangfangyang/H3Former.
Yongji Zhang, Siqi Li 0001, Kuiyang Huang, Yue Gao 0002, Yu Jiang 0006
IEEE Trans. Image Process.2
2026 3D Semantic Gaussian via Geometric-Semantic Hypergraph Computation
abstract
Semantic labels are inherently tied to geometry and luminance reconstruction, as entities with similar shapes and appearances often share categories. Traditional methods use synthesis-analysis, NeRF, or 3D Gaussian representations to encode semantics and geometry separately. However, 2D methods lack view consistency, NeRF extensions are slow, and faster 3D Gaussian methods risk spatial and channel inconsistencies between semantic and RGB. Moreover, these methods require costly manual dense semantic labels. To alleviate resource demands and achieve effective semantic reconstruction with sparse inputs while enhancing RGB rendering quality, we build upon 3D Gaussian by integrating semantic features from pre-trained models-requiring no additional ground truth input-into Gaussian features, and construct a hypergraph neural network to capture higher-order correlations across RGB and semantic information as well as between different frames. Hypergraphs use hyperedges to link multiple vertices, capturing complex relationships essential for cross-modal tasks. This higher-order structure addresses the limitations of NeRF and Gaussian methods, which lack the capacity for such advanced associations. This framework enables precise novel view synthesis and 2D semantic reconstruction without manual annotations, achieving state-of-the-art results for RGB and semantic tasks on room-scale scenes in the ScanNet and Replica datasets, while supporting real-time rendering speeds of 34 FPS.
Dejian Guo, Siqi Li 0001, Shaoyi Du, Xiangmin Han, Yue Gao 0002
IEEE Trans. Multim.4
2026 SkiTrack: An Aerial Skiing Benchmark for Human-Centric Object Tracking
abstract
Aerial skiing is a challenging human-centric sport characterized by rapid motion, large-scale variations, and frequent occlusions. Its extensive spatial range is typically captured by cameras or drones from multiple perspectives, resulting in frequent and complex viewpoint shifts. These challenges encompass nearly all difficulties inherent in human-centric tracking tasks. In this article, we introduce SkiTrack , the first dataset explicitly designed for tracking in aerial skiing. SkiTrack enhances the performance of existing tracking algorithms across a range of human-centric scenarios by providing precise annotations. We observe distinct characteristics in the tracked components, with the skis being rigid and low in visibility and the athlete’s body highly deformable but more visible. To leverage these differences, we propose a components decoupled loss that applies separate constraints to the tracking of the athlete and skis, thereby improving tracking accuracy in skiing scenes. Our experimental results validate the effectiveness of both the SkiTrack dataset and the proposed decoupled loss function, demonstrating consistent improvements in the performance of established models on human-centric tracking tasks. Data are available at https://github.com/xiaozhangfangyang/FineSkiing .
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.3
2026 GLU-Net: Global-Local Fusion Network for Event-Based Monocular Depth Estimation via Uncertainty Optimization
abstract
Event-based monocular depth estimation is crucial for applications such as autonomous driving, obstacle avoidance, and navigation under high-speed scenarios. Events exhibit a unique and irregular modality. To adapt them to neural networks, some studies convert event streams into event voxels or other frame-like representations. However, these approaches tend to lose the temporal characteristics of events. In this study, we propose a network that aggregates global voxel and per-channel temporal local features of event voxels across the temporal dimension, explicitly extracting events’ temporal information. Furthermore, as noise in events can interfere with the training process and is more difficult to predict than that in images, we utilize the uncertainty estimation module to mitigate the impact of uncertain factors and enhance the robustness of the model. Additionally, we employ multi-level depth features for supervisory training, which improves prediction performance compared to methods relying solely on ground-truth depth supervision. Experiments on open source datasets demonstrate the effectiveness of the proposed method. Our code can be found at https://github.com/WuShangjie/GLUNET .
Shangjie Wu, Jihua Zhu, Zhikuan Zhou, Siqi Li 0001, Shaoyi Du, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2025 GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph Learning
abstract
Although the fusion of images and LiDAR point clouds is crucial to many applications in computer vision, the relative poses of cameras and LiDAR scanners are often unknown. However, due to the modality and domain gap between images and LiDAR point clouds, Image-to-Point Cloud Registration is a significant challenge, especially when the image and point cloud come from non-synchronized frames. To tackle these issues, we introduce the virtual point cloud as a bridge to alleviate the cross-modality gap between images and LiDAR point clouds. In this way, the modality gap is converted to the domain gap of point clouds. Moreover, we introduce a virtual-spherical representation achieving orthogonal decoupling between pixel location and predicted depth. As for the domain gap, we propose a distribution-based adaptive sample module to generate a unified distribution of two types of point clouds. Then, we explore the correct correspondence pattern consistency and prune the false correspondences through a graph-based selection process. Experimental results demonstrate that our method outperforms the state-of-the-art methods by more than 10.77% and 12.53% performance on the KITTI Odometry and nuScenes datasets, respectively. The results demonstrate that our method can effectively solve non-synchronized random-frame registration.
Lin Bie, Shouan Pan, Siqi Li 0001, Yue Gao 0002
CVPR3
2025 Hyper-Depth: Hypergraph-Based Multi-Scale Representation Fusion for Monocular Depth Estimation
Lin Bie, Siqi Li 0001, Yifan Feng 0001, Yue Gao 0002
ICCV2
2025 ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement
abstract
Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respond to pixel-wise brightness changes asynchronously. Event cameras' high dynamic range is pivotal for visual perception in extreme low-light scenarios, surpassing traditional cameras and enabling applications in challenging dark environments. In this paper, inspired by the success of the retinex theory for traditional frame-based low-light image restoration, we introduce the first methods that combine the retinex theory with event cameras and propose a novel retinex-based lowlight image restoration framework named ERetinex. Among our contributions, the first is developing a new approach that leverages the high temporal resolution data from event cameras with traditional image information to estimate scene illumination accurately. This method outperforms traditional image-only techniques, especially in low-light environments, by providing more precise lighting information. Additionally, we propose an effective fusion strategy that combines the high dynamic range data from event cameras with the color information of traditional images to enhance image quality. Through this fusion, we can generate clearer and more detailrich images, maintaining the integrity of visual information even under extreme lighting conditions. The experimental results indicate that our proposed method outperforms state-of-theart (SOTA) methods, achieving a gain of 1.0613 dB in PSNR while reducing FLOPS by 84.28 %. The code is available at https://github.com/lodew920/ERetinex.
Xuejian Guo, Yuehang Wang, Siqi Li 0001, Yu Jiang 0006, Shaoyi Du, Yue Gao 0002
ICRA4
2025 Event-enhanced synthetic aperture imaging
Siqi Li 0001, Shaoyi Du, Jun-Hai Yong, Yue Gao 0002
Sci. China Inf. Sci.1
2025 RGB-D Visual Perception for Occluded Scenes via Event Camera
Siqi Li 0001, Zongze Wu 0001, Zhou Xue, Yu-Shen Liu, Yue Gao 0002
Int. J. Comput. Vis.1
2025 Image Matting and 3D Reconstruction in One Loop
Xinshuang Liu, Siqi Li 0001, Yue Gao 0002
Int. J. Comput. Vis.2
2025 A Real-World Animation Super-Resolution Benchmark With Color Degradation and Multi-Scale Multi-Frequency Alignment
abstract
Animation super-resolution (SR) aims to generate high-resolution (HR) animation frames from degraded low-resolution (LR) inputs, constituting an important task in real-world SR. Existing animation SR methods typically follow a photorealistic real-world SR computational paradigm. However, digital animation frames commonly suffer from compression and transmission-related degradation, distinct from degradations in camera-captured real-world images. In this paper, we introduce a novel real-world animation super-resolution benchmark designed explicitly for animation frames, named ADASR, featuring both 2D and modern 3D animation content to facilitate industry applications. Additionally, we propose a Color-Aware Animation Super-Resolution (CAASR) method. CAASR, for the first time, incorporates a color degradation simulation mechanism tailored for animations, addressing color banding, blocking, and color shift. Furthermore, we develop a multi-scale multi-frequency alignment mechanism to robustly extract degradation-invariant features. Extensive experiments conducted on both the existing AVC dataset and our newly constructed ADASR dataset demonstrate that our proposed CAASR achieves state-of-the-art performance in restoring HR frames for both 2D and 3D animations. Code and data are available at https://github.com/huangyang-666/CAASR.
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yutong Yao, Yue Gao 0002
IEEE Trans. Image Process.3
2025 EvCSLR: Event-Guided Continuous Sign Language Recognition and Benchmark
abstract
Classical continuous sign language recognition (CSLR) suffers from some main challenges in real-world scenarios: accurate inter-frame movement trajectories may fail to be captured by traditional RGB cameras due to the motion blur, and valid information may be insufficient under low-illumination scenarios. In this paper, we for the first time leverage an event camera to overcome the above-mentioned challenges. Event cameras are bio-inspired vision sensors that could efficiently record high-speed sign language movements under low-illumination scenarios and capture human information while eliminating redundant background interference. To fully exploit the benefits of the event camera for CSLR, we propose a novel event-guided multi-modal CSLR framework, which could achieve significant performance under complex scenarios. Specifically, a time redundancy correction (TRCorr) module is proposed to rectify redundant information in the temporal sequences, directing the model to focus on distinctive features. A multi-modal cross-attention interaction (MCAI) module is proposed to facilitate information fusion between events and frame domains. Furthermore, we construct the first event-based CSLR dataset, namedEvCSLR, which will be released as the first event-based CSLR benchmark. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on EvCSLR and PHOENIX-2014 T datasets.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Qianren Guo, Qi Chu 0010, Yue Gao 0002
IEEE Trans. Multim.3
2025 Multi-space Representation Fusion Enhanced Monocular Depth Estimation via Virtual Point Cloud
abstract
Monocular Depth Estimation (MDE) is a fundamental problem in computer vision with broad applications in various downstream tasks. While recent studies focus on designing increasingly complex and powerful deep learning methods to regress depth maps directly, we propose a novel approach by introducing the Virtual Point Cloud (VPC) as an intermediate representation to provide the approximate geometric prior for the MDE task. In this article, we design a multi-scale multi-space representation fusion-enhanced MDE framework to address the challenges of MDE. Specifically, to resolve the issue of scale ambiguity, we design a VPC feature extraction module to learn multi-scale 3D geometric information for the depth prior. Then, we explicitly introduce geometric constraints for global depth prediction by incorporating a multi-space representation fusion from both the texture features in 2D space and the geometric features in 3D space. To mitigate errors at object boundaries, we introduce a confidence map generated based on the quality of the VPC to refine the predicted depth map. Specifically, we construct convolution receptive fields based on 3D spatial distances in spherical coordinates, ensuring that the confidence map provides reliable geometric guidance at object boundaries. Furthermore, we propose an independent confidence geometric consistency loss to supervise the refinement process. Experimental results demonstrate that our method significantly outperforms state-of-the-art approaches across all evaluation metrics on the KITTI and NYU-Depth-v2 datasets, achieving RMSE improvements of 9.2% and 2.8%, respectively. Moreover, zero-shot evaluations on the nuScenes and SUN-RGBD datasets further validate the generalizability of our approach.
Lin Bie, Siqi Li 0001, Xiaopin Zhong, Zongze Wu 0001, Yue Gao 0012
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Arbitrary Large-Scale Scene Reconstruction without Annotated Block Partitions
abstract
Large-scale scene reconstruction is a challenging problem. As different parts of the scene could be visible from different collected image frames, previous works manually use distance or geography to decompose the scene into parts and reconstruct each part of the scene separately. However, such manual decomposition is a laborious and time-consuming task when applied to large-scale scene reconstruction in real-world applications. To address this, we propose VisibleNeRF automatically reconstructs large-scale scenes by decomposing scenes into parts based on the part visibility. More specifically, we propose a visibility judgment strategy to decompose the scenes into visible and invisible parts. Then we reconstruct the visible part with the corresponding collected images and continue to decompose the rest of the invisible parts with the proposed visibility judgment strategy. New NeRF modules are re-established for the decomposed invisible parts until the entire scene is reconstructed. To the best of our knowledge, we are the first to propose an online reconstruction of large-scale scenes without manual decomposition. Experimental results on three datasets show that our method successfully reconstructs large-scale scenes in a fully automatic manner. Besides, in the widely used Mission Bay dataset, our model outperforms other state-of-the-art methods by a large margin.
Lin Bie, Siqi Li 0001, Dejian Guo, Shaoyi Du, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2024 3D Feature Tracking via Event Camera
abstract
This paper presents the first 3D feature tracking method with the corresponding dataset. Our proposed method takes event streams from stereo event cameras as input to pre-dict 3D trajectories of the target features with high-speed motion. To achieve this, our method leverages a joint framework to predict the 2D feature motion offsets and the 3D feature spatial position simultaneously. A motion compensation module is leveraged to overcome the feature deformation. A patch matching module based on bi-polarity hypergraph modeling is proposed to robustly es-timate the feature spatial position. Meanwhile, we collect the first 3D feature tracking dataset with high-speed moving objects and ground truth 3D feature trajectories at 250 FPS, named E-3DTrack, which can be used as the first high-speed 3D feature tracking benchmark. Our code and dataset could be found at: https://github.com/lisiqi19971013/E-3DTrack.
Siqi Li 0001, Zhikuan Zhou, Zhou Xue, Shaoyi Du, Yue Gao 0002
CVPR1
2024 Image-to-Point Registration via Cross-Modality Correspondence Retrieval
abstract
Image-to-Point Cloud registration between 2D images and 3D LiDAR point clouds is a significant task in computer vision. The traditional registration pipeline first establishes correspondences between images and point clouds and then performs pose estimation based on the generated matches. However, 2D-3D correspondences are inherently difficult to be established due to the large modality gap between images and LiDAR point clouds. To this end, we build a bridge to alleviate the 2D-3D modality gap, which aligns LiDAR point clouds to the virtual points generated by images. In this way, the modality gap can be alleviated to the domain gap of different types of point clouds, i.e. original point clouds and virtual point clouds. Concretely, our framework conducts feature fusion from the LiDAR and virtual point cloud by utilizing the Transformer layer. To relieve the domain gap, a frustum points retrieval module and a combined correspondences retrieval module are proposed based on the consistency of the feature and position descriptor to select the correct correspondences among the candidates, which are generated from the simultaneous retrieval of features and position descriptors. In the implementation procedure, we design a frustum retrieval loss and a combined correspondence retrieval loss for cross-modality correspondence retrieval. Experimental results and comparison with state-of-the-art Image-to-Point Cloud methods on KITTI and nuScenes datasets demonstrate our proposed method has achieved superior performance.
Lin Bie, Siqi Li 0001
ICMR2
2024 Hypergraph-Based Multi-View Action Recognition Using Event Cameras
abstract
Action recognition from video data forms a cornerstone with wide-ranging applications. Single-view action recognition faces limitations due to its reliance on a single viewpoint. In contrast, multi-view approaches capture complementary information from various viewpoints for improved accuracy. Recently, event cameras have emerged as innovative bio-inspired sensors, leading to advancements in event-based action recognition. However, existing works predominantly focus on single-view scenarios, leaving a gap in multi-view event data exploitation, particularly in challenges like information deficit and semantic misalignment. To bridge this gap, we introduceHyperMV, multi-view event-based action recognition framework. HyperMV converts discrete event data into frame-like representations and extracts view-related features using a shared convolutional network. By treating segments as vertices and constructing hyperedges using rule-based and KNN-based strategies, a multi-view hypergraph neural network that captures relationships across viewpoint and temporal features is established. The vertex attention hypergraph propagation is also introduced for enhanced feature fusion. To prompt research in this area, we present the largest multi-view event-based action dataset$\mathbf{THU}^{\mathbf{MV-EACT}}\mathbf{-50}$, comprising 50 actions from 6 viewpoints, which surpasses existing datasets by over tenfold. Experimental results show that HyperMV significantly outperforms baselines in both cross-subject and cross-view scenarios, and also exceeds the state-of-the-arts in frame-based multi-view action recognition.
Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Shaoyi Du
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Event-Based Low-Illumination Image Enhancement
abstract
Event cameras are bio-inspired vision sensors with a high dynamic range (140 dB for event camerasvs.60 dB for traditional cameras) and can be used to tackle the image degradation problem under extremely low-illumination scenarios, which is still not well-explored yet. In this article, we propose a joint framework to compose the underexposed frames and event streams captured by the event camera to reconstruct clear images with detailed textures under almost dark conditions. A residual fusion module is proposed to reduce the domain gap between event streams and frames by using the residuals of both modalities. A multi-level reconstruction loss based on the variability of the contrast distribution is proposed to reduce the perceptual errors of the output image. In addition, we construct the first real-world low-illumination image enhancement dataset (mainly under 2 lux illumination scenes), named LIE, containing event streams and frames collected under indoor and outdoor low-light scenarios together with the ground truth clear images. Experimental results on our LIE dataset demonstrate that our proposed method could achieve significant improvements compared with existing methods.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Minghao Zhao 0003, Yue Gao 0002
IEEE Trans. Multim.3
2023 SuperFast: 200× Video Frame Interpolation via Event Camera
abstract
Traditional frame-based video frame interpolation (VFI) methods rely on the linear motion assumption and brightness invariance assumption, which may lead to fatal errors confronting the scenarios with high-speed motions. To tackle the above challenge, inspired by the advantages of event cameras on asynchronously recording brightness changes at each pixel, we propose a Fast-Slow joint synthesis framework for event-enhanced high-speed video frame interpolation, named SuperFast, in this paper, which can generate high frame rate (5000 FPS, 200× faster) video from the input low frame rate (25 FPS) video and the corresponding event stream. In our framework, the task is divided into two sub-tasks, i.e., video frame interpolation for the contents with and without high-speed motions, which are tackled by two corresponding branches, i.e., the fast synthesis pathway and the slow synthesis pathway. The fast synthesis pathway leverages a spiking neural network to encode the input event stream, and combines boundary frames to generate intermediate results through synthesis and refinement, targeting on contents with high-speed motions. The slow synthesis pathway stacks the two input boundary frames and the event stream to synthesize intermediate results, focusing on relatively slow-motion contents. Finally, a fusion module with a comparison loss is utilized to generate the final video frame interpolation results. We also build a hybrid visual acquisition system containing an event camera and a high frame rate camera, and collect the first 5000 FPS High-Speed Event-enhanced Video frame Interpolation (THU[Formula: see text]) dataset. To evaluate the performance of our proposed framework, we have conducted experiments on our THU[Formula: see text] dataset and the existing HS-ERGB dataset. Experimental results demonstrate that our proposed framework can achieve state-of-the-art 200× video frame interpolation performance under high-speed motion scenarios.
Yue Gao 0002, Siqi Li 0001, Yandong Guo, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Action Recognition and Benchmark Using Event Cameras
abstract
Recent years have witnessed remarkable achievements in video-based action recognition. Apart from traditional frame-based cameras, event cameras are bio-inspired vision sensors that only record pixel-wise brightness changes rather than the brightness value. However, little effort has been made in event-based action recognition, and large-scale public datasets are also nearly unavailable. In this paper, we propose an event-based action recognition framework calledEV-ACT. The Learnable Multi-Fused Representation (LMFR) is first proposed to integrate multiple event information in a learnable manner. The LMFR with dual temporal granularity is fed into the event-based slow-fast network for the fusion of appearance and motion features. A spatial-temporal attention mechanism is introduced to further enhance the learning capability of action recognition. To prompt research in this direction, we have collected the largest event-based action recognition benchmark namedTHUE-ACT-50and the accompanyingTHUE-ACT-50-CHLdataset under challenging environments, including a total of over 12,830 recordings from 50 action categories, which is over 4 times the size of the previous largest dataset. Experimental results show that our proposed framework could achieve improvements of over 14.5%, 7.6%, 11.2%, and 7.4% compared to previous works on four benchmarks. We have also deployed our proposed EV-ACT framework on a mobile platform to validate its practicality and efficiency.
Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Nan Ma 0012, Shaoyi Du, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 View-Guided Point Cloud Completion
abstract
This paper presents a view-guided solution for the task of point cloud completion. Unlike most existing methods directly inferring the missing points using shape priors, we address this task by introducing ViPC (view-guided point cloud completion) that takes the missing crucial global structure information from an extra single-view image. By leveraging a framework that sequentially performs effective cross-modality and cross-level fusions, our method achieves significantly superior results over typical existing solutions on a new large-scale dataset we collect for the view-guided point cloud completion task.
Xuancheng Zhang, Yutong Feng, Siqi Li 0001, Changqing Zou, Hai Wan, Xibin Zhao, Yandong Guo, Yue Gao 0002
CVPR3
2021 Event Stream Super-Resolution via Spatiotemporal Constraint Learning
abstract
Event cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consumption. However, the spatial resolution of existing event cameras is insufficient and challenging to be enhanced at the hardware level while maintaining the asynchronous philosophy of circuit design. Therefore, it is imperative to explore the algorithm of event stream super-resolution, which is a non-trivial task due to the sparsity and strong spatio-temporal correlation of the events from an event camera. In this paper, we propose an end-to-end framework based on spiking neural network for event stream super-resolution, which can generate high-resolution (HR) event stream from the input low-resolution (LR) event stream. A spatiotemporal constraint learning mechanism is proposed to learn the spatial and temporal distributions of the event stream simultaneously. We validate our method on four large-scale datasets and the results show that our method achieves state-of-the-art performance. The satisfying results on two downstream applications, i.e. object classification and image reconstruction, further demonstrate the usability of our method. To prove the application potential of our method, we deploy it on a mobile platform. The high-quality HR event stream generated by our real-time system demonstrates the effectiveness and efficiency of our method.
Siqi Li 0001, Yutong Feng, Yu Jiang 0006, Changqing Zou, Yue Gao 0002
ICCV1
2020 Attention-Based Multi-Modal Fusion Network for Semantic Scene Completion
abstract
This paper presents an end-to-end 3D convolutional network named attention-based multi-modal fusion network (AMFNet) for the semantic scene completion (SSC) task of inferring the occupancy and semantic labels of a volumetric 3D scene from single-view RGB-D images. Compared with previous methods which use only the semantic features extracted from RGB-D images, the proposed AMFNet learns to perform effective 3D scene completion and semantic segmentation simultaneously via leveraging the experience of inferring 2D semantic segmentation from RGB-D images as well as the reliable depth cues in spatial dimension. It is achieved by employing a multi-modal fusion architecture boosted from 2D semantic segmentation and a 3D semantic completion network empowered by residual attention blocks. We validate our method on both the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset and the results show that our method respectively achieves the gains of 2.5% and 2.6% on the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset against the state-of-the-art method.
Siqi Li 0001, Changqing Zou, Xibin Zhao, Yue Gao 0002
AAAI1