VLDB 2026 Research / reviewers in the wild / expert
Shijie Li 0006
dblp:141/7586-6
· DBLP profile ↗
18ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0001-6288-6984ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Velocity Space Representation Learning for GPR Keypoint Detection and MatchingabstractReliable localization under Global Positioning System-denied or visually degraded conditions remains a fundamental challenge for autonomous systems. Vision- and Light Detection and Ranging (LiDAR)-based approaches often degrade in low illumination, adverse weather, or appearance-changing environments, as they rely on stable surface texture or geometry. In contrast, ground-penetrating radar (GPR) captures subsurface electromagnetic reflections that remain relatively stable across lighting, seasonal, and weather variations, making it a promising complementary sensing modality for long-term localization. However, spatial variability in subsurface dielectric properties induces fluctuations in electromagnetic wave velocity, leading to geometric distortions in GPR echoes and unstable feature extraction. To address this challenge, we propose the Velocity-Invariant Feature Transform (VIFT), a physics-guided self-supervised learning framework for GPR keypoint detection and description. VIFT explicitly models wave-velocity-induced distortions through a continuous velocity space parameterized by a Beta distribution, and leverages velocity-conditioned wavefield migration as physically consistent data augmentation. A Siamese network is trained with velocity-consistency supervision to jointly learn repeatable keypoint score maps and discriminative local descriptors from unlabeled real GPR scans. To further enhance robustness, sparsity-aware, dispersion, distinctiveness, and orthogonality losses are incorporated to improve repeatability, spatial coverage, and descriptor discriminability. Extensive experiments on public benchmarks and large-scale real-world GPR datasets demonstrate that VIFT consistently outperforms traditional handcrafted methods and recent learning-based Vison and GPR methods, achieving a 5–10% improvement in keypoint repeatability over state-of-the-art methods, particularly under extremely sparse keypoint sampling regimes, while also improving matching accuracy and registration robustness under diverse subsurface conditions. Xieyuanli Chen, Liang Shen 0003, Xulei Yang, Bharadwaj Veeravalli, Shijie Li 0006, Tian Jin 0001, Xiaotao Huang 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Groundingabstract3D Visual Grounding (3DVG) aims to locate objects in 3D scenes based on textual descriptions, essential for applications like augmented reality and robotics. Traditional 3DVG approaches rely on annotated 3D datasets and predefined object categories, limiting scalability and adaptability. To overcome these limitations, we introduce SeeGround, a zero-shot 3DVG framework leveraging 2D Vision-Language Models (VLMs) trained on large-scale 2D data. SeeGround represents 3D scenes as a hybrid of query-aligned rendered images and spatially enriched text descriptions, bridging the gap between 3D data and 2D-VLMs input formats. We propose two modules: the Perspective Adaptation Module, which dynamically selects viewpoints for query-relevant image rendering, and the Fusion Alignment Module, which integrates 2D images with 3D spatial descriptions to enhance object localization. Extensive experiments on ScanRefer and Nr3D demonstrate that our approach outperforms existing zero-shot methods by large margins. Notably, we exceed weakly supervised methods and rival some fully supervised ones, outperforming previous SOTA by 7.7% on ScanRefer and 7.1% on Nr3D, showcasing its effectiveness in complex 3DVG tasks. Project website (with demo and code): https://seeground.github.io. Shijie Li 0006, Lingdong Kong, Xulei Yang, Junwei Liang 0001 |
CVPR | 2 |
| 2025 | Global-Aware Monocular Semantic Scene Completion with State Space Models
Shijie Li 0006, Zhongyao Cheng, Juergen Gall, Xun Xu 0002, Xulei Yang |
ICCV | 1 |
| 2025 | Future-Aware Interaction Network for Motion Forecasting
Shijie Li 0006, Xun Xu 0002, Si Yong Yeo, Xulei Yang |
ICCV | 1 |
| 2025 | Valid: Variable-Length Input Diffusion for Novel View SynthesisabstractNovel View Synthesis (NVS), which tries to produce a realistic image at the target view given source view images and their corresponding poses, is a fundamental problem in 3D Vision. As this task is heavily under-constrained, some recent work, like Zerol23 [18], tries to solve this problem with generative modeling, specifically using pre-trained diffusion models. Although this strategy generalizes well to new scenes, compared to neural radiance field-based methods, it offers low levels of flexibility. For example, it can only accept a single-view image as input, despite realistic applications often offering multiple input images. This is because the source-view images and corresponding poses are processed separately and injected into the model at different stages. Thus it is not trivial to generalize the model into multi-view source images, once they are available. To solve this issue, we try to process each pose image pair separately and then fuse them as a unified visual representation which will be injected into the model to guide image synthesis at the target-views. However, inconsistency and computation costs increase as the number of input source-view images increases. To solve these issues, the Multi-view Cross Former module is proposed which maps variable-length input data to fix-size output data. A two-stage training strategy is introduced to further improve the efficiency during training time. Qualitative and quantitative evaluation over multiple datasets demonstrates the effectiveness of the proposed method against previous approaches. The code will be released according to the acceptance. Shijie Li 0006, Farhad G. Zanjani, Haitam Ben Yahia, Yuki Markus Asano, Juergen Gall, AmirHossein Habibian |
WACV | 1 |
| 2025 | SiFH: Siamese frequency harmonization self-supervised learning for motion forecasting
Chunyu Liu 0004, Tiechui Yao, Shijie Li 0006 |
Neurocomputing | 4 |
| 2025 | Rethinking 3-D LiDAR Point Cloud SegmentationabstractMany point-based semantic segmentation methods have been designed for indoor scenarios, but they struggle if they are applied to point clouds that are captured by a light detection and ranging (LiDAR) sensor in an outdoor environment. In order to make these methods more efficient and robust such that they can handle LiDAR data, we introduce the general concept of reformulating 3-D point-based operations such that they can operate in the projection space. While we show by means of three point-based methods that the reformulated versions are between 300 and 400 times faster and achieve higher accuracy, we furthermore demonstrate that the concept of reformulating 3-D point-based operations allows to design new architectures that unify the benefits of point-based and image-based methods. As an example, we introduce a network that integrates reformulated 3-D point-based operations into a 2-D encoder-decoder architecture that fuses the information from different 2-D scales. We evaluate the approach on four challenging datasets for semantic LiDAR point cloud segmentation and show that leveraging reformulated 3-D point-based operations with 2-D image-based operations achieves very good results for all four datasets. Shijie Li 0006, Yun Liu 0011, Juergen Gall |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | MiniSeg: An Extremely Minimum Network Based on Lightweight Multiscale Learning for Efficient COVID-19 SegmentationabstractThe rapid spread of the new pandemic, i.e., coronavirus disease 2019 (COVID-19), has severely threatened global health. Deep-learning-based computer-aided screening, e.g., COVID-19 infected area segmentation from computed tomography (CT) image, has attracted much attention by serving as an adjunct to increase the accuracy of COVID-19 screening and clinical diagnosis. Although lesion segmentation is a hot topic, traditional deep learning methods are usually data-hungry with millions of parameters, easy to overfit under limited available COVID-19 training data. On the other hand, fast training/testing and low computational cost are also necessary for quick deployment and development of COVID-19 screening systems, but traditional methods are usually computationally intensive. To address the above two problems, we propose MiniSeg, a lightweight model for efficient COVID-19 segmentation from CT images. Our efforts start with the design of an attentive hierarchical spatial pyramid (AHSP) module for lightweight, efficient, effective multiscale learning that is essential for image segmentation. Then, we build a two-path (TP) encoder for deep feature extraction, where one path uses AHSP modules for learning multiscale contextual features and the other is a shallow convolutional path for capturing fine details. The two paths interact with each other for learning effective representations. Based on the extracted features, a simple decoder is added for COVID-19 segmentation. For comparing MiniSeg to previous methods, we build a comprehensive COVID-19 segmentation benchmark. Extensive experiments demonstrate that the proposed MiniSeg achieves better accuracy because its only 83k parameters make it less prone to overfitting. Its high efficiency also makes it easy to deploy and develop. The code has been released at https://github.com/yun-liu/MiniSeg. Yun Liu 0011, Shijie Li 0006, Jing Xu 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | MS-TCN++: Multi-Stage Temporal Convolutional Network for Action SegmentationabstractWith the success of deep learning in classifying short trimmed videos, more attention has been focused on temporally segmenting and classifying activities in long untrimmed videos. State-of-the-art approaches for action segmentation utilize several layers of temporal convolution and temporal pooling. Despite the capabilities of these approaches in capturing temporal dependencies, their predictions suffer from over-segmentation errors. In this paper, we propose a multi-stage architecture for the temporal action segmentation task that overcomes the limitations of the previous approaches. The first stage generates an initial prediction that is refined by the next ones. In each stage we stack several layers of dilated temporal convolutions covering a large receptive field with few parameters. While this architecture already performs well, lower layers still suffer from a small receptive field. To address this limitation, we propose a dual dilated layer that combines both large and small receptive fields. We further decouple the design of the first stage from the refining stages to address the different requirements of these stages. Extensive evaluation shows the effectiveness of the proposed model in capturing long-range dependencies and recognizing action segments. Our models achieve state-of-the-art results on three datasets: 50Salads, Georgia Tech Egocentric Activities (GTEA), and the Breakfast dataset. Shijie Li 0006, Yazan Abu Farha, Yun Liu 0011, Ming-Ming Cheng, Juergen Gall |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Dual Pyramid Generative Adversarial Networks for Semantic Image Synthesis
Shijie Li 0006, Ming-Ming Cheng, Juergen Gall |
BMVC | 1 |
| 2021 | MiniSeg: An Extremely Minimum Network for Efficient COVID-19 SegmentationabstractThe rapid spread of the new pandemic, i.e., COVID-19, has severely threatened global health. Deep-learning-based computer-aided screening, e.g., COVID-19 infected CT area segmentation, has attracted much attention. However, the publicly available COVID-19 training data are limited, easily causing overfitting for traditional deep learning methods that are usually data-hungry with millions of parameters. On the other hand, fast training/testing and low computational cost are also necessary for quick deployment and development of COVID-19 screening systems, but traditional deep learning methods are usually computationally intensive. To address the above problems, we propose MiniSeg, a lightweight deep learning model for efficient COVID-19 segmentation. Compared with traditional segmentation methods, MiniSeg has several significant strengths: i) it only has 83K parameters and is thus not easy to overfit; ii) it has high computational efficiency and is thus convenient for practical deployment; iii) it can be fast retrained by other users using their private COVID-19 data for further improving performance. In addition, we build a comprehensive COVID-19 segmentation benchmark for comparing MiniSeg to traditional methods. Yun Liu 0011, Shijie Li 0006, Jing Xu 0008 |
AAAI | 3 |
| 2021 | Spatial-Temporal Consistency Network for Low-Latency Trajectory ForecastingabstractTrajectory forecasting is a crucial step for autonomous vehicles and mobile robots in order to navigate and interact safely. In order to handle the spatial interactions between objects, graph-based approaches have been proposed. These methods, however, model motion on a frame-to-frame basis and do not provide a strong temporal model. To overcome this limitation, we propose a compact model called Spatial-Temporal Consistency Network (STC-Net). In STC-Net, dilated temporal convolutions are introduced to model long-range dependencies along each trajectory for better temporal modeling while graph convolutions are employed to model the spatial interaction among different trajectories. Furthermore, we propose a feature-wise convolution to generate the predicted trajectories in one pass and refine the forecast trajectories together with the reconstructed observed trajectories. We demonstrate that STC-Net generates spatially and temporally consistent trajectories and outperforms other graph-based methods. Since STC-Net requires only 0.7k parameters and forecasts the future with a latency of only 1.3ms, it advances the state-of-the-art and satisfies the requirements for realistic applications. Shijie Li 0006, Yanying Zhou, Jinhui Yi, Juergen Gall |
ICCV | 1 |
| 2020 | Refinedbox: Refining for fewer and high-quality object proposals
Yun Liu 0011, Shijie Li 0006, Ming-Ming Cheng |
Neurocomputing | 2 |
| 2019 | Joint salient object detection and existence prediction
Huaizu Jiang, Ming-Ming Cheng, Shijie Li 0006, Ali Borji, Jingdong Wang 0001 |
Frontiers Comput. Sci. | 3 |
| 2018 | CubemapSLAM: A Piecewise-Pinhole Monocular Fisheye SLAM System
Shaojun Cai, Shijie Li 0006, Yun Liu 0011, Yangyan Guo, Tao Li 0022, Ming-Ming Cheng |
ACCV (6) | 3 |
| 2018 | Direct Line Guidance OdometryabstractModern visual odometry algorithms utilize sparse point-based features for tracking due to their low computational cost. Current state-of-the-art methods are split between indirect methods that process features extracted from the image, and indirect methods that deal directly on pixel intensities. In recent years, line-based features have been used in SLAM and have shown an increase in performance albeit with an increase in computational cost. In this paper, we propose an extension to a point-based direct monocular visual odometry method. Here we that uses lines to guide keypoint selection rather than acting as features. Points on a line are treated as stronger keypoints than those in other parts of the image, steering point-selection away from less distinctive points and thereby increasing efficiency. By combining intensity and geometry information from a set of points on a line, accuracy may also be increased. Shijie Li 0006, Bo Ren 0003, Yun Liu 0011, Ming-Ming Cheng, Duncan P. Frost, Victor Adrian Prisacariu |
ICRA | 1 |
| 2018 | DEL: Deep Embedding Learning for Efficient Image SegmentationabstractImage segmentation has been explored for many years and still remains a crucial vision problem. Some efficient or accurate segmentation algorithms have been widely used in many vision applications. However, it is difficult to design a both efficient and accurate image segmenter. In this paper, we propose a novel method called DEL (deep embedding learning) which can efficiently transform superpixels into image segmentation. Starting with the SLIC superpixels, we train a fully convolutional network to learn the feature embedding space for each superpixel. The learned feature embedding corresponds to a similarity measure that measures the similarity between two adjacent superpixels. With the deep similarities, we can directly merge the superpixels into large segments. The evaluation results on BSDS500 and PASCAL Context demonstrate that our approach achieves a good trade-off between efficiency and effectiveness. Specifically, our DEL algorithm can achieve comparable segments when compared with MCG but is much faster than it, i.e. 11.4fps vs. 0.07fps. Yun Liu 0011, Peng-Tao Jiang, Vahan Petrosyan, Shijie Li 0006, Jiawang Bian, Le Zhang 0001, Ming-Ming Cheng |
IJCAI | 4 |
| 2018 | Structured Skip List: A Compact Data Structure for 3D ReconstructionabstractThe model produced by 3D reconstruction algorithm is usually represented by voxels. The management of these voxels is usually divided into two categories: ordered and unordered methods. The ordered method holds too many empty voxels to maintain data order which leads to a low storage efficiency. On the contrary, the unordered method keeps massive index data to only store nonempty voxels. In this paper, we design a new data management method for real-time indoor 3D reconstruction, called Structured Skip List (SSL). The SSL can be treated as a semi-ordered method, because the advantages of both the ordered and unordered methods are taken into account: 1) it only holds nonempty voxels similar to the unordered method; 2) the structured information is introduced to reduce the storage space of index data. By these designs, the SSL has a better performance on storage efficiency. To handle the data collision in voxel allocation, a hash allocation list (HAL) is proposed. The length of each Skip List is kept balanced by fusing the IMU (Inertial Measurement Unit) information for a high operation efficiency. The storage efficiency analysis of different data management methods is shown in this paper. What's more, exhaustive investigation is carried out on several datasets with these methods. The experimental result demonstrates that our design can achieve a high storage efficiency with little time loss compared to the state-of-the-art methods. Shijie Li 0006, Ming-Ming Cheng, Yun Liu 0011, Shao-Ping Lu, Victor Adrian Prisacariu |
IROS | 1 |