VLDB 2026 Research / reviewers in the wild / expert
Shengming Li
dblp:66/9084
· DBLP profile ↗
16ranked-venue papers
3as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Robust Stereo Splatting SLAM System with Inertial-Legged FusionabstractRecent progress in stereo-based 3D Gaussian Splatting (3DGS) SLAM has enabled small-scale robots, which are too small to carry depth cameras, to achieve localization and reconstruct photorealistic scenes with high-speed rendering. However, initializing 3D Gaussians from binocular vision still requires further improvement, and the potential of robot proprioception has not been fully leveraged. This work presents a robust stereo 3DGS SLAM with efficient inertial-legged fusion for small-scale quadruped robots (SaQu-SLAM). We develop a light-weight network to densely initialize the 3D Gaussians in the space. Besides, an efficient fusion method of inertial and legged encoder data based on Kalman filter is introduced. To improve the cross-platform generalization of our algorithm, multiple configuration combinations of these three types of sensors are provided. Moreover, we propose a mode-switching mechanism to handle intermittent visual failures. At last, we perform evaluation on a benchmark dataset, which includes large- and small-scale scenes, and a small quadruped robot in real-world confined-scale scenes, reducing the absolute trajectory error by an average of 19%, 13% and 25% respectively, when compared with other state-of-the-art methods in a similar context. It is also the only successful method in our self-customized confined mixed textured and textureless scene, whereas all vision-based or visual-inertial methods fail. Our system achieves real-time performance even on an embedded platform (Jetson AGX Orin). Zuowei Chen, Yulai Zhang, Shengming Li, Toshio Fukuda |
IROS | 4 |
| 2025 | Practical searchable encryption scheme against response identity attacks
Shengming Li, Xuan Jing, Yunling Wang, Jianfeng Wang 0001 |
Inf. Sci. | 1 |
| 2025 | Self-Adaptive Vision-Language Tracking With Context PromptingabstractDue to the substantial gap between vision and language modalities, along with the mismatch problem between fixed language descriptions and dynamic visual information, existing vision-language tracking methods exhibit performance on par with or slightly worse than vision-only tracking. Effectively exploiting the rich semantics of language to enhance tracking robustness remains an open challenge. To address these issues, we propose a self-adaptive vision-language tracking framework that leverages the pre-trained multi-modal CLIP model to obtain well-aligned visual-language representations. A novel context-aware prompting mechanism is introduced to dynamically adapt linguistic cues based on the evolving visual context during tracking. Specifically, our context prompter extracts dynamic visual features from the current search image and integrates them into the text encoding process, enabling self-updating language embeddings. Furthermore, our framework employs a unified one-stream Transformer architecture, supporting joint training for both vision-only and vision-language tracking scenarios. Our method not only bridges the modality gap but also enhances robustness by allowing language features to evolve with visual context. Extensive experiments on four vision-language tracking benchmarks demonstrate that our method effectively leverages the advantages of language to enhance visual tracking. Our large model can obtain 55.0% AUC on $\text {LaSOT}_{\text {EXT}}$ and 69.0% AUC on TNL2K. Additionally, our language-only tracking model achieves performance comparable to that of state-of-the-art vision-only tracking methods on TNL2K. Code is available at https://github.com/zj5559/SAVLT. Jie Zhao 0014, Xin Chen 0032, Shengming Li, Chunjuan Bo, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Image Process. | 3 |
| 2025 | Event-Assisted Recurrent Network for Arbitrary-Temporal-Scale Blurry Image UnfoldingabstractRecovering a sequence of latent sharp frames from a motion-blurred image is a challenging task. The bio-inspired event camera, which produces an event stream with high temporal resolution, has been exploited to promote the recovery performance. However, recovering sharp sequences with arbitrary temporal scales has been ignored for a long time. Existing works can only recover a fixed number of latent frames from a blurry image once they are trained. In this work, we propose an event-assisted blurry image unfolding framework that can work across arbitrary temporal scales. A bi-directional recurrent network is employed to encode events corresponding to each latent frame, which gathers information over all events in the exposure time. Features of both the blurry image and events are fused together and fed to a bi-directional latent sequence decoder (BiLSD) to produce a sequence of latent sharp frames. Extensive experiments show that the proposed method not only performs favorably against state-of-the-art methods in recovering a fixed number of frames from a blurry image but can be well generalized to arbitrary-temporal-scale blurry image unfolding. Hao Ju 0004, Weihua He, Yaoyuan Wang, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | EvSign: Sign Language Recognition and Translation with Streaming Events
Zeren Wang, Wenyue Chen, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
ECCV (5) | 5 |
| 2024 | Multi-Stage Fusion for Event-based Multimodal TrackerabstractEvent cameras are bio-inspired sensors with high dynamic range and time resolution, which are favorable properties for visual object tracking. There are already some methods that fuse the event modality and RGB modality with cross-domain feature integrator to achieve improved tracking performance. Researchers have developed some architectures for event modality processing or fusion, successfully boosting the tracking performance. In this work, we design a RGB-E tracker with multi-stage fusion. In the early stage, frames are enhanced with aid of events to mitigate blur or under/over-exposure degradation. During the middle stage, we utilize a fusion module for feature-level integration. At the late stage, we carry out decision-level fusion by predicting tracking boxes based on frame features, event features, and fused features, and the one with highest score is taken as the final estimation. Our design thoroughly integrate information from various levels, allowing each modality to contribute to the tracking process as much as possible. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art RGB-E trackers in both accuracy and efficiency. Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Wenyue Chen, Dong Wang 0004, Shengming Li, Huchuan Lu |
ICME | 6 |
| 2024 | DCPT: Darkness Clue-Prompted Tracking in Nighttime UAVsabstractExisting nighttime unmanned aerial vehicle (UAV) trackers follow an "Enhance-then-Track" architecture - first using a light enhancer to brighten the nighttime video, then employing a daytime tracker to locate the object. This separate enhancement and tracking fails to build an end-to-end trainable vision system. To address this, we propose a novel architecture called Darkness Clue-Prompted Tracking (DCPT) that achieves robust UAV tracking at night by efficiently learning to generate darkness clue prompts. Without a separate enhancer, DCPT directly encodes anti-dark capabilities into prompts using a darkness clue prompter (DCP). Specifically, DCP iteratively learns emphasizing and undermining projections for darkness clues. It then injects these learned visual prompts into a daytime tracker with fixed parameters across transformer layers. Moreover, a gated feature aggregation mechanism enables adaptive fusion between prompts and between prompts and the base model. Extensive experiments show state-of-the-art performance for DCPT on multiple dark scenario benchmarks. The unified end-to-end learning of enhancement and tracking in DCPT enables a more trainable system. The darkness clue prompting efficiently injects anti-dark knowledge without extra modules. Code is available at https://github.com/bearyi26/DCPT. Jiawen Zhu 0003, Huayi Tang, Zhi-Qi Cheng, Jun-Yan He, Bin Luo 0008, Shihao Qiu, Shengming Li, Huchuan Lu |
ICRA | 7 |
| 2024 | Event-Guided Rolling Shutter Correction with Time-Aware Cross-AttentionsabstractMany consumer cameras with rolling shutter (RS) CMOS would suffer undesired distortion and artifacts, particularly when objects experiences fast motion. The neuromorphic event camera, with high temporal resolution events, could bring much benefit to the RS correction process. In this work, we explore the characteristics of RS images and event data for the design of the rolling shutter correction (RSC) model. Specifically, the relationship between RS images and event data is modeled by incorporating time encoding to the computation of cross-attention in transformer encoder to achieve time-aware multi-modal information fusion. Features from RS images enhanced by event data are adopted as keys and values in transformer decoder, providing source for appearance, while features from event data enhanced by RS images are adopted as queries, providing spatial transition information. By embedding the time information of the desired global shutter (GS) image into the query, the transformer with deformable attention is capable of producing the target GS image.To enhance the model's generalization ability, we propose to further self-supervise the model by cycling between time coordinate systems corresponding to RS images and GS images. Extensive evaluations over both synthetic and real datasets demonstrate that the proposed method performs favorably against state-of-the-art approaches. Hefei Huang, Xu Jia 0012, Xinyu Zhang 0017, Shengming Li, Huchuan Lu |
ACM Multimedia | 4 |
| 2024 | LGTrack: Exploiting Local and Global Properties for Robust Visual TrackingabstractRe-detection is a necessary capability for long-term tracking. Target candidate proposals in the whole image can provide a chance of tracking reset when tracking fails due to tracking drift or target invisibility. In this paper, we propose a unified local-global tracker based on the same transformer architecture sharing weights, which can not only search in a continuous local region but also provide target candidates of the global image in every frame. The requirements of both long-term and short-term scenarios can be addressed using a unified model. A simple proposal selection scheme is adopted to properly select the candidate proposals of re-detection, to assist tracking and obtain better performance. The scheme performs reevaluation of all high-quality proposals based on a transformer-based embedding network, once the predicted state of the local tracking is not sufficient to be accurate. To capture appearance variations brought by online updates in minimum risks, a long-term-friendly dynamic template update scheme is also designed. Extensive experiments are conducted to demonstrate the effectiveness of our proposed tracker, including three short-term tracking benchmarks and six long-term benchmarks. Our tracker can achieve results comparable to that of the state-of-the-art. The proposed tracker can also work well in balancing the performance and speed, achieving an average speed of approximately 25 fps tested on LaSOT testing set. Chang Liu 0071, Jie Zhao 0014, Chunjuan Bo, Shengming Li, Dong Wang 0004, Huchuan Lu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Event-Assisted Blurriness Representation Learning for Blurry Image UnfoldingabstractThe goal of blurry image deblurring and unfolding task is to recover a single sharp frame or a sequence from a blurry one. Recently, its performance is greatly improved with introduction of a bio-inspired visual sensor, event camera. Most existing event-assisted deblurring methods focus on the design of powerful network architectures and effective training strategy, while ignoring the role of blur modeling in removing various blur in dynamic scenes. In this work, we propose to implicitly model blur in an image by computing blurriness representation with an event-assisted blurriness encoder. The learning of blurriness representation is formulated as a ranking problem based on specially synthesized pairs. Blurriness-aware image unfolding is achieved by integrating blur relevant information contained in the representation into a base unfolding network. The integration is mainly realized by the proposed blurriness-guided modulation and multi-scale aggregation modules. Experiments on GOPRO and HQF datasets show favorable performance of the proposed method against state-of-the-art approaches. More results on real-world data validate its effectiveness in recovering a sequence of latent sharp frames from a blurry image. Hao Ju 0004, Lei Yu 0006, Weihua He, Yaoyuan Wang, Qi Xu 0008, Shengming Li, Dong Wang 0004, Huchuan Lu, Xu Jia 0012 |
IEEE Trans. Image Process. | 8 |
| 2023 | Forgery face detection via adaptive learning from multiple experts
Xinghe Fu, Shengming Li, Yike Yuan, Bin Li 0038, Xi Li 0001 |
Neurocomputing | 2 |
| 2022 | Entropy-Driven Sampling and Training Scheme for Conditional Diffusion Generation
Guangcong Zheng, Shengming Li, Hui Wang 0107, Taiping Yao, Shouhong Ding, Xi Li 0001 |
ECCV (22) | 2 |
| 2022 | Double cross-modality progressively guided network for RGB-D salient object detection
Cuili Yao, Lin Feng 0001, Yuqiu Kong, Shengming Li |
Image Vis. Comput. | 4 |
| 2022 | Object detection network pruning with multi-task information fusion
Shengming Li, Linsong Xue, Lin Feng 0001 |
World Wide Web | 1 |
| 2019 | Design and Implementation of an IoT-Based Indoor Air Quality Detector With Multiple Communication InterfacesabstractIndoor air quality (IAQ) monitoring has attracted increasing attention with the rapid development of industrialization and urbanization in the modern society as people typically spend more than 80% of their time in indoor environments. A novel IAQ detector (IAQD) integrated with multiple communication interfaces has been designed, built, programmed, deployed, and tested in order to meet the requirements of wide variety of scenarios. The IAQD measures the IAQ data, including temperature, humidity, CO2, dust, and formaldehyde timely. With state-of-the-art Internet of Things (IoT) technologies, the IAQD is integrated with Modbus, LoRa, WiFi, general packet radio service, and NB-IoT communication interfaces, which enables to be applied to wired communications, short-range wireless communications, and remote transmission to the cloud. The designed software in cloud allows users to track the IAQ of their home or office or industries everywhere. The performance IAQD is evaluated in terms of packet loss rate and time delay. The evaluation of IAQD are demonstrated and analyzed within the office environment over a week. Experimental results show that the proposed system is effectiveness in measuring the air-quality status and provide excellent consistency and stability. Liang Zhao 0015, Wenyan Wu 0002, Shengming Li |
IEEE Internet Things J. | 3 |
| 2013 | A workload prediction-based multi-VM provisioning mechanism in cloud computing
Shengming Li, Ying Wang 0002, Xuesong Qiu 0001, Deyuan Wang |
APNOMS | 1 |