Jisheng Li

dblp:180/2684 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 STA-HRNN-BiGRU: a spatiotemporal sequence multivariate forecasting model integrating spatiotemporal information
Junyan Sun, Zhongrong Zhang, Kang Lin, Zeyu Duan, Jisheng Li
Appl. Intell.6
2026 Wavelet-based high-frequency fusion for multi-class segmentation of fecal pathological components in microscopic images
Nuo Tong, Shuiping Gou, Bianping Liang, Shaobin Deng, Jisheng Li, Mingxue Wang
Eng. Appl. Artif. Intell.6
2025 Adaptation and learning of spatio-temporal thresholds in spiking neural networks
Shuiping Gou, Peizhao Wang, Licheng Jiao, Zhang Guo 0001, Jisheng Li
Neurocomputing6
2023 DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW Images
abstract
Low-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, however, the performance of existing methods degrades in the extreme low-light scenarios in a certain degree, due to the low signal-to-noise ratio in images. To address this challenge, images in RAW format that retain more raw sensing information have been considered in recent works with a denoise-then-detect scheme. However, existing denoising methods are still insufficient for RAW images and heavily time-consuming, which limits the practical applications of such scheme. In this paper, we propose DarkFeat, a deep learning model which directly detects and describes local features from extreme low-light RAW images in an end-to-end manner. A novel noise robustness map and selective suppression constraints are proposed to effectively mitigate the influence of noise and extract more reliable keypoints. Furthermore, a customized pipeline of synthesizing dataset containing low-light RAW image matching pairs is proposed to extend end-to-end training. Experimental results show that DarkFeat achieves state-of-the-art performance on both indoor and outdoor parts of the challenging MID benchmark, outperforms the denoise-then-detect methods and significantly reduces computational costs up to 70%. Code is available at https://github.com/THU-LYJ-Lab/DarkFeat.
Yubin Hu 0001, Wang Zhao 0001, Jisheng Li, Yong-Jin Liu 0001, Yuxing Han 0001, Jiangtao Wen
AAAI4
2023 Efficient Semantic Segmentation by Altering Resolutions for Compressed Videos
abstract
Video semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the computational cost from the input side: the input resolution. In this paper, we propose an altering resolution framework called AR-Seg for compressed videos to achieve efficient VSS. AR-Seg aims to reduce the computational cost by using low resolution for non-keyframes. To prevent the performance degradation caused by downsampling, we design a Cross Resolution Feature Fusion (CR-eFF) module, and supervise it with a novel Feature Similarity Training (FST) strategy. Specifically, CReFF first makes use of motion vectors stored in a compressed video to warp features from high-resolution keyframes to low-resolution non-keyframes for better spatial alignment, and then selectively aggregates the warped features with local attention mechanism. Furthermore, the proposed FST supervises the aggregated features with high-resolution features through an explicit similarity loss and an implicit constraint from the shared decoding layer. Extensive experiments on CamVid and Cityscapes show that AR-Seg achieves state-of-the-art performance and is compatible with different segmentation backbones. On CamVid, AR-Seg saves 67% computational cost (measured in GFLOPs) with the PSPNet18 back-bone while maintaining high segmentation accuracy. Code: https://github.com/THU-LYJ-Lab/AR-Seg.
Yubin Hu 0001, Yanghao Li, Jisheng Li, Yuxing Han 0001, Jiangtao Wen, Yong-Jin Liu 0001
CVPR4
2022 Rate Control for Learned Video Compression
abstract
Rate control is a critical part for video compression, especially in bandwidth-limited tasks such as live and broadcast. The newly-rising learned video compression has shown advantageous rate-distortion (RD) performance in previous research, but lack of rate control heavily limits its usage in real coding scenarios. In this work, we present the first rate control scheme tailored for learned video compression. Specifically, we explore the inter-frame dependency of learned video compression and propose a novel R-D-λ model accordingly for efficient rate allocation. Additionally, a staged update algorithm is developed for robust parameter estimation. Experiments on public datasets show that, the proposed rate control scheme achieves low rate error while maintaining equal or even higher RD performance, without introducing coding time overhead.
Yanghao Li, Jisheng Li, Jiangtao Wen, Yuxing Han 0001, Shan Liu 0001, Xiaozhong Xu
ICASSP3
2021 Learning to Estimate Kernel Scale and Orientation of Defocus Blur with Asymmetric Coded Aperture
abstract
Consistent in-focus input imagery is an essential precondition for machine vision systems to perceive the dynamic environment. A de-focus blur severely degrades the performance of vision systems. To tackle this problem, we propose a deep-learning-based framework estimating the kernel scale and orientation of the defocus blur to ad-just lens focus rapidly. Our pipeline utilizes 3D ConvNet for a variable number of input hypotheses to select the optimal slice from the input stack. We use random shuffle and Gumbel-softmax to improve network performance. We also propose to generate synthetic defocused images with various asymmetric coded apertures to facilitate training. Experiments are conducted to demonstrate the effectiveness of our framework.
Jisheng Li, Jiangtao Wen
ICASSP1
2021 Learning To Compose 6-DOF Omnidirectional Videos Using Multi-Sphere Images
abstract
Omnidirectional video is an essential component of Virtual Reality. Although various methods have been proposed to generate content that can be viewed with six degrees of freedom (6-DoF), existing systems usually involve complex depth estimation, image inpainting or stitching pre-processing. In this paper, we propose a system that uses a 3D ConvNet to generate a multi-sphere images (MSI) representation that can be experienced in 6-DoF VR. The system utilizes conventional omnidirectional VR camera footage directly without the need for a depth map or segmentation mask, thereby significantly simplifying the overall complexity of the 6-DoF omnidirectional video composition. By using a newly designed weighted sphere sweep volume (WSSV) fusing technique, our approach is compatible with most panoramic VR camera setups. A ground truth generation approach for high-quality artifact-free 6-DoF contents is proposed and can be used by the research and development community for 6-DoF content generation.
Jisheng Li, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen
ICIP1
2021 Extending 6-DoF VR Experience Via Multi-Sphere Images Interpolation
abstract
Three-degrees-of-freedom (3-DoF) omnidirectional imaging has been widely used in various applications ranging from street maps to 3-DoF VR live broadcasting. Although allowing for navigating viewpoints rotationally inside a virtual world, it does not provide motion parallax key for human 3D perception. Recent research mitigates this problem by introducing 3 transitional degrees of freedom (6-DoF) using multi-sphere images (MSI) which is beginning to show promises in handling occlusions and reflective objects. However, the design of MSI naturally limits the range of authentic 6-DoF experiences, as existing mechanisms for MSI rendering cannot fully utilize multi-layer information when synthesizing novel views between multiple MSIs. To tackle this problem and extend the 6-DoF range, we propose an MSI interpolation pipeline that utilizes adjacent MSIs' 3D information embedded inside their layers. In this work, we describe an MSI projection scheme along with an MSI interpolation network to predict intermediate MSIs in order to facilitate the need for extended range. We demonstrate that our system significantly improves the range of 6-DoF experience compared with other MSI-based methods. With extensive experiments, we show our algorithm outperforms state-of-the-art methods both qualitatively and quantitatively in synthesizing novel view panoramas.
Jisheng Li, Jinghui Jiao, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen
ACM Multimedia1
2016 Fireworks algorithm for the satellite link scheduling problem in the navigation constellation
abstract
Global navigation satellite system (GNSS) can provide autonomous geo-spatial positioning and time synchronization services for both civil and military uses. Satellite links in GNSS are used to transmit signal for constellation management and other applications. In this work, we focus on solving the satellite link scheduling problem over dynamic satellite network with the aim of minimizing the number of participant ground-based management stations and the cost of communication between satellites in the background of GNSS networking. Firstly, we assume the navigation constellation has finite states and cope the dynamic topology with Finite State Automation method. Secondly, a Fireworks algorithm (FWA) is designed according to the characteristic of the scheduling problem. Finally, the FWA is compared with ant colony optimization (ACO). The performance analysis of different scenarios is given. The study in this paper provides technical reference for the management of future large-scale satellite network.
Liangjun Ke, Jisheng Li, Jingqi Huang
CEC3
2016 Intra Frame Flicker Reduction for Parallelized HEVC Encoding
abstract
The existing intra flicker artifact reduction approaches, targeting at one of the major artifacts in current video encoding techniques, are not compatible with the distributed encoding structure, which is increasingly important in modern computing systems. To settle this problem, we propose a flicker reduction approach, which is effective, standard compliant, and especially suitable for parallel and distributed systems. Experimental results show that the proposed approach can reduce the flicker artifact by up to 60% on x265 and 14% on HM.
Ziyu Wen, Jisheng Li, Jiangtao Wen
DCC2
2016 Novel 3D-WPP algorithms for parallel HEVC encoding
abstract
Although wavefront parallel processing (WPP) proposed in the HEVC standard and various inter frame WPP algorithms can achieve comparatively high parallelism, their scalability for its parallelism is still very limited due to various dependencies introduced in spatial and temporal prediction in HEVC. In this paper, we propose three types of 3 Dimensional WPP (3D-WPP) algorithms that can significantly improve the parallelism, while achieving good tradeoffs between implementation complexity, determinism, and rate-distortion (RD) performance. Experimental results show that the proposed algorithms can lead to up to 2.8 × speed up compared with existing inter frame WPP methods. While the Simple 3D-WPP and Static 3D-, WPP algorithm may introduce an BD rate loss between 0 to 4.9% as compared with existing algorithms, the more complex Dynamic 3D-WPP algorithm achieves better parallelism with virtually no coding performance loss.
Ziyu Wen, Bichuan Guo, Jisheng Li, Yao Lu 0006, Jiangtao Wen
ICASSP4
2016 Novel tile segmentation scheme for omnidirectional video
abstract
Regular omnidirectional video encoding technics use map projection to flatten a scene from a spherical shape into one or several 2D shapes. Common projection methods including equirectangular and cubic projection have varying levels of interpolation that create a large number of non-information-carrying pixels that lead to wasted bitrate. In this paper, we propose a tile based omnidirectional video segmentation scheme which can save up to 28% of pixel area and 20% of BD-rate averagely compared to the traditional equirectangular projection based approach.
Jisheng Li, Ziyu Wen, Bichuan Guo, Jiangtao Wen
ICIP1