EDBT 2026 Demo / reviewers in the wild / expert
Jing Li 0010
dblp:l/JingLi10
· DBLP profile ↗
24ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-9043-8633ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 since 2021Computer networks · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LD3DGS-SLAM: Long-Distance Monocular SLAM With 3-D Gaussian Splatting and GNSS-Aided Localization for UAVsabstractIntegrating neural rendering with simultaneous localization and mapping (SLAM) has shown great promise for achieving high-precision localization and photorealistic scene reconstruction. However, the dynamic perspectives of uncrewed aerial vehicles (UAVs), compounded by Global Navigation Satellite System (GNSS) signal interference and the accumulation of localization errors, pose significant challenges to existing monocular simultaneous localization and mapping (SLAM) methods in complex environments. To address these limitations, we propose LD3DGS-SLAM, a novel framework that integrates monocular SLAM, GNSS, and 3D Gaussian splatting (3DGS) [1] to enhance UAV-based localization and mapping. The system integrates traditional geometric feature-based localization with the multiview synthesis capability of 3D Gaussian rendering, achieving a breakthrough in accuracy while enhancing robustness. First, we construct a multisensor fusion graph optimization model that tightly integrates GNSS and monocular vision data, effectively mitigating the cumulative drift typical of traditional SLAM pipelines. Afterward, we introduce a 3D Gaussian-based mapping strategy that incrementally refines a dense scene representation using SLAM-generated point clouds, fusing geometric and texture information to reconstruct highly accurate and detailed maps. To further enhance localization performance in GNSS-denied scenarios, we develop a dynamic frame interpolation and tracking approach based on 3D Gaussian rendering. By synthesizing novel viewpoints, this method improves observation matching and strengthens loop closure detection, enabling robust relocalization in large-scale environments. To validate the effectiveness of the proposed framework, we implement a UAV test system and evaluate its performance on both the publicly available TUM [2] dataset and a self-collected aerial dataset spanning several kilometers. Experimental results show that LD3DGS-SLAM achieves a state-of-the-art average localization error of only 0.8 meters using just a monocular camera, even during long-range flight missions exceeding tens of kilometers and altitudes up to 500 meters. Overall, LD3DGS-SLAM effectively addresses the limitations of monocular SLAM in aerial scenarios, providing a robust, accurate, and cost-effective localization solution for urban air mobility and aerial Internet of Things (IoT) applications. Dongdong Li 0008, Tao Yang 0006, Haidong Qin, Shixiong Fan, Shuanghan Zhang, Jing Li 0010 |
IEEE Internet Things J. | 8 |
| 2026 | GaussHead: Real-Time 3-D Head Avatar Driving for Mobile With Compact Gaussian Representation
Xiaoshi Zhou, Tao Yang 0006, Baogang Song, Chengwei Cao, Hongwei Bao, Jing Li 0010 |
IEEE Internet Things J. | 9 |
| 2025 | RT3DHVC: A Real-Time Human Holographic Video Conferencing System With a Consumer RGB-D Camera ArrayabstractIn this paper, we present an end-to-end holographic video conferencing system that enables real-time high-quality free-viewpoint rendering of participants in different spatial regions, placing them in a unified virtual space for a more immersive display. Our system offers a cost-effective, complete holographic conferencing process, including multiview 3D data capture, RGB-D stream compression and transmission, high-quality rendering, and immersive display. It employs a sparse set of commodity RGB-D cameras that capture 3D geometric and textural information. We then remotely transmit color and depth maps via standard video encoding and transmission protocols. We propose a GPU-parallelized rendering pipeline based on an image-based virtual view synthesis algorithm to achieve real-time and high-quality scene rendering. This algorithm uses an on-the-fly Truncated Signed Distance Function (TSDF) approach, which marches along virtual rays within a computed precise search interval to determine surface intersections. We then design a multiweight projective texture mapping method to fuse color information from multiple views. Furthermore, we introduce a method that uses a depth confidence map to weight the rendering results from different views, which mitigates the impact of sensor noise and inaccurate measurements on the rendering results. Finally, our system places conference participants from different spaces into a virtual conference environment with a global coordinate system through coordinate transformation, which simulates a real conference scene in physical space, providing an immersive remote conferencing experience. Experimental evaluations confirm our system’s real-time, low-latency, high-quality, and immersive capabilities. Jing Li 0010, Yanran Dai, Haidong Qin, Xiaoshi Zhou, Kefan Yan, Tao Yang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | ECC-NeRF: Anti-Aliasing Neural Radiance Fields With Elliptic Cone-Casting for Diverse Camera ModelsabstractAnti-aliasing is a crucial research topic in computer graphics, which can significantly enhance the rendering quality of neural radiance fields (NeRF). Recent studies have introduced effective anti-aliasing NeRF methods, utilizing cone-casting to replace ray-casting and modeling the 3D observation area of pixels as circular cones. The cone-casting strategy has successfully reduced blurring and aliasing in novel view rendering. However, we have observed that the light cones are not standard circular cones because the camera projection model distorts them into elliptic cones of diverse sizes and shapes. This finding motivates us to model pixel light cones as anisotropic elliptic cones and propose an elliptic cone-casting-based anti-aliasing NeRF method called “ECC-NeRF". Specifically, we first derive the elliptic cone models for common pinhole, fisheye, and panoramic cameras based on their camera projection models. Then, we integrate the proposed elliptic cone-casting into two representative cone-casting-based anti-aliasing NeRF methods: Mip-NeRF and Zip-NeRF. Our experimental evaluations on multiple datasets demonstrate that our method can achieve more accurate multi-scale anisotropic representation and better novel view rendering quality with negligible additional computation cost. Haidong Qin, Tao Yang 0006, Xiaoshi Zhou, Dongdong Li 0008, Yanran Dai, Jing Li 0010 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Real-time distance field acceleration based free-viewpoint video synthesis for large sports fieldsabstractFree-viewpoint video allows the user to view objects from any virtual perspective, creating an immersive visual experience. This technology enhances the interactivity and freedom of multimedia performances. However, many free-viewpoint video synthesis methods hardly satisfy the requirement to work in real time with high precision, particularly for sports fields having large areas and numerous moving objects. To address these issues, we propose a free-viewpoint video synthesis method based on distance field acceleration. The central idea is to fuse multi-view distance field information and use it to adjust the search step size adaptively. Adaptive step size search is used in two ways: for fast estimation of multi-object three-dimensional surfaces, and synthetic view rendering based on global occlusion judgement. We have implemented our ideas using parallel computing for interactive display, using CUDA and OpenGL frameworks, and have used real-world and simulated experimental datasets for evaluation. The results show that the proposed method can render free-viewpoint videos with multiple objects on large sports fields at 25 fps. Furthermore, the visual quality of our synthetic novel viewpoint images exceeds that of state-of-the-art neural-rendering-based methods. Yanran Dai, Jing Li 0010, Haidong Qin, Bang Liang, Shikuan Hong, Haozhe Pan, Tao Yang 0006 |
Comput. Vis. Media | 2 |
| 2024 | GS-SFS: Joint Gaussian Splatting and Shape-From-Silhouette for Multiple Human Reconstruction in Large-Scale Sports ScenesabstractWe introduce GS-SFS, a method that utilizes a camera array with wide baselines for high-quality multiple human mesh reconstruction in large-scale sports scenes. Traditional human reconstruction methods in sports scenes, such as Shape-from-Silhouette (SFS), struggle with sparse camera setups and small human targets, making it challenging to obtain complete and accurate human representations. Despite advances in differentiable rendering, including 3D Gaussian Splatting (3DGS), which can produce photorealistic novel-view renderings with dense inputs, accurate depiction of surfaces and generation of detailed meshes is still challenging. Our approach uniquely combines 3DGS's view synthesis with an optimized SFS method, thereby significantly enhancing the quality of multiperson mesh reconstruction in large-scale sports scenes. Specifically, we introduce body shape priors, including the human surface point clouds extracted through SFS and human silhouettes, to constrain 3DGS to a more accurate representation of the human body only. Then, we develop an improved mesh reconstruction method based on SFS, mainly by adding additional viewpoints through 3DGS and obtaining a more accurate surface to achieve higher-quality reconstruction models. We implement a high-density scene resampling strategy based on spherical sampling of human bounding boxes and render new perspectives using 3D Gaussian Splatting to create precise and dense multi-view human silhouettes. During mesh reconstruction, we integrate the human body's 2D Signed Distance Function (SDF) into the computation of the SFS's implicit surface field, resulting in smoother and more accurate surfaces. Moreover, we enhance mesh texture mapping by blending original and rendered images with different weights, preserving high-quality textures while compensating for missing details. The experimental results from real basketball game scenarios demonstrate the significant improvements of our approach for multiple human body model reconstruction in complex sports settings. Jing Li 0010, Haidong Qin, Yanran Dai, Jing Liu 0006, Canbin Zhang, Tao Yang 0006 |
IEEE Trans. Multim. | 2 |
| 2023 | Calibration-Free Cross-Camera Target Association Using Interaction Spatiotemporal ConsistencyabstractIn this paper, we propose a novel calibration-free cross-camera target association algorithm that aims to relate local visual data of the same object across cameras with overlapping FOVs. Unlike other methods using object's own characteristics, our approach makes full use of the interactions between objects and explores their spatiotemporal consistency in projection transformation to associate cameras. It has wider applicability in deployed overlapping multi-camera systems with unknown or rarely available calibration data, especially if there is a large perspective gap between cameras. Specifically, we first extract trajectory intersection which is one of the typical object-object interactive behaviors from each camera for feature vector construction. Then, based on the consistency of object-object interactions, we propose a multi-camera spatiotemporal alignment method via wide-domain cross-correlation analysis. It realizes time synchronization and spatial calibration of the multi-camera system simultaneously. After that, we introduce a cross-camera target association approach using aligned object-object interactions. The local data of the same target are successfully associated across cameras without any additional calibration. Extensive experimental evaluations on different databases verify the effectiveness and robustness of our proposed method. Jing Li 0010, Yuguang Xie, Jiayang Nie, Tao Yang 0006, Zhaoyang Lu |
IEEE Trans. Multim. | 2 |
| 2023 | Bullet-Time Video Synthesis Based on Virtual Dynamic Target AxisabstractBullet-time videos have been widely used in movies, TV advertisements, and computer games, and can produce an immersive and smooth orbital free-viewpoint of frozen action. However, existing bullet-time video synthesis methods remain challenging in practical applications, especially in complex situations with poor camera calibration and a variety of camera array structures. This paper proposes a novel bullet-time video synthesis method based on a virtual dynamic target axis. We adopt an image similarity transformation strategy to eliminate image distortion in the bullet-time video. We use a high-order polynomial curve fitting strategy to reserve more bullet-time video frame content. The proposed dynamic target axis strategy can support various camera array structures, including camera arrays with and without a common field of view. In addition, this strategy can also tolerate poor camera calibration situations with unevenly distributed reprojection errors to some extent and synthesize smooth bullet-time videos without high-precision camera calibration. Qualitative and quantitative experiments in real environments and on simulation platforms demonstrate the high performance of our bullet-time video synthesis method. Compared with the state-of-the-art methods, the proposed method shows superiority. Haidong Qin, Jing Li 0010, Yanran Dai, Shikuan Hong, Tao Yang 0006 |
IEEE Trans. Multim. | 2 |
| 2022 | Multi-camera joint spatial self-organization for intelligent interconnection surveillance
Jing Li 0010, Yuguang Xie, Jiayang Nie, Tao Yang 0006, Zhaoyang Lu |
Eng. Appl. Artif. Intell. | 2 |
| 2021 | Image-Only Real-Time Incremental UAV Image Mosaic for Multi-Strip FlightabstractLimited by aircraft flight altitude and camera parameters, it is necessary to obtain wide-angle panoramas quickly by stitching aerial images, which is helpful in rapid disaster investigation, recovery after earthquakes, and aerial reconnaissance. However, most existing stitching algorithms do not simultaneously meet practical real-time, robustness, and accuracy requirements, especially in the case of a long-distance multistrip flight. In this paper, we propose a novel image-only real-time UAV image mosaic framework for long-distance multistrip flights that does not require any auxiliary information, such as GPS or GCPs. The framework has a complete structure, mainly consisting of the three tasks of automatic initialization, current frame tracking, and real-time mosaic generation. The stitching plane is determined in the initialization process, the homography transformation of the current image is estimated in the tracking task, and the image is mapped to the stitching plane to generate and update the panorama in the real-time mosaic process. The core idea is that, in the tracking task, we introduce and develop a keyframe insertion strategy to generate a keyframe list and, on this basis, design a homography matrix estimation based on a local optimization strategy to reduce the accumulated error when continuously stitching image sequences collected online by UAVs and to realize real-time, effective UAV image mosaic construction. In addition, this framework has good scalability, which is not limited to a specific algorithm. To evaluate the effectiveness of the proposed framework, we carry out a large number of experiments on the AirSim simulation platform and present an exhaustive evaluation in some sequences from a popular dataset. Qualitative and quantitative experimental results in simulation and real environments demonstrate that our algorithm can obtain an effective and robust mosaic image in real-time. Through strategy comparison experiments, it is proven that the keyframe insertion strategy and the local optimization strategy both improve the stitching performance. Compared with five state-of-art image stitching approaches, the mosaic effect of the proposed method is comparable or better. In terms of algorithm speed, its performance is superior to them. Additionally, experiments of illumination change and feature replacement in the framework verify the good adaptability and scalability of the algorithm. Fangbing Zhang, Tao Yang 0006, Linfeng Liu 0002, Bang Liang, Jing Li 0010 |
IEEE Trans. Multim. | 6 |
| 2020 | Unified multiple access structure based on FBMC modulation for multi-RAT coexistence in heterogeneous wireless networksabstractWith the evolution of the fifth generation (5G) and beyond heterogeneous wireless networks, telecommunications operators and researchers are driven to develop advanced transmission technologies that enable the coexistence of multiple radio access technologies (multi‐RATs). By employing the efficient implementation of filter‐bank multi‐carrier (FBMC) transceiver and the scalable matrix transformation (SMT) module, in this study, a unified multiple access architecture termed as FBMC‐SMT is proposed, which is capable of flexibly integrating with both 3G and 4G transmission schemes and improving the system performance. As a special case of FBMC‐SMT, the performance evaluation of FBMC‐CDMA is conducted by extensive simulation. It is verified that FBMC‐CDMA with 16 subbands allocated outperforms the traditional single carrier wideband code division multiple access when the number of code channels assigned is larger than 5. Moreover, the signal‐to‐interference‐plus‐noise ratio analysis of FBMC‐SMT is also given and matches the simulation results very well. Hence, the proposed FBMC‐SMT can serve as a unified architecture to flexibly integrate multi‐RATs, thus to meet variable application requirements in 5G and beyond heterogeneous wireless networks. Xin Bian, Yun Rui, Jing Li 0010, Rongfang Song |
IET Commun. | 4 |
| 2020 | Deep Image-to-Video Adaptation and Fusion Networks for Action RecognitionabstractExisting deep learning methods for action recognition in videos require a large number of labeled videos for training, which is labor-intensive and time-consuming. For the same action, the knowledge learned from different media types, e.g., videos and images, may be related and complementary. However, due to the domain shifts and heterogeneous feature representations between videos and images, the performance of classifiers trained on images may be dramatically degraded when directly deployed to videos. In this paper, we propose a novel method, named Deep Image-to-Video Adaptation and Fusion Networks (DIVAFN), to enhance action recognition in videos by transferring knowledge from images using video keyframes as a bridge. The DIVAFN is a unified deep learning model, which integrates domain-invariant representations learning and cross-modal feature fusion into a unified optimization framework. Specifically, we design an efficient cross-modal similarities metric to reduce the modality shift among images, keyframes and videos. Then, we adopt an autoencoder architecture, whose hidden layer is constrained to be the semantic representations of the action class names. In this way, when the autoencoder is adopted to project the learned features from different domains to the same space, more compact, informative and discriminative representations can be obtained. Finally, the concatenation of the learned semantic feature representations from these three autoencoders are used to train the classifier for action recognition in videos. Comprehensive experiments on four real-world datasets show that our method outperforms some state-of-the-art domain adaptation and action recognition methods. Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006 |
IEEE Trans. Image Process. | 3 |
| 2019 | Hierarchically Learned View-Invariant Representations for Cross-View Action RecognitionabstractRecognizing human actions from varied views is challenging due to huge appearance variations in different views. The key to this problem is to learn discriminant view-invariant representations generalizing well across views. In this paper, we address this problem by learning view-invariant representations hierarchically using a novel method, referred to as joint sparse representation and distribution adaptation. To obtain robust and informative feature representations, we first incorporate a sample-affinity matrix into the marginalized Stacked Denoising Autoencoder to obtain shared features that are then combined with the private features. In order to make the feature representations of videos across views transferable, we then learn a transferable dictionary pair simultaneously from pairs of videos taken at different views to encourage each action video across views to have the same sparse representation. However, the distribution difference across views still exists because a unified subspace, where the sparse representations of one action across views are the same, may not exist when the view difference is large. Therefore, we propose a novel unsupervised distribution adaptation method that learns a set of projections that project the source and target views data into respective low-dimensional subspaces, where the marginal and conditional distribution differences are reduced simultaneously. Therefore, the finally learned feature representation is view-invariant and robust for substantial distribution difference across views even though the view difference is large. Experimental results on four multi-view datasets show that our approach outperforms the state-of-the-art approaches. Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Joint Deep and Depth for Object-Level Segmentation and Stereo Tracking in CrowdsabstractTracking multiple people in crowds is a fundamental and essential task in the multimedia field. It is often hindered by difficulties, such as dynamic occlusion between objects, cluttered background, and abrupt illumination changes. To respond to this need, in this paper, we combine deep and depth to build a stereo tracking system for crowds. The core of the system is the fusion of the advantages of deep learning and depth information, which is exploited to achieve object segmentation and improve the multiobject tracking performance in severe occlusion. More specifically, first, to obtain more accurate detection observations in the tracking system, we present a novel object-level segmentation method. This method combines the effective detection results of deep learning with depth information to obtain precise object segmentation results. Then, we integrate the segmentation results and three-dimensional (3-D) information to extract 2-D and 3-D characteristics to represent the target, and design three similarity models to realize a stereo tracking method through data association in crowds. Finally, we build a diverse stereo dataset including various challenging indoor and outdoor scenes. The comprehensive experiments verify the effective and robust tracking performance of our system in various scenarios, and the system has rich output results including segmentation results, target distance, and tracking results. Moreover, the qualitative and quantitative comparison results show that the proposed algorithm not only has good object segmentation performance but also improves the tracking performance of completely and partially occluded objects, which is superior to the tested state-of-the-art tracking approaches. Jing Li 0010, Lisong Wei, Fangbing Zhang, Tao Yang 0006, Zhaoyang Lu |
IEEE Trans. Multim. | 1 |
| 2018 | Global Temporal Representation Based CNNs for Infrared Action RecognitionabstractInfrared human action recognition has many advantages, i.e., it is insensitive to illumination change, appearance variability, and shadows. Existing methods for infrared action recognition are either based on spatial or local temporal information, however, the global temporal information, which can better describe the movements of body parts across the whole video, is not considered. In this letter, we propose a novel global temporal representation named optical-flow stacked difference image (OFSDI) and extract robust and discriminative feature from the infrared action data by considering the local, global, and spatial temporal information together. Due to the small size of the infrared action dataset, we first apply convolutional neural networks on local, spatial, and global temporal stream respectively to obtain efficient convolutional feature maps from the raw data rather than train a classifier directly. Then these convolutional feature maps are aggregated into effective descriptors named three-stream trajectory-pooled deep-convolutional descriptors by trajectory-constrained pooling. Furthermore, we improve the robustness of these features by using the locality-constrained linear coding (LLC) method. With these features, a linear support vector machine (SVM) is adopted to classify the action data in our scheme. We conduct the experiments on infrared action recognition datasets InfAR and NTU RGB+D. The experimental results show that the proposed approach outperforms the representative state-of-the-art handcrafted features and deep learning features based methods for the infrared action recognition. Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006 |
IEEE Signal Process. Lett. | 3 |
| 2017 | Cube surface modeling for human detection in crowdabstractHuman detection in dense crowds poses to be a demanding task owing to complex background and serious occlusion. In this paper, we propose a novel real-time and reliable human detection system. We solve the human detection problem by presenting a novel cube surface model captured by a binocular stereo vision camera. We first propose a cube surface model to estimate the 3D background cubes in the surveillance area. We then develop a shadow-free strategy for cube surface model updating. Thereafter, we present a shadow weighted clustering method to efficiently search for human as well as remove false alarms. Ultimately, we have developed a highly robust human detection system, and we carefully evaluate our system in many real challenge indoor and outdoor scenes. Expensive experiments demonstrate our system achieves real-time performance, higher detection rate and lower face alarms in comparison with state-of-the-art human detection methods. Jing Li 0010, Fangbing Zhang, Lisong Wei, Tao Yang 0006, Zhongzhen Li |
ICME | 1 |
| 2016 | An improved Fisher discriminant vector employing updated between-scatter matrix
Zhaoyang Lu, Jing Li 0010, Jungong Han |
Neurocomputing | 3 |
| 2016 | Kinect based real-time synthetic aperture imaging through occlusion
Tao Yang 0006, Wenguang Ma, Sibing Wang, Jing Li 0010, Jingyi Yu 0001, Yanning Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2014 | All-In-Focus Synthetic Aperture Imaging
Tao Yang 0006, Yanning Zhang 0001, Jingyi Yu 0001, Jing Li 0010, Wenguang Ma, Xiaomin Tong, Lingyan Ran |
ECCV (6) | 4 |
| 2014 | A subset method for improving Linear Discriminant Analysis
Zhaoyang Lu, Jing Li 0010, Yamei Xu, Jungong Han |
Neurocomputing | 3 |
| 2010 | Fabric defect classification using radial basis function network
Zhaoyang Lu, Jing Li 0010 |
Pattern Recognit. Lett. | 3 |
| 2009 | Fabric Defect Detection and Classification Using Gabor Filters and Gaussian Mixture Model
Zhaoyang Lu, Jing Li 0010 |
ACCV (2) | 3 |
| 2005 | Real-Time Multiple Objects Tracking with Occlusion Handling in Dynamic ScenesabstractThis work presents a real-time system for multiple objects tracking in dynamic scenes. A unique characteristic of the system is its ability to cope with long-duration and complete occlusion without a prior knowledge about the shape or motion of objects. The system produces good segment and tracking results at a frame rate of 15-20 fps for image size of 320 /spl times/ 240, as demonstrated by extensive experiments performed using video sequences under different conditions indoor and outdoor with long-duration and complete occlusions in changing background. Tao Yang 0006, Stan Z. Li, Quan Pan 0001, Jing Li 0010 |
CVPR (1) | 4 |
| 2004 | Multiple layer based background maintenance in complex environmentabstractA fast and efficient multiple layer background maintenance model is built to conserve the original and the current background separately. Fusing the properties of object motion in image pixels and the changes between the input video and the multiple background layers, this method could handle various sources of scene changes, including ghosts, abandon objects and illumination changes. An intelligent video surveillance system is developed to test the performance of the algorithm. Experiments are performed using long video sequences under different conditions indoor and outdoor. The results show that the proposed algorithm is effective and efficient in real-time and accurate background maintenance in complex environment. Tao Yang 0006, Quan Pan 0001, Stan Z. Li, Jing Li 0010 |
ICIG | 4 |