Jiankun Li

dblp:40/3011 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Guarder: A Stable and Lightweight Reconfigurable RRAM-based PIM Accelerator for DNN IP Protection
abstract
Deploying deep neural networks (DNNs) on conventional digital edge devices faces significant challenges due to high energy consumption. A promising solution is the processing-inmemory (PIM) architecture with resistive random-access memory (RRAM), but RRAM-based systems suffer from imprecise weights due to programming stochasticity and cannot effectively utilize conventional weight encryption/decryption intellectual property (IP) protection schemes. To address these issues, we propose a novel software-hardware co-design Guarder. On the hardware side, we introduce 3T2R cells to achieve reliable multiply-accumulate (MAC) operations and use reconfigurable inverter operating voltages to encode keys for encrypting DNNs on RRAM. On the software side, we implement a contrastive training method that ensures high model accuracy on authorized chips while degrading performance on unauthorized ones. This approach protects DNN IP with minimal hardware overhead while significantly mitigating the effects of RRAM programming stochasticity. Extensive experiments on tasks such as image classification (using MLP, ResNet, and ViT), segmentation (using SegFormer), and image generation (using DiT) validate the effectiveness of our method. The proposed contrastive training ensures negligible performance degradation on authorized chips, while performance on unauthorized chips drops to random guessing or generation. Compared to traditional RRAM accelerators, the 3T2R-based accelerator achieves a $1.41 \times$ reduction in area overhead and a $2.28 \times$ reduction in energy consumption.
Ning Lin, Yi Li 0049, Jiankun Li, Jichang Yang, Yangu He, Yukui Luo, Dashan Shang, Xiaoming Chen 0003, Xiaojuan Qi 0001
DAC3
2025 Extended Kalman filter-based maximum likelihood estimation for dynamic soft tissue characterisation
Xinhe Zhu, Jiankun Li, Yongmin Zhong, Chengfan Gu, Kup-Sze Choi
Eng. Appl. Artif. Intell.2
2025 CLIP-GS: CLIP-Informed Gaussian Splatting for View-Consistent 3D Indoor Semantic Understanding
abstract
Exploiting 3D Gaussian Splatting (3DGS) with Contrastive Language-Image Pre-Training (CLIP) models for open-vocabulary 3D semantic understanding of indoor scenes has emerged as an attractive research focus. Existing methods typically attach high-dimensional CLIP semantic embeddings to 3D Gaussians and leverage view-inconsistent 2D CLIP semantics as Gaussian supervision, resulting in efficiency bottlenecks and deficient 3D semantic consistency. To address these challenges, we present CLIP-GS, efficiently achieving a coherent semantic understanding of 3D indoor scenes via the proposed Semantic Attribute Compactness (SAC) and 3D Coherent Regularization (3DCR). SAC approach exploits the naturally unified semantics within objects to learn compact, yet effective, semantic Gaussian representations, enabling highly efficient rendering (>100 FPS). 3DCR enforces semantic consistency in 2D and 3D domains: In 2D, 3DCR utilizes refined view-consistent semantic outcomes derived from 3DGS to establish cross-view coherence constraints; in 3D, 3DCR encourages features similar among 3D Gaussian primitives associated with the same object, leading to more precise and coherent segmentation results. Extensive experimental results demonstrate that our method remarkably suppresses existing state-of-the-art approaches, achieving mIoU improvements of 21.20% and 13.05% on ScanNet and Replica datasets, respectively, while maintaining real-time rendering speed. Furthermore, our approach exhibits superior performance even with sparse input data, substantiating its robustness.
Guibiao Liao, Jiankun Li, Zhenyu Bao, Xiaoqing Ye, Qing Li 0029, Kanglin Liu
ACM Trans. Multim. Comput. Commun. Appl.2
2024 VLM2Scene: Self-Supervised Image-Text-LiDAR Learning with Foundation Models for Autonomous Driving Scene Understanding
abstract
Vision and language foundation models (VLMs) have showcased impressive capabilities in 2D scene understanding. However, their latent potential in elevating the understanding of 3D autonomous driving scenes remains untapped. In this paper, we propose VLM2Scene, which exploits the potential of VLMs to enhance 3D self-supervised representation learning through our proposed image-text-LiDAR contrastive learning strategy. Specifically, in the realm of autonomous driving scenes, the inherent sparsity of LiDAR point clouds poses a notable challenge for point-level contrastive learning methods. This method often grapples with limitations tied to a restricted receptive field and the presence of noisy points. To tackle this challenge, our approach emphasizes region-level learning, leveraging regional masks without semantics derived from the vision foundation model. This approach capitalizes on valuable contextual information to enhance the learning of point cloud representations. First, we introduce Region Caption Prompts to generate fine-grained language descriptions for the corresponding regions, utilizing the language foundation model. These region prompts then facilitate the establishment of positive and negative text-point pairs within the contrastive loss framework. Second, we propose a Region Semantic Concordance Regularization, which involves a semantic-filtered region learning and a region semantic assignment strategy. The former aims to filter the false negative samples based on the semantic distance, and the latter mitigates potential inaccuracies in pixel semantics, thereby enhancing overall semantic consistency. Extensive experiments on representative autonomous driving datasets demonstrate that our self-supervised method significantly outperforms other counterparts. Codes are available at https://github.com/gbliao/VLM2Scene.
Guibiao Liao, Jiankun Li, Xiaoqing Ye
AAAI2
2024 LSMR: Synergy Randomness in Liquid State Machine and RRAM-based Analog-digital Accelerator
abstract
Bio-inspired event sensors are gaining popularity at the edge, such as in robots and wearable electronics. This trend necessitates learning vast amounts of sensory data on the edge, often in few-shot or even zero-shot scenarios, posing challenges in both software and hardware. This paper presents a novel software-hardware co-design to address these issues. Software-wise, we develop an SNN-ANN model, where the SNN encoder is a liquid state machine (LSM) that naturally processes events and significantly reduces learning complexity at the edge due to fixed random weights. The lightweight trainable ANN projection heads are optimized through contrastive learning, enabling zero-shot learning of multimodal events. Hardware-wise, we propose a hybrid analog (RRAM)-digital (CMOS) accelerator - LSMR. The analog in-memory computing core physically implements the LSM by leveraging RRAM stochasticity to generate fixed random weights. The digital core utilizes innovative reconfigurable systolic arrays to accelerate the contrastive learning of ANN projection heads. Extensive experimental outcomes from six neuromorphic datasets, encompassing visual, tactile, and auditory modalities, demonstrate that LSMR considerably improves energy efficiency by a range of 1.65× to 23.70×, in comparison to state-of-the-art edge devices. Simultaneously, it reduces training complexity by a range of 152.83× to 20,587.77× across various edge learning tasks.
Ning Lin, Songqi Wang, Xinyuan Zhang 0008, Shaocong Wang 0001, Yangu He, Woyu Zhang, Bo Wang 0153, Jiankun Li, Mingzi Li, Binbin Cui, Yi Li 0049, Jia Chen 0032, Chunwei Xia, Xiaoming Chen 0003, Dashan Shang
ICCAD8
2023 Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching
abstract
Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not work uniformly well across various datasets, greatly limiting their real-world applicability. To tackle this issue, this paper proposes a new perspective to dynamically calculate correlation for robust stereo matching. A novel Uncertainty Guided Adaptive Correlation (UGAC) module is introduced to robustly adapt the same model for different scenarios. Specifically, a variance-based uncertainty estimation is employed to adaptively adjust the sampling area during warping operation. Additionally, we improve the traditional non-parametric warping with learnable parameters, such that the position-specific weights can be learned. We show that by empowering the recurrent network with the UGAC module, stereo matching can be exploited more robustly and effectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance over the ETH3D, KITTI, and Middlebury datasets when employing the same fixed model over these datasets without any retraining procedure. To target real-time applications, we further design a lightweight model based on UGAC, which also outperforms other methods over KITTI benchmarks with only 0.6 M parameters.
Junpeng Jing, Jiankun Li, Pengfei Xiong, Jiangyu Liu, Shuaicheng Liu, Xin Deng 0002, Mai Xu, Lai Jiang 0004, Leonid Sigal
ICCV2
2022 Practical Stereo Matching via Cascaded Recurrent Network with Adaptive Correlation
abstract
With the advent of convolutional neural networks, stereo matching algorithms have recently gained tremendous progress. However, it remains a great challenge to accurately extract disparities from real-world image pairs taken by consumer-level devices like smartphones, due to practical complicating factors such as thin structures, non-ideal rectification, camera module inconsistencies and various hard-case scenes. In this paper, we propose a set of innovative designs to tackle the problem of practical stereo matching: 1) to better recover fine depth details, we design a hierarchical network with recurrent refinement to update disparities in a coarse-to-fine manner, as well as a stacked cascaded architecture for inference; 2) we propose an adaptive group correlation layer to mitigate the impact of erroneous rectification; 3) we introduce a new synthetic dataset with special attention to difficult cases for better generalizing to real-world scenes. Our results not only rank 1ston both Middlebury and ETH3D benchmarks, outperforming existing state-of-the-art methods by a notable margin, but also exhibit high-quality details for real-life photos, which clearly demonstrates the efficacy of our contributions.
Jiankun Li, Peisen Wang, Pengfei Xiong, Ziwei Yan, Jiangyu Liu, Haoqiang Fan, Shuaicheng Liu
CVPR1
2022 DIP: Deep Inverse Patchmatch for High-Resolution Optical Flow
abstract
Recently, the dense correlation volume method achieves state-of-the-art performance in optical flow. However, the correlation volume computation requires a lot of memory, which makes prediction difficult on high-resolution images. In this paper, we propose a novel Patchmatch-based framework to work on high-resolution optical flow estimation. Specifically, we introduce the first end-to-end Patchmatch based deep learning optical flow. It can get high-precision results with lower memory benefiting from propagation and local search of Patchmatch. Furthermore, a new inverse propagation is proposed to decouple the complex operations of propagation, which can significantly reduce calculations in multiple iterations. At the time of submission, our method ranks 1st on all the metrics on the popular KITTI2015 [28] benchmark, and ranks 2ndon EPE on the Sintel [7] clean benchmark among published optical flow methods. Experiment shows our method has a strong cross-dataset generalization ability that the F1-all achieves 13.73%, reducing 21% from the best published result 17.4% on KITTI2015. What's more, our method shows a good details preserving result on the high-resolution dataset DAVIS [1] and consumes 2× less memory than RAFT [36]. Code will be available at github.com/zihuarheng/DIP
Zihua Zheng, Ni Nie, Pengfei Xiong, Jiangyu Liu, Jiankun Li
CVPR7
2022 MSJAD: Multi-Source Joint Anomaly Detection of Web Application Access
abstract
Fixed broadband internet service can provide a stable broadband network of up to 100 megabits or even gigabits and users at home can use fixed broadband service for all kinds of internet surfing, including website and application access, watching videos, playing games, etc. Traditional maintenance for fixed broadband networks primarily uses human manual meth-ods, supplemented by some low-level semi-automation operations. Since the long processes with numerous network elements in the fixed broadband network, it is difficult for traditional operation and maintenance to support effectively with high quality. When abnormalities occur, it is quite manpower cost and time cost to monitor and locate faults. Therefore, to improve the autonomous capability of the fixed broadband network, intelligent operation and maintenance methods are necessary. First of all, a brand-new data pre-process method is proposed to detect anomalies and problems of slow access by selecting web services access commonly visited by users. Secondly, as the fixed broadband network is a multi-level and complex structure with only a small amount of anomaly sample data, we propose a multi-source joint anomaly detection model called MSJAD model on multi-dimensional features data. The model validation results on real datasets from the real fixed broadband network are state-of-the-art. The accuracy rate reaches 98 % and the recall is over 99 %. We have already begun to deploy the model on the real fixed broadband network and have achieved good feedback.
Xinxin Chen, Chengsen Wang, Guosong Lv, Jiankun Li, Dewei Chen, Lianyuan Li
MSN6
2000 Postprocessing of Compressed 3D Graphic Data
Ka Man Cheang, Wenlong Dong, Jiankun Li, C.-C. Jay Kuo
J. Vis. Commun. Image Represent.3
1999 Reversible Variable Length Codes (RVLC) for Robust Coding of 3D Topological Mesh Data
abstract
Summary form only given. In order to limit error propagation, we divide the topological data of the entire mesh into several segments. Each segment is identified by its synchronization word and header. Due to the use of the arithmetic coder, data of a whole segment would often become useless in the presence of even a single bit error. Furthermore, several adjacent segments may be corrupted simultaneously at high bit error rates (BER). As a result, a lot of data would be required to be retransmitted in the presence of errors. Retransmitted data may also in turn get corrupted in high BER conditions. This would result in a considerable loss of coding efficiency and increased delay. We propose the use of reversible variable length codes (RVLC) to solve this problem. RVLC not only prevents error propagation in one segment but also efficiently detects the distorted portion of the bitstream due to their capability of two-way decoding. This would allow the recovery of a large portion of data from a corrupted segment. The amount of retransmitted data can thus be drastically reduced. RVLC can be matched to various sources with different probability distributions by adjusting their suffix length, and have been found suitable for image and video coding. However, the application of RVLC to robust 3D mesh coding has not yet been studied. Our study of the suitability of RVLC for the topological data is presented in this research. Experiments have been carried to prove the efficiency of the proposed robust 3D graphic coding algorithm. To design an efficient pre-defined code table, a large set of 300 MPEG-4 selected 3D models have been used in our experiments. The use of predefined code tables would result in a significantly reduced computational complexity.
Zhidong Yan, Sunil Kumar 0001, Jiankun Li, C.-C. Jay Kuo
Data Compression Conference3
1999 Refinement of 3D Meshes by Selective Subdivision
abstract
An adaptive subdivision method is proposed in this work for automatic post-processing of a 3D graphic model of coarse resolution. The method is an improved version of the Modified Butterfly Scheme (MBS) developed by Zorin et al. The main contribution of this work is to exploit the local smoothness information of a surface for adaptive refinement of a coarse 3D graphic model. With this approach, we can avoid unnecessary subdivision in relative smooth legions. The new algorithm not only reduces the computational complexity but also reduces the storage space. It is demonstrated via experiments that a more visual-pleasing representation can be obtained.
Wenlong Dong, Jiankun Li, C.-C. Jay Kuo
ICIP (4)2
1999 Robust Coding of 3D Graphic Models Using Mesh Segmentation and Data Partitioning
abstract
Current coding techniques for 3D graphic models focus more on coding efficiency, which makes them extremely sensitive to channel errors due to the irregular nature of the mesh structure. In this paper, we propose a robust mesh partitioning scheme which can reduce the error propagation length and allow a piecewise reconstruction of the original mesh while maintaining a high compression ratio. Our method first segments an arbitrary 3D mesh into a set of independent connected components. Each component is further segmented into several independent 3D pieces based on the error rate. Joint boundary information is efficiently coded which will be used to stick segmented pieces together. All resulting pieces are grouped by their sizes to keep each group as a basic unit which has a near equal length, which is determined by the error rate. The data partitioning technique is finally used to keep track of each group of pieces. Errors are thus limited to each group, and can be corrected easily. The proposed technique can also be used in graphic editing, partial texture mapping and other applications.
Zhidong Yan, Sunil Kumar 0001, Jiankun Li, C.-C. Jay Kuo
ICIP (4)3
1998 A Dual Graph Approach to 3D Triangular Mesh Compression
abstract
The triangular mesh provides one of the most popular representations for 3D graphic models. A typical triangular mesh consists of two different types of data: topological data which specify the connectivity of the mesh and geometrical data which describe information associated with each individual vertex or triangle. We propose a new compression scheme which encode topological data by using the dual graph of the original mesh. It is found that the dual graph can be represented as a degraded binary tree. Furthermore, geometrical data can be coded progressively with local prediction and embedded entropy coding. Experimental results show that an acceptable quality level can be reached at a compression ratio of 60 to 1 for general test models.
Jiankun Li, C.-C. Jay Kuo
ICIP (2)1
1998 Progressive coding of 3-D graphic models
abstract
Based on state-of-the-art graphic-simplification techniques and progressive image-coding schemes, we propose a new hierarchical three-dimensional graphic-compression scheme in this research. This scheme progressively compresses an arbitrary polygonal mesh into a single bitstream. Along the encoding process, every output bit contributes to the reduction of coding distortion, and the contribution of bits decreases according to their order of position in the bitstream. At the receiver end, the decoder can stop at any point while giving a reconstruction of the original model with the best rate-distortion tradeoff. A series of models of continuous varying resolution can thus be constructed from the single bitstream. This property, which is referred to as the embedding property since the coding of a coarser model is embedded in the coding of a finer model, can be widely used in robust error control, progressive transmission and display, level-of-detail control, etc. It is demonstrated by experiments that an acceptable quality level can be achieved at a compression ratio of 20 to 1 for several test graphic models.
Jiankun Li, C.-C. Jay Kuo
Proc. IEEE1
1997 Embedded Coding of 3D Graphic Models
abstract
A progressive compression method which encodes a 3D graphic models into an embedded bit stream is investigated. The coder first encodes the coarsest resolution of the model, and then includes the information of finer details gradually. A rate-distortion model is used in the integration of different types of bit streams into a single embedded bit stream to optimize the overall coding performance. This embedding property can be applied to progressive transmission, multi-resolution editing and level-of-detail control. Numerical experiments are provided to demonstrate the excellent rate-distortion performance of the proposed method.
Jiankun Li, C.-C. Jay Kuo
ICIP (1)1
1997 Layered DCT still image compression
abstract
Motivated by Shapiro's (1993) embedded zerotree wavelet (EZW) coding and Taubman and Zakhor's (1994) layered zero coding (LZC), we propose a layered discrete cosine transform (DCT) image compression scheme, which generates an embedded bit stream for DCT coefficients according to their importance. The new method allows progressive image transmission and simplifies the rate-control problem. In addition to these functionalities, it provides a substantial rate-distortion improvement over the JPEG standard when the bit rates become low. For example, we observe a bit rate reduction with respect to the JPEG Huffman and arithmetic coders by about 60% and 20%, respectively, for a bit rate around 0.1 b/p.
Jiankun Li, Jin Li 0001, C.-C. Jay Kuo
IEEE Trans. Circuits Syst. Video Technol.1
1996 An embedded DCT approach to progressive image compression
abstract
Motivated by Shapiro's (see IEEE Trans. on Signal Processing, vol.41, no.12, p. 3445-62, 1993) embedded zerotree wavelet coding (EZW) and Taubman and Zakhor's (see IEEE Trans. on Image Processing, vol.3, no.3, p.572-88, 1994) layered zero coding (LZC), we propose a layered DCT image compression scheme, which generates an embedded bit stream for DCT coefficients according to their importance. The new method allows progressive image transmission and simplifies the rate-control problem. In addition to these functionalities, it provides a substantial rate-distortion improvement over the JPEG standard when the bit rates become low. For example, we observe a bit rate reduction with respect to the JPEG Huffman and arithmetic coders by about 60% and 20%, respectively, for a bit rate around 0.1 bpp.
Jiankun Li, Jin Li 0001, C.-C. Jay Kuo
ICIP (1)1