EDBT 2026 Demo / reviewers in the wild / expert
Jian Li 0003
dblp:33/5448-3
· DBLP profile ↗
23ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0002-5749-2734ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Task-Preferred Inference Routes for Gradient De-Conflict in Multi-Output DNNsabstractMulti-output deep neural networks (MONs) contain multiple output branches of various tasks, and these tasks typically share partial network filters, resulting in entangled inference routes between different tasks within the networks. Due to the divergent optimization objectives, the task gradients during training usually interfere with each other along the shared routes, which decreases the overall model performance. To address this issue, we propose a novel gradient de-conflict algorithm named DR-MGF (Dynamic Routes and Meta-weighted Gradient Fusion). Different from existing de-conflict methods, DR-MGF achieves gradient de-conflict in MONs by learning task-preferred inference routes. The proposed method is motivated by our experimental findings that the shared filters are not equally important for different tasks. By designing learnable task-specific importance variables, DR-MGF evaluates the importance of filters for different tasks. Through making the dominance of tasks over filters proportional to the task-specific importance of filters, DR-MGF can effectively reduce inter-task interference. These task-specific importance variables ultimately determine task-preferred inference routes at the end of training iterations. Extensive experimental results on CIFAR, ImageNet, and NYUv2 demonstrate that DR-MGF outperforms existing de-conflict methods. Furthermore, DR-MGF can be extended to general MONs without modifying the overall network structures. Xiaochang Hu, Xin Xu 0001, Jian Li 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Infrared-Guided Reconstruction of Visible Images Occluded by Localized Smoke
Yushi Quan, Kexin Fei, Jian Li 0003 |
IEEE Signal Process. Lett. | 3 |
| 2026 | Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object DetectionabstractRecent studies in remote sensing object detection have made excellent progress and shown promising performance. However, most current detectors only explore rotation-invariant feature extraction but disregard the valuable spatial and semantic prior knowledge in remote sensing images (RSIs), which limits the detection performance when encountering blurred or heavy occluded objects. To address this issue, we propose a mask-reconstruction relation learning (MRRL) framework to learn such prior knowledge among objects and a consistency-reasoning transformer over relation proposals (CTRP) to recognize objects with limited visual features via consistency reasoning. Specifically, MRRL framework applies random mask to some objects in the training dataset and performs masked objects reconstruction to guide the network to learn the distribution consistency of objects. CTRP is the core component of the MRRL framework, which models the interaction between spatial and semantic priors, and uses easy detected objects to reason hard detected objects. The trained CTRP can be integrated into the existing detector to improve the ability of object detection with limited visual features in RSIs. Extensive experiments on widely-used datasets for two distinct tasks, namely remote sensing object detection task and occluded object detection task, demonstrate the effectiveness of the proposed method. Source code is available at https://github.com/sunpeng96/CTRP_mmrotate. Yongbin Zheng, Wanying Xu, Jian Li 0003, Jiansong Yang |
IEEE Trans. Image Process. | 4 |
| 2025 | ETA: Learning Optical Flow with Efficient Temporal AttentionabstractConsidering the potential of using multi-frame information to solve the occlusion problem, we introduce a novel idea of multi-frame information integration, which uses the attention mechanism to fuse the temporal information from the previous frame. The idea can effectively improve the estimation accuracy in occluded regions and optimize the inference speed under multi-frame settings. Meanwhile, we suggest the concept of attention confidence to provide an explicit value criterion for the model to utilize useful attention information more efficiently. Furthermore, we propose an Efficient Temporal Attention network (ETA), which achieves promising results on Sintel and KITTI benchmarks, especially with a 9.4% error reduction compared to the baseline method GMA on Sintel (test) Clean. Bo Wang 0144, Zhenping Sun, Yang Yu 0014, Li Liu 0002, Jian Li 0003, Dewen Hu |
IROS | 5 |
| 2025 | SceneTracker: Long-Term Scene Flow Estimation NetworkabstractConsidering that scene flow estimation has the capability of the spatial domain to focus but lacks the coherence of the temporal domain, this study proposes long-term scene flow estimation (LSFE), a comprehensive task that can simultaneously capture the fine-grained and long-term 3D motion in an online manner. We introduce SceneTracker, the first LSFE network that adopts an iterative approach to approximate the optimal 3D trajectory. The network dynamically and simultaneously indexes and constructs appearance correlation and depth residual features. Transformers are then employed to explore and utilize long-range connections within and between trajectories. With detailed experiments, SceneTracker shows superior capabilities in addressing 3D spatial occlusion and depth noise interference, highly tailored to the needs of the LSFE task. We build a real-world evaluation dataset, LSFDriving, for the LSFE field and use it in experiments to further demonstrate the advantage of SceneTracker in generalization abilities. Bo Wang 0144, Jian Li 0003, Yang Yu 0014, Li Liu 0002, Zhenping Sun, Dewen Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Bilateral Propagation Network for Depth CompletionabstractDepth completion aims to derive a dense depth map from sparse depth measurements with a synchronized color image. Current state-of-the-art (SOTA) methods are predominantly propagation-based, which work as an iterative refinement on the initial estimated dense depth. However, the initial depth estimations mostly result from direct applications of convolutional layers on the sparse depth map. In this paper, we present a Bilateral Propagation Network (BP-Net), that propagates depth at the earliest stage to avoid directly convolving on sparse data. Specifically, our approach propagates the target depth from nearby depth measurements via a non-linear model, whose coefficients are generated through a multi-layer perceptron conditioned on both radiometric difference and spatial distance. By integrating bilateral propagation with multi-modal fusion and depth refinement in a multi-scale framework, our BP-Net demonstrates outstanding performance on both indoor and outdoor scenes. It achieves SOTA on the NYUv2 dataset and ranks 1st on the KITTI depth completion benchmark at the time of submission. Experimental results not only show the effectiveness of bilateral propagation but also emphasize the significance of early-stage propagation in contrast to the refinement stage. Our code and trained models will be available on the project page. Jie Tang 0015, Fei-Peng Tian, Boshi An, Jian Li 0003, Ping Tan 0002 |
CVPR | 4 |
| 2024 | Efficient-PIP: Large-scale Pixel-level Aligned Image Pair Generation for Cross-time Infrared-RGB TranslationabstractGenerative models are gaining momentum in both academic and industrial applications driven by the availability of large-scale datasets, especially in tasks involving Image-to-Image Translation. Meanwhile, poor human perception of nighttime environment has led to a demand for translation from night-vision infrared to day-vision RGB images. However, collecting such cross-modal training data at the same time is impossible due to the thermal imaging properties of infrared cameras, the challenge lies in constructing image pairs during the day and at night respectively, where the requirement for data alignment poses significant difficulties. In this paper, we propose a Pixel-level aligned Image Pair generation framework PIP to explore efficient colorization of high-resolution infrared images. Specifically, we first construct a 3D high-precision point cloud map for the purpose of establishing the correlation between day and night scenes. Corresponding point clouds of modal images are collected simultaneously during data acquisition to obtain image sensor poses by Global Matching with the map, which allows us to calculate the transformation relationship from infrared to RGB image coordinate systems based on the sensor parameters and depth information of the map. Leveraging the relationship, the pixel values of RGB image is projected onto the infrared image followed by optimization as the colored image. Accordingly, we present a dataset NUDT-PIP, the first of its kind containing large-scale pixel-level aligned cross-time infrared-RGB image pairs of complicated real road scenes. Experimental results demonstrate the reliability and strong applicability of our dataset in Image-to-Image Translation. Our code will be released at https://github.com/wjjjjyourFA/NUDT-PIP. Jian Li 0003, Kexin Fei, Bokai Liu, Zongtan Zhou, Yongbin Zheng, Zhenping Sun |
IROS | 1 |
| 2024 | SplatFlow: Learning Multi-frame Optical Flow via Splatting
Bo Wang 0144, Jian Li 0003, Yang Yu 0014, Zhenping Sun, Li Liu 0002, Dewen Hu |
Int. J. Comput. Vis. | 3 |
| 2022 | Meta-GF: Training Dynamic-Depth Neural Networks Harmoniously
Jian Li 0003, Xin Xu 0001 |
ECCV (11) | 2 |
| 2022 | Cross-modal Fusion-based Prior Correction for Road Detection in Off-road EnvironmentsabstractRoad detection plays a fundamental role in the visual navigation system of autonomous vehicles. However, it's still challenging to achieve robust road detection in off-road scenarios due to their complicated road appearances and ambiguous road structures. Therefore, existing image-based road detection approaches usually fail to extract the right routes due to the lack of the effective fusion of the image and prior reference paths(road guidances generated via map annotations and GPS localization). Besides, the reference paths are not always reliable because of GPS localization errors and mapping errors. To achieve robust road detection in off-road scenarios, we propose a prior-correction-based road detection network named PR-ROAD via fusing the cross-model information provided by both the reference path and the input image. These two heterogeneous data, prior and image, are deeply fused by a cross-attention module and formulate contextual inter-dependencies. We conduct experiments in our collected rural, off-road and urban datasets. The experimental results demonstrate the effectiveness of the proposed method both on unstructured and structured roads. Yuru Wang, Jian Li 0003, Meiping Shi |
IROS | 3 |
| 2021 | Learning Guided Convolutional Network for Depth CompletionabstractDense depth perception is critical for autonomous driving and other robotics applications. However, modern LiDAR sensors only provide sparse depth measurement. It is thus necessary to complete the sparse LiDAR data, where a synchronized guidance RGB image is often used to facilitate this completion. Many neural networks have been designed for this task. However, they often naïvely fuse the LiDAR data and RGB image information by performing feature concatenation or element-wise addition. Inspired by the guided image filtering, we design a novel guided network to predict kernel weights from the guidance image. These predicted kernels are then applied to extract the depth image features. In this way, our network generates content-dependent and spatially-variant kernels for multi-modal feature fusion. Dynamically generated spatially-variant kernels could lead to prohibitive GPU memory consumption and computation overhead. We further design a convolution factorization to reduce computation and memory consumption. The GPU memory reduction makes it possible for feature fusion to work in multi-stage scheme. We conduct comprehensive experiments to verify our method on real-world outdoor, indoor and synthetic datasets. Our method produces strong results. It outperforms state-of-the-art methods on the NYUv2 dataset and ranks 1st on the KITTI depth completion benchmark at the time of submission. It also presents strong generalization capability under different 3D point densities, various lighting and weather conditions as well as cross-dataset evaluations. The code will be released for reproduction. Jie Tang 0015, Fei-Peng Tian, Wei Feng 0005, Jian Li 0003, Ping Tan 0002 |
IEEE Trans. Image Process. | 4 |
| 2019 | Tracking by Animation: Unsupervised Learning of Multi-Object Attentive TrackersabstractOnline Multi-Object Tracking (MOT) from videos is a challenging computer vision task which has been extensively studied for decades. Most of the existing MOT algorithms are based on the Tracking-by-Detection (TBD) paradigm combined with popular machine learning approaches which largely reduce the human effort to tune algorithm parameters. However, the commonly used supervised learning approaches require the labeled data (e.g., bounding boxes), which is expensive for videos. Also, the TBD framework is usually suboptimal since it is not end-to-end, i.e., it considers the task as detection and tracking, but not jointly. To achieve both label-free and end-to-end learning of MOT, we propose a Tracking-by-Animation framework, where a differentiable neural model first tracks objects from input frames and then animates these objects into reconstructed frames. Learning is then driven by the reconstruction error through backpropagation. We further propose a Reprioritized Attentive Tracking to improve the robustness of data association. Experiments conducted on both synthetic and real video datasets show the potential of the proposed model. Our project page is publicly available at: https://github.com/zhen-he/tracking-by-animation Jian Li 0003, Daxue Liu, Hangen He, David Barber |
CVPR | 2 |
| 2018 | Generalized Haar Filter-Based Object Detection for Car Sharing ServicesabstractObject detection is important in car sharing services. Accuracy, efficiency, and low memory consumption are desirable for object detection in car sharing services. This paper presents a network system that satisfies all these requirements. Our approach first divides the object detection task into multiple simpler local regression tasks. Then, we propose the generalized Haar filter-based convolutional neural network to reduce the consumption of memory and computing resource. To achieve real-time performance, we introduce a sparse window generation strategy to reduce the number of input image patches without sacrificing accuracy. We perform experiments on both vehicle and pedestrian data sets. Experimental results demonstrate that our approach can accurately detect objects under challenging conditions. Note to Practitioners-Object detection is an important part of intelligent vehicle technologies, which play an important role in car sharing services. Object detection provides metadata for collision avoidance, self-driving systems, and driver-assistance systems, which can result in better safety and consumer experiences in car sharing services. Although deep learning has achieved an excellent performance in object detection, they consume a large amount of storage and computing resource, which makes them difficult to be deployed for car sharing services. This paper suggests a novel approach which is based on the generalized Haar filter and the local regression strategy. Our approach is accurate, efficient, and light. The experimental results verify the effectiveness of the proposed approach in car sharing services. Keyu Lu, Jian Li 0003, Li Zhou 0002, Xiping Hu, Xiangjing An, Hangen He |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2017 | Efficient deep network for vision-based object detection in robotic applications
Keyu Lu, Xiangjing An, Jian Li 0003, Hangen He |
Neurocomputing | 3 |
| 2014 | Hierarchical image representation via multi-level sparse codingabstractThis paper presents a hierarchical model for robust image representation. We first introduce multi-level sparse coding algorithm and normalized max pooling strategy which are designed to obtain meaningful sparse codes and robust pooled codes, respectively. With the sparse codes and pooled codes, a hierarchical architecture is built and more robust features are extracted at the second layer. The proposed method has been evaluated on two widely used datasets: Caltech-101 and Caltech-256, and experimental results demonstrate that the proposed method is both effective and robust in image representation compared with the state-of-the-art. Keyu Lu, Jian Li 0003, Xiangjing An, Hangen He |
ICIP | 2 |
| 2014 | A hierarchical approach for road detectionabstractRoad detection is a crucial problem for autonomous navigation system (ANS) and advance driver-assistance system (ADAS). In this paper, we propose a hierarchical road detection method for robust road detection in challenging scenarios. Given an on-board road image, we first train a Gaussian mixture model (GMM) to obtain road probability density map (RPDM), and next oversegment the image into superpixels. Based on RPDM and superpixels, initial seeds are selected in an unsupervised way, and the seed superpixels iteratively try to occupy their neighbors according to GrowCut framework, the road segment is obtained after convergency. Finally, we refine the road segment with a conditional random field (CRF), which enforces the shape prior on the road segmentation task. Experiments on two challenging databases demonstrate that the proposed method exhibits high robustness compared with the state-of-the-art. Keyu Lu, Jian Li 0003, Xiangjing An, Hangen He |
ICRA | 2 |
| 2013 | Visual Saliency Based on Scale-Space Analysis in the Frequency DomainabstractWe address the issue of visual saliency from three perspectives. First, we consider saliency detection as a frequency domain analysis problem. Second, we achieve this by employing the concept of nonsaliency. Third, we simultaneously consider the detection of salient regions of different size. The paper proposes a new bottom-up paradigm for detecting visual saliency, characterized by a scale-space analysis of the amplitude spectrum of natural images. We show that the convolution of the image amplitude spectrum with a low-pass Gaussian kernel of an appropriate scale is equivalent to an image saliency detector. The saliency map is obtained by reconstructing the 2D signal using the original phase and the amplitude spectrum, filtered at a scale selected by minimizing saliency map entropy. A Hypercomplex Fourier Transform performs the analysis in the frequency domain. Using available databases, we demonstrate experimentally that the proposed model can predict human fixation data. We also introduce a new image database and use it to show that the saliency detector can highlight both small and large salient regions, as well as inhibit repeated distractors in cluttered images. In addition, we show that it is able to predict salient regions on which people focus their attention. Jian Li 0003, Martin D. Levine, Xiangjing An, Xin Xu 0001, Hangen He |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Saliency Detection Based on Frequency and Spatial Domain AnalysesabstractWe propose a new saliency detection model by combining global information from frequency domain analysis and local information from spatial domain analysis. In the frequency domain analysis, instead of modeling salient regions, we model the nonsalient regions using global information; these so-called repeating patterns that are not distinctive in the scene are suppressed by using spectrum smoothing. In spatial domain analysis, we enhance those regions that are more informative by using a center-surround mechanism similar to that found in the visual cortex. Finally, the outputs from these two channels are combined to produce the saliency map. We demonstrate that the proposed model has the ability to highlight both small and large salient regions in cluttered scenes and to inhibit repeating objects. Experimental results also show that the proposed model outperforms existing algorithms in predicting objects regions where human pay more attention. 1 Jian Li 0003, Martin D. Levine, Xiangjing An, Hangen He |
BMVC | 1 |
| 2011 | Multi-scale and Multi-orientation Local Feature Extraction for Lane Detection Using High-Level InformationabstractTask-specified computer vision systems usually need to detect certain targets at multiple scales of resolution and multiple orientations. For a vision-based lane detection system, it is essential to detect the lane-markings at different scales and orientations. In this paper, we illustrate an efficient local feature extraction algorithm for the lane detection system, which is tuned by the high-level information about the lane-markings. Firstly, we deduced the explicit expression of the scale and orientation for the local feature of the lane markings. Secondly, a filter bank for local feature extraction is designed using the SVD approach for certain orientation and scale. Thirdly, the filter bank is used to tune a special lane-marking detector to expected orientation and scale at different locations of the image. Then, non-maxima suppression is performed along the corresponding direction at that location. Lastly, a hysteresis thresholding is applied to identify the exact feature points. Unlike other works in which the authors try to remove the false local feature points with the help of high-level information, we prefer to introduce the high-level information to the local feature detection stage as early as possible. Experiment results show that the proposed algorithm is very efficient for lane detection especially in very complex road seniors. Xiangjing An, Jian Li 0003, Er-Ke Shang, Hangen He |
ICIG | 2 |
| 2011 | Lane Detection Based on Visual AttentionabstractVision based lane detection is an essential task in both autonomous lane vehicles research and active safety system development. Hitherto, lane detection is, however, still a challenging issue due to the complexity of the real road scenes. In this paper, we consider lane detection as a visual attention problem. With a Bayesian attention framework, we address the issue from three perspectives: first, lane markings are assumed to be salient in the road scenes, which will pop out driven by the low level features combining with a bottom-up attention mechanism, second, the target-related features are designed guided by a top-down attention strategy, third, the location prior of the lane markings is also investigated. The experimental results show that the proposed lane detection approach is efficient and robust in real scenes. Jian Li 0003, Xiangjing An, Hangen He |
ICIG | 1 |
| 2011 | Lane Detection Using Steerable Filters and FPGA-based ImplementationabstractVision-based lane detection is a key component for Driver-Assistance (DA) systems. It is still a challenging task in road scenes with complex shadows. This paper presents a novel local edge detector, using vanishing point position as a high level information to guide the use of steerable filters in lane detection, and its implementation on a Field Programmable Gate Array (FPGA) device. The FPGA technology has the advantages of high-performances for digital image processing and low cost, both of which are the requirements of DA systems. The main contributions of this work are twofold: 1) an edge extraction algorithm for lane detection is proposed, using the estimated vanishing point as high-level information to detect lanes. Firstly, a rough estimation of the vanishing point is used for calculating the expected local edge orientations. Secondly, a steerable filter is tuned to the expected direction for edge response. 2) a framework on FPGA is designed to implement the proposed algorithm. The framework is designed by using multi-engine technology, so it works in parallel for any order of steerable filters. Experiments and comparisons show that the proposed algorithm is very efficient in dealing with the complex shadow conditions, and works in real-time on FPGA device. Er-Ke Shang, Jian Li 0003, Xiangjing An, Hangen He |
ICIG | 2 |
| 2011 | An interactive method for extrinsic parameter calibration of onboard cameraabstractIn vision-based driver assistance applications, the onboard cameras are supposed to be calibrated as simply and accurately as possible. This paper presents an interactive method for extrinsic parameter calibration of onboard camera. In the proposed method, the intrinsic camera parameters are assumed to be known. During the calibration process, we first design a calibration scene to compute the extrinsic camera parameters. There are three lines in the scene, two of which are parallel with the vehicle heading direction and the other one is perpendicular with the first two. Then, using this scene, we derive an interactive and analytical solution to obtain the extrinsic camera parameters, which can adequately compromise the calibration complexity and accuracy. We also extend the method to multiple cameras, in which case the calibration can be performed without absolute positions of the calibration lines. Theoretical and experimental results show the efficiency of the proposed calibration method. Jian Li 0003, Xiangjing An, Hangen He |
Intelligent Vehicles Symposium | 2 |
| 2010 | From edges to linear features: A PCA and graph based approachabstractLinear features such as line segments and contour fragments are important cues for object detection and scene analysis. Least square based and Hough-like approaches are quite popular and powerful. However, least square approaches are sensitive to outliers, and are unable to handle the case where there is more than one underlying line segment; while Hough-like approaches do not work well when extracting fragments which are not very `straight'. The goal of this work is to extract linear features locally in natural scenes using PCA and graph analysis. The whole process has five stages: first, for each local area, we create a graph using edge points as vertexes and distance measurements as links. Then, some edge-point sets which may contain underlying line segments are obtained by finding connected components in the graph; second, PCA is employed to estimate the existence of linear feature in each point set, and to compute the principal axis orientation; third, the relation between any two vertexes in each set is recalculated according to the principal axis orientation, and a new graph which only contains several trees is generated; fourth, one linear feature is detected by finding a best chain in each tree according to some criteria; fifth, all the line segments in the image are refined and organized according to global information. Jian Li 0003, Xiangjing An, Hangen He |
ICASSP | 1 |