Li Zhang 0023

dblp:89/5992-23 · DBLP profile ↗
← Back
50ranked-venue papers
1as first author
3since 2021 · last 2025
0000-0002-9321-3421ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 since 2021Artificial intelligence and machine learning · 25 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
Video understanding and tracking · 50% Image recognition and object detection · 16% 3D vision · 14%
Computer graphics and multimedia
5 papers
Rendering · 39% Geometric modeling and processing · 22% Visual content generation and editing · 19%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
object tracking
3.1112020
Fuzzy Least Squares Support Vector Machine With Adaptive Membership for Object Tracking · IEEE Trans. Multim. 2020
Exploiting the Anisotropy of Correlation Filter Learning for Visual Tracking · Int. J. Comput. Vis. 2019
Exploiting Spatial-Temporal Locality of Tracking via Structured Dictionary Learning · IEEE Trans. Image Process. 2018
Computer vision › Image recognition and object detection
object detection
0.912025
Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans? · CVPR 2025
Computer vision › Image recognition and object detection › object detection
prohibited item detection
0.912025
Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans? · CVPR 2025
Geometric modeling and processing
3d reconstruction
0.812024
CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses · IEEE Trans. Multim. 2024
Geometric modeling and processing
bundle adjustment
0.812024
CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses · IEEE Trans. Multim. 2024
Virtual and augmented reality › tracking
camera pose estimation
0.812024
CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses · IEEE Trans. Multim. 2024
Rendering
neural radiance fields
0.812024
CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses · IEEE Trans. Multim. 2024
Rendering
neural rendering
0.812024
CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses · IEEE Trans. Multim. 2024
Rendering
novel view synthesis
0.812024
CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses · IEEE Trans. Multim. 2024
Visual content generation and editing
style transfer
0.722019
A Closed-Form Solution to Universal Style Transfer · ICCV 2019
Decoder Network over Lightweight Reconstructed Feature for Fast Semantic Style Transfer · ICCV 2017
Computer vision › 3D vision › object pose estimation
object pose tracking
0.412020
Occlusion-Aware Region-Based 3D Pose Tracking of Objects With Temporally Consistent Polar-Based Local Partitioning · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking › object tracking › robust tracking
occlusion-robust tracking
0.412020
Occlusion-Aware Region-Based 3D Pose Tracking of Objects With Temporally Consistent Polar-Based Local Partitioning · IEEE Trans. Image Process. 2020
Computer vision › Segmentation and scene understanding
scene parsing
0.412020
Strip Pooling: Rethinking Spatial Pooling for Scene Parsing · CVPR 2020
Rendering › non-photorealistic rendering
line drawing
0.412020
Learning to Draw Sight Lines · Int. J. Comput. Vis. 2020
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning
0.432018
Robust Visual Tracking via Sparsity-Induced Subspace Learning · IEEE Trans. Image Process. 2015
Visual Tracking via Subspace Learning: A Discriminative Approach · Int. J. Comput. Vis. 2018
Discriminative Low-Rank Tracking · ICCV 2015
Computer vision › Video understanding and tracking › object tracking
3d object tracking
0.412019
A Robust Monocular 3D Object Tracking Method Combining Statistical and Photometric Constraints · Int. J. Comput. Vis. 2019
Visual content generation and editing › style transfer
arbitrary style transfer
0.412019
A Closed-Form Solution to Universal Style Transfer · ICCV 2019
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding › dictionary learning
discriminative dictionary learning
0.312018
Exploiting Spatial-Temporal Locality of Tracking via Structured Dictionary Learning · IEEE Trans. Image Process. 2018
Computer vision › 3D vision › 3d scene understanding
scene completion
0.312018
Efficient Semantic Scene Completion Network with Spatial Group Convolution · ECCV (12) 2018
Computer vision › 3D vision › 3d scene understanding
semantic scene completion
0.312018
Efficient Semantic Scene Completion Network with Spatial Group Convolution · ECCV (12) 2018
Computer vision › 3D vision › 3d scene understanding
room layout estimation
0.312017
Physics Inspired Optimization on Semantic Transfer Features: An Alternative Method for Room Layout Estimation · CVPR 2017
Computer vision › Segmentation and scene understanding
scene understanding
0.312017
Physics Inspired Optimization on Semantic Transfer Features: An Alternative Method for Room Layout Estimation · CVPR 2017
Visual content generation and editing › style transfer
semantic style transfer
0.312017
Decoder Network over Lightweight Reconstructed Feature for Fast Semantic Style Transfer · ICCV 2017
Computer vision › Video understanding and tracking › object tracking › discriminative tracking
correlation filter tracking
0.212016
Real-Time Visual Tracking: Promoting the Robustness of Correlation Filter Learning · ECCV (8) 2016
Computer vision › Video understanding and tracking › object tracking
discriminative tracking
0.212015
Discriminative Low-Rank Tracking · ICCV 2015
Computer vision › Video understanding and tracking › object tracking
single object tracking
0.212015
Single Object Tracking With Fuzzy Least Squares Support Vector Machine · IEEE Trans. Image Process. 2015
Computer vision › Video understanding and tracking › object tracking
target representation
0.212015
Robust Visual Tracking via Sparsity-Induced Subspace Learning · IEEE Trans. Image Process. 2015
Computer vision › Video understanding and tracking › multi-object tracking
tracking-by-detection
0.212015
Single Object Tracking With Fuzzy Least Squares Support Vector Machine · IEEE Trans. Image Process. 2015
Image and video processing › image restoration
image deblurring
0.212014
Efficient Patch-Wise Non-Uniform Deblurring for a Single Image · IEEE Trans. Multim. 2014
Image and video processing
image restoration
0.212014
Efficient Patch-Wise Non-Uniform Deblurring for a Single Image · IEEE Trans. Multim. 2014

Methods — techniques the papers use, named apart from their topics

auxiliary-view enhanced network · 0.9sight line estimation · 0.9generative model · 0.9density voxel grid · 0.8coarse-to-fine optimization · 0.8fuzzy least squares support vector machine · 0.7correlation filter learning · 0.6strip pooling · 0.4spatial pooling · 0.4region-based optimization · 0.4polar-based local partitioning · 0.4color histogram · 0.4whitening and coloring transform · 0.4optimal transport · 0.4adaptive instance normalization · 0.4feature reconstruction · 0.3decoder network · 0.3kernel estimation · 0.2
YearPublicationVenuePosition
2025 Dual-view X-ray Detection: Can AI Detect Prohibited Items from Dual-view X-ray Images like Humans?
abstract
To detect prohibited items in challenging categories, human inspectors typically rely on images from two distinct views (vertical and side). Can AI detect prohibited items from dual-view X-ray images in the same way humans do? Existing X-ray datasets often suffer from limitations, such as single-view imaging or insufficient sample diversity. To address these gaps, we introduce the Large-scale Dual-view X-ray (LDXray), which consists of 353,646 instances across 12 categories, providing a diverse and comprehensive resource for training and evaluating models. To emulate human intelligence in dual-view detection, we propose the Auxiliary-view Enhanced Network (AENet), a novel detection framework that leverages both the main and auxiliary views of the same object. The main-view pipeline focuses on detecting common categories, while the auxiliary-view pipeline handles more challenging categories using "expert models" learned from the main view. Extensive experiments on the LDXray dataset demonstrate that the dual-view mechanism significantly enhances detection performance, e.g., achieving improvements of up to +24.7% for the challenging category of umbrellas. Furthermore, our results show that AENet exhibits strong generalization across seven different detection models for X-ray Inspection1.
Renshuai Tao, Yuzhe Guo, Hairong Chen, Li Zhang 0023, Xianglong Liu 0001, Yunchao Wei, Yao Zhao 0001
CVPR5
2024 CBARF: Cascaded Bundle-Adjusting Neural Radiance Fields From Imperfect Camera Poses
abstract
Existing volumetric neural rendering techniques, such as Neural Radiance Fields (NeRF), face limitations in synthesizing high-quality novel views when the camera poses of input images are imperfect. To address this issue, we propose a novel 3D reconstruction framework that enables simultaneous optimization of camera poses, dubbed CBARF (Cascaded Bundle-Adjusting NeRF). In a nutshell, our framework optimizes camera poses in a coarse-to-fine manner and then reconstructs scenes based on the rectified poses. It is observed that the initialization of camera poses has a significant impact on the performance of bundle-adjustment (BA). Therefore, we cascade multiple BA modules at different scales to progressively improve the camera poses. Meanwhile, we develop a neighbor-replacement strategy to further optimize the results of BA in each stage. In this step, we introduce a novel criterion to effectively identify poorly estimated camera poses. Then we replace them with the poses of neighboring cameras, thus further eliminating the impact of inaccurate camera poses. Once camera poses have been optimized, we employ a density voxel grid to generate high-quality 3D reconstructed scenes and images in novel views. Experimental results demonstrate that our CBARF model achieves state-of-the-art performance in both pose optimization and novel view synthesis, especially in the existence of large camera pose noise.
Hongyu Fu, Xin Yu 0002, Lincheng Li, Li Zhang 0023
IEEE Trans. Multim.4
2022 Detecting prohibited objects with physical size constraint from cluttered X-ray baggage images
An Chang, Yu Zhang 0026, Shunli Zhang 0005, Leisheng Zhong, Li Zhang 0023
Knowl. Based Syst.5
2020 Strip Pooling: Rethinking Spatial Pooling for Scene Parsing
abstract
Spatial pooling has been proven highly effective to capture long-range contextual information for pixel-wise prediction tasks, such as scene parsing. In this paper, beyond conventional spatial pooling that usually has a regular shape of NxN, we rethink the formulation of spatial pooling by introducing a new pooling strategy, called strip pooling, which considers a long but narrow kernel, i.e., 1xN or Nx1. Based on strip pooling, we further investigate spatial pooling architecture design by 1) introducing a new strip pooling module that enables backbone networks to efficiently model long-range dependencies; 2) presenting a novel building block with diverse spatial pooling as a core; and 3) systematically comparing the performance of the proposed strip pooling and conventional spatial pooling techniques. Both novel pooling-based designs are lightweight and can serve as an efficient plug-and-play modules in existing scene parsing networks. Extensive experiments on Cityscapes and ADE20K benchmarks demonstrate that our simple approach establishes new state-of-the-art results. Code is available at https://github.com/Andrew-Qibin/SPNet.
Qibin Hou, Li Zhang 0023, Ming-Ming Cheng, Jiashi Feng
CVPR2
2020 Pointly-supervised scene parsing with uncertainty mixture
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yiwen Guo, Yurong Chen 0001, Li Zhang 0023
Comput. Vis. Image Underst.6
2020 Learning to Draw Sight Lines
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yurong Chen 0001, Li Zhang 0023
Int. J. Comput. Vis.5
2020 Joint Correlation Filtering for Visual Tracking
abstract
Correlation filtering-based visual tracking has achieved impressive success in terms of both tracking accuracy and computational efficiency. In this paper, a novel correlation filtering approach is proposed by means of joint learning to bridge the gap between the circulant filtering and the classical filtering methods. The circulant structure of tracking and the information from successive frames are simultaneously exploited in the proposed work. A new formulation for the correlation filter learning is proposed to enhance the discrimination of the learned filter by integrating both the kernel and the image feature domains. The proposed approach is computationally efficient since a closed-form solution is derived for the new formulation. Extensive experiments are conducted on two popular tracking benchmarks, and the experimental results demonstrate that the proposed tracker outperforms most of the state-of-the-art trackers.
Yao Sui, Guanghui Wang 0001, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.3
2020 Learning Scale-Adaptive Tight Correlation Filter for Object Tracking
abstract
In this paper, we propose a novel tracking method by formulating tracking as a correlation filtering as well as a ridge regression problem. First, we develop a tight correlation filter-based tracking framework from the signal detection perspective. In this formulation, the correlation filter is set as the same size as the target, which can make full use of the relations of the adjacent image patches and effectively exclude the influence of the background. Specifically, we point out that the novel correlation filter model can be regarded as the ridge regression model which takes into account the different importance of the samples and has the consistent objective with tracking. Second, we focus on the scale variation problem in tracking. By making use of the spatial structure of the correlation filter, the multiscale filter banks can be generated via interpolation to handle the scale estimation problem easily. Third, we present a novel distance importance-based confidence calculation model to determine the final tracking result, which not only makes use of the fine discriminability of the correlation filter but also takes the distance importance of the candidate samples into account to alleviate the impact of similar distractors. Experimental results demonstrate that our method is superior to several state-of-the-art trackers and many other correlation filter-based methods in the benchmark datasets.
Shunli Zhang 0005, Wei Lu 0010, Weiwei Xing, Li Zhang 0023
IEEE Trans. Cybern.4
2020 Occlusion-Aware Region-Based 3D Pose Tracking of Objects With Temporally Consistent Polar-Based Local Partitioning
abstract
Region-based methods have become the state-of-art solution for monocular 6-DOF object pose tracking in recent years. However, two main challenges still remain: the robustness to heterogeneous configurations (both foreground and background), and the robustness to partial occlusions. In this paper, we propose a novel region-based monocular 3D object pose tracking method to tackle these problems. Firstly, we design a new strategy to define local regions, which is simple yet efficient in constructing discriminative local color histograms. Contrary to previous methods which define multiple circular regions around the object contour, we propose to define multiple overlapped, fan-shaped regions according to polar coordinates. This local region partitioning strategy produces much less number of local regions that need to be maintained and updated, while still being temporally consistent. Secondly, we propose to detect occluded pixels using edge distance and color cues. The proposed occlusion detection strategy is seamlessly integrated into the region-based pose optimization pipeline via a pixel-wise weight function, which significantly alleviates the interferences caused by partial occlusions. We demonstrate the effectiveness of the proposed two new strategies with a careful ablation study. Furthermore, we compare the performance of our method with the most recent state-of-art region-based methods in a recently released large dataset, in which the proposed method achieves competitive results with a higher average tracking success rate. Evaluations on two real-world datasets also show that our method is capable of handling realistic tracking scenarios.
Leisheng Zhong, Yu Zhang 0026, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Image Process.5
2020 Toward Precise Osteotomies: A Coarse-to-Fine 3D Cut Plane Planning Method for Image-Guided Pelvis Tumor Resection Surgery
abstract
Surgical resection is the main clinical method for the treatment of bone tumors. A critical procedure for bone tumor resection is to plan a set of cut planes that enable resecting the bone tumor with a safe margin while preserving the maximum amount of healthy bone. Currently, the surgeons rely on manual methods to plan the cut planes, which highly depend on the surgeons' experiences and have been demonstrated to be error-prone, and in turn, increase the recurrence rate or resect much healthy bone. This study targets on improving the precision of cut plane planning for the image guided pelvis tumor resection surgeries. A semi-automatic approach to cut plane planning was proposed via a coarse-to-fine strategy. It can efficiently identify a dangerous region in the 3D space, which contains the bone tumor and its surrounding normal tissue with a safe margin. By projecting the dangerous region into an appropriate 2D space, a segmented boundary-constrained linear regression method was leveraged to plan a set of 3D cut planes that ensure the minimum area of the resected specimen in the 2D space while having the dangerous region cleared. Further, a coarse-to-fine 3D cut plane planning method was developed by incorporating a 3D cut plane refinement scheme with our 2D planning method. Extensive experiments, on the surgical data from nine previous pelvis tumor resection surgeries, demonstrated that our proposed approach substantially improved the localization precision of cut planes ( ) and decreased the amount of resected specimen ( ), as compared to the manual method.
Yu Zhang 0026, Fengzan Li, Lei Qiu 0004, Lihui Xu, Xiaohui Niu, Yao Sui, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Medical Imaging9
2020 Fuzzy Least Squares Support Vector Machine With Adaptive Membership for Object Tracking
abstract
Fuzzy learning has been introduced into tracking and achieved great success. However, the membership in the existing fuzzy learning based tracking algorithm is fixed, which lacks the adaptivity to measure the importance of the samples. To improve the tracking adaptivity and flexibility, in this paper, we propose a novel tracking method based on fuzzy least squares support vector machine with adaptive membership (FLS-SVM-AM). First, we formulate tracking as an adaptive membership based fuzzy learning problem, which addresses the issue of fixed membership in existing methods and can better measure the importance of the training samples. Second, we present the FLS-SVM-AM method to build the appearance model, and develop an iterative optimization process to solve the FLS-SVM-AM problem. Third, we define a new membership based on the PASCAL VOC overlap rate and exponential function, which is used to measure the importance of different samples more accurately. Experimental results in the benchmark datasets demonstrate that the proposed method not only outperforms the existing fuzzy learning based tracking methods, but also is comparable to many state-of-the-art methods.
Shunli Zhang 0005, Li Zhang 0023, Alex Hauptmann 0001
IEEE Trans. Multim.2
2019 A Closed-Form Solution to Universal Style Transfer
abstract
Universal style transfer tries to explicitly minimize the losses in feature space, thus it does not require training on any pre-defined styles. It usually uses different layers of VGG network as the encoders and trains several decoders to invert the features into images. Therefore, the effect of style transfer is achieved by feature transform. Although plenty of methods have been proposed, a theoretical analysis of feature transform is still missing. In this paper, we first propose a novel interpretation by treating it as the optimal transport problem. Then, we demonstrate the relations of our formulation with former works like Adaptive Instance Normalization (AdaIN) and Whitening and Coloring Transform (WCT). Finally, we derive a closed-form solution named Optimal Style Transfer (OST) under our formulation by additionally considering the content loss of Gatys. Comparatively, our solution can preserve better structure and achieve visually pleasing results. It is simple yet effective and we demonstrate its advantages both quantitatively and qualitatively. Besides, we hope our theoretical analysis can inspire future works in neural style transfer.
Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Feng Xu 0005, Li Zhang 0023
ICCV6
2019 Exploiting the Anisotropy of Correlation Filter Learning for Visual Tracking
Yao Sui, Guanghui Wang 0001, Yafei Tang, Li Zhang 0023
Int. J. Comput. Vis.5
2019 A Robust Monocular 3D Object Tracking Method Combining Statistical and Photometric Constraints
Leisheng Zhong, Li Zhang 0023
Int. J. Comput. Vis.2
2019 Sparse subspace clustering via Low-Rank structure propagation
Yao Sui, Guanghui Wang 0001, Li Zhang 0023
Pattern Recognit.3
2019 Single Image Depth Estimation With Normal Guided Scale Invariant Deep Convolutional Fields
abstract
Estimating scene depth from a single image can be widely applied to understand 3D environments due to the easy access of the images captured by consumer-level cameras. Previous works exploit conditional random fields (CRFs) to estimate image depth, where neighboring pixels (superpixels) with similar appearances are constrained to share the same depth. However, the depth may vary significantly in the slanted surface, thus leading to severe estimation errors. In order to eliminate those errors, we propose a superpixel-based normal guided scale invariant deep convolutional field by encouraging the neighboring superpixels with similar appearance to lie on the same 3D plane of the scene. In doing so, a depth-normal multitask CNN is introduced to produce the superpixel-wise depth and surface normal predictions simultaneously. To correct the errors of the roughly estimated superpiexl-wise depth, we develop a normal guided scale invariant CRF (NGSI-CRF). NGSI-CRF consists of a scale invariant unary potential that is able to measure the relative depth between superpixels as well as the absolute depth of superpixels, and a normal guided pairwise potential that constrains spatial relationships between superpixels in accordance with the 3D layout of the scene. In other words, the normal guided pairwise potential is designed to smooth the depth prediction without deteriorating the 3D structure of the depth prediction. The superpixel-wise depth maps estimated by NGSI-CRF will be fed into a pixel-wise refinement module to produce a smooth fine-grained depth prediction. Furthermore, we derive a closed-form solution for the maximum a posteriori (MAP) inference of NGSI-CRF. Thus, our proposed network can be efficiently trained in an end-to-end manner. We conduct our experiments on various datasets, such as NYU-D2, KITTI, and Make 3D. As demonstrated in the experimental results, our method achieves superior performance in both indoor and outdoor scenes.
Han Yan 0007, Xin Yu 0002, Yu Zhang 0026, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.6
2018 Efficient Semantic Scene Completion Network with Spatial Group Convolution
Hao Zhao 0002, Anbang Yao, Yurong Chen 0001, Li Zhang 0023, Hongen Liao
ECCV (12)5
2018 Visual Tracking via Subspace Learning: A Discriminative Approach
Yao Sui, Yafei Tang, Li Zhang 0023, Guanghui Wang 0001
Int. J. Comput. Vis.3
2018 Monocular depth estimation with guidance of surface normal map
Han Yan 0007, Shunli Zhang 0005, Yu Zhang 0026, Li Zhang 0023
Neurocomputing4
2018 Using fuzzy least squares support vector machine with metric learning for object tracking
Shunli Zhang 0005, Wei Lu 0010, Weiwei Xing, Li Zhang 0023
Pattern Recognit.4
2018 PMSC: PatchMatch-Based Superpixel Cut for Accurate Stereo Matching
abstract
Estimating the disparity and normal direction of one pixel simultaneously, instead of only disparity, also known as 3D label methods, can achieve much higher subpixel accuracy in the stereo matching problem. However, it is extremely difficult to assign an appropriate 3D label to each pixel from the continuous label space R3 while maintaining global consistency because of the infinite parameter space. In this paper, we propose a novel algorithm called PatchMatch-based superpixel cut to assign 3D labels of an image more accurately. In order to achieve robust and precise stereo matching between local windows, we develop a bilayer matching cost, where a bottom-up scheme is exploited to design the two layers. The bottom layer is employed to measure the similarity between small square patches locally by exploiting a pretrained convolutional neural network, and then, the top layer is developed to assemble the local matching costs in large irregular windows induced by the tangent planes of object surfaces. To optimize the spatial smoothness of local assignments, we propose a novel strategy to update 3D labels. In the procedure of optimization, both segmentation information and random refinement of PatchMatch are exploited to update candidate 3D label set for each pixel with high probability of achieving lower loss. Since pairwise energy of general candidate label sets violates the submodular property of graph cut, we propose a novel multilayer superpixel structure to group candidate label sets into candidate assignments, which thereby can be efficiently fused by α-expansion graph cut. Extensive experiments demonstrate that our method can achieve higher subpixel accuracy in different data sets, and currently ranks first on the new challenging Middlebury 3.0 benchmark among all the existing methods.
Lincheng Li, Shunli Zhang 0005, Xin Yu 0002, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.4
2018 A Direct 3D Object Tracking Method Based on Dynamic Textured Model Rendering and Extended Dense Feature Fields
abstract
We propose a novel method for robust 6-DOF pose tracking of rigid objects from monocular images. In our method, 3D object tracking is achieved by directly aligning video frames to dynamic templates rendered from a textured 3D object model. Unlike previous methods, which usually utilize a small number of discrete templates to align with video frames, we employ an online textured model, rendering to create dynamic templates in continuous pose space according to the previously estimated object pose. In this way, a pose estimator could be easily converged to the optimal state. Besides, the rendered template also helps to detect the occlusion area by comparing it with the current frame, making our method highly robust to partial occlusions. The performance of our method is further improved by introducing a generic representation of dense images features, which we call extended dense feature fields (EDFF). Different kinds of pixel-level image features can be added to the EDFF and be optimized simultaneously in a unified Gauss-Newton optimization scheme. Attributing to dynamic templates from the textured model rendering and complementary features in EDFF, our method is able to deal with poor-textured and specular objects, as well as lighting variation and heavy occlusions. While our method is quite simple and straightforward, it achieves competitive or even superior results compared with the state of the art on challenging data sets.
Leisheng Zhong, Ming Lu 0002, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.3
2018 Correlation Filter Learning Toward Peak Strength for Visual Tracking
abstract
This paper presents a novel visual tracking approach to correlation filter learning toward peak strength of correlation response. Previous methods leverage all features of the target and the immediate background to learn a correlation filter. Some features, however, may be distractive to tracking, like those from occlusion and local deformation, resulting in unstable tracking performance. This paper aims at solving this issue and proposes a novel algorithm to learn the correlation filter. The proposed approach, by imposing an elastic net constraint on the filter, can adaptively eliminate those distractive features in the correlation filtering. A new peak strength metric is proposed to measure the discriminative capability of the learned correlation filter. It is demonstrated that the proposed approach effectively strengthens the peak of the correlation response, leading to more discriminative performance than previous methods. Extensive experiments on a challenging visual tracking benchmark demonstrate that the proposed tracker outperforms most state-of-the-art methods.
Yao Sui, Guanghui Wang 0001, Li Zhang 0023
IEEE Trans. Cybern.3
2018 Towards Occlusion Handling: Object Tracking With Background Estimation
abstract
The appearance model of the target needs to be updated for online single object tracking. However, the variation of the observation can be caused by active appearance change of the target, or the occlusion from the background. For the former case, we should update the appearance model and for the latter, the current model should be preserved. In this paper, we distinguish these two cases and resist the impact from heavy occlusion by estimating the background in the scene with moving cameras, while retaining the adaptivity to stationary cameras at the same time. The proposed method formulates the background as a Gaussian model and the target is determined in a coarse-to-fine manner. Experimental results demonstrate that our method achieves competitive results in the sequences with appearance changes and outperforms the state-of-the-art algorithms in dealing with complex occlusions.
Sicong Zhao, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Cybern.3
2018 Exploiting Spatial-Temporal Locality of Tracking via Structured Dictionary Learning
abstract
In this paper, a novel spatial-temporal locality is proposed and unified via a discriminative dictionary learning framework for visual tracking. By exploring the strong local correlations between temporally obtained target and their spatially distributed nearby background neighbors, a spatial-temporal locality is obtained. The locality is formulated as a subspace model and exploited under a unified structure of discriminative dictionary learning with a subspace structure. Using the learned dictionary, the target and its background can be described and distinguished effectively through their sparse codes. As a result, the target is localized by integrating both the descriptive and the discriminative qualities. Extensive experiments on various challenging video sequences demonstrate the superior performance of proposed algorithm over the other state-of-the-art approaches.
Yao Sui, Guanghui Wang 0001, Li Zhang 0023, Ming-Hsuan Yang 0001
IEEE Trans. Image Process.3
2018 Estimating Maximum Target Registration Error Under Uniform Restriction of Fiducial Localization Error in Image Guided System
abstract
In this paper, we investigate the estimation of the maximum target registration error (TRE) magnitude of the target location while using point-based rigid registration in the image guided system. Under the uniform restriction of fiducial localization error (FLE) magnitude, we explicitly formulate the estimation as an optimization problem. Through analyzing the approximated problem which assumes the rigidity of the fiducial set holds with the perturbation of FLE, we present a strict lower bound for the maximum TRE magnitude. The simulations show that the lower bound is close to the actual maximum TRE magnitude for the target locations lying far away from the fiducial points. Unlike the expected TRE magnitude in which all fiducial points contribute, the lower bound is only related to the fiducial points serving as the vertices of the convex hull of the fiducial set. Our analysis provides a new perspective of investigating the problem of TRE estimation and is helpful for the surgeons to learn about the worst situation during using the image guided system.
Lei Qiu 0004, Yu Zhang 0026, Lihui Xu, Xiaohui Niu, Li Zhang 0023
IEEE Trans. Medical Imaging6
2017 Physics Inspired Optimization on Semantic Transfer Features: An Alternative Method for Room Layout Estimation
abstract
In this paper, we propose an alternative method to estimate room layouts of cluttered indoor scenes. This method enjoys the benefits of two novel techniques. The first one is semantic transfer (ST), which is: (1) a formulation to integrate the relationship between scene clutter and room layout into convolutional neural networks, (2) an architecture that can be end-to-end trained, (3) a practical strategy to initialize weights for very deep networks under unbalanced training data distribution. ST allows us to extract highly robust features under various circumstances, and in order to address the computation redundance hidden in these features we develop a principled and efficient inference scheme named physics inspired optimization (PIO). PIOs basic idea is to formulate some phenomena observed in ST features into mechanics concepts. Evaluations on public datasets LSUN and Hedau show that the proposed method is more accurate than state-of-the-art methods.
Hao Zhao 0002, Ming Lu 0002, Anbang Yao, Yiwen Guo, Yurong Chen 0001, Li Zhang 0023
CVPR6
2017 Decoder Network over Lightweight Reconstructed Feature for Fast Semantic Style Transfer
abstract
Recently, the community of style transfer is trying to incorporate semantic information into traditional system. This practice achieves better perceptual results by transferring the style between semantically-corresponding regions. Yet, few efforts are invested to address the computation bottleneck of back-propagation. In this paper, we propose a new framework for fast semantic style transfer. Our method decomposes the semantic style transfer problem into feature reconstruction part and feature decoder part. The reconstruction part tactfully solves the optimization problem of content loss and style loss in feature space by particularly reconstructed feature. This significantly reduces the computation of propagating the loss through the whole network. The decoder part transforms the reconstructed feature into the stylized image. Through a careful bridging of the two modules, the proposed approach not only achieves competitive results as backward optimization methods but also is about two orders of magnitude faster.
Ming Lu 0002, Hao Zhao 0002, Anbang Yao, Feng Xu 0005, Yurong Chen 0001, Li Zhang 0023
ICCV6
2017 Motion feature augmented recurrent neural network for skeleton-based dynamic hand gesture recognition
abstract
Dynamic hand gesture recognition has attracted increasing interests because of its importance for human computer interaction. In this paper, we propose a new motion feature augmented recurrent neural network for skeleton-based dynamic hand gesture recognition. Finger motion features are extracted to describe finger movements and global motion features are utilized to represent the global movement of hand skeleton. These motion features are then fed into a bidirectional recurrent neural network (RNN) along with the skeleton sequence, which can augment the motion features for RNN and improve the classification performance. Experiments demonstrate that our proposed method is effective and outperforms start-of-the-art methods.
Xinghao Chen 0001, Hengkai Guo, Guijin Wang, Li Zhang 0023
ICIP4
2017 Graph-Regularized Structured Support Vector Machine for Object Tracking
abstract
How to build a robust and accurate appearance model is a crucial problem in object tracking. However, in most existing tracking methods, the structures among the adjacent video frames and neighboring regions, which may be helpful to improve the representation capability of the appearance model, have not been fully exploited. In this paper, we propose a novel tracking method by taking into account these structures to represent the appearance model. First, we propose a novel graph-regularized structured support vector machine (GS-SVM) algorithm by combining manifold learning and structured learning. Then, the proposed GS-SVM algorithm is employed to build the appearance model and a novel tracking method is developed. This new tracker not only absorbs the advantage of structured learning that deals with the intermediate classification step existing in tracking-by-detection methods, but also exploits the geometry structures in the tracked results and neighboring regions. In addition, a hybrid update strategy is introduced to fit with the GS-SVM-based appearance model. The experimental results demonstrate that the proposed tracking algorithm can outperform several state-of-the-art tracking methods in the benchmark dataset.
Shunli Zhang 0005, Yao Sui, Sicong Zhao, Li Zhang 0023
IEEE Trans. Circuits Syst. Video Technol.4
2016 Tracking Completion
Yao Sui, Guanghui Wang 0001, Yafei Tang, Li Zhang 0023
ECCV (8)4
2016 Real-Time Visual Tracking: Promoting the Robustness of Correlation Filter Learning
Yao Sui, Guanghui Wang 0001, Yafei Tang, Li Zhang 0023
ECCV (8)5
2016 Robust Tracking via Locally Structured Representation
Yao Sui, Li Zhang 0023
Int. J. Comput. Vis.2
2016 Object Tracking With Spatial Context Model
abstract
In object tracking, building a reliable appearance model can greatly improve the performance. In this letter, we propose a novel method that uses the spatial context to help tracking based on support vector machines (SVMs). The spatial context is decomposed into different subregions that include some parts of both the target and the background. We build appearance submodels for each group of subregions and combine them to get a robust appearance model, which can help to handle some complex problems, e.g., occlusion and deformation. Besides, we add an update strategy to retain the accuracy of the appearance model. A large number of experiments on various challenging videos demonstrate that our method outperforms many other state-of-the-art methods.
Juntao Sun, Shunli Zhang 0005, Li Zhang 0023
IEEE Signal Process. Lett.3
2015 Discriminative Low-Rank Tracking
abstract
Good tracking performance is in general attributed to accurate representation over previously obtained targets or reliable discrimination between the target and the surrounding background. In this work, we exploit the advantages of the both approaches to achieve a robust tracker. We construct a subspace to represent the target and the neighboring background, and simultaneously propagate their class labels via the learned subspace. Moreover, we propose a novel criterion to identify the target from numerous target candidates on each frame, which takes into account both discrimination reliability and representation accuracy. In addition, with the proposed criterion, the ambiguity in the class labels of the neighboring background samples, which often influences the reliability of discriminative tracking model, is effectively alleviated, while the training set is still kept small. Extensive experiments demonstrate that our tracker performs favourably against many other state-of-the-art trackers.
Yao Sui, Yafei Tang, Li Zhang 0023
ICCV3
2015 Self-expressive tracking
Yao Sui, Shunli Zhang 0005, Xin Yu 0002, Sicong Zhao, Li Zhang 0023
Pattern Recognit.6
2015 Hybrid support vector machines for robust object tracking
Shunli Zhang 0005, Yao Sui, Xin Yu 0002, Sicong Zhao, Li Zhang 0023
Pattern Recognit.5
2015 Multi-local-task learning with global regularization for object tracking
Shunli Zhang 0005, Yao Sui, Sicong Zhao, Xin Yu 0002, Li Zhang 0023
Pattern Recognit.5
2015 Visual Tracking via Locally Structured Gaussian Process Regression
abstract
We propose a new target representation method, where the temporally obtained targets are jointly represented as a time series function by exploiting their spatially local structure. With this method, we propose a new tracking algorithm, where tracking is formulated as a problem of Gaussian process regression over the joint representation. Numerous experiments on various challenging video sequences demonstrate that our tracker outperforms several other state-of-the-art trackers.
Yao Sui, Li Zhang 0023
IEEE Signal Process. Lett.2
2015 Robust Visual Tracking via Sparsity-Induced Subspace Learning
abstract
Target representation is a necessary component for a robust tracker. However, during tracking, many complicated factors may make the accumulated errors in the representation significantly large, leading to tracking drift. This paper aims to improve the robustness of target representation to avoid the influence of the accumulated errors, such that the tracker only acquires the information that facilitates tracking and ignores the distractions. We observe that the locally mutual relations between the feature observations of temporally obtained targets are beneficial to the subspace representation in visual tracking. Thus, we propose a novel subspace learning algorithm for visual tracking, which imposes joint row-wise sparsity structure on the target subspace to adaptively exclude distractive information. The sparsity is induced by exploiting the locally mutual relations between the feature observations during learning. To this end, we formulate tracking as a subspace sparsity inducing problem. A large number of experiments on various challenging video sequences demonstrate that our tracker outperforms many other state-of-the-art trackers.
Yao Sui, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Image Process.3
2015 Single Object Tracking With Fuzzy Least Squares Support Vector Machine
abstract
Single object tracking, in which a target is often initialized manually in the first frame and then is tracked and located automatically in the subsequent frames, is a hot topic in computer vision. The traditional tracking-by-detection framework, which often formulates tracking as a binary classification problem, has been widely applied and achieved great success in single object tracking. However, there are some potential issues in this formulation. For instance, the boundary between the positive and negative training samples is fuzzy, and the objectives of tracking and classification are inconsistent. In this paper, we attempt to address the above issues from the fuzzy system perspective and propose a novel tracking method by formulating tracking as a fuzzy classification problem. First, we introduce the fuzzy strategy into tracking and propose a novel fuzzy tracking framework, which can measure the importance of the training samples by assigning different memberships to them and offer more strict spatial constraints. Second, we develop a fuzzy least squares support vector machine (FLS-SVM) approach and employ it to implement a concrete tracker. In particular, the primal form, dual form, and kernel form of FLS-SVM are analyzed and the corresponding closed-form solutions are derived for efficient realizations. Besides, a least squares regression model is built to control the update adaptively, retaining the robustness of the appearance model. The experimental results demonstrate that our method can achieve comparable or superior performance to many state-of-the-art methods.
Shunli Zhang 0005, Sicong Zhao, Yao Sui, Li Zhang 0023
IEEE Trans. Image Process.4
2015 Object Tracking With Multi-View Support Vector Machines
abstract
How to build an accurate and reliable appearance model to improve the performance is a crucial problem in object tracking. Since the multi-view learning can lead to more accurate and robust representation of the object, in this paper, we propose a novel tracking method via multi-view learning framework by using multiple support vector machines (SVM). The multi-view SVMs tracking method is constructed based on multiple views of features and a novel combination strategy. To realize a comprehensive representation, we select three different types of features, i.e., gray scale value, histogram of oriented gradients (HOG), and local binary pattern (LBP), to train the corresponding SVMs. These features represent the object from the perspectives of description, detection, and recognition, respectively . In order to realize the combination of the SVMs under the multi-view learning framework, we present a novel collaborative strategy with entropy criterion, which is acquired by the confidence distribution of the candidate samples. In addition, to learn the changes of the object and the scenario, we propose a novel update scheme based on subspace evolution strategy. The new scheme can control the model update adaptively and help to address the occlusion problems . We conduct our approach on several public video sequences and the experimental results demonstrate that our method is robust and accurate, and can achieve the state-of-the-art tracking performance.
Shunli Zhang 0005, Xin Yu 0002, Yao Sui, Sicong Zhao, Li Zhang 0023
IEEE Trans. Multim.5
2014 Efficient Patch-Wise Non-Uniform Deblurring for a Single Image
abstract
In this paper, we address the problem of estimating a latent sharp image from a single spatially variant blurred image. Non-uniform deblurring methods based on projective motion path models formulate the blur as a linear combination of homographic projections of a clear image. But they are computationally expensive and require large memory due to the calculation and storage of a large number of the projections. Patch-wise non-uniform deblurring algorithms have been proposed to estimate each kernel locally by a uniform deblurring algorithm, which does not require to calculate and store the projections. The key issues of these methods are the accuracy of kernel estimation and the identification of erroneous kernels. To perform accurate kernel estimation, we employ the total variation (TV) regularization to recover a latent image, in which the edges are better enhanced and the ringing artifacts are reduced, rather than Tikhonov regularization that previous algorithms adopt. Thus blur kernels can be estimated more accurately from the latent image and estimated in a closed form while previous methods cannot estimate kernels in closed forms. To identify the erroneous kernels, we develop a novel metric, which is able to measure the similarity between the neighboring kernels. After replacing the erroneous kernels with the well-estimated ones, a clear image is obtained. The experiments show that our approach can achieve better results on the real-world blurry images while using less computation and memory.
Xin Yu 0002, Feng Xu 0005, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Multim.4
2010 Managing Workplace Resources in Office Environments through Ephemeral Social Networks
Alvin Chin, Wenchang Xu, Hao Wang 0021, Li Zhang 0023
UIC6
2008 A Three Dimensional Combinative Lifting Algorithm for Wavelet Transform Using 9/7 Filter
abstract
Three dimensional-discrete wavelet transform (3D-DWT) shows its advantages in the volumetric data compression. Some algorithms and architectures have been proposed for the 3D-DWT, but most of them perform the transform by applying three separate 1D transforms, or by applying inter-frame transform and spatial transform separately. In this paper, we choose 9/7 filter along the three directions to do the transform based on lifting scheme. Using multilinear algebra (L.D. Lathauwer, 1997), this paper let Xnoperation stand for a high order tensor's linear transform in its nth dimension. Then we successfully extend spatial combinative lifting algorithm (SCLA)(H. Meng and Z. Wang, 2000) to its 3D application. Matrices A,B,C,D,E for lifting algorithm of 9/1 filter (H. Meng and Z. Wang, 2000) can be multiplied step by step to get the final result of one level decomposition of 3D-DWT. The operation on matrix A could be represented as X1A X2A X3A.
Li Zhang 0023
DCC2
2007 A Low Power, Fully Pipelined JPEG-LS Encoder for Lossless Image Compression
abstract
By analyzing the features unfit for parallel computation and low power implementation, a VLSI architecture of JPEG-LS encoder for lossless image compression is proposed in this paper. It functionally consists of four parts: Mode decision module, clock controller, three linear parallel pipelines, and a two-tier data packer. Computations are organized in a fully pipelined style in these modules, so that real time data processing can be achieved. The clock management scheme with four interlaced clock domains and a dedicated clock controller is applied to ensure the bottleneck calculation, reduce the clock frequency on non-critical paths, and shut off the working clocks of idle modules, which reduces 15.7% of overall power consumption. The proposed JPEG-LS encoder with the features of low power and high processing speed, has been applied in a wireless endoscopy system.
Xinkai Chen, Guolin Li, Li Zhang 0023, Chun Zhang 0001, Zhihua Wang 0001
ICME5
2007 Design and Implementation of a Low Complexity Near-lossless Image Compression Method for Wireless Endoscopy Capsule System
abstract
This paper proposes a new low complexity near-lossless image compression method and its VLSI design for low power and high frame rate in the wireless endoscopy capsule system. Assuring outstanding compression performance and high image quality, the proposed method with the features of low complexity, low storage overhead and real time data processing, makes it ideal for hardware implementation. The VLSI architecture consists of two pipelined parts: The preprocessor and the JPEG-LS engine. A fully pipelined VLSI structure with a dedicated clock management scheme is proposed for the JPEG-LS engine, which ensures a low power application, and real time data processing as well. The hardware implementation has been verified on FPGA and implemented in 0.18μm CMOS technology.
Xinkai Chen, Guolin Li, Li Zhang 0023, Zhihua Wang 0001, Hong Chen 0002
ISCAS5
2007 A 2-GHz 6.1-mA Fully-Differential CMOS Phase-Locked Loop
abstract
A fully-differential 2-GHz phase-locked loop (PLL) was designed and fabricated in 0.18-μm CMOS process. The PLL rejects the common noise due to fully-differential VCO and differential charge pump. The VCO has a 16.15% tuning range (from 1.8998GHz to 2.2335GHz) due to a combination of analog and digital tuning technique (4-bit binary switch-capacitor array). With the pn-junction varactors, the phase noise of the VCO varies only about 2dB in the tuning range. The current consumption of the PLL is only about 6.1mA from a 1.8V power supply. It is comparable to the results reported in recent literatures. The phase noise of the PLL at 2.033GHz can achieve -117.17dBc/Hz at 1 MHz frequency offset from the carrier.
Li Zhang 0023, Baoyong Chi, Zhihua Wang 0001, Jinke Yao, Ende Wu
ISCAS1
2004 An improved algorithm for rate distortion optimization in JPEG2000 and its integrated circuit implementation
abstract
Rate distortion optimization (RDO) plays an important role in a JPEG2000 encoder. An improved RDO algorithm is presented in this paper. The proposed algorithm is suitable for integrated circuit implementation and can reduce the computational cost. A hardware architecture which includes control unit, memory, divider, and data converter is also given to implement the algorithm. The circuit, based on the improved algorithm, has been integrated in a JPG2000 chip codec core.
Zhihua Wang 0001, Li Zhang 0023, Chun Zhang 0001
ICASSP (5)3
2004 A new approach for near-lossless and lossless image compression with Bayer color filter arrays
abstract
This paper presents a new approach for near-lossless and lossless image compression in digital colorful image sensors with Bayer color filter arrays (CFAs). In this approach, the captured CFA raw data is firstly transformed from quincunx shape to rectangular shape and then smoothed by a low-pass filter. Lastly, the filtered data are compressed directly before full color than conventional interpolation-first image compression methods and other existing similar compression-first methods. Furthermore, through altering the quality control factor, the image quality, PSNR, can be changed from 44.15 dB to infinite with the compression ratio from averagely 3.3 bits/pixel to 6.9 bits/pixel.
Guolin Li, Zhihua Wang 0001, Chun Zhang 0001, Li Zhang 0023
ICIG7