Ze-Nian Li

dblp:l/ZeNianLi · DBLP profile ↗
← Back
85ranked-venue papers
15as first author
0since 2021 · last 2019
0000-0001-9812-1537ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 4 first-authorArtificial intelligence and machine learning · 45 · 9 first-authorHuman-computer interaction and ubiquitous computing · 7 · 2 first-authorSystems, architecture and hardware · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Video understanding and tracking · 40% Image recognition and object detection · 27% 3D vision · 18%
Computer graphics and multimedia
6 papers
Multimedia systems and quality of experience · 58% Image and video processing · 25% Visualization and visual analytics · 9%
Databases, data mining, and information retrieval
3 papers
Machine learning and data management · 64% Data mining · 32% Information retrieval · 5%

Topics — the 30 heaviest of 49, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
0.532015
Object Detection Using Generalization and Efficiency Balanced Co-Occurrence Features · ICCV 2015
Basis mapping based boosting for object detection · CVPR 2015
Matching by Linear Programming and Successive Convexification · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Computer vision › Video understanding and tracking
action recognition
0.422019
Deep Attention Network for Egocentric Action Recognition · IEEE Trans. Image Process. 2019
Unsupervised Discovery of Action Classes · CVPR (2) 2006
Computer vision › Image recognition and object detection
pedestrian detection
0.422015
Object Detection Using Generalization and Efficiency Balanced Co-Occurrence Features · ICCV 2015
Basis mapping based boosting for object detection · CVPR 2015
Computer vision › Video understanding and tracking › action recognition › human action recognition
egocentric action recognition
0.412019
Deep Attention Network for Egocentric Action Recognition · IEEE Trans. Image Process. 2019
Computer vision › Video understanding and tracking
gaze prediction
0.412019
Deep Attention Network for Egocentric Action Recognition · IEEE Trans. Image Process. 2019
Computer vision › 3D vision
depth estimation
0.232015
Continuous Depth Map Reconstruction From Light Fields · IEEE Trans. Image Process. 2015
Reciprocal-Wedge Transform for Space-Variant Sensing · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Reciprocal-Wedge Transform in Motion Stereo · ICRA 1994
Computer vision › 3D vision › depth estimation › multi-view depth estimation
light field depth estimation
0.212015
Continuous Depth Map Reconstruction From Light Fields · IEEE Trans. Image Process. 2015
Computer vision › Segmentation and scene understanding › semantic segmentation
category-specific segmentation
0.212013
Optimizing Nondecomposable Loss Functions in Structured Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Machine learning and data management
structured prediction
0.212013
Optimizing Nondecomposable Loss Functions in Structured Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2013
Multimedia systems and quality of experience › objective quality assessment
temporal distortion measurement
0.212013
Efficient video quality assessment based on spacetime texture representation · ACM Multimedia 2013
Multimedia systems and quality of experience
video quality assessment
0.212013
Efficient video quality assessment based on spacetime texture representation · ACM Multimedia 2013
Machine learning › Deep learning architectures and training
two-stream network
0.112019
Deep Attention Network for Egocentric Action Recognition · IEEE Trans. Image Process. 2019
Computer vision › 3D vision
correspondence estimation
0.112007
Matching by Linear Programming and Successive Convexification · IEEE Trans. Pattern Anal. Mach. Intell. 2007
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.112015
Basis mapping based boosting for object detection · CVPR 2015
Computer vision › Video understanding and tracking
action detection
0.112006
Successive Convex Matching for Action Detection · CVPR (2) 2006
Computer vision › Segmentation and scene understanding
image segmentation
0.112006
Unsupervised Discovery of Action Classes · CVPR (2) 2006
Machine learning › Graph learning › graph clustering
spectral clustering
0.112006
Unsupervised Discovery of Action Classes · CVPR (2) 2006
Computer vision › Video understanding and tracking
activity recognition
0.112005
Activity Recognition through Goal-Based Segmentation · AAAI 2005
Image and video processing › saliency detection
video saliency detection
0.012013
Efficient video quality assessment based on spacetime texture representation · ACM Multimedia 2013
Visualization and visual analytics
visual saliency
0.012013
Efficient video quality assessment based on spacetime texture representation · ACM Multimedia 2013
Image and video processing
motion estimation
0.012004
Optimizing Motion Estimation with Linear Programming and Detail-Preserving Variational Method · CVPR (1) 2004
Computational photography and imaging
camera calibration
0.012001
On Active Camera Control and Camera Motion Recovery with Foveate Wavelet Transform · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Image and video processing
image representation
0.012001
On Active Camera Control and Camera Motion Recovery with Foveate Wavelet Transform · IEEE Trans. Pattern Anal. Mach. Intell. 2001
Computer vision › 3D vision › stereo vision
motion stereo
0.021995
Reciprocal-Wedge Transform for Space-Variant Sensing · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Reciprocal-Wedge Transform in Motion Stereo · ICRA 1994
Computer vision › Image recognition and object detection › object recognition › appearance-based object recognition
color-based object recognition
0.011998
Illumination-Invariant Color Object Recognition via Compressed Chromaticity Histograms of Color-Channel-Normalized Images · ICCV 1998
Computer vision › Image recognition and object detection › object recognition › invariant object recognition
illumination-invariant recognition
0.011998
Illumination-Invariant Color Object Recognition via Compressed Chromaticity Histograms of Color-Channel-Normalized Images · ICCV 1998
Data mining › pattern mining
association rule mining
0.011998
MultiMediaMiner: A System Prototype for Multimedia Data Mining · SIGMOD Conference 1998
Data mining
clustering
0.011998
MultiMediaMiner: A System Prototype for Multimedia Data Mining · SIGMOD Conference 1998
Data mining › multimodal data mining
multimedia data mining
0.011998
MultiMediaMiner: A System Prototype for Multimedia Data Mining · SIGMOD Conference 1998
Data mining
pattern mining
0.011998
MultiMediaMiner: A System Prototype for Multimedia Data Mining · SIGMOD Conference 1998

Methods — techniques the papers use, named apart from their topics

boosting · 0.4HOG · 0.4two-stream network · 0.4spatial attention network · 0.4bidirectional LSTM · 0.4sparse linear system · 0.2kernel approximation · 0.2haar features · 0.2conjugate gradient · 0.2LBP · 0.2spatial pooling · 0.2self-information · 0.2quadratic programming · 0.2piecewise linear approximation · 0.2optical flow · 0.2markov random field · 0.2LP relaxation · 0.2variational method · 0.1
YearPublicationVenuePosition
2019 Deep Attention Network for Egocentric Action Recognition
abstract
Recognizing a camera wearer's actions from videos captured by an egocentric camera is a challenging task. In this paper, we employ a two-stream deep neural network composed of an appearance-based stream and a motion-based stream to recognize egocentric actions. Based on the insight that human action and gaze behavior are highly coordinated in object manipulation tasks, we propose a spatial attention network to predict human gaze in the form of attention map. The attention map helps each of the two streams to focus on the most relevant spatial region of the video frames to predict actions. To better model the temporal structure of the videos, a temporal network is proposed. The temporal network incorporates bi-directional long short-term memory to model the long-range dependencies to recognize egocentric actions. The experimental results demonstrate that our method is able to predict attention maps that are consistent with human attention and achieve competitive action recognition performance with the state-of-the-art methods on the GTEA Gaze and GTEA Gaze+ datasets.
Minlong Lu, Ze-Nian Li, Yueming Wang 0001, Gang Pan 0001
IEEE Trans. Image Process.2
2017 Recognition of facial expressions based on salient geometric features and support vector machines
Deepak Ghimire, Joonwhoan Lee, Ze-Nian Li, Sunghwan Jeong
Multim. Tools Appl.3
2017 An efficient temporal distortion measure of videos based on spacetime texture
Peng Peng 0003, Danping Liao, Ze-Nian Li
Pattern Recognit.3
2016 Learning Contextual Dependencies for Optical Flow with Recurrent Neural Networks
Minlong Lu, Zhiwei Deng, Ze-Nian Li
ACCV (4)3
2016 Semisupervised manifold learning for color transfer between multiview images
abstract
In multiview image stitching, the colors of images in a scene might vary when images are taken under different illumination or camera settings. A common way to produce a seamless stitched image is to transform the colors of a target image to match that of a source image. In this paper we present a color transfer method based on two premises: first, pixels in the generated image should have similar colors with their corresponding pixels in the source image. Second, pixels with similar colors should still have similar colors after color transfer. Our method can be considered as a semisupervised manifold learning approach, where the corresponding pixels of the input images serve as the labeled data. Our goal is to learn a final image which not only shares the same colors with the source image but also has the same image structure with the target image. While manifold learning methods aim to find an embedded space to represent the data with minimum structure loss, the proposed method further constrains the solution space using the labeled data. This paper introduces a parametric linear method and a nonparametric nonlinear method to tackle different types of color changes. Experimental results show the effectiveness of our methods both quantitatively and qualitatively.
Danping Liao, Yuntao Qian, Ze-Nian Li
ICPR3
2016 Object detection using boosted local binaries
Ze-Nian Li
Pattern Recognit.2
2015 Basis mapping based boosting for object detection
abstract
We propose a novel mapping method to improve the training accuracy and efficiency of boosted classifiers for object detection. The key step of the proposed method is a non-linear mapping on original samples by referring to the basis samples before feeding into the weak classifiers, where the basis samples correspond to the hard samples in the current training stage. We show that the basis mapping based weak classifier is an approximation of kernel weak classifiers while keeping the same computation cost as linear weak classifiers. As a result, boosting with such weak classifiers is more effective. In this paper, two different non-linear mappings are shown to work well. We adopt the LogitBoost algorithm to train the weak classifiers based on the Histogram of Oriented Gradient descriptor (HOG). Experimental results show that the proposed approach significantly improves the detection accuracy and training efficiency of the boosted classifier. It also achieves high performance on public datasets for both pedestrian detection and general object detection tasks.
Ze-Nian Li
CVPR2
2015 Object Detection Using Generalization and Efficiency Balanced Co-Occurrence Features
abstract
In this paper, we propose a high-accuracy object detector based on co-occurrence features. Firstly, we introduce three kinds of local co-occurrence features constructed by the traditional Haar, LBP, and HOG respectively. Then the boosted detectors are learned, where each weak classifier corresponds to a local image region with a co-occurrence feature. In addition, we propose a Generalization and Efficiency Balanced (GEB) framework for boosting training. In the feature selection procedure, the discrimination ability, the generalization power, and the computation cost of the candidate features are all evaluated for decision. As a result, the boosted detector achieves both high accuracy and good efficiency. It also shows performance competitive with the state-of-the-art methods for pedestrian detection and general object detection tasks.
Ze-Nian Li
ICCV2
2015 Object recognition based on deformable edge set
abstract
We aim to solve the object recognition problem by a novel contour feature called Deformable Edge Set (DES). The DES consists of several Deformable Edge Features (DEF), which is deformed from an edge template to the actual object contour according to the distribution model of pixels. Then the DES is constructed based on the combination of DEF, where the arrangement and the deformable parameters are learned in a subspace. The RealAdaBoost algorithm is further utilized to select meaningful DES to localize the object. Experimental results show that the proposed approach not only locates the object bounding boxes but also captures the object contours well. It also achieves performance competitive with the commonly-used algorithms.
Ze-Nian Li
ICIP2
2015 Object tracking using structure-aware binary features
abstract
Object tracking is one of the most important components in numerous applications of computer vision. In this paper, the target is represented by a series of binary patterns, where each binary pattern consists of several rectangle pairs in variable size and location. As complementary to traditional binary descriptors, these patterns are extracted in both the intensity domain and the gradient domain. In the tracking process, the RealAdaBoost algorithm is adopted frame by frame to select the meaningful patterns while considering the discriminative ability and the robustness. This is achieved by a penalty term based on the classification margin and structural diversity. As a result, the features good at describing the target and robust to noises will be selected. Experimental results on 10 challenging video sequences demonstrate that the tracking accuracy is significantly improved compared to traditional binary descriptors. It also achieves competitive results with the commonly-used algorithms.
Ze-Nian Li
ICME2
2015 Whole-body humanoid robot imitation with pose similarity evaluation
Jie Lei 0002, Mingli Song, Ze-Nian Li, Chun Chen 0001
Signal Process.3
2015 Continuous Depth Map Reconstruction From Light Fields
abstract
In this paper, we investigate how the recently emerged photography technology--the light field--can benefit depth map estimation, a challenging computer vision problem. A novel framework is proposed to reconstruct continuous depth maps from light field data. Unlike many traditional methods for the stereo matching problem, the proposed method does not need to quantize the depth range. By making use of the structure information amongst the densely sampled views in light field data, we can obtain dense and relatively reliable local estimations. Starting from initial estimations, we go on to propose an optimization method based on solving a sparse linear system iteratively with a conjugate gradient method. Two different affinity matrices for the linear system are employed to balance the efficiency and quality of the optimization. Then, a depth-assisted segmentation method is introduced so that different segments can employ different affinity matrices. Experiment results on both synthetic and real light fields demonstrate that our continuous results are more accurate, efficient, and able to preserve more details compared with discrete approaches.
Jianqiao Li, Minlong Lu, Ze-Nian Li
IEEE Trans. Image Process.3
2014 Age Estimation Based on Complexity-Aware Features
Ze-Nian Li
ACCV (1)2
2014 Object detection using edge histogram of oriented gradient
abstract
In this paper, we address the object detection problem by a proposed gradient feature, the Edge Histogram of Oriented Gradient (Edge-HOG). Edge-HOG consists of several blocks arranged along a line or an arc, which is designed to describe the edge pattern. In addition, we propose a new feature extraction method, which extracts the structural information based on the gravity centers as complementary to traditional gradient histograms. As a result, the proposed Edge-HOG not only reflects the local shape information of objects, but also captures more significant appearance information. Experimental results show that the proposed approach significantly improves both the detection accuracy and the convergence speed compared to the traditional HOG feature. It also achieves performance competitive with some commonly-used methods on pedestrian detection and car detection tasks.
Ze-Nian Li
ICIP2
2014 Boosted local binaries for object detection
abstract
We propose a novel binary feature for object detection encoding local neighbor patterns of different sizes and locations. Each region pair of the proposed feature is selected by RealAdaBoost algorithm with a penalty term on the structure diversity. As a result, useful features that are good at describing specific objects will be chosen to build the classifier. Moreover, the encoding scheme is applied in both the gradient domain and the intensity domain, which is complementary to standard binary features (e.g. LBP and LAB). The proposed method was tested using the CMU-MIT frontal face dataset, INRIA pedestrian dataset, and UIUC car dataset respectively. Experimental results show that the proposed method outperforms traditional binary features LBP and LAB, which contributes to a significant improvement on detection accuracy and converges 2 times faster. It also achieves comparable performance with some state-of-the-art algorithms.
Ze-Nian Li
ICME2
2014 Humanoid Robot Imitation with Pose Similarity Metric Learning
abstract
Imitation is considered to be a kind of social learning that allows the transfer of information, actions, behaviours, etc. Whereas current robots are unable to perform as many tasks as human, it is a natural way for them to learn by imitations, just as human does. With the humanoid robots being more intelligent, the field of robot imitation has getting noticeable advance. In this paper, we focus on the pose imitation between a human and a humanoid robot and learning a similarity metric between human pose and robot pose. In contrast to recent approaches that capture human data using expensive motion captures or only imitate the upper body movements, our framework adopts a Kinect instead and can deal with complex, whole body motions by keeping both single pose balance and pose sequence balance. Meanwhile, different from previous work that employs subjective evaluation, we propose a pose similarity metric based on the shared structure of the motion spaces of human and robot. The qualitative and quantitative experimental results demonstrate a satisfactory imitation performance and indicate that the proposed pose similarity metric is discriminative.
Jie Lei 0002, Mingli Song, Ze-Nian Li, Chun Chen 0001, Xianghua Xu, Shiliang Pu
ICPR3
2014 Gender Recognition Using Complexity-Aware Local Features
abstract
We propose a gender classifier using two types of local features, the gradient features which have strong discrimination capability on local patterns, and the Gabor wavelets which reflect the multi-scale directional information. The Real Ad a Boost algorithm with complexity penalty term is applied to choose meaningful regions from human face for feature extraction, while balancing the discriminative capability and the computation cost at the same time. Linear SVM is further utilized to train a gender classifier based on the selected features for accuracy evaluation. Experimental results show that the proposed approach outperforms the methods using single feature. It also achieves comparable accuracy with the state-of-the-art algorithms on both controlled datasets and real-world datasets.
Ze-Nian Li
ICPR2
2014 Audio feature reduction and analysis for automatic music genre classification
abstract
Multimedia database retrieval is growing at a fast rate thereby subsequent increase in the popularity of online retrieval system. The large datasets are major challenges for searching, retrieving, and organizing the music content. Therefore, there is a need of robust automatic music genre classification method for organizing these music data into different classes according to the certain viable information. There are two fundamental components to be considered for genre classification namely audio feature extraction and classifier design. In this paper, diverse audio features set have been proposed to characterize the music contents precisely. The feature sets belong to four different groups, i.e. dynamic, rhythm, spectral, and harmony. From the features, five different statistical parameters are considered as representatives, including up to the 4thorder central moments of each feature, and covariance components. Ultimately, significant numbers of representative attributes are controlled by MRMR algorithm. The algorithm calculates the score level of all feature attributes and orders them. The high score feature attributes are only considered for genre classification. Moreover, we can visualize that which audio features and which of the different statistical parameters derived from them are important for genre classification. Among them, mel frequency cepstral coefficients (MFCCs) have higher scored level than other feature attributes. Furthermore, MRMR does not transform the feature value like as principal component analysis (PCA). Besides these, the comparison has been made based on classification accuracy between two-dimensionality reduction methodologies using support vector machine (SVM). The classification accuracy of MRMR feature reduction set outperforms than PCA. The overall classification is also higher than other existing state-of-the-art of frame base methods.
Babu Kaji Baniya, Joonwhoan Lee, Ze-Nian Li
SMC3
2014 General-purpose image quality assessment based on distortion-aware decision fusion
Peng Peng 0003, Ze-Nian Li
Neurocomputing2
2013 Continuous depth map reconstruction from light fields
abstract
Light field analysis recently received growing interest, since its rich structure information benefits many computer vision tasks. This paper presents a novel method to reconstruct continuous depth maps from light field data. Conventional approaches usually treat depth map reconstruction as an optimization problem with discrete labels. On the contrary, our proposed method can obtain continuous depth maps by solving a linear system, which preserves richer details compared with conventional discrete approaches. Structure tensor is employed to extract raw depth information and corresponding confidence levels from the light field data. We introduce a method to reduce the adverse effect of unreliable local estimations, which helps to get rid of errors in specular areas and edges where depth values are discontinuous. Experiments on both synthetic and real light field data demonstrate the effectiveness of the proposed method.
Jianqiao Li, Ze-Nian Li
ICME2
2013 Efficient video quality assessment based on spacetime texture representation
abstract
Most existing video quality metrics measure temporal distortions based on optical-flow estimation, which typically has limited descriptive power of visual dynamics and low efficiency. This paper presents a unified and efficient framework to measure temporal distortions based on a spacetime texture representation of motion. We first propose an effective motion-tuning scheme to capture temporal distortions along motion trajectories by exploiting the distributive characteristic of the spacetime texture. Then we reuse the motion descriptors to build a self-information based spatiotemporal saliency model to guide the spatial pooling. At last, a comprehensive quality metric is developed by combining the temporal distortion measure with spatial distortion measure. Our method demonstrates high efficiency and excellent correlation with the human perception of video quality.
Peng Peng 0003, Kevin J. Cannons, Ze-Nian Li
ACM Multimedia3
2013 Optimizing Nondecomposable Loss Functions in Structured Prediction
abstract
We develop an algorithm for structured prediction with nondecomposable performance measures. The algorithm learns parameters of Markov Random Fields (MRFs) and can be applied to multivariate performance measures. Examples include performance measures such as Fβ score (natural language processing), intersection over union (object category segmentation), Precision/Recall at k (search engines), and ROC area (binary classifiers). We attack this optimization problem by approximating the loss function with a piecewise linear function. The loss augmented inference forms a Quadratic Program (QP), which we solve using LP relaxation. We apply this approach to two tasks: object class-specific segmentation and human action retrieval from videos. We show significant improvement over baseline approaches that either use simple loss functions or simple scoring functions on the PASCAL VOC and H3D Segmentation datasets, and a nursing home action recognition dataset.
Mani Ranjbar, Tian Lan 0006, Yang Wang 0003, Stephen N. Robinovitch, Ze-Nian Li, Greg Mori
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 Regularization of the structural similarity index based on preservation of edge direction
abstract
The goal of this paper is to improve the performance of the SSIM indexes while retaining their computational efficiency. To this end, we first design an edge-quality term based on the preservation of edge direction, and then adaptively combine it with the SSIM indexes, yielding the regularized SSIM indexes. The proposed method is based on two assumptions: (1) as the quality of a distorted image declines, the human vision system (HVS) is more likely to judge its quality based on the difficulty of recognizing its content; and (2) the preservation of edge direction is a good measurement of this difficulty. Extensive evaluation shows that the regularized SSIM indexes achieve comparable performance to the state-of-the-art method while requiring much less computation time.
Peng Peng 0003, Ze-Nian Li
SMC2
2012 A large margin framework for single camera offline tracking with hybrid cues
Bahman Yari Saeed Khanloo, Ferdinand Stefanus, Mani Ranjbar, Ze-Nian Li, Nicolas Saunier, Tarek Sayed, Greg Mori
Comput. Vis. Image Underst.4
2010 Learning image similarities via Probabilistic Feature Matching
abstract
In this paper, we propose a novel image similarity learning approach based on Probabilistic Feature Matching (PFM).We consider the matching process as the bipartite graph matching problem, and define the image similarity as the inner product of the feature similarities and their corresponding matching probabilities, which are learned by optimizing a quadratic formulation. Further, we prove that the image similarity and the sparsity of the learned matching probability distribution will decrease monotonically with the increase of parameter C in the quadratic formulation where C ≥ 0 is a pre-defined data-dependent constant to control the sparsity of the distribution of a feature matching probability. Essentially, our approach is the generalization of a family of similarity matching approaches. We test our approach on Graz datasets for object recognition, and achieve 89.4% on Graz-01 and 87.4% on Graz-02, respectively on average, which outperform the state-of-the-art.
Ze-Nian Li, Mark S. Drew
ICIP2
2010 AdaMKL: A Novel Biconvex Multiple Kernel Learning Approach
abstract
In this paper, we propose a novel large-margin based approach for multiple kernel learning (MKL) using biconvex optimization, called Adaptive Multiple Kernel Learning (AdaMKL). To learn the weights for support vectors and the kernel coefficients, AdaMKL minimizes the objective function alternately by learning one component while fixing the other at a time, and in this way only one convex formulation needs to be solved. We also propose a family of biconvex objective functions with an arbitrary ℓp-norm (p ≥ 1) of kernel coefficients. As our experiments show, AdaMKL performs comparably with state-of-the-art convex optimization based MKL approaches, but its learning is much simpler and faster.
Ze-Nian Li, Mark S. Drew
ICPR2
2010 Action Detection in Cluttered Video With Successive Convex Matching
abstract
We propose a novel successive convex matching method for human action detection in cluttered video. Human actions are represented as sequences of poses, and specific actions are detected by matching pose sequences. Since we represent actions as the evolution of poses and shapes, the proposed method can detect actions in videos that involve fast camera motions. Template sequence to video registration is nonlinear and highly nonconvex. Instead of directly solving the hard problem, our method convexifies it into a sequence of linear programs and refines the matching by successive trust region shrinkage. The proposed scheme further simplifies the linear programs by representing the target point space with a small set of basis points. The low complexity of the proposed method enables it to search efficiently in a large range. Experiments show that successive convex matching can robustly match a sequence of coupled shape templates simultaneously to target sequences and effectively detect specific actions in cluttered videos.
Hao Jiang 0007, Mark S. Drew, Ze-Nian Li
IEEE Trans. Circuits Syst. Video Technol.3
2009 A fast method for classifying surface textures
abstract
Surface texture classification is an important aspect of computer vision and a well studied problem. In this paper, we greatly increase speed for texture classification while maintaining accuracy. We take inspiration form past work and propose a new method for texture classification which is extremely fast due to the low dimensionality of our feature space. We extract distinctive features at a very early stage, thus removing the dependency on expensive and sensitive operations such as k-Means clustering which is used by much work in this field of research. We present experimental results on the Colombia-Utrecht Reflectance and Texture Database (CURET), to date the most challenging dataset for texture classification, and show that our method achieves comparable classification accuracy in comparison with the state-of-the-art, but at a 10-fold increased speed.
Muntaseer Salahuddin, Mark S. Drew, Ze-Nian Li
ICASSP3
2009 Feature correspondence with constrained global spatial structures
abstract
In this paper, we consider the feature correspondence task as a graph matching problem. Our approach tends to maximize a similarity objective function, which consists of not only the feature vectors but also their corresponding constrained global spatial structures, by a new polynomial-time approximate optimization algorithm. This algorithm allows every node in a smaller graph to potentially be linked with any node in a larger graph, and thus it can handle one-to-one, many-to-one, and no match cases. Especially, our approach does not necessarily require a training set. We test on the “hotel” and “house” sequences. Matching a pair of frames takes on average 1.24 and 1.22 seconds respectively using a Matlab implementation without any optimization (over an order of magnitude speedup compared to [1]), and with a 2-frame interval our errors are merely 0.07% and 0.09% respectively. Even going up to a 25-frame gap, errors are only 5.66% and 5.00% respectively.
Ze-Nian Li, Mark S. Drew
ICIP2
2009 Image trimming via saliency region detection and iterative feature matching
abstract
Detection of saliency regions in images is useful for object based image understanding and object localization. In our work, we investigate a saliency region detection algorithm based on the human visual attention (HVA) model. In the first phase, we use mutual information and probability-of-boundary (PoB) for color saliency and edge detection respectively to filter SURF (speeded up robust features) key feature points found from the image. For the second phase, bipartite feature matching is deployed for further keypoint selection. We perform the two-phase keypoint filtering iteratively and give selected keypoints different weights for their importance. The final trimmed image is a rectangle region which approximates the distribution of remaining keypoints. We conduct our experiments on Corel Photo Library and MIT-CSAIL Objects and Scenes Database and demonstrate the effectiveness of our proposed algorithm.
Ze-Nian Li
ICME2
2008 A study of image-based music composition
abstract
Visual and auditory forms have some noticeable associations that can inspire similar cognitive and aesthetical experiences. This paper presents a study on the possibilities of applying low-level visual-auditory associations to music generation. A few methods are implemented to directly convert visual features extracted from images into musical elements (pitch, duration, and chord). Some initial results suggest a high potential of composing interesting music along this direction. Possible application of this study include getting new ideas for music composers and generating accompanying music in various contexts.
Ze-Nian Li
ICME2
2008 Automatic object extraction and reconstruction in active video
Ze-Nian Li
Pattern Recognit.2
2008 Activity recognition via user-trace segmentation
abstract
A major issue of activity recognition in sensor networks is automatically recognizing a user's high-level goals accurately from low-level sensor data. Traditionally, solutions to this problem involve the use of a location-based sensor model that predicts the physical locations of a user from the sensor data. This sensor model is often trained offline, incurring a large amount of calibration effort. In this article, we address the problem using a goal-based segmentation approach, in which we automatically segment the low-level user traces that are obtained cheaply by collecting the signal sequences as a user moves in wireless environments. From the traces we discover primitive signal segments that can be used for building a probabilistic activity model to recognize goals directly. A major advantage of our algorithm is that it can reduce a significant amount of human effort in calibrating the sensor data while still achieving comparable recognition accuracy. We present our theoretical framework for activity recognition, and demonstrate the effectiveness of our new approach using the data collected in an indoor wireless environment.
Jie Yin 0001, Qiang Yang 0001, Dou Shen, Ze-Nian Li
ACM Trans. Sens. Networks4
2007 A New Energy Function for Segmentation and Compression
abstract
We propose a new DCT-based energy function EDCTfor object-oriented segmentation and compression. By minimizing EDCTthe best possible split of the image blocks can be found which leads to better segmentation and compression. Consistent motion vectors can be obtained by using an extended 3D energy function, in the spatio-temporal domain, which measures block motion over multiple frames. Tests on images and videos show promising results, where still image compression achieves over 15% improvement over JPEG and video compression achieves similar improvement over MPEG-2.
M. King, Zinovi Tauber, Ze-Nian Li
ICME3
2007 Matching by Linear Programming and Successive Convexification
abstract
We present a novel convex programming scheme to solve matching problems, focusing on the challenging problem of matching in a large search range and with cluttered background. Matching is formulated as metric labeling with L1 regularization terms, for which we propose a novel linear programming relaxation method and an efficient successive convexification implementation. The unique feature of the proposed relaxation scheme is that a much smaller set of basis labels is used to represent the original label space. This greatly reduces the size of the searching space. A successive convexification scheme solves the labeling problem in a coarse to fine manner. Importantly, the original cost function is reconvexified at each stage, in the new focus region only, and the focus region is updated so as to refine the searching result. This makes the method well-suited for large label set matching. Experiments demonstrate successful applications of the proposed matching scheme in object detection, motion estimation, and tracking.
Hao Jiang 0007, Mark S. Drew, Ze-Nian Li
IEEE Trans. Pattern Anal. Mach. Intell.3
2007 Review and Preview: Disocclusion by Inpainting for Image-Based Rendering
abstract
Image-based rendering takes as input multiple images of an object and generates photorealistic images from novel viewpoints. This approach avoids explicitly modeling scenes by replacing the modeling phase with an object reconstruction phase. Reconstruction is achieved in two possible ways: recovering 3-D point locations using multiview stereo techniques, or reasoning about consistency of each voxel in a discretized object volume space. The most challenging problem for image-based reconstruction is the presence of occlusions. Occlusions make reconstruction ambiguous for object parts not visible in any input image. These parts must be reconstructed in a visually acceptable way. This paper both reviews image inpainting and argues that inpainting can provide not only attractive reconstruction but also a framework for increasing the accuracy of depth recovery. Digital image inpainting refers to any methods that fill-in holes of arbitrary topology in images so that they seem to be a part of the original image. Available methods are broadly classified as structural inpainting or textural inpainting. Structural inpainting reconstructs using prior assumptions and boundary conditions, while textural inpainting considers only the available data from texture exemplars or other templates. Of particular particular interest is research on structural inpainting applied to 3-D models, emphasizing its effectiveness for disocclusion.
Zinovi Tauber, Ze-Nian Li, Mark S. Drew
IEEE Trans. Syst. Man Cybern. Part C2
2006 Successive Convex Matching for Action Detection
abstract
We propose human action detection based on a successive convex matching scheme. Human actions are represented as sequences of postures and specific actions are detected in video by matching the time-coupled posture sequences to video frames. The template sequence to video registration is formulated as an optimal matching problem. Instead of directly solving the highly non-convex problem, our method convexifies the matching problem into linear programs and refines the matching result by successively shrinking the trust region. The proposed scheme represents the target point space with small sets of basis points and therefore allows efficient searching. This matching scheme is applied to robustly matching a sequence of coupled binary templates simultaneously in a video sequence with cluttered backgrounds.
Hao Jiang 0007, Mark S. Drew, Ze-Nian Li
CVPR (2)3
2006 Unsupervised Discovery of Action Classes
abstract
In this paper we consider the problem of describing the action being performed by human figures in still images. We will attack this problem using an unsupervised learning approach, attempting to discover the set of action classes present in a large collection of training images. These action classes will then be used to label test images. Our approach uses the coarse shape of the human figures to match pairs of images. The distance between a pair of images is computed using a linear programming relaxation technique. This is a computationally expensive process, and we employ a fast pruning method to enable its use on a large collection of images. Spectral clustering is then performed using the resulting distances. We present clustering and image labeling results on a variety of datasets.
Yang Wang 0003, Hao Jiang 0007, Mark S. Drew, Ze-Nian Li, Greg Mori
CVPR (2)4
2006 Detecting Human Action in Active Video
abstract
We propose a novel scheme to detect human actions in active video. Active videos such as movies or sports broadcasting are taken purposively by "clever" photographers. They are object and action oriented and usually involve complex camera motions. Detecting actions in active videos is both important and challenging. We study a three-step scheme to detect complex human actions in such videos. The proposed method first locates potential objects and removes clutter with a composite filter scheme. The detected object candidates in successive frames are then associated to form object trajectories based on a consistent labeling formulation, and solved with belief propagation. Finally, specific human actions are detected in video with a linear programming matching approach that can efficiently deal with matching problems having a large target point set. The proposed method has been successfully applied in action detection for general videos and TV hockey games
Hao Jiang 0007, Ze-Nian Li, Mark S. Drew
ICME2
2005 Activity Recognition through Goal-Based Segmentation
Jie Yin 0001, Dou Shen, Qiang Yang 0001, Ze-Nian Li
AAAI4
2005 Human Posture Recognition with Convex Programming
abstract
We present a novel human posture recognition method us ing convex programming based matching schemes. Instead of trying to segment the object from the background, we develop a novel multistage linear programming scheme to locate the target by searching for the best matching region based on an automatically acquired graph template. The linear programming based visual matching scheme gener ates relatively dense matching patterns and thus presents a key for robust object matching and human posture recogni tion. By matching distance transformations of edge maps, the proposed scheme is able to match figures with large ap pearance changes. We further present object recognition methods based on the similarity of the exemplar with the matching target. The proposed scheme can also be used for recognizing multiple targets in an image. Experiments show promising results for recognizing human postures in clut tered environments.
Hao Jiang 0007, Ze-Nian Li, Mark S. Drew
ICME2
2004 Optimizing Motion Estimation with Linear Programming and Detail-Preserving Variational Method
Hao Jiang 0007, Ze-Nian Li, Mark S. Drew
CVPR (1)2
2004 Active video object extraction
abstract
This paper addresses the problem of intelligently extracting objects from videos. Our method assumes that the video making process is purposive and attempts to extract those objects that the original authors of the videos intended to capture. We accomplish this by analyzing three types of actions of the author (saccadic movements, smooth pursuits, and multi-baseline pursuits) in an active vision framework using dense 2D disparity vectors computed from successive frames of the video. We demonstrate the effectiveness of our algorithm using real video sequences.
Ze-Nian Li
ICME2
2004 A survey of motion-parallax-based 3-D reconstruction algorithms
abstract
The task of recovering three-dimensional (3-D) geometry from two-dimensional views of a scene is called 3-D reconstruction. It is an extremely active research area in computer vision. There is a large body of 3-D reconstruction algorithms available in the literature. These algorithms are often designed to provide different tradeoffs between speed, accuracy, and practicality. In addition, even the output of various algorithms can be quite different. For example, some algorithms only produce a sparse 3-D reconstruction while others are able to output a dense reconstruction. The selection of the appropriate 3-D reconstruction algorithm relies heavily on the intended application as well as the available resources. The goal of this paper is to review some of the commonly used motion-parallax-based 3-D reconstruction techniques and make clear the assumptions under which they are designed. To do so efficiently, we classify the reviewed reconstruction algorithms into two large categories depending on whether a prior calibration of the camera is required. Under each category, related algorithms are further grouped according to the common properties they share.
Jason Z. Zhang, Q. M. Jonathan Wu, Ze-Nian Li
IEEE Trans. Syst. Man Cybern. Part C4
2003 A modified space frequency decomposition algorithm for visual motion
abstract
In this paper, we present a method to decompose visual motion represented by computed optical flow using an over-complete dictionary of space frequency atoms. This is accomplished by a modified matching pursuit procedure which successively selects space and frequency localized atoms from this dictionary. In effect, this decomposition reveals the structure of optical flow field such that areas with motion at different spatial location and scale are identified. In addition, our method maintains the original resolution of the image when performing the decomposition at each scale. This minimizes the amount of the information lost in the decomposition process. We perform experiments using both synthetic and real image data sets to demonstrate the effectiveness of the proposed method.
Ze-Nian Li
ICME3
2002 Illumination color covariant locale-based visual object retrieval
Mark S. Drew, Ze-Nian Li, Zinovi Tauber
Pattern Recognit.2
2002 Spatial-temporal joint probability images for video segmentation
Ze-Nian Li, Mark S. Drew
Pattern Recognit.1
2001 Locale-Based Multiple Cue Algorithm For Object Segmentation
abstract
This paper proposes a Locale-based Multiple Cue (LMC) algorithm to solve the problem of segmenting foregroundmoving objects from the background scene. The major cue used for object segmentation is the motion information obtained from a novel locale-based motion estimation and clustering algorithm. At first, locales are classified according to color and intensity. The motion estimation is applied on the tiles of 16 x 16 pixels, which are the building blocks of locales. After motion estimation, camera motions are detected using a 2D affine motion model. Then the locales are grown from the tile level to the frame level in a pyramidal way using motion, color, and centroid variance constraint. The resulting locales, with tiles homogeneous in motion and color, are post-processed to recover the object boundary. Experimental results show that LMC combines temporal and spatial information in a graceful way, which enables it to segment the moving objects under different camera motions. Future work includes object tracking over multiple frames and utilization of texture information.
Ze-Nian Li
ICME2
2001 On Active Camera Control and Camera Motion Recovery with Foveate Wavelet Transform
abstract
In this paper, a new variable resolution technique, foveate wavelet transform (FWT), is proposed to represent digital images in an effort to efficiently represent visual data. Compared to existing variable resolution techniques, the strength of the proposed scheme encompasses its linearity preservation, orientation selectivity, and flexibility while supporting interesting behaviors resembling the animate vision system. The linearity preservation of the FWT is due to the fact that only low and/or high-pass filterings are carried out in different regions of an image in the transform. The orientation selectivity indicates the fact that details along the horizontal, vertical, and diagonal directions are readily available in the FWT representation. The flexibility of this new representation technique is witnessed by the readiness of its extensions to represent foveae of different number, shape, and locations. To demonstrate the efficacy of the FWT, two applications are presented. First, an FWT-based active camera control scheme is developed, where the computer can move a camera to track the moving object in the scene. Second, an FWT-based method purporting to recover pan/tilt/zoom camera movements from video clips is developed. Experiments of these two applications have shown encouraging performances.
Ze-Nian Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Video Dissolve and Wipe Detection via Spatio-Temporal Images of Chromatic Histogram Differences
abstract
Gradual transitions represent a challenging problem for temporal segmentation of video. Here we present two new features for detecting these. Ngo et al. set out a method for edge detection in spatio-temporal images made out of the central column (or row, or diagonal) of a video. A wipe generates a diagonal edge in such an image. In this paper we make use of all available pixels to generate spatio-temporal images. For each column of the frame (using only the DC values from a video MPEG), we form a 2D histogram based on chromaticity, and then intersect that histogram with that of the previous frame (one or several frames earlier). The result is an image in which cuts and wipes appear as very strong edges, almost 1s in a background of zeroes. Dissolves require another approach; here we extend a color-distance based histogram metric due to Hafner et al. (1995) by applying the method to 2D Cb-Cr histograms and changing the definition so that the the metric displays a near-constant value during a dissolve, and zero elsewhere. We show results on videos that include fast subject motion and camera movements.
Mark S. Drew, Ze-Nian Li
ICIP2
2000 Spatio-Temporal Joint Probability Images for Video Segmentation
abstract
In this paper we propose a novel video segmentation algorithm that is capable of reliably detecting both types of dissolves in addition to cuts and wipes based on the analysis of the spatio-temporal behavior of joint probability images (JPIs). Experimental results for identifying scene transitions in various video clips that contain moderate and significant amounts of movements are demonstrated.
Ze-Nian Li
ICIP1
2000 Locale-Based Visual Object Retrieval under Illumination Change
abstract
Providing a user with an effective image search engine has been a very active research area. A search by an object model is considered to be one of the most desirable and yet difficult tasks. An added difficulty is that objects can be photographed under different lighting conditions. We have developed a feature localization scheme that finds a set of locales in an image. We make use of a diagonal model for illumination change and obtain a candidate set of lighting transformation coefficients in chromaticity space. For each pair of coefficients, elastic correlation is performed, which is a form of correlation of locale colors. A least square minimization for pose estimation is then applied, followed by a process of texture support and shape verification. Tests on a database of over 1,400 images and video clips show promising image retrieval results. Moreover it has been shown that the method is capable of recovering lighting changes.
Zinovi Tauber, Ze-Nian Li, Mark S. Drew
ICPR2
1999 On active camera control with foveate wavelet transform
abstract
In this paper, a new variable resolution technique, foveate wavelet transform (FWT), is introduced to represent video frames in an effort to emulate the animate vision systems. Compared to the existing variable resolution techniques, the strength of the proposed scheme encompasses its flexibility and conciseness while supporting interesting behavior resembling the animate vision system. With the FWT as the representation, an efficient scheme to achieve purposive vision is developed. The foveate potential motion area is first labeled, the camera motion is then determined by following the labeled moving object. Experiments based on our computer-controlled pan-tilt-zoom camera have demonstrated its efficacy in real-time active camera control.
Ze-Nian Li
IROS2
1999 Reciprocal-Wedge Transform in Active Stereo
abstract
The Reciprocal-Wedge Transform (RWT) facilitates space-variant image representation. In this paper a V-plane projection method is presented as a model for imaging using the RWT. It is then shown that space-variant sensing with this new RWT imaging model is suitable for fixation control in active stereo that exhibits vergence and versional eye movements and scanpath behaviors. A computational interpretation of stereo fusion in relation to disparity limit in space-variant imagery leads to the development of a computational model for binocular fixation. The vergence-version movement sequence is implemented as an effective fixation mechanism using the RWT imaging. A fixation system is presented to show the various modules of camera control, vergence and version.
Ze-Nian Li
Int. J. Pattern Recognit. Artif. Intell.1
1999 Illumination Invariance and Object Model in Content-Based Image and Video Retrieval
Ze-Nian Li, Osmar R. Zaïane, Zinovi Tauber
J. Vis. Commun. Image Represent.1
1999 Illumination-invariant image retrieval and video segmentation
Mark S. Drew, Ze-Nian Li
Pattern Recognit.3
1999 An efficient two-pass MAP-MRF algorithm for motion estimation based on mean field theory
abstract
This paper presents a two-pass algorithm for estimating motion vectors from image sequences. In the proposed algorithm, the motion estimation is formulated as a problem of obtaining the maximum a posteriori in the Markov random field (MAP-MRF). An optimization method based on the mean field theory (MFT) is opted to conduct the MAP search. The estimation of motion vectors is modeled by only two MRFs, namely, the motion vector field and unpredictable field. Instead of utilizing the line field, a truncation function is introduced to handle the discontinuity between the motion vectors on neighboring sites. In this algorithm, a "double threshold" preprocessing pass is first employed to partition the sites into three regions, whereby the ensuing MPT-based pass for each MRF is conducted on one or two of the three regions. With this algorithm, no significant difference exists between the block-based and pixel-based MAP searches any more. Consequently, a good compromise between precision and efficiency can be struck with ease. To render our algorithm more resilient against noise, the mean absolute difference instead of mean square error is selected as the measure of difference, which is more reliable according to the knowledge of robust statistics. This is supported by our experimental results from both synthetic and real-world image sequences. The proposed two-pass algorithm is much faster than any other MAP-MRF motion estimation method reported in the literature so far.
Ze-Nian Li
IEEE Trans. Circuits Syst. Video Technol.2
1998 Illumination-Invariant Color Object Recognition via Compressed Chromaticity Histograms of Color-Channel-Normalized Images
abstract
Several color object recognition methods that are based on image retrieval algorithms attempt to discount changes of illumination in order to increase performance when test image illumination conditions differ from those that obtained when the image database was created. Here we extend the seminal method of Swain and Ballard to discount changing illumination. The new method is based on the first stage of the simplest color indexing method, which uses angular invariants between color image and edge image channels. That method first normalizes image channels, and then effectively discards much of the remaining information. Here we adopt the color-normalization stage as an adequate color constancy step. Further, we replace 3D color histograms by 2D chromaticity histograms. Treating these as images, we implement the method in a compressed histogram-image domain using a combination of wavelet compression and Discrete Cosine Transform (DCT) to fully exploit the technique of low-pass filtering for efficiency. Results are very encouraging, with substantially better performance than other methods tested. The method is also fast, in that the indexing process is entirely carried out in the compressed domain and uses a feature vector of only 36 or 72 values.
Mark S. Drew, Ze-Nian Li
ICCV3
1998 Foveate wavelet transform for camera motion recovery from videos
abstract
In this paper, a new variable resolution technique, foveate wavelet transform (FWT), is proposed to represent video frames in order to reduce the bandwidth storage requirement through emulating the animate visual systems. As its application, an FWT-based method purporting to recover camera motions from video clips is developed. With this method, the motion vectors are first estimated for the FWT representation of each frame using a foveate two-pass algorithm. The camera motions, namely, pan, tilt, zoom, and a mixture of them, are then recovered based on the behaviors of the dense motion fields. Encouraging experimental results have been witnessed.
Ze-Nian Li
ICPR2
1998 Efficient disparity-based gaze control with foveate wavelet transform
abstract
A new variable resolution technique-foveate wavelet transform (FWT)-is proposed in this paper to represent images in an effort to emulate the animate visual systems. With the FWT representation and a simplified configuration of the stereo rig, efficient disparity-based algorithms to achieve gaze control of the rig are developed. With these algorithms, the attention of the stereo rig can be centered around the object of most interest with desired resolution for both the static and moving objects. Experiments based on the algorithms have demonstrated their efficacies in the gaze control of the stereo rig.
Ze-Nian Li
IROS2
1998 MultiMediaMiner: A System Prototype for Multimedia Data Mining
abstract
Multimedia data mining is the mining of high-level multimedia information and knowledge from large multimedia databases. A multimedia data mining system prototype, MultiMediaMiner, has been designed and developed. It includes the construction of a multimedia data cube which facilitates multiple dimensional analysis of multimedia data, primarily based on visual content, and the mining of multiple kinds of knowledge, including summarization, comparison, classification, association, and clustering.
Osmar R. Zaïane, Jiawei Han 0001, Ze-Nian Li, Sonny Han Seng Chee, Jenny Chiang
SIGMOD Conference3
1998 On illumination invariance in color object recognition
Mark S. Drew, Ze-Nian Li
Pattern Recognit.3
1998 From NOMAD to explorer: active object recognition on mobile robots
Peter Gvozdjak, Ze-Nian Li
Pattern Recognit.2
1997 Motion compensation in color video with illumination variations
abstract
Within an image sequence, it is not unusual that in certain local areas there exist some motions as well as illumination variations. A new method, block-matching with illumination variations (BMIV), is proposed for motion compensation in color video. For blocks with motion and illumination variations, the proposed method can estimate the motion vector and generate the associated 3/spl times/3 illumination matrix which will produce the minimal mean square error. Experiments using the proposed method show good performance.
Ze-Nian Li
ICIP (3)2
1997 An enhancement to MRMC scheme in video compression
abstract
Zhang and Zafar (see ibid., vol.2, no.3, p.285-96, 1992) proposed a video compression scheme based on the wavelet representation and multiresolution motion compensation (MRMC). An additional masking module is created to further enhance its efficiency. Specifically, between the modules of wavelet decomposition and MRMC, the masking module is inserted which constructs binary images based on the difference of the wavelet coefficients. The binary images serve as masks to facilitate a more efficient motion compensation. Experiments show that the processing time could be significantly reduced from that required for a full search in the original MRMC algorithm.
Ze-Nian Li
IEEE Trans. Circuits Syst. Video Technol.2
1996 Camera model for reciprocal-wedge transform
Ze-Nian Li
Image Vis. Comput.2
1996 Analysis of disparity gradient based cooperative stereo
abstract
This paper argues that the disparity gradient subsumes various constraints for stereo matching, and can thus be used as the basis of a unified cooperative stereo algorithm. Traditionally, selection of the neighborhood support function (NSF) in cooperative stereo was left as a heuristic exercise. We present an analysis and evaluation of three families of NSFs based on the disparity gradient. It is shown that an exponential decay function with a conveniently selectable parameter is well behaved in that it yields the least error, converges steadily, and produces correctly located weak-winners. The discovery of the well-behaved function facilitates the success of the disparity gradient based approach. It is suggested that this function will help a two-pass algorithm in resolving the dilemma of surface continuity and discontinuity/occlusion. In our experiments, the unified cooperative stereo-matching algorithm is tested on random-dot stereograms containing opaque and transparent surfaces. It is also shown to be applicable to both area matching and contour matching in real-world images.
Ze-Nian Li, Gongzhu Hu
IEEE Trans. Image Process.1
1995 Reciprocal-Wedge Transform for Space-Variant Sensing
abstract
The reciprocal-wedge transform (RWT) is presented as an alternative to the log-polar transform which has been a popular model for space-variant sensing in computer vision. The RWT facilitates an anisotropic variable resolution. Unlike the log-polar, its variable resolution is predominantly in one dimension. Consequently, the RWT preserves linearity of lines and translations in the original image. In this paper, a concise matrix representation of the RWT is presented. Its properties in geometrical transformations and data reduction are described. A projective model for the transform and a potential hardware RWT camera design are also illustrated. As examples of initial applications, the RWT is used for finding road directions in navigation, and for recovering depth in motion stereo. Two types of motion stereo are presented, namely the longitudinal and lateral motion stereo. In all cases, the RWT images offer much reduced and adequate data owing to the variable resolution. Preliminary experimental results from test images of road-vehicle navigation and moving objects on a miniature assembly line are demonstrated.>
Ze-Nian Li
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Reciprocal-Wedge Transform in Motion Stereo
abstract
The reciprocal-wedge transform (RWT) facilitates space-variant sensing which enables effective use of variable-resolution data and the reduction of total amount of the sensory data. This paper presents two motion stereo methods that exploit the important properties of the RWT, i.e., the anisotropic variable resolution and the preservation of linear features. It is shown that the RWT is suitable for the correspondence process in both lateral and longitudinal motion stereo which deal with disparities of corresponding features along epipolar lines. Multiple frames of motion stereo images are employed to improve precision and error rate of the depth recovery. In the lateral motion stereo the RWT is applied in both space and time domains to transform the x-t epipolar plane in ordinary motion stereo images into a new /spl omega/-/spl tau/ epipolar plane. In the longitudinal motion stereo, the reciprocity of the RWT restores the nonlinearity in the original x-t epipolar plane. Consequently, in both cases, the correspondence problem in variable-resolution motion stereo is reduced to a simpler problem of extracting collinear points in the epipolar plane. A voting algorithm for accumulating multiple evidence is developed. The proposed method is potentially applicable to active sensing for automated inspection on assembly lines, autonomous road vehicle navigation, airport runway surveillance, etc. Preliminary experimental results are demonstrated.>
Ze-Nian Li
ICRA2
1994 Stereo Correspondence Based on Line Matching in Hough Space Using Dynamic Programming
abstract
This paper presents a method of using Hough space for solving the correspondence problem in stereo vision. It is shown that the line-matching problem in image space can readily be converted into a point-matching problem in Hough (/spl rho/-/spl theta/) space. Dynamic programming can be used for searching the optimal matching, now in Hough space. The combination of multiple constraints, especially the natural embedding of the constraint of figural continuity, ensures the accuracy of the matching. The time complexity for searching in dynamic programming is O(pmn), where m and n are the numbers of the lines for each /spl theta/ in the pair of stereo images, respectively, and p is the number of all possible line orientations. Since m and n are usually fairly small, the matching process is very efficient. Experimental results from both binocular and trinocular matchings are presented and analyzed.>
Ze-Nian Li
IEEE Trans. Syst. Man Cybern. Syst.1
1994 Depth Map Construction from Range-Guided Multiresolution Stereo Matching
abstract
This paper describes a multiresolution method for the acquisition of a complete, relatively noise-free, and high-resolution depth map from a low-resolution laser range image and a stereo pair of high resolution intensity images. Depth information from the laser range data is used to constrain the initial search in stereo matching. The inter- and intralevel linkings of edges in the pyramid allow a process where the coarse laser depth information drives a multiresolution stereo-matching process to construct a high-resolution depth map. The motivation for this new approach to multi-sensor integration is to offset the advantages and disadvantages of traditional stereo matching and triangulation range finding approaches.>
Kevin Tate, Ze-Nian Li
IEEE Trans. Syst. Man Cybern. Syst.2
1993 Real-time motion stereo
abstract
A simple parallel algorithm for calculating depth from motion stereo is implemented. It makes the assumption that the incremental disparity is less than the minimum distance between edges. To relax this constraint, a more general multiscale pyramidal algorithm is developed. The two algorithms are tested on the Simon Fraser University pyramidal vision machine for real-time computer vision applications. The total processing time is 50 ms/image for the simple algorithm and approximately 0.5 s/image for the multiscale algorithm.>
John Ens, Ze-Nian Li
CVPR2
1993 The reciprocal-wedge transform for space-invariant sensing
abstract
A reciprocal-wedge transform (RWT) for space-variant sensing is presented. The RWT preserves linear features and facilitates translations in images. A concise RWT matrix representation and its applications to geometric transformations in the image plane are described. Like the log-polar transform, the RWT facilitates space-variant sensing, which enables effective use of variable-resolution data and the reduction of the total amount of sensory data. The RWT yields anisotropic variable resolution, which provides an alternative for active sensing when the shift of attention is primarily in one dimension. A projective model of the RWT is delineated and a hardware RWT projection camera is proposed. As an application of the RWT, the problem of road following in robot navigation and some initial experimental results are demonstrated.>
Ze-Nian Li
ICCV2
1993 An X-Crossing Preserving Skeletonization Algorithm
abstract
For an image consisting of wire-like patterns, skeletonization (thinning) is often necessary as the first step towards feature extraction. But a serious problem which exists is that the intersections of the lines (X-crossings) will be elongated when applying a thinning algorithm to the image. That is, X-crossings are usually difficult to be preserved as the result of thinning. In this paper, we present a non-iterative line thinning method that preserves X-crossings of the lines in the image. The skeleton is formed by the mid-points of run-length encoding of the patterns. Line intersection areas are identified via a histogram analysis of the lengths of runs, and intersections are detected at locations where the sequences of runs merge or split.
Gongzhu Hu, Ze-Nian Li
Int. J. Pattern Recognit. Artif. Intell.2
1993 Linear generalized Hough transform and its parallelization
Ze-Nian Li, B. Yao
Image Vis. Comput.1
1993 Fast line detection in a hybrid pyramid
Ze-Nian Li, Danpo Zhang
Pattern Recognit. Lett.1
1992 On edge preservation in multiresolution images
Ze-Nian Li, Gongzhu Hu
CVGIP Graph. Model. Image Process.1
1992 On Improving the Accuracy of Line Extraction in Hough Space
abstract
This paper presents an accurate line extraction technique — the Hierarchical Peak Compaction Hough Transform (HPCHT). Vote scattering in the parameter space is a problem when the Hough transform is used for line extraction. This paper investigates the effects of image size and edge data errors on the severity of vote scattering. The HPCHT uses the Hough procedure on small subimages initially, and a recursive Hough merging scheme on the extracted line segments afterwards. A bound on vote scattering has been derived which guides the image subdivision and the adaptive quantization of the parameter space. As a result, an accurate Hough transform of low ρ-scattering and high θ-precision has been achieved. The HPCHT is suitable for fast parallel implementation on pyramid computers.
Ze-Nian Li
Int. J. Pattern Recognit. Artif. Intell.2
1992 Deductive-ER: deductive entity-relationship data model and its data language
Jiawei Han 0001, Ze-Nian Li
Inf. Softw. Technol.2
1991 Parallel algorithms for line detection on a 1×N array processor
abstract
A description is given of two algorithms that compute the Hough transform for straight lines on N*N images using 1*N processing arrays. The algorithms are developed for 1*N array processors and have been implemented on the AIS-5000 parallel vision computer. The complexity of both algorithms is O(N/sup 2/+PN) on the 1*N array processor (P is the number of angles polled which determines the theta -resolution of the Hough space), which compares well with the known optimal O(N+P) algorithms for N*N mesh arrays.>
Ze-Nian Li, Robert G. Laughlin
ICRA1
1989 Uncertainty management in a pyramid vision system
Ze-Nian Li
Int. J. Approx. Reason.1
1988 Comparisons of reasoning mechanisms for computer vision
Ze-Nian Li
Int. J. Approx. Reason.1
1987 Pyramid Vision Using Key Features to Integrate Image-Driven Bottom-Up and Model-Driven Top-Down Processes
abstract
Pyramidlike parallel hierarchical structures have been shown to be suitable for many computer vision tasks and have the potential for achieving the speeds needed for the real-time processing of real-world images. Algorithms are being developed to explore the pyramid's massively parallel and shallowly serial-hierarchical computing ability in an integrated system that combines both low-level and higher level vision tasks. Micromodular transforms are used to embody the program's knowledge of the different objects it must recognize. Pyramid vision programs are described that, starting with the image, use transforms that assess key features to dynamically imply other feature-detecting and characterizing transforms and additional top-down model-driven processes to apply. Program performance is presented for four real-world images of buildings. The use of key features in pyramid vision programs and the related search and control issues are discussed. To expedite the detection of various key features, feature-adaptable windows are developed. In addition to image-driven bottom-up and model-driven top-down processing, lateral search is used and is shown to be helpful, efficient, and feasible. The results indicate that with the use of key features and the combination of a variety of powerful search patterns, the pyramidlike structure is effective and efficient for supporting parallel and hierarchical object recognition algorithms.
Ze-Nian Li, Leonard Uhr
IEEE Trans. Syst. Man Cybern.1
1986 Evidential reasoning in a computer vision system
Ze-Nian Li, Leonard Uhr
UAI1
1986 A pyramidal approach for the recognition of neurons using key features
Ze-Nian Li, Leonard Uhr
Pattern Recognit.1