VLDB 2026 Research / reviewers in the wild / expert
Zheru Chi
dblp:25/2526
· DBLP profile ↗
101ranked-venue papers
6as first author
3since 2021 · last 2024
0000-0003-0714-8713ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 62 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 since 2021Human-computer interaction and ubiquitous computing · 12Applied, interdisciplinary, general and emerging computing · 9Databases, data management, data science and information retrieval · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Video understanding and tracking · 43% Segmentation and scene understanding · 21% Face, body and person analysis · 21% | |
| Computer graphics and multimedia
2 papers |
Image and video coding · 67% Image and video processing · 33% |
Topics — the 13 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
human pose estimation |
0.8 | 1 | 2024 | Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance Minimization · IEEE Trans. Image Process. 2024 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.8 | 1 | 2024 | Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance Minimization · IEEE Trans. Image Process. 2024 |
Computer vision › Video understanding and tracking › multi-object tracking
multi-person pose tracking |
0.8 | 1 | 2024 | Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance Minimization · IEEE Trans. Image Process. 2024 |
Computer vision › Segmentation and scene understanding
human parsing |
0.4 | 1 | 2019 | A CNN Model for Semantic Person Part Segmentation With Capacity Optimization · IEEE Trans. Image Process. 2019 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 1 | 2019 | A CNN Model for Semantic Person Part Segmentation With Capacity Optimization · IEEE Trans. Image Process. 2019 |
Computer vision › Image recognition and object detection
image classification |
0.1 | 1 | 2005 | Genetic Evolution Processing of Data Structures for Image Classification · IEEE Trans. Knowl. Data Eng. 2005 |
Computer vision › Image recognition and object detection
structural pattern recognition |
0.1 | 1 | 2005 | Genetic Evolution Processing of Data Structures for Image Classification · IEEE Trans. Knowl. Data Eng. 2005 |
Algorithms and data structures › symbolic computation › computational algebra › polynomial evaluation
polynomial root finding |
0.0 | 1 | 2004 | Finding roots of arbitrary high order polynomials based on neural network recursive partitioning method · Sci. China Ser. F Inf. Sci. 2004 |
Mathematical optimization
root finding |
0.0 | 1 | 2004 | Finding roots of arbitrary high order polynomials based on neural network recursive partitioning method · Sci. China Ser. F Inf. Sci. 2004 |
Information retrieval › pattern matching
template matching |
0.0 | 1 | 2003 | Document Image Recognition Based on Template Matching of Component Block Projections · IEEE Trans. Pattern Anal. Mach. Intell. 2003 |
Image and video processing
document image analysis |
0.0 | 1 | 2002 | Extraction and Optimization of B-Spline PBD Templates for Recognition of Connected Handwritten Digit Strings · IEEE Trans. Pattern Anal. Mach. Intell. 2002 |
Image and video coding › image compression
fractal image coding |
0.0 | 1 | 2002 | A fuzzy image metric with application to fractal coding · IEEE Trans. Image Process. 2002 |
Image and video coding
image quality assessment |
0.0 | 1 | 2002 | A fuzzy image metric with application to fractal coding · IEEE Trans. Image Process. 2002 |
Methods — techniques the papers use, named apart from their topics
sparse key-point flow estimation · 0.8hierarchical graph distance minimization · 0.8graph modeling · 0.8multi-module integration · 0.4capacity adjustment · 0.4neural network recursive partitioning · 0.1neural network · 0.1genetic algorithm · 0.1backpropagation through structures · 0.1triangle inequality · 0.0template matching · 0.0nearest-neighbor classifier · 0.0fuzzy integral · 0.0fuzzy image metric · 0.0evolutionary algorithm · 0.0dynamic programming · 0.0b-spline template · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Reasonable Anomaly Detection Based on Long-Term Sequence ModelingabstractVideo anomaly detection is a challenging task due to the unpredictable nature of abnormal actions, sophisticated semantics and a lack in training data. The visual representations of most existing approaches are limited by short-term sequences which cannot provide necessary clues for achieving reasonable detections. In this paper, we propose to comprehensively represent the motion patterns in human actions by learning from long-term sequences. Firstly, a Stacked State Machine (SSM) model with distinctive basis functions is proposed to represent the temporal dependencies which are consistent across long-term observations. Secondly, the dependencies are leveraged in filtering out problematic motion estimations which are influenced by short-term observation noises, plausible motion parameters are obtained in this way. Finally, SSM model predicts future states based on past ones, the divergence between the predictions with inherent normal patterns and observed ones determines anomalies which violate normal motion patterns. To address the challenges in drone-based surveillance, a dataset which is more diversified than existing ones is built. Extensive experiments are carried out to evaluate the proposed approach on the dataset and existing ones. Improvements over state-of-the-art methods can be observed. The proposed dataset will be made publicly available. Code is available athttps://github.com/AllenYLJiang/Anomaly-Detection-in-Sequences. Yalong Jiang, Changkang Li, Wenrui Ding, Jinzhi Xiang, Zheru Chi |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance MinimizationabstractIn this paper, we propose a novel framework for multi-person pose estimation and tracking on challenging scenarios. In view of occlusions and motion blurs which hinder the performance of pose tracking, we proposed to model humans as graphs and perform pose estimation and tracking by concentrating on the visible parts of human bodies which are informative about complete skeletons under incomplete observations. Specifically, the proposed framework involves three parts: (i) A Sparse Key-point Flow Estimating Module (SKFEM) and a Hierarchical Graph Distance Minimizing Module (HGMM) for estimating pixel-level and human-level motion, respectively; (ii) Pixel-level appearance consistency and human-level structural consistency are combined in measuring the visibility scores of body joints. The scores guide the pose estimator to predict complete skeletons by observing high-visibility parts, under the assumption that visible and invisible parts are inherently correlated in human part graphs. The pose estimator is iteratively fine-tuned to achieve this capability; (iii) Multiple historical frames are combined to benefit tracking which is implemented using HGMM. The proposed approach not only achieves state-of-the-art performance on PoseTrack datasets but also contributes to significant improvements in other tasks such as human-related anomaly detection. Yalong Jiang, Wenrui Ding, Zheru Chi |
IEEE Trans. Image Process. | 4 |
| 2022 | Iterative brain tumor retrieval for MR images based on user's intention model
Mengli Sun, Nan Hu 0001, Zheru Chi |
Pattern Recognit. | 5 |
| 2020 | Precise Temporal Localization for Complete Actions with Quantified Temporal StructureabstractExisting temporal action detection algorithms cannot distinguish complete and incomplete actions while this property is essential in many applications. To tackle this challenge, we proposed the action progression networks (APN), a novel model that predicts action progression of video frames with continuous numbers. Using the progression sequence of test video, on the top of the APN, a complete action searching algorithm (CAS) was designed to detect complete actions only. With the usage of frame-level fine-grained temporal structure modeling and detecting actions according to their whole temporal context, our framework can locate actions precisely and is good at avoiding incomplete action detection. We evaluated our framework on a new dataset (DFMAD-70) collected by ourselves which contains both complete and incomplete actions. Our framework got good temporal localization results with 95.77% average precision when the IoU threshold is 0.5. On the benchmark THUMOS14, an incomplete-ignostic dataset, our framework still obtain competitive performance. The code is available online at https://github.com/MakeCent/Action-Progression-Network. Chongkai Lu, Hong Fu, Zheru Chi |
ICPR | 7 |
| 2020 | Detection of High-Risk Depression Groups Based on Eye-Tracking Data
Simeng Lu, Shen Huang, Xiujuan Zheng, Danmin Miao, Zheru Chi |
PRCV (2) | 7 |
| 2020 | A Residual Based Attention Model for EEG Based Sleep StagingabstractSleep staging is to score the sleep state of a subject into different sleep stages such as Wake and Rapid Eye Movement (REM). It plays an indispensable role in the diagnosis and treatment of sleep disorders. As manual sleep staging through well-trained sleep experts is time consuming, tedious, and subjective, many automatic methods have been developed for accurate, efficient, and objective sleep staging. Recently, deep learning based methods have been successfully proposed for electroencephalogram (EEG) based sleep staging with promising results. However, most of these methods directly take EEG raw signals as input of convolutional neural networks (CNNs) without considering the domain knowledge of EEG staging. Apart from that, to capture temporal information, most of the existing methods utilize recurrent neural networks such as LSTM (Long Short Term Memory) which are not effective for modelling global temporal context and difficult to train. Therefore, inspired by the clinical guidelines of sleep staging such as AASM (American Academy of Sleep Medicine) rules where different stages are generally characterized by EEG waveforms of various frequencies, we propose a multi-scale deep architecture by decomposing an EEG signal into different frequency bands as input to CNNs. To model global temporal context, we utilize the multi-head self-attention module of the transformer model to not only improve performance, but also shorten the training time. In addition, we choose residual based architecture which makes training end-to-end. Experimental results on two widely used sleep staging datasets, Montreal Archive of Sleep Studies (MASS) and sleep-EDF datasets, demonstrate the effectiveness and significant efficiency (up to 12 times less training time) of our proposed method over the state-of-the-art. Zhiyong Wang 0001, Hong Hong 0001, Zheru Chi, David Dagan Feng, Ronald R. Grunstein, Christopher James Gordon |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | A CNN Model for Semantic Person Part Segmentation With Capacity OptimizationabstractIn this paper, a deep learning model with an optimal capacity is proposed to improve the performance of person part segmentation. Previous efforts in optimizing the capacity of a CNN model suffer from a lack of large datasets as well as the over-dependence on a single-modality CNN which is not effective in learning. We make several efforts in addressing these problems. Firstly, other datasets are utilized to train a CNN module for pre-processing image data and a segmentation performance improvement is achieved without a time-consuming annotation process. Secondly, we propose a novel way of integrating two complementary modules to enrich the feature representations for more reliable inferences. Thirdly, the factors to determine the capacity of a CNN model are studied and two novel methods are proposed to adjust (optimize) the capacity of a CNN to match it to the complexity of a task. The over-fitting and under-fitting problems are eased by using our methods. Experimental results show that our model outperforms the state-of-the-art deep learning models with a better generalization ability and a lower computational complexity. Yalong Jiang, Zheru Chi |
IEEE Trans. Image Process. | 2 |
| 2018 | Person Part Segmentation based on Weak Supervision
Yalong Jiang, Zheru Chi |
BMVC | 2 |
| 2018 | A Novel Structure of Convolutional Layers with a Higher Performance-Complexity Ratio for Semantic SegmentationabstractIn this paper, we study an important factor that determines the capacity of a CNN model and propose a novel structure of convolutional layers with a higher performance-complexity ratio. Firstly, the relationship of the model capacity and the number of parameters versus segmentation performance is explored. Secondly, a mechanism is proposed to optimize the structure of a CNN model for a specific task. The mechanism also provides better convergence than current state-of-the-art methods for factorizing convolutional layers, such as MobileNet. Thirdly, we propose a measure based on the mutual information between hidden activations and inputs/outputs to compute the capacity of a CNN model. This measure is highly correlated with segmentation performance. Experimental results on the segmentation of the PASCAL Person Parts Dataset show that the linear dependency among convolutional kernels is an important factor determining the capacity of a CNN model. It is also demonstrated that our approach can successfully adjust the model capacity to best match to the complexity of a dataset. The optimized CNN model achieves the similar performance to Deeplab-V2 on the segmentation task with 100 × less parameters, resulting in a significantly improved performance-complexity ratio. Yalong Jiang, Zheru Chi |
ICARCV | 2 |
| 2018 | Facial Expression Recognition in Video with Multiple Feature FusionabstractVideo based facial expression recognition has been a long standing problem and attracted growing attention recently. The key to a successful facial expression recognition system is to exploit the potentials of audiovisual modalities and design robust features to effectively characterize the facial appearance and configuration changes caused by facial motions. We propose an effective framework to address this issue in this paper. In our study, both visual modalities (face images) and audio modalities (speech) are utilized. A new feature descriptor called Histogram of Oriented Gradients from Three Orthogonal Planes (HOG-TOP) is proposed to extract dynamic textures from video sequences to characterize facial appearance changes. And a new effective geometric feature derived from the warp transformation of facial landmarks is proposed to capture facial configuration changes. Moreover, the role of audio modalities on recognition is also explored in our study. We applied the multiple feature fusion to tackle the video-based facial expression recognition problems under lab-controlled environment and in the wild, respectively. Experiments conducted on the extended Cohn-Kanade (CK+) database and the Acted Facial Expression in Wild (AFEW) 4.0 database show that our approach is robust in dealing with video-based facial expression recognition problems under lab-controlled environment and in the wild compared with the other state-of-the-art methods. JunKai Chen, Zenghai Chen, Zheru Chi, Hong Fu |
IEEE Trans. Affect. Comput. | 3 |
| 2017 | A scale-invariant framework for image classification with deep learningabstractIn this paper, we propose a scale-invariant framework based on Convolutional Neural Networks (CNNs). The network exhibits robustness to scale and resolution variations in data. Previous efforts in achieving scale invariance were made on either integrating several variant-specific CNNs or data augmentation. However, these methods did not solve the fundamental problem that CNNs develop different feature representations for the variants of the same image. The topology proposed by this paper develops a uniform representation for each of the variants of the same image. The uniformity is acquired by concatenating scale-variant and scale-invariant features to enlarge the feature space so that the case when input images are of diverse variations but from the same class can be distinguished from another case when images are of different classes. Higher-order decision boundaries lead to the success of the framework. Experimental results on a challenging dataset substantiates that our framework performs better than traditional frameworks with the same number of free parameters. Our proposed framework can also achieve a higher training efficiency. Yalong Jiang, Zheru Chi |
SMC | 2 |
| 2017 | A new framework with multiple tasks for detecting and locating pain events in video
JunKai Chen, Zheru Chi, Hong Fu |
Comput. Vis. Image Underst. | 2 |
| 2017 | Smile detection in the wild with deep convolutional neural networks
JunKai Chen, Qihao Ou, Zheru Chi, Hong Fu |
Mach. Vis. Appl. | 3 |
| 2017 | Eye tracking data guided feature selection for image classification
Xin Gao 0003, Zhiyong Wang 0001, Zheru Chi |
Pattern Recognit. | 6 |
| 2016 | Smile detection in the wild with hierarchical visual featureabstractSmile detection in the wild is an interesting and challenging problem. This paper presents an efficient approach with hierarchical visual feature to handle this problem. In our approach, Gabor filters with multi-scale, multi-orientation are first applied to extract facial textures namely Gabor faces from the input face image. After this, Histograms of Oriented Gradients (HOG) are employed to encode these extracted Gabor faces to capture and characterize the facial appearance characteristics. We further adopt a pooling strategy to transform the multiple HOG features into a global visual feature called Gabor-Hog. Finally, SVM is trained to perform the classification. The experiments conducted on the GENKI4K database show that the proposed visual feature is robust to distinguish a smile face from a no-smile face. Our method also achieves a promising performance compared with the other state-of-the-art methods. Jiahuiran Li, JunKai Chen, Zheru Chi |
ICIP | 3 |
| 2016 | Facial expression recognition with dynamic Gabor volume featureabstractFacial expression recognition is a long standing problem in affective computing community. A key step is extracting effective features from face images. Gabor filters have been widely used for this purpose. However, a big challenge for Gabor filters is its high dimensionality. In this paper, we propose an efficient feature called dynamic Gabor volume feature (DGVF) based on Gabor filters while with a lower dimensionality for facial expression recognition. In our approach, we first apply Gabor filters with multi-scale and multi-orientation to extract different Gabor faces. And these Gabor faces are arranged into a 3-D volume and Histograms of Oriented Gradients from Three Orthogonal Planes (HOG-TOP) are further employed to encode the 3-D volume in a compact way. Finally, SVM is trained to perform the classification. The experiments conducted on the Extended Cohn-Kanade (CK+) Dataset show that the proposed DGVF is robust to capture and represent the facial appearance features. And our method also achieves a superior performance compared with the other state-of-the-art methods. JunKai Chen, Zheru Chi, Hong Fu |
MMSP | 2 |
| 2015 | A new approach for pain event detection in videoabstractA new approach for pain event detection in video is presented in this paper. Different from some previous works which focused on frame-based detection, we target in detecting pain events at video level. In this work, we explore the spatial information of video frames and dynamic textures of video sequences, and propose two different types of features. HOG of fiducial points (P-HOG) is employed to extract spatial features from video frames and HOG from Three Orthogonal Planes (HOG-TOP) is used to represent dynamic textures of video subsequences. After that, we apply max pooling to represent a video sequence as a global feature vector. Multiple Kernel Learning (MKL) is utilized to find an optimal fusion of the two types of features. And an SVM with multiple kernels is trained to perform the final classification. We conduct our experiments on the UNBC-McMaster Shoulder Pain dataset and achieve promising results, showing the effectiveness of our approach. JunKai Chen, Zheru Chi, Hong Fu |
ACII | 2 |
| 2015 | Dynamic texture and geometry features for facial expression recognition in videoabstractFacial expression recognition in video has attracted growing attention recently. In this paper, we propose to handle this problem with dynamic appearance and geometric features. We propose a new feature descriptor called HOG from Three Orthogonal Planes (HOG-TOP) to represent dynamic features. In addition, we introduce two types of geometry features to represent the facial rigid changes and non-rigid changes, respectively. Multiple Kernel Learning (MKL) is applied to find an optimal combination of two types of features. And finally a Support Vector Machine (SVM) with multiple kernels is trained for the facial expression classification. Extensive experiments conducted on the extended Cohn-Kanade dataset show that our method can achieve a competitive performance compared with the other state-of-the-art methods. JunKai Chen, Zenghai Chen, Zheru Chi, Hong Fu |
ICIP | 3 |
| 2015 | Eye-Tracking Aided Digital System for Strabismus DiagnosisabstractStrabismus is a common ophthalmic disease with a relatively high prevalence (4%). It is one of the most common vision disorders in preschool children. If it is not timely diagnosed and well treated, strabismus would cause amblyopia, and even permanent vision loss. The diagnosis is thus essential. However, most of the diagnosis methods, such as cover testing, are conducted manually by an ophthalmologist. The examination cost is relatively high and the results are subjective. In this paper, we propose an eye tracking method for strabismus diagnosis. This method allows us to develop an objective and automatic strabismus diagnosis system that could significantly increase the examination efficiency and reduce the cost. Experimental results demonstrate the effectiveness of the proposed eye tracking method for strabismus diagnosis. Zenghai Chen, Hong Fu, Zheru Chi |
SMC | 4 |
| 2015 | Learning to Detect Saliency with Deep StructureabstractDeep learning has shown great successes in solving various problems of computer vision. To the best of our knowledge, however, little existing work applies deep learning to saliency modeling. In this paper, a new saliency model based on convolutional neural network is proposed. The proposed model is able to produce a saliency map directly from an image's pixels. In the model, multi-level output values are adopted to simulate continuous values in a saliency map. Differing from most neural networks that use a relatively small number of output nodes, the output layer of our model has a large number of nodes. To make the training more efficient, an improved learning algorithm is adopted to train the model. Experimental results show that the proposed model succeeds in generating acceptable saliency maps after proper training. Zenghai Chen, Zheru Chi, Hong Fu |
SMC | 3 |
| 2014 | Emotion Recognition in the Wild with Feature Fusion and Multiple Kernel LearningabstractThis paper presents our proposed approach for the second Emotion Recognition in The Wild Challenge. We propose a new feature descriptor called Histogram of Oriented Gradients from Three Orthogonal Planes (HOG_TOP) to represent facial expressions. We also explore the properties of visual features and audio features, and adopt Multiple Kernel Learning (MKL) to find an optimal feature fusion. An SVM with multiple kernels is trained for the facial expression classification. Experimental results demonstrate that our method achieves a promising performance. The overall classification accuracy on the validation set and test set are 40.21% and 45.21%, respectively. JunKai Chen, Zenghai Chen, Zheru Chi, Hong Fu |
ICMI | 3 |
| 2014 | A Hybrid Holistic/Semantic Approach for Scene ClassificationabstractThere are two main strategies to tackle scene classification: holistic and semantic. The former characterizes a scene using its global features, while the latter represents a scene by modeling its internal object configuration. Holistic strategy is good at representing scenes with simple contents, but it does not represent well complex scenes that consist of multiple objects. By contrast, semantic strategy is advantageous at recognizing scenes with complex objects, but it does not work well for simple scenes. In this paper, we propose to integrate holistic and semantic strategies to cope with scene classification. In particular, we exploit a deep learning algorithm to learn features for scene representation in the holistic way. For the semantic strategy, we explore a semantic spatial pyramid to represent the spatial object configuration of scenes. The holistic and semantic strategies are integrated using a method proposed by us. Experimental results on a benchmark natural scene dataset demonstrate the effectiveness of our proposed hybrid approach for scene classification, by comparing to several state-of-the-art algorithms. Zenghai Chen, Zheru Chi, Hong Fu |
ICPR | 2 |
| 2014 | Relative Saliency Model over Multiple Images with an Application to Yarn Surface EvaluationabstractSaliency models have been developed and widely demonstrated to benefit applications in computer vision and image understanding. In most of existing models, saliency is evaluated within an individual image. That is, saliency value of an item (object/region/pixel) represents the conspicuity of it as compared with the remaining items in the same image. We call this saliency as absolute saliency, which is uncomparable among images. However, saliency should be determined in the context of multiple images for some visual inspection tasks. For example, in yarn surface evaluation, saliency of a yarn image should be measured with regard to a set of graded standard images. We call this saliency the relative saliency, which is comparable among images. In this paper, a study of visual attention model for comparison of multiple images is explored, and a relative saliency model of multiple images is proposed based on a combination of bottom-up and top-down mechanisms, to enable relative saliency evaluation for the cases where other image contents are involved. To fully characterize the differences among multiple images, a structural feature extraction strategy is proposed, where two levels of feature (high-level, low-level) and three types of feature (global, local-local, local-global) are extracted. Mapping functions between features and saliency values are constructed and their outputs reflect relative saliency for multiimage contents instead of single image content. The performance of the proposed relative saliency model is well demonstrated in a yarn surface evaluation. Furthermore, the eye tracking technique is employed to verify the proposed concept of relative saliency for multiple images. Bingang Xu, Zheru Chi, David Dagan Feng |
IEEE Trans. Cybern. | 3 |
| 2013 | Multi-instance multi-label image classification: A neural approach
Zenghai Chen, Zheru Chi, Hong Fu, David Dagan Feng |
Neurocomputing | 2 |
| 2013 | Learning realistic facial expressions from web images
Kaimin Yu, Zhiyong Wang 0001, Li Zhuo 0001, Zheru Chi, David Dagan Feng |
Pattern Recognit. | 5 |
| 2013 | Realistic Human Action Recognition With Multimodal Feature Selection and FusionabstractAlthough promising results have been achieved for human action recognition under well-controlled conditions, it is very challenging to recognize human actions in realistic scenarios due to increased difficulties such as dynamic backgrounds. In this paper, we propose to take multimodal (i.e., audiovisual) characteristics of realistic human action videos into account in human action recognition for the first time, since, in realistic scenarios, audio signals accompanying an action generally provide a cue to the nature of the action, such as phone ringing to answering the phone . In order to cope with diverse audio cues of an action in realistic scenarios, we propose to identify effective features from a large number of audio features with the generalized multiple kernel learning algorithm. The widely used space-time interest point descriptors are utilized as visual features, and a support vector machine is employed for both audio- and video-based classifications. At the final stage, fuzzy integral is utilized to fuse recognition results of both audio and visual modalities. Experimental results on the challenging Hollywood-2 Human Action data set demonstrate that the proposed approach is able to achieve better recognition performance improvement than that of integrating scene context. It is also discovered how audio context influences realistic action recognition from our comprehensive experiments. Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Zheru Chi, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2012 | Eye typing of Chinese charactersabstractEye typing is one of the most intensively investigated topics in eye tracking technology. Currently, almost all eye typing systems are developed for English typing. Some preliminary studies have been made on developing eye typing systems for inputting Chinese characters/text. In this paper, a novel eye typing system is proposed for inputting Chinese characters, where a software keyboard is specially designed based on a study of Chinese Pinyin. Experimental results show the efficiency and usability of the proposed system. Zheru Chi |
ETRA | 3 |
| 2012 | Intelligent characterization and evaluation of yarn surface appearance using saliency map analysis, wavelet transform and fuzzy ARTMAP neural network
Bingang Xu, Zheru Chi, David Dagan Feng |
Expert Syst. Appl. | 3 |
| 2012 | Salient object detection using content-sensitive hypergraph representation and partitioning
Zheru Chi, Hong Fu, David Dagan Feng |
Pattern Recognit. | 2 |
| 2012 | An Adaptive Recognition Model for Image AnnotationabstractIn this paper, an adaptive recognition model (ARM) is proposed for image annotation. The ARM consists of an adaptive classification network (CFN) and a nonlinear correlation network (CLN). The adaptive CFN aims to annotate an image with keywords, and the CLN is used to unveil the correlative information of keywords for annotation refinement. Image annotation is carried out by an ARM in two stages. In the first stage, the features extracted from regions of the input image are fed to a CFN to produce classification labels. In the second stage, the CLN uses keyword correlations learned from the training images to refine the classification result. The ARM works in a forward-propagating manner, resulting in high efficiency in image annotation. Furthermore, the computational time of an ARM is insensitive to the number of regions of the input image and the vocabulary size. In this paper, the effect of keyword correlation in image annotation is, comprehensively, investigated on a real image dataset and a synthetic image dataset. The exploitation of a controllable synthetic dataset helps to systematically study the function of keyword correlation and effectively analyze the performance of the ARM. Experimental results demonstrate the efficiency and effectiveness of the ARM. Zenghai Chen, Hong Fu, Zheru Chi, David Dagan Feng |
IEEE Trans. Syst. Man Cybern. Part C | 3 |
| 2011 | Construction of Evolutionary Multi-agent Double Auction Market for Data Mining Combinational Strategies with Stable Returns
Zheru Chi, Na Jia, Huiqun Zhao |
ICAART (2) | 3 |
| 2011 | A Time Delay Neural Network model for simulating eye gaze dataabstractHuman eye movement modelling is a new, challenging and promising research topic in computer vision. Human eye movement modelling aims at simulating the scan path in which a human being views an image, a scene or a video. The successful modelling of human eye movements potentially benefits a wide range of applications such as image retrieval, image annotation, medical image diagnosis and human visual perception. This article presents a model based on a Time Delay Neural Network (TDNN) to simulate eye gaze data. First, 120 Hz eye gaze data are acquired by a non-intrusive table-mounted eye tracker. Our proposed model is to simulate the image reading process of a single subject. Seven features are then extracted based on the knowledge of the human oculomotor system and the image contents to train a TDNN. Finally, the trained TDNN combined with a saccade control mechanism is used to simulate the scan path of a human being viewing an image. The proposed model can generate 600 points of raw eye gaze data in a 5-second eye viewing window. Both subjective and objective methods are used to evaluate the model by comparing its behaviour and characteristics with the real eye gaze data collected from an eye tracker. Qualitative assessment shows that the subject can hardly tell the differences between the scan path from the model and that from a human being. By evaluating the coincident probability Cp and coincident significance Cs , quantitative assessment shows that the results from the TDNN model are reasonable and similar to human scan paths. Hong Fu, Zheru Chi, David Dagan Feng |
J. Exp. Theor. Artif. Intell. | 5 |
| 2010 | Combined Retrieval Strategies for Images with and without Distinct Objects
Hong Fu, Zheru Chi, David Dagan Feng |
ACIVS (1) | 2 |
| 2010 | Salient-SIFT for Image Retrieval
Hong Fu, Zheru Chi, David Dagan Feng |
ACIVS (1) | 3 |
| 2010 | Content-based image retrieval using a combination of visual features and eye tracking dataabstractImage retrieval technology has been developed for more than twenty years. However, the current image retrieval techniques cannot achieve a satisfactory recall and precision. To improve the effectiveness and efficiency of an image retrieval system, a novel content-based image retrieval method with a combination of image segmentation and eye tracking data is proposed in this paper. In the method, eye tracking data is collected by a non-intrusive table mounted eye tracker at a sampling rate of 120 Hz, and the corresponding fixation data is used to locate the human's Regions of Interest (hROIs) on the segmentation result from the JSEG algorithm. The hROIs are treated as important informative segments/objects and used in the image matching. In addition, the relative gaze duration of each hROI is used to weigh the similarity measure for image retrieval. The similarity measure proposed in this paper is based on a retrieval strategy emphasizing the most important regions. Experiments on 7346 Hemera color images annotated manually show that the retrieval results from our proposed approach compare favorably with conventional content-based image retrieval methods, especially when the important regions are difficult to be located based on visual features. Hong Fu, Zheru Chi, David Dagan Feng |
ETRA | 4 |
| 2010 | Eye movement as an interaction mechanism for relevance feedback in a content-based image retrieval systemabstractRelevance feedback (RF) mechanisms are widely adopted in Content-Based Image Retrieval (CBIR) systems to improve image retrieval performance. However, there exist some intrinsic problems: (1) the semantic gap between high-level concepts and low-level features and (2) the subjectivity of human perception of visual contents. The primary focus of this paper is to evaluate the possibility of inferring the relevance of images based on eye movement data. In total, 882 images from 101 categories are viewed by 10 subjects to test the usefulness of implicit RF, where the relevance of each image is known beforehand. A set of measures based on fixations are thoroughly evaluated which include fixation duration, fixation count, and the number of revisits. Finally, the paper proposes a decision tree to predict the user's input during the image searching tasks. The prediction precision of the decision tree is over 87%, which spreads light on a promising integration of natural eye movement into CBIR systems in the future. Hong Fu, Zheru Chi, David Dagan Feng |
ETRA | 4 |
| 2010 | A neural network model with adaptive structure for image annotationabstractA neural network model with adaptive structure for image annotation is proposed in this paper. The adaptive structure enables the proposed model to utilize both global and regional visual features, as well as correlative information of annotated keywords for annotation. In order to achieve an approximate global optimum rather than a local optimum, both genetic algorithm and traditional back-propagation algorithm, are combined for model training. The neural network model is experimented on a synthetic image dataset with controllable parameters, which has not been used in previous image annotation experiments. Experimental results demonstrate the effectiveness of the proposed model. Zenghai Chen, Hong Fu, Zheru Chi, David Dagan Feng |
ICARCV | 3 |
| 2010 | An improved algorithm for segmenting and recognizing connected handwritten charactersabstractIn this paper, an improved algorithm is proposed for the segmentation and recognition of handwritten character strings. In the method, a gradient descent mechanism is used to weigh the distance measure in applying KNN for segmenting/recognizing connected characters (numerals and Chinese characters) in the left-to-right scanning direction. In recognizing connected characters, a high quality segmentation technique is essential. Conventional approaches attempt to separate the string into individual characters without recognition and apply a recognition algorithm onto each isolated character, resulting improper segmentation and poor recognition results in many situations. Our proposed algorithm simulates the human beings's process in recognizing connected character strings where segmentation and recognition is mingled with each other. Experimental results on 1959 character strings from the USPS database of postal envelopes show that the algorithm works robustly and efficiently. Zheru Chi, David Dagan Feng |
ICARCV | 2 |
| 2010 | Refining a region based attention model using eye tracking dataabstractComputational visual attention modeling is a topic of increasing importance in machine understanding of images. In this paper, we present an approach to refine a region based attention model with eye tracking data. This paper has three main contributions. (1) A concept of fixation mask is proposed to describe the region saliency of an image by weighting the segmented regions using importance measures obtained in the Human Visual System (HVS) or computational models. (2) A Genetic Algorithm (GA) scheme for refining a region based attention model is proposed. (3) An evaluation method is developed to measure the correlation between the result from the computational model and that from the HVS in terms of fixation mask. Hong Fu, Zheru Chi, David Dagan Feng |
ICIP | 3 |
| 2010 | A Novel Shape-Based Image Classification Method by Featuring Radius Histogram of Dilating Discs Filled into Regular and Irregular Shapes
Zheru Chi, David Dagan Feng |
ICONIP (1) | 3 |
| 2010 | Recognition of attentive objects with a concept association network for image annotation
Hong Fu, Zheru Chi, David Dagan Feng |
Pattern Recognit. | 2 |
| 2009 | Eye movement data modeling using a genetic algorithmabstractWe present a computational model of human eye movements based on a genetic algorithm (GA). The model can generate elemental raw eye movement data in a four-second eye viewing window with a 25 Hz sampling rate. Based on the physiology and psychology characters of human vision system, the fitness function of the GA model is constructed by taking into consideration of five factors including the saliency map, short time memory, saccades distribution, Region of Interest (ROI) map, and a retina model. Our model can produce the scan path of a subject viewing an image, not just several fixations points or artificial ROI's as in the other models. We have also developed both subjective and objective methods to evaluate the model by comparing its behavior with the real eye movement data collected from an eye tracker. Tested on 18 (9 times 2) images from both an obvious-object image group and a non-obvious-object image group, the subjective evaluations shows very close scores between the scan paths generated by the GA model and those real scan paths; for the objective evaluation, experimental results show that the distance between GA's scan paths and human scan paths of the same image has no significant difference by a probability of 78.9% on average. Hong Fu, Zheru Chi, David Dagan Feng |
IEEE Congress on Evolutionary Computation | 5 |
| 2009 | Tree structures with attentive objects for image classification using a neural networkabstractThis paper presents an image classification method based on a neural network model dealing with tree structures of attentive objects. Apart from regions provided by image segmentation, attentive objects, which are extracted from a segmented image by an attention-driven image interpretation algorithm, are used to construct the tree structure to represent an image. Three combinations of tree structures are investigated, including ldquoimage + attentive-object + segmentsrdquo, ldquoimage + attentive-objectsrdquo, as well as ldquoimage + segmentsrdquo. Structure based neural networks are trained to classify the images by using the back propagation through structure (BPTS) algorithm. Experimental results show that the ldquoimage + attentive objectsrdquo structure is more favorable, comparing with both the other two structures proposed by us and a start-of-art tree structure reported in the literature, in terms of classification rate and computational time. Hong Fu, Shuya Zhang, Zheru Chi, David Dagan Feng |
IJCNN | 3 |
| 2009 | An efficient algorithm for attention-driven image interpretation from segments
Hong Fu, Zheru Chi, David Dagan Feng |
Pattern Recognit. | 2 |
| 2008 | Improvement of Image Classification Using Wavelet Coefficients with Structured-Based Neural NetworkabstractImage classification is a challenging problem in organizing a large image database. However, an effective method for such an objective is still under investigation. A method based on wavelet analysis to extract features for image classification is presented in this paper. After an image is decomposed by wavelet, the statistics of its features can be obtained by the distribution of histograms of wavelet coefficients, which are respectively projected onto two orthogonal axes, i.e., x and y directions. Therefore, the nodes of tree representation of images can be represented by the distribution. The high level features are described in low dimensional space including 16 attributes so that the computational complexity is significantly decreased. 2,800 images derived from seven categories are used in experiments. Half of the images were used for training neural network and the other images used for testing. The features extracted by wavelet analysis and the conventional features are used in the experiments to prove the efficacy of the proposed method. The classification rate on the training data set with wavelet analysis is up to 91%, and the classification rate on the testing data set reaches 89%. Experimental results show that our proposed approach for image classification is more effective. Weibao Zou, Zheru Chi, King Chuen Lo |
Int. J. Neural Syst. | 2 |
| 2007 | Face Recognition Based on Binary Template Matching
Jiatao Song, Beijing Chen, Zheru Chi, Xuena Qiu, Wei Wang 0106 |
ICIC (1) | 3 |
| 2007 | Image Recovery from Broken Image StreamsabstractThis paper presents an approach for image recovery from broken image streams based on an image continuity model. The information to be recovered includes the header of an image (width, height, color space etc.) and the data stream which might have been partially damaged. We begin our discussion on common image formats and necessary information for the image recovery. We then define several measures on image contents to facilitate the image recovery. Using these measures, an effective recovery algorithm is proposed. Experimental results show that the algorithm can successfully recover images of different formats, even multi-frame images from broken image streams in most cases. Zheru Chi, Hong Yan 0001, David Dagan Feng, Gang Chen 0006 |
ICIP (3) | 3 |
| 2007 | Pre-classification Module for an All-Season Image Retrieval SystemabstractFrom the study of attention-driven image interpretation and retrieval, we have found that an attention-driven strategy is able to extract important objects from an image and then focus the attentive objects while retrieving images. However, besides the images with distinct objects, there are images which do not show distinct objects. In this paper, the classification of "attentive" and "non-attentive" image is proposed to be a pre-process module in an all-season image retrieval system which can tackle both kinds of images. In this pre-classification module, an image is represented by an adaptive tree structure with each node carrying normalized features that characterize the object/region with visual contrasts and spatial information. Then a neural network is trained to classify an image as an "attentive" or "non-attentive" category by using the Back Propagation Through Structure (BPTS) algorithm. Experimental results indicate the reliability and feasibility of the pre-classification module, which encourages us to conduct further investigations on the all-season image retrieval system. Hong Fu, Zheru Chi, David Dagan Feng, Weibao Zou, King Chuen Lo |
IJCNN | 2 |
| 2007 | Pattern-Oriented Agent-Based Modeling for Financial Market Simulation
Zheru Chi |
ISNN (1) | 2 |
| 2006 | Improvement of Image Classification with Wavelet and Independent Component Analysis (ICA) based on a Structured Neural NetworkabstractImage classification is a challenging problem in organizing a large image database. However, an effective method for such an objective is still under investigation. This paper presents a method based on wavelet and Independent Analysis Component (ICA) for image classification with adaptive processing of data structures. With wavelet, an image is decomposed into low frequency bands and high frequency bands. An image can be characterized by wavelet coefficients in the form of tree representation. While the histograms of low frequency wavelet bands are effective in characterizing images, the histograms of high frequency wavelet bands are similar for different images and therefore they cannot be directly used as features. We make use of ICA for feature extraction from high frequency bands to improve image classification. Two sets of features are used together to classify images using a structured neural network. In total, 2940 images generated from seven categories are used in experiments. Half of the images are used for training the neural network and the other images used for testing. The classification rate of the training set is 92%, and the classification rate of the test set reaches 89%. The experimental results show the effectiveness of the proposed method based on combining wavelet and ICA for image classification. Weibao Zou, Yan Li 0002, King Chuen Lo, Zheru Chi |
IJCNN | 4 |
| 2006 | Annotating Image Regions Using Spatial ContextabstractImage annotation plays an important role in bridging the semantic gap between low level features and high level semantic contents in image access. In this paper, such a task is tackled by annotating regions which are primitives of a visual scene. We propose a probabilistic model to characterize spatial context for region annotation. Such a model provides a unifying framework integrating both feature distribution models and spatial context models. A wide range of advanced modeling techniques can be utilized to further extend this framework. The approach is also potentially scalable to a large number of semantic concepts and a large number of images. Experimental results based on simple parametric models demonstrate promising results of our approach by investigating the impacts of neighbors, segmentation, and visual features Zhiyong Wang 0001, David Dagan Feng, Zheru Chi |
ISM | 3 |
| 2006 | Structured-Based Neural Network Classification of Images Using Wavelet Coefficients
Weibao Zou, King Lo, Zheru Chi |
ISNN (2) | 3 |
| 2006 | An Algorithm Combining Statistics-based and Rules-based for Chunk Identification of Chinese Sentences
Rongbo Wang, Zheru Chi |
PACLIC | 2 |
| 2006 | Leaf Vein Extraction Using Independent Component AnalysisabstractThe purpose of this work is to develop an interactive tool which helps botanists to extract the vein system with its hierarchical properties with as little user interaction as possible. In this paper, we present a new venation extraction method using independent component analysis (ICA). The popular and efficient FastICA algorithm is applied to patches of leaf images to learn a set of linear basis functions or features for the images and then the basis functions are used as the pattern map for vein extraction. In our experiments, the training sets are randomly generated from different leaf images. Experimental results demonstrate that ICA is a promising technique for extracting leaf veins and edges of objects. ICA, therefore, can play an important role in automatically identifying living plants. Yan Li 0002, Zheru Chi, David Dagan Feng |
SMC | 2 |
| 2006 | Attention-driven image interpretation with application to image retrieval
Hong Fu, Zheru Chi, David Dagan Feng |
Pattern Recognit. | 2 |
| 2006 | A robust eye detection method using combined binary edge and intensity information
Jiatao Song, Zheru Chi, Jilin Liu |
Pattern Recognit. | 2 |
| 2005 | An operation-based exchange mechanism in heterogeneous CAD collaborationabstractThis paper presents an operation-based exchange mechanism for interoperability in heterogeneous CAD collaboration. The mechanism different from STEP-based exchange mechanism facilitates the engineers to exchange design information in distributed environment with shared semantic operations instead of model data, and to synchronously or asynchronously reconstruct the design process by executing the semantic operations in native CAD system. The proposed architecture based on this mechanism is developed to integrate traditional CAD systems and multimedia tools, and to provide essential collaborative support functions. The operation-based collaborative design system is applicable to many interactive design activities such as remote design review and online collaborative design. Shensheng Zhang, Zheru Chi |
CSCWD (1) | 3 |
| 2005 | Locating Human Eyes Using Edge and Intensity Information
Jiatao Song, Zheru Chi, Zhengyou Wang, Wei Wang 0106 |
ICIC (2) | 2 |
| 2005 | Scene Classification Using Adaptive Processing of Tree Representation of Rectangular-Shape Partition of Images
Ken Lo, Zheru Chi |
ISNN (2) | 3 |
| 2005 | Genetic Evolution Processing of Data Structures for Image ClassificationabstractThis paper describes a method of structural pattern recognition based on a genetic evolution processing of data structures with neural networks representation. Conventionally, one of the most popular learning formulations of data structure processing is backpropagation through structures (BPTS) [C. Goller et al., (1996)]. The BPTS algorithm has been successfully applied to a number of learning tasks that involved structural patterns such as image, shape, and texture classifications. However, this BPTS typed algorithm suffers from the long-term dependency problem in learning very deep tree structures. In this paper, we propose the genetic evolution for this data structures processing. The idea of this algorithm is to tune the learning parameters by the genetic evolution with specified chromosome structures. Also, the fitness evaluation as well as the adaptive crossover and mutation for this structural genetic processing are investigated in this paper. An application to flowers image classification by a structural representation is provided for the validation of our method. The obtained results significantly support the capabilities of our proposed approach to classify and recognize flowers in terms of generalization and noise robustness. Siu-Yeung Cho, Zheru Chi |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2005 | Zeroing polynomials using modified constrained neural network approachabstractThis paper proposes new modified constrained learning neural root finders (NRFs) of polynomial constructed by backpropagation network (BPN). The technique is based on the relationships between the roots and the coefficients of polynomial as well as between the root moments and the coefficients of the polynomial. We investigated different resulting constrained learning algorithms (CLAs) based on the variants of the error cost functions (ECFs) in the constrained BPN and derived a new modified CLA (MCLA), and found that the computational complexities of the CLA and the MCLA based on the root-moment method (RMM) are the order of polynomial, and that the MCLA is simpler than the CLA. Further, we also discussed the effects of the different parameters with the CLA and the MCLA on the NRFs. In particular, considering the coefficients of the polynomials involved in practice to possibly be perturbed by noisy sources, thus, we also evaluated and discussed the effects of noises on the two NRFs. Finally, to demonstrate the advantage of our neural approaches over the nonneural ones, a series of simulating experiments are conducted. De-Shuang Huang, Horace Ho-Shing Ip, Ken Chee-keung Law, Zheru Chi |
IEEE Trans. Neural Networks | 4 |
| 2004 | Machine learning techniques for ontology-based leaf classificationabstractLeaf classification, indexing as well as retrieval is an important part of a computerized plant identification system. In this paper, an integrated approach for an ontology-based leaf classification system is proposed, wherein machine learning techniques play a crucial role for the automatization of the system. For the leaf contour classification, a scaled CCD code system is proposed to categorize the basic shape and margin type of a leaf by using the similar taxonomy principle adopted by the botanists. Then a trained neural network is employed to recognize the detailed tooth patterns. The measurement on an unlobed leaf is also conducted automatically according to the method used in botany. For the leaf vein recognition, the vein texture is extracted by employing an efficient combined thresholding and neural network approach so as to obtain more vein details of a leaf. Compared with the past studies, the proposed method integrates low-level features of an image and the specific knowledge in the domain (ontology) of botany, and therefore provides a more practical system for users to comprehend and handle. Primary experiments have shown promising results and proven the feasibility of the proposed system. Hong Fu, Zheru Chi, David Dagan Feng, Jiatao Song |
ICARCV | 2 |
| 2004 | Texture image segmentation based on entropy theoryabstractA two-stage segmentation algorithm for textured object image is proposed based on variational approaches and entropy theory. Structure tensor of the image is adopted as the texture feature for the discrimination of texture patterns. In order to make the four information channels cooperatively push the evolving contour towards the textured foreground object, a pre-processing based on mean shift algorithm is applied, which can implement edge-preserving smoothing and has a definite stopping criterion. The Dream/sup 2/s framework is extended to the structure tensor data in the first segmentation stage in which the differential entropy is adopted as regional descriptors, and Gaussian distribution for both the foreground part and the background part is assumed. In the second stage, the nonparametric density distributions are exploited, and the mutual information between the intensity random variables of different information channels and the binary label random variable is adopted as the criterion to minimize. The nonparametric representation can describe the truly nongaussianity of data, and thus can refine the initial segmentation result obtained from the first stage. Zheru Chi |
ICARCV | 2 |
| 2004 | Motion detection and tracking based on level set algorithmabstractA novel energy-minimizing model is proposed in this paper for motion detection and object tracking. In our model, the logarithmic image is employed to compensate for illumination change. Another feature of our approach is that the boundary force is no longer necessary. For solving the corresponding curve evolution equation efficiently, a local level set algorithm is adopted. When constructing the narrow band and resetting the level set function to the signed distance function, it is not needed to explicitly label the contour points during the evolution of contours. The local level set algorithm is based on partial differential equations (PDEs), which leads to a simple, flexible and stable scheme. This paper also proposes an appropriate stopping criterion for the level set algorithm without a need of explicitly extracting the locations of the evolving curve. Zheru Chi, Tianwen Zhang |
ICARCV | 2 |
| 2004 | Extraction of face image edges with application to expression analysisabstractThis paper presents a new method for the extraction of face image edges and a scheme for facial expression analysis based on the binary edge images (BEIs). Making use of the multi-resolution property of wavelet transform, the edge detection method includes two binarization steps and a noise removing step. Our method provides a robust solution to face image edge extraction, in particular, if changes of lighting conditions are encountered. Experimental results show that different components in a BEI are well separated and pixels of the same face component are connected well. Preliminary experimental results on facial expression analysis on the subsets of AR and Yale face databases show that a satisfactory recognition rate can be achieved when our facial analysis scheme is used for recognizing four action units. Jiatao Song, Zheru Chi, Jilin Liu, Hong Fu |
ICARCV | 2 |
| 2004 | Comparison of image partition methods for adaptive image categorization based on structural image representationabstractImage categorization is very helpful for organizing large image databases efficiently, however, it is yet very challenging due to lack of effective image representations. Our previous work showed that structural representations were good at characterizing image contents, since image contents could be exploited from coarse to fine scales through the structures representation and fewer visual features are required. In this paper, several popular image partition methods are investigated for adaptive image categorization based on structural representation. Experimental results on seven categories of scenery images show that both the structure and node attributes are important to categorize image contents. In addition, the more similar the structures of each category, the better the categorization performance. Zhiyong Wang 0001, David Dagan Feng, Zheru Chi |
ICARCV | 3 |
| 2004 | Image retrieval based on Sugeno fuzzy integralabstractIn this paper, we propose a fuzzy similarity measure based on the Sugeno fuzzy integral and apply it to the content-based image retrieval. Fuzzy measure is used to formulate the subjective feedback information. Experimental results show that the proposed method can improve the efficiency and reliability of image retrieval system. Our method of formulating the subjective feedback information performs better than the weighted average operator (WAO) and a Choquet-integral-based similarity measure. Zheru Chi, Gang Chen 0006 |
ICIG | 2 |
| 2004 | Automatic generation of pen-and-ink drawings from photosabstractIn this paper, a new algorithm for automatic generating pen-and-ink drawings from photos is presented. The basic processing step of our system is the edge extraction based on the multiresolution property of wavelet transform. In order to enhance the variations in texture and tone, the proposed algorithm includes two binarization steps and a noise removing step. Experimental results show that the generated pen-and-ink drawings are of high quality, which accurately illustrate the scenic contents of the original photos. Jiatao Song, Zheru Chi, Jilin Liu, Hong Fu |
ICIP | 2 |
| 2004 | Locatization of human eyes based on a series of binary imagesabstractA new method for eyes localization is presented. By incorporating some eye geometrical constraints, our method includes such procedures as eyes image segmentation, eye candidates selection and eyeball detection. In order to locate the eyes more precisely, the valley map transformed from the grayscale face image using an arithmetical morphology method is successively binarized with a series of threshold values determined adaptively, and many possible eyeball candidates are extracted from them. The final position of the two eyes is obtained using a statistical method. Experimental results on AR and Yale face databases show that a locating rate of over 94% and an average locating disparity of below 2 pixels can be achieved Jiatao Song, Jilin Liu, Zheru Chi, Wei Wang 0106 |
ICME | 3 |
| 2004 | Finding roots of arbitrary high order polynomials based on neural network recursive partitioning method
De-Shuang Huang, Zheru Chi |
Sci. China Ser. F Inf. Sci. | 2 |
| 2004 | A Neural Root Finder of Polynomials Based on Root MomentsabstractThis letter proposes a novel neural root finder based on the root moment method (RMM) to find the arbitrary roots (including complex ones) of arbitrary polynomials. This neural root finder (NRF) was designed based on feedforward neural networks (FNN) and trained with a constrained learning algorithm (CLA). Specifically, we have incorporated the a priori information about the root moments of polynomials into the conventional backpropagation algorithm (BPA), to construct a new CLA. The resulting NRF is shown to be able to rapidly estimate the distributions of roots of polynomials. We study and compare the advantage of the RMM-based NRF over the previous root coefficient method—based NRF and the traditional Muller and Laguerre methods as well as the mathematica roots function, and the behaviors, the accuracies of the resulting root finders, and their training speeds of two specific structures corresponding to this FNN root finder: the log σand the σ FNN. We also analyze the effects of the three controlling parameters {δP0 θp η} with the CLA on the two NRFs theoretically and experimentally. Finally, we present computer simulation results to support our claims. De-Shuang Huang, Horace Ho-Shing Ip, Zheru Chi |
Neural Comput. | 3 |
| 2004 | Image coding quality assessment using fuzzy integrals with a three-component image modelabstractBased on importance measures and fuzzy integrals, a new assessment method for image coding quality is presented in this paper. The proposed assessment is based on two subevaluations. In the first subevaluation, errors on edges, textures, and flat regions are computed individually. The errors are then assessed using an assessment function. A global evaluation with Sugeno fuzzy integral is then obtained based on the importance measure of edge, texture, and flat region. In the second subevaluation, an importance measure is first established depending on the types of regions where errors occur, a subtle evaluation is then obtained using Sugeno fuzzy integral on all pixels of the image. A final evaluation is obtained based on the two subevaluations. Experimental results show that this new image quality assessment closely approximates human subjective tests such as mean opinion score with a high correlation coefficient of 0.963, which is a significant improvement over peak signal-to-noise ratio, picture quality scale, and weighted mean square error, three other image coding quality assessment methods, which have the correlation coefficients of 0.821, 0.875, and 0.716, respectively. Gang Chen 0006, Zheru Chi, Chenggang Lu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2004 | Tracer kinetic modeling of 11C-acetate applied in the liver with positron emission tomographyabstractIt is well known that 40%-50% of hepatocellular carcinoma (HCC) do not show increased 18F-fluorodeoxyglucose (FDG) uptake. Recent research studies have demonstrated that 11C-acetate may be a complementary tracer to FDG in positron emission tomography (PET) imaging of HCC in the liver. Quantitative dynamic modeling is, therefore, conducted to evaluate the kinetic characteristics of this tracer in HCC and nontumor liver tissue. A three-compartment model consisting of four parameters with dual inputs is proposed and compared with that of five parameters. Twelve regions of dynamic datasets of the liver extracted from six patients are used to test the models. Estimation of the adequacy of these models is based on Akaike Information Criteria (AIC) and Schwarz Criteria (SC) by statistical study. The forward clearance K = K1 * k3/(k2 + k3) is estimated and defined as a new parameter called the local hepatic metabolic rate-constant of acetate (LHMRAct) using both the weighted nonlinear least squares (NLS) and the linear Patlak methods. Preliminary results show that the LHMRAct of the HCC is significantly higher than that of the nontumor liver tissue. These model parameters provide quantitative evidence and understanding on the kinetic basis of C-acetate for its potential role in the imaging of HCC using PET. Sirong Chen, Chilai Ho, David Dagan Feng, Zheru Chi |
IEEE Trans. Medical Imaging | 4 |
| 2003 | Region-of-interest based flower images retrievalabstractFlower image retrieval is a very important step for computer-aided plant species identification. We propose an efficient segmentation method based on color clustering and domain knowledge to extract flower regions from flower images. For flower retrieval, we use the color histogram of a flower region to characterize the color features of a flower and two shape-based sets of features, centroid-contour-distance (CCD) and angle code histogram (ACH), to characterize the shape features of a flower contour. Experimental results show that our flower region extraction approach based on color clustering and domain knowledge can achieve accurate flower regions. The retrieval results on a database of 885 flower images collected from 14 plant species show that our region-of-interest (ROI) based retrieval approach can perform better than the Swain method based on the global color histogram (Swain, M.J. and Ballard, D.H., Int. J. of Computer Vision, vol.7, no.1, p.11-32, 1991). Anxiang Hong, Zheru Chi, Gang Chen 0006, Zhiyong Wang 0001 |
ICASSP (3) | 2 |
| 2003 | An algorithm for segmenting moving vehiclesabstractAn algorithm for region-based moving object segmentation is presented. The gray-scale image segmentation based on the mean shift algorithm (MSA) is first performed to segment each frame of a sequence into connective homogeneous regions. A method applying spatio-temporal continuity constraints of the motion vector image is then carried out to detect moving pixels robustly. Finally, each homogeneous region is labeled as either a moving-object region or a non-moving-object region according to the number of moving pixels it contains. Experimental results show that our algorithm is effective and robust in segmenting moving vehicles from noisy scenes. Su Zhang 0001, Hanfeng Chen, Zheru Chi |
ICASSP (3) | 3 |
| 2003 | Efficient Learning in Adaptive Processing of Data Structures
Siu-Yeung Cho, Zheru Chi, Zhiyong Wang 0001, Wan-Chi Siu |
Neural Process. Lett. | 2 |
| 2003 | Document Image Recognition Based on Template Matching of Component Block ProjectionsabstractDocument Image Recognition (DIR), a very useful technique in office automation and digital library applications, is to find the most similar template for any input document image in a prestored template document image data set. Existing methods use both local features and global layout information. In this paper, we propose a novel algorithm based on the global matching of Component Block Projections (CBP), which are the concatenated directional projection vectors of the component blocks of a document image. Compared to those existing methods, CBP-based template-matching methods possess two major advantages: (1) The spatial relationship among the component blocks of a document image is better represented, hence a very high matching accuracy can be obtained even for a large template set and seriously distorted input images; and (2) the effective matching distance of each template and the triangle inequality are proposed to significantly reduce the computational cost. Our experimental results confirm these advantages and show that the CBP-based template-matching methods are very suitable for DIR applications. Hanchuan Peng, Fuhui Long, Zheru Chi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Hierarchical content classification and script determination for automatic document image processing
Zheru Chi, Qing Wang 0006, Wan-Chi Siu |
Pattern Recognit. | 1 |
| 2003 | Two-stage segmentation of unconstrained handwritten Chinese character
Shuyan Zhao, Zheru Chi, Hong Yan 0001 |
Pattern Recognit. | 2 |
| 2003 | An improved algorithm for learning long-term dependency problems in adaptive processing of data structuresabstractMany researchers have explored the use of neural-network representations for the adaptive processing of data structures. One of the most popular learning formulations of data structure processing is backpropagation through structure (BPTS). The BPTS algorithm has been successful applied to a number of learning tasks that involve structural patterns such as logo and natural scene classification. The main limitations of the BPTS algorithm are attributed to slow convergence speed and the long-term dependency problem for the adaptive processing of data structures. In this paper, an improved algorithm is proposed to solve these problems. The idea of this algorithm is to optimize the free learning parameters of the neural network in the node representation by using least-squares-based optimization methods in a layer-by-layer fashion. Not only can fast convergence speed be achieved, but the long-term dependency problem can also be overcome since the vanishing of gradient information is avoided when our approach is applied to very deep tree structures. Siu-Yeung Cho, Zheru Chi, Wan-Chi Siu, Ah Chung Tsoi |
IEEE Trans. Neural Networks | 2 |
| 2002 | Fuzzy integral for leaf image retrievalabstractGenerally, the more features utilized, the better the retrieval performance. However, it is a very challenging task to combine different feature sets in a way reflecting human perception. This paper presents the combination of different shape based feature sets using fuzzy integral for leaf image retrieval. The feature sets used in our system include centroid-contour distance curve, eccentricity, and angle code histogram. The fuzzy integral approach can release the user's burden from tuning the combination parameters. In order to reduce the matching time in the retrieval process, a thinning based method is proposed to locate the start point of a leaf contour. Experimental results on 440 leaf images from 44 plant species (10 samples from each plant species) show that the fuzzy integral approach can achieve a comparable retrieval performance with the best case of the weighted summation combination. The results also indicate that our approach, which are more efficient, can achieve a better retrieval performance than both the curvature scale space (CSS) method and the modified Fourier descriptor (MFD) method. Zhiyong Wang 0001, Zheru Chi, David Dagan Feng |
FUZZ-IEEE | 2 |
| 2002 | A fast 2D entropic thresholding method by wavelet decompositionabstractCompared with ID grayscale histogram analysis, 2D entropic thresholding makes use of local average as well as pixel gray level. However, it is time consuming to search the threshold vector in the 2D histogram. In the paper, a fast algorithm using wavelet decomposition is proposed, with which a set of candidates of the vector was first obtained in the decomposed histogram. The optimal threshold vector is then obtained without exhaustive searching. Experimental results have shown that our algorithm not only finds the threshold vector as well as Brink's method (1992) but also saves computation costs, using up only 0.53% of the processing time taken by exhaustive searching. Qing Wang 0006, Qiurang Wang, David Dagan Feng, Rongchun Zhao, Zheru Chi |
ICIP (3) | 5 |
| 2002 | Image Thresholding by Maximizing the Index of Nonfuzziness of the 2-D Grayscale Histogram
Qing Wang 0006, Zheru Chi, Rongchun Zhao |
Comput. Vis. Image Underst. | 2 |
| 2002 | Extraction and Optimization of B-Spline PBD Templates for Recognition of Connected Handwritten Digit StringsabstractThe recognition of connected handwritten digit strings is a challenging task due mainly to two problems: poor character segmentation and unreliable isolated character recognition. The authors first present a rational B-spline representation of digit templates based on Pixel-to-Boundary Distance (PBD) maps. We then present a neural network approach to extract B-spline PBD templates and an evolutionary algorithm to optimize these templates. In total, 1000 templates (100 templates for each of 10 classes) were extracted from and optimized on 10426 training samples from the NIST Special Database 3. By using these templates, a nearest neighbor classifier can successfully reject 90.7 percent of nondigit patterns while achieving a 96.4 percent correct classification of isolated test digits. When our classifier is applied to the recognition of 4958 connected handwritten digit strings (4555 2-digit, 355 3-digit, and 48 4-digit strings) from the NIST Special Database 3 with a dynamic programming approach, it has a correct classification rate of 82.4 percent with a rejection rate of as low as 0.85 percent. Our classifier compares favorably in terms of correct classification rate and robustness with other classifiers that are tested. Zhongkang Lu, Zheru Chi, Wan-Chi Siu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | A fuzzy image metric with application to fractal codingabstractImage quality assessment is an important issue addressed in various image processing applications such as image/video compression and image reconstruction. The peak signal-to-noise ratio (PSNR) with the L(2)-metric is commonly used in objective image quality assessment. However, the measure does not agree very well with the human visual perception in many cases. A fuzzy image metric (FIM) is defined based on Sugeno's (1977) fuzzy integral. This new objective image metric, which is to some extent a proper evaluation from the viewpoint of the judgment procedure, is closely approximates the subjective mean opinion score (MOS) with a correlation coefficient of about 0.94, as compared to 0.82 obtained using the PSNR. Compared to the L(2)-metric, we demonstrate that a better performance can be achieved in fractal coding by using the proposed FIM. Gang Chen 0006, Zheru Chi |
IEEE Trans. Image Process. | 3 |
| 2001 | A Fuzzy Metric for Image Quality AssessmentabstractImage quality assessment is an important issue addressed in various image processing applications such as image/video compression and image reconstruction. The peak signal-to-noise ratio (PSNR) with the L/sup 2/-metric is commonly used in objective image quality assessment. However, the measure does not agree very well with human visual perception in many cases. In this paper, a fuzzy image metric (FIM) is defined based on M. Sugeno's (1974) fuzzy integral. This new objective image metric, which is to some extent a proper evaluation from the viewpoint of the judgement procedure, is closely approximates the subjective mean opinion score (MOS) with a correlation coefficient of about 0.94, as compared to 0.82 obtained using PSNR. Zheru Chi, Gang Chen 0006 |
FUZZ-IEEE | 1 |
| 2001 | Match Between Normalization Schemes and Feature Sets for Handwritten Chinese Character RecognitionabstractBecause of the large number of Chinese characters and many different writing styles involved, the recognition of handwritten Chinese characters remains a very challenging task. It is well recognized that a good feature set plays a key role in a successful recognition system. Shape normalization is as well an essential step toward achieving translation, scale, and rotation invariance in recognition. Many shape normalization methods and different feature sets have been proposed in the literature. We first review five commonly used shape normalization schemes and then discuss various feature extraction techniques usually used in handwritten Chinese character recognition. Based on numerous experiments conducted on 3,755 handwritten Chinese characters (GB2312-80), we discuss the matches made between the normalization schemes and the feature sets and suggest the best match between them in terms of classification performance. The nearest neighbor classifier was adopted in our experiments with templates obtained by using the K-means clustering algorithm. Qing Wang 0006, Zheru Chi, David Dagan Feng, Rongchun Zhao |
ICDAR | 2 |
| 2001 | Handwritten Chinese Character Segmentation Using a Two-Stage ApproachabstractCorrect segmentation of handwritten Chinese characters is crucial to the successful recognition. However, because of the many difficulties involved, little work has been done in this area. In this paper, a two-stage approach is addressed to segment unconstrained handwritten Chinese character strings. A string is first coarsely segmented according to the background skeleton and vertical projection after a proper image preprocessing. At the fine segmentation stage that follows, the strokes that may contain segmentation points are first identified. The feature points are then extracted from candidate strokes and taken as segmentation point candidates through each of which a segmentation path may be formed. Geometric features are extracted and fuzzy decision rules learned from examples are used to evaluate the segmentation paths. By using this two-stage segmentation approach, we can achieve both good performance and efficiency in segmenting unconstrained handwritten Chinese characters. Shuyan Zhao, Zheru Chi, Qing Wang 0006 |
ICDAR | 2 |
| 2001 | ART2 Neural Network for Surface EMG Decomposition
ZhengQuan Xu, Shaojun Xiao, Zheru Chi |
Neural Comput. Appl. | 3 |
| 2001 | Document image template matching based on component block list
Hanchuan Peng, Fuhui Long, Zheru Chi, Wan-Chi Siu |
Pattern Recognit. Lett. | 3 |
| 2000 | A hierarchical distributed genetic algorithm for image segmentationabstractA novel hierarchical distributed genetic algorithm is proposed for image segmentation. Firstly, a technique of histogram dichotomy is proposed to explore the statistical property of input image and produce a hierarchical quantization image. Then a hierarchical distributed genetic algorithm (HDGA) is imposed on the quantized image to explore the spatial connectivity and produce final segmentation result. HDGA is a major improvement of the original distributed genetic algorithm (DGA) and multiscale distributed genetic algorithm (MDGA) in four aspects: (1) HDGA does not require the a priori number of image regions, however it can effectively and adaptively control the segmentation quality; (2) the chromosome structure is revised from the original label (multilabel)-condition-fitness format to a more compact (storage-efficient) label-fitness format; (3) the fitness function is revised to utilize the spatial connectivity, but not the original "reconstruction" error; (4) three revised genetic operations are presented to make the algorithm computation-efficient. Our experiments give proofs for the advantages of HDGA. Hanchuan Peng, Fuhui Long, Zheru Chi, Wanchi Su |
CEC | 3 |
| 2000 | Document Image Matching Based on Component BlocksabstractDocument image matching is the key technique for document registration and retrieval. In this paper, a new matching algorithm based on document component block list and component block tree is proposed. Our method can effectively make use of the local information of each page block and the global information of page layout, while it is also robust to image distortion, filled-in text, and noises. This algorithm is then refined and applied to automatic data extraction of column forms. A demonstrating software package has been developed. Hanchuan Peng, Fuhui Long, Wan-Chi Siu, Zheru Chi, David Dagan Feng |
ICIP | 4 |
| 2000 | Hidden Markov Random Field Based Approach for Off-Line Handwritten Chinese Character RecognitionabstractThis paper presents a hidden Markov mesh random field (HMMRF) based approach for off-line handwritten Chinese characters recognition using statistical observation sequences embedded in the strokes of a character. Due to a large set of Chinese characters and many different writing styles, the recognition of handwritten Chinese characters is very challenging. In our approach, the binary image is first normalized by a nonlinear shape normalization scheme to adjust the width, length, and the correlation of strokes. Two types of stroke-based features are then extracted to represent the observation sequence. The estimation of model parameters and state sequence decoding algorithms are also discussed in the paper. Experimental results on 470 isolated handwritten Chinese characters demonstrate the effectiveness of our approach. Qing Wang 0006, Rongchun Zhao, Zheru Chi, David Dagan Feng |
ICPR | 3 |
| 2000 | A Semi-Parametric Hybrid Neural Model for Nonlinear Blind Signal SeparationabstractNonlinear blind signal separation is an important but rather difficult problem. Any general nonlinear independent component analysis algorithm for such a problem should specify which solution it tries to find. Several recent neural networks for separating the post nonlinear blind mixtures are limited to the diagonal nonlinearity, where there is no cross-channel nonlinearity. In this paper, a new semi-parametric hybrid neural network is proposed to separate the post nonlinearly mixed blind signals where cross-channel disturbance is included. This hybrid network consists of two cascading modules, which are a neural nonlinear module for approximating the post nonlinearity and a linear module for separating the predicted linear blind mixtures. The nonlinear module is a semi-parametric expansion made up of two sub-networks, one of which is a linear model and the other of which is a three-layer perceptron. These two sub-networks together produce a "weak" nonlinear operator and can approach relatively strong nonlinearity by tuning parameters. A batch learning algorithm based on the entropy maximization and the gradient descent method is deduced. This model is successfully applied to a blind signal separation problem with two sources. Our simulation results indicate that this hybrid model can effectively approach the cross-channel post nonlinearity and achieve a good visual quality as well as a high signal-to-noise ratio in some cases. Hanchuan Peng, Zheru Chi, Wan-Chi Siu |
Int. J. Neural Syst. | 2 |
| 1999 | Approaching the post nonlinearity of blind mixtures by hybrid neural networkabstractIt is very difficult to approach the post nonlinearity of blind mixtures. The recent neural networks for separating the post nonlinear blind mixtures are limited to the diagonal nonlinearity. In this paper a hybrid neural network is proposed to separate the post nonlinearly mixed blind signals with cross-channel disturbance. This hybrid network consists of a new neural blind de-mixer for approximating the post nonlinearity and a common network for separating the predicted linear mixtures. The blind de-mixer is made up of two subnets, which in total produce a "weak" nonlinear operator and can approach relatively strong nonlinearity by parameter-tuning. A six-step batch learning algorithm based on the fixed-point algorithm and information backpropagation is deduced. Preliminary results on a blind signal separation problem of two sources and four different types of post nonlinearity indicate the effectiveness of our model. Hanchuan Peng, Zheru Chi |
IJCNN | 2 |
| 1999 | A background-thinning-based approach for separating and recognizing connected handwritten digit strings
Zhongkang Lu, Zheru Chi, Wan-Chi Siu |
Pattern Recognit. | 2 |
| 1998 | A background-thinning based algorithm for separating connected handwritten digit stringsabstractMost algorithms for segmenting connected handwritten digit strings are based on the analysis of the foreground pixel distributions and the features on the upper/lower contours of the image. A new approach is presented to segment connected handwritten two-digit strings based on the thinning of background regions. The algorithm first locates several feature points on the background skeleton of the digit image. Possible segmentation paths are then constructed by matching these feature points. With geometric property measures, these segmentation paths are ranked using fuzzy rules generated from a decision-tree approach. Finally, the top ranked segmentation paths are tested one by one by an optimized nearest neighbor classifier until one of these candidates is accepted based on an acceptance criterion. Experimental results on NIST special database 3 show that our approach can achieve a correct classification rate of 92.4% with only 4.7% of digit strings rejected, which compares favorably with the other techniques tested. Zhongkang Lu, Zheru Chi |
ICASSP | 2 |
| 1996 | Handwritten digit recognition using combined ID3-derived fuzzy rules and Markov chains
Zheru Chi, Mark Suters, Hong Yan 0001 |
Pattern Recognit. | 1 |
| 1996 | ID3-derived fuzzy rules and optimized defuzzification for handwritten numeral recognitionabstractPresents a technique to produce fuzzy rules based on the ID3 approach and to optimize defuzzification parameters by using a two-layer perceptron. The technique overcomes the difficulties in a conventional syntactic approach to handwritten character recognition, including problems of choosing a starting or reference point, scaling, and learning by machines. The authors' technique provides: a way to produce meaningful and simple fuzzy rules; a method to fuzzify ID3-derived rules to deal with uncertain, noisy, or fuzzy data; and a framework to incorporate fuzzy rules learned from the training data and those extracted from human recognition experience. The authors' experimental results on NIST Special Database 3 show that the technique out-performs the straight forward ID3 approach. Moreover, ID3-derived fuzzy rules can be combined with an optimized nearest neighbor classifier, which uses intensity features only, to achieve a better classification performance than either of the classifiers. The combined classifier achieves a correct classification rate of 98.6% on the test set. Zheru Chi, Hong Yan 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 1995 | Handwritten numeral recognition using a small number of fuzzy rules with optimized defuzzification parameters
Zheru Chi, Hong Yan 0001 |
Neural Networks | 1 |
| 1995 | Handwritten numeral recognition using self-organizing maps and fuzzy rules
Zheru Chi, Hong Yan 0001 |
Pattern Recognit. | 1 |