Hong Fu

dblp:14/1706 · DBLP profile ↗
← Back
50ranked-venue papers
9as first author
14since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 7Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 Physiological navigation amplifier for remote extracting PPG signals from face video clips
Hong Fu
Neurocomputing3
2025 VIOMA: Video-Based Intelligent Ocular Misalignment Assessment
abstract
The measurement of ocular alignment is critical for the diagnosis of strabismus. Current clinical methods for assessing ocular misalignment are subjective and frequently rely on the expertise of practitioners and the extent of patient cooperation. Computer-aided diagnosis methods in recent years have improved automation and precision of measurement, but still, fall short of the requirement of clinical practice. In this study, a video-based intelligent ocular misalignment assessment (VIOMA) system, was proposed to provide an objective, repeatable, user-friendly and highly-automated alternative modality for clinical ocular misalignment measurement, in which the automatic cover tests were performed under a control and motor unit, simultaneously the eye movements were tracked using a motion-capture module and assessed through video analysis techniques, determining the presence, type, and magnitude of eye deviation. For system evaluation, an automatic cover tests video dataset for strabismus (StrabismusACT-76) was established, which consists of data from 76 participants. The Bland-Altman plot, used to compare the results of the VIOMA system and human expert, showed a mean value of 1.26 prism diopter (PD) and a half-width of the 95% limit of agreement of ±7.17 PD. VIOMA system presented a mean absolute error of 3.04 PD in measuring the deviation magnitude, within a 5 PD error tolerance. Additionally, the system’s measurements were strongly correlated with that of video labeling with the mean value of -0.26 PD, a half-width of the 95% limit of agreement of ±3.56, and the average error of 1.31 PD. The experiment results indicated that the proposed method has the capability to offer accurate and efficient assessment of ocular misalignment. Note to Practitioners—The motivation behind this work stems from the need to develop an accurate and efficient system for automated ocular misalignment assessment. The subjectivity of manual cover test performed by examiners has led to variability in outcomes, and certain existing computerized methods have limitations in terms of automation, measurement accuracy, and applicability in clinical practice. Faced with these challenges, we proposed VIOMA, by establishing an apparatus for automatic implementation of cover tests and developing assessment algorithms based on strabismic video analysis. This system can objectively and precisely measure ocular misalignment, offering a promising practical solution for clinical intelligent diagnosis of strabismus. The VIOMA system’s potential applications are not limited to strabismus but may extend to other eye-related conditions and beyond.
Yang Zheng 0006, Hong Fu, Carly Siu Yin Lam, Jimin Liang, Kaitai Guo
IEEE Trans Autom. Sci. Eng.2
2025 Power Losses Optimization of MMCs Based on Quantum Genetic Algorithm for HVdc Transmission Application
abstract
Modular multilevel converters (MMCs) obtain widespread utilization in high-voltage direct current (HVdc) applications scenarios. The cost of power losses plays a significant part in MMC’s operating costs. Hence, this article proposes a quantum genetic algorithm-based power losses optimization control (QGA-PLOC). By comprehensively considering the power losses of the MMC, the quantum genetic algorithm determines the first-best value of the injected second circulating current magnitude and phase angle in the arm, as well as the optimal MMC power losses under given conditions. The quantum genetic algorithm incorporates the quantum state vector representation into genetic encoding and utilizes quantum logic gates for chromosome evolution, greatly improving the algorithm’s performance and significantly enhancing its computational efficiency and global optimization capability. Moreover, introducing the quantum genetic algorithm into the field of MMC power losses optimization offers a new path to address the acquisition of the optimal circulating current reference value in power loss optimization problems. MMC Simulation and experiment are also conducted, and the research results verify the effectiveness of the proposed QGA-PLOC for MMCs.
Jifeng Zhao, Peidong Xu, Xinyue Wu, Jia Pei, Hong Fu, Yutan Li
IEEE Trans. Circuits Syst. I Regul. Pap.5
2024 Progressive generative adversarial network for generating high-dimensional and wide-frequency signals in intelligent fault diagnosis
Zhijun Ren, Yongsheng Zhu, Ke Feng 0004, Zheng Liu 0002, Hong Fu, Jun Hong 0002, Adam Glowacz
Eng. Appl. Artif. Intell.6
2024 Strategic promotion decisions of competing mobile application suppliers in Stackelberg game context
Lulu Xia, Kai Li 0019, Nenggui Zhao, Hong Fu, Bohai Liu
Expert Syst. Appl.4
2024 Characterising Eye Movement Events With Multi-Scale Spatio-Temporal Awareness
abstract
The intricate and dynamic nature of eye movements serves as a window into the realms of cognition, emotion, and physiological responses. Event detection, in turn, is instrumental in the precise recognition and categorization of these diverse eye movements. Deep learning methods have recently been applied to event detection, yielding promising results. However, the intrinsic multi-scale attributes of events have often been overlooked in existing approaches. To address this, we introduce “GazeUNet”, a novel network based on U-Net and Bi-GRU, which classifies gaze samples into three categories: fixation, saccade, and post-saccadic oscillations. Firstly, multi-scale spatial features are captured using a U-Net model, and then a hierarchical bidirectional gated recurrent unit (Bi-GRU) is employed to extract temporal correlations, followed by classification through fully connected layers. Our results, derived from the analysis of three publicly available datasets, consistently showcase the superiority of the proposed model compared with other state-of-the-art methods across all categories.
Yang Zheng 0006, Hong Fu, Kaitai Guo, Jimin Liang
IEEE Signal Process. Lett.3
2024 PR-PL: A Novel Prototypical Representation Based Pairwise Learning Framework for Emotion Recognition Using EEG Signals
abstract
Affective brain-computer interface based on electroencephalography (EEG) is an important branch in the field of affective computing. However, the individual differences in EEG emotional data and the noisy labeling problem in the subjective feedback seriously limit the effectiveness and generalizability of existing models. To tackle these two critical issues, we propose a novel transfer learning framework with Prototypical Representation based Pairwise Learning (PR-PL). The discriminative and generalized EEG features are learned for emotion revealing across individuals and the emotion recognition task is formulated as pairwise learning for improving the model tolerance to the noisy labels. More specifically, a prototypical learning is developed to encode the inherent emotion-related semantic structure of EEG data and align the individuals' EEG features to a shared common feature space under consideration of the feature separability of both source and target domains. Based on the aligned feature representations, pairwise learning with an adaptive pseudo labeling method is introduced to encode the proximity relationships among samples and alleviate the label noises effect on modeling. Extensive results on two benchmark databases (SEED and SEED-IV) under four different cross-validation evaluation protocols validate the model reliability and stability across subjects and sessions. Compared to the literature, the average enhancement of emotion recognition across four different evaluation protocols is 2.04% (SEED) and 2.58% (SEED-IV). The source code is available athttps://github.com/KAZABANA/PR-PL.
Rushuang Zhou, Zhiguo Zhang 0001, Hong Fu, Li Zhang 0041, Linling Li, Fali Li, Xin Yang 0009, Yining Dong, Yuan-Ting Zhang
IEEE Trans. Affect. Comput.3
2023 Serial-parallel multi-scale feature fusion for anatomy-oriented hand joint detection
Bin Li 0051, Hong Fu
Neurocomputing4
2023 ECG signal reconstruction based on facial videos via combined explicit and implicit supervision
Bin Li 0051, Wei Zhang 0133, Hong Fu, Feng Xu 0003
Knowl. Based Syst.4
2023 Non-contact PPG signal and heart rate estimation with multi-hierarchical convolutional network
Bin Li 0051, Jinye Peng 0001, Hong Fu
Pattern Recognit.4
2023 Sound Event Classification Based on Frequency-Energy Feature Representation and Two-Stage Data Dimension Reduction
abstract
The classification of environmental sound events is of great significance for applications such as machine hearing and acoustic surveillance. Feature representation and feature vector dimension directly affect system performance. To better extract features and reduce computational burden, a novel frequency-energy feature representation and two-stage dimension reduction system were proposed. First, a frequency-energy diagram is generated. Based on this, the importance screening is done and only the energy bins of high importance are retained, which reduces the dimension of feature vector while extracting key information. Then the Bicubic interpolation method is used to further reduce the dimension. And the appropriate feature vector dimension is determined based on the change of information entropy. The proposed frequency-energy feature representation and two-stage dimension reduction system are evaluated with Real Word Computing Partnership sound scene database (RWCP-SSD), UrbanSound8K, and ESC-50 datasets, which demonstrate that the robustness is satisfactory under low signal-to-noise ratios (SNRs) and 15 noise types from NOISEX-92 database.
Yinggang Liu, Hong Fu, Ying Wei 0004, Hanbing Zhang
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Multi-Level Constrained Intra and Inter Subject Feature Representation for Facial Video Based BVP Signal Measurement
abstract
Facial video-based blood volume pulse (BVP) signal measurement holds great potential for remote health monitoring, while existing methods have issues with convolutional kernel perceptual field constraints. This article proposes an end-to-end multi-level constrained spatiotemporal representation structure for facial video-based BVP signal measurement. First, an intra- and inter-subject feature representation is proposed to strengthen the BVP-related features generation at high, semantic, and shallow levels, respectively. Second, the global-local association is presented to enhance BVP signal period pattern learning, and the global temporal features are introduced into the local spatial convolution of each frame by adaptive kernel weights. Finally, the multi-dimensional fused features are mapped to one-dimensional BVP signals by the task-oriented signal estimator. The experimental results on the publicly available MMSE-HR dataset demonstrate that the proposed structure overperforms state-of-the-art methods (e.g., AutoHR) in BVP signal measurement, with a 20% and 40% reduction in mean absolute error and root mean squared error, respectively. The proposed structure would be a powerful tool for telemedical and non-contact heart health monitoring.
Bin Li 0051, Wei Zhang 0133, Hong Fu, Hao Liu 0099, Feng Xu 0003
IEEE J. Biomed. Health Informatics3
2022 Adaptive cost-sensitive learning: Improving the convergence of intelligent diagnosis models under imbalanced data
Zhijun Ren, Yongsheng Zhu, Hong Fu, Qingbo Niu, Jun Hong 0002
Knowl. Based Syst.4
2021 Deep neural network oriented evolutionary parametric eye modeling
Yang Zheng 0006, Hong Fu, Richard T. C. Hsung, Zongxi Song, Desheng Wen
Pattern Recognit.2
2020 Precise Temporal Localization for Complete Actions with Quantified Temporal Structure
abstract
Existing temporal action detection algorithms cannot distinguish complete and incomplete actions while this property is essential in many applications. To tackle this challenge, we proposed the action progression networks (APN), a novel model that predicts action progression of video frames with continuous numbers. Using the progression sequence of test video, on the top of the APN, a complete action searching algorithm (CAS) was designed to detect complete actions only. With the usage of frame-level fine-grained temporal structure modeling and detecting actions according to their whole temporal context, our framework can locate actions precisely and is good at avoiding incomplete action detection. We evaluated our framework on a new dataset (DFMAD-70) collected by ourselves which contains both complete and incomplete actions. Our framework got good temporal localization results with 95.77% average precision when the IoU threshold is 0.5. On the benchmark THUMOS14, an incomplete-ignostic dataset, our framework still obtain competitive performance. The code is available online at https://github.com/MakeCent/Action-Progression-Network.
Chongkai Lu, Hong Fu, Zheru Chi
ICPR3
2020 A multi-Internet service provider game: Equilibrium, stability, and characteristics
abstract
Summary This paper proposes a multi‐Internet service provider (ISP) game model for investigating the economic interactions among ISPs who compete or cooperate with each other for customers. The equilibrium and stability properties of the multi‐ISP game in different market scenarios, that is, competitive market, collaborative market, mixed market, and winner‐take‐all market, are investigated. Some qualitative and managerial insights into market outcomes are obtained to characterize the set of equilibria. Furthermore, numerical examples are also employed to illustrate the typical characteristics of the multi‐ISP market. The simulation results provide possible explanations for the evolution of the multi‐ISP game. The results also indicate that the constructed model can predict the main characteristics of the Internet service market in different market scenarios.
Hong Fu
Concurr. Comput. Pract. Exp.1
2019 Development of a Continuous Vertical-pulling Automatic Doffing Robot for the Ring Spinning
abstract
Doffing robot is an important part of the spinning process in the textile production. This paper analyzes the doffing process of spinning machines and points out the requirements of the structure and functions of the doffer. The locking two-finger gripper, the three-dimensional circulating operation mechanism, the collaborating locating mechanism with the toothed disc and the pre-loosening mechanism by rotating spindles are designed. On this basis the continuous vertical-pulling automatic doffing robot, named CVP doffing robot, for the ring spinning is developed. The kinematics and dynamics analysis of the CVP doffing robot are carried out. The structural parameters of the CVP doffing robot are optimized by establishing kinematics and dynamics models. The forces of pulling out cops before and after the pre-loosing operation are tested. On this basis, the strength of the key components is designed and checked. Finally, the performance of the CVP doffing robot is verified by the doffing experiment.
Wenzeng Zhang, Siyun Liu, Hong Fu
IROS5
2019 Correction to: Retinal vessel extraction using dynamic multi-scale matched filtering and dynamic threshold processing based on histogram fitting
Duoduo Gou, Ying Wei 0004, Hong Fu
Mach. Vis. Appl.3
2018 When Boys Are More Generous Than Girls: Effects of Gender and Coordination Level on Prosocial Behavior in 4-year-old Chinese Children
Yingjia Wan, Hong Fu, Michael K. Tanenhaus
CogSci2
2018 Retinal vessel extraction using dynamic multi-scale matched filtering and dynamic threshold processing based on histogram fitting
Duoduo Gou, Ying Wei 0004, Hong Fu
Mach. Vis. Appl.3
2018 Facial Expression Recognition in Video with Multiple Feature Fusion
abstract
Video based facial expression recognition has been a long standing problem and attracted growing attention recently. The key to a successful facial expression recognition system is to exploit the potentials of audiovisual modalities and design robust features to effectively characterize the facial appearance and configuration changes caused by facial motions. We propose an effective framework to address this issue in this paper. In our study, both visual modalities (face images) and audio modalities (speech) are utilized. A new feature descriptor called Histogram of Oriented Gradients from Three Orthogonal Planes (HOG-TOP) is proposed to extract dynamic textures from video sequences to characterize facial appearance changes. And a new effective geometric feature derived from the warp transformation of facial landmarks is proposed to capture facial configuration changes. Moreover, the role of audio modalities on recognition is also explored in our study. We applied the multiple feature fusion to tackle the video-based facial expression recognition problems under lab-controlled environment and in the wild, respectively. Experiments conducted on the extended Cohn-Kanade (CK+) database and the Acted Facial Expression in Wild (AFEW) 4.0 database show that our approach is robust in dealing with video-based facial expression recognition problems under lab-controlled environment and in the wild compared with the other state-of-the-art methods.
JunKai Chen, Zenghai Chen, Zheru Chi, Hong Fu
IEEE Trans. Affect. Comput.4
2017 A new framework with multiple tasks for detecting and locating pain events in video
JunKai Chen, Zheru Chi, Hong Fu
Comput. Vis. Image Underst.3
2017 Smile detection in the wild with deep convolutional neural networks
JunKai Chen, Qihao Ou, Zheru Chi, Hong Fu
Mach. Vis. Appl.4
2016 Facial expression recognition with dynamic Gabor volume feature
abstract
Facial expression recognition is a long standing problem in affective computing community. A key step is extracting effective features from face images. Gabor filters have been widely used for this purpose. However, a big challenge for Gabor filters is its high dimensionality. In this paper, we propose an efficient feature called dynamic Gabor volume feature (DGVF) based on Gabor filters while with a lower dimensionality for facial expression recognition. In our approach, we first apply Gabor filters with multi-scale and multi-orientation to extract different Gabor faces. And these Gabor faces are arranged into a 3-D volume and Histograms of Oriented Gradients from Three Orthogonal Planes (HOG-TOP) are further employed to encode the 3-D volume in a compact way. Finally, SVM is trained to perform the classification. The experiments conducted on the Extended Cohn-Kanade (CK+) Dataset show that the proposed DGVF is robust to capture and represent the facial appearance features. And our method also achieves a superior performance compared with the other state-of-the-art methods.
JunKai Chen, Zheru Chi, Hong Fu
MMSP3
2015 A new approach for pain event detection in video
abstract
A new approach for pain event detection in video is presented in this paper. Different from some previous works which focused on frame-based detection, we target in detecting pain events at video level. In this work, we explore the spatial information of video frames and dynamic textures of video sequences, and propose two different types of features. HOG of fiducial points (P-HOG) is employed to extract spatial features from video frames and HOG from Three Orthogonal Planes (HOG-TOP) is used to represent dynamic textures of video subsequences. After that, we apply max pooling to represent a video sequence as a global feature vector. Multiple Kernel Learning (MKL) is utilized to find an optimal fusion of the two types of features. And an SVM with multiple kernels is trained to perform the final classification. We conduct our experiments on the UNBC-McMaster Shoulder Pain dataset and achieve promising results, showing the effectiveness of our approach.
JunKai Chen, Zheru Chi, Hong Fu
ACII3
2015 Dynamic texture and geometry features for facial expression recognition in video
abstract
Facial expression recognition in video has attracted growing attention recently. In this paper, we propose to handle this problem with dynamic appearance and geometric features. We propose a new feature descriptor called HOG from Three Orthogonal Planes (HOG-TOP) to represent dynamic features. In addition, we introduce two types of geometry features to represent the facial rigid changes and non-rigid changes, respectively. Multiple Kernel Learning (MKL) is applied to find an optimal combination of two types of features. And finally a Support Vector Machine (SVM) with multiple kernels is trained for the facial expression classification. Extensive experiments conducted on the extended Cohn-Kanade dataset show that our method can achieve a competitive performance compared with the other state-of-the-art methods.
JunKai Chen, Zenghai Chen, Zheru Chi, Hong Fu
ICIP4
2015 Eye-Tracking Aided Digital System for Strabismus Diagnosis
abstract
Strabismus is a common ophthalmic disease with a relatively high prevalence (4%). It is one of the most common vision disorders in preschool children. If it is not timely diagnosed and well treated, strabismus would cause amblyopia, and even permanent vision loss. The diagnosis is thus essential. However, most of the diagnosis methods, such as cover testing, are conducted manually by an ophthalmologist. The examination cost is relatively high and the results are subjective. In this paper, we propose an eye tracking method for strabismus diagnosis. This method allows us to develop an objective and automatic strabismus diagnosis system that could significantly increase the examination efficiency and reduce the cost. Experimental results demonstrate the effectiveness of the proposed eye tracking method for strabismus diagnosis.
Zenghai Chen, Hong Fu, Zheru Chi
SMC2
2015 Learning to Detect Saliency with Deep Structure
abstract
Deep learning has shown great successes in solving various problems of computer vision. To the best of our knowledge, however, little existing work applies deep learning to saliency modeling. In this paper, a new saliency model based on convolutional neural network is proposed. The proposed model is able to produce a saliency map directly from an image's pixels. In the model, multi-level output values are adopted to simulate continuous values in a saliency map. Differing from most neural networks that use a relatively small number of output nodes, the output layer of our model has a large number of nodes. To make the training more efficient, an improved learning algorithm is adopted to train the model. Experimental results show that the proposed model succeeds in generating acceptable saliency maps after proper training.
Zenghai Chen, Zheru Chi, Hong Fu
SMC4
2014 Emotion Recognition in the Wild with Feature Fusion and Multiple Kernel Learning
abstract
This paper presents our proposed approach for the second Emotion Recognition in The Wild Challenge. We propose a new feature descriptor called Histogram of Oriented Gradients from Three Orthogonal Planes (HOG_TOP) to represent facial expressions. We also explore the properties of visual features and audio features, and adopt Multiple Kernel Learning (MKL) to find an optimal feature fusion. An SVM with multiple kernels is trained for the facial expression classification. Experimental results demonstrate that our method achieves a promising performance. The overall classification accuracy on the validation set and test set are 40.21% and 45.21%, respectively.
JunKai Chen, Zenghai Chen, Zheru Chi, Hong Fu
ICMI4
2014 A Hybrid Holistic/Semantic Approach for Scene Classification
abstract
There are two main strategies to tackle scene classification: holistic and semantic. The former characterizes a scene using its global features, while the latter represents a scene by modeling its internal object configuration. Holistic strategy is good at representing scenes with simple contents, but it does not represent well complex scenes that consist of multiple objects. By contrast, semantic strategy is advantageous at recognizing scenes with complex objects, but it does not work well for simple scenes. In this paper, we propose to integrate holistic and semantic strategies to cope with scene classification. In particular, we exploit a deep learning algorithm to learn features for scene representation in the holistic way. For the semantic strategy, we explore a semantic spatial pyramid to represent the spatial object configuration of scenes. The holistic and semantic strategies are integrated using a method proposed by us. Experimental results on a benchmark natural scene dataset demonstrate the effectiveness of our proposed hybrid approach for scene classification, by comparing to several state-of-the-art algorithms.
Zenghai Chen, Zheru Chi, Hong Fu
ICPR3
2013 Multi-instance multi-label image classification: A neural approach
Zenghai Chen, Zheru Chi, Hong Fu, David Dagan Feng
Neurocomputing3
2012 Salient object detection using content-sensitive hypergraph representation and partitioning
Zheru Chi, Hong Fu, David Dagan Feng
Pattern Recognit.3
2012 An Adaptive Recognition Model for Image Annotation
abstract
In this paper, an adaptive recognition model (ARM) is proposed for image annotation. The ARM consists of an adaptive classification network (CFN) and a nonlinear correlation network (CLN). The adaptive CFN aims to annotate an image with keywords, and the CLN is used to unveil the correlative information of keywords for annotation refinement. Image annotation is carried out by an ARM in two stages. In the first stage, the features extracted from regions of the input image are fed to a CFN to produce classification labels. In the second stage, the CLN uses keyword correlations learned from the training images to refine the classification result. The ARM works in a forward-propagating manner, resulting in high efficiency in image annotation. Furthermore, the computational time of an ARM is insensitive to the number of regions of the input image and the vocabulary size. In this paper, the effect of keyword correlation in image annotation is, comprehensively, investigated on a real image dataset and a synthetic image dataset. The exploitation of a controllable synthetic dataset helps to systematically study the function of keyword correlation and effectively analyze the performance of the ARM. Experimental results demonstrate the efficiency and effectiveness of the ARM.
Zenghai Chen, Hong Fu, Zheru Chi, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Part C2
2011 A Time Delay Neural Network model for simulating eye gaze data
abstract
Human eye movement modelling is a new, challenging and promising research topic in computer vision. Human eye movement modelling aims at simulating the scan path in which a human being views an image, a scene or a video. The successful modelling of human eye movements potentially benefits a wide range of applications such as image retrieval, image annotation, medical image diagnosis and human visual perception. This article presents a model based on a Time Delay Neural Network (TDNN) to simulate eye gaze data. First, 120 Hz eye gaze data are acquired by a non-intrusive table-mounted eye tracker. Our proposed model is to simulate the image reading process of a single subject. Seven features are then extracted based on the knowledge of the human oculomotor system and the image contents to train a TDNN. Finally, the trained TDNN combined with a saccade control mechanism is used to simulate the scan path of a human being viewing an image. The proposed model can generate 600 points of raw eye gaze data in a 5-second eye viewing window. Both subjective and objective methods are used to evaluate the model by comparing its behaviour and characteristics with the real eye gaze data collected from an eye tracker. Qualitative assessment shows that the subject can hardly tell the differences between the scan path from the model and that from a human being. By evaluating the coincident probability Cp and coincident significance Cs , quantitative assessment shows that the results from the TDNN model are reasonable and similar to human scan paths.
Hong Fu, Zheru Chi, David Dagan Feng
J. Exp. Theor. Artif. Intell.3
2010 Combined Retrieval Strategies for Images with and without Distinct Objects
Hong Fu, Zheru Chi, David Dagan Feng
ACIVS (1)1
2010 Salient-SIFT for Image Retrieval
Hong Fu, Zheru Chi, David Dagan Feng
ACIVS (1)2
2010 Content-based image retrieval using a combination of visual features and eye tracking data
abstract
Image retrieval technology has been developed for more than twenty years. However, the current image retrieval techniques cannot achieve a satisfactory recall and precision. To improve the effectiveness and efficiency of an image retrieval system, a novel content-based image retrieval method with a combination of image segmentation and eye tracking data is proposed in this paper. In the method, eye tracking data is collected by a non-intrusive table mounted eye tracker at a sampling rate of 120 Hz, and the corresponding fixation data is used to locate the human's Regions of Interest (hROIs) on the segmentation result from the JSEG algorithm. The hROIs are treated as important informative segments/objects and used in the image matching. In addition, the relative gaze duration of each hROI is used to weigh the similarity measure for image retrieval. The similarity measure proposed in this paper is based on a retrieval strategy emphasizing the most important regions. Experiments on 7346 Hemera color images annotated manually show that the retrieval results from our proposed approach compare favorably with conventional content-based image retrieval methods, especially when the important regions are difficult to be located based on visual features.
Hong Fu, Zheru Chi, David Dagan Feng
ETRA2
2010 Eye movement as an interaction mechanism for relevance feedback in a content-based image retrieval system
abstract
Relevance feedback (RF) mechanisms are widely adopted in Content-Based Image Retrieval (CBIR) systems to improve image retrieval performance. However, there exist some intrinsic problems: (1) the semantic gap between high-level concepts and low-level features and (2) the subjectivity of human perception of visual contents. The primary focus of this paper is to evaluate the possibility of inferring the relevance of images based on eye movement data. In total, 882 images from 101 categories are viewed by 10 subjects to test the usefulness of implicit RF, where the relevance of each image is known beforehand. A set of measures based on fixations are thoroughly evaluated which include fixation duration, fixation count, and the number of revisits. Finally, the paper proposes a decision tree to predict the user's input during the image searching tasks. The prediction precision of the decision tree is over 87%, which spreads light on a promising integration of natural eye movement into CBIR systems in the future.
Hong Fu, Zheru Chi, David Dagan Feng
ETRA2
2010 A neural network model with adaptive structure for image annotation
abstract
A neural network model with adaptive structure for image annotation is proposed in this paper. The adaptive structure enables the proposed model to utilize both global and regional visual features, as well as correlative information of annotated keywords for annotation. In order to achieve an approximate global optimum rather than a local optimum, both genetic algorithm and traditional back-propagation algorithm, are combined for model training. The neural network model is experimented on a synthetic image dataset with controllable parameters, which has not been used in previous image annotation experiments. Experimental results demonstrate the effectiveness of the proposed model.
Zenghai Chen, Hong Fu, Zheru Chi, David Dagan Feng
ICARCV2
2010 Refining a region based attention model using eye tracking data
abstract
Computational visual attention modeling is a topic of increasing importance in machine understanding of images. In this paper, we present an approach to refine a region based attention model with eye tracking data. This paper has three main contributions. (1) A concept of fixation mask is proposed to describe the region saliency of an image by weighting the segmented regions using importance measures obtained in the Human Visual System (HVS) or computational models. (2) A Genetic Algorithm (GA) scheme for refining a region based attention model is proposed. (3) An evaluation method is developed to measure the correlation between the result from the computational model and that from the HVS in terms of fixation mask.
Hong Fu, Zheru Chi, David Dagan Feng
ICIP2
2010 Adaptive Energy and Location Aware Routing in Wireless Sensor Network
Hong Fu, Xiaoming Wang 0001, Yingshu Li 0001
WASA1
2010 Recognition of attentive objects with a concept association network for image annotation
Hong Fu, Zheru Chi, David Dagan Feng
Pattern Recognit.1
2009 Eye movement data modeling using a genetic algorithm
abstract
We present a computational model of human eye movements based on a genetic algorithm (GA). The model can generate elemental raw eye movement data in a four-second eye viewing window with a 25 Hz sampling rate. Based on the physiology and psychology characters of human vision system, the fitness function of the GA model is constructed by taking into consideration of five factors including the saliency map, short time memory, saccades distribution, Region of Interest (ROI) map, and a retina model. Our model can produce the scan path of a subject viewing an image, not just several fixations points or artificial ROI's as in the other models. We have also developed both subjective and objective methods to evaluate the model by comparing its behavior with the real eye movement data collected from an eye tracker. Tested on 18 (9 times 2) images from both an obvious-object image group and a non-obvious-object image group, the subjective evaluations shows very close scores between the scan paths generated by the GA model and those real scan paths; for the objective evaluation, experimental results show that the distance between GA's scan paths and human scan paths of the same image has no significant difference by a probability of 78.9% on average.
Hong Fu, Zheru Chi, David Dagan Feng
IEEE Congress on Evolutionary Computation2
2009 Tree structures with attentive objects for image classification using a neural network
abstract
This paper presents an image classification method based on a neural network model dealing with tree structures of attentive objects. Apart from regions provided by image segmentation, attentive objects, which are extracted from a segmented image by an attention-driven image interpretation algorithm, are used to construct the tree structure to represent an image. Three combinations of tree structures are investigated, including ldquoimage + attentive-object + segmentsrdquo, ldquoimage + attentive-objectsrdquo, as well as ldquoimage + segmentsrdquo. Structure based neural networks are trained to classify the images by using the back propagation through structure (BPTS) algorithm. Experimental results show that the ldquoimage + attentive objectsrdquo structure is more favorable, comparing with both the other two structures proposed by us and a start-of-art tree structure reported in the literature, in terms of classification rate and computational time.
Hong Fu, Shuya Zhang, Zheru Chi, David Dagan Feng
IJCNN1
2009 An efficient algorithm for attention-driven image interpretation from segments
Hong Fu, Zheru Chi, David Dagan Feng
Pattern Recognit.1
2007 Pre-classification Module for an All-Season Image Retrieval System
abstract
From the study of attention-driven image interpretation and retrieval, we have found that an attention-driven strategy is able to extract important objects from an image and then focus the attentive objects while retrieving images. However, besides the images with distinct objects, there are images which do not show distinct objects. In this paper, the classification of "attentive" and "non-attentive" image is proposed to be a pre-process module in an all-season image retrieval system which can tackle both kinds of images. In this pre-classification module, an image is represented by an adaptive tree structure with each node carrying normalized features that characterize the object/region with visual contrasts and spatial information. Then a neural network is trained to classify an image as an "attentive" or "non-attentive" category by using the Back Propagation Through Structure (BPTS) algorithm. Experimental results indicate the reliability and feasibility of the pre-classification module, which encourages us to conduct further investigations on the all-season image retrieval system.
Hong Fu, Zheru Chi, David Dagan Feng, Weibao Zou, King Chuen Lo
IJCNN1
2006 Attention-driven image interpretation with application to image retrieval
Hong Fu, Zheru Chi, David Dagan Feng
Pattern Recognit.1
2004 Machine learning techniques for ontology-based leaf classification
abstract
Leaf classification, indexing as well as retrieval is an important part of a computerized plant identification system. In this paper, an integrated approach for an ontology-based leaf classification system is proposed, wherein machine learning techniques play a crucial role for the automatization of the system. For the leaf contour classification, a scaled CCD code system is proposed to categorize the basic shape and margin type of a leaf by using the similar taxonomy principle adopted by the botanists. Then a trained neural network is employed to recognize the detailed tooth patterns. The measurement on an unlobed leaf is also conducted automatically according to the method used in botany. For the leaf vein recognition, the vein texture is extracted by employing an efficient combined thresholding and neural network approach so as to obtain more vein details of a leaf. Compared with the past studies, the proposed method integrates low-level features of an image and the specific knowledge in the domain (ontology) of botany, and therefore provides a more practical system for users to comprehend and handle. Primary experiments have shown promising results and proven the feasibility of the proposed system.
Hong Fu, Zheru Chi, David Dagan Feng, Jiatao Song
ICARCV1
2004 Extraction of face image edges with application to expression analysis
abstract
This paper presents a new method for the extraction of face image edges and a scheme for facial expression analysis based on the binary edge images (BEIs). Making use of the multi-resolution property of wavelet transform, the edge detection method includes two binarization steps and a noise removing step. Our method provides a robust solution to face image edge extraction, in particular, if changes of lighting conditions are encountered. Experimental results show that different components in a BEI are well separated and pixels of the same face component are connected well. Preliminary experimental results on facial expression analysis on the subsets of AR and Yale face databases show that a satisfactory recognition rate can be achieved when our facial analysis scheme is used for recognizing four action units.
Jiatao Song, Zheru Chi, Jilin Liu, Hong Fu
ICARCV4
2004 Automatic generation of pen-and-ink drawings from photos
abstract
In this paper, a new algorithm for automatic generating pen-and-ink drawings from photos is presented. The basic processing step of our system is the edge extraction based on the multiresolution property of wavelet transform. In order to enhance the variations in texture and tone, the proposed algorithm includes two binarization steps and a noise removing step. Experimental results show that the generated pen-and-ink drawings are of high quality, which accurately illustrate the scenic contents of the original photos.
Jiatao Song, Zheru Chi, Jilin Liu, Hong Fu
ICIP4