Hongying Meng

dblp:19/6594 · DBLP profile ↗
← Back
64ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0002-8836-1382ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 35 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 2 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Edge priors guided deep unrolling network for single image super-resolution
Heping Song, Hongjie Jia, Xiangjun Shen, Jianping Gou, Yuping Lai, Hongying Meng
Expert Syst. Appl.8
2026 Robust deep dictionary learning via self-expression neighbor atom enhancement
Heping Song, Yusen Qian, Sumet Mehta, Jianping Gou, Hongying Meng, Xiangjun Shen
Expert Syst. Appl.5
2026 Mitigating model coupling in semi-supervised segmentation via deep non-consistent mean teacher and fully collaborative learning
Chongdan Min, Tao Lei 0003, Xingwu Wang, Hongying Meng, Asoke K. Nandi
Neurocomputing5
2026 Vendor-Independent Design Space Exploration and Resource Optimization Framework for 3-D Networks-on-Chip Using Hypergraph-Genetic Algorithm Integration
abstract
This paper presents a novel methodology for design space exploration and resource optimisation of three-dimensional Networks-on-Chip (3D NoC) architectures using hypergraph modelling and genetic algorithms. The proposed approach combines mathematical rigour with evolutionary search capabilities to efficiently explore the vast design space of 3D NoC configurations, providing a vendor-independent solution for NoC architects. The key contribution is the development of Performance-Cost-Ratio (PCR) functions that enable quantitative evaluation of different topologies and routing algorithms, extended to include power and thermal considerations with dynamic adaptation mechanisms for runtime traffic variations. Validation through four compute-intensive use cases demonstrates significant improvements, with optimised architectures achieving up to 33% reduction in latency, 40% increase in throughput, and 30% reduction in power consumption compared to baseline implementations. Validation against published silicon implementations shows 92-96% correlation accuracy, confirming the framework’s practical applicability for developing efficient and scalable NoC solutions as processor designs advance towards kilo-core scales and beyond.
Ahmed Al-Alousi, Maozhen Li 0001, Hongying Meng
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 Representation Sampling and Hybrid Transformer Network for Image Compressed Sensing
Heping Song, Jingyao Gong, Hongjie Jia, Xiangjun Shen, Jianping Gou, Hongying Meng, Le Wang 0003
IEEE Trans. Circuits Syst. Video Technol.6
2025 LAC-Net: Feature-Corrected Location-Aware Network for Medical Image Segmentation
Youtao Jiang, Yi Wang 0069, Shaoqing Liu, Xiaogang Du, Hongying Meng, Tao Lei 0003
PRCV (13)5
2025 Semi-supervised Medical Image Segmentation Based on Uncertainty-Driven Dynamic Correction and Multi-scale Consistency Learning
Shaoqing Liu, Wenbiao Song, Xiaogang Du, Hongying Meng, Tao Lei 0003
PRCV (13)5
2025 A Lightweight Object Counting Network Based on Density Map Knowledge Distillation
abstract
Object counting aims to count the accurate number of object instances in images, and its operation efficiency is essential. However, most current CNN-based methods rely on complex network architectures, which results in them consuming a significant amount of memory, time, and other resources at runtime. This seriously limits their deployment in practical application scenarios, such as public safety and agriculture planting. Therefore, we propose a lightweight object counting method named EdgeCount to effectively balance inference speed and object counting accuracy. Specifically, we construct a network composed of a student model (EdgeCount) and a teacher model (EdgeCount-T) with the same encoder-decoder structure based on density map knowledge distillation (DMKD), allowing the EdgeCount to learn object density distribution from the EdgeCount-T. After that, we introduce spatial and channel reconstruction convolution (SCConv), composed of a spatial reconstruction unit (SRU) and a channel reconstruction unit (CRU), to decrease spatial and channel redundancy with lower computational costs. Moreover, a low parameter weighted multi-scale feature fusion module (LWMFFM) is designed to further improve the countering ability through segmenting minor structural discrepacies among multi-scale features. Extensive experiments conducted on challenging remote sensing and dense crowd object counting datasets demonstrate the effectiveness and superiority of our method. In particular, under the four NVIDIA Jetson devices, EdgeCount can accurately counter objects with only 0.12M parameters and 19.87M floating-point operations per second (FLOPs) in the size of 128, which achieves the lowest latency and fastest FPS compared with other state-of-the-art object counters.
Zhilong Shen, Guoquan Li 0001, Ruiyang Xia, Hongying Meng, Zhengwen Huang
IEEE Trans. Circuits Syst. Video Technol.4
2024 Multi-Cross Sampling and Frequency-Division Reconstruction for Image Compressed Sensing
abstract
Deep Compressed Sensing (DCS) has attracted considerable interest due to its superior quality and speed compared to traditional CS algorithms. However, current approaches employ simplistic convolutional downsampling to acquire measurements, making it difficult to retain high-level features of the original signal for better image reconstruction. Furthermore, these approaches often overlook the presence of both high- and low-frequency information within the network, despite their critical role in achieving high-quality reconstruction. To address these challenges, we propose a novel Multi-Cross Sampling and Frequency Division Network (MCFD-Net) for image CS. The Dynamic Multi-Cross Sampling (DMCS) module, a sampling network of MCFD-Net, incorporates pyramid cross convolution and dual-branch sampling with multi-level pooling. Additionally, it introduces an attention mechanism between perception blocks to enhance adaptive learning effects. In the second deep reconstruction stage, we design a Frequency Division Reconstruction Module (FDRM). This module employs a discrete wavelet transform to extract high- and low-frequency information from images. It then applies multi-scale convolution and self-similarity attention compensation separately to both types of information before merging the output reconstruction results. The MCFD-Net integrates the DMCS and FDRM to construct an end-to-end learning network. Extensive CS experiments conducted on multiple benchmark datasets demonstrate that our MCFD-Net outperforms state-of-the-art approaches, while also exhibiting superior noise robustness.
Heping Song, Jingyao Gong, Hongying Meng, Yuping Lai
AAAI3
2024 PracticalDG: Perturbation Distillation on Vision-Language Models for Hybrid Domain Generalization
abstract
Domain Generalization (DG) aims to resolve distribution shifts between source and target domains, and current DG methods are default to the setting that data from source and target domains share identical categories. Nevertheless, there exists unseen classes from target domains in practical scenarios. To address this issue, Open Set Domain Generalization (OSDG) has emerged and several methods have been exclusively proposed. However, most existing methods adopt complex architectures with slight improvement compared with DG methods. Recently, vision-language models (VLMs) have been introduced in DG following the fine-tuning paradigm, but consume huge training overhead with large vision models. Therefore, in this paper, we innovate to transfer knowledge from VLMs to lightweight vision models and improve the robustness by introducing Perturbation Distillation (PD) from three perspectives, including Score, Class and Instance (SCI), named SCI-PD. Moreover, previous methods are oriented by the benchmarks with identical and fixed splits, ignoring the divergence between source domains. These methods are revealed to suffer from sharp performance decay with our proposed new benchmark Hybrid Domain Generalization (HDG) and a novel metric H2-CV, which construct various splits to comprehensively assess the robustness of algorithms. Extensive experiments demonstrate that our method outperforms state-of-the-art algorithms on multiple datasets, especially improving the robustness when confronting data scarcity.
Zining Chen, Weiqiu Wang, Zhicheng Zhao 0001, Aidong Men, Hongying Meng
CVPR6
2024 Neonatal White Matter Damage Analysis Using DTI Super-Resolution and Multi-Modality Image Registration
abstract
Punctate White Matter Damage (PWMD) is a common neonatal brain disease, which can easily cause neurological disorder and strongly affect life quality in terms of neuromotor and cognitive performance. Especially, at the neonatal stage, the best cure time can be easily missed because PWMD is not conducive to the diagnosis based on current existing methods. The lesion of PWMD is relatively straightforward on T1-weighted Magnetic Resonance Imaging (T1 MRI), showing semi-oval, cluster or linear high signals. Diffusion Tensor Magnetic Resonance Image (DT-MRI, referred to as DTI) is a noninvasive technique that can be used to study brain microstructures in vivo, and provide information on movement and cognition-related nerve fiber tracts. Therefore, a new method was proposed to use T1 MRI combined with DTI for better neonatal PWMD analysis based on DTI super-resolution and multi-modality image registration. First, after preprocessing, neonatal DTI super-resolution was performed with the three times B-spline interpolation algorithm based on the Log-Euclidean space to improve DTIs’ resolution to fit the T1 MRIs and facilitate nerve fiber tractography. Second, the symmetric diffeomorphic registration algorithm and inverse b0 image were selected for multi-modality image registration of DTI and T1 MRI. Finally, the 3D lesion models were combined with fiber tractography results to analyze and predict the degree of PWMD lesions affecting fiber tracts. Extensive experiments demonstrated the effectiveness and super performance of our proposed method. This streamlined technique can play an essential auxiliary role in diagnosing and treating neonatal PWMD.
Yi Wang 0069, Hongying Meng
Int. J. Neural Syst.8
2024 SSRL: Self-Supervised Spatial-Temporal Representation Learning for 3D Action Recognition
abstract
For 3D action recognition, the main challenge is to extract long-range semantic information in both temporal and spatial dimensions. In this paper, in order to better excavate long-range semantic information from large number of unlabelled skeleton sequences, we propose Self-supervised Spatial-temporal Representation Learning (SSRL), a contrastive learning framework to learn skeleton representation. SSRL consists of two novel inference tasks that enable the network to learn global semantic information in the temporal and spatial dimensions, respectively. The temporal inference task learns the temporal persistence of human actions through temporally incomplete skeleton sequences. And the spatial inference task learns the spatially coordinated nature of human action through spatially partially skeleton sequence. We design two transformation modules to efficiently realize these two tasks while fitting the encoder network. To avoid the difficulty of constructing and maintaining high-quality negative samples, our proposed framework learns by maintaining consistency among positive samples without the need of any negative sample. Experiments demonstrate that our proposed method can achieve better results in comparison with state-of-the-art methods under a variety of evaluation protocols on NTU RGB+D 60, PKU-MMD and NTU RGB+D 120 datasets.
Zhihao Jin, Qicong Wang, Yehu Shen, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.5
2024 Bayesian Estimation of Inverted Beta Mixture Models With Extended Stochastic Variational Inference for Positive Vector Classification
abstract
The finite inverted beta mixture model (IBMM) has been proven to be efficient in modeling positive vectors. Under the traditional variational inference framework, the critical challenge in Bayesian estimation of the IBMM is that the computational cost of performing inference with large datasets is prohibitively expensive, which often limits the use of Bayesian approaches to small datasets. An efficient alternative provided by the recently proposed stochastic variational inference (SVI) framework allows for efficient inference on large datasets. Nevertheless, when using the SVI framework to address the non-Gaussian statistical models, the evidence lower bound (ELBO) cannot be explicitly calculated due to the intractable moment computation. Therefore, the algorithm under the SVI framework cannot directly use stochastic optimization to optimize the ELBO, and an analytically tractable solution cannot be derived. To address this problem, we propose an extended version of the SVI framework with more flexibility, namely, the extended SVI (ESVI) framework. This framework can be used in many non-Gaussian statistical models. First, some approximation strategies are applied to further lower the ELBO to avoid intractable moment calculations. Then, stochastic optimization with noisy natural gradients is used to optimize the lower bound. The excellent performance and effectiveness of the proposed method are verified in real data evaluation.
Yuping Lai, Wenbo Guan, Lijuan Luo, Yanhui Guo 0001, Heping Song, Hongying Meng
IEEE Trans. Neural Networks Learn. Syst.6
2024 A Self-Supervised Network-Based Smoke Removal and Depth Estimation for Monocular Endoscopic Videos
abstract
In minimally invasive surgery videos, label-free monocular laparoscopic depth estimation is challenging due to smoke. For this reason, we propose a self-supervised collaborative network-based depth estimation method with smoke-removal for monocular endoscopic video, which is decomposed into two steps of smoke-removal and depth estimation. In the first step, we develop a de-endoscopic smoke for cyclic GAN (DS-cGAN) to mitigate the smoke components at different concentrations. The designed generator network comprises sharpened guide encoding module (SGEM), residual dense bottleneck module (RDBM) and refined upsampling convolution module (RUCM), which restores more detailed organ edges and tissue structures. In the second step, high resolution residual U-Net (HRR-UNet) consisting of a DepthNet and two PoseNets is designed to improve the depth estimation accuracy, and adjacent frames are used for camera self-motion estimation. In particular, the proposed method requires neither manual labeling nor patient computed tomography scans during the training and inference phases. Experimental studies on the laparoscopic data set of the Hamlyn Centre show that our method can effectively achieve accurate depth information after net smoking in real surgical scenes while preserving the blood vessels, contours and textures of the surgical site. The experimental results demonstrate that the proposed method outperforms existing state-of-the-art methods in effectiveness and achieves a frame rate of 94.45fps in real time, making it a promising clinical application.
Xinbo Gao 0001, Hongying Meng, Xixi Nie
IEEE Trans. Vis. Comput. Graph.3
2023 Bi-path Combination YOLO for Real-time Few-shot Object Detection
Ruiyang Xia, Guoquan Li 0001, Zhengwen Huang, Hongying Meng
Pattern Recognit. Lett.4
2023 Self-Supervised 3D Behavior Representation Learning Based on Homotopic Hyperbolic Embedding
abstract
Behavior sequences are generated by a series of spatio-temporal interactions and have a high-dimensional nonlinear manifold structure. Therefore, it is difficult to learn 3D behavior representations without relying on supervised signals. To this end, self-supervised learning methods can be used to explore the rich information contained in the data itself. Context-context contrastive self-supervised methods construct the manifold embedded in Euclidean space by learning the distance relationship between data, and find the geometric distribution of data. However, traditional Euclidean space is difficult to express context joint features. In order to obtain an effective global representation from the relationship between data under unlabeled conditions, this paper adopts contrastive learning to compare global feature, and proposes a self-supervised learning method based on hyperbolic embedding to mine the nonlinear relationship of behavior trajectories. This method adopts the framework of discarding negative samples, which overcomes the shortcomings of the paradigm based on positive and negative samples that pull similar data away in the feature space. Meanwhile, the output of the network is embedded in a hyperbolic space, and a multi-layer perceptron is added to convert the entire module into a homotopic mapping by using the geometric properties of operations in the hyperbolic space, so as to obtain homotopy invariant knowledge. The proposed method combines the geometric properties of hyperbolic manifolds and the equivariance of homotopy groups to promote better supervised signals for the network, which improves the performance of unsupervised learning.
Jinghong Chen, Zhihao Jin, Qicong Wang, Hongying Meng
IEEE Trans. Image Process.4
2023 StrongSORT: Make DeepSORT Great Again
abstract
Recently, Multi-Object Tracking (MOT) has attracted rising attention, and accordingly, remarkable progresses have been achieved. However, the existing methods tend to use various basic models (e.g, detector and embedding model), and different training or inference tricks, etc. As a result, the construction of a good baseline for a fair comparison is essential. In this paper, a classic tracker, i.e., DeepSORT, is first revisited, and then is significantly improved from multiple perspectives such as object detection, feature embedding, and trajectory association. The proposed tracker, named StrongSORT, contributes a strong and fair baseline for the MOT community. Moreover, two lightweight and plug-and-play algorithms are proposed to address two inherent “missing” problems of MOT: missing association and missing detection. Specifically, unlike most methods, which associate short tracklets into complete trajectories at high computation complexity, we propose an appearance-free link model (AFLink) to perform global association without appearance information, and achieve a good balance between speed and accuracy. Furthermore, we propose a Gaussian-smoothed interpolation (GSI) based on Gaussian process regression to relieve the missing detection. AFLink and GSI can be easily plugged into various trackers with a negligible extra computational cost (1.7 ms and 7.1 ms per image, respectively, on MOT17). Finally, by fusing StrongSORT with AFLink and GSI, the final tracker (StrongSORT++) achieves state-of-the-art results on multiple public benchmarks, i.e., MOT17, MOT20, DanceTrack and KITTI. Codes are available athttps://github.com/dyhBUPT/StrongSORTandhttps://github.com/open-mmlab/mmtracking.
Yunhao Du, Zhicheng Zhao 0001, Yang Song 0036, Yanyun Zhao, Hongying Meng
IEEE Trans. Multim.7
2023 LTReID: Factorizable Feature Generation With Independent Components for Long-Tailed Person Re-Identification
abstract
With the rapid increase of large-scale and real-world person datasets, it is crucial to address the problem of long-tailed data distributions,i.e., head classes have large number of images while tail classes occupy extremely few samples. We observe that the imbalanced data distribution is likely to distort the overall feature space and impair the generalization capability of trained models. Nevertheless, this long-tailed problem has been rarely investigated in previous person Re-Identification (ReID) works. In this paper, we propose a novelLong-Tailed Re-Identification(LTReID) framework to simultaneously alleviate class-imbalance and hard-imbalance problems. Specifically, each real feature is decomposed into multiple independent components with two decorrelation losses. Then these components are randomly aggregated to generate more fake features for tail classes than head ones, resulting in the class-balance between head and tail classes. For the hard-balance between easy and hard samples, we utilize adversarial learning to generate more hard features than easy ones. The proposed framework can be trained in an end-to-end manner and avoids increasing the space and time complexity of inference models. Moreover, comprehensive experiments are conducted on the four ReID datasets so as to validate the effectiveness of the overall framework and the advantage of each module. Our results show that when trained with either balanced or imbalanced datasets, the LTReID achieves superior performance over the state-of-the-art methods.
Pingyu Wang, Zhicheng Zhao 0001, Hongying Meng
IEEE Trans. Multim.4
2022 Sparse signal reconstruction via generalized two-stage thresholding
Heping Song, Zehong Ai, Yuping Lai, Hongying Meng, Qirong Mao
Sci. China Inf. Sci.4
2022 Medical image segmentation using deep learning: A survey
abstract
Abstract Deep learning has been widely used for medical image segmentation and a large number of papers has been presented recording the success of deep learning in the field. A comprehensive thematic survey on medical image segmentation using deep learning techniques is presented. This paper makes two original contributions. Firstly, compared to traditional surveys that directly divide literatures of deep learning on medical image segmentation into many groups and introduce literatures in detail for each group, we classify currently popular literatures according to a multi‐level structure from coarse to fine. Secondly, this paper focuses on supervised and weakly supervised learning approaches, without including unsupervised approaches since they have been introduced in many old surveys and they are not popular currently. For supervised learning approaches, we analyse literatures in three aspects: the selection of backbone networks, the design of network blocks, and the improvement of loss functions. For weakly supervised learning approaches, we investigate literature according to data augmentation, transfer learning, and interactive segmentation, separately. Compared to existing surveys, this survey classifies the literatures very differently from before and is more convenient for readers to understand the relevant rationale and will guide them to think of appropriate improvements in medical image segmentation based on deep learning approaches.
Risheng Wang, Tao Lei 0003, Ruixia Cui, Hongying Meng, Asoke K. Nandi
IET Image Process.5
2022 Extended variational inference for Dirichlet process mixture of Beta-Liouville distributions for proportional data modeling
abstract
Bayesian estimation of parameters in the Dirichlet mixture process of the Beta-Liouville distribution (i.e., the infinite Beta-Liouville mixture model) has recently gained considerable attention due to its modeling capability for proportional data. However, applying the conventional variational inference (VI) framework cannot derive an analytically tractable solution since the variational objective function cannot be explicitly calculated. In this paper, we adopt the recently proposed extended VI framework to derive the closed-form solution by further lower bounding the original variational objective function in the VI framework. This method is capable of simultaneously determining the model's complexity and estimating the model's parameters. Moreover, due to the nature of Bayesian nonparametric approaches, it can also avoid the problems of underfitting and overfitting. Extensive experiments were conducted on both synthetic and real data, generated from two real-world challenging applications, namely, object detection and text categorization, and its superior performance and effectiveness of the proposed method have been demonstrated.
Yuping Lai, Wenbo Guan, Lijuan Luo, Qiang Ruan, Yuan Ping 0003, Heping Song, Hongying Meng
Int. J. Intell. Syst.7
2022 Unsupervised visual feature learning based on similarity guidance
Zhihao Jin, Qicong Wang, Wenming Yang, Qingmin Liao, Hongying Meng
Neurocomputing6
2022 An end-to-end heterogeneous network for graph similarity learning
Yan Huang 0030, Qicong Wang, Hongying Meng
Neurocomputing5
2022 iCGPN: Interaction-centric graph parsing network for human-object interaction detection
Zhicheng Zhao 0001, Hongying Meng
Neurocomputing5
2022 What-Where-When Attention Network for video-based person re-identification
Ping Chen 0004, Tao Lei 0003, Yangxu Wu, Hongying Meng
Neurocomputing5
2022 Self-Supervised Representation Learning for Videos by Segmenting via Sampling Rate Order Prediction
abstract
Self-supervised representation learning for videos has been very attractive recently because these methods exploit the information inherently obtained from the video itself instead of annotated labels that is quite time-consuming. However, existing methods ignore the importance of global observation while performing spatio-temporal transformation perception, which highly limits the expression capabilities of the video representation. This paper proposes a novel pretext task that combines the temporal information perception of the video with the motion amplitude perception of moving objects to learn the spatio-temporal representation of the video. Specifically, given a video clip containing several video segments, each video segment is sampled by different sampling rates and the order of video segments is disrupted. Then, the network is used to regress the sampling rate of each video segment and classify the order of input video segments. In the pre-training stage, the network can learn rich spatio-temporal semantic information where content-related contrastive learning is introduced to make the learned video representation more discriminative. To alleviate the appearance dependency caused by contrastive learning, we design a novel and robust vector similarity measurement approach, which can take feature alignment into consideration. Moreover, a view synthesis framework is proposed to further improve the performance of contrastive learning by automatically generating reasonable transformed views. We conduct benchmark experiments with several 3D backbone networks on two datasets. The results show that our proposed method outperforms the existing state-of-the-art methods across the three backbones on two downstream tasks of human action recognition and video retrieval.
Yan Huang 0030, Qicong Wang, Wenming Yang, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.5
2022 Attentive Feature Augmentation for Long-Tailed Visual Recognition
abstract
Deep neural networks have achieved great success on many visual recognition tasks. However, training data with a long-tailed distribution dramatically degenerates the performance of recognition models. In order to relieve this imbalance problem, an effective Long-Tailed Visual Recognition (LTVR) framework is proposed based on learned balance and robust features under long-tailed distribution circumstances. In this framework, a plug-and-play Attentive Feature Augmentation (AFA) module is designed to mine class-related and variation-related features of original samples via a novel hierarchical channel attention mechanism. Then, those features are aggregated to synthesize fake features to cope with the imbalance of the original dataset. Moreover, a Lay-Back Learning Schedule (LBLS) is developed to ensure a good initialization of feature embedding. Extensive experiments are conducted with a two-stage training method to verify the effectiveness of the proposed framework on both feature learning and classifier rebalancing in the long-tailed image recognition task. Experimental results show that, when trained with imbalanced datasets, the proposed framework achieves superior performance over the state-of-the-art methods.
Weiqiu Wang, Zhicheng Zhao 0001, Pingyu Wang, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.5
2022 CBASH: Combined Backbone and Advanced Selection Heads With Object Semantic Proposals for Weakly Supervised Object Detection
abstract
Most recent object detection methods have achieved growing performance on public datasets. However, enormous efforts are needed for these methods due to the extensive annotations of ground-truth boxes. Weakly Supervised Object Detection (WSOD) methods hence have been proposed to solve this problem as only image-level annotations are required and then output bounding boxes related to the objects. In order to further elevate the weakly supervised detection methods on the extraction of reasonable features, the training of potential positive proposals, and the generation of proposals before training, we propose a new Combined Backbone and Advanced Selection Heads (CBASH) method with the proposals generated from the object semantic information. Specifically, Combined Backbone will make the unobvious object features more noticeable, Advanced Selection Heads promote more potential positive proposals to get training, and the generated object semantic proposals elevate the quality and quantity of positive proposals. The proposed method is evaluated on the challenging PASCAL VOC 2007 and 2012 benchmark datasets. Experimental results show that our proposed method can achieve improved performance on both VOC 2007 and VOC 2012 datasets and outperforms the existing state-of-the-art methods.
Ruiyang Xia, Guoquan Li 0001, Zhengwen Huang, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.4
2022 Fuzzy STUDENT'S T-Distribution Model Based on Richer Spatial Combination
abstract
Fuzzy c-means (FCM) algorithms with spatial information have been widely applied in the field of image segmentation. However, most of them suffer from two challenges. One is that the introduction of fixed or adaptive single neighboring information with narrow receptive field limits contextual constraints leading to clutter segmentations. The other is that the incorporation of superpixels with wide receptive field enlarges spatial coherency leading to block effects. To address these challenges, we propose fuzzy STUDENT’S t-distribution model based on richer spatial combination (FRSC) for image segmentation. In this article, we make two significant contributions. The first is that both the narrow and wide receptive fields are integrated into the objective function of FRSC, which is convenient to mine image features and distinguish local difference. The second is that the rich spatial combination under STUDENT’S t-distribution ensures that spatial information is introduced into the updated parameters of FRSC, which is helpful in finding a balance between the noise-immunity and detail-preservation. Experimental results on synthetic and publicly available images further demonstrate that the proposed FRSC addresses successfully the limitations of FCM algorithms with spatial information, and provides better segmentation results than state-of-the-art clustering algorithms.
Tao Lei 0003, Xiaohong Jia 0002, Dinghua Xue, Qi Wang 0009, Hongying Meng, Asoke K. Nandi
IEEE Trans. Fuzzy Syst.5
2021 Lightweight Non-Local Network for Image Super-Resolution
abstract
The popular deep convolutional networks used for image super-resolution (SR) reconstruction often increase the network depth and employ attention mechanism to improve image reconstruction effect. However, these networks suffer from two problems. The first is the deeper network easily causes higher computational cost and more GPU memory usage. The second is traditional attention mechanism often misses the spatial information of images leading the loss of image detail information. To address these issues, we propose a lightweight non-local network (LNLN) for image super resolution in this paper. The proposed network makes two contributions. First, we use non-local module instead of normal attention module to obtain larger receptive field and extract more comprehensive feature information, which is helpful for improving image SR reconstruction results. Secondly, we use the depthwise separable convolution (DSC) instead of the vanilla convolution to reconstruct the residual block, which greatly reduces the number of parameters and computational cost. The proposed LNLN and comparative networks are evaluated on five commonly public datasets, and experiments demonstrate that the proposed LNLN is superior to state-of-the-art networks in terms of reconstruction performance, the number of parameters and storage space.
Risheng Wang, Tao Lei 0003, Wenzheng Zhou, Qi Wang 0009, Hongying Meng, Asoke K. Nandi
ICASSP5
2021 HNSF Log-Demons: Diffeomorphic demons registration using hierarchical neighbourhood spectral features
abstract
Abstract Many biomedical applications require accurate non‐rigid image registration that can cope with complex deformations. However, popular diffeomorphic Demons registration algorithms suffer from difficulties for complex and serious distortions since they only use image greyscale and gradient information. To address these difficulties, a new diffeomorphic Demons registration algorithm is proposed using hierarchical neighbourhood spectral features namely HNSF Log‐Demons in this paper. In view of three important properties of hierarchical neighbourhood spectral features based on line graph such as rotation invariance, invariance of linear changes of brightness, and robustness to noise, the hierarchical neighbourhood spectral features of a reference image and a moving image is first extracted and these novel spectral features are incorporated into the energy function of the diffeomorphic registration framework to improve the capability of capturing complex distortions. Secondly, the Nystrm approximation based on random singular value decomposition is employed to effectively enhance the computational efficiency of HNSF Log‐Demons. Finally, the hybrid multi‐resolution strategy based on wavelet decomposition in the registration process is utilised to further improve the registration accuracy and efficiency. Experimental results show that the proposed HNSF Log‐Demons not only effectively ensures the generation of smooth and reversible deformation field, but also achieves better performance than state‐of‐the‐art algorithms.
Xiaogang Du, Dongxin Gu, Tao Lei 0003, Xuejun Zhang 0004, Hongying Meng
IET Image Process.6
2021 Triplet interactive attention network for cross-modality person re-identification
Ping Chen 0004, Tao Lei 0003, Hongying Meng
Pattern Recognit. Lett.4
2021 Holoscopic 3D Microgesture Recognition by Deep Neural Network Model Based on Viewpoint Images and Decision Fusion
abstract
Finger microgestures have been widely used in human computer interaction (HCI), particularly for interactive applications, such as virtual reality (VR) and augmented reality (AR) technologies, to provide immersive experience. However, traditional 2D image-based microgesture recognition suffers from low accuracy due to the limitations of 2D imaging sensors, which have no depth information. In this article, we proposed an innovative 3D microgesture recognition system based on a holoscopic 3D imaging sensor. Due to the lack of holoscopic 3D datasets, a comprehensive holoscopic 3D microgesture (HoMG) database is created and used to develop a robust 3D microgesture recognition method. Then, a fast algorithm is proposed to extract multiviewpoint images from one holoscopic image. Furthermore, we applied a CNN model with an attention-based residual block to each viewpoint image to improve the algorithm performance. Finally, bagging classification tree decision-level fusion is applied to combine the predictions. The experimental results demonstrate that the proposed method outperforms state-of-the-art methods and delivers a better accuracy than existing methods.
Yi Liu 0061, Min Peng 0004, Mohammad Rafiq Swash, Tong Chen 0008, Hongying Meng
IEEE Trans. Hum. Mach. Syst.6
2020 EMOPAIN Challenge 2020: Multimodal Pain Evaluation from Facial and Bodily Expressions
abstract
The EmoPain 2020 Challenge is the first international competition aimed at creating a uniform platform for the comparison of multi-modal machine learning and multimedia processing methods of chronic pain assessment from human expressive behaviour, and also the identification of pain-related behaviours. The objective of the challenge is to promote research in the development of assistive technologies that help improve the quality of life for people with chronic pain via real-time monitoring and feedback to help manage their condition and remain physically active. The challenge also aims to encourage the use of the relatively underutilised, albeit vital bodily expression signals for automatic pain and pain-related emotion recognition. This paper presents a description of the challenge, competition guidelines, bench-marking dataset, and the baseline systems' architecture and performance on the Challenge's three sub-tasks: pain estimation from facial expressions, pain recognition from multimodal movement, and protective movement behaviour detection.
Joy Egede, Siyang Song, Temitayo A. Olugbade, Amanda C. de C. Williams, Hongying Meng, M. S. Hane Aung, Nicholas D. Lane, Michel F. Valstar, Nadia Bianchi-Berthouze
FG6
2020 Lightweight V-Net for Liver Segmentation
abstract
The V-Net based 3D fully convolutional neural networks have been widely used in liver volumetric data segmentation. However, due to the large number of parameters of these networks, 3D FCNs suffer from high computational cost and GPU memory usage. To address these issues, we design a lightweight V-Net (LV-Net) for liver segmentation in this paper. The proposed network makes two contributions. The first is that we design an inverted residual bottleneck block (IRB block) and a 3D average pooling block and apply them to the proposed LV-Net. Compared with vanilla convolution, depth-wise convolution and point-wise convolution employed by the IRB block can not only reduce the number of parameters significantly, but also extract features sufficiently well by decoupling cross-channel corrections and spatial correlations. The second is that the LV-Net employs 3D deep supervision to improve the final loss function in training phase, which makes the proposed LV-Net acquire a more powerful discrimination capability between liver areas and non-liver areas. The proposed LV-Net is evaluated on public LiTS dataset, and experiments demonstrate that the proposed LV-Net is superior to popular 2D and 3D networks in terms of segmentation performance, parameter quantity and computational cost.
Tao Lei 0003, Wenzheng Zhou, Risheng Wang, Hongying Meng, Asoke K. Nandi
ICASSP5
2020 Discovering influential factors in variational autoencoders
Shiqi Liu 0001, Qian Zhao 0002, Xiangyong Cao, Huibin Li 0001, Deyu Meng, Hongying Meng, Sheng Liu 0033
Pattern Recognit.7
2020 Automatic Fuzzy Clustering Framework for Image Segmentation
abstract
Clustering algorithms by minimizing an objective function share a clear drawback of having to set the number of clusters manually. Although density peak clustering is able to find the number of clusters, it suffers from memory overflow when it is used for image segmentation because a moderate-size image usually includes a large number of pixels leading to a huge similarity matrix. To address this issue, here we proposed an automatic fuzzy clustering framework (AFCF) for image segmentation. The proposed framework has threefold contributions. First, the idea of superpixel is used for the density peak (DP) algorithm, which efficiently reduces the size of the similarity matrix and thus improves the computational efficiency of the DP algorithm. Second, we employ a density balance algorithm to obtain a robust decision-graph that helps the DP algorithm achieve fully automatic clustering. Finally, a fuzzy c-means clustering based on prior entropy is used in the framework to improve image segmentation results. Because the spatial neighboring information of both the pixels and membership are considered, the final segmentation result is improved effectively. Experiments show that the proposed framework not only achieves automatic image segmentation, but also provides better segmentation results than state-of-the-art algorithms.
Tao Lei 0003, Xiaohong Jia 0002, Xuande Zhang, Hongying Meng, Asoke K. Nandi
IEEE Trans. Fuzzy Syst.5
2019 End-to-end Change Detection Using a Symmetric Fully Convolutional Network for Landslide Mapping
abstract
In this paper, we propose a novel approach based on a symmetric fully convolutional network within pyramid pooling (FCN-PP) for landslide mapping (LM). The proposed approach has three advantages. Firstly, this approach is automatic and insensitive to noise because multivariate morphological reconstruction (MMR) is used for image preprocessing. Secondly, it is able to take into account features from multiple convolutional layers and explore efficiently the context of images, which leads to a good tradeoff between wider receptive field and the use of context. Finally, the selected pyramid pooling module addresses the drawback of single-scale pooling employed by convolutional neural network (CNN), fully convolutional network (FCN), U-Net, etc. Experimental results show that the proposed FCN-PP is effective for LM, and it outperforms state-of-the-art approaches in terms of four metrics, Precision, Recall, F -score, and Accuracy.
Tao Lei 0003, Qi Zhang 0091, Dinghua Xue, Tao Chen 0004, Hongying Meng, Asoke K. Nandi
ICASSP5
2019 Learning Spectral and Spatial Features Based on Generative Adversarial Network for Hyperspectral Image Super-Resolution
abstract
Super-resolution (SR) of hyperspectral images (HSIs) aims to enhance the spatial/spectral resolution of hyperspectral imagery and the super-resolved results will benefit many remote sensing applications. A generative adversarial network for HSIs super-resolution (HSRGAN) is proposed in this paper. Specifically, HSRGAN constructs spectral and spatial blocks with residual network in generator to effectively learn spectral and spatial features from HSIs. Furthermore, a new loss function which combines the pixel-wise loss and adversarial loss together is designed to guide the generator to recover images approximating the original HSIs and with finer texture details. Quantitative and qualitative results demonstrate that the proposed HSRGAN is superior to the state of the art methods like SRCNN and SRGAN for HSIs spatial SR.
Ruituo Jiang, Xu Li 0010, Lixin Li 0001, Hongying Meng, Shigang Yue, Lei Zhang 0035
IGARSS5
2019 Diffusion Tensor Image segmentation based on multi-atlas Active Shape Model
abstract
Active Shape Model (ASM) has been successfully applied in the segmentation of Diffusion Tensor Magnetic Resonance Image (DT-MRI, referred to as DTI) of brain. However, due to multiple anatomical structure types, irregular shapes, small gray-scale and large amount of these images, perfect segmentation performance could not be achieved. Especially, it is sensitive to initial values with high computational complexity. In this paper, we introduce the gray information of multiple atlases and the prior information of target shapes into the ASM and propose the Multi-Atlas Active Shape Model (referred to as MA-ASM) approach for DTI segmentation. It was evaluated in a manually labeled database with 7 Region of Interest (ROI)s for each of 20 subjects. In comparison with the state of art method of STAPLE (Simultaneous Truth Performance Level Estimation), the proposed algorithm was closer to the manual segmentation shape by subjective visual effects, and had higher overlap rates and lower error detection rates on quantitative analysis than STAPLE.
Yi Wang 0069, Yangyu Fan, Hongying Meng
Multim. Tools Appl.6
2019 Superpixel-Based Fast Fuzzy C-Means Clustering for Color Image Segmentation
abstract
A great number of improved fuzzy c-means (FCM) clustering algorithms have been widely used for grayscale and color image segmentation. However, most of them are time-consuming and unable to provide desired segmentation results for color images due to two reasons. The first one is that the incorporation of local spatial information often causes a high computational complexity due to the repeated distance computation between clustering centers and pixels within a local neighboring window. The other one is that a regular neighboring window usually breaks up the real local spatial structure of images and thus leads to a poor segmentation. In this work, we propose a superpixel-based fast FCM clustering algorithm that is significantly faster and more robust than stateof-the-art clustering algorithms for color image segmentation. To obtain better local spatial neighborhoods, we first define a multiscale morphological gradient reconstruction operation to obtain a superpixel image with accurate contour. In contrast to traditional neighboring window of fixed size and shape, the superpixel image provides better adaptive and irregular local spatial neighborhoods that are helpful for improving color image segmentation. Second, based on the obtained superpixel image, the original color image is simplified efficiently and its histogram is computed easily by counting the number of pixels in each region of the superpixel image. Finally, we implement FCM with histogram parameter on the superpixel image to obtain the final segmentation result. Experiments performed on synthetic images and real images demonstrate that the proposed algorithm provides better segmentation results and takes less time than state-of-the-art clustering algorithms for color image segmentation.
Tao Lei 0003, Xiaohong Jia 0002, Yanning Zhang 0001, Shigang Liu, Hongying Meng, Asoke K. Nandi
IEEE Trans. Fuzzy Syst.5
2019 Adaptive Morphological Reconstruction for Seeded Image Segmentation
abstract
Morphological reconstruction (MR) is often employed by seeded image segmentation algorithms such as watershed transform and power watershed, as it is able to filter out seeds (regional minima) to reduce over-segmentation. However, the MR might mistakenly filter meaningful seeds that are required for generating accurate segmentation and it is also sensitive to the scale because a single-scale structuring element is employed. In this paper, a novel adaptive morphological reconstruction (AMR) operation is proposed that has three advantages. First, AMR can adaptively filter out useless seeds while preserving meaningful ones. Second, AMR is insensitive to the scale of structuring elements because multiscale structuring elements are employed. Finally, the AMR has two attractive properties: monotonic increasingness and convergence that help seeded segmentation algorithms to achieve a hierarchical segmentation. Experiments clearly demonstrate that the AMR is useful for improving performance of algorithms of seeded image segmentation and seed-based spectral segmentation. Compared to several state-of-the-art algorithms, the proposed algorithms provide better segmentation results requiring less computing time.
Tao Lei 0003, Xiaohong Jia 0002, Tongliang Liu, Shigang Liu, Hongying Meng, Asoke K. Nandi
IEEE Trans. Image Process.5
2018 Linear and Non-Linear Multimodal Fusion for Continuous Affect Estimation In-the-Wild
abstract
Automatic continuous affect recognition from multiple modality in the wild is arguably one of the most challenging research areas in affective computing. In addressing this regression problem, the advantages of the each modality, such as audio, video and text, have been frequently explored but in an isolated way. Little attention has been paid so far to quantify the relationship within these modalities. Motivated to leverage the individual advantages of each modality, this study investigates behavioral modeling of continuous affect estimation, in multimodal fusion approaches, using Linear Regression, Exponent Weighted Decision Fusion and Multi-Gene Genetic Programming. The capabilities of each fusion approach are illustrated by applying it to the formulation of affect estimation generated from multiple modality using classical Support Vector Regression. The proposed fusion methods were applied in the public Sentiment Analysis in the Wild (SEWA) multimodal dataset and the experimental results indicate that employing proper fusion can deliver a significant performance improvement for all affect estimation. The results further show that the proposed systems is competitive or outperform the other state-of-the-art approaches.
Yona Falinie Binti A. Gaus, Hongying Meng
FG2
2018 Accurate Facial Parts Localization and Deep Learning for 3D Facial Expression Recognition
abstract
Meaningful facial parts can convey key cues for both facial action unit detection and expression prediction. Textured 3D face scan can provide both detailed 3D geometric shape and 2D texture appearance cues of the face which are beneficial for Facial Expression Recognition (FER). However, accurate facial parts extraction as well as their fusion are challenging tasks. In this paper, a novel system for 3D FER is designed based on accurate facial parts extraction and deep feature fusion of facial parts. Experiments are conducted on the BU-3DFE database, demonstrating the effectiveness of combing different facial parts, texture and depth cues and reporting the state-of-the-art results in comparison with all existing methods under the same setting.
Asim Jan, Huaxiong Ding, Hongying Meng, Liming Chen 0002, Huibin Li 0001
FG3
2018 Holoscopic 3D Micro-Gesture Database for Wearable Device Interaction
abstract
With the rapid development of augmented reality (AR) and virtual reality (VR) technology, human-computer interaction (HCI) has been greatly improved for gaming interaction of AR and VR control. The finger micro-gesture is a hot research focus due to the growth of the Internet of Things (IoT) and wearable technologies and recently Google has developed a radar based micro-gesture sensor which is Google Soli. Also, there are a number of finger micro-gesture techniques have been developed using Time of Flight (ToF) imaging sensors for wearable 3D glasses such as Atheer mobile glasses. The principle of holoscopic 3D (H3D) imaging mimics fly's eye technique that captures a true 3D optical model of the scene using a microlens array, however, there is a limited progress of holoscopic 3D systems due to the lack of high quality public available database. In this paper, holoscopic 3D camera is used to capture high quality holoscopic 3D micro-gesture video images and a new unique holoscopic 3D micro-gesture (HoMG) database is produced. HoMG database recorded the image sequence of 3 conventional gestures from 40 participants under different settings and conditions. For the purpose of H3D micro-gesture recognition, HoMG has a video subset of 960 videos and a still image subset with 30635 images. Initial micro-gesture recognition on both subsets has been conducted using the traditional 2D image and video features and popular classifiers and some encouraging performance has been achieved. The database will be available for the research communities and speed up the research in the area of holoscopic 3D micro-gesture.
Yi Liu 0061, Hongying Meng, Mohammad Rafiq Swash, Yona Falinie Binti A. Gaus
FG2
2018 Emotion detection from EEG recordings based on supervised and unsupervised dimension reduction
abstract
Summary In recent years, researchers have been trying to detect human emotions from recorded brain signals such as electroencephalogram (EEG) signals. However, due to the high levels of noise from the EEG recordings, a single feature alone cannot achieve good performance. A combination of distinct features is the key for automatic emotion detection. In this paper, we present a hybrid dimension feature reduction scheme using a total of 14 different features extracted from EEG recordings. The scheme combines these distinct features in the feature space using both supervised and unsupervised feature selection processes. Maximum Relevance Minimum Redundancy (mRMR) is applied to re‐order the combined features into max‐relevance with the labels and min‐redundancy of each feature. The generated features are further reduced with principal component analysis (PCA) for extracting the principal components. Experimental results show that the proposed work outperforms the state‐of‐art methods using the same settings in the publicly available DEAP data set.
Hongying Meng, Maozhen Li 0001, Fan Zhang 0101, Asoke K. Nandi
Concurr. Comput. Pract. Exp.2
2018 Significantly Fast and Robust Fuzzy C-Means Clustering Algorithm Based on Morphological Reconstruction and Membership Filtering
abstract
As fuzzy c-means clustering (FCM) algorithm is sensitive to noise, local spatial information is often introduced to an objective function to improve the robustness of the FCM algorithm for image segmentation. However, the introduction of local spatial information often leads to a high computational complexity, arising out of an iterative calculation of the distance between pixels within local spatial neighbors and clustering centers. To address this issue, an improved FCM algorithm based on morphological reconstruction and membership filtering (FRFCM) that is significantly faster and more robust than FCM is proposed in this paper. First, the local spatial information of images is incorporated into FRFCM by introducing morphological reconstruction operation to guarantee noise-immunity and image detail-preservation. Second, the modification of membership partition, based on the distance between pixels within local spatial neighbors and clustering centers, is replaced by local membership filtering that depends only on the spatial neighbors of membership partition. Compared with state-of-the-art algorithms, the proposed FRFCM algorithm is simpler and significantly faster, since it is unnecessary to compute the distance between pixels within local spatial neighbors and clustering centers. In addition, it is efficient for noisy image segmentation because membership filtering are able to improve membership partition matrix efficiently. Experiments performed on synthetic and real-world images demonstrate that the proposed algorithm not only achieves better results, but also requires less time than the state-of-the-art algorithms for image segmentation.
Tao Lei 0003, Xiaohong Jia 0002, Yanning Zhang 0001, Lifeng He, Hongying Meng, Asoke K. Nandi
IEEE Trans. Fuzzy Syst.5
2016 The Automatic Detection of Chronic Pain-Related Expression: Requirements, Challenges and the Multimodal EmoPain Dataset
abstract
Pain-related emotions are a major barrier to effective self rehabilitation in chronic pain. Automated coaching systems capable of detecting these emotions are a potential solution. This paper lays the foundation for the development of such systems by making three contributions. First, through literature reviews, an overview of how pain is expressed in chronic pain and the motivation for detecting it in physical rehabilitation is provided. Second, a fully labelled multimodal dataset (named `EmoPain') containing high resolution multiple-view face videos, head mounted and room audio signals, full body 3D motion capture and electromyographic signals from back muscles is supplied. Natural unconstrained pain related facial expressions and body movement behaviours were elicited from people with chronic pain carrying out physical exercises. Both instructed and non-instructed exercises were considered to reflect traditional scenarios of physiotherapist directed therapy and home-based self-directed therapy. Two sets of labels were assigned: level of pain from facial expressions annotated by eight raters and the occurrence of six pain-related body behaviours segmented by four experts. Third, through exploratory experiments grounded in the data, the factors and challenges in the automated recognition of such expressions and behaviour are described, the paper concludes by discussing potential avenues in the context of these findings also highlighting differences for the two exercise scenarios addressed.
M. S. Hane Aung, Sebastian Kaltwang, Bernardino Romera-Paredes, Brais Martínez, Aneesha Singh, Matteo Cella, Michel F. Valstar, Hongying Meng, Andrew Kemp, Moshen Shafizadeh, Aaron C. Elkins, Natalie Kanakam, Amschel de Rothschild, Nick Tyler, Paul J. Watson, Amanda C. de C. Williams, Maja Pantic, Nadia Bianchi-Berthouze
IEEE Trans. Affect. Comput.8
2016 Time-Delay Neural Network for Continuous Emotional Dimension Prediction From Facial Expression Sequences
abstract
Automatic continuous affective state prediction from naturalistic facial expression is a very challenging research topic but very important in human-computer interaction. One of the main challenges is modeling the dynamics that characterize naturalistic expressions. In this paper, a novel two-stage automatic system is proposed to continuously predict affective dimension values from facial expression videos. In the first stage, traditional regression methods are used to classify each individual video frame, while in the second stage, a time-delay neural network (TDNN) is proposed to model the temporal relationships between consecutive predictions. The two-stage approach separates the emotional state dynamics modeling from an individual emotional state prediction step based on input features. In doing so, the temporal information used by the TDNN is not biased by the high variability between features of consecutive frames and allows the network to more easily exploit the slow changing dynamics between emotional states. The system was fully tested and evaluated on three different facial expression video datasets. Our experimental results demonstrate that the use of a two-stage approach combined with the TDNN to take into account previously classified frames significantly improves the overall performance of continuous emotional state estimation in naturalistic facial expressions. The proposed approach has won the affect recognition sub-challenge of the Third International Audio/Visual Emotion Recognition Challenge.
Hongying Meng, Nadia Bianchi-Berthouze, Yangdong Deng, Jinkuang Cheng, John Cosmas
IEEE Trans. Cybern.1
2015 Social Touch Gesture Recognition using Random Forest and Boosting on Distinct Feature Sets
abstract
Touch is a primary nonverbal communication channel used to communicate emotions or other social messages. Despite its importance, this channel is still very little explored in the affective computing field, as much more focus has been placed on visual and aural channels. In this paper, we investigate the possibility to automatically discriminate between different social touch types. We propose five distinct feature sets for describing touch behaviours captured by a grid of pressure sensors. These features are then combined together by using the Random Forest and Boosting methods for categorizing the touch gesture type. The proposed methods were evaluated on both the HAART (7 gesture types over different surfaces) and the CoST (14 gesture types over the same surface) datasets made available by the Social Touch Gesture Challenge 2015. Well above chance level performances were achieved with a 67% accuracy for the HAART and 59% for the CoST testing datasets respectively.
Yona Falinie Binti A. Gaus, Temitayo A. Olugbade, Asim Jan, Fan Zhang 0101, Hongying Meng, Nadia Bianchi-Berthouze
ICMI7
2014 Affective State Level Recognition in Naturalistic Facial and Vocal Expressions
abstract
Naturalistic affective expressions change at a rate much slower than the typical rate at which video or audio is recorded. This increases the probability that consecutive recorded instants of expressions represent the same affective content. In this paper, we exploit such a relationship to improve the recognition performance of continuous naturalistic affective expressions. Using datasets of naturalistic affective expressions (AVEC 2011 audio and video dataset, PAINFUL video dataset) continuously labeled over time and over different dimensions, we analyze the transitions between levels of those dimensions (e.g., transitions in pain intensity level). We use an information theory approach to show that the transitions occur very slowly and hence suggest modeling them as first-order Markov models. The dimension levels are considered to be the hidden states in the Hidden Markov Model (HMM) framework. Their discrete transition and emission matrices are trained by using the labels provided with the training set. The recognition problem is converted into a best path-finding problem to obtain the best hidden states sequence in HMMs. This is a key difference from previous use of HMMs as classifiers. Modeling of the transitions between dimension levels is integrated in a multistage approach, where the first level performs a mapping between the affective expression features and a soft decision value (e.g., an affective dimension level), and further classification stages are modeled as HMMs that refine that mapping by taking into account the temporal relationships between the output decision labels. The experimental results for each of the unimodal datasets show overall performance to be significantly above that of a standard classification system that does not take into account temporal relationships. In particular, the results on the AVEC 2011 audio dataset outperform all other systems presented at the international competition.
Hongying Meng, Nadia Bianchi-Berthouze
IEEE Trans. Cybern.1
2012 Implementation and Applications of Tri-State Self-Organizing Maps on FPGA
abstract
This paper introduces a tri-state logic self-organizing map (bSOM) designed and implemented on a field programmable gate array (FPGA) chip. The bSOM takes binary inputs and maintains tri-state weights. A novel training rule is presented. The bSOM is well suited to FPGA implementation, trains quicker than the original self-organizing map (SOM), and can be used in clustering and classification problems with binary input data. Two practical applications, character recognition and appearance-based object identification, are used to illustrate the performance of the implementation. The appearance-based object identification forms part of an end-to-end surveillance system implemented wholly on FPGA. In both applications, binary signatures extracted from the objects are processed by the bSOM. The system performance is compared with a traditional SOM with real-valued weights and a strictly binary weighted SOM.
Kofi Appiah, Andrew Hunter, Patrick Dickinson, Hongying Meng
IEEE Trans. Circuits Syst. Video Technol.4
2012 What Does Touch Tell Us about Emotions in Touchscreen-Based Gameplay?
abstract
The increasing number of people playing games on touch-screen mobile phones raises the question of whether touch behaviors reflect players’ emotional states. This prospect would not only be a valuable evaluation indicator for game designers, but also for real-time personalization of the game experience. Psychology studies on acted touch behavior show the existence of discriminative affective profiles. In this article, finger-stroke features during gameplay on an iPod were extracted and their discriminative power analyzed. Machine learning algorithms were used to build systems for automatically discriminating between four emotional states (Excited, Relaxed, Frustrated, Bored), two levels of arousal and two levels of valence. Accuracy reached between 69% and 77% for the four emotional states, and higher results (~89%) were obtained for discriminating between two levels of arousal and two levels of valence. We conclude by discussing the factors relevant to the generalization of the results to applications other than games.
Yuan Gao 0024, Nadia Bianchi-Berthouze, Hongying Meng
ACM Trans. Comput. Hum. Interact.3
2011 Naturalistic Affective Expression Classification by a Multi-stage Approach Based on Hidden Markov Models
Hongying Meng, Nadia Bianchi-Berthouze
ACII (2)1
2011 Multi-score Learning for Affect Recognition: The Case of Body Postures
Hongying Meng, Andrea Kleinsmith, Nadia Bianchi-Berthouze
ACII (1)1
2011 Emotion recognition by two view SVM_2K classifier on dynamic facial expression features
abstract
A novel emotion recognition system has been proposed for classifying facial expression in videos. Firstly, two types of basic facial appearance descriptors were extracted. The first type of descriptor, called Motion History Histogram (MHH), was used to detect temporal changes of each pixels of the face. The second type of descriptor, called Histogram of Local Binary Patterns (LBP), was applied to each frame of the video and was used to capture local textural patterns. Secondly, based on these two basic types of descriptors, two new dynamic facial expression features called MHH_EOH and LBP MCF were proposed. These two features incorporate both dynamic and local information. Finally, the Two View SVK_2K classifier was built to integrate these two dynamic features in an efficient way. The experimental results showed that this method outperformed the baseline results set by the FERA'11 challenge.
Hongying Meng, Bernardino Romera-Paredes, Nadia Bianchi-Berthouze
FG1
2010 Accelerated hardware video object segmentation: From foreground detection to connected components labelling
Kofi Appiah, Andrew Hunter, Patrick Dickinson, Hongying Meng
Comput. Vis. Image Underst.4
2010 A modified model for the Lobula Giant Movement Detector and its FPGA implementation
Hongying Meng, Kofi Appiah, Shigang Yue, Andrew Hunter, Mervyn Hobden, Nigel Priestley, Peter Hobden, Cy Pettit
Comput. Vis. Image Underst.1
2009 A binary Self-Organizing Map and its FPGA implementation
abstract
A binary Self Organizing Map (SOM) has been designed and implemented on a Field Programmable Gate Array (FPGA) chip. A novel learning algorithm which takes binary inputs and maintains tri-state weights is presented. The binary SOM has the capability of recognizing binary input sequences after training. A novel tri-state rule is used in updating the network weights during the training phase. The rule implementation is highly suited to the FPGA architecture, and allows extremely rapid training. This architecture may be used in real-time for fast pattern clustering and classification of binary features.
Kofi Appiah, Andrew Hunter, Hongying Meng, Shigang Yue, Mervyn Hobden, Nigel Priestley, Peter Hobden, Cy Pettit
IJCNN3
2009 A modified sparse distributed memory model for extracting clean patterns from noisy inputs
abstract
The sparse distributed memory (SDM) proposed by Kanerva provides a simple model for human long-term memory, with a strong underlying mathematical theory. However, there are problematic features in the original SDM model that affect its efficiency and performance in real world applications and for hardware implementation. In this paper, we propose modifications to the SDM model that improve its efficiency and performance in pattern recall. First, the address matrix is built using training samples rather than random binary sequences. This improves the recall performance significantly. Second, the content matrix is modified using a simple tri-state logic rule. This reduces the storage requirements of the SDM and simplifies the implementation logic, making it suitable for hardware implementation. The modified model has been tested using pattern recall experiments. It is found that the modified model can recall clean patterns very well from noisy inputs.
Hongying Meng, Kofi Appiah, Andrew Hunter, Shigang Yue, Mervyn Hobden, Nigel Priestley, Peter Hobden, Cy Pettit
IJCNN1
2009 A modified neural network model for Lobula Giant Movement Detector with additional depth movement feature
abstract
The lobula giant movement detector (LGMD) is a wide-field visual neuron that is located in the lobula layer of the locust nervous system. The LGMD increases its firing rate in response to both the velocity of the approaching object and its proximity. It has been found that it can respond to looming stimuli very quickly and can trigger avoidance reactions whenever a rapidly approaching object is detected. It has been successfully applied in visual collision avoidance systems for vehicles and robots. This paper proposes a modified LGMD model that provides additional movement depth direction information. The proposed model retains the simplicity of the previous neural network model, adding only a few new cells. It has been tested on both simulated and recorded video data sets. The experimental results shows that the modified model can very efficiently provide stable information on the depth direction of movement.
Hongying Meng, Shigang Yue, Andrew Hunter, Kofi Appiah, Mervyn Hobden, Nigel Priestley, Peter Hobden, Cy Pettit
IJCNN1
2009 Descriptive temporal template features for visual motion recognition
Hongying Meng, Nick E. Pears
Pattern Recognit. Lett.1
2007 A Human Action Recognition System for Embedded Computer Vision Application
abstract
In this paper, we propose a human action recognition system suitable for embedded computer vision applications in security systems, human-computer interaction and intelligent environments. Our system is suitable for embedded computer vision application based on three reasons. Firstly, the system was based on a linear support vector machine (SVM) classifier where classification progress can be implemented easily and quickly in embedded hardware. Secondly, we use compacted motion features easily obtained from videos. We address the limitations of the well known motion history image (MHI) and propose a new hierarchical motion history histogram (HMHH) feature to represent the motion information. HMHH not only provides rich motion information, but also remains computationally inexpensive. Finally, we combine MHI and HMHH together and extract a low dimension feature vector to be used in the SVM classifiers. Experimental results show that our system achieves significant improvement on the recognition performance.
Hongying Meng, Nick E. Pears, Christopher Bailey 0002
CVPR1
2005 Two view learning: SVM-2K, Theory and Practice
abstract
Kernel methods make it relatively easy to define complex highdimensional feature spaces. This raises the question of how we can identify the relevant subspaces for a particular learning task. When two views of the same phenomenon are available kernel Canonical Correlation Analysis (KCCA) has been shown to be an effective preprocessing step that can improve the performance of classification algorithms such as the Support Vector Machine (SVM). This paper takes this observation to its logical conclusion and proposes a method that combines this two stage learning (KCCA followed by SVM) into a single optimisation termed SVM-2K. We present both experimental and theoretical analysis of the approach showing encouraging results and insights.
Jason D. R. Farquhar, David R. Hardoon, Hongying Meng, John Shawe-Taylor, Sándor Szedmák
NIPS3