Xia Mao

dblp:73/3817 · DBLP profile ↗
← Back
35ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-0700-4437ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 3 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Mining Scene Structural Guidance for Thermal Images in Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation from RGB images has seen significant advancements recently, primarily because it eliminates the need for ground truth data during training. However, applying this technique to thermal images remains challenging due to their inherent characteristics, such as low contrast, low texture, and low signal-to-noise ratio, which impede accurate self-supervision. In this paper, we propose leveraging reliable and distinct scene structural information from thermal images to enhance self-supervised signals. We introduce structural losses, including explicit structural loss in the image space and implicit structural loss in the feature space, to improve self-supervised depth estimation. This approach mitigates the interference caused by the degraded characteristics of thermal images. Our method demonstrates superior performance compared to previous state-of-the-art approaches on the ViViD benchmark dataset, both quantitatively and qualitatively.
Xinchen Ye, Xia Mao, Rui Xu 0002
ICASSP2
2025 A Framework for Runtime Safety of Industrial Control Systems Through Runtime Verification
abstract
Ensuring the safety of complex industrial control systems (ICS) cannot be fully achieved during the design and development phases. Many uncertainties and unknowns only become apparent during real-world operation, especially in the context of Industry 4.0, where ICS integrate increasing characteristics of cyber-physical systems (CPS), such as openness and connectivity. Runtime verification (RV) is extensively employed to guarantee the runtime safety of systems. However, current RV methods face substantial challenges in ICS, particularly due to extensive device heterogeneity, intricate real-time constraints, and the need for coordinating multiple controllers. In this article, we propose a novel framework that incorporates stream-based RV to ensure the runtime safety of ICS. By leveraging a communication bridge based on the open platform communications unified architecture (OPC UA) standard, our framework achieves platform compatibility. This framework, coupled with its nonintrusive verification feature, is well-suited for scenarios involving heterogeneous devices and collaborative controllers. Additionally, stream-based formal specification captures complex time-sensitive constraints, such as real-time synchronizations involving various signals, including triggering, duration, and timeout. To further enhance safety, the framework offers online correction strategies for addressing runtime violations, aiming to preserve or restore system safety. Experimental results from general case studies demonstrate that our approach surpasses existing methods in managing device heterogeneity, complex real-time constraints, and multicontroller cooperation scenarios.
Qin Li 0002, Xia Mao, Ting Wang 0001, Tengfei Li 0002
IEEE Internet Things J.3
2022 Limited text speech synthesis with electroglottograph based on Bi-LSTM and modified Tacotron-2
abstract
Abstract This paper proposes a framework of applying only the EGG signal for speech synthesis in the limited categories of contents scenario. EGG is a sort of physiological signal which can reflect the trends of the vocal cord movement. Note that EGG’s different acquisition method contrasted with speech signals, we exploit its application in speech synthesis under the following two scenarios. (1) To synthesize speeches under high noise circumstances, where clean speech signals are unavailable. (2) To enable dumb people who retain vocal cord vibration to speak again. Our study consists of two stages, EGG to text and text to speech. The first is a text content recognition model based on Bi-LSTM, which converts each EGG signal sample into the corresponding text with a limited class of contents. This model achieves 91.12% accuracy on the validation set in a 20-class content recognition experiment. Then the second step synthesizes speeches with the corresponding text and the EGG signal. Based on modified Tacotron-2, our model gains the Mel cepstral distortion (MCD) of 5.877 and the mean opinion score (MOS) of 3.87, which is comparable with the state-of-the-art performance and achieves an improvement by 0.42 and a relatively smaller model size than the origin Tacotron-2. Considering to introduce the characteristics of speakers contained in EGG to the final synthesized speech, we put forward a fine-grained fundamental frequency modification method, which adjusts the fundamental frequency according to EGG signals and achieves a lower MCD of 5.781 and a higher MOS of 3.94 than that without modification.
Lijiang Chen, Xia Mao, Qi Zhao 0037
Appl. Intell.4
2022 A refinement development approach for enhancing the safety of PLC programs with Event-B
Xia Mao, Yueling Zhang, Jianqi Shi, Yanhong Huang, Qin Li 0002
Sci. Comput. Program.1
2022 Programmable Logic Controllers Past Linear Temporal Logic for Monitoring Applications in Industrial Control Systems
abstract
Programmable logic controllers (PLC), which are widely applied in modern industrial control systems (ICS), work as the controller of sensors and actuators in ICS. These systems require strict correctness, especially for safety-critical systems. Currently, increasingly ICS move to “come online” scenarios to enhance cyber-physical features, but it makes them more vulnerable due to acquiring increased interconnection accompanied by weakening physical isolation. Moreover, with the more complex controlling environment, such as hundreds of more I/O points and more diverse field buses, the incorrect executions of PLC might cause the failure of the overall ICS. In this article, we examine how the security and safety of running PLC could be enhanced in both developing and deploying stages of ICS. We propose a novel application of runtime verification to guarantee the security and safety of real-world ICS. As a variant of temporal logic, PLC past linear temporal logic (PPLTL) is proposed to specify the security and safety properties of PLC. Using PPLTL, we synthesize monitors to improve the PLC program’s security and safety as a partner of testing and static verification. Our monitors provide twofold processing in a nonintrusive manner: One is filtering abnormal input data before invading the original programs, the other is double-checking the output signals before driving the actuators. We use several case studies and benchmarks to demonstrate the efficiency of the approach. The empirical results show that the time overhead and memory occupation are tiny.
Xia Mao, Xin Li 0109, Yanhong Huang, Jianqi Shi, Yueling Zhang
IEEE Trans. Ind. Informatics1
2021 Data Flow Testing for PLC Programs via Dynamic Symbolic Execution
abstract
Programmable logic controllers (PLCs) are broadly used in the safety-critical industrial field, which requires high reliability to avoid catastrophes. Data flow testing (DFT) focuses on data flow relationships in a program and has a stronger fault-detection ability than other control flow-based testing. However, there is no automated testing tool supporting DFT for PLC programs. Hence, we propose an automated data flow testing framework for PLC programs. Our DFT framework is based on dynamic symbolic execution (DSE). Considering the cyclic execution feature of PLC programs, our approach needs reachable states which can be provided by branch testing. Besides, our approach improves testing performance through a novel guided path search algorithm. Furthermore, we evaluate our approach on several programs to demonstrate that this approach is practical and effective.
Weigang He, Xia Mao, Ting Su 0001, Yanhong Huang, Jianqi Shi
APSEC2
2021 Dynamically Detecting Invariants for Automatic Testing PLC Programs (S)
abstract
Since programmable logic controllers (PLCs) control safety-critical infrastructures, examining the PLC software satisfies the high-reliability specifications necessary to ensure the safeness of PLCs.However, prior works have limitations in finding defects in the PLC source code.Static verification techniques suffer from notable false positives without capturing runtime behavior.The symbolic execution and conformance testing technique captures the relations of inputs and outputs.It is not sufficient to consider only the data constraints as the PLC operates in real-time.In this paper, we propose a novel approach in the detection of the runtime behavior of PLC programs with incorporated time constraints.This testing approach automatically finds implementation errors in PLC programs by mining invariants from runtime traces.As the existing tools mine only data or time invariants which are inadequate to test PLC programs, our approach focuses on the interplay of data and time invariants.Dynamically detected datatime invariants are then checked with the safety specifications.We evaluate the usefulness of our approach in a real-life case.The experimental results show that the proposed approach can find errors in PLC programs effectively.
Xia Mao, Yanhong Huang, Jianqi Shi, Yang Yang 0141
SEKE2
2021 Runtime Verification of Spatio-Temporal Specification Language
Tengfei Li 0002, Jing Liu 0012, Haiying Sun, Xiaohong Chen 0007, Ling Yin 0002, Xia Mao
Mob. Networks Appl.6
2019 Sparsity Regularization Discriminant Projection for Feature Extraction
Sen Yuan, Xia Mao, Lijiang Chen
Neural Process. Lett.2
2018 Learning deep features to recognise speech emotion using merged deep CNN
abstract
This study aims at learning deep features from different data to recognise speech emotion. The authors designed a merged convolutional neural network (CNN), which had two branches, one being one‐dimensional (1D) CNN branch and another 2D CNN branch, to learn the high‐level features from raw audio clips and log‐mel spectrograms. The building of the merged deep CNN consists of two steps. First, one 1D CNN and one 2D CNN architectures were designed and evaluated; then, after the deletion of the second dense layers, the two CNN architectures were merged together. To speed up the training of the merged CNN, transfer learning was introduced in the training. The 1D CNN and 2D CNN were trained first. Then, the learned features of the 1D CNN and 2D CNN were repurposed and transferred to the merged CNN. Finally, the merged deep CNN initialised with transferred features was fine‐tuned. Two hyperparameters of the designed architectures were chosen through Bayesian optimisation in the training. The experiments conducted on two benchmark datasets show that the merged deep CNN can improve emotion classification performance significantly.
Jianfeng Zhao 0005, Xia Mao, Lijiang Chen
IET Signal Process.2
2018 Exponential elastic preserving projections for facial expression recognition
Sen Yuan, Xia Mao
Neurocomputing2
2018 Elastic preserving projections based on L1-norm maximization
Sen Yuan, Xia Mao, Lijiang Chen
Multim. Tools Appl.2
2018 A closed-form solution to the graph total variation problem for continuous emotion profiling in noisy environment
Shaoling Jing, Xia Mao, Lijiang Chen, Maria Colomba Comes, Arianna Mencattini, Grazia Raguso, Fabien Ringeval, Björn W. Schuller, Corrado Di Natale, Eugenio Martinelli
Speech Commun.2
2018 Learning deep facial expression features from image and optical flow sequences using 3D CNN
Jianfeng Zhao 0005, Xia Mao, Jian Zhang 0002
Vis. Comput.2
2017 Decomposition and Collaboration of Industrial Control System with Resource Constraints
abstract
With the development of "Industry 4.0", the scale and complexity of industrial control system grow rapidly. Hence, the analysis and verification of such systems face really big challenges. Industry requires a reliable approach for decomposing the existing complex system model to multiple fine-grained and interactive models. In this paper, we propose a general event-triggered language named IMCL for modeling industrial control systems. IMCL can describe the physical resources and system in one unified model. Following the given physical resource constraints, we present the reliable and efficient decomposition and collaboration algorithms based on IMCL models to meet the industrial requirements. In particular, we have implemented these algorithms in a tool and get same encouraging results.
Jiawen Xiong, Xia Mao, Jianqi Shi, Yanhong Huang
ICECCS3
2017 Human interaction recognition fusing multiple features of depth sequences
abstract
Human interaction recognition has played a major role in building intelligent video surveillance systems. Recently, depth data captured by the emerging RGB‐D sensors began to show its importability in human interaction recognition. This study proposes a novel framework for human interaction recognition using depth information including an algorithm to reconstruct depth sequence with as few key frames as possible. The proposed framework includes two essential modules. First, key frames extraction by sparse constraint, then the fusion multi‐feature, is constructed by using two types of available features and Max‐pooling, respectively. Finally, multiple features are directly sent to the SVM for the recognition of the human activity. This study explores the static and dynamic feature fusion method to improve the recognition performance with contextual relevance of continuous frames. A weight is used to fuse shape and optical flow features, which not only enhance the description capability of human behavioural characteristics in the spatiotemporal domain, but also effectively reduces the adverse impact of certain distortion point of interest for target recognition. Experimental results show that the proposed approach yields considerable performance improvement over the state‐of‐the‐art approaches with respect to accuracy on a public action dataset.
Xia Mao, Lijiang Chen
IET Comput. Vis.2
2017 Multimodal data fusion for SB-JPALS status prediction under antenna motion fault mode
Xia Mao
Neurocomputing2
2017 Multimodal Data fusion for SRGPS antenna motion error reduction
Xia Mao
Multim. Tools Appl.2
2017 Multilinear Spatial Discriminant Analysis for Dimensionality Reduction
abstract
In the last few years, great efforts have been made to extend the linear projection technique (LPT) for multidimensional data (i.e., tensor), generally referred to as the multilinear projection technique (MPT). The vectorized nature of LPT requires high-dimensional data to be converted into vector, and hence may lose spatial neighborhood information of raw data. MPT well addresses this problem by encoding multidimensional data as general tensors of a second or even higher order. In this paper, we propose a novel multilinear projection technique, called multilinear spatial discriminant analysis (MSDA), to identify the underlying manifold of high-order tensor data. MSDA considers both the nonlocal structure and the local structure of data in the transform domain, seeking to learn the projection matrices from all directions of tensor data that simultaneously maximize the nonlocal structure and minimize the local structure. Different from multilinear principal component analysis (MPCA) that aims to preserve the global structure and tensor locality preserving projection (TLPP) that is in favor of preserving the local structure, MSDA seeks a tradeoff between the nonlocal (global) and local structures so as to drive its discriminant information from the range of the non-local structure and the range of the local structure. This spatial discriminant characteristic makes MSDA have more powerful manifold preserving ability than TLPP and MPCA. Theoretical analysis shows that traditional MPTs, such as multilinear linear discriminant analysis, TLPP, MPCA, and tensor maximum margin criterion, could be derived from the MSDA model by setting different graphs and constraints. Extensive experiments on face databases (ORL, CMU PIE, and the extended Yale-B) and the Weizmann action database demonstrate the effectiveness of the proposed MSDA method.
Sen Yuan, Xia Mao, Lijiang Chen
IEEE Trans. Image Process.2
2016 Human action recognition based on tensor shape descriptor
abstract
Human action recognition is an important task. This study presents an efficient framework for recognising action with a 3D skeleton kinematic joint model in less computational time for practical usage. First, a tensor shape descriptor (TSD) is proposed in this study, which takes advantage of the spatial independence of body joints, avoids a lot of difficult problem of the explicit motion estimation required in traditional methods, reserves the spatial information of each frame. Thus, the new TSD is a complete and view‐invariant descriptor. Second, a novel tensor dynamic time warping (TDTW) method is proposed to measure joint‐to‐joint similarity of 3D skeletal body joints locally in the temporal extent, which is implemented by extending DTW to that of two multiway data arrays (or tensors). Then, a multi‐linear projection process is employed to map the TSD to a low‐dimensional tensor subspace, which is classified by the nearest neighbour classifier. The experiment results on the public action data set (MSR‐Action3D) and motion capture data set (CMU_Mocap) show that the proposed method can achieve a comparable or better performance in recognition accuracy compared with the state‐of‐the‐art approaches.
Xia Mao, Xiao-Geng Liang
IET Comput. Vis.2
2016 Illumination compensation for facial feature point localization in a single 2D face image
Jizheng Yi, Xia Mao, Lijiang Chen, Alberto Rovetta
Neurocomputing2
2016 Text-Independent Phoneme Segmentation Combining EGG and Speech Data
abstract
A new approach for text-independent phoneme segmentation at sampling point level is proposed in this paper. The algorithm consists of two phases: First, the voiced sections in speech data are detected using the information of vocal folds vibration contained in electroglottograph (EGG). A Hilbert envelope feature is adopted to achieve sampling point level detection accuracy. Second, the voiced sections and other sections are treated separately. Each voiced section is divided into several candidate phonemes using the Viterbi algorithm. Then adjacent candidate phonemes are merged based on a Hotellings T-square test method. For other sections, the unvoiced consonants are detected from silence based on a singularity exponent feature. Comparison experiments show that the proposed method has better performance than the existing ones for a variety of tolerances, and is more robust to noise.
Lijiang Chen, Xia Mao, Hong Yan 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Trajectory-based view-invariant hand gesture recognition by fusing shape and orientation
abstract
Traditional studies in vision‐based hand gesture recognition remain rooted in view‐dependent representations, and hence users are forced to be fronto‐parallel to the camera. To solve this problem, view‐invariant gesture recognition aims to make the recognition result independent of viewpoint changes. However, in current works the view‐invariance is achieved at the price of mixing different gesture patterns that have similar trajectory curve shape but different semantic meanings. For example, the gesture ‘push’ can be mistaken as ‘drag’ from another viewpoint. To address this shortcoming, in this study, the authors use a shape descriptor to extract the view‐invariant features of a three‐dimensional (3D) trajectory. As the shape features are invariant to omnidirectional viewpoint changes, the orientation features are then added into weight different rotation angles so that similar trajectory shapes are better separated. The proposed method was conducted on two different databases, including a popular Australian Sign Language database and a challenging Kinect Hand Trajectory database. Experimental results show that the proposed algorithm achieves a higher average recognition rate than the state‐of‐the‐art approaches, and can better distinguish confusing gestures while meeting the view‐invariant condition.
Xia Mao, Lijiang Chen, Yu-Li Xue
IET Comput. Vis.2
2015 Kernel optimization using nonparametric Fisher criterion in the subspace
Xia Mao, Lijiang Chen, Yu-Li Xue, Alberto Rovetta
Pattern Recognit. Lett.2
2014 View-Invariant Gesture Recognition Using Nonparametric Shape Descriptor
abstract
In this paper we propose a new method for view-invariant gesture recognition, based on what we call nonparametric shape descriptor. We represent gestures as 3D motion trajectories and then we prove that the shape of a trajectory is equivalent to the Euclidean distances between all its points. The set of point-to-point distances description is mapped to a high-dimensional kernel space by kernel principal component analysis (KPCA), and then nonparametric discriminant analysis (NDA) is used to extract the view-invariant shape features as the input for pattern classification. The algorithm is performed on a public dataset, and shows better view-invariant performance than other state-of-the-art methods.
Xia Mao, Lijiang Chen, Yu-Li Xue, Angelo Compare
ICPR2
2014 Facial expression recognition considering individual differences in facial structure and texture
abstract
Facial expression recognition (FER) plays an important role in human–computer interaction. The recent years have witnessed an increasing trend of various approaches for the FER, but these approaches usually do not consider the effect of individual differences to the recognition result. When the face images change from neutral to a certain expression, the changing information constituted of the structural characteristics and the texture information can provide rich important clues not seen in either face image. Therefore it is believed to be of great importance for machine vision. This study proposes a novel FER algorithm by exploiting the structural characteristics and the texture information hiding in the image space. Firstly, the feature points are marked by an active appearance model. Secondly, three facial features, which are feature point distance ratio coefficient, connection angle ratio coefficient and skin deformation energy parameter, are proposed to eliminate the differences among the individuals. Finally, a radial basis function neural network is utilised as the classifier for the FER. Extensive experimental results on the Cohn–Kanade database and the Beihang University (BHU) facial expression database show the significant advantages of the proposed method over the existing ones.
Jizheng Yi, Xia Mao, Lijiang Chen, Yu-Li Xue, Angelo Compare
IET Comput. Vis.2
2013 Speech Emotional Features Extraction Based on Electroglottograph
abstract
This study proposes two classes of speech emotional features extracted from electroglottography (EGG) and speech signal. The power-law distribution coefficients (PLDC) of voiced segments duration, pitch rise duration, and pitch down duration are obtained to reflect the information of vocal folds excitation. The real discrete cosine transform coefficients of the normalized spectrum of EGG and speech signal are calculated to reflect the information of vocal tract modulation. Two experiments are carried out. One is of proposed features and traditional features based on sequential forward floating search and sequential backward floating search. The other is the comparative emotion recognition based on support vector machine. The results show that proposed features are better than those commonly used in the case of speaker-independent and content-independent speech emotion recognition.
Lijiang Chen, Xia Mao, Pengfei Wei 0001, Angelo Compare
Neural Comput.2
2012 Emphasizing on the Timing and Type - Enhancing the Backchannel Performance of Virtual Agent
Xia Mao, Yu-Li Xue
ICAART (2)1
2012 Mandarin emotion recognition combining acoustic and emotional point information
Lijiang Chen, Xia Mao, Pengfei Wei 0001, Yu-Li Xue, Mitsuru Ishizuka
Appl. Intell.2
2012 EEMML: the emotional eye movement animation toolkit
Xia Mao
Multim. Tools Appl.2
2011 Combined pattern search optimization of feature extraction and classification parameters in facial recognition
Catalin-Daniel Caleanu, Xia Mao, Gilbert Pradel, Sorin Moga, Yu-Li Xue
Pattern Recognit. Lett.2
2010 A PCA-based approach for exploring space-time structure of urban mobility dynamics
abstract
Understanding of urban mobility dynamics benefits both aggregate human mobility in wireless communications, and the planning and provision of urban facilities and services. Due to the high penetration of cell phones, the cellular networks provide information for urban dynamics with large spatial extent and continuous temporal coverage. In this paper, a novel approach is proposed to explore the space-time structure of urban dynamics, based on the original data collected by cellular networks in a southern city of China, recording population distribution by dividing the city into thousands of pixels. By applying principal component analysis, the intrinsic dimensionality is revealed. The structure of all the pixel population variations could be well captured by a small set of eigen pixel population variations. According to the classification of eigen pixel population variations, each pixel population variation can be decomposed into three constitutions: deterministic trends, short-lived spikes, and noise. Moreover, the most significant eigen pixel population variations are utilized in the applications of forecasting and anomaly detection.
Yue Wang 0007, Hongbo Si, Xia Mao, Xiuming Shan
IWCMC4
2009 Providing expressive eye movement to virtual agents
abstract
Non-verbal behavior, particularly eye movement, plays a fundamental role in nonverbal communication among people. In order to realize natural and intuitive human-agent interaction, the virtual agents need to employ this communicative channel effectively. Against this background, our research addresses the problem of emotionally expressive eye movement manner by describing a preliminary approach based on the parameters picked from real-time eye movement data (pupil size, blink rate and saccade).
Xia Mao
ICMI2
2008 An Extension of MPML with Emotion Recognition Functions Attached
Xia Mao, Haiyan Bao
IVA1
2008 Describing and Generating Web-Based Affective Human-Agent Interaction
Xia Mao, Haiyan Bao
KES (1)1