Xi Zhao 0001

dblp:95/6322-1 · DBLP profile ↗
← Back
41ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0001-9983-6366ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 2 since 2021Computer networks · 8 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 SegAuth: Semantic-Aware Multimotion Behavioral Biometric-Based Implicit Authentication
abstract
Multi-motion behavioral biometrics based implicit authentication leverages individual unique behavioral patterns for user authentication. However, traditional methods using fixed-sized time window segmentation often disrupt local temporal structures and overlook behavioral semantics. This paper investigates a semantic-aware segmentation-based implicit authentication approach, yet challenges persist in achieving semantically consistent segmentation, fixed-dimension representations for variable-length data, and robust modeling under large intra-class variance. Towards this end, we propose SegAuth, a novel semantic-aware multi-motion behavioral biometrics based implicit authentication system. Specifically, given the input raw multi-motion data, SegAuth first adopts a data-driven semantic-aware segmentation method to adaptively generate variable-sized segments, capturing fine-grained behavioral patterns for authentication. Next, SegAuth proposes a causal temporal convolutional network, which allows to learn the effective embedding of varied-sized multi-motion data segments. Finally, a multi-center deep one-class classifier–based authentication model is developed to capture behavioral representations from genuine user data characterized by high intra-class variance, allowing it to identify and authenticate behaviors that differ from normal patterns. Extensive experiments are conducted on a large-scale uncontrolled evaluation dataset. The experimental results demonstrate the state-of-the-art authentication performance of SegAuth.
Zhihao Shen 0001, Chengmei Zhao, Xi Zhao 0001, Cong Zou, Jiakun Zhao
IEEE Internet Things J.3
2026 Multi-Motion Spatio-Temporal Graph-Based User Behavior Representation for Enhanced Smartphone Security
abstract
As central hubs of the Internet of Everything, smart phones integrate essential functions such as payments, navigation, and IoT connectivity. However, this expanded functionality also heightens security risks. Motion dynamics biometrics, which utilizes motion patterns from user-phone interactions captured via multi-sensor data, has emerged as a promising solution for smartphone security. Offering continuous and unobtrusive protection by analyzing natural user interactions, it still faces challenges in effectively modeling the complex spatio-temporal dynamics between the user and the phone within multi-sensor data. This paper focuses on leveraging graph neural networks (GNNs) to enhance user behavior modeling for smartphone security protection by capturing the relationships within motion sensor data, but it is non-trivial due to the characteristics of complexity, asynchrony, and temporal dependencies of multi motion sensor data. Towards this end, we propose MotionGNN, a multi-motion spatial-temporal graph based behavior modeling framework for user identification and authentication. Specifically, MotionGNN first divides the input multi-motion sensor data into a sequence of segments adaptively by developing a context-aware data segmentation method. Then, MotionGNN constructs fully connected spatio-temporal graphs to model sensor dependencies and temporal dynamics. Finally, windowing graph convolutions are adopted to learn user behavior representations. To evaluate the performance of MotionGNN, we collect a large-scale dataset from real-world scenarios. Extensive experiments demonstrate the state-of-the-art performance of MotionGNN in user identi fication and authentication tasks. We also test MotionGNN for 7 days on smartphones, showing high authentication accuracy with minimal battery and memory usage, making it a reliable solution for smartphone security protection.
Zhihao Shen 0001, Chengmei Zhao, Cong Zou, Xi Zhao 0001, Jiakun Zhao, Jianhua Zou
IEEE Trans. Dependable Secur. Comput.4
2025 CryptoMixer: Fine-grained market information-aware MLP Networks for Individual Cryptocurrency Trading Prediction
abstract
Accurately predicting user trading behavior in decentralized exchanges is essential for investors to mitigate risks and optimize their trading strategies. While existing research primarily focuses on predicting trading behavior in stock markets, these methods often struggle to adapt to the distinct nature of cryptocurrency trading. Specifically, they face issues such as limited adaptivity to high-frequency and algorithmic trading, as well as an insufficient consideration of fine-grained real-time market participants' behavior.Thanks to the pending mechanism of blockchain, it becomes possible to capture traders' interactions before transactions are finalized, providing valuable insights into market state. However, accurately modeling and predicting trading behavior in decentralized exchanges presents challenges, including limited adaptability to high-frequency trading, a lack of fine-grained transaction data, and high computational costs. This work proposes CryptoMixer, a lightweight fine-grained market information-aware multilayer perceptron (MLP)-based model for high-frequency cryptocurrency trading behavior prediction. Specifically, to overcome the sparsity and asynchrony of user behavior data, CryptoMixer develops a Market Information Augmenter that aggregates historical transaction data of users. Furthermore, CryptoMixer designs a Market Information Mixer as well as a Two-stream MLP Fusion Mixer to capture fine-grained user trading behavior patterns. We evaluate CryptoMixer on real-world user trading datasets from the Uniswap decentralized finance platform. Experimental results demonstrate that CryptoMixer outperforms traditional prediction models while maintaining low computational overheads, providing a practical solution for real-time cryptocurrency trading behavior prediction. The code is available at https://github.com/aqua111000/CryptoMixer.
Tingsheng Feng, Zhihao Shen 0001, Xi Zhao 0001, Xiaoni Lu
KDD (2)3
2024 The effect of fixed physical usage patterns on the engagement of physical activity apps: a real-world data analysis
abstract
Physical activity applications (PA apps) offer low-cost, time-space-independent interventions that make it possible to promote public health. To increase users’ stickiness, the commercially available PA apps usually provide various services to adapt to different app usage patterns of users, thus helping them develop the habit of using apps. However, evidence was rare about whether different usage patterns are associated with the maintenance of PA app use. In this study, we introduce dual process theories and quantify users’ app usage patterns in two dimensions: time and space. We analyse the impact of fixed usage patterns on app engagement, by collecting usage data from several commercial-available PA apps in China, which includes 9,175 users. Results show that repeatedly using the PA app at the same time and space could reduce the decline of future app engagement. Moreover, a high degree of self-regulation capacity can mitigate the negative effect of non-fixed usage patterns. This study extends the understanding of health behaviour intervention from the behaviour change level to the behaviour maintenance level. In addition, it provides practical insights for PA app developers in terms of designing behaviour change technologies and for policymakers in terms of facilitating the public’s daily physical activities.
Xi Zhao 0001
Behav. Inf. Technol.2
2024 Cognitive process-driven model design: A deep learning recommendation model with textual review and context
abstract
Online reviews play a crucial role in comprehending user rating behavior and improving personalized recommendations in e-commerce. However, existing review-based recommendation systems ignore the influence of theory-driven and context information. This paper proposes the Deep Learning Recommendation Model with Textual Review and Context (DeepRM-TC), which is built upon a cognitive process-driven approach to improve the quality and interpretability of recommendations. The DeepRM-TC framework mimics the human brain's cognitive process for predicting user rating behavior. Essentially, the simulation of human cognitive processing is manifested in several aspects of the designed neural network , including treating rating prediction to an attitude question, mapping raw data in latent space as the user's belief, injecting attention mechanisms for rendering judgment and predicting rating as an answer. Furthermore, we design a three-layer coattention mechanism to adaptively match product information based on users' preferences in various contexts. This mechanism extracts finer-grained interaction information from user–product–context pairs. Extensive experiments on real datasets demonstrate that our proposed model outperforms existing state-of-the-art models. We demonstrate the importance of context information and the three-layer coattention mechanism in enhancing recommendation accuracy through ablation and hierarchic analysis, respectively. Additionally, we further validate the performance of our model through data sparseness analysis, scalability analysis, other datasets, classification analysis , and user study.
Xi Zhao 0001, Ningning Liu, Zhihao Shen 0001, Cong Zou
Decis. Support Syst.2
2024 IncreAuth: Incremental-Learning-Based Behavioral Biometric Authentication on Smartphones
abstract
Touch behavior biometric has been widely studied for continuous authentication on mobile devices, which provides a more secure authentication in an implicit process. However, the existing touch behavior biometric-based authentication systems suffer from two issues. First, the existing touch behavior representation methods are hard to characterize touch operations under complex usage context. Second, the authentication accuracy of existing authentication models is inclined to degrade over time in a long-term real-life usage scenario due to change in data distribution caused by varying touch behavior. Toward this end, in this article, we develop IncreAuth, an incremental learning-based continuous authentication framework, which allows to provide effective stable authentication performance in the long-term smartphone usage scenario. Specifically, we first propose a novel context-aware feature set to characterize touch behavior patterns in complex usage context. Then, we develop an authentication model GBDTNN, which integrates the advantages of a gradient boosting decision tree model for processing our high-dimensional feature set and neural network model for efficient online updating. A behavior drift-based online updating mechanism is also designed to learn both long-term and short-term touch behavior patterns. To evaluate our framework, we construct a large-scale smartphone usage data set over two months collected from the unconstrained environment. Extensive experiments demonstrate that IncreAuth achieves the state-of-the-art and stable authentication accuracy over time and low system overheads.
Zhihao Shen 0001, Xi Zhao 0001, Jianhua Zou
IEEE Internet Things J.3
2024 CT-Auth: Capacitive Touchscreen-Based Continuous Authentication on Smartphones
abstract
Continuous authentication, which provides identity verification using behavioral biometrics in an implicit and transparent manner, has shown potentials for protecting privacy. As the most common way of human-computer interaction, touch behavior pattern of each user has been proven distinctive and widely adopted for continuous authentication. However, most touch based solutions rely on the touchscreen signals obtained from high-level application programming interfaces, which are hard to characterize fine-grained appearance and contour profile of contact fingertips as well as dynamic sliding information in a touch gesture. In this paper, we propose a continuous authentication framework called CT-Auth, which leverages raw capacitive value collected from capacitive touchscreen on smartphone as a descriptor of touch behavior for authentication. Specifically, we first develop a three-dimensional convolution neural network model for capturing intra-gesture spatial-temporal feature and a structure extraction model for capturing structural information between moving fingertips of a touch gesture and touchscreen. A recurrent neural network based model is also applied for capturing temporal patterns among a sequence of touch gestures. To evaluate the effectiveness of our framework, we recruit 100 volunteers over 2 months and collect a large-scale dataset in the unconstrained conditions. Extensive experiments reveal that CT-Auth provides the state-of-the-art authentication accuracy.
Zhihao Shen 0001, Xi Zhao 0001, Jianhua Zou
IEEE Trans. Knowl. Data Eng.3
2024 DMM: A Deep Reinforcement Learning Based Map Matching Framework for Cellular Data
abstract
This paper presents a novel map matching framework that adopts deep learning techniques to map a sequence of cell tower locations to a trajectory on a road network. Map matching is an essential pre-processing step for many applications, such as traffic optimization and human mobility analysis. However, most recent approaches are based on hidden Markov models (HMMs) or neural networks that are hard to consider high-order location information or heuristics observed from real driving scenarios. In this paper, we develop a deep reinforcement learning based map matching framework for cellular data, named as DMM, which adopts a recurrent neural network (RNN) coupled with a reinforcement learning scheme to identify the most-likely trajectory of roads given a sequence of cell towers. To transform DMM into a practical system, several challenges are addressed by developing a set of techniques, including spatial-aware representation of input cell tower sequences, an encoder-decoder based RNN network for map matching model with variable-length input and output, and a global heuristics-driven reinforcement learning based scheme for optimizing the parameters of the encoder-decoder map matching model. Extensive experiments on a large-scale anonymized cellular dataset reveal that DMM provides high map matching accuracy and fast inference time.
Zhihao Shen 0001, Kang Yang 0005, Xi Zhao 0001, Jianhua Zou, Wan Du, Junjie Wu 0002
IEEE Trans. Knowl. Data Eng.3
2024 Retrieving Similar Trajectories from Cellular Data of Multiple Carriers at City Scale
abstract
Retrieving similar trajectories aims to search for the trajectories that are close to a query trajectory in spatio-temporal domain from a large trajectory dataset. This is critical for a variety of applications, like transportation planning and mobility analysis. Unlike previous studies that perform similar trajectory retrieval on fine-grained GPS data or single cellular carrier, we investigate the feasibility of finding similar trajectories from cellular data of multiple carriers, which provide more comprehensive coverage of population and space. To handle the issues of spatial bias of cellular data from multiple carriers, coarse spatial granularity, and irregular sparse temporal sampling, we develop a holistic system cellSim . Specifically, to avoid the issue of spatial bias, we first propose a novel map matching approach, which transforms the cell tower sequences from multiple carriers to routes on a unified road map. Then, to address the issue of temporal sparse sampling, we generate multiple routes with different confidences to increase the probability of finding truly similar trajectories. Finally, a new trajectory similarity measure is developed for similar trajectory search by calculating the similarities between the irregularly-sampled trajectories. Extensive experiments on a large-scale cellular dataset from two carriers and real-world 1,701 km query trajectories reveal that cellSim provides state-of-the-art performance for similar trajectory retrieval.
Zhihao Shen 0001, Wan Du, Xi Zhao 0001, Jianhua Zou
ACM Trans. Sens. Networks3
2023 GinApp: An Inductive Graph Learning based Framework for Mobile Application Usage Prediction
Zhihao Shen 0001, Xi Zhao 0001, Jianhua Zou
INFOCOM2
2023 DeepAPP: A Deep Reinforcement Learning Framework for Mobile Application Usage Prediction
abstract
This paper aims to predict a set of apps a user will open on her mobile device in the next time slot. Such an information is essential for many smartphone operations, e.g., app pre-loading and content pre-caching, to improve user experience. However, it is hard to build an explicit model that accurately captures the complex environment context and predicts a set of apps at one time. This paper presents a deep reinforcement learning framework, named as DeepAPP, which learns a model-free predictive neural network from historical app usage data. Meanwhile, an online updating strategy is designed to adapt the predictive network to the time-varying app usage behavior. To transform DeepAPP into a practical deep reinforcement learning system, several challenges are addressed by developing a context representation method for complex contextual environment, a general agent for overcoming data sparsity and a lightweight personalized agent for minimizing the prediction time. Extensive experiments on a large-scale anonymized app usage dataset reveal that DeepAPP provides high accuracy (precision 70.6 percent and recall of 62.4 percent) and reduces the prediction time of the state-of-the-art by 6.58×. A field experiment of 29 participants demonstrates DeepAPP can effectively reduce launch time of apps.
Zhihao Shen 0001, Kang Yang 0005, Xi Zhao 0001, Jianhua Zou, Wan Du
IEEE Trans. Mob. Comput.3
2023 ATPP: A Mobile App Prediction System Based on Deep Marked Temporal Point Processes
abstract
Predicting the next application (app) a user will open is essential for improving the user experience, e.g., app pre-loading and app recommendation. Unlike previous solutions that only predict which app the user will open, this article predicts both the next app and the time to open it. Time prediction is essential to avoid loading the next app too early and consuming unnecessary resources on smartphones. To predict the next app and open time jointly, we model the app usage sequence as a marked temporal point process (MTPP), whose conditional intensity function can capture the probability of a new app usage event. We develop a novel data-driven MTPP-based app prediction system, named ATPP (App Temporal Point Process), which adopts a recurrent neural network architecture to learn the MTPP conditional intensity function for app prediction. ATPP adopts a set of techniques to incorporate the unique features of app prediction in our RNN architecture, including learning the correlated usage behavior of different apps by representation learning, the temporal dependency of app usage by an attention mechanism, and the location-related app usage behavior by feature extraction and fusion layer. We conduct extensive experiments on a large-scale anonymized app usage dataset to verify ATPP’s effectiveness.
Kang Yang 0005, Xi Zhao 0001, Jianhua Zou, Wan Du
ACM Trans. Sens. Networks2
2022 MMAuth: A Continuous Authentication Framework on Smartphones Using Multiple Modalities
abstract
With the wide use of smartphones, more private data are collected and saved in the smartphones. This raises higher requirements for secure and effective user authentication scheme. Continuous authentication leverages behavioral biometrics as identity information and shows promising characteristics for user verification in a continuous and passive means. However, most studies require users to operate the smartphones in a specific mobile application or perform user-defined touch operations. This paper studies the continuous authentication on smartphones in the wild, where it is hard to characterize touching behavior accurately due to the complexity of usage context and cross-use of various types of touch gestures. Towards this end, in this paper, we propose a continuous authentication framework using multiple modalities, named as MMAuth, which integrates the heterogeneous information of user identity from multiple modalities (e.g., motion movement pattern, touch dynamics, usage context). A time-extended behavioral feature set (TEB) and a deep learning based one-class classifier (DeSVDD) are developed for performing more accurate authentication. Evaluations are conducted using a novel unconstrained smartphone usage dataset collected from 100 volunteers in real world as well as a public laboratory dataset. Extensive experimental results demonstrate that the state-of-the-art authentication performance of MMAuth in both unconstrained and laboratory environment, and the effectiveness of its two proposed modules (the TEB feature set and the DeSVDD classifier). Additional experiments on system robustness, in terms of usability to different touch gestures, sensitivity to various mobile applications, and scalability to user space, are also provided to examine the applicability of MMAuth.
Zhihao Shen 0001, Xi Zhao 0001, Jianhua Zou
IEEE Trans. Inf. Forensics Secur.3
2021 ATPP: A Mobile App Prediction System Based on Deep Marked Temporal Point Processes
abstract
Predicting the app that a user will open next is essential for improving user experience, e.g., app pre-loading. Unlike previous solutions that only predict next app’s ID, this work also predicts the time to open next app. Time prediction is important to avoid loading the next app too early, consuming too much memory and energy on smartphones. To predict next app’s ID and open time jointly, we model app usage as a marked temporal point process (MTPP), whose conditional intensity function can capture the probability of a new app usage event. We develop a novel data-driven MTPP-based app prediction system, named as ATPP (App Temporal Point Process), which adopts a recurrent neural network architecture to learn the MTPP conditional intensity function for app prediction. ATPP adopts a set of techniques to incorporate the unique features of app prediction in our RNN architecture, including learning the correlated usage behavior of different apps by representation learning, the temporal dependency of app usage events by an attention mechanism, and the location-related app usage behavior by a feature extraction and fusion layer. We conduct extensive experiments on a large-scale anonymized app usage dataset from 443 users over 21 days. The experiment results demonstrate that ATPP outperforms the state-of-the-art app prediction method by 6.0% in the prediction accuracy of next app ID and 2.09× reduction in the prediction error of next app open time. A field experiment of 22 users reveals that ATPP can reduce the app loading time by 78%.
Kang Yang 0005, Xi Zhao 0001, Jianhua Zou, Wan Du
DCOSS2
2021 A Hybrid of Deep Reinforcement Learning and Local Search for the Vehicle Routing Problems
abstract
Different variants of the Vehicle Routing Problem (VRP) have been studied for decades. State-of-the-art methods based on local search have been developed for VRPs, while still facing problems of slow running time and poor solution quality in the case of large problem size. To overcome these problems, we first propose a novel deep reinforcement learning (DRL) model, which is composed of an actor, an adaptive critic and a routing simulator. The actor, based on the attention mechanism, is designed to generate routing strategies. The adaptive critic is devised to change the network structure adaptively, in order to accelerate the convergence rate and improve the solution quality during training. The routing simulator is developed to provide graph information and reward with the actor and adaptive cirtic. Then, we combine this DRL model with a local search method to further improve the solution quality. The output of the DRL model can serve as the initial solution for the following local search method, from where the final solution of the VRP is obtained. Tested on three datasets with customer points of 20, 50 and 100 respectively, experimental results demonstrate that the DRL model alone finds better solutions compared to construction algorithms and previous DRL approaches, while enabling a 5- to 40-fold speedup. We also observe that combining the DRL model with various local search methods yields excellent solutions at a superior generation speed, comparing to that of other initial solutions.
Jiuxia Zhao, Minjia Mao, Xi Zhao 0001, Jianhua Zou
IEEE Trans. Intell. Transp. Syst.3
2020 City Metro Network Expansion with Reinforcement Learning
abstract
City metro network expansion, included in the transportation network design, aims to design new lines based on the existing metro network. Existing methods in the field of transportation network design either (i) can hardly formulate this problem efficiently, (ii) depend on expert guidance to produce solutions, or (iii) appeal to problem-specific heuristics which are difficult to design. To address these limitations, we propose a reinforcement learning based method for the city metro network expansion problem. In this method, we formulate the metro line expansion as a Markov decision process (MDP), which characterizes the problem as a process of sequential station selection. Then, we train an actor-critic model to design the next metro line on the basis of the existing metro network. The actor is an encoder-decoder network with an attention mechanism to generate the parameterized policy which is used to select the stations. The critic estimates the expected cumulative reward to assist the training of the actor by reducing training variance. The proposed method does not require expert guidance during design, since the learning procedure only relies on the reward calculation to tune the policy for better station selection. Also, it avoids the difficulty of heuristics designing by the policy formalizing the station selection. Considering origin-destination (OD) trips and social equity, we expand the current metro network in Xi'an, China, based on the real mobility information of 24,770,715 mobile phone users in the whole city. The results demonstrate the advantages of our method compared with existing approaches.
Minjia Mao, Xi Zhao 0001, Jianhua Zou
KDD3
2020 DMM: fast map matching for cellular data
abstract
Map matching for cellular data is to transform a sequence of cell tower locations to a trajectory on a road map. It is an essential processing step for many applications, such as traffic optimization and human mobility analysis. However, most current map matching approaches are based on Hidden Markov Models (HMMs) that have heavy computation overhead to consider high-order cell tower information. This paper presents a fast map matching framework for cellular data, named as DMM, which adopts a recurrent neural network (RNN) to identify the most-likely trajectory of roads given a sequence of cell towers. Once the RNN model is trained, it can process cell tower sequences as making RNN inference, resulting in fast map matching speed. To transform DMM into a practical system, several challenges are addressed by developing a set of techniques, including spatial-aware representation of input cell tower sequences, an encoder-decoder framework for map matching model with variable-length input and output, and a reinforcement learning based model for optimizing the matched outputs. Extensive experiments on a large-scale anonymized cellular dataset reveal that DMM provides high map matching accuracy (precision 80.43% and recall 85.42%) and reduces the average inference time of HMM-based approaches by 46.58×.
Zhihao Shen 0001, Wan Du, Xi Zhao 0001, Jianhua Zou
MobiCom3
2020 mHealth App recommendation based on the prediction of suitable behavior change techniques
Xiaoxin Mao, Xi Zhao 0001
Decis. Support Syst.2
2020 Impact of technostress on productivity from the theoretical perspective of appraisal and coping processes
Xi Zhao 0001, Qihui Xia, Wayne Huang 0001
Inf. Manag.1
2019 Auto Data Augmentation for Testing Set
Wanshun Gao, Xi Zhao 0001
PRCV (2)2
2019 DeepAPP: a deep reinforcement learning framework for mobile application usage prediction
abstract
This paper aims to predict the apps a user will open on her mobile device next. Such an information is essential for many smartphone operations, e.g., app pre-loading and content pre-caching, to save mobile energy. However, it is hard to build an explicit model that accurately depicts the affecting factors and their affecting mechanism of time-varying app usage behavior. This paper presents a deep reinforcement learning framework, named as DeepAPP, which learns a model-free predictive neural network from historical app usage data. Meanwhile, an online updating strategy is designed to adapt the predictive network to the time-varying app usage behavior. To transform DeepAPP into a practical deep reinforcement learning system, several challenges are addressed by developing a context representation method for complex contextual environment, a general agent for overcoming data sparsity and a lightweight personalized agent for minimizing the prediction time. Extensive experiments on a large-scale anonymized app usage dataset reveal that DeepAPP provides high accuracy (precision 70.6% and recall of 62.4%) and reduces the prediction time of the state-of-the-art by 6.58×. A field experiment of 29 participants also demonstrates DeepAPP can effectively reduce time of loading apps.
Zhihao Shen 0001, Kang Yang 0005, Wan Du, Xi Zhao 0001, Jianhua Zou
SenSys4
2019 Utilizing multi-source data in popularity prediction for shop-type recommendation
Xiaoxin Mao, Xi Zhao 0001, Jun Lin 0002, Enrique Herrera-Viedma
Knowl. Based Syst.2
2019 A complementing preference based method for location recommendation with cellular data
Xi Zhao 0001, Jianhua Zou, Enrique Herrera-Viedma
Knowl. Based Syst.2
2019 Improving Shadow Suppression for Illumination Robust Face Recognition
abstract
2D face analysis techniques, such as face landmarking, face recognition and face verification, are reasonably dependent on illumination conditions which are usually uncontrolled and unpredictable in the real world. The current massive data-driven approach, e.g., deep learning-based face recognition, requires a huge amount of labeled training face data that hardly cover the infinite lighting variations that can be encountered in real-life applications. An illumination robust preprocessing method thus remains a very interesting but also a significant challenge in reliable face analysis. In this paper we propose a novel model driven approach to improve lighting normalization of face images. Specifically, we propose to build the underlying reflectance model which characterizes interactions between skin surface, lighting source and camera sensor, and elaborate the formation of face color appearance. The proposed illumination processing pipeline enables generation of the Chromaticity Intrinsic Image (CII) in a log chromaticity space which is robust to illumination variations. Moreover, as an advantage over most prevailing methods, a photo-realistic color face image is subsequently reconstructed, which eliminates a wide variety of shadows whilst retaining the color information and identity details. Experimental results under different scenarios and using various face databases show the effectiveness of the proposed approach in dealing with lighting variations, including both soft and hard shadows, in face recognition.
Wuming Zhang, Xi Zhao 0001, Jean-Marie Morvan, Liming Chen 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 A new aspect on P2P online lending default prediction using meta-level phone usage data in China
Xi Zhao 0001
Decis. Support Syst.2
2018 CoC: A Unified Distributed Ledger Based Supply Chain Management System
Zhimin Gao, Lei Xu 0012, Lin Chen 0009, Xi Zhao 0001, Yang Lu 0010, Larry Shi
J. Comput. Sci. Technol.4
2017 3D-2D face recognition with pose and illumination normalization
Ioannis A. Kakadiaris, George Toderici, Georgios Evangelopoulos, Georgios Passalis, Dat Chu, Xi Zhao 0001, Shishir Shah 0001, Theoharis Theoharis
Comput. Vis. Image Underst.6
2016 Automatic 2.5-D Facial Landmarking and Emotion Annotation for Social Interaction Assistance
abstract
People with low vision, Alzheimer's disease, and autism spectrum disorder experience difficulties in perceiving or interpreting facial expression of emotion in their social lives. Though automatic facial expression recognition (FER) methods on 2-D videos have been extensively investigated, their performance was constrained by challenges in head pose and lighting conditions. The shape information in 3-D facial data can reduce or even overcome these challenges. However, high expenses of 3-D cameras prevent their widespread use. Fortunately, 2.5-D facial data from emerging portable RGB-D cameras provide a good balance for this dilemma. In this paper, we propose an automatic emotion annotation solution on 2.5-D facial data collected from RGB-D cameras. The solution consists of a facial landmarking method and a FER method. Specifically, we propose building a deformable partial face model and fit the model to a 2.5-D face for localizing facial landmarks automatically. In FER, a novel action unit (AU) space-based FER method has been proposed. Facial features are extracted using landmarks and further represented as coordinates in the AU space, which are classified into facial expressions. Evaluated on three publicly accessible facial databases, namely EURECOM, FRGC, and Bosphorus databases, the proposed facial landmarking and expression recognition methods have achieved satisfactory results. Possible real-world applications using our algorithms have also been discussed.
Xi Zhao 0001, Jianhua Zou, Huibin Li 0001, Emmanuel Dellandréa, Ioannis A. Kakadiaris, Liming Chen 0002
IEEE Trans. Cybern.1
2015 An efficient multimodal 2D + 3D feature-based approach to automatic facial expression recognition
Huibin Li 0001, Huaxiong Ding, Di Huang 0001, Yunhong Wang 0001, Xi Zhao 0001, Jean-Marie Morvan, Liming Chen 0002
Comput. Vis. Image Underst.5
2014 Fault resilient physical neural networks on a single chip
abstract
Device scaling engineering is facing major challenges in producing reliable transistors for future electronic technologies. With shrinking device sizes, the total circuit sensitivity to both permanent and transient faults has increased significantly. Research for fault tolerant processors has primarily focused on the conventional processor architectures. Neural network computing has been employed to solve a wide range of problems. This paper presents a design and implementation of a physical neural network that is resilient to permanent hardware faults. To achieve scalability, it uses tiled neuron clusters where neuron outputs are efficiently forwarded to the target neurons using source based spanning tree routing. To achieve fault resilience in the face of increasing number of permanent hardware failures, the design pro-actively preserves neural network computing performance by selectively replicating performance critical neurons. Furthermore, the paper presents a spanning tree recovery solution that mitigates disruption to distribution of neuron outputs caused by failed neuron clusters. The proposed neuron cluster design is implemented in Verilog. We studied the fault resilience performance of the described design using a RBM neural network trained for classifying handwritten digit images. Results demonstrate that our approach can achieve improved fault resilience performance by replicating only 5% most important neurons.
Larry Shi, Yuanfeng Wen, Ziyi Liu 0002, Xi Zhao 0001, Dainis Boumber, Ricardo Vilalta, Lei Xu 0012
CASES4
2014 Minimizing Illumination Differences for 3D to 2D Face Recognition Using Lighting Maps
abstract
Asymmetric 3D to 2D face recognition has gained attention from the research community since the real-world application of 3D to 3D recognition is limited by the unavailability of inexpensive 3D data acquisition equipment. A 3D to 2D face recognition system explicitly relies on 3D facial data to account for uncontrolled image conditions related to head pose or illumination. We build upon such a system, which matches relit gallery textures with pose-normalized probe images, using the gallery facial meshes. The relighting process, however, is based on an assumption of indoor lighting conditions and limits recognition performance on outdoor images. In this paper, we propose a novel method for minimizing illumination difference by unlighting a 3D face texture via albedo estimation using lighting maps. The algorithm is evaluated on challenging databases (UHDB30, UHDB11, FRGC v2.0) with drastic lighting and pose variations. The experimental results demonstrate the robustness of our method for estimating the albedo from both indoor and outdoor captured images, and the effectiveness and efficiency for illumination normalization in face recognition.
Xi Zhao 0001, Georgios Evangelopoulos, Dat Chu, Shishir Shah 0001, Ioannis A. Kakadiaris
IEEE Trans. Cybern.1
2014 Mobile User Authentication Using Statistical Touch Dynamics Images
abstract
Behavioral biometrics have recently begun to gain attention for mobile user authentication. The feasibility of touch gestures as a novel modality for behavioral biometrics has been investigated. In this paper, we propose applying a statistical touch dynamics image (aka statistical feature model) trained from graphic touch gesture features to retain discriminative power for user authentication while significantly reducing computational time during online authentication. Systematic evaluation and comparisons with state-of-the-art methods have been performed on touch gesture data sets. Implemented as an Android App, the usability and effectiveness of the proposed method have also been evaluated.
Xi Zhao 0001, Tao Feng 0011, Larry Shi, Ioannis A. Kakadiaris
IEEE Trans. Inf. Forensics Secur.1
2013 Back to the Future: Using Magnetic Tapes in Cloud Based Storage Infrastructures
Varun Prakash, Xi Zhao 0001, Yuanfeng Wen, Larry Shi
Middleware2
2013 A unified probabilistic framework for automatic 3D facial expression analysis based on a Bayesian belief inference and statistical feature models
Xi Zhao 0001, Emmanuel Dellandréa, Jianhua Zou, Liming Chen 0002
Image Vis. Comput.1
2012 3D/4D facial expression analysis: An advanced annotated face model approach
Tianhong Fang, Xi Zhao 0001, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris
Image Vis. Comput.2
2011 3D facial expression recognition: A perspective on promises and challenges
abstract
This survey focuses on discrete expression classification and facial action unit recognition performed using 3D face data, possibly including a corresponding 2D texture image. Research trends to date are summarized and the limitations of current methods are discussed. The challenges towards the development of more accurate and automated 3D facial expression recognition methods are identified. We also call for standardized experimental protocols in order to draw fair and meaningful comparisons between different systems.
Tianhong Fang, Xi Zhao 0001, Omar Ocegueda, Shishir Shah 0001, Ioannis A. Kakadiaris
FG2
2011 Accurate Landmarking of Three-Dimensional Facial Data in the Presence of Facial Expressions and Occlusions Using a Three-Dimensional Statistical Facial Feature Model
abstract
Three-dimensional face landmarking aims at automatically localizing facial landmarks and has a wide range of applications (e.g., face recognition, face tracking, and facial expression analysis). Existing methods assume neutral facial expressions and unoccluded faces. In this paper, we propose a general learning-based framework for reliable landmark localization on 3-D facial data under challenging conditions (i.e., facial expressions and occlusions). Our approach relies on a statistical model, called 3-D statistical facial feature model, which learns both the global variations in configurational relationships between landmarks and the local variations of texture and geometry around each landmark. Based on this model, we further propose an occlusion classifier and a fitting algorithm. Results from experiments on three publicly available 3-D face databases (FRGC, BU-3-DFE, and Bosphorus) demonstrate the effectiveness of our approach, in terms of landmarking accuracy and robustness, in the presence of expressions and occlusions.
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002, Ioannis A. Kakadiaris
IEEE Trans. Syst. Man Cybern. Part B1
2010 Automatic 3D Facial Expression Recognition Based on a Bayesian Belief Net and a Statistical Facial Feature Model
abstract
Automatic facial expression recognition on 3D face data is still a challenging problem. In this paper we propose a novel approach to perform expression recognition automatically and flexibly by combining a Bayesian Belief Net (BBN) and Statistical facial feature models (SFAM). A novel BBN is designed for the specific problem with our proposed parameter computing method. By learning global variations in face landmark configuration (morphology) and local ones in terms of texture and shape around landmarks, morphable Statistic Facial feature Model (SFAM) allows not only to perform an automatic landmarking but also to compute the belief to feed the BBN. Tested on the public 3D face expression database BU-3DFE, our automatic approach allows to recognize expressions successfully, reaching an average recognition rate over 82%.
Xi Zhao 0001, Di Huang 0001, Emmanuel Dellandréa, Liming Chen 0002
ICPR1
2009 A 3D Statistical Facial Feature Model and Its Application on Locating Facial Landmarks
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002
ACIVS1
2009 A People Counting System Based on Face Detection and Tracking in a Video
abstract
Vision-based people counting systems have wide potential applications including video surveillance and public resources management. Most works in the literature rely on moving object detection and tracking, assuming that all moving objects are people. In this paper, we present our people counting approach based on face detection, tracking and trajectory classification. While we have used a standard face detector, we achieve face tracking combining a new scale invariant Kalman filter with kernel based tracking algorithm. From each potential face trajectory an angle histogram of neighboring points is then extracted. Finally, an Earth Mover's Distance-based K-NN classification discriminates true face trajectories from the false ones. Experimented on a video dataset of more than 160 potential people trajectories, our approach displays an accuracy rate up to 93%.
Xi Zhao 0001, Emmanuel Dellandréa, Liming Chen 0002
AVSS1
2009 Precise 2.5D facial landmarking via an analysis by synthesis approach
abstract
3D face landmarking aims at automatic localization of 3D facial features and has a wide range of applications, including face recognition, face tracking, facial expression analysis. Methods so far developed for pure 2D texture images were shown sensitive to lighting condition changes. In this paper, we present a statistical model-based technique for accurate 3D face landmarking, thus using an ¿analysis by synthesis¿ approach. Our model learns from a training set both variations of global face shapes as well as the local ones in terms of scale-free texture and range patches around each landmark. Given a shape instance, local regions of a new face can be approximated by synthesizing texture and range instances using respectively the texture and range models. By optimizing an objective function describing the similarity of the new face and instances, we can optimize the best shape in order to locate the landmarks. Experimented on more than 1860 face models from FRGC datasets, our method achieves an average of locating errors less than 7 mm for 15 feature points. Compared with a curvature analysis-based method also developed within our team, this learning-based method enables localization of more facial landmarks with a general better accuracy at the cost of a learning step.
Xi Zhao 0001, Przemyslaw Szeptycki, Emmanuel Dellandréa, Liming Chen 0002
WACV1