Neng Gao

dblp:53/11537 · DBLP profile ↗
← Back
83ranked-venue papers
0as first author
25since 2021 · last 2026
0000-0002-0870-5692ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 36 · 10 since 2021Security and privacy · 17 · 1 since 2021Databases, data management, data science and information retrieval · 10 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 7 since 2021Computer networks · 6Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 NumCoKE: Ordinal-Aware Numerical Reasoning over Knowledge Graphs with Mixture-of-Experts and Contrastive Learning
abstract
Knowledge graphs (KGs) serve as a vital backbone for a wide range of AI applications, including natural language understanding and recommendation. A promising yet underexplored direction is numerical reasoning over KGs, which involves inferring new facts by leveraging not only symbolic triples but also numerical attribute values (e.g., length, weight). However, existing methods fall short in two key aspects: (1) Incomplete semantic integration: Most models struggle to jointly encode entities, relations, and numerical attributes in a unified representation space, limiting their ability to extract relation-aware semantics from numeric information. (2) Ordinal indistinguishability: Due to subtle differences between close values and sampling imbalance, models often fail to capture fine-grained ordinal relationships (e.g., longer, heavier), especially in the presence of hard negatives. To address these challenges, we propose NumCoKE—a numerical reasoning framework for KGs based on Mixture-of-Experts and Ordinal Contrastive Embedding. To overcome (C1), we introduce a Mixture-of-Experts Knowledge-Aware (MoEKA) encoder that jointly aligns symbolic and numeric components into a shared semantic space, while dynamically routing attribute features to relation-specific experts. To handle (C2), we propose Ordinal Knowledge Contrastive Learning (OKCL), which constructs ordinal-aware positive and negative samples using prior knowledge, enabling the model to better discriminate subtle semantic shifts. Extensive experiments on three public KG benchmarks demonstrate that NumCoKE consistently outperforms competitive baselines across diverse attribute distributions, validating its superiority in both semantic integration and ordinal reasoning.
Ming Yin 0016, Zongsheng Cao, Qiqing Xia, Chenyang Tu, Neng Gao
AAAI5
2025 Confront Insider Threat: Precise Anomaly Detection in Behavior Logs Based on LLM Fine-Tuning
abstract
Anomaly-based detection is effective against evolving insider threats but still suffers from low precision. Current data processing can result in information loss, and models often struggle to distinguish between benign anomalies and actual threats. Both issues hinder precise detection. To address these issues, we propose a precise anomaly detection solution for behavior logs based on Large Language Model (LLM) fine-tuning. By representing user behavior in natural language, we minimize information loss. We fine-tune the LLM with a user behavior pattern contrastive task for anomaly detection, using a two-stage strategy: first learning general behavior patterns, then refining with user-specific data to improve differentiation between benign anomalies and threats. We also implement a fine-grained threat tracing mechanism to provide behavior-level audit trails. To the best of our knowledge, our solution is the first to apply LLM fine-tuning in insider threat detection, achieving an F1 score of 0.8941 on the CERT v6.2 dataset, surpassing all baselines.
Neng Gao
COLING3
2025 NeuRAG: Retrieval-Augmented Generation based on Dynamic Neural Matching
Neng Gao
ICIC (12)3
2025 Towards effective privacy preservation in federated learning with automatic gradient clipping and gradient transformation perturbation
abstract
Abstract Differential privacy can effectively help federated learning resist privacy attacks from various parties. However, existing approaches that use differential privacy for privacy protection greatly decrease the model performance of federated learning, especially in scenarios with complex model structures and large parameters. In this paper, we propose a novel privacy preservation scheme for federated learning that combines automatic gradient clipping and gradient transformation perturbation. Our approach primarily reduces the impact of differential privacy on federated learning from two aspects. Firstly, we efficiently control the gradient sensitivity by using automatic gradient clipping instead of traditional threshold clipping. Secondly, we utilize the space transformation technique to alleviate the dramatic accuracy degradation of the model caused by the insertion noise. Extensive experiments on various benchmark datasets demonstrate that our approach achieves a good trade-off between data privacy and effectiveness under the same privacy budget.
Chuanyin Wang, Neng Gao
Comput. J.3
2024 Enhancing Document-Level Event Extraction via Structure-Aware Heterogeneous Graph with Multi-Granularity Subsentences
abstract
Document-level Event Extraction aims to identify events from an entire article. It is quite a challenging task because event arguments scatter across several sentences and multiple events in a document may have influence on each other. Previous methods, however, did not take advantage of document structures that have been proved to be effective for sentence-level event extraction. In this work, we propose a structure-aware heterogeneous graph with subsentences for document-level event extraction. Firstly, we build a syntactic graph to capture long-range dependencies between cross-sentence event arguments. Then, multi-granularity sub-sentences are added into the graph to acquire fine-grained understanding. Finally, a global memory stores extracted events so that interactions among multiple events can be captured. Extensive experiments demonstrate that our model outperforms the state of the art models on a widely used large-scale document-level event extraction dataset.
Yuhan Liu 0012, Neng Gao, Zhe Kong
ICASSP2
2024 Elevating Privacy in Federated Learning: An Efficient Approach with SAM Optimization for Personalized Local Models
Chuanyin Wang, Neng Gao
ICIC (8)3
2024 Enhancing Generative Generalized Zero Shot Learning via Multi-Space Constraints and Adaptive Integration
Zhe Kong, Neng Gao, Yuhan Liu 0012
MMM (1)2
2024 BRITD: behavior rhythm insider threat detection with time awareness and user adaptation
abstract
Abstract Researchers usually detect insider threats by analyzing user behavior. The time information of user behavior is an important concern in internal threat detection. Existing works on insider threat detection fail to make full use of the time information, which leads to their poor detection performance. In this paper, we propose a novel behavioral feature extraction scheme: we implicitly encode absolute time information in the behavioral feature sequences and use a feature sequence construction method taking covariance into account to make our scheme adaptive to users. We select Stacked Bidirectional LSTM and Feedforward Neural Network to build a deep learning-based insider threat detection model: Behavior Rhythm Insider Threat Detection (BRITD). BRITD is universally applicable to various insider threat scenarios, and it has good insider threat detection performance: it achieves an AUC of 0.9730 and a precision of 0.8072 with the CMU CERT dataset, which exceeds all baselines. Graphical Abstract
Neng Gao, Cunqing Ma
Cybersecur.2
2023 Learning to Select Prototypical Parts for Interpretable Sequential Data Modeling
abstract
Prototype-based interpretability methods provide intuitive explanations of model prediction by comparing samples to a reference set of memorized exemplars or typical representatives in terms of similarity. In the field of sequential data modeling, similarity calculations of prototypes are usually based on encoded representation vectors. However, due to highly recursive functions, there is usually a non-negligible disparity between the prototype-based explanations and the original input. In this work, we propose a Self-Explaining Selective Model (SESM) that uses a linear combination of prototypical concepts to explain its own predictions. The model employs the idea of case-based reasoning by selecting sub-sequences of the input that mostly activate different concepts as prototypical parts, which users can compare to sub-sequences selected from different example inputs to understand model decisions. For better interpretability, we design multiple constraints including diversity, stability, and locality as training objectives. Extensive experiments in different domains demonstrate that our method exhibits promising interpretability and competitive accuracy.
Neng Gao, Cunqing Ma
AAAI2
2023 A Pure Hardware Design and Implementation on FPGA of WireGuard-based VPN Gateway
abstract
In the face of rising dangers to the internal network due to remote cooperation, VPN gateways are an important tool for organisational network administrators, and the appropriate execution of VPN gateway functions is a vital component in safeguarding the internal network. The VPN gateway confronts security risks from the underlying cryptographic algorithm library, the current operating system, and the central processor as the number of attackers grows and attack methods evolve. In this paper, we propose a pure hardware logic VPN gateway to address security threats from the cryptographic algorithm library, operating system, and CPU by independently implementing the WireGuard protocol’s underlying cryptographic algorithm and building the WireGuard protocol’s hardware logic circuit on the FPGA platform. Actual testing on the NetFPGA-1G-CML platform reveals that the system’s network throughput can reach 35Mbps/s, whereas the network throughput of the software’s WireGuard VPN is 23Mbps/s under the same network settings. Simultaneously, the delay statistics of 300 repetitions of data packet encryption were performed. The encryption latency was less than 20 microseconds when the data packet size was the default MTU.098
Jihong Liu, Neng Gao, Chenyang Tu, Yongjuan Sun
CSCWD2
2023 A Multi-view Knowledge Graph Embedding Model Considering Structure and Semantics
abstract
The essence of knowledge representation learning is to embed the knowledge graph into a low-dimensional vector space to make knowledge computable and deductible. Semantic indiscriminate knowledge representation models usually focus more on the scalability on real world knowledge graphs. They assume that the vector representations of entities and relations are consistent in any semantic environment. Semantic discriminate knowledge representation models focus more on precision. They assume that the vector representations should depend on the specific semantic environment. However, both the two kinds only consider knowledge embedding in semantic space, ignoring the rich features of network structure contained between triplet entities. The MulSS model proposed in this paper is a joint embedding learning method across network structure space and semantic space. By synchronizing the Deepwalk network representation learning method into the semantic indiscriminate model TransE, MulSS achieves better performance than TransE and some semantic discriminate knowledge representation models on triplet classification task. This shows that it is of great significance to extend knowledge representation learning from the single semantic space to the network structure and semantic joint space.
Jia Peng, Neng Gao
CSCWD2
2023 Variational Autoencoder with Stochastic Masks to Solve Exposure Bias in Recommendation (S)
abstract
In the recommended scenario, an unobserved interaction may exists in two cases: the user is not interested in the item or the user is not aware of the item at all.This phenomenon leads to a serious exposure bias problem in recommendation system.To solve this problem, we propose Masked Variational Autoencoder (MVAE).Firstly, we predict the missing values in the sparse user-item interaction matrix by matrix completion.Then we randomly mask the elements in the obtained matrix and use them as the input to the variational autoencoder.The decoder can reconstruct the user interaction matrix closer to the true distribution and fully exploit the potential preferences of users in uninteracted items.In particular, we use a combined dual VAE to tackle the exposure bias problem from the user side and the item side respectively.Extensive experiments on three real-world datasets also illustrate the effectiveness of MVAE for solving exposure bias in recommendation.
Yingshuai Kou, Neng Gao
SEKE2
2022 GSDM: A Gated Semantic Discriminating Model for Knowledge Graph Completion
abstract
Knowledge representation learning is an automatic learning technique that can embeds a knowledge graph into a low-dimensional vector space. With use of this, knowledge becomes computable and various intelligent applications can be realized. Traditional semantic discriminating models suggest that the embeddings of entities should depend on the specific semantic environment. We find that the multiple latent information of relations has not been put to use by these models. In this paper, a gated semantic discriminating model (GSDM) is proposed to select useful latent information and neglect useless information according to the specific semantic environment for both entities and relations. Experiments show that GSDM achieves better performance than related state-of-the-art baselines on most indicators. The better trade-off between the discriminate parameter pressure and the model performance has proved the correctness and feasibility of semantic discriminating mechanism to some extent.
Neng Gao, Nan Mu, Yao Dong 0003, Lei Wang 0135, Yuanye He
CSCWD2
2022 Multi-level Fusion of Multi-modal Semantic Embeddings for Zero Shot Learning
abstract
Zero shot learning aims to recognize objects whose instances may not be covered by the training data. To generalize knowledge from seen classes to the novel ones, semantic space is built to embed knowledge from various views into multi-modal semantic embeddings. Existing semantic embeddings neglect the relationships between classes which are essential to transfer knowledge between classes. Moreover, existing zero shot learning models ignore the complementarity between semantic embeddings from different modalities. To tackle these problems, in this work, we resort to graph theory to explicitly model the interdependence between classes and then obtain new modal semantic embeddings. Furthermore, we pioneer to propose a multi-level fusion model to effectively combine knowledge encoded in multi-modal semantic embeddings together. By the virtue of subsequent fusion block, the results of multi-level fusion can be furtherly enriched and fused. Experiments show that our model could achieve promising results on various datasets. Ablation study suggests that our method is well suited for zero shot learning.
Zhe Kong, Xin Wang 0086, Neng Gao, Yuhan Liu 0012, Chenyang Tu
ICMI3
2022 BEFSR: A Multiple Attention-Based Model Considering Bidirectional Entity Information Flows and Few-Shot Relations
abstract
The traditional knowledge representation learning (KRL) models treat each triplet in a knowledge base independently, so they can not make full use of the neighborhood information across triplets. KRL models based on graph attention networks (GAT) can not only capture feature interactions across triplets, but also further distinguish the importance of neighbor entities. Recently, we find that there are two flaws in GAT-based KRL models: (1) Ignoring the bidirectionality of information flows leads to insufficient utilization of entity neighborhood information. When encapsulating the neighborhood information, only the forward information flows flowing into the target entity are considered, but the backward information flows flowing out are neglected. (2) The unified update process for all relations causes the useful information related the few-shot relations to be diluted. We propose a multiple attention-based model considering bidirectional entity information flows and few-shot relations (BEFSR). In our model, a GAT-based attention framework is used to integrate forward information flows and backward information flows of each triplet respectively to capture feature interactions across triplets, and an LSTM-based attention framework is adopted to gradually aggregate the entity-pair information for few-shot relations’ updating. In BEFSR, entities and relations can be updated more appropriately. Experiments demonstrate that BEFSR outperforms state-of-the-art KRL models in knowledge base completion task.
Neng Gao, Fali Wang, Nan Mu, Lei Wang 0135, Yao Dong 0003
ICPR2
2022 BiGNN: A Bilateral-Branch Graph Neural Network to Solve Popularity Bias in Recommendation
abstract
Traditional recommendation methods aim to recom-mend personalized items by analyzing user's history interaction data. They ignore the fact that the data follows a long-tail distribution, which means that a small number of popular items account for most of the interaction records. This phenomenon causes the model to recommend more popular items, resulting in a severe popularity bias. In order to pay more attention to the long-tail items and debias the popular bias, we propose a Bilateral-Branch Graph Neural Network(BiGNN). In the long- tail branch, we construct a separate long-tail sub graph by eliminating the popular items with high degree. When the Graph Neural Network(GNN) aggregates information layer by layer in the subgraph, the receptive field of the single hop becomes larger, which increases the exposure of the long-tail items. Besides, another branch takes the original interaction graph as input to learn the general data distribution and generate the global embeddings of users and items. The two branches use the same GNN structure and share parameters. We employ the point-wise mutual information (PMI) strategy to indicate interaction between users and reconstruct the long-tail sub graph. The two branches are aggregated through an accumulated learning module, which makes the model first learn the conventional patterns and then pay attention to the long-tail data gradually. Extensive experiments on three real-world datasets show that BiGNN evidently outperforms the state-of- the-art methods consistently.
Yingshuai Kou, Neng Gao, Chenyang Tu, Cunqing Ma
ICTAI2
2022 Hierarchical Graph Attention Network with Heterogeneous Tripartite Graph for Numerical Reasoning over Text
abstract
Numerical reasoning machine reading comprehension (MRC) is a challenging natural language understanding task, which requires machines to have the numerical reasoning ability to add, subtract, sort and count over text to predict answers. Previous state-of-the-art models typically use graph neural networks (GNNs) to perform reasoning over heterogeneous graphs containing numbers and entities, to enhance their numerical reasoning ability. However, their constructed graphs contain only number nodes or only consider related entities in the same sentence, ignoring the relationship between entities. To alleviate the issue, in this paper, we propose a hierarchical graph attention network with heterogeneous tripartite graph (HGAT-HTG) for the numerical reasoning MRC task. It introduces additional sentence nodes as the intermediary between number nodes and entity nodes to construct the heterogeneous tripartite graph, then its hierarchical graph attention network (GAT) reasoning module first performs coarse-grained reasoning on the whole graph, then performs iterative fine-grained reasoning on the number-sentence subgraph and entity-sentence subgraph respectively, to enhance its numerical reasoning ability for the task. Experimental results indicate that HGAT-HTG model offers significant and consistent improvements, outperforming previous state-of-the-art QDGAT on DROP dataset.
Shoukang Han, Neng Gao, Yiwei Shan
IJCNN2
2022 RuleSQLova: Improving Text-to-SQL with Logic Rules
abstract
Text-to-SQL aims to map natural language questions to SQL queries. The sketch-based SQLova model combined with execution-guided (EG) decoding strategy has achieved a super-human performance on the WikiSQL dataset. However, through our fine-grained error analysis, we find that SQLova cannot handle well the aggregation operator selection for nu-meric columns, due to the lack of column type information to distinguish between textual and numeric columns. Besides, most predicted value spans in the WHERE clause have the same meaning with the ground truth, but they do not match exactly, leading to unnecessary errors. Therefore we propose RuleSQLova model, which enhances the SQLova base model with logic rules to deal with these two major weaknesses of SQLova. It first incorporates four logic rules into the model to constrain the aggregation operator prediction for numeric columns, using the general framework of iterative rule knowledge distillation. Then it leverages another logic rule for post-processing before EG decoding to ensure the consistency between predicted value spans and values in the table column. Experimental results indicate that RuleSQLova model offers significant and consistent improvements over SQLova, outperforming competitive sketch-based models on the WikiSQL dataset, and our method also brings improvements to the sketch-based models on the Spider dataset.
Shoukang Han, Neng Gao, Yiwei Shan
IJCNN2
2021 Incorporating Attributes Semantics into Knowledge Graph Embeddings
abstract
More and more work has focused on incorporating different kinds of literals into Knowledge Graph to promote the performance of knowledge embedding. These literals contain numeric literals, text literals, image literals and so on. These additional descriptions are connected to the entities through certain attributes. To incorporate numeric literals, some methods combine the embeddings of literals part with the traditional part - embeddings of entities. However, in the construction of literals embeddings, these existing methods consider the differences of these attributes: one dimension represents one attribute. But they ignore semantic meanings of attributes themselves. In this paper, we propose two methods to incorporate attributes semantics into knowledge graph embeddings from two perspectives: LiteralEAN and literalE-AT. They concatenate with the embeddings of numeric literals by different ways. Furthermore, their extension model LiteralE-C is also proposed as having a more comprehensive representation of attributes semantics. In an empirical study over two standard datasets FB15k and FB15k-237, we evaluate our models for link prediction. We demonstrate that they show an effective way to improve LiteralE and achieve state-of-the-art results. In ablation experiments, we find combined models do better than their singular counterparts in most cases.
Neng Gao, Chenyang Tu, Jia Peng
CSCWD2
2021 Learning from Audience Interaction: Multi-Instance Multi-Label Topic Model for Video Shots Annotating
abstract
In recent years, audiences can find their interested TV play or movie videos by labels easily. However, for finding shots with certain semantic content in these videos, it is still a problem to annotate video shots by labels. Some existing approaches train models with annotated shots which cost a lot in labeling manually. Some other methods in solving this kind of task assume that the content of a video is only limited in the labels of the video. They ignore that the labels of a video are too coarse-grained to cover all content of the video. In this paper, we propose a multi-label, multi-instance topic model to annotate video shots by video labels. In a multi-label, multi-instance framework, video shots can be regarded as instances and shot labels are learned from labels in video level which makes the cost of labeling cheaper. On the other hand, our model learns label semantics by controlling the relationship between video labels and shots to solve coarse-grained problem. Furthermore, we also learn keywords for every video. The experiments on a large-scale real-world dataset show that our model outperforms other baseline models substantially.
Zehua Zeng, Neng Gao, Yuanye He
CSCWD2
2021 TEA-RNN: Topic-Enhanced Attentive RNN for Attribute Inference Attacks via User Behaviors
abstract
Obtaining demographic attributes of online users is of great significance for retail marketing, targeted advertisement and many other scenarios. Users' wanderings on various websites and applications contains user preference on different items, and can be leveraged to infer one's private attributes. Existing studies usually focus on manually defined features, relationships in online social networks, or modeling global user preferences. However, attribute inference from the most common behavioral data (e.g., browsing history, shopping cart) is recently overlooked, and still requires further research. In this work, we propose a Topic-Enhanced Attentive Recurrent Neural Network (TEA-RNN) model to capture both local neighborhood-based features (with attentive RNN) and global patterns (with topic model) within user behaviors, and apply multi-task learning mechanism with weighted losses to further leverage the latent relationships within demographics. Experimental results on real-world datasets demonstrates the effectiveness of TEA-RNN by comparing with several commonly used baselines.
Junsha Chen, Neng Gao, Jiameng Bai
CSCWD4
2021 Image-Enhanced Multi-Modal Representation for Local Topic Detection from Social Media
Junsha Chen, Neng Gao, Chenyang Tu
DASFAA (2)2
2021 CMVCG: Non-autoregressive Conditional Masked Live Video Comments Generation Model
abstract
The blooming of live comment videos leads to the need of automatic live video comment generating task. Previous works focus on autoregressive live video comments generation and can only generate comments by giving the first word of the target comment. However, in some scenes, users need to generate comments by their given prompt keywords, which can't be solved by the traditional live video comment generation methods. In this paper, we propose a Transformer based non-autoregressive conditional masked live video comments generation model called CMVCG model. Our model considers not only the visual and textual context of the comments, but also time and color information. To predict the position of the given prompt keywords, we also introduce a keywords position predicting module. By leveraging the conditional masked language model, our model achieves non-autoregressive live video comment generation. Furthermore, we collect and introduce a large-scale real-world live video comment dataset called Bili-22 dataset. We evaluate our model in two live comment datasets and the experiment results present that our model outperforms the state-of-the-art models in most of the metrics.
Zehua Zeng, Chenyang Tu, Neng Gao, Cunqing Ma, Yiwei Shan
IJCNN3
2021 Incorporating Common Knowledge and Specific Entity Linking Knowledge for Machine Reading Comprehension
Shoukang Han, Neng Gao, Yiwei Shan
KSEM2
2021 PLVCG: A Pretraining Based Model for Live Video Comment Generation
Zehua Zeng, Neng Gao, Chenyang Tu
PAKDD (2)2
2020 FGCRec: Fine-Grained Geographical Characteristics Modeling for Point-of-Interest Recommendation
abstract
With the popularity of location-based social networks (LBSNs), Point-of-Interest (POI) recommendation has become an essential location-based service to help people explore novel locations. Although the massive check-in data bring a good opportunity, there are still many challenges in building personalized POI recommender systems based on geographical information. First, current coarse-grained geographical models provide considerably limited improvements on POI recommendations and fail to capture the overall impact of fine-grained geographical characteristics in LBSNs. Second, previous methods such as matrix factorization always give equal weight to each positive example and may not distinguish between their different contributions in learning the objective function. To cope with these challenges, we develop a fine-grained POI recommendation framework that makes full use of the geographical characteristics from both users’ and locations’ perspectives. For capturing the fine-grained geographical influence, we present a unified probability distribution model based on four key geographical characteristics. For mining more contribution information from positive examples, we assign a higher weight to highlight the contribution of a higher check-in frequency by employing a logistic matrix factorization. Finally, experimental results on two real-world datasets demonstrate the effectiveness and superiority of the proposed method.
Yijun Su, Xiang Li 0045, Baoping Liu, Daren Zha, Ji Xiang, Neng Gao
ICC7
2020 Knowledge Graph Embedding Based on Relevance and Inner Sequence of Relations
Jia Peng, Neng Gao, Jun Yuan 0008
ICONIP (4)2
2020 Leveraging Knowledge Context Information to Enhance Personalized Recommendation
Yingshuai Kou, Neng Gao, Chenyang Tu
ICONIP (3)4
2020 PrivRec: User-Centric Differentially Private Collaborative Filtering Using LSH and KD
Neng Gao, Junsha Chen, Chenyang Tu
ICONIP (4)2
2020 A Practical Defense against Attribute Inference Attacks in Session-based Recommendations
abstract
When users in various web and mobile applications enjoy the convenience of recommendation systems, they are vulnerable to attribute inference attacks. The accumulating online behaviors of users (e.g., clicks, searches, ratings) naturally brings out user preferences, and poses an inevitable threat of privacy that adversaries can infer one's private profiles (e.g., gender, sexual orientation, political view) with AI-based algorithms. Existing defense methods assume the existence of a trusted third party, rely on computationally intractable algorithms, or have impact on recommendation utility. These imperfections make them impractical for privacy preservation in real-life scenarios. In this work, we introduce BiasBooster, a practical proactive defense method based on behavior segmentation, to protect user privacy against attribute inference attacks from user behaviors, while retaining recommendation utility with a heuristic recommendation aggregation module. BiasBooster is a user-centric approach from client side, which proactively divides a user's behaviors into weakly related segments and perform them with several dummy identities, then aggregates real-time recommendations for user from different dummy identities. We estimate its effectiveness of preservation on both privacy and recommendation utility through extensive evaluations on two real-world datasets. A Chrome extension is conducted to demonstrate the feasibility of applying BiasBooster in real world. Experimental results show that compared to existing defenses, BiasBooster substantially reduces the averaged accuracy of attribute inference attacks, with minor utility loss of recommendations.
Neng Gao, Junsha Chen
ICWS2
2020 SECL: Separated Embedding and Correlation Learning for Demographic Prediction in Ubiquitous Sensor Scenario
abstract
Knowing exact demographic attributes of users is crucial for human-computer interaction, intelligent marketing and automatic advertising. Ubiquitous sensor devices yield massive volumes of temporal data which hide a lot of valuable demographic information. In this paper, we bridge the gap between sensor data and demographic prediction to obtain real attributes of users from popular sensor devices: pedometer, which is widely used in mobile devices. We propose a novel model named Separated Embedding and Correlation Learning (SECL) for demographic prediction. Specifically, SECL first process the input data with a separated embedding layer to disentangle task-specific features for interference eliminating, and then capture the hidden correlations between different tasks via a correlation learning layer, finally the refined task-specific features are fed into a multi-task prediction layer to predict demographic attributes. Experimental results show impressive performance of our model on a real-world pedometer dataset, which is made publicly available on https://github.com/deepdeed/SECL.
Yiwen Jiang, Neng Gao, Chenyang Tu, Jia Peng
IJCNN3
2020 User Alignment with Jumping Seed Alignment Information Propagation
abstract
User Alignment is to find users belonging to a same real person on different social networks and has become a fundamental task for many sequent applications such as cross-network recommendation systems. When matching users in multiple social networks, existing approaches always know some correctly matched users, which can be called seeds. Then, existing methods strongly depend on the neighboring users of each user to propagate alignment information from seeds and align probable matching users implicitly. However, the completeness and validity of original alignment information among seeds cannot be fully preserved when learning and aligning multiple user spaces. In this paper, we propose a unified framework named Jumping Seed Alignment Information Propagation (JSAIP) to flexibly leverage, for each user, complete and correct alignment information from seeds. Specifically, JSAIP learns a reasonable user space for each social network by preserving enough original network and label information. Then, JSAIP ensures the correct alignment among seeds and shared labels to reduce the diversity between different user spaces. Finally, JSAIP constructs jumping links from seeds to each user in each social network and ultilizes original seed alignment information to enhance or rectify the alignment information propagated from neighbors. Experiments on real world datasets demonstrate the effectiveness of our proposed JSAIP method compared to several state-of-the-art methods.
Xiang Li 0045, Yijun Su, Neng Gao, Ji Xiang, Yuewu Wang
IJCNN3
2020 FGRec: A Fine-Grained Point-of-Interest Recommendation Framework by Capturing Intrinsic Influences
abstract
Point-of-interest (POI) recommendation has become an important service to help users discover attractive locations. A variety of available check-in data make it possible to build a personalized POI recommender system, but the extreme sparsity of check-in data poses a severe challenge for POI recommendation. Recent studies mainly utilize social information, categorical information and/or geographical information to supplement the highly sparse check-in data. However, these studies often apply shallow methods for the extra information and provide considerably limited improvements on POI recommendation. In this paper, we propose a fine-grained POI recommendation framework, called FGRec to capture the intrinsic influences of social, categorical and geographical information on the check-in behaviors of users. First, we study the social influence in depth by exploiting the multi-hop social friends and top-n nearest neighbor friends, not only the direct friends (i.e., 1-hop friends). Second, we investigate the categorical influence by factorizing both user-POI and user-category matrices simultaneously over the same user embedding space, rather than simply using the popularity of POI categories. Third, we explore the geographical influence by integrating two types of distance (i.e., the distance between user homes and POIs and the distance among POIs) into a unified probability distribution over check-in POIs, instead of modeling them separately. Finally, experimental results on two large-scale real-world datasets demonstrate the effectiveness and superiority of the proposed method.
Yijun Su, Jia-Dong Zhang, Xiang Li 0045, Daren Zha, Ji Xiang, Neng Gao
IJCNN7
2020 Binarized Attributed Network Embedding via Neural Networks
Hangyu Xia, Neng Gao, Jia Peng, Jingjie Mo
IJCNN2
2020 Flush-Detector: More Secure API Resistant to Flush-Based Spectre Attacks on ARM Cortex-A9
abstract
ARM series processors are increasingly used in IoT and cloud services because of their high performance and flexibility of hardware design, especially Cortex-A9 MPCore processor. However, they also suffer from various types of security threats, typically such as flush-based cache attacks. Among these attacks, flush-based Spectre attacks(using Flush + Reload for Spectre attacks) represent a serious threat to system. They usually induce the victim to speculatively perform operations that would not occur during the correct program execution, and then leak the victim’s confidential information to the adversary via cache side channel attacks. So far, there is no widely accepted solution to defend against Spectre attacks. The proposed solutions either lead to large performance losses or sacrifice transparency. In this paper, we propose a secure flush operation API named Flush-Detector to mitigate flush-based Spectre attacks. We present the design and implement of Flush-Detector to detect and defend against flush-based Spectre attacks on ARM Cortex-A9 MPCore. The attack experimental results show that Flush-Detector can detect flush-based Spectre attacks in real time and reduce the attack success rate to less than 1%. Moreover, performance test results demonstrate that the time consumption of Flush-Detector API is about 17.7% longer than the original cache flush API.
Cunqing Ma, Jingquan Ge, Neng Gao, Chenyang Tu
ISCC4
2020 A Hardware/Software Collaborative SM4 Implementation Resistant to Side-channel Attacks on ARM-FPGA Embedded SoC
abstract
The SM4 algorithm is the first commercial cryptographic algorithm officially announced in China for wireless local area network products. It is suitable for scenarios that require high real-time performance, such as wireless communication and IoT sensor nodes. It can be seen that the security research of the SM4 algorithm is of great significance to wireless devices in the IoT. Like other symmetric encryption algorithms, the SM4 algorithm faces some security threats, such as side-channel attacks. Among them, cache timing attacks and power/electromagnetic analysis attacks are becoming more and more threatening due to their low execution difficulty and powerful attack capabilities. Most implementations of anti-side channel attacks against the SM4 algorithm can only resist one of above two attacks. However, side-channel leakages associated with above attacks often coexist.Therefore in this paper, we present a hardware/software collaborative SM4 implementation on ARM-FPGA embedded SoC which can resist above two types of attacks simultaneously. It randomly divides the 32 rounds of SM4 encryption into three stages: the beginning software stage, the middle hardware stage, and the final software stage. Besides, we shuffle the order of some independent operations in each round of the software stages and add dummy rounds to the hardware stage. Finally, we conduct above two types of attacks on unprotected software/hardware SM4, shuffled software SM4 and our scheme, then evaluate their performance respectively. The data throughput of our scheme is 0.86 times that of the original software SM4, while the FPGA resource requirements of our scheme are 0.87 times that of the unprotected hardware implementation.
Ping Peng, Cunqing Ma, Jingquan Ge, Neng Gao, Chenyang Tu
ISCC4
2020 TransBidiFilter: Knowledge Embedding Based on a Bidirectional Filter
Neng Gao, Jun Yuan 0008, Lin Zhao 0006, Lei Wang 0135, Sibo Cai
NLPCC (1)2
2020 Multiple Demographic Attributes Prediction in Mobile and Sensor Devices
Yiwen Jiang, Neng Gao, Ji Xiang, Chenyang Tu
PAKDD (1)3
2020 TransMVG: Knowledge Graph Embedding Based on Multiple-Valued Gates
Neng Gao, Jun Yuan 0008, Xin Wang 0086, Lei Wang 0135
WISE (1)2
2020 MACM: How to Reduce the Multi-Round SCA to the Single-Round Attack on the Feistel-SP Networks
abstract
Since the master key length becomes longer and longer in ciphers, an adversary often needs to preform the multi-round side channel analysis (SCA) in order to recover the master key by enough round keys. Traditional multi-round SCA is launched by adaptive manner in practice, which means that the input of each round is calculated in an on-the-fly way based on all round keys of anterior rounds. However, compared to the classical single-round SCA, the multi-round SCA in adaptive manner is severely limited in several practical scenarios, because all round keys of anterior rounds must be properly recovered before the attack against the next round. In this paper, we focus on the Feistel-SP networks, break the interdependency between the alternating measurement and analysis phases, propose a Multi-round non-Adaptive Chosen Message (MACM) approach, which can reduce the multi-round SCA to the single-round attack. In MACM, the set of plaintexts applied to multiple rounds is calculated in an off-line way. We also prove that the revealed round keys by MACM are adequate to recover the master key. Furthermore, we carefully analyze the advantages of MACM regarding to robustness and compatibility. In order to further manifest the validity of MACM, we perform extensive experiments on three typical Feistel-SP ciphers, Camellia, CLEFIA and SM4, the master keys are recovered as expected, and the number of traces in MACM is at least 25% less than that in the adaptive manner.
Chenyang Tu, Zeyi Liu 0002, Neng Gao, Cunqing Ma, Jingquan Ge, Lingchen Zhang
IEEE Trans. Inf. Forensics Secur.3
2020 Tainting-Assisted and Context-Migrated Symbolic Execution of Android Framework for Vulnerability Discovery and Exploit Generation
abstract
Android Application Framework is an integral and foundational part of the Android system. Each of the two billion (as of 2017) Android devices relies on the system services of Android Framework to manage applications and system resources. Given its critical role, a vulnerability in the framework can be exploited to launch large-scale cyber attacks and cause severe harms to user security and privacy. Recently, many vulnerabilities in Android Framework were exposed, showing that it is indeed vulnerable and exploitable. While there is a large body of studies on Android application analysis, research on Android Framework analysis is very limited. In particular, to our knowledge, there is no prior work that investigates how to enable symbolic execution of the framework, an approach that has proven to be very powerful for vulnerability discovery and exploit generation. We design and build the first system, Centaur, that enables symbolic execution of Android Framework. Due to the middleware nature and technical peculiarities of the framework that impinge on the analysis, many unique challenges arise and are addressed in Centaur. The system has been applied to discovering new vulnerability instances, which can be exploited by recently uncovered attacks against the framework, and to generating PoC exploits.
Lannan Luo, Qiang Zeng 0001, Chen Cao 0004, Kai Chen 0012, Jian Liu 0008, Neng Gao, Min Yang 0002, Xinyu Xing 0001, Peng Liu 0005
IEEE Trans. Mob. Comput.7
2019 TransGate: Knowledge Graph Embedding with Shared Gate Structure
abstract
Embedding knowledge graphs (KGs) into continuous vector space is an essential problem in knowledge extraction. Current models continue to improve embedding by focusing on discriminating relation-specific information from entities with increasingly complex feature engineering. We noted that they ignored the inherent relevance between relations and tried to learn unique discriminate parameter set for each relation. Thus, these models potentially suffer from high time complexity and large parameters, preventing them from efficiently applying on real-world KGs. In this paper, we follow the thought of parameter sharing to simultaneously learn more expressive features, reduce parameters and avoid complex feature engineering. Based on gate structure from LSTM, we propose a novel model TransGate and develop shared discriminate mechanism, resulting in almost same space complexity as indiscriminate models. Furthermore, to develop a more effective and scalable model, we reconstruct the gate with weight vectors making our method has comparative time complexity against indiscriminate model. We conduct extensive experiments on link prediction and triplets classification. Experiments show that TransGate not only outperforms state-of-art baselines, but also reduces parameters greatly. For example, TransGate outperforms ConvE and RGCN with 6x and 17x fewer parameters, respectively. These results indicate that parameter sharing is a superior way to further optimize embedding and TransGate finds a better trade-off between complexity and expressivity.
Jun Yuan 0008, Neng Gao, Ji Xiang
AAAI2
2019 More Secure Collaborative APIs Resistant to Flush+Reload and Flush+Flush Attacks on ARMv8-A
abstract
With the popularity of smart devices such as mobile phones and tablets, the security problem of the widely used ARMv8-A processor has received more and more attention. Flush+Reload and Flush+Flush cache attacks have become two of the most important security threats due to their low noise and high resolution. In order to resist Flush+Reload and Flush+Flush attacks, researchers proposed many defense methods. However, these existing methods have various shortcomings. The runtime defense methods using hardware performance counters cannot detect attacks fast enough, effectively detect Flush+Flush or avoid a high false positive rate. Static code analysis schemes are powerless for obfuscation techniques. The approaches of permanently reducing the resolution can only be utilized on browser products and cannot be applied in the system. In this paper, we design two more secure collaborative APIs-flush operation API and high resolution time API-which can resist Flush+Reload and Flush+Flush attacks. When the flush operation API is called, the high resolution time API temporarily reduces its resolution and automatically restores. Moreover, the flush operation API also has the ability to detect and handle suspected Flush+Reload and Flush+Flush attacks. The attack and performance comparison experiments prove that the two APIs we designed are safer and the performance losses are acceptable.
Jingquan Ge, Neng Gao, Chenyang Tu, Ji Xiang, Zeyi Liu 0002
APSEC2
2019 Perceiving Topic Bubbles: Local Topic Detection in Spatio-Temporal Tweet Stream
Junsha Chen, Neng Gao, Chenyang Tu, Daren Zha
DASFAA (2)2
2019 DCAR: Deep Collaborative Autoencoder for Recommendation with Implicit Feedback
Neng Gao, Jia Peng, Jingjie Mo
ICANN (4)2
2019 Dynamic Graph Link Prediction by Semantic Evolution
abstract
Dynamic graph link prediction has attracted increasing attention in various fields such as social networks, paper citation networks and knowledge graphs. Many models have been developed to predict the future graph structure. In this paper, we propose a link prediction model with semantic evolution (LISE), to predict links in a sequence of graph over time. Our approach is based on the discovery of non-random initialization dynamic word embedding which is a kind of method to study semantic evolution. It can help us train node embedding in the same space and introduce temporal context into the embedding training of nodes. Based on node embedding in the same space, LISE can unify historical behavior, graph snapshots structure information and dynamic attributes into a frame. We evaluate our proposed method and various comparing methods on two real-world datasets. The experimental results prove the effectiveness of the link prediction made by LISE model.
Yujing Zhou, Yuanye He, Jingjie Mo, Neng Gao
ICC6
2019 AdapTimer: Hardware/Software Collaborative Timer Resistant to Flush-Based Cache Attacks on ARM-FPGA Embedded SoC
abstract
ARM-FPGA embedded SoCs have been widely used in the fields of drones, embedded and IoT devices due to its high performance and hardware design flexibility. However, ARM-FPGA embedded SoC suffers various types of security threats, one of which is flush-based cache attack. The proposed defense schemes either lead to a high false positive rate or a large performance loss. Due to the importance of high resolution time APIs in the system, schemes that permanently reduce the resolution of time APIs can only be implemented in specific applications such as browsers. Moreover, the method of protecting high resolution timers in software cannot defend against an attacker with root privileges. In this paper, we propose a more secure timer which is a hardware/software co-design on ARM-FPGA embedded SoC. When a software process calls the flush operation, the timer adaptively reduces its resolution and recover after a short period of time. In the case that the flush operation is not called, the impact of the timer on system performance is almost negligible. This hardware/software co-design guarantees the availability of a high resolution time API while defend against attackers with root privileges. The results of the attack experiments show that the success rates of Flush+Reload and flush-based Spectre attacks can be reduced to less than 1% when using the timer. Performance test results show that the timer access latency is 9.5% slower than the fastest PMCCNTR but 5% faster than the global timer of Cortex-A9 MPCore. The modified flush operation API for the design only increases the time consumption by about 12%.
Jingquan Ge, Neng Gao, Chenyang Tu, Ji Xiang, Zeyi Liu 0002
ICCD2
2019 Local Topic Detection Using Word Embedding from Spatio-Temporal Social Media
Junsha Chen, Neng Gao, Chenyang Tu
ICONIP (5)2
2019 Demographic Prediction from Purchase Data Based on Knowledge-Aware Embedding
Yiwen Jiang, Neng Gao, Ji Xiang, Yijun Su
ICONIP (5)3
2019 Aligning Users Across Social Networks by Joint User and Label Consistence Representation
Xiang Li 0045, Yijun Su, Neng Gao, Ji Xiang, Yuewu Wang
ICONIP (2)3
2019 Anchor User Oriented Accordant Embedding for User Identity Linkage
Xiang Li 0045, Yijun Su, Neng Gao, Ji Xiang, Yuewu Wang
ICONIP (5)3
2019 Node-Edge Bilateral Attributed Network Embedding
Jingjie Mo, Neng Gao, Ji Xiang, Daren Zha
ICONIP (5)2
2019 HRec: Heterogeneous Graph Embedding-Based Personalized Point-of-Interest Recommendation
Yijun Su, Xiang Li 0045, Daren Zha, Yiwen Jiang, Ji Xiang, Neng Gao
ICONIP (3)7
2019 SCS: Style and Content Supervision Network for Character Recognition with Unseen Font Style
Yiwen Jiang, Neng Gao, Ji Xiang, Yijun Su, Xiang Li 0045
ICONIP (5)3
2019 STNet: A Style Transformation Network for Deep Image Steganography
Neng Gao, Xin Wang 0086, Ji Xiang, Guanqun Liu 0002
ICONIP (2)2
2019 The Application of Network Based Embedding in Local Topic Detection from Social Media
abstract
Detecting local topic from social media is an important task for many applications, such as local event discovery and activity recommendation. Recent years have witnessed growing interest in utilizing spatio-temporal social media for local topic detection. However, conventional topic models consider keywords as independent items, which suffer great limitations in modeling short texts from social media. Therefore, some studies introduce embedding into topic models to preserve the semantic correlation among keywords of short texts. Nevertheless, due to the lack of rich contexts in social media, the performance of these embedding based topic models still remain unsatisfactory. In order to enrich the contexts of keywords, we propose two network based embedding methods, both of which can generate rich contexts for keywords by random walks and produce coherent keyword embeddings for topic modeling. Besides, processing continuous spatio-temporal information in social media is also very challenging. Most of the existing methods simply split time and location into equal-size units, which fall short in capturing the continuity of spatio-temporal information. To address this issue, we present a hotspot detection algorithm to identify spatial and temporal hotspots, which can address spatio-temporal continuity and alleviate data sparsity. Finally, the experiments show that the performance of our methods has been improved significantly compared to the state-of-the-art methods.
Junsha Chen, Neng Gao, Chenyang Tu
ICTAI2
2019 Personalized Point-of-Interest Recommendation on Ranking with Poisson Factorization
abstract
The increasing prevalence of location-based social networks (LBSNs) poses a wonderful opportunity to build per-sonalized point-of-interest (POI) recommendations, which aim at recommending a top-N ranked list of POIs to users according to their preferences. Although previous studies on collaborative filtering are widely applied for POI recommendation, there are two significant challenges have not been solved perfectly. (1) These approaches cannot effectively and efficiently exploit unobserved feedback and are also unable to learn useful information from it. (2) How to seamlessly integrate multiple types of context information into these models is still under exploration. To cope with the aforementioned challenges, we develop a new Personalized pairwise Ranking Framework based on Poisson Factor factorization (PRFPF) that follows the assumption that users’ preferences for visited POIs are preferred over potential POIs, unvisited POIs are less preferred than potential POIs. The framework PRFPF is composed of two modules: candidate module and ranking module. Specifically, the candidate module is used to generate a series of potential POIs from unvisited POIs by incorporating multiple types of context information (e.g., social and geographical information). The ranking module learns the ultimate order of users’ preference by leveraging the potential POIs. Experimental results evaluated on two large-scale real-world datasets show that our framework outperforms other state-of-the-art approaches in terms of various metrics.
Yijun Su, Xiang Li 0045, Daren Zha, Ji Xiang, Neng Gao
IJCNN6
2019 TransI: Translating Infinite Dimensional Embeddings Based on Trend Smooth Distance
Neng Gao, Lei Wang 0135, Xin Wang 0086
KSEM (1)2
2019 Knowledge Graph Embedding with Order Information of Triplets
Jun Yuan 0008, Neng Gao, Ji Xiang, Chenyang Tu, Jingquan Ge
PAKDD (3)2
2019 HidingGAN: High Capacity Information Hiding with Generative Adversarial Network
abstract
Abstract Image steganography is the technique of hiding secret information within images. It is an important research direction in the security field. Benefitting from the rapid development of deep neural networks, many steganographic algorithms based on deep learning have been proposed. However, two problems remain to be solved in which the most existing methods are limited by small image size and information capacity. In this paper, to address these problems, we propose a high capacity image steganographic model named HidingGAN. The proposed model utilizes a new secret information preprocessing method and Inception‐ResNet block to promote better integration of secret information and image features. Meanwhile, we introduce generative adversarial networks and perceptual loss to maintain the same statistical characteristics of cover images and stego images in the high‐dimensional feature space, thereby improving the undetectability. Through these manners, our model reaches higher imperceptibility, security, and capacity. Experiment results show that our HidingGAN achieves the capacity of 4 bits‐per‐pixel (bpp) at 256 × 256 pixels, improving over the previous best result of 0.4 bpp at 32 × 32 pixels.
Neng Gao, Xin Wang 0086, Ji Xiang, Daren Zha
Comput. Graph. Forum2
2018 Combination of Hardware and Software: An Efficient AES Implementation Resistant to Side-Channel Attacks on All Programmable SoC
Jingquan Ge, Neng Gao, Chenyang Tu, Ji Xiang, Zeyi Liu 0002, Jun Yuan 0008
ESORICS (1)2
2018 CNN-Based Chinese Character Recognition with Skeleton Feature
Yijun Su, Xiang Li 0045, Daren Zha, Weiyu Jiang, Neng Gao, Ji Xiang
ICONIP (5)6
2018 SSteGAN: Self-learning Steganography Based on Generative Adversarial Networks
Neng Gao, Xin Wang 0086, Xuexin Qu
ICONIP (2)2
2018 MultNet: An Efficient Network Representation Learning for Large-Scale Social Relation Extraction
Jun Yuan 0008, Neng Gao, Lei Wang 0135, Zeyi Liu 0002
ICONIP (3)2
2018 Learning from Audience Intelligence: Dynamic Labeled LDA Model for Time-Sync Commented Video Tagging
Zehua Zeng, Neng Gao, Lei Wang 0135, Zeyi Liu 0002
ICONIP (3)3
2018 Translation-Based Attributed Network Embedding
abstract
Attributed network embedding, which aims to map the structural and attribute information into a latent vector space jointly, has attracted a surge of research attention in recent years. However, a vast majority of existing work explores the correlation between node structure and attribute values whereas the attribute type information which can be potentially complementary is ignored. How to effectively model the nodes, attribute types and attribute values as well as their relations in a unified framework is an open yet challenging problem. To this end, we propose a translation-based attributed network embedding method named TransANE. In our approach, the whole attributed network is considered as a coupled network which consists of two components, i.e., node relation network and attribute correlation network. We construct attribute correlation network by the co-occurrence of attribute values. Each node-attribute relation is regarded as an attributional triple, e.g., (Tom, Gender, Male). We introduce knowledge representation method to model the mapping between nodes, attribute types and attribute values. Empirically, experiments on two real-world datasets including node multi-class classification and network visualization are conducted to evaluate the effectiveness of our method TransANE in this paper. Our method achieves significant performance compared with state-of-the-art baselines.
Jingjie Mo, Neng Gao, Yujing Zhou
ICTAI2
2018 CryptMe: Data Leakage Prevention for Unmodified Programs on ARM Devices
Chen Cao 0004, Le Guan, Ning Zhang 0017, Neng Gao, Jingqiang Lin 0001, Bo Luo, Peng Liu 0005, Ji Xiang, Wenjing Lou
RAID4
2018 User Identity Linkage with Accumulated Information from Neighbouring Anchor Links
Xiang Li 0045, Yijun Su, Neng Gao, Ji Xiang
WISE (2)4
2018 NANE: Attributed Network Embedding with Local and Global Information
Jingjie Mo, Neng Gao, Yujing Zhou
WISE (1)2
2017 A Practical Chosen Message Power Analysis Approach Against Ciphers with the Key Whitening Layers
Chenyang Tu, Lingchen Zhang, Zeyi Liu 0002, Neng Gao
ACNS4
2017 System Service Call-oriented Symbolic Execution of Android Framework with Applications to Vulnerability Discovery and Exploit Generation
abstract
Android Application Framework is an integral and foundational part of the Android system. Each of the 1.4 billion Android devices relies on the system services of Android Framework to manage applications and system resources. Given its critical role, a vulnerability in the framework can be exploited to launch large-scale cyber attacks and cause severe harms to user security and privacy. Recently, many vulnerabilities in Android Framework were exposed, showing that it is vulnerable and exploitable. However, most of the existing research has been limited to analyzing Android applications, while there are very few techniques and tools developed for analyzing Android Framework. In particular, to our knowledge, there is no previous work that analyzes the framework through symbolic execution, an approach that has proven to be very powerful for vulnerability discovery and exploit generation. We design and build the first system, Centaur, that enables symbolic execution of Android Framework. Due to some unique characteristics of the framework, such as its middleware nature and extraordinary complexity, many new challenges arise and are tackled in Centaur. In addition, we demonstrate how the system can be applied to discovering new vulnerability instances, which can be exploited by several recently uncovered attacks against the framework, and to generating PoC exploits.
Lannan Luo, Qiang Zeng 0001, Chen Cao 0004, Kai Chen 0012, Jian Liu 0008, Neng Gao, Min Yang 0002, Xinyu Xing 0001, Peng Liu 0005
MobiSys7
2016 Leakage Fingerprints: A Non-negligible Vulnerability in Side-Channel Analysis
abstract
Low-entropy masking schemes and shuffling technique are two common countermeasures against traditional side-channel analysis. Improved Rotating S-box Masking (RSM) is a combination of both countermeasures and is implemented by DPA contest committee to improve the software security level of AES-128. Compared with the original version, improved RSM mainly introduces both the offset and shuffle array as security foundations to counteract the existing attacks. In this paper, we first point out a general vulnerability referred to as "leakage fingerprints" and make use of it to successfully crack the offset array with 100% accuracy, which breaks down the masking countermeasure in the first step. Then, we show that cracking the shuffle array is still feasible but not necessary since several other vulnerabilities in the implementation level can be exploited to bypass the shuffle countermeasure directly. By selectively combining all these vulnerabilities, a dozen of attacks can be put forward, and we perform two of them as examples to verify their effectiveness. Official evaluation results show that, both attacks submitted by us are practical and feasible, and also operate with high efficiency. In terms of two major performance metrics, our best scheme requires 4 traces to reveal the AES master key with 80% Global Success Rate (GSR) and only 2 traces are enough to reduce the Maximum Partial Guessing Entropy (PGE) under 10.
Zeyi Liu 0002, Neng Gao, Chenyang Tu
AsiaCCS2
2016 Novel MITM Attacks on Security Protocols in SDN: A Feasibility Study
Xin Wang 0086, Neng Gao, Lingchen Zhang, Zongbin Liu, Lei Wang 0135
ICICS2
2016 Detecting Side Channel Vulnerabilities in Improved Rotating S-Box Masking Scheme - Presenting Four Non-profiled Attacks
Zeyi Liu 0002, Neng Gao, Chenyang Tu, Zongbin Liu
SAC2
2015 Towards Analyzing the Input Validation Vulnerabilities associated with Android System Services
abstract
Although the input validation vulnerabilities play a critical role in web application security, such vulnerabilities are so far largely neglected in the Android security research community. We found that due to the unique Framework Code layer, Android devices do need specific input validation vulnerability analysis in system services. In this work, we take the first steps to analyze Android specific input validation vulnerabilities. In particular, a) we take the first steps towards measuring the corresponding attack surface and reporting the current input validation status of Android system services. b) We developed a new input validation vulnerability scanner for Android devices. This tool fuzzes all the Android system services by sending requests with malformed arguments to them. Through comprehensive evaluation of Android system with over 90 system services and over 1,900 system service methods, we identified 16 vulnerabilities in Android system services. We have reported all the issues to Google and Google has confirmed them.
Chen Cao 0004, Neng Gao, Peng Liu 0005, Ji Xiang
ACSAC2
2015 QRL: A High Performance Quadruple-Rail Logic for Resisting DPA on FPGA Implementations
Chenyang Tu, Neng Gao, Zeyi Liu 0002, Zongbin Liu
ICICS3
2015 RIKE+ : using revocable identities to support key escrow in public key infrastructures with flexibility
abstract
Public key infrastructures (PKIs) are proposed to provide various security services. Some security services such as confidentiality require key escrow in certain scenarios, whereas some others such as non‐repudiation and authentication usually prohibit key escrow. Moreover, these two conflicting requirements can coexist for one PKI user. The popular solution in which each user has two different certificates and an escrow authority backs up all escrowed private keys faces the problems of efficiency and scalability. In this study, a novel key management infrastructure called RIKE + is proposed to integrate the ‘inherent key escrow’ of identity‐based encryption (IBE) into PKIs. In RIKE+, (the hash value of) a user's PKI certificate also serves as a ‘revocable identity’ to derive the user's IBE public key, and the revocation of this IBE key pair is achieved by the certificate revocation of PKIs. Therefore the certificate binds the user with two key pairs, one of which is escrowed inherently and the other is not. Furthermore, RIKE+ employs chameleon hash to flexibly control the relationship between the certificate and the IBE key pair. In the case of certificate renewal and revocation, chameleon hash enables RIKE+ to manipulate the hash value of the new certificate, so the user's IBE key pair is not unconditionally changed unless it is necessary. RIKE+ is an effective certificate‐based solution compatible with traditional PKIs and can be built on existing X.509 PKIs.
Jingqiang Lin 0001, Wen Tao Zhu, Qiongxiao Wang, Jiwu Jing, Neng Gao
IET Inf. Secur.6
2014 Remotely wiping sensitive data on stolen smartphones
abstract
Smartphones are playing an increasingly important role in personal life and carrying massive private data. Unfortunately, once the smartphones are stolen, all the sensitive information, such as contacts, messages, photos, credit card information and passwords, may fall into the hands of malicious people. In order to protect the private data, remote deletion mechanism is required to allow owners to wipe the sensitive data on the stolen phone remotely. Existing remote deletion techniques rely on the availability of either WiFi for Internet connection or SIM card for cellular network connection; however, these requirements may not be satisfied when the phones are stolen by some sophisticated adversaries. In this paper, we propose a new remote deletion mechanism that allows the phone owner to delete the private data remotely even if the WiFi is disabled and the SIM card is unplugged. The basic idea is to use emergency call mechanisms to establish a communication connection with a service provider to verify the state of the phone and perform remote deletion. We present a case study of our mechanism with the Universal Mobile Telecommunications System (UMTS) network.
Xingjie Yu, Kun Sun 0001, Wen Tao Zhu, Neng Gao, Jiwu Jing
AsiaCCS5
2014 A Progressive Dual-Rail Routing Repair Approach for FPGA Implementation of Crypto Algorithm
Chenyang Tu, Neng Gao, Eduardo de la Torre, Zeyi Liu 0002
ISPEC3
2014 Towards Efficient Update of Access Control Policy for Cryptographic Cloud Storage
Weiyu Jiang, Neng Gao
SecureComm (2)4
2013 An Efficient Reconfigurable II-ONB Modular Multiplier
Liangsheng He, Tongjie Yang, Neng Gao, Zongbin Liu
SecureComm4
2012 RIKE: Using Revocable Identities to Support Key Escrow in PKIs
Jingqiang Lin 0001, Jiwu Jing, Neng Gao
ACNS4
2011 Towards Attack Resilient Social Network Based Threshold Signing
Ji Xiang, Neng Gao
Inscrypt3