Ji Zhang 0001

dblp:86/1953-1 · DBLP profile ↗
← Back
94ranked-venue papers in the field
17as first author
53since 2021 · last 2026
0000-0001-7167-6970ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 31 (10 first)Data Mining & Knowledge Discovery · 28 (4 first)Information Retrieval & Web Search · 14 (2 first)Big Data, Cloud & Distributed Data Systems · 11Knowledge Engineering, Semantic Web & Information Systems · 5 (1 first)Other / Interdisciplinary · 5
YearPublicationVenuePosition
2026 GNN-Based Item Indexing for LLM-Enhanced Recommendation
abstract
Large language models (LLMs) have transformed recommender systems through strong semantic understanding and generalization. However, the design of item identifiers remains a critical bottleneck that directly affects recommendation quality. Traditional metadata-based identifiers introduce length variability and semantic ambiguity, whereas existing collaborative indexing (CID) approaches often neglect item attributes, show limited cross-dataset generalizability, and incur high computational cost at scale. To address these limitations, we propose a Graph Neural Network (GNN)–based item indexing framework with three coordinated innovations. First, we construct attribute-enriched co-occurrence graphs and use a GNN encoder to fuse item features with collaborative signals, yielding semantically informed representations that work well for attribute-rich catalogs. Second, we replace recursive spectral clustering with hierarchical agglomerative clustering on GNN embeddings, enabling direct control of index length via tree depth and reducing hyperparameter tuning across datasets. Third, we exploit localized message passing rather than global eigendecomposition, which provides considerably better runtime efficiency and is amenable to mini-batch training, supporting online index updates as interactions evolve. Across five benchmarks, GID achieves strong average ranking performance, showing larger improvements on sparse and attribute-rich datasets while remaining competitive in dense settings. The framework is robust under both seen and unseen prompt templates, which supports practical LLM-based recommendation. On sequential recommendation, GID improves HR@10 by 7.9% on average over the strongest baseline in each dataset.
Senlin Mao, Ji Zhang 0001, Peng Zhang 0001, Ze Wang 0016, Xiaoyao Zheng, Jia Wang 0009
SIGIR2
2026 PPA++: Preference Prototype-Aware Learning with Large Language Model for Universal Cross-Domain Recommendation
abstract
While user preferences are important to cross-domain recommendation (CDR), existing methods primarily discover preferences under specific, yet possibly redundant, item features. To this end, we first propose a novel Preference Prototype-Aware (PPA) learning method to quantitatively learn user preferences while minimizing disturbances from the source domain. It introduces a mix-encoder and a proto-decoder. On the one hand, the mix-encoder learns better general representations of interacted items and captures the intrinsic relationships between items across different domains. On the other hand, the proto-decoder implements a learnable prototype matching mechanism to quantitatively perceive user preferences, avoiding disturbances caused by item features from the source domain. Moreover, through experiments on PPA, we observe another two issues that affect existing CDR methods’ performance, i.e., the semantic deficiency caused by sparse item categories and the imbalance weights caused by different user-item distributions. Thus, we further propose a LoRA-based extractor and a domain cross-attention module to alleviate the two issues, respectively. The PPA incorporating with new extractor and attention module is called PPA++. Extensive experiments show that PPA++ outperforms the other state-of-the-art counterparts in four different CDR scenarios.
Ji Zhang 0001, Feiyang Xu, Lvying Chen, Bohan Li 0001, Ning Wang 0005, Huawei Tu, Lei Guo 0008, Hongzhi Yin
Data Sci. Eng.2
2026 Enhancing Knowledge Tracing via Breakpoint-Aware Sequence Augmentation
abstract
Knowledge Tracing (KT) aims to predict learners' future responses by modeling historical learning sequences. However, existing works often overlook the severe disruptions in temporal and interaction continuity arising from large time intervals between interactions, which leads to significant prediction errors. Therefore, we identify and validate thisbreakpointeffect through empirical studies, and try to address it via data augmentation (DA). Existing DA methods usually emphasize sequence diversity or similarity while neglecting the breakpoint effect in KT. Hence, we strive to enhance knowledge tracing via breakpoint-aware sequence augmentation and propose Collaborative Learning Path Augmentation (CLPA), a systematic three-phase method to enhance models' representational capacity and alleviate the breakpoint effects. First, identifying boundary interaction pairs that exhibit the temporal gap and response inconsistency. Second, leveraging cross-sequence collaborative information to infer functional pseudo-paths between identified boundaries. Third, reconstructing and inserting pseudo-learning paths into target sequences to restore sequence continuity; moreover, two refinements of the sampling quantity constraint and the distance-based sampling strategy are proposed to ensure augmentation quality. Comprehensive experiments on four real-world datasets validate the effectiveness of CLPA, which significantly enhances the target KT model's representational capacity and alleviates the breakpoint effects.
Chunli Huang, Kenli Li 0001, Ji Zhang 0001, Jie Wu 0001
IEEE Trans. Knowl. Data Eng.4
2026 Horizontal Multi-Party Data Publishing Under Differential Privacy via Weight-Aware Bidirectional Generative Adversarial Networks
Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Lihua Yin, Puning Zhao, Zhiquan Liu 0001, Li Sun 0008, Lei Shi 0030, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.10
2026 Locally Differentially Private Truth Discovery for Sparse Crowdsensing
abstract
Truth discovery has emerged as an effective tool to mitigate data inconsistency in crowdsensing by prioritizing data from high-quality responders. While local differential privacy (LDP) has emerged as a crucial privacy-preserving paradigm, existing studies under LDP rarely explore a worker's participation in specific tasks for sparse scenarios, which may also reveal sensitive information such as individual preferences and behaviors. Existing LDP mechanisms, when applied to truth discovery in sparse settings, may create undesirable dense distributions, provide insufficient privacy protection, and introduce excessive noise, compromising the efficacy of subsequent non-private truth discovery. Additionally, the interplay between noise injection and truth discovery remains insufficiently explored in the current literature. To address these issues, we propose a lOcally differentially private truth diSCovery approach for spArse cRowdsensing, namely OSCAR. The main idea is to use advanced optimization techniques to reconstruct the sparse data distribution and re-formalize truth discovery by considering the statistical characteristics of injected Laplacian noise while protecting the privacy of both the tasks being completed and the corresponding sensory data. Specifically, to address the data density concerns while alleviating noise, we design a randomized response based Bernoulli matrix factorization method BerRR. To recover the sparse structures from densified, perturbed data, we formalize a 0-1 integer programming problem and develop a sparse recovery solving method SpaIE based on implicit enumeration. We further devise a Laplacian-sensitive truth discovery method LapCRH that leverages maximum likelihood estimation to re-formalize truth discovery by measuring differences between noisy values and truths based on the statistical characteristic of Laplacian noise. Our comprehensive theoretical analysis establishes OSCAR's privacy guarantees, utility bounds, and computational complexity. Experimental results show that OSCAR surpasses the state-of-the-arts by at least 30% in accuracy improvement.
Pengfei Zhang 0010, Zhikun Zhang 0001, Yang Cao 0011, Xiang Cheng 0003, Youwen Zhu, Zhiquan Liu 0001, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.7
2025 eBaaS: AIoT-Enabled eBike Battery-Swap as a Service for Last-Mile Delivery
abstract
In China, the number of riders in the on-demand delivery industry has surpassed ten million. Ensuring that these riders earn a decent income can enhance their financial security, reduce poverty, and promote social equity and stability. Due to ease of use, lower-cost maintenance and environmental friendliness, electric bicycles (e-bikes) are the primary mode of transportation for delivery riders. However, these riders frequently encounter depleted batteries due to limited capacity and prolonged charging times, necessitating inconvenient swaps or recharges during deliveries. To address this issue, we propose the e-bike Battery Swap-as-a-Service (eBaaS), an innovative battery-swapping system that leverages an intelligent AIoT network for seamless battery swapping at distributed locations across urban areas. eBaaS integrates edge-cloud collaboration, battery resource allocation, battery anomaly detection, and battery range prediction to minimize downtime and reduce unnecessary mileage. While eBaaS's potential benefits are evident, there has been a lack of robust methods to quantify its impact. Thus, we further developed the eBaaS Impact Evaluation Method (EIEM), the first comprehensive model to address this gap. EIEM analyzes data from approximately 260,000 delivery riders and 5 million riding trajectories. Findings indicate that eBaaS reduces average invalid mileage by 6 km and increases the order volume by an average of over 20% daily per e-bike rider. Meanwhile, the annual electricity savings result in a reduction of 2.74 million kilograms of carbon emissions for 260,000 riders. The eBaaS system is therefore significantly beneficial for environmental conservation and sustainable urban development.
Donghui Ding, Zhao Li 0007, Jiarun Zhang, Xuanwu Liu, Ji Zhang 0001, Yuchen Li 0001, Peng Cai 0001, Jianxun Liu 0001, Guodong Long
WWW5
2025 Few-shot learning-based human pose estimation model
Jia-Ching Ying, Ji Zhang 0001
Inf. Sci.3
2024 Preference Prototype-Aware Learning for Universal Cross-Domain Recommendation
abstract
Cross-domain recommendation (CDR) aims to suggest items from new domains that align with potential user preferences, based on their historical interactions. Existing methods primarily focus on acquiring item representations by discovering user preferences under specific, yet possibly redundant, item features. However, user preferences may be more strongly associated with interacted items at higher semantic levels, rather than specific item features. Consequently, this item feature-focused recommendation approach can easily become suboptimal or even obsolete when conducting CDR with disturbances of these redundant features. In this paper, we propose a novel Preference Prototype-Aware (PPA) learning method to quantitatively learn user preferences while minimizing disturbances from the source domain. The PPA framework consists of two complementary components: a mix-encoder and a preference prototype-aware decoder, forming an end-to-end unified framework suitable for various real-world scenarios. The mix-encoder employs a mix-network to learn better general representations of interacted items and capture the intrinsic relationships between items across different domains. The preference prototype-aware decoder implements a learnable prototype matching mechanism to quantitatively perceive user preferences, which can accurately capture user preferences at a higher semantic level. This decoder can also avoid disturbances caused by item features from the source domain. The experimental results on public benchmark datasets in different scenarios demonstrate the superiority of the proposed PPA learning method compared to state-of-the-art counterparts. PPA excels not only in providing accurate recommendations but also in offering reliable preference prototypes. Our code is available at https://github.com/zyx-nuaa/PPA-for-CDR.
Ji Zhang 0001, Feiyang Xu, Lvying Chen, Bohan Li 0001, Lei Guo 0008, Hongzhi Yin
CIKM2
2024 A Novel Multi-scale Spatiotemporal Graph Neural Network for Epidemic Prediction
Zenghui Xu, Mingzhang Li, Ting Yu 0004, Linlin Hou, Peng Zhang 0001, R. Uday Kiran, Zhao Li 0007, Ji Zhang 0001
DEXA (2)8
2024 Cross-Domain Sequential Recommendation with Temporal Encoding and Projection-Based Learning
Lvying Chen, Ji Zhang 0001, Sujie Yu, Bohan Li 0001
WISE (3)2
2024 Equivariant Diffusion-Based Sequential Hypergraph Neural Networks with Co-attention Fusion for Information Diffusion Prediction
Ji Zhang 0001, Ting Yu 0004, Gaoming Yang
WISE (4)2
2024 Discriminative boundary generation for effective outlier detection
Ji Zhang 0001, Qiliang Liang, Mohamed Jaward Bah, Hongzhou Li, Liang Chang 0003, R. Uday Kiran
Knowl. Inf. Syst.1
2024 Variate Associated Domain Adaptation for Unsupervised Multivariate Time Series Anomaly Detection
abstract
Multivariate Time Series Anomaly Detection (MTS-AD) is crucial for the effective management and maintenance of devices in complex systems, such as server clusters, spacecrafts, and financial systems, and so on. However, upgrade or cross-platform deployment of these devices will introduce the issue of cross-domain distribution shift, which leads to the prototypical problem of domain adaptation for MTS-AD. Compared with general domain adaptation problems, MTS-AD domain adaptation presents two peculiar challenges: (1) the dimensions of data from the source domain and the target domain are usually different, so alignment without losing any information is necessary; and (2) the association between different variates plays a vital role in the MTS-AD task, which is overlooked by traditional domain adaptation approaches. Aiming at addressing the above issues, we propose a Variate Associated Domain Adaptation Method Combined with a Graph Deviation Network (VANDA) for MTS-AD, which includes two major contributions. First, we characterize the intra-domain variate associations of the source domain by a graph deviation network (GDN), which can share parameters across domains without dimension alignment. Second, we propose a sliding similarity to measure the inter-domain variate associations and perform joint training by minimizing the optimal transport distance between source and target data for transferring variate associations across domains. VANDA achieves domain adaptation by transferring both variate associations and GDN parameters from the source domain to the target domain. We construct two pairs of MTS-AD datasets from existing MTS-AD data and combine three domain adaptation strategies with six MTS-AD backbones as the benchmark methods for experimental evaluation and comparison. Extensive experiments demonstrate the effectiveness of our approach, which outperforms the benchmark methods, and significantly improves the AD performance of the target domain by effectively utilizing the source domain knowledge.
Yifan He 0005, Yatao Bian, Bingzhe Wu, Jihong Guan, Ji Zhang 0001, Shuigeng Zhou
ACM Trans. Knowl. Discov. Data6
2024 Attribute Diversity Aware Community Detection on Attributed Graphs Using Three-View Graph Attention Neural Networks
abstract
Community detection is a fundamental yet important task for characterizing and understanding the structure of attributed graphs. Existing methods mainly focus on the structural tightness and attribute similarity among nodes in a community. However, grouping numerous semantically homogeneous nodes will result in information cocoons and thus reduce the robustness of community structure and the efficiency of node collaboration in real-world applications, such as recommendation systems and collaboration networks. Since nodes with closer connections tend to be more similar, finding communities with dense structures and diverse attributes poses great challenges to mining latent relationships between the graph structure and attribute distribution. To our best knowledge, very little research has been conducted to address this challenge. In this article, we propose a novel three-view graph attention neural networks (TvGANN) model to formally address the attribute diversity aware community detection problem. TvGANN reveals correlations between the graph structure and attributes distribution from the perspective of node organization, attribute co-occurrence, and the node-attribute interaction. It effectively captures structural features and attributes distribution by feeding a structural network and an attribute co-occurrence network into graph attention modules through the encoder–decoder framework. It also learns heterogeneous information by feeding a network into a meta-node attention module. Then, it fuzes the three modules and clusters the embedding representations through a Student's t -distribution approach, which iteratively refines the clustering results. The experiments show that our method not only improves the quality in dense community detection but also performs efficiently for attributed graphs.
Yang Zhang 0042, Ting Yu 0004, Shengqiang Chi, Zhen Wang 0037, Yue Gao 0002, Ji Zhang 0001
ACM Trans. Knowl. Discov. Data6
2024 Incremental Maximal Clique Enumeration for Hybrid Edge Changes in Large Dynamic Graphs
abstract
Incremental maximal clique enumeration (IMCE), which maintains maximal cliques in dynamic graphs, is a fundamental problem in graph analysis. A maximal clique has a solid descriptive power of dense structures in graphs. Real-world graph data is often large and dynamic. Studies on IMCE face significant challenges in the efficiency of incremental batch computation and hybrid edge changes. Moreover, with growing graph sizes, new requirements occur on indexing global maximal cliques and obtaining maximal cliques under specific vertex scope constraints. This work presents a new data structure SOMEi to maintain intermediate maximal cliques during construction. SOMEi serves as a space-efficient index to retrieve scope-constrained maximal cliques on the fly. Based on SOMEi, we design a procedure-oriented IMCE algorithm to deal with hybrid edge changes within a unified algorithm framework. In particular, the algorithm is able to process a large batch of edge changes and significantly improve the average processing time of a single edge change through an efficient pruning strategy. Experimental results on real and synthetic graph data demonstrate that the proposed algorithm outperforms all the baselines and achieves good efficiency through pruning.
Ting Yu 0004, Ting Jiang 0006, Mohamed Jaward Bah, Chen Zhao 0019, Hao Huang 0001, Mengchi Liu, Shuigeng Zhou, Zhao Li 0007, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.9
2024 Penalized Flow Hypergraph Local Clustering
abstract
In recent years, hypergraph analysis have attracted increasing attention due to their ability to model complex data correlation, with hypergraph clustering being one of the most important tasks. However, when the scale of hypergraph is large enough, clustering is difficult based on global consistency. Existing flow-based hypergraph local clustering methods have good theoretical cut improvements and runtime guarantees. However, these methods exhibit poor performance when the initial reference node set is small and are prone to causing the output set to shrink into a small subset, resulting in local minima. To address this issue, we propose the Penalized Flow Hypergraph Local Clustering(PFHLC) and provide new conductance guarantees and runtime analyses for our method. First, we use the random walk method to grow the initial seed set, and introduce the random walk information of nodes as penalized flow into the flow-based framework to optimize the output. Second, we propose a generalized objective function containing random walk information, which takes full advantage of the semi-supervised information of the target cluster to protect important nodes. This feature can avoid the local minima of previous flow-based methods. Importantly, our method is strongly-local and can run efficiently on large-scale hypergraphs. We contribute a real-world dataset and the experiments on real-world large-scale datasets show that PFHLC achieves the state-of-the-art significantly.
Yubo Zhang 0006, Chenggang Yan 0001, Zuxing Xuan, Ting Yu 0004, Ji Zhang 0001, Shihui Ying, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.6
2023 CoSaR: Combating Label Noise Using Collaborative Sample Selection and Adversarial Regularization
abstract
Learning with noisy labels is nontrivial for deep learning models. Sample selection is a widely investigated research topic for handling noisy labels. However, most existing methods face challenges such as imprecise selection, a lack of global selection capabilities, and the need for tedious hyperparameter tuning. In this paper, we propose CoSaR (Collaborative Selection and adversarial Regularization ), a twin-networks based model that performs globally adaptive sample selection to tackle label noise. Specifically, the collaborative selection estimates the average distribution distances between predictions and generation labels through the collaboration of two networks to address the bias of the average distribution distances and the manual tuning of hyperparameters. Adversarial regularization is integrated into CoSaR to restrict the network's tendency to fit and memorize noisy labels, thereby enhancing its collaborative selection capability. In addition, we employ a label smoothing regularization and two types of data augmentation to enhance the robustness of the model further. Extensive experiments on both synthetic and real-world noisy datasets demonstrate that the proposed model outperforms baseline methods remarkably, with an accuracy improvement ranging between +0.56% and +15.14%.
Hao Wang 0068, Wei Wang 0011, Panpan Ni, Ji Zhang 0001
CIKM6
2023 An Efficient Embedding Framework for Uncertain Attribute Graph
Ting Jiang 0006, Ting Yu 0004, Xueting Qiao, Ji Zhang 0001
DEXA (2)4
2023 Efficient Cross Dynamic Task Assignment in Spatial Crowdsourcing
abstract
As a novel intelligent sensing paradigm, spatial crowdsourcing has received extensive attention. Task assignment is a key issue in spatial crowdsourcing. In practice, tasks are unevenly distributed in time and space. Accordingly, the problem of cross task assignment attracts growing attention in both industry and academia. Although there has been a research on this problem, it focuses only on maximizing total revenues for inner platforms. Therefore, it can also be improved to bring a multi-win situation for outer workers and task requesters as well as the inner platform. Inspired by this, we first formulate a new cross dynamic task assignment (CDTA) problem by introducing the reputation scores of workers, and prove it to be NP-hard. For the CDTA problem, a hybrid batch-based framework is presented on the basis of a new cross-platform incentive mechanism and a hybrid batch processing strategy, which are efficient in solving the problem of uneven spatial and time distribution of tasks, respectively. After that, a KM-based algorithm and a density-aware greedy algorithm are proposed to gain an accurate assignment result of tasks in each batch and good performance, respectively. Furthermore, the CDTA problem is modeled as a potential game that is proven to have at least a pure Nash Equilibrium theoretically. Last but not least, a game-theoretic approach is developed to maximize the revenues of the inner platform and outer workers at the same time. Extensive experiments on both real and synthetic datasets are conducted to demonstrate the effectiveness and efficiency of the proposed algorithms.
Tianyue Ren, Xu Zhou 0001, Kenli Li 0001, Yunjun Gao, Ji Zhang 0001, Keqin Li 0001
ICDE5
2023 Uncovering Multivariate Structural Dependency for Analyzing Irregularly Sampled Time Series
Zhen Wang 0037, Ting Jiang 0006, Zenghui Xu, Jianliang Gao, Ou Wu 0001, Ke Yan 0001, Ji Zhang 0001
ECML/PKDD (5)7
2023 The Impact on Employability by COVID-19 Pandemic - AI Case Studies
Venkata Bharath Bandi, Xiaohui Tao 0001, Thanveer Shaik, Jianming Yong, Ji Zhang 0001
WISE5
2023 Exploring developments of the AI field from the perspective of methods, datasets, and metrics
Rujing Yao, Yingchun Ye, Ji Zhang 0001, Shuxiao Li, Ou Wu 0001
Inf. Process. Manag.3
2023 A fast approximate method for k-edge connected component detection in graphs with high accuracy
Ting Yu 0004, Mengchi Liu, Zujie Ren, Ji Zhang 0001
Inf. Sci.4
2023 SAGES: Scalable Attributed Graph Embedding With Sampling for Unsupervised Learning
abstract
Unsupervised graph embedding method generates node embeddings to preserve structural and content features in a graph without human labeling. However, most unsupervised graph representation learning methods suffer issues like poor scalability or limited utilization of content/structural relationships, especially on attributed graphs. In this paper, we propose SAGES, a graph sampling based autoencoder framework, which can alleviate these issues. Specifically, we propose a graph sampler considering both structural and content features, in which nodes with greater influence on each other have more chances to be sampled in the same subgraph. In addition, an unbiased Graph Autoencoder (GAE) with structure-level, content-level, and community-level reconstruction loss is built from the properly sampled subgraph each iteration. The time and space complexity analysis is carried out to show the scalability of SAGES. We conducted experiments on three medium-size attributed graphs and three large attributed graphs. Experimental results illustrate that SAGES achieves the competitive performance in unsupervised attributed graph learning on various downstream tasks including node classification, link prediction, and node clustering.
Xiaoru Qu, Jinze Bai, Zhao Li 0007, Ji Zhang 0001, Jun Gao 0003
IEEE Trans. Knowl. Data Eng.5
2023 EGraph: Efficient Concurrent GPU-Based Dynamic Graph Processing
abstract
In many applications of the analysis of dynamic graph, manyTiming iterative Graph Processing(TGP) jobs usually need to be generated for the processing of the corresponding snapshots of the dynamic graph to obtain the results at different points of time. For high throughput of such applications, it is expected to run the TGP jobs on the GPU concurrently. Although many GPU-based systems have been recently developed, for out-of-GPU-memory dynamic graph processing, this concurrent way suffers from significant data access overhead due to a large volume of data transfer between CPU and GPU and the interference between these concurrently running jobs, which eventually incurs low GPU utilization ratio. In this work, we observed that the TGP jobs have strong temporal and spatial similarity when they access different snapshots for their own processing as most parts of the snapshots are the same and only a few parts are changing with time. It creates ideal opportunities for efficient concurrent execution of the TGP jobs by dramatically reducing CPU-GPU graph data transfer cost. Based on this observation, we develop the first GPU-based dynamic graph processing systemEGraph, which can be integrated into the existing out-of-GPU-memory static graph processing systems to enable them to efficiently support concurrent execution of TGP jobs on dynamic graphs with the help of GPU accelerators. Different from the existing approaches, we propose in EGraph an effectiveLoading-Processing-Switching(LPS) execution model. It is able to effectively reduce the overhead of CPU-GPU data transfer and ensures a higher GPU utilization ratio for efficient execution of the TGP jobs by fully utilizing the data access similarity between the TGP jobs. Experimental results show that the existing GPU-accelerated systems achieve performance improvements of 2.3-3.5 times after being integrated with EGraph.
Yu Zhang 0027, Jin Zhao 0003, Fubing Mao, Lin Gu 0002, Xiaofei Liao, Hai Jin 0001, Haikun Liu, Song Guo 0001, Yangqing Zeng, Hang Hu 0018, Chen Li 0078, Ji Zhang 0001
IEEE Trans. Knowl. Data Eng.13
2022 Event Detection from Web Data in Chinese Based on Bi-LSTM with Attention
Zenghui Xu, Hongzhou Li, Yuquan Gan, Jia-Ching Ying, Ting Yu 0004, Ji Zhang 0001
ADMA (1)7
2022 A Deep Learning Framework for Removing Bias from Single-Photon Emission Computerized Tomography
Jia-Ching Ying, Wan-Ju Yang, Ji Zhang 0001, Yu-Ching Ni, Chia-Yu Lin, Fan-Pin Tseng, Xiaohui Tao 0001
ADMA (1)3
2022 A Parallel Framework for Streaming Graphs Computing
abstract
Streaming computation for large graphs on parallel systems faces challenges in task decomposition, data skew, and resource scheduling. In this work, we propose a general parallel streaming framework for the node-centered graph algorithms to improve the computation efficiency. We construct the parallel procedure of the incremental maximal clique enumeration (IMCE) task and accelerate the incremental Candidate Map Constructor (CMC) algorithm through the framework for large-scale streaming graphs. Experimental results on three large real-world graphs show the framework’s positive effect on the algorithm’s execution time.
Ting Jiang 0006, Ting Yu 0004, Zexian Hong, Zujie Ren, Ji Zhang 0001
IEEE Big Data5
2022 Fast Fourier Transform and Ensemble Model to Classify Epileptic EEG Signals
abstract
The analysis of electroencephalogram (EEG) signals can provide valuable insights to the nature of many diseases such as Alzheimer, epilepsy, sleep problems and thus can improve our understating and treatment about them. One of the major EEG signal applications is related to Epilepsy. The main contribution of this work is the proposal of an effective scheme for classifying the EEG signals for the study of epilepsy based on Fast Fourier Transform (FFT) under an ensemble model. The EEG signals are decomposed into frequency bands by using Fast Fourier Transform to extract the key statistical features, which are then fed into an ensemble model to classify the epileptic patients. Three base classifiers – Naïve Bayes, Least Square Support Vector Machine and Neural Networks - are utilized to construct the ensemble framework. The final decision of classification is dependent on the aggregation of the three classifiers decisions. The experimental results demonstrated that the proposed technique is a promising tool in accurately classifying the epileptic EEG signals.
Raid Lafta, Hisham Alshaheen, Xiaohui Tao 0001, Lingling Li 0004, Ji Zhang 0001
IEEE Big Data7
2022 Debiased Learning of Self-Labeled Twitter Data for User Demographic Prediction
abstract
Labeling sufficient data for supervised learning remains an open challenge in social network analysis. An alternative is to collect self-labeled data, i.e. the data labeled by their owners. Emmery et al show that standard models can be trained and perform well on self-labeled data, suggesting the effectiveness of this approach. In this paper, we argue self-labeled data may not be representative of the population. Taking Twitter demographic prediction as an example, we show the popular FastText model standardly trained on self-labeled data does not generalize well on random testing samples. We then present a new learner DeFastText that aims to correct data bias using the kernel means matching technique. In experiment, we show it achieves lower generalization errors than FastText. This research raises an attention of the data bias problem when learning from self-labeled data in social network analysis.
Zhen Wang 0037, Madison Cooley, Yang Zhang 0042, Chao Lan, Ji Zhang 0001
IEEE Big Data5
2022 IDGMS: a One-Stop Graph Mining System for Infectious Diseases
abstract
Data mining in infectious disease pandemic scenarios is a complex giant task involving data from various fields and requirements of real-time and dynamic. In this paper, we propose a graph mining system for the infectious disease pandemic, IDGMS, with one-stop, dynamic, and interactive characteristics. The system has been applied to solve problems from three view scales and performs well. The system is constructed as a loose coupling structure at the front and back ends and can be extended to more graph mining issues. To the best of our knowledge, we are the first graph system especially targeting data mining of infectious diseases.
Zenghui Xu, Ting Yu 0004, Xingyun Hong, Mingzhang Li, Yang Zhang 0042, Zujie Ren, Ji Zhang 0001
IEEE Big Data7
2022 Knowledge Tracing Based on Gated Heterogeneous Graph Convolutional Networks
abstract
The advancement of science and technology provides the possibility of personalized intelligent education. Representation learning of students’ behavior data is challenging because whether time sequences and interactive behaviors or the correlation between knowledge points and students carrying important information. Some researchers propose knowledge tracing to provide ideas for solving this dilemma. However, existing knowledge tracing methods are divided into machine learning and deep learning. Machine learning-based methods require manual feature extraction and a large amount of prior knowledge. Although deep learning-based methods can automatically extract features, most methods either only use the time series information of the data, or use the association between knowledge points. All the methods ignore the association between knowledge points and students. To fill this gap, we propose a Gated Heterogeneous Graph Convolutional Network (GHGCN) model. We utilize the encoder-decoder framework to predict student performance using the representations of nodes, which is learned from heterogeneous convolutional networks and gate recurrent unit. To validate the effectiveness of the proposed GHGCN model, we conduct the experiments on three public datasets: Simulated Data, Assistments 2009, and Assistments 2015. The results indicate that our method can achieve better performance compared with state-of-the-art algorithms.
Yang Zhang 0042, Zhen Wang 0037, Ting Yu 0004, Mingming Lu, Zujie Ren, Ji Zhang 0001
IEEE Big Data6
2022 Skeleton-Based Mutual Action Recognition Using Interactive Skeleton Graph and Joint Attention
Xiangze Jia, Ji Zhang 0001, Zhen Wang 0037, Yonglong Luo, Fulong Chen 0002, Gaoming Yang
DEXA (2)2
2022 Effective and Robust Boundary-Based Outlier Detection Using Generative Adversarial Networks
Qiliang Liang, Ji Zhang 0001, Mohamed Jaward Bah, Hongzhou Li, Liang Chang 0003, R. Uday Kiran
DEXA (2)2
2022 Towards Efficient Discovery of Periodic-Frequent Patterns in Dense Temporal Databases Using Complements
Veena Pamalla, Tarun Sreepada, R. Uday Kiran, Minh-Son Dao, Koji Zettsu, Yutaka Watanobe, Ji Zhang 0001
DEXA (2)7
2022 Graph Decipher: A transparent dual-attention graph neural network to understand the message-passing mechanism for the node classification
abstract
Graph neural networks (GNNs) can be effectively applied to solve many real-world problems across widely diverse fields. Their success is inseparable from the message-passing mechanisms evolving over the years. However, current mechanisms treat all node features equally at the macro-level (node-level), and the optimal aggregation method has not yet been explored. In this paper, we propose a new GNN called Graph Decipher (GD), which transparentizes the message flows of node features from micro-level (feature-level) to global-level and boosts the performance on node classification tasks. Besides, to reduce the computational burden caused by investigating message-passing, only the relevant representative node attributes are extracted by graph feature filters, allowing calculations to be performed in a category-oriented manner. Experiments on 10 node classification data sets show that GD achieves state-of-the-art performance while imposing a substantially lower computational cost. Additionally, since GD has the ability to explore the representative node attributes by category, it can also be applied to imbalanced node classification on multiclass graph data sets.
Teng Huang 0001, Zhen Wang 0037, Poorya Hosseini, Ji Zhang 0001, Chao Liu 0037, Shan Ai
Int. J. Intell. Syst.6
2022 Sparse-Dyn: Sparse dynamic graph multirepresentation learning via event-based sparse temporal attention network
abstract
Dynamic graph neural networks (DGNNs) have been widely used in modeling and representation learning of graph structure data. Current dynamic representation learning focuses on either discrete learning which results in temporal information loss, or continuous learning which involves heavy computation. In this study, we proposed a novel DGNN, sparse dynamic (Sparse-Dyn). It adaptively encodes temporal information into a sequence of patches with an equal amount of temporal-topological structure. Therefore, while avoiding using snapshots which cause information loss, it also achieves a finer time granularity, which is close to what continuous networks could provide. In addition, we also designed a lightweight module, Sparse Temporal Transformer, to compute node representations through structural neighborhoods and temporal dynamics. Since the fully connected attention conjunction is simplified, the computation cost is far lower than the current state-of-the-art. Link prediction experiments are conducted on both continuous and discrete graph data sets. By comparing several state-of-the-art graph embedding baselines, the experimental results demonstrate that Sparse-Dyn has a faster inference speed while having competitive performance.
Ai Shan, Zhen Wang 0037, Ji Zhang 0001, Teng Huang 0001, Chao Liu 0037
Int. J. Intell. Syst.6
2022 Network Public Opinion Detection During the Coronavirus Pandemic: A Short-Text Relational Topic Model
abstract
Online social media provides rich and varied information reflecting the significant concerns of the public during the coronavirus pandemic. Analyzing what the public is concerned with from social media information can support policy-makers to maintain the stability of the social economy and life of the society. In this article, we focus on the detection of the network public opinions during the coronavirus pandemic. We propose a novel Relational Topic Model for Short texts (RTMS) to draw opinion topics from social media data. RTMS exploits the feature of texts in online social media and the opinion propagation patterns among individuals. Moreover, a dynamic version of RTMS (DRTMS) is proposed to capture the evolution of public opinions. Our experiment is conducted on a real-world dataset which includes 67,592 comments from 14,992 users. The results demonstrate that, compared with the benchmark methods, the proposed RTMS and DRTMS models can detect meaningful public opinions by leveraging the feature of social media data. It can also effectively capture the evolution of public concerns during different phases of the coronavirus pandemic.
Yuan-Chun Jiang, Ruicheng Liang, Ji Zhang 0001, Jianshan Sun, Ye-Zheng Liu 0001, Yang Qian 0001
ACM Trans. Knowl. Discov. Data3
2021 VAGA: Towards Accurate and Interpretable Outlier Detection Based on Variational Auto-Encoder and Genetic Algorithm for High-Dimensional Data
abstract
The curse of dimensionality in high-dimensional data makes it difficult to capture the abnormality of data points in full data space. To deal with this problem, we propose an outlier detection model based on Variational Autoencoder and Genetic Algorithm for subspace outlier analysis of high-dimensional data (VAGA). The proposed VAGA model constructs a variational autoencoder (VAE) to preliminarily detect outliers. Then the genetic algorithm (GA) is used to search the abnormal subspace of the outliers obtained by the VAE layer to provide a basis for subspace outlier analysis. The subsequent clustering of the abnormal subspaces help filter out the false positives which are fed back to the VAE layer to adjust network weights. The comparative experiments performed on three public benchmark datasets show that the outlier detection results of the proposed VAGA model are highly interpretable and have better accuracy performance than the state-of-the-art outlier detection methods.
Jiamu Li, Ji Zhang 0001, Jian Wang 0038, Youwen Zhu, Mohamed Jaward Bah, Gaoming Yang, Yuquan Gan
IEEE BigData2
2021 Improving Irregularly Sampled Time Series Learning with Time-Aware Dual-Attention Memory-Augmented Networks
abstract
Irregularly, asynchronously and sparsely sampled multivariate time series (IASS-MTS) are characterized by sparse non-uniform time intervals between successive observations and different sampling rates amongst series. Those properties pose substantial challenges to mainstream machine learning models for learning complicated relations within and across IASS-MTS. This is because that most of the models assume that the time series in question are even, complete (fixed-dimensional features) and synchronous. To address these challenges, we present a novel time-aware Dual-Attention and Memory-Augmented Network (DAMA-Net). The proposed model can leverage both time irregularity, multi-sampling rates and global temporal patterns information inherent in IASS-MTS so as to learn more effective representations for improving prediction performance. Comprehensive experiments on real datasets show that the DAMA-Net outperforms the state-of-the-art methods in multivariate time series classification task.
Zhen Wang 0037, Yang Zhang 0042, Ai Jiang, Ji Zhang 0001, Zhao Li 0007, Jun Gao 0003, Ke Li 0044, Chenhao Lu, Zujie Ren
CIKM4
2021 LSTM Based Sentiment Analysis for Cryptocurrency Prediction
Xin Huang 0005, Wenbin Zhang 0002, Xuejiao Tang, Jayachander Surbiryala, Vasileios Iosifidis, Zhen Liu 0017, Ji Zhang 0001
DASFAA (3)8
2021 Cognitive Visual Commonsense Reasoning Using Dynamic Working Memory
Xuejiao Tang, Xin Huang 0005, Wenbin Zhang 0002, Travers B. Child, Zhen Liu 0017, Ji Zhang 0001
DaWaK7
2021 An Effective Algorithm for Classification of Text with Weak Sequential Relationships
Qiqiang Xu, Ji Zhang 0001, Ting Yu 0004, Wenbin Zhang 0002, Yonglong Luo, Fulong Chen 0002, Zhen Liu 0017
DEXA (2)2
2021 AutoEncoder for Neuroimage
Fan Zhang 0045, Jianxin Zhang 0001, Ahmad Chaddad, Fenghua Guo, Wenbin Zhang 0002, Ji Zhang 0001, Alan C. Evans
DEXA (2)7
2021 Large-scale Fake Click Detection for E-commerce Recommendation Systems
abstract
With the development of e-commerce platforms, e-commerce recommendation systems are playing an increasingly important role for the purpose of product recommendation. As a new attack model against e-commerce recommendation systems, the "Ride Item's Coattails" attack creates fake click information to establish the deceptive correlation between popular products and low-quality products in order to mislead the recommendation system of e-commerce platform to boost the sales of low-quality products. This attack is characterized by high concealment and strong destructiveness, which can cause great damage to e-commerce recommendation systems, and adversely affect the usability of the e-commerce platform and users' shopping experience. It is therefore of great practical significance to study how to quickly and effectively identify the false click information and the corresponding "Ride Item's Coattails" attack to better safeguard e-commerce recommendation systems. At present, there is no previously reported relevant research work conducted specifically for addressing the detection of the "Ride Item's Coattails" attack. In this work, we carried out pioneering work in analyzing and summarizing the characteristics of the false click information produced by attackers on the target products in the "Ride Item's Coattails" attack and designed a set of attack detection techniques suitable for e-commerce recommendation systems. Experimental results on real e-commerce datasets show that our proposed techniques can quickly and effectively detect the large-scale fake click information as well as the associated "Ride Item's Coattails" attack in e-commerce recommendation systems.
Jingdong Li, Zhao Li 0007, Ji Zhang 0001, Xiaoling Wang 0004, Xingjian Lu, Jingren Zhou 0001
ICDE4
2021 AdaBoosting Clusters on Graph Neural Networks
abstract
Graph Neural Networks (GNNs), combining node features and structure information flexibly, have been widely studied and applied in many fields. The growth of graph size and rich features generates a considerable demand for achieving scalability while maintaining good classification performance in the research of GNNs. Graph partition technique, as used in a recent work ClusterGCN, which divides the graph into several sub-graphs, has become an important strategy to achieve the scalability, but the loss of information still affects the results. In this paper, AdClusterGCN is proposed to establish the interaction between graph partition and node classification, in which they can promote each other, and the effectiveness and efficiency of the model can be ensured at the same time. AdClusterGCN combines GNN models trained on a sequence of graph partitions to capture different features, where the current partition is affected using adjusted node/edge weights computed from the results of GNN models on previous partitions. The PageRank and resampling techniques are adopted to keep sufficient attention on important nodes in different models. We implement our method with TensorFlow and experimental studies show that AdClusterGCN achieves state-of-the-art performance on several public benchmarks.
Jun Gao 0003, Zhao Li 0007, Ji Zhang 0001
ICDM4
2021 A Generic Knowledge Based Medical Diagnosis Expert System
abstract
In this paper, we design and implement a generic medical knowledge based system (MKBS) for identifying diseases from several symptoms. In this system, some important aspects like knowledge bases system, knowledge representation, inference engine have been addressed. The system asks users different questions and inference engines will use the certainty factor to prune out low possible solutions. The proposed disease diagnosis system also uses a graphical user interface (GUI) to facilitate users to interact with the expert system. Our expert system is generic and flexible, which can be integrated with any rule bases system in disease diagnosis.
Xin Huang 0005, Xuejiao Tang, Wenbin Zhang 0002, Ji Zhang 0001, Wensheng Gan, Shichao Pei, Zhen Liu 0017, Yiyi Huang
iiWAS4
2021 Live-Streaming Fraud Detection: A Heterogeneous Graph Neural Network Approach
abstract
Live-streaming platforms have recently gained significant popularity by attracting an increasing number of young users and have become a very promising form of online shopping. Similar to the traditional online shopping platforms such as Taobao, live-streaming platforms also suffer from online malicious fraudulent behaviors where many transactions are not genuine. The existing anti-fraud models proposed to recognize fraudulent transactions on traditional online shopping platforms are inapplicable on live-streaming platforms. This is mainly because live-streaming platforms are characterized by a unique type of heterogeneous live-streaming networks where multiple heterogeneous types of nodes such as users, live-streamers, and products are connected with multiple different types of edges associated with edge features. In this paper, we propose a new approach based on a heterogeneous graph neural network for LIve-streaming Fraud dEtection (called LIFE). LIFE designs an innovative heterogeneous graph learning model that fully utilizes various heterogeneous information of shopping transactions, users, streamers, and items from a given live-streaming platform. Moreover, a label propagation algorithm is employed within our LIFE framework to handle the limited number of labeled fraudulent transactions for model training. Extensive experimental results on a large-scale Taobao live-streaming platform demonstrate that the proposed method is superior to the baseline models in terms of fraud detection effectiveness on live-streaming platforms. Furthermore, we conduct a case study to show that the proposed method is able to effectively detect fraud communities for live-streaming e-commerce platforms.
Haishuai Wang, Zhao Li 0007, Peng Zhang 0001, Pengrui Hui, Jian Liao 0001, Ji Zhang 0001, Jiajun Bu
KDD7
2021 Learning Probabilistic Latent Structure for Outlier Detection from Multi-view Data
Zhen Wang 0037, Ji Zhang 0001, Yizheng Chen 0003, Chenhao Lu, Jerry Chun-Wei Lin, Jing Xiao 0005, R. Uday Kiran
PAKDD (1)2
2021 ATJ-Net: Auto-Table-Join Network for Automatic Learning on Relational Databases
abstract
A relational database, consisting of multiple tables, provides heterogeneous information across various entities, widely used in real-world services. This paper studies the supervised learning task on multiple tables, aiming to predict one label column with the help of multiple-tabular data. However, classical ML techniques mainly focus on single-tabular data. Multiple-tabular data refers to many-to-many mapping among joinable attributes and n-ary relations, which cannot be utilized directly by classical ML techniques. Besides, current graph techniques, like heterogeneous information network (HIN) and graph neural networks (GNN), are infeasible to be deployed directly and automatically in a multi-table environment, which limits the learning on databases.
Jinze Bai, Zhao Li 0007, Donghui Ding, Ji Zhang 0001, Jun Gao 0003
WWW5
2021 Combined cause inference: Definition, model and performance
abstract
In recent years, many methods have been developed for discovering causal relationships from observed data. However, as an important kind of causes existing in many causal systems, combined causes (e.g. multi-factor causes consisting of two or more component variables that individually might not be a cause) have not received enough attention. The existing approach includes both individual and combined variables in the causal discovery process using constraint-based methods, can neither distinguish a set of Markov equivalence classes nor identify a combined cause containing one (or more) individual cause(s), therefore can output only some combined causes, instead of all combined causes. In this paper, we first subsume all possible combined causes into three types and give them formal definitions, then extend the additive noise model (ANM) to infer combined causes. We show that if a candidate variable set X w.r.t. a target Y satisfies: (1) allowing ANM for only the forward direction X→Y, and (2) no disturbance variable is contained in X, i.e., removing any component of X will weaken the causal relationship between X and Y, then X forms a combined cause. Based on this finding, we develop an efficient method to discover combined causes. Furthermore, we also conduct extensive experiments to validate the proposed method on both synthetic and real-world data sets.
Hao Zhang 0079, Chuanxu Yan, Shuigeng Zhou, Jihong Guan, Ji Zhang 0001
Inf. Sci.5
2021 ICS-GNN: Lightweight Interactive Community Search via Graph Neural Network
abstract
Searching a community containing a given query vertex in an online social network enjoys wide applications like recommendation, team organization, etc. When applied to real-life networks, the existing approaches face two major limitations. First, they usually take two steps, i.e. , crawling a large part of the network first and then finding the community next, but the entire network is usually too big and most of the data are not interesting to end users. Second, the existing methods utilize hand-crafted rules to measure community membership, while it is very difficult to define effective rules as the communities are flexible for different query vertices. In this paper, we propose an Interactive Community Search method based on Graph Neural Network (shortened by ICS-GNN) to locate the target community over a subgraph collected on the fly from an online network. Specifically, we recast the community membership problem as a vertex classification problem using GNN, which captures similarities between the graph vertices and the query vertex by combining content and structural features seamlessly and flexibly under the guide of users' labeling. We then introduce a k -sized Maximum-GNN-scores (shortened by kMG ) community to describe the target community. We next discover the target community iteratively and interactively. In each iteration, we build a candidate subgraph using the crawled pages with the guide of the query vertex and labeled vertices, infer the vertex scores with a GNN model trained on the subgraph, and discover the kMG community which will be evaluated by end users to acquire more feedback. Besides, two optimization strategies are proposed to combine ranking loss into the GNN model and search more space in the target community location. We conduct the experiments in both offline and online real-life data sets, and demonstrate that ICS-GNN can produce effective communities with low overhead in communication, computation, and user labeling.
Jun Gao 0003, Jiazun Chen, Zhao Li 0007, Ji Zhang 0001
Proc. VLDB Endow.4
2021 TARA-Net: A Fusion Network for Detecting Takeaway Rider Accidents
abstract
In the emerging business of food delivery, rider traffic accidents raise financial cost and social traffic burden. Although there has been much effort on traffic accident forecasting using temporal-spatial prediction models, none of the existing work studies the problem of detecting the takeaway rider accidents based on food delivery trajectory data. In this article, we aim to detect whether a takeaway rider meets an accident on a certain time period based on trajectories of food delivery and riders’ contextual information. The food delivery data has a heterogeneous information structure and carries contextual information such as weather and delivery history, and trajectory data are collected as a spatial-temporal sequence. In this article, we propose a TakeAway Rider Accident detection fusion network TARA-Net to jointly model these heterogeneous and spatial-temporal sequence data. We utilize the residual network to extract basic contextual information features and take advantage of a transformer encoder to capture trajectory features. These embedding features are concatenated into a pyramidal feed-forward neural network. We jointly train the above three components to combine the benefits of spatial-temporal trajectory data and sparse basic contextual data for early detecting traffic accidents. Furthermore, although traffic accidents rarely happen in food delivery, we propose a sampling mechanism to alleviate the imbalance of samples when training the model. We evaluate the model on a transportation mode classification dataset Geolife and a real-world Ele.me dataset with over 3 million riders. The experimental results show that the proposed model is superior to the state-of-the-art.
Yifan He 0005, Zhao Li 0007, Anhui Wang, Peng Zhang 0001, Shuigeng Zhou, Ji Zhang 0001, Ting Yu 0004
ACM Trans. Intell. Syst. Technol.7
2020 Context-aware Adaptive Outlier Detection in Trajectory Data
abstract
With the advent of data mining and business processes automation, outlier detection has evolved into a major problem attracting significant research in relation to several application domains. Further advances in Global Positioning system, tracking of anomalous events based on data enhances effective decision making and pro-active measures to overcome risks and avoid unwarranted outputs. Significant work has been done in trajectory outlier detection although no singular approach fits all the domains. By including position and collective outliers on the same visualizations will enhance understanding of an outlier behavior. As such, we have leveraged Hidden Markov Method for prediction-based point outlier detection and pattern mining to identify points or segments of outliers in trajectory data.
Srinivas Danda, Ji Zhang 0001, Xiaohui Tao 0001, Jerry Chun-Wei Lin, Wenbin Zhang 0002
IEEE BigData2
2020 A Data-driven Human Responsibility Management System
abstract
An ideal safe workplace is described as a place where staffs fulfill responsibilities in a well-organized order, potential hazardous events are being monitored in real-time, as well as the number of accidents and relevant damages are minimized. However, occupational-related death and injury are still increasing and have been highly attended in the last decades due to the lack of comprehensive safety management. A smart safety management system is therefore urgently needed, in which the staffs are instructed to fulfill responsibilities as well as automating risk evaluations and alerting staffs and departments when needed. In this paper, a smart system for safety management in the workplace based on responsibility big data analysis and the internet of things (IoT) are proposed. The real world implementation and assessment demonstrate that the proposed systems have superior accountability performance and improve the responsibility fulfillment through real-time supervision and self-reminder.
Xuejiao Tang, Jiong Qiu, Wenbin Zhang 0002, Vasileios Iosifidis, Zhen Liu 0017, Ji Zhang 0001
IEEE BigData9
2020 Effective Tuple-based Anonymization for Massive Streaming Categorical Data
abstract
In this poster, we propose a novel, effective tuple-based anonymization technique for categorical data over the Internet. By utilizing a new structure, called Candidate Encoding Sequence with Frequency, and a set of new rules for generating such a sequence for each domain value of the categorical data, we can effectively solve the key limitation of the existing methods. Our experimental results demonstrate the superiority of our method against the existing method in terms of the strength of privacy protection.
Qiqiang Xu, Ji Zhang 0001, Zenghui Xu, Yonglong Luo, Fulong Chen 0002, Xiaoyao Zheng, Gaoming Yang
IEEE BigData2
2020 A Block-Level RNN Model for Resume Block Classification
abstract
Resume block classification is the most significant step in resume information extraction. However, the existing algorithms applied to resume block classification are all the general text classification algorithms, which failed to consider the contextual order of each block within a resume. In order to improve the performance of resume block classification, we propose in this paper a block-level bidirectional recurrent neural network model that makes full use of the contextual order relationship among different resume blocks. The experimental results show that the average F1-score value of our model on three 1,400 real resume datasets is 6% to 9% higher than the existing methods.
Qiqiang Xu, Ji Zhang 0001, Youwen Zhu, Bohan Li 0001, Donghai Guan, Xin Wang 0030
IEEE BigData2
2020 Category-aware Graph Neural Networks for Improving E-commerce Review Helpfulness Prediction
abstract
Helpful reviews in e-commerce sites can help customers acquire detailed information about a certain item, thus affecting customers' buying decisions. Predicting review helpfulness automatically in Taobao is an essential but challenging task for two reasons: (1) whether a review is helpful not only relies on its text, but also is related with the corresponding item and the user who posts the review, (2) the criteria of classifying review helpfulness under different items are not the same. To handle these two challenges, we propose CA-GNN (Category Aware Graph Neural Networks), which uses graph neural networks (GNNs) to identify helpful reviews in a multi-task manner --- we employ GNNs with one shared and many item-specific graph convolutions to learn the common features and each item's specific criterion for classifying reviews simultaneously. To reduce the number of parameters in CA-GNN and further boost its performance, we partition the items into several clusters according to their category information, such that items in one cluster share a common graph convolution.We conduct solid experiments on two public datasets and demonstrate that CA-GNN outperforms existing methods by up to 10.9% in AUC. We also deployed our system in Taobao with online A/B Test and verify that CA-GNN still outperforms the baseline system in most cases.
Xiaoru Qu, Zhao Li 0007, Pengcheng Zou, Junxiao Jiang, Rong Xiao 0005, Ji Zhang 0001, Jun Gao 0003
CIKM9
2020 Recommendation on Heterogeneous Information Network with Type-Sensitive Sampling
Jinze Bai, Zhao Li 0007, Donghui Ding, Pengrui Hui, Jun Gao 0003, Ji Zhang 0001, Zujie Ren
DASFAA (3)8
2020 Predicting Workplace Injuries Using Machine Learning Algorithms
abstract
Predicting workplace injury using automated techniques opens newer possibilities in evidence-based research. This paper presents our preliminary research in a PhD project in predicting workplace incidents using machine learning algorithms. The analysis on the model performance using several mainstream machine learning algorithms including random forest, k-nearest neighbor and decision tree indicated that the general performance of the decision tree model was found to be statistically higher than that of the other two algorithms.
Divya Sukumar, Ji Zhang 0001, Xiaohui Tao 0001, Xin Wang 0030, Wenbin Zhang 0002
DSAA2
2020 High average-utility sequential pattern mining based on uncertain databases
Jerry Chun-Wei Lin, Ting Li 0011, Matin Pirouz, Ji Zhang 0001, Philippe Fournier-Viger
Knowl. Inf. Syst.4
2019 Method and Dataset Mining in Scientific Papers
abstract
Literature analysis facilitates researchers better understanding the development of science and technology. The conventional literature analysis focuses on the topics, authors, abstracts, keywords, references, etc., and rarely pays attention to the content of papers. In the field of machine learning, the involved methods (M) and datasets (D) are key information in papers. The extraction and mining of M and D are useful for discipline analysis and algorithm recommendation. In this paper, we propose a novel entity recognition model, called MDER, and constructe datasets from the papers of the PAKDD conferences (2009-2019). Some preliminary experiments are conducted to assess the extraction performance and the mining results are visualized.
Rujing Yao, Linlin Hou, Yingchun Ye, Ji Zhang 0001, Jian Wu 0006
IEEE BigData4
2019 Learning Relational Fractals for Deep Knowledge Graph Embedding in Online Social Networks
Ji Zhang 0001, Leonard Tan, Xiaohui Tao 0001, Dianwei Wang, Jia-Ching Ying, Xin Wang 0030
WISE1
2019 The 1st International Workshop on Context-Aware Recommendation Systems with Big Data Analytics (CARS-BDA)
abstract
With the explosive growth of online service platforms, increasing number of people and enterprises are doing everything online. In order for organizations, governments, and individuals to understand their users, and promote their products or services, it is necessary for them to analyse big data and recommend the media or online services in real time. Effective recommendation of items of interest to consumers has become critical for enterprises in domains such as retail, e-commerce, and online media. Driven by the business successes, academic research in this field has also been active for many years. Through many scientific breakthroughs have been achieved, there are still tremendous challenges in developing effective and scalable recommendation systems for real-world industrial applications. Existing solutions focus on recommending items based on pre-set contexts, such as time, location, weather etc. The big data sizes and complex contextual information add further challenges to the deployment of advanced recommender systems. This workshop aims to bring together researchers with wide-ranging backgrounds to identify important research questions, to exchange ideas from different research disciplines, and, more generally, to facilitate discussion and innovation in the area of context-aware recommender systems and big data analytics.
Xiangmin Zhou, Ji Zhang 0001, Yanchun Zhang
WSDM2
2018 A Genetic Algorithm Based Technique for Outlier Detection with Fast Convergence
Ji Zhang 0001, Zewen Hu, Hongzhou Li, Liang Chang 0003, Youwen Zhu, Jerry Chun-Wei Lin, Yongrui Qin
ADMA2
2018 SLIND: Identifying Stable Links in Online Social Networks
Ji Zhang 0001, Leonard Tan, Xiaohui Tao 0001, Xiaoyao Zheng, Yonglong Luo, Jerry Chun-Wei Lin
DASFAA (2)1
2018 Anonymization of Multiple and Personalized Sensitive Attributes
Jerry Chun-Wei Lin, Qiankun Liu 0002, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DaWaK5
2018 A Recommender System with Advanced Time Series Medical Data Analysis for Diabetes Patients in a Telehealth Environment
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Jerry Chun-Wei Lin, Fulong Chen 0002, Yonglong Luo, Xiaoyao Zheng
DEXA (2)2
2018 A Metaheuristic Algorithm for Hiding Sensitive Itemsets
Jerry Chun-Wei Lin, Yuyu Zhang, Philippe Fournier-Viger, Youcef Djenouri, Ji Zhang 0001
DEXA (2)5
2018 On Link Stability Detection for Online Social Networks
Ji Zhang 0001, Xiaohui Tao 0001, Leonard Tan, Jerry Chun-Wei Lin, Hongzhou Li, Liang Chang 0003
DEXA (1)1
2018 A tourism destination recommender system using users' sentiment and temporal dynamics
Xiaoyao Zheng, Yonglong Luo, Ji Zhang 0001, Fulong Chen 0002
J. Intell. Inf. Syst.4
2018 Exploiting highly qualified pattern with frequency and weight occupancy
Wensheng Gan, Jerry Chun-Wei Lin, Philippe Fournier-Viger, Han-Chieh Chao, Justin Zhijun Zhan, Ji Zhang 0001
Knowl. Inf. Syst.6
2018 FrauDetector+: An Incremental Graph-Mining Approach for Efficient Fraudulent Phone Call Detection
abstract
In recent years, telecommunication fraud has become more rampant internationally with the development of modern technology and global communication. Because of rapid growth in the volume of call logs, the task of fraudulent phone call detection is confronted with big data issues in real-world implementations. Although our previous work, FrauDetector , addressed this problem and achieved some promising results, it can be further enhanced because it focuses only on fraud detection accuracy, whereas the efficiency and scalability are not top priorities. Other known approaches for fraudulent call number detection suffer from long training times or cannot accurately detect fraudulent phone calls in real time. However, the learning process of FrauDetector is too time-consuming to support real-world application. Although we have attempted to accelerate the the learning process of FrauDetector by parallelization, the parallelized learning process, namely PFrauDetector , still cannot afford the computing cost. In this article, we propose a highly efficient incremental graph-mining-based fraudulent phone call detection approach, namely FrauDetector + , which can automatically label fraudulent phone numbers with a “fraud” tag a crucial prerequisite for distinguishing fraudulent phone call numbers from nonfraudulent ones. FrauDetector + initially generates smaller, more manageable subnetworks from original graph and performs a parallelized weighted HITS algorithm for a significant speed increase in the graph learning module. It adopts a novel aggregation approach to generate a trust (or experience) value for each phone number (or user) based on their respective local values. After the initial procedure, we can incrementally update the trust (or experience) value for each phone number (or user) while a new fraud phone number is identified. An efficient fraud-centric hash structure is constructed to support fast real-time detection of fraudulent phone numbers in the detection module. We conduct a comprehensive experimental study based on real datasets collected through an antifraud mobile application called Whoscall . The results demonstrate a significantly improved efficiency of our approach compared with FrauDetector as well as superior performance against other major classifier-based methods.
Jia-Ching Ying, Ji Zhang 0001, Che-Wei Huang, Kuan-Ta Chen, Vincent S. Tseng
ACM Trans. Knowl. Discov. Data2
2017 Mining Drug Properties for Decision Support in Dental Clinics
WeePheng Goh, Xiaohui Tao 0001, Ji Zhang 0001, Jianming Yong
PAKDD (2)3
2017 A Fast Fourier Transform-Coupled Machine Learning-Based Ensemble Model for Disease Risk Prediction Using a Real-Life Dataset
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Wessam Abbas, Yonglong Luo, Fulong Chen 0002, Vincent S. Tseng
PAKDD (1)2
2017 Coupling topic modelling in opinion mining for social media analysis
abstract
Many of social media platforms such as Facebook and Twitter make it easy for everyone to share their thoughts on literally anything. Topic and opinion detection in social media facilitates the identification of emerging societal trends, analysis of public reactions to policies and business products. In this paper, we proposed a new method that combines the opining mining and context-based topic modelling to analyse public opinions on social media data. Context based topic modelling is used to categorise data in groups and discover hidden communities in data group. The unwanted data group discovered by the topic model then will be discarded. A lexicon based opinion mining method will be applied to the remaining data groups to spot out the public sentiment about the entities. A set of Tweets data on Australian Federal Election 2010 was used in our experiments. Our experimental results demonstrate that, with the help of topic modelling, our social media analysis model is accurate and effective.
Xujuan Zhou, Xiaohui Tao 0001, Md Mostafijur Rahman, Ji Zhang 0001
WI4
2017 A two-phase approach to mine short-period high-utility itemsets in transactional databases
Jerry Chun-Wei Lin, Jiexiong Zhang, Philippe Fournier-Viger, Tzung-Pei Hong, Ji Zhang 0001
Adv. Eng. Informatics5
2016 Adopting Hybrid Descriptors to Recognise Leaf Images for Automatic Plant Specie Identification
Ali A. Al-kharaz, Xiaohui Tao 0001, Ji Zhang 0001, Raid Lafta
ADMA3
2016 IRS-HD: An Intelligent Personalized Recommender System for Heart Disease Patients in a Tele-Health Environment
Raid Lafta, Ji Zhang 0001, Xiaohui Tao 0001, Yan Li 0002, Vincent S. Tseng
ADMA2
2016 Sentiment Analysis for Depression Detection on Social Networks
Xiaohui Tao 0001, Xujuan Zhou, Ji Zhang 0001, Jianming Yong
ADMA3
2015 Detecting anomalies from big network traffic data using an adaptive detection approach
Ji Zhang 0001, Hongzhou Li, Qigang Gao, Hai H. Wang, Yonglong Luo
Inf. Sci.1
2013 Minimising K-Dominating Set in Arbitrary Network Graphs
Guangyuan Wang, Hua Wang 0002, Xiaohui Tao 0001, Ji Zhang 0001
ADMA (2)4
2013 An efficient and robust privacy protection technique for massive streaming choice-based information
abstract
Protecting users' privacy when transmitting a large amount of data over the Internet is becoming increasingly important nowadays. In this paper, we focus on the streaming choice-based information and propose a novel anonymization technique for providing a strong privacy protection to safeguard against privacy disclosure and information tampering. Our technique utilizes an innovative two-phase encoding-and-decoding approach which is very easy to implement, highly efficient in terms of speed and communication, and is robust against possible tampering from adversaries. The experimental evaluation demonstrates the promising performance of our technique.
Ji Zhang 0001, Yonglong Luo
CIKM1
2013 SODIT: An innovative system for outlier detection using multiple localized thresholding and interactive feedback
abstract
Outlier detection is an important long-standing research problem in data mining and has enjoyed applications in a wide range of applications in business, engineering, biology and security, etc. However, the traditional outlier detection methods inevitably need to use different parameters for detection such as those used to specify the distance or density cutoff for distinguish outliers from normal data points. Using the trial and error approach, the traditional outlier detection methods are rather tedious in parameter tuning. In this demo proposal, we introduce an innovative outlier detection system, called SODIT, that uses localized thresholding to assist the value specification of the thresholds that reflect closely the local data distribution. In addition, easy-to-use user feedback are employed to further facilitate the determination of optimal parameter values. SODIT is able to make outlier detection much easier to operate and produce more accurate, intuitive and informative results than before.
Ji Zhang 0001, Hua Wang 0002, Xiaohui Tao 0001, Lili Sun
ICDE1
2010 DISTRO: A System for Detecting Global Outliers from Distributed Data Streams with Privacy Protection
Ji Zhang 0001, Stijn Dekeyser, Hua Wang 0002, Yanfeng Shu
DASFAA (2)1
2009 Detecting Projected Outliers in High-Dimensional Data Streams
Ji Zhang 0001, Qigang Gao, Hai H. Wang, Qing Liu 0001, Kai Xu 0003
DEXA1
2008 SPOT: A System for Detecting Projected Outliers From High-dimensional Data Streams
abstract
In this paper, we present a new technique, called stream projected ouliter detector (SPOT), to deal with outlier detection problem in high-dimensional data streams. SPOT is unique in a number of aspects. First, SPOT employs a novel window-based time model and decaying cell summaries to capture statistics from the data stream. Second, sparse subspace template (SST), a set of top sparse subspaces obtained by unsupervised and/or supervised learning processes, is constructed in SPOT to detect projected outliers effectively. Multi-Objective genetic algorithm (MOGA) is employed as an effective search method in unsupervised learning for finding outlying subspaces from training data. Finally, SST is able to carry out online self- evolution to cope with dynamics of data streams. This paper provides details on the motivation and technical challenges of detecting outliers from high-dimensional data streams, present an overview of SPOT, and give the plans for system demonstration of SPOT.
Ji Zhang 0001, Qigang Gao, Hai H. Wang
ICDE1
2006 A Novel Method for Detecting Outlying Subspaces in High-dimensional Databases Using Genetic Algorithm
abstract
Detecting outlying subspaces is a relatively new research problem in outlier-ness analysis for high-dimensional data. An outlying subspace for a given data point p is the subspace in which p is an outlier. Outlying subspace detection can facilitate a better characterization process for the detected outliers. It can also enable outlier mining for highdimensional data to be performed more accurately and efficiently. In this paper, we proposed a new method using genetic algorithm paradigm for searching outlying subspaces efficiently. We developed a technique for efficiently computing the lower and upper bounds of the distance between a given point and its kth nearest neighbor in each possible subspace. These bounds are used to speed up the fitness evaluation of the designed genetic algorithm for outlying subspace detection. We also proposed a random sampling technique to further reduce the computation of the genetic algorithm. The optimal number of sampling data is specified to ensure the accuracy of the result. We show that the proposed method is efficient and effective in handling outlying subspace detection problem by a set of experiments conducted on both synthetic and real-life datasets.
Ji Zhang 0001, Qigang Gao, Hai H. Wang
ICDM1
2006 A Framework for Efficient Association Rule Mining in XML Data
abstract
In this article, we propose a framework, called XAR-Miner, for mining ARs from XML documents efficiently. In XAR-Miner, raw data in the XML document first are preprocessed to transform either to an Indexed XML Tree (IX-tree) or to Multirelational Databases (Multi-DB), depending on the size of the XML document and the memory constraint of the system, for efficient data selection and AR mining. Concepts that are relevant to the AR mining task are generalized to produce generalized metapatterns. A suitable metric is devised for measuring the degree of concept generalization in order to prevent undergeneralization or overgeneralization. Resulting generalized metapatterns are used to generate large ARs that meet the support and confidence levels. A greedy algorithm is also presented in order to integrate data selection and large itemset generation to enhance the efficiency of the AR mining process. The experiments conducted show that XAR-Miner is more efficient in performing a large number of AR mining tasks from XML documents than the state-of-the-art method of repetitively scanning through XML documents in order to perform each of the mining tasks.
Ji Zhang 0001, Han Liu 0001, Tok Wang Ling, Robert M. Bruckner, A Min Tjoa
J. Database Manag.1
2006 Detecting outlying subspaces for high-dimensional data: the new task, algorithms, and performance
Ji Zhang 0001, Hai H. Wang
Knowl. Inf. Syst.1
2004 PC-Filter: A Robust Filtering Technique for Duplicate Record Detection in Large Databases
Ji Zhang 0001, Tok Wang Ling, Robert M. Bruckner, Han Liu 0001
DEXA1
2004 On Efficient and Effective Association Rule Mining from XML Data
Ji Zhang 0001, Tok Wang Ling, Robert M. Bruckner, A Min Tjoa, Han Liu 0001
DEXA1
2004 HOS-Miner: A System for Detecting Outlying Subspaces of High-dimensional Data
Ji Zhang 0001, Meng Lou, Tok Wang Ling, Hai H. Wang
VLDB1
2003 Building XML Data Warehouse Based on Frequent Patterns in User Queries
Ji Zhang 0001, Tok Wang Ling, Robert M. Bruckner, A Min Tjoa
DaWaK1