VLDB 2026 Research / reviewers in the wild / expert
Jilian Zhang
dblp:51/1086
· DBLP profile ↗
59ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0001-7117-3951ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 22 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Security and privacy · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | E2E-PP: End-to-End Privacy Protection via compressive sensing and personalized differential privacy for mobile crowdsensing
Xingyu Zheng, Kaimin Wei, Zhiquan Liu 0001, Jinpeng Chen 0001, Chengkun Jia, Jilian Zhang |
Comput. Secur. | 6 |
| 2026 | RNN-based optimal consensus of high-order heterogeneous nonlinear MAS with input constraints
Yinyan Zhang, Yuxuan Xiong, Jilian Zhang, Guanggang Geng, Shuai Li 0002 |
Neurocomputing | 3 |
| 2026 | A Robust Reversible Watermarking scheme using DC prediction and histogram shifting
Jiancheng Xiao, Shuaichao Wu, Bingwen Feng, Jilian Zhang, Bing Chen 0004, Zhihua Xia, Wei Lu 0001 |
Signal Process. | 4 |
| 2026 | Unveiling the Centralized Security Risks in Decentralized EcosystemsabstractThe decentralized ecosystem is claimed to avoid security risks caused by centralization. Decentralized services, such as crypto wallets and decentralized applications (DApps), are purported to offer more reliable security and better protect user privacy. However, our research suggests a different reality: centralized components or scenarios are still prevalent within decentralized ecosystems, introducing security risks typically associated with centralization. This work systematically investigated the centralized security risks in crypto wallets and DApps. We found seven security risks and developed a series of methods to identify these risks. The detection results indicate that centralized security risks are widespread in the decentralized ecosystem. Among the 28 Ethereum-recommended crypto wallets, 96.4% have security risks. Of the 78 Web3 sites (frontends of DApps), 100% contain third-party scripts, and 44.9% expose the user's address to third parties. Furthermore, we developed a high-precision automated tool and inspected 110,506 on-chain smart contracts (backends of DApps), discovering that 83.5% contain at least one security risk. These risks affect 260 well-known tokens with a combined market capitalization exceeding${\$}$98 billion. Kailun Yan, Jilian Zhang, Wenrui Diao |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | Medical Image Privacy in Federated Learning: Segmentation-Reorganization and Sparsified Gradient Matching AttacksabstractIn modern medicine, the widespread use of medical imaging has greatly improved diagnostic and treatment efficiency. However, these images contain sensitive personal information, and any leakage could seriously compromise patient privacy, leading to ethical and legal issues. Federated learning (FL), an emerging privacy-preserving technique, transmits gradients rather than raw data for model training. Yet, recent studies reveal that gradient inversion attacks can exploit this information to reconstruct private data, posing a significant threat to FL. Current attacks remain limited in image resolution, similarity, and batch processing, and thus do not yet pose a significant risk to FL. To address this, we propose a novel gradient inversion attack based on sparsified gradient matching and segmentation reorganization (SR) to reconstruct high-resolution, high-similarity medical images in batch mode. Specifically, an $L_{1}$ loss function optimises the gradient sparsification process, while the SR strategy enhances image resolution. An adaptive learning rate adjustment mechanism is also employed to improve optimisation stability and avoid local optima. Experimental results demonstrate that our method significantly outperforms state-of-the-art approaches in both visual quality and quantitative metrics, achieving up to a 146% improvement in similarity. Kaimin Wei, Chengkun Jia, Jinpeng Chen 0001, Jilian Zhang, Yongdong Wu |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Sanitizing Backdoored Graph Neural Networks: A Multidimensional ApproachabstractGraph Neural Networks (GNNs) are known to be prone to adversarial attacks, among which backdoor attack is a major security threat. By injecting backdoor triggers into a graph and assigning a target class label to nodes attached to the triggers, the attacker can mislead the GNN model trained on the poisoned graph to classify test nodes attached with a trigger to the target class. To defend against backdoor attacks, existing defense methods rely on anomaly detection in feature distribution or label transformation. However, these approaches are incapable of detecting in-distribution triggers or clean-label attacks that do not alter the class label of target nodes. To tackle these threats, we empirically analyze triggers from a multidimensional aspect, and our analysis shows that there are clear distinctions between trigger nodes and normal ones in terms of node feature values, node embeddings, and class prediction probabilities. Based on these findings, we propose a Multidimensional Anomaly Detection framework (MAD) that can effectively minimize the impact of triggers by pruning away anomalous nodes and edges. Extensive experiments show that at the cost of slight loss in clean classification accuracy, MAD achieves considerably lower attack success rate as compared to state-of-the-art backdoor defense methods. Jilian Zhang, Yinyan Zhang, Jian Weng 0001 |
IJCAI | 2 |
| 2025 | JPEG Compression-Resistant Generative Image Hiding Utilizing Cascaded Invertible NetworksabstractGenerative steganography is renowned for its exceptional undetectability. However, prevalent generative methods often have insufficient capacity for concealing secret images. Furthermore, the sensitivity of commonly utilized generative models exacerbates the challenge of ensuring robustness against channel distortions such as JPEG compression. In this paper, we introduce a generative image hiding network that employs two invertible generators to transform secret images into stego images within a disparate image domain. Additionally, we seamlessly integrate an up-and-down sampling module (UDM) within these generators to facilitate efficient decoupling of the intermediate representations obtained by each generator. The UDM serves multiple purposes: preserving coherence between the intermediate representations, enhancing resilience against JPEG compression, and safeguarding the confidentiality of the concealed images. To address the complexity of mapping both uncompressed and compressed stego images to a unified intermediary representation, we implement two distinct flows for the forward and backward processes of the generator associated with the stego images. The experimental results show that our scheme offers concurrent advantages in terms of full-size image hiding ability, undetectability, confidentiality, and robustness. Tiewei Qin, Bingwen Feng, Bingbing Zhou, Jilian Zhang, Zhihua Xia, Jian Weng 0001, Wei Lu 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Verifiable Graph-Based Approximate Nearest Neighbor Search
Chenzhao Wang, Jilian Zhang, Kaimin Wei, Bingwen Feng |
ADMA (3) | 2 |
| 2024 | GI-SMN: Gradient Inversion Attack Against Federated Learning Without Prior Knowledge
Kaimin Wei, Yongdong Wu, Jilian Zhang, Jinpeng Chen 0001, Huan Bao |
ICIC (8) | 4 |
| 2024 | Privacy-Preserving k-core Decomposition for Graphs
Bingwen Feng, Jilian Zhang |
WISE (5) | 4 |
| 2024 | Noise-resistant graph neural networks with manifold consistency and label consistency
Zhengyu Lu, Guoqiu Wen, Jilian Zhang |
Expert Syst. Appl. | 6 |
| 2024 | Graph augmentation for node-level few-shot learning
Zongqian Wu, Peng Zhou 0011, Junbo Ma, Jilian Zhang, Guoqin Yuan, Xiaofeng Zhu 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Quantum KNN Classification With K Value Selection and Neighbor SelectionabstractThe KNN (K-nearest neighbors) algorithm is one of Top-10 data mining algorithms and is widely used in various fields of artificial intelligence. This leads to that quantum KNN algorithms have developed and achieved certain speed improvements, denoted as Q-KNN. However, these Q-KNN methods must face two key problems as follows. The first one is that they are mainly focused on neighbor selection without paying attention to the influence of K value on the algorithm. The second is that only the neighbor selection process is quantized, and the selection of K value is not quantized. To solve these problems, this paper designs a novel quantum circuit for KNN classification, so as to simultaneously quantumize the neighbor selection and K value selection process. Specifically, the least squares loss and sparse regularization term are first used to construct the objective function of the proposed quantum KNN, so that it can simultaneously obtain the optimal K value and K nearest neighbors of the testing data. And then, a new quantum circuit is proposed to quantumize the process through quantum phase estimation, controlled rotation, and inverse phase estimation techniques. Finally, experiments are conducted with qiskit and matlab to output the quantum and classical results of the algorithm, verifying that the proposed algorithm can output the optimal K value and K nearest neighbors for each testing data. Jiaye Li 0001, Jian Zhang 0048, Jilian Zhang, Shichao Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Bad Apples: Understanding the Centralized Security Risks in Decentralized EcosystemsabstractThe blockchain-powered decentralized applications and systems have been widely deployed in recent years. The decentralization feature promises users anonymity, security, and non-censorship, which is especially welcomed in the areas of decentralized finance and digital assets. From the perspective of most common users, a decentralized ecosystem means every service follows the principle of decentralization. However, we find that the services in a decentralized ecosystem still may contain centralized components or scenarios, like third-party SDKs and privileged operations, which violate the promise of decentralization and may cause a series of centralized security risks. In this work, we systematically study the centralized security risks existing in decentralized ecosystems. Specifically, we identify seven centralized security risks in the deployment of two typical decentralized services – crypto wallets and DApps, such as anonymity loss and overpowered owner. Also, to measure these risks in the wild, we designed an automated detection tool called Naga and carried out large-scale experiments. Based on the measurement of 28 Ethereum crypto wallets (Android version) and 110,506 on-chain smart contracts, the result shows that the centralized security risks are widespread. Up to 96.4% of wallets and 83.5% of contracts exist at least one security risk, including 260 well-known tokens with a total market cap of over $98 billion. Kailun Yan, Jilian Zhang, Wenrui Diao, Shanqing Guo |
WWW | 2 |
| 2023 | Improving Bitcoin Transaction Propagation Efficiency through Local Clique NetworkabstractAbstract Bitcoin is a popular decentralized cryptocurrency, and the Bitcoin network is essentially an unstructured peer-to-peer (P2P) network that can synchronize distributed database of replicated ledgers through message broadcasting. In the Bitcoin network, the average clustering coefficient of nodes is very high, resulting in low message propagation efficiency. In addition, average node degree in the Bitcoin network is also considerably large, causing high message redundancy when nodes use the gossip protocol to broadcast messages. These may affect message propagation speed, hindering Bitcoin from being applied to scenarios of high transactional throughputs. To illustrate, we have collected single-hop propagation data of transactions of 366 blocks from Bitcoin Core. The analysis results show that transaction verification and network delay are two major causes of low transaction propagation efficiency. In this paper, we propose a novel P2P network structure, called local clique network (LCN), for message broadcasting in the Bitcoin network. Specifically, to reduce transaction validation latency and message redundancy, in LCN local nodes (logically) form cliques, and only a few nodes in a clique broadcast messages to the other cliques, instead of each node sending messages to its neighboring nodes. We have conducted extensive experiments, and the results show that message redundancy is low in LCN, and message propagation speed increases significantly. Meanwhile, LCN exhibits excellent robustness when average node degree remains high in the Bitcoin network. Kailun Yan, Jilian Zhang, Yongdong Wu |
Comput. J. | 2 |
| 2023 | FePN: A robust feature purification network to defend against adversarial examples
Dongliang Cao, Kaimin Wei, Yongdong Wu, Jilian Zhang, Bingwen Feng, Jinpeng Chen 0001 |
Comput. Secur. | 4 |
| 2023 | Dynamic Moore-Penrose Inversion With Unknown Derivatives: Gradient Neural Network ApproachabstractFinding dynamic Moore-Penrose inverses (DMPIs) in real-time is a challenging problem due to the time-varying nature of the inverse. Traditional numerical methods for static Moore-Penrose inverse are not efficient for calculating DMPIs and are restricted by serial processing. The current state-of-the-art method for finding DMPIs is called the zeroing neural network (ZNN) method, which requires that the time derivative of the associated matrix is available all the time during the solution process. However, in practice, the time derivative of the associated dynamic matrix may not be available in a real-time manner or be subject to noises caused by differentiators. In this article, we propose a novel gradient-based neural network (GNN) method for computing DMPIs, which does not need the time derivative of the associated dynamic matrix. In particular, the neural state matrix of the proposed GNN converges to the theoretical DMPI in finite time. The finite-time convergence is kept by simply setting a large parameter when there are additive noises in the implementation of the GNN model. Simulation results demonstrate the efficacy and superiority of the proposed GNN method. Yinyan Zhang, Jilian Zhang, Jian Weng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | High-Performance UAV Crowdsensing: A Deep Reinforcement Learning ApproachabstractPath planning is critical to realizing a high-performance unmanned aerial vehicle (UAV) crowdsensing system, which can be deployed to carry out large-scale tasks in the physical world, especially in emergency scenarios, such as earthquakes and mudslides. Deep reinforcement learning (DRL) has recently proven its superiority in path design. However, it is often applied under the assumption that the entire status of the target region is available, which is hard to achieve in practice. Instead, efforts should be made to ensure the efficient flight of several UAVs in order to collect data with incomplete observations in specified places. In this work, we set out to create a high-performance UAV crowdsensing system by combining DRL with partial observations. We present a novel DRL-based path-planning algorithm called DRL-PP. Specifically, we integrate an attention mechanism into the actor–critic technique to assist UAV swarm collaboration to collect data. We also design an incentive mechanism to ease the problem of sparse reward. Furthermore, we provide a dilemma detection system to prevent the generation of overlapping flight paths. Experimental results from extensive simulations prove that compared with the state-of-the-art approaches, the proposed DRL-PP can significantly improve the efficiency of data collection. Kaimin Wei, Yongdong Wu, Zhetao Li, Hongliang He 0004, Jilian Zhang, Jinpeng Chen 0001, Song Guo 0001 |
IEEE Internet Things J. | 6 |
| 2022 | Multimodality Alzheimer's Disease Analysis in Deep Riemannian Manifold
Junbo Ma, Jilian Zhang |
Inf. Process. Manag. | 2 |
| 2022 | Attention-Emotion-Enhanced Convolutional LSTM for Sentiment AnalysisabstractLong short-term memory (LSTM) neural networks and attention mechanism have been widely used in sentiment representation learning and detection of texts. However, most of the existing deep learning models for text sentiment analysis ignore emotion's modulation effect on sentiment feature extraction, and the attention mechanisms of these deep neural network architectures are based on word- or sentence-level abstractions. Ignoring higher level abstractions may pose a negative effect on learning text sentiment features and further degrade sentiment classification performance. To address this issue, in this article, a novel model named AEC-LSTM is proposed for text sentiment detection, which aims to improve the LSTM network by integrating emotional intelligence (EI) and attention mechanism. Specifically, an emotion-enhanced LSTM, named ELSTM, is first devised by utilizing EI to improve the feature learning ability of LSTM networks, which accomplishes its emotion modulation of learning system via the proposed emotion modulator and emotion estimator. In order to better capture various structure patterns in text sequence, ELSTM is further integrated with other operations, including convolution, pooling, and concatenation. Then, topic-level attention mechanism is proposed to adaptively adjust the weight of text hidden representation. With the introduction of EI and attention mechanism, sentiment representation and classification can be more effectively achieved by utilizing sentiment semantic information hidden in text topic and context. Experiments on real-world data sets show that our approach can improve sentiment classification performance effectively and outperform state-of-the-art deep learning-based methods significantly. Faliang Huang, Xuelong Li 0001, Chang-an Yuan 0001, Shichao Zhang 0001, Jilian Zhang, Shaojie Qiao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Non-negative Matrix Factorization: A SurveyabstractAbstract Non-negative matrix factorization (NMF) is a powerful tool for data science researchers, and it has been successfully applied to data mining and machine learning community, due to its advantages such as simple form, good interpretability and less storage space. In this paper, we give a detailed survey on existing NMF methods, including a comprehensive analysis of their design principles, characteristics and drawbacks. In addition, we also discuss various variants of NMF methods and analyse properties and applications of these variants. Finally, we evaluate the performance of nine NMF methods through numerical experiments, and the results show that NMF methods perform well in clustering tasks. Jiangzhang Gan, Jilian Zhang |
Comput. J. | 4 |
| 2021 | DeepChain: Auditable and Privacy-Preserving Deep Learning with Blockchain-Based IncentiveabstractDeep learning can achieve higher accuracy than traditional machine learning algorithms in a variety of machine learning tasks. Recently, privacy-preserving deep learning has drawn tremendous attention from information security community, in which neither training data nor the training model is expected to be exposed. Federated learning is a popular learning mechanism, where multiple parties upload local gradients to a server and the server updates model parameters with the collected gradients. However, there are many security problems neglected in federated learning, for example, the participants may behave incorrectly in gradient collecting or parameter updating, and the server may be malicious as well. In this article, we present a distributed, secure, and fair deep learning framework named DeepChain to solve these problems. DeepChain provides a value-driven incentive mechanism based on Blockchain to force the participants to behave correctly. Meanwhile, DeepChain guarantees data privacy for each participant and provides auditability for the whole training process. We implement a prototype of DeepChain and conduct experiments on a real dataset for different settings, and the results show that our DeepChain is promising. Jia-Si Weng 0001, Jian Weng 0001, Jilian Zhang, Ming Li 0049, Yue Zhang 0025, Weiqi Luo 0002 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2020 | Unsupervised nonlinear feature selection algorithm via kernel function
Jiaye Li 0001, Shichao Zhang 0001, Leyuan Zhang, Cong Lei, Jilian Zhang |
Neural Comput. Appl. | 5 |
| 2020 | Heuristic algorithms for diversity-aware balanced multi-way number partitioning
Jilian Zhang, Kaimin Wei, Xuelian Deng |
Pattern Recognit. Lett. | 1 |
| 2020 | Sparse Graph Connectivity for Image SegmentationabstractIt has been demonstrated that the segmentation performance is highly dependent on both subspace preservation and graph connectivity. In the literature, the full connectivity method linearly represents each data point ( e.g., a pixel in one image) by all data points for achieving subspace preservation, while the sparse connectivity method was designed to linearly represent each data point by a set of data points for achieving graph connectivity. However, previous methods only focused on either subspace preservation or graph connectivity. In this article, we propose a Sparse Graph Connectivity (SGC) method for image segmentation to automatically learn the affinity matrix from the low-dimensional space of original data, which aims at simultaneously achieving subspace preservation and graph connectivity. To do this, the proposed SGC simultaneously learns a self-representation affinity matrix for subspace preservation and a sparse affinity matrix for graph connectivity, from the intrinsic low-dimensional feature space of high-dimensional original data. Meanwhile, the self-representation affinity matrix is pushed to be similar to the sparse affinity as well as be the final segmentation results. Experimental result on synthetic and real-image datasets showed that our SGC method achieved the best segmentation performance, compared to state-of-the-art segmentation methods. Xiaofeng Zhu 0001, Shichao Zhang 0001, Jilian Zhang, Guangquan Lu, Yang Yang 0002 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2019 | Feature selection for text classification: A review
Xuelian Deng, Jian Weng 0001, Jilian Zhang |
Multim. Tools Appl. | 4 |
| 2019 | Nonlinear sparse feature selection algorithm via low matrix rank constraint
Leyuan Zhang, Yangding Li, Jilian Zhang, Pengqing Li, Jiaye Li 0001 |
Multim. Tools Appl. | 3 |
| 2019 | Low-Rank Sparse Subspace for Spectral ClusteringabstractTraditional graph clustering methods consist of two sequential steps, i.e., constructing an affinity matrix from the original data and then performing spectral clustering on the resulting affinity matrix. This two-step strategy achieves optimal solution for each step separately, but cannot guarantee that it will obtain the globally optimal clustering results. Moreover, the affinity matrix directly learned from the original data will seriously affect the clustering performance, since high-dimensional data are usually noisy and may contain redundancy. To address the above issues, this paper proposes a Low-rank Sparse Subspace (LSS) clustering method via dynamically learning the affinity matrix from low-dimensional space of the original data. Specifically, we learn a transformation matrix to project the original data to their low-dimensional space, by conducting feature selection and subspace learning in the sample self-representation framework. Then, we utilize the rank constraint and the affinity matrix directly obtained from the original data to construct a dynamic and intrinsic affinity matrix. Moreover, each of these three matrices is updated iteratively while fixing the other two. In this way, the affinity matrix learned from the low-dimensional space is the final clustering results. Extensive experiments are conducted on both synthetic and real datasets to show that our proposed LSS method outperforms the state-of-the-art clustering methods. Xiaofeng Zhu 0001, Shichao Zhang 0001, Jilian Zhang, Lifeng Yang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2018 | Pattern discovery from multi-source data
Xiaofeng Zhu 0001, Jie Shao 0001, Jilian Zhang |
Pattern Recognit. Lett. | 3 |
| 2018 | Harmonious Genetic ClusteringabstractTo automatically determine the number of clusters and generate more quality clusters while clustering data samples, we propose a harmonious genetic clustering algorithm, named HGCA, which is based on harmonious mating in eugenic theory. Different from extant genetic clustering methods that only use fitness, HGCA aims to select the most suitable mate for each chromosome and takes into account chromosomes gender, age, and fitness when computing mating attractiveness. To avoid illegal mating, we design three mating prohibition schemes, i.e., no mating prohibition, mating prohibition based on lineal relativeness, and mating prohibition based on collateral relativeness, and three mating strategies, i.e., greedy eugenics-based mating strategy, eugenics-based mating strategy based on weighted bipartite matching, and eugenics-based mating strategy based on unweighted bipartite matching, for harmonious mating. In particular, a novel single-point crossover operator called variable-length-and-gender-balance crossover is devised to probabilistically guarantee the balance between population gender ratio and dynamics of chromosome lengths. We evaluate the proposed approach on real-life and artificial datasets, and the results show that our algorithm outperforms existing genetic clustering methods in terms of robustness, efficiency, and effectiveness. Faliang Huang, Xuelong Li 0001, Shichao Zhang 0001, Jilian Zhang |
IEEE Trans. Cybern. | 4 |
| 2018 | Editorial: Deep Mining Big Social Data
Xiaofeng Zhu 0001, Gerard Sanroma, Jilian Zhang, Brent C. Munsell |
World Wide Web | 3 |
| 2017 | Supervised Feature Selection Algorithm Based on Low-Rank and Manifold Learning
Jilian Zhang, Shichao Zhang 0001, Cong Lei |
ADMA | 2 |
| 2017 | Multimodal learning for topic sentiment analysis in microblogging
Faliang Huang, Shichao Zhang 0001, Jilian Zhang, Ge Yu 0001 |
Neurocomputing | 3 |
| 2017 | Low-rank feature selection for multi-view regression
Rongyao Hu, Debo Cheng, Wei He 0017, Guoqiu Wen, Yonghua Zhu, Jilian Zhang, Shichao Zhang 0001 |
Multim. Tools Appl. | 6 |
| 2017 | Overlapping Community Detection for Multimedia Social NetworksabstractFinding overlapping communities from multimedia social networks is an interesting and important problem in data mining and recommender systems. However, extant overlapping community discovery with swarm intelligence often generates overlapping community structures with superfluous small communities. To deal with the problem, in this paper, an efficient algorithm (LEPSO) is proposed for overlapping communities discovery, which is based on line graph theory, ensemble learning, and particle swarm optimization (PSO). Specifically, a discrete PSO, consisting of an encoding scheme with ordered neighbors and a particle updating strategy with ensemble clustering, is devised for improving the optimization ability to search communities hidden in social networks. Then, a postprocessing strategy is presented for merging the finer-grained and suboptimal overlapping communities. Experiments on some real-world and synthetic datasets show that our approach is superior in terms of robustness, effectiveness, and automatically determination of the number of clusters, which can discover overlapping communities that have better quality than those computed by state-of-the-art algorithms for overlapping communities detection. Faliang Huang, Xuelong Li 0001, Shichao Zhang 0001, Jilian Zhang, Zhi-nian Zhai |
IEEE Trans. Multim. | 4 |
| 2016 | Supervised Feature Selection by Robust Sparse Reduced-Rank Regression
Rongyao Hu, Xiaofeng Zhu 0001, Wei He 0017, Jilian Zhang, Shichao Zhang 0001 |
ADMA | 4 |
| 2015 | Maximum Rank QueryabstractThe top-kquery is a common means to shortlist a number of options from a set of alternatives, based on the user's preferences. Typically, these preferences are expressed as a vector of query weights, defined over the options' attributes. The query vector implicitly associates each alternative with a numeric score, and thus imposes a ranking among them. The top-kresult includes thekoptions with the highest scores. In this context, we define themaximum rankquery (MaxRank). Given a focal option in a set of alternatives, theMaxRankproblem is to compute the highest rank this option may achieve under any possible user preference, and furthermore, to report all the regions in the query vector's domain where that rank is achieved.MaxRankfinds application in market impact analysis, customer profiling, targeted advertising, etc. We propose a methodology forMaxRankprocessing and evaluate it with experiments on real and benchmark synthetic datasets. Kyriakos Mouratidis, Jilian Zhang, HweeHwa Pang |
Proc. VLDB Endow. | 2 |
| 2014 | Global immutable region computationabstractA top-k query shortlists the k records in a dataset that best match the user's preferences. To indicate her preferences, the user typically determines a numeric weight for each data dimension (i.e., attribute). We refer to these weights collectively as the query vector. Based on this vector, each data record is implicitly mapped to a score value (via a weighted sum function). The records with the k largest scores are reported as the result. In this paper we propose an auxiliary feature to standard top-k query processing. Specifically, we compute the maximal locus within which the query vector incurs no change in the current top-k result. In other words, we compute all possible query weight settings that produce exactly the same top-k result as the user's original query. We call this locus the global immutable region (GIR). The GIR can be used as a guide to query vector readjustments, as a sensitivity measure for the top-k result, as well as to enable effective result caching. We develop efficient algorithms for GIR computation, and verify their robustness using a variety of real and synthetic datasets. Jilian Zhang, Kyriakos Mouratidis, HweeHwa Pang |
SIGMOD Conference | 1 |
| 2014 | Direct neighbor search
Jilian Zhang, Kyriakos Mouratidis, HweeHwa Pang |
Inf. Syst. | 1 |
| 2013 | Mining Item Popularity for Recommender Systems
Jilian Zhang, Xiaofeng Zhu 0001, Xianxian Li, Shichao Zhang 0001 |
ADMA (2) | 1 |
| 2013 | Mixed-Norm Regression for Visual Classification
Xiaofeng Zhu 0001, Jilian Zhang, Shichao Zhang 0001 |
ADMA (1) | 2 |
| 2013 | Enhancing Access Privacy of Range Retrievals over (𝔹+)-TreesabstractUsers of databases that are hosted on shared servers cannot take for granted that their queries will not be disclosed to unauthorized parties. Even if the database is encrypted, an adversary who is monitoring the I/O activity on the server may still be able to infer some information about a user query. For the particular case of a B+-tree that has its nodes encrypted, we identify properties that enable the ordering among the leaf nodes to be deduced. These properties allow us to construct adversarial algorithms to recover the B+-tree structure from the I/O traces generated by range queries. Combining this structure with knowledge of the key distribution (or the plaintext database itself), the adversary can infer the selection range of user queries. To counter the threat, we propose a privacy-enhancing PB+-tree index which ensures that there is high uncertainty about what data the user has worked on, even to a knowledgeable adversary who has observed numerous query executions. The core idea in PB+-tree is to conceal the order of the leaf nodes in an encrypted B+-tree. In particular, it groups the nodes of the tree into buckets, and employs homomorphic encryption techniques to prevent the adversary from pinpointing the exact nodes retrieved by range queries. PB+-tree can be tuned to balance its privacy strength with the computational and I/O overheads incurred. Moreover, it can be adapted to protect access privacy in cases where the attacker additionally knows a priori the access frequencies of key values. Experiments demonstrate that PB+-tree effectively impairs the adversary's ability to recover the B+-tree structure and deduce the query ranges in all considered scenarios. HweeHwa Pang, Jilian Zhang, Kyriakos Mouratidis |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | Heuristic Algorithms for Balanced Multi-Way Number PartitioningabstractBalanced multi-way number partitioning (BMNP) seeks to split a collection of numbers into subsets with (roughly) the same cardinality and subset sum. The problem is NP-hard, and there are several exact and approximate algorithms for it. However, existing exact algorithms solve only the simpler, balanced two-way number partitioning variant, whereas the most effective approximate algorithm, BLDM, may produce widely varying subset sums. In this paper, we introduce the LRM algorithm that lowers the expected spread in subset sums to one third that of BLDM for uniformly distributed numbers and odd subset cardinalities. We also propose Meld, a novel strategy for skewed number distributions. A combination of LRM and Meld leads to a heuristic technique that consistently achieves a narrower spread of subset sums than BLDM. 1 Jilian Zhang, Kyriakos Mouratidis, HweeHwa Pang |
IJCAI | 1 |
| 2009 | Setting discrete bid levels adaptively in repeated auctionsabstractThe success of an auction design often hinges on its ability to set parameters such as reserve price and bid levels that will maximize an objective function such as the auctioneer revenue. Works on designing adaptive auction mechanisms have emerged recently, and the challenge is in learning different auction parameters by observing the bidding in previous auctions. In this paper, we propose a non-parametric method for determining discrete bid levels dynamically so as to maximize the auctioneer revenue. First, we propose a non-parametric kernel method for estimating the probabilities of closing price with past auction data. Then a greedy strategy has been devised to determine the discrete bid levels based on the estimated probability information of closing price. We show experimentally that our non-parametric method is robust to changes in parameters such as the distributions of participating bidders as well as the individual bidder evaluation, and it consistently outperforms different competitors with various settings with respect to auctioneer revenue maximization. Jilian Zhang, Hoong Chuin Lau, Jialie Shen 0001 |
ICEC | 1 |
| 2009 | POP algorithm: Kernel-based imputation to treat missing values in knowledge discovery from databases
Yongsong Qin, Shichao Zhang 0001, Xiaofeng Zhu 0001, Jilian Zhang, Chengqi Zhang |
Expert Syst. Appl. | 4 |
| 2009 | Estimating confidence intervals for structural differences between contrast groups with missing data
Yongsong Qin, Shichao Zhang 0001, Xiaofeng Zhu 0001, Jilian Zhang, Chengqi Zhang |
Expert Syst. Appl. | 4 |
| 2009 | A decremental algorithm of frequent itemset maintenance for mining updated databases
Shichao Zhang 0001, Jilian Zhang, Zhi Jin 0001 |
Expert Syst. Appl. | 2 |
| 2009 | Scalable Verification for Outsourced Dynamic DatabasesabstractQuery answers from servers operated by third parties need to be verified, as the third parties may not be trusted or their servers may be compromised. Most of the existing authentication methods construct validity proofs based on the Merkle hash tree (MHT). The MHT, however, imposes severe concurrency constraints that slow down data updates. We introduce a protocol, built upon signature aggregation, for checking the authenticity, completeness and freshness of query answers. The protocol offers the important property of allowing new data to be disseminated immediately , while ensuring that outdated values beyond a pre-set age can be detected. We also propose an efficient verification technique for ad-hoc equijoins, for which no practical solution existed. In addition, for servers that need to process heavy query workloads, we introduce a mechanism that significantly reduces the proof construction time by caching just a small number of strategically chosen aggregate signatures. The efficiency and efficacy of our proposed mechanisms are confirmed through extensive experiments. HweeHwa Pang, Jilian Zhang, Kyriakos Mouratidis |
Proc. VLDB Endow. | 2 |
| 2008 | Mining follow-up correlation patterns from time-related databases
Shichao Zhang 0001, Zifang Huang, Jilian Zhang, Xiaofeng Zhu 0001 |
Knowl. Inf. Syst. | 3 |
| 2007 | Measuring the Uncertainty of Differences for Contrasting Groups
Jilian Zhang, Shichao Zhang 0001, Xiaofeng Zhu 0001, Xindong Wu 0001, Chengqi Zhang |
AAAI | 1 |
| 2007 | Cost-Sensitive Imputing Missing Values with Ordering
Xiaofeng Zhu 0001, Shichao Zhang 0001, Jilian Zhang, Chengqi Zhang |
AAAI | 3 |
| 2007 | Cost-Time Sensitive Decision Tree with Missing Values
Shichao Zhang 0001, Xiaofeng Zhu 0001, Jilian Zhang, Chengqi Zhang |
KSEM | 3 |
| 2007 | GBKII: An Imputation Method for Missing Values
Chengqi Zhang, Xiaofeng Zhu 0001, Jilian Zhang, Yongsong Qin, Shichao Zhang 0001 |
PAKDD | 3 |
| 2007 | Semi-parametric optimization for missing data imputation
Yongsong Qin, Shichao Zhang 0001, Xiaofeng Zhu 0001, Jilian Zhang, Chengqi Zhang |
Appl. Intell. | 4 |
| 2007 | EDUA: An efficient algorithm for dynamic database mining
Shichao Zhang 0001, Jilian Zhang, Chengqi Zhang |
Inf. Sci. | 2 |
| 2006 | Difference Detection Between Two Contrast Sets
Huijing Huang, Yongsong Qin, Xiaofeng Zhu 0001, Jilian Zhang, Shichao Zhang 0001 |
DaWaK | 4 |
| 2006 | Identifying Follow-Correlation Itemset-PairsabstractAn association rule ArarrB is useful to predict that B will likely occur when A occurs. This is a classical association rule. In real world applications, such as bioinformatics and medical research, there are many follow correlations between itemsets A and B: B likely occurs n times after A occurred m times, wrote tom, BN>. We refer to this follow-correlation as P3.1 itemset-pairs because3, B1> like that in the example ( Example 2) should be uninterested in association analysis. This paper designs an efficient algorithm for identifying P3.1 itemset-pairs in sequential data. We experimentally evaluate our approach, and demonstrate that the proposed approach is efficient and promising. Shichao Zhang 0001, Jilian Zhang, Xiaofeng Zhu 0001, Zifang Huang |
ICDM | 2 |
| 2006 | Optimized Parameters for Missing Data Imputation
Shichao Zhang 0001, Yongsong Qin, Xiaofeng Zhu 0001, Jilian Zhang, Chengqi Zhang |
PRICAI | 4 |
| 2005 | A Decremental Algorithm for Maintaining Frequent Itemsets in Dynamic Databases
Shichao Zhang 0001, Xindong Wu 0001, Jilian Zhang, Chengqi Zhang |
DaWaK | 3 |