Xiaodong Lin 0004

dblp:59/554-4 · DBLP profile ↗
← Back
22ranked-venue papers
3as first author
8since 2021 · last 2024
0009-0000-0686-5206ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorComputer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 Compositional Text-to-Image Synthesis with Attention Map Control of Diffusion Models
abstract
Recent text-to-image (T2I) diffusion models show outstanding performance in generating high-quality images conditioned on textual prompts. However, they fail to semantically align the generated images with the prompts due to their limited compositional capabilities, leading to attribute leakage, entity leakage, and missing entities. In this paper, we propose a novel attention mask control strategy based on predicted object boxes to address these issues. In particular, we first train a BoxNet to predict a box for each entity that possesses the attribute specified in the prompt. Then, depending on the predicted boxes, a unique mask control is applied to the cross- and self-attention maps. Our approach produces a more semantically accurate synthesis by constraining the attention regions of each token in the prompt to the image. In addition, the proposed method is straightforward and effective and can be readily integrated into existing cross-attention-based T2I generators. We compare our approach to competing methods and demonstrate that it can faithfully convey the semantics of the original text to the generated content and achieve high availability as a ready-to-use plugin. Please refer to https://github.com/OPPO-Mente-Lab/attention-mask-control.
Zekang Chen, Chen Chen 0015, Jian Ma 0010, Haonan Lu, Xiaodong Lin 0004
AAAI6
2023 Interactive Interior Design Recommendation via Coarse-to-fine Multimodal Reinforcement Learning
abstract
Personalized interior decoration design often incurs high labor costs. Recent efforts in developing intelligent interior design systems have focused on generating textual requirement-based decoration designs while neglecting the problem of how to mine homeowner's hidden preferences and choose the proper initial design. To fill this gap, we propose an Interactive Interior Design Recommendation System (IIDRS) based on reinforcement learning (RL). IIDRS aims to find an ideal plan by interacting with the user, who provides feedback on the gap between the recommended plan and their ideal one. To improve decision-making efficiency and effectiveness in large decoration spaces, we propose a Decoration Recommendation Coarse-to-Fine Policy Network (DecorRCFN). Additionally, to enhance generalization in online scenarios, we propose an object-aware feedback generation method that augments model training with diversified and dynamic textual feedback. Extensive experiments on a real-world dataset demonstrate our method outperforms traditional methods by a large margin in terms of recommendation accuracy. Further user studies demonstrate that our method reaches higher real-world user satisfaction than baseline methods.
He Zhang 0030, Ying Sun 0006, Weiyu Guo, Haonan Lu, Xiaodong Lin 0004, Hui Xiong 0001
ACM Multimedia6
2023 VPiP: Values Packing in Paillier for Communication Efficient Oblivious Linear Computations
abstract
The technique of packing multiple values into one message without losing homomorphic computation properties is the main workhorse that drives many exciting advances in applying lattice-based homomorphic encryption schemes to privacy-preserving Machine-Learning-as-a-Service (MLaaS). However, this technique does not directly work for the classic Paillier homomorphic encryption scheme, limiting the use of the Paillier scheme in the privacy-preserving MLaaS. To enrich the applications of Paillier in privacy-preserving MLaaS, we present a set of new methods for efficient linear computations over packed values under the Paillier scheme, such as vector multiplication, matrix multiplication, and convolutional calculation between ciphertexts and plaintexts. Different from the packing methods of lattice-based schemes, the Paillier packing method naturally allows higher packing capability for values in lower bit-length. This property can significantly benefit privacy-preserving MLaaS, as the values of user inputs and parameters of machine learning models are often quantized into low bits (e.g., 1-8 bits). We conduct comparisons based on different linear computation tasks, the proposed methods under the Paillier scheme clearly outperform the state-of-the-art in terms of communication and computational efficiency, especially in realistic scenarios. For example, compared to one of the recent arts CrypTFlow2 [1], the communication cost of our solution can be 21.7× smaller at best. Thanks to the reduction of communication cost, the runtime can be 2.46× faster than CrypTFlow2 at the median-country-speed of current global mobile broadband.
Weibin Wu 0003, Jun Wang 0020, Yangpan Zhang, Zhe Liu 0001, Lu Zhou 0002, Xiaodong Lin 0004
IEEE Trans. Inf. Forensics Secur.6
2022 Learning to Walk with Dual Agents for Knowledge Graph Reasoning
abstract
Graph walking based on reinforcement learning (RL) has shown great success in navigating an agent to automatically complete various reasoning tasks over an incomplete knowledge graph (KG) by exploring multi-hop relational paths. However, existing multi-hop reasoning approaches only work well on short reasoning paths and tend to miss the target entity with the increasing path length. This is undesirable for many reasoning tasks in real-world scenarios, where short paths connecting the source and target entities are not available in incomplete KGs, and thus the reasoning performances drop drastically unless the agent is able to seek out more clues from longer paths. To address the above challenge, in this paper, we propose a dual-agent reinforcement learning framework, which trains two agents (Giant and Dwarf) to walk over a KG jointly and search for the answer collaboratively. Our approach tackles the reasoning challenge in long paths by assigning one of the agents (Giant) searching on cluster-level paths quickly and providing stage-wise hints for another agent (Dwarf). Finally, experimental results on several KG reasoning benchmarks show that our approach can search answers more accurately and efficiently, and outperforms existing RL-based methods for long path queries by a large margin.
Zixuan Yuan, Hao Liu 0026, Xiaodong Lin 0004, Hui Xiong 0001
AAAI4
2022 GammaE: Gamma Embeddings for Logical Queries on Knowledge Graphs
abstract
Embedding knowledge graphs (KGs) for multihop logical reasoning is a challenging problem due to massive and complicated structures in many KGs.Recently, many promising works projected entities and queries into a geometric space to efficiently find answers.However, it remains challenging to model the negation and union operator.The negation operator has no strict boundaries, which generates overlapped embeddings and leads to obtaining ambiguous answers.An additional limitation is that the union operator is non-closure, which undermines the model to handle a series of union operators.To address these problems, we propose a novel probabilistic embedding model, namely Gamma Embeddings (GammaE), for encoding entities and queries to answer different types of FOL queries on KGs.We utilize the linear property and strong boundary support of the Gamma distribution to capture more features of entities and queries, which dramatically reduces model uncertainty.Furthermore, Gam-maE implements the Gamma mixture method to design the closed union operator.The performance of GammaE is validated on three large logical query datasets.Experimental results show that GammaE significantly outperforms state-of-the-art models on public benchmarks.
Peijun Qing, Haonan Lu, Xiaodong Lin 0004
EMNLP5
2022 DensE: An enhanced non-commutative representation for knowledge graph embedding with adaptive semantic hierarchy
Haonan Lu, Hailin Hu 0002, Xiaodong Lin 0004
Neurocomputing3
2021 An outer-inner linearization method for non-convex and nondifferentiable composite regularization problems
Xiaodong Lin 0004, Andrzej Ruszczynski, Yu Du 0003
J. Glob. Optim.2
2021 Anomaly detection in dynamic attributed networks
Ruizhi Zhou, Qin Zhang 0011, Peng Zhang 0001, Lingfeng Niu, Xiaodong Lin 0004
Neural Comput. Appl.5
2020 In-Bed Body Motion Detection and Classification System
abstract
In-bed motion detection and classification are important techniques that can enable an array of applications, among which are sleep monitoring and abnormal movement detection. In this article, we present a low-cost, low-overhead, and highly robust system for in-bed movement detection and classification that uses low-end load cells. To detect movements, we have designed a feature that we refer to as Log-Peak, which can be extracted from load cell data that is collected through wireless links in an energy-efficient manner. After detection, we set out to achieve a precise body motion classification. Toward this goal, we define nine classes of movements, and design a machine learning algorithm using Support Vector Machine, Random Forest, and XGBoost techniques to classify a movement into one of nine classes. For every movement, we have extracted 24 features and used them in our model. This movement detection/classification system was evaluated on data collected from 40 subjects who performed 35 predefined movements in each experiment. We have applied multiple tree topologies for each technique to reach their best results. After examining various combinations, we have achieved a final classification accuracy of 91.5%. This system can be used conveniently for long-term home monitoring.
Musaab Alaziz, Zhenhua Jia, Richard E. Howard, Xiaodong Lin 0004, Yanyong Zhang
ACM Trans. Sens. Networks4
2017 Leakage of signal function with reused keys in RLWE key exchange
abstract
In this paper, we show that the signal function used in Ring-Learning with Errors (RLWE) key exchange could leak information to find the secret s of a reused public key p = as+2e. This work is motivated by an attack proposed in [1] and gives an insight into how public keys reused for long term in RLWE key exchange protocols can be exploited. This work specifically focuses on the attack on the KE protocol in [2] by initiating multiple sessions with the honest party and analyze the output of the signal function. Experiments have confirmed the success of our attack in recovering the secret.
Jintai Ding, Saed Alsayigh, R. V. Saraswathy, Scott R. Fluhrer, Xiaodong Lin 0004
ICC5
2014 Alternating linearization for structured regularization problems
Xiaodong Lin 0004, Andrzej Ruszczynski
J. Mach. Learn. Res.1
2012 Improving RF-based device-free passive localization in cluttered indoor environments through probabilistic classification methods
abstract
Radio frequency based device-free passive localization has been proposed as an alternative to indoor localization because it does not require subjects to wear a radio device. This technique observes how people disturb the pattern of radio waves in an indoor space and derives their positions accordingly. The well-known multipath effect makes this problem very challenging, because in a complex environment it is impractical to have enough knowledge to be able to accurately model the effects of a subject on the surrounding radio links. In addition, even minor changes in the environment over time change radio propagation sufficiently to invalidate the datasets needed by simple fingerprint-based methods. In this paper, we develop a fingerprinting-based method using probabilistic classification approaches based on discriminant analysis. We also devise ways to mitigate the error caused by multipath effect in data collection, further boosting the classification likelihood.
Chenren Xu, Bernhard Firner, Yanyong Zhang, Richard E. Howard, Jun Li 0034, Xiaodong Lin 0004
IPSN6
2012 A covariance-free iterative algorithm for distributed principal component analysis on vertically partitioned data
Yue-Fei Guo, Xiaodong Lin 0004, Zhou Teng, Xiangyang Xue 0001, Jianping Fan 0001
Pattern Recognit.2
2010 Reachability Analysis in Privacy-Preserving Perturbed Graphs
abstract
Many real world phenomena can be naturally modeled as graph structures whose nodes representing entities and whose edges representing interactions or relationships between entities. The analysis of the graph data have many practical implications. However, the release of the data often poses considerable privacy risk to the individuals involved. In this paper, we address the edge privacy problem in graphs. In particular, we explore random perturbation for privacy preservation in graph data, and propose an iterative derivation process to analyze node reachability within the graph. We specifically focus on deriving the probability that the shortest path linking two nodes in a directed graph is of a particular length. This allows us to determine the expected length of the shortest path between two nodes, and determine whether they are linked or not. The performance of the proposed method is demonstrated via extensive experiments on both real and synthetic datasets.
Xiaoyun He, Jaideep Vaidya, Basit Shafiq, Nabil R. Adam, Xiaodong Lin 0004
Web Intelligence5
2010 Active Learning From Stream Data Using Optimal Weight Classifier Ensemble
abstract
In this paper, we propose a new research problem on active learning from data streams, where data volumes grow continuously, and labeling all data is considered expensive and impractical. The objective is to label a small portion of stream data from which a model is derived to predict future instances as accurately as possible. To tackle the technical challenges raised by the dynamic nature of the stream data, i.e., increasing data volumes and evolving decision concepts, we propose a classifier-ensemble-based active learning framework that selectively labels instances from data streams to build a classifier ensemble. We argue that a classifier ensemble's variance directly corresponds to its error rate, and reducing a classifier ensemble's variance is equivalent to improving its prediction accuracy. Because of this, one should label instances toward the minimization of the variance of the underlying classifier ensemble. Accordingly, we introduce a minimum-variance (MV) principle to guide the instance labeling process for data streams. In addition, we derive an optimal-weight calculation method to determine the weight values for the classifier ensemble. The MV principle and the optimal weighting module are combined to build an active learning framework for data streams. Experimental results on synthetic and real-world data demonstrate the performance of the proposed work in comparison with other approaches.
Xingquan Zhu 0001, Peng Zhang 0001, Xiaodong Lin 0004, Yong Shi 0001
IEEE Trans. Syst. Man Cybern. Part B3
2009 A distributed approach to enabling privacy-preserving model-based classifier training
Hangzai Luo, Jianping Fan 0001, Xiaodong Lin 0004, Aoying Zhou, Elisa Bertino
Knowl. Inf. Syst.3
2007 Active Learning from Data Streams
abstract
In this paper, we address a new research problem on active learning from data streams where data volumes grow continuously and labeling all data is considered expensive and impractical. The objective is to label a small portion of stream data from which a model is derived to predict newly arrived instances as accurate as possible. In order to tackle the challenges raised by data streams' dynamic nature, we propose a classifier ensembling based active learning framework which selectively labels instances from data streams to build an accurate classifier. A minimal variance principle is introduced to guide instance labeling from data streams. In addition, a weight updating rule is derived to ensure that our instance labeling process can adaptively adjust to dynamic drifting concepts in the data. Experimental results on synthetic and real-world data demonstrate the performances of the proposed efforts in comparison with other simple approaches.
Xingquan Zhu 0001, Peng Zhang 0001, Xiaodong Lin 0004, Yong Shi 0001
ICDM3
2007 Information Conversion, Effective Samples, and Parameter Size
abstract
Consider the relative entropy between a posterior density for a parameter given a sample and a second posterior density for the same parameter, based on a different model and a different data set. Then the relative entropy can be minimized over the second sample to get a virtual sample that would make the second posterior as close as possible to the first in an informational sense. If the first posterior is based on a dependent dataset and the second posterior uses an independence model, the effective inferential power of the dependent sample is transferred into the independent sample by the optimization. Examples of this optimization are presented for models with nuisance parameters, finite mixture models, and models for correlated data. Our approach is also used to choose the effective parameter size in a Bayesian hierarchical model.
Xiaodong Lin 0004, Jennifer Pittman, Bertrand S. Clarke
IEEE Trans. Inf. Theory1
2006 Gene selection using support vector machines with non-convex penalty
abstract
MOTIVATION: With the development of DNA microarray technology, scientists can now measure the expression levels of thousands of genes simultaneously in one single experiment. One current difficulty in interpreting microarray data comes from their innate nature of 'high-dimensional low sample size'. Therefore, robust and accurate gene selection methods are required to identify differentially expressed group of genes across different samples, e.g. between cancerous and normal cells. Successful gene selection will help to classify different cancer types, lead to a better understanding of genetic signatures in cancers and improve treatment strategies. Although gene selection and cancer classification are two closely related problems, most existing approaches handle them separately by selecting genes prior to classification. We provide a unified procedure for simultaneous gene selection and cancer classification, achieving high accuracy in both aspects. RESULTS: In this paper we develop a novel type of regularization in support vector machines (SVMs) to identify important genes for cancer classification. A special nonconvex penalty, called the smoothly clipped absolute deviation penalty, is imposed on the hinge loss function in the SVM. By systematically thresholding small estimates to zeros, the new procedure eliminates redundant genes automatically and yields a compact and accurate classifier. A successive quadratic algorithm is proposed to convert the non-differentiable and non-convex optimization problem into easily solved linear equation systems. The method is applied to two real datasets and has produced very promising results. AVAILABILITY: MATLAB codes are available upon request from the authors.
Hao Helen Zhang 0001, Jeongyoun Ahn, Xiaodong Lin 0004, Cheolwoo Park
Bioinform.3
2005 Privacy-preserving clustering with distributed EM mixture modeling
Xiaodong Lin 0004, Chris Clifton, Michael Yu Zhu
Knowl. Inf. Syst.1
2004 Privacy preserving regression modelling via distributed computation
abstract
Reluctance of data owners to share their possibly confidential or proprietary data with others who own related databases is a serious impediment to conducting a mutually beneficial data mining analysis. We address the case of vertically partitioned data -- multiple data owners/agencies each possess a few attributes of every data record. We focus on the case of the agencies wanting to conduct a linear regression analysis with complete records without disclosing values of their own attributes. This paper describes an algorithm that enables such agencies to compute the exact regression coefficients of the global regression equation and also perform some basic goodness-of-fit diagnostics while protecting the confidentiality of their data. In more general settings beyond the privacy scenario, this algorithm can also be viewed as method for the distributed computation for regression analyses.
Ashish P. Sanil, Alan F. Karr, Xiaodong Lin 0004, Jerome P. Reiter
KDD3
2004 Learning a complex metabolomic dataset using random forests and support vector machines
abstract
Metabolomics is the "omics" science of biochemistry. The associated data include the quantitative measurements of all small molecule metabolites in a biological sample. These datasets provide a window into dynamic biochemical networks and conjointly with other "omic" data, genes and proteins, have great potential to unravel complex human diseases. The dataset used in this study has 63 individuals, normal and diseased, and the diseased are drug treated or not, so there are three classes. The goal is to classify these individuals using the observed metabolite levels for 317 measured metabolites. There are a number of statistical challenges: non-normal data, the number of samples is less than the number of metabolites; there are missing data and the fact that data are missing is informative (assay values below detection limits can point to a specific class); also, there are high correlations among the metabolites. We investigate support vector machines (SVM), and random forest (RF), for outlier detection, variable selection and classification. We use the variables selected with RF in SVM and visa versa. The benefit of this study is insight into interplay of variable selection and classification methods. We link our selected predictors to the biochemistry of the disease.
Young Truong, Xiaodong Lin 0004, Chris Beecher
KDD2