EDBT 2026 Demo / reviewers in the wild / expert
Xiaolin Li 0001
dblp:07/6728-1 · also Xiaolin Andy Li
· DBLP profile ↗
83ranked-venue papers
13as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 6 first-author · 1 since 2021Computer networks · 24 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 since 2021Security and privacy · 5 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CLM-Access: A Specialized Foundation Model for High-Dimensional Single-Cell ATAC-Seq AnalysisabstractInspired by the success of large language models (LLMs) in natural language processing, cell language models (CLMs) have emerged as a promising paradigm to learn cell representations from high-dimensional single-cell data—particularly transcriptomic profiles from scRNA-seq. These foundation models have shown remarkable potential across a variety of downstream applications. However, there remains a lack of foundation models for scATAC-seq data, which measures chromatin accessibility at single-cell level and is critical for decoding epigenetic regulation. Developing such model is considerably more challenging due to the unique characteristics of scATAC-seq data, including the vast number of chromatin regions, lack of standardized annotations, extreme sparsity, and near-binary distributions. To address these challenges, we systematically explore various strategies and propose CLM-Access, a specialized foundation model for scATAC-seq data. CLM-Access incorporates three main innovations: (1) an unified data processing pipeline that maps 2.8 million cells onto an unified reference of over 1 million chromatin regions; (2) a specialized patching and embedding strategy to effectively manage high-dimensional inputs; and (3) a tailored masking and loss function design that preserves fine-grained regional information while enhancing training efficiency and representation quality. With comprehensive benchmarks, we show that CLM-Access significantly outperforms existing methods in key downstream tasks, including batch effect correction, cell type annotation, RNA expression prediction, and multi-modal integration. This work establishes a scalable and interpretable foundation model for single-cell epigenomic analysis and expands the application of CLMs in single-cell research. Chulin Sha, Xiaolin Li 0001 |
AAAI | 7 |
| 2025 | QiMLP: Quantum-inspired Multilayer Perceptron with Strong Correlation Mining and Parameter CompressionabstractMultilayer Perceptron (MLP) is a simple practice of Neural Network (NN) and the cornerstone of research and development of deep learning. Each neuron is connected to all neurons in the previous layer and implements a non-linear mapping through activation functions. MLP can learn complex non-linear relationships among features through the superposition of multiple hidden layers, but it still cannot discover the inherent strong correlation among features. The reason is that each neuron uses a simple weighted summation method to organize all the neurons in the previous layer. Inspired by quantum theory, this paper builds a non-linear NN layer that can mine strong correlations among features based on multi-body quantum systems, and then constructs a multi-layer perceptron, called Quantum-inspired MLP (QiMLP). It is conceivable that QiMLP will have important inspirational significance in reshaping machine learning, deep learning and large language models. We theoretically analyzed the basis for QiMLP to mine strong correlations among features, and implemented experiments on multiple classic deep learning datasets. Experimental results verify that QiMLP not only learns strong correlations among features, but also significantly reduces the number of parameters with hundreds of times improvement. Tianheng Wang, Pengju Yan, Xiaolin Li 0001 |
AAAI | 5 |
| 2025 | RNADiffFold: generative RNA secondary structure prediction using discrete diffusion modelsabstractRibonucleic acid (RNA) molecules are essential macromolecules that perform diverse biological functions in living beings. Precise prediction of RNA secondary structures is instrumental in deciphering their complex three-dimensional architecture and functionality. Traditional methodologies for RNA structure prediction, including energy-based and learning-based approaches, often depict RNA secondary structures from a static perspective and rely on stringent a priori constraints. Inspired by the success of diffusion models, in this work, we introduce RNADiffFold, an innovative generative prediction approach of RNA secondary structures based on multinomial diffusion. We reconceptualize the prediction of contact maps as akin to pixel-wise segmentation and accordingly train a denoising model to refine the contact maps starting from a noise-infused state progressively. We also devise a potent conditioning mechanism that harnesses features extracted from RNA sequences to steer the model toward generating an accurate secondary structure. These features encompass one-hot encoded sequences, probabilistic maps generated from a pre-trained scoring network, and embeddings and attention maps derived from RNA foundation model. Experimental results on both within- and cross-family datasets demonstrate RNADiffFold's competitive performance compared with current state-of-the-art methods. Additionally, RNADiffFold has shown a notable proficiency in capturing the dynamic aspects of RNA structures, a claim corroborated by its performance on datasets comprising multiple conformations. Zhen Wang 0056, Yizhen Feng, Qingwen Tian, Pengju Yan, Xiaolin Li 0001 |
Briefings Bioinform. | 6 |
| 2025 | M4: Multi-proxy multi-gate mixture of experts network for multiple instance learning in histopathology image analysisabstractMultiple instance learning (MIL) has been successfully applied for whole slide images (WSIs) analysis in computational pathology, enabling a wide range of prediction tasks from tumor subtyping to inferring genetic mutations and multi-omics biomarkers. However, existing MIL methods predominantly focus on single-task learning, resulting in not only overall low efficiency but also the overlook of inter-task relatedness. To address these issues, we proposed an adapted architecture of Multi-gate Mixture-of-experts with Multi-proxy for Multiple instance learning (M4), and applied this framework for simultaneous prediction of multiple genetic mutations from WSIs. The proposed M4 model has two main innovations: (1) adopting a multi-gate mixture-of-experts strategy for multiple genetic mutation simultaneous prediction on a single WSI; (2) introducing a multi-proxy CNN construction on the expert and gate networks to effectively and efficiently capture patch-patch interactions from WSI. Our model achieved significant improvements across five tested TCGA datasets in comparison to current state-of-the-art single-task methods. The code is available at: https://github.com/Bigyehahaha/M4. Ye Zhang 0028, Wen Shu, Pengju Yan, Xiaolin Li 0001, Chulin Sha |
Medical Image Anal. | 7 |
| 2024 | DrugMetric: quantitative drug-likeness scoring based on chemical space distanceabstractThe process of drug discovery is widely known to be lengthy and resource-intensive. Artificial Intelligence approaches bring hope for accelerating the identification of molecules with the necessary properties for drug development. Drug-likeness assessment is crucial for the virtual screening of candidate drugs. However, traditional methods like Quantitative Estimation of Drug-likeness (QED) struggle to distinguish between drug and non-drug molecules accurately. Additionally, some deep learning-based binary classification models heavily rely on selecting training negative sets. To address these challenges, we introduce a novel unsupervised learning framework called DrugMetric, an innovative framework for quantitatively assessing drug-likeness based on the chemical space distance. DrugMetric blends the powerful learning ability of variational autoencoders with the discriminative ability of the Gaussian Mixture Model. This synergy enables DrugMetric to identify significant differences in drug-likeness across different datasets effectively. Moreover, DrugMetric incorporates principles of ensemble learning to enhance its predictive capabilities. Upon testing over a variety of tasks and datasets, DrugMetric consistently showcases superior scoring and classification performance. It excels in quantifying drug-likeness and accurately distinguishing candidate drugs from non-drugs, surpassing traditional methods including QED. This work highlights DrugMetric as a practical tool for drug-likeness scoring, facilitating the acceleration of virtual drug screening, and has potential applications in other biochemical fields. Zhen Wang 0056, Yanxin Tao, Chulin Sha, Xiaolin Li 0001 |
Briefings Bioinform. | 7 |
| 2024 | BatmanNet: bi-branch masked graph transformer autoencoder for molecular representationabstractAlthough substantial efforts have been made using graph neural networks (GNNs) for artificial intelligence (AI)-driven drug discovery, effective molecular representation learning remains an open challenge, especially in the case of insufficient labeled molecules. Recent studies suggest that big GNN models pre-trained by self-supervised learning on unlabeled datasets enable better transfer performance in downstream molecular property prediction tasks. However, the approaches in these studies require multiple complex self-supervised tasks and large-scale datasets , which are time-consuming, computationally expensive and difficult to pre-train end-to-end. Here, we design a simple yet effective self-supervised strategy to simultaneously learn local and global information about molecules, and further propose a novel bi-branch masked graph transformer autoencoder (BatmanNet) to learn molecular representations. BatmanNet features two tailored complementary and asymmetric graph autoencoders to reconstruct the missing nodes and edges, respectively, from a masked molecular graph. With this design, BatmanNet can effectively capture the underlying structure and semantic information of molecules, thus improving the performance of molecular representation. BatmanNet achieves state-of-the-art results for multiple drug discovery tasks, including molecular properties prediction, drug-drug interaction and drug-target interaction, on 13 benchmark datasets, demonstrating its great potential and superiority in molecular representation learning. Zhen Wang 0056, Zheng Feng, Yanjun Li 0005, Yongrui Wang, Chulin Sha, Xiaolin Li 0001 |
Briefings Bioinform. | 8 |
| 2024 | AptaDiff: de novo design and optimization of aptamers based on diffusion modelsabstractAptamers are single-stranded nucleic acid ligands, featuring high affinity and specificity to target molecules. Traditionally they are identified from large DNA/RNA libraries using $in vitro$ methods, like Systematic Evolution of Ligands by Exponential Enrichment (SELEX). However, these libraries capture only a small fraction of theoretical sequence space, and various aptamer candidates are constrained by actual sequencing capabilities from the experiment. Addressing this, we proposed AptaDiff, the first in silico aptamer design and optimization method based on the diffusion model. Our Aptadiff can generate aptamers beyond the constraints of high-throughput sequencing data, leveraging motif-dependent latent embeddings from variational autoencoder, and can optimize aptamers by affinity-guided aptamer generation according to Bayesian optimization. Comparative evaluations revealed AptaDiff's superiority over existing aptamer generation methods in terms of quality and fidelity across four high-throughput screening data targeting distinct proteins. Moreover, surface plasmon resonance experiments were conducted to validate the binding affinity of aptamers generated through Bayesian optimization for two target proteins. The results unveiled a significant boost of $87.9\%$ and $60.2\%$ in RU values, along with a 3.6-fold and 2.4-fold decrease in KD values for the respective target proteins. Notably, the optimized aptamers demonstrated superior binding affinity compared to top experimental candidates selected through SELEX, underscoring the promising outcomes of our AptaDiff in accelerating the discovery of superior aptamers. Zhen Wang 0056, Yanjun Li 0005, Yizhen Feng, Shaokang Lv, Han Diao, Zhaofeng Luo, Pengju Yan, Xiaolin Li 0001 |
Briefings Bioinform. | 11 |
| 2024 | A Coarse-Fine Collaborative Learning Model for Three Vessel Segmentation in Fetal Cardiac Ultrasound ImagesabstractCongenital heart disease (CHD) is the most frequent birth defect and a leading cause of infant mortality, emphasizing the crucial need for its early diagnosis. Ultrasound is the primary imaging modality for prenatal CHD screening. As a complement to the four-chamber view, the three-vessel view (3VV) plays a vital role in detecting anomalies in the great vessels. However, the interpretation of fetal cardiac ultrasound images is subjective and relies heavily on operator experience, leading to variability in CHD detection rates, particularly in resource-constrained regions. In this study, we propose an automated method for segmenting the pulmonary artery, ascending aorta, and superior vena cava in the 3VV using a novel deep learning network named CoFi-Net. Our network incorporates a coarse-fine collaborative strategy with two parallel branches dedicated to simultaneous global localization and fine segmentation of the vessels. The coarse branch employs a partial decoder to leverage high-level semantic features, enabling global localization of objects and suppression of irrelevant structures. The fine branch utilizes attention-parameterized skip connections to improve feature representations and improve boundary information. The outputs of the two branches are fused to generate accurate vessel segmentations. Extensive experiments conducted on a collected dataset demonstrate the superiority of CoFi-Net compared to state-of-the-art segmentation models for 3VV segmentation, indicating its great potential for enhancing CHD diagnostic efficiency in clinical practice. Furthermore, CoFi-Net outperforms other deep learning models in breast lesion segmentation on a public breast ultrasound dataset, despite not being specifically designed for this task, demonstrating its potential and robustness for various segmentation tasks. Shan Ling, Laifa Yan, Rongsong Mao, Jizhou Li, Haoran Xi, Fei Wang 0138, Xiaolin Li 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Communication-Efficient and Attack-Resistant Federated Edge Learning With Dataset DistillationabstractFederated Edge Learning considers a large amount of distributed edge nodes collectively train a global gradient-based model for edge computing in the Artificial Internet of Things, which significantly promotes the development of cloud computing. However, current federated learning algorithms take tens of communication rounds transmitting unwieldy model weights under ideal circumstances and hundreds when data is poorly distributed. This drawback directly results in expensive communication overhead for edge devices. Inspired by recent work on dataset distillation and distributed one-shot learning, we propose Distilled One-Shot Federated Learning (DOSFL) to significantly reduce the communication cost while achieving comparable performance. In just one round, each client distills their private dataset, sends the synthetic data to the server, and collectively trains a global model. The distilled data look like noise and are only useful to the specific model weights,i.e.,become useless after the model updates. With this weight-less and gradient-less design, the total communication cost of DOSFL is up to three orders of magnitude less than FedAvg while preserving up to 99% performance of centralized training on both vision and language tasks with different models including CNN, LSTM, Transformer,etc. We demonstrate that an eavesdropping attacker cannot properly train a good model using the leaked distilled data, without knowing the initial model weights. DOSFL serves as an inexpensive method to quickly converge on a performant pre-trained model with less than 0.1% communication cost of traditional methods. Xiyao Ma, Dapeng Oliver Wu, Xiaolin Li 0001 |
IEEE Trans. Cloud Comput. | 4 |
| 2022 | Deep Learning in Drug Design: Protein-Ligand Binding Affinity PredictionabstractComputational drug design relies on the calculation of binding strength between two biological counterparts especially a chemical compound, i.e., a ligand, and a protein. Predicting the affinity of protein-ligand binding with reasonable accuracy is crucial for drug discovery, and enables the optimization of compounds to achieve better interaction with their target protein. In this paper, we propose a data-driven framework named DeepAtom to accurately predict the protein-ligand binding affinity. With 3D Convolutional Neural Network (3D-CNN) architecture, DeepAtom could automatically extract binding related atomic interaction patterns from the voxelized complex structure. Compared with the other CNN based approaches, our light-weight model design effectively improves the model representational capacity, even with the limited available training data. We carried out validation experiments on the PDBbind v.2016 benchmark and the independent Astex Diverse Set. We demonstrate that the less feature engineering dependent DeepAtom approach consistently outperforms the other baseline scoring methods. We also compile and propose a new benchmark dataset to further improve the model performances. With the new dataset as training input, DeepAtom achieves Pearson's R=0.83 and RMSE=1.23 pK units on the PDBbind v.2016 core set. The promising results demonstrate that DeepAtom models can be potentially adopted in computational drug development protocols such as molecular docking and virtual screening. Mohammad A. Rezaei, Yanjun Li 0005, Dapeng Oliver Wu, Xiaolin Li 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2022 | A Praise for Defensive Programming: Leveraging Uncertainty for Effective Malware MitigationabstractA promising avenue for improving the effectiveness of behavioral-based malware detectors is to leverage two-phase detection mechanisms. Existing problem in two-phase detection is that after the first phase produces borderline decision, suspicious behaviors are not well contained before the second phase completes. This article improvesChameleon, a framework to realize the uncertain environment.Chameleonoffers two environments: standard—for software identified as benign by the first phase, and uncertain—for software received borderline classification from the first phase. The uncertain environment adds obstacles to software execution through random perturbations applied probabilistically. We introduce a dynamic perturbation threshold that can target malware disproportionately more than benign software. We analyzed the effects of the uncertain environment by manually studying 113 software and 100 malware, and found that 92 percent malware and 10 percent benign software disrupted during execution. The results were then corroborated by an extended dataset (5,679 Linux malware samples) on a newer system. Finally, a careful inspection of the benign software crashes revealed some software bugs, highlightingChameleon's potential as a practical complementary anti-malware solution. Ruimin Sun, Marcus Botacin, Nikolaos Sapountzis, Xiaoyong Yuan, Matt Bishop, Donald E. Porter, Xiaolin Li 0001, André Ricardo Abed Grégio, Daniela Oliveira 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2022 | Learning Fast and Slow: Propedeutica for Real-Time Malware DetectionabstractExisting malware detectors on safety-critical devices have difficulties in runtime detection due to the performance overhead. In this article, we introduce Propedeutica, a framework for efficient and effective real-time malware detection, leveraging the best of conventional machine learning (ML) and deep learning (DL) techniques. In Propedeutica, all software start executions are considered as benign and monitored by a conventional ML classifier for fast detection. If the software receives a borderline classification from the ML detector (e.g., the software is 50% likely to be benign and 50% likely to be malicious), the software will be transferred to a more accurate, yet performance demanding DL detector. To address spatial-temporal dynamics and software execution heterogeneity, we introduce a novel DL architecture (DeepMalware) for Propedeutica with multistream inputs. We evaluated Propedeutica with 9115 malware samples and 1338 benign software from various categories for the Windows OS. With a borderline interval of [30%, 70%], Propedeutica achieves an accuracy of 94.34% and a false-positive rate of 8.75%, with 41.45% of the samples moved for DeepMalwareanalysis. Even using only CPU, Propedeutica can detect malware within less than 0.1 s. Ruimin Sun, Xiaoyong Yuan, Pan He, Qile Zhu, Aokun Chen, André Ricardo Abed Grégio, Daniela Oliveira 0001, Xiaolin Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2020 | Improving Question Generation with Sentence-Level Semantic Matching and Answer Position InferringabstractTaking an answer and its context as input, sequence-to-sequence models have made considerable progress on question generation. However, we observe that these approaches often generate wrong question words or keywords and copy answer-irrelevant words from the input. We believe that lacking global question semantics and exploiting answer position-awareness not well are the key root causes. In this paper, we propose a neural question generation model with two general modules: sentence-level semantic matching and answer position inferring. Further, we enhance the initial state of the decoder by leveraging the answer-aware gated fusion mechanism. Experimental results demonstrate that our model outperforms the state-of-the-art (SOTA) models on SQuAD and MARCO datasets. Owing to its generality, our work also improves the existing models significantly. Xiyao Ma, Qile Zhu, Xiaolin Li 0001 |
AAAI | 4 |
| 2020 | A Batch Normalized Inference Network Keeps the KL Vanishing AwayabstractVariational Autoencoder (VAE) is widely used as a generative model to approximate a model's posterior on latent variables by combining the amortized variational inference and deep neural networks.However, when paired with strong autoregressive decoders, VAE often converges to a degenerated local optimum known as "posterior collapse".Previous approaches consider the Kullback-Leibler divergence (KL) individual for each datapoint.We propose to let the KL follow a distribution across the whole dataset, and analyze that it is sufficient to prevent posterior collapse by keeping the expectation of the KL's distribution positive.Then we propose Batch Normalized-VAE (BN-VAE), a simple but effective approach to set a lower bound of the expectation by regularizing the distribution of the approximate posterior's parameters.Without introducing any new model component or modifying the objective, our approach can avoid the posterior collapse effectively and efficiently.We further show that the proposed BN-VAE can be extended to conditional VAE (CVAE).Empirically, our approach surpasses strong autoregressive baselines on language modeling, text classification and dialogue generation, and rivals more complex approaches while keeping almost the same training time as VAE. Qile Zhu, Wei Bi, Xiaojiang Liu, Xiyao Ma, Xiaolin Li 0001, Dapeng Oliver Wu |
ACL | 5 |
| 2020 | Connecting Web Event Forecasting with Anomaly Detection: A Case Study on Enterprise Web Applications Using Self-supervised Neural Networks
Xiaoyong Yuan, Lei Ding 0003, Xiaolin Li 0001, Dapeng Oliver Wu |
SecureComm (1) | 4 |
| 2019 | Generalized Batch Normalization: Towards Accelerating Deep Neural NetworksabstractUtilizing recently introduced concepts from statistics and quantitative risk management, we present a general variant of Batch Normalization (BN) that offers accelerated convergence of Neural Network training compared to conventional BN. In general, we show that mean and standard deviation are not always the most appropriate choice for the centering and scaling procedure within the BN transformation, particularly if ReLU follows the normalization step. We present a Generalized Batch Normalization (GBN) transformation, which can utilize a variety of alternative deviation measures for scaling and statistics for centering, choices which naturally arise from the theory of generalized deviation measures and risk theory in general. When used in conjunction with the ReLU non-linearity, the underlying risk theory suggests natural, arguably optimal choices for the deviation measure and statistic. Utilizing the suggested deviation measure and statistic, we show experimentally that training is accelerated more so than with conventional BN, often with improved error rate as well. Overall, we propose a more flexible BN transformation supported by a complimentary theoretical framework that can potentially guide design choices. Xiaoyong Yuan, Zheng Feng, Matthew Norton 0001, Xiaolin Li 0001 |
AAAI | 4 |
| 2019 | DeepAtom: A Framework for Protein-Ligand Binding Affinity PredictionabstractThe cornerstone of computational drug design is the calculation of binding affinity between two biological counterparts especially a chemical compound, i.e. a ligand, and a protein. Predicting the strength of protein-ligand binding with reasonable accuracy is critical for drug discovery. In this paper, we propose a data-driven framework named DeepAtom to accurately predict the protein-ligand binding affinity. With 3D Convolutional Neural Network (3D-CNN) architecture, DeepAtom could automatically extract binding related atomic interaction patterns from the voxelized complex structure. Compared with the other CNN based approaches, our light-weight model design effectively improves the model representational capacity, even with the limited available training data. With validation experiments on the PDBbind v.2016 benchmark and the independent Astex Diverse Set, we demonstrate that the less feature engineering dependent DeepAtom approach consistently outperforms the other state-of-the-art scoring methods. We also compile and propose a new benchmark dataset to further improve the model performances. With the new dataset as training input, DeepAtom achieves Pearson's$\mathrm{R}=0.83$and$\text{RMSE}=1.23\ \ pK$units on the PDBbind v.2016 core set. The promising results demonstrate that DeepAtom models can be potentially adopted in computational drug development protocols such as molecular docking and virtual screening. Yanjun Li 0005, Mohammad A. Rezaei, Xiaolin Li 0001 |
BIBM | 4 |
| 2019 | Adaptive Leader-Follower Formation Control and Obstacle Avoidance via Deep Reinforcement LearningabstractWe propose a deep reinforcement learning (DRL) methodology for the tracking, obstacle avoidance, and formation control of nonholonomic robots. By separating vision-based control into a perception module and a controller module, we can train a DRL agent without sophisticated physics or 3D modeling. In addition, the modular framework averts daunting retrains of an image-to-action end-to-end neural network, and provides flexibility in transferring the controller to different robots. First, we train a convolutional neural network (CNN) to accurately localize in an indoor setting with dynamic foreground/background. Then, we design a new DRL algorithm named Momentum Policy Gradient (MPG) for continuous control tasks and prove its convergence. We also show that MPG is robust at tracking varying leader movements and can naturally be extended to problems of formation control. Leveraging reward shaping, features such as collision and obstacle avoidance can be easily integrated into a DRL controller. George Pu, Xiyao Ma, Runhan Sun, Hsi-Yuan Chen, Xiaolin Li 0001 |
IROS | 7 |
| 2019 | Adversarial Examples: Attacks and Defenses for Deep LearningabstractWith rapid progress and significant successes in a wide spectrum of applications, deep learning is being applied in many safety-critical environments. However, deep neural networks (DNNs) have been recently found vulnerable to well-designed input samples called adversarial examples. Adversarial perturbations are imperceptible to human but can easily fool DNNs in the testing/deploying stage. The vulnerability to adversarial examples becomes one of the major risks for applying DNNs in safety-critical environments. Therefore, attacks and defenses on adversarial examples draw great attention. In this paper, we review recent findings on adversarial examples for DNNs, summarize the methods for generating adversarial examples, and propose a taxonomy of these methods. Under the taxonomy, applications for adversarial examples are investigated. We further elaborate on countermeasures for adversarial examples. In addition, three major challenges in adversarial examples and the potential solutions are discussed. Xiaoyong Yuan, Pan He, Qile Zhu, Xiaolin Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | GraphBTM: Graph Enhanced Autoencoded Variational Inference for Biterm Topic ModelabstractDiscovering the latent topics within texts has been a fundamental task for many applications.However, conventional topic models suffer different problems in different settings.The Latent Dirichlet Allocation (LDA) may not work well for short texts due to the data sparsity (i.e., the sparse word co-occurrence patterns in short documents).The Biterm Topic Model (BTM) learns topics by modeling the word-pairs named biterms in the whole corpus.This assumption is very strong when documents are long with rich topic information and do not exhibit the transitivity of biterms.In this paper, we propose a novel way called GraphBTM to represent biterms as graphs and design Graph Convolutional Networks (GCNs) with residual connections to extract transitive features from biterms.To overcome the data sparsity of LDA and the strong assumption of BTM, we sample a fixed number of documents to form a mini-corpus as a training instance.We also propose a dataset called All N ews extracted from (Thompson, 2017), in which documents are much longer than 20 Newsgroups.We present an amortized variational inference method for GraphBTM.Our method generates more coherent topics compared with previous approaches.Experiments show that the sampling strategy improves performance by a large margin. Qile Zhu, Zheng Feng, Xiaolin Li 0001 |
EMNLP | 3 |
| 2018 | GRAM-CNN: a deep learning approach with local context for named entity recognition in biomedical textabstractMotivation: Best performing named entity recognition (NER) methods for biomedical literature are based on hand-crafted features or task-specific rules, which are costly to produce and difficult to generalize to other corpora. End-to-end neural networks achieve state-of-the-art performance without hand-crafted features and task-specific knowledge in non-biomedical NER tasks. However, in the biomedical domain, using the same architecture does not yield competitive performance compared with conventional machine learning models. Results: We propose a novel end-to-end deep learning approach for biomedical NER tasks that leverages the local contexts based on n-gram character and word embeddings via Convolutional Neural Network (CNN). We call this approach GRAM-CNN. To automatically label a word, this method uses the local information around a word. Therefore, the GRAM-CNN method does not require any specific knowledge or feature engineering and can be theoretically applied to a wide range of existing NER problems. The GRAM-CNN approach was evaluated on three well-known biomedical datasets containing different BioNER entities. It obtained an F1-score of 87.26% on the Biocreative II dataset, 87.26% on the NCBI dataset and 72.57% on the JNLPBA dataset. Those results put GRAM-CNN in the lead of the biological NER methods. To the best of our knowledge, we are the first to apply CNN based structures to BioNER problems. Availability and implementation: The GRAM-CNN source code, datasets and pre-trained model are available online at: https://github.com/valdersoul/GRAM-CNN. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Qile Zhu, Xiaolin Li 0001, Ana Conesa, Cécile Pereira |
Bioinform. | 2 |
| 2018 | Maximizing positive influence spread in online social networks via fluid dynamics
Feng Wang 0051, Xiaolin Li 0001, Guojun Wang 0001 |
Future Gener. Comput. Syst. | 3 |
| 2018 | Enhancing Localization Scalability and Accuracy via Opportunistic Sensing
Xiaolin Li 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2017 | Single Shot Text Detector with Regional AttentionabstractWe present a novel single-shot text detector that directly outputs word-level bounding boxes in a natural image. We propose an attention mechanism which roughly identifies text regions via an automatically learned attentional map. This substantially suppresses background interference in the convolutional features, which is the key to producing accurate inference of words, particularly at extremely small sizes. This results in a single model that essentially works in a coarse-to-fine manner. It departs from recent FCN-based text detectors which cascade multiple FCN models to achieve an accurate prediction. Furthermore, we develop a hierarchical inception module which efficiently aggregates multi-scale inception features. This enhances local details, and also encodes strong context information, allowing the detector to work reliably on multi-scale and multi-orientation text with single-scale images. Our text detector achieves an F-measure of 77% on the ICDAR 2015 benchmark, advancing the state-of-the-art results in [18, 28]. Demo is available at: http://sstd.whuang.org/. Pan He, Tong He 0001, Qile Zhu, Yu Qiao 0001, Xiaolin Li 0001 |
ICCV | 6 |
| 2017 | DeepDefense: Identifying DDoS Attack via Deep LearningabstractDistributed Denial of Service (DDoS) attacks grow rapidly and become one of the fatal threats to the Internet. Automatically detecting DDoS attack packets is one of the main defense mechanisms. Conventional solutions monitor network traffic and identify attack activities from legitimate network traffic based on statistical divergence. Machine learning is another method to improve identifying performance based on statistical features. However, conventional machine learning techniques are limited by the shallow representation models. In this paper, we propose a deep learning based DDoS attack detection approach (DeepDefense). Deep learning approach can automatically extract high-level features from low-level ones and gain powerful representation and inference. We design a recurrent deep neural network to learn patterns from sequences of network traffic and trace network attack activities. The experimental results demonstrate a better performance of our model compared with conventional machine learning models. We reduce the error rate from 7.517% to 2.103% compared with conventional machine learning method in the larger data set. Xiaoyong Yuan, Chuanhuang Li, Xiaolin Li 0001 |
SMARTCOMP | 3 |
| 2017 | My Privacy My Decision: Control of Photo Sharing on Online Social NetworksabstractPhoto sharing is an attractive feature which popularizes online social networks (OSNs). Unfortunately, it may leak users' privacy if they are allowed to post, comment, and tag a photo freely. In this paper, we attempt to address this issue and study the scenario when a user shares a photo containing individuals other than himself/herself (termed co-photo for short). To prevent possible privacy leakage of a photo, we design a mechanism to enable each individual in a photo be aware of the posting activity and participate in the decision making on the photo posting. For this purpose, we need an efficient facial recognition (FR) system that can recognize everyone in the photo. However, more demanding privacy setting may limit the number of the photos publicly available to train the FR system. To deal with this dilemma, our mechanism attempts to utilize users' private photos to design a personalized FR system specifically trained to differentiate possible photo co-owners without leaking their privacy. We also develop a distributed consensus-based method to reduce the computational complexity and protect the private training set. We show that our system is superior to other possible approaches in terms of recognition ratio and efficiency. Our mechanism is implemented as a proof of concept Android application on Facebook's platform. Kaihe Xu, Yuanxiong Guo, Linke Guo, Yuguang Fang, Xiaolin Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2016 | Enhancing Smartphone Indoor Localization via Opportunistic SensingabstractUsing a mobile phone for fine-grained indoor localization remains an open problem. Low-complexity approaches without infrastructure could not achieve accurate and reliable results due to various restrictions. Accurate solutions relying on dense anchor nodes are inconvenient and cumbersome in deployment. The anchor blockage problem would further reduce the effective coverages. In this paper, we investigate the problems associated with improving indoor localization of a mobile phone via opportunistic anchor sensing, a new sensing paradigm leveraging multiple anchors without minimum number or constellation requirement. One key motivation is that the location results could be improved by exploring more data types rather than deploying more anchor nodes. To enable this high scalability and accuracy design, we leverage low-coupling hybrid ranging by our low cost anchor nodes with centimeter-level relative distance estimation. Activity pattern extracted in local smartphone is utilized for accurate displacement and direction estimation. Finer localization resolution could be achieved with sufficient anchor access. We introduce the delay-constraint robust semidefinite programming in trilateration calculation with the potential of centimeter-level location resolution. We conduct extensive experiments in various scenarios. Compared with other approaches, opportunistic sensing could improve the location accuracy, scalability as well as robustness under various anchor accessibilities. Di Wu 0002, Xiaolin Li 0001 |
SECON | 3 |
| 2016 | Guoguo: Enabling Fine-Grained Smartphone Localization via Acoustic AnchorsabstractModern smartphones and location-based services and apps are poised to transform our daily life. However, current smartphone-based localization solutions are limited mainly to outdoor, mostly missing practical, robust and accurate indoor location solutions. Despite significant efforts on indoor localization in both academia and industry in the past two decades, highly accurate and practical smartphone-based indoor localization remains an open problem. To enable indoor location-based services (ILBS), e.g., step-by-step navigation for the Blind and visually impaired, there are several stringent requirements: highly accurate (foot-level); no additional hardware components or extensions on users' smartphones; scalable to massive concurrent users. Current GPS, Radio RSS (e.g. Wi-Fi, Bluetooth, ZigBee), or Fingerprinting based solutions can only achieve meter-level or room-level accuracy. In this paper, we propose a practical and accurate solution that fills the long-lasting gap of smartphone-based fine-grained indoor localization. Specifically, we design and implement an indoor localization ecosystem Guoguo. Guoguo consists of an anchor network with a coordination protocol to transmit modulated localization beacons using high-band acoustic signals, a realtime processing app in a smartphone, and a backend server for indoor contexts and location-based services. We further propose approaches to improve its coverage, accuracy, and location update rate with low-power consumption. Our prototype shows centimeter-level localization accuracy in several typical indoor environments. Such precise indoor localization is expected to have high impact in the future ILBS and our daily activities. Xinxin Liu 0006, Xiaolin Li 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2015 | Hiding Media Data via Shaders: Enabling Private Sharing in the CloudsabstractIn the era of Cloud and Social Networks, mobile devices exhibit much more powerful abilities for big media data storage and sharing. However, many users are still reluctant to share/store their data via clouds due to the potential leakage of confidential or private information. Although some cloud services provide storage encryption and access protection, privacy risks are still high since the protection is not always adequately conducted from end-to-end. Most customers are aware of the danger of letting data control out of their hands, e.g., Storing them to YouTube, Flickr, Facebook, Google+. Because of substantial practical and business needs, existing cloud services are restricted to the desired formats, e.g., Video and photo, without allowing arbitrary encrypted data. In this paper, we propose a format-compliant end-to-end privacy-preserving scheme for media sharing/storage issues with considerations for big data, clouds, and mobility. To realize efficient encryption for big media data, we jointly achieve format-compliant, compression-independent and correlation-preserving via multi-channel chained solutions under the guideline of Markov cipher. The encryption and decryption process is integrated into an image/video filter via GPU Shader for display-to-display full encryption. The proposed scheme makes big media data sharing/storage safer and easier in the clouds. Min Li 0012, Xiaolin Li 0001 |
CLOUD | 3 |
| 2015 | Taming Non-local Stragglers Using Efficient Prefetching in MapReduceabstractMapReduce has been widely adopted as a programming model to process big data. However, parallel jobs in MapReduce are prone to be plagued by stragglers caused by non-local tasks for two reasons: first, system logs from production clusters show that a non-local task can be two times slower than a local task; second, a job's completion time is bottlenecked by its slowest parallel tasks. As a result, even one single non-local task can become the straggler of the whole job, causing significant delay of the whole job. In this paper, we propose to alleviate this problem by proactively prefetching input data for non-local tasks. However, performing such prefetching efficiently in MapReduce is difficult, because it requires both application-level information to generate accurate prefetching requests at runtime, and an appropriate network flow scheduling mechanism to guarantee the timeliness of prefetching flows. To address these challenges, we design and implement FlexFetch, which 1) leverages a novel mechanism called speculative scheduling to accurately generate prefetching flows, 2) explicitly allocates network resources to prefetching flows using a criticality-aware deadline-driven flow scheduling algorithm. We evaluate FlexFetch through both testbed experiments and large-scale simulations using production workloads. The results show that FlexFetch reduces the completion time by 41.8% for small jobs and 26.8% on average, compared with the default MapReduce implementation in Hadoop. Min Li 0012, Xin Yang 0006, Han Zhao 0001, Xiaolin Li 0001 |
CLUSTER | 5 |
| 2015 | Exploring Fine-Grained Resource Rental Planning in Cloud ComputingabstractApplication services based on cloud computing infrastructure are proliferating over the Internet. In this paper, we investigate the problem of how to minimize cloud resource rental cost associated with hosting such cloud-based application services, while meeting the projected service demand. This problem arises when applications generate high volume of data that incurs significant cost on storage and transfer. As a result, an application service provider (ASP) needs to carefully evaluate various resource rental options before finalizing the application deployment. We choose Amazon EC2 marketplace as a case of study, and analyze the economical trade-off for on-demand resource rental strategies. Given fixed resource pricing, we first develop a deterministic model, using a mixed integer linear program, to facilitate resource rental decision making. Evaluation results show that our planning optimization model reduces resource rental cost by as much as 50 percent compared with a baseline strategy. Next, we further investigate planning solutions to resource market featuring time-varying pricing (Amazon Spot Instance Market). We perform time-series analysis over the spot price trace and examine its predictability using auto-regressive integrated moving-average (ARIMA). We also develop a stochastic planning model based on multistage recourse. By comparing these two approaches, we discover that spot price forecasting does not provide our planning model with a crystal ball due to the weak correlation of past and future price, and the stochastic planning model better hedges against resource pricing uncertainty than resource rental planning using forecast prices. Han Zhao 0001, Miao Pan, Xinxin Liu 0006, Xiaolin Li 0001, Yuguang Fang |
IEEE Trans. Cloud Comput. | 4 |
| 2015 | Enabling Context-Aware Indoor Augmented Reality via Smartphone Sensing and Vision TrackingabstractAugmented reality (AR) aims to render the world that users see and overlay information that reflects the real physical dynamics. The digital view could be potentially projected near the Point-of-Interest (POI) in a way that makes the virtual view attached to the POI even when the camera moves. Achieving smooth support for movements is a subject of extensive studies. One of the key problems is where the augmented information should be added to the field of vision in real time. Existing solutions either leverage GPS location for rendering outdoor AR views (hundreds of kilometers away) or rely on image markers for small-scale presentation (only for the marker region). To realize AR applications under various scales and dynamics, we propose a suite of algorithms for fine-grained AR view tracking to improve the accuracy of attitude and displacement estimation, reduce the drift, eliminate the marker, and lower the computation cost. Instead of requiring extremely high, accurate, absolute locations, we propose multimodal solutions according to mobility levels without additional hardware requirement. Experimental results demonstrate significantly less error in projecting and tracking the AR view. These results are expected to make users excited to explore their surroundings with enriched content. Xiaolin Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2014 | Taming Computation Skews of Block-Oriented Iterative Scientific Applications in MapReduce SystemsabstractNowadays, scientists are embracing big data techniques for exploring significant discoveries from large volumes of scientific data quickly. Properly partitioning workloads is essential for fully exploiting the benefit of parallelism, but is difficult for applications whose computations change iteratively. Computation skews are inevitable when executing block-oriented iterative scientific applications in MapReduce systems. This paper proposes iPart, an autonomic workload partitioning system for taming computation skews of block-oriented iterative scientific applications in MapReduce systems. iPart introduces a workload control loop into the conventional execution of MapReduce jobs. Workload estimates in terms of execution time are collected in the reduce phase and fed back to the partition phase to update partitioning plans. Computation skews are detected and addressed by adapting partitioning to computation changes iteratively. Two adaptive partitioning methods based on the binary partitioning method are presented. Experimental evaluations with two simulated applications and the synthetic and real-world data prove that iPart responds to computation changes and adapts partitioning quickly and accurately. Xin Yang 0006, Min Li 0012, Xiaolin Li 0001 |
IEEE CLOUD | 4 |
| 2014 | Palantir: Reseizing Network Proximity in Large-Scale Distributed Computing Frameworks Using SDNabstractParallel/Distributed computing frameworks, such as MapReduce and Dryad, have been widely adopted to analyze massive data. Traditionally, these frameworks depend on manual configuration to acquire network proximity information to optimize the data placement and task scheduling. However, this approach is cumbersome, inflexible or even infeasible in largescale deployments, for example, across multiple datacenters. In this paper, we address this problem by utilizing the Software-Defined Networking (SDN) capability. We build Palantir, an SDN service specific for parallel/distributed computing frameworks to abstract the proximity information out of the network. Palantir frees the framework developers/ administrators from having to manually configure the network. In addition, Palantir is flexible because it allows different frameworks to define the proximity according to the framework-specific metrics. We design and implement a datacenter-aware MapReduce to demonstrate Palantir's usefullness. Our evaluation shows that, based on Palantir, datacenter-aware MapReduce achieves siginficant performance improvement. Min Li 0012, Xin Yang 0006, Xiaolin Li 0001 |
IEEE CLOUD | 4 |
| 2014 | Control of photo sharing over Online Social NetworksabstractPhoto sharing is an attractive feature which popularizes Online Social Networks (OSNs). Unfortunately, it may leak users' privacy if they are allowed to post, comment, and tag a photo freely. In this paper, we attempt to address this issue and study the scenario when a user shares a photo containing individuals other than himself/herself (termed co-photo for short). To prevent possible leakage of a photo privacy, we design a mechanism to enable each individual in a photo be aware of the posting activity and participate in the decision making on the photo posting. For this purpose, we need an efficient facial recognition (FR) system that can recognize everyone in the photo. However, more demanding privacy setting may limit the number of the photos publicly available to train the FR system. To deal with this dilemma, our mechanism attempts to utilize users' private photos to design a personalized FR system specifically trained to differentiate possible photo co-owners without leaking his/her privacy. We have also developed a distributed consensus-based method to not only reduce the computational complexity, but also preserve the privacy during the training. We show that our system is superior to other possible approaches in terms of recognition ratio and efficiency. Our mechanism is implemented as an Android application on Facebook's platform. Kaihe Xu, Yuanxiong Guo, Linke Guo, Yuguang Fang, Xiaolin Li 0001 |
GLOBECOM | 5 |
| 2014 | Finding Nemo: Finding Your Lost Child in Crowds via Mobile Crowd SensingabstractMobile Crowd Sourcing/Sensing (MCS), as a new paradigm for participatory sensing, is suitable for large-scale hard tasks that are costly, or infeasible with conventional methods. Utilizing the ubiquitousness of "crowds" of sensor-rich smartphones, MCS has enormous potential to truly unleash the power of collaborative locating and searching at a societal scale. In this paper, we target the application of finding and locating the lost child in crowds via MCS. Conventional localization approaches require fixed anchor networks or fingerprinting points as references. It is not effective for locating the child in open and uncontrolled areas. We propose MCS-based collaborative localization via nearby opportunistically connected participators. To obtain sufficient measurements, we utilize one-hop and multi-hop assistants to reach more participators. Semidefinite Programming (SDP) based global optimization approaches are proposed to leverage all the location and ranging measurements in a best-effort way. We conduct extensive experiments and simulations in various scenarios. Compared with other classic algorithms, our proposed approach achieves significant accuracy improvement and could locate the "unlocalizable" child. Xiaolin Li 0001 |
MASS | 2 |
| 2014 | Towards efficient and fair resource trading in community-based cloud computing
Han Zhao 0001, Xinxin Liu 0006, Xiaolin Li 0001 |
J. Parallel Distributed Comput. | 3 |
| 2013 | A game-theoretic approach for achieving k-anonymity in Location Based ServicesabstractLocation Based Service (LBS), although it greatly benefits the daily life of mobile device users, has introduced significant threats to privacy. In an LBS system, even under the protection of pseudonyms, users may become victims of inference attacks, where an adversary reveals a user's real identity and complete moving trajectory with the aid of side information, e.g., accidental identity disclosure through personal encounters. To enhance privacy protection for LBS users, a common approach is to include extra fake location information associated with different pseudonyms, known as dummy users, in normal location reports. Due to the high cost of dummy generation using resource constrained mobile devices, self-interested users may free-ride on others' efforts. The presence of such selfish behaviors may have an adverse effect on privacy protection. In this paper, we study the behaviors of self-interested users in the LBS system from a game-theoretic perspective. We model the distributed dummy user generation as Bayesian games in both static and timing-aware contexts, and analyze the existence and properties of the Bayesian Nash Equilibria for both models. Based on the analysis, we propose a strategy selection algorithm to help users achieve optimized payoffs. Leveraging a beta distribution generalized from real-world location privacy data traces, we perform simulations to assess the privacy protection effectiveness of our approach. The simulation results validate our theoretical analysis for the dummy user generation game models. Xinxin Liu 0006, Linke Guo, Xiaolin Li 0001, Yuguang Fang |
INFOCOM | 4 |
| 2013 | Towards accurate acoustic localization on a smartphoneabstractSince our daily activities are dominantly indoor, as smart phones emerge as the most popular personal computing companions, major IT companies recently launched aggressive investment on mobile indoor location services and positioning systems, e.g., on iOS or Android mobile devices. However, one major hurdle has not been conquered yet: smart phone-based high-resolution indoor localization. In this paper, we propose a practical solution for accurate ranging and localization based on acoustic communication between anchor nodes with speakers and the microphone on a smartphone. To identify different anchor nodes and enable time-of-arrival (TOA) ranging, we propose approaches for signal modulation, symbol detection and demodulation, synchronization and ranging. Experimental results show that the communication bit-error-rate and ranging accuracy is sufficient for our target applications. The preliminary results of localization demonstrate that our algorithm could achieve highaccuracy of 23cm in the offline mode with a promising potential for realtime smartphone-based indoor localization. Xinxin Liu 0006, Lulu Xie, Xiaolin Li 0001 |
INFOCOM | 4 |
| 2013 | Improving GPS Service via Social CollaborationabstractThe popularity of GPS-enabled smartphones enables a wide variety of new location-based or location-aware services and applications. However, the GPS module in a smartphone produces inaccurate position estimates and incurs high energy consumption, which inhibits the wide use of location-aware applications. To address this, we propose a social-aided cooperative location optimization (Coloc) scheme, which is capable of improving positioning accuracy and achieving low energy consumption. Specifically, our scheme enhances positioning accuracy by fusing the GPS positions of multiple co-located smartphones in a social network, or by neighborhood-based weighted least-squares estimation when relative distances between smartphones are available. The energy efficiency is achieved by sharing location information among co-located users and lower the update rate of the GPS module without sacrificing the accuracy. To validate our proposed approach, we conduct experiments in stationary and moving scenarios. Experimental results show that our proposed cooperative localization scheme can achieve sufficient performance gains in both indoor and outdoor environments. Qiuyuan Huang, Jiecong Wang, Xiaolin Li 0001, Dapeng Oliver Wu |
MASS | 4 |
| 2013 | Guoguo: enabling fine-grained indoor localization via smartphoneabstractUsing smartphones for accurate indoor localization opens a new frontier of mobile services, offering enormous opportunities to enhance users' experiences in indoor environments. Despite significant efforts on indoor localization in both academia and industry in the past two decades, highly accurate and practical smartphone-based indoor localization remains an open problem. To enable indoor location-based services (ILBS), there are several stringent requirements for an indoor localization system: highly accurate that can differentiate massive users' locations (foot-level); no additional hardware components or extensions on users' smartphones; scalable to massive concurrent users. Current GPS, Radio RSS (e.g. WiFi, Bluetooth, ZigBee), or Fingerprinting based solutions can only achieve meter-level or room-level accuracy. In this paper, we propose a practical and accurate solution that fills the long-lasting gap of smartphone-based indoor localization. Specifically, we design and implement an indoor localization ecosystem Guoguo. Guoguo consists of an anchor network with a coordination protocol to transmit modulated localization beacons using high-band acoustic signals, a realtime processing app in a smartphone, and a backend server for indoor contexts and location-based services. We further propose approaches to improve its coverage, accuracy, and location update rate with low-power consumption. Our prototype shows centimeter-level localization accuracy in an office and classroom environment. Such precise indoor localization is expected to have high impact in the future ILBS and our daily activities. Xinxin Liu 0006, Xiaolin Li 0001 |
MobiSys | 3 |
| 2013 | Special issue of JCSS on UbiSafe computing and communications
Guojun Wang 0001, Jianhua Ma 0002, Xiaolin Li 0001, Athanasios V. Vasilakos |
J. Comput. Syst. Sci. | 3 |
| 2013 | SinkTrail: A Proactive Data Reporting Protocol for Wireless Sensor NetworksabstractIn large-scale Wireless Sensor Networks (WSNs), leveraging data sinks' mobility for data gathering has drawn substantial interests in recent years. Current researches either focus on planning a mobile sink's moving trajectory in advance to achieve optimized network performance, or target at collecting a small portion of sensed data in the network. In many application scenarios, however, a mobile sink cannot move freely in the deployed area. Therefore, the precalculated trajectories may not be applicable. To avoid constant sink location update traffics when a sink's future locations cannot be scheduled in advance, we propose two energy-efficient proactive data reporting protocols, SinkTrail and SinkTrail-S, for mobile sink-based data collection. The proposed protocols feature low-complexity and reduced control overheads. Two unique aspects distinguish our approach from previous ones: 1) we allow sufficient flexibility in the movement of mobile sinks to dynamically adapt to various terrestrial changes; and 2) without requirements of GPS devices or predefined landmarks, SinkTrail establishes a logical coordinate system for routing and forwarding data packets, making it suitable for diverse application scenarios. We systematically analyze the impact of several design factors in the proposed algorithms. Both theoretical analysis and simulation results demonstrate that the proposed algorithms reduce control overheads and yield satisfactory performance in finding shorter routing paths. Xinxin Liu 0006, Han Zhao 0001, Xin Yang 0006, Xiaolin Li 0001 |
IEEE Trans. Computers | 4 |
| 2013 | VectorTrust: trust vector aggregation scheme for trust management in peer-to-peer networks
Huanyu Zhao, Xiaolin Li 0001 |
J. Supercomput. | 2 |
| 2012 | IncMR: Incremental Data Processing Based on MapReduceabstractMapReduce programming model is widely used for large scale and one-time data-intensive distributed computing, but lacks flexibility and efficiency of processing small incremental data. IncMR framework is proposed in this paper for incrementally processing new data of a large data set, which takes state as implicit input and combines it with new data. Map tasks are created according to new splits instead of entire splits while reduce tasks fetch their inputs including the state and the intermediate results of new map tasks from designate nodes or local nodes. Data locality is considered as one of the main optimization means for job scheduling. It is implemented based on Hadoop, compatible with the original MapReduce interfaces and transparent to users. Experiments show that non-iterative algorithms running in MapReduce framework can be migrated to IncMR directly to get efficient incremental and continuous processing without any modification. IncMR is competitive and in all studied cases runs faster than that processing the entire data set. Cairong Yan, Xin Yang 0006, Min Li 0012, Xiaolin Li 0001 |
IEEE CLOUD | 5 |
| 2012 | Affinity-aware Virtual Cluster Optimization for MapReduce ApplicationsabstractInfrastructure-as-a-Service clouds are becoming ubiquitous for provisioning virtual machines on demand. Cloud service providers expect to use least resources to deliver best services. As users frequently request virtual machines to build virtual clusters and run MapReduce-like jobs for big data processing, cloud service providers intend to place virtual machines closely to minimize network latency and subsequently reduce data movement cost. In this paper we focus on the virtual machine placement issue for provisioning virtual clusters with minimum network latency in clouds. We define distance as the latency between virtual machines and use it to measure the affinity of virtual clusters. Such metric of distance indicates the considerations of virtual machine placement and topology of physical nodes in clouds. Then we formulate our problem as the classical shortest distance problem and solve it by modeling to integer programming problem. A greedy virtual machine placement algorithm is designed to get a compact virtual cluster. Furthermore, an improved heuristic algorithm is also presented for achieving a global resource optimization. The simulation results verify our algorithms and the experiment results validate the improvement achieved by our approaches. Cairong Yan, Ming Zhu 0004, Xin Yang 0006, Min Li 0012, Youqun Shi, Xiaolin Li 0001 |
CLUSTER | 7 |
| 2012 | User-centric private matching for eHealth networks - A social perspectiveabstractThe widely deployed electronic health (eHealth) systems changed people's daily life due to the extraordinary benefits, such as more efficiency, higher accuracy and broader availability. Patients in the eHealth network use their personal health records (PHRs) to communicate with their physicians and obtain medical services. As a matter of fact, patients who share the same diseases or symptoms want to communicate with each other not only for treatment, but also for psychological therapy. However, without sufficient knowledge of the authenticity of other patients' PHRs, patients are reluctant to share their medical information. On the other hand, patients would accept the patient-to-patient interaction only if their privacy issues of PHR are also well preserved. In this paper, we design a privacy-preserving user-centric private matching scheme from a social perspective in eHealth networks, where patients use verified PHR to find other patients who share the same situations and derive different user-centric results based on each one's own policy. In our scheme, the matching process guarantees both the verifiability and the privacy of patients' PHRs. Based on security and efficiency analysis, we show that our work satisfies both the privacy preservation and practicality requirements. Linke Guo, Xinxin Liu 0006, Yuguang Fang, Xiaolin Li 0001 |
GLOBECOM | 4 |
| 2012 | Acoustic ranging and communication via microphone channelabstractAbsolute GPS-coordinates are typically inaccessible indoor. Pervasive smartphones offer new opportunities for relative indoor localization. The low-complexity and high-accuracy ranging under the communication-link is a crucial requirement for enabling location sensing in indoor environments. In this paper, we propose a Time-of-Arrival (TOA) estimation and communication scheme utilizing the ubiquitous speaker/microphone pair to achieve accurate ranging. To compensate for the performance loss caused by this simple device, an optimized TOA estimation method is proposed for enhanced ranging reliability and accuracy. To overcome the strong channel fading in acoustic communication, a dynamic demodulation method that jointly uses frequency and amplitude information has been proposed. Experimental results show that our proposed scheme can achieve near 8cm TOA ranging accuracy and 0.55% communication bit-error-rate (BER) with 70% probability. Xinxin Liu 0006, Xiaolin Li 0001 |
GLOBECOM | 3 |
| 2012 | Blind Spots: Unveiling users' true willingness in online social networksabstractAlthough online social networks reflect real world social relationships, in many cases, online data is too scarce or implicit to reveal a user's true willingness. This causes the Blind Spot problem in socially-rendered willingness inference systems. Blind spots are the undervalued online contacts in willingness inference because of insufficient explicit evidences. To the best of our knowledge, this is the first time to introduce and address the blind spot problem. In this paper, we propose a scheme to detect blind spots, by contradicting explicit evidences and implicit inferences. The proposed scheme uses interaction history as the explicit evidence, and social circles for implicit inference. Real world experiments and surveys demonstrate that our scheme can detect blind spots. Xinxin Liu 0006, Xiaolin Li 0001 |
GLOBECOM | 3 |
| 2012 | Traffic-aware multiple mix zone placement for protecting location privacyabstractPrivacy protection is of critical concern to Location-Based Service (LBS) users in mobile networks. Long-term pseudonyms, although appear to be anonymous, in fact empower third-party service providers to continuously track users' movements. Researchers have proposed the mix zone model to allow pseudonym changes in protected areas. In this paper, we investigate a new form of privacy attack to the LBS system that an adversary reveals a user's true identity and complete moving trajectory with the aid of side information. We propose a new metric to quantify the system's resilience to such attacks, and suggest using multiple mix zones to tackle this problem. A mathematical model is presented that treats the deployment of multiple mix zones as a cost constrained optimization problem. Furthermore, the influence of traffic density is also taken into account to enhance the protection effectiveness. The placement optimization problem is NP-hard. We therefore design two heuristic algorithms as practical and effective means to strategically select mix zone locations, and consequently reduce the privacy risks of mobile users trajectories. The effectiveness of our proposed solutions is demonstrated through extensive simulations on real-world mobile user data traces. Xinxin Liu 0006, Han Zhao 0001, Miao Pan, Hao Yue 0001, Xiaolin Li 0001, Yuguang Fang |
INFOCOM | 5 |
| 2012 | Optimal Resource Rental Planning for Elastic Applications in Cloud MarketabstractThis paper studies the optimization problem of minimizing resource rental cost for running elastic applications in cloud while meeting application service requirements. Such a problem arises when excessive generated data incurs significant monetary cost on transfer and inventory in cloud. The goal of planning is to make resource rental decisions in response to varying application progress in the most cost-effective way. To address this problem, we first develop a Deterministic Resource Rental Planning (DRRP) model, using a mixed integer linear program, to generate optimal rental decisions given fixed cost parameters. Next, we systematically analyze the predictability of the time-varying spot instance prices in Amazon EC2 and find that the best achievable prediction is insufficient to provide a close approximation to the actual prices. This fact motivates us to propose a Stochastic Resource Rental Planning (SRRP) model that explicitly considers the price uncertainty in rental decision making. Using empirical spot price data sets and realistic cost parameters, we conduct simulations over a wide range of experimental scenarios. Results show that DRRP achieves as much as 50% cost reduction compared to the no-planning scheme. Moreover, SRRP consistently outperforms its DRRP counterpart in terms of cost saving, which demonstrates that SRRP is highly adaptive to the unpredictable nature of spot price in cloud resource market. Han Zhao 0001, Miao Pan, Xinxin Liu 0006, Xiaolin Li 0001, Yuguang Fang |
IPDPS | 4 |
| 2012 | An incentive mechanism to reinforce truthful reports in reputation systems
Huanyu Zhao, Xin Yang 0006, Xiaolin Li 0001 |
J. Netw. Comput. Appl. | 3 |
| 2011 | EPC: Energy-Aware Probability-Based Clustering Algorithm for Correlated Data Gathering in Wireless Sensor NetworksabstractThis paper addresses energy-efficient data gathering issues in wireless sensor networks (WSNs). Leveraging data correlation in densely-deployed sensor networks, we propose an Energy-aware Probability-based Clustering algorithm (EPC), featuring high scalability and flexibility particularly suitable for large-scale WSNs. Unlike most existing data gathering schemes that construct static routing structures or only consider spatial correlation among sensed data, EPC establishes energy-efficient routes on the fly during the data gathering process, and dynamically organizes sensor nodes into clusters based on a probability factor determined by both spatial and temporal data correlations. Redundant data transmissions are suppressed within a cluster and energy consumption is balanced to prolong network lifetime. To verify the effectiveness of EPC, extensive simulations are conducted on a network of 625 randomly deployed sensor nodes. Results show that EPC balances the energy consumption of the whole network and reduces up to 71% of the transmission costs with near negligible error rates for representative aggregation functions. Xinxin Liu 0006, Han Zhao 0001, Xiaolin Li 0001 |
AINA | 3 |
| 2010 | cTrust: Trust Aggregation in Cyclic Mobile Ad Hoc Networks
Huanyu Zhao, Xin Yang 0006, Xiaolin Li 0001 |
Euro-Par (2) | 3 |
| 2010 | WIM: A Wage-Based Incentive Mechanism for Reinforcing Truthful Feedbacks in Reputation SystemsabstractThe success of current trust and reputation systems is on the premise that truthful feedbacks are obtained. However, without appropriate mechanisms, silent and lying strategies usually yield higher payoffs for peers than truthful feedback strategies. Thus, to ensure trustworthiness, incentive mechanisms are critically needed for a reputation system to encourage rational peers to provide truthful feedbacks. In this paper, we model the feedback reporting process in reputation system as a reporting game. We propose a Wage-based Incentive Mechanism (WIM) for enforcing truthful report in self-interested P2P networks. We design, implement, and analyze incentive mechanisms and players' strategies. The extensive simulation results demonstrate that the proposed incentive mechanisms reinforce truthful feedbacks and achieve optimal welfare. Huanyu Zhao, Xin Yang 0006, Xiaolin Li 0001 |
GLOBECOM | 3 |
| 2010 | Hypergraph-based task-bundle scheduling towards efficiency and fairness in heterogeneous distributed systemsabstractThis paper investigates scheduling loosely coupled task-bundles in highly heterogeneous distributed systems. Two allocation quality metrics are used in pay-per-service distributed applications: efficiency in terms of social welfare, and fairness in terms of envy-freeness. The first contribution of this work is that we build a unified hypergraph scheduling model under which efficiency and fairness are compatible with each other. Second, in the scenario of budget-unawareness, we formulate a strategic algorithm design for distributed negotiations among autonomous self-interested computing peers and prove its convergence to complete local efficiency and envy-freeness. Third, we add budget limitation to the allocation problem and propose a class of hill-climbing heuristics in favor of different performance metrics. Finally we conduct extensive simulations to validate the performance of all the proposed algorithms. The results show that the decentralized hypergraph scheduling method is scalable, and yields desired allocation performance in various scenarios. Han Zhao 0001, Xinxin Liu 0006, Xiaolin Li 0001 |
IPDPS | 3 |
| 2010 | Trailing mobile sinks: A proactive data reporting protocol for Wireless Sensor NetworksabstractIn Wireless Sensor Networks (WSN), data gathering using mobile sinks typically incurs constant propagation of sink location indication messages to guide the direction of data reporting. Such behavior is undesirable, especially when the sensor network scale increases, as frequent message flooding will cause serious congestion in network communication and significantly impair the sensor network lifetime. In this paper, we propose a proactive data reporting protocol, SinkTrail, which achieves energy efficient data forwarding to multiple mobile sinks, and effectively reduces the number of sink location broadcasting messages. SinkTrail is unique in two aspects: (1) it allows sufficient flexibility in the movement of mobile sinks to dynamically adapt to unknown terrestrial changes; and (2) without assistance of GPS or predefined landmarks, SinkTrail establishes a logical coordinate system for predicting and tracking mobile sinks' locations, thereby significantly saves energy consumed during the data reporting process. We systematically analyze the impact of several design factors in SinkTrail and explore potential design improvements. The simulation results demonstrate that SinkTrail outperforms the Frequent Flooding Method (FFM) in finding shorter routing path and consumes less energy during data gathering process. Xinxin Liu 0006, Han Zhao 0001, Xin Yang 0006, Xiaolin Li 0001 |
MASS | 4 |
| 2009 | Efficient Grid Task-Bundle Allocation Using Bargaining Based Self-Adaptive AuctionabstractTo address coordination and complexity issues, we formulate a grid task allocation problem as a bargaining based self-adaptive auction and propose the BarSAA grid task-bundle allocation algorithm. During the auction, prices are iteratively negotiated and dynamically adjusted until market equilibrium is reached. The BarSAA algorithm features decentralized bidding decision making in a heterogeneous distributed environment so that scheduler can offload its duty onto participating computing nodes and significantly reduces scheduling overheads. When a BarSAA auction converges, the equilibrium point is Pareto Optimal and achieves social efficient outcome and double-sided revenue maximization. In addition, BarSAA promotes truthful behavior among selfish nodes. Through game theoretical analysis, we demonstrate that truthful revelation is beneficial to bidders in making bidding strategies. Extensive simulation results are presented to demonstrate the efficiency of the BarSAA strategy and validate several important analytical properties. Han Zhao 0001, Xiaolin Li 0001 |
CCGRID | 2 |
| 2009 | Load Scheduling Strategies for Parallel DNA Sequencing ApplicationsabstractThis paper studies a divisible load scheduling strategy with near-optimal processing time leveraging the computational characteristics of parallel DNA sequence alignment algorithms, specifically, the Needleman-Wunsch algorithm. Following the divisible load scheduling theory, an efficient load scheduling strategy is designed in large-scale networks so that the overall processing time of the sequencing tasks is minimized. In this study, the load distribution depends on the length of the sequence and number of processors in the network. Since we consider both of computation and communication overheads, the total processing time is also affected by communication link speed. Several cases have been considered in the study by varying the sequences, communication and computation speeds, and number of processors. Through simulation and numerical analysis, this study demonstrates that for a constant sequence length as the numbers of processors increase in the network the processing time for the job decreases and minimum overall processing time is achieved. Sudha Gunturu, Xiaolin Li 0001, Laurence T. Yang |
HPCC | 2 |
| 2009 | VectorTrust: Trust Vector Aggregation Scheme for Trust Management in Peer-to-Peer NetworksabstractWith emerging Internet-scale open content and resource sharing, social networks, and complex cyber-physical systems, trust issues become prominent. In this paper, we propose a trust vector based scheme (VectorTrust) for aggregation of distributed trust scores. Leveraging a Bellman-Ford based algorithm for fast trust score aggregation, VectorTrust features localized and distributed concurrent communication. A Vector Trust-enabled system is decentralized by nature and does not rely on any centralized server or centralized trust aggregation. We design, implement, and analyze trust aggregation and trust management strategies. To evaluate the performance, we design and implement a VectorTrust simulator (VTSim) in an unstructured P2P network. The analysis and simulation results demonstrate the efficiency, accuracy, scalability and robustness of VectorTrust scheme. On average, VectorTrust converges faster and involves less computational complexity than most existing trust schemes. VectorTrust remains robust and tolerant to malicious peers and malicious behaviors. With dynamic growth of P2P network scales and topology complexities, VectorTrust scales well with reasonable overheads (O(lgN) communication overheads) and fast convergence speed (about O(lgN) steps). Huanyu Zhao, Xiaolin Li 0001 |
ICCCN | 2 |
| 2009 | TinyBee: Mobile-Agent-Based Data Gathering System in Wireless Sensor NetworksabstractThis paper proposes a mobile-agent-based data gathering system (called TinyBee) in wireless sensor networks. Most existing mobile-agent-based systems consider only static sinks/servers. In this paper, we consider both mobile servers and lightweight mobile agents. We aim to design a data gathering system using a special kind of mobile agent called TinyBee to collect data all over a network. TinyBee migrates from node to node after being dispatched from a mobile server in order to collect data so that physical movement of mobile servers is greatly reduced. Mobile-agent-based approaches outperform traditional client/server paradigms in terms of execution time and power consumption. Extensive simulation results demonstrate that our proposed schemes achieve significant performance gains. Kaoru Ota, Mianxiong Dong, Xiaolin Li 0001 |
NAS | 3 |
| 2009 | H-Trust: A Group Trust Management System for Peer-to-Peer Desktop Grid
Huanyu Zhao, Xiaolin Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2008 | GridMate: A Portable Simulation Environment for Large-Scale Adaptive Scientific ApplicationsabstractIn this paper, we present a portable simulation environment GridMate for large-scale adaptive scientific applications in multi-site Grid environments. GridMate is a discrete-event based simulator, consisting of abstractions of trace-based applications, computing resources, partitioners and schedulers, a 3D visualization tool, and user interfaces. It supports the analysis of runtime management strategies that address spatial and temporal heterogeneity in both adaptive scientific applications and geographically distributed resources in Grid computing environments. The targeted applications are a class of emerging large-scale dynamic Grid applications that require large amount of computational resources typically spanning multiple sites and exhibit long execution times. The underlying partitioning and scheduling algorithms are based on our previous work on the hybrid space-time runtime management strategy (HRMS). HRMS defines a set of flexible mechanisms and policies to adapt to state transitions of both applications and resources. The major components of GridMate are developed in Java, making GridMate highly portable and extensible. The design of GridMate and simulation results using GridMate are presented. Xiaolin Li 0001, Manish Parashar |
CCGRID | 1 |
| 2008 | Autonomic Management of Hybrid Sensor Grid Systems and ApplicationsabstractIn this paper, we propose an autonomic management framework (ASGrid) to address the requirements of emerging large-scale applications in hybrid grid and sensor network systems. To the best of our knowledge, we are the first who proposed the autonomic sensor grid system concept in a holistic manner targeted at non-trivial large applications. To bridge the gap between the physical world and the digital world and facilitate information analysis and decision making, ASGrid is designed to smooth the integration of sensor networks and grid systems and efficiently use both on demand. Under the blueprint of ASGrid, we present several building blocks that fulfill the following major features: (1) Self-configuration through content-based aggregation and associative rendezvous mechanisms; (2) Self-optimization through utility-based sensor selection and model-driven hierarchical sensing task scheduling; (3) Self-protection through ActiveKey dynamic key management and S3Trust trust management mechanisms. Experimental and simulation results on these aspects are presented. Xiaolin Li 0001, Xinxin Liu 0006, Huanyu Zhao, Nanyan Jiang, Manish Parashar |
ICCCN | 1 |
| 2008 | A Personalized Group Trust Management System for Collaborative ServicesabstractIn this paper, we present a group trust management system S3Trust for collaborative services in peer-to-peer grid systems. S3Trust is based on personalized trust rating and selective aggregation algorithms. Leveraging the robustness of a simplistic but elegant co-constraint aggregation algorithm (inspired by H-index) under incomplete and uncertain circumstances, S3Trust offers a robust and lightweight reputation evaluation mechanism for both individual and group trusts with minimal communication and computation overheads. The five phases of S3Trust scheme are presented in detail, including trust recording, local trust evaluation, trust query phase, spatial-temporal update phase, and group reputation evaluation phases. Simulation results demonstrate that S3Trust is robust and can efficiently aggregate cooperative groups in systems with malicious users. Xiaolin Li 0001, Huanyu Zhao |
ICCCN | 1 |
| 2008 | DLBEM: Dynamic load balancing using expectation-maximizationabstractThis paper proposes a dynamic load balancing strategy called DLBEM based on maximum likelihood estimation methods for parallel and distributed applications. A mixture Gaussian model is employed to characterize workload in data- intensive applications. Using a small subset of workload information in systems, the DLBEM strategy reduces considerable communication overheads caused by workload information exchange and job migration. In the meantime, based on the Expectation-Maximization algorithm, DLBEM achieves near accurate estimation of the global system state with significantly less communication overheads and results in efficient workload balancing. Simulation results for some representative cases on a two-dimensional 16*16 grid demonstrate that DLBEM approach achieves even resource utilization and over 90% accuracy in the estimation of the global system state information with over 70% reduction on communication overheads compared to a baseline strategy. Han Zhao 0001, Xinxin Liu 0006, Xiaolin Li 0001 |
IPDPS | 3 |
| 2008 | Coordinated Workload Scheduling in Hierarchical Sensor Networks for Data Fusion Applications
Xiaolin Li 0001, Jiannong Cao 0001 |
J. Comput. Sci. Technol. | 1 |
| 2007 | Minimizing Distribution Cost of Distributed Neural Networks in Wireless Sensor NetworksabstractThis paper presents a novel study on how to distribute neural networks in a wireless sensor networks (WSNs) such that the energy consumption is minimized while improving the accuracy and training efficiency. Artificial neural network (ANN) learning has been shown robust to noisy and uncertain sensory data for function approximation and pattern classification applications. With the advances of miniature hardware technologies for powerful sensor nodes, embedded neural networks will emerge as important decision-making brains for WSNs and vast surveillance applications to enable adaptive data quality and self-managing capabilities. To distribute neural networks in WSNs in an energy-efficient manner, we propose parallel transmission and adaptive neural selection algorithms(ANSA) in multilayer backpropagation(MLBP) learning process of neural networks, which is a popular supervised learning technique used for training feedforward artificial neural networks. We further analyze the energy consumption components in the online training process and evaluate the reduced energy consumption using our proposed algorithms. Xiaolin Li 0001 |
GLOBECOM | 2 |
| 2007 | Sensing Workload Scheduling in Sensor Networks Using Divisible Load TheoryabstractThis paper presents scheduling strategies for sensing workload in wireless sensor networks using Divisible Load Theory (DLT), which offers a tractable model and realistic approach to investigate optimal scheduling issues in distributed systems. Due to the limited energy resource it is desirable that a sensor network can complete tasks as fast as possible. Sensor nodes are coordinated to perform measuring, transmitting, and processing data. Two closely related network models are presented to illustrate how the workload is scheduled among sensor nodes so that the finish time is minimized. Closed-form solutions are derived to achieve the optimization if the source node satisfies certain utility rate, which is used to evaluate the informative ratio of sensory data. Furthermore, we present the energy model for sensor nodes in the multi-hop multi-source network topology. Finally, simulation results are presented to demonstrate the effects of different parameters such as the number of sensor nodes, measurement, communication, and processing speed on the finish time and energy consumption. Xiaolin Li 0001, Xinxin Liu 0006 |
GLOBECOM | 1 |
| 2007 | PARMI: A Publish/Subscribe Based Asynchronous RMI Framework for Cluster Computing
Heejin Son, Xiaolin Li 0001 |
HPCC | 2 |
| 2007 | Sensing workload scheduling in hierarchical sensor networks for data fusion applicationsabstractWe consider a sensing task scheduling problem in two-level hierarchical sensor networks. To minimize the execution time of a given task, we propose efficient scheduling strategies following the divisible load scheduling paradigm. The proposed scheduling strategies minimize the finish time by eliminating transmission collisions and idle gaps between two successive data transmissions. In-network data aggregation for sensor data is further considered at data fusion nodes. Fused data are produced by some fusion functions on original data from local clusters. The scheduling strategies consist of two phases: intra-cluster scheduling and inter-cluster scheduling. Intra-cluster scheduling deals with assigning different fractions of a sensing workload among source nodes in each cluster; inter-cluster scheduling involves the distribution of fused data among all fusion nodes. Closed-form solutions to the problem of task scheduling are derived. Energy model is described for each kind of sensor nodes, considering data acquisition, communication, and processing. Finally, simulation results are presented to demonstrate the impacts of different system parameters such as the number of sensor nodes, measurement, communication, and processing speed, on the finish time and energy consumption. Xiaolin Li 0001, Hsiao-Hwa Chen |
IWCMC | 1 |
| 2007 | Coordinated Workload Scheduling in Hierarchical Sensor Networks for Data Fusion ApplicationsabstractTo minimize the execution time of a sensing task over a multi-hop hierarchical sensor network, we present a coordinated scheduling method following the divisible load scheduling paradigm. The proposed scheduling strategy builds from eliminating transmission collisions and idle gaps between two successive data transmissions. We consider a sensor network consisting of several clusters. In a cluster, after related raw data measured by source nodes are collected at the fusion node, in-network data aggregation is further considered. The scheduling strategies consist of two phases: intra-cluster scheduling and inter-cluster scheduling. Intra-cluster scheduling deals with assigning different fractions of a sensing workload among source nodes in each cluster; inter-cluster scheduling involves the distribution of fused data among all fusion nodes. Closed-form solutions to the problem of task scheduling are derived. Finally, numerical examples are presented to demonstrate the impacts of different system parameters such as the number of sensor nodes, measurement, communication, and processing speed, on the finish time and energy consumption. Xiaolin Li 0001, Jiannong Cao 0001 |
MASS | 1 |
| 2007 | Power-Aware Markov Chain Based Tracking Approach for Wireless Sensor NetworksabstractWe propose a novel measure method of information utility for tracking and localization in wireless sensor networks (WSNs). The target moving arbitrarily in WSNs is modeled by Markov chains using a transition matrix. The proposed information utility measurement allows us to expect the next state of the target and identify the informative sensors. Further, compared with existing localization methods, the proposed power-aware sensor selection considers the energy constraint of WSNs. To conserve energy, subsets of sensor nodes are activated based on a combinative measurement including information utility, communication cost, and residual energy. We have implemented the proposed localization system on real motes and experimented in an obstacle-free environment. The experimental results demonstrate that the proposed method outperforms two popular baseline schemes, k-nearest-neighbor and stochastic schemes, at extending the network lifetime. In addition, it balances the energy level of sensors in the network so that energy consumption is spread uniformly over all the sensors. Xiaolin Li 0001, Patrick J. Moran |
WCNC | 2 |
| 2007 | Enabling scalable parallel implementations of structured adaptive mesh refinement applications
Sumir Chandra, Xiaolin Li 0001, Taher Saif, Manish Parashar |
J. Supercomput. | 2 |
| 2007 | Hybrid Runtime Management of Space-Time Heterogeneity for Parallel Structured Adaptive ApplicationsabstractStructured adaptive mesh refinement (SAMR) techniques provide an effective means for dynamically concentrating computational effort and resources to appropriate regions in the application domain. However, due to their dynamism and space-time heterogeneity, scalable parallel implementation of SAMR applications remains a challenge. This paper investigates hybrid runtime management strategies and presents an adaptive hierarchical multipartitioner (AHMP) framework. AHMP dynamically applies multiple partitioners to different regions of the domain, in a hierarchical manner, to match the local requirements of the regions. Key components of the AHMP framework include a segmentation-based clustering algorithm (SBC) that can efficiently identify regions in the domain with relatively homogeneous partitioning requirements, mechanisms for characterizing the partitioning requirements of these regions, and a runtime system for selecting, configuring, and applying the most appropriate partitioner to each region. Further, to address dynamic resource situations for long-running applications, AHMP provides a hybrid partitioning strategy (HPS) that involves application-level pipelining, trading space for time when resources are sufficiently large and underutilized, and an application-level out-of-core strategy (ALOC), trading time for space when resources are scarce in order to enhance the survivability of applications. The AHMP framework has been implemented and experimentally evaluated on up to 1,280 processors of the IBM SP4 cluster at the San Diego Supercomputer Center. Xiaolin Li 0001, Manish Parashar |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2006 | Autonomic and Trusted Computing Paradigms
Xiaolin Li 0001, Patrick Harrington, Johnson Thomas |
ATC | 1 |
| 2006 | Autonomic Sensor Networks: A New Paradigm for Collaborative Information ProcessingabstractWireless sensor networks (WSNs) are severely constrained in computation and communication capabilities due to the cost and size of available sensors. On the other hand, autonomic computing (AC) offers a promising solution to manage large-scale computing systems without human intervention. Realizing the similarity between WSNs and AC applications, this paper proposes an autonomic sensor network framework to enable self-managing wireless sensor network systems for collaborative information processing. In particular, a preliminary power-aware self-configuring and self-optimizing sensor selection scheme is developed to improve the performance and extend the lifetime of sensor networks. The simulation results confirm that the proposed power-aware scheme prolongs the network lifetime and balances the energy in sensor nodes Xiaolin Li 0001, Patrick J. Moran |
DASC | 2 |
| 2005 | Using Clustering to Address Heterogeneity and Dynamism in Parallel Scientific Applications
Xiaolin Li 0001, Manish Parashar |
HiPC | 1 |
| 2004 | Hierarchical Partitioning Techniques for Structured Adaptive Mesh Refinement Applications
Xiaolin Li 0001, Manish Parashar |
J. Supercomput. | 1 |
| 2003 | Dynamic Load Partitioning Strategies for Managing Data of Space and Time Heterogeneity in Parallel SAMR Applications
Xiaolin Li 0001, Manish Parashar |
Euro-Par | 1 |
| 2001 | Divisible Load Scheduling on a Hypercube Cluster with Finite-Size Buffers and Granularity ConstraintsabstractIn this paper we address the problem of scheduling a large size divisible load on a hypercube cluster of processors. Unlike in earlier studies in the divisible load theory (DLT) literature, here, we assume that the processors have finite-size buffers. Further we impose constraints on the extent to which the load can be divided, referred to as granularity constraint. We first present the closed-form solutions for the case with infinite-size buffers. For the case with load granularity constraint, we propose a simple algorithm to find the sub-optimal solution and then we analyze the case when these buffer are of finite size. For this case, we present an elegant strategy, referred to as incremental balancing strategy (IBS), to obtain an optimal load distribution. Based on the rigorous mathematical analysis, a number of interesting and useful properties exhibited by the algorithm are proven. Numerical examples are presented for the ease of understanding. Xiaolin Li 0001, Bharadwaj Veeravalli, Chi Chung Ko |
CCGRID | 1 |
| 2000 | Efficient partitioning and scheduling of computer vision and image processing data on bus networks using divisible load analysis
Bharadwaj Veeravalli, Xiaolin Li 0001, Chi Chung Ko |
Image Vis. Comput. | 2 |
| 2000 | On the Influence of Start-Up Costs in Scheduling Divisible Loads on Bus NetworksabstractOptimal distribution of divisible loads in bus networks is considered in this paper. The problem of minimizing the processing time is investigated by including all the overhead components that could penalize the performance of the system, in addition to the inherent communication and computation delays. These overheads are considered to be constant additive factors to the respective communication and computation components. Closed-form solution for the processing time is derived and the influence of overheads on the optimal processing time is analyzed. We derive a necessary and sufficient condition for the existence of the optimal processing time. We then study the effect of changing the load distribution sequence on the time performance. Through rigorous analysis, an optimal sequence to distribute the load among the processors is identified, whenever it exists. In case such an optimal sequence fails to exist, we present a greedy algorithm to obtain a suboptimal sequence based on some important properties of the overhead factors. Then, the effect of granularity of the data that is divisible is considered in the analysis for the case of homogeneous networks. An integer approximation algorithm capable of generating integer values of the load fractions in time O(m), where m is the number of processors in the network, is proposed. We then show that the upper bound on the suboptimal solution generated by our algorithm lies within a radius given by the sum of the computation and communication delays. Several numerical examples are presented to illustrate the concepts. Bharadwaj Veeravalli, Xiaolin Li 0001, Chi Chung Ko |
IEEE Trans. Parallel Distributed Syst. | 2 |