EDBT 2026 Demo / reviewers in the wild / expert
Ning An 0001
dblp:98/6171-1
· DBLP profile ↗
61ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0003-3317-5299ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 15 since 2021Databases, data management, data science and information retrieval · 16 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 4 since 2021Computer networks · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EaNet: Enhanced Multimodal Awareness Alignment Network for Multimodal Aspect-Based Sentiment Analysis
Aoqiang Zhu, Min Hu 0010, Xiaohua Wang 0002, Yan Xing 0002, Yiming Tang 0001, Jiaoyun Yang, Ning An 0001, Fuji Ren |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | SCImputation: Mitigating Feature Confounding From a Structural Causal Perspective for Data Imputation
Jiaoyun Yang, Ning An 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete DataabstractMultimodal Sentiment Analysis (MSA) with incomplete data has gained significant attention recently.Existing studies focus on optimizing model structures to handle modality missingness, but models still face challenges in robustness when dealing with uncertain missingness.To this end, we propose a data-centric robust multimodal sentiment analysis method, Proxy-Driven Robust Multimodal Fusion (P-RMF).First, we map unimodal data to the latent space of Gaussian distributions to capture core features and structure, thereby learn stable modality representation.Then, we combine the quantified modality intrinsic uncertainty to learn stable multimodal joint representation (i.e., proxy modality), which is further enhanced through multi-layer dynamic cross-modal injection to increase its diversity.Extensive experimental results show that P-RMF outperforms existing models in noise resistance and achieves state-of-the-art performance on multiple benchmark datasets. Aoqiang Zhu, Min Hu 0010, Xiaohua Wang 0002, Jiaoyun Yang, Yiming Tang 0001, Ning An 0001 |
ACL (1) | 6 |
| 2025 | Measuring Ageism in Large Language ModelsabstractAs large language models gain prominence, there is increasing concern about the potential biases they may perpetuate. While various biases have been studied, ageism in language models remains underexplored. According to the World Health Organization, ageism can significantly impact the physical and mental well-being of older adults, an impact that could grow as the global aging population increases. To address this research gap, we developed AgeismSet, a comprehensive Chinese dataset comprising 6,444 sentences, enhanced with neutral impression options to provide a balanced evaluation framework. Our study then used AgeismSet to investigate ageism in large language models across cognitive, affective, and behavioral dimensions, using AgeismSet to evaluate models such as GPT-4, GPT-3.5, GLM-3-Turbo, ERNIE Bot, Gemini Pro, and DeepSeek-V3. Our findings, quantified by the Ageism Score (AS), reveal that while some models perform well, there is considerable room for improvement in mitigating ageism. This work underscores the necessity for targeted interventions to ensure more equitable AI systems. Jiaoyun Yang, Hongtu Chen, Ning An 0001 |
ECAI | 5 |
| 2025 | SynNER: Synergizing Large and Small Language Models for Few-Shot Nested NERabstractLarge language models (LLMs) encounter challenges when addressing few-shot nested named entity recognition (NER) tasks. Traditional LLM-based approaches typically either prompt the model to generate entity words or types in sentences based on entity categories or word spans, or directly extract all entities of specific types present in the sentences. These methods often suffer from issues such as low query efficiency or suboptimal accuracy. This paper introduces an innovative framework, SynNER, which synergizes small and large language models to address these limitations. Initially, a small language model identifies low-confidence word spans, which are then refined and refined by a large language model. To simultaneously ensure recognition accuracy and improve the query efficiency of the LLM, we propose a Batch-Prompt strategy and an Entity Indexing method. These techniques enable the LLM to process multiple test instances simultaneously while maintaining precise correction results. Experimental results demonstrate that our method achieves significant performance gains on benchmark datasets, offering a cost-effective solution for few-shot nested NER tasks. Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001 |
IJCNN | 5 |
| 2025 | SynNER: Synergizing Large and Small Language Models for Few-Shot Nested NERabstractLarge language models (LLMs) encounter challenges when addressing few-shot nested named entity recognition (NER) tasks. Traditional LLM-based approaches typically either prompt the model to generate entity words or types in sentences based on entity categories or word spans, or directly extract all entities of specific types present in the sentences. These methods often suffer from issues such as low query efficiency or suboptimal accuracy. This paper introduces an innovative framework, SynNER, which synergizes small and large language models to address these limitations. Initially, a small language model identifies low-confidence word spans, which are then refined and refined by a large language model. To simultaneously ensure recognition accuracy and improve the query efficiency of the LLM, we propose a Batch-Prompt strategy and an Entity Indexing method. These techniques enable the LLM to process multiple test instances simultaneously while maintaining precise correction results. Experimental results demonstrate that our method achieves significant performance gains on benchmark datasets, offering a cost-effective solution for few-shot nested NER tasks. Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001 |
IJCNN | 5 |
| 2025 | Harnessing high-quality pseudo-labels for robust few-shot nested named entity recognition
Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Mitigating prototype shift: Few-shot nested named entity recognition with prototype-attention contrastive learning
Hong Ming, Jiaoyun Yang, Lili Jiang 0002, Ning An 0001 |
Expert Syst. Appl. | 5 |
| 2025 | CDC-FSL: A Causal De-Confounding framework for Few-Shot Learning
Jiaoyun Yang, Ning An 0001, Lian Li 0003 |
Knowl. Based Syst. | 3 |
| 2024 | LPNER: Label Prompt for Few-shot Nested Named Entity Recognition
Jiaoyun Yang, Zhihan Zhu, Hong Ming, Lili Jiang 0002, Ning An 0001 |
ACML | 5 |
| 2024 | CauImputation: Utilizing Structural Causal Model in Data ImputationabstractMissing data pose significant obstacles in data analysis. Many imputation methods, operating under the assumption that similar instances exhibit similar feature values, often overlook the essential role of the feature containing the missing value. These methods tend to apply a uniform strategy to all missing data within a single instance. This study proposes selecting neighbor instances for imputation based on both instance and missing feature information. To support this approach, a Structural Causal Model (SCM) delineates the relationships among missing data, features, and instances. The model reveals that the feature with the target missing value acts as a confounder, influencing both the selection of neighbor instances and the prediction of the missing value. Building upon these findings, this study introduces a new imputation method called the Causal Imputation strategy (CauImputation). CauImputation utilizes both feature and instance information to select neighbor instances and then applies causal interventions to compute the missing value. This strategy effectively leverages feature information while mitigating confounding effects. Tests on the National Alzheimer’s Coordinating Center (NACC) dataset and the National Center for Biotechnology Information (NCBI) microarray datasets demonstrate that CauImputation outperforms nine established methods, achieving a 3% to 15% increase in imputation accuracy and reducing root mean squared errors by 0.05 to 0.2. By utilizing the Structural Causal Model, this research introduces a novel perspective on data imputation, enhancing the robustness and reliability of biomedical data analyses. Jiaoyun Yang, Ning An 0001 |
BIBM | 3 |
| 2024 | Deep Frequency Derivative Learning for Non-stationary Time Series Forecasting
Wei Fan 0010, Kun Yi 0001, Hangting Ye, Zhiyuan Ning 0001, Qi Zhang 0020, Ning An 0001 |
IJCAI | 6 |
| 2024 | Few-shot nested named entity recognition
Hong Ming, Jiaoyun Yang, Fang Gui, Lili Jiang 0002, Ning An 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Parallel Multiscale Bridge Fusion Network for Audio-Visual Automatic Depression AssessmentabstractDepression is a prevalent and severe mental illness that significantly impacts patients’ physical health and daily life. Recent studies have focused on multimodal depression assessment, aiming to objectively and conveniently evaluate depression using multimodal data. However, existing methods based on audio–visual modalities struggle to capture the dynamic variations in depression clues and cannot fully explore multimodal data over a long time. In addition, they rely heavily on insufficient single-stage multimodal fusion, which limits the accuracy of depression assessment. To address these limitations, we propose a novel parallel multiscale bridge fusion network (PMBFN) for audio–visual depression assessment. PMBFN comprehensively captures subtle multilevel dynamic changes in depression expression through parallel multiscale dynamic convolutions and long short-term memories (LSTMs) and effectively solves the problem of long-term audio–visual sequence information loss by using spatiotemporal attention pooling modules. Furthermore, the multimodal bridge fusion module is proposed in PMBFN to achieve multistage interactive recursive multimodal fusion, enhancing the expressive capacity of multimodal depression-related features to improve the accuracy of assessment. Extensive experiments on the DAIC-WOZ and E-DAIC datasets demonstrate that our method outperforms current state-of-the-art methods and clearly shows our method's effectiveness eventually. Min Hu 0010, Xiaohua Wang 0002, Yiming Tang 0001, Jiaoyun Yang, Ning An 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | ER-C3D: Enhancing R-C3-D Network With Adaptive Shrinkage and Symmetrical Multiscale for Behavior DetectionabstractBehavior detection receives considerable attention in real-life human–computer interaction, where the complexity of background information and the variable durations of movements are two major factors affecting the accuracy of behavior detection. To overcome the inadequacy of these factors, this article proposes an enhancing region convolutional 3-D (ER-C3D) network with adaptive shrinkage and symmetrical multiscale for behavior detection. The improved ER-C3D network includes a feature subnet, a proposal subnet, and a classification subnet. First, a 3D-RSST unit is constructed by embedding an adaptive shrinkage structure and a soft thresholding operation. Meanwhile, a residual adaptive shrinkage mechanism, composed of multiple cascaded 3D-RSST units with different parameters, is designed to reduce redundant information of video streams in the feature subnet. Second, a spatiotemporal symmetrical multiscale structure is substituted for the single-layer convolution and embedded into the proposal subnet. Specially, contextual symmetrical multiscale motion characteristics with different levels and granularities are acquired by expanding the spatiotemporal receptive field of candidate temporal proposals. Finally, a soft-nonmaximal suppression strategy is introduced to filter high-quality temporal proposals in the classification subnet. The experimental results on the THUMOS’14 and ActivityNet1.2 datasets indicate that the [email protected] of the improved ER-C3D network reaches 39.4% and 42.2%, respectively, which is 10.5% and 15.4% higher than R-C3D. Compared with related methods, the proposed method shows improvement in both the positional precision of behavioral boundary and the accuracy of behavioral classification. Mengyuan Tao, Ning An 0001, Min Hu 0010, Fuji Ren |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Deep Coupling Network for Multivariate Time Series ForecastingabstractMultivariate time series (MTS) forecasting is crucial in many real-world applications. To achieve accurate MTS forecasting, it is essential to simultaneously consider both intra- and inter-series relationships among time series data. However, previous work has typically modeled intra- and inter-series relationships separately and has disregarded multi-order interactions present within and between time series data, which can seriously degrade forecasting accuracy. In this article, we reexamine intra- and inter-series relationships from the perspective of mutual information and accordingly construct a comprehensive relationship learning mechanism tailored to simultaneously capture the intricate multi-order intra- and inter-series couplings. Based on the mechanism, we propose a novel deep coupling network for MTS forecasting, named DeepCN, which consists of a coupling mechanism dedicated to explicitly exploring the multi-order intra- and inter-series relationships among time series data concurrently, a coupled variable representation module aimed at encoding diverse variable patterns, and an inference module facilitating predictions through one forward step. Extensive experiments conducted on seven real-world datasets demonstrate that our proposed DeepCN achieves superior performance compared with the state-of-the-art baselines. Kun Yi 0001, Qi Zhang 0020, Kaize Shi, Liang Hu 0004, Ning An 0001, Zhendong Niu |
ACM Trans. Inf. Syst. | 6 |
| 2023 | A Multiscale Dynamic Temporal Convolution Network For Continuous Dimensional Emotion RecognitionabstractVideo-based dimensional emotion recognition, especially long sequence modeling, is a challenging task. This paper proposes a Multiscale Dynamic Temporal Convolution Network(MDTCN) for long sequence modeling of dimensional emotion recognition. Although Temporal Convolution Network(TCN) has performed well in sequence modeling, there are still some limitations that the feature scale of each layer is single and the receptive field of the upper layer is too large to capture the short-term dependence. Therefore, Dynamic Dilated Block(DDB) is constructed to better capture the temporal dependence by using dilated convolutions of different kernel size. In order to obtain both long-term and short-term temporal features, Multiscale TCN Structure is proposed. Extensive experiments are conducted on the AFFWILD2 and SEWA Dataset and the competitive result is shown compared to state-of-the-art methods. Min Hu 0010, Jialu Sun, Xiaohua Wang 0002, Ning An 0001 |
IJCNN | 4 |
| 2023 | FourierGNN: Rethinking Multivariate Time Series Forecasting from a Pure Graph PerspectiveabstractMultivariate time series (MTS) forecasting has shown great importance in numerous industries. Current state-of-the-art graph neural network (GNN)-based forecasting methods usually require both graph networks (e.g., GCN) and temporal networks (e.g., LSTM) to capture inter-series (spatial) dynamics and intra-series (temporal) dependencies, respectively. However, the uncertain compatibility of the two networks puts an extra burden on handcrafted model designs. Moreover, the separate spatial and temporal modeling naturally violates the unified spatiotemporal inter-dependencies in real world, which largely hinders the forecasting performance. To overcome these problems, we explore an interesting direction of directly applying graph networks and rethink MTS forecasting from a pure graph perspective. We first define a novel data structure, hypervariate graph, which regards each series value (regardless of variates or timestamps) as a graph node, and represents sliding windows as space-time fully-connected graphs. This perspective considers spatiotemporal dynamics unitedly and reformulates classic MTS forecasting into the predictions on hypervariate graphs.
Then, we propose a novel architecture Fourier Graph Neural Network (FourierGNN) by stacking our proposed Fourier Graph Operator (FGO) to perform matrix multiplications in Fourier space. FourierGNN accommodates adequate expressiveness and achieves much lower complexity, which can effectively and efficiently accomplish {the forecasting}. Besides, our theoretical analysis reveals FGO's equivalence to graph convolutions in the time domain, which further verifies the validity of FourierGNN. Extensive experiments on seven datasets have demonstrated our superior performance with higher efficiency and fewer parameters compared with state-of-the-art methods. Code is available at this repository: https://github.com/aikunyi/FourierGNN. Kun Yi 0001, Qi Zhang 0020, Wei Fan 0010, Liang Hu 0004, Pengyang Wang, Ning An 0001, Longbing Cao, Zhendong Niu |
NeurIPS | 7 |
| 2023 | Frequency-domain MLPs are More Effective Learners in Time Series ForecastingabstractTime series forecasting has played the key role in different industrial, including finance, traffic, energy, and healthcare domains. While existing literatures have designed many sophisticated architectures based on RNNs, GNNs, or Transformers, another kind of approaches based on multi-layer perceptrons (MLPs) are proposed with simple structure, low complexity, and superior performance. However, most MLP-based forecasting methods suffer from the point-wise mappings and information bottleneck, which largely hinders the forecasting performance. To overcome this problem, we explore a novel direction of applying MLPs in the frequency domain for time series forecasting. We investigate the learned patterns of frequency-domain MLPs and discover their two inherent characteristic benefiting forecasting, (i) global view: frequency spectrum makes MLPs own a complete view for signals and learn global dependencies more easily, and (ii) energy compaction: frequency-domain MLPs concentrate on smaller key part of frequency components with compact signal energy. Then, we propose FreTS, a simple yet effective architecture built upon Frequency-domain MLPs for Time Series forecasting. FreTS mainly involves two stages, (i) Domain Conversion, that transforms time-domain signals into complex numbers of frequency domain; (ii) Frequency Learning, that performs our redesigned MLPs for the learning of real and imaginary part of frequency components. The above stages operated on both inter-series and intra-series scales further contribute to channel-wise and time-wise dependency learning. Extensive experiments on 13 real-world benchmarks (including 7 benchmarks for short-term forecasting and 6 benchmarks for long-term forecasting) demonstrate our consistent superiority over state-of-the-art methods. Code is available at this repository: https://github.com/aikunyi/FreTS. Kun Yi 0001, Qi Zhang 0020, Wei Fan 0010, Shoujin Wang, Pengyang Wang, Ning An 0001, Defu Lian, Longbing Cao, Zhendong Niu |
NeurIPS | 7 |
| 2023 | Toward Better Understanding Older Adults: A Biography Brief Timeline Extraction ApproachabstractStudies have shown that life stories can help caregivers better understand older adults, leading to better care. The original life stories are often redundant and disordered, hindering the discovery of valuable information. There has been little research on organizing older adults’ life stories automatically. This article proposes the ALBERT Based Text Extraction Network (ABTE-NET) to generate valuable event timelines to address this problem. To evaluate the proposed method, we created an Older Adults’ Life Story dataset with 80 older adults’ life stories. In experiments, we verify that the timelines generated by ABTE-NET have good readability and summarization for life stories. A survey of 33 caregivers from two nursing homes shows that timelines can help caregivers understand older adults and build positive relationships, just like life stories. More importantly, timelines are better organized and more concise than original life stories, reducing the cognitive load and helping caregivers form a preliminary understanding of the older adult quickly. Ning An 0001, Fang Gui, Liuqi Jin, Hong Ming, Jiaoyun Yang |
Int. J. Hum. Comput. Interact. | 1 |
| 2022 | Causal Discovery on Non-Euclidean DataabstractResearchers recently started developing deep learning models capable of handling non-Euclidean data. However, because of existing framework limitations on model representations and learning algorithms, few have explored causal discovery on non-Euclidean data. This paper is the first attempt to do so. We start by proposing the Non-Euclidean Causal Model (NECM) which describes the causal generative relationship of non-Euclidean data and creates a new tensor data type along with a mapping process for the non-Euclidean causal mechanism. Second, within the NECM, we propose the non-Euclidean Hybrid Learning (NEHL) method, a causal discovery algorithm relying on the concept of the ball covariance recently introduced in the statistics field. Third, we generate two types of non-Euclidean datasets: Functional Data and Symmetric Positive Definite manifold data in conformity with the NECM. Finally, experimental results on the generated data and real-world data demonstrate the effectiveness of the proposed NEHL method. Jing Yang 0008, Ning An 0001 |
KDD | 3 |
| 2022 | The Case for FPGA-Based Edge ComputingabstractEdge Computing has emerged as a new computing paradigm dedicated for mobile performance enhancement and energy efficiency purposes. Specifically, it benefits today’s interactive applications on power-constrained devices by offloading compute-intensive tasks to the edge nodes in close proximity. Meanwhile, FPGA is well known for its excellence in accelerating (domain-specific) compute-intensive tasks such as deep learning algorithms in a high performance and energy-efficient manner due to its hardware-customizable nature. In this paper, we make the first attempt to leverage and combine the advantages of these two, and proposed a new network-assisted computing model, namely FPGA-based edge computing. As a case study, we choose three computer vision (CV)-based mobile interactive applications, and implement their back-end computation engines on FPGA. By deploying such application-customized accelerator modules for computation offloading at the network edge, we experimentally demonstrate that this approach can effectively reduce response time for the applications and energy consumption for the entire system in comparison with traditional CPU-based edge/cloud offloading approach. Chenren Xu, Shuang Jiang, Guojie Luo, Guangyu Sun 0003, Ning An 0001, Gang Huang 0001, Xuanzhe Liu |
IEEE Trans. Mob. Comput. | 5 |
| 2021 | Feasibility study of practical vital sign detection using millimeter-wave radios
Zhenhua Jia, Chenren Xu, Guojie Luo, Daqing Zhang 0001, Ning An 0001, Yanyong Zhang |
CCF Trans. Pervasive Comput. Interact. | 6 |
| 2020 | Personal Cognitive Health Library
Ning An 0001, Huitong Ding, Gil Alterovitz |
AMIA | 2 |
| 2020 | Deep ensemble learning for Alzheimer's disease classification
Ning An 0001, Huitong Ding, Jiaoyun Yang, Rhoda Au, Ting Fang Alvin Ang |
J. Biomed. Informatics | 1 |
| 2020 | An efficient causal structure learning algorithm for linear arbitrarily distributed continuous data
Jing Yang 0008, Ning An 0001, Yu Chen 0018, Gil Alterovitz |
J. Supercomput. | 3 |
| 2019 | Microarray Missing Value Imputation: A Regularized Local Learning MethodabstractMicroarray experiments on gene expression inevitably generate missing values, which impedes further downstream biological analysis. Therefore, it is key to estimate the missing values accurately. Most of the existing imputation methods tend to suffer from the over-fitting problem. In this study, we propose two regularized local learning methods for microarray missing value imputation. Motivated by the grouping effect of $L_{2}$L2 regularization, after selecting the target gene, we train an $L_{2}$L2 Regularized Local Least Squares imputation model (RLLSimpute_L2) on the target gene and its neighbors to estimate the missing values of the target gene. Furthermore, RLLSimpute_L2 imputes the missing values in an ascending order based on the associated missing rate with each target gene. This contributes to fully utilizing the previously estimated values. Besides $L_{2}$L2, we further explore $L_{1}$L1 regularization and propose an $L_{1}$L1 Regularized Local Least Squares imputation model (RLLSimpute_L1). To evaluate their effectiveness, we conducted extensive experimental studies on six benchmark datasets covering both time series and non-time series cases. Nine state-of-the-art imputation methods are compared with RLLSimpute_L2 and RLLSimpute_L1 in terms of three performance metrics. The comparative experimental results indicate that RLLSimpute_L2 outperforms its competitors by achieving smaller imputation errors and better structure preservation of differentially expressed genes. Aiguo Wang 0002, Ye Chen 0011, Ning An 0001, Jing Yang 0008, Lian Li 0001, Lili Jiang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Streaming feature-based causal structure learning algorithm with symmetrical uncertainty
Jing Yang 0008, Xiaoxue Guo, Ning An 0001, Aiguo Wang 0002, Kui Yu |
Inf. Sci. | 3 |
| 2018 | Subtype dependent biomarker identification and tumor classification from gene expression profiles
Aiguo Wang 0002, Ning An 0001, Guilin Chen, Li Liu 0001, Gil Alterovitz |
Knowl. Based Syst. | 2 |
| 2018 | uGait: A platform for automated quantitative gait analysis and its application to Parkinson's diseaseabstractGait analysis is the systematic study of human motion and aims at quantifying the gait characteristics with various temporal and spatial gait parameters. Gait has proven to be linked to certain diseases, such as hemiplegic paralysis and Parkinson’s disease (PD). In this work, we present our designe d and developed gait sensing platform that can capture human movement, called uGait, and we further present its application in discriminating patients with Parkinson’s disease from healthy subjects. First, we show the overall system architecture of uGait. Second, we detail how to extract and select meaningful spatial-temporal gait parameters and present an effective wrapper-based feature selection method. Typically, we consider the gait parameters associated with the situation where one walks on a u-shape walkway rather than only on a straight walkway, and we consider the influence of human body height on gait parameters extraction. We then apply our proposed feature selector to identify a discriminant subset of gait features and construct a classifier with better generalization ability for PD detection. Finally, we collect experimental data with uGait and conduct extensive experiments with three different classifiers and three different feature selectors. Experimental results demonstrate power of the quantitative gait analysis in classifying PD patients. Xi Wu 0003, Shengqiang Xu, Ning An 0001 |
Web Intell. | 6 |
| 2017 | A study on gait-based Parkinson's disease detection using a force sensitive platformabstractGait analysis aims to study human motion and its potential association with chronic diseases, such as Parkinson's disease and hemiplegic paralysis, by extracting various gait characteristics. It has been a challenging problem to accurately extract temporal and spatial gait parameter and to explore the relationship between gait signal and a disease of interest. In this study, we introduce a gait sensing platform that can capture human movement and classify patients with Parkinson's disease from healthy subjects. Specifically, we first show the platform that consists of force sensitive pressure sensors. Second, we extract gait features from the gait signal collected from the platform. Finally, we collect experimental data from 386 volunteers, including 218 healthy subjects and 168 patients with Parkinson's disease, and conduct extensive experiments to show the possibility of classifying Parkinson's disease patients at a high confidence level. Experimental results of nine different classifiers show that the random forest model outperforms the other eight competitors and obtains an accuracy of 92.49%, demonstrating the power of quantitative gait analysis in the early detection of Parkinson's disease. Xi Wu 0003, You Duan, Shengqiang Xu, Ning An 0001 |
BIBM | 6 |
| 2017 | Optimizing mobility and energy charging for mobile cloudletabstractMobile cloudlet has recently received considerable attentions in mobile cloud computing because of its reduction in energy consumption and communication latency. It faces, however, two major challenges: first, its mobility path could significantly affect mobile application services it can offer to other mobile devices; second, its resources are very limited including its battery capacity. In this paper, we aim to optimize the mobility and energy charging for mobile cloudlet at the same time. Specifically, we design an efficient mobility path for mobile cloudlet to maximize the overall revenue from serving other mobile devices while preserving its own services. With this new mobility path, we further propose a dynamic programming technique to minimize the energy recharging cost of mobile cloudlet, by appropriately selecting the wireless charging stations along its path. Finally, simulations are conducted to evaluate the effectiveness of the proposed schemes on optimizing both the mobility and the energy charging of a mobile cloudlet. Xiumin Wang 0004, Min Peng 0001, Ning An 0001 |
ICC | 4 |
| 2017 | Nonlinear dimensionality reduction methods for synthetic biology biobricks' visualizationabstractBACKGROUND: Visualizing data by dimensionality reduction is an important strategy in Bioinformatics, which could help to discover hidden data properties and detect data quality issues, e.g. data noise, inappropriately labeled data, etc. As crowdsourcing-based synthetic biology databases face similar data quality issues, we propose to visualize biobricks to tackle them. However, existing dimensionality reduction methods could not be directly applied on biobricks datasets. Hereby, we use normalized edit distance to enhance dimensionality reduction methods, including Isomap and Laplacian Eigenmaps. RESULTS: By extracting biobricks from synthetic biology database Registry of Standard Biological Parts, six combinations of various types of biobricks are tested. The visualization graphs illustrate discriminated biobricks and inappropriately labeled biobricks. Clustering algorithm K-means is adopted to quantify the reduction results. The average clustering accuracy for Isomap and Laplacian Eigenmaps are 0.857 and 0.844, respectively. Besides, Laplacian Eigenmaps is 5 times faster than Isomap, and its visualization graph is more concentrated to discriminate biobricks. CONCLUSIONS: By combining normalized edit distance with Isomap and Laplacian Eigenmaps, synthetic biology biobircks are successfully visualized in two dimensional space. Various types of biobricks could be discriminated and inappropriately labeled biobricks could be determined, which could help to assess crowdsourcing-based synthetic biology databases' quality, and make biobricks selection. Jiaoyun Yang, Huitong Ding, Ning An 0001, Gil Alterovitz |
BMC Bioinform. | 4 |
| 2017 | Redundancy reduction for indoor device-free localization
Jinjun Liu, Ning An 0001, Md. Tanbir Hassan, Min Peng 0001 |
Pers. Ubiquitous Comput. | 2 |
| 2017 | Electric Vehicle Charging Station Placement for Urban Public Bus SystemsabstractDue to the low pollution and sustainable properties, using electric buses for public transportation systems has attracted considerable attention, whereas how to recharge the electric buses with long continuous service hours remains an open problem. In this paper, we consider the problem of placing electric vehicle (EV) charging stations at selected bus stops, to minimize the total installation cost of charging stations. Specifically, we study two EV charging station placement cases, with and without considering the limited battery size, which are called ECSP_LB and ECSP problems, respectively. The solution of the ECSP problem achieves the lower bound compared with the solution of the ECSP_LB problem, and the larger the battery size of the EV, the lower the overall cost of the charging station installation. For both cases, we prove that the placement problems under consideration are NP-hard and formulate them into integer linear programming. Specifically, for the ECSP problem we design a linear programming relaxation algorithm to get a suboptimal solution and derive an approximation ratio of the algorithm. Moreover, we derive the condition of the battery size when the ECSP problem can be applied. For the ECSP_LB problem, we show that, for a single bus route, the problem can be optimally solved with a backtracking algorithm, whereas for multiple bus routes we propose two heuristic algorithms, namely, multiple backtracking and greedy algorithms. Finally, simulation results show the effectiveness of the proposed schemes. Xiumin Wang 0004, Chau Yuen, Naveed Ul Hassan, Ning An 0001, Weiwei Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | Stochastic Multiview Hashing for Large-Scale Near-Duplicate Video RetrievalabstractNear-duplicate video retrieval (NDVR) has been a significant research task in multimedia given its high impact in applications, such as video search, recommendation, and copyright protection. In addition to accurate retrieval performance, the exponential growth of online videos has imposed heavy demands on the efficiency and scalability of the existing systems. Aiming at improving both the retrieval accuracy and speed, we propose a novel stochastic multiview hashing algorithm to facilitate the construction of a large-scale NDVR system. Reliable mapping functions, which convert multiple types of keyframe features, enhanced by auxiliary information such as video-keyframe association and ground truth relevance to binary hash code strings, are learned by maximizing a mixture of the generalized retrieval precision and recall scores. A composite Kullback-Leibler divergence measure is used to approximate the retrieval scores, which aligns stochastically the neighborhood structures between the original feature and the relaxed hash code spaces. The efficiency and effectiveness of the proposed method are examined using two public near-duplicate video collections and are compared against various classical and state-of-the-art NDVR systems. Yanbin Hao, Tingting Mu, Richang Hong, Meng Wang 0001, Ning An 0001, John Yannis Goulermas |
IEEE Trans. Multim. | 5 |
| 2016 | HB-Phone: A Bed-Mounted Geophone-Based Heartbeat Monitoring SystemabstractHeartbeat monitoring during sleep is critically important to ensuring the well-being of many people, ranging from patients to elderly. Technologies that support heartbeat monitoring should be unobtrusive, and thus solutions that are accurate and can be easily applied to existing beds is an important need that has been unfulfilled. We tackle the challenge of accurate, low-cost and easy to deploy heartbeat monitoring by investigating whether off-the- shelf analog geophone sensors can be used to detect heartbeats when installed under a bed. Geophones have the desirable property of being insensitive to lower-frequency movements, which lends itself to heartbeat monitoring as the heartbeat signal has harmonic frequencies that are easily captured by the geophone. At the same time, lower-frequency movements such as respiration, can be naturally filtered out by the geophone. With carefully-designed signal processing algorithms, we show it is possible to detect and extract heartbeats in the presence of environmental noise and other body movements a person may have during sleep. We have built a prototype sensor and conducted detailed experiments that involve 43 subjects (with IRB approval), which demonstrate that the geophone sensor is a compelling solution to long-term at-home heartbeat monitoring. We compared the average heartbeat rate estimated by our prototype and that reported by a pulse oximeter. The results revealed that the average error rate is around 1.30% over 500 data samples when the subjects were still on the bed, and 3.87% over 300 data samples when the subjects had different types of body movements while lying on the bed. We also deployed the prototype in the homes of 9 subjects for a total of 25 nights, and found that the average estimation error rate was 8.25% over more than 181 hours' data. Zhenhua Jia, Musaab Alaziz, Xiang Chi, Richard E. Howard, Yanyong Zhang, Pei Zhang 0001, Wade Trappe, Anand Sivasubramaniam, Ning An 0001 |
IPSN | 9 |
| 2016 | Poster Abstract: Node Deployment Mechanism for Quick, Indoor, and Device-Free LocalizationabstractRadio frequency based device-free passive localization has been actively studied. An important issue with this technique is that a large amount of sensor nodes are usually deployed in localization environments, causing high deployment cost, high computation complexity and other drawbacks. To reduce the number of nodes while maintaining high localization accuracy, this paper proposes a fingerprint-based node deployment mechanism, which saves hardware cost, speeds up node deployment process as well as maintains a similar accuracy. The experiments are conducted in two environments with considerably different sizes of 150 m2and 25 m2. Maintaining the accuracies above 95%, the proposed mechanism can efficiently reduce the total number of nodes in large room by 54.55% and in small room by 68.75%. Especially, the localization response time sharply reduces by 93.08% in the large room and 79.87% in the small room. Jinjun Liu, Ning An 0001, Md. Tanbir Hassan, Guilin Chen, Yanyong Zhang |
IPSN | 2 |
| 2016 | A Partial Correlation Statistic Structure Learning Algorithm Under Linear Structural Equation ModelsabstractA new algorithm, the Partial Correlation Statistic (PCS) algorithm, is presented for structure learning under linear Structural Equation Models. The PCS algorithm can deal with continuous data following linear arbitrary distribution rather than only a Gaussian distribution. This paper makes two specific contributions. First, for linear arbitrarily distributed datasets, which are generated by the linear structural equation models, if the sample size is sufficiently large, partial correlation coefficient statistic is proved to follow a Student's t-distribution. Second, the PCS algorithm combines hypothesis testing of partial correlation statistic and local learning to select potential neighbors of the target node. This significantly reduces the search space and achieves good time performance. The PCS algorithm does not need to choose optimal threshold of partial correlation by large amount of experiments. Especially, the PCS algorithm redefines the relevance from statistic theory and measure the relevance of the variables based on$p$-value. The effectiveness of the algorithm is compared with current state of the art methods on seven networks. A simulation shows that the PCS algorithm outperforms existing algorithms in terms of both accuracy and time performance on average. Jing Yang 0008, Ning An 0001, Gil Alterovitz |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | Harmony search for feature selection in speech emotion recognitionabstractFeature selection is a significant aspect of speech emotion recognition system. How to select a small subset out of the thousands of speech data is important for accurate classification of speech emotion. In this paper we investigate heuristic algorithm Harmony search (HS) for feature selection. We extract 3 feature sets, including MFCC, Fourier Parameters (FP), and features extracted with The Munich open Speech and Music Interpretation by Large Space Extraction (openSMILE) toolkit, from Berlin German emotion database (EMODB) and Chinese Elderly emotion database (EESDB). And combine MFCC with FP as the fourth feature set. We use Harmony search to select subsets and decrease the dimension space, and employ 10-fold cross validation in LIBSVM to evaluate the change of accuracy between selected subsets and original sets. Experimental results show that each subset's size reduced by about 50%, however, there is no sharp degeneration on accuracy and the accuracy almost maintains the original ones. Yongsen Tao, Kunxia Wang, Jing Yang 0008, Ning An 0001, Lian Li 0001 |
ACII | 4 |
| 2015 | Predicting hypertension without measurement: A non-invasive, questionnaire-based approach
Aiguo Wang 0002, Ning An 0001, Guilin Chen, Lian Li 0001, Gil Alterovitz |
Expert Syst. Appl. | 2 |
| 2015 | Accelerating wrapper-based feature selection with K-nearest-neighbor
Aiguo Wang 0002, Ning An 0001, Guilin Chen, Lian Li 0001, Gil Alterovitz |
Knowl. Based Syst. | 2 |
| 2015 | Speech Emotion Recognition Using Fourier ParametersabstractRecently, studies have been performed on harmony features for speech emotion recognition. It is found in our study that the first- and second-order differences of harmony features also play an important role in speech emotion recognition. Therefore, we propose a new Fourier parameter model using the perceptual content of voice quality and the first- and second-order differences for speaker-independent speech emotion recognition. Experimental results show that the proposed Fourier parameter (FP) features are effective in identifying various emotional states in speech signals. They improve the recognition rates over the methods using Mel frequency cepstral coefficient (MFCC) features by 16.2, 6.8 and 16.6 points on the German database (EMODB), Chinese language database (CASIA) and Chinese elderly emotion database (EESDB). In particular, when combining FP with MFCC, the recognition rates can be further improved on the aforementioned databases by 17.5, 10 and 10.5 points, respectively. Kunxia Wang, Ning An 0001, Bing Nan Li, Yanyong Zhang, Lian Li 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | Accelerating incremental wrapper based gene selection with K-Nearest-NeighborabstractWrapper based gene selection methods tend to obtain better classification accuracy than filter methods, while it is much more time consuming. Accelerating this process without degrading the high accuracy is of great value for researchers to better analyze gene expression profiles. In this paper, we explore to reduce the time complexity of wrapper based gene selection method with K-Nearest-Neighbor (KNN) classifier embedded. Instead of taking KNN as a black box, we incrementally construct and maintain a classifier distance matrix to speed up the gene selection process. Experiments on six publicly available microarrays were first conducted to show the effectiveness of incremental wrapper based gene selection method with KNN. Then, to demonstrate the performance gain in time cost reduction, we analyzed the time complexity and experimentally evaluated it. Both theoretical analysis and experimental results prove that the proposed approach greatly accelerates the gene selection process without degrading the classification accuracy. Aiguo Wang 0002, Ning An 0001, Guilin Chen, Lian Li 0001, Gil Alterovitz |
BIBM | 2 |
| 2014 | Incremental wrapper based gene selection with Markov blanketabstractGene selection plays a crucial role in the analysis of microarray data with high dimensionality and small sample size. Incremental wrapper based feature subset selection (FSS) methods, among various feature selection approaches, tend to obtain high quality feature subset and better classification accuracy than filter methods, while it is much more time consuming since the interdependence and redundancy between features is evaluated in a wrapper way. In this paper, we explore to introduce Markov Blanket (MB) into incremental wrapper based FSS process. Rather than evaluate the quality of all the features ranked by a filter method, our proposal eliminates features that are redundant to the newly selected one via MB during the wrapper evaluation process to reduce the number of wrappers, enabling us to select the relevant features and eliminate redundant ones efficiently. To verify the effectiveness and efficiency of the proposed approach, experimental comparisons on six publicly available microarray data are conducted with two typical classifiers with different metrics, Naïve Bayes and 1-Nearest-Neighbor. Experimental results demonstrate that our approach greatly speeds up the feature selection process, obtains more compact feature subset and achieves better classification accuracy compared to that without MB for both two-category and multi-category problems. Aiguo Wang 0002, Ning An 0001, Guilin Chen, Jing Yang 0008, Lian Li 0001, Gil Alterovitz |
BIBM | 2 |
| 2014 | Comprehensive assessment of gait signals using multiple time scale featuresabstractIt is a challenging problem to detect and analyze gait signals for health evaluation. In this article, we propose a comprehensive assessment method using multiple time scale features to extract gait signal characteristics. Multi-resolution wavelet transform, together with logic regression and correlation analysis, was adapted for statistical analysis. The results show that the primary period and autocorrelation of gait signals vary substantially in three cohorts of people, namely normal young people, healthy old people and those with Parkinson's diseases. Furthermore, it is found that there is a correlation between the periodicity of gait sequences and the degree of Parkinson's diseases. In conclusion, these multiple scale features are very useful for health evaluation. Xi Wu 0003, Huitong Ding, Bing Nan Li, Ning An 0001 |
BIBM | 4 |
| 2014 | Toward detection of aliases without string similarity
Ning An 0001, Lili Jiang 0002, Jianyong Wang 0001, Ping Luo 0001, Min Wang 0001, Bing Nan Li |
Inf. Sci. | 1 |
| 2014 | Optimal parameters selection for BP neural network based on particle swarm optimization: A case study of wind speed forecasting
Chao Ren 0009, Ning An 0001, Lian Li 0001, Bin Hu 0001, Duo Shang |
Knowl. Based Syst. | 2 |
| 2013 | Causal discovery based on healthcare informationabstractCorrectly discovering causal relations from healthcare information can help people to understand disease mechanisms and discover disease causes. In some cases, the healthcare data do not follow a multivariate Gaussian distribution. We design a new causal structure learning algorithm. The algorithm can effectively combines ideas from local learning with simultaneous equations models techniques. In the first phase of the algorithm we select potential neighbors for each variable based on simultaneous equations models, and then perform a constrained hill-climbing search to orient the edges. Using the algorithm without prior knowledge, we analyze causal relations in the real data from the National Health and Nutrition Examination Survey. Jing Yang 0008, Ning An 0001, Gil Alterovitz, Lian Li 0001, Aiguo Wang 0002 |
BIBM | 2 |
| 2013 | GRIAS: An Entity-Relation Graph Based Framework for Discovering Entity AliasesabstractRecognizing the various aliases of an entity is a critical task for many applications, including Web search, recommendation system, and e-discovery. The goal of this paper is to accurately identify entity aliases, especially the long tail ones in the unstructured data. Our solution GRIAS (abbr. for a Graph-based framework for discovering entity Aliases) is motivated by the entity relationships collected from both the structured and unstructured data. These relationships help to build an entity-relation graph, and the graph-based similarity is calculated between an entity and its alias candidates which are first chosen by our proposed candidate selection method. Extensive experimental results on two real-world datasets demonstrate both the effectiveness and efficiency of the proposed framework. Lili Jiang 0002, Ping Luo 0001, Jianyong Wang 0001, Yuhong Xiong, Bingduan Lin, Min Wang 0001, Ning An 0001 |
ICDM | 7 |
| 2013 | SCPL: indoor device-free multi-subject counting and localization using radio signal strengthabstractRadio frequency based device-free passive (DfP) localization techniques have shown great potentials in localizing individual human subjects, without requiring them to carry any radio devices. In this study, we extend the DfP technique to count and localize multiple subjects in indoor environments. To address the impact of multipath on indoor radio signals, we adopt a fingerprinting based approach to infer subject locations from observed signal strengths through profiling the environment. When multiple subjects are present, our objective is to use the profiling data collected by a single subject to count and localize multiple subjects without any extra effort. In order to address the non-linearity of the impact of multiple subjects, we propose a successive cancellation based algorithm to iteratively determine the number of subjects. We model indoor human trajectories as a state transition process, exploit indoor human mobility constraints and integrate all information into a conditional random field (CRF) to simultaneously localize multiple subjects. As a result, we call the proposed algorithm SCPL -- sequential counting, parallel localizing. We test SCPL with two different indoor settings, one with size 150 m2 and the other 400 m2. In each setting, we have four different subjects, walking around in the deployed areas, sometimes with overlapping trajectories. Through extensive experimental results, we show that SCPL can count the present subjects with 86% counting percentage when their trajectories are not completely overlapping. Our localization algorithms are also highly accurate, with an average localization error distance of 1.3 m. Chenren Xu, Bernhard Firner, Robert S. Moore, Yanyong Zhang, Wade Trappe, Richard E. Howard, Feixiong Zhang, Ning An 0001 |
IPSN | 8 |
| 2013 | Human Dynamics Revealed through Log Analytics in a Cloud Computing Environment
Sixi Chen, Ning An 0001, Lian Li 0001, Yongwei Wu 0001 |
WAIM | 2 |
| 2012 | Towards alias detection without string similarity: an active learning based approachabstractEntity aliases commonly exist and accurately detecting these aliases plays a vital role in various applications. In this paper, we use an active-learning-based method to detect aliases without string similarity. To minimize the cost on pairwise comparison, a subset-based method restricts the alias selection within a small-scale entity set. Within each generated entity set, an active learning based logistic regression classifier is employed to predict whether a candidate is the alias of a given entity. The experimental results on three datasets clearly demonstrate that our proposed approach can effectively detect this kind of entity aliases. Lili Jiang 0002, Jianyong Wang 0001, Ping Luo 0001, Ning An 0001, Min Wang 0001 |
SIGIR | 4 |
| 2010 | GRAPE: a system for disambiguating and tagging people names in web searchabstractName ambiguity is a big challenge in people information retrieval and has received considerable attention, especially with the increasing volume of Web data in recent years. In this demo, we present a system, GRAPE, which is capable of finding people related information over the Web. The salient features of our system are people name disambiguation and people tag presentation, which effectively distinguish different people entities sharing the same name and uniquely represent each namesake with a cluster of tags, such as occupation, birthdate, and organization. Lili Jiang 0002, Wei Shen 0004, Jianyong Wang 0001, Ning An 0001 |
WWW | 4 |
| 2009 | GRAPE: A Graph-Based Framework for Disambiguating People Appearances in Web SearchabstractFinding information about people using search engines is one of the most common activities on the Web. However, search engines usually return a long list of Web pages, which may be relevant to many namesakes, especially given the explosive growth of Web data. To address the challenge caused by name ambiguity in Web people search, this paper proposes a novel graph-based framework, GRAPE (abbr. a Graph-based fRamework for disAmbiguating People appEarances in Web search). In GRAPE, people tag information (e.g., people name, organization, and email address) surrounding the queried people name is extracted from the search results, a graph-based unsupervised algorithm is then developed to cluster the extracted tags, where a new method, Cohesion, is introduced to measure the importance of a tag for clustering, and each final cluster of tags represents a unique people entity. Experimental results show that our proposed framework outperforms the state-of-the-art Web people name disambiguation approaches. Lili Jiang 0002, Jianyong Wang 0001, Ning An 0001, Shengyuan Wang 0001, Jian Zhan, Lian Li 0001 |
ICDM | 3 |
| 2009 | Two birds with one stone: a graph-based framework for disambiguating and tagging people names in web searchabstractThe ever growing volume of Web data makes it increasingly challenging to accurately find relevant information about a specific person on the Web. To address the challenge caused by name ambiguity in Web people search, this paper explores a novel graph-based framework to both disambiguate and tag people entities in Web search results. Experimental results demonstrate the effectiveness of the proposed framework in tag discovery and name disambiguation. Lili Jiang 0002, Jianyong Wang 0001, Ning An 0001, Shengyuan Wang 0001, Jian Zhan, Lian Li 0001 |
WWW | 3 |
| 2003 | Toward an Accurate Analysis of Range Queries on Spatial DataabstractAnalysis of range queries on spatial (multidimensional) data is both important and challenging. Most previous analysis attempts have made certain simplifying assumptions about the data sets and/or queries to keep the analysis tractable. As a result, they may not be universally applicable. This paper proposes a set of five analysis techniques to estimate the selectivity and number of index nodes accessed in serving a range query. The underlying philosophy behind these techniques is to maintain an auxiliary data structure, called a density file, whose creation is a one-time cost, which can be quickly consulted when the query is given. The schemes differ in what information is kept in the density file, how it is maintained, and how this information is looked up. It is shown that one of the proposed schemes, called cumulative density (CD), gives very accurate results (usually less than 5 percent error) using a diverse suite of point and rectangular data sets, that are uniform or skewed, and a wide range of query window parameters. The estimation takes a constant amount of time, which is typically lower than 1 percent of the time that it would take to execute the query, regardless of data set or query window parameters. Ning An 0001, Ji Jin, Anand Sivasubramaniam |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2002 | Energy-performance trade-offs for spatial access methods on memory-resident data
Ning An 0001, Sudhanva Gurumurthi, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin |
VLDB J. | 1 |
| 2001 | Selectivity Estimation for Spatial JoinsabstractSpatial joins are important and time consuming operations in spatial database management systems. It is crucial to be able to accurately estimate the performance of these operations so that one can derive efficient query execution plans, and even develop/refine data structures to improve their performance. While estimation techniques for analyzing the performance of other operations, such as range queries, on spatial data has come under scrutiny, the problem of estimating selectivity for spatial joins has been little explored. The limited forays into this area have used parametric techniques, which are largely restrictive on the datasets that they can be used for since they tend to make simplifying assumptions about the nature of the datasets to be joined. Sampling and histogram based techniques, on the other hand, are much less restrictive. However, there has been no prior attempt at understanding the accuracy of sampling techniques, or developing histogram based techniques to estimate the selectivity of spatial joins. Apart from extensively evaluating the accuracy of sampling techniques for the very first time, this paper presents two novel histogram based solutions for spatial join estimation. Using a wide spectrum of both real and synthetic datasets, it is shown that one of our proposed schemes, called Geometric Histograms (GH), can accurately quantify the selectivity of spatial joins. Ning An 0001, Zhen-Yu Yang, Anand Sivasubramaniam |
ICDE | 1 |
| 2001 | Analyzing energy behavior of spatial access methods for memory-resident data
Ning An 0001, Anand Sivasubramaniam, Narayanan Vijaykrishnan, Mahmut T. Kandemir, Mary Jane Irwin, Sudhanva Gurumurthi |
VLDB | 1 |
| 2000 | Analyzing Range Queries on Spatial DataabstractAnalysis of range queries on spatial (multidimensional) data is both important and challenging. Most previous analysis attempts have made certain simplifying assumptions about the data sets and/or queries to keep the analysis tractable. As a result, they may no be universally applicable. This paper proposes a set of five analysis techniques to estimate the selectivity and number of index nodes accessed in serving a range query. The underlying philosophy behind these techniques is to maintain an auxiliary data structure called a density file, whose creation is a one-time cost, which can be quickly consulted when the query is given. The schemes differ in what information is kept in the density file, how it is maintained and how this information is looked up. It is shown that one of the proposed schemes, called "cumulative density" (CD), gives very accurate results (usually less then 5% error) using a diverse suite of point and rectangular data sets, that are uniform or skewed, and a wide range of query window parameters. The estimation takes a constant amount of time, which is typically lower than 1% of the time that it would take to execute the query, regardless of data set or query window parameters. Ji Jin, Ning An 0001, Anand Sivasubramaniam |
ICDE | 2 |