EDBT 2026 Demo / reviewers in the wild / expert
Kotagiri Ramamohanarao
dblp:r/KRamamohanarao · also Ramamohanarao Kotagiri, Rao Kotagiri
· DBLP profile ↗
296ranked-venue papers
16as first author
16since 2021 · last 2024
0000-0003-3304-9268ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 157 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 102 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 34 · 3 first-author · 5 since 2021Systems, architecture and hardware · 25 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21Software engineering, systems software and programming languages · 15 · 2 first-authorComputer networks · 10 · 1 since 2021Theory of computation · 8 · 3 first-authorSecurity and privacy · 7Human-computer interaction and ubiquitous computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Collaborative Knowledge Distillation via Multiknowledge TransferabstractKnowledge distillation (KD), as an efficient and effective model compression technique, has received considerable attention in deep learning. The key to its success is about transferring knowledge from a large teacher network to a small student network. However, most existing KD methods consider only one type of knowledge learned from either instance features or relations via a specific distillation strategy, failing to explore the idea of transferring different types of knowledge with different distillation strategies. Moreover, the widely used offline distillation also suffers from a limited learning capacity due to the fixed large-to-small teacher-student architecture. In this article, we devise a collaborative KD via multiknowledge transfer (CKD-MKT) that prompts both self-learning and collaborative learning in a unified framework. Specifically, CKD-MKT utilizes a multiple knowledge transfer framework that assembles self and online distillation strategies to effectively: 1) fuse different kinds of knowledge, which allows multiple students to learn knowledge from both individual instances and instance relations, and 2) guide each other by learning from themselves using collaborative and self-learning. Experiments and ablation studies on six image datasets demonstrate that the proposed CKD-MKT significantly outperforms recent state-of-the-art methods for KD. Jianping Gou, Liyuan Sun 0005, Baosheng Yu, Lan Du 0002, Kotagiri Ramamohanarao, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | An Optimal Online Semi-connected PLA Algorithm with Maximum Error Bound (Extended Abstract)abstractPiecewise Linear Approximation (PLA) is one of the most widely used approaches for representing a time series with a set of approximated line segments. With this compressed form of representation, many large complicated time series can be efficiently stored, transmitted and analyzed. In this article, with the introduced concept of "semi-connection" that allowing two representation lines to be connected at a point between two consecutive time stamps, we propose a new optimal linear-time PLA algorithm SemiOptConnAlg for generating the least number of semi-connected line segments with guaranteed maximum error bound. With extended experimental tests, we demonstrate that the proposed algorithm is very efficient in execution and achieves better performances than the state-of-art solutions. Huanyu Zhao, Chaoyi Pang, Kotagiri Ramamohanarao, Christopher Kuo Pang, Jian Yang 0001, Tongliang Li |
ICDE | 3 |
| 2022 | Modelling Zeros in Blockmodelling
Laurence Anthony F. Park, Mohadeseh Ganji, Emir Demirovic, Jeffrey Chan, Peter J. Stuckey, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao |
PAKDD (2) | 8 |
| 2022 | MurTree: Optimal Decision Trees via Dynamic Programming and SearchabstractDecision tree learning is a widely used approach in machine learning, favoured in applications that require concise and interpretable models. Heuristic methods are traditionally used to quickly produce models with reasonably high accuracy. A commonly criticised point, however, is that the resulting trees may not necessarily be the best representation of the data in terms of accuracy and size. In recent years, this motivated the development of optimal classification tree algorithms that globally optimise the decision tree in contrast to heuristic methods that perform a sequence of locally optimal decisions. We follow this line of work and provide a novel algorithm for learning optimal classification trees based on dynamic programming and search. Our algorithm supports constraints on the depth of the tree and number of nodes. The success of our approach is attributed to a series of specialised techniques that exploit properties unique to classification trees. Whereas algorithms for optimal classification trees have traditionally been plagued by high runtimes and limited scalability, we show in a detailed experimental study that our approach uses only a fraction of the time required by the state-of-the-art and can handle datasets with tens of thousands of instances, providing several orders of magnitude improvements and notably contributing towards the practical use of optimal decision trees. Emir Demirovic, Anna Lukina, Emmanuel Hebrard, Jeffrey Chan, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Peter J. Stuckey |
J. Mach. Learn. Res. | 7 |
| 2022 | MRMondrian: Scalable Multidimensional Anonymisation for Big Data Privacy PreservationabstractScalable data processing platforms built on cloud computing becomes increasingly attractive as infrastructure for supporting big data applications. But privacy concerns are one of the major obstacles to making use of public cloud platforms. Multidimensional anonymisation, a global-recoding generalisation scheme for privacy-preserving data publishing, has been a recent focus due to its capability of balancing data obfuscation and usability. Existing multidimensional anonymisation methods suffer from scalability problems when handling big data due to the impractical serial I/O cost. Given the recursive feature of multidimensional anonymisation, parallelisation is an ideal solution to scalability issues. However, it is still a challenge to use existing distributed and parallel paradigms directly for recursive computation. In this paper, we propose a scalable approach for big data multidimensional anonymisation based on MapReduce, a state-of-the-art data processing paradigm. Our basic idea is to partition a data set recursively into smaller partitions using MapReduce until all partitions can fit in the memory of a computing node. A tree indexing structure is proposed to achieve recursive computation. Moreover, we show the applicability of our approach to differential privacy. Experimental results on real-life data demonstrate that our approach can significantly improve the scalability of multidimensional anonymisation over existing methods. Xuyun Zhang, Lianyong Qi, Wan-Chun Dou, Qiang He 0001, Christopher Leckie, Kotagiri Ramamohanarao, Zoran A. Salcic |
IEEE Trans. Big Data | 6 |
| 2022 | An Optimal Online Semi-Connected PLA Algorithm With Maximum Error BoundabstractPiecewise Linear Approximation (PLA) is one of the most widely used approaches for representing a time series with a set of approximated line segments. With this compressed form of representation, many large complicated time series can be efficiently stored, transmitted and analyzed. In this article, with the introduced concept of “semi-connection” that allowing two representation lines to be connected at a point between two consecutive time stamps, we propose a new optimal linear-time PLA algorithm SemiOptConnAlg for generating the least number of semi-connected line segments with guaranteed maximum error bound. With extended experimental tests, we demonstrate that the proposed algorithm is very efficient in execution time and achieves better performances than the state-of-art solutions. Huanyu Zhao, Chaoyi Pang, Kotagiri Ramamohanarao, Christopher Kuo Pang, Jian Yang 0001, Tongliang Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Dynamic Scheduling for Stochastic Edge-Cloud Computing Environments Using A3C Learning and Residual Recurrent Neural NetworksabstractThe ubiquitous adoption of Internet-of-Things (IoT) based applications has resulted in the emergence of the Fog computing paradigm, which allows seamlessly harnessing both mobile-edge and cloud resources. Efficient scheduling of application tasks in such environments is challenging due to constrained resource capabilities, mobility factors in IoT, resource heterogeneity, network hierarchy, and stochastic behaviors. Existing heuristics and Reinforcement Learning based approaches lack generalizability and quick adaptability, thus failing to tackle this problem optimally. They are also unable to utilize the temporal workload patterns and are suitable only for centralized setups. However, asynchronous-advantage-actor-critic (A3C) learning is known to quickly adapt to dynamic scenarios with less data and residual recurrent neural network (R2N2) to quickly update model parameters. Thus, we propose an A3C based real-time scheduler for stochastic Edge-Cloud environments allowing decentralized learning, concurrently across multiple agents. We use the R2N2 architecture to capture a large number of host and task parameters together with temporal patterns to provide efficient scheduling decisions. The proposed model is adaptive and able to tune different hyper-parameters based on the application requirements. We explicate our choice of hyper-parameters through sensitivity analysis. The experiments conducted on real-world data set show a significant improvement in terms of energy consumption, response time, Service-Level-Agreement and running cost by 14.4, 7.74, 31.9, and 4.64 percent, respectively when compared to the state-of-the-art algorithms. Shreshth Tuli, Shashikant Ilager, Kotagiri Ramamohanarao, Rajkumar Buyya |
IEEE Trans. Mob. Comput. | 3 |
| 2021 | Effective Traffic Forecasting with Multi-Resolution LearningabstractTraffic forecasting plays a vital role in traffic management systems. Recently, deep learning models have been applied to citywide traffic forecasting. However, the existing work models and predicts traffic at a single (dense) resolution, making it challenging to capture long-range spatial dependencies or high-level traffic dynamics. This shortcoming limits the accuracy of prediction and results in computationally expensive models. We propose a traffic forecasting model based on deep convolutional networks to improve the accuracy of citywide traffic forecasting. Our model uses a hierarchical architecture that captures traffic dynamics at multiple spatial resolutions. Based on this architecture, we apply a multi-task learning scheme, which trains the model to predict traffic at different resolutions. Our model helps provide a coherent understanding of traffic dynamics by capturing spatial dependencies between different regions of a city. Experimental results on multiple real datasets show that our model can achieve competitive results compared to complex state-of-the-art approaches while being more computationally efficient. Abdullah AlDwyish, Egemen Tanin, Hairuo Xie, Shanika Karunasekera, Kotagiri Ramamohanarao |
SSTD | 5 |
| 2021 | eQTLHap: a tool for comprehensive eQTL analysis considering haplotypic and genotypic effectsabstractMOTIVATION: The high accuracy of recent haplotype phasing tools is enabling the integration of haplotype (or phase) information more widely in genetic investigations. One such possibility is phase-aware expression quantitative trait loci (eQTL) analysis, where haplotype-based analysis has the potential to detect associations that may otherwise be missed by standard SNP-based approaches. RESULTS: We present eQTLHap, a novel method to investigate associations between gene expression and genetic variants, considering their haplotypic and genotypic effect. Using multiple simulations based on real data, we demonstrate that phase-aware eQTL analysis significantly outperforms typical SNP-based methods when the causal genetic architecture involves multiple SNPs. We show that phase-aware eQTL analysis is robust to phasing errors, showing only a minor impact ($<4\%$) on sensitivity. Applying eQTLHap to real GEUVADIS and GTEx datasets detects numerous novel eQTLs undetected by a single-SNP approach, with 22 eQTLs replicating across studies or tissue types, highlighting the utility of phase-aware eQTL analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/ziadbkh/eQTLHap. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Briefings in Bioinformatics online. Ziad Al Bkhetan, Gursharan Chana, Cheng Soon Ong, Benjamin Goudey, Kotagiri Ramamohanarao |
Briefings Bioinform. | 5 |
| 2021 | Evaluation of consensus strategies for haplotype phasingabstractHaplotype phasing is a critical step for many genetic applications but incorrect estimates of phase can negatively impact downstream analyses. One proposed strategy to improve phasing accuracy is to combine multiple independent phasing estimates to overcome the limitations of any individual estimate. However, such a strategy is yet to be thoroughly explored. This study provides a comprehensive evaluation of consensus strategies for haplotype phasing. We explore the performance of different consensus paradigms, and the effect of specific constituent tools, across several datasets with different characteristics and their impact on the downstream task of genotype imputation. Based on the outputs of existing phasing tools, we explore two different strategies to construct haplotype consensus estimators: voting across outputs from multiple phasing tools and multiple outputs of a single non-deterministic tool. We find that the consensus approach from multiple tools reduces SE by an average of 10% compared to any constituent tool when applied to European populations and has the highest accuracy regardless of population ethnicity, sample size, variant density or variant frequency. Furthermore, the consensus estimator improves the accuracy of the downstream task of genotype imputation carried out by the widely used Minimac3, pbwt and BEAGLE5 tools. Our results provide guidance on how to produce the most accurate phasing estimates and the trade-offs that a consensus approach may have. Our implementation of consensus haplotype phasing, consHap, is available freely at https://github.com/ziadbkh/consHap. Supplementary information: Supplementary data are available at Briefings in Bioinformatics online. Ziad Al Bkhetan, Gursharan Chana, Kotagiri Ramamohanarao, Karin Verspoor, Benjamin Goudey |
Briefings Bioinform. | 3 |
| 2021 | Route intersection reduction with connected autonomous vehicles
Sadegh Motallebi, Hairuo Xie, Egemen Tanin, Jianzhong Qi 0001, Kotagiri Ramamohanarao |
GeoInformatica | 5 |
| 2021 | A differentially private algorithm for range queries on trajectories
Soheila Ghane, Lars Kulik, Kotagiri Ramamohanarao |
Knowl. Inf. Syst. | 3 |
| 2021 | Hedonic Pricing of Cloud Computing ServicesabstractCloud service providers (CSP) and cloud consumers often need to forecast the cloud price to optimize their business strategy. However, pricing of cloud services is a challenging task due to its services complexity and dynamic nature of the ever-changing environment. Moreover, the cloud pricing based on consumers' willingness to pay (W2P) becomes even more challenging due to the subjectiveness of consumers' experiences and implicit values of some non-marketable features, such as burstable CPU, dedicated server, and cloud data center global footprints. Unfortunately, many existing pricing models often cannot support value-based pricing. In this paper, we propose a novel solution based on value-based pricing, which does not only consider how much does the service cost (or intrinsic values) to a CSP but also how much a customer is willing to pay (or extrinsic values) for the service. We demonstrate that the cloud extrinsic values would not only become one of the competitive advantages for CSPs to lead the cloud market but also increase the profit margin. Our approach is often referred to as a hedonic pricing model. We show that our model can capture the value of non-marketable features. This value is about 43.4 percent on average above the baseline, which is often ignored by many traditional cloud pricing models. We also show that Average Annual Growth Rate (AAGR) of Amazon Web Services' (AWS) is about -20.0 percent per annum between 2008 and 2017, ceteris paribus. In comparison with Moore's law (-50 percent per annum), it is at a far slower pace. We argue this value is Moore's law equivalent in the cloud. The primary goal of this research is to provide a less biased pricing model for cloud decision makers to develop their optimizing investment strategy. Caesar Wu, Adel Nadjaran Toosi, Rajkumar Buyya, Kotagiri Ramamohanarao |
IEEE Trans. Cloud Comput. | 4 |
| 2021 | Preserving Privacy in the Internet of Connected VehiclesabstractToday's vehicles are advancing from stand-alone transportation means to vehicle-to-vehicle, and vehicle-to-infrastructure communications enabled devices which are able to exchange data through the transportation communication infrastructure. As the IoT and data remain intrinsically linked together, the fast-changing mobility landscape of intent-based networking for the Internet of connected vehicles comes with a great risk of data security and privacy violations. This paper considers the privacy issues in the distributed edge computing, in which the data is communicated between a number of vehicles in the IoT layer and potentially untrusted edge controllers at the edge of the network. The sensory data communicated by the vehicles contain sensitive information, such as location and speed, which could violate the users' privacy if they are leaked with no perturbation. Recent studies suggest mechanisms for randomizing the stream of data to ensure individuals' privacy. Although the past works on differential privacy provide a strong privacy guarantee, they are limited to applications where communication parties are trusted and/or there is no correlation between the users or the featured of sensory data. In this paper, we address this gap by proposing a differentially private data streaming system that adds a correlated noise in the vehicle's side (IoT layer) rather than the transportation infrastructure. Also, our system is able to ensure a strong privacy level over time. The proposed mechanism is data-adaptive and scales the noise with respect to the data correlation. Our extensive experiments demonstrate that the utility of the output generated by our method outperforms the recent approaches. Soheila Ghane, Alireza Jolfaei, Lars Kulik, Kotagiri Ramamohanarao, Deepak Puthal |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Thermal Prediction for Efficient Energy Management of Clouds Using Machine LearningabstractThermal management in the hyper-scale cloud data centers is a critical problem. Increased host temperature creates hotspots which significantly increases cooling cost and affects reliability. Accurate prediction of host temperature is crucial for managing the resources effectively. Temperature estimation is a non-trivial problem due to thermal variations in the data center. Existing solutions for temperature estimation are inefficient due to their computational complexity and lack of accurate prediction. However, data-driven machine learning methods for temperature prediction is a promising approach. In this regard, we collect and study data from a private cloud and show the presence of thermal variations. We investigate several machine learning models to accurately predict the host temperature. Specifically, we propose a gradient boosting machine learning model for temperature prediction. The experiment results show that our model accurately predicts the temperature with the average RMSE value of 0.05 or an average prediction error of 2.38 °C, which is 6 °C less as compared to an existing theoretical model. In addition, we propose a dynamic scheduling algorithm to minimize the peak temperature of hosts. The results show that our algorithm reduces the peak temperature by 6.5 °C and consumes 34.5 percent less energy as compared to the baseline algorithm. Shashikant Ilager, Kotagiri Ramamohanarao, Rajkumar Buyya |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2021 | ADRL: A Hybrid Anomaly-Aware Deep Reinforcement Learning-Based Resource Scaling in CloudsabstractThe virtualization concept and elasticity feature of cloud computing enable users to request resources on-demand and in the pay-as-you-go model. However, the high flexibility of the model makes the on-time resource scaling problem more complex. A variety of techniques such as threshold-based rules, time series analysis, or control theory are utilized to increase the efficiency of dynamic scaling of resources. However, the inherent dynamicity of cloud-hosted applications requires autonomic and adaptable systems that learn from the environment in real-time. Reinforcement Learning (RL) is a paradigm that requires some agents to monitor the surroundings and regularly perform an action based on the observed states. RL has a weakness to handle high dimensional state space problems. Deep-RL models are a recent breakthrough for modeling and learning in complex state space problems. In this article, we propose a Hybrid Anomaly-aware Deep Reinforcement Learning-based Resource Scaling (ADRL) for dynamic scaling of resources in the cloud. ADRL takes advantage of anomaly detection techniques to increase the stability of decision-makers by triggering actions in response to the identified anomalous states in the system. Two levels of global and local decision-makers are introduced to handle the required scaling actions. An extensive set of experiments for different types of anomaly problems shows that ADRL can significantly improve the quality of service with less number of actions and increased stability of the system. Sara Kardani-Moghaddam, Rajkumar Buyya, Kotagiri Ramamohanarao |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Dynamic Programming for Predict+OptimiseabstractWe study the predict+optimise problem, where machine learning and combinatorial optimisation must interact to achieve a common goal. These problems are important when optimisation needs to be performed on input parameters that are not fully observed but must instead be estimated using machine learning. We provide a novel learning technique for predict+optimise to directly reason about the underlying combinatorial optimisation problem, offering a meaningful integration of machine learning and optimisation. This is done by representing the combinatorial problem as a piecewise linear function parameterised by the coefficients of the learning model and then iteratively performing coordinate descent on the learning coefficients. Our approach is applicable to linear learning functions and any optimisation problem solvable by dynamic programming. We illustrate the effectiveness of our approach on benchmarks from the literature. Emir Demirovic, Peter J. Stuckey, Tias Guns, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jeffrey Chan |
AAAI | 6 |
| 2020 | A Data-Driven Frequency Scaling Approach for Deadline-aware Energy Efficient Scheduling on Graphics Processing Units (GPUs)abstractModern computing paradigms, such as cloud computing, are increasingly adopting GPUs to boost their computing capabilities primarily due to the heterogeneous nature of AI/ML/deep learning workloads. However, the energy consumption of GPUs is a critical problem. Dynamic Voltage Frequency Scaling (DVFS) is a widely used technique to reduce the dynamic power of GPUs. Yet, configuring the optimal clock frequency for essential performance requirements is a non-trivial task due to the complex nonlinear relationship between the application's runtime performance characteristics, energy, and execution time. It becomes more challenging when different applications behave distinctively with similar clock settings. Simple analytical solutions and standard GPU frequency scaling heuristics fail to capture these intricacies and scale the frequencies appropriately. In this regard, we propose a data-driven frequency scaling technique by predicting the power and execution time of a given application over different clock settings. We collect the data from application profiling and train the models to predict the outcome accurately. The proposed solution is generic and can be easily extended to different kinds of workloads and GPU architectures. Furthermore, using this frequency scaling by prediction models, we present a deadline-aware application scheduling algorithm to reduce energy consumption while simultaneously meeting their deadlines. We conduct real extensive experiments on NVIDIA GPUs using several benchmark applications. The experiment results have shown that our prediction models have high accuracy with the average RMSE values of 0.38 and 0.05 for energy and time prediction, respectively. Also, the scheduling algorithm consumes 15.07% less energy as compared to the baseline policies. Shashikant Ilager, Rajeev Muralidhar, Kotagiri Ramamohanarao, Rajkumar Buyya |
CCGRID | 3 |
| 2020 | Short-Term and Long-Term Context Aggregation Network for Video Inpainting
Ang Li 0008, Shanshan Zhao 0001, Xingjun Ma, Mingming Gong, Jianzhong Qi 0001, Rui Zhang 0003, Dacheng Tao, Kotagiri Ramamohanarao |
ECCV (4) | 8 |
| 2020 | Learning with Bounded Instance and Label-dependent Label NoiseabstractInstance- and Label-dependent label Noise (ILN) widely exists in real-world datasets but has been rarely studied. In this paper, we focus on Bounded Instance- and Label-dependent label Noise (BILN), a particular case of ILN where the label noise rates—the probabilities that the true labels of examples flip into the corrupted ones—have upper bound less than $1$. Specifically, we introduce the concept of distilled examples, i.e. examples whose labels are identical with the labels assigned for them by the Bayes optimal classifier, and prove that under certain conditions classifiers learnt on distilled examples will converge to the Bayes optimal classifier. Inspired by the idea of learning with distilled examples, we then propose a learning algorithm with theoretical guarantees for its robustness to BILN. At last, empirical evaluations on both synthetic and real-world datasets show effectiveness of our algorithm in learning with BILN. Tongliang Liu, Kotagiri Ramamohanarao, Dacheng Tao |
ICML | 3 |
| 2020 | TimeSAN: A Time-Modulated Self-Attentive Network for Next Point-of-Interest RecommendationabstractNext Point-of-Interest (POI) recommendation aims to rank a list of POIs by their attractiveness to users based on the users' historical records of POI visits. This task is challenging, because user preferences may be influenced by various contextual factors. In this paper, we consider the temporal contextual factor, i.e., the time of users' POI visits. Previous attempts for modelling the impact of temporal contexts can be categorized into two groups: factorization based methods and recurrent neural network based methods. The first group adds a time dimension to their latent recommendation spaces, which may suffer from the data sparsity problem due to the additional dimension. The second group uses time-aware contextual gates to update the hidden and cell states in RNNs, which may have limited capability in capturing long-range temporal dynamics. In this paper, we propose a time -modulated s elf-attentive network (TimeSAN) for next POI recommendation. This model learns the relevance between a user's next POI visit and her historical visits via the self-attention mechanism, where the relevance is modulated by the temporal contextual influence. The learned time-aware relevance is further fused with users' long-term interests to provide final recommendations. We conduct extensive experiments on real-world datasets. The results confirm that TimeSAN outperforms previous methods consistently and significantly in recommendation accuracy, while attaining a high model training efficiency. Estrid He, Jianzhong Qi 0001, Kotagiri Ramamohanarao |
IJCNN | 3 |
| 2020 | Segmented Pairwise Distance for Time Series with Large DiscontinuitiesabstractTime series with large discontinuities are common in many scenarios. However, existing distance-based algorithms (e.g., DTW and its derivative algorithms) may perform poorly in measuring distances between these time series pairs. In this paper, we propose the segmented pairwise distance (SPD) algorithm to measure distances between time series with large discontinuities. SPD is orthogonal to distance-based algorithms and can be embedded in them. We validate advantages of SPD-embedded algorithms over corresponding distance-based ones on both open datasets and a proprietary dataset of surgical time series (of surgeons performing a temporal bone surgery in a virtual reality surgery simulator). Experimental results demonstrate that SPD-embedded algorithms outperform corresponding distance-based ones in distance measurement between time series with large discontinuities, measured by the Silhouette index (SI). Jiabo He, Sarah M. Erfani, Sudanthi N. R. Wijewickrema, Stephen J. O'Leary, Kotagiri Ramamohanarao |
IJCNN | 5 |
| 2020 | Learning Non-Unique Segmentation with Reward-Penalty Dice LossabstractSemantic segmentation is one of the key problems in the field of computer vision, as it enables computer image understanding. However, most research and applications of semantic segmentation focus on addressing unique segmentation problems, where there is only one gold standard segmentation result for every input image. This may not be true in some problems, e.g., medical applications. We may have non-unique segmentation annotations as different surgeons may perform successful surgeries for the same patient in slightly different ways. To comprehensively learn non-unique segmentation tasks, we propose the reward-penalty Dice loss (RPDL) function as the optimization objective for deep convolutional neural networks (DCNN). RPDL is capable of helping DCNN learn non-unique segmentation by enhancing common regions and penalizing outside ones. Experimental results show that RPDL improves the performance of DCNN models by up to 18.4% compared with other loss functions on our collected surgical dataset. Jiabo He, Sarah M. Erfani, Sudanthi N. R. Wijewickrema, Stephen J. O'Leary, Kotagiri Ramamohanarao |
IJCNN | 5 |
| 2020 | Instance-Based Ensemble Selection Using Deep Reinforcement LearningabstractEnsemble selection is a very active research topic in machine learning area. It aims to achieve a better performance by selecting a proper subset of the original ensemble, which is essentially a searching problem in large combinatorial spaces. In this paper, we propose an instance-based reinforcement learning (IBRL) model, that selects distinct subsets for different instances. Specifically, we use deep Q-network to approximate the optimal policy. Rather than considering the overall performance of each classifier, the network learns from the feedback of classifiers on individual instance, so that it generates non-static subsets for different instances. Experiments are conducted to compare our model against state-of-the-art approaches for both selection and combination. The proposed method generates promising results and it shows exceptional advantage in large scale distributed environment. Due to the environment-free characteristic of reinforcement learning, our model is adaptable to various real world tasks with minimal changes. Zhengshang Liu, Kotagiri Ramamohanarao |
IJCNN | 2 |
| 2020 | Improving Single and Multi-View Blockmodelling by Algebraic SimplificationabstractBlockmodelling is an important technique in social network analysis for discovering the latent structures and groupings in graphs. State-of-the-art approaches approximate the graph using matrix factorisation, which can discover both the latent graph structures and vertex groupings. However, factorisation is a one-way approximation, in that it only approximates the graph with a lossy model that removes the background noise. Traditional Blockmodelling methods rely on an alternating 2-step optimization that involves iteratively updating the matrix representing membership while fixing the matrix representing the graph's underlying structure, and then updating the structure matrix while keeping the membership matrix fixed. We propose a single step optimization method, which uses algebraic simplifi-cation to directly update the lower dimensional, latent structure representation. This helps improve both the convergence and accuracy of blockmodelling. We also show that this approach can solve multi-view blockmodelling problems, involving multiple graphs over the same vertices. We use real datasets to show that our approach has much higher accuracy and comparable running times to competing approaches. Rishabh Ramteke, Peter J. Stuckey, Jeffrey Chan, Kotagiri Ramamohanarao, James Bailey 0001, Christopher Leckie, Emir Demirovic |
IJCNN | 4 |
| 2020 | Heterogeneous Task Co-location in Containerized Cloud Computing EnvironmentsabstractAlthough cloud computing became a mainstream industrial computing paradigm, low resource utilization remains a common problem that most warehouse-scale datacenters suffer from. This leads to a significant waste of hardware resources, infrastructure investment, and energy consumption. As the diversity in application workloads grows into an essential characteristic in modern datacenters, task co-location of different workloads to the same compute cluster has gained immense popularity as a heuristic solution for resource utilization optimization. Although the existing co-location methodologies manage to improve resource efficiency to a certain degree, application QoS is usually sacrificed as a trade-off when dealing with resource interference between different applications. This paper proposes a containerized task co-location (CTCL) scheduler to improve resource utilization and minimize task eviction rate. Our CTCL scheduler (1) applies an elastic task co-location strategy to improve resource utilization; and (2) supports a dynamic task rescheduling mechanism to prevent severe QoS degradation from frequent task evictions. We evaluate our approach in terms of resource efficiency and rescheduling cost through the ContainerCloudSim simulator. Our experiments with the Alibaba 2018 workload traces demonstrate that CTCL could improve overall resource efficiency and reduce rescheduling rate by 38% and 99% respectively. Zhiheng Zhong, Jiabo He, Maria Rodriguez Read, Sarah M. Erfani, Kotagiri Ramamohanarao, Rajkumar Buyya |
ISORC | 5 |
| 2020 | Modeling cloud business customers' utility functions
Caesar Wu, Rajkumar Buyya, Kotagiri Ramamohanarao |
Future Gener. Comput. Syst. | 3 |
| 2020 | TGM: A Generative Mechanism for Publishing Trajectories With Differential PrivacyabstractWe describe a new generative algorithm called trajectory generative mechanism (TGM) for publishing trajectory datasets with ε-differential privacy guarantee, which achieves substantially higher computational efficiency and utility (practical) than the state-of-the-art algorithms. Our algorithm first encodes (models) the data as a graphical generative model and accurately captures the statistics of moving object trajectories. Using this model, TGM then privately generates synthetic trajectories such that the noise is optimally added to capture the movement direction of an object. Our algorithm preserves both the spatial and temporal information of trajectories in the generated dataset, requires less memory and computation than the competing approaches, and preserves the properties of real trajectory data in terms of traveled distance and stay location. We demonstrate the performance of TGM on both real and simulated datasets with a wide range of settings. Our experimental results show that TGM achieves high utility and efficiency by using the properties of the data. Soheila Ghane, Lars Kulik, Kotagiri Ramamohanarao |
IEEE Internet Things J. | 3 |
| 2020 | Profit-aware application placement for integrated Fog-Cloud computing environments
Md. Redowan Mahmud, Satish Narayana Srirama, Kotagiri Ramamohanarao, Rajkumar Buyya |
J. Parallel Distributed Comput. | 3 |
| 2020 | Exploiting patterns to explain individual predictions
Yunzhe Jia, James Bailey 0001, Kotagiri Ramamohanarao, Christopher Leckie, Xingjun Ma |
Knowl. Inf. Syst. | 3 |
| 2020 | Unsupervised online change point detection in high-dimensional time series
Masoomeh Zameni, Amin Sadri, Zahra Ghafoori, Masud Moshtaghi, Flora D. Salim, Christopher Leckie, Kotagiri Ramamohanarao |
Knowl. Inf. Syst. | 7 |
| 2020 | Context-Aware Placement of Industry 4.0 Applications in Fog Computing EnvironmentsabstractThe fourth industrial revolution, widely known as Industry 4.0, is realizable through widespread deployment of Internet of Things (IoT) devices across the industrial ambiance. Due to communication latency and geographical distribution, Cloud-centric IoT models often fail to satisfy the Quality of Service requirements of different IoT applications assisting Industry 4.0 in real time. Therefore, Fog computing focuses on harnessing edge resources to place and execute these applications in the proximity of data sources. Since most of the Fog nodes are heterogeneous, distributed, and resource-constrained, it is challenging to place Industry 4.0-oriented applications (I4OAs) over them ensuring time-optimized service delivery. Diversified data sensing frequency of different industrial IoT devices and their data size further intensify the application placement problem. To address this issue, in this article we propose a context-aware application placement policy for Fog environments. Our policy coordinates the IoT device-level contexts with the capacity of Fog nodes and minimizes the service delivery time of various I4OAs such as image processing and robot navigation applications. It also ensures that the streams of input data flowing toward the placed applications neither congest the network nor increase the computing overhead of host Fog nodes significantly. Performance of the proposed policy is evaluated in both real-world and simulated Fog environments and compared with the existing placement policies. The experiment results show that our policy offers overall 16% improvement in service latency, network relaxation, and computing overhead management compared to other placement policies. Md. Redowan Mahmud, Adel Nadjaran Toosi, Kotagiri Ramamohanarao, Rajkumar Buyya |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | A Scalable Multi-Data Sources Based Recursive Approximation Approach for Fast Error Recovery in Big Sensing Data on CloudabstractBig sensing data is commonly encountered from various surveillance or sensing systems. Sampling and transferring errors are commonly encountered during each stage of sensing data processing. How to recover from these errors with accuracy and efficiency is quite challenging because of high sensing data volume and unrepeatable wireless communication environment. While Cloud provides a promising platform for processing big sensing data, however scalable and accurate error recovery solutions are still need. In this paper, we propose a novel approach to achieve fast error recovery in a scalable manner on cloud. This approach is based on the prediction of a recovery replacement data by making multiple data sources based approximation. The approximation process will use coverage information carried by data units to limit the algorithm in a small cluster of sensing data instead of a whole data spectrum. Specifically, in each sensing data cluster, a Euclidean distance based approximation is proposed to calculate a time series prediction. With the calculated time series, a detected error can be recovered with a predicted data value. Through the experiment with real world meteorological data sets on cloud, we demonstrate that the proposed error recovery approach can achieve high accuracy in data approximation to replace the original data error. At the same time, with MapReduce based implementation for scalability, the experimental results also show significant efficiency on time saving. Chi Yang, Xianghua Xu, Kotagiri Ramamohanarao, Jinjun Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Real-Time Recurrent Tactile Recognition: Momentum Batch-Sequential Echo State NetworksabstractTactile recognition aims at identifying target objects according to tactile sensory readings. Tactile data have two salient properties: 1) sequentially real-time and 2) temporally correlated, which essentially calls for a real-time (i.e., online fixed-budget) and recurrent recognition procedure. Based on an efficient and robust spatio-temporal feature representation for tactile sequences, we handle the problem of real-time recurrent tactile recognition by proposing a bounded online-sequential learning framework, and incorporates the strength of batch-regularization bootstrapping, bounded recursive reservoir, and momentum-based estimation. Experimental evaluations show that it outperforms the state-of-the-art methods by a large margin on test accuracy; and its training performance is superior to most compared models from aspects of average online training error, computational complexity, and storage efficiency. Le-le Cao, Fuchun Sun 0001, Kotagiri Ramamohanarao, Wenbing Huang 0001, Weihao Cheng 0001, Xiaolong Liu 0010 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2019 | An Investigation into Prediction + Optimisation for the Knapsack Problem
Emir Demirovic, Peter J. Stuckey, James Bailey 0001, Jeffrey Chan, Christopher Leckie, Kotagiri Ramamohanarao, Tias Guns |
CPAIOR | 6 |
| 2019 | Streaming Route Assignment for Connected Autonomous Vehicles (Systems Paper)abstractIn the coming era of connected autonomous vehicles, data-driven traffic optimization will reach its full potential. By collecting highly detailed real-time traffic data from sensors and vehicles, a traffic management system will have the full view of the entire road network, allowing it to plan traffic in a virtual world that replicates the real road network. This will bring significant innovations to transport-domain applications. We prototype a traffic management system that can perform traffic optimization with connected autonomous vehicles. We propose two route assignment algorithms that aim to reduce traffic delays by reducing intersecting routes. The proposed algorithms and two state-of-the-art route assignment algorithms are implemented in the prototype system. We evaluate the algorithms with both synthetic and real road networks. The experimental results show that the proposed algorithms outperform competitors in terms of the travel times of the routes. Sadegh Motallebi, Hairuo Xie, Egemen Tanin, Jianzhong Qi 0001, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 5 |
| 2019 | A Joint Context-Aware Embedding for Trip RecommendationsabstractTrip recommendation is an important location-based service that helps relieve users from the time and efforts for trip planning. It aims to recommend a sequence of places of interest (POIs) for a user to visit that maximizes the user's satisfaction. When adding a POI to a recommended trip, it is essential to understand the context of the recommendation, including the POI popularity, other POIs co-occurring in the trip, and the preferences of the user. These contextual factors are learned separately in existing studies, while in reality, they jointly impact on a user's choice of POI visits. In this study, we propose a POI embedding model to jointly learn the impact of these contextual factors. We call the learned POI embedding a context-aware POI embedding. To showcase the effectiveness of this embedding, we apply it to generate trip recommendations given a user and a time budget. We propose two trip recommendation algorithms based on our context-aware POI embedding. The first algorithm finds the exact optimal trip by transforming and solving the trip recommendation problem as an integer linear programming problem. To achieve a high computation efficiency, the second algorithm finds a heuristically optimal trip based on adaptive large neighborhood search. We perform extensive experiments on real datasets. The results show that our proposed algorithms consistently outperform state-of-the-art algorithms in trip recommendation quality, with an advantage of up to 43% in F_1-score. Estrid He, Jianzhong Qi 0001, Kotagiri Ramamohanarao |
ICDE | 3 |
| 2019 | 2ED: An Efficient Entity Extraction Algorithm Using Two-Level Edit-DistanceabstractEntity extraction is fundamental to many text mining tasks such as organisation name recognition. A popular approach to entity extraction is based on string matching against a dictionary of known entities. For approximate entity extraction from free text, considering solely character-based or solely token-based similarity cannot simultaneously deal with minor name variations at token-level and typos at character-level. Moreover, the tolerance of mismatch in character-level may be different from that in token-level, and the tolerance thresholds of the two levels should be able to be customised individually. In this paper, we propose an efficient character-level and token-level edit-distance based algorithm called FuzzyED. To improve the efficiency of FuzzyED, we develop various novel techniques including (i) a spanning-based candidate sub-string producing technique, (ii) a lower bound dissimilarity to determine the boundaries of candidate sub-strings, (iii) a core token based technique that makes use of the importance of tokens to reduce the number of unpromising candidate sub-strings, and (iv) a shrinking technique to reuse computation. Empirical results on real world datasets show that FuzzyED can efficiently extract entities and produce a high F1score in the range of [0.91, 0.97]. Zeyi Wen, Dong Deng 0001, Rui Zhang 0003, Kotagiri Ramamohanarao |
ICDE | 4 |
| 2019 | Predict+Optimise with Ranking Objectives: Exhaustively Learning Linear FunctionsabstractWe study the predict+optimise problem, where machine learning and combinatorial optimisation must interact to achieve a common goal. These problems are important when optimisation needs to be performed on input parameters that are not fully observed but must instead be estimated using machine learning. Our contributions are two-fold: 1) we provide theoretical insight into the properties and computational complexity of predict+optimise problems in general, and 2) develop a novel framework that, in contrast to related work, guarantees to compute the optimal parameters for a linear learning function given any ranking optimisation problem. We illustrate the applicability of our framework for the particular case of the unit-weighted knapsack predict+optimise problem and evaluate on benchmarks from the literature. Emir Demirovic, Peter J. Stuckey, James Bailey 0001, Jeffrey Chan, Christopher Leckie, Kotagiri Ramamohanarao, Tias Guns |
IJCAI | 6 |
| 2019 | Generative Image Inpainting with Submanifold AlignmentabstractImage inpainting aims at restoring missing regions of corrupted images, which has many applications such as image restoration and object removal. However, current GAN-based generative inpainting models do not explicitly exploit the structural or textural consistency between restored contents and their surrounding contexts. To address this limitation, we propose to enforce the alignment (or closeness) between the local data submanifolds (subspaces) around restored images and those around the original (uncorrupted) images during the learning process of GAN-based inpainting models. We exploit Local Intrinsic Dimensionality (LID) to measure, in deep feature space, the alignment between data submanifolds learned by a GAN model and those of the original data, from a perspective of both images (denoted as iLID) and local patches (denoted as pLID) of images. We then apply iLID and pLID as regularizations for GAN-based inpainting models to encourage two different levels of submanifold alignments: 1) an image-level alignment to improve structural consistency, and 2) a patch-level alignment to improve textural details. Experimental results on four benchmark datasets show that our proposed model can generate more accurate results than state-of-the-art models. Ang Li 0008, Jianzhong Qi 0001, Rui Zhang 0003, Xingjun Ma, Kotagiri Ramamohanarao |
IJCAI | 5 |
| 2019 | Boosted GAN with Semantically Interpretable Information for Image InpaintingabstractImage inpainting aims at restoring missing regions of corrupted images, which has many applications such as image restoration and object removal. However, current GAN-based inpainting models fail to explicitly consider the semantic consistency between restored images and original images. For example, given a male image with image region of one eye missing, current models may restore it with a female eye. This is due to the ambiguity of GAN-based inpainting models: these models can generate many possible restorations given a missing region. To address this limitation, our key insight is that semantically interpretable information (such as attribute and segmentation information) of input images (with missing regions) can provide essential guidance for the inpainting process. Based on this insight, we propose a boosted GAN with semantically interpretable information for image inpainting that consists of an inpainting network and a discriminative network. The inpainting network utilizes two auxiliary pretrained networks to discover the attribute and segmentation information of input images and incorporates them into the inpainting process to provide explicit semantic-level guidance. The discriminative network adopts a multi-level design that can enforce regularizations not only on overall realness but also on attribute and segmentation consistency with the original images. Experimental results show that our proposed model can preserve consistency on both attribute and segmentation level, and significantly outperforms the state-of-the-art models. Ang Li 0008, Jianzhong Qi 0001, Rui Zhang 0003, Kotagiri Ramamohanarao |
IJCNN | 4 |
| 2019 | Improving the Quality of Explanations with Local Embedding PerturbationsabstractClassifier explanations have been identified as a crucial component of knowledge discovery. Local explanations evaluate the behavior of a classifier in the vicinity of a given instance. A key step in this approach is to generate synthetic neighbors of the given instance. This neighbor generation process is challenging and it has considerable impact on the quality of explanations. To assess quality of generated neighborhoods, we propose a local intrinsic dimensionality (LID) based locality constraint. Based on this, we then propose a new neighborhood generation method. Our method first fits a local embedding/subspace around a given instance using the LID of the test instance as the target dimensionality, then generates neighbors in the local embedding and projects them back to the original space. Experimental results show that our method generates more realistic neighborhoods and consequently better explanations. It can be used in combination with existing local explanation algorithms. Yunzhe Jia, James Bailey 0001, Kotagiri Ramamohanarao, Christopher Leckie, Michael E. Houle |
KDD | 3 |
| 2019 | Query-Aware Bayesian Committee Machine for Scalable Gaussian Process RegressionabstractThe Gaussian process (GP) model is a powerful tool for regression problems. However, the high computational costs of the GP model has constrained its applications over large-scale data sets. To overcome this limitation, aggregation models employ distributed GP submodels (experts) for parallel training and predicting, and then merge the predictions of all submodels to produce an approximated result. The state-of-the-art aggregation models are based on Bayesian committee machines, where a prior is assumed at the start and then updated by each submodel. In this paper, we investigate the impact of the prior on the accuracy of aggregations. We propose a query-aware Bayesian committee machine (QBCM). The QBCM model partitions the testing data (i.e., queries) into subsets, and incorporates a query-aware prior when merging the predictions of submodels. This model improves the prediction accuracy, while retaining the advantages of aggregation models, i.e., closed-form inference and parallelizability. We conduct both theoretical analysis and empirical experiments on real data. The results confirm the effectiveness and efficiency of the proposed model QBCM. Estrid He, Jianzhong Qi 0001, Kotagiri Ramamohanarao |
SDM | 3 |
| 2019 | ETAS: Energy and thermal-aware dynamic virtual machine consolidation in cloud data center with proactive hotspot mitigationabstractSummary Data centers consume an enormous amount of energy to meet the ever‐increasing demand for cloud resources. Computing and Cooling are the two main subsystems that largely contribute to energy consumption in a data center. Dynamic Virtual Machine (VM) consolidation is a widely adopted technique to reduce the energy consumption of computing systems. However, aggressive consolidation leads to the creation of local hotspots that has adverse effects on energy consumption and reliability of the system. These issues can be addressed through efficient and thermal‐aware consolidation methods. We propose an Energy and Thermal‐Aware Scheduling (ETAS) algorithm that dynamically consolidates VMs to minimize the overall energy consumption while proactively preventing hotspots. ETAS is designed to address the trade‐off between time and the cost savings and it can be tuned based on the requirement. We perform extensive experiments by using the real‐world traces with precise power and thermal models. The experimental results and empirical studies demonstrate that ETAS outperforms other state‐of‐the‐art algorithms by reducing overall energy without any hotspot creation. Shashikant Ilager, Kotagiri Ramamohanarao, Rajkumar Buyya |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Performance anomaly detection using isolation-trees in heterogeneous workloads of web applications in computing cloudsabstractSummary Cloud computing is a model for on‐demand access to shared resources based on the pay‐per‐use policy. In order to efficiently manage the resources, a continuous analysis of the operational state of the system is required to be able to detect the performance degradations and malfunctioned resources as soon as possible. Every change in the workload, hardware condition, or software code can change the state of the system from normal to abnormal, which causes the performance and quality of service degradations. These changes or anomalies vary from a simple gradual increase in the load to flash crowds, hardware faults, software bugs, etc. In this paper, we propose Isolation‐Forest based anomaly detection (IFAD) framework based on the unsupervised Isolation technique for anomaly detection in a multi‐attribute space of performance indicators for web‐based applications. Unsupervised nature of the algorithm and its fast execution make this algorithm most suitable for the environments with dynamic nature where the patterns of data change frequently. The experiment results demonstrate that IFAD can achieve good detection accuracy especially in terms of precision for multiple types of the anomaly. Moreover, we show the importance of validating the accuracy of anomaly detection algorithms with regard to both Area Under the Curve (AUC) and Precision‐Recall AUC (PRAUC) in an extensive set of comparisons including multiple unsupervised algorithms. The demonstration of the effectiveness of each algorithm shown by PRAUC results indicates the importance of PRAUC in selecting suitable anomaly detection algorithm, which is largely ignored in the literature. Sara Kardani-Moghaddam, Rajkumar Buyya, Kotagiri Ramamohanarao |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | Value-based cloud price modeling for segmented business to business market
Caesar Wu, Rajkumar Buyya, Kotagiri Ramamohanarao |
Future Gener. Comput. Syst. | 3 |
| 2019 | Differential privacy for renewable energy resources based smart metering
Muneeb Ul Hassan 0001, Mubashir Husain Rehmani, Kotagiri Ramamohanarao, Jiekui Zhang, Jinjun Chen |
J. Parallel Distributed Comput. | 3 |
| 2019 | ACAS: An anomaly-based cause aware auto-scaling framework for clouds
Sara Kardani-Moghaddam, Rajkumar Buyya, Kotagiri Ramamohanarao |
J. Parallel Distributed Comput. | 3 |
| 2019 | Quality of Experience (QoE)-aware placement of applications in Fog computing environments
Md. Redowan Mahmud, Satish Narayana Srirama, Kotagiri Ramamohanarao, Rajkumar Buyya |
J. Parallel Distributed Comput. | 3 |
| 2019 | Holistic resource management for sustainable and reliable cloud computing: An innovative solution to global challenge
Sukhpal Singh, Peter Garraghan, Vlado Stankovski, Giuliano Casale, Ruppa K. Thulasiram, Soumya K. Ghosh 0001, Kotagiri Ramamohanarao, Rajkumar Buyya |
J. Syst. Softw. | 7 |
| 2019 | ROMIR: Robust Multi-View Image Re-RankingabstractIn multi-view re-ranking, multiple heterogeneous visual features are usually projected onto a low-dimensional subspace, and thus the resulting latent representation can be used for the subsequent similarity-based ranking. Albeit effective, this standard mechanism underplays the intrinsic structure underlying the latent subspace and does not take into account the substantial noise in the original spaces. In this paper, we propose a robust multi-view image re-ranking strategy. Due to the dramatic variability in image visual appearance, it is necessary to uncover the shared components underlying those query-related instances that are visually unlike for improving the re-ranking accuracy. Consequently, it is reasonable to assume the latent subspace enjoys the low-rank property and thus the subspace recovery can be achieved via the low-rank modeling accordingly. In addition, since the real-world data are usually partially contaminated, we employ `2;1-norm based sparsity constraint to appropriately model the sample-specific mapping noise for enhancing the model robustness. In order to produce discriminative representations, we encode a similarity preserving term in our multi-view embedding framework. As a result, the sample separability is maximally maintained in the latent subspace with sufficient discriminative power. The extensive evaluations on public landmark benchmarks demonstrate the efficacy and superiority of the proposed method. Jun Li 0033, Chang Xu 0002, Wankou Yang, Changyin Sun 0001, Kotagiri Ramamohanarao, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2019 | Latency-Aware Application Module Management for Fog Computing EnvironmentsabstractThe fog computing paradigm has drawn significant research interest as it focuses on bringing cloud-based services closer to Internet of Things (IoT) users in an efficient and timely manner. Most of the physical devices in the fog computing environment, commonly named fog nodes, are geographically distributed, resource constrained, and heterogeneous. To fully leverage the capabilities of the fog nodes, large-scale applications that are decomposed into interdependent Application Modules can be deployed in an orderly way over the nodes based on their latency sensitivity. In this article, we propose a latency-aware Application Module management policy for the fog environment that meets the diverse service delivery latency and amount of data signals to be processed in per unit of time for different applications. The policy aims to ensure applications’ Quality of Service (QoS) in satisfying service delivery deadlines and to optimize resource usage in the fog environment. We model and evaluate our proposed policy in an iFogSim-simulated fog environment. Results of the simulation studies demonstrate significant improvement in performance over alternative latency-aware strategies. Md. Redowan Mahmud, Kotagiri Ramamohanarao, Rajkumar Buyya |
ACM Trans. Internet Techn. | 2 |
| 2019 | Exploiting GPUs for Efficient Gradient Boosting Decision Tree TrainingabstractIn this paper, we present a novel parallel implementation for training Gradient Boosting Decision Trees (GBDTs) on Graphics Processing Units (GPUs). Thanks to the excellent results on classification/regression and the open sourced libraries such as XGBoost, GBDTs have become very popular in recent years and won many awards in machine learning and data mining competitions. Although GPUs have demonstrated their success in accelerating many machine learning applications, it is challenging to develop an efficient GPU-based GBDT algorithm. The key challenges include irregular memory accesses, many sorting operations with small inputs and varying data parallel granularities in tree construction. To tackle these challenges on GPUs, we propose various novel techniques including (i) Run-length Encoding compression and thread/block workload dynamic allocation, (ii) data partitioning based on stable sort, and fast and memory efficient attribute ID lookup in node splitting, (iii) finding approximate split points using two-stage histogram building, (iv) building histograms with the aware of sparsity and exploiting histogram subtraction to reduce histogram building workload, (v) reusing intermediate training results for efficient gradient computation, and (vi) exploiting multiple GPUs to handle larger data sets efficiently. Our experimental results show that our algorithm named ThunderGBM can be 10x times faster than the state-of-the-art libraries (i.e., XGBoost, LightGBM and CatBoost) running on a relatively high-end workstation of 20 CPU cores. In comparison with the libraries on GPUs, ThunderGBM can handle higher dimensional problems which the libraries become extremely slow or simply fail. For the data sets the existing libraries on GPUs can handle, ThunderGBM achieves up to 10 times speedup on the same hardware, which demonstrates the significance of our GPU optimizations. Moreover, the models trained by ThunderGBM are identical to those trained by XGBoost, and have similar quality as those trained by LightGBM and CatBoost. Zeyi Wen, Jiashuai Shi, Bingsheng He, Jian Chen 0011, Kotagiri Ramamohanarao, Qinbin Li |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | Protecting privacy for distance and rank based group nearest neighbor queries
Tanzima Hashem, Lars Kulik, Kotagiri Ramamohanarao, Rui Zhang 0003, Subarna Chowdhury Soma |
World Wide Web | 3 |
| 2018 | Learning Datum-Wise Sampling Frequency for Energy-Efficient Human Activity RecognitionabstractContinuous Human Activity Recognition (HAR) is an important application of smart mobile/wearable systems for providing dynamic assistance to users. However, HAR in real-time requires continuous sampling of data using built-in sensors (e.g., accelerometer), which significantly increases the energy cost and shortens the operating span. Reducing sampling rate can save energy but causes low recognition accuracy. Therefore, choosing adaptive sampling frequency that balances accuracy and energy efficiency becomes a critical problem in HAR. In this paper, we formalize the problem as minimizing both classification error and energy cost by choosing dynamically appropriate sampling rates. We propose Datum-Wise Frequency Selection (DWFS) to solve the problem via a continuous state Markov Decision Process (MDP). A policy function is learned from the MDP, which selects the best frequency for sampling an incoming data entity by exploiting a datum related state of the system. We propose a method for alternative learning the parameters of an activity classification model and the MDP that improves both the accuracy and the energy efficiency. We evaluate DWFS with three real-world HAR datasets, and the results show that DWFS statistically outperforms the state-of-the-arts regarding a combined measurement of accuracy and energy efficiency. Weihao Cheng 0001, Sarah M. Erfani, Rui Zhang 0003, Kotagiri Ramamohanarao |
AAAI | 4 |
| 2018 | Activity-based ride-sharing in action (demo paper)abstractActivity-Based ride-sharing is a new paradigm which enhances the current model based on fixed origins and destinations, namely trip-based ride-sharing. In this new model, a user issues a ride-sharing request with his origin and the activity he wants to perform at any convenient destination. Then, the system computes the travel plans and users will be suggested the optimal destinations, which may be common to many users. In this way, the set of possible destinations for each user is expanded and further distance savings can be made as we have already shown in our previous work [1, 3]. In this paper, we show Activity-Based ride-sharing in action through our web-service-based framework, which is able to suggest routes and meeting points for many users in a city-scale scenario. Oscar Correa, Egemen Tanin, Lars Kulik, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 4 |
| 2018 | Studying transportation problems with the SMARTS simulator (demo paper)abstractMicroscopic traffic simulators play a major role to carry research on transportation problems. Microscopic traffic simulation is powerful because it enables efficient analysis of complex traffic problems to the highest level of detail. We developed Scalable Microscopic Adaptive Road Traffic Simulator (SMARTS) [14] that can perform large-scale simulations at a high speed by utilizing distributed computing resources. Previous results show that SMARTS can run 1.14 times faster than real time when simulating one million vehicles for the city of Melbourne on 30 distributed processors, while producing highly accurate simulation results. SMARTS' pluggable architecture allows it to be easily extended to simulate specific scenarios of interest to users. In this demonstration we show how SMARTS can be used to simulate an intersection design, the P-turn, in a major intersection of Melbourne. Our simulation shows the impact of the design on the traffic flow, confirming the justification for introduction of the particular intersection. The demo can be used as a template for future use of the simulator for other traffic problems. Hairuo Xie, Egemen Tanin, Shanika Karunasekera, Lars Kulik, Rui Zhang 0003, Jianzhong Qi 0001, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 7 |
| 2018 | Automatic Optical Coherence Tomography Imaging Analysis for Retinal Disease Screening Using Machine Learning Techniques
Kotagiri Ramamohanarao |
ICDM | 1 |
| 2018 | Cloud Computing Market Segmentation
Caesar Wu, Rajkumar Buyya, Kotagiri Ramamohanarao |
ICSOFT | 3 |
| 2018 | Predicting Complex Activities from Ongoing Multivariate Time SeriesabstractThe rapid development of sensor networks enables recognition of complex activities (CAs) using multivariate time series. However, CAs are usually performed over long periods of time, which causes slow recognition by models based on fully observed data. Therefore, predicting CAs at early stages becomes an important problem. In this paper, we propose Simultaneous Complex Activities Recognition and Action Sequence Discovering (SimRAD), an algorithm which predicts a CA over time by mining a sequence of multivariate actions from sensor data using a Deep Neural Network. SimRAD simultaneously learns two probabilistic models for inferring CAs and action sequences, where the estimations of the two models are conditionally dependent on each other. SimRAD continuously predicts the CA and the action sequence, thus the predictions are mutually updated until the end of the CA. We conduct evaluations on a real-world CA dataset consisting of a rich amount of sensor data, and the results show that SimRAD outperforms state-of-the-art methods by average 7.2% in prediction accuracy with high confidence. Weihao Cheng 0001, Sarah M. Erfani, Rui Zhang 0003, Kotagiri Ramamohanarao |
IJCAI | 4 |
| 2018 | Efficient Gradient Boosted Decision Tree Training on GPUsabstractIn this paper, we present a novel parallel implementation for training Gradient Boosting Decision Trees (GBDTs) on Graphics Processing Units (GPUs). Thanks to the wide use of the open sourced XGBoost library, GBDTs have become very popular in recent years and won many awards in machine learning and data mining competitions. Although GPUs have demonstrated their success in accelerating many machine learning applications, there are a series of key challenges of developing a GPU-based GBDT algorithm, including irregular memory accesses, many small sorting operations and varying data parallel granularities in tree construction. To tackle these challenges on GPUs, we propose various novel techniques (including Run-length Encoding compression and thread/block workload dynamic allocation, and reusing intermediate training results for efficient gradient computation). Our experimental results show that our algorithm named GPU-GBDT is often 10 to 20 times faster than the sequential version of XGBoost, and achieves 1.5 to 2 times speedup over a 40 threaded XGBoost running on a relatively high-end workstation of 20 CPU cores. Moreover, GPU-GBDT outperforms its CPU counterpart by 2 to 3 times in terms of performance-price ratio. Zeyi Wen, Bingsheng He, Kotagiri Ramamohanarao, Shengliang Lu, Jiashuai Shi |
IPDPS | 3 |
| 2018 | Fast Manifold Landmarking Using Locality-Sensitive Hashing
Zay Maung Maung Aye, Benjamin I. P. Rubinstein, Kotagiri Ramamohanarao |
PAKDD (3) | 3 |
| 2018 | A Joint Optimization Approach for Personalized Recommendation Diversification
Xiaojie Wang 0003, Jianzhong Qi 0001, Kotagiri Ramamohanarao, Yu Sun 0021, Bo Li 0026, Rui Zhang 0003 |
PAKDD (3) | 3 |
| 2018 | Semi-supervised Blockmodelling with Pairwise Guidance
Mohadeseh Ganji, Jeffrey Chan, Peter J. Stuckey, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Laurence Anthony F. Park |
ECML/PKDD (2) | 6 |
| 2018 | Urban Sensing for Anomalous Event Detection: - Distinguishing Between Legitimate Traffic Changes and Abnormal Traffic Variability
Masoomeh Zameni, Mengyi He, Masud Moshtaghi, Zahra Ghafoori, Christopher Leckie, James C. Bezdek, Kotagiri Ramamohanarao |
ECML/PKDD (3) | 7 |
| 2018 | Image Constrained Blockmodelling: A Constraint Programming ApproachabstractBlockmodelling is an important technique for detecting underlying patterns in graphs. However, existing blockmodelling algorithms do not provide the user with any explicit control to specify which patterns might be of interest. Furthermore, existing algorithms focus on finding standard community structures in graphs, and are likely to overlook informative but more complex patterns, such as hierarchical or ring blockmodel structures. In this paper, we propose a generic constraint programming framework for blockmodelling, which allows a user to specify and search for complex blockmodel patterns in graphs. Our proposed framework can be incorporated into existing iterative blockmodelling algorithms, operating as a hybrid optimization scheme that provides high flexibility and expressiveness. We demonstrate the power of our framework for discovering complex patterns, via experiments over a range of synthetic and real data sets. Mohadeseh Ganji, Jeffrey Chan, Peter J. Stuckey, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Ian Davidson |
SDM | 6 |
| 2018 | Publishing spatial histograms under differential privacyabstractStudying trajectories of individuals has received growing interest. The aggregated movement behaviour of people provides important insights about their habits, interests, and lifestyles. Understanding and utilizing trajectory data is a crucial part of many applications such as location based services, urban planning, and traffic monitoring systems. Spatial histograms and spatial range queries are key components in such applications to efficiently store and answer queries on trajectory data. A spatial histogram maintains the sequentiality of location points in a trajectory by a strong sequential dependency among histogram cells. This dependency is an essential property in answering spatial range queries. However, the trajectories of individuals are unique and even aggregating them in spatial histograms cannot completely ensure an individual's privacy. A key technique to ensure privacy for data publishing ϵ-differential privacy as it provides a strong guarantee on an individual's provided data. Our work is the first that guarantees ϵ-differential privacy for spatial histograms on trajectories, while ensuring the sequentiality of trajectory data, i.e., its consistency. Consistency is key for any database and our proposed mechanism, PriSH, synthesizes a spatial histogram and ensures the consistency of published histogram with respect to the strong dependency constraint. In extensive experiments on real and synthetic datasets, we show that (1) PriSH is highly scalable with the dataset size and granularity of the space decomposition, (2) the distribution of aggregate trajectory information in the synthesized histogram accurately preserves the distribution of original histogram, and (3) the output has high accuracy in answering arbitrary spatial range queries. Soheila Ghane, Lars Kulik, Kotagiri Ramamohanarao |
SSDBM | 3 |
| 2018 | Density Biased Sampling with Locality Sensitive Hashing for Outlier Detection
Xuyun Zhang, Mahsa Salehi, Christopher Leckie, Qiang He 0001, Rui Zhou 0001, Kotagiri Ramamohanarao |
WISE (2) | 7 |
| 2018 | Detecting performance anomalies in scientific workflows using hierarchical temporal memory
Maria Rodriguez Read, Kotagiri Ramamohanarao, Rajkumar Buyya |
Future Gener. Comput. Syst. | 2 |
| 2018 | Dealing with Inliers in Feature Vector DataabstractInliers (bridge points) between clusters degrade the ability of many algorithms to find clusters in numerical data. We present three new approaches to the detection and removal of inliers. Two approaches are based on Local Outlier Factor (LOF) scores. We also discuss using LOF scores for an isolation Nearest Neighbour Ensemble (iNNE) approach to inlier detection. The third approach uses MaxiMin (MM) sampling to remove both inliers and outliers. We compare the three approaches on a synthetic and two real-life datasets. The failure of single linkage clustering due to the existence of bridging points is used as a means for evaluating the relative effectiveness of the three methods. We also show how inliers can degrade the quality of images built by the improved Visual Assessment of Tendency (iVAT) algorithm, which provides a visual representation of potential single linkage clusters in the data. Zahra Ghafoori, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao, Marimuthu Palaniswami |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 5 |
| 2018 | Spectral-based fault localization using hyperbolic functionabstractSummary Debugging is crucial for producing reliable software. One of the effective bug localization techniques is spectral‐based fault localization. It tries to locate a buggy statement by applying an evaluation metric to program spectra and ranking program components on the basis of the score it computes. Here, we propose a restricted class of “hyperbolic” metrics, with a small number of numeric parameters. This class of functions is based on past theoretical and empirical results. We show that optimization methods such as genetic programming and simulated annealing can reliably discover effective metrics over a wide range of data sets of program spectra. We evaluate the performance for both real programs and model programs with single bugs, multiple bugs, “deterministic” bugs, and nondeterministic bugs and find that the proposed class of metrics performs as well as or better than the previous best‐performing metrics over a broad range of data. Neelofar, Lee Naish, Kotagiri Ramamohanarao |
Softw. Pract. Exp. | 3 |
| 2018 | Scalable and fast SVM regression using modern hardware
Zeyi Wen, Rui Zhang 0003, Kotagiri Ramamohanarao |
World Wide Web | 3 |
| 2017 | From Shared Subspaces to Shared Landmarks: A Robust Multi-Source Classification ApproachabstractTraining machine leaning algorithms on augmented data fromdifferent related sources is a challenging task. This problemarises in several applications, such as the Internet of Things(IoT), where data may be collected from devices with differentsettings. The learned model on such datasets can generalizepoorly due to distribution bias. In this paper we considerthe problem of classifying unseen datasets, given several labeledtraining samples drawn from similar distributions. Weexploit the intrinsic structure of samples in a latent subspaceand identify landmarks, a subset of training instances fromdifferent sources that should be similar. Incorporating subspacelearning and landmark selection enhances generalizationby alleviating the impact of noise and outliers, as well asimproving efficiency by reducing the size of the data. However,since addressing the two issues simultaneously resultsin an intractable problem, we relax the objective functionby leveraging the theory of nonlinear projection and solve atractable convex optimisation. Through comprehensive analysis,we show that our proposed approach outperforms stateof-the-art results on several benchmark datasets, while keepingthe computational complexity low. Sarah M. Erfani, Mahsa Baktash, Masud Moshtaghi, Vinh Nguyen 0003, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
AAAI | 7 |
| 2017 | Improving Efficiency of SVM k-Fold Cross-Validation by Alpha SeedingabstractThe k-fold cross-validation is commonly used to evaluate the effectiveness of SVMs with the selected hyper-parameters. It is known that the SVM k-fold cross-validation is expensive, since it requires training k SVMs. However, little work has explored reusing the h-th SVM for training the (h+1)-th SVM for improving the efficiency of k-fold cross-validation. In this paper, we propose three algorithms that reuse the h-th SVM for improving the efficiency of training the (h+1)-th SVM. Our key idea is to efficiently identify the support vectors and to accurately estimate their associated weights (also called alpha values) of the next SVM by using the previous SVM. Our experimental results show that our algorithms are several times faster than the k-fold cross-validation which does not make use of the previously trained SVM. Moreover, our algorithms produce the same results (hence same accuracy) as the k-fold cross-validation which does not make use of the previously trained SVM. Zeyi Wen, Bin Li 0073, Kotagiri Ramamohanarao, Jian Chen 0011, Rui Zhang 0003 |
AAAI | 3 |
| 2017 | Search Result Personalization in Twitter Using Neural Word Embeddings
Sameendra Samarawickrama, Shanika Karunasekera, Aaron Harwood, Kotagiri Ramamohanarao |
DaWaK | 4 |
| 2017 | Using a Traffic Simulator for Navigation ServiceabstractTraffic congestion is a serious problem that is only expected to get worse in the future. Statistics shows that half of traffic congestion is caused by temporary disruptions like accidents. These events have dramatic impact on road network availability and cause huge delays for commuters. Also, they are usually unexpected and hard to manage by traffic authorities. State-of-the-art navigation systems started to provide real-time information about traffic conditions to help users make better routing decisions. However, traffic in the road network changes rapidly and the advice calculated now may not be valid after few minutes. This is especially critical in the presence of traffic incidents, where the impact of the incident could cause traffic to propagate to nearby roads. Thus, it is important for navigation systems to consider the evolution and future impact of traffic events. In this work, we present a navigation system that uses faster than realtime simulations to predict the evolution of traffic events and help drivers proactively avoid congestion caused by events. The system can subscribe to real-time traffic information and forecast the traffic conditions using fast simulations. We evaluate our approach through extensive experiments to test the performance and accuracy of the simulator with real data obtained from TomTom Traffic API. Also, we test the quality of navigation advice in realistic settings and show that our solution is able to help drivers avoid congested areas in cases where even real-time update methods lead drivers to congested routes. Abdullah AlDwyish, Hairuo Xie, Egemen Tanin, Shanika Karunasekera, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 5 |
| 2017 | Ride-sharing is About Agreeing on a DestinationabstractRide-sharing is rapidly becoming an alternative form of transportation mainly due to its economic benefits. Existing research on ridesharing aims to optimally match trajectories between people with pre-selected destinations. In this paper, we show better ride-sharing arrangements are possible when users are presented with more destinations and agree on a common destination. Given a set of points of interest (POIs) and a set of users, our approach presents destination POIs and computes ride-sharing plans. Each arrangement for a subset of users that fit in a car can be presented as a minimum Steiner tree (MST) problem. An optimal solution of the overall problem minimizes the total length of all the MSTs. The problem is a version of the set cover problem and is NP-hard. We first develop a series of baseline methods which use a popular MST algorithm. Then, we propose our method which uses constraints on intermediary points where users can meet to share rides. These constraints reduce the time complexity significantly and our method is up to two orders of magnitude faster than the best baseline method. Since our algorithm finds the subsets of users and POIs for each arrangement, we define and solve a new type of MST problem as a first step. Our experiments show that our method can provide a fast and readily deployable solution for real world large city scenarios. A. K. M. Mustafizur Rahman Khan, Oscar Correa, Egemen Tanin, Lars Kulik, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 5 |
| 2017 | Exploiting Data Dependency to Mitigate Stragglers in Distributed Spatial SimulationabstractDistributed spatial simulations commonly employ Bulk Synchronous Parallel model (BSP) implementation. However, implementations using BSP are usually fraught with the straggler problem, where the delay of any worker slows down the entire system. Random stragglers commonly occur due to many reasons: imbalanced workload, operating system scheduling, or communication delays. The straggler problem is further exasperated with increasing parallelism. To reduce the straggler problem and preserve simplicity and scalability advantages of the BSP model, we propose a new parallel model, which we call Priority Asynchronous Parallel (PAP) model. PAP exploits data dependencies of parallel processes to be computed and synchronized based on data priority to the other workers. For further computational improvement, we develop a load balancing and partitioning method, called GridGraph that utilizes the spatial and connectivity properties of the simulation space to reduce the size of exchanged data in addition to balancing the workload among workers. The proposed schemes are implemented and evaluated in a microscopic traffic simulator. Running traffic simulation for Melbourne, Beijing, and New York cities on 80 workers, the simulation achieves a performance speedup of around 47.4% for Melbourne, 52.18% for Beijing, and 65.84% for New York, using PAP model combined with GridGraph partitioning compared to BSP model. Eman Bin Khunayn, Shanika Karunasekera, Hairuo Xie, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 4 |
| 2017 | From How to Where: Traffic Optimization in the Era of Automated VehiclesabstractA large number of self-driving cars will be on roads in the near future. They will change traffic significantly. Self-driving cars can infer and decide travel paths from passenger input. Passengers do not need to involve in route planning. This provides great opportunities for traffic management systems to collaborate and achieve more efficient traffic management. By knowing most source-destination pairs of the passengers, we envisage an increasingly integrated system that can optimize routes and traffic lights to minimize travel time. By optimally scheduling time of travel and traffic light switching timings, such systems can also provide simultaneously emergency corridors for high priority vehicles such as police cars, fire engines, and ambulances when required. Kotagiri Ramamohanarao, Jianzhong Qi 0001, Egemen Tanin, Sadegh Motallebi |
SIGSPATIAL/GIS | 1 |
| 2017 | Straggler Mitigation for Distributed Behavioral SimulationabstractRunning large-scale behavioral simulations requires high computational power, which can be acquired by distributing computation workload to multiple computing nodes (i.e., workers) that run in parallel. The implementations of such systems commonly follow the Bulk Synchronous Parallel (BSP) model. However, implementations using BSP usually suffer from the straggler problem, where the delay of any worker slows down the entire simulation. The problem usually occurs due to communication delays or imbalanced workload among workers. To mitigate the straggler problem, we propose a novel parallel computational model, called Priority Synchronous Parallel (PSP) model. PSP exploits data dependencies of parallel processes to determine high priority data to be computed and synchronized while computing the remaining data. PSP is implemented and evaluated using traffic simulations for three large cities. The proposed technique shows significant performance improvements over the BSP model. Eman Bin Khunayn, Shanika Karunasekera, Hairuo Xie, Kotagiri Ramamohanarao |
ICDCS | 4 |
| 2017 | LSHiForest: A Generic Framework for Fast Tree Isolation Based Ensemble Anomaly AnalysisabstractAnomaly or outlier detection is a major challenge in big data analytics because anomaly patterns provide valuable insights for decision-making in a wide range of applications. Recently proposed anomaly detection methods based on the tree isolation mechanism are very fast due to their logarithmic time complexity, making them capable of handling big data sets efficiently. However, the underlying similarity or distance measures in these methods have not been well understood. Contrary to the claims that these methods never rely on any distance measure, we find that they have close relationships with certain distance measures. This implies that the current use of this fast isolation mechanism is only limited to these distance measures and fails to generalise to other commonlyused measures. In this paper, we propose a generic framework named LSHiForest for fast tree isolation based ensemble anomaly analysis with the use of a Locality-Sensitive Hashing (LSH) forest. Being generic, the proposed framework can be instantiated with a diverse range of LSH families, and the fast isolation mechanism can be extended to any distance measures, data types and data spaces where an LSH family is defined. In particular, the instances of our framework with kernelised LSH families or learning based hashing schemes can detect complicated anomalies like local or surrounded anomalies. We also formally show that the existing tree isolation based detection methods are special cases of our framework with the corresponding distance measures. Extensive experiments on both synthetic and real-world benchmark data sets show that the framework can achieve both high time efficiency and anomaly detection quality. Xuyun Zhang, Wan-Chun Dou, Qiang He 0001, Rui Zhou 0001, Christopher Leckie, Kotagiri Ramamohanarao, Zoran A. Salcic |
ICDE | 6 |
| 2017 | Fix-Budget and Recurrent Data Mining for Online Haptic Perception
Le-le Cao, Fuchun Sun 0001, Xiaolong Liu 0010, Wenbing Huang 0001, Weihao Cheng 0001, Kotagiri Ramamohanarao |
ICONIP (5) | 6 |
| 2017 | The Hitchhiker's Guide to the Optimal Route PlanningabstractHitchhiking is the oldest ridesharing process without prior arrangements by the ride sharers. It usually involves uncertain waiting times and various combinations of lifts on roads. For this way of traveling, the problem of finding an optimal route is extremely important and has not been studied. We propose the concept of a hitchhiking graph to represent all possible decisions that a hitchhiker can consider on a road network. We develop an efficient pruning technique for a faster computation of the optimal route with the least expected journey time. The effectiveness of our methods is evaluated on road networks of selected countries. Oleksii Vedernikov, Lars Kulik, Kotagiri Ramamohanarao |
MDM | 3 |
| 2017 | Markov Dynamic Subsequence Ensemble for Energy-Efficient Activity RecognitionabstractUbiquitous mobile computing technology provides opportunities for accurate Activity Recognition (AR). Recently, ensemble models using multiple feature representations based on time series subsequences have demonstrated excellent performance on recognition accuracy. However, these models can significantly increase the energy overhead and shorten battery lifespans of the mobile devices. We formalize a dynamic subsequence selection problem that minimizes the computational cost while persevering a high recognition accuracy. To solve the problem, we propose Markov Dynamic Subsequence Ensemble (MDSE), an algorithm for the selection of the subsequences via a Markov Decision Process (MDP), where a policy is learned for choosing the best subsequence given the state of prediction. Regarding MDSE, we derive an upper bound of the expected ensemble size, so that the energy consumption caused by the computations of the proposed method is guaranteed. Extensive experiments are conducted on 6 real AR datasets to evaluate the effectiveness of MDSE. Compared to the state-of-the-art methods, MDSE reduces 70.8% computational cost which is 3.42 times more energy efficient, and achieves a comparably high accuracy. Weihao Cheng 0001, Sarah M. Erfani, Rui Zhang 0003, Kotagiri Ramamohanarao |
MobiQuitous | 4 |
| 2017 | From Ride-Sourcing to Ride-Sharing through Hot-SpotsabstractSmartphones have allowed us to make ad-hoc travel arrangements. Ride-sharing is emerging as one of the new types of transportation enabled by smartphone revolution. Ride-sharing aims to alleviate current environmental, social and economical issues many big cities are facing due to low vehicle occupancy rates. Although ride-sharing companies have millions of users around the world, some of them do not offer true ride-sharing but a similar service called ride-sourcing where private car owners provide for-hire rides. Ride-sharing uptake has not been wide due to lack of convenience and incentives. We propose an enhanced ride-sharing model through the inclusion of proper places to meet that we call hot-spots. Hot-spots are shown to increase the convenience by solving the round-trip ride-sharing problem. As we represent our enhanced model through graphs, we introduce a new graph problem that we call Constrained Variable Steiner Tree, which is NP-hard. An effective and readily deployable heuristic solution to this problem is presented which is up to two orders of magnitude faster than the state-of-the-art solution as combinatorial explosion is avoided by the usage of a novel monotonic nondecreasing function. Oscar Correa, Kotagiri Ramamohanarao, Egemen Tanin, Lars Kulik |
MobiQuitous | 2 |
| 2017 | Accurate Recognition of the Current Activity in the Presence of Multiple Activities
Weihao Cheng 0001, Sarah M. Erfani, Rui Zhang 0003, Kotagiri Ramamohanarao |
PAKDD (2) | 4 |
| 2017 | A Simulation Study of Emergency Vehicle Prioritization in Intelligent Transportation SystemsabstractEmergency vehicle prioritization is important to the efficiency of emergency services. To address certain challenges in emergency vehicle prioritization, we perform microscopic simulations of an intelligent transportation system, where emergency vehicles broadcast certain information about their routes to nearby vehicles and traffic lights. Our study shows that broadcasting the route information can help reduce the response time of emergency vehicles significantly. In certain case, travel time of emergency vehicles can be as low as 37.1% of that of non-priority vehicles. Hairuo Xie, Shanika Karunasekera, Lars Kulik, Egemen Tanin, Rui Zhang 0003, Kotagiri Ramamohanarao |
VTC Spring | 6 |
| 2017 | On the effectiveness of isolation-based anomaly detection in cloud data centersabstractSummary The high volume of monitoring information generated by large‐scale cloud infrastructures poses a challenge to the capacity of cloud providers in detecting anomalies in the infrastructure. Traditional anomaly detection methods are resource‐intensive and computationally complex for training and/or detection, what is undesirable in very dynamic and large‐scale environment such as clouds. Isolation‐based methods have the advantage of low complexity for training and detection and are optimized for detecting failures. In this work, we explore the feasibility of Isolation Forest, an isolation‐based anomaly detection method, to detect anomalies in large‐scale cloud data centers. We propose a method to code time‐series information as extra attributes that enable temporal anomaly detection and establish its feasibility to adapt to seasonality and trends in the time‐series and to be applied online and in real‐time. Rodrigo N. Calheiros, Kotagiri Ramamohanarao, Rajkumar Buyya, Christopher Leckie, Steve Versteeg |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Septic shock prediction for ICU patients via coupled HMM walking on sequential contrast patterns
Shameek Ghosh, Jinyan Li 0001, Longbing Cao, Kotagiri Ramamohanarao |
J. Biomed. Informatics | 4 |
| 2017 | KRNN: k Rare-class Nearest Neighbour classification
Xiuzhen Zhang 0001, Yuxuan Li 0001, Kotagiri Ramamohanarao, Lifang Wu, Zahir Tari, Mohamed Cheriet |
Pattern Recognit. | 3 |
| 2017 | Improving spectral-based fault localization using static analysisabstractSummary Debugging is crucial for producing reliable software. One of the effective bug localization techniques is spectral‐based fault localization (SBFL). It helps to locate a buggy statement by applying an evaluation metric to program spectra and ranking program components on the basis of the score it computes. SBFL is an example of a dynamic analysis – an analysis of computer program that is performed by executing it with sufficient number of test cases. Static analysis, on the other hand, is performed in a non‐runtime environment. We introduce a weighting technique by combining these two kinds of program analysis. Static analysis is performed to categorize program statements into different classes and giving them weights based on the likelihood of being buggy statement. Statements are finally ranked on the basis of the weights computed by statements' categorization (static analysis) and scores computed by SBFL metrics (dynamic analysis). We evaluate the performance of our technique on Siemens test suite and Flex (having seeded bugs seeded by expert developers), Sed (having mixture of real and seeded bugs), and Space (having real bugs). In our evaluation, proposed weighting technique improves the performance of a wide variety of fault localization metrics up to 20% on single bug datasets and up to 42% on multi‐bug datasets. Copyright © 2017 John Wiley & Sons, Ltd. Neelofar, Lee Naish, Kotagiri Ramamohanarao |
Softw. Pract. Exp. | 4 |
| 2017 | Optimal Pick up Point Selection for Effective Ride SharingabstractCar occupancy rates (travelers per vehicle) are currently very low in most developed countries, for example, on average between 1.15 and 1.25 in Australia. Enabling shared rides on short notice can be an effective solution to counter the problem of increasing traffic through the use of the untapped transportation capacity. Common inhibitors for the uptake of ride sharing services are privacy and safety concerns. We present an approach to ride sharing where the pick up/drop off locations for passengers are selected from a fixed set, which has the advantage of increased safety through video surveillance. We present a scheme that chooses optimally fixed locations of Pick up Points (PuPs) and aims to maximize the car occupancy rates while preserving user privacy and safety. Our method enhances privacy as the users do not need to provide their precise home/work locations. We have extended the well studied 1-coverage problem, i.e., to cover an area with the minimum number of circles of a given radius [1], to road networks. The challenges for road networks are the varying population densities of suburbs which requires circles of different radii. The aim is to ensure that every point of a city's area is covered by at least one PuP while minimizing the total number of PuPs. By ensuring that we have different circle radii for PuPs the anonymity of individuals is the same throughout. Using Voronoi diagrams we present a k-anonymity model that guarantees a minimum number of individuals covered by every PuP. Our problem is a multi objective problem where we aim to maximize coverage, k-anonymity and privacy provided by the system to its users while facilitating ride sharing. Through greedy randomized adaptive search procedure (GRASP) we find out the Pareto front of solutions and evaluate their impact on ride sharing. Preeti Goel, Lars Kulik, Kotagiri Ramamohanarao |
IEEE Trans. Big Data | 3 |
| 2017 | SMARTS: Scalable Microscopic Adaptive Road Traffic SimulatorabstractMicroscopic traffic simulators are important tools for studying transportation systems as they describe the evolution of traffic to the highest level of detail. A major challenge to microscopic simulators is the slow simulation speed due to the complexity of traffic models. We have developed the Scalable Microscopic Adaptive Road Traffic Simulator (SMARTS), a distributed microscopic traffic simulator that can utilize multiple independent processes in parallel. SMARTS can perform fast large-scale simulations. For example, when simulating 1 million vehicles in an area the size of Melbourne, the system runs 1.14 times faster than real time with 30 computing nodes and 0.2s simulation timestep. SMARTS supports various driver models and traffic rules, such as the car-following model and lane-changing model, which can be driver dependent. It can simulate multiple vehicle types, including bus and tram. The simulator is equipped with a wide range of features that help to customize, calibrate, and monitor simulations. Simulations are accurate and confirm with real traffic behaviours. For example, it achieves 79.1% accuracy in predicting traffic on a 10km freeway 90 minutes into the future. The simulator can be used for predictive traffic advisories as well as traffic management decisions as simulations complete well ahead of real time. SMARTS can be easily deployed to different operating systems as it is developed with the standard Java libraries. Kotagiri Ramamohanarao, Hairuo Xie, Lars Kulik, Shanika Karunasekera, Egemen Tanin, Rui Zhang 0003, Eman Bin Khunayn |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | Efficient Spatio-Temporal Tactile Object Recognition with Randomized Tiling Convolutional Networks in a Hierarchical Fusion StrategyabstractRobotic tactile recognition aims at identifying target objects or environments from tactile sensory readings. The advancement of unsupervised feature learning and biological tactile sensing inspire us proposing the model of 3T-RTCN that performs spatio-temporal feature representation and fusion for tactile recognition. It decomposes tactile data into spatial and temporal threads, and incorporates the strength of randomized tiling convolutional networks. Experimental evaluations show that it outperforms some state-of-the-art methods with a large margin regarding recognition accuracy, robustness, and fault-tolerance; we also achieve an order-of-magnitude speedup over equivalent networks with pretraining and finetuning. Practical suggestions and hints are summarized in the end for effectively handling the tactile data. Le-le Cao, Kotagiri Ramamohanarao, Fuchun Sun 0001, Hongbo Li 0001, Wenbing Huang 0001, Zay Maung Maung Aye |
AAAI | 2 |
| 2016 | Efficient Mining of Pan-Correlation Patterns from Time Course Data
Qian Liu 0014, Jinyan Li 0001, Limsoon Wong, Kotagiri Ramamohanarao |
ADMA | 4 |
| 2016 | Automatic Generation and Validation of Road Maps from GPS Trajectory Data SetsabstractWith the popularity of mobile GPS devices such as on-board navigation systems and smart phones, users can contribute their GPS trajectory data for creating geo-volunteered road maps. However, the quality of these road maps cannot be guaranteed due to the lack of expertise among contributing users. Therefore, important challenges are (i) to automatically generate accurate roads from GPS traces and (ii) to validate the correctness of existing road maps. To address these challenges, we propose a novel Spatial-Linear Clustering (SLC) technique to infer road segments from GPS traces. In our algorithm, we propose the use of spatial-linear clusters to appropriately represent the linear nature of GPS points collected from the same road segment. Through inferring road segments our algorithm can detect missing roads and checking the correctness of existing road network. For our evaluation, we conduct extensive experiments that compare our method to the state-of-the-art methods on two real data sets. The experimental results show that the F1 score of our algorithm is on average 10.7% higher than the best state-of-the-art method. Hengfeng Li, Lars Kulik, Kotagiri Ramamohanarao |
CIKM | 3 |
| 2016 | Scalable Local-Recoding Anonymization using Locality Sensitive Hashing for Big Data Privacy PreservationabstractWhile cloud computing has become an attractive platform for supporting data intensive applications, a major obstacle to the adoption of cloud computing in sectors such as health and defense is the privacy risk associated with releasing datasets to third-parties in the cloud for analysis. A widely-adopted technique for data privacy preservation is to anonymize data via local recoding. However, most existing local-recoding techniques are either serial or distributed without directly optimizing scalability, thus rendering them unsuitable for big data applications. In this paper, we propose a highly scalable approach to local-recoding anonymization in cloud computing, based on Locality Sensitive Hashing (LSH). Specifically, a novel semantic distance metric is presented for use with LSH to measure the similarity between two data records. Then, LSH with the MinHash function family can be employed to divide datasets into multiple partitions for use with MapReduce to parallelize computation while preserving similarity. By using our efficient LSH-based scheme, we can anonymize each partition through the use of a recursive agglomerative $k$-member clustering algorithm. Extensive experiments on real-life datasets show that our approach significantly improves the scalability and time-efficiency of local-recoding anonymization by orders of magnitude over existing approaches. Xuyun Zhang, Christopher Leckie, Wan-Chun Dou, Jinjun Chen, Kotagiri Ramamohanarao, Zoran A. Salcic |
CIKM | 5 |
| 2016 | Tensor canonical correlation analysis for multi-view dimension reductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multiview canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
ICDE | 3 |
| 2016 | Training robust models using Random ProjectionabstractRegularization plays an important role in machine learning systems. We propose a novel methodology for model regularization using random projection. We demonstrate the technique on neural networks, since such models usually comprise a very large number of parameters, calling for strong regularizers. It has been shown recently that neural networks are sensitive to two kinds of samples: (i) adversarial samples, which are generated by imperceptible perturbations of previously correctly-classified samples—yet the network will misclassify them; and (ii) fooling samples, which are completely unrecognizable, yet the network will classify them with extremely high confidence. In this paper, we show how robust neural networks can be trained using random projection. We show that while random projection acts as a strong regularizer, boosting model accuracy similar to other regularizers, such as weight decay and dropout, it is far more robust to adversarial noise and fooling samples. We further show that random projection also helps to improve the robustness of traditional classifiers, such as Random Forrest and Gradient Boosting Machines. Xuan Vinh Nguyen, Sarah M. Erfani, Sakrapee Paisitkriangkrai, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao |
ICPR | 6 |
| 2016 | Robust Domain Generalisation by Enforcing Distribution Invariance
Sarah M. Erfani, Mahsa Baktash, Masud Moshtaghi, Vinh Nguyen 0003, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
IJCAI | 7 |
| 2016 | Large Scale Metric learningabstractMany machine learning and pattern recognition algorithms rely heavily on good distance metrics to achieve competitive performance. While distance metrics can be learned, the computational expense of doing so is currently infeasible on large datasets. In this paper, we propose two efficient-and-effective approaches for selecting the training dataset using Locality-Sensitive Hashing (LSH) with discriminative information, and with K-Means clustering inside LSH buckets, for accelerating metric learning. Our methods yield a speedup factor of (N/C)2, where N is training set size and C ≪ N is the user-selected compressed set size, achieving quadratic speedup to metric learning often realized as a 1–2 or more orders of magnitude improvement with little degradation to accuracy. For example, our generic filter approach enables the current overall fastest Large Margin Nearest Neighbor (LMNN) to learn metrics on one million samples in 6.8 minutes down from 5.4hrs—a 48x speedup. LMNN and similar state-of-the-art methods use tree data structures to speed up nearest-neighbor queries—an advantage that degrades at higher dimensions. Our approach does not share this limitation. Zay Maung Maung Aye, Kotagiri Ramamohanarao, Benjamin I. P. Rubinstein |
IJCNN | 2 |
| 2016 | Fast trajectory clustering using Hashing methodsabstractThere has been an explosion in the usage of trajectory data. Clustering is one of the simplest and most powerful approaches for knowledge discovery from trajectories. In order to produce meaningful clusters, well-defined metrics are required to capture the essence of similarity between trajectories. One such distance function is Dynamic Time Warping (DTW), which aligns two trajectories together in order to determine similarity. DTW has been widely accepted as a very good distance measure for trajectory data. However, trajectory clustering is very expensive due to the complexity of the similarity functions, for example, DTW has a high computational cost O(n2), where n is the average length of the trajectory, which makes the clustering process very expensive. In this paper, we propose the use of hashing techniques based on Distance-Based Hashing (DBH) and Locality Sensitive Hashing (LSH) to produce approximate clusters and speed up the clustering process. Zay Maung Maung Aye, Benjamin I. P. Rubinstein, Kotagiri Ramamohanarao |
IJCNN | 4 |
| 2016 | Node Re-Ordering as a Means of Anomaly Detection in Time-Evolving Graphs
Lida Rashidi, Andrey Kan, James Bailey 0001, Jeffrey Chan, Christopher Leckie, Wei Liu 0007, Sutharshan Rajasegarar, Kotagiri Ramamohanarao |
ECML/PKDD (2) | 8 |
| 2016 | R1STM: One-class Support Tensor Machine with Randomised KernelabstractIdentifying unusual or anomalous patterns in an underlying dataset is an important but challenging task in many applications. The focus of the unsupervised anomaly detection literature has mostly been on vectorised data. However, many applications are more naturally described using higher-order tensor representations. Approaches that vectorise tensorial data can destroy the structural information encoded in the high-dimensional space, and lead to the problem of the curse of dimensionality. In this paper we present the first unsupervised tensorial anomaly detection method, along with a randomised version of our method. Our anomaly detection method, the One-class Support Tensor Machine (1STM), is a generalisation of conventional one-class Support Vector Machines to higher-order spaces. 1STM preserves the multiway structure of tensor data, while achieving significant improvement in accuracy and efficiency over conventional vectorised methods. We then leverage the theory of nonlinear random projections to propose the Randomised 1STM (R1STM). Our empirical analysis on several real and synthetic datasets shows that our R1STM algorithm delivers comparable or better accuracy to a state-of-the-art deep learning method and traditional kernelised approaches for anomaly detection, while being approximately 100 times faster in training and testing. Sarah M. Erfani, Mahsa Baktash, Sutharshan Rajasegarar, Vinh Nguyen 0003, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
SDM | 7 |
| 2016 | WTEN: An Advanced Coupled Tensor Factorization Strategy for Learning from Imbalanced Data
Thanh Pham, Wei Liu 0007, Kotagiri Ramamohanarao |
WISE (1) | 4 |
| 2016 | Discovering outlying aspects in large datasets
Xuan Vinh Nguyen, Jeffrey Chan, Simone Romano 0003, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jian Pei 0001 |
Data Min. Knowl. Discov. | 6 |
| 2016 | Enhancing Reliability of Workflow Execution Using Task Replication and Spot InstancesabstractCloud environments offer low-cost computing resources as a subscription-based service. These resources are elastically scalable and dynamically provisioned. Furthermore, cloud providers have also pioneered new pricing models like spot instances that are cost-effective. As a result, scientific workflows are increasingly adopting cloud computing. However, spot instances are terminated when the market price exceeds the users bid price. Likewise, cloud is not a utopian environment. Failures are inevitable in such large complex distributed systems. It is also well studied that cloud resources experience fluctuations in the delivered performance. These challenges make fault tolerance an important criterion in workflow scheduling. This article presents an adaptive, just-in-time scheduling algorithm for scientific workflows. This algorithm judiciously uses both spot and on-demand instances to reduce cost and provide fault tolerance. The proposed scheduling algorithm also consolidates resources to further minimize execution time and cost. Extensive simulations show that the proposed heuristics are fault tolerant and are effective, especially under short deadlines, providing robust schedules with minimal makespan and cost. Deepak Poola, Kotagiri Ramamohanarao, Rajkumar Buyya |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2016 | Visual Assessment of Clustering Tendency for Incomplete DataabstractThe iVAT (asiVAT) algorithms reorder symmetric (asymmetric) dissimilarity data so that an image of the data may reveal cluster substructure. Images formed from incomplete data don't offer a very rich interpretation of cluster structure. In this paper, we examine four methods for completing the input data with imputed values before imaging. We choose a best method using contaminated versions of the complete Iris data, for which the desired results are known. Then, we analyze two real world data sets from social networks that are incomplete using the best imputation method chosen in the juried trials with Iris: (i) Sampson's monastery data, an incomplete, asymmetric relation matrix; and (ii) the karate club data, comprising a symmetric similarity matrix that is about 86 percent incomplete. Laurence Anthony F. Park, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao, James Bailey 0001, Marimuthu Palaniswami |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2015 | Traffic forecasting in complex urban networks: Leveraging big data and machine learningabstractAccurate network-wide real time traffic forecasting is essential for next generation smart cities. In this context, we study a novel and complex traffic data set and explore the potential to apply big data and machine learning analysis. We evaluate several hypotheses and find that the availability of big data is able to facilitate more accurate predictions. Furthermore, we find that spatial aspects have more influence than temporal ones and that careful choice of thresholding parameters is crucial for high performance classification. Florin Schimbinschi, Xuan Vinh Nguyen, James Bailey 0001, Christopher Leckie, Hai Le Vu 0001, Kotagiri Ramamohanarao |
IEEE BigData | 6 |
| 2015 | Disc segmentation and BMO-MRW measurement from SD-OCT image using graph search and tracing of three bench mark reference layers of retinaabstractSpectral Domain Optical Coherence Tomography (SD-OCT) imaging is a robust diagnostic tool for visualizing pathology in the retina for glaucoma and other diseases. Bruch's Membrane Opening minimum rim width (BMO-MRW) of retina is a good measurable marker for distinguishing glaucoma and normal individuals. In this paper, we have proposed a novel automatic method to compute BMO-MRW from segmenting the Inner limiting membrane (ILM) and Bruch's Membrane Opening (BMO). Our technique uses the location tracking of three bench mark reference (TBMR) layers of the retina that reduce search space with improving accuracy. This also allows locating the boundary of the optic disc. The accuracy was tested against manually detected disc boundary for 13 subjects. The precision, recall and F1 score shows more than 95% for disc area and unsigned mean error rate of BMO-MRW is 58.62 db 43.12 (Mean ± SD) μm. Md. Akter Hussain, Alauddin Bhuiyan, Kotagiri Ramamohanarao |
ICIP | 3 |
| 2015 | Detection of Deception in the Mafia Party GameabstractThe problem of deception detection is very challenging. Only trained people with specialist knowledge are able to demonstrate an accuracy that is sufficiently higher than random predictions. We present a multi-stage automatic system for extracting features from facial cues and evaluate it on the Mafia game database which we have collected. It is a large database of truthful and deceptive people, recorded in conditions more variable and realistic than many other databases of similar kind. We demonstrate that using the extracted features we are able to correctly classify instances with an average AUC (area under the ROC curve) equal to 0.61, significantly better than random predictions. Sergey Demyanov, James Bailey 0001, Kotagiri Ramamohanarao, Christopher Leckie |
ICMI | 3 |
| 2015 | Big Data Analytics-Enhanced Cloud Computing: Challenges, Architectural Elements, and Future DirectionsabstractThe emergence of cloud computing has made dynamic provisioning of elastic capacity to applications on-demand. Cloud data centers contain thousands of physical servers hosting orders of magnitude more virtual machines that can be allocated on demand to users in a pay-as-you-go model. However, not all systems are able to scale up by just adding more virtual machines. Therefore, it is essential, even for scalable systems, to project workloads in advance rather than using a purely reactive approach. Given the scale of modern cloud infrastructures generating real time monitoring information, along with all the information generated by operating systems and applications, this data poses the issues of volume, velocity, and variety that are addressed by Big Data approaches. In this paper, we investigate how utilization of Big Data analytics helps in enhancing the operation of cloud computing environments. We discuss diverse applications of Big Data analytics in clouds, open issues for enhancing cloud operations via Big Data analytics, and architecture for anomaly detection and prevention in clouds along with future research directions. Rajkumar Buyya, Kotagiri Ramamohanarao, Christopher Leckie, Rodrigo N. Calheiros, Amir Vahid Dastjerdi, Steve Versteeg |
ICPADS | 2 |
| 2015 | SLA-Based Resource Scheduling for Big Data Analytics as a Service in Cloud Computing EnvironmentsabstractData analytics plays a significant role in gaining insight of big data that can benefit in decision making and problem solving for various application domains such as science, engineering, and commerce. Cloud computing is a suitable platform for Big Data Analytic Applications (BDAAs) that can greatly reduce application cost by elastically provisioning resources based on user requirements and in a pay as you go model. BDAAs are typically catered for specific domains and are usually expensive. Moreover, it is difficult to provision resources for BDAAs with fluctuating resource requirements and reduce the resource cost. As a result, BDAAs are mostly used by large enterprises. Therefore, it is necessary to have a general Analytics as a Service (AaaS) platform that can provision BDAAs to users in various domains as consumable services in an easy to use way and at lower price. To support the AaaS platform, our research focuses on efficiently scheduling Cloud resources for BDAAs to satisfy Quality of Service (QoS) requirements of budget and deadline for data analytic requests and maximize profit for the AaaS platform. We propose an admission control and resource scheduling algorithm, which not only satisfies QoS requirements of requests as guaranteed in Service Level Agreements (SLAs), but also increases the profit for AaaS providers by offering a cost-effective resource scheduling solution. We propose the architecture and models for the AaaS platform and conduct experiments to evaluate the proposed algorithm. Results show the efficiency of the algorithm in SLA guarantee, profit enhancement, and cost saving. Yali Zhao, Rodrigo N. Calheiros, Graeme Gange, Kotagiri Ramamohanarao, Rajkumar Buyya |
ICPP | 4 |
| 2015 | Scalable Outlying-Inlying Aspects Discovery via Feature Ranking
Xuan Vinh Nguyen, Jeffrey Chan, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jian Pei 0001 |
PAKDD (2) | 5 |
| 2015 | Robust inferences of travel paths from GPS trajectoriesabstractMonitoring and predicting traffic conditions are of utmost importance in reacting to emergency events in time and for computing the real-time shortest travel-time path. Mobile sensors, such as GPS devices and smartphones, are useful for monitoring urban traffic due to their large coverage area and ease of deployment. Many researchers have employed such sensed data to model and predict traffic conditions. To do so, we first have to address the problem of associating GPS trajectories with the road network in a robust manner. Existing methods rely on point-by-point matching to map individual GPS points to a road segment. However, GPS data is imprecise due to noise in GPS signals. GPS coordinates can have errors of several meters and, therefore, direct mapping of individual points is error prone. Acknowledging that every GPS point is potentially noisy, we propose a radically different approach to overcome inaccuracy in GPS data. Instead of focusing on a point-by-point approach, our proposed method considers the set of relevant GPS points in a trajectory that can be mapped together to a road segment. This clustering approach gives us a macroscopic view of the GPS trajectories even under very noisy conditions. Our method clusters points based on the direction of movement as a spatial-linear cluster, ranks the possible route segments in the graph for each group, and searches for the best combination of segments as the overall path for the given set of GPS points. Through extensive experiments on both synthetic and real datasets, we demonstrate that, even with highly noisy GPS measurements, our proposed algorithm outperforms state-of-the-art methods in terms of both accuracy and computational cost. Hengfeng Li, Lars Kulik, Kotagiri Ramamohanarao |
Int. J. Geogr. Inf. Sci. | 3 |
| 2015 | Revenue Maximization with Optimal Capacity Control in Infrastructure as a Service Cloud MarketsabstractInfrastructure-as-a-Service cloud providers offer diverse purchasing options and pricing plans, namely on-demand, reservation, and spot market plans. This allows them to efficiently target a variety of customer groups with distinct preferences and to generate more revenue accordingly. An important consequence of this diversification however, is that it introduces a non-trivial optimization problem related to the allocation of the provider's available data center capacity to each pricing plan. The complexity of the problem follows from the different levels of revenue generated per unit of capacity sold, and the different commitments consumers and providers make when resources are allocated under a given plan. In this work, we address a novel problem of maximizing revenue through an optimization of capacity allocation to each pricing plan by means of admission control for reservation contracts, in a setting where aforementioned plans are jointly offered to customers. We devise both an optimal algorithm based on a stochastic dynamic programming formulation and two heuristics that trade-off optimality and computational complexity. Our evaluation, which relies on an adaptation of a large-scale real-world workload trace of Google, shows that our algorithms can significantly increase revenue compared to an allocation without capacity control given that sufficient resource contention is present in the system. In addition, we show that our heuristics effectively allow for online decision making and quantify the revenue loss caused by the assumptions made to render the optimization problem tractable. Adel Nadjaran Toosi, Kurt Vanmechelen, Kotagiri Ramamohanarao, Rajkumar Buyya |
IEEE Trans. Cloud Comput. | 3 |
| 2015 | Tensor Canonical Correlation Analysis for Multi-View Dimension ReductionabstractCanonical correlation analysis (CCA) has proven an effective tool for two-view dimension reduction due to its profound theoretical foundation and success in practical applications. In respect of multi-view learning, however, it is limited by its capability of only handling data represented by two-view features, while in many real-world applications, the number of views is frequently many more. Although the ad hoc way of simultaneously exploring all possible pairs of features can numerically deal with multi-view data, it ignores the high order statistics (correlation information) which can only be discovered by simultaneously exploring all features. Therefore, in this work, we develop tensor CCA (TCCA) which straightforwardly yet naturally generalizes CCA to handle the data of an arbitrary number of views by analyzing the covariance tensor of the different views. TCCA aims to directly maximize the canonical correlation of multiple (more than two) views. Crucially, we prove that the main problem of multi-view canonical correlation maximization is equivalent to finding the best rank-1 approximation of the data covariance tensor, which can be solved efficiently using the well-known alternating least squares (ALS) algorithm. As a consequence, the high order correlation information contained in the different views is explored and thus a more reliable common subspace shared by all features can be obtained. In addition, a non-linear extension of TCCA is presented. Experiments on various challenge tasks, including large scale biometric structure prediction, internet advertisement classification, and web image annotation, demonstrate the effectiveness of the proposed method. Yong Luo 0002, Dacheng Tao, Kotagiri Ramamohanarao, Chao Xu 0006, Yonggang Wen 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2015 | Indexable online time series segmentation with error bound guarantee
Jianzhong Qi 0001, Rui Zhang 0003, Kotagiri Ramamohanarao, Hongzhi Wang 0001, Zeyi Wen |
World Wide Web | 3 |
| 2014 | Robust Scheduling of Scientific Workflows with Deadline and Budget Constraints in CloudsabstractDynamic resource provisioning and the notion of seemingly unlimited resources are attracting scientific workflows rapidly into Cloud computing. Existing works on workflow scheduling in the context of Clouds are either on deadline or cost optimization, ignoring the necessity for robustness. Robust scheduling that handles performance variations of Cloud resources and failures in the environment is essential in the context of Clouds. In this paper, we present a robust scheduling algorithm with resource allocation policies that schedule workflow tasks on heterogeneous Cloud resources while trying to minimize the total elapsed time (make span) and the cost. Our results show that the proposed resource allocation policies provide robust and fault-tolerant schedule while minimizing make span. The results also show that with the increase in budget, our policies increase the robustness of the schedule. Deepak Poola, Saurabh Kumar Garg 0001, Rajkumar Buyya, Yun Yang 0001, Kotagiri Ramamohanarao |
AINA | 5 |
| 2014 | Topic-specific post identification in microblog streamsabstractThe tracking of microblog discussion, on a given topic, is useful for a wide range of higher level applications. Microblog services like Twitter provide a simple keyword based tracking capability, where any tweet containing a keyword is returned. Due to the short length of microblog posts, using a small number of topic specific query words for tracking, would impact recall. Use of a larger number of keywords (compared to regular document retrieval) is generally required in order to obtain good recall, but this would result in a large number of off-topic posts, resulting in low precision. In our work, we consider the scenario of using a large number of query terms to maintain high recall, for automated tracking of a microblog streams. The challenge we address is how to score each of the returned microblogs, with respect to the query, on-line, in an unsupervised manner, so as to identify those that are on topic. To this end, we proposed a new term-scoring expression, which we call Adjusted Information Gain (AIG), and we compare this to other term-scoring expressions: inverse document frequency, Dice, Jaccard and keyword frequency. Our comparisons consider a selection of document-scoring functions applied to roughly 40 million tweets collects over a 20 day period for each of two topics. Our results show significant improvements (from 8%-40% of the area under the ROC curves) to existing term-scoring expressions, depending on topic and specificity, and provide insight into further work in query expansion techniques. Shanika Karunasekera, Aaron Harwood, Sameendra Samarawickrama, Kotagiri Ramamohanarao, Garry Robins |
IEEE BigData | 4 |
| 2014 | Volatility homogenisation decomposition for forecastingabstractWe explore the idea that by modeling a financial time series at regular points in space (i.e. price) rather than regular points in time, more predictive power can be extracted from the time series. We will term this concept of modeling time series at regular points in space as “volatility homogenisation”. Our hypothesis is that if we select the correct quantum in terms of regular steps in space, we replace noise which can normally interfere with prediction methods and thus uncover the underlying patterns in the time series. Furthermore, this technique can also be viewed a way of decoupling spatial and temporal dependence, which again, can replace unnecessary noise. We apply this decomposition to nine different financial time series and then apply support vector classification in order to make our predictions on the decomposed time series. Our results show that in the majority of cases, this technique yields better predictions than applications to data that has regular points in time, with applications of techniques such as support vector regression and Autoregressive Integrated Moving Averages models. The contribution of this paper is that it demonstrates the efficacy of this new methodology known as “volatility homogenisation”. Adam W. Kowalewski, Owen D. Jones, Kotagiri Ramamohanarao |
CIFEr | 3 |
| 2014 | Using Local Information to Significantly Improve Classification PerformanceabstractIn this research we propose to derive new features based on data samples' local information with the aim of improving the performance of general supervised learning algorithms. The creation of new features is inspired by the measure of average precision which is known to be a robust measure that is insensitive to the number of retrieved items in information retrieval. We use the idea of average precision to weight the neighbours of an instance and show that this weighting strategy is insensitive to the number of neighbours in the locality. Information captured in the new features allows a general classifier to learn additional useful peripheral knowledge that are helpful in building effective classification models. We comprehensively evaluate our method on real datasets and the results show substantial improvements in the performance of classifiers including SVM, Bayesian networks, random forest, and C4.5. Wei Liu 0007, Dong Lee, Kotagiri Ramamohanarao |
CIKM | 3 |
| 2014 | Enabling Precision/Recall Preferences for Semi-supervised SVM TrainingabstractSemi-supervised learning is an essential approach to classification when the available labeled data is insufficient and we need to also make use of unlabeled data in the learning process. Numerous research efforts have focused on designing algorithms to improve the F1 score, but have any mechanism to control precision or recall individually. However, many applications have precision/recall preferences. For instance, an email spam classifier requires a precision of 0.9 to mitigate the false dismissal of useful emails. In this paper, we propose a method that allows to specify a precision/recall preference while maximising the F1 score. Our key idea is that we divide the semi-supervised learning process into multiple rounds of supervised learning, and the classifier learned at each round is calibrated using a sub-set of the labeled dataset before we use it on the unlabeled dataset for enlarging the training dataset. Our idea is applicable to a number of learning models such as Support Vector Machines (SVMs), Bayesian networks and neural networks. We focus our research and the implementation of our idea on SVMs. We conduct extensive experiments to validate the effectiveness of our method. The experimental results show that our method can train classifiers with a precision/recall preference, while the popular semi-supervised SVM training algorithm (which we use as the baseline) cannot. When we specify the precision preference and the recall preference to be the same, which indicates to maximise the F1 score only as the baseline does, our method achieves better or similar F1 scores to the baseline. An additional advantage of our method is that it converges much faster than the baseline. Zeyi Wen, Rui Zhang 0003, Kotagiri Ramamohanarao |
CIKM | 3 |
| 2014 | Spatio-temporal trajectory simplification for inferring travel pathsabstractMining GPS trajectories of moving vehicles has led to many research directions, such as traffic modeling and driving predication. An important challenge is how to map GPS traces to a road network accurately under noisy conditions. However, to the best of our knowledge, there is no existing work that first simplifies a trajectory to improve map matching. In this paper we propose three trajectory simplification algorithms that can deal with both offline and online trajectory data. We use weighting functions to incorporate spatial knowledge, such as segment lengths and turning angles, into our simplification algorithms. In addition, we measure the noise degree of a GPS point based on its spatio-temporal relationship to its neighbors. The effectiveness of our algorithms is comprehensively evaluated on real trajectory datasets with varying the noise levels and sampling rates. Our evaluation shows that under highly noisy conditions, our proposed algorithms considerably improve map matching accuracy and reduce computational costs compared to the state-of-the-art methods. Hengfeng Li, Lars Kulik, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 3 |
| 2014 | TRIBAC: Discovering Interpretable Clusters and Latent Structures in GraphsabstractGraphs are a powerful representation of relational data, such as social and biological networks. Often, these entities form groups and are organised according to a latent structure. However, these groupings and structures are generally unknown and it can be difficult to identify them. Graph clustering is an important type of approach used to discover these vertex groups and the latent structure within graphs. One type of approach for graph clustering is non-negative matrix factorisation However, the formulations of existing factorisation approaches can be overly relaxed and their groupings and results consequently difficult to interpret, may fail to discover the true latent structure and groupings, and converge to extreme solutions. In this paper, we propose a new formulation of the graph clustering problem that results in clusterings that are easy to interpret. Combined with a novel algorithm, the clusterings are also more accurate than state-of-the-art algorithms for both synthetic and real datasets. Jeffrey Chan, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
ICDM | 4 |
| 2014 | MASCOT: Fast and Highly Scalable SVM Cross-Validation Using GPUs and SSDsabstractCross-validation is a commonly used method for evaluating the effectiveness of Support Vector Machines (SVMs). However, existing SVM cross-validation algorithms are not scalable to large datasets because they have to (i) hold the whole dataset in memory and/or (ii) perform a very large number of kernel value computation. In this paper, we propose a scheme to dramatically improve the scalability and efficiency of SVM cross-validation through the following key ideas. (i) To avoid holding the whole dataset in the memory and avoid performing repeated kernel value computation, we precompute the kernel values and reuse them. (ii) We store the precomputed kernel values to a high-speed storage framework, consisting of CPU memory extended by solid state drives (SSDs) and GPU memory as a cache, so that reusing (i.e., Reading) kernel values takes much lesser time than computing them on-the-fly. (iii) To further improve the efficiency of the SVM training, we apply a number of techniques for the extreme example search algorithm, design a parallel kernel value read algorithm, propose a caching strategy well-suited to the characteristics of the storage framework, and parallelize the tasks on the GPU and the CPU. For datasets of sizes that existing algorithms can handle, our scheme achieves several orders of magnitude of speedup. More importantly, our scheme enables SVM cross-validation on datasets of very large scale that existing algorithms are unable to handle. Zeyi Wen, Rui Zhang 0003, Kotagiri Ramamohanarao, Jianzhong Qi 0001, Kerry L. Taylor |
ICDM | 3 |
| 2014 | Structure-Aware Distance Measures for Comparing Clusterings in Graphs
Jeffrey Chan, Xuan Vinh Nguyen, Wei Liu 0007, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao, Jian Pei 0001 |
PAKDD (1) | 6 |
| 2014 | A Robust Classifier for Imbalanced Datasets
Sori Kang, Kotagiri Ramamohanarao |
PAKDD (1) | 2 |
| 2014 | A spatiotemporal compression based approach for efficient big data processing on Cloud
Chi Yang, Xuyun Zhang, Changmin Zhong, Chang Liu 0001, Jian Pei 0001, Kotagiri Ramamohanarao, Jinjun Chen |
J. Comput. Syst. Sci. | 6 |
| 2014 | Authorized Public Auditing of Dynamic Big Data Storage on Cloud with Efficient Verifiable Fine-Grained UpdatesabstractCloud computing opens a new era in IT as it can provide various elastic and scalable IT services in a pay-as-you-go fashion, where its users can reduce the huge capital investments in their own IT infrastructure. In this philosophy, users of cloud storage services no longer physically maintain direct control over their data, which makes data security one of the major concerns of using cloud. Existing research work already allows data integrity to be verified without possession of the actual data file. When the verification is done by a trusted third party, this verification process is also called data auditing, and this third party is called an auditor. However, such schemes in existence suffer from several common drawbacks. First, a necessary authorization/authentication process is missing between the auditor and cloud service provider, i.e., anyone can challenge the cloud service provider for a proof of integrity of certain file, which potentially puts the quality of the so-called ‘auditing-as-a-service’ at risk; Second, although some of the recent work based on BLS signature can already support fully dynamic data updates over fixed-size data blocks, they only support updates with fixed-sized blocks as basic unit, which we call coarse-grained updates. As a result, every small update will cause re-computation and updating of the authenticator for an entire file block, which in turn causes higher storage and communication overheads. In this paper, we provide a formal analysis for possible types of fine-grained data updates and propose a scheme that can fully support authorized auditing and fine-grained update requests. Based on our scheme, we also propose an enhancement that can dramatically reduce communication overheads for verifying small updates. Theoretical analysis and experimental results demonstrate that our scheme can offer not only enhanced security and flexibility, but also significantly lower overhead for big data applications with a large number of frequent small updates, such as applications in social media and business transactions. Chang Liu 0001, Jinjun Chen, Laurence T. Yang, Xuyun Zhang, Chi Yang, Rajiv Ranjan 0001, Kotagiri Ramamohanarao |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2013 | Spatio-temporal event detection using probabilistic graphical models (PGMs)abstractEvent detection concerns identifying occurrence of interesting events which are meaningful and understandable. In dynamic fields, as time passes the attribute of phenomenon varies in spatial locations. Detecting events in dynamic fields requires an approach to deal with the highly granular data arriving in real time. This paper proposes a spatiotemporal event detection algorithm in dynamic fields which are monitored by wireless sensor networks (WSNs). The algorithm provides a method using probabilistic graphical models (PGMs) in WSNs to cope with the uncertainty of sensor readings. The algorithm incorporates the ability of Markov chains in temporal dependency modelling and Markov random fields theory to model the spatial dependency of sensors in a distributed fashion. Experimental evaluation of the proposed algorithm demonstrates that the decentralized approach improves the F1-score to 82% and 29% better precision than simple threshold technique. In addition, the performance of the algorithm was evaluated and compared with respect to the scalability (in terms of communication complexity). In comparison with the centralized approach the decentralized algorithm can substantially improve the scalability of communication in wireless sensor networks. Azadeh Mousavi, Matt Duckham, Kotagiri Ramamohanarao, Abbas Rajabifard |
CIDM | 3 |
| 2013 | Discovering latent blockmodels in sparse and noisy graphs using non-negative matrix factorisationabstractBlockmodelling is an important technique in social network analysis for discovering the latent structure in graphs. A blockmodel partitions the set of vertices in a graph into groups, where there are either many edges or few edges between any two groups. For example, in the reply graph of a question and answer forum, blockmodelling can identify the group of experts by their many replies to questioners, and the group of questioners by their lack of replies among themselves but many replies from experts. Jeffrey Chan, Wei Liu 0007, Andrey Kan, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
CIKM | 6 |
| 2013 | Clustering and visualization of fuzzy communities in social networksabstractWe discuss a new formulation of a fuzzy validity index that generalizes the Newman-Girvan (NG) modularity function. The NG function serves as a cluster validity functional in community detection studies. The input data is an undirected graph G = (V, E) that represents a social network. Clusters in V correspond to socially similar substructures in the network. We compare our fuzzy modularity to an existing modularity function using the well-studied Karate Club data set. Timothy C. Havens, James C. Bezdek, Christopher Leckie, Jeffrey Chan, Wei Liu 0007, James Bailey 0001, Kotagiri Ramamohanarao, Marimuthu Palaniswami |
FUZZ-IEEE | 7 |
| 2013 | Automatic detection of retinal vascular landmark features for colour fundus image matching and patient longitudinal studyabstractRetinal vascular landmark points such as branching points and crossovers are important features for automatic retinal image matching and vascular abnormality detection. These landmark points can enable automatic screening of large dataset through the detection of vascular network abnormalities (i.e., arteriovenous nicking, retinal vein occlusion) which are important for hypertension and cardiovascular disease prediction. Existing methods for crossover point detection use only local information at each image pixel without considering vascular features to detect crossover positions. This leads to the misclassification of very acute crossovers which are represented by two bifurcation points in the skeleton image. In this article, we propose a robust method that utilizes both local information and vascular geometrical features at the crossing to distinguish crossover from non-crossover points in a retinal image. The proposed method was validated on fifteen high resolution retinal images and the results show that our method achieves higher accuracy than any existing methods. In particular, the proposed method can discover more than 74% (recall) of crossovers with a detection accuracy (fraction of detected crossover points that are correct) of 83% (precision). The detected crossovers provide essential results for the automatic detection of vascular network abnormalities, such as arteriovenous nicking, neovascularization, and retinal vein occlusion. Uyen T. V. Nguyen, Alauddin Bhuiyan, Laurence Anthony F. Park, Ryo Kawasaki, Tien Yin Wong, Kotagiri Ramamohanarao |
ICIP | 6 |
| 2013 | Automated segmentation of multiple sclerosis lesion in intensity enhanced flair MRI using texture features and support vector machineabstractIn this paper, a fully automated segmentation method is proposed to identify Multiple Sclerosis (MS) related white matter lesions from brain magnetic resonance imaging (MRI) data. The main contribution of this paper is to obtain a new texture feature set for MS Lesion segmentation that is a combination of local and global neighbourhood information. The proposed method adopts a robust intensity normalization technique and lesion contrast enhancementfilter for enhancing the region of interest. We use a Support Vector Machine (SVM) to classify lesion pixels and level set based active contour and morphological filtering to achieve higher accuracy on lesion pixel identification. Quantitative evaluation of the proposed method is carried on real MRI data set provided by MS Lesion Challenge 2008. The results obtained from our method indicate significant improvement in performance compare to three state of the art methods that shows the proposed method's high suitability for assisting the neurologist to detect the MS in clinical practice. Pallab Kanti Roy, Alauddin Bhuiyan, Kotagiri Ramamohanarao |
ICIP | 3 |
| 2013 | A k-leader fuel-efficient traffic modelabstractOptimizing travel time and energy consumption without compromising safety to attain efficient road traffic is a key goal in transport telematics. Microscopic traffic simulations are important tools to study the impact of new algorithms on road traffic. The aim of these simulations is to achieve a high degree of realism through the use of microscopic car-following models, which characterize real-time interaction among individual vehicles. These models play a vital role in Advanced Vehicle Control and Safety Systems (AVCSS) such as collision warning, adaptive cruise control, or lane guidance and in modelling simulation of safety studies and capacity analysis in transportation science. Although a range of models have been proposed to model the longitudinal interaction between adjacent vehicles due to its importance, surprisingly few comparative evaluations of the models exist. In this paper, we first identify limitations of the prominent car-following models. We then propose a k-leader fuel-efficient car-following model and show that our model is effective in terms of safety, trip times, flow and fuel efficiency. We also highlight new research challenges and important directions for further research. Tanveer Awal, Lars Kulik, Kotagiri Ramamohanarao |
Intelligent Vehicles Symposium | 3 |
| 2013 | A Bayesian Classifier for Learning from Tensorial Data
Wei Liu 0007, Jeffrey Chan, James Bailey 0001, Christopher Leckie, Fang Chen 0001, Kotagiri Ramamohanarao |
ECML/PKDD (2) | 6 |
| 2013 | Mining Labelled Tensors by Discovering both their Common and Discriminative SubspacesabstractConventional non-negative tensor factorization (NTF) methods assume there is only one tensor that needs to be decomposed to low-rank factors. However, in practice data are usually generated from different time periods or by different class labels, which are represented by a sequence of multiple tensors associated with different labels. This raises the problem that when one needs to analyze and compare multiple tensors, existing NTF is unsuitable for discovering all potentially useful patterns: 1) if one factorizes each tensor separately, the common information shared by the tensors is lost in the factors, and 2) if one concatenates these tensors together and forms a larger tensor to factorize, the intrinsic discriminative subspaces that are unique to each tensor are not captured. The cause of such an issue is from the fact that conventional factorization methods handle data observations in an unsupervised way, which only considers features and not labels of the data. To tackle this problem, in this paper we design a novel factorization algorithm called CDNTF (common and discriminative subspace non-negative tensor factorization), which takes both features and class labels into account in the factorization process. CDNTF uses a set of labelled tensors as input and computes both their common and discriminative subspaces simultaneously as output. We design an iterative algorithm that solves the common and discriminative subspace factorization problem with a proof of convergence. Experiment results on solving graph classification problems demonstrate the power and the effectiveness of the subspaces discovered by our method. James Bailey 0001, Jeffrey Chan, Kotagiri Ramamohanarao, Christopher Leckie, Wei Liu 0007 |
SDM | 3 |
| 2013 | A HMM-based adaptive fuzzy inference system for stock market forecasting
Md. Rafiul Hassan, Kotagiri Ramamohanarao, Joarder Kamruzzaman, Mustafizur Rahman 0003, M. Maruf Hossain |
Neurocomputing | 2 |
| 2013 | An effective retinal blood vessel segmentation method using multi-scale line detection
Uyen T. V. Nguyen, Alauddin Bhuiyan, Laurence Anthony F. Park, Kotagiri Ramamohanarao |
Pattern Recognit. | 4 |
| 2013 | A Soft Modularity Function For Detecting Fuzzy Communities in Social NetworksabstractWe discuss a new formulation of a fuzzy validity index that generalizes the Newman-Girvan (NG) modularity function. The NG function serves as a cluster validity functional in community detection studies. The input data is an undirected weighted graph that represents, e.g., a social network. Clusters correspond to socially similar substructures in the network. We compare our fuzzy modularity with two existing modularity functions using the well-studied Karate Club and American College Football datasets. Timothy C. Havens, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao, Marimuthu Palaniswami |
IEEE Trans. Fuzzy Syst. | 4 |
| 2012 | Utilizing common substructures to speedup tensor factorization for mining dynamic graphsabstractIn large and complex graphs of social, chemical/biological, or other relations, frequent substructures are commonly shared by different graphs or by graphs evolving through different time periods. Tensors are natural representations of these complex time-evolving graph data. A factorization of a tensor provides a high-quality low-rank compact basis for each dimension of the tensor, which facilitates the interpretation of frequent substructures of the original graphs. However, the high computational cost of tensor factorization makes it infeasible for conventional tensor factorization methods to handle large graphs that evolve frequently with time. To address this problem, in this paper we propose a novel iterative tensor factorization (ITF) method whose time complexity is linear in the cardinalities of all dimensions of a tensor. This low time complexity means that when using tensors to represent dynamic graphs, the computational cost of ITF is linear in the size (number of edges/vertices) of graphs and is also linear in the number of time periods over which the graph evolves. More importantly, an error estimation of ITF suggests that its factorization correctness is comparable to that of the standard factorization method. We empirically evaluate our method on publication networks and chemical compound graphs, and demonstrate that ITF is an order of magnitude faster than the conventional method and at the same time preserves factorization quality. To the best of our knowledge, this research is the first work that uses important frequent substructures to speed up tensor factorizations for mining dynamic graphs. Wei Liu 0007, Jeffrey Chan, James Bailey 0001, Christopher Leckie, Kotagiri Ramamohanarao |
CIKM | 5 |
| 2012 | On compressing weighted time-evolving graphsabstractExisting graph compression techniquesmostly focus on static graphs. However for many practical graphs such as social networks the edge weights frequently change over time. This phenomenon raises the question of how to compress dynamic graphs while maintaining most of their intrinsic structural patterns at each time snapshot. In this paper we show that the encoding cost of a dynamic graph is proportional to the heterogeneity of a three dimensional tensor that represents the dynamic graph. We propose an effective algorithm that compresses a dynamic graph by reducing the heterogeneity of its tensor representation, and at the same time also maintains a maximum lossy compression error at any time stamp of the dynamic graph. The bounded compression error benefits compressed graphs in that they retain good approximations of the original edge weights, and hence properties of the original graph (such as shortest paths) are well preserved. To the best of our knowledge, this is the first work that compresses weighted dynamic graphs with bounded lossy compression error at any time snapshot of the graph. Wei Liu 0007, Andrey Kan, Jeffrey Chan, James Bailey 0001, Christopher Leckie, Jian Pei 0001, Kotagiri Ramamohanarao |
CIKM | 7 |
| 2012 | An adaptive algorithm for online time series segmentation with error bound guaranteeabstractThe volume of time series data grows rapidly in various applications such as network traffic management, telecommunications, finance and sensor network. To reduce the cost of storage, transmission and processing of time series data, the need for more compact representations of time series data is compelling. Segmentation is one of the most commonly used methods to meet this requirement. Both PLA and PPA are common segmentation methods which divide a time series into segments and use a linear function or a polynomial function to approximate each segment, respectively. However, while most of the current PLA and PPA methods aim to minimize the holistic error between the approximation and the original time series, few works try to represent time series as compact as possible with an error bound guarantee on each data point. Furthermore, in many real world situations, the patterns of the time series do not follow a constant rule such that using only one type of functions may not yield the best compaction. Rui Zhang 0003, Kotagiri Ramamohanarao, Parampalli Udaya |
EDBT | 3 |
| 2012 | Privacy aware trajectory determination in road traffic networksabstractOrigin-destination matrices are important for effective real time traffic management. These matrices contain the spatial and temporal distribution of traffic demand, which is a vital input for transportation planning processes. Trajectories represent the different traffic flow routes between the source destination pairs taken by the travelling vehicles. We present a privacy aware model to compute origin-destination (OD) matrix and the corresponding trajectories. Our main contribution is the use of partial vehicle information to compute the trajectories and OD matrix while maintaining accuracy and privacy. We propose local re-identification of vehicles to build local transition matrices at every node of the network. We introduce k-anonymous l-grouping for our trajectory estimation to provide a trade-off between accuracy and privacy. We present a Trajectory Estimation algorithm to determine trajectories and estimate the OD matrix. Preeti Goel, Lars Kulik, Kotagiri Ramamohanarao |
SIGSPATIAL/GIS | 3 |
| 2012 | SeqiBloc: mining multi-time spanning blockmodels in dynamic graphsabstractBlockmodelling is an important technique for decomposing graphs into sets of roles. Vertices playing the same role have similar patterns of interactions with vertices in other roles. These roles, along with the role to role interactions, can succinctly summarise the underlying structure of the studied graphs. As the underlying graphs evolve with time, it is important to study how their blockmodels evolve too. This will enable us to detect role changes across time, detect different patterns of interactions, for example, weekday and weekend behaviour, and allow us to study how the structure in the underlying dynamic graph evolves. To date, there has been limited research on studying dynamic blockmodels. They focus on smoothing role changes between adjacent time instances. However, this approach can overfit during stationary periods where the underling structure does not change but there is random noise in the graph. Therefore, an approach to a) find blockmodels across spans of time and b) to find the stationary periods is needed. In this paper, we propose an information theoretic framework, SeqiBloc, combined with a change point detection approach to achieve a) and b). In addition, we propose new vertex equivalence definitions that include time, and show how they relate back to our information theoretic approach. We demonstrate their usefulness and superior accuracy over existing work on synthetic and real datasets. Jeffrey Chan, Wei Liu 0007, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
KDD | 5 |
| 2012 | Sentiment Analysis by Augmenting Expectation Maximisation with Lexical Knowledge
Xiuzhen Zhang 0001, Yun Zhou 0002, James Bailey 0001, Kotagiri Ramamohanarao |
WISE | 4 |
| 2012 | Probabilistic Voronoi diagrams for probabilistic moving nearest neighbor queries
Mohammed Eunus Ali, Egemen Tanin, Rui Zhang 0003, Kotagiri Ramamohanarao |
Data Knowl. Eng. | 4 |
| 2012 | Continuous Detour Queries in Spatial NetworksabstractWe study the problem of finding the shortest route between two locations that includes a stopover of a given type. An example scenario of this problem is given as follows: “On the way to Bob's place, Alice searches for a nearby take-away Italian restaurant to buy a pizza.” Assuming that Alice is interested in minimizing the total trip distance, this scenario can be modeled as a query where the current Alice's location (start) and Bob's place (destination) function as query points. Based on these two query points, we find the minimum detour object (MDO), i.e., a stopover that minimizes the sum of the distances: 1) from the start to the stopover, and 2) from the stopover to the destination. In a realistic location-based application environment, a user can be indecisive about committing to a particular detour option. The user may wish to browse multiple (k) MDOs before making a decision. Furthermore, when a user moves, the k{\rm MDO} results at one location may become obsolete. We propose a method for continuous detour query (CDQ) processing based on incremental construction of a shortest path tree. We conducted experimental studies to compare the performance of our proposed method against two methods derived from existing k-nearest neighbor querying techniques using real road-network data sets. Experimental results show that our proposed method significantly outperforms the two competitive techniques. Sarana Nutanong, Egemen Tanin, Jie Shao 0001, Rui Zhang 0003, Kotagiri Ramamohanarao |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2011 | ScoreTree: A Decentralised Framework for Credibility Management of User-Generated Content
Yang Liao, Aaron Harwood, Kotagiri Ramamohanarao |
DAIS | 3 |
| 2011 | The role of KL divergence in anomaly detectionabstractWe study the role of Kullback-Leibler divergence in the framework of anomaly detection, where its abilities as a statistic underlying detection have never been investigated in depth. We give an in-principle analysis of network attack detection, showing explicitly attacks may be masked at minimal cost through 'camouflage'. We illustrate on both synthetic distributions and ones taken from real traffic. Darryl Veitch, Kotagiri Ramamohanarao |
SIGMETRICS | 3 |
| 2011 | Approximate pairwise clustering for large data sets via sampling plus extension
Liang Wang 0001, Christopher Leckie, Kotagiri Ramamohanarao, James C. Bezdek |
Pattern Recognit. | 3 |
| 2011 | Handoff Optimization Using Hidden Markov ModelabstractThis letter establishes the similarity between the sensor scheduling problem and the handoff (i.e., base station assignment) problem in cellular networks. A mobile user behavior is then modelled by a Hidden Markov Model (HMM). The handoff problem is formulated as an optimization problem of base station scheduling that minimizes a cost function that involves the HMM state estimation error and base station measurement costs. The optimization problem can be solved using algorithms known as partially observed Markov decision processes. Malka N. Halgamuge, Kotagiri Ramamohanarao, Moshe Zukerman, Hai Le Vu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2011 | Multiresolution Web Link Analysis Using Generalized Link RelationsabstractWeb link analysis methods such as PageRank, HITS, and SALSA have focused on obtaining global popularity or authority of the set of Web pages in question. Although global popularity is useful for general queries, we find that global popularity is not as useful for queries in which the global population has less knowledge of. By examining the many different communities that appear within a Web page graph, we are able to compute the popularity or authority from a specific community. Multiresolution popularity lists allow us to observe the popularity of Web pages with respect to communities at different resolutions within the Web. Multiresolution popularity lists have been shown to have high potential when compared against PageRank. In this paper, we generalize the multiresolution popularity analysis to use any form of Web page link relations. We provide results for both the PageRank relations and the In-degree relations. By utilizing the multiresolution popularity lists, we achieve a 13 percent and 25 percent improvement in mean average precision over In-degree and PageRank, respectively. Laurence Anthony F. Park, Kotagiri Ramamohanarao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2011 | A model for spectra-based software diagnosisabstractThis article presents an improved approach to assist diagnosis of failures in software (fault localisation) by ranking program statements or blocks in accordance with to how likely they are to be buggy. We present a very simple single-bug program to model the problem. By examining different possible execution paths through this model program over a number of test cases, the effectiveness of different proposed spectral ranking methods can be evaluated in idealised conditions. The results are remarkably consistent to those arrived at empirically using the Siemens test suite and Space benchmarks. The model also helps identify groups of metrics that are equivalent for ranking. Due to the simplicity of the model, an optimal ranking method can be devised. This new method out-performs previously proposed methods for the model program, the Siemens test suite and Space. It also helps provide insight into other ranking methods. Lee Naish, Hua Jie Lee, Kotagiri Ramamohanarao |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2010 | Contrast Pattern Mining and Its Application for Building Robust Classifiers
Kotagiri Ramamohanarao |
ALT | 1 |
| 2010 | Statements versus Predicates in Spectral Bug LocalizationabstractThis paper investigates the relationship between the use of predicate-based and statement-based program spectra for bug localization. Branch and path spectra are also considered. Although statement and predicate spectra can be based on the same raw data, the way the data is aggregated results in different information being lost. We propose a simple and cheap modification to the statement-based approach which retains strictly more information. This allows us to compare statement and predicate ''metrics'' (functions used to rank the statements, predicates or paths). We show that improved bug localization performance is possible using single-bug models and benchmarks. Lee Naish, Hua Jie Lee, Kotagiri Ramamohanarao |
APSEC | 3 |
| 2010 | Effective Software Bug Localization Using Spectral Frequency Weighting FunctionabstractThis paper presents an approach of bug localization using a frequency weighting function. In an existing approach, only binary information of execution count from test executions is used. Information of each program statement being executed and not executed by a particular test is used; indicated by 1 and 0 respectively. In our proposed approach, frequency execution count of each program statement executed by a respective test is used. We evaluate several well-known spectra metrics using our proposed approach and the existing approach (using binary information of execution count) on two test suites; Siemens Test Suite and Unix datasets. We show that the bug localization performance is improved by using our proposed approach. We conduct statistical test and show that the improved bug localization performance using our approach (using frequency execution count) is statistically significant than using the existing approach (using binary information of execution count). Hua Jie Lee, Lee Naish, Kotagiri Ramamohanarao |
COMPSAC | 3 |
| 2010 | Contrast Pattern Mining and Its Application for Building Robust Classifiers
Kotagiri Ramamohanarao |
Discovery Science | 1 |
| 2010 | ScoreFinder: A method for collaborative quality inference on user-generated contentabstractUser-generated content is quickly becoming the greatest source of information on the World Wide Web. Shared content items are initially considered unconfirmed in the sense that their credibility has not yet been established. Conventional, centralized confirmation of credibility is infeasible at the Internet scale and so making use of the annotators themselves to evaluate each item is essential. However, users usually differ in opinions to the same item, and the existence of bias, variance and maliciousness makes the problem of aggregating opinions more difficult. Addressing this problem, we propose the use of an Author-Annotator model with an iterative algorithm, called ScoreFinder, for inferring credibility by ranking shared items. In order to reduce the influence from a variety of error sources, we identify reliable users on each topic, and adaptively aggregate scores from them. Moreover, we transform the users' input to remove errors/anomalies, by identifying patterns of misbehaviour learned from a real data set. We show how our algorithm performs on both real data sets and synthetic data sets, and a significant improvement was achieved in the experiment. Yang Liao, Aaron Harwood, Kotagiri Ramamohanarao |
ICDE | 3 |
| 2010 | Mining distribution change in stock order streamsabstractDetecting changes in stock prices is a well known problem in finance with important implications for monitoring and business intelligence. Forewarning of changes in stock price, can be made by the early detection of changes in the distributions of stock order numbers. In this paper, we address the change detection problem for streams of stock order numbers and propose a novel incremental detection algorithm. Our algorithm gains high accuracy and low delay by employing a natural Poisson distribution assumption about the nature of stock order streams. We establish that our algorithm is highly scalable and has linear complexity. We also experimentally demonstrate its effectiveness for detecting change points, via experiments using both synthetic and real-world datasets. Xindong Wu 0001, Huaiqing Wang, Rui Zhang 0003, James Bailey 0001, Kotagiri Ramamohanarao |
ICDE | 6 |
| 2010 | Combining Real and Virtual Graphs to Enhance Data ClusteringabstractFusion of multiple information sources can yield significant benefits to accomplishing certain learning tasks. This paper exploits the sparse representation of signals for the problem of data clustering. The method is built within the framework of spectral clustering algorithms, which convexly combines a real graph constructed from the given physical features with a virtual graph constructed from sparse reconstructive coefficients. The experimental results on several real-world data sets have shown that fusion of both real and virtual graphs can obtain better (or at least comparable) results than using either graph alone. Liang Wang 0001, Christopher Leckie, Kotagiri Ramamohanarao |
ICPR | 3 |
| 2010 | A Novel Scalable Multi-class ROC for Effective Visualization and Computation
Md. Rafiul Hassan, Kotagiri Ramamohanarao, Chandan K. Karmakar, M. Maruf Hossain, James Bailey 0001 |
PAKDD (1) | 2 |
| 2010 | Decentralisation of ScoreFinder: A Framework for Credibility Management on User-Generated Contents
Yang Liao, Aaron Harwood, Kotagiri Ramamohanarao |
PAKDD (2) | 3 |
| 2010 | iVAT and aVAT: Enhanced Visual Analysis for Cluster Tendency Assessment
Liang Wang 0001, Uyen T. V. Nguyen, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao |
PAKDD (1) | 5 |
| 2010 | A fast indexing approach for protein structure comparisonabstractBACKGROUND: Protein structure comparison is a fundamental task in structural biology. While the number of known protein structures has grown rapidly over the last decade, searching a large database of protein structures is still relatively slow using existing methods. There is a need for new techniques which can rapidly compare protein structures, whilst maintaining high matching accuracy. RESULTS: We have developed IR Tableau, a fast protein comparison algorithm, which leverages the tableau representation to compare protein tertiary structures. IR tableau compares tableaux using information retrieval style feature indexing techniques. Experimental analysis on the ASTRAL SCOP protein structural domain database demonstrates that IR Tableau achieves two orders of magnitude speedup over the search times of existing methods, while producing search results of comparable accuracy. CONCLUSION: We show that it is possible to obtain very significant speedups for the protein structure comparison problem, by employing an information retrieval style approach for indexing proteins. The comparison accuracy achieved is also strong, thus opening the way for large scale processing of very large protein structure databases. James Bailey 0001, Arun Siddharth Konagurthu, Kotagiri Ramamohanarao |
BMC Bioinform. | 4 |
| 2010 | Classifying proteins using gapped Markov feature pairs
Xiaonan Ji, James Bailey 0001, Kotagiri Ramamohanarao |
Neurocomputing | 3 |
| 2010 | Optimized algorithms for predictive range and KNN queries on moving objects
Rui Zhang 0003, H. V. Jagadish, Bing Tian Dai, Kotagiri Ramamohanarao |
Inf. Syst. | 4 |
| 2010 | Layered Approach Using Conditional Random Fields for Intrusion DetectionabstractIntrusion detection faces a number of challenges; an intrusion detection system must reliably detect malicious activities in a network and must perform efficiently to cope with the large amount of network traffic. In this paper, we address these two issues of Accuracy and Efficiency using Conditional Random Fields and Layered Approach. We demonstrate that high attack detection accuracy can be achieved by using Conditional Random Fields and high efficiency by implementing the Layered Approach. Experimental results on the benchmark KDD '99 intrusion data set show that our proposed system based on Layered Conditional Random Fields outperforms other well-known methods such as the decision trees and the naive Bayes. The improvement in attack detection accuracy is very high, particularly, for the U2R attacks (34.8 percent improvement) and the R2L attacks (34.5 percent improvement). Statistical Tests also demonstrate higher confidence in detection accuracy for our method. Finally, we show that our system is robust and is able to handle noisy data without compromising performance. Kapil Kumar Gupta, Baikunth Nath, Kotagiri Ramamohanarao |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2010 | Enhanced Visual Analysis for Cluster Tendency Assessment and Data PartitioningabstractVisual methods have been widely studied and used in data cluster analysis. Given a pairwise dissimilarity matrix {\schmi D} of a set of n objects, visual methods such as the VAT algorithm generally represent {\schmi D} as an n\times n image {\rm I}(\tilde{{\schmi D}}) where the objects are reordered to reveal hidden cluster structure as dark blocks along the diagonal of the image. A major limitation of such methods is their inability to highlight cluster structure when {\schmi D} contains highly complex clusters. This paper addresses this limitation by proposing a Spectral VAT algorithm, where {\schmi D} is mapped to {\schmi D}^{\prime } in a graph embedding space and then reordered to {{\tilde{\schmi D}^{\prime }}} using the VAT algorithm. A strategy for automatic determination of the number of clusters in {\rm I}({\tilde{{\schmi D}^{\prime }}}) is then proposed, as well as a visual method for cluster formation from {\rm I}({\tilde{{\schmi D}^{\prime }}}) based on the difference between diagonal blocks and off-diagonal blocks. A sampling-based extended scheme is also proposed to enable visual cluster analysis for large data sets. Extensive experimental results on several synthetic and real-world data sets validate our algorithms. Liang Wang 0001, Xin Geng 0001, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2009 | RoleVAT: Visual Assessment of Practical Need for Role Based Access ControlabstractRole based access control (RBAC) is a powerful security administration concept that can simplify permission assignment management. Migration to and maintenance of RBAC requires role engineering, the identification of a set of roles that offer administrative benefit. However, establishing that RBAC is desirable in a given enterprise is lacking in current role engineering processes. To help identify the practical need for RBAC, we propose RoleVAT, a Role engineering tool for the Visual Assessment of user and permission Tendencies. User and permission clusters can be visually identified as potential user groups or roles. The benefit and impact of this visual analysis in enterprise environments is discussed and demonstrated through testing on real life as well as synthetic datasets. Our experimental results show the effectiveness of RoleVAT as well as interesting user and role tendencies in real enterprise environments. Dana Zhang, Kotagiri Ramamohanarao, Steve Versteeg, Rui Zhang 0003 |
ACSAC | 2 |
| 2009 | Spectral Debugging with Weights and Incremental RankingabstractSoftware faults can be diagnosed using program spectra. The program spectra considered here provide information about which statements are executed in each one of a set of test cases. This information is used to compute a value for each statement which indicates how likely it is to be buggy, and the statements are ranked according to these values. We present two improvements to this method. First, we associate varying weights with failed test cases --- test cases which execute fewer statements are given more weight and have more influence on the ranking. This generally improves diagnosis accuracy, with little additional cost. Second, the ranking is computed incrementally. After the top-ranked statement is identified, the weights are adjusted in order to compute the rest of the ranking. This further improves accuracy. The cost is more significant, but not prohibitive. Lee Naish, Hua Jie Lee, Kotagiri Ramamohanarao |
APSEC | 3 |
| 2009 | Kernel latent semantic analysis using an information retrieval based kernelabstractHidden term relationships can be found within a document collection using Latent semantic analysis (LSA) and can be used to assist in information retrieval. LSA uses the inner product as its similarity function, which unfortunately introduces bias due to document length and term rarity into the term relationships. In this article, we present the novel kernel based LSA method, which uses separate document and query kernel functions to compute document and query similarities, rather than the inner product. We show that by providing an appropriate kernel function, we are able to provide a better fit of our data and hence produce more effective term relationships. Laurence Anthony F. Park, Kotagiri Ramamohanarao |
CIKM | 2 |
| 2009 | Robust traffic merging strategies for sensor-enabled cars using time geographyabstractWe present two novel merging algorithms that optimize traffic flow on highways, particularly at intersections of ramps and main roads. In our work, cars are equipped with sensors that can detect distance to neighboring cars, and communicate their velocity and acceleration readings with one another. Sensor-enabled cars can locally exchange sensed information about traffic and adapt their behavior much earlier than regular cars. However, the accuracy level of sensors is a major challenge for merging algorithms, because inaccuracies can potentially lead to unsafe merging behaviors. In this paper, we investigate how the accuracy of sensors impacts merging algorithms, and design robust merging algorithms that tolerate sensor errors. Experimental results show that our main proposed merging algorithm, which is based on concepts from time geography, is able to guarantee safe merging while tolerating four times more imprecise positioning information, and can double the road capacity and increase the traffic flow by 25%. Ziyuan Wang 0003, Lars Kulik, Kotagiri Ramamohanarao |
GIS | 3 |
| 2009 | Grouped ECOC Conditional Random Fields for Prediction of Web User Behavior
Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park |
PAKDD | 2 |
| 2009 | Approximate Spectral Clustering
Liang Wang 0001, Christopher Leckie, Kotagiri Ramamohanarao, James C. Bezdek |
PAKDD | 3 |
| 2009 | The Sensitivity of Latent Dirichlet Allocation for Information Retrieval
Laurence Anthony F. Park, Kotagiri Ramamohanarao |
ECML/PKDD (2) | 2 |
| 2009 | Feature Weighted SVMs Using Receiver Operating CharacteristicsabstractSupport Vector Machines (SVMs) are a leading tool in classification and pattern recognition and the kernel function is one of its most important components.This function is used to map the input space into a high dimensional feature space.However, it can perform rather poorly when there are too many dimensions (e.g. for gene expression data) or when there is a lot of noise.In this paper, we investigate the suitability of using a new feature weighting scheme for SVM kernel functions, based on receiver operating characteristics (ROC).This strategy is clean, simple and surprisingly effective.We experimentally demonstrate that it can significantly and substantially boost classification performance, across a range of datasets. Shaoyi Zhang, M. Maruf Hossain, Md. Rafiul Hassan, James Bailey 0001, Kotagiri Ramamohanarao |
SDM | 5 |
| 2009 | A voting approach to identify a small number of highly predictive genes using multiple classifiersabstractBACKGROUND: Microarray gene expression profiling has provided extensive datasets that can describe characteristics of cancer patients. An important challenge for this type of data is the discovery of gene sets which can be used as the basis of developing a clinical predictor for cancer. It is desirable that such gene sets be compact, give accurate predictions across many classifiers, be biologically relevant and have good biological process coverage. RESULTS: By using a new type of multiple classifier voting approach, we have identified gene sets that can predict breast cancer prognosis accurately, for a range of classification algorithms. Unlike a wrapper approach, our method is not specialised towards a single classification technique. Experimental analysis demonstrates higher prediction accuracies for our sets of genes compared to previous work in the area. Moreover, our sets of genes are generally more compact than those previously proposed. Taking a biological viewpoint, from the literature, most of the genes in our sets are known to be strongly related to cancer. CONCLUSION: We show that it is possible to obtain superior classification accuracy with our approach and obtain a compact gene set that is also biologically relevant and has good coverage of different biological processes. Md. Rafiul Hassan, M. Maruf Hossain, James Bailey 0001, Geoff MacIntyre, Joshua W. K. Ho, Kotagiri Ramamohanarao |
BMC Bioinform. | 6 |
| 2009 | Trust-based robust scheduling and runtime adaptation of scientific workflowabstractAbstract Robustness and reliability with respect to the successful completion of a schedule are crucial requirements for scheduling in scientific workflow management systems because service providers are becoming autonomous. We introduce a model to incorporate trust, which indicates the probability that a service agent will comply with its commitments to improve the predictability and stability of the schedule. To deal with exceptions during the execution of a schedule, we adapt and evolve the schedule at runtime by interleaving the processes of evaluating, scheduling, executing and monitoring in the life cycle of the workflow management. Experiments show that schedules maximizing participants' trust are more likely to survive and succeed in open and dynamic environments. The results also prove that the proposed approach of workflow evaluation can find the most robust execution flow efficiently, thus avoiding the need of scheduling every possible execution path in the workflow definition. Copyright © 2009 John Wiley & Sons, Ltd. Mingzhong Wang, Kotagiri Ramamohanarao, Jinjun Chen |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Automatically Determining the Number of Clusters in Unlabeled Data SetsabstractClustering is a popular tool for exploratory data analysis. One of the major problems in cluster analysis is the determination of the number of clusters in unlabeled data, which is a basic input for most clustering algorithms. In this paper we investigate a new method called DBE (dark block extraction) for automatically estimating the number of clusters in unlabeled data sets, which is based on an existing algorithm for visual assessment of cluster tendency (VAT) of a data set, using several common image and signal processing techniques. Basic steps include: 1) generating a VAT image of an input dissimilarity matrix; 2) performing image segmentation on the VAT image to obtain a binary image, followed by directional morphological filtering; 3) applying a distance transform to the filtered binary image and projecting the pixel values onto the main diagonal axis of the image to form a projection signal; 4) smoothing the projection signal, computing its first-order derivative, and then detecting major peaks and valleys in the resulting signal to decide the number of clusters. Our new DBE method is nearly "automatic", depending on just one easy-to-set parameter. Several numerical and real-world examples are presented to illustrate the effectiveness of DBE. Liang Wang 0001, Christopher Leckie, Kotagiri Ramamohanarao, James C. Bezdek |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2009 | An analysis of latent semantic term self-correlationabstractLatent semantic analysis (LSA) is a generalized vector space method that uses dimension reduction to generate term correlations for use during the information retrieval process. We hypothesized that even though the dimension reduction establishes correlations between terms, the dimension reduction is causing a degradation in the correlation of a term to itself (self-correlation). In this article, we have proven that there is a direct relationship to the size of the LSA dimension reduction and the LSA self-correlation. We have also shown that by altering the LSA term self-correlations we gain a substantial increase in precision, while also reducing the computation required during the information retrieval process. Laurence Anthony F. Park, Kotagiri Ramamohanarao |
ACM Trans. Inf. Syst. | 2 |
| 2009 | Efficient storage and retrieval of probabilistic latent semantic information for information retrieval
Laurence Anthony F. Park, Kotagiri Ramamohanarao |
VLDB J. | 2 |
| 2009 | Building more robust multi-agent systems using a log-based approachabstractIn an agent system, the ability to handle problems and recover from them is important in sustaining stability and providing robustness. We claim that execution logging is essential to support agent system robustness, and that agents should have archi Amy Unruh, James Bailey 0001, Kotagiri Ramamohanarao |
Web Intell. Agent Syst. | 3 |
| 2008 | Permission Set Mining: Discovering Practical and Useful RolesabstractRole based access control is an efficient and effective way to manage and govern permissions to a large number of users. However, defining a role infrastructure that accurately reflects the internal functionalities and workings of a large enterprise is a challenging task. Recent research has focused on the theoretical components of automated role identification while practical applications for identifying roles remain unsolved.This research proposes a practical data mining heuristic method that is fast, scalable and capable of identifying comprehensive roles and placing them into a hierarchy. Permission set pattern data mining can be used to identify the roles with partial orderings that cover the largest portion of user permissions within a system. We test the algorithm on real user permission assignments as well as on generated data sets. Roles identified in test sets cover up to 85% of user permissions and analysis show the roles offer significant administrative benefit. We find interesting correlations between roles and their relationships and analyse the tradeoffs between identifying roles with complete coverage to identifying roles that are most effective and offer significant administrative benefit. Dana Zhang, Kotagiri Ramamohanarao, Tim Ebringer, Trevor Yann |
ACSAC | 2 |
| 2008 | Moving shape dynamics: A signal processing perspectiveabstractThis paper provides a new perspective on human motion analysis, namely regarding human motions in video as general discrete time signals. While this seems an intuitive idea, research on human motion analysis has attracted little attention from the signal processing community. Sophisticated signal processing techniques create important opportunities for new solutions to the problem of human motion analysis. This paper investigates how the deformations of human silhouettes (or shapes) during articulated motion can be used as discriminating features to implicitly capture motion dynamics. In particular, we demonstrate the applicability of two widely used signal transform methods, namely the discrete Fourier transform (DFT) and discrete wavelet transform (DWT), for characterization and recognition of human motion sequences. Experimental results show the effectiveness of the proposed method on two state-of-the-art data sets. Liang Wang 0001, Xin Geng 0001, Christopher Leckie, Kotagiri Ramamohanarao |
CVPR | 4 |
| 2008 | Web Page Prediction Based on Conditional Random FieldsabstractWeb page prefetching is used to reduce the access latency of the Internet. However, if most prefetched Web pages are not visited by the users in their subsequent accesses, the limited network bandwidth and server resources will not be used efficiently and may worsen the access delay problem. Therefore, it is critical that we have an accurate prediction method during prefetching. Conditional Random Fields (CRFs), which are popular sequential learning models, have already been successfully used for many Natural Language Processing (NLP) tasks such as POS tagging, name entity recognition (NER) and segmentation. In this paper, we propose the use of CRFs in the field of Web page prediction. We treat the accessing sessions of previous Web users as observation sequences and label each element of these observation sequences to get the corresponding label sequences, then based on these observation and label sequences we use CRFs to train a prediction model and predict the probable subsequent Web pages for the current users. Our experimental results show that CRFs can produce higher Web page prediction accuracy effectively when compared with other popular techniques like plain Markov Chains and Hidden Markov Models (HMMs). Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park |
ECAI | 2 |
| 2008 | Continuous Intersection Joins Over Moving ObjectsabstractThe continuous intersection join query is computationally expensive yet important for various applications on moving objects. No previous study has specifically addressed this query type. We can adopt a naive algorithm or extend an existing technique (TP-Join) to process the query. However, they compute the answer for either too long or too short a time interval, which results in either a very large computation cost per object update or too frequent answer updates, respectively. This motivates us to optimize the query processing in the time dimension. In this study, we achieve this optimization by introducing the new concept of time-constrained (TC) processing. Further, TC processing enables a set of effective improvement techniques on traditional intersection join algorithms. With a thorough experimental study, we show that our algorithm outperforms the best adapted existing solution by several orders of magnitude. Rui Zhang 0003, Dan Lin 0001, Kotagiri Ramamohanarao, Elisa Bertino |
ICDE | 3 |
| 2008 | SpecVAT: Enhanced Visual Cluster AnalysisabstractGiven a pairwise dissimilarity matrix D of a set of objects, visual methods such as the VAT algorithm (for visual analysis of cluster tendency) represent (D macr )as an image (D macr ) where the objects are reordered to highlight cluster structure as dark blocks along the diagonal of the image. A major limitation of such visual methods is their inability to highlight cluster structure in 1(D macr ) when D contains clusters with highly complex structure. In this paper, we address this limitation by proposing a Spectral VAT (SpecVAT) algorithm, where D is mapped to D' in an embedding space by spectral decomposition of the Laplacian matrix, and then reordered to D' using the VAT algorithm. We also propose a strategy to automatically determine the number of clusters in (D macr '), as well as a method for cluster formation from (D macr ') based on the difference between diagonal blocks and off-diagonal blocks. We demonstrate the effectiveness of our algorithms on several synthetic and real-world data sets that are not amenable to analysis via traditional VAT. Liang Wang 0001, Xin Geng 0001, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao |
ICDM | 5 |
| 2008 | Query Expansion for the Language Modelling Framework Using the Naïve Bayes Assumption
Laurence Anthony F. Park, Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2008 | Characteristic-Based Descriptors for Motion Sequence Recognition
Liang Wang 0001, Xiaozhe Wang, Christopher Leckie, Kotagiri Ramamohanarao |
PAKDD | 4 |
| 2008 | Improving k-Nearest Neighbour Classification with Distance Functions Based on Receiver Operating Characteristics
Md. Rafiul Hassan, M. Maruf Hossain, James Bailey 0001, Kotagiri Ramamohanarao |
ECML/PKDD (1) | 4 |
| 2008 | User Session Modeling for Effective Application Intrusion Detection
Kapil Kumar Gupta, Baikunth Nath, Kotagiri Ramamohanarao |
SEC | 3 |
| 2008 | The Effect of Weighted Term Frequencies on Probabilistic Latent Semantic Term Relationships
Laurence Anthony F. Park, Kotagiri Ramamohanarao |
SPIRE | 2 |
| 2008 | Error Correcting Output Coding-Based Conditional Random Fields for Web Page PredictionabstractWeb page prefetching has been used efficiently to reduce the access latency problem of the Internet, its success mainly relies on the accuracy of Web page prediction. As powerful sequential learning models, conditional random fields (CRFs) have been used successfully to improve the Web page prediction accuracy when the total number of unique Web pages is small. However, because the training complexity of CRFs is quadratic to the number of labels, when applied to a Web site with a large number of unique pages, the training of CRFs may become very slow and even intractable. In this paper, we decrease the training time and computational resource requirements of CRFs training by integrating error correcting output coding (ECOC) method. Moreover, since the performance of ECOC-based methods crucially depends on the ECOC code matrix in use, we employ a coding method, search coding, to design the code matrix of good quality. Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park |
Web Intelligence | 2 |
| 2008 | PConPy - a Python module for generating 2D protein mapsabstractUNLABELLED: PConPy is an open-source Python module for generating protein contact maps, distance maps and hydrogen bond plots. These maps can be generated in a number of publication-quality vector and raster image formats. Contact maps can be annotated with secondary structure and hydrogen bond assignments. PConPy offers a more flexible choice of contact definition parameters than existing toolkits, most notably a greater choice of inter-residue distance metrics. PConPy can be used as a stand-alone application or imported into existing source code. A web-interface to PConPy is also available for use. AVAILABILITY: The PConPy web-interface and source code can be accessed from its website at http://www.csse.unimelb.edu.au/~hohkhkh1/pconpy/. CONTACT: [email protected] Hui Kian Ho, Michael J. Kuiper, Kotagiri Ramamohanarao |
Bioinform. | 3 |
| 2008 | Selective sampling for approximate clustering of very large data setsabstractA key challenge in pattern recognition is how to scale the computational efficiency of clustering algorithms on large data sets. The extension of non-Euclidean relational fuzzy c-means (NERF) clustering to very large (VL = unloadable) relational data is called the extended NERF (eNERF) clustering algorithm, which comprises four phases: (i) finding distinguished features that monitor progressive sampling; (ii) progressively sampling from a N × N relational matrix RN to obtain a n × n sample matrix Rn; (iii) clustering Rn with literal NERF; and (iv) extending the clusters in Rn to the remainder of the relational data. Previously published examples on several fairly small data sets suggest that eNERF is feasible for truly large data sets. However, it seems that phases (i) and (ii), i.e., finding Rn, are not very practical because the sample size n often turns out to be roughly 50% of n, and this over-sampling defeats the whole purpose of eNERF. In this paper, we examine the performance of the sampling scheme of eNERF with respect to different parameters. We propose a modified sampling scheme for use with eNERF that combines simple random sampling with (parts of) the sampling procedures used by eNERF and a related algorithm sVAT (scalable visual assessment of clustering tendency). We demonstrate that our modified sampling scheme can eliminate over-sampling of the original progressive sampling scheme, thus enabling the processing of truly VL data. Numerical experiments on a distance matrix of a set of 3,000,000 vectors drawn from a mixture of 5 bivariate normal distributions demonstrate the feasibility and effectiveness of the proposed sampling method. We also find that actually running eNERF on a data set of this size is very costly in terms of computation time. Thus, our results demonstrate that further modification of eNERF, especially the extension stage, will be needed before it is truly practical for VL data. © 2008 Wiley Periodicals, Inc. Liang Wang 0001, James C. Bezdek, Christopher Leckie, Kotagiri Ramamohanarao |
Int. J. Intell. Syst. | 4 |
| 2007 | Mining web multi-resolution community-based popularity for information retrievalabstractThe PageRank algorithm is used in Web information retrieval to calculate a single list of popularity scores for each page in the Web. These popularity scores are used to rank query results when presented to the user. By using the structure of the entire Web to calculate one score per document, we are calculating a general popularity score, not particular to any community. Therefore, the PageRank scores are more suited to general queries. In this paper, we introduce a more general form of PageRank, using Web multi-resolution community-based popularity scores, where each document obtains a popularity score dependent on a given Web community. When a query is related to a specific community, we choose the associated set of popularity scores and order the query results accordingly. Using Web-community based popularity scores, we achieved an 11% increase in precision over PageRank. Laurence Anthony F. Park, Kotagiri Ramamohanarao |
CIKM | 2 |
| 2007 | Blood Vessel Segmentation from Color Retinal Images using Unsupervised Texture ClassificationabstractAutomated blood vessel segmentation is an important issue for assessing retinal abnormalities and diagnoses of many diseases. The segmentation of vessels is complicated by huge variations in local contrast, particularly in case of the minor vessels. In this paper, we propose a new method of texture based vessel segmentation to overcome this problem. We use Gaussian and L*a*b* perceptually uniform color spaces with original RGB for texture feature extraction on retinal images. A bank of Gabor energy filters are used to analyze the texture features from which a feature vector is constructed for each pixel. The fuzzy C-means (FCM) clustering algorithm is used to classify the feature vectors into vessel or non-vessel based on the texture properties. From the FCM clustering output we attain the final output segmented image after a post processing step. We compare our method with hand-labeled ground truth segmentation of five images and achieve 84.37% sensitivity and 99.61% specificity. Alauddin Bhuiyan, Baikunth Nath, Joselíto J. Chua, Kotagiri Ramamohanarao |
ICIP (5) | 4 |
| 2007 | Multiple Self-Splitting and Merging Competitive Learning Algorithm
Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2007 | Query Expansion Using a Collection Dependent Probabilistic Latent Semantic Thesaurus
Laurence Anthony F. Park, Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2007 | Incorporating Fault Tolerance with Replication on Very Large Scale GridsabstractProviding fault tolerance for message passing parallel application on a distributed environment is a rule rather than an exception. A node failure can cause the whole computation to stop and has to be restarted from the beginning if no fault tolerance is available. However, introducing fault tolerance has some overhead on speedup that can be achieved. In this paper, we introduce a new technique called replication with cross-over packets for reliability and to increase fault tolerance over Very Large Scale Grids (VLSG). This technique has two pronged effect of avoiding single point of failure and single link of failure. We incorporate this new technique into the L-BSP model and show the possible speedup of parallel process. We also derive the achievable speedup for some fundamental parallel algorithms using this technique. Elankovan Sundararajan, Aaron Harwood, Kotagiri Ramamohanarao |
PDCAT | 3 |
| 2007 | Role engineering using graph optimisationabstractRole engineering is one of the fundamental phases for migrating existing enterprises to Role Based Access Control. In organisations with a large number of users and permissions, this task can be time consuming and costly if a top down approach is used. Existing bottom up approaches are not sufficient in producing a comprehensive set of roles for hierarchical Role Based Access Control. In this research, we propose a predominately bottom up approach that uses Graph Optimisation to identify appropriate role hierarchies. Additional partial role specifications can be incorporated to produce a hybrid approach. Using rules that reduce administration requirements, roles and their hierarchies are automatically extracted from large numbers of permission assignments. The results of the Graph Optimisation approach are hierarchical Role Based Access Control infrastructures that offer improved access control administration for the system. Dana Zhang, Kotagiri Ramamohanarao, Tim Ebringer |
SACMAT | 2 |
| 2007 | Personalized PageRank for Web Page Prediction Based on Access Time-Length and FrequencyabstractWeb page prefetching techniques are used to address the access latency problem of the Internet. To perform successful prefetching, we must be able to predict the next set of pages that will be accessed by users. The PageRank algorithm used by Google is able to compute the popularity of a set of Web pages based on their link structure. In this paper, a novel PageRank-like algorithm is proposed for conducting Web page prediction. Two biasing factors are adopted to personalize PageRank, so that it favors the pages that are more important to users. One factor is the length of time spent on visiting a page and the other is the frequency that a page was visited. The experiments conducted show that using these two factors simultaneously to bias PageRank results in more accurate Web page prediction than other methods that use only one of these two factors. Yong Zhen Guo, Kotagiri Ramamohanarao, Laurence Anthony F. Park |
Web Intelligence | 2 |
| 2007 | Information sharing for distributed intrusion detection systems
Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
J. Netw. Comput. Appl. | 3 |
| 2006 | Structure-based querying of proteins using waveletsabstractThe ability to retrieve molecules based on structural similarity has use in many applications, from disease diagnosis and treatment to drug discovery and design. In this paper, we present a method to represent protein molecules that allows for the fast, flexible and efficient retrieval of similar structures, based on either global or local attributes. We begin by computing the pair-wise distance between amino acids, transforming each 3D structure into a 2D distance matrix. We normalize this matrix to a specific size and apply a 2D wavelet decomposition to generate a set of approximation coefficients, which serves as our global feature vector. This transformation reduces the overall dimensionality of the data while still preserving spatial features and correlations. We test our method by running queries on three different protein data sets that have been used previously in the literature, basing our comparisons on labels taken from the SCOP database. We find that our method significantly outperforms existing approaches, in terms of retrieval accuracy, memory utilization and execution time. Specifically, using a k-d tree and running a 10-nearest-neighbor search on a dataset of 33,000 proteins against itself, we see an average accuracy of 89% at the SCOP SuperFamily level and a total query time that is up to 350 times faster than previously published techniques. In addition to processing queries based on global similarity, we also propose innovative extensions to effectively match proteins based solely on shared local substructures, allowing for a more flexible query interface. Keith Marsolo, Srinivasan Parthasarathy 0001, Kotagiri Ramamohanarao |
CIKM | 3 |
| 2006 | Computing Iceberg Quotient Cubes with Bounding
Xiuzhen Zhang 0001, Pauline Lin, Kotagiri Ramamohanarao |
DaWaK | 3 |
| 2006 | Attacking Confidentiality: An Agent Based Approach
Kapil Kumar Gupta, Baikunth Nath, Kotagiri Ramamohanarao, Ashraf U. Kazi |
ISI | 3 |
| 2006 | Further Improving Emerging Pattern Based Classifiers Via Bagging
Hongjian Fan, Kotagiri Ramamohanarao, Mengxu Liu |
PAKDD | 3 |
| 2006 | Evaluation of handoff algorithms using a call quality measure with signal based penaltiesabstractThis paper proposes a new call quality measure based on mobile signal strength measurements to evaluate performance of handoff algorithms in wireless cellular networks. The proposed measure allows the quantification of the impact of the handoff algorithms of performance. Using the proposed measure we compare existing handoff algorithms to identify the trade-off between signal quality and required number of handoffs. Our results indicate that a handoff method based on a threshold with 2 dB hysteresis provides better performance compared to the conventional wisdom of 3 dB hysteresis. We provide a benchmark value for handoff algorithms based on an off-line heuristic method using the new measure. Our benchmark shows that there is substantial room for improvement of the existing handoff algorithm Malka N. Halgamuge, Kotagiri Ramamohanarao, Hai Le Vu 0001, Moshe Zukerman |
WCNC | 2 |
| 2006 | Approximate clustering in very large relational dataabstractDifferent extensions of fuzzy c-means (FCM) clustering have been developed to approximate FCM clustering in very large (unloadable) image (eFFCM) and object vector (geFFCM) data. Both extensions share three phases: (1) progressive sampling of the VL data, terminated when a sample passes a statistical goodness of fit test; (2) clustering with (literal or exact) FCM; and (3) noniterative extension of the literal clusters to the remainder of the data set. This article presents a comparable method for the remaining case of interest, namely, clustering in VL relational data. We will propose and discuss each of the four phases of eNERF and our algorithm for this last case: (1) finding distinguished features that monitor progressive sampling, (2) progressively sampling a square N × N relation matrix RN until an n × n sample relation Rn passes a statistical test, (3) clustering Rn with literal non-Euclidean relational fuzzy c-means, and (4) extending the clusters in Rn to the remainder of the relational data. The extension phase in this third case is not as straightforward as it was in the image and object data cases, but our numerical examples suggest that eNERF has the same approximation qualities that eFFCM and geFFCM do. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 817–841, 2006. James C. Bezdek, Richard J. Hathaway, Jacalyn M. Huband, Christopher Leckie, Kotagiri Ramamohanarao |
Int. J. Intell. Syst. | 5 |
| 2006 | Using Emerging Patterns to Construct Weighted Decision TreesabstractDecision trees (DTs) represent one of the most important and popular solutions to the problem of classification. They have been shown to have excellent performance in the field of data mining and machine learning. However, the problem of DTs is that they are built using data instances assigned to crisp classes. In this paper, we generalize decision trees so that they can take into account weighted classes assigned to the training data instances. Moreover, we propose a novel method for discovering weights for the training instances. Our method is based on emerging patterns (EPs). EPs are those itemsets whose supports (probabilities) in one class are significantly higher than their supports (probabilities) in the other classes. Our experimental evaluation shows that the new proposed method has good performance and excellent noise tolerance. Hamad Alhammady, Kotagiri Ramamohanarao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2006 | Fast Discovery and the Generalization of Strong Jumping Emerging Patterns for Building Compact and Accurate ClassifiersabstractClassification of large data sets is an important data mining problem that has wide applications. Jumping emerging patterns (JEPs) are those itemsets whose supports increase abruptly from zero in one data set to nonzero in another data set. In this paper, we propose a fast, accurate, and less complex classifier based on a subset of JEPs, called strong jumping emerging patterns (SJEPs). The support constraint of SJEP removes potentially less useful JEPs while retaining those with high discriminating power. Previous algorithms based on the manipulation of border as well as consEPMiner cannot directly mine SJEPs. In this paper, we present a new tree-based algorithm for their efficient discovery. Experimental results show that: 1) the training of our classifier is typically 10 times faster than earlier approaches, 2) our classifier uses much fewer patterns than the JEP-classifier to achieve a similar (and, often, improved) accuracy, and 3) in many cases, it is superior to other state-of-the-art classification systems such as naive Bayes, CBA, C4.5, and bagged and boosted versions of C4.5. We argue that SJEPs are high-quality patterns which possess the most differentiating power. As a consequence, they represent sufficient information for the construction of accurate classifiers. In addition, we generalize these patterns by introducing noise-tolerant emerging patterns (NEPs) and generalized noise-tolerant emerging patterns (GNEPs). Our tree-based algorithms can be adopted to easily discover these variations. We experimentally demonstrate that SJEPs, NEPs, and GNEPs are extremely useful for building effective classifiers that can deal well with noise. Hongjian Fan, Kotagiri Ramamohanarao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2005 | Broadening Vector Space Schemes for Improving the Quality of Information Retrieval
Kotagiri Ramamohanarao, Laurence Anthony F. Park |
APWeb | 1 |
| 2005 | Exploiting Traffic Localities for Efficient Flow State Lookup
Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
NETWORKING | 3 |
| 2005 | Improved Self-splitting Competitive Learning Algorithm
Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2005 | Expanding the Training Data Space Using Emerging Patterns and Genetic MethodsabstractClassification is a major problem in machine learning. Many classifiers have been developed recently. However, the performance of these classifiers is proportional to the knowledge obtained from the training data. As a result, traditional classifiers can not perform very well when the training data space is very limited. In this paper, we propose a new approach to expand the training data space (ETDS) using emerging patterns (EPs) [4] and genetic methods (GMs) [7]. EPs are those itemsets whose supports in one class are significantly higher than their supports in the other classes. GMs are evolutionary methods that incorporate computational techniques inspired by biology [8]. We combine the power of EPs and GMs to expand the training data space before applying standard classifiers. The expansion process is performed by generating more training instances using four techniques. An extensive experimental evaluation carried out on a number of datasets shows that our approach has a great impact on the performance of many traditional classifiers. Hamad Alhammady, Kotagiri Ramamohanarao |
SDM | 2 |
| 2005 | Mining Emerging Patterns and Classification in Data StreamsabstractA data stream model has been proposed recently for those data intensive applications such as financial applications, manufacturing, and others (Babcock et al., 2002). In this model, data arrives in multiple, continuous, rapid, time-varying data streams. These characteristics make it infeasible for traditional classification and mining techniques to deal with data streams. In this paper, we propose a novel method for mining emerging patterns (EPs) in data streams. Moreover, we show how these EPs can be used to classify data streams. EPs (Dong and Li, 1999) are those itemsets whose supports in one class are significantly higher than their supports in the other classes. The experimental evaluation shows that our proposed method can achieve up to 10% increase in accuracy compared to the other methods. Hamad Alhammady, Kotagiri Ramamohanarao |
Web Intelligence | 2 |
| 2005 | A Novel Document Ranking Method Using the Discrete Cosine TransformabstractWe propose a new Spectral text retrieval method using the Discrete Cosine Transform (DCT). By taking advantage of the properties of the DCT and by employing the fast query and compression techniques found in vector space methods (VSM), we show that we can process queries as fast as VSM and achieve a much higher precision. Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2005 | Incremental maintenance of shortest distance and transitive closure in first-order logic and SQLabstractGiven a database, the view maintenance problem is concerned with the efficient computation of the new contents of a given view when updates to the database happen. We consider the view maintenance problem for the situation when the database contains a weighted graph and the view is either the transitive closure or the answer to the all-pairs shortest-distance problem ( APSD ). We give incremental algorithms for APSD , which support both edge insertions and deletions. For transitive closure, the algorithm is applicable to a more general class of graphs than those previously explored. Our algorithms use first-order queries, along with addition (+) and less-than (<) operations ( FO (+,<)); they store O ( n 2 ) number of tuples, where n is the number of vertices, and have AC 0 data complexity for integer weights. Since FO (+,<) is a sublanguage of SQL and is supported by almost all current database systems, our maintenance algorithms are more appropriate for database applications than nondatabase query types of maintenance algorithms. Chaoyi Pang, Guozhu Dong, Kotagiri Ramamohanarao |
ACM Trans. Database Syst. | 3 |
| 2005 | A novel document retrieval method using the discrete wavelet transformabstractCurrent information retrieval methods either ignore the term positions or deal with exact term positions; the former can be seen as coarse document resolution, the latter as fine document resolution. We propose a new spectral-based information retrieval method that is able to utilize many different levels of document resolution by examining the term patterns that occur in the documents. To do this, we take advantage of the multiresolution analysis properties of the wavelet transform. We show that we are able to achieve higher precision when compared to vector space and proximity retrieval methods, while producing fast query times and using a compact index. Laurence Anthony F. Park, Kotagiri Ramamohanarao, Marimuthu Palaniswami |
ACM Trans. Inf. Syst. | 2 |
| 2004 | Using Emerging Patterns and Decision Trees in Rare-Class ClassificationabstractThe problem of classifying rarely occurring cases is faced in many real life applications. The scarcity of the rare cases makes it difficult to classify them correctly using traditional classifiers. In this paper, we propose an approach to use emerging patterns (EPs) (G. Dong and J. Li, 1999) and decision trees (DTs) in rare-class classification (EPDT). EPs are those itemsets whose supports in one class are significantly higher than their supports in the other classes. EPDT employs the power of EPs to improve the quality of rare-case classification. To achieve this aim, we first introduce the idea of generating nonexisting rare-class instances, and then we over-sample the most important rare-class instances. Our experiments show that EPDT outperforms many classification methods. Hamad Alhammady, Kotagiri Ramamohanarao |
ICDM | 2 |
| 2004 | Hybrid Pre-Query Term Expansion using Latent Semantic AnalysisabstractLatent semantic retrieval methods (unlike vector space methods) take the document and query vectors and map them into a topic space to cluster related terms and documents. This produces a more precise retrieval but also a long query time. We present a new method of document retrieval which allows us to process the latent semantic information into a hybrid latent semantic-vector space query mapping. This mapping automatically expands the users query based on the latent semantic information in the document set. This expanded query is processed using a fast vector space method. Since we have the latent semantic data in a mapping, we are able to store and retrieve vector information in the same fast manner that the vector space method offers. Multiple mappings are combined to produce hybrid latent semantic retrieval which provide precision results 5% greater than the vector space method and fast query times. Laurence Anthony F. Park, Kotagiri Ramamohanarao |
ICDM | 2 |
| 2004 | Automatic extraction of semantic concepts in medical imagesabstractA novel automatic system for extracting the semantic descriptions of medical image content and concept in text form is presented. We first extract and analyse image features and the features are mapped to semantic descriptions by fuzzy functions. Based on these semantic descriptions, our system facilitates knowledge base construction using a machine learning scheme. The result will be useful for other researchers in medical image retrieval area, who can take advantage of both text-based queries and image queries. Mira Park 0001, Kotagiri Ramamohanarao |
ICIP | 2 |
| 2004 | Proactively Detecting Distributed Denial of Service Attacks Using Source IP Address Monitoring
Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
NETWORKING | 3 |
| 2004 | The Application of Emerging Patterns for Improving the Quality of Rare-Class Classification
Hamad Alhammady, Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2004 | Noise Tolerant Classification by Chi Emerging Patterns
Hongjian Fan, Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2004 | A Tree-Based Approach to the Discovery of Diagnostic Biomarkers for Ovarian Cancer
Jinyan Li 0001, Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2004 | ParaDualMiner: An Efficient Parallel Implementation of the DualMiner Algorithm
Roger Ming Hieng Ting, James Bailey 0001, Kotagiri Ramamohanarao |
PAKDD | 3 |
| 2004 | Incremental Maintenance on the Border of the Space of Emerging Patterns
Jinyan Li 0001, Thomas Manoukian, Guozhu Dong, Kotagiri Ramamohanarao |
Data Min. Knowl. Discov. | 4 |
| 2004 | DeEPs: A New Instance-Based Lazy Discovery and Classification System
Jinyan Li 0001, Guozhu Dong, Kotagiri Ramamohanarao, Limsoon Wong |
Mach. Learn. | 3 |
| 2004 | On the decidability of the termination problem of active database systems
James Bailey 0001, Guozhu Dong, Kotagiri Ramamohanarao |
Theor. Comput. Sci. | 3 |
| 2004 | Fourier Domain Scoring: A Novel Document Ranking MethodabstractCurrent document retrieval methods use a vector space similarity measure to give scores of relevance to documents when related to a specific query. The central problem with these methods is that they neglect any spatial information within the documents in question. We present a new method, called Fourier Domain Scoring (FDS), which takes advantage of this spatial information, via the Fourier transform, to give a more accurate ordering of relevance to a document set. We show that FDS gives an improvement in precision over the vector space similarity measures for the common case of Web like queries, and it gives similar results to the vector space measures for longer queries. Laurence Anthony F. Park, Kotagiri Ramamohanarao, Marimuthu Palaniswami |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2003 | Detecting Distributed Denial of Service Attacks by Sharing Distributed Beliefs
Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
ACISP | 3 |
| 2003 | Detecting reflector attacks by sharing beliefsabstractIn this paper, we present a distributed approach to detecting a type of distributed denial of service attack known as reflector attacks. In our approach, every potential reflector monitors the incoming packets and broadcasts a warning message to other potential reflectors if any abnormal traffic is observed. The warning message contains a description of the abnormal traffic it has observed. A detection decision can be made based on the information from multiple potential reflectors. We present a learning algorithm to decide when to broadcast the warning message. Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
GLOBECOM | 3 |
| 2003 | Protection from distributed denial of service attacks using history-based IP filteringabstractIn this paper, we introduce a practical scheme to defend against distributed denial of service (DDoS) attacks based on IP source address filtering. The edge router keeps a history of all the legitimate IP addresses which have previously appeared in the network. When the edge router is overloaded, this history is used to decide whether to admit an incoming Ip packet. Unlike other proposals to defend against DDoS attacks, our scheme works well during highly-distributed DDoS attacks, i.e., from a large number of sources. We present several heuristic methods to make the IP address database accurate and robust, and we present experimental results that demonstrate the effectiveness of our scheme in defending against highly-distributed DDoS attacks. Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
ICC | 3 |
| 2003 | A Fast Algorithm for Computing Hypergraph Transversals and its Application in Mining Emerging PatternsabstractComputing the minimal transversals of a hypergraph is an important problem in computer science that has significant applications in data mining. We present a new algorithm for computing hypergraph transversals and highlight their close connection to an important class of patterns known as emerging patterns. We evaluate our technique on a number of large datasets and show that it outperforms previous approaches by a factor of 9-29 times. James Bailey 0001, Thomas Manoukian, Kotagiri Ramamohanarao |
ICDM | 3 |
| 2003 | Classification Using Constrained Emerging Patterns
James Bailey 0001, Thomas Manoukian, Kotagiri Ramamohanarao |
WAIM | 3 |
| 2003 | Efficiently Mining Interesting Emerging Patterns
Hongjian Fan, Kotagiri Ramamohanarao |
WAIM | 2 |
| 2002 | A new implementation technique for fast Spectral based document retrieval systemsabstractThe traditional methods of spectral text retrieval (FDS,CDS) create an index of spatial data and convert the data to its spectral form at query time. We present a new method of implementing and querying an index containing spectral data which will conserve the high precision performance of the spectral methods, reduce the time needed to resolve the query, and maintain an acceptable size for the index. This is done by taking advantage of the properties of the discrete cosine transform and by applying ideas from vector space document ranking methods. Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao |
ICDM | 3 |
| 2002 | Learning to Share Distributed Probabilistic Beliefs
Christopher Leckie, Kotagiri Ramamohanarao |
ICML | 2 |
| 2002 | Sparse Bayesian Learning for Regression and Classification using Markov Chain Monte Carlo
Shien-Shin Tham, Arnaud Doucet, Kotagiri Ramamohanarao |
ICML | 3 |
| 2002 | Adjusted Probabilistic Packet Marking for IP Traceback
Tao Peng 0002, Christopher Leckie, Kotagiri Ramamohanarao |
NETWORKING | 3 |
| 2002 | A probabilistic approach to detecting network scansabstractThis paper presents a probabilistic approach for detecting network scans in real-time. Unlike previous approaches, our model takes into consideration both the number of destinations or ports accessed by a source, as well as how unusual these accesses are. We demonstrate the effectiveness of our approach in terms of accuracy and throughput, based on an analysis of the unusual sources that were found in real-life packet trace files. Christopher Leckie, Kotagiri Ramamohanarao |
NOMS | 2 |
| 2002 | An Efficient Single-Scan Algorithm for Mining Essential Jumping Emerging Patterns for Classification
Hongjian Fan, Kotagiri Ramamohanarao |
PAKDD | 2 |
| 2002 | Fast Algorithms for Mining Emerging Patterns
James Bailey 0001, Thomas Manoukian, Kotagiri Ramamohanarao |
PKDD | 3 |
| 2002 | Long-Term Learning for Web Search EnginesabstractThis paper considers how web search engines can learn from the successful searches recorded in their user logs. Document Transformation is a feasible approach that uses these logs to improve document representations. Existing test collections do not allow an adequate investigation of Document Transformation, but we show how a rigorous evaluation of this method can be carried out using the referer logs kept by web servers. We also describe a new strategy for Document Transformation that is suitable for long-term incremental learning. Our experiments show that Document Transformation improves retrieval performance over a medium sized collection of webpages. Commercial search engines may be able to achieve similar improvements by incorporating this approach. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Charles Kemp, Kotagiri Ramamohanarao |
PKDD | 2 |
| 2002 | A Novel Web Text Mining Method Using the Discrete Cosine Transform
Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao |
PKDD | 3 |
| 2001 | Transaction Oriented Computational Models for Multi-Agent SystemsabstractBDI (Belief, Desire, Intention) is a mature and commonly adopted architecture for intelligent agents. However, the current computational model adopted by BDI has a number of problems with concurrency control, recoverability and predictability. This has hindered the construction of agents having robust and predictable behaviour. Indeed the conceptual and practical tools needed for building dependable agent systems, resilient to faults and other unexpected situations, are still at an early stage of research. To this end, we propose to use ideas from established database technology as an adjunct to BDI agent systems. In the long term, we envisage a two layer model for distributed system development. The upper layer, called "behavioural" in our scheme, consists of concepts and techniques coming from agent research. The lower layer, called "control", is based on more traditional computer technologies and implements the actions decided by the upper layer. The emphasis of the upper behavioural layer is on flexibility of the software development process, in order to be effective in application domains that are difficult to analyse or rapidly changing. By contrast, the main focus of the lower control layer is rigorous and efficient implementation within wellknown boundaries. We discuss the development of an agent system having a computational model with well-defined correctness criteria. Instead of hardwiring robustness and fault-tolerant behaviour into agent plans, well defined notions of correctness exist at the semantic level. Verification can then be undertaken at the desired level of abstraction. Kotagiri Ramamohanarao, James Bailey 0001, Paolo Busetta |
ICTAI | 1 |
| 2001 | Combining the Strength of Pattern Frequency and Distance for Classification
Jinyan Li 0001, Kotagiri Ramamohanarao, Guozhu Dong |
PAKDD | 2 |
| 2001 | Building Behaviour Knowledge Space to Make Classification Decision
Xiuzhen Zhang 0001, Guozhu Dong, Kotagiri Ramamohanarao |
PAKDD | 3 |
| 2001 | Internet Document Filtering Using Fourier Domain Scoring
Laurence Anthony F. Park, Marimuthu Palaniswami, Kotagiri Ramamohanarao |
PKDD | 3 |
| 2001 | Making Use of the Most Expressive Jumping Emerging Patterns for Classification
Jinyan Li 0001, Guozhu Dong, Kotagiri Ramamohanarao |
Knowl. Inf. Syst. | 3 |
| 2000 | The Space of Jumping Emerging Patterns and Its Incremental Maintenance Algorithms
Jinyan Li 0001, Kotagiri Ramamohanarao, Guozhu Dong |
ICML | 2 |
| 2000 | Information-Based Classification by Aggregating Emerging Patterns
Xiuzhen Zhang 0001, Guozhu Dong, Kotagiri Ramamohanarao |
IDEAL | 3 |
| 2000 | Exploring constraints to efficiently mine emerging patterns from large high-dimensional datasetsabstractEmerging patterns (EPs) were proposed recently to capture changes or dierences betw een datasets: an EP is a multivariate feature whose support increases sharply from a background dataset to a target dataset, and the support ratio is called its gro wth rate.Interesting long EPs often have l o w support; mining suc h EPs from high-dimensional datasets is a great challenge due to the combinatorial explosion of the number of candidates.We propose a Constraint-based EP Miner, ConsEPMiner, that utilizes tw o types of constraints for eectively pruning the search space: External constrain tsare user-giv en minimums on support, growth rate, and growth-rate improvement to con ne the resulting EP set.Inheren t constrain ts | same subset support, top growth rate, and same origin | are deriv ed from the propertiesof EPs and datasets, and are solely for pruning the search space and saving computation.ConsEPMiner can eÆciently mine all EPs at low support on large highdimensional datasets, with low minimums on growth rate and growth-rate improvement.In comparison, the widely known Apriori-like approach is ineective on high-dimensional data.While ConsEPMiner adopts several ideas from Dense-Miner [4], a recent constrain t-based association rule miner, its main new contributions are the introduction of inherent constrain ts and the w ays to use them together with externalconstrain ts for eÆcien t EP mining from dense datasets.Experiments on dense data sho w that, at low support, Con-sEPMiner outperforms the Apriori-like approach b y o r d e r s of magnitude and is more than twice as fast as the Dense-Miner approach.D, supp(X), is jft2DjXtgj jDj . Xiuzhen Zhang 0001, Guozhu Dong, Kotagiri Ramamohanarao |
KDD | 3 |
| 2000 | Making Use of the Most Expressive Jumping Emerging Patterns for Classification
Jinyan Li 0001, Guozhu Dong, Kotagiri Ramamohanarao |
PAKDD | 3 |
| 2000 | Instance-Based Classification by Emerging Patterns
Jinyan Li 0001, Guozhu Dong, Kotagiri Ramamohanarao |
PKDD | 3 |
| 1999 | Incremental FO(+, <) Maintenance of All-Pairs Shortest Paths for Undirected Graphs after Insertions and Deletions
Chaoyi Pang, Kotagiri Ramamohanarao, Guozhu Dong |
ICDT | 2 |
| 1999 | Efficient Mining of High Confidience Association Rules without Support Thresholds
Jinyan Li 0001, Xiuzhen Zhang 0001, Guozhu Dong, Kotagiri Ramamohanarao |
PKDD | 4 |
| 1998 | Classifying Inheritance Mechanisms in Concurrent Object Oriented Programming
Lobel Crnogorac, Anand S. Rao, Kotagiri Ramamohanarao |
ECOOP | 3 |
| 1998 | Decidability and Undecidability Results for the Termination Problem of Active Database RulesabstractActive database systems enhance the functionality of traditional databases through the use of active rules or `triggers'. One of the principal questions for such systems is that of termination - is it possible for the rules to recursively activate one another indefinitely, given an initial triggering event. In this paper, we study the decidability of the termination problem, our aim being to delimit the boundary between the decidable and the undecidable. We present two families of rule languages, the one literal languages where each update is permitted to have just one atom in its body, and the unary languages where only unary relations may be updated, but higher arity relations may be accessed through views. Within each of these, we identify members close to the boundary of (un)decidability. Our context is similar to the while query language and the dynamics gives an interesting contrast to Datalog with negation; our results shed insights on the power of triggers as well as comparison of the termination problem to boundedness and query containment. James Bailey 0001, Guozhu Dong, Kotagiri Ramamohanarao |
PODS | 3 |
| 1998 | Efficient Recursive Aggregation and Negation in Deductive DatabasesabstractWe present an efficient evaluation technique for modularly stratified deductive database programs for which the local strata level mappings are known at compile time. We present an important subclass of these programs (called EMS-programs) in which one can easily express problems, such as shortest distance, company ownership, bill of materials, and preferential vote counting. Programs written in this style have an easy-to-understand semantics and can be efficiently computed. Another important virtue of these programs is that their modular-stratification properties are independent of the extensional database. David B. Kemp, Kotagiri Ramamohanarao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1998 | Inverted Files Versus Signature Files for Text IndexingabstractTwo well-known indexing methods are inverted files and signature files. We have undertaken a detailed comparison of these two approaches in the context of text indexing, paying particular attention to query evaluation speed and space requirements. We have examined their relative performance using both experimentation and a refined approach to modeling of signature files, and demonstrate that inverted files are distinctly superior to signature files. Not only can inverted files be used to evaluate typical queries in less time than can signature files, but inverted files require less space and provide greater functionality. Our results also show that a synthetic text database can provide a realistic indication of the behavior of an actual text database. The tools used to generate the synthetic database have been made publicly available Justin Zobel, Alistair Moffat, Kotagiri Ramamohanarao |
ACM Trans. Database Syst. | 3 |
| 1997 | Database Transactions in a Purely Declarative Logic Programming Language
David B. Kemp, Thomas C. Conway, Evan P. Harris, Fergus Henderson, Kotagiri Ramamohanarao, Zoltan Somogyi |
DASFAA | 5 |
| 1997 | Abstract Interpretation of Active Rules and its Use in Termination Analysis
James Bailey 0001, Lobel Crnogorac, Kotagiri Ramamohanarao, Harald Søndergaard |
ICDT | 3 |
| 1997 | Structural Issues in Active Rule Systems
James Bailey 0001, Guozhu Dong, Kotagiri Ramamohanarao |
ICDT | 3 |
| 1997 | Analysis of Inheritance Mechanisms in Agent-Oriented Programming
Lobel Crnogorac, Anand S. Rao, Kotagiri Ramamohanarao |
IJCAI (1) | 3 |
| 1997 | Optimal Clustering of Relations to Improve Sorting and Partitioning for JoinsabstractThe sorting or partitioning of relations is very common in relational database systems. Implementations of the join operation include the sort–merge join algorithm, which sorts both relations, and the hash join algorithm, which usually partitions both relations. We describe how clustering records using an optimal multi-attribute hash (MAH) file, taking the query pattern and distribution into account, reduces the average cost of sorting or partitioning. We demonstrate that maintaining multiple copies of a data file, each with a different clustering organization, further reduces the average cost of sorting or partitioning. We describe an inexpensive method for determining a good partitioning index (MAH file organization). Our analysis and experiments show that the partitioning indexes we find are usually optimal and can often partition a relation more than ten times faster than by not using any clustering. We also show that a significant change in the query pattern or distribution is required before a reorganization of a data file is necessary, and that such a reorganization is, in general, an inexpensive operation. Evan P. Harris, Kotagiri Ramamohanarao |
Comput. J. | 2 |
| 1996 | Join Algorithm Costs Revisited
Evan P. Harris, Kotagiri Ramamohanarao |
VLDB J. | 2 |
| 1995 | Atlas: A Nested Relational Database System for Text ApplicationsabstractAdvanced database applications require facilities such as text indexing, image storage, and the ability to store data with a complex structure. However, these facilities are not usually included in traditional database systems. In this paper we describe Atlas, a nested relational database system that has been designed for text-based applications. The Atlas query language is TQL, an SQL-like query language with text operators. The query language is supported by signature file text indexing techniques, and by a parser that can be configured for different text formats and even some foreign languages. Atlas can also be used to store images and audio.> Ron Sacks-Davis, Alan J. Kent, Kotagiri Ramamohanarao, James A. Thom, Justin Zobel |
IEEE Trans. Knowl. Data Eng. | 3 |
| 1994 | Algebraic Equivalences Among Nested Relational ExpressionsabstractAlgebraic optimization is both theoretically and practically important for query processing in (nested) relational databases. In this paper, we consider this issue and investigate some algebraic properties concerning the nested relational operators. We also outline a heuristic optimization algorithm for nested relational expressions by adopting algebraic transformation rules developed in this paper and previous related work. Hong-Cheu Liu, Kotagiri Ramamohanarao |
CIKM | 2 |
| 1994 | Subsumption-Free Bottom-up Evaluation of Logic Programs with Partially Instantiated Data Structures
Zoltan Somogyi, David B. Kemp, James Harland, Kotagiri Ramamohanarao |
EDBT | 4 |
| 1994 | An Introduction to Deductive Database Languages and Systems
Kotagiri Ramamohanarao, James Harland |
VLDB J. | 1 |
| 1994 | The Aditi Deductive Database System
Jayen Vaghani, Kotagiri Ramamohanarao, David B. Kemp, Zoltan Somogyi, Peter J. Stuckey, Tim S. Leask, James Harland |
VLDB J. | 2 |
| 1993 | Constraint Propagation for Linear Recursive Rules
James Harland, Kotagiri Ramamohanarao |
ICLP | 2 |
| 1993 | Status of the Aditi Deductive Database System
Jayen Vaghani, Kotagiri Ramamohanarao, David B. Kemp, Zoltan Somogyi, Peter J. Stuckey, Tim S. Leask, James Harland |
ICLP | 2 |
| 1991 | The Aditi Deductive Database System (Extented Abstract)
Kotagiri Ramamohanarao |
DASFAA | 1 |
| 1991 | Design Overview of the Aditi Deductive Database SystemabstractAn overview of the structure of Aditi, a disk-based deductive database system under continuous development at the University of Melbourne, is presented. The aim of the project is to find out what implementation methods and optimization techniques would make deductive databases competitive with current commercial relational databases. The structure of the Aditi prototype is based on a variant of the client-server model. The front end of Aditi interacts with the user exclusively in a logical language that has more expressive power than relational query languages. The back end uses relational technology for efficiency in the management of disk-based data and uses some optimization algorithms especially developed for the bottom-up evaluation of logical queries involving recursion. The system has been functional for almost two years now, and has already proven its worth as a research tool.> Jayen Vaghani, Kotagiri Ramamohanarao, David B. Kemp, Zoltan Somogyi, Peter J. Stuckey |
ICDE | 2 |
| 1990 | Right-, left- and multi-linear rule transformations that maintain context information
David B. Kemp, Kotagiri Ramamohanarao, Zoltan Somogyi |
VLDB | 2 |
| 1990 | A signature file scheme based on multiple organizations for indexing very large text databasesabstractA new signature file method for accessing information from large databases containing both formatted and free text data is presented. The new method, called the multiorganizational scheme is proposed for indexing very large databases containing hundreds of thousands or possibly millions of records. With this method, records are grouped into blocks and signatures are formed for each block of records. These signatures are stored in a block descriptor file using a storage device called the bit slice organization. By forming multiple block descriptor files, each based on a possibly different grouping of records into blocks, it is possible to efficiently determine record matches on query. Both computational results based on a mathematical model as well as experimental results using a library database are presented. These results show that the method provides effective access to large text databases. © 1990 John Wiley & Sons, Inc. Alan J. Kent, Ron Sacks-Davis, Kotagiri Ramamohanarao |
J. Am. Soc. Inf. Sci. | 3 |
| 1989 | Automatic Synthesis of Boolean Equations Using Programmable Array LogicabstractArticle Automatic synthesis of Boolean equations using programmable array logic Share on Authors: R. P. Goré Dept. of Computer Science, University of Melbourne, Parkville, 3052, Australia Dept. of Computer Science, University of Melbourne, Parkville, 3052, AustraliaView Profile , K. Ramaamohanarao Dept. of Computer Science, University of Melbourne, Parkville, 3052, Australia Dept. of Computer Science, University of Melbourne, Parkville, 3052, AustraliaView Profile Authors Info & Claims DAC '89: Proceedings of the 26th ACM/IEEE Design Automation ConferenceJune 1989 Pages 283–289https://doi.org/10.1145/74382.74430Online:01 June 1989Publication History 6citation208DownloadsMetricsTotal Citations6Total Downloads208Last 12 Months3Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Rajeev Goré, Kotagiri Ramamohanarao |
DAC | 2 |
| 1989 | Partial-match Retrieval using Multiple-Key Hashing with Multiple File Copies
Kotagiri Ramamohanarao, John Shepherd 0001, Ron Sacks-Davis |
DASFAA | 1 |
| 1988 | A Superimposed Coding Scheme Based on Multiple Block Descriptor Files for Indexing Very Large Data Bases
Alan J. Kent, Ron Sacks-Davis, Kotagiri Ramamohanarao |
VLDB | 3 |
| 1987 | Concurrent Database Updates in PROLOG
Lee Naish, James A. Thom, Kotagiri Ramamohanarao |
ICLP | 3 |
| 1987 | Answering Queries in Deductive Database Systems
Kotagiri Ramamohanarao, John Shepherd 0001 |
ICLP | 1 |
| 1987 | Multikey Access Methods Based on Superimposed Coding TechniquesabstractBoth single-level and two-level indexed descriptor schemes for multikey retrieval are presented and compared. The descriptors are formed using superimposed coding techniques and stored using a bit-inversion technique. A fast-batch insertion algorithm for which the cost of forming the bit-inverted file is less than one disk access per record is presented. For large data files, it is shown that the two-level implementation is generally more efficient for queries with a small number of matching records. For queries that specify two or more values, there is a potential problem with the two-level implementation in that costs may accrue when blocks of records match the query but individual records within these blocks do not. One approach to overcoming this problem is to set bits in the descriptors based on pairs of indexed terms. This approach is presented and analyzed. Ron Sacks-Davis, Alan J. Kent, Kotagiri Ramamohanarao |
ACM Trans. Database Syst. | 3 |
| 1986 | A Superimposed Codeword Indexing Scheme for Very Large Prolog Databases
Kotagiri Ramamohanarao, John Shepherd 0001 |
ICLP | 1 |
| 1986 | A Superjoin Algorithm for Deductive Databases
James A. Thom, Kotagiri Ramamohanarao, Lee Naish |
VLDB | 2 |
| 1984 | Recursive Linear HashingabstractA modification of linear hashing is proposed for which the conventional use of overflow records is avoided. Furthermore, an implementation of linear hashing is presented for which the amount of physical storage claimed is only fractionally more than the minimum required. This implementation uses a fixed amount of in-core space. Simulation results are given which indicate that even for storage utilizations approaching 95 percent, the average successful search cost for this method is close to one disk access. Kotagiri Ramamohanarao, Ron Sacks-Davis |
ACM Trans. Database Syst. | 1 |
| 1983 | A two level superimposed coding scheme for partial match retrieval
Ron Sacks-Davis, Kotagiri Ramamohanarao |
Inf. Syst. | 2 |
| 1983 | Partial-Match Retrieval Using Hashing and DescriptorsabstractThis paper studies a partial-match retrieval scheme based on hash functions and descriptors. The emphasis is placed on showing how the use of a descriptor file can improve the performance of the scheme. Records in the file are given addresses according to hash functions for each field in the record. Furthermore, each page of the file has associated with it a descriptor, which is a fixed-length bit string, determined by the records actually present in the page. Before a page is accessed to see if it contains records in the answer to a query, the descriptor for the page is checked. This check may show that no relevant records are on the page and, hence, that the page does not have to be accessed. The method is shown to have a very substantial performance advantage over pure hashing schemes, when some fields in the records have large key spaces. A mathematical model of the scheme, plus an algorithm for optimizing performance, is given. Kotagiri Ramamohanarao, John W. Lloyd |
ACM Trans. Database Syst. | 1 |
| 1982 | On Synchronizing Readers and Writers with SemaphoresabstractA weakness in the reader priority solution proposed by Curtois, Heymans and Parnas for the problem of synchronizing concurrent readers and writers is described and an improvement is explained. The difficulties of solving complex synchronizing problems by using standard semaphore primitives, as illustrated by this example, lead us to propose that special-purpose synchronization techniques should be supported by a judicious combination of hardware/microcode and software routines. We then describe an efficient solution for the reader/writer problem which is easy to understand, to implement and to use. James Leslie Keedy, John Rosenberg, Kotagiri Ramamohanarao |
Comput. J. | 3 |
| 1982 | Dynamic Hashing SchemesabstractIn this paper, we study two new dynamic hashing schemes for primary key retrieval. The schemes are related to those of Scholl, Litwin and Larson. The first scheme is simple and elegant and has certain performance advantages over earlier schemes. We give a detailed mathematical analysis of this scheme and also present simulation results. The second scheme is essentially that of Larson. However, we have made a number of changes which simplify his scheme. Kotagiri Ramamohanarao, John W. Lloyd |
Comput. J. | 1 |
| 1981 | Hardware Address Translation for Machines with a Large Virtual Memory
Kotagiri Ramamohanarao, Ron Sacks-Davis |
Inf. Process. Lett. | 1 |
| 1979 | On Implementing Semaphores with SetsabstractIt is proposed that semaphores should in some circumstances be extended by associating with each semaphore, in addition to the usual integer, a set (bit string). There are two separate uses for this set, which are considered separately and which can be implemented independently of each other. The first, the available resources set, has a bit identifying each resource controlled by the semaphore, which will indicate when the resource is free. When no resources are free the set may then be used to identify those processes waiting on a resource. The advantages and disadvantages of both sets are discussed, including the possibility of eliminating MUTEX semaphores from situations such as producer/consumer activities, and the possibility of entirely ‘automating’ (i.e. controlling by hardware semaphore instructions) the synchronisation and scheduling of processes. James Leslie Keedy, Kotagiri Ramamohanarao, John Rosenberg |
Comput. J. | 2 |