EDBT 2026 Demo / reviewers in the wild / expert
Kannan Govindarajan
dblp:15/6502
· DBLP profile ↗
20ranked-venue papers
5as first author
3since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Stochastic gradient descent-based support vector machines training optimization on Big Data and HPC frameworksabstractSummary Support vector machines (SVM) is a widely used machine learning algorithm. With the increasing amount of research data nowadays, understanding how to do efficient training is more important than ever. This article discusses the performance optimizations and benchmarks related to providing high‐performance support for SVM training. In this research, we have focused on a highly scalable gradient descent‐based approach to implementing the core SVM algorithm. In providing a scalable solution, we have designed optimized high‐performance computing and dataflow‐oriented SVM implementations. A high‐performance computing approach means the algorithm is implemented with the bulk synchronous parallel (BSP) model. In addition, we analyzed the language level optimizations and math kernel optimizations on a prominent HPC modeling programming language (C++) and dataflow modeling programming language (Java). In the experiments, we compared the performance of classic HPC models, classic dataflow models, and hybrid models designed on classic HPC and dataflow programming models. Our research illustrates a scientific approach in designing the SVM algorithm at scale in classic HPC, dataflow, and hybrid systems. Vibhatha Abeykoon, Geoffrey C. Fox, Saliya Ekanayake, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Niranda Perera, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahatin Akkas |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | Twister2 Cross-platform resource scheduler for big dataabstractAbstract Twister2 is an open‐source big data hosting environment designed to process both batch and streaming data at scale. Twister2 runs jobs in both high‐performance computing (HPC) and big data clusters. It provides a cross‐platform resource scheduler to run jobs in diverse environments. Twister2 is designed with a layered architecture to support various clusters and big data problems. In this paper, we present the cross‐platform resource scheduler of Twister2. We identify required services and explain implementation details. We present job startup delays for single jobs and multiple concurrent jobs in Kubernetes and OpenMPI clusters. We compare job startup delays for Twister2 and Spark at a Kubernetes cluster. In addition, we compare the performance of terasort algorithm on Kubernetes and bare metal clusters at AWS cloud. Ahmet Uyar, Gurhan Gunduz, Supun Kamburugamuve, Pulasthi Wickramasinghe, Chathura Widanage, Kannan Govindarajan, Niranda Perera, Vibhatha Abeykoon, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 6 |
| 2022 | High-performance iterative dataflow abstractions in Twister2: TSetabstractSummary The dataflow model is gradually becoming the de facto standard for big data applications. While many popular frameworks are built around this model, very little research has been done on understanding its inner workings, which in turn has led to inefficiencies in existing frameworks. It is important to note that understanding the relationship between dataflow and high performance computing (HPC) building blocks allows us to address and alleviate many of these fundamental inefficiencies by learning from the extensive research literature in the HPC community. In this article, we present TSets, the dataflow abstraction of Twister2, which is a big data framework designed for high‐performance dataflow and iterative computations. We discuss the dataflow model adopted by TSets and the rationale behind implementing iteration handling at the worker level. Finally, we evaluate TSets to show the performance of the framework and the importance of the worker level iteration model. Pulasthi Wickramasinghe, Niranda Perera, Supun Kamburugamuve, Kannan Govindarajan, Vibhatha Abeykoon, Chathura Widanage, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | Twister2: Design of a big data toolkitabstractSummary Data‐driven applications are essential to handle the ever‐increasing volume, velocity, and veracity of data generated by sources such as the Web and Internet of Things (IoT) devices. Simultaneously, an event‐driven computational paradigm is emerging as the core of modern systems designed for database queries, data analytics, and on‐demand applications. Modern big data processing runtimes and asynchronous many task (AMT) systems from high performance computing (HPC) community have adopted dataflow event‐driven model. The services are increasingly moving to an event‐driven model in the form of Function as a Service (FaaS) to compose services. An event‐driven runtime designed for data processing consists of well‐understood components such as communication, scheduling, and fault tolerance. Different design choices adopted by these components determine the type of applications a system can support efficiently. We find that modern systems are limited to specific sets of applications because they have been designed with fixed choices that cannot be changed easily. In this paper, we present a loosely coupled component‐based design of a big data toolkit where each component can have different implementations to support various applications. Such a polymorphic design would allow services and data analytics to be integrated seamlessly and expand from edge to cloud to HPC environments. Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Vibhatha Abeykoon, Geoffrey C. Fox |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Streaming Machine Learning Algorithms with Big Data SystemsabstractDesigning low latency applications that can process large volumes data with higher efficiency is a challenging problem. With the limited time to process data, usage of online algorithms are becoming important in the big-data applications. Stream processing is a well-known area that has been studied for a long time. In this research, our objective is to use state of the art big-data analytic engines to implement online algorithms and compare the strengths and weaknesses in each system. We use a streaming version of Support Vector Machines (SVM) and KMeans to do the analysis. Apache Flink, Apache Storm and Twister2 streaming frameworks are used to implement these algorithms. Our study focuses on the efficiency of online training of these algorithms and the results show higher performance in Twister2 framework for these algorithms. Vibhatha Abeykoon, Gregor von Laszewski, Supun Kamburugamuve, Kannan Govindarajan, Pulasthi Wickramasinghe, Chathura Widanage, Niranda Perera, Ahmet Uyar, Gurhan Gunduz, Selahattin Akkas |
IEEE BigData | 4 |
| 2019 | Operating Enterprise AI as a Service
Fabio Casati, Kannan Govindarajan, Baskar Jayaraman, Aniruddha Thakur, Sriram Palapudi, Firat Karakusoglu, Debu Chatterjee |
ICSOC | 2 |
| 2018 | Twister: Net - Communication Library for Big Data Processing in HPC and Cloud EnvironmentsabstractStreaming processing and batch data processing are the dominant forms of big data analytics today, with numerous systems such as Hadoop, Spark, and Heron designed to process the ever-increasing explosion of data. Generally, these systems are developed as single projects with aspects such as communication, task management, and data management integrated together. By contrast, we take a component-based approach to big data by developing the essential features of a big data system as independent components with polymorphic implementations to support different requirements. Consequently, we recognize the requirements of both dataflow used in popular Apache Systems and the Bulk Synchronous Processing communication style common in High-Performance Computing (HPC) for different applications. Message Passing Interface (MPI) implementations are dominant in HPC but there are no such standard libraries available for big data. Twister:Net is a stand-alone, highly optimized dataflow style parallel communication library which can be used by big data systems or advanced users. Twister:Net can work both in cloud environments using TCP or HPC environments using MPI implementations. This paper introduces Twister:Net and compares it with existing systems to highlight its design and performance. Supun Kamburugamuve, Pulasthi Wickramasinghe, Kannan Govindarajan, Ahmet Uyar, Gurhan Gunduz, Vibhatha Abeykoon, Geoffrey C. Fox |
IEEE CLOUD | 3 |
| 2016 | Swarm Intelligence (SI) based profiling and scheduling of big data applicationsabstractPersonalization targets a user's software, hardware, and QoS requirements at any given moment in the cloud environment for the big data applications. However, the individualization aims to target the daily needs of an individual user in a dynamic manner. The proposed research work aims to design a system which will be able to optimize user's applications towards a specified target goal. Furthermore, it is integrated with a Particle Swarm Optimization (PSO) based application profiling and resource selection mechanism which comes from the family of Swarm Intelligence (SI). The proposed algorithms create an application profile template and preferred resource list for each submitted big data applications and select the cloud resources from the preferred resource list which is based on the application preferences and availability of cloud resources in an optimal manner. From the experimental results, it is evident that the proposed research work maximizes the application success ratio, scheduling success rate, utilization of cloud resources, and user satisfaction. Thamarai Selvi Somasundaram, Kannan Govindarajan, Vive Kumar |
IEEE BigData | 2 |
| 2015 | Parallel Particle Swarm Optimization (PPSO) clustering for learning analyticsabstractAnalytics is all about insights. Learning-oriented insights are the targets for Learning Analytics researchers. Insights could be detected, analysed, or created in the context of variables such as the quality of interactions with the content, study habits, engagement, competence growth, sentiments, learning efficiency, and instructional effectiveness. Clustering techniques offer an effective solution for grouping learners using observed patterns. For instance, learners could be clustered based on the effectiveness of learners' self-regulation initiatives in reaching the target learning outcomes. Each learner could belong to a number of clusters that target different types of insights. One could also analyse the distance between clusters as a means to guide learners towards better performance. Further, one could analyse the effectiveness of cohesive peer groups within and among clusters. Traditional clustering techniques only cope with numerical or categorical data and are not readily applicable in offering learning analytics solutions. In addressing this gap, this research aims to design a Parallel Particle Swarm Optimization (PPSO) algorithm for the purposes of learning analytics, where the arrival of data is continuous, the types of data is both structured and unstructured, and the volume of data can be significantly large. The research will also describe the application of the PPSO algorithm to detect, analyse, and generate learning-oriented insights. Kannan Govindarajan, David Boulanger, Vive Kumar, Kinshuk |
IEEE BigData | 1 |
| 2015 | Performance Analysis of Parallel Particle Swarm Optimization Based Clustering of StudentsabstractWhile accurate computational models that embody learning efficiency remain a distant and elusive goal, big data learning analytics approaches this goal by recognizing competency growth of learners, at various levels of granularity, using a combination of continuous, formative, and summative assessments. Our earlier research employed the conventional Particle Swarm Optimization (PSO) based clustering mechanism to cluster large numbers of learners based on their observed study habits and the consequent growth of subject knowledge competencies. This paper describes a Parallel Particle Swarm Optimization (PPSO) based clustering mechanism to cluster learners. Using a simulation study, performance measures of quality of clusters such as the Inter Cluster Distance, the Intra Cluster Distance, the processing time and the acceleration values are estimated and compared. Kannan Govindarajan, David Boulanger, Jeremie Seanosky, Jason Bell, Colin Pinnell, Vive Kumar, Kinshuk, Thamarai Selvi Somasundaram |
ICALT | 1 |
| 2014 | CLOUDRB: A framework for scheduling and managing High-Performance Computing (HPC) applications in science cloud
Thamarai Selvi Somasundaram, Kannan Govindarajan |
Future Gener. Comput. Syst. | 2 |
| 2014 | DDoS defense system for web services in a cloud environment
Thomas Vissers, Thamarai Selvi Somasundaram, Luc Pieters, Kannan Govindarajan, Peter Hellinckx |
Future Gener. Comput. Syst. | 4 |
| 2014 | Semantic-enabled CARE Resource Broker (SeCRB) for managing grid and cloud environment
Thamarai Selvi Somasundaram, Kannan Govindarajan, Usha Kiruthika, Rajkumar Buyya |
J. Supercomput. | 2 |
| 2013 | Particle Swarm Optimization (PSO)-Based Clustering for Improving the Quality of Learning using Cloud ComputingabstractVirtual Learning is a key enabler for giving equal opportunity to all throughout the globe. However, the pedagogical approach preferred by a group of learners may differ from another set of learners. By providing different pedagogical approaches through virtual learning, it is possible to satisfy the need of the learners, thereby improving the quality of learning. To identify the preference or choice of the pedagogy, the behavior of the learners is captured and analyzed. According to the understanding capability, the appropriate pedagogy is adopted for that learner. The conventional Learning Management System (LMS) plays a major role for achieving effective teaching and learning process. However, the conventional LMS fails to address the effective teaching and learning process by not providing the contents based on individual user's ability. The proposed work mainly intends to capture the data from students, analyze and cluster the data based on their individual performances in terms of accuracy, efficiency and quality. The clustering process is carried out by employing the population-based metaheuristic algorithm of Particle Swarm Optimization (PSO). The simulation process is carried out by generating the data. The generated data is based on the real data collected from engineering undergraduate students. The proposed PSO-based clustering is compared with existing K-means algorithm for analyze the performance of inter cluster and intra cluster distances. Finally, the processed data is effectively stored in the Cloud resources using Hadoop Distributed File System (HDFS). Kannan Govindarajan, Thamarai Selvi Somasundaram, Vive Kumar, Kinshuk |
ICALT | 1 |
| 2010 | CARE Resource Broker: A framework for scheduling and supporting virtual resource management
Thamarai Selvi Somasundaram, Balachandar R. Amarnath, Rangasamy Kumar, Ponnuram Balakrishnan, Kandan Rajendar, R. Rajiv, Kannan Govindarajan, Gnanapragasam Rajesh Britto, E. Mahendran, B. Madusudhanan |
Future Gener. Comput. Syst. | 7 |
| 2006 | Enabling Outsourced Service Providers to Think Globally While Acting Locally
Kevin Wilkinson, Harumi A. Kuno, Kannan Govindarajan, Kei Yuasa, Kevin Smathers, Jyotirmaya Nanda, Umeshwar Dayal |
EDBT | 3 |
| 1998 | Preference Logic Grammars
Bharat Jayaraman, Kannan Govindarajan, Surya Mantha |
Comput. Lang. | 2 |
| 1996 | Optimization and Relaxation in Constraint Logic LanguagesabstractOptimization and relaxation are two important operations that naturally arise in many applications involving constraints, e.g., engineering design, scheduling, decision support, etc. In optimization, we are interested in finding the optimal (i.e., best) solutions to a set of constraints with respect to an objective function. In many applications, optimal solutions may be difficult or impossible to obtain, and hence we are interested in finding suboptimal solutions, by either relaxing the constraints or relaxing the objective function. The contribution of this paper lies in providing a logical framework for performing optimization and relaxation in a constraint logic programming language. Our proposed framework is called preference logic programming (PLP), and its use for optimization was discussed in [8]. Essentially, in PLP we can designate certain predicates as optimization predicates, and we can specify the objective function by stating preference criteria for determining the optimal solutions to these predicates. This paper extends the PLP paradigm with facilities to formulate relaxation problems in a natural manner. We introduce the concept of a relaxation goal, and discuss its use for preference relaxation. Our model-theoretic semantics of relaxation is based on simple concepts from modal logic: Essentially, each world in the possible-worlds semantics for a preference logic program is a model for the constraints of the program, and an ordering over these worlds is determined by the objective function. Optimization can then be expressed as truth in strongly optimal worlds, while relaxation becomes truth in suitably-defined suboptimal worlds. We also present an operational semantics for relaxation as well as correctness results. Our conclusion is that the concept of preference provides a unifying framework for formulating optimization as well as relaxation problems. Kannan Govindarajan, Bharat Jayaraman, Surya Mantha |
POPL | 1 |
| 1996 | Comparing SIMD and MIMD Programming Modes
Ravikanth Ganesan, Kannan Govindarajan, Min-You Wu |
J. Parallel Distributed Comput. | 2 |
| 1995 | Preference Logic Programming
Kannan Govindarajan, Bharat Jayaraman, Surya Mantha |
ICLP | 1 |