EDBT 2026 Demo / reviewers in the wild / expert
Nikos Chrisochoides
dblp:41/5222 · also Nikos P. Chrisochoides
· DBLP profile ↗
44ranked-venue papers
6as first author
2since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 6 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Human-computer interaction and ubiquitous computing · 5Applied, interdisciplinary, general and emerging computing · 4Theory of computation · 3Artificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Parallel and multicore computing · 66% High-performance computing · 32% Electronic design automation · 1% | |
| Computer graphics and multimedia
2 papers |
Geometric modeling and processing · 92% Visualization and visual analytics · 8% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › parallel computing › parallel scientific computing
parallel mesh generation |
0.6 | 3 | 2018 | A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018 Scalable 3D hybrid parallel Delaunay image-to-mesh conversion algorithm for distributed shared memory architectures · Comput. Aided Des. 2017 Guaranteed: quality parallel delaunay refinement for restricted polyhedral domains · SCG 2002 |
Geometric modeling and processing
mesh generation |
0.6 | 2 | 2018 | A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018 23rd International Meshing Roundtable - Mesh modeling for simulations and visualization · Comput. Aided Des. 2016 |
Geometric modeling and processing › mesh generation
delaunay triangulation |
0.3 | 1 | 2018 | A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018 |
High-performance computing
distributed memory systems |
0.3 | 1 | 2018 | A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018 |
Parallel and multicore computing › load balancing
dynamic load balancing |
0.1 | 2 | 2004 | A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004 An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel Applications · SC 2003 |
Parallel and multicore computing › parallel programming models
shared-memory parallelization |
0.1 | 1 | 2017 | Scalable 3D hybrid parallel Delaunay image-to-mesh conversion algorithm for distributed shared memory architectures · Comput. Aided Des. 2017 |
Visualization and visual analytics › scientific visualization › geometric visualization
mesh visualization |
0.1 | 1 | 2016 | 23rd International Meshing Roundtable - Mesh modeling for simulations and visualization · Comput. Aided Des. 2016 |
Medical and health informatics › image-guided intervention
image-guided neurosurgery |
0.1 | 1 | 2006 | Imaging and visual analysis - Toward real-time image guided neurosurgery using distributed and grid computing · SC 2006 |
Parallel and multicore computing
parallel programming runtimes |
0.1 | 2 | 2004 | A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004 An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel Applications · SC 2003 |
Parallel and multicore computing
load balancing |
0.0 | 1 | 2004 | A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004 |
High-performance computing › scientific computing systems
adaptive mesh refinement |
0.0 | 1 | 2003 | An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel Applications · SC 2003 |
Parallel and multicore computing › parallel computing › parallel applications
irregular applications |
0.0 | 1 | 2003 | An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel Applications · SC 2003 |
High-performance computing
scientific computing systems |
0.0 | 1 | 2003 | An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel Applications · SC 2003 |
Computational geometry › mesh generation
delaunay refinement |
0.0 | 1 | 2002 | Guaranteed: quality parallel delaunay refinement for restricted polyhedral domains · SCG 2002 |
Computational geometry
mesh generation |
0.0 | 1 | 2002 | Guaranteed: quality parallel delaunay refinement for restricted polyhedral domains · SCG 2002 |
Parallel and multicore computing › parallel architecture
distributed-memory parallel computing |
0.0 | 1 | 2004 | A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004 |
Electronic design automation
mesh generation |
0.0 | 1 | 2004 | A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004 |
Distributed systems › distributed object systems
object migration |
0.0 | 1 | 2003 | An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel Applications · SC 2003 |
Methods — techniques the papers use, named apart from their topics
parallel computing · 0.7delaunay refinement · 0.7hybrid parallel algorithm · 0.3parallel implementation · 0.1grid computing · 0.1stop-and-repartition load balancing · 0.0runtime software system · 0.0object migration · 0.0global name space · 0.0active messages · 0.0distributed memory parallelism · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Toward runtime support for unstructured and dynamic exascale-era applications
Polykarpos Thomadakis, Nikos Chrisochoides |
J. Supercomput. | 2 |
| 2022 | Tasking framework for adaptive speculative parallel mesh generation
Christos Tsolakis, Polykarpos Thomadakis, Nikos Chrisochoides |
J. Supercomput. | 3 |
| 2018 | A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters
Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides |
Comput. Aided Des. | 3 |
| 2017 | Scalable 3D hybrid parallel Delaunay image-to-mesh conversion algorithm for distributed shared memory architectures
Daming Feng, Christos Tsolakis, Andrey N. Chernikov, Nikos Chrisochoides |
Comput. Aided Des. | 4 |
| 2016 | 23rd International Meshing Roundtable - Mesh modeling for simulations and visualization
Matthew L. Staten, Per-Olof Persson, Nikos Chrisochoides, Franck Ledoux, David Martineau, Katherine Lewis, Scott A. Canann, Jean-François Remacle, Kathy Loeppky |
Comput. Aided Des. | 3 |
| 2016 | Two-level locality-aware parallel Delaunay image-to-mesh conversion
Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides |
Parallel Comput. | 3 |
| 2015 | Evolutionary soft co-clustering: formulations, algorithms, and applications
Wenlu Zhang, Rongjian Li, Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides, Christopher Osgood, Shuiwang Ji |
Data Min. Knowl. Discov. | 5 |
| 2014 | Challenges and perspectives in an undergraduate flipped classroom experience: Looking through the lens of learning analyticsabstractRecent technical and infrastructural developments posit flipped classroom approaches ripe for exploration. Flipped classroom approaches have students use technology to access the lecture and other instructional resources outside the classroom in order to engage them in active learning during in-class time. Scholars and educators have reported a variety of positive outcomes of a flipped (or inverted) approach to instruction. Although, flipped classroom practices have been used in a number of education studies, the detailed framework and data obtained from students' interaction with the technology materials are typically not described. In this paper, we present a flipped classroom framework and the first captured results of such data. The framework incorporates basic e-learning tools and traditional learning practices, making it accessible to anyone wanting to implement a flipped classroom experience in his/her course. The framework is structured on open-source and easy-to-use tools, allowing for the incorporation of any additional specificities of a course. This work-in-progress can provide insights for other scholars and practitioners to further validate, examine, and extend the proposed approach. This approach can be used for those interested in incorporating flipped classroom in their teaching, since it is a flexible procedure that may be adapted to meet their needs. Michail N. Giannakos, Nikos Chrisochoides |
FIE | 2 |
| 2014 | Collecting and making sense of video learning analyticsabstractTeachers have employed online video as an element of their instructional media portfolio, alongside with books, slides, notes, etc. In comparison to other instructional media, online video affords more opportunities for recording of student navigation on a video lecture. Video analytics might provide insights into student learning performance and inform the improvement of teaching tactics. Nevertheless, those analytics are not accessible to learning stakeholders, such as researchers and educators, mainly because online video platforms do not share broadly the interactions of the users with their systems. As a remedy, we have designed an open-access video analytics system and employed it in a video-assisted course. In this paper, we present a longitudinal study, which provides valuable insights through the lens of the collected video analytics. In particular, we collected and analyzed students' video navigation, learning performance, and attitudes, and we provide the lessons learned for further development and refinement of video-assisted courses and practices. Michail N. Giannakos, Konstantinos Chorianopoulos, Nikos Chrisochoides |
FIE | 3 |
| 2014 | Examining and mapping CS teachers' technological, pedagogical and content knowledge (TPACK) in K-12 schoolsabstractComputer Science (CS) teachers' training and profile is crucial to ensure students have access to quality Computer Science Education (CSE). The aim of this study is to examine the profile of CS teachers in Greece and map it using the technique of persona. This study examines a national sample of 1127 CS teachers who teach algorithms and programming in upper secondary education. The building of the persona is based on teachers' abilities and needs regarding the central aspects of their knowledge in respect to three key domains in the Technological, Pedagogical, and Content Knowledge (TPACK) framework. According to the results in the TPACK subscales, teachers' state that their Content Knowledge scales is sufficient and Pedagogical, Content Knowledge needs to be improved on. In addition, teachers feel that they need further training in how to incorporate technology in their teaching as well as how to teach algorithms; which are two areas that relate to Pedagogical Content Knowledge and TPACK By mapping the knowledge, abilities and needs of CS teachers, we will be able to recognize the challenges they face during teaching and consider strategies and policies for addressing these challenges. Michail N. Giannakos, Spyros Doukakis, Helen Crompton, Nikos Chrisochoides, Nikos Adamopoulos, Panagiota Giannopoulou |
FIE | 4 |
| 2014 | Open Service for Video Learning AnalyticsabstractVideo learning analytics are not open to education stakeholders, such as researchers and teachers, because online video platforms do not share the interactions of the users with their systems. Nevertheless, video learning analytics are necessary to all researchers and teachers that need to understand and improve the effectiveness of the video lecture pedagogy. In this paper, we present an open video learning analytics service, which is freely accessible online. The video learning analytics service (named Social Skip) facilitates the analysis of video learning behavior by capturing learners' interactions with the video player (e.g., seek/scrub, play, pause). The service empowers any researcher or teacher to create a custom video-based experiment by selecting: 1) a video lecture from You Tube, 2) quiz questions from Google Drive, and 3) custom video player buttons. The open video analytics system has been validated through dozens of user studies, which produced thousands of video interactions. In this study, we present an indicative example, which highlights the usability and usefulness of the system. In addition to interaction frequencies, the system models the captured data as a learner activity time series. Further research should consider user modeling and personalization in order to dynamically respond to the interactivity of students with video lectures. Konstantinos Chorianopoulos, Michail N. Giannakos, Nikos Chrisochoides, Scott Reed 0002 |
ICALT | 3 |
| 2014 | Open system for video learning analyticsabstractVideo lectures are nowadays widely used by growing numbers of learners all over the world. Nevertheless, learners' interactions with the videos are not readily available, because online video platforms do not share them. In this paper, we present an open-source video learning analytics system, which is also available as a free service to researchers. Our system facilitates the analysis of video learning behavior by capturing learners' interactions with the video player (e.g, seek/scrub, play, pause). In an empirical user study, we captured hundreds of user interactions with the video player by analyzing the interactions as a learner activity time series. We found that learners employed the replaying activity to retrieve the video segments that contained the answers to the survey questions. The above findings indicate the potential of video analytics to represent learner behavior. Further research, should be able to elaborate on learner behavior by collecting large-scale data. In this way, the producers of online video pedagogy will be able to understand the use of this emerging medium and proceed with the appropriate amendments to the current video-based learning systems and practices. Konstantinos Chorianopoulos, Michail N. Giannakos, Nikos Chrisochoides |
L@S | 3 |
| 2014 | Guaranteed quality tetrahedral Delaunay meshing for medical images
Panagiotis A. Foteinos, Andrey N. Chernikov, Nikos Chrisochoides |
Comput. Geom. | 3 |
| 2014 | High quality real-time Image-to-Mesh conversion for finite element simulations
Panagiotis A. Foteinos, Nikos Chrisochoides |
J. Parallel Distributed Comput. | 2 |
| 2013 | High quality real-time image-to-mesh conversion for finite element simulationsabstractIn this paper, we present a parallel Image-to-Mesh Conversion (I2M) algorithm with quality and fidelity guarantees achieved by dynamic point insertions and removals. Starting directly from an image, it is able to recover the isosurface and mesh the volume with tetrahedra of good shape. Our tightly-coupled shared-memory parallel speculative execution paradigm employs carefully designed contention managers, load balancing, synchronization and optimizations schemes which boost the parallel efficiency with little overhead: our single-threaded performance is faster than CGAL, the state of the art sequential mesh generation software we are aware of. The effectiveness of our method is shown on Blacklight, the Pittsburgh Supercomputing Center's cache-coherent NUMA machine, via a series of case studies justifying our choices. We observe a more than 82% strong scaling efficiency for up to 64 cores, and a more than 95% weak scaling efficiency for up to 144 cores, reaching a rate of 14.7 Million Elements per second. To the best of our knowledge, this is the fastest and most scalable 3D Delaunay refinement algorithm. Panagiotis A. Foteinos, Nikos Chrisochoides |
ICS | 2 |
| 2013 | Multi-layered unstructured mesh generationabstractFinite Element Mesh Generation is a critical component for many (bio-)engineering and science applications. In this project we will develop a novel framework for guaranteed quality mesh generation for 3D and 4D Finite Element (FE) analysis, able to scale to thousands of cores. Panagiotis A. Foteinos, Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides |
ICS | 4 |
| 2013 | How students estimate the effects of ICT and programming coursesabstractThe curricula for Computer Science Education (CSE) of many countries comprise both Programming and Information and Communication Technology (ICT); however these two areas have substantial differences, inter alia the attitudes and beliefs of the students regarding the intended learning content. In this study, variables from the Unified Theory of Acceptance and Use of Technology and Social Cognitive Theory were chosen as important factors in students' behavior and attitude towards CSE. This hybrid framework aims to measure the level of the selected key variables on CSE and identify potential differences among ICT and Programming courses. Responses from the total of 126 Greek students, (71 attending ICT courses and 55 attending Programming Courses) were used to measure the variables and to identify the differences between ICT and Programming students. The results revealed several differences in the measured variables. The overall outcomes are expected to contribute to the understanding of students' likelihood to pursue computing related careers and promote the acceptance of CSE. Michail N. Giannakos, Peter Hubwieser, Nikos Chrisochoides |
SIGCSE | 3 |
| 2013 | A mesh generation and machine learning framework for Drosophila gene expression pattern image analysisabstractBACKGROUND: Multicellular organisms consist of cells of many different types that are established during development. Each type of cell is characterized by the unique combination of expressed gene products as a result of spatiotemporal gene regulation. Currently, a fundamental challenge in regulatory biology is to elucidate the gene expression controls that generate the complex body plans during development. Recent advances in high-throughput biotechnologies have generated spatiotemporal expression patterns for thousands of genes in the model organism fruit fly Drosophila melanogaster. Existing qualitative methods enhanced by a quantitative analysis based on computational tools we present in this paper would provide promising ways for addressing key scientific questions. RESULTS: We develop a set of computational methods and open source tools for identifying co-expressed embryonic domains and the associated genes simultaneously. To map the expression patterns of many genes into the same coordinate space and account for the embryonic shape variations, we develop a mesh generation method to deform a meshed generic ellipse to each individual embryo. We then develop a co-clustering formulation to cluster the genes and the mesh elements, thereby identifying co-expressed embryonic domains and the associated genes simultaneously. Experimental results indicate that the gene and mesh co-clusters can be correlated to key developmental events during the stages of embryogenesis we study. The open source software tool has been made available at http://compbio.cs.odu.edu/fly/. CONCLUSIONS: Our mesh generation and machine learning methods and tools improve upon the flexibility, ease-of-use and accuracy of existing methods. Wenlu Zhang, Daming Feng, Rongjian Li, Andrey N. Chernikov, Nikos Chrisochoides, Christopher Osgood, Charlotte Konikoff, Stuart J. Newfeld, Sudhir Kumar 0001, Shuiwang Ji |
BMC Bioinform. | 5 |
| 2011 | The Evaluation of an Effective Out-of-Core Run-Time System in the Context of Parallel Mesh GenerationabstractWe present an out-of-core run-time system that supports effective parallel computation of large irregular and adaptive problems, in particular unstructured mesh generation (PUMG). PUMG is a highly challenging application due to intensive memory accesses, unpredictable communication patterns, and variable and irregular data dependencies reflecting the unstructured spatial connectivity of mesh elements. Our runtime system allows to transform the footprint of parallel applications from wide and shallow into narrow and deep by extending the memory utilization to the out-of-core level. It simplifies and streamlines the development of otherwise highly time consuming out-of-core applications as well as the converting of existing applications. It utilizes disk, network and memory hierarchy to achieve high utilization of computing resources without sacrificing performance with PUMG. The runtime system combines different programming paradigms: multi-threading within the nodes using industrial strength software framework, one-sided active messages among the nodes, and an out-of-core subsystem for managing large datasets. We performed an evaluation on traditional parallel platforms to stress test all layers of the run-time system using three different PUMG methods with significantly varying communication and synchronization patterns. We demonstrated high overlap in computation, communication, and disk I/O which results in good performance when computing large out-of-core problems. The runtime system adds very small overhead (up to 18% on most configurations) when computing in-core which means performance is not compromised. Andriy Kot, Andrey N. Chernikov, Nikos Chrisochoides |
IPDPS | 3 |
| 2009 | Real-Time Non-rigid Registration of Medical Images on a Cooperative Parallel ArchitectureabstractUnacceptable execution time of Non-rigid registration (NRR) often presents a major obstacle to its routine clinical use. Parallel computing is an effective way to accelerate NRR. However, development of efficient parallel NRR codes is a very challenging task. One desirable approach is to map the existing sequential algorithm to the parallel architecture to gain speedup instead of designing a new parallel algorithm. Multicores and GPU provide us a cooperative architecture, in which both Single Instruction Multiple Data (SIMD) and Single Program Multiple Data (SPMD) programming models can co-exist and complement each other. We present a method to parallelize a NRR on this cooperative architecture. Our approach is first to separate the sequential algorithm into regular and irregular parts. We then map the regular part on GPU following SIMD paradigm and irregular part on multicores in a SPMD fashion. Unlike the approaches that use multicores or GPU alone, our approach leads to desirable speedup for the whole application by taking advantage of all components of the cooperative parallel architecture, for all individual parts of the application. This helps us to get closer to our goal: cheaper and faster NRR that leads to its more widespread use. The results on clinical brain MRI data show that the GPU-based Block Matching (regular part) can run at least 1.9 times faster than on a typical cluster of workstations with eight high-performance nodes. The multicores-based implementation of the incremental finite element solver (irregular part) achieves speedup of up to 7 times compared to its sequential version. As a result, the total run time of the NRR code can be reduced to less than 1 minute therefore satisfying the real time requirement for its clinical application. Yixun Liu, Andriy Fedorov, Ron Kikinis, Nikos Chrisochoides |
BIBM | 4 |
| 2009 | Modeling class cohesion as mixtures of latent topicsabstractThe paper proposes a new measure for the cohesion of classes in object-oriented software systems. It is based on the analysis of latent topics embedded in comments and identifiers in source code. The measure, named as maximal weighted entropy, utilizes the latent Dirichlet allocation technique and information entropy measures to quantitatively evaluate the cohesion of classes in software. This paper presents the principles and the technology that stand behind the proposed measure. Two case studies on a large open source software system are presented. They compare the new measure with an extensive set of existing metrics and use them to construct models that predict software faults. The case studies indicate that the novel measure captures different aspects of class cohesion compared to the existing cohesion measures and improves fault prediction for most metrics, which are combined with maximal weighted entropy. Yixun Liu, Denys Poshyvanyk, Rudolf Ferenc, Tibor Gyimóthy, Nikos Chrisochoides |
ICSM | 5 |
| 2009 | A multigrain Delaunay mesh generation method for multicore SMT-based architectures
Christos D. Antonopoulos, Filip Blagojevic, Andrey N. Chernikov, Nikos Chrisochoides, Dimitrios S. Nikolopoulos |
J. Parallel Distributed Comput. | 4 |
| 2009 | Algorithm, software, and hardware optimizations for Delaunay mesh generation on simultaneous multithreaded architectures
Christos D. Antonopoulos, Filip Blagojevic, Andrey N. Chernikov, Nikos Chrisochoides, Dimitrios S. Nikolopoulos |
J. Parallel Distributed Comput. | 4 |
| 2008 | Three-dimensional delaunay refinement for multi-core processorsabstractWe develop the first ever fully functional three-dimensional guaranteed quality parallel graded Delaunay mesh generator. First, we prove a criterion and a sufficient condition of Delaunay-independence of Steiner points in three dimensions. Based on these results, we decompose the iteration space of the sequential Delaunay refinement algorithm by selecting independent subsets from the set of the candidate Steiner points without resorting to rollbacks. We use an octree which overlaps the mesh for a coarse-grained decomposition of the set of candidate Steiner points based on their location. We partition the worklist containing poor quality tetrahedra into independent lists associated with specific separated leaves of the octree. Finally, we describe an example parallel implementation using a publicly available state-of-the art sequential Delaunay library (Tetgen). This work provides a case study for the design of abstractions and parallel frameworks for the use of complex labor intensive sequential codes on multicore architectures. Andrey N. Chernikov, Nikos Chrisochoides |
ICS | 2 |
| 2008 | Toward improved tumor targeting for image guided neurosurgery with intra-operative parametric search using distributed and grid computingabstractWe describe a high-performance distributed software environment for real-time nonrigid registration during Image- Guided Neurosurgery (IGNS). The implementation allows to perform volumetric non-rigid registration of MRI data within two minutes of computation time, and can enable large-scale parametric studies of the registration algorithm. We explore some of the parameters of the non-rigid registration, and evaluate their impact on registration accuracy using ground truth. The results of the evaluation motivate running the registration process within the Grid environment, particularly when searching for optimal parameters of intraoperative registration. Based on the results of our evaluation, distributed parametric searching of optimal registration settings can significantly improve the registration accuracy in some regions of the brain, as compared to the accuracy achieved using the default parameters. Andriy Fedorov, Nikos Chrisochoides |
IPDPS | 2 |
| 2008 | Algorithm 872: Parallel 2D constrained Delaunay mesh generationabstractDelaunay refinement is a widely used method for the construction of guaranteed quality triangular and tetrahedral meshes. We present an algorithm and a software for the parallel constrained Delaunay mesh generation in two dimensions. Our approach is based on the decomposition of the original mesh generation problem into N smaller subproblems which are meshed in parallel. The parallel algorithm is asynchronous with small messages which can be aggregated and exhibits low communication costs. On a heterogeneous cluster of more than 100 processors our implementation can generate over one billion triangles in less than 3 minutes, while the single-node performance is comparable to that of the fastest to our knowledge sequential guaranteed quality Delaunay meshing library (the Triangle). Andrey N. Chernikov, Nikos Chrisochoides |
ACM Trans. Math. Softw. | 2 |
| 2008 | Algorithm 870: A static geometric Medial Axis domain decomposition in 2D Euclidean spaceabstractWe present a geometric domain decomposition method and its implementation, which produces good domain decompositions in terms of three basic criteria: (1) The boundary of the subdomains create good angles, that is, angles no smaller than a given tolerance Φ 0 , where the value of Φ 0 is determined by the application which will use the domain decomposition. (2) The size of the separator should be relatively small compared to the area of the subdomains. (3) The maximum area of the subdomains should be close to the average subdomain area. The domain decomposition method uses an approximation of a Medial Axis as an auxiliary structure for constructing the boundary of the subdomains (separators). The N -way decomposition is based on the “divide and conquer” algorithmic paradigm and on a smoothing procedure that eliminates the creation of any new artifacts in the subdomains. This approach produces well shaped uniform and graded domain decompositions, which are suitable for parallel mesh generation. Leonidas Linardakis, Nikos Chrisochoides |
ACM Trans. Math. Softw. | 2 |
| 2007 | Evaluation of Remote Memory Access Communication on the Cray XT3abstractThis paper evaluates remote memory access (RMA) communication capabilities and performance on the Cray XT3. We discuss properties of the network hardware and portals networking software layer and corresponding implementation issues for SHMEM and ARMCI portable RMA interfaces. The performance of these interfaces is studied and compared to MPI performance. Vinod Tipparaju, Andriy Kot, Jarek Nieplocha, Monika ten Bruggencate, Nikos Chrisochoides |
IPDPS | 5 |
| 2006 | Effective out-of-core parallel Delaunay mesh refinement using off-the-shelf softwareabstractWe present two cost-effective and high-performance out-of-core parallel mesh generation algorithms and their implementation on cluster of workstations (CoWs). The total wall-clock time including wait-in-queue delays for the out-of-core methods on a small cluster (16 processors) is three times shorter than the total wall-clock time for the in-core generation of the same size mesh (about a billion elements) using 121 processors. Our best out-of-core method, for mesh sizes that fit completely in the core of the CoWs, is about 5% slower than its in-core parallel counterpart method. This is a modest performance penalty for savings of many hours in response time. Both the in-core and out-of-core methods use the best publicly available off-the-shelf sequential in-core Delaunay mesh generator Andriy Kot, Andrey N. Chernikov, Nikos Chrisochoides |
IPDPS | 3 |
| 2006 | Imaging and visual analysis - Toward real-time image guided neurosurgery using distributed and grid computingabstractNeurosurgical resection is a therapeutic intervention in the treatment of brain tumors. Precision of the resection can be improved by utilizing Magnetic Resonance Imaging (MRI) as an aid in decision making during Image Guided Neurosurgery (IGNS). Image registration adjusts pre-operative data according to intra-operative tissue deformation. Some of the approaches increase the registration accuracy by tracking image landmarks through the whole brain volume. High computational cost used to render these techniques inappropriate for clinical applications. In this paper we present a parallel implementation of a state of the art registration method, and a number of needed incremental improvements. Overall, we reduced the response time for registration of an average dataset from about an hour and for some cases more than an hour to less than seven minutes, which is within the time constraints imposed by neurosurgeons. For the first time in clinical practice we demonstrated, that with the help of distributed computing non-rigid MRI registration based on volume tracking can be computed intra-operatively. Nikos Chrisochoides, Andriy Fedorov, Andriy Kot, Neculai Archip, Peter M. Black, Olivier Clatz, Alexandra J. Golby, Ron Kikinis, Simon K. Warfield |
SC | 1 |
| 2005 | Multigrain parallel Delaunay Mesh generation: challenges and opportunities for multithreaded architecturesabstractGiven the importance of parallel mesh generation in large-scale scientific applications and the proliferation of multilevel SMT-based architectures, it is imperative to obtain insight on the interaction between meshing algorithms and these systems. We focus on Parallel Constrained Delaunay Mesh (PCDM) generation. We exploit coarse-grain parallelism at the subdomain level and fine-grain at the element level. This multigrain data parallel approach targets clusters built from low-end, commercially available SMTs. Our experimental evaluation shows that current SMTs are not capable of executing fine-grain parallelism in PCDM. However, experiments on a simulated SMT indicate that with modest hardware support it is possible to exploit fine-grain parallelism opportunities. The exploitation of fine-grain parallelism results to higher performance than a pure MPI implementation and closes the gap between the performance of PCDM and the state-of-the-art sequential mesher on a single physical processor. Our findings extend to other adaptive and irregular multigrain, parallel algorithms. Christos D. Antonopoulos, Xiaoning Ding, Andrey N. Chernikov, Filip Blagojevic, Dimitrios S. Nikolopoulos, Nikos Chrisochoides |
ICS | 6 |
| 2004 | Location management in object-based distributed computingabstractDistributed systems which use nonstationary communicating objects have to address the problem of managing locations of these objects. We study location management (LM) and its impact on application performance in the context of dynamic load-balancing for parallel distributed computing. We summarize our experience with LM in distributed object-based systems by comparing six location management policies (LMPs). The LMPs are studied within the PREMA framework used mainly for parallel mesh generation and refinement applications. Our experimental study is using a synthetic tunable microbenchmark and two mesh generation applications. We explain why the commonly adopted in practice Jump Update LMP is efficient for most of our applications and we compare it with the other location management techniques. We show how the performance of a particular LMP can be affected by the application properties and its data layout. Finally, we identify the conditions under which certain LMPs are more beneficial than jump update. Andriy Fedorov, Nikos Chrisochoides |
CLUSTER | 2 |
| 2004 | Practical and efficient point insertion scheduling method for parallel guaranteed quality delaunay refinementabstractWe describe a parallel scheduler, for guaranteed quality parallel mesh generation and refinement methods. We prove a sufficient condition for the new points to be independent, which permits the concurrent insertion of more than two points without destroying the conformity and Delaunay properties of the mesh. The scheduling technique we present is much more efficient than existing coloring methods and thus it is suitable for practical use. The condition for concurrent point insertion is based on the comparison of the distance between the candidate points against the upper bound on triangle circumradius in the mesh. Our experimental data show that the scheduler introduces a small overhead (in the order of 1--2% of the total execution time) it requires local and structured communication compared to irregular, variable and unpredictable communication of the other existing practical parallel guaranteed quality mesh generation and refinement method. Finally, on a cluster of more than 100 workstations using a simple (block) decomposition our data show that we can generate about 900 million elements in less than 300 seconds. Andrey N. Chernikov, Nikos Chrisochoides |
ICS | 2 |
| 2004 | Guaranteed-quality parallel Delaunay refinement for restricted polyhedral domains
Démian Nave, Nikos Chrisochoides, L. Paul Chew |
Comput. Geom. | 2 |
| 2004 | A Load Balancing Framework for Adaptive and Asynchronous ApplicationsabstractWe describe the design of a flexible load balancing framework and runtime software system for supporting the development of adaptive applications on distributed-memory parallel computers. The runtime system supports a global namespace, transparent object migration, automatic message forwarding and routing, and automatic load balancing. These features can be used at the discretion of the application developer in order to simplify program development and to eliminate complex bookkeeping associated with mobile data objects. An evaluation of this system in the context of a three-dimensional tetrahedral advancing front parallel mesh generator shows that overall runtime improvements of 15 percent compared to common stop-and-repartition load balancing methods, 30 percent compared to explicit intrusive load balancing methods, and 42 percent compared to no load balancing are possible on large processor configurations. At the same time, the overheads attributable to the runtime system are a fraction of 1 percent of the total runtime. The parallel advancing front method is a coarse-grained and highly adaptive application and therefore exercises all of the features of the runtime system. Kevin J. Barker, Andrey N. Chernikov, Nikos Chrisochoides, Keshav Pingali |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2003 | An Evaluation of a Framework for the Dynamic Load Balancing of Highly Adaptive and Irregular Parallel ApplicationsabstractWe present an evaluation of a flexible framework and runtime software system for the dynamic load balancing of asynchronous and highly adaptive and irregular applications. These applications, which include parallel unstructured and adaptive mesh refinement, serve as building blocks for a large class of scientific applications. Extensive study has lead to the development of solutions to the dynamic load balancing problem for loosely synchronous and computation intensive programs; however, these methods are not suitable for asynchronous and highly adaptive applications. We evaluate a new software framework which includes support for an Active Messages style communication mechanism, global name space, transparent object migration, and preemptive decision making. Our results from both a 3-dimensional parallel advancing front mesh generation program, as well as a synthetic microbenchmark, indicate that this new framework out-performs two existing general-purpose, well-known, and widely used software systems for the dynamic load balancing of adpative and irregular parallel applications. Kevin J. Barker, Nikos Chrisochoides |
SC | 2 |
| 2002 | Guaranteed: quality parallel delaunay refinement for restricted polyhedral domainsabstractWe describe a distributed memory parallel Delaunay refinement algorithm for polyhedral domains which can generate meshes containing tetrahedra with circumradius to shortest edge ratio less than 2, as long as the angle separating any two incident segments and/or facets is between 90° and 270° degrees. Input to our implementation is an element--wise partitioned, conforming Delaunay mesh of a restricted polyhedral domain which has been distributed to the processors of a parallel system. The submeshes of the distributed mesh are then independently refined by concurrently inserting new mesh vertices.Our algorithm allows a new mesh vertex to affect both the submesh tetrahedralizations and the submesh interfaces induced by the partitioning. This flexibility is crucial to ensure mesh quality, but it introduces unpredictable and variable latencies due to long delays in gathering remote data required for updating mesh data structures. In our experiments, more than 80% of this latency was masked with computation due to the fine--grained concurrency of our algorithm.Our experiments also show that the algorithm is efficient in practice, even for certain domains whose boundaries do not conform to the theoretical limits imposed by the algorithm. The algorithm we describe is the first step in the development of much more sophisticated guaranteed--quality parallel mesh generation algorithms. Démian Nave, Nikos Chrisochoides, L. Paul Chew |
SCG | 2 |
| 2002 | Date movement and control substrate for parallel adaptive applicationsabstractAbstract In this paper, we present the Data Movement and Control Substrate (DMCS), a library which implements low‐latency one‐sided communication primitives for use in parallel adaptive and irregular applications. DMCS is built on top of low‐level, vendor‐specific communication subsystems such as LAPI (Low‐level Application Programme Interface) for IBM SP machines, as well as on widely available message‐passing libraries like MPI for clusters of workstations and PCs. DMCS adds a small overhead to the communication operations provided by the lower communication system. In return, DMCS provides a flexible and easy to understand application program interface for one‐sided communication operations. Furthermore, DMCS is designed so that it can be easily ported and maintained by non‐experts. Copyright © 2002 John Wiley & Sons, Ltd. Kevin J. Barker, Nikos Chrisochoides, Jeffrey Dobbelaere, Démian Nave, Keshav Pingali |
Concurr. Comput. Pract. Exp. | 2 |
| 1997 | Compiler and Run-Time Support for Semi-Structured ApplicationsabstractAdaptive mesh refinement (AMR) is a very important scientific application. Several libraries implementing specific distribution policies have been written for AMR. In this paper, we present a "fully general block distribution " which subsumes these distributions, and discuss compiler and run-time tools for supporting these distributions efficiently in the context of a restructuring compiler. We also present performance numbers which suggest that in comparison with library code written for a particular distribution policy, the overhead arising from the generality of our approach is small. 1 Introduction Semi-structured methods such as adaptive mesh refinement and multigrid are used in applications which are computationally intensive. It is difficult to implement these methods efficiently even on a sequential machine; parallelism adds an order of magnitude overhead to the complexity. The computation in semi-structured methods is characterized by irregularly organized regular computatio... Nikos Chrisochoides, Induprakas Kodukula, Keshav Pingali |
International Conference on Supercomputing | 1 |
| 1997 | A comparison of optimization heuristics for the data mapping problemabstractIn the paper we compare the performance of six heuristics with suboptimal solutions for the data distribution of two dimensional meshes that are used for the numerical solution of partial differential equations (PDEs) on multicomputers. The data mapping heuristics are evaluated with respect to seven criteria covering load balancing, interprocessor communication, flexibility and ease of use for a class of single-phase iterative PDE solvers. Our evaluation suggests that the simple and fast block distribution heuristic can be as effective as the other five complex and computational expensive algorithms. © 1997 by John Wiley & Sons, Ltd. Nikos Chrisochoides, Nashat Mansour, Geoffrey C. Fox |
Concurr. Pract. Exp. | 1 |
| 1994 | Mapping Algorithms and Software Environment for Data Parallel PDE Iterative SolversabstractWe consider computations associated with data parallel iterative solvers used for the numerical solution of partial differential equations (PDEs). The mapping of such computations into load balanced tasks requiring minimum synchronization and communication is a difficult combinatorial optimization problem. Its optimal solution is essential for the efficient parallel processing of PDE computations. Determining data mappings that optimize a number of criteria, like workload balance, synchronization, and local communication, often involves the solution of an NP-Complete problem. Although data mapping algorithms have been known for a few years, there is lack of qualitative and quantitative comparisons based on the actual performance of the parallel computation. In this paper we present two new data mapping algorithms and evaluate them together with a large number of existing ones using the actual performance of data parallel iterative PDE solvers on the nCUBE II. Comparisons on the performance of data parallel iterative PDE solvers on medium and large scale problems demonstrate that some computationally inexpensive data block partitioning algorithms are as effective as the computationally expensive deterministic optimization algorithms. Also, these comparisons demonstrate that the existing approach in solving the data partitioning problem is inefficient for large scale problems. Finally, a software environment for the solution of the partitioning problem of data parallel iterative solvers is presented. Nikos Chrisochoides, Elias N. Houstis, John R. Rice |
J. Parallel Distributed Comput. | 1 |
| 1991 | Geometry based mapping strategies for PDE computationsabstractIn this paper we formulate new mapping strategies for partial differential equations Nikos Chrisochoides, Elias N. Houstis, Catherine E. Houstis |
ICS | 1 |
| 1990 | PELLPACK: a numerical simulation programming environment for parallel MIMD machines
Elias N. Houstis, John R. Rice, Nikos Chrisochoides, H. C. Karathanasis, P. N. Papachiou, M. K. Samartzis, Emmanuel A. Vavalis, Ko-Yang Wang, Sanjiva Weerawarana |
ICS | 3 |
| 1989 | Automatic load balanced paritioning strategies for PDE computationsabstractIn this paper we study the partitioning and allocation of computations associated with the numerical solution of partial differential equations (PDEs). Strategies for the mapping of such computations to parallel MIMD architectures can be applied to different levels of the solution process. We introduce and study heuristic approaches defined on the associated geometric data structures (meshes). Specifically, we study methods for decomposing finite element and finite difference meshes into balanced, nonoverlapping subdomains which guarantee minimum communication and synchronization among the underlying associated subcomputations. Two types of algorithms are considered: clustering techniques based on sequential orderings of the discrete geometric data and optimization based techniques involving geometric or graphical metric criteria. These algorithms support the automatic mode of a geometry decomposition tool developed in the parallel ELLPACK environment which is implemented under X11-window systems. A brief description of this tool is presented. Nikos Chrisochoides, Catherine E. Houstis, Elias N. Houstis, S. K. Kortesis, John R. Rice |
ICS | 1 |