Andrey N. Chernikov

dblp:37/854 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 71% High-performance computing · 27% Electronic design automation · 1%
Computer graphics and multimedia
1 paper
Geometric modeling and processing · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › parallel computing › parallel scientific computing
parallel mesh generation
0.622018
A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018
Scalable 3D hybrid parallel Delaunay image-to-mesh conversion algorithm for distributed shared memory architectures · Comput. Aided Des. 2017
Geometric modeling and processing › mesh generation
delaunay triangulation
0.312018
A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018
Geometric modeling and processing
mesh generation
0.312018
A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018
High-performance computing
distributed memory systems
0.312018
A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters · Comput. Aided Des. 2018
Parallel and multicore computing › parallel programming models
shared-memory parallelization
0.112017
Scalable 3D hybrid parallel Delaunay image-to-mesh conversion algorithm for distributed shared memory architectures · Comput. Aided Des. 2017
Parallel and multicore computing › load balancing
dynamic load balancing
0.012004
A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004
Parallel and multicore computing
load balancing
0.012004
A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004
Parallel and multicore computing
parallel programming runtimes
0.012004
A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004
Parallel and multicore computing › parallel architecture
distributed-memory parallel computing
0.012004
A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004
Electronic design automation
mesh generation
0.012004
A Load Balancing Framework for Adaptive and Asynchronous Applications · IEEE Trans. Parallel Distributed Syst. 2004

Methods — techniques the papers use, named apart from their topics

parallel computing · 0.7delaunay refinement · 0.7hybrid parallel algorithm · 0.3stop-and-repartition load balancing · 0.0runtime software system · 0.0
YearPublicationVenuePosition
2019 Algorithm 995: An Efficient Parallel Anisotropic Delaunay Mesh Generator for Two-Dimensional Finite Element Analysis
abstract
A bottom-up approach to parallel anisotropic mesh generation is presented by building a mesh generator starting from the basic operations of vertex insertion and Delaunay triangles. Applications focusing on high-lift design or dynamic stall, or numerical methods and modeling test cases, still focus on two-dimensional domains. This automated parallel mesh generation approach can generate high-fidelity unstructured meshes with anisotropic boundary layers for use in the computational fluid dynamics field. The anisotropy requirement adds a level of complexity to a parallel meshing algorithm by making computation depend on the local alignment of elements, which in turn is dictated by geometric boundaries and the density functions— one-dimensional spacing functions generated from an exponential distribution. This approach yields computational savings in mesh generation and flow solution through well-shaped anisotropic triangles instead of isotropic triangles. The validity of the meshes is shown through solution characteristic comparisons to verified reference solutions. A 79% parallel weak scaling efficiency on 1,024 distributed memory nodes, and a 72% parallel efficiency over the fastest sequential isotropic mesh generator on 512 distributed memory nodes, is shown through numerical experiments.
Juliette Pardue, Andrey N. Chernikov
ACM Trans. Math. Softw.2
2018 A hybrid parallel Delaunay image-to-mesh conversion algorithm scalable on distributed-memory clusters
Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides
Comput. Aided Des.2
2017 Scalable 3D hybrid parallel Delaunay image-to-mesh conversion algorithm for distributed shared memory architectures
Daming Feng, Christos Tsolakis, Andrey N. Chernikov, Nikos Chrisochoides
Comput. Aided Des.3
2016 Parallel Two-Dimensional Unstructured Anisotropic Delaunay Mesh Generation of Complex Domains for Aerospace Applications
abstract
In this paper, we present a bottom-up approach to parallel anisotropic mesh generation by building a mesh generator from principles. Applications focusing on high-lift design or dynamic stall, or numerical methods and modeling test cases still focus on the two-dimensions. Our push-button parallel mesh generation approach can generate high-fidelity unstructured meshes with anisotropic boundary layers for use in the computational fluid dynamics field. The anisotropy requirement adds a level of complexity to a parallel meshing algorithm by making computation depend on the local alignment of elements, which in turn is dictated by geometric boundaries and the density functions. Our experimental results show 70% parallel efficiency over the fastest sequential isotropic mesh generator on 256 distributed memory nodes.
Juliette Pardue, Andrey N. Chernikov
ICPP2
2016 Two-level locality-aware parallel Delaunay image-to-mesh conversion
Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides
Parallel Comput.2
2015 Curvilinear Triangular Discretization of Biomedical Images
Andrey N. Chernikov
ISBRA2
2015 Evolutionary soft co-clustering: formulations, algorithms, and applications
Wenlu Zhang, Rongjian Li, Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides, Christopher Osgood, Shuiwang Ji
Data Min. Knowl. Discov.4
2014 Guaranteed quality tetrahedral Delaunay meshing for medical images
Panagiotis A. Foteinos, Andrey N. Chernikov, Nikos Chrisochoides
Comput. Geom.2
2013 Multi-layered unstructured mesh generation
abstract
Finite Element Mesh Generation is a critical component for many (bio-)engineering and science applications. In this project we will develop a novel framework for guaranteed quality mesh generation for 3D and 4D Finite Element (FE) analysis, able to scale to thousands of cores.
Panagiotis A. Foteinos, Daming Feng, Andrey N. Chernikov, Nikos Chrisochoides
ICS3
2013 A mesh generation and machine learning framework for Drosophila gene expression pattern image analysis
abstract
BACKGROUND: Multicellular organisms consist of cells of many different types that are established during development. Each type of cell is characterized by the unique combination of expressed gene products as a result of spatiotemporal gene regulation. Currently, a fundamental challenge in regulatory biology is to elucidate the gene expression controls that generate the complex body plans during development. Recent advances in high-throughput biotechnologies have generated spatiotemporal expression patterns for thousands of genes in the model organism fruit fly Drosophila melanogaster. Existing qualitative methods enhanced by a quantitative analysis based on computational tools we present in this paper would provide promising ways for addressing key scientific questions. RESULTS: We develop a set of computational methods and open source tools for identifying co-expressed embryonic domains and the associated genes simultaneously. To map the expression patterns of many genes into the same coordinate space and account for the embryonic shape variations, we develop a mesh generation method to deform a meshed generic ellipse to each individual embryo. We then develop a co-clustering formulation to cluster the genes and the mesh elements, thereby identifying co-expressed embryonic domains and the associated genes simultaneously. Experimental results indicate that the gene and mesh co-clusters can be correlated to key developmental events during the stages of embryogenesis we study. The open source software tool has been made available at http://compbio.cs.odu.edu/fly/. CONCLUSIONS: Our mesh generation and machine learning methods and tools improve upon the flexibility, ease-of-use and accuracy of existing methods.
Wenlu Zhang, Daming Feng, Rongjian Li, Andrey N. Chernikov, Nikos Chrisochoides, Christopher Osgood, Charlotte Konikoff, Stuart J. Newfeld, Sudhir Kumar 0001, Shuiwang Ji
BMC Bioinform.4
2011 The Evaluation of an Effective Out-of-Core Run-Time System in the Context of Parallel Mesh Generation
abstract
We present an out-of-core run-time system that supports effective parallel computation of large irregular and adaptive problems, in particular unstructured mesh generation (PUMG). PUMG is a highly challenging application due to intensive memory accesses, unpredictable communication patterns, and variable and irregular data dependencies reflecting the unstructured spatial connectivity of mesh elements. Our runtime system allows to transform the footprint of parallel applications from wide and shallow into narrow and deep by extending the memory utilization to the out-of-core level. It simplifies and streamlines the development of otherwise highly time consuming out-of-core applications as well as the converting of existing applications. It utilizes disk, network and memory hierarchy to achieve high utilization of computing resources without sacrificing performance with PUMG. The runtime system combines different programming paradigms: multi-threading within the nodes using industrial strength software framework, one-sided active messages among the nodes, and an out-of-core subsystem for managing large datasets. We performed an evaluation on traditional parallel platforms to stress test all layers of the run-time system using three different PUMG methods with significantly varying communication and synchronization patterns. We demonstrated high overlap in computation, communication, and disk I/O which results in good performance when computing large out-of-core problems. The runtime system adds very small overhead (up to 18% on most configurations) when computing in-core which means performance is not compromised.
Andriy Kot, Andrey N. Chernikov, Nikos Chrisochoides
IPDPS2
2009 A multigrain Delaunay mesh generation method for multicore SMT-based architectures
Christos D. Antonopoulos, Filip Blagojevic, Andrey N. Chernikov, Nikos Chrisochoides, Dimitrios S. Nikolopoulos
J. Parallel Distributed Comput.3
2009 Algorithm, software, and hardware optimizations for Delaunay mesh generation on simultaneous multithreaded architectures
Christos D. Antonopoulos, Filip Blagojevic, Andrey N. Chernikov, Nikos Chrisochoides, Dimitrios S. Nikolopoulos
J. Parallel Distributed Comput.3
2008 Three-dimensional delaunay refinement for multi-core processors
abstract
We develop the first ever fully functional three-dimensional guaranteed quality parallel graded Delaunay mesh generator. First, we prove a criterion and a sufficient condition of Delaunay-independence of Steiner points in three dimensions. Based on these results, we decompose the iteration space of the sequential Delaunay refinement algorithm by selecting independent subsets from the set of the candidate Steiner points without resorting to rollbacks. We use an octree which overlaps the mesh for a coarse-grained decomposition of the set of candidate Steiner points based on their location. We partition the worklist containing poor quality tetrahedra into independent lists associated with specific separated leaves of the octree. Finally, we describe an example parallel implementation using a publicly available state-of-the art sequential Delaunay library (Tetgen). This work provides a case study for the design of abstractions and parallel frameworks for the use of complex labor intensive sequential codes on multicore architectures.
Andrey N. Chernikov, Nikos Chrisochoides
ICS1
2008 Algorithm 872: Parallel 2D constrained Delaunay mesh generation
abstract
Delaunay refinement is a widely used method for the construction of guaranteed quality triangular and tetrahedral meshes. We present an algorithm and a software for the parallel constrained Delaunay mesh generation in two dimensions. Our approach is based on the decomposition of the original mesh generation problem into N smaller subproblems which are meshed in parallel. The parallel algorithm is asynchronous with small messages which can be aggregated and exhibits low communication costs. On a heterogeneous cluster of more than 100 processors our implementation can generate over one billion triangles in less than 3 minutes, while the single-node performance is comparable to that of the fastest to our knowledge sequential guaranteed quality Delaunay meshing library (the Triangle).
Andrey N. Chernikov, Nikos Chrisochoides
ACM Trans. Math. Softw.1
2006 Effective out-of-core parallel Delaunay mesh refinement using off-the-shelf software
abstract
We present two cost-effective and high-performance out-of-core parallel mesh generation algorithms and their implementation on cluster of workstations (CoWs). The total wall-clock time including wait-in-queue delays for the out-of-core methods on a small cluster (16 processors) is three times shorter than the total wall-clock time for the in-core generation of the same size mesh (about a billion elements) using 121 processors. Our best out-of-core method, for mesh sizes that fit completely in the core of the CoWs, is about 5% slower than its in-core parallel counterpart method. This is a modest performance penalty for savings of many hours in response time. Both the in-core and out-of-core methods use the best publicly available off-the-shelf sequential in-core Delaunay mesh generator
Andriy Kot, Andrey N. Chernikov, Nikos Chrisochoides
IPDPS2
2005 Multigrain parallel Delaunay Mesh generation: challenges and opportunities for multithreaded architectures
abstract
Given the importance of parallel mesh generation in large-scale scientific applications and the proliferation of multilevel SMT-based architectures, it is imperative to obtain insight on the interaction between meshing algorithms and these systems. We focus on Parallel Constrained Delaunay Mesh (PCDM) generation. We exploit coarse-grain parallelism at the subdomain level and fine-grain at the element level. This multigrain data parallel approach targets clusters built from low-end, commercially available SMTs. Our experimental evaluation shows that current SMTs are not capable of executing fine-grain parallelism in PCDM. However, experiments on a simulated SMT indicate that with modest hardware support it is possible to exploit fine-grain parallelism opportunities. The exploitation of fine-grain parallelism results to higher performance than a pure MPI implementation and closes the gap between the performance of PCDM and the state-of-the-art sequential mesher on a single physical processor. Our findings extend to other adaptive and irregular multigrain, parallel algorithms.
Christos D. Antonopoulos, Xiaoning Ding, Andrey N. Chernikov, Filip Blagojevic, Dimitrios S. Nikolopoulos, Nikos Chrisochoides
ICS3
2004 Practical and efficient point insertion scheduling method for parallel guaranteed quality delaunay refinement
abstract
We describe a parallel scheduler, for guaranteed quality parallel mesh generation and refinement methods. We prove a sufficient condition for the new points to be independent, which permits the concurrent insertion of more than two points without destroying the conformity and Delaunay properties of the mesh. The scheduling technique we present is much more efficient than existing coloring methods and thus it is suitable for practical use. The condition for concurrent point insertion is based on the comparison of the distance between the candidate points against the upper bound on triangle circumradius in the mesh. Our experimental data show that the scheduler introduces a small overhead (in the order of 1--2% of the total execution time) it requires local and structured communication compared to irregular, variable and unpredictable communication of the other existing practical parallel guaranteed quality mesh generation and refinement method. Finally, on a cluster of more than 100 workstations using a simple (block) decomposition our data show that we can generate about 900 million elements in less than 300 seconds.
Andrey N. Chernikov, Nikos Chrisochoides
ICS1
2004 A Load Balancing Framework for Adaptive and Asynchronous Applications
abstract
We describe the design of a flexible load balancing framework and runtime software system for supporting the development of adaptive applications on distributed-memory parallel computers. The runtime system supports a global namespace, transparent object migration, automatic message forwarding and routing, and automatic load balancing. These features can be used at the discretion of the application developer in order to simplify program development and to eliminate complex bookkeeping associated with mobile data objects. An evaluation of this system in the context of a three-dimensional tetrahedral advancing front parallel mesh generator shows that overall runtime improvements of 15 percent compared to common stop-and-repartition load balancing methods, 30 percent compared to explicit intrusive load balancing methods, and 42 percent compared to no load balancing are possible on large processor configurations. At the same time, the overheads attributable to the runtime system are a fraction of 1 percent of the total runtime. The parallel advancing front method is a coarse-grained and highly adaptive application and therefore exercises all of the features of the runtime system.
Kevin J. Barker, Andrey N. Chernikov, Nikos Chrisochoides, Keshav Pingali
IEEE Trans. Parallel Distributed Syst.2