Yili Zheng

dblp:85/4075 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
1since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Parallel and multicore computing · 39% High-performance computing · 28% GPUs and heterogeneous computing · 11%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming models
0.222012
Designing a unified programming model for heterogeneous machines · SC 2012
Shared memory programming for large scale machines · PLDI 2006
High-performance computing › parallel numerical algorithms
communication-avoiding algorithms
0.112012
Communication avoiding and overlapping for numerical linear algebra · SC 2012
Distributed systems › communication optimization
communication-computation overlap
0.112012
Communication avoiding and overlapping for numerical linear algebra · SC 2012
GPUs and heterogeneous computing
heterogeneous programming models
0.112012
Designing a unified programming model for heterogeneous machines · SC 2012
High-performance computing
numerical linear algebra
0.112012
Communication avoiding and overlapping for numerical linear algebra · SC 2012
Parallel and multicore computing › parallel programming models
unified programming model
0.112012
Designing a unified programming model for heterogeneous machines · SC 2012
Parallel and multicore computing › parallel programming models › distributed memory programming models
partitioned global address space
0.122012
Shared memory programming for large scale machines · PLDI 2006
Communication avoiding and overlapping for numerical linear algebra · SC 2012
Compilers and program optimization
optimizing compiler
0.112006
Shared memory programming for large scale machines · PLDI 2006
Storage systems › file systems
distributed file system
0.012004
Kosha: A Peer-to-Peer Enhancement for the Network File System · SC 2004
Storage systems › file systems › distributed file system
file replication
0.012004
Kosha: A Peer-to-Peer Enhancement for the Network File System · SC 2004
Parallel and multicore computing
load balancing
0.012004
Kosha: A Peer-to-Peer Enhancement for the Network File System · SC 2004
Storage systems › distributed storage
peer-to-peer storage
0.012004
Kosha: A Peer-to-Peer Enhancement for the Network File System · SC 2004
High-performance computing
distributed memory systems
0.012006
Shared memory programming for large scale machines · PLDI 2006

Methods — techniques the papers use, named apart from their topics

performance modeling · 0.1one-sided communication · 0.1GASNet runtime · 0.12.5d algorithm · 0.1asynchronous message · 0.1affinity test elimination · 0.1
YearPublicationVenuePosition
2026 A new enhanced lightweight detection model to identify stored grain insects on grain bulk surfaces
abstract
Rapid and accurate detection of stored grain insects is essential for minimizing insect damage. To address challenges of stored grain insect identification, an enhanced lightweight detection model was developed. The developed model integrated Channel-Transposed Attention (CTA) with the C3K module to improve fine-grained feature representation and reduce background interference. To further strengthen detection robustness under poor illumination, a novel image preprocessing component, named CPA-Enhancer, was embedded into the backbone network. This developed module adaptively adjusted image contrast and exposure to enhance feature visibility under dim and uneven light conditions. In addition, to overcome the number imbalance of insect images provided in the dataset, an Adaptive Threshold Focal Loss (ATFL) function was introduced. This function increased sensitivity to minor classes while maintaining overall model stability. To evaluate the performance of the developed model, subset and ablation experiments and comparison were conducted among different models under the same configuration and by employing an Edge device. The developed model attained a precision of 91.1%, recall of 93.0%, F1 score of 92.0%, and mean Average Precision (mAP) of 94.2% when all dataset was used. Ablation study verified the individual and synergized effectiveness of the CTA, CPA-Enhancer, and ATFL modules. Moreover, deployment evaluations on edge device confirmed its capability for real-time insect detection under resource-constrained field environments.
Jinhui Zhao, Yili Zheng, Fuji Jian, Xueyan Zhu
Eng. Appl. Artif. Intell.4
2016 A Hartree-Fock Application Using UPC++ and the New DArray Library
abstract
The Hartree-Fock (HF) method is the fundamental first step for incorporating quantum mechanics into many-electron simulations of atoms and molecules, and it is an important component of computational chemistry toolkits like NWChem. The GTFock code is an HF implementation that, while it does not have all the features in NWChem, represents crucial algorithmic advances that reduce communication and improve load balance by doing an up-front static partitioning of tasks, followed by work stealing whenever necessary. To enable innovations in algorithms and exploit next generation exascale systems, it is crucial to support quantum chemistry codes using expressive and convenient programming models and runtime systems that are also efficient and scalable. This paper presents an HF implementation similar to GTFock using UPC++, a partitioned global address space model that includes flexible communication, asynchronous remote computation, and a powerful multidimensional array library. UPC++ offers runtime features that are useful for HF such as active messages, a rich calculus for array operations, hardware-supported fetch-and-add, and functions for ensuring asynchronous runtime progress. We present a new distributed array abstraction, DArray, that is convenient for the kinds of random-access array updates and linear algebra operations on block-distributed arrays with irregular data ownership. We analyze the performance of atomic fetch-and-add operations (relevant for load balancing) and runtime attentiveness, then compare various techniques and optimizations for each. Our optimized implementation of HF using UPC++ and the DArrays library shows up to 20% improvement over GTFock with Global Arrays at scales up to 24,000 cores.
David Ozog, Amir Kamil, Yili Zheng, Paul Hargrove, Jeff R. Hammond, Allen D. Malony, Wibe de Jong, Katherine A. Yelick
IPDPS3
2015 Parallel Hessian Assembly for Seismic Waveform Inversion Using Global Updates
abstract
We present the design and evaluation of a distributed matrix-assembly abstraction for large-scale inverse problems in HPC environments: namely, physics-based Hessian estimation in full-waveform seismic inversion at the scale of the entire globe. Our solution to this data-assimilation problem relies on UPC++, a new PGAS extension to the C++language, to implement one-sided asynchronous updates to distributed matrix elements, and allows us to tackle inverse problems well beyond our previous capabilities. Our evaluation includes scaling results for Hessian estimation on up to 12, 288 cores, typical of current production scientific runs and next-generation inversions. We also present comparisons with an alternative implementation based on MPI-3 remote memory access (RMA) operations, focusing on performance and code complexity. Interoperability between UPC++and other parallel programming tools (e.g. MPI, OpenMP) allowed for incremental adoption of the PGAS model where most beneficial. Further, we note that this model of asynchronous assembly can generalize to other data-assimilation applications that accumulate updates into shared global state.
Scott French, Yili Zheng, Barbara Romanowicz, Katherine A. Yelick
IPDPS2
2014 UPC++: A PGAS Extension for C++
abstract
Partitioned Global Address Space (PGAS) languages are convenient for expressing algorithms with large, random-access data, and they have proven to provide high performance and scalability through lightweight one-sided communication and locality control. While very convenient for moving data around the system, PGAS languages have taken different views on the model of computation, with the static Single Program Multiple Data (SPMD) model providing the best scalability. In this paper we present UPC++, a PGAS extension for C++ that has three main objectives: 1) to provide an object-oriented PGAS programming model in the context of the popular C++ language, 2) to add useful parallel programming idioms unavailable in UPC, such as asynchronous remote function invocation and multidimensional arrays, to support complex scientific applications, 3) to offer an easy on-ramp to PGAS programming through interoperability with other existing parallel programming systems (e.g., MPI, OpenMP, CUDA). We implement UPC++ with a "compiler-free" approach using C++ templates and runtime libraries. We borrow heavily from previous PGAS languages and describe the design decisions that led to this particular set of language features, providing significantly more expressiveness than UPC with very similar performance characteristics. We evaluate the programmability and performance of UPC++ using five benchmarks on two representative supercomputers, demonstrating that UPC++ can deliver excellent performance at large scale up to 32K cores while offering PGAS productivity features to C++ applications.
Yili Zheng, Amir Kamil, Michael B. Driscoll, Hongzhang Shan, Katherine A. Yelick
IPDPS1
2012 PGAS for Distributed Numerical Python Targeting Multi-core Clusters
abstract
In this paper we propose a parallel programming model that combines two well-known execution models: Single Instruction, Multiple Data (SIMD) and Single Program, Multiple Data (SPMD). The combined model supports SIMD-style data parallelism in global address space and supports SPMD-style task parallelism in local address space. One of the most important features in the combined model is that data communication is expressed by global data assignments instead of message passing. We implement this combined programming model into Python, making parallel programming with Python both highly productive and performing on distributed memory multi-core systems. We base the SIMD data parallelism on DistNumPy, an auto-parallel zing version of the Numerical Python (NumPy) package that allows sequential NumPy programs to run on distributed memory architectures. We implement the SPMD task parallelism as an extension to DistNumPy that enables each process to have direct access to the local part of a shared array. To harvest the multi-core benefits in modern processors we exploit multi-threading in both SIMD and SPMD execution models. The multi-threading is completely transparent to the user -- it is implemented in the runtime with Open MP and by using multi-threaded libraries when available. We evaluate the implementation of the combined programming model with several scientific computing benchmarks using two representative multi-core distributed memory systems -- an Intel Nehalem cluster with Infini band interconnects and a Cray XE-6 supercomputer -- up to 1536 cores. The benchmarking results demonstrate scalable good performance.
Mads Ruben Burgdorff Kristensen, Yili Zheng, Brian Vinter
IPDPS2
2012 Designing a unified programming model for heterogeneous machines
abstract
While high-efficiency machines are increasingly embracing heterogeneous architectures and massive multithreading, contemporary mainstream programming languages reflect a mental model in which processing elements are homogeneous, concurrency is limited, and memory is a flat undifferentiated pool of storage. Moreover, the current state of the art in programming heterogeneous machines tends towards using separate programming models, such as OpenMP and CUDA, for different portions of the machine. Both of these factors make programming emerging heterogeneous machines unnecessarily difficult. We describe the design of the Phalanx programming model, which seeks to provide a unified programming model for heterogeneous machines. It provides constructs for bulk parallelism, synchronization, and data placement which operate across the entire machine. Our prototype implementation is able to launch and coordinate work on both CPU and GPU processors within a single node, and by leveraging the GASNet runtime, is able to run across all the nodes of a distributed-memory machine.
Michael Garland, Manjunath Kudlur, Yili Zheng
SC3
2012 Communication avoiding and overlapping for numerical linear algebra
abstract
To efficiently scale dense linear algebra problems to future exascale systems, communication cost must be avoided or overlapped. Communication-avoiding 2.5D algorithms improve scalability by reducing inter-processor data transfer volume at the cost of extra memory usage. Communication overlap attempts to hide messaging latency by pipelining messages and overlapping with computational work. We study the interaction and compatibility of these two techniques for two matrix multiplication algorithms (Cannon and SUMMA), triangular solve, and Cholesky factorization. For each algorithm, we construct a detailed performance model that considers both critical path dependencies and idle time. We give novel implementations of 2.5D algorithms with overlap for each of these problems. Our software employs UPC, a partitioned global address space (PGAS) language that provides fast one-sided communication. We show communication avoidance and overlap provide a cumulative benefit as core counts scale, including results using over 24K cores of a Cray XE6 system.
Evangelos Georganas, Jorge González-Domínguez, Edgar Solomonik, Yili Zheng, Juan Touriño, Katherine A. Yelick
SC4
2011 Cosmic microwave background map-making at the petascale and beyond
abstract
The analysis of Cosmic Microwave Background (CMB) observations is a long-standing computational challenge, driven by the exponential growth in the size of the data sets being gathered. Since this growth is projected to continue for at least the next decade, it will be critical to extend the analysis algorithms and their implementations to peta-scale high performance computing (HPC) systems and beyond. The most computationally intensive part of the analysis is generating and reducing Monte Carlo realizations of an experiment’s data. In this work we take the current stateof-the-art simulation and mapping software and investigate its performance when pushed to tens of thousands of cores on a range of leading HPC systems, in particular focusing on the communication bottleneck that emerges at high concurrencies. We present a new communication strategy that removes this bottleneck, allowing for CMB analyses of unprecedented scale and hence fidelity. Experimental results show a communication speedup of up to 116 × using our alternative strategy. 1.
Rajesh Sudarsan, Julian Borrill, Christopher Cantalupo, Theodore Kisner, Kamesh Madduri, Leonid Oliker, Yili Zheng, Horst D. Simon
ICS7
2011 Tuning collective communication for Partitioned Global Address Space programming models
Rajesh Nishtala, Yili Zheng, Paul Hargrove, Katherine A. Yelick
Parallel Comput.2
2010 Oversubscription on multicore processors
abstract
Existing multicore systems already provide deep levels of thread parallelism; hybrid programming models and composability of parallel libraries are very active areas of research within the scientific programming community. As more applications and libraries become parallel, scenarios where multiple threads compete for a core are unavoidable. In this paper we evaluate the impact of task oversubscription on the performance of MPI, OpenMP and UPC implementations of the NAS Parallel Benchmarks on UMA and NUMA multi-socket architectures. We evaluate explicit thread affinity management against the default Linux load balancing and discuss sharing and partitioning system management techniques. Our results indicate that oversubscription provides beneficial effects for applications running in competitive environments. Sharing all the available cores between applications provides better throughput than explicit partitioning. Modest levels of oversubscription improve system throughput by 27% and provide better performance isolation of applications from their co-runners: best overall throughput is always observed when applications share cores and each is executed with multiple threads per core. Rather than “resource” symbiosis, our results indicate that the determining behavioral factor when applications share a system is the granularity of the synchronization operations.
Costin Iancu, Steven Hofmeyr, Filip Blagojevic, Yili Zheng
IPDPS4
2008 A parallel software toolkit for statistical 3-D virus reconstructions from cryo electron microscopy images using computer clusters with multi-core shared-memory nodes
abstract
A statistical approach for computing 3-D reconstructions of virus particles from cryo electron microscope images and minimal prior information has been developed which can solve a range of specific problems. The statistical approach causes high computation and storage complexity in the software implementation. A parallel software toolkit is described which allows the construction of software targeted at commodity PC clusters which is modular, reusable, and user-transparent while at the same time delivering nearly linear speedup on practical problems.
Yili Zheng, Peter C. Doerschuk
IPDPS1
2006 Shared memory programming for large scale machines
abstract
This paper describes the design and implementation of a scalable run-time system and an optimizing compiler for Unified Parallel C (UPC). An experimental evaluation on BlueGene/L®, a distributed-memory machine, demonstrates that the combination of the compiler with the runtime system produces programs with performance comparable to that of efficient MPI programs and good performance scalability up to hundreds of thousands of processors.Our runtime system design solves the problem of maintaining shared object consistency efficiently in a distributed memory machine. Our compiler infrastructure simplifies the code generated for parallel loops in UPC through the elimination of affinity tests, eliminates several levels of indirection for accesses to segments of shared arrays that the compiler can prove to be local, and implements remote update operations through a lower-cost asynchronous message. The performance evaluation uses three well-known benchmarks --- HPC RandomAccess, HPC STREAM and NAS CG --- to obtain scaling and absolute performance numbers for these benchmarks on up to 131072 processors, the full BlueGene/L machine. These results were used to win the HPC Challenge Competition at SC05 in Seattle WA, demonstrating that PGAS languages support both productivity and performance.
Christopher Barton, Calin Cascaval, Gheorghe Almási 0001, Yili Zheng, Montse Farreras, Siddhartha Chatterjee, José Nelson Amaral
PLDI4
2006 Kosha: A Peer-to-Peer Enhancement for the Network File System
Ali Raza Butt, Troy A. Johnson, Yili Zheng, Y. Charlie Hu
J. Grid Comput.3
2005 Optimization of MPI collective communication on BlueGene/L systems
abstract
BlueGene/L is currently the world's fastest supercomputer. It consists of a large number of low power dual-processor compute nodes interconnected by high speed torus and collective networks, Because compute nodes do not have shared memory, MPI is the the natural programming model for this machine. The BlueGene/L MPI library is a port of MPICH2.In this paper we discuss the implementation of MPI collectives on BlueGene/L. The MPICH2 implementation of MPI collectives is based on point-to-point communication primitives. This turns out to be suboptimal for a number of reasons. Machine-optimized MPI collectives are necessary to harness the performance of BlueGene/L. We discuss these optimized MPI collectives, describing the algorithms and presenting performance results measured with targeted micro-benchmarks on real BlueGene/L hardware with up to 4096 compute nodes.
Gheorghe Almási 0001, Philip Heidelberger, Charles Archer, Xavier Martorell, C. Christopher Erway, José E. Moreira, Burkhard D. Steinmacher-Burow, Yili Zheng
ICS8
2004 Kosha: A Peer-to-Peer Enhancement for the Network File System
abstract
This paper presents Kosha, a peer-to-peer (p2p) enhancement for the widely-used Network File System (NFS). Kosha harvests redundant storage space on cluster nodes and user desktops to provide a reliable, shared file system that acts as a large storage with normal NFS semantics. P2p storage systems provide location transparency, mobility transparency, load balancing, and file replication - features that are not available in NFS. On the other hand, NFS provides hierarchical file organization, directory listings, and file permissions, which are missing from p2p storage systems. By blending the strengths of NFS and p2p storage systems, Kosha provides a low overhead storage solution. Our experiments show that compared to unmodified NFS, Kosha introduces a 4.1% fixed overhead and 1.5% additional overhead as nodes are increased from one to eight. For larger number of nodes, the additional overhead increases slowly. Kosha achieves load balancing in distributed directories, and guarantees 99.99% or better file availability.
Ali Raza Butt, Troy A. Johnson, Yili Zheng, Y. Charlie Hu
SC3
2002 Robustness of 3-D maximum likelihood reconstructions of viruses from cryo electron microscope images
abstract
A statistical method for computing 3-D reconstructions of virus particles from cryo electron microscope images and minimal prior information-particle symmetry and radii-is described and demonstrated numerically focusing on the robustness of the approach.
Zhye Yin, Yili Zheng, Peter C. Doerschuk
ICASSP2
2002 3-D maximum likelihood reconstructions of viruses from cryo electron microscope images and parallel computation
abstract
A statistical method for computing 3-D reconstructions of virus particles from cryo electron microscope images and minimal prior information - particle symmetry and radii - is described. Two different approaches for parallel implementation of the algorithm are described based on parallel processing of different images and based on parallel numerical integration and numerical results showing nearly linear speedup are presented.
Yili Zheng, Zhye Yin, Peter C. Doerschuk
ICIP (2)1