VLDB 2026 Research / reviewers in the wild / expert
Ping Xiang
dblp:60/6730
· DBLP profile ↗
23ranked-venue papers
5as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 6Artificial intelligence and machine learning · 3 · 1 since 2021Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
Electronic design automation · 29% GPUs and heterogeneous computing · 23% Cloud and datacenter computing · 19% | |
| Software engineering, system software, and programming languages
4 papers |
Compilers and program optimization · 100% | |
| Computer networks
1 paper |
Routing and switching · 100% |
Topics — the 21 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
high-level synthesis |
0.6 | 1 | 2022 | Latency-driven Optimization of Switching Pipeline Design in Network Chips · RTSS 2022 |
Electronic design automation
physical design |
0.6 | 1 | 2022 | Latency-driven Optimization of Switching Pipeline Design in Network Chips · RTSS 2022 |
Processor architecture and microarchitecture › pipelining
pipeline optimization |
0.6 | 1 | 2022 | Latency-driven Optimization of Switching Pipeline Design in Network Chips · RTSS 2022 |
Cloud and datacenter computing › resource allocation
resource mapping |
0.6 | 1 | 2022 | Latency-driven Optimization of Switching Pipeline Design in Network Chips · RTSS 2022 |
Compilers and program optimization › accelerator compilation
GPU compiler |
0.3 | 2 | 2012 | A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012 An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
GPUs and heterogeneous computing
GPU programming |
0.3 | 2 | 2012 | A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012 A GPGPU compiler for memory optimization and parallelism management · PLDI 2010 |
GPUs and heterogeneous computing
GPU architecture |
0.2 | 1 | 2014 | Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014 |
Cloud and datacenter computing
resource management |
0.2 | 1 | 2014 | Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014 |
Parallel and multicore computing
thread-level parallelism |
0.2 | 1 | 2014 | Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.2 | 1 | 2014 | Warp-level divergence in GPUs: Characterization, impact, and mitigation · HPCA 2014 |
Routing and switching
switch architecture |
0.2 | 1 | 2022 | Latency-driven Optimization of Switching Pipeline Design in Network Chips · RTSS 2022 |
Compilers and program optimization › memory optimization
memory hierarchy optimization |
0.1 | 1 | 2012 | A unified optimizing compiler framework for different GPGPU architectures · ACM Trans. Archit. Code Optim. 2012 |
Memory systems
cache |
0.1 | 1 | 2012 | CPU-assisted GPGPU on fused CPU-GPU architectures · HPCA 2012 |
GPUs and heterogeneous computing › CPU-GPU heterogeneous computing
fused CPU-GPU architecture |
0.1 | 1 | 2012 | CPU-assisted GPGPU on fused CPU-GPU architectures · HPCA 2012 |
Compilers and program optimization › accelerator compilation
GPU compiler optimization |
0.1 | 1 | 2010 | A GPGPU compiler for memory optimization and parallelism management · PLDI 2010 |
Compilers and program optimization › memory optimization
memory access optimization |
0.1 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
Memory systems › data locality
data reuse |
0.1 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
GPUs and heterogeneous computing
GPU computing |
0.1 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
Memory systems › memory hierarchy
memory hierarchy optimization |
0.1 | 1 | 2010 | A GPGPU compiler for memory optimization and parallelism management · PLDI 2010 |
Compilers and program optimization › compiler construction
compiler algorithms |
0.0 | 1 | 2012 | CPU-assisted GPGPU on fused CPU-GPU architectures · HPCA 2012 |
GPUs and heterogeneous computing › GPU memory access
memory coalescing |
0.0 | 1 | 2010 | An optimizing compiler for GPGPU programs with input-data sharing · PPoPP 2010 |
Methods — techniques the papers use, named apart from their topics
multi-objective tabu search · 1.1greedy algorithm · 1.1NSGA-II · 1.1pre-execution · 0.3l2 prefetcher · 0.3kernel generation · 0.3instruction-level parallelism · 0.3architecture-specific optimization · 0.3warp-level resource allocation · 0.2dynamic resource release · 0.2parallelism management · 0.1memory hierarchy optimization · 0.1data prefetching · 0.1auto-tuning · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A novel vehicle-bridge interaction framework for long-span bridges considering multi-scale heterogeneous traffic pattern
Zhou Huang 0003, Xinfeng Yin, Yang Quan, Ping Xiang |
Adv. Eng. Informatics | 5 |
| 2024 | Fusing multi-source quality statistical data for construction risk assessment and warning based on deep learning
Binwei Gao, Zhehao Ma, Jianan Gu, Xueqiao Han, Ping Xiang, Xiaoyue Lv |
Knowl. Based Syst. | 5 |
| 2022 | Latency-driven Optimization of Switching Pipeline Design in Network ChipsabstractA network switch implements multiple services and each service is formed by a number of match-action operations through several pipeline stages. These services running in the switch equipment are to process various packets based on standard internet protocols to decide the route of each packet. Data packets come in serial to a port, where each packet is processed by a service according to the contents of the packet headers and then send out via another port. Design of the switch, i.e., mapping services to physical resources in the pipeline stages, aims to achieve low switching latency with small chip area while respecting data-flow dependencies and hardware constraints. The current practice relies on expertise of engineers empirically, which is laborious and generates mediocre results. In this paper, we propose a switching pipeline design optimizatton technique, called SPOT. Our main contributions are as follows: (i) We first formulate the bi-objective (latency and chip area) constrained design optimization problem; (ii) SPOT quickly spots a feasible solution from a largely unfeasible design space using a dependency-aware greedy algorithm; (iii) Based on the above feasible seed, SPOT explores the design space with hundreds of decision dimensions towards Pareto optimal solutions using non-dominated sorting genetic algorithm II (NSGA-II) and multi-objective tabu search (MOTS), both adapted to be deployed in this problem setting. We apply SPOT on three sets of real-world network services. In comparison to the design sheets prepared by expert engineers, experiments show that SPOT offers 20.63% shorter service latency and 4.55% smaller chip area on average. As a by-product, the power consumption is lowered by 23.72% on average, which is correlated to the chip area. For hard real-time scenarios, the longest service latency a data packet may experience is the major concern. SPOT reduces the worst-case service latency by 12.65% on average. SPOT is the first automated optimization solution for switching pipeline design in network chips, being utilized in millions of network products of various kinds and saving manual efforts from days to minutes. Debayan Roy, Hui Chen 0016, Ping Xiang, Yuhong Feng, Wanli Chang 0001 |
RTSS | 5 |
| 2020 | NOMA based VR Video Transmissions Exploiting User Behavioral CoherenceabstractIn this work, we study the cooperative and non-cooperative transmission schemes design for live VR video broadcast scenarios by utilizing non-orthogonal multiple access (NOMA), considering that users’ viewports partly overlap due to behavioral coherence. To characterize the performance of the proposed cooperative and non-cooperative transmission schemes, the exact and asymptotic expressions of outage probability, as well as the average outage capacity under imperfect successive interference cancellation (SIC), are derived, respectively. Based on the asymptotic outage probability results, we optimize the power allocation to maximize the average outage capacity of the proposed schemes. Finally, simulation results demonstrate that both of the proposed schemes can achieve a considerable performance gain over the traditional orthogonal multiple access (OMA) scheme in average outage capacity, and each of the proposed schemes has its advantages and applicable scenarios. Ping Xiang, Hangguan Shan, Zhaoyang Zhang 0001, Lu Yu 0003, Tony Q. S. Quek |
WCNC | 1 |
| 2015 | Revisiting ILP Designs for Throughput-Oriented GPGPU ArchitectureabstractMany-core architectures such as graphics processing units (GPUs) rely on thread-level parallelism (TLP)to overcome pipeline hazards. Consequently, each core in a many-core processor employs a relatively simple in-order pipeline with limited capability to exploit instruction-level parallelism (ILP). In this paper, we study the ILP impact on the throughput-oriented many-core architecture, including data bypassing, score boarding and branch prediction. We show that these ILP techniques significantly reduce the performance dependency on TLP. This is especially useful for applications, whose resource usage limits the hardware to run a high number of threads concurrently. Furthermore, ILP techniques reduce the demand on on-chip resource to support high TLP. Given the workload-dependent impact from ILP, we propose heterogeneous GPGPU architecture, consisting of both the cores designed for high TLP and those customized with ILPtechniques. Our results show that our heterogeneous GPUarchitecture achieves high throughput as well as high energy and area-efficiency compared to homogenous designs. Ping Xiang, Yi Yang 0018, Mike Mantor, Norman Rubin, Huiyang Zhou |
CCGRID | 1 |
| 2014 | Warp-level divergence in GPUs: Characterization, impact, and mitigationabstractHigh throughput architectures rely on high thread-level parallelism (TLP) to hide execution latencies. In state-of-art graphics processing units (GPUs), threads are organized in a grid of thread blocks (TBs) and each TB contains tens to hundreds of threads. With a TB-level resource management scheme, all the resource required by a TB is allocated/released when it is dispatched to / finished in a streaming multiprocessor (SM). In this paper, we highlight that such TB-level resource management can severely affect the TLP that may be achieved in the hardware. First, different warps in a TB may finish at different times, which we refer to as `warp-level divergence'. Due to TB-level resource management, the resources allocated to early finished warps are essentially wasted as they need to wait for the longest running warp in the same TB to finish. Second, TB-level management can lead to resource fragmentation. For example, the maximum number of threads to run on an SM in an NVIDIA GTX 480 GPU is 1536. For an application with a TB containing 1024 threads, only 1 TB can run on the SM even though it has sufficient resource for a few hundreds more threads. To overcome these inefficiencies, we propose to allocate and release resources at the warp level. Warps are dispatched to an SM as long as it has sufficient resource for a warp rather than a TB. Furthermore, whenever a warp is completed, its resource is released and can accommodate a new warp. This way, we effectively increase the number of active warps without actually increasing the size of critical resources. We present our lightweight architectural support for our proposed warp-level resource management. The experimental results show that our approach achieves up to 76.0% and an average of 16.0% performance gains and up to 21.7% and an average of 6.7% energy savings at minor hardware overhead. Ping Xiang, Yi Yang 0018, Huiyang Zhou |
HPCA | 1 |
| 2014 | A Case for a Flexible Scalar Unit in SIMT ArchitectureabstractThe wide availability and the Single-Instruction Multiple-Thread (SIMT)-style programming model have made graphics processing units (GPUs) a promising choice for high performance computing. However, because of the SIMT style processing, an instruction will be executed in every thread even if the operands are identical for all the threads. To overcome this inefficiency, the AMD's latest Graphics Core Next (GCN) architecture integrates a scalar unit into a SIMT unit. In GCN, both the SIMT unit and the scalar unit share a single SIMT style instruction stream. Depending on its type, an instruction is issued to either a scalar or a SIMT unit. In this paper, we propose to extend the scalar unit so that it can either share the instruction stream with the SIMT unit or execute a separate instruction stream. The program to be executed by the scalar unit is referred to as a scalar program and its purpose is to assist SIMT-unit execution. The scalar programs are either generated from SIMT programs automatically by the compiler or manually developed by expert developers. We make a case for our proposed flexible scalar unit through three collaborative execution paradigms: data prefetching, control divergence elimination, and scalar-workload extraction. Our experimental results show that significant performance gains can be achieved using our proposed approaches compared to the state-of-art SIMT style processing. Yi Yang 0018, Ping Xiang, Mike Mantor, Norman Rubin, Lisa R. Hsu, Qunfeng Dong, Huiyang Zhou |
IPDPS | 2 |
| 2013 | Exploiting uniform vector instructions for GPGPU performance, energy efficiency, and opportunistic reliability enhancementabstractState-of-art graphics processing units (GPUs) employ the single-instruction multiple-data (SIMD) style execution to achieve both high computational throughput and energy efficiency. As previous works have shown, there exists significant computational redundancy in SIMD execution, where different execution lanes operate on the same operand values. Such value locality is referred to as uniform vectors. In this paper, we first show that besides redundancy within a uniform vector, different vectors can also have the identical values. Then, we propose detailed architecture designs to exploit both types of redundancy. For redundancy within a uniform vector, we propose to either extend the vector register file with token bits or add a separate small scalar register file to eliminate redundant computations as well as redundant data storage. For redundancy across different uniform vectors, we adopt instruction reuse, proposed originally for CPU architectures, to detect and eliminate redundancy. The elimination of redundant computations and data storage leads to both significant energy savings and performance improvement. Furthermore, we propose to leverage such redundancy to protect arithmetic-logic units (ALUs) and register files against hardware errors. Our detailed evaluation shows that our proposed design has low hardware overhead and achieves performance gains, up to 23.9% and 12.0% on average, along with energy savings, up to 24.8% and 12.6% on average, as well as a 21.1% and 14.1% protection coverage for ALUs and register files, respectively. Ping Xiang, Yi Yang 0018, Mike Mantor, Norman Rubin, Lisa R. Hsu, Huiyang Zhou |
ICS | 1 |
| 2013 | Locality principle revisited: A probability-based quantitative approach
Saurabh Gupta 0002, Ping Xiang, Yi Yang 0018, Huiyang Zhou |
J. Parallel Distributed Comput. | 2 |
| 2012 | Many-thread aware instruction-level parallelism: architecting shader cores for GPU computingabstractNo abstract available. Ping Xiang, Yi Yang 0018, Mike Mantor, Norman Rubin, Huiyang Zhou |
PACT | 1 |
| 2012 | Shared memory multiplexing: a novel way to improve GPGPU throughputabstractOn-chip shared memory (a.k.a. local data share) is a critical resource to many GPGPU applications. In current GPUs, the shared memory is allocated when a thread block (also called a workgroup) is dispatched to a streaming multiprocessor (SM) and is released when the thread block is completed. As a result, the limited capacity of shared memory becomes a bottleneck for a GPU to host a high number of thread blocks, limiting the otherwise available thread-level parallelism (TLP). In this paper, we propose software and/or hardware approaches to multiplex the shared memory among multiple thread blocks. Yi Yang 0018, Ping Xiang, Mike Mantor, Norman Rubin, Huiyang Zhou |
PACT | 2 |
| 2012 | CPU-assisted GPGPU on fused CPU-GPU architecturesabstractThis paper presents a novel approach to utilize the CPU resource to facilitate the execution of GPGPU programs on fused CPU-GPU architectures. In our model of fused architectures, the GPU and the CPU are integrated on the same die and share the on-chip L3 cache and off-chip memory, similar to the latest Intel Sandy Bridge and AMD accelerated processing unit (APU) platforms. In our proposed CPU-assisted GPGPU, after the CPU launches a GPU program, it executes a pre-execution program, which is generated automatically from the GPU kernel using our proposed compiler algorithms and contains memory access instructions of the GPU kernel for multiple thread-blocks. The CPU pre-execution program runs ahead of GPU threads because (1) the CPU pre-execution thread only contains memory fetch instructions from GPU kernels and not floating-point computations, and (2) the CPU runs at higher frequencies and exploits higher degrees of instruction-level parallelism than GPU scalar cores. We also leverage the prefetcher at the L2-cache on the CPU side to increase the memory traffic from CPU. As a result, the memory accesses of GPU threads hit in the L3 cache and their latency can be drastically reduced. Since our pre-execution is directly controlled by user-level applications, it enjoys both high accuracy and flexibility. Our experiments on a set of benchmarks show that our proposed pre-execution improves the performance by up to 113% and 21.4% on average. Yi Yang 0018, Ping Xiang, Mike Mantor, Huiyang Zhou |
HPCA | 2 |
| 2012 | Fixing Performance Bugs: An Empirical Study of Open-Source GPGPU ProgramsabstractGiven the extraordinary computational power of modern graphics processing units (GPUs), general purpose computation on GPUs (GPGPU) has become an increasingly important platform for high performance computing. To better understand how well the GPU resource has been utilized by application developers and then to facilitate them to develop high performance GPGPU code, we conduct an empirical study on GPGPU programs from ten open-source projects. These projects span a wide range of disciplines and many are designed as high performance libraries. Among these projects, we found various performance 'bugs', i.e., code segments leading to inefficient use of GPU hardware. We characterize these performance bugs, and propose the bug fixes. Our experiments confirm both significant performance gains and energy savings from our fixes and reveal interesting insights on different GPUs. Yi Yang 0018, Ping Xiang, Mike Mantor, Huiyang Zhou |
ICPP | 2 |
| 2012 | Locality Principle Revisited: A Probability-Based Quantitative ApproachabstractThis paper revisits the fundamental concept of the locality of references and proposes to quantify it as a conditional probability: in an address stream, given the condition that an address is accessed, how likely the same address (temporal locality) or an address within its neighborhood (spatial locality) will be accessed in the near future. Based on this definition, spatial locality is a function of two parameters, the neighborhood size and the scope of near future, and can be visualized with a 3D mesh. Temporal locality becomes a special case of spatial locality with the neighborhood size being zero byte. Previous works on locality analysis use stack/reuse distances to compute distance histograms as a measure of temporal locality. For spatial locality, some ad-hoc metrics have been proposed as a quantitative measure. In contrast, our conditional probability-based locality measure has a clear mathematical meaning, offers justification for distance histograms, and provides a theoretically sound and unified way to quantify both temporal and spatial locality. The proposed locality measure clearly exhibits the inherent application characteristics, from which we can easily derive information such as the sizes of the working data sets and how locality can be exploited. We showcase that our quantified locality visualized in 3D-meshes can be used to evaluate compiler optimizations, to analyze the locality at different levels of memory hierarchy, to optimize the cache architecture to effectively leverage the locality, and to examine the effect of data prefetching mechanisms. A GPU-based parallel algorithm is also presented to accelerate the locality computation for large address traces. Saurabh Gupta 0002, Ping Xiang, Yi Yang 0018, Huiyang Zhou |
IPDPS | 2 |
| 2012 | A unified optimizing compiler framework for different GPGPU architecturesabstractThis article presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism. The input to our compiler is a naïve GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler generates two kernels, one optimized for global memories and the other for texture memories. The proposed compilation process is effective for both AMD/ATI and NVIDIA GPUs. The experiments show that our optimized code achieves very high performance, either superior or very close to highly fine-tuned libraries. Yi Yang 0018, Ping Xiang, Jingfei Kong, Mike Mantor, Huiyang Zhou |
ACM Trans. Archit. Code Optim. | 2 |
| 2011 | A new shape based segmentation framework using statistical and variational methodsabstractIn this paper, we propose a new shape based segmentation and registration of the vertebral bodies (VBs) in clinical computed tomography (CT) images. The VB and surrounding organs have very close gray level information and there are no strong edges in some CT images. To overcome these challenges, image appearance and shape information of VBs are used. There are three phases of our experiments: i) the detection of the VB region using the Matched filter, ii) initial segmentation using the graph cuts which integrates the intensity and spatial interaction models, iii) registration of the shape priors and initially segmented region to obtain the final segmentation. Preliminary results show that our proposed algorithm gives very encouraging results and can solve many segmentation and registration problems. Melih S. Aslan, Hossam E. Abdelmunim, Aly A. Farag, Ben Arnold, Eslam A. Mostafa, Ping Xiang |
ICIP | 6 |
| 2010 | A novel, fast, and complete 3D segmentation of vertebral bonesabstractBone mineral density (BMD) measurements and fracture analysis of the spine bones are restricted to the Vertebral bodies (VBs), especially the trabecular bones (TBs). In this paper, we propose a novel, fast, and robust 3D framework to segment VBs and trabecular bones in clinical computed tomography (CT) images without any user intervention. The Matched filter is employed to detect the VB region automatically. To segment the whole VB, the graph cuts method which integrates a linear combination of Gaussians (LCG) and Markov Gibbs Random Field (MGRF) is used. Then, the cortical and trabecular bones are segmented using local volume growing methods. Validity was analyzed using ground truths of data sets (expert segmentation) and the European Spine Phantom (ESP) as a known reference. Experiments on the data sets show that the proposed segmentation approach is more accurate than other known alternatives. Melih S. Aslan, Asem M. Ali, Ham M. Rara, Ben Arnold, Rachid Fahmi, Aly A. Farag, Ping Xiang |
ICASSP | 7 |
| 2010 | 3D vertebrae segmentation using graph cuts with shape prior constraintsabstractOsteoporosis is a bone disease characterized by a reduction in bone mass, resulting in an increased risk of fractures. To diagnose the osteoporosis accurately, bone mineral density (BMD) measurements and fracture analysis (FA) of the Vertebral bodies (VBs) are required. In this paper, we propose a robust and 3D shape based method to segment VBs in clinical computed tomography (CT) images in order to make BMD measurements and FA accurately. In this experiment, image appearance and shape information of VBs are used. In the training step, 3D shape information is obtained from a set of data sets. Then, we estimate the shape variations using a distance probabilistic model which approximates the marginal densities of the VB and background in the variability region. In the segmentation step, the Matched filter is used to detect the VB region automatically. We align the detected volume with 3D shape prior in order to be used in distance probabilistic model. Then, the graph cuts method which integrates the linear combination of Gaussians (LCG), Markov Gibbs Random Field (MGRF), and distance probabilistic model obtained from 3D shape prior is used. Melih S. Aslan, Asem M. Ali, Dongqing Chen, Ben Arnold, Aly A. Farag, Ping Xiang |
ICIP | 6 |
| 2010 | 3D Vertebrae Segmentation in CT Images with Random NoisesabstractExposure levels (X-ray tube amperage and peak kilovoltage) are associated with various noise levels and radiation dose. When higher exposure levels are applied, the images have higher signal to noise ratio (SNR) in the CT images. However, the patient receives higher radiation dose in this case. In this paper, we use our robust 3D framework to segment vertebral bodies (VBs) in clinical computed tomography (CT) images with different noise levels. The Matched filter is employed to detect the VB region automatically. In the graph cuts method, a VB (object) and surrounding organs (background) are represented using a gray level distribution models which are approximated by a linear combination of Gaussians (LCG). Initial segmentation based on the LCG models is then iteratively refined by using Markov Gibbs random field(MGRF) with analytically estimated potentials. Experiments on the data sets show that the proposed segmentation approach is more accurate and robust than other known alternatives. Melih S. Aslan, Asem M. Ali, Aly A. Farag, Ben Arnold, Dongqing Chen, Ping Xiang |
ICPR | 6 |
| 2010 | 3D Vertebral Body Segmentation Using Shape Based Graph CutsabstractBone mineral density (BMD) measurements and fracture analysis of the spine bones are restricted to the Vertebral bodies (VBs). In this paper, we propose a novel 3D shape based method to segment VBs in clinical computed tomography (CT) images without any user intervention. The proposed method depends on both image appearance and shape information. 3D shape information is obtained from a set of training data sets. Then, we estimate the shape variations using a distance probabilistic model which approximates the marginal densities of the VB and background in the variability region. To segment a VB, the Matched filter is used to detect the VB region automatically. We align the detected volume with 3D shape prior in order to be used in distance probabilistic model. Then, the graph cuts method which integrates the linear combination of Gaussians (LCG), Markov Gibbs Random Field (MGRF), and distance probabilistic model obtained from 3D shape prior is used. Experiments on the data sets show that the proposed segmentation approach is more accurate than other known alternatives. Melih S. Aslan, Asem M. Ali, Aly A. Farag, Ham M. Rara, Ben Arnold, Ping Xiang |
ICPR | 6 |
| 2010 | A GPGPU compiler for memory optimization and parallelism managementabstractThis paper presents a novel optimizing compiler for general purpose computation on graphics processing units (GPGPU). It addresses two major challenges of developing high performance GPGPU programs: effective utilization of GPU memory hierarchy and judicious management of parallelism. Yi Yang 0018, Ping Xiang, Jingfei Kong, Huiyang Zhou |
PLDI | 2 |
| 2010 | An optimizing compiler for GPGPU programs with input-data sharingabstractDeveloping high performance GPGPU programs is challenging for application developers since the performance is dependent upon how well the code leverages the hardware features of specific graphics processors. To solve this problem and relieve application developers of low-level hardware-specific optimizations, we introduce a novel compiler to optimize GPGPU programs. Our compiler takes a naive GPU kernel function, which is functionally correct but without any consideration for performance optimization. The compiler then analyzes the code, identifies memory access patterns, and generates optimized code. The proposed compiler optimizations target at one category of scientific and media processing algorithms, which has the characteristics of input-data sharing when computing neighboring output pixels/elements. Many commonly used algorithms, such as matrix multiplication, convolution, etc., share such characteristics. For these algorithms, novel approaches are proposed to enforce memory coalescing and achieve effective data reuse. Data prefetching and hardware-specific tuning are also performed automatically with our compiler framework. The experimental results based on a set of applications show that our compiler achieves very high performance, either superior or very close to the highly fine-tuned library, NVIDIA CUBLAS 2.1. Yi Yang 0018, Ping Xiang, Jingfei Kong, Huiyang Zhou |
PPoPP | 2 |
| 2009 | Segmentation of trabecular bones from Vertebral bodies in volumetric CT spine imagesabstractWe present a 3D segmentation technique of trabecular (cancellous) bones in CT images of Vertebral bodies (VBs). In order to be used for Bone Mineral Density (BMD) measurements, the cortical and trabecular bones are subsequently segmented using graph cuts method and local volume growing methods separately. In the final step, we measure our segmentation accuracy for each method. Validity was analyzed using ground truths of data sets and the European Spine Phantom (ESP). Preliminary results are very encouraging and a reproducibility of the results was achieved for 16 data sets. The average segmentation error is below 2.0% for both methods. Melih S. Aslan, Asem M. Ali, Ben Arnold, Rachid Fahmi, Aly A. Farag, Ping Xiang |
ICIP | 6 |