VLDB 2026 Research / reviewers in the wild / expert
Wenjing Ma
dblp:27/5028
· DBLP profile ↗
54ranked-venue papers
15as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 8 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TBF: A Tunable Blocking-and-Fusion Algorithm for Efficient GPU Symmetric Rank-2K Updates
Lijuan Hu, Xinzhe Chen, Hongyaoxing Gu, Wenjing Ma, Fangfang Liu 0004 |
Euro-Par (1) | 5 |
| 2025 | Integrating Epigenetic and Phenotypic Features for Biological Age Estimation in Cancer Patients via Multimodal LearningabstractBiological age, which may be older or younger than chronological age due to factors such as genetic predisposition, environmental exposures, serves as a meaningful biomarker of aging processes and can inform risk stratification, treatment planning, and survivorship care in cancer patients. We propose EpiCAge, a multimodal framework that integrates epigenetic and phenotypic data to improve biological age prediction. Evaluated on eight internal and four external cancer cohorts, EpiCAge consistently outperforms existing epigenetic and phenotypic age clocks. Our analyses show that EpiCAge identifies biologically relevant markers, and its derived age acceleration is significantly associated with mortality risk. These results highlight EpiCAge as a promising multimodal machine learning tool for biological age assessment in oncology. Shuyue Jiang, Wenjing Ma, Shaojun Yu, Runze Yan, Jiaying Lu 0001 |
BIBM | 2 |
| 2025 | ScPanKD: Distilling Pan-Cancer Knowledge for Enhanced T Cell Subtypes Annotation in Single-Cell Transcriptomics DataabstractSingle-cell RNA sequencing (scRNA-seq) enables high-resolution characterization of cellular heterogeneity, and annotating major cell types has become a standard practice in scRNA -seq analysis pipelines. However, accurately identifying fine-grained subtypes within major cell types remains challenging, particularly in heterogeneous tissues such as cancer samples. Here, we present ScPanKD, a computational frame-work for accurate and robust classification of fine-grained T cell subtypes in cancer samples. Unlike existing methods that suffer from cell type mismatches between reference and query datasets due to cancer heterogeneity, ScPanKD leverages knowledge distillation (KD) to accurately identify T cell subtypes even when the reference dataset contains more subtype diversity than the query. ScPanKD learns a cancer-invariant feature space and employs a two-step strategy, anchor cell selection followed by KD, to mitigate distribution shifts between reference and query datasets. Across extensive experiments using pan-cancer level CD4+ or CD8+ T cell atlases as references, we demonstrate that ScPanKD outperforms conventional annotation methods and single-cell foundation models, achieving more accurate and robust T cell subtype classification. ScPanKD and all reproducible scripts are available at https://github.com/marvinquiet/ScPanKD. Wenjing Ma, Xiaoqing Yu, Jiaying Lu 0001 |
BIBM | 1 |
| 2025 | High-Accuracy prediction and efficient adjustment of surface shape distortion in optical elements: Model correction based on uncertainty quantification-driven transfer learning
Zhihao Fan, Xiaokai Mu, Rongxuan Zhao, Kangcheng Yin, Qingchao Sun, Kaike Yang, Wenjing Ma |
Adv. Eng. Informatics | 8 |
| 2025 | A review on knowledge graphs for healthcare: Resources, applications, and promises
Hejie Cui, Jiaying Lu 0001, Ran Xu 0002, Shiyu Wang 0002, Wenjing Ma, Yue Yu 0001, Shaojun Yu, Xuan Kan, Chen Ling 0003, Liang Zhao 0002, Zhaohui S. Qin, Joyce C. Ho, Tianfan Fu, Jing Ma 0005, Mengdi Huai, Carl Yang 0001 |
J. Biomed. Informatics | 5 |
| 2025 | Optimization of Generalized Eigensolver for Dense Symmetric Matrices on AMD GPU
Zitong Su, Wenjing Ma, Leisheng Li |
J. Comput. Sci. Technol. | 5 |
| 2024 | Target-oriented Reference Construction for supervised cell type identification in scRNA-seqabstractCell type identification is a crucial step in single-cell RNA-seq (scRNA-seq) data analysis. The supervised cell type identification method is a preferred solution due to its accuracy and efficiency. The performance of such methods highly depends on the quality of the reference data. Although there are many supervised cell type identification tools, no method currently exists for constructing reference data. Here, we develop Target-Oriented Reference Construction (TORC), a widely applicable strategy for constructing references based on target datasets in scRNA-seq supervised cell type identification. TORC focuses on alleviating differences in cell type composition between the reference and target sets. Extensive benchmarks on simulated and real data analyses demonstrate consistent improvements in cell type identification with TORC. Wenjing Ma, Zhijin Wu, Hao Wu 0003 |
BIBM | 2 |
| 2024 | Real-time scheduling for two-stage assembly flowshop with dynamic job arrivals by deep reinforcement learning
Jian Chen 0022, Hanlei Zhang, Wenjing Ma, Gangyan Xu |
Adv. Eng. Informatics | 3 |
| 2024 | cypress: an R/Bioconductor package for cell-type-specific differential expression analysis power assessmentabstractSUMMARY: Recent methodology advances in computational signal deconvolution have enabled bulk transcriptome data analysis at a finer cell-type level. Through deconvolution, identifying cell-type-specific differentially expressed (csDE) genes is drawing increasing attention in clinical applications. However, researchers still face a number of difficulties in adopting csDE genes detection methods in practice, especially in their experimental design. Here we present cypress, the first experimental design and statistical power analysis tool in csDE genes identification. This tool can reliably model purified cell-type-specific (CTS) profiles, cell-type compositions, biological and technical variations, offering a high-fidelity simulator for bulk RNA-seq convolution and deconvolution. cypress conducts simulation and evaluates the impact of multiple influencing factors, by various statistical metrics, to help researchers optimize experimental design and conduct power analysis. AVAILABILITY AND IMPLEMENTATION: cypress is an open-source R/Bioconductor package at https://bioconductor.org/packages/cypress/. Shilin Yu, Guanqun Meng, Wen Tang 0003, Wenjing Ma, Xiongwei Zhu, Hao Feng 0005 |
Bioinform. | 4 |
| 2023 | GFFT: a Task Graph Based Fast Fourier Transform Optimization FrameworkabstractFast Fourier Transform (FFT) is a widely used mathematical tool in scientific and engineering applications, and optimizing its performance remains a challenging problem. This paper introduces GFFT, a novel task-graph-based FFT optimization framework that leverages modern hardware and software techniques to achieve high-performance computation. GFFT features a tuning model that uses hardware parameters to optimize FFT decomposition, a bi-directional recursive FFT algorithm that avoids strided load in SIMD implementation, and several graph optimizers inspired by deep learning frameworks to enhance performance. In addition, GFFT utilizes task-based parallelism to exploit performance on multi-core processors and provide potential compatibility with heterogeneous systems. Experimental results demonstrate that GFFT outperforms popular FFT frameworks, achieving an average speedup of 1.17x to FFTW and 1.27x to oneMKL on the Intel Xeon processor, 1.18x to AOCL-FFTW on the AMD EPYC processor, and 2.11x to FFTW on the Sunway multi-core processor with a single thread. Additionally, GFFT achieves an average speedup of 11.48x to FFTW and 1.41x to oneMKL on the Intel Xeon processor, 9.87x to AOCL-FFTW on the AMD EPYC processor with 16-threads. Qinglin Lu, Wenjing Ma, Daokun Chen, Fangfang Liu 0004 |
ICPP | 3 |
| 2023 | HiPrompt: Few-Shot Biomedical Knowledge Fusion via Hierarchy-Oriented PromptingabstractMedical decision-making processes can be enhanced by comprehensive biomedical knowledge bases, which require fusing knowledge graphs constructed from different sources via a uniform index system. The index system often organizes biomedical terms in a hierarchy to provide the aligned entities with fine-grained granularity. To address the challenge of scarce supervision in the biomedical knowledge fusion (BKF) task, researchers have proposed various unsupervised methods. However, these methods heavily rely on ad-hoc lexical and structural matching algorithms, which fail to capture the rich semantics conveyed by biomedical entities and terms. Recently, neural embedding models have proved effective in semantic-rich tasks, but they rely on sufficient labeled data to be adequately trained. To bridge the gap between the scarce-labeled BKF and neural embedding models, we propose HiPrompt, a supervision-efficient knowledge fusion framework that elicits the few-shot reasoning ability of large language models through hierarchy-oriented prompts. Empirical results on the collected KG-Hi-BKF benchmark datasets demonstrate the effectiveness of HiPrompt. Jiaying Lu 0001, Bo Xiong 0001, Wenjing Ma, Steffen Staab, Carl Yang 0001 |
SIGIR | 4 |
| 2023 | Logic-based Benders decomposition for order acceptance and scheduling in distributed manufacturing
Jian Chen 0022, Wenjing Ma, Xudong Ye, Zhiheng Zhao |
Adv. Eng. Informatics | 2 |
| 2023 | xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002 |
CCF Trans. High Perform. Comput. | 2 |
| 2023 | Publisher Correction: xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002 |
CCF Trans. High Perform. Comput. | 2 |
| 2023 | Editorial for the special issue on new algorithms and software for E-scale high performance computing
Jiachang Sun, Wenjing Ma |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | Evolving the HPL benchmark towards multi-GPGPU clusters
Qiao Sun 0005, Wenjing Ma, Jiachang Sun, Huiyuan Li 0002 |
CCF Trans. High Perform. Comput. | 2 |
| 2023 | An Optimized Framework for Matrix Factorization on the New Sunway Many-core PlatformabstractMatrix factorization functions are used in many areas and often play an important role in the overall performance of the applications. In the LAPACK library, matrix factorization functions are implemented with blocked factorization algorithm, shifting most of the workload to the high-performance Level-3 BLAS functions. But the non-blocked part, the panel factorization, becomes the performance bottleneck, especially for small- and medium-size matrices that are the common cases in many real applications. On the new Sunway many-core platform, the performance bottleneck of panel factorization can be alleviated by keeping the panel in the LDM for the panel factorization. Therefore, we propose a new framework for implementing matrix factorization functions on the new Sunway many-core platform, facilitating the in-LDM panel factorization. The framework provides a template class with wrapper functions, which integrates inter-CPE communication for the Level-1 and Level-2 BLAS functions with flexible interfaces and can accommodate different partitioning schemes. With the framework, writing panel factorization code with data residing in the LDM space can be done with much higher productivity. We implemented three functions ( dgetrf , dgeqrf , and dpotrf ) based on the framework and compared our work with a CPE_BLAS version, which uses the original LAPACK implementation linked with optimized BLAS library that runs on the CPE mesh. Using the most favorable partitioning, the panel factorization part achieves speedup of up to 26.3, 19.1, and 18.2 for the three matrix factorization functions. For the whole function, our implementation is based on a carefully tuned recursion framework, and we added specific optimization to some subroutines used in the factorization functions. Overall, we obtained average speedup of 9.76 on dgetrf , 10.12 on dgeqrf , and 4.16 on dpotrf , compared to the CPE_BLAS version. Based on the current template class, our work can be extended to support more categories of linear algebra functions. Wenjing Ma, Fangfang Liu 0004, Daokun Chen, Qinglin Lu, Hongsen Wang, Xinhui Yuan |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | MFFT: A GPU Accelerated Highly Efficient Mixed-Precision Large-Scale FFT FrameworkabstractFast Fourier transform (FFT) is widely used in computing applications in large-scale parallel programs, and data communication is the main performance bottleneck of FFT and seriously affects its parallel efficiency. To tackle this problem, we propose a new large-scale FFT framework, MFFT, which optimizes parallel FFT with a new mixed-precision optimization technique, adopting the “high precision computation, low precision communication” strategy. To enable “low precision communication”, we propose a shared-exponent floating-point number compression technique, which reduces the volume of data communication, while maintaining higher accuracy. In addition, we apply a two-phase normalization technique to further reduce the round-off error. Based on the mixed-precision MFFT framework, we apply several optimization techniques to improve the performance, such as streaming of GPU kernels, MPI message combination, kernel optimization, and memory optimization. We evaluate MFFT on a system with 4,096 GPUs. The results show that shared-exponent MFFT is 1.23 × faster than that of double-precision MFFT on average, and double-precision MFFT achieves performance 3.53× and 9.48× on average higher than open source library 2Decomp&FFT (CPU-based version) and heFFTe (AMD GPU-based version), respectively. The parallel efficiency of double-precision MFFT increased from 53.2% to 78.1% compared with 2Decomp&FFT, and shared-exponent MFFT further increases the parallel efficiency to 83.8%. Fangfang Liu 0004, Wenjing Ma, Huiyuan Li 0002, Yuanchi Peng |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | EasyView: Enabling and Scheduling Tensor Views in Deep Learning CompilersabstractIn recent years, memory-intensive operations are becoming dominant in efficiency of running novel neural networks. Just-in-time operator fusion on accelerating devices like GPU proves an effective method for optimizing memory-intensive operations, and suits the numerous varying model structures. In particular, we find memory-intensive operations on tensor views are ubiquitous in neural network implementations. Tensors are the de facto representation for numerical data in deep learning areas, while tensor views cover a bunch of sophisticated syntax, which allow various interpretations on the underlying tensor data without memory copy. The support of views in deep learning compilers could greatly enlarge operator fusion scope, and appeal to optimizing novel neural networks. Nevertheless, mainstream solutions in state-of-the-art deep learning compilers exhibit imperfections either in view syntax representations or operator fusion. In this article, we propose EasyView, which enables and schedules tensor views in an end-to-end workflow from neural networks onto devices. Aiming at maximizing memory utilization and reducing data movement, we categorize various view contexts in high-level language, and lower views in accordance with different scenarios. Reference-semantic in terms of views are kept in the lowering from native high-level language features to intermediate representations. Based on the reserved reference-semantics, memory activities related to data dependence of read and write are tracked for further compute and memory optimization. Besides, ample operator fusion is applied to memory-intensive operations with views. In our tests, the proposed work could get average 5.63X, 2.44X, and 4.67X speedup compared with the XLA, JAX, and TorchScript, respectively for hotspot Python functions. In addition, operation fusion with views could bring 8.02% performance improvement in end-to-end neural networks. Lijuan Jiang, Qianchao Zhu, Shengen Yan, Xingcheng Zhang, Dahua Lin, Wenjing Ma, Zhouyang Li, Minxi Jin, Chao Yang 0002 |
ICPP | 8 |
| 2022 | LRcell: detecting the source of differential expression at the sub-cell-type level from bulk RNA-seq dataabstractGiven most tissues are consist of abundant and diverse (sub-)cell types, an important yet unaddressed problem in bulk RNA-seq analysis is to identify at which (sub-)cell type(s) the differential expression occurs. Single-cell RNA-sequencing (scRNA-seq) technologies can answer the question, but they are often labor-intensive and cost-prohibitive. Here, we present LRcell, a computational method aiming to identify specific (sub-)cell type(s) that drives the changes observed in a bulk RNA-seq experiment. In addition, LRcell provides pre-embedded marker genes computed from putative scRNA-seq experiments as options to execute the analyses. We conduct a simulation study to demonstrate the effectiveness and reliability of LRcell. Using three different real datasets, we show that LRcell successfully identifies known cell types involved in psychiatric disorders. Applying LRcell to bulk RNA-seq results can produce a hypothesis on which (sub-)cell type(s) contributes to the differential expression. LRcell is complementary to cell type deconvolution methods. Wenjing Ma, Sumeet Sharma, Shannon L. Gourley, Zhaohui S. Qin |
Briefings Bioinform. | 1 |
| 2020 | Enabling Highly Efficient Batched Matrix Multiplications on SW26010 Many-core ProcessorabstractWe present a systematic methodology for optimizing batched matrix multiplications on SW26010 many-core processor of the Sunway TaihuLight supercomputer. Five surrogate algorithms and a machine learning–based algorithm selector are proposed to fully exploit the computing capability of SW26010 and cope with the sophisticated algorithm characteristics of batched matrix multiplications. Experiment results show that the algorithm selector is able to adaptively choose the appropriate algorithm for various matrix shapes and batch sizes with low overhead and high accuracy. In particular, the optimized batched matrix multiplications can substantially outperform the non-batched version and reach around 84.8% of the performance upper bound. Lijuan Jiang, Chao Yang 0002, Wenjing Ma |
ACM Trans. Archit. Code Optim. | 3 |
| 2019 | Enabling Highly Efficient k-Means Computations on the SW26010 Many-Core Processor of Sunway TaihuLight
Chao Yang 0002, Qiao Sun 0005, Wenjing Ma, Wenlong Cao, Yulong Ao |
J. Comput. Sci. Technol. | 4 |
| 2018 | Extreme-Scale Realistic Stencil Computations on Sunway TaihuLight with Ten Million CoresabstractStencil computation arises from a large variety of scientific and engineering applications and often plays a critical role in the performance of extreme-scale simulations. Due to the memory bound nature, it is a challenging task to optimize stencil computation kernels on many leadership supercomputers, such as Sunway TaihuLight, which has relatively high computing throughput whilst relatively low data-moving capability. In this white paper, we show the efforts we have been making during the past two years in developing end-to-end implementation and optimization techniques for extreme-scale stencil computations on Sunway TaihuLight. We started with a work on optimizing the 3-D 2nd-order 13-point stencil for nonhydrostatic atmospheric dynamics simulation, which is an important part of the 2016 ACM Gordon Bell Prize winning work, and extended it in ways that can handle a broader range of realistic and challenging problems, such as the HPGMG benchmark that consists of memory-hungry stencils and the gaseous wave detonation simulation that relies on complex high-order stencils. The presented stencil computation paradigm on Sunway TaihuLight includes not only multilevel parallelization to exploit the parallelism on different hardware levels, but also systematic performance optimization techniques for communication, memory access, and computation. We show by extreme-scale tests that the proposed systematic stencil computation paradigm can successfully deliver remarkable performance on Sunway TaihuLight with ten million heterogeneous cores. In particular, we achieve an aggregate performance of 23.12 Pflops for the 3-D 5th order WENO stencil computation in gaseous wave detonation simulation, which is the highest performance result for high-order stencil computations as far as we know, and an aggregate performance of solving over one trillion unknowns per second in the HPGMG benchmark, which ranks the first place in the HPGMG List of Nov 2017. Chao Yang 0002, Wenjing Ma, Yulong Ao |
CCGrid | 3 |
| 2018 | Extreme-Scale High-Order WENO Simulations of 3-D Detonation Wave with 10 Million CoresabstractHigh-order stencil computations, frequently found in many applications, pose severe challenges to emerging many-core platforms due to the complexities of hardware architectures as well as the sophisticated computing and data movement patterns. In this article, we tackle the challenges of high-order WENO computations in extreme-scale simulations of 3D gaseous waves on Sunway TaihuLight. We design efficient parallelization algorithms and present effective optimization techniques to fully exploit various parallelisms with reduced memory footprints, enhanced data reuse, and balanced computation load. Test results show the optimized code can scale to 9.98 million cores, solving 12.74 trillion unknowns with 23.12 Pflops double-precision performance. Yulong Ao, Chao Yang 0002, Wenjing Ma |
ACM Trans. Archit. Code Optim. | 4 |
| 2017 | Towards Highly Efficient DGEMM on the Emerging SW26010 Many-Core ProcessorabstractThe matrix-matrix multiplication is an essential building block that can be found in various scientific and engineering applications. High-performance implementations of the matrix-matrix multiplication on state-of-the-art processors may be of great importance for both the vendors and the users. In this paper, we present a detailed methodology of implementing and optimizing the double-precision general format matrix-matrix multiplication (DGEMM) kernel on the emerging SW26010 processor, which is used to build the Sunway TaihuLight supercomputer. We propose a three level blocking algorithm to orchestrate data on the memory hierarchy and expose parallelism on different hardware levels, and design a collective data sharing scheme by using the register communication mechanism to exchange data efficiently among different cores. On top of those, further optimizations are done based on a data-thread mapping method for efficient data distribution, a double buffering scheme for asynchronous DMA data transfer, and an instruction scheduling method for maximizing the pipeline usage. Experiment results show that the proposed DGEMM implementation can fully exploit the unique hardware features provided by SW26010 and can sustain up to 95% of the peak performance. Lijuan Jiang, Chao Yang 0002, Yulong Ao, Wanwang Yin, Wenjing Ma, Qiao Sun 0005, Fangfang Liu 0004, Rongfen Lin |
ICPP | 5 |
| 2017 | 26 PFLOPS Stencil Computations for Atmospheric Modeling on Sunway TaihuLightabstractStencil computation arises from a broad set of scientific and engineering applications and often plays a critical role in the performance of extreme-scale simulations. Due to the memory bound nature, it is a challenging task to opti- mize stencil computation kernels on modern supercomputers with relatively high computing throughput whilst relatively low data-moving capability. This work serves as a demon- stration on the details of the algorithms, implementations and optimizations of a real-world stencil computation in 3D nonhydrostatic atmospheric modeling on the newly announced Sunway TaihuLight supercomputer. At the algorithm level, we present a computation-communication overlapping technique to reduce the inter-process communication overhead, a locality- aware blocking method to fully exploit on-chip parallelism with enhanced data locality, and a collaborative data accessing scheme for sharing data among different threads. In addition, a variety of effective hardware specific implementation and optimization strategies on both the process- and thread-level, from the fine-grained data management to the data layout transformation, are developed to further improve the per- formance. Our experiments demonstrate that a single-process many-core speedup of as high as 170x can be achieved by using the proposed algorithm and optimization strategies. The code scales well to millions of cores in terms of strong scalability. And for the weak-scaling tests, the code can scale in a nearly ideal way to the full system scale of more than 10 million cores, sustaining 25.96 PFLOPS in double precision, which is 20% of the peak performance. Yulong Ao, Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Fangfang Liu 0004, Lin Gan 0001, Wenjing Ma |
IPDPS | 9 |
| 2017 | Localized Fault Recovery for Nested Fork-Join ProgramsabstractNested fork-join programs scheduled using work stealing can automatically balance load and adapt to changes in the execution environment. In this paper, we design an approach to efficiently recover from faults encountered by these programs. Specifically, we focus on localized recovery of the task space in the presence of fail-stop failures. We present an approach to efficiently track, under work stealing, the relationships between the work executed by various threads. This information is used to identify and schedule the tasks to be re-executed without interfering with normal task execution. The algorithm precisely computes the work lost, incurs minimal re-execution overhead, and can recover from an arbitrary number of failures. Experimental evaluation demonstrates low overheads in the absence of failures, recovery overheads on the same order as the lost work, and much lower recovery costs than alternative strategies. Gokcen Kestor, Sriram Krishnamoorthy, Wenjing Ma |
IPDPS | 3 |
| 2016 | Multi-Scale Fully Convolutional Network for Fast Face Detection
Yancheng Bai, Wenjing Ma, Yucheng Li 0002, Liangliang Cao, Luwei Yang |
BMVC | 2 |
| 2016 | Online variational Bayesian Support Vector RegressionabstractTraditional Support Vector Regression (SVR) solvers require user pre-specified penalty (regularization) parameter as input and typically model the training data with maximum a posterior (MAP) principle. The resultant point estimates can be affected seriously by inappropriate regularization, outliers and noise, especially when training online. In this paper, we address the aforementioned problems by developing a Bayesian SVR model with the pseudo-likelihood and data augmentation idea. Then we perform variational posterior inference in an augmented variable space and the approximate posterior of model weights, rather than point estimates as in traditional SVR, are used to make robust predictions. Besides, once the approximate posterior is obtained from a given set of data, we can regard it as model prior when dealing with new arrival data, which leads to a natural way to extend our batch model to the online scenario. Experiments on several benchmark regression problems as well as a real vehicle accident rate prediction task show that our models have superior performance while inferring penalty parameter automatically. Siqi Deng, Kan Gao, Changying Du, Wenjing Ma, Guoping Long, Yucheng Li 0002 |
IJCNN | 4 |
| 2016 | GPU-FV: Realtime Fisher Vector and Its Applications in Video MonitoringabstractFisher vector has been widely used in many multimedia retrieval and visual recognition applications with good performance. However, the computation complexity prevents its usage in real-time video monitoring. In this work, we proposed and implemented GPU-FV, a fast Fisher vector extraction method with the help of modern GPUs. The challenge of implementing Fisher vector on GPUs lies in the data dependency in feature extraction and expensive memory access in Fisher vector computing. To handle these challenges, we carefully designed GPU-FV in a way that utilizes the computing power of GPU as much as possible, and applied optimizations such as loop tiling to boost the performance. GPU-FV is about 12 times faster than the CPU version, and 50\% faster than a non-optimized GPU implementation. For standard video input (320*240), GPU-FV can process each frame within 34ms on a model GPU. Our experiments show that GPU-FV obtains a similar recognition accuracy as traditional FV on VOC 2007 and Caltech 256 image sets. We also applied GPU-FV for realtime video monitoring tasks and found that GPU-FV outperforms a number of previous works. Especially, when the number of training examples are small, GPU-FV outperforms the recent popular deep CNN features borrowed from ImageNet. Wenjing Ma, Liangliang Cao, Lei Yu 0012, Guoping Long, Yucheng Li 0002 |
ICMR | 1 |
| 2016 | HPSVM: Heterogeneous Parallel SVM with Factorization Based IPM Algorithm on CPU-GPU ClusterabstractSupport vector machine (SVM) is a supervised method widely used in the statistical classification and regression analysis. SVM training can be solved via the interior point method (IPM) with the advantages of low storage, fast convergence and easy parallelization. However, it is still confronted with the challenges of training speed and memory use. In this paper, we propose a parallel primal-dual IPM algorithm based on the incomplete Cholesky factorization (ICF) for efficiently training large-scale SVMs, named HPSVM, on CPU-GPU cluster. Our approach is distinguished from earlier work in that it is specifically designed to take maximal advantage of the CPU-GPU collaborative computation with the dual buffers 3-stage pipeline mechanism, and efficiently handles large-scale training datasets. In HPSVM, the heterogeneous hierarchical memory is fully explored to alleviate the bottleneck for optimizing data transfer, and the programming paradigm is presented to build an efficient collaboration mechanism between CPU and GPU. Comprehensive experiments show that HPSVM is up to 11 times faster than the CPU version on real datasets. Tao Li 0022, Xuechen Liu 0002, Qiankun Dong, Wenjing Ma, Kai Wang 0001 |
PDP | 4 |
| 2016 | Bridging Semantic Gap Between App Names: Collective Matrix Factorization for Similar Mobile App Recommendation
Ning Bu, Shuzi Niu, Lei Yu 0012, Wenjing Ma, Guoping Long |
WISE (2) | 4 |
| 2016 | Highly Optimized Code Generation for Stencil Codes with Computation Reuse for GPUs
Wenjing Ma, Kan Gao, Guoping Long |
J. Comput. Sci. Technol. | 1 |
| 2015 | PE-TLD: Parallel Extended Tracking-Learning-Detection for Multi-target Tracking
Chenggang Zhou, Qiankun Dong, Wenjing Ma, Guoping Long, Tao Li 0022 |
ICA3PP (2) | 3 |
| 2015 | Global transformations for legacy parallel applications via structural analysis and rewriting
Daniel G. Chavarría-Miranda, Ajay Panyala, Wenjing Ma, Adrian Prantl, Sriram Krishnamoorthy |
Parallel Comput. | 3 |
| 2012 | GMProf: A low-overhead, fine-grained profiling approach for GPU programsabstractDriven by the cost-effectiveness and the power-efficiency, GPUs are being increasingly used to accelerate computations in many domains. However, developing highly efficient GPU implementations requires a lot of expertise and effort. Thus, tool support for tuning GPU programs is urgently needed, and more specifically, low-overhead mechanisms for collecting fine-grained runtime information are critically required. Unfortunately, profiling tools and mechanisms available today either collect very coarse-grained information, or have prohibitive overheads. This paper presents a low-overhead and fine-grained profiling technique developed specifically for GPUs, which we refer to as GMProf. GMProf uses two ideas to help reduce the overheads of collecting fine-grained information. The first idea involves exploiting a number of GPU architectural features to collect reasonably accurate information very efficiently, and the second idea is to use simple static analysis methods to reduce the overhead of runtime profiling. The specific implementation of GMProf we report in this paper focuses on shared memory usage. Particularly, we help programmers understand (1) which locations in shared memory are infrequently accessed? and (2) which data elements in device memory are frequently accessed? We have evaluated GMProf using six popular GPU kernels with different characteristics. Our experimental results show that GMProf, with all optimizations, incurs a moderate overhead, e.g., 1.36 times on average for shared memory profiling. Furthermore, for three of the six evaluated kernels, GMProf verified that shared memory is effectively used, and for the remaining three kernels, it not only helped accurately identify the inefficient use of shared memory, but also helped tune the implementations. The resulting tuned implementations had a speedup of 15.18 times on average. Mai Zheng, Vignesh T. Ravi, Wenjing Ma, Gagan Agrawal |
HiPC | 3 |
| 2012 | Data-driven fault tolerance for work stealing computationsabstractWork stealing is a promising technique to dynamically tolerate variations in the execution environment, including faults, system noise, and energy constraints. In this paper, we present fault tolerance mechanisms for task parallel computations, a popular computation idiom, employing work stealing. The computation is organized as a collection of tasks with data in a global address space. The completion of data operations, rather than the actual messages, is tracked to derive an idempotent data store. This information is also used to accurately identify the tasks to be re-executed in the presence of random work stealing. We consider three recovery schemes that present distinct trade-offs --- lazy recovery with potentially increased re-execution cost, immediate collective recovery with associated synchronization overheads, and noncollective recovery enabled by additional communication. We employ distributed-memory work stealing to dynamically rebalance the tasks onto the live processes and evaluate the three schemes using candidate application benchmarks. We demonstrate that the overheads (space and time) of the fault tolerance mechanism are low, the costs incurred due to failures are small, and the overheads decrease with per-process work at scale. Wenjing Ma, Sriram Krishnamoorthy |
ICS | 1 |
| 2012 | Water quality model parameters inversion based on improved stochastic optimizationabstractAs inherent optical properties (IOPs) are directly related to the constituents in the water, the condition of water quality can be reflected by fundamental IOPs absorption and scattering coefficients. And these values can be derived by analytically inverting the remote sensing spectral reflectance. In this paper, the relations between the remote sensing reflectance and water quality information are established, and the model parameters of water quality are obtained by stochastic optimization. Based on Threshold Accepting algorithm, a method with the improved searching strategy and new optimization criteria is proposed to find optimal parameters for the inversion model. The experiments conducted on the simulated data and real data, which indicate that through the division of optimization parameters and the use of different search methods, the accuracy of inversion and operational efficiency can be improved. Junping Zhang, Wenjing Ma, Jiaguo Qi |
IGARSS | 2 |
| 2012 | SHALE: an efficient algorithm for allocation of guaranteed display advertisingabstractMotivated by the problem of optimizing allocation in guaranteed display advertising, we develop an efficient, lightweight method of generating a compact allocation plan that can be used to guide ad server decisions. The plan itself uses just O(1) state per guaranteed contract, is robust to noise, and allows us to serve (provably) nearly optimally. Vijay Bharadwaj, Peiji Chen, Wenjing Ma, Chandrashekhar Nagarajan, John A. Tomlin, Sergei Vassilvitskii, Erik Vee, Jian Yang 0002 |
KDD | 3 |
| 2012 | Ad serving using a compact allocation planabstractA large fraction of online display advertising is sold via guaranteed contracts: a publisher guarantees to the advertiser a certain number of user visits satisfying the targeting predicates of the contract. The publisher is then tasked with solving the ad serving problem ---given a user visit, which of the thousands of matching contracts should be displayed, so that by the expiration time every contract has obtained the requisite number of user visits. The challenges of the problem come from (1) the sheer size of the problem being solved, with tens of thousands of contracts and billions of user visits, (2) the unpredictability of user behavior, since these contracts are sold months ahead of time, when only a forecast of user visits is available and (3) the minute amount of resources available online, as an ad server must respond with a matching contract in a fraction of a second. Peiji Chen, Wenjing Ma, Srinath Mandalapu, Chandrashekhar Nagarajan, Jayavel Shanmugasundaram, Sergei Vassilvitskii, Erik Vee, Manfai Yu, Jason Y. Zien |
EC | 2 |
| 2012 | Compiler and runtime support for enabling reduction computations on heterogeneous systemsabstractSUMMARY A trend that has materialized, and has given rise to much attention, is of the increasingly heterogeneous computing platforms. Presently, it has become very common for a desktop or a notebook computer to come equipped with both a multi‐core CPU and a graphics processing unit (GPU). Capitalizing on the maximum computational power of such architectures (i.e., by simultaneously exploiting both the multi‐core CPU and the GPU), starting from a high‐level API, is a critical challenge. We believe that it would be highly desirable to support a simple way for programmers to realize the full potential of today's heterogeneous machines. This paper describes a compiler and runtime framework that can map a class of applications, namely those characterized bygeneralized reductions, to a system with a multi‐core CPU and GPU. Starting with simple C functions with added annotations, we automatically generate the middleware API code for the multi‐core, as well as CUDA code to exploit the GPU simultaneously. The runtime system provides efficient schemes for dynamically partitioning the work between CPU cores and the GPU. Our experimental results from two applications, for example, k‐means clustering and principal component analysis, show that, through effectively harnessing the heterogeneous architecture, we can achieve significantly higher performance compared with using only the GPU or the multi‐core CPU. In k‐means clustering, the heterogeneous version with eight CPU cores and a GPU achieved a speedup of about 32.09x relative to one‐thread CPU. When compared with the faster of CPU‐only and GPU‐only executions, we were able to achieve a performance gain of about 60%. In principal component analysis, the heterogeneous version attained a speedup of 10.4x relative to the one‐thread CPU version. When compared with the faster of CPU‐only and GPU‐only versions, the heterogeneous version achieved a performance gain of about 63.8%. Copyright © 2011 John Wiley & Sons, Ltd. Vignesh T. Ravi, Wenjing Ma, David Chiu 0001, Gagan Agrawal |
Concurr. Comput. Pract. Exp. | 2 |
| 2011 | Parameterized Micro-benchmarking: An Auto-tuning Approach for Complex ApplicationsabstractAuto-tuning has emerged as an important practical method for creating highly optimized code. However, the growing complexity of architectures and applications has resulted in a prohibitively large search space that preclude empirical auto-tuning. Here, we focus on the challenge to auto-tuning presented by applications that require auto-tuning of not just a small number of distinct kernels, but a large number of kernels that exhibit similar computation and memory access characteristics and require optimization over similar problem spaces. We propose an auto-tuning method for tensor contraction functions on GPUs, based on parameterized micro-benchmarks. Using our parameterized micro-benchmarking approach, we obtain a speedup of up to 2 over the version that used default optimizations without auto-tuning. Wenjing Ma, Sriram Krishnamoorthy, Gagan Agrawal |
PACT | 1 |
| 2011 | Practical Loop Transformations for Tensor Contraction Expressions on Multi-level Memory Hierarchies
Wenjing Ma, Sriram Krishnamoorthy, Gagan Agrawal |
CC | 1 |
| 2011 | An execution strategy and optimized runtime support for parallelizing irregular reductions on modern GPUsabstractGPUs have rapidly emerged as a very significant player in high performance computing. However, despite the popularity of CUDA, there are significant challenges in porting different classes of HPC applications on modern GPUs. This paper focuses on the challenges of implementing irregular applications arising from unstructured grids on modern NVIDIA GPUs. Considering the importance of irregular reductions in scientific and engineering codes, substantial effort was made in developing compiler and runtime support for parallelization or optimization of these codes in the previous two decades, with different efforts targeting distributed memory machines, distributed shared memory machines, shared memory machines, or cache performance improvement on uniprocessor machines. However, there have not been any systematic studies on parallelizing these applications on modern GPUs. There are at least two significant challenges associated with porting this class of applications on modern GPUs. The first is related to correct and efficient parallelization while using a large number of threads. The second challenge is effective use of shared memory. Since data accesses cannot be determined statically, runtime partitioning methods are needed for effectively using the shared memory. This paper describes an execution methodology that can address the above two challenges. We have also developed optimized runtime modules to support our execution methodology. Our approach and runtime methods have been extensively evaluated using two indirection array based applications. Xin Huo, Vignesh T. Ravi, Wenjing Ma, Gagan Agrawal |
ICS | 3 |
| 2010 | An integer programming framework for optimizing shared memory use on GPUsabstractGeneral purpose computing using GPUs is becoming increasingly popular, because of GPU's extremely favorable performance/price ratio. Like standard processors, GPUs also have a memory hierarchy, which must be carefully optimized for in order to achieve efficient execution. Specifically, modern NVIDIA GPUs have a very small programmable cache, referred to as shared memory, accesses to which are nearly 100 to 150 times faster than accesses to the regular device memory. An automatically generated or hand-written CUDA program can explicitly control what variables and array sections are allocated on the shared memory at any point during the execution. This, however, leads to a difficult optimization problem. Wenjing Ma, Gagan Agrawal |
PACT | 1 |
| 2010 | Pricing guaranteed contracts in online display advertisingabstractWe consider the problem of pricing guaranteed contracts in online display advertising. This problem has two key characteristics that when taken together distinguish it from related offline and online pricing problems: (1) the guaranteed contracts are sold months in advance, and at various points in time, and (2) the inventory that is sold to guaranteed contracts - user visits - is very high-dimensional, having hundreds of possible attributes, and advertisers can potentially buy any of the very large number (many trillions) of combinations of these attributes. Consequently, traditional pricing methods such as real-time or combinatorial auctions, or optimization-based pricing based on self- and cross-elasticities are not directly applicable to this problem. We hence propose a new pricing method, whereby the price of a guaranteed contract is computed based on the prices of the individual user visits that the contract is expected to get. The price of each individual user visit is in turn computed using historical sales prices that are negotiated between a sales person and an advertiser, and we propose two different variants in this context. Our evaluation using real guaranteed contracts shows that the proposed pricing method is accurate in the sense that it can effectively predict the prices of other (out-of-sample) historical contracts. Vijay Bharadwaj, Wenjing Ma, Michael Schwarz 0002, Jayavel Shanmugasundaram, Erik Vee, Jack Xie, Jian Yang 0002 |
CIKM | 2 |
| 2010 | Acceleration of Streamed Tensor Contraction Expressions on GPGPU-Based ClustersabstractTensor contractions are generalized multidimensional matrix multiplication operations that widely occur in quantum chemistry. Efficient execution of tensor contractions on GPUs requires tackling several challenges to be addressed, including index permutation and small dimension-sizes reducing thread block utilization. In this paper, we present our approach to automatically generate CUDA code to execute tensor contractions on GPUs, including management of data movement between CPU and GPU. GPU-enabled code is generated for the most expensive contractions in CCSD(T), a key coupled cluster method, and incorporated into NW Chem, a popular computational chemistry suite. We demonstrate speedup over a factor of 8.4 using one core per node and over 2.6 when utilizing the entire system using hybrid CPU+GPU solution with 2 GPUs and 5 cores. Finally, we analyze the implementation behavior on future GPU systems. Wenjing Ma, Sriram Krishnamoorthy, Oreste Villa, Karol Kowalski |
CLUSTER | 1 |
| 2010 | Approaches for parallelizing reductions on modern GPUsabstractGPU hardware and software has been evolving rapidly. CUDA versions 1.1 and higher started supporting atomic operations on device memory, and CUDA versions 1.2 and higher started supporting atomic operations on shared memory. This paper focuses on parallelizing applications involving reductions on GPUs. Prior to the availability of support for locking, these applications could only be parallelized using full replication, i.e., by creating a copy of the reduction object for each thread. However, CUDA 1.1 (1.2) onwards, use of atomic operations (on shared memory) is another option, though some effort is still required in supporting locking on floating point numbers and for supporting coarse-grained locking. Based on the tradeoffs between locking and full replication, we also introduce a hybrid approach, in which a group of threads use atomic operations to update one copy of the reduction object. Using three data mining algorithms that follow the reduction structure - k-means clustering, Principal Component Analysis (PCA) and k-nearest neighbor search (kNN), we evaluate the relative performance of these three approaches. We show how the relative performance of these techniques can vary depending upon the application and its parameters. The hybrid approach we have introduced clearly outperforms other approaches in several cases. Xin Huo, Vignesh T. Ravi, Wenjing Ma, Gagan Agrawal |
HiPC | 3 |
| 2010 | An integer programming framework for optimizing shared memory use on GPUsabstractGeneral purpose computing using GPUs is becoming increasingly popular, because of GPU's extremely favorable performance/price ratio. Besides application development using CUDA, automatic code generation for GPUs is also receiving attention. Like standard processors, GPUs also have a memory hierarchy, which must be carefully optimized for in order to achieve efficient execution. Specifically, modern NVIDIA GPUs have a very small programmable cache, referred to as shared memory, accesses to which are nearly 100 to 150 times faster than accesses to the regular device memory. An automatically generated or hand-written CUDA program can explicitly control what variables and array sections are allocated on the shared memory at any point during the execution. This, however, leads to a difficult optimization problem. In this paper, we formulate and solve the shared memory allocation problem as an integer programming problem. We present a global (intraprocedural) framework which can model structured control flow, and is not restricted to a single loop nest. We consider allocation of scalars, arrays, and array sections on shared memory. We also briefly show how our framework can suggest useful loop transformations to further improve performance. Our experiments using several non-scientific application show that our integer programming framework outperforms a recently published heuristic method, and our loop transformations also improve performance for many applications. Wenjing Ma, Gagan Agrawal |
HiPC | 1 |
| 2010 | Compiler and runtime support for enabling generalized reduction computations on heterogeneous parallel configurationsabstractA trend that has materialized, and has given rise to much attention, is of the increasingly heterogeneous computing platforms. Presently, it has become very common for a desktop or a notebook computer to come equipped with both a multi-core CPU and a GPU. Capitalizing on the maximum computational power of such architectures (i.e., by simultaneously exploiting both the multi-core CPU and the GPU) starting from a high-level API is a critical challenge. We believe that it would be highly desirable to support a simple way for programmers to realize the full potential of today's heterogeneous machines. Vignesh T. Ravi, Wenjing Ma, David Chiu 0001, Gagan Agrawal |
ICS | 2 |
| 2010 | A Light-Size AKA Mechanism for Optimal Distributed AAA authorization ArchitectureabstractAccording to the different identities of end users and attributes and roles that they have been granted, this paper presents an optimal distributed AAA authorization architecture to assign them different network resources or services. To improve the performance, this paper thus gives a detailed analysis about its security issues and proposes a light-size key agreement mechanism, including three kinds of keys in different domains and anonymous identity verification to ensure communication security between Pac and NAS in wireless network. Final tests prove the feasibility of optimal authorization mechanism and good performance of this light-size key agreement mechanism. Wenjing Ma |
VTC Spring | 1 |
| 2009 | A translation system for enabling data mining applications on GPUsabstractModern GPUs offer much computing power at a very modest cost. Even though CUDA and other related recent developments are accelerating the use of GPUs for general purpose applications, several challenges still remain in programming the GPUs. Thus, it is clearly desirable to be able to program GPUs using a higher-level interface.In this paper, we offer a solution that targets a specific class of applications, which are the data mining and scientific data analysis applications. Our work is driven by the observation that a common processing structure, that of generalized reductions, fits a large number of popular data mining algorithms. In our solution, the programmers simply need to specify the sequential reduction loop(s) with some additional information about the parameters. We use program analysis and code generation to map the applications to a GPU. Several additional optimizations are also performed by the system.We have evaluated our system using three popular data mining applications, k-means clustering, EM clustering, and Principal Component Analysis (PCA). The main observations from our experiments are as follows. The speedup that each of these applications achieve over a sequential CPU version ranges between 20 and 50. The automatically generated version did not have any noticeable overheads compared to hand written codes. Finally, the optimizations performed in the system resulted in significant performance improvements. Wenjing Ma, Gagan Agrawal |
ICS | 1 |
| 2009 | A compiler and runtime system for enabling data mining applications on gpusabstractWith increasing need for accelerating data mining and scientific data analysis on large data sets, and less chance to improve processor performance by simply increasing clock frequencies, multi-core architectures and accelerators like FPGAs and GPUs have become popular. A recent development in using GPU for general computing has been the release of CUDA (Compute Unified Device Architecture) by NVIDIA. CUDA allows GPU programming with Clanguage-like features, thus easing the development of non-graphics applications on a GPU. However, several challenges still remain in programming the GPUs with CUDA, because CUDA involves explicit parallel programming and management of its complex memory hierarchy, as well as allocating device memory, moving data between CPU anddevice memory, and specification of thread grid configurations. Wenjing Ma, Gagan Agrawal |
PPoPP | 1 |
| 2008 | An Optimization Method to Develop AAA Architectures with MIPv6 Mobility SupportabstractWith the development of mobile network and computer technology, MIPv6 is brought to the internet. Taking care of the security concerns about network connection, we bring AAA system into the mobile network. In order to be permitted in the integrated architecture of MIPv6 and AAA systems, the users have to get network access permission and AAA response from AAAH. This paper presents an optimization method to enhance handover performance. Above all, we build up a hierarchical AAA architecture and temporarily store AAA credentials at the AAASL. So that mobile user does not have the need to send request to AAAH. Then we encapsulate BU or HoT1 into authentication/authorization request information and save time needed for the BU to travel from MN to HA. Also an improved efficient security association is considered to solve the network access problem. Finally, Experiments indicate that compared with current MIPv6, this optimization method could shorten handover time, especially when the distance between MN and AAAS is long. Wenjing Ma, Yong Zhang 0025 |
APSCC | 1 |