EDBT 2026 Demo / reviewers in the wild / expert
Zhonghua Lu
dblp:09/4333
· DBLP profile ↗
22ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WindStencil: Unleashing GPU Potential for High-Order Stencil Computation in High-Performance Inviscid CFD Simulations
Xiazhen Liu, Runfeng Jin, Jian Zhang 0070, Wu Yuan 0002, Shan Liang 0005, Zhonghua Lu |
ICS | 9 |
| 2025 | An Integrated Topology-Aware Mapping Scheme for Large-Scale CFD Applications on the ORISE Supercomputer
Lin Gao 0004, Wu Yuan 0002, Xiazhen Liu, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu |
ICA3PP (6) | 6 |
| 2025 | FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUsabstractEfficiently solving large-scale linear systems is a critical challenge in electromagnetic simulations, particularly when using the Crank-Nicolson Finite-Difference Time-Domain method. Existing iterative solvers are commonly employed to handle the resulting sparse systems but suffer from slow convergence due to the ill-conditioned nature of the double-curl operator. Approximate preconditioners, like SOR and Incomplete LU decomposition (ILU) provide insufficient convergence, while direct solvers are impractical due to excessive memory requirements. To address this, we propose FlashMP, a novel preconditioning system that designs a subdomain exact solver based on discrete transforms. FlashMP provides an efficient GPU implementation that achieves multi-GPU scalability through domain decomposition. Evaluations on AMD MI60 GPU clusters (up to$\mathbf{1 0 0 0 ~ G P U s}$) show that FlashMP reduces iteration counts by up to$16 \times$and achieves speedups of$2.5 \times$to$4.9 \times$compared to baseline implementations in state-of-the-art libraries Hypre. Weak scalability tests show parallel efficiencies up to 84.1 %. Yaqian Gao, Runfeng Jin, Yidong Chen 0014, Wu Yuan 0002, Wenpeng Ma, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu |
ICCD | 12 |
| 2025 | A Simplified Input Strategy for Predicting Multi-Type Associations in miRNA-LncRNA-Disease Network via Stacked Deep Matrix FactorizationabstractUnderstanding the associations among microRNAs (miRNAs), long non-coding RNAs (lncRNAs), and various diseases as biomarkers holds significant biological importance. Developing efficient, straightforward prediction models is essential to reduce the high cost of experimental research. However, most existing methods typically predict miRNA-disease associations (MDAs), lncRNA-disease associations (LDAs), and lncRNA miRNA interactions (LMIs) separately, often relying on both similarities and associations as inputs. These approaches complicate their application across diverse biological and medical domains. Moreover, few models are capable of simultaneously predicting all three types of associations in a unified framework. In this work, we propose a novel and simplified model, called Simplified input strategy for Multiple Associations Prediction (SimpleMAP). Unlike previous approaches, SimpleMAP eliminates the need for similarity networks or external biological data and instead uses only known associations as input, reducing feature contamination and ensuring better generalization. SimpleMAP is designed to predict MDAs, LDAs, and LMIs concurrently, by constructing a three-layer heterogeneous biomolecular network that captures the associations among miRNAs, lncRNAs, and diseases. Our method employs a single, end-to-end architecture based on stacked deep matrix factorization (SDMF) to process sparse input data and learn latent features effectively. SimpleMAP is designed to concurrently predict MDAs, LDAs, and LMIs by constructing a three-layer heterogeneous biomolecular network that captures multi-relational associations among miRNAs, lncRNAs, and diseases. To enhance predictive performance, we incorporate multiple feature integration strategies to fuse representations extracted by SDMF. This streamlined design makes SimpleMAP one of the first models to predict multiple bio-entity associations jointly using only minimal input data, offering a highly scalable and biologically meaningful solution. SimpleMAP demonstrates superior performance against strong baselines. Further validation on two additional datasets involving miRNA-circRNA-disease associations confirms the models robustness and adaptability. Finally, biologically validated case studies underscore the realworld applicability of SimpleMAP for biomarker discovery in complex biological systems. Overall, SimpleMAP introduces a new paradigm in bio-entity association predictionłachieving multi-type, high-performance prediction with minimal input complexityłmaking it a valuable tool for computational biology and biomedical research. Ning Ai, Zhonghua Lu, Yong Liang 0001, Qi Hong Lai, Loi Lei Lai, Hongmin Cai, Dong Ouyang |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | Generative Evolution Attacks Portfolio SelectionabstractIt is agreed that portfolio selection is of great importance for the financial market. Numerous outstanding exact and heuristic algorithms have been proposed in the past decades. However, their development always demands meticulous human ideas and could be time-consuming. Moreover, most of them tend to suffer from performance degradation when exposed to new portfolio selection models and different investment environments. Learning-enabled approaches have recently yielded impressive results, but these methods still grapple with challenges in model design and training. In this paper, we explore the mutual facilitation of large language models (LLMs) and huristic approaches in portfolio selection, and propose a novel LLM-based multi-objective evolutionary algorithm (MOEA) named IlmPC-NSGA-II. In this algorithm, the LLM with carefully-designed well-structured prompts serves as a straightforward yet effective engine for generating new solutions, non-dominated sorting and crowding distance calculation are adopted to enable the LLM and the evolutionary process to mutually guide toward the optimal region of the solution space. Experimental results on various scales of constrained multi-objective portfolio selection models and four benchmark problems demonstrate that our proposed approach can achieve a more competitive performance compared to widely-used MOEAs and the LLM-only method. Chen Li 0068, Jinrong Jiang, Lian Zhao, Yidi Bai, Zhonghua Lu, Xuebin Chi |
CEC | 6 |
| 2024 | MIST: Efficient Mixed-Precision Preconditioning Through Iterative Sparse- Triangular Solver DesignabstractExact sparse-triangular solvers are highly sequential and difficult to implement efficiently on GPUs with ILU preconditioning. Lower precision is crucial for reducing data movement and storage demands in memory-bound problems. However, current mixed-precision systems struggle to achieve performance gains for ILU preconditioning on multi-GPU platforms due to challenges in (1) utilizing two levels of parallelism and (2) minimizing off-chip memory bandwidth while maintaining accuracy. Additionally, these systems focus on scalar operations and lack support for point-block matrices, which arise naturally in multiphysics problems and require tailored algorithm designs. To address these challenges, we propose MIST, a novel Mixed-precision Iterative Sparse-Triangular solver optimized for GPUs to accelerate preconditioning in Krylov methods. We (1) implement an efficient mixed-precision Jacobi iterative local solver to harness single-GPU parallelism and scale it to multi- GPU via domain decomposition, and (2) design a BSpMVA kernel to reduce bandwidth while achieving high double-precision accuracy. Integrated into a widely-used numerical library, MIST offers end-to-end support for solving sparse linear systems, balancing efficiency and convergence. Experimental results show that MIST provides a 3.38× average speedup over cuSPARSE�s exact sparse-triangular solver, with an additional 1.37× speedup when using low-precision, while maintaining robustness Yidong Chen 0014, Wenpeng Ma, Wu Yuan 0002, Jian Zhang 0070, Zhonghua Lu |
ICCD | 6 |
| 2024 | MixQ: Taming Dynamic Outliers in Mixed-Precision Quantization by Online PredictionabstractMixed-precision quantization has shown to be a promising method for enhancing the efficiency of LLMs. This technique boosts computational efficiency by processing most values with low-precision, high-throughput compute units and maintains accuracy by processing outliers in high-precision. However, due to the dynamic, irregular, and sparse nature of outliers, this approach is far from using hardware efficiently. In this work, we propose MixQ, an efficient mixed-precision quantization system. Through our in-depth analysis of outlier distribution, we introduce a locality-based outlier prediction algorithm that can predict all outliers of 95.8% of tokens. Based on this accurate prediction, we propose a quantization ahead of detection (QAD) technique that can verify the correctness of prediction. A new data structure is proposed for efficient outlier processing. Evaluation shows that MixQ achieves $1.52 \times$ and $1.78 \times$ speedup over FP16 and Bitsandbytes on 8-bit quantization; plus $1.48 \times 1.93 \times$ and $6 \times$ speedup over QUIK, FP16, and AWQ on 4-bit quantization.11Our code is available on:https://github.com/Qcompiler/MIXQ Yidong Chen 0003, Chen Zhang 0001, Rongchao Dong, Zhonghua Lu, Jidong Zhai |
SC | 6 |
| 2024 | BSPADMM: block splitting proximal ADMM for sparse representation with strong scalability
Yidong Chen 0003, Jingshan Pan, Yonghong Hu, Zhonghua Lu |
CCF Trans. High Perform. Comput. | 6 |
| 2024 | Mixed-precision block incomplete sparse approximate preconditioner on Tensor core
Wenpeng Ma, Wu Yuan 0002, Jian Zhang 0070, Zhonghua Lu |
CCF Trans. High Perform. Comput. | 5 |
| 2023 | Distributed Generative Adversarial Networks for Fuzzy Portfolio Optimization
Xueying Yang, Zhonghua Lu |
ICA3PP (4) | 4 |
| 2023 | A parallel non-convex approximation framework for risk parity portfolio designabstractIn this paper, we propose a parallel non-convex approximation framework (NCAQ) for optimization problems where the objective is to minimize a convex function plus the sum of non-convex functions. Based on the structure of the objective function, our framework transforms the non-convex constraints to the logarithmic barrier function and approximates the non-convex problem by a parallel quadratic approximation scheme, which will allow the original problem to be solved by accelerated inexact gradient descent in the parallel environment. Moreover, we give a detailed convergence analysis for the proposed framework. The numerical experiments show that our framework outperforms the state-of-art approaches in terms of accuracy and computation time on the high dimension non-convex Rosenbrock test functions and the risk parity problems. In particular, we implement the proposed framework on CUDA, showing a more than 25 times speed-up ratio and removing the computational bottleneck for non-convex risk-parity portfolio design. Finally, we construct the high dimension risk parity portfolio in China’s stock market that consistently outperforms the equal weight portfolio. Yidong Chen 0003, Chen Li 0068, Yonghong Hu, Zhonghua Lu |
Parallel Comput. | 4 |
| 2022 | Computing Wasserstein-$p$ Distance Between Images with Linear CostabstractWhen the images are formulated as discrete measures, computing Wasserstein-p distance between them is challenging due to the complexity of solving the corresponding Kantorovich's problem. In this paper, we propose a novel algorithm to compute the Wasserstein-p distance between discrete measures by restricting the optimal transport (OT) problem on a subset. First, we define the restricted OT problem and prove the solution of the restricted problem converges to Kantorovich's OT solution. Second, we propose the SparseSinkhorn algorithm for the restricted problem and provide a multi-scale algorithm to estimate the subset. Finally, we implement the proposed algorithm on CUDA and illustrate the linear computational cost in terms of time and memory requirements. We compute Wasserstein-p distance, estimate the transport mapping, and transfer color between color images with size ranges from$64\times 64$to$1920\times 1200$. (Our code is available at https://github.com/ucascnic/CudaOT) Yidong Chen 0003, Chen Li 0068, Zhonghua Lu |
CVPR | 3 |
| 2021 | A Multiperiod Multiobjective Portfolio Selection Model With Fuzzy Random Returns for Large Scale Securities DataabstractIt is agreed that portfolio selection models are of great importance for the financial market. In this article, a constrained multiperiod multiobjective portfolio model is established. This model introduces several constraints to reflect the trading restrictions and quantifies future security returns by fuzzy random variables to capture fuzzy and random uncertainties in the financial market. Meanwhile, it considers terminal wealth, conditional value at risk (CVaR), and skewness as tricriteria for decision making. Obviously, the proposed model is computationally challenging. This situation gets worse when investors are interested in a larger financial market since the data they need to analyze may constitute typical big data. Whereafter, a novel intelligent hybrid algorithm is devised to solve the presented model. In this algorithm, the uncertain objectives of the model are approximated by a simulated annealing resilient back propagation (SARPROP) neural network which is trained on the data provided by fuzzy random simulation. An improved imperialist competitive algorithm, named IFMOICA, is designed to search the solution space. The intelligent hybrid algorithm is compared with the one obtained by combining NSGA-II, SARPROP neural network, and fuzzy random simulation. The results demonstrate that the proposed algorithm significantly outperforms the compared one not only in the running time but also in the quality of obtained Pareto frontier. To improve the computational efficiency and handle the large scale securities data, the algorithm is parallelized using MPI. The conducted experiments illustrate that the parallel algorithm is scalable and can solve the model with the size of securities more than 400 in an acceptable time. Chen Li 0068, Yulei Wu, Zhonghua Lu, Jue Wang 0013, Yonghong Hu |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | CNNPruner: Pruning Convolutional Neural Networks with Visual AnalyticsabstractConvolutional neural networks (CNNs) have demonstrated extraordinarily good performance in many computer vision tasks. The increasing size of CNN models, however, prevents them from being widely deployed to devices with limited computational resources, e.g., mobile/embedded devices. The emerging topic of model pruning strives to address this problem by removing less important neurons and fine-tuning the pruned networks to minimize the accuracy loss. Nevertheless, existing automated pruning solutions often rely on a numerical threshold of the pruning criteria, lacking the flexibility to optimally balance the trade-off between efficiency and accuracy. Moreover, the complicated interplay between the stages of neuron pruning and model fine-tuning makes this process opaque, and therefore becomes difficult to optimize. In this paper, we address these challenges through a visual analytics approach, named CNNPruner. It considers the importance of convolutional filters through both instability and sensitivity, and allows users to interactively create pruning plans according to a desired goal on model size or accuracy. Also, CNNPruner integrates state-of-the-art filter visualization techniques to help users understand the roles that different filters played and refine their pruning plans. Through comprehensive case studies on CNNs with real-world sizes, we validate the effectiveness of CNNPruner. Guan Li 0002, Junpeng Wang 0001, Han-Wei Shen, Kaixin Chen 0004, Guihua Shan, Zhonghua Lu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Distribution-based Particle Data Reduction for In-situ Analysis and Visualization of Large-scale N-body Cosmological SimulationsabstractCosmological N-body simulation is an important tool for scientists to study the evolution of the universe. With the increase of computing power, billions of particles of high space-time fidelity can be simulated by supercomputers. However, limited computer storage can only hold a small subset of the simulation output for analysis, which makes the understanding of the underlying cosmological phenomena difficult. To alleviate the problem, we design an in-situ data reduction method for large-scale unstructured particle data. During the data generation phase, we use a combined k-dimensional partitioning and Gaussian mixture model approach to reduce the data by utilizing probability distributions. We offer a model evaluation criterion to examine the quality of the probabilistic distribution models, which allows us to identify and improve low-quality models. After the in-situ processing, the particle data size is greatly reduced, which satisfies the requirements from the domain experts. By comparing the astronomical attributes and visualizations of the reconstructed data with the raw data, we demonstrate the effectiveness of our in-situ particle data reduction technique. Guan Li 0002, Jiayi Xu 0001, Tianchi Zhang 0003, Guihua Shan, Han-Wei Shen, Ko-Chih Wang, Shihong Liao, Zhonghua Lu |
PacificVis | 8 |
| 2020 | Digital Currency Investment Strategy Framework Based on Ranking
Chuangchuang Dai, Xueying Yang, Meikang Qiu, Xiaobing Guo, Zhonghua Lu, Beifang Niu |
ICA3PP (3) | 5 |
| 2020 | Performance Optimization for Feature Extraction Section of DeepChem
Ke Zhan, Zhonghua Lu, Yunquan Zhang |
ICA3PP (1) | 2 |
| 2009 | A Task-Based Fault-Tolerance Mechanism to Hierarchical Master/Worker with Divisible TasksabstractThe master/worker API of the ProActive middleware provides with an easy way to use framework for parallelizing embarrassingly parallel applications. However, the traditional master/worker model faces great challenges as the development of the scalability of the distributed computing. A single-layer hierarchical master/worker has been implemented as a solution to the scalability issues of the MW API. In the new framework, the mainmaster only communicates with some submasters, and each submaster manages a set of workers. A ldquobully election algorithmrdquo and an ldquoobject discovery mechanismrdquo are implemented to solve the fault-tolerance problems of the submasters. An automatic load-balancing mechanism is implemented for the hierarchical master/worker to solve divisible tasks. Moreover, an optimization has been done to make the fault-tolerance mechanism more efficient. Zhihui Dai, Fabien Viale, Xuebin Chi, Denis Caromel, Zhonghua Lu |
HPCC | 5 |
| 2009 | A Parallel Refined Block Arnoldi Algorithm for Large Unsymmetric MatricesabstractThis paper proposed a parallel refined block Arnoldi method for computing a few eigenvalues with largest or smallest real parts. The method accelerated by Chebyshev iteration is also investigated. We report some numerical results and compare the parallel refined block methods with single vector counterparts. The results show that the proposed method is more efficient than single vector counterparts. Xuebin Chi, Jinrong Jiang, Jun Liu 0059, Zhonghua Lu |
HPCC | 5 |
| 2009 | The minimal Laplacian spectral radius of trees with a given diameter
Ruifang Liu, Zhonghua Lu, Jinlong Shu |
Theor. Comput. Sci. | 2 |
| 2006 | Deploying Scientific Applications to the PRAGMA Grid Testbed: Strategies and LessonsabstractRecent advances in grid infrastructure and middleware development have enabled various types of applications in science and engineering to be deployed on the grid. The characteristics of these applications and the diverse infrastructure and middleware solutions developed, utilized or adapted by PRAGMA member institutes are summarized. The applications include those for climate modeling, computational chemistry, bioinformatics and computational genomics, remote control of instruments, and distributed databases. Many of the applications are deployed to the PRAGMA grid testbed in routine basis experiments. Strategies for deploying applications without modifications, and those taking advantage of new programming models on the grid are explored and valuable lessons learned are reported. Comprehensive end to end solutions from PRAGMA member institutes that provide important grid middleware components and generalized models of integrating applications and instruments on the grid are also described. David Abramson 0001, Amanda Lynch, Hiroshi Takemiya, Yusuke Tanimura, Susumu Date, Haruki Nakamura, Karpjoo Jeong, Suntae Hwang, Zhonghua Lu, Céline Amoreira, Kim K. Baldridge, Hurng-Chun Lee, Chi-Wei Wang, Horng-Liang Shih, Tomas E. Molina, Wilfred W. Li, Peter W. Arzberger |
CCGRID | 10 |
| 2006 | Analysis of the Bioinformatics Grid Technique Applications in China
Ang Guo, Zhonghua Lu, Yongwei Wu 0001, Xuebin Chi |
CCGRID | 3 |