VLDB 2026 Research / reviewers in the wild / expert
Fangfang Liu 0004
dblp:44/6976-4
· DBLP profile ↗
20ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0001-7344-7493ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLRQ: Faster LLM Quantization with Flexible Low-Rank Matrix SketchingabstractTraditional post-training quantization (PTQ) is considered an effective approach to reduce model size and accelerate inference of large-scale language models (LLMs). However, existing low-rank PTQ methods require costly fine-tuning to determine a compromise rank for diverse data and layers in large models, failing to exploit their full potential. Additionally, the current SVD-based low-rank approximation compounds the computational overhead. In this work, we thoroughly analyze the varying effectiveness of low-rank approximation across different layers in representative models. Accordingly, we introduce Flexible Low-Rank Quantization (FLRQ), a novel solution designed to quickly identify the accuracy-optimal ranks and aggregate them to achieve minimal storage combinations. FLRQ comprises two powerful components, Rank1-Sketch-based Flexible Rank Selection (R1-FLR) and Best Low-rank Approximation under Clipping (BLC). R1-FLR applies the R1-Sketch with Gaussian projection for the fast low-rank approximation, enabling outlier-aware rank extraction for each layer. Meanwhile, BLC aims at minimizing the low-rank quantization error under the scaling and clipping strategy through an iterative method. FLRQ demonstrates strong effectiveness and robustness in comprehensive experiments, achieving state-of-the-art performance in both quantization quality and algorithm efficiency. Hongyaoxing Gu, Lijuan Hu, Shuzi Niu, Fangfang Liu 0004 |
AAAI | 4 |
| 2026 | Self-supervised Learning for Sparse Matrix Reordering
Fangfang Liu 0004, Shuzi Niu, Huiyuan Li 0002, Wenjia Wu |
DASFAA (6) | 3 |
| 2026 | TBF: A Tunable Blocking-and-Fusion Algorithm for Efficient GPU Symmetric Rank-2K Updates
Lijuan Hu, Xinzhe Chen, Hongyaoxing Gu, Wenjing Ma, Fangfang Liu 0004 |
Euro-Par (1) | 6 |
| 2025 | Implementation and optimization of batch 3D FFT for sunway many-core processor
Mian Huo, Fangfang Liu 0004 |
CCF Trans. High Perform. Comput. | 3 |
| 2023 | GFFT: a Task Graph Based Fast Fourier Transform Optimization FrameworkabstractFast Fourier Transform (FFT) is a widely used mathematical tool in scientific and engineering applications, and optimizing its performance remains a challenging problem. This paper introduces GFFT, a novel task-graph-based FFT optimization framework that leverages modern hardware and software techniques to achieve high-performance computation. GFFT features a tuning model that uses hardware parameters to optimize FFT decomposition, a bi-directional recursive FFT algorithm that avoids strided load in SIMD implementation, and several graph optimizers inspired by deep learning frameworks to enhance performance. In addition, GFFT utilizes task-based parallelism to exploit performance on multi-core processors and provide potential compatibility with heterogeneous systems. Experimental results demonstrate that GFFT outperforms popular FFT frameworks, achieving an average speedup of 1.17x to FFTW and 1.27x to oneMKL on the Intel Xeon processor, 1.18x to AOCL-FFTW on the AMD EPYC processor, and 2.11x to FFTW on the Sunway multi-core processor with a single thread. Additionally, GFFT achieves an average speedup of 11.48x to FFTW and 1.41x to oneMKL on the Intel Xeon processor, 9.87x to AOCL-FFTW on the AMD EPYC processor with 16-threads. Qinglin Lu, Wenjing Ma, Daokun Chen, Fangfang Liu 0004 |
ICPP | 6 |
| 2023 | xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002 |
CCF Trans. High Perform. Comput. | 1 |
| 2023 | Publisher Correction: xMath2.0: a high-performance extended math library for SW26010-Pro many-core processor
Fangfang Liu 0004, Wenjing Ma, Daokun Chen, Qinglin Lu, Wanwang Yin, Xinhui Yuan, Lijuan Jiang, Hongsen Wang, Chao Yang 0002 |
CCF Trans. High Perform. Comput. | 1 |
| 2023 | An Optimized Framework for Matrix Factorization on the New Sunway Many-core PlatformabstractMatrix factorization functions are used in many areas and often play an important role in the overall performance of the applications. In the LAPACK library, matrix factorization functions are implemented with blocked factorization algorithm, shifting most of the workload to the high-performance Level-3 BLAS functions. But the non-blocked part, the panel factorization, becomes the performance bottleneck, especially for small- and medium-size matrices that are the common cases in many real applications. On the new Sunway many-core platform, the performance bottleneck of panel factorization can be alleviated by keeping the panel in the LDM for the panel factorization. Therefore, we propose a new framework for implementing matrix factorization functions on the new Sunway many-core platform, facilitating the in-LDM panel factorization. The framework provides a template class with wrapper functions, which integrates inter-CPE communication for the Level-1 and Level-2 BLAS functions with flexible interfaces and can accommodate different partitioning schemes. With the framework, writing panel factorization code with data residing in the LDM space can be done with much higher productivity. We implemented three functions ( dgetrf , dgeqrf , and dpotrf ) based on the framework and compared our work with a CPE_BLAS version, which uses the original LAPACK implementation linked with optimized BLAS library that runs on the CPE mesh. Using the most favorable partitioning, the panel factorization part achieves speedup of up to 26.3, 19.1, and 18.2 for the three matrix factorization functions. For the whole function, our implementation is based on a carefully tuned recursion framework, and we added specific optimization to some subroutines used in the factorization functions. Overall, we obtained average speedup of 9.76 on dgetrf , 10.12 on dgeqrf , and 4.16 on dpotrf , compared to the CPE_BLAS version. Based on the current template class, our work can be extended to support more categories of linear algebra functions. Wenjing Ma, Fangfang Liu 0004, Daokun Chen, Qinglin Lu, Hongsen Wang, Xinhui Yuan |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | MFFT: A GPU Accelerated Highly Efficient Mixed-Precision Large-Scale FFT FrameworkabstractFast Fourier transform (FFT) is widely used in computing applications in large-scale parallel programs, and data communication is the main performance bottleneck of FFT and seriously affects its parallel efficiency. To tackle this problem, we propose a new large-scale FFT framework, MFFT, which optimizes parallel FFT with a new mixed-precision optimization technique, adopting the “high precision computation, low precision communication” strategy. To enable “low precision communication”, we propose a shared-exponent floating-point number compression technique, which reduces the volume of data communication, while maintaining higher accuracy. In addition, we apply a two-phase normalization technique to further reduce the round-off error. Based on the mixed-precision MFFT framework, we apply several optimization techniques to improve the performance, such as streaming of GPU kernels, MPI message combination, kernel optimization, and memory optimization. We evaluate MFFT on a system with 4,096 GPUs. The results show that shared-exponent MFFT is 1.23 × faster than that of double-precision MFFT on average, and double-precision MFFT achieves performance 3.53× and 9.48× on average higher than open source library 2Decomp&FFT (CPU-based version) and heFFTe (AMD GPU-based version), respectively. The parallel efficiency of double-precision MFFT increased from 53.2% to 78.1% compared with 2Decomp&FFT, and shared-exponent MFFT further increases the parallel efficiency to 83.8%. Fangfang Liu 0004, Wenjing Ma, Huiyuan Li 0002, Yuanchi Peng |
ACM Trans. Archit. Code Optim. | 2 |
| 2018 | Performance Optimization of the HPCG Benchmark on the Sunway TaihuLight SupercomputerabstractIn this article, we present some key techniques for optimizing HPCG on Sunway TaihuLight and demonstrate how to achieve high performance in memory-bound applications by exploiting specific characteristics of the hardware architecture. In particular, we utilize a block multicoloring approach for parallelization and propose methods such as requirement-based data mapping and customized gather collective to enhance the effective memory bandwidth. Experiments indicate that the optimized HPCG code can sustain 77% of the theoretical memory bandwidth and scale to the full system of more than 10 million cores, with an aggregated performance of 480.8 Tflop/s and a weak scaling efficiency of 87.3%. Yulong Ao, Chao Yang 0002, Fangfang Liu 0004, Wanwang Yin, Lijuan Jiang, Qiao Sun 0005 |
ACM Trans. Archit. Code Optim. | 3 |
| 2017 | Towards Highly Efficient DGEMM on the Emerging SW26010 Many-Core ProcessorabstractThe matrix-matrix multiplication is an essential building block that can be found in various scientific and engineering applications. High-performance implementations of the matrix-matrix multiplication on state-of-the-art processors may be of great importance for both the vendors and the users. In this paper, we present a detailed methodology of implementing and optimizing the double-precision general format matrix-matrix multiplication (DGEMM) kernel on the emerging SW26010 processor, which is used to build the Sunway TaihuLight supercomputer. We propose a three level blocking algorithm to orchestrate data on the memory hierarchy and expose parallelism on different hardware levels, and design a collective data sharing scheme by using the register communication mechanism to exchange data efficiently among different cores. On top of those, further optimizations are done based on a data-thread mapping method for efficient data distribution, a double buffering scheme for asynchronous DMA data transfer, and an instruction scheduling method for maximizing the pipeline usage. Experiment results show that the proposed DGEMM implementation can fully exploit the unique hardware features provided by SW26010 and can sustain up to 95% of the peak performance. Lijuan Jiang, Chao Yang 0002, Yulong Ao, Wanwang Yin, Wenjing Ma, Qiao Sun 0005, Fangfang Liu 0004, Rongfen Lin |
ICPP | 7 |
| 2017 | 26 PFLOPS Stencil Computations for Atmospheric Modeling on Sunway TaihuLightabstractStencil computation arises from a broad set of scientific and engineering applications and often plays a critical role in the performance of extreme-scale simulations. Due to the memory bound nature, it is a challenging task to opti- mize stencil computation kernels on modern supercomputers with relatively high computing throughput whilst relatively low data-moving capability. This work serves as a demon- stration on the details of the algorithms, implementations and optimizations of a real-world stencil computation in 3D nonhydrostatic atmospheric modeling on the newly announced Sunway TaihuLight supercomputer. At the algorithm level, we present a computation-communication overlapping technique to reduce the inter-process communication overhead, a locality- aware blocking method to fully exploit on-chip parallelism with enhanced data locality, and a collaborative data accessing scheme for sharing data among different threads. In addition, a variety of effective hardware specific implementation and optimization strategies on both the process- and thread-level, from the fine-grained data management to the data layout transformation, are developed to further improve the per- formance. Our experiments demonstrate that a single-process many-core speedup of as high as 170x can be achieved by using the proposed algorithm and optimization strategies. The code scales well to millions of cores in terms of strong scalability. And for the weak-scaling tests, the code can scale in a nearly ideal way to the full system scale of more than 10 million cores, sustaining 25.96 PFLOPS in double precision, which is 20% of the peak performance. Yulong Ao, Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Fangfang Liu 0004, Lin Gan 0001, Wenjing Ma |
IPDPS | 6 |
| 2016 | Fast Parallel Stream Compaction for IA-Based Multi/many-core ProcessorsabstractStream compaction, frequently found in a large variety of applications, serves as a general primitive that reduces an input stream to a subset containing only the wanted elements so that the follow-on computation can be done efficiently. In this paper, we propose a fast parallel stream compaction for IA-based multi-/many-core processors. Unlike the previously studied algorithms that depend heavily on a black-box parallel scan, we open the black-box in the proposed algorithm and manually tailor it so that both the workload and the memory footprint is significantly reduced. By further eliminating the conditional statements and applying automatic code generation/optimization for performance-critical kernels, the proposed parallel stream compaction achieves high performance in different cases and for various data types across different IA-based multi/manycore platforms. Experimental results on three typical IA-based processors, including a quad-core Core-i7 CPU, a dual-socket 8-core Xeon CPU, and a 61-core Xeon Phi accelerator show that the proposed implementation outperforms the referenced parallel counterpart in the state-of-art library Thrust. On top of the above, we apply it in the random forest based data classifier to show its potential to boost the performance of real-world applications. Qiao Sun 0005, Chao Yang 0002, Changmao Wu, Leisheng Li, Fangfang Liu 0004 |
CCGrid | 5 |
| 2016 | Accelerating the Simulation of Thermal Convection in the Earth's Outer Core on Tianhe-2abstractNumerical simulation of thermal convection in the Earth's outer core requires extreme-scale computing due to the large temporal and spatial disparity, extreme physical parameters, rapid rotation and spherical geometry. In this work, the numerical simulation of the thermal convection in the Earth's outer core for CPU-MIC heterogeneous many-core systems is studied. Firstly, starting from a legacy parallel code based on the PETSc software package, a framework of the numerical simulation built on CPU-MIC heterogeneous many-core systems has been developed. Secondly, a sparse linear solver for CPUMIC heterogeneous many-core systems, which focuses on solving the two linear systems of the simulation, is presented and optimized. Thirdly, some computational kernels of the simulation, including sparse matrix-vector multiplication (SpMV) and polynomial preconditioner on distributed memory Xeon Phiaccelerated systems are implemented and optimized. In addition, in order to reduce the cost of data movement, we use methods to minimize the memory access, the PCI-E data transfer, and the MPI communication. Finally, some optimized measures are taken to the extended code. Experiments on Tianhe-2 Supercomputer show that as compared to the original code, our Xeon Phiaccelerated design is able to deliver 6.93x and 6.00x speedups for single MIC device and 64 MIC devices, respectively. Changmao Wu, Fangfang Liu 0004, Chao Yang 0002, Ligang Li, Yutong Lu, Leisheng Li, Yunfei Du 0001 |
ICPADS | 2 |
| 2016 | 10M-core scalable fully-implicit solver for nonhydrostatic atmospheric dynamicsabstractAn ultra-scalable fully-implicit solver is developed for stiff time-dependent problems arising from the hyperbolic conservation laws in nonhydrostatic atmospheric dynamics. In the solver, we propose a highly efficient hybrid domain-decomposed multigrid preconditioner that can greatly accelerate the convergence rate at the extreme scale. For solving the overlapped subdomain problems, a geometry-based pipelined incomplete LU factorization method is designed to further exploit the on-chip fine-grained concurrency. We perform systematic optimizations on different hardware levels to achieve best utilization of the heterogeneous computing units and substantial reduction of data movement cost. The fully-implicit solver successfully scales to the entire system of the Sunway TaihuLight supercomputer with over 10.5M heterogeneous cores, sustaining an aggregate performance of 7.95 PFLOPS in double-precision, and enables fast and accurate atmospheric simulations at the 488-m horizontal resolution (over 770 billion unknowns) with 0.07 simulated-years-per-day. This is, to our knowledge, the largest fully-implicit simulation to date. Chao Yang 0002, Wei Xue 0003, Haohuan Fu, Hongtao You, Yulong Ao, Fangfang Liu 0004, Lin Gan 0001, Lanning Wang, Guangwen Yang 0002 |
SC | 7 |
| 2016 | The Sunway TaihuLight supercomputer: system and applications
Haohuan Fu, Junfeng Liao, Jinzhe Yang, Lanning Wang, Zhenya Song, Xiaomeng Huang, Chao Yang 0002, Wei Xue 0003, Fangfang Liu 0004, Fangli Qiao, Xunqiang Yin, Chaofeng Hou, Jian Zhang 0070, Yangang Wang 0002, Chunbo Zhou, Guangwen Yang 0002 |
Sci. China Inf. Sci. | 9 |
| 2015 | Performance Evaluation of HPGMG on Tianhe-2: Early Experience
Yulong Ao, Yiqung Liu 0005, Chao Yang 0002, Fangfang Liu 0004, Yutong Lu, Yunfei Du 0001 |
ICA3PP (4) | 4 |
| 2015 | Pattern-Driven Hybrid Multi- and Many-Core Acceleration in the MPAS Shallow-Water ModelabstractThere is an urgent demand in studying efficient methodologies to enable hybrid multi- and many-core accelerations in global climate simulations. The Model for Prediction Across Scales (MPAS) is a family of earth-system component models that receives increasingly more attention. Like many other models, MPAS, though features some emerging numerical algorithms, employs a pure MPI approach for parallel computing, which, to date, is in lack of support for multi-threaded parallelism, especially on many-core accelerated systems. In this work, we extend the shallow-water model in MPAS to demonstrate a pattern-driven approach for hybrid multi- and many-core accelerations of climate models. We first identify all basic computation patterns through a rigorous analysis of the MPAS code. Then for the whole model, we use the identified patterns as building blocks to draw a data-flow diagram, which serves as a perfect indicator to recognize data dependencies and exploit inherent parallelism. And finally, based on the data-flow diagram, a hybrid algorithm is designed to support concurrent computations done on both multi-core CPUs and many-core accelerators. We implement the algorithm and optimize it on an x86-based heterogeneous supercomputer equipped with both Intel Xeon CPUs and Intel Xeon Phi devices. Experiments show that our hybrid design is able to deliver an 8.35x speedup as compared to the original code and scales up to 64 processes with a nearly ideal parallel efficiency. Yulong Ao, Chao Yang 0002, Yiqung Liu 0005, Fangfang Liu 0004, Changmao Wu |
ICPP | 5 |
| 2014 | Optimizing and Scaling HPCG on Tianhe-2: Early Experience
Xianyi Zhang, Chao Yang 0002, Fangfang Liu 0004, Yiqung Liu 0005, Yutong Lu |
ICA3PP (1) | 3 |
| 2014 | Accelerating HPCG on Tianhe-2: A hybrid CPU-MIC algorithmabstractIn this paper, we propose a hybrid algorithm to enable and accelerate the High Performance Conjugate Gradient (HPCG) benchmark on a heterogeneous node with an arbitrary number of accelerators. In the hybrid algorithm, each subdomain is assigned to a node after a three-dimensional domain decomposition. The subdomain is further divided to several regular inner blocks and an outer part with a flexible inner-outer partitioning strategy. Each inner task is assigned to a MIC device and the size is adjustable to adapt the accelerator's computational power. The only outer part is assigned to CPU and the thickness of boundary size is also adjustable to maintain load balance between CPU and MICs. By properly fusing the computational kernels with preceding ones, we present an asynchronous data transfer scheme to better overlap local computation with the PCI-express data transfer. All basic HPCG kernels, especially the time-consuming sparse matrix-vector multiplication (SpMV) and the symmetric Gauss-Seidel relaxation (SymGS), are extensively optimized for both CPU and MIC, on both algorithmic and architectural levels. On a single node of Tianhe-2 which is composed of an Intel Xeon processor and three Intel Xeon Phi coprocessors, we successfully obtain an aggregated performance of 50.2 Gflops, which is around 1.5% of the peak performance. Yiqung Liu 0005, Xianyi Zhang, Chao Yang 0002, Fangfang Liu 0004, Yutong Lu |
ICPADS | 4 |