Lihua Chi

dblp:17/9750 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
7since 2021 · last 2025
0009-0004-3216-4189ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Cross-scale hierarchical spatio-temporal transformer for video enhancement
Lihua Chi, Jie Liu 0002
Knowl. Based Syst.3
2024 Detailed Analysis and Optimization of Irregular-Shaped Matrix Multiplication on Multi-Core DSPs
abstract
Irregular-shaped General Matrix Multiplication (GEMM) is extensively used in diverse workloads, such as scientific simulation and deep learning. In response to energy efficiency constraints, low-power multi-core digital signal processors (DSPs) have emerged as a viable alternative architecture in HPC systems. This study examines the performance of existing GEMM implementations working on irregular-shaped matrices and observes sub-optimal outcomes due to deficiencies in memory optimization and core-level parallelism extraction. For multi-core DSPs in FT-M7032, a CPU-DSP heterogeneous processor for HPC, we introduce dspIMM - a new multi-core parallel implementation for irregular-shaped matrix multiplications. dspIMM incorporates a new loop ordering, stack-space optimization, multi-dimensional core-level parallelization, a communication-computation overlap implementation with data prefetch across loops, and blocking optimization. These optimizations effectively enhance memory access and the core-level parallelism in irregular-shaped GEMMs. Experimental results demonstrate that the proposed communication-computation overlap optimization achieves the highest performance improvement in dspIMM, and the average speedup of dspIMM over previous implementations finally achieves up to 3.34 times.
Haotian Mo, Linyu Liao, Biao Li 0009, Lihua Chi, Jie Liu 0002
ICPP5
2024 Implicit Neural Alignment Network for Arbitrary-scale Space-Time Video Super-Resolution
abstract
Space-time video super-resolution (STVSR) is a comprehensive task comprising two subtasks: video super resolution in space dimension and video frame interpolation in time dimension. Conventional decoupled two-stage approaches tend to overlook the intrinsic correlation between the two tasks. Overcoming this challenge requires the development of a unified model capable of simultaneously implementing space-time super-resolution across arbitrary scales. Most existing models are confined to training on fixed space upsampling scales or specific frame-rate videos, resulting in limited generalization capabilities for flexible space-time super-resolution scenarios. In response to this limitation, our approach draws inspiration from continuous implicit neural representation. We propose an enhanced Implicit Neural Alignment Network (INAN) based on the VideoINR framework, encompassing feature refinement, precise motion flow estimation, and multi-scale feature fusion to optimize the final implicit neural decoding. Our extensive experimental evaluations on multiple benchmarks underscore the efficacy of the INAN model, indicate its superior performance compared to prior STVSR methods.
Lihua Chi
IJCNN3
2024 DeepEnhancer: Temporally Consistent Focal Transformer for Comprehensive Video Enhancement
abstract
Restoring and colorizing old films is a comprehensive video enhancement task, marked by the presence of heterogeneous and structured degradations. Our DeepEnhancer addresses this challenge through a unified pipeline that combines restoration and colorization. In this workflow, we implement a bidirectional propagation strategy. Specifically, we incorporate second-order feature alignment to reduce the accumulation of inaccuracies in optical flow estimation. Simultaneously, we utilize cross-scale long-term attention mechanisms to model correlations within hidden states, thereby ensuring spatial and temporal consistency. To address the notable content loss in aging films, we introduce a temporally consistent focal transformer guided by global information. This transformer utilizes various window levels with distinct sub-window sizes to seamlessly integrate fine-grained and coarse-grained features. Comprehensive experimental results conclusively demonstrate the superiority of our model in both quantitative and qualitative comparisons when compared to existing approaches.
Lihua Chi, Wentao Ma 0003, Feng Li 0037, Jie Liu 0002
ICMR3
2024 TempDiff: Enhancing Temporal-awareness in Latent Diffusion for Real-World Video Super-Resolution
abstract
Abstract Latent diffusion models (LDMs) have demonstrated remarkable success in generative modeling. It is promising to leverage the potential of diffusion priors to enhance performance in image and video tasks. However, applying LDMs to video super‐resolution (VSR) presents significant challenges due to the high demands for realistic details and temporal consistency in generated videos, exacerbated by the inherent stochasticity in the diffusion process. In this work, we propose a novel diffusion‐based framework, Temporal‐awareness Latent Diffusion Model (TempDiff), specifically designed for real‐world video super‐resolution, where degradations are diverse and complex. TempDiff harnesses the powerful generative prior of a pre‐trained diffusion model and enhances temporal awareness through the following mechanisms: 1) Incorporating temporal layers into the denoising U‐Net and VAE‐Decoder, and fine‐tuning these added modules to maintain temporal coherency; 2) Estimating optical flow guidance using a pre‐trained flow net for latent optimization and propagation across video sequences, ensuring overall stability in the generated high‐quality video. Extensive experiments demonstrate that TempDiff achieves compelling results, outperforming state‐of‐the‐art methods on both synthetic and real‐world VSR benchmark datasets. Code will be available at https://github.com/jiangqin567/TempDiff
Lihua Chi, X. H. Chen, Q. Y. Zhang, Z. Q. Deng, J. S. Deng, B. B. Tang, S. H. Lv, Jie Liu 0002
Comput. Graph. Forum3
2023 TH-Allreduce: Optimizing Small Data Allreduce Operation on Tianhe System
abstract
Scaling up parallel applications can be challenging, especially when dealing with large volumes of data that need to be distributed across multiple nodes. In this paper, we explore the system architecture of Tianhe and propose a cutting-edge solution for optimizing global data communication. Our optimized small data blocking/non-blocking allreduce method (TH-Allreduce) is tailored for scientific applications like solving linear systems Ax=b, which often require massive data processing capabilities. To address intra-node communication challenges, we introduce the Ping-Pong Small Data Shared Memory (PPSDSM) framework, which utilizes ping-pong communication patterns to minimize Round-Trip Time (RTT) and reduce computational costs. We further present a latency-aware allreduce algorithm (PP-LA) based on PPSDSM that optimizes communication overheads and computational costs. For inter-node communication, we leverage the power of the Tianhe offloading engine to propose a topology-aware offloading allreduce method. Our experimental results show that our state-of-the-art library outperforms typical MPI implementations on different CPUs, achieving a speedup of 1.5-12x for intra-node allreduce and 1.32-3.34x for multi-node small data allreduce on the Tianhe Exascale Prototype Upgrade System at scale. These findings demonstrate that our proposed methods can significantly improve communication efficiency and scalability for distributed computing on Tianhe systems, opening up new avenues for a wide range of scientific applications.
Jintao Peng, Lihua Chi
ICPADS4
2022 A novel neural network approach for airfoil mesh quality evaluation
Xinhai Chen 0001, Chunye Gong, Jie Liu 0002, Yufei Pang, Liang Deng, Lihua Chi, Kenli Li 0001
J. Parallel Distributed Comput.6
2018 TAMM: A New Topology-Aware Mapping Method for Parallel Applications on the Tianhe-2A Supercomputer
Xinhai Chen 0001, Jie Liu 0002, Shengguo Li, Peizhen Xie, Lihua Chi
ICA3PP (1)5
2018 Customizing the HPL for China accelerator
Xinbiao Gan, Yikun Hu 0001, Jie Liu 0002, Lihua Chi, Han Xu 0008, Chunye Gong, Shengguo Li, Yihui Yan
Sci. China Inf. Sci.4
2018 An efficient SIMD compression format for sparse matrix-vector multiplication
abstract
Summary Sparse matrix‐vector multiplication (SpMV) is an essential kernel in sparse linear algebra and has been studied extensively on all modern processor and accelerator architectures. Compressed Sparse Row (CSR) is a frequently used format for sparse matrices storage. However, CSR‐based SpMV has poor performance on processors with vector units. In order to take full advantage of SIMD acceleration technology in SpMV, we proposed a new matrix storage format called CSR‐SIMD. The new storage format compresses the non‐zero elements into many variable‐length data fragments with consecutive memory access addresses. Thus, the data locality of sparse matrix A and dense vector x expands and the floating‐point operations for each fragment can be completely calculated by vectorized implementation on wide SIMD units. Our experimental results indicate that CSR‐SIMD has better storage efficiency and low‐overhead for format conversion. Besides, the new format achieves high scalability on wide SIMD units. In comparison with the CSR‐based and BCSR‐based SpMV, CSR‐SIMD obtains better performance on FT1500A, Intel Xeon, and Intel Xeon Phi.
Xinhai Chen 0001, Peizhen Xie, Lihua Chi, Jie Liu 0002, Chunye Gong
Concurr. Comput. Pract. Exp.3
2012 High-Performance Matrix Multiply on a Massively Multithreaded Fiteng1000 Processor
Jie Liu 0002, Lihua Chi, Chunye Gong, Han Xu 0008, Yihui Yan, Qingfeng Hu
ICA3PP (2)2