Chao Li 0070

dblp:66/190-70 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-8721-4826ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FracMix: A Fractional Fourier Based Augmentation for Generalizable Person Re-Identification
abstract
Domain generalization in person re-identification (DG re-ID) aims to build models that can accurately retrieve a targeted person across cameras in arbitrary unseen domains, without access to those domains during training. Data augmentation (DA) has become a de facto solution to DG re-ID, and one line of approaches realizes this goal from a frequency-centric perspective, where the Fourier Transform (FT) is employed to mix amplitudes for novel sample generation. However, FT-based strategies are fundamentally limited by the assumption of signal stationarity, which may be inadequate to explore a broader perturbation space. To overcome this limitation, this work investigates the potential of the Fractional Fourier Transform (FrFT) for DA and proposes FracMix (Fractional Fourier based MixUp), which yields an intermediate representation that bridges spatial and frequency domains with a fractional order α, enabling the generation of richer samples. Furthermore, an Adaptive Token Pruning strategy (ATP) is designed to dynamically select the top-K salient tokens and perform perturbations exclusively to corresponding patches. This simple yet effective mechanism alleviates over-reliance on a small set of dominant patches and encourages broader contextual utilization, thereby improving robustness and generalization. Experiments on multiple DG re-ID benchmarks demonstrate the effectiveness of FracMix.
Jieru Jia, Huidi Xie, Yantao Song, Lichao Zhang 0001, Chao Li 0070
IEEE Signal Process. Lett.7
2026 GroupSD: Self-Distillation From Intermediate ViT Layers for Generalizable Person Re-Identification
Jieru Jia, Jianchao Yang, Chao Li 0070, Qiuqi Ruan
IEEE Trans. Circuits Syst. Video Technol.5
2025 GCPNet: An interpretable Generic Crystal Pattern graph neural Network for predicting material properties
Hengda Gao, Genglin Li, Chao Li 0070, Canqun Yang
Neural Networks4
2024 A Motion Trace Decomposition-based overset grid method for parallel CFD simulations with moving boundaries
abstract
The overset grid method is widely employed to solve moving boundary problems in numerical simulations. However, the heavy and inevitable communication resulting from boundary movements severely impedes the improvement of parallel efficiency. This paper proposes a Motion Trace Decomposition (MTD) method to alleviate this issue. The MTD method minimizes communication overhead between processors by decomposing sub-grids and distributing them according to the object motion trajectory, negating the need to reproduce communication areas when boundaries move. Various tests were conducted to evaluate the MTD method, incorporating diverse motion types, such as displacement and rotation. Results from experimental simulations with 1.9 × 106 grid cells indicate that the proposed method enhances the parallel efficiency of the assembly process by up to 20.35% using 72 processors. These findings showcase the significant potential of the MTD method in alleviating communication challenges associated with simulating moving boundary problems using overset grids.
Chao Li 0070, Xi Yang 0020, Tao Tang 0001, Canqun Yang
ICPP2
2023 An Improved Parallel Overset Grid Method for Fluid Simulation with Moving Boundary
abstract
The Overset Grid method is a promising computational approach for tackling the challenging moving boundary problems in Computational Fluid Dynamics (CFD) simulations. The computational efficiency and accuracy of the method are critically dependent on the effectiveness of the Overset Grid Assembly (OGA) process. However, the OGA process is plagued by unavoidable issues of load imbalance and communication overheads, which adversely impact the parallel efficiency of the method, particularly when dealing with sub-grids in motion. This paper proposes an improved parallel assembly approach as an effective alternative to address these challenges. Specifically, we introduce a Balanced Merging After Decomposition (BMAD) approach, which ensures that each processor possesses a uniform number of cells from each sub-grid after partitioning and a consistent donor search time. In addition, we deploy a fine-grained list to reduce the data transfer domain, thereby minimizing communication redundancy and cost. We validate the efficiency of our approach in the case of a moving Autonomous Underwater Vehicle (AUV). Experimental results in 3 × 106 grid cells indicate that the proposed approach reduces the parallel computational cost of the OGA process by an average of 21.9% and the speedup has increased by 23.9% with 128 processors. Additionally, it demonstrated equally effective and stable performance in tests using 6 × 106 grid cells, especially achieving the highest speedup of 55.0 with 256 processors.
Chao Li 0070, Yi Liu 0083, Canqun Yang
ICPP2
2023 A large scale parallel fluid-structure interaction computing platform for simulating structural responses to a detonation shock
abstract
Abstract Due to the intrinsic nature of multi‐physics, it is prohibitively complex to design and implement a simulation software platform for study of structural responses to a detonation shock. In this article, a partitioned fluid‐structure interaction computing platform is designed for parallel simulating structural responses to a detonation shock. The detonation and wave propagation are modeled in an open‐source multi‐component solver based on OpenFOAM and blastFoam, and the structural responses are simulated through the finite element library deal.II. To capture the interaction dynamics between the fluid and the structure, both solvers are adapted to preCICE. For improving the parallel performance of the computing platform, the inter‐solver data is exchanged by peer‐to‐peer communications and the intermediate server in conventional multi‐physics software is eliminated. Furthermore, the coupled solver with detonation support has been deployed on a computing cluster after considering the distributed data storage and load‐balancing between solvers. The 3D numerical result of structural responses to a detonation shock is presented and analyzed. On 256 processor cores, the speedup ratio of the simulations for a detonation shock reach 178.0 with 5.1 million of mesh cells and the parallel efficiency achieve 69.5%. The results demonstrate good potential of massively parallel simulations. Overall, a general‐purpose fluid‐structure interaction software platform with detonation support is proposed by integrating open source codes. And this work has important practical significance for engineering application in fields of construction blasting, mining, and so forth.
Chao Li 0070, Yi Liu 0083, Sijiang Fan, Canqun Yang
Softw. Pract. Exp.3
2022 ParallelDualSPHysics: supporting efficient parallel fluid simulations through MPI-enabled SPH method
abstract
Smoothed Particle Hydrodynamics (SPH) is a classical mesh-free particle method which has been successfully applied in the field of Computational Fluid Dynamics (CFD). Its advantages over traditional mesh-based methods have made it very popular in simulating problems involving large deformation and free-surface flow. The high computational cost of the SPH method has obstructed its vast application. A lot of research effort has been devoted to accelerating the SPH method using GPU and multi threading. However, developing efficient parallel SPH algorithms on modern high-performance computers (HPCs) remains significantly challenging, especially for simulating real-world engineering problems involving hundreds of millions of particles. In this paper, we proposed an MPI-enabled parallel SPH algorithm and developed the ParallelDualSPHysics1, an open-source software supporting efficient parallel fluid simulations. Based on an efficient domain decomposition scheme, the essential data structure and algorithms of DualSPHysics were refactored to build the parallel version. For collaborating with evenly distributed particles on a distributed-memory HPC system, the parallel particle interaction and particle update modules were introduced, which enabled the SPH solver to synchronize computations among multiple processors using MPI. In addition, the redesigned pre-processing and post-processing capabilities of the ParallelDualSPHysics supported the applications of this software in a wide range of areas. Real-life test cases with up to 120 million particles were simulated and analyzed on a modern HPC system. The results showed that the parallel efficiency of ParallelDualSPHysics exceeds 90 with up to 1024 CPU cores. It indicated that ParallelDualSPHysics has the potential for large-scale engineering applications.
Xiaokang Fan, Chao Li 0070, Kelvin K. L. Wong, Yi Liu 0083, Canqun Yang
ICPP4
2019 The Communication-Overlapped Hybrid Decomposition Parallel Algorithm for Multi-Scale Fluid Simulations
abstract
The MCDPar (Parallel algorithm for multi-scale simulations based on Mesh and BCF Decomposition) algorithm significantly reduced the execution time and improved the parallel scalability for the multi-scale fluid simulations. However, the performance bottleneck still exists for extremely large-scale parallel simulations. In this paper, we designed a communication-overlapped hybrid decomposition parallel algorithm to improve the performance of the original MCDPar on large-scale clusters. Through non-blocking communication and code scheduling, the communication overhead between the master and slave groups have been overlapped with the computation of more microscopic configuration fields for the master process. Thus the parallel efficiency and scalability of the multi-scale solver could be improved on large-scale parallel simulations. In the test case with the number of configuration fields NBCF = 1000 and mesh cells Ncell = 64000, the communication percentage between the corresponding master and slave processes is reduced by 39.71%. In the test case with NBCF = 3000 and Ncell = 64000, the time cost of the fastest execution is reduced by 31.13% using the communication-overlapped algorithm, which offers a better parallel scaling on 256 cores compared to original 128 cores.
Yi Liu 0083, Chao Li 0070, Canqun Yang, Xinbiao Gan, Peng Zhang 0061, Sijiang Fan
ICPP3