VLDB 2026 Research / reviewers in the wild / expert
Bingchao Li
dblp:166/7170
· DBLP profile ↗
14ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-8629-6265ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Systems, architecture and hardware · 5 · 5 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unlabeled Data-Driven Airport Pavement Crack Segmentation Algorithm
Bingchao Li, Tongyu Shi, Nansha Li |
ICIC (21) | 1 |
| 2025 | BASH: A Bandwidth Sharing Mechanism for Subpartition Scheduling in GPU L2 CachesabstractModern GPUs employ a multi-level cache hierarchy where several requests missing in L1 Caches can be redirected to L2 Cache simultaneously. Therefore, the memory space of L2 Cache is divided into multiple partitions. A partition corresponding to a DRAM channel is composed of two subpartitions, each of which featuring a L2 Cache bank with a 32-byte/cycle data port. Given that the NoC datapath between L2 and L1 Caches also has a 32-byte bandwidth. Therefore, a request associating 128-byte data that are obtained from the L2 cache-line requires four cycles for completion, leading to bandwidth fragmentation and low resource utilization when request distribution is uneven. To address this issue, we propose the BASH architecture, which aggregates four subpartition data ports into a logical 128-byte channel and incorporates intra-group round-robin scheduling. This enables full processing of a request data transmission within a single cycle for a subpartition. Without modifying the existing model, our solution significantly improves L2 bandwidth utilization and system performance. Experimental results demonstrate that BASH can increase data port utilization by 15.7% compared to the baseline GPU, with an average performance improvement of 23.1%. Bingchao Li, Jizeng Wei |
HPCC | 1 |
| 2025 | Spatiotemporal Optimization of GPR Full Waveform Inversion Based on Super-Resolution TechnologyabstractTheoretical advancements in full waveform inversion (FWI) of ground-penetrating radar (GPR) data have shown promising potential for enhancing the accuracy of GPR data interpretation. However, the widespread implementation of FWI faces significant challenges due to its low-computational efficiency and high memory consumption, primarily attributed to the gradient operation stage. To address these issues, we propose a spatiotemporal optimization approach for GPR FWI based on super-resolution (SR) technology. The proposed method focuses on three optimization directions: adopting a storage strategy that only preserves the forward wavefield while synchronizing the gradient operation and adjoint wavefield operation, compressing the time dimension of the GPR wavefield based on the Nyquist sampling law, and obtaining a fuzzy gradient in the spatial dimension by sampling the wavefield at each moment and restoring it using an SR network to complete the FWI. Experimental results demonstrate that the proposed optimization method achieves a nearly 50% acceleration in computational efficiency without compromising the original inversion architecture. Moreover, it reduces the memory usage to approximately 4.17% of the original memory, while maintaining the effectiveness of the inversion process. This method exhibits practicality and effectiveness through several numerical and measured data experiments, providing a solid foundation for the widespread application of FWI on commonly available microcomputers. Xun Wang 0011, Tianxiao Yu, Deshan Feng, Bingchao Li, Siyuan Ding |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | A Novel Proactive Fault Tolerance Loss Function for Crack SegmentationabstractOptimizing the generalization performance of road surface crack models in practical applications represents a challenging task. Especially for thin and irregular cracks with random and expansive topologies, the loss functions used in current deep learning-based crack segmentation models are sensitive to single pixels, which tends to cause the model to overfit the training data, diminishing its generalization ability in real-world scenarios. Therefore, we take the loss function as a starting point and explore the introduction of proactive fault tolerance mechanisms into the training process of the crack segmentation model, which is called Proactive Fault Tolerance Loss (PFT Loss), to enhance the generalization capability of model in actual applications. Specifically, the PFT Loss function establishes correlations between the segmentation prediction pixels and the corresponding labeled pixels within the neighborhood window using Markov Random Fields (MRFs). The correlation is used as a reference for predicting relative shifts in segmented pixels. Proactive Fault Tolerance is performed on the loss between labeling and prediction to achieve a more natural and adaptive training method for crack segmentation. Full experiments are conducted on five public crack datasets and one self-constructed dataset. The experimental results indicate that the model trained with PFT Loss has better segmentation performance compared to other loss functions. Bingchao Li, Jianping Zong, Huaichao Wang, Nansha Li, Haifeng Li 0008 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Pseudo-Cache: Extending the Access Scope of Requests with Global Perspective in GPUsabstractGPUs are comprised of numerous streaming multiprocessors (SMs) tailored for high performance computing. SMs incorporate private L1 caches to facilitate swift data access for thousands of threads running concurrently inside SMs. Requests that miss in L1 caches are directed to L2 caches that are shared by all SMs to retrieve the desired data, which takes much longer time compared with the delay of accessing L1 cache. Due to the locality among tasks running on SMs, the same data block might be accessed by requests from various SMs, leading to data replication across multiple L1 caches. Notably, data replication is prevalent in applications exhibiting high data locality. In order to capitalize on the data replications among L1 caches, we propose routing requests destined for L2 cache to other L1 caches that are predicted to contain the desired data, extending the access scope of requests. Consequently, we introduce a Pseudo-Cache positioned adjacent to the network on chip on L2 cache side, offering a global perspective of all SMs to manage these predictions effectively. Moreover, we implement the data path for request forwarding in a cost-effective manner, leveraging the existing network structure. Experimental results underscore the efficacy of our approach, showcasing an average 16.3% performance enhancement for applications characterized by substantial data replication on GPUs. Bingchao Li, Jizeng Wei |
HPCC | 1 |
| 2024 | Efficient Common Offset Ground Penetrating Radar Reverse Time Migration Based on Finite Domain and Optimized Multitraces Cross Correlation WindowabstractReverse time migration (RTM) is an important technology for imaging ground penetrating radar (GPR) data. To address the problem of artifacts flooding of imaging results and high memory consumption of RTM, we propose an optimized multitraces cross correlation window (MCW) to increase the order of magnitude difference between the signals and artifacts for more obvious separation effect, but it also exacerbates the problem of computational cost. With the high sampling rate and high efficiency of collection method, common offset GPR is convenient to acquire large amounts of data, which consumes more numerous cost for RTM. Due to the attenuation property of high-frequency radar waves, most of the signals of common offset GPR originate from a small region below the antenna. Inspired by the footprint in airborne electromagnetic method, we propose the finite domain (FD) strategy, which limits the calculation of single trace to FD, and combine it with optimized MCW. It can reduce the computational cost of RTM and MCW significantly at the same time, especially for long profile data. Numerical experiments show that the FD reduces the computation by 77.22% with speedup 11.01. The optimized MCW retains the effective information separated from artifacts. The migration of the measured data proves the advantages and practicality of this method in engineering practical exploration. Deshan Feng, Zhengyang Fang, Xun Wang 0011, Tianxiao Yu, Siyuan Ding, Bingchao Li |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | GPR Least-Squares Reverse Time Migration Based on the Improved Cross Correlation WindowabstractGround penetrating radar (GPR) migration is a crucial imaging method to obtain the spatial position, size, and shape of the underground structures. However, Kirchhoff migration, finite-difference migration, F-K migration, and reverse time migration (RTM) focus on geometric structure imaging and cannot provide realistic reflection coefficients. Least-squares reverse time migration (LSRTM) regards imaging as an inversion problem in the sense of least squares. It continuously corrects the imaging results by minimizing the residual between the simulated data and the observed data to obtain realistic reflection coefficients. In order to enhance the accuracy of the LSRTM result, we introduce the cross-correlation window to suppress artifacts and noise. Although the non-interface information in the gradient is effectively suppressed, the cross-correlation window will cause new noise to appear. This makes the LSRTM result unsatisfactory because the window is used multiple times in the calculation. Therefore, we proposed the improved cross-correlation window that utilizes the Block-matching and 3D Filtering (BM3D). This improvement preserves the ability of eliminating artifacts while preventing the window from introducing new noise. Experiments results with the synthetic data and the measured data demonstrate that compared with the traditional methods, the LSRTM based on the improved cross-correlation window suppresses noise, reduces artifacts, enhances clarity of the interfaces, and achieves higher imaging accuracy. Deshan Feng, Bingchao Li, Xun Wang 0011, Xiaoyong Tai, Tianxiao Yu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Reverse Time Migration of Ground Penetrating Radar With Optimized Full Wavefield Separation Based on Poynting Vector Imaging Condition and TV-L1-Based Artifacts SuppressionabstractReverse time migration (RTM) has the advantage of high-precision imaging, and it can converge the radar wave back to its actual position, making it widely used in radar exploration. However, there are artifacts, low-frequency noise and fuzzy deep imaging in RTM results. Researchers have proposed full wavefield separation imaging condition and total variation (TV) technique, both of which could suppress noise and artifacts. However, the original wavefield separation method was considerably limited by its extensive calculation, and it cannot solve the problem of weak energy of imaging in the deep zone; the conventional TV technique was likely to be affected by artifacts due to the inevitable over-smoothing-suppression of anomaly edges. To address these issues, this paper improves the RTM methodology by combining an optimized full wavefield separation based on Poynting vector imaging condition and TV-L1 based artifacts suppressing technique. Specifically, the physical significance of the Poynting vector is introduced to separate the wavefield for reducing the calculation burden; the compensation function is integrated with the imaging condition to compensate for the deep energy; the TV-L1 based artifacts suppressing method is used to resolve the imaging problem of loss of specific and edge details. Synthetic data and laboratory data experiments are carried out to verify the effectiveness and practicability of the proposed RTM methodology. Deshan Feng, Zheng Feng, Xun Wang 0011, Deru Xu, Bingchao Li, Tianxiao Yu, Siyuan Ding |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | An Efficient Dual-Parameter Full Waveform Inversion for GPR Data Using Data EncodingabstractGround penetrating radar (GPR) is an important shallow electromagnetic non-destructive detection technology. The full waveform inversion (FWI) of GPR data utilizes all information including dynamics and kinematics, theoretically has the highest imaging accuracy, and meets the increasingly sophisticated needs of engineering exploration imaging. However, the bottleneck restricting the FWI is the low calculation efficiency, which cannot meet the requirements of rapid reconstruction of underground medium in actual engineering. In order to improve the calculation efficiency, we introduce the data encoding into the GPR dual-parameter FWI. Data encoding often brings crosstalk noise, and the noise is closely related to the encoding methods and data types. For this reason, we select the encoding of the crosshole data, wide-angle reflection and refraction data, and common-offset data for inversion. Experiments show that data encoding can effectively reduce computing time, and three different GPR data require different encoding methods due to their different redundancies. Total variation (TV) regularization can suppress the noise caused by data encoding. Although it will slightly increase the calculation time, it can significantly improve the inversion quality. Deshan Feng, Bingchao Li, Xun Wang 0011, Siyuan Ding, Xiaoyong Tai, Liqiong Cai, Xuan Su |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Multiparameter Elastic Full Waveform Inversion Based on Random Source-Encoding and Projection RegularizationabstractMulti-parameter elastic full waveform inversion (FWI) makes full use of the dynamic and kinematic information of all seismic wavefield. Through the mutual constraint and verification of the three parameters of P-wave velocity, S-wave velocity, and density, the joint evaluation is carried out, which is helpful to understand the structural and lithologic information of underground media more comprehensively. The bottleneck restricting the multi-parameter FWI is the large amount of calculation and low efficiency. To improve this problem, multiple shots are directly superimposed to form super shots. While it usually results in an unstable inversion due to that a large amount of crosstalk noise will be easily generated between adjacent shots. In this paper, we introduce the random source-encoding strategy to improve the inversion efficiency and load the total-variation (TV) regularization term to suppress the crosstalk noise, but it also brings the problem of regularization parameters selection for multi-parameter FWI. Thus, the projection method is applied to directly load the regularization term into the model as a constraint, which avoids the unsatisfactory results caused by the improper selection of regularization parameters and effectively improves the ill-posedness of inversion. Finally, three examples of the graben, the 1994BP, and the overthrust model are used to prove that the proposed algorithm based on random source-encoding and projection regularization can effectively improve the inversion efficiency, suppress noise, and has good practicability and adaptability. Deshan Feng, Bingchao Li, Xun Wang 0011, Deru Xu, Cen Cao, Tianxiao Yu, Zheng Feng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Inspection and Imaging of Tree Trunk Defects Using GPR Multifrequency Full-Waveform Dual-Parameter InversionabstractGround-penetrating radar (GPR) has been regarded as a potentially efficient way of evaluating the growth status of trees and preventing deterioration associated with trunk defects. The majority of current GPR data inversions, however, focused on imaging the macroscale location of defects. As the first attempt to seek a preferable quantitative inversion methodology for specifying tree protection and remedies, this article proposes a full-waveform inversion (FWI) approach involving dual-parameter attributes applied to common-offset GPR data from a commercial antenna. Specifically, the synchronous inversion of both dielectric constant and conductivity improves the identification accuracy of certain defect types. In particular, both a multifrequency strategy and total-variation (TV) regularization are seamlessly introduced to assure inversion stability by overcoming local minima and cycle skipping. Through an irregular trunk model test, the effectiveness of the optimized inversion is initially verified by presenting the precise features of the crack, hollow, and decay with the dual-parameter inversion results. In addition, several other synthetic trunk models and in-site trunk model tests further demonstrate the robustness and practicability of the proposed algorithm, which can offer more specific and comprehensive guidance for the formulation of tree protection and restoration measures. Deshan Feng, Xun Wang 0011, Bin Zhang 0034, Siyuan Ding, Tianxiao Yu, Bingchao Li, Zheng Feng |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | REMOC: efficient request managements for on-chip memories of GPUsabstractThe on-chip memories of GPUs, including the register file, shared memory and L1 cache, can provide high bandwidth and low latency access for the temporary storage of data. The capacity of L1 cache can be increased by using the registers/shared memory that are unassigned to any warps/thread blocks or released after warps/thread blocks are finished as cache-lines. In this paper, we propose two techniques to manage requests for on-chip memories to improve the efficiency of L1 cache on the base of leveraging registers and shared memory as cache-lines. Specifically, we develop a data transferring policy which is triggered when cache-lines are recalled by the first register or shared memory accesses of warps that are newly launched to prevent the data locality from being destroyed. Additionally, we design a parallel issue scheme by exploring the parallel feature of requests of an instruction accessing the register file, shared memory and L1 cache to decrease the processing latency and hence increase the throughput of instructions. The experimental results demonstrate that our approach improves the performance by 15% over prior work. Bingchao Li, Jizeng Wei |
CF | 1 |
| 2019 | An Efficient GPU Cache Architecture for Applications with Irregular Memory Access PatternsabstractGPUs provide high-bandwidth/low-latency on-chip shared memory and L1 cache to efficiently service a large number of concurrent memory requests. Specifically, concurrent memory requests accessing contiguous memory space are coalesced into warp-wide accesses. To support such large accesses to L1 cache with low latency, the size of L1 cache line is no smaller than that of warp-wide accesses. However, such L1 cache architecture cannot always be efficiently utilized when applications generate many memory requests with irregular access patterns especially due to branch and memory divergences that make requests uncoalesced and small. Furthermore, unlike L1 cache, the shared memory of GPUs is not often used in many applications, which essentially depends on programmers. In this article, we propose Elastic-Cache, which can efficiently support both fine- and coarse-grained L1 cache line management for applications with both regular and irregular memory access patterns to improve the L1 cache efficiency. Specifically, it can store 32- or 64-byte words in non-contiguous memory space to a single 128-byte cache line. Furthermore, it neither requires an extra memory structure nor reduces the capacity of L1 cache for tag storage, since it stores auxiliary tags for fine-grained L1 cache line managements in the shared memory space that is not fully used in many applications. To improve the bandwidth utilization of L1 cache with Elastic-Cache for fine-grained accesses, we further propose Elastic-Plus to issue 32-byte memory requests in parallel, which can reduce the processing latency of memory instructions and improve the throughput of GPUs. Our experiment result shows that Elastic-Cache improves the geometric-mean performance of applications with irregular memory access patterns by 104% without degrading the performance of applications with regular memory access patterns. Elastic-Plus outperforms Elastic-Cache and improves the performance of applications with irregular memory access patterns by 131%. Bingchao Li, Jizeng Wei, Murali Annavaram, Nam Sung Kim |
ACM Trans. Archit. Code Optim. | 1 |
| 2017 | Elastic-Cache: GPU Cache Architecture for Efficient Fine- and Coarse-Grained Cache-Line ManagementabstractGPUs provide high-bandwidth/low-latency on-chip shared memory and L1 cache to efficiently service a large number of concurrent memory requests (to contiguous memory space). To support warp-wide accesses to L1 cache, GPU L1 cache lines are very wide. However, such L1 cache architecture cannot always be efficiently utilized when applications generate many memory requests with irregular access patterns especially due to branch and memory divergences. In this paper, we propose Elastic-Cache that can efficiently support both fine- and coarse-grained L1 cache-line management for applications with both regular and irregular memory access patterns. Specifically, it can store 32- or 64-byte words in non-contiguous memory space to a single 128-byte cache line. Furthermore, it neither requires an extra tag storage structure nor reduces the capacity of L1 cache since it stores auxiliary tags for fine-grained L1 cache-line managements in sharedmemory space that is not fully used in many applications. Our experiment shows that Elastic-Cache improves the geo-mean performance of applications with irregular memory access patterns by 58% without degrading performance of applications with regular memory access patterns. Bingchao Li, Murali Annavaram, Nam Sung Kim |
IPDPS | 1 |