Kangping Wang

dblp:29/6707 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
10since 2021 · last 2025
0000-0003-4402-3346ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 WST: Wavelet-Based Multi-scale Tuning for Visual Transfer Learning
abstract
Large-scale pre-trained Vision Transformer (ViT) models have demonstrated remarkable performance on visual tasks but are computationally expensive to transfer to downstream tasks. Parameter-Efficient Fine-Tuning (PEFT) offers a promising transferring approach by updating only a subset of parameters. However, PEFT's effectiveness is hindered by discrepancies between pre-training and downstream tasks in terms of object scale and granularity. Downstream tasks often focus on finer-grained and more specialized recognition, requiring more detailed features. The diversity of feature scales of existing PEFT methods for ViT is limited. To address this, we propose a novel PEFT method named Wavelet-based multi-Scale Tuning (WST), which learns multi-scale features in a simple and efficient way. WST introduces a parallel fine-tuning patch embedding branch with a smaller patch size than the pre-trained model to capture finer-grained features. Furthermore, to handle the computational challenge from the resulting longer token sequence, WST designs wavelet fine-tuning blocks that balance both efficiency and performance. In the block, wavelet transform enables invertible and lossless down-sampling of the longer token sequence, aligning it with that of the backbone, and two lightweight linear mappings are employed to learn task-specific features. This design facilitates efficient multi-scale information exchange between the pre-trained backbone and fine-tuning branch. Extensive experiments on transfer learning demonstrate the promising performance and efficiency of our WST.
Kangping Wang
AAAI3
2025 Efficient feature selection for pre-trained vision transformers
Lan Huang 0002, Mengqiang Yu, Weiping Ding 0001, Xingyu Bai, Kangping Wang
Comput. Vis. Image Underst.6
2025 MCSSA: A Stream-Based Multiconcurrency Systolic Sorting Array Combining Merge Tree
abstract
The exploration of utilizing reconfigurable circuits with parallel computing capabilities has been conducted to enhance sorting performance and reduce power consumption. However, most sorting algorithms using dedicated processors are based on parallelization designs of serial algorithms without considering the design method of large-scale integrated circuits. This results in various issues, including the overuse of$I/O$interface resources, on-chip storage resources, and complex layout wiring. In this article, we extend the 2-tuple relation in the uniform recurrence equation (URE) structure used to define the systolic array to n-tuples, and the extended structure is flexible in defining$I/O$bandwidth and concurrency. Then we define the multiconcurrency systolic sorter array (MCSSA) algorithm based on the extended URE structure, which has a flexible$4N/n$time complexity based on the n-tuple relation. Moreover, this systolic array can simultaneously sort two independent sequences, increasing the reuse of resources. Afterwards, we encapsulate each n-tuple into a processing element (PE) cell. The entire MCSSA consists of these interconnected PE cells, each of which can be customized in terms of data bit width and type. Last but not least, we have improved the merge tree structure called MC-merge tree. The concurrency of this algorithm can also be flexibly defined, we use this algorithm combined with MCSSA to cope with large-scale sorting scenarios. In our experiments, we have demonstrated the speed-up ratio of MCSSA relative to other state of the art (SOTA) sorting algorithms. Inheriting the unity and simplicity from the Systolic Array architecture, MCSSA achieves a maximum$73.17\times $acceleration ratio on the U200. In addition, the MC-merge tree expands the MCSSA sorting scale with a maximum of 450.56 times while maintaining the advantage of the acceleration ratio. The results of our study demonstrate that MCSSA and MC-merge tree have better acceleration, throughput and scalability advantages over other SOTA algorithms.
Lan Huang 0002, Teng Gao, Kangping Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 FAM: Improving columnar vision transformer with feature attention mechanism
Lan Huang 0002, Xingyu Bai, Mengqiang Yu, Wei Pang 0001, Kangping Wang
Comput. Vis. Image Underst.6
2024 Prompt-based learning framework for zero-shot cross-lingual text classification
Lan Huang 0002, Kangping Wang, Wei Wei 0042, Rui Zhang 0084
Eng. Appl. Artif. Intell.3
2024 Neural Architecture Search for Text Classification With Limited Computing Resources Using Efficient Cartesian Genetic Programming
abstract
Cartesian Genetic Programming (CGP) has often been applied for Neural Architecture Search (NAS). However, the performance of CGP is less than ideal when searching for architectures with limited computing resources. To better facilitate NAS with limited computing resources, this paper proposes a crossover operator, a light-weighted age mechanism, and two adaptive mutation operators as the novel components in our Efficient Cartesian Genetic Programming (ECGP) method. To assess the performance of ECGP, we conduct extensive experiments on three text classification task datasets. The experimental results demonstrate that ECGP outperforms other NAS methods, requiring only hundreds of fitness evaluations to find architectures with competitive accuracy compared with human-designed models. Additionally, the ECGP-evolved architectures are shown as converging fast and stably, and having high-level transferability with merely a 1-2% accuracy drop. Ablation studies demonstrate the effectiveness of the proposed operators and age mechanism, and identify GRU as the most critical function in the text classification task. Finally, we summarize three design principles observed from the ECGP-evolved architectures that are in line with human-design strategies. To the best of our knowledge, this work introduces the first attention-derived NAS benchmark for the text classification task.
Xuan Wu 0004, Di Wang 0004, Huanhuan Chen 0001, Lele Yan, Yubin Xiao, Chunyan Miao, Hong-Wei Ge, Dong Xu 0002, Yanchun Liang 0001, Kangping Wang, Chunguo Wu, You Zhou 0008
IEEE Trans. Evol. Comput.10
2024 SSA: A Uniformly Recursive Bidirection-Sequence Systolic Sorter Array
abstract
The use of reconfigurable circuits with parallel computing capabilities has been explored to enhance sorting performance and reduce power consumption. Nonetheless, most sorting algorithms utilizing dedicated processors are designed solely based on the parallelization of the algorithm, lacking considerations of specialized hardware structures. This leads to problems, including but not limited to the consumption of excessive I/O interface resources, on-chip storage resources, and complex layout wiring. In this paper, we propose a Systolic Sorter Array, implemented by a Uniform Recurrence Equation (URE) with highly parameterised in terms of data size, bit width and type. Leveraging this uniformly recursive structure, the sorter can simultaneously sort two independent sequences. In addition, we implemented global and local control modes on the FPGA to achieve higher computational frequencies. In our experiments, we have demonstrated the speed-up ratio of SSA relative to other state of the art (SOTA) sorting algorithms using C++$std$::$sort()$as benchmark. Inheriting the benefits from the Systolic Array architecture, the SSA reaches up to 810 Mhz computing frequency on the U200. The results of our study show that SSA outperforms other sorting algorithms in terms of throughput, speed-up ratio, and computation frequency.
Teng Gao, Lan Huang 0002, Kangping Wang
IEEE Trans. Parallel Distributed Syst.4
2023 U-DARTS: Uniform-space differentiable architecture search
Lan Huang 0002, Wencong Wang, Wei Pang 0001, Kangping Wang
Inf. Sci.6
2022 Restorable-inpainting: A novel deep learning approach for shoeprint restoration
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Kangping Wang, Daixi Li, You Zhou 0008, Dong Xu 0002
Inf. Sci.5
2022 Research on Reverse Skyline Query Algorithm Based on Decision Set
abstract
Reverse skyline query is an extension of the classical skyline query, widely used in the decision support in e-business. The vast burst of big data in e-business challenges the classical algorithms for such queries. This paper provides a novel definition of decision set and a decision set based reverse skyline query method called DRS on the double-layer R tree indexing in a map-reduce manner. Theoretical proofs are provided for the correctness and complexity of the DRS algorithm. Experiments made using several large data sets are presented and analyzed to illustrate the applicability and the outperformance of DRS over the state-of-the-art reverse skyline query methods.
Lan Huang 0002, Yuanwei Zhao, Pedro Mestre, Laipeng Han, Kangping Wang, Wenjuan Gao, Rui Zhang 0084
J. Database Manag.5
2020 An Incremental Learning Network Model Based on Random Sample Distribution Fitting
Wencong Wang, Lan Huang 0002, Kainuo Li, Kangping Wang
KSEM (2)7
2020 A Survey on Performance Optimization of High-Level Synthesis Tools
Lan Huang 0002, Dalin Li, Kangping Wang, Teng Gao, Adriano Tavares
J. Comput. Sci. Technol.3
2020 An Extended Nonstrict Partially Ordered Set-Based Configurable Linear Sorter on FPGAs
abstract
Sorting is essential for many scientific and data processing problems. It is significant to improve the efficiency of sorting. Taking advantage of specialized hardware, parallel sorting, e.g., sorting networks and linear sorters, implements sorting in lower time complexity. However, most of them are designed based on the parallelization of algorithms, lacking consideration of specialized hardware structures. In this article, we propose an extended nonstrict partially ordered set-based configurable linear sorter on field-programmable gate arrays (FPGAs). First, we extend nonstrict partial order to the binary tuple and n-tuple nonstrict partial orders. Then, the linear sorting algorithm is defined based on them, with the consideration of hardware performance. It has 4N/n time complexity varying from 4 to 2 N as the tuple size varies. The number of comparisons reduces to N/2 in binary tuple-based sorting, which is half of the state-of-the-art insertion linear sorting. Finally, we implement the linear sorter on FPGAs. It consists of multiple customizable micro-cores, named sorting units (SUs). The SU packages the storage and comparison of the tuple. All the SUs are connected into a chain with simple communication, which makes the sorter fully configurable in length, bandwidth, and throughput. They also act the same in each clock cycle, so that the achieved frequency of the sorter improves. In our experiment, the sorter achieves at most 660-MHz frequency, 5.6 Gb/s throughput, and 87 times speed-up compared with the quick sort algorithm on general processors.
Dalin Li, Lan Huang 0002, Teng Gao, Adriano Tavares, Kangping Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6