Dekui Wang

dblp:223/6273 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-4783-446XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Electronic design automation · 79% Parallel and multicore computing · 13% Interconnection networks and networks-on-chip · 8%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation › physical design › routing
FPGA routing
2.142023
FCRoute: A Fast FPGA Connection Router Using Soft Routing-Space Pruning Algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
A Fast FPGA Connection Router Using Prerouting-Based Parallel Local Routing Algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
ParRA: A Shared Memory Parallel FPGA Router Using Hybrid Partitioning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Electronic design automation
physical design
2.142023
FCRoute: A Fast FPGA Connection Router Using Soft Routing-Space Pruning Algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
A Fast FPGA Connection Router Using Prerouting-Based Parallel Local Routing Algorithm · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
ParRA: A Shared Memory Parallel FPGA Router Using Hybrid Partitioning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Interconnection networks and networks-on-chip › routing algorithms
parallel routing
0.412020
ParRA: A Shared Memory Parallel FPGA Router Using Hybrid Partitioning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Parallel and multicore computing › parallel algorithms
shared-memory parallel algorithms
0.412020
ParRA: A Shared Memory Parallel FPGA Router Using Hybrid Partitioning Approach · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Parallel and multicore computing
runtime optimization
0.312018
A Runtime Optimization Approach for FPGA Routing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Electronic design automation › physical design › routing
timing-driven routing
0.312018
A Runtime Optimization Approach for FPGA Routing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018

Methods — techniques the papers use, named apart from their topics

soft routing-space pruning · 0.7prerouting · 0.7parallel local search · 0.7backtracking · 0.7a-star maze search · 0.7a-star maze expansion · 0.7parallel routing strategies · 0.4hybrid partitioning · 0.4conflict-free subset partitioning · 0.4maze expansion · 0.3
YearPublicationVenuePosition
2026 Toward Safe Driving: Efficient Detection of Small Blurred Signs in Real-World Scenarios
abstract
Accurate traffic sign recognition is critical for safe driving, as over half of traffic accidents stem from drivers’ negligence of traffic signs. Thus, developing robust traffic sign detection methods is essential to improve road safety. While existing object detection methods have achieved remarkable success, their performance in traffic sign detection is often limited by small object sizes and low-resolution appearances. To address these issues, this study proposes a novel traffic sign detection framework with three innovative components for precise localization and classification: 1) a hierarchical feature aggregation module that emphasizes high-level semantic information for traffic sign localization; 2) a cross-layer semantic residual network that enhances recognition of small and blurred signs via hierarchical feature interaction and fusion; 3) a lightweight feature alignment unit that bridges semantic gaps between cross-level representations. These components jointly tackle the challenges of detecting small and blurred traffic signs in real-world driving scenarios. Experiments were conducted on three public datasets (TT100K, CCTSDB2021, GTSDB). Results show the proposed model outperforms other methods with comparable parameters. Additionally, dynamic motion blur augmentation was applied to datasets to simulate real driving scenarios, and experiments confirm the proposed method achieves state-of-the-art performance under such challenging conditions. Code is publicly available athttps://github.com/Mo7nex/SAttFusion-YOLO
Dekui Wang, Jun Feng 0003, Qirong Bo, Yaqiong Xing, Wei Zhou 0012, Xingxing Hao
IEEE Trans. Intell. Transp. Syst.2
2025 Exploring multi-scale and cross-type features in 3D point cloud learning with CCMNET
Wei Zhou 0012, Weiwei Jin, Dekui Wang, Xingxing Hao, Yongxiang Yu, Caiwen Ma
Expert Syst. Appl.3
2023 Speech Emotion Recognition via Heterogeneous Feature Learning
abstract
Speech emotion recognition (SER) based on multi-view learning has made some progress on speaker-independent scenarios. How-ever, the existing SER methods always rely on excessive feature views and ignore the importance of heterogeneous feature learning. In this paper, we propose a novel multi-level attention method to effectively learn the heterogeneous information from the hand-crafted feature (MFCC) and the feature (W2V2) extracted from the pre-trained model. Specifically, we first design an Attention based Multi-scale Low-level Feature (A-MLF) extractor to extract scale-specific emotion-related regions from MFCC. Then, the Multi-Unit Attention (MUA) module is used to simultaneously learn discriminative features in three different dimensions. Finally, a two-stage feature fusion strategy is used for joint representation space learning. We demonstrate our method on two speaker-independent validation strategies and interpret the SOTA performance by visualizing the feature distribution.
Dongya Wu, Dekui Wang, Jun Feng 0003
ICASSP3
2023 Speech Emotion Recognition Via Two-Stream Pooling Attention With Discriminative Channel Weighting
abstract
Multi-view Speech Emotion Recognition (SER) based on the pre-trained model has achieved success in speaker-independent scenarios. However, the existing SER methods rely on excessive feature views and have complicated feature fusion strategies. In this paper, we propose a novel method to learn effective emotion-related information from two feature views. First, we present a Discriminative Channel Weighting (DCW) module to weight the channel dimension of the features produced by a set of multi-scale convolution layers. This module allows for discriminative weighting of complex channel dimensions. Second, a concise Two-stream Pooling Attention (TsPA) strategy is proposed to generate two groups of fusion features based on different channel-level embeddings with different emphasis. Finally, the SER task is completed by three consecutive fully connected layers. The effectiveness of the proposed method has been demonstrated on two speaker-independent validation strategies, outperforming other state-of-the-art approaches.
Dekui Wang, Dongya Wu, Jun Feng 0003
ICASSP2
2023 Multi Point-Voxel Convolution (MPVConv) for deep learning on point clouds
Wei Zhou 0012, Xingxing Hao, Dekui Wang, Ying He 0001
Comput. Graph.4
2023 A Fast FPGA Connection Router Using Prerouting-Based Parallel Local Routing Algorithm
abstract
Routing is one of the most time-consuming steps in the field-programmable gate array (FPGA) design process. Even if unceasing efforts have been made to accelerate FPGA routing, the existing work seldom pays attention to the underlying FPGA connection router. In this article, we present a fast FPGA connection router called PRoute which implements a novel prerouting-based parallel local routing algorithm. Basically, PRoute precomputes the potential routing solutions for various connection patterns on FPGAs, which can be directly used in the later practical routing. On the whole, PRoute is composed of A-star maze expansion and parallel local search. In the first part, PRoute gradually expands the maze wavefront toward the lowest-cost node to search for the target sink. For a wire-type node, PRoute invokes a fast parallel local search instead taking advantage of the prerouting results, and hence the time-expensive maze expansion can be reduced. Particularly, it allows PRoute to call one another between A-star maze expansion and parallel local search. This enables the runtime efficiency of PRoute while ensuring its global search ability. In addition, we put forward an engineering improvement to further speed up PRoute by avoiding the exploration of block output pins. To our best knowledge, this work is the first to apply the idea of prerouting for FPGAs. Experimental results show that PRoute achieves speedups of$1.8\times $,$2.4\times $,$3.2\times $,$4.1\times $, and$5.1\times $with 1, 4, 8, 16, and 32 threads relative to the baseline versatile place to route’s connection router, respectively, without degrading the quality of results.
Dekui Wang, Jun Feng 0003, Wei Zhou 0012, Xingxing Hao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 FCRoute: A Fast FPGA Connection Router Using Soft Routing-Space Pruning Algorithm
abstract
Routing is one of the most time-consuming stages in the field-programmable gate array (FPGA) design flow. Even if various attempts have been made to reduce route time, the existing work rarely focuses on improving the underlying A*-based FPGA connection router. In this article, we present a fast FPGA connection router called FCRoute based on a novel soft routing-space pruning algorithm. Within FCRoute, a routing resource priority mechanism is applied to classify the routing resource nodes into high-priority nodes and low-priority ones. On the whole, FCRoute is composed of fast maze search and backtracking process. During the fast maze search, we explore only the high-priority nodes in the routing space. In this way, a great deal of unnecessary work can be avoided. When the fast maze search fails to find the target sink, it allows the backtracking process to explore the low-priority nodes promising to be on the best path, after which a new fast maze search is called. By avoiding the exploration of the majority of low-priority nodes, FCRoute maintains runtime efficiency while ensuring global search ability. In addition, we further accelerate FCRoute with an engineering enhancement which simplifies the cost computations of nodes. Runtime and quality of results are compared with the state-of-the-art connection router in VPR 8. Experimental results show that on average FCRoute explores less than half the number of routing resource nodes, and therefore reduces runtime by 38% while enabling the quality of results. When combined with the enhancement, FCRoute achieves an average 45% reduction on runtime without sacrificing the quality of results.
Dekui Wang, Jun Feng 0003, Wei Zhou 0012, Xingxing Hao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Speech Emotion Recognition via Multi-Level Attention Network
abstract
Aiming to improve the performance of human speech emotion recognition (SER), the existing work has made great progress based on the popular mel-scale frequency cepstral coefficient (MFCC). However, the existing work rarely pays attention to the low-level emotion related features in MFCC, such as the underlying interactive relations. In this letter, we propose a novel multi-level attention network (MLAnet), which contains a multi-scale low-level feature (MLF) extractor and a multi-unit attention (MUA) module. Within the MLF extractor, we minimize the task-irrelevant information which harms the performance of SER by applying the attention mechanism. Since the features extracted by the MLF extractor contain rich domain-specific emotion information, we further present a MUA module to simultaneously weight the features in terms of time, frequency and channel dimensions. In this way, the discriminative emotion features in different dimensions can be extracted by corresponding weighting blocks. Experimental results on two benchmark datasets demonstrate that the proposed method outperforms other state-of-the-art approaches.
Dekui Wang, Dongya Wu, Jun Feng 0003
IEEE Signal Process. Lett.2
2020 ParRA: A Shared Memory Parallel FPGA Router Using Hybrid Partitioning Approach
abstract
In this paper, we propose a shared-memory parallel field-programmable gate array (FPGA) router called ParRA. Basically, ParRA is composed of hybrid partitioning and parallel routing. During the hybrid partitioning, first an FPGA is split into multiple subregions and nets are geographically partitioned into local subsets. As the intersubregion nets usually overlap each other, these nets cannot be routed in parallel. Second, the intersubregion nets are further partitioned into conflict-free subsets. Since each conflict-free subset consists of intersubregion nets do not overlap each other, the nets in the same conflict-free subset can be routed in parallel. In this way, we significantly increase the number of nets that have potential to be routed in parallel. During the parallel routing process, two novel parallel routing strategies are applied to route the nets in conflict-free and local subsets, respectively. With conflict-free subsets, sinks in the same conflict-free subset are routed in parallel while conflict-free subsets are routed one by one. On the contrast, local subsets are routed in parallel while the nets in the same local subset are routed sequentially. With the two different parallel routing strategies, we reduce the interference between threads and balance the workload of threads, which contributes to gain more parallelism. The proposed parallel router provides deterministic routing results. The experimental results show that ParRA achieves an average speedup of $24.3 {\times }$ with 16 threads compared to VPR 7.0, has no negative impact on the quality of results.
Dekui Wang, Cong Tian 0001, Bohu Huang, Nan Zhang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 A Runtime Optimization Approach for FPGA Routing
abstract
In this paper, we present a new field-programmable gate array (FPGA) routing approach on the basis of the PathFinder routing algorithm. During each routing iteration, our approach applies a novel timing-based rerouting strategy to only reroute the illegal paths. At a lower level, each maze expansion is started from the relatively close part of current routing tree to search for the target sink on the routing resource graph. Experimental results demonstrate that on average the proposed approach reduces the routing runtime by 68.5% compared with the timing-driven router in versatile place and route FPGA placement and routing framework, with reduction of 2.5% and 1.4% in critical path delay and wirelength, respectively.
Dekui Wang, Cong Tian 0001, Bohu Huang, Nan Zhang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1