Songwen Pei

dblp:61/763 · DBLP profile ↗
← Back
49ranked-venue papers
19as first author
36since 2021 · last 2026
0000-0003-0810-1458ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 7 first-author · 16 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 12 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Revisiting Global Value Prediction: A Resurgent Complement to Local Predictors
Ling Yang 0008, Libo Huang 0002, Bingcai Sui, Sheng Ma, Yongwen Wang, Li Shen 0007, Qianming Yang, Songwen Pei
ISCA11
2026 An efficient three-stage network via Multi-Scale Orthogonal Complementary Transformer for low-light image enhancement
Junhao Tan, Songwen Pei, Kai Cong
Comput. Vis. Image Underst.2
2026 FedACE: A Federated Adaptive Components Exfoliation method for medical image segmentation in non-IID scenarios
Sheng Liang, Songwen Pei, Chunhua Gu, Zelei Liu, Lixin Fan
Expert Syst. Appl.2
2026 A Dynamic Application Configuration and Capacity-Aware Offloading System for MEC Optimization
abstract
Mobile Edge Computing (MEC) is a key technology for enabling energy-efficient and low-latency processing of computation-intensive applications through task offloading. However, current frameworks typically model dependent tasks using static Directed Acyclic Graphs (DAGs), which are poorly suited to dynamic edge environments characterized by fluctuating resources and diverse QoS demands. These fixed DAG structures often fail to adapt to runtime changes in bandwidth or node workload, leading to frequent task failures or inefficient resource usage. To overcome these limitations, we propose Dynamic Application Configuration, a paradigm that allows applications to switch between multiple predefined DAG variants at runtime. This enables the system to dynamically balance accuracy and resource efficiency at critical decision points by adapting to current network and computing conditions. Based on this concept, we design the Dynamic Application Configuration and Capacityaware task offloading System (DACCS), which employs a two-step offloading strategy: (i) a Dynamic Graph Selection (DGS) algorithm that adaptively adjusts application configurations during key subtask execution based on real-time resource states, and (ii) a dependency-aware offloading algorithm (DAP) that optimizes offloading assignments by jointly optimizing the application completion rate, processing delay and effective utilization rate. Experiments and real-system validations demonstrate that DGS improves the service performance across multiple strategies. Furthermore, the proposed DGS-DAP strategy outperforms other benchmark approaches under dynamic conditions.
Bobo Ju, Songwen Pei, Azzedine Boukerche, Peng Sun 0007
IEEE Internet Things J.4
2026 Multimodal human video generation with uncertainty-aware pose guidance
Kun Yang 0010, Yuanyuan Meng, Juncen Guo, Yanda Meng, Songwen Pei, Jing Liu 0050, Yang Liu 0246
Pattern Recognit.7
2025 MEVLERG: Medical Vision-Language Encoder for Report Generation
Songwen Pei, Kai Cong, Xing Jia
IEEE Big Data1
2025 TGRPO-SRM: Token-Aware GRPO with Semantic Reward Modeling for Referring Expression Grounding
Songwen Pei, Kai Cong
IEEE Big Data1
2025 Mitigating Privacy Issues in RAG through Causal Disentanglement
Songwen Pei, Kai Cong, Joel Rodrigues
IEEE Big Data2
2025 Facial Features Enhanced Multi-branch Graph Network for Driver Drowsiness Detection
Songwen Pei, Huichen Zhang
DASFAA (1)1
2025 SONet: Towards Practical Online Neural Network for Enhancing Hard-to-Predict Branches
Zhenxuan Xiong, Libo Huang 0002, Ling Yang 0008, Hui Guo 0004, Songwen Pei, Gang Chen 0023, Yongwen Wang
Euro-Par (2)7
2025 DropNaE: Alleviating irregularity for large-scale graph representation learning
Xin Liu 0073, Xunbin Xiong, Mingyu Yan, Runzhen Xue, Shirui Pan, Songwen Pei, Lei Deng 0003, Xiaochun Ye, Dongrui Fan
Neural Networks6
2024 Why Misinformation is Created? Detecting them by Integrating Intent Features
abstract
Various social media platforms, e.g., Twitter and Reddit, allow people to disseminate a plethora of information more efficiently and conveniently. However, they are inevitably full of misinformation, causing damage to diverse aspects of our daily lives. To reduce the negative impact, timely identification of misinformation, namely Misinformation Detection (MD), has become an active research topic receiving widespread attention. As a complex phenomenon, the veracity of an article is influenced by various aspects. In this paper, we are inspired by the opposition of intents between misinformation and real information. Accordingly, we propose to reason the intent of articles and form the corresponding intent features to promote the veracity discrimination of article features. To achieve this, we build a hierarchy of a set of intents for both misinformation and real information by referring to the existing psychological theories, and we apply it to reason the intent of articles by progressively generating binary answers with an encoder-decoder structure. We form the corresponding intent features and integrate it with the token features to achieve more discriminative article features for MD. Upon these ideas, we suggest a novel MD method, namely Detecting Misinformation by Integrating Intent featuRes (DM-INTER). To evaluate the performance of DM-INTER, we conduct extensive experiments on benchmark MD datasets. The experimental results validate that DM-INTER can outperform the existing baseline MD methods.
Bing Wang 0018, Ximing Li 0002, Changchun Li, Bo Fu 0001, Songwen Pei, Sheng-Sheng Wang 0001
CIKM5
2024 INSPIRE: Accelerating Deep Neural Networks via Hardware-friendly Index-Pair Encoding
abstract
Deep Neural Network (DNN) inference consumes significant computing resources and development efforts due to the growing model size. Quantization is a promising technique to reduce the computation and memory cost of DNNs. Most existing quantization methods rely on fixed-point integers or floating-point types, which require more bits to maintain model accuracy. In contrast, variable-length quantization, which combines high precision for values with significant magnitudes (i.e., outliers) and low precision for normal values, offers algorithmic advantages but introduces significant hardware overhead due to variable-length encoding and decoding. Also, existing quantization methods are less effective for both (dynamic) activations and (static) weights due to the presence of outliers.
Fangxin Liu, Ning Yang 0012, Zhiyan Song, Zongwu Wang, Haomin Li 0002, Shiyuan Huang 0004, Zhuoran Song, Songwen Pei, Li Jiang 0002
DAC8
2024 EOS: An Energy-Oriented Attack Framework for Spiking Neural Networks
abstract
Spiking neural networks (SNNs) are emerging as energy-efficient alternatives to traditional artificial neural networks (ANNs). Their event-driven information processing significantly reduces computational demands while maintaining competitive performance. However, as SNNs are increasingly deployed in edge devices, security concerns have emerged. While significant research efforts have been dedicated to addressing the security vulnerabilities stemming from malicious input, often referred to as adversarial examples, the security of SNN parameters remains relatively unexplored. This work introduces a novel attack methodology for SNNs known as Energy-Oriented SNN attack (EOS). EOS is designed to increase the energy consumption of SNNs through the malicious manipulation of binary bits within their memory systems (i.e., DRAM), where neuronal information is stored. The key insight of EOS lies in the observation that energy consumption in SNN implementations is intricately linked to spiking activity. The bit-flip operation, the well-known Row Hammer technique, is employed in EOS. It achieves this by identifying the most robust neurons in the SNN based on the spiking activity, particularly those related to the firing threshold, which is stored as binary bits in memory. EOS employs a combination of spiking activity analysis and a progressive search strategy to pinpoint the target neurons for bit-flip attacks. The primary objective is to incrementally increase the energy consumption of the SNN while ensuring that accuracy remains intact. With the implementation of EOS, successful attacks on SNNs can lead to an average of 43% energy increase with no drop in accuracy.
Ning Yang 0012, Fangxin Liu, Zongwu Wang, Haomin Li 0002, Zhuoran Song, Songwen Pei, Li Jiang 0002
DAC6
2024 SPARK: Scalable and Precision-Aware Acceleration of Neural Networks via Efficient Encoding
abstract
Deep Neural Networks (DNNs) have demonstrated remarkable success; however, their increasing model size poses a challenge due to the widening gap between model size and hardware capacity. To address this, model compression techniques have been proposed, but existing compression methods struggle to effectively handle the significant parameter variations (activations and weights) within the model. Moreover, current variance-aware encoding solutions for compression introduce complex logic, leading to limited compression benefits and hardware efficiency. In this context, we present SPARK, a novel algorithm/architecture co-designed solution that utilizes variable-length data representation for local parameter value processing, offering low hardware overhead and high-performance gains. Our key insight is that the high-order part in quantized values are often sparse, allowing us to employ an identity bit to assign the appropriate encoding length, thereby eliminating redundant bit-length footprints. This reduction in data representation based on data characteristics enables a serialized structured data encoding scheme that seamlessly integrates with existing hardware accelerators, such as systolic arrays. We evaluate SPARK-based accelerators against some existing encoding-based accelerator, and our results demonstrate significant improvements. The SPARK-based accelerator achieves up to 4.65 × speedup and 74.7% energy reduction, while maintaining superior model accuracy.
Fangxin Liu, Ning Yang 0012, Haomin Li 0002, Zongwu Wang, Zhuoran Song, Songwen Pei, Li Jiang 0002
HPCA6
2024 DiffSenseNet: Integrating Hierarchical Features and Angular Diffusion for Remote Sensing Object Detection
abstract
Object detection in remote sensing images is a significant and challenging task. The detection performance is difficult to further improve due to the wide variation in object scale and unpredictable orientations. However, traditional methods typically rely on Region Proposal Networks (RPN) with fixed anchor boxes, which makes it difficult to locate objects of multi-scale and multi-orientation. In this paper, we propose DiffSenseNet, a detection network that integrates Hierarchical Feature Fusion (HFF) and Angular Diffusion Augmentation (ADA). The HFF architecture integrates both bottom-up and top-down pathways to effectively merge features at different scales. The ADA strategy introduces directional information into noisy boxes by adding Gaussian noise to the ground truth boxes. On the DOTA and HRSC2016 datasets, DiffSenseNet achieves the accuracy of 75.89% and 86.24% mAP, respectively. Comprehensive experiments show that DiffSenseNet achieves relatively superior performance compared to previous state-of-the-art jobs.
Songwen Pei, Hongli Ma, Xinyun Qiu, Libo Huang 0002
ISPA1
2024 Global Color-Aware Arbitrary Style Transfer with Discrete Wavelet Transform
Bochao Chen, Junhao Tan, Songwen Pei
NPC (1)4
2024 A simple rapid sample-based clustering for large-scale data
Yewang Chen, Songwen Pei, Yi Chen 0007, Jixiang Du
Eng. Appl. Artif. Intell.3
2024 DPQ: dynamic pseudo-mean mixed-precision quantization for pruned neural network
Songwen Pei, Bingxue Zhang, Hai Xue, Xiaochun Ye, Mingsong Chen 0001
Mach. Learn.1
2024 Crowded pedestrian detection with optimal bounding box relocation
Ren Han, Meiqi Xu, Songwen Pei
Multim. Tools Appl.3
2024 Exploiting Temporal-Unrolled Parallelism for Energy-Efficient SNN Acceleration
abstract
Event-driven spiking neural networks (SNNs) have demonstrated significant potential for achieving high energy and area efficiency. However, existing SNN accelerators suffer from issues such as high latency and energy consumption due to serial accumulation-comparison operations. This is mainly because SNN neurons integrate spikes, accumulate membrane potential, and generate output spikes when the potential exceeds a threshold. To address this, one approach is to leverage the sparsity of SNN spikes to reduce the number of time steps. However, this method can result in imbalanced workloads among neurons and limit the utilization of processing elements (PEs). In this paper, we present SATO, a temporal-parallel SNN accelerator that enables parallel accumulation of membrane potential for all time steps. SATO adopts a two-stage pipeline methodology, effectively decoupling neuron computations. This not only maintains accuracy but also unveils opportunities for fine-grained parallelism. By dividing the neuron computation into distinct stages, SATO enables the concurrent execution of spike accumulation for each time step, leveraging the parallel processing capabilities of modern hardware architectures. This not only enhances the overall efficiency of the accelerator but also reduces latency by exploiting parallelism at a granular level. The architecture of SATO includes a novel binary adder-search tree for generating the output spike train, effectively decoupling the chronological dependence in the accumulation-comparison operation. Furthermore, SATO employs a bucket-sort-based method to evenly distribute compressed workloads to all PEs, maximizing data locality of input spike trains. Experimental results on various SNN models demonstrate that SATO outperforms the well-known accelerator, the 8-bit version of “Eyeriss” by$20.7\times$in terms of speedup and$6.0\times$energy-saving, on average. Compared to the state-of-the-art SNN accelerator “SpinalFlow”, SATO can also achieve$4.6\times$performance gain and$3.1\times$energy reduction on average, which is quite impressive for inference.
Fangxin Liu, Zongwu Wang, Wenbo Zhao 0005, Ning Yang 0012, Yongbiao Chen, Shiyuan Huang 0004, Haomin Li 0002, Tao Yang 0031, Songwen Pei, Xiaoyao Liang, Li Jiang 0002
IEEE Trans. Parallel Distributed Syst.9
2023 Multi-view Neighbor-Enriched Contrastive Learning Framework for Bundle Recommendation
Sheng Liang, Songwen Pei
ICA3PP (3)3
2023 Realistic Sketch Face Generation via Sketch-Guided Incomplete Restoration
abstract
Generating realistic sketches of human faces is an important research direction in the fields of computer vision and computer graphics. However, generating high-quality sketch faces remains challenging, especially when dealing with incomplete input sketch images. In this paper, we propose a framework based on the U-Net generative network, which progressively optimizes network performance through a three-stage training strategy. We introduce the Sketch Guidance and Completion Network (SGCN), which is capable of restoring missing features, completing blurry edges, and handling occluded regions in sketch images. By employing multiple discriminators and adversarial training strategies, we enhance the realism and preserve fine details in the generated results. Through experiments on publicly available sketch face datasets, we demonstrate significant improvements in reconstructing incomplete sketch faces using our approach. The generated images exhibit enhanced completeness, realism, and retain fine details and important facial features. Quantitative evaluations validate the effectiveness and superiority of our method.
Chunhua Gu, Huichen Zhang, Songwen Pei
ICPADS5
2023 Knowledge Enhancement and Feature Purification for Single-stage Joint Entity and Relation Extraction
abstract
Joint entity and relation extraction aim to achieve named entity recognition and relation extraction in unstructured text. We use the form of triples (subject, relation, object) to describe entity and relation. Joint entity and relation extraction play an important role in knowledge graph construction, question and answering, data analysis, and other natural language processing domains. Most of the existing works suffered from insufficient interaction between entity features and relation features due to the extraction order. Moreover, there is a prevalence of heterogeneous representations of entities and relations. They not only increase the amount of model design and the complexity of training, but also lead to the problem of exposure bias due to the extraction order. Therefore, in this paper, we propose KFRel, where we enhance the entity and relation representations by encoding text and relation. Then, the features are fused to enhance the entity and relation representations, and adopt a feature purification module, where the features are enhanced by a feature purification module that removes feature information irrelevant to the joint entity and relation extraction while retaining feature information of relevance. In addition, we adopt a single-stage joint entity and relation extraction module to address the issue of overlapping triples. This module aims to help improve the effectiveness of joint entity and relation extraction. Comprehensive experiments are conducted on two widely-used datasets, and the experimental results demonstrate that the proposed method is effective and outperforms the state-of-the-art baselines.
Guohao Zhai, Songwen Pei
ICPADS2
2023 SLSNet: Weakly-Supervised Skin Lesion Segmentation Network with Self-attentions
Songwen Pei
PRICAI (3)1
2023 Carbon Emissions Reduction of Neural Network by Discrete Rank Pruning
Songwen Pei, Sheng Liang, Haonan Ding, Xiaochun Ye, Mingsong Chen 0001
CCF Trans. High Perform. Comput.1
2023 Attention-based efficient robot grasp detection network
abstract
To balance the inference speed and detection accuracy of a grasp detection algorithm, which are both important for robot grasping tasks, we propose an encoder–decoder structured pixel-level grasp detection neural network named the attention-based efficient robot grasp detection network (AE-GDN). Three spatial attention modules are introduced in the encoder stages to enhance the detailed information, and three channel attention modules are introduced in the decoder stages to extract more semantic information. Several lightweight and efficient DenseBlocks are used to connect the encoder and decoder paths to improve the feature modeling capability of AE-GDN. A high intersection over union (IoU) value between the predicted grasp rectangle and the ground truth does not necessarily mean a high-quality grasp configuration, but might cause a collision. This is because traditional IoU loss calculation methods treat the center part of the predicted rectangle as having the same importance as the area around the grippers. We design a new IoU loss calculation method based on an hourglass box matching mechanism, which will create good correspondence between high IoUs and high-quality grasp configurations. AEGDN achieves the accuracy of 98.9% and 96.6% on the Cornell and Jacquard datasets, respectively. The inference speed reaches 43.5 frames per second with only about 1.2 × 106 parameters. The proposed AE-GDN has also been deployed on a practical robotic arm grasping system and performs grasping well. Codes are available at https://github.com/robvincen/robot_gradet .
Xiaofei Qin, Wenkai Hu, Chen Xiao, Changxiang He, Songwen Pei, Xuedian Zhang
Frontiers Inf. Technol. Electron. Eng.5
2023 RISAT: real-time instance segmentation with adversarial training
Songwen Pei, Bo Ni, Tianma Shen, Zhenling Zhou, Yewang Chen, Meikang Qiu
Multim. Tools Appl.1
2022 Wavelet-Based Mamba with Fourier Adjustment for Low-Light Image Enhancement
Junhao Tan, Songwen Pei, Bo Fu 0001, Ximing Li 0002
ACCV (4)2
2022 DRP: Discrete Rank Pruning for Neural Network
Songwen Pei, Sheng Liang
NPC1
2022 TransMigrator: A Transformer-Based Predictive Page Migration Mechanism for Heterogeneous Memory
Songwen Pei, Yihuan Qian, Jie Tang 0003, Jean-Luc Gaudiot
NPC1
2022 Neural Network Pruning by Recurrent Weights for Finance Market
abstract
Convolutional Neural Networks (CNNs) and deep learning technology are applied in current financial market to rapidly promote the development of finance market and Internet economy. The continuous development of neural networks with more hidden layers improves the performance but increases the computational complexity. Generally, channel pruning methods are useful to compact neural networks. However, typical channel pruning methods would remove layers by mistake due to the static pruning ratio of manual setting, which could destroy the whole structure of neural networks. It is difficult to improve the ratio of compressing neural networks only by pruning channels while maintaining good network structures. Therefore, we propose a novel neural Networks Pruning by Recurrent Weights ( NPRW ) that can repeatedly evaluate the significance of weights and adaptively adjust them to compress neural networks within acceptable loss of accuracy. The recurrent weights with low sensitivity are compulsorily set to zero by evaluating the magnitude of weights, and pruned network only uses a few significant weights. Then, we add the regularization to the scaling factors on neural networks, in which recurrent weights with high sensitivity can be dynamically updated and weights of low sensitivity stay at zero invariably. By this way, the significance of channels can be quantitatively evaluated by recurrent weights. It has been verified with typical neural networks of LeNet, VGGNet, and ResNet on multiple benchmark datasets involving stock index futures, digital recognition, and image classification. The pruned LeNet-5 achieves the 58.9% reduction amount of parameters with 0.29% loss of total accuracy for Shanghai and Shenzhen 300 stock index futures. As for the CIFAR-10, the pruned VGG-19 reduces more than 50% FLOPs, and the decrease of network accuracy is less than 0.5%. In addition, the pruned ResNet-164 tested on the SVHN reduces more than 58% FLOPs with relative improvement on accuracy by 0.11%.
Songwen Pei, Yusheng Wu, Meikang Qiu
ACM Trans. Internet Techn.1
2021 STARS: Spatial Temporal Graph Convolution Network for Action Recognition System on FPGAs
abstract
Graph convolution neural network is one of the hot demanding research areas in the last few years. Due to the issues of data irregularity and computation complexity driven by typical GNN networks, we propose a spatial temporal graph convolution network for action recognition system on FPGA(STARS). STARS has redesigned several computing kernels based on the original layers of ST-GCN and adopted specific algorithm with optimization strategies for different kernels. To optimize the performance of accelerator, the ping-pong buffers for data transmission and dynamic quantification for model inference are implemented. The effectiveness of STARS driven accelerator is verified on Xilinx Pynq-Z1 prototyping board.
Songwen Pei, Xianrong Wang, Sheng Liang
COMPSAC1
2021 Genetic scheduling policy on codelet model
abstract
Summary The Codelet Model is a fine‐grained event‐driven hybrid parallel model inspired by dataflow, whose computing performance depends on the scheduling policy. An approximate optimal codelet scheduling policy based on the features of the task graphs is important to accelerate the performance of dataflow computer system. Therefore, we have proposed an adaptive genetic scheduling policy (GSP) for codelet by improving a “pure” genetic algorithm (IPGA) for given tasks with complex dependencies. It is verified that the genetic scheduling policy is effective according to bunches of experimental results.
Songwen Pei, Linhua Jiang, Naixue Xiong, Jean-Luc Gaudiot
Concurr. Comput. Pract. Exp.1
2021 Heterogeneous computation in specific domain accelerations
Chen Liu 0001, Songwen Pei
Future Gener. Comput. Syst.2
2021 KNN-BLOCK DBSCAN: Fast Clustering for Large-Scale Data
abstract
Large-scale data clustering is an essential key for big data problem. However, no current existing approach is “optimal” for big data due to high complexity, which remains it a great challenge. In this article, a simple but fast approximate DBSCAN, namely, KNN-BLOCK DBSCAN, is proposed based on two findings: 1) the problem of identifying whether a point is a core point or not is, in fact, a kNN problem and 2) a point has a similar density distribution to its neighbors, and neighbor points are highly possible to be the same type (core point, border point, or noise). KNN-BLOCK DBSCAN uses a fast approximate kNN algorithm, namely, FLANN, to detect core-blocks (CBs), noncore-blocks, and noise-blocks within which all points have the same type, then a fast algorithm for merging CBs and assigning noncore points to proper clusters is also invented to speedup the clustering process. The experimental results show that KNN-BLOCK DBSCAN is an effective approximate DBSCAN algorithm with high accuracy, and outperforms other current variants of DBSCAN, including ρ-approximate DBSCAN and AnyDBC.
Yewang Chen, Lida Zhou, Songwen Pei, Zhiwen Yu 0002, Yi Chen 0007, Xin Liu 0011, Jixiang Du, Naixue Xiong
IEEE Trans. Syst. Man Cybern. Syst.3
2020 Neural Network Compression and Acceleration by Federated Pruning
Songwen Pei, Yusheng Wu, Meikang Qiu
ICA3PP (2)1
2020 An efficient dataflow accelerator for scientific applications
Xiaochun Ye, Xu Tan 0001, Meng Wu 0006, Yujing Feng, Hao Zhang 0009, Songwen Pei, Dongrui Fan
Future Gener. Comput. Syst.7
2020 3DACN: 3D Augmented convolutional network for time series data
Songwen Pei, Tianma Shen, Xianrong Wang, Chunhua Gu, Zhong Ning, Xiaochun Ye, Naixue Xiong
Inf. Sci.1
2019 DA-BERT: Enhancing Part-of-Speech Tagging of Aspect Sentiment Analysis Using BERT
Songwen Pei, Lulu Wang 0007, Tianma Shen, Zhong Ning
APPT1
2018 Teaching Autonomous Driving Using a Modular and Integrated Approach
abstract
Introduction: Teaching autonomous driving is a challenging task. Indeed, most existing autonomous driving teaching activities focus on a few of the technologies involved. This not only fails to provide a comprehensive coverage, but also sets a high entry barrier for students with different backgrounds. Objective: The primary objective of this study is to present a modular, integrated approach towards teaching autonomous driving. Methods: We organize the technologies used in autonomous driving into modules. This is described in the textbook we have developed as well as a series of multimedia online lectures designed to provide technical overview for each module. Once the students have understood these modules, the experimental platforms for integration we have developed allow the students to fully understand how the modules interact with each other. Results: To verify this teaching approach, we present three case studies: an introductory class on autonomous driving for students with only a basic technology background; a new session in an existing embedded systems class to demonstrate how embedded system technologies can be applied towards autonomous driving; and an industry professional training session to quickly bring up experienced engineers to work in autonomous driving. The results show that students can maintain a high interest level and make great progress by starting with familiar concepts before moving onto other modules. Conclusions: Autonomous driving is not one single technology, but rather a complex system integrating many technologies. Our modular and integrated approach is an effective method in teaching autonomous driving.
Jie Tang 0003, Shaoshan Liu, Songwen Pei, Stéphane Zuckerman, Chen Liu 0001, Weisong Shi, Jean-Luc Gaudiot
COMPSAC (1)3
2018 Localized Traffic Sign Detection with Multi-scale Deconvolution Networks
abstract
Autonomous driving is becoming a future practical lifestyle greatly driven by deep learning. Specifically, an effective traffic sign detection by deep learning plays a critical role for it. However, different countries have different sets of traffic signs, making localized traffic sign recognition model training a tedious and daunting task. To address the issues of taking amount of time to compute complicate algorithm and low ratio of detecting blurred and sub-pixel images of localized traffic signs, we propose Multi-Scale Deconvolution Networks (MDN), which flexibly combines multi-scale convolutional neural network with deconvolution sub-network, leading to efficient and reliable localized traffic sign recognition model training. It is demonstrated that the proposed MDN is effective compared with classical algorithms on the benchmarks of the localized traffic sign, such as Chinese Traffic Sign Dataset (CTSD), and the German Traffic Sign Benchmarks (GTSRB).
Songwen Pei, Fuwu Tang, Yanfei Ji, Zhong Ning
COMPSAC (1)1
2018 Decentralized Clustering by Finding Loose and Distributed Density Cores
Yewang Chen, Shengyu Tang, Lida Zhou, Cheng Wang 0020, Jixiang Du, Tian Wang 0001, Songwen Pei
Inf. Sci.7
2018 DHeat: A Density Heat-Based Algorithm for Clustering With Effective Radius
abstract
Density-based clustering is one of the most popular paradigms of existing clustering approaches, most approaches of this kind, such as DBSCAN, recognize clusters of data characterized by a fixed scanning radius. However, some flaws are caused by the fixed scanning radius, e.g., the determination of a proper scanning radius is nontrivial. In order to solve these problems, we revise DBSCAN, Meanshift, DPeak, etc. based on two new features, i.e., effective radius and density heat (DHeat). Generally, we name these revised clustering algorithms as DHeat. The underlying idea is based on two assumptions: 1) the existence of clusters is raised by the nonuniformity of data distribution, and the density of one data point within its r-neighborhood is proportional to the volume of the neighborhood provided the density distribution is uniform and 2) each cluster can be divided into different density layers, such as edges, shallow inner, deep inner, etc.; the deeper inner of a point locates, the higher density of that point. The experiments conducted on various test cases show that the advantage of DHeat lies in its good performance and the self-adapting scanning radius.
Yewang Chen, Shengyu Tang, Songwen Pei, Cheng Wang 0020, Jixiang Du, Naixue Xiong
IEEE Trans. Syst. Man Cybern. Syst.3
2010 A GPU-based computing framework for CSCW
abstract
Graphics processing units (GPUs) have evolved from fixed graphics pipeline processors into more flexible and powerful data-parallel processors. Their ever-increasing computing power makes them an attractive platform for high performance computing at a low cost. Up to the present, most efforts that exploit GPUs are graphical and scientific applications. Nevertheless, little attention has been paid to harnessing these highly parallel devices to support collaborative work and design. In this paper, we propose a GPU-based framework based on stream computing technology, which aims at providing the computation service more efficiently among the collaborators in the CSCW system. The framework consists of two parts: A CPU-based data service mainly presiding over management and a GPU-based compute service primarily focus on computation.
Gang Chen 0012, Guobo Li, Baifeng Wu, Songwen Pei
CSCWD4
2009 GPGPU supported cooperative acceleration in molecular dynamics
abstract
Molecular dynamics simulations have become a significant computational approach to study complicated physical phenomena at the atomic level. Nevertheless, accurate simulations are limited in size and timescale by the available computing resources, which make the simulations very time-consuming. This consequentially leads to tremendous computational requirements. Therefore, the need for speeding up this process is crucial. In this paper, we present a novel implementation to accelerate molecular dynamics simulations with GPGPU (general purpose graphics processing unit). Our goal is to reduce the total computational time of MD simulations at a very high performance/cost ratio with the introduction of the GPGPU algorithm. This is motivated by their enhanced programmability, attractive cost/performance ratio and incredible growth in speed. To demonstrate that GPGPUs already provide an inexpensive alternative to scientific applications, we have used AMD's Brook+ streaming programming environment to implement a new parallel algorithm. Our experimental results show the novel approach achieves speedup by the factor of fifteen compared to the corresponding sequential implementation.
Gang Chen 0012, Guobo Li, Songwen Pei, Baifeng Wu
CSCWD3
2009 A self-embedded watermarking scheme based on relationship function of corresponding inter-blocks DCT coefficient
abstract
In the realm of computer supported cooperative work in design (CSCWD), how to ensure the authenticity and integrity of an image plays important roles. This paper presents a novel semi-fragile image watermarking scheme for authenticating and recovering image content. The scheme also shows strong robustness on those images' content reserved operations. It can precisely detect and locate malicious operations and tampers, and recover main content of tampered image region. Firstly, it assigns the exclusive precursor block and successor block for every block to be a block circle link, and regards the function mapping result of DCT coefficients of neighboring blocks as watermark, and embeds it into middle-frequency domain of DCT coefficients of successor block to form a watermarked image. Secondly, it verifies the blocks, which have been maliciously operated and tampered, by detecting whether the DCT coefficients of neighboring blocks satisfy the relationship function after operations. Thirdly, it uses corresponding relationship function of neighboring blocks in the link to estimate DCT direct current coefficients (DC) and low-frequency coefficients of the tampered blocks and recovers their main content. The experimental results show that the scheme is effective and feasible.
Guobo Li, Songwen Pei, Gang Chen 0012, Wenjun Cao, Baifeng Wu
CSCWD2
2008 Prototyping system of codec for novel 2-D continuous barcode
abstract
To address the limits of capacity and the disadvantages of motional scanning capability of traditional two-dimensional (2-D) barcodes, a novel 2- D continuous barcode (2-D CoBe) with improving the structure of traditional 2-D barcodes is proposed. The novel barcode not only enable to store a large mount of data with indefinitive capacity, but also support realtime decoding. Furthermore, a prototyping system of codec including encoding sub-system and decoding sub-system is examined and implemented. In this codec system, 2-D CoBe can be generated and printed on papers by encoding software sub-system. Besides, it can be recognized and decoded by a hand-held device with decoding sub-system in real-time. Based on the experimental results of encoding binary voice data into 2-D CoBe and decoding it with a hand-held device, it is verified that the codec system for 2-D CoBe is practical and effective.
Songwen Pei, Guobo Li, Gang Chen 0012, Baifeng Wu
CSCWD1
2007 Novel Collaborative Automated Testing Framework Using DDF*
Songwen Pei, Baifeng Wu, Kun Zhu 0005
CDVE1