Xiaofeng Zou

dblp:150/6859 · DBLP profile ↗
← Back
35ranked-venue papers
7as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RVLLM-Bench: A Comprehensive Benchmark for Large Language Model Inference with RISC-V Vector Extension
Zhilu Pan, Xiaofeng Zou, Panfeng Chen, Hui Li 0046, Yanhao Wang 0001
DASFAA (6)2
2026 Binocular dual attention interaction siamese network for diabetic retinopathy grading
Yanfei Guo, Yuncui Wang, Fei Ma 0004, Jing Meng 0001, Xiaofeng Zou
Inf. Sci.7
2025 DSPlacer: DSP Placement for FPGA-based CNN Accelerator
abstract
Deploying convolutional neural networks (CNNs) on hardware platforms like Field Programmable Gate Arrays (FPGAs) has garnered significant attention due to their inherent flexibility and parallelism. Achieving optimal timing closure remains a critical challenge, as placement directly impacts clock frequency and throughput. Existing approaches often face scalability issues with large designs or fail to formalize placement rules into automated algorithms. In this paper, we propose DSPlacer, a novel DSP placement framework designed for diverse CNN accelerator architectures in the context of FPGA design. The proposed approach iteratively optimizes the placement of datapath DSPs to enhance timing performance. To achieve this, DSPlacer integrates several advanced techniques, including graph convolutional network-based datapath DSP identification, DSP graph construction, min-cost-flow DSP assignment, and integer linear programming (ILP)-based cascade constraint legalization. These techniques collectively address two key requirements for datapath DSP placement: (1) cascading datapath DSPs to achieve a compact layout, and (2) preserving direct datapath information between the processing system and programmable logic. The framework has been evaluated on multiple academic benchmarks and compared against AMD Xilinx Vivado 2020.2 and AMF-Placer 2.0. Experimental results demonstrate that DSPlacer improves Worst Negative Slack (WNS) by 32% and 65%, respectively, highlighting its efficacy and superiority.
Baohui Xie, Xinrui Zhu, Yuan Pu 0001, Tongkai Wu, Xiaofeng Zou, Bei Yu 0001, Tinghuan Chen
DAC6
2025 A Compact Full Pipeline Architecture of SM3 Algorithm with High throughput and High Efficiency
abstract
This paper presents a compact full pipeline architecture of SM3 algorithm with high throughput and high efficiency, which is applied in high-performance security computing and large-scale data stream processing. The algorithm is reconstructed and a full pipeline computing architecture is proposed to achieve high throughput. In order to optimize the timing of algorithm, a compact architecture of compression function is further proposed, in which the critical path is reorganized and the adder delay of the path is optimized using CSA. Registers multiplexing is employed on the architecture of message word expansion for reducing area cost. Additionally, block RAM is employed to replace logic of polling and shifting, further optimizing the timing and reducing the LUT resources cost. The proposed architecture is implemented on a Xilinx Virtex-7 FPGA, and experimental results show that the designed SM3 achieves a peak frequency of 243 MHz, a high throughput of 124.42 Gbps, and a high efficiency of 15.22 Mbps/slice.
Ruiting Wang, Xiaofeng Zou, Dunshan Yu
ISCAS5
2025 Analytic Continual Test-Time Adaptation for Multi-Modality Corruption
Hongxin Wei, Zhiping Lin 0001, Xiaofeng Zou, Cen Chen 0002, Huiping Zhuang
ACM Multimedia5
2025 PCSViT: Efficient and hardware friendly Pyramid Vision Transformer with channel and spatial self-attentions
Xiaofeng Zou, Yuanxi Peng, Xinye Cao
Neurocomputing1
2025 An Adaptive and Scalable Framework for Resource-Efficient Deployment of Mixture of Experts in LLM-Based Intelligent IoT Networks
abstract
The exponential growth of the Internet of Things (IoT) necessitates the deployment of large-scale models capable of processing the complex and diverse data generated by IoT devices. However, the substantial memory requirements of these models pose significant challenges, especially in scenarios where rapid decision-making and low-latency responses are critical. To address these challenges, we propose three innovative strategies for optimizing large model usage in IoT environments. The first strategy is an adaptive loading scheme, which enables dynamic loading of individual model experts. The second strategy involves an expert-by-expert loading approach, further enhancing the ability to load experts as needed, which optimizes memory usage and accelerates computations. The third strategy employs an interlayer expert reuse mechanism, facilitating the efficient reuse of experts across different layers, thus enhancing response rates without compromising model accuracy. Importantly, these strategies can be directly applied to Mixture of Experts (MoE) large language models without requiring additional training, thereby providing a seamless and efficient solution for leveraging these models in memory-constrained, high-performance IoT environments.
Chengxu Liu 0003, Yangfan Li 0001, Cen Chen 0002, Hailan Kuang, Xiaolin Ma, Xiaofeng Zou, Jing Liu 0032, Zhaoyuan Zhang
IEEE Internet Things J.6
2025 CST-ViT: Cascaded Spatio-Temporal Redundancy Elimination for Efficient Vision Transformers on Edge IoT Devices
abstract
Transformer-based models have demonstrated outstanding performance in video understanding tasks due to their capacity to capture long-range dependencies. However, their high computational cost, along with the massive volume of streaming video data, presents significant challenges for real-time deployment on resource-constrained edge devices integrated into internet of things (IoT) systems. Existing approaches typically eliminate spatial or temporal redundancy in isolation, failing to fully exploit the inherent spatio-temporal similarity in video data. To address this limitation, we propose CST-ViT, a cascaded spatio-temporal redundancy elimination framework that jointly reduces dynamic temporal and intra-frame spatial redundancy. CST-ViT incorporates three gating modules: the direct temporal gate for matching unchanged backgrounds, the offset temporal gate for capturing motion-related changes, and the spatial gate for intra-frame similarity matching. Together with a spatiotemporal caching and token reuse mechanism, CST-ViT enables efficient token filtering and computation reuse. Experimental results show that CST-ViT reduces computation by 55.88% with no loss in accuracy, and achieves up to a 74.75% reduction in computation with less than 1% accuracy degradation, outperforming state-of-the-art methods in terms of accuracy–efficiency trade-off for video transformers.
Qinyu Wang 0002, Xiaofeng Zou, Chuang Li 0004, Yujie Peng, Heshi Wang, Yanhua Wen, Minaer Yeerlan, Cen Chen 0002
IEEE Internet Things J.2
2025 SFP: Similarity-based filter pruning for deep neural networks
RenGang Li, Chaoyao Shen, Xiaofeng Zou, Jiuyang Wang, Nanjun Li
Inf. Sci.5
2025 DRViT: A dynamic redundancy-aware vision transformer accelerator via algorithm and architecture co-design on FPGA
Xiangfeng Sun, Yuan-Ting Zhang, Xiaofeng Zou, Ziqian Zeng, Huiping Zhuang
J. Parallel Distributed Comput.4
2025 Seesaw: A 4096-bit vector processor for accelerating Kyber based on RISC-V ISA extensions
Xiaofeng Zou, Yuanxi Peng, Lingjun Kong
Parallel Comput.1
2025 Subkv: Quantizing Long Context KV Cache for Sub-Billion Parameter Language Models on Edge Devices
abstract
ABSTRACT Background Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their substantial computational and memory requirements present significant challenges for widespread deployment on edge devices. Motivation In long‐context scenarios, even sub‐billion parameter LLMs face unavoidable memory and performance bottlenecks due to inefficient KV Cache utilization. Existing quantization methods fail to address these challenges effectively. Method This paper addresses these challenges by introducing advanced quantization techniques tailored for sub‐billion parameter LLMs. It specifically targets reducing memory consumption through the conversion of the model's KV Cache to lower‐bit integers. We present SubKV, a quantization method specifically designed to optimize the KV Cache in sub‐billion parameter LLMs. Our analysis reveals distinct distributional differences in the magnitude of key and value caches. Leveraging this insight, we apply Per‐Channel Quantization to the key cache and Per‐Token Quantization to the value cache. Furthermore, we introduce the Dynamic Window Quantization method to enhance attention computations. To mitigate the extreme sensitivity of the first token, we also introduce Attention Sink‐Aware Quantization. Results Experimental results demonstrate that SubKV significantly reduces the KV Cache size during long context inference while maintaining model performance, offering superior results to existing KV Cache quantization methods.
Ziqian Zeng, Tao Zhang 0019, Zhengdong Lu, Huiping Zhuang, Hongen Shao, Sin G. Teo, Xiaofeng Zou
Softw. Pract. Exp.8
2025 ReViT: Vision Transformer Accelerator With Reconfigurable Semantic-Aware Differential Attention
abstract
While vision transformers (ViTs) have continued to achieve new milestones in computer vision, their complicated network architectures with high computation and memory costs have hindered their deployment on resource-limited edge devices. Some customized accelerators have been proposed to accelerate the execution of ViTs, achieving improved performance with reduced energy consumption. However, these approaches utilize flattened attention mechanisms and ignore the inherent hierarchical visual semantics in images. In this work, we conduct a thorough analysis of hierarchical visual semantics in real-world images, revealing opportunities and challenges of leveraging visual semantics to accelerate ViTs. We propose ReViT, a systematic algorithm and architecture co-design approach, which aims to exploit the visual semantics to accelerate ViTs. Our proposed algorithm can leverage the same semantic class with strong feature similarity to reduce computation and communication in a differential attention mechanism, and support the semantic-aware attention efficiently. A novel dedicated architecture is designed to support the proposed algorithm and translate it into performance improvements. Moreover, we propose an efficient execution dataflow to alleviate workload imbalance and maximize hardware utilization. ReViT opens new directions for accelerating ViTs by exploring the underlying visual semantics of images. ReViT gains an average of 2.3$\boldsymbol{\times}$speedup and 3.6$\boldsymbol{\times}$energy efficiency over state-of-the-art ViT accelerators.
Xiaofeng Zou, Cen Chen 0002, Hongen Shao, Qinyu Wang 0002, Xiaobin Zhuang, Yangfan Li 0001, Keqin Li 0001
IEEE Trans. Computers1
2025 SimDiff: Point Cloud Acceleration by Utilizing Spatial Similarity and Differential Execution
abstract
Point cloud neural networks are gaining increasing attention in emerging 3-D computer vision applications, such as autonomous driving, robotics, and virtual reality. Many customized accelerators for 3-D point clouds have been developed to pursue superior time and energy efficiencies. In this work, we reveal that spatially adjacent points in a 3-D point cloud show similar feature values and relationships, implying substantial redundant computations and memory accesses, while which have been previously ignored. To reduce such redundancies, we propose SimDiff, an algorithm-accelerator co-design framework that boosts 3-D point cloud processing by cleverly leveraging spatial similarity toward excellent speedup and energy efficiency. On the algorithm side, we design a novel similarity-aware differential point cloud neural network (dubbed SD-PCNet). Differing from the standard flow of mainstream point cloud networks, it abstracts a brand-new execution flow for point cloud processing by utilizing spatial similarity among points and dynamic differential execution. On the accelerator side, we propose SD-PCAcc, a supporting accelerator to convert algorithm-level redundancy reductions into performance enhancements. On the deployment side, we propose efficient strategies for network-to-accelerator mapping and scheduling, high-bandwidth memory (HBM) channel allocation, and core component reconfiguration, facilitating the proposed methodologies into practical implementation. Extensive evaluation results show that, with preserved accuracy, our SimDiff gains an average of$3.2\times $speedup and$3.1\times $energy efficiency compared to the state-of-the-art competitors.
Yangfan Li 0001, Mengquan Li, Cen Chen 0002, Xiaofeng Zou, Hongen Shao, Fengxiao Tang, Kenli Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 HEN: a novel hybrid explainable neural network based framework for robust network intrusion detection
Wei Wei 0006, Sijin Chen, Cen Chen 0002, Heshi Wang, Jing Liu 0032, Zhongyao Cheng, Xiaofeng Zou
Sci. China Inf. Sci.7
2023 Point Cloud Acceleration by Exploiting Geometric Similarity
abstract
Deep learning on point clouds has attracted increasing attention for various emerging 3D computer vision applications, such as autonomous driving, robotics, and virtual reality. These applications interact with people in real-time on edge devices and thus require low latency and low energy. To accelerate the execution of deep neural networks (DNNs) on point clouds, some customized accelerators have been proposed, which achieved a significantly higher performance with reduced energy consumption than GPUs and existing DNN accelerators.
Cen Chen 0002, Xiaofeng Zou, Hongen Shao, Yangfan Li 0001, Kenli Li 0001
MICRO2
2023 DGSLN: Differentiable graph structure learning neural network for robust graph representations
Xiaofeng Zou, Kenli Li 0001, Cen Chen 0002, Xulei Yang, Wei Wei 0006, Keqin Li 0001
Inf. Sci.1
2023 CoIn: Correlation Induced Clustering for Cognition of High Dimensional Bioinformatics Data
abstract
Analysis of high dimensional biomedical data such as microarray gene expression data and mass spectrometry images, is crucial to provide better medical services including cancer subtyping, protein homology detection, etc. Clustering is a fundamental cognitive task which aims to group unlabeled data into multiple clusters based on their intrinsic similarities. However, for most clustering methods, including the most widely used K-means algorithm, all features of the high dimensional data are considered equally in relevance, which distorts the performance when clustering high-dimensional data where there exist many redundant variables and correlated variables. In this paper, we aim at addressing the problem of the high dimensional bioinformatics data clustering and propose a new correlation induced clustering, CoIn, to capture complex correlations among high dimensional data and guarantee the correlation consistency within each cluster. We evaluate the proposed method on a high dimensional mass spectrometry dataset of liver cancer tumor to explore the metabolic differences on tissues and discover the intra-tumor heterogeneity (ITH). By comparing the results of baselines and ours, it has been found that our method produces more explainable and understandable results for clinical analysis, which demonstrates the proposed clustering paradigm has the potential with application to knowledge discovery in high dimensional bioinformatics data.
Zeng Zeng, Ziyuan Zhao, Kaixin Xu, Yangfan Li 0001, Cen Chen 0002, Xiaofeng Zou, Yulan Wang 0004, Wei Wei 0006, Pierce K. H. Chow, Xiaoli Li 0001
IEEE J. Biomed. Health Informatics6
2022 ReGNN: A Redundancy-Eliminated Graph Neural Networks Accelerator
abstract
Graph neural networks (GNNs), which extend conventional deep learning technologies to process graph-structured data, have shown its powerful graph representation learning ability. Existing typical GNNs utilize neighborhood message passing mechanism based on neural networks that updates target vertex representations by aggregating feature messages from neighboring source vertices. To accelerate the computations of GNNs, some customized accelerators, which follow the neighborhood aggregation computation pattern for each vertex, have been proposed. Through analysis, we observe that a naive implementation of the neighborhood aggregation results in redundant computations and communications.In this paper, we propose a novel redundancy-eliminated GNN accelerator, shortly termed as ReGNN. ReGNN is supported by an algorithm and architecture co-design. We first propose a dynamic redundancy-eliminated neighborhood message passing algorithm for GNNs. Then a novel architecture is designed to support the proposed algorithm and transform the redundancy elimination into performance improvement. ReGNN is also a configurable pipelined architecture that can be configured to support different GNN variants. In terms of the same computations, ReGNN provides the same accuracy as traditional GNNs. To the best of our knowledge, ReGNN is the first accelerator that can eliminate computation redundancy in GNNs. Our proposed ReGNN system gains an average of 9.1× speedup and 8.9× energy efficiency over state-of-the-art GNN accelerators.
Cen Chen 0002, Kenli Li 0001, Yangfan Li 0001, Xiaofeng Zou
HPCA4
2022 Latent Vector Prototypes Guided Conditional Face Synthesis
abstract
Recent advances in deep neural networks, especially in generative adversarial networks (GAN), have shown remarkable progress in face image generations. However, most of the existing face image generators can only synthesize random face images, but are not able to control the attributes of the generated face images. Though conditional GAN based methods can manipulate the attributes to some extent, but can only generate low-resolution face images up to 256 × 256. In this study, based on StyleGAN, one of the state-of-the-art image generators for synthesizing high-quality face images, we propose a simple but efficient approach to generate high-resolution and hyper-realistic face images with any desired attribute. By training an attribute classifier to assign attribute labels to given synthesized face images, we build the links between latent vectors and face attributes. In such a way, the latent vectors can be grouped into different clusters, one cluster corresponding to one face attribute, respectively. We then extract the prototypes for the clusters, which are used to control the attribute of the generated face image. Extensive experiments demonstrate the effectiveness of the proposed approach for high-quality face image generation with predefined attributes.
Qiyu Wei, Xulei Yang, Tong Sang, Huijiao Wang, Xiaofeng Zou, Zhongyao Cheng, Ziyuan Zhao, Zeng Zeng
ICIP5
2022 Multifunctional Module Design Based on Hybrid CMOS-Memristor Logic Circuit
abstract
With the rapid development of computer technology, logical circuits play an important role in many fields. Semiconductor transistors, as the basic logic circuit unit of modern computer hardware structure, also encounter more and more technical limitations in the highly integrated field such as computer chips. In order to continue Moore's Law effectively, higher requirements need to be put forward for the basic logic unit. The emergence of memristor opens up a new way to explore advanced computing architecture. This paper presents a MRL- based logic circuit in which input and output logic is provided by high and low voltage. Logic gate circuit elements such as XOR, NAND, NOR are implemented in conjunction with CMOS inverters. Through the construction of basic ratio logic gate, the adder and multiplier is simplified, and the result of signal simulation using LTspice software is completely correct. Finally, a multifunctional logic module is innovatively designed.
Xiaofeng Zou
ICSS3
2022 Algorithms and architecture support of degree-based quantization for graph neural networks
Yilong Guo, Xiaofeng Zou, Xulei Yang, Yuandong Gu
J. Syst. Archit.3
2022 Exploring Structural Knowledge for Automated Visual Inspection of Moving Trains
abstract
Deep learning methods are becoming the de-facto standard for generic visual recognition in the literature. However, their adaptations to industrial scenarios, such as visual recognition for machines, product streamlines, etc., which consist of countless components, have not been investigated well yet. Compared with the generic object detection, there is some strong structural knowledge in these scenarios (e.g., fixed relative positions of components, component relationships, etc.). A case worth exploring could be automated visual inspection for trains, where there are various correlated components. However, the dominant object detection paradigm is limited by treating the visual features of each object region separately without considering common sense knowledge among objects. In this article, we propose a novel automated visual inspection framework for trains exploring structural knowledge for train component detection, which is called SKTCD. SKTCD is an end-to-end trainable framework, in which the visual features of train components and structural knowledge (including hierarchical scene contexts and spatial-aware component relationships) are jointly exploited for train component detection. We propose novel residual multiple gated recurrent units (Res-MGRUs) that can optimally fuse the visual features of train components and messages from the structural knowledge in a weighted-recurrent way. In order to verify the feasibility of SKTCD, a dataset that contains high-resolution images captured from moving trains has been collected, in which 18 590 critical train components are manually annotated. Extensive experiments on this dataset and on the PASCAL VOC dataset have demonstrated that SKTCD outperforms the existing challenging baselines significantly. The dataset as well as the source code can be downloaded online (https://github.com/smartprobe/SKCD).
Cen Chen 0002, Xiaofeng Zou, Zeng Zeng, Zhongyao Cheng, Le Zhang 0001, Steven C. H. Hoi
IEEE Trans. Cybern.2
2022 Multilevel Attention Based U-Shape Graph Neural Network for Point Clouds Learning
abstract
With the popularity of 3-D sensors in industrial Internet of Things (IIoT), point clouds learning is increasingly important. In this article, we propose a novel multilevel attention based U-shape graph neural network (MAUGNN) for point clouds learning, which can effectively learn the features from low-level to high-level and fuse multiple-level features based on the graph neural networks and attention mechanism. There are three parts in MAUGNN: encoder, decoder, and connections. In the encoder and decoder, we design an attention-based graph convolution to explore the structural information for point clouds. During the encoder, a structure-aware attention pooling is proposed to support down-sampling on point cloud data. To adaptively fuse coarse-grained features from the encoder and fine-grained features from the decoder together, we also propose a structure-aware attention skip connection mechanism. Extensive experiments on popular point cloud datasets demonstrate the superior performance of our MAUGNN over state-of-the-art baselines.
Xiaofeng Zou, Kenli Li 0001, Cen Chen 0002
IEEE Trans. Ind. Informatics1
2022 Multi-Task Y-Shaped Graph Neural Network for Point Cloud Learning in Autonomous Driving
abstract
Point cloud, an efficient 3D object representation, plays an indispensable role in autonomous driving technologies, such as object avoidance, localization, and map building. The analysis of point clouds (e.g., 3D segmentation) is essential to exploit the informative value of point clouds for such applications. The main challenge remains to effectively and completely extract high-level point cloud feature representations. To this end, we present a novel multi-task Y-shaped graph neural network to explore 3D point clouds, referred to as MTYGNN. By extending the conventional U-Net, MTYGNN contains two main branches to simultaneously perform classification and segmentation tasks in point clouds. Meanwhile, the classification prediction is fused together with the semantic features as the scene context to make the segmentation task more accurate. Furthermore, we consider the homoscedastic uncertainty of each task to calculate the weights of multiple loss functions to ensure that tasks do not negatively interfere with each other. The proposed MTYGNN is evaluated on popular point cloud datasets in traffic scenarios. Experimental results demonstrate that our framework outperforms the state-of-the-art baseline methods.
Xiaofeng Zou, Kenli Li 0001, Yangfan Li 0001, Wei Wei 0006, Cen Chen 0002
IEEE Trans. Intell. Transp. Syst.1
2022 Hierarchical Semantic Graph Reasoning for Train Component Detection
abstract
Recently, deep learning-based approaches have achieved superior performance on object detection applications. However, object detection for industrial scenarios, where the objects may also have some structures and the structured patterns are normally presented in a hierarchical way, is not well investigated yet. In this work, we propose a novel deep learning-based method, hierarchical graphical reasoning (HGR), which utilizes the hierarchical structures of trains for train component detection. HGR contains multiple graphical reasoning branches, each of which is utilized to conduct graphical reasoning for one cluster of train components based on their sizes. In each branch, the visual appearances and structures of train components are considered jointly with our proposed novel densely connected dual-gated recurrent units (Dense-DGRUs). To the best of our knowledge, HGR is the first kind of framework that explores hierarchical structures among objects for object detection. We have collected a data set of 1130 images captured from moving trains, in which 17 334 train components are manually annotated with bounding boxes. Based on this data set, we carry out extensive experiments that have demonstrated our proposed HGR outperforms the existing state-of-the-art baselines significantly. The data set and the source code can be downloaded online at https://github.com/ChengZY/HGR.
Cen Chen 0002, Kenli Li 0001, Xiaofeng Zou, Zhongyao Cheng, Wei Wei 0006, Qi Tian 0001, Zeng Zeng
IEEE Trans. Neural Networks Learn. Syst.3
2022 Determinantal point process-based new radio unlicensed link scheduling for multi-access edge computing
Chigang Xing, Yangfan Li 0001, Cen Chen 0002, Fangmin Li, Zeng Zeng, Xiaofeng Zou
World Wide Web6
2021 DyGNN: Algorithm and Architecture Support of Dynamic Pruning for Graph Neural Networks
abstract
Recently, graph neural networks (GNNs) have achieved great success for graph representation learning tasks. Enlightened by the fact that numerous message passing redundancies exist in GNNs, we propose DyGNN, which speeds up GNNs by reducing redundancies. DyGNN is supported by an algorithm and architecture co-design. The proposed algorithm can dynamically prune vertices and edges during execution without accuracy loss. An architecture is designed to support dynamic pruning and transform it into performance improvement. DyGNN opens new directions for accelerating GNNs by pruning vertices and edges. DyGNN gains average $2\times$ speedup with accuracy improvement of 4% compared with state-of-the-art GNN accelerators.
Cen Chen 0002, Kenli Li 0001, Xiaofeng Zou, Yangfan Li 0001
DAC3
2021 Multiple local 3D CNNs for region-based prediction in smart cities
Yibi Chen, Xiaofeng Zou, Kenli Li 0001, Keqin Li 0001, Xulei Yang, Cen Chen 0002
Inf. Sci.2
2020 Multiple Balance Subsets Stacking for Imbalanced Healthcare Datasets
abstract
Accurate prediction is highly important for clinical decision making and early treatment. In this paper, we study the imbalanced data problem in prediction, a key challenge existing in the healthcare area. Imbalanced datasets bias classifiers towards the majority class, leading to an unsatisfied classification prediction performance on the minority class, which is known as imbalance problem. Existing imbalance learning methods may suffer from issues like information loss, overfitting, and high training time cost. To tackle these issues, we propose a novel ensemble learning method called Multiple bAlance Subsets Stacking (MASS) by exploiting a multiple balance subsets construction strategy. Furthermore, we improve MASS with introducing parallelism (Parallel MASS) to reduce the training time cost. We evaluate MASS on three real-world healthcare datasets, and experimental results demonstrate that its prediction performance outperforms the state-of-art methods in terms of AUC, F1-score and MCC. Through the speedup analysis, Parallel MASS reduces the training time cost greatly on large dataset, and its speedup increases as the data size grows.
Yachao Shao, Tao Zhao 0007, Xiaofeng Zou, Xiaoming Fu 0001
ICPADS4
2020 Multi-task cascade deep convolutional neural networks for large-scale commodity recognition
Xiaofeng Zou, Liqian Zhou, Kenli Li 0001, Aijia Ouyang, Cen Chen 0002
Neural Comput. Appl.1
2020 Citywide Traffic Flow Prediction Based on Multiple Gated Spatio-temporal Convolutional Neural Networks
abstract
Traffic flow prediction is crucial for public safety and traffic management, and remains a big challenge because of many complicated factors, e.g., multiple spatio-temporal dependencies, holidays, and weather. Some work leveraged 2D convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) to explore spatial relations and temporal relations, respectively, which outperformed the classical approaches. However, it is hard for these work to model spatio-temporal relations jointly. To tackle this, some studies utilized LSTMs to connect high-level layers of CNNs, but left the spatio-temporal correlations not fully exploited in low-level layers. In this work, we propose novel spatio-temporal CNNs to extract spatio-temporal features simultaneously from low-level to high-level layers, and propose a novel gated scheme to control the spatio-temporal features that should be propagated through the hierarchy of layers. Based on these, we propose an end-to-end framework, multiple gated spatio-temporal CNNs (MGSTC), for citywide traffic flow prediction. MGSTC can explore multiple spatio-temporal dependencies through multiple gated spatio-temporal CNN branches, and combine the spatio-temporal features with external factors dynamically. Extensive experiments on two real traffic datasets demonstrates that MGSTC outperforms other state-of-the-art baselines.
Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Xiaofeng Zou, Keqin Li 0001, Zeng Zeng
ACM Trans. Knowl. Discov. Data4
2019 Gated Residual Recurrent Graph Neural Networks for Traffic Prediction
abstract
Traffic prediction is of great importance to traffic management and public safety, and very challenging as it is affected by many complex factors, such as spatial dependency of complicated road networks and temporal dynamics, and many more. The factors make traffic prediction a challenging task due to the uncertainty and complexity of traffic states. In the literature, many research works have applied deep learning methods on traffic prediction problems combining convolutional neural networks (CNNs) with recurrent neural networks (RNNs), which CNNs are utilized for spatial dependency and RNNs for temporal dynamics. However, such combinations cannot capture the connectivity and globality of traffic networks. In this paper, we first propose to adopt residual recurrent graph neural networks (Res-RGNN) that can capture graph-based spatial dependencies and temporal dynamics jointly. Due to gradient vanishing, RNNs are hard to capture periodic temporal correlations. Hence, we further propose a novel hop scheme into Res-RGNN to utilize the periodic temporal dependencies. Based on Res-RGNN and hop Res-RGNN, we finally propose a novel end-to-end multiple Res-RGNNs framework, referred to as “MRes-RGNN”, for traffic prediction. Experimental results on two traffic datasets have demonstrated that the proposed MRes-RGNN outperforms state-of-the-art methods significantly.
Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Xiaofeng Zou, Jie Wang 0042, Zeng Zeng
AAAI4
2018 Exploiting Spatio-Temporal Correlations with Multiple 3D Convolutional Neural Networks for Citywide Vehicle Flow Prediction
abstract
Predicting vehicle flows is of great importance to traffic management and public safety in smart cities, and very challenging as it is affected by many complex factors, such as spatio-temporal dependencies with external factors (e.g., holidays, events and weather). Recently, deep learning has shown remarkable performance on traditional challenging tasks, such as image classification, due to its powerful feature learning capabilities. Some works have utilized LSTMs to connect the high-level layers of 2D convolutional neural networks (CNNs) to learn the spatio-temporal features, and have shown better performance as compared to many classical methods in traffic prediction. However, these works only build temporal connections on the high-level features at the top layer while leaving the spatio-temporal correlations in the low-level layers not fully exploited. In this paper, we propose to apply 3D CNNs to learn the spatio-temporal correlation features jointly from low-level to high-level layers for traffic data. We also design an end-to-end structure, named as MST3D, especially for vehicle flow prediction. MST3D can learn spatial and multiple temporal dependencies jointly by multiple 3D CNNs, combine the learned features with external factors and assign different weights to different branches dynamically. To the best of our knowledge, it is the first framework that utilizes 3D CNNs for traffic prediction. Experiments on two vehicle flow datasets Beijing and New York City have demonstrated that the proposed framework, MST3D, outperforms the state-of-the-art methods.
Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Guizi Chen, Xiaofeng Zou, Xulei Yang, Ramaseshan C. Vijay, Jiashi Feng, Zeng Zeng
ICDM5
2014 Benefits of Adding Hardware Support for Broadcast and Reduce Operations in MPSoC Applications
abstract
MPI has been used as a parallel programming model for supercomputers and clusters and recently in MultiProcessor Systems-on-Chip (MPSoC). One component of MPI is collective communication and its performance is key for certain parallel applications to achieve good speedups. Previous work showed that, with synthetic communication-only benchmarks, communication improvements of up to 11.4-fold and 22-fold for broadcast and reduce operations, respectively, can be achieved by providing hardware support at the network level in a Network-on-Chip (NoC). However, these numbers do not provide a good estimation of the advantage for actual applications, as there are other factors that affect performance besides communications, such as computation. To this end, we extend our previous work by evaluating the impact of hardware support over a set of five parallel application kernels of varying computation-to-communication ratios. By introducing some useful computation to the performance evaluation, we obtain more representative results of the benefits of adding hardware support for broadcast and reduce operations. The experiments show that applications with lower computation-to-communication ratios benefit the most from hardware support as they highly depend on efficient collective communications to achieve better scalability. We also extend our work by doing more analysis on clock frequency, resource usage, power, and energy. The results show reasonable scalability for resource utilization and power in the network interfaces as the number of channels increases and that, even though more power is dissipated in the network interfaces due to the added hardware, the total energy used can still be less if the actual speedup is sufficient. The application kernels are executed in a 24-embedded-processor system distributed across four FPGAs.
Yuanxi Peng, Manuel Saldaña, Christopher A. Madill, Xiaofeng Zou, Paul Chow
ACM Trans. Reconfigurable Technol. Syst.4