Cen Chen 0002

dblp:152/6215-2 · DBLP profile ↗
← Back
73ranked-venue papers
15as first author
59since 2021 · last 2026
0000-0003-1389-0148ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 4 first-author · 21 since 2021Systems, architecture and hardware · 18 · 7 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 9 since 2021Computer networks · 7 · 6 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior
abstract
Large Language Models (LLMs) with long chain-of-thought (CoT) capability, termed Reasoning Models, demonstrate superior intricate problem-solving abilities through multi-step long CoT reasoning. To create a dual-capability model with long CoT capability and domain-specific knowledge without substantial computational and data costs, model merging emerges as a highly resource-efficient method. However, significant challenges lie in merging domain-specific LLMs with long CoT ones since nowadays merging methods suffer from reasoning capability degradation, even gibberish output and output collapse. To overcome this, we introduce RCP-Merging: Merging Long Chain-of-Thought Models with Domain-Specific Models by Considering Reasoning Capability as Prior, a novel merging framework designed to integrate domain-specific LLMs with long CoT capability, meanwhile maintaining model performance in the original domain. Treating reasoning model weights as foundational prior, our method utilizes a reasoning capability indicator to preserve core long CoT capability model weights while selectively merging essential domain-specific weights. We conducted extensive experiments on Qwen2.5-7B, Llama3.1-8B, and Qwen2.5-1.5B models in BioMedicine and Finance domains. Our results show that RCP-Merging successfully merges a reasoning model with domain-specific ones, improving domain task performance by 9.5% and 9.2% over state-of-the-art methods, without significantly harming the original long CoT reasoning capability.
Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
AAAI4
2026 MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
abstract
Long Chain-of-Thought (CoT) reasoning has significantly advanced the capabilities of Large Language Models (LLMs), but this progress is accompanied by substantial memory and latency overhead from the extensive Key-Value (KV) cache.Although KV cache quantization is a promising compression technique, existing low-bit quantization methods often exhibit severe performance degradation on complex reasoning tasks.Fixed-precision quantization struggles to handle outlier channels in the key cache, while current mixed-precision strategies fail to accurately identify components requiring high-precision representation.We find that an effective low-bit KV cache quantization strategy must consider two factors: a key channel's intrinsic quantization difficulty and its relevance to the query.Based on this insight, we propose MixKVQ, a novel quantization method that introduces a lightweight, query-aware algorithm to identify and preserve critical key channels that need higher precision, while applying per-token quantization for value cache.Experiments on complex reasoning datasets demonstrate that our approach significantly outperforms existing low-bit methods, achieving performance comparable to a fullprecision baseline at a substantially reduced memory footprint.The source code is available at https://github.com/ZeroNLP/MixKVQ.
Tao Zhang 0019, Ziqian Zeng, Huiping Zhuang, Cen Chen 0002
ACL (1)5
2026 PointShuffler: Accelerating Point Cloud Neural Networks on General-Purpose GPUs
abstract
Point Cloud Neural Networks (PCNNs) have emerged as a vital tool for latency-sensitive 3D perception applications, such as autonomous driving and AR/VR. However, their inherent computational redundancy—arising from excessive global sampling/search operations and repeated feature updates/aggregations caused by shared neighbors—severely constrains execution efficiency. More critically, conventional redundancy elimination methods usually introduce operations that are highly GPU-unfriendly, resulting in high memory overhead, increased branch divergence, irregular memory access, and serial dependencies, which together pose a significant challenge to PCNN acceleration.
Yangfan Li 0001, Zhengjie Jin, Mengquan Li, Fengxiao Tang, Ming Zhao 0007, Cen Chen 0002
EuroSys7
2026 Task completion-oriented service migration for connected autonomous vehicles in multi-server edge computing
Jing Liu 0032, Jieyi Deng, Longxin Zhang, Qiushi Cao, Wei Hu 0001, Cen Chen 0002, Keqin Li 0001
Comput. Networks6
2026 Game-Theoretic Bandwidth Allocation and Task Offloading in Cloud-Edge Collaboration
abstract
The rapid growth of Internet of Things (IoT) devices has imposed higher demands on computational capabilities, which traditional cloud computing struggles to meet in real-time scenarios due to latency issues. Mobile edge computing (MEC) addresses these challenges by processing data at the network edge, thereby reducing latency and enhancing computational efficiency. However, MEC alone is insufficient for handling complex tasks, requiring more robust solutions. This paper proposes a hybrid cloud-edge computing framework that enhances system performance by integrating cloud and edge computing. A game-theoretic model is used to optimize wireless bandwidth allocation, and a Stackelberg game mechanism is introduced to incentivize task offloading. This approach orchestrates resource allocation and task offloading dynamics through game theory, ensuring cost minimization and delay requirements are met while fostering cloud-edge collaboration. Theoretical analysis demonstrates the existence of Nash equilibria in both layers of the game, ensuring the system’s stability and effectiveness in complex environments. Based on this, the GA-based resource allocation and offloading (GRAO) algorithm, and the iterative game-theoretic offloading (IGTO) algorithm are proposed. Experimental results validate the proposed algorithms, showing that the IGTO algorithm reduces the average cost for mobile devices (MDs) by 49.8% compared to the best baseline, while enhancing overall performance for both MEC servers and the cloud.
Zhao Tong 0001, Yuanyang Zhang, Jing Mei, Cen Chen 0002, Keqin Li 0001
IEEE Internet Things J.4
2025 SEA: Low-Resource Safety Alignment for Multimodal Large Language Models via Synthetic Embeddings
abstract
Multimodal Large Language Models (MLLMs) have serious security vulnerabilities.While safety alignment using multimodal datasets consisting of text and data of additional modalities can effectively enhance MLLM's security, it is costly to construct these datasets.Existing low-resource security alignment methods, including textual alignment, have been found to struggle with the security risks posed by additional modalities.To address this, we propose Synthetic Embedding augmented safety Alignment (SEA), which optimizes embeddings of additional modality through gradient updates to expand textual datasets.This enables multimodal safety alignment training even when only textual data is available.Extensive experiments on image, video, and audio-based MLLMs demonstrate that SEA can synthesize a high-quality embedding on a single RTX3090 GPU within 24 seconds.SEA significantly improves the security of MLLMs when faced with threats from additional modalities.To assess the security risks introduced by video and audio, we also introduced a new benchmark called VA-SafetyBench.High attack success rates across multiple MLLMs validate its challenge.Our code and data will be available at https://github.com/ZeroNLP/SEA.This paper contains harmful data and modelgenerated content that can be offensive in nature.
Weikai Lu, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
ACL (1)4
2025 PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration
abstract
Ziqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziqian Zeng, Junyao Yang, Zhengdong Lu, Huiping Zhuang, Cen Chen 0002
ACL (1)7
2025 GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models
abstract
Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a “chosen” and a “rejected” response. Compared to the “rejected” responses, the “chosen” responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the “rejected” responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs.
Tao Zhang 0019, Ziqian Zeng, YuxiangXiao YuxiangXiao, Huiping Zhuang, Cen Chen 0002, James R. Foulds, Shimei Pan
ACL (1)5
2025 RewardDS: Privacy-Preserving Fine-Tuning for Large Language Models via Reward Driven Data Synthesis
abstract
The success of large language models (LLMs) has attracted many individuals to fine-tune them for domain-specific tasks by uploading their data.However, in sensitive areas like healthcare and finance, privacy concerns often arise.One promising solution is to generate synthetic data with Differential Privacy (DP) guarantees to replace private data.However, these synthetic data contain significant flawed data, which are considered as noise.Existing solutions typically rely on naive filtering by comparing ROUGE-L scores or embedding similarities, which are ineffective in addressing the noise.To address this issue, we propose RewardDS, a novel privacy-preserving framework that fine-tunes a reward proxy model and uses reward signals to guide the synthetic data generation.Our RewardDS introduces two key modules, Reward Guided Filtering and Self-Optimizing Refinement, to both filter and refine the synthetic data, effectively mitigating the noise.Extensive experiments across medical, financial, and code generation domains demonstrate the effectiveness of our method.
Chengming Shi, Junyao Yang, Huiping Zhuang, Cen Chen 0002, Ziqian Zeng
EMNLP7
2025 Semantic Shift Estimation via Dual-Projection and Classifier Reconstruction for Exemplar-Free Class-Incremental Learning
abstract
Exemplar-Free Class-Incremental Learning (EFCIL) aims to sequentially learn from distinct categories without retaining exemplars but easily suffers from catastrophic forgetting of learned knowledge. While existing EFCIL methods leverage knowledge distillation to alleviate forgetting, they still face two critical challenges: semantic shift and decision bias. Specifically, the embeddings of old tasks shift in the embedding space after learning new tasks, and the classifier becomes biased towards new tasks due to training solely with new data, hindering the balance between old and new knowledge. To address these issues, we propose the Dual-Projection Shift Estimation and Classifier Reconstruction (DPCR) approach for EFCIL. DPCR effectively estimates semantic shift through a dual-projection, which combines a learnable transformation with a row-space projection to capture both task-wise and category-wise shifts. Furthermore, to mitigate decision bias, DPCR employs ridge regression to reformulate a classifier reconstruction process. This reconstruction exploits previous in covariance and prototype of each class after calibration with estimated shift, thereby reducing decision bias. Extensive experiments demonstrate that, on various datasets, DPCR effectively balances old and new tasks, outperforming state-of-the-art EFCIL methods. Our codes are available at https://github.com/RHe502/ICML25-DPCR.
Run He, Di Fang 0004, Yawen Cui, Ming Li 0011, Cen Chen 0002, Ziqian Zeng, Huiping Zhuang
ICML6
2025 L3A: Label-Augmented Analytic Adaptation for Multi-Label Class Incremental Learning
abstract
Class-incremental learning (CIL) enables models to learn new classes continually without forgetting previously acquired knowledge. Multi-label CIL (MLCIL) extends CIL to a real-world scenario where each sample may belong to multiple classes, introducing several challenges: label absence, which leads to incomplete historical information due to missing labels, and class imbalance, which results in the model bias toward majority classes. To address these challenges, we propose Label-Augmented Analytic Adaptation (L3A), an exemplar-free approach without storing past samples. L3A integrates two key modules. The pseudo-label (PL) module implements label augmentation by generating pseudo-labels for current phase samples, addressing the label absence problem. The weighted analytic classifier (WAC) derives a closed-form solution for neural networks. It introduces sample-specific weights to adaptively balance the class contribution and mitigate class imbalance. Experiments on MS-COCO and PASCAL VOC datasets demonstrate that L3A outperforms existing methods in MLCIL tasks. Our code is available at https://github.com/scut-zx/L3A.
Run He, Chen Jiao, Di Fang 0004, Ming Li 0073, Ziqian Zeng, Cen Chen 0002, Huiping Zhuang
ICML7
2025 GSDNet: Revisiting Incomplete Multimodality-Diffusion Emotion Recognition from the Perspective of Graph Spectrum
abstract
Multimodal Emotion Recognition (MER) combines technologies from multiple fields (e.g., computer vision, natural language processing, and audio signal processing), aiming to infer an individual's emotional state by analyzing information from different sources (i.e., video, audio, and text). Compared with single modality, by fusing complementary semantic information from different modalities, the model can obtain more robust knowledge representation. However, the modality missing problem limits the performance of MERC in practical scenarios. Recent work has achieved impressive performance on modality completion using graph neural networks and diffusion models, respectively. This inspires us to combine these two dimensions in the completion network to obtain more powerful representation capabilities. However, we argue that directly running a full-rank score-based diffusion model on the entire graph adjacency matrix space may adversely affect the learning process of the diffusion model. This is because the model assumes a direct relationship between each pair of nodes and ignores local structural features and sparse connections between nodes, thereby significantly reducing the quality of the generated data. Based on the above ideas, we propose a novel Graph Spectral Diffusion Network (GSDNet), which utilizes a low-rank score-based diffusion model to map Gaussian noise to the graph spectral distribution space of missing modalities and recover the missing data according to its original distribution. Extensive experiments have demonstrated that GSDNet achieves state-of-the-art emotion recognition performance in various modality loss scenarios.
Yuntao Shou, Wei Ai 0001, Cen Chen 0002, Keqin Li 0001
IJCAI5
2025 Analytic Continual Test-Time Adaptation for Multi-Modality Corruption
Hongxin Wei, Zhiping Lin 0001, Xiaofeng Zou, Cen Chen 0002, Huiping Zhuang
ACM Multimedia6
2025 Efficient Resource Allocation Algorithm for Maximizing Operator Profit in 5G Edge Computing Network
Jing Liu 0032, Chunhua Deng, Longxin Zhang, Cen Chen 0002, Keqin Li 0001
J. Grid Comput.5
2025 An Adaptive and Scalable Framework for Resource-Efficient Deployment of Mixture of Experts in LLM-Based Intelligent IoT Networks
abstract
The exponential growth of the Internet of Things (IoT) necessitates the deployment of large-scale models capable of processing the complex and diverse data generated by IoT devices. However, the substantial memory requirements of these models pose significant challenges, especially in scenarios where rapid decision-making and low-latency responses are critical. To address these challenges, we propose three innovative strategies for optimizing large model usage in IoT environments. The first strategy is an adaptive loading scheme, which enables dynamic loading of individual model experts. The second strategy involves an expert-by-expert loading approach, further enhancing the ability to load experts as needed, which optimizes memory usage and accelerates computations. The third strategy employs an interlayer expert reuse mechanism, facilitating the efficient reuse of experts across different layers, thus enhancing response rates without compromising model accuracy. Importantly, these strategies can be directly applied to Mixture of Experts (MoE) large language models without requiring additional training, thereby providing a seamless and efficient solution for leveraging these models in memory-constrained, high-performance IoT environments.
Chengxu Liu 0003, Yangfan Li 0001, Cen Chen 0002, Hailan Kuang, Xiaolin Ma, Xiaofeng Zou, Jing Liu 0032, Zhaoyuan Zhang
IEEE Internet Things J.3
2025 CST-ViT: Cascaded Spatio-Temporal Redundancy Elimination for Efficient Vision Transformers on Edge IoT Devices
abstract
Transformer-based models have demonstrated outstanding performance in video understanding tasks due to their capacity to capture long-range dependencies. However, their high computational cost, along with the massive volume of streaming video data, presents significant challenges for real-time deployment on resource-constrained edge devices integrated into internet of things (IoT) systems. Existing approaches typically eliminate spatial or temporal redundancy in isolation, failing to fully exploit the inherent spatio-temporal similarity in video data. To address this limitation, we propose CST-ViT, a cascaded spatio-temporal redundancy elimination framework that jointly reduces dynamic temporal and intra-frame spatial redundancy. CST-ViT incorporates three gating modules: the direct temporal gate for matching unchanged backgrounds, the offset temporal gate for capturing motion-related changes, and the spatial gate for intra-frame similarity matching. Together with a spatiotemporal caching and token reuse mechanism, CST-ViT enables efficient token filtering and computation reuse. Experimental results show that CST-ViT reduces computation by 55.88% with no loss in accuracy, and achieves up to a 74.75% reduction in computation with less than 1% accuracy degradation, outperforming state-of-the-art methods in terms of accuracy–efficiency trade-off for video transformers.
Qinyu Wang 0002, Xiaofeng Zou, Chuang Li 0004, Yujie Peng, Heshi Wang, Yanhua Wen, Minaer Yeerlan, Cen Chen 0002
IEEE Internet Things J.8
2025 REAL: Representation enhanced analytic learning for exemplar-free class-incremental learning
Run He, Di Fang 0004, Yizhu Chen, Kai Tong, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau, Huiping Zhuang
Knowl. Based Syst.5
2025 ReViT: Vision Transformer Accelerator With Reconfigurable Semantic-Aware Differential Attention
abstract
While vision transformers (ViTs) have continued to achieve new milestones in computer vision, their complicated network architectures with high computation and memory costs have hindered their deployment on resource-limited edge devices. Some customized accelerators have been proposed to accelerate the execution of ViTs, achieving improved performance with reduced energy consumption. However, these approaches utilize flattened attention mechanisms and ignore the inherent hierarchical visual semantics in images. In this work, we conduct a thorough analysis of hierarchical visual semantics in real-world images, revealing opportunities and challenges of leveraging visual semantics to accelerate ViTs. We propose ReViT, a systematic algorithm and architecture co-design approach, which aims to exploit the visual semantics to accelerate ViTs. Our proposed algorithm can leverage the same semantic class with strong feature similarity to reduce computation and communication in a differential attention mechanism, and support the semantic-aware attention efficiently. A novel dedicated architecture is designed to support the proposed algorithm and translate it into performance improvements. Moreover, we propose an efficient execution dataflow to alleviate workload imbalance and maximize hardware utilization. ReViT opens new directions for accelerating ViTs by exploring the underlying visual semantics of images. ReViT gains an average of 2.3$\boldsymbol{\times}$speedup and 3.6$\boldsymbol{\times}$energy efficiency over state-of-the-art ViT accelerators.
Xiaofeng Zou, Cen Chen 0002, Hongen Shao, Qinyu Wang 0002, Xiaobin Zhuang, Yangfan Li 0001, Keqin Li 0001
IEEE Trans. Computers2
2025 SimDiff: Point Cloud Acceleration by Utilizing Spatial Similarity and Differential Execution
abstract
Point cloud neural networks are gaining increasing attention in emerging 3-D computer vision applications, such as autonomous driving, robotics, and virtual reality. Many customized accelerators for 3-D point clouds have been developed to pursue superior time and energy efficiencies. In this work, we reveal that spatially adjacent points in a 3-D point cloud show similar feature values and relationships, implying substantial redundant computations and memory accesses, while which have been previously ignored. To reduce such redundancies, we propose SimDiff, an algorithm-accelerator co-design framework that boosts 3-D point cloud processing by cleverly leveraging spatial similarity toward excellent speedup and energy efficiency. On the algorithm side, we design a novel similarity-aware differential point cloud neural network (dubbed SD-PCNet). Differing from the standard flow of mainstream point cloud networks, it abstracts a brand-new execution flow for point cloud processing by utilizing spatial similarity among points and dynamic differential execution. On the accelerator side, we propose SD-PCAcc, a supporting accelerator to convert algorithm-level redundancy reductions into performance enhancements. On the deployment side, we propose efficient strategies for network-to-accelerator mapping and scheduling, high-bandwidth memory (HBM) channel allocation, and core component reconfiguration, facilitating the proposed methodologies into practical implementation. Extensive evaluation results show that, with preserved accuracy, our SimDiff gains an average of$3.2\times $speedup and$3.1\times $energy efficiency compared to the state-of-the-art competitors.
Yangfan Li 0001, Mengquan Li, Cen Chen 0002, Xiaofeng Zou, Hongen Shao, Fengxiao Tang, Kenli Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2024 ConsistentEE: A Consistent and Hardness-Guided Early Exiting Method for Accelerating Language Models Inference
abstract
Early Exiting is one of the most popular methods to achieve efficient inference. Current early exiting methods adopt the (weighted) sum of the cross entropy loss of all internal classifiers as the objective function during training, imposing all these classifiers to predict all instances correctly. However, during inference, as long as one internal classifier predicts an instance correctly, it can accelerate without losing accuracy. Thus, there is a notable gap between training and inference. We propose ConsistentEE, an early exiting method that is consistent in training and inference. ConsistentEE formulates the early exiting process as a reinforcement learning problem. A policy network is added to decide whether an instance should exit or continue. The training objective of ConsistentEE only requires each instance to be predicted correctly by one internal classifier. Additionally, we introduce the concept "Memorized Layer" to measure the hardness of an instance. We incorporate the memorized layer into reward function design, which allows "easy'' instances to focus more on acceleration while ``hard'' instances to focus more on accuracy. Experimental results show that our method outperforms other baselines on various natural language understanding and generation tasks using PLMs and LLMs as backbones respectively.
Ziqian Zeng, Yihuai Hong, Huiping Zhuang, Cen Chen 0002
AAAI5
2024 DS-AL: A Dual-Stream Analytic Learning for Exemplar-Free Class-Incremental Learning
abstract
Class-incremental learning (CIL) under an exemplar-free constraint has presented a significant challenge. Existing methods adhering to this constraint are prone to catastrophic forgetting, far more so than replay-based techniques that retain access to past samples. In this paper, to solve the exemplar-free CIL problem, we propose a Dual-Stream Analytic Learning (DS-AL) approach. The DS-AL contains a main stream offering an analytical (i.e., closed-form) linear solution, and a compensation stream improving the inherent under-fitting limitation due to adopting linear mapping. The main stream redefines the CIL problem into a Concatenated Recursive Least Squares (C-RLS) task, allowing an equivalence between the CIL and its joint-learning counterpart. The compensation stream is governed by a Dual-Activation Compensation (DAC) module. This module re-activates the embedding with a different activation function from the main stream one, and seeks fitting compensation by projecting the embedding to the null space of the main stream's linear mapping. Empirical results demonstrate that the DS-AL, despite being an exemplar-free technique, delivers performance comparable with or better than that of replay-based methods across various datasets, including CIFAR-100, ImageNet-100 and ImageNet-Full. Additionally, the C-RLS' equivalent property allows the DS-AL to execute CIL in a phase-invariant manner. This is evidenced by a never-before-seen 500-phase CIL ImageNet task, which performs on a level identical to a 5-phase one. Our codes are available at https://github.com/ZHUANGHP/Analytic-continual-learning.
Huiping Zhuang, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Zhiping Lin 0001
AAAI5
2024 Hybrid Explainable Network Intrusion Detection Framework Based on Shapley Additive Explanations
abstract
With the rapid advancements in network technology and automation processes, the threats posed by cyberattacks have become increasingly significant. To address these threats, numerous researchers have developed various network intrusion detection systems (NIDS) to monitor network traffic. However, with the continuous complexification of 5G networks and the exponential increase in network traffic, the emergence of new attacks alongside the lack of interpretability in NIDS posed challenges to the performance and efficiency of network intrusion detection. To tackle these issues, this paper proposes a hybrid explainable network intrusion detection framework that combines the strengths of supervised and unsupervised learning, enabling effective detection of emerging attacks within the network. Specifically, we utilize a Light Gradient Boosting Machine (LightGBM) model for supervised learning, followed by the SHapley Additive exPlanations (SHAP) method for model explanation and feature selection. Additionally, we employ a network utilizing Convolutional Neural Network and Long short term memory for encoding and decoding (ACLNet) purposes in unsupervised learning. Finally, the results from these two learning processes are integrated for anomaly detection. The simulation experimental results on the NSL-KDD dataset verify that our approach generates detection performance on par with state-of-the-art methods while offering significantly enhanced interpretability and improving detection efficiency.
Sijin Chen, Jing Liu 0032, Cen Chen 0002, Songyu Xie, Zhongyao Cheng
ISPA3
2024 GACL: Exemplar-Free Generalized Analytic Continual Learning
abstract
Class incremental learning (CIL) trains a network on sequential tasks with separated categories in each task but suffers from catastrophic forgetting, where models quickly lose previously learned knowledge when acquiring new tasks. The generalized CIL (GCIL) aims to address the CIL problem in a more real-world scenario, where incoming data have mixed data categories and unknown sample size distribution. Existing attempts for the GCIL either have poor performance or invade data privacy by saving exemplars. In this paper, we propose a new exemplar-free GCIL technique named generalized analytic continual learning (GACL). The GACL adopts analytic learning (a gradient-free training technique) and delivers an analytical (i.e., closed-form) solution to the GCIL scenario. This solution is derived via decomposing the incoming data into exposed and unexposed classes, thereby attaining a weight-invariant property, a rare yet valuable property supporting an equivalence between incremental learning and its joint training. Such an equivalence is crucial in GCIL settings as data distributions among different tasks no longer pose challenges to adopting our GACL. Theoretically, this equivalence property is validated through matrix analysis tools. Empirically, we conduct extensive experiments where, compared with existing GCIL methods, our GACL exhibits a consistently leading performance across various datasets and GCIL settings. Source code is available at https://github.com/CHEN-YIZHU/GACL.
Huiping Zhuang, Yizhu Chen, Di Fang 0004, Run He, Kai Tong, Hongxin Wei, Ziqian Zeng, Cen Chen 0002
NeurIPS8
2024 F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental Learning
abstract
Online Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods show competitive results but requires extra memory for storing exemplars, while exemplar-free (i.e., data need not be stored for replay in production) methods are resource friendly but often lack accuracy. In this paper, we propose an exemplar-free approach—Forward-only Online Analytic Learning (F-OAL). Unlike traditional methods, F-OAL does not rely on back-propagation and is forward-only, significantly reducing memory usage and computational time. Cooperating with a pre-trained frozen encoder with Feature Fusion, F-OAL only needs to update a linear classifier by recursive least square. This approach simultaneously achieves high accuracy and low resource consumption. Extensive experiments on bench mark datasets demonstrate F-OAL’s robust performance in OCIL scenarios. Code is available at: https://github.com/liuyuchen-cz/F-OAL
Huiping Zhuang, Yuchen Liu 0001, Run He, Kai Tong, Ziqian Zeng, Cen Chen 0002, Yi Wang 0068, Lap-Pui Chau
NeurIPS6
2024 HEN: a novel hybrid explainable neural network based framework for robust network intrusion detection
Wei Wei 0006, Sijin Chen, Cen Chen 0002, Heshi Wang, Jing Liu 0032, Zhongyao Cheng, Xiaofeng Zou
Sci. China Inf. Sci.3
2024 Toward Compact and Robust Model Learning Under Dynamically Perturbed Environments
abstract
Network pruning has been widely studied to reduce the complexity of deep neural networks (DNNs) and hence speed up their inference. Unfortunately, most existing pruning methods ignore the changes in the model’s robustness before and after pruning, which makes pruned models vulnerable under dynamically perturbed environments (e.g., autonomous driving). Only a few works have explored the robustness of pruned models against adversarial attacks that significantly differ from perturbations in real-world scenarios. To bridge the gap between real-world applications and existing studies, in this work, we propose an adversarial pruning scheme, which automatically identifies and preserves robust channels to obtain robust pruned models that are suitable for practical deployment in dynamically perturbed environments. Specifically, to simulate real-world perturbations, we first employ multi-type adversarial attack samples and adversarial perturbation samples generated by an adversarial perturbation generator to create mixed noise samples. Then, we propose a plug-and-play feature scoring module and a novel contribution difference loss to evaluate the robustness of intermediate features dynamically. Next, to leverage robust intermediate features to identify robust channels, we have developed a simple but effective gating mechanism that evaluates the robustness of channels and preserves robust channels during training. Lastly, we compress the model in a layer-wise or block-wise manner. Compared to existing methods, our scheme enhances the robustness of the pruned model in a broader sense, making it better able to against dynamic perturbations in the real world. Extensive experimental results on well-known dataset benchmarks and popular network architectures demonstrate the effectiveness of our method.
Hui Luo 0002, Zhuangwei Zhuang, Yuanqing Li 0001, Mingkui Tan, Cen Chen 0002, Jianlin Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 CPCM: Contextual Point Cloud Modeling for Weakly-supervised Point Cloud Semantic Segmentation
abstract
We study the task of weakly-supervised point cloud semantic segmentation with sparse annotations (e.g., less than 0.1% points are labeled), aiming to reduce the expensive cost of dense annotations. Unfortunately, with extremely sparse annotated points, it is very difficult to extract both contextual and object information for scene understanding such as semantic segmentation. Motivated by masked modeling (e.g., MAE) in image and video representation learning, we seek to endow the power of masked modeling to learn contextual information from sparsely-annotated points. However, directly applying MAE to 3D point clouds with sparse annotations may fail to work. First, it is nontrivial to effectively mask out the informative visual context from 3D point clouds. Second, how to fully exploit the sparse annotations for context modeling remains an open question. In this paper, we propose a simple yet effective Contextual Point Cloud Modeling (CPCM) method that consists of two parts: a region-wise masking (Region-Mask) strategy and a contextual masked training (CMT) method. Specifically, RegionMask masks the point cloud continuously in geometric space to construct a meaningful masked prediction task for subsequent context learning. CMT disentangles the learning of supervised segmentation and unsupervised masked context prediction for effectively learning the very limited labeled points and mass unlabeled points, respectively. Extensive experiments on the widely-tested ScanNet V2 and S3DIS benchmarks demonstrate the superiority of CPCM over the state-of-the-art.
Lizhao Liu, Zhuangwei Zhuang, Shangxin Huang, Xunlong Xiao, Tianhang Xiang, Cen Chen 0002, Jingdong Wang 0001, Mingkui Tan
ICCV6
2023 COCO-TEACH: A Contrastive Co-Teaching Network For Incremental 3D Object Detection
abstract
Deep learning (DL) models for 3D object detection from point clouds have shown remarkable progress in various autonomous perception scenarios. However, the issue of catastrophic forgetting seriously hinders the deployment of these models in real-world applications where new classes are encountered over time. In order to address this issue, we present the Contrastive Co-Teaching Network (COCO-TEACH) framework for class-incremental 3D object detection. Our proposed framework consists of two teacher networks: a primary teacher network that detects old class objects in new data and provides them with pseudo-labels and an auxiliary teacher network that leverages the unlabelled objects in new data. The two teacher models transfer their learned knowledge to the target student model through a class-aware consistency loss. To enhance this transfer, a supervised contrastive loss is further incorporated into the loss function. We evaluate the performance of our proposed method against baseline methods through extensive experiments on two benchmark datasets. The results show that our proposed framework achieves state-of-the-art performance on incremental 3D object detection.
Zhongyao Cheng, Cen Chen 0002, Ziyuan Zhao, Peisheng Qian, Xiaoli Li 0001, Xulei Yang
ICIP2
2023 Spatial and Temporal Dual-Scale Adaptive Pruning for Point Cloud Videos
abstract
Point clouds, characterized by irregularity and disorder, are widely utilized in the domain of the Internet of Things, including applications such as terrain exploration and autonomous driving. 3D action recognition and semantic segmentation widely employ them. However, point cloud videos often exhibit massive data volume and contain substantial data redundancy. This situation is highly detrimental to the real-time applications of point clouds such as autonomous driving and virtual reality. Network pruning is imperative to mitigate the redundancy. This paper proposes a spatial and temporal dual-scale adaptive pruning (STDAP) method for point cloud videos based on the attention mechanism to reduce redundancy. This method operates in both temporal and spatial dimensions, enabling adaptive pruning of point cloud videos and cutting down the inference time and memory usage. Experimental results demonstrate that the proposed method outperforms state-of-the-art techniques regarding accuracy and inference time. It achieves notable acceleration while maintaining efficacy.
Songyu Xie, Jing Liu 0032, Cen Chen 0002, Zhongyao Cheng
ICPADS3
2023 Point Cloud Acceleration by Exploiting Geometric Similarity
abstract
Deep learning on point clouds has attracted increasing attention for various emerging 3D computer vision applications, such as autonomous driving, robotics, and virtual reality. These applications interact with people in real-time on edge devices and thus require low latency and low energy. To accelerate the execution of deep neural networks (DNNs) on point clouds, some customized accelerators have been proposed, which achieved a significantly higher performance with reduced energy consumption than GPUs and existing DNN accelerators.
Cen Chen 0002, Xiaofeng Zou, Hongen Shao, Yangfan Li 0001, Kenli Li 0001
MICRO1
2023 Learning discriminative multi-relation representations for multimodal sentiment analysis
Zemin Tang, Xu Zhou 0001, Yangfan Li 0001, Cen Chen 0002, Kenli Li 0001
Inf. Sci.5
2023 DGSLN: Differentiable graph structure learning neural network for robust graph representations
Xiaofeng Zou, Kenli Li 0001, Cen Chen 0002, Xulei Yang, Wei Wei 0006, Keqin Li 0001
Inf. Sci.3
2023 Intelligent energy-efficient scheduling with ant colony techniques for heterogeneous edge computing
Jing Liu 0032, Cen Chen 0002
J. Parallel Distributed Comput.3
2023 Toward Communication-Efficient Digital Twin via AI-Powered Transmission and Reconstruction
abstract
Digital twin technology has recently gathered pace in engineering communities as it allows for the convergence of the real structure and its digital counterpart. 3D point cloud data is a more effective way to describe the real world and to reconstruct the digital counterpart than the conventional 2D images or 360-degree images. Large-scale, e.g., city-scale digital twins, typically collect point cloud data via internet-of-things (IoT) devices and transmit it over wireless networks. However, the existing wireless transmission technology can not carry real-time point cloud transmission for digital twin reconstruction due to mass data volume, high processing overheads, and low delay-tolerance. We propose a novel artificial intelligence (AI) powered end-to-end framework, termed AIRec, for efficient digital twin communication from point cloud compression, wireless channel coding, and digital twin reconstruction. AIRec adopts the encoder-decoder architecture. In the encoder, a novel importance-aware pooling scheme is designed to adaptively select important points with learnable thresholds to reduce the transmission volume. We also design a novel noise-aware joint source and channel coding is proposed to adaptively adjust the transmission strategy based on SNR and map the features to error-resilient channel symbols for wireless transmission to achieve a good tradeoff between the transmission rate and reconstruction quality. The decoder can accurately reconstruct the digital twins from the received symbols. Extensive experiments of typical datasets and comparison with baselines show that we achieve a good reconstruction quality under$24\times $compression ratio.
Cen Chen 0002, Xulei Yang, Joey Tianyi Zhou, Tao Zhang 0019, Yangfan Li 0001
IEEE J. Sel. Areas Commun.2
2023 Maximizing the number of completed tasks in MEC considering time and energy constraints
Haijian Yu, Jing Liu 0032, Chunhua Deng, Cen Chen 0002, Keqin Li 0001
Soft Comput.4
2023 Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array
abstract
Transformer model architectures have recently received great interest in natural language, machine translation, and computer vision, where attention mechanisms are their building blocks. However, the attention mechanism is expensive because of its intensive matrix computations and complicated data flow. The existing hardware architecture has some disadvantages for the computing structure of attention, such as inflexibility and low efficiency. Most of the existing papers accelerate attention by reducing the amount of computation through various pruning algorithms, which will affect the results to a certain extent with different sparsity. This paper proposes the hardware accelerator for the multi-head attention (MHA) on field-programmable gate arrays (FPGAs) with reconfigurable architecture, efficient systolic array, and hardware-friendly radix-2 softmax. We propose a novel method called Four inputs Processing Element (FPE) to double the computation rate of the data-aware systolic array (SA) and make it efficient and load balance. Especially, the computation framework is well designed to ensure the utilization of SA efficiently. Our design is evaluated on a Xilinx Alveo U250 card, and the proposed architecture achieves 51.3×, 17.3× improvement in latency, and 54.4×, 17.9× energy savings compared to CPU and GPU.
Wenhua Ye, Xu Zhou 0001, Joey Tianyi Zhou, Cen Chen 0002, Kenli Li 0001
ACM Trans. Embed. Comput. Syst.4
2023 CoIn: Correlation Induced Clustering for Cognition of High Dimensional Bioinformatics Data
abstract
Analysis of high dimensional biomedical data such as microarray gene expression data and mass spectrometry images, is crucial to provide better medical services including cancer subtyping, protein homology detection, etc. Clustering is a fundamental cognitive task which aims to group unlabeled data into multiple clusters based on their intrinsic similarities. However, for most clustering methods, including the most widely used K-means algorithm, all features of the high dimensional data are considered equally in relevance, which distorts the performance when clustering high-dimensional data where there exist many redundant variables and correlated variables. In this paper, we aim at addressing the problem of the high dimensional bioinformatics data clustering and propose a new correlation induced clustering, CoIn, to capture complex correlations among high dimensional data and guarantee the correlation consistency within each cluster. We evaluate the proposed method on a high dimensional mass spectrometry dataset of liver cancer tumor to explore the metabolic differences on tissues and discover the intra-tumor heterogeneity (ITH). By comparing the results of baselines and ours, it has been found that our method produces more explainable and understandable results for clinical analysis, which demonstrates the proposed clustering paradigm has the potential with application to knowledge discovery in high dimensional bioinformatics data.
Zeng Zeng, Ziyuan Zhao, Kaixin Xu, Yangfan Li 0001, Cen Chen 0002, Xiaofeng Zou, Yulan Wang 0004, Wei Wei 0006, Pierce K. H. Chow, Xiaoli Li 0001
IEEE J. Biomed. Health Informatics5
2023 Cascade Graph Neural Networks for Few-Shot Learning on Point Clouds
abstract
Point cloud data, a flexible 3D object representation, is critical for various applications such as autonomous driving, robotics and remote sensing. Despite the recent success of deep neural networks (DNNs) on supervised point cloud analysis tasks, they still rely on tedious manual annotation of point clouds and cannot make predictions for new classes. Unlike few-shot learning for 2D images with the advantages of large-scale datasets and high-quality deep pre-trained models like ResNet, for 3D few-shot learning, obtaining discriminative representations of unseen classes with high intra-class similarity and inter-class difference is very challenging. To address this issue, this work proposes a novel cascade graph neural network for few-shot learning on point clouds, termed as CGNN, in which two cascade GNNs are adopted to extract the intra-object topological information and learn the inter-object relations respectively. To further increase the discriminability of point cloud features, we first design a novel discriminative edge label to model the intra-class similarity and inter-class dissimilarity based on channel-wise feature variance and class consistency. Second, we propose a novel few-shot circle loss which classifies the nodes into two subsets, i.e., support to support pairs and support to query pairs, and optimizes the pair-wise similarity on two subsets independently. Extensive experiments on benchmark CAD and real LiDAR point cloud datasets have demonstrated that CGNN improves accuracy by 5.98% over the state-of-the-art GNN-based few-shot classification methods.
Yangfan Li 0001, Cen Chen 0002, Weiquan Yan, Zhongyao Cheng, Hui Li Tan, Wenjie Zhang 0004
IEEE Trans. Intell. Transp. Syst.2
2023 CoRec: An Efficient Internet Behavior-based Recommendation Framework with Edge-cloud Collaboration on Deep Convolution Neural Networks
abstract
Both accurate and fast mobile recommendation systems based on click behaviors analysis are crucial in e-business. Deep learning has achieved state-of-the-art accuracy and the traditional wisdom often hosts these computation-intensive models in powerful cloud centers. However, the cloud-only approaches put significant computational pressure on cloud servers and increase the latency in heavy-load scenarios. Moreover, existing work often adopts RNN structures to model behaviors that suffer from low processing speed for under-utilization of parallel devices such as GPUs. In this work, we propose an efficient internet behavior-based recommendation framework with edge-cloud collaboration on deep CNNs (CoRec) to improve both the accuracy and speed for mobile recommendation. A novel convolutional interest network (CIN) improves the accuracy by modeling the long- and short-term interests and accelerates the prediction through parallel-friendly convolutions. To further improve the serving throughput and latency, a novel device-cloud collaboration strategy reduces workloads by pre-computing and caching long-term interests in the cloud offline and real-time computation of short-term interests in devices. Extensive experiments on real-world datasets show that CoRec significantly outperforms the state-of-the-art methods in accuracy and has achieved at least an order of magnitude improvement in latency and throughput compared to cloud-only RNN-based approaches for long behaviors.
Yangfan Li 0001, Kenli Li 0001, Wei Wei 0006, Joey Tianyi Zhou, Cen Chen 0002
ACM Trans. Sens. Networks5
2022 ReGNN: A Redundancy-Eliminated Graph Neural Networks Accelerator
abstract
Graph neural networks (GNNs), which extend conventional deep learning technologies to process graph-structured data, have shown its powerful graph representation learning ability. Existing typical GNNs utilize neighborhood message passing mechanism based on neural networks that updates target vertex representations by aggregating feature messages from neighboring source vertices. To accelerate the computations of GNNs, some customized accelerators, which follow the neighborhood aggregation computation pattern for each vertex, have been proposed. Through analysis, we observe that a naive implementation of the neighborhood aggregation results in redundant computations and communications.In this paper, we propose a novel redundancy-eliminated GNN accelerator, shortly termed as ReGNN. ReGNN is supported by an algorithm and architecture co-design. We first propose a dynamic redundancy-eliminated neighborhood message passing algorithm for GNNs. Then a novel architecture is designed to support the proposed algorithm and transform the redundancy elimination into performance improvement. ReGNN is also a configurable pipelined architecture that can be configured to support different GNN variants. In terms of the same computations, ReGNN provides the same accuracy as traditional GNNs. To the best of our knowledge, ReGNN is the first accelerator that can eliminate computation redundancy in GNNs. Our proposed ReGNN system gains an average of 9.1× speedup and 8.9× energy efficiency over state-of-the-art GNN accelerators.
Cen Chen 0002, Kenli Li 0001, Yangfan Li 0001, Xiaofeng Zou
HPCA1
2022 Latency-driven Model Placement for Efficient Edge Intelligence Service
abstract
Deep learning services are extensively required and have powerful expected effects in a wide range of applications, such as auto-self driving, voice assistant and so on.Traditionally, deep learning services are mainly provided based on cloud computing, referred to as cloud intelligence services, where the deep learning model is deployed in the cloud, and endusers need to upload data through wireless and core network when requesting training and inference services. However, the cloud computing-based deep learning services have deficiencies in latency, privacy, etc. For example, the deep learning service users do not want their private data to leak, and the privacy problem is difficult to solve when the data is uploading to the cloud. With the widespread use of the Internet of Things (IoT) and the rapid development of mobile devices such as smartphones and IoT sensors, large amounts of data need to be used for a variety of real-time deep learning services, such as target recognition and voice recognition for smart cities, smart medical care, and the Internet of Vehicles (IoVs). With a certain network bandwidth, a large amount of data uploaded to the cloud will cause network congestion and greatly increase the response time. To meet the requirements of low latency, researchers have begun to consider the deployment of deep learning services in edges, i.e., edge intelligence service.In edge intelligence services, the computation capability and memory of processors (or devices) are different from a large. At the same time, the requirement of memory size of deep neural network (DNN) models is increasing, such as the memory usage for Alexnet and Resnet are 2.12G and 16.20G separately. Also, in DNN model design, the branches are becoming common, which brings the parallelism. Deploying deep learning models on multiple processors can support the large-scale DNN models and the parallel implementation of DNN model, where the computation of a deep learning model can be conducted in parallel is a possible solution to improve the efficiency of edge intelligence services. The key point in edge intelligence services is how to partition and assign the implementation of the DNN model.In this paper, we propose a novel latency-driven deep learning model placement method for efficient edge intelligence service. Model placement contains two procedures: model partition and sub-models assignment. In our method, we first convert a DNN model into an execution graph, which is a directed acyclic graph (DAG), and propose a novel latency-driven multilevel graph partition for the model. Then the partitioned submodels are heuristically assigned to available processors. To the best of our knowledge, it is the first work that proposes latency-driven graph partition algorithms for model placement. Extensive experiments on several commonly used DNN models and synthetic datasets show that our method can achieve the lowest execution latency with low complexity compared with other state-of-the-art model placement methods.
Peiying Lin, Zhichen Shi, Cen Chen 0002, Kenli Li 0001
SERVICES4
2022 Introduction to the Special Issue on edge intelligence: Neurocomputing meets edge computing
Zeng Zeng, Cen Chen 0002, Bharadwaj Veeravalli, Keqin Li 0001, Joey Tianyi Zhou
Neurocomputing2
2022 Hierarchical Graph Neural Networks for Few-Shot Learning
abstract
Recent graph neural network (GNN) based methods for few-shot learning (FSL) represent the samples of interest as a fully-connected graph and conduct reasoning on the nodes flatly, which ignores the hierarchical correlations among nodes. However, real-world categories may have hierarchical structures, and for FSL, it is important to extract the distinguishing features of the categories from individual samples. To explore this, we propose a novel hierarchical graph neural network (HGNN) for FSL, which consists of three parts, i.e., bottom-up reasoning, top-down reasoning, and skip connections, to enable the efficient learning of multi-level relationships. For the bottom-up reasoning, we design intra-class k-nearest neighbor pooling (intra-class knnPool) and inter-class knnPool layers, to conduct hierarchical learning for both the intra- and inter-class nodes. For the top-down reasoning, we propose to utilize graph unpooling (gUnpool) layers to restore the down-sampled graph into its original size. Skip connections are proposed to fuse multi-level features for the final node classification. The parameters of HGNN are learned by episodic training with the signal of node losses, which aims to train a well-generalizable model for recognizing unseen classes with few labeled data. Experimental results on benchmark datasets have demonstrated that HGNN outperforms other state-of-the-art GNN based methods significantly, for both transductive and non-transductive FSL tasks. The dataset as well as the source code can be downloaded online1
Cen Chen 0002, Kenli Li 0001, Wei Wei 0006, Joey Tianyi Zhou, Zeng Zeng
IEEE Trans. Circuits Syst. Video Technol.1
2022 Exploring Structural Knowledge for Automated Visual Inspection of Moving Trains
abstract
Deep learning methods are becoming the de-facto standard for generic visual recognition in the literature. However, their adaptations to industrial scenarios, such as visual recognition for machines, product streamlines, etc., which consist of countless components, have not been investigated well yet. Compared with the generic object detection, there is some strong structural knowledge in these scenarios (e.g., fixed relative positions of components, component relationships, etc.). A case worth exploring could be automated visual inspection for trains, where there are various correlated components. However, the dominant object detection paradigm is limited by treating the visual features of each object region separately without considering common sense knowledge among objects. In this article, we propose a novel automated visual inspection framework for trains exploring structural knowledge for train component detection, which is called SKTCD. SKTCD is an end-to-end trainable framework, in which the visual features of train components and structural knowledge (including hierarchical scene contexts and spatial-aware component relationships) are jointly exploited for train component detection. We propose novel residual multiple gated recurrent units (Res-MGRUs) that can optimally fuse the visual features of train components and messages from the structural knowledge in a weighted-recurrent way. In order to verify the feasibility of SKTCD, a dataset that contains high-resolution images captured from moving trains has been collected, in which 18 590 critical train components are manually annotated. Extensive experiments on this dataset and on the PASCAL VOC dataset have demonstrated that SKTCD outperforms the existing challenging baselines significantly. The dataset as well as the source code can be downloaded online (https://github.com/smartprobe/SKCD).
Cen Chen 0002, Xiaofeng Zou, Zeng Zeng, Zhongyao Cheng, Le Zhang 0001, Steven C. H. Hoi
IEEE Trans. Cybern.1
2022 Robust Traffic Prediction From Spatial-Temporal Data Based on Conditional Distribution Learning
abstract
Traffic prediction based on massive speed data collected from traffic sensors plays an important role in traffic management. However, it is still challenging to obtain satisfactory performance due to the complex and dynamic spatial-temporal correlations among the data. Recently, many research works have demonstrated the effectiveness of graph neural networks (GNNs) for spatial-temporal modeling. However, such models are restricted by conditional distribution during training, and may not perform well when the target is outside the primary region of interest in the distribution. In this article, we address this problem with a stagewise learning mechanism, in which we redefine speed prediction as a conditional distribution learning followed by speed regression. We first perform a conditional distribution learning for each observed speed class, and then obtain speed prediction by optimizing regression learning, based on the learned conditional distribution. To effectively learn the conditional distribution, we introduce a mean-residue loss, consisting of two parts: 1) a mean loss, which penalizes the differences between the mean of the estimated conditional distribution and the ground truth and 2) a residue loss, which penalizes residue errors of the long tails in the distribution. To optimize the subsequent regression based on distribution information, we combine the mean absolute error (MAE) as another part of the loss function. We also incorporate a GNN-based architecture with our proposed learning mechanism. Mean-residue loss is employed to supervise the hidden speed representation in the network at each time interval, followed by a shared layer to recalibrate the hidden temporal dependencies in the conditional distribution. The experimental results based on three public traffic datasets have demonstrated that the effectiveness of the proposed method outperforms state-of-the-art methods.
Zeng Zeng, Wei Zhao 0035, Peisheng Qian, Yingjie Zhou 0001, Ziyuan Zhao, Cen Chen 0002, Cuntai Guan
IEEE Trans. Cybern.6
2022 Memory-Assistant Collaborative Language Understanding for Artificial Intelligence of Things
abstract
Artificial intelligence shows promising efforts in collaborating the language models with the artificial intelligence of things (AIoT), promoting the edging intelligence on natural language understanding. To adapt to the limited computational resources in AIoT, the large language models (e.g., transformer) are compressed into light-weight models, which always results in poor feature representation and unsatisfactory performance on downstream tasks, especially on those low-resource language understanding tasks. To address the above issues, we propose a method named memory-assistant multi-task learning (MAMT), where an auxiliary memory module is introduced to promote multitask learning (MT), which serves as a surrogate of target domain representation and performs instance-level weighted MT. More importantly, our MAMT module is in a plug-and-play fashion. Thus, researchers can plug in it to conduct collaborative training and plug it out for AIoT model inference without extra computation burdens. Experiments demonstrate that MAMT significantly improves the performance of light-weight transformer models and show its superiority over the state-of-the-arts on eight GLUE subtasks.
Ming Yan 0007, Cen Chen 0002, Jiawei Du 0002, Xi Peng 0001, Joey Tianyi Zhou, Zeng Zeng
IEEE Trans. Ind. Informatics2
2022 Multilevel Attention Based U-Shape Graph Neural Network for Point Clouds Learning
abstract
With the popularity of 3-D sensors in industrial Internet of Things (IIoT), point clouds learning is increasingly important. In this article, we propose a novel multilevel attention based U-shape graph neural network (MAUGNN) for point clouds learning, which can effectively learn the features from low-level to high-level and fuse multiple-level features based on the graph neural networks and attention mechanism. There are three parts in MAUGNN: encoder, decoder, and connections. In the encoder and decoder, we design an attention-based graph convolution to explore the structural information for point clouds. During the encoder, a structure-aware attention pooling is proposed to support down-sampling on point cloud data. To adaptively fuse coarse-grained features from the encoder and fine-grained features from the decoder together, we also propose a structure-aware attention skip connection mechanism. Extensive experiments on popular point cloud datasets demonstrate the superior performance of our MAUGNN over state-of-the-art baselines.
Xiaofeng Zou, Kenli Li 0001, Cen Chen 0002
IEEE Trans. Ind. Informatics3
2022 A Hybrid Deep Learning Based Framework for Component Defect Detection of Moving Trains
abstract
Defect detection of trains is of great significance for operation safety and maintenance efficiency for railway maintenance. Nowadays, China railway system utilizes high-speed line scan cameras to capture images of critical parts of moving trains. The visual inspection on the images still heavily relies on manual interpretation. To reduce the labor requirements, we propose a novel two-stage deep learning based framework for component defect detection of moving trains. The proposed framework is composed of two major successive stages: detecting train components by using our proposed hierarchical object detection scheme (HOD), and detecting component defects based on multiple neural networks and image processing methods. Our proposed HOD can effectively detect and localize train components from large to small in a hierarchical way. Furthermore, a gated feature fusion method that can extract and combine the hierarchical contextual features and spatial contexts is also proposed to improve the performance. To the best of our knowledge, it is the first time in the literature that component defect detection of moving trains is systematically analyzed. Extensive experiments on real images from China railway system have demonstrated that our framework outperforms the state-of-the-art baselines significantly.
Cen Chen 0002, Kenli Li 0001, Zhongyao Cheng, Francesco Piccialli, Steven C. H. Hoi, Zeng Zeng
IEEE Trans. Intell. Transp. Syst.1
2022 Multi-Task Y-Shaped Graph Neural Network for Point Cloud Learning in Autonomous Driving
abstract
Point cloud, an efficient 3D object representation, plays an indispensable role in autonomous driving technologies, such as object avoidance, localization, and map building. The analysis of point clouds (e.g., 3D segmentation) is essential to exploit the informative value of point clouds for such applications. The main challenge remains to effectively and completely extract high-level point cloud feature representations. To this end, we present a novel multi-task Y-shaped graph neural network to explore 3D point clouds, referred to as MTYGNN. By extending the conventional U-Net, MTYGNN contains two main branches to simultaneously perform classification and segmentation tasks in point clouds. Meanwhile, the classification prediction is fused together with the semantic features as the scene context to make the segmentation task more accurate. Furthermore, we consider the homoscedastic uncertainty of each task to calculate the weights of multiple loss functions to ensure that tasks do not negatively interfere with each other. The proposed MTYGNN is evaluated on popular point cloud datasets in traffic scenarios. Experimental results demonstrate that our framework outperforms the state-of-the-art baseline methods.
Xiaofeng Zou, Kenli Li 0001, Yangfan Li 0001, Wei Wei 0006, Cen Chen 0002
IEEE Trans. Intell. Transp. Syst.5
2022 Modeling Temporal Patterns with Dilated Convolutions for Time-Series Forecasting
abstract
Time-series forecasting is an important problem across a wide range of domains. Designing accurate and prompt forecasting algorithms is a non-trivial task, as temporal data that arise in real applications often involve both non-linear dynamics and linear dependencies, and always have some mixtures of sequential and periodic patterns, such as daily, weekly repetitions, and so on. At this point, however, most recent deep models often use Recurrent Neural Networks (RNNs) to capture these temporal patterns, which is hard to parallelize and not fast enough for real-world applications especially when a huge amount of user requests are coming. Recently, CNNs have demonstrated significant advantages for sequence modeling tasks over the de-facto RNNs, while providing high computational efficiency due to the inherent parallelism. In this work, we propose HyDCNN, a novel hybrid framework based on fully Dilated CNN for time-series forecasting tasks. The core component in HyDCNN is a proposed hybrid module, in which our proposed position-aware dilated CNNs are utilized to capture the sequential non-linear dynamics and an autoregressive model is leveraged to capture the sequential linear dependencies. To further capture the periodic temporal patterns, a novel hop scheme is introduced in the hybrid module. HyDCNN is then composed of multiple hybrid modules to capture the sequential and periodic patterns. Each of these hybrid modules targets on either the sequential pattern or one kind of periodic patterns. Extensive experiments on five real-world datasets have shown that the proposed HyDCNN is better compared with state-of-the-art baselines and is at least 200% better than RNN baselines. The datasets and source code will be published in Github to facilitate more future work.
Yangfan Li 0001, Kenli Li 0001, Cen Chen 0002, Xu Zhou 0001, Zeng Zeng, Keqin Li 0001
ACM Trans. Knowl. Discov. Data3
2022 Hierarchical Semantic Graph Reasoning for Train Component Detection
abstract
Recently, deep learning-based approaches have achieved superior performance on object detection applications. However, object detection for industrial scenarios, where the objects may also have some structures and the structured patterns are normally presented in a hierarchical way, is not well investigated yet. In this work, we propose a novel deep learning-based method, hierarchical graphical reasoning (HGR), which utilizes the hierarchical structures of trains for train component detection. HGR contains multiple graphical reasoning branches, each of which is utilized to conduct graphical reasoning for one cluster of train components based on their sizes. In each branch, the visual appearances and structures of train components are considered jointly with our proposed novel densely connected dual-gated recurrent units (Dense-DGRUs). To the best of our knowledge, HGR is the first kind of framework that explores hierarchical structures among objects for object detection. We have collected a data set of 1130 images captured from moving trains, in which 17 334 train components are manually annotated with bounding boxes. Based on this data set, we carry out extensive experiments that have demonstrated our proposed HGR outperforms the existing state-of-the-art baselines significantly. The data set and the source code can be downloaded online at https://github.com/ChengZY/HGR.
Cen Chen 0002, Kenli Li 0001, Xiaofeng Zou, Zhongyao Cheng, Wei Wei 0006, Qi Tian 0001, Zeng Zeng
IEEE Trans. Neural Networks Learn. Syst.1
2022 Latency-Driven Model Placement for Efficient Edge Intelligence Service
abstract
Deep learning services based on cloud computing have deficiencies in latency, privacy, etc. To meet the requirements of low latency, researchers have begun to consider the deployment of deep learning services in edges, i.e., edge intelligence service. Deploying deep learning models on multiple processors or devices so that the computation of a deep learning model can be conducted in parallel is a possible solution to improve the efficiency of edge intelligence services. In this article, we propose a novel latency-driven deep learning model placement method for efficient edge intelligence service. Model placement contains two procedures: model partition and sub-models assignment. In our method, we first convert the model into execution graphs and propose a novel latency-driven multilevel graph partition for the model. Then the partitioned sub-models are heuristically assigned to available processors. To the best of our knowledge, it is the first work that proposes latency-driven graph partition algorithms for model placement. Extensive experiments on several commonly used DNN (deep neural network) models and synthetic datasets show that our method can achieve the lowest execution latency with low complexity compared with other state-of-the-art model placement methods.
Peiying Lin, Zhichen Shi, Cen Chen 0002, Kenli Li 0001
IEEE Trans. Serv. Comput.4
2022 Determinantal point process-based new radio unlicensed link scheduling for multi-access edge computing
Chigang Xing, Yangfan Li 0001, Cen Chen 0002, Fangmin Li, Zeng Zeng, Xiaofeng Zou
World Wide Web3
2021 DyGNN: Algorithm and Architecture Support of Dynamic Pruning for Graph Neural Networks
abstract
Recently, graph neural networks (GNNs) have achieved great success for graph representation learning tasks. Enlightened by the fact that numerous message passing redundancies exist in GNNs, we propose DyGNN, which speeds up GNNs by reducing redundancies. DyGNN is supported by an algorithm and architecture co-design. The proposed algorithm can dynamically prune vertices and edges during execution without accuracy loss. An architecture is designed to support dynamic pruning and transform it into performance improvement. DyGNN opens new directions for accelerating GNNs by pruning vertices and edges. DyGNN gains average $2\times$ speedup with accuracy improvement of 4% compared with state-of-the-art GNN accelerators.
Cen Chen 0002, Kenli Li 0001, Xiaofeng Zou, Yangfan Li 0001
DAC1
2021 Work in Progress: Topology-based Multilevel Algorithm for Large-scale Task Scheduling in Clouds
abstract
Task scheduling in cloud environments is the problem of assigning and executing computational tasks on the available cloud resources. Effective task scheduling can improve processor utilization, reduce processor energy consumption, and improve user experience. Large-scale task scheduling under multiple constraints is an NP-complete problem. The traditional task scheduling algorithm cannot be applied to large-scale scheduling, either because of high time complexity or because its heuristic algorithm cannot be applied to complexly large-scale scenarios. The problem of large-scale task scheduling is gradually becoming a challenge in cloud computing. The article proposes a topology-based multilevel algorithm for large-scale task scheduling in clouds. Based on the topological order of the graph, multi-level coarsening is performed on the large-scale graphs, and then uses the traditional scheduling algorithm for the initial scheduling of the coarse graph, and then refine the initial scheduling result during its uncoarsen phrase. It can perform fast and efficient scheduling of large-scale task graphs. At the same time, it has good compatibility, which can be combined with excellent traditional scheduling algorithms.
Minjia Li, Yikun Hu 0001, Cen Chen 0002, Chubo Liu, Kenli Li 0001
RTAS3
2021 Work in Progress: Path-based Graph Partition for Parallel Hardware-accelerated Functional Verification
abstract
Functional verification of large scale circuit design is a basic problem in Very Large Scale Integrated (VLSI) design. With the increasing scale of the circuit, it is urgent to divide the whole large scale circuit into some smaller sub-circuits so as to perform parallel functional verification on multiple hardware processors. The partition problem of hardware-accelerated functional verification can be regarded as a graph partition problem. However, unlike the traditional graph partition requirements for minimum cutting, the hardware-accelerated functional verification partition needs to reduce the simulation depth and improve the parallelism of the simulation. Therefore, partition for hardware-accelerated functional verification is a problem combined with graph partitioning and schedule. While the traditional schedule algorithms have high complexity and cannot handle large scale Directied Acyclic Graph (DAG) scheduling. To tackle the parallelism, depth, and cut edge problem, we design a new method, called path-metis. Path-metis combines the scheduling idea, such as the critical path information and task priority of the DAG, into the traditional multilevel partitioning method. Our preliminary experiments on real circuits show the effectiveness of the method, and the simulation depth can be reduced by about 11.35% on average compared with metis only with 27.58% cut size increasing.
Peiying Lin, Kenli Li 0001, Cen Chen 0002, Siyang Yu
RTAS4
2021 Introduction to the Special issue on Advances of neurocomputing for smart cities (NEUROCOM for smart cities)
Kenli Li 0001, Keqin Li 0001, Cen Chen 0002, Xiaokang Wang 0001
Neurocomputing3
2021 Multiple local 3D CNNs for region-based prediction in smart cities
Yibi Chen, Xiaofeng Zou, Kenli Li 0001, Keqin Li 0001, Xulei Yang, Cen Chen 0002
Inf. Sci.6
2021 Attention-Aware Encoder-Decoder Neural Networks for Heterogeneous Graphs of Things
abstract
Recent trend focuses on using heterogeneous graph of things (HGoT) to represent things and their relations in the Internet of Things, thereby facilitating the applying of advanced learning frameworks, i.e., deep learning (DL). Nevertheless, this is a challenging task since the existing DL models are hard to accurately express the complex semantics and attributes for those heterogeneous nodes and links in HGoT. To address this issue, we develop attention-aware encoder-decoder graph neural networks for HGoT, termed as HGAED. Specifically, we utilize the attention-based separate-and-merge method to improve the accuracy, and leverage the encoder-decoder architecture for implementation. In the heart of HGAED, the separate-and-merge processes can be encapsulated into encoding and decoding blocks. Then, blocks are stacked for constructing an encoder-decoder architecture to jointly and hierarchically fuse heterogeneous structures and contents of nodes. Extensive experiments on three real-world datasets demonstrate the superior performance of HGAED over state-of-the-art baselines.
Yangfan Li 0001, Cen Chen 0002, Mingxing Duan, Zeng Zeng, Kenli Li 0001
IEEE Trans. Ind. Informatics2
2020 A two-stage attention aware method for train bearing shed oil inspection based on convolutional neural networks
Kenli Li 0001, Jing Liu 0032, Keqin Li 0001, Zeng Zeng, Cen Chen 0002
Neurocomputing6
2020 WiFi-Based Indoor Robot Positioning Using Deep Fuzzy Forests
abstract
Addressing the positioning problem of a mobile robot remains challenging to date despite many years of research. Indoor robot positioning strategies developed in the literature either rely on sophisticated computer vision techniques to handle visual inputs or require strong domain knowledge for nonvisual sensors. Although some systems have been deployed, the former may be lacking due to the intrinsic limitation of cameras (such as calibration, data association, system initialization, etc.) and the latter usually only works under certain environment layouts and additional equipment. To cope with those issues, we design a lightweight indoor robot positioning system which operates on cost-effective WiFi-based received signal strength (RSS) and could be readily pluggable into any existing WiFi network infrastructures. Moreover, a novel deep fuzzy forest is proposed to inherit the merits of decision trees and deep neural networks within an end-to-end trainable architecture. Real-world indoor localization experiments are conducted and results demonstrate the superiority of the proposed method over the existing approaches.
Le Zhang 0001, Zhenghua Chen, Wei Cui 0002, Bing Li 0002, Cen Chen 0002, Zhiguang Cao, Kai-Zhou Gao
IEEE Internet Things J.5
2020 A hierarchical deep convolutional neural network and gated recurrent unit framework for structural damage detection
Jianxi Yang, Cen Chen 0002, Yangfan Li 0001, Guiping Wang, Shixin Jiang, Zeng Zeng
Inf. Sci.3
2020 Multi-task cascade deep convolutional neural networks for large-scale commodity recognition
Xiaofeng Zou, Liqian Zhou, Kenli Li 0001, Aijia Ouyang, Cen Chen 0002
Neural Comput. Appl.5
2020 Citywide Traffic Flow Prediction Based on Multiple Gated Spatio-temporal Convolutional Neural Networks
abstract
Traffic flow prediction is crucial for public safety and traffic management, and remains a big challenge because of many complicated factors, e.g., multiple spatio-temporal dependencies, holidays, and weather. Some work leveraged 2D convolutional neural networks (CNNs) and long short-term memory networks (LSTMs) to explore spatial relations and temporal relations, respectively, which outperformed the classical approaches. However, it is hard for these work to model spatio-temporal relations jointly. To tackle this, some studies utilized LSTMs to connect high-level layers of CNNs, but left the spatio-temporal correlations not fully exploited in low-level layers. In this work, we propose novel spatio-temporal CNNs to extract spatio-temporal features simultaneously from low-level to high-level layers, and propose a novel gated scheme to control the spatio-temporal features that should be propagated through the hierarchy of layers. Based on these, we propose an end-to-end framework, multiple gated spatio-temporal CNNs (MGSTC), for citywide traffic flow prediction. MGSTC can explore multiple spatio-temporal dependencies through multiple gated spatio-temporal CNN branches, and combine the spatio-temporal features with external factors dynamically. Extensive experiments on two real traffic datasets demonstrates that MGSTC outperforms other state-of-the-art baselines.
Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Xiaofeng Zou, Keqin Li 0001, Zeng Zeng
ACM Trans. Knowl. Discov. Data1
2019 Gated Residual Recurrent Graph Neural Networks for Traffic Prediction
abstract
Traffic prediction is of great importance to traffic management and public safety, and very challenging as it is affected by many complex factors, such as spatial dependency of complicated road networks and temporal dynamics, and many more. The factors make traffic prediction a challenging task due to the uncertainty and complexity of traffic states. In the literature, many research works have applied deep learning methods on traffic prediction problems combining convolutional neural networks (CNNs) with recurrent neural networks (RNNs), which CNNs are utilized for spatial dependency and RNNs for temporal dynamics. However, such combinations cannot capture the connectivity and globality of traffic networks. In this paper, we first propose to adopt residual recurrent graph neural networks (Res-RGNN) that can capture graph-based spatial dependencies and temporal dynamics jointly. Due to gradient vanishing, RNNs are hard to capture periodic temporal correlations. Hence, we further propose a novel hop scheme into Res-RGNN to utilize the periodic temporal dependencies. Based on Res-RGNN and hop Res-RGNN, we finally propose a novel end-to-end multiple Res-RGNNs framework, referred to as “MRes-RGNN”, for traffic prediction. Experimental results on two traffic datasets have demonstrated that the proposed MRes-RGNN outperforms state-of-the-art methods significantly.
Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Xiaofeng Zou, Jie Wang 0042, Zeng Zeng
AAAI1
2019 Multiple convolutional neural networks for multivariate time series prediction
Kenli Li 0001, Liqian Zhou, Yikun Hu 0001, Zhongyao Cheng, Jing Liu 0032, Cen Chen 0002
Neurocomputing7
2018 Exploiting Spatio-Temporal Correlations with Multiple 3D Convolutional Neural Networks for Citywide Vehicle Flow Prediction
abstract
Predicting vehicle flows is of great importance to traffic management and public safety in smart cities, and very challenging as it is affected by many complex factors, such as spatio-temporal dependencies with external factors (e.g., holidays, events and weather). Recently, deep learning has shown remarkable performance on traditional challenging tasks, such as image classification, due to its powerful feature learning capabilities. Some works have utilized LSTMs to connect the high-level layers of 2D convolutional neural networks (CNNs) to learn the spatio-temporal features, and have shown better performance as compared to many classical methods in traffic prediction. However, these works only build temporal connections on the high-level features at the top layer while leaving the spatio-temporal correlations in the low-level layers not fully exploited. In this paper, we propose to apply 3D CNNs to learn the spatio-temporal correlation features jointly from low-level to high-level layers for traffic data. We also design an end-to-end structure, named as MST3D, especially for vehicle flow prediction. MST3D can learn spatial and multiple temporal dependencies jointly by multiple 3D CNNs, combine the learned features with external factors and assign different weights to different branches dynamically. To the best of our knowledge, it is the first framework that utilizes 3D CNNs for traffic prediction. Experiments on two vehicle flow datasets Beijing and New York City have demonstrated that the proposed framework, MST3D, outperforms the state-of-the-art methods.
Cen Chen 0002, Kenli Li 0001, Sin G. Teo, Guizi Chen, Xiaofeng Zou, Xulei Yang, Ramaseshan C. Vijay, Jiashi Feng, Zeng Zeng
ICDM1
2018 FlinkCL: An OpenCL-Based In-Memory Computing Architecture on Heterogeneous CPU-GPU Clusters for Big Data
abstract
Research on in-memory big data management and processing has been prompted by the increase in main memory capacity and the explosion in big data. By offering an efficient in-memory distributed execution model, existing in-memory cluster computing platforms such as Flink and Spark have been proven to be outstanding for processing big data. This paper proposes FlinkCL, an in-memory computing architecture on heterogeneous CPU-GPU clusters based on OpenCL that enables Flink to utilize GPU's massive parallel processing ability. Our proposed architecture utilizes four techniques: a heterogeneous distributed abstract model (HDST), a Just-In-Time (JIT) compiling schema, a hierarchical partial reduction (HPR) and a heterogeneous task management strategy. Using FlinkCL, programmers only need to write Java code with simple interfaces. The Java code can be compiled to OpenCL kernels and executed on CPUs and GPUs automatically. In the HDST, a novel memory mapping scheme is proposed to avoid serialization or deserialization between Java Virtual Machine (JVM) objects and OpenCL structs. We have comprehensively evaluated FlinkCL with a set of representative workloads to show its effectiveness. Our results show that FlinkCL improve the performance by up to$11 \times$for some computationally heavy algorithms and maintains minor performance improvements for a I/O bound algorithm.
Cen Chen 0002, Kenli Li 0001, Aijia Ouyang, Keqin Li 0001
IEEE Trans. Computers1
2018 GFlink: An In-Memory Computing Architecture on Heterogeneous CPU-GPU Clusters for Big Data
abstract
The increasing main memory capacity and the explosion of big data have fueled the development of in-memory big data management and processing. By offering an efficient in-memory parallel execution model which can eliminate disk I/O bottleneck, existing in-memory cluster computing platforms (e.g., Flink and Spark) have already been proven to be outstanding platforms for big data processing. However, these platforms are merely CPU-based systems. This paper proposes GFlink, an in-memory computing architecture on heterogeneous CPU-GPU clusters for big data. Our proposed architecture extends the original Flink from CPU clusters to heterogeneous CPU-GPU clusters, greatly improving the computational power of Flink. Furthermore, we have proposed a programming framework based on Flink's abstract model, i.e., DataSet (DST), hiding the programming complexity of GPUs behind the simple and familiar high-level interfaces. To achieve high performance and good load-balance, an efficient JVM-GPU communication strategy, a GPU cache scheme, and an adaptive locality-aware scheduling scheme for three-stage pipelining execution are proposed. Extensive experiment results indicate that the high computational power of GPUs can be efficiently utilized, and the implementation on GFlink outperforms that on the original CPU-based Flink.
Cen Chen 0002, Kenli Li 0001, Aijia Ouyang, Zeng Zeng, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.1
2017 DHCRF: A Distributed Conditional Random Field Algorithm on a Heterogeneous CPU-GPU Cluster for Big Data
abstract
As one of the most recognized models in machine learning, the conditional random fields (CRF) has been widely used in many applications. As the parameter estimation of CRF is highly time-consuming, how to improve the performance of CRF has received significant attention, in particular in the big data environment. To deal with large-scale data, CPU-based or GPU-based parallelization solutions have been proposed to improve performance. However, the problem is an ongoing one. In this paper, we focus on the big data environment and propose a distributed CRF on a heterogeneous CPU-GPU cluster called DHCRF. Our approach differs from previous work. Specifically, it leverages a three-stage heterogeneous Map and Reduce operation to improve the performance, making full use of CPU-GPU collaborative computing capabilities in a big data environment. Furthermore, by combining elastic data partition and intermediate results multiplexing method, the distributed CRF is optimized. Elastic data partition is performed to keep the load balanced, and the intermediate results multiplexing method is adopted to reduce data communication. Experimental results show that the DHCRF outperforms the baseline CRF algorithm and the CPU-based parallel CRF algorithm with notable performance improvement while maintaining competitive correctness at the same time.
Wei Ai 0001, Kenli Li 0001, Cen Chen 0002, Jiwu Peng, Keqin Li 0001
ICDCS3
2017 A parallel approximate SS-ELM algorithm based on MapReduce for large-scale datasets
Cen Chen 0002, Kenli Li 0001, Aijia Ouyang, Keqin Li 0001
J. Parallel Distributed Comput.1
2017 GPU-Accelerated Parallel Hierarchical Extreme Learning Machine on Flink for Big Data
abstract
The extreme learning machine (ELM) has become one of the most important and popular algorithms of machine learning, because of its extremely fast training speed, good generalization, and universal approximation/classification capability. The proposal of hierarchical ELM (H-ELM) extends ELM from single hidden layer feedforward networks to multilayer perceptron, greatly strengthening the applicability of ELM. Generally speaking, during training H-ELM, large-scale datasets (DSTs) are needed. Therefore, how to make use of H-ELM framework in processing big data is worth further exploration. This paper proposes a parallel H-ELM algorithm based on Flink, which is one of the in-memory cluster computing platforms, and graphics processing units (GPUs). Several optimizations are adopted to improve the performance, such as cache-based scheme, reasonable partitioning strategy, memory mapping scheme for mapping specific Java virtual machine objects to buffers. Most importantly, our proposed framework for utilizing GPUs to accelerate Flink for big data is general. This framework can be utilized to accelerate many other variants of ELM and other machine learning algorithms. To the best of our knowledge, it is the first kind of library, which combines in-memory cluster computing with GPUs to parallelize H-ELM. The experimental results have demonstrated that our proposed GPU-accelerated parallel H-ELM named as GPH-ELM can efficiently process large-scale DSTs with good performance of speedup and scalability, leveraging the computing power of both CPUs and GPUs in the cluster.
Cen Chen 0002, Kenli Li 0001, Aijia Ouyang, Zhuo Tang, Keqin Li 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2016 GFlink: An In-Memory Computing Architecture on Heterogeneous CPU-GPU Clusters for Big Data
abstract
The increasing main memory capacity and the explosion of big data has fueled the development of in-memory big data management and processing. By offering efficient in-memory parallel execution model which eliminates disk I/O bottleneck, existing in-memory cluster computing platforms (e.g., Flink and Spark) have already been proven to be outstanding platforms for big data processing. However, these platforms are merely CPU-based systems now. This paper has proposed GFlink, an in-memory computing architecture on heterogeneous CPU-GPU clusters for big data. Our proposed architecture extends the original Flink from CPU clusters to heterogeneous CPU-GPU clusters, greatly improving the computational power of Flink. Furthermore, we proposed an abstract GPU-based model named as GDST, hiding the programming complexity of GPUs behind the simple and familiar high-level interfaces, and automatically managing task partitioning, device memory, and parallelization on MultiCore GPUs. To achieve high performance and good load-balance, an efficient JVM-GPU communication strategy and an adaptive locality-aware scheduling scheme for three-stage pipeline execution are proposed. Extensive experiment results indicate that not only the high computational power of GPUs can be efficiently utilized, but also the implementations on GFlink outperforms that on the original CPU-based Flink.
Cen Chen 0002, Kenli Li 0001, Aijia Ouyang, Zhuo Tang, Keqin Li 0001
ICPP1