Yangdong Deng

dblp:90/5987 · also Yangdong Steve Deng · DBLP profile ↗
← Back
60ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-8257-693XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 39 · 5 first-author · 4 since 2021Artificial intelligence and machine learning · 13 · 9 since 2021Software engineering, systems software and programming languages · 7Graphics, computer vision, multimedia, augmented reality and games · 5Computer networks · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language Models
abstract
Yan Liu, Feng Zhang, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Han Liu, Yangdong Deng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhanyu Ma, Jun Xu 0001, Jiuchong Gao, Jinghua Hao, Renqing He, Yangdong Deng
ACL (1)9
2024 CQIL: Inference Latency Optimization with Concurrent Computation of Quasi-Independent Layers
abstract
The fast-growing large scale language models are delivering unprecedented performance on almost all natural language processing tasks.However, the effectiveness of large language models are reliant on an exponentially increasing number of parameters.The overwhelming computation complexity incurs a high inference latency that negatively affects user experience.Existing methods to improve inference efficiency, such as tensor parallelism and quantization, target to reduce per-layer computing latency, yet overlook the cumulative latency due to the number of layers.Recent works on reducing the cumulative latency through layer removing, however, lead to significant performance drop.Motivated by the similarity of inputs among adjacent layers, we propose to identify quasi-independent layers, which can be concurrently computed to significantly decrease inference latency.We also introduce a bypassing technique to mitigate the effect of information loss.Empirical experiments of the proposed approach on the LLaMA models confirm that Concurrent Computation of Quasi-Independent Layers (CQIL) can reduce latency by up to 48.3% on LLaMA-33B, while maintaining a close level of performance.
Longwei Zou, Jiangangkong Jiangangkong, Yangdong Deng
ACL (1)6
2024 A Multi-Level Framework for Accelerating Training Transformer Models
abstract
The fast growing capabilities of large-scale deep learning models, such as Bert, GPT and ViT, are revolutionizing the landscape of NLP, CV and many other domains. Training such models, however, poses an unprecedented demand for computing power, which incurs exponentially increasing energy cost and carbon dioxide emissions. It is thus critical to develop efficient training solutions to reduce the training costs. Motivated by a set of key observations of inter- and intra-layer similarities among feature maps and attentions that can be identified from typical training processes, we propose a multi-level framework for training acceleration. Specifically, the framework is based on three basic operators, Coalescing, De-coalescing and Interpolation, which can be orchestrated to build a multi-level training framework. The framework consists of a V-cycle training process, which progressively down- and up-scales the model size and projects the parameters between adjacent levels of models via coalescing and de-coalescing. The key idea is that a smaller model that can be trained for fast convergence and the trained parameters provides high-qualities intermediate solutions for the next level larger network. The interpolation operator is designed to break the symmetry of neurons incurred by de-coalescing for better convergence performance. Our experiments on transformer-based language models (e.g. Bert, GPT) as well as a vision model (e.g. DeiT) prove that the proposed framework reduces the computational cost by about 20% on training BERT/GPT-Base models and up to 51.6% on training the BERT-Large model while preserving the performance.
Longwei Zou, Yangdong Deng
ICLR3
2023 Beat LLMs at Their Own Game: Zero-Shot LLM-Generated Text Detection via Querying ChatGPT
abstract
Biru Zhu, Lifan Yuan, Ganqu Cui, Yangyi Chen, Chong Fu, Bingxiang He, Yangdong Deng, Zhiyuan Liu, Maosong Sun, Ming Gu. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Biru Zhu, Lifan Yuan, Ganqu Cui, Yangyi Chen, Bingxiang He, Yangdong Deng, Zhiyuan Liu 0001, Maosong Sun 0001, Ming Gu 0001
EMNLP7
2023 A spatiotemporal deep neural network for fine-grained multi-horizon wind prediction
Fanling Huang, Yangdong Deng
Data Min. Knowl. Discov.2
2023 TCGAN: Convolutional Generative Adversarial Network for time series classification and clustering
abstract
Recent works have demonstrated the superiority of supervised Convolutional Neural Networks (CNNs) in learning hierarchical representations from time series data for successful classification. These methods require sufficiently large labeled data for stable learning, however acquiring high-quality labeled time series data can be costly and potentially infeasible. Generative Adversarial Networks (GANs) have achieved great success in enhancing unsupervised and semi-supervised learning. Nonetheless, to our best knowledge, it remains unclear how effectively GANs can serve as a general-purpose solution to learn representations for time series recognition, i.e., classification and clustering. The above considerations inspire us to introduce a Time-series Convolutional GAN (TCGAN). TCGAN learns by playing an adversarial game between two one-dimensional CNNs (i.e., a generator and a discriminator) in the absence of label information. Parts of the trained TCGAN are then reused to construct a representation encoder to empower linear recognition methods. We conducted comprehensive experiments on synthetic and real-world datasets. The results demonstrate that TCGAN is faster and more accurate than existing time-series GANs. The learned representations enable simple classification and clustering methods to achieve superior and stable performance. Furthermore, TCGAN retains high efficacy in scenarios with few-labeled and imbalanced-labeled data. Our work provides a promising path to effectively utilize abundant unlabeled time series data.
Fanling Huang, Yangdong Deng
Neural Networks2
2023 Removing Backdoors in Pre-trained Models by Regularized Continual Pre-training
abstract
Abstract Recent research has revealed that pre-trained models (PTMs) are vulnerable to backdoor attacks before the fine-tuning stage. The attackers can implant transferable task-agnostic backdoors in PTMs, and control model outputs on any downstream task, which poses severe security threats to all downstream applications. Existing backdoor-removal defenses focus on task-specific classification models and they are not suitable for defending PTMs against task-agnostic backdoor attacks. To this end, we propose the first task-agnostic backdoor removal method for PTMs. Based on the selective activation phenomenon in backdoored PTMs, we design a simple and effective backdoor eraser, which continually pre-trains the backdoored PTMs with a regularization term in an end-to-end approach. The regularization term removes backdoor functionalities from PTMs while the continual pre-training maintains the normal functionalities of PTMs. We conduct extensive experiments on pre-trained models across different modalities and architectures. The experimental results show that our method can effectively remove backdoors inside PTMs and preserve benign functionalities of PTMs with a few downstream-task-irrelevant auxiliary data, e.g., unlabeled plain texts. The average attack success rate on three downstream datasets is reduced from 99.88% to 8.10% after our defense on the backdoored BERT. The codes are publicly available at https://github.com/thunlp/RECIPE.
Biru Zhu, Ganqu Cui, Yangyi Chen, Yujia Qin, Lifan Yuan, Yangdong Deng, Zhiyuan Liu 0001, Maosong Sun 0001, Ming Gu 0001
Trans. Assoc. Comput. Linguistics7
2022 Pass off Fish Eyes for Pearls: Attacking Model Selection of Pre-trained Models
abstract
Biru Zhu, Yujia Qin, Fanchao Qi, Yangdong Deng, Zhiyuan Liu, Maosong Sun, Ming Gu. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Biru Zhu, Yujia Qin, Fanchao Qi, Yangdong Deng, Zhiyuan Liu 0001, Maosong Sun 0001, Ming Gu 0001
ACL (1)4
2022 Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language Models
abstract
Despite the great success of pre-trained language models (PLMs) in a large set of natural language processing (NLP) tasks, there has been a growing concern about their security in real-world applications. Backdoor attack, which poisons a small number of training samples by inserting backdoor triggers, is a typical threat to security. Trained on the poisoned dataset, a victim model would perform normally on benign samples but predict the attacker-chosen label on samples containing pre-defined triggers. The vulnerability of PLMs under backdoor attacks has been proved with increasing evidence in the literature. In this paper, we present several simple yet effective training strategies that could effectively defend against such attacks. To the best of our knowledge, this is the first work to explore the possibility of backdoor-free adaptation for PLMs. Our motivation is based on the observation that, when trained on the poisoned dataset, the PLM's adaptation follows a strict order of two stages: (1) a moderate-fitting stage, where the model mainly learns the major features corresponding to the original task instead of subsidiary features of backdoor triggers, and (2) an overfitting stage, where both features are learned adequately. Therefore, if we could properly restrict the PLM's adaptation to the moderate-fitting stage, the model would neglect the backdoor triggers but still achieve satisfying performance on the original task. To this end, we design three methods to defend against backdoor attacks by reducing the model capacity, training epochs, and learning rate, respectively. Experimental results demonstrate the effectiveness of our methods in defending against several representative NLP backdoor attacks. We also perform visualization-based analysis to attain a deeper understanding of how the model learns different features, and explore the effect of the poisoning ratio. Finally, we explore whether our methods could defend against backdoor attacks for the pre-trained CV model. The codes are publicly available at https://github.com/thunlp/Moderate-fitting.
Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Yangdong Deng, Zhiyuan Liu 0001, Jingang Wang, Wei Wu 0014, Maosong Sun 0001, Ming Gu 0001
NeurIPS7
2022 Agglomerative Memory and Thread Scheduling for High-Performance Ray-Tracing on GPUs
abstract
Ray-tracing rendering has long been considered as a promising technology to enable a higher level of visual experience. The democratization of the ray-tracing rendering to consumer platforms, however, poses significant challenges to rendering hardware and software due to its highly irregular computing patterns. In fact, modern ray-tracing techniques typically depend on a tree-based acceleration structure to reduce the computing complexity of intersection testing of rays and graphics primitives. The traversal by a massive number of rays on a graphics processing unit (GPU) incurs a significant amount of irregular memory traffic, which turns out to be a major stumbling block for real-time performance. In this work, a scheduling mechanism, so-called Agglomerative Memory and Thread Scheduling, is proposed to unleash the inherence parallelism in the ray-tracing process on GPUs. It is associated with a tile-based ray-tracing framework in which the acceleration structure (i.e., KD-tree in this work) is partitioned into subtrees that can be completely loaded into the on-chip L1 cache inside a streaming multiprocessor. An effective scheduling mechanism collects threads with regard to the subtrees hit by their respective rays and regroup threads into warps for dispatching. In addition, subtrees are dynamically preloaded into the L1 cache of multiprocessors in an on-demand fashion. The proposed scheduler can be integrated on today’s high-end GPUs with only minor overhead. Microarchitecture simulation results prove that the proposed framework significantly improves memory efficiency and outperforms a traditional GPU microarchitecture by 47.4% for average.
Yufei Ni, Yangdong Deng, Zonghui Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 Duplicacy: A New Generation of Cloud Backup Tool Based on Lock-Free Deduplication
abstract
The pervasive deployment of cloud services poses an ever-increasing demand for cross-client deduplication solutions to save network bandwidth, lower storage costs, and improve backup speeds. However, existing solutions typically depend on lock-based approaches relying on a centralized chunk database, which tends to hinder performance scalability. In this article, we present a new cross-client cloud backup solution, named Duplicacy, based on a Lock-Free Deduplication approach. Lock-Free Deduplication stores chunks to network or cloud storage using content hashes as file names. It then adopts a two-step fossil deletion algorithm to solve the hard problem of deleting unreferenced chunks in the presence of concurrent backups, without the need for any locks. Experiments demonstrate that Duplicacy enables significant performance improvement for backups over previous well-known backup tools. In addition, Duplicacy can work with many general-purpose network or cloud storage services which only support a basic set of file operations, and turn them into sophisticated deduplication-aware storage servers without server-side changes.
Zonghui Li, Gilbert Chen, Yangdong Deng
IEEE Trans. Cloud Comput.3
2021 Knowledge Enhanced Fact Checking and Verification
abstract
As the Internet and social media offer increasing opportunities for organizations and individuals to publicize online contents, it has become essential to develop effective means to identify misinformation like fake news. Recently, fact checking systems have been regarded as a promising tool to automatically deal with large amounts of information. How to effectively take advantage of existing unstructured document knowledge bases and structured knowledge graphs to build robust fact checking systems, however, remains to be a challenge. In this paper, we propose a knowledge enhanced fact checking system, which leverages the Wikidata5M knowledge graph and Wikipedia documents to incorporate external knowledge into the claim to be checked for more robust and accurate fact checking. First, we devise a contextualized knowledge graph selection method to identify the most relevant sub-graph with the checked claim from the large knowledge graph. We then construct a novel claim-evidence-knowledge graph and use a graph attention network to integrate natural language evidence with structured knowledge triplets by allowing them to propagate information among each other. By integrating the claim, retrieved evidence and selected knowledge triplets in a unified claim-evidence-knowledge graph, our method improves the label accuracy of predicted claims by more than 4% on the FEVER dataset over state-of-the-art fact checking models.
Biru Zhu, Xingyao Zhang 0003, Ming Gu 0001, Yangdong Deng
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 An Elastic Task Scheduling Scheme on Coarse-Grained Reconfigurable Architectures
abstract
Coarse-grained reconfigurable architectures (CGRAs) are increasingly employed as domain-specific accelerators due to their efficiency and flexibility. A CGRA typically relies on compilers to perform task scheduling. The longstanding problem of static scheduling is that it suffers from insufficient parallelism in handling irregularities due to over-serialization and workload imbalance, which leads to severe resource underutilization and performance loss. To counteract the limitations of static scheduling in CGRAs, it is essential to exploit dynamic parallelism automatically and manage hardware resources adaptively. However, existing dynamic scheduling mechanisms, e.g., work stealing, often reschedule aggressively for instant performance but sacrifice efficiency, which is unfavorable to CGRAs that emphasize efficiency and fewer reconfigurations. This article proposes an elastic task scheduling scheme that enables lightweight dynamic scheduling in CGRAs. Tasks are rescheduled at runtime according to the classic tagged-token dataflow paradigm to enable dynamic task-level parallelism. Meanwhile, tasks are dynamically resized according to run-time throughputs via duplication, combination, and substitution operators for balanced multitask execution. We implement the elastic task scheduling scheme on a well-known reconfigurable architecture - triggered instruction architecture (TIA). Evaluation on the MachSuite benchmarks shows that the proposed scheme is effective in improving performance and energy efficiency. The average speedup is 2× over the baseline. Also, our design attains a 57 percent improvement in the area-normalized performance and a 49 percent better energy efficiency. Compared with a state-of-the-art dynamic scheduling method, our scheme achieves 1.6× speedup and 1.6× energy efficiency than work-stealing mechanism on the same substrate.
Longlong Chen, Jianfeng Zhu 0001, Yangdong Deng, Zhaoshi Li, Xiaowei Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu
IEEE Trans. Parallel Distributed Syst.3
2021 A GPU Acceleration Framework for Motif and Discord Based Pattern Mining
abstract
With the fast digitalization of our society, mining patterns from large time series data is increasingly becoming a critical problem for a wide range of big data applications. Motif and discord discovery algorithms, which offer effective solutions to identify repeatedly appearing and abnormal patterns, respectively, are fundamental building blocks for time series processing. Both approaches, however, can be time extremely consuming when handling large time series due to the subsequence-based computations of distance similarity metrics. In this article, we show that the highly involved subsequence-based computations can actually be decomposed into a few fine-grained computing patterns for efficient data parallel computing. By developing highly efficient GPU algorithms for such basic patterns and effectively composing such patterns, we are able to solve both motif and discord discovery problems under euclidean and DTW distance metrics in a unified GPU acceleration framework. Extensive experiments prove that the proposed framework outperforms pruned CPU algorithms by up to three orders of magnitude. Our work paves the foundation of building GPU acceleration frameworks for large-scale time series datasets.
Biru Zhu, Youyou Jiang, Ming Gu 0001, Yangdong Deng
IEEE Trans. Parallel Distributed Syst.4
2020 GraphABCD: Scaling Out Graph Analytics with Asynchronous Block Coordinate Descent
abstract
It is of vital importance to efficiently process large graphs for many data-intensive applications. As a result, a large collection of graph analytic frameworks has been proposed to improve the per-iteration performance on a single kind of computation resource. However, heavy coordination and synchronization overhead make it hard to scale out graph analytic frameworks from single platform to heterogeneous platforms. Furthermore, increasing the convergence rate, i.e. reducing the number of iterations, which is equally vital for improving the overall performance of iterative graph algorithms, receives much less attention. In this paper, we introduce the Block Coordinate Descent (BCD) view of graph algorithms and propose an asynchronous heterogeneous graph analytic framework, GraphABCD, using the BCD view. The BCD view offers key insights and trade-offs on achieving high convergence rate of iterative graph algorithms. GraphABCD features fast convergence under the algorithm design options suggested by BCD. GraphABCD offers algorithm and architectural supports for asynchronous execution, without undermining its fast convergence properties. With minimum synchronization overhead, GraphABCD is able to scale out to heterogeneous and distributed accelerators efficiently. To demonstrate GraphABCD, we prototype its whole system on Intel HARPv2 CPU-FPGA heterogeneous platform. Evaluations on HARPv2 show that GraphABCD achieves geo-mean speedups of 4.8x and 2.0x over GraphMat, a state-of-the-art framework in terms of convergence rate and execution time, respectively.
Zhaoshi Li, Yangdong Deng, Shouyi Yin, Shaojun Wei, Leibo Liu
ISCA3
2020 Time-Triggered Switch-Memory-Switch Architecture for Time-Sensitive Networking Switches
abstract
Time-sensitive networking (TSN) is a set of extended standards for the IEEE 802.3 Ethernet under development by the IEEE 802.1 TSN task group. TSN depends on two key components, scheduling and fault tolerance, to provide realtime and reliable transmission. There is a strong motivation to replace the widely used field-buses with TSNs in industrial networking applications. However, industrial network devices are typical application-specific embedded systems with limited memory resources. Time-sensitive (TS) transmission certainly prefers on-chip memory, which is even more scarce for embedded systems. As a result, it is critical for TSNs to develop memory-efficient switching techniques with scalable schedulability and elegant fault-tolerance support. This paper proposes a time-triggered switch-memory-switch (SMS) architecture for memory-efficient TSN switches. First, based on the SMS shared memory, our architecture makes it possible to statically schedule memory allocation with full utilization for TS traffic and the remaining memory for other traffic. Compared with perport memory, the shared memory achieves a ratio of (nn/n!) (≈ (en/√(2πn)), n → ∞), where n is the port number, in the feasible solution space under memory constraints and thus significantly improves scheduling memory ability and flexibility. Moreover, we develop a fault-tolerance scheme for reliable transmission. It facilitates a memory-efficient implementation of the popular multiline redundancy in industrial networks. The scheme is validated by five classes of memory conflicts and a case study on two-line redundancy.
Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Model-Based Adaptation of Mixed-Criticality Multiservice Systems for Extreme Physical Environments
abstract
An increasingly important trend in the design of industry-strength embedded systems is the integration of multiple services with varying criticality levels into a common computing platform. Such systems are characterized as mixed-criticality multiservice systems (MCMSs). An MCMS has to survive in rigorous environments posed by industry-level requirements. Such survival, however, is becoming continuously more challenging due to the growing system complexity and integrating more and more services. While existing works typically target reliability-driven design optimization to improve the system robustness rather than deal with the surviving problem of the system in extreme physical environments, this paper addresses the problem by enabling the service capability transitions of an MCMS to adapt to the environments. This paper proposes a service capability model to capture the importance of functional modules for the criticality of different services. A model-based service-capability transition mechanism is designed to automatically identify the maximum allowed service capability under a given physical environment. A case study of the proposed techniques was performed on an industrial Ethernet switch which is a typical MCMS, to validate the capability of adaptation to high and low temperatures. The experimental results demonstrate the significant potential of our approach to improve system survivability under extreme physical environments.
Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2019 FPGA-Accelerated Optimistic Concurrency Control for Transactional Memory
abstract
Transactional Memory (TM) has been considered as a promising alternative to existing synchronization operations, which are often the largest stumbling block to unleashing parallelism of applications. Efficient implementations of TM, however, are challenging due to the tension between lowering performance overhead and avoiding unnecessary aborts.
Zhaoshi Li, Leibo Liu, Yangdong Deng, Shouyi Yin, Shaojun Wei
MICRO3
2019 An Enhanced Reconfiguration for Deterministic Transmission in Time-Triggered Networks
abstract
The emerging momentum of digital transformation of industry, i.e. Industry 4.0, poses strong demands for integrating industrial control networks, and Ethernet to enable the real-time Internet of Things (RT-IoT). Time-triggered (TT) networks provide a cost-efficient integrated solution while RT-IoT arouses the reconfiguration challenges: the network has to be flexible enough to adapt to changes and yet provides deterministic transmission persistently during network reconfiguration. Software defined network benefits the flexible industrial control by configuring the rules handling frames. However, previous reconfiguration mechanisms are mostly oriented to the context of data centers and wide area networks and thus do not consider the deterministic transmission in TT networks. This paper focuses on the reconfiguration (i.e., updates) for the deterministic transmission. To minimize the overhead during updates, namely the minimum number of loss frames and the minimum duration time of updates, we first establish an update theory based on the dependence relationship derived by the conflicts during updates. In addition then the reconfiguration problem is modeled with the dependence graph built by the relationship. On such a basis, we present a reconfiguration mechanism and its implementation to solve the problem. Finally, we evaluate the proposed reconfiguration mechanism in two real industrial network topologies. The experimental results demonstrate that compared with previous methods, our mechanism significantly reduces the number of loss frames and achieves zero loss in almost all cases.
Zonghui Li, Hai Wan, Zaiyu Pang, Qiubo Chen, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE/ACM Trans. Netw.5
2018 Energy-Efficient Automatic Train Driving by Learning Driving Patterns
abstract
Railway is regarded as the most sustainable means of modern transportation. With the fast-growing of fleet size and the railway mileage, the energy consumption of trains is becoming a serious concern globally. The nature of railway offers a unique opportunity to optimize the energy efficiency of locomotives by taking advantage of the undulating terrains along a route. The derivation of an energy-optimal train driving solution, however, proves to be a significant challenge due to the high dimension, nonlinearity, complex constraints, and time-varying characteristic of the problem. An optimized solution can only be attained by considering both the complex environmental conditions of a given route and the inherent characteristics of a locomotive. To tackle the problem, this paper employs a high-order correlation learning method for online generation of the energy optimized train driving solutions. Based on the driving data of experienced human drivers, a hypergraph model is used to learn the optimal embedding from the specified features for the decision of a driving operation. First, we design a feature set capturing the driving status. Next all the training data are formulated as a hypergraph and an inductive learning process is conducted to obtain the embedding matrix. The hypergraph model can be used for real-time generation of driving operation. We also proposed a reinforcement updating scheme, which offers the capability of sustainable enhancement on the hypergraph model in industrial applications. The learned model can be used to determine an optimized driving operation in real-time tested on the Hardware-in-Loop platform. Validation experiments proved that the energy consumption of the proposed solution is around 10% lower than that of average human drivers.
Jin Huang 0002, Yue Gao 0002, Xibin Zhao, Yangdong Deng, Ming Gu 0001
AAAI5
2018 R-FCN++: Towards Accurate Region-Based Fully Convolutional Networks for Object Detection
abstract
Region based detectors like Faster R-CNN and R-FCN have achieved leading performance on object detection benchmarks. However, in Faster R-CNN, RoI pooling is used to extract feature of each region, which might harm the classification as the RoI pooling loses spatial resolution. Also it gets slow when a large number of proposals are utilized. R-FCN is a fully convolutional structure that uses a position-sensitive pooling layer to extract prediction score of each region, which speeds up network by sharing computation of RoIs and prevents the feature map from losing information in RoI-pooling. But R-FCN can not benefit from fully connected layer (or global average pooling), which enables Faster R-CNN to utilize global context information. In this paper, we propose R-FCN++ to address this issue in two-fold: first we involve Global Context Module to improve the classification score maps by adopting large, separable convolutional kernels. Second we introduce a new pooling method to better extract scores from the score maps, by using row-wise or column-wise max pooling. Our approach achieves state-of-the-art single-model results on both Pascal VOC and MS COCO object detection benchmarks, 87.3% on Pascal VOC 2012 test dataset and 42.3% on COCO 2015 test-dev dataset. Code will be made publicly available.
Gang Yu 0002, Yangdong Deng
AAAI4
2018 DetNet: Design Backbone for Object Detection
Chao Peng 0001, Gang Yu 0002, Xiangyu Zhang 0005, Yangdong Deng, Jian Sun 0001
ECCV (9)5
2018 Work-in-Progress: A Flattened Priority Framework for Mixed-Criticality Real-Time Systems
abstract
Recent years witnessed a fast growing popularity of mixed-criticality real-time applications on smart devices. Priority schedulers are typically the central component to provide differential quality of service (QoS) for mixed-criticality tasks. The increasing deployment of such applications on smart devices, however, poses new challenges for the design of effective schedulers. First, the scheduling algorithms for mixed-criticality tasks are generally NP-complete. Second, the scheduling algorithms have to be effective enough under the limited computing resource of smart devices. This paper presents a work-in-progress report on a novel technique to design efficient and effective mixed-criticality schedulers. We propose a flattened priority framework to transform a non-priority scheduler into a priority one. The framework is typically a iterative framework based on feedback loops. Given an optimal nonpriority scheduler, for P priorities, the transformed scheduler converges with P iterations in the worst case. With the proposed framework, the design of priority schedulers is simplified into the design of non-priority schedulers. Such a simplification dramatically lowers the design effort and system complexity. A case study was performed on FPGA-based Industrial Ethernet switches. The proposed method achieves a 30%~50% reduction in the usage of look-up tables (LUTs) without performance loss.
Zonghui Li, Hai Wan, Yangdong Deng, Ming Gu 0001
RTAS3
2018 Triggered-Issuance and Triggered-Execution: A Control Paradigm to Minimize Pipeline Stalls in Distributed Controlled Coarse-Grained Reconfigurable Arrays
abstract
Distributed controlled coarse-grained reconfigurable arrays (CGRAs) enable efficient execution of irregular control flows by reconciling divergence in the processing elements (PEs). To further improve performance by better exploiting spatial parallelism, the triggered instruction architecture (TIA) eliminates the program counter and branch instructions by converting control flows into predicate dependencies as triggers. However, pipeline stalls, which occur in pipelines composed of both intra and inter-PEs, remain a major obstacle to the overall performance. In fact, the stalls in distributed controlled CGRAs pose a unique problem that is difficult to resolve by previous techniques. This work presents a triggered-issuance and triggered-execution (TITE) paradigm in which the issuance and execution of instructions are separately triggered to further relax the predicate dependencies in TIA. In this paradigm, instructions are paired as dual instructions to eliminate stalls caused by control divergence. Tags that identify the data transmitted between PEs are forwarded for acceleration. As a result, pipeline stalls of both intra- and inter-PEs can be significantly minimized. Experiments show that TITE improves performance by 21 percent, energy efficiency by 17 percent, and area efficiency by 12 percent compared with a baseline TIA.
Yanan Lu, Leibo Liu, Yangdong Deng, Jian Weng 0004, Shouyi Yin, Yiyu Shi 0001, Shaojun Wei
IEEE Trans. Parallel Distributed Syst.3
2017 Human experience knowledge induction based intelligent train driving
abstract
As the most sustainable means of modern transportation, the railway trains are eagerly approaching autonomous driving due to their congenital advantages on operating environments compare to, e.g., road traffics. The intelligent automatic train driving aims at train control with a goal of energy efficiency, punctuality and safety. The derivation of an optimized train driving solution by taking advantage of the undulating terrains along a route, however, proves to be a significant challenge due to the high dimension, nonlinearity, complex constraints, and time-varying characteristic of the problem. To tackle the problem, we propose a two-level human driving experience learning framework and employ the fuzzy rule induction method for online generation of the optimized driving solutions. Based on the records of experienced human drivers, a FURIA model was built to learn the driving rules indicating the correlation between the specified features to the decision of a driving sequence. The fuzzy rules can generally find the best-match driving operation under certain running circumstances. The learned model can be used to determine an optimized driving operation in real-time. Validation experiments show that the energy consumption of the proposed solution is around 8.93% lower than that of average human drivers.
Jin Huang 0002, Yangdong Deng, Xibin Zhao, Ming Gu 0001
ICIS3
2017 Minimizing Pipeline Stalls in Distributed-Controlled Coarse-Grained Reconfigurable Arrays with Triggered Instruction Issue and Execution
abstract
The pipeline stall in distributed-controlled coarse-grained reconfigurable arrays is a major source stumbling performance. This work presents a Triggered-Issue and Triggered-Execution (TITE) paradigm motivated from the Triggered Instruction Architecture (TIA) which converts control and data dependencies into predicate dependencies as triggers for spatial parallelism. TITE separately triggers the issuing and execution of instructions to further relax the predicate dependencies in TIA. Triggered dual instructions and tag forwarding are proposed to minimize pipeline stalls of both intra and inter-processing elements. Experiments show that TITE improves performance, energy efficiency, and area efficiency by 21%, 17%, and 12%, respectively, compared with TIA.
Yanan Lu, Leibo Liu, Yangdong Deng, Jian Weng 0004, Zhaoshi Li, Chenchen Deng, Shaojun Wei
DAC3
2017 Aggressive Pipelining of Irregular Applications on Reconfigurable Hardware
Zhaoshi Li, Leibo Liu, Yangdong Deng, Shouyi Yin, Shaojun Wei
ISCA3
2017 Path compression kd-trees with multi-layer parallel construction a case study on ray tracing
abstract
Kd-tree is a fundamental data structure with extensive applications in computer graphics. The performance of many interactive applications such as real-time ray tracing hinges on the construction and traversal efficiency of kd-trees. In recent years, there is a pressing demand for accelerating the construction process due to the fast-growing need of handling dynamic scenes. Existing construction algorithms typically follow a layer-by-layer scheme, which significantly limits the efficiency on the use of multi-core CPUs and GPUs. In this paper, we propose a concurrent multi-layer kd-tree construction algorithm to unleash the inherent parallelism. For a given scene, the algorithm uses Morton code to split its bounding box and orders primitives by Morton curve. A path compression procedure is then concurrently executed on all essential nodes that contain primitives to generate the hierarchy in the target kd-tree. All redundant nodes that have no primitives along the compression paths are collapsed to fast slip empty space. The fully parallel algorithmic scheme adapts variable primitives space and drastically shortens the construction time. A case study on ray tracing benchmarks demonstrates that our kd-tree construction method outperforms the state of art work by an average factor of over 10 and still enables high performance traversal.
Zonghui Li, Yangdong Deng, Ming Gu 0001
I3D2
2017 Toward Robust Vehicle Platooning With Bounded Spacing Error
abstract
Intelligent transportation has become an essential field of cyber-physical systems. Among various intelligent transportation technologies, the automated highway system (AHS) has its unique advantage of being able to coordinate a platoon of vehicles as a whole unit. The major challenge of building a robust AHS is the nonlinear and (potentially) fast time-varying uncertainty induced by parameter variations and external disturbances. Finally, reflected as the spacing between neighboring vehicles, such uncertainties can be a serious concern for maintaining safety. This paper addresses the problem by proposing a mathematical transformation scheme to bound the spacing error and build a distributed control algorithm on such a basis. The propose algorithm achieves a spacing error satisfying both uniformly boundedness and uniformly ultimate boundedness. Our decentralized algorithm is communication efficient in the sense that it only requires the state information of the preceding car and the acceleration feedback and does not need to communicate with all other cars.
Jin Huang 0002, Qingmin Huang, Yangdong Deng, Ye-Hwa Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 An Energy-Efficient Train Control Framework for Smart Railway Transportation
abstract
Railway transportation systems are the backbone of smart cities. With the rapid increasing of railway mileage, the energy consumption of train becomes a major concern. The uniqueness of train operations is that the geographic characteristics of each route is known a priori. On the other hand, the parameters (e.g., loads) of a train varies from trip to trip. Such a specialty determines that an energy-optimal driving profile for each train operation has to be pursued by considering both the geographic information and the inherent train conditions. The solution of the optimization problem, however, is hard due to its high dimension, nonlinearity, complex constraints and time-varying characteristics of a control sequence. As a result, an energy-saving solution to the train control optimization problem has to address the dilemma of optimization quality and computing time. This work proposes an energy-efficient train control framework by integrating both offline and onboard optimization techniques. The offline processing builds a decision tree based sketchy solution through a complete flow of sequence mining, optimization and machine learning. The onboard system feeds the train parameters into the decision tree to derive an optimized control sequence. A key innovation of this work is the identification of optimal patterns of control sequence by data mining the driving behaviors of the experienced train drivers and then apply the patterns to online trip planning. The proposed framework efficiently find an optimized driving solution by leveraging the training results derived with a compute-intensive offline learning flow. The framework was already testified in a smart freight train system. It was demonstrated an average of$9.84$percent energy-saving can be achieved.
Jin Huang 0002, Yangdong Deng, Qinwen Yang, Jia-Guang Sun 0001
IEEE Trans. Computers2
2016 Time-Delay Neural Network for Continuous Emotional Dimension Prediction From Facial Expression Sequences
abstract
Automatic continuous affective state prediction from naturalistic facial expression is a very challenging research topic but very important in human-computer interaction. One of the main challenges is modeling the dynamics that characterize naturalistic expressions. In this paper, a novel two-stage automatic system is proposed to continuously predict affective dimension values from facial expression videos. In the first stage, traditional regression methods are used to classify each individual video frame, while in the second stage, a time-delay neural network (TDNN) is proposed to model the temporal relationships between consecutive predictions. The two-stage approach separates the emotional state dynamics modeling from an individual emotional state prediction step based on input features. In doing so, the temporal information used by the TDNN is not biased by the high variability between features of consecutive frames and allows the network to more easily exploit the slow changing dynamics between emotional states. The system was fully tested and evaluated on three different facial expression video datasets. Our experimental results demonstrate that the use of a two-stage approach combined with the TDNN to take into account previously classified frames significantly improves the overall performance of continuous emotional state estimation in naturalistic facial expressions. The proposed approach has won the affect recognition sub-challenge of the Third International Audio/Visual Emotion Recognition Challenge.
Hongying Meng, Nadia Bianchi-Berthouze, Yangdong Deng, Jinkuang Cheng, John Cosmas
IEEE Trans. Cybern.3
2015 FastTree: a hardware KD-tree construction acceleration engine for real-time ray tracing
Yangdong Deng, Yufei Ni, Zonghui Li
DATE2
2015 RadixBoost: A hardware acceleration structure for scalable radix sort on graphic processors
abstract
In this paper, we propose RadixBoost, a hardware acceleration structure for scalable 32-bit integer radix sort on GPU. The whole structure is integrated into a GPU microarchitecture as a special functional unit and can be started by new instructions. Our design enables a significantly faster sorting procedure for general purpose GPU computing. The RadixBoost architecture was validated by an FPGA prototype integrated in FPGA-based GPU microarchitecture simulator, Fastlanes. An ASIC evaluation of RadixBoost was also performed. Our results proved that RadixBoost outperformed its GPU software equivalent by a factor of over 6 with an 1% and 3% increase in area and power respectively in cutting-edge Fermi GPU.
Shikai Li, Kuan Fang, Yufei Ni, Zonghui Li, Yangdong Deng
ISCAS6
2015 GPU accelerated sparse matrix-vector multiplication and sparse matrix-transpose vector multiplication
abstract
Summary Many high performance computing applications require computing both sparse matrix‐vector product (SMVP) and sparse matrix‐transpose vector product (SMTVP) for better overall performance. Under such a circumstance, it is critical to maintain a similarly high throughput for these two computing patterns with the underlying sparse matrix encoded in a single storage format. The compressed sparse block (CSB) format proposed by Buluç et al. allows computing both problems on multi‐core CPUs with nearly identical throughputs. On the other hand, a direct porting of CSB to graphics processing units (GPUs), which have been recently recognized as a powerful general purpose computing platform, turns out to be inefficient. In this work, we propose a new data structure, designated as expanded CSB (eCSB), to minimize the throughput gap between SMVP and SMTVP computations on GPUs, while at the same time enable a high computing throughput. We also use a hybrid storage format to store elements in each block, which can be selected dynamically at runtime. Experimental results show that the proposed techniques implemented on a Kepler GPU delivers similar throughput on both SMVP and SMTVP and the throughput is up to 13 times faster than that of the CPU‐based CSB implementation. In addition, our eCSB procedure outperforms the previous GPU results by up to 188% and 914% in computing SMVP and SMTVP, and we validate the effectiveness of eCSB by means of wall‐clock time of bi‐conjugate gradient algorithm; our eCSB is 25% faster than Compressed Sparse Rows (CSR) and 6% faster than HYB, respectively. Copyright © 2014 John Wiley & Sons, Ltd.
Yangdong Deng, Shuai Mu 0002, Zhenzhong Zhang, Mingfa Zhu
Concurr. Comput. Pract. Exp.2
2014 Atomic reduction based sparse matrix-transpose vector multiplication on GPUs
abstract
Sparse Matrix-Transpose Vector Product (SMTVP) is a frequently used computation pattern in High Performance Computing applications. It is typically solved by transposition followed by a Sparse Matrix-Vector Product (SMVP) in current linear algebra packages. However, the transposition process can be a serious bottleneck on modern parallel computing platforms. A previous work proposed a relatively complex data structure for efficiently computing SMTVP with multi-core CPUs, but it proved to be inefficient on GPUs. In this work, we show that the Compressed Sparse Row (CSR) based SMVP algorithm can also be efficient for SMTVP computation on modern GPUs. The proposed method exploits atomic operations to perform the reduce operation in the computation of each inner product of a row in the transposed matrix and the vector. Experimental results show that the simple technique can outperform the SMTVP flow of transposition plus SMVP released in the CUSPARSE package by up to 405-fold.
Yangdong Deng, Shuai Mu 0002, Mingfa Zhu, Zhibin Huang
ICPADS2
2014 Fully parallel kd-tree construction for real-time ray tracing
abstract
This work proposes a fully parallel kd-tree construction algorithm, which depends on the Morton code to identify all candidate split planes and derive their exact positions in parallel. Our techniques drastically shorten construction process. Experimental results on a set of frequently used scenes prove that the proposed kd-tree construction algorithm outperforms a state-of-the-art algorithm kd-tree construction algorithm by over one order of magnitude.
Zonghui Li, Tong Wang 0036, Yangdong Deng
I3D3
2014 A Fast and Accurate Segmentation Method for Ordered LiDAR Point Cloud of Large-Scale Scenes
abstract
This letter proposes an efficient two-step segmentation method for large-scale 3-D point cloud data collected by the mobile laser scanners. First, a new scan-line-based ground segmentation algorithm is designed to filter the points corresponding to the ground with high accuracy. Second, we propose a selfadaptive Euclidean clustering algorithm to further separate the off-ground points corresponding to different objects. Experiments show that our method delivers superior segmentation results on scanned data. In fact, the proposed method can be used in complex scenes including slope and bumpy road at an error rate of 0.674% and a computing throughput of over 20 million points/s.
Dan Wang 0020, Yiyi Ren, Guolin Li, Yangdong Deng, Zhihua Wang 0001
IEEE Geosci. Remote. Sens. Lett.6
2014 Orchestrating Cache Management and Memory Scheduling for GPGPU Applications
abstract
Modern graphics processing units (GPUs) are delivering tremendous computing horsepower by running tens of thousands of threads concurrently. The massively parallel execution model has been effective to hide the long latency of off-chip memory accesses in graphics and other general computing applications exhibiting regular memory behaviors. With the fast-growing demand for general purpose computing on GPUs (GPGPU), GPU workloads are becoming highly diversified, and thus requiring a synergistic coordination of both computing and memory resources to unleash the computing power of GPUs. Accordingly, recent graphics processors begin to integrate an on-die level-2 (L2) cache. The huge number of threads on GPUs, however, poses significant challenges to L2 cache design. The experiments on a variety of GPGPU applications reveal that the L2 cache may or may not improve the overall performance depending on the characteristics of applications. In this paper, we propose efficient techniques to improve GPGPU performance by orchestrating both L2 cache and memory in a unified framework. The basic philosophy is to exploit the temporal locality among the massive number of concurrent memory requests and minimize the impact of memory divergence behaviors among simultaneously executed groups of threads. Our major contributions are twofold. First, a priority-based cache management is proposed to maximize the chance of frequently revisited data to be kept in the cache. Second, an effective memory scheduling is introduced to reorder memory requests in the memory controller according to the divergence behavior for reducing average waiting time of warps. Simulation results reveal that our techniques enhance the overall performance by 10% on average for memory intensive benchmarks, whereas the maximum gain can be up to 30%.
Shuai Mu 0002, Yangdong Deng, Yubei Chen, Huaiming Li, Jianming Pan, Zhihua Wang 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2013 FastLanes: An FPGA accelerated GPU microarchitecture simulator
abstract
Graphic Processing Units (GPUs) have emerged as a new general purpose computing platform that attracts significant research efforts. Currently, GPU architecture research resorts to time-consuming software simulations to evaluate microarchitecture innovations. In this paper, we propose FastLanes, an FPGA based simulator for a generic GPU microarchitecture, to enable hardware-accelerated simulation. FastLanes consists of a function model and a timing model, both implemented on FPGA. The functional model implements the full functionality of a multiprocessor of GPU and emulates multiple multiprocessors via time-division multiplexing. We develop a hybrid implementation strategy in which certain GPU logic is directly mapped to FPGA while the other logic is simulated by reusing the same FPGA logic. A corresponding context shifting mechanism is proposed to store execution states of threads from FPGA to external on-board memory, and vice versa. Such a mechanism makes it possible to simulate hundreds of GPU cores on a single FPGA evaluation board. Driven by the functional simulation results, the timing model considers the detailed configuration of GPU microarchitecture to derive the performance evaluation. A compiler tool-chain is also developed to allow the execution of NVIDIA GPU binary on FastLanes. Experimental results prove that FastLanes outperforms its software equivalent by up to 2 orders of magnitude.
Kuan Fang, Yufei Ni, Jiayuan He 0005, Zonghui Li, Shuai Mu 0002, Yangdong Deng
ICCD6
2013 Design and optimization of multi-clocked embedded systems using formal technique
abstract
Today’s system-on-chip and distributed systems are commonly equipped with multiple clocks. The key challenge in designing such systems is that heterogenous control-oriented and data-oriented behaviors within one clock domain, and asynchronous communications between two clock domains have to be captured and evaluated in a single framework. In this paper, we propose to use timed automata and synchronous dataflow to capture the dynamic behaviors of multi-clock embedded systems. A timed automata and synchronous dataflow based modeling and analyzing framework is constructed to evaluate and optimize the performance of multiclock embedded systems. Data-oriented behaviors are captured by synchronous dataflow, while synchronous control-oriented behaviors are captured by timed automata, and inter clock-domain asynchronous communication can be modeled in an interface timed automaton or a synchronous dataflow module with the CSP mechanism. The behaviors of synchronous dataflow are interpreted by some equivalent timed automata to maintain the semantic consistency of the mixed model. Then, various functional properties can be simulated and verified within the framework. We apply this framework in the design process of a sub-system that is used in real world subway communication control system
Yu Jiang 0001, Zonghui Li, Hehua Zhang, Yangdong Deng, Ming Gu 0001, Jia-Guang Sun 0001
ESEC/SIGSOFT FSE4
2012 A Thermal-Driven Test Application Scheme for 3-Dimensional ICs
abstract
In this work, we propose a novel scan architecture for 3-D ICs by considering the interconnection overhead of through-silicon-vias (TSVs). Since hotspots in 3-DICs often cause performance and reliability issues, we also developed a new test ordering scheme to avoid applying test vectors that could worsen the temperature distribution. Experimental results show that the peak temperature can be lowered by 20% by the 3-D scan tree architecture. When combined with the test ordering scheme, the 3-D scan tree can further reduce peak temperature by over 30%.
Kele Shen, Yangdong Deng
Asian Test Symposium3
2012 A theoretical and empirical error analysis of mobile 3D data acquisition system
abstract
This paper proposes a theoretic error analysis framework as well as corresponding calibration methods for a mobile 3D data acquisition system equipped with LiDAR and GPS/IMU. Our framework considers inherent errors of variable sensors, installation errors, and time synchronization errors among different sensors. Our framework is enhanced by an empirical error analysis method based on a triangulated irregular network structure. Experimental results prove that the proposed error analysis methodology can provide accurate and efficient evaluation for mobile 3D data acquisition systems.
Yiyi Ren, Wenshou Chen, Guolin Li, Yangdong Deng, Enbo Shi, Zhihua Wang 0001
ISCAS5
2012 Towards accelerating irregular EDA applications with GPUs
Hao Qian 0001, Yangdong Deng, Bo D. Wang, Shuai Mu 0002
Integr.2
2012 A Framework for Layout-Dependent STI Stress Analysis and Stress-Aware Circuit Optimization
abstract
With the continuous shrinking of the feature size, the effect of stress on the performance of the IC device and circuit can no longer be ignored. In fact, stress engineering is becoming more and more widely used today in advanced IC manufacture processes to improve device performance. Different from the intentionally introduced stresses to improve circuit performance, the shallow-trench-isolation (STI) stress, which is exerted by STI wells on the active area of devices, is a by-product of the fabrication process and has increasingly significant impact on the circuit behavior. This paper proposes a complete flow to characterize the influence of STI stress on the performance of RF/analog circuits by considering detailed layout and process information. An accurate and efficient finite-element method-based stress simulator has been developed to extract stress distribution from layouts of IC designs. The existing MOSFET model is also enhanced to capture the effects of stress on mobility, threshold voltage. With the enhanced model, we are able to study the influence of layout-dependent STI stress on the performance of real circuits and establish corresponding optimization strategies. The proposed flow has been applied to a series of RF/analog IC designs based on a 90-nm CMOS technology.
Jiying Xue, Yangdong Deng, Zuochang Ye, Zhiping Yu
IEEE Trans. Very Large Scale Integr. Syst.2
2011 Hermes: an integrated CPU/GPU microarchitecture for IP routing
abstract
With the constantly increasing Internet traffic and fast changing network protocols, future routers have to simultaneously satisfy the requirements for throughput, QoS, flexibility, and scalability. In this work, we propose a novel integrated CPU/GPU microarchitecture, Hermes, for QoS-aware high speed routing. We also develop a new thread scheduling mechanism, which significantly improves all QoS metrics.
Yuhao Zhu 0001, Yangdong Deng, Yubei Chen
DAC2
2011 Scalable packet classification via GPU metaprogramming
abstract
Packet classification has been a fundamental processing pattern of modern networking devices. Today's high-performance routers use specialized hardware for packet classification, but such solutions suffer from prohibitive cost, high power consumption, and poor extensibility. On the other hand, software-based routers offer the best flexibility, but could only deliver limited performance (<;10Gbps). Recently, graphics processing units (GPUs) have been proved to be an efficient accelerator for software routers. In this work, we propose a GPU-based linear search framework for packet classification. The core of our framework is a metaprogramming technique that dramatically enhances the execution efficiency. Experimental results prove that our solution could outperform a CPU-based solution by a factor of 17, in terms of classification throughput. Our technique is scalable to large rule sets consisting of over 50K rules and thus provides a solid foundation for future applications of packet context inspection.
Kang Kang, Yangdong Deng
DATE2
2011 Evaluating the potential of graphics processors for high performance embedded computing
abstract
Today's high performance embedded computing applications are posing significant challenges for processing throughout. Traditionally, such applications have been realized on application specific integrated circuits (ASICs) and/or digital signal processors (DSP). However, ASICs' advantage in performance and power often could not justify the fast increasing fabrication cost, while current DSP offers a limited processing throughput that is usually lower than 100GFLOPS. On the other hand, current multi-core processors, especially graphics processing units (GPUs), deliver very high computing throughput, and at the same time maintain high flexibility and programmability. It is thus appealing to study the potential of GPUs for high performance embedded computing. In this work, we perform a comprehensive performance evaluation on GPUs with the high performance embedded computing (HPEC) benchmark suite, which consist a broad range of signal processing benchmarks with an emphasis on radar processing applications. We develop efficient GPU implementations that could outperform previous results for all the benchmarks. In addition, a systematic instruction level analysis for the GPU implementations is conducted with a GPU micro-architecture simulator. The results provide key insights on optimizing GPU hardware and software. Meanwhile, we also compared the performance and power efficiency between GPU and DSP with the HPEC benchmarks. The comparison reveals that the major hurdle for GPU's applications in embedded computing is its relatively low power efficiency.
Shuai Mu 0002, Maohua Zhu, Yangdong Deng
DATE8
2011 Accelerating RTL simulation with GPUs
abstract
With the fast increasing complexity of integrated circuits, verification has become the bottleneck of today's IC design flow. In fact, over 70% of the IC design turn-around time can be spent on the verification process in a typical IC design project. Among various verification tasks, Register Transfer Level (RTL) simulation is the most widely used method to validate the correctness of digital IC designs. When simulating a large IC design with complicated internal behaviors (e.g., CPU cores running embedded software), RTL simulation can be extremely time consuming. Since RTL-to-layout is still the most prevalent IC design methodology, it is essential to speedup the RTL simulation process. Recently, General Purpose computing on Graphics Processing Units (GPGPU) is becoming a promising paradigm to accelerate computing-intensive workloads. A few recent works have demonstrated the effectiveness of using GPU to expedite gate and system level simulation tasks. In this work, we proposed an efficient GPU-accelerated RTL simulation framework. We introduce a methodology to translate Verilog RTL description into equivalent GPU source code so as to simulate circuit behavior on GPUs. In addition, a CMB based parallel simulation protocol is also adopted to provide a sufficient level of parallelism. Because RTL simulation lacks data-level parallelism, we also present a novel solution to use GPU as an efficient task-level parallel processor. Experimental results prove that our GPU based simulator outperforms a commercial sequential RTL simulator by over 20 fold.
Hao Qian 0001, Yangdong Deng
ICCAD2
2011 Exploiting graphics processors for high-performance IP lookup in software routers
abstract
As the physical link speeds grow and the size of routing table continues to increase, IP address lookup has been a challenging problem at routers. There have been growing demands in achieving high-performance IP lookup cost-effectively. Existing approaches typically resort to specialized hardwares, such as TCAM. While these approaches can take advantage of hardware parallelism to achieve high-performance IP lookup, they also have the disadvantage of high cost. This paper investigates a new way to build a cost-effective IP lookup scheme using graphics processor units (GPU). Our contribution here is to design a practical architecture for high-performance IP lookup engine with GPU, and to develop efficient algorithms for routing prefix update operations such as deletion, insertion, and modification. Leveraging GPU's many-core parallelism, the proposed schemes addressed the challenges in designing IP lookup at GPU-based software routers. Our experimental results on real-world route traces show promising gains in IP lookup and update operations.
Jin Zhao 0001, Xinya Zhang, Xin Wang 0002, Yangdong Deng, Xiaoming Fu 0001
INFOCOM4
2011 Massively Parallel Logic Simulation with GPUs
abstract
In this article, we developed a massively parallel gate-level logical simulator to address the ever-increasing computing demand for VLSI verification. To the best of the authors’ knowledge, this work is the first one to leverage the power of modern GPUs to successfully unleash the massive parallelism of a conservative discrete event-driven algorithm, CMB algorithm. A novel data-parallel strategy is proposed to manipulate the fine-grain message passing mechanism required by the CMB protocol. To support robust and complete simulation for real VLSI designs, we establish both a memory paging mechanism and an adaptive issuing strategy to efficiently utilize the GPU memory with a limited capacity. A set of GPU architecture-specific optimizations are performed to further enhance the overall simulation performance. On average, our simulator outperforms a CPU baseline event-driven simulator by a factor of 47.4X. This work proves that the CMB algorithm can be efficiently and effectively deployed on modern GPUs without the performance overhead that had hindered its successful applications on previous parallel architectures.
Yuhao Zhu 0001, Bo D. Wang, Yangdong Deng
ACM Trans. Design Autom. Electr. Syst.3
2010 Distributed time, conservative parallel logic simulation on GPUs
abstract
Logical simulation is the primary method to verify the correctness of IC designs. However, today's complex VLSI designs pose ever higher demand for the throughput of logic simulators. In this work, a parallel logic simulator was developed by leveraging the computing power of modern graphics processing units (GPUs). To expose more parallelism, we implemented a conservative parallel simulation approach, the CMB algorithm, on NVidia GPUs. The simulation processing is mapped to GPU hardware at the finest granularity. With carefully designed data structures and data flow organizations, our GPU based simulator could overcome many problems that hindered efficient implementations of the CMB algorithm on traditional parallel computers. In order to efficiently use the relatively limited capacity of GPU memory, a novel memory management mechanism was proposed to dynamically allocate and recycle GPU memory during simulation. We also introduced a CPU/GPU co-processing strategy for the best usage of computing resources. Experimental results showed that our GPU based simulator could outperform a CPU baseline event driven simulator by a factor of 29.2.
Bo D. Wang, Yuhao Zhu 0001, Yangdong Deng
DAC3
2010 IP routing processing with graphic processors
abstract
Throughput and programmability have always been the central, but generally conflicting concerns for modern IP router designs. Current high performance routers depend on proprietary hardware solutions, which make it difficult to adapt to ever-changing network protocols. On the other hand, software routers offer the best flexibility and programmability, but could only achieve a throughput one order of magnitude lower. Modern GPUs are offering significant computing power, and its data-parallel computing model well matches the typical patterns of packet processing on routers. Accordingly, in this research we investigate the potential of CUDA-enabled GPUs for IP routing applications. As a first step toward exploring the architecture of a GPU based software router, we developed GPU solutions for a series of core IP routing applications such as IP routing table lookup and pattern match. For the deep packet inspection application, we implemented both a Bloom-filter based string matching algorithm and a finite automata based regular expression matching algorithm. A GPU based routing table lookup solution is also proposed in this work. Experimental results proved that GPU could accelerate the routing processing by one order of magnitude. Our work suggests that, with proper architectural modifications, GPU based software routers could deliver significant higher throughput than previous CPU based solutions.
Shuai Mu 0002, Xinya Zhang, Nairen Zhang, Yangdong Deng
DATE5
2010 Full-chip leakage analysis for 65 nm CMOS technology and beyond
Jiying Xue, Yangdong Deng, Zhiping Yu
Integr.3
2009 Taming irregular EDA applications on GPUs
abstract
Recently general purpose computing on graphic processing units (GPUs) is rising as an exciting new trend in high-performance computing. Thus it is appealing to study the potential of GPU for Electronic Design Automation (EDA) applications. However, EDA generally involves irregular data structures such as sparse matrix and graph operations, which pose significant challenges for efficient GPU implementations. In this paper, we propose highperformance GPU implementations for two important irregular EDA computing patterns, Sparse-Matrix Vector Product (SMVP) and graph traversal. On a wide range of EDA problem instances, our SMVP implementations outperform all published work and achieve a speedup of one order of magnitude over the CPU baseline. Upon such a basis, both timing analysis and linear system solution can be considerably accelerated. We also introduce a SMVP based formulation for Breadth-First Search and observe considerable speedup on GPU implementations. Our results suggest that the power of GPU computing can be successfully unleashed through designing GPU-friendly algorithms and/or re-organizing computing structures of current algorithms.
Yangdong Deng, Bo D. Wang, Shuai Mu 0002
ICCAD1
2009 Layout-dependent STI stress analysis and stress-aware RF/analog circuit design optimization
abstract
With the continuous shrinking of feature size, various effects due to shallow-trench-isolation (STI) stress are becoming more and more significant. The resulting nonuniform distribution of stress affects the MOSFET characteristics and hence changes the circuit behavior. This paper proposes a complete flow to characterize the influence of STI stress on performance of RF/analog circuits based on layout design and process information. An accurate and efficient FEM-based stress simulator has been developed to handle the layout dependence. A comprehensive MOSFET model is also proposed to capture the effects of STI stress on mobility, threshold voltage, and leakage current. The influence of layout-dependent STI stress on the circuit performance is further studied, and the corresponding optimization strategies to circuit design are discussed. A realistic PLL design realized using 90nm CMOS technology is used as a test case for the proposed approach.
Jiying Xue, Zuochang Ye, Yangdong Deng, Zhiping Yu
ICCAD3
2005 Temperature-Dependent Optimization of Cache Leakage Power Dissipation
abstract
Leakage power consists of an increasing portion of the total power consumption for modern IC designs. Due to the strong inter-dependency between leakage and temperature, it becomes imperative to consider the thermal effects while optimizing the leakage power. In this paper, we present a temperature-dependent optimization methodology for on-chip caches. By integrating fast yet accurate coupled thermal-leakage simulations into an optimization flow, we are able to optimally tradeoff between the cache performance and leakage power while considering realistic on-chip temperature distribution. Our analysis indicates that for future memory intensive designs, the lack of chip temperature information can cause a significant error in the leakage power estimation, thus leading to non-optimal cache designs. Our results further imply that the optimization of cache performance and leakage power shall be attacked as part of the whole system design task in which chip-level floor planning and its thermal impacts are fully addressed.
Peng Li 0001, Yangdong Deng, Lawrence T. Pileggi
ICCD2
2005 2.5-dimensional VLSI system integration
abstract
The excessive interconnection delay and fast increasing development cost, as well as complexity of the single-chip integration of different technologies, are likely to become the major stumbling blocks for the success of monolithic system-on-chips. To address the above problems, this paper investigates a new VLSI integration paradigm, the so-called 2.5-dimensional (2.5-D) integration scheme. Using this scheme, a VLSI system is implemented as a three-dimensional stacking of monolithic chips. A cost analysis framework was developed to justify the 2.5-D integration scheme from an economic point of view. Enabling technologies for the new integration scheme are also reviewed.
Yangdong Deng, Wojciech Maly
IEEE Trans. Very Large Scale Integr. Syst.1
2004 2.5D system integration: a design driven system implementation schema
Yangdong Deng, Wojciech Maly
ASP-DAC1
2003 Physical Design of the "2.5D" Stacked System
abstract
Excessive on-chip wire length and fast increasing fabrication cost have been the main factors impairing the effectiveness of monolithic system-on-chip. We investigate a die stacking based system integration strategy (2.5D system integration) to address these problems. The new scheme is design-tools-enabled rather than technology-driven. We developed a layout design framework, which is able to floorplan, place and route a VLSI design into stacked chips. Our results show that this new scheme has a potential to outperform its monolithic equivalent.
Yangdong Deng, Wojciech Maly
ICCD1
2001 Interconnect characteristics of 2.5-D system integration scheme
abstract
Growing number of excessively long on-chip wires in modern monolithic ICs is a byproduct of growing chip size. To address this problem instead of placing all systems components in one layer (i.e. in 2-D space) one can use a stack of single layer monolithic ICs (called here a 2.5-D integrated IC). To assess the potential benefits of such a 2.5-D integration schema this paper compares wire length distributions, obtained for 2-D and 2.5-D implementations of benchmark circuits. In the assessment two newly developed floorplanning and placement tools were used. Significant reductions in both total wirelength and worst-case wirelength was observed for the systems implemented as 2.5-D ICs.
Yangdong Deng, Wojciech Maly
ISPD1