EDBT 2026 Demo / reviewers in the wild / expert
Dajiang Liu
dblp:123/8216
· DBLP profile ↗
44ranked-venue papers
9as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 9 first-author · 17 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Software engineering, systems software and programming languages · 6 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Data Transfer Optimization for Loop Mapping on CGRAs via Polyhedral TransformationabstractCoarse-Grained Reconfigurable Arrays (CGRAs) play an important role in accelerating computation-intensive applications due to their flexibility and power efficiency. Most existing mapping works for CGRA assume that all the data of a loop kernel can be previously moved from DRAM to ScratchPad Memory (SPM) with one data transfer. However, when targeting low-power CGRAs, the SPM can hardly hold the entire data even for a small loop kernel, and it inevitably incurs multiple data transfers, leading to notable costs. In this case, where and how to perform data transfer considering the DRAM access behavior is crucial to the overall execution performance. To this end, this paper proposes a data transfer optimization method to simultaneously optimize loop pipelining and data transfer. Since there is an interplay between loop structure and data transfer, we establish an optimization problem considering both polyhedral transformation and transfer strategy exploration, offering opportunities to find the near-optimal data transfer strategy within the search space. The experimental results show that compared with existing methods, our method can achieve 1.36×–4.55× speedup for loop kernels with different on-chip buffer sizes, and only increase the compilation time by a small amount. Zhaorui Chen, Liao Huang, Xiao Xiong, Dajiang Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | DynaX: Sparse Attention Acceleration with Dynamic X: M Fine-Grained Structured Pruning
Xiao Xiong, Zhaorui Chen, Yue Liang 0004, Minghao Tian, Jiaxing Shang, Dajiang Liu |
ASPLOS (2) | 7 |
| 2025 | SDVD: Self-supervised dual-view modeling of user and cascade dynamics for information diffusion prediction
Haoyu Xiong, Jiaxing Shang, Fei Hao 0001, Dajiang Liu, Geyong Min |
Knowl. Based Syst. | 4 |
| 2025 | CoSpMV: Towards Agile Software and Hardware Co-Design for SpMV ComputationabstractSparse Matrix-Vector multiplication (SpMV) is a widely used kernel in scientific or engineering applications and it is commonly implemented in FPGAs for acceleration. Existing works on FPGA usually pre-process the sparse matrix for data compression from the software perspective, and then design a unified architecture from the hardware perspective. However, as different SpMV kernels expose different levels of data parallelism after software processing, a unified architecture may not efficiently tap the underlying parallelism exposed in a specific kernel, leading to poor bandwidth utilization (BU) or poor resource utilization. To this end, this paper proposes an agile software and hardware co-design framework, CoSpMV, that employs design space exploration on both software and hardware for a specific kernel. Specifically, by providing a scalable compressed data format and a highly pipelined hardware template, CoSpMV can select the most suitable software and hardware configurations for different kernels and generate the accelerator instantly. The experimental results show that CoSpMV can achieve 3.91$\times$speedup on GFLOPs, and 1.31$\times$speedup on BU compared to the state-of-the-art work. Minghao Tian, Yue Liang 0004, Dajiang Liu |
IEEE Trans. Computers | 4 |
| 2025 | Optimizing Data Reuse for Loop Mapping on CGRAs With Joint Affine and Nonaffine TransformationsabstractCoarse-grained reconfigurable arrays (CGRAs) can provide high energy efficiency while maintaining flexibility, which is promising to keep pace with the power requirements and the frequent updates of applicants. With flexible register chains, modern CGRAs enable data reuse within the processing element array (PEA) to reduce on-chip memory accesses and improve pipelining performance. However, existing works pay little attention to comprehensive loop transformations, such as affine and nonaffine transformation, to obtain a data reuse-friendly loop structure. Therefore, this article proposes a data-reuse-friendly loop mapping approach using joint affine and nonaffine transformations. With affine transformations, the distance of loop dependencies could be reduced and then handled by in-PEA routes. With nonaffine transformations (i.e., loop unrolling), small loop kernels could be unrolled and expose more memory accesses for data reuse. To efficiently solve the loop transformation problem, we first establish a reduced polyhedral formulation and then propose a divide-and-conquer-based solution to find optimized transformations with moderate compilation time. Experimental results demonstrate that our approach can achieve a speedup up to$1.74\times $compared to state-of-the-art methods. Liao Huang, Dajiang Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | PMP: Pattern Morphing-based Memory Partitioning in High-Level SynthesisabstractMemory partitioning is a widely used technique to reduce access conflicts on multi-bank memory in high-level synthesis. Previous memory partitioning methods mainly focus on a given access pattern extracted from stencil applications. Restricted by the pattern shape, these methods are prone to sub-optimal bank numbers or large overhead on address generation. In this work, we propose a pattern-morphing-based memory partitioning method, PMP, that only requires reduced hyperplane families to achieve the minimal bank number. To reduce the side effect of extra data padding, an integer linear programming problem is formulated for pattern morphing. Compared to the previous hyperplane-based memory partitioning, the experimental results show that our approach could achieve the optimal partition factor while saving 22% in LUTs, 21% in FlipFlops, 10% in DSPs, and 40% in memory overhead, on average. Dajiang Liu, Decai Pan, Xiao Xiong, Jiaxing Shang, Shouyi Yin |
DAC | 1 |
| 2024 | Optimizing Imperfectly-Nested Loop Mapping on CGRAs via Polyhedral-Guided FlatteningabstractCoarse-Grained Reconfigurable Arrays (CGRAs) offer a promising balance between high performance and power efficiency. To reduce the invocation overhead when mapping an imperfectly nested loop, loop flattening is used to transform the nested loop into a single-level loop. However, loop flattening not only leads to a big loop body but also has a narrow application scope. To this end, this work proposes a polyhedral model-based loop flattening approach for imperfectly nested loop mapping. By exploring loop structures via polyhedral transformation, we can find a flattening-friendly loop structure with more data reuse opportunities and reduced sibling loops, resulting in improved loop pipelining performance. Experimental results demonstrate a remarkable$(1.37-1.62\times)$speedup compared to the state-of-the-art approaches while maintaining short compilation times. Xingyu Mo, Dajiang Liu |
DATE | 3 |
| 2024 | E2EMap: End-to-End Reinforcement Learning for CGRA Compilation via Reverse MappingabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a promising architecture to cope with the challenges of increasing demand for high performance and high energy efficiency. However, the actual achieved performance of CGRA is highly dependent on the mappers. Traditional mappers using heuristics or combinatorial optimization can hardly learn from past experience, suffering from poor quality and portability. Recently, machine learning has been introduced to partial components in CGRA compilers, leaving other components to traditional heuristics, which is also prone to a sub-optimum, To this end, this paper proposes an end-to-end learning framework, E2EMap, for CGRA mapping that can cover the full mapping process. To reduce the complexity of the learning model, a reverse mapping problem is formulated, where various routing strategies can be thoroughly explored. To solve the problem, policy gradient reinforcement learning is introduced to learn from scratch. Experimental results demonstrate that E2EMap can achieve up to 2.23 x mapping quality across different CGRA settings while consuming even less compilation time as compared to state-of-the-art works. Dajiang Liu, Jiaxing Shang, Shouyi Yin |
HPCA | 1 |
| 2024 | DISC: Exploiting Data Parallelism of Non-Stencil Computations on CGRAs via Dynamic Iteration SchedulingabstractMemory partitioning is commonly used to enhance data parallelism of Coarse-grain reconfigurable arrays (CGRAs), typically targeting stencil computation with regular memory access patterns. However, many important workloads, such as linear algebra and signal processing, include non-stencil computation with irregular memory access patterns where memory partitioning is not feasible, leading to memory access conflicts and poor data parallelism. In this paper, we propose a Dynamic Iteration Scheduling CGRA (DISC) that can dynamically exploit data parallelism from non-stencil computations. Via dynamic scheduling on loop iterations, DISC can select conflict-free iterations from an iteration buffer for parallel data access, while holding data dependence. To further enhance the ability to find conflict-free iterations, dynamic data reuse is also introduced to reduce the number of memory references. The experimental results show that DISC can achieve 1.41× performance and 2.75× energy efficiency while consuming much less area and power overhead, as compared to dynamic-scheduling CGRA which supports dynamic operator scheduling. Yue Liang 0004, Di Mou, Dajiang Liu |
ICCAD | 3 |
| 2024 | BALQUE: Batch active learning by querying unstable examples with calibrated confidenceabstractActive learning alleviates labeling costs by selecting and labeling the most informative examples from an unlabeled pool. However, most existing active learning approaches estimate informativeness with uncalibrated confidence, resulting in unreliable informativeness estimation. These approaches generally ignored two significant issues caused by uncalibrated confidence methods. Firstly, the average uncalibrated confidence generated by modern neural networks is usually higher than the accuracy. Secondly, examples located near the decision boundaries are unstable during prediction when the target model updates parameters in the last several epochs, even throughout the training process. This phenomenon, caused by the forgetting characteristic of neural networks , has a significant impact on some specific models that estimate the informativeness by predicted probability vectors or pseudo labels. To address these issues, in this paper, we propose a novel active learning approach to reliably estimate informativeness with calibrated confidence. Specifically, we integrate the intermediate predictions for each unlabeled example , generated by the target model during the training process, to generate calibrated confidence. The calibrated confidence can capture a tendentious label from an indecisive subset of the class space. We show that the calibrated confidence with tendentiousness can maintain the ability of correct predictions. The empirical results demonstrate that our approach outperforms the state-of-the-art active learning methods on image classification tasks. Yincheng Han, Dajiang Liu, Jiaxing Shang, Linjiang Zheng, Wu Xie |
Pattern Recognit. | 2 |
| 2024 | IMTCN: An Interpretable Flight Safety Analysis and Prediction Model Based on Multi-Scale Temporal Convolutional NetworksabstractFlight safety is a key issue in the aviation industry. Recently, with the prevalence of flight data recording systems, some deep learning-based studies have been devoted to predicting safety incidents based on flight data. However, these studies, although they exhibit higher prediction accuracy, have largely neglected the interpretability analysis of safety incidents which is of great concern to airlines and pilots. To address this issue, we define flight safety prediction as a multiscale time series classification problem and propose an interpretable model named IMTCN to provide both accurate predictions and high interpretability of flight safety. First, multiple temporal convolutional networks (TCNs) are utilized to capture local representations and long effective histories from multivariate flight data. Because different flight parameters are collected with diverse sampling frequencies, multiple TCNs are used to handle these parameters separately. Then, we creatively adapt the class activation mapping (CAM) method, which has been used for interpretation in image classification, and combine it with the TCN to provide flight data interpretability. The established model can pinpoint key flight parameters and corresponding moments that contribute most to safety incidents. Experimental results on a real-world dataset with 37,943 Airbus A320 aircraft flights show that our model outperforms the baselines on the task of exceedance classification and prediction 2 seconds and 4 seconds in advance, and case studies demonstrate its superb interpretability for flight safety analysis. Xu Li 0014, Jiaxing Shang, Linjiang Zheng, Qixing Wang, Dajiang Liu, Fan Li 0020 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | SC-CGRA: An Energy-Efficient CGRA Using Stochastic ComputingabstractStochastic Computing (SC) offers a promising computing paradigm for low-power and cost-effective applications, with the added advantage of high error tolerance. In parallel, Coarse-Grained Reconfigurable Arrays (CGRA) prove to be a highly promising platform for domain-specific applications due to their combination of energy efficiency and flexibility. Intuitively, introducing SC to CGRA would significantly reinforce the strengths of both paradigms. However, existing SC-based architectures often encounter inherent computation errors, while the stochastic number generators employed in SC result in exponentially growing latency, which is deemed unacceptable in CGRA. In this work, we propose an SC-based CGRA by replacing the exact multiplication in traditional CGRA with an SC-based multiplication. To improve the accuracy of SC and shorten the latency of Stochastic Number Generators (SNG), we introduce the leading zero shifting and comparator truncation, while keeping the length of bitstream fixed. In addition, due to the flexible interconnections among PEs, we propose a quality scaling strategy that combines neighbor PEs to achieve high-accuracy operations without switching costs like power-gating. Compared to the state-of-the-art approximate computing design of CGRA, our proposed CGRA can averagely achieve a 65.3% reduction in output error while having a 21.2% reduction in energy consumption and a noteworthy 28.37% area savings. Di Mou, Dajiang Liu |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Towards High-Bandwidth-Utilization SpMV on FPGAs via Partial Vector DuplicationabstractSparse matrix-vector multiplication (SpMV) is widely used in many fields and usually dominates the execution time of a task. With large off-chip memory bandwidth, customizable on-chip resources and high-performance float-point operation, FPGA is a potential platform to accelerate SpMV tasks. However, as compressed data formats for SpMV usually introduce irregular memory access while it is also memory-intensive, implementing an SpMV accelerator on FPGA to achieve a high bandwidth utilization (BU) is a challenging work. Existing works either eliminate irregular memory access at the sacrifice of increasing data redundancy or try to locally reduce the port conflicts introduced by irregular memory access, leading to a limited BU improvement. To this end, this paper proposes a high-bandwidth-utilization SpMV accelerator on FPGAs using partial vector duplication, where read-conflict-free vector buffer, writing-conflict-free adder tree, and ping-pong-like accumulator registers are well elaborated. The FPGA implementation results show that the proposed design can achieve an average of 1.10x performance speedup compared to the state-of-the-art work. Dajiang Liu |
ASP-DAC | 2 |
| 2023 | Optimizing Data Reuse for CGRA Mapping Using Polyhedral-based Loop TransformationsabstractCoarse-Grained Reconfigurable Arrays (CGRA) can provide high energy efficiency while keeping moderate flexibility. With flexible connections, modern CGRAs are allowed to construct register chains on demand such that data reuse could be achieved. However, existing works put little effort into loop transformations for better data reuse. Therefore, this paper proposes an efficient loop transformation approach considering data reuse for the overall performance. Using reduced polyhedral formulation and Dynamical Programming (DP) based searching, loop structures could be thoroughly and efficiently explored for optimized solutions. The experimental results show that our approach can achieve 1.11-1.15 × speedup compared to the state-of-the-art approach. Liao Huang, Dajiang Liu |
DAC | 2 |
| 2023 | DARIC: A Data Reuse-Friendly CGRA for Parallel Data Access via Elastic FIFOsabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a promising architecture for data-intensive applications. For parallel data accesses, uniform memory partitioning is usually introduced to CGRA for better pipelining performance. However, uniform memory partitioning not only suffers from a local minimum, but also introduces non-negligible overhead for banking function, which may greatly degrade the performance of CGRA. To this end, this paper introduces non-uniform memory partitioning and proposes a data-reuse-friendly CGRA (DARIC). With well elaborated configurable bank groups cooperated with register chains, elastic FIFOs can be achieved for non-uniform memory partitioning. Based on the resource graph of DARIC, a mapping algorithm supporting path sharing is proposed. Finally, the experimental results show that DARIC can achieve 2.35 × throughput and 2.59 × energy efficiency while having even less area and power overhead, as compared to the state-of-the-art. Dajiang Liu, Di Mou, Yan Zhuang 0003, Jiaxing Shang, Shouyi Yin |
DAC | 1 |
| 2023 | Optimizing Memory Allocation for Multi-Subgraph Mapping on Spatial AcceleratorsabstractSpatial accelerators enable the pervasive use of energy-efficient solutions for computation-intensive applications. In the mapping of spatial accelerators, a large kernel is usually partitioned into multiple subgraphs for resource constraints, leading to more memory accesses and access conflicts. To minimize the access conflicts, existing works either neglect the interference of multiple subgraphs or pay little attention to data's life cycle along the execution order. To this end, this paper proposes an optimized memory allocation approach for multi-subgraph mapping on spatial accelerators by constructing an optimization problem using Integer Linear Programming (ILP). The experimental results demonstrate that our work can find conflict-free solutions for most kernels and achieve 1.15× speedup, as compared to the state-of-the-art approach. Decai Pan, Dajiang Liu, Xueliang Du |
SYSTOR | 3 |
| 2023 | HMSG: Heterogeneous graph neural network based on Metapath SubGraph learning
Mengya Guan, Xinjun Cai, Jiaxing Shang, Fei Hao 0001, Dajiang Liu, Xianlong Jiao, Wancheng Ni |
Knowl. Based Syst. | 5 |
| 2022 | CollaborateCas: Popularity Prediction of Information Cascades Based on Collaborative Graph Attention Networks
Xianren Zhang, Jiaxing Shang, Xueqi Jia, Dajiang Liu, Fei Hao 0001 |
DASFAA (1) | 4 |
| 2022 | Towards Energy-Efficient CGRAs via Stochastic ComputingabstractStochastic computing (SC) is a promising computing paradigm for low-power and low-cost applications with the added benefit of high error tolerance. Meanwhile, Coarse-Grained Re-configurable Architecture (CGRA) is also a promising platform for domain-specific applications for its combination of energy efficiency and flexibility. Intuitively, introducing SC to CGRA would synergistically reinforce the strengths of both paradigms. Accordingly, this paper proposes an SC-based CGRA by replacing the exact multiplication in traditional CGRA with an SC-based multiplication, where the problem of accuracy and latency are both improved using parallel stochastic sequence generators and leading zero shifters. In addition, with the flexible connections among PEs, the high-accuracy operation can be easily achieved by combing neighbor PEs without switching costs like power-gating. Compared to the state-of-the-art approximate computing design of CGRA, our proposed CGRA has 16% more energy reduction and 34% energy efficiency improvement while keeping high configuration flexibility. Jiaxing Shang, Dajiang Liu |
DATE | 4 |
| 2022 | RF-CGRA: A Routing-Friendly CGRA with Hierarchical Register ChainsabstractCGRAs are promising architectures to accelerate domain-specific applications as they combine high energy-efficiency and flexibility. With either isolated register files (RFs) or link-consuming distributed registers in each processing element (PE), existing CGRAs are all not friendly to data routing for data-flow graphs (DFGs) with a high edge/node ratio since there are many multi-cycle dependences. To this end, this paper proposes a Routing-Friendly CGRA (RF-CGRA) where hierarchical (intra-PE or inter-PE) register chains could be flexibly (wide range of chain length) and compactly (consuming fewer links among PEs) achieved for data routing, resulting in a new mapping problem that requires the improvement of a compiler. Experimental results show that RF-CGRA gets 1.19× performance and 1.14× energy efficiency of the state-of-the-art CGRA with single-cycle multi-hop connections (HyCUBE) while keeping a moderate compilation time. Dajiang Liu |
DATE | 3 |
| 2022 | Towards High-Quality CGRA Mapping with Graph Neural Networks and Reinforcement LearningabstractCoarse-Grained Reconfigurable Architectures (CGRA) is a promising solution to accelerate domain applications due to its good combination of energy-efficiency and flexibility. Loops, as computation-intensive parts of applications, are often mapped onto CGRA and modulo scheduling is commonly used to improve the execution performance. However, the actual performance using modulo scheduling is highly dependent on the mapping ability of the Data Dependency Graph (DDG) extracted from a loop. As existing approaches usually separate routing exploration of multi-cycle dependence from mapping for fast compilation, they may easily suffer from poor mapping quality. In this paper, we integrate the routing explorations into the mapping process and make it have more opportunities to find a globally optimized solution. Meanwhile, with a reduced resource graph defined, the searching space of the new mapping problem is not greatly increased. To efficiently solve the problem, we introduce graph neural network based reinforcement learning to predict a placement distribution over different resource nodes for all operations in a DDG. Using the routing connectivity as the reward signal, we optimize the parameters of neural network to find a valid mapping solution with a policy gradient method. Without much engineering and heuristic designing, our approach achieves 1.57× mapping quality, as compared to the state-of-the-art heuristic. Yan Zhuang 0003, Dajiang Liu |
ICCAD | 3 |
| 2022 | ConCas: Cascade Popularity Prediction Based on Topic-Aware Graph Contrastive Learning
Xianren Zhang, Jiaxing Shang, Dajiang Liu, Wu Xie, Baohua Qiang |
KSEM (1) | 4 |
| 2022 | Social Information Popularity Prediction based on Heterogeneous Diffusion Attention NetworkabstractInformation popularity prediction on social media platforms is a valuable and challenging issue.However, existing studies either neglect the correlation among different cascades, or lack a comprehensive consideration of user behavioral proximity and preference with respect to different messages.In this paper we propose a graph neural network-based framework named HeDAN (heterogeneous diffusion attention network), which comprehensively considers various factors affecting the information diffusion to predict the information popularity more accurately.Specifically, we first construct a heterogeneous diffusion graph with two types of nodes (user and message) and three types of relations (Friendship, Interaction, and Interest).Among them, Friendship reflects the strength of social relationship between users, Interaction reflects the behavioral proximity between users, and Interest reflects user preference to messages.Next, a graph neural network model with hierarchical attention mechanism is proposed to learn from these relations.Specifically, at the nodelevel, we utilize the graph attention network to learn the subgraph structure and generate the representations of nodes under each specific relationship.At the semantic-level, we distinguish the importance of different nodes in different relations via multihead self-attention mechanism.Extensive experimental results on three datasets show the superior performance of our proposed model over the state-of-the-arts. Xueqi Jia, Jiaxing Shang, Linjiang Zheng, Dajiang Liu |
SEKE | 4 |
| 2022 | IM2Vec: Representation learning-based preference maximization in geo-social networks
Ziwei Jin, Jiaxing Shang, Wancheng Ni, Liang Zhao 0004, Dajiang Liu, Baohua Qiang, Wu Xie, Geyong Min |
Inf. Sci. | 5 |
| 2022 | HeDAN: Heterogeneous diffusion attention network for popularity prediction of online content
Xueqi Jia, Jiaxing Shang, Dajiang Liu, Wancheng Ni |
Knowl. Based Syst. | 3 |
| 2021 | Polyhedral-based Pipelining of Imperfectly-Nested Loop for CGRAsabstractCoarse-Grained Reconfigurable Architectures (CGRAs) are promising architectures with high energy efficiency and flexibility. The computation-intensive portions of an application (e.g. loops) are often executed on CGRAs for acceleration and modulo scheduling is commonly used for loop mapping. However, for imperfectly-nested loops, existing methods don't fully explore the structure of the loops before performing modulo scheduling, resulting in poor execution performance. To tackle this problem, we propose a polyhedral-based pipelining approach for mapping imperfectly-nested loops on CGRA. By efficiently exploring the transformation space for imperfectly-nested loops using the polyhedral model and taking total execution time as an optimization metric, our approach could improve the execution performance greatly. On a$4\times 4$mesh-connected CGRA, the experimental results show that our approach can reduce the total execution time of nested loop by 50.1 % on average, as compared to the state-of-the-art techniques. Moreover, the compilation time is moderate in practice. Dajiang Liu, Xingyu Mo, Jiaxing Shang, Shouyi Yin |
ICCAD | 1 |
| 2021 | Dynamic Convolution Pruning Using Pooling Characteristic in Convolution Neural Networks
Dajiang Liu, Yongkang Xing |
ICONIP (5) | 2 |
| 2021 | BaCIM: Balanced Competitive Influence Maximization based on Blocked Reverse Influence SamplingabstractInfluence maximization, which seeks to find top influential individuals from a social network, has been extensively investigated in recent years. However, previous studies mainly focused on single diffusion or the diffusion of positive and negative messages, in which a competitor dominates the diffusion process. However, in a more realistic scenario, there is a level playing field between similar competitors. To cope with this, we introduce a new Balanced Competitive Influence Maximization (BaCIM) problem which considers the balance in information dissemination. We propose a Balanced Competitive Independent Cascade (BCIC) model to describe how two similar competitive products spread and compete in the same mobile social network. Given the competitor's seeding strategy, BaCIM aims to find a size-k seed set to maximize its own influence spread. We prove that the problem is NP-hard and the objective function is submodular, based on which a greedy algorithm is proposed with (1-1/e-ε) approximation guarantee. To handle large networks, we further propose a Blocked Reverse Influence Sampling algorithm named BRIS, in which we redesign the reverse influence sampling procedure to support the diffusion model. Experimental results on two location-based social networks and several large-scale real datasets validate effectiveness and efficiency of our algorithm. Wu Xie, Jiaxing Shang, Dajiang Liu, Baohua Qiang |
MDM | 4 |
| 2021 | Accepted Influence Maximization under Linear Threshold Model on Large-Scale Social NetworksabstractThe influence maximization (IM) problem, which aims to find$k$most influential individuals from a social network to maximize the influence spread, has been extensively studied. Existing works all rely on the assumption that influenced individuals will definitely try to propagate the information to their neighbors through social trust. However, in real-world this assumption can be over-simplistic since trust-levels among different individuals usually exhibit high diversity. As a result, an influenced individual may choose not to further propagate the information to his neighbors. Motivated by the observation, in this paper we propose a new accepted influence maximization (AIM) problem where the influenced individuals are further divided into two subgroups, i.e., accepted and active, where only active individuals will continue to propagate the information. We prove this problem is NP-hard and the objective function is submodular, based on which a greedy algorithm is proposed with ($1-1/e-\epsilon$) approximation guarantee. Considering the low computational efficiency of the greedy algorithm, we further propose a scalable path-based algorithm ALDAG. We conduct experiments on real datasets and the results demonstrate the effectiveness and efficiency of our method. Xiaojuan Yang, Jiaxing Shang, Linjiang Zheng, Dajiang Liu, Shu Fu, Baohua Qiang |
TrustCom | 4 |
| 2020 | A Deep Sequence-to-Sequence Method for Aircraft Landing Speed Prediction Based on QAR Data
Zongwei Kang, Jiaxing Shang, Yong Feng 0002, Linjiang Zheng, Dajiang Liu, Baohua Qiang |
WISE (2) | 5 |
| 2020 | FRWCAE: joint faster-RCNN and Wasserstein convolutional auto-encoder for instance retrieval
Yong Feng 0002, Dajiang Liu, Jiaxing Shang, Baohua Qiang |
Appl. Intell. | 3 |
| 2019 | Data-Flow Graph Mapping Optimization for CGRA With Deep Reinforcement LearningabstractCoarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their flexibility and energy efficiency. Data flow graphs (DFGs) are often mapped onto CGRAs for acceleration. The problem of DFG mapping is challenging due to the diverse structures from DFGs and constrained hardware from CGRAs. Consequently, it is difficult to find a valid and high quality solution simultaneously. Inspired from the great progress in deep reinforcement learning (RL) for AI problems, we consider building methods that learn to map DFGs onto spatially programmed CGRAs directly from experiences. We propose RLMap, a solution that formulates DFG mapping on CGRA as an agent in RL, which unifies placement, routing and processing element insertion by interchange actions of the agent. Experimental results show that RLMap performs comparably to state-of-the-art heuristics in mapping quality, adapts to different architecture, and converges quickly. Dajiang Liu, Shouyi Yin, Guojie Luo, Jiaxing Shang, Leibo Liu, Shaojun Wei, Yong Feng 0002, Shangbo Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2017 | Learning Convolutional Neural Networks for Data-Flow Graph Mapping on Spatial Programmable Architectures (Abstract Only)
Shouyi Yin, Dajiang Liu, Lifeng Sun, Xinhan Lin, Leibo Liu, Shaojun Wei |
FPGA | 2 |
| 2017 | DFGNet: Mapping dataflow graph onto CGRA by a deep learning approachabstractThe coarse-grained reconfigurable architecture (C-GRA) is a promising platform that provides both high performance and high power-efficiency. Dataflow graph (DFG) mapping is critical to tap the potentials of CGRAs. Inspired from the great progress made in tree search game using deep neural network, we proposed a frame work for learning convolutional neural network for mapping DFGs onto spatial programmable CGRAs. Considering the mapping process, we present a dual-input neural network capturing the features from both DFGs in applications and Process Element Array (PEA) in CGRA. In order to train the neural network, algorithms are designed to automatically generate a data set from PEA intermediate states of preprocessed DFG. Finally, experimental results demonstrate that our proposed mapping approach is competitive with state-of-the-art DFG mapping algorithms in performance while the compilation time is greatly reduced. Shouyi Yin, Dajiang Liu, Lifeng Sun, Leibo Liu, Shaojun Wei |
ISCAS | 2 |
| 2017 | Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank MemoryabstractCoarse-grained reconfigurable architecture (CGRA) is a promising architecture with high performance, high power-efficiency and attraction of flexibility. The computation-intensive parts of an application (e.g., loops) are often mapped on CGRA for acceleration. Due to the high parallel data access demands, the architecture with multi-bank memory is proposed to improve parallelism. For CGRA with multi-bank memory, a joint solution, which simultaneously considers the memory partitioning and modulo scheduling, is proposed to achieve a valid mapping with better performance. In this solution, the modulo scheduling and operator scheduling are used to achieve a valid loop mapping and a valid data placement without any memory access conflicts. By avoiding the pipelining stalls caused by conflicts, the performance of loop mapping is greatly improved. The experimental results on benchmarks of the Livermore, Polybench and Mediabench show that our approach can improve the performance of loops on CGRA to 1.89×, 1.49× and 1.37× compared with REGIMap, HTDM and REGIMap with memory partitioning, at cost of an acceptable increase in compilation time. Shouyi Yin, Xianqing Yao, Dajiang Liu, Jiangyuan Gu, Leibo Liu, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Joint Modulo Scheduling and Vdd Assignment for Loop Mapping on Dual- Vdd CGRAsabstractCoarse-grained reconfigurable architecture (CGRA) is becoming an increasingly attractive platform because of its high performance and power (or energy) efficiency. To reduce energy consumption, the dual-Vddtechnique has been employed in CGRAs, and the modulo scheduling technique is widely used to improve performance of applications. To achieve both high performance and energy-efficiency simultaneously, this paper formulates the solution as a biobjective optimization problem of energy consumption and initiation interval of loop pipelines on CGRAs, and proposes a joint modulo scheduling and dual-Vddassignment approach. The experimental results show that the proposed approach can bring a significant energy reduction of 24.8% and kernel energy efficiency acceleration of 1.41× on average, while the performance is maintained. Shouyi Yin, Jiangyuan Gu, Dajiang Liu, Leibo Liu, Shaojun Wei |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2016 | Improving Nested Loop Pipelining on Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architecture (CGRA) is a promising architecture with high performance, high power efficiency, and attraction of flexibility. The computation-intensive portions of applications, i.e., loops, are often implemented on CGRAs for acceleration. The loop pipelining techniques are usually used to exploit the parallelism of loops. However, for nested loops, the existing loop pipelining methods often result in poor hardware utilization and low execution performance. To tackle this problem, this paper makes three contributions: 1) we propose the use of affine transformation to facilitate nested loop pipelining; 2) based on polyhedral model, we present a precise and general formulation of the nested loop pipelining problem on a CGRA; and 3) using the insights from problem formulation, we design a joint affine transformation and multipipeline merging approach to improve the performance of nested loop on CGRA. The experimental results show that our approach can improve the performance of nested loops up to 35% on average, compared with the state-of-the-art techniques. Shouyi Yin, Dajiang Liu, Leibo Liu, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Memory-Aware Loop Mapping on Coarse-Grained Reconfigurable ArchitecturesabstractThe coarse-grained reconfigurable architectures (CGRAs) are a promising class of architectures with the advantages of high performance and high power efficiency. The compute-intensive parts of an application (e.g., loops) are often mapped onto the CGRA for acceleration. Due to the extra overhead of memory access and the limited communication bandwidth between the processing element (PE) array and local memory, previous works trying to solve the routing problem are mainly confined in the internal resources of PE arrays (e.g., PEs and registers). Inevitably, routing with PEs or registers will consume a lot of computational resources and cause the increase of the initiation interval. To solve this problem, this paper makes two contributions: 1) establishing a precise formulation for the CGRA mapping problem while using shared local data memory as a routing resource and 2) extracting an effective approach for mapping loops to CGRAs. The experimental results on loops of the SPEC2006, Livermore, and MiBench show that our approach (called MEMMap) can improve the performance of the kernels on CGRA up to 1.62×, 1.58×, 1.28×, and 1.23× compared with the edge-centric modulo scheduling, EPIMap, REGIMap, and force-directed map, respectively, with an acceptable increase in compilation time. Shouyi Yin, Xianqing Yao, Dajiang Liu, Leibo Liu, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Joint affine transformation and loop pipelining for mapping nested loop on CGRAs
Shouyi Yin, Dajiang Liu, Leibo Liu, Shaojun Wei, Yike Guo |
DATE | 2 |
| 2015 | Optimizing Spatial Mapping of Nested Loop for Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) have drawn increasing attention due to their flexibility and efficiency. Loops in applications are often mapped onto CGRAs for acceleration, and the mapping of loops onto CGRA is quite a challenging work due to the parallel execution paradigm and constrained hardware resource. To map loops onto CGRAs efficiently, it is important to transform loops into pieces that obey hardware resource constraints with less overhead (e.g., communication and configuration overhead). In this paper, we tackle this problem by establishing a performance optimization problem, including loop transformation and back- end placing and routing. A novel searching strategy is also designed to find the optimal result efficiently. Finally, we built a complete flow of mapping loop nests onto CGRA. Experiment results on most kernels of the Polybench show that our proposed approach can improve the performance of the kernels by 42% on average, as compared with the state-of-the-art methods. The runtime complexity of our approach is also acceptable. Dajiang Liu, Shouyi Yin, Leibo Liu, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Exploiting Outer Loop Parallelism of Nested Loop on Coarse-Grained Reconfigurable ArchitecturesabstractA coarse-grained reconfigurable architecture is a promising architecture with high power efficiency, which is typically composed of a host controller and a processing element array (PEA). Loops are often mapped onto PEAs for acceleration. In previous work, innermost loop is pipelined, and the the maximal number of concurrently executable operators (CEOs) in the kernel is limited by the inner loop. The loop body DFG of the input 2D nested loop with a inner loop carried dependence ([0,1]) and outer loop carried dependence ([1,1]). We would map this loop onto a 4×4 PEA with pipelining. We assume that the latency of executing one loop iteration is Lb, and the number of iterations involved at one cycle in the kernel phase of pipelining is Wk. As there is a inner loop dependence ([0,1]), the initiation interval (IIi) of inner loop pipelining could be minimized to 1 and we get Wk= 4. We also note that the angle α is contained by two sides in Figure 1(b), which could be written as follow: tan(α) = Wk/Lb = 1/IIi. Dajiang Liu, Shouyi Yin, Leibo Liu, Shaojun Wei |
FCCM | 1 |
| 2014 | RAREMETAL: fast and powerful meta-analysis for rare variantsabstractSUMMARY: RAREMETAL is a computationally efficient tool for meta-analysis of rare variants genotyped using sequencing or arrays. RAREMETAL facilitates analyses of individual studies, accommodates a variety of input file formats, handles related and unrelated individuals, executes both single variant and burden tests and performs conditional association analyses. AVAILABILITY AND IMPLEMENTATION: http://genome.sph.umich.edu/wiki/RAREMETAL for executables, source code, documentation and tutorial. Dajiang Liu, Xiaowei Zhan, Mary Kate Wing, Gonçalo R. Abecasis |
Bioinform. | 2 |
| 2013 | Polyhedral model based mapping optimization of loop nests for CGRAsabstractThe coarse-grained reconfigurable architecture (CGRA) is a promising platform that provides both high performance and high power-efficiency. The compute-intensive portions of an application (e.g. loops) are often mapped onto CGRA for acceleration. To optimize the mapping of loop nests to CGRA, this paper makes two contributions: i) Establishing a precise CGRA performance model and formulating the loop nests mapping as a nonlinear optimization problem based on polyhedral model, ii) Extracting an efficient heuristic loop transformation and mapping algorithm (PolyMAP) to improve mapping performance. Experiment results on most kernels of the PolyBench and real-life applications show that our proposed approach can improve the performance of the kernels by 21% on average, as compared to one of the best existing mapping algorithm, EPIMap. The runtime complexity of PolyMAP is also acceptable. Dajiang Liu, Shouyi Yin, Leibo Liu, Shaojun Wei |
DAC | 1 |
| 2013 | Affine transformations for communication and reconfiguration optimization of loops on CGRAsabstractA coarse-grained reconfigurable architecture (CGRA) is typically a hybrid architecture, which is composed of a reconfigurable processing unit (RPU) and a host microprocessor. Many compute-intensive applications (e.g., loop nests) are often mapped onto RPUs to speed up the execution of programs. However, communication volume and reconfiguration cost are two bottlenecks for the performance of RPUs. Therefore, loop transformations to break through the bottlenecks and tap the potentials of RPU would be of much significance. In this paper, an automatic loop transformation approach for RPUs is proposed, where the communication cost and reconfiguration cost are under a joint consideration. Experimental results show that our scheme can save up to 22.7% of execution time on average on partial differential equation (PDE) solver kernels compared with the approach just considering communication cost, and performs much better than the loop unrolling scheme on a great majority of loop kernels. Also, run-time complexity is acceptable for the practical cases. Dajiang Liu, Shouyi Yin, Leibo Liu, Shaojun Wei |
ISCAS | 1 |