VLDB 2026 Research / reviewers in the wild / expert
Yan Ding 0004
dblp:57/4533-4
· DBLP profile ↗
26ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0001-6956-9260ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anchored Maximum Communities over Large Directed Graphs
Xu Zhou 0001, Yan Ding 0004, Qing Liu 0002, Haoxian Xu, Kenli Li 0001 |
Proc. VLDB Endow. | 3 |
| 2026 | CAMVA: An Extension Architecture of CNN Accelerators for Multi-View AccelerationabstractAutopilot vehicles integrate additional views to capture comprehensive feature information, enhancing target detection accuracy. However, this integration imposes a higher computational burden on CNN accelerators in autopilot systems, potentially increasing the response latency of autopilot systems. It is a challenge for safety-critical autopilot systems. Traditionally, researchers have addressed this challenge by employing complex designs and processes to enhance the arithmetic capabilities of chips. In contrast, the paper proposes a simple technology that is an extension architecture of CNN accelerators for multiview acceleration (CAMVA). The extension architecture utilizes the characteristic of multi-view approximation to enhance the sparsity of input features and dynamically reuse the approximated feature extraction results, accelerating CNN accelerators. This paper uses multi-view autonomous driving datasets (KITTI and nuScenes) to create two tasks, and evaluates the impact of CAMVA on 2D and 3D object detection networks by simulating their data flows. Results of 2D experiments show that for 2D object detection, the accuracy of the KITTI task decreases by 2.29%~4.26%, while that of the nuScenes task decreases by 0.37%~1.11%. For 3D object detection, the accuracy of the KITTI task decreases by 1.97%~-0.22%. Then, the paper utilizes the pruning operation of CAMVA to create two subtasks, which are subsets of the KITTI with sparsity of 9.6% and 19.8%, respectively. It evaluates the performance, energy consumption, and RTL of CAMVA at the hardware architecture level based on the two subtasks. The results show that CAMVA enhances the performance of CNN accelerator by 1.04~1.08× and reduces energy consumption by 4.23%~8.01% in the sparsity 9.6% subtask; additionally, it improves performance by 1.13~1.34× and decreases energy consumption by 14.13%~25.97% in the sparsity 19.8% subtask. CAMVA increases CNN accelerator’;s area by a mere 0.75%. Yan Ding 0004, Chubo Liu, Keqin Li 0001 |
IEEE Trans. Computers | 2 |
| 2026 | AEIS: A New Energy Efficiency Improvement Scheme for MLC STT-MRAMabstractSpin Transfer Torque-Magnetic Random Access Memory (STT-MRAM), as a new non-volatile memory technology with lower leakage power and higher density, is widely considered to be a new generation of memory technology that may replace SRAM in the cache. STT-MRAM is divided into Single-Level Cell (SLC) STT-MRAM and Multi-Level Cell (MLC) STT-MRAM. Compared with SLC STT-MRAM, MLC STT-MRAM has further improved its storage density. However, MLC STT-MRAM has a high energy consumption and write latency due to its unique two-step state transitions (TTs) issue. State-of-the-art approaches mitigate this issue by eliminating TTs with expansion coding methods. Unfortunately, they focus more on eliminating TTs and have limited improvement in reducing energy consumption. To this end, we propose a new scheme, AEIS, which further reduces energy consumption while eliminating TTs. Our work begins with exploring the general rules of (M,N)-based expansion coding methods that eliminate TTs. Based on the discovered rules, the minimum energy coding method is found. To further improve energy efficiency, we segment the cache lines according to the data pattern. We only apply the expansion coding to those flipping segments to reduce the expansion coding overhead. The evaluation results show that AEIS can eliminate TTs in MLC STT-MRAM, reduce energy consumption by 28.5%, and increase the lifetime by 24.8%, while the total number of bits used for the cache only increases by 5.7%. Huizhang Luo, Yan Ding 0004, Chubo Liu, Wenchao Zhao, Kenli Li 0001 |
IEEE Trans. Computers | 3 |
| 2026 | Automated Screening Network for Fetal Closed Spina Bifida With Semantic Enhancement and Projected AttentionabstractClosed spina bifida is a high-incidence developmental disorder among rare fetal diseases. Its signs in ultrasound imaging are subtle, making it prone to misdiagnosis and heavily reliant on sonographers' experience. Therefore, we propose a novel semantic enhancement framework incorporating projected attention for the automated screening of closed spina bifida through precise landmark detection. In this method, we utilize a multi-granularity deep supervision and voting mechanism to generate point-specific features and reconstruct saliency maps for each landmark, effectively reducing interference from homogeneous high-echogenic noise in ultrasound images while preserving rich semantic information. Additionally, a coordinate attention projection module is designed to convert the 2D landmark probability maps into one-dimensional vectors, ensuring low computational complexity along with precise coordinate regression. The clinical application potential of this intelligent system is significant, as it facilitates automated fetal spine counting and anatomical measurement, enabling early warnings of diseases based on identified anomalies. Extensive experiments comparing our method with advanced baselines on an in-house dataset and two public datasets demonstrate its clear advantage in computational complexity and accuracy. Yan Ding 0004, Ningbo Zhu, Chunlian Wang, Shengli Li 0001, Kenli Li 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | Enhancing Large Language Models Reasoning via Multi-Path Optimization on Knowledge Graph
Jiyong Liao, Chubo Liu, Yan Ding 0004, Haotian Wang 0006, Zhuo Tang, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Property-Induced Partitioning for Graph Pattern Queries on Distributed RDF SystemsabstractGraph pattern queries (GPQ) over RDF graphs extend basic graph patterns to support variable-length paths (VLP), thereby enabling complex knowledge retrieval and navigation. Generally, variable-length paths describe the reachability between two vertices via a given property within a specified range. With the increasing scale of RDF graphs, it is necessary to design a partitioning method to achieve efficient distributed queries. Although many partitioning strategies have been proposed for large RDF graphs, most existing methods result in numerous inter-partition joins when processing GPQs, which impacts query performance. In this paper, we formulate a new partitioning problem, MaxLocJoin, aims to minimize inter-partition joins during distributed GPQ processing. For MaxLocJoin, we propose a partitioning framework (PIP) based on property-induced subgraphs, which consist of edges with a specific set of properties. The framework first finds a locally joinable property set using a cost-driven algorithm, LJPS, where the cost depends on the sizes of weakly connected components within its property-induced subgraphs. Subsequently, the graph is partitioned according to the weakly connected components. The framework can achieve two key objectives: first, it enables complete local processing of all variable-length path queries (eliminating inter-partition joins); second, it can minimize the number of inter-partition joins required for traditional graph pattern queries. Moreover, we identify two types of independently executable queries (IEQ): the locally joinable IEQ and the single-property IEQ. After that, a query decomposition algorithm is designed to transform all GPQ into one of them for independent execution in distributed environments. In experiments, we implement two prototype systems based on Jena and Virtuoso, and evaluate them over both real and synthetic RDF graphs. The results show that MaxLocJoin achieves performance improvements from 2.8x to 10.7x over existing methods. Shidan Ma, Yan Ding 0004, Xu Zhou 0001, Peng Peng 0001, Youhuan Li, Zhibang Yang, Kenli Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | Win-Win Approaches for Cross Dynamic Task Assignment in Spatial CrowdsourcingabstractSpatial crowdsourcing (SC) is becoming increasingly popular recently. As a critical issue in SC, task assignment currently faces challenges due to the imbalanced spatiotemporal distribution of tasks. Hence, many related studies and applications focusing on cross-platform task allocation in SC have emerged. Existing work primarily focuses on the maximization of total revenue for inner platform in cross task assignment. In this work, we formulate a SC problem called Cross Dynamic Task Assignment (CDTA) to maximize the overall utility and propose improved solutions aiming at creating a win-win situation for inner platform, task requesters, and outer workers. We first design a hybrid batch processing framework and a novel cross-platform incentive mechanism. Then, with the purpose of allocating tasks to both inner and outer workers, we present a KM-based algorithm that gets the accurate assignment result in each batch and a density-aware greedy algorithm with high efficiency. To maximize the revenue of inner platform and outer workers simultaneously, we model the competition among outer workers as a potential game that is shown to have at least one pure Nash equilibrium and develop a game-theoretic method. Additionally, a simulated annealing-based improved algorithm is proposed to avoid falling into local optima. Last but not least, since random thresholds lead to unstable results when picking tasks that are preferentially assigned to inner workers, we devise an adaptive threshold selection algorithm based on multi-armed bandit to further improve the overall utility. Extensive experiments demonstrate the effectiveness and efficiency of our proposed algorithms on both real and synthetic datasets. Tianyue Ren, Zhibang Yang, Yan Ding 0004, Xu Zhou 0001, Kenli Li 0001, Yunjun Gao, Keqin Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2026 | Polarity-Aware and Adaptive Sparse Aggregation for Implicit Heterophilic Graph ClassificationabstractGraph-structured data appears in domains such as molecular analysis, social networks, and program optimization, where graphs often exhibit implicit heterogeneity, as nodes may look homogeneous in type yet differ significantly in semantics or functionality. Graph Neural Networks (GNNs), while powerful on homophilic graphs, tend to degrade in such settings due to polarity confusion, over-smoothing, and inefficiency caused by dense propagation. We propose a polarity-aware framework for graph classification that addresses these challenges through adaptive directional sparse aggregation. The framework introduces a polarity-aware propagation mechanism that adaptively reinforces or inverts neighbor signals, mitigating contamination under heterophily. A polarity-guided sparse aggregation operator further alleviates over-smoothing, improves scalability by constraining redundant connections, and condenses information flow into more effective representations, while maintaining unbiased estimation with controlled variance. We provide theoretical analyses that characterize the computational complexity, stability properties, and expressive behavior of signed directional aggregation, offering theoretical insights into its computational, stability, and expressive properties. Extensive experiments on molecular and social graph benchmarks with implicit heterophily demonstrate consistent improvements in graph classification accuracy and efficiency. Our method achieves a 2.36% improvement when compared with the strongest baseline on each dataset. In addition, it improves accuracy by 4.53% on average on program optimization strategy recognition tasks, reaching 80.12% overall. Haotian Wang 0006, Yan Ding 0004, Wangdong Yang, Zhuo Tang, Chubo Liu, Kenli Li 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2026 | On Imbalance in Case Types: Evaluating and Enhancing PLMs for Criminal Court View GenerationabstractThe criminal court view generation (CCVG) task aims to produce succinct and coherent summaries of fact descriptions, providing interpretable opinions for verdicts. Traditional text generation evaluation metrics, such as ROUGE, BLEU, and BERTSCORE, are extensively employed for this task and measure performance by averaging the assessment scores of all samples within the test set. However, these sample-averaged metrics encounter two primary dilemmas: 1) they fail to fairly assess overall evaluation scores across different case types and 2) they overlook the measurement of the degree of performance imbalance between case types. To fill this research gap, we propose two novel case-type-oriented evaluation metrics: Case-type-oriented Text Generation (CTG) and Case-type-oriented Imbalance Performance (CIP). First, CTG mitigates the unfair assessment among different case types by assigning equal weight to each type. Second, CIP evaluates performance imbalance by measuring the distance between the performance of each case type and the overall performance. We provide three theorems to elucidate the properties of CIP, demonstrating that CIP can effectively identify the extent to which a CCVG model achieves balanced generation performance across different case types. Furthermore, we propose an embarrassingly simple and effective charge-guided encoder-decoder (CGED) framework to enhance performance fairly across different case types in encoder-decoder pretrained language models (PLMs). Code is available at https://yuquanle.github.io/Case-type-oriented-metrics-homepage/. Yuquan Le, Yan Ding 0004, Chng Eng Siong, Kenli Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2026 | EBFL: An Efficient Blockchain Framework for Federated Learning ServicesabstractFederated Learning (FL) has emerged as a key framework to deliver AI services, recognized for its capability to construct global models while ensuring individual data. Nevertheless, FL heavily relies on a central server, which introduces significant challenges for participants to collaborate effectively and substantially limits the scalability of FL. Blockchain-based FL (BFL) offers a promising solution by replacing the central server with a decentralized blockchain system, thereby establishing a secure and trustworthy environment for FL. However, current BFL approaches face challenges in balancing high computational overhead, consistency, and security. In view of this, this paper introduces EBFL, an efficient blockchain framework for FL services. EBFL incorporates both asynchronous and synchronous advantages. A DAG-based (Directed Acyclic Graph) asynchronous computation enhances computational efficiency by mitigating delays caused by slow devices and reducing unnecessary waiting due to frequent synchronized consensus. Simultaneously, a periodic synchronized consensus mechanism is introduced during asynchronous training to ensure consistency, thereby improving security and model accuracy. Additionally, taking into account the unique characteristics of FL, we have designed a series of operations tailored for EBFL to further enhance the performance. Experimental results demonstrate that, compared to traditional synchronous BFL (TBFL) approaches, EBFL achieved a maximum speedup of up to 2.38× while retaining 92% of their accuracy. Subsequently, in-depth analytical experiments show that EBFL excels in both convergence speed and security, thereby confirming its potential to balance computational efficiency, consistency, and security. Ze Yin, Haotian Wang 0006, Chubo Liu, Yan Ding 0004, Keqin Li 0001, Kenli Li 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2025 | SSpMV: A Sparsity-aware SpMV Framework Empowered by Multimodal Machine LearningabstractSparse Matrix-Vector Multiplication (SpMV) is an essential sparse operation in scientific computing and artificial intelligence. Efficiently adapting SpMV algorithms to diverse matrices and architectures requires a framework capable of accurately recognizing sparse patterns and selecting the optimal implementation. In this work, we introduce Sparsity-aware SpMV (SSpMV), a framework that integrates expert-designed features with multimodal representations to adaptively predict the best-performing algorithm and parameters. For this purpose, we design a multimodal neural network called MM-Adapter, to capture diverse modalities to represent the computational features of SpMV. Experimental results demonstrate that MMAdapter achieves the highest accuracy of $81.05 \%$, outperforming existing SpMV prediction models. Furthermore, SSpMV consistently delivers substantial performance improvements over state-of-the-art sparse libraries across various multi-core platforms. Shengle Lin, Chubo Liu, Yan Ding 0004, Joey Tianyi Zhou, Kenli Li 0001, Wangdong Yang |
DAC | 3 |
| 2025 | A Post-Implementation Performance Prediction Method with HLS Optimization DirectivesabstractHigh-Level Synthesis (HLS) offers various optimization directives that enable designers to flexibly adjust hardware microarchitecture. However, existing HLS performance prediction methods typically rely on control data flow graphs generated (CDFG) from original HLS C/C++, which struggle to capture the complex interactions between directives and the resulting hardware resource reuse issues. To address these issues, this paper proposes a post-implementation performance prediction method tailored for directive-optimized circuits design, which utilizes a Graph Builder to integrate directive optimization and resource reuse information into the graph representation. In addition, the performance prediction model, which integrates a TransformerConv-based graph neural network (GNN) and an aggregation pool module, effectively captures key features related to post-implementation performance. Experimental results show that our method can reduce the prediction error of critical path delay (CP), power, and resource utilization to 3.87% $\sim$ 8.08%, significantly outperforming existing state-of-the-art methods. It also demonstrates excellent generalization on unseen kernels, providing a more effective and accurate performance prediction tool for HLS. Yan Ding 0004, Kenli Li 0001, Chubo Liu |
DAC | 2 |
| 2025 | MoNeRF: Deformable Neural Rendering for Talking Heads via Latent Motion NavigationabstractAbstract Novel view synthesis for talking heads presents significant challenges due to the complex and diverse motion transformations involved. Conventional methods often resort to reliance on structure priors, like facial templates, to warp observed images into a canonical space conducive to rendering. However, the incorporation of such priors introduces a trade‐off‐while aiding in synthesis, they concurrently amplify model complexity, limiting generalizability to other deformable scenes. Departing from this paradigm, we introduce a pioneering solution: the motion‐conditioned neural radiance field, MoNeRF, designed to model talking heads through latent motion navigation. At the core of MoNeRF lies a novel approach utilizing a compact set of latent codes to represent orthogonal motion directions. This innovative strategy empowers MoNeRF to efficiently capture and depict intricate scene motion by linearly combining these latent codes. In an extended capability, MoNeRF facilitates motion control through latent code adjustments, supports view transfer based on reference videos, and seamlessly extends its applicability to model human bodies without necessitating structural modifications. Rigorous quantitative and qualitative experiments unequivocally demonstrate MoNeRF's superior performance compared to state‐of‐the‐art methods in talking head synthesis. We will release the source code upon publication. Yan Ding 0004, Ruihui Li, Zhuo Tang, Kenli Li 0001 |
Comput. Graph. Forum | 2 |
| 2025 | BEAST-GNN: A United Bit Sparsity-Aware Accelerator for Graph Neural NetworksabstractGraph Neural Networks (GNNs) excel in processing graph-structured data, making them attractive and promising for tasks such as recommender systems and traffic forecasting. However, GNNs’ irregular computational patterns limit their ability to achieve low latency and high energy efficiency, particularly in edge computing environments. Current GNN accelerators predominantly focus on value sparsity, underutilizing the potential performance gains from bit-level sparsity. However, applying existing bit-serial accelerators to GNNs presents several challenges. These challenges arise from GNNs’ more complex data flow compared to conventional neural networks, as well as difficulties in data localization and load balancing with irregular graph data.To address these challenges, we propose BEAST-GNN, a bit-serial GNN accelerator that fully exploits bit-level sparsity. BEAST-GNN introduces streamlined sparse-dense bit matrix multiplication for optimized data flow, a column-overlapped graph partitioning method to enhance data locality by reducing memory access inefficiencies, and a sparse bit-counting strategy to ensure balanced workload distribution across processing elements (PEs). Compared to state-of-the-art accelerators, including HyGCN, GCNAX, Laconic, GROW, I-GCN, SGCN, and MEGA, BEAST-GNN achieves speedups of 21.7 ×, 6.4×, 10.5×, 3.7×, 4.0×, 3.3×, and 1.4× respectively, while also reducing DRAM access by 36.3×, 7.9×, 6.6×, 3.9×, 5.38×, 3.37×, and 1.44×. Additionally, BEAST-GNN consumes only 4.8%, 12.4%, 19.6%, 27.7%, 17.0%, 26.5%, and 82.8% of the energy required by these architectures. Yunzhen Luo, Yan Ding 0004, Zhuo Tang, Keqin Li 0001, Kenli Li 0001, Chubo Liu |
IEEE Trans. Computers | 2 |
| 2025 | A Context-Awareness and Hardware-Friendly Sparse Matrix Multiplication Kernel for CNN Inference AccelerationabstractSparsification technology is crucial for deploying convolutional neural networks in resource-constrained environments. However, the efficiency of sparse models is hampered by irregular memory access patterns in sparse matrix multiplication kernels. Hardware-level support for 2:4 granularity in sparse tensor cores presents an opportunity for designing efficient sparse matrix multiplication kernels. Existing approaches often involve adjusting sparse structures or secondary sparsification, introducing additional computational errors. To tackle this challenge, we introduce a flexible 2:4 structured adaptive sparse matrix multiplication (FS-AMM) method, a hardware-friendly sparse matrix multiplication kernel that leverages model context to accelerate convolutional neural networks. First, we propose a model context-aware matrix pre-processing method that employs heuristic algorithms to estimate a loss of accuracy due to weight sparsity at each layer. Second, we design a hardware-friendly sparse storage format that combines 2:4 sparse and dense storage formats, enabling more versatile sparsity ratio selection. Third, we implement efficient matrix multiplication kernels to optimize GPU utilization. Finally, experimental results on A100 GPUs show that our method effectively utilizes the sparse tensor kernel and obtains an average 3.09 times speedup ratio compared to other sparse methods while maintaining a high accuracy. Haotian Wang 0006, Yan Ding 0004, Weichen Liu 0001, Chubo Liu, Wangdong Yang, Kenli Li 0001 |
IEEE Trans. Computers | 2 |
| 2025 | MixSSC: Forward-Backward Mixture for Vision-Based 3D Semantic Scene CompletionabstractVision-based semantic scene completion task aims to predict dense geometric and semantic 3D scene representations from 2D images. However, 3D modeling from a single view is an ill-posed problem, limited by the field of view and occlusion problems caused by image input. Moreover, existing methods tend to produce erroneous scene hallucinations and overly smooth boundary segmentation due to a lack of information. To address this problem, we propose MixSSC, which mixes the sparsity of forward projection with the denseness of depth-prior backward projection. The aim is to use sparse features to fill information-poor regions and dense features to enhance visible regions. Specifically, we develop the forward-backward mixture module, which enables the generation of scene mixture voxel representation by leveraging the benefits of both forward and backward projection. Subsequently, we design the semantic-spatial fusion module, which utilizes a coarse-to-fine approach to process mixture voxel features at the semantic-spatial level. Extensive experimental results on the SemanticKITTI, SSCBench-KITTI-360 and nuScenes datasets demonstrate the superiority of MixSSC. Our code is available on https://github.com/willemeng/MixSSC. Meng Wang 0040, Yan Ding 0004, Yunchuan Qin, Ruihui Li, Zhuo Tang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | A Completing Missing Pedestrian Trajectories Method Driven by Prior-Posterior Knowledge and Interactive InformationabstractPedestrian historical trajectory completion significantly bolsters the predictive accuracy of models. However, traditional statistical models such as Hidden Markov Models (HMM), which focus solely on individual pedestrian trajectories, often fall short in terms of generalization. Conversely, data-driven deep learning approaches demand extensive and meticulous data annotation as well as large datasets. Additionally, leveraging sequential historical data and uncovering the correlation between neighboring pedestrians during absences presents a significant challenge. To address these issues, we introduce a novel trajectory completion method that harnesses prior-posterior knowledge and interactive information, termed CMPT. Our approach commences with the design of a Neighbor Pedestrian Selection module (NPS), adept at identifying neighboring pedestrians through a composite scoring system that evaluates feature similarity and proximity. Subsequently, we employ a Top-Graph Attention Network (T-GAT) to extract multiple correlation sets between preceding and succeeding moments within the scenario. These correlations are then fed into the Markov-Inverse Recovery module (MR), which utilizes prior and posterior insights to flesh out the neighbor influence at the unobserved intervals. Culminating in the Trajectory Reconstruction module (TR), we integrate the completed neighbor influence data with the historical trajectory of the missing pedestrian to finalize the missing trajectory reconstruction. Empirical evidence from our experiments indicates that the Final Distance Error (FDE) of the trajectories completed by CMPT is a commendable 0.30. The source code for CMPT is available from https://github.com/ZYueliang/CMPT-Net. Mingxing Duan, Xinyue Zheng, Huilong Pi, Yan Ding 0004, Zhuo Tang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | A Real-time Execution System of Multimodal Transformer through PIM-GPU CollaborationabstractMultimodal transformer excels in various applications, but faces great challenges such as high memory consumption and limited data reuse that hinder real-time performance. To address these issues, we propose a processing-in-memory (PIM)-GPU collaboration oriented compiler to accelerate the multimodal transformers. The PIM-GPU collaboration adapts well to multimodal transformers and significantly accelerates model inference. In addition, we introduce a tailored PIM allocation algorithm for variable-length inputs to further improve computation efficiency. Experimental results show that our scheme can achieve an average 15x end-to-end speedup. Shengyi Ji, Chubo Liu, Yan Ding 0004, Qing Liao 0001, Zhuo Tang |
DAC | 3 |
| 2024 | X-EDF: An Efficient Defensive Deception Framework against Reconnaissance AttacksabstractDeception techniques are increasingly recognized as trans-formative in the realm of cyber defense. With the advent of sophisticated, large-scale scanning technologies such as ZMap, attackers can swiftly pinpoint active and vulnerable ports on edge nodes. Given the diversity of these nodes, a versatile security tool adaptable to various deployment environments is essential. Moreover, edge nodes often encounter performance constraints, necessitating a defense strategy that balances cost-effectiveness for defenders. In response to these challenges, we introduce the X-EDF: an eXpress Data Path (XDP)-based Efficient Defensive De-ception Framework. This framework facilitates an efficient and lightweight deceptive defense leveraging XDP technology. The X-EDF can efficiently respond to attackers' scanning requests with deceptive messages before these requests enter the protocol stack, thus achieving deception defense at a minimal cost. We have validated the effectiveness of our defense strategy through game-theoretic proofs and real-world network deployments. Zhihang Zhang, Chenlin Huang, Yan Ding 0004, Jinzhu Kong, Qing Liao 0001, Pan Dong, Haifang Zhou |
MSN | 3 |
| 2024 | CBANA: A Lightweight, Efficient, and Flexible Cache Behavior Analysis FrameworkabstractCache miss analysis has become one of the most important things to improve the execution performance of a program. Generally, the approaches for analyzing cache misses can be categorized into dynamic analysis and static analysis. The former collects sampling statistics during program execution but is limited to specialized hardware support and incurs expensive execution overhead. The latter avoids the limitations but faces two challenges: inaccurate execution path prediction and inefficient analysis resulted by the explosion of the program state graph. To overcome these challenges, we propose CBANA, an LLVM- and process address space-based lightweight, efficient, and flexible cache behavior analysis framework. CBANA significantly improves the prediction accuracy of the execution path with awareness of inputs. To improve analysis efficiency and utilize the program preprocessing, CBANA refactors loop structures to reduce search space and dynamically splices intermediate results to reduce unnecessary or redundant computations. CBANA also supports configurable hardware parameter settings, and decouples the module of cache replacement policy from other modules. Thus, its flexibility is established. We evaluate CBANA by using the popular open benchmark PolyBench, graph workloads, and our synthetic workloads with good and poor data locality. Compared with the popular dynamic cache analysis tools Perf and Valgrind, the cache miss gap is less than 3.79% and 2.74% respectively with over ten thousand data accesses for the synthetic workloads, and the time reduction is up to 92.38% and 97.51% for the multiple-path workloads. Compared with the popular static cache analysis tool Heptane, CBANA achieves a time reduction of 97.71% while ensuring accuracy at the same time. Qilin Hu, Yan Ding 0004, Chubo Liu, Keqin Li 0001, Kenli Li 0001, Albert Y. Zomaya |
IEEE Trans. Computers | 2 |
| 2023 | HAIMA: A Hybrid SRAM and DRAM Accelerator-in-Memory Architecture for TransformerabstractThrough the attention mechanism, Transformer-based large-scale deep neural networks (LSDNNs) have demonstrated remarkable achievements in artificial intelligence applications such as natural language processing and computer vision. The matrix-matrix multiplication operation (MMMO) in Transformer makes data movement dominate the inference overhead over computation. A solution for efficient data movement during Transformer inference is to embed arithmetic logic units (ALUs) into the memory array, hence an accelerator-in-memory architecture (AIMA). Existing work along this direction has not considered the heterogeneity of parallelism and resource requirements among Transformer layers. This increases the inference latency and lowers the resource utilization, which is critical for the embedded systems domain. To this end, we propose HAIMA, a hybrid AIMA and the parallel dataflow for Transformer, which exploit the cooperation between SRAM and DRAM to accelerate different MMMOs. Compared to the state-of-the-art Newton and TransPIM, our proposed hardware-software co-design achieves 1.4x-1.5x speedup, and solves the problem of resource under-utilization when DRAM-based AIMA performs the light-weight MMMOs. Yan Ding 0004, Chubo Liu, Mingxing Duan, Wanli Chang 0001, Keqin Li 0001, Kenli Li 0001 |
DAC | 1 |
| 2023 | Budget-Constrained Service Allocation Optimization for Mobile Edge ComputingabstractThe service resource allocation strategy optimization problem has always been a hot issue in mobile edge computing (MEC). In this paper, the problem is formulated as a long-term quality of service (QoS) improvement problem while satisfying the budget of MEC service provider (MSP). Since it is very unrealistic to accurately obtain the request information of user equipments (UEs) over a long time, we first transform the original problem into a series of real-time linear programing sub-problems by using Lyapunov optimization method, and propose a centralized algorithm to determine the resource allocation strategies. However, since the sub-problems are still NP-hard problems, it is a huge challenge to determine the strategies for all UEs with the centralized algorithm in a large scale MEC environment. Thus, we then formulate the sub-problems as an N players non-cooperative game, prove that there exists a Nash equilibrium, and develop two iterative algorithms to find the Nash equilibrium while determining the strategies. Experimental results show that the algorithms can take into account QoS and budget of MSP at the same time, and perform better compared to five other common schemes. Yan Ding 0004, Kenli Li 0001, Chubo Liu, Zhuo Tang, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2022 | A Potential Game Theoretic Approach to Computation Offloading Strategy Optimization in End-Edge-Cloud ComputingabstractIntegrating user ends (UEs), edge servers (ESs), and the cloud into end-edge-cloud computing (EECC) can enhance the utilization of resources and improve quality of experience (QoE). However, the performance of EECC is significantly affected by its architecture. In this article, we classify EECC into two computing architectures types according to the visibility and accessibility of the cloud to UEs, i.e., hierarchical end-edge-cloud computing (Hi-EECC) and horizontal end-edge-cloud computing (Ho-EECC). In Hi-EECC, UEs can offload their tasks only to ESs. When the resources of ESs are exhausted, the ESs request the cloud to provide resources to UEs. In Ho-EECC, UEs can offload their tasks directly to ESs and the cloud. In this article, we construct a potential game for the EECC environment, in which each UE selfishly minimizes its payoff, study the computation offloading strategy optimization problems, and develop two potential game-based algorithms in Hi-EECC and Ho-EECC. Extensive experiments with real-world data are conducted to demonstrate the performance of the proposed algorithms. Moreover, the scalability and applicability of the two computing architectures are comprehensively analyzed. The conclusions of our work can provide useful suggestions for choosing specific computing architectures under different application environments to improve the performance of EECC and QoE. Yan Ding 0004, Kenli Li 0001, Chubo Liu, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Short- and long-term cost and performance optimization for mobile user equipments
Yan Ding 0004, Kenli Li 0001, Chubo Liu, Zhuo Tang, Keqin Li 0001 |
J. Parallel Distributed Comput. | 1 |
| 2020 | A Code-Oriented Partitioning Computation Offloading Strategy for Multiple Users and Multiple Mobile Edge Computing ServersabstractIn this article, we investigate code-oriented partitioning computation offloading strategy for multiple user equipments (UEs) and multiple mobile edge computing servers with limited resources (i.e., limited computing power and waiting task queues with finite capacity). This article aims to develop an offloading strategy to decide the execution location, CPU frequency, and transmission power for UE while minimizing the execution overhead (i.e., a weighted sum of energy consumption and computational time) of UE's applications, which is an NP-hard problem. To achieve the objective, first, we transform the problem into a convex optimization problem and find the optimal solution. Second, we propose a decentralized computation offloading strategy (DCOS) algorithm for UE, and define a dictionary data structure for recording the strategy of the UE to reduce the algorithm complexity. Finally, the effectiveness of DCOS, and the impact of various key parameters on the strategy and overhead are demonstrated by simulation experiments. Yan Ding 0004, Chubo Liu, Xu Zhou 0001, Zhao Liu 0006, Zhuo Tang |
IEEE Trans. Ind. Informatics | 1 |
| 2019 | A keyword-based combination approach for detecting phishing webpages
Yan Ding 0004, Nurbol Luktarhan, Keqin Li 0001, Wushour Slamu |
Comput. Secur. | 1 |