EDBT 2026 Demo / reviewers in the wild / expert
Jiawen Sun
dblp:176/5001
· DBLP profile ↗
12ranked-venue papers
6as first author
5since 2021 · last 2026
0000-0002-1458-3484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Memory systems · 39% Reconfigurable computing and FPGAs · 20% High-performance computing · 20% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 50% Algorithms and data structures · 50% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems
memory access latency |
0.5 | 1 | 2021 | Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-Design · HPCA 2021 |
Memory systems › cache
prefetching |
0.5 | 1 | 2021 | Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-Design · HPCA 2021 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.5 | 1 | 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph Analytics · DAC 2021 |
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication |
0.5 | 1 | 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph Analytics · DAC 2021 |
Graph algorithms and graph theory
graph partitioning |
0.4 | 1 | 2019 | VEBO: a vertex- and edge-balanced ordering heuristic to load balance parallel graph processing · PPoPP 2019 |
Algorithms and data structures
load balancing |
0.4 | 1 | 2019 | VEBO: a vertex- and edge-balanced ordering heuristic to load balance parallel graph processing · PPoPP 2019 |
Parallel and multicore computing
graph processing |
0.1 | 1 | 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph Analytics · DAC 2021 |
Hardware accelerators and domain-specific architectures
graph processing accelerator |
0.1 | 1 | 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph Analytics · DAC 2021 |
Electronic design automation
hardware/software co-design |
0.1 | 1 | 2021 | Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-Design · HPCA 2021 |
Parallel and multicore computing › parallel graph algorithms
shared-memory parallel graph algorithm |
0.1 | 1 | 2019 | VEBO: a vertex- and edge-balanced ordering heuristic to load balance parallel graph processing · PPoPP 2019 |
Methods — techniques the papers use, named apart from their topics
hardware-software reconfiguration · 0.5dynamic algorithm selection · 0.5data indirection graph · 0.5compiler pass · 0.5adaptive prefetch distance · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Attention and Residual Structure Fusion for Photovoltaic Cell Defect Detection in Electroluminescence Images
Honghu Li, Jinhai Sa, Jiawen Sun, Huizi Li |
ICIC (20) | 3 |
| 2026 | SEAL: Semantic-Aware Contrastive Learning for scRNA-Seq ClusteringabstractThe development of single-cell RNA sequencing (scRNA-seq) technology has enabled the exploration of biological processes at the cellular level. A critical task in scRNA-seq data analysis is the unsupervised clustering of cells to distinguish different cell types. While various clustering methods have been successfully developed for scRNA-seq data, they still face limitations, particularly in terms of unstable clustering performance. This is often due to their inability to fully capture the intrinsic properties of cells, especially in the presence of high dropout rates and noise in the data. In this work, we propose a SEmantic-Aware contrastive Learning (SEAL) approach for scRNA-seq clustering. Specifically, we randomly mask the gene expression of each cell to generate two different augmentations of the cell data, and then apply semantic-aware contrastive learning to capture semantically invariant representations across these augmentations by leveraging semantic information from generated pseudo-labels. Experimental results demonstrate that our method effectively learns biologically meaningful representations and accurately identifies cell types. Yixuan Ye, Jiawen Sun, Jiajun Xian, Cheng Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | LLM-Assisted IDOR Detection in Hospital Mini-Programs: Risks to PII and PHIabstractHospital mini-programs have become widely adopted as lightweight portals for medical services, handling large volumes of personally identifiable information (PII) and protected health information (PHI). Among the most critical threats to such systems is the insecure direct object reference (IDOR) vulnerability, which allows unauthorized access to sensitive resources due to improper object–level access control. However, systematic detection of IDOR in the wild, especially within hospital mini-programs, remains underexplored due to restricted server access and stringent ethical regulations. To address this challenge, we propose a black–box detection framework designed for hospital mini-programs operating in sensitive data environments. Our framework introduces a novel token–substitution probing strategy that adheres to ethical standards and pioneers the use of Large Language Models (LLMs) for automated IDOR vulnerability detection in API endpoints, enabling token field identification, request classification and differential response analysis. We evaluated the framework on 80 real-world mini-programs and identified 114 vulnerable endpoints across 38 applications. Among these, 55 involved sensitive data disclosure, and 34 enabled unauthorized execution of sensitive operations. All findings were responsibly disclosed to the CNVD, and 15 cases have been officially confirmed. Jiawen Sun, Shangru Zhao, Xiangming Zhou, He Wang 0014, Yuqing Zhang 0001 |
TrustCom | 1 |
| 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph AnalyticsabstractSparse matrix-vector multiplication (SpMV) is a critical building block for iterative graph analytics algorithms. Typically, such algorithms have a varying active vertex set across iterations. This variability has been used to improve performance by either dynamically switching algorithms between iterations (software) or designing custom accelerators (hardware) for graph analytics algorithms. In this work, we propose a novel framework, CoSPARSE, that employs hardware and software reconfiguration as a synergistic solution to accelerate SpMV-based graph analytics algorithms. Building on previously proposed general-purpose reconfigurable hardware, we implement CoSPARSE as a software layer, abstracting the hardware as a specialized SpMV accelerator. CoSPARSE dynamically selects software and hardware configurations for each iteration and achieves a maximum speedup of 2.0 × compared to the naïve implementation with no reconfiguration. Across a suite of graph algorithms, CoSPARSE outperforms a state-of-the-art shared memory framework, Ligra, on a Xeon CPU with up to 3.51 × better performance and 877 × better energy efficiency. Siying Feng, Jiawen Sun, Subhankar Pal, Xin He 0011, Kuba Kaszyk, Dong-Hyeon Park, John Magnus Morton, Trevor N. Mudge, Murray Cole, Michael F. P. O'Boyle, Chaitali Chakrabarti, Ronald G. Dreslinski |
DAC | 2 |
| 2021 | Prodigy: Improving the Memory Latency of Data-Indirect Irregular Workloads Using Hardware-Software Co-DesignabstractIrregular workloads are typically bottlenecked by the memory system. These workloads often use sparse data representations, e.g., compressed sparse row/column (CSR/CSC), to conserve space at the cost of complicated, irregular traversals. Such traversals access large volumes of data and offer little locality for caches and conventional prefetchers to exploit. This paper presents Prodigy, a low-cost hardware-software codesign solution for intelligent prefetching to improve the memory latency of several important irregular workloads. Prodigy targets irregular workloads including graph analytics, sparse linear algebra, and fluid mechanics that exhibit two specific types of data-dependent memory access patterns. Prodigy adopts a “best of both worlds” approach by using static program information from software, and dynamic run-time information from hardware. The core of the system is the Data Indirection Graph (DIG)-a proposed compact representation used to express program semantics such as the layout and memory access patterns of key data structures. The DIG representation is agnostic to a particular data structure format and is demonstrated to work with several sparse formats including CSR and CSC. Program semantics are automatically captured with a compiler pass, encoded as a DIG, and inserted into the application binary. The DIG is then used to program a low-cost hardware prefetcher to fetch data according to an irregular algorithm's data structure traversal pattern. We equip the prefetcher with a flexible prefetching algorithm that maintains timeliness by dynamically adapting its prefetch distance to an application's execution pace. We evaluate the performance, energy consumption, and transistor cost of Prodigy using a variety of algorithms from the GAP, HPCG, and NAS benchmark suites. We compare the performance of Prodigy against a non-prefetching baseline as well as state-of-the-art prefetchers. We show that by using just 0.8KB of storage, Prodigy outperforms a non-prefetching baseline by $2.6 \times$ and saves energy by $1.6 \times$, on average. Prodigy also outperforms modern data prefetchers by $1.5- 2.3 \times$. Nishil Talati, Kyle May, Armand Behroozi, Yichen Yang 0005, Kuba Kaszyk, Christos Vasiladiotis, Tarunesh Verma, Brandon Nguyen, Jiawen Sun, John Magnus Morton, Agreen Ahmadi, Todd M. Austin, Michael F. P. O'Boyle, Scott A. Mahlke, Trevor N. Mudge, Ronald G. Dreslinski |
HPCA | 10 |
| 2020 | Transmuter: Bridging the Efficiency Gap using Memory and Dataflow ReconfigurationabstractWith the end of Dennard scaling and Moore's law, it is becoming increasingly difficult to build hardware for emerging applications that meet power and performance targets, while remaining flexible and programmable for end users. This is particularly true for domains that have frequently changing algorithms and applications involving mixed sparse/dense data structures, such as those in machine learning and graph analytics. To overcome this, we present a flexible accelerator called Transmuter, in a novel effort to bridge the gap between General-Purpose Processors (GPPs) and Application-Specific Integrated Circuits (ASICs). Transmuter adapts to changing kernel characteristics, such as data reuse and control divergence, through the ability to reconfigure the on-chip memory type, resource sharing and dataflow at run-time within a short latency. This is facilitated by a fabric of light-weight cores connected to a network of reconfigurable caches and crossbars. Transmuter addresses a rapidly growing set of algorithms exhibiting dynamic data movement patterns, irregularity, and sparsity, while delivering GPU-like efficiencies for traditional dense applications. Finally, in order to support programmability and ease-of-adoption, we prototype a software stack composed of low-level runtime routines, and a high-level language library called TransPy, that cater to expert programmers and end-users, respectively. Subhankar Pal, Siying Feng, Dong-Hyeon Park, Aporva Amarnath, Chi-Sheng Yang, Xin He 0011, Jonathan Beaumont, Kyle May, Yan Xiong 0002, Kuba Kaszyk, John Magnus Morton, Jiawen Sun, Michael F. P. O'Boyle, Murray Cole, Chaitali Chakrabarti, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski |
PACT | 13 |
| 2020 | DelayRepay: delayed execution for kernel fusion in PythonabstractPython is a popular, dynamic language for data science and scientific computing. To ensure efficiency, significant numerical libraries are implemented in static native languages. However, performance suffers when switching between native and non-native code, especially if data has to be converted between native arrays and Python data structures. As GPU accelerators are increasingly used, this problem becomes particularly acute. Data and control has to be repeatedly transferred between the accelerator and the host. John Magnus Morton, Kuba Kaszyk, Jiawen Sun, Christophe Dubach, Michel Steuwer, Murray Cole, Michael F. P. O'Boyle |
DLS | 4 |
| 2020 | Fast load balance parallel graph analytics with an automatic graph data structure selection algorithm
Jiawen Sun, Hans Vandierendonck, Dimitrios S. Nikolopoulos |
Future Gener. Comput. Syst. | 1 |
| 2019 | VEBO: a vertex- and edge-balanced ordering heuristic to load balance parallel graph processingabstractThis work proposes Vertex- and Edge-Balanced Ordering (VEBO): balance the number of edges and the number of unique destinations of those edges. VEBO balances edges and vertices for graphs with a power-law degree distribution, and ensures an equal degree distribution between partitions. Experimental evaluation on three shared-memory graph processing systems (Ligra, Polymer and GraphGrind) shows that VEBO achieves excellent load balance and improves performance by 1.09× over Ligra, 1.41× over Polymer and 1.65× over GraphGrind, compared to their respective partitioning algorithms, averaged across 8 algorithms and 7 graphs. VEBO improves GraphGrind performance with a speedup of 2.9× over Ligra on average. Jiawen Sun, Hans Vandierendonck, Dimitrios S. Nikolopoulos |
PPoPP | 1 |
| 2017 | Accelerating Graph Analytics by Utilising the Memory Locality of Graph PartitioningabstractThis paper investigates how to improve the memory locality of graph-structured analytics on large-scale shared memory systems. We demonstrate that a graph partitioning where all in-edges for a vertex are placed in the same partition improves memory locality. However, realising performance improvement through such graph partitioning poses several challenges and requires rethinking the classification of graph algorithms and preferred data structures. We introduce the notion of medium dense frontiers, a type of frontier that is sufficiently dense for a bitmap representation, yet benefits from an indexed graph layout. Using three types of frontiers, and three graph layout schemes optimized to each frontier type, we design an edge traversal algorithm that autonomously decides which type to use. The distinction of forward vs. backward graph traversal folds into this decision and need no longer be specified by the programmer.We have implemented our techniques in a NUMA-aware graph analytics framework derived from Ligra and demonstrate a speedup of up to 4.34× over Ligra and up to 2.93× over Polymer. Jiawen Sun, Hans Vandierendonck, Dimitrios S. Nikolopoulos |
ICPP | 1 |
| 2017 | GraphGrind: addressing load imbalance of graph partitioningabstractWe investigate how graph partitioning adversely affects the performance of graph analytics. We demonstrate that graph partitioning induces extra work during graph traversal and that graph partitions have markedly different connectivity than the original graph. By consequence, increasing the number of partitions reaches a tipping point after which overheads quickly dominate performance gains. Moreover, we show that the heuristic to balance CPU load between graph partitions by balancing the number of edges is inappropriate for a range of graph analyses. However, even when it is appropriate, it is sub-optimal due to the skewed degree distribution of social networks. Based on these observations, we propose GraphGrind, a new graph analytics system that addresses the limitations incurred by graph partitioning. We moreover propose a NUMA-aware extension to the Cilk programming language and obtain a scale-free yet NUMA-aware parallel programming environment which underpins NUMA-aware scheduling in GraphGrind. We demonstrate that Graph-Grind outperforms state-of-the-art graph analytics systems for shared memory including Ligra, Polymer and Galois. Jiawen Sun, Hans Vandierendonck, Dimitrios S. Nikolopoulos |
ICS | 1 |
| 2016 | Student Research Poster: A Scalable General Purpose System for Large-Scale Graph ProcessingabstractGraph analytics is an important and computationally demanding class of data analytics. It is essential to balance scalability, ease-of-use and high performance in large scale graph analytics. As such, it is necessary to hide the complexity of parallelism, data distribution and memory locality behind an abstract interface. Jiawen Sun |
PACT | 1 |