VLDB 2026 Research / reviewers in the wild / expert
Jiaxian Chen
dblp:256/1838
· DBLP profile ↗
16ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empirical knowledge-driven remaining useful life prediction method under unknown failure pattern
Jiaxian Chen, Shuhan Deng, Guolin He, Zhuyun Chen 0001, Weihua Li 0004 |
Adv. Eng. Informatics | 1 |
| 2026 | Registry: Enhancing Vertex Reusability for GCN Inference on Hybrid Stacked Memory
Zhaoyu Zhong, Jiaxian Chen, Yunhao Dong, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | GALC: Guided Amplified Learning With Lipschitz Constraint for Robust Trajectory GenerationabstractOffline reinforcement learning (RL) demonstrated remarkable performance in learning valid policies by benefiting from high-quality offline datasets. However, collecting such a dataset is labor-intensive, especially for humanoid locomotion. For this reason, many data augmentation techniques have been proposed to improve the quality of offline datasets through noise injection or data synthesis. However, existing data augmentation methods are noise-sensitive, resulting in limited capability in complex robotic environments. To address these issues, we propose guided amplified learning with Lipschitz constraint (GALC), a novel trajectory augmentation method that employs the reward-amplification-guided conditional diffusion model for noise-insensitive data augmentation. Specifically, we introduce a local Lipschitz continuity constraint to regulate the reverse denoising process from the offline dataset. Consequently, the exploration of the diffusion model can be restricted within the local continuity region of the original dataset, thereby generating high-reward trajectories. Moreover, the generated trajectories are also enforced to be noise-insensitive to perturbations, thus enjoying robustness. Notably, our proposed method can prevent the generation of unsafe actions that do not align with the environment dynamics. Extensive experiments on sparse reward scenarios and high-dimensional robotic tasks show that our proposed GALC achieves significant improvements in both the augmented trajectories and policy performance. Zhiliang Lin, Zhuangzhuang Chen, Guanming Zhu, Jiaxian Chen, Jianqiang Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2025 | Unifying Two Operators with One PIM: Leveraging Hybrid Bonding for Efficient LLM Inference
Jiaxian Chen, Yuxuan Qi, Kaoyi Sun, Zhiliang Lin, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
APPT | 1 |
| 2025 | Move Less, Retrieve Fast: A Retrieval-in-Memory Architecture for Language ModelsabstractRetrieval-augmented language models (RALMs) have attracted widespread attention for addressing the limitations of traditional large language models. However, challenges involved in retrieval, including substantial data movement and irregular access patterns, seriously impact the efficiency and deployment of RALMs. The emerging 3D-stacked processing-in-memory (PIM) architecture, characterized by its high memory bandwidth and near-data computing capabilities, presents a promising solution for efficient retrieval. To support large-scale retrieval in RALMs, the PIM architecture should be carefully designed with joint software and hardware optimization. This paper presents Rimast, a retrieval-in-memory architecture for fast retrieval in RALMs. The objective is to minimize data movement and improve overall performance through hardwaresoftware co-design. At the hardware level, a hierarchical PIM architecture with a retrieval-in-memory dataflow is designed to reduce unnecessary data transfer. At the software level, skew-free data mapping and adaptive offloading strategies are proposed to address the irregular access patterns associated with retrieval in RALMs. We demonstrate the effectiveness of the proposed Rimast using extensive experiments. The experimental results demonstrate that Rimast effectively reduces data movement, achieving average speedups of $273 \times 55 \times$, and $2.41 \times$ over CPUs, GPUs, and prior art accelerators, respectively. Jiaxian Chen, Yuxuan Qi, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
DAC | 1 |
| 2025 | Anchor First, Accelerate Next: Revolutionizing GNNs with PIM by Harnessing Stationary DataabstractSubstantial data movement caused by irregular graph topologies hinders the efficient processing of graph neural networks (GNNs). Although the emerging near-bank processing-in-memory (PIM) architecture offers a promising solution to reduce data transfer between memory and computing units, cross-bank communication remains a critical challenge, limiting the benefits of PIM architectures. Our findings indicate that only $35.6 \%$ of the data can stay stationary within PIM units on average, with the rest requiring movement due to graph dependencies. This situation worsens as the number of PIM units increases, reducing the ratio to $18.7 \%$. In this paper, we argue that to fully leverage PIM architectures, systems must maximize stationary data and minimize the movement of non-stationary data. Following this principle, we propose Anchor, a scalable PIM architecture that exploits stationary data for GNNs through a hardware-software co-design approach. To maximize stationary data, we introduce the graph partitioning algorithm Mastav, which carefully allocates vertices and edges to preserve data locality. To minimize the movement of non-stationary data, we employ a two-step strategy. First, a customized dataflow ensures that non-stationary data is accessed and distributed exactly once. Second, an optimized communication mechanism reduces redundant data transfers through critical paths. Our extensive experiments demonstrate that Anchor significantly reduces processing latency and data movement compared to representative schemes. Jiaxian Chen, Yuxuan Qi, Yongbiao Zhu, Jianan Yuan, Kaoyi Sun, Tianyu Wang 0009, Chenlin Ma, Yi Wang 0003 |
DAC | 1 |
| 2025 | MiniWear: Minimizing Flash Wear via Hybrid Persistent Cache for Extended EF-SMR LifetimeabstractAs the huge discrepancy between traffic and capacity persists, the lifetime of flash in EF-SMR systems faces a grave issue. EF-SMR systems combine NAND flash with Shingled Magnetic Recording (SMR) disks to achieve both low cost and high performance. However, previous research has primarily focused on issues such as write amplification and tail-latency in EF-SMR disks, overlooking the critical issue of flash lifetime. Studying the durability of EF-SMR systems is essential for developing future high-performance, low-cost storage solutions.This paper presents MiniWear, a hybrid persistent cache (PC) design aimed at extending the lifetime of EF-SMR systems. MiniWear adopts a hybrid medium persistent cache and proposes a customized scheduling strategy to reduce flash wear without impacting the EF-SMR system performance. At the hardware level, the hybrid PC of EF-SMR, composed of flash and SMR disk, is organized into Flash-PC and SMR-PC. At the software level, a fine-grained scheduling strategy is proposed to better manage PC resources. Additionally, we introduce a proactive balancing strategy to address PC resource idleness. Experimental results show that, compared to existing methods, MiniWear can reduce flash wear by up to 66.67%. Chenlin Ma, Kaoyi Sun, Yuxuan Qi, Jiaxian Chen, Xiaochuan Zheng, Tianyu Wang 0009, Yi Wang 0003 |
DAC | 4 |
| 2024 | Leanor: A Learning-Based Accelerator for Efficient Approximate Nearest Neighbor Search via Reduced Memory AccessabstractApproximate Nearest Neighbor Search (ANNS) is a classical problem in data science. ANNS is both computationally-intensive and memory-intensive. As a typical implementation of ANNS, Inverted File with Product Quantization (IVFPQ) has the properties of high precision and rapid processing. However, the traversal of non-nearest neighbor vectors in IVFPQ leads to redundant memory accesses. This significantly impacts retrieval efficiency. A promising approach involves the utilization of learned indexes, leveraging insights from data distribution to optimize search efficiency. Existing learned indexes are primarily customized for low-dimensional data. How to tackle ANNS in high-dimensional vectors is a challenging issue. Yi Wang 0003, Jianan Yuan, Jiaxian Chen, Tianyu Wang 0009, Chenlin Ma, Rui Mao 0001 |
DAC | 4 |
| 2024 | Boosting Write Performance of KV Stores: An NVM - Enabled Storage Collaboration ApproachabstractAs the most common data structure for key-value stores, LogStructured Merge Tree (LSM-tree) can eliminate random write operations and keep acceptable read performance. However, write stall and write amplification introduced by the leveled compaction of LSM-tree significantly degrade the system performance. The emerging non-volatile memory (NVM) provides byte-addressable access and low-latency data persistence. Integrating DIMM-interface NVM in the design of the LSM-tree can potentially alleviate the write stall and write amplification issue, as the access speed of NVM is several orders of magnitude faster than hard disk drives or flash memory-based solid-state drives. This hybrid storage should be carefully designed, requiring new architectural and key-value structural support. This paper presents ZigZagDB, an NVM-enabled data man-agement scheme for LSM-tree-based key-value stores. ZigZagDB adds additional layers of key-value stores and uses non-volatile memory as the storage media to hold these additional layers of data. The newly designed key-value stores alternately access the data from either SSD or NVM. This ‘ZigZag’ shape of storage collaboration and synchronization can benefit write efficiency and space utilization. By utilizing the NVM with very limited capacity, the redesigned organization of LSM-tree can effectively solve the write stall and write amplification issue. We demonstrate the viability of the proposed ZigZagDB using a set of extensive experiments. Experimental results show that ZigZagDB can significantly reduce the write amplification and boost the throughput in comparison with representative schemes. Yi Wang 0003, Jiajian He, Kaoyi Sun, Yunhao Dong, Jiaxian Chen, Chenlin Ma, Amelie Chi Zhou, Rui Mao 0001 |
ICDE | 5 |
| 2024 | LeaderKV: Improving Read Performance of KV Stores via Learned Index and Decoupled KV TableabstractLog-structured merge-tree (LSM-tree) is a storage architecture widely used in key-value (KV) stores. To enhance the read efficiency of LSM-tree, recent works utilize the learned index to learn the mapping between keys and locations. However, in existing learned-index-aided KV stores, inefficient design of the learned index and disk access significantly impact the read performance. How to design a learned KV store to improve index efficiency and minimize disk access remains a critical problem. This paper presents LeaderKV, a read-optimized LSM-tree-based KV store. LeaderKV employs decoupled KV tables (DK-Table) and efficient learned indexes for data retrieval. DKTables are storage files in Leader Kvbecause they avoid reading irrelevant data in collaboration with learned indexes during queries. A learned index called Leader is proposed to accelerate data retrieval within DKTable. Leader is composed of precise models and approximate models. A redirect mechanism is designed to reduce the cost of mispredictions in Leader. We integrate DKTable and Leader into LeaderKV and demonstrate its effectiveness using a variety of datasets and workloads. Experimental results show that LeaderKV significantly improves the read performance compared to representative schemes. Yi Wang 0003, Jianan Yuan, Shangyu Wu, Jiaxian Chen, Chenlin Ma, Jianbin Qin |
ICDE | 5 |
| 2024 | A digital twin-driven approach for partial domain fault diagnosis of rotating machinery
Jingyan Xia, Zhuyun Chen 0001, Jiaxian Chen, Guolin He, Ruyi Huang, Weihua Li 0004 |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Knowledge Embedded Autoencoder Network for Harmonic Drive Fault Diagnosis Under Few-Shot Industrial ScenariosabstractThe development of Internet of Things technology provides abundant data resources for prognostics health management of industrial machinery, and data-driven methods have shown their powerful ability in the field of fault diagnosis. However, these methods have several limitations: 1) Using less labeled data to obtain higher accuracy is a challenging task, which limits the application of diagnostic models in practical applications. 2) Physics-informed knowledge is largely ignored during the modeling process, which contains a wealth of information that can reflect the harmonic drive’s health status. To address these challenges, a self-supervised fault diagnosis framework is developed by integrating prior knowledge with deep learning to improve the accuracy and reliability of diagnosis models in industrial applications. Specifically, the physics-based knowledge including 32-dimensional time domain, frequency domain, and time-frequency domain features, is first designed to provide fault information and significantly reduce the amount of data required for deep learning. Furthermore, a self-supervised knowledge embedded auto-encoder network is built by employing the prior knowledge in the multi-scale convolutional auto-encoder. With the ability to integrate prior knowledge and the self-supervised learning mechanism, the proposed method can provide a strong tool for knowledge representation and an effective solution for fault diagnosis under a few-shot industrial scenario. The experimental results conducted on a real harmonic drive fault dataset prove that the proposed network framework provides effective insights on fault diagnosis and has excellent generalizability in practical industrial applications. Jiaxian Chen, Kairu Wen, Jingyan Xia, Ruyi Huang, Zhuyun Chen 0001, Weihua Li 0004 |
IEEE Internet Things J. | 1 |
| 2023 | Lift: Exploiting Hybrid Stacked Memory for Energy-Efficient Processing of Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) are powerful learning approaches for graph-structured data. GCNs are both computing- and memory-intensive. The emerging 3D-stacked computation-in-memory (CIM) architecture provides a promising solution to process GCNs efficiently. The CIM architecture can provide near-data computing, thereby reducing data movement between computing logic and memory. However, previous works do not fully exploit the CIM architecture in both dataflow and mapping, leading to significant energy consumption.This paper presents Lift, an energy-efficient GCN accelerator based on 3D CIM architecture using software and hardware co-design. At the hardware level, Lift introduces a hybrid architecture to process vertices with different characteristics. Lift adopts near-bank processing units with a push-based dataflow to process vertices with strong re-usability. A dedicated unit is introduced to reduce massive data movement caused by high-degree vertices. At the software level, Lift adopts a hybrid mapping to further exploit data locality and fully utilize the hybrid computing resources. The experimental results show that the proposed scheme can significantly reduce data movement and energy consumption compared with representative schemes. Jiaxian Chen, Zhaoyu Zhong, Kaoyi Sun, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
DAC | 1 |
| 2023 | Self-Supervised Deep Domain-Adversarial Regression Adaptation for Online Remaining Useful Life Prediction of Rolling Bearing Under Unknown Working ConditionabstractThis article proposes a novel deep transfer learning-based online remaining useful life (RUL) approach for rolling bearings under unknown working condition. This approach solves the following concerns: the drift of online working condition would block data accumulation and raise bias in the prediction model, and online bearing merely has early fault data when activating RUL prediction, failing to conduct transfer learning from offline data. First, a new transfer learning-based time series recursive forecasting model is constructed to generate online RUL pseudovalues via fusing prior degradation information from offline whole-life data. With such supervised information, a new deep domain-adversarial regression network with multilevel adaptation is further built to transfer prognostic knowledge from offline data to online scenario and evaluate the RUL values of online data batch. Experimental results on the IEEE PHM Challenge 2012 bearing dataset and XJTU-SY bearing dataset validate the effectiveness of the proposed approach. Wentao Mao, Jiaxian Chen, Jing Liu 0067, Xihui Liang |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | GCIM: Toward Efficient Processing of Graph Convolutional Networks in 3D-Stacked MemoryabstractGraph convolutional networks (GCNs) have become a powerful deep learning approach for graph-structured data. Different from traditional neural networks such as convolutional neural networks, GCNs handle irregular input graph data, and GCNs are both computation-bound and memory-bound. How to efficiently utilize the underlying computation and memory resource becomes a critical issue. The emerging 3D-stacked computation-in-memory (CIM) architecture can reduce the data movement between computing logic and memory, thereby presenting a promising solution for the processing of GCNs. An unsolved key challenge is how to allocate GCNs to take advantage of fast near-data processing of the 3D-stacked CIM architecture. This article presents GCIM, a software–hardware co-design approach to exploit the efficient processing of GCNs on the CIM architecture. At the level of hardware design, GCIM integrates lightweight computing units near memory banks to fully exploit bank-level bandwidth and parallelism. At the level of software design, a locality-aware data mapping algorithm is proposed to partition the input graph and achieve workload balancing. GCIM is evaluated through a set of representative GCN models and standard graph datasets. The experimental results show that GCIM can significantly reduce the processing latency and data movement overhead compared with representative schemes. Jiaxian Chen, Yiquan Lin, Kaoyi Sun, Jiexin Chen, Chenlin Ma, Rui Mao 0001, Yi Wang 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Towards efficient allocation of graph convolutional networks on hybrid computation-in-memory architecture
Jiaxian Chen, Guanquan Lin, Jiexin Chen, Yi Wang 0003 |
Sci. China Inf. Sci. | 1 |