VLDB 2026 Research / reviewers in the wild / expert
Ren Chen
dblp:87/2517
· DBLP profile ↗
13ranked-venue papers
8as first author
3since 2021 · last 2025
0000-0003-3501-7630ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 8 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Reconfigurable computing and FPGAs · 34% Processor architecture and microarchitecture · 23% Hardware accelerators and domain-specific architectures · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Recommender systems · 100% | |
| Artificial intelligence
1 paper |
Graph learning · 100% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Recommender systems › sequential recommendation
efficient sequential recommendation |
0.9 | 1 | 2025 | Efficient Sequential Recommendation for Long Term User Interest Via Personalization · ICDM 2025 |
Recommender systems
sequential recommendation |
0.9 | 1 | 2025 | Efficient Sequential Recommendation for Long Term User Interest Via Personalization · ICDM 2025 |
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based sorting |
0.5 | 2 | 2017 | Computer Generation of High Throughput and Memory Efficient Sorting Designs on FPGA · IEEE Trans. Parallel Distributed Syst. 2017 Energy and Memory Efficient Mapping of Bitonic Sorting on FPGA · FPGA 2015 |
Machine learning › Graph learning › graph neural network
expressive power |
0.5 | 1 | 2021 | Decoupling the Depth and Scope of Graph Neural Networks · NeurIPS 2021 |
Machine learning › Graph learning
graph neural network |
0.5 | 1 | 2021 | Decoupling the Depth and Scope of Graph Neural Networks · NeurIPS 2021 |
Machine learning › Graph learning › graph neural network
scalable graph neural network |
0.5 | 1 | 2021 | Decoupling the Depth and Scope of Graph Neural Networks · NeurIPS 2021 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator |
0.3 | 1 | 2018 | A Framework for Generating High Throughput CNN Implementations on FPGAs · FPGA 2018 |
Processor architecture and microarchitecture
bit reversal |
0.3 | 1 | 2017 | Optimal Circuits for Parallel Bit Reversal · DAC 2017 |
Integrated circuit design
digital circuit design |
0.3 | 1 | 2017 | Optimal Circuits for Parallel Bit Reversal · DAC 2017 |
Processor architecture and microarchitecture
instruction set architecture |
0.3 | 1 | 2017 | Optimal Circuits for Parallel Bit Reversal · DAC 2017 |
Interconnection networks and networks-on-chip
sorting network |
0.3 | 1 | 2017 | Computer Generation of High Throughput and Memory Efficient Sorting Designs on FPGA · IEEE Trans. Parallel Distributed Syst. 2017 |
Machine learning › Graph learning › graph analytics › graph mining
subgraph extraction |
0.1 | 1 | 2021 | Decoupling the Depth and Scope of Graph Neural Networks · NeurIPS 2021 |
Electronic design automation › hardware verification and test
verilog generation |
0.1 | 1 | 2018 | A Framework for Generating High Throughput CNN Implementations on FPGAs · FPGA 2018 |
Parallel and multicore computing
data permutation |
0.1 | 1 | 2017 | Computer Generation of High Throughput and Memory Efficient Sorting Designs on FPGA · IEEE Trans. Parallel Distributed Syst. 2017 |
Reconfigurable computing and FPGAs
FPGA power reduction |
0.1 | 1 | 2015 | Energy and Memory Efficient Mapping of Bitonic Sorting on FPGA · FPGA 2015 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.9token compression · 0.9clos network folding · 0.5topological learning · 0.5graph signal processing · 0.5function approximation · 0.5overlap-and-add · 0.3frequency domain loop tiling · 0.3design space exploration · 0.3concatenate-and-pad · 0.3bitonic sorting network · 0.3RAM-based design · 0.3streaming permutation network · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Sequential Recommendation for Long Term User Interest Via PersonalizationabstractRecent years have witnessed success of sequential modeling, generative recommender, and large language model for recommendation. Though the scaling law has been validated for sequential models, it showed inefficiency in computational capacity when considering real-world applications like recommendation, due to the non-linear(quadratic) increasing nature of the transformer model. To improve the efficiency of the sequential model, we introduced a novel approach to sequential recommendation that leverages personalization techniques to enhance efficiency and performance. Our method compresses long user interaction histories into learnable tokens, which are then combined with recent interactions to generate recommendations. This approach significantly reduces computational costs while maintaining high recommendation accuracy. Our method could be applied to existing transformer based recommendation models, e.g., HSTU and HLLM. Extensive experiments on multiple sequential models demonstrate its versatility and effectiveness. Source code is available at https://github.com/facebookresearch/PerSRec. Hanchao Yu, Ivan Ji, Chen Yuan 0001, Chihuang Liu, Christopher E. Lambert, Ren Chen, Chen Kovacs, Xinzhu Bei, Renqin Cai, Lizhu Zhang, Xiangjun Fan, Qunshu Zhang, Benyu Zhang |
ICDM | 9 |
| 2022 | Deep Learning for Image and Point Cloud Fusion in Autonomous Driving: A ReviewabstractAutonomous vehicles were experiencing rapid development in the past few years. However, achieving full autonomy is not a trivial task, due to the nature of the complex and dynamic driving environment. Therefore, autonomous vehicles are equipped with a suite of different sensors to ensure robust, accurate environmental perception. In particular, the camera-LiDAR fusion is becoming an emerging research theme. However, so far there has been no critical review that focuses on deep-learning-based camera-LiDAR fusion methods. To bridge this gap and motivate future research, this article devotes to review recent deep-learning-based data fusion approaches that leverage both image and point cloud. This review gives a brief overview of deep learning on image and point cloud data processing. Followed by in-depth reviews of camera-LiDAR fusion methods in depth completion, object detection, semantic segmentation, tracking and online cross-sensor calibration, which are organized based on their respective fusion levels. Furthermore, we compare these methods on publicly available datasets. Finally, we identified gaps and over-looked challenges between current academic researches and real-world applications. Based on these observations, we provide our insights and point out promising research directions. Yaodong Cui, Ren Chen, Wenbo Chu, Long Chen 0005, Daxin Tian, Ying Li 0036, Dongpu Cao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Decoupling the Depth and Scope of Graph Neural NetworksabstractState-of-the-art Graph Neural Networks (GNNs) have limited scalability with respect to the graph and model sizes. On large graphs, increasing the model depth often means exponential expansion of the scope (i.e., receptive field). Beyond just a few layers, two fundamental challenges emerge: 1. degraded expressivity due to oversmoothing, and 2. expensive computation due to neighborhood explosion. We propose a design principle to decouple the depth and scope of GNNs – to generate representation of a target entity (i.e., a node or an edge), we first extract a localized subgraph as the bounded-size scope, and then apply a GNN of arbitrary depth on top of the subgraph. A properly extracted subgraph consists of a small number of critical neighbors, while excluding irrelevant ones. The GNN, no matter how deep it is, smooths the local neighborhood into informative representation rather than oversmoothing the global graph into “white noise”. Theoretically, decoupling improves the GNN expressive power from the perspectives of graph signal processing (GCN), function approximation (GraphSAGE) and topological learning (GIN). Empirically, on seven graphs (with up to 110M nodes) and six backbone GNN architectures, our design achieves significant accuracy improvement with orders of magnitude reduction in computation and hardware cost. Hanqing Zeng, Muhan Zhang, Yinglong Xia, Ajitesh Srivastava, Andrey Malevich, Rajgopal Kannan, Viktor Prasanna 0001, Ren Chen |
NeurIPS | 9 |
| 2018 | A Framework for Generating High Throughput CNN Implementations on FPGAsabstractWe propose a framework to generate highly efficient accelerators for inferencing on FPGAs. Our framework consists of multiple algorithmic optimizations for computation complexity and communication volume reduction, a mapping methodology for efficient resource utilization, and a tool for automatic \textttVerilog generation. The algorithmic optimizations improve throughput of frequency domain convolution so as to satisfy a given set of hardware constraints. While the Overlap-and-Add (OaA) technique has been known, it performs "wasted" computation at the edges. We propose a novel Concatenate-and-Pad (CaP) technique, which improves OaA significantly by reducing the "wasted" computation on the padded pixels. The proposed CaP used in conjunction with OaA enables us to choose a fixed FFT size at design time, and achieve low computation complexity for layers with various image sizes and kernel window sizes. We also develop a novel frequency domain loop tiling technique to further boost throughput by improving data reuse. Our mapping methodology optimizes the architecture for the target device by fast design space exploration. We quantitatively categorize FPGAs by capturing their DSP resources, on-chip memory size and external memory bandwidth into a device coefficient. We identify the optimal architectural parameters based on the tradeoff between computation and communication cost. Our framework includes a tool to automatically generate fully synthesizable \textttVerilog. We demonstrate the framework by generating high throughput accelerators for state-of-the-art CNN models on Intel HARP heterogeneous platform. Using our framework, we achieve throughput of $780.6$ $GOPS$, $669.1$ $GOPS$ and $552.1$ $GOPS$ for AlexNet, VGG16 and FCN-16s respectively. These correspond to $6.8\times$ (AlexNet) and $4.9\times$ (VGG16) improvement compared with the state-of-the-art implementations. Hanqing Zeng, Ren Chen, Chi Zhang 0022, Viktor Prasanna 0001 |
FPGA | 2 |
| 2018 | C-Graph: A Highly Efficient Concurrent Graph Reachability Query FrameworkabstractMany big data analytics applications explore a set of related entities, which are naturally modeled as graph. However, graph processing is notorious for its performance challenges due to random data access patterns, especially for large data volumes. Solving these challenges is critical to the performance of industry-scale applications. In contrast to most prior works, which focus on accelerating a single graph processing task, in industrial practice we consider multiple graph processing tasks running concurrently, such as a group of queries issued simultaneously to the same graph. In this paper, we present an edge-set based graph traversal framework called C-Graph (i.e. Concurrent Graph), running on a distributed infrastructure, that achieves both high concurrency and efficiency for k-hop reachability queries. The proposed framework maintains global vertex states to facilitate graph traversals, and supports both synchronous and asynchronous communication. In this study, we decompose a set of graph processing tasks into local traversals and analyze their performance on C-Graph. More specifically, we optimize the organization of the physical edge-set and explore the shared subgraphs. We experimentally show that our proposed framework outperforms several baseline methods. Li Zhou 0012, Ren Chen, Yinglong Xia, Radu Teodorescu |
ICPP | 2 |
| 2017 | Optimal Circuits for Parallel Bit ReversalabstractIn this paper, we develop novel parallel circuit designs for calculating the bit reversal. To perform bit reversal on 2n data words, the designs take 2k (k Ren Chen, Viktor Prasanna 0001 |
DAC | 1 |
| 2017 | Optimal dynamic data layouts for 2D FFT on 3D memory integrated FPGA
Ren Chen, Shreyas G. Singapura, Viktor Prasanna 0001 |
J. Supercomput. | 1 |
| 2017 | Computer Generation of High Throughput and Memory Efficient Sorting Designs on FPGAabstractAccelerating sorting using dedicated hardware to fully utilize the memory bandwidth for Big Data applications has gained much interest in the research community. Recently, parallel sorting networks have been widely employed in hardware implementations due to their high data parallelism and low control overhead. In this paper, we propose a systematic methodology for mapping large-scale bitonic sorting networks onto FPGA. To realize data permutations in the sorting network, we develop a novel RAM-based design by vertically “folding” the classic Clos network. By utilizing the proposed design for data permutation, we develop a hardware generator to automatically build bitonic sorting architectures on FPGAs. For given input size, data width and data parallelism, the hardware generator specializes both the datapath and the control unit for sorting and generates a design in high level hardware description language. We demonstrate trade-offs among throughput, latency and area using two illustrative sorting designs including a high throughput design and a resource efficient design. With a data parallelism of p (2 ≤ p ≤ N/2), the high throughput design sorts an N-key sequence with latency 6N=p + o(N), throughputp results per cycle and uses 6N + o(N) memory. This achieves optimal memory efficiency (defined as the ratio of throughput to the amount of on-chip memory used by the design) and outperforms the state-of-the-art. Experimental results show that the designs obtained by our proposed hardware generator achieve 49 to 112 percent improvement in energy efficiency and 56 to 430 percent higher memory efficiency compared with the state-of-the-art. Ren Chen, Viktor Prasanna 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Accelerating Equi-Join on a CPU-FPGA Heterogeneous PlatformabstractAccelerating database applications using FPGAs has recently been an area of growing interest in both academia and industry. Equi-join is one of the key database operations whose performance highly depends on sorting, which exhibits high memory usage on FPGA. A fully pipelined N-key merge sorter consists of log N sorting stages using O(N) memory totally. For large data sets, external memory has to be employed to perform data buffering between the sorting stages. This introduces pipeline stalls as well as several iterations between FPGA and external memory, causing significant performance degradation. In this paper, we speed-up equi-join using a hybrid CPU-FPGA heterogeneous platform. To alleviate the performance impact of limited memory, we propose a merge sort based hybrid design where the first few sorting stages in the merge sort tree are replaced with "folded" bitonic sorting networks. These "folded" bitonic sorting networks operate in parallel on the FPGA. The partial results are then merged on the CPU to produce the final sorted result. Based on this hybrid sorting design, we develop two streaming join algorithms by optimizing the classic CPU-based nested-loop join and sort-merge join algorithms. On a rangeof data set sizes, our design achieves throughput improvement of 3.1x and 1.9x compared with software-only and FPGA only implementations, respectively. Our design sustains 21.6% of thepeak bandwidth, which is 3.9x utilization obtained by the state-of-the-art FPGA equi-join implementation. Ren Chen, Viktor Prasanna 0001 |
FCCM | 1 |
| 2016 | Optimizing interconnection complexity for realizing fixed permutation in data and signal processing algorithmsabstractIn hardware implementation of several widely used data and signal processing algorithms, data permutations need to be performed between the consecutive computation stages consisting of parallel computational units. Recently, some highly data parallel streaming architectures for data permutation have been proposed to achieve high throughput. However, the interconnection complexity of these designs increases dramatically with the problem size and data parallelism. In this paper, we develop a hardware structure to perform data permutation with optimized interconnection complexity, denned as the interconnection area per throughput. We propose a novel design technique such that the required interconnection logic is highly reduced for realizing a fixed permutation on streaming data. Our experimental results show that the proposed design technique reduces interconnection complexity by 27.3% to 75.8%, and improves the throughput by 5.3%~129% and the energy efficiency by 1.2×~3.5× compared with the state-of-the-art. Ren Chen, Viktor Prasanna 0001 |
FPL | 1 |
| 2015 | Energy and Memory Efficient Mapping of Bitonic Sorting on FPGAabstractParallel sorting networks are widely employed in hardware implementations for sorting due to their high data parallelism and low control overhead. In this paper, we propose an energy and memory efficient mapping methodology for implementing bitonic sorting network on FPGA. Using this methodology, the proposed sorting architecture can be built for a given data parallelism while supporting continuous data streams. We propose a streaming permutation network (SPN) by "folding" the classic Clos network. We prove that the SPN is programmable to realize all the interconnection patterns in the bitonic sorting network. A low cost design for sorting with minimal resource usage is obtained by reusing one SPN . We also demonstrate a high throughput design by trading off area for performance. With a data parallelism of p (2 ≤ p ≤ N/ log2 N), the high throughput design sorts an N-key sequence with latency O(N/p), throughput (# of keys sorted per cycle) O(p) and uses O(N) memory. This achieves optimal memory efficiency (defined as the ratio of throughput to the amount of on-chip memory used by the design) of O(p/N). Another noteworthy feature of the high throughput design is that only single-port memory rather than dual-port memory is required for processing continuous data streams. This results in 50% reduction in memory consumption. Post place-and-route results show that our architecture demonstrates 1.3x ∼1.6x improvment in energy efficiency and 1.5x ∼ 5.3x better memory efficiency compared with the state-of-the-art designs. Ren Chen, Sruja Siriyal, Viktor Prasanna 0001 |
FPGA | 1 |
| 2015 | Automatic generation of high throughput energy efficient streaming architectures for arbitrary fixed permutationsabstractDue to their high data-rate and simple control, streaming architectures have become popular for hardware implementation of data intensive applications. A key problem in designing such architectures is to permute streaming data. In this paper, we present a technique to realize arbitrary fixed permutation on streaming data. We develop a parameterized architecture which accepts data streams as input and generates the permuted data after a certain amount of delay. Our design accepts continuous input at a fixed rate of p per cycle, where p is the data parallelism of the architecture. To construct the streaming architecture for a given fixed permutation, we develop a mapping approach by configuring the classic Benes network to obtain the datapath and the control logic. We demonstrate a complete design automation tool which takes as input design parameters including the permutation pattern and the data parallelism p, and produces register-transfer level Verilog description of the design. We evaluate the generated designs on Xilinx Virtex-7 FPGA using post place-and-route results. Ren Chen, Viktor Prasanna 0001 |
FPL | 1 |
| 2013 | Energy efficient parameterized FFT architectureabstractIn this paper, we revisit the classic Fast Fourier Transform (FFT) for energy efficient designs on FPGAs. A parameterized FFT architecture is proposed to identify the design trade-offs in achieving energy efficiency. We first perform design space exploration by varying the algorithm mapping parameters, such as the degree of vertical and horizontal parallelism, that characterize decomposition based FFT algorithms. Then we explore an energy efficient design by empirical selection on the values of the chosen architecture parameters, including the type of memory elements, the type of interconnection network and the number of pipeline stages. The trade offs between energy, area, and time are analyzed using two performance metrics: the energy efficiency (defined as the number of operations per Joule) and the Energy×Area×Time (EAT) composite metric. From the experimental results, a design space is generated to demonstrate the effect of these parameters on the various performance metrics. For N-point FFT (16 ≤ N ≤ 1024), our designs achieve up to 28% and 38% improvement in the energy efficiency and EAT, respectively, compared with a state-of-the-art design. Ren Chen, Hoang Le, Viktor Prasanna 0001 |
FPL | 1 |