VLDB 2026 Research / reviewers in the wild / expert
Shengqi Chen 0001
dblp:156/3418
· DBLP profile ↗
13ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-2310-5249ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Parallel MIP Solving with Dynamic Task Decomposition
Peng Lin 0005, Shaowei Cai 0001, Mengchuan Zou, Shengqi Chen 0001 |
CP | 4 |
| 2025 | MEPipe: Democratizing LLM Training with Memory-Efficient Slice-Level Pipeline Scheduling on Cost-Effective AcceleratorsabstractThe training of large language models (LLMs) typically needs costly GPUs, such as NVIDIA A100 or H100. They possess substantial high-bandwidth on-chip memory and rapid interconnects like NVLinks. The exorbitant expenses associated with LLM training pose not just an economic challenge but also a societal one, as it restricts the ability to train LLMs from scratch to a selected few organizations. Zhenbo Sun, Shengqi Chen 0001, Yuanwei Wang, Jian Sha, Guanyu Feng |
EuroSys | 2 |
| 2025 | Iterative Design of a Teaching Assistant Training Program in Computer Science Using the Agile MethodabstractFacing soaring enrollment and disruptive educational technologies, computing education increasingly relies on the contributions of teaching assistants (TAs), hence the critical importance of high-quality TA training. However, the design and implementation of TA training in computer science face substantial barriers, such as the lack of experienced TA trainers and the scarcity of relevant training materials. Runda Liu, Shengqi Chen 0001, Songjie Niu, Yuchun Ma, Xiaofeng Tang |
SIGCSE (1) | 2 |
| 2025 | HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender Models
Jiaao He, Shengqi Chen 0001, Kezhao Huang, Jidong Zhai |
USENIX ATC | 2 |
| 2024 | AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningabstractLarge language models (LLMs) have demonstrated powerful capabilities, requiring huge memory with their increasing sizes and sequence lengths, thus demanding larger parallel systems. The broadly adopted pipeline parallelism introduces even heavier and unbalanced memory consumption. Recomputation is a widely employed technique to mitigate the problem but introduces extra computation overhead. Zhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng, Shengqi Chen 0001, Haojie Wang 0004 |
ASPLOS (3) | 5 |
| 2024 | POSTER: Pattern-Aware Sparse Communication for Scalable Recommendation Model TrainingabstractRecommendation models are an important category of deep learning models whose size is growing enormous. They consist of a sparse part with TBs of memory footprint and a dense part that demands PFLOPs of computing capability to train. Unfortunately, the high sparse communication cost to re-organize data for different parallel strategies of the two parts impedes the scalability in training. Jiaao He, Shengqi Chen 0001, Jidong Zhai |
PPoPP | 2 |
| 2023 | TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory ProgramsabstractOut-of-core systems rely on high-performance cache sub-systems to reduce the number of I/O operations. Although the page cache in modern operating systems enables transparent access to memory and storage devices, it suffers from efficiency and scalability issues on cache misses, forcing out-of-core systems to design and implement their own cache components, which is a non-trivial task. This study proposes TriCache, a cache mechanism that enables in-memory programs to efficiently process out-of-core datasets without requiring any code rewrite. It provides a virtual memory interface on top of the conventional block interface to simultaneously achieve user transparency and sufficient out-of-core performance. A multi-level block cache design is proposed to address the challenge of per-access address translations required by a memory interface. It can exploit spatial and temporal localities in memory or storage accesses to render storage-to-memory address translation and page-level concurrency control adequately efficient for the virtual memory interface. Our evaluation shows that in-memory systems operating on top of TriCache can outperform Linux OS page cache by more than one order of magnitude, and can deliver performance comparable to or even better than that of corresponding counterparts designed specifically for out-of-core scenarios. Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001 |
ACM Trans. Storage | 7 |
| 2022 | Efficiently emulating high-bitwidth computation with low-bitwidth hardwareabstractDomain-Specific Accelerators (DSAs) are being rapidly developed to support high-performance domain-specific computation. Although DSAs provide massive computation capability, they often only support limited native data types. To mitigate this problem, previous works have explored software emulation for certain data types, which provides some compensation for hardware limitations. However, how to efficiently design more emulated data types and choose a high-performance one without hurting correctness or precision for a given application still remains an open problem. Zixuan Ma, Haojie Wang 0004, Guanyu Feng, Chen Zhang 0001, Jiaao He, Shengqi Chen 0001, Jidong Zhai |
ICS | 7 |
| 2022 | TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs
Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001 |
OSDI | 7 |
| 2021 | PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective FeedbackabstractSecond language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency. Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002 |
CHI | 6 |
| 2021 | RisGraph: A Real-Time Streaming System for Evolving Graphs to Support Sub-millisecond Per-update Analysis at Millions Ops/sabstractEvolving graphs in the real world are large-scale and constantly changing, as hundreds of thousands of updates may come every second. Monotonic algorithms such as Reachability and Shortest Path are widely used in real-time analytics to gain both static and temporal insights and can be accelerated by incremental computing. Existing streaming systems adopt the incremental computing model and achieve either low latency or high throughput, but not both. However, both high throughput and low latency are required in real scenarios such as financial fraud detection. This paper presents RisGraph, a real-time streaming system that provides low-latency analysis for each update with high throughput. RisGraph addresses the challenge with localized data access and inter-update parallelism. We propose a data structure named Indexed Adjacency Lists and use sparse arrays and Hybrid Parallel Mode to enable localized data access. To achieve inter-update parallelism, we propose a domain-specific concurrency control mechanism based on the classification of safe and unsafe updates. Experiments show that RisGraph can ingest millions of updates per second for graphs with several hundred million vertices and billions of edges, and the P999 processing time latency is within 20 milliseconds. RisGraph achieves orders-of-magnitude improvement on throughput when analyses are executed for each update without batching and performs better than existing systems with batches of up to 20 million updates. Guanyu Feng, Zixuan Ma, Daixuan Li, Shengqi Chen 0001, Xiaowei Zhu 0001 |
SIGMOD Conference | 4 |
| 2021 | Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Tsinghua UniversityabstractIn this article we present our results from the SC19 Student Cluster Competition Reproducibility Challenge. The challenge entails reproducing the article entitled “Computing Planetary Interior Normal Modes with A Highly Parallel Polynomial Filtering Eigensolver” presented at SC'18, which proposes a parallel polynomial filtered Lanczos algorithm to directly calculate the planetary normal modes of heterogeneous planets. The proposed algorithm showed excellent performance with relatively low memory consumption and high parallel efficiency. In this work, we reproduce the scaling tests in that article on a cluster using Intel Cascade Lake architecture and use the proposed algorithm to illustrate specific normal modes of Mars. We compare the results obtained on our cluster with those in the original article. We also design a new metric to better analyze the results. In addition, we use the profiling tool Intel VTune Amplifier to explain our discoveries. Our results demonstrate that the given models show great scalability, which is similar to the original article. The required normal modes of Mars are also successfully calculated and visualized. Chen Zhang 0001, Chenggang Zhao, Jiaao He, Shengqi Chen 0001, Liyan Zheng 0001, Kezhao Huang, Jidong Zhai |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Visual-speech Synthesis of Exaggerated Corrective FeedbackabstractTo provide more discriminative feedback for the second language (L2) learners to better identify their mispronunciation, we propose a method for exaggerated visual-speech feedback in computer-assisted pronunciation training (CAPT). The speech exaggeration is realized by an emphatic speech generation neural network based on Tacotron, while the visual exaggeration is accomplished by ADC Viseme Blending, namely increasing Amplitude of movement, extending the phone's Duration and enhancing the color Contrast. User studies show that exaggerated feedback outperforms non-exaggerated version on helping learners with pronunciation identification and pronunciation improvement. Yaohua Bu, Shengqi Chen 0001, Jia Jia 0001, Kun Li 0003, Xiaobo Lu |
ACM Multimedia | 4 |