Shengqi Chen 0001

dblp:156/3418 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-2310-5249ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 7 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Parallel MIP Solving with Dynamic Task Decomposition
Peng Lin 0005, Shaowei Cai 0001, Mengchuan Zou, Shengqi Chen 0001
CP4
2025 MEPipe: Democratizing LLM Training with Memory-Efficient Slice-Level Pipeline Scheduling on Cost-Effective Accelerators
abstract
The training of large language models (LLMs) typically needs costly GPUs, such as NVIDIA A100 or H100. They possess substantial high-bandwidth on-chip memory and rapid interconnects like NVLinks. The exorbitant expenses associated with LLM training pose not just an economic challenge but also a societal one, as it restricts the ability to train LLMs from scratch to a selected few organizations.
Zhenbo Sun, Shengqi Chen 0001, Yuanwei Wang, Jian Sha, Guanyu Feng
EuroSys2
2025 Iterative Design of a Teaching Assistant Training Program in Computer Science Using the Agile Method
abstract
Facing soaring enrollment and disruptive educational technologies, computing education increasingly relies on the contributions of teaching assistants (TAs), hence the critical importance of high-quality TA training. However, the design and implementation of TA training in computer science face substantial barriers, such as the lack of experienced TA trainers and the scarcity of relevant training materials.
Runda Liu, Shengqi Chen 0001, Songjie Niu, Yuchun Ma, Xiaofeng Tang
SIGCSE (1)2
2025 HypeReca: Distributed Heterogeneous In-Memory Embedding Database for Training Recommender Models
Jiaao He, Shengqi Chen 0001, Kezhao Huang, Jidong Zhai
USENIX ATC2
2024 AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and Partitioning
abstract
Large language models (LLMs) have demonstrated powerful capabilities, requiring huge memory with their increasing sizes and sequence lengths, thus demanding larger parallel systems. The broadly adopted pipeline parallelism introduces even heavier and unbalanced memory consumption. Recomputation is a widely employed technique to mitigate the problem but introduces extra computation overhead.
Zhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng, Shengqi Chen 0001, Haojie Wang 0004
ASPLOS (3)5
2024 POSTER: Pattern-Aware Sparse Communication for Scalable Recommendation Model Training
abstract
Recommendation models are an important category of deep learning models whose size is growing enormous. They consist of a sparse part with TBs of memory footprint and a dense part that demands PFLOPs of computing capability to train. Unfortunately, the high sparse communication cost to re-organize data for different parallel strategies of the two parts impedes the scalability in training.
Jiaao He, Shengqi Chen 0001, Jidong Zhai
PPoPP2
2023 TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs
abstract
Out-of-core systems rely on high-performance cache sub-systems to reduce the number of I/O operations. Although the page cache in modern operating systems enables transparent access to memory and storage devices, it suffers from efficiency and scalability issues on cache misses, forcing out-of-core systems to design and implement their own cache components, which is a non-trivial task. This study proposes TriCache, a cache mechanism that enables in-memory programs to efficiently process out-of-core datasets without requiring any code rewrite. It provides a virtual memory interface on top of the conventional block interface to simultaneously achieve user transparency and sufficient out-of-core performance. A multi-level block cache design is proposed to address the challenge of per-access address translations required by a memory interface. It can exploit spatial and temporal localities in memory or storage accesses to render storage-to-memory address translation and page-level concurrency control adequately efficient for the virtual memory interface. Our evaluation shows that in-memory systems operating on top of TriCache can outperform Linux OS page cache by more than one order of magnitude, and can deliver performance comparable to or even better than that of corresponding counterparts designed specifically for out-of-core scenarios.
Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001
ACM Trans. Storage7
2022 Efficiently emulating high-bitwidth computation with low-bitwidth hardware
abstract
Domain-Specific Accelerators (DSAs) are being rapidly developed to support high-performance domain-specific computation. Although DSAs provide massive computation capability, they often only support limited native data types. To mitigate this problem, previous works have explored software emulation for certain data types, which provides some compensation for hardware limitations. However, how to efficiently design more emulated data types and choose a high-performance one without hurting correctness or precision for a given application still remains an open problem.
Zixuan Ma, Haojie Wang 0004, Guanyu Feng, Chen Zhang 0001, Jiaao He, Shengqi Chen 0001, Jidong Zhai
ICS7
2022 TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs
Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001
OSDI7
2021 PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
abstract
Second language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency.
Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002
CHI6
2021 RisGraph: A Real-Time Streaming System for Evolving Graphs to Support Sub-millisecond Per-update Analysis at Millions Ops/s
abstract
Evolving graphs in the real world are large-scale and constantly changing, as hundreds of thousands of updates may come every second. Monotonic algorithms such as Reachability and Shortest Path are widely used in real-time analytics to gain both static and temporal insights and can be accelerated by incremental computing. Existing streaming systems adopt the incremental computing model and achieve either low latency or high throughput, but not both. However, both high throughput and low latency are required in real scenarios such as financial fraud detection. This paper presents RisGraph, a real-time streaming system that provides low-latency analysis for each update with high throughput. RisGraph addresses the challenge with localized data access and inter-update parallelism. We propose a data structure named Indexed Adjacency Lists and use sparse arrays and Hybrid Parallel Mode to enable localized data access. To achieve inter-update parallelism, we propose a domain-specific concurrency control mechanism based on the classification of safe and unsafe updates. Experiments show that RisGraph can ingest millions of updates per second for graphs with several hundred million vertices and billions of edges, and the P999 processing time latency is within 20 milliseconds. RisGraph achieves orders-of-magnitude improvement on throughput when analyses are executed for each update without batching and performs better than existing systems with batches of up to 20 million updates.
Guanyu Feng, Zixuan Ma, Daixuan Li, Shengqi Chen 0001, Xiaowei Zhu 0001
SIGMOD Conference4
2021 Critique of "Planetary Normal Mode Computation: Parallel Algorithms, Performance, and Reproducibility" by SCC Team From Tsinghua University
abstract
In this article we present our results from the SC19 Student Cluster Competition Reproducibility Challenge. The challenge entails reproducing the article entitled “Computing Planetary Interior Normal Modes with A Highly Parallel Polynomial Filtering Eigensolver” presented at SC'18, which proposes a parallel polynomial filtered Lanczos algorithm to directly calculate the planetary normal modes of heterogeneous planets. The proposed algorithm showed excellent performance with relatively low memory consumption and high parallel efficiency. In this work, we reproduce the scaling tests in that article on a cluster using Intel Cascade Lake architecture and use the proposed algorithm to illustrate specific normal modes of Mars. We compare the results obtained on our cluster with those in the original article. We also design a new metric to better analyze the results. In addition, we use the profiling tool Intel VTune Amplifier to explain our discoveries. Our results demonstrate that the given models show great scalability, which is similar to the original article. The required normal modes of Mars are also successfully calculated and visualized.
Chen Zhang 0001, Chenggang Zhao, Jiaao He, Shengqi Chen 0001, Liyan Zheng 0001, Kezhao Huang, Jidong Zhai
IEEE Trans. Parallel Distributed Syst.4
2020 Visual-speech Synthesis of Exaggerated Corrective Feedback
abstract
To provide more discriminative feedback for the second language (L2) learners to better identify their mispronunciation, we propose a method for exaggerated visual-speech feedback in computer-assisted pronunciation training (CAPT). The speech exaggeration is realized by an emphatic speech generation neural network based on Tacotron, while the visual exaggeration is accomplished by ADC Viseme Blending, namely increasing Amplitude of movement, extending the phone's Duration and enhancing the color Contrast. User studies show that exaggerated feedback outperforms non-exaggerated version on helping learners with pronunciation identification and pronunciation improvement.
Yaohua Bu, Shengqi Chen 0001, Jia Jia 0001, Kun Li 0003, Xiaobo Lu
ACM Multimedia4