EDBT 2026 Demo / reviewers in the wild / expert
Sen Peng
dblp:72/7845
· DBLP profile ↗
14ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GA-Avatar: Geometry-Aware Human Avatar Modeling Based on Gaussian Splatting From Monocular VideosabstractThe reconstruction of realistic 3-D digital humans from monocular videos plays a crucial role in socially intelligent systems, enabling human–machine interaction, behavioral analysis, and immersive communication in virtual environments. However, achieving geometrically consistent and visually faithful avatars under limited views remains challenging. Existing Gaussian-based approaches lack effective geometric constraints and resulting in structural drift and appearance degradation in novel poses. To address the above problems, we propose GA-Avatar, a dual-branch reconstruction network via decoupled geometry and appearance modeling, which improves the geometric accuracy as well as rendering quality. First, we construct a high-resolution mesh in canonical space, bind Gaussian primitives to mesh vertices, and achieve personalized geometric adjustments by optimizing Gaussian positions. Based on canonical features of the Gaussians, we design a dual-branch modeling network for geometry and appearance, learning vertex deformation and color features, respectively. To improve the modeling capabilities of nonrigid deformations, we introduce pose-dependent deformation modules in both the geometry and appearance branches. In the optimization stage, we proposed photometric loss and geometric optimization based on depth and normal constraints, which effectively enhanced the structure preservation and appearance consistency of the model in different poses. We conduct rich experiments and comparisons, and GA-Avatar outperforms existing reconstruction methods on multiple datasets. GA-Avatar provides a foundation for computational modeling of human appearance and behavior in intelligent systems. Sen Peng, Zhiyang Deng, Rong Jin 0003 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | Buffer Prospector: Discovering and Exploiting Untapped Buffer Resources in Many-Core DNN AcceleratorsabstractIn large-scale DNN inference accelerators, the many-core architecture has emerged as a predominant design, with layer-pipeline (LP) mapping being a mainstream mapping approach. However, our experimental findings and theoretical justifications uncover a hardware-independent and prevalent flaw in employing layer-pipeline mapping on many-core accelerators: a significant underutilization of buffer space across numerous cores, indicating substantial potential for optimization. Building on this discovery, we develop a universal and efficient buffer allocation strategy, BufferProspector, which includes a Buffer Requirement Calculator and Buffer Allocator, to capitalize on these unused buffers, addressing the timing mismatch challenge inherent in LP mapping. Compared to the state-of-the-art (SOTA) open-source LP mapping framework Tangram, BufferProspector averages a simultaneous increase in energy efficiency and performance by 1.44× and 2.26×, respectively. Moreover, we conduct some case studies on architecture and mapping. BufferProspector will be open-sourced. Jingwei Cai, Mingyu Gao 0001, Sen Peng, Zuotong Wu, Guiming Shi, Kaisheng Ma |
DAC | 4 |
| 2025 | SoMa: Identifying, Exploring, and Understanding the DRAM Communication Scheduling Space for DNN AcceleratorsabstractModern Deep Neural Network (DNN) accelerators are equipped with increasingly larger on-chip buffers to provide more opportunities to alleviate the increasingly severe DRAM bandwidth pressure. However, most existing research on buffer utilization still primarily focuses on single-layer dataflow scheduling optimization. As buffers grow large enough to accommodate most single-layer weights in most networks, the impact of single-layer dataflow optimization on DRAM communication diminishes significantly. Therefore, developing new paradigms that fuse multiple layers to fully leverage the increasingly abundant onchip buffer resources to reduce DRAM accesses has become particularly important, yet remains an open challenge.To address this challenge, we first identify the optimization opportunities in DRAM communication scheduling by analyzing the drawbacks of existing works on the layer fusion paradigm and recognizing the vast optimization potential in scheduling the timing of data prefetching from and storing to DRAM. To fully exploit these optimization opportunities, we develop a Tensor-centric Notation and its corresponding parsing method to represent different DRAM communication scheduling schemes and depict the overall space of DRAM communication scheduling. Then, to thoroughly and efficiently explore the space of DRAM communication scheduling for diverse accelerators and workloads, we develop an end-to-end scheduling framework, SoMa, which has already been developed into a compiler for our commercial accelerator product. Compared with the state-of-the-art (SOTA) Cocco framework, SoMa achieves, on average, a 2.11× performance improvement and a 37.3% reduction in energy cost simultaneously. Then, we leverage SoMa to study optimizations for LLM, perform design space exploration (DSE), and analyze the DRAM communication scheduling space through a practical example, yielding some interesting insights. Moreover, SoMa has been open-sourced at https://github.com/SET-Scheduling-Project/SoMa-HPCA2025. Jingwei Cai, Mingyu Gao 0001, Sen Peng, Zuotong Wu, Kaisheng Ma |
HPCA | 4 |
| 2025 | CAT: Contrastive Adversarial Training for Evaluating the Robustness of Protective Perturbations in Latent Diffusion ModelsabstractLatent diffusion models have recently demonstrated superior capabilities in many downstream image synthesis tasks. However, customization of latent diffusion models using unauthorized data can severely compromise the privacy and intellectual property rights of data owners. Adversarial examples as protective perturbations have been developed to defend against unauthorized data usage by introducing imperceptible noise to customization samples, preventing diffusion models from effectively learning them. In this paper, we first reveal that the primary reason adversarial examples are effective as protective perturbations in latent diffusion models is the distortion of their latent representations, as demonstrated through qualitative and quantitative experiments. We then propose the Contrastive Adversarial Training (CAT) utilizing lightweight adapters as an adaptive attack against these protection methods, highlighting their lack of robustness. Extensive experiments demonstrate that our CAT method significantly reduces the effectiveness of protective perturbations in customization, urging the community to reconsider and improve the robustness of existing protective perturbations. The code is available at https://github.com/senp98/CAT. Sen Peng, Jianfei He, Jijia Yang, Xiaohua Jia |
ICML | 1 |
| 2025 | Improve Fluency Of Neural Machine Translation Using Large Language ModelsabstractLarge language models (LLMs) demonstrate significant capabilities in many natural language processing. However, their performance in machine translation is still behind the models that are specially trained for machine translation with an encoder-decoder architecture. This paper investigates how to improve neural machine translation (NMT) with LLMs. Our proposal is based on an empirical insight that NMT gets worse fluency than human translation. We propose to use LLMs to enhance the fluency of NMT’s generation by integrating a language model at the target side. we use contrastive learning to constrain fluency so that it does not exceed the LLMs. Our experiments on three language pairs show that this method can improve the performance of NMT. Our empirical analysis further demonstrates that this method improves the fluency at the target side. Our experiments also show that some straightforward post-processing methods using LLMs, such as re-ranking and refinement, are not effective. Jianfei He, Wenbo Pan 0001, Jijia Yang, Sen Peng, Xiaohua Jia |
MTSummit (1) | 4 |
| 2025 | RMAvatar: Photorealistic human avatar reconstruction from monocular video based on rectified mesh-embedded GaussiansabstractWe introduce RMAvatar, a novel human avatar representation with Gaussian splatting embedded on mesh to learn clothed avatar from a monocular video. We utilize the explicit mesh geometry to represent motion and shape of a virtual human and implicit appearance rendering with Gaussian Splatting. Our method consists of two main modules: Gaussian initialization module and Gaussian rectification module. We embed Gaussians into triangular faces and control their motion through the mesh, which ensures low-frequency motion and surface deformation of the avatar. Due to the limitations of LBS formula, the human skeleton is hard to control complex non-rigid transformations. We then design a pose-related Gaussian rectification module to learn fine-detailed non-rigid deformations, further improving the realism and expressiveness of the avatar. We conduct extensive experiments on public datasets, and RMAvatar shows state-of-the-art performance on both rendering quality and quantitative evaluations. Please see our project page at https://rm-avatar.github.io . Sen Peng, Weixing Xie, Xiaohu Guo, Zhonggui Chen, Baorong Yang |
Graph. Model. | 1 |
| 2025 | 4D Gaussian Splatting for high-fidelity dynamic reconstruction of single-view scenes
Weixing Xie, Sen Peng, Yihang Fu, Wentao Fan 0001, Baorong Yang |
Neurocomputing | 3 |
| 2025 | GenericAvatar: generic human modeling from monocular video based on mesh-guided Gaussians
Sen Peng, Yihang Fu, Runjie Miu, Tianyi Lv, Baorong Yang |
Vis. Comput. | 1 |
| 2024 | Gemini: Mapping and Architecture Co-exploration for Large-scale DNN Chiplet AcceleratorsabstractChiplet technology enables the integration of an increasing number of transistors on a single accelerator with higher yield in the post-Moore era, addressing the immense computational demands arising from rapid AI advancements. However, it also introduces more expensive packaging costs and costly Die-to-Die (D2D) interfaces, which require more area, consume higher power, and offer lower bandwidth than onchip interconnects. Maximizing the benefits and minimizing the drawbacks of chiplet technology is crucial for developing largescale DNN chiplet accelerators, which poses challenges to both architecture and mapping. Despite its importance in the post-Moore era, methods to address these challenges remain scarce. To bridge the gap, we first propose a layer-centric encoding method to encode Layer-Pipeline (LP) spatial mapping for largescale DNN inference accelerators and depict the optimization space of it. Based on it, we analyze the unexplored optimization opportunities within this space, which play a more crucial role in chiplet scenarios. Based on the encoding method and a highly configurable and universal hardware template, we propose an architecture and mapping co-exploration framework, Gemini, to explore the design and mapping space of large-scale DNN chiplet accelerators while taking monetary cost (MC), performance, and energy efficiency into account. Compared to the state-of-the-art (SOTA) Simba architecture with SOTA Tangram LP Mapping, Gemini's co-optimized architecture and mapping achieve, on average, 1.98 × performance improvement and 1.41 × energy efficiency improvement simultaneously across various DNNs and batch sizes, with only a 14.3% increase in monetary cost. Moreover, we leverage Gemini to uncover intriguing insights into the methods for utilizing chiplet technology in architecture design and mapping DNN workloads under chiplet scenarios. The Gemini framework is open-sourced at https://github.com/SETScheduling-Project/GEMINI-HPCA2024. Jingwei Cai, Zuotong Wu, Sen Peng, Zhanhong Tan, Guiming Shi, Mingyu Gao 0001, Kaisheng Ma |
HPCA | 3 |
| 2024 | Intellectual Property Protection of Diffusion Models via the Watermark Diffusion Process
Sen Peng, Yufei Chen 0001, Cong Wang 0001, Xiaohua Jia |
WISE (2) | 1 |
| 2023 | Inter-layer Scheduling Space Definition and Exploration for Tiled AcceleratorsabstractWith the continuous expansion of the DNN accelerator scale, inter-layer scheduling, which studies the allocation of computing resources to each layer and the computing order of all layers in a DNN, plays an increasingly important role in maintaining a high utilization rate and energy efficiency of DNN inference accelerators. However, current inter-layer scheduling is mainly conducted based on some heuristic patterns. The space of inter-layer scheduling has not been clearly defined, resulting in significantly limited optimization opportunities and a lack of understanding on different inter-layer scheduling choices and their consequences. Jingwei Cai, Zuotong Wu, Sen Peng, Kaisheng Ma |
ISCA | 4 |
| 2023 | AdaptChain: Adaptive Scaling Blockchain With Transaction DeduplicationabstractAlthough existing schemes improve blockchain throughput by allowing concurrent blocks to be appended to the blockchain, little attention has been devoted to adjusting blockchain throughput dynamically and deduplicating transactions between concurrent blocks. In this article, we propose AdaptChain, an adaptive scaling blockchain with transaction deduplication. When the transaction demand of users in the network is high, the blockchain expands to meet the demand; when the transaction demand is low, the blockchain shrinks to save communication and storage costs. Our transaction deduplication mechanism ensures that no duplicate transactions are added to the blockchain, thereby improving bandwidth utilization and achieving higher effective throughput. Besides, we randomly split the mining power of the system to achieve mining power load balancing and resist attacks. We formally analyze the blockchain security and implement the proposed prototype on Amazon EC2. Experimental results show that AdaptChain achieves dynamic and higher effective blockchain throughput. Jie Xu 0031, Qingyuan Xie, Sen Peng, Cong Wang 0001, Xiaohua Jia |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2023 | Intellectual property protection of DNN models
Sen Peng, Yufei Chen 0001, Jie Xu 0031, Zizhuo Chen, Cong Wang 0001, Xiaohua Jia |
World Wide Web (WWW) | 1 |
| 2019 | Probing glioblastoma and its microenvironment using single-nucleus and single-cell sequencingabstractSingle-cell (scSeq) and single-nucleus sequencing (snSeq) are powerful tools to investigate cancer genomics at single cell resolution. Multiple studies have recently illuminated intratumoral heterogeneity in glioblastoma, however, the majority focused on molecular complexity of tumor cells, without considering unexplored host cell types that contribute to the microenvironment around tumor. To address the glioblastoma microenvironment composition and potential tumor-host interactions, we performed deep coverage sequencing of freshly resected primary GBM patient tissue without implementing any tumor enrichment strategies. The sequencing resulted in 902 cells and 1186 nuclei, respectively, passing quality control and with low mitochondrial gene percentage. We customized reference transcriptome by listing gene transcript loci as exons to take into account immature RNA, which greatly improved the alignment rate for single-nucleus data. We applied Cell Ranger pipelines (Version 3.0.2) and Seurat package (Version 2.3.1) and discovered 10 clusters in both scSeq and snSeq. Pathway analysis of each cluster signature in scSeq data along with known GBM microenvironment cell signatures revealed glioma tumor population along with surrounding microglia/macrophages, astrocytes, pericytes, oligodendrocytes, T cells and endothelial cells. The analysis of snSeq was able to capture the majority of cell types from patient tissues (tumor and microenvironment cells), but interestingly presented different cell type composition in microenvironment cell types such as microglia/macrophages. Integrating single-cell and single-nucleus transcriptomic data using canonical correlation analysis facilitated a comparison of snSeq and scSeq, contrasting depiction for certain cell types (e.g. NKX6-2 gene in Oligodendrocytes). Differential analysis of pathways between tumor and microenvironment cells unveiled potentially rewired pathways such as double strand break repair pathway. Our results demonstrate the cellular diversity of brain tumor microenvironment and lay a foundation to further investigate the individual tumor and host cell transcriptomes that are influenced not only by their cell identity but also by their interaction with surrounding microenvironment. Sen Peng, Seungchan Kim, Harshil Dhruv, Sanhita Rath, Connor Vuong, Saumya Bollam, Jenny Eschbacher, Xishuang Dong, Shwetal Mehta, Nader Sanai, Michael E. Berens |
BIBM | 1 |