EDBT 2026 Demo / reviewers in the wild / expert
Jiahui Peng
dblp:297/5645
· DBLP profile ↗
9ranked-venue papers
3as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 49% Efficient and distributed learning · 31% Representation and self-supervised learning · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Hardware reliability and fault tolerance · 50% Embedded and real-time systems · 50% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.7 | 2 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
1.7 | 2 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.9 | 1 | 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language Models · ACL (1) 2025 |
Machine learning › Representation and self-supervised learning › pre-training › data-centric pre-training
data selection for pre-training |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Natural language and speech › Language models and text generation › large language model training
pretraining data selection |
0.9 | 1 | 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration · ACL (1) 2025 |
Embedded and real-time systems
cyber-physical system platforms |
0.9 | 1 | 2025 | FT-MUX: A Fault-Tolerant Microfluidic Multiplexer Design · DAC 2025 |
Hardware reliability and fault tolerance
fault-tolerant design |
0.9 | 1 | 2025 | FT-MUX: A Fault-Tolerant Microfluidic Multiplexer Design · DAC 2025 |
Natural language and speech › Language models and text generation › hallucination mitigation
hallucination correction |
0.8 | 1 | 2024 | VIGC: Visual Instruction Generation and Correction · AAAI 2024 |
Natural language and speech › Language models and text generation › instruction tuning
instruction data generation |
0.8 | 1 | 2024 | VIGC: Visual Instruction Generation and Correction · AAAI 2024 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.8 | 1 | 2024 | VIGC: Visual Instruction Generation and Correction · AAAI 2024 |
Methods — techniques the papers use, named apart from their topics
multi-dimensional data selection · 0.9multi-actor collaboration · 0.9curriculum learning · 0.9binary constant weight code · 0.9visual instruction generation · 0.8iterative update mechanism · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Built-In Self-Test for Locating Leakage Defects on Continuous-Flow Microfluidic Chips
Jiahui Peng, Mengchu Li, Tsun-Ming Tseng, Ulf Schlichtmann |
ASP-DAC | 1 |
| 2026 | MedP-CLIP: Medical CLIP with region-aware prompt integration
Jiahui Peng, He Yao, Yanzhou Su, Sibo Ju, Hongchun Lu, Xue Li 0008, Lincheng Jiang, Min Zhu 0005, Junlong Cheng |
Medical Image Anal. | 1 |
| 2025 | Efficient Pretraining Data Selection for Language Models via Multi-Actor CollaborationabstractEfficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https://github.com/Relaxed-System-Lab/multi-actor-data-selection. Tianyi Bai, Ling Yang 0006, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Lijun Wu 0003, Jiantao Qiu, Wentao Zhang 0001, Binhang Yuan, Conghui He |
ACL (1) | 6 |
| 2025 | Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsabstractXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Qiu Jiantao, Chi Zhang, Ying Qian, Conghui He. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Xinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang, Tianyi Bai, Xingjian Wei, Jiantao Qiu, Conghui He |
ACL (1) | 2 |
| 2025 | FT-MUX: A Fault-Tolerant Microfluidic Multiplexer DesignabstractContinuous-flow microfluidic chips are multilayered miniaturized platforms to manipulate small volumes of fluids with valves. There are two types of channels on a chip: flow channels for the reaction of fluids, and control channels for the actuation of valves. Multiplexers (MUXes) are essential microfluidic components for individually addressing many flow channels with few control channels. As the integration scale of microfluidic chips increases, the reliability of MUXes becomes a critical concern, as a single defective control channel in a MUX will affect a large part of the flow channels addressed by the MUX. This paper formally analyzes and identifies the design rules for a MUX to tolerate n defective control channels, and model the fault-tolerant MUX (FT-MUX) design problem as a binary constant weight code problem to minimize resource overheads. We demonstrate that FT-MUX improves resource efficiency by up to hundreds of times compared to the conventional fault-tolerant design method. Besides, given no less than 10 control channels, FT-MUX tolerates at least one defective control channel and addresses even more flow channels with equal or fewer resources than a standard MUX. The advantages become more significant as the integration scale increases. Mengchu Li, Jiahui Peng, Tsun-Ming Tseng, Ulf Schlichtmann |
DAC | 2 |
| 2024 | VIGC: Visual Instruction Generation and CorrectionabstractThe integration of visual encoders and large language models (LLMs) has driven recent progress in multimodal large language models (MLLMs). However, the scarcity of high-quality instruction-tuning data for vision-language tasks remains a challenge. The current leading paradigm, such as LLaVA, relies on language-only GPT-4 to generate data, which requires pre-annotated image captions and detection bounding boxes, suffering from understanding image details. A practical solution to this problem would be to utilize the available multimodal large language models to generate instruction data for vision-language tasks. However, it's worth noting that the currently accessible MLLMs are not as powerful as their LLM counterparts, as they tend to produce inadequate responses and generate false information. As a solution for addressing the current issue, this paper proposes the Visual Instruction Generation and Correction (VIGC) framework that enables multimodal large language models to generate instruction-tuning data and progressively enhance its quality on-the-fly. Specifically, Visual Instruction Generation (VIG) guides the vision-language model to generate diverse instruction-tuning data. To ensure generation quality, Visual Instruction Correction (VIC) adopts an iterative update mechanism to correct any inaccuracies in data produced by VIG, effectively reducing the risk of hallucination. Leveraging the diverse, high-quality data generated by VIGC, we finetune mainstream models and validate data quality based on various evaluations. Experimental results demonstrate that VIGC not only compensates for the shortcomings of language-only data generation methods, but also effectively enhances the benchmark performance. The models, datasets, and code are available at https://opendatalab.github.io/VIGC Bin Wang 0065, Fan Wu 0006, Jiahui Peng, Huaping Zhong, Pan Zhang 0001, Xiaoyi Dong, Wei Li 0320, Jiaqi Wang 0003, Conghui He |
AAAI | 4 |
| 2024 | Triple Temporal Vision Transformer for the Coverage Classification with Multi-Temporal Polsar ImagesabstractMulti-temporal SAR and Polarimetric SAR (PolSAR) images can provide the scattering change characteristics caused by vegetation growth to help the classifier capture phenological characteristics. To effectively utilize multi-temporal PolSAR data, a triple temporal vision transformer (TriTempoViT) model is proposed to capture correlation from multi-dimensional features. The method uses a three-branch network architecture to extract spatial-temporal, spatial-polarimetric, and temporal-polarization features respectively, and then the features from the three branches will be integrated into the vision transformer (ViT) for information interaction. Additionally, a 3D channel-spatial attention module (3DCSAM) is tailored to automatically weight the importance of the multi-dimensional feature maps. Moreover, a temporal interaction feature extraction module (TIFEM) is designed to comprehensively consider correlations between different temporal sequences. Compared with the recently developed state-of-the-art approach, the proposed method can improve OA of the classification accuracy in a Radarsat-2 dataset from Flevoland by about 0.84%, which proves the effectiveness of the proposed method. Jiahui Peng, Dapeng Tao, Carlos López-Martínez |
IGARSS | 1 |
| 2024 | Achieving fair and accountable data trading for educational multimedia data based on blockchain
Xianxian Li, Jiahui Peng, Shiqi Gao, Zhenkui Shi, Chunpei Li |
Wirel. Networks | 2 |
| 2021 | Achieving Fair and Accountable Data Trading Scheme for Educational Multimedia Data Based on Blockchain
Xianxian Li, Jiahui Peng, Zhenkui Shi, Chunpei Li |
QSHINE | 2 |