Mingcheng Chen

dblp:72/11186 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0003-2590-3196ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Emerging computing paradigms · 53% High-performance computing · 27% Parallel and multicore computing · 9%
Artificial intelligence
2 papers
Kernel, tree and ensemble methods · 69% Learning paradigms · 21% Language models and text generation · 10%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%
Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 100%

Topics — the 11 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Emerging computing paradigms
quantum computing
0.612022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
Emerging computing paradigms › quantum computing
quantum simulation
0.612022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
High-performance computing
supercomputing
0.612022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosting decision tree
0.512021
Task-wise Split Gradient Boosting Trees for Multi-center Diabetes Prediction · KDD 2021
Visualization and visual analytics
flow visualization
0.212016
Fast Coherent Particle Advection through Time-Varying Unstructured Flow Datasets · IEEE Trans. Vis. Comput. Graph. 2016
Performance modeling and evaluation
benchmarking
0.212022
Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Learning paradigms
multi-task learning
0.112021
Task-wise Split Gradient Boosting Trees for Multi-center Diabetes Prediction · KDD 2021
Distributed and cloud data management › parallel data processing
distributed matrix computation
0.112012
MadLINQ: large-scale distributed matrix computation for the cloud · EuroSys 2012
Parallel and multicore computing
data-parallel programming
0.112012
MadLINQ: large-scale distributed matrix computation for the cloud · EuroSys 2012
GPUs and heterogeneous computing
GPU computing
0.112016
Fast Coherent Particle Advection through Time-Varying Unstructured Flow Datasets · IEEE Trans. Vis. Comput. Graph. 2016
Parallel and multicore computing
parallel programming models
0.012012
MadLINQ: large-scale distributed matrix computation for the cloud · EuroSys 2012

Methods — techniques the papers use, named apart from their topics

multi-task learning · 1.0gradient boosting decision tree · 1.0parallel framework · 0.6multiple-precision fixed-point arithmetic · 0.6instruction scheduling · 0.6spatially coherent bundling · 0.5one-shot learning · 0.5latent attention · 0.5block advection · 0.5fault tolerance · 0.3data-parallel execution · 0.3
YearPublicationVenuePosition
2026 VDMPAGR: A vulnerability detection model based on pointer analysis and graph representation
Yukun Dong, Xiaoshan Liu, Mingcheng Chen, Yinzhou Feng
Inf. Softw. Technol.4
2025 Offline Model-based Optimization with Fisher Divergence Regularization
abstract
The goal of offline model-based optimization (MBO) is to seek designs that maximize a black-box function with only offline historical function evaluations, which has a wide range of applications. Since the black-box function cannot be queried directly, a surrogate model is usually learned to approximate the objective function. To overcome the challenge of distribution shift between the empirical offline data distribution and the underlying data distribution, the literature typically applies regularization to the surrogate. In this work, we propose Fisher Regularized Model (FiRM), a novel forward approach to offline MBO. FiRM decomposes a forward surrogate into an offline density model and an offset network, and applies gradient regularization on the offset to constrain the optimized designs to stay close to the offline data support, so as to make a proper trade-off between conservative estimation and generalization. In addition, from the perspective of energy-based model (EBM), we prove that the gradient regularization is equivalent to a Fisher-divergence constraint to the surrogate. Further theoretical analysis establishes connections between FiRM and existing work. Experimental results on five different tasks in a standard benchmark show that our method outperforms all the baselines, achieving the highest mean score and mean rank.
Mingcheng Chen, Minkai Xu, Yichen Zhu 0002, Weinan Zhang 0001, Yong Yu 0001
IJCNN1
2025 A search-and-fill strategy to code generation for complex software requirements
Yukun Dong, Lingjie Kong, Xiaoshan Liu, Mingcheng Chen
Inf. Softw. Technol.7
2024 Adaptive Retrieval-Based Gradient Planning for Offline Multi-context Model-Based Optimization
Mingcheng Chen, Weinan Zhang 0001, Yong Yu 0001
ICONIP (2)2
2023 ROMO: Retrieval-enhanced Offline Model-based Optimization
abstract
Data-driven black-box model-based optimization (MBO) problems arise in a great number of practical application scenarios, where the goal is to find a design over the whole space maximizing a black-box target function based on a static offline dataset. In this work, we consider a more general but challenging MBO setting, named constrained MBO (CoMBO), where only part of the design space can be optimized while the rest is constrained by the environment. A new challenge arising from CoMBO is that most observed designs that satisfy the constraints are mediocre in evaluation. Therefore, we focus on optimizing these mediocre designs in the offline dataset while maintaining the given constraints rather than further boosting the best observed design in the traditional MBO setting. We propose retrieval-enhanced offline model-based optimization (ROMO), a new derivable forward approach that retrieves the offline dataset and aggregates relevant samples to provide a trusted prediction, and use it for gradient-based optimization. ROMO is simple to implement and outperforms state-of-the-art approaches in the CoMBO setting. Empirically, we conduct experiments on a synthetic Hartmann (3D) function dataset, an industrial CIO dataset, and a suite of modified tasks in the Design-Bench benchmark. Results show that ROMO performs well in a wide range of constrained optimization tasks.
Mingcheng Chen, Hulei Fan, Hongqiao Gao, Yong Yu 0001, Zheng Tian 0002
DAI1
2022 A gradient boosting tree model for multi-department venous thromboembolism risk assessment with imbalanced data
abstract
Venous thromboembolism (VTE) is the world's third most common cause of vascular mortality and a serious complication from multiple departments. Risk assessment of VTE guides clinical intervention in time and is of great importance to in-hospital patients. Traditional VTE risk assessment methods based on scaling tools, which always require rules carefully designed by human experts, are difficult to apply to large-population scenarios since the manually designed rules are not guaranteed to be accurate to all populations. In contrast, with the development of the electronic health record (EHR) datasets, data-driven machine-learning-based risk assessment methods have proven superior predictability in many studies in recent years. This paper uses the gradient boosting tree model to study the VTE risk assessment problem with multi-department data. There exist two distinct characteristics of VTE data collected at the level of the entire hospital: its wide distribution and heterogeneity across multiple departments. To this end, we consider the prediction task over multiple departments as a multi-task learning process, and introduce the algorithm of a task-aware tree-based method TSGB to tackle the multi-task prediction problem. Although the introduction of multi-task learning improves overall across-department performance, we reveal the problem of task-wise performance decline while dealing with imbalanced VTE data volume. According to the analysis, we finally propose two variants of TSGB to alleviate the problems and further boost the prediction performance. Compared with state-of-the-art rule-based and multi-task tree-based methods, the experimental results show the proposed methods not only improve the overall across-department AUC performance effectively, but also ensure the improvement of performance over every single department prediction.
Handong Ma, Zhecheng Dong, Mingcheng Chen, Wenbo Sheng, Weinan Zhang 0001, Shaodian Zhang, Yong Yu 0001
J. Biomed. Informatics3
2022 Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight
abstract
Boson sampling is expected to be an important milestone that will demonstrate quantum computational advantage (or quantum supremacy). This work establishes the benchmarking of Gaussian boson sampling (GBS) with threshold detection based on the Sunway TaihuLight supercomputer. To achieve the best performance and provide a competitive scenario for future quantum computing studies, the selected simulation algorithm is fully optimized based on a set of innovative approaches, including a parallel framework with almost perfect load balance and an instruction-level optimizing scheme based on a shortest-path-based instruction scheduling. In addition, data precision is carefully processed by an integer-instruction-based and multiple-precision fixed-point implementation, including 128- and 256-bit precison mode, which can be appropriately selected based on an adaptive precision optimizing scheme. Based on these methods, a highly efficient parallel quantum sampling algorithm is designed. The largest run enables us to obtain one Torontonian function of a$100\times 100$submatrix from 50-photon GBS within 20 hours in 128-bit precision and 2 days in 256-bit precision. To our knowledge, this was the largest quantum computing simulation based on Boson Sampling by using modern supercomputers.
Lin Gan 0001, Mingcheng Chen, Yaojian Chen, Haitian Lu, Chao-Yang Lu, Jian-Wei Pan, Haohuan Fu, Guangwen Yang 0002
IEEE Trans. Parallel Distributed Syst.3
2021 Task-wise Split Gradient Boosting Trees for Multi-center Diabetes Prediction
abstract
Diabetes prediction is an important data science application in the social healthcare domain. There exist two main challenges in the diabetes prediction task: data heterogeneity since demographic and metabolic data are of different types, data insufficiency since the number of diabetes cases in a single medical center is usually limited. To tackle the above challenges, we employ gradient boosting decision trees (GBDT) to handle data heterogeneity and introduce multi-task learning (MTL) to solve data insufficiency. To this end, Task-wise Split Gradient Boosting Trees (TSGB) is proposed for the multi-center diabetes prediction task. Specifically, we firstly introduce task gain to evaluate each task separately during tree construction, with a theoretical analysis of GBDT's learning objective. Secondly, we reveal a problem when directly applying GBDT in MTL, i.e., the negative task gain problem. Finally, we propose a novel split method for GBDT in MTL based on the task gain statistics, named task-wise split, as an alternative to standard feature-wise split to overcome the mentioned negative task gain problem. Extensive experiments on a large-scale real-world diabetes dataset and a commonly used benchmark dataset demonstrate TSGB achieves superior performance against several state-of-the-art methods. Detailed case studies further support our analysis of negative task gain problems and provide insightful findings. The proposed TSGB method has been deployed as an online diabetes risk assessment software for early diagnosis.
Mingcheng Chen, Zhenghui Wang, Zhiyun Zhao, Weinan Zhang 0001, Xiawei Guo, Jian Shen 0003, Yanru Qu, Jieli Lu, Wei-Wei Tu, Yong Yu 0001, Yufang Bi, Guang Ning
KDD1
2021 Model-Based Offline Policy Optimization with Distribution Correcting Regularization
Jian Shen 0003, Mingcheng Chen, Zhengyu Yang 0002, Weinan Zhang 0001, Yong Yu 0001
ECML/PKDD (1)2
2018 A Dynamic Resource Overbooking Mechanism in Fog Computing
abstract
Fog Computing (FC - similarly edge computing) as new computing paradigm can support distributed domain-specific or area-specific applications with cloud-like quality of service (QoS). This promising paradigm thus can find its wide applications in various industrial scenarios and smart cities in which the resource requirements will be divided into peak-hour or non-peak-hour. To deal with such features of applications, a flexible resource allocation approach based on pricing model can be critical for the success of such paradigm. To the best of our knowledge, we have not seen such pricing based resource allocation approach ever been reported for FC scenarios. In this paper, we propose a novel pricing based dynamic resource allocation model through overbooking mechanism, and it is realized through three steps: 1) According to different QoS requirements of user tasks, methods of on-demand billing, daily billing, and auction billing are designed, in which we allow the resource to be overbooked; 2) For auction billing, we design an auction approach including pricing rule and winner determination rule. We prove that our auction approach guarantees individual rationality, computational efficiency, and truthfulness. 3) To overbook as much resource as possible with a high degree of QoS satisfaction of on-demand and daily billing, we overbook the resource based on a resource utilization prediction using neural network and service level agreement violation feedback. In the end, we validate the mechanism with real-world data trace. Experimental results show that our auction approach achieves desirable properties, and our dynamic resource overbooking mechanism maximizes the profit of nodes with a high degree of QoS satisfaction of on-demand and daily billing and a high resource utilization prediction accuracy rate.
Fuming Zhang, Zhiqing Tang, Mingcheng Chen, Weijia Jia 0001
MASS3
2016 Latent Attention For If-Then Program Synthesis
abstract
Automatic translation from natural language descriptions into programs is a long-standing challenging problem. In this work, we consider a simple yet important sub-problem: translation from textual descriptions to If-Then programs. We devise a novel neural network architecture for this task which we train end-to-end. Specifically, we introduce Latent Attention, which computes multiplicative weights for the words in the description in a two-stage process with the goal of better leveraging the natural language structures that indicate the relevant parts for predicting program elements. Our architecture reduces the error rate by 28.57% compared to prior art. We also propose a one-shot learning scenario of If-Then program synthesis and simulate it with our existing dataset. We demonstrate a variation on the training procedure for this scenario that outperforms the original procedure, significantly closing the gap to the model trained with all data.
Chang Liu 0021, Richard Shin, Mingcheng Chen, Dawn Song
NIPS4
2016 Fast Coherent Particle Advection through Time-Varying Unstructured Flow Datasets
abstract
Tracing the paths of collections of particles through a flow field is a key step for many flow visualization and analysis methods. When a flow field is interpolated from the nodes of an unstructured mesh, the process of advecting a particle must first find which cell in the unstructured mesh contains the particle. Since the paths of nearby particles often diverge, the parallelization of particle advection quickly leads to incoherent memory accesses of the unstructured mesh. We have developed a new block advection GPU approach that reorganizes particles into spatially coherent bundles as they follow their advection paths, which greatly improves memory coherence and thus shared-memory GPU performance. This approach works best for flows that meet the CFL criterion on unstructured meshes of uniformly sized elements, small enough to fit at least two timesteps in GPU memory.
Mingcheng Chen, Shawn C. Shadden, John C. Hart
IEEE Trans. Vis. Comput. Graph.1
2012 MadLINQ: large-scale distributed matrix computation for the cloud
abstract
The computation core of many data-intensive applications can be best expressed as matrix computations. The MadLINQ project addresses the following two important research problems: the need for a highly scalable, efficient and fault-tolerant matrix computation system that is also easy to program, and the seamless integration of such specialized execution engines in a general purpose data-parallel computing system.
Zhengping Qian, Xiuwei Chen, Nanxi Kang, Mingcheng Chen, Thomas Moscibroda, Zheng Zhang 0001
EuroSys4