EDBT 2026 Demo / reviewers in the wild / expert
Mingxuan He
dblp:220/3788
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2025
0009-0000-8646-5279ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 41% Memory systems · 24% Hardware accelerators and domain-specific architectures · 24% | |
| Databases, data mining, and information retrieval
1 paper |
Data stream processing · 100% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data stream processing
stream join |
0.8 | 1 | 2024 | Low-Latency Adaptive Distributed Stream Join System Based on a Flexible Join Model · Proc. ACM Manag. Data 2024 |
Distributed systems › stream processing
distributed stream processing |
0.8 | 1 | 2024 | Low-Latency Adaptive Distributed Stream Join System Based on a Flexible Join Model · Proc. ACM Manag. Data 2024 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine Learning · MICRO 2020 |
Memory systems
processing-in-memory |
0.4 | 1 | 2020 | Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine Learning · MICRO 2020 |
Performance modeling and evaluation
queueing analysis |
0.2 | 1 | 2024 | Low-Latency Adaptive Distributed Stream Join System Based on a Flexible Join Model · Proc. ACM Manag. Data 2024 |
Methods — techniques the papers use, named apart from their topics
queuing theory · 1.5adaptive scheduling · 1.5multiply-accumulate units · 0.9interleaved matrix layout · 0.9command ganging · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Gaussian Temporal Based Graph Convolutional Network for Traffic Operation Flow ForecastingabstractHigh-Quality traffic flow data may provide planners with a basis for designing road capacity, pavement and intersection control, etc., thus assisting in the construction of a more rational traffic network. This study developed a new Gaussian Temporal Network Module, which is a module based on a Gaussian process that utilizes a kernel Gaussian convolution kernel to assist in the extraction of temporal features. Which is aim to solve the common gradient vanishing and gradient explosion problems in temporal neural networks Further, this study combines the GTNM with the GCN module as a GT-GCN model and tests its performance on a publicly available dataset. Experiment results show that GT-GCN demonstrates superior performance, attaining state-of-the-art or second best outcomes across different test dataset. Several of these results surpass baseline benchmarks by more than 10%, which then well illustrates the effectiveness and stability of GT-GCN model. Shulan Guo, Shangda Xiao, Mingxuan He, Hequn Xian, Zesheng Cheng |
ICCCN | 4 |
| 2024 | Low-Latency Adaptive Distributed Stream Join System Based on a Flexible Join ModelabstractStream join is a fundamental operation in stream processing and has attracted extensive research due to its large resource consumption and serious impact on system performance. As the theoretical basis of stream join systems, the stream join model greatly affects system performance. State-of-the-art stream join models either consume too much computing resources or too much storage resources, thus resulting in lower throughput or higher latency. In this paper, we propose a new stream join model for processing arbitrary join predicates, called CoModel, which offers a flexible trade-off between memory and computing resource consumption. More importantly, CoModel can achieve the minimum sum of the number of store operations and join operations among all existing join models, and thus can achieve the lowest latency and highest throughput when the overheads associated with the local stream join for each input tuple are approximately constant. We give a trade-off strategy for CoModel and theoretically prove its performance advantages based on queuing theory. Furthermore, we design and implement an adaptive distributed stream join system, CoStream, based on CoModel. CoStream can adaptively adjust its structure according to resource constraints and statistics of input data. We conduct extensive experiments for CoStream to evaluate its performance and adaptivity, and the results show that CoStream has the lowest latency and highest throughput in various scenarios. De-Cheng Zuo, Zhan Zhang 0006, Yanjun Shu, Mingxuan He |
Proc. ACM Manag. Data | 6 |
| 2022 | Booster: An Accelerator for Gradient Boosting Decision Trees Training and InferenceabstractRecent breakthroughs in machine learning (ML) have sparked hardware innovation for efficient execution of the emerging ML workloads. For instance, due to recent refine-ments and high-performance implementations, well-established gradient boosting decision tree (GBT) models (e.g., XGBoost) have demonstrated their dominance in commercially-important contexts, such as table-based datasets (e.g., relational databases and spreadsheets). Unfortunately, GBT training and inference are time-consuming (e.g., several hours of training for large datasets). Despite their importance, GBTs have not been targeted for hardware acceleration as much as neural networks. We propose Booster, a novel accelerator for GBTs based on their unique characteristics. We observe that the dominant steps of GBT training and inference (accounting for 90-98% of time) involve simple, fine-grained, independent operations on small-footprint data structures (e.g., histograms and shallow trees) - i.e., GBT is on-chip memory bandwidth-bound. Unfortunately, existing multicores and GPUs do not support massively-parallel data structure accesses that are irregular and data-dependent. By employing a scalable sea-of-small-SRAMs approach and an SRAM bandwidth-preserving mapping of data record fields to the SRAMs called group-by-field mapping, Booster achieves significantly more parallelism (e.g., 3200-way parallelism) than multicores and GPUs. In addition, Booster employs a redun-dant data representation that significantly lowers the memory bandwidth demand. Our simulations reveal that Booster achieves 11.4x and 6.4x speedups for training, and 45x and 22x (21x and 11x) speedups for offline (online) inference, over an ideal 32-core multicore and an ideal GPU, respectively. Based on ASIC synthesis of FPGA-validated RTL using 45 nm technology, we estimate a Booster chip to occupy 60 mm2of area and dissipate 23 W when operating at 1-G Hz clock speed. Mingxuan He, Mithuna Thottethodi, T. N. Vijaykumar |
IPDPS | 1 |
| 2021 | Attention cutting and padding learning for fine-grained image recognition
Xiaolin Duan, Xiangyan Zeng, Mingxuan He |
Multim. Tools Appl. | 5 |
| 2020 | Newton: A DRAM-maker's Accelerator-in-Memory (AiM) Architecture for Machine LearningabstractAdvances in machine learning (ML) have ignited hardware innovations for efficient execution of the ML models many of which are memory-bound (e.g., long short-term memories, multi-level perceptrons, and recurrent neural networks). Specifically, inference using these ML models with small batches, as would be the case at the Cloud edge, has little reuse of the large filters and is deeply memory-bound. Simultaneously, processing-in or -near memory (PIM or PNM) is promising unprecedented high-bandwidth connection between compute and memory. Fortunately, the memory-bound ML models are a good fit for PIM. We focus on digital PIM which provides higher bandwidth than PNM and does not incur the reliability issues of analog PIM. Previous PIM and PNM approaches advocate full processor cores which do not conform to PIM's severe area and power constraints. We describe Newton, a major DRAM maker's upcoming accelerator-in-memory (AiM) product for machine learning, which makes the following contributions: (1) To satisfy PIM's area constraints, Newton (a) places a minimal compute of only multiply-accumulate units and buffers in the DRAM which avoids the full-core area and power overheads of previous work and thus makes PIM feasible for the first time, and (b) employs a DRAM-like interface for the host to issue commands to the PIM compute. The PIM compute is rate-matched to the internal DRAM bandwidth and employs a non-intuitive, global input vector buffer shared by the entire channel to capture input reuse while amortizing buffer area cost. To the host, Newton's interface is indistinguishable from regular DRAM without any offloading overheads and PIM/non-PIM mode switching, and with the same deterministic latencies even for floating-point commands. (2) To prevent the PIM-host interface from becoming a bottleneck, we include three optimizations: commands which gang multiple compute operations both within a bank and across banks; complex, multi-step compute commands - both of which save critical command bandwidth; and targeted reduction of tFAWoverhead. (3) To capture output vector reuse with reasonable buffering, Newton employs an unusually-wide interleaved layout for the matrix. Our simulations running state-of-the-art neural networks show that building on a realistic HBM2E-like DRAM, Newton achieves 10x and 54x average speedup over a non-PIM system with infinite compute that perfectly uses the external DRAM bandwidth and a realistic GPU, respectively. Mingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong, Seho Kim, Il Park 0001, Mithuna Thottethodi, T. N. Vijaykumar |
MICRO | 1 |