Meichen Dong

dblp:285/3538 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-9437-9290ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 61% Processor architecture and microarchitecture · 30% Hardware accelerators and domain-specific architectures · 9%
Artificial intelligence
1 paper
Generative modeling · 56% Deep learning architectures and training · 28% Vision and language · 17%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture
dataflow architecture
1.012026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
GPUs and heterogeneous computing › GPU computing › tensor cores
sparse tensor core
1.012026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
GPUs and heterogeneous computing › GPU computing
tensor cores
1.012026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
Machine learning › Generative modeling
diffusion model
0.912025
End-to-End Multi-Modal Diffusion Mamba · ICCV 2025
Machine learning › Generative modeling › diffusion model
multimodal diffusion model
0.912025
End-to-End Multi-Modal Diffusion Mamba · ICCV 2025
Hardware accelerators and domain-specific architectures
sparse computation
0.312026
Uni-STC: Unified Sparse Tensor Core · HPCA 2026
Computer vision › Vision and language
image captioning
0.312025
End-to-End Multi-Modal Diffusion Mamba · ICCV 2025
Computer vision › Vision and language
visual question answering
0.312025
End-to-End Multi-Modal Diffusion Mamba · ICCV 2025

Methods — techniques the papers use, named apart from their topics

unified sparse format · 1.0fine-grained task partitioning · 1.0variational autoencoder · 0.9mamba · 0.9diffusion model · 0.9
YearPublicationVenuePosition
2026 Uni-STC: Unified Sparse Tensor Core
abstract
Modern processors are increasingly adopting tensor cores as key computational units. Compared to existing designs for dense and structured sparsity, recent dual-side sparse tensor cores have evolved to support general sparsity. However, existing methods still face limitations on generality (incomplete sparse kernel support prevents broad applicability) and performance (outer-product/row-row schemes yield unsatisfactory hardware utilisation, data reuse, and energy efficiency). In this paper, we propose Uni-STC, a unified sparse tensor core that delivers high-performance dataflows for four key sparse kernels: sparse matrix-vector multiplication (SpMV), sparse matrixsparse vector multiplication (SpMSpV), sparse matrix-multiple vector multiplication (SpMM), and sparse general matrix-matrix multiplication (SpGEMM). To efficiently support these diverse sparse workloads, we first introduce BBC, a unified sparse format co-designed with Uni-STC's dataflow. We then design UniSTC's architecture supporting (1) fine-grained task partitioning to improve resource utilisation, (2) parallel sparse-tile processing to enhance data reuse, and (3) a dynamic network to reduce intermediate data movement and energy consumption. Evaluated across 2893 SuiteSparse and 302 DLMC matrices, Uni-STC demonstrates significant improvements, outperforming the state-of-the-art RM-STC with a$2.21 \times$geomean speedup and$2.96 \times$higher energy efficiency.
Haocheng Lian, Meichen Dong, Yijie Nie, Junzhong Shen, Chun Huang 0006, Bingcai Sui, Weifeng Liu 0002
HPCA4
2025 End-to-End Multi-Modal Diffusion Mamba
abstract
Current end-to-end multi-modal models utilize different encoders and decoders to process input and output information. This separation hinders the joint representation learning of various modalities. To unify multi-modal processing, we propose a novel architecture called MDM (Multi-modal Diffusion Mamba). MDM utilizes a Mamba-based multi-step selection diffusion model to progressively generate and refine modality-specific information through a unified variational autoencoder for both encoding and decoding. This innovative approach allows MDM to achieve superior performance when processing high-dimensional data, particularly in generating high-resolution images and extended text sequences simultaneously. Our evaluations in areas such as image generation, image captioning, visual question answering, text comprehension, and reasoning tasks demonstrate that MDM significantly outperforms existing end-to-end models (MonoFormer, LlamaGen, and Chameleon etc.) and competes effectively with SOTA models like GPT-4V, Gemini Pro, and Mistral. Our results validate MDM's effectiveness in unifying multi-modal processes while maintaining computational efficiency, establishing a new direction for end-to-end multi-modal architectures.
Chunhao Lu, Qiang Lu 0005, Meichen Dong, Jake Luo
ICCV3
2021 TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs
abstract
With the extensive use of GPUs in modern supercomputers, accelerating sparse matrix-vector multiplication (SpMV) on GPUs received much attention in the last couple of decades. A number of techniques, such as increasing utilization of wide vector units, reducing load imbalance and selecting the best formats, have been developed. However, the 2D spatial sparsity structure has not been well exploited in the existing work for SpMV on GPUs. In this paper, we propose an efficient tiled algorithm called TileSpMV for optimizing SpMV on GPUs through exploiting 2D spatial structure of sparse matrices. We first implement seven warp-level SpMV methods for calculating sparse tiles stored in a variety of formats, and then design a selection method to find the best format and SpMV implementation for each tile. We also adaptively extract nonzeros in the very sparse tiles into a separate matrix to maximize the overall performance. The experimental results show that our method is faster than state-of-the-art SpMV methods such as Merge-SpMV, CSR5 and BSR in most matrices of the full SuiteSparse Matrix Collection and delivers up to 2.61x, 3.96x and 426.59x speedups, respectively.
Yuyao Niu, Zhengyang Lu 0003, Meichen Dong, Zhou Jin 0001, Weifeng Liu 0002, Guangming Tan
IPDPS3
2021 SCDC: bulk gene expression deconvolution by multiple single-cell RNA sequencing references
abstract
Recent advances in single-cell RNA sequencing (scRNA-seq) enable characterization of transcriptomic profiles with single-cell resolution and circumvent averaging artifacts associated with traditional bulk RNA sequencing (RNA-seq) data. Here, we propose SCDC, a deconvolution method for bulk RNA-seq that leverages cell-type specific gene expression profiles from multiple scRNA-seq reference datasets. SCDC adopts an ENSEMBLE method to integrate deconvolution results from different scRNA-seq datasets that are produced in different laboratories and at different times, implicitly addressing the problem of batch-effect confounding. SCDC is benchmarked against existing methods using both in silico generated pseudo-bulk samples and experimentally mixed cell lines, whose known cell-type compositions serve as ground truths. We show that SCDC outperforms existing methods with improved accuracy of cell-type decomposition under both settings. To illustrate how the ENSEMBLE framework performs in complex tissues under different scenarios, we further apply our method to a human pancreatic islet dataset and a mouse mammary gland dataset. SCDC returns results that are more consistent with experimental designs and that reproduce more significant associations between cell-type proportions and measured phenotypes.
Meichen Dong, Aatish Thennavan, Eugene Urrutia, Charles M. Perou
Briefings Bioinform.1