EDBT 2026 Demo / reviewers in the wild / expert
Jun Shi 0007
dblp:31/626-7
· DBLP profile ↗
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-9888-6238ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CISim: ISA-Agnostic Custom Instruction Simulation for General-Purpose ProcessorabstractPre-RTL ISA-agnostic simulators have been established for designing heterogeneous systems, but few of them are suitable for evaluating a general-purpose processor (GPP) with custom instructions (CIs). MosaicSim [1], a state-of-the-art ISA-agnostic simulator, still has several limitations for CI design and simulation. First, it shows inaccuracy in simulating GPPs due to an oversimplified performance model. Second, as designed for kernel simulation, it lacks support for running complex real-world benchmarks. Third, it cannot evaluate fine-grained irregular CIs due to the lack of the ability to represent or define them in benchmarks. To this end, we propose CISim, a new ISA-agnostic simulation framework containing an offloader that generates and integrates CIs into benchmarks, along with a simulator capable of executing benchmarks with CIs. Evaluations show that CISim is accurate by validating against Gem5 [2] and achieves higher accuracy than MosaicSim. A case study evaluating CI exploration methods highlights the strength and flexibility of CISim. Jun Shi 0007, Junshi Chen 0003, Hong An |
DATE | 4 |
| 2026 | AtSpMV: Model-Guided Adaptive Tiling and Load Balancing for SpMV on GPUs
Junshi Chen 0003, Longsheng Song, Jun Shi 0007, Hong An |
Euro-Par (1) | 6 |
| 2026 | RVDeformer: Sparse Point Cloud-Guided Right Ventricle 3-D Reconstruction in Echocardiogramsabstract3D reconstruction of the Right Ventricle (RV) from echocardiograms is crucial for accurate clinical evaluation of cardiac function. However, existing methods are hindered by the complex RV anatomy and the incomplete spatial information inherent in 2D multi-view echocardiograms. Therefore, we propose RVDeformer, a sparse point cloud-guided framework for RV 3D reconstruction. RVDeformer reformulates the reconstruction task as a mesh deformation problem, learning to deform a predefined template mesh to match the target structure under the guidance of the sparse anatomical point cloud. Specifically, this framework employs the end-to-end neural network RVDeformNet to extract the features of the point cloud and template mesh for predicting the displacement of each mesh vertex. We design a point cloud-mesh fusion module that can effectively align and fuse features from the two modalities to enhance the representation ability of the model. We conduct extensive validation on a clinical dataset of 1,278 cases and demonstrate that RVDeformer outperforms existing state-of-the-art methods, achieving a Chamfer Distance (CD) of $2.24\pm 0.55$ mm, an F1-score of $0.74\pm 0.10$ at the 3 mm threshold, and a Volumetric Similarity (VS) of $91.53\pm 2.28$ %, with significant potential for clinical applications. The code is available at https://github.com/onezh95/RVDeformer. Zhaohui Wang 0002, Jun Shi 0007, Minfan Zhao, Hong An |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Pruner: A Draft-then-Verify Exploration Mechanism to Accelerate Tensor Program TuningabstractTensor program tuning is essential for the efficient deployment of deep neural networks. Search-based approaches have demonstrated scalability and effectiveness in automatically finding high-performance programs for specific hardware. However, the search process is often inefficient, taking hours or even days to discover optimal programs due to the exploration mechanisms guided by an accurate but slow-learned cost model. Meanwhile, the learned cost model trained on one platform cannot seamlessly adapt online to another, which we call cross-platform online unawareness. In this work, we propose Pruner and MoA-Pruner. Pruner is a ''Draft-then-Verify'' exploration mechanism that accelerates the schedule search process. Instead of applying the complex learned cost model to all explored candidates, Pruner drafts small-scale potential candidates by introducing a naive Symbol-based Analyzer (draft model), then identifies the best candidates by the learned cost model. MoA-Pruner introduces a Momentum online Adaptation strategy to address the cross-platform online unawareness. Jun Shi 0007, Minfan Zhao, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan |
ASPLOS (2) | 2 |
| 2025 | Carver: Learning to Reconstruct Right Ventricle from Sparse Multi-View 2D EchocardiogramsabstractAccurate 3D reconstruction of the right ventricle from multi-view echocardiograms is crucial for the quantitative diagnosis of cardiac diseases. However, existing methods often fail to deliver satisfactory results due to the structural complexity of the right ventricle and the sparsity of non-parallel ultrasound views. In this paper, we propose an efficient reconstruction method named Carver, which redefines 3D reconstruction of the right ventricle as a voxel-wise dense prediction task for the first time. The core idea lies in using a deep neural network to learn the end-to-end deformation from a coarse geometric convex hull to a fine structure of the right ventricle, similar to carving. To improve the reconstruction accuracy and robustness, we design a dual-aware network incorporating prior contour information to enhance learning representation. We conduct extensive experiments on an echocardiography dataset containing 1,278 instances to validate the effectiveness of the proposed method. Experimental results demonstrate that Carver outperforms existing state-of-the-art methods, achieving a Volume Similarity (VS) of 98.75%, a Dice Similarity Coefficient (DSC) of 97.80%, a Hausdorff Distance (HD) of 4.96, and a Root Mean Square Error (RMSE) of 0.013 for the ejection fraction, while maintaining considerable robustness even with sparser inputs. The code is available at https://github.com/ustclyd/Carver. Jun Shi 0007, Zhaohui Wang 0002, Tiantong Wang, Minfan Zhao, Junshi Chen 0003, Hong An |
ICASSP | 2 |
| 2025 | PromptSeg: Learning to Segment Medical Image via Visual PromptsabstractDeep learning has made remarkable medical image segmentation advancements, yet its generalization capability across tasks remains challenging. The variety of task objectives, disease-dependent labeling variations, and multi-center data contribute to the poor generalization capacity of task-specific models on unseen tasks, necessitating domain adaptation or fine-tuning. This typically involves data annotation and network retraining, limiting the application of deep learning in clinical practice. In this study, we propose PromptSeg, an innovative Transformer-based unified segmentation framework for general medical image segmentation tasks. PromptSeg aims to utilize the provided visual prompts to recognize task patterns and learn contextual representations, thereby breaking the restrictions of the task-specific paradigm. When faced with unseen datasets or segmentation targets during inference, our method only requires a few annotated prompt pairs to understand the task and segment the query images without retraining, alleviating the need for large-scale annotation data. The experimental results demonstrate that our method outperforms existing state-of-the-art methods and exhibits high generalization capability on multiple unseen tasks. The source code is available at https://github.com/MinfanZhao/PromptSeg. Minfan Zhao, Jun Shi 0007, Zhaohui Wang 0002, Junshi Chen 0003, Hong An |
ICASSP | 3 |
| 2025 | WinRS: Accelerate Winograd Backward-Filter Convolution with Tiny WorkspaceabstractWinograd algorithm powerfully accelerates Convolutional Neural Networks. However, for backward-filter convolution (BFC), existing implementations often struggle to achieve both high throughput and low memory usage, due to challenges from large filters and small outputs. We propose WinRS, a fast, memory-efficient, and flexible BFC algorithm. WinRS reduces N-D large filters into 1D formats and precisely splits them to match the fastest kernels. These fully-fused kernels execute BFC in on-chip memory with tiny workspace, leveraging the superior acceleration potential of 1D Winograd. WinRS adaptively balances workloads into an optimal number of block groups, maximizing hardware utilization in small-output cases. When ported to FP16 on Tensor Cores, WinRS achieves 3.27 × throughput of its FP32 CUDA-Core version. In experiments, WinRS achieves 1.05 × to 4.7 × speedup over cuDNN GEMM using comparable workspace; WinRS uses less than 4% workspace of cuDNN FFT and Winograd, and exhibits higher throughput with memory- and FLOP-bound workloads. Junshi Chen 0003, Jingwei Sun 0001, Zhuopin Xu, Jun Shi 0007, Qi Wang 0131 |
ICPP | 6 |
| 2025 | CIExplorer: Microarchitecture-Aware Exploration for Tightly Integrated Custom InstructionabstractExtending existing architectures with customized instruction extensions is emerging to achieve high performance and energy efficiency for specific applications.Automated discovery of custom instructions (CIs) is well-studied nowadays, which requires exploring combinations of different types and quantities of operations, resulting in a vast search space.However, previous works typically use microarchitectureagnostic cost models, leading to suboptimal CIs that may degrade performance.They leverage graph isomorphism to reduce area overhead, but few of them consider its potential to benefit performance-oriented exploration.To this end, we present CIExplorer, a framework for adaptive CI exploration. Qingcai Jiang, Jun Shi 0007, Junshi Chen 0003, Hong An, Xulong Tang, Honghui Yuan |
ICS | 5 |
| 2024 | Predictive Accuracy-Based Active Learning for Medical Image Segmentation
Jun Shi 0007, Shulan Ruan, Minfan Zhao, Hong An, Xudong Xue |
IJCAI | 1 |
| 2023 | H-DenseFormer: An Efficient Hybrid Densely Connected Transformer for Multimodal Tumor Segmentation
Jun Shi 0007, Hongyu Kan, Shulan Ruan, Minfan Zhao, Zhaohui Wang 0002, Hong An, Xudong Xue |
MICCAI (4) | 1 |
| 2023 | Multi-View Attention Learning for Residual Disease Prediction of Ovarian CancerabstractIn the treatment of ovarian cancer, precise residual disease prediction is significant for clinical and surgical decision-making. However, traditional methods are either invasive (e.g., laparoscopy) or time-consuming (e.g., manual analysis). Recently, deep learning methods make many efforts in automatic analysis of medical images. Despite the remarkable progress, most of them underestimated the importance of 3D image information of disease, which might brings a limited performance for residual disease prediction, especially in small-scale datasets. To this end, in this paper, we propose a novel Multi-View Attention Learning (MuVAL) method for residual disease prediction, which focuses on the comprehensive learning of 3D Computed Tomography (CT) images in a multi-view manner. Specifically, we first obtain multi-view of 3D CT images from transverse, coronal and sagittal views. To better represent the image features in a multi-view manner, we further leverage attention mechanism to help find the more relevant slices in each view. Extensive experiments on a dataset of 111 patients show that our method outperforms existing deep-learning methods. Xiangneng Gao, Shulan Ruan, Jun Shi 0007, Guoqing Hu |
SMC | 3 |
| 2021 | DARNet: Dual-Attention Residual Network for Automatic Diagnosis of COVID-19 via CT ImagesabstractThe ongoing global pandemic of Coronavirus Disease 2019 (COVID-19) poses a serious threat to public health and the economy. Rapid and accurate diagnosis of COVID-19 is essential to prevent the further spread of the disease and reduce its mortality. Chest Computed tomography (CT) is an effective tool for the early diagnosis of lung diseases including pneumonia. However, detecting COVID-19 from CT is demanding and prone to human errors as some early-stage patients may have negative findings on images. Recently, many deep learning methods have achieved impressive performance in this regard. Despite their effectiveness, most of these methods underestimate the rich spatial information preserved in the 3D structure or suffer from the propagation of errors. To address this problem, we propose a Dual-Attention Residual Network (DARNet) to automatically identify COVID-19 from other common pneumonia (CP) and healthy people using 3D chest CT images. Specifically, we design a dual-attention module consisting of channel-wise attention and depth-wise attention mechanisms. The former is utilized to enhance channel independence, while the latter is developed to recalibrate the depth-level features. Then, we integrate them in a unified manner to extract and refine the features at different levels to further improve the diagnostic performance. We evaluate DARNet on a large public CT dataset and obtain superior performance. Besides, the ablation study and visualization analysis prove the effectiveness and interpretability of the proposed method. Jun Shi 0007, Huite Yi, Shulan Ruan, Zhaohui Wang 0002, Hong An |
BIBM | 1 |