VLDB 2026 Research / reviewers in the wild / expert
Haolin Pan
dblp:310/5648
· DBLP profile ↗
12ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Achieving Fast and High Throughput Data Exchange for Serverless Computing Systems via Switch-Native Serialization/Deserialization
Gonglong Chen, Baiyan Ke, Yuxin Xu, Shenghong Xiong, Haolin Pan, Jiamei Lv, Wenxing Ge, Cheng-Zhong Xu 0001, Kejiang Ye |
IWQoS | 5 |
| 2026 | FlowXpert: Context-Aware Flow Embedding for Enhanced Traffic Detection in IoT NetworkabstractIn the Internet of Things (IoT) environment, continuous interaction among a large number of devices generates complex and dynamic network traffic, which poses significant challenges to rule-based detection approaches. Machine learning (ML)-based traffic detection technology, capable of identifying anomalous patterns and potential threats within this traffic, serves as a critical component in ensuring network security. This study first identifies a significant issue with widely adopted feature extraction tools (e.g., CICFlowMeter): the extensive use of time- and length-related features leads to high sparsity, which adversely affects model convergence. Furthermore, existing traffic detection methods generally lack an embedding mechanism capable of efficiently and comprehensively capturing the semantic characteristics of network traffic. To address these challenges, we propose a novel feature extraction tool that eliminates traditional time and length features in favor of context-aware semantic features related to the source host, thus improving the generalizability of the model. In addition, we design an embedding training framework that integrates the unsupervised DBSCAN clustering algorithm with a contrastive learning strategy to effectively capture fine-grained semantic representations of traffic. Extensive empirical evaluations are conducted on the real-world dataset (Mawi) and two simulated datasets (CICIDS-2017 and UNSW-NB15) to validate the proposed method in terms of detection accuracy, robustness, and generalization. Comparative experiments against several state-of-the-art (SOTA) models demonstrate the superior performance of our approach. Furthermore, we confirm its applicability and deployability in real-time scenarios. Chao Zha, Haolin Pan, Ruyun Zhang 0001 |
IEEE Trans. Mob. Comput. | 2 |
| 2025 | Restabilizing Diffusion Models with Predictive Noise Fusion Strategy for Image Super-ResolutionabstractDiffusion models are prominent in image generation for producing detailed and realistic images from Gaussian noises. However, they often encounter instability issues in image restoration tasks, e.g., super-resolution. Existing methods typically rely on multiple runs to find an initial noise that produces a reasonably restored image. Unfortunately, these methods are computationally expensive and time-consuming without guaranteeing stable and consistent performance. To address these challenges, we propose a novel Predictive Noise Fusion Strategy (PNFS) that predicts pixel-wise errors in the restored image and combines different noises to generate a more effective noise. Extensive experiments show that PNFS significantly improves the stability and performance of diffusion models in super-resolution, both quantitatively and qualitatively. Furthermore, PNFS can be flexibly integrated into various diffusion models to enhance their stability. Luoqian Jiang, Bingna Xu, Haolin Pan, Jiezhang Cao, Wenbo Li 0002, Jian Chen 0011 |
AAAI | 4 |
| 2025 | Towards Efficient Compiler Auto-tuning: Leveraging Synergistic Search SpacesabstractDetermining the optimal sequence of compiler optimization passes is challenging due to the extensive and intricate search space. Traditional auto-tuning techniques, such as iterative compilation and machine learning methods, are often limited by high computational costs and difficulties in generalizing to new programs. These approaches can be inefficient and may not fully address the varying optimization needs across different programs. This paper introduces a novel approach that leverages the synergistic relationships between optimization passes to effectively reduce the search space. By focusing on chained synergy pass pairs that jointly optimize a specific target, our method uses K-means clustering to capture common optimization patterns across programs and forms these pairs into coresets. Leveraging a supervised learning model trained on these coresets, we effectively predict the most beneficial coreset for new programs, streamlining the search for optimal sequences. By integrating various search strategies, our method quickly converges to near-optimal solutions. Our approach achieves state-of-the-art performance on ten benchmark datasets, including MiBench, CBench, NPB, and CHStone, demonstrating an average reduction of 7.5% in Intermediate Representation (IR) instruction count compared to Oz. Furthermore, this set of chained synergy pass pairs is also well-suited for iterative search studies by other researchers, as it enables achieving an average codesize reduction of 13.9% compared to Oz with a simple search strategy that takes only about 5 seconds, outperforming existing search-based techniques in the initial pass search space across five datasets. Haolin Pan, Yuanyu Wei, Mingjie Xing, Chen Zhao 0024 |
CGO | 1 |
| 2025 | Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic TeacherabstractKnowledge distillation aims to transfer knowledge from a large teacher model to a compact student counterpart, often coming with a significant performance gap between them. Interestingly, we find that a too-large performance gap can hamper the training process.
To alleviate this, we propose a **Gap Preserving Distillation (GPD)** method that trains an additional dynamic teacher model from scratch along with the student to maintain a reasonable performance gap. To further strengthen distillation, we develop a hard strategy by enforcing both models to share parameters. Besides, we also build the soft bidirectional mappings between them through ***Inverse Reparameterization (IR)*** and ***Channel-Branch Reparameterization (CBR)***.
IR initializes a larger dynamic teacher with approximately the same accuracy as the student to avoid a too large gap in early stage of training. CBR enables direct extraction of an effective student model from the dynamic teacher without post-training.
In experiments, GPD significantly outperforms existing distillation methods on top of both CNNs and transformers, achieving up to 1.58\% accuracy improvement.
Interestingly, GPD also generalizes well to the scenarios without a pre-trained teacher, including training from scratch and fine-tuning, yielding a large improvement of 1.80\% and 0.89\% on ResNet18, respectively. Shulian Zhang, Haolin Pan, Jing Liu 0048, Yulun Zhang 0001, Jian Chen 0011 |
ICLR | 3 |
| 2025 | HybridSIMD: A Super C++ SIMD Library with Integrated Auto-tuning CapabilitiesabstractSingle Instruction, Multiple Data (SIMD) technology is crucial for enhancing computational efficiency in High-Performance Computing (HPC). While C++ SIMD libraries abstract away low-level complexities, their proliferation has led to a fragmented set of libraries, creating significant challenges in both performance and usability for developers. To overcome these library-level limitations, this paper introduces a new collaborative concept for SIMD library design. We present HybridSIMD, a C++ library to embody this principle, resolving fragmentation through a unified interface and an operator-level collaborative back-end that leverages the collective strengths of existing libraries. A built-in auto-tuning engine, featuring a hierarchical search strategy, automatically navigates the rich optimization space created by this collaborative approach to deliver maximum performance without manual intervention. Experimental results across six real-world HPC benchmarks on AVX2, AVX512, and NEON architectures demonstrate HybridSIMD’s superiority. Notably, the highest speedups achieved are 185.34× on AVX2, 97.80× on AVX512, and 71.32× on NEON, showcasing its effectiveness in resolving fragmentation while delivering state-of-the-art performance. Our artifact is available at https://github.com/Panhaolin2001/HybridSIMD. Haolin Pan, Xulin Zhou, Mingjie Xing |
ASE | 1 |
| 2025 | Compiler-R1: Towards Agentic Compiler Auto-tuning with Reinforcement LearningabstractCompiler auto-tuning optimizes pass sequences to improve performance metrics such as Intermediate Representation (IR) instruction count. Although recent advances leveraging Large Language Models (LLMs) have shown promise in automating compiler tuning, two significant challenges still remain: the absence of high-quality reasoning datasets for agents training, and limited effective interactions with the compilation environment. In this work, we introduce Compiler-R1, the first reinforcement learning (RL)-driven framework specifically augmenting LLM capabilities for compiler auto-tuning. Compiler-R1 features a curated, high-quality reasoning dataset and a novel two-stage end-to-end RL training pipeline, enabling efficient environment exploration and learning through an outcome-based reward. Extensive experiments across seven datasets demonstrate Compiler-R1 achieving an average 8.46\% IR instruction count reduction compared to opt -Oz, showcasing the strong potential of RL-trained LLMs for compiler optimization. Our code and datasets are publicly available at https://github.com/Panhaolin2001/Compiler-R1. Haolin Pan, Kaichun Yao, Libo Zhang 0001, Mingjie Xing |
NeurIPS | 1 |
| 2025 | Navigating the SIMD Optimization Maze: A Reinforcement Learning Approach to Library and Compiler Co-OptimizationabstractSingle Instruction Multiple Data (SIMD) programs are crucial for performance, yet their optimization is complicated by hardware diversity and the varied behaviors of C++ SIMD libraries used to ensure portability.Standard compiler heuristics often struggle with the complex interactions between SIMD library implementations and optimization pass sequences, sometimes even leading to performance degradation compared to scalar code.The vast search space created by the need to cooptimize both SIMD library selection and compiler pass ordering makes manual tuning infeasible.To address this joint optimization challenge, we propose a Reinforcement Learning (RL) based auto-tuning framework specifically for SIMD programs.Our approach introduces three core contributions: (1) A unified C++ template interface seamlessly integrating seven distinct C++ SIMD libraries, facilitating switching between them.(2) An RL agent operating within a joint action space that simultaneously selects both the C++ SIMD libraries and LLVM optimization passes, enabling their co-optimization.(3) An assembly-level program representation designed to capture low-level SIMD characteristics, providing effective guidance for the RL agent.Using LLVM Intermediate Representation(IR) instruction count reduction as a stable proxy metric, experiments on nine SIMD benchmarks demonstrate that our method achieves an average 14.3% improvement over the LLVM opt -Oz baseline.This result highlights the efficacy of our approach in navigating the complex joint optimization space, outperforming traditional auto-tuning techniques. Haolin Pan, Mingjie Xing |
SEKE | 1 |
| 2024 | Boosting semi-supervised learning with Contrastive Complementary Labeling
Qinyi Deng, Zhibang Yang, Haolin Pan, Jian Chen 0011 |
Neural Networks | 4 |
| 2024 | Enhanced Long-Tailed Recognition With Contrastive CutMix AugmentationabstractReal-world data often follows a long-tailed distribution, where a few head classes occupy most of the data and a large number of tail classes only contain very limited samples. In practice, deep models often show poor generalization performance on tail classes due to the imbalanced distribution. To tackle this, data augmentation has become an effective way by synthesizing new samples for tail classes. Among them, one popular way is to use CutMix that explicitly mixups the images of tail classes and the others, while constructing the labels according to the ratio of areas cropped from two images. However, the area-based labels entirely ignore the inherent semantic information of the augmented samples, often leading to misleading training signals. To address this issue, we propose a Contrastive CutMix (ConCutMix) that constructs augmented samples with semantically consistent labels to boost the performance of long-tailed recognition. Specifically, we compute the similarities between samples in the semantic space learned by contrastive learning, and use them to rectify the area-based labels. Experiments show that our ConCutMix significantly improves the accuracy on tail classes as well as the overall performance. For example, based on ResNeXt-50, we improve the overall accuracy on ImageNet-LT by 3.0% thanks to the significant improvement of 3.3% on tail classes. We highlight that the improvement also generalizes well to other benchmarks and models. Our code and pretrained models are available at https://github.com/PanHaulin/ConCutMix. Haolin Pan, Mianjie Yu, Jian Chen 0011 |
IEEE Trans. Image Process. | 1 |
| 2023 | Improving fine-tuning of self-supervised models with Contrastive Initialization
Haolin Pan, Qinyi Deng, Haomin Yang, Jian Chen 0011 |
Neural Networks | 1 |
| 2021 | Multi-scale Space-time Registration of Growing PlantsabstractIn this paper, we introduce a new method for the space-time registration of a growing plant that is based on matching the plant at different geometric scales. The proposed method starts with the creation of a topological skeleton of the plant at each time step. This skeleton is then used to segment the plant into parts that we call branches. Then these branches are further divided into smaller segments that possess a simple geometric structure. These segments are matched between two time steps using a random forest classifier based on their topological and geometric features. Then, for each pair of segments matched, a point-wise registration is devised using a non-rigid registration method based on a local ICP. We applied our method to various types of plants, including arabidopsis, tomato plant and maize. We established three different metrics for 3D point-wise shape correspondence to test the accuracy, continuity, and cycle consistency of the mapping. We then compared our method with the state-of-the-art. Our results show that our approach achieves better or similar results with a shorter running time. Haolin Pan, Franck Hétroy-Wheeler, Julie Charlaix, David Colliaux |
3DV | 1 |