EDBT 2026 Demo / reviewers in the wild / expert
Yuhao Qing
dblp:285/8994
· DBLP profile ↗
16ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CHTracker: Confidence-Guided Hierarchical Association Paradigm for Multi-Object TrackingabstractMulti-object tracking (MOT) has garnered considerable attention due to its relevance in practical applications such as automated devices in smart cities. However, under complex conditions, existing trackers often fail to accurately capture or characterize target motion patterns, exhibiting limitations in flexibility and interpretability. To address these challenges, this paper introduces CHTracker, a confidence-guided hierarchical association paradigm for MOT. By integrating spatial features with varying confidence levels, CHTracker enhances the granularity of motion pattern modeling in edge-case scenarios where conventional trackers are prone to association ambiguity. Our paradigm adaptively utilizes distinct tracking cues and assignment metrics tailored to hierarchical target structures, thereby enabling collaborative tracking. Additionally, CHTracker incorporates the diagonal length of the target bounding box as a state variable during position prediction, which significantly improves the robustness against diverse motion noise. Extensive experimental results on multiple benchmarks, including Dance-Track, MOT17, MOT20, and Singapore Maritime Dataset (SMD), demonstrate that CHTracker achieves the state-of-the-art performance in accuracy, robustness, and generalization. Furthermore, our association paradigm is extended to a visible-infrared fusion version for evaluation on the multimodal CAMEL dataset, underscoring its practical potential to fulfill heterogeneous modality requirements in real-world scenarios. Our code will be available at https://github.com/ZyanChenyang/CHTracker. Chenyang Yan, Yueying Wang, Yuhao Qing, Weidong Zhang 0007, Xin Xu 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | MCFINet: A Cost-Efficient Multi-Channel Feature Integration Network for Surface Scenarios Image Super-ResolutionabstractConvolutional Neural Network (CNN) and Vision Transformer (ViT) have revolutionized the field of image super-resolution (SR). However, their complexity poses challenges for resource—constrained scenarios, particularly due to the high computational demands of Transformers and their excessive reliance on global information. To tackle these challenges, we propose a Multi-Channel Feature Integration Network (MCFINet), designed to maximize input pixel utilization while minimizing computational overhead. It integrates both local and global features within the channels, thereby exploiting their complementary advantages. First, the designed Feature Integration Block (FIB) effectively captures local information and improves visual quality by enhancing the mapping of non-local features. Subsequently, we utilize the Adaptive Channel Fusion Block (ACFB), which strengthens the interaction between features and channels while maintaining computational efficiency. Finally, for SR task on resource-constrained surface scenarios, we propose a more suitable pre-training method, which further boosts the model’s learning ability. Evaluation results indicate that the proposed MCFINet achieves a better balance between lightweight design and high-quality restoration on both standard evaluation datasets and water surface target datasets. Specifically, compared to the traditional SwinIR-L, MCFINet reduces model training time and runtime by 12% on the test set, while also decreasing model complexity by 43%. Our codes are available at https://github.com/Lcasjz/MCFINet . Liangcheng Zhao, Yueying Wang, Yuhao Qing, Dan Zeng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Integrating Low-Level Visual Cues for Enhanced Unsupervised Semantic SegmentationabstractUnsupervised semantic segmentation algorithms aim to identify meaningful semantic groups without annotations. Recent approaches leveraging self-supervised transformers as pre-training backbones have successfully obtained high-level dense features that effectively express semantic coherence. However, these methods often overlook local semantic coherence and low-level features such as color and texture. We propose integrating low-level visual cues to complement high-level visual cues derived from self-supervised pre-training branches. Our findings indicate that low-level visual cues provide a more coherent recognition of color-texture aspects, ensuring the continuity of spatial structures within classes. This insight led us to develop IL2Vseg, an unsupervised semantic segmentation method that leverages the complementation of low-level visual cues. The core of IL2Vseg is a spatially-constrained fuzzy clustering algorithm based on color affinities, which preserves the intra-class affinity of spatially-adjacent and similarly-colored pixels in low-level visual cues. Additionally, to effectively couple low-level and high-level visual cues, we introduce a feature similarity loss function to optimize the feature representation of fused visual cues. To further enhance consistent feature learning, we incorporate contrast loss functions based on color invariance and luminosity invariance, which improve the learning of features from different semantic categories. Extensive experiments on multiple datasets, including COCO-Stuff-27, Cityscapes, Potsdam, and MaSTr1325, demonstrate that IL2Vseg achieves state-of-the-art results. Yuhao Qing, Dan Zeng 0001, Shaorong Xie, Kaer Huang, Yueying Wang |
AAAI | 1 |
| 2025 | FoldMoE: Efficient Long Sequence MoE Training via Attention-MoE PipeliningabstractGuichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu, Fanxin Li, Dong Huang, Zekai Sun, Heming Cui. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Guichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu, Fanxin Li, Dong Huang 0005, Zekai Sun, Heming Cui |
ACL (1) | 3 |
| 2025 | EffiCoder: Enhancing Code Generation in Large Language Models through Efficiency-Aware Fine-tuningabstractAs large language models (LLMs) play an increasingly important role in code generation, enhancing both correctness and efficiency has become crucial. Current methods primarily focus on correctness, often overlooking efficiency. To address this gap, we introduce SWIFTCODE to improve both aspects by fine-tuning LLMs on a high-quality dataset comprising correct and efficient code samples. Our methodology involves leveraging multiple LLMs to generate diverse candidate code solutions for various tasks across different programming languages. We then evaluate these solutions by directly measuring their execution time and memory usage through local execution. The code solution with the lowest execution time and memory consumption is selected as the final output for each task. Experimental results demonstrate significant improvements when fine-tuning with SWIFTCODE. For instance, Qwen2.5-Coder-7B-Instruct’s pass@1 score increases from 44.8% to 57.7%, while the average execution time for correct tasks decreases by 48.4%. SWIFTCODE offers a scalable and effective solution for advancing AI-driven code generation, benefiting both software development and computational problem-solving. Dong Huang 0005, Guangtao Zeng, Jianbo Dai, Meng Luo 0010, Han Weng, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050 |
ICML | 6 |
| 2025 | Two Heads are Better than One: Robust Learning Meets Multi-branch ModelsabstractDeep neural networks (DNNs) are vulnerable to adversarial examples, in which DNNs are misled to false outputs due to inputs containing imperceptible perturbations. Adversarial training, a reliable and effective method of defense, may significantly reduce the vulnerability of neural networks and becomes the de facto standard for robust learning. While many recent works practice the data-centric philosophy, such as how to generate better adversarial examples or use generative models to produce additional training data, we look back to the models themselves and revisit the adversarial robustness from the perspective of deep feature distribution as an insightful complementarity. In this paper, we propose Branch Orthogonality adveRsarial Training (BORT) to obtain state-of-the-art performance with solely the original dataset for adversarial training. To practice our design idea of integrating multiple orthogonal solution spaces, we leverage a simple multi-branch neural network and propose a corresponding loss function, branch-orthogonal loss, to make each solution space of the multi-branch model orthogonal. We evaluate our approach on CIFAR-10, CIFAR-100 and SVHN against$\ell_{\infty}$norm-bounded perturbations of size$\epsilon=8 / 255$, respectively. Exhaustive experiments are conducted to show that our method goes beyond all state-of-the-art methods without any tricks. Compared to all methods that do not use additional data for training, our models achieve 67.3 % and 41.5 % robust accuracy on CIFAR-10 and CIFAR-100 (improving upon the state-of-the-art by$\mathbf{+7.23 \%}$and$\mathbf{+9.07 \%}$). Zongyuan Zhang, Qingwen Bu, Tianyang Duan, Zheng Lin 0001, Yuhao Qing, Zihan Fang 0003, Heming Cui, Dong Huang 0005 |
ICPADS | 5 |
| 2025 | Afterburner: Reinforcement Learning Facilitates Self-Improving Code Efficiency OptimizationabstractLarge Language Models (LLMs) generate functionally correct solutions but often fall short in code efficiency, a critical bottleneck for real-world deployment. In this paper, we introduce a novel test-time iterative optimization framework to address this, employing a closed-loop system where LLMs iteratively refine code based on empirical performance feedback from an execution sandbox. We explore three training strategies: Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Group Relative Policy Optimization~(GRPO). Experiments on our Venus dataset and the APPS benchmark show that SFT and DPO rapidly saturate in efficiency gains. In contrast, GRPO, using reinforcement learning (RL) with execution feedback, continuously optimizes code performance, significantly boosting both pass@1 (from 47% to 62%) and the likelihood of outperforming human submissions in efficiency (from 31% to 45%). Our work demonstrates effective test-time code efficiency improvement and critically reveals the power of RL in teaching LLMs to truly self-improve code efficiency. We released our code and data at https://github.com/Elfsong/Afterburner. Mingzhe Du, Anh Tuan Luu, Yue Liu 0008, Yuhao Qing, Dong Huang 0005, Qian Liu 0033, Zejun Ma 0001, See-Kiong Ng |
NeurIPS | 4 |
| 2025 | EffiBench-X: A Multi-Language Benchmark for Measuring Efficiency of LLM-Generated CodeabstractExisting code generation benchmarks primarily evaluate functional correctness, with limited attention to code efficiency, and they are often restricted to a single language such as Python. To address this gap, we introduce EffiBench‑X, the first large‑scale multi‑language benchmark specifically designed for robust efficiency evaluation of LLM‑generated code. EffiBench‑X supports Python, C++, Java, JavaScript, Ruby, and Go, and comprises competitive programming tasks paired with human‑expert solutions as efficiency baselines. Evaluating state‑of‑the‑art LLMs on EffiBench‑X reveals that while models frequently generate functionally correct code, they consistently underperform human experts in efficiency. Even the most efficient LLM‑generated solutions (e.g., Qwen3‑32B) achieve only around 62% of human efficiency on average, with significant language‑specific variation: models tend to perform better in Python, Ruby, and JavaScript than in Java, C++, and Go (e.g., DeepSeek‑R1’s Python code is markedly more efficient than its Java code). These findings highlight the need for research into optimization‑oriented methods to improve the efficiency of LLM‑generated code across diverse languages. The dataset and evaluation infrastructure are publicly available at https://github.com/EffiBench/EffiBench-X.git and https://huggingface.co/datasets/EffiBench/effibench-x. Yuhao Qing, Boyu Zhu, Mingzhe Du, Zhijiang Guo, Terry Yue Zhuo, Qianru Zhang, Jie Zhang 0050, Heming Cui, Siu-Ming Yiu, Dong Huang 0005, See-Kiong Ng, Anh Tuan Luu |
NeurIPS | 1 |
| 2025 | DiffUIE: Learning Latent Global Priors in Diffusion Models for Underwater Image EnhancementabstractUnderwater imagery often suffers from light attenuation and color distortion, resulting in images with low contrast and blurriness. Enhancing these images is crucial yet challenging due to the complex degradation and noise inherent in underwater environments. In this study, we introduce a novel diffusion model, termed Underwater Image Enhancement(UIE) Diffusion, which leverages a global feature prior for effective underwater image enhancement. To our knowledge, this is the inaugural application of a diffusion model to the task of underwater image enhancement, setting a new benchmark in performance. Our approach begins with the introduction of a global feature prior to augment the diffusion model, mitigating the impact of noise and distortion during training. We then incorporate an underwater image degradation model to facilitate the learning of mappings between high-quality and degraded underwater images. To address over-enhancement caused by high-frequency components, we employ scaling factors to modulate the influence of frequency features during diffusion. Additionally, we enhance the model's stability during inference by integrating a backward diffusion process into its training. Comprehensive evaluations on multiple public datasets demonstrate that UIE Diffusion surpasses existing state-of-the-art methods in both subjective outcomes and objective assessments. Yuhao Qing, Si Liu 0001, Hai Wang 0004, Yueying Wang |
IEEE Trans. Multim. | 1 |
| 2025 | PipeMesh: Achieving Memory-Efficient Computation-Communication Overlap for Training Large Language ModelsabstractEfficiently training large language models (LLMs) on commodity cloud resources remains challenging due to limitations in network bandwidth and accelerator memory capacity. Existing training systems can be categorized based on their pipeline schedules. Depth-first scheduling, employed by systems like Megatron, prioritizes memory efficiency but restricts the overlap between communication and computation, causing accelerators to remain idle for over 20% of the training time. Conversely, breadth-first scheduling maximizes communication overlap but generates excessive intermediate activations, exceeding memory capacity and slowing computation by more than 34%. To address these limitations, we propose a novel elastic pipeline schedule that enables fine-grained control over the trade-off between communication overlap and memory consumption. Our approach determines the number of micro-batches scheduled together according to the communication time and the memory available. Furthermore, we introduce a mixed sharding strategy and a pipeline-aware selective recomputation technique to reduce memory usage. Experimental results demonstrate that our system eliminates most of the 28% all-accelerator idle time caused by communication, with recomputation accounting for less than 1.9% of the training time. Compared to existing baselines,PIPEMESHimproves training throughput on commodity clouds by 20.1% to 33.8%. Fanxin Li, Shixiong Zhao, Yuhao Qing, Jianyu Jiang, Xusheng Chen, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | EffiLearner: Enhancing Efficiency of Generated Code via Self-OptimizationabstractLarge language models (LLMs) have shown remarkable progress in code generation, but their generated code often suffers from inefficiency, resulting in longer execution times and higher memory consumption. To address this issue, we propose EffiLearner, a self-optimization framework that utilizes execution overhead profiles to improve the efficiency of LLM-generated code. EffiLearner first generates code using an LLM, then executes it locally to capture execution time and memory usage profiles. These profiles are fed back to the LLM, which then revises the code to reduce overhead. To evaluate the effectiveness of EffiLearner, we conduct extensive experiments on EffiBench and two commonly used code generation benchmarks with 16 open-source and 6 closed-source models. Our evaluation results demonstrate that through iterative self-optimization, EffiLearner significantly enhances the efficiency of LLM-generated code. For example, the execution time (ET) of StarCoder2-15B for the EffiBench decreases from 0.93 (s) to 0.12 (s) which reduces 87.1\% execution time requirement compared with the initial code. The total memory usage (TMU) of StarCoder2-15B also decreases from 22.02 (Mb*s) to 2.03 (Mb*s), which decreases 90.8\% total memory consumption during the execution process. Dong Huang 0005, Jianbo Dai, Han Weng, Puzhen Wu, Yuhao Qing, Heming Cui, Zhijiang Guo, Jie Zhang 0050 |
NeurIPS | 5 |
| 2024 | EffiBench: Benchmarking the Efficiency of Automatically Generated CodeabstractCode generation models have increasingly become integral to aiding software development. Although current research has thoroughly examined the correctness of the code produced by code generation models, a vital aspect that plays a pivotal role in greencomputing and sustainability efforts — the efficiency of the generated code — has often been neglected. This paper presents Effibench, a benchmark with 1,000 efficiency-critical coding problems to assess the efficiency of code generated by code generation models. EffiBench contains a diverse set of LeetCode coding problems. Each problem is paired with an executable human-written canonical solution, which obtains the SOTA efficiency on the LeetCode solution leaderboard. With EffiBench, we empirically examine the ability of 42 large language models (35 open-source and 7 closed-source) to generate efficient code. Our evaluation results demonstrate that the efficiency of the code generated by LLMs is generally worse than the efficiency of human-written canonical solutions. For example, GPT-4 generated code has an average \textbf{3.12} times execution time that of the human-written canonical solutions. In the most extreme cases, the execution time and total memory usage of GPT-4 code are \textbf{13.89} and \textbf{43.92} times that of the canonical solutions. The source code of EffiBench is released on https://github.com/huangd1999/EffiBench. We also provide the LeaderBoard in https://huggingface.co/spaces/EffiBench/effibench-leaderboard. Dong Huang 0005, Yuhao Qing, Weiyi Shang, Heming Cui, Jie Zhang 0050 |
NeurIPS | 2 |
| 2024 | Neuron Sensitivity-Guided Test Case SelectionabstractDeep neural networks (DNNs) have been widely deployed in software to address various tasks (e.g., autonomous driving, medical diagnosis). However, they can also produce incorrect behaviors that result in financial losses and even threaten human safety. To reveal and repair incorrect behaviors in DNNs, developers often collect rich, unlabeled datasets from the natural world and label them to test DNN models. However, properly labeling a large number of datasets is a highly expensive and time-consuming task. To address the above-mentioned problem, we propose neuron sensitivity-guided test case selection (NSS), which can reduce the labeling time by selecting valuable test cases from unlabeled datasets. NSS leverages the information of the internal neuron induced by the test cases to select valuable test cases, which have high confidence in causing the model to behave incorrectly. We evaluated NSS with four widely used datasets and four well-designed DNN models compared to the state-of-the-art (SOTA) baseline methods. The results show that NSS performs well in assessing the probability of failure triggering in test cases and in the improvement capabilities of the model. Specifically, compared to the baseline approaches, NSS achieves a higher fault detection rate (e.g., when selecting 5% of the test cases from the unlabeled dataset in the MNIST and LeNet1 experiment, NSS can obtain an 81.8% fault detection rate, which is a 20% increase compared with SOTA baseline strategies). Dong Huang 0005, Qingwen Bu, Yichao Fu, Yuhao Qing, Xiaofei Xie, Junjie Chen 0003, Heming Cui |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Fold3D: Rethinking and Parallelizing Computational and Communicational Tasks in the Training of Large DNN ModelsabstractTraining a large DNN (e.g., GPT3) efficiently on commodity clouds is challenging even with the latest 3D parallel training systems (e.g., Megatron v3.0). In particular, along the pipeline parallelism dimension, computational tasks that produce a whole DNN's gradients with multiple input batches should be concurrently activated; along the data parallelism dimension, a set of heavy-weight communications (for aggregating the accumulated outputs of computational tasks) isinevitably serializedafter the pipelined tasks, undermining the training performance (e.g., in Megatron, data parallelism caused all GPUs idle for over 44% of the training time) over commodity cloud networks. To deserialize these communicational and computational tasks, we propose the AIAO scheduling (for 3D parallelism) which slices a DNN into multiple segments, so that the computational tasks processing the same DNN segment can be scheduled together, and the communicational tasks that synchronize this segment can be launched and overlapped (deserialized) with other segments’ computational tasks. We realized this idea in ourFold3Dtraining system. Extensive evaluation showsFold3Deliminated most of the all-GPU 44% idle time in Megatron (caused by data parallelism), leading to 25.2%–42.1% training throughput improvement compared to four notable baselines over various settings;Fold3D's high performance scaled to many GPUs. Fanxin Li, Shixiong Zhao, Yuhao Qing, Xusheng Chen, Xiuxian Guan, Sen Wang 0004, Gong Zhang 0001, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Somatic variant analysis suite: copy number variation clonal visualization online platform for large-scale single-cell genomicsabstractThe recent advance of single-cell copy number variation (CNV) analysis plays an essential role in addressing intratumor heterogeneity, identifying tumor subgroups and restoring tumor-evolving trajectories at single-cell scale. Informative visualization of copy number analysis results boosts productive scientific exploration, validation and sharing. Several single-cell analysis figures have the effectiveness of visualizations for understanding single-cell genomics in published articles and software packages. However, they almost lack real-time interaction, and it is hard to reproduce them. Moreover, existing tools are time-consuming and memory-intensive when they reach large-scale single-cell throughputs. We present an online visualization platform, single-cell Somatic Variant Analysis Suite (scSVAS), for real-time interactive single-cell genomics data visualization. scSVAS is specifically designed for large-scale single-cell genomic analysis that provides an arsenal of unique functionalities. After uploading the specified input files, scSVAS deploys the online interactive visualization automatically. Users may conduct scientific discoveries, share interactive visualizations and download high-quality publication-ready figures. scSVAS provides versatile utilities for managing, investigating, sharing and publishing single-cell CNV profiles. We envision this online platform will expedite the biological understanding of cancer clonal evolution in single-cell resolution. All visualizations are publicly hosted at https://sc.deepomics.org. Lingxi Chen, Yuhao Qing, Ruikang Li, Chaohui Li, Hechen Li, Xikang Feng, Shuaicheng Li 0001 |
Briefings Bioinform. | 2 |
| 2022 | vPipe: A Virtualized Acceleration System for Achieving Efficient and Scalable Pipeline Parallel DNN TrainingabstractThe increasing computational complexity of DNNs achieved unprecedented successes in various areas such as machine vision and natural language processing (NLP), e.g., the recent advanced Transformer has billions of parameters. However, as large-scale DNNs significantly exceed GPU's physical memory limit, they cannot be trained by conventional methods such as data parallelism. Pipeline parallelism that partitions a large DNN into small subnets and trains them on different GPUs is a plausible solution. Unfortunately, the layer partitioning and memory management in existing pipeline parallel systems are fixed during training, making them easily impeded by out-of-memory errors and the GPU under-utilization. These drawbacks amplify when performing neural architecture search (NAS) such as the evolved Transformer, where different network architectures of Transformer needed to be trained repeatedly. vPipe is the first system that transparently provides dynamic layer partitioning and memory management for pipeline parallelism. vPipe has two unique contributions, including (1) an online algorithm for searching a near-optimal layer partitioning and memory management plan, and (2) a live layer migration protocol for re-balancing the layer distribution across a training pipeline. vPipe improved the training throughput of two notable baselines (Pipedream and GPipe) by 61.4-463.4 percent and 24.8-291.3 percent on various large DNNs and training settings. Shixiong Zhao, Fanxin Li, Xusheng Chen, Xiuxian Guan, Jianyu Jiang, Dong Huang 0005, Yuhao Qing, Sen Wang 0004, Peng Wang 0037, Gong Zhang 0001, Cheng Li 0001, Ping Luo 0002, Heming Cui |
IEEE Trans. Parallel Distributed Syst. | 7 |