VLDB 2026 Research / reviewers in the wild / expert
Hanyang Peng
dblp:162/0123
· DBLP profile ↗
15ranked-venue papers
8as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Medverse: A Universal Model for Full-Resolution 3D Medical Image Segmentation, Transformation and EnhancementabstractIn-context learning (ICL) offers a promising paradigm for universal medical image analysis, enabling models to perform diverse image processing tasks without retraining. However, current ICL models for medical imaging remain limited in two critical aspects: they cannot simultaneously achieve high-fidelity predictions and global anatomical understanding, and there is no unified model trained across diverse medical imaging tasks (e.g., segmentation and enhancement) and anatomical regions. As a result, the full potential of ICL in medical imaging remains underexplored. Thus, we present Medverse, a universal ICL model for 3D medical imaging, trained on 22 datasets covering diverse tasks in universal image segmentation, transformation, and enhancement across multiple organs, imaging modalities, and clinical centers. Medverse employs a next-scale autoregressive in-context learning framework that progressively refines predictions from coarse to fine, generating consistent, full-resolution volumetric outputs and enabling multi-scale anatomical awareness. We further propose a blockwise cross-attention module that facilitates long-range interactions between context and target inputs while preserving computational efficiency through spatial sparsity. Medverse is extensively evaluated on a broad collection of held-out datasets covering previously unseen clinical centers, organs, species, and imaging modalities. Results demonstrate that Medverse substantially outperforms existing ICL baselines and establishes a novel paradigm for in-context learning. Jiesi Hu, Jianfeng Cao, Yanwu Yang 0001, Chenfei Ye, Hanyang Peng, Heather Ting Ma |
AAAI | 6 |
| 2026 | Symphony-MoE: Harmonizing Disparate Pre-trained Models into a Coherent Mixture-of-ExpertsabstractMixture-of-Experts (MoE) models enable scalable performance by activating large parameter sets sparsely, minimizing computational overhead. To mitigate the prohibitive cost of training MoEs from scratch, recent work employs upcycling, reusing a single pre-trained dense model by replicating its feed-forward network (FFN) layers into experts. However, this limits expert diversity, as all experts originate from a single pre-trained dense model. This paper addresses this limitation by constructing powerful MoE models using experts sourced from multiple identically-architected but disparate pre-trained models (e.g., Qwen2.5-Coder and Qwen2). A key challenge lies in the fact that these source models occupy disparate, dissonant regions of the parameter space, making direct upcycling prone to severe performance degradation. To overcome this, we propose Symphony-MoE, a novel two-stage framework designed to harmonize these models into a single, coherent expert mixture. First, we establish this harmony in a training-free manner: we construct a shared backbone via a layer-aware fusion strategy and, crucially, alleviate parameter misalignment among experts using activation-based functional alignment. Subsequently, a stage of post-training coordinates the entire architecture. Experiments demonstrate that our method successfully integrates experts from heterogeneous sources, achieving an MoE model that significantly surpasses baselines in multi-domain tasks and out-of-distribution generalization. Hanyang Peng |
AAAI | 2 |
| 2026 | RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image InterpretationabstractThe rapid advancement of foundation models has revolutionized visual representation learning in a self-supervised manner. However, their application in remote sensing (RS) remains constrained by a fundamental gap: existing models predominantly handle single or limited modalities, overlooking the inherently multi-modal nature of RS observations. Optical, synthetic aperture radar (SAR), and multi-spectral data offer complementary insights that significantly reduce the inherent ambiguity and uncertainty in single-source analysis. To bridge this gap, we introduce RingMoE, a unified multi-modal RS foundation model with 14.7 billion parameters, pre-trained on 400 million multi-modal RS images from nine satellites. RingMoE incorporates three key innovations: 1) A hierarchical Mixture-of-Experts (MoE) architecture comprising modal-specialized, collaborative, and shared experts, effectively modeling intra-modal knowledge while capturing cross-modal dependencies to mitigate conflicts between modal representations; 2) Physics-informed self-supervised learning, explicitly embedding sensor-specific radiometric characteristics into the pre-training objectives; 3) Dynamic expert pruning, enabling adaptive model compression from 14.7B to 1B parameters while maintaining performance, facilitating efficient deployment in Earth observation applications. Evaluated across 23 benchmarks spanning six key RS tasks (i.e., classification, detection, segmentation, tracking, change detection, and depth estimation), RingMoE outperforms existing foundation models and sets new SOTAs, demonstrating remarkable adaptability from single-modal to multi-modal scenarios. Beyond theoretical progress, it has been deployed and trialed in multiple sectors, including emergency response, land management, marine sciences, and urban planning. Hanbo Bi, Yingchao Feng, Boyuan Tong, Haichen Yu, Yongqiang Mao, Wenhui Diao, Peijin Wang, Yue Yu 0001, Hanyang Peng, Yehong Zhang, Kun Fu 0001, Xian Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2026 | High-Frequency Prioritized Sparse Attention Network for Image RestorationabstractImage restoration aims to restore high-quality images from degraded inputs caused by factors such as motion blur, defocus blur, and rain, where the primary difference between degraded and high-quality images lies in their high-frequency components. Despite the critical role of high frequencies in restoration, few methods explicitly prioritize computational resources for high frequencies over low frequencies. To address this issue, we propose a High-Frequency Prioritized Sparse Attention Network (HFP-SAN), a novel architecture for image restoration tasks. We explicitly prioritize high-frequency components by designing a symmetric encoder-decoder framework integrated with High-Frequency Selective Sparse Attention (HFSSA) modules while handling low-frequency components with a smaller residual network, thereby proportionally allocating computational resources based on their relative importance. HFSSA incorporates a Frequency-Selective Matching (FSM) algorithm to focus attention on strongly correlated high-frequency regions, mitigating computation on areas with weak correlations and irrelevant areas. Additionally, we introduce a dynamically adjustable high-frequency mask that guides the network to focus on the severely degraded regions, further refining restoration quality. The above designs ensure the final reconstructed image is a high-quality product. Experiments demonstrate that our HFP-SAN achieves state-of-the-art performance across multiple image restoration tasks, both quantitatively and qualitatively. Shuting Dong, Zhe Wu 0006, Hongyang Wei, Mingzhi Chen 0003, Guanghao Li 0003, Haolong Qian, Hanyang Peng, Chun Yuan 0003 |
IEEE Trans. Multim. | 7 |
| 2025 | Correcting Large Language Model Behavior via Influence FunctionabstractRecent advancements in AI alignment techniques have significantly improved the alignment of large language models (LLMs) with static human preferences. However, the dynamic nature of human preferences can render some prior training data outdated or even erroneous, ultimately causing LLMs to deviate from contemporary human preferences and societal norms. Existing methodologies, either curation of new data for continual alignment or manual correction of outdated data for re-alignment, demand costly human resources. To address this, we propose a novel approach, LLM BehAvior Correction with INfluence FunCtion REcall and Post-Training (LANCET), which needs no human involvement. LANCET consists of two phases: (1) using a new method LinFAC to efficiently identify the training data that significantly impact undesirable model outputs, and (2) applying an novel Influence-driven Bregman Optimization (IBO) technique to adjust the model’s outputs based on these influence distributions. Our experiments show that LANCET effectively and efficiently corrects inappropriate behaviors of LLMs while preserving model utility. Further more, LANCET exhibits stronger generalization ability than all baselines under out-of-distribution harmful prompts, offering better interpretability and compatibility with real-world applications of LLMs. Han Zhang 0025, Zhuo Zhang 0007, Yi Zhang 0127, Yuanzhao Zhai, Hanyang Peng, Yue Yu 0001, Hui Wang 0030, Bin Liang 0004, Lin Gui 0003, Ruifeng Xu 0001 |
AAAI | 5 |
| 2025 | Neuroverse3D: Developing in-Context Learning Universal Model for Neuroimaging in 3DabstractIn-context learning (ICL), a type of universal model, demonstrates exceptional generalization across a wide range of tasks without retraining by leveraging task-specific guidance from context, making it particularly effective for the intricate demands of neuroimaging. However, current ICL models, limited to 2D inputs and thus exhibiting suboptimal performance, struggle to extend to 3D inputs due to the high memory demands of ICL. In this regard, we introduce Neuroverse3D, an ICL model capable of performing multiple neuroimaging tasks in 3D (e.g., segmentation, denoising, inpainting). Neuroverse3D overcomes the large memory consumption associated with 3D inputs through adaptive parallel-sequential context processing and a U-shaped fusion strategy, allowing it to handle an unlimited number of context images. Additionally, we propose an optimized loss function to balance multi-task training and enhance focus on anatomical boundaries. Our study incorporates 43,674 3D multi-modal scans from 19 neuroimaging datasets and evaluates Neuroverse3D on 14 diverse tasks using held-out test sets. The results demonstrate that Neuroverse3D significantly outperforms existing ICL models and closely matches task-specific models, enabling flexible adaptation to medical center variations without retraining. The code and model weights are publicly available at https://github.com/jiesihu/Neuroverse3D. Jiesi Hu, Hanyang Peng, Yanwu Yang 0001, Xutao Guo, Yang Shang, Chenfei Ye, Heather Ting Ma |
ICCV | 2 |
| 2025 | VoiceMark: Zero-Shot Voice Cloning-Resistant Watermarking Approach Leveraging Speaker-Specific Latents
Haiyun Li, Jingran Xie, Yaoxun Xu, Hanyang Peng |
INTERSPEECH | 6 |
| 2024 | Re-Thinking the Effectiveness of Batch Normalization and BeyondabstractBatch normalization (BN) is used by default in many modern deep neural networks due to its effectiveness in accelerating training convergence and boosting inference performance. Recent studies suggest that the effectiveness of BN is due to the Lipschitzness of the loss and gradient, rather than the reduction of internal covariate shift. However, questions remain about whether Lipschitzness is sufficient to explain the effectiveness of BN and whether there is room for vanilla BN to be further improved. To answer these questions, we first prove that when stochastic gradient descent (SGD) is applied to optimize a general non-convex problem, three effects will help convergence to be faster and better: (i) reduction of the gradient Lipschitz constant, (ii) reduction of the expectation of the square of the stochastic gradient, and (iii) reduction of the variance of the stochastic gradient. We demonstrate that vanilla BN only with ReLU can induce the three effects above, rather than Lipschitzness, but vanilla BN with other nonlinearities like Sigmoid, Tanh, and SELU will result in degraded convergence performance. To improve vanilla BN, we propose a new normalization approach, dubbed complete batch normalization (CBN), which changes the placement position of normalization and modifies the structure of vanilla BN based on the theory. It is proven that CBN can elicit all the three effects above, regardless of the nonlinear activation used. Extensive experiments on benchmark datasets CIFAR10, CIFAR100, and ILSVRC2012 validate that CBN makes the training convergence faster, and the training loss converges to a smaller local minimum than vanilla BN. Moreover, CBN helps networks with multiple nonlinear activations (Sigmoid, Tanh, ReLU, SELU, and Swish) achieve higher test accuracy steadily. Specifically, benefitting from CBN, the classification accuracies for networks with Sigmoid, Tanh, and SELU are boosted by more than 15.0%, 4.5%, and 4.0% on average, respectively, which is even comparable to the performance for ReLU. Hanyang Peng, Yue Yu 0001, Shiqi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Birder: Communication-Efficient 1-bit Adaptive Optimizer for Practical Distributed DNN TrainingabstractVarious gradient compression algorithms have been proposed to alleviate the communication bottleneck in distributed learning, and they have demonstrated effectiveness in terms of high compression ratios and theoretical low communication complexity. However, when it comes to practically training modern deep neural networks (DNNs), these algorithms have yet to match the inference performance of uncompressed SGD-momentum (SGDM) and adaptive optimizers (e.g.,Adam). More importantly, recent studies suggest that these algorithms actually offer no speed advantages over SGDM/Adam when used with common distributed DNN training frameworks ( e.g., DistributedDataParallel (DDP)) in the typical settings, due to heavy compression/decompression computation or incompatibility with the efficient All-Reduce or the requirement of uncompressed warmup at the early stage. For these reasons, we propose a novel 1-bit adaptive optimizer, dubbed *Bi*nary *r*andomization a*d*aptive optimiz*er* (**Birder**). The quantization of Birder can be easily and lightly computed, and it does not require warmup with its uncompressed version in the beginning. Also, we devise Hierarchical-1-bit-All-Reduce to further lower the communication volume. We theoretically prove that it promises the same convergence rate as the Adam. Extensive experiments, conducted on 8 to 64 GPUs (1 to 8 nodes) using DDP, demonstrate that Birder achieves comparable inference performance to uncompressed SGDM/Adam, with up to ${2.5 \times}$ speedup for training ResNet-50 and ${6.3\times}$ speedup for training BERT-Base. Code is publicly available at https://openi.pcl.ac.cn/c2net_optim/Birder. Hanyang Peng, Shuang Qin |
NeurIPS | 1 |
| 2021 | Beyond softmax loss: Intra-concentration and inter-separability loss for classificationabstractIn the past years, most works have focused on designing an indigenous network architecture to advance progress in classification, but another potential opportunity for improvement, that is, research on classification losses, is underdeveloped. Although some new losses have been proposed, most of them either are variants of softmax loss or should combine with softmax loss. Hence, the inherent deficiencies of softmax loss, such as sensitiveness, class-balanced restriction, closed-set limitation, non-scale-invariance and incoordination between the intraclass distance and interclass distance, cannot be completely overcome. In light of this, we pave a new way to design a loss that has no relation to softmax loss and can avoid its weaknesses. We also propose an efficient algorithm to optimize the new loss that can circumvent computing the complicated gradients of a fraction, and the convergence is theoretically ensured. Extensive experimental results on benchmark datasets demonstrate that the new loss is competitive with state-of-the-art losses for classification. Additionally, other specially designed experiments show that the new loss is also effective at handling class-imbalanced problems, is robust in addressing outliers and can discover samples of unseen classes in open-set cases. Hanyang Peng, Shiqi Yu 0001 |
Neurocomputing | 1 |
| 2021 | A Systematic IoU-Related Method: Beyond Simplified Regression for Better LocalizationabstractFour-variable-independent-regression localization losses, such as Smooth-l1Loss, are used by default in modern detectors. Nevertheless, this kind of loss is oversimplified so that it is inconsistent with the final evaluation metric, intersection over union (IoU). Directly employing the standard IoU is also not infeasible, since the constant-zero plateau in the case of non-overlapping boxes and the non-zero gradient at the minimum may make it not trainable. Accordingly, we propose a systematic method to address these problems. Firstly, we propose a new metric, the extended IoU (EIoU), which is well-defined when two boxes are not overlapping and reduced to the standard IoU when overlapping. Secondly, we present the convexification technique (CT) to construct a loss on the basis of EIoU, which can guarantee the gradient at the minimum to be zero. Thirdly, we propose a steady optimization technique (SOT) to make the fractional EIoU loss approaching the minimum more steadily and smoothly. Fourthly, to fully exploit the capability of the EIoU based loss, we introduce an interrelated IoU-predicting head to further boost localization accuracy. With the proposed contributions, the new method incorporated into Faster R-CNN with ResNet50+FPN as the backbone yields 4.2 mAP gain on VOC2007 and 2.3 mAP gain on COCO2017 over the baseline Smooth-l1Loss, at almost no training and inferencing computational cost. Specifically, the stricter the metric is, the more notable the gain is, improving 8.2 mAP on VOC2007 and 5.4 mAP on COCO2017 at metric AP90. Hanyang Peng, Shiqi Yu 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Discriminative Feature Selection via Employing Smooth and Robust Hinge LossabstractA wide variety of sparsity-inducing feature selection methods have been developed in recent years. Most of the loss functions of these approaches are built upon regression since it is general and easy to optimize, but regression is not well suitable for classification. In contrast, the hinge loss (HL) of support vector machines has proved to be powerful to handle classification tasks, but a model with existing multiclass HL and sparsity regularization is difficult to optimize. In view of that, we propose a new loss, called smooth and robust HL, which gathers the merits of regression and HL but overcome their drawbacks, and apply it to our sparsity regularized feature selection model. To optimize the model, we present a new variant of accelerated proximal gradient (APG) algorithm, which boosts the discriminative margins among different classes, compared with standard APG algorithms. We further propose an efficient optimization technique to solve the proximal projection problem at each iteration step, which is a key component of the new APG algorithm. We theoretically prove that the new APG algorithm converges at rate O(1/k2) if it is convex (k is the iteration counter), which is the optimal convergence rate for smooth problems. Experimental results on nine publicly available data sets demonstrate the effectiveness of our method. Hanyang Peng, Cheng-Lin Liu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | A General Framework for Sparsity Regularized Feature Selection via Iteratively Reweighted Least Square MinimizationabstractA variety of feature selection methods based on sparsity regularization have been developed with different loss functions and sparse regularization functions. Capitalizing on the existing sparsity regularized feature selection methods, we propose a general sparsity feature selection (GSR-FS) algorithm that optimizes a ℓ2,r (0 < r ≤ 2) based loss function with a ℓ2,p-norm (0 < p ≤ 2) sparse regularization. The ℓ2,r-norm (0 < Hanyang Peng |
AAAI | 1 |
| 2017 | Feature selection by optimizing a lower bound of conditional mutual information
Hanyang Peng, Yong Fan 0001 |
Inf. Sci. | 1 |
| 2016 | Direct Sparsity Optimization Based Feature Selection for Multi-Class Classification
Hanyang Peng |
IJCAI | 1 |