Hongyang He

dblp:03/8056 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
abstract
Multimodal in-context learning (ICL) is becoming a key capability that allows large vision-language models (LVLMs) to adapt to novel tasks without parameter updates, which expands their usefulness in many real-world applications. However, ICL performance remains unstable even when the in-context demonstrations (ICDs) are well matched, showing that LVLMs still struggle to make full use of the provided context. While existing work mainly focuses on prompt engineering or post-hoc logit calibration, we study the attention mechanisms inside LVLMs to address their inherent limitations. We identify two important weaknesses in their self-attention that hinder effective ICL. To address these weaknesses, we propose Context-Aware Modulated Attention (CAMA), a training-free and plug-and-play method that dynamically adjusts attention logits based on the input in-context sequence. CAMA uses a two-stage modulation process that strengthens attention to semantically important tokens, especially visual ones. Across four LVLMs and seven benchmarks, CAMA consistently outperforms vanilla models and baselines, showing clear effectiveness and generalization. It can also activate the intended benefits of prompt engineering methods and remains robust across different sequence configurations. Therefore, CAMA opens up new directions for improving multimodal reasoning through a deeper understanding of attention dynamics.
Yanshu Li, Jianjiang Yang, Ziteng Yang, Bozheng Li, Ligong Han, Hongyang He, Zhengtao Yao, Victor Y. Chen, Songlin Fei, Dongfang Liu, Ruixiang Tang
AAAI6
2026 MoA: Heterogeneous Mixture of Adapters for Parameter-Efficient Fine-Tuning of Large Language Models
abstract
Jie Cao, Tianwei Lin, Bo Yuan, Rolan Yan, Hongyang He, Wenqiao Zhang, Juncheng Li, Dongping Zhang, Siliang Tang, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianwei Lin 0001, Rolan Yan, Hongyang He, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang
ACL (1)5
2026 Reinforcement Learning-based Adaptive Control of Classifier-Free Guidance and Timestep Embeddings in Diffusion Models
Haochen You, Baojing Liu, Hongyang He
WACV3
2026 MAKIMA: Tuning-free multi-attribute open-domain video editing via mask-guided attention modulation
Wenqiao Zhang, Hongyang He, Juncheng Li 0006, Zheqi Lv, Siliang Tang, Yueting Zhuang
Expert Syst. Appl.4
2025 Semi-ViM: Bidirectional State Space Model for Mitigating Label Imbalance in Semi-Supervised Learning
Hongyang He, Hongyang Xie, Haochen You, Victor Sanchez
ICCV1
2025 MSDet: Receptive Field Enhanced Multiscale Detection for Tiny Pulmonary Nodule
abstract
Pulmonary nodules are critical for early lung cancer diagnosis, but traditional CT imaging methods suffer from low detection rates and poor localization. Small nodule detection is challenging due to subtle differences in density and issues like occlusion. Existing methods such as FPN, with its fixed feature fusion and limited receptive field, struggle to effectively overcome these issues. To address these challenges, our paper proposed three key contributions: Firstly, we proposed MSDet, a multiscale attention and receptive field network for detecting tiny pulmonary nodules. Secondly, we proposed the extended receptive domain (ERD) strategy to capture richer contextual information and reduce false positives caused by nodule occlusion. We also proposed the position channel attention mechanism (PCAM) to optimize feature learning and reduce multiscale detection errors, and designed the tiny object detection block (TODB) to enhance the detection of tiny nodules. Experiments on the LUNA16 dataset show an 8.8% improvement in mAP over YOLOv8, achieving state-of-the-art performance. The code is available at https://github.com/CaiGuoHui123/MSDet.
Guohui Cai, Ruicheng Zhang, Hongyang He, Zeyu Zhang 0006, Daji Ergu, Yuanzhouhan Cao, Jinman Zhao, Binbin Hu, Zhibin Liao, Yang Zhao 0019, Ying Cai 0002
ICME3
2025 4S-Classifier: Empowering Conservation through Semi-Supervised Learning for Rare and Endangered Species
abstract
The survival of numerous endangered wildlife species is increasingly jeopardized by drastic climate changes, ecological disturbances, and human activities, leading to rapid population declines. Identifying endangered and rare species using computer vision is an important task that can aid in biodiversity conservation. However, many rare species are recognizable only by specialized biologists, making the labeling process of images both challenging and costly. Moreover, even with labeled data, training models with imbalanced datasets, i.e., datasets with under-represented categories, often results in inherent learning biases. To address high labeling costs and challenges associated with sparse label learning, we propose a novel semi-supervised learning framework, 4S-Classifier, which comprises two key modules: Rare Species Bank (RSBank) and Attention-based Rare Species Embedding (RSEmbed). The RSBank module stores embeddings of rare species across multiple training epochs, using clustering, kernel density estimation, and confidence scores to enhance the learning of under-represented categories. The RSEmbed module acts as a fusion-based augmentation approach that employs embeddings to improve model performance on these sparse and rare species data. By integrating both modules, our framework achieves a classification accuracy of 88.37% on an endangered species dataset (iNaturalist) and 91.56% on a wildlife dataset (Wildlife Insights) with only 25% of the labeled data, demonstrating its outstanding performance. The code is available at https://github.com/SIPLab24/4S-Classifier
Hongyang He, Hongyang Xie, Guodong Shen, Boyang Fu, Haochen You, Victor Sanchez
IJCNN1
2025 Twin Co-Adaptive Dialogue for Progressive Image Generation
abstract
Modern text-to-image generation systems have enabled the creation of remarkably realistic and high-quality visuals, yet they often falter when handling the inherent ambiguities in user prompts. In this work, we present Twin-Co, a framework that leverages synchronized, co-adaptive dialogue to progressively refine image generation. Instead of a static generation process, Twin-Co employs a dynamic, iterative workflow where an intelligent dialogue agent continuously interacts with the user. Initially, a base image is generated from the user's prompt. Then, through a series of synchronized dialogue exchanges, the system adapts and optimizes the image according to evolving user feedback. The co-adaptive process allows the system to progressively narrow down ambiguities and better align with user intent. Experiments demonstrate that Twin-Co not only enhances user experience by reducing trial-and-error iterations but also improves the quality of the generated images, streamlining creative process across various applications.
Jianhui Wang 0001, Yangfan He, Yan Zhong 0001, Xinyuan Song 0002, Jiayi Su, Yuheng Feng, Hongyang He, Wenyu Zhu, Xinhang Yuan, Miao Zhang 0010, Tianyu Shi 0003, Xueqian Wang 0001
ACM Multimedia8
2025 TRiCo: Triadic Game-Theoretic Co-Training for Robust Semi-Supervised Learning
abstract
We introduce TRiCo, a novel triadic game-theoretic co-training framework that rethinks the structure of semi-supervised learning by incorporating a teacher, two students, and an adversarial generator into a unified training paradigm. Unlike existing co-training or teacher-student approaches, TRiCo formulates SSL as a structured interaction among three roles: (i) two student classifiers trained on frozen, complementary representations, (ii) a meta-learned teacher that adaptively regulates pseudo-label selection and loss balancing via validation-based feedback, and (iii) a non-parametric generator that perturbs embeddings to uncover decision boundary weaknesses. Pseudo-labels are selected based on mutual information rather than confidence, providing a more robust measure of epistemic uncertainty. This triadic interaction is formalized as a Stackelberg game, where the teacher leads strategy optimization and students follow under adversarial perturbations. By addressing key limitations in existing SSL frameworks—such as static view interactions, unreliable pseudo-labels, and lack of hard sample modeling—TRiCo provides a principled and generalizable solution. Extensive experiments on CIFAR-10, SVHN, STL-10, and ImageNet demonstrate that TRiCo consistently achieves state-of-the-art performance in low-label regimes, while remaining architecture-agnostic and compatible with frozen vision backbones.
Hongyang He, Xinyuan Song 0002, Yangfan He, Yanshu Li, Haochen You, Lifan Sun, Wenqiao Zhang
NeurIPS1
2025 Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language Models
abstract
Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable achievements in recent years, they remain vulnerable to adversarial examples that result in harmful responses. Existing attacks typically focus on optimizing adversarial perturbations for a certain multimodal image-prompt pair or fixed training dataset, which often leads to overfitting. Consequently, these perturbations fail to remain malicious once transferred to attack unseen image-prompt pairs, suffering from significant resource costs to cover the diverse multimodal inputs in complicated real-world scenarios. To alleviate this issue, this paper proposes a novel adversarial attack on MLLMs based on distribution approximation theory, which models the potential image-prompt input distribution and adds the same distribution-fitting adversarial perturbation on multimodal input pairs to achieve effective cross-image/prompt transfer attacks. Specifically, we exploit the Laplace approximation to model the Gaussian distribution of the image and prompt inputs for the MLLM, deriving an estimate of the mean and covariance parameters. By sampling from this approximated distribution with Monte Carlo mechanism, we efficiently optimize and fit a single input‑agnostic perturbation over diverse image‑prompt pairs, yielding strong universality and transferability. Extensive experiments are conducted to verify the strong adversarial capabilities of our proposed attack against prevalent MLLMs spanning a spectrum of images/prompts.
Hai Yan, Haijian Ma, Xiaowen Cai 0001, Daizong Liu, Zenghui Yuan, Xiaoye Qu, Jianfeng Dong, Runwei Guan, Hongyang He, Yulai Xie 0002, Pan Zhou 0001
NeurIPS10
2025 Modular MeanFlow: Towards Stable and Scalable One-Step Generative Modeling
Haochen You, Baojing Liu, Hongyang He
PRCV (3)3
2024 Intelligent Video Monitoring and Analysis System for Power Grid Construction Site Safety Using Wireless Power Transfer
abstract
Power grid construction significantly enhances power grid management and risk control. Inconsistent operations at construction sites can jeopardize grid stability and crew safety. Traditional power lines are less favored due to mobility limitations, while batteries add burden and impracticality. To address this, a wireless power transfer and video monitoring system is developed using RF technology and Yolo V3 model. This enables continuous monitoring and employee safety analysis. The system's detection performance is optimized using HPSO, surpassing existing methods in accuracy and speed. It ensures real-time monitoring and improves the detection of potential risk sources, crucial for construction site safety.
Hongyang He
Int. J. Inf. Secur. Priv.2
2024 Adaptive Gravity-Aided Inertial Navigation Based on Characteristic Analysis of Marine Gravity Anomaly From Satellite Altimetry
abstract
With the improvement of resolution and precision of marine gravity anomaly data from satellite altimetry, the gravity matching navigation of underwater vehicles has been supported by background field data. To improve the accuracy, a multi-index model for analyzing the gravity background field characteristics is constructed, and an adaptive matching navigation algorithm is designed. First, the common single-index feature analysis methods are introduced, and their advantages and disadvantages are analyzed. The multi-index average statistical parameter (MASP) model applicable to the characterization of marine gravity anomaly background field is designed, and the reliability is verified by two satellite altimetric gravity anomaly datasets from the northwest Indian Ocean and the western Pacific Ocean. Then, an adaptive matching navigation algorithm is proposed by introducing the MASP model in the filtering process, which can adjust the filter structure in real time. Finally, the performance of the method is verified by about 30 h of field test. The matching results show that the proposed method can adjust the filter structure adaptively and has better comprehensive performance.
Jiangning Xu, Fangneng Li, Fangjun Qin, Hongyang He
IEEE Trans. Geosci. Remote. Sens.6
2023 Marine Gravity Anomaly From Satellite Altimetry: Interpolation and Matching Navigation
abstract
The background field of marine gravity anomaly from satellite altimetry is an important data support for gravity-aided navigation of underwater vehicles. To achieve long-endurance and high-precision navigation and positioning, an interpolation algorithm based on ensemble learning is proposed, and an adaptive matching navigation method is designed. First, the common interpolation theories of marine gravity anomaly background field are introduced, and the advantages and disadvantages of these theories are analyzed. The BP-Bagging ensemble learning algorithm for marine gravity anomaly background field interpolation is designed, and the reliability is verified by two satellite altimetric gravity anomaly datasets from the western Pacific Ocean. Then, a matching navigation algorithm based on the calculation method of gravity gradient deviation on survey line is proposed, which can evaluate the matching capability online according to the position of the underwater vehicle and adaptively adjusts the filter feedback coefficients. Finally, the performance of the method is verified by field tests. The matching results show that the proposed method improves navigation accuracy while enhancing the adaptive capability of the matching algorithm and providing better overall performance. Meanwhile, the gravity background field with high resolution yields better matching results.
Jiangning Xu, Fangjun Qin, Hongyang He
IEEE Trans. Geosci. Remote. Sens.6
2010 Polynomial root finding technique for joint DOA DOD estimation in bistatic MIMO radar
Mohamed Laid Bencheikh, Yide Wang, Hongyang He
Signal Process.3
2009 Focusing-based approach for wide-band source localization in near-field
abstract
A wide-band near-field source localization method is presented in this paper. Based on a pre-estimated source location, we form a focusing matrix which is able to compensate the wavefront distortion (with respect to far-field wavefront) due to near-field propagation and frequency dependent phase shift simultaneously. The focused covariance matrix has been proved to have a partially far-field narrow band structure, which allows us to estimate the bearings of the sources by the well studied far-field DOA estimators. The range estimation is then carried out via peak searching of the 1D MUSIC spectral with the estimated bearings. The performance of the algorithm is tested by simulations.
Hongyang He, Yide Wang, Joseph Saillard
ICASSP1