Weiquan Huang

dblp:18/8078 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation
abstract
CLIP is a seminal multimodal model that maps images and text into a shared representation space by contrastive learning on billions of image–caption pairs. Inspired by the rapid progress of large language models (LLMs), we investigate how the superior linguistic understanding and broad world knowledge of LLMs can further strengthen CLIP—particularly in handling long, complex captions. We introduce an efficient fine-tuning framework that embeds an LLM into a pretrained CLIP while incurring almost the same training cost as regular CLIP fine-tuning. Our method first “embedding-izes” the LLM for the CLIP setting, then couples it to the pretrained CLIP vision encoder through a lightweight adaptor trained on only a few million image–caption pairs. With this strategy we achieve large performance gains—without large-scale retraining—over state-of-the-art CLIP variants such as EVA02 and SigLIP-2. The LLM-enhanced CLIP delivers consistent improvements across a wide spectrum of downstream tasks, including linear-probe classification, zero-shot image–text retrieval with both short and long captions (in English and other languages), zero-shot/supervised image segmentation, object detection, and used as tokenizer for multimodal large-model benchmarks.
Weiquan Huang, Aoqi Wu, Yifan Yang 0004, Xufang Luo, Yuqing Yang 0001, Usman Naseem, Chunyu Wang 0001, Qi Dai 0001, Xiyang Dai, Dongdong Chen 0001, Chong Luo 0001, Lili Qiu, Liang Hu 0004
AAAI1
2026 AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information Minimization
abstract
The escalating energy consumption and carbon footprint of training large-scale Web AI models pose urgent challenges for sustainable development. Dataset Distillation (DD) offers a promising avenue for green AI by compressing large datasets into small synthetic ones for efficient training. However, most existing DD methods overfit to the inductive biases of specific source architectures (e.g., CNNs or ViTs), resulting in poor cross-model generalization. This limitation necessitates redundant re-distillation processes for different architectures, severely undermining the energy-saving potential of DD. To address this, we introduce Adversarial Mutual Information Distillation (AMID), a rigorous framework designed to create highly reusable and robust synthetic datasets. From an information-theoretic perspective, we cast model-agnosticism as minimizing the mutual information (MI) between the synthetic data and the specific identity of the distillation model. We convert this intractable objective into a tractable two-player adversarial game, which unifies knowledge preservation with adversarial unlearning of architectural bias. Extensive experiments on CIFAR-10 and Tiny ImageNet demonstrate that AMID achieves state-of-the-art cross-architecture generalization across diverse CNNs and ViTs. Crucially, our analysis confirms that AMID significantly reduces the computational overhead and CO2 emissions of downstream training while maintaining robust performance, paving the way for energy-efficient, transferable, and sustainable Web AI ecosystems.
Aoqi Wu, Weiquan Huang, Liang Hu 0004, Yifan Yang 0004, Qi Zhang 0020, Jiaxing Miao, Yuhan Tang, Zhongyuan Lai
WWW4
2026 Distributed Event-Driven Optimal Control for Networked Unmanned Surface Vehicles With Long Transmission Delays: A Memory-Based Estimation and Communication Assignment Approach
abstract
The Internet of Things (IoT), comprising networked unmanned surface vehicles (NUSVs), enables significant advancements in distributed optimal coordination techniques for marine activities. However, constrained communication bandwidth and long transmission delays pose substantial challenges to achieving optimal operations of NUSVs. To address these challenges, this paper proposes a delay-tolerant distributed optimation method (DTDOM) based on a time assignment event-driven communication mechanism (TAEDCM). In this approach, the memory-based state estimator (MSE) is developed to provide reliable data support for implementing the TAEDCM-based DTDOM. Specifically, the MSE within DTDOM estimates neighbors’ states using historical data pairs during packet gaps, while the one within TAEDCM predicts possible neighbor estimates from locally recorded data. Additionally, a communication assignment approach (CAA) integrated into TAEDCM is proposed to allocate interaction activation slots for each vehicle, preventing channel competition caused by simultaneous data exchanges. Compared with existing methods, the proposed framework offers advantages in actively compensating the delay-induced lag errors, determining appropriate communication timings under long delays, and ensuring competition-free inter-vehicle interactions. Theoretical analysis and semi-physical simulations validate the effectiveness and superiority of proposed solution.
Weiquan Huang
IEEE Internet Things J.4
2026 Sea-Trial Validated Adaptive Control for Uncrewed Surface Vehicles: Integrating Robust Line-of-Sight With Dual-Mode Model Predictive Control
abstract
This paper addresses path following control for underactuated unmanned surface vehicle (USV) equipped with towed side-scan sonars, where dynamic coupling and nonlinear disturbances pose significant challenges. Existing methods often lack experimental validation in realistic sea trials. To overcome these problems, we propose an adaptive robust line-of-sight model predictive control (ARLOS-MPC) scheme integrated with dual-mode artificial bee colony (ABC) optimization. Our approach enhances robustness against kinematic discrepancies and hydrodynamic disturbances under operational constraints. Key innovations include an adaptive Line-of-Sight (LOS) method employing sliding mode concepts with a speed-dependent lookahead distance, a dual-mode predictive control algorithm that incorporates path geometry and environmental disturbances via curvature change rates and dynamic weights, and a collaborative mechanism coupling step size and prediction horizon for synergistic global-local search. Empirical validation through simulations and sea trials under varied conditions shows that ARLOS-MPC reduces path following errors, accelerates convergence, improves robustness, and enhances adaptability, outperforming existing methods in computational efficiency and error reduction.
Lihui Deng, Hongjian Wang 0003, Weiquan Huang, Zhikang Chi, Jingfei Ren
IEEE Trans Autom. Sci. Eng.3
2025 A Reinforcement Learning-Enhanced Dung Beetle Optimization Approach for Agile Earth Observation Satellite Scheduling
abstract
Due to the increasing demand for remote sensing imaging products, the agile earth observation satellite scheduling problem (AEOSSP) has garnered significant attention. In response, this paper proposes a reinforcement learning-based dung beetle optimization (RLDBO) algorithm to address AEOSSP. The proposed method dynamically adjusts the proportions of four types of dung beetles (ball-rolling beetle, brood ball beetle, small dung beetle, and thief beetle), enabling adaptive optimization of the scheduling scheme to better handle the complexities and uncertainties of the search space. The Q-learning mechanism guides the adjustment of these proportions, effectively balancing global exploration and local exploitation at different stages of the search process. Experimental results demonstrate that the RLDBO algorithm effectively solves AEOSSP across multiple instances, and it outperforms other algorithms in various aspects, including optimization performance, convergence speed, and scheduling effectiveness. The experimental validation confirms that RLDBO significantly enhances the efficiency and effectiveness of agile earth observation satellite scheduling.
Weiquan Huang, He Wang 0039, Junyu Wu, Haoyu Hou, Yanjie Song 0001
IEEE Geosci. Remote. Sens. Lett.1
2025 Exploring Causal Information Bottleneck for Adversarial Defense
abstract
Information bottleneck (IB) is a promising defense solution against adversarial attacks on deep neural networks. However, these methods often suffer from spurious correlations. A correlation exists between the prediction and the non-robust features, yet it does not reflect the causal relationship well. Such spurious correlations induce the neural networks to learn fragile and incomprehensible (non-robust) features. This issue limits its potential for further improving adversarial robustness. This paper addresses this issue by incorporating causal inference into the IB-based defense framework. Specifically, we propose a novel defense method that use the instrumental variables to enhance the adversarial robustness. Our proposed method divides the features into two parts for causal effect estimation: robust and non-robust features. The robust features relate to understanding semantic information, and the non-robust features link to the vulnerable style information. By employing this framework, the IB method can mitigate the influence of non-robust features and extract the robust features linking to the semantic information of objects. We conduct a thorough analysis of the effectiveness of our proposed method. Notably, the experiments on MNIST, FashionMNIST, CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that our method significantly boosts the adversarial robustness against multiple adversarial attacks compared to previous methods. Our regularization method can improve adversarial robustness in both natural and adversarial training frameworks. Besides, CausalIB can be applied to both Convolutional Neural Networks and Vision Transformers as a plug-and-play module. Our code is available at https://github.com/HydrogenWasser/CausalIB.
Jun Yan 0013, Huan Hua, Weiquan Huang, Wancheng Ge, Jiancheng Yang
IEEE Trans. Inf. Forensics Secur.3
2024 Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
abstract
Jiaxing Sun, Weiquan Huang, Jiang Wu, Chenya Gu, Wei Li, Songyang Zhang, Hang Yan, Conghui He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Jiaxing Sun 0001, Weiquan Huang, Jiang Wu 0003, Chenya Gu, Wei Li 0320, Songyang Zhang 0001, Hang Yan 0001, Conghui He
ACL (1)2
2024 Optimal Observation Planning Method for Agile Earth Observation Satellites Based on Improved Ant Colony Algorithm
abstract
This study focuses on generating the optimal sequence of point target observations in dense areas, adopting a new optical agile satellite imaging approach: staring observation, and proposes an efficient and high-yield observation target clustering method along with a heuristic ant colony optimization algorithm. Firstly, aiming at the staring observation characteristics of agile satellites, this paper designs a group division strategy based on the earliest end time and, by synthesizing dense point targets to reduce the satellite’s maneuvering frequency and energy consumption, thereby enhancing observation efficiency. Secondly, considering constraints such as the satellite’s visible time windows, maneuverability, resource limitations, and the solar elevation angle, we constructed a multi-objective optimization model targeting observation benefits and energy consumption. To solve this model, an improved heuristic ant colony optimization (IHACO) algorithm is proposed, which takes task intervals, priorities, and feasible start time lengths as heuristic information and introduces Lévy flight to improve the pheromone evaporation mechanism, thus enhancing the algorithm’s ability to avoid falling into local optima. Finally, the efficiency of the algorithm is verified through simulation experiments and application scenarios of large-scale dense point targets.
Weiquan Huang, He Wang 0039, Dongmei Bai, Yanjie Song 0003
IECON1
2024 Sparse Query Dense: Enhancing 3D Object Detection with Pseudo Points
abstract
Current LiDAR-only 3D detection methods are limited by the sparsity of point clouds. The previous method used pseudo points generated by depth completion to supplement the LiDAR point cloud, but the pseudo points sampling process was complex, and the distribution of pseudo points was uneven. Meanwhile, due to the imprecision of depth completion, the pseudo points suffer from noise and local structural ambiguity, which limit the further improvement of detection accuracy. This paper presents SQDNet, a novel framework designed to address these challenges. SQDNet incorporates two key components: the SQD, which achieves sparse-to-dense matching via grid position indices, allowing for rapid sampling of large-scale pseudo points on the dense depth map directly, thus streamlining the data preprocessing pipeline. And use the density of LiDAR points within these grids to alleviate the uneven distribution and noise problems of pseudo points. Meanwhile, the sparse 3D Backbone is designed to capture long-distance dependencies, thereby improving voxel feature extraction and mitigating local structural blur in pseudo points. The experimental results validate the effectiveness of SQD and achieve considerable detection performance for difficult-to-detect instances on the KITTI test.
Yujian Mo, Yan Wu 0011, Junqiao Zhao, Zhenjie Hou, Weiquan Huang, Jun Yan 0009
ACM Multimedia5
2024 Multistage Compression Optimization Strategies for Accelerating Diffusion Models
Weiquan Huang
PRCV (1)1
2024 A Strategy Fusion-Based Multiobjective Optimization Approach for Agile Earth Observation Satellite Scheduling Problem
abstract
Agile satellite imaging scheduling plays a vital role in improving emergency response, urban planning, national defense, and resource management. With the rise in the number of in-orbit satellites and observation windows, the need for diverse agile Earth observation satellite (AEOS) scheduling has surged. However, current research seldom addresses multiple optimization objectives, which are crucial in many engineering practices. This article tackles a multiobjective AEOS scheduling problem (MOAEOSSP) that aims to optimize total observation task profit, satellite energy consumption, and load balancing. To address this intricate problem, we propose a strategy-fused multiobjective dung beetle optimization (SFMODBO) algorithm. This novel algorithm harnesses the position update characteristics of various dung beetle populations and integrates multiple high-adaptability strategies. Consequently, it strikes a better balance between global search capability and local exploitation accuracy, making it more effective at exploring the solution space and avoiding local optima. The SFMODBO algorithm enhances global search capabilities through diverse strategies, ensuring thorough coverage of the search space. Simultaneously, it significantly improves local optimization precision by fine-tuning solutions in promising regions. This dual approach enables more robust and efficient problem-solving. Simulation experiments confirm the effectiveness and efficiency of the SFMODBO algorithm. Results indicate that it significantly outperforms competitors across multiple metrics, achieving superior scheduling schemes. In addition to these enhanced metrics, the proposed algorithm also exhibits advantages in computation time and resource utilization. This not only demonstrates the algorithm’s robustness but also underscores its efficiency and speed in solving the MOAEOSSP.
He Wang 0039, Weiquan Huang, Sindri Magnússon, Tony Lindgren, Yanjie Song 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Semi-Supervised Semantic Segmentation with Structured Output Space Adaption
abstract
Semi-supervised semantic segmentation methods rely on dense pixel-level classification with limited data and can thus be developed to adapt source ground truth labels to a target domain. In this paper, we creatively propose a method for semi-supervised semantic segmentation. The key innovation is our adversarial learning method for space adaptation in context, which can be regarded as a structured output that contains spatial similarities between unlabeled data and labeled data. To achieve this, we construct an adversarial learning network to efficiently adapt to the structural output space in labeled and unlabeled similar samples. Furthermore, we introduce two learning strategies into semi-supervised semantic segmentation, one that can selectively capture intra-category and inter-category context dependencies, resulting in robust feature representations. While the other explicitly concatenates the shape information of objects as a separate processing branch to produce sharper predicted boundaries of objects. Experimental results on two well-known benchmark datasets show that our method achieves better performance compared to the previous competitive models.
Weiquan Huang, Fu Zhang 0001
ICASSP1
2023 Attentive Mask CLIP
abstract
In vision-language modeling, image token removal is an efficient augmentation technique to reduce the cost of encoding image features. The CLIP-style models, however, have been found to be negatively impacted by this technique. We hypothesize that removing a large portion of image tokens may inadvertently destroy the semantic information associated to a given text description, resulting in misaligned paired data in CLIP training. To address this issue, we propose an attentive token removal approach, which retains a small number of tokens that have a strong semantic correlation to the corresponding text description. The correlation scores are dynamically evaluated through an EMA-updated vision encoder. Our method, termed attentive mask CLIP, outperforms original CLIP and CLIP variant with random token removal while saving the training time. In addition, our approach also enables efficient multi-view contrastive learning. Experimentally, by training ViT-B on YFCC-15M dataset, our approach achieves 43.9% top-1 accuracy on ImageNet-1K zero-shot classification, 62.7/42.1 and 38.0/23.2 I2T/T2I retrieval accuracy on Flickr30K and MS COCO, outperforming SLIP by +1.1%, +5.5/+0.9, and +4.4/+1.3, respectively, while being 2.30× faster. An efficient version of our approach runs 1.16× faster than the plain CLIP model, while achieving significant gains of +5.3%, +11.3/+8.0, and +9.5/+4.9 on these benchmarks, respectively. Code will be release in https://github.com/microsoft/A-CLIP.
Yifan Yang 0004, Weiquan Huang, Yixuan Wei, Houwen Peng, Xinyang Jiang, Huiqiang Jiang, Fangyun Wei, Han Hu 0001, Lili Qiu, Yuqing Yang 0001
ICCV2
2022 Seeing Speech: Magnetic Resonance Imaging-Based Vocal Tract Deformation Visualization Using Cross-Modal Transformer
abstract
As an essential component to advance speech science, understanding of speech production can be greatly helpful to improve our understanding of motor control, dynamical systems of humans during natural speech. Different medical imaging modalities have been leveraged to visualize the dynamic process, in which Magnetic resonance imaging (MRI) provides a valuable tool for evaluating static postures. In this demo, we present our solution to visualize the vocal tract deformation, leveraging the correlation between the MRI and the acoustical signals. We first formulate the problem as a cross-modal prediction task and a novel cross-modal Transformer network is proposed. Thus, we can infer the deformation of the vocal tract by only utilizing the acoustical signals. Then, we present an interactive framework, which can be used to visualize the deformation utilizing the aforementioned network. We hope our solution can also be helpful in pronunciation training for children with sound speech disorders and second language learning.
Kele Xu, Ming Feng, Weiquan Huang
ACM Multimedia3
2020 Intra-Camera Supervised Person Re-ID by Tracklet Level Classifier
Weiquan Huang
PRCV (2)2