VLDB 2026 Research / reviewers in the wild / expert
Yunqing Hu
dblp:151/4913
· DBLP profile ↗
16ranked-venue papers
4as first author
15since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Dual-granularity image-text alignment for zero-shot composed image retrieval
Wenjie Peng, Shuangping Huang, Yunqing Hu, Tianshui Chen |
Expert Syst. Appl. | 3 |
| 2026 | An block-diagonal elastic weight consolidation-attention mechanism LSTM enabled lifelong learning approach for lifetime prediction of insulated gate bipolar transistor
Tao Li 0068, Enyu Wang, Xingping Liu, Rongjun Ding, Yunqing Hu |
Expert Syst. Appl. | 5 |
| 2026 | Knowledge-embedded graph representation learning for document-level relation extraction
Jinglin Liang 0001, Yutao Qin, Shuangping Huang, Yunqing Hu, Xinwu Liu, Tianshui Chen |
Expert Syst. Appl. | 4 |
| 2026 | The unambiguous structure representation of tabular data for recognition
Fan Yang 0082, Junwen Tan, Tianshui Chen, Shuangping Huang, Yunqing Hu |
Neural Networks | 5 |
| 2026 | Improving Pseudo-Labeling by Dynamic Confidence Calibration for Semi-Supervised Sequence RecognitionabstractSequence recognition models under a fully supervised paradigm require large-scale training data, incurring substantial annotation costs. Pseudo-labeling is one of the most effective techniques in semi-supervised learning, which leverages predicted confidence to filter pseudo-labels on unlabeled data for model training. However, recent studies indicate that the performance of semi-supervised learning is compromised by overconfident models, as the predicted unreliable confidences will filter noisy samples into training. In this work, we discover that the overconfidence in sequence recognition models is influenced by the linguistic properties of a sequence, where the tail character classes are prone to be mispredicted as the head ones that frequently appear in the language with high confidence. And this overconfidence continuously intensifies throughout the semi-supervised training process. To address this limitation, we propose a Dynamic Sequential Class-Aware Smoothing (DSCS) method that calibrates the overconfidence of the head class to alleviate the inaccurate pseudo-labeling caused by overconfident misprediction to improve the quality of pseudo-labels. Specifically, we design a sequential class-aware smoothing module that incorporates token class frequency information to regularize the model and prevent it from becoming overconfident toward the head class. Meanwhile, to address the overconfidence problem intensifying throughout the semi-supervised learning processes, we introduce a dynamic regularization module to adjust the calibration strength dynamically for the coordination between the calibration and semi-supervised learning processes. Extensive experiments demonstrate the effectiveness and generality of our method, which significantly reduces annotation efforts while maintaining competitive recognition performance. Keke Xu, Zhenghua Peng, Shuangping Huang, Yunqing Hu, Wenjie Peng |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | A coordinate LVRT control for High-Frequency Link Matrix Converters with enhanced efficiencyabstractHigh-frequency link matrix converter (HFLMC) is attractive for renewable grid-connection system because of the advantages of its high power density and reliability. Due to eliminate the intermediate DC link, the grid side fault of the HFLMC will directly reflect to the high-frequency link side, which will result in the severe voltage and current oscillations in the high-frequency transformer. This paper proposes a coordinate low voltage ride-through (LVRT) control for HFLMC. The proposed control strategy actively enables the HFLMC to inject reactive power into the grid to provide voltage support for the grid. In addition, the proposed control strategy coordinate controls the phase shift angles of the HFLMC with the objective of the minimal RMS value of the current on high-frequency transformer (HFT). Simulation studies with the time-domain professional tool PLECS are conducted to confirm the effectiveness of the proposed control strategy. Yunqing Hu, Fujin Deng, Sayed Abulanwar, Abdelhady Ghanem |
IECON | 1 |
| 2025 | Order-Level Attention Similarity Across Language Models: A Latent CommonalityabstractIn this paper, we explore an important yet previously neglected question: Do context aggregation patterns across Language Models (LMs) share commonalities?
While some works have investigated context aggregation or attention weights in LMs, they typically focus on individual models or attention heads, lacking a systematic analysis across multiple LMs to explore their commonalities.
In contrast, we focus on the commonalities among LMs, which can deepen our understanding of LMs and even facilitate cross-model knowledge transfer.
In this work, we introduce the Order-Level Attention (OLA) derived from the order-wise decomposition of Attention Rollout and reveal that the OLA at the same order across LMs exhibits significant similarities.
Furthermore, we discover an implicit mapping between OLA and syntactic knowledge.
Based on these two findings, we propose the Transferable OLA Adapter (TOA), a training-free cross-LM adapter transfer method.
Specifically, we treat the OLA as a unified syntactic feature representation and train an adapter that takes OLA as input.
Due to the similarities in OLA across LMs, the adapter generalizes to unseen LMs without requiring any parameter updates.
Extensive experiments demonstrate that TOA's cross-LM generalization effectively enhances the performance of unseen LMs.
Code is available at \url{https://github.com/jinglin-liang/OLAS}. Jinglin Liang 0001, Jin Zhong 0001, Shuangping Huang, Yunqing Hu, Lixin Fan, Hanlin Gu |
NeurIPS | 4 |
| 2025 | A Low-Rank Enhanced Lightweight Multimodal LLM Framework for Efficient Edge Power Visual DetectionabstractWith the development of artificial intelligence, many visual detection methods have been applied to power systems. However, diverse tasks cause existing solutions to face low detection accuracy and resource constraints in edge power scenarios. In this paper, we propose a low-rank enhanced lightweight multimodal LLM framework for efficient edge power visual detection. The framework can effectively balance the model performance with the resource constraints of edge deployment by introducing multimodal LLM and fusing data augmentation strategies and low-rank optimization. First, to enhance the adaptability of diverse visual detection tasks, we design a diverse data augmentation strategy for multimodal LLM to solve the problem of insufficient power scene data. Then, we reduce the number of training parameters by a low-rank optimization technique to enable the model to run efficiently on resource-limited edge devices. Experimental results demonstrate that our proposed framework can improve the accuracy by more than 3% on average under different power visual detection tasks. It can also save 76.49% of graphics memory consumption and accelerate training by 81.88%. Zheming Yang, Yunqing Hu, Jingce Xu, Wancai Zhang |
SMC | 4 |
| 2025 | Heterogeneous Correlation Aware Regularization for Sequential Confidence CalibrationabstractDespite notable advancements across various tasks, deep sequence recognition models are shown to grapple with the dilemma of over-confidence, leading to unreliable predicted confidence, necessitating the need for calibration. Current efforts predominantly focus on classification model calibration, leaving the sequence recognition model calibration analysis underexplored and challenging. In this work, we discover that the primary reason for over-confidence in sequence recognition models stems from the one-hot encoding target sequence training paradigm and identify two distinct manifestations of over-confidence: perception and semantic context over-confidence. To address these challenges, we propose a heterogeneous correlation aware sequence regularization (HCSR) method that adaptively incorporates correlated sequences into training alongside the target sequence as additional supervision to regularize the probability of the target sequence from arbitrarily escalating. Specifically, a correlated sequence mining (CSM) model is designed, capable of efficiently mining heterogeneous correlated sequences, which can be flexibly customized to search for specific types of correlated sequences in demand to facilitate the calibration of corresponding types of over-confidence in the calibrating model, thereby achieving fine-grained calibration. Meanwhile, an adaptive calibration module is introduced to adaptively coordinate the optimization weights between the target sequence and correlated sequences, enabling the co-calibration among different samples. Comprehensive experiments conducted on several widely employed sequence recognition tasks demonstrate that the proposed method outperforms the current competing methods by a substantial margin. Zhenghua Peng, Tianshui Chen, Shuangping Huang, Yunqing Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Task Difficulty Aware Parameter Allocation & Regularization for Lifelong LearningabstractParameter regularization or allocation methods are effective in overcoming catastrophic forgetting in lifelong learning. However, they solve all tasks in a sequence uniformly and ignore the differences in the learning difficulty of different tasks. So parameter regularization methods face significant forgetting when learning a new task very different from learned tasks, and parameter allocation methods face unnecessary parameter overhead when learning simple tasks. In this paper, we propose the Parameter Allocation & Regularization (PAR), which adaptively select an appropriate strategy for each task from parameter allocation and regularization based on its learning difficulty. A task is easy for a model that has learned tasks related to it and vice versa. We propose a divergence estimation method based on the Nearest-Prototype distance to measure the task relatedness using only features of the new task. Moreover, we propose a time-efficient relatedness-aware sampling-based architecture search strategy to reduce the parameter overhead for allocation. Experimental results on multiple benchmarks demonstrate that, compared with SOTAs, our method is scalable and significantly reduces the model's redundancy while improving the model's performance. Further qualitative analysis indicates that PAR obtains reasonable task-relatedness. Wenjin Wang 0003, Yunqing Hu, Qianglong Chen, Yin Zhang 0006 |
CVPR | 2 |
| 2023 | Dual Collaborative Visual-Semantic Mapping for Multi-Label Zero-Shot Image RecognitionabstractMulti-label zero-shot learning (ML-ZSL), with the difficulty of both multi-label learning and zero-shot learning, aims to recognize various unseen objects that are not observed during training. Previous methods mainly use a single directional visual-semantic mapping to associate the visual and semantic embedding space, which is not sufficient to adequately realize knowledge transfer from seen to unseen classes. In this paper, we propose a novel dual collaborative visual-semantic mapping framework, constructing abundant connection relationships by exploring two aspects of mapping streams, i.e., the visual-to-semantic (V2S) mapping and the semantic-to-visual (S2V) mapping. Through the collaborative learning of these two effective mappings, our method achieves state-of-the-art performance on the MS-COCO and PASCAL-VOC, two benchmarks for ML-ZSL. Yunqing Hu, Xuan Jin, Yin Zhang 0006 |
ICASSP | 1 |
| 2022 | Diverse Instance Discovery: Vision-Transformer for Instance-Aware Multi-Label Image RecognitionabstractPrevious works on multi-label image recognition (MLIR) usually use CNNs as a starting point for research. In this paper, we take pure Vision Transformer (ViT) as the research base and make full use of the advantages of Transformer with long-range dependency modeling to circumvent the disadvantages of CNNs limited to local receptive field. However, for multi-label images containing multiple objects from different categories, scales, and spatial relations, it is not optimal to use global information alone. Our goal is to leverage ViT's patch tokens and self-attention mechanism to mine rich instances in multi-label images, named diverse instance discovery (DiD). To this end, we propose a semantic category-aware module and a spatial relationship-aware module, respectively, and then combine the two by a re-constraint strategy to obtain instance-aware attention maps. Finally, we propose a weakly supervised object localization-based approach to extract multi-scale local features, to form a multi-view pipeline. Our method requires only weakly supervised information at the label level, no additional knowledge injection or other strongly supervised information is required. Experiments on three benchmark datasets show that our method significantly outperforms previous works and achieves state-of-the-art results under fair experimental comparisons. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Feihu Yan, Yuan He 0011, Hui Xue 0001 |
ICME | 1 |
| 2021 | DRDF: Determining the Importance of Different Multimodal Information with Dual-Router Dynamic FrameworkabstractIn multimodal tasks, the importance of text and image modal information often varies for different input cases. To model the difference of importance of different modal information, we propose a high-performance and highly general Dual-Router Dynamic Framework (DRDF), consisting of Dual-Router, MWF-Layer, experts and expert fusion unit. The text router and image router in Dual-Router take text modal information and image modal information respectively, and MWF-Layer is responsible to determine the importance of modal information. Based on the result of the determination, MWF-Layer generates fused weights for the subsequent experts fusion. Experts can adopt a variety of backbones that match the current multimodal or unimodal task. DRDF features high generality and modularity, and we test 12 backbones such as Visual BERT and their corresponding DRDF instances on the multimodal dataset Hateful memes, and unimodal datasets CIFAR10, CIFAR100, and TinyImagenet. Our DRDF instance outperforms those backbones. We also validate the effectiveness of components of DRDF by ablation studies, and discuss the reasons and ideas of DRDF design. Haiwen Hong, Xuan Jin, Yin Zhang 0006, Yunqing Hu, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 4 |
| 2021 | RAMS-Trans: Recurrent Attention Multi-scale Transformer for Fine-grained Image RecognitionabstractIn fine-grained image recognition (FGIR), the localization and amplification of region attention is an important factor, which has been explored extensively convolutional neural networks (CNNs) based approaches. The recently developed vision transformer (ViT) has achieved promising results in computer vision tasks. Compared with CNNs, Image sequentialization is a brand new manner. However, ViT is limited in its receptive field size and thus lacks local attention like CNNs due to the fixed size of its patches, and is unable to generate multi-scale features to learn discriminative region attention. To facilitate the learning of discriminative region attention without box/part annotations, we use the strength of the attention weights to measure the importance of the patch tokens corresponding to the raw images. We propose the recurrent attention multi-scale transformer (RAMS-Trans), which uses the transformer's self-attention to recursively learn discriminative region attention in a multi-scale manner. Specifically, at the core of our approach lies the dynamic patch proposal module (DPPM) responsible for guiding region amplification to complete the integration of multi-scale image patches. The DPPM starts with the full-size image patches and iteratively scales up the region attention to generate new patches from global to local by the intensity of the attention weights generated at each scale as an indicator. Our approach requires only the attention weights that come with ViT itself and can be easily trained end-to-end. Extensive experiments demonstrate that RAMS-Trans performs better than exising works, in addition to efficient CNN models, achieving state-of-the-art results on three benchmark datasets. Yunqing Hu, Xuan Jin, Yin Zhang 0006, Haiwen Hong, Jingfeng Zhang, Yuan He 0011, Hui Xue 0001 |
ACM Multimedia | 1 |
| 2021 | A Vibration Control Method for Hybrid-Structured Flexible Manipulator Based on Sliding Mode Control and Reinforcement LearningabstractThe hybrid-structured flexible manipulator has a complex structure and strong coupling between state variables. Meanwhile, the natural frequency of the hybrid-structured flexible manipulator varies with the motion of the telescopic joint, so it is difficult to suppress the vibration quickly. In this article, the tip state signal of the hybrid-structured flexible manipulator is decomposed into elastic vibration signal and tip vibration equilibrium position signal, and a combined control method is proposed to improve tip positioning accuracy and trajectory tracking accuracy. In the proposed combined control method, an improved nominal model-based sliding mode controller (NMBSMC) is used as the main controller to output the driving torque, and an actor-critic-based reinforcement learning controller (ACBRLC) is used as an auxiliary controller to output small compensation torque. The improved NMBSMC can be divided into a nominal model-based sliding mode robust controller and a practical model-based integral sliding mode controller. Two sliding mode controllers with different structures make full use of the mathematical model and the measured data of the actual system to improve the vibration equilibrium position tracking accuracy. The ACBRLC uses the tip elastic vibration signal and the prioritized experience replay method to obtain the small reverse compensation torque, which is superimposed with the output of the NMBSMC to suppress tip vibration and improve the positioning accuracy of the hybrid-structured flexible manipulator. Finally, several groups of experiments are designed to verify the effectiveness and robustness of the proposed combined control method. En Li 0001, Yunqing Hu, Lei Yang 0053, Junfeng Fan, Zi-ze Liang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | ChamNet: Towards Efficient Network Design Through Platform-Aware Model AdaptationabstractThis paper proposes an efficient neural network (NN) architecture design methodology called Chameleon that honors given resource constraints. Instead of developing new building blocks or using computationally-intensive reinforcement learning algorithms, our approach leverages existing efficient network building blocks and focuses on exploiting hardware traits and adapting computation resources to fit target latency and/or energy constraints. We formulate platform-aware NN architecture search in an optimization framework and propose a novel algorithm to search for optimal architectures aided by efficient accuracy and resource (latency and/or energy) predictors. At the core of our algorithm lies an accuracy predictor built atop Gaussian Process with Bayesian optimization for iterative sampling. With a one-time building cost for the predictors, our algorithm produces state-of-the-art model architectures on different platforms under given constraints in just minutes. Our results show that adapting computation resources to building blocks is critical to model performance. Without the addition of any special features, our models achieve significant accuracy improvements relative to state-of-the-art handcrafted and automatically designed architectures. We achieve 73.8% and 75.3% top-1 accuracy on ImageNet at 20ms latency on a mobile CPU and DSP. At reduced latency, our models achieve up to 8.2% (4.8%) and 6.7% (9.3%) absolute top-1 accuracy improvements compared to MobileNetV2 and MnasNet, respectively, on a mobile CPU (DSP), and 2.7% (4.6%) and 5.6% (2.6%) accuracy gains over ResNet-101 and ResNet-152, respectively, on an Nvidia GPU (Intel CPU). Xiaoliang Dai, Peizhao Zhang, Bichen Wu, Hongxu Yin, Fei Sun 0002, Yanghan Wang, Marat Dukhan, Yunqing Hu, Yiming Wu 0013, Yangqing Jia, Peter Vajda, Matthew Uyttendaele, Niraj K. Jha |
CVPR | 8 |