Qiushi Huang

dblp:204/2933 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 KICGPTv2: Large Language Model With Knowledge in Context for Knowledge Graph Completion
abstract
Knowledge Graph Completion (KGC) is an essential task aimed at mitigating the issue of incompleteness in knowledge graphs, thereby enhancing their utility for various downstream applications. Existing KGC models predominantly fall into two categories: structure-based and semantic-based approaches. Structure-based methods often encounter challenges with long-tail entities due to the scarcity of structural information and imbalanced entity distributions. Conversely, semantic-based methods, while addressing those limitations, necessitate extensive training of language models and specific finetuning for each knowledge graph, thus constraining their practical efficiency. To alleviate those limitations in both approaches, in this paper, we propose KICGPTv2, an innovative framework that synergizes a large language model (LLM) with traditional KGC methods. This integration effectively mitigates the long-tail entity problem without incurring significant additional training overhead. Central to the KICGPTv2 model is a novel in-context learning strategy, termed Knowledge Prompt, which encodes structural knowledge into demonstrations to effectively guide the LLM. Comprehensive evaluations on various KGC tasks, including link prediction, relation prediction, and triple classification, underscore the efficacy of the KICGPTv2 model, highlighting its ability to achieve competitive performance with reduced training demands and without the need for finetuning
Yanbin Wei, Qiushi Huang, James T. Kwok, Yu Zhang 0006
IEEE Trans. Knowl. Data Eng.2
2025 ComLoRA: A Competitive Learning Approach for Enhancing LoRA
abstract
We propose a Competitive Low-Rank Adaptation (ComLoRA) framework to address the limitations of the LoRA method, which either lacks capacity with a single rank-$r$ LoRA or risks inefficiency and overfitting with a larger rank-$Kr$ LoRA, where $K$ is an integer larger than 1. The proposed ComLoRA method initializes $K$ distinct LoRA components, each with rank $r$, and allows them to compete during training. This competition drives each LoRA component to outperform the others, improving overall model performance. The best-performing LoRA is selected based on validation metrics, ensuring that the final model outperforms a single rank-$r$ LoRA and matches the effectiveness of a larger rank-$Kr$ LoRA, all while avoiding extra computational overhead during inference. To the best of our knowledge, this is the first work to introduce and explore competitive learning in the context of LoRA optimization. The ComLoRA's code is available at https://github.com/hqsiswiliam/comlora.
Qiushi Huang, Tom Ko, Lilian Tang, Yu Zhang 0006
ICLR1
2025 HiRA: Parameter-Efficient Hadamard High-Rank Adaptation for Large Language Models
abstract
We propose Hadamard High-Rank Adaptation (HiRA), a parameter-efficient fine-tuning (PEFT) method that enhances the adaptability of Large Language Models (LLMs). While Low-rank Adaptation (LoRA) is widely used to reduce resource demands, its low-rank updates may limit its expressiveness for new tasks. HiRA addresses this by using a Hadamard product to retain high-rank update parameters, improving the model capacity. Empirically, HiRA outperforms LoRA and its variants on several tasks, with extensive ablation studies validating its effectiveness. Our code is available at https://github.com/hqsiswiliam/hira.
Qiushi Huang, Tom Ko, Zhan Zhuang, Lilian Tang, Yu Zhang 0006
ICLR1
2025 Come Together, But Not Right Now: A Progressive Strategy to Boost Low-Rank Adaptation
abstract
Low-rank adaptation (LoRA) has emerged as a leading parameter-efficient fine-tuning technique for adapting large foundation models, yet it often locks adapters into suboptimal minima near their initialization. This hampers model generalization and limits downstream operators such as adapter merging and pruning. Here, we propose CoTo, a progressive training strategy that gradually increases adapters’ activation probability over the course of fine-tuning. By stochastically deactivating adapters, CoTo encourages more balanced optimization and broader exploration of the loss landscape. We provide a theoretical analysis showing that CoTo promotes layer-wise dropout stability and linear mode connectivity, and we adopt a cooperative-game approach to quantify each adapter’s marginal contribution. Extensive experiments demonstrate that CoTo consistently boosts single-task performance, enhances multi-task merging accuracy, improves pruning robustness, and reduces training overhead, all while remaining compatible with diverse LoRA variants. Code is available at https://github.com/zwebzone/coto.
Zhan Zhuang, Xiequn Wang, Yulong Zhang 0005, Qiushi Huang, Shuhao Chen, Xuehao Wang, Yanbin Wei, Yuhe Nie, Kede Ma, Yu Zhang 0006, Ying Wei 0001
ICML5
2025 Reproducibility Companion Paper: AdOCTeRA - Adaptive Optimization Constraints for Improved text-guided Retrieval of Apartments
abstract
This reproducibility Companion paper supports the approach presented in our paper titled ''AdOCTeRA: Adaptive Optimization Constraints for Improved Text-Guided Retrieval of Apartments.'' In that work, we addressed the problem of apartment retrieval using textual descriptions. Specifically, we proposed a novel adaptive approach that leverages the similarity between apartment descriptions in the dataset to enforce different levels of distance-a small distance between similar elements, a slightly larger distance between less similar elements, and an even larger distance between dissimilar elements.
Ali Abdari, Alex Falcon, Giuseppe Serra 0001, Qiushi Huang
ICMR4
2025 SMPV: Social Media Prediction for Videos
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng
ACM Multimedia3
2024 Retrieval-Augmented Text-to-Audio Generation
abstract
Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes while underperforming in the rare ones, thus degrading the overall generation performance. We refer to this problem as long-tailed text-to-audio generation. To address this issue, we propose a simple retrieval-augmented approach for TTA models. Specifically, given an input text prompt, we first leverage a Contrastive Language Audio Pretraining (CLAP) model to retrieve relevant text-audio pairs. The features of the retrieved audio-text data are then used as additional conditions to guide the learning of TTA models. We enhance AudioLDM with our proposed approach and denote the resulting augmented system as Re-AudioLDM. On the AudioCaps dataset, Re-AudioLDM achieves a state-of-the-art Frechet Audio Distance (FAD) of 1.37, outperforming the existing approaches by a large margin. Furthermore, we show that Re-AudioLDM can generate realistic audio for complex scenes, rare audio classes, and even unseen audio types, indicating its potential in TTA tasks.
Haohe Liu, Xubo Liu 0001, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001
ICASSP4
2024 Nemesis: Normalizing the Soft-prompt Vectors of Vision-Language Models
abstract
With the prevalence of large-scale pretrained vision-language models (VLMs), such as CLIP, soft-prompt tuning has become a popular method for adapting these models to various downstream tasks. However, few works delve into the inherent properties of learnable soft-prompt vectors, specifically the impact of their norms to the performance of VLMs. This motivates us to pose an unexplored research question: ``Do we need to normalize the soft prompts in VLMs?'' To fill this research gap, we first uncover a phenomenon, called the $\textbf{Low-Norm Effect}$ by performing extensive corruption experiments, suggesting that reducing the norms of certain learned prompts occasionally enhances the performance of VLMs, while increasing them often degrades it. To harness this effect, we propose a novel method named $\textbf{N}$ormalizing th$\textbf{e}$ soft-pro$\textbf{m}$pt v$\textbf{e}$ctors of vi$\textbf{si}$on-language model$\textbf{s}$ ($\textbf{Nemesis}$) to normalize soft-prompt vectors in VLMs. To the best of our knowledge, our work is the first to systematically investigate the role of norms of soft-prompt vector in VLMs, offering valuable insights for future research in soft-prompt tuning.
Xiequn Wang, Qiushi Huang, Yu Zhang 0006
ICLR3
2024 Reproducibility Companion Paper: Recommendation of Mix-and-Match Clothing by Modeling Indirect Personal Compatibility
abstract
ICMR '24: International Conference on Multimedia Retrieval, Phuket, Thailand, June 10-14, 2024
Shuiying Liao, Yujuan Ding, P. Y. Mok 0001, Qiushi Huang, Jialun Cao
ICMR4
2024 SMP Challenge Summary: Social Media Prediction Challenge
Bo Wu 0018, Peiye Liu, Qiushi Huang, Zhaoyang Zeng, Jia Wang 0020, Bei Liu 0001, Jiebo Luo 0001, Wen-Huang Cheng
ACM Multimedia3
2023 Personalized Dialogue Generation with Persona-Adaptive Attention
abstract
Persona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs.
Qiushi Huang, Yu Zhang 0006, Tom Ko, Xubo Liu 0001, Bo Wu 0018, Wenwu Wang 0001, Lilian Tang
AAAI1
2023 Learning Retrieval Augmentation for Personalized Dialogue Generation
abstract
Personalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications.However, persona profiles, a prevalent setting in current personalized dialogue datasets, typically composed of merely four to five sentences, may not offer comprehensive descriptions of the persona about the agent, posing a challenge to generate truly personalized dialogues.To handle this problem, we propose Learning Retrieval Augmentation for Personalized DialOgue Generation (LAPDOG), which studies the potential of leveraging external knowledge for persona dialogue generation.Specifically, the proposed LAPDOG model consists of a story retriever and a dialogue generator.The story retriever uses a given persona profile as queries to retrieve relevant information from the story document, which serves as a supplementary context to augment the persona profile.The dialogue generator utilizes both the dialogue history and the augmented persona profile to generate personalized responses.For optimization, we adopt a joint training framework that collaboratively learns the story retriever and dialogue generator, where the story retriever is optimized towards desired ultimate metrics (e.g., BLEU) to retrieve content for the dialogue generator to generate personalized responses.Experiments conducted on the CONVAI2 dataset with ROCStory as a supplementary data source show that the proposed LAPDOG method substantially outperforms the baselines, indicating the effectiveness of the proposed method.The LAPDOG model code is publicly available for further exploration.
Qiushi Huang, Xubo Liu 0001, Wenwu Wang 0001, Tom Ko, Yu Zhang 0006, Lilian Tang
EMNLP1
2023 Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention
abstract
Audio captioning aims to generate text descriptions of audio clips.In the real world, many objects produce similar sounds.How to accurately recognize ambiguous sounds is a major challenge for audio captioning.In this work, inspired by inherent human multimodal perception, we propose visuallyaware audio captioning, which makes use of visual information to help the description of ambiguous sounding objects.Specifically, we introduce an off-the-shelf visual encoder to extract video features and incorporate the visual features into an audio captioning system.Furthermore, to better exploit complementary audio-visual contexts, we propose an audio-visual attention mechanism that adaptively integrates audio and visual context and removes the redundant information in the latent space.Experimental results on AudioCaps, the largest audio captioning dataset, show that our proposed method achieves state-of-theart results on machine translation metrics.
Xubo Liu 0001, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang 0006, Lilian Tang, Mark D. Plumbley, Volkan Kilic, Wenwu Wang 0001
INTERSPEECH2
2023 SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge
abstract
Social Media Popularity Prediction (SMPP) is a crucial task that involves automatically predicting future popularity values of online posts, leveraging vast amounts of multimodal data available on social media platforms. Studying and investigating social media popularity becomes central to various online applications and requires novel methods of comprehensive analysis, multimodal comprehension, and accurate prediction.
Bo Wu 0018, Peiye Liu, Wen-Huang Cheng, Bei Liu 0001, Zhaoyang Zeng, Jia Wang 0020, Qiushi Huang, Jiebo Luo 0001
ACM Multimedia7
2022 Separate What You Describe: Language-Queried Audio Source Separation
abstract
In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language query of the target source (e.g., "a man tells a joke followed by people laughing"). A unique challenge in LASS is associated with the complexity of natural language description and its relation with the audio sources. To address this issue, we proposed LASS-Net, an end-to-end neural network that is learned to jointly process acoustic and linguistic information, and separate the target source that is consistent with the language query from an audio mixture. We evaluate the performance of our proposed system with a dataset created from the AudioCaps dataset. Experimental results show that LASS-Net achieves considerable improvements over baseline methods. Furthermore, we observe that LASS-Net achieves promising generalization results when using diverse human-annotated descriptions as queries, indicating its potential use in real-world scenarios. The separated audio samples and source code are available at https://liuxubo717.github.io/LASS-demopage.
Xubo Liu 0001, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001
INTERSPEECH6
2021 Token-Level Supervised Contrastive Learning for Punctuation Restoration
abstract
Punctuation is critical in understanding natural language text. Currently, most automatic speech recognition (ASR) systems do not generate punctuation, which affects the performance of downstream tasks, such as intent detection and slot filling. This gives rise to the need for punctuation restoration. Recent work in punctuation restoration heavily utilizes pre-trained language models without considering data imbalance when predicting punctuation classes. In this work, we address this problem by proposing a token-level supervised contrastive learning method that aims at maximizing the distance of representation of different punctuation marks in the embedding space. The result shows that training with token-level supervised contrastive learning obtains up to 3.2% absolute F1 improvement on the test set.
Qiushi Huang, Tom Ko, Lilian Tang, Xubo Liu 0001, Bo Wu 0018
Interspeech1
2021 Focus and retain: Complement the Broken Pose in Human Image Synthesis
abstract
Given a target pose, how to generate an image of a specific style with that target pose remains an ill-posed and thus complicated problem. Most recent works treat the human pose synthesis tasks as an image spatial transformation problem using flow warping techniques. However, we observe that, due to the inherent ill-posed nature of many complicated human poses, former methods fail to generate body parts. To tackle this problem, we propose a feature-level flow attention module and an Enhancer Network. The flow attention module produces a flow attention mask to guide the combination of the flow-warped features and the structural pose features. Then, we apply the Enhancer Network to re-fine the coarse image by injecting the pose information. We present our experimental evaluation both qualitatively and quantitatively on DeepFashion, Market-1501, and Youtube dance datasets. Quantitative results show that our method has 12.995 FID at DeepFashion, 25.459 FID at Market-1501, 14.516 FID at Youtube dance datasets, which outperforms some state-of-the-arts including Guide-Pixe2Pixe, Global-Flow-Local-Attn, and CocosNet.
Pu Ge, Qiushi Huang, Xue Jing, Yule Li, Yiyong Li, Zhun Sun
WACV2
2020 A Feature Generalization Framework for Social Media Popularity Prediction
abstract
Social media is an indispensable part in modern life and social media popularity prediction can be applied to many aspects of sociality. In this paper, we propose a novel combined framework for social media popularity prediction, which accomplishes feature generalization and temporal modeling based on multi-modal feature extraction. On the one hand, in order to address the generalization problem caused by massive missing data, we train two CatBoost models with different datasets and integrate their outputs with a linear combination. On the other hand, sliding window average is employed to mine potential short-term dependency for each user's post sequence. Extensive experiments show that our proposed framework has superiorities in both feature generalization and temporal modeling. Besides, our approach achieves the 1st place on the leader board of the SMP Challenge in 2020, which proves the effectiveness of our proposed framework.
Qiushi Huang, Zhendong Mao 0001, Yongdong Zhang 0001
ACM Multimedia4
2017 Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks
abstract
Prediction of popularity has profound impact for social media, since it offers opportunities to reveal individual preference and public attention from evolutionary social systems. Previous research, although achieves promising results, neglects one distinctive characteristic of social data, i.e., sequentiality. For example, the popularity of online content is generated over time with sequential post streams of social media. To investigate the sequential prediction of popularity, we propose a novel prediction framework called Deep Temporal Context Networks (DTCN) by incorporating both temporal context and temporal attention into account. Our DTCN contains three main components, from embedding, learning to predicting. With a joint embedding network, we obtain a unified deep representation of multi-modal user-post data in a common embedding space. Then, based on the embedded data sequence over time, temporal context learning attempts to recurrently learn two adaptive temporal contexts for sequential popularity. Finally, a novel temporal attention is designed to predict new popularity (the popularity of a new user-post pair) with temporal coherence across multiple time-scales. Experiments on our released image dataset with about 600K Flickr photos demonstrate that DTCN outperforms state-of-the-art deep prediction algorithms, with an average of 21.51% relative performance improvement in the popularity prediction (Spearman Ranking Correlation).
Bo Wu 0018, Wen-Huang Cheng, Yongdong Zhang 0001, Qiushi Huang, Jintao Li 0001, Tao Mei 0001
IJCAI4