Qihan Huang

dblp:270/3041 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Multisource Graphs and Dual KAN-Transformers for Next POI Recommendation
Jing Zhang 0040, Zhenhan Huang, Tian Wang 0001, Qihan Huang, Li Xu 0002, Xiucai Ye
IEEE Internet Things J.4
2025 Resolving Multi-Condition Confusion for Finetuning-Free Personalized Image Generation
abstract
Personalized text-to-image generation methods can generate customized images based on the reference images, which have garnered wide research interest. Recent methods propose a finetuning-free approach with a decoupled cross-attention mechanism to generate personalized images requiring no test-time finetuning. However, when multiple reference images are provided, the current decoupled cross-attention mechanism encounters the object confusion problem and fails to map each reference image to its corresponding object, thereby seriously limiting its scope of application. To address the object confusion problem, in this work we investigate the relevance of different positions of the latent image features to the target object in diffusion model, and accordingly propose a weighted-merge method to merge multiple reference image features into the corresponding objects. Next, we integrate this weighted-merge method into existing pre-trained models and continue to train the model on a multi-object dataset constructed from the open-sourced SA-1B dataset. To mitigate object confusion and reduce training costs, we propose an object quality score to estimate the image quality for the selection of high-quality training samples. Furthermore, our weighted-merge training framework can be employed on single-object generation when a single object has multiple reference images. The experiments verify that our method achieves superior performance to the state-of-the-arts on the Concept101 dataset and DreamBooth dataset of multi-object personalized image generation, and remarkably improves the performance on single-object personalized image generation.
Qihan Huang, Siming Fu, Hao Jiang 0014, Yipeng Yu, Jie Song 0011
AAAI1
2025 PatchDPO: Patch-level DPO for Finetuning-free Personalized Image Generation
abstract
Finetuning-free personalized image generation can synthesize customized images without test-time finetuning, attracting wide research interest owing to its high efficiency. Current finetuning-free methods simply adopt a single training stage with a simple image reconstruction task, and they typically generate low-quality images inconsistent with the reference images during test-time. To mitigate this problem, inspired by the recent DPO (i.e., direct preference optimization) technique, this work proposes an additional training stage to improve the pre-trained personalized generation models. However, traditional DPO only determines the overall superiority or inferiority of two samples, which is not suitable for personalized image generation because the generated images are commonly inconsistent with the reference images only in some local image patches. To tackle this problem, this work proposes PatchDPO that estimates the quality of image patches within each generated image and accordingly trains the model. To this end, PatchDPO first leverages the pre-trained vision model with a proposed self-supervised training method to estimate the patch quality. Next, PatchDPO adopts a weighted training approach to train the model with the estimated patch quality, which rewards the image patches with high quality while penalizing the image patches with low quality. Experiment results demonstrate that PatchDPO significantly improves the performance of multiple pre-trained personalized generation models, and achieves state-of-the-art performance on both single-object and multi-object personalized image generation. Our code is available at https://github.com/hqhQAQ/PatchDPO.
Qihan Huang, Long Chan, Wanggui He, Hao Jiang 0062, Mingli Song, Jie Song 0011
CVPR1
2025 TFCustom: Customized Image Generation with Time-Aware Frequency Feature Guidance
abstract
Subject-driven image personalization has seen notable advancements, especially with the ReferenceNet paradigm, which excels in integrating reference image features for creative and commercial applications. However, current ReferenceNet implementations mainly function as latent-level feature extractors, limiting their potential. This restricts the delivery of suitable features to the denoising backbone across timesteps, resulting in suboptimal image consistency. In this paper, we revisit reference feature extraction and propose TFCustom, a framework that focuses on reference image features at different temporal and frequency levels. We introduce synchronized ReferenceNet to extract reference features while optimizing noise injection and denoising. We also propose a time-aware frequency refinement module that uses high- and low-frequency filters with time embeddings to adaptively select reference feature injection. Additionally, we introduce a reward-based loss to improve the similarity between reference objects and generated images. Experimental results show that TFCustom outperforms existing methods in single-object and multi-object reference generation, with significant improvements in textual details.
Mushui Liu, Dong She, Jingxuan Pang, Qihan Huang, Jiacheng Ying, Wanggui He, Yuanlei Hou, Siming Fu
CVPR4
2025 Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
abstract
MLLM reasoning has drawn widespread research for its excellent problem-solving capability. Current reasoning methods fall into two types: PRM, which supervises the intermediate reasoning steps, and ORM, which supervises the final results. Recently, DeepSeek-R1 has challenged the traditional view that PRM outperforms ORM, which demonstrates strong generalization performance using an ORM method (i.e., GRPO). However, current MLLM's GRPO algorithms still struggle to handle challenging and complex multimodal reasoning tasks (e.g., mathematical reasoning). In this work, we reveal two problems that impede the performance of GRPO on the MLLM: Low data utilization and Text-bias. Low data utilization refers to that GRPO cannot acquire positive rewards to update the MLLM on difficult samples, and text-bias is a phenomenon that the MLLM bypasses image condition and solely relies on text condition for generation after GRPO training. To tackle these problems, this work proposes Hint-GRPO that improves data utilization by adaptively providing hints for samples of varying difficulty, and text-bias calibration that mitigates text-bias by calibrating the token prediction logits with image condition in test-time. Experiment results on three base MLLMs across eleven datasets demonstrate that our proposed methods advance the reasoning capability of original MLLM by a large margin, exhibiting superior performance to existing MLLM reasoning methods. Our code is available at https://github.com/hqhQAQ/Hint-GRPO.
Qihan Huang, Weilong Dai, Wanggui He, Hao Jiang 0062, Mingli Song, Jingyuan Chen 0003, Chang Yao 0001, Jie Song 0011
ICCV1
2025 MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
abstract
Recent advancements in text-to-image generation models have dramatically enhanced the generation of photorealistic images from textual prompts, leading to an increased interest in personalized text-to-image applications, particularly in multi-subject scenarios. However, these advances are hindered by two main challenges: firstly, the need to accurately maintain the details of each referenced subject in accordance with the textual descriptions; and secondly, the difficulty in achieving a cohesive representation of multiple subjects in a single image without introducing inconsistencies. To address these concerns, our research introduces the MS-Diffusion framework for layout-guided zero-shot image personalization with multi-subjects. This innovative approach integrates grounding tokens with the feature resampler to maintain detail fidelity among subjects. With the layout guidance, MS-Diffusion further improves the cross-attention to adapt to the multi-subject inputs, ensuring that each subject condition acts on specific areas. The proposed multi-subject cross-attention orchestrates harmonious inter-subject compositions while preserving the control of texts. Comprehensive quantitative and qualitative experiments affirm that this method surpasses existing models in both image and text fidelity, promoting the development of personalized text-to-image generation.
Xierui Wang, Siming Fu, Qihan Huang, Wanggui He, Hao Jiang 0062
ICLR3
2024 On the Concept Trustworthiness in Concept Bottleneck Models
abstract
Concept Bottleneck Models (CBMs), which break down the reasoning process into the input-to-concept mapping and the concept-to-label prediction, have garnered significant attention due to their remarkable interpretability achieved by the interpretable concept bottleneck. However, despite the transparency of the concept-to-label prediction, the mapping from the input to the intermediate concept remains a black box, giving rise to concerns about the trustworthiness of the learned concepts (i.e., these concepts may be predicted based on spurious cues). The issue of concept untrustworthiness greatly hampers the interpretability of CBMs, thereby hindering their further advancement. To conduct a comprehensive analysis on this issue, in this study we establish a benchmark to assess the trustworthiness of concepts in CBMs. A pioneering metric, referred to as concept trustworthiness score, is proposed to gauge whether the concepts are derived from relevant regions. Additionally, an enhanced CBM is introduced, enabling concept predictions to be made specifically from distinct parts of the feature map, thereby facilitating the exploration of their related regions. Besides, we introduce three modules, namely the cross-layer alignment (CLA) module, the cross-image alignment (CIA) module, and the prediction alignment (PA) module, to further enhance the concept trustworthiness within the elaborated CBM. The experiments on five datasets across ten architectures demonstrate that without using any concept localization annotations during training, our model improves the concept trustworthiness by a large margin, meanwhile achieving superior accuracy to the state-of-the-arts. Our code is available at https://github.com/hqhQAQ/ProtoCBM.
Qihan Huang, Jie Song 0011, Haofei Zhang, Mingli Song
AAAI1
2024 ProtoPFormer: Concentrating on Prototypical Parts in Vision Transformers for Interpretable Image Recognition
Mengqi Xue, Qihan Huang, Haofei Zhang, Jie Song 0011, Mingli Song, Canghong Jin
IJCAI2
2024 LG-CAV: Train Any Concept Activation Vector with Language Guidance
abstract
Concept activation vector (CAV) has attracted broad research interest in explainable AI, by elegantly attributing model predictions to specific concepts. However, the training of CAV often necessitates a large number of high-quality images, which are expensive to curate and thus limited to a predefined set of concepts. To address this issue, we propose Language-Guided CAV (LG-CAV) to harness the abundant concept knowledge within the certain pre-trained vision-language models (e.g., CLIP). This method allows training any CAV without labeled data, by utilizing the corresponding concept descriptions as guidance. To bridge the gap between vision-language model and the target model, we calculate the activation values of concept descriptions on a common pool of images (probe images) with vision-language model and utilize them as language guidance to train the LG-CAV. Furthermore, after training high-quality LG-CAVs related to all the predicted classes in the target model, we propose the activation sample reweighting (ASR), serving as a model correction technique, to improve the performance of the target model in return. Experiments on four datasets across nine architectures demonstrate that LG-CAV achieves significantly superior quality to previous CAV methods given any concept, and our model correction method achieves state-of-the-art performance compared to existing concept-based methods. Our code is available at https://github.com/hqhQAQ/LG-CAV.
Qihan Huang, Jie Song 0011, Mengqi Xue, Haofei Zhang, Bingde Hu, Huiqiong Wang, Hao Jiang 0014, Xingen Wang, Mingli Song
NeurIPS1
2023 Evaluation and Improvement of Interpretability for Self-Explainable Part-Prototype Networks
abstract
Part-prototype networks (e.g., ProtoPNet, ProtoTree, and ProtoPool) have attracted broad research interest for their intrinsic interpretability and comparable accuracy to non-interpretable counterparts. However, recent works find that the interpretability from prototypes is fragile, due to the semantic gap between the similarities in the feature space and that in the input space. In this work, we strive to address this challenge by making the first attempt to quantitatively and objectively evaluate the interpretability of the part-prototype networks. Specifically, we propose two evaluation metrics, termed as "consistency score" and "stability score", to evaluate the explanation consistency across images and the explanation robustness against perturbations, respectively, both of which are essential for explanations taken into practice. Furthermore, we propose an elaborated part-prototype network with a shallow-deep feature alignment (SDFA) module and a score aggregation (SA) module to improve the interpretability of prototypes. We conduct systematical evaluation experiments and provide substantial discussions to uncover the interpretability of existing part-prototype networks. Experiments on three benchmarks across nine architectures demonstrate that our model achieves significantly superior performance to the state of the art, in both the accuracy and interpretability. Our code is available at https://github.com/hqhQAQ/EvalProtoPNet.
Qihan Huang, Mengqi Xue, Wenqi Huang 0002, Haofei Zhang, Jie Song 0011, Yongcheng Jing, Mingli Song
ICCV1
2023 Generative adversarial learning for optimizing ontology alignment
abstract
Abstract The edge computing in knowledge‐defined network (KDN) is a kind of distributed computing architecture, and the format of edge resources stored in different edge computing nodes are different, which yields the data heterogeneity problem and hampers the interaction between edge nodes. Ontology is considered as the solution of data heterogeneity on Semantic Web, and matching ontologies is a high‐efficiency method of addressing the data heterogeneity problem. Ontology meta‐matching investigates how to determine the optimal weights to aggregate multiple similarity measures to achieve high‐quality ontology alignment, which is a challenge about nonlinear mathematical problem in ontology matching domain. To face this challenge, unsupervised learning method such as generative adversarial network (GAN) becomes an effective methodology. GAN consists of two models of different targets that are opposed to each other in training to produce the final best result. To improve the GAN's efficiency, this work further proposes a GAN with simulated annealing algorithm (SA‐GAN), where the stagnation counter is introduced to accelerate GAN's the convergence speed. The experiment uses the famous benchmark in the ontology domain, and the comparisons with the advanced ontology matching systems shows that SA‐GAN is able to find high‐quality alignments to help build bridges between edge nodes on edge computing.
Xingsi Xue, Qihan Huang
Expert Syst. J. Knowl. Eng.2
2023 DP-TrajGAN: A privacy-aware trajectory generation model with differential privacy
Jing Zhang 0040, Qihan Huang, Yirui Huang, Pei-Wei Tsai
Future Gener. Comput. Syst.2
2023 Hasse sensitivity level: A sensitivity-aware trajectory privacy-enhanced framework with Reinforcement Learning
Jing Zhang 0040, Yi-rui Huang, Qihan Huang, Yan-zi Li, Xiucai Ye
Future Gener. Comput. Syst.3
2023 Dimension-aware under spatiotemporal constraints: an efficient privacy-preserving framework with peak density clustering
Jing Zhang 0040, Qihan Huang, Jian-Yu Hu, Xiucai Ye
J. Supercomput.2
2022 Learn decision trees with deep visual primitives
Mengqi Xue, Haofei Zhang, Qihan Huang, Jie Song 0011, Mingli Song
J. Vis. Commun. Image Represent.3
2020 Effects of holding postures on user-defined touch gestures for tablet interaction
Huawei Tu, Qihan Huang, Yanchao Zhao, Boyu Gao 0003
Int. J. Hum. Comput. Stud.2