Hongxiang Li 0004

dblp:62/442-4 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
23since 2021 · last 2025
0009-0000-7710-8835ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
abstract
Controllable human image animation aims to generate videos from reference images using driving videos. Due to the limited control signals provided by sparse guidance (e.g., skeleton pose), recent works have attempted to introduce additional dense conditions (e.g., depth map) to ensure motion alignment. However, such strict dense guidance impairs the quality of the generated video when the body shape of the reference character differs significantly from that of the driving video. In this paper, we present DisPose to mine more generalizable and effective control signals without additional dense input, which disentangles the sparse skeleton pose in human image animation into motion field guidance and keypoint correspondence. Specifically, we generate a dense motion field from a sparse motion field and the reference image, which provides region-level dense guidance while maintaining the generalization of the sparse pose control. We also extract diffusion features corresponding to pose keypoints from the reference image, and then these point features are transferred to the target pose to provide distinct identity information. To seamlessly integrate into existing models, we propose a plug-and-play hybrid ControlNet that improves the quality and consistency of generated videos while freezing the existing model parameters. Extensive qualitative and quantitative experiments demonstrate the superiority of DisPose compared to current methods. Project page: https://github.com/lihxxx/DisPose.
Hongxiang Li 0004, Yaowei Li 0001, Zhihong Zhu 0001, Xuxin Cheng, Long Chen 0016
ICLR1
2025 BlobCtrl: Taming Controllable Blob for Element-level Image Editing
abstract
As user expectations for image editing continue to rise, the demand for flexible, fine-grained manipulation of specific visual elements presents a challenge for current diffusion-based methods. In this work, we present BlobCtrl, a framework for element-level image editing based on a probabilistic blob-based representation. Treating blobs as visual primitives, BlobCtrl disentangles layout from appearance, affording fine-grained, controllable object-level elements manipulation. Our key contributions are twofold: 1) an in-context dual-branch diffusion model that separates foreground and background processing, incorporating blob representations to explicitly decouple layout and appearance; and 2) a self-supervised disentangle-then-reconstruct training paradigm with an identity-preserving loss function, along with tailored strategies to efficiently leverage blob-image pairs. To foster further research, we introduce BlobData for large-scale training, and BlobBench, a benchmark for systematic evaluation. Experimental results demonstrate that BlobCtrl achieves state-of-the-art performance in a variety of element-level editing tasks—such as object addition, removal, scaling, and replacement—while maintaining computational efficiency.
Yaowei Li 0001, Lingen Li, Zhaoyang Zhang 0004, Xiaoyu Li 0002, Guangzhi Wang, Hongxiang Li 0004, Xiaodong Cun, Ying Shan, Yuexian Zou
SIGGRAPH Asia6
2025 DAT: Dual-Branch Adapter-Tuning for Few-Shot Recognition
abstract
Parameter-Efficient Fine-Tuning methods based on vision-language models (such as CLIP) for few-shot learning have recently received considerable attention. However, previous works only fine-tune either the image or text branch, breaking the alignment of the original two branches, meanwhile fine-tuning both branches of the CLIP would inevitably introduce more trainable parameters and likely cause more severe over-fitting due to the limited training data. In this study, we propose a novel Dual-branch Adapter-Tuning framework (DAT), which collaboratively trains the visual adapter and textual adapter added to the two branches of the original CLIP with multiple consistency constraints. By effectively utilizing the semantically detailed class-specific prompts and outputs of the original CLIP to guide the fine-tuning of both branches, our method gains exceptional adaptation ability to the downstream few-shot learning tasks and alleviates the over-fitting issue, meanwhile maximally preserving the generalization ability of the original CLIP model. Our proposed framework has achieved superior performance on diverse datasets under various few-shot learning settings compared to the existing approaches. The source code is available athttps://github.com/SandyXi/DAT.
Junxi Chen, Guangxing Wu, Hongxiang Li 0004, Jiankang Chen, Wentao Zhang 0005, Wei-Shi Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Cross-Modal Conditioned Reconstruction for Language-Guided Medical Image Segmentation
abstract
Recent developments underscore the potential of textual information in enhancing learning models for a deeper understanding of medical visual semantics. However, language-guided medical image segmentation still faces a challenging issue. Previous works employ implicit architectures to embed textual information. This leads to segmentation results that are inconsistent with the semantics represented by the language, sometimes even diverging significantly. To this end, we propose a novel cross-modal conditioned Reconstruction for Language-guided Medical Image Segmentation (RecLMIS) to explicitly capture cross-modal interactions, which assumes that well-aligned medical visual features and medical notes can effectively reconstruct each other. We introduce conditioned interaction to adaptively predict patches and words of interest. Subsequently, they are utilized as conditioning factors for mutual reconstruction to align with regions described in the medical notes. Extensive experiments demonstrate the superiority of our RecLMIS, surpassing LViT by 3.74% mIoU on the MosMedData+ dataset and 1.89% mIoU on the QATA-CoV19 dataset. More importantly, we achieve a relative reduction of 20.2% in parameter count and a 55.5% decrease in computational load. The code will be available at https://github.com/ShawnHuang497/RecLMIS.
Xiaoshuang Huang, Hongxiang Li 0004, Meng Cao 0002, Long Chen 0016, Chenyu You, Dong An 0001
IEEE Trans. Medical Imaging2
2024 Towards Multi-Intent Spoken Language Understanding via Hierarchical Attention and Optimal Transport
abstract
Multi-Intent spoken language understanding (SLU) can handle complicated utterances expressing multiple intents, which has attracted increasing attention from researchers. Although existing models have achieved promising performance, most of them still suffer from two leading problems: (1) each intent has its specific scope and the semantic information outside the scope might potentially hinder accurate predictions, i.e. scope barrier; (2) only the guidance from intent to slot is modeled but the guidance from slot to intent is often neglected, i.e. unidirectional guidance. In this paper, we propose a novel Multi-Intent SLU framework termed HAOT, which utilizes hierarchical attention to divide the scopes of each intent and applies optimal transport to achieve the mutual guidance between slot and intent. Experiments demonstrate that our model achieves state-of-the-art performance on two public Multi-Intent SLU datasets, obtaining the 3.4 improvement on MixATIS dataset compared to the previous best models in overall accuracy.
Xuxin Cheng, Zhihong Zhu 0001, Hongxiang Li 0004, Yaowei Li 0001, Xianwei Zhuang, Yuexian Zou
AAAI3
2024 Exploiting Auxiliary Caption for Video Grounding
abstract
Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the sparsity dilemma in video annotations, which fails to provide the context information between potential events and query sentences in the dataset. In this paper, we contend that exploiting easily available captions which describe general actions, i.e., auxiliary captions defined in our paper, will significantly boost the performance. To this end, we propose an Auxiliary Caption Network (ACNet) for video grounding. Specifically, we first introduce dense video captioning to generate dense captions and then obtain auxiliary captions by Non-Auxiliary Caption Suppression (NACS). To capture the potential information in auxiliary captions, we propose Caption Guided Attention (CGA) project the semantic relations between auxiliary captions and query sentences into temporal space and fuse them into visual representations. Considering the gap between auxiliary captions and ground truth, we propose Asymmetric Cross-modal Contrastive Learning (ACCL) for constructing more negative pairs to maximize cross-modal mutual information. Extensive experiments on three public datasets (i.e., ActivityNet Captions, TACoS and ActivityNet-CG) demonstrate that our method significantly outperforms state-of-the-art methods.
Hongxiang Li 0004, Meng Cao 0002, Xuxin Cheng, Yaowei Li 0001, Zhihong Zhu 0001, Yuexian Zou
AAAI1
2024 Aligner²: Enhancing Joint Multiple Intent Detection and Slot Filling via Adjustive and Forced Cross-Task Alignment
abstract
Multi-intent spoken language understanding (SLU) has garnered growing attention due to its ability to handle multiple intent utterances, which closely mirrors practical scenarios. Unlike traditional SLU, each intent in multi-intent SLU corresponds to its designated scope for slots, which occurs in certain fragments within the utterance. As a result, establishing precise scope alignment to mitigate noise impact emerges as a key challenge in multi-intent SLU. More seriously, they lack alignment between the predictions of the two sub-tasks due to task-independent decoding, resulting in a limitation on the overall performance. To address these challenges, we propose a novel framework termed Aligner² for multi-intent SLU, which contains an Adjustive Cross-task Aligner (ACA) and a Forced Cross-task Aligner (FCA). ACA utilizes the information conveyed by joint label embeddings to accurately align the scope of intent and corresponding slots, before the interaction of the two subtasks. FCA introduces reinforcement learning, to enforce the alignment of the task-specific hidden states after the interaction, which is explicitly guided by the prediction. Extensive experiments on two public multi-intent SLU datasets demonstrate the superiority of our Aligner² over state-of-the-art methods. More encouragingly, the proposed method Aligner² can be easily integrated into existing multi-intent SLU frameworks, to further boost performance.
Zhihong Zhu 0001, Xuxin Cheng, Yaowei Li 0001, Hongxiang Li 0004, Yuexian Zou
AAAI4
2024 Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup
abstract
Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently.It has been verified that visual information brings greater performance gains when the textual information is limited.However, most previous works ignore to take advantage of the complete textual inputs and the limited textual inputs at the same time, which limits the overall performance.To solve this issue, we propose a mixup method termed Soul-Mix to enhance MMT by using visual information more effectively.We mix the predicted translations of complete textual input and the limited textual inputs.Experimental results on the Multi30K dataset of three translation directions show that our Soul-Mix significantly outperforms existing approaches and achieves new state-of-the-art performance with fewer parameters than some previous models.Besides, the strength of Soul-Mix is more obvious on more challenging MSCOCO dataset which includes more out-of-domain instances with lots of ambiguous verbs.
Xuxin Cheng, Ziyu Yao 0001, Yifei Xin, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
ACL (1)5
2024 KDProR: A Knowledge-Decoupling Probabilistic Framework for Video-Text Retrieval
Xianwei Zhuang, Hongxiang Li 0004, Xuxin Cheng, Zhihong Zhu 0001, Yuxin Xie 0004, Yuexian Zou
ECCV (34)2
2024 KC-Prompt: End-To-End Knowledge-Complementary Prompting for Rehearsal-Free Continual Learning
abstract
Continuous learning requires adapting quickly to incoming tasks while avoiding catastrophic forgetting. Typical solutions resort to a rehearsal buffer to replay old data, which is intractable to apply in real-world scenarios with limited memory and inaccessible privacy. Recently, with the emergence of large-scale pre-trained models, prompting methods have rapidly become a popular rehearsal-free alternative to rehearsal-based methods. The core of prmopting is to encode knowledge leveraging a set of parameters, however, knowledge decoupling and complementarity still remain some challenges. To tackle these challenges, this paper presents a KnowledgeComplementary Prompting approach, KC-Prompt, which end-to-end integrates and releases the task-invariant and task-specific knowledge for the ViT backbone. KC-Prompt designs knowledge maintenance and knowledge sharing mechanisms to form complementary prompt generators. In addition, we employ a components weighting method to instantiate prompt generators, making the training process fully differentiable. Sufficient experiments on CIFAR-100 and Split ImageNet-R benchmarks demonstrate the superiority of KC-Prompt in the challenging and realistic class-incremental learning setting.
Yaowei Li 0001, Xuxin Cheng, Zhihong Zhu 0001, Hongxiang Li 0004, Bang Yang, Zhiqi Huang 0001
ICASSP5
2024 Textual Inversion and Self-supervised Refinement for Radiology Report Generation
Yuanjiang Luo, Hongxiang Li 0004, Meng Cao 0002, Xiaoshuang Huang, Zhihong Zhu 0001, Peixi Liao
MICCAI (5)2
2024 Multivariate Cooperative Game for Image-Report Pairs: Hierarchical Semantic Alignment for Medical Report Generation
Zhihong Zhu 0001, Xuxin Cheng, Yunyan Zhang, Zhaorun Chen, Qingqing Long, Hongxiang Li 0004, Zhiqi Huang 0001, Xian Wu 0001, Yefeng Zheng 0001
MICCAI (3)6
2024 Towards Multimodal-augmented Pre-trained Language Models via Self-balanced Expectation-Maximization Iteration
abstract
Pre-trained language models (PLMs) that rely solely on textual corpus may present limitations in multimodal semantics comprehension. Existing studies attempt to alleviate this issue by incorporating additional modal information through image retrieval or generation. However, these methods: (1) inevitably encounter modality gaps and noise; (2) treat all modalities indiscriminately; and (3) ignore visual or acoustic semantics of key entities. To tackle these challenges, we propose a novel principled iterative framework for multimodal-augmented PLMs termed MASE, which achieves efficient and balanced injection of multimodal semantics under the proposed Expectation Maximization (EM) based iterative algorithm. Initially, MASE utilizes multimodal proxies instead of explicit data to enhance PLMs, which avoids noise and modality gaps. In E-step, MASE adopts a novel information-driven self-balanced strategy to estimate allocation weights. Furthermore, MASE employs heterogeneous graph attention to capture entity-level fine-grained semantics on the proposed multimodal-semantic scene graph. In M-step, MASE injects global multimodal knowledge into PLMs through a cross-modal contrastive loss. Experimental results show that MASE consistently outperforms competitive baselines on multiple tasks across various architectures. More impressively, MASE is compatible with existing efficient parameter fine-tuning methods, such as prompt learning.
Xianwei Zhuang, Xuxin Cheng, Zhihong Zhu 0001, Zhanpeng Chen, Hongxiang Li 0004, Yuexian Zou
ACM Multimedia5
2024 Dance with Labels: Dual-Heterogeneous Label Graph Interaction for Multi-intent Spoken Language Understanding
abstract
Multi-intent spoken language understanding (SLU) has garnered increasing attention since it can handle complex utterances expressing multiple intents in real-world scenarios. However, existing joint models are disturbed by label statistical frequencies, or adopt homogeneous graphs to capture interactions between the different types (e.g., intent and slot) of label nodes, thereby limiting the performance. To overcome these limitations, we propose Dual Heterogeneous Graph Label Interaction for multi-intent SLU, named DHLG. Concretely, we propose a global static heterogeneous label graph interaction layer to model both intra- and inter-label statistical dependencies across the entire training corpus. Based on this, we introduce a local dynamic heterogeneous label graph layer to further facilitate adaptive interactions between intents and slots for each utterance. Extensive experiments and analyses on two widely-used benchmark datasets demonstrate the superiority of our proposed DHLG over state-of-the-art methods.
Zhihong Zhu 0001, Xuxin Cheng, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
WSDM3
2023 Towards Spoken Language Understanding via Multi-level Multi-grained Contrastive Learning
abstract
Spoken language understanding (SLU) is a core task in task-oriented dialogue systems, which aims at understanding user's current goal through constructing semantic frames. SLU usually consists of two subtasks, including intent detection and slot filling. Although there are some SLU frameworks joint modeling the two subtasks and achieve the high performance, most of them still overlook the inherent relationships between intents and slots, and fail to achieve mutual guidance between the two subtasks. To solve the problem, we propose a multi-level multi-grained SLU framework MMCL to apply contrastive learning at three levels, including utterance level, slot level, and word level to enable intent and slot to mutually guide each other. For the utterance level, our framework implements coarse granularity contrastive learning and fine granularity contrastive learning simultaneously. Besides, we also apply the self-distillation method to improve the robustness of the model. Experimental results and further analysis demonstrate that our proposed model achieves new state-of-the-art results on two public multi-intent SLU datasets, obtaining a 2.6 overall accuracy improvement on MixATIS dataset compared to previous best models.
Xuxin Cheng, Wanshi Xu, Zhihong Zhu 0001, Hongxiang Li 0004, Yuexian Zou
CIKM4
2023 DAS-CL: Towards Multimodal Machine Translation via Dual-Level Asymmetric Contrastive Learning
abstract
Multimodal machine translation (MMT) aims to exploit visual information to improve neural machine translation (NMT). It has been demonstrated that image captioning and object detection can further improve MMT. In this paper, to leverage image captioning and object detection more effectively, we propose a Dual-level ASymmetric Contrastive Learning (DAS-CL) framework. Specifically, we leverage image captioning and object detection to generate more pairs of visual inputs and textual inputs. At the utterance level, we introduce an image captioning model to generate more coarse-grained pairs. At the word level, we introduce an object detection model to generate more fine-grained pairs. To mitigate the negative impact of noise in generated pairs, we apply asymmetric contrastive learning at these two levels. Experiments on the Multi30K dataset of three translation directions demonstrate that DAS-CL significantly outperforms existing MMT frameworks and achieves new state-of-the-art performance. More encouragingly, further analysis displays that DAS-CL is more robust to irrelevant visual information.
Xuxin Cheng, Zhihong Zhu 0001, Yaowei Li 0001, Hongxiang Li 0004, Yuexian Zou
CIKM4
2023 SSVMR: Saliency-Based Self-Training for Video-Music Retrieval
abstract
With the rise of short videos, the demand for selecting appropriate background music (BGM) for a video has increased significantly, video-music retrieval (VMR) task gradually draws much attention by research community. As other cross-modal learning tasks, existing VMR approaches usually attempt to measure the similarity between the video and music in the feature space. However, they (1) neglect the inevitable label noise; (2) neglect to enhance the ability to capture critical video clips. In this paper, we propose a novel saliency-based self-training framework, which is termed SSVMR. Specifically, we first explore to fully make use of the information containing in the training dataset by applying a semi-supervised method to suppress the adverse impact of label noise problem, where a self-training approach is adopted. In addition, we propose to capture the saliency of the video by mixing two videos at span level and preserving the locality of the two original videos. Inspired by back translation in NLP, we also conduct back retrieval to obtain more training data. Experimental results on MVD dataset show that our SSVMR achieves the state-of-the-art performance by a large margin, obtaining a relative improvement of 34.8% over the previous best model in terms of R@1.
Xuxin Cheng, Zhihong Zhu 0001, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
ICASSP3
2023 G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory
abstract
The recent video grounding works attempt to introduce vanilla contrastive learning into video grounding. However, we claim that this naive solution is suboptimal. Contrastive learning requires two key properties: (1) alignment of features of similar samples, and (2) uniformity of the induced distribution of the normalized features on the hypersphere. Due to two annoying issues in video grounding: (1) the coexistence of some visual entities in both ground truth and other moments, i.e. semantic overlapping; (2) only a few moments in the video are annotated, i.e. sparse annotation dilemma, vanilla contrastive learning is unable to model the correlations between temporally distant moments and learned inconsistent video representations. Both characteristics lead to vanilla contrastive learning being unsuitable for video grounding. In this paper, we introduce Geodesic and Game Localization (G2L), a semantically aligned and uniform video grounding framework via geodesic and game theory. We quantify the correlations among moments leveraging the geodesic distance that guides the model to learn the correct cross-modal representations. Furthermore, from the novel perspective of game theory, we propose semantic Shapley interaction based on geodesic distance sampling to learn fine-grained semantic alignment in similar moments. Experiments on three benchmarks demonstrate the effectiveness of our method.
Hongxiang Li 0004, Meng Cao 0002, Xuxin Cheng, Yaowei Li 0001, Zhihong Zhu 0001, Yuexian Zou
ICCV1
2023 Unify, Align and Refine: Multi-Level Semantic Alignment for Radiology Report Generation
abstract
Automatic radiology report generation has attracted enormous research interest due to its practical value in reducing the workload of radiologists. However, simultaneously establishing global correspondences between the image (e.g., Chest X-ray) and its related report and local alignments between image patches and keywords remains challenging. To this end, we propose an Unify, Align and then Refine (UAR) approach to learn multi-level cross-modal alignments and introduce three novel modules: Latent Space Unifier (LSU), Cross-modal Representation Aligner (CRA) and Text-to-Image Refiner (TIR). Specifically, LSU unifies multimodal data into discrete tokens, making it flexible to learn common knowledge among modalities with a shared network. The modality-agnostic CRA learns discriminative features via a set of orthonormal basis and a dual-gate mechanism first and then globally aligns visual and textual representations under a triplet contrastive loss. TIR boosts token-level local alignment via calibrating text-to-image attention with a learnable mask. Additionally, we design a two-stage training procedure to make UAR gradually grasp cross-modal alignments at different levels, which imitates radiologists’ workflow: writing sentence by sentence first and then checking word by word. Extensive experiments and analyses on IU-Xray and MIMIC-CXR benchmark datasets demonstrate the superiority of our UAR against varied state-of-the-art methods.
Yaowei Li 0001, Bang Yang, Xuxin Cheng, Zhihong Zhu 0001, Hongxiang Li 0004, Yuexian Zou
ICCV5
2023 C²A-SLU: Cross and Contrastive Attention for Improving ASR Robustness in Spoken Language Understanding
Xuxin Cheng, Ziyu Yao 0001, Zhihong Zhu 0001, Yaowei Li 0001, Hongxiang Li 0004, Yuexian Zou
INTERSPEECH5
2023 FC-MTLF: A Fine- and Coarse-grained Multi-Task Learning Framework for Cross-Lingual Spoken Language Understanding
Xuxin Cheng, Wanshi Xu, Ziyu Yao 0001, Zhihong Zhu 0001, Yaowei Li 0001, Hongxiang Li 0004, Yuexian Zou
INTERSPEECH6
2023 GhostT5: Generate More Features with Cheap Operations to Improve Textless Spoken Question Answering
Xuxin Cheng, Zhihong Zhu 0001, Ziyu Yao 0001, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
INTERSPEECH4
2023 Mix before Align: Towards Zero-shot Cross-lingual Sentiment Analysis via Soft-Mix and Multi-View Learning
Zhihong Zhu 0001, Xuxin Cheng, Zhiqi Huang 0001, Hongxiang Li 0004, Yuexian Zou
INTERSPEECH5