EDBT 2026 Demo / reviewers in the wild / expert
Guoxing Yang
dblp:271/9521
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image RetrievalabstractComposed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion strategies for addressing the task. However, these methods typically fail to simultaneously meet two key requirements of CIR: comprehensively extracting visual information and faithfully following the user intent. In this work, we propose CIR-LVLM, a novel framework that leverages the large vision-language model (LVLM) as the powerful user intent-aware encoder to better meet these requirements. Our motivation is to explore the advanced reasoning and instruction-following capabilities of LVLM for accurately understanding and responding the user intent. Furthermore, we design a novel hybrid intent instruction module to provide explicit intent guidance at two levels: (1) The task prompt clarifies the task requirement and assists the model in discerning user intent at the task level. (2) The instance-specific soft prompt, which is adaptively selected from the learnable prompt pool, enables the model to better comprehend the user intent at the instance level compared to a universal prompt for all instances. CIR-LVLM achieves state-of-the-art performance across three prominent benchmarks with acceptable inference efficiency. We believe this study provides fundamental insights into CIR-related fields. Zelong Sun, Dong Jing, Guoxing Yang, Nanyi Fei, Zhiwu Lu 0001 |
AAAI | 3 |
| 2024 | MoVL: Exploring Fusion Strategies for the Domain-Adaptive Application of Pretrained Models in Medical Imaging TasksabstractMedical images are often more difficult to acquire than natural images due to the specialized equipment and technology required, leading to fewer available medical image datasets. This limitation poses challenges in training robust pretrained medical vision models. How to best leverage natural pretrained vision models and adapt them to the medical domain remains an open question. For image classification, linear probing (LP) is a commonly used technique. However, LP primarily focuses on the output after feature extraction, without addressing the inherent differences between medical images and natural image-based pretrained models. To bridge this gap, we introduce visual prompting (VP) and investigate strategies for integrating LP and VP. We propose a joint learning framework with a composite loss function that includes a categorization loss and a discrepancy loss, capturing the variance between prompted and plain images, naming this joint training strategy MoVL (Mixture of Visual Prompting and Linear Probe). We experiment on four medical image classification datasets, with two mainstream architectures, ResNet and CLIP. Results shows that without changing the parameters and architecture of backbone model and with less parameters, there is potential for MoVL to achieve full finetune (FF) accuracy (on four medical datasets, average 90.91% for MoVL and 91.13% for FF). On out of distribution medical dataset, our method (90.33%) can outperform FF (85.15%) with absolute 5.18 % lead. Haijiang Tian, Jingkun Yue, Xiaohong Liu 0007, Guoxing Yang |
BIBM | 4 |
| 2024 | DIFFSC: Semantic Communication Framework With Enhanced Denoising Through Diffusion Probabilistic ModelsabstractIn communication systems, the challenge of ensuring accurate data transmission across noisy channels remains paramount. While semantic communication shows potential in improving image transmission and reconstruction, existing methods still suffer from perceptual quality degradation in high-noise environments. To address these issues, we introduce DiffSC, a novel semantic communication framework that integrates the Diffusion Probabilistic Model (DPM). Within DPM, Gaussian noise modeling is leveraged to facilitate enhanced image generation. DiffSC is trained to recover semantic information compromised during transmission, amplifying its capabilities in both image reconstruction and denoising. Additionally, we propose a Multi-Dimensional Feature Extraction Module (MFM) that employs multiple convolution kernels along with channel and spatial attention mechanisms to enrich encoding and decoding. Experimental results demonstrate that DiffSC significantly outperforms existing systems, improving SSIM by more than 19% and PSNR by more than 11% compared to Deep JSCC. Further ablation studies demonstrate the effectiveness of our proposed denoiser and feature extraction modules. Guoxing Yang, Weizhi Li, Aini Li |
ICASSP | 3 |
| 2024 | Image Retrieval with Composed Query by Multi-Scale Multi-Modal FusionabstractImage retrieval with composed query (IR-CQ) is a challenging task since it aims to retrieve the target image according to a hybrid-modality query which consists of a reference image and a text modifier. Previous approaches mainly focus on designing various multi-modal fusion modules to fuse the hybrid-modality query, but these fusion modules are often suboptimal without considering sufficient fusion between the two modalities. In this paper, we propose a general fusion block by taking three fusion strategies: weighted summing, concatenating, and bilinear pooling. Importantly, this general fusion block can be deployed to fuse not only the hybrid-modality query but also the multi-scale features of the reference image. Specifically, we first fuse the multi-scale features of the reference image with the Multi-Scale Fusion (MSF) block and then fuse the features of the reference image and text modifier with the Multi-Modal Fusion (MMF) block, where both MSF and MMF are instantiations of our general fusion block. Extensive experiments on three benchmark datasets show that our proposed model significantly outperforms existing approaches. Zelong Sun, Guoxing Yang, Zhiwu Lu 0001, Hao Jiang 0022, Guojie Zhu, Zhao Cao |
ICASSP | 2 |
| 2024 | Progressive Image Synthesis from Semantics to Details with Denoising Diffusion GANabstractAlthough denoising diffusion probabilistic models (DDPMs) have shown remarkable progress in image generation, they typically face two main challenges: the time-expensive sampling process and the semantically meaningless latent space, which are often addressed separately in previous works. In particular, the latest representative work Denoising Diffusion GAN reduces the sampling steps to as few as two but ignores the semantics of the latent space. To address the two challenges simultaneously, we propose a two-stage framework to make the latent space of Denoising Diffusion GAN more semantically meaningful while enjoying its efficiency. Extensive results on three benchmark datasets demonstrate that our proposed diffusion model achieves competitive results with only two sampling steps in unconditional image generation. More importantly, the latent space of our diffusion model trained for unconditional image generation is shown to be semantically meaningful, which can be exploited on various downstream tasks (e.g., attribute editing) without further training. Guoxing Yang, Haoyu Lu, Chongxuan Li, Guang Zhou, Zhiwu Lu 0001 |
ICASSP | 1 |
| 2024 | UniAdapter: Unified Parameter-Efficient Transfer Learning for Cross-modal ModelingabstractLarge-scale vision-language pre-trained models have shown promising transferability to various downstream tasks. As the size of these foundation models and the number of downstream tasks grow, the standard full fine-tuning paradigm becomes unsustainable due to heavy computational and storage costs. This paper proposes UniAdapter, which unifies unimodal and multimodal adapters for parameter-efficient cross-modal adaptation on pre-trained vision-language models. Specifically, adapters are distributed to different modalities and their interactions, with the total number of tunable parameters reduced by partial weight sharing. The unified and knowledge-sharing design enables powerful cross-modal representations that can benefit various downstream tasks, requiring only 1.0%-2.0% tunable parameters of the pre-trained model. Extensive experiments on 7 cross-modal downstream benchmarks (including video-text retrieval, image-text retrieval, VideoQA, VQA and Caption) show that in most cases, UniAdapter not only outperforms the state-of-the-arts, but even beats the full fine-tuning strategy. Particularly, on the MSRVTT retrieval task, UniAdapter achieves 49.7% recall@1 with 2.2% model parameters, outperforming the latest competitors by 2.0%. The code and models are available at https://github.com/RERV/UniAdapter. Haoyu Lu, Yuqi Huo, Guoxing Yang, Zhiwu Lu 0001, Masayoshi Tomizuka, Mingyu Ding |
ICLR | 3 |
| 2024 | VDT: General-purpose Video Diffusion Transformers via Mask ModelingabstractThis work introduces Video Diffusion Transformer (VDT), which pioneers the use of transformers in diffusion-based video generation.
It features transformer blocks with modularized temporal and spatial attention modules to leverage the rich spatial-temporal representation inherited in transformers. Additionally, we propose a unified spatial-temporal mask modeling mechanism, seamlessly integrated with the model, to cater to diverse video generation scenarios.
VDT offers several appealing benefits. (1) It excels at capturing temporal dependencies to produce temporally consistent video frames and even simulate the physics and dynamics of 3D objects over time. (2) It facilitates flexible conditioning information, e.g., simple concatenation in the token space, effectively unifying different token lengths and modalities. (3) Pairing with our proposed spatial-temporal mask modeling mechanism, it becomes a general-purpose video diffuser for harnessing a range of tasks, including unconditional generation, video prediction, interpolation, animation, and completion, etc. Extensive experiments on these tasks spanning various scenarios, including autonomous driving, natural weather, human action, and physics-based simulation, demonstrate the effectiveness of VDT. Moreover, we provide a comprehensive study on the capabilities of VDT in capturing accurate temporal dependencies, handling conditioning information, and the spatial-temporal mask modeling mechanism. Additionally, we present comprehensive studies on how VDT handles conditioning information with the mask modeling mechanism, which we believe will benefit future research and advance the field. Codes and models are available at the https://VDT-2023.github.io. Haoyu Lu, Guoxing Yang, Nanyi Fei, Yuqi Huo, Zhiwu Lu 0001, Ping Luo 0002, Mingyu Ding |
ICLR | 2 |
| 2024 | FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingabstractContrastive Language-Image Pre-training (CLIP) achieves impressive performance on tasks like image classification and image-text retrieval by learning on large-scale image-text datasets. However, CLIP struggles with dense prediction tasks due to the poor grasp of the fine-grained details. Although existing works pay attention to this issue, they achieve limited improvements and usually sacrifice the important visual-semantic consistency. To overcome these limitations, we propose FineCLIP, which keeps the global contrastive learning to preserve the visual-semantic consistency and further enhances the fine-grained understanding through two innovations: 1) A real-time self-distillation scheme that facilitates the transfer of representation capability from global to local features. 2) A semantically-rich regional contrastive learning paradigm with generated region-text pairs, boosting the local representation capabilities with abundant fine-grained knowledge.
Both cooperate to fully leverage diverse semantics and multi-grained complementary information.
To validate the superiority of our FineCLIP and the rationality of each design, we conduct extensive experiments on challenging dense prediction and image-level tasks.
All the observations demonstrate the effectiveness of FineCLIP. Dong Jing, Xiaolong He 0003, Yutian Luo, Nanyi Fei, Guoxing Yang, Huiwen Zhao, Zhiwu Lu 0001 |
NeurIPS | 5 |
| 2023 | Song-to-Video Translation: Writing a Video from Song Lyrics Based on Multimodal Pre-training
Feifei Fu, Zelong Sun, Guoxing Yang, Xiaolong He 0003, Zhiwu Lu 0001 |
ADMA (2) | 3 |
| 2023 | Enhancing Medical Language Understanding: Adapting LLMs to the Medical Domain through Hybrid Granularity Mask LearningabstractLarge Language models have made remarkable strides in natural language understanding and generation. However, their performance in specialized fields like medicine often falls short due to the lack of domain-specific knowledge during pre-training. While fine-tuning on labeled medical data is a common approach for task adaptation, it may not capture the comprehensive medical knowledge required. In this paper, we proposed a Hybrid Granularity Mask Learning (HGM) method for domain adaptation in the medical field. Our method incorporates multi-level linguistic characteries including token, entity, and subsentence to enable the model to acquire medical knowledge comprehensively. We fine-tune a medical-specific language model derived from ChatGLM-6B and Bloom-7B on downstream medical tasks and evaluate its performance. The results demonstrate a significant improvement compared to the baseline, thus affirming the effectiveness of our proposed method. Longjun Fan, Xiaohong Liu 0007, Guoxing Yang, Zongxin Du |
BIBM | 4 |
| 2023 | DiST-GAN: Distillation-based Semantic Transfer for Text-Guided Face GenerationabstractRecently, large-scale pre-training has achieved great success in multi-modal tasks and shown powerful generalization ability due to superior semantic comprehension. In the field of text-to-image synthesis, recent works induce large-scale pre-training with VQ-VAE as a discrete visual tokenizer, which can synthesize realistic images from arbitrary text inputs. However, the quality of images generated by these methods is still inferior to that of images generated by GAN-based methods, especially in some specific domains. To leverage both the superior semantic comprehension of large-scale pre-training models and the powerful ability of GAN-based models in photorealistic image generation, we propose a novel knowledge distillation framework termed DiST-GAN to transfer the semantic knowledge of large-scale visual-language pre-training models (e.g., CLIP) to GAN-based generator for text-guided face image generation. Our DiST-GAN consists of two key components: (1) A new CLIP-based adaptive contrastive loss is devised to ensure the generated images are consistent with the input texts. (2) A language-to-vision (L2V) transformation module is learned to transform token embeddings of each text into an intermediate embedding that is aligned with the image embedding extracted by CLIP. With these two novel components, the semantic knowledge contained in CLIP can thus be transferred to GAN-based generator which preserves the superior ability of photorealistic image generation in the mean time. Extensive results on the Multi-Modal CelebA-HQ dataset show that our DiST-GAN achieves significant improvements over the state-of-the-arts. Guoxing Yang, Feifei Fu, Nanyi Fei, Ruitao Ma, Zhiwu Lu 0001 |
ICME | 1 |
| 2023 | Shot Retrieval and Assembly with Text Script for Video Montage GenerationabstractWith the development of video sharing websites, numerous users desire to create their own attractive video montages. However, it is difficult for inexperienced users to create well-edited video montages due to the lack of professional expertise. In the meantime, it is time-consuming even for experts to create video montages of high quality, which requires effectively selecting shots from abundant candidates and assembling them together. Instead of manual creation, various automatic methods have been proposed for video montage generation, which typically take a single sentence as input for text-to-shot retrieval, and ignore the semantic cross-sentence coherence given complicated text script of multiple sentences. To overcome this drawback, we propose a novel model for video montage generation by retrieving and assembling shots with arbitrary text scripts. To this end, a sequence consistency transformer is devised for cross-sentence coherence modeling. More importantly, with this transformer, two novel sequence-level tasks are defined for sentence-shot alignment in sequence-level: Cross-Modal Sequence Matching (CMSM) task, and Chaotic Sequence Recovering (CSR) task. To facilitate the research on video montage generation, we construct a new, highly-varied dataset which collects thousands of video-script pairs in documentary. Extensive experiments on the constructed dataset demonstrate the superior performance of the proposed model. The dataset and generated video demos are available at https://github.com/RATVDemo/RATV. Guoxing Yang, Haoyu Lu, Zelong Sun, Zhiwu Lu 0001 |
ICMR | 1 |
| 2022 | Enhanced CT Image Generation by GAN for Improving Thyroid Anatomy DetectionabstractComputed tomography (CT) is one of the most imaging methods widely used to locate lesions such as nodules, tumors, and cysts, and make primary diagnosis. For clearer imaging of anatomical or lesions, contrast-enhanced CT (CECT) scans are imaging with injecting a contrast agent into a patient during examination. But there are limits to iodine contrast injections so that CECT scans are not convenient like non-contrast enhanced CT (NECT). Recently, deep learning models bring impressive results in computer vision, including image translation. So, we would like to apply image translation methods to generate CECT images from the more accessible NECT images, and evaluate the effects of generated images on image detection tasks. In this study, we propose a method called cross-modal enhancement training strategy for thyroid anatomy detection, which employs CycleGAN to translate non-constrast enhanced CT images to enhanced CT style images with content reserved. The experiments are conducted on thyroid CT images with anatomy object annotation. The experimental results show that by adding translated images into the training dataset, the performance of thyroid anatomy detection can be effectively improved. We achieve the best mAP of 82.5% compared to 73.2% in the along non-contrast enhanced CT training. Jianyu Shi, Xiaohong Liu 0007, Guoxing Yang |
BIBM | 3 |
| 2022 | AIAT: Adaptive Iteration Adversarial Training for Robust Pulmonary Nodule DetectionabstractLung cancer is one of the leading causes of death worldwide. Early diagnosis through cancer screening can significantly improve lung cancer patients’ survival. Recently, deep learning based diagnostic systems for nodule detection have shown great potential in assisting radiologists to screen cancer more efficiently. However, studies have found that deep learning models lack robustness against imperceptible crafted adversarial attacks and few studied improving the robustness of pulmonary nodule detection. Therefore, making pulmonary nodule detection models robust remains challenges. Moreover, traditional adversarial training methods either hurt the natural generalization or need expensive computational cost. To address these challenges, here we propose a novel adversarial training method called, Adaptive Iteration Adversarial Training (AIAT). AIAT generates adversarial samples by adding adversarial noise with an adaptive iteration strategy, so that it can stably and fast train models with improving robustness. Extensive experiments on the LUNA 16 dataset show that AIAT improves robustness for pulmonary nodule detection without compromising the natural generalization, and largely reduces training time. Guoxing Yang, Xiaohong Liu 0007, Jianyu Shi, Xianchao Zhang 0002 |
BIBM | 1 |
| 2022 | Visual Prompt Tuning for Few-Shot Text ClassificationabstractDeploying large-scale pre-trained models in the prompt-tuning paradigm has demonstrated promising performance in few-shot learning. Particularly, vision-language pre-training models (VL-PTMs) have been intensively explored in various few-shot downstream tasks. However, most existing works only apply VL-PTMs to visual tasks like image classification, with few attempts being made on language tasks like text classification. In few-shot text classification, a feasible paradigm for deploying VL-PTMs is to align the input samples and their category names via the text encoders. However, it leads to the waste of visual information learned by the image encoders of VL-PTMs. To overcome this drawback, we propose a novel method named Visual Prompt Tuning (VPT). To our best knowledge, this method is the first attempt to deploy VL-PTM in few-shot text classification task. The main idea is to generate the image embeddings w.r.t. category names as visual prompt and then add them to the aligning process. Extensive experiments show that our VPT can achieve significant improvements under both zero-shot and few-shot settings. Importantly, our VPT even outperforms the most recent prompt-tuning methods on five public text classification datasets. Jingyuan Wen, Yutian Luo, Nanyi Fei, Guoxing Yang, Zhiwu Lu 0001, Hao Jiang 0022, Zhao Cao |
COLING | 4 |
| 2021 | L2M-GAN: Learning To Manipulate Latent Space Semantics for Facial Attribute EditingabstractA deep facial attribute editing model strives to meet two requirements: (1) attribute correctness – the target attribute should correctly appear on the edited face image; (2) irrelevance preservation – any irrelevant information (e.g., identity) should not be changed after editing. Meeting both requirements challenges the state-of-the-art works which resort to either spatial attention or latent space factorization. Specifically, the former assume that each attribute has well-defined local support regions; they are often more effective for editing a local attribute than a global one. The latter factorize the latent space of a fixed pretrained GAN into different attribute-relevant parts, but they cannot be trained end-to-end with the GAN, leading to sub-optimal solutions. To overcome these limitations, we propose a novel latent space factorization model, called L2M-GAN, which is learned end-to-end and effective for editing both local and global attributes. The key novel components are: (1) A latent space vector of the GAN is factorized into an attribute-relevant and irrelevant codes with an orthogonality constraint imposed to ensure disentanglement. (2) An attribute-relevant code transformer is learned to manipulate the attribute value; crucially, the transformed code are subject to the same orthogonality constraint. By forcing both the original attribute-relevant latent code and the edited code to be disentangled from any attribute-irrelevant code, our model strikes the perfect balance between attribute correctness and irrelevance preservation. Extensive experiments on CelebA-HQ show that our L2M-GAN achieves significant improvements over the state-of-the-arts. Guoxing Yang, Nanyi Fei, Mingyu Ding, Guangzhen Liu, Zhiwu Lu 0001, Tao Xiang 0002 |
CVPR | 1 |
| 2021 | Complex Action Segmentation in Compressed VideosabstractComplex action segmentation aims to detect what actions and when they happen in fine-grained level from long videos. Despite the fact that videos are often stored in a compressed format (e.g., MPEG-4), most existing approaches are proposed to directly model raw RGB videos: when only compressed videos are accessible, they have to first decode these videos, which is very time-consuming. In this paper, by explicitly leveraging the ‘compressed’ characteristic of compressed videos, we are the first to address the challenging task of complex action segmentation in compressed videos. To extract meaningful representations for complex action segmentation, we introduce the GOP-Level Compressed features (Golec), which can be obtained directly from compressed videos without video decompression. Importantly, by taking GOPs as the atomic units of actions, our Golec representation is intrinsically suitable for fine-grained action segmentation. Moreover, to remedy the coarser motion vectors (compared with optical flows which are computed from raw frames) used in our Golec representation for capturing the temporal context, we propose a new Bi-path knowledge distillation strategy. Extensive experiments show the effectiveness of our Golec representation and the Bi-path strategy. Importantly, our proposed model for complex action detection not only runs 5.2 times faster but also achieves significantly better results than the state-of-the-art alternatives using raw videos. Hongfeng Han, Guoxing Yang, Yuqi Huo, Zhiwu Lu 0001, Ji-Rong Wen |
ICME | 2 |