EDBT 2026 Demo / reviewers in the wild / expert
Fangxiang Feng
dblp:132/9036
· DBLP profile ↗
39ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0002-4798-4233ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 18 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase LocalizationabstractWeakly supervised phrase localization (WSPL) aims to localize visual objects mentioned by given phrases, but it learns without human-annotated bounding boxes. Previous works struggle in multi-object scenarios where objects in the background often appear simultaneously with the target objects. To this end, we propose a Diffusion-Assisted PrOgressive learning framework (i.e., DAPO) for WSPL task in this paper. Specifically, we score the difficulty of training samples based on the quantity of objects and the level of semantic alignment. These samples are then used progressively during training, in an order by their difficulty scores. To address the sample imbalance problem, we propose a Generation-Assisted Tuning (GAT) method for the grounding network. First, to enrich the samples from few-object scenarios, we leverage Stable Diffusion (SD) to generate images with phrases. Second, we introduce an attention-driven scheme to direct SD's attention on the mentioned objects. Finally, we design a diffusion-guided loss, which helps the grounding network learn the objects' layouts. Extensive experiments show that our DAPO framework outperforms the strong baselines on benchmark datasets. Pengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia 0006, Fangxiang Feng, Ruifan Li |
AAAI | 5 |
| 2026 | Towards balancing the efficiency and effectiveness: a unified edit-based framework for automatic image captioning
Ruifan Li, Siwei Xu, Pengyue Lin, Fangxiang Feng, Zhangyu Ma |
Multim. Syst. | 4 |
| 2026 | A fine-grained entity understanding network for weakly supervised phrase grounding
Pengyue Lin, Ruifan Li, Fangxiang Feng, Lun Ke, Zhanyu Ma, Xiaojie Wang 0006 |
Pattern Recognit. | 3 |
| 2025 | Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image SynthesisabstractThe customization of text-to-image models has seen significant advancements, yet generating multiple personalized concepts remains a challenging task. Current methods struggle with attribute leakage and layout confusion when handling multiple concepts, leading to reduced concept fidelity and semantic consistency. In this work, we introduce a novel training-free framework, Concept Conductor, designed to ensure visual fidelity and correct layout in multi-concept customization. Concept Conductor isolates the sampling processes of multiple customized models to prevent attribute leakage between different concepts and corrects erroneous layouts through self-attention-based spatial guidance. Additionally, we present a concept injection technique that employs shape-aware masks to specify the generation area for each concept. This technique injects the structure and appearance of personalized concepts through feature fusion in the attention layers, ensuring harmony in the final image. Extensive qualitative and quantitative experiments demonstrate that Concept Conductor can consistently generate composite images with accurate layouts while preserving the visual details of each concept. Compared to existing baselines, Concept Conductor shows significant performance improvements. Our method supports the combination of any number of concepts and maintains high fidelity even when dealing with visually similar concepts. The code and trained models will be made publicly available. Zebin Yao, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
AAAI | 2 |
| 2025 | VL-DynaRefine: A Vision-Language Dynamic Refinement Approach for Visual ReasoningabstractVisual reasoning is a key capability that significantly impacts the performance of multimodal tasks, such as compositional visual question answering and visual grounding. These tasks often require complex, multi-step reasoning processes. In recent years, several training-free methods for Vision-Language Models (VLMs) have emerged, with visual programming methods being proposed to enhance the capability of VLMs in visual reasoning tasks. While these methods have made some progress, they still face two primary challenges due to the lack of verification and refinement mechanisms for each action's output during the reasoning process: error accumulation and feedback delay, as well as insufficient utilization of multimodal contextual information. To address these challenges, we propose VL-DynaRefine, a training-free approach consisting of three modules: a planner, a verifier, and a refiner. The planner generates programmatic actions to solve the problem and executes each action in sequence, which is inspected by a verifier that reassesses the actions via confidence scores and determines whether refinement is necessary based on the evaluation results. In the refiner module, we incorporate a context-aware local refinement mechanism and a global refinement mechanism based on visual and action trajectories to reduce the impact of reasoning errors on the outcome. We evaluate our approach on multiple visual reasoning datasets, and the experimental results show that our method outperforms existing visual programming methods in both reasoning accuracy and efficiency, further validating its effectiveness in visual reasoning tasks. Zeyuan Zang, Fangxiang Feng, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006 |
ACM Multimedia | 4 |
| 2025 | Exploiting Prior Tacit Knowledge to Enhance Alignment and Verification in zero-shot video grounding
Jing Wang 0169, Xianbing Zhao, Xiaojie Wang 0006, Fangxiang Feng |
Neurocomputing | 4 |
| 2024 | Visual Prompt Tuning for Weakly Supervised Phrase GroundingabstractPrevious works on the task of weakly supervised phrase grounding (WSG) rely heavily on object detectors providing RoIs for the localization. However, such methods cannot be applied effectively to real-world scenarios largely because that the detectors are trained with limited categories. In this paper, we propose a refinement-based approach to WSG through fine-tuning a detector-free phrase grounding model with a visual prompt. This visual prompt is extracted from the text-related representations in CLIP. Furthermore, we combine the visual prompt with learnable features and then fine-tune the grounding network. Our experimental results significantly outperform state-of-the-art methods on the WSG task and shows the effectiveness of our method. Pengyue Lin, Zhihan Yu, Mingcong Lu, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
ICASSP | 4 |
| 2024 | DiffHarmony: Latent Diffusion Model Meets Image HarmonizationabstractImage harmonization, which involves adjusting the foreground of a composite image to attain a unified visual consistency with the background, can be conceptualized as an image-to-image translation task. Diffusion models have recently promoted the rapid development of image-to-image translation tasks . However, training diffusion models from scratch is computationally intensive. Fine-tuning pre-trained latent diffusion models entails dealing with the reconstruction error induced by the image compression autoencoder, making it unsuitable for image generation tasks that involve pixel-level evaluation metrics. To deal with these issues, in this paper, we first adapt a pre-trained latent diffusion model to the image harmonization task to generate the harmonious but potentially blurry initial images. Then we implement two strategies: utilizing higher-resolution images during inference and incorporating an additional refinement stage, to further enhance the clarity of the initially harmonized images. Extensive experiments on iHarmony4 datasets demonstrate the superiority of our proposed method. The code is available at \hrefhttps://github.com/nicecv/DiffHarmony https://github.com/nicecv/DiffHarmony. Fangxiang Feng, Xiaojie Wang 0006 |
ICMR | 2 |
| 2024 | Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak SupervisionabstractPhrase Grounding, i.e., PG aims to locate objects referred by noun phrases. Recently, PG under weak supervision (i.e., grounding without region-level annotations) and zero-shot PG (i.e., grounding from seen categories to unseen ones) are proposed, respectively. However, for real-world applications these two approaches are limited due to slight annotations and numerable categories during training. In this paper, we propose a framework of zero-shot PG under weak supervision. Specifically, our PG framework is built on triple alignment strategies. Firstly, we propose a region-text alignment (RTA) strategy to build region-level attribute associations via CLIP. Secondly, we propose a domain alignment (DomA) strategy by minimizing the difference between distributions of seen classes in the training and those of the pre-training. Thirdly, we propose a category alignment (CatA) strategy by considering both category semantics and region-category relations. Extensive experimental results show that our proposed PG framework outperforms previous zero-shot methods and weakly-supervised methods. Our code is available at https://github.com/LinPengyue/ZS-WSG. Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006 |
ACM Multimedia | 5 |
| 2024 | Q-MoE: Connector for MLLMs with Text-Driven RoutingabstractMultimodal Large Language Models (MLLMs) have showcased remarkable advances in handling various vision-language tasks. These models typically consist of a Large Language Model (LLM), a vision encoder and a connector structure, which is used to bridge the modality gap between vision and language. It is challenging for the connector to filter the right visual information for LLM according to the task in hand. Most of previous connectors, such as light-weight projection and Q-former, treat visual information for diverse tasks uniformly, therefore lacking task-specific visual information extraction capabilities. To address the issue, this paper proposes Q-MoE, a query-based connector with Mixture-of-Experts (MoE) to extract task-specific information with text-driven routing. Furthermore, an optimal path based training strategy is proposed to find an optimal expert combination. Extensive experiments on two popular open-source LLMs and several different visual-language tasks demonstrate the effectiveness of the Q-MoE connecter. Hanzi Wang, Jiamin Ren, Huixing Jiang, Fangxiang Feng, Xiaojie Wang 0006 |
ACM Multimedia | 7 |
| 2024 | DiffHarmony++: Enhancing Image Harmonization with Harmony-VAE and Inverse Harmonization ModelabstractLatent diffusion model has demonstrated impressive efficacy in image generation and editing tasks. Recently, it has also promoted the advancement of image harmonization. However, methods involving latent diffusion model all face a common challenge: the severe image distortion introduced by the VAE component, while image harmonization is a low-level image processing task that relies on pixel-level evaluation metrics. In this paper, we propose Harmony-VAE, leveraging the input of the harmonization task itself to enhance the quality of decoded images. The input involving composite image contains the precise pixel level information, which can complement the correct foreground appearance and color information contained in denoised latents. Meanwhile, the inherent generative nature of diffusion models makes it naturally adapt to inverse image harmonization, i.e. generating synthetic composite images based on real images and foreground masks. We train an inverse harmonization diffusion model to perform data augmentation on two subsets of iHarmony4 and construct a new human harmonization dataset with prominent foreground objects. Extensive experiments demonstrate the effectiveness of our proposed Harmony-VAE and inverse harmonization model. Code and pretrained models are available at https://github.com/nicecv/DiffHarmony. Fangxiang Feng, Guang Liu 0006, Ruifan Li, Xiaojie Wang 0006 |
ACM Multimedia | 2 |
| 2024 | LGR-NET: Language Guided Reasoning Network for Referring Expression ComprehensionabstractReferring Expression Comprehension(REC) is a fundamental task in the vision and language domain, which aims to locate an image region according to a natural language expression. REC requires the models to capture key clues in the text and perform accurate cross-modal reasoning. A recent trend employs transformer-based methods to address this problem. However, most of these methods typically treat image and text equally. They usually perform cross-modal reasoning in a crude way, and utilize textual features as a whole without detailed considerations (e.g., spatial information). This insufficient utilization of textual features will lead to sub-optimal results. In this paper, we propose aLanguage Guided Reasoning Network(LGR-NET) to fully utilize the guidance of the referring expression. To localize the referred object, we set a prediction token to capture cross-modal features. Furthermore, to sufficiently utilize the textual features, we extend them by our Textual Feature Extender (TFE) from three aspects.First, we design a novel coordinate embedding based on textual features. The coordinate embedding is incorporated to the prediction token to promote its capture of language-related visual features.Second, we employ the extracted textual features for Text-guided Cross-modal Alignment (TCA) and Fusion (TCF), alternately.Third, we devise a novel cross-modal loss to enhance cross-modal alignment between the referring expression and the learnable prediction token. We conduct extensive experiments on five benchmark datasets, and the experimental results show that our LGR-NET achieves a new state-of-the-art. Source code is available at https://github.com/lmc8133/LGR-NET. Mingcong Lu, Ruifan Li, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DualGCN: Exploring Syntactic and Semantic Information for Aspect-Based Sentiment AnalysisabstractThe task of aspect-based sentiment analysis aims to identify sentiment polarities of given aspects in a sentence. Recent advances have demonstrated the advantage of incorporating the syntactic dependency structure with graph convolutional networks (GCNs). However, their performance of these GCN-based methods largely depends on the dependency parsers, which would produce diverse parsing results for a sentence. In this article, we propose a dual GCN (DualGCN) that jointly considers the syntax structures and semantic correlations. Our DualGCN model mainly comprises four modules: 1) SynGCN: instead of explicitly encoding syntactic structure, the SynGCN module uses the dependency probability matrix as a graph structure to implicitly integrate the syntactic information; 2) SemGCN: we design the SemGCN module with multihead attention to enhance the performance of the syntactic structure with the semantic information; 3) Regularizers: we propose orthogonal and differential regularizers to precisely capture semantic correlations between words by constraining attention scores in the SemGCN module; and 4) Mutual BiAffine: we use the BiAffine module to bridge relevant information between the SynGCN and SemGCN modules. Extensive experiments are conducted compared with up-to-date pretrained language encoders on two groups of datasets, one including Restaurant14, Laptop14, and Twitter and the other including Restaurant15 and Restaurant16. The experimental results demonstrate that the parsing results of various dependency parsers affect their performance of the GCN-based models. Our DualGCN model achieves superior performance compared with the state-of-the-art approaches. The source code and preprocessed datasets are provided and publicly available on GitHub (see https://github.com/CCChenhao997/DualGCN-ABSA). Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Whether you can locate or not? Interactive Referring Expression GenerationabstractReferring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object. Existing methods construct REG models independently by using only the REs as ground truth for model training, without considering the potential interaction between REG and REC models. In this paper, we propose an Interactive REG (IREG) model that can interact with a real REC model, utilizing signals indicating whether the object is located and the visual region located by the REC model to gradually modify REs. Our experimental results on three RE benchmark datasets, RefCOCO, RefCOCO+, and RefCOCOg show that IREG outperforms previous state-of-the-art methods on popular evaluation metrics. Furthermore, a human evaluation shows that IREG generates better REs with the capability of interaction. Fulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie Wang 0006 |
ACM Multimedia | 3 |
| 2022 | Enhanced Multi-Channel Graph Convolutional Network for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) is an emerging sentiment analysis task.Most of the existing studies focus on devising a new tagging scheme that enables the model to extract the sentiment triplets in an end-to-end fashion.However, these methods ignore the relations between words for ASTE task.In this paper, we propose an Enhanced Multi-Channel Graph Convolutional Network model (EMC-GCN) to fully utilize the relations between words.Specifically, we first define ten types of relations for ASTE task, and then adopt a biaffine attention module to embed these relations as an adjacent tensor between words in a sentence.After that, our EMC-GCN transforms the sentence into a multi-channel graph by treating words and the relation adjacent tensor as nodes and edges, respectively.Thus, relationaware node representations can be learnt.Furthermore, we consider diverse linguistic features to enhance our EMC-GCN model.Finally, we design an effective refining strategy on EMC-GCN for word-pair representation refinement, which considers the implicit results of aspect and opinion extraction when determining whether word pairs match or not.Extensive experimental results on the benchmark datasets demonstrate that the effectiveness and robustness of our proposed model, which outperforms state-of-the-art methods significantly. Hao Chen 0041, Zepeng Zhai, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
ACL (1) | 3 |
| 2022 | A Simple Model for Distantly Supervised Relation ExtractionabstractDistantly supervised relation extraction is challenging due to the noise within data. Recent methods focus on exploiting bag representations based on deep neural networks with complex de-noising scheme to achieve remarkable performance. In this paper, we propose a simple but effective BERT-based Graph convolutional network Model (i.e., BGM). Our BGM comprises of an instance embedding module and a bag representation module. The instance embedding module uses a BERT-based pretrained language model to extract key information from each instance. The bag representaion module constructs the corresponding bag graph then apply a convolutional operation to obtain the bag representation. Our BGM model achieves a considerable improvement on two benchmark datasets, i.e., NYT10 and GDS. Ziqin Rao, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
COLING | 2 |
| 2022 | COM-MRC: A COntext-Masked Machine Reading Comprehension Framework for Aspect Sentiment Triplet ExtractionabstractAspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from sentences, which was recently formalized as an effective machine reading comprehension (MRC) based framework.However, when facing multiple aspect terms, the MRC-based methods could fail due to the interference from other aspect terms.In this paper, we propose a novel COntext-Masked MRC (COM-MRC) framework for ASTE.Our COM-MRC framework comprises three closely-related components: a context augmentation strategy, a discriminative model, and an inference method.Specifically, a context augmentation strategy is designed by enumerating all masked contexts for each aspect term.The discriminative model comprises four modules, i.e., aspect and opinion extraction modules, sentiment classification and aspect detection modules.In addition, a two-stage inference method first extracts all aspects and then identifies their opinions and sentiment through iteratively masking the aspects.Extensive experimental results on benchmark datasets show the effectiveness of our proposed COM-MRC framework, which outperforms state-of-the-art methods consistently 1 . Zepeng Zhai, Hao Chen 0041, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
EMNLP | 3 |
| 2022 | Improving Image Paragraph Captioning with Dual RelationsabstractImage paragraph captioning aims to generate multiple de-scriptive sentences for an image. However, most previous methods ignore the explicit relations among objects resulting in unsatisfactory performance. In this paper, we propose a novel model (i.e., DualRel) to capture spatial and seman-tic relations among objects. Specifically, the spatial relation embedding is obtained solely from images using a predefined geometry pattern. With the help of captions, the semantic relation embedding is learned in a weakly supervised man-ner. These two relation embeddings are then interacted with regional features of objects through a relation-aware attention interaction. It first obtains a visual context vector using regional features. Then with the visual context vector, we obtain the corresponding spatial and semantic relation-aware vectors using attentions. These three vectors are fused with two gates for language decoding to further generate a para-graph. Experimental results on Stanford benchmark dataset show that DualRel achieves remarkable improvements11Code released at https://github.com/fuyunll07/DualRel. Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006 |
ICME | 3 |
| 2022 | Question-Driven Graph Fusion Network for Visual Question AnsweringabstractExisting Visual Question Answering (VQA) models have ex-plored various visual relationships between objects in the im-age to answer complex questions, which inevitably introduces irrelevant information brought by inaccurate object detection and text grounding. To address the problem, we propose a Question-Driven Graph Fusion Network (QD-GFN). It first models semantic, spatial, and implicit visual relations in images by three graph attention networks, then question in-formation is utilized to guide the aggregation process of the three graphs, further, our QD-GFN adopts an object filtering mechanism to remove question-irrelevant objects contained in the image. Experiment results demonstrate that our QD-GFN outperforms the prior state-of-the-art on both VQA 2.0 and VQA-CP v2 datasets. Further analysis shows that both the novel graph aggregation method and object filtering mecha-nism play a significant role in improving the performance of the model. Yuxi Qian, Yuncong Hu, Fangxiang Feng, Xiaojie Wang 0006 |
ICME | 4 |
| 2022 | GR-GAN: Gradual Refinement Text-To-Image GenerationabstractA good Text-to-Image model should not only generate high quality images, but also ensure the consistency between the text and the generated image. Previous models failed to simultaneously fix both sides well. This paper proposes a Gradual Refinement Generative Adversarial Network (GR-GAN) to alleviates the problem efficiently. A GRG module is designed to generate images from low resolution to high resolution with the corresponding text constraints from coarse granularity (sentence) to fine granularity (word) stage by stage, a ITM module is designed to provide image-text matching losses at both sentence-image level and word-region level for corresponding stages. We also introduce a new metric Cross-Model Distance (CMD) for simultaneously evaluating image quality and image-text consistency. Experimental results show GR-GAN significant outperform previous models, and achieve new state-of-the-art on both FID and CMD. A detailed analysis demonstrates the efficiency of different generation stages in GR-GAN. Fangxiang Feng, Xiaojie Wang 0006 |
ICME | 2 |
| 2022 | A Region-based Document VQAabstractPractical Document Visual Question Answering (DocVQA) needs not only to recognize and extract the document contents, but also reason on them for answering questions. However, previous DocVQA data mainly focuses on in-line questions, where the answers could be directly extracted after locating keywords in the documents, which needs less reasoning. This paper therefore builds a large-scale dataset named Region-based Document VQA (RDVQA), which includes more practical questions for DocVQA. We then propose a novel Reason-over-In-region-Question-answering (ReIQ) model for addressing the problems. It is a pre-training-based model, where a Spatial-Token Pre-trained Model (STPM) is employed as the backbone. Two novel pre-training tasks, Masked Text Box Regression and Shuffled Triplet Reconstruction, are proposed to learn the entailment relationship between text blocks and tokens as well as contextual information, respectively. Moreover, a DocVQA State Tracking Module (DocST) is also proposed to track the DocVQA state in the fine-tuning stage. Experimental results show that our model improves the performance onRDVQA significantly, although more work should be done for practical DocVQA as shown inRDVQA. Xinya Wu, Duo Zheng, Jiashen Sun, Minzhen Hu, Fangxiang Feng, Xiaojie Wang 0006, Huixing Jiang, Fan Yang 0087 |
ACM Multimedia | 6 |
| 2022 | Visual Dialog for Spotting the Differences between Pairs of Similar ImagesabstractVisual dialog has witnessed great progress after introducing various vision-oriented goals into the conversation. Much of previous work focuses on tasks where only one image can be accessed by two interlocutors, such as VisDial and GuessWhat. The work on situations where two interlocutors access different images has received less attention. Those situations are common in real world and bring some different challenges compared with one-image tasks. The lack of such types of dialog tasks and corresponding large-scale datasets makes it impossible to carry out in-depth research. This paper therefore first proposes a new visual dialog task named Dial-the-Diff, where two interlocutors accessing two similar images respectively try to spot the difference between the images through conversing in natural language. The task raises new challenges to the dialog strategy and the ability of categorizing objects. We then build a large-scale multi-modal dataset for the task, named DialDiff, which contains 87k Virtual Reality images and 78k dialogs. Some details of the data are given and analyzed to highlight the challenges behind the task. Finally, we propose benchmark models for this task, and conduct extensive experiments to evaluate their performance as well as its problems remained. Duo Zheng, Fandong Meng, Qingyi Si, Hairun Fan, Zipeng Xu, Jie Zhou 0016, Fangxiang Feng, Xiaojie Wang 0006 |
ACM Multimedia | 7 |
| 2022 | A deformable CNN-based triplet model for fine-grained sketch-based image retrieval
Xianlin Zhang, Mengling Shen, Xueming Li 0002, Fangxiang Feng |
Pattern Recognit. | 4 |
| 2022 | Modality Disentangled Discriminator for Text-to-Image SynthesisabstractText-to-image (T2I) synthesis aims at generating photo-realistic images from text descriptions, which is a particularly important task in bridging vision and language. Each generated image consists of two parts: the content part related to the text and the style part irrelevant to the text. The existing discriminator does not distinguish between the content part and the style part. This not only precludes the T2I synthesis models from generating the content part effectively but also makes it difficult to manipulate the style of the generated image. In this paper, we propose a modality disentangled discriminator that distinguishes between the content part and the style part at a specific layer. Specifically, we enforce the early layers of a certain number in the discriminator to become the disentangled representation extractor through two losses. The extracted common representation for the content part can make the discriminator more effective for capturing the text-image correlation, while the extracted modality-specific representation for the style part can be directly transferred to other images. The combination of these two representations can also improve the quality of the generated images. Our proposed discriminator is used to substitute the discriminator of each stage in the representative model AttnGAN and the SOTA model DM-GAN. Extensive experiments are conducted on three widely used datasets, i.e. CUB, Oxford-102, and COCO, for the T2I synthesis task, demonstrating the superior performance of the modality disentangled discriminator over the base models. Code for DM-GAN with our modality disentangled discriminator is available athttps://github.com/FangxiangFeng/DM-GAN-MDD. Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006 |
IEEE Trans. Multim. | 1 |
| 2021 | Dual Graph Convolutional Networks for Aspect-based Sentiment AnalysisabstractRuifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
ACL/IJCNLP (1) | 3 |
| 2021 | Multi-stage Pre-training over Simplified Multimodal Pre-training ModelsabstractTongtong Liu, Fangxiang Feng, Xiaojie Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Fangxiang Feng, Xiaojie Wang 0006 |
ACL/IJCNLP (1) | 2 |
| 2021 | S2TD: A Tree-Structured Decoder for Image Paragraph CaptioningabstractImage paragraph captioning, a task to generate the paragraph description for a given image, usually requires mining and organizing linguistic counterparts from abundant visual clues. Limited by sequential decoding perspective, previous methods have difficulty in organizing the visual clues holistically or capturing the structural nature of linguistic descriptions. In this paper, we propose a novel tree-structured visual paragraph decoder network, called Splitting to Tree Decoder (S2TD) to address this problem. The key idea is to model the paragraph decoding process as a top-down binary tree expansion. S2TD consists of three modules: a split module, a score module, and a word-level RNN. The split module iteratively splits ancestral visual representations into two parts through a gating mechanism. To determine the tree topology, the score module uses cosine similarity to evaluate the nodes splitting. A novel tree structure loss is proposed to enable end-to-end learning. After the tree expansion, the word-level RNN decodes leaf nodes into sentences forming a coherent paragraph. Extensive experiments are conducted on the Stanford benchmark dataset. The experimental results show promising performance of our proposed S2TD. Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006 |
MMAsia | 3 |
| 2020 | Image Synthesis from Locally Related TextsabstractText-to-image synthesis refers to generating photo-realistic images from text descriptions. Recent works focus on generating images with complex scenes and multiple objects. However, the text inputs to these models are the only captions that always describe the most apparent object or feature of the image and detailed information (e.g. visual attributes) for regions and objects are often missing. Quantitative evaluation of generation performances is still an unsolved problem, where traditional image classification- or retrieval-based metrics fail at evaluating complex images. To address these problems, we propose to generate images conditioned on locally-related texts, i.e., descriptions of local image regions or objects instead of the whole image. Specifically, questions and answers (QAs) are chosen as locally-related texts, which makes it possible to use VQA accuracy as a new evaluation metric. The intuition is simple: higher image quality and image-text consistency (both globally and locally) can help a VQA model answer questions more correctly. We purposed VQA-GAN model with three key modules: hierarchical QA encoder, QA-conditional GAN and external VQA loss. These modules help leverage the new inputs effectively. Thorough experiments on two public VQA datasets demonstrate the effectiveness of the model and the newly proposed metric. Tianrui Niu, Fangxiang Feng, Lingxuan Li, Xiaojie Wang 0006 |
ICMR | 2 |
| 2020 | Learning Visual Features from Product Title for Image RetrievalabstractThere is a huge market demand for searching for products by images in e-commerce sites. Visual features play the most important role in solving this content-based image retrieval task. Most existing methods leverage pre-trained models on other large-scale datasets with well-annotated labels, e.g. the ImageNet dataset, to extract visual features. However, due to the large difference between the product images and the images in ImageNet, the feature extractor trained on ImageNet is not efficient in extracting the visual features of product images. And retraining the feature extractor on the product images is faced with the dilemma of lacking the annotated labels. In this paper, we utilize the easily accessible text information, that is, the product title, as a supervised signal to learn the features of the product image. Specifically, we use the n-grams extracted from the product title as the label of the product image to construct a dataset for image classification. This dataset is then used to fine-tuned a pre-trained model. Finally, the basic max-pooling activation of convolutions (MAC) feature is extracted from the fine-tuned model. As a result, we achieve the fourth position in the Grand Challenge of AI Meets Beauty in 2020 ACM Multimedia by using only a single ResNet-50 model without any human annotations and pre-processing or post-processing tricks. Our code is available at: \urlhttps://github.com/FangxiangFeng/AI-Meets-Beauty-2020. Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006, Huixing Jiang |
ACM Multimedia | 1 |
| 2020 | Answer-Driven Visual State Estimator for Goal-Oriented Visual DialogueabstractA goal-oriented visual dialogue involves multi-turn interactions between two agents, Questioner and Oracle. During which, the answer given by Oracle is of great significance, as it provides golden response to what Questioner concerns. Based on the answer, Questioner updates its belief on target visual content and further raises another question. Notably, different answers drive into different visual beliefs and future questions. However, existing methods always indiscriminately encode answers after much longer questions, resulting in a weak utilization of answers. In this paper, we propose an Answer-Driven Visual State Estimator (ADVSE) to impose the effects of different answers on visual states. First, we propose an Answer-Driven Focusing Attention (ADFA) to capture the answer-driven effect on visual attention by sharpening question-related attention and adjusting it by answer-based logical operation at each turn. Then based on the focusing attention, we get the visual state estimation by Conditional Visual Information Fusion (CVIF), where overall information and difference information are fused conditioning on the question-answer state. We evaluate the proposed ADVSE to both question generator and guesser tasks on the large-scale GuessWhat?! dataset and achieve the state-of-the-art performances on both tasks. The qualitative results indicate that the ADVSE boosts the agent to generate highly efficient questions and obtains reliable visual attentions during the reasonable question generation and guess processes. Zipeng Xu, Fangxiang Feng, Xiaojie Wang 0006, Yushu Yang, Huixing Jiang, Zhongyuan Wang 0006 |
ACM Multimedia | 2 |
| 2020 | Referring Expression Generation via Visual Dialogue
Lingxuan Li, Tianrui Niu, Fangxiang Feng, Xiaojie Wang 0006 |
NLPCC (2) | 5 |
| 2020 | Dual-CNN: A Convolutional language decoder for paragraph image captioning
Ruifan Li, Yihui Shi, Fangxiang Feng, Xiaojie Wang 0006 |
Neurocomputing | 4 |
| 2020 | Exploring Global and Local Linguistic Representations for Text-to-Image SynthesisabstractThe task of text-to-image synthesis is to generate photographic images conditioned on given textual descriptions. This challenging task has recently attracted considerable attention from the multimedia community due to its potential applications. Most of the up-to-date approaches are built based on generative adversarial network (GAN) models, and they synthesize images conditioned on the global linguistic representation. However, the sparsity of the global representation results in training difficulties on GANs and a shortage of fine-grained information in the generated images. To address this problem, we propose cross-modal global and local linguistic representations-based generative adversarial networks (CGL-GAN) by incorporating the local linguistic representation into the GAN. In our CGL-GAN, we construct a generator to synthesize the target images and a discriminator to judge whether the generated images conform with the text description. In the discriminator, we construct the cross-modal correlation by projecting the image representations at high and low levels onto the global and local linguistic representations, respectively. We design the hinge loss function to train our CGL-GAN model. We evaluate the proposed CGL-GAN on two publicly available datasets, the CUB and the MS-COCO. The extensive experiments demonstrate that incorporating fine-grained local linguistic information with cross-modal correlation can greatly improve the performance of text-to-image synthesis, even when generating high-resolution images. Ruifan Li, Fangxiang Feng, Guangwei Zhang 0003, Xiaojie Wang 0006 |
IEEE Trans. Multim. | 3 |
| 2019 | A survey on freehand sketch recognition and retrieval
Xianlin Zhang, Xueming Li 0002, Yang Liu 0132, Fangxiang Feng |
Image Vis. Comput. | 4 |
| 2015 | Deep correspondence restricted Boltzmann machine for cross-modal retrieval
Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006 |
Neurocomputing | 1 |
| 2015 | Challenges in representation learning: A report on three machine learning contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio |
Neural Networks | 13 |
| 2015 | Correspondence Autoencoders for Cross-Modal RetrievalabstractThis article considers the problem of cross-modal retrieval, such as using a text query to search for images and vice-versa. Based on different autoencoders, several novel models are proposed here for solving this problem. These models are constructed by correlating hidden representations of a pair of autoencoders. A novel optimal objective, which minimizes a linear combination of the representation learning errors for each modality and the correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimizing the correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimizing the representation learning error makes hidden representations good enough to reconstruct inputs of each modality. To balance the two kind of errors induced by representation learning and correlation learning, we set a specific parameter in our models. Furthermore, according to the modalities the models attempt to reconstruct they are divided into two groups. One group including three models is named multimodal reconstruction correspondence autoencoder since it reconstructs both modalities. The other group including two models is named unimodal reconstruction correspondence autoencoder since it reconstructs a single modality. The proposed models are evaluated on three publicly available datasets. And our experiments demonstrate that our proposed correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multimodal deep models on cross-modal retrieval tasks. Fangxiang Feng, Xiaojie Wang 0006, Ruifan Li |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2014 | Cross-modal Retrieval with Correspondence AutoencoderabstractThe problem of cross-modal retrieval, e.g., using a text query to search for images and vice-versa, is considered in this paper. A novel model involving correspondence autoencoder (Corr-AE) is proposed here for solving this problem. The model is constructed by correlating hidden representations of two uni-modal autoencoders. A novel optimal objective, which minimizes a linear combination of representation learning errors for each modality and correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimization of correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimization of representation learning error makes hidden representations are good enough to reconstruct input of each modality. A parameter $\alpha$ is used to balance the representation learning error and the correlation learning error. Based on two different multi-modal autoencoders, Corr-AE is extended to other two correspondence models, here we called Corr-Cross-AE and Corr-Full-AE. The proposed models are evaluated on three publicly available data sets from real scenes. We demonstrate that the three correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multi-modal deep models on cross-modal retrieval tasks. Fangxiang Feng, Xiaojie Wang 0006, Ruifan Li |
ACM Multimedia | 1 |
| 2013 | Challenges in Representation Learning: A Report on Three Machine Learning Contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio |
ICONIP (3) | 13 |