Xiaojie Wang 0006

dblp:99/7033-6 · DBLP profile ↗
← Back
98ranked-venue papers
2as first author
49since 2021 · last 2026
0000-0003-0314-8951ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 60 · 2 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 30 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Computer networks · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 A fine-grained entity understanding network for weakly supervised phrase grounding
Pengyue Lin, Ruifan Li, Fangxiang Feng, Lun Ke, Zhanyu Ma, Xiaojie Wang 0006
Pattern Recognit.6
2025 A Systematic Exploration of Knowledge Graph Alignment with Large Language Models in Retrieval Augmented Generation
abstract
Retrieval Augmented Generation (RAG) with Knowledge Graphs (KGs) is an effective way to enhance Large Language Models (LLMs). Due to the natural discrepancy between structured KGs and sequential LLMs, KGs must be linearized to text before being inputted into LLMs, leading to the problem of KG Alignment with LLMs (KGA). However, recent KG+RAG methods only consider KGA as a simple step without comprehensive and in-depth explorations, leaving three essential problems unclear: (1) What are the factors and their effects in KGA? (2) How do LLMs understand KGs? (3) How to improve KG+RAG by KGA? To fill this gap, we conduct systematic explorations on KGA, where we first define the problem of KGA and subdivide it into the graph transformation phase (graph-to-graph) and the linearization phase (graph-to-text). In the graph transformation phase, we study graph features at the node, edge, and full graph levels from low to high granularity. In the linearization phase, we study factors on formats, orders, and templates from structural to token levels. We conduct substantial experiments on 15 typical LLMs and three common datasets. Our main findings include: (1) The centrality of the KG affects the final generation; formats have the greatest impact on KGA; orders are model-dependent, without an optimal order adapting for all models; the templates with special token separators are better. (2) LLMs understand KGs by a unique mechanism, different from processing natural sentences, and separators play an important role. (3) We achieved 7.3% average performance improvements on four common LLMs on the KGQA task by combining the optimal factors to enhance KGA.
Shiyu Tian, Shuyue Xing, Xingrui Li, Yangyang Luo, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006
AAAI8
2025 Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis
abstract
The customization of text-to-image models has seen significant advancements, yet generating multiple personalized concepts remains a challenging task. Current methods struggle with attribute leakage and layout confusion when handling multiple concepts, leading to reduced concept fidelity and semantic consistency. In this work, we introduce a novel training-free framework, Concept Conductor, designed to ensure visual fidelity and correct layout in multi-concept customization. Concept Conductor isolates the sampling processes of multiple customized models to prevent attribute leakage between different concepts and corrects erroneous layouts through self-attention-based spatial guidance. Additionally, we present a concept injection technique that employs shape-aware masks to specify the generation area for each concept. This technique injects the structure and appearance of personalized concepts through feature fusion in the attention layers, ensuring harmony in the final image. Extensive qualitative and quantitative experiments demonstrate that Concept Conductor can consistently generate composite images with accurate layouts while preserving the visual details of each concept. Compared to existing baselines, Concept Conductor shows significant performance improvements. Our method supports the combination of any number of concepts and maintains high fidelity even when dealing with visually similar concepts. The code and trained models will be made publicly available.
Zebin Yao, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
AAAI4
2025 Multimodal Aspect-Based Sentiment Analysis under Conditional Relation
abstract
Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to extract aspect terms from text-image pairs and identify their sentiments. Previous methods are based on the premise that the image contains the objects referred by the aspects within the text. However, this condition cannot always be met, resulting in a suboptimal performance. In this paper, we propose COnditional Relation based Sentiment Analysis framework (CORSA). Specifically, we design a conditional relation detector (CRD) to mitigate the impact of the unmet conditional image. Moreover, we design a visual object localizer (VOL) to locate the exact condition-related visual regions associated with the aspects. With CRD and VOL, our CORSA framework takes a multi-task form. In addition, to effectively learn CORSA we conduct two types of annotations. One is the conditional relation using a pretrained referring expression comprehension model; the other is the bounding boxes of visual objects by a pretrained object detection model. Experiments on our built C-MABSA dataset show that CORSA consistently outperforms existing methods. The code and data are available at https://github.com/Liuxj-Anya/CORSA.
Xinjing Liu, Ruifan Li, Shuqin Ye, Guangwei Zhang 0003, Xiaojie Wang 0006
COLING5
2025 A Weighted Cross-entropy Loss for Mitigating LLM Hallucinations in Cross-lingual Continual Pretraining
abstract
Recently, due to the explosive advances of large language models (LLMs) on English, cross-lingual continual pretraining has been widely applied in obtaining Chinese LLMs. However, previous studies showed that these LLMs have suffered severe hallucinations, mainly caused by noisy tokens. To this aim, we propose a novel loss function, InfoLoss for continual pretraining. Specifically, our loss function takes into account the co-occurrence of noisy and normal tokens, and uses point-wise mutual information to reduce the impact of noisy tokens. We use InfoLoss to continually pretrain 30 billion tokens on Llama 2-7B with 64 A100 GPUs for 24 days, obtaining C-Llama. We then conduct experiments on 12 benchmarks for evaluations. The results show the effectiveness of our proposed InfoLoss. Our datasets and codes are publicly available at https://github.com/Fluxation996/C-Llama.
Yuantao Fan, Ruifan Li, Guangwei Zhang 0003, Xiaojie Wang 0006
ICASSP5
2025 VL-DynaRefine: A Vision-Language Dynamic Refinement Approach for Visual Reasoning
abstract
Visual reasoning is a key capability that significantly impacts the performance of multimodal tasks, such as compositional visual question answering and visual grounding. These tasks often require complex, multi-step reasoning processes. In recent years, several training-free methods for Vision-Language Models (VLMs) have emerged, with visual programming methods being proposed to enhance the capability of VLMs in visual reasoning tasks. While these methods have made some progress, they still face two primary challenges due to the lack of verification and refinement mechanisms for each action's output during the reasoning process: error accumulation and feedback delay, as well as insufficient utilization of multimodal contextual information. To address these challenges, we propose VL-DynaRefine, a training-free approach consisting of three modules: a planner, a verifier, and a refiner. The planner generates programmatic actions to solve the problem and executes each action in sequence, which is inspected by a verifier that reassesses the actions via confidence scores and determines whether refinement is necessary based on the evaluation results. In the refiner module, we incorporate a context-aware local refinement mechanism and a global refinement mechanism based on visual and action trajectories to reduce the impact of reasoning errors on the outcome. We evaluate our approach on multiple visual reasoning datasets, and the experimental results show that our method outperforms existing visual programming methods in both reasoning accuracy and efficiency, further validating its effectiveness in visual reasoning tasks.
Zeyuan Zang, Fangxiang Feng, Caixia Yuan, Huixing Jiang, Xiaojie Wang 0006
ACM Multimedia9
2025 Multi-task Contrastive Learning Enhanced Instruction Tuning for Dialog Understanding
Zimeng Bai, Xiying Zhao, Zhuoxin Han, Lujie Niu, Caixia Yuan, Xiaojie Wang 0006
NLPCC (3)7
2025 Exploiting Prior Tacit Knowledge to Enhance Alignment and Verification in zero-shot video grounding
Jing Wang 0169, Xianbing Zhao, Xiaojie Wang 0006, Fangxiang Feng
Neurocomputing3
2024 Visual Prompt Tuning for Weakly Supervised Phrase Grounding
abstract
Previous works on the task of weakly supervised phrase grounding (WSG) rely heavily on object detectors providing RoIs for the localization. However, such methods cannot be applied effectively to real-world scenarios largely because that the detectors are trained with limited categories. In this paper, we propose a refinement-based approach to WSG through fine-tuning a detector-free phrase grounding model with a visual prompt. This visual prompt is extracted from the text-related representations in CLIP. Furthermore, we combine the visual prompt with learnable features and then fine-tune the grounding network. Our experimental results significantly outperform state-of-the-art methods on the WSG task and shows the effectiveness of our method.
Pengyue Lin, Zhihan Yu, Mingcong Lu, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
ICASSP6
2024 DiffHarmony: Latent Diffusion Model Meets Image Harmonization
abstract
Image harmonization, which involves adjusting the foreground of a composite image to attain a unified visual consistency with the background, can be conceptualized as an image-to-image translation task. Diffusion models have recently promoted the rapid development of image-to-image translation tasks . However, training diffusion models from scratch is computationally intensive. Fine-tuning pre-trained latent diffusion models entails dealing with the reconstruction error induced by the image compression autoencoder, making it unsuitable for image generation tasks that involve pixel-level evaluation metrics. To deal with these issues, in this paper, we first adapt a pre-trained latent diffusion model to the image harmonization task to generate the harmonious but potentially blurry initial images. Then we implement two strategies: utilizing higher-resolution images during inference and incorporating an additional refinement stage, to further enhance the clarity of the initially harmonized images. Extensive experiments on iHarmony4 datasets demonstrate the superiority of our proposed method. The code is available at \hrefhttps://github.com/nicecv/DiffHarmony https://github.com/nicecv/DiffHarmony.
Fangxiang Feng, Xiaojie Wang 0006
ICMR3
2024 Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak Supervision
abstract
Phrase Grounding, i.e., PG aims to locate objects referred by noun phrases. Recently, PG under weak supervision (i.e., grounding without region-level annotations) and zero-shot PG (i.e., grounding from seen categories to unseen ones) are proposed, respectively. However, for real-world applications these two approaches are limited due to slight annotations and numerable categories during training. In this paper, we propose a framework of zero-shot PG under weak supervision. Specifically, our PG framework is built on triple alignment strategies. Firstly, we propose a region-text alignment (RTA) strategy to build region-level attribute associations via CLIP. Secondly, we propose a domain alignment (DomA) strategy by minimizing the difference between distributions of seen classes in the training and those of the pre-training. Thirdly, we propose a category alignment (CatA) strategy by considering both category semantics and region-category relations. Extensive experimental results show that our proposed PG framework outperforms previous zero-shot methods and weakly-supervised methods. Our code is available at https://github.com/LinPengyue/ZS-WSG.
Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006
ACM Multimedia7
2024 Q-MoE: Connector for MLLMs with Text-Driven Routing
abstract
Multimodal Large Language Models (MLLMs) have showcased remarkable advances in handling various vision-language tasks. These models typically consist of a Large Language Model (LLM), a vision encoder and a connector structure, which is used to bridge the modality gap between vision and language. It is challenging for the connector to filter the right visual information for LLM according to the task in hand. Most of previous connectors, such as light-weight projection and Q-former, treat visual information for diverse tasks uniformly, therefore lacking task-specific visual information extraction capabilities. To address the issue, this paper proposes Q-MoE, a query-based connector with Mixture-of-Experts (MoE) to extract task-specific information with text-driven routing. Furthermore, an optimal path based training strategy is proposed to find an optimal expert combination. Extensive experiments on two popular open-source LLMs and several different visual-language tasks demonstrate the effectiveness of the Q-MoE connecter.
Hanzi Wang, Jiamin Ren, Huixing Jiang, Fangxiang Feng, Xiaojie Wang 0006
ACM Multimedia8
2024 DiffHarmony++: Enhancing Image Harmonization with Harmony-VAE and Inverse Harmonization Model
abstract
Latent diffusion model has demonstrated impressive efficacy in image generation and editing tasks. Recently, it has also promoted the advancement of image harmonization. However, methods involving latent diffusion model all face a common challenge: the severe image distortion introduced by the VAE component, while image harmonization is a low-level image processing task that relies on pixel-level evaluation metrics. In this paper, we propose Harmony-VAE, leveraging the input of the harmonization task itself to enhance the quality of decoded images. The input involving composite image contains the precise pixel level information, which can complement the correct foreground appearance and color information contained in denoised latents. Meanwhile, the inherent generative nature of diffusion models makes it naturally adapt to inverse image harmonization, i.e. generating synthetic composite images based on real images and foreground masks. We train an inverse harmonization diffusion model to perform data augmentation on two subsets of iHarmony4 and construct a new human harmonization dataset with prominent foreground objects. Extensive experiments demonstrate the effectiveness of our proposed Harmony-VAE and inverse harmonization model. Code and pretrained models are available at https://github.com/nicecv/DiffHarmony.
Fangxiang Feng, Guang Liu 0006, Ruifan Li, Xiaojie Wang 0006
ACM Multimedia5
2024 Improving Causal Inference of Large Language Models with SCM Tools
Zhenyang Hua, Shuyue Xing, Huixing Jiang, Xiaojie Wang 0006
NLPCC (3)5
2024 Enhancing Document Information Selection Through Multi-Granularity Responses for Dialogue Generation
abstract
Abstract Document information selection is an essential part of document-grounded dialogue tasks, and more accurate information selection results can provide more appropriate dialogue responses. Existing works have achieved excellent results by employing multi-granularity of dialogue history information, indicating the effectiveness of multi-level historical information. However, these works often focus on exploring the hierarchical information of dialogue history, while neglecting the multi-granularity utilization in response, important information that holds an impact on the decoding process. Therefore, this paper proposes a model for document information selection based on multi-granularity responses. By integrating the document selection results at the response word level and semantic unit level, the model enhances its capability in knowledge selection and produces better responses. For the division at the semantic unit level of the response, we propose two semantic unit division methods, static and dynamic. Experiments on two public datasets show that our models combining static or dynamic semantic unit levels significantly outperform baseline models.
Kangyu Qiao, Shuyue Xing, Caixia Yuan, Xiaojie Wang 0006
Neural Process. Lett.5
2024 LGR-NET: Language Guided Reasoning Network for Referring Expression Comprehension
abstract
Referring Expression Comprehension(REC) is a fundamental task in the vision and language domain, which aims to locate an image region according to a natural language expression. REC requires the models to capture key clues in the text and perform accurate cross-modal reasoning. A recent trend employs transformer-based methods to address this problem. However, most of these methods typically treat image and text equally. They usually perform cross-modal reasoning in a crude way, and utilize textual features as a whole without detailed considerations (e.g., spatial information). This insufficient utilization of textual features will lead to sub-optimal results. In this paper, we propose aLanguage Guided Reasoning Network(LGR-NET) to fully utilize the guidance of the referring expression. To localize the referred object, we set a prediction token to capture cross-modal features. Furthermore, to sufficiently utilize the textual features, we extend them by our Textual Feature Extender (TFE) from three aspects.First, we design a novel coordinate embedding based on textual features. The coordinate embedding is incorporated to the prediction token to promote its capture of language-related visual features.Second, we employ the extracted textual features for Text-guided Cross-modal Alignment (TCA) and Fusion (TCF), alternately.Third, we devise a novel cross-modal loss to enhance cross-modal alignment between the referring expression and the learnable prediction token. We conduct extensive experiments on five benchmark datasets, and the experimental results show that our LGR-NET achieves a new state-of-the-art. Source code is available at https://github.com/lmc8133/LGR-NET.
Mingcong Lu, Ruifan Li, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006
IEEE Trans. Circuits Syst. Video Technol.5
2024 DualGCN: Exploring Syntactic and Semantic Information for Aspect-Based Sentiment Analysis
abstract
The task of aspect-based sentiment analysis aims to identify sentiment polarities of given aspects in a sentence. Recent advances have demonstrated the advantage of incorporating the syntactic dependency structure with graph convolutional networks (GCNs). However, their performance of these GCN-based methods largely depends on the dependency parsers, which would produce diverse parsing results for a sentence. In this article, we propose a dual GCN (DualGCN) that jointly considers the syntax structures and semantic correlations. Our DualGCN model mainly comprises four modules: 1) SynGCN: instead of explicitly encoding syntactic structure, the SynGCN module uses the dependency probability matrix as a graph structure to implicitly integrate the syntactic information; 2) SemGCN: we design the SemGCN module with multihead attention to enhance the performance of the syntactic structure with the semantic information; 3) Regularizers: we propose orthogonal and differential regularizers to precisely capture semantic correlations between words by constraining attention scores in the SemGCN module; and 4) Mutual BiAffine: we use the BiAffine module to bridge relevant information between the SynGCN and SemGCN modules. Extensive experiments are conducted compared with up-to-date pretrained language encoders on two groups of datasets, one including Restaurant14, Laptop14, and Twitter and the other including Restaurant15 and Restaurant16. The experimental results demonstrate that the parsing results of various dependency parsers affect their performance of the GCN-based models. Our DualGCN model achieves superior performance compared with the state-of-the-art approaches. The source code and preprocessed datasets are provided and publicly available on GitHub (see https://github.com/CCChenhao997/DualGCN-ABSA).
Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy
IEEE Trans. Neural Networks Learn. Syst.5
2023 SPRING: Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph
abstract
Existing multimodal conversation agents have shown impressive abilities to locate absolute positions or retrieve attributes in simple scenarios, but they fail to perform well when complex relative positions and information alignments are involved, which poses a bottleneck in response quality. In this paper, we propose a Situated Conversation Agent Pretrained with Multimodal Questions from Incremental Layout Graph (SPRING) with abilities of reasoning multi-hops spatial relations and connecting them with visual attributes in crowded situated scenarios. Specifically, we design two types of Multimodal Question Answering (MQA) tasks to pretrain the agent. All QA pairs utilized during pretraining are generated from novel Increment Layout Graphs (ILG). QA pair difficulty labels automatically annotated by ILG are used to promote MQA-based Curriculum Learning. Experimental results verify the SPRING's effectiveness, showing that it significantly outperforms state-of-the-art approaches on both SIMMC 1.0 and SIMMC 2.0 datasets. We release our code and data at https://github.com/LYX0501/SPRING.
Yuxing Long, Binyuan Hui, Fulong Ye, Yanyang Li, Zhuoxin Han, Caixia Yuan, Yongbin Li 0001, Xiaojie Wang 0006
AAAI8
2023 USSA: A Unified Table Filling Scheme for Structured Sentiment Analysis
abstract
Most previous studies on Structured Sentiment Analysis (SSA) have cast it as a problem of bi-lexical dependency parsing, which cannot address issues of overlap and discontinuity simultaneously.In this paper, we propose a nichetargeting and effective solution.Our approach involves creating a novel bi-lexical dependency parsing graph, which is then converted to a unified 2D table-filling scheme, namely USSA.The proposed scheme resolves the kernel bottleneck of previous SSA methods by utilizing 13 different types of relations.In addition, to closely collaborate with the USSA scheme, we have developed a model that includes a proposed bi-axial attention module to effectively capture the correlations among relations in the rows and columns of the table.Extensive experimental results on benchmark datasets demonstrate the effectiveness and robustness of our proposed framework, outperforming state-ofthe-art methods consistently 1 .
Zepeng Zhai, Hao Chen 0041, Ruifan Li, Xiaojie Wang 0006
ACL (1)4
2023 FATRER: Full-Attention Topic Regularizer for Accurate and Robust Conversational Emotion Recognition
abstract
This paper concentrates on the understanding of interlocutors’ emotions evoked in conversational utterances. Previous studies in this literature mainly focus on more accurate emotional predictions, while ignoring model robustness when the local context is corrupted by adversarial attacks. To maintain robustness while ensuring accuracy, we propose an emotion recognizer augmented by a full-attention topic regularizer, which enables an emotion-related global view when modeling the local context in a conversation. A joint topic modeling strategy is introduced to implement regularization from both representation and loss perspectives. To avoid over-regularization, we drop the constraints on prior distributions that exist in traditional topic modeling and perform probabilistic approximations based entirely on attention alignment. Experiments show that our models obtain more favorable results than state-of-the-art models, and gain convincing robustness under three types of adversarial attacks. Code: https://github.com/ludybupt/FATRER.
Yuzhao Mao, Xiaojie Wang 0006
ECAI4
2023 Joint Modeling for ASR Correction and Dialog State Tracking
abstract
In spoken dialog system, transcription errors in Automated Speech Recognition (ASR) impact downstream task, especially dialog state tracking (DST). Approaches to alleviate such errors involve using richer information such as word-lattices and word confusion networks. However, in some cases, this information may not be easily obtained. In addition, the large pre-trained language model is trained on plain text, leading to the gap between spoken DST and original pretrained model. In this paper, we propose a multi-task method which performs DST jointly with ASR correction to improve the performance of both tasks. To do so, we build a MultiWOZ-ASR dataset containing ASR noise in DST and mitigate the gap by utilizing a multi-task pre-training framework. Moreover, curriculum learning is adopted to alleviate the phenomenon that the correction task is difficult to converge at the initial stage of pre-training. Experimental results show that our model achieves significant improvements on DSTC2 and MultiWOZ-ASR dataset.
Deyuan Wang, Caixia Yuan, Xiaojie Wang 0006
ICASSP4
2023 An Asynchronous Updating Reinforcement Learning Framework for Task-Oriented Dialog System
abstract
Reinforcement learning has been applied to train the dialog systems in many works. Previous approaches divide the dialog system into multiple modules including DST (dialog state tracking) and DP (dialog policy), and train these modules simultaneously. However, different modules influence each other during training. The errors from DST might misguide the dialog policy, and the system action brings extra difficulties for the DST module. To alleviate this problem, we propose Asynchronous Updating Reinforcement Learning framework (AURL) that updates the DST module and the DP module asynchronously under a cooperative setting. Furthermore, curriculum learning is implemented to address the problem of unbalanced data distribution during reinforcement learning sampling, and multiple user models are introduced to increase the dialog diversity. Results on the public SSD-PHONE dataset show that our method achieves a compelling result with a 31.37% improvement on the dialog success rate. The code is publicly available via https://github.com/shunjiu/AURL.
Xiaojie Wang 0006, Caixia Yuan
ICASSP3
2023 Whether you can locate or not? Interactive Referring Expression Generation
abstract
Referring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object. Existing methods construct REG models independently by using only the REs as ground truth for model training, without considering the potential interaction between REG and REC models. In this paper, we propose an Interactive REG (IREG) model that can interact with a real REC model, utilizing signals indicating whether the object is located and the visual region located by the REC model to gradually modify REs. Our experimental results on three RE benchmark datasets, RefCOCO, RefCOCO+, and RefCOCOg show that IREG outperforms previous state-of-the-art methods on popular evaluation metrics. Furthermore, a human evaluation shows that IREG generates better REs with the capability of interaction.
Fulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie Wang 0006
ACM Multimedia4
2023 A Task-Oriented Dialog Model with Task-Progressive and Policy-Aware Pre-training
Lucen Zhong, Hengtong Lu, Caixia Yuan, Xiaojie Wang 0006, Jiashen Sun, Guanglu Wan
NLPCC (1)4
2023 Hierarchical history based information selection for document grounded dialogue generation
Shiyu Tian, Ziwei Bai, Caixia Yuan, Xiaojie Wang 0006
Appl. Intell.5
2023 Dual-Lens HDR using Guided 3D Exposure CNN and Guided Denoising Transformer
abstract
We study the high dynamic range (HDR) imaging problem in dual-lens systems. Existing methods usually treat the HDR imaging problem as an image fusion problem and the HDR result is estimated by fusing the aligned short exposure image and long exposure image. However, the image fusion pipeline depends highly on the image alignment, which is difficult to be perfect. We propose to transfer the dual-lens HDR imaging problem into the disentangled enhancement of exposure correction and denoising for the short exposure image, guided by the long exposure image. In the guided exposure correction module, we make use of the guidance image and 3D color transformation to propose a guided 3D exposure CNN (GEC) to get the rough HDR result from the short exposure image. Then, in the guided denoising module, we make use of the cross-attention mechanism to propose a guided denoising transformer (GDT) to directly use the long exposure image as guidance to denoise the rough HDR result in a pyramid way. And in both modules, we bypass the difficult image alignment processing. Experimental results demonstrate the superiority of our method over the state-of-the-art ones.
Weixin Li 0001, Chang Liu 0071, Xue Tian, Ya Li 0001, Xiaojie Wang 0006, Xuan Dong 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2023 CUR Transformer: A Convolutional Unbiased Regional Transformer for Image Denoising
abstract
Image denoising is a fundamental problem in computer vision and multimedia computation. Non-local filters are effective for image denoising. But existing deep learning methods that use non-local computation structures are mostly designed for high-level tasks, and global self-attention is usually adopted. For the task of image denoising, they have high computational complexity and have a lot of redundant computation of uncorrelated pixels. To solve this problem and combine the marvelous advantages of non-local filter and deep learning, we propose a Convolutional Unbiased Regional (CUR) transformer. Based on the prior that, for each pixel, its similar pixels are usually spatially close, our insights are that (1) we partition the image into non-overlapped windows and perform regional self-attention to reduce the search range of each pixel, and (2) we encourage pixels across different windows to communicate with each other. Based on our insights, the CUR transformer is cascaded by a series of convolutional regional self-attention (CRSA) blocks with U-style short connections. In each CRSA block, we use convolutional layers to extract the query, key, and value features, namely Q , K , and V , of the input feature. Then, we partition the Q , K , and V features into local non-overlapped windows and perform regional self-attention within each window to obtain the output feature of this CRSA block. Among different CRSA blocks, we perform the unbiased window partition by changing the partition positions of the windows. Experimental results show that the CUR transformer outperforms the state-of-the-art methods significantly on four low-level vision tasks, including real and synthetic image denoising, JPEG compression artifact reduction, and low-light image enhancement.
Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Xuan Dong 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2023 AABLSTM: A Novel Multi-task Based CNN-RNN Deep Model for Fashion Analysis
abstract
With the rapid growth of online commerce and fashion-related applications, visual clothing analysis and recognition has become a hotspot in computer vision. In this paper, we propose a novel AABLSTM network, which is based on deep CNN-RNN, to solve the visual fashion analysis of clothing category classification, attribute detection, and landmark localization. The designed fashion model is leveraged with the multi-task driven mechanism as follows: firstly, a bidirectional LSTM (Bi-LSTM) branch is proposed for efficiently mining the semantic association between related attributes so as to improve the precision of clothing category classification and attribute detection; then, an imitated hourglass sub-network of “down-up sampling” is constructed for boosting the accuracy of fashion landmark localization; and finally, a specially designed multi-loss function is constructed to better optimize the network training. Extensive experimental results on large-scale fashion datasets demonstrate the superior performance of our approach.
Xianlin Zhang, Mengling Shen, Xueming Li 0002, Xiaojie Wang 0006
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Enhanced Multi-Channel Graph Convolutional Network for Aspect Sentiment Triplet Extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE) is an emerging sentiment analysis task.Most of the existing studies focus on devising a new tagging scheme that enables the model to extract the sentiment triplets in an end-to-end fashion.However, these methods ignore the relations between words for ASTE task.In this paper, we propose an Enhanced Multi-Channel Graph Convolutional Network model (EMC-GCN) to fully utilize the relations between words.Specifically, we first define ten types of relations for ASTE task, and then adopt a biaffine attention module to embed these relations as an adjacent tensor between words in a sentence.After that, our EMC-GCN transforms the sentence into a multi-channel graph by treating words and the relation adjacent tensor as nodes and edges, respectively.Thus, relationaware node representations can be learnt.Furthermore, we consider diverse linguistic features to enhance our EMC-GCN model.Finally, we design an effective refining strategy on EMC-GCN for word-pair representation refinement, which considers the implicit results of aspect and opinion extraction when determining whether word pairs match or not.Extensive experimental results on the benchmark datasets demonstrate that the effectiveness and robustness of our proposed model, which outperforms state-of-the-art methods significantly.
Hao Chen 0041, Zepeng Zhai, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
ACL (1)5
2022 Learn to Adapt for Generalized Zero-Shot Text Classification
abstract
Generalized zero-shot text classification aims to classify textual instances from both previously seen classes and incrementally emerging unseen classes.Most existing methods generalize poorly since the learned parameters are only optimal for seen classes rather than for both classes, and the parameters keep stationary in predicting procedures.To address these challenges, we propose a novel Learn to Adapt (LTA) network using a variant meta-learning framework.Specifically, LTA trains an adaptive classifier by using both seen and virtual unseen classes to simulate a generalized zero-shot learning (GZSL) scenario in accordance with the test time, and simultaneously learns to calibrate the class prototypes and sample representations to make the learned parameters adaptive to incoming unseen classes.We claim that the proposed model is capable of representing all prototypes and samples from both classes to a more consistent distribution in a global space.Extensive experiments on five text classification datasets show that our model outperforms several competitive previous approaches by large margins.
Caixia Yuan, Xiaojie Wang 0006, Ziwei Bai
ACL (1)3
2022 A Simple Model for Distantly Supervised Relation Extraction
abstract
Distantly supervised relation extraction is challenging due to the noise within data. Recent methods focus on exploiting bag representations based on deep neural networks with complex de-noising scheme to achieve remarkable performance. In this paper, we propose a simple but effective BERT-based Graph convolutional network Model (i.e., BGM). Our BGM comprises of an instance embedding module and a bag representation module. The instance embedding module uses a BERT-based pretrained language model to extract key information from each instance. The bag representaion module constructs the corresponding bag graph then apply a convolutional operation to obtain the bag representation. Our BGM model achieves a considerable improvement on two benchmark datasets, i.e., NYT10 and GDS.
Ziqin Rao, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
COLING4
2022 COM-MRC: A COntext-Masked Machine Reading Comprehension Framework for Aspect Sentiment Triplet Extraction
abstract
Aspect Sentiment Triplet Extraction (ASTE) aims to extract sentiment triplets from sentences, which was recently formalized as an effective machine reading comprehension (MRC) based framework.However, when facing multiple aspect terms, the MRC-based methods could fail due to the interference from other aspect terms.In this paper, we propose a novel COntext-Masked MRC (COM-MRC) framework for ASTE.Our COM-MRC framework comprises three closely-related components: a context augmentation strategy, a discriminative model, and an inference method.Specifically, a context augmentation strategy is designed by enumerating all masked contexts for each aspect term.The discriminative model comprises four modules, i.e., aspect and opinion extraction modules, sentiment classification and aspect detection modules.In addition, a two-stage inference method first extracts all aspects and then identifies their opinions and sentiment through iteratively masking the aspects.Extensive experimental results on benchmark datasets show the effectiveness of our proposed COM-MRC framework, which outperforms state-of-the-art methods consistently 1 .
Zepeng Zhai, Hao Chen 0041, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
EMNLP5
2022 Towards Unifying Reference Expression Generation and Comprehension
abstract
Reference Expression Generation (REG) and Comprehension (REC) are two highly correlated tasks. Modeling REG and REC simultaneously for utilizing the relation between them is a promising way to improve both. However, the problem of distinct inputs, as well as building connections between them in a single model, brings challenges to the design and training of the joint model. To address the problems, we propose a unified model for REG and REC, named UniRef. It unifies these two tasks with the carefully-designed Image-Region-Text Fusion layer (IRTF), which fuses the image, region and text via the image cross-attention and region cross-attention. Additionally, IRTF could generate pseudo input regions for the REC task to enable a uniform way for sharing the identical representation space across the REC and REG. We further propose Vision-conditioned Masked Language Modeling (VMLM) and Text-Conditioned Region Prediction (TRP) to pre-train UniRef model on multi-granular corpora. The VMLM and TRP are directly related to REG and REC, respectively, but could help each other. We conduct extensive experiments on three benchmark datasets, RefCOCO, RefCOCO+ and RefCOCOg. Experimental results show that our model outperforms previous state-of-the-art methods on both REG and REC.
Duo Zheng, Tao Kong, Ya Jing, Jiaan Wang, Xiaojie Wang 0006
EMNLP5
2022 Improving Image Paragraph Captioning with Dual Relations
abstract
Image paragraph captioning aims to generate multiple de-scriptive sentences for an image. However, most previous methods ignore the explicit relations among objects resulting in unsatisfactory performance. In this paper, we propose a novel model (i.e., DualRel) to capture spatial and seman-tic relations among objects. Specifically, the spatial relation embedding is obtained solely from images using a predefined geometry pattern. With the help of captions, the semantic relation embedding is learned in a weakly supervised man-ner. These two relation embeddings are then interacted with regional features of objects through a relation-aware attention interaction. It first obtains a visual context vector using regional features. Then with the visual context vector, we obtain the corresponding spatial and semantic relation-aware vectors using attentions. These three vectors are fused with two gates for language decoding to further generate a para-graph. Experimental results on Stanford benchmark dataset show that DualRel achieves remarkable improvements11Code released at https://github.com/fuyunll07/DualRel.
Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006
ICME6
2022 Question-Driven Graph Fusion Network for Visual Question Answering
abstract
Existing Visual Question Answering (VQA) models have ex-plored various visual relationships between objects in the im-age to answer complex questions, which inevitably introduces irrelevant information brought by inaccurate object detection and text grounding. To address the problem, we propose a Question-Driven Graph Fusion Network (QD-GFN). It first models semantic, spatial, and implicit visual relations in images by three graph attention networks, then question in-formation is utilized to guide the aggregation process of the three graphs, further, our QD-GFN adopts an object filtering mechanism to remove question-irrelevant objects contained in the image. Experiment results demonstrate that our QD-GFN outperforms the prior state-of-the-art on both VQA 2.0 and VQA-CP v2 datasets. Further analysis shows that both the novel graph aggregation method and object filtering mecha-nism play a significant role in improving the performance of the model.
Yuxi Qian, Yuncong Hu, Fangxiang Feng, Xiaojie Wang 0006
ICME5
2022 GR-GAN: Gradual Refinement Text-To-Image Generation
abstract
A good Text-to-Image model should not only generate high quality images, but also ensure the consistency between the text and the generated image. Previous models failed to simultaneously fix both sides well. This paper proposes a Gradual Refinement Generative Adversarial Network (GR-GAN) to alleviates the problem efficiently. A GRG module is designed to generate images from low resolution to high resolution with the corresponding text constraints from coarse granularity (sentence) to fine granularity (word) stage by stage, a ITM module is designed to provide image-text matching losses at both sentence-image level and word-region level for corresponding stages. We also introduce a new metric Cross-Model Distance (CMD) for simultaneously evaluating image quality and image-text consistency. Experimental results show GR-GAN significant outperform previous models, and achieve new state-of-the-art on both FID and CMD. A detailed analysis demonstrates the efficiency of different generation stages in GR-GAN.
Fangxiang Feng, Xiaojie Wang 0006
ICME3
2022 A Region-based Document VQA
abstract
Practical Document Visual Question Answering (DocVQA) needs not only to recognize and extract the document contents, but also reason on them for answering questions. However, previous DocVQA data mainly focuses on in-line questions, where the answers could be directly extracted after locating keywords in the documents, which needs less reasoning. This paper therefore builds a large-scale dataset named Region-based Document VQA (RDVQA), which includes more practical questions for DocVQA. We then propose a novel Reason-over-In-region-Question-answering (ReIQ) model for addressing the problems. It is a pre-training-based model, where a Spatial-Token Pre-trained Model (STPM) is employed as the backbone. Two novel pre-training tasks, Masked Text Box Regression and Shuffled Triplet Reconstruction, are proposed to learn the entailment relationship between text blocks and tokens as well as contextual information, respectively. Moreover, a DocVQA State Tracking Module (DocST) is also proposed to track the DocVQA state in the fine-tuning stage. Experimental results show that our model improves the performance onRDVQA significantly, although more work should be done for practical DocVQA as shown inRDVQA.
Xinya Wu, Duo Zheng, Jiashen Sun, Minzhen Hu, Fangxiang Feng, Xiaojie Wang 0006, Huixing Jiang, Fan Yang 0087
ACM Multimedia7
2022 Visual Dialog for Spotting the Differences between Pairs of Similar Images
abstract
Visual dialog has witnessed great progress after introducing various vision-oriented goals into the conversation. Much of previous work focuses on tasks where only one image can be accessed by two interlocutors, such as VisDial and GuessWhat. The work on situations where two interlocutors access different images has received less attention. Those situations are common in real world and bring some different challenges compared with one-image tasks. The lack of such types of dialog tasks and corresponding large-scale datasets makes it impossible to carry out in-depth research. This paper therefore first proposes a new visual dialog task named Dial-the-Diff, where two interlocutors accessing two similar images respectively try to spot the difference between the images through conversing in natural language. The task raises new challenges to the dialog strategy and the ability of categorizing objects. We then build a large-scale multi-modal dataset for the task, named DialDiff, which contains 87k Virtual Reality images and 78k dialogs. Some details of the data are given and analyzed to highlight the challenges behind the task. Finally, we propose benchmark models for this task, and conduct extensive experiments to evaluate their performance as well as its problems remained.
Duo Zheng, Fandong Meng, Qingyi Si, Hairun Fan, Zipeng Xu, Jie Zhou 0016, Fangxiang Feng, Xiaojie Wang 0006
ACM Multimedia8
2022 Domain adaptive multi-task transformer for low-resource machine reading comprehension
Ziwei Bai, Baoxun Wang, Zongsheng Wang, Caixia Yuan, Xiaojie Wang 0006
Neurocomputing5
2022 Modality Disentangled Discriminator for Text-to-Image Synthesis
abstract
Text-to-image (T2I) synthesis aims at generating photo-realistic images from text descriptions, which is a particularly important task in bridging vision and language. Each generated image consists of two parts: the content part related to the text and the style part irrelevant to the text. The existing discriminator does not distinguish between the content part and the style part. This not only precludes the T2I synthesis models from generating the content part effectively but also makes it difficult to manipulate the style of the generated image. In this paper, we propose a modality disentangled discriminator that distinguishes between the content part and the style part at a specific layer. Specifically, we enforce the early layers of a certain number in the discriminator to become the disentangled representation extractor through two losses. The extracted common representation for the content part can make the discriminator more effective for capturing the text-image correlation, while the extracted modality-specific representation for the style part can be directly transferred to other images. The combination of these two representations can also improve the quality of the generated images. Our proposed discriminator is used to substitute the discriminator of each stage in the representative model AttnGAN and the SOTA model DM-GAN. Extensive experiments are conducted on three widely used datasets, i.e. CUB, Oxford-102, and COCO, for the T2I synthesis task, demonstrating the superior performance of the modality disentangled discriminator over the base models. Code for DM-GAN with our modality disentangled discriminator is available athttps://github.com/FangxiangFeng/DM-GAN-MDD.
Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006
IEEE Trans. Multim.4
2022 A Colorization Framework for Monochrome-Color Dual-Lens Systems Using a Deep Convolutional Network
abstract
In monochrome-color dual-lens systems, the monochrome camera can capture images with higher quality than the color camera. To obtain high quality color images, a better approach is to colorize the gray images from the monochrome camera with the color images from the color camera serving as a reference. In addition, the colorization may fail in some cases, which makes the estimation of the colorization quality a necessary step before outputting the colorization result. To solve these problems, we propose a deep convolutional network based framework. 1) In the colorization module, the proposed colorization CNN uses deep feature representations, attention operation, 3-D regulation and color correction to make use of colors of multiple pixels in the reference image for colorizing each pixel in the input gray image. 2) In the colorization quality estimation module, based on the symmetry property of colorization, we propose to utilize the colorization CNN again to colorize the gray map of the original reference color image using the first-time colorization result from the colorization module as reference. Then, the quality loss of the second-time colorization result can be used for estimating the colorization quality. Experimental results show that our method can largely outperform the state-of-the-art colorization methods and estimate the colorization quality accurately as well.
Xuan Dong 0001, Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Yunhong Wang 0001
IEEE Trans. Vis. Comput. Graph.4
2021 MIEHDR CNN: Main Image Enhancement based Ghost-Free High Dynamic Range Imaging using Dual-Lens Systems
abstract
We study the High Dynamic Range (HDR) imaging problem using two Low Dynamic Range (LDR) images that are shot from dual-lens systems in a single shot time with different exposures. In most of the related HDR imaging methods, the problem is usually solved by Multiple Images Merging, i.e. the final HDR image is fused from pixels of all the input LDR images. However, ghost artifacts can be hardly avoided using this strategy. Instead of directly merging the multiple LDR inputs, we use an indirect way which enhances the main image, i.e. the short exposure image IS, using the long exposure image IL serving as guidance. In detail, we propose a new model, named MIEHDR CNN model, which consists of three subnets, i.e. Soft Warp CNN, 3D Guided Denoising CNN and Fusion CNN. The Soft Warp CNN aligns IL to get the aligned result ILA using the soft exposed result of IS as reference. The 3D Guided Denoising CNN denoises the soft exposed result of IS using ILA as guidance, whose result are fed into the Fusion CNN with IS to get the HDR result. The MIEHDR CNN model is implemented by MindSpore and experimental results show that we can outperform related methods largely and avoid ghost artifacts.
Xuan Dong 0001, Xiaoyan Hu 0006, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001
AAAI4
2021 Converse, Focus and Guess - Towards Multi-Document Driven Dialogue
Caixia Yuan, Xiaojie Wang 0006, Yushu Yang, Huixing Jiang, Zhongyuan Wang 0006
AAAI3
2021 Dual Graph Convolutional Networks for Aspect-based Sentiment Analysis
abstract
Ruifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy
ACL/IJCNLP (1)5
2021 Multi-stage Pre-training over Simplified Multimodal Pre-training Models
abstract
Tongtong Liu, Fangxiang Feng, Xiaojie Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Fangxiang Feng, Xiaojie Wang 0006
ACL/IJCNLP (1)3
2021 Modeling Explicit Concerning States for Reinforcement Learning in Visual Dialogue
Zipeng Xu, Fandong Meng, Xiaojie Wang 0006, Duo Zheng, Chenxu Lv, Jie Zhou 0016
BMVC3
2021 S2TD: A Tree-Structured Decoder for Image Paragraph Captioning
abstract
Image paragraph captioning, a task to generate the paragraph description for a given image, usually requires mining and organizing linguistic counterparts from abundant visual clues. Limited by sequential decoding perspective, previous methods have difficulty in organizing the visual clues holistically or capturing the structural nature of linguistic descriptions. In this paper, we propose a novel tree-structured visual paragraph decoder network, called Splitting to Tree Decoder (S2TD) to address this problem. The key idea is to model the paragraph decoding process as a top-down binary tree expansion. S2TD consists of three modules: a split module, a score module, and a word-level RNN. The split module iteratively splits ancestral visual representations into two parts through a gating mechanism. To determine the tree topology, the score module uses cosine similarity to evaluate the nodes splitting. A novel tree structure loss is proposed to enable end-to-end learning. After the tree expansion, the word-level RNN decodes leaf nodes into sentences forming a coherent paragraph. Extensive experiments are conducted on the Stanford benchmark dataset. The experimental results show promising performance of our proposed S2TD.
Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006
MMAsia6
2021 Pyramid convolutional network for colorization in monochrome-color multi-lens camera system
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006
Neurocomputing3
2021 Self-Supervised Colorization Towards Monochrome-Color Camera Systems Using Cycle CNN
abstract
Colorization in monochrome-color camera systems aims to colorize the gray image IGfrom the monochrome camera using the color image RCfrom the color camera as reference. Since monochrome cameras have better imaging quality than color cameras, the colorization can help obtain higher quality color images. Related learning based methods usually simulate the monochrome-color camera systems to generate the synthesized data for training, due to the lack of ground-truth color information of the gray image in the real data. However, the methods that are trained relying on the synthesized data may get poor results when colorizing real data, because the synthesized data may deviate from the real data. We present a self-supervised CNN model, named Cycle CNN, which can directly use the real data from monochrome-color camera systems for training. In detail, we use the Weighted Average Colorization (WAC) network to do the colorization twice. First, we colorize IGusing RCas reference to obtain the first-time colorization result IC. Second, we colorize the de-colored map of RC, i.e. RG, using the concatenated image of IGand Cb/Cr channels of the first-time colorization result IC, i.e. ICCband ICCr, as reference to obtain the second-time colorization result RC'. In this way, for the second-time colorization result RC', we use the Cb and Cr channels of the original color map RCas ground-truth and introduce the cycle consistency loss to push RC'Cb/Cr≈ RCCb/Cr. Also, for the Y channel of the first-time colorization result ICY, we propose the Global Curve Adjustment (GCA) network and the structure similarity loss to encourage the structure similarity between ICYand IG. In addition, we introduce a spatial smoothness loss within the WAC network to encourage spatial smoothness of the colorization result. Combining all these losses, we could train the Cycle CNN using the real data in the absence of the ground-truth color information of IG. Experimental results show that we can outperform related methods largely for colorizing real data.
Xuan Dong 0001, Chang Liu 0071, Weixin Li 0001, Xiaoyan Hu 0006, Xiaojie Wang 0006, Yunhong Wang 0001
IEEE Trans. Image Process.5
2020 Cycle-CNN for Colorization towards Real Monochrome-Color Camera Systems
abstract
Colorization in monochrome-color camera systems aims to colorize the gray image IG from the monochrome camera using the color image RC from the color camera as reference. Since monochrome cameras have better imaging quality than color cameras, the colorization can help obtain higher quality color images. Related learning based methods usually simulate the monochrome-color camera systems to generate the synthesized data for training, due to the lack of ground-truth color information of the gray image in the real data. However, the methods that are trained relying on the synthesized data may get poor results when colorizing real data, because the synthesized data may deviate from the real data. We present a new CNN model, named cycle CNN, which can directly use the real data from monochrome-color camera systems for training. In detail, we use the colorization CNN model to do the colorization twice. First, we colorize IG using RC as reference to obtain the first-time colorization result IC. Second, we colorize the de-colored map of RC, i.e. RG, using the first-time colorization result IC as reference to obtain the second-time colorization result R′C. In this way, for the second-time colorization result R′C, we use the original color map RC as ground-truth and introduce the cycle consistency loss to push R′C ≈ RC. Also, for the first-time colorization result IC, we propose a structure similarity loss to encourage the luminance maps between IG and IC to have similar structures. In addition, we introduce a spatial smoothness loss within the colorization CNN model to encourage spatial smoothness of the colorization result. Combining all these losses, we could train the colorization CNN model using the real data in the absence of the ground-truth color information of IG. Experimental results show that we can outperform related methods largely for colorizing real data.
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001
AAAI3
2020 Visual Dialogue State Tracking for Question Generation
abstract
GuessWhat?! is a visual dialogue task between a guesser and an oracle. The guesser aims to locate an object supposed by the oracle oneself in an image by asking a sequence of Yes/No questions. Asking proper questions with the progress of dialogue is vital for achieving successful final guess. As a result, the progress of dialogue should be properly represented and tracked. Previous models for question generation pay less attention on the representation and tracking of dialogue states, and therefore are prone to asking low quality questions such as repeated questions. This paper proposes visual dialogue state tracking (VDST) based method for question generation. A visual dialogue state is defined as the distribution on objects in the image as well as representations of objects. Representations of objects are updated with the change of the distribution on objects. An object-difference based attention is used to decode new question. The distribution on objects is updated by comparing the question-answer pair and objects. Experimental results on GuessWhat?! dataset show that our model significantly outperforms existing methods and achieves new state-of-the-art performance. It is also noticeable that our model reduces the rate of repeated questions from more than 50% to 21.9% compared with previous state-of-the-art methods.
Xiaojie Wang 0006
AAAI2
2020 Guessing State Tracking for Visual Dialogue
Xiaojie Wang 0006
ECCV (16)2
2020 Multi-scale Two-way Deep Neural Network for Stock Trend Prediction
abstract
Stock Trend Prediction(STP) has drawn wide attention from various fields, especially Artificial Intelligence. Most previous studies are single-scale oriented which results in information loss from a multi-scale perspective. In fact, multi-scale behavior is vital for making intelligent investment decisions. A mature investor will thoroughly investigate the state of a stock market at various time scales. To automatically learn the multi-scale information in stock data, we propose a Multi-scale Two-way Deep Neural Network. It learns multi-scale patterns from two types of scale-information, wavelet-based and downsampling-based, by eXtreme Gradient Boosting and Recurrent Convolutional Neural Network, respectively. After combining the learned patterns from the two-way, our model achieves state-of-the-art performance on FI-2010 and CSI-2016, where the latter is our published long-range stock dataset to help future studies for STP task. Extensive experimental results on the two datasets indicate that multi-scale information can significantly improve the STP performance and our model is superior in capturing such information.
Guang Liu 0007, Yuzhao Mao, Hailong Huang 0003, Weiguo Gao, Jianping Shen, Ruifan Li, Xiaojie Wang 0006
IJCAI9
2020 Multi-Domain Dialogue State Tracking with Hierarchical Task Graph
abstract
Multi-domain dialogue state tracking (DST), which tracks user goals and intentions across multiple domains, is a core task for multi-domain task-oriented dialogue system. Previous works in multi-domain DST focus on the open-vocabulary setting to alleviate the over-dependence on pre-defined ontology. However, they come up short of modeling the relationships among domains and slots in an explicit and efficient way. In this paper, we propose a multi-domain dialogue state tracker with hierarchical task graph (DST-HTG) to address the above issues. DST-HTG uses a copy mechanism to perform DST under the open-vocabulary setting, which makes our model eliminate the dependence on pre-defined full ontology. Moreover, we extend our DST model with a hierarchical task graph which has simple structure and rich semantic information to incorporate the relationships among domains and slots into DST process explicitly and efficiently. Empirical results show that DST-HTG achieves the state-of-the-art joint goal accuracy and slot accuracy in MultiWOZ 2.0, a recently proposed multi-domain task-oriented dialogue dataset, which indicates the effectiveness of our proposed model.
Tianhao Shen, Xiaojie Wang 0006
IJCNN2
2020 Image Synthesis from Locally Related Texts
abstract
Text-to-image synthesis refers to generating photo-realistic images from text descriptions. Recent works focus on generating images with complex scenes and multiple objects. However, the text inputs to these models are the only captions that always describe the most apparent object or feature of the image and detailed information (e.g. visual attributes) for regions and objects are often missing. Quantitative evaluation of generation performances is still an unsolved problem, where traditional image classification- or retrieval-based metrics fail at evaluating complex images. To address these problems, we propose to generate images conditioned on locally-related texts, i.e., descriptions of local image regions or objects instead of the whole image. Specifically, questions and answers (QAs) are chosen as locally-related texts, which makes it possible to use VQA accuracy as a new evaluation metric. The intuition is simple: higher image quality and image-text consistency (both globally and locally) can help a VQA model answer questions more correctly. We purposed VQA-GAN model with three key modules: hierarchical QA encoder, QA-conditional GAN and external VQA loss. These modules help leverage the new inputs effectively. Thorough experiments on two public VQA datasets demonstrate the effectiveness of the model and the newly proposed metric.
Tianrui Niu, Fangxiang Feng, Lingxuan Li, Xiaojie Wang 0006
ICMR4
2020 Learning Visual Features from Product Title for Image Retrieval
abstract
There is a huge market demand for searching for products by images in e-commerce sites. Visual features play the most important role in solving this content-based image retrieval task. Most existing methods leverage pre-trained models on other large-scale datasets with well-annotated labels, e.g. the ImageNet dataset, to extract visual features. However, due to the large difference between the product images and the images in ImageNet, the feature extractor trained on ImageNet is not efficient in extracting the visual features of product images. And retraining the feature extractor on the product images is faced with the dilemma of lacking the annotated labels. In this paper, we utilize the easily accessible text information, that is, the product title, as a supervised signal to learn the features of the product image. Specifically, we use the n-grams extracted from the product title as the label of the product image to construct a dataset for image classification. This dataset is then used to fine-tuned a pre-trained model. Finally, the basic max-pooling activation of convolutions (MAC) feature is extracted from the fine-tuned model. As a result, we achieve the fourth position in the Grand Challenge of AI Meets Beauty in 2020 ACM Multimedia by using only a single ResNet-50 model without any human annotations and pre-processing or post-processing tricks. Our code is available at: \urlhttps://github.com/FangxiangFeng/AI-Meets-Beauty-2020.
Fangxiang Feng, Tianrui Niu, Ruifan Li, Xiaojie Wang 0006, Huixing Jiang
ACM Multimedia4
2020 Weakly Supervised Real-time Image Cropping based on Aesthetic Distributions
abstract
Image cropping is an effective tool to edit and manipulate images to achieve better aesthetic quality. Most existing cropping approaches rely on the two-step paradigm where multiple candidate cropping areas are proposed initially and the optimal cropping window is determined based on some quality criteria for these candidates afterwards. The obvious disadvantage of this mechanism is its low efficiency due to the huge searching space of candidate crops. In order to tackle this problem, a weakly supervised cropping framework is proposed, where the distribution dissimilarity between high quality images and cropped images is used to guide the coordinate predictor's training and the ground truths of cropping windows are not required by the proposed method. Meanwhile, to improve the cropping performance, a saliency loss is also designed in the proposed framework to force the neural network to focus more on the interested objects in the image. Under this framework, the images can be cropped effectively by the trained coordinate predictor in a one-pass favor without multiple candidates proposals, which ensures the high efficiency of the proposed system . Also, based on the proposed framework, many existing distribution dissimilarity measurements can be applied to train the image cropping system with high flexibility, such as likelihood based and divergence based distribution dissimilarity measure proposed in this work. The experiments on the public databases show that the proposed cropping method achieves the state-of-the-art accuracy, and the high computation efficiency as fast as 285 FPS is also obtained.
Peng Lu 0007, Xujun Peng, Xiaojie Wang 0006
ACM Multimedia4
2020 Gray2ColorNet: Transfer More Colors from Reference Image
abstract
Image colorization is an effective approach to provide plausible colors for grayscale images, which can achieve better and pleasing visual qualities. Although exemplar based colorization approaches provide promising results, they are relied on semantic colors or global colors only from the reference images. For the former situation, when the correspondence between the input grayscale image and reference image is not established, the colors of the reference image cannot be transferred to the input grayscale image successfully. With the later circumstance, because only global colors are considered, it is hard to produce a color image whose objects have the same color as the reference image when they are semantically related. Thus, an end-to-end colorization network Gray2ColorNet is proposed in this work, where an attention gating mechanism based color fusion network is designed to accomplish the colorization tasks. Relied on the proposed method, the semantic colors and global color distribution from the reference image are fused effectively, which are transferred to the final color images along with the prior knowledge of colors contained in the training data. The experimental results demonstrate the superior colorization performances of the proposed method compared to other state-of-the-art approaches.
Peng Lu 0007, Jinbei Yu, Xujun Peng, Zhaoran Zhao, Xiaojie Wang 0006
ACM Multimedia5
2020 Answer-Driven Visual State Estimator for Goal-Oriented Visual Dialogue
abstract
A goal-oriented visual dialogue involves multi-turn interactions between two agents, Questioner and Oracle. During which, the answer given by Oracle is of great significance, as it provides golden response to what Questioner concerns. Based on the answer, Questioner updates its belief on target visual content and further raises another question. Notably, different answers drive into different visual beliefs and future questions. However, existing methods always indiscriminately encode answers after much longer questions, resulting in a weak utilization of answers. In this paper, we propose an Answer-Driven Visual State Estimator (ADVSE) to impose the effects of different answers on visual states. First, we propose an Answer-Driven Focusing Attention (ADFA) to capture the answer-driven effect on visual attention by sharpening question-related attention and adjusting it by answer-based logical operation at each turn. Then based on the focusing attention, we get the visual state estimation by Conditional Visual Information Fusion (CVIF), where overall information and difference information are fused conditioning on the question-answer state. We evaluate the proposed ADVSE to both question generator and guesser tasks on the large-scale GuessWhat?! dataset and achieve the state-of-the-art performances on both tasks. The qualitative results indicate that the ADVSE boosts the agent to generate highly efficient questions and obtains reliable visual attentions during the reasonable question generation and guess processes.
Zipeng Xu, Fangxiang Feng, Xiaojie Wang 0006, Yushu Yang, Huixing Jiang, Zhongyuan Wang 0006
ACM Multimedia3
2020 Referring Expression Generation via Visual Dialogue
Lingxuan Li, Tianrui Niu, Fangxiang Feng, Xiaojie Wang 0006
NLPCC (2)6
2020 Label-Wise Document Pre-training for Multi-label Text Classification
Caixia Yuan, Xiaojie Wang 0006
NLPCC (1)3
2020 Dual-CNN: A Convolutional language decoder for paragraph image captioning
Ruifan Li, Yihui Shi, Fangxiang Feng, Xiaojie Wang 0006
Neurocomputing5
2020 Multi-negative samples with Generative Adversarial Networks for image retrieval
Ruifan Li, Xuesen Zhang, Yuzhao Mao, Xiaojie Wang 0006
Neurocomputing5
2020 Exploring Global and Local Linguistic Representations for Text-to-Image Synthesis
abstract
The task of text-to-image synthesis is to generate photographic images conditioned on given textual descriptions. This challenging task has recently attracted considerable attention from the multimedia community due to its potential applications. Most of the up-to-date approaches are built based on generative adversarial network (GAN) models, and they synthesize images conditioned on the global linguistic representation. However, the sparsity of the global representation results in training difficulties on GANs and a shortage of fine-grained information in the generated images. To address this problem, we propose cross-modal global and local linguistic representations-based generative adversarial networks (CGL-GAN) by incorporating the local linguistic representation into the GAN. In our CGL-GAN, we construct a generator to synthesize the target images and a discriminator to judge whether the generated images conform with the text description. In the discriminator, we construct the cross-modal correlation by projecting the image representations at high and low levels onto the global and local linguistic representations, respectively. We design the hinge loss function to train our CGL-GAN model. We evaluate the proposed CGL-GAN on two publicly available datasets, the CUB and the MS-COCO. The extensive experiments demonstrate that incorporating fine-grained local linguistic information with cross-modal correlation can greatly improve the performance of text-to-image synthesis, even when generating high-resolution images.
Ruifan Li, Fangxiang Feng, Guangwei Zhang 0003, Xiaojie Wang 0006
IEEE Trans. Multim.5
2019 Learning a Deep Convolutional Network for Colorization in Monochrome-Color Dual-Lens System
abstract
In the monochrome-color dual-lens system, the gray image captured by the monochrome camera has better quality than the color image from the color camera, but does not have color information. To get high-quality color images, it is desired to colorize the gray image with the color image as reference. Related works usually use hand-crafted methods to search for the best-matching pixel in the reference image for each pixel in the input gray image, and copy the color of the best-matching pixel as the result. We propose a novel deep convolution network to solve the colorization problem in an end-to-end way. Based on our observation that, for each pixel in the input image, there usually exist multiple pixels in the reference image that have the correct colors, our method performs weighted average of colors of the candidate pixels in the reference image to utilize more candidate pixels with correct colors. The weight values between pixels in the input image and the reference image are obtained by learning a weight volume using deep feature representations, where an attention operation is proposed to focus on more useful candidate pixels and a 3-D regulation is performed to learn with context information. In addition, to correct wrongly colorized pixels in occlusion regions, we propose a color residue joint learning module to correct the colorization result with the input gray image as guidance. We evaluate our method on the Scene Flow, Cityscapes, Middlebury, and Sintel datasets. Experimental results show that our method largely outperforms the state-of-the-art methods.
Xuan Dong 0001, Weixin Li 0001, Xiaojie Wang 0006, Yunhong Wang 0001
AAAI3
2019 Differential Networks for Visual Question Answering
abstract
The task of Visual Question Answering (VQA) has emerged in recent years for its potential applications. To address the VQA task, the model should fuse feature elements from both images and questions efficiently. Existing models fuse image feature element vi and question feature element qi directly, such as an element product viqi. Those solutions largely ignore the following two key points: 1) Whether vi and qi are in the same space. 2) How to reduce the observation noises in vi and qi. We argue that two differences between those two feature elements themselves, like (vi − vj) and (qi −qj), are more probably in the same space. And the difference operation would be beneficial to reduce observation noise. To achieve this, we first propose Differential Networks (DN), a novel plug-and-play module which enables differences between pair-wise feature elements. With the tool of DN, we then propose DN based Fusion (DF), a novel model for VQA task. We achieve state-of-the-art results on four publicly available datasets. Ablation studies also show the effectiveness of difference operations in DF model.
Chenfei Wu, Jinlai Liu, Xiaojie Wang 0006, Ruifan Li
AAAI3
2019 MrMep: Joint Extraction of Multiple Relations and Multiple Entity Pairs Based on Triplet Attention
abstract
This paper focuses on how to extract multiple relational facts from unstructured text.Neural encoder-decoder models have provided a viable new approach for jointly extracting relations and entity pairs.However, these models either fail to deal with entity overlapping among relational facts, or neglect to produce the whole entity pairs.In this work, we propose a novel architecture that augments the encoder and decoder in two elegant ways.First, we apply a binary CNN classifier for each relation, which identifies all possible relations maintained in the text, while retaining the target relation representation to aid entity pair recognition.Second, we perform a multihead attention over the text and a triplet attention with the target relation interacting with every token of the text to precisely produce all possible entity pairs in a sequential manner.Experiments on three benchmark datasets show that our proposed method successfully addresses the multiple relations and multiple entity pairs even with complex overlapping and significantly outperforms the state-of-theart methods.All source code and documentations are available at https://github. com/chenjiayu1502/MrMep.
Caixia Yuan, Xiaojie Wang 0006, Ziwei Bai
CoNLL3
2019 A new metric for individual stock trend prediction
Guang Liu 0006, Xiaojie Wang 0006
Eng. Appl. Artif. Intell.2
2019 Cascaded deep neural network models for dialog state tracking
Guohua Yang, Xiaojie Wang 0006
Multim. Tools Appl.2
2019 Hierarchical Dialog State Tracking with Unknown Slot Values
Guohua Yang, Xiaojie Wang 0006, Caixia Yuan
Neural Process. Lett.2
2018 Show and Tell More: Topic-Oriented Multi-Sentence Image Captioning
abstract
Image captioning aims to generate textual descriptions for images. Most previous work generates a single-sentence description for each image. However, a picture is worth a thousand words. Single-sentence can hardly give a complete view of an image even by humans. In this paper, we propose a novel Topic-Oriented Multi-Sentence (\emph{TOMS}) captioning model, which can generate multiple topic-oriented sentences to describe an image. Different from object instances or attributes, topics mined by the latent Dirichlet allocation reflect hidden thematic structures in reference sentences of an image. In our model, each topic is integrated to a caption generator with a Fusion Gate Unit (FGU) to guide the generation of a sentence towards a certain topic perspective. With multiple sentences from different topics, our \emph{TOMS} provides a complete description of an image. Experimental results on both sentence and paragraph datasets demonstrate the effectiveness of our \emph{TOMS} in terms of topical consistency and descriptive completeness.
Yuzhao Mao, Xiaojie Wang 0006, Ruifan Li
IJCAI3
2018 Differentiated Attentive Representation Learning for Sentence Classification
abstract
Attention-based models have shown to be effective in learning representations for sentence classification. They are typically equipped with multi-hop attention mechanism. However, existing multi-hop models still suffer from the problem of paying much attention to the most frequently noticed words, which might not be important to classify the current sentence. And there is a lack of explicitly effective way that helps the attention to be shifted out of a wrong part in the sentence. In this paper, we alleviate this problem by proposing a differentiated attentive learning model. It is composed of two branches of attention subnets and an example discriminator. An explicit signal with the loss information of the first attention subnet is passed on to the second one to drive them to learn different attentive preference. The example discriminator then selects the suitable attention subnet for sentence classification. Experimental results on real and synthetic datasets demonstrate the effectiveness of our model.
Qianrong Zhou, Xiaojie Wang 0006, Xuan Dong 0001
IJCAI2
2018 Object-Difference Attention: A Simple Relational Attention for Visual Question Answering
abstract
Attention mechanism has greatly promoted the development of Visual Question Answering (VQA). Attention distribution, which weights differently on objects (such as image regions or bounding boxes) in an image according to their importance for answering a question, plays a crucial role in attention mechanism. Most of the existing work focuses on fusing image features and text features to calculate the attention distribution without comparisons between different image objects. As a major property of attention, selectivity depends on comparisons between different objects. Comparisons provide more information for assigning attentions better. For achieving this, we propose an object-difference attention (ODA) which calculates the probability of attention by implementing difference operator between different image objects in an image under the guidance of questions in hand. Experimental results on three publicly available datasets show our ODA based VQA model achieves the state-of-the-art results. Furthermore, a general form of relational attention is proposed. Besides ODA, several other relational attentions are given. Experimental results show those relational attentions have strengths on different types of questions.
Chenfei Wu, Jinlai Liu, Xiaojie Wang 0006, Xuan Dong 0001
ACM Multimedia3
2018 Chain of Reasoning for Visual Question Answering
abstract
Reasoning plays an essential role in Visual Question Answering (VQA). Multi-step and dynamic reasoning is often necessary for answering complex questions. For example, a question "What is placed next to the bus on the right of the picture?" talks about a compound object "bus on the right," which is generated by the relation . Furthermore, a new relation including this compound object is then required to infer the answer. However, previous methods support either one-step or static reasoning, without updating relations or generating compound objects. This paper proposes a novel reasoning model for addressing these problems. A chain of reasoning (CoR) is constructed for supporting multi-step and dynamic reasoning on changed relations and objects. In detail, iteratively, the relational reasoning operations form new relations between objects, and the object refining operations generate new compound objects from relations. We achieve new state-of-the-art results on four publicly available datasets. The visualization of the chain of reasoning illustrates the progress that the CoR generates new compound objects that lead to the answer of the question step by step.
Chenfei Wu, Jinlai Liu, Xiaojie Wang 0006, Xuan Dong 0001
NeurIPS3
2018 Supervised latent Dirichlet allocation with a mixture of sparse softmax
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002
Neurocomputing6
2018 Corrigendum to "Supervised latent Dirichlet allocation with a mixture of sparse softmax" [Neurocomputing, volume 312, 27 October 2018, Pages 324-335]
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002
Neurocomputing6
2017 Cascaded LSTMs Based Deep Reinforcement Learning for Goal-Driven Dialogue
Xiaojie Wang 0006, Zhenjiang Dong
NLPCC2
2017 Jointly Modeling Intent Identification and Slot Filling with Contextual and Hierarchical Information
Liyun Wen, Xiaojie Wang 0006, Zhenjiang Dong
NLPCC2
2016 Image color harmony modeling through neighbored co-occurrence colors
Peng Lu 0007, Xujun Peng, Caixia Yuan, Ruifan Li, Xiaojie Wang 0006
Neurocomputing5
2015 Measuring the External Influence in Information Diffusion
abstract
Information flow in social network is assumed to be transmitted from node to node through the edges of network. In real world, however, people are influenced not only by local social neighbors but also by out-of-network services and sources, such as mass media and external websites. As a consequence, in addition to spreading by social edges, information can also reach a long-distance node by "jumping" cross the network. Then one of important issues coming out of the phenomenon is: how do these external services affect the diffusion process in social network? In this paper we develop an algorithm which allows us to distinguish the effects of external influence in diffusion process. By applying the algorithm to millions of diffusion cascades, we find that, although only a small portion of reshare activities arise from external influence directly, external influence plays a significant role in information diffusion. In particular, external influence affects nearly 50% to 70% of cascade node in average, and the effects become stronger as the cascade becomes larger. In addition, we characterize external influence as two categories, and show that one category mainly affects the size of diffusion tree and the other focuses on affecting the depth. Finally, we find that, due to the external services, the influentials become less important and more large cascades can be triggered by ordinary people. Together, these observations suggest new directions for modeling diffusion process and constructing more useful viral marketing strategy.
Jiagui Xiong, Xiaojie Wang 0006
MDM (2)3
2015 Recognition of Person Relation Indicated by Predicates
abstract
This paper focuses on recognizing person relations indicated by predicates from large scale of free texts. In order to determine whether a sentence contains a potential relation between persons, we cast this problem to a classification task. Dynamic Convolution Neural Network (DCNN) is improved for this task. It uses frame convolution for making uses of more features efficiently. Experimental results on Chinese person relation recognition show that the proposed model is superior when compared to the original DCNN and several strong baseline models. We also explore employing large scale unlabeled data to achieve further improvements.
Zhongping Liang, Caixia Yuan, Bing Leng, Xiaojie Wang 0006
NLPCC4
2015 Stochastic Language Generation Using Situated PCFGs
abstract
This paper presents a purely data-driven approach for generating natural language (NL) expressions from its corresponding semantic representations. Our aim is to exploit a parsing paradigm for natural language generation (NLG) task, which first encodes semantic representations with a situated probabilistic context-free grammar (PCFG), then decodes and yields natural sentences at the leaves of the optimal parsing tree. We deployed our system in two different domains, one is response generation for a Chinese spoken dialogue system, and the other is instruction generation for a virtual environment in English language, obtaining results comparable to state-of-the-art systems both in terms of BLEU scores and human evaluation.
Caixia Yuan, Xiaojie Wang 0006, Ziming Zhong
NLPCC2
2015 Cross-lingual Pseudo Relevance Feedback Based on Weak Relevant Topic Alignment
Xuwen Wang, Xiaojie Wang 0006, Junlian Li
PACLIC3
2015 Deep correspondence restricted Boltzmann machine for cross-modal retrieval
Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006
Neurocomputing3
2015 Challenges in representation learning: A report on three machine learning contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
Neural Networks15
2015 Towards aesthetics of image: A Bayesian framework for color harmony modeling
Peng Lu 0007, Xujun Peng, Ruifan Li, Xiaojie Wang 0006
Signal Process. Image Commun.4
2015 Correspondence Autoencoders for Cross-Modal Retrieval
abstract
This article considers the problem of cross-modal retrieval, such as using a text query to search for images and vice-versa. Based on different autoencoders, several novel models are proposed here for solving this problem. These models are constructed by correlating hidden representations of a pair of autoencoders. A novel optimal objective, which minimizes a linear combination of the representation learning errors for each modality and the correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimizing the correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimizing the representation learning error makes hidden representations good enough to reconstruct inputs of each modality. To balance the two kind of errors induced by representation learning and correlation learning, we set a specific parameter in our models. Furthermore, according to the modalities the models attempt to reconstruct they are divided into two groups. One group including three models is named multimodal reconstruction correspondence autoencoder since it reconstructs both modalities. The other group including two models is named unimodal reconstruction correspondence autoencoder since it reconstructs a single modality. The proposed models are evaluated on three publicly available datasets. And our experiments demonstrate that our proposed correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multimodal deep models on cross-modal retrieval tasks.
Fangxiang Feng, Xiaojie Wang 0006, Ruifan Li
ACM Trans. Multim. Comput. Commun. Appl.2
2014 Object Ranking on Deformable Part Models with Bagged LambdaMART
Chaobo Sun, Xiaojie Wang 0006, Peng Lu 0007
ACCV (2)2
2014 Cross-modal Retrieval with Correspondence Autoencoder
abstract
The problem of cross-modal retrieval, e.g., using a text query to search for images and vice-versa, is considered in this paper. A novel model involving correspondence autoencoder (Corr-AE) is proposed here for solving this problem. The model is constructed by correlating hidden representations of two uni-modal autoencoders. A novel optimal objective, which minimizes a linear combination of representation learning errors for each modality and correlation learning error between hidden representations of two modalities, is used to train the model as a whole. Minimization of correlation learning error forces the model to learn hidden representations with only common information in different modalities, while minimization of representation learning error makes hidden representations are good enough to reconstruct input of each modality. A parameter $\alpha$ is used to balance the representation learning error and the correlation learning error. Based on two different multi-modal autoencoders, Corr-AE is extended to other two correspondence models, here we called Corr-Cross-AE and Corr-Full-AE. The proposed models are evaluated on three publicly available data sets from real scenes. We demonstrate that the three correspondence autoencoders perform significantly better than three canonical correlation analysis based models and two popular multi-modal deep models on cross-modal retrieval tasks.
Fangxiang Feng, Xiaojie Wang 0006, Ruifan Li
ACM Multimedia2
2013 Challenges in Representation Learning: A Report on Three Machine Learning Contests
Ian J. Goodfellow, Dumitru Erhan, Pierre Luc Carrier, Aaron C. Courville, Mehdi Mirza, Benjamin Hamner, Will Cukierski, Yichuan Tang, Dave Thaler, Yingbo Zhou 0002, Chetan Ramaiah, Fangxiang Feng, Ruifan Li, Xiaojie Wang 0006, Dimitris Athanasakis, John Shawe-Taylor, Maxim Milakov, John Park, Radu Tudor Ionescu, Marius Popescu, Cristian Grozea, James Bergstra, Jingjing Xie, Lukasz Romaszko, Yoshua Bengio
ICONIP (3)15
2011 An Improved SalBayes Model with GMM
Hairu Guo, Xiaojie Wang 0006, Yixin Zhong, Song Bi
CAIP (2)2
2011 Can Persona Facilitate Ideation? A Comparative Study on Effects of Personas in Brainstorming
Xiantao Chen, Ying Liu 0041, Xiaojie Wang 0006
INTERACT (4)4
2010 Second-Order HMM for Event Extraction from Short Message
Huixing Jiang, Xiaojie Wang 0006, Jilei Tian
NLDB2
2010 Dependency Relation Based Detection of Lexicalized User Goals
Ruixue Duan, Xiaojie Wang 0006, Rile Hu, Jilei Tian
UIC2
2009 Injecting Structured Data to Generative Topic Model in Enterprise Settings
Han Xiao 0002, Xiaojie Wang 0006
ACML2
2008 BUPT Systems in the SIGHAN Bakeoff 2007
Caixia Yuan, Jiashen Sun, Xiaojie Wang 0006
IJCNLP4
2005 Chinese-Japanese Clause Alignment
Xiaojie Wang 0006, Fuji Ren
CICLing1
1999 A new way to conceptual meaning representation
Xiaojie Wang 0006, Yixin Zhong
MTSummit1