VLDB 2026 Research / reviewers in the wild / expert
Xinyu Xiao
dblp:219/1913
· DBLP profile ↗
30ranked-venue papers
13as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AWMA-MoE: Attention-Guided Watermark Adapter with MoE for Latent Diffusion ModelsabstractWith the evolving generative models, generated images are closer to reality, raising concerns about information authenticity and malicious misuse. Invisible watermarks offer a practical approach to detecting and tracing them. However, while image watermarking inevitably introduces quality degradation, most existing methods primarily focus on improving watermark robustness. To address this limitation, we propose AWMA-MoE, a framework that enhances the quality of generated images while preserving strong watermark robustness. Specifically, we design an attention-based adapter that adaptively embeds watermarks with spatially varying strengths across image regions. Building upon this, we introduce an MoE architecture that leverages diverse experts to further improve image quality while retaining watermark robustness. Experiments demonstrate that AWMA-MoE can reduce the distortion of generated images and exhibit competitive watermark performance, thus striking an improved balance for watermarking generated image tasks and better linking post-hoc and in-generation methods. Xinyu Xiao, Jian Zhang 0019, Shuhan Qi, Yulin Wu 0001, Xuan Wang 0002 |
WWW | 2 |
| 2026 | PruneRAG: Confidence-Guided Query Decomposition Trees for Efficient Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) has become a powerful framework for enhancing large language models in knowledge-intensive and reasoning tasks. However, as reasoning chains deepen or search trees expand, RAG systems often face two persistent failures: evidence forgetting, where retrieved knowledge is not effectively used, and inefficiency, caused by uncontrolled query expansions and redundant retrieval. These issues reveal a critical gap between retrieval and evidence utilization in current RAG architectures. We propose PruneRAG, a confidence-guided query decomposition framework that builds a structured query decomposition tree to perform stable and efficient reasoning. PruneRAG introduces three key mechanisms: adaptive node expansion that regulates tree width and depth, confidence-guided decisions that accept reliable answers and prune uncertain branches, and fine-grained retrieval that extracts entity-level anchors to improve retrieval precision. Together, these components preserve salient evidence throughout multi-hop reasoning while significantly reducing retrieval overhead. To better analyze evidence misuse, we define the Evidence Forgetting Rate as a metric to quantify cases where golden evidence is retrieved but not correctly used. Extensive experiments across various multi-hop QA benchmarks show that PruneRAG achieves superior accuracy and efficiency over state-of-the-art baselines. The code is publicly available. Shuguang Jiao, Xinyu Xiao, Yunfan Wei, Shuhan Qi, Chengkai Huang, Quan Z. Sheng, Lina Yao 0001 |
WWW | 2 |
| 2026 | Uncertainty-aware mixture of experts for robust multimodal sentiment analysis
Xinyu Xiao, Xuan Wang 0002, Shuhan Qi, Huale Li |
Pattern Recognit. | 1 |
| 2026 | Dual Feature Fusion for Incomplete Multi-View Multi-Label LearningabstractMulti-view Multi-label Learning (MVML) aims to leverage multi-view information from input samples to achieve accurate predictions of multiple labels. Unfortunately, most existing MVML methods operate under the assumption of data completeness, which makes them ineffective in practical scenarios involving missing views or uncertain labels. Recent methods address incomplete data, but few approaches handle scenarios where both views and labels are missing. To address this challenge, we propose a Dual-view Feature-guided Fusion Learning (DFFL) framework. DFFL considers both view-specific unique features and inter-view consistent features. Specifically, DFFL constructs view-uniqueness contrastive learning to ensure that features within the same view maintain high semantic relevance under the condition of view missing, while the semantics between different views are different. Unlike previous methods, DFFL assumes that label relevance can be reversely mapped to high-dimensional features. By establishing View-consistency learning, the mutual information in the shared embedding space is maximized to achieve consistent feature alignment. In particular, DFFL minimizes the conditional entropy of the marginal distribution of multi-view features through dual prediction, thereby deriving the maximum joint distribution of feature fusion and combining the missing view index matrix to achieve feature fusion. This process can effectively alleviate the fusion feature suppression existing in previous methods. Finally, the missing label index matrix is combined with the fusion feature to complete the classification task. We validate the framework on five widely used datasets, and experimental results demonstrate that our approach achieves superior performance compared to state-of-the-art methods. Ablation studies further validated the effectiveness of each component in DFFL. Xinyu Xiao, Shuhan Qi, Yulin Wu 0001, Bin Chen 0011, Xuan Wang 0002 |
IEEE Trans. Multim. | 1 |
| 2025 | Privacy-Preserving V2X Collaborative Perception Integrating Unknown CollaboratorsabstractVehicle-to-everything (V2X) collaborative perception has recently gained increasing attention in autonomous driving due to its ability to enhance scene understanding by integrating information from other collaborators, e.g. vehicles or infrastructure. Existing algorithms usually share deep features to achieve a trade-off between accuracy and bandwidth. However, most of these methods require joint training of all agents, which results in privacy leakage and is impractical and unacceptable in the real world. Sharing prediction results seems to be a direct solution, but its performance is suboptimal and sensitive to localization noise and communication delay. In this paper, we propose a privacy-preserving collaborative perception framework, where each agent is separately trained with its own dataset and the ego vehicle needs to integrate with completely unknown collaborators. Specifically, we propose MSD, a multi-scale feature fusion method combined with deformable attention, to better fuse features of different agents. We also propose a plug-in domain adapter to align the features from unknown collaborators to ego-domain. Extensive experiments on the challenging DAIR-V2X and V2V4Real demonstrate that: 1) MSD achieves remarkable performance, outperforming others by at least 2.8% and 6.7% in AP0.7 on DAIR-V2X and V2V4Real, respectively; 2) After domain adaptation, it significantly outperforms the No Fusion, Late Fusion scenarios and can approach or even surpass the performance of joint training. We truly achieves privacy-preserving collaboration, providing a new paradigm for the study of collaborative perception, which is crucial for practical applications. Xinyu Xiao, Changzhou Zhang, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
AAAI | 2 |
| 2025 | Towards Building Human-like Smart Agents in Modern 3D Video Games (Student Abstract)abstractIn recent years, reinforcement learning has been widely applied in the field of games. However, most studies focus on assisting agents to achieve victory, with less attention paid to whether the agents exhibit human-like characteristics. In order to build human-like agents with high performance, we propose a method for learning the strategies of human players in modern three-dimensional video games. Our method utilizes a hierarchical framework, learning basic behaviors and intentions of human players at the lower level through imitation learning, and generalized policies at the high level through reinforcement learning. Compared with other existing methods, our method demonstrates significant advantages in learning human-like strategies in complex environments. Zhihang Sun, Shuhan Qi, Xinhao Huang, Xinyu Xiao, Jiajia Zhang 0001, Xuan Wang 0002, Peixi Peng |
AAAI | 4 |
| 2025 | Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMsabstractRecent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities.Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Continual Learning (MCL).The exploration of MCL is in its early stages.In this work, we dive into the causes of performance degradation in MCL.We uncover that it suffers not only from forgetting as in traditional continual learning, but also from misalignment between the modality-agnostic and modality-specific components.To this end, we propose an elegantly simple MCL paradigm called "MErge then ReAlign" (MERA) to address both forgetting and misalignment.MERA avoids introducing heavy model budgets or modifying model architectures, hence is easy to deploy and highly reusable in the MLLM community.Extensive experiments demonstrate the impressive performance of MERA, holding an average of 99.84% Backward Relative Gain when extending to four modalities, achieving nearly lossless MCL performance.Our findings underscore the misalignment issue in MCL.More broadly, our work showcases how to adjust different components of MLLMs during continual learning. Dingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen, Xuan Wang 0002 |
EMNLP | 3 |
| 2025 | Unified Visual Generation via Next-Set Prediction in Continuous Domain
Zhanzhou Feng, Qingpei Guo, Xinyu Xiao, Ruihan Xu 0002, Ming Yang 0007, Shiliang Zhang |
ICCV | 3 |
| 2025 | Multi-faceted Complementary Learning for Incomplete Multi-view Multi-label Classification
Xinyu Xiao, Peixi Peng, Qiang Wang 0022, Shuhan Qi |
ACM Multimedia | 1 |
| 2025 | RayFusion: Ray Fusion Enhanced Collaborative Visual PerceptionabstractCollaborative visual perception methods have gained widespread attention in the autonomous driving community in recent years due to their ability to address sensor limitation problems. However, the absence of explicit depth information often makes it difficult for camera-based perception systems, e.g., 3D object detection, to generate accurate predictions. To alleviate the ambiguity in depth estimation, we propose RayFusion, a ray-based fusion method for collaborative visual perception. Using ray occupancy information from collaborators, RayFusion reduces redundancy and false positive predictions along camera rays, enhancing the detection performance of purely camera-based collaborative perception systems. Comprehensive experiments show that our method consistently outperforms existing state-of-the-art models, substantially advancing the performance of collaborative visual perception. Our code will be made publicly available. Shaohong Wang, Lu Bin, Xinyu Xiao, Hanzhi Zhong, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
NeurIPS | 3 |
| 2025 | Exploring Fine-Grained Image-Text Alignment for Referring Remote Sensing Image SegmentationabstractGiven a language expression, referring remote sensing image segmentation (RRSIS) aims to identify ground objects and assign pixelwise labels within the imagery. One of the key challenges for this task is to capture discriminative multimodal features via image-text alignment. However, the existing RRSIS methods use one vanilla and coarse alignment, where the language expression is directly extracted to be fused with the visual features. In this article, we argue that a “fine-grained image-text alignment” can improve the extraction of multimodal information. To this point, we propose a new RRSIS method to fully exploit the visual and linguistic representations. Specifically, the original referring expression is regarded as context text, which is further decoupled into the ground object and spatial position texts. The proposed fine-grained image-text alignment module (FIAM) would simultaneously leverage the features of the input image and the corresponding texts, obtaining better discriminative multimodal representation. Meanwhile, to handle the various scales of ground objects in remote sensing, we introduce a text-aware multiscale enhancement module (TMEM) to adaptively perform cross-scale fusion and intersections. We evaluate the effectiveness of the proposed method on two public referring remote sensing datasets including RefSegRS and RRSIS-D, and our method obtains superior performance over several state-of-the-art methods. The code will be publicly available athttps://github.com/Shaosifan/FIANet. Sen Lei, Xinyu Xiao, Heng-Chao Li 0001, Zhenwei Shi 0001, Qing Zhu 0012 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | IFTR: An Instance-Level Fusion Transformer for Visual Collaborative Perception
Shaohong Wang, Lu Bin, Xinyu Xiao, Zhiyu Xiang, Hangguan Shan, Eryun Liu |
ECCV (87) | 3 |
| 2024 | Local-to-Global Self-Consistency Learning for Temporal Action LocalizationabstractThe object of temporal action localization (TAL) is to predict the predefined action labels and the corresponding temporal boundary in a video. It can be found that TAL is a task of multi-modal modeling and highly dependent on the effect of temporal context representation. Inspired by this property, we propose an end-to-end local-to-global modeling architecture to learn the contextual consistency information in temporal sequence and cross-modal. Specifically, a local-to-global encoding Transformer is applied to model the video sequence to obtain video representation of different time scales. To achieve a reasonable balance between the specificity and correlation of different modalities, a cross semantic alignment (CSA) module is proposed to re-weight the encoded multi-model features by whether attending to the semantic correlations or specificity in different modalities. Further, to learn the trans-modal consistency from local to global and the uni-modal consistency belonging to the same category, the self-consistency learning (SCL) is designed to train the network. The experimental results demonstrate the significance of our method in major improvements upon prior works. Our model achieves 68.3% and 37.1% average mAPs on THUMOS14 and ActivityNet 1.3, outperforming state-of-the-art multi-stage and one-stage models. Xinyu Xiao, Yun Hu 0003, Eryun Liu |
ICME | 1 |
| 2024 | GTPAN: Global Target Preference Attention Network for session-based recommendation
Tingwei Lu, Xinyu Xiao, Yin Xiao, Junhao Wen 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Preformer: Simple and Efficient Design for Precipitation Nowcasting With TransformersabstractThe primary objective of precipitation nowcasting is to predict precipitation patterns several hours in advance. Recent studies have emphasized the potential of deep learning methods for this task. To harness the correlations among various meteorological elements, existing frameworks project multiple meteorological elements into a latent space and then utilize convolutional-recurrent networks for future precipitation prediction. Although effective, the escalating model complexity may impede practical applications. This letter develops the Preformer, a streamlined Transformer framework for precipitation nowcasting that efficiently captures global spatiotemporal dependencies among multiple meteorological elements. The Preformer implements an encoder-translator-decoder architecture, where the encoder integrates spatial features of multiple elements, the translator models spatiotemporal dynamics, and the decoder combines spatiotemporal information to forecast future precipitation. Without introducing complex structures or strategies, the Preformer achieves state-of-the-art performance even with the least parameters. Qizhao Jin, Xinbang Zhang, Xinyu Xiao, Ying Wang 0008, Shiming Xiang, Chunhong Pan |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Unsupervised Modality-Transferable Video Highlight Detection With Representation Activation Sequence LearningabstractIdentifying highlight moments of raw video materials is crucial for improving the efficiency of editing videos that are pervasive on internet platforms. However, the extensive work of manually labeling footage has created obstacles to applying supervised methods to videos of unseen categories. The absence of an audio modality that contains valuable cues for highlight detection in many videos also makes it difficult to use multimodal strategies. In this paper, we propose a novel model with cross-modal perception for unsupervised highlight detection. The proposed model learns representations with visual-audio level semantics from image-audio pair data via a self-reconstruction task. To achieve unsupervised highlight detection, we investigate the latent representations of the network and propose the representation activation sequence learning (RASL) module with k-point contrastive learning to learn significant representation activations. To connect the visual modality with the audio modality, we use the symmetric contrastive learning (SCL) module to learn the paired visual and audio representations. Furthermore, an auxiliary task of masked feature vector sequence (FVS) reconstruction is simultaneously conducted during pretraining for representation enhancement. During inference, the cross-modal pretrained model can generate representations with paired visual-audio semantics given only the visual modality. The RASL module is used to output the highlight scores. The experimental results show that the proposed framework achieves superior performance compared to other state-of-the-art approaches. Tingtian Li, Zixun Sun, Xinyu Xiao |
IEEE Trans. Image Process. | 3 |
| 2023 | LSIAN: Exploiting interval interests for session-based recommendation via sparse attention network
Xinyu Xiao, Wei Zhou 0028, Junhao Wen 0001 |
Inf. Sci. | 1 |
| 2022 | Spatiotemporal Contextual Consistency Network for Precipitation NowcastingabstractPrecipitation nowcasting is forecasting rainfall in the short-term conditioned by the known meteorological parameters. Recently, deep neural networks (DNNs) have shown outstanding performance in this task. But, there are several challenges imposed by the multiple meteorological elements, including the multimodal modeling, the considerable variation in scales of precipitation region, as well as the long-tailed distribution of rainfall data. To solve these problems, this paper proposes Spatiotemporal Contextual Consistency Network (SCCN) for learning from the multi meteorological elements. Architecturally, a parameter-shared multimodal fusion CNN encoder, which dynamically exchanges features between different modalities, is used to encode the multimodal meteorological data. To improve the spatial modeling of the multiple meteorological features, we compose the multi-scale filters and deconstruction convolution to modify the gate operators in ConvLSTM to propose a spatial contextual consistency ConvLSTM (SCC-ConvLSTM). Furthermore, considering the temporal consistency in rainfall, a temporal consistency module (TCM) is designed to gear to long-tailed distribution. Under this module, different long-tailed meteorological elements are calculated to encode features and residuals fused with the previous precipitation distribution in sequence. The experimental results of precipitation nowcasting demonstrate the effectiveness of our method on the ERA5 dataset and WeatherBench dataset. Xinyu Xiao, Qizhao Jin, Gaofeng Meng, Shiming Xiang, Chunhong Pan |
ICDM | 1 |
| 2022 | Relational Graph Reasoning Transformer for Image CaptioningabstractThe current published methods of image captioning are directly inputting the features of objects in image into model, and introduced a variety of attention mechanisms to capture the associations between the objects and specific words. But the relationships of vision and semantic between objects are not sufficiently concerned. In this paper, we propose a relational graph reasoning Transformer which explicitly incorporates the relationships of vision and semantic between objects to construct an object relational graph in Transformer. Specifically, besides the detected object features, the global spatial relationships and the semantic context between different objects is attended. Meanwhile, a graph structures feature which correlates object features, their spatial and semantic information is reasoned by a learned grafting mechanism. Finally, the contextual graph feature is integrated into the proposed Transformer decoder. Experimental results demonstrate the significance of our relationship reasoning Transformer model. Xinyu Xiao, Zixun Sun, Tingtian Li, Yipeng Yu |
ICME | 1 |
| 2022 | Improving Graph Neural Network For Session-based Recommendation System Via Time SessionsabstractSession-based recommendations (SBR) have attracted much attention because of their high commercial value. SBR aims to predict a user's next click based on the sequence of previous sessions. However, existing graph neural network (GNN)-based recommendation methods only focus on session graphs and ignore the time frame in the session which contains valuable temporal information for reducing the impact of user's unintentional clicks. Meanwhile, the expression ability of hidden factors in GNNs is neglected. In order to solve the above limitations, this paper propose an improving Graph Neural Network via Time sessions (TGNN) model. For the data preprocessing aspect, TGNN builds temporal sessions from the time frame in the dataset. Dwell time recurrent network explores complex temporal influence on items by learning time features. Multi-head attention is introduced to enhance the expression ability of hidden factors. Target-aware attention is utilized to activate different user interests for different target items adaptively, which can obtain interest expression changes of target items. Comparative experiments are conducted on two public datasets, and experimental results show that our method outperforms state-of-the-art methods. Ablation experiments demonstrate the effectiveness of the proposed model. Xuqi Yang, Xinyu Xiao, Wei Zhou 0028, Junhao Wen 0001 |
IJCNN | 2 |
| 2022 | CALM: Constrastive Cross-modal Speaking Style Modeling for Expressive Text-to-Speech Synthesis
Xiang Li 0105, Zhiyong Wu 0001, Tingtian Li, Zixun Sun, Xinyu Xiao, Chi Sun, Hui Zhan, Helen M. Meng |
INTERSPEECH | 6 |
| 2022 | Multi-interaction fusion collaborative filtering for social recommendation
Xinyu Xiao, Junhao Wen 0001, Wei Zhou 0028, Fengji Luo, Min Gao 0001, Jun Zeng 0003 |
Expert Syst. Appl. | 1 |
| 2021 | Reinforcement Stacked Learning with Semantic-Associated Attention for Visual Question AnsweringabstractThe task of visual question answering (VQA) is to generate an answer for a question according to the content of an image being asked. In this process, the critical problems of effectively embedding the question feature and image feature as well as transforming the features to the prediction of answer are still faithfully unresolved. In this paper, depending on these problems, a semantic-associated attention method and a reinforcement stacked learning mechanism are proposed. Firstly, within the associations of high-level semantics, a visual spatial attention model (VSA) and a multi-semantic attention model (MSA) are proposed to extract the low-level image feature and high-level semantic feature, respectively. Furthermore, we develop a reinforcement stacked learning architecture, which splits the transformation process into multiple stages, to gradually approach the answers. At each stage, a new reinforcement learning (RL) method is introduced to directly criticize inappropriate answers to optimize the model. The extensive experiments on the VQA task show that our method can achieve state-of-the-art performance. Xinyu Xiao, Chunxia Zhang 0001, Shiming Xiang, Chunhong Pan |
ICASSP | 1 |
| 2021 | Relational Attention with Textual Enhanced Transformer for Image Captioning
Lifei Song, Yiwen Shi, Xinyu Xiao, Chunxia Zhang 0001, Shiming Xiang |
PRCV (3) | 3 |
| 2019 | What and Where the Themes Dominate in ImageabstractThe image captioning is to describe an image with natural language as human, which has benefited from the advances in deep neural network and achieved substantial progress in performance. However, the perspective of human description to scene has not been fully considered in this task recently. Actually, the human description to scene is tightly related to the endogenous knowledge and the exogenous salient objects simultaneously, which implies that the content in the description is confined to the known salient objects. Inspired by this observation, this paper proposes a novel framework, which explicitly applies the known salient objects in image captioning. Under this framework, the known salient objects are served as the themes to guide the description generation. According to the property of the known salient object, a theme is composed of two components: its endogenous concept (what) and the exogenous spatial attention feature (where). Specifically, the prediction of each word is dominated by the concept and spatial attention feature of the corresponding theme in the process of caption prediction. Moreover, we introduce a novel learning method of Distinctive Learning (DL) to get more specificity of generated captions like human descriptions. It formulates two constraints in the theme learning process to encourage distinctiveness between different images. Particularly, reinforcement learning is introduced into the framework to address the exposure bias problem between the training and the testing modes. Extensive experiments on the COCO and Flickr30K datasets achieve superior results when compared with the state-of-the-art methods. Xinyu Xiao, Lingfeng Wang 0002, Shiming Xiang, Chunhong Pan |
AAAI | 1 |
| 2019 | Guiding the Flowing of Semantics: Interpretable Video Captioning via POS TagabstractXinyu Xiao, Lingfeng Wang, Bin Fan, Shinming Xiang, Chunhong Pan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Xinyu Xiao, Lingfeng Wang 0002, Bin Fan 0001, Shiming Xiang, Chunhong Pan |
EMNLP/IJCNLP (1) | 1 |
| 2019 | DetNAS: Backbone Search for Object DetectionabstractObject detectors are usually equipped with backbone networks designed for image classification. It might be sub-optimal because of the gap between the tasks of image classification and object detection. In this work, we present DetNAS to use Neural Architecture Search (NAS) for the design of better backbones for object detection. It is non-trivial because detection training typically needs ImageNetpre-training while NAS systems require accuracies on the target detection task as supervisory signals. Based on the technique of one-shot supernet, which contains all possible networks in the search space, we propose a framework for backbone search on object detection. We train the supernet under the typical detector training schedule: ImageNet pre-training and detection fine-tuning. Then, the architecture search is performed on the trained supernet, using the detection task as the guidance. This framework makes NAS on backbones very efficient. In experiments, we show the effectiveness of DetNAS on various detectors, for instance, one-stage RetinaNetand the two-stage FPN. We empirically find that networks searched on object detection shows consistent superiority compared to those searched on ImageNet classification. The resulting architecture achieves superior performance than hand-crafted networks on COCO with much less FLOPs complexity. Yukang Chen, Tong Yang 0005, Xiangyu Zhang 0005, Gaofeng Meng, Xinyu Xiao, Jian Sun 0001 |
NeurIPS | 5 |
| 2019 | Dense semantic embedding network for image captioning
Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
Pattern Recognit. | 1 |
| 2019 | Deep Hierarchical Encoder-Decoder Network for Image CaptioningabstractEncoder-decoder models have been widely used in image captioning, and most of them are designed via single long short term memory (LSTM). The capacity of single-layer network, whose encoder and decoder are integrated together, is limited for such a complex task of image captioning. Moreover, how to effectively increase the “vertical depth” of encoder-decoder remains to be solved. To deal with these problems, a novel deep hierarchical encoder-decoder network is proposed for image captioning, where a deep hierarchical structure is explored to separate the functions of encoder and decoder. This model is capable of efficiently exerting the representation capacity of deep networks to fuse high level semantics of vision and language in generating captions. Specifically, visual representations in top levels of abstraction are simultaneously considered, and each of these levels is associated to one LSTM. The bottom-most LSTM is applied as the encoder of textual inputs. The application of the middle layer in encoder-decoder is to enhance the decoding ability of top-most LSTM. Furthermore, depending on the introduction of semantic enhancement module of image feature and distribution combine module of text feature, variants of architectures of our model are constructed to explore the impacts and mutual interactions among the visual representation, textual representations, and the output of the middle LSTM layer. Particularly, the framework is training under a reinforcement learning method to address the exposure bias problem between the training and the testing by the policy gradient optimization. Qualitative analyses indicate the process that our model “translates” image to sentence and further visualization presents the evolution of the hidden states from different hierarchical LSTMs over time. Extensive experiments demonstrate that our model outperforms current state-of-the-art models on three benchmark datasets: Flickr8K, Flickr30K, and MSCOCO. On both image captioning and retrieval tasks, our method achieves the best results. On MSCOCO captioning Leaderboard, our method also achieves superior performance. Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan |
IEEE Trans. Multim. | 1 |
| 2017 | PUED: A Social Spammer Detection Method Based on PU Learning and Ensemble Learning
Yuqi Song, Min Gao 0001, Junliang Yu, Wentao Li 0001, Lulan Yu, Xinyu Xiao |
CollaborateCom | 6 |