Yiling Wu

dblp:127/4819 · DBLP profile ↗
← Back
27ranked-venue papers
12as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Asymmetric Frequency-Adaptive State-Space Model for Roadside Cooperative Perception
abstract
Accurate and efficient roadside cooperative perception is crucial for reducing blind spots and extending sensing ranges. However, it faces challenges in modeling long-short range cooperative dependencies and representing the heterogeneous-density distribution of cross-infrastructure data. While CNNs, Transformers, and State-Space Models have demonstrated superior performance, they inherently struggle to balance the flexibility of long-short range receptive fields with computational costs. Additionally, frequency-domain decomposition remains underutilized for heterogeneous-density data representation. In this work, we propose an innovative Asymmetric Multi-Frequency Scale-Adaptive Mamba (AsymMamba) framework, performing lightweight heterogeneous-density data decomposition to support scalable long-short range cooperative representation. First, an Asymmetric Multi-Frequency Decomposition (AsymFreq) module is designed with wavelet transforms, which unifies the spatial distribution representation of heterogeneous-density data in the frequency domain while mitigating information loss through asymmetric scale partitioning. Subsequently, AsymMamba designs a Scale-Adaptive State-Space Model (AdaSSM) module with a spatial compression and channel expansion mechanism. It not only effectively captures local short-range semantic information but also efficiently models global long-range cooperative dependencies with linear complexity. Experiments on real-world DAIR-V2X and RCooper datasets demonstrate that AsymMamba outperforms state-of-the-art methods, including the Transformer-based CoBEVT and recent Mamba-based variants. Specifically, it achieves 3.4%, 4.3%, and 0.6% 3D object detection improvements at [email protected] in vehicle-to-infrastructure cooperation, complex intersection, and long-range corridor roadside cooperative perception scenarios, respectively. Moreover, AsymMamba also achieves superior real-time efficiency with 4x faster inference latency than CoBEVT in a 100m sensing range, and 7x faster in a 200m long-range scenario. Code will available upon acceptance.
Yiling Wu, Mingkai Qiu, Yaowei Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 FreqBEV-V2I: Frequency-Domain BEV-Enhanced Vehicle-to-Infrastructure Cooperative 3D Detection
abstract
Accurate vehicle-to-infrastructure (V2I) cooperation can significantly enhance the perception performance of autonomous vehicles by leveraging information from infrastructure. However, existing cooperation methods based on spatial bird’s-eye-view (BEV) representation struggle with asynchronous temporal misalignment and heterogeneous feature collaboration, leading to 3D detection performance degradation and compromised safety. In this paper, we explore a frequency domain BEV representation to address these challenges and propose the FreqBEV-V2I framework that incorporates FreqBEVFlow and FreqBEVFusion blocks. In FreqBEVFlow, we design global filter spatial differential matching and wavelet-enhanced Fourier channel refinement networks to capture global motion variations via self-supervised learning, which effectively addresses transmission asynchronous latency. Meanwhile, FreqBEVFusion integrates features from vehicle and infrastructure with a frequency adaptive convolution network for V2I heterogeneous feature collaboration. Experimental results on the real-world DAIR-V2X dataset demonstrate that FreqBEV-V2I significantly outperforms current state-of-the-art methods, achieving superior 3D object detection performance and robustness across various latency conditions. Specifically, under ideal V2I communication conditions, FreqBEV-V2I achieves 61.59% mAP@3D (IoU=0.5), surpassing individual no-fusion and existing state-of-the-art methods by 16.77% and 5.78%, respectively. Even in latency-aware scenarios, FreqBEV-V2I maintains high accuracy with an mAP@3D (IoU=0.5) of 61.20% at 200 ms latency, significantly outperforming other methods. The code is available athttps://github.com/DeepPhysicVision/FreqBEV-V2I.
Yiling Wu, Mingkai Qiu, Zhihui Ye
IEEE Trans. Intell. Transp. Syst.2
2026 DMutDE: Dual-View Mutual Distillation Framework for Knowledge Graph Embeddings
abstract
Knowledge graphs (KGs) have caught more and more attention in recent years. Currently, in some practical scenarios, KG embedding (KGE) models are expected to reduce their spatial complexity without losing much performance to address the challenges of storage limitations and knowledge reasoning efficiency. To achieve this, existing works use one or more large and high-performance teacher models to improve the performance of a lightweight student model via knowledge distillation (KD), thus meeting the requirements of some practical complicated applications. However, in resource-constrained scenarios, obtaining high-performance teacher models is challenging due to high training costs and significant storage requirements. Thus, enhancing the student model's performance without large teacher models is crucial. To address this issue, we propose Dual-View Mutual Distillation Framework for Knowledge Graph Embeddings (DMutDE), a distillation framework leveraging mutual learning for peer-to-peer distillation between two KGE models with different architectures. In KGE models, we notice that the way of modeling relational directed edges determines the model view of KGE model for learning KG data. Thus, integrating the model views from two different KGE models by KD into a student KGE model can improve its generalization, so as to increase its performance. To identify an effective dual-view fusion method, we design two modules in the DMutDE framework. Specifically, we design a novel soft-label fusion (SLF) module for noise filtering and response knowledge transfer. Then, we propose an entity embedding distillation (EED) module to distill structural features from each other. Finally, we conduct several comprehensive experiments on the standard open-source benchmarks to demonstrate that our framework achieves the state-of-the-art results. The code is available at https://github.com/RuizhouLiu/DMutDE.
Ruizhou Liu, Zhe Wu 0006, Yiling Wu, Zongsheng Cao, Qianqian Xu 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.3
2025 Mixture-of-KAN for Multivariate Time Series Forecasting
abstract
Multivariate time series forecasting is a crucial task that predicts the future states based on historical inputs. Although current deep learning-based methods have made significant advancements, they still face the criticism of lacking interpretability. The rise of the Kolmogorov-Arnold Network (KAN) provides a new perspective to implement an efficient and interpretable deep learning-based method for forecasting time series. However, we find there are two main challenges in the application of KAN in time series forecasting: how to select the appropriate one from various KAN variants and how to train the deep KAN-based network. To this end, we propose the multi-layer mixture-of-KAN network, which achieves excellent performance while retaining KAN's ability to be transformed into a combination of symbolic functions. The core module is the mixture-of-KAN layer, which uses a mixture-of-experts structure to assign variables to best-matched KAN experts. Then, we analyze the shortcomings of parameter initialization in the original KAN and provide an effective initialization method to alleviate training instability. Extensive experimental results demonstrate that our proposed method is effective in multivariate time series forecasting. Codes are released in https://github.com/2448845600/EasyTSF.
Zhenduo Zhang, Xinfeng Zhang 0001, Yiling Wu, Zhe Wu 0006
CIKM4
2025 CDFNet: Collaborative Decomposition and Forecasting Network for Time Series
Zhenduo Zhang, Yiling Wu, Xinfeng Zhang 0001, Qingming Huang
ICIC (7)3
2025 Multimodal-LLM Agent For Text-Driven Multi-Attribute Face Editing
abstract
Facial attribute editing aims to manipulate specific attributes of a face image while the other attributes are not affected. The development of models like StyleGAN and CLIP has significantly advanced text-driven facial attribute editing. However, single models frequently struggle with complex textual descriptions involving multiple attributes, and the absence of verification and self-correction mechanisms, making the results unreliable. To address these challenges, we propose MMFE, a text-driven multi-attribute facial editing system. MMFE integrates various facial editing models and leverages a large language model (LLM) as an agent to select appropriate models based on natural language requests. Finally, a multimodal large language model (MLLM) is introduced for result verification. Our experiments demonstrate that MMFE significantly improves performance in text-driven multi-attribute facial editing tasks.
Lang Yue, Yiling Wu
ICIP2
2025 Learning Fine-Grained Representations through Textual Token Disentanglement in Composed Video Retrieval
abstract
With the explosive growth of video data, finding videos that meet detailed requirements in large datasets has become a challenge. To address this, the composed video retrieval task has been introduced, enabling users to retrieve videos using complex queries that involve both visual and textual information. However, the inherent heterogeneity between the modalities poses significant challenges. Textual data are highly abstract, while video content contains substantial redundancy. The modality gap in information representation makes existing methods struggle with the modality fusion and alignment required for fine-grained composed retrieval. To overcome these challenges, we first introduce FineCVR-1M, a fine-grained composed video retrieval dataset containing 1,010,071 video-text triplets with detailed textual descriptions. This dataset is constructed through an automated process that identifies key concept changes between video pairs to generate textual descriptions for both static and action concepts. For fine-grained retrieval methods, the key challenge lies in understanding the detailed requirements. Text description serves as clear expressions of intent, but it requires models to distinguish subtle differences in the description of video semantics. Therefore, we propose a textual Feature Disentanglement and Cross-modal Alignment framework (FDCA) that disentangles features at both the sentence and token levels. At the sequence level, we separate text features into retained and injected features. At the token level, an Auxiliary Token Disentangling mechanism is proposed to disentangle texts into retained, injected, and excluded tokens. The disentanglement at both levels extracts fine-grained features, which are aligned and fused with the reference video to extract global representations for video retrieval. Experiments on FineCVR-1M dataset demonstrate the superior performance of FDCA. Our code and dataset are available at: https://may2333.github.io/FineCVR/.
Zhaobo Qi, Yiling Wu, Junshu Sun, Yaowei Wang 0001, Shuhui Wang
ICLR3
2025 Probabilistic Decision-Making for Virtually Coupled Trains Under Uncertainty
abstract
The virtually coupled trains present challenges due to the risks associated with small separation. The following train must make appropriate decisions in uncertain and high-risk situations to ensure safety and efficiency. We address uncertainty from sensor noise and unknown preceding train behaviors, establishing a probabilistic decision-making model using a partially observable Markov decision process (POMDP). Collision and risk envelopes are defined using the emergency and full-service braking distance of the following train, considering the “worst-case separation” scenarios, leading to a novel chance-constrained POMDP (CC-POMDP) for the virtually coupled trains. We propose an online planning algorithm, Partially Observable Monte Carlo Planning based on Double Shielding and Progressive Widening (POMCP-DS-PW). The results demonstrate that even without a pre-planned recommended speed curve, the following train can optimize acceleration based on incomplete observations and unknown behaviors of the preceding train, thereby ensuring safety while achieving operational efficiency.
Yiling Wu
IEEE Trans. Intell. Transp. Syst.1
2025 SGKGE: Semantically Guided Knowledge Graph Embeddings via Complementary Latent Representations
abstract
Knowledge graph (KG) completion is a challenging yet essential task that has attracted increasing attention in recent years. While entities in KGs typically present complex semantics (a phenomenon known as polysemy), previous works primarily focus on holistic but often inaccurate representations of entities, neglecting the diversity of their semantics. This limitation results in suboptimal representations for entities within KGs. To address this issue, we propose a new method termed semantically guided KG embeddings (SGKGE), which captures the precise semantics of entities in KGs from a semantics-guided perspective. Specifically, SGKGE first guides the learning of holistic semantics of entities through a hyperbolic manifold with learnable shared curvature and a geometric attention-fusion module, facilitating efficient reasoning. Subsequently, SGKGE captures fine-grained semantics through a set of Cartesian product Riemannian manifolds with distinct curvatures, coupled with a semantic interactions module. This approach enables SGKGE to produce more accurate entity semantics and enhance downstream applications. Experimental results demonstrate that our model achieves state-of-the-art performance on six well-established KG completion benchmarks. The release code is available at https://github.com/RuizhouLiu/SGKGE.
Ruizhou Liu, Zongsheng Cao, Zhe Wu 0006, Yiling Wu, Qianqian Xu 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.4
2024 Event Traffic Forecasting with Sparse Multimodal Data
abstract
With the development of deep learning, traffic forecasting technology has made significant progress and is being applied in many practical scenarios. However, various events held in cities, such as sporting events, exhibitions, concerts, etc., have a significant impact on traffic patterns of surrounding areas, causing current advanced prediction models to fail in this case. In this paper, to broaden the applicable scenarios of traffic forecasting, we focus on modeling the impact of events on traffic patterns and propose an event traffic forecasting problem with multimodal inputs. We outline the main challenges of this problem: diversity and sparsity of events, as well as insufficient data. To address these issues, we first use textual modal data containing rich semantics to describe the diverse characteristics of events. Then, we propose a simple yet effective multi-modal event traffic forecasting model that uses pre-trained text and traffic encoders to extract the embeddings and fuses the two embeddings for prediction. Encoders pre-trained on large-scale data have powerful generalization abilities to cope with the challenge of sparse data. Next, we design an efficient large language model-based event description text generation pipeline to build multi-modal event traffic forecasting datasets, ShenzhenCEC and SuzhouIEC. Experiments on two real-world datasets show that our method achieves state-of-the-art performance compared with eight baselines, reducing mean absolute error during the event peak period by 4.26%. Code is available at: https://github.com/2448845600/EventTrafficForecasting.
Zhenduo Zhang, Yiling Wu, Xinfeng Zhang 0001, Zhe Wu 0006
ACM Multimedia3
2024 MLFA: Toward Realistic Test Time Adaptive Object Detection by Multi-Level Feature Alignment
abstract
Object detection methods have achieved remarkable performances when the training and testing data satisfy the assumption of i.i.d. However, the training and testing data may be collected from different domains, and the gap between the domains can significantly degrade the detectors. Test Time Adaptive Object Detection (TTA-OD) is a novel online approach that aims to adapt detectors quickly and make predictions during the testing procedure. TTA-OD is more realistic than the existing unsupervised domain adaptation and source-free unsupervised domain adaptation approaches. For example, self-driving cars need to improve their perception of new environments in the TTA-OD paradigm during driving. To address this, we propose a multi-level feature alignment (MLFA) method for TTA-OD, which is able to adapt the model online based on the steaming target domain data. For a more straightforward adaptation, we select informative foreground and background features from image feature maps and capture their distributions using probabilistic models. Our approach includes: i) global-level feature alignment to align all informative feature distributions, thereby encouraging detectors to extract domain-invariant features, and ii) cluster-level feature alignment to match feature distributions for each category cluster across different domains. Through the multi-level alignment, we can prompt detectors to extract domain-invariant features, as well as align the category-specific components of image features from distinct domains. We conduct extensive experiments to verify the effectiveness of our proposed method. Our code is accessible at https://github.com/yaboliudotug/MLFA.
Chao Huang 0008, Yiling Wu, Yong Xu 0001, Xiaochun Cao
IEEE Trans. Image Process.4
2024 Knowledge-Based Multiple Relations Modeling for Traffic Forecasting
abstract
Traffic forecasting is a critical task in intelligent transportation systems. In recent years, lots of methods have been proposed and achieved significant progress in modeling highly nonlinear and complex spatiotemporal pattern for traffic forecasting. However, most methods neglect the specific internal and external factors of the traffic system, such as the road connections, buildings surrounding each place, transfer stations, etc. The main challenges of utilizing the diverse knowledge from internal and external factors are to represent and fuse the impact of various factors. Few works use distance-based adjacency matrices to represent factors and the element-wise multiplication to fuse them, which may lead to even worse performances. In this paper, we propose to utilize knowledge graph to represent traffic system factors and design a novel neural network to exploit them for traffic foresting. First, we model the relations between factors and traffic conditions from the perspective of the knowledge graph and express them as unified triplets. Then, we generate multi-hop path features with embeddings learned from the knowledge representation model and multi-hop paths searched from the graph structure. Next, we present a knowledge-based multi-hop network (KMHNet) that uses an attention-based module to learn the correlation from multi-hop path features. Finally, to evaluate the performance of the proposed method, we build two real-world datasets both containing a traffic condition sub-dataset and a traffic knowledge graph. Experiments on two datasets demonstrate that our proposed KMHNet outperforms eight well-known methods. The code is publicly available at https://github.com/2448845600/KMHNet.
Xinfeng Zhang 0001, Yiling Wu, Zhenduo Zhang, Yaowei Wang 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Spatial-Temporal Correlation Learning for Traffic Demand Prediction
abstract
Traffic demand prediction has been drawing increasing research interest due to its critical role in intelligent transportation systems. However, conventional deep learning methods for traffic demand forecast ignore the correlations between the pick-up and drop-off demands, thus not fully exploring the patterns of demand evolution. In this work, the pick-up and drop-off demands are treated as two modalities, and an architecture is designed to explicitly model the interactions between the pick-up and drop-off demands both spatially and temporally. Specifically, the self-attention mechanism is adopted to automatically discover spatio-temporal patterns without manual designation for each demand. Then, the cross-attention mechanism is utilized to let the two demands attend to each other, resulting in information exchange between the two demands. The self-attention and cross-attention are combined to capture spatio-temporal correlations simultaneously. Finally, experiments are carried out on three real-world datasets, NYC Citi Bike, NYC Taxi, and BJ Subway, and the results show that this newly proposed method outperforms the state-of-the-art methods.
Yiling Wu, Yingping Zhao, Xinfeng Zhang 0001, Yaowei Wang 0001
IEEE Trans. Intell. Transp. Syst.1
2023 Adaptive Graph Neural Diffusion for Traffic Demand Forecasting
abstract
This paper studies the problem of spatial-temporal modeling for traffic demand forecasting. In practice, the temporal-spatial dependencies are complex. Conventional methods using graph convolutional networks and gated recurrent units cannot fully explore the patterns of demand evolution. Therefore, we propose Adaptive Graph Neural Diffusion (AGND) for spatial-temporal graph modeling. Specifically, complex spatial relations are modeled with a diffusion process by the graph neural diffusion. The spatial attention mechanism and a data-driven semantic adjacency matrix are used to describe the diffusivity function in the graph neural diffusion, which provides both local and global spatial information. Long-term temporal dependencies are modeled by the temporal attention mechanism. The proposed method is applied to two real-world datasets, and the results show that the proposed method outperforms state-of-the-art methods.
Yiling Wu, Xinfeng Zhang 0001, Yaowei Wang 0001
CIKM1
2022 Span-based Audio-Visual Localization
abstract
This paper focuses on the audio-visual event localization task that aims to match both visible and audible components in a video to identify the event of interest. Existing methods primarily ignore the continuity of audio-visual events and classify each segment separately. They either classify the event category score of each segment separately or calculate the event-relevant score of each segment separately. However, events in video are often continuous and last several segments. Motivated by these, we propose a span-based framework that considers consecutive segments jointly. The span-based framework handles the audio-visual localization task by predicting the event class and extracting the event span. Specifically, a [CLS] token is applied to collect the global information with self-attention mechanisms to predict the event class. Relevance scores and positional embeddings are inserted into the span predictor to estimate the start and end boundaries of the event. Multi-modal Mixup are further used to improve the robustness and generalization of the model. Experiments conducted on the AVE dataset demonstrate that the proposed method outperforms state-of-the-art methods.
Yiling Wu, Xinfeng Zhang 0001, Yaowei Wang 0001, Qingming Huang
ACM Multimedia1
2021 Talking Face Generation Based on Information Bottleneck and Complementary Representations
abstract
Audio-driven talking face generation is an active research direction in the field of virtual reality. The main challenge is that the generated lip shape of the speaker is out of sync with the input audio. To address this challenge, we propose a novel solution to synthesize lip-synchronized, high-quality, realistic video given input audio. We first decompose the target person's video frames into 3D face model parameters, and the information bottleneck is inserted into the audio-to-expression network to learn the mapping between audio features and expression parameters. Then, we replace the expression parameters in the target video frame with the extracted expression parameters from audio and re-render the face. Finally, we add high-level audio embedding extracted from the raw audio and lip landmarks embedding in the neural rendering network. The 3D face shapes, 2D landmarks, and audio embedding provide complementary information for the neural rendering network which guarantees the generation of lip-synchronized high-quality video portraits from the synthesized rendered faces. Experimental results show that compared with other talking face generation methods, our method is the best concerning lip synchronization with high video definition.
Yiling Wu, Minglei Li 0001
CIKM2
2021 VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and Adaptation
abstract
For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The representations of existing PAD works with simple global pooling method, however, lose the local feature discriminability. In this paper, the VLAD aggregation method is adopted to quantize local features with visual vocabulary locally partitioning the feature space, and hence preserve the local discriminability. We further propose the vocabulary separation and adaptation method to modify VLAD for cross-domain PAD task. The proposed vocabulary separation method divides vocabulary into domain-shared and domain-specific visual words to cope with the diversity of live and attack faces under the cross-domain scenario.The proposed vocabulary adaptation method imitates the maximization step of the k-means algorithm in the end-to-end training, which guarantees the visual words be close to the center of assigned local features and thus brings robust similarity measurement. We give illustrations and extensive experiments to demonstrate the effectiveness of VLAD with the proposed vocabulary separation and adaptation method on standard cross-domain PAD benchmarks. The codes are available at https://github.com/Liubinggunzu/VLAD-VSA.
Zhou Zhao 0001, Weike Jin, Xinyu Duan, Zhen Lei 0001, Baoxing Huai, Yiling Wu, Xiaofei He 0001
ACM Multimedia7
2021 Combining cross-modal knowledge transfer and semi-supervised learning for speech emotion recognition
Min Chen 0003, Jincai Chen, Yuan-Fang Li, Yiling Wu, Minglei Li 0001, Chuanbo Zhu 0002
Knowl. Based Syst.5
2021 Augmented Adversarial Training for Cross-Modal Retrieval
abstract
Cross-modal retrieval has received considerable attention in recent years. The core of cross-modal retrieval is to find a representation space to align data from different modalities according to their semantics. In this paper, we propose a cross-modal retrieval method that aligns data from different modalities by transferring one source modality to another target modality with augmented adversarial training. To preserve the semantic meaning in the modality transfer process, we employ the idea of conditional GANs and augment it. The key idea is to incorporate semantic information from the label space into the adversarial training process by sampling more semantic relevant and irrelevant source-target sample pairs. The augmented sample pairs improve the alignment from two aspects. First, relevant source-target sample pairs provide more training samples, leading to a better guidance of the alignment of fake targets and true paired targets. Second, relevant and irrelevant source-target sample pairs teach the discriminator to better distinguish true relevant pairs from fake relevant pairs, which guides the generator to better transfer from the source modality to the target modality. Extensive experiments compared with state-of-the-art methods show the promising power of our approach.
Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang
IEEE Trans. Multim.1
2020 Online Fast Adaptive Low-Rank Similarity Learning for Cross-Modal Retrieval
abstract
The semantic similarity among cross-modal data objects, e.g., similarities between images and texts, are recognized as the bottleneck of cross-modal retrieval. However, existing batch-style correlation learning methods suffer from prohibitive time complexity and extra memory consumption in handling large-scale high dimensional cross-modal data. In this paper, we propose a Cross-Modal Online Low-Rank Similarity function learning (CMOLRS) method, which learns a low-rank bilinear similarity measurement for cross-modal retrieval. We model the cross-modal relations by relative similarities on the training data triplets and formulate the relative relations as convex hinge loss. By adapting the margin in hinge loss with pair-wise distances in feature space and label space, CMOLRS effectively captures the multi-level semantic correlation and adapts to the content divergence among cross-modal data. Imposed with a low-rank constraint, the similarity function is trained by online learning in the manifold of low-rank matrices. The low-rank constraint not only endows the model learning process with faster speed and better scalability, but also improves the model generality. We further propose fast-CMOLRS combining multiple triplets for each query instead of standard process using single triplet at each model update step, which further reduces the times of gradient updates and retractions. Extensive experiments are conducted on four public datasets, and comparisons with state-of-the-art methods show the effectiveness and efficiency of our approach.
Yiling Wu, Shuhui Wang, Qingming Huang
IEEE Trans. Multim.1
2019 Learning Fragment Self-Attention Embeddings for Image-Text Matching
abstract
In image-text matching task, the key to good matching quality is to capture the rich contextual dependencies between fragments of image and text. However, previous works either simply aggregate the similarity of all possible pairs of image regions and words, or take multi-step cross attention to attend to image regions and words with each other as context, which requires exhaustive similarity computation between all image region and word pairs. In this paper, we propose Self-Attention Embeddings (SAEM) to exploit fragment relations in images or texts by self-attention mechanism, and aggregate fragment information into visual and textual embeddings. Specifically, SAEM extracts salient image regions based on bottom-up attention, and takes WordPiece tokens as sentence fragments. The self-attention layers are built to model subtle and fine-grained fragment relation in image and text respectively, which consists of multi-head self-attention sub-layer and position-wise feed-forward network sub-layer. Consequently, the fragment self-attention mechanism can discover the fragment relations and identify the semantically salient regions in images or words in sentences, and capture their interaction more accurately. By simultaneously exploiting the fine-grained fragment relation in both visual and textual modalities, our method produces more semantically consistent embeddings for representing images and texts, and demonstrates promising image-text matching accuracy and high efficiency on Flickr30K and MSCOCO datasets.
Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang
ACM Multimedia1
2019 Multi-modal semantic autoencoder for cross-modal retrieval
Yiling Wu, Shuhui Wang, Qingming Huang
Neurocomputing1
2019 Online Asymmetric Metric Learning With Multi-Layer Similarity Aggregation for Cross-Modal Retrieval
abstract
Cross-modal retrieval has attracted intensive attention in recent years, where a substantial yet challenging problem is how to measure the similarity between heterogeneous data modalities. Despite using modality-specific representation learning techniques, most existing shallow or deep models treat different modalities equally and neglect the intrinsic modality heterogeneity and information imbalance among images and texts. In this paper, we propose an online similarity function learning framework to learn the metric that can well reflect the cross-modal semantic relation. Considering that multiple CNN feature layers naturally represent visual information from low-level visual patterns to high-level semantic abstraction, we propose a new asymmetric image-text similarity formulation which aggregates the layer-wise visual-textual similarities parameterized by different bilinear parameter matrices. To effectively learn the aggregated similarity function, we develop three different similarity combination strategies, i.e., average kernel, multiple kernel learning, and layer gating. The former two kernel-based strategies assign uniform weights on different layers to all data pairs; the latter works on the original feature representation and assigns instance-aware weights on different layers to different data pairs, and they are all learned by preserving the bi-directional relative similarity expressed by a large number of cross-modal training triplets. The experiments conducted on three public datasets well demonstrate the effectiveness of our methods.
Yiling Wu, Shuhui Wang, Guoli Song, Qingming Huang
IEEE Trans. Image Process.1
2018 Learning Semantic Structure-preserved Embeddings for Cross-modal Retrieval
abstract
This paper learns semantic embeddings for multi-label cross-modal retrieval. Our method exploits the structure in semantics represented by label vectors to guide the learning of embeddings. First, we construct a semantic graph based on label vectors which incorporates data from both modalities, and enforce the embeddings to preserve the local structure of this semantic graph. Second, we enforce the embeddings to well reconstruct the labels, i.e., the global semantic structure. In addition, we encourage the embeddings to preserve local geometric structure of each modality. Accordingly, the local and global semantic structure consistencies as well as the local geometric structure consistency are enforced, simultaneously. The mappings between inputs and embeddings are designed to be nonlinear neural network with larger capacity and more flexibility. The overall objective function is optimized by stochastic gradient descent to gain the scalability on large datasets. Experiments conducted on three real world datasets clearly demonstrate the superiority of our proposed approach over the state-of-the-art methods.
Yiling Wu, Shuhui Wang, Qingming Huang
ACM Multimedia1
2017 Online Asymmetric Similarity Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval has attracted intensive attention in recent years. Measuring the semantic similarity between heterogeneous data objects is an essential yet challenging problem in cross-modal retrieval. In this paper, we propose an online learning method to learn the similarity function between heterogeneous modalities by preserving the relative similarity in the training data, which is modeled as a set of bi-directional hinge loss constraints on the cross-modal training triplets. The overall online similarity function learning problem is optimized by the margin based Passive-Aggressive algorithm. We further extend the approach to learn similarity function in reproducing kernel Hilbert spaces by kernelizing the approach and combining multiple kernels derived from different layers of the CNN features using the Hedging algorithm. Theoretical mistake bounds are given for our methods. Experiments conducted on real world datasets well demonstrate the effectiveness of our methods.
Yiling Wu, Shuhui Wang, Qingming Huang
CVPR1
2017 Online low-rank similarity function learning with adaptive relative margin for cross-modal retrieval
abstract
This paper presents a Cross-Modal Online Low-Rank Similarity function learning method (CMOLRS) for cross-modal retrieval, which learns a low-rank bilinear similarity measure on data from different modalities. CMOLRS models the cross-modal relations by relative similarities on a set of training data triplets and formulates the relative relations as convex hinge loss functions. By adapting the margin of hinge loss using information from feature space and label space for each triplet, CMOLRS effectively captures the multi-level semantic correlation among cross-modal data. The similarity function is learned by online learning in the manifold of low-rank matrices, thus good scalability is gained when processing large scale datasets. Extensive experiments are conducted on three public datasets. Comparisons with the state-of-the-art methods show the effectiveness and efficiency of our approach.
Yiling Wu, Shuhui Wang, Weigang Zhang, Qingming Huang
ICME1
2015 Improving cross-modal correlation learning with hyperlinks
abstract
We propose a new cross-modal correlation learning framework which boosts the performance of correlation learning models using the hyperlink information. First, we design a neighborhood selection paradigm using the hyperlink structure and content similarities to identify a set of semantically related documents for each multi-modal document in both training and testing stage. Based on the neighborhood structure, we revise two well-established content-based correlation learning models, i.e., canonical correlation analysis (CCA) and kernel canonical correlation analysis (KCCA) with a structure coding matrix. Third, we develop a correlation score aggregation technique to discover more semantically relevant cross-modal documents. To our best knowledge, this is the first to introduce hyperlink information into cross-modal correlation learning. Experimental results demonstrate that our proposed framework can significantly improve the model generality towards real-world cross-modal retrieval.
Shuhui Wang, Yiling Wu, Qingming Huang
ICME2