Yi Yu 0001

dblp:99/111-1 · DBLP profile ↗
← Back
125ranked-venue papers
17as first author
66since 2021 · last 2026
0000-0002-0294-6620ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 94 · 14 first-author · 50 since 2021Artificial intelligence and machine learning · 28 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 4 since 2021Computer networks · 10 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion Model
abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). Recently, several SSME methods have been proposed and achieved remarkable successes. However, existing methods are still facing two critical issues: firstly, there is a lack of an effective data augmentation method for SSME, which results in insufficient utilization of unlabeled data. Secondly, existing SSME methods discards too much unlabeled data in the stage of consistency regularization, which hinders the further improvements of SSME task. In this paper, we present \emph{ELH-SME}, a novel framework that better utilizes the unlabeled musical data for SSME task. Specifically, our proposed ELH-SME framework consists of three modules: (1) we first propose a diffusion-based multi-bands augmentation (DMA) method to increase the amounts of training data. The proposed DMA methods employs a diffusion model to generate perturbation at the specific frequency bands in an end-to-end manner, thereby avoiding sharply perturbations to the spectrogram. (2) To improve the utilization rate of unlabeled data, we suggest a global-class confidence (GCC) module. During the phase of consistency regularization, we consider both the global-wise and class-wise confidence values, improving the utilization rate of unlabeled data. (3) To further improve the utilization of unlabeled data, we also propose to enhance the representation capability of unlabeled data by extracting channel-level features from labeled data via channel cross attention (CCA). We evaluate our proposed framework on several well-known public available datasets, and the conducted experiments demonstrate the effectiveness of our method.
Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001
AAAI4
2026 Towards boundary confusion for volumetric medical image segmentation
Xin You 0002, Junyang Wu, Yi Yu 0001, Jie Yang 0002, Yun Gu
Medical Image Anal.6
2025 Dialogue-AV: A Dialogue-Attended Audiovisual Dataset
abstract
This work introduces Dialogue-AV, a benchmarking dataset for Audio-Video-Language (AVL). We propose using dialogue to describe video content instead of single captions, capturing nuances and shared meanings between audio and visual elements. This approach contributes significantly to improving the diversity of video descriptions and enables comprehensive evaluation of AVL learning across different downstream tasks, such as Cross-Modal Retrieval, Visual Question-Answering, and Video Captioning. Our dataset comprises approximately 258k audiovisual samples accompanied by dialogue-based descriptions for benchmarking. Dialogue-AV builds upon existing State-of-the-Art (SOTA) datasets that feature human-generated descriptions, enhancing them with model-generated ones that describe all modalities. We also present zero-shot baseline results utilising SOTA Visual-Language Models (VLMs), demonstrating that Dialogue-AV is capable of benchmarking a variety of downstream tasks with diverse inputs. Our key contributions include: 1) Dialogue-AV, a benchmark dataset for dialogue-based AVL models; and 2) benchmarks that expose the limitations of current SOTA VLMs. The code and dataset are accessible at: github.com/lvilaca16/dialogue-av.
Luís Vilaça, Paula Viana, Yi Yu 0001
CBMI3
2025 Chain-of-Thought Prompting with Causal Intervention for Multimodal Aspect-Based Sentiment Analysis
Zhuopan Yang, Haoran Xie 0001, Lap-Kei Lee, Fu Lee Wang, Yi Yu 0001, Zhenguo Yang
DASFAA (2)6
2025 A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
abstract
Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequency-supervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method1.
Xiaoliang He, Kangjie Dong, Jingkai Cao, Shuai Yu 0002, Wei Li 0012, Yi Yu 0001
ICASSP6
2025 Enhancing Video-Text Matching via Sparse Stratified Sampling
abstract
Video-text matching is a critical task in multimedia retrieval, but traditional methods often fail to capture the diversity and depth of video content due to inefficient and inaccurate frame sampling. We propose a novel sparse stratified sampling technique that can substantially improve the video-text matching process by segmenting video content into clusters based on relevant features and selectively sampling representative frames. Our method further introduces a threshold for the feature metric used to divide clusters, eliminating video frames with low relevance. We propose two variants of our approach: an offline approach that performs sampling before training, and an online approach that dynamically conducts sampling based on the relevance between video frames and the text query during training. Extensive experiments on datasets like MSRVTT and AVSD for video retrieval and multiple-choice VideoQA datasets, including AVQA and Music-AVQA, demonstrate the superiority of our method over previous state-of-the-art approaches. Our sparse stratified sampling technique achieves improvements of over 1.2% on MSRVTT and 1.7% on AVSD for R@1 in video retrieval tasks. For multiple-choice VideoQA tasks, our approach achieves significant improvements of 1.8% accuracy on AVQA and 3.9% on Music-AVQA, strongly supporting its effectiveness in enhancing video-text matching systems.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Liting Zhou, Pintu Lohar, Yi Yu 0001, Longyue Wang
ICASSP6
2025 Visual Entity-Centric Prompting for Knowledge Retrieval in Knowledge-based VQA
abstract
External knowledge provides critical clues for knowledge-based visual question answering (KB-VQA), while the implicit knowledge in images is difficult to capture in order to construct effective queries for knowledge bases. To this end, we propose a visual entity-centric prompting for knowledge retrieval (VEPR) to bridge the gap between the implicit and explicit knowledge driven by visual entities via large language models. More specifically, a visual entity question answering (EQ) module is devised to localize the critical entities from the given images and questions to generate entity-centric questions via a large language model. In particular, EQ obtains several entity-centric question-answer pairs via a visual language model. Furthermore, a question-answer-enhanced retrieval (ER) module is devised to construct a query by summarizing the text including question-answer pairs, captions and questions, in order to require explicit knowledge items. Finally, a multi-branch reader (MR) module is designed to encode the given questions, visual content and retrieved knowledge items, which are decoded to make answer predictions. Extensive experiments conducted on two public datasets demonstrate the effectiveness of the VEPR.
Jiuxiang You, Ziyue Qiu, Guobo Xie, Yi Yu 0001, Zhenguo Yang
ICASSP5
2025 KCE-Unet: A novel music denoising method with KANConv ECA Unet
abstract
During concerts, people often spontaneously record memorable moments with their phones. However, these recordings are frequently accompanied by noise, such as cheering and applause, which diminishes the playback experience. In this paper, we introduce a novel task specifically designed for denoising music in concert environments, a challenge that has been largely overlooked in previous research. To support this task, we created a new concert denoising dataset that includes songs performed in various major languages at concerts, with noise segments like cheering and applause. Building on this, we propose KANConv ECA Unet (KCE-Unet), a method that combines the U-Net network, efficient channel attention (ECA), and the recently proposed KAN network to flexibly remove noise in the mid-to-high frequency range of spectrograms. Extensive experiments demonstrate that our method outperforms previous models in denoising performance and effectively restore disrupted musical structures.
Yulun Wu 0002, Ganghui Ru, Yi Yu 0001, Wei Li 0012
ICASSP4
2025 HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking
abstract
Fine-tuning pre-trained foundation models has made significant progress in music information retrieval. However, applying these models to beat tracking tasks remains unexplored as the limited annotated data renders conventional fine-tuning methods ineffective. To address this challenge, we propose HingeNet, a novel and general parameter-efficient fine-tuning method specifically designed for beat tracking tasks. HingeNet is a lightweight and separable network, visually resembling a hinge, designed to tightly interface with pre-trained foundation models by using their intermediate feature representations as input. This unique architecture grants HingeNet broad generalizability, enabling effective integration with various pre-trained foundation models. Furthermore, considering the significance of harmonics in beat tracking, we introduce harmonic-aware mechanism during the fine-tuning process to better capture and emphasize the harmonic structures in musical signals. Experiments on benchmark datasets demonstrate that HingeNet achieves state-of-the-art performance in beat and downbeat tracking.
Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012
ICME5
2025 BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
abstract
Beat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles and accurately capture complex rhythmic structures. To overcome these challenges, we propose a novel beat tracking paradigm BeatFM, which introduces a pre-trained music foundation model and leverages its rich semantic knowledge to improve beat tracking performance. Pre-training on diverse music datasets endows music foundation models with a robust understanding of music, thereby effectively addressing these challenges. To further adapt it for beat tracking, we design a plug-and-play multi-dimensional semantic aggregation module, which is composed of three parallel sub-modules, each focusing on semantic aggregation in the temporal, frequency, and channel domains, respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance in beat and downbeat tracking across multiple benchmark datasets.
Ganghui Ru, Jieying Wang, Yulun Wu 0002, Yi Yu 0001, Nannan Jiang, Wei Wang 0414, Wei Li 0012
ICME5
2025 Scene-guided Attention Network for Spatial Understanding in 3D Scenes
abstract
3D Visual Question Answering (3D-VQA) aims to understand the positional relationship, object attributes and layout in 3D scenes. A challenging issue is to align the representations between the entities associated with questions and the relevant 3D objects while diminishing the fine-grained vision-language relations. To this end, we propose a Scene-guided Attention Network for Spatial Understanding, denoted as SceSU, to perceive spatial information and attribute information based on different types of questions in 3D scenes. More specifically, a Scene-Driven Spatial Understanding (SDSU) mechanism is designed to obtain the crucial entities relevant to the question, and construct the fine-grained scene description via a large language model. Furthermore, a 3D Perception Attention (3D-PA) module is devised to fuse natural language and 3D features while understanding the detailed relationship between them by employing a dual-branch attention network. Finally, SceSU utilizes the 3D-PA to comprehend the fine-grained scene description generated by the SDSU mechanism, bridging the gap between natural language and 3D domains. Extensive experiments conducted on SQA3D and ScanQA datasets demonstrate the effectiveness of the SceSU for 3D-VQA.
Yunqi Jiang, Chaoyang Lin, Yi Yu 0001, Zhenguo Yang
ICMR4
2025 DUDA: A Two-stage Decoupling Unsupervised Domain Adaptation Framework for Semi-supervised Singing Melody Extraction from Polyphonic Music
abstract
Semi-supervised singing melody extraction (SSME) is one of the key tasks in the field of music information retrieval (MIR). However, there are two critical issues that remain to be addressed in data limited scenarios. Firstly, the prior unsupervised domain adaptation methods for SSME typically rely on learning domain-agnostic features at holistic level, which ignores the associations between holistic information (i.e., fundamental frequency) and fine-grained information (i.e., tone and octave). Secondly, the fine-grained information can be utilized to judge the availability of unlabeled data, which is ignored by prior methods. There is a lack of a consistency regularization method that utilizes fine-grained information to validate the availability of unlabeled data. To address these issues, in this paper, we propose a novel two-stage decoupling unsupervised domain adaptation framework for semi-supervised singing melody extraction, termed as DUDA. Specifically, in the first stage, we decouple the holistic information into fine-grained information: tone and octave, and narrow the domain gap at the tone and octave level, respectively. This enables the model to align the tone-octave information between source and target domains for better feature distribution. Then, we leverage the learned domain-agnostic fine-grained features as additional information to obtain domain-agnostic holistic features. We also suggest to align intra-domain, inter-domain, and sample-level features to further improve the performances. In the second stage, we propose a novel tone-octave consistency regularization method by leveraging the extracted fine-grained information to judge the availability of unlabeled data. We evaluate our proposed framework on several well-known public datasets, and the conducted experiments demonstrate the effectiveness of our method.
Shuai Yu 0002, Xiaoliang He, Kangjie Dong, Yi Yu 0001
ACM Multimedia4
2025 Enhancing semantic audio-visual representation learning with supervised multi-scale attention
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Guo-Jun Qi, Haiyuan Wu, Hirotaka Hachiya
Pattern Anal. Appl.2
2025 Semantic Frame Aggregation-Based Transformer for Live Video Comment Generation
Anam Fatima, Yi Yu 0001, Janak Kapuriya, Julien Lalanne, Jainendra Shukla
IEEE Trans. Multim.2
2024 Scalable Motion Style Transfer with Constrained Diffusion Generation
abstract
Current training of motion style transfer systems relies on consistency losses across style domains to preserve contents, hindering its scalable application to a large number of domains and private data. Recent image transfer works show the potential of independent training on each domain by leveraging implicit bridging between diffusion models, with the content preservation, however, limited to simple data patterns. We address this by imposing biased sampling in backward diffusion while maintaining the domain independence in the training stage. We construct the bias from the source domain keyframes and apply them as the gradient of content constraints, yielding a framework with keyframe manifold constraint gradients (KMCGs). Our validation demonstrates the success of training separate models to transfer between as many as ten dance motion styles. Comprehensive experiments find a significant improvement in preserving motion contents in comparison to baseline and ablative diffusion-based style transfer models. In addition, we perform a human study for a subjective assessment of the quality of generated dance motions. The results validate the competitiveness of KMCGs.
Yi Yu 0001, Hang Yin 0001, Danica Kragic, Mårten Björkman
AAAI2
2024 Semantic Enrichment for Video Question Answering with Gated Graph Neural Networks
abstract
Video Question Answering (VideoQA) is a complex task that requires a deep understanding of a video to accurately answer questions. Existing methods often struggle to effectively integrate the visual and language-based semantic information, subsequently leading to an incomplete understanding of video content and sub-optimal performance. To address the challenge, we introduce a novel approach in this paper to enrich the semantics of video frames, questions, and answer candidates. Specifically, we parse video frames and questions into semantic graphs - visual semantic graph and question semantic graph, which captures information about objects, their attributes, and relationships. These graphs are then encoded using a Gated Graph Neural Network (GGNN). For answer candidates, we propose to verbalize them using Large Language Models (LLMs) to further inject more semantic information from visual and acoustic aspects. We evaluate our approach on benchmark VideoQA datasets: AVQA and Music-AVQA. Experimental results show that our approach outperforms competitive baseline models, achieving state-of-the-art performance on various question types.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Yi Yu 0001, Longyue Wang
ICASSP4
2024 A Scalable Sparse Transformer Model for Singing Melody Extraction
abstract
Extracting the melody of a singing voice is an essential task within the realm of music information retrieval (MIR). Recently, transformer based models have drawn great attention in the field of MIR. However, due to the expensive computation cost and extensive parameters, it is difficult to train and deploy a transformer-based model for practical singing melody extraction. In this paper, we propose a simple yet effective scalable sparse transformer for singing melody extraction. To be specific, we first propose to employ a sparse transformer to reduce computation cost and the amount of parameters. Then, we proposed to scale the self-attention region of the sparse transformer in the spectrogram to obtain more accurate performance. Moreover, we propose to combine a scalable sparse transformer (S2Former) with CNN-based model to extract global and local features in the spectrogram. The proposed scalable transformer model can achieve a better balance between a standard transformer and a sparse transformer. To better fuse the features from transformer and CNN, we further propose a transformer-CNN fusion (TCF) module to combine significant features from transformer and CNN. The proposed model obtains state-of-the-art results on several public datasets. The conducted experiments confirm the effectiveness of the model we proposed.
Shuai Yu 0002, Yi Yu 0001, Wei Li 0012
ICASSP3
2024 Prototype-Guided Prior Enhancement and Rectification in Few-shot Semantic Segmentation
abstract
Few-shot segmentation aims to recognize unseen objects from a few examples per class. It divides the training process into episodes with a support set and a query set. The support set provides annotated examples, and the query set contains images for annotation. Most models follow a prototypical approach, which extracts representative prototypes from support images to compare against query image features. PFENet proposed using prior maps to focus models on likely foreground areas, which effectively improves the model’s performance and is widely used in subsequent methods. However, existing prior-based techniques solely derive priors from high-level features. High-level features tend to lack detailed structure information due to the cumulative effects of convolutions and pooling operations, which leads to their poor performance in recognizing the boundaries of foreground areas. The uncertainty of high-level priors near boundaries may erroneously introduce category-irrelevant noise, as shown in Fig. 1. Moreover, high-level priors often overly focus on the most relevant parts, potentially overlooking other representative structures and losing semantic cues.
Yiming Tang 0004, Yi Yu 0001, Yan Qiu Chen
ICME2
2024 Harmonic Frequency-Separable Transformer for Instrument-Agnostic Music Transcription
abstract
Automatic Music Transcription (AMT) aims to convert music audio into symbolic representations. Recently, transformer-based methods have been successfully applied to instrument-agnostic music transcription. This allows transcription models can no longer focus on specific characteristics for an instrument class. However, these transformer-based methods designs for AMT were mainly motivated by other research fields and uses additional large-scale datasets, without considering the intrinsic features and patterns of the music signals. In this paper, we propose the Harmonic Frequency-Separable Transformer (HFSFormer), providing effective prior information based on music knowledge for instrument-agnostic transcription. The HFSFormer can capture the harmonic structure of music and separate time-frequency representations to decouple multiple pitches and different timbres, which can better explicitly model the note’s onset/offset and pitch. Experimental results show that our proposed method outperforms state-of-the-art peers on public datasets while having an order of magnitude fewer parameters.
Yulun Wu 0002, Weixing Wei, Dichucheng Li, Mengbo Li, Yi Yu 0001, Yongwei Gao, Wei Li 0012
ICME5
2024 Anchor-aware Deep Metric Learning for Audio-visual Retrieval
abstract
Metric learning minimizes the gap between similar (positive) pairs of data points and increases the separation of dissimilar (negative) pairs, aiming at capturing the underlying data structure and enhancing the performance of tasks like audio-visual cross-modal retrieval (AV-CMR). Recent works employ sampling methods to select impactful data points from the embedding space during training. However, the model training fails to fully explore the space due to the scarcity of training data points, resulting in an incomplete representation of the overall positive and negative distributions. In this paper, we propose an innovative Anchor-aware Deep Metric Learning (AADML) method to address this challenge by uncovering the underlying correlations among existing data points, which enhances the quality of the shared embedding space. Specifically, our method establishes a correlation graph-based manifold structure by considering the dependencies between each sample as the anchor and its semantically similar samples. Through dynamic weighting of the correlations within this underlying manifold structure using an attention-driven mechanism, Anchor Awareness (AA) scores are obtained for each anchor. These AA scores serve as data proxies to compute relative distances in metric learning approaches. Extensive experiments conducted on two audio-visual benchmark datasets demonstrate the effectiveness of our proposed AADML method, significantly surpassing state-of-the-art models. Furthermore, we investigate the integration of AA proxies with various metric learning methods, further highlighting the efficacy of our approach.
Donghuo Zeng, Yanan Wang 0002, Kazushi Ikeda, Yi Yu 0001
ICMR4
2024 Generalized News Event Discovery via Dynamic Augmentation and Entropy Optimization
abstract
News event discovery refers to the identification and detection of news events using multimodal data on social media. Currently, most works assume that the test set consists of known events. However, in real life, the emergence of new events is more frequent, which invalidates this assumption. In this paper, we propose a Dynamic Augmentation and Entropy Optimization (DAEO) model to address the scenario of generalized news event discovery, which requires the model to not only identify known events but also distinguish various new events. Specifically, we first introduce a multimodal augmentation module, which utilizes adversarial learning to enhance the multimodal representation capability. Secondly, we design an adaptive entropy optimization strategy combined with a self-distillation method, which uses multi-view pseudo-label consistency to improve the model's performance on both known and new events. In addition, we collect a multimodal news event discovery (MNED) dataset of 161,350 samples annotated with 66 real-world events. Extensive experimental results on the MNED dataset demonstrate the effectiveness of our proposed method. Our dataset is available on https://github.com/RetrainIt/MNED.
Zehang Lin, Jiayuan Xie, Zhenguo Yang, Yi Yu 0001, Qing Li 0001
ACM Multimedia4
2024 HKDSME: Heterogeneous Knowledge Distillation for Semi-supervised Singing Melody Extraction Using Harmonic Supervision
abstract
Singing melody extraction is a key task in the field of music information retrieval (MIR). However, decades of research works have uncovered two difficult issues. First, binary classification on frequency-domain audio features (e.g., spectrogram) is regarded as the primary method, which ignores the potential associations of musical information at different frequency bins, as well as their varying significance for output decisions. Second, the existing semi-supervised singing melody extraction models ignore the accuracy of the generated pseudo labels by semi-supervised models, which largely limits the further improvements of the model. To solve the two issues, in this paper, we propose a heterogeneous knowledge distillation framework for semi-supervised singing melody extraction using harmonic supervision, termed as HKDSME. We begin by proposing a four-class classification paradigm for determining the results of singing melody extraction using harmonic supervision. This enables the model to capture more information regarding melodic relations in spectrograms. To improve the accuracy issue of pseudo labels, we then build a semi-supervised method by leveraging the extracted harmonics as a consistent regularization. Different from previous methods, it judges the availability of unlabeled data in terms of the inner positional relations of extracted harmonics. To further build a light-weight semi-supervised model, we propose a heterogeneous knowledge distillation (HKD) module, which enables the prior knowledge to transfer between heterogeneous models. We also propose a novel confidence guided loss, which incorporates with the proposed HKD module to reduce the wrong pseudo labels. We evaluate our proposed method using several well-known public available datasets, and the findings demonstrate the efficacy of our proposed method.
Shuai Yu 0002, Xiaoliang He, Ke Chen 0021, Yi Yu 0001
ACM Multimedia4
2024 Semantic dependency network for lyrics generation from melody
Wei Duan 0004, Yi Yu 0001, Keizo Oyama
Neural Comput. Appl.2
2024 A Progressive Placeholder Learning Network for Multimodal Zero-Shot Learning
abstract
It is challenging to eliminate the domain shift between seen and unseen classes in multimodal zero-shot learning tasks due to the underlying disparity between the data distributions in the seen and unseen domains. In this paper, we propose a progressive placeholder learning network with mixup hallucination and an alternating mixer, denoted as MHAM, to maintain embedding spaces for unseen classes. Utilizing mixup hallucination (MH) on the visual and textual features obtained by BERT and a vision transformer, MHAM generates visual and textual hallucinated representations with pseudo class embeddings as placeholders for the unseen classes. Furthermore, a number of alternating mixer (AM) blocks are stacked to obtain modality-shared representations for the seen classes and hallucinated representations of progressive placeholders for the unseen classes. In particular, modality-shared representations are obtained by a mixer in an AM block by reversing the dimensionality of the modality-specific and raw representations to model intermodal interactions. MHAM exploits a freezing strategy by fixing the weights over the unseen classes in the last fully connected layer; this step acts as a projection from the raw and modality-shared representations to the embedding space of the seen and unseen classes. Experiments conducted on zero-shot datasets and news event datasets demonstrate the superior performance of the proposed MHAM method.
Zhuopan Yang, Zhenguo Yang, Xiaoping Li 0001, Yi Yu 0001, Qing Li 0001, Wenyin Liu
IEEE Trans. Multim.4
2024 Controllable Syllable-Level Lyrics Generation From Melody With Prior Attention
abstract
Melody-to-lyrics generation, which is based on syllable-level generation, is an intriguing and challenging topic in the interdisciplinary field of music, multimedia, and machine learning. Many previous research projects generate word-level lyrics sequences due to the lack of alignments between syllables and musical notes. Moreover, controllable lyrics generation from melody is also less explored but important for facilitating humans to generate diverse desired lyrics. In this work, we propose a controllable melody-to-lyrics model that is able to generate syllable-level lyrics with user-desired rhythm. An explicit n-gram (EXPLING) loss is proposed to train the Transformer-based model to capture the sequence dependency and alignment relationship between melody and lyrics and predict the lyrics sequences at the syllable level. A prior attention mechanism is proposed to enhance the controllability and diversity of lyrics generation. Experiments and evaluation metrics verified that our proposed model has the ability to generate higher-quality lyrics than previous methods and the feasibility of interacting with users for controllable and diverse lyrics generation. We believe this work provides valuable insights into human-centered AI research in music generation tasks. The source codes for this work will be made publicly available for further reference and exploration.
Zhe Zhang 0050, Yi Yu 0001, Atsuhiro Takasu
IEEE Trans. Multim.2
2023 Frame-Level Multi-Label Playing Technique Detection Using Multi-Scale Network and Self-Attention Mechanism
abstract
Instrument playing technique (IPT) is a key element of musical presentation. However, most of the existing works for IPT detection only concern monophonic music signals, yet little has been done to detect IPTs in polyphonic instrumental solo pieces with overlapping IPTs or mixed IPTs. In this paper, we formulate it as a frame-level multi-label classification problem and apply it to Guzheng, a Chinese plucked string instrument. We create a new dataset, Guzheng Tech99, containing Guzheng recordings and onset, offset, pitch, IPT annotations of each note. Because different IPTs vary a lot in their lengths, we propose a new method to solve this problem using multi-scale network and self-attention. The multi-scale network extracts features from different scales, and the self-attention mechanism applied to the feature maps at the coarsest scale further enhances the long-range feature extraction. Our approach outperforms existing works by a large margin, indicating its effectiveness in IPT detection.
Dichucheng Li, Mingjin Che, Wenwu Meng, Yulun Wu 0002, Yi Yu 0001, Wei Li 0012
ICASSP5
2023 LC-Beating: An Online System for Beat and Downbeat Tracking using Latency-Controlled Mechanism
abstract
Beat and downbeat tracking is to predict beat and downbeat time steps from a given music piece. Some deep learning models with a dilated structure such as Temporal Convolutional Network (TCN) and Dilated Self-Attention Network (DSAN) have achieved promising performance for this task. However, most of them have to see the whole music context during inference, which limits their deployment to online systems. In this paper, we propose LC-Beating, a novel latency-controlled (LC) mechanism for online beat and downbeat tracking, in which the model only looks ahead a few frames. By appending limited future information, the model can better capture the activity of relevant musical beats, which significantly boosts the performance of online algorithms with limited latency. Moreover, LC-Beating applies a novel real-time implementation of the LC mechanism to TCN and DSAN. The experimental results show that our proposed method outperforms the previous online models by a large margin and is close to the results of the offline models.
Xinlu Liu, Jiale Qian, Qiqi He, Yi Yu 0001, Wei Li 0012
ICME4
2023 MFAE: Masked frame-level autoencoder with hybrid-supervision for low-resource music transcription
abstract
Automantic Music Transcription (AMT) is an essential topic in music information retrieval (MIR), and it aims to transcribe audio recordings into symbolic representations. Recently, large-scale piano datasets with high-quality notations have been proposed for high-resolution piano transcription, which resulted in domain-specific AMT models achieved state-of- the-art results. However, those methods are hardly generalized to other ’low-resource’ instruments (such as guitar, cello, clarinet, etc.) transcription. In this paper, we propose a hybrid-supervised framework, the masked frame-level autoencoder (MFAE), to solve this issue. The proposed MFAE reconstructs the frame-level features of low-resource data to understand generic representations of low-resource instruments and improves low-resource transcription performance. Experimental results on several low- resource datasets (MAPS, MusicNet, and Guitarset) show that our framework achieves state-of-the-art performance in note-wise scores (Note F1 83.4%\64.1%\86.7%, Note-with-offset F1 59.8%\41.4%\71.6%). Moreover, our framework can be well generalized to various genres of instrument transcription, both in data-plentiful and data-limited scenarios.
Yulun Wu 0002, Yi Yu 0001, Wei Li 0012
ICME3
2023 Detecting Dialogue Hallucination Using Graph Neural Networks
abstract
Even though large language models (LLMs) accumulate tremendous knowledge, dialogue systems built with LLMs induce hallucinations, leading to the generation of non-factual responses. How to provide proper references to achieve interpretable hallucination detection is a key issue that needs to be addressed. In this paper, we propose a graph neural network (GNN)-based method to achieve high-performance and interpretable hallucination detection for domain-specific dialogue systems. The method involves performing graph matching between a reference knowledge graph obtained from a knowledge database and a response knowledge graph extracted from the response to detect non-factual responses. By comparing with strong baselines, our method achieves a recall improvement of up to 11% and infers the cause of hallucinations with a probability of over 79%.
Kazuaki Furumai, Yanan Wang 0002, Makoto Shinohara, Kazushi Ikeda, Yi Yu 0001, Tsuneo Kato
ICMLA5
2023 Semantic Difference Guidance for the Uncertain Boundary Segmentation of CT Left Atrial Appendage
Xin You 0002, Yangqian Wu, Yi Yu 0001, Yun Gu, Jie Yang 0002
MICCAI (7)5
2023 Graph-Based Video-Language Learning with Multi-Grained Audio-Visual Alignment
abstract
Video-language learning has attracted significant attention in the fields of multimedia, computer vision and natural language processing in recent years. One of the key challenges in this area is how to effectively integrate visual and linguistic information to enable machines to understand video content and query information. In this work, we leverage graph-based representations and multi-grained audio-visual alignment to address this challenge. First, our approach starts by transforming video and query inputs into visual-scene graphs and semantic role graphs using a visual-scene parser and semantic role labeler respectively. These graphs are then encoded using graph neural networks to obtain enriched representations and combined to obtain a video-query joint representation that enhances the semantic expressivity of the inputs. Second, to achieve accurate matching of relevant parts of audio and visual features, we propose a multi-grained alignment module that aligns the audio and visual features at multiple scales. This enables us to effectively fuse the audio and visual information in a way that is consistent with the semantic-level information captured by the graph-based representations. Experiments on five representative datasets collected for Video Retrieval and Video Question Answering tasks show that our approach outperforms the literature on several metrics. Our extensive ablation studies demonstrate the effectiveness of graph-based representation and multi-grained audio-visual alignment.
Chenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang, Liting Zhou, Cathal Gurrin, Linyi Yang, Yi Yu 0001, Yvette Graham, Jennifer Foster
ACM Multimedia8
2023 A neural harmonic-aware network with gated attentive fusion for singing melody extraction
Shuai Yu 0002, Yi Yu 0001, Xiaoheng Sun, Wei Li 0012
Neurocomputing2
2023 Conditional hybrid GAN for melody generation from lyrics
Yi Yu 0001, Zhe Zhang 0050, Wei Duan 0004, Rajiv Ratn Shah
Neural Comput. Appl.1
2023 Controllable lyrics-to-melody generation
Zhe Zhang 0050, Yi Yu 0001, Atsuhiro Takasu
Neural Comput. Appl.2
2023 Multi-scale network with shared cross-attention for audio-visual correlation learning
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Wei Li 0012
Neural Comput. Appl.2
2023 Context-Patch Representation Learning With Adaptive Neighbor Embedding for Robust Face Image Super-Resolution
abstract
Representation learning steered robust face image super-resolution (FSR) methods have attracted extensive attention in the past few decades. Most previous methods were devoted to exploiting the local position patches in the training set for FSR. However, they usually overlooked the sufficient usage of the contextual information around the testing patches, which are useful for stable representation learning. In this article, we attempt to utilize the context-patch around the testing patch and propose a method named context-patch representation learning with adaptive neighbor embedding (CRL-ANE) for FSR. On one hand, we simultaneously use the testing position patch and its adjacent ones for stable representation weight learning. This contextual information can compensate for recovering missing details in the target patch. On the other hand, for each input patch set, due to its inherent facial structural properties, we design an adaptive neighbor embedding strategy to elaborately and adaptively choose primary candidates for more accurate reconstruction. These two improvements enable the proposed method to achieve better SR performance than some of the other methods. Qualitative and quantitative experiments on some benchmarks have validated the superiority of the proposed method over some state-of-the-art methods.
Guangwei Gao, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003, Dong Yue 0001
IEEE Trans. Multim.2
2023 FBSNet: A Fast Bilateral Symmetrical Network for Real-Time Semantic Segmentation
abstract
Real-time semantic segmentation, which can be visually understood as the pixel-level classification task on the input image, currently has broad application prospects, especially in the fast-developing fields of autonomous driving and drone navigation. However, the huge burden of calculation together with redundant parameters are still the obstacles to its technological development. In this article, we propose a Fast Bilateral Symmetrical Network (FBSNet) to alleviate the above challenges. Specifically, FBSNet employs a symmetrical encoder-decoder structure with two branches, semantic information branch and spatial detail branch. The Semantic Information Branch (SIB) is the main branch with semantic architecture to acquire the contextual information of the input image and meanwhile acquire sufficient receptive field. While the Spatial Detail Branch (SDB) is a shallow and simple network used to establish local dependencies of each pixel for preserving details, which is essential for restoring the original resolution during the decoding phase. Meanwhile, a Feature Aggregation Module (FAM) is designed to effectively combine the output of these two branches. Experimental results of Cityscapes and CamVid show that the proposed FBSNet can strike a good balance between accuracy and efficiency. Specifically, it obtains 70.9% and 68.9% mIoU along with the inference speed of 90 fps and 120 fps on these two test datasets, respectively, with only 0.62 million parameters on a single RTX 2080Ti GPU. The code is available athttps://github.com/IVIPLab/FBSNet.
Guangwei Gao, Guoan Xu, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001, Jian Yang 0003
IEEE Trans. Multim.4
2023 Variational Autoencoder with CCA for Audio-Visual Cross-modal Retrieval
abstract
Cross-modal retrieval is to utilize one modality as a query to retrieve data from another modality, which has become a popular topic in information retrieval, machine learning, and databases. Finding a method to effectively measure the similarity between different modality data is the major challenge of cross-modal retrieval. Although several research works have calculated the correlation between different modality data via learning a common subspace representation, the encoder’s ability to extract features from multi-modal information is not satisfactory. In this article, we present a novel variational autoencoder architecture for audio–visual cross-modal retrieval by learning paired audio–visual correlation embedding and category correlation embedding as constraints to reinforce the mutuality of audio–visual information. On the one hand, audio encoder and visual encoder separately encode audio data and visual data into two different latent spaces. Further, two mutual latent spaces are respectively constructed by canonical correlation analysis. On the other hand, probabilistic modeling methods are used to deal with possible noise and missing information in the data. Additionally, in this way, the cross-modal discrepancies from intra-modal and inter-modal information are simultaneously eliminated in the joint embedding subspace. We conduct extensive experiments over two benchmark datasets. The experimental results confirm that the proposed architecture is effective in learning audio–visual correlation and is appreciably better than the existing cross-modal retrieval methods.
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Wei Li 0012
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Melody Generation from Lyrics with Local Interpretability
abstract
Melody generation aims to learn the distribution of real melodies to generate new melodies conditioned on lyrics, which has been a very interesting topic in the area of artificial intelligence and music. However, a challenging issue still limits the quality and reliability of melody generation conditioned on lyrics: how to enhance the interpretability between the input lyrics and generated melodies so humans can understand their relationships. To solve this issue, in this article, we propose a model for melody generation from lyrics with local interpretability, which contains two significant contributions: (i) Mutual information between input lyrics and generated melody is exploited to instruct the training of the network, which avoids the loss of content consistency during the training stage. (ii) Transformer is explored to efficiently extract semantic features from lyrics sequences, which provides more interpretable correlations between different syllables in lyrics. Experiments on a large-scale dataset with paired lyrics-melodies demonstrate that the proposed approach can generate higher-quality melodies from lyrics compared with existing methods.
Wei Duan 0004, Yi Yu 0001, Xulong Zhang 0001, Suhua Tang, Wei Li 0012, Keizo Oyama
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Query-Guided Prototype Learning with Decoder Alignment and Dynamic Fusion in Few-Shot Segmentation
abstract
Few-shot segmentation aims to segment objects belonging to a specific class under the guidance of a few annotated examples. Most existing approaches follow the prototype learning paradigm and generate category prototypes by squeezing masked feature maps extracted from images in the support set. These support prototypes may lead to inaccurate predictions when directly compared with features extracted from the query set due to the considerable distribution discrepancy between support and query features. We propose a query-guided prototype learning architecture to address this problem from two aspects: (i) We propose a cross-alignment loss for training the segmentation decoder. This loss function will help the decoder improve its robustness against the distribution discrepancy between support and query features. (ii) We build a dynamic fusion module to strengthen the original support prototype with another prototype extracted from query features. Experiments show that our method achieves promising results compared to previous prototype learning methods on PASCAL-5 i and COCO-20 i datasets.
Yiming Tang 0004, Yi Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 Learning Explicit and Implicit Dual Common Subspaces for Audio-visual Cross-modal Retrieval
abstract
Audio-visual tracks in video contain rich semantic information with potential in many applications and research. Since the audio-visual data have inconsistent distributions and because of the heterogeneous nature of representations, the heterogeneous gap between modalities makes them impossible to compare directly. To bridge the modality gap, a frequently adopted approach is to simultaneously project audio-visual data into a common subspace to capture the commonalities and characteristics of modalities for measurement, which has been extensively studied in relation to the issues of modality-common and modality-specific feature learning in previous research. However, it is difficult for existing methods to address the tradeoff between both issues; e.g., the modality-common feature is learned from the latent commonalities of audio-visual data or the correlated features as aligned projections, in which the modality-specific feature can be lost. To solve the tradeoff, we propose a novel end-to-end architecture, which synchronously projects audio-visual data into the explicit and the implicit dual common subspaces. The explicit subspace is used to learn modality-common features and reduce the modality gap of explicitly paired audio-visual data, where the representation-specific details are abandoned to retain the common underlying structure of audio-visual data. The implicit subspace is used to learn modality-specific features, where each modality privately pulls apart the feature distances between different categories to maintain the category-based distinctions, by minimizing the distance between audio-visual features and corresponding labels. The comprehensive experimental results on two audio-visual datasets, VEGAS and AVE, demonstrate that our proposed model for using two different common subspaces for audio-visual cross-modal learning is effective and significantly outperforms the state-of-the-art cross-modal models that learn features from a single common subspace by 4.30% and 2.30% in terms of average MAP on the VEGAS and AVE datasets, respectively.
Donghuo Zeng, Gen Hattori, Yi Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Feature Distillation Interaction Weighting Network for Lightweight Image Super-resolution
abstract
Convolutional neural networks based single-image superresolution (SISR) has made great progress in recent years. However, it is difficult to apply these methods to real-world scenarios due to the computational and memory cost. Meanwhile, how to take full advantage of the intermediate features under the constraints of limited parameters and calculations is also a huge challenge. To alleviate these issues, we propose a lightweight yet efficient Feature Distillation Interaction Weighted Network (FDIWN). Specifically, FDIWN utilizes a series of specially designed Feature Shuffle Weighted Groups (FSWG) as the backbone, and several novel mutual Wide-residual Distillation Interaction Blocks (WDIB) form an FSWG. In addition, Wide Identical Residual Weighting (WIRW) units and Wide Convolutional Residual Weighting (WCRW) units are introduced into WDIB for better feature distillation. Moreover, a Wide-Residual Distillation Connection (WRDC) framework and a Self-Calibration Fusion (SCF) unit are proposed to interact features with different scales more flexibly and efficiently. Extensive experiments show that our FDIWN is superior to other models to strike a good balance between model performance and efficiency. The code is available at https://github.com/IVIPLab/FDIWN.
Guangwei Gao, Juncheng Li 0003, Fei Wu 0004, Huimin Lu 0001, Yi Yu 0001
AAAI6
2022 Deepchorus: A Hybrid Model of Multi-Scale Convolution And Self-Attention for Chorus Detection
abstract
Chorus detection is a challenging problem in musical signal processing as the chorus often repeats more than once in popular songs, usually with rich instruments and complex rhythm forms. Most of the existing works focus on the receptiveness of chorus sections based on some explicit features such as loudness and occurrence frequency. These pre-assumptions for chorus limit the generalization capacity of these methods, causing misdetection on other repeated sections such as verse. To solve the problem, in this paper we propose an end-to-end chorus detection model DeepChorus, reducing the engineering effort and the need for prior knowledge. The proposed model includes two main structures: i) a Multi-Scale Network to derive preliminary representations of chorus segments, and ii) a Self-Attention Convolution Network to further process the features into probability curves representing chorus presence. To obtain the final results, we apply an adaptive threshold to binarize the original curve. The experimental results show that DeepChorus outperforms existing state-of-the-art methods in most cases.
Qiqi He, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012
ICASSP3
2022 HarmoF0: Logarithmic Scale Dilated Convolution for Pitch Estimation
abstract
Sounds, especially music, contain various harmonic components scattered in the frequency dimension. It is difficult for normal convolutional neural networks to ob-serve these overtones. This paper introduces a multiple rates dilated causal convolution (MRDC-Conv) method to capture the harmonic structure in logarithmic scale spectrograms efficiently. The harmonic is helpful for pitch estimation, which is important for many sound processing applications. We propose HarmoF0, a fully convolutional network, to evaluate the MRDC-Conv and other dilated convolutions in pitch estimation. The re-sults show that this model outperforms the DeepF0, yields state-of-the-art performance in three datasets, and simultaneously reduces more than 90% parameters. We also find that it has stronger noise resistance and fewer octave errors.
Weixing Wei, Yi Yu 0001, Wei Li 0012
ICME3
2022 Multimodal Music Emotion Recognition with Hierarchical Cross-Modal Attention Network
abstract
Computational music emotion recognition is to recognize the emotional content in music tracks. In computational music emotion recognition studies, researchers have paid close attention to the audio content of the music tracks. Although lyrics content and music context contribute greatly to the perceived emotion, these kinds of emotional information are usually ignored. Based on this finding, we propose a multimodal music emotion recognition method jointly predicting the valence and arousal values by combining the audio, lyrics, track name, and artist of a given track. Audio features, lyrics features and context features are extracted separately and fused by a cross-modal attention mechanism, forming a hierarchical structure. Our proposed model outperforms two baselines by a large margin and achieves state-of-the-art performance on two public datasets.
Ganghui Ru, Yi Yu 0001, Yulun Wu 0002, Dichucheng Li, Wei Li 0012
ICME3
2022 Lightweight Bimodal Network for Single-Image Super-Resolution via Symmetric CNN and Recursive Transformer
abstract
Single-image super-resolution (SISR) has achieved significant breakthroughs with the development of deep learning. However, these methods are difficult to be applied in real-world scenarios since they are inevitably accompanied by the problems of computational and memory costs caused by the complex operations. To solve this issue, we propose a Lightweight Bimodal Network (LBNet) for SISR. Specifically, an effective Symmetric CNN is designed for local feature extraction and coarse image reconstruction. Meanwhile, we propose a Recursive Transformer to fully learn the long-term dependence of images thus the global information can be fully used to further refine texture details. Studies show that the hybrid of CNN and Transformer can build a more efficient model. Extensive experiments have proved that our LBNet achieves more prominent performance than other state-of-the-art methods with a relatively low computational cost and memory consumption. The code is available at https://github.com/IVIPLab/LBNet.
Guangwei Gao, Zhengxue Wang, Juncheng Li 0003, Yi Yu 0001, Tieyong Zeng
IJCAI5
2022 Deep Attention-Based Alignment Network for Melody Generation from Incomplete Lyrics
abstract
We propose a deep attention-based alignment network, which aims to automatically predict lyrics and melody with given incomplete lyrics as input in a way similar to the music creation of humans. Most importantly, a deep neural lyrics-to-melody net is trained in an encoder-decoder way to predict possible pairs of lyrics-melody when given incomplete lyrics (few keywords). The attention mechanism is exploited to align the predicted lyrics with the melody during the lyrics-to-melody generation. The qualitative and quantitative evaluation metrics reveal that the proposed method is indeed capable of generating proper lyrics and corresponding melody for composing new songs given a piece of incomplete seed lyrics.
Gurunath Reddy M, Zhe Zhang 0050, Yi Yu 0001, Florian Harscoët, Simon Canales, Suhua Tang
ISM3
2022 Interpretable Melody Generation from Lyrics with Discrete-Valued Adversarial Training
abstract
Generating melody from lyrics is an interesting yet challenging task in the area of artificial intelligence and music. However, the difficulty of keeping the consistency between input lyrics and generated melody limits the generation quality of previous works. In our proposal, we demonstrate our proposed interpretable lyrics-to-melody generation system which can interact with users to understand the generation process and recreate the desired songs. To improve the reliability of melody generation that matches lyrics, mutual information is exploited to strengthen the consistency between lyrics and generated melodies. Gumbel-Softmax is exploited to solve the non-differentiability problem of generating discrete music attributes by Generative Adversarial Networks (GANs). Moreover, the predicted probabilities output by the generator is utilized to recommend music attributes. Interacting with our lyrics-to-melody generation system, users can listen to the generated AI song as well as recreate a new song by selecting from recommended music attributes.
Wei Duan 0004, Zhe Zhang 0050, Yi Yu 0001, Keizo Oyama
ACM Multimedia3
2022 Emotional Talking Faces: Making Videos More Expressive and Realistic
abstract
Lip synchronization and talking face generation have gained a specific interest from the research community with the advent and need of digital communication in different fields. Prior works propose several elegant solutions to this problem. However, they often fail to create realistic-looking videos that account for people's expressions and emotions. To mitigate this, we build a talking face generation framework conditioned on a categorical emotion to generate videos with appropriate expressions, making them more real-looking and convincing. With a broad range of six emotions i.e., anger, disgust, fear, happiness, neutral, and sad, we show that our model generalizes across identities, emotions, and languages.
Sahil Goyal, Shagun Uppal, Sarthak Bhagat, Dhroov Goel, Sakshat Mali, Yi Yu 0001, Yifang Yin, Rajiv Ratn Shah
MMAsia6
2022 Melody Generation from Lyrics Using Three Branch Conditional LSTM-GAN
Wei Duan 0004, Rajiv Ratn Shah, Suhua Tang, Wei Li 0012, Yi Yu 0001
MMM (1)7
2022 Hierarchical Deep CNN Feature Set-Based Representation Learning for Robust Cross-Resolution Face Recognition
abstract
Cross-resolution face recognition (CRFR), which is important in intelligent surveillance and biometric forensics, refers to the problem of matching a low-resolution (LR) probe face image against high-resolution (HR) gallery face images. Existing shallow learning-based and deep learning-based methods focus on mapping the HR-LR face pairs into a joint feature space where the resolution discrepancy is mitigated. However, little works consider how to extract and utilize the intermediate discriminative features from the noisy LR query faces to further mitigate the resolution discrepancy due to the resolution limitations. In this study, we desire to fully exploit the multi-level deep convolutional neural network (CNN) feature set for robust CRFR. In particular, our contributions are threefold. (i) To learn more robust and discriminative features, we desire to adaptively fuse the contextual features from different layers. (ii) To fully exploit these contextual features, we design a feature set-based representation learning (FSRL) scheme to collaboratively represent the hierarchical features for more accurate recognition. Moreover, FSRL utilizes the primitive form of feature maps to keep the latent structural information, especially in noisy cases. (iii) To further promote the recognition performance, we desire to fuse the hierarchical recognition outputs from different stages. Meanwhile, the discriminability from different scales can also be fully integrated. By exploiting these advantages, the efficiency of the proposed method can be delivered. Experimental results on several face datasets have verified the superiority of the presented algorithm to the other competitive CRFR approaches.
Guangwei Gao, Yi Yu 0001, Jian Yang 0003, Guo-Jun Qi, Meng Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 MSCFNet: A Lightweight Network With Multi-Scale Context Fusion for Real-Time Semantic Segmentation
abstract
In recent years, how to strike a good trade-off between accuracy, inference speed, and model size has become the core issue for real-time semantic segmentation applications, which plays a vital role in real-world scenarios such as autonomous driving systems and drones. In this study, we devise a novel lightweight network using a multi-scale context fusion (MSCFNet) scheme, which explores an asymmetric encoder-decoder architecture to alleviate these problems. More specifically, the encoder adopts some developed efficient asymmetric residual (EAR) modules, which are composed of factorization depth-wise convolution and dilation convolution. Meanwhile, instead of complicated computation, simple deconvolution is applied in the decoder to further reduce the amount of parameters while still maintaining the high segmentation accuracy. Also, MSCFNet has branches with efficient attention modules from different stages of the network to well capture multi-scale contextual information. Then we combine them before the final classification to enhance the expression of the features and improve the segmentation efficiency. Comprehensive experiments on challenging datasets have demonstrated that the proposed MSCFNet, which contains only 1.15M parameters, achieves 71.9% Mean IoU on the Cityscapes testing dataset and can run at over 50 FPS on a single Titan XP GPU configuration.
Guangwei Gao, Guoan Xu, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Dong Yue 0001
IEEE Trans. Intell. Transp. Syst.3
2022 Towards Multi-Domain Face Synthesis Via Domain-Invariant Representations and Multi-Level Feature Parts
abstract
Cross-domain face synthesis plays a positive role in the real world. It is challenging to synthesize high-quality faces across multiple domains based on limited paired data because the multiple mappings between different domains may interfere with each other. Cognitive science investigates that the brain can recognize the same person with multiple different expressions by extracting invariant information on the face and we humans perceive instances by decomposing them into parts. Motivated by these cognition, we propose a unified semi-supervised framework for multi-domain face synthesis by extracting a domain-invariant representation and exploiting parts of multi-level features. Specifically, realized by adversarial training with additional ability to utilize domain-specific information, a encoder is trained to remove domain-specific information and extract the domain-invariant representation from multiple inputs. Then, we utilize the multi-level feature parts extracted from inputs and reconstructed faces via a pre-trained recognition model to ensure that the domain-invariant representation contains enough useful semantic information. we also utilize the feature parts extracted from inputs and limited paired data to compose pseudo features in target domain for supervising the synthesis, which makes our framework suitable for large amounts of unpaired training data. By exploiting this framework, we can achieve face synthesis between multiple domains using some paired data together with a large training database without ground truth target faces. Experimental results demonstrate our framework achieves great performances on qualitative and quantitative evaluations under both artificial and uncontrolled environments, and our framework has competitive performances in single translation compared with specialized methods for translation between two specific domains.
Dawei Zhou 0004, Nannan Wang 0001, Chunlei Peng, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001
IEEE Trans. Multim.4
2022 Leaning compact and representative features for cross-modality person re-identification
Guangwei Gao, Hao Shao, Fei Wu 0004, Meng Yang 0001, Yi Yu 0001
World Wide Web5
2021 Long-short Term Prediction for Occluded Multiple Object Tracking
abstract
Online multiple object tracking (MOT) is a challenging problem in complex scenes due to frequent occlusions. Most of the existing MOT methods tend to focus on addressing an individual type of occlusion, which cannot meet the requirements of real complex scenes. In this paper, we propose a unified MOT framework that combines long- and short-term prediction models for online multiple object tracking. Basically, The short-term prediction model consists of an appearance-based model and a motion-based model, aiming at exploiting the appearance and motion of objects to handle different types of occlusions jointly. Furthermore, we adopt a cubic spline interpolation as a long-term prediction model to estimate the trajectory of the target in occluded frames. To handle different lengths of occlusions, an adaptive weighted fusion model is proposed to combine the short-term prediction model, and the long-term prediction model. Experimental results on several challenging datasets demonstrate that the proposed method outperforms state-of-the-art methods.
Jun Chen 0001, Mithun Mukherjee 0001, Weijian Ruan, Chao Liang 0001, Yi Yu 0001
GLOBECOM6
2021 Interpretable Visual Understanding with Cognitive Attention Network
Xuejiao Tang, Wenbin Zhang 0002, Yi Yu 0001, Kea Turner, Tyler Derr, Eirini Ntoutsi
ICANN (1)3
2021 Frequency-Temporal Attention Network for Singing Melody Extraction
abstract
Musical audio is generally composed of three physical properties: frequency, time and magnitude. Interestingly, human auditory periphery also provides neural codes for each of these dimensions to perceive music. Inspired by these intrinsic characteristics, a frequency-temporal attention network is proposed to mimic human auditory for singing melody extraction. In particular, the proposed model contains frequency-temporal attention modules and a selective fusion module corresponding to these three physical properties. The frequency attention module is used to select the same activation frequency bands as did in cochlear and the temporal attention module is responsible for analyzing temporal patterns. Finally, the selective fusion module is suggested to recalibrate magnitudes and fuse the raw information for prediction. In addition, we propose to use another branch to simultaneously predict the presence of singing voice melody. The experimental results show that the proposed model outperforms existing state-of-the-art methods1.
Shuai Yu 0002, Xiaoheng Sun, Yi Yu 0001, Wei Li 0012
ICASSP3
2021 Singer Identification Using Deep Timbre Feature Learning with KNN-NET
abstract
In this paper, we study the issue of automatic singer identification (SID) in popular music recordings, which aims to recognize who sang a given piece of song. The main challenge for this investigation lies in the fact that a singer’s singing voice changes and intertwines with the signal of background accompaniment in time domain. To handle this challenge, we propose the KNN-Net for SID, which is a deep neural network model with the goal of learning local timbre feature representation from the mixture of singer voice and background music. Unlike other deep neural networks using the softmax layer as the output layer, we instead utilize the KNN as a more interpretable layer to output target singer labels. Moreover, attention mechanism is first introduced to highlight crucial timbre features for SID. Experiments on the existing artist20 dataset show that the proposed approach outperforms the state-of-the-art method by 4%. We also create singer32 and singer60 datasets consisting of Chinese pop music to evaluate the reliability of the proposed method. The more extensive experiments additionally indicate that our proposed model achieves a significant performance improvement compared to the state-of-the-art methods.
Xulong Zhang 0001, Jiale Qian, Yi Yu 0001, Yifu Sun, Wei Li 0012
ICASSP3
2021 Lightweight Image Super-Resolution with Multi-Scale Feature Interaction Network
abstract
Recently, the single image super-resolution (SISR) approaches with deep and complex convolutional neural network structures have achieved promising performance. However, those methods improve the performance at the cost of higher memory consumption, which is difficult to be applied for some mobile devices with limited storage and computing resources. To solve this problem, we present a lightweight multi-scale feature interaction network (MSFIN). For lightweight SISR, MSFIN expands the receptive field and adequately exploits the informative features of the low-resolution observed images from various scales and interactive connections. In addition, we design a lightweight recurrent residual channel attention block (RRCAB) so that the network can benefit from the channel attention mechanism while being sufficiently lightweight. Extensive experiments on some benchmarks have confirmed that our proposed MSFIN can achieve comparable performance against the state-of-the-arts with a more lightweight model.
Zhengxue Wang, Guangwei Gao, Juncheng Li 0003, Yi Yu 0001, Huimin Lu 0001
ICME4
2021 Adversarial Learning with Mask Reconstruction for Text-Guided Image Inpainting
abstract
Text-guided image inpainting aims to complete the corrupted patches coherent with both visual and textual context. On one hand, existing works focus on surrounding pixels of the corrupted patches without considering the objects in the image, resulting in the characteristics of objects described in text being painted on non-object regions. On the other hand, the redundant information in text may distract the generation of objects of interest in the restored image. In this paper, we propose an adversarial learning framework with mask reconstruction (ALMR) for image inpainting with textual guidance, which consists of a two-stage generator and dual discriminators. The two-stage generator aims to restore coarse-grained and fine-grained images, respectively. In particular, we devise a dual-attention module (DAM) to incorporate the word-level and sentence-level textual features as guidance on generating the coarse-grained and fine-grained details in the two stages. Furthermore, we design a mask reconstruction module (MRM) to penalize the restoration of the objects of interest with the given textual descriptions about the objects. For adversarial training, we exploit global and local discriminators for the whole image and corrupted patches, respectively. Extensive experiments conducted on CUB-200-2011, Oxford-102 and CelebA-HQ show the outperformance of the proposed ALMR (e.g., FID value is reduced from 29.69 to 14.69 compared with the state-of-the-art approach on CUB-200-2011). Codes are available at https://github.com/GaranWu/ALMR
Xingcai Wu, Yucheng Xie, Jiaqi Zeng, Zhenguo Yang, Yi Yu 0001, Qing Li 0001, Wenyin Liu
ACM Multimedia5
2021 LBAN-IL: A novel method of high discriminative representation for facial expression recognition
Hangyu Li 0001, Nannan Wang 0001, Yi Yu 0001, Xi Yang 0011, Xinbo Gao 0001
Neurocomputing3
2021 Constructing multilayer locality-constrained matrix regression framework for noise robust face super-resolution
Guangwei Gao, Yi Yu 0001, Jin Xie 0001, Jian Yang 0003, Meng Yang 0001, Jian Zhang 0002
Pattern Recognit.2
2021 HANME: Hierarchical Attention Network for Singing Melody Extraction
abstract
Singing melody extraction in polyphonic musical audio is a very critical and challenging task in music information retrieval (MIR). Contextual frame-level information has proven its effectiveness in this task. However, existing works assign equal weight to each contextual frame, which may hinder the further improvement of the performance. To this end, we propose a hierarchical attention network for singing melody extraction (HANME) to extract the discriminative attention-aware features and alleviate the workload of the convolutional recurrent neural network (CRNN) for extracting local spatial and temporal features. Specifically, the first attention layer learns the context vector based on local spatial features extracted by residual convolutional neural network (CNN), and the second attention layer learns the temporal context vector based on long-term features extracted by Bidirectional Gated Recurrent Units (BiGRU). Due to the scarcity of labeled training data, we further propose a partial parameter adaptation approach to address the imbalance distribution of the labels for this task. We use the RWC dataset and part of vocal tracks of the MedleyDB dataset for training the model and evaluate the performance on the ADC2004, MIREX 05 and MedleyDB datasets. The experimental study demonstrates the superiority of our method compared with other state-of-the-art ones.
Shuai Yu 0002, Yi Yu 0001, Wei Li 0012
IEEE Signal Process. Lett.2
2021 Robust Facial Image Super-Resolution by Kernel Locality-Constrained Coupled-Layer Regression
abstract
Super-resolution methods for facial image via representation learning scheme have become very effective methods due to their efficiency. The key problem for the super-resolution of facial image is to reveal the latent relationship between the low-resolution ( LR ) and the corresponding high-resolution ( HR ) training patch pairs. To simultaneously utilize the contextual information of the target position and the manifold structure of the primitive HR space, in this work, we design a robust context-patch facial image super-resolution scheme via a kernel locality-constrained coupled-layer regression (KLC2LR) scheme to obtain the desired HR version from the acquired LR image. Here, KLC2LR proposes to acquire contextual surrounding patches to represent the target patch and adds an HR layer constraint to compensate the detail information. Additionally, KLC2LR desires to acquire more high-frequency information by searching for nearest neighbors in the HR sample space. We also utilize kernel function to map features in original low-dimensional space into a high-dimensional one to obtain potential nonlinear characteristics. Our compared experiments in the noisy and noiseless cases have verified that our suggested methodology performs better than many existing predominant facial image super-resolution methods.
Guangwei Gao, Huimin Lu 0001, Yi Yu 0001, Heyou Chang, Dong Yue 0001
ACM Trans. Internet Techn.4
2021 Correlation Discrepancy Insight Network for Video Re-identification
abstract
Video-based person re-identification (ReID) aims at re-identifying a specified person sequence from videos that were captured by disjoint cameras. Most existing works on this task ignore the quality discrepancy across frames by using all video frames to develop a ReID method. Additionally, they adopt only the person self-characteristic as the representation, which cannot adapt to cross-camera variation effectively. To that end, we propose a novel correlation discrepancy insight network for video-based person ReID, which consists of an unsupervised correlation insight model (CIM) for video purification and a discrepancy description network (DDN) for person representation. Concretely, CIM is constructed by using kernelized correlation filters to encode person half-parts, which evaluates the frame quality by the cross correlation across frames for selecting discriminative video fragments. Furthermore, DDN exploits the selected video fragments to generate a discrepancy descriptor using a compression network, which aims at employing the discrepancies with other persons’ to facilitate the representation of the target person rather than only using the self-characteristic. Due to the advantage in handling cross-domain variation, the discrepancy descriptor is expected to provide a new pattern for the object representation in cross-camera tasks. Experimental results on three public benchmarks demonstrate that the proposed method outperforms several state-of-the-art methods.
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Wu Liu 0005, Jun Chen 0001, Jiayi Ma 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Conditional LSTM-GAN for Melody Generation from Lyrics
abstract
Melody generation from lyrics has been a challenging research issue in the field of artificial intelligence and music, which enables us to learn and discover latent relationships between interesting lyrics and accompanying melodies. Unfortunately, the limited availability of a paired lyrics–melody dataset with alignment information has hindered the research progress. To address this problem, we create a large dataset consisting of 12,197 MIDI songs each with paired lyrics and melody alignment through leveraging different music sources where alignment relationship between syllables and music attributes is extracted. Most importantly, we propose a novel deep generative model, conditional Long Short-Term Memory (LSTM)–Generative Adversarial Network for melody generation from lyrics, which contains a deep LSTM generator and a deep LSTM discriminator both conditioned on lyrics. In particular, lyrics-conditioned melody and alignment relationship between syllables of given lyrics and notes of predicted melody are generated simultaneously. Extensive experimental results have proved the effectiveness of our proposed lyrics-to-melody generative model, where plausible and tuneful sequences can be inferred from lyrics.
Yi Yu 0001, Simon Canales
ACM Trans. Multim. Comput. Commun. Appl.1
2020 End-to-End Named Entity Recognition from English Speech
abstract
Named entity recognition (NER) from text has been a widely studied problem and usually extracts semantic information from text. Until now, NER from speech is mostly studied in a two-step pipeline process that includes first applying an automatic speech recognition (ASR) system on an audio sample and then passing the predicted transcript to a NER tagger. In such cases, the error does not propagate from one step to another as both the tasks are not optimized in an end-to-end (E2E) fashion. Recent studies confirm that integrated approaches (e.g., E2E ASR) outperform sequential ones (e.g., phoneme based ASR). In this paper, we introduce a first publicly available NER annotated dataset for English speech and present an E2E approach, which jointly optimizes the ASR and NER tagger components. Experimental results show that the proposed E2E approach outperforms the classical two-step approach. We also discuss how NER from speech can be used to handle out of vocabulary (OOV) words in an ASR system.
Hemant Yadav, Sreyan Ghosh, Yi Yu 0001, Rajiv Ratn Shah
INTERSPEECH3
2020 C3VQG: category consistent cyclic visual question generation
abstract
Visual Question Generation (VQG) is the task of generating natural questions based on an image. Popular methods in the past have explored image-to-sequence architectures trained with maximum likelihood which have demonstrated meaningful generated questions given an image and its associated ground-truth answer. VQG becomes more challenging if the image contains rich contextual information describing its different semantic categories. In this paper, we try to exploit the different visual cues and concepts in an image to generate questions using a variational autoencoder (VAE) without ground-truth answers. Our approach solves two major shortcomings of existing VQG systems: (i) minimize the level of supervision and (ii) replace generic questions with category relevant generations. Most importantly, by eliminating expensive answer annotations, the required supervision is weakened. Using different categories enables us to exploit different concepts as the inference requires only the image and the category. Mutual information is maximized between the image, question, and answer category in the latent space of our VAE. A novel category consistent cyclic loss is proposed to enable the model to generate consistent predictions with respect to the answer category, reducing redundancies and irregularities. Additionally, we also impose supplementary constraints on the latent space of our generative model to provide structure based on categories and enhance generalization by encapsulating decorrelated features within each dimension. Through extensive experiments, the proposed model, C3VQG outperforms state-of-the-art VQG methods with weak supervision.
Shagun Uppal, Anish Madan, Sarthak Bhagat, Yi Yu 0001, Rajiv Ratn Shah
MMAsia4
2020 Lyrics-Conditioned Neural Melody Generation
Yi Yu 0001, Florian Harscoët, Simon Canales, Gurunath Reddy M, Suhua Tang, Junjun Jiang
MMM (2)1
2020 A Relation Learning Hierarchical Framework for Multi-label Charge Prediction
Wei Duan 0004, Yi Yu 0001
PAKDD (2)3
2020 Image super-resolution via multi-view information fusion networks
Nannan Wang 0001, Jingwei Xin, Xi Yang 0011, Yi Yu 0001, Xinbo Gao 0001
Neurocomputing5
2020 SIST: Online Scale-Adaptive Object tracking with Stepwise Insight
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Jun Chen 0001, Ruimin Hu
Neurocomputing3
2020 Cross-resolution face recognition with pose variations via multilayer locality-constrained structural orthogonal procrustes regression
Guangwei Gao, Yi Yu 0001, Meng Yang 0001, Heyou Chang, Dong Yue 0001
Inf. Sci.2
2020 Context-Patch Face Hallucination Based on Thresholding Locality-Constrained Representation and Reproducing Learning
abstract
Face hallucination is a technique that reconstructs high-resolution (HR) faces from low-resolution (LR) faces, by using the prior knowledge learned from HR/LR face pairs. Most state-of-the-arts leverage position-patch prior knowledge of the human face to estimate the optimal representation coefficients for each image patch. However, they focus only the position information and usually ignore the context information of the image patch. In addition, when they are confronted with misalignment or the small sample size (SSS) problem, the hallucination performance is very poor. To this end, this paper incorporates the contextual information of the image patch and proposes a powerful and efficient context-patch-based face hallucination approach, namely, thresholding locality-constrained representation and reproducing learning (TLcR-RL). Under the context-patch-based framework, we advance a thresholding-based representation method to enhance the reconstruction accuracy and reduce the computational complexity. To further improve the performance of the proposed algorithm, we propose a promotion strategy called reproducing learning. By adding the estimated HR face to the training set, which can simulate the case that the HR version of the input LR face is present in the training set, it thus iteratively enhances the final hallucination result. Experiments demonstrate that the proposed TLcR-RL method achieves a substantial increase in the hallucinated results, both subjectively and objectively. In addition, the proposed framework is more robust to face misalignment and the SSS problem, and its hallucinated HR face is still very good when the LR test face is from the real world. The MATLAB source code is available at https://github.com/junjun-jiang/TLcR-RL.
Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Akiko Aizawa, Kiyoharu Aizawa
IEEE Trans. Cybern.2
2020 Ensemble Super-Resolution With a Reference Dataset
abstract
By developing sophisticated image priors or designing deep(er) architectures, a variety of image super-resolution (SR) approaches have been proposed recently and achieved very promising performance. A natural question that arises is whether these methods can be reformulated into a unifying framework and whether this framework assists in SR reconstruction? In this paper, we present a simple but effective single image SR method based on ensemble learning, which can produce a better performance than that could be obtained from any of SR methods to be ensembled (or called component super-resolvers). Based on the assumption that better component super-resolver should have larger ensemble weight when performing SR reconstruction, we present a maximum a posteriori (MAP) estimation framework for the inference of optimal ensemble weights. Especially, we introduce a reference dataset, which is composed of high-resolution (HR) and low-resolution (LR) image pairs, to measure the SR abilities (prior knowledge) of different component super-resolvers. To obtain the optimal ensemble weights, we propose to incorporate the reconstruction constraint, which states that the degenerated HR estimation should be equal to the LR observation one, as well as the prior knowledge of ensemble weights into the MAP estimation framework. Moreover, the proposed optimization problem can be solved by an analytical solution. We study the performance of the proposed method by comparing with different competitive approaches, including four state-of-the-art nondeep learning-based methods, four latest deep learning-based methods, and one ensemble learning-based method, and prove its effectiveness and superiority on some general image datasets and face image datasets.
Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Suhua Tang, Ruimin Hu, Jiayi Ma 0001
IEEE Trans. Cybern.2
2020 Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-Modal Retrieval
abstract
Cross-modal retrieval aims to retrieve data in one modality by a query in another modality, which has been a very interesting research issue in the field of multimedia, information retrieval, and computer vision, and database. Most existing works focus on cross-modal retrieval between text-image, text-video, and lyrics-audio. Little research addresses cross-modal retrieval between audio and video due to limited audio-video paired datasets and semantic information. The main challenge of the audio-visual cross-modal retrieval task focuses on learning joint embeddings from a shared subspace for computing the similarity across different modalities, where generating new representations is to maximize the correlation between audio and visual modalities space. In this work, we propose TNN-C-CCA, a novel deep triplet neural network with cluster canonical correlation analysis, which is an end-to-end supervised learning architecture with an audio branch and a video branch. We not only consider the matching pairs in the common space but also compute the mismatching pairs when maximizing the correlation. In particular, two significant contributions are made. First, a better representation by constructing a deep triplet neural network with triplet loss for optimal projections can be generated to maximize correlation in the shared subspace. Second, positive examples and negative examples are used in the learning stage to improve the capability of embedding learning between audio and video. Our experiment is run over fivefold cross validation, where average performance is applied to demonstrate the performance of audio-video cross-modal retrieval. The experimental results achieved on two different audio-visual datasets show that the proposed learning architecture with two branches outperforms existing six canonical correlation analysis–based methods and four state-of-the-art-based cross-modal retrieval methods.
Donghuo Zeng, Yi Yu 0001, Keizo Oyama
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Face hallucination through differential evolution parameter map learning with facial structure prior
Junjun Jiang, Jiayi Ma 0001, Suhua Tang, Yi Yu 0001, Kiyoharu Aizawa
Inf. Sci.4
2019 Graph-Regularized Locality-Constrained Joint Dictionary and Residual Learning for Face Sketch Synthesis
abstract
Face sketch synthesis is a crucial issue in digital entertainment and law enforcement. It can bridge the considerable texture discrepancy between face photos and sketches. Most of the current face sketch synthesis approaches directly to learn the relationship between the photos and sketches, and it is very difficult for them to generate the individual specific features, which we call rare characteristics. In this paper, we propose a novel face sketch synthesis approach through residual learning. In contrast to traditional approaches, which aim to reconstruct a sketch image directly (i.e., learn the mapping relationship between the photo and sketch), we aim to predict the residual image by learning the mapping relationship between the photo and residual, i.e., the difference between the photo and sketch, given an observed photo. This technique will render optimizing the residual mapping easier than optimizing the original mapping and deriving rare characteristic information. We also introduce a joint dictionary learning algorithm by preserving the local geometry structure of a data space. Through the learned joint dictionary, we transform the face sketch synthesis from an image space to a new and compact space; the new and compact space is spanned by learned dictionary atoms, where the manifold assumption can be further guaranteed. Results show that the proposed method demonstrates an impressive performance in the face sketch synthesis task on three public face sketch datasets and various real-world photos. These results are derived by comparing the proposed method with several state-of-the-art techniques, including certain recently proposed deep learning-based approaches.
Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Xianming Liu 0005, Jiayi Ma 0001
IEEE Trans. Image Process.2
2019 Incremental Re-Identification by Cross-Direction and Cross-Ranking Adaption
abstract
Person re-identification is widely applied in video surveillance and criminal investigation applications. To achieve better performance, an additional re-ranking step is often exploited. Related methods attempt to optimize the result according to every single query independently. However, in a practical scene, as the investigation process goes on, the other queries, in particular, the gradually accumulated logs, can be used to guide or regularize the current query. In this paper, we propose to optimize the result according to not only the current query itself but also the other queries and historical logs. We respectively investigate the cross-direction and the cross-ranking constraints among different queries. Based on the investigations, we propose a reciprocal optimization method to refine multiple ranking lists reciprocally. Experiments on the VIPeR, new-protocol CUHK03, and Market-1501 datasets confirm the effectiveness of our method. In particular, on the Market-1501 dataset, with full utilization of the other queries, the method achieves an accuracy rate of 94.66% at rank-1 and a very high mAP of 75.12%, and significantly outperforms the state-of-the-art methods.
Zheng Wang 0007, Junjun Jiang, Yi Yu 0001, Shin'ichi Satoh 0001
IEEE Trans. Multim.3
2019 Category-Based Deep CCA for Fine-Grained Venue Discovery From Multimodal Data
abstract
In this work, travel destinations and business locations are taken as venues. Discovering a venue by a photograph is very important for visual context-aware applications. Unfortunately, few efforts paid attention to complicated real images such as venue photographs generated by users. Our goal is fine-grained venue discovery from heterogeneous social multimodal data. To this end, we propose a novel deep learning model, category-based deep canonical correlation analysis. Given a photograph as input, this model performs: 1) exact venue search (find the venue where the photograph was taken) and 2) group venue search (find relevant venues that have the same category as the photograph), by the cross-modal correlation between the input photograph and textual description of venues. In this model, data in different modalities are projected to a same space via deep networks. Pairwise correlation (between different modality data from the same venue) for exact venue search and category-based correlation (between different modality data from different venues with the same category) for group venue search are jointly optimized. Because a photograph cannot fully reflect rich text description of a venue, the number of photographs per venue in the training phase is increased to capture more aspects of a venue. We build a new venue-aware multimodal data set by integrating Wikipedia featured articles and Foursquare venue photographs. Experimental results on this data set confirm the feasibility of the proposed method. Moreover, the evaluation over another publicly available data set confirms that the proposed method outperforms state of the arts for cross-modal retrieval between image and text.
Yi Yu 0001, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa
IEEE Trans. Neural Networks Learn. Syst.1
2019 Deep Cross-Modal Correlation Learning for Audio and Lyrics in Music Retrieval
abstract
Deep cross-modal learning has successfully demonstrated excellent performance in cross-modal multimedia retrieval, with the aim of learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal structures of different data modalities, such as audio and lyrics, should be taken into account. Stemming from the characteristic of temporal structures of music in nature, we are motivated to learn the deep sequential correlation between audio and lyrics. In this work, we propose a deep cross-modal correlation learning architecture involving two-branch deep neural networks for audio modality and text modality (lyrics). Data in different modalities are converted to the same canonical space where intermodal canonical correlation analysis is utilized as an objective function to calculate the similarity of temporal structures. This is the first study that uses deep architectures for learning the temporal correlation between audio and lyrics. A pretrained Doc2Vec model followed by fully connected layers is used to represent lyrics. Two significant contributions are made in the audio branch, as follows: (i) We propose an end-to-end network to learn cross-modal correlation between audio and lyrics, where feature extraction and correlation learning are simultaneously performed and joint representation is learned by considering temporal structures. (ii) And, as for feature extraction, we further represent an audio signal by a short sequence of local summaries (VGG16 features) and apply a recurrent neural network to compute a compact feature that better learns the temporal structures of music audio. Experimental results, using audio to retrieve lyrics or using lyrics to retrieve audio, verify the effectiveness of the proposed deep correlation learning architectures in cross-modal music retrieval.
Yi Yu 0001, Suhua Tang, Francisco Raposo 0001, Lei Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Video-Based Person Re-Identification via Self Paced Weighting
abstract
Person re-identification (re-id) is a fundamental technique to associate various person images, captured by differentsurveillance cameras, to the same person. Compared to the single image based person re-id methods, video-based personre-id has attracted widespread attentions because extra space-time information and more appearance cues that can beused to greatly improve the matching performance. However, most existing video-based person re-id methods equally treatall video frames, ignoring their quality discrepancy caused by object occlusion and motions, which is a common phenomenonin real surveillance scenario. Based on this finding, we propose a novel video-based person re-id method via self paced weighting (SPW). Firstly, we propose a self paced outlier detection method to evaluate the noise degree of video sub sequences. Thereafter, a weighted multi-pair distance metric learning approach is adopted to measure the distance of two person image sequences. Experimental results on two public datasets demonstrate the superiority of the proposed method over current state-of-the-art work.
Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Weijian Ruan, Ruimin Hu
AAAI3
2018 Residual Learning for Face Sketch Synthesis
abstract
Face sketch synthesis plays an important role in both digital entertainment and law enforcement. It can bridge the great texture discrepancy between face photos and sketches. Most of the current face sketch synthesis approaches directly learn the relationship between the photos and sketches, and it is very difficult for them to generate the individual specific details, which we call rare features. To address this problem, in this paper we propose a novel face sketch synthesis through residual learning. In contrast the traditional approaches, which try to construct the sketch image directly, we aim at predicting the residual image (between the photo and sketch), given the photo observation. In addition, we also introduce a couple dictionary learning algorithm through preserving the local geometry structure of data space, which is usually ignored by existing methods. Our proposed method shows impressive results on the face sketch synthesis task, when compared with some state-of-the-arts including some recent proposed deep learning based approaches.
Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Jiayi Ma 0001
ICASSP2
2018 Deep Knowledge Tracing and Dynamic Student Classification for Knowledge Tracing
abstract
In Intelligent Tutoring System (ITS), tracing the student's knowledge state during learning has been studied for several decades in order to provide more supportive learning instructions. In this paper, we propose a novel model for knowledge tracing that i) captures students' learning ability and dynamically assigns students into distinct groups with similar ability at regular time intervals, and ii) combines this information with a Recurrent Neural Network architecture known as Deep Knowledge Tracing. Experimental results confirm that the proposed model is significantly better at predicting student performance than well known state-of-the-art techniques for student modelling.
Sein Minn, Yi Yu 0001, Michel C. Desmarais, Feida Zhu 0001, Jill-Jênn Vie
ICDM2
2018 Deep CNN Denoiser and Multi-layer Neighbor Component Embedding for Face Hallucination
abstract
Most of the current face hallucination methods, whether they are shallow learning-based or deep learning-based, all try to learn a relationship model between Low-Resolution (LR) and High-Resolution (HR) spaces with the help of a training set. They mainly focus on modeling image prior through either model-based optimization or discriminative inference learning. However, when the input LR face is tiny, the learned prior knowledge is no longer effective and their performance will drop sharply. To solve this problem, in this paper we propose a general face hallucination method that can integrate model-based optimization and discriminative inference. In particular, to exploit the model based prior, the Deep Convolutional Neural Networks (CNN) denoiser prior is plugged into the super-resolution optimization model with the aid of image-adaptive Laplacian regularization. Additionally, we further develop a high-frequency details compensation method by dividing the face image to facial components and performing face hallucination in a multi-layer neighbor embedding manner. Experiments demonstrate that the proposed method can achieve promising super-resolution results for tiny input LR faces.
Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001
IJCAI2
2018 Deep Learning of Human Perception in Audio Event Classification
abstract
In this paper, we introduce our recent studies on human perception in audio event classification. In particular, the pre-trained model VGGish is used as feature extractor to process audio data, and DenseNet is trained by and used as feature extractor for our electroencephalography (EEG) data. The correlation between audio stimuli and EEG is learned in a shared space. In the experiments, we record brain activities (EEG signals) of several subjects while they are listening to music events of 8 audio categories selected from Google AudioSet. Our experimental results demonstrate that i) audio event classification can be improved by exploiting the power of human perception, and ii) the correlation between audio stimuli and EEG can be learned to complement audio event understanding.
Yi Yu 0001, Samuel Beuret, Donghuo Zeng, Keizo Oyama
ISM1
2018 Audio-Visual Embedding for Cross-Modal Music Video Retrieval through Supervised Deep CCA
abstract
Deep learning has successfully shown excellent performance in learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal structures of different data modalities, such as audio and video, should be taken into account. Music video retrieval by a given musical audio is a natural way to search and interact with music contents. In this work, we study cross-modal music video retrieval in terms of emotion similarity. Particularly, an audio of an arbitrary length is used to retrieve a longer or full-length music video. To this end, we propose a novel audio-visual embedding algorithm by Supervised Deep Canonical Correlation Analysis (S-DCCA) that projects audio and video into a shared space to bridge the semantic gap between audio and video. This also preserves the similarity among audio and visual contents from different videos with the same class label and the temporal structure. The contribution of our approach is mainly manifested in the two aspects: i) We propose to select top k audio chunks by attention-based Long Short-Term Memory (LSTM) model, which can represent good audio summarization with local properties. ii) We propose an end-to-end deep model for crossmodal audio-visual learning where S-DCCA is trained to learn the semantic correlation between audio and visual modalities. Due to the lack of music video dataset, we construct 10K music video dataset from YouTube 8M dataset. Some promising results such as MAP and precision-recall show that our proposed model can be applied to music video retrieval.
Donghuo Zeng, Yi Yu 0001, Keizo Oyama
ISM2
2018 Person Reidentification via Discrepancy Matrix and Matrix Metric
abstract
Person reidentification (re-id), as an important task in video surveillance and forensics applications, has been widely studied. Previous research efforts toward solving the person re-id problem have primarily focused on constructing robust vector description by exploiting appearance's characteristic, or learning discriminative distance metric by labeled vectors. Based on the cognition and identification process of human, we propose a new pattern, which transforms the feature description from characteristic vector to discrepancy matrix. In particular, in order to well identify a person, it converts the distance metric from vector metric to matrix metric, which consists of the intradiscrepancy projection and interdiscrepancy projection parts. We introduce a consistent term and a discriminative term to form the objective function. To solve it efficiently, we utilize a simple gradient-descent method under the alternating optimization process with respect to the two projections. Experimental results on public datasets demonstrate the effectiveness of the proposed pattern as compared with the state-of-the-art approaches.
Zheng Wang 0007, Ruimin Hu, Chen Chen 0001, Yi Yu 0001, Junjun Jiang, Chao Liang 0001, Shin'ichi Satoh 0001
IEEE Trans. Cybern.4
2017 Taichi distance for person re-identification
abstract
Metric learning is an important issue in person re-identification, and Mahalanobis-distance based metric learning methods prevail in this field. All of these approaches can be considered as equivalently projecting all samples to a new metric space and calculating the Euclidean distance there. However, the performance of distinguishing similar samples from dissimilar ones via absolute distance is limited. In this paper, we suggest using relative distance instead. We adopt a bi-target perspective. The core idea is to construct a virtual opposite target for each original target. Then, the similarity between a sample and the others is judged by using both the original and opposite targets of the sample. In this way, we propose a bi-target metric method, named TAICHI distance. Considering simplicity and efficiency, we follow the KISSME metric in this paper. Extensive evaluations on challenging datasets confirm the effectiveness of the proposed method.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Chen Chen 0001
ICASSP3
2017 Compact LBP and WLBP descriptor with magnitude and direction difference for face recognition
abstract
In this paper, we propose a novel descriptor for face recognition on grayscale images, depth images and 2D+depth images. It is a compact and effective descriptor computed from the magnitude and the direction difference. It can be concatenated with conventional descriptors such as well-known Local Binary Pattern (LBP) and Weber Local Binary Pattern (WLBP), to enhance their discrimination capability. To evaluate the performance of our descriptor, we conducted extensive experiments on three types of images using four different databases. The experimental results demonstrate the robustness and superiority of our approach, and the performances of our new descriptor surpass that without magnitude and direction difference. At the end, we further compare our descriptor with Convolution Neural Network (CNN) to show the compactness and effectiveness of the proposed approach.
Soo-Chang Pei, Mei-Shuo Chen, Yi Yu 0001, Suhua Tang, Chunlin Zhong
ICIP3
2017 Context-patch based face hallucination via thresholding locality-constrained representation and reproducing learning
abstract
Face hallucination, which refers to predicting a HighResolution (HR) face image from an observed Low-Resolution (LR) one, is a challenging problem. Most state-of-the-arts employ local face structure prior to estimate the optimal representations for each patch by the training patches of the same position, and achieve good reconstruction performance. However, they do not take into account the contextual information of image patch, which is very useful for the expression of human face. Different from position-patch based methods, in this paper we leverage the contextual information and develop a robust and efficient context-patch face hallucination algorithm, called Thresholding Locality-constrained Representation with Reproducing learning (TLcR-RL). In TLcR-RL, we use a thresholding strategy to enhance the stability of patch representation and the reconstruction accuracy. Additionally, we develop a reproducing learning to iteratively enhance the estimated result by adding the estimated HR face to the training set. Experiments demonstrate that the performance of our proposed framework has a substantial increase when compared to state-of-the-arts, including recently proposed deep learning based method.
Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Guo-Jun Qi, Akiko Aizawa
ICME2
2017 VenueNet: Fine-Grained Venue Discovery by Deep Correlation Learning
abstract
Venue photos, as a new type of multimedia contents, are exploding on the Internet because users like to take photos and share with their friends in which venue they spent time and what impressed them there. Discovering a venue by a social photo is very useful for supplementing venue retrieval and recommendation. However, little research focused on fine-grained venue discovery by leveraging multimodal venue dataset. In this paper, we present the first multimodal dataset specially built for venue discovery, which includes venue photos, descriptions, and categories. Using this dataset, we propose a novel framework for fine-grained venue discovery through correlating venue photos and descriptions, aiming to learn a VenueNet representing a knowledge base and association for venues and their properties in different modalities. In the training phase, visual and textual features of the same venues, by two sub-networks, are respectively mapped to a same semantic space, in which canonical correlation analysis (CCA) is applied to these features to train the two sub-networks. In the query phase, given a photo, its correlation with textual features in the dataset is analyzed to find the most similar venue. Experimental results verify the practicability of the Deep CCA model for fine-grained venue discovery from large-scale multimodal dataset.
Yi Yu 0001, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa
ISM1
2017 Statistical Inference of Gaussian-Laplace Distribution for Person Verification
abstract
Metric learning is an important issue in the person verification problem, which is to identify whether a pair of face or human body images is about the same person. Due to low running cost, the non-iterative statistical inference methods for metric learning show their efficiency and effectiveness to large scale datasets and on-line updating person verification applications. The KISSME method is a typical one that constructs the metric based on two assumptions that both of the discrepancy spaces of negative pairs and positive pairs should be Gaussian structures. However, we find that, in fact, the distribution of discrepancies of positive pairs might tend to the Laplace distribution rather than the Gaussian distribution. Based on this finding, we propose a metric learning method by exploiting Gaussian-Laplace distribution statistical inference, where the Gaussian distribution of negative discrepancies and the Laplace distribution of positive discrepancies are considered together. Experiments conducted on two human body datasets (VIPeR and Market-1501) and one face dataset (LFW) show its superiority in terms of effectiveness and efficiency as compared with the state-of-the-art approaches, no matter the appearance description is handcrafted or deep learned.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Junjun Jiang, Jiayi Ma 0001, Shin'ichi Satoh 0001
ACM Multimedia3
2017 Spatial-Aware Collaborative Representation for Hyperspectral Remote Sensing Image Classification
abstract
Representation-residual-based classifiers have attracted much attention in recent years in hyperspectral image (HSI) classification. How to obtain the optimal representa-tion coefficients for the classification task is the key problem of these methods. In this letter, spatial-aware collaborative representation (CR) is proposed for HSI classification. In order to make full use of the spatial-spectral information, we propose a closed-form solution, in which the spatial and spectral features are both utilized to induce the distance-weighted regularization terms. Different from traditional CR-based HSI classification algorithms, which model the spatial feature in a preprocessing or postprocessing stage, we directly incorporate the spatial information by adding a spatial regularization term to the representation objective function. The experimental results on three HSI data sets verify that our proposed approach outperforms the state-of-the-art classifiers.
Junjun Jiang, Chen Chen 0001, Yi Yu 0001, Xinwei Jiang, Jiayi Ma 0001
IEEE Geosci. Remote. Sens. Lett.3
2017 A query refinement framework for xml keyword search
Zhifeng Bao, Yi Yu 0001, Jian Shen 0001, Zhangjie Fu 0001
World Wide Web2
2016 Scale-Adaptive Low-Resolution Person Re-Identification via Learning a Discriminating Surface
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Junjun Jiang, Chao Liang 0001, Jinqiao Wang
IJCAI3
2016 PROMPT: Personalized User Tag Recommendation for Social Media Photos Leveraging Personal and Social Contexts
abstract
Social media platforms such as Flickr allow users to annotate photos with descriptive keywords, called, tags with the goal of making multimedia content easily understandable, searchable, and discoverable. However, manual annotation is very time-consuming and cumbersome for most users, which makes it difficult to search relevant photos. Moreover, predicted tags for a photo are not necessarily relevant to users' interests. Thus, it necessitates for an automatic tag prediction system that considers users' interests and describes objective aspects of the photo such as visual content and activities. To this end, this paper presents a tag recommendation system, called, PROMPT, that recommends personalized tags for a given photo leveraging personal and social contexts. Specifically, first, we determine a group of users who have similar tagging behavior as the user of the photo, which is very useful in recommending personalized tags. Next, we find candidate tags from visual content, textual metadata, and tags of neighboring photos, and recommends five most suitable tags. We initialize scores of the candidate tags using asymmetric tag co-occurrence probabilities and normalized scores of tags after neighbor voting, and later perform random walk to promote the tags that have many close neighbors and weaken isolated tags. Finally, we recommend top five user tags to the given photo. Experimental results on a Flickr dataset (46,700 photos in the test set and 28 million photos in the train set) with 1,540 unique user tags confirm that the proposed algorithm outperforms state-of-the-arts.
Rajiv Ratn Shah, Anupam Samanta, Yi Yu 0001, Suhua Tang, Roger Zimmermann
ISM4
2016 Videopedia: Lecture Video Recommendation for Educational Blogs Using Topic Modeling
Subhasree Basu, Yi Yu 0001, Vivek K. Singh 0003, Roger Zimmermann
MMM (1)2
2016 Camera Network Based Person Re-identification by Leveraging Spatial-Temporal Constraint and Multiple Cameras Relations
Wenxin Huang, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Xian Zhong, Chunjie Zhang 0001
MMM (1)4
2016 NEWSMAN: Uploading Videos over Adaptive Middleboxes to News Servers in Weak Network Infrastructures
Rajiv Ratn Shah, Mohamed Hefeeda, Roger Zimmermann, Khaled A. Harras, Cheng-Hsin Hsu, Yi Yu 0001
MMM (1)6
2016 Leveraging multimodal information for event summarization and concept-level sentiment analysis
Rajiv Ratn Shah, Yi Yu 0001, Akshay Verma, Suhua Tang, Anwar Dilawar Shaikh, Roger Zimmermann
Knowl. Based Syst.2
2016 Zero-Shot Person Re-identification via Cross-View Consistency
abstract
Person re-identification, aiming to identify images of the same person from various cameras configured in different places, has attracted much attention in the multimedia retrieval community. In this problem, choosing a proper distance metric is a crucial aspect, and many classic methods utilize a uniform learnt metric. However, their performance is limited due to ignoring the zero-shot and fine-grained characteristics presented in real person re-identification applications. In this paper, we investigate two consistencies across two cameras, which are cross-view support consistency and cross-view projection consistency. The philosophy behind it is that, in spite of visual changes in two images of the same person under two camera views, the support sets in their respective views are highly consistent, and after being projected to the same view, their context sets are also highly consistent. Based on the above phenomena, we propose a data-driven distance metric (DDDM) method, re-exploiting the training data to adjust the metric for each query-gallery pair. Experiments conducted on three public data sets have validated the effectiveness of the proposed method, with a significant improvement over three baseline metric learning methods. In particular, on the public VIPeR dataset, the proposed method achieves an accuracy rate of 42.09% at rank-1, which outperforms the state-of-the-art methods by 4.29%.
Zheng Wang 0007, Ruimin Hu, Chao Liang 0001, Yi Yu 0001, Junjun Jiang, Mang Ye, Jun Chen 0001, Qingming Leng
IEEE Trans. Multim.4
2016 Person Reidentification via Ranking Aggregation of Similarity Pulling and Dissimilarity Pushing
abstract
Person reidentification is a key technique to match different persons observed in nonoverlapping camera views. Many researchers treat it as a special object-retrieval problem, where ranking optimization plays an important role. Existing ranking optimization methods mainly utilize the similarity relationship between the probe and gallery images to optimize the original ranking list, but seldom consider the important dissimilarity relationship. In this paper, we propose to use both similarity and dissimilarity cues in a ranking optimization framework for person reidentification. Its core idea is that the true match should not only be similar to those strongly similar galleries of the probe, but also be dissimilar to those strongly dissimilar galleries of the probe. Furthermore, motivated by the philosophy of multiview verification, a ranking aggregation algorithm is proposed to enhance the detection of similarity and dissimilarity based on the following assumption: the true match should be similar to the probe in different baseline methods. In other words, if a gallery blue image is strongly similar to the probe in one method, while simultaneously strongly dissimilar to the probe in another method, it will probably be a wrong match of the probe. Extensive experiments conducted on public benchmark datasets and comparisons with different baseline methods have shown the great superiority of the proposed ranking optimization method.
Mang Ye, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Qingming Leng, Chunxia Xiao, Jun Chen 0001, Ruimin Hu
IEEE Trans. Multim.3
2015 TRACE: Linguistic-Based Approach for Automatic Lecture Video Segmentation Leveraging Wikipedia Texts
abstract
In multimedia-based e - learning systems, the accessibility and searchability of most lecture video content is still insufficient due to the unscripted and spontaneous speech of the speakers. Moreover, this problem becomes even more challenging when the quality of such lecture videos is not sufficiently high. To extract the structural knowledge of a multi-topic lecture video and thus make it easily accessible it is very desirable to divide each video into shorter clips by performing an automatic topic-wise video segmentation. To this end, this paper presents the TRACE system to automatically perform such a segmentation based on a linguistic approach using Wikipedia texts. TRACE has two main contributions: (i) the extraction of a novel linguistic-based Wikipedia feature to segment lecture videos efficiently, and (ii) the investigation of the late fusion of video segmentation results derived from state-of-the-art algorithms. Specifically for the late fusion, we combine confidence scores produced by the models constructed from visual, transcriptional, and Wikipedia features. According to our experiments on lecture videos from VideoLectures.NET and NPTEL, the proposed algorithm segments knowledge structures more accurately compared to existing state-of-the-art algorithms. The evaluation results are very encouraging and thus confirm the effectiveness of TRACE.
Rajiv Ratn Shah, Yi Yu 0001, Anwar Dilawar Shaikh, Roger Zimmermann
ISM2
2015 EventBuilder: Real-time Multimedia Event Summarization by Visualizing Social Media
abstract
Due to the ubiquitous availability of smartphones and digital cameras, the number of photos/videos online has increased rapidly. Therefore, it is challenging to efficiently browse multimedia content and obtain a summary of an event from a large collection of photos/videos aggregated in social media sharing platforms such as Flickr and Instagram. To this end, this paper presents the EventBuilder system that enables people to automatically generate a summary for a given event in real-time by visualizing different social media such as Wikipedia and Flickr. EventBuilder has two novel characteristics: (i) leveraging Wikipedia as event background knowledge to obtain more contextual information about an input event, and (ii) visualizing an interesting event in real-time with a diverse set of social media activities. According to our initial experiments on the YFCC100M dataset from Flickr, the proposed algorithm efficiently summarizes knowledge structures based on the metadata of photos/videos and Wikipedia articles.
Rajiv Ratn Shah, Anwar Dilawar Shaikh, Yi Yu 0001, Wenjing Geng, Roger Zimmermann, Gangshan Wu
ACM Multimedia3
2015 Multi-Level Fusion for Person Re-identification with Incomplete Marks
abstract
Most video surveillance suspect investigation systems rely on the videos taken in different camera views. Actually, besides the videos, in the investigation process, investigators also manually label some marks, which, albeit incomplete, can be quite accurate and helpful in identifying persons. This paper studies the problem of Person Re-identification with Incomplete Marks (PRIM), aiming at ranking the persons in the gallery according to both the videos and incomplete marks. This problem is solved by a multi-step fusion algorithm, which consists of three key steps: (i) The early fusing step exploits both visual features and marked attributes to predict a complete and precise attribute vector. (ii) Based on the statistical attribute d ominance and saliency phenomena, a dominance-saliency matching model is suggested for measuring the distance between attribute vectors. (iii) The gallery is ranked separately by using visual features and attribute vectors, and the overall ranking list is the result of a late fusion. Experiments conducted on VIPeR dataset have validated the effectiveness of the proposed method in all the three key steps. The results also show that through introducing marks, the retrieval accuracy is significantly improved.
Zheng Wang 0007, Ruimin Hu, Yi Yu 0001, Chao Liang 0001, Wenxin Huang
ACM Multimedia3
2015 On Generating Content-Oriented Geo Features for Sensor-Rich Outdoor Video Search
abstract
Advanced technologies in consumer electronics products have enabled individual users to record, share, and view videos on mobile devices. With the volume of videos increasing tremendously on the Internet, fast and accurate video search has attracted much research attention. A good similarity measure is a key component in a video retrieval system. Most of the existing solutions only rely on either the low-level visual features or the surrounding textual annotations. Those approaches often suffer from low recall as they are highly susceptible to changes in viewpoint, illumination, and noisy tags. By leveraging geo-metadata , more reliable and precise search results can be obtained. However, two issues remain challenging: (1) how to quantify the spatial relevance of videos with the visual similarity to generate a pertinent ranking of results according to users' needs, and (2) how to design a compact video representation that supports efficient indexing for fast video retrieval. In this study, we propose a novel video description which consists of (a) determining the geographic coverage of a video based on the camera's field-of-view and a pre-constructed geo-codebook, and (b) fusing video spatial relevance and region-aware visual similarities to achieve a robust video similarity measure. Toward a better encoding of a video's geo-coverage, we construct a geo-codebook by semantically segmenting a map into a collection of coherent regions. To evaluate the proposed technique we developed a video retrieval prototype. Experiments show that our proposed method improves the mean average precision by 4.6% ~ 10.5%, compared with existing approaches.
Yifang Yin, Yi Yu 0001, Roger Zimmermann
IEEE Trans. Multim.2
2014 ATLAS: Automatic Temporal Segmentation and Annotation of Lecture Videos Based on Modelling Transition Time
abstract
The number of lecture videos available is increasing rapidly, though there is still insufficient accessibility and traceability of lecture video contents. Specifically, it is very desirable to enable people to navigate and access specific slides or topics within lecture videos. To this end, this paper presents the ATLAS system for the VideoLectures.NET challenge (MediaMixer, transLectures) to automatically perform the temporal segmentation and annotation of lecture videos. ATLAS has two main novelties: (i) a SVMhmm model is proposed to learn temporal transition cues and (ii) a fusion scheme is suggested to combine transition cues extracted from heterogeneous information of lecture videos. According to our initial experiments on videos provided by VideoLectures.NET, the proposed algorithm is able to segment and annotate knowledge structures based on fusing temporal transition cues and the evaluation results are very encouraging, which confirms the effectiveness of our ATLAS system.
Rajiv Ratn Shah, Yi Yu 0001, Anwar Dilawar Shaikh, Suhua Tang, Roger Zimmermann
ACM Multimedia2
2014 ADVISOR: Personalized Video Soundtrack Recommendation by Late Fusion with Heuristic Rankings
abstract
Capturing videos anytime and anywhere, and then instantly sharing them online, has become a very popular activity. However, many outdoor user-generated videos (UGVs) lack a certain appeal because their soundtracks consist mostly of ambient background noise. Aimed at making UGVs more attractive, we introduce ADVISOR, a personalized video soundtrack recommendation system. We propose a fast and effective heuristic ranking approach based on heterogeneous late fusion by jointly considering three aspects: venue categories, visual scene, and user listening history. Specifically, we combine confidence scores, produced by SVMhmm models constructed from geographic, visual, and audio features, to obtain different types of video characteristics. Our contributions are threefold. First, we predict scene moods from a real-world video dataset that was collected from users' daily outdoor activities. Second, we perform heuristic rankings to fuse the predicted confidence scores of multiple models, and third we customize the video soundtrack recommendation functionality to make it compatible with mobile devices. A series of extensive experiments confirm that our approach performs well and recommends appealing soundtracks for UGVs to enhance the viewing experience.
Rajiv Ratn Shah, Yi Yu 0001, Roger Zimmermann
ACM Multimedia2
2014 Emerging Topics on Personalized and Localized Multimedia Information Systems
abstract
We are experiencing an era with a rapid increase of data relevant to different aspects of users' daily life. On the one hand, such data contains personal information of each individual user. On the other hand, it also reflects user behaviors related to the society as data of more users is aggregated. These data could not only be very beneficial for studying various lifestyle patterns, but also be used to generate more descriptive and explanatory analysis across the landscape of diverse multimedia data. Using personal mobile devices and web services to systematically explore interesting aspects of people world has attracted much attention recently. This is a full-day tutorial that addresses emerging topics on personalized and localized multimedia technologies and applications and emphasizes knowledge sensing and discovery in multimedia landscape. This tutorial aims to deliver anoverall introduction to multimedia landscapes with multimedia processing, contextual data acquisition, people activity logs, data analytics, geographic-aware multimedia sharing and delivery, and serves as an important lecture on fundamental and advanced research areas of personalized and localized multimedia information systems.
Yi Yu 0001, Kiyoharu Aizawa, Toshihiko Yamasaki, Roger Zimmermann
ACM Multimedia1
2014 WISMM'14 - First ACM International Workshop on Internet-Scale Multimedia Management
abstract
Advanced technologies in consumer electronic products have enabled individual users to record, transmit and receive images and videos with mobile devices. Every day people create and consume massive amounts of multimedia information and data by engaging with various mobile Internet services. With a wide variety of multimedia information and data around us being aggregated over time, the Internet is getting increasingly information-centric. We are experiencing an age of increasing demands on how we host various people's online engagements and how to augment people's lives in the physical world with more personalized smart services. This workshop addresses a focused but broad research theme with an emphasis on how to manage and derive value from multimedia data in the social Internet landscape to facilitate the connections between users' physical world and their online activities.
Roger Zimmermann, Yi Yu 0001
ACM Multimedia2
2014 User preference-aware music video generation based on modeling scene moods
abstract
Due to technical advances in mobile devices (e.g., smartphones, tablets) and wireless communications, people now can easily capture user-generated videos (UGVs) anywhere, anytime and instantly share their real-life experiences via social web sites. Enjoying videos has become very popular entertainment. One challenge is that many mobile videos do not have very appealing audio that was captured with the video. In this demonstration, to overcome this issue we propose a music video generation/creation system (Android app and backend system) that aims to make UGVs more attractive by generating scene-adaptive and user-preference aware music tracks. In our system, we take geographic categories, visual content and user listening history into account. In particular, the sequences of geographic categories and visual features are integrated into a SVMhmm model to predict video scene moods. The music genre, as a user preference is also exploited to personalize the recommended songs. We believe this is the first work that predicts scene moods from a real-world video dataset collected by users' daily outdoor recordings to facilitate user-preference aware music video generation. Our experiments confirm that our system can effectively combine objective scene moods and individual music tastes to recommend appealing soundtracks for videos. Our Android app only sends recorded sensor data and a few keyframes of a UGV to a cloud service (backend system) to retrieve recommended music tracks, therefore it is bandwidth efficient since the transmission of video data is not required for analysis.
Rajiv Ratn Shah, Yi Yu 0001, Roger Zimmermann
MMSys2
2014 A Probabilistic Associative Model for Segmenting Weakly Supervised Images
abstract
Weakly-supervised image segmentation is an important yet challenging task in image processing and pattern recognition fields. It is defined as: in the training stage, semantic labels are only at the image-level, without regard to their specific object/scene location within the image. Given a test image, the goal is to predict the semantics of every pixel/superpixel. In this paper, we propose a new weakly-supervised image segmentation model, focusing on learning the semantic associations between superpixel sets (graphlets in this work). In particular, we first extract graphlets from each image, where a graphlet is a small-sized graph measures the potential of multiple spatially neighboring superpixels (i.e., the probability of these superpixels sharing a common semantic label, such as the "sky" or the "sea"). To compare dierent-sized graphlets and to incorporate image-level labels, a manifold embedding algorithm is designed to transform all graphlets into equal-length feature vectors. Finally, we present a hierarchical Bayesian network (BN) to capture the semantic associations between post-embedding graphlets, based on which the semantics of each superpixel is inferred accordingly. Experimental results demonstrate that: 1) our approach performs competitively compared with the state-of-the-art approaches on three public data sets, and 2) considerable performance enhancement is achieved when using our approach on segmentation-based photo cropping and image categorization.
Yi Yang 0001, Yue Gao 0002, Yi Yu 0001, Changbo Wang, Xuelong Li 0001
IEEE Trans. Image Process.4
2013 Edge-based locality sensitive hashing for efficient geo-fencing application
abstract
Geo-fencing is a promising technique for emerging location-based services. Its two basic spatial predicates, INSIDE and WITHIN pairings between points and polygons, can be addressed by state-of-the-art methods such as the crossing number algorithm. In the era of big-data, however, geo-fencing has to process millions of points and hundreds of polygons or even more in real-time. In this paper, we propose an efficient algorithm to improve the scalability of geo-fencing, which consists of two main stages. At the first stage, an R-tree is used to quickly detect whether a point is inside the minimum bounding rectangle of a polygon. In the second stage, instead of an exhaustive search, we design an edge-based locality sensitive hashing scheme adapted to the crossing number algorithm. As for the case of WITHIN detection, a probing scheme is suggested to locate adjacent buckets so as to check all edges near to a target point. By further exploiting batch processing and multi-threading programming, our algorithm can achieve a fast speed while retaining 100% accuracy over all training datasets provided by the GIS Cup 2013 organizers.
Yi Yu 0001, Suhua Tang, Roger Zimmermann
SIGSPATIAL/GIS1
2013 Social interactions over geographic-aware multimedia systems
abstract
User-centric Internet multimedia scenes challenge us to discover more interesting events and topics based on matching users' needs associated with personal preferences, geographic interests and social norms. Geotagged multimedia contents from online social sites, e.g., Flickr and Twitter, provide large volumes of data about many given locations. Hence, location is one of the most important user-generated contexts and contains rich information about an individual's interests and behavior. Location-based social multimedia streams (e.g., tweets, videos, images) can provide us with socially complementary information to predict users' needs.
Roger Zimmermann, Yi Yu 0001
ACM Multimedia2
2013 Query-Document-Dependent Fusion: A Case Study of Multimodal Music Retrieval
abstract
In recent years, multimodal fusion has emerged as a promising technology for effective multimedia retrieval. Developing the optimal fusion strategy for different modalities (e.g., content, metadata) has been the subject of intensive research. Given a query, existing methods derive a unified fusion strategy for all documents with the underlying assumption that the relative significance of a modality remains the same across all documents. However, this assumption is often invalid. We thus propose a general multimodal fusion framework, query-document-dependent fusion (QDDF), which derives the optimal fusion strategy for each query-document pair via intelligent content analysis of both queries and documents. By investigating multimodal fusion strategies adaptive to both queries and documents, we demonstrate that existing multimodal fusion approaches are special cases of QDDF and propose two QDDF approaches to derive fusion strategies. The dual-phase QDDF explicitly derives and fuses query- and document-dependent weights, and the regression-based QDDF determines the fusion weight for a query-document pair via a regression model derived from training data. To evaluate the proposed approaches, comprehensive experiments have been conducted using a multimedia data set with around 17 K full songs and over 236 K social queries. Results indicate that the regression-based QDDF is superior in handling single-dimension queries. In comparison, the dual-phase QDDF outperforms existing approaches for most query types. We found that document-dependent weights are instrumental in enhancing multimedia fusion performance. In addition, efficiency analysis demonstrates the scalability of QDDF over large data sets.
Bingjun Zhang, Yi Yu 0001, Jialie Shen 0001, Ye Wang 0007
IEEE Trans. Multim.3
2013 Scalable Content-Based Music Retrieval Using Chord Progression Histogram and Tree-Structure LSH
abstract
With more and more multimedia content made available on the Internet, music information retrieval is becoming a critical but challenging research topic, especially for real-time online search of similar songs from websites. In this paper we study how to quickly and reliably retrieve relevant songs from a large-scale dataset of music audio tracks according to melody similarity. Our contributions are two-fold: (i) Compact and accurate representation of audio tracks by exploiting music semantics. Chord progressions are recognized from audio signals based on trained music rules, and the recognition accuracy is improved by multi-probing. A concise chord progression histogram (CPH) is computed from each audio track as a mid-level feature, which retains the discriminative capability in describing audio content. (ii) Efficient organization of audio tracks according to their CPHs by using only one locality sensitive hash table with a tree-structure. A set of dominant chord progressions of each song is used as the hash key. Average degradation of ranks is further defined to estimate the similarity of two songs in terms of their dominant chord progressions, and used to control the number of probing in the retrieval stage. Experimental results on a large dataset with 74,055 music audio tracks confirm the scalability of the proposed retrieval algorithm. Compared to state-of-the-art methods, our algorithm improves the accuracy of summarization and indexing, and makes a further step towards the optimal performance determined by an exhaustive sequence comparison.
Yi Yu 0001, Roger Zimmermann, Ye Wang 0007, Vincent Oria
IEEE Trans. Multim.1
2012 Recognition and Summarization of Chord Progressions and Their Application to Music Information Retrieval
abstract
Accurate and compact representation of music signals is a key component of large-scale content-based music applications such as music content management and near duplicate audio detection. This problem is not well solved yet despite many research efforts in this field. In this paper, we suggest mid-level summarization of music signals based on chord progressions. More specially, in our proposed algorithm, chord progressions are recognized from music signals based on a supervised learning model, and recognition accuracy is improved by locally probing n-best candidates. By investigating the properties of chord progressions, we further calculate a histogram from the probed chord progressions as a summary of the music signal. We show that the chord progression-based summarization is a powerful feature descriptor for representing harmonic progressions and tonal structures of music signals. The proposed algorithm is evaluated with content-based music retrieval as a typical application. The experimental results on a dataset with more than 70,000 songs confirm that our algorithm can effectively improve summarization accuracy of musical audio contents and retrieval performance, and enhance music retrieval applications on large-scale audio databases.
Yi Yu 0001, Roger Zimmermann, Ye Wang 0007, Vincent Oria
ISM1
2012 Automatic music soundtrack generation for outdoor videos from contextual sensor information
abstract
We present a system to automatically generate soundtracks for user-generated outdoor videos (UGV) based on concurrently captured contextual sensor information with mobile apps for the ACM Multimedia 2012 Google challenge: Automatic Music Video Generation. Our method addresses the use case of making "a video much more attractive for sharing by adding a matching soundtrack to it." Our system correlates viewable scene information from sensors with geographic contextual tags from OpenStreetMap. The co-occurance of geo-tags and mood tags are investigated from a set of categories of the web site Foursquare.com and a mapping from geo-tags to mood tags is obtained. Finally, a music retrieval component returns music based on matching mood tags. The experimental results show that our system can successfully create soundtracks that are related to the mood and situation of UGVs and therefore enhance the enjoyment of viewers. Our system sends only sensor data to a cloud service and is therefore bandwidth efficient since video data does not need to be transmitted for analysis.
Yi Yu 0001, Zhijie Shen, Roger Zimmermann
ACM Multimedia1
2010 Combining multi-probe histogram and order-statistics based LSH for scalable audio content retrieval
abstract
In order to improve the reliability and the scalability of content-based retrieval of variant audio tracks from large music databases, we suggest a new multi-stage LSH scheme that consists in (i) extracting compact but accurate representations from audio tracks by exploiting the LSH idea to summarize audio tracks, and (ii) adequately organizing the resulting representations in LSH tables, retaining almost the same accuracy as an exact kNN retrieval. In the first stage, we use major bins of successive chroma features to calculate a multi-probe histogram (MPH) that is concise but retains the information about local temporal correlations. In the second stage, based on the order statistics (OS) of the MPH, we propose a new LSH scheme, OS-LSH, to organize and probe the histograms. The representation and organization of the audio tracks are storage efficient and support robust and scalable retrieval. Extensive experiments over a large dataset with 30,000 real audio tracks confirm the effectiveness and efficiency of the proposed scheme.
Yi Yu 0001, Michel Crucianu, Vincent Oria, Ernesto Damiani
ACM Multimedia1
2009 Local summarization and multi-level LSH for retrieving multi-variant audio tracks
abstract
In this paper we study the problem of detecting and grouping multi-variant audio tracks in large audio datasets. To address this issue, a fast and reliable retrieval method is necessary. But reliability requires elaborate representations of audio content, which challenges fast retrieval by similarity from a large audio database. To find a better tradeoff between retrieval quality and efficiency, we put forward an approach relying on local summarization and multi-level Locality-Sensitive Hashing (LSH). More precisely, each audio track is divided into multiple Continuously Correlated Periods (CCP) of variable length according to spectral similarity. The description for each CCP is calculated based on its Weighted Mean Chroma (WMC). A track is thus represented as a sequence of WMCs. Then, an adapted two-level LSH is employed for efficiently delineating a narrow relevant search region. The "coarse" hashing level restricts search to items having a non-negligible similarity to the query. The subsequent, "refined" level only returns items showing a much higher similarity. Experimental evaluations performed on a real multi-variant audio dataset confirm that our approach supports fast and reliable retrieval of audio track variants.
Yi Yu 0001, Michel Crucianu, Vincent Oria, Lei Chen 0002
ACM Multimedia1
2008 Indexing high-dimensional data in dual distance spaces: a symmetrical encoding approach
abstract
Due to the well-known dimensionality curse problem, search in a high-dimensional space is considered as a "hard" problem. In this paper, a novel symmetrical encoding-based index structure, which is called EHD-Tree (for symmetrical Encoding-based Hybrid Distance Tree), is proposed to support fast k-Nearest-Neighbor (k-NN) search in high-dimensional spaces. In an EHD-Tree, all data points are first grouped into clusters by a k-Means clustering algorithm. Then the uniform ID number of each data point is obtained by a dual-distance-driven encoding scheme in which each cluster sphere is partitioned twice according to the dual distances of start- and centroid-distance. Finally, the uniform ID number and the centroid-distance of each data point are combined to get a uniform index key, the latter is then indexed through a partition-based B+-tree. Thus, given a query point, its k-NN search in high-dimensional spaces can be transformed into search in a single dimensional space with the aid of the EHD-Tree index. Extensive performance studies are conducted to evaluate the effectiveness and efficiency of our proposed scheme, and the results demonstrate that this method outperforms the state-of-the-art high dimensional search techniques such as the X-Tree, VA-file, iDistance and NB-Tree, especially when the query radius is not very large.
Yi Zhuang 0001, Yueting Zhuang, Qing Li 0001, Lei Chen 0002, Yi Yu 0001
EDBT5
2008 Using Exact Locality Sensitive Mapping to Group and Detect Audio-Based Cover Songs
abstract
Cover song detection is becoming a very hot research topic when plentiful personal music recordings or performance are released on the Internet. A nice cover song recognizer helps us group and detect cover songs to improve the searching experience. The traditional detection is to match two musical audio sequences by exhaustive pairwise comparisons. Different from the existing work, our aim is to generate a group of concatenated feature sets based on regression modeling and arrange them by indexing-based approximate techniques to avoid complicated audio sequence comparisons. We mainly focus on using exact locality sensitive mapping (ELSM) to join the concatenated feature sets and soft hash values. Similarity-invariance among audio sequence comparison is applied to define an optimal combination of several audio features. Soft hash values are pre-calculated to help locate searching range more accurately. Furthermore, we implement our algorithms in analyzing the real audio cover songs and grouping and detecting a batch of relevant cover songs embedded in large audio datasets.
Yi Yu 0001, J. Stephen Downie, Fabian Mörchen, Lei Chen 0002, Kazuki Joe
ISM1
2008 COSIN: content-based retrieval system for cover songs
abstract
We develop a content-based audio COver Song IdeNtification (COSIN) system to detect/group cover songs.The COSIN takes music audio content as input and performs similarity searching to locate variants of the input (i.e., cover versions). Identified cover songs are returned in the rank order according to their similarity to the input.The COSIN also incorporates a set of tools to evaluate retrieval performance so researchers can explore different retrieval schemes and parameters (e.g. recall, precision).The COSIN utilizes a suite of techniques to detect cover songs including: Pitch + Dynamic Programming (DP), Chroma + DP, and Semantic Feature Summarization (SFS) + Hash-Based Approximate Matching (HBAM). Demonstration system shows that COSIN is a very potential music content retrieval tool. Running some music retrieval schemes on COSIN platform, recent experiments with SFS + LSH Variants demonstrate a nicely balanced efficiency (search speed) v. performance (search accuracy) tradeoff.
Yi Yu 0001, J. Stephen Downie, Fabian Mörchen, Lei Chen 0002, Kazuki Joe, Vincent Oria
ACM Multimedia1
2007 Similarity Searching Techniques in Content-Based Audio Retrieval Via Hashing
Yi Yu 0001, Masami Takata, Kazuki Joe
MMM (1)1