EDBT 2026 Demo / reviewers in the wild / expert
Ying Cheng 0005
dblp:54/4536-5
· DBLP profile ↗
22ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-8964-3998ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EviMMQA: Multimodal question answering for medical evidence extraction in systematic reviews
Changkai Ji, Yingwen Wang, Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001 |
Pattern Recognit. | 5 |
| 2025 | AS-Det: Active Sampling for Adaptive 3D Object Detection in Point Cloudsabstract3D object detection in point clouds is critical in 3D computer vision, autonomous driving, and robotics. Existing point-based detectors, tailored to handle unstructured raw point clouds, often rely on simplistic sampling strategies to select a subset of points for local representation learning and detection. However, the diverse patterns exhibited by multiple types of point cloud data present a significant challenge to the universality of current detectors, particularly those captured by varied sensors (e.g., LiDAR and 4D Imaging Radar). In response to this challenge, we introduce an adaptable point-based single-stage 3D detector, AS-Det, engineered to excel on both LiDAR and 4D Radar point clouds. Specifically, we propose a novel active sampling strategy that actively mines object-related information to achieve efficient sampling and representation across different types of point clouds through end-to-end training. Additionally, we introduce a lightweight multi-scale center feature aggregation module to exploit multi-scale object context for precise and low-cost detection. By integrating the abovementioned modules, AS-Det achieves highly adaptive detection on various point clouds, encompassing different sensors and scales. Experimental results demonstrate the superior performance and adaptability of AS-Det on both LiDAR and 4D Radar point clouds. Ziheng Ding, Xiaze Zhang, Ying Cheng 0005, Rui Feng 0001 |
AAAI | 4 |
| 2025 | Fine-Grained Knowledge-Guided Alignment for Medical Vision-Language Pre-TrainingabstractMedical contrastive Vision-Language Pre-training (VLP) has emerged as a promising approach, enabling models to learn joint representations from paired medical images and radiology reports. Despite existing methods exploring local visual representation learning techniques, they often fall short in local alignment and knowledge infusion, e.g., uniform token treatment and isolated knowledge assignment. To address these issues, we propose a novel Fine-grained Knowledge-Guided Alignment (FKGA) framework for medical VLP. Specifically, we propose a Fine-grained Disease Knowledge Integration (FDKI) module to inject detailed disease descriptions into corresponding disease tokens in reports. Based on these semantic-enriched tokens, we introduce global instance-wise and local token-wise contrastive learning to further align the semantically related visual and textual modalities. In contrast to previous local visual representation learning methods, our design of semantic-enriched token alignment and context-preserved knowledge infusion enhances the semantic understanding of diseases. Extensive experimental results on five downstream tasks demonstrate that our proposed method outperforms other state-of-the-art methods across seven datasets. Yaning Pan, Ying Cheng 0005, Qingqiu Li, Runtian Yuan, Rui Feng 0001 |
BIBM | 2 |
| 2025 | RoBGuard: Enhancing LLMs to Assess Risk of Bias in Clinical Trial DocumentsabstractRandomized Controlled Trials (RCTs) are rigorous clinical studies crucial for reliable decision-making, but their credibility can be compromised by bias. The Cochrane Risk of Bias tool (RoB 2) assesses this risk, yet manual assessments are time-consuming and labor-intensive. Previous approaches have employed Large Language Models (LLMs) to automate this process. However, they typically focus on manually crafted prompts and a restricted set of simple questions, limiting their accuracy and generalizability. Inspired by the human bias assessment process, we propose RoBGuard, a novel framework for enhancing LLMs to assess the risk of bias in RCTs. Specifically, RoBGuard integrates medical knowledge-enhanced question reformulation, multimodal document parsing, and multi-expert collaboration to ensure both completeness and accuracy. Additionally, to address the lack of suitable datasets, we introduce two new datasets: RoB-Item and RoB-Domain. Experimental results demonstrate RoBGuard’s effectiveness on the RoB-Item dataset, outperforming existing methods. Changkai Ji, Yingwen Wang, Yuejie Zhang, Ying Cheng 0005, Rui Feng 0001 |
COLING | 6 |
| 2025 | SplitOcc: Multi-Resolution Sparse Voxel for Efficient LiDAR-Based Semantic Scene CompletionabstractLiDAR-based Semantic Scene Completion (SSC) is crucial for enhancing environmental perception and ensuring safety in autonomous driving. However, current methods face challenges in balancing accuracy and computational efficiency. On the one hand, projection-based methods reduce complexity but often suffer from spatial information loss. On the other hand, voxel-based methods preserve 3D structures but are computationally expensive. To address these limitations, we introduce SplitOcc, a novel multi-resolution approach that utilizes low-resolution voxels to represent large structures (e.g., road) and high-resolution voxels for detailed objects (e.g., bicycle). By employing multi-resolution sparse semantic voxels, SplitOcc can understand and represent the environment efficiently and accurately. Furthermore, the proposed Multi-Label Loss and Delayed-Drop strategies improve accuracy by preserving key semantic details during reconstruction. Extensive experiments demonstrate that our SplitOcc outperforms existing state-of-the-art methods across multiple evaluation metrics, showing notable improvements in both perception accuracy and detail preservation. Chaoyi Sun, Xiaze Zhang, Ziheng Ding, Ying Cheng 0005, Rui Feng 0001 |
ECAI | 4 |
| 2025 | Uncertainty-Aware Dynamic Fusion for Multimodal Clinical Prediction TasksabstractMultimodal fusion offers significant potential for enhancing medical diagnosis, particularly in the Intensive Care Unit (ICU), where integrating diverse data sources is crucial. Traditional static fusion models often fail to account for sample-wise variations in modality importance, which can impact prediction accuracy. To address this issue, we propose a dynamic Uncertainty-Aware Weighting (UAW) strategy that adaptively adjusts the importance of different modalities based on their reliability. This strategy is coupled with an Expert Ensemble Fusion (EEF) module, which leverages self-attention mechanisms and modality-specific FeedForward Networks (FFNs) to preserve and integrate critical information from various modalities. The proposed method demonstrates its efficacy through extensive experiments on phenotype classification and mortality prediction tasks, showing improved accuracy and robustness in handling diverse clinical data. Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001 |
ICASSP | 2 |
| 2025 | ConTrack3D: Contrastive Learning Contributes Concise 3D Multi-Object TrackingabstractOnline object detection and tracking are crucial for embodied intelligence systems, including autonomous vehicles and robotics. Traditional approaches employ a pipeline structure to perform detection and tracking separately, which can not fully leverage information from the detector. Moreover, most prior tracking methods rely on motion models such as constant velocity for state updates, which can lead to incorrect associations when the velocity estimates are inaccurate. To address these limitations, we propose ConTrack3D, an online tracking approach that jointly performs detection and tracking in an end-to-end manner. Specifically, ConTrack3D incorporates a Joint Encoder module to capture detection embeddings and a Temporal Extender module for data-driven state updates. By employing contrastive learning, ConTrack3D learns discriminative tracking representation for more accurate association. ConTrack3D is evaluated on the nuScenes benchmark, and the experimental results demonstrate its significant improvements in tracking performance. Ruibin Du, Ziheng Ding, Xiaze Zhang, Ying Cheng 0005, Rui Feng 0001 |
ICRA | 5 |
| 2025 | Semantic-Aware Hard Negative Mining for Medical Vision-Language Contrastive PretrainingabstractExisting medical vision-language contrastive pretraining methods aim to bring the paired image-report embeddings close together while pushing the unpaired ones apart. However, medical images often exhibit high inter-class visual similarity with only subtle differences, leading to the presence of hard negative samples that are semantically distinct from the anchor but incorrectly close to it in the embedding space, making it challenging to distinguish semantically dissimilar samples. Previous methods consider only the embedding similarity between samples to identify hard negatives, often wrongly treating false negatives as hard negatives. To address this issue, we design a simple yet effective approach called Semantic-Aware Hard Negative mining (SAHN), distinguishing hard negatives from false negatives and encouraging the model to pay greater attention to hard negatives. Specifically, hard negatives are identified as samples with high embedding similarity but low semantic similarity to the anchor and assigned greater importance weights. By integrating these importance weights into the InfoNCE loss, SAHN enhances the model's ability to separate semantically dissimilar samples while clustering semantically similar ones. We further conduct a gradient-based theoretical analysis to validate the effectiveness of SAHN. Extensive experimental results on four downstream medical tasks covering image classification, object detection, semantic segmentation, and cross-modal retrieval demonstrate the superiority of our approach. Ying Cheng 0005, Yaning Pan, Rui Feng 0001 |
ACM Multimedia | 2 |
| 2024 | ADSNet: Cross-Domain LTV Prediction with an Adaptive Siamese Network in AdvertisingabstractAdvertising platforms have evolved in estimating Lifetime Value (LTV) to better align with advertisers' true performance metric which considers cumulative sum of purchases a customer contributes over a period. Accurate LTV estimation is crucial for the precision of the advertising system and the effectiveness of advertisements. However, the sparsity of real-world LTV data presents a significant challenge to LTV predictive model(i.e., pLTV), severely limiting the their capabilities. Therefore, we propose to utilize external data, in addition to the internal data of advertising platform, to expand the size of purchase samples and enhance the LTV prediction model of the advertising platform. To tackle the issue of data distribution shift between internal and external platforms, we introduce an Adaptive Difference Siamese Network (ADSNet), which employs cross-domain transfer learning to prevent negative transfer. Specifically, ADSNet is designed to learn information that is beneficial to the target domain. We introduce a gain evaluation strategy to calculate information gain, aiding the model in learning helpful information for the target domain and providing the ability to reject noisy samples, thus avoiding negative transfer. Additionally, we also design a Domain Adaptation Module as a bridge to connect different domains, reduce the distribution distance between them, and enhance the consistency of representation space distribution. We conduct extensive offline experiments and online A/B tests on a real advertising platform. Our proposed ADSNet method outperforms other methods, improving GINI by 2%. The ablation study highlights the importance of the gain evaluation strategy in negative gain sample rejection and improving model performance. Additionally, ADSNet significantly improves long-tail prediction. The online A/B tests confirm ADSNet's efficacy, increasing online LTV by 3.47% and GMV by 3.89%. Ying Cheng 0005, Qi He 0011, Xing Zhou 0003, Rui Feng 0001, Jie Jiang 0008 |
KDD | 3 |
| 2024 | DeepPointMap2: Accurate and Robust LiDAR-Visual SLAM with Neural DescriptorsabstractSimultaneous Localization and Mapping (SLAM) plays a pivotal role in autonomous driving and robotics. Existing methods often rely on hand-craft feature extraction and cross-modal fusion techniques, resulting in limited feature representation capability and reduced robustness. To address this challenge, we introduce DeepPointMap2, a novel learning-based LiDAR-Visual SLAM architecture that leverages neural descriptors to tackle multiple SLAM sub-tasks in a unified manner. Our approach employs neural networks to extract multi-modal tokens, which are then adaptively fused by the Visual-Point Fusion Module to generate sparse 3D neural descriptors, ensuring precise and robust performance. As a pioneering work, our method achieves state-of-the-art localization performance among various Visual-, LiDAR-, and Visual-LiDAR-based methods in widely-used benchmarks, as shown in the experiment results. Furthermore, the approach proves to be robust in scenarios involving camera failure and LiDAR obstruction. Xiaze Zhang, Ziheng Ding, Ying Cheng 0005, Wenchao Ding 0001, Rui Feng 0001 |
ACM Multimedia | 4 |
| 2024 | CT2C-QA: Multimodal Question Answering over Chinese Text, Table and ChartabstractMultimodal Question Answering (MMQA) is crucial as it enables comprehensive understanding and accurate responses by integrating insights from diverse data representations such as tables, charts, and text. Most existing researches in MMQA only focus on two modalities such as image-text QA, table-text QA and chart-text QA, and there remains a notable scarcity in studies that investigate the joint analysis of text, tables, and charts. In this paper, we present CT2C-QA, a pioneering Chinese reasoning-based QA dataset that includes an extensive collection of text, tables, and charts, meticulously compiled from 200 selectively sourced webpages. Our dataset simulates real webpages and serves as a great test for the capability of the model to analyze and reason with multimodal data, because the answer to a question could appear in various modalities, or even potentially not exist at all. Additionally, we present AED (Allocating, Expert and Decision), a multi-agent system implemented through collaborative deployment, information interaction, and collective decision-making among different agents. Specifically, the Assignment Agent is in charge of selecting and activating expert agents, including those proficient in text, tables, and charts. The Decision Agent bears the responsibility of delivering the final verdict, drawing upon the analytical insights provided by these expert agents. We execute a comprehensive analysis, comparing AED with various state-of-the-art models in MMQA, including GPT-4. The experimental outcomes demonstrate that current methodologies, including GPT-4, are yet to meet the benchmarks set by our dataset. Tianhao Cheng, Yuejie Zhang, Ying Cheng 0005, Rui Feng 0001 |
ACM Multimedia | 4 |
| 2024 | Mixtures of Experts for Audio-Visual LearningabstractWith the rapid development of multimedia technology, audio-visual learning has emerged as a promising research topic within the field of multimodal analysis. In this paper, we explore parameter-efficient transfer learning for audio-visual learning and propose the Audio-Visual Mixture of Experts (\ourmethodname) to inject adapters into pre-trained models flexibly. Specifically, we introduce unimodal and cross-modal adapters as multiple experts to specialize in intra-modal and inter-modal information, respectively, and employ a lightweight router to dynamically allocate the weights of each expert according to the specific demands of each task. Extensive experiments demonstrate that our proposed approach \ourmethodname achieves superior performance across multiple audio-visual tasks,
including AVE, AVVP, AVS, and AVQA. Furthermore, visual-only experimental results also indicate that our approach can tackle challenging scenes where modality information is missing.
The source code is available at \url{https://github.com/yingchengy/AVMOE}. Ying Cheng 0005, Rui Feng 0001 |
NeurIPS | 1 |
| 2024 | Learning Music-Dance Representations Through Explicit-Implicit Rhythm SynchronizationabstractAlthough audio-visual representation has been proven to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of the dancer and music rhythm, we introduceMuDaR, a novelMusic-DanceRepresentation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin. Jiashuo Yu, Junfu Pu, Ying Cheng 0005, Rui Feng 0001, Ying Shan |
IEEE Trans. Multim. | 3 |
| 2022 | Self-Supervised Video Representation Learning with Motion-Contrastive PerceptionabstractVisual-only self-supervised learning has achieved significant improvement in video representation learning. Existing related methods encourage models to learn video representations by utilizing contrastive learning or designing specific pretext tasks. However, some models are likely to focus on the background, which is unimportant for learning video representations. To alleviate this problem, we propose a new view called long-range residual frame to obtain more motion-specific information. Based on this, we propose the Motion-Contrastive Perception Network (MCPNet), which consists of two branches, namely, Motion Information Perception (MIP) and Contrastive Instance Perception (CIP), to learn generic video representations by focusing on the changing areas in videos. Specifically, the MIP branch aims to learn fine-grained motion features, and the CIP branch performs contrastive learning to learn overall semantics information for each instance. Experiments on two benchmark datasets UCF-101 and HMDB-51 show that our method outperforms current state-of-the-art visual-only self-supervised approaches. Ying Cheng 0005, Yuejie Zhang, Rui Feng 0001 |
ICME | 2 |
| 2022 | IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-trainingabstractVision-Language Pre-training (VLP) with large-scale image-text pairs has demonstrated superior performance in various fields. However, the image-text pairs co-occurrent on the Internet typically lack explicit alignment information, which is suboptimal for VLP. Existing methods proposed to adopt an off-the-shelf object detector to utilize additional image tag information. However, the object detector is time-consuming and can only identify the pre-defined object categories, limiting the model capacity. Inspired by the observation that the texts incorporate incomplete fine-grained image information, we introduce IDEA, which stands for increasing text diversity via online multi-label recognition for VLP. IDEA shows that multi-label learning with image tags extracted from the texts can be jointly optimized during VLP. Moreover, IDEA can identify valuable image tags online to provide more explicit textual supervision. Comprehensive experiments demonstrate that IDEA can significantly boost the performance on multiple downstream datasets with a small extra computational cost. Youcai Zhang, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang, Yandong Guo |
ACM Multimedia | 3 |
| 2022 | MM-Pyramid: Multimodal Pyramid Attentional Network for Audio-Visual Event Localization and Video ParsingabstractRecognizing and localizing events in videos is a fundamental task for video understanding. Since events may occur in auditory and visual modalities, multimodal detailed perception is essential for complete scene comprehension. Most previous works attempted to analyze videos from a holistic perspective. However, they do not consider semantic information at multiple scales, which makes the model difficult to localize events in different lengths. In this paper, we present a Multimodal Pyramid Attentional Network (MM-Pyramid ) for event localization. Specifically, we first propose the attentive feature pyramid module. This module captures temporal pyramid features via several stacking pyramid units, each of them is composed of a fixed-size attention block and dilated convolution block. We also design an adaptive semantic fusion module, which leverages a unit-level attention block and a selective fusion block to integrate pyramid features interactively. Extensive experiments on audio-visual event localization and weakly-supervised audio-visual video parsing tasks verify the effectiveness of our approach. Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang |
ACM Multimedia | 2 |
| 2022 | Modality-aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence DetectionabstractWeakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early or intermediate manner, yet overlooking the modality heterogeneousness over the weakly-supervised setting. In this paper, we analyze the modality asynchrony and undifferentiated instances phenomena of the multiple instance learning (MIL) procedure, and further investigate its negative impact on weakly-supervised audio-visual learning. To address these issues, we propose a modality-aware contrastive instance learning with self-distillation (MACIL-SD) strategy . Specifically, we leverage a lightweight two-stream network to generate audio and visual bags, in which unimodal background, violent, and normal instances are clustered into semi-bags in an unsupervised way. Then audio and visual violent semi-bag representations are assembled as positive pairs, and violent semi-bags are combined with background and normal instances in the opposite modality as contrastive negative pairs. Furthermore, a self-distillation module is applied to transfer unimodal visual knowledge to the audio-visual model, which alleviates noises and closes the semantic gap between unimodal and multimodal features. Experiments show that our framework outperforms previous methods with lower complexity on the large-scale XD-Violence dataset. Results also demonstrate that our proposed approach can be used as plug-in modules to enhance other networks. Codes are available at https://github.com/JustinYuu/MACIL_SD. Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang |
ACM Multimedia | 3 |
| 2021 | Improving Multimodal Speech Enhancement by Incorporating Self-Supervised and Curriculum LearningabstractSpeech enhancement in realistic scenarios still remains many challenges, such as complex background signals and data limitations. In this paper, we present a co-attention based framework that incorporates self-supervised and curriculum learning to derive the target speech in noisy environments. Specifically, we first leverage self-supervision to pre-train the co-attention model on the task of audio-visual synchronization. The pre-trained model can focus on the lip of speakers automatically, and then the self-supervised features from the model are combined with a u-net regression network to separate the spectrograms of sound mixtures. To make the training process easier and further improve the performance, we introduce the curriculum learning scheme for the training stage of speech enhancement. Extensive experiments show that our model achieves superior performance over previous self-supervised method for speech enhancement, and demonstrate the generalizability of our approach to the transferred dataset. Ying Cheng 0005, Mengyu He, Jiashuo Yu, Rui Feng 0001 |
ICASSP | 1 |
| 2021 | MPN: Multimodal Parallel Network for Audio-Visual Event LocalizationabstractAudio-visual event localization aims to localize an event that is both audible and visible in the wild, which is a widespread audio-visual scene analysis task for unconstrained videos. To address this task, we propose a Multimodal Parallel Network (MPN), which can perceive global semantics and unmixed local information parallelly. Specifically, our MPN framework consists of a classification subnetwork to predict event categories and a localization subnetwork to predict event boundaries. The classification subnetwork is constructed by the Multimodal Co-attention Module (MCM) and obtains global contexts. The localization subnetwork consists of Multimodal Bottleneck Attention Module (MBAM), which is designed to extract fine-grained segment-level contents. Extensive experiments demonstrate that our framework achieves the state-of-the-art performance both in fully supervised and weakly supervised settings on the Audio-Visual Event (AVE) dataset. Jiashuo Yu, Ying Cheng 0005, Rui Feng 0001 |
ICME | 2 |
| 2021 | Exploring Logical Reasoning for Referring Expression ComprehensionabstractReferring expression comprehension aims to localize the target object in an image referred by a natural language expression. Most existing approaches neglect the implicit logical correlations among fine-grained cues, e.g., categories, attributes, which are beneficial for distinguishing objects. In this paper, we propose a logic-guided approach to explore logical knowledge for referring expression comprehension in a hierarchical modular-based framework. Specifically, we propose to extract fine-grained cues in visual and textual domains and perform logical reasoning over them with explicit logical expressions to regularize the matching process without extra parameters. Besides, we propose to improve existing modular-based methods by introducing context information of objects in the relationship module. Extensive experiments are conducted on three referring expression datasets, and the results demonstrate that our model can produce more consistent predictions and further achieve superior performance compared with previous methods. Ying Cheng 0005, Jiashuo Yu, Yuejie Zhang, Rui Feng 0001 |
ACM Multimedia | 1 |
| 2020 | Keep it Consistent: Topic-Aware Storytelling from an Image Stream via Iterative Multi-agent CommunicationabstractVisual storytelling aims to generate a narrative paragraph from a sequence of images automatically.Existing approaches construct text description independently for each image and roughly concatenate them as a story, which leads to the problem of generating semantically incoherent content.In this paper, we propose a new way for visual storytelling by introducing a topic description task to detect the global semantic context of an image stream.A story is then constructed with the guidance of the topic description.In order to combine the two generation tasks, we propose a multi-agent communication framework that regards the topic description generator and the story generator as two agents and learn them simultaneously via iterative updating mechanism.We validate our approach on VIST dataset, where quantitative results, ablations, and human evaluation demonstrate our method's good ability in generating stories with higher quality compared to state-of-the-art methods. Zhongyu Wei, Ying Cheng 0005, Piji Li, Haijun Shan, Ji Zhang 0011, Qi Zhang 0001, Xuanjing Huang 0001 |
COLING | 3 |
| 2020 | Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation LearningabstractWhen watching videos, the occurrence of a visual event is often accompanied by an audio event, e.g., the voice of lip motion, the music of playing instruments. There is an underlying correlation between audio and visual events, which can be utilized as free supervised information to train a neural network by solving the pretext task of audio-visual synchronization. In this paper, we propose a novel self-supervised framework with co-attention mechanism to learn generic cross-modal representations from unlabelled videos in the wild, and further benefit downstream tasks. Specifically, we explore three different co-attention modules to focus on discriminative visual regions correlated to the sounds and introduce the interactions between them. Experiments show that our model achieves state-of-the-art performance on the pretext task while having fewer parameters compared with existing methods. To further evaluate the generalizability and transferability of our approach, we apply the pre-trained model on two downstream tasks, i.e., sound source localization and action recognition. Extensive experiments demonstrate that our model provides competitive results with other self-supervised methods, and also indicate that our approach can tackle the challenging scenes which contain multiple sound sources. Ying Cheng 0005, Zhihao Pan, Rui Feng 0001, Yuejie Zhang |
ACM Multimedia | 1 |