Yanjie Sun

dblp:73/10353 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Face, body and person analysis · 33% Vision and language · 19% Language models and text generation · 19%
Computer graphics and multimedia
3 papers
Audio and music processing · 54% Visualization and visual analytics · 38% Visual content generation and editing · 9%
Software engineering, system software, and programming languages
1 paper
Software testing · 100%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
audio classification
1.422024
Automated Data Augmentation for Audio Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024
Automatic Audio Augmentation for Requests Sub-Challenge · ACM Multimedia 2023
Natural language and speech › Language models and text generation › natural language understanding
emotion understanding
0.912025
UniEmotion: A Unified Framework for Multimodal Emotion Recognition with Iterative Consensus-based Training · ACM Multimedia 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.912025
UniEmotion: A Unified Framework for Multimodal Emotion Recognition with Iterative Consensus-based Training · ACM Multimedia 2025
Computer vision › Face, body and person analysis
face recognition
0.812024
Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association · ACM Multimedia 2024
Computer vision › Face, body and person analysis › person identification
voice-face association
0.812024
Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association · ACM Multimedia 2024
Visualization and visual analytics › geospatial visualization
terrain visualization
0.812024
Adaptive Color Transfer From Images to Terrain Visualizations · IEEE Trans. Vis. Comput. Graph. 2024
Machine learning › Deep learning architectures and training
data augmentation
0.712023
Automatic Audio Augmentation for Requests Sub-Challenge · ACM Multimedia 2023
Software testing
combinatorial testing
0.612022
Enhance Combinatorial Testing With Metamorphic Relations · IEEE Trans. Software Eng. 2022
Software testing › metamorphic testing
metamorphic relations
0.612022
Enhance Combinatorial Testing With Metamorphic Relations · IEEE Trans. Software Eng. 2022
Software testing
metamorphic testing
0.612022
Enhance Combinatorial Testing With Metamorphic Relations · IEEE Trans. Software Eng. 2022
Software testing › test oracle
test oracle generation
0.612022
Enhance Combinatorial Testing With Metamorphic Relations · IEEE Trans. Software Eng. 2022
Machine learning › Representation and self-supervised learning
contrastive learning
0.212024
Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association · ACM Multimedia 2024
Machine learning › Representation and self-supervised learning › contrastive learning
multimodal contrastive learning
0.212024
Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association · ACM Multimedia 2024
Visualization and visual analytics › visual encoding
color mapping
0.212024
Adaptive Color Transfer From Images to Terrain Visualizations · IEEE Trans. Vis. Comput. Graph. 2024
Visual content generation and editing › style transfer
color transfer
0.212024
Adaptive Color Transfer From Images to Terrain Visualizations · IEEE Trans. Vis. Comput. Graph. 2024
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised audio representation learning
0.212023
Automatic Audio Augmentation for Requests Sub-Challenge · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.3multi-channel fusion · 1.3pseudo-labeling · 0.9iterative consensus-based training · 0.9consistency regularization · 0.9waveform-level augmentation · 0.8spectrogram-level augmentation · 0.8heuristic search · 0.8dual-objective optimization · 0.8contrastive learning · 0.8color grid organization · 0.8clustering · 0.8bayesian optimization · 0.8t-way coverage · 0.6metamorphic group generation · 0.6
YearPublicationVenuePosition
2025 UniEmotion: A Unified Framework for Multimodal Emotion Recognition with Iterative Consensus-based Training
abstract
Traditional emotion recognition methods struggle with complex emotional dynamics including multi-emotion states, transitions, and contextual reasoning. While multimodal large language models demonstrate great potential for understanding such complex scene dynamics, they still face challenges in adapting to emotion recognition tasks. We propose UniEmotion, a unified framework that simultaneously addresses conventional categorical emotion recognition, open-vocabulary fine-grained emotion recognition, and descriptive emotion understanding. Our approach leverages an iterative consensus-based training pipeline where pseudo-labels and model parameters co-evolve, maximizing large models' utility while mitigating downstream limitations. The framework integrates a selector module that identifies high-quality samples through prediction variance analysis, coupled with a pseudo-labeling module employing consistency regularization and class-wise adaptive mapping. This dual mechanism reduces error accumulation during self-training while aligning open-vocabulary output with task-specific labels. Experimental results demonstrate the effectiveness of our framework, achieving state-of-the-art performance across all three tracks, including 1st place on the MER-SEMI track with a significant improvement of 11.97% over the best baseline, and 2nd place on the MER-DES track.
Yanjie Sun, Wuyang Chen 0002, Yong Dou
ACM Multimedia1
2024 Self-Supervised Learning-Based General Fine-tuning Framework For Audio Classification and Event Detection
abstract
Recently, self-supervised learning (SSL) has made remarkable progress in signal representation and has become a de facto solution for different audio processing tasks. Generally, the SSL consists of the foundation pre-training and downstream fine-tuning phases. However, fine-tuning frameworks may lack universality due to the distinct learning paradigms and model designs employed in audio signal processing tasks. Furthermore, the varying degrees of dataset labeling across different tasks challenge unifying a fine-tuning framework. To address these issues, we propose vec2task, a cross-task general fine-tuning framework based on the SSL pre-trained model. It employs a semantic-aware module and an alternating training strategy, enabling the framework to generalize across various audio signal processing tasks. Additionally, the framework employs automatic audio augmentation strategies, eliminating the requirement for individually tailored algorithms to improve task performance. Experimental validations of the vec2task framework outperformed previous methods in audio classification and event detection tasks, showcasing its generalization ability across tasks.
Yanjie Sun, Kele Xu, Yong Dou
ICME1
2024 Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
Wuyang Chen 0002, Yanjie Sun, Kele Xu, Yong Dou
ACM Multimedia2
2024 Automated Data Augmentation for Audio Classification
abstract
Audio classification is a challenging task that requires categorizing audio data based on its content or characteristics. Existing approaches for audio classification rely either on supervised learning or fine-tuning based on self-supervised learning, both of which require manually labeled data. However, manually labeling audio datasets is a time-consuming and expensive process that limits the dataset's size. Moreover, the diversity of sound categories and class imbalances can further impede classification performance. To overcome these challenges, researchers have proposed various audio data augmentation methods. However, most of these methods focus less on augmentations combination and design and rely solely on waveform-based or spectrogram-based approaches. This paper presents an Automated Audio Augmentation (AAA) method for audio classification, which generates learnable and composable augmentation policies suitable for the audio classification task and can be employed in a plug-and-play manner. This method leverages both waveform-level and spectrogram-level augmentation, and a Bayesian optimization algorithm is proposed to search for composed augmentation policies. To the best of our knowledge, this is the first attempt to propose an automatic data augmentation method for audio classification tasks. Through large-scale empirical studies, we demonstrate that the proposed method outperforms previous competitive methods by a significant margin. We improve the average performance of multiple datasets by 6.421% and by 7.330% on few-shot scenarios, respectively.
Yanjie Sun, Kele Xu, Chaorun Liu, Yong Dou, Huaimin Wang 0001, Bo Ding 0001, Qinghua Pan
IEEE ACM Trans. Audio Speech Lang. Process.1
2024 Adaptive Color Transfer From Images to Terrain Visualizations
abstract
Terrain mapping is not only dedicated to communicating how high or steep a landscape is but can also help to indicate how we feel about a place. However, crafting effective and expressive elevation colors is challenging for both nonexperts and experts. In this article, we present a two-step image-to-terrain color transfer method that can transfer color from arbitrary images to diverse terrain models. First, we present a new image color organization method that organizes discrete, irregular image colors into a continuous, regular color grid that facilitates a series of color operations, such as local and global searching, categorical color selection and sequential color interpolation. Second, we quantify a series of cartographic concerns about elevation color crafting, such as the "lower, higher" principle, color conventions, and aerial perspectives. We also define color similarity between images and terrain visualizations with aesthetic quality. We then mathematically formulate image-to-terrain color transfer as a dual-objective optimization problem and offer a heuristic searching method to solve the problem. Finally, we compare elevation colors from our method with a standard color scheme and a representative color scale generation tool based on four test terrains. The evaluations show that the elevation colors from the proposed method are most effective and that our results are visually favorable. We also showcase that our method can transfer emotion from images to terrain visualizations.
Mingguang Wu, Yanjie Sun, Shangjing Jiang
IEEE Trans. Vis. Comput. Graph.2
2023 Automatic Audio Augmentation for Requests Sub-Challenge
abstract
This paper presents our solution for the Requests Sub-challenge of the ACM Multimedia 2023 Computational Paralinguistics Challenge. Drawing upon the framework of self-supervised learning, we put forth an automated data augmentation technique for audio classification, accompanied by a multi-channel fusion strategy aimed at enhancing overall performance. Specifically, to tackle the issue of imbalanced classes in complaint classification, we propose an audio data augmentation method that generates appropriate augmentation strategies for the challenge dataset. Furthermore, recognizing the distinctive characteristics of the dual-channel HC-C dataset, we individually evaluate the classification performance of the left channel, right channel, channel difference, and channel sum, subsequently selecting the optimal integration approach. Our approach yields a significant improvement in performance when compared to the competitive baselines, particularly in the context of the complaint task. Moreover, our method demonstrates noteworthy cross-task transferability.
Yanjie Sun, Kele Xu, Chaorun Liu, Yong Dou, Kun Qian 0003
ACM Multimedia1
2022 A Wall-Following Navigation Method for Autonomous Driving Based on Lidar in Tunnel Scenes
abstract
Wall following navigation uses the wall to guide the robot or vehicle to move from one position to another and always keep a certain distance from the wall in this process. It is of great significance in some aspects, such as rapid disease detection and auxiliary equipment detection of tunnel. This paper proposes an automatic wall following navigation method based on lidar to solve the existing problems in the tunnel scenes. Firstly, the mathematical model of spatial coordinate transformation among vehicle, point cloud, and the wall is established, and the RANSAC algorithm is used to improve the quality of point cloud of lidar. Then, according to the kinematic model of the autonomous vehicle, an improved pure pursuit wall following algorithm is proposed for lateral vehicle control, and the wall following mathematical model is established. Finally, the algorithm is verified in the two wall-following navigation cases under the tunnel wall, which shows that this method has a good tracking effect and stability.
Xiaobo Che, Yanjie Sun, Yanqiang Li
CSCWD3
2022 Enhance Combinatorial Testing With Metamorphic Relations
abstract
Due to the effectiveness and efficiency in detecting defects caused by interactions of multiple factors, Combinatorial Testing (CT) has received considerable scholarly attention in the last decades. Despite numerous practical test case generation techniques being developed, there remains a paucity of studies addressing the automated oracle generation problem, which holds back the overall automation of CT. As a consequence, much human intervention is inevitable, which is time-consuming and error-prone. This costly manual task also restricts the application of higher testing strength, inhibiting the full exploitation of CT in the industrial practice. To bridge the gap between test designs and fully automated test flows, and to extend the applicability of CT, this paper presents a novel CT methodology, named COMER, to enhance the traditional CT by accounting for Metamorphic Relations (MRs). COMER puts a high priority on generating pairs of test cases which match the input rules of MRs, i.e., the Metamorphic Group (MG), such that the correctness can be automatically determined by verifying whether the outputs of these test cases violate their MRs. As a result, COMER can not only satisfy the t-way coverage as what CT does, but also automatically check test oracle as many violations as possible. Several empirical studies conducted on 31 real-world software projects have shown that COMER increased the number of metamorphic groups by an average factor of 75.9 and also increased the failure detection rate by an average factor of 11.3, when compared with CT, while the overall number of test cases generated by COMER barely increased.
Xintao Niu, Yanjie Sun, Huayao Wu, Changhai Nie, Yu Lei 0001, Xiaoyin Wang
IEEE Trans. Software Eng.2
2012 A trust-augmented voting scheme for collaborative privacy management
abstract
Social networking sites have sprung up and become a hot issue of current society. In spite of the fact that these sites provide users with a variety of attractive features, much to users' dismay, however, they are prone to expose users' private information. In this paper, we propose an approach whi ch addresses the problem of collaboratively deciding privacy policies for, but not limited to, shared photos. Our approach utilizes trust relations in social networks and combines them with Condorcet's preferential voting scheme. We study properties of our trust-augmented voting scheme and develop two approximations to improve its efficiency. Our algorithms are compared and justified by experimental results, which support the usability of our trust-augmented voting scheme.
Yanjie Sun, Chenyi Zhang 0001, Jun Pang 0001, Baptiste Alcalde, Sjouke Mauw
J. Comput. Secur.1