Yue Yao 0001

dblp:78/10015-1 · DBLP profile ↗
← Back
18ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-9852-4667ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 9 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents
Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao 0001, Dongliang Xu
AAAI5
2026 Bipartite Mode Matching for Vision Training Set Search from a Hierarchical Data Server
abstract
We explore a situation in which the target domain is accessible, but real-time data annotation is not feasible. Instead, we would like to construct an alternative training set from a large-scale data server so that a competitive model can be obtained. For this problem, because the target domain usually exhibits distinct modes (i.e., semantic clusters representing data distribution), if the training set does not contain these target modes, the model performance would be compromised. While prior existing works improve algorithms iteratively, our research explores the often-overlooked potential of optimizing the structure of the data server. Inspired by the hierarchical nature of web search engines, we introduce a hierarchical data server, together with a bipartite mode matching algorithm (BMM) to align source and target modes. For each target mode, we look in the server data tree for the best mode match, which might be large or small in size. Through bipartite matching, we aim for all target modes to be optimally matched with source modes in a one-on-one fashion. Compared with existing training set search algorithms, we show that the matched server modes constitute training sets that have consistently smaller domain gaps with the target domain across object re-identification (re-ID) and detection tasks. Consequently, models trained on our searched training sets have higher accuracy than those trained otherwise. BMM allows data-centric unsupervised domain adaptation (UDA) orthogonal to existing model-centric UDA methods. By combining the BMM with existing UDA methods like pseudo-labeling, further improvement is observed.
Yue Yao 0001, Ruining Yang, Tom Gedeon
AAAI1
2026 Thought graph traversal for test-time scaling in chest X-ray VLLMs
Yue Yao 0001, Zelin Wen, Xuqing Li, Dongliang Xu, Tom Gedeon
Pattern Recognit.1
2025 Unsupervised Search for Ethnic Minorities' Medical Segmentation Training Set
abstract
This paper investigates the critical issue of dataset bias in medical imaging, with a particular emphasis on racial disparities caused by uneven population distribution in dataset collection. Our analysis reveals that medical segmentation datasets are significantly biased, primarily influenced by the demographic composition of their collection sites. For instance, Scanning Laser Ophthalmoscopy (SLO) fundus datasets collected in the United States predominantly feature images of White individuals, with minority racial groups underrepresented. This imbalance can result in biased model performance and inequitable clinical outcomes, particularly for minority populations. To address this challenge, we propose a novel training set search strategy aimed at reducing these biases by focusing on underrepresented racial groups. Our approach utilizes existing datasets and employs a simple greedy algorithm to identify source images that closely match the target domain distribution. By selecting training data that aligns more closely with the characteristics of minority populations, our strategy improves the accuracy of medical segmentation models on specific minorities, i.e., Black. Our experimental results demonstrate the effectiveness of this approach in mitigating bias. We also discuss the broader societal implications, highlighting how addressing these disparities can contribute to more equitable healthcare outcomes. Our code is available at https://github.com/yorkeyao/SnP.
Yue Yao 0001, Ruining Yang, Ashu Gupta, Tom Gedeon
ICASSP2
2024 Open-Set Facial Expression Recognition
abstract
Facial expression recognition (FER) models are typically trained on datasets with a fixed number of seven basic classes. However, recent research works (Cowen et al. 2021; Bryant et al. 2022; Kollias 2023) point out that there are far more expressions than the basic ones. Thus, when these models are deployed in the real world, they may encounter unknown classes, such as compound expressions that cannot be classified into existing basic classes. To address this issue, we propose the open-set FER task for the first time. Though there are many existing open-set recognition methods, we argue that they do not work well for open-set FER because FER data are all human faces with very small inter-class distances, which makes the open-set samples very similar to close-set samples. In this paper, we are the first to transform the disadvantage of small inter-class distance into an advantage by proposing a new way for open-set FER. Specifically, we find that small inter-class distance allows for sparsely distributed pseudo labels of open-set samples, which can be viewed as symmetric noisy labels. Based on this novel observation, we convert the open-set FER to a noisy label detection problem. We further propose a novel method that incorporates attention map consistency and cycle training to detect the open-set samples. Extensive experiments on various FER datasets demonstrate that our method clearly outperforms state-of-the-art open-set recognition methods by large margins. Code is available at https://github.com/zyh-uaiaaaa.
Yue Yao 0001, Xuannan Liu, Lixiong Qin, Weihong Deng
AAAI2
2024 Alice Benchmarks: Connecting Real World Re-Identification with the Synthetic
abstract
For object re-identification (re-ID), learning from synthetic data has become a promising strategy to cheaply acquire large-scale annotated datasets and effective models, with few privacy concerns. Many interesting research problems arise from this strategy, e.g., how to reduce the domain gap between synthetic source and real-world target. To facilitate developing more new approaches in learning from synthetic data, we introduce the Alice benchmarks, large-scale datasets providing benchmarks as well as evaluation protocols to the research community. Within the Alice benchmarks, two object re-ID tasks are offered: person and vehicle re-ID. We collected and annotated two challenging real-world target datasets: AlicePerson and AliceVehicle, captured under various illuminations, image resolutions, etc. As an important feature of our real target, the clusterability of its training set is not manually guaranteed to make it closer to a real domain adaptation test scenario. Correspondingly, we reuse existing PersonX and VehicleX as synthetic source domains. The primary goal is to train models from synthetic data that can work effectively in the real world. In this paper, we detail the settings of Alice benchmarks, provide an analysis of existing commonly-used domain adaptation methods, and discuss some interesting future directions. An online server has been set up for the community to evaluate methods conveniently and fairly. Datasets and the online server details are available at https://sites.google.com/view/alice-benchmarks.
Xiaoxiao Sun 0002, Yue Yao 0001, Shengjin Wang, Hongdong Li, Liang Zheng 0001
ICLR2
2024 Attribute Descent: Simulating Object-Centric Datasets on the Content Level and Beyond
abstract
This article aims to use graphic engines to simulate a large number of training data that have free annotations and possibly strongly resemble to real-world data. Between synthetic and real, a two-level domain gap exists, involving content level and appearance level. While the latter is concerned with appearance style, the former problem arises from a different mechanism, i.e., content mismatch in attributes such as camera viewpoint, object placement and lighting conditions. In contrast to the widely-studied appearance-level gap, the content-level discrepancy has not been broadly studied. To address the content-level misalignment, we propose an attribute descent approach that automatically optimizes engine attributes to enable synthetic data to approximate real-world data. We verify our method on object-centric tasks, wherein an object takes up a major portion of an image. In these tasks, the search space is relatively small, and the optimization of each attribute yields sufficiently obvious supervision signals. We collect a new synthetic asset VehicleX, and reformat and reuse existing the synthetic assets ObjectX and PersonX. Extensive experiments on image classification and object re-identification confirm that adapted synthetic data can be effectively used in three scenarios: training with synthetic data only, training data augmentation and numerically understanding dataset content.
Yue Yao 0001, Liang Zheng 0001, Xiaodong Yang 0001, Milind Napthade, Tom Gedeon
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Large-scale Training Data Search for Object Re-identification
abstract
We consider a scenario where we have access to the target domain, but cannot afford on-the-fly training data annotation, and instead would like to construct an alternative training set from a large-scale data pool such that a competitive model can be obtained. We propose a search and pruning (SnP) solution to this training data search problem, tailored to object re-identification (re-ID), an application aiming to match the same object captured by different cameras. Specifically, the search stage identifies and merges clusters of source identities which exhibit similar distributions with the target domain. The second stage, subject to a budget, then selects identities and their images from the Stage I output, to control the size of the resulting training set for efficient training. The two steps provide us with training sets 80% smaller than the source pool while achieving a similar or even higher re-ID accuracy. These training sets are also shown to be superior to a few existing search methods such as random sampling and greedy sampling under the same budget on training data size. If we release the budget, training sets resulting from the first stage alone allow even higher re-ID accuracy. We provide interesting discussions on the specificity of our method to the re-ID problem and particularly its role in bridging the re-ID domain gap. The code is available at https://github.com/yorkeyao/SnP
Yue Yao 0001, Tom Gedeon, Liang Zheng 0001
CVPR1
2021 EEG Feature Significance Analysis
Yue Yao 0001, Shafin Rahman, Tom Gedeon
ICONIP (6)2
2020 Simulating Content Consistent Vehicle Datasets with Attribute Descent
Yue Yao 0001, Liang Zheng 0001, Xiaodong Yang 0001, Milind Naphade, Tom Gedeon
ECCV (6)1
2020 Disguising Personal Identity Information in EEG Signals
Shiya Liu, Yue Yao 0001, Chaoyue Xing, Tom Gedeon
ICONIP (5)2
2020 Pairwise-GAN: Pose-Based View Synthesis Through Pair-Wise Training
Xuyang Shen, Jo Plested, Yue Yao 0001, Tom Gedeon
ICONIP (4)3
2020 Information-preserving feature filter for short-term EEG signals
Yue Yao 0001, Jo Plested, Tom Gedeon
Neurocomputing1
2019 Spotting Visual Keywords from Temporal Sliding Windows
abstract
Visual Keyword Spotting (KWS), as a newly proposed task deriving from visual speech recognition, has plenty of room for improvements. This paper details our Visual Keyword Spotting system used in the first Mandarin Audio-Visual Speech Recognition Challenge (MAVSR 2019). With the assumption that the vocabularies of target dataset are a subset of the vocabulary of the training set, we proposed a simple and scalable classification based strategy that achieves 19.0% mean average precision (mAP) on this challenge. Our method is based on the idea of using sliding windows to bridge between the word-level dataset and the sentence-level dataset, showing that a strong word level classifier can be directly used in building sentence embedding, thereby making it possible to build a KWS system.
Yue Yao 0001, Heming Du, Liang Zheng 0001, Tom Gedeon
ICMI1
2019 Generalized Alignment for Multimodal Physiological Signal Learning
abstract
Revealing the correspondences and relationships between physiological signals is attractive for bioinformatics and human-computer interaction. Time alignment is a straightforward way to figure out correspondences between time sequential data. However, alignment between multimodal physiological signals is hard to achieve because the similarity metrics are difficult to define if the two physiological signals being investigated are non-linearly correlated, misaligned or quite different in morphology. In this paper, we propose a generalized time alignment method for multimodal physiological signals which (i) learns the feature extractions on physiological signals in a generalized way, and (ii) enables learned features to be in a coordinated space where the similarity between sub-components from two signals can be defined. Furthermore, we applied our alignment based multimodal feature fusion on an evaluation model to perform emotion recognition tasks on the DEAP multimodal physiological signal dataset. The experimental results show that the alignment based feature fusion outperforms the non-aligned feature fusion in most cases.
Yuchi Liu, Yue Yao 0001, Jo Plested, Tom Gedeon
IJCNN2
2019 Improved Techniques for Building EEG Feature Filters
abstract
Recent advances in the generative adversarial network (GAN) based image translation have shown its potential of being an image style transformer. Similarly, defined as a style transformer for physiological signals, a feature filter is used to filter privacy-related features while still keeping useful features. However, existing feature filter techniques have three problems: (1) the privacy-related features cannot be filtered out to the extent we need through a simple Conv-Deconv generator structure, and (2) the generator cannot control the semantics (maintain desired features) of given physiological signals. To address these problems, we utilize deeper neural networks and adopt techniques from domain adaptation. This includes semantic loss and a GAN based model structure with two generators, two discriminators and a classifier to form a game of five. Our results on the UCI EEG dataset demonstrate that our model can simultaneously (1) achieve the state-of-the-art accuracy removal for the privacy-related feature, (2) reduce the desired feature removal accuracy drop, and (3) make the filtered signals can be interpreted or visually checked.
Yue Yao 0001, Jo Plested, Tom Gedeon, Yuchi Liu
IJCNN1
2018 Deep Feature Learning and Visualization for EEG Recording Using Autoencoders
Yue Yao 0001, Jo Plested, Tom Gedeon
ICONIP (7)1
2018 A Feature Filter for EEG Using Cycle-GAN Structure
Yue Yao 0001, Jo Plested, Tom Gedeon
ICONIP (7)1