Yu Yin 0001

dblp:83/4081-1 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-9588-5854ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Assessing LLMs for Serendipity Discovery in Knowledge Graphs: A Case for Drug Repurposing
abstract
Large Language Models (LLMs) have greatly advanced knowledge graph question answering (KGQA), yet existing systems are typically optimized for returning highly relevant but predictable answers. A missing yet desired capacity is to exploit LLMs to suggest surprise and novel ("serendipitious") answers. In this paper, we formally define the serendipity-aware KGQA task and propose the SerenQA framework to evaluate LLMs' ability to uncover unexpected insights in scientific KGQA tasks. SerenQA includes a rigorous serendipity metric based on relevance, novelty, and surprise, along with an expert-annotated benchmark derived from the Clinical Knowledge Graph for drug repurposing. Additionally, it features a structured evaluation pipeline encompassing three subtasks: knowledge retrieval, subgraph reasoning, and serendipity exploration. Our experiments reveal that while state-of-the-art LLMs perform well on retrieval, they still struggle to identify genuinely surprising and valuable discoveries, underscoring a significant room for future research.
Mengying Wang 0001, Chenhui Ma, Ao Jiao, Tuo Liang, Pengjun Lu, Shrinidhi Hegde, Yu Yin 0001, Evren Gurkan-Cavusoglu, Yinghui Wu 0001
AAAI7
2026 When 'Yes' Meets 'But': Can AI Comprehend Contradictory Humor in Comics?
abstract
Understanding humor, especially when it involves complex and contradictory narratives requiring comparative reasoning ability, remains a significant challenge for large vision-language models (VLMs). This limitation hinders AI's ability to engage in human-like reasoning and cultural expression. In this paper, we investigate this challenge through an in-depth analysis of comics that juxtapose panels to create humor through contradictions. We introduce the YesBut, a novel benchmark with 1,262 comic images from diverse multilingual and multicultural contexts, featuring comprehensive annotations that capture various aspects of narrative understanding. Using this benchmark, we systematically evaluate a wide range of VLMs through four complementary tasks spanning from surface content comprehension to deep narrative reasoning, with particular emphasis on comparative reasoning between contradictory elements. Our extensive experiments reveal that even the most advanced models significantly underperform compared to humans, with common failures in visual perception, key element identification, comparative analysis, and hallucinations. We further investigate text-based training strategies and social knowledge augmentation methods to enhance model performance. Our findings not only highlight critical weaknesses in VLMs' understanding of cultural and creative expressions but also provide pathways toward developing context-aware models capable of deeper narrative understanding through comparative reasoning.
Tuo Liang, Jing Li 0049, Yiren Lu 0002, Yunlai Zhou, Yiran Qiao 0001, Disheng Liu, Jierui Peng, Jing Ma 0002, Yu Yin 0001
IEEE Trans. Pattern Anal. Mach. Intell.11
2025 Certified Causal Defense with Generalizable Robustness
abstract
While machine learning models have proven effective across various scenarios, it is widely acknowledged that many models are vulnerable to adversarial attacks. Recently, numerous efforts have emerged in adversarial defense. Among them, certified defense is well known for its theoretical guarantees against arbitrary adversarial perturbations on input within a certain range. However, most existing works in this line struggle to generalize their certified robustness in other data domains with distribution shifts. This issue is rooted in the difficulty of eliminating the negative impact of spurious correlations on robustness in different domains. To address this problem, in this work, we propose a novel certified defense framework GLEAN, which incorporates a causal perspective into the generalization problem in certified defense. More specifically, our framework integrates a certifiable causal factor learning component to disentangle the causal relations and spurious correlations between input and label, thereby excluding the negative effect of spurious correlations on defense. On top of that, we design a causally certified defense strategy to handle adversarial attacks on latent causal factors. In this way, our framework is not only robust against malicious noises on data in the training distribution but also can generalize its robustness across domains with distribution shifts. Extensive experiments on benchmark datasets validate the superiority of our framework in certified robustness generalization in different data domains.
Yiran Qiao 0001, Yu Yin 0001, Chen Chen 0022, Jing Ma 0002
AAAI2
2025 Debate-to-Write: A Persona-Driven Multi-Agent Framework for Diverse Argument Generation
abstract
Writing arguments is a challenging task for both humans and machines. It entails incorporating high-level beliefs from various perspectives on the topic, along with deliberate reasoning and planning to construct a coherent narrative. Current language models often generate outputs autoregressively, lacking explicit integration of these underlying controls, resulting in limited output diversity and coherence. In this work, we propose a persona-based multi-agent framework for argument writing. Inspired by the human debate, we first assign each agent a persona representing its high-level beliefs from a unique perspective, and then design an agent interaction process so that the agents can collaboratively debate and discuss the idea to form an overall plan for argument writing. Such debate process enables fluid and nonlinear development of ideas. We evaluate our framework on argumentative essay writing. The results show that our framework generates more diverse and persuasive arguments by both automatic and human evaluations.
Hou Pong Chan, Jing Li 0049, Yu Yin 0001
COLING4
2025 BARD-GS: Blur-Aware Reconstruction of Dynamic Scenes via Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has shown remarkable potential for static scene reconstruction, and recent advancements have extended its application to dynamic scenes. However, the quality of reconstructions depends heavily on high-quality input images and precise camera poses, which are not that trivial to fulfill in real-world scenarios. Capturing dynamic scenes with handheld monocular cameras, for instance, typically involves simultaneous movement of both the camera and objects within a single exposure. This combined motion frequently results in image blur that existing methods cannot adequately handle. To address these challenges, we introduce BARD-GS, a novel approach for robust dynamic scene reconstruction that effectively handles blurry inputs and imprecise camera poses. Our method comprises two main components: 1) camera motion deblurring and 2) object motion deblurring. By explicitly decomposing motion blur into camera motion blur and object motion blur and modeling them separately, we achieve significantly improved rendering results in dynamic regions. In addition, we collect a real-world motion blur dataset of dynamic scenes to evaluate our approach. Extensive experiments demonstrate that BARD-GS effectively reconstructs high-quality dynamic scenes under realistic conditions, significantly outperforming existing methods.
Yiren Lu 0002, Yunlai Zhou, Disheng Liu, Tuo Liang, Yu Yin 0001
CVPR5
2025 Towards Open-set Face Anti-spoofing with Unseen Attack Synthesis
abstract
Existing face anti-spoofing (FAS) methods have primarily focused on close-set or cross-domain settings with a few pre-defined presentation attacks (PAs). However, with the continual emergence of diverse PAs, we argue that developing a generalizable FAS detector for unseen PAs deserves more attention from the FAS community. In this work, we investigate the open-set FAS setting, where the spoof types in testing are unseen. The key challenge is that the unseen PAs are designed to deceive the FAS model and thus appear similar to live faces both at the pixel level and in the feature space. To address this issue, we propose a novel framework for synthesizing unseen PAs and pushing the generated samples toward an open category space. Our approach is motivated by empirical findings that unseen PAs are more likely to be compactly clustered by spoof type and located at the boundary of the live distribution in the spoof-type-aware feature space derived from multi-class optimization. Lastly, we evaluate our method on the SiW-Mv2 cross-type benchmark using both fine-grained and coarse-grained protocols. Compared to the baselines and existing top competitors in close-set or cross-domain settings, our method outperforms them significantly on both protocols.
Chang Liu 0022, Yu Yin 0001, Yun Fu 0001
FG3
2025 Praxis-VLM: Vision-Grounded Decision Making via Text-Driven Reinforcement Learning
abstract
Vision Language Models exhibit impressive performance for various tasks, yet they often lack the sophisticated situational reasoning required for complex decision-making. This paper shows that VLMs can achieve surprisingly strong decision-making performance when visual scenes are replaced by textual descriptions, suggesting foundational reasoning can be effectively learned from language. Motivated by this insight, we propose Praxis-VLM, a reasoning VLM for vision-grounded decision-making. Praxis-VLM employs the GRPO algorithm on textual scenarios to instill robust reasoning capabilities, where models learn to evaluate actions and their consequences. These reasoning skills, acquired purely from text, successfully transfer to multimodal inference with visual inputs, significantly reducing reliance on scarce paired image-text training data. Experiments across diverse decision-making benchmarks demonstrate that Praxis-VLM substantially outperforms standard supervised fine-tuning, exhibiting superior performance and generalizability. Further analysis confirms that our models engage in explicit and effective reasoning, underpinning their enhanced performance and adaptability.
Jing Li 0049, Zhongzhu Pu, Hou Pong Chan, Yu Yin 0001
NeurIPS5
2025 Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
abstract
Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to multi-view inconsistencies and poor 3D object retrieval. Moreover, they are limited to static scenes and struggle with dynamic scenes due to the complexities of motion modeling. In this paper, we propose Segment then Splat, a 3D-aware open vocabulary segmentation approach for both static and dynamic scenes based on Gaussian Splatting. Segment then Splat reverses the long established approach of "segmentation after reconstruction'' by dividing Gaussians into distinct object sets before reconstruction. Once reconstruction is complete, the scene is naturally segmented into individual objects, achieving true 3D segmentation. This design eliminates both geometric and semantic ambiguities, as well as Gaussian–object misalignment issues in dynamic scenes. It also accelerates the optimization process, as it eliminates the need for learning a separate language field. After optimization, a CLIP embedding is assigned to each object to enable open-vocabulary querying. Extensive experiments on various datasets demonstrate the effectiveness of our proposed method in both static and dynamic scenarios.
Yiren Lu 0002, Yunlai Zhou, Yiran Qiao 0001, Chaoda Song, Tuo Liang, Jing Ma 0002, Huan Wang 0014, Yu Yin 0001
NeurIPS8
2024 VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
abstract
Large vision language models (VLMs) have demonstrated significant potential for integration into daily life, making it crucial for them to incorporate human values when making decisions in real-world situations.This paper introduces VIVA, a benchmark for VIsiongrounded decision-making driven by human VAlues.While most large VLMs focus on physical-level skills, our work is the first to examine their multimodal capabilities in leveraging human values to make decisions under a vision-depicted situation.VIVA contains 1,240 images depicting diverse real-world situations and the manually annotated decisions grounded in them.Given an image there, the model should select the most appropriate action to address the situation and provide the relevant human values and reason underlying the decision.Extensive experiments based on VIVA show the limitation of VLMs in using human values to make multimodal decisions.Further analyses indicate the potential benefits of exploiting action consequences and predicted human values.Our code and dataset are available at https://github.com/Derekkk/ VIVA_EMNLP24.
Yixiao Ren, Jing Li 0049, Yu Yin 0001
EMNLP4
2024 AMERICANO: Argument Generation with Discourse-driven Decomposition and Agent Interaction
abstract
Argument generation is a challenging task in natural language processing, which requires rigorous reasoning and proper content organization.Inspired by recent chain-of-thought prompting that breaks down a complex task into intermediate steps, we propose AMERI-CANO, a novel framework with agent interaction for argument generation.Our approach decomposes the generation process into sequential actions grounded on argumentation theory, which first executes actions sequentially to generate argumentative discourse components, and then produces a final argument conditioned on the components.To further mimic the human writing process and improve the left-to-right generation paradigm of current autoregressive language models, we introduce an argument refinement module that automatically evaluates and refines argument drafts based on feedback received.We evaluate our framework on the task of counterargument generation using a subset of Reddit/CMV dataset.The results show that our method outperforms both end-to-end and chain-of-thought prompting methods and can generate more coherent and persuasive arguments with diverse and rich contents.
Hou Pong Chan, Yu Yin 0001
INLG3
2024 View-consistent Object Removal in Radiance Fields
abstract
Radiance Fields (RFs) have emerged as a crucial technology for 3D scene representation, enabling the synthesis of novel views with remarkable realism. However, as RFs become more widely used, the need for effective editing techniques that maintain coherence across different perspectives becomes evident. Current methods primarily depend on per-frame 2D image inpainting, which often fails to maintain consistency across views, thus compromising the realism of edited RF scenes. In this work, we introduce a novel RF editing pipeline that significantly enhances consistency by requiring the inpainting of only a single reference image. This image is then projected across multiple views using a depth-based approach, effectively reducing the inconsistencies observed with per-frame inpainting. However, projections typically assume photometric consistency across views, which is often impractical in real-world settings. To accommodate realistic variations in lighting and viewpoint, our pipeline adjusts the appearance of the projected views by generating multiple directional variants of the inpainted image, thereby adapting to different photometric conditions. Additionally, we present an effective and robust multi-view object segmentation approach as a valuable byproduct of our pipeline. Extensive experiments demonstrate that our method significantly surpasses existing frameworks in maintaining content consistency across views and enhancing visual quality. More results are available at https://vulab-ai.github.io/View-consistent_Object_Removal_in_Radiance_Fields/.
Yiren Lu 0002, Jing Ma 0002, Yu Yin 0001
ACM Multimedia3
2024 Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions
abstract
Recent advancements in large vision language models have demonstrated remarkable proficiency across a wide range of tasks. Yet, these models still struggle with understanding the nuances of human humor through juxtaposition, particularly when it involves nonlinear narratives that underpin many jokes and humor cues. This paper investigates this challenge by focusing on comics with contradictory narratives, where each comic consists of two panels that create a humorous contradiction. We introduce the YesBut benchmark, which comprises tasks of varying difficulty aimed at assessing AI's capabilities in recognizing and interpreting these comics, ranging from literal content comprehension to deep narrative reasoning. Through extensive experimentation and analysis of recent commercial or open-sourced large vision language models, we assess their capability to comprehend the complex interplay of the narrative humor inherent in these comics. Our results show that even the state-of-the-art models still struggle with this task. Our findings offer insights into the current limitations and potential improvements for AI in understanding human creative expressions.
Tuo Liang, Jing Li 0049, Yiren Lu 0002, Yunlai Zhou, Yiran Qiao 0001, Jing Ma 0002, Yu Yin 0001
NeurIPS8
2023 NeRFInvertor: High Fidelity NeRF-GAN Inversion for Single-Shot Real Image Animation
abstract
Nerf-based Generative models have shown impressive capacity in generating high-quality images with consistent 3D geometry. Despite successful synthesis of fake identity images randomly sampled from latent space, adopting these models for generating face images of real subjects is still a challenging task due to its so-called inversion issue. In this paper, we propose a universal method to surgically finetune these NeRF-GAN models in order to achieve high-fidelity animation of real subjects only by a single image. Given the optimized latent code for an out-of-domain real image, we employ 2D loss functions on the rendered image to reduce the identity gap. Furthermore, our method leverages explicit and implicit 3D regularizations using the in-domain neighborhood samples around the optimized latent code to remove geometrical and visual artifacts. Our experiments confirm the effectiveness of our method in realistic, high-fidelity, and 3D consistent animation of real faces on multiple NeRF-GAN models across different datasets.
Yu Yin 0001, Kamran Ghasedi, HsiangTao Wu, Jiaolong Yang, Xin Tong 0001, Yun Fu 0001
CVPR1
2023 Concentric Ring Loss for Face Forgery Detection
abstract
The issue of detecting face forgeries has garnered significant interest in the field of computer vision, primarily driven by the growing social concerns of indistinguishable deepfake images. One of the primary obstacles encountered in the field of deepfake detection is enhancing the discriminative power of learned features. In this paper, we propose a Concentric Ring Loss (CRL) that aims to promote the learning of compressed intra-class features and separated inter-class features inside a model. Specifically, we apply margin penalties in both Euclidean and angular space separately, which serve to increase the separation between real and fake images. Moreover, we introduce a frequency-aware triplet network with a self-developed sample generation strategy, which provides efficient hard triplets for model training. Extensive experiments demonstrate the superiority of our methods over multiple datasets. We show that CRL consistently outperforms the state-of-the-art by a large margin.
Yu Yin 0001, Yizhou Wang 0006, Yun Fu 0001
ICDM1
2023 Human Motion Segmentation via Velocity-Sensitive Dual-Side Auto-Encoder
abstract
Human motion segmentation (HMS) aims to segment a long human action video into a bunch of short and meaningful action clips. Existing supervised learning approaches need a large amount of training data which may be costly in real-world scenario, while most unsupervised clustering methods cannot fully explore the temporal correlations among human motions and hard to achieve promising performances. In our paper, we design a novel unsupervised framework, called Velocity-Sensitive Dual-Side Auto-Encoder (VSDA), for HMS task. Specifically, a multi-neighbor auto-encoder (MNA) is proposed to extract informative temporal features, which fully explores the local temporal patterns of human motions. In addition, a long-short distance encoding (LSE) strategy is designed. It constrains the encoded representations of close (short-distance) frames becoming similar while the representations of far-away (long-distance) frames becoming distinctive. Similarly, this strategy is also deployed on the decoded outputs as the long-short distance decoding (LSD) module. The LSE/LSD guides the learning process explicitly and implicitly to achieve the dual-side structure. Moreover, we consider the energy variations during the human motion to propose the velocity-sensitive (VS) guidance mechanism for further model improvement. VSDA leverages the temporal characteristics of human motion and derives promising HMS performance. Comprehensive experiments on six real-world human motion datasets illustrate the effectiveness of our proposed model.
Lichen Wang, Yunyu Liu, Yu Yin 0001, Hang Di, Yun Fu 0001
IEEE Trans. Image Process.4
2022 Generating Topological Structure of Floorplans from Room Attributes
abstract
Analysis of indoor spaces requires topological information. In this paper, we propose to extract topological information from room attributes using what we call Iterative and adaptive graph Topology Learning (ITL). ITL progressively predicts multiple relations between rooms; at each iteration, it improves node embeddings, which in turn facilitates the generation of a better topological graph structure. This notion of iterative improvement of node embeddings and topological graph structure is in the same spirit as [5]. However, while [5] computes the adjacency matrix based on node similarity, we learn the graph metric using a relational decoder to extract room correlations. Experiments using a new challenging indoor dataset validate our proposed method. Qualitative and quantitative evaluation for layout topology prediction and floorplan generation applications also demonstrate the effectiveness of ITL.
Yu Yin 0001, Will Hutchcroft, Naji Khosravan, Ivaylo Boyadzhiev, Yun Fu 0001, Sing Bing Kang
ICMR1
2022 Multimodal In-bed Pose and Shape Estimation under the Blankets
abstract
Advancing technology to monitor our bodies and behavior while sleeping and resting are essential for healthcare. However, keen challenges arise from our tendency to rest under blankets. We present a multimodal approach to uncover the subjects and view bodies at rest without the blankets obscuring the view. For this, we introduce a channel-based fusion scheme to effectively fuse different modalities in a way that best leverages the knowledge captured by the multimodal sensors, including visual- and non-visual-based. The channel-based fusion scheme enhances the model's flexibility in the input at inference: one-to-many input modalities required at test time. Nonetheless, multimodal data or not, detecting humans at rest in bed is still a challenge due to the extreme occlusion when covered by a blanket. To mitigate the negative effects of blanket occlusion, we use an attention-based reconstruction module to explicitly reduce the uncertainty of occluded parts by generating uncovered modalities, which further update the current estimation via a cyclic fashion. Extensive experiments validate the proposed model's superiority over others.
Yu Yin 0001, Joseph P. Robinson, Yun Fu 0001
ACM Multimedia1
2022 Collaborative Attention Mechanism for Multi-Modal Time Series Classification
abstract
Multi-modal time series classification (MTC) uses complementary information from different modalities to improve the learning performance. Obtaining informative modality-specific representation plays an essential role in MTC. Attention mechanism has been widely adopted as an effective strategy for discovering discriminative cues underlying temporal data. However, most existing MTC methods only utilize attention to balance the feature weights within or cross modalities but ignore digging latent patterns from mutual-support information in attention space. Specifically, the attention distributions are different for multiple modalities which are supportive and instructional with each other. To this end, we propose a collaborative attention mechanism (CAM) for MTC based on a novel perspective to utilize attention module. CAM detects the attention differences among multi-modal time series, and adaptively integrates different attention information to benefit each other. We extend the long short-term memory (LSTM) to a Mutual-Aid RNN (MAR) for multi-modal collaboration. CAM takes advantages of modality-specific attention to guide another modality and discover potential information which is hard to be explored by itself. It paves a novel way of employing attention to enhance the capacity of multi-modal representations. Extensive experiments on four multi-modal time series datasets illustrate the CAM effectiveness to improve the single-modal and also boost multi-modal performances.
Zhiqiang Tao, Lichen Wang, Sheng Li 0001, Yu Yin 0001, Yun Fu 0001
SDM5
2022 Semi-Supervised Domain Adaptive Structure Learning
abstract
Semi-supervised domain adaptation (SSDA) is quite a challenging problem requiring methods to overcome both 1) overfitting towards poorly annotated data and 2) distribution shift across domains. Unfortunately, a simple combination of domain adaptation (DA) and semi-supervised learning (SSL) methods often fail to address such two objects because of training data bias towards labeled samples. In this paper, we introduce an adaptive structure learning method to regularize the cooperation of SSL and DA. Inspired by the multi-views learning, our proposed framework is composed of a shared feature encoder network and two classifier networks, trained for contradictory purposes. Among them, one of the classifiers is applied to group target features to improve intra-class density, enlarging the gap of categorical clusters for robust representation learning. Meanwhile, the other classifier, serviced as a regularizer, attempts to scatter the source features to enhance the smoothness of the decision boundary. The iterations of target clustering and source expansion make the target features being well-enclosed inside the dilated boundary of the corresponding source points. For the joint address of cross-domain features alignment and partially labeled data learning, we apply the maximum mean discrepancy (MMD) distance minimization and self-training (ST) to project the contradictory structures into a shared view to make the reliable final decision. The experimental results over the standard SSDA benchmarks, including DomainNet and Office-home, demonstrate both the accuracy and robustness of our method over the state-of-the-art approaches.
Can Qin, Lichen Wang, Qianqian Ma, Yu Yin 0001, Huan Wang 0014, Yun Fu 0001
IEEE Trans. Image Process.4
2022 Families in Wild Multimedia: A Multimodal Database for Recognizing Kinship
abstract
Kinship, a soft biometric detectable in media, is fundamental for a myriad of use-cases. Despite the difficulty of detecting kinship, annual data challenges using still-images have consistently improved performances and attracted new researchers. Now, systems reach performance levels unforeseeable a decade ago, closing in on performances acceptable to deploy in practice. Like other biometric tasks, we expect systems can receive help from other modalities. We hypothesize that adding modalities toFamilies In the Wild(FIW), which has only still-images, will improve performance. Thus, to narrow the gap between research and reality and enhance the power of kinship recognition systems, we extend FIW with multimedia (MM) data (i.e., video, audio, and text captions). Specifically, we introduce the first publicly available multi-task MM kinship dataset. To buildFIW in Multimedia(FIW MM), we developed machinery to automatically collect, annotate, and prepare the data, requiring minimal human input and no financial cost. The proposed MM corpus allows the problem statements to be more realistic template-based protocols. We show significant improvements in all benchmarks with the added modalities. The results highlight edge cases to inspire future research with different areas of improvement. FIW MM supplies the data needed to increase the potential of automated systems to detect kinship in MM. It also allows experts from diverse fields to collaborate in novel ways.
Joseph P. Robinson, Zaid Khan 0001, Yu Yin 0001, Ming Shao, Yun Fu 0001
IEEE Trans. Multim.3
2021 Context-Aware Interaction Network for Question Matching
abstract
Impressive milestones have been achieved in text matching by adopting a cross-attention mechanism to capture pertinent semantic connections between two sentence representations.However, regular cross-attention focuses on word-level links between the two input sequences, neglecting the importance of contextual information.We propose a context-aware interaction network (COIN) to properly align two sequences and infer their semantic relationship.Specifically, each interaction block includes (1) a context-aware cross-attention mechanism to effectively integrate contextual information when aligning two sequences, and (2) a gate fusion layer to flexibly interpolate aligned representations.We apply multiple stacked interaction blocks to produce alignments at different levels and gradually refine the attention results.Experiments on two question matching datasets and detailed analyses demonstrate the effectiveness of our model.
Zuohui Fu, Yu Yin 0001, Gerard de Melo
EMNLP (1)3
2021 SuperFront: From Low-resolution to High-resolution Frontal Face Synthesis
abstract
Even the most impressive achievement in frontal face synthesis is challenged by large poses and low-quality data given one single side-view face. We propose a synthesizer called SuperFront GAN (SF-GAN) to accept one or more low-resolution (LR) faces at the input to then output a high-resolution (HR) frontal face with various poses and such to preserve identity information. SF-GAN includes intra-class and inter-class constraints, which allow it to learn an identity-preserving representation from multiple LR faces in an improved, comprehensive manner. We adopt an orthogonal loss as the intra-class constraint that diversifies the learned feature-space per subject. Hence, each sample is made to complement the others to its max ability. Additionally, a triplet loss is used as the inter-class constraint: it improves the discriminative power of the new representation, which, hence, maintains the identity information. Furthermore, we integrate a super-resolution (SR) side-view module as part of the SF-GAN to help preserve the finer details of HR side-views. This helps the model reconstruct the high-frequency parts of the face (i.e. periocular region, nose, and mouth regions). Quantitative and qualitative results demonstrate the superiority of SF-GAN. SF-GAN holds promise as a pre-processing step to normalize and align faces before passing to CV system for processing.
Yu Yin 0001, Joseph P. Robinson, Songyao Jiang, Can Qin, Yun Fu 0001
ACM Multimedia1
2021 Contradictory Structure Learning for Semi-supervised Domain Adaptation
abstract
Current adversarial adaptation methods attempt to align the cross-domain features, whereas two challenges remain unsolved: 1) the conditional distribution mismatch and 2) the bias of the decision boundary towards the source domain.To solve these challenges, we propose a novel framework for semi-supervised domain adaptation by unifying the learning of opposite structures (UODA).UODA consists of a generator and two classifiers (i.e., the sourcescattering classifier and the target-clustering classifier), which are trained for contradictory purposes.The target-clustering classifier attempts to cluster the target features to improve intra-class density and enlarge inter-class divergence.Meanwhile, the source-scattering classifier is designed to scatter the source features to enhance the decision boundary's smoothness.Through the alternation of source-feature expansion and target-feature clustering procedures, the target features are well-enclosed within the dilated boundary of the corresponding source features.This strategy can make the cross-domain features to be precisely aligned against the source bias simultaneously.Moreover, to overcome the model collapse through training, we progressively update the measurement of feature's distance and their representation via an adversarial training paradigm.Extensive experiments on the benchmarks of DomainNet and Office-home datasets demonstrate the superiority of our approach over the state-of-the-art methods.
Can Qin, Lichen Wang, Qianqian Ma, Yu Yin 0001, Huan Wang 0014, Yun Fu 0001
SDM4
2020 Joint Super-Resolution and Alignment of Tiny Faces
abstract
Super-resolution (SR) and landmark localization of tiny faces are highly correlated tasks. On the one hand, landmark localization could obtain higher accuracy with faces of high-resolution (HR). On the other hand, face SR would benefit from prior knowledge of facial attributes such as landmarks. Thus, we propose a joint alignment and SR network to simultaneously detect facial landmarks and super-resolve tiny faces. More specifically, a shared deep encoder is applied to extract features for both tasks by leveraging complementary information. To exploit representative power of the hierarchical encoder, intermediate layers of a shared feature extraction module are fused to form efficient feature representations. The fused features are then fed to task-specific modules to detect landmarks and super-resolve face images in parallel. Extensive experiments demonstrate that the proposed model significantly outperforms the state-of-the-art in both landmark localization and SR of faces. We show a large improvement for landmark localization of tiny faces (i.e., 16 × 16). Furthermore, the proposed framework yields comparable results for landmark localization on low-resolution (LR) faces (i.e., 64 × 64) to existing methods on HR (i.e., 256 × 256). As for SR, the proposed method recovers sharper edges and more details from LR face images than other state-of-the-art methods, which we demonstrate qualitatively and quantitatively.
Yu Yin 0001, Joseph P. Robinson, Yulun Zhang 0001, Yun Fu 0001
AAAI1
2020 Recognizing Families In the Wild (RFIW): The 4th Edition
abstract
Recognizing Families In the Wild (RFIW)- an annual large-scale, multi-track automatic kinship recognition evaluation- supports various visual kin-based problems on scales much higher than ever before. Organized in conjunction with the as a Challenge, RFIW provides a platform for publishing original work and the gathering of experts for a discussion of the next steps. This paper summarizes the supported tasks (i.e., kinship verification, tri-subject verification, and search & retrieval of missing children) in the evaluation protocols, which include the practical motivation, technical background, data splits, metrics, and benchmark results. Furthermore, top submissions (i.e., leader-board stats) are listed and reviewed as a high-level analysis on the state of the problem. In the end, the purpose of this paper is to describe the 2020 RFIW challenge, end-to-end, along with forecasts in promising future directions.
Joseph P. Robinson, Yu Yin 0001, Zaid Khan 0001, Ming Shao, Si-Yu Xia, Michael Stopa, Samson Timoner, Matthew Turk 0001, Rama Chellappa, Yun Fu 0001
FG2
2020 Dual-Attention GAN for Large-Pose Face Frontalization
abstract
Face frontalization provides an effective and efficient way for face data augmentation and further improves the face recognition performance in extreme pose scenario. Despite recent advances in deep learning-based face synthesis approaches, this problem is still challenging due to significant pose and illumination discrepancy. In this paper, we present a novel Dual-Attention Generative Adversarial Network (DA-GAN) for photo-realistic face frontalization by capturing both contextual dependencies and local consistency during GAN training. Specifically, a self-attention-based generator is introduced to integrate local features with their long-range dependencies yielding better feature representations, and hence generate faces that preserves identities better, especially for larger pose angles. Moreover, a novel face-attention-based discriminator is applied to emphasize local features of face regions, and hence reinforce the realism of synthetic frontal faces. Guided by semantic segmentation, four independent discriminators are used to distinguish between different aspects of a face (i.e., skin, keypoints, hairline, and frontalized face). By introducing these two complementary attention mechanisms in generator and discriminator separately, we can learn a richer feature representation and generate identity preserving inference of frontal views with much finer details (i.e., more accurate facial appearance and textures) comparing to the state-of-the-art. Quantitative and qualitative experimental results demonstrate the effectiveness and efficiency of our DA-GAN approach.
Yu Yin 0001, Songyao Jiang, Joseph P. Robinson, Yun Fu 0001
FG1
2020 Dual-Side Auto-Encoder for High-Dimensional Time Series Segmentation
abstract
High-dimensional time series segmentation aims to segment a long temporal sequence into several short and meaningful subsequences. The high-dimensionality makes it challenging due to the complicated correlations among the sequential features. A large number of labeled data is required in existing supervised methods, and unsupervised methods mainly deploy clustering approaches, which are sensitive to outliers and hard to guarantee high performance. Also, most existing methods mainly rely on hand-craft features to deal with regular time series segmentation and achieve promising results. However, these approaches cannot effectively handle high-dimensional time series and will result in a high computational cost. In our work, we propose a novel unsupervised representation learning framework called Dual-Side Auto-Encoder (DSAE). It mainly focuses on high-dimensional time series segmentation by effectively capturing the temporal correlative patterns. Specifically, a single-to-multiple auto-encoder is designed to capture local sequential information. Besides, a long-shot distance encoding strategy is proposed. It aims to explicitly guide the learning process to obtain distinctive representations for segmentation. Furthermore, the long-short distance strategy is also executed in the decoded feature space, which implicitly directs the representation learning. Substantial experiments on six datasets illustrate the model effectiveness.
Lichen Wang, Yunyu Liu, Yu Yin 0001, Yun Fu 0001
ICDM4
2020 Analysis of multimodal physiological signals within and between individuals to predict psychological challenge vs. threat
Aya Khalaf, Mohsen Nabian, Miaolin Fan, Yu Yin 0001, Jolie B. Wormwood, Erika Siegel, Karen S. Quigley, Lisa Feldman Barrett, Murat Akçakaya, Chun-An Chou, Sarah Ostadabbas
Expert Syst. Appl.4