Yawei Luo

dblp:160/7852 · DBLP profile ↗
← Back
55ranked-venue papers
10as first author
42since 2021 · last 2026
0000-0002-7037-1806ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 4 first-author · 31 since 2021Artificial intelligence and machine learning · 32 · 9 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Beyond Pedagogical Principles: Multi-Horizon Preference Optimization for Efficient Socratic Tutoring
abstract
The development of LLM-based tutor agents faces challenges in simultaneously ensuring adherence to pedagogical principles and achieving optimal pedagogical effectiveness, particularly in dynamic, multi-turn interactions.Existing methods are often constrained by static data or sparse reward signals in online settings.To address this gap, we propose Multi-Horizon Preference Optimization (MHPO), a novel framework that iteratively refines tutor agents using a multi-horizon reward function within a dynamic teacher-student simulation environment.Specifically, this reward function is designed to capture both turn-level pedagogical quality and trajectory-level pedagogical effectiveness, which is estimated via Monte Carlo rollouts.We further investigate two distinct strategies to aggregate these rewards for policy optimization.Our experiments demonstrate that MHPO significantly enhances base model performance, achieving a superior balance between principles and effectiveness compared to various baselines.
Xueqiao Zhang, Yawei Luo
ACL (1)5
2026 22DEditor: Zero-Shot Text-Driven Volumetric Video Editing With Spatiotemporal Consistency via Single 2D Diffusion
abstract
Volumetric video represents 4D (i.e., dynamic 3D) scenes captured from multiple viewpoints, offering rich spatial and temporal information for immersive applications. Despite advances in diffusion-based editing for images, videos, and 3D objects, achieving precise, spatiotemporally coherent edits in volumetric videos remains challenging due to complex geometry and dynamic motion. In this work, we propose 22DEditor, a novel zero-shot volumetric video editing framework that enforces temporal-spatial coherence using a single 2D diffusion model. A key component, Spatiotemporal Consistent Null-Text Optimization (SCNO), jointly models continuity across viewpoints and timestamps, mitigating inconsistencies in multi-frame editing. To further enhance fidelity, the Source-Attention-Guided Editing (SAGE) module leverages self-attention maps to preserve spatial structures, reuses cross-attention maps for semantic consistency, and aggregates multi-layer source attention to prevent unintended edits when target objects exit the scene. Building upon these components, we develop a round-by-round editing pipeline, enabling diverse local and global modifications—including semantic, stylistic, and attribute-based adjustments—while maintaining coherent dynamic scene structure. Extensive experiments on multi-view volumetric video datasets demonstrate that our approach significantly improves editing precision, semantic fidelity, and spatiotemporal consistency, outperforming state-of-the-art zero-shot methods and establishing a new benchmark for high-fidelity volumetric video editing.
Qiaowei Miao, Kehan Li 0008, Yawei Luo, Yi Yang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Counterfactual Co-Occurring Learning for Bias Mitigation in Weakly-Supervised Object Localization
abstract
Contemporary weakly-supervised object localization (WSOL) methods have primarily focused on addressing the challenge of localizing the most discriminative region while largely overlooking the relatively less explored issue of biased activation—incorrectly spotlighting co-occurring background with the foreground feature. In this paper, we conduct a thorough causal analysis to investigate the origins of biased activation. Based on our analysis, we attribute this phenomenon to the presence of co-occurring background confounders. Building upon this profound insight, we introduce a pioneering paradigm known as Counterfactual Co-occurring Learning (CCL), meticulously engendering counterfactual representations by adeptly disentangling the foreground from the co-occurring background elements. Furthermore, we propose an innovative network architecture known as Counterfactual-CAM. This architecture seamlessly incorporates a perturbation mechanism for counterfactual representations into the vanilla CAM-based model. By training the WSOL model with these perturbed representations, we guide the model to prioritize the consistent foreground content while concurrently reducing the influence of distracting co-occurring backgrounds. To the best of our knowledge, this study represents the initial exploration of this research direction. Our extensive experiments conducted across multiple benchmarks validate the effectiveness of the proposed Counterfactual-CAM in mitigating biased activation.
Feifei Shao, Yawei Luo, Lei Chen 0082, Ping Liu 0004, Wei Yang 0034, Yi Yang 0001, Jun Xiao 0001
IEEE Trans. Multim.2
2025 TSD-SR: One-Step Diffusion with Target Score Distillation for Real-World Image Super-Resolution
abstract
Pre-trained text-to-image diffusion models are increasingly applied to real-world image super-resolution (Real-ISR) task. Given the iterative refinement nature of diffusion models, most existing approaches are computationally expensive. While methods such as SinSR and OSEDiff have emerged to condense inference steps via distillation, their performance in image restoration or details recovery is not satisfied. To address this, we propose TSD-SR, a novel distillation framework specifically designed for real-world image super-resolution, aiming to construct an efficient and effective one-step model. We first introduce the Target Score Distillation, which leverages the priors of diffusion models and real image references to achieve more realistic image restoration. Secondly, we propose a Distribution-Aware Sampling Module to make detail-oriented gradients more readily accessible, addressing the challenge of recovering fine details. Extensive experiments demonstrate that our TSD-SR has superior restoration results (most of the metrics perform the best) and the fastest inference speed (e.g. 40 times faster than SeeSR) compared to the past Real-ISR approaches based on pre-trained diffusion priors.
Linwei Dong, Qingnan Fan, Yihong Guo, Yawei Luo, Changqing Zou
CVPR7
2025 SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
abstract
Video-based Large Language Models (Video-LLMs) have witnessed substantial advancements in recent years, propelled by the advancement in multi-modal LLMs. Although these models have demonstrated proficiency in providing the overall description of videos, they struggle with fine-grained understanding, particularly in aspects such as visual dynamics and video details inquiries. To tackle these shortcomings, we find that fine-tuning Video-LLMs on self-supervised fragment tasks, greatly improve their fine-grained video understanding abilities. Hence we propose two key contributions: (1) Self-Supervised Fragment Fine-Tuning (SF2T), a novel effortless fine-tuning method, employs the rich inherent characteristics of videos for training, while unlocking more fine-grained understanding ability of Video-LLMs. Moreover, it relieves researchers from labor-intensive annotations and smartly circumvents the limitations of natural language, which often fails to capture the complex spatiotemporal variations in videos; (2) A novel benchmark dataset, namely FineVidBench, for rigorously assessing Video-LLMs’ performance at both the scene and fragment levels, offering a comprehensive evaluation of their capabilities. We assessed multiple models and validated the effectiveness of SF2T on them. Experimental results reveal that our approach improves their ability to capture and interpret spatiotemporal details.
Yangliu Hu, Zikai Song, Na Feng, Yawei Luo, Junqing Yu, Yi-Ping Phoebe Chen, Wei Yang 0034
CVPR4
2025 MICAS: Multi-grained In-Context Adaptive Sampling for 3D Point Cloud Processing
abstract
Point cloud processing (PCP) encompasses tasks like reconstruction, denoising, registration, and segmentation, each often requiring specialized models to address unique task characteristics. While in-context learning (ICL) has shown promise across tasks by using a single model with task-specific demonstration prompts, its application to PCP reveals significant limitations. We identify inter-task and intra-task sensitivity issues in current ICL methods for PCP, which we attribute to inflexible sampling strategies lacking context adaptation at the point and prompt levels. To address these challenges, we propose MICAS, an advanced ICL framework featuring a multi-grained adaptive sampling mechanism tailored for PCP. MICAS introduces two core components: task-adaptive point sampling, which leverages inter-task cues for point-level sampling, and query-specific prompt sampling, which selects optimal prompts per query to mitigate intra-task sensitivity. To our knowledge, this is the first approach to introduce adaptive sampling tailored to the unique requirements of point clouds within an ICL framework. Extensive experiments show that MICAS not only efficiently handles various PCP tasks but also significantly outperforms existing methods. Notably, it achieves a remarkable 4.1% improvement in the part segmentation task and delivers consistent gains across various PCP applications.
Feifei Shao, Ping Liu 0004, Yawei Luo, Jun Xiao 0001
CVPR4
2025 Ref-GS: Directional Factorization for 2D Gaussian Splatting
abstract
In this paper, we introduce Ref-GS, a novel approach for directional light factorization in 2D Gaussian splatting [8], which enables photorealistic view-dependent appearance rendering and precise geometry recovery. Ref-GS builds upon the deferred rendering of Gaussian splatting and applies directional encoding to the deferred-rendered surface, effectively reducing the ambiguity between orientation and viewing angle. Next, we introduce a spherical Mip-grid to capture varying levels of surface roughness, enabling roughness-aware Gaussian shading. Additionally, we propose a simple yet efficient geometry-lighting factorization that connects geometry and lighting via the vector outer product, significantly reducing renderer overhead when integrating volumetric attributes. Our method achieves superior photorealistic rendering for a range of open-world scenes while also accurately recovering geometry. See our interactive project page.
Youjia Zhang, Anpei Chen, Yumin Wan, Zikai Song, Junqing Yu, Yawei Luo, Wei Yang 0034
CVPR6
2025 Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
abstract
Role-playing agents (RPAs) have attracted growing interest for their ability to simulate immersive and interactive characters. However, existing approaches primarily focus on static role profiles, overlooking the dynamic perceptual abilities inherent to humans. To bridge this gap, we introduce the concept of dynamic role profiles by incorporating video modality into RPAs. To support this, we construct Role-playing-Video60k, a large-scale, high-quality dataset comprising 60k videos and 700k corresponding dialogues. Based on this dataset, we develop a comprehensive RPA framework that combines adaptive temporal sampling with both dynamic and static role profile representations. Specifically, the dynamic profile is created by adaptively sampling video frames and feeding them to the LLM in temporal order, while the static profile consists of (1) character dialogues from training videos during fine-tuning, and (2) a summary context from the input video during inference. This joint integration enables RPAs to generate greater responses. Furthermore, we propose a robust evaluation method covering eight metrics. Experimental results demonstrate the effectiveness of our framework, highlighting the importance of dynamic role profiles in developing RPAs.
Xueqiao Zhang, Jingtao Xu, Yi Yang 0001, Yawei Luo
EMNLP7
2025 DICS: Find Domain-Invariant and Class-Specific Features for Out-of-Distribution Generalization
abstract
While deep neural networks have made remarkable progress in various tasks, their performance typically deteriorates and faces insecurity when tested in out-of-distribution (OOD) scenarios. Many OOD methods focus on extracting domain-invariant features but neglect whether these features are unique to each class. Even if some features are domain-invariant, they cannot serve as key criteria if shared across different classes. In OOD tasks, both domain-related and class-shared features act as confounders that hinder generalization. In this paper, we propose a DICS model to extract Domain-Invariant and Class-Specific features, including Domain Invariance Testing (DIT) and Class Specificity Testing (CST), which mitigate the effects of spurious correlations introduced by confounders. DIT learns domain-related features of each source domain and removes them from inputs to isolate domain-invariant class-related features. DIT ensures domain invariance by aligning same-class features across different domains. Then, CST calculates soft labels for those features by comparing them with features learned in previous steps. We optimize the cross-entropy between the soft labels and their true labels, which enhances same-class similarity and different-class distinctiveness, thereby reinforcing class specificity. Extensive experiments on widely-used benchmarks demonstrate the effectiveness of our proposed algorithm. Additional visualizations further demonstrate that DICS effectively identifies the key features of each class in target domains.
Qiaowei Miao, Yawei Luo, Yi Yang 0001
ICASSP2
2025 MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-Adsorbed Gaussian Splatting
abstract
3D reconstruction and simulation, although interrelated, have distinct objectives: reconstruction requires a flexible 3D representation that can adapt to diverse scenes, while simulation needs a structured representation to model motion principles effectively. This paper introduces the Mesh-adsorbed Gaussian Splatting (MaGS) method to address this challenge. MaGS constrains 3D Gaussians to roam near the mesh, creating a mutually adsorbed mesh-Gaussian 3D representation. Such representation harnesses both the rendering flexibility of 3D Gaussians and the structured property of meshes. To achieve this, we introduce RMD-Net, a network that learns motion priors from video data to refine mesh deformations, alongside RGD-Net, which models the relative displacement between the mesh and Gaussians to enhance rendering fidelity under mesh constraints. To generalize to novel, user-defined deformations beyond input video without reliance on temporal data, we propose MPE-Net, which leverages inherent mesh information to bootstrap RMD-Net and RGD-Net. Due to the universality of meshes, MaGS is compatible with various deformation priors such as ARAP, SMPL, and soft physics simulation. Extensive experiments on the D-NeRF, DG-Mesh, and PeopleSnapshot datasets demonstrate that MaGS achieves state-of-the-art performance in both reconstruction and simulation.
Shaojie Ma, Yawei Luo, Wei Yang 0011, Yi Yang 0001
ICCV2
2025 Grounding Creativity in Physics: A Brief Survey of Physical Priors in AIGC
abstract
Recent advancements in AI-generated content have significantly improved the realism of 3D and 4D generation. However, most existing methods prioritize appearance consistency while neglecting underlying physical principles, leading to artifacts such as unrealistic deformations, unstable dynamics, and implausible objects interactions. Incorporating physics priors into generative models has become a crucial research direction to enhance structural integrity and motion realism. This survey provides a review of physics-aware generative methods, systematically analyzing how physical constraints are integrated into 3D and 4D generation. First, we examine recent works in incorporating physical priors into static and dynamic 3D generation, categorizing methods based on representation types, including vision-based, NeRF-based, and Gaussian Splatting-based approaches. Second, we explore emerging techniques in 4D generation, focusing on methods that model temporal dynamics with physical simulations. Finally, we conduct a comparative analysis of major methods, highlighting their strengths, limitations, and suitability for different materials and motion dynamics. By presenting an in-depth analysis of physics-grounded AIGC, this survey aims to bridge the gap between generative models and physical realism, providing insights that inspire future research in physically consistent content generation.
Siwei Meng, Yawei Luo, Ping Liu 0004
IJCAI2
2025 Optimized View and Geometry Distillation from Multi-view Diffuser
abstract
Generating multi-view images from a single input view using image-conditioned diffusion models is a recent advancement and has shown considerable potential. However, issues such as the lack of consistency in synthesized views and over-smoothing in extracted geometry persist. Previous methods integrate multi-view consistency modules or impose additional supervisory to enhance view consistency while compromising on the flexibility of camera positioning and limiting the versatility of view synthesis. In this study, we consider the radiance field optimized during geometry extraction as a more rigid consistency prior, compared to volume and ray aggregation used in previous works. We further identify and rectify a critical bias in the traditional radiance field optimization process through score distillation from a multi-view diffuser. We introduce an Unbiased Score Distillation (USD) that utilizes unconditioned noises from a 2D diffusion model, greatly refining the radiance field fidelity. We leverage the rendered views from the optimized radiance field as the basis and develop a two-step specialization process of a 2D diffusion model, which is adept at conducting object-specific denoising and generating high-quality multi-view images. Finally, we recover faithful geometry and texture directly from the refined multi-view images. Empirical evaluations demonstrate that our optimized geometry and view distillation technique generates comparable results to the state-of-the-art models trained on extensive datasets, all while maintaining freedom in camera positioning. Source code of our work is publicly available at: https://youjiazhang.github.io/USD/.
Youjia Zhang, Zikai Song, Junqing Yu, Yawei Luo, Wei Yang 0034
IJCAI4
2025 SonicGauss: Position-Aware Physical Sound Synthesis for 3D Gaussian Representations
abstract
While 3D Gaussian representations (3DGS) have proven effective for modeling the geometry and appearance of objects, their potential for capturing other physical attributes-such as sound-remains largely unexplored. In this paper, we present a novel framework dubbed SonicGauss for synthesizing impact sounds from 3DGS representations by leveraging their inherent geometric and material properties. Specifically, we integrate a diffusion-based sound synthesis model with a PointTransformer-based feature extractor to infer material characteristics and spatial-acoustic correlations directly from Gaussian ellipsoids. Our approach supports spatially varying sound responses conditioned on impact locations and generalizes across a wide range of object categories. Experiments on the ObjectFolder dataset and real-world recordings demonstrate that our method produces realistic, position-aware auditory feedback. The results highlight the framework's robustness and generalization ability, offering a promising step toward bridging 3D visual representations and interactive sound synthesis.
Chunshi Wang, Yawei Luo
ACM Multimedia3
2025 PanoExplorer: From Single Panorama to Immersive Walkthrough via Structure-Aware Completion and Refinement
Yawei Luo, Yi Yang 0001
PRCV (10)2
2025 Kill Two Birds with One Stone: Domain Generalization for Semantic Segmentation via Network Pruning
Yawei Luo, Ping Liu 0004, Yi Yang 0001
Int. J. Comput. Vis.1
2025 ExpAvatar: High-Fidelity Avatar Generation of Unseen Expressions with 3D Face Priors
abstract
The reconstruction of dynamic head avatars has gained increasing significance, giving rise to various downstream applications such as visual dubbing and digital human creation. Despite recent advancements, generating novel, unseen expressions for a given identity remains challenging in concurrently achieving (1) accurate expression and consistent appearance and (2) high-quality and realistic faces. This article introduces ExpAvatar, a novel approach crafted to address these challenges. ExpAvatar elaborately leverages the appearance consistency capabilities inherent in 3DMMs-based models along with the robust generalization ability of DDPMs-based models to alleviate appearance drift issues and enhance the generation of unseen expressions. Specifically, ExpAvatar introduces a Face Priors-Conditioned Diffusion (FPDiff) model to inject 3D face priors into generation models through fine-tuning. Furthermore, a Face Priors-Conditioned Catalyst (FPCatalyst) is employed to enhance the inference efficiency and generation quality. Moreover, we propose a unique confidence-based regularizer function to mitigate the effect of imperfect face-tracking estimates, thereby improving the quality of dynamic neural head avatars. Experimental results demonstrate that ExpAvatar surpasses current state-of-the-art solutions in generating unseen expressions, marking an advancement in the realm of dynamic head avatar synthesis. Code: https://github.com/yuangan/ExpAvatar .
Yuan Gan, Ruijie Quan, Yawei Luo
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Balancing Humans and Machines: A Study on Integration Scale and Its Impact on Collaborative Performance
abstract
In the evolving artificial intelligence domain, hybrid human-machine systems have emerged as a transformative research area. While many studies have concentrated on individual human-machine interactions, there is a lack of focus on multi-human and multi-machine dynamics. This paper delves into these nuances by introducing a novel statistical framework that discerns integration accuracy in terms of precision and diversity. Empirical studies reveal that performance surges consistently with scale, either in human or machine settings. However, hybrid systems present complexities. Their performance is intricately tied to the human-to-machine ratio. Interestingly, as the scale expands, integration performance growth isn't limitless. It reaches a threshold influenced by model diversity. This introduces a pivotal `knee point', signifying the optimal balance between performance and scale. This knowledge is vital for resource allocation in practical applications. Grounded in rigorous evaluations using public datasets, our findings emphasize the framework's robustness in refining integrated systems.
Sannyuya Liu, Yawei Luo, Jintian Feng, Mengqi Wei
AAAI3
2024 Entangled View-Epipolar Information Aggregation for Generalizable Neural Radiance Fields
abstract
Generalizable NeRF can directly synthesize novel views across new scenes, eliminating the need for scene-specific re-training in vanilla NeRF. A critical enabling factor in these approaches is the extraction of a generalizable 3D representation by aggregating source-view features. In this paper, we propose an Entangled View-Epipolar Information Aggregation method dubbed EVE-NeRF. Differentfrom existing methods that consider cross-view and along-epipolar information independently, EVE-NeRF conducts the view-epipolar feature aggregation in an entangled manner by injecting the scene-invariant appearance continuity and geometry consistency priors to the aggregation process. Our approach effectively mitigates the potential lack of inherent geometric and appearance constraints resulting from one-dimensional interactions, thus further boosting the 3D representation generalizability. EVE-NeRF attains state-of-the-art performance across various evaluation scenarios. Extensive experiments demonstrate that, compared to pre-vailing single-dimensional aggregation, the entangled network excels in the accuracy of 3D scene geometry and appearance reconstruction. Our code is publicly available at https://github.com/tatakai1/EVENeRF.
Zhiyuan Min, Yawei Luo, Wei Yang 0011, Yuesong Wang 0001, Yi Yang 0001
CVPR2
2024 Epipolar-Free 3D Gaussian Splatting for Generalizable Novel View Synthesis
abstract
Generalizable 3D Gaussian splitting (3DGS) can reconstruct new scenes from sparse-view observations in a feed-forward inference manner, eliminating the need for scene-specific retraining required in conventional 3DGS. However, existing methods rely heavily on epipolar priors, which can be unreliable in complex real-world scenes, particularly in non-overlapping and occluded regions. In this paper, we propose eFreeSplat, an efficient feed-forward 3DGS-based model for generalizable novel view synthesis that operates independently of epipolar line constraints. To enhance multiview feature extraction with 3D perception, we employ a self-supervised Vision Transformer (ViT) with cross-view completion pre-training on large-scale datasets. Additionally, we introduce an Iterative Cross-view Gaussians Alignment method to ensure consistent depth scales across different views. Our eFreeSplat represents a new paradigm for generalizable novel view synthesis. We evaluate eFreeSplat on wide-baseline novel view synthesis tasks using the RealEstate10K and ACID datasets. Extensive experiments demonstrate that eFreeSplat surpasses state-of-the-art baselines that rely on epipolar priors, achieving superior geometry reconstruction and novel view synthesis quality.
Zhiyuan Min, Yawei Luo, Yi Yang 0001
NeurIPS2
2024 COMET : "cone of experience" enhanced large multimodal model for mathematical problem generation
Sannyuya Liu, Jintian Feng, Zongkai Yang, Yawei Luo, Qian Wan 0007, Xiaoxuan Shen
Sci. China Inf. Sci.4
2024 Large language model and domain-specific model collaboration for smart education
abstract
提出旨在增强智能教育的大型语言与领域特定模型协作(LDMC)框架。LDMC框架充分利用大型领域通用模型的综合全面知识,将其与小型领域特定模型的专业和学科知识相结合,并融入来自学习理论模型的教育学知识。这种整合产生的多重知识表达促进了个性化和自适应的教育体验。在智能教育背景下探讨了LDMC框架的各种应用,包括群体学习、个性化辅导、课堂管理等。LDMC融合了多种规模模型的智能,代表了一种先进而全面的教育辅助框架。随着人工智能的不断发展,该框架有望在智慧教育领域展现较大潜力。
Yawei Luo, Yi Yang 0001
Frontiers Inf. Technol. Electron. Eng.1
2024 Knowledge-Guided Causal Intervention for Weakly-Supervised Object Localization
abstract
Previous weakly-supervised object localization (WSOL) methods aim to expand activation map discriminative areas to cover the whole objects, yet neglect two inherent challenges when relying solely on image-level labels. First, the “entangled context” issue arises from object-context co-occurrence (e.g., fish and water), making the model inspection hard to distinguish object boundaries clearly. Second, the “C-L dilemma” issue results from the information decay caused by the pooling layers, which struggle to retain both the semantic information for precise classification and those essential details for accurate localization, leading to a trade-off in performance. In this paper, we propose a knowledge-guided causal intervention method, dubbed KG-CI-CAM, to address these two under-explored issues in one go. More specifically, we tackle the co-occurrence context confounder problem via causal intervention, which explores the causalities among image features, contexts, and categories to eliminate the biased object-context entanglement in the class activation maps. Based on the disentangled object feature, we introduce a multi-source knowledge guidance framework to strike a balance between absorbing classification knowledge and localization knowledge during model training. Extensive experiments conducted on several benchmark datasets demonstrate the effectiveness of KG-CI-CAM in learning distinct object boundaries amidst confounding contexts and mitigating the dilemma between classification and localization performance.
Feifei Shao, Yawei Luo, Fei Gao 0014, Yi Yang 0001, Jun Xiao 0001
IEEE Trans. Knowl. Data Eng.2
2024 Taking a Closer Look At Visual Relation: Unbiased Video Scene Graph Generation With Decoupled Label Learning
abstract
Current video-based scene graph generation (VidSGG) methods have been found to perform poorly in predicting predicates that are less represented due to the inherently biased distribution of the training data. In this paper, we take a closer look at the inherent characteristics of predicates and identify that most visual relations (e.g.sit_above) involve both actional pattern (sit) and spatial pattern (above), while the distribution bias is much less severe at the pattern level. Based on this insight, we propose a decoupled label learning (DLL) paradigm to address the intractable visual relation prediction from the pattern-level perspective. Specifically, DLL decouples the predicate labels and adopts separate classifiers to learn actional and spatial patterns respectively. The patterns are then combined and mapped back to the predicate. Moreover, we propose a knowledge-level label decoupling method to transfer non-target knowledge from head predicates to tail predicates within the same pattern to calibrate the distribution of tail classes. We validate the effectiveness of DLL on the commonly used VidSGG benchmark, i.e. VidVRD. Extensive experiments demonstrate that the DLL offers a remarkably simple but highly effective solution to the long-tailed problem, achieving the state-of-the-art VidSGG performance.
Yawei Luo, Zhiqing Chen, Tao Jiang 0042, Yi Yang 0001, Jun Xiao 0001
IEEE Trans. Multim.2
2023 Adaptive Patch Deformation for Textureless-Resilient Multi-View Stereo
abstract
In recent years, deep learning-based approaches have shown great strength in multi-view stereo because of their outstanding ability to extract robust visual features. However, most learning-based methods need to build the cost volume and increase the receptive field enormously to get a satisfactory result when dealing with large-scale textureless regions, consequently leading to prohibitive memory consumption. To ensure both memory-friendly and textureless-resilient, we innovatively transplant the spirit of deformable convolution from deep learning into the traditional PatchMatch-based method. Specifically, for each pixel with matching ambiguity (termed unreliable pixel), we adaptively deform the patch centered on it to extend the receptive field until covering enough correlative reliable pixels (without matching ambiguity) that serve as anchors. When performing PatchMatch, constrained by the anchor pixels, the matching cost of an unreliable pixel is guaranteed to reach the global minimum at the correct depth and therefore increases the robustness of multi-view stereo significantly. To detect more anchor pixels to ensure better adaptive patch deformation, we propose to evaluate the matching ambiguity of a certain pixel by checking the convergence of the estimated depth as optimization proceeds. As a result, our method achieves state-of-the-art performance on ETH3D and Tanks and Temples while preserving low memory consumption.
Yuesong Wang 0001, Zhaojie Zeng, Wei Yang 0011, Zhuo Chen 0054, Luoyuan Xu, Yawei Luo
CVPR8
2023 Unsupervised Domain Adaptive Person Re-Identification with Adaptive Structure Learning
abstract
Unsupervised domain adaptive person re-identification has garnered considerable attention due to its practical significance. Previous research has focused on utilizing the teacher-student framework within the clustering and fine-tuning paradigm to reduce the domain gap between different person re-identification datasets. Building upon these recent advancements, we propose a novel approach to further explore the structural information of the teacher and student networks from multiple perspectives, including adaptive sample updating, global feature arrangement, and local discrepancy learning. By integrating these three components, we introduce the Adaptive Structure Learning Framework (ASL) for unsupervised domain adaptive person re-identification. Compared to the baseline method, our approach achieves a significant improvement of 7.3% in mean Average Precision (mAP) on the Market2MSMT task. Furthermore, our experimental results across three benchmark datasets provide further evidence of the effectiveness of our proposed method.
Ping Liu 0004, Xiyuan Yang, Pan Zhou 0001, Yawei Luo, Jingen Liu
ICIP5
2023 Addressing Predicate Overlap in Scene Graph Generation with Semantic Granularity Controller
abstract
Semantic overlap between predicates (e.g., riding versus on) occurs inevitably when describing a scene. However, most existing Scene Graph Generation (SGG) works sidestep it by modeling the semantic overlap at category-level and assigning merely one-hot target to each sample, which hurt the performance on other reasonable predicates. In this paper, we argue that semantic overlap between predicates tends to vary in different abstract patterns, and a subject-object pair should retain multiple reasonable predicates. To this end, we make an early attempt to reformulate SGG as a partial multi-label learning problem and accordingly propose a model-agnostic Semantic Granularity Controller (SGC). SGC consists of a pattern-specific controller, partial multi-label learning, and controllable inference. The former two solve semantic confusion during training, while the latter makes the semantic granularity of prediction controllable. Extensive experiments demonstrate that SGC can improve the performance of SGG and guide the model to predict coarse/fine-grained predicates.
Guikun Chen, Lin Li 0065, Yawei Luo, Jun Xiao 0001
ICME3
2023 Dark Knowledge Balance Learning for Unbiased Scene Graph Generation
abstract
One of the major obstacles that hinders the current scene graph generation (SGG) performance lies in the severe predicate annotation bias. Conventional solutions to this problem are mainly based on reweighting/resampling heuristics. Despite achieving some improvements on tail classes, these methods are prone to cause serious performance degradation of head predicates. In this paper, we propose to tackle this problem from a brand-new perspective of dark knowledge. In consideration of the unique nature of SGG that requires a large number of negative samples to be employed for predicate learning, we design to capitalize on the dark knowledge contained in negative samples for debiasing the predicate distribution. Along such vein, we propose a novel SGG method dubbed Dark Knowledge Balance Learning (DKBL). In DKBL, we first design a dark knowledge balancing loss, which helps the model learn to balance head and tail predicates while maintaining the overall performance. We further introduce a dark knowledge semantic enhancement module to better encode the semantics of predicates. DKBL is orthogonal to existing SGG methods and can be easily plugged into their training process for further improvement. Extensive experiments on VG dataset show that the proposed DKBL can consistently achieve well trade-off performance between head and tail predicates, which is significantly better than previous state-of-the-art methods. The code is available in https://github.com/chenzqing/DKBL.
Zhiqing Chen, Yawei Luo, Jian Shao 0001, Yi Yang 0001, Chunping Wang 0001, Lei Chen 0082, Jun Xiao 0001
ACM Multimedia2
2023 Adversarial Bootstrapped Question Representation Learning for Knowledge Tracing
abstract
Knowledge tracing (KT), which estimates and traces the degree of learners' mastery of concepts based on students' responses to learning resources, has become an increasingly relevant problem in intelligent education. The accuracy of predictions greatly depends on the quality of question representations. While contrastive learning has been commonly used to generate high-quality representations, the selection of positive and negative samples for knowledge tracing remains a challenge. To address this issue, we propose an adversarial bootstrapped question representation (ABQR) model, which can generate robust and high-quality question representations without requiring negative samples. Specifically, ABQR introduces the bootstrap self-supervised learning framework, which learns question representations from different views of the skill-informed question interaction graph and facilitates question representations between each view to predict one another, thereby circumventing the need for negative sample selection. Moreover, we propose a multi-objective multi-round feature adversarial graph augmentation method to obtain a higher-quality target view, while preserving the structural information of the original graph. ABQR is versatile and can be easily integrated with any base KT model as a plug-in to enhance the quality of question representation. Extensive experiments demonstrate that ABQR significantly improves the performance of the base KT model and outperforms state-of-the-art models. Ablation experiments confirm the effectiveness of each module of ABQR. The code is available at https://github.com/lilstrawberry/ABQR.
Fenghua Yu, Sannyuya Liu, Yawei Luo, Ruxia Liang, Xiaoxuan Shen
ACM Multimedia4
2023 Triple Correlations-Guided Label Supplementation for Unbiased Video Scene Graph Generation
abstract
Video-based scene graph generation (VidSGG) is an approach that aims to represent video content in a dynamic graph by identifying visual entities and their relationships. Due to the inherently biased distribution and missing annotations in the training data, current VidSGG methods have been found to perform poorly on less-represented predicates. In this paper, we propose an explicit solution to address this under-explored issue by supplementing missing predicates that should be included in the ground-truth annotations. Dubbed Trico, our method seeks to supplement the missing predicates that are supposed to appear in the ground-truth annotations, by exploring three complementary spatio-temporal correlations. Guided by these correlations, the missing labels can be effectively supplemented thus achieving an unbiased predicate predictions. We validate the effectiveness of Trico on the most widely used VidSGG datasets, i.e., VidVRD and VidOR. Extensive experiments demonstrate the state-of-the-art performance achieved by Trico, particularly on those tail predicates. The code is available in the supplementary material.
Kaifeng Gao, Yawei Luo, Tao Jiang 0042, Fei Gao 0014, Jian Shao 0001, Jun Xiao 0001
ACM Multimedia3
2023 Personas-based Student Grouping using reinforcement learning and linear programming
Shaojie Ma, Yawei Luo, Yi Yang 0001
Knowl. Based Syst.2
2023 Local-Global Context Aware Transformer for Language-Guided Video Segmentation
abstract
We explore the task of language-guided video segmentation (LVS). Previous algorithms mostly adopt 3D CNNs to learn video representation, struggling to capture long-term context and easily suffering from visual-linguistic misalignment. In light of this, we present Locater (local-global context aware Transformer), which augments the Transformer architecture with a finite memory so as to query the entire video with the language expression in an efficient manner. The memory is designed to involve two components - one for persistently preserving global video content, and one for dynamically gathering local temporal context and segmentation history. Based on the memorized local-global context and the particular content of each frame, Locater holistically and flexibly comprehends the expression as an adaptive query vector for each frame. The vector is used to query the corresponding frame for mask generation. The memory also allows Locater to process videos with linear time complexity and constant size memory, while Transformer-style self-attention computation scales quadratically with sequence length. To thoroughly examine the visual grounding capability of LVS models, we contribute a new LVS dataset, A2D-S +, which is built upon A2D-S dataset but poses increased challenges in disambiguating among similar objects. Experiments on three LVS datasets and our A2D-S + show that Locater outperforms previous state-of-the-arts. Further, we won the 1st place in the Referring Video Object Segmentation Track of the 3rd Large-scale Video Object Segmentation Challenge, where Locater served as the foundation for the winning solution.
Wenguan Wang, Tianfei Zhou, Jiaxu Miao, Yawei Luo, Yi Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 DUDA: Online-Offline Dual Domain Adaption for Semantic Segmentation
An-tao Pan, Yawei Luo, Yi Yang 0001, Jun Xiao 0001
BMVC2
2022 Bidirectional Self-Training with Multiple Anisotropic Prototypes for Domain Adaptive Semantic Segmentation
abstract
A thriving trend for domain adaptive segmentation endeavors to generate the high-quality pseudo labels for target domain and retrain the segmentor on them. Under this self-training paradigm, some competitive methods have sought to the latent-space information, which establishes the feature centroids (a.k.a prototypes) of the semantic classes and determines the pseudo label candidates by their distances from these centroids. In this paper, we argue that the latent space contains more information to be exploited thus taking one step further to capitalize on it. Firstly, instead of merely using the source-domain prototypes to determine the target pseudo labels as most of the traditional methods do, we bidirectionally produce the target-domain prototypes to degrade those source features which might be too hard or disturbed for the adaptation. Secondly, existing attempts simply model each category as a single and isotropic prototype while ignoring the variance of the feature distribution, which could lead to the confusion of similar categories. To cope with this issue, we propose to represent each category with multiple and anisotropic prototypes via Gaussian Mixture Model, in order to fit the de facto distribution of source domain and estimate the likelihood of target samples based on the probability density. We apply our method on GTA5->Cityscapes and Synthia->Cityscapes tasks and achieve 61.2% and 62.8% respectively in terms of mean IoU, substantially outperforming other competitive self-training methods. Noticeably, in some categories which severely suffer from the categorical confusion such as "truck" and "bus", our method achieves 56.4% and 68.8% respectively, which further demonstrates the effectiveness of our design. The code and model are available at https://github.com/luyvlei/BiSMAPs.
Yulei Lu, Yawei Luo, Zheyang Li, Yi Yang 0001, Jun Xiao 0001
ACM Multimedia2
2022 Active Learning for Point Cloud Semantic Segmentation via Spatial-Structural Diversity Reasoning
abstract
The expensive annotation cost is notoriously known as the main constraint for the development of the point cloud semantic segmentation technique. Active learning methods endeavor to reduce such cost by selecting and labeling only a subset of the point clouds, yet previous attempts ignore the spatial-structural diversity of the selected samples, inducing the model to select clustered candidates with similar shapes in a local area while missing other representative ones in the global environment. In this paper, we propose a new 3D region-based active learning method to tackle this problem. Dubbed SSDR-AL, our method groups the original point clouds into superpoints and incrementally selects the most informative and representative ones for label acquisition. We achieve the selection mechanism via a graph reasoning network that considers both the spatial and structural diversities of superpoints. To deploy SSDR-AL in a more practical scenario, we design a noise-aware iterative labeling strategy to confront the "noisy annotation'' problem introduced by the previous "dominant labeling'' strategy in superpoints. Extensive experiments on two point cloud benchmarks demonstrate the effectiveness of SSDR-AL in the semantic segmentation task. Particularly, SSDR-AL significantly outperforms the baseline method and reduces the annotation cost by up to $63.0%$ and $24.0%$ when achieving $90%$ performance of fully supervised learning, respectively. Code is available at https://github.com/shaofeifei11/SSDR-AL.
Feifei Shao, Yawei Luo, Ping Liu 0004, Yi Yang 0001, Yulei Lu, Jun Xiao 0001
ACM Multimedia2
2022 Self-Supervised Multi-view Stereo via Adjacent Geometry Guided Volume Completion
abstract
Existing self-supervised multi-view stereo (MVS) approaches largely rely on photometric consistency for geometry inference, and hence suffer from low-texture or non-Lambertian appearances. In this paper, we observe that adjacent geometry shares certain commonality that can help to infer the correct geometry of the challenging or low-confident regions. Yet exploiting such property in a non-supervised MVS approach remains challenging for the lacking of training data and necessity of ensuring consistency between views. To address the issues, we propose a novel geometry inference training scheme by selectively masking regions with rich textures, where geometry can be well recovered and used for supervisory signal, and then lead a deliberately designed cost volume completion network to learn how to recover geometry of the masked regions. During inference, we then mask the low-confident regions instead and use the cost volume completion network for geometry correction. To deal with the different depth hypotheses of the cost volume pyramid, we design a three-branch volume inference structure for the completion network. Further, by considering plane as a special geometry, we first identify planar regions from pseudo labels and then correct the low-confident pixels by high-confident labels through plane normal consistency. Extensive experiments on DTU and Tanks & Temples demonstrate the effectiveness of the proposed framework and the state-of-the-art performance.
Luoyuan Xu, Yuesong Wang 0001, Yawei Luo, Zhuo Chen 0054, Wei Yang 0011
ACM Multimedia4
2022 Category-Level Adversarial Adaptation for Semantic Segmentation Using Purified Features
abstract
We target the problem named unsupervised domain adaptive semantic segmentation. A key in this campaign consists in reducing the domain shift, so that a classifier based on labeled data from one domain can generalize well to other domains. With the advancement of adversarial learning method, recent works prefer the strategy of aligning the marginal distribution in the feature spaces for minimizing the domain discrepancy. However, based on the observance in experiments, only focusing on aligning global marginal distribution but ignoring the local joint distribution alignment fails to be the optimal choice. Other than that, the noisy factors existing in the feature spaces, which are not relevant to the target task, entangle with the domain invariant factors improperly and make the domain distribution alignment more difficult. To address those problems, we introduce two new modules, Significance-aware Information Bottleneck (SIB) and Category-level alignment (CLA), to construct a purified embedding-based category-level adversarial network. As the name suggests, our designed network, CLAN, can not only disentangle the noisy factors and suppress their influences for target tasks but also utilize those purified features to conduct a more delicate level domain calibration, i.e., global marginal distribution and local joint distribution alignment simultaneously. In three domain adaptation tasks, i.e., GTA5 → Cityscapes, SYNTHIA → Cityscapes and Cross Season, we validate that our proposed method matches the state of the art in segmentation accuracy.
Yawei Luo, Ping Liu 0004, Liang Zheng 0001, Junqing Yu, Yi Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Recursive Copy and Paste GAN: Face Hallucination From Shaded Thumbnails
abstract
Existing face hallucination methods based on convolutional neural networks (CNNs) have achieved impressive performance on low-resolution (LR) faces in a normal illumination condition. However, their performance degrades dramatically when LR faces are captured in non-uniform illumination conditions. This paper proposes a Recursive Copy and Paste Generative Adversarial Network (Re-CPGAN) to recover authentic high-resolution (HR) face images while compensating for non-uniform illumination. To this end, we develop two key components in our Re-CPGAN: internal and recursive external Copy and Paste networks (CPnets). Our internal CPnet exploits facial self-similarity information residing in the input image to enhance facial details; while our recursive external CPnet leverages an external guided face for illumination compensation. Specifically, our recursive external CPnet stacks multiple external Copy and Paste (EX-CP) units in a compact model to learn normal illumination and enhance facial details recursively. By doing so, our method offsets illumination and upsamples facial details progressively in a coarse-to-fine fashion, thus alleviating the ambiguity of correspondences between LR inputs and external guided inputs. Furthermore, a new illumination compensation loss is developed to capture illumination from the external guided face image effectively. Extensive experiments demonstrate that our method achieves authentic HR face images in a uniform illumination condition with a 16× magnification factor and outperforms state-of-the-art methods qualitatively and quantitatively.
Yang Zhang 0067, Ivor W. Tsang, Yawei Luo, Changhui Hu 0001, Xiaobo Lu, Xin Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Temporal Cross-Layer Correlation Mining for Action Recognition
abstract
Neighboring frames are more correlated compared to frames from further temporal distances. In this paper, we aim to explore the temporal correlations among neighboring frames and exploit cross-layer multi-scale features for action recognition. First, we present a Temporal Cross-Layer Correlation (TCLC) framework for temporal correlation learning. The unified framework uncovers both local and global structures from video data, enabling a better exploration of temporal context and assisting cross-layer spatio-temporal feature learning. Second, we propose a novel cross-layer attention and a center-guided attention mechanism to integrate features with contextual knowledge from multiple scales. Our method is a two-stage process for effective cross-layer feature learning. The first stage incorporates the cross-layer attention module to decide the importance weight of the convolutional layers. The second stage leverages the center-guided attention mechanism to aggregate local features from each layer for the generation of a final video representation. We leverage global centers to extract shared semantic knowledge among videos. We evaluate TCLC on three action recognition datasets, i.e., UCF-101, HMDB-51 and Kinetics. Our experimental results demonstrate the superiority of our proposed temporal correlation mining method.
Linchao Zhu, Hehe Fan, Yawei Luo, Mingliang Xu 0001, Yi Yang 0001
IEEE Trans. Multim.3
2021 Improving Weakly Supervised Object Localization via Causal Intervention
abstract
The recently emerged weakly-supervised object localization (WSOL) methods can learn to localize an object in the image only using image-level labels. Previous works endeavor to perceive the interval objects from the small and sparse discriminative attention map, yet ignoring the co-occurrence confounder (e.g., duck and water), which makes the model inspection (e.g., CAM) hard to distinguish between the object and context. In this paper, we make an early attempt to tackle this challenge via causal intervention (CI). Our proposed method, dubbed CI-CAM, explores the causalities among image features, contexts, and categories to eliminate the biased object-context entanglement in the class activation maps thus improving the accuracy of object localization. Extensive experiments on several benchmarks demonstrate the effectiveness of CI-CAM in learning the clear object boundary from confounding contexts. Particularly, on the CUB-200-2011 which severely suffers from the co-occurrence confounder, CI-CAM significantly outperforms the traditional CAM-based baseline (58.39% vs 52.4% in Top-1 localization accuracy). While in more general scenarios such as ILSVRC 2016, CI-CAM can also perform on par with the state of the arts.
Feifei Shao, Yawei Luo, Lu Ye, Siliang Tang, Yi Yang 0001, Jun Xiao 0001
ACM Multimedia2
2021 A visibility-based surface reconstruction method on the GPU
Hailong Pan, Keyang Luo, Yawei Luo, Junqing Yu
Comput. Aided Geom. Des.4
2021 Discriminative Feature Learning for Thorax Disease Classification in Chest X-ray Images
abstract
This paper focuses on the thorax disease classification problem in chest X-ray (CXR) images. Different from the generic image classification task, a robust and stable CXR image analysis system should consider the unique characteristics of CXR images. Particularly, it should be able to: 1) automatically focus on the disease-critical regions, which usually are of small sizes; 2) adaptively capture the intrinsic relationships among different disease features and utilize them to boost the multi-label disease recognition rates jointly. In this paper, we propose to learn discriminative features with a two-branch architecture, named ConsultNet, to achieve those two purposes simultaneously. ConsultNet consists of two components. First, an information bottleneck constrained feature selector extracts critical disease-specific features according to the feature importance. Second, a spatial-and-channel encoding based feature integrator enhances the latent semantic dependencies in the feature space. ConsultNet fuses these discriminative features to improve the performance of thorax disease classification in CXRs. Experiments conducted on the ChestX-ray14 and CheXpert dataset demonstrate the effectiveness of the proposed method.
Qingji Guan, Yawei Luo, Ping Liu 0004, Mingliang Xu 0001, Yi Yang 0001
IEEE Trans. Image Process.3
2021 Few-Shot Common-Object Reasoning Using Common-Centric Localization Network
abstract
In the few-shot common-localization task, given few support images without bounding box annotations at each episode, the goal is to localize the common object in the query image of unseen categories. The few-shot common-localization task involves common object reasoning from the given images, predicting the spatial locations of the object with different shapes, sizes, and orientations. In this work, we propose a common-centric localization (CCL) network for few-shot common-localization. The motivation of our common-centric localization network is to learn the common object features by dynamic feature relation reasoning via a graph convolutional network with conditional feature aggregation. First, we propose a local common object region generation pipeline to reduce background noises due to feature misalignment. Each support image predicts more accurate object spatial locations by replacing the query with the images in the support set. Second, we introduce a graph convolutional network with dynamic feature transformation to enforce the common object reasoning. To enhance the discriminability during feature matching and enable a better generalization in unseen scenarios, we leverage a conditional feature encoding function to alter visual features according to the input query adaptively. Third, we introduce a common-centric relation structure to model the correlation between the common features and the query image feature. The generated common features guide the query image feature towards a more common object-related representation. We evaluate our common-centric localization network on four datasets, i.e., CL-VOC-07, CL-VOC-12, CL-COCO, CL-VID. We obtain significant improvements compared to state-of-the-art. Our quantitative results confirm the effectiveness of our network.
Linchao Zhu, Hehe Fan, Yawei Luo, Mingliang Xu 0001, Yi Yang 0001
IEEE Trans. Image Process.3
2020 Attention-Aware Multi-View Stereo
abstract
Multi-view stereo is a crucial task in computer vision, that requires accurate and robust photo-consistency among input images for depth estimation. Recent studies have shown that learning-based feature matching and confidence regularization can play a vital role in this task. Nevertheless, how to design good matching confidence volumes as well as effective regularizers for them are still under in-depth study. In this paper, we propose an attention-aware deep neural network “AttMVS” for learning multi-view stereo. In particular, we propose a novel attention-enhanced matching confidence volume, that combines the raw pixel-wise matching confidence from the extracted perceptual features with the contextual information of local scenes, to improve the matching robustness. Furthermore, we develop an attention-guided regularization module, which consists of multilevel ray fusion modules, to hierarchically aggregate and regularize the matching confidence volume into a latent depth probability volume.Experimental results show that our approach achieves the best overall performance on the DTU dataset and the intermediate sequences of Tanks & Temples benchmark over many state-of-the-art MVS algorithms.
Keyang Luo, Lili Ju, Yuesong Wang 0001, Zhuo Chen 0054, Yawei Luo
CVPR6
2020 Mesh-Guided Multi-View Stereo With Pyramid Architecture
abstract
Multi-view stereo (MVS) aims to reconstruct 3D geometry of the target scene by using only information from 2D images. Although much progress has been made, it still suffers from textureless regions. To overcome this difficulty, we propose a mesh-guided MVS method with pyramid architecture, which makes use of the surface mesh obtained from coarse-scale images to guide the reconstruction process. Specifically, a PatchMatch-based MVS algorithm is first used to generate depth maps for coarse-scale images and the corresponding surface mesh is obtained by a surface reconstruction algorithm. Next we project the mesh onto each of depth maps to replace unreliable depth values and the corrected depth maps are fed to fine-scale reconstruction for initialization. To alleviate the influence of possible erroneous faces on the mesh, we further design and train a convolutional neural network to remove incorrect depths. In addition, it is often hard for the correct depth values for low-textured regions to survive at the fine-scale, thus we also develop an efficient method to seek out these regions and further enforce the geometric consistency in these regions. Experimental results on the ETH3D high-resolution dataset demonstrate that our method achieves state-of-the-art performance, especially in completeness.
Yuesong Wang 0001, Zhuo Chen 0054, Yawei Luo, Keyang Luo, Lili Ju
CVPR4
2020 Copy and Paste GAN: Face Hallucination From Shaded Thumbnails
abstract
Existing face hallucination methods based on convolutional neural networks (CNN) have achieved impressive performance on low-resolution (LR) faces in a normal illumination condition. However, their performance degrades dramatically when LR faces are captured in low or non-uniform illumination conditions. This paper proposes a Copy and Paste Generative Adversarial Network (CPGAN) to recover authentic high-resolution (HR) face images while compensating for low and non-uniform illumination. To this end, we develop two key components in our CPGAN: internal and external Copy and Paste nets (CPnets). Specifically, our internal CPnet exploits facial information residing in the input image to enhance facial details; while our external CPnet leverages an external HR face for illumination compensation. A new illumination compensation loss is thus developed to capture illumination from the external guided face image effectively. Furthermore, our method offsets illumination and upsamples facial details alternatively in a coarse-to-fine fashion, thus alleviating the correspondence ambiguity between LR inputs and external HR inputs. Extensive experiments demonstrate that our method manifests authentic HR face images in a uniform illumination condition and outperforms state-of-the-art methods qualitatively and quantitatively.
Yang Zhang 0067, Ivor W. Tsang, Yawei Luo, Changhui Hu 0001, Xiaobo Lu, Xin Yu 0002
CVPR3
2020 PC-Net: A Deep Network for 3D Point Clouds Analysis
abstract
Due to the irregularity and sparsity of 3D point clouds, applying convolutional neural networks directly on them can be nontrivial. In this work, we propose a simple but effective approach for 3D Point Clouds analysis, named PC-Net. PC-Net directly learns on point sets and is equipped with three new operations: first, we apply a novel scale-aware neighbor search for adaptive neighborhood extracting; second, for each neighboring point, we learn a local spatial feature as a complement to their associated features; finally, at the end we use a distance re-weighted pooling to aggregate all the features from local structure. With this module, we design hierarchical neural network for point cloud understanding. For both classification and segmentation tasks, our architecture proves effective in the experiments and our models demonstrate state-of-the-art performance over existing deep learning methods on popular point cloud benchmarks.
Zhuo Chen 0054, Yawei Luo, Yuesong Wang 0001, Keyang Luo, Luoyuan Xu
ICPR3
2020 Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation
abstract
We aim at the problem named One-Shot Unsupervised Domain Adaptation. Unlike traditional Unsupervised Domain Adaptation, it assumes that only one unlabeled target sample can be available when learning to adapt. This setting is realistic but more challenging, in which conventional adaptation approaches are prone to failure due to the scarce of unlabeled target data. To this end, we propose a novel Adversarial Style Mining approach, which combines the style transfer module and task-specific module into an adversarial manner. Specifically, the style transfer module iteratively searches for harder stylized images around the one-shot target sample according to the current learning state, leading the task model to explore the potential styles that are difficult to solve in the almost unseen target domain, thus boosting the adaptation performance in a data-scarce scenario. The adversarial learning framework makes the style transfer module and task-specific module benefit each other during the competition. Extensive experiments on both cross-domain classification and segmentation benchmarks verify that ASM achieves state-of-the-art adaptation performance under the challenging one-shot setting.
Yawei Luo, Ping Liu 0004, Junqing Yu, Yi Yang 0001
NeurIPS1
2020 Every node counts: Self-ensembling graph convolutional networks for semi-supervised learning
Yawei Luo, Rongrong Ji, Junqing Yu, Ping Liu 0004, Yi Yang 0001
Pattern Recognit.1
2019 Taking a Closer Look at Domain Shift: Category-Level Adversaries for Semantics Consistent Domain Adaptation
abstract
We consider the problem of unsupervised domain adaptation in semantic segmentation. The key in this campaign consists in reducing the domain shift, i.e., enforcing the data distributions of the two domains to be similar. A popular strategy is to align the marginal distribution in the feature space through adversarial learning. However, this global alignment strategy does not consider the local category-level feature distribution. A possible consequence of the global movement is that some categories which are originally well aligned between the source and target may be incorrectly mapped. To address this problem, this paper introduces a category-level adversarial network, aiming to enforce local semantic consistency during the trend of global alignment. Our idea is to take a close look at the category-level data distribution and align each class with an adaptive adversarial loss. Specifically, we reduce the weight of the adversarial loss for category-level aligned features while increasing the adversarial force for those poorly aligned. In this process, we decide how well a feature is category-level aligned between source and target by a co-training approach. In two domain adaptation tasks, i.e., GTA5 → Cityscapes and SYNTHIA → Cityscapes, we validate that the proposed method matches the state of the art in segmentation accuracy.
Yawei Luo, Liang Zheng 0001, Junqing Yu, Yi Yang 0001
CVPR1
2019 P-MVSNet: Learning Patch-Wise Matching Confidence Aggregation for Multi-View Stereo
abstract
Learning-based methods are demonstrating their strong competitiveness in estimating depth for multi-view stereo reconstruction in recent years. Among them the approaches that generate cost volumes based on the plane-sweeping algorithm and then use them for feature matching have shown to be very prominent recently. The plane-sweep volumes are essentially anisotropic in depth and spatial directions, but they are often approximated by isotropic cost volumes in those methods, which could be detrimental. In this paper, we propose a new end-to-end deep learning network of P-MVSNet for multi-view stereo based on isotropic and anisotropic 3D convolutions. Our P-MVSNet consists of two core modules: a patch-wise aggregation module learns to aggregate the pixel-wise correspondence information of extracted features to generate a matching confidence volume, from which a hybrid 3D U-Net then infers a depth probability distribution and predicts the depth maps. We perform extensive experiments on the DTU and Tanks & Temples benchmark datasets, and the results show that the proposed P-MVSNet achieves the state-of-the-art performance over many existing methods on multi-view stereo.
Keyang Luo, Lili Ju, Haipeng Huang, Yawei Luo
ICCV5
2019 Significance-Aware Information Bottleneck for Domain Adaptive Semantic Segmentation
abstract
For unsupervised domain adaptation problems, the strategy of aligning the two domains in latent feature space through adversarial learning has achieved much progress in image classification, but usually fails in semantic segmentation tasks in which the latent representations are overcomplex. In this work, we equip the adversarial network with a “significance-aware information bottleneck (SIB)”, to address the above problem. The new network structure, called SIBAN, enables a significance-aware feature purification before the adversarial adaptation, which eases the feature alignment and stabilizes the adversarial training course. In two domain adaptation tasks, i.e., GTA5 → Cityscapes and SYNTHIA → Cityscapes, we validate that the proposed method can yield leading results compared with other feature-space alternatives. Moreover, SIBAN can even match the state-of-the-art output-space methods in segmentation accuracy, while the latter are often considered to be better choices for domain adaptive segmentation task.
Yawei Luo, Ping Liu 0004, Junqing Yu, Yi Yang 0001
ICCV1
2018 Macro-Micro Adversarial Network for Human Parsing
Yawei Luo, Zhedong Zheng, Liang Zheng 0001, Junqing Yu, Yi Yang 0001
ECCV (9)1
2016 Accurate localization for mobile device using a multi-planar city model
abstract
This paper presents a novel method for estimating the unknown 6DOF pose of a mobile device. The method is based on matching between the mobile image and the virtual city model which is merely composed of 3D points on planar building facade. The main contributions of this paper are as follows: firstly, we design a new plane generation strategy which fuses the 3D model points, photo homography and the orientation of the buildings together within RANSAC framework. Secondly, we propose a novel energy-based method which can parallel solve the mobile poses as well as the best 2D-3D matches. Thirdly, a client/server mode is established to support a speedy localization experience on the mobile devices. To the best of our knowledge, this is the first implementation that uses such a multi-planar model to accurately locate the mobile device in a scene of city scale. Experiment shows that the localization performance becomes faster and more robust comparing with other methods when adding the multi-planar information of the buildings to localization algorithm.
Yawei Luo, Hailong Pan, Yuesong Wang 0001, Junqing Yu
ICPR1
2016 Dense 3D reconstruction combining depth and RGB information
Hailong Pan, Yawei Luo, Liya Duan, Liu Yi, Yizhu Zhao, Junqing Yu
Neurocomputing3
2015 Fast terrain mapping from low altitude digital imagery
Yawei Luo, Benchang Wei, Hailong Pan, Junqing Yu
Neurocomputing1