EDBT 2026 Demo / reviewers in the wild / expert
Yunfei Liu 0001
dblp:136/3330-1
· DBLP profile ↗
27ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0001-6898-0058ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 19 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identity-Preserving Video Dubbing Using Motion Warping
Runzhen Liu, Qinjie Lin, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Yu Li 0003, Chuhua Xian, Fa-Ting Hong |
Int. J. Comput. Vis. | 3 |
| 2025 | AnyTalk: Multi-modal Driven Multi-domain Talking Head GenerationabstractCross-domain talking head generation, such as animating a static cartoon animal photo with real human video, is crucial for personalized content creation. However, prior works typically rely on domain-specific frameworks and paired videos, limiting its utility and complicating its architecture with additional motion alignment modules. Addressing these shortcomings, we propose Anytalk, a unified framework that eliminates the need for paired data and learns a shared motion representation across different domains. The motion is represented by canonical 3D keypoints extracted using an unsupervised 3D keypoint detector. Further, we propose an expression consistency loss to improve the accuracy of facial dynamics in video generation. Additionally, we present AniTalk, a comprehensive dataset designed for advanced multi-modal cross-domain generation. Our experiments demonstrate that Anytalk excels at generating high-quality, multi-modal talking head videos, showcasing remarkable generalization capabilities across diverse domains. Yu Wang 0027, Yunfei Liu 0001, Fa-Ting Hong, Lijian Lin, Yu Li 0003 |
AAAI | 2 |
| 2025 | HRAvatar: High-Quality and Relightable Gaussian Head AvatarabstractReconstructing animatable and high-quality 3D head avatars from monocular videos, especially with realistic relighting, is a valuable task. However, the limited information from single-view input, combined with the complex head poses and facial movements, makes this challenging. Previous methods achieve real-time performance by combining 3D Gaussian Splatting with a parametric head model, but the resulting head quality suffers from inaccurate face tracking and limited expressiveness of the deformation model. These methods also fail to produce realistic effects under novel lighting conditions. To address these issues, we propose HRAvatar, a 3DGS-based method that reconstructs high-fidelity, relightable 3D head avatars. HRA-vatar reduces tracking errors through end-to-end optimization and better captures individual facial deformations using learnable blendshapes and learnable linear blend skinning. Additionally, it decomposes head appearance into several physical properties and incorporates physically-based shading to account for environmental lighting. Extensive experiments demonstrate that HRAvatar not only reconstructs superior-quality heads but also achieves realistic visual effects under varying lighting conditions. Video results and code are available at the project page. Dongbin Zhang, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Kangjie Chen, Minghan Qin, Yu Li 0003, Haoqian Wang |
CVPR | 2 |
| 2025 | Canonswap: High-Fidelity and Consistent Video Face Swapping Via Canonical Space Modulation
Ye Zhu 0003, Yunfei Liu 0001, Lijian Lin, Cong Wan, Zijian Cai, Yu Li 0003, Shao-Lun Huang |
ICCV | 3 |
| 2025 | GUAVA: Generalizable Upper Body 3D Gaussian AvatarabstractReconstructing a high-quality, animatable 3D human avatar with expressive facial and hand motions from a single image has gained significant attention due to its broad application potential. 3D human avatar reconstruction typically requires multi-view or monocular videos and training on individual IDs, which is both complex and time-consuming. Furthermore, limited by SMPLX's expressiveness, these methods often focus on body motion but struggle with facial expressions. To address these challenges, we first introduce an expressive human model (EHM) to enhance facial expression capabilities and develop an accurate tracking method. Based on this template model, we propose GUAVA, the first framework for fast animatable upper-body 3D Gaussian avatar reconstruction. We leverage inverse texture mapping and projection sampling techniques to infer Ubody (upper-body) Gaussians from a single image. The rendered images are refined through a neural refiner. Experimental results demonstrate that GUAVA significantly outperforms previous methods in rendering quality and offers significant speed improvements, with reconstruction times in the sub-second range (0.1s), and supports real-time animation and rendering. Dongbin Zhang, Yunfei Liu 0001, Lijian Lin, Ye Zhu 0003, Minghan Qin, Yu Li 0003, Haoqian Wang |
ICCV | 2 |
| 2025 | Qffusion: Controllable Portrait Video Editing via Quadrant-Grid Attention LearningabstractThis paper presents Qffusion, a dual-frame-guided framework for portrait video editing. Specifically, we consider a design principle of "animation for editing", and train Qffusion as a general animation framework from two still reference images while we can use it for portrait video editing easily by applying modified start and end frames as references during inference. Leveraging the powerful generative power of Stable Diffusion, we propose a Quadrant-grid Arrangement (QGA) scheme for latent re-arrangement, which arranges the latent codes of two reference images and that of four facial conditions into a four-grid fashion, separately. Then, we fuse features of these two modalities and use self-attention for both appearance and temporal learning, where representations at different times are jointly modeled under QGA. Our Qffusion can achieve stable video editing without additional networks or complex training stages, where only the input format of Stable Diffusion is modified. Further, we propose a Quadrant-grid Propagation (QGP) inference strategy, which enjoys a unique advantage on stable arbitrary-length video generation by processing reference and condition frames recursively. Through extensive experiments, Qffusion consistently outperforms state-of-the-art techniques on portrait video editing. Maomao Li, Lijian Lin, Yunfei Liu 0001, Ye Zhu 0003, Yu Li 0003 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-Driven Holistic 3D Expression and Gesture GenerationabstractWe propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length. While previous works focused on co-speech gesture or expression generation individually, the joint generation of synchronized expressions and gestures remains barely explored. To address this, our diffusion-based co-speech motion generation transformer enables uni-directional information flow from expression to gesture, facilitating improved matching of joint expression-gesture distributions. Furthermore, we introduce an outpainting-based sampling strategy for arbitrary long sequence generation in diffusion models, offering flexibility and computational efficiency. Our method provides a practical solution that produces high-quality synchronized expression and gesture generation driven by speech. Evaluated on two public datasets, our approach achieves state-of-the-art performance both quantitatively and qualitatively. Additionally, a user study confirms the superiority of DiffSHEG over prior approaches. By enabling the real-time generation of expressive and synchronized motions, DiffSHEG showcases its potential for various applications in the development of digital humans and embodied agents. Yunfei Liu 0001, Ailing Zeng, Yu Li 0003, Qifeng Chen 0001 |
CVPR | 2 |
| 2024 | A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video EditingabstractThis paper presents a video inversion approach for zero-shot video editing, which models the input video with low-rank representation during the inversion process. The existing video editing methods usually apply the typical 2D DDIM inversion or naï ve spatial-temporal DDIM inversion before editing, which leverages time-varying representation for each frame to derive noisy latent. Unlike most existing approaches, we propose a Spatial-Temporal Expectation-Maximization (STEM) inversion, which formulates the dense video feature under an expectation-maximization manner and iteratively estimates a more compact basis set to represent the whole video. Each frame applies the fixed and global representation for inversion, which is more friendly for temporal consistency during reconstruction and editing. Extensive qualitative and quantitative experiments demonstrate that our STEM inversion can achieve consistent improvement on two state-of-the-art video editing methods. Project page: https://steminv.github.io/page/. Maomao Li, Yu Li 0003, Tianyu Yang 0003, Yunfei Liu 0001, Dongxu Yue, Zhihui Lin |
CVPR | 4 |
| 2024 | AddMe: Zero-Shot Group-Photo Synthesis by Inserting People Into Scenes
Dongxu Yue, Maomao Li, Yunfei Liu 0001, Ailing Zeng, Tianyu Yang 0003, Yu Li 0003 |
ECCV (20) | 3 |
| 2024 | A Cross-Consistency Strategy for Clearer Perception in Low-Light Haze
Sijia Wen, Chaoqun Zhuang, Yunfei Liu 0001, Feng Lu 0005 |
PRCV (8) | 3 |
| 2024 | PnP-GA+: Plug-and-Play Domain Adaptation for Gaze Estimation Using Model VariantsabstractAppearance-based gaze estimation has garnered increasing attention in recent years. However, deep learning-based gaze estimation models still suffer from suboptimal performance when deployed in new domains, e.g., unseen environments or individuals. In our previous work, we took this challenge for the first time by introducing a plug-and-play method (PnP-GA) to adapt the gaze estimation model to new domains. The core concept of PnP-GA is to leverage the diversity brought by a group of model variants to enhance the adaptability to diverse environments. In this article, we propose the PnP-GA+ by extending our approach to explore the impact of assembling model variants using three additional perspectives: color space, data augmentation, and model structure. Moreover, we propose an intra-group attention module that dynamically optimizes pseudo-labeling during adaptation. Experimental results demonstrate that by directly plugging several existing gaze estimation networks into the PnP-GA+ framework, it outperforms state-of-the-art domain adaptation approaches on four standard gaze domain adaptation tasks on public datasets. Our method consistently enhances cross-domain performance, and its versatility is improved through various ways of assembling the model group. Ruicong Liu, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | MODA: Mapping-Once Audio-driven Portrait Animation with Dual AttentionsabstractAudio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of applications. Previous methods have attempted to capture different motion modes and generate high-fidelity portrait videos by training different models or sampling signals from given videos. However, lacking correlation learning between lip-sync and other movements (e.g., head pose/eye blinking) usually leads to unnatural results. In this paper, we propose a unified system for multi-person, diverse, and high-fidelity talking portrait generation. Our method contains three stages, i.e., 1) Mapping-Once network with Dual Attentions (MODA) generates talking representation from given audio. In MODA, we design a dual-attention module to encode accurate mouth movements and diverse modalities. 2) Facial composer network generates dense and detailed face landmarks, and 3) temporal-guided renderer syntheses stable videos. Extensive evaluations demonstrate that the proposed system produces more natural and realistic video portraits compared to previous methods. Yunfei Liu 0001, Lijian Lin, F. Richard Yu, Changyin Zhou, Yu Li 0003 |
ICCV | 1 |
| 2023 | Accurate 3D Face Reconstruction with Facial Component TokensabstractAccurately reconstructing 3D faces from monocular images and videos is crucial for various applications, such as digital avatar creation. However, the current deep learning-based methods face significant challenges in achieving accurate reconstruction with disentangled facial parameters and ensuring temporal stability in single-frame methods for 3D face tracking on video data. In this paper, we propose TokenFace, a transformer-based monocular 3D face reconstruction model. TokenFace uses separate tokens for different facial components to capture information about different facial parameters and employs temporal transformers to capture temporal information from video data. This design can naturally disentangle different facial components and is flexible to both 2D and 3D training data. Trained on hybrid 2D and 3D data, our model shows its power in accurately reconstructing faces from images and producing stable results for video data. Experimental results on popular benchmarks NoWand Stirling demonstrate that TokenFace achieves state-of-the-art performance, outperforming existing methods on all metrics by a large margin. Tianke Zhang, Xuangeng Chu, Yunfei Liu 0001, Lijian Lin, Zhendong Yang, Zhengzhuo Xu, Chengkun Cao, F. Richard Yu, Changyin Zhou, Chun Yuan 0003, Yu Li 0003 |
ICCV | 3 |
| 2023 | Discriminative feature encoding for intrinsic image decompositionabstractIntrinsic image decomposition is an important and long-standing computer vision problem. Given an input image, recovering the physical scene properties is ill-posed. Several physically motivated priors have been used to restrict the solution space of the optimization problem for intrinsic image decomposition. This work takes advantage of deep learning, and shows that it can solve this challenging computer vision problem with high efficiency. The focus lies in the feature encoding phase to extract discriminative features for different intrinsic layers from an input image. To achieve this goal, we explore the distinctive characteristics of different intrinsic components in the high-dimensional feature embedding space. We define feature distribution divergence to efficiently separate the feature vectors of different intrinsic components. The feature distributions are also constrained to fit the real ones through a feature distribution consistency. In addition, a data refinement approach is provided to remove data inconsistency from the Sintel dataset, making it more suitable for intrinsic image decomposition. Our method is also extended to intrinsic video decomposition based on pixel-wise correspondences between adjacent frames. Experimental results indicate that our proposed network structure can outperform the existing state-of-the-art. Zongji Wang, Yunfei Liu 0001, Feng Lu 0005 |
Comput. Vis. Media | 2 |
| 2023 | First- And Third-Person Video Co-Analysis By Learning Spatial-Temporal Joint AttentionabstractRecent years have witnessed a tremendous increase of first-person videos captured by wearable devices. Such videos record information from different perspectives than the traditional third-person view, and thus show a wide range of potential usages. However, techniques for analyzing videos from different views can be fundamentally different, not to mention co-analyzing on both views to explore the shared information. In this paper, we take the challenge of cross-view video co-analysis and deliver a novel learning-based method. At the core of our method is the notion of "joint attention", indicating the shared attention regions that link the corresponding views, and eventually guide the shared representation learning across views. To this end, we propose a multi-branch deep network, which extracts cross-view joint attention and shared representation from static frames with spatial constraints, in a self-supervised and simultaneous manner. In addition, by incorporating the temporal transition model of the joint attention, we obtain spatial-temporal joint attention that can robustly capture the essential information extending through time. Our method outperforms the state-of-the-art on the standard cross-view video matching tasks on public datasets. Furthermore, we demonstrate how the learnt joint information can benefit various applications through a set of qualitative and quantitative experiments. Huangyue Yu, Minjie Cai, Yunfei Liu 0001, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Generalizing Gaze Estimation with Rotation ConsistencyabstractRecent advances of deep learning-based approaches have achieved remarkable performance on appearance-based gaze estimation. However, due to the shortage of target domain data and absence of target labels, generalizing gaze estimation algorithm to unseen environments is still challenging. In this paper, we discover the rotation-consistency property in gaze estimation and introduce the ‘sub-label’ for unsupervised domain adaptation. Consequently, we propose the Rotation-enhanced Unsupervised Domain Adaptation (RUDA) for gaze estimation. First, we rotate the original images with different angles for training. Then we conduct domain adaptation under the constraint of rotation consistency. The target domain images are assigned with sub-labels, derived from relative rotation angles rather than untouchable real labels. With such sub-labels, we propose a novel distribution loss that facilitates the domain adaptation. We evaluate the RUDA framework on four cross-domain gaze estimation tasks. Experimental results demonstrate that it improves the performance over the baselines with gains ranging from 12.2% to 30.5%. Our framework has the potential to be used in other computer vision tasks with physical constraints. Yiwei Bao, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
CVPR | 2 |
| 2022 | GazeOnce: Real-Time Multi-Person Gaze EstimationabstractAppearance-based gaze estimation aims to predict the 3D eye gaze direction from a single image. While recent deep learning-based approaches have demonstrated excellent performance, they usually assume one calibrated face in each input image and cannot output multi-person gaze in real time. However, simultaneous gaze estimation for multiple people in the wild is necessary for real-world applications. In this paper, we propose the first one-stage end-to-end gaze estimation method, GazeOnce, which is capable of simultaneously predicting gaze directions for multiple faces (> 10) in an image. In addition, we design a sophisticated data generation pipeline and propose a new dataset, MPSGaze, which contains full images of multiple people with 3D gaze ground truth. Experimental results demonstrate that our unified framework not only offers a faster speed, but also provides a lower gaze estimation error compared with state-of-the-art methods. This technique can be useful in real-time applications with multiple users. Mingfang Zhang 0002, Yunfei Liu 0001, Feng Lu 0005 |
CVPR | 2 |
| 2022 | Reconstructing 3D Virtual Face with Eye Gaze from a Single ImageabstractReconstructing 3D virtual face from a single image has a wide range of applications in virtual reality. Existing approaches synthesize plausible reconstructed virtual faces, however, eye gaze information is usually ignored, which is critical in human-computer interaction. In this paper, we propose to reconstruct 3D virtual face with eye gaze information from a single image. The main challenges lie in two aspects, one is the low reconstruction quality in the eye region, the other one is the lack of an efficient method to obtain precise eye gaze information. To address these problems, we decompose this task into two key steps, i.e., 3D face reconstruction with precise eye region and eye contact guided facial-rotation for eye gaze information. The first step is designed for precise eye region reconstruction through joint optimization on 3D face/eye shapes and textures. The second step consists of two parts: eye contact discriminator and automatic eye contact search algorithm via gradient-based optimization to perform both eye contact and gaze estimation simultaneously. Extensive experiments on different tasks demonstrate the significant gain of the proposed approach, achieving an MSE of (30%), an SSIM of (17.85%), and a PSNR of (8.4%). It also produces lower angular errors (63.01%) in the gaze estimation task compared with human annotations. Jiadong Liang, Yunfei Liu 0001, Feng Lu 0005 |
VR | 2 |
| 2022 | Semantic Guided Single Image Reflection RemovalabstractReflection is common when we see through a glass window, which not only is a visual disturbance but also influences the performance of computer vision algorithms. Removing the reflection from a single image, however, is highly ill-posed since the color at each pixel needs to be separated into two values belonging to the clear background and the reflection, respectively. To solve this, existing methods use additional priors such as reflection layer smoothness, double reflection effect, and color consistency to distinguish the two layers. However, these low-level priors may not be consistently valid in real cases. In this paper, inspired by the fact that human beings can separate the two layers easily by recognizing the objects and understanding the scene, we propose to use the object semantic cue, which is high-level information, as the guidance to help reflection removal. Based on the data analysis, we develop a multi-task end-to-end deep learning method with a semantic guidance component, to solve reflection removal and semantic segmentation jointly. Extensive experiments on different datasets show significant performance gain when using high-level object-oriented information. We also demonstrate the application of our method to other computer vision tasks. Yunfei Liu 0001, Yu Li 0003, Shaodi You, Feng Lu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Separating Content and Style for Unsupervised Image-to-Image Translation
Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
BMVC | 1 |
| 2021 | Generalizing Gaze Estimation with Outlier-guided Collaborative AdaptationabstractDeep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation framework (PnP-GA), which is an ensemble of networks that learn collaboratively with the guidance of outliers. Since our proposed framework does not require ground-truth labels in the target domain, the existing gaze estimation networks can be directly plugged into PnPGA and generalize the algorithms to new domains. We test PnP-GA on four gaze domain adaptation tasks, ETH-to-MPII, ETH-to-EyeDiap, Gaze360-to-MPII, and Gaze360to-EyeDiap. The experimental results demonstrate that the PnP-GA framework achieves considerable performance improvements of 36.9%, 31.6%, 19.4%, and 11.8% over the baseline system. The proposed framework also outperforms the state-of-the-art domain adaptation approaches on gaze domain adaptation tasks. Code has been released at https://github.com/DreamtaleCore/PnP-GA. Yunfei Liu 0001, Ruicong Liu, Haofei Wang 0001, Feng Lu 0005 |
ICCV | 1 |
| 2021 | Edge-Guided Near-Eye Image Analysis for Head Mounted DisplaysabstractEye tracking provides an effective way for interaction in Augmented Reality (AR) Head Mounted Displays (HMDs). Current eye tracking techniques for AR HMDs require eye segmentation and ellipse fitting under near-infrared illumination. However, due to the low contrast between sclera and iris regions and unpredictable reflections, it is still challenging to accomplish accurate iris/pupil segmentation and the corresponding ellipse fitting tasks. In this paper, inspired by the fact that most essential information is encoded in the edge areas, we propose a novel near-eye image analysis method with edge maps as guidance. Specifically, we first utilize an Edge Extraction Network ($E^{2}-$Net) to predict high-quality edge maps, which only contain eyelids and iris/pupil contours without other undesired edges. Then we feed the edge maps into an Edge-Guided Segmentation and Fitting Network (ESF-Net) for accurate segmentation and ellipse fitting. Extensive experimental results demonstrate that our method outperforms current state-of-the-art methods in near-eye image segmentation and ellipse fitting tasks, based on which we present applications of eye tracking with AR HMD. Zhimin Wang 0001, Yunfei Liu 0001, Feng Lu 0005 |
ISMAR | 3 |
| 2020 | Separate in Latent Space: Unsupervised Single Image Layer SeparationabstractMany real world vision tasks, such as reflection removal from a transparent surface and intrinsic image decomposition, can be modeled as single image layer separation. However, this problem is highly ill-posed, requiring accurately aligned and hard to collect triplet data to train the CNN models. To address this problem, this paper proposes an unsupervised method that requires no ground truth data triplet in training. At the core of the method are two assumptions about data distributions in the latent spaces of different layers, based on which a novel unsupervised layer separation pipeline can be derived. Then the method can be constructed based on the GANs framework with self-supervision and cycle consistency constraints, etc. Experimental results demonstrate its successfulness in outperforming existing unsupervised methods in both synthetic and real world tasks. The method also shows its ability to solve a more challenging multi-layer separation task. Yunfei Liu 0001, Feng Lu 0005 |
AAAI | 1 |
| 2020 | Unsupervised Learning for Intrinsic Image Decomposition From a Single ImageabstractIntrinsic image decomposition, which is an essential task in computer vision, aims to infer the reflectance and shading of the scene. It is challenging since it needs to separate one image into two components. To tackle this, conventional methods introduce various priors to constrain the solution, yet with limited performance. Meanwhile, the problem is typically solved by supervised learning methods, which is actually not an ideal solution since obtaining ground truth reflectance and shading for massive general natural scenes is challenging and even impossible. In this paper, we propose a novel unsupervised intrinsic image decomposition framework, which relies on neither labeled training data nor hand-crafted priors. Instead, it directly learns the latent feature of reflectance and shading from unsupervised and uncorrelated data. To enable this, we explore the independence between reflectance and shading, the domain invariant content constraint and the physical constraint. Extensive experiments on both synthetic and real image datasets demonstrate consistently superior performance of the proposed method. Yunfei Liu 0001, Yu Li 0003, Shaodi You, Feng Lu 0005 |
CVPR | 1 |
| 2020 | Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
Yunfei Liu 0001, Xingjun Ma, James Bailey 0001, Feng Lu 0005 |
ECCV (10) | 1 |
| 2020 | Adaptive Feature Fusion Network for Gaze Tracking in Mobile TabletsabstractRecently, many multi-stream gaze estimation methods have been proposed. They estimate gaze from eye and face appearances and achieve reasonable accuracy. However, most of the methods simply concatenate the features extracted from eye and face appearance. The feature fusion process has been ignored. In this paper, we propose a novel Adaptive Feature Fusion Network (AFF-Net), which performs gaze tracking task in mobile tablets. We stack two-eye feature maps and utilize Squeeze-and-Excitation layers to adaptively fuse two-eye features according to their similarity on appearance. Meanwhile, we also propose Adaptive Group Normalization to recalibrate eye features with the guidance of facial feature. Extensive experiments on both GazeCapture and MPIIFaceGaze datasets demonstrate consistently superior performance of the proposed method. Yiwei Bao, Yihua Cheng, Yunfei Liu 0001, Feng Lu 0005 |
ICPR | 3 |
| 2019 | What I See Is What You See: Joint Attention Learning for First and Third Person Video Co-analysisabstractIn recent years, more and more videos are captured from the first-person viewpoint by wearable cameras. Such first-person video provides additional information besides the traditional third-person video, and thus has a wide range of applications. However, techniques for analyzing the first-person video can be fundamentally different from those for the third-person video, and it is even more difficult to explore the shared information from both viewpoints. In this paper, we propose a novel method for first- and third-person video co-analysis. At the core of our method is the notion of "joint attention'', indicating the learnable representation that corresponds to the shared attention regions in different viewpoints and thus links the two viewpoints. To this end, we develop a multi-branch deep network with a triplet loss to extract the joint attention from the first- and third-person videos via self-supervised learning. We evaluate our method on the public dataset with cross-viewpoint video matching tasks. Our method outperforms the state-of-the-art both qualitatively and quantitatively. We also demonstrate how the learned joint attention can benefit various applications through a set of additional experiments. Huangyue Yu, Minjie Cai, Yunfei Liu 0001, Feng Lu 0005 |
ACM Multimedia | 3 |