VLDB 2026 Research / reviewers in the wild / expert
Feng Lu 0005
dblp:97/6483-5
· DBLP profile ↗
106ranked-venue papers
16as first author
61since 2021 · last 2026
0000-0001-9064-7964ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 77 · 10 first-author · 40 since 2021Artificial intelligence and machine learning · 62 · 11 first-author · 32 since 2021Human-computer interaction and ubiquitous computing · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy UnfoldingabstractLearning procedural-aware video representations is a key step towards building agents that can reason about and execute complex tasks. Existing methods typically address this problem by aligning visual content with textual descriptions at the task and step levels to inject procedural semantics into video representations. However, due to their high level of abstraction, "task" and "step" descriptions fail to form a robust alignment with the concrete, observable details in visual data. To address this, we introduce "states", i.e., textual snapshots of object configurations, as a visually-grounded semantic layer that anchors abstract procedures to what a model can actually see. We formalize this insight in a novel Task-Step-State (TSS) framework, where tasks are achieved via steps that drive transitions between observable states. To enforce this structure, we propose a progressive pre-training strategy that unfolds the TSS hierarchy, forcing the model to first ground representations in states before associating them with steps and, ultimately, high-level tasks. Extensive experiments on the COIN and CrossTask datasets show that our method outperforms baseline models on multiple downstream tasks, including task recognition, step recognition, and next step prediction. Ablation studies show that introducing state supervision is a key driver of performance gains across all tasks. Additionally, our progressive pretraining strategy proves more effective than standard joint training, as it better enforces the intended hierarchical structure. Jinghan Zhao, Yifei Huang 0002, Feng Lu 0005 |
AAAI | 3 |
| 2026 | Towards cobodied/symbodied AI: concept and eight scientific and technical problems
Feng Lu 0005, Qinping Zhao |
Sci. China Inf. Sci. | 1 |
| 2026 | Assessing Flow State in Virtual Reality: A Multi-Channel Physiological Framework With Self-Supervised Pre-TrainingabstractFlow, a state of complete immersion, focus, and enjoyment, has important implications for learning, productivity, and well-being. While research on flow is growing, there remains a need for refined methods to elicit and assess flow states specifically in immersive VR environments. In this work, we address this need by developingBeat Flow, a VR music rhythm game tailored to participants' skill levels, designed to evoke distinct flow states. Instead of using traditional physiological apparatus that limits mobility, we employ lightweight wearable devices to capture electroencephalogram (EEG), galvanic skin response (GSR), photoplethysmography (PPG), and eye blink. Furthermore, we propose a VR-native self-report tool, the 3D Flow Scale (3FS), for efficient flow assessment in VR environments. Besides, we develop a deep learning model with a self-supervised pre-training strategy to classify flow states, achieving 78.4% accuracy within a 10-second time window using 5-fold cross-validation. The model demonstrates state-of-the-art performance, especially in cross-user scenarios, advancing flow detection in VR and paving the way for flow-aware VR applications across various domains. Bo Liu 0112, Jinghan Zhao, Feng Lu 0005 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | SIAgent: Spatial Interaction Agent via LLM-Powered Eye-Hand Motion Intent Understanding in VRabstractEye-hand coordinated interaction is becoming a mainstream interaction modality in Virtual Reality (VR) user interfaces. Current paradigms for this multimodal interaction require users to learn predefined gestures and memorize multiple gesture-task associations, which can be summarized as an "Operation-to-Intent" paradigm. This paradigm increases users' learning costs and has low interaction error tolerance. In this paper, we propose SIAgent, a novel "Intent-to-Operation" framework allowing users to express interaction intents through natural eye-hand motions based on common sense and habits. Our system features two main components: (1) intent recognition that translates spatial interaction data into natural language and infers user intent, and (2) agent-based execution that generates an agent to execute corresponding tasks. This eliminates the need for gesture memorization and accommodates individual motion preferences with high error tolerance. We conduct two user studies across over 60 interaction tasks, comparing our method with two "Operation-to-Intent" techniques. Results show our method achieves higher intent recognition accuracy than gaze + pinch interaction (97.2% versus 93.1%) while reducing arm fatigue and improving usability, and user preference. Another study verifies the function of eye gaze and hand motion channels in intent recognition. Our work offers valuable insights into enhancing VR interaction intelligence through intent-driven design. Zhimin Wang 0001, Chenyu Gu, Feng Lu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | From embodied AI to cobodied AI: Foundations and frontiersabstractThis paper presents a systematic review and technical revisit of recent developments in embodied AI. We comprehensively synthesize current dominant paradigms, representative architectures, and pivotal advancements around five core modules: (1) perception and understanding, (2) reasoning and decision making, (3) control and action, (4) modeling and learning (e.g., VLA (vision-language-action) and WM (world models)), and (5) data and simulation. This analysis establishes a structured and cutting-edge panoramic view of embodied AI technology. Building on this foundation, we explore the transition from embodied AI to a new paradigm of deep human-AI collaboration. Anchored in the original concept of cobodied AI, we adopt a human-centered technical perspective to systematically investigate this paradigm shift. Specifically, we discuss breakthroughs in critical dimensions such as perceptual alignment, collaborative decision-making, action guidance, and bidirectional mutual learning. These efforts aim to realize three core characteristics: human-centered egocentric grounding, dual-mode cognitive integration, and physical co-embodiment. To our knowledge, this work constitutes the first systematic construction of a technical framework and implementation roadmap for realizing cobodied AI, building upon advances in embodied AI but fundamentally reorienting intelligence around the human body and intent, providing a foundational reference for future research in this domain. Feng Lu 0005, Ruijia Pang, Bosong Qi, Qinping Zhao |
Virtual Real. Intell. Hardw. | 1 |
| 2025 | GazeSwipe: Enhancing Mobile Touchscreen Reachability through Seamless Gaze and Finger-Swipe IntegrationabstractSmartphones with large screens provide users with increased display and interaction space but pose challenges in reaching certain areas with the thumb when using the device with one hand. To address this, we introduce GazeSwipe, a multimodal interaction technique that combines eye gaze with finger-swipe gestures, enabling intuitive and low-friction reach on mobile touchscreens. Specifically, we design a gaze estimation method that eliminates the need for explicit gaze calibration. Our approach also avoids the use of additional eye-tracking hardware by leveraging the smartphone's built-in front-facing camera. Considering the potential decrease in gaze accuracy without dedicated eye trackers, we use finger-swipe gestures to compensate for any inaccuracies in gaze estimation. Additionally, we introduce a user-unaware auto-calibration method that improves gaze accuracy during interaction. Through extensive experiments on smartphones and tablets, we compare our technique with various methods for touchscreen reachability and evaluate the performance of our auto-calibration strategy. The results demonstrate that our method achieves high success rates and is preferred by users. The findings also validate the effectiveness of the auto-calibration strategy. Zhuojiang Cai, Jingkai Hong, Zhimin Wang 0001, Feng Lu 0005 |
CHI | 4 |
| 2025 | GazeGene: Large-scale Synthetic Gaze Dataset with 3D Eyeball AnnotationsabstractThanks to the introduction of large-scale datasets, deep-learning has become the mainstream approach for appearance-based gaze estimation problems. However, current large-scale datasets contain annotation errors and provide only a single vector for gaze annotation, lacking key information such as 3D eyeball structures. Limitations in annotation accuracy and variety have constrained the progress in research and development of deep-learning methods for appearance-based gaze-related tasks. In this paper, we present GazeGene, a new large-scale synthetic gaze dataset with photo-realistic samples. More importantly, GazeGene not only provides accurate gaze annotations, but also offers 3D annotations of vital eye structures such as the pupil, iris, eyeball, optical and visual axes for the first time. Experiments show that GazeGene achieves comparable quality and generalization ability with real-world datasets, even outperforms most existing datasets on high-resolution images. Furthermore, its 3D eyeball annotations expand the application of deep-learning methods on various gaze-related tasks, offering new insights into this field. The dataset is available at: https://phiai.buaa.edu.cn/GazeGene/. Yiwei Bao, Feng Lu 0005 |
CVPR | 3 |
| 2025 | 3D Prior Is All You Need: Cross-Task Few-shot 2D Gaze Estimationabstract3D and 2D gaze estimation share the fundamental objective of capturing eye movements but are traditionally treated as two distinct research domains. In this paper, we introduce a novel cross-task few-shot 2D gaze estimation approach, aiming to adapt a pre-trained 3D gaze estimation network for 2D gaze prediction on unseen devices using only a few training images. This task is highly challenging due to the domain gap between 3D and 2D gaze, unknown screen poses, and limited training data. To address these challenges, we propose a novel framework that bridges the gap between 3D and 2D gaze. Our framework contains a physics-based differentiable projection module with learnable parameters to model screen poses and project 3D gaze into 2D gaze. The framework is fully differentiable and can integrate into existing 3D gaze networks without modifying their original architecture. Additionally, we introduce a dynamic pseudo-labelling strategy for flipped images, which is particularly challenging for 2D labels due to unknown screen poses. To overcome this, we reverse the projection process by converting 2D labels to 3D space, where flipping is performed. Notably, this 3D space is not aligned with the camera coordinate system, so we learn a dynamic transformation matrix to compensate for this misalignment. We evaluate our method on MPIIGaze, EVE, and GazeCapture datasets, collected respectively on laptops, desktop computers, and mobile devices. The superior performance highlights the effectiveness of our approach, and demonstrates its strong potential for real-world applications. Yihua Cheng, Hengfei Wang, Zhongqun Zhang, Boeun Kim, Feng Lu 0005, Hyung Jin Chang |
CVPR | 6 |
| 2025 | Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation
Jiadong Liang, Feng Lu 0005 |
ICXR | 2 |
| 2025 | 3DPE-Gaze: Unlocking the Potential of 3D Facial Priors for Generalized Gaze EstimationabstractIn recent years, face-based deep-learning gaze estimation methods have achieved significant advancements. However, while face images provide supplementary information beneficial for gaze inference, the substantial extraneous information they contain also increases the risk of overfitting during model training and compromises generalization capability. To alleviate this problem, we propose the 3DPE-Gaze framework, explicitly modeling 3D facial priors for feature decoupling and generalized gaze estimation. The 3DPE-Gaze framework consists of two core modules: the 3D Geometric Prior Module (3DGP) incorporating the FLAME model to parameterize facial structures and gaze-irrelevant facial appearances while extracting gaze features; the Semantic Concept Alignment Module (SCAM) separates gaze-related and unrelated concepts through CLIP-guided contrastive learning. Finally, the 3DPE-Gaze framework combines 3D facial landmark as prior for generalized gaze estimation. Experimental results show that 3DPE-Gaze outperforms existing state-of-the-art methods on four major cross-domain tasks, with particularly outstanding performance in challenging scenarios such as lighting variations, extreme head poses, and glasses occlusion. Yangshi Ge, Yiwei Bao, Feng Lu 0005 |
NeurIPS | 3 |
| 2025 | Gam360: sensing gaze activities of multi-persons in 360 degrees
Zhuojiang Cai, Haofei Wang 0001, Yuhao Niu, Feng Lu 0005 |
CCF Trans. Pervasive Comput. Interact. | 4 |
| 2025 | From Gaze Jitter to Domain Adaptation: Generalizing Gaze Estimation by Manipulating High-Frequency Components
Ruicong Liu, Haofei Wang 0001, Feng Lu 0005 |
Int. J. Comput. Vis. | 3 |
| 2025 | Polarization State Attention Dehazing Network With a Simulated Polar-Haze DatasetabstractImage dehazing under harsh weather conditions remains a challenging and ill-posed problem. In addition, acquiring real-time haze-free counterparts of hazy images poses difficulties. Existing approaches commonly synthesize hazy data by relying on estimated depth information, which is prone to errors due to its physical unreliability. While generative networks can transfer some hazy features to clear images, the resulting hazy images still exhibit an artificial appearance. In this paper, we introduce polarization cues to propose a haze simulation strategy to synthesize hazy data, ensuring visually pleasing results that adhere to physical laws. Leveraging on the simulated Polar-Haze dataset, we present a polarization state attention dehazing network (PSADNet), which consists of a polarization extraction module and a polarization dehazing module. The proposed polarization extraction model incorporates an attention mechanism to capture high-level image features related to polarization and chromaticity. The polarization dehazing module utilizes these features derived from the polarization analysis to enhance image dehazing capabilities while preserving the accuracy of the polarization information. Promising results are observed in both qualitative and quantitative experiments, supporting the effectiveness of the proposed PSADNet and the validity of polarization-based haze simulation strategy. Sijia Wen, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Multim. | 3 |
| 2025 | Dominant-Eye-Aware Asymmetric Foveated Rendering for Virtual RealityabstractTo address the increasing computational demands of high-resolution virtual reality headsets, foveated rendering reduces pixel sampling in the peripheral regions of the visual field. However, existing methods have not fully leveraged binocular vision, particularly the dominant eye theory. In our prior work, we proposed Dominant-Eye-Aware foveated rendering optimized with Multi-Parameter foveation (DEAMP), which divided each eye's visual field into fixed eccentricity layers ([0, 10]$^{\circ }$∘, [10, 22.5]$^{\circ }$∘, [22.5, 45]$^{\circ }$∘). The non-dominant eye received greater foveation within the same layers compared to the dominant eye. We further argue that the eccentricity ranges should vary between the eyes due to inter-eye, individual, and scene-specific differences. In this article, we introduce an enhanced method, Dominant-Eye-aware Asymmetric Foveated Rendering (DEA-FoR). This treats eccentricity as a new variable, allowing users to select eccentricity sets tailored to their eyes and supporting asymmetric configurations between the two eyes. Experimental results demonstrate significant improvements in rendering speed over our previous method while maintaining perceptual quality. Additionally, we found that individual differences and scene texture complexity significantly influence the eccentricity settings. This work offers new insights into perceptual differences in binocular vision and contributes to optimizing virtual reality experiences. Zhimin Wang 0001, Xiangyuan Gu, Feng Lu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | UVAGaze: Unsupervised 1-to-2 Views Adaptation for Gaze EstimationabstractGaze estimation has become a subject of growing interest in recent research. Most of the current methods rely on single-view facial images as input. Yet, it is hard for these approaches to handle large head angles, leading to potential inaccuracies in the estimation. To address this issue, adding a second-view camera can help better capture eye appearance. However, existing multi-view methods have two limitations. 1) They require multi-view annotations for training, which are expensive. 2) More importantly, during testing, the exact positions of the multiple cameras must be known and match those used in training, which limits the application scenario. To address these challenges, we propose a novel 1-view-to-2-views (1-to-2 views) adaptation solution in this paper, the Unsupervised 1-to-2 Views Adaptation framework for Gaze estimation (UVAGaze). Our method adapts a traditional single-view gaze estimator for flexibly placed dual cameras. Here, the "flexibly" means we place the dual cameras in arbitrary places regardless of the training data, without knowing their extrinsic parameters. Specifically, the UVAGaze builds a dual-view mutual supervision adaptation strategy, which takes advantage of the intrinsic consistency of gaze directions between both views. In this way, our method can not only benefit from common single-view pre-training, but also achieve more advanced dual-view gaze estimation. The experimental results show that a single-view estimator, when adapted for dual views, can achieve much higher accuracy, especially in cross-dataset settings, with a substantial improvement of 47.0%. Project page: https://github.com/MickeyLLG/UVAGaze. Ruicong Liu, Feng Lu 0005 |
AAAI | 2 |
| 2024 | Gaze from Origin: Learning for Generalized Gaze Estimation by Embedding the Gaze Frontalization ProcessabstractGaze estimation aims to accurately estimate the direction or position at which a person is looking. With the development of deep learning techniques, a number of gaze estimation methods have been proposed and achieved state-of-the-art performance. However, these methods are limited to within-dataset settings, whose performance drops when tested on unseen datasets. We argue that this is caused by infinite and continuous gaze labels. To alleviate this problem, we propose using gaze frontalization as an auxiliary task to constrain gaze estimation. Based on this, we propose a novel gaze domain generalization framework named Gaze Frontalization-based Auxiliary Learning (GFAL) Framework which embeds the gaze frontalization process, i.e., guiding the feature so that the eyeball can rotate and look at the front (camera), without any target domain information during training. Experimental results show that our proposed framework is able to achieve state-of-the-art performance on gaze domain generalization task, which is competitive with or even superior to the SOTA gaze unsupervised domain adaptation methods. Mingjie Xu, Feng Lu 0005 |
AAAI | 2 |
| 2024 | Gaze Target Detection by Merging Human Attention and Activity CuesabstractDespite achieving impressive performance, current methods for detecting gaze targets, which depend on visual saliency and spatial scene geometry, continue to face challenges when it comes to detecting gaze targets within intricate image backgrounds. One of the primary reasons for this lies in the oversight of the intricate connection between human attention and activity cues. In this study, we introduce an innovative approach that amalgamates the visual saliency detection with the body-part & object interaction both guided by the soft gaze attention. This fusion enables precise and dependable detection of gaze targets amidst intricate image backgrounds. Our approach attains state-of-the-art performance on both the Gazefollow benchmark and the GazeVideoAttn benchmark. In comparison to recent methods that rely on intricate 3D reconstruction of a single input image, our approach, which solely leverages 2D image information, still exhibits a substantial lead across all evaluation metrics, positioning it closer to human-level performance. These outcomes underscore the potent effectiveness of our proposed method in the gaze target detection task. Yaokun Yang, Yihan Yin, Feng Lu 0005 |
AAAI | 3 |
| 2024 | From Feature to Gaze: A Generalizable Replacement of Linear Layer for Gaze EstimationabstractDeep-Learning-based gaze estimation approaches often suffer from notable performance degradation in unseen target domains. One of the primary reasons is that the Fully Connected layer is highly prone to overfitting when mapping the high-dimensional image feature to 3D gaze. In this paper, we propose Analytical Gaze Generalization framework (AGG) to improve the generalization ability of gaze estimation models without touching target domain data. The AGG consists of two modules, the Geodesic Projection Module (GPM) and the Sphere-Oriented Training (SOT). GPM is a generalizable replacement of FC layer, which projects high-dimensional image features to 3D space analytically to extract the principle components of gaze. Then, we propose Sphere-Oriented Training (SOT) to incorporate the GPM into the training process and further improve cross-domain performances. Experimental results demonstrate that the AGG effectively alleviate the overfitting problem and consistently improves the cross-domain gaze estimation accuracy in 12 cross-domain settings, without requiring any target domain data. The insight from the Analytical Gaze Generalization framework has the potential to benefit other regression tasks with physical meanings. Yiwei Bao, Feng Lu 0005 |
CVPR | 2 |
| 2024 | Unsupervised Gaze Representation Learning from Multi-view Face ImagesabstractAnnotating gaze is an expensive and time-consuming endeavor, requiring costly eye-trackers or complex geometric calibration procedures. Although some eye-based unsupervised gaze representation learning methods have been proposed, the quality of gaze representation extracted by these methods degrades severely when the head pose is large. In this paper, we present the Multi-View Dual-Encoder (MVDE), a framework designed to learn gaze representations from unlabeled multi-view face images. Through the proposed Dual-Encoder architecture and the multi-view gaze representation swapping strategy, the MV-DE successfully disentangles gaze from general facial information and derives gaze representations closely tied to the subject's eye-ball rotation without gaze label. Experimental results illustrate that the gaze representations learned by the MV-DE can be used in downstream tasks, including gaze estimation and redirection. Gaze estimation results indicates that the proposed MV-DE displays notably higher robustness to uncontrolled head movements when compared to state-of-the-art (SOTA) unsupervised learning methods. Yiwei Bao, Feng Lu 0005 |
CVPR | 2 |
| 2024 | Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang 0001, Nie Lin, Zhishan Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Feng Lu 0005, Karim Abou Zeid, Bastian Leibe, Jeongwan On, Seungryul Baek, Saurabh Gupta 0001, Yoichi Sato 0001, Otmar Hilliges, Hyung Jin Chang, Angela Yao |
ECCV (25) | 13 |
| 2024 | De-confounded Gaze Estimation
Ziyang Liang, Yiwei Bao, Feng Lu 0005 |
ECCV (23) | 3 |
| 2024 | Gaze Target Detection Based on Head-Local-Global Coordination
Yaokun Yang, Feng Lu 0005 |
ECCV (27) | 2 |
| 2024 | Enhancing Text Entry in Mixed Reality with Tangible Feedback
Haofei Wang 0001, Feng Lu 0005 |
ICXR | 3 |
| 2024 | GazeRing: Enhancing Hand-Eye Coordination with Pressure Ring in Augmented RealityabstractHand-eye coordination techniques find widespread utility in augmented reality and virtual reality headsets, as they retain the speed and intuitiveness of eye gaze while leveraging the precision of hand gestures. However, in contrast to obvious interactive gestures, users prefer less noticeable interactions in public settings due to concerns about social acceptance. To address this, we propose GazeRing, a multimodal interaction technique that combines eye gaze with a smart ring, enabling private and subtle hand-eye coordination while allowing users’ hands complete freedom of movement. Specifically, we design a pressure-sensitive ring that supports sliding interactions in eight directions to facilitate efficient 3D object manipulation. Additionally, we introduce two control modes for the ring: finger-tap and finger-slide, to accommodate diverse usage scenarios. Through user studies involving object selection and translation tasks under two eye-tracking accuracy conditions, with two degrees of occlusion, GazeRing demonstrates significant advantages over existing techniques that do not require obvious hand gestures (e.g., gaze-only and gaze-speech interactions). Our GazeRing technique achieves private and subtle interactions, potentially improving the user experience in public settings. A demo video can be found at zhimin-wang.github.io/GazeRing.html. Zhimin Wang 0001, Mingwei Hu, Maohang Rao, Feng Lu 0005 |
ISMAR | 6 |
| 2024 | SRMAE: Masked Image Modeling for Scale-Invariant Deep Representations
Lin Gu 0003, Feng Lu 0005 |
PRCV (2) | 3 |
| 2024 | A Cross-Consistency Strategy for Clearer Perception in Low-Light Haze
Sijia Wen, Chaoqun Zhuang, Yunfei Liu 0001, Feng Lu 0005 |
PRCV (8) | 4 |
| 2024 | Matching Compound Prototypes for Few-Shot Action RecognitionabstractAbstract The task of few-shot action recognition aims to recognize novel action classes using only a small number of labeled training samples. How to better describe the action in each video and how to compare the similarity between videos are two of the most critical factors in this task. Directly describing the video globally or by its individual frames cannot well represent the spatiotemporal dependencies within an action. On the other hand, naively matching the global representations of two videos is also not optimal since action can happen at different locations in a video with different speeds. In this work, we propose a novel approach that describes each video using multiple types of prototypes and then computes the video similarity with a particular matching strategy for each type of prototypes. To better model the spatiotemporal dependency, we describe the video by generating prototypes that model the multi-level spatiotemporal relations via transformers. There are a total of three types of prototypes. The first type of prototypes are trained to describe specific aspects of the action in the video e.g., the start of the action, regardless of its timestamp. These prototypes are directly matched one-to-one between two videos to compare their similarity. The second type of prototypes are the timestamp-centered prototypes that are trained to focus on specific timestamps of the video. To deal with the temporal variation of actions in a video, we apply bipartite matching to allow the matching of prototypes of different timestamps. The third type of prototypes are generated from the timestamp-centered prototypes, which regularize their temporal consistency while serving as an auxiliary summarization of the whole video. Experiments demonstrate that our proposed method achieves state-of-the-art results on multiple benchmarks. Yifei Huang 0002, Lijin Yang, Guo Chen 0006, Hongjie Zhang 0002, Feng Lu 0005, Yoichi Sato 0001 |
Int. J. Comput. Vis. | 5 |
| 2024 | Appearance-Based Gaze Estimation With Deep Learning: A Review and BenchmarkabstractHuman gaze provides valuable information on human focus and intentions, making it a crucial area of research. Recently, deep learning has revolutionized appearance-based gaze estimation. However, due to the unique features of gaze estimation research, such as the unfair comparison between 2D gaze positions and 3D gaze vectors and the different pre-processing and post-processing methods, there is a lack of a definitive guideline for developing deep learning-based gaze estimation algorithms. In this paper, we present a systematic review of the appearance-based gaze estimation methods using deep learning. First, we survey the existing gaze estimation algorithms along the typical gaze estimation pipeline: deep feature extraction, deep learning model design, personal calibration and platforms. Second, to fairly compare the performance of different approaches, we summarize the data pre-processing and post-processing methods, including face/eye detection, data rectification, 2D/3D gaze conversion and gaze origin conversion. Finally, we set up a comprehensive benchmark for deep learning-based gaze estimation. We characterize all the public datasets and provide the source code of typical gaze estimation algorithms. This paper serves not only as a reference to develop deep learning-based gaze estimation methods, but also a guideline for future gaze estimation research. Yihua Cheng, Haofei Wang 0001, Yiwei Bao, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | PnP-GA+: Plug-and-Play Domain Adaptation for Gaze Estimation Using Model VariantsabstractAppearance-based gaze estimation has garnered increasing attention in recent years. However, deep learning-based gaze estimation models still suffer from suboptimal performance when deployed in new domains, e.g., unseen environments or individuals. In our previous work, we took this challenge for the first time by introducing a plug-and-play method (PnP-GA) to adapt the gaze estimation model to new domains. The core concept of PnP-GA is to leverage the diversity brought by a group of model variants to enhance the adaptability to diverse environments. In this article, we propose the PnP-GA+ by extending our approach to explore the impact of assembling model variants using three additional perspectives: color space, data augmentation, and model structure. Moreover, we propose an intra-group attention module that dynamically optimizes pseudo-labeling during adaptation. Experimental results demonstrate that by directly plugging several existing gaze estimation networks into the PnP-GA+ framework, it outperforms state-of-the-art domain adaptation approaches on four standard gaze domain adaptation tasks on public datasets. Our method consistently enhances cross-domain performance, and its versatility is improved through various ways of assembling the model group. Ruicong Liu, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | cbPPGGAN: A Generic Enhancement Framework for Unpaired Pulse Waveforms in Camera-Based PhotoplethysmographyabstractCamera-based photoplethysmography (cbP PG) is a non-contact technique that measures cardiac-related blood volume alterations in skin surface vessels through the analysis of facial videos. While traditional approaches can estimate heart rate (HR) under different illuminations, their accuracy can be affected by motion artifacts, leading to poor waveform fidelity and hindering further analysis of heart rate variability (HRV); deep learning-based approaches reconstruct high-quality pulse waveform, yet their performance significantly degrades under illumination variations. In this work, we aim to leverage the strength of these two methods and propose a framework that possesses favorable generalization capabilities while maintaining waveform fidelity. For this purpose, we propose the cbPPGGAN, an enhancement framework for cbPPG that enables the flexible incorporation of both unpaired and paired data sources in the training process. Based on the waveforms extracted by traditional approaches, the cbPPGGAN reconstructs high-quality waveforms that enable accurate HR estimation and HRV analysis. In addition, to address the lack of paired training data problems in real-world applications, we propose a cycle consistency loss that guarantees the time-frequency consistency before/after mapping. The method enhances the waveform quality of traditional POS approaches in different illumination tests (BH-rPPG) and cross-datasets (UBFC-rPPG) with mean absolute error (MAE) values of 1.34 bpm and 1.65 bpm, and average beat-to-beat (AVBB) values of 27.46 ms and 45.28 ms, respectively. Experimental results demonstrate that the cbPPGGAN enhances cbPPG signal quality and outperforms the state-of-the-art approaches in HR estimation and HRV analysis. The proposed framework opens a new pathway toward accurate HR estimation in an unconstrained environment. Haofei Wang 0001, Bo Liu 0112, Feng Lu 0005 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Real-Time Gaze Tracking via Head-Eye Cues on Head Mounted DevicesabstractGaze is a crucial element in human-computer interaction and plays an increasingly vital role in promoting the adoption of head-mounted devices (HMDs). Existing gaze tracking methods for HMDs either demand user calibration or face challenges in balancing accuracy and speed, compromising the overall user experience. In this paper, we introduce a novel strategy for real-time, calibration-free gaze tracking using joint head-eye cues on HMDs. Initially, we create a multimodal gaze tracking dataset named HE-Gaze, encompassing synchronized eye images and 6DoF head movement data, addressing a gap in the current data landscape. Statistical analyses unveil the correlation between head movements and gaze positions. Building on these insights, we introduce the hierarchical head-eye coordinated gaze tracking model (HHE-Tracker), which incorporates two lightweight branches to encode input eye images and head sequences efficiently. It combines encoded head velocity and posture features with eye features across various scales to infer gaze position. HHE-Tracker was implemented on a commercial HMD, and its performance was assessed in unconstrained scenarios. The results demonstrate the HHE-Tracker's capability to accurately estimate gaze positions in real-time. In comparison to the state-of-the-art gaze tracking algorithm, HHE-Tracker exhibits commendable accuracy (3.47$^{\circ }$) and a 40-fold speedup (81FPSon a Snapdragon 845 SoC). Yingxi Li, Xiaowei Bai, Liang Xie 0012, Feng Lu 0005, Feitian Zhang, Ye Yan 0001, Erwei Yin |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Robust Dual-Modal Speech Keyword Spotting for XR HeadsetsabstractWhile speech interaction finds widespread utility within the Extended Reality (XR) domain, conventional vocal speech keyword spotting systems continue to grapple with formidable challenges, including suboptimal performance in noisy environments, impracticality in situations requiring silence, and susceptibility to inadvertent activations when others speak nearby. These challenges, however, can potentially be surmounted through the cost-effective fusion of voice and lip movement information. Consequently, we propose a novel vocal-echoic dual-modal keyword spotting system designed for XR headsets. We devise two different modal fusion approches and conduct experiments to test the system's performance across diverse scenarios. The results show that our dual-modal system not only consistently outperforms its single-modal counterparts, demonstrating higher precision in both typical and noisy environments, but also excels in accurately identifying silent utterances. Furthermore, we have successfully applied the system in real-time demonstrations, achieving promising results. The code is available at https://github.com/caizhuojiang/VE-KWS. Zhuojiang Cai, Feng Lu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Tasks Reflected in the Eyes: Egocentric Gaze-Aware Visual Task Type Recognition in Virtual RealityabstractWith eye tracking finding widespread utility in augmented reality and virtual reality headsets, eye gaze has the potential to recognize users' visual tasks and adaptively adjust virtual content displays, thereby enhancing the intelligence of these headsets. However, current studies on visual task recognition often focus on scene-specific tasks, like copying tasks for office environments, which lack applicability to new scenarios, e.g., museums. In this paper, we propose four scene-agnostic task types for facilitating task type recognition across a broader range of scenarios. We present a new dataset that includes eye and head movement data recorded from 20 participants while they engaged in four task types across 15 360-degree VR videos. Using this dataset, we propose an egocentric gaze-aware task type recognition method, TRCLP, which achieves promising results. Additionally, we illustrate the practical applications of task type recognition with three examples. Our work offers valuable insights for content developers in designing task-aware intelligent applications. Our dataset and source code are available at zhimin-wang.github.io/TaskTypeRecognition.html. Zhimin Wang 0001, Feng Lu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | Learning a Generalized Gaze Estimator from Gaze-Consistent FeatureabstractGaze estimator computes the gaze direction based on face images. Most existing gaze estimation methods perform well under within-dataset settings, but can not generalize to unseen domains. In particular, the ground-truth labels in unseen domain are often unavailable. In this paper, we propose a new domain generalization method based on gaze-consistent features. Our idea is to consider the gaze-irrelevant factors as unfavorable interference and disturb the training data against them, so that the model cannot fit to these gaze-irrelevant factors, instead, only fits to the gaze-consistent features. To this end, we first disturb the training data via adversarial attack or data augmentation based on the gaze-irrelevant factors, i.e., identity, expression, illumination and tone. Then we extract the gaze-consistent features by aligning the gaze features from disturbed data with non-disturbed gaze features. Experimental results show that our proposed method achieves state-of-the-art performance on gaze domain generalization task. Furthermore, our proposed method also improves domain adaption performance on gaze estimation. Our work provides new insight on gaze domain generalization task. Mingjie Xu, Haofei Wang 0001, Feng Lu 0005 |
AAAI | 3 |
| 2023 | SynBlink and BlinkFormer: A Synthetic Dataset and Transformer-Based Method for Video Blink Detection
Bo Liu 0112, Feng Lu 0005 |
BMVC | 3 |
| 2023 | Joint Low-light Enhancement and Super Resolution with Image Underexposure Level Guidance
Mingjie Xu, Chaoqun Zhuang, Feifan Lv, Feng Lu 0005 |
BMVC | 4 |
| 2023 | DVGaze: Dual-View Gaze EstimationabstractGaze estimation methods estimate gaze from facial appearance with a single camera. However, due to the limited view of a single camera, the captured facial appearance cannot provide complete facial information and thus complicate the gaze estimation problem. Recently, camera devices are rapidly updated. Dual cameras are affordable for users and have been integrated in many devices. This development suggests that we can further improve gaze estimation performance with dual-view gaze estimation. In this paper, we propose a dual-view gaze estimation network (DV-Gaze). DV-Gaze estimates dual-view gaze directions from a pair of images. We first propose a dual-view interactive convolution (DIC) block in DV-Gaze. DIC blocks exchange dual-view information during convolution in multiple feature scales. It fuses dual-view features along epipolar lines and compensates for the original feature with the fused feature. We further propose a dual-view transformer to estimate gaze from dual-view features. Camera poses are encoded to indicate the position information in the transformer. We also consider the geometric relation between dual-view gaze directions and propose a dual-view gaze consistency loss for DV-Gaze. DV-Gaze achieves state-of-the-art performance on ETH-XGaze and EVE datasets. Our experiments also prove the potential of dual-view gaze estimation. We release codes in https://github.com/yihuacheng/DVGaze. Yihua Cheng, Feng Lu 0005 |
ICCV | 2 |
| 2023 | DEAMP: Dominant-Eye-Aware Foveated Rendering with Multi-Parameter optimizationabstractThe increasing use of high-resolution displays and the demand for interactive frame rates presents a major challenge to widespread adoption of virtual reality. Foveated rendering address this issue by lowering pixel sampling rate at the periphery of the display. How-ever, existing techniques do not fully exploit the feature of human binocular vision, i.e., the dominant eye. In this paper, we propose a Dominant-Eye-Aware foveated rendering method optimized with Multi-Parameter foveation (DEAMP). Specifically, we control the level of foveation for both eyes with two distinct sets of foveation parameters. To achieve this, each eye’s visual field is divided into three nested layers based on eccentricity. Multiple parameters govern the level of foveation of each layer, respectively. We conduct user studies to evaluate our method. Experimental results demonstrate that DEAMP is superior in terms of rendering time and reduces the disparity between pixel sampling rate and the visual acuity fall-off model while maintaining the perceptual quality. Zhimin Wang 0001, Xiangyuan Gu, Feng Lu 0005 |
ISMAR | 3 |
| 2023 | Exploring 3D Interaction with Gaze Guidance in Augmented RealityabstractRecent research based on hand-eye coordination has shown that gaze could improve object selection and translation experience under certain scenarios in AR. However, several limitations still exist. Specifically, we investigate whether gaze could help object selection with heavy 3D occlusions and help 3D object translation in the depth dimension. In addition, we also investigate the possibility of reducing the gaze calibration burden before use. Therefore, we develop new methods with proper gaze guidance for 3D interaction in AR, and also an implicit online calibration method. We conduct two user studies to evaluate different interaction methods and the results show that our methods not only improve the effectiveness of occluded objects selection but also alleviate the arm fatigue problem significantly in the depth translation task. We also evaluate the proposed implicit online calibration method and find its accuracy comparable to standard 9 points explicit calibration, which makes a step towards practical use in the real world. Yiwei Bao, Zhimin Wang 0001, Feng Lu 0005 |
VR | 4 |
| 2023 | Discriminative feature encoding for intrinsic image decompositionabstractIntrinsic image decomposition is an important and long-standing computer vision problem. Given an input image, recovering the physical scene properties is ill-posed. Several physically motivated priors have been used to restrict the solution space of the optimization problem for intrinsic image decomposition. This work takes advantage of deep learning, and shows that it can solve this challenging computer vision problem with high efficiency. The focus lies in the feature encoding phase to extract discriminative features for different intrinsic layers from an input image. To achieve this goal, we explore the distinctive characteristics of different intrinsic components in the high-dimensional feature embedding space. We define feature distribution divergence to efficiently separate the feature vectors of different intrinsic components. The feature distributions are also constrained to fit the real ones through a feature distribution consistency. In addition, a data refinement approach is provided to remove data inconsistency from the Sintel dataset, making it more suitable for intrinsic image decomposition. Our method is also extended to intrinsic video decomposition based on pixel-wise correspondences between adjacent frames. Experimental results indicate that our proposed network structure can outperform the existing state-of-the-art. Zongji Wang, Yunfei Liu 0001, Feng Lu 0005 |
Comput. Vis. Media | 3 |
| 2023 | First- And Third-Person Video Co-Analysis By Learning Spatial-Temporal Joint AttentionabstractRecent years have witnessed a tremendous increase of first-person videos captured by wearable devices. Such videos record information from different perspectives than the traditional third-person view, and thus show a wide range of potential usages. However, techniques for analyzing videos from different views can be fundamentally different, not to mention co-analyzing on both views to explore the shared information. In this paper, we take the challenge of cross-view video co-analysis and deliver a novel learning-based method. At the core of our method is the notion of "joint attention", indicating the shared attention regions that link the corresponding views, and eventually guide the shared representation learning across views. To this end, we propose a multi-branch deep network, which extracts cross-view joint attention and shared representation from static frames with spatial constraints, in a self-supervised and simultaneous manner. In addition, by incorporating the temporal transition model of the joint attention, we obtain spatial-temporal joint attention that can robustly capture the essential information extending through time. Our method outperforms the state-of-the-art on the standard cross-view video matching tasks on public datasets. Furthermore, we demonstrate how the learnt joint information can benefit various applications through a set of qualitative and quantitative experiments. Huangyue Yu, Minjie Cai, Yunfei Liu 0001, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Layout-Bridging Text-to-Image SynthesisabstractThe crux of text-to-image synthesis stems from the difficulty of preserving the cross-modality semantic consistency between the input text and the synthesized image. Typical methods, which seek to model the text-to-image mapping directly, could only capture keywords in the text that indicates common objects or actions but fail to learn their spatial distribution patterns. An effective way to circumvent this limitation is to generate an image layout as guidance, which is attempted by a few methods. Nevertheless, these methods fail to generate practically effective layouts due to the diversity of input text and object location. In this paper we push for effective modeling in both text-to-layout generation and layout-to-image synthesis. Specifically, we formulate the text-to-layout generation as a sequence-to-sequence modeling task, and build our model upon Transformer to learn the spatial relationships between objects by modeling the sequential dependencies between them. In the stage of layout-to-image synthesis, we focus on learning the textual-visual semantic alignment per object in the layout to precisely incorporate the input text into the layout-to-image synthesizing process. To evaluate the quality of generated layout, we design a new metric specifically, dubbed Layout Quality Score, which considers both the absolute distribution errors of bounding boxes in the layout and the mutual spatial relationships between them. Extensive experiments on three datasets demonstrate the superior performance of our method over state-of-the-art methods on both predicting the layout and synthesizing the image from the given text. Jiadong Liang, Wenjie Pei, Feng Lu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | PureGaze: Purifying Gaze Feature for Generalizable Gaze EstimationabstractGaze estimation methods learn eye gaze from facial features. However, among rich information in the facial image, real gaze-relevant features only correspond to subtle changes in eye region, while other gaze-irrelevant features like illumination, personal appearance and even facial expression may affect the learning in an unexpected way. This is a major reason why existing methods show significant performance degradation in cross-domain/dataset evaluation. In this paper, we tackle the cross-domain problem in gaze estimation. Different from common domain adaption methods, we propose a domain generalization method to improve the cross-domain performance without touching target samples. The domain generalization is realized by gaze feature purification. We eliminate gaze-irrelevant factors such as illumination and identity to improve the cross-domain performance. We design a plug-and-play self-adversarial framework for the gaze feature purification. The framework enhances not only our baseline but also existing gaze estimation methods directly and significantly. To the best of our knowledge, we are the first to propose domain generalization methods in gaze estimation. Our method achieves not only state-of-the-art performance among typical gaze estimation methods but also competitive results among domain adaption methods. The code is released in https://github.com/yihuacheng/PureGaze. Yihua Cheng, Yiwei Bao, Feng Lu 0005 |
AAAI | 3 |
| 2022 | Generalizing Gaze Estimation with Rotation ConsistencyabstractRecent advances of deep learning-based approaches have achieved remarkable performance on appearance-based gaze estimation. However, due to the shortage of target domain data and absence of target labels, generalizing gaze estimation algorithm to unseen environments is still challenging. In this paper, we discover the rotation-consistency property in gaze estimation and introduce the ‘sub-label’ for unsupervised domain adaptation. Consequently, we propose the Rotation-enhanced Unsupervised Domain Adaptation (RUDA) for gaze estimation. First, we rotate the original images with different angles for training. Then we conduct domain adaptation under the constraint of rotation consistency. The target domain images are assigned with sub-labels, derived from relative rotation angles rather than untouchable real labels. With such sub-labels, we propose a novel distribution loss that facilitates the domain adaptation. We evaluate the RUDA framework on four cross-domain gaze estimation tasks. Experimental results demonstrate that it improves the performance over the baselines with gains ranging from 12.2% to 30.5%. Our framework has the potential to be used in other computer vision tasks with physical constraints. Yiwei Bao, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
CVPR | 4 |
| 2022 | GazeOnce: Real-Time Multi-Person Gaze EstimationabstractAppearance-based gaze estimation aims to predict the 3D eye gaze direction from a single image. While recent deep learning-based approaches have demonstrated excellent performance, they usually assume one calibrated face in each input image and cannot output multi-person gaze in real time. However, simultaneous gaze estimation for multiple people in the wild is necessary for real-world applications. In this paper, we propose the first one-stage end-to-end gaze estimation method, GazeOnce, which is capable of simultaneously predicting gaze directions for multiple faces (> 10) in an image. In addition, we design a sophisticated data generation pipeline and propose a new dataset, MPSGaze, which contains full images of multiple people with 3D gaze ground truth. Experimental results demonstrate that our unified framework not only offers a faster speed, but also provides a lower gaze estimation error compared with state-of-the-art methods. This technique can be useful in real-time applications with multiple users. Mingfang Zhang 0002, Yunfei Liu 0001, Feng Lu 0005 |
CVPR | 3 |
| 2022 | Gaze Estimation using TransformerabstractRecent work has proven the effectiveness of transformers in many computer vision tasks. However, the performance of transformers in gaze estimation is still unexplored. In this paper, we employ transformers and assess their effectiveness for gaze estimation. We consider two forms of vision transformer which are pure transformers and hybrid transformers. We first follow the popular ViT and employ a pure transformer to estimate gaze from images. On the other hand, we preserve the convolutional layers and integrate CNNs as well as transformers. The transformer serves as a component to complement CNNs. We compare the performance of the two transformers in gaze estimation. The Hybrid transformer significantly outperforms the pure transformer in all evaluation datasets with fewer parameters. We further conduct experiments to assess the effectiveness of the hybrid transformer and explore the advantage of the self-attention mechanism. Experiments show the hybrid transformer can achieve state-of-the-art performance in all benchmarks with pre-training. To facilitate further research, we release codes and models in https://github.com/yihuacheng/GazeTR. Yihua Cheng, Feng Lu 0005 |
ICPR | 2 |
| 2022 | Reconstructing 3D Virtual Face with Eye Gaze from a Single ImageabstractReconstructing 3D virtual face from a single image has a wide range of applications in virtual reality. Existing approaches synthesize plausible reconstructed virtual faces, however, eye gaze information is usually ignored, which is critical in human-computer interaction. In this paper, we propose to reconstruct 3D virtual face with eye gaze information from a single image. The main challenges lie in two aspects, one is the low reconstruction quality in the eye region, the other one is the lack of an efficient method to obtain precise eye gaze information. To address these problems, we decompose this task into two key steps, i.e., 3D face reconstruction with precise eye region and eye contact guided facial-rotation for eye gaze information. The first step is designed for precise eye region reconstruction through joint optimization on 3D face/eye shapes and textures. The second step consists of two parts: eye contact discriminator and automatic eye contact search algorithm via gradient-based optimization to perform both eye contact and gaze estimation simultaneously. Extensive experiments on different tasks demonstrate the significant gain of the proposed approach, achieving an MSE of (30%), an SSIM of (17.85%), and a PSNR of (8.4%). It also produces lower angular errors (63.01%) in the gaze estimation task compared with human annotations. Jiadong Liang, Yunfei Liu 0001, Feng Lu 0005 |
VR | 3 |
| 2022 | Patch-Based Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illumination without any calibration objects or initial guess of the target shape. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary orthogonal ambiguity. We further build the patch connections by extracting consistent surface normal pairs via spatial overlaps among patches and intensity profiles. Guided by these connections, the local ambiguities are unified to a global orthogonal one through Markov Random Field optimization and rotation averaging. After applying the integrability constraint, our solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Heng Guo 0003, Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Ping Tan 0002, Yasuyuki Matsushita |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Optical Flow in the DarkabstractOptical flow estimation in low-light conditions is a challenging task for existing methods and current optical flow datasets lack low-light samples. Even if the dark images are enhanced before estimation, which could achieve great visual perception, it still leads to suboptimal optical flow results because information like motion consistency may be broken during the enhancement. We propose to apply a novel training policy to learn optical flow directly from new synthetic and real low-light images. Specifically, first, we design a method to collect a new optical flow dataset in multiple exposures with shared optical flow pseudo labels. Then we apply a two-step process to create a synthetic low-light optical flow dataset, based on an existing bright one, by simulating low-light raw features from the multi-exposure raw images we collected. To extend the data diversity, we also include published low-light raw videos without optical flow labels. In our training pipeline, with the three datasets, we create two teacher-student pairs to progressively obtain optical flow labels for all data. Finally, we apply a mix-up training policy with our diversified datasets to produce low-light-robust optical flow models for release. The experiments show that our method can relatively maintain the optical flow accuracy as the image exposure descends and the generalization ability of our method is tested with different cameras in multiple practical scenes. Mingfang Zhang 0002, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Assessment of Deep Learning-Based Heart Rate Estimation Using Remote Photoplethysmography Under Different IlluminationsabstractRemote photoplethysmography (rPPG) monitors heart rate (HR) without requiring physical contact, which has applications. Deep learning based rPPG has demonstrated superior performance over the traditional approaches in controlled context. However, the lighting situation in indoor space is typically complex, with uneven light distribution and frequent variations in illumination. It lacks a fair comparison of different methods under different illuminations using the same dataset. In this article, we present a public dataset, namely the BeiHang University remote photoplethysmography (BH-rPPG) dataset, which contains data from 35 subjects under three illuminations: 1) low; 2) medium; and 3) high illumination. We also provide the ground truth HR measured by an oximeter. We evaluate the performance of three deep learning-based methods (Deepphys, rPPGNet, and Physnet) to that of four traditional methods (CHROM, GREEN, ICA, and POS) using two public datasets: 1) UBFC-rPPG; 2) the BH-rPPG. The experimental results demonstrate that traditional methods are more resistant to fluctuating illuminations. We found that the Physnet achieves lowest mean absolute error among deep learning based method under medium illumination, whereas the CHROM achieves 1.04 beats per minute, outperforming the Physnet by 80$\%$. Additionally, we investigate potential methods for improving performance of deep learning based methods. We find that brightness augmentation make model more robust to variation illumination. These findings suggest that while developing deep learning based HR estimation algorithms, illumination variation should be taken into account. This work serves as a benchmark for rPPG performance evaluation and it opens a pathway for future investigation into deep learning based rPPG under illumination variations. Haofei Wang 0001, Feng Lu 0005 |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2022 | Explainable Diabetic Retinopathy Detection and Retinal Image GenerationabstractThough deep learning has shown successful performance in classifying the label and severity stage of certain diseases, most of them give few explanations on how to make predictions. Inspired by Koch's Postulates, the foundation in evidence-based medicine (EBM) to identify the pathogen, we propose to exploit the interpretability of deep learning application in medical diagnosis. By isolating neuron activation patterns from a diabetic retinopathy (DR) detector and visualizing them, we can determine the symptoms that the DR detector identifies as evidence to make prediction. To be specific, we first define novel pathological descriptors using activated neurons of the DR detector to encode both spatial and appearance information of lesions. Then, to visualize the symptom encoded in the descriptor, we propose Patho-GAN, a new network to synthesize medically plausible retinal images. By manipulating these descriptors, we could even arbitrarily control the position, quantity, and categories of generated lesions. We also show that our synthesized images carry the symptoms directly related to diabetic retinopathy diagnosis. Our generated images are both qualitatively and quantitatively superior to the ones by previous methods. Besides, compared to existing methods that take hours to generate an image, our second level speed endows the potential to be an effective solution for data augmentation. Yuhao Niu, Lin Gu 0003, Yitian Zhao, Feng Lu 0005 |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Semantic Guided Single Image Reflection RemovalabstractReflection is common when we see through a glass window, which not only is a visual disturbance but also influences the performance of computer vision algorithms. Removing the reflection from a single image, however, is highly ill-posed since the color at each pixel needs to be separated into two values belonging to the clear background and the reflection, respectively. To solve this, existing methods use additional priors such as reflection layer smoothness, double reflection effect, and color consistency to distinguish the two layers. However, these low-level priors may not be consistently valid in real cases. In this paper, inspired by the fact that human beings can separate the two layers easily by recognizing the objects and understanding the scene, we propose to use the object semantic cue, which is high-level information, as the guidance to help reflection removal. Based on the data analysis, we develop a multi-task end-to-end deep learning method with a semantic guidance component, to solve reflection removal and semantic segmentation jointly. Extensive experiments on different datasets show significant performance gain when using high-level object-oriented information. We also demonstrate the application of our method to other computer vision tasks. Yunfei Liu 0001, Yu Li 0003, Shaodi You, Feng Lu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Gaze-Vergence-Controlled See-Through Vision in Augmented RealityabstractAugmented Reality (AR) see-through vision is an interesting research topic since it enables users to see through a wall and see the occluded objects. Most existing research focuses on the visual effects of see-through vision, while the interaction method is less studied. However, we argue that using common interaction modalities, e.g., midair click and speech, may not be the optimal way to control see-through vision. This is because when we want to see through something, it is physically related to our gaze depth/vergence and thus should be naturally controlled by the eyes. Following this idea, this paper proposes a novel gaze-vergence-controlled (GVC) see-through vision technique in AR. Since gaze depth is needed, we build a gaze tracking module with two infrared cameras and the corresponding algorithm and assemble it into the Microsoft HoloLens 2 to achieve gaze depth estimation. We then propose two different GVC modes for see-through vision to fit different scenarios. Extensive experimental results demonstrate that our gaze depth estimation is efficient and accurate. By comparing with conventional interaction modalities, our GVC techniques are also shown to be superior in terms of efficiency and more preferred by users. Finally, we present four example applications of gaze-vergence-controlled see-through vision. Zhimin Wang 0001, Feng Lu 0005 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2021 | Separating Content and Style for Unsupervised Image-to-Image Translation
Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005 |
BMVC | 4 |
| 2021 | Generalizing Gaze Estimation with Outlier-guided Collaborative AdaptationabstractDeep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation framework (PnP-GA), which is an ensemble of networks that learn collaboratively with the guidance of outliers. Since our proposed framework does not require ground-truth labels in the target domain, the existing gaze estimation networks can be directly plugged into PnPGA and generalize the algorithms to new domains. We test PnP-GA on four gaze domain adaptation tasks, ETH-to-MPII, ETH-to-EyeDiap, Gaze360-to-MPII, and Gaze360to-EyeDiap. The experimental results demonstrate that the PnP-GA framework achieves considerable performance improvements of 36.9%, 31.6%, 19.4%, and 11.8% over the baseline system. The proposed framework also outperforms the state-of-the-art domain adaptation approaches on gaze domain adaptation tasks. Code has been released at https://github.com/DreamtaleCore/PnP-GA. Yunfei Liu 0001, Ruicong Liu, Haofei Wang 0001, Feng Lu 0005 |
ICCV | 4 |
| 2021 | Edge-Guided Near-Eye Image Analysis for Head Mounted DisplaysabstractEye tracking provides an effective way for interaction in Augmented Reality (AR) Head Mounted Displays (HMDs). Current eye tracking techniques for AR HMDs require eye segmentation and ellipse fitting under near-infrared illumination. However, due to the low contrast between sclera and iris regions and unpredictable reflections, it is still challenging to accomplish accurate iris/pupil segmentation and the corresponding ellipse fitting tasks. In this paper, inspired by the fact that most essential information is encoded in the edge areas, we propose a novel near-eye image analysis method with edge maps as guidance. Specifically, we first utilize an Edge Extraction Network ($E^{2}-$Net) to predict high-quality edge maps, which only contain eyelids and iris/pupil contours without other undesired edges. Then we feed the edge maps into an Edge-Guided Segmentation and Fitting Network (ESF-Net) for accurate segmentation and ellipse fitting. Extensive experimental results demonstrate that our method outperforms current state-of-the-art methods in near-eye image segmentation and ellipse fitting tasks, based on which we present applications of eye tracking with AR HMD. Zhimin Wang 0001, Yunfei Liu 0001, Feng Lu 0005 |
ISMAR | 4 |
| 2021 | Attention Guided Low-Light Image Enhancement with a Large Scale Low-Light Simulation Dataset
Feifan Lv, Yu Li 0003, Feng Lu 0005 |
Int. J. Comput. Vis. | 3 |
| 2021 | Understanding adversarial attacks on deep learning based medical image analysis systems
Xingjun Ma, Yuhao Niu, Lin Gu 0003, Yisen Wang 0001, Yitian Zhao, James Bailey 0001, Feng Lu 0005 |
Pattern Recognit. | 7 |
| 2021 | Interaction With Gaze, Gesture, and Speech in a Flexibly Configurable Augmented Reality SystemabstractMultimodal interaction has become a recent research focus since it offers better user experience in augmented reality (AR) systems. However, most existing works only combine two modalities at a time, e.g., gesture and speech. Multimodal interactive system integrating gaze cue has rarely been investigated. In this article, we propose a multimodal interactive system that integrates gaze, gesture, and speech in a flexibly configurable AR system. Our lightweight head-mounted device supports accurate gaze tracking, hand gesture recognition, and speech recognition simultaneously. The system can be easily configured into various modality combinations, which enables us to investigate the effects of different interaction techniques. We evaluate the efficiency of these modalities using two tasks: the lamp brightness adjustment task and the cube manipulation task. We also collect subjective feedback when using such systems. The experimental results demonstrate that theGaze+Gesture+Speechmodality is superior in terms of efficiency, and theGesture+Speechmodality is more preferred by users. Our system opens the pathway toward a multimodal interactive AR system that enables flexible configuration. Zhimin Wang 0001, Haofei Wang 0001, Huangyue Yu, Feng Lu 0005 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2021 | A Sparse Representation Based Joint Demosaicing Method for Single-Chip Polarized Color SensorabstractThe emergence of the single-chip polarized color sensor now allows for simultaneously capturing chromatic and polarimetric information of the scene on a monochromatic image plane. However, unlike the usual camera with an embedded demosaicing method, the latest polarized color camera is not delivered with an in-built demosaicing tool. For demosaicing, the users have to down-sample the captured images or to use traditional interpolation techniques. Neither of them can perform well since the polarization and color are interdependent. Therefore, joint chromatic and polarimetric demosaicing is the key to obtaining high-quality polarized color images. In this paper, we propose a joint chromatic and polarimetric demosaicing model to address this challenging problem. Instead of mechanically demosaicing for the multi-channel polarized color image, we further present a sparse representation-based optimization strategy that utilizes chromatic information and polarimetric information to jointly optimize the model. To avoid the interaction between color and polarization during demosaicing, we separately construct the corresponding dictionaries. We also build an optical data acquisition system to collect a dataset, which contains various sources of polarization, such as illumination, reflectance and birefringence. Results of both qualitative and quantitative experiments have shown that our method is capable of faithfully recovering full RGB information of four polarization angles for each pixel from a single mosaic input image. Moreover, the proposed method can perform well not only on the synthetic data but the real captured data. Sijia Wen, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Image Process. | 3 |
| 2021 | Polarization Guided Specular Reflection SeparationabstractSince specular reflection often exists in the real captured images and causes deviation between the recorded color and intrinsic color, specular reflection separation can bring advantages to multiple applications that require consistent object surface appearance. However, due to the color of an object is significantly influenced by the color of the illumination, the existing researches still suffer from the near-duplicate challenge, that is, the separation becomes unstable when the illumination color is close to the surface color. In this paper, we derive a polarization guided model to incorporate the polarization information into a designed iteration optimization separation strategy to separate the specular reflection. Based on the analysis of polarization, we propose a polarization guided model to generate a polarization chromaticity image, which is able to reveal the geometrical profile of the input image in complex scenarios, e.g., diversity of illumination. The polarization chromaticity image can accurately cluster the pixels with similar diffuse color. We further use the specular separation of all these clusters as an implicit prior to ensure that the diffuse component will not be mistakenly separated as the specular component. With the polarization guided model, we reformulate the specular reflection separation into a unified optimization function which can be solved by the ADMM strategy. The specular reflection will be detected and separated jointly by RGB and polarimetric information. Both qualitative and quantitative experimental results have shown that our method can faithfully separate the specular reflection, especially in some challenging scenarios. Sijia Wen, Yinqiang Zheng, Feng Lu 0005 |
IEEE Trans. Image Process. | 3 |
| 2020 | A Coarse-to-Fine Adaptive Network for Appearance-Based Gaze EstimationabstractHuman gaze is essential for various appealing applications. Aiming at more accurate gaze estimation, a series of recent works propose to utilize face and eye images simultaneously. Nevertheless, face and eye images only serve as independent or parallel feature sources in those works, the intrinsic correlation between their features is overlooked. In this paper we make the following contributions: 1) We propose a coarse-to-fine strategy which estimates a basic gaze direction from face image and refines it with corresponding residual predicted from eye images. 2) Guided by the proposed strategy, we design a framework which introduces a bi-gram model to bridge gaze residual and basic gaze direction, and an attention component to adaptively acquire suitable fine-grained feature. 3) Integrating the above innovations, we construct a coarse-to-fine adaptive network named CA-Net and achieve state-of-the-art performances on MPIIGaze and EyeDiap. Yihua Cheng, Shiyao Huang, Fei Wang 0032, Chen Qian 0006, Feng Lu 0005 |
AAAI | 5 |
| 2020 | Separate in Latent Space: Unsupervised Single Image Layer SeparationabstractMany real world vision tasks, such as reflection removal from a transparent surface and intrinsic image decomposition, can be modeled as single image layer separation. However, this problem is highly ill-posed, requiring accurately aligned and hard to collect triplet data to train the CNN models. To address this problem, this paper proposes an unsupervised method that requires no ground truth data triplet in training. At the core of the method are two assumptions about data distributions in the latent spaces of different layers, based on which a novel unsupervised layer separation pipeline can be derived. Then the method can be constructed based on the GANs framework with self-supervision and cycle consistency constraints, etc. Experimental results demonstrate its successfulness in outperforming existing unsupervised methods in both synthetic and real world tasks. The method also shows its ability to solve a more challenging multi-layer separation task. Yunfei Liu 0001, Feng Lu 0005 |
AAAI | 2 |
| 2020 | An Integrated Enhancement Solution for 24-Hour Colorful ImagingabstractThe current industry practice for 24-hour outdoor imaging is to use a silicon camera supplemented with near-infrared (NIR) illumination. This will result in color images with poor contrast at daytime and absence of chrominance at nighttime. For this dilemma, all existing solutions try to capture RGB and NIR images separately. However, they need additional hardware support and suffer from various drawbacks, including short service life, high price, specific usage scenario, etc. In this paper, we propose a novel and integrated enhancement solution that produces clear color images, whether at abundant sunlight daytime or extremely low-light nighttime. Our key idea is to separate the VIS and NIR information from mixed signals, and enhance the VIS signal adaptively with the NIR signal as assistance. To this end, we build an optical system to collect a new VIS-NIR-MIX dataset and present a physically meaningful image processing algorithm based on CNN. Extensive experiments show outstanding results, which demonstrate the effectiveness of our solution. Feifan Lv, Yinqiang Zheng, Feng Lu 0005 |
AAAI | 4 |
| 2020 | Real-Time Semantic Segmentation via Multiply Spatial Fusion Network
Haiyang Si, Feng Lu 0005 |
BMVC | 3 |
| 2020 | Generalizing Hand Segmentation in Egocentric Videos With Uncertainty-Guided Model AdaptationabstractAlthough the performance of hand segmentation in egocentric videos has been significantly improved by using CNNs, it still remains a challenging issue to generalize the trained models to new domains, e.g., unseen environments. In this work, we solve the hand segmentation generalization problem without requiring segmentation labels in the target domain. To this end, we propose a Bayesian CNN-based model adaptation framework for hand segmentation, which introduces and considers two key factors: 1) prediction uncertainty when the model is applied in a new domain and 2) common information about hand shapes shared across domains. Consequently, we propose an iterative self-training method for hand segmentation in the new domain, which is guided by the model uncertainty estimated by a Bayesian CNN. We further use an adversarial component in our framework to utilize shared information about hand shapes to constrain the model adaptation process. Experiments on multiple egocentric datasets show that the proposed method significantly improves the generalization performance of hand segmentation. Minjie Cai, Feng Lu 0005, Yoichi Sato 0001 |
CVPR | 2 |
| 2020 | Unsupervised Learning for Intrinsic Image Decomposition From a Single ImageabstractIntrinsic image decomposition, which is an essential task in computer vision, aims to infer the reflectance and shading of the scene. It is challenging since it needs to separate one image into two components. To tackle this, conventional methods introduce various priors to constrain the solution, yet with limited performance. Meanwhile, the problem is typically solved by supervised learning methods, which is actually not an ideal solution since obtaining ground truth reflectance and shading for massive general natural scenes is challenging and even impossible. In this paper, we propose a novel unsupervised intrinsic image decomposition framework, which relies on neither labeled training data nor hand-crafted priors. Instead, it directly learns the latent feature of reflectance and shading from unsupervised and uncorrelated data. To enable this, we explore the independence between reflectance and shading, the domain invariant content constraint and the physical constraint. Extensive experiments on both synthetic and real image datasets demonstrate consistently superior performance of the proposed method. Yunfei Liu 0001, Yu Li 0003, Shaodi You, Feng Lu 0005 |
CVPR | 4 |
| 2020 | Optical Flow in the DarkabstractMany successful optical flow estimation methods have been proposed, but they become invalid when tested in dark scenes because low-light scenarios are not considered when they are designed and current optical flow benchmark datasets lack low-light samples. Even if we preprocess to enhance the dark images, which achieves great visual perception, it still leads to poor optical flow results or even worse ones, because information like motion consistency may be broken while enhancing. We propose an end-to-end data-driven method that avoids error accumulation and learns optical flow directly from low-light noisy images. Specifically, we develop a method to synthesize large-scale low-light optical flow datasets by simulating the noise model on dark raw images. We also collect a new optical flow dataset in raw format with a large range of exposure to be used as a benchmark. The models trained on our synthetic dataset can relatively maintain optical flow accuracy as the image brightness descends and they outperform the existing methods greatly on low-light images. Yinqiang Zheng, Mingfang Zhang 0002, Feng Lu 0005 |
CVPR | 3 |
| 2020 | CPGAN: Content-Parsing Generative Adversarial Networks for Text-to-Image Synthesis
Jiadong Liang, Wenjie Pei, Feng Lu 0005 |
ECCV (4) | 3 |
| 2020 | Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
Yunfei Liu 0001, Xingjun Ma, James Bailey 0001, Feng Lu 0005 |
ECCV (10) | 4 |
| 2020 | Adaptive Feature Fusion Network for Gaze Tracking in Mobile TabletsabstractRecently, many multi-stream gaze estimation methods have been proposed. They estimate gaze from eye and face appearances and achieve reasonable accuracy. However, most of the methods simply concatenate the features extracted from eye and face appearance. The feature fusion process has been ignored. In this paper, we propose a novel Adaptive Feature Fusion Network (AFF-Net), which performs gaze tracking task in mobile tablets. We stack two-eye feature maps and utilize Squeeze-and-Excitation layers to adaptively fuse two-eye features according to their similarity on appearance. Meanwhile, we also propose Adaptive Group Normalization to recalibrate eye features with the guidance of facial feature. Extensive experiments on both GazeCapture and MPIIFaceGaze datasets demonstrate consistently superior performance of the proposed method. Yiwei Bao, Yihua Cheng, Yunfei Liu 0001, Feng Lu 0005 |
ICPR | 4 |
| 2020 | Fast Enhancement for Non-Uniform Illumination Images using Light-weight CNNsabstractThis paper proposes a new light-weight convolutional neural network (~5k params) for non-uniform illumination image enhancement to handle color, exposure, contrast, noise and artifacts, etc., simultaneously and effectively. More concretely, the input image is first enhanced using Retinex model from dual different aspects (enhancing under-exposure and suppressing over-exposure), respectively. Then, these two enhanced results and the original image are fused to obtain an image with satisfactory brightness, contrast and details. Finally, the extra noise and compression artifacts are removed to get the final result. To train this network, we propose a semi-supervised retouching solution and construct a new dataset (~82k images) that contains various scenes and light conditions. Our model can enhance 0.5 mega-pixel (like 600×800) images in real-time (~50 fps), which is faster than existing enhancement methods. Extensive experiments show that our solution is fast and effective to deal with non-uniform illumination images. Feifan Lv, Bo Liu 0112, Feng Lu 0005 |
ACM Multimedia | 3 |
| 2020 | Gaze Estimation by Exploring Two-Eye AsymmetryabstractEye gaze estimation is increasingly demanded by recent intelligent systems to facilitate a range of interactive applications. Unfortunately, learning the highly complicated regression from a single eye image to the gaze direction is not trivial. Thus, the problem is yet to be solved efficiently. Inspired by the two-eye asymmetry as two eyes of the same person may appear uneven, we propose the face-based asymmetric regression-evaluation network (FARE-Net) to optimize the gaze estimation results by considering the difference between left and right eyes. The proposed method includes one face-based asymmetric regression network (FAR-Net) and one evaluation network (E-Net). The FAR-Net predicts 3D gaze directions for both eyes and is trained with the asymmetric mechanism, which asymmetrically weights and sums the loss generated by two-eye gaze directions. With the asymmetric mechanism, the FAR-Net utilizes the eyes that can achieve high performance to optimize network. The E-Net learns the reliabilities of two eyes to balance the learning of the asymmetric mechanism and symmetric mechanism. Our FARENet achieves leading performances on MPIIGaze, EyeDiap and RT-Gene datasets. Additionally, we investigate the effectiveness of FARE-Net by analyzing the distribution of errors and ablation study. Yihua Cheng, Xucong Zhang, Feng Lu 0005, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Mutual Context Network for Jointly Estimating Egocentric Gaze and ActionabstractIn this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context: the information from gaze prediction facilitates action recognition and vice versa. Our assumption is that during the procedure of performing a manipulation task, on the one hand, what a person is doing determines where the person is looking at. On the other hand, the gaze location reveals gaze regions which contain important and information about the undergoing action and also the non-gaze regions that include complimentary clues for differentiating some fine-grained actions. We propose a novel mutual context network (MCN) that jointly learns action-dependent gaze prediction and gaze-guided action recognition in an end-to-end manner. Experiments on multiple egocentric video datasets demonstrate that our MCN achieves state-of-the-art performance of both gaze prediction and action recognition. The experiments also show that action-dependent gaze patterns could be learned with our method. Yifei Huang 0002, Minjie Cai, Zhenqiang Li 0002, Feng Lu 0005, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 39 |
| 2020 | VoxSegNet: Volumetric CNNs for Semantic Part Segmentation of 3D ShapesabstractVolumetric representation has been widely used for 3D deep learning in shape analysis due to its generalization ability and regular data format. However, for fine-grained tasks like part segmentation, volumetric data has not been widely adopted compared to other representations. Aiming at delivering an effective volumetric method for 3D shape part segmentation, this paper proposes a novel volumetric convolutional neural network. Our method can extract discriminative features encoding detailed information from voxelized 3D data under limited resolution. To this purpose, a spatial dense extraction (SDE) module is designed to preserve spatial resolution during feature extraction procedure, alleviating the loss of details caused by sub-sampling operations such as max pooling. An attention feature aggregation (AFA) module is also introduced to adaptively select informative features from different abstraction levels, leading to segmentation with both semantic consistency and high accuracy of details. Experimental results demonstrate that promising results can be achieved by using volumetric data, with part segmentation accuracy comparable or superior to state-of-the-art non-volumetric methods. Zongji Wang, Feng Lu 0005 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2019 | Pathological Evidence Exploration in Deep Retinal Image DiagnosisabstractThough deep learning has shown successful performance in classifying the label and severity stage of certain disease, most of them give few evidence on how to make prediction. Here, we propose to exploit the interpretability of deep learning application in medical diagnosis. Inspired by Koch’s Postulates, a well-known strategy in medical research to identify the property of pathogen, we define a pathological descriptor that can be extracted from the activated neurons of a diabetic retinopathy detector. To visualize the symptom and feature encoded in this descriptor, we propose a GAN based method to synthesize pathological retinal image given the descriptor and a binary vessel segmentation. Besides, with this descriptor, we can arbitrarily manipulate the position and quantity of lesions. As verified by a panel of 5 licensed ophthalmologists, our synthesized images carry the symptoms that are directly related to diabetic retinopathy diagnosis. The panel survey also shows that our generated images is both qualitatively and quantitatively superior to existing methods. Yuhao Niu, Lin Gu 0003, Feng Lu 0005, Feifan Lv, Zongji Wang, Imari Sato, Zijian Zhang 0004, Yangyan Xiao, Xunzhang Dai |
AAAI | 3 |
| 2019 | Turn a Silicon Camera Into an InGaAs CameraabstractShort-wave infrared (SWIR) imaging has a wide range of applications for both industry and civilian. However, the InGaAs sensors commonly used for SWIR imaging suffer from a variety of drawbacks, including high price, low resolution, unstable quality, and so on. In this paper, we propose a novel solution for SWIR imaging using a common Silicon sensor, which has cheaper price, higher resolution and better technical maturity compared with the specialized InGaAs sensor. Our key idea is to approximate the response of the InGaAs sensor by exploiting the largely ignored sensitivity of a Silicon sensor, weak as it is, in the SWIR range. To this end, we build a multi-channel optical system to collect a new SWIR dataset and present a physically meaningful three-stage image processing algorithm on the basis of CNN. Both qualitative and quantitative experiments show promising experimental results, which demonstrate the effectiveness of the proposed method. Feifan Lv, Yinqiang Zheng, Feng Lu 0005 |
CVPR | 4 |
| 2019 | Unsupervised Ensemble Strategy for Retinal Vessel Segmentation
Bo Liu 0112, Lin Gu 0003, Feng Lu 0005 |
MICCAI (1) | 3 |
| 2019 | What I See Is What You See: Joint Attention Learning for First and Third Person Video Co-analysisabstractIn recent years, more and more videos are captured from the first-person viewpoint by wearable cameras. Such first-person video provides additional information besides the traditional third-person video, and thus has a wide range of applications. However, techniques for analyzing the first-person video can be fundamentally different from those for the third-person video, and it is even more difficult to explore the shared information from both viewpoints. In this paper, we propose a novel method for first- and third-person video co-analysis. At the core of our method is the notion of "joint attention'', indicating the learnable representation that corresponds to the shared attention regions in different viewpoints and thus links the two viewpoints. To this end, we develop a multi-branch deep network with a triplet loss to extract the joint attention from the first- and third-person videos via self-supervised learning. We evaluate our method on the public dataset with cross-viewpoint video matching tasks. Our method outperforms the state-of-the-art both qualitatively and quantitatively. We also demonstrate how the learned joint attention can benefit various applications through a set of additional experiments. Huangyue Yu, Minjie Cai, Yunfei Liu 0001, Feng Lu 0005 |
ACM Multimedia | 4 |
| 2019 | Desktop Action Recognition From First-Person Point-of-ViewabstractDesktop action recognition from first-person view (egocentric) video is an important task due to its omnipresence in our daily life, and the ideal first-person viewing perspective for observing hand-object interactions. However, no previous research efforts have been dedicated on the benchmark of the task. In this paper, we first release a dataset of daily desktop actions recorded with a wearable camera and publish it as a benchmark for desktop action recognition. Regular desktop activities of six participants were recorded in egocentric video with a wide-angle head-mounted camera. In particular, we focus on five common desktop actions in which hands are involved. We provide original video data, action annotations at frame-level, and hand masks at pixel-level. We also propose a feature representation for the characterization of different desktop actions based on the spatial and temporal information of hands. In experiments, we illustrate the statistical information about the dataset, and evaluate the action recognition performance of different features as a baseline. The proposed method achieves promising performance for five action classes. Minjie Cai, Feng Lu 0005, Yue Gao 0002 |
IEEE Trans. Cybern. | 2 |
| 2019 | Makeup Removal via Bidirectional Tunable De-Makeup NetworkabstractWe present a deep learning-based method for removing makeup effects (de-makeup) in a face image. This problem poses a major challenge due to obscuring of the underlying facial features by cosmetics, which is very important in multimedia applications in the field of security, entertainment, and social networking. To address this task, we propose the bidirectional tunable de-makeup network (BTD-Net), which jointly learns the makeup process to aid in learning the de-makeup process. For tractable learning of the makeup process, which is a one-to-many mapping determined by the cosmetics that are applied, we introduce a latent variable that reflects the makeup style. This latent variable is extracted in the de-makeup process and used as a condition on the makeup process to constrain the one-to-many mapping to a specific solution. Through extensive experiments, our proposed BTD-Net is found to surpass the state-of-art techniques in estimating realistic non-makeup faces that correspond to the input makeup images. We additionally show that applications such as tuning the amount of makeup can be enhanced through the use of this method. Feng Lu 0005, Chen Li 0031, Stephen Lin 0001, Xukun Shen |
IEEE Trans. Multim. | 2 |
| 2018 | Detail Preserving Depth Estimation from a Single Image Using Attention Guided NetworksabstractConvolutional Neural Networks have demonstrated superior performance on single image depth estimation in recent years. These works usually use stacked spatial pooling or strided convolution to get high-level information which are common practices in classification task. However, depth estimation is a dense prediction problem and low-resolution feature maps usually generate blurred depth map which is undesirable in application. In order to produce high quality depth map, say clean and accurate, we propose a network consists of a Dense Feature Extractor (DFE) and a Depth Map Generator (DMG). The DFE combines ResNet and dilated convolutions. It extracts multi-scale information from input image while keeping the feature maps dense. As for DMG, we use attention mechanism to fuse multi-scale features produced in DFE. Our Network is trained end-to-end and does not need any post-processing. Hence, it runs fast and can predict depth map in about 15 fps. Experiment results show that our method is competitive with the state-of-the-art in quantitative evaluation, but can preserve better structural details of the scene depth. Zhixiang Hao, Yu Li 0003, Shaodi You, Feng Lu 0005 |
3DV | 4 |
| 2018 | MBLLEN: Low-Light Image/Video Enhancement Using CNNs
Feifan Lv, Feng Lu 0005, Chongsoon Lim |
BMVC | 2 |
| 2018 | Uncalibrated Photometric Stereo Under Natural IlluminationabstractThis paper presents a photometric stereo method that works with unknown natural illuminations without any calibration object. To solve this challenging problem, we propose the use of an equivalent directional lighting model for small surface patches consisting of slowly varying normals, and solve each patch up to an arbitrary rotation ambiguity. Our method connects the resulting patches and unifies the local ambiguities to a global rotation one through angular distance propagation defined over the whole surface. After applying the integrability constraint, our final solution contains only a binary ambiguity, which could be easily removed. Experiments using both synthetic and real-world datasets show our method provides even comparable results to calibrated methods. Zhipeng Mo, Boxin Shi, Feng Lu 0005, Sai-Kit Yeung, Yasuyuki Matsushita |
CVPR | 3 |
| 2018 | Appearance-Based Gaze Estimation via Evaluation-Guided Asymmetric Regression
Yihua Cheng, Feng Lu 0005, Xucong Zhang |
ECCV (14) | 2 |
| 2018 | Reconstructing non-rigid object with large movement using a single depth camera
Feixiang Lu, Feng Lu 0005, Yu Zhang 0035, Xiaowu Chen 0001, Qinping Zhao |
Comput. Aided Geom. Des. | 3 |
| 2018 | SymPS: BRDF Symmetry Guided Photometric Stereo for Shape and Light Source EstimationabstractWe propose uncalibrated photometric stereo methods that address the problem due to unknown isotropic reflectance. At the core of our methods is the notion of "constrained half-vector symmetry" for general isotropic BRDFs. We show that such symmetry can be observed in various real-world materials, and it leads to new techniques for shape and light source estimation. Based on the 1D and 2D representations of the symmetry, we propose two methods for surface normal estimation; one focuses on accurate elevation angle recovery for surface normals when the light sources only cover the visible hemisphere, and the other for comprehensive surface normal optimization in the case that the light sources are also non-uniformly distributed. The proposed robust light source estimation method also plays an essential role to let our methods work in an uncalibrated manner with good accuracy. Quantitative evaluations are conducted with both synthetic and real-world scenes, which produce the state-of-the-art accuracy for all of the non-Lambertian materials in MERL database and the real-world datasets. Feng Lu 0005, Xiaowu Chen 0001, Imari Sato, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Look, Perceive and Segment: Finding the Salient Objects in Images via Two-stream Fixation-Semantic CNNsabstractRecently, CNN-based models have achieved remarkable success in image-based salient object detection (SOD). In these models, a key issue is to find a proper network architecture that best fits for the task of SOD. Toward this end, this paper proposes two-stream fixation-semantic CNNs, whose architecture is inspired by the fact that salient objects in complex images can be unambiguously annotated by selecting the pre-segmented semantic objects that receive the highest fixation density in eye-tracking experiments. In the two-stream CNNs, a fixation stream is pre-trained on eye-tracking data whose architecture well fits for the task of fixation prediction, and a semantic stream is pre-trained on images with semantic tags that has a proper architecture for semantic perception. By fusing these two streams into an inception-segmentation module and jointly fine-tuning them on images with manually annotated salient objects, the proposed networks show impressive performance in segmenting salient objects. Experimental results show that our approach outperforms 10 state-of-the-art models (5 deep, 5 non-deep) on 4 datasets. Xiaowu Chen 0001, Anlin Zheng, Jia Li 0003, Feng Lu 0005 |
ICCV | 4 |
| 2017 | Teaching robots to do object assembly using multi-modal 3D vision
Weiwei Wan, Feng Lu 0005, Zepei Wu, Kensuke Harada |
Neurocomputing | 2 |
| 2017 | Appearance-Based Gaze Estimation via Uncalibrated Gaze Pattern RecoveryabstractAiming at reducing the restrictions due to person/scene dependence, we deliver a novel method that solves appearance-based gaze estimation in a novel fashion. First, we introduce and solve an "uncalibrated gaze pattern" solely from eye images independent of the person and scene. The gaze pattern recovers gaze movements up to only scaling and translation ambiguities, via nonlinear dimension reduction and pixel motion analysis, while no training/calibration is needed. This is new in the literature and enables novel applications. Second, our method allows simple calibrations to align the gaze pattern to any gaze target. This is much simpler than conventional calibrations which rely on sufficient training data to compute person and scene-specific nonlinear gaze mappings. Through various evaluations, we show that: 1) the proposed uncalibrated gaze pattern has novel and broad capabilities; 2) the proposed calibration is simple and efficient, and can be even omitted in some scenarios; and 3) quantitative evaluations produce promising results under various conditions. Feng Lu 0005, Xiaowu Chen 0001, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | 3-Points Convex Hull Matching (3PCHM) for fast and robust point set registration
Jingfan Fan, Jian Yang 0009, Feng Lu 0005, Danni Ai, Yitian Zhao, Yongtian Wang |
Neurocomputing | 3 |
| 2016 | Person-independent eye gaze prediction from eye images using patch-based features
Feng Lu 0005, Xiaowu Chen 0001 |
Neurocomputing | 1 |
| 2016 | Visual facial expression modeling and early predicting from 3D data via subtle feature enhancing
Lumei Su, Feng Lu 0005 |
Multim. Tools Appl. | 2 |
| 2016 | Error-tolerant manipulation by caging
Weiwei Wan, Feng Lu 0005, Rui Fukui |
Signal Process. | 2 |
| 2016 | Estimating 3D Gaze Directions Using Unlabeled Eye Images via Synthetic Iris Appearance FittingabstractEstimating three-dimensional (3D) human eye gaze by capturing a single eye image without active illumination is challenging. Although the elliptical iris shape provides a useful cue, existing methods face difficulties in ellipse fitting due to unreliable iris contour detection. These methods may fail frequently especially with low resolution eye images. In this paper, we propose a synthetic iris appearance fitting (SIAF) method that is model-driven to compute 3D gaze direction from iris shape. Instead of fitting an ellipse based on exactly detected iris contour, our method first synthesizes a set of physically possible iris appearances and then optimizes inside this synthetic space to find the best solution to explain the captured eye image. In this way, the solution is highly constrained and guaranteed to be physically feasible. In addition, the proposed advanced image analysis techniques also help the SIAF method be robust to the unreliable iris contour detection. Furthermore, with multiple eye images, we propose a SIAF-joint method that can further reduce the gaze error by half, and it also resolves the binary ambiguity which is inevitable in conventional methods based on simple ellipse fitting. Feng Lu 0005, Yue Gao 0002, Xiaowu Chen 0001 |
IEEE Trans. Multim. | 1 |
| 2015 | Uncalibrated photometric stereo based on elevation angle recovery from BRDF symmetry of isotropic materialsabstractThis paper addresses the problem of uncalibrated photometric stereo with isotropic reflectances. Existing methods face difficulty in solving for the elevation angles of surface normals when the light sources only cover the visible hemisphere. Here, we introduce the notion of “constrained half-vector symmetry” for general isotropic BRDFs and show its capability of elevation angle recovery. This sort of symmetry can be observed in a 1D BRDF slice from a subset of surface normals with the same azimuth angle, and we use it to devise an efficient modeling and solution method to constrain and recover the elevation angles of surface normals accurately. To enable our method to work in an uncalibrated manner, we further solve for light sources in the case of general isotropic BRDFs. By combining this method with the existing ones for azimuth angle estimation, we can get state-of-the-art results for uncalibrated photometric stereo with general isotropic reflectances. Feng Lu 0005, Imari Sato, Yoichi Sato 0001 |
CVPR | 1 |
| 2015 | From Intensity Profile to Surface Normal: Photometric Stereo for Unknown Light Sources and Isotropic ReflectancesabstractWe propose an uncalibrated photometric stereo method that works with general and unknown isotropic reflectances. Our method uses a pixel intensity profile, which is a sequence of radiance intensities recorded at a pixel under unknown varying directional illumination. We show that for general isotropic materials and uniformly distributed light directions, the geodesic distance between intensity profiles is linearly related to the angular difference of their corresponding surface normals, and that the intensity distribution of the intensity profile reveals reflectance properties. Based on these observations, we develop two methods for surface normal estimation; one for a general setting that uses only the recorded intensity profiles, the other for the case where a BRDF database is available while the exact BRDF of the target scene is still unknown. Quantitative and qualitative evaluations are conducted using both synthetic and real-world scenes, which show the state-of-the-art accuracy of smaller than 10 degree without using reference data and 5 degree with reference data for all 100 materials in MERL database. Feng Lu 0005, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Gaze Estimation From Eye Appearance: A Head Pose-Free Method via Eye Image SynthesisabstractIn this paper, we address the problem of free head motion in appearance-based gaze estimation. This problem remains challenging because head motion changes eye appearance significantly, and thus, training images captured for an original head pose cannot handle test images captured for other head poses. To overcome this difficulty, we propose a novel gaze estimation method that handles free head motion via eye image synthesis based on a single camera. Compared with conventional fixed head pose methods with original training images, our method only captures four additional eye images under four reference head poses, and then, precisely synthesizes new training images for other unseen head poses in estimation. To this end, we propose a single-directional (SD) flow model to efficiently handle eye image variations due to head motion. We show how to estimate SD flows for reference head poses first, and then use them to produce new SD flows for training image synthesis. Finally, with synthetic training images, joint optimization is applied that simultaneously solves an eye image alignment and a gaze estimation. Evaluation of the method was conducted through experiments to assess its performance and demonstrate its effectiveness. Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Learning gaze biases with head motion for head pose-free gaze estimation
Feng Lu 0005, Takahiro Okabe, Yusuke Sugano, Yoichi Sato 0001 |
Image Vis. Comput. | 1 |
| 2014 | Adaptive Linear Regressionfor Appearance-Based Gaze EstimationabstractWe investigate the appearance-based gaze estimation problem, with respect to its essential difficulty in reducing the number of required training samples, and other practical issues such as slight head motion, image resolution variation, and eye blinking. We cast the problem as mapping high-dimensional eye image features to low-dimensional gaze positions, and propose an adaptive linear regression (ALR) method as the key to our solution. The ALR method adaptively selects an optimal set of sparsest training samples for the gaze estimation via ℓ(1)-optimization. In this sense, the number of required training samples is significantly reduced for high accuracy estimation. In addition, by adopting the basic ALR objective function, we integrate the gaze estimation, subpixel alignment and blink detection into a unified optimization framework. By solving these problems simultaneously, we successfully handle slight head motion, image resolution variation and eye blinking in appearance-based gaze estimation. We evaluated the proposed method by conducting experiments with multiple users and variant conditions to verify its effectiveness. Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Uncalibrated Photometric Stereo for Unknown Isotropic ReflectancesabstractWe propose an uncalibrated photometric stereo method that works with general and unknown isotropic reflectances. Our method uses a pixel intensity profile, which is a sequence of radiance intensities recorded at a pixel across multi-illuminance images. We show that for general isotropic materials, the geodesic distance between intensity profiles is linearly related to the angular difference of their surface normals, and that the intensity distribution of an intensity profile conveys information about the reflectance properties, when the intensity profile is obtained under uniformly distributed directional lightings. Based on these observations, we show that surface normals can be estimated up to a convex/concave ambiguity. A solution method based on matrix decomposition with missing data is developed for a reliable estimation. Quantitative and qualitative evaluations of our method are performed using both synthetic and real-world scenes. Feng Lu 0005, Yasuyuki Matsushita, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
CVPR | 1 |
| 2012 | Head pose-free appearance-based gaze sensing via eye image synthesis
Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001 |
ICPR | 1 |
| 2011 | A Head Pose-free Approach for Appearance-based Gaze EstimationabstractTo infer human gaze from eye appearance, various methods have been proposed. However, most of them assume a fixed head pose because allowing free head motion adds 6 degrees of freedom to the problem and requires a prohibitively large number of training samples. In this paper, we aim at solving the appearance-based gaze estimation problem under free head motion without significantly increasing the cost of training. The idea is to decompose the problem into subproblems, including initial estimation under fixed head pose and subsequent compensations for estimation biases caused by head rotation and eye appearance distortion. Then each subproblem is solved by either learning-based method or geometric-based calculation. Specifically, the gaze estimation bias caused by eye appearance distortion is learnt effectively from a 5-seconds video clip. Extensive experiments were conducted to verify the effectiveness of the proposed approach. 1 Feng Lu 0005, Takahiro Okabe, Yusuke Sugano, Yoichi Sato 0001 |
BMVC | 1 |
| 2011 | Inferring human gaze from appearance via adaptive linear regressionabstractThe problem of estimating human gaze from eye appearance is regarded as mapping high-dimensional features to low-dimensional target space. Conventional methods require densely obtained training samples on the eye appearance manifold, which results in a tedious calibration stage. In this paper, we introduce an adaptive linear regression (ALR) method for accurate mapping via sparsely collected training samples. The key idea is to adaptively find the subset of training samples where the test sample is most linearly representable. We solve the problem via l1-optimization and thoroughly study the key issues to seek for the best solution for regression. The proposed gaze estimation approach based on ALR is naturally sparse and low-dimensional, giving the ability to infer human gaze from variant resolution eye images using much fewer training samples than existing methods. Especially, the optimization procedure in ALR is extended to solve the subpixel alignment problem simultaneously for low resolution test eye images. Performance of the proposed method is evaluated by extensive experiments against various factors such as number of training samples, feature dimensionality and eye image resolution to verify its effectiveness. Feng Lu 0005, Yusuke Sugano, Takahiro Okabe, Yoichi Sato 0001 |
ICCV | 1 |
| 2010 | Multi-View Stereo Reconstruction with High Dynamic Range Texture
Feng Lu 0005, Xiangyang Ji, Qionghai Dai, Guihua Er |
ACCV (2) | 1 |