EDBT 2026 Demo / reviewers in the wild / expert
Hao Chen 0034
dblp:175/3324-34
· DBLP profile ↗
32ranked-venue papers
12as first author
21since 2021 · last 2026
0000-0002-3138-505XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 14 since 2021Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIR-DR: Adaptive Image Retargeting with Instance Relocation and Dual-guidance RepaintingabstractImage retargeting aims to adjust the aspect ratio of images to accommodate various display devices. While existing methods consider both foreground semantics and background inpainting, their Seam-carving-based framework is inherently destructive, often compromising the structural integrity of foreground instances. Furthermore, conventional inpainting models struggle to achieve pixel-level accuracy with global-only guidance, leading to local inconsistencies and background distortions. To address these challenges, we reformulate image retargeting as a instance-level re-layout task. By Adaptive Instance Relocation and Dual-guidance Repainting (AIR-DR), our method preserves the structural integrity of the foreground and recovers the background with consistent details. Additionally, we introduce an adaptive retargeting decision that maintains robustness across challenging retargeting scenarios and any ratios. Extensive experiments on multiple public datasets across various aspect ratios demonstrate that our approach consistently outperforms existing methods in both objective metrics and subjective evaluations. Comprehensive ablation studies further validate the effectiveness of each component. Zhitong Dong, Yongjian Deng, Hao Chen 0034 |
AAAI | 4 |
| 2026 | Dissecting RGB-D Learning for Improved Multi-Modal FusionabstractIn the RGB-D vision community, extensive research has been focused on designing multi-modal learning strategies and fusion structures. However, the complementary and fusion mechanisms in RGB-D models remain a opaque box. In this paper, we present an analytical framework and a novel score to dissect the RGB-D vision community. Our approach involves measuring proposed semantic variance and feature similarity across modalities and levels, conducting visual and quantitative analyzes on multi-modal learning through comprehensive experiments. Specifically, we investigate the consistency and specialty of features across modalities, evolution rules within each modality, and the collaboration logic used when optimizing a RGB-D model. Our studies reveal/verify several important findings, such as the discrepancy in cross-modal features and the hybrid multi-modal cooperation rule, which highlights consistency and specialty simultaneously for complementary inference. We also showcase the versatility of the proposed RGB-D dissection method and introduce a straightforward fusion strategy based on our findings, which delivers significant enhancements across various tasks and even other multi-modal data. Hao Chen 0034, Yunshu Zhang, Zheng Lin 0005, Yongjian Deng |
IEEE Trans. Image Process. | 1 |
| 2026 | Multi-Task-Driven Adapter-Based Foundation Model for Locomotion Prediction in Virtual RealityabstractServing as a fundamental interaction in Virtual Reality, Locomotion technology defines how users navigate and explore the immersive virtual environments with high degree of freedom. Accurate prediction of locomotion not only enhances the sense of presence and ease of movement in virtual environment but also benefits VR applications through context-aware optimization such as pre-rendering and scene streaming. To leverage the superior understanding and causal modeling capabilities of Foundation Models (FMs) in the domain of numerical prediction, we apply FMs to time-series data, enabling more accurate estimation of users’ future spatial coordinates based on historical motion and gaze data. In this article, we introduce LoCoFoMo, an Adapter-based FM architecture specifically designed for future trajectory prediction in VR contexts. We conduct extensive experiments to evaluate the effectiveness of LoCoFoMo and its components. The proposed model demonstrates strong competitiveness when compared with several trajectory prediction and temporal reasoning models, evidenced by a performance gain of over 25% against the best baseline on datasets with interaction paradigms like Touchpad and Arm-Swing, coupled with superior stability in mitigating error accumulation. Ding Ding 0002, Hao Chen 0034 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | EvSAM: Segment Anything Model with Event-based AssistanceabstractThe general-purpose Segment Anything Model (SAM) is limited by the inherent constraints of RGB sensors, which render it inadequate for challenging real-world scenarios such as adverse lighting conditions and rapid motion. In contrast, event cameras, a novel type of bio-inspired visual sensor, offer distinct imaging advantages, including high temporal resolution and a high dynamic range. The event streams generated by these cameras provide spatiotemporal dynamic cues that are often absent in conventional image frames. To overcome the limitations of RGB-based models, we propose SAM with Event-based Assistance (EvSAM) , a novel RGB-event multi-modal semantic segmentation framework. EvSAM leverages the strong generalization capabilities of SAM while incorporating the complementary characteristics of event data to enhance scene comprehension, particularly under adverse conditions. To address the challenges of fusing two modals (image and event) with large data format discrepancy, we introduce two core components: the Multi-spatiotemporal-scale Patch Alignment Block (MS 2 PAB) and the Event-based Feature Injector (EFInj) for SAM. Specifically, the MS \({}^{2}\) PAB captures spatiotemporal semantic coherence from the event stream and transforms it into a frame-based complementary representation using a multi-spatiotemporal alignment strategy. The EFInj introduces a dynamic event feature update mechanism, wherein the fused features at a given layer guide the adaptive generation of deeper event representations. This process facilitates the integration of RGB spatial semantics with event-based motion cues. Owing to these core designs, EvSAM demonstrates superior performance on event-based semantic segmentation datasets, thereby fully validating its distinct advantages in handling extreme visual scenarios. Furthermore, we extend our model to the task of depth estimation, which further demonstrates its strong generalization ability and scalability for various downstream applications. Yuhan Liu 0021, Hao Chen 0034, Ding Ding 0002, Zhen Yang 0004, Youfu Li 0001, Yongjian Deng |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | Multimodal Large Language Model for Virtual Object GroundingabstractWe propose a novel task, Virtual Object Grounding (VOG) . It aims to predict plausible locations in an image for inserting virtual objects that align with a given textual description. This VOG task can address the challenge of providing region constraints for object insertion in image editing, thereby ensuring the consistency of irrelevant areas in the image. To support this task, we construct Virtual Segmentation (VirtualSeg) dataset, a dataset of over 92,000 samples automatically generated from VrR-VG via a four-step dataset construction pipeline. This pipeline employs CLIP to automatically filter out low-quality data samples, ensuring the quality of VirtualSeg. Furthermore, we propose the VirLLaVA model, a novel VOG framework built upon LLaVA-7B. By equipping the MLLM backbone with two sequences of learnable tokens and a dual grounding module, and by guiding the model during training to learn step-by-step how to locate virtual objects, our method enables it to reason about their positions from textual and visual inputs. Experiments show that VirLLaVA significantly improves performance in VOG, while also offering a promising direction for consistent and automated image editing. The code and dataset are available at https://github.com/Royxia0818/MLLM_for_VOG . Ziheng Xia, Ding Ding 0002, Hao Chen 0034 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Know Where You Are From: Event-Based Segmentation via Spatio-Temporal PropagationabstractEvent cameras have gained attention in segmentation due to their higher temporal resolution and dynamic range compared to traditional cameras. However, they struggle with issues like lack of color perception and triggering only at motion edges, making it hard to distinguish objects with similar contours or segment spatially continuous objects. Our work aims to address these often overlooked issues. Based on the assumption that various objects exhibit different motion patterns, we believe that embedding the historical motion states of objects into segmented scenes can effectively address these challenges. Inspired by this, we propose the ESS framework ``Know Where You Are From" (KWYAF), which incorporates past motion cues through spatio-temporal propagation embedding. This framework features two core components: the Sequential Motion Encoding Module (SME) and the Event-Based Reliable Region Selection Mechanism (ER²SM). SMEs construct prior motion features through spatio-temporal correlation modeling for boosting final segmentation, while ER²SM adapts to identify high-confidence regions, embedding motion more precisely through local window masks and reliable region selection. A large number of experiments have demonstrated the effectiveness of our proposed framework in terms of both quantity and quality. Gengyu Lyu, Hao Chen 0034, Bochen Xie, Zhen Yang 0004, Youfu Li 0001, Yongjian Deng |
AAAI | 3 |
| 2025 | ESEG: Event-Based Segmentation Boosted by Explicit Edge-Semantic GuidanceabstractEvent-based semantic segmentation (ESS) has attracted researchers' attention recently, as event cameras can solve problems such as under/over-exposure or motion blur that are difficult for RGB cameras to handle. However, event data are noisy and sparse, resulting in difficulties for the model to locate and extract reliable cues from their sparse representations, especially when performing pixel-level tasks. In this paper, we propose a novel framework ESEG to alleviate the dilemma. Given that event signals relate closely to moving edges, instead of proposing complex structures to expect them to recognize those reliable edge regions behind event signals on their own, we introduce the explicit edge-semantic supervision as a reference to let the ESS model globally optimize semantics, considering the high confidence of event data in edge regions. In addition, we propose a fusion module named Density-Aware Dynamic-Window Cross Attention Fusion (D\textsuperscript{2}CAF), in which the density perception, cross-attention, and dynamic window masking mechanisms are jointly imposed to optimize edge-dense feature fusion, leveraging the characteristics of event cameras. Experimental results on DSEC and DDD17 datasets demonstrate the efficacy of the ESEG framework and its core designs. Gengyu Lyu, Hao Chen 0034, Zhen Yang 0004, Yongjian Deng |
AAAI | 5 |
| 2025 | Separation for Better Integration: Disentangling Edge and Motion in Event-Based Deblurring
Hao Chen 0034, Yongjian Deng |
ICCV | 2 |
| 2025 | Improving Multimodal Learning Balance and Sufficiency through Data RemixingabstractDifferent modalities hold considerable gaps in optimization trajectories, including speeds and paths, which lead to *modality laziness* and *modality clash* when jointly training multimodal models, resulting in insufficient and imbalanced multimodal learning.
Existing methods focus on enforcing the weak modality by adding modality-specific optimization objectives, aligning their optimization speeds, or decomposing multimodal learning to enhance unimodal learning. These methods fail to achieve both unimodal sufficiency and multimodal balance.
In this paper, we, for the first time, address both concerns by proposing multimodal Data Remixing, including decoupling multimodal data and filtering hard samples for each modality to mitigate modality imbalance; and then batch-level reassembling to align the gradient directions and avoid cross-modal interference, thus enhancing unimodal learning sufficiency.
Experimental results demonstrate that our method can be seamlessly integrated with existing approaches, improving accuracy by approximately **6.50\%$\uparrow$** on CREMAD and **3.41\%$\uparrow$** on Kinetic-Sounds, without training set expansion or additional computational overhead during inference. The source code is available at Data Remixing. Hao Chen 0034, Yongjian Deng |
ICML | 2 |
| 2025 | EPA: Boosting Event-based Video Frame Interpolation with Perceptually Aligned LearningabstractEvent cameras, with their capacity to provide high temporal resolution information between frames, are increasingly utilized for video frame interpolation (VFI) in challenging scenarios characterized by high-speed motion and significant occlusion. However, prevalent issues of blur and distortion within the keyframes and ground truth data used for training and inference in these demanding conditions are frequently overlooked. This oversight impedes the perceptual realism and multi-scene generalization capabilities of existing event-based VFI (E-VFI) methods when generating interpolated frames. Motivated by the observation that semantic-perceptual discrepancies between degraded and pristine images are considerably smaller than their image-level differences, we introduce EPA. This novel E-VFI framework diverges from approaches reliant on direct image-level supervision by constructing multilevel, degradation-insensitive semantic perceptual supervisory signals to enhance the perceptual realism and multi-scene generalization of the model's predictions. Specifically, EPA operates in two phases: it first employs a DINO-based perceptual extractor, a customized style adapter, and a reconstruction generator to derive multi-layered, degradation-insensitive semantic-perceptual features ($\mathcal{S}$). Second, a novel Bidirectional Event-Guided Alignment (BEGA) module utilizes deformable convolutions to align perceptual features from keyframes to ground truth with inter-frame temporal guidance extracted from event signals. By decoupling the learning process from direct image-level supervision, EPA enhances model robustness against degraded keyframes and unreliable ground truth information. Extensive experiments demonstrate that this approach yields interpolated frames more consistent with human perceptual preferences. *The code will be released upon acceptance.* Yuhan Liu 0021, Linghui Fu, Zhen Yang 0004, Hao Chen 0034, Youfu Li 0001, Yongjian Deng |
NeurIPS | 4 |
| 2025 | Event-based video interpolation via complementary motion information
Yuhan Liu 0021, Linghui Fu, Hao Chen 0034, Zhen Yang 0004, Youfu Li 0001, Yongjian Deng |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | A Dynamic GCN with Cross-Representation Distillation for Event-Based LearningabstractRecent advances in event-based research prioritize sparsity and temporal precision. Approaches learning sparse point-based representations through graph CNNs (GCN) become more popular. Yet, these graph techniques hold lower performance than their frame-based counterpart due to two issues: (i) Biased graph structures that don't properly incorporate varied attributes (such as semantics, and spatial and temporal signals) for each vertex, resulting in inaccurate graph representations. (ii) A shortage of robust pretrained models. Here we solve the first problem by proposing a new event-based GCN (EDGCN), with a dynamic aggregation module to integrate all attributes of vertices adaptively. To address the second problem, we introduce a novel learning framework called cross-representation distillation (CRD), which leverages the dense representation of events as a cross-representation auxiliary to provide additional supervision and prior knowledge for the event graph. This frame-to-graph distillation allows us to benefit from the large-scale priors provided by CNNs while still retaining the advantages of graph-based models. Extensive experiments show our model and learning framework are effective and generalize well across multiple vision tasks. Yongjian Deng, Hao Chen 0034, Youfu Li 0001 |
AAAI | 2 |
| 2024 | Video Frame Interpolation via Direct Synthesis with the Event-based ReferenceabstractVideo Frame Interpolation (VFI) has witnessed a surge in popularity due to its abundant downstream applications. Event-based VFI (E-VFI) has recently propelled the ad-vancement of VFI. Thanks to the high temporal resolution benefits, event cameras can bridge the informational void present between successive video frames. Most state-of-the-art E-VFI methodologies follow the conventional VFI paradigm, which pivots on motion estimation between consecutive frames to generate intermediate frames through a process of warping and refinement. However, this reliance engenders a heavy dependency on the quality and consis-tency of keyframes, rendering these methods susceptible to challenges in extreme real-world scenarios, such as missing moving objects and severe occlusion dilemmas. This study proposes a novel E-VFI framework that directly synthesize intermediate frames leveraging event-based reference, obviating the necessity for explicit motion estimation and substantially enhancing the capacity to handle motion occlusion. Given the sparse and inher-ently noisy nature of event data, we prioritize the relia-bility of the event-based reference, leading to the development of an innovative event-aware reconstruction strategy for accurate reference generation. Besides, we implement a bi-directional event-guided alignment from keyframes to the reference using the introduced E-PCD module. Finally, a transformer-based decoder is adopted for prediction re-finement. Comprehensive experimental evaluations on both synthetic and real-world datasets underscore the superiority of our approach and its potential to execute high-quality VFI tasks. Yuhan Liu 0021, Yongjian Deng, Hao Chen 0034, Zhen Yang 0004 |
CVPR | 3 |
| 2024 | SAM-Event-Adapter: Adapting Segment Anything Model for Event-RGB Semantic SegmentationabstractSemantic segmentation, a fundamental visual task ubiquitously employed in sectors ranging from transportation and robotics to healthcare, has always captivated the research community. In the wake of rapid advancements in large model research, the foundation model for semantic segmentation tasks, termed the Segment Anything Model (SAM), has been introduced. This model substantially addresses the dilemma of poor generalizability of previous segmentation models and the disadvantage in requiring to retrain the whole model on variant datasets. Nonetheless, segmentation models developed on SAM remain constrained by the inherent limitations of RGB sensors, particularly in scenarios characterized by complex lighting conditions and high-speed motion. Motivated by these observations, a natural recourse is to adapt SAM to additional visual modalities without compromising its robust generalizability. To achieve this, we introduce a lightweight SAM-Event-Adapter (SE-Adapter) module, which incorporates event camera data into a cross-modal learning architecture based on SAM, with only limited tunable parameters incremental. Capitalizing on the high dynamic range and temporal resolution afforded by event cameras, our proposed multi-modal Event-RGB learning architecture effectively augments the performance of semantic segmentation tasks. In addition, we propose a novel paradigm for representing event data in a patch format compatible with transformer-based models, employing multi-spatiotemporal scale encoding to efficiently extract motion and semantic correlations from event representations. Exhaustive empirical evaluations conducted on the DSEC-Semantic and DDD17 datasets provide validation of the effectiveness and rationality of our proposed approach. Yongjian Deng, Yuhan Liu 0021, Hao Chen 0034, Youfu Li 0001, Zhen Yang 0004 |
ICRA | 4 |
| 2024 | A Motion-aware Spatio-temporal Graph for Video Salient Object RankingabstractVideo salient object ranking aims to simulate the human attention mechanism by dynamically prioritizing the visual attraction of objects in a scene over time. Despite its numerous practical applications, this area remains underexplored. In this work, we propose a graph model for video salient object ranking. This graph simultaneously explores multi-scale spatial contrasts and intra-/inter-instance temporal correlations across frames to extract diverse spatio-temporal saliency cues. It has two advantages: 1. Unlike previous methods that only perform global inter-frame contrast or compare all proposals across frames globally, we explicitly model the motion of each instance by comparing its features with those in the same spatial region in adjacent frames, thus obtaining more accurate motion saliency cues. 2. We synchronize the spatio-temporal saliency cues in a single graph for joint optimization, which exhibits better dynamics compared to the previous stage-wise methods that prioritize spatial cues followed by temporal cues. Additionally, we propose a simple yet effective video retargeting method based on video saliency ranking. Extensive experiments demonstrate the superiority of our model in video salient object ranking and the effectiveness of the video retargeting method. Our codes/models are released at [https://github.com/zyf-815/VSOR/tree/main](https://github.com/zyf-815/VSOR/tree/main). Hao Chen 0034, Yongjian Deng |
NeurIPS | 1 |
| 2024 | Prune and Repaint: Content-Aware Image Retargeting for any RatioabstractImage retargeting is the task of adjusting the aspect ratio of images to suit different display devices or presentation environments. However, existing retargeting methods often struggle to balance the preservation of key semantics and image quality, resulting in either deformation or loss of important objects, or the introduction of local artifacts such as discontinuous pixels and inconsistent regenerated content. To address these issues, we propose a content-aware retargeting method called PruneRepaint. It incorporates semantic importance for each pixel to guide the identification of regions that need to be pruned or preserved in order to maintain key semantics. Additionally, we introduce an adaptive repainting module that selects image regions for repainting based on the distribution of pruned pixels and the proportion between foreground size and target aspect ratio, thus achieving local smoothness after pruning. By focusing on the content and structure of the foreground, our PruneRepaint approach adaptively avoids key content loss and deformation, while effectively mitigating artifacts with local repainting. We conduct experiments on the public RetargetMe benchmark and demonstrate through objective experimental results and subjective user studies that our method outperforms previous approaches in terms of preserving semantics and aesthetics, as well as better generalization across diverse aspect ratios. Codes will be available at
https://github.com/fhshen2022/PruneRepaint. Feihong Shen, Yifeng Geng, Yongjian Deng, Hao Chen 0034 |
NeurIPS | 5 |
| 2024 | Disentangled Cross-Modal Transformer for RGB-D Salient Object Detection and BeyondabstractPrevious multi-modal transformers for RGB-D salient object detection (SOD) generally directly connect all patches from two modalities to model cross-modal correlation and perform multi-modal combination without differentiation, which can lead to confusing and inefficient fusion. Instead, we disentangle the cross-modal complementarity from two views to reduce cross-modal fusion ambiguity: 1) Context disentanglement. We argue that modeling long-range dependencies across modalities as done before is uninformative due to the severe modality gap. Differently, we propose to disentangle the cross-modal complementary contexts to intra-modal self-attention to explore global complementary understanding, and spatial-aligned inter-modal attention to capture local cross-modal correlations, respectively. 2) Representation disentanglement. Unlike previous undifferentiated combination of cross-modal representations, we find that cross-modal cues complement each other by enhancing common discriminative regions and mutually supplement modal-specific highlights. On top of this, we divide the tokens into consistent and private ones in the channel dimension to disentangle the multi-modal integration path and explicitly boost two complementary ways. By progressively propagate this strategy across layers, the proposed Disentangled Feature Pyramid module (DFP) enables informative cross-modal cross-level integration and better fusion adaptivity. Comprehensive experiments on a large variety of public datasets verify the efficacy of our context and representation disentanglement and the consistent improvement over state-of-the-art models. Additionally, our cross-modal attention hierarchy can be plug-and-play for different backbone architectures (both transformer and CNN) and downstream tasks, and experiments on a CNN-based model and RGB-D semantic segmentation verify this generalization ability. Hao Chen 0034, Feihong Shen, Ding Ding 0002, Yongjian Deng |
IEEE Trans. Image Process. | 1 |
| 2022 | A Voxel Graph CNN for Object Classification with Event CamerasabstractEvent cameras attract researchers' attention due to their low power consumption, high dynamic range, and extremely high temporal resolution. Learning models on event-based object classification have recently achieved massive success by accumulating sparse events into dense frames to apply traditional 2D learning methods. Yet, these approaches necessitate heavy-weight models and are with high computational complexity due to the redundant information introduced by the sparse-to-dense conversion, limiting the potential of event cameras on real-life applications. This study aims to address the core problem of balancing accuracy and model complexity for event-based classification models. To this end, we introduce a novel graph representation for event data to exploit their sparsity better and customize a lightweight voxel graph convolutional neural network (EV-VGCNN) for event-based classification. Specifically, (1) using voxel-wise vertices rather than previous point-wise inputs to explicitly exploit regional 2D semantics of event streams while keeping the sparsity; (2) proposing a multi-scale feature relational layer (MFRL) to extract spatial and motion cues from each vertex discriminatively concerning its distances to neighbors. Comprehensive experiments show that our model can advance state-of-the-art classification accuracy with extremely low model complexity (merely 0.84M parameters). Yongjian Deng, Hao Chen 0034, Hai Liu 0004, Youfu Li 0001 |
CVPR | 2 |
| 2022 | MVF-Net: A Multi-View Fusion Network for Event-Based Object ClassificationabstractEvent-based object recognition has drawn increasing attention for event cameras’ distinguished advantages of low power consumption and high dynamic range. For this new modality, previous works based on customizing low-level descriptors are vulnerable to noise and with limited generalizability. Although recent works turn to design various deep neural networks to extract event features, they either suffer from data insufficiency to fully train the event-based model or fail to encode spatial and temporal cues simultaneously with their single view network. In this work, we address these limitations by proposing a multi-view attention-aware network, in which an event stream is projected to multi-view 2D maps to utilize well-trained 2D models and explore spatio-temporal complements. Besides, the attention mechanism is used to boost the complements in different streams for better joint inference. Comprehensive experiments show the large superiority of our model over state-of-the-art methods as well as the efficacy of our multi-view fusion framework for event data. Yongjian Deng, Hao Chen 0034, Youfu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | CNN-Based RGB-D Salient Object Detection: Learn, Select, and Fuse
Hao Chen 0034, Youfu Li 0001, Yongjian Deng, Guosheng Lin |
Int. J. Comput. Vis. | 1 |
| 2021 | Learning From Images: A Distillation Learning Framework for Event CamerasabstractEvent cameras have recently drawn massive attention in the computer vision community because of their low power consumption and high response speed. These cameras produce sparse and non-uniform spatiotemporal representations of a scene. These characteristics of representations make it difficult for event-based models to extract discriminative cues (such as textures and geometric relationships). Consequently, event-based methods usually perform poorly compared to their conventional image counterparts. Considering that traditional images and event signals share considerable visual information, this paper aims to improve the feature extraction ability of event-based models by using knowledge distilled from the image domain to additionally provide explicit feature-level supervision for the learning of event data. Specifically, we propose a simple yet effective distillation learning framework, including multi-level customized knowledge distillation constraints. Our framework can significantly boost the feature extraction process for event data and is applicable to various downstream tasks. We evaluate our framework on high-level and low-level tasks, i.e., object classification and optical flow prediction. Experimental results show that our framework can effectively improve the performance of event-based models on both tasks by a large margin. Furthermore, we present a 10K dataset (CEP-DVS) for event-based object classification. This dataset consists of samples recorded under random motion trajectories that can better evaluate the motion robustness of the event-based model and is compatible with multi-modality vision tasks. Yongjian Deng, Hao Chen 0034, Youfu Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Discriminative Cross-Modal Transfer Learning and Densely Cross-Level Feedback Fusion for RGB-D Salient Object DetectionabstractThis article addresses two key issues in RGB-D salient object detection based on the convolutional neural network (CNN). 1) How to bridge the gap between the "data-hungry" nature of CNNs and the insufficient labeled training data in the depth modality? 2) How to take full advantages of the complementary information among two modalities. To solve the first problem, we model the depth-induced saliency detection as a CNN-based cross-modal transfer learning problem. Instead of directly adopting the RGB CNN as initialization, we additionally train a modality classification network (MCNet) to encourage discriminative modality-specific representations in minimizing the modality classification loss. To solve the second problem, we propose a densely cross-level feedback topology, in which the cross-modal complements are combined in each level and then densely fed back to all shallower layers for sufficient cross-level interactions. Compared to traditional two-stream frameworks, the proposed one can better explore, select, and fuse cross-modal cross-level complements. Experiments show the significant and consistent improvements of the proposed CNN framework over other state-of-the-art methods. Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Cross-Validated Locally Polynomial Modeling for 2-D/3-D Gaze Tracking With Head-Worn DevicesabstractIn the context of wearable gaze tracking techniques, the problems of two-dimensional (2-D) and three-dimensional (3-D) gaze estimation can be viewed as inferring 2-D epipolar lines and 3-D visual axes from eye monitoring cameras. To this end, in this article, a simple local polynomial model is proposed to back-project a pupil center onto its corresponding visual axis. Based on this approximation, a homographylike relation is derived in a local manner, and via the Leave-One-Out cross-validation criterion, training gaze samples at one certain depth is leveraged to partition entire input space into multiple overlapping subregions. Then, the gaze data at another depth are utilized to recover the epipolar point, i.e., the image eyeball center. Thus, given a pupil image, the corresponding epipolar line can be determined by the resolved homographylike mapping and the epipolar point. By using the same partition structure, 3-D gaze prediction model can be inferred by solving a nonlinear optimization problem, which aims to minimize the angular disparities between training visual directions and prediction ones. Meanwhile, it is necessary to form a good starting point and suitable constraints for the optimization problem. Otherwise, it may end up with trivial solutions, i.e., faraway eye positions. To facilitate the practical implementation of our proposed method, we also analyze how the spatial distribution of calibration points impacts the model learning accuracy. The experiment results justify the effectiveness of our proposed gaze estimation method for both the normal vision and eyewear users. Dan Su 0001, Youfu Li 0001, Hao Chen 0034 |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | RGBD Salient Object Detection via Disentangled Cross-Modal FusionabstractDepth is beneficial for salient object detection (SOD) for its additional saliency cues. Existing RGBD SOD methods focus on tailoring complicated cross-modal fusion topologies, which although achieve encouraging performance, are with a high risk of over-fitting and ambiguous in studying cross-modal complementarity. Different from these conventional approaches combining cross-modal features entirely without differentiating, we concentrate our attention on decoupling the diverse cross-modal complements to simplify the fusion process and enhance the fusion sufficiency. We argue that if cross-modal heterogeneous representations can be disentangled explicitly, the cross-modal fusion process can hold less uncertainty, while enjoying better adaptability. To this end, we design a disentangled cross-modal fusion network to expose structural and content representations from both modalities by cross-modal reconstruction. For different scenes, the disentangled representations allow the fusion module to easily identify, and incorporate desired complements for informative multi-modal fusion. Extensive experiments show the effectiveness of our designs and a large outperformance over state-of-the-art methods. Hao Chen 0034, Yongjian Deng, Youfu Li 0001, Tzu-Yi Hung, Guosheng Lin |
IEEE Trans. Image Process. | 1 |
| 2019 | Region-wise Polynomial Regression for 3D Mobile Gaze EstimationabstractIn the context of mobile gaze tracking techniques, a 3D gaze point can be calculated as the middle point between two 3D visual axes. To infer gaze directions and eyeball positions, a nonlinear optimization problem is typically formulated to minimize the angular disparities between the training gaze directions and prediction ones. Nonetheless, the experimental results reported by some previous works show that this kind of approaches are very likely to yield large prediction errors hence considered less useful for human-machine interactions. In this study, we aim to address this widespread issue in three aspects. At first, instead of using a global regression model, a simple local polynomial model is proposed to back-project a pupil center onto its corresponding visual axis. Based on the Leave-One-Out cross-validation criterion, the partition structure is automatically learned in the process of resolving a homography-like relationship. Secondly, a good starting point for nonlinear-optimization is obtained by the image eyeball center, which can be estimated by systematic parallax errors. Meanwhile, it is necessary to add the suitable constraints for 3D eye positions. Otherwise, the optimization may end up with trivial solutions, i.e., faraway eye positions. Thirdly, we explore a strategy for designing the spatial distribution of calibration points in a principled manner. The experiment results demonstrate that an encouraging gaze estimation accuracy can be achieved by our proposed framework for both the normal vision and eyewear users. Dan Su 0001, Youfu Li 0001, Hao Chen 0034 |
IROS | 3 |
| 2019 | Multi-modal fusion network with multi-scale multi-path and cross-modal interactions for RGB-D salient object detection
Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
Pattern Recognit. | 1 |
| 2019 | Toward Precise Gaze Estimation for Mobile Head-Mounted Gaze Tracking SystemsabstractThe gaze estimation in the mobile scenario often suffers from the extrapolation and parallax errors. In this paper, we propose a novel calibration framework to achieve the precise gaze estimation for head-mounted gaze trackers. Our proposed framework consists of two steps to learn a point-to-point and a point-to-line relations, respectively. The aim of step I is to infer the relation between pupil centers and spatially constrained points of regard. By adopting the “CalibMe” gaze data acquisition method, a sparse Gaussian Process using pseudo-inputs is used to capture the smooth residual field unmodeled by the polynomial function. Meanwhile, a distraction detection criterion is introduced to identify the moment when user's attention is taken away from the calibration point thereby removing outliers. By combining with the point-to-point relation inferred in step I, the observed parallax errors are leveraged in step II to obtain a point-to-line relation, i.e., each pupil center will correspond to an epipolar line. Thus, the real image gaze point projected from different depths is predicted as the intersection of two epipolar lines inferred from binocular data. The simulation and experimental results show the effectiveness of our proposed calibration framework for head-mounted gaze trackers. Dan Su 0001, Youfu Li 0001, Hao Chen 0034 |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Three-Stream Attention-Aware Network for RGB-D Salient Object DetectionabstractPrevious RGB-D fusion systems based on convolutional neural networks (CNNs) typically employ a two-stream architecture, in which RGB and depth inputs are learnt independently. The multi-modal fusion stage is typically performed by concatenating the deep features from each stream in the inference process. The traditional two-stream architecture might experience insufficient multi-modal fusion due to two following limitations: (1) The cross-modal complementarity is rarely studied in the bottom-up path, wherein we believe the crossmodal complements can be combined to learn new discriminative features to enlarge the RGB-D representation community; (2) The cross-modal channels are typically combined by undifferentiated concatenation, which appears ambiguous to select cross-modal complementary features. In this work, we address these two limitations by proposing a novel three-stream attention-aware multi-modal fusion network. In the proposed architecture, a cross-modal distillation stream, accompanying the RGB-specific and depth-specific streams, is introduced to extract new RGB-D features in each level in the bottom-up path. Furthermore, the channel-wise attention mechanism is innovatively introduced to the cross-modal cross-level fusion problem to adaptively select complementary feature maps from each modality in each level. Extensive experiments report the effectiveness of the proposed architecture and the significant improvement over state-of-theart RGB-D salient object detection methods. Hao Chen 0034, Youfu Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2018 | Progressively Complementarity-Aware Fusion Network for RGB-D Salient Object DetectionabstractHow to incorporate cross-modal complementarity sufficiently is the cornerstone question for RGB-D salient object detection. Previous works mainly address this issue by simply concatenating multi-modal features or combining unimodal predictions. In this paper, we answer this question from two perspectives: (1) We argue that if the complementary part can be modelled more explicitly, the cross-modal complement is likely to be better captured. To this end, we design a novel complementarity-aware fusion (CA-Fuse) module when adopting the Convolutional Neural Network (CNN). By introducing cross-modal residual functions and complementarity-aware supervisions in each CA-Fuse module, the problem of learning complementary information from the paired modality is explicitly posed as asymptotically approximating the residual function. (2) Exploring the complement across all the levels. By cascading the CA-Fuse module and adding level-wise supervision from deep to shallow densely, the cross-level complement can be selected and combined progressively. The proposed RGB-D fusion network disambiguates both cross-modal and cross-level fusion processes and enables more sufficient fusion results. The experiments on public datasets show the effectiveness of the proposed CA-Fuse module and the RGB-D salient object detection network. Hao Chen 0034, Youfu Li 0001 |
CVPR | 1 |
| 2018 | Attention-Aware Cross-Modal Cross-Level Fusion Network for RGB-D Salient Object DetectionabstractConvolutional neural networks have achieved wide success in RGB saliency detection. Recently, the advent of RGB-D sensors such as Kinect provide additional geometric saliency cues. However, the key challenge for RGB-D salient object detection that how to fuse RGB and depth information sufficiently is still under-studied. Traditional works mainly follow the two-stream architecture and combine RGB and depth features/decisions in an early or late point. The multi-modal fusion stage is performed by directly concatenating the features from two modalities without selection. In this work, we address this question by proposing a novel network with a distinguished insight: A selection module is significantly helpful for more informative and sufficient cross-modal cross-level combination. To this end, we introduce a top-down RGB-D fusion network which integrates an attention-aware cross-modal cross-level fusion block in each level to select discriminative features from each level and each modality. Extensive experiments on public datasets show that the proposed network is able to solve the key problems in RGB-D fusion and achieves state-of-the-art performance on RGB-D salient object detection. Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
IROS | 1 |
| 2017 | RGB-D Saliency Detection by Multi-stream Late Fusion Network
Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
ICVS | 1 |
| 2017 | M3Net: Multi-scale multi-path multi-modal fusion network and example application to RGB-D salient object detectionabstractFusing RGB and depth data is compelling in boosting performance for various robotic and computer vision tasks. Typically, the streams of RGB and depth information are merged into a single fusion point in an early or late stage to generate combined features or decisions. The single fusion point also means single fusion path, which is congested and inflexible to fuse all the information from different modalities. As a result, the fusion process is brute-force and consequently insufficient. To address this problem, we propose a multi-scale multi-path multi-modal fusion network (M3Net), in which the fusion path is scattered to diversify the contributions of each modality from global and local perspectives. Specially, the CNN streams of each modality are fused with a global understanding path and meanwhile a local capturing path. By filtering and regulating information flow in a multi-path way, the M3Net is equipped with more adaptive and flexible fusion mechanism, thus easing the gradient-based learning process, improving the directness and transparency of the fusion process and simultaneously facilitating the fusion process with multi-scale perspectives. Comprehensive experiments demonstrate the significant and consistent improvements of the proposed approach over state-of-the-art methods. Hao Chen 0034, Youfu Li 0001, Dan Su 0001 |
IROS | 1 |