Dakai Ren

dblp:285/6935 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0003-0442-6861ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Geometry-Insensitive RPN Prototypes for Domain Adaptive 3D Object Detection
abstract
The region proposal network (RPN) plays a critical role in object detection for a two-stage domain adaptive 3D object detector. However, current methods usually minimize the disparity between source and target domains by reducing the bias in intrinsic geometric information or by undertaking feature alignment according to the geometric disparity but ignore the transferability of RPN-related features and neglect the discriminability between foreground and background, resulting in generating low-quality RPN proposals. Thus, we propose a novel domain adaptation method to distinguish the discriminability between foreground and background. It could implicitly avoid the geometric disparity of objects in feature alignment. Specifically, we first construct learnable and geometry-insensitive foreground RPN prototype and background RPN prototype. Then, we enforce the foreground RPN features and background RPN features to align with the foreground RPN prototype and background RPN prototype, respectively. By this way, the distributional discrepancy is effectively decreased and the adaptability is promoted for existing 3D detectors. We demonstrate that our approach achieves promising results compared with other domain adaptation works on multiple cross-domain detection scenarios.
Jiazhong Chen, Dakai Ren, Zian Fu, Furui Liu
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Asymmetry-aware bilinear pooling in multi-modal data for head pose estimation
Jiazhong Chen, Dakai Ren, Hua Cao
Signal Process. Image Commun.3
2022 Multi-view facial action unit detection via DenseNets and CapsNets
Dakai Ren, Xiangming Wen, Jiazhong Chen, Shiqi Zhang 0002
Multim. Tools Appl.1
2021 Video Saliency Prediction via Deep Eye Movement Learning
abstract
Existing methods often utilize temporal motion information and spatial layout information in video to predict video saliency. However, the fixations are not always consistent with the moving object of interest, because human eye fixations are determined not only by the spatio-temporal information, but also by the velocity of eye movement. To address this issue, a new saliency prediction method via deep eye movement learning (EML) is proposed in this paper. Compared with previous methods that use human fixations as ground truth, our method uses the optical flow of fixations between successive frames as an extra ground truth for the purpose of eye movement learning. Experimental results on DHF1K, Hollywood2, and UCF-sports datasets show the proposed EML model achieves a promising result across a wide of metrics.
Jiazhong Chen, Jie Chen 0058, Dakai Ren, Shiqi Zhang 0002, Zongyi Li
MMAsia4
2021 Video saliency prediction via spatio-temporal reasoning
Jiazhong Chen, Zongyi Li, Yi Jin 0001, Dakai Ren
Neurocomputing4
2021 Audiovisual saliency prediction via deep learning
Jiazhong Chen, Dakai Ren, Ping Duan
Neurocomputing4
2021 Gaze estimation via bilinear pooling-based attention networks
Dakai Ren, Jiazhong Chen, Zhaoming Lu, Zongyi Li
J. Vis. Commun. Image Represent.1
2021 Saliency detection via cross-scale deep inference
Dakai Ren, Xiangming Wen, Jiazhong Chen, Zongyi Li
J. Vis. Commun. Image Represent.1
2020 Cross-modal learning for saliency prediction in mobile environment
abstract
The existing researches reveal that a significant impact is introduced by viewing conditions for visual perception when viewing media on mobile screens. This brings two issues in the area of visual saliency that we need to address: how the saliency models perform in mobile conditions, and how to consider the mobile conditions when designing a saliency model. To investigate the performance of saliency models in mobile environment, eye fixations in four typical mobile conditions are collected as the mobile ground truth in this work. To consider the mobile conditions when designing a saliency model, we combine viewing factors and visual stimuli as two modalities, and a cross-modal based deep learning architecture is proposed for visual attention prediction. Experimental results demonstrate the model with the consideration of mobile viewing factors often outperforms the models without such consideration.
Dakai Ren, Xiangming Wen, Xiaoya Liu, Jiazhong Chen
MMAsia1