Ruicong Liu

dblp:184/6964 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-8460-8763ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Leveraging RGB Images for Pre-training of Event-Based Hand Pose Estimation
Ruicong Liu, Takehiko Ohkawa, Tze Ho Elden Tse, Mingfang Zhang 0002, Angela Yao, Yoichi Sato 0001
ICPR (8)1
2025 Egocentric Action-Aware Inertial Localization in Point Clouds with Vision-Language Guidance
abstract
This paper presents a novel inertial localization framework named Egocentric Action-aware Inertial Localization (EAIL), which leverages egocentric action cues from head-mounted IMU signals to localize the target individual within a 3D point cloud. Human inertial localization is challenging due to IMU sensor noise that causes trajectory drift over time. The diversity of human actions further complicates IMU signal processing by introducing various motion patterns. Nevertheless, we observe that some actions captured by the head-mounted IMU correlate with spatial environmental structures (e.g., bending down to look inside an oven, washing dishes next to a sink), thereby serving as spatial anchors to compensate for the localization drift. The proposed EAIL framework learns such correlations via hierarchical multi-modal alignment with vision-language guidance. By assuming that the 3D point cloud of the environment is available, it contrastively learns modality encoders that align short-term egocentric action cues in IMU signals with local environmental features in the point cloud. The learning process is enhanced using concurrently collected vision and language signals to improve multimodal alignment. The learned encoders are then used in reasoning the IMU data and the point cloud over time and space to perform inertial localization. Interestingly, these encoders can further be utilized to recognize the corresponding sequence of actions as a by-product. Extensive experiments demonstrate the effectiveness of the proposed framework over state-of-the-art inertial localization and inertial action recognition baselines.
Mingfang Zhang 0002, Ryo Yonetani, Yifei Huang 0002, Liangyang Ouyang, Ruicong Liu, Yoichi Sato 0001
ICCV5
2025 From Gaze Jitter to Domain Adaptation: Generalizing Gaze Estimation by Manipulating High-Frequency Components
Ruicong Liu, Haofei Wang 0001, Feng Lu 0005
Int. J. Comput. Vis.1
2024 UVAGaze: Unsupervised 1-to-2 Views Adaptation for Gaze Estimation
abstract
Gaze estimation has become a subject of growing interest in recent research. Most of the current methods rely on single-view facial images as input. Yet, it is hard for these approaches to handle large head angles, leading to potential inaccuracies in the estimation. To address this issue, adding a second-view camera can help better capture eye appearance. However, existing multi-view methods have two limitations. 1) They require multi-view annotations for training, which are expensive. 2) More importantly, during testing, the exact positions of the multiple cameras must be known and match those used in training, which limits the application scenario. To address these challenges, we propose a novel 1-view-to-2-views (1-to-2 views) adaptation solution in this paper, the Unsupervised 1-to-2 Views Adaptation framework for Gaze estimation (UVAGaze). Our method adapts a traditional single-view gaze estimator for flexibly placed dual cameras. Here, the "flexibly" means we place the dual cameras in arbitrary places regardless of the training data, without knowing their extrinsic parameters. Specifically, the UVAGaze builds a dual-view mutual supervision adaptation strategy, which takes advantage of the intrinsic consistency of gaze directions between both views. In this way, our method can not only benefit from common single-view pre-training, but also achieve more advanced dual-view gaze estimation. The experimental results show that a single-view estimator, when adapted for dual views, can achieve much higher accuracy, especially in cross-dataset settings, with a substantial improvement of 47.0%. Project page: https://github.com/MickeyLLG/UVAGaze.
Ruicong Liu, Feng Lu 0005
AAAI1
2024 Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
abstract
The pursuit of accurate 3D hand pose estimation stands as a keystone for understanding human activity in the realm of egocentric vision. The majority of existing estimation methods still rely on single-view images as input, leading to potential limitations, e.g., limited field-of-view and ambiguity in depth. To address these problems, adding another camera to better capture the shape of hands is a practi-cal direction. However, existing multi-view hand pose estimation methods suffer from two main drawbacks: 1) Re-quiring multi-view annotations for training, which are ex-pensive. 2) During testing, the model becomes inapplica-ble if camera parameters/layout are not the same as those used in training. In this paper, we propose a novel Single-to-Dual-view adaptation (S2DHand) solution that adapts a pretrained single-view estimator to dual views. Compared with existing multi-view training methods, 1) our adaptation process is unsupervised, eliminating the need for multi-view annotation. 2) Moreover, our method can handle arbitrary dual-view pairs with unknown camera parameters, making the model applicable to diverse camera settings. Specifically, S2DHand is built on certain stereo constraints, including pairwise cross-view consensus and invariance of transformation between both views. These two stereo constraints are used in a complementary manner to gen-erate pseudo-labels, allowing reliable adaptation. Evalu-ation results reveal that S2DHand achieves significant improvements on arbitrary camera pairs under both in-dataset and cross-dataset settings, and outperforms existing adaptation methods with leading performance. Project page: https://github.com/ut-vision/S2DHand.
Ruicong Liu, Takehiko Ohkawa, Mingfang Zhang 0002, Yoichi Sato 0001
CVPR1
2024 ActionVOS: Actions as Prompts for Video Object Segmentation
Liangyang Ouyang, Ruicong Liu, Yifei Huang 0002, Ryosuke Furuta, Yoichi Sato 0001
ECCV (10)2
2024 Masked Video and Body-Worn IMU Autoencoder for Egocentric Action Recognition
Mingfang Zhang 0002, Yifei Huang 0002, Ruicong Liu, Yoichi Sato 0001
ECCV (18)3
2024 PnP-GA+: Plug-and-Play Domain Adaptation for Gaze Estimation Using Model Variants
abstract
Appearance-based gaze estimation has garnered increasing attention in recent years. However, deep learning-based gaze estimation models still suffer from suboptimal performance when deployed in new domains, e.g., unseen environments or individuals. In our previous work, we took this challenge for the first time by introducing a plug-and-play method (PnP-GA) to adapt the gaze estimation model to new domains. The core concept of PnP-GA is to leverage the diversity brought by a group of model variants to enhance the adaptability to diverse environments. In this article, we propose the PnP-GA+ by extending our approach to explore the impact of assembling model variants using three additional perspectives: color space, data augmentation, and model structure. Moreover, we propose an intra-group attention module that dynamically optimizes pseudo-labeling during adaptation. Experimental results demonstrate that by directly plugging several existing gaze estimation networks into the PnP-GA+ framework, it outperforms state-of-the-art domain adaptation approaches on four standard gaze domain adaptation tasks on public datasets. Our method consistently enhances cross-domain performance, and its versatility is improved through various ways of assembling the model group.
Ruicong Liu, Yunfei Liu 0001, Haofei Wang 0001, Feng Lu 0005
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Generalizing Gaze Estimation with Outlier-guided Collaborative Adaptation
abstract
Deep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons. In this paper, we propose a plug-and-play gaze adaptation framework (PnP-GA), which is an ensemble of networks that learn collaboratively with the guidance of outliers. Since our proposed framework does not require ground-truth labels in the target domain, the existing gaze estimation networks can be directly plugged into PnPGA and generalize the algorithms to new domains. We test PnP-GA on four gaze domain adaptation tasks, ETH-to-MPII, ETH-to-EyeDiap, Gaze360-to-MPII, and Gaze360to-EyeDiap. The experimental results demonstrate that the PnP-GA framework achieves considerable performance improvements of 36.9%, 31.6%, 19.4%, and 11.8% over the baseline system. The proposed framework also outperforms the state-of-the-art domain adaptation approaches on gaze domain adaptation tasks. Code has been released at https://github.com/DreamtaleCore/PnP-GA.
Yunfei Liu 0001, Ruicong Liu, Haofei Wang 0001, Feng Lu 0005
ICCV2
2020 Predicting Liquidity Ratio of Mutual Funds via Ensemble Learning
abstract
We are entering a new era of AI, in which the core technology is machine learning. However, many machine learning models are opaque and not intuitive enough, making it difficult for users to understand how AI systems make decisions. In some scenarios, especially in the fields of finance, healthcare, automatic driving, users have a strong demand for the interpretability of the model. Although AI systems provide a lot of benefits, if the decision and behavior cannot be explained to users or regulators, the effectiveness and development of the systems will be restricted. To gain the trust of users, Explainable AI is necessary.Daily prediction of mutual fund holdings can be a very useful tool. If we can predict the daily holding positions of large mutual funds, we can gain insights into the sentiments of institutional investors which shed lights into the outlook of the market. How to get daily holding from the delayed disclosure information? We leverage on another source of key information - the price of a mutual fund is updated daily, often released a few hours after the market closes. Therefore, we can utilize the daily price fluctuation, combined with quarterly revealed holdings, to make daily predictions of mutual fund holdings. In this paper, we proposed an Ensemble Learning model to predict liquidity ratio of mutual funds. The model has strong interpretability, which is beneficial to users, developers and regulators and all parties involved. Compared with the real fund position data only disclosed once a quarter, our model effectuates timely and efficient high-frequency calculation. In the process of modeling, we creatively apply the framework of Ensemble Learning to portfolio decomposition for the first time. The Ensemble Learning model leverages the diversity of base learners to improve the overall prediction performance. Extensive empirical results on China A-Shares market show that our model can achieve superior accuracy, robustness, and generalization ability.
Kun Kong, Ruicong Liu
IEEE BigData2