Ruolin Wang

dblp:34/7464 · DBLP profile ↗
← Back
25ranked-venue papers
9as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 9 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Gather to Glow: an Autonomy-Supportive Game for Parent-child Co-Play
abstract
Parent–child co-play is often framed as beneficial for family interaction, yet caregiver involvement can unintentionally become parental dominance, potentially undermining children’s autonomy. We present Gather to Glow, a two-controller co-play game that explores how autonomy-supportive design can structure caregiver participation while preserving child-led play. Grounded in Self-Determination Theory (SDT) and parental autonomy support (PAS), the system combines three mechanisms: (1) choices-as-collectibles with AI customization, (2) asymmetric roles that preserve child decision authority while guiding caregiver actions to support, and (3) visual feedback designed to sustain intrinsic motivation. The demo invites a child–caregiver dyad to experience a four-phase co-play loop—exploration, conversation, cooperation, and reflection—through a hands-on interaction of approximately 15 minutes. We also report early pilot observations that informed the demo’s design and highlight how autonomy-supportive interaction can be embedded in co-play.
Junhong Jiang, Xinyun Luo, Hanxiang Xu, Tzu-Jung Tseng, Ruolin Wang, Yuhua Jin
IDC8
2026 A hybrid model for carbon prices: Integrating investor attention, mixed-frequency data, quantile regression and deep learning
Di Sha, Arne Johannssen, Xianyi Zeng, Ruolin Wang, Kim Phuc Tran
Expert Syst. Appl.4
2025 Traffic Scenario Logic: A Spatial-Temporal Logic for Modeling and Reasoning of Urban Traffic Scenarios
abstract
Formal representations of traffic scenarios can be used to generate test cases for the safety verification of autonomous driving. However, most existing methods are limited to highway or highly simplified intersection scenarios due to the intricacy and diversity of traffic scenarios. In response, we propose Traffic Scenario Logic (TSL), which is a spatial-temporal logic designed for modeling and reasoning of urban pedestrian-free traffic scenarios. TSL provides a formal representation of the urban road network that can be derived from OpenDRIVE, i.e., the de facto industry standard of high-definition maps for autonomous driving, enabling the representation of a broad range of traffic scenarios without discretization approximations. We implemented the reasoning of TSL using Telingo, i.e., a solver for temporal programs based on Answer Set Programming, and tested it on different urban road layouts. Demonstrations show the effectiveness of TSL in test scenario generation and its potential value in areas like decision-making and control verification of autonomous driving. The code for TSL reasoning has been open-sourced.
Ruolin Wang, Yuejiao Xu, Jianmin Ji
AAAI1
2025 CoSight: Exploring Viewer Contributions to Online Video Accessibility Through Descriptive Commenting
abstract
Figure 1: CoSight, a Chrome extension developed as a design probe to explore how lightweight interface nudges might encourage accessibility contributions from sighted video viewers when watching and commenting.The prototype augments YouTube video pages with features inspired by Fogg's Behavior Model [24], including color labels to highlight accessibility gaps (sparks), hints and references to guide contributions (facilitators), and reminders at key moments (signals).
Ruolin Wang, Xingyu Liu 0002, Wayne Zhang 0004, Ziqian Liao, Ziwen Li 0001, Amy Pavel, Xiang 'Anthony' Chen
UIST1
2024 LDIP: Real-time on-road object detection with depth estimation from a single image
abstract
Detecting on-road objects with absolute depth information is one of the most crucial tasks in autonomous driving to ensure safety. Traditional 2D object detection aims to classify and locate objects in image space, but it cannot acquire in-depth information. While 3D object detection and pixel-level depth detection tasks can provide accurate depth information for objects, they are challenging to deploy in real-world scenarios due to their significant inference overhead. This paper proposes a novel deep learning-based model named the Location and Depth Information Perceptron (LDIP), designed to provide positional, categorical, and absolute depth information for given objects in the images.We first conducted model training and validation on the vehicle-side autonomous driving dataset—KITTI. The experimental results show that we achieved a 68.6% mAP in object recognition tasks and an RMSE of 0.101 and AbsRel of 2.327 in depth estimation tasks, all of which represent state-of-the-art performance in comparable tasks. Subsequently, we fine-tuned the trained model on DAIR, where the validated mAP, AbsRel, and RMSE reached 65.4%, 0.092, and 2.461 respectively. This demonstrates the robustness and generalization of our model across different types of road datasets.Moreover, in comparison to other models, our model is more compact while maintaining accuracy, achieving an inference speed of 70 frames per second on an NVIDIA 4060 GPU, thus making it deployable in practical scenarios. Relevant code is available at https://github.com/xcp-ustc/LDIP.
Chengpeng Xu, Xiao Sun 0003, Yangyang Xu 0002, Ruolin Wang
IROS4
2024 From Text to Pixels: Enhancing User Understanding through Text-to-Image Model Explanations
abstract
Recent progress in Text-to-Image (T2I) models promises transformative applications in art, design, education, medicine, and entertainment. These models, exemplified by Dall-e, Imagen, and Stable Diffusion, have the potential to revolutionize various industries. However, a primary concern is their operation as a ‘black-box’ for many users. Without understanding the underlying mechanics, users are unable to harness the full potential of these models. This study focuses on bridging this gap by developing and evaluating explanation techniques for T2I models, targeting inexperienced end users. While prior works have delved into Explainable AI (XAI) methods for classification or regression tasks, T2I generation poses distinct challenges. Through formative studies with experts, we identified unique explanation goals and subsequently designed tailored explanation strategies. We then empirically evaluated these methods with a cohort of 473 participants from Amazon Mechanical Turk (AMT) across three tasks. Our results highlight users’ ability to learn new keywords through explanations, a preference for example-based explanations, and challenges in comprehending explanations that significantly shift the image’s theme. Moreover, findings suggest users benefit from a limited set of concurrent explanations. Our main contributions include a curated dataset for evaluating T2I explainability techniques, insights from a comprehensive AMT user study, and observations critical for future T2I model explainability research.
Noyan Evirgen, Ruolin Wang, Xiang 'Anthony' Chen
IUI2
2023 A²CoST: An ASP-based Avoidable Collision Scenario Testbench for Autonomous Vehicles
abstract
This paper addresses the challenge of generating safety-critical scenarios with multiple adversarial vehicles for testing autonomous vehicles. Such scenarios must be plausible and collision-avoidable while resulting in a collision with the vehicle-under-test. However, the tremendous number of scenarios and the low ratio of plausible scenarios makes previous methods squander primary resources on implausible scenarios, degenerating their efficiency. We propose a two-stage framework called the ASP-based Avoidable Collision Scenario Testbench (A²CoST) to overcome this obstacle and improve efficiency. In the former stage, we apply Answer Set Programming (ASP) for generating plausible logical scenarios. In the latter stage, we use a search algorithm to refine logical scenarios into safety-critical concrete scenarios. We also compute collision-free trajectories in these concrete scenarios while the vehicle-under-test fails to avoid the collision. We empirically show the A²CoST significantly decreases the time consumption for simple scenarios while still effectively generating complex critical scenarios. The comparison with real-world traffic data further demonstrates the value of A²CoST in generating plausible scenarios. The source codes of our method and the baselines are opened at https://github.com/Autonomous-Driving-Safety-Project/AACoST.
Ruolin Wang, Yuejiao Xu, Jie Peng 0002, Jianmin Ji
KR1
2023 A mobile intelligent guide system for visually impaired pedestrian
Zimiao Xie, Pengxin Yuan, Ruolin Wang
J. Syst. Softw.4
2023 Semantic Relevance Learning for Video-Query Based Video Moment Retrieval
abstract
The task of video-query based video moment retrieval (VQ-VMR) aims to localize the segment in the reference video, which matches semantically with a short query video. This is a challenging task due to the rapid expansion and massive growth of online video services. With accurate retrieval of the target moment, we propose a new metric to effectively assess the semantic relevance between the query video and segments in the reference video. We also develop a new VQ-VMR framework to discover the intrinsic semantic relevance between a pair of input videos. It comprises two key components: a Fine-grained Feature Interaction (FFI) module and a Semantic Relevance Measurement (SRM) module. Together they can effectively deal with both the spatial and temporal dimensions of videos. First, the FFI module computes the semantic similarity between videos at a local frame level, mainly considering the spatial information in the videos. Subsequently, the SRM module learns the similarity between videos from a global perspective, taking into account the temporal information. We have conducted extensive experiments on two key datasets which demonstrate noticeable improvements of the proposed approach over the state-of-the-art methods.
Shuwei Huo, Yuan Zhou 0006, Ruolin Wang, Wei Xiang 0001, Sun-Yuan Kung
IEEE Trans. Multim.3
2022 The Transition Law of Sepsis Patients' Illness States Based on Complex Network
Ruolin Wang, Jingming Liu, Minghui Gong, Chunping Li
AIME1
2022 TypeOut: Leveraging Just-in-Time Self-Affirmation for Smartphone Overuse Reduction
abstract
Smartphone overuse is related to a variety of issues such as lack of sleep and anxiety. We explore the application of Self-Affirmation Theory on smartphone overuse intervention in a just-in-time manner. We present TypeOut, a just-in-time intervention technique that integrates two components: an in-situ typing-based unlock process to improve user engagement, and self-affirmation-based typing content to enhance effectiveness. We hypothesize that the integration of typing and self-affirmation content can better reduce smartphone overuse. We conducted a 10-week within-subject field experiment (N=54) and compared TypeOut against two baselines: one only showing the self-affirmation content (a common notification-based intervention), and one only requiring typing non-semantic content (a state-of-the-art method). TypeOut reduces app usage by over 50%, and both app opening frequency and usage duration by over 25%, all significantly outperforming baselines. TypeOut can potentially be used in other domains where an intervention may benefit from integrating self-affirmation exercises with an engaging just-in-time mechanism.
Xuhai Xu, Tianyuan Zou, Yanzhang Li, Ruolin Wang, Tianyi Yuan, Yuntao Wang 0001, Yuanchun Shi, Jennifer Mankoff, Anind K. Dey
CHI5
2022 CrossA11y: Identifying Video Accessibility Issues via Cross-modal Grounding
abstract
Authors make their videos visually accessible by adding audio descriptions (AD), and auditorily accessible by adding closed captions (CC). However, creating AD and CC is challenging and tedious, especially for non-professional describers and captioners, due to the difficulty of identifying accessibility problems in videos. A video author will have to watch the video through and manually check for inaccessible information frame-by-frame, for both visual and auditory modalities. In this paper, we present CrossA11y, a system that helps authors efficiently detect and address visual and auditory accessibility issues in videos. Using cross-modal grounding analysis, CrossA11y automatically measures accessibility of visual and audio segments in a video by checking for modality asymmetries. CrossA11y then displays these segments and surfaces visual and audio accessibility issues in a unified interface, making it intuitive to locate, review, script AD/CC in-place, and preview the described and captioned video immediately. We demonstrate the effectiveness of CrossA11y through a lab study with 11 participants, comparing to existing baseline.
Xingyu Liu 0002, Ruolin Wang, Dingzeyu Li, Xiang 'Anthony' Chen, Amy Pavel
UIST2
2022 Stochastic Geometry Analysis of LEO Constellation Coverage under Atmospheric Attenuation
abstract
The development of 6G communication is now putting higher performance requirements on the mobile satellite system. How to achieve comprehensive coverage by adding satellites to mobile communication system has become a popular topic on integrated satellite-ground network. However, using traditional methods to analyze coverage performance is limited by the topology of the satellite constellation. In this paper we use the stochastic geometry to model the LEO constellation. The stochastic geometry analysis method weakens the influence of constellation topology and provides a new representation of the coverage performance characteristics. This paper considers the models of beam coverage angle and atmospheric attenuation on coverage performance and proposes a new interference simplification method. It provides a new tool for the future optimization analysis of coverage performance of LEO satellite constellation and gives a new general expression for the coverage probability. The simulation shows that the coverage analysis model established in this paper can clearly represent the characteristics of the variation of the coverage probability with different parameters, which is conforms to the changing trends of the actual satellite constellation. For example, it can obtain the number of satellites with optimal coverage probability at different orbital altitudes.
Ruolin Wang, Pinyi Ren, Dongyang Xu 0003
VTC Fall1
2021 LightWrite: Teach Handwriting to The Visually Impaired with A Smartphone
abstract
Learning to write is challenging for blind and low vision (BLV) people because of the lack of visual feedback. Regardless of the drastic advancement of digital technology, handwriting is still an essential part of daily life. Although tools designed for teaching BLV to write exist, many are expensive and require the help of sighted teachers. We propose LightWrite, a low-cost, easy-to-access smartphone application that uses voice-based descriptive instruction and feedback to teach BLV users to write English lowercase letters and Arabian digits in a specifically designed font. A two-stage study with 15 BLV users with little prior writing knowledge shows that LightWrite can successfully teach users to learn handwriting characters in an average of 1.09 minutes for each letter. After initial training and 20-minute daily practice for 5 days, participants were able to write an average of 19.9 out of 26 letters that are recognizable by sighted raters.
Zihan Wu 0002, Chun Yu, Xuhai Xu, Tianyuan Zou, Ruolin Wang, Yuanchun Shi
CHI6
2021 Revamp: Enhancing Accessible Information Seeking Experience of Online Shopping for Blind or Low Vision Users
abstract
Online shopping has become a valuable modern convenience, but blind or low vision (BLV) users still face significant challenges using it, because of: 1) inadequate image descriptions and 2) the inability to filter large amounts of information using screen readers. To address those challenges, we propose Revamp, a system that leverages customer reviews for interactive information retrieval. Revamp is a browser integration that supports review-based question-answering interactions on a reconstructed product page. From our interview, we identified four main aspects (color, logo, shape, and size) that are vital for BLV users to understand the visual appearance of a product. Based on the findings, we formulated syntactic rules to extract review snippets, which were used to generate image descriptions and responses to users’ queries. Evaluations with eight BLV users showed that Revamp 1) provided useful descriptive information for understanding product appearance and 2) helped the participants locate key information efficiently.
Ruolin Wang, Mingrui Ray Zhang, Zhaoheng Li, Zhixiu Liu, Zihan Dang, Chun Yu, Xiang 'Anthony' Chen
CHI1
2021 Voicemoji: Emoji Entry Using Voice for Visually Impaired People
abstract
Keyboard-based emoji entry can be challenging for people with visual impairments: users have to sequentially navigate emoji lists using screen readers to find their desired emojis, which is a slow and tedious process. In this work, we explore the design and benefits of emoji entry with speech input, a popular text entry method among people with visual impairments. After conducting interviews to understand blind or low vision (BLV) users’ current emoji input experiences, we developed Voicemoji, which (1) outputs relevant emojis in response to voice commands, and (2) provides context-sensitive emoji suggestions through speech output. We also conducted a multi-stage evaluation study with six BLV participants from the United States and six BLV participants from China, finding that Voicemoji significantly reduced entry time by 91.2% and was preferred by all participants over the Apple iOS keyboard. Based on our findings, we present Voicemoji as a feasible solution for voice-based emoji entry.
Mingrui Ray Zhang, Ruolin Wang, Xuhai Xu, Qisheng Li, Ather Sharif, Jacob O. Wobbrock
CHI2
2021 Temporal Action Localization Using Long Short-Term Dependency
abstract
Temporal action localization in untrimmed videos is an important but difficult task. Difficulties are encountered in the application of existing methods when modeling the temporal structures of videos. In the present study, we develop a novel method, referred to as the Gemini Network, for effective modeling of temporal structures and achieving high-performance temporal action localization. The significant improvements afforded by the proposed method are due to three major factors. First, temporal dependencies are explicitly distinguished as long-term temporal dependencies and short-term temporal dependencies and are separately captured by two dedicated subnets. Second, a long-range temporal dependency capture module combined with a self-adaptive pooling module is proposed to capture long-term temporal dependency. Third, the proposed method uses auxiliary supervision, with the auxiliary classifier losses affording additional constraints for improving the modeling capability of the network. As a demonstration of its effectiveness, the Gemini Network is used to achieve a state-of-the-art temporal action localization performance on two challenging datasets, namely, THUMOS14 and ActivityNet.
Yuan Zhou 0006, Ruolin Wang, Sun-Yuan Kung
IEEE Trans. Multim.2
2020 A Feature Pair Fusion And Hierarchical Learning Framework For Video Re-Localization
abstract
Video re-localization has become an emerging research topic nowadays but existing methods still have many deficiencies. The existing deficiencies mainly lie in the interference caused by the irrelevant information in the input reference video and the ignorance of the correlation between query and reference video features. Therefore, we present a novel framework named Semantic Relevance Learning Network to address these shortcomings. First, we extract effective proposals from reference video as new inputs to reduce interference from irrelevant video frames. Second, two key components of our proposed model, the Attention-based Fusion Tensor and Semantic Relevance Measurement, jointly explore the intrinsic correlation between video feature pairs and finally get a score as measurement. To better evaluate our proposed model, we reorganize Thumos14 to obtain another new dataset for the video re-localization task. For both ActivityNet and Thumos14, our model achieves the best performance reported so far.
Ruolin Wang, Yuan Zhou 0006
ICIP1
2020 Dynamic MRI reconstruction exploiting blind compressed sensing combined transform learning regularization
Ruolin Wang, Yixue Wang
Neurocomputing2
2020 Adaptively weighted nonlocal means and TV minimization for speckle reduction in SAR images
Ruolin Wang, Yixue Wang, Ke Lu 0002
Multim. Tools Appl.1
2019 EarTouch: Facilitating Smartphone Use for Visually Impaired People in Mobile and Public Scenarios
abstract
Interacting with a smartphone using touch input and speech output is challenging for visually impaired people in mobile and public scenarios, where only one hand may be available for input (e.g., while holding a cane) and using the loudspeaker for speech output is constrained by environmental noise, privacy, and social concerns. To address these issues, we propose EarTouch, a one-handed interaction technique that allows the users to interact with a smartphone using the ear to perform gestures on the touchscreen. Users hold the phone to their ears and listen to speech output from the ear speaker privately. We report how the technique was designed, implemented, and evaluated through a series of studies. Results show that EarTouch is easy, efficient, fun and socially acceptable to use.
Ruolin Wang, Chun Yu, Xing-Dong Yang, Yuanchun Shi
CHI1
2018 Improving Person Re-Identification by Adaptive Hard Sample Mining
abstract
The field of person reidentification has made significant advances riding on the wave of deep learning. However, owing to the fact that there are much more easy examples than those meaningful hard examples in dataset, the training tends to stagnate quickly and the model may suffer from over-fitting. Therefore, the hard sample mining method is fateful to optimize the model and improve the learning efficiency. In this paper, an Adaptive Hard Sample Mining algorithm is proposed for training a robust person re-identification model. No need for hand-picking the images in the batch or designing the loss function for both positive and negative pairs, we can briefly calculate the hard level by comparing the prediction result with the true label of the sample. Meanwhile, an adaptive threshold of hard level can make the algorithm not only stay in step with training process harmoniously but also alleviate the under-fitting and over-fitting problem simultaneously. Besides, the designed network to implement the approach has good generalization performance that can be combined with various of existing models readily. Experimental results on Market-1501 and DukeMTMC-reID datasets clearly demonstrate the effectiveness of the proposed algorithm.
Kezhou Chen, Chuchu Han, Nong Sang, Changxin Gao, Ruolin Wang
ICIP6
2015 Fusion-based edge-sensitive interpolation method for deinterlacing
Hao Zhang 0031, Ruolin Wang, Wenjiang Liu, Mengtian Rong
Multim. Tools Appl.2
2013 A Transductive Transfer Learning Method for Ship Target Recognition
abstract
Ship target recognition in infrared image remains a difficult problem, due to the projection or silhouette of a three-dimensional ship target being variable in shape, orientation and scale to make its recognizability unstable. In this paper, a transductive transfer learning framework is proposed to solve the problem. Hu moments is firstly extracted as feature vectors of target. Then the transductive transfer learning method is used to find the common parameters between the feature spaces of the training ship samples and the detected ship targets, and transfer the similar knowledge from those data with different distributions. According to the experiment result in simulation infrared images, it shows that the ship targets can be recognized highly and reliably by our proposed framework. It demonstrates the robustness and effectiveness of our method for infrared images.
Zhiping Dan, Nong Sang, Ruolin Wang, Yanfei Chen
ICIG3
2013 Novel License Plate Detection Method for Complex Scenes
abstract
In this paper, a novel license plate detection method is proposed. There are three key steps in our method, i.e. image preprocessing, license plate detection and license plate confirmation. First, the noises are removed and the diversities of license plate forms are unified through image preprocessing. And then, the license plates are detected roughly by using the cascade AdaBoost classifier. Finally, the gradient images are binarized, and the connected component analysis will be adopted to remove some false plates. Meanwhile, the offline trained Support Vector Machine (SVM) classifier is adopted to confirm the license plate candidates in further. The promising results of the proposed method is verified by experiments on a challenging database.
Nong Sang, Ruolin Wang, Xiaoqin Kuang
ICIG3