VLDB 2026 Research / reviewers in the wild / expert
Rui Yu 0002
dblp:43/4940-2
· DBLP profile ↗
24ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-0946-6769ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding Parents' Perspectives on Responsible AI for Children's Self-Directed LearningabstractGenerative AI is increasingly present in children’s learning environments, yet little is known about how families navigate this technology in middle childhood (ages 7–13), when parental guidance remains strong but children seek independence. Drawing on self-directed learning (SDL), we explore how parents in our exploratory sample perceived children’s emerging self-directness and agency. Through focus groups with 13 parent–child pairs, we examine parents’ views on children’s AI literacy development, readiness factors, and mediation strategies. Parents described emergent pathways shaped by screen time, self-directness, and knowledge growth. They often confined AI to learning-only contexts, positioning it as a tutor while overlooking non-learning uses and risks such as privacy and infrastructural embedding. Many acknowledged limited AI literacy and turned to joint engagement as opportunities for co-learning. Our findings surface possible parental pathways of children’s AI literacy, highlight gaps between pragmatic expectations and critical literacies, and offer situated design considerations for AI systems that scaffold SDL while balancing oversight with autonomy. Jingyi Xie 0001, Chuhao Wu, Ge Wang 0004, Rui Yu 0002, He Zhang 0033, Ronald A. Metoyer, Si Chen 0006 |
CHI | 4 |
| 2025 | Beyond Visual Perception: Insights from Smartphone Interaction of Visually Impaired Users with Large Multimodal ModelsabstractLarge multimodal models (LMMs) have enabled new AI-powered applications that help people with visual impairments (PVI) receive natural language descriptions of their surroundings through audible text. We investigated how this emerging paradigm of visual assistance transforms how PVI perform and manage their daily tasks. Moving beyond basic usability assessments, we examined both the capabilities and limitations of LMM-based tools in personal and social contexts, while exploring design implications for their future development. Through interviews with 14 visually impaired users and analysis of image descriptions from both participants and social media using Be My AI (an LMM-based application), we identified two key limitations. First, these systems' context awareness suffers from hallucinations and misinterpretations of social contexts, styles, and human identities. Second, their intent-oriented capabilities often fail to grasp and act on users' intentions. Based on these findings, we propose design strategies for improving both human-AI and AI-AI interactions, contributing to the development of more effective, interactive, and personalized assistive technologies. Jingyi Xie 0001, Rui Yu 0002, He Zhang 0033, Syed Masum Billah, Sooyeon Lee, John M. Carroll 0001 |
CHI | 2 |
| 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Imageabstract3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data. In this work, we introduce a novel framework dubbed ZeroPlane, a Transformer-based model targeting zero-shot 3D plane detection and reconstruction from a single image, over diverse domains and environments. To enable data-driven models across multiple domains, we have curated a large-scale planar benchmark, comprising over 14 datasets and 560,000 high-resolution, dense planar annotations for diverse indoor and outdoor scenes. To address the challenge of achieving desirable planar geometry on multi-dataset training, we propose to disentangle the representation of plane normal and offset, and employ an exemplar-guided, classification-then-regression paradigm to learn plane and offset respectively. Additionally, we employ advanced backbones as image encoder, and present an effective pixel-geometry-enhanced plane embedding module to further facilitate planar reconstruction. Extensive experiments across multiple zero-shot evaluation datasets have demonstrated that our approach significantly outperforms previous methods on both reconstruction accuracy and generalizability, especially over in-the-wild data. Our code and data are available at: https://github.com/jcliu0428/ZeroPlane. Rui Yu 0002, Sili Chen, Sharon X. Huang, Hengkai Guo |
CVPR | 2 |
| 2025 | PlanarSplatting: Accurate Planar Surface Reconstruction in 3 MinutesabstractThis paper presents PlanarSplatting, an ultra-fast and accurate surface reconstruction approach for multi-view indoor images. We take the 3D planes as the main objective due to their compactness and structural expressiveness in indoor scenes, and develop an explicit optimization framework that learns to fit the expected surface of indoor scenes by splatting the 3D planes into 2.5D depth and normal maps. As our PlanarSplatting operates directly on the 3D plane primitives, it eliminates the dependencies on 2D/3D plane detection and plane matching/tracking for planar surface reconstruction. Furthermore, with the essential merits of plane-based representation coupled with CUDA-based implementation of planar splatting functions, Planar-Splatting reconstructs an indoor scene in 3 minutes while having significantly better geometric accuracy. Thanks to our ultra-fast reconstruction speed, the largest quantitative evaluation on the ScanNet and ScanNet++ datasets over hundreds of scenes clearly demonstrated the advantages of our method. We believe that our accurate and ultrafast planar surface reconstruction method will be applied in the structured data curation for surface reconstruction in the future. The code of our CUDA implementation will be publicly available. Bin Tan 0002, Rui Yu 0002, Yujun Shen, Nan Xue 0001 |
CVPR | 2 |
| 2025 | Describe, Adapt and Combine: Empowering CLIP Encoders for Open-Set 3D Object Retrieval
Yang Zhou 0007, Zhe Liu 0033, Rui Yu 0002, Song Bai 0001, Yulong Wang 0002, Xinwei He 0001, Xiang Bai |
ICCV | 4 |
| 2025 | Top2Pano: Learning to Generate Indoor Panoramas from Top-Down View
Zitong Zhang 0003, Suranjan Gautam, Rui Yu 0002 |
ICCV | 3 |
| 2025 | Noise-Robust Tuning of SAM for Domain Generalized Ultrasound Image Segmentation
Zhikai Wei, Hanyu Du, Rui Yu 0002, Bo Du 0001, Yongchao Xu |
MICCAI (5) | 4 |
| 2025 | Fairness-Aware Graph Representation Learning with Limited Demographic Information
Zichong Wang, Zhipeng Yin, Liping Yang 0002, Jun Zhuang 0004, Rui Yu 0002, Qingzhao Kong, Wenbin Zhang 0002 |
ECML/PKDD (1) | 5 |
| 2025 | Computer-Aided Layout Generation for Building Design: A ReviewabstractGenerating realistic building layouts for automatic building design has been studied in both computer vision and architectural domains. Traditional approaches in the latter, which are based on optimization techniques or heuristic design guidelines, can synthesize desirable layouts, but usually require post-processing and involve human interaction in the design pipeline, making them costly and time-consuming. The advent of deep generative models has significantly improved the fidelity and diversity of the generated architecture layouts, reducing the workload of designers and making the process much more efficient. This paper presents a comprehensive review of three major research topics in architectural layout design and generation: floorplan layout generation, scene layout synthesis, and generation of various other formats of building layouts. For each topic, we overview the leading paradigms, categorized either by research domains (architecture or machine learning) or by user input conditions or constraints. We then introduce commonly-adopted benchmark datasets used to verify the effectiveness of the methods, as well as corresponding evaluation metrics. Finally, we identify the well-solved problems and limitations of existing approaches, and then propose promising directions for future research. This survey has an associated project which aims to maintain the resources, at https://github.com/jcliu0428/awesome-building-layout-generation. Yuan Xue 0002, Haomiao Ni, Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang |
Comput. Vis. Media | 4 |
| 2024 | BubbleCam: Engaging Privacy in Remote Sighted AssistanceabstractRemote sighted assistance (RSA) offers prosthetic support to people with visual impairments (PVI) through image- or video-based conversations with remote sighted assistants. While useful, RSA services introduce privacy concerns, as PVI may reveal private visual content inadvertently. Solutions have emerged to address these concerns on image-based asynchronous RSA, but exploration into solutions for video-based synchronous RSA remains limited. In this study, we developed BubbleCam, a high-fidelity prototype allowing PVI to conceal objects beyond a certain distance during RSA, granting them privacy control. Through an exploratory field study with 24 participants, we found that 22 appreciated the privacy enhancements offered by BubbleCam. The users gained autonomy, reducing embarrassment by concealing private items, messy areas, or bystanders, while assistants could avoid irrelevant content. Importantly, BubbleCam maintained RSA’s primary function without compromising privacy. Our study highlighted a cooperative approach to privacy preservation, transitioning the traditionally individual task of maintaining privacy into an interactive, engaging privacy preserving experience. Jingyi Xie 0001, Rui Yu 0002, He Zhang 0033, Sooyeon Lee, Syed Masum Billah, John M. Carroll 0001 |
CHI | 2 |
| 2024 | NeRF-Enhanced Outpainting for Faithful Field-of-View ExtrapolationabstractIn various applications, such as robotic navigation and remote visual assistance, expanding the field of view (FOV) of the camera proves beneficial for enhancing environmental perception. Unlike image outpainting techniques aimed solely at generating aesthetically pleasing visuals, these applications demand an extended view that faithfully represents the scene. To achieve this, we formulate a new problem of faithful FOV extrapolation that utilizes a set of pre-captured images as prior knowledge of the scene. To address this problem, we present a simple yet effective solution called NeRF-Enhanced Outpainting (NEO) that uses extended-FOV images generated through NeRF to train a scene-specific image outpainting model. To assess the performance of NEO, we conduct comprehensive evaluations on three photorealistic datasets and one real-world dataset. Extensive experiments on the benchmark datasets showcase the robustness and potential of our method in addressing this challenge. We believe our work lays a strong foundation for future exploration within the research community. Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang |
ICRA | 1 |
| 2024 | Spatial-Aware Attention Generative Adversarial Network for Semi-supervised Anomaly Detection in Medical Image
Zhichao Sun 0004, Zelong Liu, Rui Yu 0002, Bo Du 0001, Yongchao Xu |
MICCAI (5) | 5 |
| 2024 | MoreStyle: Relax Low-Frequency Constraint of Fourier-Based Image Reconstruction in Generalizable Medical Image Segmentation
Rui Yu 0002, Bo Du 0001, Yongchao Xu |
MICCAI (8) | 3 |
| 2024 | WIA-LD2ND: Wavelet-Based Image Alignment for Self-supervised Low-Dose CT Denoising
Yuliang Gu, Bo Du 0001, Yongchao Xu, Rui Yu 0002 |
MICCAI (7) | 6 |
| 2023 | Are Two Heads Better than One? Investigating Remote Sighted Assistance with Paired VolunteersabstractRemote Sighted Assistance (RSA) is a popular smartphone-mediated aid for people with blindness, where a sighted individual converses with a blind individual in a one-on-one (1:1) session. Since sighted assistants outnumber blind individuals (13:1), this paper investigates what happens when more than one sighted individual assists a single blind individual in a session. Specifically, we propose paired-volunteer RSA, a new paradigm where two sighted volunteers assist a single user with blindness. We investigate the feasibility, desirability, and challenges of this paradigm and explore its opportunities. Our study with 8 sighted volunteers and 9 blind users reveals that the proposed paradigm extends the one-on-one RSA to cover a broader range of more intellectual and experiential tasks, providing new and distinctive opportunities in supporting complex, open-ended tasks (e.g., pursuing hobbies, appreciating arts, and seeking entertainment). These opportunities can not only enrich the blind users' quality of life and independence but also offer a fun and engaging experience for the sighted volunteers. The study also reveals the costs of extended collaboration in this paradigm. Finally, we synthesize a taxonomy of tasks where the proposed RSA paradigm can succeed and outline how HCI researchers and system designers can realize this paradigm. Jingyi Xie 0001, Rui Yu 0002, Kaiming Cui, Sooyeon Lee, John M. Carroll 0001, Syed Masum Billah |
Conference on Designing Interactive Systems | 2 |
| 2023 | Be Real in Scale: Swing for True Scale in Dual Camera ModeabstractMany mobile AR apps that use the front-facing camera can benefit significantly from knowing the metric scale of the user’s face. However, the true scale of the face is hard to measure because monocular vision suffers from a fundamental ambiguity in scale. The methods based on prior knowledge about the scene either have a large error or are not easily accessible. In this paper, we propose a new method to measure the face scale by a simple user interaction: the user only needs to swing the phone to capture two selfies while using the recently popular Dual Camera mode. This mode allows simultaneous streaming of the front camera and the rear cameras and has become a key feature in many social apps. A computer vision method is applied to first estimate the absolute motion of the phone from the images captured by two rear cameras, and then calculate the point cloud of the face by triangulation. We develop a prototype mobile app to validate the proposed method. Our user study shows that the proposed method is favored compared to existing methods because of its high accuracy and ease of use. Our method can be built into Dual Camera mode and can enable a wide range of applications (e.g., virtual try-on for online shopping, true-scale 3D face modeling, gaze tracking, and face anti-spoofing) by introducing true scale to smartphone-based XR. The code is available at https://github.com/ruiyu0/Swing-for-True-Scale. Rui Yu 0002, Jian Wang 0100, Sizhuo Ma, Sharon X. Huang, Gurunandan Krishnan |
ISMAR | 1 |
| 2022 | Helping Helpers: Supporting Volunteers in Remote Sighted Assistance with Augmented Reality Mapsabstract., agents, provide real-time assistance to blind users via video-chat-like communication. Prior work identified several challenges for the agents to provide navigational assistance to users and proposed computer vision-mediated RSA service to address those challenges. We present an interactive system implementing a high-fidelity prototype of RSA service using augmented reality (AR) maps with localization and virtual elements placement capabilities. The paper also presents a confederate-based study design to evaluate the effects of AR maps with 13 untrained agents. The study revealed that, compared to baseline RSA, agents were significantly faster in providing indoor navigational assistance to a confederate playing the role of users, and agents' mental workload was significantly reduced-all indicate the feasibility and scalability of AR maps in RSA services. Jingyi Xie 0001, Rui Yu 0002, Sooyeon Lee, Yao Lyu, Syed Masum Billah, John M. Carroll 0001 |
Conference on Designing Interactive Systems | 2 |
| 2022 | Cascade Transformers for End-to-End Person SearchabstractThe goal of person search is to localize a target person from a gallery set of scene images, which is extremely challenging due to large scale variations, pose/viewpoint changes, and occlusions. In this paper, we propose the Cascade Occluded Attention Transformer (COAT) for end-to-end person search. Our three-stage cascade design focuses on detecting people in the first stage, while later stages simultaneously and progressively refine the representation for person detection and re-identification. At each stage the occluded attention transformer applies tighter intersection over union thresholds, forcing the network to learn coarse-to-fine pose/scale invariant features. Meanwhile, we calculate each detection's occluded attention to differentiate a person's tokens from other people or the background. In this way, we simulate the effect of other objects occluding a person of interest at the token-level. Through comprehensive experiments, we demonstrate the benefits of our method by achieving state-of-the-art performance on two benchmark datasets. Rui Yu 0002, Dawei Du, Rodney LaLonde, Daniel Davila, Christopher Funk, Anthony Hoogs, Brian Clipp |
CVPR | 1 |
| 2022 | Opportunities for Human-AI Collaboration in Remote Sighted AssistanceabstractRemote sighted assistance (RSA) has emerged as a conversational assistive technology for people with visual impairments (VI), where remote sighted agents provide realtime navigational assistance to users with visual impairments via video-chat-like communication. In this paper, we conducted a literature review and interviewed 12 RSA users to comprehensively understand technical and navigational challenges in RSA for both the agents and users. Technical challenges are organized into four categories: agents' difficulties in orienting and localizing the users; acquiring the users' surroundings and detecting obstacles; delivering information and understanding user-specific situations; and coping with a poor network connection. Navigational challenges are presented in 15 real-world scenarios (8 outdoor, 7 indoor) for the users. Prior work indicates that computer vision (CV) technologies, especially interactive 3D maps and realtime localization, can address a subset of these challenges. However, we argue that addressing the full spectrum of these challenges warrants new development in Human-CV collaboration, which we formalize as five emerging problems: making object recognition and obstacle avoidance algorithms blind-aware; localizing users under poor networks; recognizing digital content on LCD screens; recognizing texts on irregular surfaces; and predicting the trajectory of out-of-frame pedestrians or objects. Addressing these problems can advance computer vision research and usher into the next generation of RSA service. Sooyeon Lee, Rui Yu 0002, Jingyi Xie 0001, Syed Masum Billah, John M. Carroll 0001 |
IUI | 2 |
| 2021 | Towards Robust Human Trajectory Prediction in Raw VideosabstractHuman trajectory prediction has received increased attention lately due to its importance in applications such as autonomous vehicles and indoor robots. However, most existing methods make predictions based on human-labeled trajectories and ignore the errors and noises in detection and tracking. In this paper, we study the problem of human trajectory forecasting in raw videos, and show that the prediction accuracy can be severely affected by various types of tracking errors. Accordingly, we propose a simple yet effective strategy to correct the tracking failures by enforcing prediction consistency over time. The proposed "re-tracking" algorithm can be applied to any existing tracking and prediction pipelines. Experiments on public benchmark datasets demonstrate that the proposed method can improve both tracking and prediction performance in challenging real-world scenarios. The code and data are available at https://git.io/retracking-prediction. Rui Yu 0002, Zihan Zhou 0001 |
IROS | 1 |
| 2020 | Data-driven Distributed State Estimation and Behavior Modeling in Sensor NetworksabstractNowadays, the prevalence of sensor networks has enabled tracking of the states of dynamic objects for a wide spectrum of applications from autonomous driving to environmental monitoring and urban planning. However, tracking realworld objects often faces two key challenges: First, due to the limitation of individual sensors, state estimation needs to be solved in a collaborative and distributed manner. Second, the objects' movement behavior model is unknown, and needs to be learned using sensor observations. In this work, for the first time, we formally formulate the problem of simultaneous state estimation and behavior learning in a sensor network. We then propose a simple yet effective solution to this new problem by extending the Gaussian process-based Bayes filters (GPBayesFilters) to an online, distributed setting. The effectiveness of the proposed method is evaluated on tracking objects with unknown movement behaviors using both synthetic data and data collected from a multi-robot platform. Rui Yu 0002, Zhenyuan Yuan, Zihan Zhou 0001 |
IROS | 1 |
| 2020 | Deep-Person: Learning discriminative deep features for person Re-Identification
Xiang Bai, Tengteng Huang, Zhiyong Dou, Rui Yu 0002, Yongchao Xu |
Pattern Recognit. | 5 |
| 2018 | Hard-Aware Point-to-Set Deep Metric for Person Re-identification
Rui Yu 0002, Zhiyong Dou, Song Bai 0001, Zhaoxiang Zhang 0001, Yongchao Xu, Xiang Bai |
ECCV (16) | 1 |
| 2017 | Divide and Fuse: A Re-ranking Approach for Person Re-identification
Rui Yu 0002, Song Bai 0001, Xiang Bai |
BMVC | 1 |