Shuchang Xu

dblp:47/6190 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Wearable AR for Restorative Breaks: How Interactive Narrative Experiences Support Relaxation for Young People
abstract
Young adults often take breaks from screen-intensive work by consuming digital content on mobile phones, which undermines rest through visual fatigue and inactivity. We introduce a design framework that embeds light break activities into media content on AR smart glasses, balancing engagement and recovery, which employs three strategies: (1) seamlessly guiding users by embedding activity cues aligned with media elements; (2) transitioning to audio-centric formats to reduce visual load while sustaining immersion; and (3) structuring sessions with "rise-peak-closure"pacing for smooth transitions. In a within-subjects study (N=16) comparing passive viewing, reminder-based breaks, and non-narrative activities, InteractiveBreak instantiated from our framework seamlessly guided activities, sustained engagement, and enhanced break quality. These findings demonstrate wearable AR's potential to support restorative relaxation by transforming breaks into engaging, meaningful experiences. © 2026 the owner/author(s).
Jin-Du Wang, Runze Cai, Shuchang Xu, Tianrui Hu, Huamin Qu, Shengdong Zhao 0001, Linping Yuan
CHI3
2026 Accurate facade parsing based on new facade dataset
abstract
Facade parsing is a vital technology for applications such as urban modeling, urban planning, and digital twin city construction. High-resolution facade images are particularly important for achieving fine-grained building reconstruction. However, existing facade datasets rarely contain images with resolutions exceeding 2K × 2K. In this paper, we introduce a new dataset featuring much higher-resolution images captured from various angles and with denser window distributions. We also propose a novel method to automatically compute the perspective transformation matrix for generating corrected facade images, which is used to create a twin version of the dataset after perspective correction. Furthermore, we introduce a new network GLNet, designed to achieve superior facade parsing results using high-resolution images as input. Experimental results on three public datasets (CMP, CFP, and ETRIMS) as well as our own dataset demonstrate that GLNet outperform existing methods in facade segmentation. The dataset and code are available at: https://github.com/OctAne0113/GLNet .
Junjie Cheng, Weijing Qin, Haichi Ma, Shuchang Xu
Adv. Eng. Informatics6
2026 SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition
abstract
Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data inevitably introduces significant performance bottlenecks. Furthermore, in the medical domain, many surgical video tasks encounter additional challenges posed by the limited availability of video data and the necessity for comprehensive spatiotemporal modeling. Recently, Parameter-Efficient Image-to-Video Transfer Learning (PEIVTL) has emerged as an efficient and effective paradigm for video action recognition tasks, which employs image-level pre-trained models with promising feature transferability and involves cross-modality temporal modeling with minimal fine-tuning. Nevertheless, the effectiveness and generalizability of this paradigm within intricate surgical domain remain unexplored. In this paper, we delve into a novel problem of efficiently adapting image-level pre-trained models to specialize in fine-grained surgical phase recognition, termed Parameter-Efficient Image-to-Surgical-Video Transfer Learning. First, we develop SurgPETL, a parameter-efficient transfer learning framework for surgical phase recognition, and conduct extensive experiments with three advanced methods based on ViTs of two distinct scales pre-trained on five large-scale natural and medical datasets. Then, we introduce the Adaptive Spatiotemporal Representation Modulation (ASRM) module, integrating a standard spatial adapter with a novel temporal adapter to capture detailed spatial features and establish connections across temporal sequences for robust spatiotemporal modeling. Extensive experiments on three challenging datasets spanning various surgical procedures demonstrate the effectiveness of SurgPETL with ASRM. SurgPETL-ASRM outperforms both parameter-efficient alternatives and state-of-the-art surgical phase recognition methods while maintaining parameter efficiency and minimizing overhead.
Shu Yang 0004, Zhiyuan Cai, Luyang Luo, Shuchang Xu, Hao Chen 0011
IEEE Trans. Medical Imaging5
2025 DanmuA11y: Making Time-Synced On-Screen Video Comments (Danmu) Accessible to Blind and Low Vision Users via Multi-Viewer Audio Discussions
abstract
By overlaying time-synced user comments on videos, Danmu creates a co-watching experience for online viewers. However, its visual-centric design poses significant challenges for blind and low vision (BLV) viewers. Our formative study identified three primary challenges that hinder BLV viewers' engagement with Danmu: the lack of visual context, the speech interference between comments and videos, and the disorganization of comments. To address these challenges, we present DanmuA11y, a system that makes Danmu accessible by transforming it into multi-viewer audio discussions. DanmuA11y incorporates three core features: (1) Augmenting Danmu with visual context, (2) Seamlessly integrating Danmu into videos, and (3) Presenting Danmu via multi-viewer discussions. Evaluation with twelve BLV viewers demonstrated that DanmuA11y significantly improved Danmu comprehension, provided smooth viewing experiences, and fostered social connections among viewers. We further highlight implications for enhancing commentary accessibility in video-based social media and live-streaming platforms.
Shuchang Xu, Xiaofu Jin, Huamin Qu, Yukang Yan
CHI1
2025 RhythmTA: A Visual-Aided Interactive System for ESL Rhythm Training via Dubbing Practice
Chang Chen 0005, Sicheng Song, Shuchang Xu, Huamin Qu, Yanna Lin
UIST3
2025 Branch Explorer: Leveraging Branching Narratives to Support Interactive 360° Video Viewing for Blind and Low Vision Users
abstract
Figure 1: Branch Explorer transforms 360° videos into branching narratives-stories that dynamically unfold based on viewer choices-to create an engaging experience for blind and low vision (BLV) users.It employs a multi-modal machine learning pipeline to generate diverse narrative paths, enabling BLV users to make choices at key branching points and explore each storyline through immersive audio guidance.The figure shows the 360° video HELP (available at: https://youtu.be/G-XZhKqQAHU).
Shuchang Xu, Xiaofu Jin, Huamin Qu, Yukang Yan
UIST1
2025 NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification
abstract
Conversational LLMs have been widely adopted by domain users with limited programming experience to solve domain problems.However, these users often face misalignment between their intent and generated code, resulting in frustration and rounds of clarification.This work first investigates the cause of this misalignment, which dues to bidirectional ambiguity: both user intents and coding tasks are inherently nonlinear, yet must be expressed and interpreted through linear prompts and code sequences.To address this, we propose direct intent-task matching, a new human-LLM interaction paradigm that externalizes and enables direct manipulation of the LLM understanding, i.e., the coding tasks and their relationships inferred by the LLM prior to code generation.As a proof-of-concept, this paradigm is then implemented in NeuroSync, which employs a knowledge distillation pipeline to extract LLM understanding, user intents, and their mappings, and enhances the alignment by allowing users to intuitively inspect and edit them via visualizations.We evaluate the algorithmic components of NeuroSync via technical experiments, and assess its overall usability and effectiveness via a user study (N=12).The results show that it enhances intent-task alignment, lowers cognitive effort, and improves coding efficiency.
Leixian Shen, Shuchang Xu, Jin-Du Wang, Jian Zhao 0010, Huamin Qu, Linping Yuan
UIST3
2025 Advancing neural aesthetic assessment of artistic images based on bundle features integration
Simin Yan, Shuchang Xu, Aiping Lei, Sanyuan Zhang
Vis. Comput.2
2025 Hybrid annotation alignment-based multi-region crop model for high-resolution image
Yuetao Yuan, Shuchang Xu, Junjie Cheng, Shudong Lin
Vis. Comput.2
2024 Memory Reviver: Supporting Photo-Collection Reminiscence for People with Visual Impairment via a Proactive Chatbot
abstract
Reminiscing with photo collections offers significant psychological benefits but poses challenges for people with visual impairment (PVI). Their current reliance on sighted help restricts the flexibility of this activity. In response, we explored using a chatbot in a preliminary study. We identified two primary challenges that hinder effective reminiscence with a chatbot: the scattering of information and a lack of proactive guidance. To address these limitations, we present Memory Reviver, a proactive chatbot that helps PVI reminisce with a photo collection through natural language communication. Memory Reviver incorporates two novel features: (1) a Memory Tree, which uses a hierarchical structure to organize the information in a photo collection; and (2) a Proactive Strategy, which actively delivers information to users at proper conversation rounds. Evaluation with twelve PVI demonstrated that Memory Reviver effectively facilitated engaging reminiscence, enhanced understanding of photo collections, and delivered natural conversational experiences. Based on our findings, we distill implications for supporting photo reminiscence and designing chatbots for PVI.
Shuchang Xu, Chang Chen 0005, Xiaofu Jin, Linping Yuan, Yukang Yan, Huamin Qu
UIST1
2024 Self-Driven Dual-Path Learning for Reference-Based Line Art Colorization Under Limited Data
abstract
Synthesizing color images based on line arts while considering the styles of reference photos is a flexible form of artistic creation that has recently attracted public attention. Previous approaches usually require large datasets at training, causing great inconvenience to the application. Besides, the sparsity of line art pictures often leads to a failure in learning valid mappings. To this end, we present SDL, a self-driven dual-path framework for reference-based line art colorization under limited data. Given small training sets containing sketch-image pairs, SDL first utilizes a novel Dynamic Pseudo Sample Generator (DPSG) to produce quantities of fake samples. Then, we introduce a dual-path network to achieve better visual effects, in which the Content-Generation Path reconstructs reliable content features to help establish multi-level correspondence in the Content-Color Aggregation Module (CCAM) of the Color-Transfer Path. Furthermore, we develop a Region-aware Contrastive Scheme (RCS) to focus on fine-grained details and a Style-augmented Contrastive Scheme (SCS) to encourage style consistency. Experiments verify the superiority of our model compared with existing works. We also demonstrate SDL outperforms state-of-the-art self-driven methods even though they adopt much more data than us ($30\times $on CelebA-HQ Dataset and$17\times $on ASCP Dataset).
Shukai Wu, Weiming Liu 0005, Shuchang Xu, Sanyuan Zhang
IEEE Trans. Circuits Syst. Video Technol.4
2024 An art-oriented pixelation method for cartoon images
Shuchang Xu, Sanyuan Zhang
Vis. Comput.2
2024 Jigsaw puzzle difficulty assessment and analysis of influencing factors based on deep learning method
Yuetao Yuan, Shuchang Xu, Shudong Lin
Vis. Comput.2
2023 Research on Strategies for Tripeaks Variant with Various Layouts
Yijie Gao, Shuchang Xu, Shunpeng Du
ICIG (4)2
2023 FlexIcon: Flexible Icon Colorization via Guided Images and Palettes
abstract
Automatic icon colorization systems show great potential value as they can serve as a source of inspiration for designers. Despite yielding promising results, previous reference-guided approaches ignore how to effectively fuse icon structure and style, leading to unpleasant color effects. Meanwhile, they cannot take free-style palettes as inputs, which is less user-friendly. To this end, we present FlexIcon, a Flexible Icon colorization model based on guided images and palettes. To promote visual quality, our model first leverages a Hybrid Multi-expert Module to aggregate better structural features, followed by dynamically integrating the global style with each individual pixel of the structure map via the Pixel-Style Aggregation Layer. We also introduce an efficient learning scheme for free-style palette-based colorization, editing, interpolation, and diverse generation. Extensive experiments demonstrate the superiority of our framework compared with state-of-the-art approaches. In addition, we contribute a Mandala dataset to the multimedia community and further validate the application value of the proposed model.
Shukai Wu, Shuchang Xu, Weiming Liu 0005, Sanyuan Zhang
ACM Multimedia3
2023 Unsupervised industrial anomaly detection with diffusion models
Haohao Xu, Shuchang Xu, Wenzhen Yang
J. Vis. Commun. Image Represent.2
2023 Image recoloring based on fast and flexible palette extraction
Simin Yan, Shuchang Xu, Wenzhen Yang, Sanyuan Zhang
Multim. Tools Appl.2
2022 Improving Reference-Based Image Colorization For Line Arts Via Feature Aggregation And Contrastive Learning
abstract
The tremendous semantic discrepancy between the line art drawings without texture and the reference pictures containing rich color challenges current image-to-image translation models. Previous works attempt to establish cross-domain correspondence. However, they fail to capture more detailed features. A Reference-based Line art Translation Network (RLTN) is introduced with a Multi-level Feature Aggregation Module (MFAM) to improve the performance. The MFAM concentrates on more meaningful information for feature matching by utilizing the Multi-stream High Frequency Block (MHFB) and the Pixel-wise Correlation Block (PCB). We also employ the Channel-level Attention Block (CAB) and the Spatial-level Attention Block (SAB) for a better fusion of features. Moreover, a Style-based Contrastive Loss (SCL) is proposed to maintain the style similarity between the synthesized images and the reference examples. Experiments conducted on three datasets demonstrate the effectiveness of our model in producing more pleasing visual effects compared with state-of-the-art approaches.
Shukai Wu, Qingqin Wang, Shuchang Xu, Sanyuan Zhang
ICASSP3
2022 RefFaceNet: Reference-based Face Image Generation from Line Art Drawings
Shukai Wu, Weiming Liu 0005, Qingqin Wang, Sanyuan Zhang, Zhenjie Hong, Shuchang Xu
Neurocomputing6
2021 Tactile Compass: Enabling Visually Impaired People to Follow a Path with Continuous Directional Feedback
abstract
Accurate and effective directional feedback is crucial for an electronic traveling aid device that guides visually impaired people in walking through paths. This paper presents Tactile Compass, a hand-held device that provides continuous directional feedback with a rotatable needle pointing toward the planned direction. We conducted two lab studies to evaluate the effectiveness of the feedback solution. Results showed that, using Tactile Compass, participants could reach the target direction in place with a mean deviation of 3.03° and could smoothly navigate along paths of 60cm width, with a mean deviation from the centerline of 12.1cm. Subjective feedback showed that Tactile Compass was easy to learn and use.
Guanhong Liu, Tianyu Yu 0001, Chun Yu, Haiqing Xu 0001, Shuchang Xu, Ciyuan Yang, Haipeng Mi, Yuanchun Shi
CHI5
2019 Separating Skin Surface Reflection Component from Single Color Image
Shuchang Xu, Zhengwei Yao
ICIG (2)1
2019 Accurate and Low-Latency Sensing of Touch Contact on Any Surface with Finger-Worn IMU Sensor
abstract
Head-mounted Mixed Reality (MR) systems enable touch in­teraction on any physical surface. However, optical methods (i.e., with cameras on the headset) have difficulty in determin­ing the touch contact accurately. We show that a finger ring with Inertial Measurement Unit (IMU) can substantially im­prove the accuracy of contact sensing from 84.74% to 98.61% (f1 score), with a low latency of 10 ms. We tested different ring wearing positions and tapping postures (e.g., with different fingers and parts). Results show that an IMU-based ring worn on the proximal phalanx of the index finger can accurately sense touch contact of most usable tapping postures. Partici­pants preferred wearing a ring for better user experience. Our approach can be used in combination with the optical touch sensing to provide robust and low-latency contact detection.
Yizheng Gu, Chun Yu, Zhipeng Li 0001, Shuchang Xu, Xiaoying Wei, Yuanchun Shi
UIST5
2018 A machine learning framework to identify detailed routing short violations from a placed netlist
abstract
Detecting and preventing routing violations has become a critical issue in physical design, especially in the early stages. Lack of correlation between global and detailed routing congestion estimations and the long runtime required to frequently consult a global router adds to the problem. In this paper, we propose a machine learning framework to predict detailed routing short violations from a placed netlist. Factors contributing to routing violations are determined and a supervised neural network model is implemented to detect these violations. Experimental results show that the proposed method is able to predict on average 90% of the shorts with only 7% false alarms and considerably reduced computational time.
Aysa Fakheri Tabrizi, Nima Karimpour Darav, Shuchang Xu, Logan Rakai, Ismail Bustany, Andrew A. Kennings, Laleh Behjat
DAC3
2008 Automatic skin decomposition based on single image
Shuchang Xu, Xiuzi Ye, Franck Giron, Jean Luc Lévêque, Bernard Querleux
Comput. Vis. Image Underst.1
2005 Uniform color transfer
abstract
We present in this paper a general algorithm called uniform color transfer for transferring dominant colors in one image to another. Our algorithm frees users from interactions required for colorizing grayscale image or transferring color between images. Different from existing algorithms, we cluster the source color image into regions using the GMM-EM method, and the target image using the K-means algorithm. We then impose chromatic mean values on corresponding source image regions in the target image. Experiments showed that our algorithm can give better results.
Shuchang Xu, Yin Zhang 0006, Sanyuan Zhang, Xiuzi Ye
ICIP (3)1