Liuxin Zhang

dblp:90/7642 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
14since 2021 · last 2025
0000-0002-6779-4920ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Chorus of the Past: Toward Designing a Multi-agent Conversational Reminiscence System with Digital Artifacts for Older Adults
abstract
Reminiscence has been shown to provide benefits for older adults, but traditionally relies on personal photos as memory cues and interactions with real people who may not always be available. We present ReminiBuddy, a novel LLM-powered multi-agent conversational system, which allows older adults to engage with two distinct agents - one embodying an older identity and the other a younger identity - while using not only personal photos but also 3D models of generic nostalgic objects as memory cues. Our study, with older adult participants, found that the conversational approach both enjoyable and beneficial for reminiscence. While the younger agent was perceived as more emotionally engaging, the older one fostered greater resonance in content. Personal photos prompted autobiographical memories, whereas 3D generic nostalgic objects evoked shared memories of an era, contributing to a more multifaceted reminiscence experience. We further present design implications for better supporting older adults in reminiscing with LLM-powered conversational agents.
Jingwei Sun 0005, Nianlong Li, Zhangwei Lu, Liuxin Zhang, Yu Zhang 0124, Qianying Wang 0002, Mingming Fan 0001
CHI7
2025 GBC-Splat: Generalizable Gaussian-Based Clothed Human Digitalization under Sparse RGB Cameras
abstract
We present an efficient approach for generalizable clothed human digitalization, termed GBC-Splat. Unlike previous methods that necessitate per-subject optimizations or discount watertight geometry, the proposed method is dedicated to reconstructing complete human shapes and Gaussian Splatting via sparse view RGB inputs in a feed-forward manner. We first extract a fine-grained mesh using a combination of implicit occupancy field regression and explicit disparity estimation between views. The reconstructed high-quality geometry allows us to easily anchor Gaussian primitives to mesh surface according to surface normal and texture, which allows 6-DoF photorealistic novel view synthesis. In addition, we introduce a simple yet effective algorithm to subdivide Gaussian primitives in high-frequency areas to further enhance the visual quality. Without the assistance of human parametric models, our method can tackle loose garments, such as dresses and costumes. Our method outperforms state-of-the-art methods in terms of novel view synthesis while keeping high efficiency, enabling the potential of deployment in real-time applications.
Hanzhang Tu, Zhanfeng Liao, Boyao Zhou, Shunyuan Zheng, Liuxin Zhang, Qianying Wang 0002, Yebin Liu
CVPR6
2025 A Dual-Stick Controller for Enhancing Raycasting Interactions with Virtual Objects
abstract
This work presents Dual-Stick, a novel controller with two sticks connected at the end that innovates a Dual-Ray interaction paradigm to enrich raycasting input in Virtual Reality (VR). Dual-Stick leverages the inherent human dexterity in using everyday tools such as clamps and tweezers to adjust the relative angle between two sticks. This design supports Dual-Ray interactions that provide with a heuristics-based enhanced mechanism. It also offers more flexible manipulation by taking advantages of additional degrees of freedom provided by clamping angle. We conducted two studies to evaluate the effectiveness of Dual-Ray in target selection and manipulation tasks. The results indicated that Dual-Ray significantly improved efficiency in target selection compared to single-ray input but did not outperform the enhanced single-ray technique. In terms of manipulation, Dual-Ray effectively reduced completion time and mode switching compared to single-ray input.
Nianlong Li, Zhenxuan He, Luyao Shen, Tianren Luo, Teng Han, Boyu Gao 0003, Yu Zhang 0199, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002
VR9
2024 See Widely, Think Wisely: Toward Designing a Generative Multi-agent System to Burst Filter Bubbles
abstract
The proliferation of AI-powered search and recommendation systems has accelerated the formation of “filter bubbles” that reinforce people’s biases and narrow their perspectives. Previous research has attempted to address this issue by increasing the diversity of information exposure, which is often hindered by a lack of user motivation to engage with. In this study, we took a human-centered approach to explore how Large Language Models (LLMs) could assist users in embracing more diverse perspectives. We developed a prototype featuring LLM-powered multi-agent characters that users could interact with while reading social media content. We conducted a participatory design study with 18 participants and found that multi-agent dialogues with gamification incentives could motivate users to engage with opposing viewpoints. Additionally, progressive interactions with assessment tasks could promote thoughtful consideration. Based on these findings, we provided design implications with future work outlooks for leveraging LLMs to help users burst their filter bubbles.
Yu Zhang 0124, Jingwei Sun 0005, Cen Yao, Mingming Fan 0001, Liuxin Zhang, Qianying Wang 0002, Xin Geng 0001, Yong Rui
CHI6
2024 Low-Complexity 3D-Vision Conferencing System based on Accelerated RIFE Model
abstract
Recent advancements in telecommunication technologies have exceeded the requirements of numerous video-based Real-Time Communication (RTC) applications. Mean-while, in light of the growing demand for immersive 3D visual experience, researchers are currently focusing on developing next-generation telepresence systems. This paper presents a novel immersive conferencing system that offers a smooth, high-fidelity, and life-size autostereoscopic display of remote user portraits. To generate the binocular stereo vision in real time, an adaptive low-complexity view synthesis method based on an accelerated Real-time Intermediate Flow Estimation (RIFE) model is employed, which performs direct cross-view generation based on decoded multi-view videos and tracked eye positions of the watching user. Thanks to eliminating the complex 3D modeling procedure that relies on depth images, the proposed system requires significantly fewer computational resources and lower video transmission bandwidth compared to existing immersive conferencing systems. Therefore, the proposed system is low-cost and flexible when accommodating diverse conferencing environments.
Hongyue Huang, Hongbo Ning, Haopeng Lu, Qi Zhang 0042, Yanpeng Liang, Wanjun Lyu, Chuanmin Jia, Xinfeng Zhang 0001, Liuxin Zhang, Siwei Ma 0001
PCS10
2024 Towards Workplace Metaverse: A Human-Centered Approach for Designing and Evaluating XR Virtual Displays
abstract
Work is becoming more and more flexible nowadays. It can take place in the office, at home, or even on the go; tasks may encompass activities in the form of documents, multimedia, or 3D models in virtual space. Under such circumstances, personal computers (PC), the most widely used productivity devices for work today, cannot well address the diverse needs such as screen size, privacy, and flexibility to display diverse content formats (e.g., 2D to 3D), due to their fixed hardware specs. We believe the solution lies in Extended Reality (XR). In this article, we explored how XR glasses can be used for PC’s virtual extended displays, and conducted user interviews and usability tests to propose a systematic user experience design and evaluation framework. We discovered that the design space encompasses four dimensions general placement, display specs, operating system integration, and interaction behaviors) and summarized users’ corresponding preferences. We proposed a quality-of-experience (QoE) evaluation framework for XR virtual displays consisting of visual quality, visual fatigue and discomfort, as well as immersiveness, and identified clarity as the most significant factor that affects user satisfaction. Our design and evaluation frameworks could serve as a resource for both practitioners and scholars with an interest in the design and evaluation of virtual displays.
Yu Zhang 0124, Jingwei Sun 0005, Qicheng Ding, Liuxin Zhang, Qianying Wang 0002, Xin Geng 0001, Yong Rui
Int. J. Hum. Comput. Interact.4
2023 Side-by-Side vs Face-to-Face: Evaluating Colocated Collaboration via a Transparent Wall-sized Display
abstract
Traditional wall-sized displays mostly only support side-by-side co-located collaboration, while transparent displays naturally support face-to-face interaction. Many previous works assume transparent displays support collaboration. Yet it is unknown how exactly its afforded face-to-face interaction can support loose or close collaboration, especially compared to the side-by-side configuration offered by traditional large displays. In this paper, we used an established experimental task that operationalizes different collaboration coupling and layout locality, to compare pairs of participants collaborating side-by-side versus face-to-face in each collaborative situation. We compared quantitative measures and collected interview and observation data to further illustrate and explain our observed user behavior patterns. The results showed that the unique face-to-face collaboration brought by transparent display can result in more efficient task performance, different territorial behavior, and both positive and negative collaborative factors. Our findings provided empirical understanding about the collaborative experience supported by wall-sized transparent displays and shed light on its future design.
Jiangtao Gong, Mengdi Chu, Minghao Luo, Liuxin Zhang, Yaqiang Wu, Qianying Wang 0002, Can Liu 0003
Proc. ACM Hum. Comput. Interact.7
2022 Remote Co-teaching in Rural Classroom: Current Practices, Impacts, and Challenges
abstract
The shortage of high-quality teachers is one of the biggest educational problems faced by underdeveloped areas. With the development of information and communication technologies (ICTs), China has begun a remote co-teaching intervention program using ICTs for rural classes, forming a unique “co-teaching classroom”. We conducted semi-structured interviews with nine remote urban teachers and twelve local rural teachers. We identified the remote co-teaching classes’ standard practices and co-teachers’ collaborative work process. We also found that remote teachers’ high-quality class directly impacted local teachers and students. Furthermore, interestingly, local teachers were also actively involved in making indirect impacts on their students by deeply coordinating with remote teachers and adapting the resources offered by the remote teachers. We conclude by summarizing and discussing the challenges faced by teachers, lessons learned from the current program, and related design implications to achieve a more adaptive and sustainable ICT4D program design.
Siling Guo, Tianchen Sun, Jiangtao Gong, Zhicong Lu, Liuxin Zhang, Qianying Wang 0002
CHI5
2022 A Novel Distilled Generative Essay Polish System via Hierarchical Pre-Training
abstract
In language processing tasks, the most important process in automated text polishment always consists of text correction and text supplementation. Finding that text polishment is a necessary step in the field of English essay reviewing, we are motivated to be the first of building an end-to-end automated English essay polish system, to support writing instruction. There were independent methods for text correction tasks and text supplementation tasks, but when combining them for essay polishment tasks, conflicts arise from their interplay. In this paper, we propose a polish system that elegantly performs text correction and text supplementation at the same time, achieving an improved revision quality. Furthermore, we design a closed-loop essay polishing process, made up of a Rewriting Model and a Scoring Model, which refers to modified GPT2 and ensembled Bert respectively. The rewriting process targets the deficiencies of the essays by a threshold controlled mechanism. Lastly, the performance of our proposed system is further enhanced by an optimization method. Intensive experiments on both real data and simulated data have shown score improvements on full essays by Scoring Model, as well as higher text correction accuracy and longer text supplementation length.
Qichuan Yang, Liuxin Zhang, Yang Zhang 0002, Jinghua Gao
IJCNN2
2021 All in One Group: Current Practices, Lessons and Challenges of Chinese Home-School Communication in IM Group Chat
abstract
When schools and families form a good partnership, children benefit. With the recent flourishing of communication apps, families and schools in China have shifted their primary communication channels to chat groups hosted on popular instant-messenger(IM) tools such as WeChat and QQ. With an interview study consisting of 18 parents and 9 teachers, followed by a survey study with 210 teachers, we found that IM group chat has become the most popular way that the majority of parents and teachers communicate, from among the many different channels available. While there are definite advantages to this kind of group chat, we also found a number of problematic issues, including a lack of privacy and repeated negative feedback shared by both parents and teachers. We discuss our results on how IM-based group chat could affect Chinese teachers’ authoritative figures, affect Chinese teacher’s work-life balance and potentially compromise Chinese students’ privacy.
Jiangtao Gong, Zhicong Lu, Qicheng Ding, Yu Zhang 0124, Liuxin Zhang, Qianying Wang 0002
CHI6
2021 A Bi-modal Automated Essay Scoring System for Handwritten Essays
abstract
During the past few decades, Automated Essay Scoring (AES) technology has been widely used to alleviate the workload of teachers and improve the feedback cycle in educational systems. However, the scoring of handwritten essays poses great challenges for existing systems, since most of them only take text as input without consideration of errors or bias which may be introduced by Optical Character Recognition (OCR) as a necessary pre-processing step. This paper proposes VisualAES, a bimodal automated essay scoring system that utilizes both textual and visual features for handwritten essay scoring. Specifically, we first employ three powerful pre-trained transformer-based models as the backbone, and extend them to take both textual features and visual features extracted by Faster R-CNN. Then, a stacking ensemble model is subsequently adopted to robustly map their outputs to a final score. We evaluate VisualAES with the public Automated Student Assessment Prize (ASAP) dataset and our proposed handwritten Chinese Students' handwritten essay dataset (ChnStd). Results show that the proposed VisualAES outperforms all state-of-the-art methods on both datasets. More importantly, by incorporating handwritten image information, we also achieve a further performance improvement on ChnStd and reduce the side effects of OCR.
Jinghua Gao, Qichuan Yang, Yang Zhang 0002, Liuxin Zhang
IJCNN4
2021 VisDG: Towards Spontaneous Visual Dialogue Generation
abstract
Current vision and language understanding tasks can be generally divided into two categories: image (video) description and visual question answering. Both of them aim to train a model that either straightforwardly generates a description or answers the predefined questions based on an image or a sequence of images from a spectators perspective. However, a large proportion of real-world human interactions also involve spontaneous dialogue exchanges among multiple speakers as well as dynamic visual-textual context, which requires an AI agent to hold a natural and open-ended dialog with humans in a first-person manner based on both visual and textual context. To move closer towards achieving such a spontaneous multimodal conversation, we introduce a new visual dialogue generation dataset (VisDG) based on keyframes and corresponding subtitles extracted from Friends - an American sitcom television series. Specifically, given a start frame and its corresponding dialogue text, the agent has to generate both a meaningful textual response as well as a correct image candidate for the latter part of the dialogue turn. Furthermore, we also propose an end-to-end image-text synergistic network (ITSN) for the task, which outperforms several sophisticated baselines on the proposed VisDG.
Qichuan Yang, Liuxin Zhang, Yang Zhang 0002, Jinghua Gao
IJCNN2
2021 HoloBoard: a Large-format Immersive Teaching Board based on pseudo HoloGraphics
abstract
In this paper, we present HoloBoard, an interactive large-format pseduo-holographic display system for lecture based classes. With its unique properties of immersive visual display and transparent screen, we designed and implemented a rich set of novel interaction techniques like immersive presentation, role-play, and lecturing behind the scene that are potentially valuable for lecturing in class. We conducted a controlled experimental study to compare a HoloBoard class with a normal class through measuring students’ learning outcomes and three dimensions of engagement (i.e., behavioral, emotional, and cognitive engagement). We used pre-/post- knowledge tests and multimodal learning analytics to measure students’ learning outcomes and learning experiences. Results indicated that the lecture-based class utilizing HoloBoard lead to slightly better learning outcomes and a significantly higher level of student engagement. Given the results, we discussed the impact of HoloBoard as an immersive media in the classroom setting and suggest several design implications for deploying HoloBoard in immersive teaching practices.
Jiangtao Gong, Teng Han, Siling Guo, Jiannan Li, Siyu Zha, Liuxin Zhang, Feng Tian 0001, Qianying Wang 0002, Yong Rui
UIST6
2021 Grabbing the Long Tail: A data normalization method for diverse and informative dialogue generation
Zhiqiang Zhan, Yang Zhang 0002, Jiangtao Gong, Qianying Wang 0002, Liuxin Zhang
Neurocomputing7
2018 Adaptive Learning of Local Semantic and Global Structure Representations for Text Classification
abstract
Representation learning is a key issue for most Natural Language Processing (NLP) tasks. Most existing representation models either learn little structure information or just rely on pre-defined structures, leading to degradation of performance and generalization capability. This paper focuses on learning both local semantic and global structure representations for text classification. In detail, we propose a novel Sandwich Neural Network (SNN) to learn semantic and structure representations automatically without relying on parsers. More importantly, semantic and structure information contribute unequally to the text representation at corpus and instance level. To solve the fusion problem, we propose two strategies: Adaptive Learning Sandwich Neural Network (AL-SNN) and Self-Attention Sandwich Neural Network (SA-SNN). The former learns the weights at corpus level, and the latter further combines attention mechanism to assign the weights at instance level. Experimental results demonstrate that our approach achieves competitive performance on several text classification tasks, including sentiment analysis, question type classification and subjectivity classification. Specifically, the accuracies are MR (82.1%), SST-5 (50.4%), TREC (96%) and SUBJ (93.9%).
Zhiqiang Zhan, Qichuan Yang, Yang Zhang 0002, Changjian Hu, Zhensheng Li, Liuxin Zhang, Zhiqiang He 0002
COLING7
2011 Multiview Visibility Estimation for Image-Based Modeling
Liuxin Zhang, Ming-Tao Pei, Yunde Jia
J. Comput. Sci. Technol.1
2010 Visibility of Multiple Cameras in a Scene with Unknown Geometry
abstract
In this paper, we investigate the problem of determining the visible regions of multiple cameras in a 3D scene without a priori knowledge of the scene geometry. Our approach is based on a variational energy functional where both the unresolved visibility information of multiple cameras and the unknown scene geometry are included. We cast visibility estimation and scene geometry reconstruction as an optimization of the variational energy functional amenable for minimization with the Euler-Lagrange driven evolution. Starting from any initial value, the accurate visibility of multiple cameras as well as the true scene geometry can be obtained at the end of the evolution. Experimental results show the validity of our approach.
Liuxin Zhang, Yunde Jia
ICPR1
2010 Layer-Constraint-Based Visibility for Volumetric Multi-view Reconstruction
Yumo Yang, Liuxin Zhang, Yunde Jia
MMM2
2010 Surface Reconstruction from Images Using a Variational Formulation
Liuxin Zhang, Yunde Jia
MMM1
2007 A Plane-based Calibration for Multi-camera Systems
abstract
In this paper, a plane-based calibration algorithm with use for calibrating large complexes of linear projective cameras of varying intrinsic and extrinsic parameters is proposed. A planar pattern with known reference points placed at a few different locations is only required as a calibration object. All the cameras do not have to see this planar pattern at all locations simultaneously, and only reasonable overlap between them is necessary. We divide these cameras into groups according to their positions and orientations, to make sure cameras in each group have a common field of view, and estimate the intrinsic parameters of each camera via a plane-based technique. Common views of planes are used to represent each camera in the world coordinate system of its own group and determine the relationships between these world coordinate systems. Rodrigues formula is adopted to deal with the inconsistency when recovering the rigid displacements between cameras. At last, all these cameras can be completely calibrated in a uniform world coordinate system robustly and accurately.
Liuxin Zhang, Yunde Jia
CAD/Graphics1