EDBT 2026 Demo / reviewers in the wild / expert
Zhaoyang Xia
dblp:293/5465
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0003-3536-5387ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Large Sign Language Models: Toward 3D American Sign Language TranslationabstractWe present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals’ virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language. Xiaoxiao He, Di Liu 0003, Zhaoyang Xia, Chaowei Tan, Vivian Li, Bo Liu 0005, Dimitris N. Metaxas, Mubbasir Kapadia |
WACV | 4 |
| 2024 | ProxEdit: Improving Tuning-Free Real Image Editing with Proximal GuidanceabstractDDIM inversion has revealed the remarkable potential of real image editing within diffusion-based methods. However, the accuracy of DDIM reconstruction degrades as larger classifier-free guidance (CFG) scales being used for enhanced editing. Null-text inversion (NTI) optimizes null embeddings to align the reconstruction and inversion trajectories with larger CFG scales, enabling real image editing with cross-attention control. Negative-prompt inversion (NPI) further offers a training-free closed-form solution of NTI. However, it may introduce artifacts and is still constrained by DDIM reconstruction quality. To overcome these limitations, we propose proximal guidance and incorporate it to NPI with cross-attention control. We enhance NPI with a regularization term and inversion guidance, which reduces artifacts while capitalizing on its training-free nature. Additionally, we extend the concepts to incorporate mutual self-attention control, enabling geometry and layout alterations in the editing process. Our method provides an efficient and straightforward approach, effectively addressing real image editing tasks with minimal computational overhead. Ligong Han, Song Wen 0001, Kunpeng Song, Mengwei Ren, Ruijiang Gao, Anastasis Stathopoulos, Xiaoxiao He, Yuxiao Chen 0002, Di Liu 0003, Qilong Zhangli, Jindong Jiang, Zhaoyang Xia, Akash Srivastava, Dimitris N. Metaxas |
WACV | 14 |
| 2023 | SCPNet: Semantic Scene Completion on Point CloudabstractTraining deep models for semantic scene completion (SSC) is challenging due to the sparse and incomplete input, a large quantity of objects of diverse scales as well as the inherent label noise for moving objects. To address the above-mentioned problems, we propose the following three solutions: 1) Redesigning the completion sub-network. We design a novel completion sub-network, which consists of several Multi-Path Blocks (MPBs) to aggregate multi-scale features and is free from the lossy downsampling operations. 2) Distilling rich knowledge from the multi-frame model. We design a novel knowledge distillation objective, dubbed Dense-to-Sparse Knowledge Distillation (DSKD). It transfers the dense, relation-based semantic knowledge from the multi-frame teacher to the single-frame student, significantly improving the representation learning of the single-frame model. 3) Completion label rectification. We propose a simple yet effective label rectification strategy, which uses off-the-shelf panoptic segmentation labels to remove the traces of dynamic objects in completion labels, greatly improving the performance of deep models especially for those moving objects. Extensive experiments are conducted in two public SSC benchmarks, i.e., SemanticKITTI and SemanticPOSS. Our SCPNet ranks 1st on SemanticKITTI semantic scene completion challenge and surpasses the competitive S3CNet [3] by 7.2 mIoU. SCP-Net also outperforms previous completion algorithms on the SemanticPOSS dataset. Besides, our method also achieves competitive results on SemanticKITTI semantic segmentation tasks, showing that knowledge learned in the scene completion is beneficial to the segmentation task. Zhaoyang Xia, Youquan Liu, Xin Li 0110, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yuenan Hou, Yu Qiao 0001 |
CVPR | 1 |
| 2023 | UniSeg: A Unified Multi-Modal LiDAR Segmentation Network and the OpenPCSeg CodebaseabstractPoint-, voxel-, and range-views are three representative forms of point clouds. All of them have accurate 3D measurements but lack color and texture information. RGB images are a natural complement to these point cloud views and fully utilizing the comprehensive information of them benefits more robust perceptions. In this paper, we present a unified multi-modal LiDAR segmentation network, termed UniSeg, which leverages the information of RGB images and three views of the point cloud, and accomplishes semantic segmentation and panoptic segmentation simultaneously. Specifically, we first design the Learnable cross-Modal Association (LMA) module to automatically fuse voxel-view and range-view features with image features, which fully utilize the rich semantic information of images and are robust to calibration errors. Then, the enhanced voxel-view and range-view features are transformed to the point space, where three views of point cloud features are further fused adaptively by the Learnable cross-View Association module (LVA). Notably, UniSeg achieves promising results in three public benchmarks, i.e., SemanticKITTI, nuScenes, and Waymo Open Dataset (WOD); it ranks 1st on two challenges of two benchmarks, including the LiDAR semantic segmentation challenge of nuScenes and panoptic segmentation challenges of SemanticKITTI. Besides, we construct the OpenPCSeg codebase, which is the largest and most comprehensive outdoor LiDAR segmentation codebase. It contains most of the popular outdoor LiDAR segmentation algorithms and provides reproducible implementations. The OpenPCSeg codebase will be made publicly available at https://github.com/PJLab-ADG/PCSeg. Youquan Liu, Runnan Chen, Xin Li 0110, Lingdong Kong, Yuchen Yang 0003, Zhaoyang Xia, Yeqi Bai, Xinge Zhu, Yuexin Ma, Yikang Li 0002, Yu Qiao 0001, Yuenan Hou |
ICCV | 6 |
| 2023 | Learning Causality-inspired Representation Consistency for Video Anomaly DetectionabstractVideo anomaly detection is an essential yet challenging task in the multimedia community, with promising applications in smart cities and secure communities. Existing methods attempt to learn abstract representations of regular events with statistical dependence to model the endogenous normality, which discriminates anomalies by measuring the deviations to the learned distribution. However, conventional representation learning is only a crude description of video normality and lacks an exploration of its underlying causality. The learned statistical dependence is unreliable for diverse regular events in the real world and may cause high false alarms due to over generalization. Inspired by causal representation learning, we think that there exists a causal variable capable of adequately representing the general patterns of regular events in which anomalies will present significant variations. Therefore, we design a causality-inspired representation consistency (CRC) framework to implicitly learn the unobservable causal variables of normality directly from available normal videos and detect abnormal events with the learned representation consistency. Extensive experiments show that the causality-inspired normality is robust to regular events with label-independent shifts, and the proposed CRC framework can quickly and accurately detect various complicated anomalies from real-world surveillance videos. Yang Liu 0246, Zhaoyang Xia, Mengyang Zhao 0002, Donglai Wei 0002, Siao Liu, Bobo Ju, Gaoyun Fang, Jing Liu 0050 |
ACM Multimedia | 2 |
| 2022 | Hierarchically Self-supervised Transformer for Human Skeleton Representation Learning
Yuxiao Chen 0002, Long Zhao 0003, Yu Tian 0003, Zhaoyang Xia, Shijie Geng, Ligong Han, Dimitris N. Metaxas |
ECCV (26) | 5 |
| 2022 | TransFusion: Multi-view Divergent Fusion for Medical Image Segmentation with Transformers
Di Liu 0003, Yunhe Gao, Qilong Zhangli, Ligong Han, Xiaoxiao He, Zhaoyang Xia, Song Wen 0001, Zhennan Yan, Mu Zhou, Dimitris N. Metaxas |
MICCAI (5) | 6 |
| 2022 | Region Proposal Rectification Towards Robust Instance Segmentation of Biological Images
Qilong Zhangli, Jingru Yi, Di Liu 0003, Xiaoxiao He, Zhaoyang Xia, Ligong Han, Yunhe Gao, Song Wen 0001, Haiming Tang, He Wang 0016, Mu Zhou, Dimitris N. Metaxas |
MICCAI (4) | 5 |
| 2022 | Person Identification With Millimeter-Wave Radar in Realistic Smart Home ScenariosabstractCompared with visual sensors that have light dependence and privacy intrusion issues, non-intrusion millimeter-wave (mmW) radars are more suitable for the daily person identification. In a realistic home scenario, there are new challenges that are not taken into account in the existing research. This letter attempts to address these issues such as multipath interference, complex walking process, and recognition robustness in smart home scenarios and designs a lightweight multi-branch convolutional neural network (CNN) with an Inception-Pool module and a Residual-Pool module to learn and classify gait Doppler features. The experimental results in a home living room scenario indicate that the designed mmW radar person identification system can achieve accurate and robust real-time identification performance. Zhaoyang Xia, Genming Ding, Feng Xu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | American Sign Language Video Anonymization to Support Online Participation of Deaf and Hard of Hearing UsersabstractWithout a commonly accepted writing system for American Sign Language (ASL), Deaf or Hard of Hearing (DHH) ASL signers who wish to express opinions or ask questions online must post a video of their signing, if they prefer not to use written English, a language in which they may feel less proficient. Since the face conveys essential linguistic meaning, the face cannot simply be removed from the video in order to preserve anonymity. Thus, DHH ASL signers cannot easily discuss sensitive, personal, or controversial topics in their primary language, limiting engagement in online debate or inquiries about health or legal issues. We explored several recent attempts to address this problem through development of “face swap” technologies to automatically disguise the face in videos while preserving essential facial expressions and natural human appearance. We presented several prototypes to DHH ASL signers (N=16) and examined their interests in and requirements for such technology. After viewing transformed videos of other signers and of themselves, participants evaluated the understandability, naturalness of appearance, and degree of anonymity protection of these technologies. Our study revealed users’ perception of key trade-offs among these three dimensions, factors that contribute to each, and their views on transformation options enabled by this technology, for use in various contexts. Our findings guide future designers of this technology and inform selection of applications and design features. Sooyeon Lee, Abraham Glasser, Becca Dingman, Zhaoyang Xia, Dimitris N. Metaxas, Carol Neidle, Matt Huenerfauth |
ASSETS | 4 |
| 2021 | Multidimensional Feature Representation and Learning for Robust Hand-Gesture Recognition on Commercial Millimeter-Wave RadarabstractThis article presents a robust hand-gesture recognition method via multidimensional feature representation and learning specifically designed for commercial frequency-modulated continuous wave (FMCW) multi-input multi-output (MIMO) millimeter-wave radar. First, the optimal configuration of the radar system parameters for the hand-gesture recognition scenario is investigated and a standard procedure to determine the system configuration is given. Then a moving scattering center model is proposed to represent the 3-D point cloud in the range-Doppler (RD)-angular multidimensional feature space. A scattering point detection and tracking algorithm is presented based on a set of motion constraints in terms of position, velocity, and acceleration. It is derived from the space-time continuity of a nonrigid target. Finally, a lightweight multichannel convolutional neural network (CNN) is designed to learn and classify multidimensional gesture features including radial RD and tangential azimuth-elevation. Extensive experiments are carried out with the developed system and a large data set is obtained to train and test the classifier. The results show that the proposed gesture recognition method can effectively distinguish gestures that are easily confused in the RD domain and achieve robust performances under various conditions. Zhaoyang Xia, Yixiang Luomei, Chenglong Zhou, Feng Xu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |