Xiaoxue Chen

dblp:159/8776 · DBLP profile ↗
← Back
22ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Hoodie: Hierarchical point cloud and latent code diffusion for joint and conditional generation
Zhenyu Ding, Guiyu Zhang, Huan-ang Gao, Xiaoxue Chen, Zhaoxin Fan, Ning Ding 0006, Hao Zhao 0002
Neurocomputing4
2026 Ultraman: ultra-fast and high-resolution texture generation for 3D human reconstruction from a single image
Mingjin Chen, Huan-ang Gao, Xiaoxue Chen, Zhaoxin Fan, Hao Zhao 0002
Mach. Vis. Appl.4
2025 InvRGB+L: Inverse Rendering of Complex Scenes with Unified Color and LiDAR Reflectance Modeling
abstract
We present InvRGB+L, a novel inverse rendering model that reconstructs large, relightable, and dynamic scenes from a single RGB+LiDAR sequence. Conventional inverse graphics methods rely primarily on RGB observations and use LiDAR mainly for geometric information, often resulting in suboptimal material estimates due to visible light interference. We find that LiDAR's intensity values-captured with active illumination in a different spectral range-offer complementary cues for robust material estimation under variable lighting. Inspired by this, InvRGB+L leverages LiDAR intensity cues to overcome challenges inherent in RGB-centric inverse graphics through two key innovations: (1) a novel physics-based LiDAR shading model and (2) RGB-LiDAR material consistency losses. The model produces novel-view RGB and LiDAR renderings of urban and indoor scenes and supports relighting, night simulations, and dynamic object insertions, achieving results that surpass current state-of-the-art methods in both scene-level urban inverse rendering and LiDAR simulation.
Xiaoxue Chen, Bhargav Chandaka, Chih-Hao Lin, Ya-Qin Zhang, David A. Forsyth, Shenlong Wang
ICCV1
2025 CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting
abstract
Vehicle-to-everything (V2X) communication plays a crucial role in autonomous driving, enabling cooperation between vehicles and infrastructure. While simulation has significantly contributed to various autonomous driving tasks, its potential for data generation and augmentation in V2X scenarios remains underexplored. In this paper, we introduce CRUISE, a comprehensive reconstruction-and-synthesis framework designed for V2X driving environments. CRUISE employs decomposed Gaussian Splatting to accurately reconstruct real-world scenes while supporting flexible editing. By decomposing dynamic traffic participants into editable Gaussian representations, CRUISE allows for seamless modification and augmentation of driving scenes. Furthermore, the framework renders images from both ego-vehicle and infrastructure views, enabling large-scale V2X dataset augmentation for training and evaluation. Our experimental results demonstrate that: 1) CRUISE reconstructs real-world V2X driving scenes with high fidelity; 2) using CRUISE improves 3D detection across ego-vehicle, infrastructure, and cooperative views, as well as cooperative 3D tracking on the V2X-Seq benchmark; and 3) CRUISE effectively generates challenging corner cases. The code will be publicly available at https://github.com/SainingZhang/CRUISE.
Haoran Xu 0003, Saining Zhang, Peishuo Li, Baijun Ye, Xiaoxue Chen, Huan-ang Gao, Jv Zheng, Ziqiao Peng, Run Miao, Jinrang Jia, Yifeng Shi, Guangqi Yi, Hang Zhao 0021, Hao Tang 0005, Hongyang Li 0001, Kaicheng Yu, Hao Zhao 0002
IROS5
2024 Locate N' Rotate: Two-Stage Openable Part Detection with Foundation Model Priors
Siqi Li 0009, Xiaoxue Chen, Haoyu Cheng, Guyue Zhou, Hao Zhao 0002, Guanzhong Tian
ACCV (7)2
2024 Time-conditioned Illumination for Inverse Rendering of Outdoor Scenes
Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang
BMVC1
2024 Drone-assisted Road Gaussian Splatting with Cross-view Uncertainty
Saining Zhang, Baijun Ye, Xiaoxue Chen, Yuantao Chen, Zongzheng Zhang, Yongliang Shi, Hao Zhao 0002
BMVC3
2024 ECT: Fine-grained edge detection with learned cause tokens
Shaocong Xu, Xiaoxue Chen, Yuhang Zheng 0004, Guyue Zhou, Yurong Chen 0001, Hongbin Zha, Hao Zhao 0002
Image Vis. Comput.2
2023 DPF: Learning Dense Prediction Fields with Weak Supervision
abstract
Nowadays, many visual scene understanding problems are addressed by dense prediction networks. But pixel-wise dense annotations are very expensive (e.g., for scene parsing) or impossible (e.g., for intrinsic image decomposition), motivating us to leverage cheap point-level weak supervision. However, existing pointly-supervised methods still use the same architecture designed for full supervision. In stark contrast to them, we propose a new paradigm that makes predictions for point coordinate queries, as inspired by the recent success of implicit representations, like distance or radiance fields. As such, the method is named as dense prediction fields (DPFs). DPFs generate expressive intermediate features for continuous sub-pixel locations, thus allowing outputs of an arbitrary resolution. DPFs are naturally compatible with point-level supervision. We showcase the effectiveness of DPFs using two substantially different tasks: high-level semantic parsing and low-level intrinsic image decomposition. In these two cases, supervision comes in the form of single-point semantic category and two-point relative reflectance, respectively. As benchmarked by three large-scale public datasets PASCALContext, ADE20K and IIW, DPFs set new state-of-the-art performance on all of them with significant margins. Code can be accessed at https://github.com/cxx226/DPF.
Xiaoxue Chen, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang
CVPR1
2023 INT2: Interactive Trajectory Prediction at Intersections
abstract
Motion forecasting is an important component in autonomous driving systems. One of the most challenging problems in motion forecasting is interactive trajectory prediction, whose goal is to jointly forecasts the future trajectories of interacting agents. To this end, we present a large-scale interactive trajectory prediction dataset named INT2 for INTeractive trajectory prediction at INTersections. INT2 includes 612,000 scenes, each lasting 1 minute, containing up to 10,200 hours of data. The agent trajectories are auto-labeled by a high-performance offline temporal detection and fusion algorithm, whose quality is further inspected by human judges. Vectorized semantic maps and traffic light information are also included in INT2. Additionally, the dataset poses an interesting domain mismatch challenge. For each intersection, we treat rush-hour and non-rush-hour segments as different domains. We benchmark the best open-sourced interactive trajectory prediction method on INT2 and Waymo Open Motion, under in-domain and cross-domain settings. The dataset, code and models are publicly available at https://github.com/AIRDISCOVER/INT2.
Zhijie Yan, Pengfei Li 0007, Zheng Fu, Shaocong Xu, Yongliang Shi, Xiaoxue Chen, Yuhang Zheng 0004, Yang Li 0178, Tianyu Liu 0008, Chuxuan Li, Nairui Luo, Zuoxu Wang, Yifeng Shi, Zhengxiao Han, Jirui Yuan, Jiangtao Gong, Guyue Zhou, Hang Zhao 0021, Hao Zhao 0002
ICCV6
2023 Understanding Embodied Reference with Touch-Line Transformer
Yang Li 0178, Xiaoxue Chen, Hao Zhao 0002, Jiangtao Gong, Guyue Zhou, Federico Rossano, Yixin Zhu 0001
ICLR2
2023 From Semi-supervised to Omni-supervised Room Layout Estimation Using Point Clouds
abstract
Room layout estimation is a long-existing robotic vision task that benefits both environment sensing and motion planning. However, layout estimation using point clouds (PCs) still suffers from data scarcity due to annotation difficulty. As such, we address the semi-supervised setting of this task based upon the idea of model exponential moving averaging. But adapting this scheme to the state-of-the-art (SOTA) solution for PC-based layout estimation is not straightforward. To this end, we define a quad set matching strategy and several consistency losses based upon metrics tailored for layout quads. Besides, we propose a new online pseudo-label harvesting algorithm that decomposes the distribution of a hybrid distance measure between quads and PC into two components. This technique does not need manual threshold selection and intuitively encourages quads to align with reliable layout points. Surprisingly, this framework also works for the fully-supervised setting, achieving a new SOTA on the ScanNet benchmark. Last but not least, we also push the semi-supervised setting to the realistic omni-supervised setting, demonstrating significantly promoted performance on a newly annotated ARKitScenes testing set. Our codes, data and models are made publicly available**Code: https://github.com/AIR-DISCOVER/Omni-PQ.
Huan-ang Gao, Beiwen Tian, Pengfei Li 0007, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Yurong Chen 0001, Hongbin Zha
ICRA4
2023 Fine-grained Pseudo Labels for Scene Text Recognition
abstract
Pseudo-Labeling based semi-supervised learning has shown promising advantages in Scene Text Recognition (STR). Most of them usually use a pre-trained model to generate sequence-level pseudo labels for text images and then re-train the model. Recently, conducting Pseudo-Labeling in a teacher-student framework (a student model is supervised by the pseudo labels from a teacher model) has become increasingly popular, which trains in an end-to-end manner and yields outstanding performance in semi-supervised learning. However, applying this framework directly to Pseudo-Labeling STR exhibits unstable convergence, as generating pseudo labels at the coarse-grained sequence-level leads to inefficient utilization of unlabelled data. Furthermore, the inherent domain shift between labeled and unlabeled data results in low quality of derived pseudo labels. To mitigate the above issues, we propose a novel Cross-domain Pseudo-Labeling (CPL) approach for scene text recognition, which makes better utilization of unlabeled data at the character-level and provides more accurate pseudo labels. Specifically, our proposed Pseudo-Labeled Curriculum Learning dynamically adjusts the thresholds for different character classes according to the model's learning status. Moreover, an Adaptive Distribution Regularizer is employed to bridge the domain gap and improve the quality of pseudo labels. Extensive experiments show that CPL boosts those representative STR models to achieve state-of-the-art results on six challenging STR benchmarks. Besides, it can be effectively generalized to handwritten text.
Xiaoxue Chen, Zuming Huang, Lele Xie, Jingdong Chen, Ming Yang 0007
ACM Multimedia2
2022 Cerberus Transformer: Joint Semantic, Affordance and Attribute Parsing
abstract
Multi-task indoor scene understanding is widely considered as an intriguing formulation, as the affinity of different tasks may lead to improved performance. In this paper, we tackle the new problem of Joint semantic, affordance and attribute parsing. However, successfully resolving it requires a model to capture long-range dependency, learn from weakly aligned data and properly balance sub-tasks during training. To this end, we propose an attention-based architecture named Cerberus and a tailored training framework. Our method effectively addresses aforementioned challenges and achieves state-of-the-art performance on all three tasks. Moreover, an in-depth analysis shows concept affinity consistent with human cognition, which inspires us to explore the possibility of weakly supervised learning. Surprisingly, Cerberus achieves strong results using only 0.1%–1% annotation. Visualizations further confirm that this success is credited to common attention maps across tasks. Code and models can be accessed at https://github.com/OPEN-AIR-SUN/Cerberus.
Xiaoxue Chen, Tianyu Liu 0008, Hao Zhao 0001, Guyue Zhou, Ya-Qin Zhang
CVPR1
2022 TOIST: Task Oriented Instance Segmentation Transformer with Noun-Pronoun Distillation
abstract
Current referring expression comprehension algorithms can effectively detect or segment objects indicated by nouns, but how to understand verb reference is still under-explored. As such, we study the challenging problem of task oriented detection, which aims to find objects that best afford an action indicated by verbs like sit comfortably on. Towards a finer localization that better serves downstream applications like robot interaction, we extend the problem into task oriented instance segmentation. A unique requirement of this task is to select preferred candidates among possible alternatives. Thus we resort to the transformer architecture which naturally models pair-wise query relationships with attention, leading to the TOIST method. In order to leverage pre-trained noun referring expression comprehension models and the fact that we can access privileged noun ground truth during training, a novel noun-pronoun distillation framework is proposed. Noun prototypes are generated in an unsupervised manner and contextual pronoun features are trained to select prototypes. As such, the network remains noun-agnostic during inference. We evaluate TOIST on the large-scale task oriented dataset COCO-Tasks and achieve +10.7% higher $\rm{mAP^{box}}$ than the best-reported results. The proposed noun-pronoun distillation can boost $\rm{mAP^{box}}$ and $\rm{mAP^{mask}}$ by +2.6% and +3.6%. Codes and models are publicly available.
Pengfei Li 0007, Beiwen Tian, Yongliang Shi, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou, Ya-Qin Zhang
NeurIPS4
2022 SNAKE: Shape-aware Neural 3D Keypoint Field
abstract
Detecting 3D keypoints from point clouds is important for shape reconstruction, while this work investigates the dual question: can shape reconstruction benefit 3D keypoint detection? Existing methods either seek salient features according to statistics of different orders or learn to predict keypoints that are invariant to transformation. Nevertheless, the idea of incorporating shape reconstruction into 3D keypoint detection is under-explored. We argue that this is restricted by former problem formulations. To this end, a novel unsupervised paradigm named SNAKE is proposed, which is short for shape-aware neural 3D keypoint field. Similar to recent coordinate-based radiance or distance field, our network takes 3D coordinates as inputs and predicts implicit shape indicators and keypoint saliency simultaneously, thus naturally entangling 3D keypoint detection and shape reconstruction. We achieve superior performance on various public benchmarks, including standalone object datasets ModelNet40, KeypointNet, SMPL meshes and scene-level datasets 3DMatch and Redwood. Intrinsic shape awareness brings several advantages as follows. (1) SNAKE generates 3D keypoints consistent with human semantic annotation, even without such supervision. (2) SNAKE outperforms counterparts in terms of repeatability, especially when the input point clouds are down-sampled. (3) the generated keypoints allow accurate geometric registration, notably in a zero-shot setting. Codes and models are available at https://github.com/zhongcl-thu/SNAKE.
Chengliang Zhong, Peixing You, Xiaoxue Chen, Hao Zhao 0002, Fuchun Sun 0001, Guyue Zhou, Xiaodong Mu, Chuang Gan 0001, Wenbing Huang 0001
NeurIPS3
2022 A novel compression framework of the dense point-cloud model for cultural heritage artifacts
Kang Li 0005, Jiaojiao Kou, Xiaoxue Chen, Linqi Hai, Guohua Geng, Shunli Zhang 0002
Multim. Tools Appl.4
2022 Distance-Aware Occlusion Detection With Focused Attention
abstract
For humans, understanding the relationships between objects using visual signals is intuitive. For artificial intelligence, however, this task remains challenging. Researchers have made significant progress studying semantic relationship detection, such as human-object interaction detection and visual relationship detection. We take the study of visual relationships a step further from semantic to geometric. In specific, we predict relative occlusion and relative distance relationships. However, detecting these relationships from a single image is challenging. Enforcing focused attention to task-specific regions plays a critical role in successfully detecting these relationships. In this work, (1) we propose a novel three-decoder architecture as the infrastructure for focused attention; 2) we use the generalized intersection box prediction task to effectively guide our model to focus on occlusion-specific regions; 3) our model achieves a new state-of-the-art performance on distance-aware relationship detection. Specifically, our model increases the distance F1-score from 33.8% to 38.6% and boosts the occlusion F1-score from 34.4% to 41.2%. Our code and data will be publicly available.
Yang Li 0178, Yucheng Tu, Xiaoxue Chen, Hao Zhao 0002, Guyue Zhou
IEEE Trans. Image Process.3
2020 Decoupled Attention Network for Text Recognition
abstract
Text recognition has attracted considerable research interests because of its various applications. The cutting-edge text recognition methods are based on attention mechanisms. However, most of attention methods usually suffer from serious alignment problem due to its recurrency alignment operation, where the alignment relies on historical decoding results. To remedy this issue, we propose a decoupled attention network (DAN), which decouples the alignment operation from using historical decoding results. DAN is an effective, flexible and robust end-to-end text recognizer, which consists of three components: 1) a feature encoder that extracts visual features from the input image; 2) a convolutional alignment module that performs the alignment operation based on visual features from the encoder; and 3) a decoupled text decoder that makes final prediction by jointly using the feature map and attention maps. Experimental results show that DAN achieves state-of-the-art performance on multiple text recognition tasks, including offline handwritten text recognition and regular/irregular scene text recognition. Codes will be released.1
Canjie Luo, Xiaoxue Chen, Yaqiang Wu, Qianying Wang 0002, Mingxiang Cai
AAAI5
2020 Adaptive embedding gate for attention-based scene text recognition
abstract
Scene text recognition has attracted particular research interest because it is a very challenging problem and has various applications. The most cutting-edge methods are attentional encoder-decoder frameworks that learn the alignment between the input image and output sequences. In particular, the decoder recurrently outputs predictions, using the prediction of the previous step as a guidance for every time step. In this study, we point out that the inappropriate use of previous predictions in existing attentional decoders restricts the recognition performance and brings instability. To handle this problem, we propose a novel module, namely adaptive embedding gate (AEG). The proposed AEG focuses on introducing high-order character language models to attentional decoders by controlling the information transmission between adjacent characters. AEG is a flexible module and can be easily integrated into the state-of-the-art attentional decoders for scene text recognition. We evaluate its effectiveness as well as robustness on a number of standard benchmarks, including the IIIT5K, SVT, SVT-P, CUTE80, and ICDAR datasets. Experimental results demonstrate that AEG can significantly boost recognition performance and bring better robustness.
Xiaoxue Chen, Canjie Luo
Neurocomputing1
2019 Data linkages between patient-powered research networks and health plans: a foundation for collaborative research
abstract
OBJECTIVE: Patient-powered research networks (PPRNs) are a valuable source of patient-generated information. Diagnosis code-based algorithms developed by PPRNs can be used to query health plans' claims data to identify patients for research opportunities. Our objective was to implement privacy-preserving record linkage processes between PPRN members' and health plan enrollees' data, compare linked and nonlinked members, and measure disease-specific confirmation rates for specific health conditions. MATERIALS AND METHODS: This descriptive study identified overlapping members from 4 PPRN registries and 14 health plans. Our methods for the anonymous linkage of overlapping members used secure Health Insurance Portability and Accountability Act-compliant, 1-way, cryptographic hash functions. Self-reported diagnoses by PPRN members were compared with claims-based computable phenotypes to calculate confirmation rates across varying durations of health plan coverage. RESULTS: Data for 21 616 PPRN members were hashed. Of these, 4487 (21%) members were linked, regardless of any expected overlap with the health plans. Linked members were more likely to be female and younger than nonlinked members were. Irrespective of duration of enrollment, the confirmation rates for the breast or ovarian cancer, rheumatoid or psoriatic arthritis or psoriasis, multiple sclerosis, or vasculitis PPRNs were 72%, 50%, 75%, and 67%, increasing to 91%, 67%, 93%, and 80%, respectively, for members with ≥5 years of continuous health plan enrollment. CONCLUSIONS: This study demonstrated that PPRN membership and health plan data can be successfully linked using privacy-preserving record linkage methodology, and used to confirm self-reported diagnosis. Identifying and confirming self-reported diagnosis of members can expedite patient selection for research opportunities, shorten study recruitment timelines, and optimize costs.
Abiy Agiro, Xiaoxue Chen, Biruk Eshete, Rebecca Sutphen, Elizabeth Bourquardez Clark, Cristina M. Burroughs, William Benjamin Nowell, Jeffrey R. Curtis, Sara Loud, Robert N. McBurney, Peter A. Merkel, Antoine G. Sreih, Kalen Young, Kevin Haynes
J. Am. Medical Informatics Assoc.2
2015 An audio-visual human attention analysis approach to abrupt change detection in videos
Minglong Song, Lixia Xue, Xiaoxue Chen
Signal Process.4