EDBT 2026 Demo / reviewers in the wild / expert
Yuan Xiong
dblp:38/2840
· DBLP profile ↗
20ranked-venue papers
4as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly DetectionabstractIndustrial anomaly detection is a critical component of modern manufacturing, yet the scarcity of defective samples restricts traditional detection methods to scenario-specific applications. Although Vision-Language Models (VLMs) demonstrate significant advantages in generalization capabilities, their performance in industrial anomaly detection remains limited. To address this challenge, we propose IAD-R1, a universal post-training framework applicable to VLMs of different architectures and parameter scales, which substantially enhances their anomaly detection capabilities. IAD-R1 employs a two-stage training strategy: the Perception Activation Supervised Fine-Tuning (PA-SFT) stage utilizes a meticulously constructed high-quality Chain-of-Thought dataset (Expert-AD) for training, enhancing anomaly perception capabilities and establishing reasoning-to-answer correlations; the Structured Control Group Relative Policy Optimization (SC-GRPO) stage employs carefully designed reward functions to achieve a capability leap from "Anomaly Perception" to "Anomaly Interpretation". Experimental results demonstrate that IAD-R1 achieves significant improvements across 7 VLMs, the largest improvement was on the DAGM dataset, with average accuracy 43.3% higher than the 0.5B baseline. Notably, the 0.5B parameter model trained with IAD-R1 surpasses commercial models including GPT-4.1 and Claude-Sonnet-4 in zero-shot settings, demonstrating the effectiveness and superiority of IAD-R1. Yunkang Cao, Chengliang Liu 0003, Yuan Xiong, Xinghui Dong, Chao Huang 0008 |
AAAI | 4 |
| 2026 | Detecting Fake News in Short Videos Through Multi-View AggregationabstractThe increasing prominence of short video platforms has positioned them as a primary channel for public awareness of current events, while also facilitating the widespread dissemination of fake news, thus highlighting the critical need for automated detection technologies. In contrast to fake news confined to text and images, short video news encompasses multiple modalities and extensive information, presenting heightened challenges. Most existing research emphasizes the analysis of news content or user comments alone, while overlooking the crucial role of publishers, leading to poor model performance when handling fake news lacking obvious false signals. Therefore, we propose a Publisher Profiling Module to identify new false signals. To enable a more comprehensive detection of misinformation, we design a Multi-View Aggregation (MVA) model, simultaneously evaluating news from three distinct perspectives: sentiment analysis, content understanding, and publisher profiling. Late fusion is applied at the decision level to leverage the complementary strengths of these perspectives, addressing the limitations of single-view methods. Our experiments conducted on the FakeSV and FVC datasets demonstrate the superior performance of the proposed method. Yuan Xiong, Chengliang Liu 0003, Jie Wen 0001, Chao Huang 0008 |
AAAI | 2 |
| 2026 | Response Attack: Exploiting Contextual Priming to Jailbreak Large Language ModelsabstractContextual priming, where earlier stimuli covertly bias later judgments, offers an unexplored attack surface for large language models (LLMs). We uncover a contextual priming vulnerability in which the previous response in the dialogue can steer its subsequent behavior toward policy-violating content. While existing jailbreak attacks largely rely on single-turn or multi-turn prompt manipulations, or inject static in-context examples, these methods suffer from limited effectiveness, inefficiency, or semantic drift. We introduce Response Attack (RA), a novel framework that strategically leverages intermediate, mildly harmful responses as contextual primers within a dialogue. By reformulating harmful queries and injecting these intermediate responses before issuing a targeted trigger prompt, RA exploits a previously overlooked vulnerability in LLMs. Extensive experiments across eight state-of-the-art LLMs show that RA consistently achieves significantly higher attack success rates than nine leading jailbreak baselines. Our results demonstrate that the success of RA is directly attributable to the strategic use of intermediate responses, which induce models to generate more explicit and relevant harmful content while maintaining stealth, efficiency, and fidelity to the original query. Ziqi Miao, Yuan Xiong |
AAAI | 3 |
| 2026 | Leveraging Multi-Text Joint Prompts in SAM for Robust Medical Image SegmentationabstractThe Segment Anything Model (SAM) has attracted considerable attention due to its impressive performance and demonstrates potential in medical image segmentation. Compared to SAM's native point andbounding box prompts, text prompts offer a simpler and more efficient alternative in the medical field, yet this approach remains relatively underexplored. In this paper, we propose a SAM-based framework that integrates a pre-trained vision-language model to generate referring prompts, with SAM handling the segmentation task. The outputs from multimodal models such as CLIP serve as input to SAM's prompt encoder. A critical challenge stems from the inherent complexity of medical text descriptions: they typically encompass anatomical characteristics, imaging modalities, and diagnostic priorities, resulting in information redundancy and semantic ambiguity. To address this, we propose a text decomposition-recomposition strategy. First, clinical narratives are parsed into atomic semantic units (appearance, location, pathology, and so on). These elements are then recombined into optimized text expressions. We employ a cross-attention module among multiple texts to interact with the joint features, ensuring that the model focuses on features corresponding to effective descriptions. To validate the effectiveness of our method, we conducted experiments on several datasets. Compared to the native SAM based on geometric prompts, our model shows improved performance and usability. Xu Zhang 0044, Huangxuan Zhao, Lefei Zhang, Yuan Xiong |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Enhancing open-vocabulary scene understanding via push-pull alignment in gaussian splatting
Shengjia Liang, Yuan Xiong, Qichuan Geng, Zhong Zhou |
Vis. Comput. | 3 |
| 2026 | CleanSplat: curriculum structural Gaussian splatting for spot-free novel view synthesisabstract3D Gaussian splatting (3DGS) has gained significant attention for its real-time, photorealistic rendering in novel view synthesis. However, its performance degrades severely when applied to real-world scenes with transients that break cross-view consistency. Existing methods typically attempt to identify and mask out these transients, but inherent masking errors often leave behind incoherent floating Gaussians, resulting in spotted artifacts that degrade the final scene quality. To address these limitations, we propose CleanSplat, a curriculum structural Gaussian splatting method that purifies the scene Gaussians through curriculum 3DGS optimization for transient removal and structural pruning of spotted artifacts. We introduce a curriculum-guided masking paradigm that generates coarse-to-fine transient masks from multi-scale features. The progressive optimization is driven by modulating the masking supervision based on current training state. To clear the spots, we propose a structure-aware handling strategy that employs a superpoint graph (SPG) partitioning of the Gaussians to perform principled identification and hierarchical pruning. This allows for the filtering of both intra-superpoint outliers and entire spurious superpoints based on local 3D coherence instead of only simple photometric consistency. By integrating curriculum 3DGS optimization and structural pruning, our method effectively separates the transients and purifies the static scene Gaussians. Extensive experiments on challenging datasets demonstrate that CleanSplat significantly outperforms state-of-the-art methods, delivering more detailed and cleaner novel view synthesis. Shengjia Liang, Qichuan Geng, Yuan Xiong, Zhong Zhou |
Virtual Real. Intell. Hardw. | 4 |
| 2025 | RobustSplat: Decoupling Densification and Dynamics for Transient-Free 3DGS
Chuanyu Fu, Kunbin Yao, Guanying Chen, Yuan Xiong, Chuan Huang 0001, Shuguang Cui, Xiaochun Cao |
ICCV | 5 |
| 2025 | A Differentiable Optimization Framework for Camera Pose and Depth Supervision in 3D Gaussian Splatting
Yuan Xiong, Changliang Li, Yiqi Wu |
ICIG (3) | 2 |
| 2025 | LangLoc: Language-Driven Localization via Formatted Spatial Description GenerationabstractExisting localization methods commonly employ vision to perceive scene and achieve localization in GNSS-denied areas, yet they often struggle in environments with complex lighting conditions, dynamic objects or privacy-preserving areas. Humans possess the ability to describe various scenes using natural language, effectively inferring their location by leveraging the rich semantic information in these descriptions. Harnessing language presents a potential solution for robust localization. Thus, this study introduces a new task, Language-driven Localization, and proposes a novel localization framework, LangLoc, which determines the user's position and orientation through textual descriptions. Given the diversity of natural language descriptions, we first design a Spatial Description Generator (SDG), foundational to LangLoc, which extracts and combines the position and attribute information of objects within a scene to generate uniformly formatted textual descriptions. SDG eliminates the ambiguity of language, detailing the spatial layout and object relations of the scene, providing a reliable basis for localization. With generated descriptions, LangLoc effortlessly achieves language-only localization using text encoder and pose regressor. Furthermore, LangLoc can add one image to text input, achieving mutual optimization and feature adaptive fusion across modalities through two modality-specific encoders, cross-modal fusion, and multimodal joint learning strategies. This enhances the framework's capability to handle complex scenes, achieving more accurate localization. Extensive experiments on the Oxford RobotCar, 4-Seasons, and Virtual Gallery datasets demonstrate LangLoc's effectiveness in both language-only and visual-language localization across various outdoor and indoor scenarios. Notably, LangLoc achieves noticeable performance gains when using both text and image inputs in challenging conditions such as overexposure, low lighting, and occlusions, showcasing its superior robustness. Changhao Chen, Kaige Li, Yuan Xiong, Xiaochun Cao, Zhong Zhou |
IEEE Trans. Image Process. | 4 |
| 2025 | FDCPNet:feature discrimination and context propagation network for 3D shape representationabstractThree-dimensional (3D) shape representation using mesh data is essential in various applications, such as virtual reality and simulation technologies. Current methods for extracting features from mesh edges or faces struggle with complex 3D models because edge-based approaches miss global contexts and face-based methods overlook variations in adjacent areas, which affects the overall precision. To address these issues, we propose the Feature Discrimination and Context Propagation Network (FDCPNet), which is a novel approach that synergistically integrates local and global features in mesh datasets. FDCPNet is composed of two modules: (1) the Feature Discrimination Module, which employs an attention mechanism to enhance the identification of key local features, and (2) the Context Propagation Module, which enriches key local features by integrating global contextual information, thereby facilitating a more detailed and comprehensive representation of crucial areas within the mesh model. Experiments on popular datasets validated the effectiveness of FDCPNet, showing an improvement in the classification accuracy over the baseline MeshNet. Furthermore, even with reduced mesh face numbers and limited training data, FDCPNet achieved promising results, demonstrating its robustness in scenarios of variable complexity. Yuan Xiong, Zhong Zhou |
Virtual Real. Intell. Hardw. | 2 |
| 2024 | SADNet: Generating immersive virtual reality avatars by real-time monocular pose estimationabstractSummary Generating immersive virtual reality avatars is a challenging task in VR/AR applications, which maps physical human body poses to avatars in virtual scenes for an immersive user experience. However, most existing work is time‐consuming and limited by datasets, which does not satisfy immersive and real‐time requirements of VR systems. In this paper, we aim to generate 3D real‐time virtual reality avatars based on a monocular camera to solve these problems. Specifically, we first design a self‐attention distillation network (SADNet) for effective human pose estimation, which is guided by a pre‐trained teacher. Secondly, we propose a lightweight pose mapping method for human avatars that utilizes the camera model to map 2D poses to 3D avatar keypoints, generating real‐time human avatars with pose consistency. Finally, we integrate our framework into a VR system, displaying generated 3D pose‐driven avatars on Helmet‐Mounted Display devices for an immersive user experience. We evaluate SADNet on two publicly available datasets. Experimental results show that SADNet achieves a state‐of‐the‐art trade‐off between speed and accuracy. In addition, we conducted a user experience study on the performance and immersion of virtual reality avatars. Results show that pose‐driven 3D human avatars generated by our method are smooth and attractive. Ling Jiang 0001, Yuan Xiong, Qianqian Wang 0011, Wei Wu 0008, Zhong Zhou |
Comput. Animat. Virtual Worlds | 2 |
| 2024 | DreamWalk: Dynamic remapping and multiperspectivity for large-scale redirected walkingabstractSummary Redirected walking (RDW) provides an immersive user experience in virtual reality applications. In RDW, the size of the physical play area is limited, which makes it challenging to design the virtual path in a larger virtual space. Mainstream RDW approaches rigidly manipulate gains to guide the user to follow predetermined rules. However, these methods may cause simulator sickness, boundary collision, and reset. Static mapping approaches warp the virtual path through expensive vertex replacement in the stage of model pre‐processing. They are restricted to narrow spaces with non‐looping pathways, partition walls, and planar surfaces. These methods fail to provide a smooth walking experience for large‐scale open scenes. To tackle these problems, we propose a novel approach that dynamically redirects the user to walk in a non‐linear virtual space. More specifically, we propose a Bezier‐curve‐based mapping algorithm to warp the virtual space dynamically and apply multiperspective fusion for visualization augmentation. We conduct comparable experiments to show its superiority over state‐of‐the‐art large‐scale redirected walking approaches on our self‐collected photogrammetry dataset. Yuan Xiong, Tianjing Li, Zhong Zhou |
Comput. Animat. Virtual Worlds | 1 |
| 2024 | QR-CLIP: Introducing Explicit Knowledge for Location and Time ReasoningabstractThis article focuses on reasoning about the location and time behind images. Given that pre-trained vision-language models (VLMs) exhibit excellent image and text understanding capabilities, most existing methods leverage them to match visual cues with location and time-related descriptions. However, these methods cannot look beyond the actual content of an image, failing to produce satisfactory reasoning results, as such reasoning requires connecting visual details with rich external cues (e.g., relevant event contexts). To this end, we propose a novel reasoning method, QR-CLIP , that aims at enhancing the model’s ability to reason about location and time through interaction with external explicit knowledge such as Wikipedia. Specifically, QR-CLIP consists of two modules: (1) The Quantity module abstracts the image into multiple distinct representations and uses them to search and gather external knowledge from different perspectives that are beneficial to model reasoning. (2) The Relevance module filters the visual features and the searched explicit knowledge and dynamically integrates them to form a comprehensive reasoning result. Extensive experiments demonstrate the effectiveness and generalizability of QR-CLIP . On the WikiTiLo dataset, QR-CLIP boosts the accuracy of location (country) and time reasoning by 7.03% and 2.22%, respectively, over previous SOTA methods. On the more challenging TARA dataset, it improves the accuracy for location and time reasoning by 3.05% and 2.45%, respectively. The source code is at https://github.com/Shi-Wm/QR-CLIP . Dehong Gao, Yuan Xiong, Zhong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | VirtualLoc: Large-scale Visual Localization Using Virtual ImagesabstractRobust and accurate camera pose estimation is fundamental in computer vision. Learning-based regression approaches acquire six-degree-of-freedom camera parameters accurately from visual cues of an input image. However, most are trained on street-view and landmark datasets. These approaches can hardly be generalized to overlooking use cases, such as the calibration of the surveillance camera and unmanned aerial vehicle. Besides, reference images captured from the real world are rare and expensive, and their diversity is not guaranteed. In this article, we address the problem of using alternative virtual images for visual localization training. This work has the following principle contributions: First, we present a new challenging localization dataset containing six reconstructed large-scale three-dimensional scenes, 10,594 calibrated photographs with condition changes, and 300k virtual images with pixelwise labeled depth, relative surface normal, and semantic segmentation. Second, we present a flexible multi-feature fusion network trained on virtual image datasets for robust image retrieval. Third, we propose an end-to-end confidence map prediction network for feature filtering and pose estimation. We demonstrate that large-scale rendered virtual images are beneficial to visual localization. Using virtual images can solve the diversity problem of real images and leverage labeled multi-feature data for deep learning. Experimental results show that our method achieves remarkable performance surpassing state-of-the-art approaches. To foster research on improvement for visual localization using synthetic images, we release our benchmark at https://github.com/YuanXiong/contributions . Yuan Xiong, Zhong Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Incorporating Multiple Features to Predict Bug Fixing Time with Neural NetworksabstractDebugging is a well-known time-consuming task, and knowing how long it would take to resolve bugs is of great importance for allocating the limited resources in a software development team. However, it is challenging to predict bug fixing time since fixing bugs is subject to a plethora of uncertain factors such as types of bugs, program complexity and developers' abilities. Existing work mainly focuses on developers' activities in a bug lifecycle and ignores other important factors. In light of the limitations of existing work, we propose a novel approach to predicting the bug fixing time by incorporating a comprehensive set of relevant features. Specifically, we consider four types of features including developers' activities, developers' sentiments, semantics of bugs, and efforts caused by understanding and analyzing source code, and design particular neural networks to take advantage of these features and make them work efficiently. Experimental results on four real application datasets demonstrate that on the one hand, our approach outperforms the state-of-the-art by over 5% in accuracy and 7.3% in F1-score on average; on the other hand, each type of the features considered in our approach plays an important part in the prediction. Wei Yuan 0011, Yuan Xiong, Hailong Sun 0001, Xudong Liu 0001 |
ICSME | 2 |
| 2021 | Long-Term Visual Localization with Semantic Enhanced Global RetrievalabstractVisual localization under varying conditions such as changes in illumination, season and weather is a fundamental task for applications such as autonomous navigation. In this paper, we present a novel method of using semantic information for global image retrieval. By exploiting the distribution of different classes in a semantic scene, the discriminative features of the scene’s structure layout is embedded into a normalized vector that can be used for retrieval, i.e. semantic retrieval. Color image retrieval is based on low-level visual features extracted by algorithms or Convolutional Neural Networks (CNNs), while semantic retrieval is based on high-level semantic features which are robust in scene appearance variations. By combining semantic retrieval with color image retrieval in the global retrieval step, we show that these two methods can complement with each other and significantly improve the localization performance. Experiments on the challenging CMU Seasons dataset show that our method is robust across large variations of appearance and achieves state-of-the-art localization performance. Hongrui Chen, Yuan Xiong, Zhong Zhou |
MSN | 2 |
| 2021 | DSNet: Deep Shadow Network for Illumination EstimationabstractIllumination consistency has applications to modeling and rendering in virtual reality. In 3D reconstruction and Mixed Reality(MR) fusion, the appearance of a large-scale outdoor scene may change in response to lighting and seasons, for example. Since 3D reconstruction from scratch is costly, it is helpful to be able to update existing models with recently captured photographs. However, the illumination conditions of the captured photograph can be arbitrary, making it challenging to fit to the existing model. To tackle this problem, this paper proposes a novel approach that can precisely estimate the illumination of the input image. Our Deep Shadow Network (DSNet) collaboratively utilizes illumination-based data augmentation for sun position estimation, along with a dataset of illumination-based augmented renderings. Our run-time rendering and optimization strategy is also discussed. We show that accurate simulation of illumination can improve the performance of visual applications including place recognition and long-term localization. Experimental results validate the effectiveness of the proposed approach, and show its superiority over the state-of-the-art. Yuan Xiong, Hongrui Chen, Zhe Zhu, Zhong Zhou |
VR | 1 |
| 2020 | Background Noise Filtering and Distribution Dividing for Crowd CountingabstractCrowd counting is a challenging problem due to the diverse crowd distribution and background interference. In this paper, we propose a new approach for head size estimation to reduce the impact of different crowd scale and background noise. Different from just using local information of distance between human heads, the global information of the people distribution in the whole image is also under consideration. We obey the order of far- to near-region (small to large) to spread head size, and ensure that the propagation is uninterrupted by inserting dummy head points. The estimated head size is further exploited, such as dividing the crowd into parts of different densities and generating a high-fidelity head mask. On the other hand, we design three different head mask usage mechanisms and the corresponding head masks to analyze where and which mask could lead to better background filtering1. Based on the learned masks, two competitive models are proposed which can perform robust crowd estimation against background noise and diverse crowd scale. We evaluate the proposed method on three public crowd counting datasets of ShanghaiTech [2], UCFQNRF [3] and UCFCC_50 [4]. Experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art crowd counting approaches. Hong Mo, Wenqi Ren, Yuan Xiong, Xiaoqi Pan, Zhong Zhou, Xiaochun Cao, Wei Wu 0008 |
IEEE Trans. Image Process. | 3 |
| 2006 | A novel white blood cell segmentation scheme based on feature space clustering
Kan Jiang, Qing-Min Liao, Yuan Xiong |
Soft Comput. | 3 |
| 2004 | A learning-based tracking for diving motionsabstractA learning-based tracking algorithm for diving motions is presented in this paper. In this algorithm, a complex diving motion is considered as the combination of several simple sub-motions. The contour of the athlete in each sub-motion is represented by B-spline snake, which can be fitted to the real body contour by a recursive curve-fitting algorithm. By learning from the videos in a training set, the initial contour templates for each sub-motion are set up and each possible frame where a new sub-motion begins is found out, which allows the possibility of whole motion tracking. Experiments demonstrate that the proposed algorithm is robust and efficient in diving motions tracking. Yuan Xiong |
ICIG | 1 |