EDBT 2026 Demo / reviewers in the wild / expert
Xiaotian Qiao
dblp:166/2445
· DBLP profile ↗
17ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0002-5351-8335ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RSA-CR: Resisting Shilling Attacks in Citation Recommendation via Dumbbell Inductive LearningabstractCitation recommendation aims to provide researchers with the most relevant references for their manuscripts, helping them swiftly discover pertinent studies and bolster the reliability of their arguments. However, some individuals manipulate these recommendation systems by injecting false information, such as deliberately inflating the citation count of their own papers, to obtain favorable recommendations and ratings. This form of attack, commonly termed “shilling attack”, is not only highly concealed but also has an unimaginable impact on all scientific research. To address this problem, we theoretically reveal the impact of shilling attacks on citation recommendation and propose three feasible resistance strategies: historical collaborations, significant citations and content constraints. Based on these insights, we introduce RSA-CR, a robust and hybrid citation recommendation algorithm resistant to shilling attacks. The algorithm constructs a two-layer academic graph and uses random and content generation strategies to initialize author and paper embeddings. Confidence-guided inductive aggregations based on collaboration and citation relationships are then performed at the author and paper sides, where author aggregation results directly influences the paper aggregation strength. Finally, recommendations are made by measuring the distances between the fused paper embeddings. The entire learning process resembles a dumbbell, hence termed “dumbbell inductive learning”. Experiments on four academic datasets demonstrate that our method outperforms baselines in both effectiveness and robustness. Xiyue Gao, Zhuoqi Ma, Xiaotian Qiao, Hui Li 0005, Kunhua Zhang, Jiangtao Cui |
AAAI | 4 |
| 2025 | HDLayout: Hierarchical and Directional Layout Planning for Arbitrary Shaped Visual Text GenerationabstractVisual text generation, which aims to generate photo-realistic images with coherent and well-formed scene text being rendered, has attracted widespread attention. Although recent works have achieved promising performance, the limited flexibility and controllability hinder their practical applications. We observe that different from natural objects, visual text in real scenes often has an arbitrarily shaped structure with different granularities (i.e., character, word, or line). In this paper, we consider the modality gap between image and text, and propose a new separation and composition pipeline for flexible and controllable visual text generation from only text prompts. At the core of our framework is a novel Hierarchical and Directional Layout representation, i.e., HDLayout, which can model the sequential and multi-granularity nature of the visual text. Under this formulation, we are able to generate arbitrarily shaped visual text automatically. Extensive experiments demonstrate that our method outperforms several strong baselines in a variety of scenarios both qualitatively and quantitatively, yielding state-of-the-art performances on arbitrarily shaped visual text generation. Tonghui Feng, Chunsheng Yan, Qianru Wang, Jiangtao Cui, Xiaotian Qiao |
AAAI | 5 |
| 2025 | CollageNoter: Real-Time and Adaptive Collage Layout Design for Screenshot-Based E-Note-TakingabstractTo enhance the processing of complex multi-modal documents (e.g. e-books, long web pages, etc.), it is an efficient way for users to take digital screenshots of key parts and reorganize them into a new collage E-Note. Existing methods for assisting collage layout design primarily employ a semantic relevance-first strategy, with arranging related contents together. Though capable, it can not ensure the visual readability of screenshots and may conflict with human natural reading patterns. In this paper, we introduce CollageNoter for real-time collage layout design that adapts to various devices (e.g. laptop, tablet, phone, etc.), offering users with visually and cognitively well-organized screenshot-based E-Notes. Specifically, we construct a novel two-stage pipeline for collage design, including 1) readability-first layout generation and 2) cognitive-driven layout adjustment. In addition, to achieve real-time response and adaptive model training, we propose a cascade transformer-based layout generator named CollageFormer and a size-aware collage layout builder for automatic dataset construction. Extensive experimental results have confirmed the effectiveness of our CollageNoter. Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Xiaotian Qiao, Ying Zhang 0047, Zhiwen Yu 0001 |
AAAI | 4 |
| 2025 | Less is More: Efficient Image Vectorization with Adaptive ParameterizationabstractImage vectorization aims to convert raster images to vector ones, allowing for easy scaling and editing. Existing works mainly rely on preset parameters (i.e., a fixed number of paths and control points), ignoring the complexity of the image and posing significant challenges to practical applications. We demonstrate that such an assumption is often incorrect, as the preset paths or control points may be neither essential nor enough to achieve accurate and editable vectorization results. Based on this key insight, in this paper, we propose AdaVec, an efficient image vectorization method with adaptive parametrization, where the paths and control points can be adjusted dynamically based on the complexity of the input raster image. In particular, we first decompose the input raster image into a set of pure-colored layers that are aligned with human perception. For each layer with varying shape complexity, we propose a novel allocation mechanism to adaptively adjust the control point distribution. We further adopt a differentiable rendering process to compose and optimize the shape and color parameters of each layer iteratively. Extensive experiments demonstrate that AdaVec outperforms the baselines qualitatively and quantitatively, in terms of computational efficiency, vectorization accuracy, and editing flexibility. Kaibo Zhao 0001, Liang Bao, Xu Su, Xiaotian Qiao |
CVPR | 6 |
| 2025 | Under the Shadow: Exploiting Opacity Variation for Fine-grained Shadow DetectionabstractShadow characteristics are of great importance for scene understanding.
Existing works mainly consider shadow regions as binary masks, often leading to imprecise detection results and suboptimal performance for scene understanding.
We demonstrate that such an assumption oversimplifies light-object interactions in the scene, as the scene details under either hard or soft shadows remain visible to a certain degree.
Based on this insight, we aim to reformulate the shadow detection paradigm from the opacity perspective, and introduce a new fine-grained shadow detection method.
In particular, given an input image, we first propose a shadow opacity augmentation module to generate realistic images with varied shadow opacities.
We then introduce a shadow feature separation module to learn the shadow position and opacity representations separately, followed by an opacity mask prediction module that fuses these representations and predicts fine-grained shadow detection results.
In addition, we construct a new dataset with opacity-annotated shadow masks across varied scenarios.
Extensive experiments demonstrate that our method outperforms the baselines qualitatively and quantitatively, enhancing a wide range of applications, including shadow removal, shadow editing, and 3D reconstruction. Xiaotian Qiao, Xianglong Yang, Ruijie Dong, Xiaofang Xia, Jiangtao Cui |
NeurIPS | 1 |
| 2025 | Deep multi-negative supervised hashing for large-scale image retrieval
Yingfan Liu, Xiaotian Qiao, Zhaoqing Liu, Xiaofang Xia, Yinlong Zhang, Jiangtao Cui |
Expert Syst. Appl. | 2 |
| 2025 | Design Graph Guided Element Importance-Aware Layout Generation With Multimodality Cascade TransformerabstractGraphic designs are pervasive in our daily lives and widely used to communicate information hierarchically to humans. To achieve this, the layout plays an essential role in guiding readers to understand the importance of different elements and comprehend the content. To deal with the rapidly increasing demands of graphic designs, recent studies attempt to automatically generate layouts based on category information and spatial relations, often resulting in layouts with poor communication quality. In this article, we make the first attempt to explore element importance-aware layout generation under the guidance of a novel design graph, which attracts readers’ attention to a layout by formulating aesthetic relations implicitly involved in graphic designs between element pairs. The core of our approach is a learning-based framework with a new multimodality cascade transformer (MCT) in a coarse-to-fine manner. A hierarchical multimodality fusion (HMF) mechanism and two new losses are introduced to guide the training process progressively. We further collect a new fine-grained advertisement poster layout dataset containing more than 30 K layouts labeled with 91 element labels. Both qualitative and quantitative experiments demonstrate the effectiveness of our approach against existing works. We also conduct user studies and cognitive experiments to evaluate the direct adaptability and attractiveness of generated layouts. Qiuyun Zhang, Bin Guo 0001, Lina Yao 0001, Xiaotian Qiao, Hao Wang 0182, Ying Zhang 0047, Zhiwen Yu 0001 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2023 | Design Order Guided Visual Note Layout OptimizationabstractWith the goal of making contents easy to understand, memorize and share, a clear and easy-to-follow layout is important for visual notes. Unfortunately, since visual notes are often taken by the designers in real time while watching a video or listening to a presentation, the contents are usually not carefully structured, resulting in layouts that may be difficult for others to follow. In this article, we address this problem by proposing a novel approach to automatically optimize the layouts of visual notes. Our approach predicts the design order of a visual note and then warps the contents along the predicted design order such that the visual note can be easier to follow and understand. At the core of our approach is a learning-based framework to reason about the element-wise design orders of visual notes. In particular, we first propose a hierarchical LSTM-based architecture to predict a grid-based design order of the visual note, based on the graphical and textual information. We then derive the element-wise order from the grid-based prediction. Such an idea allows our network to be weakly-supervised, i.e., making it possible to predict dense grid-based orders from visual notes with only coarse annotations. We evaluate the effectiveness of our approach on visual notes with diverse content densities and layouts. The results show that our network can predict plausible design orders for various types of visual notes and our approach can effectively optimize their layouts in order for them to be easier to follow. Xiaotian Qiao, Ying Cao 0001, Rynson W. H. Lau |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | Learning Object Context for Novel-view Scene Layout GenerationabstractNovel-view prediction of a scene has many applications. Existing works mainly focus on generating novel-view images via pixel-wise prediction in the image space, often resulting in severe ghosting and blurry artifacts. In this paper, we make the first attempt to explore novel-view prediction in the layout space, and introduce the new problem of novel-view scene layout generation. Given a single scene layout and the camera transformation as inputs, our goal is to generate a plausible scene layout for a specified viewpoint. Such a problem is challenging as it involves accurate understanding of the 3D geometry and semantics of the scene from as little as a single 2D scene layout. To tackle this challenging problem, we propose a deep model to capture contextualized object representation by explicitly modeling the object context transformation in the scene. The contextualized object representation is essential in generating geometrically and semantically consistent scene layouts of different views. Experiments show that our model outperforms several strong baselines on many indoor and outdoor scenes, both qualitatively and quantitatively. We also show that our model enables a wide range of applications, including novel-view image synthesis, novel-view image editing, and amodal object estimation. Xiaotian Qiao, Gerhard P. Hancke 0002, Rynson W. H. Lau |
CVPR | 1 |
| 2022 | Instance-Aware Scene Layout Forecasting
Xiaotian Qiao, Quanlong Zheng, Ying Cao 0001, Rynson W. H. Lau |
Int. J. Comput. Vis. | 1 |
| 2022 | Correction to: Instance-Aware Scene Layout Forecasting
Xiaotian Qiao, Quanlong Zheng, Ying Cao 0001, Rynson W. H. Lau |
Int. J. Comput. Vis. | 1 |
| 2022 | Object-Level Scene Context PredictionabstractContextual information plays an important role in solving various image and scene understanding tasks. Prior works have focused on the extraction of contextual information from an image and use it to infer the properties of some object(s) in the image or understand the scene behind the image, e.g., context-based object detection, recognition and semantic segmentation. In this paper, we consider an inverse problem, i.e., how to hallucinate the missing contextual information from the properties of standalone objects. We refer to it as object-level scene context prediction. This problem is difficult, as it requires extensive knowledge of the complex and diverse relationships among objects in the scene. We propose a deep neural network, which takes as input the properties (i.e., category, shape, and position) of a few standalone objects to predict an object-level scene layout that compactly encodes the semantics and structure of the scene context where the given objects are. Quantitative experiments and user studies demonstrate that our model can generate more plausible scene contexts than the baselines. Our model also enables the synthesis of realistic scene images from partial scene layouts. Finally, we validate that our model internally learns useful features for scene recognition and fake scene detection. Xiaotian Qiao, Quanlong Zheng, Ying Cao 0001, Rynson W. H. Lau |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Light Source Guided Single-Image Flare Removal from Unpaired DataabstractCausally-taken images often suffer from flare artifacts, due to the unintended reflections and scattering of light inside the camera. However, as flares may appear in a variety of shapes, positions, and colors, detecting and removing them entirely from an image is very challenging. Existing methods rely on predefined intensity and geometry priors of flares, and may fail to distinguish the difference between light sources and flare artifacts. We observe that the conditions of the light source in the image play an important role in the resulting flares. In this paper, we present a deep framework with light source aware guidance for single-image flare removal (SIFR). In particular, we first detect the light source regions and the flare regions separately, and then remove the flare artifacts based on the light source aware guidance. By learning the underlying relationships between the two types of regions, our approach can remove different kinds of flares from the image. In addition, instead of using paired training data which are difficult to collect, we propose the first unpaired flare removal dataset and new cycle-consistency constraints to obtain more diverse examples and avoid manual annotations. Extensive experiments demonstrate that our method outperforms the baselines qualitatively and quantitatively. We also show that our model can be applied to flare effect manipulation (e.g., adding or changing image flares). Xiaotian Qiao, Gerhard P. Hancke 0002, Rynson W. H. Lau |
ICCV | 1 |
| 2019 | Tell Me Where I Am: Object-Level Scene Context PredictionabstractContextual information has been shown to be effective in helping solve various image understanding tasks. Previous works have focused on the extraction of contextual information from an image and use it to infer the properties of some object(s) in the image. In this paper, we consider an inverse problem of how to hallucinate missing contextual information from the properties of a few standalone objects. We refer to it as scene context prediction. This problem is difficult as it requires an extensive knowledge of complex and diverse relationships among different objects in natural scenes. We propose a convolutional neural network, which takes as input the properties (i.e., category, shape, and position) of a few standalone objects to predict an object-level scene layout that compactly encodes the semantics and structure of the scene context where the given objects are. Our quantitative experiments and user studies show that our model can generate more plausible scene context than the baseline approach. We demonstrate that our model allows for the synthesis of realistic scene images from just partial scene layouts and internally learns useful features for scene recognition. Xiaotian Qiao, Quanlong Zheng, Ying Cao 0001, Rynson W. H. Lau |
CVPR | 1 |
| 2019 | Distraction-Aware Shadow DetectionabstractShadow detection is an important and challenging task for scene understanding. Despite promising results from recent deep learning based methods. Existing works still struggle with ambiguous cases where the visual appearances of shadow and non-shadow regions are similar (referred to as distraction in our context). In this paper, we propose a Distraction-aware Shadow Detection Network (DSDNet) by explicitly learning and integrating the semantics of visual distraction regions in an end-to-end framework. At the core of our framework is a novel standalone, differentiable Distraction-aware Shadow (DS) module, which allows us to learn distraction-aware, discriminative features for robust shadow detection, by explicitly predicting false positives and false negatives. We conduct extensive experiments on three public shadow detection datasets, SBU, UCF and ISTD, to evaluate our method. Experimental results demonstrate that our model can boost shadow detection performance, by effectively suppressing the detection of false positives and false negatives, achieving state-of-the-art results. Quanlong Zheng, Xiaotian Qiao, Ying Cao 0001, Rynson W. H. Lau |
CVPR | 2 |
| 2019 | Content-aware generative modeling of graphic design layoutsabstractLayout is fundamental to graphic designs. For visual attractiveness and efficient communication of messages and ideas, graphic design layouts often have great variation, driven by the contents to be presented. In this paper, we study the problem of content-aware graphic design layout generation. We propose a deep generative model for graphic design layouts that is able to synthesize layout designs based on the visual and textual semantics of user inputs. Unlike previous approaches that are oblivious to the input contents and rely on heuristic criteria, our model captures the effect of visual and textual contents on layouts, and implicitly learns complex layout structure variations from data without the use of any heuristic rules. To train our model, we build a large-scale magazine layout dataset with fine-grained layout annotations and keyword labeling. Experimental results show that our model can synthesize high-quality layouts based on the visual semantics of input images and keyword-based summary of input text. We also demonstrate that our model internally learns powerful features that capture the subtle interaction between contents and layouts, which are useful for layout-aware design retrieval. Xinru Zheng, Xiaotian Qiao, Ying Cao 0001, Rynson W. H. Lau |
ACM Trans. Graph. | 2 |
| 2015 | Undersampled Dynamic MRI Reconstruction by Double Sparse Spatiotemporal Dictionary
Juerong Wu, Xiaotian Qiao, Lianghao Wang, Ming Zhang 0001 |
ICIG (3) | 3 |