EDBT 2026 Demo / reviewers in the wild / expert
Xiaochao Qu
dblp:23/9252
· DBLP profile ↗
18ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-3181-3704ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Draw Like an Artist: Complex Scene Generation With Diffusion Model via Composition, Painting, and RetouchingabstractRecent advances in text-to-image diffusion models have demonstrated impressive capabilities in image quality. However, complex scene generation remains relatively unexplored, and even the definition of ‘complex scene’ itself remains unclear. In this paper, we address this gap by providing a precise definition of complex scenes and introducing a set of Complex Decomposition Criteria (CDC) based on this definition. Inspired by the artist’s painting process, we propose a training-free diffusion framework called Complex Diffusion (CxD), which divides the process into three stages: composition, painting and retouching. Our method leverages the powerful chain-of-thought capabilities of large language models (LLMs) to decompose complex prompts based on CDC and to manage composition and layout. We then develop an attention modulation method that guides simple prompts to specific regions to complete the complex scene painting. Finally, we inject the detailed output of the LLM into a retouching model to enhance the image details, thus implementing the retouching stage. Extensive experiments demonstrate that our method outperforms previous SOTA approaches, significantly improving the generation of high-quality, semantically consistent, and visually diverse images for complex scenes, even with intricate prompts. Minghao Liu 0022, Yingjie Tian 0001, Xiaochao Qu, Luoqi Liu, Ting Liu 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Memory Efficient Matting with Adaptive Token RoutingabstractTransformer-based models have recently achieved outstanding performance in image matting. However, their application to high-resolution images remains challenging due to the quadratic complexity of global self-attention. To address this issue, we propose MEMatte, a memory-efficient matting framework for processing high-resolution images. MEMatte incorporates a router before each global attention block, directing informative tokens to the global attention while routing other tokens to a Lightweight Token Refinement Module (LTRM). Specifically, the router employs a local-global strategy to predict the routing probability of each token, and the LTRM utilizes efficient modules to simulate global attention. Additionally, we introduce a Batch-constrained Adaptive Token Routing (BATR) mechanism, which allows each router to dynamically route tokens based on image content and the stages of attention block in the network. Furthermore, we construct an ultra high-resolution image matting dataset, UHR-395, comprising 35,500 training images and 1,000 test images, with an average resolution of 4872 × 6017. This dataset is created by compositing 395 different alpha mattes across 11 categories onto various backgrounds, all with high-quality manual annotation. Extensive experiments demonstrate that MEMatte outperforms existing methods on both high-resolution and real-world datasets, significantly reducing memory usage by approximately 88% and latency by 50% on the Composition-1K benchmark. Yiheng Lin 0002, Yihan Hu 0004, Chenyi Zhang 0004, Ting Liu 0018, Xiaochao Qu, Luoqi Liu, Yao Zhao 0001, Yunchao Wei |
AAAI | 5 |
| 2025 | NTClick: Achieving Precise Interactive Segmentation With Noise-tolerant ClicksabstractInteractive segmentation is a pivotal task in computer vision, focused on predicting precise masks with minimal user input. Although the click has recently become the most prevalent form of interaction due to its flexibility and efficiency, its advantages diminish as the complexity and details of target objects increase because it’s time-consuming and user-unfriendly to precisely locate and click on narrow, fine regions. To tackle this problem, we propose NTClick, a powerful click-based interactive segmentation method capable of predicting accurate masks even with imprecise user clicks when dealing with intricate targets. We first introduce a novel interaction form called noise-tolerant click, a type of click that does not require user’s precise localization when selecting fine regions. Then, we design a two-stage workflow, consisting of an Explicit Coarse Perception network for initial estimation and a High Resolution Refinement network for final classification. Quantitative results across extensive datasets demonstrate that NTClick not only maintains an efficient and user-friendly interaction mode but also significantly outperforms existing methods in segmentation accuracy. Chenyi Zhang 0004, Ting Liu 0018, Xiaochao Qu, Luoqi Liu, Yao Zhao 0001, Yunchao Wei |
CVPR | 3 |
| 2025 | EVPGS: Enhanced View Prior Guidance for Splatting-based Extrapolated View SynthesisabstractGaussian Splatting (GS)-based methods rely on sufficient training view coverage and perform synthesis on interpolated views. In this work, we tackle the more challenging and underexplored Extrapolated View Synthesis (EVS) task. Here we enable GS-based models trained with limited view coverage to generalize well to extrapolated views. To achieve our goal, we propose a view augmentation framework to guide training through a coarse-to-fine process. At the coarse stage, we reduce rendering artifacts due to insufficient view coverage by introducing a regularization strategy at both appearance and geometry levels. At the fine stage, we generate reliable view priors to provide further training guidance. To this end, we incorporate an occlusion awareness into the view prior generation process, and refine the view priors with the aid of coarse stage output. We call our framework Enhanced View Prior Guidance for Splatting (EVPGS). To comprehensively evaluate EVPGS on the EVS task, we collect a real-world dataset called Merchandise3D dedicated to the EVS scenario. Experiments on three datasets including both real and synthetic demonstrate EVPGS achieves state-of-the-art performance, while improving synthesis quality at extrapolated views for GS-based methods both qualitatively and quantitatively. Our code and dataset are available on the EVPGS Homepage. Jiahe Li 0006, Xiaochao Qu, Chengjing Wu, Luoqi Liu, Ting Liu 0018 |
CVPR | 3 |
| 2025 | MTADiffusion: Mask Text Alignment Diffusion Model for Object InpaintingabstractAdvancements in generative models have enabled image inpainting models to generate content within specific regions of an image based on provided prompts and masks. However, existing inpainting methods often suffer from problems such as semantic misalignment, structural distortion, and style inconsistency. In this work, we present MTADiffusion, a Mask-Text Alignment diffusion model designed for object inpainting. To enhance the semantic capabilities of the inpainting model, we introduce MTAPipeline, an automatic solution for annotating masks with detailed descriptions. Based on the MTAPipeline, we construct a new MTADataset comprising 5 million images and 25 million mask-text pairs. Furthermore, we propose a multi-task training strategy that integrates both inpainting and edge prediction tasks to improve structural stability. To promote style consistency, we present a novel inpainting style-consistency loss using a pre-trained VGG network and the Gram matrix. Comprehensive evaluations on BrushBench and EditBench demonstrate that MTADiffusion achieves state-of-the-art performance compared to other methods. Ting Liu 0018, Yihang Wu, Xiaochao Qu, Luoqi Liu |
CVPR | 4 |
| 2025 | GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text EditingabstractScene text editing, a subfield of image editing, requires modifying texts in images while preserving style consistency and visual coherence with the surrounding environment. While diffusion-based methods have shown promise in text generation, they still struggle to produce high-quality results. These methods often generate distorted or unrecognizable characters, particularly when dealing with complex characters like Chinese. In such systems, characters are composed of intricate stroke patterns and spatial relationships that must be precisely maintained. We present Glyph-Mastero, a specialized glyph encoder designed to guide the latent diffusion model for generating texts with stroke-level precision. Our key insight is that existing methods, despite using pretrained OCR models for feature extraction, fail to capture the hierarchical nature of text structures - from individual strokes to stroke-level interactions to overall character-level structure. To address this, our glyph encoder explicitly models and captures the cross-level interactions between local-level individual characters and global-level text lines through our novel glyph attention module. Meanwhile, our model implements a feature pyramid network to fuse the multi-scale OCR backbone features at the global-level. Through these cross-level and multi-scale fusions, we obtain more detailed glyph-aware guidance, enabling precise control over the scene text generation process. Our method achieves an 18.02% improvement in sentence accuracy over the state-of-the-art multi-lingual scene text editing baseline, while simultaneously reducing the text-region Fréchet inception distance by 53.28%. Ting Liu 0018, Xiaochao Qu, Chengjing Wu, Luoqi Liu |
CVPR | 3 |
| 2025 | SAM-REF: Introducing Image-Prompt Synergy during Interaction for Detail Enhancement in the Segment Anything ModelabstractInteractive segmentation is to segment the mask of the target object according to the user’s interactive prompts. There are two mainstream strategies: early fusion and late fusion. Current specialist models utilize the early fusion strategy that encodes the combination of images and prompts to target the prompted objects, yet repetitive complex computations on the images result in high latency. Late fusion models extract image embeddings once and merge them with the prompts in later interactions. This strategy avoids redundant image feature extraction and improves efficiency significantly. A recent milestone is the Segment Anything Model (SAM). However, this strategy limits the models’ ability to extract detailed information from the prompted target zone. To address this issue, we propose SAM-REF, a two-stage refinement framework that fully integrates images and prompts by using a lightweight refiner into the interaction of late fusion, which combines the accuracy of early fusion and maintains the efficiency of late fusion. Through extensive experiments, we show that our SAM-REF model outperforms the current state-of-the-art method in most metrics on segmentation quality without compromising efficiency. Chongkai Yu, Ting Liu 0018, Xiaochao Qu, Chengjing Wu, Luoqi Liu |
CVPR | 4 |
| 2022 | Distribution-Aware Single-Stage Models for Multi-Person 3D Pose EstimationabstractIn this paper, we present a novel Distribution-Aware Single-stage (DAS) model for tackling the challenging multi-person 3D pose estimation problem. Different from existing top-down and bottom-up methods, the proposed DAS model simultaneously localizes person positions and their corresponding body joints in the 3D camera space in a one-pass manner. This leads to a simplified pipeline with enhanced efficiency. In addition, DAS learns the true distribution of body joints for the regression of their positions, rather than making a simple Laplacian or Gaussian assumption as previous works. This provides valuable priors for model prediction and thus boosts the regression-based scheme to achieve competitive performance with volumetric-base ones. Moreover, DAS exploits a recur-sive update strategy for progressively approaching to regression target, alleviating the optimization difficulty and further lifting the regression performance. DAS is implemented with a fully Convolutional Neural Network and end-to-end learnable. Comprehensive experiments on benchmarks CMU Panoptic and MuPoTS-3D demonstrate the superior efficiency of the proposed DAS model, specifically 1.5x speedup over previous best model, and its stat-of-the-art accuracy for multi-person 3D pose estimation. Zitian Wang, Xuecheng Nie, Xiaochao Qu, Yunpeng Chen, Si Liu 0001 |
CVPR | 3 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 20 |
| 2019 | Reversible Data Hiding With Automatic Brightness Preserving Contrast EnhancementabstractReversible data hiding with automatic contrast enhancement methods provide an interoperable way to reduce the storage requirement for automatic image enhancement applications: original image can be recovered from the enhanced image without any additional information. Unlike the previous work, where the goal was to maximize the contrast, the proposed method increases the contrast to an appropriate level using an idea called brightness preservation. This is achieved by using an adaptive bin selection process based on the original brightness. Extensive experimental results verify that the enhanced images produced using the proposed method are visually and quantitatively superior than the existing work. Suah Kim, Rolf Lussi, Xiaochao Qu, Fangjun Huang, Hyoung Joong Kim |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Skewed Histogram Shifting for Reversible Data Hiding Using a Pair of Extreme PredictionsabstractReversible data hiding hides data in an image such that the original image is recoverable. This paper presents a novel embedding framework with reduced distortion called skewed histogram shifting using a pair of extreme predictions. Unlike traditional prediction error histogram shifting schemes, where only one good prediction is used to generate a prediction error histogram, the proposed scheme uses a pair of extreme predictions to generate two skewed histograms. By exploiting the structure of the skewed histogram, only the pixels from the peak and the short tail are used for embedding, which decreases the distortion from the lesser number of pixels being shifted. Detailed experiments and analysis are provided using several image databases. Suah Kim, Xiaochao Qu, Vasiliy Sachnev, Hyoung Joong Kim |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Reversible Data Hiding in JPEG ImagesabstractAmong various digital image formats used in daily life, the Joint Photographic Experts Group (JPEG) is the most popular. Therefore, reversible data hiding (RDH) in JPEG images is important and useful for many applications such as archive management and image authentication. However, RDH in JPEG images is considerably more difficult than that in uncompressed images because there is less information redundancy in JPEG images than that in uncompressed images, and any modification in the compressed domain may introduce more distortion in the host image. Furthermore, along with the embedding capacity and fidelity (visual quality), which have to be considered for uncompressed images, the storage size of the marked JPEG file should be considered. In this paper, based on the philosophy behind the JPEG encoder and the statistical properties of discrete cosine transform (DCT) coefficients, we present some basic insights into how to select quantized DCT coefficients for RDH. Then, a new histogram shifting-based RDH scheme for JPEG images is proposed, in which the zero coefficients remain unchanged and only coefficients with values 1 and -1 are expanded to carry message bits. Moreover, a block selection strategy based on the number of zero coefficients in each 8 × 8 block is proposed, which can be utilized to adaptively choose DCT coefficients for data hiding. Experimental results demonstrate that by using the proposed method we can easily realize high embedding capacity and good visual quality. The storage size of the host JPEG file can also be well preserved. Fangjun Huang, Xiaochao Qu, Hyoung Joong Kim, Jiwu Huang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Weighted sparse representation using a learned distance metric for face recognitionabstractThis paper presents a novel weighted sparse representation classification for face recognition with a learned distance metric (WSRC-LDM) which learns a Mahalanobis distance to calculate the weight and code the testing face. The Mahalanobis distance is learned by using the information-theoretic metric learning (ITML) which helps to define a better weight used in WSRC. In the meantime, the learned distance metric takes advantage of the classification rule of SRC which helps the proposed method classify more accurately. Extensive experiments verify the effectiveness of the proposed method. Xiaochao Qu, Suah Kim, Dessalegn Atnafu, Hyoung Joong Kim |
ICIP | 1 |
| 2015 | Local pixel patternsabstractIn this paper, a new class of image texture operators is proposed. We firstly determine that the number of gray levels in each B × B subblock is a fundamental property of the local image texture. Thus, an occurrence histogram for each B × B sub-block can be utilized to describe the texture of the image. Moreover, using a new multi-bit plane strategy, i.e., representing the image texture with the occurrence histogram of the first one or more significant bit-planes of the input image, more powerful operators for describing the image texture can be obtained. The proposed approach is invariant to gray scale variations since the operators are, by definition, invariant under any monotonic transformation of the gray scale, and robust to rotation. They can also be used as supplementary operators to local binary patterns (LBP) to improve their capability to resist illuminance variation, surface transformations, etc. Fangjun Huang, Xiaochao Qu, Hyoung Joong Kim, Jiwu Huang |
Comput. Vis. Media | 2 |
| 2015 | Linear collaborative discriminant regression classification for face recognition
Xiaochao Qu, Suah Kim, Run Cui, Hyoung Joong Kim |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | Pixel-based pixel value ordering predictor for high-fidelity reversible data hiding
Xiaochao Qu, Hyoung Joong Kim |
Signal Process. | 1 |
| 2014 | Reversible Data Hiding Based on Combined Predictor and Prediction Error Expansion
Xiaochao Qu, Suah Kim, Run Cui, Fangjun Huang, Hyoung Joong Kim |
IWDW | 1 |
| 2012 | Automated Motion Correction for In Vivo Optical Projection TomographyabstractIn in vivo optical projection tomography (OPT), object motion will significantly reduce the quality and resolution of the reconstructed image. Based on the well-known Helgason-Ludwig consistency condition (HLCC), we propose a novel method for motion correction in OPT under parallel beam illumination. The method estimates object motion from projection data directly and does not require any other additional information, which results in a straightforward implementation. We decompose object movement into translation and rotation, and discuss how to correct for both translation and general motion simultaneously. Since finding the center of rotation accurately is critical in OPT, we also point out that the system's geometrical offset can be considered as object translation and therefore also calibrated through the translation estimation method. In order to verify the algorithm effectiveness, both simulated and in vivo OPT experiments are performed. Our results demonstrate that the proposed approach is capable of decreasing movement artifacts significantly thus providing high quality reconstructed images in the presence of object motion. Shouping Zhu, Di Dong, Udo Birk, Matthias Rieckher, Nektarios Tavernarakis, Xiaochao Qu, Jimin Liang, Jie Tian 0001, Jorge Ripoll |
IEEE Trans. Medical Imaging | 6 |