EDBT 2026 Demo / reviewers in the wild / expert
Xianbiao Qi
dblp:118/3741
· DBLP profile ↗
43ranked-venue papers
13as first author
23since 2021 · last 2026
0000-0002-8493-1966ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 11 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A cyclic diffusion framework for structure-authentic and annotation-disentangled anomaly generation
Linchun Wu, Qin Zou 0001, Xianbiao Qi, Zhongyuan Wang 0001, Qingquan Li 0001 |
Neurocomputing | 3 |
| 2026 | HCAttention: Extreme KV cache compression via heterogeneous attention computing for LLMs
Dongquan Yang, Xiaotian Yu, Xianbiao Qi, Rong Xiao 0003 |
Neurocomputing | 4 |
| 2025 | CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and CompatibilityabstractVideo inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited number of text-guided video inpainting techniques, and these techniques struggle with maintaining visual quality and exhibit poor semantic representation capabilities. In this paper, we introduce CoCoCo, a text-guided video inpainting diffusion framework. To address the aforementioned challenges, we enhance both the training data and model structure. Specifically, we devise an instance-aware region selection strategy for masked area sampling and develop a novel motion block that incorporates efficient 3D full attention and textual cross attention. Additionally, our CoCoCo framework can be seamlessly integrated with various personalized text-to-image diffusion models through a delicate training-free transfer mechanism. Comprehensive experiments demonstrate that CoCoCo can create high-quality visual content with enhanced temporal consistency, improved text controllability, and better compatibility with personalized image models. Bojia Zi, Xianbiao Qi, Yukai Shi, Bin Liang 0004, Rong Xiao 0003, Kam-Fai Wong, Lei Zhang 0001 |
AAAI | 3 |
| 2025 | BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation CapabilitiesabstractWe introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same framework.
BiGR features a binary tokenizer, a masked modeling mechanism, and a binary transcoder for binary code prediction.
Additionally, we introduce a novel entropy-ordered sampling method to enable efficient image generation.
Extensive experiments validate BiGR's superior performance in generation quality, as measured by FID-50k, and representation capabilities, as evidenced by linear-probe accuracy.
Moreover, BiGR showcases zero-shot generalization across various vision tasks, enabling applications such as image inpainting, outpainting, editing, interpolation, and enrichment, without the need for structural modifications. Our findings suggest that BiGR unifies generative and discriminative tasks effectively, paving the way for further advancements in the field. We further enable BiGR to perform text-to-image generation, showcasing its potential for broader applications. Shaozhe Hao, Xuantong Liu, Xianbiao Qi, Bojia Zi, Rong Xiao 0003, Kai Han 0001, Kwan-Yee Kenneth Wong |
ICLR | 3 |
| 2025 | Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain ModelabstractRoom layout estimation from multiple-perspective images is poorly investigated due to the complexities that emerge from multi-view geometry, which requires muti-step solutions such as camera intrinsic and extrinsic estimation, image matching, and triangulation. However, in 3D reconstruction, the advancement of recent 3D foundation models such as DUSt3R has shifted the paradigm from the traditional multi-step structure-from-motion process to an end-to-end single-step approach.
To this end, we introduce Plane-DUSt3R, a novel method for multi-view room layout estimation leveraging the 3D foundation model DUSt3R. Plane-DUSt3R incorporates the DUSt3R framework and fine-tunes on a room layout dataset (Structure3D) with a modified objective to estimate structural planes. By generating uniform and parsimonious results, Plane-DUSt3R enables room layout estimation with only a single post-processing step and 2D detection results.
Unlike previous methods that rely on single-perspective or panorama image, Plane-DUSt3R extends the setting to handle multiple-perspective images. Moreover, it offers a streamlined, end-to-end solution that simplifies the process and reduces error accumulation.
Experimental results demonstrate that Plane-DUSt3R not only outperforms state-of-the-art methods on the synthetic dataset but also proves robust and effective on in the wild data with different image styles such as cartoon. Our code is available at: https://github.com/justacar/Plane-DUSt3R Yaxuan Huang, Xili Dai, Xianbiao Qi, Yixing Yuan, Xiangyu Yue 0001 |
ICLR | 4 |
| 2025 | Exploring a Principled Framework for Deep Subspace ClusteringabstractSubspace clustering is a classical unsupervised learning task, built on a basic assumption that high-dimensional data can be approximated by a union of subspaces (UoS). Nevertheless, the real-world data are often deviating from the UoS assumption. To address this challenge, state-of-the-art deep subspace clustering algorithms attempt to jointly learn UoS representations and self-expressive coefficients. However, the general framework of the existing algorithms suffers from feature collapse and lacks a theoretical guarantee to learn desired UoS representation. In this paper, we present a Principled fRamewOrk for Deep Subspace Clustering (PRO-DSC), which is designed to learn structured representations and self-expressive coefficients in a unified manner. Specifically, in PRO-DSC, we incorporate an effective regularization on the learned representations into the self-expressive model, prove that the regularized self-expressive model is able to prevent feature space collapse, and demonstrate that the learned optimal representations under certain condition lie on a union of orthogonal subspaces. Moreover, we provide a scalable and efficient approach to implement our PRO-DSC and conduct extensive experiments to verify our theoretical findings and demonstrate the superior performance of our proposed deep subspace clustering approach. Xianghan Meng, Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li |
ICLR | 4 |
| 2025 | Taming Transformer Without Using Learning Rate WarmupabstractScaling Transformer to a large scale without using some technical tricks such as learning rate warump and an obviously lower learning rate, is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal a key problem behind model crash phenomenon in the training process, termed *spectral energy concentration* of ${W_q}^{\top} W_k$, which is the reason for a malignant entropy collapse, where ${W_q}$ and $W_k$ are the projection matrices for the query and the key in Transformer, respectively.
To remedy this problem, motivated by *Weyl's Inequality*, we present a novel optimization strategy, \ie, making the weight updating in successive steps steady---if the ratio $\frac{\sigma_{1}(\nabla W_t)}{\sigma_{1}(W_{t-1})}$ is larger than a threshold, we will automatically bound the learning rate to a weighted multiple of $\frac{\sigma_{1}(W_{t-1})}{\sigma_{1}(\nabla W_t)}$, where $\nabla W_t$ is the updating quantity in step $t$. Such an optimization strategy can prevent spectral energy concentration to only a few directions, and thus can avoid malignant entropy collapse which will trigger the model crash. We conduct extensive experiments using ViT, Swin-Transformer and GPT, showing that our optimization strategy can effectively and stably train these (Transformer) models without using learning rate warmup. Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li, Bojia Zi, Xili Dai, Qin Zou 0001, Rong Xiao 0003 |
ICLR | 1 |
| 2025 | Elucidating the design space of language models for image generationabstractThe success of large language models (LLMs) in text generation has inspired their application to image generation. However, existing methods either rely on specialized designs with inductive biases or adopt LLMs without fully exploring their potential in vision tasks. In this work, we systematically investigate the design space of LLMs for image generation and demonstrate that LLMs can achieve near state-of-the-art performance without domain-specific designs, simply by making proper choices in tokenization methods, modeling approaches, scan patterns, vocabulary design, and sampling strategies. We further analyze autoregressive models' learning and scaling behavior, revealing how larger models effectively capture more useful information than the smaller ones. Additionally, we explore the inherent differences between text and image modalities, highlighting the potential of LLMs across domains. The exploration provides valuable insights to inspire more effective designs when applying LLMs to other domains. With extensive experiments, our proposed model, **ELM** achieves an FID of 1.54 on 256$\times$256 ImageNet and an FID of 3.29 on 512$\times$512 ImageNet, demonstrating the powerful generative potential of LLMs in vision tasks. Xuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu 0001, Jun Wang 0123, Rong Xiao 0003, Yuan Yao 0011 |
ICML | 3 |
| 2025 | BrokenVideos: A Benchmark Dataset for Fine-Grained Artifact Localization in AI-Generated VideosabstractThe field of video generation has witnessed remarkable advances in recent years, driven by innovations in deep generative models. Nevertheless, the fidelity of AI-generated videos remains far from perfect, with synthesized content frequently exhibiting visual artifacts, such as temporally inconsistent motion, physically implausible trajectories, unnatural object deformations, and local blurring, that undermine realism and user trust. Precise detection and spatial localization of these artifacts are of critical importance: not only are they essential for automatic quality control pipelines that improves user experience, but they also provide actionable diagnostic signals for researchers and practitioners to guide model development and evaluation. Despite its significance, the research community currently lacks a comprehensive benchmark tailored for artifact localization in AI-generated videos. Existing datasets either focus solely on detection at the video or frame level, or lack fine-grained spatial annotations necessary for developing and benchmarking localization methods. To fill this gap, we present BrokenVideos, a benchmark dataset comprising ~3,254 AI-generated videos with carefully-annotated, pixel-level masks indicating regions of visual corruption. Each annotation is the result of careful human inspection, ensuring high-quality ground truth for artifact localization tasks. We demonstrate that training existing video artifact detection models and multi-modal large language models (MLLMs) on BrokenVideos substantially enhances their ability to localize corrupted regions within generated content. Through extensive experiments and cross-model evaluations, we show that BrokenVideos provides a critical foundation for both benchmarking and advancing artifact localization research. We hope our dataset can catalyze further innovation in both video generation and its quality assurance. The dataset is available at: https://broken-video-detection-datetsets.github.io/Broken-Video-Detection-Datasets.github.io/. Weixuan Peng, Bojia Zi, Yifeng Gao 0002, Xianbiao Qi, Xingjun Ma, Yu-Gang Jiang 0001 |
ACM Multimedia | 5 |
| 2025 | MiniMax-Remover: Taming Bad Noise Helps Video Object RemovalabstractRecent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated objects and visual artifacts. Furthermore, existing methods often rely on computationally expensive sampling procedures and classifier-free guidance (CFG), resulting in slow inference. To address these limitations, we propose **MiniMax-Remover**, a novel two-stage video object removal approach. Motivated by the observation that text condition is not best suited for this task, we simplify the pretrained video generation model by removing textual input and cross-attention layers. In this way, we obtain a more lightweight and efficient model architecture in the first stage.
In the second stage, we proposed a minimax optimization strategy to further distill the remover with the successful videos produced by stage-1 model. Specifically, the inner maximization identifies adversarial input noise ("bad noise'') that leads to failure removals, while the outer minimization trains the model to generate high-quality removal results even under such challenging conditions. As a result, our method achieves a state-of-the-art video object removal results using as few as 6 sampling steps without CFG usage. Extensive experiments demonstrate the effectiveness and superiority of MiniMax-Remover compared to existing methods. Codes and Videos are available at: **https://minimax-remover.github.io**. Bojia Zi, Weixuan Peng, Xianbiao Qi, Rong Xiao 0003, Kam-Fai Wong |
NeurIPS | 3 |
| 2025 | Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video SpecialistsabstractVideo content editing has a wide range of applications. With the advancement of diffusion-based generative models, video editing techniques have made remarkable progress, yet they still remain far from practical usability. Existing inversion-based video editing methods are time-consuming and struggle to maintain consistency in unedited regions. Although instruction-based methods have high theoretical potential, they face significant challenges in constructing high-quality training datasets - current datasets suffer from issues such as editing correctness, frame consistency, and sample diversity. To bridge these gaps, we introduce the Señorita-2M dataset, a large-scale, diverse, and high-quality video editing dataset. We systematically categorize editing tasks into 2 classes consisting of 18 subcategories. To build this dataset, we design four new task specialists and employ or modify 14 existing task experts to generate data samples for each subclass. In addition, we design a filtering pipeline at both the visual content and instruction levels to further enhance data quality. This approach ensures the reliability of constructed data. Finally, the Señorita-2M dataset comprises 2 million high-fidelity samples with diverse resolutions and frame counts. We trained multiple models using different base video models, i.e., Wan2.1 and CogVideoX-5B, on Señorita-2M, and the results demonstrate that the models exhibit superior visual quality, robust frame-to-frame consistency, and strong instruction following capability. More videos are available at: https://senorita-2m-dataset.github.io. Bojia Zi, Penghui Ruan, Marco Chen, Xianbiao Qi, Shaozhe Hao, Youze Huang, Bin Liang 0004, Rong Xiao 0003, Kam-Fai Wong |
NeurIPS | 4 |
| 2025 | Neural Normalized Cut: A differential and generalizable approach for spectral clustering
Shangzhi Zhang, Chun-Guang Li, Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002 |
Pattern Recognit. | 4 |
| 2024 | Graph Cut-Guided Maximal Coding Rate Reduction for Learning Image Embedding and Clustering
Xianghan Meng, Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li |
ACCV (10) | 4 |
| 2024 | DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D GenerationabstractText-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models often exhibit two limitations: (a) quality concerns such as missing attributes and distorted shape and texture; (b) extremely low diversity comparing to text-guided image synthesis. In this paper, we show that the conflict between the 3D optimization process and uniform timestep sampling in score distillation is the main reason for these limitations. To resolve this conflict, we propose to prioritize timestep sampling with monotonically non-increasing functions, which aligns the 3D optimization process with the sampling process of diffusion model. Extensive experiments show that our simple redesign significantly improves 3D content creation with faster convergence, better quality and diversity. Yukai Shi, Boshi Tang, Xianbiao Qi, Lei Zhang 0001 |
ICLR | 5 |
| 2024 | TOSS: High-quality Text-guided Novel View Synthesis from a Single ImageabstractIn this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image.
While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often result in implausible NVS generations.
To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space.
TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details.
Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with higher-quality NVS results and faster convergence. We further support these results with comprehensive ablations that underscore the effectiveness and potential of
the introduced semantic guidance and architecture design. Yukai Shi, He Cao, Boshi Tang, Xianbiao Qi, Tianyu Yang 0003, Shilong Liu 0004, Lei Zhang 0001, Harry Shum |
ICLR | 5 |
| 2023 | DisCo-CLIP: A Distributed Contrastive Loss for Memory Efficient CLIP TrainingabstractWe propose DisCo-CLIP, a distributed memory-efficient CLIP training approach, to reduce the memory consumption of contrastive loss when training contrastive learning models. Our approach decomposes the contrastive loss and its gradient computation into two parts, one to calculate the intra-GPU gradients and the other to compute the inter-GPU gradients. According to our decomposition, only the intra-GPU gradients are computed on the current GPU, while the inter-GPU gradients are collected via all_reduce from other GPUs instead of being repeatedly computed on every GPU. In this way, we can reduce the GPU memory consumption of contrastive loss computation from$\mathcal{O}(B^{2})$to$\mathcal{O}(\frac{B^{2}}{N})$, where$B$and$N$are the batch size and the number of GPUs used for training. Such a distributed solution is mathematically equivalent to the original non-distributed contrastive loss computation, without sacrificing any computation accuracy. It is particularly efficient for large-batch CLIP training. For instance, DisCo-CLIP can enable contrastive training of a ViT-B/32 model with a batch size of 32K or 196K using 8 or 64 A100 40GB GPUs, compared with the original CLIP solution which requires 128 A100 40GB GPUs to train a ViT-B/32 model with a batch size of 32K. Xianbiao Qi, Lei Zhang 0001 |
CVPR | 2 |
| 2023 | LipsFormer: Introducing Lipschitz Continuity to Vision Transformers
Xianbiao Qi, Yukai Shi, Lei Zhang 0001 |
ICLR | 1 |
| 2023 | DreamWaltz: Make a Scene with Complex 3D Animatable AvatarsabstractWe present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains challenging. To create high-quality 3D avatars, DreamWaltz proposes 3D-consistent occlusion-aware Score Distillation Sampling (SDS) to optimize implicit neural representations with canonical poses. It provides view-aligned supervision via 3D-aware skeleton conditioning which enables complex avatar generation without artifacts and multiple faces. For animation, our method learns an animatable 3D avatar representation from abundant image priors of diffusion model conditioned on various poses, which could animate complex non-rigged avatars given arbitrary poses without retraining. Extensive evaluations demonstrate that DreamWaltz is an effective and robust approach for creating 3D avatars that can take on complex shapes and appearances as well as novel poses for animation. The proposed framework further enables the creation of complex scenes with diverse compositions, including avatar-avatar, avatar-object and avatar-scene interactions. See https://dreamwaltz3d.github.io/ for more vivid 3D avatar and animation results. Ailing Zeng, He Cao, Xianbiao Qi, Yukai Shi, Zhengjun Zha, Lei Zhang 0001 |
NeurIPS | 5 |
| 2022 | DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Shilong Liu 0004, Feng Li 0040, Hao Zhang 0097, Xiao Yang 0028, Xianbiao Qi, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ICLR | 5 |
| 2022 | Learning graph normalization for graph neural networks
Xianbiao Qi, Chun-Guang Li, Rong Xiao 0003 |
Neurocomputing | 3 |
| 2022 | EMU: Effective Multi-Hot Encoding Net for Lightweight Scene Text Recognition With a Large Character SetabstractDeploying a lightweight deep model for scene text recognition task on mobile devices has great commercial value. However, the conventional softmax-based one-hot classification module becomes a cumbersome obstacle when handling multi-languages or languages with large character set (e.g., Chinese) due to the rapid expansion of model parameters with the number of classes. To this end, we propose an Effective Multi-hot encoding and classification modUle (EMU) for scene text recognition in the scenario of multi-languages or languages with large character set. Specifically, EMU generates a binary multi-hot label for each class with a real-valued sub-network in training stage and produces the prediction by calculating the inner product between the multi-hot code and the multi-hot label. Compared to the softmax-based one-hot classifier, EMU reduces the storage requirement and the time cost in inference stage significantly, retaining similar performance. Furthermore, we design a convolution feature basedLightweight TransFormerto learn the effective features for EMU and consequently develop a lightweight scene text recognition framework, termedLight-Former-EMU. We conduct extensive experiments on seven public English benchmarks and two real-world Chinese challenge benchmarks. Experimental results verify the effectiveness of the proposed EMU and demonstrate the promising performance of the proposed Light-Former-EMU. Bingcong Li, Xianbiao Qi, Chun-Guang Li, Rong Xiao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | MASTER: Multi-aspect non-local network for scene text recognition
Ning Lu 0003, Wenwen Yu, Xianbiao Qi, Ping Gong 0003, Rong Xiao 0003, Xiang Bai |
Pattern Recognit. | 3 |
| 2021 | Detail-Preserving Multi-Exposure Fusion With Edge-Preserving Structural Patch DecompositionabstractThe multi-exposure fusion (MEF) methods have received much attention in recent years due to the importance of constructing high dynamic range images. Among most of the existing studies, multi-scale structural-patch-decomposition-based MEF (MSPD-MEF) has achieved state-of-the-art fusion quality and the fastest running time. However, this method still suffers from detail loss in the fused images. To tackle this issue, we first incorporate the edge-preserving factors into this method to preserve the details in the fused images in a single-scale setting. Then, we develop the novel and flexible bell curve function, which can further preserve the details in both bright and dark regions. After that, we also show that our method can seamlessly plug in to this multi-scale framework. Extensive experimental results indicate that the proposed method can produce pleasing fusion results with little artifacts and low computational cost in both static and dynamic scenes. Hui Li 0029, Tsz Nam Chan, Xianbiao Qi, Wuyuan Xie |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional NetworksabstractComputer vision with state-of-the-art deep learning models has achieved huge success in the field of Optical Character Recognition (OCR) including text detection and recognition tasks recently. However, Key Information Extraction (KIE) from documents as the downstream task of OCR, having a large number of use scenarios in real-world, remains a challenge because documents not only have textual features extracting from OCR systems but also have semantic visual features that are not fully exploited and play a critical role in KIE. Too little work has been devoted to efficiently make full use of both textual and visual features of the documents. In this paper, we introduce PICK, a framework that is effective and robust in handling complex documents layout for KIE by combining graph learning with graph convolution operation, yielding a richer semantic representation containing the textual and visual features and global layout without ambiguity. Extensive experiments on realworld datasets have been conducted to show that our method outperforms baselines methods by significant margins. Our code is available at https://github.com/wenwenyu/PICK-pytorch. Wenwen Yu, Ning Lu 0003, Xianbiao Qi, Ping Gong 0003, Rong Xiao 0003 |
ICPR | 3 |
| 2020 | Learning Convolution Feature Aggregation via Edge Attention Convolution Network for Person Re-IdentificationabstractPerson Re-Identification (Re-ID) is a challenging task of matching pedestrian images collected from nonoverlapping multiple camera views due to huge variations from pose changes, occlusions, varying illumination and clutter background. Recently, graph convolution network or graph neural network increasingly gains a lot of research attention in person Re-ID. However, the existing methods have not fully exploit the available features on the graph. In this paper, we propose an efficient and effective end-to-end trainable framework, termed Edge Attention Convolution Network (EACN), to perform convolution feature learning and attentive feature aggregation for person Re-ID, in which the learned convolution features on vertex and its edges are attentively aggregated on a dynamic graph. We conduct extensive experiments on two large benchmark datasets, Market-1501 and DukeMTMC. Experimental results validate the efficiency and effectiveness of our proposal. Chaoqun Lin, Ruo-Pei Guo, Xianbiao Qi, Chun-Guang Li |
VCIP | 4 |
| 2020 | Fast infrared and visible image fusion with structural decomposition
Hui Li 0029, Xianbiao Qi, Wuyuan Xie |
Knowl. Based Syst. | 2 |
| 2019 | Self-Supervised Convolutional Subspace Clustering NetworkabstractSubspace clustering methods based on data self-expression have become very popular for learning from data that lie in a union of low-dimensional linear subspaces. However, the applicability of subspace clustering has been limited because practical visual data in raw form do not necessarily lie in such linear subspaces. On the other hand, while Convolutional Neural Network (ConvNet) has been demonstrated to be a powerful tool for extracting discriminative features from visual data, training such a ConvNet usually requires a large amount of labeled data, which are unavailable in subspace clustering applications. To achieve simultaneous feature learning and subspace clustering, we propose an end-to-end trainable framework, called Self-Supervised Convolutional Subspace Clustering Network (S$^2$ConvSCN), that combines a ConvNet module (for feature learning), a self-expression module (for subspace clustering) and a spectral clustering module (for self-supervision) into a joint optimization framework. Particularly, we introduce a dual self-supervision that exploits the output of spectral clustering to supervise the training of the feature learning module (via a classification loss) and the self-expression module (via a spectral clustering loss). Our experiments on four benchmark datasets show the effectiveness of the dual self-supervision and demonstrate superior performance of our proposed approach. Chun-Guang Li, Chong You, Xianbiao Qi, Honggang Zhang 0002, Jun Guo 0002, Zhouchen Lin |
CVPR | 4 |
| 2019 | Homocentric Hypersphere Feature Embedding for Person Re-IdentificationabstractTriplet loss and softmax loss are two widely used loss functions in Person Re-Identification (Person ReID). However, previous works that try to apply these two loss functions have measure inconsistency during training and testing stage and among different parts of the total loss function, which would cause inferior performance of models. To address this issue, we propose a novel homocentric hypersphere embedding scheme to decouple magnitude and orientation information for both feature and weight vectors, and reformulate the triplet loss and the softmax loss to their angular versions and combine them into an angular discriminative loss. We evaluate our proposed method extensively on the widely used Person ReID benchmarks. Our method demonstrates leading performance on all datasets. Wangmeng Xiang, Jianqiang Huang 0001, Xianbiao Qi, Xian-Sheng Hua 0001, Lei Zhang 0006 |
ICIP | 3 |
| 2019 | DeepCrack: Learning Hierarchical Convolutional Features for Crack DetectionabstractCracks are typical line structures that are of interest in many computer-vision applications. In practice, many cracks, e.g., pavement cracks, show poor continuity and low contrast, which brings great challenges to image-based crack detection by using low-level features. In this paper, we propose DeepCrack - an end-to-end trainable deep convolutional neural network for automatic crack detection by learning high-level features for crack representation. In this method, multi-scale deep convolutional features learned at hierarchical convolutional stages are fused together to capture the line structures. More detailed representations are made in larger-scale feature maps and more holistic representations are made in smaller-scale feature maps. We build DeepCrack net on the encoder-decoder architecture of SegNet, and pairwisely fuse the convolutional features generated in the encoder network and in the decoder network at the same scale. We train DeepCrack net on one crack dataset and evaluate it on three others. The experimental results demonstrate that DeepCrack achieves F-Measure over 0.87 on the three challenging datasets in average and outperforms the current state-of-the-art methods. Qin Zou 0001, Zheng Zhang 0036, Qingquan Li 0001, Xianbiao Qi, Qian Wang 0002, Song Wang 0002 |
IEEE Trans. Image Process. | 4 |
| 2017 | 3D Surface Detail Enhancement from a Single Normal MapabstractIn 3D reconstruction, the obtained surface details are mainly limited to the visual sensor due to sampling and quantization in the digitalization process. How to get a fine-grained 3D surface with low-cost is still a challenging obstacle in terms of experience, equipment and easyto-obtain. This work introduces a novel framework for enhancing surfaces reconstructed from normal map, where the assumptions on hardware (e.g., photometric stereo setup) and reflection model (e.g., Lambertion reflection) are not necessarily needed. We propose to use a new measure, angle profile, to infer the hidden micro-structure from existing surfaces. In addition, the inferred results are further improved in the domain of discrete geometry processing (DGP) which is able to achieve a stable surface structure under a selectable enhancement setting. Extensive simulation results show that the proposed method obtains significantly improvements over uniform sharpening method in terms of both subjective visual assessment and objective quality metric. Wuyuan Xie, Miaohui Wang, Xianbiao Qi, Lei Zhang 0006 |
ICCV | 3 |
| 2017 | HEp-2 Cell Classification via Combining Multiresolution Co-Occurrence Texture and Large Region Shape InformationabstractIndirect immunofluorescence imaging of human epithelial type 2 (HEp-2) cell image is an effective evidence to diagnose autoimmune diseases. Recently, computer-aided diagnosis of autoimmune diseases by the HEp-2 cell classification has attracted great attention. However, the HEp-2 cell classification task is quite challenging due to large intraclass and small interclass variations. In this paper, we propose an effective approach for the automatic HEp-2 cell classification by combining multiresolution co-occurrence texture and large regional shape information. To be more specific, we propose to: 1) capture multiresolution co-occurrence texture information by a novel pairwise rotation-invariant co-occurrence of local Gabor binary pattern descriptor; 2) depict large regional shape information by using an improved Fisher vector model with RootSIFT features, which are sampled from large image patches in multiple scales; and 3) combine both features. We evaluate systematically the proposed approach on the IEEE International Conference on Pattern Recognition (ICPR) 2012, the IEEE International Conference on Image Processing (ICIP) 2013, and the ICPR 2014 contest datasets. The proposed method based on the combination of the introduced two features outperforms the winners of the ICPR 2012 contest using the same experimental protocol. Our method also greatly improves the winner of the ICIP 2013 contest under four different experimental setups. Using the leave-one-specimen-out evaluation strategy, our method achieves comparable performance with the winner of the ICPR 2014 contest that combined four features. Xianbiao Qi, Guoying Zhao 0001, Chun-Guang Li, Jun Guo 0002, Matti Pietikäinen |
IEEE J. Biomed. Health Informatics | 1 |
| 2016 | Dynamic texture and scene classification by transferring deep image features
Xianbiao Qi, Chun-Guang Li, Guoying Zhao 0001, Xiaopeng Hong, Matti Pietikäinen |
Neurocomputing | 1 |
| 2016 | LOAD: Local orientation adaptive descriptor for texture and material classification
Xianbiao Qi, Guoying Zhao 0001, LinLin Shen, Qingquan Li 0001, Matti Pietikäinen |
Neurocomputing | 1 |
| 2016 | Exploring illumination robust descriptors for human epithelial type 2 cell classification
Xianbiao Qi, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen |
Pattern Recognit. | 1 |
| 2016 | HEp-2 cell classification: The role of Gaussian Scale Space Theory as a pre-processing approach
Xianbiao Qi, Guoying Zhao 0001, Jie Chen 0001, Matti Pietikäinen |
Pattern Recognit. Lett. | 1 |
| 2015 | Discriminative regional color co-occurrence descriptorabstractTraditional color feature descriptors are focused on color-value distributions in the color space, e.g., color histograms, color bag-of-words, which ignore the spatial location and contextual information of different colors. In this paper, a new regional color co-occurrence feature descriptor (RCC) is proposed to reflect spatial relations of colors in an image. First, we partition an image into a number of disjoint regions using superpixel techniques. Then, we construct a color histogram for each region, based on which we construct a color co-occurrence matrix for each pair of neighboring regions. Finally, all the constructed co-occurrence matrices from an image are summed up and normalized as a color descriptor to represent this image. This new color descriptor reflects the color-collocation patterns in the image. We use this new color descriptor for image/object classification and find that it leads to higher classification accuracies than other competing color descriptors. Qin Zou 0001, Xianbiao Qi, Qingquan Li 0001, Song Wang 0002 |
ICIP | 2 |
| 2015 | Globally rotation invariant multi-scale co-occurrence local binary pattern
Xianbiao Qi, LinLin Shen, Guoying Zhao 0001, Qingquan Li 0001, Matti Pietikäinen |
Image Vis. Comput. | 1 |
| 2014 | Pairwise Rotation Invariant Co-Occurrence Local Binary PatternabstractDesigning effective features is a fundamental problem in computer vision. However, it is usually difficult to achieve a great tradeoff between discriminative power and robustness. Previous works shown that spatial co-occurrence can boost the discriminative power of features. However the current existing co-occurrence features are taking few considerations to the robustness and hence suffering from sensitivity to geometric and photometric variations. In this work, we study the Transform Invariance (TI) of co-occurrence features. Concretely we formally introduce a Pairwise Transform Invariance (PTI) principle, and then propose a novel Pairwise Rotation Invariant Co-occurrence Local Binary Pattern (PRICoLBP) feature, and further extend it to incorporate multi-scale, multi-orientation, and multi-channel information. Different from other LBP variants, PRICoLBP can not only capture the spatial context co-occurrence information effectively, but also possess rotation invariance. We evaluate PRICoLBP comprehensively on nine benchmark data sets from five different perspectives, e.g., encoding strategy, rotation invariance, the number of templates, speed, and discriminative power compared to other LBP variants. Furthermore we apply PRICoLBP to six different but related applications-texture, material, flower, leaf, food, and scene classification, and demonstrate that PRICoLBP is efficient, effective, and of a well-balanced tradeoff between the discriminative power and robustness. Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li, Yu Qiao 0001, Jun Guo 0002, Xiaoou Tang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Exploring Motion Boundary based Sampling and Spatial-Temporal Context Descriptors for Action RecognitionabstractFeature representation is important for human action recognition.Recently, Wang et al. [25] proposed dense trajectory (DT) based features for action video representation and achieved state-of-the-art performance on several action datasets.In this paper, we improve the DT method in two folds.Firstly, we introduce a motion boundary based dense sampling strategy, which greatly reduces the number of valid trajectories while preserves the discriminative power.Secondly, we develop a set of new descriptors which describe the spatial-temporal context of motion trajectories.To evaluate the performance of the proposed methods, we conduct extensive experiments on three benchmarks including K-TH, YouTube and HMDB51.The results show that our sampling strategy significantly reduces the computational cost of point tracking without degrading performance.Meanwhile, we achieve superior performance than the state-of-the-art methods by utilizing our spatial-temporal context descriptors. Xiaojiang Peng, Yu Qiao 0001, Qiang Peng, Xianbiao Qi |
BMVC | 4 |
| 2013 | Multi-scale Joint Encoding of Local Binary Patterns for Texture and Material Classification
Xianbiao Qi, Yu Qiao 0001, Chun-Guang Li, Jun Guo 0002 |
BMVC | 1 |
| 2013 | Exploring Cross-Channel Texture Correlation for Color Texture ClassificationabstractThis paper proposes a novel approach to encode cross-channel texture correlation for color texture classification task. Firstly, we quantitatively study the correlation between different color channels using Local Binary Pattern (LBP) as the texture descriptor and using Shannon’s information theory to measure the correlation. We find that (R, G) channel pair exhibits stronger correlation than (R, B) and (G, B) channel pairs. Secondly, we propose a novel descriptor to encode the cross-channel texture correlation. The proposed descriptor can capture well the relative variance of texture patterns between different channels. Meanwhile, our descriptor is computationally efficient and robust to image rotation. We conduct extensive experiments on four challenging color texture databases to validate the effectiveness of the proposed approach. The experimental results show that the proposed approach significantly outperforms its mostly relevant counterpart (Multichannel color LBP), and achieves the state-of-the-art performance. Xianbiao Qi, Yu Qiao 0001, Chun-Guang Li, Jun Guo 0002 |
BMVC | 1 |
| 2012 | Pairwise Rotation Invariant Co-occurrence Local Binary Pattern
Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002, Lei Zhang 0001 |
ECCV (6) | 1 |
| 2012 | A rapid flower/leaf recognition systemabstractIn this work, we introduce a rapid and accurate flower/leaf recognition system. The system could process one query in less than 0.35s with users' simple interaction. Meanwhile, high accuracy and recall is achieved. Furthermore, low computational resource and memory cost are required by the system. Now, the system is demonstrated on 172 categories of flowers, the largest flower dataset until now, and 220 categories of leaves. Xianbiao Qi, Rong Xiao 0003, Lei Zhang 0001, Chun-Guang Li, Jun Guo 0002 |
ACM Multimedia | 1 |