Rong Xiao 0003

dblp:75/5560-3 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
17since 2021 · last 2026
0000-0003-2207-5698ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3Applied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 UVLM: Benchmarking Video Language Model for Underwater World Understanding
abstract
Recently, video-language models (VidLMs) have gained widespread attention and adoption. However, existing works primarily focus on terrestrial scenarios, overlooking the highly demanding application needs of underwater observation. To overcome this gap, we introduce UVLM, an under water observation benchmark which is build through a collaborative approach combining human expertise and AI models. To ensure data quality, we have conducted in-depth considerations from multiple perspectives. First, to address the unique challenges of underwater environments, we selected videos that represent typical underwater challenges including light variations, water turbidity, and diverse viewing angles to construct the dataset. Second, to ensure data diversity, the dataset covers a wide range of frame rates, resolutions, 419 classes of marine animals, and various static plants and terrains. Next, for task diversity, we adopted a structured design where observation targets are categorized into two major classes: biological and environmental. Each category includes content observation and change/action observation, totaling 20 subtask types. Finally, we designed several challenging evaluation metrics to enable quantitative comparison and analysis of different methods. Experiments on two representative VidLMs demonstrate that fine-tuning VidLMs on UVLM significantly improves underwater world understanding while also showing potential for slight improvements on existing in-air VidLM benchmarks.
Xizhe Xue, Dawei Yan 0001, Lijie Tao, Ying Li 0017, Haokui Zhang, Rong Xiao 0003
AAAI8
2026 HCAttention: Extreme KV cache compression via heterogeneous attention computing for LLMs
Dongquan Yang, Xiaotian Yu, Xianbiao Qi, Rong Xiao 0003
Neurocomputing5
2026 SA-BCT: Self-Adapting Backward-Compatible Training
abstract
Backward-compatible training enables the deployment of advanced models without requiring updates to old gallery databases. However, existing methods, including old-prototype-based (i.e., those relying on prototypes from the old model) and instance-based approaches, often overlook the impact of the old model's quality. High-quality old models exhibit compact intra-class feature distributions, which facilitate effective alignment between old and new models across various methods. In contrast, low-quality old models produce dispersed features, making it difficult for old-prototype-based methods to extract sufficient information. Additionally, instance-based methods are overly restrictive, limiting the flexibility of new models. In this work, we propose SA-BCT, an extremely simple yet effective backward-compatible training method that offers a unified framework for accommodating old models of varying quality. SA-BCT employs a single loss function applied to both old and new features, self-adaptively adjusting the constraint space for new features based on the distribution of old features. Extensive experiments in diverse settings demonstrate the effectiveness of SA-BCT. Code is available athttps://github.com/yuleung/SA-BCT.
Yufeng Zhang 0001, Shiliang Zhang, Sheng Xiao, Rong Xiao 0003, Xiaoyu Wang 0002, Kenli Li 0001
IEEE Trans. Multim.5
2025 CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and Compatibility
abstract
Video inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited number of text-guided video inpainting techniques, and these techniques struggle with maintaining visual quality and exhibit poor semantic representation capabilities. In this paper, we introduce CoCoCo, a text-guided video inpainting diffusion framework. To address the aforementioned challenges, we enhance both the training data and model structure. Specifically, we devise an instance-aware region selection strategy for masked area sampling and develop a novel motion block that incorporates efficient 3D full attention and textual cross attention. Additionally, our CoCoCo framework can be seamlessly integrated with various personalized text-to-image diffusion models through a delicate training-free transfer mechanism. Comprehensive experiments demonstrate that CoCoCo can create high-quality visual content with enhanced temporal consistency, improved text controllability, and better compatibility with personalized image models.
Bojia Zi, Xianbiao Qi, Yukai Shi, Bin Liang 0004, Rong Xiao 0003, Kam-Fai Wong, Lei Zhang 0001
AAAI8
2025 BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation Capabilities
abstract
We introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same framework. BiGR features a binary tokenizer, a masked modeling mechanism, and a binary transcoder for binary code prediction. Additionally, we introduce a novel entropy-ordered sampling method to enable efficient image generation. Extensive experiments validate BiGR's superior performance in generation quality, as measured by FID-50k, and representation capabilities, as evidenced by linear-probe accuracy. Moreover, BiGR showcases zero-shot generalization across various vision tasks, enabling applications such as image inpainting, outpainting, editing, interpolation, and enrichment, without the need for structural modifications. Our findings suggest that BiGR unifies generative and discriminative tasks effectively, paving the way for further advancements in the field. We further enable BiGR to perform text-to-image generation, showcasing its potential for broader applications.
Shaozhe Hao, Xuantong Liu, Xianbiao Qi, Bojia Zi, Rong Xiao 0003, Kai Han 0001, Kwan-Yee Kenneth Wong
ICLR6
2025 Exploring a Principled Framework for Deep Subspace Clustering
abstract
Subspace clustering is a classical unsupervised learning task, built on a basic assumption that high-dimensional data can be approximated by a union of subspaces (UoS). Nevertheless, the real-world data are often deviating from the UoS assumption. To address this challenge, state-of-the-art deep subspace clustering algorithms attempt to jointly learn UoS representations and self-expressive coefficients. However, the general framework of the existing algorithms suffers from feature collapse and lacks a theoretical guarantee to learn desired UoS representation. In this paper, we present a Principled fRamewOrk for Deep Subspace Clustering (PRO-DSC), which is designed to learn structured representations and self-expressive coefficients in a unified manner. Specifically, in PRO-DSC, we incorporate an effective regularization on the learned representations into the self-expressive model, prove that the regularized self-expressive model is able to prevent feature space collapse, and demonstrate that the learned optimal representations under certain condition lie on a union of orthogonal subspaces. Moreover, we provide a scalable and efficient approach to implement our PRO-DSC and conduct extensive experiments to verify our theoretical findings and demonstrate the superior performance of our proposed deep subspace clustering approach.
Xianghan Meng, Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li
ICLR5
2025 Taming Transformer Without Using Learning Rate Warmup
abstract
Scaling Transformer to a large scale without using some technical tricks such as learning rate warump and an obviously lower learning rate, is an extremely challenging task, and is increasingly gaining more attention. In this paper, we provide a theoretical analysis for the process of training Transformer and reveal a key problem behind model crash phenomenon in the training process, termed *spectral energy concentration* of ${W_q}^{\top} W_k$, which is the reason for a malignant entropy collapse, where ${W_q}$ and $W_k$ are the projection matrices for the query and the key in Transformer, respectively. To remedy this problem, motivated by *Weyl's Inequality*, we present a novel optimization strategy, \ie, making the weight updating in successive steps steady---if the ratio $\frac{\sigma_{1}(\nabla W_t)}{\sigma_{1}(W_{t-1})}$ is larger than a threshold, we will automatically bound the learning rate to a weighted multiple of $\frac{\sigma_{1}(W_{t-1})}{\sigma_{1}(\nabla W_t)}$, where $\nabla W_t$ is the updating quantity in step $t$. Such an optimization strategy can prevent spectral energy concentration to only a few directions, and thus can avoid malignant entropy collapse which will trigger the model crash. We conduct extensive experiments using ViT, Swin-Transformer and GPT, showing that our optimization strategy can effectively and stably train these (Transformer) models without using learning rate warmup.
Xianbiao Qi, Yelin He, Jiaquan Ye, Chun-Guang Li, Bojia Zi, Xili Dai, Qin Zou 0001, Rong Xiao 0003
ICLR8
2025 Elucidating the design space of language models for image generation
abstract
The success of large language models (LLMs) in text generation has inspired their application to image generation. However, existing methods either rely on specialized designs with inductive biases or adopt LLMs without fully exploring their potential in vision tasks. In this work, we systematically investigate the design space of LLMs for image generation and demonstrate that LLMs can achieve near state-of-the-art performance without domain-specific designs, simply by making proper choices in tokenization methods, modeling approaches, scan patterns, vocabulary design, and sampling strategies. We further analyze autoregressive models' learning and scaling behavior, revealing how larger models effectively capture more useful information than the smaller ones. Additionally, we explore the inherent differences between text and image modalities, highlighting the potential of LLMs across domains. The exploration provides valuable insights to inspire more effective designs when applying LLMs to other domains. With extensive experiments, our proposed model, **ELM** achieves an FID of 1.54 on 256$\times$256 ImageNet and an FID of 3.29 on 512$\times$512 ImageNet, demonstrating the powerful generative potential of LLMs in vision tasks.
Xuantong Liu, Shaozhe Hao, Xianbiao Qi, Tianyang Hu 0001, Jun Wang 0123, Rong Xiao 0003, Yuan Yao 0011
ICML6
2025 MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
abstract
Recent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated objects and visual artifacts. Furthermore, existing methods often rely on computationally expensive sampling procedures and classifier-free guidance (CFG), resulting in slow inference. To address these limitations, we propose **MiniMax-Remover**, a novel two-stage video object removal approach. Motivated by the observation that text condition is not best suited for this task, we simplify the pretrained video generation model by removing textual input and cross-attention layers. In this way, we obtain a more lightweight and efficient model architecture in the first stage. In the second stage, we proposed a minimax optimization strategy to further distill the remover with the successful videos produced by stage-1 model. Specifically, the inner maximization identifies adversarial input noise ("bad noise'') that leads to failure removals, while the outer minimization trains the model to generate high-quality removal results even under such challenging conditions. As a result, our method achieves a state-of-the-art video object removal results using as few as 6 sampling steps without CFG usage. Extensive experiments demonstrate the effectiveness and superiority of MiniMax-Remover compared to existing methods. Codes and Videos are available at: **https://minimax-remover.github.io**.
Bojia Zi, Weixuan Peng, Xianbiao Qi, Rong Xiao 0003, Kam-Fai Wong
NeurIPS6
2025 Señorita-2M: A High-Quality Instruction-based Dataset for General Video Editing by Video Specialists
abstract
Video content editing has a wide range of applications. With the advancement of diffusion-based generative models, video editing techniques have made remarkable progress, yet they still remain far from practical usability. Existing inversion-based video editing methods are time-consuming and struggle to maintain consistency in unedited regions. Although instruction-based methods have high theoretical potential, they face significant challenges in constructing high-quality training datasets - current datasets suffer from issues such as editing correctness, frame consistency, and sample diversity. To bridge these gaps, we introduce the Señorita-2M dataset, a large-scale, diverse, and high-quality video editing dataset. We systematically categorize editing tasks into 2 classes consisting of 18 subcategories. To build this dataset, we design four new task specialists and employ or modify 14 existing task experts to generate data samples for each subclass. In addition, we design a filtering pipeline at both the visual content and instruction levels to further enhance data quality. This approach ensures the reliability of constructed data. Finally, the Señorita-2M dataset comprises 2 million high-fidelity samples with diverse resolutions and frame counts. We trained multiple models using different base video models, i.e., Wan2.1 and CogVideoX-5B, on Señorita-2M, and the results demonstrate that the models exhibit superior visual quality, robust frame-to-frame consistency, and strong instruction following capability. More videos are available at: https://senorita-2m-dataset.github.io.
Bojia Zi, Penghui Ruan, Marco Chen, Xianbiao Qi, Shaozhe Hao, Youze Huang, Bin Liang 0004, Rong Xiao 0003, Kam-Fai Wong
NeurIPS9
2025 BiTA: Bi-directional tuning for lossless acceleration in large language models
Feng Lin 0009, Hanling Yi, Xiaotian Yu, Guangming Lu 0002, Rong Xiao 0003
Expert Syst. Appl.7
2025 Neural Normalized Cut: A differential and generalizable approach for spectral clustering
Shangzhi Zhang, Chun-Guang Li, Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002
Pattern Recognit.5
2024 Graph Cut-Guided Maximal Coding Rate Reduction for Learning Image Embedding and Clustering
Xianghan Meng, Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li
ACCV (10)5
2023 NAR-Former V2: Rethinking Transformer for Universal Neural Network Representation Learning
abstract
As more deep learning models are being applied in real-world applications, there is a growing need for modeling and learning the representations of neural networks themselves. An effective representation can be used to predict target attributes of networks without the need for actual training and deployment procedures, facilitating efficient network design and deployment. Recently, inspired by the success of Transformer, some Transformer-based representation learning frameworks have been proposed and achieved promising performance in handling cell-structured models. However, graph neural network (GNN) based approaches still dominate the field of learning representation for the entire network. In this paper, we revisit the Transformer and compare it with GNN to analyze their different architectural characteristics. We then propose a modified Transformer-based universal neural network representation learning model NAR-Former V2. It can learn efficient representations from both cell-structured networks and entire networks. Specifically, we first take the network as a graph and design a straightforward tokenizer to encode the network into a sequence. Then, we incorporate the inductive representation learning capability of GNN into Transformer, enabling Transformer to generalize better when encountering unseen architecture. Additionally, we introduce a series of simple yet effective modifications to enhance the ability of the Transformer in learning representation from graph structures. In encoding entire networks and then predicting the latency, our proposed method surpasses the GNN-based method NNLP by a significant margin on the NNLQP dataset. Furthermore, regarding accuracy prediction on the cell-structured NASBench101 and NASBench201 datasets, our method achieves highly comparable performance to other state-of-the-art methods. The code is available at https://github.com/yuny220/NAR-Former-V2.
Yun Yi, Haokui Zhang, Rong Xiao 0003, Nannan Wang 0001, Xiaoyu Wang 0002
NeurIPS3
2022 Learning graph normalization for graph neural networks
Xianbiao Qi, Chun-Guang Li, Rong Xiao 0003
Neurocomputing5
2022 EMU: Effective Multi-Hot Encoding Net for Lightweight Scene Text Recognition With a Large Character Set
abstract
Deploying a lightweight deep model for scene text recognition task on mobile devices has great commercial value. However, the conventional softmax-based one-hot classification module becomes a cumbersome obstacle when handling multi-languages or languages with large character set (e.g., Chinese) due to the rapid expansion of model parameters with the number of classes. To this end, we propose an Effective Multi-hot encoding and classification modUle (EMU) for scene text recognition in the scenario of multi-languages or languages with large character set. Specifically, EMU generates a binary multi-hot label for each class with a real-valued sub-network in training stage and produces the prediction by calculating the inner product between the multi-hot code and the multi-hot label. Compared to the softmax-based one-hot classifier, EMU reduces the storage requirement and the time cost in inference stage significantly, retaining similar performance. Furthermore, we design a convolution feature basedLightweight TransFormerto learn the effective features for EMU and consequently develop a lightweight scene text recognition framework, termedLight-Former-EMU. We conduct extensive experiments on seven public English benchmarks and two real-world Chinese challenge benchmarks. Experimental results verify the effectiveness of the proposed EMU and demonstrate the promising performance of the proposed Light-Former-EMU.
Bingcong Li, Xianbiao Qi, Chun-Guang Li, Rong Xiao 0003
IEEE Trans. Circuits Syst. Video Technol.6
2021 MASTER: Multi-aspect non-local network for scene text recognition
Ning Lu 0003, Wenwen Yu, Xianbiao Qi, Ping Gong 0003, Rong Xiao 0003, Xiang Bai
Pattern Recognit.6
2020 PICK: Processing Key Information Extraction from Documents using Improved Graph Learning-Convolutional Networks
abstract
Computer vision with state-of-the-art deep learning models has achieved huge success in the field of Optical Character Recognition (OCR) including text detection and recognition tasks recently. However, Key Information Extraction (KIE) from documents as the downstream task of OCR, having a large number of use scenarios in real-world, remains a challenge because documents not only have textual features extracting from OCR systems but also have semantic visual features that are not fully exploited and play a critical role in KIE. Too little work has been devoted to efficiently make full use of both textual and visual features of the documents. In this paper, we introduce PICK, a framework that is effective and robust in handling complex documents layout for KIE by combining graph learning with graph convolution operation, yielding a richer semantic representation containing the textual and visual features and global layout without ambiguity. Extensive experiments on realworld datasets have been conducted to show that our method outperforms baselines methods by significant margins. Our code is available at https://github.com/wenwenyu/PICK-pytorch.
Wenwen Yu, Ning Lu 0003, Xianbiao Qi, Ping Gong 0003, Rong Xiao 0003
ICPR5
2019 High Frequency Residual Learning for Multi-Scale Image Classification
Bowen Cheng, Rong Xiao 0003, Thomas S. Huang, Lei Zhang 0001
BMVC2
2014 Pairwise Rotation Invariant Co-Occurrence Local Binary Pattern
abstract
Designing effective features is a fundamental problem in computer vision. However, it is usually difficult to achieve a great tradeoff between discriminative power and robustness. Previous works shown that spatial co-occurrence can boost the discriminative power of features. However the current existing co-occurrence features are taking few considerations to the robustness and hence suffering from sensitivity to geometric and photometric variations. In this work, we study the Transform Invariance (TI) of co-occurrence features. Concretely we formally introduce a Pairwise Transform Invariance (PTI) principle, and then propose a novel Pairwise Rotation Invariant Co-occurrence Local Binary Pattern (PRICoLBP) feature, and further extend it to incorporate multi-scale, multi-orientation, and multi-channel information. Different from other LBP variants, PRICoLBP can not only capture the spatial context co-occurrence information effectively, but also possess rotation invariance. We evaluate PRICoLBP comprehensively on nine benchmark data sets from five different perspectives, e.g., encoding strategy, rotation invariance, the number of templates, speed, and discriminative power compared to other LBP variants. Furthermore we apply PRICoLBP to six different but related applications-texture, material, flower, leaf, food, and scene classification, and demonstrate that PRICoLBP is efficient, effective, and of a well-balanced tradeoff between the discriminative power and robustness.
Xianbiao Qi, Rong Xiao 0003, Chun-Guang Li, Yu Qiao 0001, Jun Guo 0002, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2012 Pairwise Rotation Invariant Co-occurrence Local Binary Pattern
Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002, Lei Zhang 0001
ECCV (6)2
2012 A rapid flower/leaf recognition system
abstract
In this work, we introduce a rapid and accurate flower/leaf recognition system. The system could process one query in less than 0.35s with users' simple interaction. Meanwhile, high accuracy and recall is achieved. Furthermore, low computational resource and memory cost are required by the system. Now, the system is demonstrated on 172 categories of flowers, the largest flower dataset until now, and 220 categories of leaves.
Xianbiao Qi, Rong Xiao 0003, Lei Zhang 0001, Chun-Guang Li, Jun Guo 0002
ACM Multimedia2
2011 Rank-SIFT: Learning to rank repeatable local interest points
abstract
Scale-invariant feature transform (SIFT) has been well studied in recent years. Most related research efforts focused on designing and learning effective descriptors to characterize a local interest point. However, how to identify stable local interest points is still a very challenging problem. In this paper, we propose a set of differential features, and based on them we adopt a data-driven approach to learn a ranking function to sort local interest points according to their stabilities across images containing the same visual objects. Compared with the handcrafted rule-based method used by the standard SIFT algorithm, our algorithm substantially improves the stability of detected local interest point on a very challenging benchmark dataset, in which images were generated under very different imaging conditions. Experimental results on the Oxford and PASCAL databases further demonstrate the superior performance of the proposed algorithm on both object image retrieval and category recognition.
Rong Xiao 0003, Zhiwei Li 0006, Rui Cai 0002, Bao-Liang Lu, Lei Zhang 0001
CVPR2
2011 On theme location discovery for travelogue services
abstract
In this paper, we aim to develop a travelogue service that discovers and conveys various travelogue digests, in form of theme locations, geographical scope, traveling trajectory and location snippet, to users. In this service, theme locations in a travelogue are the core information to discover. Thus we aim to address the problem of theme location discovery to enable the above travelogue services. Due to the inherent ambiguity of location relevance, we perform location relevance mining (LRM) in two complementary angles, relevance classification and relevance ranking, to provide comprehensive understanding of locations. Furthermore, we explore the textual (e.g., surrounding words) and geographical (e.g., geographical relationship among locations) features of locations to develop a co-training model for enhancement of classification performance. Built upon the mining result of LRM, we develop a series of techniques for provisioning of the aforementioned travelogue digests in our travelogue system. Finally, we conduct comprehensive experiments on collected travelogues to evaluate the performance of our location relevance mining techniques and demonstrate the effectiveness of the travelogue service.
Mao Ye 0002, Rong Xiao 0003, Wang-Chien Lee, Xing Xie 0001
SIGIR2
2010 An efficient location extraction algorithm by leveraging web contextual information
abstract
A typical location extraction approach consists of two steps, location name detection and location entity disambiguation. Promising results have been obtained in the last decade based on natural language processing technologies. However, there are still two challenges which requires further investigation: 1)How to leverage the prior and contextual evidence to improve the location extraction performance, and 2) How to utilize the interdependence information between the named entity recognition step and disambiguation step. In this paper, we propose an iterative detection-ranking framework to address these problems as well as a set of novel features to mine contextual information from web resources. Experimental results show that our solution outperforms the state-of-the-art approaches, including Metacarta GeoTagger and Yahoo Placemaker.
Teng Qin, Rong Xiao 0003, Lei Fang 0004, Xing Xie 0001, Lei Zhang 0001
GIS2
2010 Equip tourists with knowledge mined from travelogues
abstract
With the prosperity of tourism and Web 2.0 technologies, more and more people have willingness to share their travel experiences on the Web (e.g., weblogs, forums, or Web 2.0 communities). These so-called travelogues contain rich information, particularly including location-representative knowledge such as attractions (e.g., Golden Gate Bridge), styles (e.g., beach, history), and activities (e.g., diving, surfing). The location-representative information in travelogues can greatly facilitate other tourists' trip planning, if it can be correctly extracted and summarized. However, since most travelogues are unstructured and contain much noise, it is difficult for common users to utilize such knowledge effectively. In this paper, to mine location-representative knowledge from a large collection of travelogues, we propose a probabilistic topic model, named as Location-Topic model. This model has the advantages of (1) differentiability between two kinds of topics, i.e., local topics which characterize locations and global topics which represent other common themes shared by various locations, and (2) representation of locations in the local topic space to encode both location-representative knowledge and similarities between locations. Some novel applications are developed based on the proposed model, including (1) destination recommendation for on flexible queries, (2) characteristic summarization for a given destination with representative tags and snippets, and (3) identification of informative parts of a travelogue and enriching such highlights with related images. Based on a large collection of travelogues, the proposed framework is evaluated using both objective and subjective evaluation methods and shows promising results.
Rui Cai 0002, Changhu Wang, Rong Xiao 0003, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001
WWW4
2009 TravelScope: standing on the shoulders of dedicated travelers
abstract
In this paper, we propose a system called TravelScope that helps users experience virtual tours by presenting information mined from user-generated travelogues and photos. The system can (1) recommend popular places for a given region; (2) characterize comprehensive aspects (e.g., landmarks, styles, activities) for a location; and (3) show representative images for a landmark. A novel user interface is designed to provide a better user experience, by organizing both the textual and visual information generated by dedicated travelers in an attractive way.
Rui Cai 0002, Jiang-Ming Yang, Rong Xiao 0003, Like Liu, Lei Zhang 0001
ACM Multimedia4
2008 Face Alignment Via Component-Based Discriminative Search
Rong Xiao 0003, Fang Wen 0001, Jian Sun 0001
ECCV (2)2
2008 3D Face Recognition by Local Shape Difference Boosting
Yueming Wang 0001, Xiaoou Tang, Jianzhuang Liu, Gang Pan 0001, Rong Xiao 0003
ECCV (1)5
2007 EasyAlbum: an interactive photo annotation system based on face clustering and re-ranking
abstract
Digital photo management is becoming indispensable for the explosively growing family photo albums due to the rapid popularization of digital cameras and mobile phone cameras. In an effective photo management system photo annotation is the most challenging task. In this paper, we develop several innovative interaction techniques for semi-automatic photo annotation. Compared with traditional annotation systems, our approach provides the following new features: "cluster annotation" puts similar faces or photos with similar scene together, and enables user label them in one operation; "contextual re-ranking" boosts the labeling productivity by guessing the user intention; "ad hoc annotation" allows user label photos while they are browsing or searching, and improves system performance progressively through learning propagation. Our results show that these technologies provide a more user friendly interface for the annotation of person name, location, and event, and thus substantially improve the annotation performance especially for a large photo album.
Jingyu Cui, Fang Wen 0001, Rong Xiao 0003, Yuandong Tian, Xiaoou Tang
CHI3
2007 A Face Annotation Framework with Partial Clustering and Interactive Labeling
abstract
Face annotation technology is important for a photo management system. In this paper, we propose a novel interactive face annotation framework combining unsupervised and interactive learning. There are two main contributions in our framework. In the unsupervised stage, a partial clustering algorithm is proposed to find the most evident clusters instead of grouping all instances into clusters, which leads to a good initial labeling for later user interaction. In the interactive stage, an efficient labeling procedure based on minimization of both global system uncertainty and estimated number of user operations is proposed to reduce user interaction as much as possible. Experimental results show that the proposed annotation framework can significantly reduce the face annotation workload and is superior to existing solutions in the literature.
Yuandong Tian, Wei Liu 0026, Rong Xiao 0003, Fang Wen 0001, Xiaoou Tang
CVPR3
2007 Linear Laplacian Discrimination for Feature Extraction
abstract
Discriminant feature extraction plays a fundamental role in pattern recognition. In this paper, we propose the linear Laplacian discrimination (LLD) algorithm/or discriminant feature extraction. LLD is an extension of linear discriminant analysis (LDA). Our motivation is to address the issue that LDA cannot work well in cases where sample spaces are non-Euclidean. Specifically, we define the within-class scatter and the between-class scatter using similarities which are based on pairwise distances in sample spaces. Thus the structural information of classes is contained in the within-class and the between-class Laplacian matrices which are free from metrics of sample spaces. The optimal discriminant subspace can be derived by controlling the structural evolution of Laplacian matrices. Experiments are performed on the facial database for FRGC version 2. Experimental results show that LLD is effective in extracting discriminant features.
Deli Zhao, Zhouchen Lin, Rong Xiao 0003, Xiaoou Tang
CVPR3
2007 Dynamic Cascades for Face Detection
abstract
In this paper, we propose a novel method, called "dynamic cascade", for training an efficient face detector on massive data sets. There are three key contributions. The first is a new cascade algorithm called "dynamic cascade ", which can train cascade classifiers on massive data sets and only requires a small number of training parameters. The second is the introduction of a new kind of weak classifier, called "Bayesian stump", for training boost classifiers. It produces more stable boost classifiers with fewer features. Moreover, we propose a strategy for using our dynamic cascade algorithm with multiple sets of features to further improve the detection performance without significant increase in the detector's computational cost. Experimental results show that all the new techniques effectively improve the detection performance. Finally, we provide the first large standard data set for face detection, so that future researches on the topic can be compared on the same training and testing set.
Rong Xiao 0003, Huaiyi Zhu, He Sun 0003, Xiaoou Tang
ICCV1
2006 Joint Boosting Feature Selection for Robust Face Recognition
abstract
A fundamental challenge in face recognition lies in determining what facial features are important for the identification of faces. In this paper, a novel face recognition framework is proposed to address this problem. In our framework, 3D face models are used to synthesize a huge database of realistic face images which covers wide appearance variations of faces due to various pose, illumination, and expression changes. A novel feature selection algorithm which we call Joint Boosting is developed to extract discriminative face features using this massive database. The major contributions of this paper are: (1) With the help of 3D face models, a massive database of realistic virtual face images is generated to achieve robust feature selection; (2)Because the huge database covers a wide range of face variations, our feature selection procedure only needs to be trained once, and the selected feature set can be generalized to other face database without re-training; (3) We propose a new learning algorithm, Joint Boosting Algorithm, which is effective and efficient in learning directly from a massive database without having to convert face images to intra-personal and extra-personal difference images. This property is important for applying our algorithm to other general pattern recognition problems. Experimental results show that our method significantly improves recognition performance.
Rong Xiao 0003, Wu-Jun Li, Yuandong Tian, Xiaoou Tang
CVPR (2)1