VLDB 2026 Research / reviewers in the wild / expert
Yijuan Lu
dblp:30/6535
· DBLP profile ↗
63ranked-venue papers
12as first author
9since 2021 · last 2025
0000-0002-9855-8365ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 43 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 21 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorComputer networks · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
12 papers |
Vision and language · 42% Image recognition and object detection · 16% Language models and text generation · 12% | |
| Databases, data mining, and information retrieval
11 papers |
Information retrieval · 91% Machine learning and data management · 7% Data mining · 2% | |
| Computer graphics and multimedia
5 papers |
Multimedia analysis and retrieval · 68% Geometric modeling and processing · 17% Image and video coding · 15% |
Topics — the 30 heaviest of 63, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
image retrieval |
1.1 | 7 | 2018 | Cascaded Feature Augmentation with Diffusion for Image Retrieval · ACM Multimedia 2018 BSIFT: Toward Data-Independent Codebook for Large Scale Image Search · IEEE Trans. Image Process. 2015 Query Difficulty Prediction for Web Image Search · IEEE Trans. Multim. 2012 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language › visual reasoning
visual chain-of-thought |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Vision and language
visual reasoning |
0.9 | 1 | 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding · ICML 2025 |
Computer vision › Image recognition and object detection › text recognition
optical character recognition |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Computer vision › Image recognition and object detection
text recognition |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Natural language and speech › Language models and text generation › text generation › neural text generation
transformer-based text generation |
0.7 | 1 | 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained Models · AAAI 2023 |
Information retrieval › image retrieval
large-scale image retrieval |
0.6 | 4 | 2018 | BSIFT: Toward Data-Independent Codebook for Large Scale Image Search · IEEE Trans. Image Process. 2015 Scalar quantization for large scale image search · ACM Multimedia 2012 Large scale image search with geometric coding · ACM Multimedia 2011 |
Computer vision › Video understanding and tracking › temporal localization
moment localization |
0.6 | 1 | 2022 | Multi-Scale 2D Temporal Adjacency Networks for Moment Localization With Natural Language · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Vision and language
natural language moment localization |
0.6 | 1 | 2022 | Multi-Scale 2D Temporal Adjacency Networks for Moment Localization With Natural Language · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Vision and language › visual grounding
video-language grounding |
0.6 | 1 | 2022 | Multi-Scale 2D Temporal Adjacency Networks for Moment Localization With Natural Language · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Natural language and speech › Information extraction and text analysis
document AI |
0.5 | 1 | 2021 | DI-2021: The Second Document Intelligence Workshop · KDD 2021 |
Natural language and speech › Information extraction and text analysis
document understanding |
0.5 | 1 | 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document Understanding · ACL/IJCNLP (1) 2021 |
Computer vision › Vision and language › visual question answering
text-based visual question answering |
0.5 | 1 | 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption · CVPR 2021 |
Computer vision › Vision and language
vision-language pretraining |
0.5 | 1 | 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-Caption · CVPR 2021 |
Information retrieval › retrieval models › neural retrieval
diffusion-based retrieval |
0.3 | 1 | 2018 | Cascaded Feature Augmentation with Diffusion for Image Retrieval · ACM Multimedia 2018 |
Machine learning and data management › feature transformation
feature augmentation |
0.3 | 1 | 2018 | Cascaded Feature Augmentation with Diffusion for Image Retrieval · ACM Multimedia 2018 |
Information retrieval
reranking |
0.3 | 1 | 2018 | Cascaded Feature Augmentation with Diffusion for Image Retrieval · ACM Multimedia 2018 |
Computer vision › Video understanding and tracking
object tracking |
0.3 | 1 | 2017 | Weighted Sparse Representation Regularized Graph Learning for RGB-T Object Tracking · ACM Multimedia 2017 |
Computer vision › Video understanding and tracking › object tracking › multi-modal tracking
RGBT tracking |
0.3 | 1 | 2017 | Weighted Sparse Representation Regularized Graph Learning for RGB-T Object Tracking · ACM Multimedia 2017 |
Information retrieval › image retrieval › instance retrieval
partial-duplicate image search |
0.3 | 2 | 2012 | Principal Visual Word Discovery for Automatic License Plate Detection · IEEE Trans. Image Process. 2012 Large scale image search with geometric coding · ACM Multimedia 2011 |
Computer vision › 3D vision
shape matching |
0.2 | 1 | 2016 | Shape Retrieval of Non-rigid 3D Human Models · Int. J. Comput. Vis. 2016 |
Geometric modeling and processing
shape analysis |
0.2 | 1 | 2016 | Shape Retrieval of Non-rigid 3D Human Models · Int. J. Comput. Vis. 2016 |
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
shape retrieval |
0.2 | 1 | 2016 | Shape Retrieval of Non-rigid 3D Human Models · Int. J. Comput. Vis. 2016 |
Multimedia analysis and retrieval
image retrieval |
0.2 | 2 | 2010 | Large scale partially duplicated web image retrieval · ACM Multimedia 2010 Spatial coding for large scale partial-duplicate web image search · ACM Multimedia 2010 |
Multimedia analysis and retrieval › image retrieval › instance-level image retrieval
partial-duplicate image retrieval |
0.2 | 2 | 2010 | Large scale partially duplicated web image retrieval · ACM Multimedia 2010 Spatial coding for large scale partial-duplicate web image search · ACM Multimedia 2010 |
Image and video coding › image compression
spatial coding |
0.2 | 2 | 2010 | Large scale partially duplicated web image retrieval · ACM Multimedia 2010 Spatial coding for large scale partial-duplicate web image search · ACM Multimedia 2010 |
Computer vision › Video understanding and tracking › temporal modeling
temporal context modeling |
0.2 | 1 | 2022 | Multi-Scale 2D Temporal Adjacency Networks for Moment Localization With Natural Language · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Information retrieval
retrieval evaluation |
0.2 | 2 | 2012 | Learning to judge image search results · ACM Multimedia 2011 Query Difficulty Prediction for Web Image Search · IEEE Trans. Multim. 2012 |
Methods — techniques the papers use, named apart from their topics
visual editing · 0.9code generation · 0.9bag-of-words · 0.7pre-trained text transformer · 0.7pre-trained image transformer · 0.7multi-scale temporal adjacency · 0.62d temporal maps · 0.6SIFT · 0.5optical character recognition · 0.5multimodal pretraining · 0.5masked language modeling · 0.5layout-aware transformers · 0.5contrastive matching · 0.5cascaded feature augmentation · 0.3cascaded cluster diffusion · 0.3visual feature analysis · 0.3shape descriptors · 0.2benchmark evaluation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReFocus: Visual Editing as a Chain of Thought for Structured Image UnderstandingabstractStructured image understanding, such as interpreting tables and charts, requires strategically refocusing across various structures and texts within an image, forming a reasoning sequence to arrive at the final answer. However, current multimodal large language models (LLMs) lack this multihop selective attention capability. In this work, we introduce ReFocus, a simple yet effective framework that equips multimodal LLMs with the ability to generate ``visual thoughts'' by performing visual editing on the input image through code, shifting and refining their visual focuses. Specifically, ReFocus enables multimodal LLMs to generate Python codes to call tools and modify the input image, sequentially drawing boxes, highlighting sections, and masking out areas, thereby enhancing the visual reasoning process. We experiment upon a wide range of structured image understanding tasks involving tables and charts. ReFocus largely improves performance on all tasks over GPT-4o without visual editing, yielding an average gain of 11.0% on table tasks and 6.8% on chart tasks. We present an in-depth analysis of the effects of different visual edits, and reasons why ReFocus can improve the performance without introducing additional information. Further, we collect a 14k training set using ReFocus, and prove that such visual chain-of-thought with intermediate information offers a better supervision than standard VQA data, reaching a 8.0% average gain over the same model trained with QA pairs and 2.6% over CoT. Minqian Liu, Zhengyuan Yang, John Corring, Yijuan Lu, Dan Roth 0001, Dinei A. F. Florêncio, Cha Zhang |
ICML | 5 |
| 2023 | TrOCR: Transformer-Based Optical Character Recognition with Pre-trained ModelsabstractText recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-processing step. In this paper, we propose an end-to-end text recognition approach with pre-trained image Transformer and text Transformer models, namely TrOCR, which leverages the Transformer architecture for both image understanding and wordpiece-level text generation. The TrOCR model is simple but effective, and can be pre-trained with large-scale synthetic data and fine-tuned with human-labeled datasets. Experiments show that the TrOCR model outperforms the current state-of-the-art models on the printed, handwritten and scene text recognition tasks. The TrOCR models and code are publicly available at https://aka.ms/trocr. Minghao Li 0004, Tengchao Lv, Jingye Chen, Lei Cui 0001, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Zhoujun Li 0001, Furu Wei |
AAAI | 5 |
| 2023 | Diffusion-Based Document Layout Generation
Yijuan Lu, John Corring, Dinei A. F. Florêncio, Cha Zhang |
ICDAR (1) | 2 |
| 2022 | Unsupervised adversarial image retrieval
Ling Huang 0003, Cong Bai, Yijuan Lu, Shaobo Zhang 0005, Shengyong Chen |
Multim. Syst. | 3 |
| 2022 | Multi-Scale 2D Temporal Adjacency Networks for Moment Localization With Natural LanguageabstractWe address the problem of retrieving a specific moment from an untrimmed video by natural language. It is a challenging problem because a target moment may take place in the context of other temporal moments in the untrimmed video. Existing methods cannot tackle this challenge well since they do not fully consider the temporal contexts between temporal moments. In this paper, we model the temporal context between video moments by a set of predefined two-dimensional maps under different temporal scales. For each map, one dimension indicates the starting time of a moment and the other indicates the duration. These 2D temporal maps can cover diverse video moments with different lengths, while representing their adjacent contexts at different temporal scales. Based on the 2D temporal maps, we propose a Multi-Scale Temporal Adjacency Network (MS-2D-TAN), a single-shot framework for moment localization. It is capable of encoding the adjacent temporal contexts at each scale, while learning discriminative features for matching video moments with referring expressions. We evaluate the proposed MS-2D-TAN on three challenging benchmarks, i.e., Charades-STA, ActivityNet Captions, and TACoS, where our MS-2D-TAN outperforms the state of the art. Songyang Zhang 0004, Houwen Peng, Jianlong Fu, Yijuan Lu, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | LayoutLMv2: Multi-modal Pre-training for Visually-rich Document UnderstandingabstractYang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, Lidong Zhou. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yang Xu 0049, Yiheng Xu, Tengchao Lv, Lei Cui 0001, Furu Wei, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Wanxiang Che, Min Zhang 0005, Lidong Zhou |
ACL/IJCNLP (1) | 7 |
| 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionabstractIn this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to conventional vision-language pretraining that fails to capture scene text and its relationship with the visual and text modalities, TAP explicitly incorporates scene text (generated from OCR engines) during pretraining. With three pre-training tasks, including masked language modeling (MLM), image-text (contrastive) matching (ITM), and relative (spatial) position prediction (RPP), pre-training with scene text effectively helps the model learn a better aligned representation among the three modalities: text word, visual object, and scene text. Due to this aligned representation learning, even pre-trained on the same downstream task dataset, TAP already boosts the absolute accuracy on the TextVQA dataset by +5:4%, compared with a non-TAP baseline. To further improve the performance, we build a large-scale scene text-related imagetext dataset based on the Conceptual Caption dataset, named OCR-CC, which contains 1:4 million images with scene text. Pre-trained on this OCR-CC dataset, our approach outperforms the state of the art by large margins on multiple tasks, i.e., +8:3% accuracy on TextVQA, +8:6% accuracy on ST-VQA, and +10:2 CIDEr score on TextCaps. Zhengyuan Yang, Yijuan Lu, Xi Yin 0006, Dinei A. F. Florêncio, Cha Zhang, Lei Zhang 0001, Jiebo Luo 0001 |
CVPR | 2 |
| 2021 | DI-2021: The Second Document Intelligence WorkshopabstractBusiness documents are central to the operation of all organizations, and they come in all shapes and sizes: project reports, planning documents, technical specifications, financial statements, meeting minutes, legal agreements, contracts, resumes, purchase orders, invoices, and many more. The ability to read, understand and interpret these documents, referred to here as Document Intelligence (DI), is challenging due to not only many domains of knowledge involved, but also their complex formats and structures, internal and external cross references deployed, and even less-than-ideal quality of scans and OCR oftentimes performed on them. This workshop aims to explore and advance the current state of research and practice in answering these challenges. Benjamin Han, Douglas Burdick, Dave Lewis 0003, Yijuan Lu, Hamid R. Motahari Nezhad, Sandeep Tata |
KDD | 4 |
| 2021 | 3D sketching for 3D object retrieval
Bo Li 0013, Juefei Yuan, Yuxiang Ye, Yijuan Lu, Qi Tian 0001 |
Multim. Tools Appl. | 4 |
| 2020 | A comparison of methods for 3D scene shape retrieval
Juefei Yuan, Hameed Abdul-Rashid, Bo Li 0013, Yijuan Lu, Tobias Schreck, Song Bai 0001, Xiang Bai, Ngoc-Minh Bui, Minh N. Do, Trong-Le Do, Anh Duc Duong, Xinwei He 0001, Mike Holenderski, Dmitri Jarnikov, Tu-Khiem Le, Wenhui Li 0001, Anan Liu |
Comput. Vis. Image Underst. | 4 |
| 2019 | A sketch recognition method based on transfer deep learning with the fusion of multi-granular sketches
Peng Zhao 0010, Yijuan Lu, Benpeng Xu |
Multim. Tools Appl. | 3 |
| 2019 | RGB-T object tracking: Benchmark and baseline
Chenglong Li 0002, Xinyan Liang, Yijuan Lu, Jin Tang 0001 |
Pattern Recognit. | 3 |
| 2018 | Cascaded Feature Augmentation with Diffusion for Image RetrievalabstractRecently, as an effective re-ranking technique, diffusion has attracted considerable attention in research on image retrieval. It inherits from random surfer model and is effective to deeply explore data manifold structure. However, as a common practice, diffusion is performed at query time which relies heavily on initial retrieval shortlists and suffers the bottleneck of online time-efficiency. To this end, in this paper, we present a more generalized method named CFA (cascaded feature augmentation) based on diffusion. First of all, we transfer diffusion process from online stage to offline stage and innovatively utilize output of diffusion to augment database features in a cascaded mode, which can eliminate iteration process at query time radically. Second, to scale the diffusion method to large image database, we propose a cascaded cluster diffusion technique for feature augmentation which largely reduces computational cost. Third, we extend our cascaded feature augmentation scheme to cases with multiple features without involving extra memory and time cost. Our CFA is compatible with other re-ranking methods. Extensive experiments on four public datasets demonstrate the effectiveness of our proposed algorithm. Yuanqiang Fang, Wengang Zhou 0001, Yijuan Lu, Jinhui Tang 0001, Qi Tian 0001, Houqiang Li |
ACM Multimedia | 3 |
| 2018 | Session details: System-2 (Smart Multimedia Systems)
Yijuan Lu |
ACM Multimedia | 1 |
| 2018 | Transfer robust sparse coding based on graph and joint distribution adaption for image representation
Peng Zhao 0010, Yijuan Lu, Huiting Liu 0001, Sheng Yao 0001 |
Knowl. Based Syst. | 3 |
| 2017 | Weighted Sparse Representation Regularized Graph Learning for RGB-T Object TrackingabstractIn this paper, we propose a novel graph model, called weighted sparse representation regularized graph, to learn a robust object representation using multispectral (RGB and thermal) data for visual tracking. In particular, the tracked object is represented with a graph with image patches as nodes. This graph is dynamically learned from two aspects. First, the graph affinity (i.e., graph structure and edge weights) that indicates the appearance compatibility of two neighboring nodes is optimized based on the weighted sparse representation, in which the modality weight is introduced to leverage RGB and thermal information adaptively. Second, each node weight that indicates how likely it belongs to the foreground is propagated from others along with graph affinity. The optimized patch weights are then imposed on the extracted RGB and thermal features, and the target object is finally located by adopting the structured SVM algorithm. Moreover, we also contribute a comprehensive dataset for RGB-T tracking purpose. Comparing with existing ones, the new dataset has the following advantages: 1) Its size is sufficiently large for large-scale performance evaluation (total frame number: 210K, maximum frames per video pair: 8K). 2) The alignment between RGB-T video pairs is highly accurate, which does not need pre- and post-processing. 3) The occlusion levels are annotated for analyzing the occlusion-sensitive performance of different methods. Extensive experiments on both public and newly created datasets demonstrate the effectiveness of the proposed tracker against several state-of-the-art tracking methods. Chenglong Li 0002, Yijuan Lu, Chengli Zhu, Jin Tang 0001 |
ACM Multimedia | 3 |
| 2017 | Sketch-based 3D model retrieval utilizing adaptive view clustering and semantic information
Bo Li 0013, Yijuan Lu, Henry Johan, Ribel Fares |
Multim. Tools Appl. | 2 |
| 2016 | A semantic tree-based approach for sketch-based 3D model retrievalabstractSketch-based 3D model retrieval is to retrieve 3D models given a user's hand-drawn sketch. Due to the big semantic gap between rough sketch representation and accurate 3D model coordinates, sketch-based 3D model retrieval (SBR) is one of the most challenging research topics in the field of 3D model retrieval. To bridge the semantic gap, a novel semantic tree-based SBR algorithm is proposed in this paper. Given a 2D sketch query and a collection of 3D models, a 3D semantic tree is built up first based on the ontology structure of WordNet. Every leaf node in the tree contains a set of 3D models assigned to this class according to their semantic classification/label information. Then, sketch components of the 2D query sketch are identified by sketch segmentation and annotation. Finally, by measuring the semantic relatedness between the sketch components' annotations and tree nodes in the 3D semantic tree, the similarities between the 2D sketch and 3D models are computed to find out the most relevant 3D models. Experimental results demonstrate the effectiveness and promising potentials of our approach on sketch-based 3D model retrieval. Bo Li 0013, Yijuan Lu |
ICPR | 2 |
| 2016 | 3D sketch-based 3D model retrieval with convolutional neural networkabstract3D sketch-based 3D model retrieval is to retrieve similar 3D models using users' hand-drawn 3D sketches as input. Compared with traditional 2D sketch-based retrieval, 3D sketch-based 3D model retrieval is a brand new and challenging research topic. In this paper, we employ advanced deep learning method and propose a novel 3D sketch based 3D model retrieval system. Our system has been comprehensively tested on two benchmark datasets and compared with other existing 3D model retrieval algorithms. The experimental results reveal our approach outperforms other competing state-of-the-arts and demonstrate promising potential of our approach on 3D sketch based applications. Yuxiang Ye, Bo Li 0013, Yijuan Lu |
ICPR | 3 |
| 2016 | Human's Scene Sketch UnderstandingabstractHuman's sketch understanding is important. It has many applications in human computer interaction, multimedia, and computer vision. Recognizing human sketches is also challenging. Previous methods focus on single-object sketch recognition. Understanding human's scene sketch that involves multiple objects and their complex interactions has not been explored. In this paper, we tackle this new problem. We create the first scene sketch dataset "Scene250" and propose a deep learning method to understand human scene sketches. We propose "Scene-Net", a new deep convolutional neural network (CNN) structure, based on which we build a novel scene sketch recognition system. Our system has been tested on the collected scene sketch dataset and compared with other state-of-the-art CNNs and sketch recognition approaches. Our experimental results demonstrate that our method achieves the state of art. Yuxiang Ye, Yijuan Lu, Hao Jiang 0007 |
ICMR | 2 |
| 2016 | Efficient 3D reflection symmetry detection: A view-based approach
Bo Li 0013, Henry Johan, Yuxiang Ye, Yijuan Lu |
Graph. Model. | 4 |
| 2016 | Shape Retrieval of Non-rigid 3D Human Modelsabstract3D models of humans are commonly used within computer graphics and vision, and so the ability to distinguish between body shapes is an important shape retrieval problem. We extend our recent paper which provided a benchmark for testing non-rigid 3D shape retrieval algorithms on 3D human models. This benchmark provided a far stricter challenge than previous shape benchmarks. We have added 145 new models for use as a separate training set, in order to standardise the training data used and provide a fairer comparison. We have also included experiments with the FAUST dataset of human scans. All participants of the previous benchmark study have taken part in the new tests reported here, many providing updated results using the new data. In addition, further participants have also taken part, and we provide extra analysis of the retrieval results. A total of 25 different shape retrieval methods are compared. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhouhui Lian, Masaki Aono, A. Ben Hamza, Alexander M. Bronstein, Michael M. Bronstein, S. Bu, Umberto Castellani, S. Cheng, Valeria Garro, Andrea Giachetti 0001, Afzal Godil, Luca Isaia, Henry Johan, Long Lai, Bo Li 0013, Chenfeng Li, Hai-Sheng Li 0002, Roee Litman, Yijuan Lu, Li Sun 0004, Gary K. L. Tam, Atsushi Tatsuma, Jianbo Ye |
Int. J. Comput. Vis. | 27 |
| 2016 | A novel hand-drawn sketch descriptor based on the fusion of multiple features
Peng Zhao 0010, Guoqin Wu, Yijuan Lu, Xianwen Wu, Sheng Yao 0001 |
Neurocomputing | 3 |
| 2015 | 3D Sketch-Based 3D Model RetrievalabstractHow to draw 3D sketches in a three-dimensional space and how to use a hand-drawn 3D sketch to search similar 3D models are brand new and challenging research topics. In this paper, we make an initial study on 3D sketching and propose a novel 3D sketch-based 3D model retrieval system. Our system allows users to freely draw 3D sketches in the air as well as to find similar 3D models given human-drawn 3D sketches. Promising retrieval performance has been achieved in experiments based on 300 collected 3D sketches and a recent large scale sketch-based 3D shape retrieval benchmark. Bo Li 0013, Yijuan Lu, Azeem Ghumman, Bradley Strylowski, Mario A. Gutierrez, Safiyah Sadiq, Scott Forster, Natacha Feola, Travis Bugerin |
ICMR | 2 |
| 2015 | KinectSBR: A Kinect-Assisted 3D Sketch-Based 3D Model Retrieval SystemabstractHow to draw 3D sketches and how to search 3D models based on a hand-drawn 3D sketch are interesting but challenging questions. In this demonstration, we try to answer them by developing a novel Kinect-assisted 3D sketch-based 3D model retrieval system which also allows users to freely draw 3D sketches in a 3D space. We demonstrate its promising potentials in both collecting 3D sketch data and conducting 3D sketch-based 3D model retrieval. Bo Li 0013, Yijuan Lu, Azeem Ghumman, Bradley Strylowski, Mario A. Gutierrez, Safiyah Sadiq, Scott Forster, Natacha Feola, Travis Bugerin |
ICMR | 2 |
| 2015 | A comparison of 3D shape retrieval methods based on a large-scale benchmark supporting multimodal queries
Bo Li 0013, Yijuan Lu, Chunyuan Li, Afzal Godil, Tobias Schreck, Masaki Aono, Martin Burtscher, Nihad Karim Chowdhury, Hongbo Fu 0001, Takahiko Furuya, Hai-Sheng Li 0002, Jianzhuang Liu, Henry Johan, Ryuichi Kosaka, Hitoshi Koyanagi, Ryutarou Ohbuchi, Atsushi Tatsuma, Yajuan Wan, Changqing Zou |
Comput. Vis. Image Underst. | 2 |
| 2015 | Visual word expansion and BSIFT verification for large-scale image search
Wengang Zhou 0001, Houqiang Li, Yijuan Lu, Meng Wang 0001, Qi Tian 0001 |
Multim. Syst. | 3 |
| 2015 | Exploration of Image Search Results Quality AssessmentabstractImage retrieval plays an increasingly important role in our daily lives. There are many factors which affect the quality of image search results, including chosen search algorithms, ranking functions, and indexing features. Applying different settings for these factors generates search result lists with varying levels of quality. However, no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper aims to solve this problem and makes four main innovations. First, a preference learning model is proposed to quantitatively study and formulate the best image search result list identification problem. Second, a set of valuable preference learning related features is proposed by exploring the visual characters of returned images. Third, a query-dependent preference learning model is further designed for building a more precise and query-specific model. Fourth, the proposed approach has been tested on a variety of applications including re-ranking ability assessment, optimal search engine selection, and synonymous query suggestion. Extensive experimental results on three image search datasets demonstrate the effectiveness and promising potential of the proposed method. Xinmei Tian 0001, Yijuan Lu, Nate Stender, Linjun Yang, Dacheng Tao |
IEEE Trans. Big Data | 2 |
| 2015 | Image Search Reranking With Hierarchical Topic AwarenessabstractWith much attention from both academia and industrial communities, visual search reranking has recently been proposed to refine image search results obtained from text-based image search engines. Most of the traditional reranking methods cannot capture both relevance and diversity of the search results at the same time. Or they ignore the hierarchical topic structure of search result. Each topic is treated equally and independently. However, in real applications, images returned for certain queries are naturally in hierarchical organization, rather than simple parallel relation. In this paper, a new reranking method "topic-aware reranking (TARerank)" is proposed. TARerank describes the hierarchical topic structure of search results in one model, and seamlessly captures both relevance and diversity of the image search results simultaneously. Through a structured learning framework, relevance and diversity are modeled in TARerank by a set of carefully designed features, and then the model is learned from human-labeled training samples. The learned model is expected to predict reranking results with high relevance and diversity for testing queries. To verify the effectiveness of the proposed method, we collect an image search dataset and conduct comparison experiments on it. The experimental results demonstrate that the proposed TARerank outperforms the existing relevance-based and diversified reranking methods. Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Qi Tian 0001, Dacheng Tao |
IEEE Trans. Cybern. | 3 |
| 2015 | BSIFT: Toward Data-Independent Codebook for Large Scale Image SearchabstractBag-of-Words (BoWs) model based on Scale Invariant Feature Transform (SIFT) has been widely used in large-scale image retrieval applications. Feature quantization by vector quantization plays a crucial role in BoW model, which generates visual words from the high- dimensional SIFT features, so as to adapt to the inverted file structure for the scalable retrieval. Traditional feature quantization approaches suffer several issues, such as necessity of visual codebook training, limited reliability, and update inefficiency. To avoid the above problems, in this paper, a novel feature quantization scheme is proposed to efficiently quantize each SIFT descriptor to a descriptive and discriminative bit-vector, which is called binary SIFT (BSIFT). Our quantizer is independent of image collections. In addition, by taking the first 32 bits out from BSIFT as code word, the generated BSIFT naturally lends itself to adapt to the classic inverted file structure for image indexing. Moreover, the quantization error is reduced by feature filtering, code word expansion, and query sensitive mask shielding. Without any explicit codebook for quantization, our approach can be readily applied in image search in some resource-limited scenarios. We evaluate the proposed algorithm for large scale image search on two public image data sets. Experimental results demonstrate the index efficiency and retrieval accuracy of our approach. Wengang Zhou 0001, Houqiang Li, Richang Hong, Yijuan Lu, Qi Tian 0001 |
IEEE Trans. Image Process. | 4 |
| 2014 | A comparison of methods for sketch-based 3D shape retrieval
Bo Li 0013, Yijuan Lu, Afzal Godil, Tobias Schreck, Benjamin Bustos, Alfredo Ferreira, Takahiko Furuya, Manuel J. Fonseca, Henry Johan, Takahiro Matsuda 0003, Ryutarou Ohbuchi, Pedro B. Pascoal, José M. Saavedra |
Comput. Vis. Image Underst. | 2 |
| 2014 | Encoding Spatial Context for Large-Scale Partial-Duplicate Web Image Retrieval
Wengang Zhou 0001, Houqiang Li, Yijuan Lu, Qi Tian 0001 |
J. Comput. Sci. Technol. | 3 |
| 2014 | A benchmark of simulated range images for partial shape retrieval
Ivan Sipiran, Rafael Meruane, Benjamin Bustos, Tobias Schreck, Bo Li 0013, Yijuan Lu, Henry Johan |
Vis. Comput. | 6 |
| 2013 | Human movement summarization and depiction from videosabstractHuman movement summarization and depiction from videos is to automatically turn an input video into high level action illustrations, in which the movements of the body parts are visualized using arrows and motion particles. Motion depiction compactly illustrates how specific movements are performed. Previous action summarization methods reply on 3D motion capture or manually labeled data, without which depicting actions is a challenging task. In this paper, we propose a novel scheme to automatically summarize and depict human movements from 2D videos without 3D motion capture or manually labeled data. The proposed method first segments videos into sub-actions with an effective streamline matching scheme. Then, to estimate human movement, we propose a novel trajectory following method to track points by using both body part detection and optical flow. With the estimated movement, we depict the human articulated motion with arrows and motion particles. Our experiments on a variety of videos show that the proposed method is effective in summarizing complex human movements and generating compact depictions. Yijuan Lu, Hao Jiang 0007 |
ICME | 1 |
| 2013 | Semantic-Spatial Matching for image classificationabstractSpatial Pyramid Matching (SPM) has been proven a simple but effective extension to bag-of-visual-words image representation for spatial layout information compensation. SPM describes image in coarse-to-fine scale by partitioning the image into blocks over multiple levels and the features extracted from each block are concatenated into a long vector representation. Based on the assumption that images from the same class have similar spatial configurations, SPM matches the blocks from different images according to their spatial layout, by aligning all blocks from an image in a fixed spatial order. However, target objects may appear at any location in the image with various backgrounds. Therefore, the fixed spatial matching in SPM fails to match similar objects located different locations. To solve this problem, we propose an effective and efficient block matching method, Semantic-Spatial Matching (SSM). In this method, not only the spatial layout but also the semantic content is considered for block matching. The experiments on two benchmark image classification datasets demonstrate the effectiveness of SSM. Yupeng Yan, Xinmei Tian 0001, Linjun Yang, Yijuan Lu, Houqiang Li |
ICME | 4 |
| 2013 | Discriminative codebook learning for Web image search
Xinmei Tian 0001, Yijuan Lu |
Signal Process. | 2 |
| 2013 | SIFT match verification by geometric coding for large-scale partial-duplicate web image searchabstractMost large-scale image retrieval systems are based on the bag-of-visual-words model. However, the traditional bag-of-visual-words model does not capture the geometric context among local features in images well, which plays an important role in image retrieval. In order to fully explore geometric context of all visual words in images, efficient global geometric verification methods have been attracting lots of attention. Unfortunately, current existing methods on global geometric verification are either computationally expensive to ensure real-time response, or cannot handle rotation well. To solve the preceding problems, in this article, we propose a novel geometric coding algorithm, to encode the spatial context among local features for large-scale partial-duplicate Web image retrieval. Our geometric coding consists of geometric square coding and geometric fan coding, which describe the spatial relationships of SIFT features into three geo-maps for global verification to remove geometrically inconsistent SIFT matches. Our approach is not only computationally efficient, but also effective in detecting partial-duplicate images with rotation, scale changes, partial-occlusion, and background clutter. Experiments in partial-duplicate Web image search, using two datasets with one million Web images as distractors, reveal that our approach outperforms the baseline bag-of-visual-words approach even following a RANSAC verification in mean average precision. Besides, our approach achieves comparable performance to other state-of-the-art global geometric verification methods, for example, spatial coding scheme, but is more computationally efficient. Wengang Zhou 0001, Houqiang Li, Yijuan Lu, Qi Tian 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Scalar quantization for large scale image searchabstractBag-of-Words (BoW) model based on SIFT has been widely used in large scale image retrieval applications. Feature quantization plays a crucial role in BoW model, which generates visual words from the high dimensional SIFT features, so as to adapt to the inverted file structure for indexing. Traditional feature quantization approaches suffer several problems: 1) high computational cost---visual words generation (codebook construction) is time consuming especially with large amount of features; 2) limited reliability---different collections of images may produce totally different codebooks and quantization error is hard to be controlled; 3) update inefficiency--once the codebook is constructed, it is not easy to be updated. In this paper, a novel feature quantization algorithm, scalar quantization, is proposed. With scalar quantization, a SIFT feature is quantized to a descriptive and discriminative bit-vector, of which the first tens of bits are taken out as code word. Our quantizer is independent of collections of images. In addition, the result of scalar quantization naturally lends itself to adapt to the classic inverted file structure for image indexing. Moreover, the quantization error can be flexibly reduced and controlled by efficiently enumerating nearest neighbors of code words. Wengang Zhou 0001, Yijuan Lu, Houqiang Li, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2012 | Genre identification for office document search and browsing
Francine Chen 0001, Andreas Girgensohn, Matthew Cooper 0002, Yijuan Lu, Gerry Filby |
Int. J. Document Anal. Recognit. | 4 |
| 2012 | Principal Visual Word Discovery for Automatic License Plate DetectionabstractLicense plates detection is widely considered a solved problem, with many systems already in operation. However, the existing algorithms or systems work well only under some controlled conditions. There are still many challenges for license plate detection in an open environment, such as various observation angles, background clutter, scale changes, multiple plates, uneven illumination, and so on. In this paper, we propose a novel scheme to automatically locate license plates by principal visual word (PVW), discovery and local feature matching. Observing that characters in different license plates are duplicates of each other, we bring in the idea of using the bag-of-words (BoW) model popularly applied in partial-duplicate image search. Unlike the classic BoW model, for each plate character, we automatically discover the PVW characterized with geometric context. Given a new image, the license plates are extracted by matching local features with PVW. Besides license plate detection, our approach can also be extended to the detection of logos and trademarks. Due to the invariance virtue of scale-invariant feature transform feature, our method can adaptively deal with various changes in the license plates, such as rotation, scaling, illumination, etc. Promising results of the proposed approach are demonstrated with an experimental study in license plate detection. Wengang Zhou 0001, Houqiang Li, Yijuan Lu, Qi Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2012 | Query Difficulty Prediction for Web Image SearchabstractImage search plays an important role in our daily life. Given a query, the image search engine is to retrieve images related to it. However, different queries have different search difficulty levels. For some queries, they are easy to be retrieved (the search engine can return very good search results). While for others, they are difficult (the search results are very unsatisfactory). Thus, it is desirable to identify those “difficult” queries in order to handle them properly. Query difficulty prediction (QDP) is an attempt to predict the quality of the search result for a query over a given collection. QDP problem has been investigated for many years in text document retrieval, and its importance has been recognized in the information retrieval (IR) community. However, little effort has been conducted on the image query difficulty prediction problem for image search. Compared with QDP in document retrieval, QDP in image search is more challenging due to the noise of textual features and the well-known semantic gap of visual features. This paper aims to investigate the QDP problem in Web image search. A novel method is proposed to automatically predict the quality of image search results for an arbitrary query. This model is built based on a set of valuable features that are designed by exploring the visual characteristic of images in the search results. The experiments on two real image search datasets demonstrate the effectiveness of the proposed query difficulty prediction method. Two applications, including optimal image search engine selection and search results merging, are presented to show the promising applicability of QDP. Xinmei Tian 0001, Yijuan Lu, Linjun Yang |
IEEE Trans. Multim. | 2 |
| 2011 | Spatial pooling for transformation invariant image representationabstractSpatial Pyramid Matching (SPM) [2] has been proposed to extend the Bag-of-Word (BoW) model for object classification. By re-serving the finer level information, it makes image matching more accurate. However, for not well-aligned images, where the object is rotated, flipped or translated, SPM may lose its discrimination power. To tackle this problem, we propose novel spatial pooling layouts to address various transformations, and generate a more general image representation. To evaluate the effectiveness of the proposed approach, we conduct extensive experiments on three transformation emphasized datasets for object classification task. Experimental results demonstrate its superiority over the state-of-the-arts. Besides, the proposed image representation is compact and consistent with the BoW model, which makes it applicable to image retrieval task as well. Yan Song 0001, Yijuan Lu, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2011 | Learning to judge image search resultsabstractGiven the explosive growth of the Web and the popularity of image sharing Web sites, image retrieval plays an increasingly important role in our daily lives. Search engines aim to provide beneficial image search results to users in response to queries. The quality of image search results depends on many factors: chosen search algorithms, ranking functions, indexing features, the base image database, etc. Applying different settings for these factors generates search result lists with varying levels of quality. Previous research has shown that no setting can always perform optimally for all queries. Therefore, given a set of search result lists generated by different settings, it is crucial to automatically determine which result list is the best in order to present it to users. This paper proposes a novel method to automatically identify the best search result list from a number of candidates. There are three main innovations in this paper. First, we propose a preference learning model to quantitatively study the best image search result identification problem. Second, we propose a set of valuable preference learning related features by exploring the visual characters of returned images. Third, our method shows promising potential in applications such as reranking ability assessment and optimal search engine selection. Experiments on two image search datasets show that our method achieves about 80% prediction accuracy for reranking ability assessment, and selects optimal search engine for about 70% queries correctly. Xinmei Tian 0001, Yijuan Lu, Linjun Yang, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2011 | Large scale image search with geometric codingabstractBag-of-Visual-Words model is popular in large-scale image search. However, traditional Bag-of-Visual-Words model does not capture the geometric context among local features in images. To fully explore geometric context of all visual words in images, efficient global geometric verification methods are demanded. In this paper, we propose a novel geometric coding algorithm to encode the spatial context among local features of an image for large scale partial duplicate image retrieval. Our approach is not only computationally efficient, but also can effectively detect duplicate images with rotation, scale changes, occlusion, and background clutter with low computational cost. Experiments show the promising results of our approach. Wengang Zhou 0001, Houqiang Li, Yijuan Lu, Qi Tian 0001 |
ACM Multimedia | 3 |
| 2011 | Personalization in multimedia retrieval: A survey
Yijuan Lu, Nicu Sebe, Ross Hytnen, Qi Tian 0001 |
Multim. Tools Appl. | 1 |
| 2011 | Latent visual context learning for web image applications
Wengang Zhou 0001, Qi Tian 0001, Yijuan Lu, Linjun Yang, Houqiang Li |
Pattern Recognit. | 3 |
| 2010 | Building pair-wise visual word tree for efficent image re-rankingabstractBag-of-visual Words (BoW) image representation is getting popular in computer vision and multimedia communities. However, experiments show that the traditional BoW representation is not as effective as it is desired. One of the most important reasons for its ineffectiveness is that, the traditional BoW representation lost the spatial information in images. To overcome this problem, we propose the pair-wise visual word tree, within which each visual word keeps both the appearance and spatial information between two interest points in image. Thus, the corresponding novel BoW representation preserves the spatial structure in image. Based on the pair-wise visual word tree, we propose an efficient topic word selection algorithm, which utilizes the Latent Semantic Analysis to discover the most expressive visual words for different image categories. An efficient strategy is then utilized to combine the selected topic words for image re-ranking. Massive experiments show that the novel BoW representation shows promising performance. Meanwhile, the proposed image re-ranking strategy shows the state-of-the-art precision and promising efficiency. Shiliang Zhang, Qingming Huang, Yijuan Lu, Wen Gao 0001, Qi Tian 0001 |
ICASSP | 3 |
| 2010 | Large scale partial-duplicate image retrieval with bi-space quantization and geometric consistencyabstractThe state-of-the-art image retrieval approaches represent image with a high dimensional vector of visual words by quantizing local features, such as SIFT, solely in descriptor space. The resulting visual words usually suffer from the dilemma of discrimination and ambiguity. Besides, geometric relationships among visual words are usually ignored or only used for post-processing such as re-ranking. In this paper, to improve the discriminative power and reduce the ambiguity of visual word, we propose a novel bispace quantization strategy. Local features are quantized to visual words first in descriptor space and then in orientation space. Moreover, geometric consistency constraints are embedded into the relevance formulation. Experiments in web image search with a database of one million images show that our approach achieves an improvement of 65.4% over the baseline bag-of-words approach. Wengang Zhou 0001, Houqiang Li, Yijuan Lu, Qi Tian 0001 |
ICASSP | 3 |
| 2010 | Canonical Image Selection by Visual Context LearningabstractCanonical image selection is to select a subset of photos that best summarize a photo collection. In this paper, we define the canonical image as those that contain most important and distinctive visual words. We propose to use visual context learning to discover visual word significance and develop Weighted Set Coverage algorithm to select canonical images containing distinctive visual words. Experiments with web image datasets demonstrate that the canonical images selected by our approach are not only representatives of the collected photos, but also exhibit a diverse set of views with minimal redundancy. Wengang Zhou 0001, Yijuan Lu, Houqiang Li, Qi Tian 0001 |
ICPR | 2 |
| 2010 | Spatial coding for large scale partial-duplicate web image searchabstractThe state-of-the-art image retrieval approaches represent images with a high dimensional vector of visual words by quantizing local features, such as SIFT, in the descriptor space. The geometric clues among visual words in an image is usually ignored or exploited for full geometric verification, which is computationally expensive. In this paper, we focus on partial-duplicate web image retrieval, and propose a novel scheme, spatial coding, to encode the spatial relationships among local features in an image. Our spatial coding is both efficient and effective to discover false matches of local features between images, and can greatly improve retrieval performance. Experiments in partial-duplicate web image search, using a database of one million images, reveal that our approach achieves a 53% improvement in mean average precision and 46% reduction in time cost over the baseline bag-of-words approach. Wengang Zhou 0001, Yijuan Lu, Houqiang Li, Yibing Song, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2010 | Large scale partially duplicated web image retrievalabstractThe state-of-the-art image retrieval approaches represent images with a high dimensional vector of visual words by quantizing local features, such as SIFT, in the descriptor space. The geometric clues among visual words in an image is usually ignored or exploited for full geometric verification, which is computationally expensive. In recent years, partially duplicated images are prevalent on the web. In this demo, we focus on partial-duplicated web image retrieval, and propose a retrieval system based on a novel scheme, spatial coding, to encode the spatial information among local features in an image. Our spatial coding is both efficient and effective to discover false matches of local features between images, and can greatly improve retrieval performance. Wengang Zhou 0001, Yijuan Lu, Houqiang Li, Yibing Song, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2010 | Constructing Concept Lexica With Small Semantic GapsabstractIn recent years, constructing mathematical models for visual concepts by using content features, i.e., color, texture, shape, or local features, has led to the fast development of concept-based multimedia retrieval. In concept-based multimedia retrieval, defining a good lexicon of high-level concepts is the first and important step. However, which concepts should be used for data collection and model construction is still an open question. People agree that concepts that can be easily described by low-level visual features can construct a good lexicon. These concepts are called concepts with small semantic gaps. Unfortunately, there is very little research found on semantic gap analysis and on automatically choosing multimedia concepts with small semantic gaps, even though differences of semantic gaps among concepts are well worth investigating. Yijuan Lu, Lei Zhang 0001, Jiemin Liu, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2009 | A lexica family with small semantic gapabstractDefining a lexicon of high-level concepts is the first step for data collection and model construction in concept-based image retrieval. Differences of semantic gaps among concepts are well worth considering. By measuring consistency in visual space and textual space, concepts with small semantic gap can be obtained. Considering so many diverse concepts in large-scale image dataset, we construct a lexica family of high-level concepts with small semantic gap based on different low-level features and different consistency measurements. In this lexica family, the lexica are independent to each other and mutually complementary. It provides helpful suggestions about data collection, feature selection and search model construction for large-scale image retrieval. Jiemin Liu, Qi Tian 0001, Yijuan Lu, Changhu Wang, Lei Zhang 0001, Xiaokang Yang 0001, Shipeng Li 0001 |
ICME | 3 |
| 2009 | Refining image retrieval using one-class classificationabstractCan we take advantage of the huge number of online images to improve image search quality? Motivated by this question, we propose a novel model to re-rank Google image search results by exploring the latent characteristic of massive unrelated images as a clue to filter them in the reranking. Inspired by the characteristic of the intrinsic diversity and the unwanted availability of the unrelated images, in our model, we adopt one-class classification to build a hyper-sphere for the target objects, unrelated images, and construct a robust boundary to distinguish them from the related images effectively. Then the Google results can be easily re-ranked by filtering the unrelated images with the built-up model. Extensive experiments demonstrate our approach outperforms Google image search engine's results, even if its baseline is high. Yun Fu 0001, Yijuan Lu, Qi Tian 0001 |
ICME | 3 |
| 2009 | Discriminant Subspace Analysis: An Adaptive Approach for Image ClassificationabstractLinear discriminant analysis (LDA) and biased discriminant analysis (BDA) are two effective techniques for dimension reduction, which pay attention to different roles of the positive and negative samples in finding discriminating subspace. However, the drawbacks of these two methods are obvious: LDA has limited efficiency in classifying sample data from subclasses with different distributions, and BDA does not account for the underlying distribution of negative samples. In order to effectively exploit favorable attributes of both BDA and LDA and avoid their unfavorable ones, we propose a novel adaptive discriminant analysis (ADA) for image classification. ADA can find an optimal discriminative subspace with adaptation to different sample distributions. In addition, three novel variants and extensions of ADA are further proposed: 1) integrated boosting (i.Boosting), which enhances and combines a set of ADA classifiers into a more powerful one. i.Boosting integrates feature re-weighting, relevance feedback, and AdaBoost into one framework. With affordable computational cost, i.Boosting can provide a unified and stable solution to ADA prediction result. 2) Fast adaptive discriminant analysis (FADA). Instead of searching parameters, FADA can directly find a close-to-optimal projection very fast based on different sample distributions. 3) Two-dimensional adaptive discriminant analysis (2DADA). As opposed to ADA, 2DADA is based on 2-D image matrix representation rather than 1-D vector. So it is simpler, more straightforward, and has lower time complexity to use for image feature extraction. Extensive experiments on synthetic data, UCI benchmark data sets, hand-digit data set, four facial image data sets, and COREL color image data sets show the superior performance of our proposed approaches. Yijuan Lu, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2008 | What are the high-level concepts with small semantic gaps?abstractConcept-based multimedia search has become more and more popular in Multimedia Information Retrieval (MIR). However, which semantic concepts should be used for data collection and model construction is still an open question. Currently, there is very little research found on automatically choosing multimedia concepts with small semantic gaps. In this paper, we propose a novel framework to develop a lexicon of high-level concepts with small semantic gaps (LCSS) from a large-scale web image dataset. By defining a confidence map and content-context similarity matrix, images with small semantic gaps are selected and clustered. The final concept lexicon is mined from the surrounding descriptions (titles, categories and comments) of these images. This lexicon offers a set of high-level concepts with small semantic gaps, which is very helpful for people to focus for data collection, annotation and modeling. It also shows a promising application potential for image annotation refinement and rejection. The experimental results demonstrate the validity of the developed concepts lexicon. Yijuan Lu, Lei Zhang 0001, Qi Tian 0001, Wei-Ying Ma |
CVPR | 1 |
| 2008 | Adaptive discriminant analysis for microarray-based classificationabstractMicroarray technology has generated enormous amounts of high-dimensional gene expression data, providing a unique platform for exploring gene regulatory networks. However, the curse of dimensionality plagues effort to analyze these high throughput data. Linear Discriminant Analysis (LDA) and Biased Discriminant Analysis (BDA) are two popular techniques for dimension reduction, which pay attention to different roles of the positive and negative samples in finding discriminating subspace. However, the drawbacks of these two methods are obvious: LDA has limited efficiency in classifying sample data from subclasses with different distributions, and BDA does not account for the underlying distribution of negative samples. In this paper, we propose a novel dimension reduction technique for microarray analysis: Adaptive Discriminant Analysis (ADA), which effectively exploits favorable attributes of both BDA and LDA and avoids their unfavorable ones. ADA can find a good discriminative subspace with adaptation to different sample distributions. It not only alleviates the problem of high dimensionality, but also enhances the classification performance in the subspace with naïve Bayes classifier. To learn the best model fitting the real scenario, boosted Adaptive Discriminant Analysis is further proposed. Extensive experiments on the yeast cell cycle regulation data set, and the expression data of the red blood cell cycle in malaria parasite Plasmodium falciparum demonstrate the superior performance of ADA and boosted ADA. We also present some putative genes of specific functional classes predicted by boosted ADA. Their potential functionality is confirmed by independent predictions based on Gene Ontology, demonstrating that ADA and boosted ADA are effective dimension reduction methods for microarray-based classification. Yijuan Lu, Qi Tian 0001, Jennifer L. Neary, Yufeng Wang 0002 |
ACM Trans. Knowl. Discov. Data | 1 |
| 2007 | Two-Dimensional Adaptive Discriminant AnalysisabstractIn this paper, we develop a new feature extraction and dimension reduction technique: 2-dimensional adaptive discriminant analysis (2DADA) based on 2DLDA and our proposed 2DBDA. It effectively exploits favorable attributes of both 2DBDA and 2DLDA and avoids their unfavorable ones. 2DADA can easily find an optimal discriminative subspace with adaptation to different sample distributions. It not only alleviates the problem of high dimensionality, but also enhances the classification performance in the subspace with KNN classifier. Experimental results on hand-written digit database and face databases show an improvement of 2DADA over other traditional dimension reduction techniques. Yijuan Lu, Jie Yu 0001, Nicu Sebe, Qi Tian 0001 |
ICASSP (1) | 1 |
| 2007 | Integrating Relevance Feedback in Boosting for Content-Based Image RetrievalabstractMany content-based image retrieval applications suffer from small sample set and high dimensionality problems. Relevance feedback is often used to alleviate those problems. In this paper, we propose a novel interactive boosting framework to integrate user feedback into boosting scheme and bridge the gap between high-level semantic concept and low-level image features. Our method achieves more performance improvement from the relevance feedback than AdaBoost does because human judgment is accumulated iteratively to facilitate learning process. It also has obvious advantage over the classic relevance feedback method in that the classifiers are trained to pay more attention to wrongfully predicted samples in user feedback through a reinforcement training process. An interactive boosting scheme called i.Boost is implemented and tested using adaptive discriminant projection (ADP) as base classifiers, which not only combines but also enhances a set of ADP classifiers into a more powerful one. To evaluate its performance, several applications are designed on UCI benchmark data sets, Harvard, UMIST, ATT facial image data sets and COREL color image data sets. The proposed method is compared to normal AdaBoost, classic relevance feedback and the state-of-the-art projection-based classifiers. The experiment results show the superior performance of i.Boost and the interactive boosting framework. Jie Yu 0001, Yijuan Lu, Yuning Xu, Nicu Sebe, Qi Tian 0001 |
ICASSP (1) | 2 |
| 2007 | i.Boosting for Image ClassificationabstractTraditional boosting method like adaboost, boosts a weak learning algorithm by updating the sample weights (the relative importance of the training samples) iteratively. In this paper, we propose to integrate feature re-weighting into boosting scheme, which not only weights the samples but also weights the feature elements iteratively. To avoid overfitting problem caused by feature re-weighting on a small training data set, we also incorporate relevance feedback into boosting and propose an interactive boosting called i.Boosting. It merges adaboost, feature re-weighting and relevance feedback into one framework and exploits the favorable attributes of these methods. In this paper, i.Boosting is implemented using Adaptive Discriminant Analysis (ADA) as base classifiers. It not only enhances but also combines a set of ADA classifiers into a more powerful one. A specific feature re-weighting method for ADA is also proposed and integrated in i.Boosting. Extensive experiments show the superior performance of i.Boosting over AdaBoost and other state-of-the-art projection-based classifiers. Yijuan Lu, Tong Zhang 0007, Qi Tian 0001 |
ICME | 1 |
| 2007 | Feature selection using principal feature analysisabstractDimensionality reduction of a feature set is a common preprocessing step used for pattern recognition and classification applications. Principal Component Analysis (PCA) is one of the popular methods used, and can be shown to be optimal using different optimality criteria. However, it has the disadvantage that measurements from all the original features are used in the projection to the lower dimensional space. This paper proposes a novel method for dimensionality reduction of a feature set by choosing a subset of the original features that contains most of the essential information, using the same criteria as PCA. We call this method Principal Feature Analysis (PFA). The proposed method is successfully applied for choosing the principal features in face tracking and content-based image retrieval (CBIR) problems. Automated annotation of digital pictures has been a highly challenging problem for computer scientists since the invention of computers. The capability of annotating pictures by computers can lead to breakthroughs in a wide range of applications including Web image search, online picture-sharing communities, and scientific experiments. In our work, by advancing statistical modeling and optimization techniques, we can train computers about hundreds of semantic concepts using example pictures from each concept. The ALIPR (Automatic Linguistic Indexing of Pictures - Real Time) system of fully automatic and high speed annotation for online pictures has been constructed. Thousands of pictures from an Internet photo-sharing site, unrelated to the source of those pictures used in the training process, have been tested. The experimental results show that a single computer processor can suggest annotation terms in real-time and with good accuracy. Yijuan Lu, Ira Cohen, Xiang Sean Zhou, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2007 | Interactive Semisupervised Learning for Microarray AnalysisabstractMicroarray technology has generated vast amounts of gene expression data with distinct patterns. Based on the premise that genes of correlated functions tend to exhibit similar expression patterns, various machine learning methods have been applied to capture these specific patterns in microarray data. However, the discrepancy between the rich expression profiles and the limited knowledge of gene functions has been a major hurdle to the understanding of cellular networks. To bridge this gap so as to properly comprehend and interpret expression data, we introduce Relevance Feedback to microarray analysis and propose an interactive learning framework to incorporate the expert knowledge into the decision module. In order to find a good learning method and solve two intrinsic problems in microarray data, high dimensionality and small sample size, we also propose a semisupervised learning algorithm: Kernel Discriminant-EM (KDEM). This algorithm efficiently utilizes a large set of unlabeled data to compensate for the insufficiency of a small set of labeled data and it extends the linear algorithm in Discriminant-EM (DEM) to a kernel algorithm to handle nonlinearly separable data in a lower dimensional space. The Relevance Feedback technique and KDEM together construct an efficient and effective interactive semisupervised learning framework for microarray analysis. Extensive experiments on the yeast cell cycle regulation data set and Plasmodium falciparum red blood cell cycle data set show the promise of this approach. Yijuan Lu, Qi Tian 0001, Maribel Sanchez, Yufeng Wang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2005 | Hybrid PCA and LDA Analysis of Microarray Gene Expression Data
Yijuan Lu, Qi Tian 0001, Maribel Sanchez, Yufeng Wang 0002 |
CIBCB | 1 |