VLDB 2026 Research / reviewers in the wild / expert
Duy-Dinh Le
dblp:48/4683
· DBLP profile ↗
45ranked-venue papers
9as first author
7since 2021 · last 2026
0000-0003-0356-5501ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 7 first-author · 5 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NII-UIT at VBS2026: Towards Effective Visual Question Answering for Interactive and Multimodal Video Retrieval
Bao Tran, Tien Do, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
MMM (4) | 4 |
| 2026 | Skeleton-guided artistic text recognition
Tien Do, Thuyen Tran 0001, Khiem Le 0001, Duy-Dinh Le, Thanh Duc Ngo |
Int. J. Document Anal. Recognit. | 4 |
| 2025 | Skeleton-Guided Artistic Text Recognition
Tien Do, Thuyen Tran 0001, Khiem Le 0001, Duy-Dinh Le, Thanh Duc Ngo |
ICDAR (5) | 4 |
| 2025 | NII-UIT at VBS2025: Multimodal Video Retrieval with LLM Integration and Dynamic Temporal Search
Bao Tran Gia, Tuong Bui Cong Khanh, Tam Le Thi Thanh, Thuyen Tran 0001, Khiem Le 0001, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
MMM (5) | 9 |
| 2024 | Text Query to Web Image to Video: A Comprehensive Ad-Hoc Video Search
Nhat-Minh Nguyen, Tien-Dung Mai, Duy-Dinh Le |
ACCV (3) | 3 |
| 2023 | Abstraction-perception preserving cartoon face synthesis
Sy-Tuyen Ho, Manh-Khanh Ngo Huu, Thanh-Danh Nguyen, Nguyen Phan, Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 7 |
| 2022 | UIT at VBS 2022: An Unified and Interactive Video Retrieval System with Temporal Search
Khanh Ho, Vu Xuan Dinh, Khiem Le 0001, Khang Dinh Tran, Tien Do, Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le |
MMM (2) | 9 |
| 2019 | You always look again: Learning to detect the unseen objects
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | YADA: you always dream again for better object detection
Khanh-Duy Nguyen, Khang Nguyen 0001, Duy-Dinh Le, Duc Anh Duong, Tam V. Nguyen 0002 |
Multim. Tools Appl. | 3 |
| 2018 | Video Search Based on Semantic Extraction and Locally Regional Object Proposal
Thanh-Dat Truong, Vinh-Tiep Nguyen, Minh-Triet Tran, Trang-Vinh Trieu, Tien Do, Thanh Duc Ngo, Duy-Dinh Le |
MMM (2) | 7 |
| 2017 | Evaluation of Deep Models for Real-Time Small Object Detection
Phuoc Pham, Tien Do, Thanh Duc Ngo, Duy-Dinh Le |
ICONIP (3) | 5 |
| 2017 | Video Indexing, Search, Detection, and Description with Focus on TRECVIDabstractThere has been a tremendous growth in video data the last decade. People are using mobile phones and tablets to take, share or watch videos more than ever before. Video cameras are around us almost everywhere in the public domain (e.g. stores, streets, public facilities, ...etc). Efficient and effective retrieval methods are critically needed in different applications. The goal of TRECVID is to encourage research in content-based video retrieval by providing large test collections, uniform scoring procedures, and a forum for organizations interested in comparing their results. In this tutorial, we present and discuss some of the most important and fundamental content-based video retrieval problems such as recognizing predefined visual concepts, searching in videos for complex ad-hoc user queries, searching by image/video examples in a video dataset to retrieve specific objects, persons, or locations, detecting events, and finally bridging the gap between vision and language by looking into how can systems automatically describe videos in a natural language. A review of the state of the art, current challenges, and future directions along with pointers to useful resources will be presented by different regular TRECVID participating teams. Each team will present one of the following tasks: George Awad, Duy-Dinh Le, Chong-Wah Ngo, Vinh-Tiep Nguyen, Georges Quénot, Cees Snoek, Shin'ichi Satoh 0001 |
ICMR | 2 |
| 2017 | Semantic Extraction and Object Proposal for Video Search
Vinh-Tiep Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Minh-Triet Tran, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 3 |
| 2017 | Efficient large-scale multi-class image classification by learning balanced trees
Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001 |
Comput. Vis. Image Underst. | 3 |
| 2017 | Evaluation of multiple features for violent scenes detection
Vu Lam, Sang Phan Le, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Scalable Face Track Retrieval in Video Archives Using Bag-of-Faces Sparse RepresentationabstractHuge video archives consisting of news programs, dramas, movies, and Web videos (e.g., YouTube) are available in our daily life. In all these videos, human is usually one of the most important subjects. Using state-of-the-art techniques, we can efficiently detect and track faces in the videos. In order to organize large-scale face tracks, containing sequences of (detected) consecutive faces in the videos, we propose an efficient method to retrieve human face tracks using bag-of-faces sparse representation (BoF-SR). Using the proposed method, a face track is encoded as a single BoF-SR, therefore allowing an efficient indexing method to handle large-scale data. To further consider the possible variations in face tracks, we generalize our method to find multiple SRs, in an unsupervised manner, to represent a bag of faces and balance the tradeoff between performance and retrieval time. The experimental results on two real-world (million-scale) data sets confirm that the proposed methods achieve significant performance gains compared with different state-of-the-art methods. Bor-Chun Chen, Yan-Ying Chen, Yin-Hsi Kuo, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Winston H. Hsu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | Video Event Detection by Exploiting Word Dependencies from Image CaptionsabstractVideo event detection is a challenging problem in information and multimedia retrieval. Different from single action detection, event detection requires a richer level of semantic information from video. In order to overcome this challenge, existing solutions often represent videos using high level features such as concepts. However, concept-based representation can be confusing because it does not encode the relationship between concepts. This issue can be addressed by exploiting the co-occurrences of the concepts, however, it often leads to a very huge number of possible combinations. In this paper, we propose a new approach to obtain the relationship between concepts by exploiting the syntactic dependencies between words in the image captions. The main advantage of this approach is that it significantly reduces the number of informative combinations between concepts. We conduct extensive experiments to analyze the effectiveness of using the new dependency representation for event detection on two large-scale TRECVID Multimedia Event Detection 2013 and 2014 datasets. Experimental results show that i) Dependency features are more discriminative than concept-based features. ii) Dependency features can be combined with our current event detection system to further improve the performance. For instance, the relative improvement can be as far as 8.6% on the MEDTEST14 10Ex setting. Sang Phan Le, Yusuke Miyao, Duy-Dinh Le, Shin'ichi Satoh 0001 |
COLING | 3 |
| 2016 | Efficient Large Scale Image Classification via Prediction Score Decomposition
Duy-Dinh Le, Tien-Dung Mai, Shin'ichi Satoh 0001, Thanh Duc Ngo, Duc Anh Duong |
ECCV (6) | 1 |
| 2016 | Using node relationships for hierarchical classificationabstractHierarchical classification is a computational efficient approach for large-scale image classification. The main challenging issue of this approach is to deal with error propagation. Irrelevant branching decision made at a parent node cannot be corrected at its child nodes in traversing the tree for classification. This paper presents a novel approach to reduce branching error at a node by taking its relative relationship into account. Given a node on the tree, we model each candidate branch by considering classification response of its child nodes, grandchild nodes and their differences with siblings. A maximum margin classifier is then applied to select the most discriminating candidate. Our proposed approach outperforms related approaches on Caltech-256, SUN-397 and ILSVRC2010-1K. Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001 |
ICIP | 3 |
| 2016 | News Archive Exploration Combining Face Detection and Tracking with Network Visual AnalyticsabstractVisual analytics helps analytical reasoning and exploration of complex systems, for which it combines the means of interactive visualization, with the power of data analytics. The recent progress in computer vision techniques opens wide applications in real world video archives. Particularly, recent advances in face detection and recognition have been put under the spotlight. The applications of such techniques are often concern intelligence, or peer recognition in photo posted in social networks. We propose to combine those two domains by demonstrating a visual exploration of over a decade of the news program from the Japanese broadcaster NHK News 7. We derive social networks from face detection and tracking of this large dataset. With the help of a little domain knowledge, we monitor the activity of political public figures and explore the archive. This allows understanding and comparison of the politico-media scene presented by NHK under different Prime Minister's governance. The social networks are interactive, and also allow to explore the multimedia database and explore its video content. Benjamin Renoust, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
ACM Multimedia | 3 |
| 2016 | Visual Analytics of Political Networks From Face-Tracking of News VideoabstractThe rich nature of news makes it a classic subject of visual analytics research. Such analysis is often based on rich textual data. However, we want to test how much we can understand the news from video information through face detection and tracking. Towards this goal, we propose a visual analytics system and discuss its design and implementation to support media experts in understanding political interactions in an archive of 12 years of the Japanese public broadcaster NHK's News 7 program. After identifying the tasks and abstraction required for our analysis, we construct links from face detection and tracking to derive multiple political networks. Our proposed design embeds this rich data into a visual analytics framework that presents four levels of abstraction: time period, network, timeline, and face-tracks within video. We present how the exploration of the archive with our system results in good understanding of the Japanese politico-media scene during these 12 years while finding evidence of “presidentialization” of the media. Benjamin Renoust, Duy-Dinh Le, Shin'ichi Satoh 0001 |
IEEE Trans. Multim. | 2 |
| 2015 | Multimedia Event Detection Using Event-Driven Multiple Instance LearningabstractA complex event can be recognized by observing necessary evidences. In the real world scenarios, this is a difficult task because the evidences can happen anywhere in a video. A straightforward solution is to decompose the video into several segments and search for the evidences in each segment. This approach is based on the assumption that segment annotation can be assigned from its video label. However, this is a weak assumption because the importance of each segment is not considered. On the other hand, the importance of a segment to an event can be obtained by matching its detected concepts against the evidential description of that event. Leveraging this prior knowledge, we propose a new method, Event-driven Multiple Instance Learning (EDMIL), to learn the key evidences for event detection. We treat each segment as an instance and quantize the instance-event similarity into different levels of relatedness. Then the instance label is learned by jointly optimizing the instance classifier and its related level. The significant performance improvement on the TRECVID Multimedia Event Detection (MED) 2012 dataset proves the effectiveness of our approach. Sang Phan Le, Duy-Dinh Le, Shin'ichi Satoh 0001 |
ACM Multimedia | 2 |
| 2015 | NII-UIT Browser: A Multimodal Video Search System
Thanh Duc Ngo, Vinh-Tiep Nguyen, Vu Hoang Nguyen, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 4 |
| 2015 | AttRel: An Approach to Person Re-Identification by Exploiting Attribute Relationships
Ngoc-Bao Nguyen, Vu Hoang Nguyen, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong |
MMM (2) | 4 |
| 2015 | Human Action recognition from depth videos using multi-projection based representationabstractIn this paper, a novel method for human action recognition from depth videos is proposed. We project 3D data on to multiple 2D-planes from which dense trajectories features are extracted. In the training stage, for each projection, a classifier is trained using the training data. In the testing stage, for each test video, the multiple trained classifiers are applied and the predicted scores are combined for final decision. We propose a greedy-based method to select a subset of the trained classifiers for optimal combination. Experiments on the MSR Action 3D dataset show that the proposed method outperforms the baseline method that does not use multi-projection-based features. Chien-Quang Le, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Duc Anh Duong |
MMSP | 3 |
| 2015 | Large scale multi-class classification using latent classifiersabstractWe study the problem of multi-class image classification with large number of classes, of which the one-vs-all based approach is prohibitive in practical applications. Recent state-of-the-art approaches rely on label tree to reduce classification complexity. However, building optimal tree structures and learning precise classifiers to optimize tree loss is challenging. In this paper, we introduce a novel approach using latent classifiers that can achieve comparable speed but better performance. The key idea is that instead of using C one-vs-all classifiers (C is the number of classes) to generate the score matrix for label prediction, a much smaller number of classifiers are used. These classifiers, called latent classifiers, are generated by analyzing the correlation among classes and removing redundancy. Experiments on several large datasets including ImageNet-1K, SUN-397, and Caltech-256 show the efficiency of our approach. Tien-Dung Mai, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Kiem Hoang, Shin'ichi Satoh 0001 |
MMSP | 3 |
| 2015 | Query-adaptive late fusion with neural network for instance searchabstractBag-of-Word based model is one of the state-of-the-art approaches for object retrieval or also known as instance search problem. Although this model and its extensions are good for rich-textured objects, it is still unsolved for searching on textureless ones. In this paper, we propose to combine this model with Deformable Part Models object detector using late fusion technique to improve final result. To find the optimal weights for each type of query objects, we further propose to use a neural network to learn query features including object area, number of shared visual words to get optimal weights for each model. Experimental results on TRECVID Instance Search (INS) dataset with queries in INS2013 and INS2014 show that our proposed method significantly improves 18.48% and 14.63% in mAP respectively comparing to standard BOW model and outperform other state-of-the-art methods. This method opens a new way of adaptively combining DPM, an object detector, in a hybrid model for visual instance search. Vinh-Tiep Nguyen, Dinh-Luan Nguyen, Minh-Triet Tran, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMSP | 4 |
| 2015 | Cross-View Action Recognition by Projection-Based Augmentation
Chien-Quang Le, Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001, Duc Anh Duong |
PSIVT | 3 |
| 2014 | Sum-max video pooling for complex event recognitionabstractA video can be viewed as a layered structure where the lowest layer are frames, the top layer is the entire video, and the middle layers are the sequences of consecutive frames or the concatenation of lower layers. While it is easy to find local discriminative features in video from lower layers, it is non-trivial to aggregate these features into a discriminative video representation. In literature, people often use sum pooling to obtain reasonable recognition performance on artificial videos. However, the sum pooling technique does not work well on complex videos because the region of interests may reside within some middle layers. In this paper, we leverage the layered structure of video to propose a new pooling method, named sum-max video pooling, to handle this problem. Basically, we apply sum pooling at the low layer representation while using max pooling at the high layer representation. Sum pooling is used to keep sufficient relevant features at the low layer, while max pooling is used to retrieve the most relevant features at the high layer, therefore it can discard irrelevant features in the final video representation. Experimental results on the TRECVID Multimedia Event Detection 2010 dataset shows the effectiveness of our method. Sang Phan Le, Duy-Dinh Le, Shin'ichi Satoh 0001 |
ICIP | 2 |
| 2014 | Integrating Spatial Information into Inverted Index for Large-Scale Image RetrievalabstractIn recent years, large-scale image retrieval has been shown remarkable potential in real-life applications. To reduce retrieval time as searched database may contain thousands of images, Inverted Indexing is the basic technique, given images are represented by Bag-of-Words model. However, one major limitation of both standard Inverted Index and Bag-of-Words model is that they ignore spatial information of the visual words in images. This might reduce retrieval accuracy. In this paper, we introduce an approach to integrate spatial information into inverted index to improve accuracy while maintaining short retrieval time. Experiments conducted on several benchmark datasets (Oxford Building 5K, Paris 6K and Oxford Building 5K+100K) demonstrate the effectiveness of our proposed approach. Bien-Van Nguyen, Duy Pham, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong |
ISM | 4 |
| 2014 | NII-UIT: A Tool for Known Item Search by Sequential Pattern Filtering
Thanh Duc Ngo, Vu Hoang Nguyen, Vu Lam, Sang Phan Le, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 5 |
| 2013 | Efficient Traffic Sign Detection Using Bag of Visual Words and Multi-scales SIFT
Khanh-Duy Nguyen, Duy-Dinh Le, Duc Anh Duong |
ICONIP (3) | 2 |
| 2013 | Person Re-identification Using Deformable Part Models
Vu Hoang Nguyen, Kien Nguyen 0002, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
ICONIP (3) | 3 |
| 2013 | A Classification-Based Approach for Retake and Scene Detection in Rushes Video
Quang-Vinh Tran, Duy-Dinh Le, Duc Anh Duong, Shin'ichi Satoh 0001 |
ICONIP (3) | 2 |
| 2013 | NII-UIT-VBS: A Video Browsing Tool for Known Item Search
Duy-Dinh Le, Vu Lam, Thanh Duc Ngo, Vinh Quang Tran, Vu Hoang Nguyen, Duc Anh Duong, Shin'ichi Satoh 0001 |
MMM (2) | 1 |
| 2012 | Auto face re-ranking by mining the web and video archivesabstractIt is necessary to utilize visual information to improve the efficiency of retrieval in image-search engines that use textual information for indexing. One popular approach has been to learn visual consistency between images returned by these search engines. Most state-of-the-art methods of learning visual consistency usually learn one specific classifier for each query to re-rank the returned images. The main drawback with these query-specific based methods is that they require computational cost and processing time that are unsuitable for handling a large number of queries. Another approach has been to learn one generic classifier once and then use for all queries. Pursuing the generic classifier based approach, we study the problem of re-ranking faces returned by existing search engines to improve retrieval performance. Learning a generic classifier involves finding good query-dependent feature representation and collecting sufficient large number of training samples. Existing work [9, 15] studies query-dependent features for general objects rather than faces. In addition, training samples are usually collected manually. The key contribution of this research is to introduce a query-dependent feature for faces and an unsupervised method of automatically collecting training samples to learn the generic classifier. The experimental results demonstrated that the proposed method performed very well in various datasets. Duy-Dinh Le, Shin'ichi Satoh 0001 |
CVPR | 1 |
| 2012 | Robust eye localization in video by combining eye detector and eye tracker
Chi Nhan Duong, Thang Cap Pham Dinh, Thanh Duc Ngo, Duy-Dinh Le, Duc Anh Duong, Bac Le, Shin'ichi Satoh 0001 |
ICPR | 4 |
| 2011 | Boosting global scene classification accuracy by discriminative region localizationabstractCombining global scene classification with object detection has helped in improving the classification accuracy. However, training an object detector requires a large amount of manual annotation. The object detector may also fail when the object is occluded. Meanwhile, the presence of the object is not only indicated by the entire object region but any of its parts or its correlations with other regions in the image. To overcome these limitations, we propose using discriminative region localization instead of object detection in the combination. Our contribution is two-fold, a) a complete framework that combines global scene classification with discriminative region localization for image classification and b) a weakly supervised discriminative region localization approach that utilizes spatial context to improve the learning accuracy. Our experimental results on benchmark datasets demonstrated that the proposed discriminative region localization approach outperforms the state-of-the-art approach. In addition, the combination significantly increases the classification performance. Thanh Duc Ngo, Duy-Dinh Le, Shin'ichi Satoh 0001 |
ICIP | 2 |
| 2011 | Fast face sequence matching in large-scale video databasesabstractThere have recently been many methods proposed for matching face sequences in the field of face retrieval. However, most of them have proven to be inefficient in large-scale video databases because they frequently require a huge amount of computational cost to obtain a high degree of accuracy. We present an efficient matching method that is based on the face sequences (called face tracks) in large-scale video databases. The key idea is how to capture the distribution of a face track in the fewest number of low-computational steps. In order to do that, each face track is represented by a vector that approximates the first principal component of the face track distribution and the similarity of face tracks bases on the similarity of these vectors. Our experimental results from a large-scale database of 457,320 human faces extracted from 370 hours of TRECVID videos from 2004-2006 show that the proposed method easily handles the scalability by maintaining a good balance between the speed and the accuracy. Hung Thanh Vu, Thanh Duc Ngo, Thao Ngoc Nguyen, Duy-Dinh Le, Shin'ichi Satoh 0001, Bac Le, Duc Anh Duong |
ICIP | 4 |
| 2008 | Unsupervised Face Annotation by Mining the WebabstractSearching for images of people is an essential task for image and video search engines. However, current search engines have limited capabilities for this task since they rely on text associated with images and video, and such text is likely to return many irrelevant results. We propose a method for retrieving relevant faces of one person by learning the visual consistency among results retrieved from text correlation-based search engines. The method consists of two steps. In the first step, each candidate face obtained from a text-based search engine is ranked with a score that measures the distribution of visual similarities among the faces. Faces that are possibly very relevant or irrelevant are ranked at the top or bottom of the list, respectively. The second step improves this ranking by treating this problem as a classification problem in which input faces are classified as psilaperson-Xpsila or psilanon-person-Xpsila; and the faces are re-ranked according to their relevant score inferred from the classifierpsilas probability output. To train this classifier, we use a bagging-based framework to combine results from multiple weak classifiers trained using different subsets. These training subsets are extracted and labeled automatically from the rank list produced from the classifier trained from the previous step. In this way, the accuracy of the ranked list increases after a number of iterations. Experimental results on various face sets retrieved from captions of news photos show that the retrieval performance improved after each iteration, with the final performance being higher than those of the existing algorithms. Duy-Dinh Le, Shin'ichi Satoh 0001 |
ICDM | 1 |
| 2008 | A text segmentation based approach to video shot boundary detectionabstractVideo shot boundary detection is one of the fundamental tasks of video indexing and retrieval applications. Although many methods have been proposed for this task, finding a general and robust shot boundary method that is able to handle the various transition types caused by photo flashes, rapid camera movement and object movement is still challenging. We present a novel approach for detecting video shot boundaries in which we cast the problem of shot boundary detection into the problem of text segmentation in natural language processing. This is possible by assuming that each frame is a word and then the shot boundaries are treated as text segment boundaries (e.g. topics). The text segmentation based approaches in natural language processing can be used. The experimental results from various long video sequences have proved the effectiveness of our approach. Duy-Dinh Le, Shin'ichi Satoh 0001, Thanh Duc Ngo, Duc Anh Duong |
MMSP | 1 |
| 2007 | Boosting Face Retrieval by using Relevant Set Correlation ClusteringabstractWe present a method to improve the performance of face retrieval in news videos by using the relevant-set correlation (RSC) clustering model. In this method, faces of a person are firstly retrieved by finding in shots whose associated transcripts contain that person's name. Then, by using the RSC clustering, these faces are organized into clusters and only representative faces of these clusters are introduced to users. As a result, the retrieval performance is significantly increased since only a small number of faces belonging to the clusters relevant to the target person are returned instead of all initial retrieved faces. The contribution of new RSC clustering model for this problem is two-fold: First, it can automatically determine an appropriate number of clusters and discards a large number of irrelevant faces. Second, since the precision of clusters can be controlled easily by a threshold, high precision clusters can be obtained and treated as labeled sets so that they can be used for feature selection by using the linear discriminant analysis (LDA) method. Experiments on the TRECVID 2004 dataset showed that the retrieval performance of the proposed method outperforms other existing methods. Duy-Dinh Le, Shin'ichi Satoh 0001, Michael E. Houle |
ICME | 1 |
| 2007 | Ent-Boost: Boosting using entropy measures for robust object detection
Duy-Dinh Le, Shin'ichi Satoh 0001 |
Pattern Recognit. Lett. | 1 |
| 2006 | Robust Object Detection using Fast Feature Selection from Huge Feature SetsabstractThis paper describes an efficient feature selection method that quickly selects a small subset out of a given huge feature set; for building robust object detection systems. In this filter-based method, features are selected not only to maximize their relevance with the target class but also to minimize their mutual dependency. As a result, the selected feature set contains only highly informative and non-redundant features, which significantly improve classification performance when combined. The relevance and mutual dependency of features are measured by using conditional mutual information (CMI) in which features and classes are treated as discrete random variables. Experiments on different huge feature sets have shown that the proposed CMI-based feature selection can both reduce the training time significantly and achieve high accuracy. Duy-Dinh Le, Shin'ichi Satoh 0001 |
ICIP | 1 |
| 2005 | Multi-Stage Approach to Fast Face Detection
Duy-Dinh Le, Shin'ichi Satoh 0001 |
BMVC | 1 |