VLDB 2026 Research / reviewers in the wild / expert
Hangzai Luo
dblp:58/3640
· DBLP profile ↗
62ranked-venue papers
12as first author
14since 2021 · last 2025
0000-0003-3354-5739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 11 first-author · 3 since 2021Artificial intelligence and machine learning · 15 · 7 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task LearningabstractGenerative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in downstream visual tasks. This paper introduces the Iterative Self-Training with Class-Aware Text-to-Image Synthesis (IST-CATS) framework, which addresses these challenges by integrating a class-aware text-to-image synthesis (CATS) component with an iterative self-training (IST) strategy. CATS innovatively introduces a class-aware chain approach to generate detailed descriptions. These descriptions act as prompts for a diffusion model, enabling the creation of a diverse of images accompanied by distinguishable objects against the background. The generated images can be easily pseudo-labeled by an unsupervised instance segmentation method, and then noisy pseudo labels can be effectively purified by a novel feature similarity-based filtering mechanism. The generated images underpin our IST, which progressively enhances vision models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt meticulously improves the quality of pseudo labels by employing class-adaptive techniques at both the pixel and object levels, ensuring refined pseudo-label accuracy. IST-CATS demonstrates superior performance in object detection and semantic segmentation compared to traditional synthetic and semi/weakly-supervised methods, effectively addressing data collection and annotation challenges. Xiang Zhang 0018, Wanqing Zhao, Pengyang Li, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 5 |
| 2025 | Semantic image segmentation via dynamic curriculum learning
Xiang Zhang 0018, Wanqing Zhao, Chenji Wang, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Appl. Intell. | 4 |
| 2024 | Enabling Near-Zero Cost Object Detection in Remote Sensing Imagery via Progressive Self-TrainingabstractDeep learning-based object detection models rely heavily on large-scale and precise annotations for training. However, manually annotating bounding-box annotations for such data is both time-consuming and costly, especially when dealing with high-resolution satellite imagery containing densely packed small-sized objects. To alleviate the burden of manual annotation, we propose a simple yet effective approach, called progressive self-training object detection (PSTDet), to enable accurate object detection in remote sensing imagery without relying on manual annotations. Our PSTDet framework consists of two main components: initial pseudo label generation (IPLG) and progressive self-training with relabeling (PST-R). In IPLG, we leverage unsupervised image clustering, unsupervised instance detection, and geometric constraints to automatically generate high-quality bounding-box annotations for the initial training dataset. This innovative approach significantly reduces the time and expense associated with data annotation, laying a solid foundation for the subsequent progressive self-training stage. The annotations produced by IPLG serve as the training data for PST-R, which enhances the detector and pseudo labels through progressive self-training and our proposed noisy pseudo label filtering strategy (NPLFilter). Our NPLFilter purifies the quality of pseudo labels by integrating geometric constraints, prior knowledge, and category-adaptive thresholds. Experimental results demonstrate that our method achieves significant performance improvement on challenging NWPU VHR-10.v2 and DIOR datasets. Notably, our method far outperforms state-of-the-art weakly supervised methods and compares favorably with fully supervised methods. Xiang Zhang 0018, Xiangteng Jiang, Qiyao Hu, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Hardware-Based Satellite Network Broadcast Storm Suppression Method
Keran Zhang, Hangzai Luo, Sheng Zhong 0006 |
Mob. Networks Appl. | 3 |
| 2023 | Learning to recover lost details from the dark
Maomei Liu, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | Toward Blind-Adaptive Remote Sensing Image RestorationabstractWhile deep convolutional neural networks (CNNs) have substantially boosted the performance of low-level vision tasks, they remain largely under-explored in CNN-based remote sensing image restoration. This paper studies the JPEG-LS compressed remote sensing image restoration that faces the following problems. It requires a trade-off in preserving local context information and expanding spatial receptive fields. It needs blind restoration while achieving flexible performance. To this end, we propose a blind-adaptive restoration network, called TBANet, that integrates three modules into an end-to-end network to remedy these problems separately. Specifically, we build a scale-invariant wise-skip ResNet as the baseline to extract more context information. We present a receptive field expansion module by using scale-wise convolution for removing banding artifacts. We design a blind-adaptive controller to provide a deterministic result meanwhile meeting the needs of the user’s preference. In experiments, we compare the restoration accuracy among our model and many different variants of restoration methods on our collected remote sensing image dataset. The proposed network achieves superior performance against state-of-the-art methods in terms of both quantitative metrics and visual quality. Code and models are available at: https://github.com/lmmhh/TBANet. Maomei Liu, Lijia Fan, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Movable Object Detection in Remote Sensing Images via Dynamic Automatic LearningabstractThe performance of deep networks for object detection in remote sensing images (RSIs) largely depends on the availability of large-scale training images whose labels are given at the bounding-box level through a labor-intensive manual labeling process. To alleviate the huge burden of providing bounding-box annotations manually for movable objects, we propose a new approach, called dynamic automatic learning (DAL), to progressively learn object detectors. Specifically, a novel initial annotation generation (IAG) strategy is first designed to produce bounding-box annotations for movable objects in multi-temporal remote sensing images. During this process, image-level labels need to be manually labeled for the generated candidates. Next, a detection network learns the detection knowledge from multi-temporal remote sensing images with bounding-box annotations and then transfers the knowledge to generate pseudo boxes for the unlabeled data. Finally, with these pseudo boxes, the object detector can be optimized for generating accurate pseudo boxes iteratively. Furthermore, we introduce a pseudo box filtering (PBF) strategy to purify the quality of pseudo boxes to obtain accurate supervision. Our experiments on the challenging NWPU VHR-10.v2 and DIOR datasets have demonstrated that our DAL approach can achieve competitive results compared to state-of-the-art methods. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Boundary-Aware Bilateral Fusion Network for Cloud DetectionabstractCloud detection is one of the key technologies in the field of remote sensing. Although extensive deep learning-based cloud detection methods achieve good performance, their detection results in confusing areas such as cloud boundaries and thin clouds are often not satisfactory due to the potential inter-class similarity and intra-class inconsistency of objects. To this end, we propose a Boundary-Aware Bilateral Fusion network (BABFNet), which effectively enhances cloud detection in confusing areas by introducing a boundary prediction branch as an auxiliary. To avoid the loss of details, the boundary prediction branch is designed to run at full resolution with a shallow architecture, while some Semantic Enhancement Modules (SEMs) are used to supplement high-level semantic information by introducing multi-level encoder features of the cloud detection branch. This feature sharing in turn drives the cloud detection branch to focus more on cloud boundaries during training. At the end of the network, a Bilateral Fusion Module (BFM) is added for information complementarity between features from these two branches. The features from the cloud detection branch provide multi-scale features to the boundary prediction branch for more accurate boundary prediction, while the features from the boundary prediction branch further serve as prior knowledge to help the cloud detection branch aggregate contextual information. To verify the effectiveness of the proposed method, we select four different networks as cloud detection branches and conduct comparative experiments on two public datasets, GF-1 WFV and MODIS. The experimental results show that the proposed method significantly enhances cloud detection in confusing areas. Xiang Zhang 0018, Nailiang Kuang, Hangzai Luo, Sheng Zhong 0006, Jianping Fan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Class Guided Channel Weighting Network for Fine-Grained Semantic SegmentationabstractDeep learning has achieved promising performance on semantic segmentation, but few works focus on semantic segmentation at the fine-grained level. Fine-grained semantic segmentation requires recognizing and distinguishing hundreds of sub-categories. Due to the high similarity of different sub-categories and large variations in poses, scales, rotations, and color of the same sub-category in the fine-grained image set, the performance of traditional semantic segmentation methods will decline sharply. To alleviate these dilemmas, a new approach, named Class Guided Channel Weighting Network (CGCWNet), is developed in this paper to enable fine-grained semantic segmentation. For the large intra-class variations, we propose a Class Guided Weighting (CGW) module, which learns the image-level fine-grained category probabilities by exploiting second-order feature statistics, and use them as global information to guide semantic segmentation. For the high similarity between different sub-categories, we specially build a Channel Relationship Attention (CRA) module to amplify the distinction of features. Furthermore, a Detail Enhanced Guided Filter (DEGF) module is proposed to refine the boundaries of object masks by using an edge contour cue extracted from the enhanced original image. Experimental results on PASCAL VOC 2012 and six fine-grained image sets show that our proposed CGCWNet has achieved state-of-the-art results. Xiang Zhang 0018, Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
AAAI | 3 |
| 2022 | Automatic learning for object detection
Xiang Zhang 0018, Hangzai Luo, Wanqing Zhao, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 3 |
| 2022 | Detail-Aware Multiscale Context Fusion Network for Cloud DetectionabstractIn recent years, a large number of convolutional neural networks-based cloud detection algorithms have been proposed for remote sensing image preprocessing and most of them have an encoder-decoder structure. However, downsampling and upsampling operations, as the basic components of these methods, inevitably lead to the loss of detailed information in high-level features, which affects cloud detection performance. At the same time, the physical characteristics of the cloud, such as the variable size and irregular structure, also put forward requirements for the multi-scale feature representation ability of the network. To this end, we propose a novel cloud detection network named DMNet, which contains a Dense Feature Enhancement Module (DFEM) and a Multi-scale Context Fusion Spatial Attention Module (MCFSAM). DFEM aims to achieve information complementarity by exploiting the different properties of the features at different levels of the encoder, so as to strengthen the detailed information of high-level features and make the low-level features have more semantics. MCFSAM introduces a Multi-scale Context Fusion Block (MCFB) in spatial attention, which enables the network to densely capture contextual information at different scales and further emphasize useful features in the spatial dimension. Extensive experiments on GF-1 wide field-of-view satellite imagery (GF-1 WFV) dataset and Moderate-Resolution Imaging Spectroradiometer (MODIS) dataset demonstrate that our method outperforms other state-of-the-art cloud detection algorithms. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Hierarchical bilinear convolutional neural network for image classificationabstractAbstract Image classification is one of the mainstream tasks of computer vision. However, the most existing methods use labels of the same granularity level for training. This leads to ignoring the hierarchy that may help to differentiate different visual objects better. Embedding hierarchical information into the convolutional neural networks (CNNs) can effectively regulate the semantic space and thus reduce the ambiguity of prediction. To this end, a multi‐task learning framework, named as Hierarchical Bilinear Convolutional Neural Network (HB‐CNN), is developed by seamlessly integrating CNNs with multi‐task learning over the hierarchical visual concept structures. Specifically, the labels with a tree structure are used as the supervision to hierarchically train multiple branch networks. In this way, the model can not only learn additional information (e.g. context information) as the coarse‐level category features, but also focus the learned fine‐level category features on the object properties. To smoothly pass hierarchical conceptual information and encourage feature reuse, a connectivity pattern is proposed to connect features at different levels. Furthermore, a bilinear module is embedded to generalise various orderless texture feature descriptors so that our model can capture more discriminative features. The proposed method is extensively evaluated on the CIFAR‐10, CIFAR‐100, and ‘Orchid’ Plant image sets. The experimental results show the effectiveness and superiority of our method. Xiang Zhang 0018, Hangzai Luo, Sheng Zhong 0006, Ziyu Guan, Long Chen 0007, Jinye Peng 0001, Jianping Fan 0001 |
IET Comput. Vis. | 3 |
| 2021 | Learning noise-decoupled affine models for extreme low-light image enhancement
Maomei Liu, Sheng Zhong 0006, Hangzai Luo, Jinye Peng 0001 |
Neurocomputing | 4 |
| 2021 | Deep Multiple Instance Hashing for Fast Multi-Object Image SearchabstractMulti-keyword query is widely supported in text search engines. However, an analogue in image retrieval systems, multi-object query, is rarely studied. Meanwhile, traditional object-based image retrieval methods often involve multiple steps separately. In this work, we propose a weakly-supervised Deep Multiple Instance Hashing (DMIH) approach for multi-object image retrieval. Our DMIH approach, which leverages a popular CNN model to build the end-to-end relation between a raw image and the binary hash codes of its multiple objects, can support multi-object queries effectively and integrate object detection with hashing learning seamlessly. We treat object detection as a binary multiple instance learning (MIL) problem and such instances are automatically extracted from multi-scale convolutional feature maps. We also design a conditional random field (CRF) module to capture both the semantic and spatial relations among different class labels. For hashing training, we sample image pairs to learn their semantic relationships in terms of hash codes of the most probable proposals for owned labels as guided by object predictors. The two objectives benefit each other in a multi-task learning scheme. Finally, a two-level inverted index method is proposed to further speed up the retrieval of multi-object queries. Our DMIH approach outperforms state-of-the-arts on public benchmarks for object-based image retrieval and achieves promising results for multi-object queries. Wanqing Zhao, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | 6D object pose estimation via viewpoint relation reasoning
Wanqing Zhao, Shaobo Zhang 0006, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 4 |
| 2019 | Multilayer feature descriptors fusion CNN models for fine-grained visual recognitionabstractAbstract Fine‐grained image classification is a challenging topic in the field of computer vision. General models based on first‐order local features cannot achieve acceptable performance because the features are not so efficient in capturing fine‐grained difference. A bilinear convolutional neural network (CNN) model exhibits that a second‐order statistical feature is more efficient in capturing fine‐grained difference than a first‐order local feature. However, this framework only considers the extraction of a second‐order feature descriptor, using a single convolutional layer. The potential effective classification features of other convolutional layers are ignored, resulting in loss of recognition accuracy. In this paper, a multilayer feature descriptors fusion CNN model is proposed. It fully considers the second‐order feature descriptors and the first‐order local feature descriptor generated by different layers. Experimental verification was carried out on fine‐grained classification benchmark data sets, CUB‐200‐2011, Stanford Cars, and FGVC‐aircraft. Compared with the bilinear CNN model, the proposed method has improved accuracy by 0.8%, 1.1%, and 5.5%. Compared with the compact bilinear pooling model, there is an accuracy increase of 0.64%, 1.63%, and 1.45%, respectively. In addition, the proposed model effectively uses multiple 1×1 convolution kernels to reduce dimension. The experimental results show that the multilayer low‐dimensional second‐order feature descriptors fusion model has comparable recognition accuracy of the original model. Yong Hou, Hangzai Luo, Wanqing Zhao, Xiang Zhang 0018, Jun Wang 0078, Jinye Peng 0001 |
Comput. Animat. Virtual Worlds | 2 |
| 2019 | Plant recognition via leaf shape and margin features
Xiang Zhang 0018, Wanqing Zhao, Hangzai Luo, Long Chen 0007, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Locally linear spatial pyramid hash for large-scale image search
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Deep Multiple Instance Hashing for Object-based Image RetrievalabstractMulti-keyword query is widely supported in text search engines. However, an analogue in image retrieval systems, multi-object query, is rarely studied. Meanwhile, traditional object-based image retrieval methods often involve multiple steps separately and need expensive location labeling for detecting objects. In this work, we propose a weakly-supervised Deep Multiple Instance Hashing (DMIH) framework for object-based image retrieval. DMIH integrates object detection and hashing learning on the basis of a popular CNN model to build the end-to-end relation between a raw image and the binary hashing codes of multiple objects in it. Specifically, we cast the object detection of each object class as a binary multiple instance learning problem where instances are object proposals extracted from multi-scale convolutional feature maps. For hashing training, we sample image pairs to learn their semantic relationships in terms of hash codes of the most probable proposals for owned labels as guided by object predictors. The two objectives benefit each other in learning. DMIH outperforms state-of-the-arts on public benchmarks for object-based image retrieval and achieves promising results for multi-object queries. Wanqing Zhao, Ziyu Guan, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
IJCAI | 3 |
| 2017 | Spatial pyramid deep hashing for large-scale image retrieval
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Neurocomputing | 2 |
| 2017 | MapReduce-based clustering for near-duplicate image identification
Wanqing Zhao, Hangzai Luo, Jinye Peng 0001, Jianping Fan 0001 |
Multim. Tools Appl. | 2 |
| 2016 | A novel framework for semantic entity identification and relationship integration in large scale text data
Dingxian Wang, Xiao Liu 0004, Hangzai Luo, Jianping Fan 0001 |
Future Gener. Comput. Syst. | 3 |
| 2016 | Statistical modeling for automatic image indexing and retrieval
Baopeng Zhang, Hangzai Luo, Jianping Fan 0001 |
Neurocomputing | 2 |
| 2013 | Semantic Entity Identification in Large Scale Data via Statistical Features and DT-SVM
Dingxian Wang, Xiao Liu 0004, Hangzai Luo, Jianping Fan 0001 |
WISE (1) | 3 |
| 2011 | Saliency-based visualization for image searchabstractIn this paper, we propose a novel algorithm for improving and visualizing image search results. The proposed algorithm improves user's image search experience by three steps: (1) re-rank the initial image search results by the random walk refinement based on visual consistency and saliency cues, (2) project the re-ranked images into a 2-dimentional panel according to their saliency information and correlations, (3) detect and extract the saliency regions in each image for final visualization. To evaluate the performance of our algorithm, user study has been conducted. Experimental results demonstrate that our visualization algorithm provides more pleasing image search experience than the conventional image search methods. Jiajie Hu, Bin Jin, Weiyao Lin, Hangzai Luo, Zhenzhong Chen 0001, Hongxiang Li 0001 |
MMSP | 5 |
| 2011 | A new package-group-transmission-based algorithm for human activity recognition in videosabstractIn this paper, a new package-group-transmission-based algorithm is proposed for human activity recognition in videos. The proposed algorithm first models the entire scene as a network where each node in the network corresponds to a segmentation of the scene. Based on this network, we further model people in the scene as groups of packages. Thus, various human activities can be modeled as the process of "package group transmission" in the network and these activities can be efficiently recognized by suitably analyzing the "package transmission" process. Our proposed algorithm can not only detect activities under the challenging multiple camera scenario, but also be able to recognize various complex group activities among people. Experimental results demonstrate the effectiveness of our proposed algorithm. Yuanzhe Chen, Weiyao Lin, Hongxiang Li 0001, Hangzai Luo, Yisi Tao, Donghua Liu |
VCIP | 4 |
| 2011 | Multimedia news exploration and retrieval by integrating keywords, relations and visual features
Hangzai Luo, Jianping Fan 0001, Youjie Zhou |
Multim. Tools Appl. | 1 |
| 2010 | Semantic Entity-Relationship Model for Large-Scale Multimedia News Exploration and Recommendation
Hangzai Luo, Jianping Fan 0001 |
MMM | 1 |
| 2010 | Generating Visual Concept Network from Large-Scale Weakly-Tagged Images
Chunlei Yang, Hangzai Luo, Jianping Fan 0001 |
MMM | 2 |
| 2010 | Semantic Entity Detection by Integrating CRF and SVM
Peng Cai 0001, Hangzai Luo, Aoying Zhou |
WAIM | 2 |
| 2010 | Constructing distributed hippocratic video databases for privacy-preserving online patient training and counselingabstractDigital video now plays an important role in supporting more profitable online patient training and counseling, and integration of patient training videos from multiple competitive organizations in the health care network will result in better offerings for patients. However, privacy concerns often prevent multiple competitive organizations from sharing and integrating their patient training videos. In addition, patients with infectious or chronic diseases may not want the online patient training organizations to identify who they are or even which video clips they are interested in. Thus, there is an urgent need to develop more effective techniques to protect both video content privacy and access privacy . In this paper, we have developed a new approach to construct a distributed Hippocratic video database system for supporting more profitable online patient training and counseling. First, a new database modeling approach is developed to support concept-oriented video database organization and assign a degree of privacy of the video content for each database level automatically. Second, a new algorithm is developed to protect the video content privacy at the level of individual video clip by filtering out the privacy-sensitive human objects automatically. In order to integrate the patient training videos from multiple competitive organizations for constructing a centralized video database indexing structure, a privacy-preserving video sharing scheme is developed to support privacy-preserving distributed classifier training and prevent the statistical inferences from the videos that are shared for cross-validation of video classifiers. Our experiments on large-scale video databases have also provided very convincing results. Jinye Peng 0001, Noboru Babaguchi, Hangzai Luo, Yuli Gao, Jianping Fan 0001 |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2009 | Extracting informative images from web news pages via imbalanced classificationabstractIn this paper we propose an imbalanced classification algorithm to extract informative images from web news pages. Our algorithm resolve the difficult problem based on two approaches. First, we limit our dataset to a specific application area so that the patterns of the informative images can be captured by existing classification algorithms. Second, we propose an automatic negative samples filtering algorithm to eliminate most negative samples, so that the classification training data is rebalanced. Because most classification algorithms have reduced performance on imbalanced training data, our algorithm improves the overall performance significantly. In addition, our approach is inherently robust to new web sites and style/layout change of web sites. Hangzai Luo, Jianping Fan 0001 |
ACM Multimedia | 2 |
| 2009 | Incorporating camera metadata for attended region detection and consumer photo classificationabstractPhotos taken by human beings significantly differ from the pictures that are taken by a surveillance camera or a vision sensor on a robot, e.g., human beings may intentionally capture photos to express his/her feeling or record a memorial scene. Such a creative photo capture process is accomplished by adjusting two factors: (1) the parameters setting of a camera; and (2) the position between the camera and the interesting objects or scenes. To enable automatic understanding and interpretation of the semantics of photos, it is very important to take all these factors into account. Unfortunately, most existing algorithms for image understanding focus on only the content of the images while completely ignoring these two important factors. In this paper, we have developed a new algorithm to calculate what the interestingness of the photographer is and what the core content of a photo is. The gained information (i.e., attended regions and attention of the photographer) is further used to support more effective photo classification and retrieval. Our experiments on 70,000+ photos taken by 200+ different models of cameras have obtained very positive results. Hangzai Luo, Jianping Fan 0001 |
ACM Multimedia | 2 |
| 2009 | Personalized Image Recommendation
Yuli Gao, Hangzai Luo, Jianping Fan 0001 |
MMM | 2 |
| 2009 | Personalized News Video Recommendation
Hangzai Luo, Jianping Fan 0001, Daniel A. Keim, Shin'ichi Satoh 0001 |
MMM | 1 |
| 2009 | A distributed approach to enabling privacy-preserving model-based classifier training
Hangzai Luo, Jianping Fan 0001, Xiaodong Lin 0004, Aoying Zhou, Elisa Bertino |
Knowl. Inf. Syst. | 1 |
| 2009 | JustClick: Personalized Image Recommendation via Exploratory Search From Large-Scale Flickr ImagesabstractIn this paper, we have developed a novel framework calledJustClickto enable personalized image recommendation via exploratory search from large-scale collections of Flickr images. First, a topic network is automatically generated to summarize large-scale collections of Flickr images at a semantic level. Hyperbolic visualization is further used to enable interactive navigation and exploration of the topic network, so that users can gain insights of large-scale image collections at the first glance, build up their mental query models interactively and specify their queries (i.e., image needs) more precisely by selecting the image topics on the topic network directly. Thus, our personalized query recommendation framework can effectively address both the problem of query formulation and the problem of vocabulary discrepancy and null returns. Second, a small set of most representative images are recommended for the given image topic according to their representativeness scores. Kernel principal component analysis and hyperbolic visualization are seamlessly integrated to organize and layout the recommended images (i.e., most representative images) according to their nonlinear visual similarity contexts, so that users can assess the relevance between the recommended images and their real query intentions interactively. An interactive interface is implemented to allow users to express their time-varying query intentions precisely and to direct ourJustClicksystem to more relevant images according to their personal preferences. Our experiments on large-scale collections of Flickr images show very positive results. Jianping Fan 0001, Daniel A. Keim, Yuli Gao, Hangzai Luo, Zongmin Li |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | An Interactive Approach for Filtering Out Junk Images From Keyword-Based Google Search ResultsabstractThe keyword-based Google images search engine is now becoming very popular for online image search. Unfortunately, only the text terms that are explicitly or implicitly linked with the images are used for image indexing but the associated text terms may not have exact correspondence with the underlying image semantics, thus the keyword-based Google images search engine may return large amounts of junk images which are irrelevant to the given keyword-based queries. Based on this observation, we have developed an interactive approach to filter out the junk images from keyword-based Google images search results and our approach consists of the following major components. a) A kernel-based image clustering technique is developed to partition the returned images into multiple clusters and outliers. b) Hyperbolic visualization is incorporated to display large amounts of returned images according to their nonlinear visual similarity contexts, so that users can assess the relevance between the returned images and their real query intentions interactively and select one or multiple images to express their query intentions and personal preferences precisely. c) An incremental kernel learning algorithm is developed to translate the users' query intentions and personal preferences for updating the mixture-of-kernels and generating better hypotheses to achieve more accurate clustering of the returned images and filter out the junk images more effectively. Experiments on diverse keyword-based queries from Google images search engine have obtained very positive results. Our junk image filtering system is released for public evaluation at: http://www.cs.uncc.edu/~jfan/google-demo/. Yuli Gao, Jinye Peng 0001, Hangzai Luo, Daniel A. Keim, Jianping Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2008 | Personalized news video recommendationabstractIn this paper, we have developed an interactive system to enable personalized news video recommendation. First, multi-modal information channels (audio, video and closed captions) are seamlessly integrated and synchronized to achieve more reliable news topic detection, and the contextual relationships between the news topics are extracted automatically. Second, topic network and hyperbolic visualization are seamlessly integrated to achieve interactive navigation and exploration of large-scale collections of news videos at the topic level, so that users can have a good global overview of large-scale collections of news videos at the first glance. In such interactive topic network navigation and exploration process, the users' personal background knowledge can be taken into consideration for obtaining the news topics of interest interactively, building up their mental models of news needs precisely and formulating their searches easily by selecting the visible news topics on the screen directly. Our system can further recommend the relevant web news, the new search directions, and the most relevant news videos according to their importance and representativeness scores. Hangzai Luo, Jianping Fan 0001, Daniel A. Keim |
ACM Multimedia | 1 |
| 2008 | New Approach for Hierarchical Classifier Training and Multi-level Image Annotation
Jianping Fan 0001, Yuli Gao, Hangzai Luo, Shin'ichi Satoh 0001 |
MMM | 3 |
| 2008 | A Novel Approach for Filtering Junk Images from Google Search Results
Yuli Gao, Jianping Fan 0001, Hangzai Luo, Shin'ichi Satoh 0001 |
MMM | 3 |
| 2008 | Hypergraph partitioning for document clustering: a unified clique perspectiveabstractHypergraph partitioning has been considered as a promising method to address the challenges of high dimensionality in document clustering. With documents modeled as vertices and the relationship among documents captured by the hyperedges, the goal of graph partitioning is to minimize the edge cut. Therefore, the definition of hyperedges is vital to the clustering performance. While several definitions of hyperedges have been proposed, a systematic understanding of desired characteristics of hyperedges is still missing. To that end, in this paper, we first provide a unified clique perspective of the definition of hyperedges, which serves as a guide to define hyperedges. With this perspective, based on the concepts of hypercliques and shared (reverse) nearest neighbors, we propose three new types of clique hyperedges and analyze their properties regarding purity and size issues. Finally, we present an extensive evaluation using real-world document datasets. The experimental results show that, with shared (reverse) nearest neighbor based hyperedges, the clustering performance can be improved significantly in terms of various external validation measures without the need for fine tuning of parameters. Tianming Hu, Hui Xiong 0001, Wenjun Zhou 0001, Sam Yuan Sung, Hangzai Luo |
SIGIR | 5 |
| 2008 | Integrating multi-modal content analysis and hyperbolic visualization for large-scale news video retrieval and exploration
Hangzai Luo, Jianping Fan 0001, Shin'ichi Satoh 0001, Jiahang Yang, William Ribarsky |
Signal Process. Image Commun. | 1 |
| 2008 | Integrating Concept Ontology and Multitask Learning to Achieve More Effective Classifier Training for Multilevel Image AnnotationabstractIn this paper, we have developed a new scheme for achieving multilevel annotations of large-scale images automatically. To achieve more sufficient representation of various visual properties of the images, both the global visual features and the local visual features are extracted for image content representation. To tackle the problem of huge intraconcept visual diversity, multiple types of kernels are integrated to characterize the diverse visual similarity relationships between the images more precisely, and a multiple kernel learning algorithm is developed for SVM image classifier training. To address the problem of huge interconcept visual similarity, a novel multitask learning algorithm is developed to learn the correlated classifiers for the sibling image concepts under the same parent concept and enhance their discrimination and adaptation power significantly. To tackle the problem of huge intraconcept visual diversity for the image concepts at the higher levels of the concept ontology, a novel hierarchical boosting algorithm is developed to learn their ensemble classifiers hierarchically. In order to assist users on selecting more effective hypotheses for image classifier training, we have developed a novel hyperbolic framework for large-scale image visualization and interactive hypotheses assessment. Our experiments on large-scale image collections have also obtained very positive results. Jianping Fan 0001, Yuli Gao, Hangzai Luo |
IEEE Trans. Image Process. | 3 |
| 2008 | Mining Multilevel Image Semantics via Hierarchical ClassificationabstractIn this paper, we have proposed a novel framework for mining multilevel image semantics via hierarchical classification. To bridge the semantic gap more successfully, salient objects are used to characterize the intermediate image semantics effectively. The salient objects are defined as the connected image regions that capture the dominant visual properties linked to the corresponding physical objects in an image. To achieve a more reliable and tractable concept learning in high-dimensional feature space, a novel algorithm calledproduct of mixture-experts(PoM) is proposed to reduce the size of training images and speed up concept learning. A novel hierarchical concept learning algorithm is proposed by incorporating concept ontology and multitask learning to enhance the discrimination power of the concept models and reduce the computational complexity for learning the concept models for large amount of image concepts, which may have huge intra-concept variations and inter-concept similarities on their visual properties. A hyperbolic image visualization algorithm has been developed for allowing users to specify their queries easily and assess the query results interactively. Our experiments on large-scale image collections have also obtained very positive results. Jianping Fan 0001, Yuli Gao, Hangzai Luo, Ramesh Jain 0001 |
IEEE Trans. Multim. | 3 |
| 2008 | Incorporating feature hierarchy and boosting to achieve more effective classifier training and concept-oriented video summarization and skimmingabstractFor online medical education purposes, we have developed a novel scheme to incorporate the results of semantic video classification to select the most representative video shots for generating concept-oriented summarization and skimming ofsurgery education videos. First, salient objects are used as the video patterns for feature extraction to achieve a good representation of the intermediate video semantics. The salient objects are defined as the salient video compounds that can be used to characterize the most significant perceptual properties of the corresponding real world physical objects in a video, and thus the appearances of such salient objects can be used to predict the appearances of the relevant semantic video concepts in a specific video domain. Second, a novelmulti-modal boostingalgorithm is developed to achieve more reliable video classifier training by incorporating feature hierarchy and boosting to dramatically reduce both the training cost and the size of training samples, thus it can significantly speed up SVM (support vector machine) classifier training. In addition, the unlabeled samples are integrated to reduce the human efforts on labeling large amount of training samples. Finally, the results of semantic video classification are incorporated to enable concept-oriented video summarization and skimming. Experimental results in a specific domain ofsurgery education videosare provided. Hangzai Luo, Yuli Gao, Xiangyang Xue 0001, Jinye Peng 0001, Jianping Fan 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2007 | Hierarchical classification for automatic image annotationabstractIn this paper, a hierarchical classification framework has been proposed for bridging the semantic gap effectively and achieving multi-level image annotation automatically. First, the semantic gap between the low-level computable visual features and users' real information needs is partitioned into four smaller gaps, and multiple approachesallare proposed to bridge these smaller gaps more effectively. To learn more reliable contextual relationships between the atomic image concepts and the co-appearances of salient objects, a multi-modal boosting algorithm is proposed. To enable hierarchical image classification and avoid inter-level error transmission, a hierarchical boosting algorithm is proposed by incorporating concept ontology and multi-task learning to achieve hierarchical image classifier training with automatic error recovery. To bridge the gap between the computable image concepts and the users' real information needs, a novel hyperbolic visualization framework is seamlessly incorporated to enable intuitive query specification and evaluation by acquainting the users with a good global view of large-scale image collections. Our experiments on large-scale image databases have also obtained very positive results. Jianping Fan 0001, Yuli Gao, Hangzai Luo |
SIGIR | 3 |
| 2007 | Incorporating Concept Ontology for Hierarchical Video Classification, Annotation, and VisualizationabstractMost existing content-based video retrieval (CBVR) systems are now amenable to support automatic low-level feature extraction, but they still have limited effectiveness from a user's perspective because of the semantic gap. Automatic video concept detection via semantic classification is one promising solution to bridge the semantic gap. To speed up SVM video classifier training in high-dimensional heterogeneous feature space, a novel multimodal boosting algorithm is proposed by incorporating feature hierarchy and boosting to reduce both the training cost and the size of training samples significantly. To avoid the inter-level error transmission problem, a novel hierarchical boosting scheme is proposed by incorporating concept ontology and multitask learning to boost hierarchical video classifier training through exploiting the strong correlations between the video concepts. To bridge the semantic gap between the available video concepts and the users' real needs, a novel hyperbolic visualization framework is seamlessly incorporated to enable intuitive query specification and evaluation by acquainting the users with a good global view of large-scale video collections. Our experiments in one specific domain of surgery education videos have also provided very convincing results. Jianping Fan 0001, Hangzai Luo, Yuli Gao, Ramesh Jain 0001 |
IEEE Trans. Multim. | 2 |
| 2006 | Searching and browsing large scale image database using keywords and ontologyabstractAutomatic image annotation is a promising solution to enable keyword-based semantic image retrieval. In this demo, we present our image search engine I-Search implemented using a multi-level semantic image annotation algorithm. By incorporating a domain-specific ontology into the autogenerated annotations, we are able to organize large scale natural image databases into hierarchical structures for browsing and keyword-based searching without referring to external text information. Yuli Gao, Hangzai Luo, Jianping Fan 0001 |
ACM Multimedia | 2 |
| 2006 | Building concept ontology for medical video annotationabstractMost existing systems for content-based video retrieval (CBVR) are now amenable to support automatic low-level video content analysis and feature extraction, but they have limited effectiveness from a user's perspective. To support semantic video retrieval via keywords, we have proposed a novel framework by incorporating the concept ontology to enable more effective modeling and representation of semantic video concepts. Specifically, this novel framework includes: (a) Using the salient objects to achieve a middle-level understanding of the semantics of video contents; (b) Building a domain dependent concept ontology to enable multi-level modeling and representation of semantic video concepts; (c) Developing a multi-task boosting technique to achieve hierarchical video classifier training for automatic multi-level video annotation. The experimental results in a certain domain of medical education videos are also provided. Hangzai Luo, Jianping Fan 0001 |
ACM Multimedia | 1 |
| 2006 | Large-scale video retrieval via semantic classificationabstractMotivated by Google's great success on text document retrieval and recent progresses of semantic video understanding, researchers begin to build new generation of video retrieval systems that are able to support semantic sensitive video retrieval via keywords. Unfortunately, these systems are not able to provide satisfactory results for the masses because of several inter-related challenging problems. We have proposed novel algorithms to resolve some of these problems. Firstly, the salient object based semantic classification algorithm is proposed to extract semantic concepts of video clips. Secondly, the video visualization based interactive retrieval framework is proposed to help users input semantic and visual queries efficiently and effectively. Finally, the concept-oriented skimming algorithm is proposed to help users efficiently check search results. Hangzai Luo, Jianping Fan 0001 |
ACM Multimedia | 1 |
| 2006 | Large-scale news video retrieval via visualizationabstractAs the content of everyday news reports is unpredictable, keyword based news search engine can't provide effective services to audiences because the audiences may not be able to figure out proper keywords to search. In this paper, a novel framework is proposed to help audiences browse and retrieve news video clips without the need of keywords. Interesting keyframes and keywords are automatically extracted from news video clips and visually represented according to their interestingness and informativeness measurement. A computational approach is also developed to quantify the interestingness measurement of video clips. The keyframes and keywords are carefully organized so that the audiences can find news stories of interest at first glance. Hangzai Luo, Jianping Fan 0001, Yuli Gao, William Ribarsky, Shin'ichi Satoh 0001 |
ACM Multimedia | 1 |
| 2005 | A novel approach for privacy-preserving video sharingabstractInternational audience Jianping Fan 0001, Hangzai Luo, Mohand-Said Hacid, Elisa Bertino |
CIKM | 2 |
| 2005 | Learning the Semantics of Images by Using Unlabeled SamplesabstractIn this paper, we have proposed a novel framework to achieve more effective classifier training by using unlabeled samples. By integrating concept hierarchy for semantic image concept organization, a hierarchical mixture model is proposed to enable multi-level image concept modeling and hierarchical classifier training. To effectively learn the base-level classifiers for the atomic image concepts at the first level of the concept hierarchy, we have proposed a novel adaptive EM algorithm to achieve more effective classifier training with higher prediction accuracy. To effectively learn the classifiers for the higher-level semantic image concepts, we have also proposed a novel technique for classifier combining by using hierarchical mixture model. The experimental results on two large-scale image databases are also provided. Jianping Fan 0001, Hangzai Luo, Yuli Gao |
CVPR (2) | 2 |
| 2005 | Mining images on semantics via statistical learningabstractInternational audience Jianping Fan 0001, Hangzai Luo, Mohand-Said Hacid |
KDD | 2 |
| 2005 | Statistical modeling and conceptualization of natural images
Jianping Fan 0001, Yuli Gao, Hangzai Luo, Guangyou Xu |
Pattern Recognit. | 3 |
| 2004 | Multi-level annotation of natural scenes using dominant image components and semantic conceptsabstractAutomatic image annotation is a promising solution to enable semantic image retrieval via keywords. In this paper, we propose a multi-level approach to annotate the semantics of natural scenes by using both the dominant image components (salient objects) and the relevant semantic concepts. To achieve automatic image annotation at the content level, we use salient objects as the dominant image components for image content representation and feature extraction. To support automatic image annotation at the concept level, a novel image classification technique is developed to map the images into the most relevant semantic image concepts. In addition, Support Vector Machine (SVM) classifiers are used to learn the detection functions for the pre-defined salient objects and finite mixture models are used for semantic concept interpretation and modeling. An adaptive EM algorithm has been proposed to determine the optimal model structure and model parameters simultaneously. We have also demonstrated that our algorithms are very effective to enable multi-level annotation of natural scenes in a large-scale image dataset. Jianping Fan 0001, Yuli Gao, Hangzai Luo |
ACM Multimedia | 3 |
| 2004 | Concept-oriented video skimming via semantic video classificationabstractEffective video skimming requires a good understanding of the semantics of video contents. However, more existing systems for content-based video retrieval (CBVR) can only support low-level video analysis, but they have limited effectiveness on achieving semantic-sensitive video understanding. In this paper, we have developed a novel framework to achieve concept-oriented video skimming and it consists of three parts: (a) using salient objects for semantic-sensitive video content representation; (b) using finite mixture models for semantic video concept modeling and classification; (c) enabling concept-oriented video skimming via semantic video classification. Hangzai Luo, Jianping Fan 0001 |
ACM Multimedia | 1 |
| 2004 | Automatic image annotation by using concept-sensitive salient objects for image content representationabstractMulti-level annotation of images is a promising solution to enable more effective semantic image retrieval by using various keywords at different semantic levels. In this paper, we propose a multi-level approach to annotate the semantics of natural scenes by using both the dominant image components and the relevant semantic concepts. In contrast to the well-known image-based and region-based approaches, we use the salient objects as the dominant image components to achieve automatic image annotation at the content level. By using the salient objects for image content representation, a novel image classification technique is developed to achieve automatic image annotation at the concept level. To detect the salient objects automatically, a set of detection functions are learned from the labeled image regions by using Support Vector Machine (SVM) classifiers with an automatic scheme for searching the optimal model parameters. To generate the semantic concepts, finite mixture models are used to approximate the class distributions of the relevant salient objects. An adaptive EM algorithm has been proposed to determine the optimal model structure and model parameters simultaneously. We have also demonstrated that our algorithms are very effective to enable multi-level annotation of natural scenes in a large-scale dataset. Jianping Fan 0001, Yuli Gao, Hangzai Luo, Guangyou Xu |
SIGIR | 3 |
| 2004 | Semantic video classification by integrating unlabeled samples for classifier trainingabstractSemantic video classification has become an active research topic to enable more effective video retrieval and knowledge discovery from large-scale video databases. However, most existing techniques for classifier training require a large number of hand-labeled samples to learn correctly. To address this problem, we have proposed a semi-supervised framework to achieve incremental classifier training by integrating a limited number of labeled samples with a large number of unlabeled samples. Specifically, this emi-supervised framework includes: (a) Modeling the semantic video concepts by using the finite mixture models to approximate the class distributions of the relevant salient objects; (b) Developing an adaptive EM algorithm to integrate the unlabeled samples to achieve parameter estimation and model selection simultaneously; The experimental results in a certain domain of medical videos are also provided. Jianping Fan 0001, Hangzai Luo |
SIGIR | 2 |
| 2004 | Concept-oriented indexing of video databases: toward semantic sensitive retrieval and browsingabstractDigital video now plays an important role in medical education, health care, telemedicine and other medical applications. Several content-based video retrieval (CBVR) systems have been proposed in the past, but they still suffer from the following challenging problems: semantic gap, semantic video concept modeling, semantic video classification, and concept-oriented video database indexing and access. In this paper, we propose a novel framework to make some advances toward the final goal to solve these problems. Specifically, the framework includes: 1) a semantic-sensitive video content representation framework by using principal video shots to enhance the quality of features; 2) semantic video concept interpretation by using flexible mixture model to bridge the semantic gap; 3) a novel semantic video-classifier training framework by integrating feature selection, parameter estimation, and model selection seamlessly in a single algorithm; and 4) a concept-oriented video database organization technique through a certain domain-dependent concept hierarchy to enable semantic-sensitive video retrieval and browsing. Jianping Fan 0001, Hangzai Luo, Ahmed K. Elmagarmid |
IEEE Trans. Image Process. | 2 |
| 2003 | Semantic principal video shot classification via mixture GaussianabstractAs digital cameras become more affordable, digital video now plays an important role in medical education and healthcare. In this paper, we propose a novel framework to facilitate semantic classification of surgery education videos. Specifically, the framework includes: (a) semantic-sensitive video content characterization via principal video shots, (b) semantic video classification via a mixture Gaussian model to bridge the semantic gap between low-level visual features and semantic visual concepts in a specific surgery education video domain. Hangzai Luo, Jianping Fan 0001, Jing Xiao 0001, Xingquan Zhu 0001 |
ICME | 1 |