VLDB 2026 Research / reviewers in the wild / expert
Sujuan Hou
dblp:150/1756
· DBLP profile ↗
30ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ComAdPro: compositional learning with prototype adaptation for logo few-shot class-incremental recognition (ChinaMM 2025)
Jianxin Zhan, Wentai Chen, Sujuan Hou, Weiqing Min |
Multim. Syst. | 3 |
| 2026 | Large-Scale Logo DetectionabstractLogo detection is crucial for trademark compliance and media monitoring, enabling companies to monitor online trademark usage and evaluate brand visibility on social media and advertisements. The use of large datasets significantly improves accuracy and generalization, emphasizing the need for high-quality datasets to optimize performance and enhance reasoning abilities in visual detection models. This drove us to create Logo4500, an unparalleled dataset featuring 4,500 logo categories and over 293,000 meticulously labeled images. To ensure the dataset's quality, we meticulously designed the construction and annotation process, with detailed information provided in our paper. Compared to existing logo datasets, Logo4500 offers greater diversity and class imbalance, making it more reflective of real-world distribution. Leveraging this high-quality dataset, we introduce a benchmark called Frequency-Aware Learnable Dual Reweighting Network (FALDR-Net), which enhances the representation of ambiguous features and addresses class imbalance for large-scale logo detection. We conducted extensive experiments, evaluating various recent methods on this new dataset and several existing publicly available logo datasets, demonstrating its effectiveness. Additionally, we verified Logo4500's generalization ability in several tasks. We anticipate that Logo4500 and the benchmark will inspire further exploration in the logo-related research community, facilitating the advancement of visual foundation models. Sujuan Hou, Weiqing Min, Jianxin Zhan, Mengmeng Zhang 0008, Peng Li 0081, Shuqiang Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | CondFoodGen: A Conditional Two-Stream Network for Controllable Food Image GenerationabstractFood image generation is an important research direction in food computing, aiming to produce highly realistic images that accurately capture the visual characteristics of various dishes while adhering to specified input conditions. Existing methods that rely solely on textual descriptions struggle to handle the large intra-class variability of food, often resulting in limited diversity and accuracy. Although some approaches incorporate additional conditions, they generally lack optimizations for food-specific challenges, leading to inconsistencies in texture, shape, and color fidelity. To address these limitations, we propose CondFoodGen, a diffusion-based two-stream network for controllable food image generation. The architecture consists of a control stream and a generation stream, where the control stream provides conditional guidance to regulate the generation process. To optimize bidirectional interactions between the two streams, we introduce the Bidirectional Adaptive Gating (BAG) mechanism, which not only guides synthesis but also adaptively refines control representations through feedback from the generation stream. In addition, we propose the Wavelet-Guided Hierarchical Attention (WGHA) module, which combines wavelet-based multi-frequency analysis with hierarchical attention to enhance fine-grained texture fidelity and structural realism. A progressive multi-stage training strategy further stabilizes optimization and enables seamless integration of conditional guidance with bidirectional interaction. Extensive experiments on three food image datasets demonstrate that CondFoodGen consistently generates high-quality and diverse images. Compared with the best existing food image generation methods, our approach achieves an average improvement of about 11.0% across three evaluation metrics and compared to the leading conditional generation approaches, the average improvement reaches 16.2%. The source code, trained models, and supplementary materials are publicly available at https://github.com/housujuan123/CondFoodGen. Mengyao Zhao, Hao Xiong 0001, Weiqing Min, Sujuan Hou, Mengmeng Zhang 0008, Shuqiang Jiang |
IEEE Trans. Image Process. | 4 |
| 2025 | DSDGF-Nutri: A Decoupled Self-Distillation Network with Gating Fusion For Food Nutritional AssessmentabstractAccurate assessment of food nutrition is essential for promoting healthy eating habits. While recent deep learning approaches have enhanced vision-based nutritional estimation through RGB-D multi-modal fusion, they often overlook fine-grained surface components (e.g., oil and sugar) that significantly influence nutritional values. Some recent approaches have improved accuracy by incorporating ingredient data, but their reliance on such input during inference limits practical applicability, as ingredient details are often unavailable in real-world settings. To address this limitation, we propose DSDGF-Nutri, a novel Decoupled Self-Distillation network with Gating Fusion for food Nutri tional assessment. Our method leverages ingredient knowledge during training but relies solely on RGB-D inputs at inference. Specifically, DSDGF-Nutri introduces: (1) a self-distillation mechanism with gating fusion that transfers ingredient-aware features to the RGB-D network, enabling robust prediction without test-time ingredient input, and (2) a multi-task decoupling architecture with task-specific decoders to minimize cross-task interference. Extensive evaluations on two benchmark datasets demonstrate DSDGF-Nutri outperforms existing methods, achieving state-of-the-art results. This work establishes a new paradigm of multimodal fusion in nutritional assessment by unifying scientific measurements with scalable computer vision applications. Sujuan Hou, Zhihui Feng, Hao Xiong 0001, Weiqing Min, Peng Li 0081, Shuqiang Jiang |
ACM Multimedia | 1 |
| 2024 | Multi-Stage Progressive Refinement and RoI Context Enhancement Network for Small Logo DetectionabstractLogo detection is a critical task in computer vision with a wide range of applications. In logo detection tasks, small logos occupy a limited number of pixels in the image due to their small size. Additionally, background clutter may have textures, colors, or shapes that are similar to small logos, further making it difficult for the detection algorithm to distinguish between the object and the background. To address this problem, we propose a Multi-stage Progressive Refinement and RoI (Region of Interest) Context Enhancement Network (MPRRCENet) for small logo detection. Specifically, a Multi-stage Progressive Refinement (MPR) module is proposed to progressive refine features with discriminative capability and promote the interaction of feature information. Furthermore, a RoI Context Enhancement (RCE) module is proposed that utilizes contextual information and channel modeling to enhance RoI features. Extensive experiments on four publicly available logo datasets demonstrate the effectiveness of our proposed method. Songhui Zhao, Sujuan Hou |
ICASSP | 2 |
| 2024 | RD-FGM: A novel model for high-quality and diverse food image generation and ingredient classification
Jing Wang 0138, Yuanjie Zheng, Junxia Wang, Sujuan Hou |
Expert Syst. Appl. | 6 |
| 2024 | Context-based modeling for accurate logo detection in complex environments
Zhixiang Jia, Sujuan Hou, Peng Li 0081 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Cross-View Representation Learning: A Superior ContextIB Method for Logo ClassificationabstractLogo classification systems have become increasingly important in various industries for tasks, such as infringement detection and industrial production. However, challenges still exist in logo classification due to real-world image background interference, the high similarity between classes, labeling difficulties, and the insufficient representation of occlusion in single-view logos. Many existing algorithms fail to consider the data characteristics and the intrinsic information of multiple views, which limits their performance. To overcome these limitations, we developed a novel Cross-View Information Awareness Network (CVIA-Net) for logo classification. To differentiate between similar logo categories, the CVIA-Net novel learns context-shared features of the same category via a self-supervised way without labeled, which solves the problem of insufficient features due to occlusion. For single-view images, CVIA-Net establishes a “bottleneck” representation to address background interference. Extensive experiments on three datasets demonstrate that it outperforms state-of-the-art methods. The method is expected to advance the development of cross-view representation learning. Jing Wang 0138, Yuanjie Zheng, Zeyu Han, Mei Lv, Sujuan Hou |
IEEE Signal Process. Lett. | 5 |
| 2024 | Deep Learning for Logo Detection: A SurveyabstractLogo detection has gradually become a research hotspot in the field of computer vision and multimedia for its various applications, such as social media monitoring, intelligent transportation, and video advertising recommendation. Recent advances in this area are dominated by deep learning-based solutions, where many datasets, learning strategies, network architectures, and loss functions have been employed. This article reviews the advance in applying deep learning techniques to logo detection. First, we discuss a comprehensive account of public datasets designed to facilitate performance evaluation of logo detection algorithms, which tend to be more diverse, more challenging, and more reflective of real life. Next, we perform an in-depth analysis of the existing logo detection strategies and their strengths and weaknesses of each learning strategy. Subsequently, we summarize the applications of logo detection in various fields, from intelligent transportation and brand monitoring to copyright and trademark compliance. Finally, we analyze the potential challenges and present the future directions for the development of logo detection. This study aims better to inform readers about the current state of logo detection and encourage more researchers to get involved in logo detection. Sujuan Hou, Weiqing Min, Yanna Zhao, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | A Cross-direction Task Decoupling Network for Small Logo DetectionabstractLogo detection plays an integral role in many applications. However, handling small logos is still difficult since they occupy too few pixels in the image, which burdens the extraction of discriminative features. The aggregation of small logos also brings a great challenge to the classification and localization of logos. To solve these problems, we creatively propose Cross-direction Task Decoupling Network (CTDNet) for small logo detection. We first introduce Cross-direction Feature Pyramid (CFP) to realize cross-direction feature fusion by adopting horizontal transmission and vertical transmission. In addition, Multi-frequency Task Decoupling Head (MTDH) decouples the classification and localization tasks into two branches. A multi-frequency attention convolution branch is designed to achieve more accurate regression by combining discrete cosine transform and convolution creatively. Comprehensive experiments on four logo datasets demonstrate the effectiveness and efficiency of the proposed method. Sujuan Hou, Xingzhuo Li, Weiqing Min, Jing Wang 0138, Yuanjie Zheng, Shuqiang Jiang |
ICME | 1 |
| 2023 | A Decoupled Cross-layer Fusion Network with Bidirectional Guidance for Detecting Small LogosabstractLogo detection involves the use of machine learning algorithms to recognize and locate logos in images and videos, which has applications in a wide range of industries, including e-commerce, advertising, and entertainment. However, detecting small logos is still a challenging task due to their limited coverage of pixels and unclear details resulting in insufficient feature information for detection. Therefore, they are often easily confused by complex backgrounds and have lower perturbation tolerance to the bounding box, making them more difficult to detect compared to medium and large-scale logos. To address this problem, we propose a Decoupled Cross-layer Fusion Network (DCFNet) that enhances the feature representation of small logo objects, resulting in excellent detection performance. Specifically, the proposed DCFNet first adopts a bidirectional cross-layer connection mechanism to capture complementary information between different layers. Next, a two-phase feature averaging and enhancement strategy is used to further enhance the features. In the detection phase, DCFNet decouples the classification and boundary box regression branches into two identical Fully Connected (FC) heads, improving the accuracy of small logo classification and localization by avoiding mutual interference between the branches. Extensive experiments conducted on three publicly available logo datasets demonstrate that DCFNet achieves state-of-the-art performance in detecting small logos. Songhui Zhao, Sujuan Hou, Baisong Zhang |
MMAsia | 2 |
| 2023 | Few-shot logo detectionabstractAbstract The proliferation of deep learning has driven research into deep learning‐based logo detection, which usually needs a large number of annotated data to train the model. However, due to the occasional appearance of new brands or the high cost of annotation, the number of training data is limited. Against this backdrop, the authors adapt the few‐shot object detection into logo detection, and thus present a cutting‐edge method called Double Classification Head (DCH) for Few‐Shot Logo Detection (DCH‐FSLogo), which aims at detecting the unseen logo classes using few annotated data. Unlike the traditional few‐shot detection, some logo objects are similar to their backgrounds and have diverse shapes as well. For this reason, the authors adopt balanced feature pyramid and deformable Region of Interest pooling in DCH‐FSLogo, this enhances the feature extraction capability and adapts to the different logo shapes. In addition, we introduce the DCH for few‐shot logo detection to detect logo objects using few annotated data. Specifically, we use an extra classification head for the base classes to ease the influence from the novel classes. The experimental results on four datasets, namely: FlickrLogos‐32, FoodLogoDet‐1500‐100, LogoDet‐3K‐100 and QMUL‐OpenLogo‐100, demonstrate that our method achieves better performance. Sujuan Hou, Wenjie Liu 0001, Karim Awudu, Zhixiang Jia, Weikuan Jia, Yuanjie Zheng |
IET Comput. Vis. | 1 |
| 2023 | Lightweight Seizure Detection Based on Multi-Scale Channel AttentionabstractEpilepsy is one kind of neurological disease characterized by recurring seizures. Recurrent seizures can cause ongoing negative mental and cognitive damage to the patient. Therefore, timely diagnosis and treatment of epilepsy are crucial for patients. Manual electroencephalography (EEG) signals analysis is time and energy consuming, making automatic detection using EEG signals particularly important. Many deep learning algorithms have thus been proposed to detect seizures. These methods rely on expensive and bulky hardware, which makes them unsuitable for deployment on devices with limited resources due to their high demands on computer resources. In this paper, we propose a novel lightweight neural network for seizure detection using pure convolutions, which is composed of inverted residual structure and multi-scale channel attention mechanism. Compared with other methods, our approach significantly reduces the computational complexity, making it possible to deploy on low-cost portable devices for seizures detection. We conduct experiments on the CHB-MIT dataset and achieves 98.7% accuracy, 98.3% sensitivity and 99.1% specificity with 2.68[Formula: see text]M multiply-accumulate operations (MACs) and only 88[Formula: see text]K parameters. Sujuan Hou, Tiantian Xiao, Yongfeng Zhang 0001, Hongbin Lv, Yanna Zhao |
Int. J. Neural Syst. | 2 |
| 2023 | Image Matting With Deep Gaussian ProcessabstractWe observe a common characteristic between the classical propagation-based image matting and the Gaussian process (GP)-based regression. The former produces closer alpha matte values for pixels associated with a higher affinity, while the outputs regressed by the latter are more correlated for more similar inputs. Based on this observation, we reformulate image matting as GP and find that this novel matting-GP formulation results in a set of attractive properties. First, it offers an alternative view on and approach to propagation-based image matting. Second, an application of kernel learning in GP brings in a novel deep matting-GP technique, which is pretty powerful for encapsulating the expressive power of deep architecture on the image relative to its matting. Third, an existing scalable GP technique can be incorporated to further reduce the computational complexity to$\mathcal {O}(n)$from$\mathcal {O}(n^{3})$of many conventional matting propagation techniques. Our deep matting-GP provides an attractive strategy toward addressing the limit of widespread adoption of deep learning techniques to image matting for which a sufficiently large labeled dataset is lacking. A set of experiments on both synthetically composited images and real-world images show the superiority of the deep matting-GP to not only the classical propagation-based matting techniques but also modern deep learning-based approaches. Yuanjie Zheng, Yunshuai Yang, Tongtong Che, Sujuan Hou, Wenhui Huang 0002, Yue Gao 0002, Ping Tan 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | LogoDet-3K: A Large-scale Image Dataset for Logo DetectionabstractLogo detection has been gaining considerable attention because of its wide range of applications in the multimedia field, such as copyright infringement detection, brand visibility monitoring, and product brand management on social media. In this article, we introduce LogoDet-3K, the largest logo detection dataset with full annotation, which has 3,000 logo categories, about 200,000 manually annotated logo objects, and 158,652 images. LogoDet-3K creates a more challenging benchmark for logo detection, for its higher comprehensive coverage and wider variety in both logo categories and annotated objects compared with existing datasets. We describe the collection and annotation process of our dataset and analyze its scale and diversity in comparison to other datasets for logo detection. We further propose a strong baseline method Logo-Yolo, which incorporates Focal loss and CIoU loss into the basic YOLOv3 framework for large-scale logo detection. It obtains about 4% improvement on the average performance compared with YOLOv3, and greater improvements compared with reported several deep detection models on LogoDet-3K. We perform extensive evaluation on three other existing datasets to further verify on both logo detection and retrieval tasks, and we demonstrate better generalization ability of LogoDet-3K on logo detection and retrieval tasks. The LogoDet-3K dataset is used to promote large-scale logo-related research. The code and LogoDet-3K can be found at https://github.com/Wangjing1551/LogoDet-3K-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Shuqiang Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2021 | Exploiting Probabilistic Siamese Visual Tracking with a Conditional Variational AutoencoderabstractVisual tracking is a fundamental capability for robots tasked with humans and environment interaction. However, state-of-the-art visual tracking methods are still prone to failures and are imprecise when applied to challenging stereos, and their results are generally confidence agonistic. These methods depend on an embedded deep learning model to provide deterministic features or regression maps. A deterministic output with low confidence can result in disastrous consequences and lacks evidence needed for subsequent operations. Moreover, training data ambiguities or noise in the observations (so-called data uncertainty) can also lead to inherent uncertainty. In this paper, we focus on exploiting probabilistic Siamese visual tracking with a conditional variational autoencoder (CVAE). First, we build a bridge between the Siamese architecture and the CVAE and propose a novel Bayesian visual tracking method. Second, the proposed method generates a complete probability distribution that enables the production of multiple plausible tracking outputs. Third, CVAE conditioned by ground truth data encodes a low-dimensional latent space and conducts noise-injection training to prevent overfitting. Our proposed tracking method outperformed the state-of-the-art trackers on the VOT2016, VOT2018 and TColor-128 datasets. Wenhui Huang 0002, Jason Gu, Peiyong Duan, Sujuan Hou, Yuanjie Zheng |
ICRA | 4 |
| 2021 | FoodLogoDet-1500: A Dataset for Large-Scale Food Logo Detection via Multi-Scale Feature Decoupling NetworkabstractFood logo detection plays an important role in the multimedia for its wide real-world applications, such as food recommendation of the self-service shop and infringement detection on e-commerce platforms. A large-scale food logo dataset is urgently needed for developing advanced food logo detection algorithms. However, there are no available food logo datasets with food brand information. To support efforts towards food logo detection, we introduce the dataset FoodLogoDet-1500, a new large-scale publicly available food logo dataset, which has 1,500 categories, about 100,000 images and about 150,000 manually annotated food logo objects. We describe the collection and annotation process of FoodLogoDet-1500, analyze its scale and diversity, and compare it with other logo datasets. To the best of our knowledge, FoodLogoDet-1500 is the first largest publicly available high-quality dataset for food logo detection. The challenge of food logo detection lies in the large-scale categories and similarities between food logo categories. For that, we propose a novel food logo detection method Multi-scale Feature Decoupling Network (MFDNet), which decouples classification and regression into two branches and focuses on the classification branch to solve the problem of distinguishing multiple food logo categories. Specifically, we introduce the feature offset module, which utilizes the deformation-learning for optimal classification offset and can effectively obtain the most representative features of classification in detection. In addition, we adopt a balanced feature pyramid in MFDNet, which pays attention to global information, balances the multi-scale feature maps, and enhances feature extraction capability. Comprehensive experiments on FoodLogoDet-1500 and other two popular benchmark logo datasets demonstrate the effectiveness of the proposed method. The code and FoodLogoDet-1500 can be found at https://github.com/hq03/FoodLogoDet-1500-Dataset. Weiqing Min, Jing Wang 0138, Sujuan Hou, Yuanjie Zheng, Shuqiang Jiang |
ACM Multimedia | 4 |
| 2021 | Cross-View Representation Learning for Multi-View Logo Classification with Information BottleneckabstractMulti-view logo classification is a challenging task due to the cross-view misalignment of logo image varies under different viewpoints, large intra-classes and small inter-classes variation of logo appearance. Cross-view data can represent objects from different views and thus provide complementary information for data analysis. However, most existing multi-view algorithms usually maximize the correlation between different views for consistency. Those methods ignore the interaction among different views and may cause semantic bias during the process of common feature learning. In this paper, we investigate the information bottleneck (IB) to the multi-view learning for extracting the different view common features of one category, named Dual-View Information Bottleneck representation (Dual-view IB). To the best of our knowledge, this is the first cross-view learning method for logo classification. Specifically, we maximize the mutual information between the representations of the two views to achieve the preservation of key features in the classification task, while eliminating the redundant information that is not shared between the two views. In addition, due to the unbalance of samples and limited computing resources, we further introduce a novel Pair Batch Data Augmentation (PB) algorithm for Dual-view IB model, which applies augmentations from a learned policy based on replicates instances of two samples within the same batch. Comprehensive experiments on three existing benchmark datasets, which demonstrate the effectiveness of the proposed method that outperforms the methods in the state of the art. The proposed method is expected to further the development of cross-view representation learning. Jing Wang 0138, Yuanjie Zheng, Jingqi Song, Sujuan Hou |
ACM Multimedia | 4 |
| 2021 | Kirsch Direction Template Despeckling Algorithm of High-Resolution SAR Images-Based on Structural Information DetectionabstractIn order to overcome the drawback of the traditional Kirsch template despeckling usings fixed windows, an improved Kirsch direction template despeckling algorithm, based on structural information detection, is proposed for high-resolution synthetic aperture radar (SAR) images. First, the point targets are detected and preserved in the current region. Second, the window is enlarged adaptively based on the statistical characteristics of the local region. Finally, the window finally obtained is classified. The averaged filter is directly adopted if the region is homogeneous, or else the Kirsch template filter is used. Combining point target detection, adaptive windowing, and region classification, altogether the proposed algorithm can effectively improve the performance of the traditional Kirsch direction template despeckling. Despeckling experiments on simulated and real high-resolution SAR images demonstrate that the Kirsch direction template despeckling algorithm based on structural information detection can not only sufficiently suppress speckle in homogenous and edge regions, but also effectively preserve point targets and edge information, leading to good despeckling results. Sujuan Hou, Zengguo Sun, Yunjing Song |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2021 | SDOF-GAN: Symmetric Dense Optical Flow Estimation With Generative Adversarial NetworksabstractThere is a growing consensus in computer vision that symmetric optical flow estimation constitutes a better model than a generic asymmetric one for its independence of the selection of source/target image. Yet, convolutional neural networks (CNNs), that are considered the de facto standard vision model, deal with the asymmetric case only in most cutting-edge CNNs-based optical flow techniques. We bridge this gap by introducing a novel model named SDOF-GAN: symmetric dense optical flow with generative adversarial networks (GANs). SDOF-GAN realizes a consistency between the forward mapping (source-to-target) and the backward one (target-to-source) by ensuring that they are inverse of each other with an inverse network. In addition, SDOF-GAN leverages a GAN model for which the generator estimates symmetric optical flow fields while the discriminator differentiates the "real" ground-truth flow field from a "fake" estimation by assessing the flow warping error. Finally, SDOF-GAN is trained in a semi-supervised fashion to enable both the precious labeled data and large amounts of unlabeled data to be fully-exploited. We demonstrate significant performance benefits of SDOF-GAN on five publicly-available datasets in contrast to several representative state-of-the-art models for optical flow estimation. Tongtong Che, Yuanjie Zheng, Yunshuai Yang, Sujuan Hou, Weikuan Jia, Jie Yang 0002, Chen Gong 0002 |
IEEE Trans. Image Process. | 4 |
| 2021 | Solving Jigsaw Puzzles via Nonconvex Quadratic Programming With the Projected Power MethodabstractJigsaw puzzles consist of reconstructing a picture that has been divided into many interlocking pieces. This paper describes an automatic global method for solving the square-piece jigsaw puzzle problem in which neither the orientations nor the locations of the jigsaw pieces are known. This hard combinatorial sorting task is formulated as a nonconvex quadratic programming problem that is solved via the projected power method. Specifically, this work aims to specify the locations and orientations of puzzle pieces by maximizing a constrained quadratic function that resolves an optimized permutation matrix composed of the noisy pairwise affinities between jigsaw pieces. The experimental results obtained in the MIT, McGill and Pomeranz datasets indicate that our method outperforms state-of-the-art techniques. Fang Yan 0003, Yuanjie Zheng, Jinyu Cong, Liu Liu 0014, Dacheng Tao, Sujuan Hou |
IEEE Trans. Multim. | 6 |
| 2020 | Logo-2K+: A Large-Scale Logo Dataset for Scalable Logo ClassificationabstractLogo classification has gained increasing attention for its various applications, such as copyright infringement detection, product recommendation and contextual advertising. Compared with other types of object images, the real-world logo images have larger variety in logo appearance and more complexity in their background. Therefore, recognizing the logo from images is challenging. To support efforts towards scalable logo classification task, we have curated a dataset, Logo-2K+, a new large-scale publicly available real-world logo dataset with 2,341 categories and 167,140 images. Compared with existing popular logo datasets, such as FlickrLogos-32 and LOGO-Net, Logo-2K+ has more comprehensive coverage of logo categories and larger quantity of logo images. Moreover, we propose a Discriminative Region Navigation and Augmentation Network (DRNA-Net), which is capable of discovering more informative logo regions and augmenting these image regions for logo classification. DRNA-Net consists of four sub-networks: the navigator sub-network first selected informative logo-relevant regions guided by the teacher sub-network, which can evaluate its confidence belonging to the ground-truth logo class. The data augmentation sub-network then augments the selected regions via both region cropping and region dropping. Finally, the scrutinizer sub-network fuses features from augmented regions and the whole image for logo classification. Comprehensive experiments on Logo-2K+ and other three existing benchmark datasets demonstrate the effectiveness of proposed method. Logo-2K+ and the proposed strong baseline DRNA-Net are expected to further the development of scalable logo image recognition, and the Logo-2K+ dataset can be found at https://github.com/msn199959/Logo-2k-plus-Dataset. Jing Wang 0138, Weiqing Min, Sujuan Hou, Shengnan Ma, Yuanjie Zheng, Haishuai Wang, Shuqiang Jiang |
AAAI | 3 |
| 2020 | Singular value decomposition-based virtual representation for face recognition
Shigang Liu, Sujuan Hou, Keyou Zhang, Xiaojun Wu 0002 |
Mach. Vis. Appl. | 4 |
| 2020 | Online Multi-Expert Learning for Visual TrackingabstractThe correlation filters based trackers have achieved an excellent performance for object tracking in recent years. However, most existing methods use only one filter but ignore the information of the previous filters. In this paper, we propose a novel online multi-expert learning algorithm for visual tracking. In our proposed scheme, there are former trackers which retain the previous filters, and those trackers will give their predictions in each frame. The current tracker represents the filter of current frame, and both the current tracker and the former trackers constitute our expert ensemble. We use an adaptive Second-order Quantile strategy to learn the weights of each expert, which can take full advantage of all the experts. To simplify our model and remove some bad experts, we prune our models via a minimum entropy criterion. Finally, we propose a new update strategy to avoid the model corruption problem. Extensive experimental results on both OTB2013 and OTB2015 benchmarks demonstrate that our proposed tracker performs favorably against state-of-the-art methods. Zhetao Li, Tianzhu Zhang 0001, Meng Wang 0001, Sujuan Hou, Xin Peng 0002 |
IEEE Trans. Image Process. | 5 |
| 2019 | A novel optimized GA-Elman neural network algorithm
Weikuan Jia, Dean Zhao, Yuanjie Zheng, Sujuan Hou |
Neural Comput. Appl. | 4 |
| 2018 | Deblurring retinal optical coherence tomography via a convolutional neural network with anisotropic and double convolution layerabstractVarious image pre‐processing tasks in optical coherence tomography (OCT) systems involve reversing degradation effects (e.g. deblurring). Current deblurring research mainly focuses on how to build suitable degradation models using deconvolution operators. However, model‐based solutions may not work well in many scenarios. To solve this problem, the authors propose a non‐model architecture, called a deep convolutional neural network, to address parameter‐free situations. The proposed solution employs a deep learning strategy to bridge the gap between traditional model‐based methods and neural network architectures. Experiments on retinal OCT images demonstrate that the proposed approach achieves superior performance compared with the state‐of‐the‐art model‐based OCT deblurring methods. Jian Lian, Sujuan Hou, Xiaodan Sui, Fangzhou Xu, Yuanjie Zheng |
IET Comput. Vis. | 2 |
| 2018 | Classifying advertising video by topicalizing high-level semantic concepts
Sujuan Hou, Shangbo Zhou, Wenjie Liu 0001, Yuanjie Zheng |
Multim. Tools Appl. | 1 |
| 2017 | Multi-layer multi-view topic model for classifying advertising video
Sujuan Hou, Ling Chen 0006, Dacheng Tao, Shangbo Zhou, Wenjie Liu 0001, Yuanjie Zheng |
Pattern Recognit. | 1 |
| 2016 | Multi-label learning with label relevance in advertising video
Sujuan Hou, Shangbo Zhou, Ling Chen 0006, Karim Awudu |
Neurocomputing | 1 |
| 2014 | A compressed sensing approach for query by example video retrieval
Sujuan Hou, Shangbo Zhou, Muhammad Abubakar Siddique |
Multim. Tools Appl. | 1 |