VLDB 2026 Research / reviewers in the wild / expert
Nong Sang
dblp:10/1545
· DBLP profile ↗
227ranked-venue papers
3as first author
103since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 145 · 1 first-author · 65 since 2021Artificial intelligence and machine learning · 132 · 3 first-author · 66 since 2021Databases, data management, data science and information retrieval · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning to Tell Apart: Weakly Supervised Video Anomaly Detection via Disentangled Semantic AlignmentabstractRecent advancements in weakly-supervised video anomaly detection have achieved remarkable performance by applying the multiple instance learning paradigm based on multimodal foundation models such as CLIP to highlight anomalous instances and classify categories. However, their objectives may tend to detect the most salient response segments, while neglecting to mine diverse normal patterns separated from anomalies, and are prone to category confusion due to similar appearance, leading to unsatisfactory fine-grained classification results. Therefore, we propose a novel Disentangled Semantic Alignment Network (DSANet) to explicitly separate abnormal and normal features from coarse-grained and fine-grained aspects, enhancing the distinguishability. Specifically, at the coarse-grained level, we introduce a self-guided normality modeling branch that reconstructs input video features under the guidance of learned normal prototypes, encouraging the model to exploit normality cues inherent in the video, thereby improving the temporal separation of normal patterns and anomalous events. At the fine-grained level, we present a decoupled contrastive semantic alignment mechanism, which first temporally decomposes each video into event-centric and background-centric components using frame-level anomaly scores and then applies visual-language contrastive learning to enhance class-discriminative representations. Comprehensive experiments on two standard benchmarks, namely XD-Violence and UCF-Crime, demonstrate that DSANet outperforms existing state-of-the-art methods. Wenti Yin, Huaxin Zhang, Xiang Wang 0012, Yuqing Lu, Bingquan Gong, Jialong Zuo, Li Yu 0003, Changxin Gao, Nong Sang |
AAAI | 10 |
| 2026 | Few-shot action recognition with captioning foundation models
Xiang Wang 0012, Shiwei Zhang 0001, Hangjie Yuan, Yingya Zhang, Changxin Gao, Deli Zhao, Nong Sang |
Comput. Vis. Image Underst. | 7 |
| 2026 | High-Resolution Photo Enhancement in Real-Time: A Laplacian Pyramid NetworkabstractPhoto enhancement plays a crucial role in augmenting the visual aesthetics of a photograph. In recent years, photo enhancement methods have either focused on enhancement performance, producing powerful models that cannot be deployed on edge devices, or prioritized computational efficiency, resulting in inadequate performance for real-world applications. To this end, this paper introduces a pyramid network called LLF-LUT++, which integrates global and local operators through closed-form Laplacian pyramid decomposition and reconstruction. This approach enables fast processing of high-resolution images while also achieving excellent performance. Specifically, we utilize an image-adaptive 3D LUT that capitalizes on the global tonal characteristics of downsampled images, while incorporating two distinct weight fusion strategies to achieve coarse global image enhancement. To implement this strategy, we designed a spatial-frequency transformer weight predictor that effectively extracts the desired distinct weights by leveraging frequency features. Additionally, we apply local Laplacian filters to adaptively refine edge details in high-frequency components. After meticulously redesigning the network structure and transformer model, LLF-LUT++ not only achieves a 2.64 dB improvement in PSNR on the HDR+ dataset, but also further reduces runtime, with 4 K resolution images processed in just 13 ms on a single GPU. Extensive experimental results on two benchmark datasets further show that the proposed approach performs favorably compared to state-of-the-art methods. Feng Zhang 0039, Haoyou Deng, Lida Li, Qingbo Lu, Zisheng Cao, Minchen Wei, Changxin Gao, Nong Sang, Xiang Bai |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | Mitigating distribution shift via adaptive reweighting for robust visual question answering
Xingdong Song, Congzhen Yu, Zukun Wan, Tianming Ma, Changxin Gao, Nong Sang |
Pattern Recognit. | 8 |
| 2026 | Weakly supervised crowd counting with joint CNN and transformer network
Fusen Wang, Kai Liu 0054, Nong Sang, Xiaofeng Xia, Jun Sang |
Pattern Recognit. | 4 |
| 2026 | A2HA: Attribute-aware hierarchical alignment for text-image person re-identification
Qiuju Dai, Lingxin Cui, Xingdong Song, Congzhen Yu, Changxin Gao, Nong Sang |
Pattern Recognit. | 10 |
| 2026 | MSINet: A Mask Structure Inference Network for Scene Text Image Super-ResolutionabstractUnderstanding the structure of characters is crucial for recovering clear and readable high-resolution scene text images in Scene Text Image Super-Resolution (STISR). Recently, many existing STISR methods inject the character structure information implicit in the recognition priors into the super-resolution network to guide the super-resolution process, thereby facilitating the generation of more legible text images. However, the recognition priors obtained from low resolution are inaccurate, which means that directly embedding these priors into the network easily misleads the super-resolution process. To address this problem, we draw inspiration from Masked Image Modeling (MIM) and propose the Mask Structure Inference Network (MSINet), which can generate scene text images with accurate character structures without directly embedding recognition priors. To make STISR compatible with MIM, we also propose a Mask-and-Inference Paradigm (MIP), which consists of a mask image pre-training stage for character structure learning and a fine-tuning stage for character structure inference. In addition, a novel mask strategy named Text Confidence Mask (TCM) is proposed to avoid recovery errors by masking legible character regions. With MIP and TCM, MSINet impressively improves the clarity and readability of the degraded scene text images. Specifically, MSINet-B outperforms recent state-of-the-art methods by about +3.7% on the TextZoom and average +3.6% on six manually degraded scene text recognition datasets in recognition accuracy. The code will be released at https://github.com/Yuanssr/MSINet. Shengrong Yuan, Shan Ye, Xingdong Song, Shengyou Qian, Changxin Gao, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | Toward Robust Alignment for Video Dehazing With Temporal Lookup TableabstractVideo dehazing aims to restore clean scenarios from a sequence of hazy frames, where frame alignment is a critical stage for leveraging temporal information. However, haze degrades contrast and obscures details, making alignment challenging. Existing methods ignore the impairment of haze on alignment and thus struggle to align frames accurately. To address this challenge, we propose an alignment network with the temporal lookup table (temporal-LUT), which effectively enhances the haze-degraded frames and provides vivid cues for precise alignment. Specifically, to tackle the color degradation of haze, we employ a learnable lookup table (LUT) to enhance hazy color. The color mapping nature of LUT favorably preserves the naturalness of enhanced outcomes. Besides, we introduce a temporal weight prediction strategy to strengthen inter-frame interaction, which ensures temporal consistency across enhanced results and thereby benefits alignment. Extensive experimental results on two widely used benchmarks and real-world scenes demonstrate the superiority of our method. Haoyou Deng, Feng Zhang 0039, Qingbo Lu, Changxin Gao, Nong Sang |
IEEE Trans. Image Process. | 7 |
| 2026 | REPAIR: Rank Correlation and Noisy Pair Half-Replacing With Memory for Noisy CorrespondenceabstractThe presence of noise in acquired data invariably leads to performance degradation in cross-modal matching. Unfortunately, obtaining precise annotations in the multimodal field is expensive, which has prompted some methods to tackle the mismatched data pair issue in cross-modal matching contexts, termed as noisy correspondence. However, most of these existing noisy correspondence methods exhibit the following limitations: a) the problem of self-reinforcing error accumulation, and b) improper handling of noisy data pair. To tackle the two problems, we propose a generalized framework termed as Rank corrElation and noisy Pair hAlf-replacing wIth memoRy (REPAIR), which benefits from maintaining a memory bank for features of matched pairs. Specifically, we calculate the distances between the features in the memory bank and those of the target pair for each respective modality, and use the rank correlation of these two sets of distances to estimate the soft correspondence label of the target pair. Estimating soft correspondence based on memory bank features rather than using a similarity network can avoid the accumulation of errors due to incorrect network identifications. For pairs that are completely mismatched, REPAIR searches the memory bank for the most matching feature to replace one feature of one modality, instead of using the original pair directly or merely discarding the mismatched pair. We conduct experiments on three cross-modal datasets,i.e., Flickr30 K, MS-COCO, and CC152 K, proving the effectiveness and robustness of our REPAIR on synthetic and real-world noise. Ruochen Zheng, Jiahao Hong, Changxin Gao, Nong Sang |
IEEE Trans. Multim. | 4 |
| 2025 | Structural Pruning via Spatial-aware Information Redundancy for Semantic SegmentationabstractIn recent years, semantic segmentation has flourished in various applications. However, the high computational cost remains a significant challenge that hinders its further adoption. The filter pruning method for structured network slimming offers a direct and effective solution for the reduction of segmentation networks. Nevertheless, we argue that most existing pruning methods, originally designed for image classification, overlook the fact that segmentation is a location-sensitive task, which consequently leads to their suboptimal performance when applied to segmentation networks. To address this issue, this paper proposes a novel approach, denoted as Spatial-aware Information Redundancy Filter Pruning (SIRFP), which aims to reduce feature redundancy between channels. First, we formulate the pruning process as a maximum edge weight clique problem (MEWCP) in graph theory, thereby minimizing the redundancy among the remaining features after pruning. Within this framework, we introduce a spatial-aware redundancy metric based on feature maps, thus endowing the pruning process with location sensitivity to better adapt to pruning segmentation networks. Additionally, based on the MEWCP, we propose a low computational complexity greedy strategy to solve this NP-hard problem, making it feasible and efficient for structured pruning. To validate the effectiveness of our method, we conducted extensive comparative experiments on various challenging datasets. The results demonstrate the superior performance of SIRFP for semantic segmentation tasks. Dongyue Wu, Zilin Guo, Li Yu 0003, Nong Sang, Changxin Gao |
AAAI | 4 |
| 2025 | Adaptive Prototype Replay for Class Incremental Semantic SegmentationabstractClass incremental semantic segmentation (CISS) aims to segment new classes during continual steps while preventing the forgetting of old knowledge. Existing methods alleviate catastrophic forgetting by replaying distributions of previously learned classes using stored prototypes or features. However, they overlook a critical issue: in CISS, the representation of class knowledge is updated continuously through incremental learning, whereas prototype replay methods maintain fixed prototypes. This mismatch between updated representation and fixed prototypes limits the effectiveness of the prototype replay strategy. To address this issue, we propose the Adaptive prototype replay (Adapter) for CISS in this paper. Adapter comprises an adaptive deviation compensation (ADC) strategy and an uncertainty-aware constraint (UAC) loss. Specifically, the ADC strategy dynamically updates the stored prototypes based on the estimated representation shift distance to match the updated representation of old class. The UAC loss reduces prediction uncertainty, aggregating discriminative features to aid in generating compact prototypes. Additionally, we introduce a compensation-based prototype similarity discriminative (CPD) loss to ensure adequate differentiation between similar prototypes, thereby enhancing the efficiency of the adaptive prototype replay strategy. Extensive experiments on Pascal VOC and ADE20K datasets demonstrate that Adapter achieves state-of-the-art results and proves effective across various CISS tasks, particularly in challenging multi-step scenarios. Guilin Zhu, Dongyue Wu, Changxin Gao, Nong Sang |
AAAI | 6 |
| 2025 | L-Man: A Large Multi-modal Model Unifying Human-centric TasksabstractLarge language models (LLMs) have recently shown notable progress in unifying various visual tasks with an open-ended form. However, when transferred to human-centric tasks, despite their remarkable multi-modal understanding ability in general domains, they lack further human-related domain knowledge and show unsatisfactory performance. Meanwhile, current human-centric unified models are mostly restricted to a pre-defined form and lack open-ended task capability. Therefore, it is necessary to propose a large multi-modal model which utilizes LLMs to unify various human-centric tasks. We forge ahead along this path from the aspects of dataset and model. Specifically, we first construct a large-scale language-image instruction-following dataset named HumanIns based on existing 20 open datasets from 6 diverse downstream tasks, which provides sufficient and diverse data to implement multi-modal training. Then, a model named L-Man including a query adapter is designed to extract the multi-grained semantics of image and align the cross-modal information between image and text. In practice, we introduce a two-stage training strategy, where the first stage extracts generic text-relevant visual information, and the second stage maps the visual features to the embedding space of the LLM. By tuning on HumanIns, our model shows significant superiority on human-centric tasks compared with existing large multi-modal models, and also achieves even better results on downstream datasets compared with respective task-specific models. Jialong Zuo, Tianyu Guo 0001, Huaxin Zhang, Jiahao Hong, Nong Sang, Changxin Gao, Kai Han 0002 |
AAAI | 6 |
| 2025 | Customizing Visual-Language Foundation Models for Multi-Modal Anomaly Detection and ReasoningabstractAnomaly detection is vital in various industrial scenarios, including the identification of unusual patterns in production lines and the detection of manufacturing defects for quality control. Existing techniques tend to be specialized in individual scenarios and lack generalization capacities. In this study, our objective is to develop a generic anomaly detection model that can be applied in multiple scenarios. To achieve this, we custom-build generic visual language foundation models that possess extensive knowledge and robust reasoning abilities as anomaly detectors and reasoners. Specifically, we introduce a multi-modal prompting strategy that incorporates domain knowledge from experts as conditions to guide the models. Our approach considers diverse prompt types, including task descriptions, class context, normality rules, and reference images. In addition, we unify the input representation of multi-modality into a 2D image format, enabling multi-modal anomaly detection and reasoning. Our preliminary studies demonstrate that combining visual and language prompts as conditions for customizing the models enhances anomaly detection performance. The customized models showcase the ability to detect anomalies across different data modalities such as images, point clouds, and videos. Qualitative case studies further highlight the anomaly detection and reasoning capabilities, particularly for multi-object scenes and temporal data. Our code is publicly available at https://github.com/Xiaohac-Xu/Customizable-VLM.11More insights of customized foundation models for broader anomaly detection settings are available at Github repo: https://github.com/caoyunkang/GPT4V-for-Generic-Anomaly-Detection. Xiaohao Xu, Yunkang Cao, Huaxin Zhang, Nong Sang, Xiaonan Huang |
CSCWD | 4 |
| 2025 | Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any GranularityabstractHow can we enable models to comprehend video anomalies occurring over varying temporal scales and contexts? Traditional Video Anomaly Understanding (VAU) methods focus on frame-level anomaly prediction, often missing the interpretability of complex and diverse real-world anomalies. Recent multimodal approaches leverage visual and textual data but lack hierarchical annotations that capture both short-term and long-term anomalies. To address this challenge, we introduce HIVAU-70k, a large-scale benchmark for hierarchical video anomaly understanding across any granularity. We develop a semi-automated annotation engine that efficiently scales high-quality annotations by combining manual video segmentation with recursive free-text annotation using large language models (LLMs). This results in over 70,000 multi-granular annotations organized at clip-level, event-level, and video-level segments. For efficient anomaly detection in long videos, we propose the Anomaly-focused Temporal Sampler (ATS). ATS integrates an anomaly scorer with a density-aware sampler to adaptively select frames based on anomaly scores, ensuring that the multimodal LLM concentrates on anomaly-rich regions, which significantly enhances both efficiency and accuracy. Extensive experiments demonstrate that our hierarchical instruction data markedly improves anomaly comprehension. The integrated ATS and visual-language model outperform traditional methods in processing long videos. Our benchmark and model are publicly available at https://github.com/pipixin321/HolmesVAU. Huaxin Zhang, Xiaohao Xu, Xiang Wang 0012, Jialong Zuo, Xiaonan Huang, Changxin Gao, Shanjun Zhang, Li Yu 0003, Nong Sang |
CVPR | 9 |
| 2025 | Partial Forward Blocking: A Novel Data Pruning Paradigm for Lossless Training AccelerationabstractThe ever-growing size of training datasets enhances the generalization capability of modern machine learning models but also incurs exorbitant computational costs. Existing data pruning approaches aim to accelerate training by removing those less important samples. However, they often rely on gradients or proxy models, leading to prohibitive additional costs of gradient back-propagation and proxy model training. In this paper, we propose Partial Forward Blocking (PFB), a novel framework for lossless training acceleration. The efficiency of PFB stems from its unique adaptive pruning pipeline: sample importance is assessed based on features extracted from the shallow layers of the target model. Less important samples are then pruned, allowing only the retained ones to proceed with the subsequent forward pass and loss back-propagation. This mechanism significantly reduces the computational overhead of deep-layer forward passes and back-propagation for pruned samples, while also eliminating the need for auxiliary backward computations and proxy model training. Moreover, PFB introduces probability density as an indicator of sample importance. Combined with an adaptive distribution estimation module, our method dynamically prioritizes relatively rare samples, aligning with the constantly evolving training state. Extensive experiments demonstrate the significant superiority of PFB in performance and speed. On ImageNet, PFB achieves a 0.5% accuracy improvement and 33% training time reduction with 40% data pruned. Dongyue Wu, Zilin Guo, Jialong Zuo, Nong Sang, Changxin Gao |
ICCV | 4 |
| 2025 | StyleSRN: Scene Text Image Super-Resolution with Text Style Embedding
Shengrong Yuan, Ke Hao, Xuqi Ma, Changxin Gao, Li Liu 0002, Nong Sang |
ICCV | 7 |
| 2025 | Boosting Portrait Matting with Spatial-Frequency Harmony
Rongsheng Luo, Changxin Gao, Nong Sang |
ICIG (2) | 3 |
| 2025 | MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane RepresentationabstractHuman instance matting aims to estimate an alpha matte for each human instance in an image, which is challenging as it easily fails in complex cases requiring disentangling mingled pixels belonging to multiple instances along hairy and thin boundary structures. In this work, we address this by introducing MP-Mat, a novel 3D-and-instance-aware matting framework with multiplane representation, where the multiplane concept is designed from two different perspectives: scene geometry level and instance level. Specifically, we first build feature-level multiplane representations to split the scene into multiple planes based on depth differences. This approach makes the scene representation 3D-aware, and can serve as an effective clue for splitting instances in different 3D positions, thereby improving interpretability and boundary handling ability especially in occlusion areas. Then, we introduce another multiplane representation that splits the scene in an instance-level perspective, and represents each instance with both matte and color. We also treat background as a special instance, which is often overlooked by existing methods. Such an instance-level representation facilitates both foreground and background content awareness, and is useful for other down-stream tasks like image editing. Once built, the representation can be reused to realize controllable instance-level image editing with high efficiency. Extensive experiments validate the clear advantage of MP-Mat in matting task. We also demonstrate its superiority in image editing tasks, an area under-explored by existing matting-focused methods, where our approach under zero-shot inference even outperforms trained specialized image editing techniques by large margins. Code is open-sourced at https://github.com/JiaoSiyi/MPMat.git. Siyi Jiao, Wenzheng Zeng, Yerong Li, Changxin Gao, Nong Sang, Zheng Shou 0001 |
ICLR | 6 |
| 2025 | GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly DetectionabstractIn recent years, video anomaly detection has been extensively investigated in both unsupervised and weakly supervised settings to alleviate costly temporal labeling. Despite significant progress, these methods still suffer from unsatisfactory results such as numerous false alarms, primarily due to the absence of precise temporal anomaly annotation. In this paper, we present a novel labeling paradigm, termed "glance annotation", to achieve a better balance between anomaly detection accuracy and annotation cost. Specifically, glance annotation is a random frame within each abnormal event, which can be easily accessed and is cost-effective. To assess its effectiveness, we manually annotate the glance annotations for two standard video anomaly detection datasets: UCF-Crime and XD-Violence. Additionally, we propose a customized GlanceVAD method, that leverages gaussian kernels as the basic unit to compose the temporal anomaly distribution, enabling the learning of diverse and robust anomaly representations from the glance annotations. Through comprehensive analysis and experiments, we verify that the proposed labeling paradigm can achieve an excellent trade-off between annotation cost and model performance. Extensive experimental results also demonstrate the effectiveness of our GlanceVAD approach, which significantly outperforms existing advanced unsupervised and weakly supervised methods. Our annotations and code are publicly available at https://github.com/pipixin321/GlanceVAD. Huaxin Zhang, Xiang Wang 0012, Xiaohao Xu, Xiaonan Huang, Changxin Gao, Yuehuan Wang, Shanjun Zhang, Nong Sang |
ICME | 8 |
| 2025 | Towards Reliable and Holistic Visual In-Context Learning Prompt SelectionabstractVisual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedded in in-context examples, which can be formulated as a global ranking problem of potential candidates. Current VICL methods, such as Partial2Global and VPR, are grounded in the similarity-priority assumption that images more visually similar to a query image serve as better in-context examples. This foundational assumption, while intuitive, lacks sufficient justification for its efficacy in selecting optimal in-context examples. Furthermore, Partial2Global constructs its global ranking from a series of randomly sampled pairwise preference predictions. Such a reliance on random sampling can lead to incomplete coverage and redundant samplings of comparisons, thus further adversely impacting the final global ranking. To address these issues, this paper introduces an enhanced variant of Partial2Global designed for reliable and holistic selection of in-context examples in VICL. Our proposed method, dubbed RH-Partial2Global, leverages a jackknife conformal prediction-guided strategy to construct reliable alternative sets and a covering design-based sampling approach to ensure comprehensive and uniform coverage of pairwise preferences. Extensive experiments demonstrate that RH-Partial2Global achieves excellent performance and outperforms Partial2Global across diverse visual tasks. Wenxiao Wu, Jing-Hao Xue, Chengming Xu 0001, Chen Liu 0030, Xinwei Sun 0001, Changxin Gao, Nong Sang, Yanwei Fu 0001 |
NeurIPS | 7 |
| 2025 | Continual Gaussian Mixture Distribution Modeling for Class Incremental Semantic SegmentationabstractClass incremental semantic segmentation (CISS) enables a model to continually segment new classes from non-stationary data while preserving previously learned knowledge. Recent top-performing approaches are prototype-based methods that assign a prototype to each learned class to reproduce previous knowledge. However, modeling each class distribution relying on only a single prototype, which remains fixed throughout the incremental process, presents two key limitations: (i) a single prototype is insufficient to accurately represent the complete class distribution when incoming data stream for a class is naturally multimodal; (ii) the features of old classes may exhibit anisotropy during the incremental process, preventing fixed prototypes from faithfully reproducing the matched distribution. To address the aforementioned limitations, we propose a Continual Gaussian Mixture Distribution (CoGaMiD) modeling method. Specifically, the means and covariance matrices of the Gaussian Mixture Models (GMMs) are estimated to model the complete feature distributions of learned classes. These GMMs are stored to generate pseudo-features that support the learning of novel classes in incremental steps. Moreover, we introduce a Dynamic Adjustment (DA) strategy that utilizes the features of previous classes within incoming data streams to update the stored GMMs. This adaptive update mitigates the mismatch between fixed GMMs and continually evolving distributions. Furthermore, a Gaussian-based Representation Constraint (GRC) loss is proposed to enhance the discriminability of new classes, avoiding confusion between new and old classes. Extensive experiments on Pascal VOC and ADE20K show that our method achieves superior performance compared to previous methods, especially in more challenging long-term incremental scenarios. Guilin Zhu, Yuanjie Shao, Nong Sang, Changxin Gao |
NeurIPS | 5 |
| 2025 | VideoLucy: Deep Memory Backtracking for Long Video UnderstandingabstractRecent studies have shown that agent-based systems leveraging large language models (LLMs) for key information retrieval and integration have emerged as a promising approach for long video understanding. However, these systems face two major challenges. First, they typically perform modeling and reasoning on individual frames, struggling to capture the temporal context of consecutive frames. Second, to reduce the cost of dense frame-level captioning, they adopt sparse frame sampling, which risks discarding crucial information. To overcome these limitations, we propose VideoLucy, a deep memory backtracking framework for long video understanding. Inspired by the human recollection process from coarse to fine, VideoLucy employs a hierarchical memory structure with progressive granularity. This structure explicitly defines the detail level and temporal scope of memory at different hierarchical depths. Through an agent-based iterative backtracking mechanism, VideoLucy systematically mines video-wide, question-relevant deep memories until sufficient information is gathered to provide a confident answer. This design enables effective temporal understanding of consecutive frames while preserving critical details. In addition, we introduce EgoMem, a new benchmark for long video understanding. EgoMem is designed to comprehensively evaluate a model's ability to understand complex events that unfold over time and capture fine-grained details in extremely long videos. Extensive experiments demonstrate the superiority of VideoLucy. Built on open-source models, VideoLucy significantly outperforms state-of-the-art methods on multiple long video understanding benchmarks, achieving performance even surpassing the latest proprietary models such as GPT-4o. Our code and dataset will be made publicly available. Jialong Zuo, Yongtai Deng, Lingdong Kong, Nong Sang, Liang Pan, Ziwei Liu 0002, Changxin Gao |
NeurIPS | 7 |
| 2025 | ReID5o: Achieving Omni Multi-modal Person Re-identification in a Single ModelabstractIn real-word scenarios, person re-identification (ReID) expects to identify a person-of-interest via the descriptive query, regardless of whether the query is a single modality or a combination of multiple modalities. However, existing methods and datasets remain constrained to limited modalities, failing to meet this requirement. Therefore, we investigate a new challenging problem called Omni Multi-modal Person Re-identification (OM-ReID), which aims to achieve effective retrieval with varying multi-modal queries. To address dataset scarcity, we construct ORBench, the first high-quality multi-modal dataset comprising 1,000 unique identities across five modalities: RGB, infrared, color pencil, sketch, and textual description. This dataset also has significant superiority in terms of diversity, such as the painting perspectives and textual information. It could serve as an ideal platform for follow-up investigations in OM-ReID. Moreover, we propose ReID5o, a novel multi-modal learning framework for person ReID. It enables synergistic fusion and cross-modal alignment of arbitrary modality combinations in a single model, with a unified encoding and multi-expert routing mechanism proposed. Extensive experiments verify the advancement and practicality of our ORBench. A range of models have been compared on it, and our proposed ReID5o gives the best performance. Jialong Zuo, Yongtai Deng, Mengdan Tan, Dongyue Wu, Nong Sang, Liang Pan, Changxin Gao |
NeurIPS | 6 |
| 2025 | Clean Feature Distillation for Low-Light Detection
Haoyou Deng, Rongsheng Luo, Xirong Song, Changxin Gao, Nong Sang |
PRCV (16) | 6 |
| 2025 | DMPT: Decoupled Modality-Aware Prompt Tuning for Multi-Modal Object Re-IdentificationabstractCurrent multi-modal object re-identification approaches based on large-scale pre-trained backbones (i.e., ViT) have displayed remarkable progress and achieved excellent performance. However, these methods usually adopt the standard full fine-tuning paradigm, which requires the optimization of considerable backbone parameters, causing extensive computational and storage requirements. In this work, we propose an efficient prompt-tuning framework tailored for multi-modal object re-identification; dubbed DMPT, which freezes the main backbone and only optimizes several newly added decoupled modality-aware parameters. Specifically, we explicitly decouple the visual prompts into modality-specific prompts which leverage prior modality knowledge from a powerful text encoder and modality-independent semantic prompts which extract semantic information from multi-modal inputs, such as visible, near-infrared, and thermal-infrared. Built upon the extracted features, we further design a Prompt Inverse Bind (PromptIBind) strategy that employs bind prompts as a medium to connect the semantic prompt tokens of different modalities and facilitates the exchange of complementary multi-modal information, boosting final re-identification results. Experimental results on multiple common benchmarks demonstrate that our DMPT can achieve competitive results to existing state-of-the-art methods while requiring only 6.5% fine-tuning of the backbone parameters. Minghui Lin, Xiang Wang 0012, Jianhua Tang, Longbin Fu, Zhengrong Zuo, Nong Sang |
WACV | 7 |
| 2025 | CTR-Driven Advertising Image Generation with Multimodal Large Language ModelsabstractIn web data, advertising images are crucial for capturing user attention and improving advertising effectiveness. Most existing methods generate background for products primarily focus on the aesthetic quality, which may fail to achieve satisfactory online performance. To address this limitation, we explore the use of Multimodal Large Language Models (MLLMs) for generating advertising images by optimizing for Click-Through Rate (CTR) as the primary objective. Firstly, we build targeted pre-training tasks, and leverage a large-scale e-commerce multimodal dataset to equip MLLMs with initial capabilities for advertising image generation tasks. To further improve the CTR of generated images, we propose a novel reward model to fine-tune pre-trained MLLMs through Reinforcement Learning (RL), which can jointly utilize multimodal features and accurately reflect user click preferences. Meanwhile, a product-centric preference optimization strategy is developed to ensure that the generated background content aligns with the product characteristics after fine-tuning, enhancing the overall relevance and effectiveness of the advertising images. Extensive experiments have demonstrated that our method achieves state-of-the-art performance in both online and offline metrics. Our code and pre-trained models are publicly available at: https://github.com/Chenguoz/CAIG. Xingye Chen, Zhenbang Du, Yanyin Chen, Haohan Wang, Linkai Liu 0002, Jinyuan Zhao, Jingjing Lv, Junjie Shen 0008, Zhangang Lin, Jingping Shao, Yuanjie Shao, Xinge You, Changxin Gao, Nong Sang |
WWW | 19 |
| 2025 | UniAnimate: taming unified video diffusion models for consistent human image animation
Xiang Wang 0012, Shiwei Zhang 0001, Changxin Gao, Xiaoqiang Zhou, Yingya Zhang, Luxin Yan, Nong Sang |
Sci. China Inf. Sci. | 8 |
| 2025 | Spatial cascaded clustering and weighted memory for unsupervised person re-identification
Jiahao Hong, Jialong Zuo, Chuchu Han, Ruochen Zheng, Ming Tian, Changxin Gao, Nong Sang |
Image Vis. Comput. | 7 |
| 2025 | Exploring sample relationship for few-shot classification
Xingye Chen, Wenxiao Wu, Li Ma 0005, Xinge You, Changxin Gao, Nong Sang, Yuanjie Shao |
Pattern Recognit. | 6 |
| 2025 | T2EA: Target-Aware Taylor Expansion Approximation Network for Infrared and Visible Image FusionabstractIn the image fusion mission, the crucial task is to generate high-quality images for highlighting the key objects while enhancing the scenes to be understood. To complete this task and provide a powerful interpretability as well as a strong generalization ability in producing enjoyable fusion results which are comfortable for vision tasks (such as objects detection and their segmentation), we present a novel interpretable decomposition scheme and develop a target-aware Taylor expansion approximation (T2EA) network for infrared and visible image fusion, where our T2EA includes the following key procedures: Firstly, visible and infrared images are both decomposed into feature maps through a designed Taylor expansion approximation (TEA) network. Then, the Taylor feature maps are hierarchically fused by a dual-branch feature fusion (DBFF) network. Next, the fused map of each layer is contributed to synthesize an enjoyable fusion result by the inverse Taylor expansion. Finally, a segmentation network is jointed to refine the fusion network parameters which can promote the pleasing fusion results to be more suitable for segmenting the objects. To validate the effectiveness of our reported T2EA network, we first discuss the selection of Taylor expansion layers and fusion strategies. Then, both quantitatively and qualitatively experimental results generated by the selected SOTA approaches on three datasets (MSRS, TNO, andLLVIP) are compared in testing, generalization, and target detection and segmentation, demonstrating that our T2EA can produce more competitive fusion results for vision tasks and is more powerful for image adaption. The code will be available at https://github.com/MysterYxby/T2EA. Zhenghua Huang, Biyun Xu, Menghan Xia, Qian Li 0019, Yansheng Li 0001, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Linear Feature Source Prediction and Recombination Network for Noisy Label LearningabstractCollecting training data for deep models from the Internet is a common data acquisition approach. However, there are challenges in using these data directly, as they often contain inaccurate annotations. This situation has increased the attention and importance of noisy label learning, the process of training a deep model with unreliable annotations. The typical strategy in noisy label learning is to identify potential mislabeled samples and assign pseudo-labels generated by the network to them, replacing the original labels. However, existing methods encounter the following problems: 1) they typically do not evaluate the pseudo-labels and directly use all of them, and 2) empirical parameter settings are often dataset-specific. These shortcomings limit the application of these methods in real-world scenarios. In this paper, we propose the Linear Feature Source Prediction and Recombination Network (LFSPR), trying to solve the problem above by proposing a new pretext task. The pretext task is designed to build the linear connection between the high-dimensional feature and the low-dimensional feature. The source of the latter is regarded as the high-dimensional feature, which follows a non-linear head network to obtain the low-dimensional feature. The pretext task is designed in low-dimensional space by predicting the linear composition weights of the potential source. Based on the pretext task, our method can generate pseudo-labels for uncertain samples while dynamically evaluating and selecting them, rather than simply using all pseudo-labels or discarding a fixed proportion of pseudo-labels for a given dataset. To the best of our knowledge, this is the first approach in the noisy label learning domain to employ pretext task for the pseudo-labels generation, evaluation and selection. The experiments on CIFAR-10, CIFAR-100 and Clothing1M demonstrate the effectiveness of our method. Ruochen Zheng, Chuchu Han, Changxin Gao, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | S3INet: Semantic-Information Space Sharing Interaction Network for Arbitrary Shape Text DetectionabstractThe detecting arbitrary shape text is a challenging task due to the significant variation in text shape, size, and aspect ratio, as well as the complexity of scene backgrounds. The enhancing feature extraction capabilities is essential for the boosting text detection accuracy. However, traditional text feature extraction methods face several issues, including insufficient multiscale feature fusion, limited information transfer between different feature levels, and constrained receptive field expansion when using asymmetric convolutional kernels for long text detection. To address these challenges, this article introduces an arbitrarily shaped scene text detector called the semantic-information space sharing interaction network (S3INet). The proposed network leverages the semantic-information space sharing module (S3M) to generate a single-level feature map capable of capturing multiscale features with rich semantic information and prominent foreground elements. In addition, we propose the multibranch parallel asymmetric convolutional module (MPACM) group to enhance the representation of text features, thereby further enhancing text detection performance. Extensive experimental evaluations on five publicly available natural scene text datasets (CTW-1500, Total-Text, MSRA-TD500, ICDAR2015, and ICDAR2017-MLT) and two traffic text datasets (CTST-1600 and TPD) demonstrate the superiority of our method. The results indicate that S3INet significantly outperforms most existing state-of-the-art methods in both accuracy and robustness. The code will be released at: https://github.com/runminwang/S3INet. Yanbin Zhu, Xiaofei Cao, Zhenlin Zhu, Shengyou Qian, Changxin Gao, Li Liu 0002, Nong Sang |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2024 | SCTNet: Single-Branch CNN with Transformer Semantic Information for Real-Time SegmentationabstractRecent real-time semantic segmentation methods usually adopt an additional semantic branch to pursue rich long-range context. However, the additional branch incurs undesirable computational overhead and slows inference speed. To eliminate this dilemma, we propose SCTNet, a single branch CNN with transformer semantic information for real-time segmentation. SCTNet enjoys the rich semantic representations of an inference-free semantic branch while retaining the high efficiency of lightweight single branch CNN. SCTNet utilizes a transformer as the training-only semantic branch considering its superb ability to extract long-range context. With the help of the proposed transformer-like CNN block CFBlock and the semantic information alignment module, SCTNet could capture the rich semantic information from the transformer branch in training. During the inference, only the single branch CNN needs to be deployed. We conduct extensive experiments on Cityscapes, ADE20K, and COCO-Stuff-10K, and the results show that our method achieves the new state-of-the-art performance. The code and model is available at https://github.com/xzz777/SCTNet. Zhengze Xu, Dongyue Wu, Changqian Yu, Xiangxiang Chu, Nong Sang, Changxin Gao |
AAAI | 5 |
| 2024 | HR-Pro: Point-Supervised Temporal Action Localization via Hierarchical Reliability PropagationabstractPoint-supervised Temporal Action Localization (PSTAL) is an emerging research direction for label-efficient learning. However, current methods mainly focus on optimizing the network either at the snippet-level or the instance-level, neglecting the inherent reliability of point annotations at both levels. In this paper, we propose a Hierarchical Reliability Propagation (HR-Pro) framework, which consists of two reliability-aware stages: Snippet-level Discrimination Learning and Instance-level Completeness Learning, both stages explore the efficient propagation of high-confidence cues in point annotations. For snippet-level learning, we introduce an online-updated memory to store reliable snippet prototypes for each class. We then employ a Reliability-aware Attention Block to capture both intra-video and inter-video dependencies of snippets, resulting in more discriminative and robust snippet representation. For instance-level learning, we propose a point-based proposal generation approach as a means of connecting snippets and instances, which produces high-confidence proposals for further optimization at the instance level. Through multi-level reliability-aware learning, we obtain more reliable confidence scores and more accurate temporal boundaries of predicted proposals. Our HR-Pro achieves state-of-the-art performance on multiple challenging benchmarks, including an impressive average mAP of 60.3% on THUMOS14. Notably, our HR-Pro largely surpasses all previous point-supervised methods, and even outperforms several competitive fully-supervised methods. Code will be available at https://github.com/pipixin321/HR-Pro. Huaxin Zhang, Xiang Wang 0012, Xiaohao Xu, Zhiwu Qing, Changxin Gao, Nong Sang |
AAAI | 6 |
| 2024 | DFIMat: Decoupled Flexible Interactive Matting in Multi-person Scenarios
Siyi Jiao, Wenzheng Zeng, Changxin Gao, Nong Sang |
ACCV (10) | 4 |
| 2024 | Real-Time Exposure Correction via Collaborative Transformations and Adaptive SamplingabstractMost of the previous exposure correction methods learn dense pixel-wise transformations to achieve promising results, but consume huge computational resources. Recently, Learnable 3D lookup tables (3D LUTs) have demon-strated impressive performance and efficiency for image enhancement. However, these methods can only perform global transformations and fail to finely manipulate local regions. Moreover, they uniformly downsample the input image, which loses the rich color information and limits the learning of color transformation capabilities. In this paper, we present a collaborative transformation framework (CoTF) for real-time exposure correction, which integrates global transformation with pixel-wise transformations in an efficient manner. Specifically, the global transformation adjusts the overall appearance using image-adaptive 3D LUTs to provide decent global contrast and sharp details, while the pixel transformation compensates for local context. Then, a relation-aware modulation module is designed to combine these two components effectively. In addition, we propose an adaptive sampling strategy to preserve more color information by predicting the sampling intervals, thus providing higher quality input data for the learning of 3D LUTs. Extensive experiments demonstrate that our method can process high-resolution images in real-time on GPUs while achieving comparable performance against current state-of-the-art methods. The code is avail-able at https://github.com/HUST-IAL/CoTF. Ziwen Li 0005, Feng Zhang 0039, Jinpu Zhang, Yuanjie Shao, Yuehuan Wang, Nong Sang |
CVPR | 7 |
| 2024 | Hierarchical Spatio-temporal Decoupling for Text-to- Video GenerationabstractDespite diffusion models having shown powerful abilities to generate photorealistic images, generating videos that are realistic and diverse still remains in its infancy. One of the key reasons is that current methods intertwine spatial content and temporal dynamics together, leading to a notably increased complexity of text-to-video generation (T2V). In this work, we propose HiGen, a diffusion model-based method that improves performance by decoupling the spatial and temporal factors of videos from two perspectives, i.e., structure level and content level. At the structure level, we decompose the T2V task into two steps, including spatial reasoning and temporal reasoning, using a unified denoiser. Specifically, we generate spatially coherent priors using text during spatial reasoning and then generate temporally coherent motions from these priors during temporal reasoning. At the content level, we extract two subtle cues from the content of the input video that can express motion and appearance changes, respectively. These two cues then guide the model's training for generating videos, enabling flexible content variations and enhancing temporal stability. Through the decoupled paradigm, HiGen can effectively reduce the complexity of this task and generate realistic videos with semantics accuracy and motion stability. Extensive experiments demonstrate the superior performance of HiGen over the state-of-the-art T2V methods. We have released our source code and models. Zhiwu Qing, Shiwei Zhang 0001, Xiang Wang 0012, Yujie Wei 0001, Yingya Zhang, Changxin Gao, Nong Sang |
CVPR | 8 |
| 2024 | Open-Vocabulary Semantic Segmentation with Image Embedding BalancingabstractOpen-vocabulary semantic segmentation is a challenging task, which requires the model to output semantic masks of an image beyond a close-set vocabulary. Although many efforts have been made to utilize powerful CLIP models to accomplish this task, they are still easily overfitting to training classes due to the natural gaps in semantic information between training and new classes. To overcome this challenge, we propose a novel framework for open-vocabulary semantic segmentation called EBSeg, incorpo-rating an Adaptively Balanced Decoder (AdaB Decoder) and a Semantic Structure Consistency loss (SSC Loss). The AdaB Decoder is designed to generate different image embeddings for both training and new classes. Subsequently, these two types of embeddings are adaptively balanced to fully exploit their ability to recognize training classes and generalization ability for new classes. To learn a consistent semantic structure from CLIP, the SSC Loss aligns the inter-classes affinity in the image feature space with that in the text feature space of CLIP, thereby improving the generalization ability of our model. Furthermore, we employ a frozen SAM image encoder to complement the spatial information that CLIP features lack due to the low training image resolution and image-level supervision inherent in CLIP. Extensive experiments conducted across various benchmarks demonstrate that the proposed EBSeg outperforms the state-of-the-art methods. Our code and trained models will be here: https://github.com/slonetime/EBSeg. Xiangheng Shan, Dongyue Wu, Guilin Zhu, Yuanjie Shao, Nong Sang, Changxin Gao |
CVPR | 5 |
| 2024 | A Recipe for Scaling up Text-to-Video Generation with Text-free VideosabstractDiffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs in WebVid10m vs. 5B image-text pairs in LAION), considering the high cost of video captioning. Instead, it could be far easier to collect unlabeled clips from video platforms like YouTube. Motivated by this, we come up with a novel text-to-video generation framework, termed TF-T2V, which can directly learn with text-free videos. The rationale behind is to separate the process of text decoding from that of temporal modeling. To this end, we employ a content branch and a motion branch, which are jointly optimized with weights shared. Following such a pipeline, we study the effect of doubling the scale of training set (i.e., video-only WebVid10M) with some randomly collected text-free videos and are encouraged to observe the performance improvement (FID from 9.67 to 8.19 and FVD from 484 to 441), demonstrating the scalability of our approach. We also find that our model could enjoy sustainable performance gain (FID from 8.19 to 7.64 and FVD from 441 to 366) after reintroducing some text labels for training. Finally, we validate the effectiveness and generalizability of our ideology on both native text-to-video generation and compositional video synthesis paradigms. Code and models will be publicly available at here. Xiang Wang 0012, Shiwei Zhang 0001, Hangjie Yuan, Zhiwu Qing, Biao Gong, Yingya Zhang, Yujun Shen, Changxin Gao, Nong Sang |
CVPR | 9 |
| 2024 | UFineBench: Towards Text-based Person Retrieval with Ultra-fine GranularityabstractExisting text-based person retrieval datasets often have relatively coarse-grained text annotations. This hinders the model to comprehend the fine-grained semantics of query texts in real scenarios. To address this problem, we con-tribute a new benchmark named UFineBench for text-based person retrieval with ultra-fine granularity. Firstly, we construct a new dataset named UFine6926. We collect a large number of person images and manually annotate each image with two detailed textual descriptions, averaging 80.8 words each. The average word count is three to four times that of the previous datasets. In addition of standard in-domain evaluation, we also propose a spe-cial evaluation paradigm more representative of real sce-narios. It contains a new evaluation set with cross domains, cross textual granularity and cross textual styles, named UFine3C, and a new evaluation metric for accurately mea-suring retrieval ability, named mean Similarity Distribution (mSD). Moreover, we propose CFAM, a more efficient al-gorithm especially designed for text-based person retrieval with ultra fine-grained texts. It achieves fine granularity mining by adopting a shared cross-modal granularity de-coder and hard negative match mechanism. With standard in-domain evaluation, CFAM establishes competitive performance across various datasets, espe-cially on our ultra fine-grained UFine6926. Furthermore, by evaluating on UFine3C, we demonstrate that training on our UFine6926 significantly improves generalization to real scenarios compared with other coarse-grained datasets. The dataset and code will be made publicly available at https://github.com/Zplusdragon/UFineBench. Jialong Zuo, Hanyu Zhou, Feng Zhang 0039, Tianyu Guo 0001, Nong Sang, Yunhe Wang 0001, Changxin Gao |
CVPR | 6 |
| 2024 | Multi-Task Affinity Propagation Based Natural Image MattingabstractImage matting, aiming to accurately extract foreground objects by estimating their opacity against the background, has made remarkable progress through deep-learning approaches. Nevertheless, the majority of these methods require a user-defined auxiliary input, such as a trimap, which limits their applications in real-world scenarios. There are many auxiliary input-free methods that have been proposed by now, and some of them adopt a multi-task learning framework that includes a shared encoder and two separate decoders. However, these methods lack interactions between the two decoders, or interactions are implemented through simple summation or concatenation. Unfortunately, the integration of different features may cause negative transfer and limit the model performance due to the invisible information transmission process. To address the issue, we introduce the Pattern-Affinitive Propagation Module (PAP) to explicitly model cross-task propagation and task-specific propagation. Furthermore, image matting not only requires high-resolution detail features, but also semantic features. However, current CNN-based methods have limited receptive fields, making it challenging to capture global semantic features. Therefore, we design a module that integrates Dilated Convolution and Spectral Transformer (DSM), which can effectively capture global features and enhance global-local feature fusion. Extensive experiments on AM-2k and P3M-10k datasets demonstrate the superiority of our method. Renkai Zhang, Nong Sang |
ICIP | 2 |
| 2024 | Self-distilled Dual-Network with Pixel Screening Loss for Blind Image Deblurring
Ming Tian, Changxin Gao, Nong Sang |
ICPR (32) | 4 |
| 2024 | Complementary Dual-Branch Network for Space-Time Video Super-Resolution
Ming Tian, Rongsheng Luo, Changxin Gao, Nong Sang |
ICPR (32) | 5 |
| 2024 | Tunnel Try-on: Excavating Spatial-temporal Tunnels for High-quality Virtual Try-on in VideosabstractVideo try-on is challenging and has not been well tackled in previous works. The main obstacle lies in preserving the clothing details and modeling the coherent motions simultaneously. Faced with those difficulties, we address video try-on by proposing a diffusion-based framework named ''Tunnel Try-on.'' The core idea is excavating a ''focus tunnel'' in the input video that gives close-up shots around the clothing regions. We zoom in on the region in the tunnel to better preserve the fine details of the clothing. To generate coherent motions, we leverage the Kalman filter to smooth the tunnel and inject its position embedding into attention layers to improve the continuity of the generated videos. In addition, we develop an environment encoder to extract the context information outside the tunnels. Equipped with these techniques, Tunnel Try-on keeps fine clothing details and synthesizes stable and smooth videos. Demonstrating significant advancements, Tunnel Try-on could be regarded as the first attempt toward the commercial-level application of virtual try-on in videos. The project page is https://mengtingchen.github.io/tunnel-try-on-page/. Zhengze Xu, Linyu Xing, Zhonghua Zhai, Nong Sang, Jinsong Lan, Shuai Xiao 0002, Changxin Gao |
ACM Multimedia | 6 |
| 2024 | PLIP: Language-Image Pre-training for Person Representation LearningabstractLanguage-image pre-training is an effective technique for learning powerful representations in general domains. However, when directly turning to person representation learning, these general pre-training methods suffer from unsatisfactory performance. The reason is that they neglect critical person-related characteristics, i.e., fine-grained attributes and identities. To address this issue, we propose a novel language-image pre-training framework for person representation learning, termed PLIP. Specifically, we elaborately design three pretext tasks: 1) Text-guided Image Colorization, aims to establish the correspondence between the person-related image regions and the fine-grained color-part textual phrases. 2) Image-guided Attributes Prediction, aims to mine fine-grained attribute information of the person body in the image; and 3) Identity-based Vision-Language Contrast, aims to correlate the cross-modal representations at the identity level rather than the instance level. Moreover, to implement our pre-train framework, we construct a large-scale person dataset with image-text pairs named SYNTH-PEDES by automatically generating textual annotations. We pre-train PLIP on SYNTH-PEDES and evaluate our models by spanning downstream person-centric tasks. PLIP not only significantly improves existing methods on all these tasks, but also shows great ability in the zero-shot and domain generalization settings. The code, dataset and weight will be made publicly available. Jialong Zuo, Jiahao Hong, Feng Zhang 0039, Changqian Yu, Hanyu Zhou, Changxin Gao, Nong Sang, Jingdong Wang 0001 |
NeurIPS | 7 |
| 2024 | Cross-video Identity Correlating for Person Re-identification Pre-trainingabstractRecent researches have proven that pre-training on large-scale person images extracted from internet videos is an effective way in learning better representations for person re-identification. However, these researches are mostly confined to pre-training at the instance-level or single-video tracklet-level. They ignore the identity-invariance in images of the same person across different videos, which is a key focus in person re-identification. To address this issue, we propose a Cross-video Identity-cOrrelating pre-traiNing (CION) framework. Defining a noise concept that comprehensively considers both intra-identity consistency and inter-identity discrimination, CION seeks the identity correlation from cross-video images by modeling it as a progressive multi-level denoising problem. Furthermore, an identity-guided self-distillation loss is proposed to implement better large-scale pre-training by mining the identity-invariance within person images. We conduct extensive experiments to verify the superiority of our CION in terms of efficiency and performance. CION achieves significantly leading performance with even fewer training samples. For example, compared with the previous state-of-the-art ISR, CION with the same ResNet50-IBN achieves higher mAP of 93.3% and 74.3% on Market1501 and MSMT17, while only utilizing 8% training samples. Finally, with CION demonstrating superior model-agnostic ability, we contribute a model zoo named ReIDZoo to meet diverse research and application needs in this field. It contains a series of CION pre-trained models with spanning structures and parameters, totaling 32 models with 10 different structures, including GhostNet, ConvNext, RepViT, FastViT and so on. The code and models will be open-sourced. Jialong Zuo, Hanyu Zhou, Huaxin Zhang, Haoyu Wang 0003, Tianyu Guo 0001, Nong Sang, Changxin Gao |
NeurIPS | 7 |
| 2024 | EfficientMatting: Bilateral Matting Network for Real-Time Human Matting
Rongsheng Luo, Rukai Wei, Huaxin Zhang, Ming Tian, Changxin Gao, Nong Sang |
PRCV (12) | 6 |
| 2024 | CLIP-guided Prototype Modulating for Few-shot Action Recognition
Xiang Wang 0012, Shiwei Zhang 0001, Jun Cen, Changxin Gao, Yingya Zhang, Deli Zhao, Nong Sang |
Int. J. Comput. Vis. | 7 |
| 2024 | HyRSM++: Hybrid relation guided temporal set matching for few-shot action recognition
Xiang Wang 0012, Shiwei Zhang 0001, Zhiwu Qing, Zhengrong Zuo, Changxin Gao, Rong Jin 0001, Nong Sang |
Pattern Recognit. | 7 |
| 2024 | Query-centric distance modulator for few-shot classification
Wenxiao Wu, Yuanjie Shao, Changxin Gao, Jing-Hao Xue, Nong Sang |
Pattern Recognit. | 5 |
| 2024 | Difficulty-Aware Dynamic Network for Lightweight Exposure CorrectionabstractRecently, deep learning-based methods have been successfully applied to the field of exposure correction. However, most of the existing methods treat different locations of an image in the same way, ignoring the inhomogeneous recovery difficulty and spatially-varying visual patterns in the image, which is sub-optimal and not perfectly efficient. In this paper, we propose a difficulty-aware dynamic network (DDNet) for lightweight exposure correction. Specifically, we propose a difficulty-aware strategy that determines the difficulty of feature patches according to a difficulty mask. Then, only the difficult patches are further refined instead of the whole features, which greatly reduces the overall computational complexity. Moreover, in order to achieve spatially-varying processing with a minimal computational burden, we design a spatial-aware dynamic convolution (SDConv), which is generated by predicting a set of basic kernels and a spatial-aware weight map. Benefiting from these designs, our method can strike a good trade-off between performance and complexity. Extensive experiments on several datasets demonstrate that our approach outperforms the state-of-the-art methods both qualitatively and quantitatively while requiring cheaper computational costs. Ziwen Li 0005, Yuanjie Shao, Feng Zhang 0039, Jinpu Zhang, Yuehuan Wang, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | An Adaptive Post-Processing Network With the Global-Local Aggregation for Semantic SegmentationabstractCurrent semantic segmentation methods mainly focus on modeling the context of the global image to obtain high-quality segmentation results. However, they ignore the role of local image patches, which contain complementary and effective context information. In this paper, we propose an adaptive post-processing network (APPNet) for semantic segmentation based on the predictions of current methods in the global image and local image patches. The key point of APPNet is the global-local aggregation module, which models the context between global predictions and local predictions to generate accurate pixel-wise representation. Furthermore, we develop an adaptive points replacement module to compensate for the lack of fine detail in global prediction and the overconfidence in local predictions. Our method can be readily integrated into existing segmentation methods (i.e., ConvNeXt, HRNet, ViT-Adapter) with little memory and without extra modification in current models. We empirically demonstrate our method brings performance improvements across diverse datasets (i.e., Cityscapes, ADE20K, PASCAL-Context, COCO-Stuff). The code and models will be publicly available athttps://github.com/zhu-gl-ux/APPN. Guilin Zhu, Zhenlin Zhu, Changxin Gao, Li Liu 0002, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Toward Blind Flare Removal Using Knowledge-Driven Flare-Level EstimatorabstractLens flare is a common phenomenon when strong light rays arrive at the camera sensor and a clean scene is consequently mixed up with various opaque and semi-transparent artifacts. Existing deep learning methods are always constrained with limited real image pairs for training. Though recent synthesis-based approaches are found effective, synthesized pairs still deviate from the real ones as the mixing mechanism of flare artifacts and scenes in the wild always depends on a line of undetermined factors, such as lens structure, scratches, etc. In this paper, we present a new perspective from the blind nature of the flare removal task in a knowledge-driven manner. Specifically, we present a simple yet effective flare-level estimator to predict the corruption level of a flare-corrupted image. The estimated flare-level can be interpreted as additive information of the gap between corrupted images and their flare-free correspondences to facilitate a network at both training and testing stages adaptively. Besides, we utilize a flare-level modulator to better integrate the estimations into networks. We also devise a flare-aware block for more accurate flare recognition and reconstruction. Additionally, we collect a new real-world flare dataset for benchmarking, namely WiderFlare. Extensive experiments on three benchmark datasets demonstrate that our method outperforms state-of-the-art methods quantitatively and qualitatively. Haoyou Deng, Lida Li, Feng Zhang 0039, Qingbo Lu, Changxin Gao, Nong Sang |
IEEE Trans. Image Process. | 8 |
| 2024 | TTDNet: An End-to-End Traffic Text Detection Framework for Open Driving EnvironmentsabstractTraffic text detection is crucial for traffic scene understanding in intelligent transportation systems (ITS). Although natural scene text detection has been extensively studied, yielding noteworthy results, little research has focused on traffic text detection. Traffic text, as a special type of natural scene text, faces not only the general challenges of natural scene text detection but also the significant impact of false alarms from non-traffic texts on system performance. In light of these challenges, we propose an end-to-end traffic text detection framework that can effectively detect traffic text captured by in-vehicle cameras in various driving scenarios. The key contributions of our proposed approach are: (1) an Image Enhancement Module designed to remove fog and enhance low-quality images; (2) a plug-and-play Text Feature Enhancement Module; (3) a Joint Loss Function; and (4) the creation of an open driving environment traffic text dataset (named ODETT-3000) containing various traffic environments. Comprehensive experimental studies have verified that our method achieves state-of-the-art performance on traffic text datasets such as CTST-1600, TPD, and ODETT-3000, as well as promising results on public natural scene text datasets such as MSRA-TD500, ICDAR 2015, and CTW 1500, thereby showcasing the superior performance and adaptability of our method. The code and our dataset (ODETT-3000) will be available athttps://github.com/yanbin-zhu/TTDNet. Yanbin Zhu, Zhenlin Zhu, Yajun Ding, Shengyou Qian, Changxin Gao, Li Liu 0002, Nong Sang |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2024 | MAR: Masked Autoencoders for Efficient Action RecognitionabstractStandard approaches for video action recognition usually operate on full input videos, which is inefficient due to the widespread spatio-temporal redundancy in videos. The recent progress in masked video modelling, specifically VideoMAE, has shown the ability of vanilla Vision Transformers (ViT) to complement spatio-temporal contexts using limited visual content. Inspired by this, we propose Masked Action Recognition (MAR), which reduces redundant computation by discarding a proportion of patches and operating only on a portion of the videos. MAR includes two essential components:cell running maskingandbridging classifier. Specifically, to enable the ViT to perceive the details beyond the visible patches, cell running masking is used to preserve the spatio-temporal correlations in videos. This ensures that the patches at the same spatial location can be observed in turn for easy reconstructions. Additionally, we notice that, although the partially observed features can reconstruct semantically explicit invisible patches, they fail to achieve accurate classification. To address this issue, we propose a bridging classifier that can help fill the semantic gap between the ViT encoded features used for reconstruction and the specialized features used for classification. Our proposed MAR can reduce the computational cost of ViT by 53%. Extensive experiments have demonstrated that MAR consistently outperforms existing ViT models by a notable margin. Notably, we found that a ViT-Large model fine-tuned by MAR achieves comparable performance to a ViT-Huge model fine-tuned by standard training methods on both Kinetics-400 and Something-Something v2 datasets. Moreover, the computation overhead of our ViT-Large model is only 14.5% of that of the ViT-Huge model. Codes have been made availablehttps://github.com/alibaba-mmai-research/Masked-Action-Recognition. Zhiwu Qing, Shiwei Zhang 0001, Ziyuan Huang 0003, Xiang Wang 0012, Yuehuan Wang, Yiliang Lv, Changxin Gao, Nong Sang |
IEEE Trans. Multim. | 8 |
| 2024 | DIMGNet: A Transformer-Based Network for Pedestrian Reidentification With Multi-Granularity Information Mutual GainabstractPedestrian reidentification (ReID) is a challenging task that involves identifying and retrieving specific pedestrians across different cameras and scenes. This problem has significant implications for security surveillance, and has thus received substantial attention in recent years. However, traditional convolutional neural networks (CNNs) have limited receptive fields and cannot capture global information. Moreover, transformer networks, which excel in long-range feature capture, are prone to accuracy degradation due to loss of details. To address these limitations, we propose a transformer-based pedestrian ReID network with double-branch information mutual gain (DIMGNet), which leverages hierarchical parallel levels to support multi-granularity feature information mutual gain. Our model also incorporates an auxiliary camera information (ACI) module to improve feature representation ability. We further embed a cross-attention mechanism into the architecture to enhance mutual gain between multi-granularity features and improve feature discrimination. Finally, we introduce a shuffling technique to increase the robustness of the extracted features. We evaluate the proposed method on several benchmark datasets, including Market-1501[1], MSMT17[2], DukeMTMC-reID [3], and Occluded-Duke [4], achieving$mAP$values of 90.7%, 68.4%, 83.7%, and 60.6%, respectively. Our method outperforms most state-of-the-art methods, demonstrating the effectiveness of our method. The code will be publicly released athttps://github.com/ZhenlinZhu/DIMGNet. Zhenlin Zhu, Yanbin Zhu, Yongzhong Liao, Yajun Ding, Changxin Gao, Nong Sang |
IEEE Trans. Multim. | 9 |
| 2023 | MoLo: Motion-Augmented Long-Short Contrastive Learning for Few-Shot Action RecognitionabstractCurrent state-of-the-art approaches for few-shot action recognition achieve promising performance by conducting frame-level matching on learned visual features. However, they generally suffer from two limitations: i) the matching procedure between local frames tends to be inaccurate due to the lack of guidance to force long-range temporal perception; ii) explicit motion learning is usually ignored, leading to partial information loss. To address these issues, we develop a Motion-augmented Long-short Contrastive Learning (MoLo) method that contains two crucial components, including a long-short contrastive objective and a motion autodecoder. Specifically, the long-short contrastive objective is to endow local frame features with long-form temporal awareness by maximizing their agreement with the global token of videos belonging to the same class. The motion autodecoder is a lightweight architecture to reconstruct pixel motions from the differential features, which explicitly embeds the network with motion dynamics. By this means, MoLo can simultaneously learn long-range temporal context and motion cues for comprehensive few-shot matching. To demonstrate the effectiveness, we evaluate MoLo on five standard benchmarks, and the results show that MoLo favorably outperforms recent advanced methods. The source code is available at https://github.com/alibaba-mmai-research/MoLo. Xiang Wang 0012, Shiwei Zhang 0001, Zhiwu Qing, Changxin Gao, Yingya Zhang, Deli Zhao, Nong Sang |
CVPR | 7 |
| 2023 | Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer LearningabstractRecently, large-scale pre-trained language-image models like CLIP have shown extraordinary capabilities for understanding spatial contents, but naively transferring such models to video recognition still suffers from unsatisfactory temporal modeling capabilities. Existing methods insert tunable structures into or in parallel with the pre-trained model, which either requires back-propagation through the whole pre-trained model and is thus resource-demanding, or is limited by the temporal reasoning capability of the pre-trained structure. In this work, we present DiST, which disentangles the learning of spatial and temporal aspects of videos. Specifically, DiST uses a dual-encoder structure, where a pre-trained foundation model acts as the spatial encoder, and a lightweight network is introduced as the temporal encoder. An integration branch is inserted between the encoders to fuse spatio-temporal information. The disentangled spatial and temporal learning in DiST is highly efficient because it avoids the back-propagation of massive pre-trained parameters. Meanwhile, we empirically show that disentangled learning with an extra network for integration benefits both spatial and temporal understanding. Extensive experiments on five benchmarks show that DiST delivers better performance than existing state-of-the-art methods by convincing gaps. When pre-training on the large-scale Kinetics-710, we achieve 89.7% on Kinetics-400 with a frozen ViT-L model, which verifies the scalability of DiST. Codes and models can be found in https://github.com/alibaba-mmai-research/DiST. Zhiwu Qing, Shiwei Zhang 0001, Ziyuan Huang 0003, Yingya Zhang, Changxin Gao, Deli Zhao, Nong Sang |
ICCV | 7 |
| 2023 | Towards General Low-Light Raw Noise Synthesis and ModelingabstractModeling and synthesizing low-light raw noise is a fundamental problem for computational photography and image processing applications. Although most recent works have adopted physics-based models to synthesize noise, the signal-independent noise in low-light conditions is far more complicated and varies dramatically across camera sensors, which is beyond the description of these models. To address this issue, we introduce a new perspective to synthesize the signal-independent noise by a generative model. Specifically, we synthesize the signal-dependent and signal-independent noise in a physics-and learning-based manner, respectively. In this way, our method can be considered as a general model, that is, it can simultaneously learn different noise characteristics for different ISO levels and generalize to various sensors. Subsequently, we present an effective multi-scale discriminator termed Fourier transformer discriminator (FTD) to distinguish the noise distribution accurately. Additionally, we collect a new low-light raw denoising (LRD) dataset for training and benchmarking. Qualitative validation shows that the noise generated by our proposed noise model can be highly similar to the real noise in terms of distribution. Furthermore, extensive denoising experiments demonstrate that our method performs favorably against state-of-the-art methods on different sensors. Feng Zhang 0039, Qingbo Lu, Changxin Gao, Nong Sang |
ICCV | 7 |
| 2023 | Lookup Table meets Local Laplacian Filter: Pyramid Reconstruction Network for Tone MappingabstractTone mapping aims to convert high dynamic range (HDR) images to low dynamic range (LDR) representations, a critical task in the camera imaging pipeline. In recent years, 3-Dimensional LookUp Table (3D LUT) based methods have gained attention due to their ability to strike a favorable balance between enhancement performance and computational efficiency. However, these methods often fail to deliver satisfactory results in local areas since the look-up table is a global operator for tone mapping, which works based on pixel values and fails to incorporate crucial local information. To this end, this paper aims to address this issue by exploring a novel strategy that integrates global and local operators by utilizing closed-form Laplacian pyramid decomposition and reconstruction. Specifically, we employ image-adaptive 3D LUTs to manipulate the tone in the low-frequency image by leveraging the specific characteristics of the frequency information. Furthermore, we utilize local Laplacian filters to refine the edge details in the high-frequency components in an adaptive manner. Local Laplacian filters are widely used to preserve edge details in photographs, but their conventional usage involves manual tuning and fixed implementation within camera imaging pipelines or photo editing tools. We propose to learn parameter value maps progressively for local Laplacian filters from annotated data using a lightweight network. Our model achieves simultaneous global tone manipulation and local edge detail preservation in an end-to-end manner. Extensive experimental results on two benchmark datasets demonstrate that the proposed method performs favorably against state-of-the-art methods. Feng Zhang 0039, Ming Tian, Qingbo Lu, Changxin Gao, Nong Sang |
NeurIPS | 7 |
| 2023 | Self-supervised Low-Light Image Enhancement via Histogram Equalization Prior
Feng Zhang 0039, Yuanjie Shao, Yishi Sun, Changxin Gao, Nong Sang |
PRCV (11) | 5 |
| 2023 | Camera distance helps 3D hand pose estimated from a single RGB imageabstractMost existing methods for RGB hand pose estimation use root-relative 3D coordinates for supervision. However, such supervision neglects the distance between the camera and the object (i.e., the hand). The camera distance is especially important under a perspective camera, which controls the depth-dependent scaling of the perspective projection. As a result, the same hand pose, with different camera distances can be projected into different 2D shapes by the same perspective camera. Neglecting such important information results in ambiguities in recovering 3D poses from 2D images. In this article, we propose a camera projection learning module (CPLM) that uses the scale factor contained in the camera distance to associate 3D hand pose with 2D UV coordinates, which facilities to further optimize the accuracy of the estimated hand joints. Specifically, following the previous work, we use a two-stage RGB-to-2D and 2D-to-3D method to estimate 3D hand pose and embed a graph convolutional network in the second stage to leverage the information contained in the complex non-Euclidean structure of 2D hand joints. Experimental results demonstrate that our proposed method surpasses state-of-the-art methods on the benchmark dataset RHD and obtains competitive results on the STB and D+O datasets. Moran Li, Yuan Gao 0015, Changxin Gao, Nong Sang |
Graph. Model. | 8 |
| 2023 | Cross-domain few-shot action recognition with unlabeled videos
Xiang Wang 0012, Shiwei Zhang 0001, Zhiwu Qing, Yiliang Lv, Changxin Gao, Nong Sang |
Comput. Vis. Image Underst. | 6 |
| 2023 | Dual Attention Mechanism Based Outline Loss for Image Stylization
Pengqi Tu, Nong Sang |
Neural Process. Lett. | 2 |
| 2023 | DMRNet++: Learning Discriminative Features With Decoupled Networks and Enriched Pairs for One-Step Person SearchabstractPerson search aims at localizing and recognizing query persons from raw video frames, which is a combination of two sub-tasks, i.e., pedestrian detection and person re-identification. The dominant fashion is termed as the one-step person search that jointly optimizes detection and identification in a unified network, exhibiting higher efficiency. However, there remain major challenges: (i) conflicting objectives of multiple sub-tasks under the shared feature space, (ii) inconsistent memory bank caused by the limited batch size, (iii) underutilized unlabeled identities during the identification learning. To address these issues, we develop an enhanced decoupled and memory-reinforced network (DMRNet++). First, we simplify the standard tightly coupled pipelines and establish a task-decoupled framework (TDF). Second, we build a memory-reinforced mechanism (MRM), with a slow-moving average of the network to better encode the consistency of the memorized features. Third, considering the potential of unlabeled samples, we model the recognition process as semi-supervised learning. An unlabeled-aided contrastive loss (UCL) is developed to boost the identification feature learning by exploiting the aggregation of unlabeled identities. Experimentally, the proposed DMRNet++ obtains the mAP of 94.5% and 52.1% on CUHK-SYSU and PRW datasets, which exceeds most existing methods. Chuchu Han, Zhedong Zheng, Dongdong Yu, Zehuan Yuan, Changxin Gao, Nong Sang, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2023 | Self-Supervised Learning from Untrimmed Videos via Hierarchical ConsistencyabstractNatural untrimmed videos provide rich visual content for self-supervised learning. Yet most previous efforts to learn spatio-temporal representations rely on manually trimmed videos, such as Kinetics dataset (Carreira and Zisserman 2017), resulting in limited diversity in visual patterns and limited performance gains. In this work, we aim to improve video representations by leveraging the rich information in natural untrimmed videos. For this purpose, we propose learning a hierarchy of temporal consistencies in videos, i.e., visual consistency and topical consistency, corresponding respectively to clip pairs that tend to be visually similar when separated by a short time span, and clip pairs that share similar topics when separated by a long time span. Specifically, we present a Hierarchical Consistency (HiCo++) learning framework, in which the visually consistent pairs are encouraged to share the same feature representations by contrastive learning, while topically consistent pairs are coupled through a topical classifier that distinguishes whether they are topic-related, i.e., from the same untrimmed video. Additionally, we impose a gradual sampling algorithm for the proposed hierarchical consistency learning, and demonstrate its theoretical superiority. Empirically, we show that HiCo++ can not only generate stronger representations on untrimmed videos, but also improve the representation quality when applied to trimmed videos. This contrasts with standard contrastive learning, which fails to learn powerful representations from untrimmed videos. Source code will be made available here. Zhiwu Qing, Shiwei Zhang 0001, Ziyuan Huang 0003, Yi Xu 0008, Xiang Wang 0012, Changxin Gao, Rong Jin 0001, Nong Sang |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Improving the Generalization of MAML in Few-Shot Classification via Bi-Level ConstraintabstractFew-shot classification (FSC), which aims to identify novel classes in the presence of a few labeled samples, has drawn vast attention in recent years. One of the representative few-shot classification methods is model-agnostic meta-learning (MAML), which focuses on learning an initialization that can quickly adapt to novel categories with a few annotated samples. However, due to insufficient samples, MAML can easily fall into the dilemma of overfitting. Most existing MAML-based methods either improve the inner-loop update rule to achieve better generalization or constrain the outer-loop optimization to learn a more desirable initialization, without considering improving the two optimization processes jointly, resulting in unsatisfactory performance. In this paper, we propose a bi-level constrained MAML (BLC-MAML) method for few-shot classification. Specifically, in the inner-loop optimization, we introduce a supervised contrastive loss to constrain the adaptation procedure, which can effectively increase the intra-class aggregation and inter-class separability, thus improving the generalization of the adapted model. In the case of the outer loop, we propose a cross-task metric (CTM) loss to constrain the adapted model to perform well on the different few-shot task. The CTM loss can enforce the adapted model to learn more discriminative and generalized feature representations, further boosting the generalization of the learned initialization. By simultaneously constraining the bi-level optimization procedure, the proposed BLC-MAML can learn an initialization with better generalization. Extensive experiments on several FSC benchmarks show that our method can effectively improve the performance of MAML under both the within-domain and cross-domain settings, and also perform favorably against the state-of-the-art FSC algorithms. Yuanjie Shao, Wenxiao Wu, Xinge You, Changxin Gao, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Pose-Guided Hierarchical Semantic Decomposition and Composition for Human ParsingabstractHuman parsing is a fine-grained semantic segmentation task, which needs to understand human semantic parts. Most existing methods model human parsing as a general semantic segmentation, which ignores the inherent relationship among hierarchical human parts. In this work, we propose a pose-guided hierarchical semantic decomposition and composition framework for human parsing. Specifically, our method includes a semantic maintained decomposition and composition (SMDC) module and a pose distillation (PC) module. SMDC progressively disassembles the human body to focus on the more concise regions of interest in the decomposition stage and then gradually assembles human parts under the guidance of pose information in the composition stage. Notably, SMDC maintains the atomic semantic labels during both stages to avoid the error propagation issue of the hierarchical structure. To further take advantage of the relationship of human parts, we introduce pose information as explicit guidance for the composition. However, the discrete structure prediction in pose estimation is against the requirement of the continuous region in human parsing. To this end, we design a PC module to broadcast the maximum responses of pose estimation to form the continuous structure in the way of knowledge distillation. The experimental results on the look-into-person (LIP) and PASCAL-Person-Part datasets demonstrate the superiority of our method compared with the state-of-the-art methods, that is, 55.21% mean Intersection of Union (mIoU) on LIP and 69.88% mIoU on PASCAL-Person-Part. Changqian Yu, Jin-Gang Yu, Changxin Gao, Nong Sang |
IEEE Trans. Cybern. | 5 |
| 2023 | Conditional Boundary Loss for Semantic SegmentationabstractImproving boundary segmentation results has recently attracted increasing attention in the field of semantic segmentation. Since existing popular methods usually exploit the long-range context, the boundary cues are obscure in the feature space, leading to poor boundary results. In this paper, we propose a novel conditional boundary loss (CBL) for semantic segmentation to improve the performance of the boundaries. The CBL creates a unique optimization goal for each boundary pixel, conditioned on its surrounding neighbors. The conditional optimization of the CBL is easy yet effective. In contrast, most previous boundary-aware methods have difficult optimization goals or may cause potential conflicts with the semantic segmentation task. Specifically, the CBL enhances the intra-class consistency and inter-class difference, by pulling each boundary pixel closer to its unique local class center and pushing it away from its different-class neighbors. Moreover, the CBL filters out noisy and incorrect information to obtain precise boundaries, since only surrounding neighbors that are correctly classified participate in the loss calculation. Our loss is a plug-and-play solution that can be used to improve the boundary segmentation performance of any semantic segmentation network. We conduct extensive experiments on ADE20K, Cityscapes, and Pascal Context, and the results show that applying the CBL to various popular segmentation networks can significantly improve the mIoU and boundary F-score performance. Dongyue Wu, Zilin Guo, Aoyan Li, Changqian Yu, Changxin Gao, Nong Sang |
IEEE Trans. Image Process. | 6 |
| 2023 | ParamCrop: Parametric Cubic Cropping for Video Contrastive LearningabstractThe central idea of contrastive learning is to discriminate between different instances and force different views from the same instance to share the same representation. To avoid trivial solutions, augmentation plays an important role in generating different views, among which random cropping is shown to be effective for the model to learn a generalized and robust representation. Commonly used random crop operation keeps the distribution of the difference between two views unchanged along the training process. In this work, we show that adaptively controlling the disparity between two augmented views along the training process enhances the quality of the learned representations. Specifically, we present a parametric cubic cropping operation, ParamCrop, for video contrastive learning, which automatically crops a 3D cubic by differentiable 3D affine transformations. ParamCrop is trained simultaneously with the video backbone using an adversarial objective, so that it learns to increase the contrastive loss and thus gradually reduces the shared contents between two cropped views. Experiments show that this adaptive and gradual increase in the disparity yielded by ParamCrop is beneficial to learning a strong and generalized representation for downstream tasks, which is shown to be effective on multiple contrastive learning frameworks and video backbones. Zhiwu Qing, Ziyuan Huang 0003, Shiwei Zhang 0001, Mingqian Tang, Changxin Gao, Rong Jin 0001, Marcelo H. Ang, Nong Sang |
IEEE Trans. Multim. | 8 |
| 2022 | Multi-Centroid Representation Network for Domain Adaptive Person Re-IDabstractRecently, many approaches tackle the Unsupervised Domain Adaptive person re-identification (UDA re-ID) problem through pseudo-label-based contrastive learning. During training, a uni-centroid representation is obtained by simply averaging all the instance features from a cluster with the same pseudo label. However, a cluster may contain images with different identities (label noises) due to the imperfect clustering results, which makes the uni-centroid representation inappropriate. In this paper, we present a novel Multi-Centroid Memory (MCM) to adaptively capture different identity information within the cluster. MCM can effectively alleviate the issue of label noises by selecting proper positive/negative centroids for the query image. Moreover, we further propose two strategies to improve the contrastive learning process. First, we present a Domain-Specific Contrastive Learning (DSCL) mechanism to fully explore intra-domain information by comparing samples only from the same domain. Second, we propose Second-Order Nearest Interpolation (SONI) to obtain abundant and informative negative samples. We integrate MCM, DSCL, and SONI into a unified framework named Multi-Centroid Representation Network (MCRN). Extensive experiments demonstrate the superiority of MCRN over state-of-the-art approaches on multiple UDA re-ID tasks and fully unsupervised re-ID tasks. Tengteng Huang, Chi Zhang 0026, Yuanjie Shao, Chuchu Han, Changxin Gao, Nong Sang |
AAAI | 8 |
| 2022 | Learning from Untrimmed Videos: Self-Supervised Video Representation Learning with Hierarchical ConsistencyabstractNatural videos provide rich visual contents for selfsupervised learning. Yet most existing approaches for learning spatio-temporal representations rely on manually trimmed videos, leading to limited diversity in visual patterns and limited performance gain. In this work, we aim to learn representations by leveraging more abundant information in untrimmed videos. To this end, we propose to learn a hierarchy of consistencies in videos, i.e., visual consistency and topical consistency, corresponding respectively to clip pairs that tend to be visually similar when separated by a short time span and share similar topics when separated by a long time span. Specifically, a hierarchical consistency learning framework HiCo is presented, where the visually consistent pairs are encouraged to have the same representation through contrastive learning, while the topically consistent pairs are coupled through a topical classifier that distinguishes whether they are topicrelated. Further, we impose a gradual sampling algorithm for proposed hierarchical consistency learning, and demonstrate its theoretical superiority. Empirically, we show that not only HiCo can generate stronger representations on untrimmed videos, it also improves the representation quality when applied to trimmed videos. This is in contrast to standard contrastive learning that fails to learn appropriate representations from untrimmed videos. Zhiwu Qing, Shiwei Zhang 0001, Ziyuan Huang 0003, Yi Xu 0008, Xiang Wang 0012, Mingqian Tang, Changxin Gao, Rong Jin 0001, Nong Sang |
CVPR | 9 |
| 2022 | Hybrid Relation Guided Set Matching for Few-shot Action RecognitionabstractCurrent few-shot action recognition methods reach impressive performance by learning discriminative features for each video via episodic training and designing various temporal alignment strategies. Nevertheless, they are limited in that (a) learning individual features without considering the entire task may lose the most relevant information in the current episode, and (b) these alignment strategies may fail in misaligned instances. To overcome the two limitations, we propose a novel Hybrid Relation guided Set Matching (HyRSM) approach that incorporates two key components: hybrid relation module and set matching metric. The purpose of the hybrid relation module is to learn task-specific embeddings by fully exploiting associated relations within and cross videos in an episode. Built upon the task-specific features, we reformulate distance measure between query and support videos as a set matching problem and further design a bidirectional Mean Hausdorff Metric to improve the resilience to misaligned instances. By this means, the proposed HyRSM can be highly informative and flexible to predict query categories under the few-shot settings. We evaluate HyRSM on six challenging benchmarks, and the experimental results show its superiority over the state-of-the-art methods by a convincing margin. Project page: https://hyrsm-cvpr2022.github.io/. Xiang Wang 0012, Shiwei Zhang 0001, Zhiwu Qing, Mingqian Tang, Zhengrong Zuo, Changxin Gao, Rong Jin 0001, Nong Sang |
CVPR | 8 |
| 2022 | Hierarchical Feature Embedding for Visual Tracking
Zhixiong Pi, Weitao Wan, Changxin Gao, Nong Sang, Chen Li 0025 |
ECCV (22) | 5 |
| 2022 | Applying Deep Learning to Known-Plaintext Attack on Chaotic Image Encryption SchemesabstractIn this paper, we demonstrate that traditional chaotic encryption schemes are vulnerable to the known-plaintext attack (KPA) with deep learning. Considering the decryption process as image restoration based on deep learning, we apply Convolutional Neural Network to perform known-plaintext attack on chaotic cryptosystems. We design a network to learn the operation mechanism of chaotic cryptosystems, and utilize the trained network as the decryption system. To prove the effectiveness, we select three existing chaotic encryption schemes as the attacked targets. The experimental results demonstrate that deep learning can be applied to known-plaintext attack against chaotic cryptosystems successfully. Compared with traditional attack methods for chaotic cryptosystems, the proposed method shows obvious advantages: (1) One neural network may be applied to cryptanalysis of various chaotic cryptosystems, not limited to specific one; (2) the proposed method is significantly convenient and cost-efficient. This paper provides a new idea for the cryptanalysis of chaotic cryptosystems. Fusen Wang, Jun Sang, Chunlin Huang, Hong Xiang, Nong Sang |
ICASSP | 6 |
| 2022 | RFNet: A Refinement Network for Semantic SegmentationabstractAs one of the basic tasks of computer vision, semantic segmentation is widely used in many fields, e.g., medical images parsing, scene parsing, autonomous driving, etc. In the current mainstream approaches, downsampling or patching operation is required to ensure that the GPU memory is not overloaded for dealing with the high-resolution input images. However, the corresponding cost is the lack of details in the final segmentation map. In this work, we proposed RFNet, a refinement network which resolves the lack of detailed information in coarse predictions by fusing the coarse predictions and the fine predictions gained by fine input image patches. There are three key characteristics: (i) designing a spatial information extraction module which can efficiently process information in coarse and fine feature maps at spatial level. (ii) proposing an auxiliary-fusion information branch calculated from the prediction maps, which contribute to refine predictions. (iii) designing a boundary auxiliary loss function in the training process, which makes the model pay more attention to those pixels belonging to the boundary of objects. We show the superiority of the proposed RFNet on the Cityscapes dataset, the experimental results illustrate that ours RFNet performance outperforms other state-of-the-art approaches with low computation consumption. The codes will be available at: https://github.com/zhu-gl-ux/RFNet. Guilin Zhu, Chang Han, Yajun Ding, Minghao Liu 0014, Nong Sang |
ICPR | 8 |
| 2022 | Learning a Condensed Frame for Memory-Efficient Video Class-Incremental LearningabstractRecent incremental learning for action recognition usually stores representative videos to mitigate catastrophic forgetting. However, only a few bulky videos can be stored due to the limited memory. To address this problem, we propose FrameMaker, a memory-efficient video class-incremental learning approach that learns to produce a condensed frame for each selected video. Specifically, FrameMaker is mainly composed of two crucial components: Frame Condensing and Instance-Specific Prompt. The former is to reduce the memory cost by preserving only one condensed frame instead of the whole video, while the latter aims to compensate the lost spatio-temporal details in the Frame Condensing stage. By this means, FrameMaker enables a remarkable reduction in memory but keep enough information that can be applied to following incremental tasks. Experimental results on multiple challenging benchmarks, i.e., HMDB51, UCF101 and Something-Something V2, demonstrate that FrameMaker can achieve better performance to recent advanced methods while consuming only 20% memory. Additionally, under the same memory consumption conditions, FrameMaker significantly outperforms existing state-of-the-arts by a convincing margin. Yixuan Pei, Zhiwu Qing, Jun Cen, Xiang Wang 0012, Shiwei Zhang 0001, Yaxiong Wang, Mingqian Tang, Nong Sang, Xueming Qian |
NeurIPS | 8 |
| 2022 | Unsupervised Image Translation with GAN Prior
Pengqi Tu, Changxin Gao, Nong Sang |
PRCV (1) | 3 |
| 2022 | MGSNet: A multi-scale and gated spatial attention network for crowd counting
Jun Sang, Zhongyuan Wu, Fusen Wang, Xiaofeng Xia, Nong Sang |
Appl. Intell. | 7 |
| 2022 | Implicit Neural Deformation for Sparse-View Face ReconstructionabstractAbstract In this work, we present a new method for 3D face reconstruction from sparse‐view RGB images. Unlike previous methods which are built upon 3D morphable models (3DMMs) with limited details, we leverage an implicit representation to encode rich geometric features. Our overall pipeline consists of two major components, including a geometry network, which learns a deformable neural signed distance function (SDF) as the 3D face representation, and a rendering network, which learns to render on‐surface points of the neural SDF to match the input images via self‐supervised optimization. To handle in‐the‐wild sparse‐view input of the same target with different expressions at test time, we propose residual latent code to effectively expand the shape space of the learned implicit face representation as well as a novel view‐switch loss to enforce consistency among different views. Our experimental results on several benchmark datasets demonstrate that our approach outperforms alternative baselines and achieves superior face reconstruction results compared to state‐of‐the‐art methods. Moran Li, Mengtian Li 0003, Nong Sang, Chongyang Ma |
Comput. Graph. Forum | 5 |
| 2022 | Hybrid attention network based on progressive embedding scale-context for crowd counting
Fusen Wang, Jun Sang, Zhongyuan Wu, Nong Sang |
Inf. Sci. | 5 |
| 2022 | Single image based 3D human pose estimation via uncertainty learning
Chuchu Han, Xin Yu 0002, Changxin Gao, Nong Sang, Yi Yang 0001 |
Pattern Recognit. | 4 |
| 2022 | D2T: A Framework For transferring detection to tracking
Huai Qin, Changqian Yu, Changxin Gao, Nong Sang |
Pattern Recognit. | 4 |
| 2022 | Norm-Aware Margin Assignment for Person Re-IdentificationabstractMargin-based metric losses have shown great success in Person Re-identification and Face Verification. But most existing works adopt a fixed class-level margin regardless of the difference between each training sample. This paper proposes a Norm-Aware Margin Assignment (NAMA) scheme to dynamically adjust the weight of each sample during training. Combined with the existing margin-based classification losses, NAMA improves the robustness of feature embedding by assigning larger margins to more recognizable samples. NAMA is a fully trainable module that automatically models the correlation between the optimal margin and image quality during back-propagation without supervision. To stabilize the training and make the assigned margin more controllable, we introduce a margin re-balance mechanism to align the expectation of learned margins to a pre-defined value. Extensive experiments on three popular ReID benchmarks validate the effectiveness of our NAMA method. Code will be publicly available at: https://github.com/huangzongheng/NAMA. Zongheng Huang, Botao He, Changxin Gao, Nong Sang |
IEEE Signal Process. Lett. | 5 |
| 2022 | Instance-Based Feature Pyramid for Visual Object TrackingabstractThe deep learning based methods have improved the visual tracking precision significantly. However, the background distraction and the high precise localization remain challenging problems. Despite that some methods have fused the deep and shallow layer features to solve these problems, the existing fusion methods, like simply concatenating or adding the features from the different layers, cannot take the advantage of both the deep and shallow layer features fully. In this paper, we propose a new adaptive feature fusion method, called the instance-based feature pyramid (IBFP) to obtain the discriminative high-resolution feature, which not only inherits the discriminative information from the deep layer feature, but also keeps the high precision localization information of the shallow layer feature. For utilizing the deep and shallow features effectively, we design an instance-based upsampling (IBU) module to fuse them, and a compressed space channel selection (CSCS) module to re-weight the feature channels adaptively. We insert the IBU and CSCS modules in the Siamese tracker for end-to-end training and testing. By using the proposed IBU and CSCS modules, we fuse the deep and shallow features in a series manner. Experiments on large-scale benchmark datasets demonstrate that the proposed modules boost the capabilities of distinguishing the targets and the similar distractors and perform favorably against the state-of-the-art. Zhixiong Pi, Yuanjie Shao, Changxin Gao, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Decoupled and Memory-Reinforced Networks: Towards Effective Feature Learning for One-Step Person SearchabstractThe goal of person search is to localize and match query persons from scene images. For high efficiency, one-step methods have been developed to jointly handle the pedestrian detection and identification sub-tasks using a single network. There are two major challenges in the current one-step approaches. One is the mutual interference between the optimization objectives of multiple sub-tasks. The other is the sub-optimal identification feature learning caused by small batch size when end-to-end training. To overcome these problems, we propose a decoupled and memory-reinforced network (DMRNet). Specifically, to reconcile the conflicts of multiple objectives, we simplify the standard tightly coupled pipelines and establish a deeply decoupled multi-task learning framework. Further, we build a memory-reinforced mechanism to boost the identification feature learning. By queuing the identification features of recently accessed instances into a memory bank, the mechanism augments the similarity pair construction for pairwise metric learning. For better encoding consistency of the stored features, a slow-moving average of the network is applied for extracting these features. In this way, the dual networks reinforce each other and converge to robust solution states. Experimentally, the proposed method obtains 93.2% and 46.9% mAP on CUHK-SYSU and PRW datasets, which exceeds all the existing one-step methods. Chuchu Han, Zhedong Zheng, Changxin Gao, Nong Sang, Yi Yang 0001 |
AAAI | 4 |
| 2021 | Exploiting Learnable Joint Groups for Hand Pose EstimationabstractIn this paper, we propose to estimate 3D hand pose by recovering the 3D coordinates of joints in a group-wise manner, where less-related joints are automatically categorized into different groups and exhibit different features. This is different from the previous methods where all the joints are considered holistically and share the same feature. The benefits of our method are illustrated by the principle of multi-task learning (MTL), i.e., by separating less-related joints into different groups (as different tasks), our method learns different features for each of them, therefore efficiently avoids the negative transfer (among less related tasks/groups of joints). The key of our method is a novel binary selector that automatically selects related joints into the same group. We implement such a selector with binary values stochastically sampled from a Concretedistribution, which is constructed using Gumbel softmax on trainable parameters. This enables us to preserve the differentiable property of the whole network. We further exploit features from those less-related groups by carrying out an additional feature fusing scheme among them, to learn more discriminative features. This is realized by implementing multiple 1x1 convolutions on the concatenated features, where each joint group contains a unique 1x1convolution for feature fusion. The detailed ablation analysis and the extensive experiments on several benchmark datasets demonstrate the promising performance of the proposed method over the state-of-the-art (SOTA) methods. Besides, our method achieves top-1 among all the methods that do not exploit the dense 3D shape labels on the most recently released FreiHAND competition at the submission date. The source code and models are available at https://github.com/moranli-aca/LearnableGroups-Hand. Moran Li, Nong Sang |
AAAI | 3 |
| 2021 | Temporal Context Aggregation Network for Temporal Action Proposal RefinementabstractTemporal action proposal generation aims to estimate temporal intervals of actions in untrimmed videos, which is a challenging yet important task in the video understanding field. The proposals generated by current methods still suffer from inaccurate temporal boundaries and inferior confidence used for retrieval owing to the lack of efficient temporal modeling and effective boundary context utilization. In this paper, we propose Temporal Context Aggregation Network (TCANet) to generate high-quality action proposals through "local and global" temporal context aggregation and complementary as well as progressive boundary refinement. Specifically, we first design a Local-Global Temporal Encoder (LGTE), which adopts the channel grouping strategy to efficiently encode both "local and global" temporal inter-dependencies. Furthermore, both the boundary and internal context of proposals are adopted for frame-level and segment-level boundary regressions, respectively. Temporal Boundary Regressor (TBR) is designed to combine these two regression granularities in an end-to-end fashion, which achieves the precise boundaries and reliable confidence of proposals through progressive refinement. Extensive experiments are conducted on three challenging datasets: HACS, ActivityNet-v1.3, and THUMOS-14, where TCANet can generate proposals with high precision and recall. By combining with the existing action classifier, TCANet can obtain remarkable temporal action detection performance compared with other methods. Not surprisingly, the proposed TCANet won the 1stplace in the CVPR 2020 - HACS challenge leaderboard on temporal action localization task. Zhiwu Qing, Haisheng Su, Weihao Gan, Wei Wu 0021, Xiang Wang 0012, Yu Qiao 0001, Changxin Gao, Nong Sang |
CVPR | 10 |
| 2021 | Self-Supervised Learning for Semi-Supervised Temporal Action ProposalabstractSelf-supervised learning presents a remarkable performance to utilize unlabeled data for various video tasks. In this paper, we focus on applying the power of self-supervised methods to improve semi-supervised action proposal generation. Particularly, we design an effective Self-supervised Semi-supervised Temporal Action Proposal (SSTAP) framework. The SSTAP contains two crucial branches, i.e., temporal-aware semi-supervised branch and relation-aware self-supervised branch. The semi-supervised branch improves the proposal model by introducing two temporal perturbations, i.e., temporal feature shift and temporal feature flip, in the mean teacher framework. The self-supervised branch defines two pretext tasks, including masked feature reconstruction and clip-order prediction, to learn the relation of temporal clues. By this means, SSTAP can better explore unlabeled videos, and improve the discriminative abilities of learned action features. We extensively evaluate the proposed SSTAP on THUMOS14 and ActivityNet v1.3 datasets. The experimental results demonstrate that SSTAP significantly outperforms state-of-the-art semi-supervised methods and even matches fully-supervised methods. Code is available at https://github.com/wangxiang1230/SSTAP. Xiang Wang 0012, Shiwei Zhang 0001, Zhiwu Qing, Yuanjie Shao, Changxin Gao, Nong Sang |
CVPR | 6 |
| 2021 | Lite-HRNet: A Lightweight High-Resolution NetworkabstractWe present an efficient high-resolution network, Lite-HRNet, for human pose estimation. We start by simply applying the efficient shuffle block in ShuffleNet to HRNet (high-resolution network), yielding stronger performance over popular lightweight networks, such as MobileNet, ShuffleNet, and Small HRNet. We find that the heavily-used pointwise (1 × 1) convolutions in shuffle blocks become the computational bottleneck. We introduce a lightweight unit, conditional channel weighting, to replace costly pointwise (1 × 1) convolutions in shuffle blocks. The complexity of channel weighting is linear w.r.t the number of channels and lower than the quadratic time complexity for pointwise convolutions. Our solution learns the weights from all the channels and over multiple resolutions that are readily available in the parallel branches in HRNet. It uses the weights as the bridge to exchange information across channels and resolutions, compensating the role played by the pointwise (1 × 1) convolution. Lite-HRNet demonstrates superior results on human pose estimation over popular lightweight networks. Moreover, Lite-HRNet can be easily applied to semantic segmentation task in the same lightweight manner. The code and models have been publicly available at https://github.com/HRNet/Lite-HRNet. Changqian Yu, Bin Xiao 0004, Changxin Gao, Lu Yuan 0001, Lei Zhang 0001, Nong Sang, Jingdong Wang 0001 |
CVPR | 6 |
| 2021 | GC-MRNet: Gated Cascade Multi-stage Regression Network for Crowd Counting
Jun Sang, Jinghan Tan, Zhongyuan Wu, Nong Sang |
ICANN (2) | 6 |
| 2021 | Weakly Supervised Person Search with Region Siamese NetworksabstractSupervised learning is dominant in person search, but it requires elaborate labeling of bounding boxes and identities. Large-scale labeled training data is often difficult to collect, especially for person identities. A natural question is whether a good person search model can be trained without the need of identity supervision. In this paper, we present a weakly supervised setting where only bounding box annotations are available. Based on this new setting, we provide an effective baseline model termed Region Siamese Networks (R-SiamNets). Towards learning useful representations for recognition in the absence of identity labels, we supervise the R-SiamNet with instance-level consistency loss and cluster-level contrastive loss. For instance-level consistency learning, the R-SiamNet is constrained to extract consistent features from each person region with or without out-of-region context. For cluster-level contrastive learning, we enforce the aggregation of closest instances and the separation of dissimilar ones in feature space. Extensive experiments validate the utility of our weakly supervised method. Our model achieves the rank-1 of 87.1% and mAP of 86.0% on CUHK-SYSU benchmark, which surpasses several fully supervised methods, such as OIM [36] and MGTS [4], by a clear margin. More promising performance can be reached by incorporating extra training data. We hope this work could encourage the future research in this field. Chuchu Han, Dongdong Yu, Zehuan Yuan, Changxin Gao, Nong Sang, Yi Yang 0001, Changhu Wang |
ICCV | 6 |
| 2021 | OadTR: Online Action Detection with TransformersabstractMost recent approaches for online action detection tend to apply Recurrent Neural Network (RNN) to capture long-range temporal structure. However, RNN suffers from non-parallelism and gradient vanishing, hence it is hard to be optimized. In this paper, we propose a new encoder-decoder framework based on Transformers, named OadTR, to tackle these problems. The encoder attached with a task token aims to capture the relationships and global inter-actions between historical observations. The decoder extracts auxiliary information by aggregating anticipated future clip representations. Therefore, OadTR can recognize current actions by encoding historical information and predicting future context simultaneously. We extensively evaluate the proposed OadTR on three challenging datasets: HDD, TVSeries, and THUMOS14. The experimental results show that OadTR achieves higher training and inference speeds than current RNN based approaches, and significantly outperforms the state-of-the-art methods in terms of both mAP and mcAP. Code is available at https://github.com/wangxiang1230/OadTR. Xiang Wang 0012, Shiwei Zhang 0001, Zhiwu Qing, Yuanjie Shao, Zhengrong Zuo, Changxin Gao, Nong Sang |
ICCV | 7 |
| 2021 | Weakly Supervised Text-based Person Re-IdentificationabstractThe conventional text-based person re-identification methods heavily rely on identity annotations. However, this labeling process is costly and time-consuming. In this paper, we consider a more practical setting called weakly supervised text-based person re-identification, where only the text-image pairs are available without the requirement of annotating identities during the training phase. To this end, we propose a Cross-Modal Mutual Training (CMMT) framework. Specifically, to alleviate the intra-class variations, a clustering method is utilized to generate pseudo labels for both visual and textual instances. To further re-fine the clustering results, CMMT provides a Mutual Pseudo Label Refinement module, which leverages the clustering results in one modality to refine that in the other modality constrained by the text-image pairwise relationship. Mean-while, CMMT introduces a Text-IoU Guided Cross-Modal Projection Matching loss to resolve the cross-modal matching ambiguity problem. A Text-IoU Guided Hard Sample Mining method is also proposed for learning discriminative textual-visual joint embeddings. We conduct extensive experiments to demonstrate the effectiveness of the proposed CMMT, and the results show that CMMT performs favorably against existing text-based person re-identification methods. Our code will be available at https://github.com/X-BrainLab/WS_Text-ReID. Shizhen Zhao, Changxin Gao, Yuanjie Shao, Wei-Shi Zheng 0001, Nong Sang |
ICCV | 5 |
| 2021 | Multi-scale Deformable Deblurring Kernel Prediction for Dynamic Scene Deblurring
Nong Sang |
ICIG (3) | 2 |
| 2021 | CRANet: Cascade Residual Attention Network for Crowd CountingabstractThe existing approaches for crowd counting usually estimate a density map with deep convolutional neural network to obtain the crowd counts. Influenced by the background noises, some approaches may result in incorrect pedestrian heads recognition. Therefore, some approaches try to estimate an attention map to mask background noises. However, since the background noises are complex and stochastic, single attention is of incompetence to recognize them. Consequently, we proposed softer, and more reasonable Cascade Residual Attention Network (CRANet), which cascades several effective residual attention modules to mask background noises. Also, due to pixel-level isolation of Euclidean loss, we designed a novel Pyramid Structural Similarity Loss to train our CRANet. The proposed approach was evaluated on three crowd datasets. Experimental results demonstrated that our approach achieves the state-of-the-art. Zhongyuan Wu, Jun Sang, Nong Sang |
ICME | 5 |
| 2021 | Scale-Aware Multi-stage Fusion Network for Crowd Counting
Jun Sang, Fusen Wang, Xiaofeng Xia, Nong Sang |
ICONIP (6) | 6 |
| 2021 | Attribute-specific Control Units in StyleGAN for Fine-grained Image ManipulationabstractImage manipulation with StyleGAN has been an increasing concern in recent years. Recent works have achieved tremendous success in analyzing several semantic latent spaces to edit the attributes of the generated images. However, due to the limited semantic and spatial manipulation precision in these latent spaces, the existing endeavors are defeated in fine-grained StyleGAN image manipulation, i.e., local attribute translation. To address this issue, we discover attribute-specific control units, which consist of multiple channels of feature maps and modulation styles. Specifically, we collaboratively manipulate the modulation style channels and feature maps in control units rather than individual ones to obtain the semantic and spatial disentangled controls. Furthermore, we propose a simple yet effective method to detect the attribute-specific control units. We move the modulation style along a specific sparse direction vector and replace the filter-wise styles used to compute the feature maps to manipulate these control units. We evaluate our proposed method in various face attribute manipulation tasks. Extensive qualitative and quantitative results demonstrate that our proposed method performs favorably against the state-of-the-art methods. The manipulation results of real images further show the effectiveness of our method. Rui Wang 0099, Gang Yu 0002, Li Sun 0012, Changqian Yu, Changxin Gao, Nong Sang |
ACM Multimedia | 7 |
| 2021 | BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation
Changqian Yu, Changxin Gao, Jingbo Wang 0003, Gang Yu 0002, Chunhua Shen, Nong Sang |
Int. J. Comput. Vis. | 6 |
| 2021 | CSENet: Cascade semantic erasing network for weakly-supervised semantic segmentation
Changqian Yu, Changxin Gao, Nong Sang |
Neurocomputing | 5 |
| 2021 | Viewpoint Transform Matching model for person re-identification
Ruochen Zheng, Changxin Gao, Nong Sang |
Neurocomputing | 3 |
| 2021 | CondNet: Conditional Classifier for Scene SegmentationabstractThe fully convolutional network (FCN) has achieved tremendous success in dense visual recognition tasks, such as scene segmentation. The last layer of FCN is typically a global classifier (1×1 convolution) to recognize each pixel to a semantic label. We empirically show that this global classifier, ignoring the intra-class distinction, may lead to sub-optimal results. In this work, we present a conditional classifier to replace the traditional global classifier, where the kernels of the classifier are generated dynamically conditioned on the input. The main advantages of the new classifier consist of: (i) it attends on the intra-class distinction, leading to stronger dense recognition capability; (ii) the conditional classifier is simple and flexible to be integrated into almost arbitrary FCN architectures to improve the prediction. Extensive experiments demonstrate that the proposed classifier performs favourably against the traditional classifier on the FCN architecture. The framework equipped with the conditional classifier (called CondNet) achieves new state-of-the-art performances on two datasets. The code and models are available at https://git.io/CondNet. Changqian Yu, Yuanjie Shao, Changxin Gao, Nong Sang |
IEEE Signal Process. Lett. | 4 |
| 2021 | Latent Distribution-Based 3D Hand Pose Estimation From Monocular RGB ImagesabstractIn this article, we propose a novel compressed latent distribution representation for 3D hand pose estimation from monocular RGB images to alleviate the channel correspondence problem. The channel correspondence problem occurs when the 2D and depth coordinates are estimated from independent feature maps, which means the 2D and depth channel sequences may not match during the cross-dataset inference. In contrast, we propose a compressed latent distribution representation that the 2D and depth feature maps for each joint are interconnected and inter-constrained more directly, effectively alleviating the channel correspondence problem and improving cross-dataset performance. Moreover, we design an efficient encoder-decoder network that can maintain the resolution of feature maps to enable better hand feature extraction from monocular RGB images. In this work, the overall pipeline contains two branches: one is the 2D hand pose estimation branch based on a latent heatmap representation (LHR); the other is the 3D hand pose estimation branch based on our proposed latent distribution representation (LDR). In this way, the 2D estimation branch serves as guidance for the 3D branch, which simplifies the optimization of the overall network and results in a more rapid convergence during training. The results on several benchmark datasets (including STB, RHD, and the most recently released InterHand2.6M) demonstrate that our proposed method achieves state-of-the-art (SOTA) performance. Moran Li, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | GTNet: Generative Transfer Network for Zero-Shot Object DetectionabstractWe propose a Generative Transfer Network (GTNet) for zero-shot object detection (ZSD). GTNet consists of an Object Detection Module and a Knowledge Transfer Module. The Object Detection Module can learn large-scale seen domain knowledge. The Knowledge Transfer Module leverages a feature synthesizer to generate unseen class features, which are applied to train a new classification layer for the Object Detection Module. In order to synthesize features for each unseen class with both the intra-class variance and the IoU variance, we design an IoU-Aware Generative Adversarial Network (IoUGAN) as the feature synthesizer, which can be easily integrated into GTNet. Specifically, IoUGAN consists of three unit models: Class Feature Generating Unit (CFU), Foreground Feature Generating Unit (FFU), and Background Feature Generating Unit (BFU). CFU generates unseen features with the intra-class variance conditioned on the class semantic embeddings. FFU and BFU add the IoU variance to the results of CFU, yielding class-specific foreground and background features, respectively. We evaluate our method on three public datasets and the results demonstrate that our method performs favorably against the state-of-the-art ZSD approaches. Shizhen Zhao, Changxin Gao, Yuanjie Shao, Lerenhan Li, Changqian Yu, Zhong Ji, Nong Sang |
AAAI | 7 |
| 2020 | Domain Adaptation for Image DehazingabstractImage dehazing using learning-based methods has achieved state-of-the-art performance in recent years. However, most existing methods train a dehazing model on synthetic hazy images, which are less able to generalize well to real hazy images due to domain shift. To address this issue, we propose a domain adaptation paradigm, which consists of an image translation module and two image dehazing modules. Specifically, we first apply a bidirectional translation network to bridge the gap between the synthetic and real domains by translating images from one domain to another. And then, we use images before and after translation to train the proposed two image dehazing networks with a consistency constraint. In this phase, we incorporate the real hazy image into the dehazing training via exploiting the properties of the clear image (e.g., dark channel prior and image gradient smoothing) to further improve the domain adaptivity. By training image translation and dehazing network in an end-to-end manner, we can obtain better effects of both image translation and dehazing. Experimental results on both synthetic and real-world images demonstrate that our model performs favorably against the state-of-the-art dehazing algorithms. Yuanjie Shao, Lerenhan Li, Wenqi Ren, Changxin Gao, Nong Sang |
CVPR | 5 |
| 2020 | Context Prior for Scene SegmentationabstractRecent works have widely explored the contextual dependencies to achieve more accurate segmentation results. However, most approaches rarely distinguish different types of contextual dependencies, which may pollute the scene understanding. In this work, we directly supervise the feature aggregation to distinguish the intra-class and interclass context clearly. Specifically, we develop a Context Prior with the supervision of the Affinity Loss. Given an input image and corresponding ground truth, Affinity Loss constructs an ideal affinity map to supervise the learning of Context Prior. The learned Context Prior extracts the pixels belonging to the same category, while the reversed prior focuses on the pixels of different classes. Embedded into a conventional deep CNN, the proposed Context Prior Layer can selectively capture the intra-class and inter-class contextual dependencies, leading to robust feature representation. To validate the effectiveness, we design an effective Context Prior Network (CPNet). Extensive quantitative and qualitative evaluations demonstrate that the proposed model performs favorably against state-of-the-art semantic segmentation approaches. More specifically, our algorithm achieves 46.3% mIoU on ADE20K, 53.9% mIoU on PASCAL-Context, and 81.3% mIoU on Cityscapes. Code is available at https://git.io/ContextPrior. Changqian Yu, Jingbo Wang 0003, Changxin Gao, Gang Yu 0002, Chunhua Shen, Nong Sang |
CVPR | 6 |
| 2020 | Adversarial Semantic Data Augmentation for Human Pose Estimation
Yanrui Bin, Xuan Cao, Xinya Chen, Yanhao Ge, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Changxin Gao, Nong Sang |
ECCV (19) | 10 |
| 2020 | Representative Graph Neural Network
Changqian Yu, Yifan Liu 0001, Changxin Gao, Chunhua Shen, Nong Sang |
ECCV (7) | 5 |
| 2020 | Do Not Disturb Me: Person Re-identification Under the Interference of Other Pedestrians
Shizhen Zhao, Changxin Gao, Jun Zhang 0018, Hao Cheng 0012, Chuchu Han, Xinyang Jiang, Wei-Shi Zheng 0001, Nong Sang, Xing Sun 0001 |
ECCV (6) | 9 |
| 2020 | Keypoint-Based Feature Matching For Partial Person Re-IdentificationabstractAs a derivative of person re-identification (re-ID), partial re- ID aims to retrieve a partial pedestrian across holistic person images captured by non-overlapping cameras. This task is more challenging and closer to real-world applications. Since we cannot locate the part of the partial image, the misaligned region compromises the performance greatly when directly (a) compare a partial pedestrian with a holistic one. To alleviate this issue, we propose a Keypoint-Based Feature Matching (KBFM) network, which constructs a simple and effective framework for partial re-ID. Specifically, our architecture explicitly leverages the keypoints generated by pose estimation. Based on the visible keypoints, coordinates of the corresponding visible region can be computed. And the keypoint-based feature embeddings can be generated by bilinear sampling. When matching two images, we extract their features on the basis of shared visible keypoints, avoiding the misalignment and disturbance. Moreover, considering the triplet loss cannot be flexibly built in the partial re-ID pipeline, we improve the original sampling method and achieve significant performance. Extensive experimental results on two widely used benchmarks demonstrate significant performance improvements of our method over most state-of-the-art methods. Chuchu Han, Changxin Gao, Nong Sang |
ICIP | 3 |
| 2020 | End-to-End Blurry Template Matching Method Based on Siamese Networks
Nong Sang |
PRCV (3) | 2 |
| 2020 | Multi-level Temporal Pyramid Network for Action Detection
Xiang Wang 0012, Changxin Gao, Shiwei Zhang 0001, Nong Sang |
PRCV (2) | 4 |
| 2020 | Relevant region prediction for crowd counting
Xinya Chen, Yanrui Bin, Changxin Gao, Nong Sang, Hao Tang 0005 |
Neurocomputing | 4 |
| 2020 | Hard sample mining makes person re-identification more efficient and accurate
Kezhou Chen, Chuchu Han, Nong Sang, Changxin Gao |
Neurocomputing | 4 |
| 2020 | Jointly detecting and multiple people tracking by semantic and scene information
Zhixiong Pi, Huai Qin, Changxin Gao, Nong Sang |
Neurocomputing | 4 |
| 2020 | Pose-guided spatiotemporal alignment for video-based person Re-identification
Changxin Gao, Jin-Gang Yu, Nong Sang |
Inf. Sci. | 4 |
| 2020 | Structure-aware human pose estimation with graph convolutional networks
Yanrui Bin, Xiu-Shen Wei, Xinya Chen, Changxin Gao, Nong Sang |
Pattern Recognit. | 6 |
| 2020 | Joint image deblurring and matching with feature-based sparse representation prior
Juncai Peng, Yuanjie Shao, Nong Sang, Changxin Gao |
Pattern Recognit. | 3 |
| 2020 | Joint image restoration and matching method based on distance-weighted sparse representation prior
Yuanjie Shao, Nong Sang, Changxin Gao |
Pattern Recognit. Lett. | 2 |
| 2020 | Complementation-Reinforced Attention Network for Person Re-IdentificationabstractFine-grained information has been proved helpful for person re-identification, and multi-head attention mechanism offers a feasible solution for this. However, we observe severe redundancy among the multiple branches, which might make the learned representation over-emphasize certain discriminative regions and correspondingly ignore other potentially informative regions. Therefore, we tackle this issue by two aspects yielding the so-called Complementation-Reinforced Attention Network (CRAN). One is the redundancy among branches, and we propose to impose complementing constraints among multiple attention heads. The constraints are two-fold: on the one hand, it encourages each branch to attend to complementary attention regions; on the other hand, it enforces orthogonality among the learned features of different regions in the embedding space. The other is the redundancy among query positions for each attention head. So we simplify the attention block by sparsifying the query positions. Besides, in order to achieve efficient retrieval, we propose an adaptive feature fusion method for dimensional reduction. Compared with the commonly used feature ensemble, our method effectively reduces the dimensionality while keeping the discriminative ability. We demonstrate the effectiveness of our method on MSMT17, Market-1501, DukeMTMC-reID, and CUHK03 datasets. Chuchu Han, Ruochen Zheng, Changxin Gao, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Joint Analysis and Weighted Synthesis Sparsity Priors for Simultaneous Denoising and Destriping Optical Remote Sensing ImagesabstractStripe and random noise are two different degradation phenomena that commonly coexist in optical remote sensing images, and they are often modeled as inverse problems. In model-based inverse problems, analysis and synthesis sparse representations (SSRs) are used as regularization terms to obtain approximate solutions due to their respective merits, i.e., the nonzero coefficients in SSR are usually used to describe an image, while the indexes of zeros in analysis sparse representation (ASR) are used to characterize the stripe. Inspired by these merits, we propose a unified variational framework, called a joint analysis and weighted synthesis (JAWS) sparsity model, to simultaneously separate the clean image and the stripe from a single optical remote sensing image. To solve the JAWS sparsity model efficiently, an alternating minimization optimization strategy is first employed to separate it into two subproblems that are used for different tasks. One called as weighted SSR (WSSR) is the main for optical remote sensing image denoising, which can be effectively solved by employing the weighted singular value thresholding operator, while the other called as ASR is the main approach for optical remote sensing image destriping, which is optimized by adopting the split Bregman iteration. By minimizing the two subproblems alternatively, the proposed JAWS sparsity model is efficiently solved. Finally, both quantitative and qualitative results of experiments on synthetic and real-world optical remote sensing images validate that the proposed approach is effective and even better than the state of the arts. Zhenghua Huang, Yaozong Zhang, Qian Li 0019, Tianxu Zhang, Nong Sang, Hanyu Hong |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2020 | Semi-Supervised Image DehazingabstractWe present an effective semi-supervised learning algorithm for single image dehazing. The proposed algorithm applies a deep Convolutional Neural Network (CNN) containing a supervised learning branch and an unsupervised learning branch. In the supervised branch, the deep neural network is constrained by the supervised loss functions, which are mean squared, perceptual, and adversarial losses. In the unsupervised branch, we exploit the properties of clean images via sparsity of dark channel and gradient priors to constrain the network. We train the proposed network on both the synthetic data and real-world images in an end-to-end manner. Our analysis shows that the proposed semi-supervised learning algorithm is not limited to synthetic training datasets and can be generalized well to real-world images. Extensive experimental results demonstrate that the proposed algorithm performs favorably against the state-of-the-art single image dehazing algorithms on both benchmark datasets and real-world images. Lerenhan Li, Yunlong Dong, Wenqi Ren, Jinshan Pan, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Dynamic Scene Deblurring by Depth Guided ModelabstractDynamic scene blur is usually caused by object motion, depth variation as well as camera shake. Most existing methods usually solve this problem using image segmentation or fully end-to-end trainable deep convolutional neural networks by considering different object motions or camera shakes. However, these algorithms are less effective when there exist depth variations. In this work, we propose a deep neural convolutional network that exploits the depth map for dynamic scene deblurring. Given a blurred image, we first extract the depth map and adopt a depth refinement network to restore the edges and structure in the depth map. To effectively exploit the depth map, we adopt the spatial feature transform layer to extract depth features and fuse with the image features through scaling and shifting. Our image deblurring network thus learns to restore a clear image under the guidance of the depth map. With substantial experiments and analysis, we show that the depth information is crucial to the performance of the proposed model. Finally, extensive quantitative and qualitative evaluations demonstrate that the proposed model performs favorably against the state-of-the-art dynamic scene deblurring approaches as well as conventional depth-based deblurring algorithms. Lerenhan Li, Jinshan Pan, Wei-Sheng Lai, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Learning Nonclassical Receptive Field Modulation for Contour DetectionabstractThis work develops a biologically inspired neural network for contour detection in natural images by combining the nonclassical receptive field modulation mechanism with a deep learning framework. The input image is first convolved with the local feature detectors to produce the classical receptive field responses, and then a corresponding modulatory kernel is constructed for each feature map to model the nonclassical receptive field modulation behaviors. The modulatory effects can activate a larger cortical area and thus allow cortical neurons to integrate a broader range of visual information to recognize complex cases. Additionally, to characterize spatial structures at various scales, a multiresolution technique is used to represent visual field information from fine to coarse. Different scale responses are combined to estimate the contour probability. Our method achieves state-of-the-art results among all biologically inspired contour detection models. This study provides a method for improving visual modeling of contour detection and inspires new ideas for integrating more brain cognitive mechanisms into deep neural networks. Qiling Tang, Nong Sang, Haihua Liu |
IEEE Trans. Image Process. | 2 |
| 2020 | GLNet: Global Local Network for Weakly Supervised Action LocalizationabstractIn this paper, we address the challenging problem of weakly supervised spatio-temporal action localization for which only video-level action labels are available during training. To solve this problem, we propose an end-to-end Global Local Network (GLNet) to predict the probability distribution simultaneously in both spatial and temporal space. The proposed GLNet model includes two key components: a local spatial module and a global temporal module. The local spatial module aims to predict the frame-level spatial distribution by encoding short-term temporal information. In particular, we propose a Region Actionness Network (RAN) to select the target region boxes from the precomputed exhaustive proposals. The global temporal module can predict temporal distribution by a long-term temporal structure modelling. Specifically, we design a temporal fusion-and-excitation architecture on the top of several clips, and trained by a sparse loss function. Therefore, the proposed GLNet model can perform spatio-temporal action localization in an end-to-end manner. We evaluate the performance of GLNet on the J-HMDB and UCF101-24 datasets. The experimental results demonstrate GLNet achieves a significant margin against other state-of-the-art weakly supervised methods and even some fully supervised methods in terms of frame mean Average Precision (mAP) and the video mAP (called frame-mAP and video-mAP, respectively). Shiwei Zhang 0001, Lin Song 0002, Changxin Gao, Nong Sang |
IEEE Trans. Multim. | 4 |
| 2019 | Camera Style and Identity Disentangling Network for Person Re-identification
Ruochen Zheng, Lerenhan Li, Chuchu Han, Changxin Gao, Nong Sang |
BMVC | 5 |
| 2019 | Re-ID Driven Localization Refinement for Person SearchabstractPerson search aims at localizing and identifying a query person from a gallery of uncropped scene images. Different from person re-identification (re-ID), its performance also depends on the localization accuracy of a pedestrian detector. The state-of-the-art methods train the detector individually, and the detected bounding boxes may be sub-optimal for the following re-ID task. To alleviate this issue, we propose a re-ID driven localization refinement framework for providing the refined detection boxes for person search. Specifically, we develop a differentiable ROI transform layer to effectively transform the bounding boxes from the original images. Thus, the box coordinates can be supervised by the re-ID training other than the original detection task. With this supervision, the detector can generate more reliable bounding boxes, and the downstream re-ID model can produce more discriminative embeddings based on the refined person localizations. Extensive experimental results on the widely used benchmarks demonstrate that our proposed method performs favorably against the state-of-the-art person search methods. Chuchu Han, Jiacheng Ye, Yunshan Zhong, Xin Tan 0002, Chi Zhang 0026, Changxin Gao, Nong Sang |
ICCV | 7 |
| 2019 | Blurred Template Matching Based on Cascaded Network
Juncai Peng, Nong Sang, Changxin Gao, Lerenhan Li |
ICIG (1) | 2 |
| 2019 | Densenet-Based Multi-scale Recurrent Network for Video Restoration with Gaussian Blur
Liyou Wu, Nong Sang, Lihong Jing, Changxin Gao, Lerenhan Li |
ICIG (1) | 2 |
| 2019 | Joint Image Restoration and Matching Based on Hierarchical Sparse RepresentationabstractImage matching is widely used in visual-based navigation systems, and most matching methods simply assume the ideal inputs without considering the degradation of real world, such as image blur, which is very common in real-time images. Joint image restoration and matching, such as JRM-DSR is a good way to deal with the degradation of real-time images, which utilizes the sparse representation of the real-time image on the dictionary constructed from the reference image. However, once the size of the reference image is much bigger than that of the real-time image, the size of the dictionary would be so huge that it becomes time-consuming and tough to get the sparse representation. In this paper, we propose a joint image restoration and matching method based on hierarchical sparse representation (JRM-HSR), which shrinks the size of the dictionary with the help of clustering to perform the coarse matching, and then performs the fine matching in a subset of the original dictionary. JRM-HSR is a practical model benefits from the hierarchical structure. In contrast to JRM-DSR, the speed of JRM-HSR is 16 times faster in single sparse representation and 2 times faster in single complete algorithm flow while maintaining the same accuracy. Nong Sang, Changxin Gao, Yuanjie Shao |
ICIP | 2 |
| 2019 | Learning What and Where from Attributes to Improve Person Re-IdentificationabstractDue to high-level semantic cues (what) and spatial properties (where) of person attribute, some recent works try to introduce it into person re-identification. However, jointly learning attributes and identity by directly combining their loss function does not work, because of the significant difference between these two tasks. To address this problem, we propose an Attribute-identity Feature Fusion Network (AFFNet) for person re-ID, which fuses attribute and identity recognition tasks not only on loss level, but also on feature level. Specifically, to learn different features for attribute and identity, we split them into two branches to avoid the interference effects between each other. These two types of features are then concatenated to form the final representation. In the attribute branch, we propose to combine hierarchical features and use a Feature Attention Block (FAB), to mining high-level semantic and spatial information, respectively. The experimental results on two public datasets show that the proposed method performs favorably against state-of-the-art methods. Jinghao Luo, Changxin Gao, Nong Sang |
ICIP | 4 |
| 2019 | Orientation Adaptive YOLOv3 for Object Detection in Remote Sensing Images
Jiahui Lei, Chongjun Gao, Changxin Gao, Nong Sang |
PRCV (1) | 5 |
| 2019 | Graph-Based Scale-Aware Network for Human Parsing
Changqian Yu, Changxin Gao, Nong Sang |
PRCV (2) | 5 |
| 2019 | Scale Pyramid Network for Crowd CountingabstractCrowd counting is a concerned yet challenging task in computer vision. The difficulty is particularly pronounced by scale variations in crowd images. Most state-of-art approaches tackle the multi-scale problem by adopting multi-column CNN architectures where different columns are designed with different filter sizes to adapt to variable pedestrian/object sizes. However, the structure is bloated and inefficient, and it is infeasible to adopt multiple deep columns due to the huge resource cost. We instead propose a Scale Pyramid Network (SPN) which adopts a shared single deep column structure and extracts multi-scale information in high layers by Scale Pyramid Module. In Scale Pyramid Module, we specifically employ different rates of dilated convolutions in parallel instead of traditional convolutions with different sizes. Compared to other methods of coping with scale issues, our single column structure with Scale Pyramid Module can get more accurate estimation with simpler structure and less complexity of training. And our Scale Pyramid Module can be easily applied to a deep network. Experimental results on four datasets show that our method achieves state-of-the-art performance. On ShanghaiTech Part_A dataset which is challenging for its highly congested scenes and scale variation, we achieve 9.5% lower MAE and 13.5% lower MSE than the previous state-of-the-art method. We also extend our model on TRANCOS vehicle counting dataset and significantly achieve 5.9% lower GAME(0), 10% lower GAME(1), 24.5% lower GAME(2), 38.7% lower GAME(3) than the previous state-of-the-art method. The experimental results prove the robustness of our model for crowd counting, especially with scale variations. Xinya Chen, Yanrui Bin, Nong Sang, Changxin Gao |
WACV | 3 |
| 2019 | Blind Image Deblurring via Deep Discriminative Priors
Lerenhan Li, Jinshan Pan, Wei-Sheng Lai, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 5 |
| 2019 | Motion-blur kernel size estimation via learning a convolutional neural network
Lerenhan Li, Nong Sang, Luxin Yan, Changxin Gao |
Pattern Recognit. Lett. | 2 |
| 2018 | Multiple Object Tracking by Learning Feature Representation and Distance Metric Jointly
Guoshuai Zhang, Nong Sang, Rui Huang 0001, Jianhua Hou |
BMVC | 3 |
| 2018 | Learning a Discriminative Prior for Blind Image DeblurringabstractWe present an effective blind image deblurring method based on a data-driven discriminative prior. Our work is motivated by the fact that a good image prior should favor clear images over blurred ones. In this work, we formulate the image prior as a binary classifier which can be achieved by a deep convolutional neural network (CNN). The learned prior is able to distinguish whether an input image is clear or not. Embedded into the maximum a posterior (MAP) framework, it helps blind deblurring in various scenarios, including natural, face, text, and low-illumination images. However, it is difficult to optimize the deblurring method with the learned image prior as it involves a non-linear CNN. Therefore, we develop an efficient numerical approach based on the half-quadratic splitting method and gradient decent algorithm to solve the proposed model. Furthermore, the proposed model can be easily extended to non-uniform deblurring. Both qualitative and quantitative experimental results show that our method performs favorably against state-of-the-art algorithms as well as domain-specific image deblurring approaches. Lerenhan Li, Jinshan Pan, Wei-Sheng Lai, Changxin Gao, Nong Sang, Ming-Hsuan Yang 0001 |
CVPR | 5 |
| 2018 | Learning a Discriminative Feature Network for Semantic SegmentationabstractMost existing methods of semantic segmentation still suffer from two aspects of challenges: intra-class inconsistency and inter-class indistinction. To tackle these two problems, we propose a Discriminative Feature Network (DFN), which contains two sub-networks: Smooth Network and Border Network. Specifically, to handle the intra-class inconsistency problem, we specially design a Smooth Network with Channel Attention Block and global average pooling to select the more discriminative features. Furthermore, we propose a Border Network to make the bilateral features of boundary distinguishable with deep semantic boundary supervision. Based on our proposed DFN, we achieve state-of-the-art performance 86.2% mean IOU on PASCAL VOC 2012 and 80.3% mean IOU on Cityscapes dataset. Changqian Yu, Jingbo Wang 0003, Chao Peng 0001, Changxin Gao, Gang Yu 0002, Nong Sang |
CVPR | 6 |
| 2018 | BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation
Changqian Yu, Jingbo Wang 0003, Chao Peng 0001, Changxin Gao, Gang Yu 0002, Nong Sang |
ECCV (13) | 6 |
| 2018 | Multi-Scale YOLOv2 for Hand Detection in Complex ScenesabstractThis paper presents a model named Multi-Scale YOLOv2 (MS-YOLOv2) for hand detection in complex scenes. The proposed MS-YOLOv2 is implemented by introducing three modules to YOLOv2, including a Multi-Scale Feature Refinement Module to acquire fine-grained features, a Channel Importance Evaluation Module to recalibrate feature channels and a Hard Example Punishment Module to get rid of hand interference areas. Experiment results show that the proposed MS-YOLOv2 makes much performance improvement to YOLOv2, but with little computational complexity gain. On our dataset, the proposed MS-YOLOv2 can achieve 98.2% of AP and 97.9% of AR. Moreover, on the VIVA challenge, the proposed MS-YOLOv2 achieves AP/AR of 85.1%/45.8% at Level-1 and 80.1%/45.9% at Level-2. Zihan Ni, Nong Sang |
ICARCV | 3 |
| 2018 | Improving Person Re-Identification by Adaptive Hard Sample MiningabstractThe field of person reidentification has made significant advances riding on the wave of deep learning. However, owing to the fact that there are much more easy examples than those meaningful hard examples in dataset, the training tends to stagnate quickly and the model may suffer from over-fitting. Therefore, the hard sample mining method is fateful to optimize the model and improve the learning efficiency. In this paper, an Adaptive Hard Sample Mining algorithm is proposed for training a robust person re-identification model. No need for hand-picking the images in the batch or designing the loss function for both positive and negative pairs, we can briefly calculate the hard level by comparing the prediction result with the true label of the sample. Meanwhile, an adaptive threshold of hard level can make the algorithm not only stay in step with training process harmoniously but also alleviate the under-fitting and over-fitting problem simultaneously. Besides, the designed network to implement the approach has good generalization performance that can be combined with various of existing models readily. Experimental results on Market-1501 and DukeMTMC-reID datasets clearly demonstrate the effectiveness of the proposed algorithm. Kezhou Chen, Chuchu Han, Nong Sang, Changxin Gao, Ruolin Wang |
ICIP | 4 |
| 2018 | Light YOLO for High-Speed Gesture RecognitionabstractThis paper proposes an efficient model named Light YOLO for hand gesture recognition on the embedded platforms. Light YOLO improves accuracy, speed, and model size, in three aspects. To deal with the small scale gestures in practical applications, we strengthen the YOLOv2 with a spatial refinement module to obtain fine-grained features. To accelerate the refined network, we propose a selective-dropout channel pruning approach to prune the redundancy convolution kernels in the network. Moreover, we introduce a dataset for hand gesture recognition in complex scenes. The experimental results on this dataset show that the proposed Light YOLO significantly improve the YOLOv2 network, i.e., accuracy from 96.80% to 98.06%, speed form 40PFS to 125FPS, and size form 250M to 4MB. Zihan Ni, Nong Sang, Changxin Gao, Leyuan Liu 0001 |
ICIP | 3 |
| 2018 | Spatially Attentive Correlation Filters for Visual TrackingabstractAlthough correlation filter based trackers have recently demonstrated excellent performance, they still suffer from the boundary effects. The cosine window is introduced to alleviate the boundary affects, which however may result in poor performance in case of occlusion or fast motion. To address this problem, we propose a simple yet effective framework, which builds a spatially attentive model with multiple features to guide the detection of the correlation filter based trackers. The proposed method not only can breakthrough the spatial extent of cosine window, but also can provides prior information about the target object. Moreover, to model a robust object prior, we propose a generic strategy for adaptive fusion and update of multiple features. Extensive experiments over multiple tracking benchmarks demonstrate the superior accuracy and real-time performance of our methods compared to the state-of-the-art trackers. Huai Qin, Zhixiong Pi, Changqian Yu, Changxin Gao, Jin-Gang Yu, Nong Sang |
ICIP | 6 |
| 2018 | Joint Image Restoration and Matching Based on Distance-Weighted Sparse RepresentationabstractImage matching is widely used in visual-based navigation systems, most of which simply assume the ideal inputs without considering the degradation of the real world, such as image blur. In presence of such situation, the traditional matching methods first resort to image restoration and then perform image matching with the restored image. However, by treating the restoration and matching separately, the accuracy of image matching will be reduced by the defective output of the image restoration. In this paper, we propose a joint image restoration and matching method based on distance-weighted sparse representation (JRM-DSR), which utilizes the sparse representation prior to exploit the correlation between restoration and matching. This prior assumes that the blurry image, if correctly restored, can be well represented as a sparse linear combination of the dictionary constructed by the reference image. In order to achieve more accurate matching results to help restoration, we consider both local and sparse information and adopt distance-weighted sparse representation to obtain better representation coefficients. By iteratively restoring the input image in pursuit of the sparest representation, our approach can achieve restoration and matching simultaneous, and these two tasks can benefit greatly from each other. matching, we give a coarse to fine matching strategy to further improve the matching accuracy. Experiments demonstrate the effectiveness of our method compared with conventional methods. Yuanjie Shao, Nong Sang, Changxin Gao |
ICPR | 2 |
| 2018 | Re-ranking Person Re-identification with Adaptive Hard Sample Mining
Chuchu Han, Kezhou Chen, Jin Wang 0019, Changxin Gao, Nong Sang |
PRCV (1) | 5 |
| 2018 | Global Feature Learning with Human Body Region Guided for Person Re-identification
Nong Sang, Kezhou Chen, Chuchu Han, Changxin Gao |
PRCV (1) | 2 |
| 2018 | Center-Level Verification Model for Person Re-identification
Ruochen Zheng, Changqian Yu, Chuchu Han, Changxin Gao, Nong Sang |
PRCV (1) | 6 |
| 2018 | Iterative weighted sparse representation for X-ray cardiovascular angiogram image denoising over learned dictionaryabstractNon‐local self‐similar patch‐based denoising techniques have been viewed as the most popular denoising approaches in computer vision. This study has proposed a novel iterative weighted sparse representation (IWSR) scheme for X‐ray cardiovascular angiogram image denoising. The main procedures of this scheme include four parts. First, a maximum a posterior (MAP) distribution by the Bayes’ theory is adopted to simultaneously estimate the estimated image and sparse representation with different Gaussian distributions approximating to likelihood prior, non‐local self‐similar patch prior and sparse representation prior. Second, the MAP problem is converted to minimise an energy function using the logarithmic transformation. Third, the function is efficiently solved by the single and effective alternating directions method of multipliers algorithm along with singular value decomposition (SVD) algorithm. Finally, owing to learned dictionary by K‐SVD algorithm, the qualitative and quantitative results of widely synthetic experiments demonstrate that the proposed IWSR denoising method performs effectively and can obtain competitive denoising performance and high‐quality images compared with those advanced denoising methods. The results of extensive experiments on clinical X‐ray angiogram images further illustrate that the IWSR method performs well on noise reduction and vascular structures including edges and capillaries preservation, integral cardiovascular trees of which are beneficial for clinicians to diagnose and analyse cardiovascular diseases. Zhenghua Huang, Qian Li 0019, Tianxu Zhang, Nong Sang, Hanyu Hong |
IET Image Process. | 4 |
| 2018 | Framelet regularization for uneven intensity correction of color images with illumination and reflectance estimation
Zhenghua Huang, Likun Huang, Qian Li 0019, Tianxu Zhang, Nong Sang |
Neurocomputing | 5 |
| 2018 | Progressive Dual-Domain Filter for Enhancing and Denoising Optical Remote-Sensing ImagesabstractEnhancement and denoising have always been a pair of conflicting problems in image processing of computer vision. Inspired by an earlier dual-domain filter (DDF), this letter proposes a progressive DDF to simultaneously enhance and denoise low-quality optical remote-sensing images. The main procedure of the proposed enhancement filter has two parts. First, a bilateral filter is exploited as a guide filter to obtain high-contrast images, which are enhanced by a histogram modification method. Then, low-contrast useful structures are restored by a short-time Fourier transform and are enhanced using an adaptive correction parameter. Both the quantitative and qualitative results of experiments on synthetic and real-world low-quality remote-sensing images demonstrate that the proposed method performs well on contrast enhancement, structure preservation, and noise reduction. Moreover, its satisfactory computation time resulting from its simple implementation makes it suitable for extensive application. Zhenghua Huang, Yaozong Zhang, Qian Li 0019, Tianxu Zhang, Nong Sang, Hanyu Hong |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Spatial and class structure regularized sparse representation graph for semi-supervised hyperspectral image classification
Yuanjie Shao, Nong Sang, Changxin Gao, Li Ma 0005 |
Pattern Recognit. | 2 |
| 2018 | Equidistance constrained metric learning for person re-identification
Jin Wang 0019, Zheng Wang 0007, Chao Liang 0001, Changxin Gao, Nong Sang |
Pattern Recognit. | 5 |
| 2018 | Representation Space-Based Discriminative Graph Construction for Semisupervised Hyperspectral Image ClassificationabstractGraph-based semisupervised learning methods have been successfully applied in hyperspectral image (HSI) classification with limited labeled samples. The critical step of graph-based methods is to learn a similarity graph, and numerous graph construction methods have been developed in recent years. However, existing approaches usually return a similarity matrix from the raw data space. In this letter, we propose a representation space-based discriminative graph for semisupervised HSI classification, which can learn the representations of samples and the similarity matrix of representations simultaneously. Moreover, we explicitly incorporate the probabilistic class relationship between sample and class, which can be estimated by the partial label information, into the above model to further boost the discriminability of graph. The experimental results on Hyperion and AVIRIS hyperspectral data demonstrate the effectiveness of the proposed approach. Yuanjie Shao, Nong Sang, Changxin Gao |
IEEE Signal Process. Lett. | 2 |
| 2018 | Discriminative Part Selection for Human Action RecognitionabstractSemantic parts have shown a powerful discriminative capacity for action recognition. However, many existing methods select parts according to predefined heuristic rules, which may cause the correlation among parts to be lost, or do not appropriately consider the cluttered candidate part space, which may result in weak generalizability of the resulting action labels. Therefore, better consideration of the correlation among parts and refinement of the candidate space will lead to a more discriminative action representation. This paper achieves improved performance by more elegantly addressing these two factors. First, considering the cluttered nature of the candidate space, we propose a recursive part elimination strategy for iterative refinement of the candidate parts. In each iteration, we eliminate the parts with the lowest weights, which are deemed to be noise. Second, we measure the discriminative capabilities of the candidates and select the top-ranked parts by applying a maximum margin model, which can alleviate overfitting while simultaneously improving generalizability and correlation extraction. Finally, using the selected parts, we extract mid-level features. We report experiments conducted on four datasets (KTH, Olympic Sports, UCF50, and HMDB51). The proposed method can achieve significant improvements compared with other recent methods, including a lower computational cost, a faster speed, and higher accuracy. Shiwei Zhang 0001, Changxin Gao, Nong Sang |
IEEE Trans. Multim. | 5 |
| 2017 | Graph coloring based surveillance video synopsis
Yi He 0004, Changxin Gao, Nong Sang, Zhiguo Qu |
Neurocomputing | 3 |
| 2017 | A discriminant sparse representation graph-based semi-supervised learning for hyperspectral image classification
Yuanjie Shao, Changxin Gao, Nong Sang |
Multim. Tools Appl. | 3 |
| 2017 | Regularized max-min linear discriminant analysis
Guowan Shao, Nong Sang |
Pattern Recognit. | 2 |
| 2017 | Probabilistic class structure regularized sparse representation graph for semi-supervised hyperspectral image classification
Yuanjie Shao, Nong Sang, Changxin Gao, Li Ma 0005 |
Pattern Recognit. | 2 |
| 2017 | A low-cost real-time face tracking system for ITSs and SDASsabstractSummary It is important to track people's face efficiently and accurately in many Intelligent Transportation Systems (ITSs) and Safety Driving Assistant Systems (SDASs). This paper presents a high‐performance and low‐cost real‐time face tracking system, which runs on general onboard computer with very low CPU consumption. The proposed face tracking system is composed of four modules: the motion detector, face detector, face tracker, and face validator. The motion detector extracts motion areas by using a spatial‐temporal bi‐differential method with a very low computational cost. The face detector integrates motions into a cascade face detection framework to reject most of non‐face scanning‐windows to ensure efficient face localization. The face tracker fuses motion feature with color feature to alleviate the drifting problem during tracking. The face validator builds face appearance models online and identifies each specific tracked face to avoid confusion. Experimental results on three challenging video sequences show that the proposed face tracking system outperforms the state‐of‐the‐art face tracker and consumes only 5–13% CPU resources of a low‐spec onboard computer while processing in real time. Copyright © 2016 John Wiley & Sons, Ltd. Leyuan Liu 0001, Jingying Chen 0001, Changxin Gao, Nong Sang |
Softw. Pract. Exp. | 4 |
| 2017 | Fast Online Video Synopsis Based on Potential Collision GraphabstractVideo synopsis is a smart solution to fast browsing and retrieval of raw surveillance data, in which tube rearrangement plays a key role. However, conventional methods for tube rearrangement are based on minimizing a global energy function, which is computational intensive and time consuming. In this letter, we propose a novel tube rearrangement strategy for online video synopsis by analyzing collision relationship between tubes. A potential collision graph (PCG) is constructed to represent the tubes and their potential collision relationship. Based on the PCG, tube rearrangement is achieved by filling the tubes into synopsis video in a deterministic way, which decreases computational complexity. Finally, we incorporate the proposed tube rearrangement into an online framework to generate video synopsis and validate its efficiency with extensive experiments. Yi He 0004, Zhiguo Qu, Changxin Gao, Nong Sang |
IEEE Signal Process. Lett. | 4 |
| 2017 | Robust Visual Tracking Using Exemplar-Based DetectorsabstractTracking by detection has become an attractive tracking technique, which treats tracking as an object detection problem and trains a detector to separate the target object from the background in each frame. While this strategy is effective to some extent, we argue that the task in tracking should be searching for a specific object instance instead of an object category. Based on this viewpoint, a novel framework based on object exemplar detectors is proposed for visual tracking. To build a specific and discriminative model to separate the object instance from the background, the proposed method trains an exemplar-based linear discriminant analysis (ELDA) classifier for the object exemplar, using the current tracked instance as the positive sample and massive negative samples obtained both offline and online. To improve the trackers' adaptivity, we use an ensemble of the above ELDA detectors and update them during the tracking to cover the variation in object appearance. Extensive experimental results on a large benchmark data set show that the proposed method outperforms many state-of-the-art trackers, demonstrating the effectiveness and robustness of the ELDA tracker. Changxin Gao, Jin-Gang Yu, Rui Huang 0001, Nong Sang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | DeepList: Learning Deep Features With Adaptive Listwise Constraint for Person ReidentificationabstractPerson reidentification (re-id) aims to match a specific person across nonoverlapping cameras, which is an important but challenging task in video surveillance. Conventional methods mainly focus either on feature constructing or metric learning. Recently, some deep learning-based methods have been proposed to learn image features and similarity measures jointly. However, current deep models for person re-id are usually trained with eitherpairwise loss, where the number of negative pairs greatly outnumbering that of positive pairs may lead the training model to be biased toward negative pairs orconstant margin hinge loss, without considering the fact that hard negative samples should be paid more attention in the training stage. In this paper, we propose to learn deep representations with an adaptive margin listwise loss. First, ranking lists instead of image pairs are used as training samples, in this way, the problem of data imbalance is relaxed. Second, by introducing an adaptive margin parameter in the listwise loss function, it can assign larger margins to harder negative samples, which can be interpreted as an implementation of the automatic hard negative mining strategy. To gain robustness against changes in poses and part occlusions, our architecture combines four convolutional neural networks, each of which embeds images from different scales or different body parts. The final combined model performs much better than each single model. The experimental results show that our approach achieves very promising results on the challenging CUHK03, CUHK01, and VIPeR data sets. Jin Wang 0019, Zheng Wang 0007, Changxin Gao, Nong Sang, Rui Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Group Sparse-Based Mid-Level Representation for Action RecognitionabstractMid-level parts are shown to be effective for human action recognition in videos. Typically, these semantic parts are first mined with some heuristic rules, then videos are represented via volumetric max-pooling (VMP) method. However, these methods have two issues: 1) the VMP strategy divides videos by static grids. In this case, a semantic part may occur in different localizations in different videos. That means the VMP strategy loses the space-time invariance. To solve this problem, we propose to apply a saliency-driven max-pooling scheme to represent a video. We extract the video semantic cues by the saliency map, and dynamically pool the local maximum responses. This scheme can be considered as a semantic content-based feature alignment method and 2) the parts discovered by heuristic rules may be intuitive but not discriminative enough for action classification because they neglect the relations between the detectors. For this issue, we propose to apply a sparse classifier model to select discriminative parts. Moreover, to further improve the discriminative ability of the representation, we propose to conduct feature selection by the corresponding entry magnitude of the model coefficients. We conduct experiments on four challenging datasets-KTH, Olympic Sports, UCF50, and HMDB51. The results show that the proposed method significantly outperforms the state-of-the-art methods. Shiwei Zhang 0001, Changxin Gao, Sihui Luo 0001, Nong Sang |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2016 | Local Fractional Order Derivative Vector Quantization Pattern for Face Recognition
Jing Li 0007, Nong Sang, Changxin Gao |
ACCV (3) | 2 |
| 2016 | Data Association Based Multi-target Tracking Using a Joint Formulation
Jianhua Hou, Changxin Gao, Nong Sang |
ACCV (4) | 4 |
| 2016 | Temporally aligned pooling representation for video-based person re-identificationabstractThis paper proposes an effective Temporally Aligned Pooling Representation (TAPR) for video-based person re-identification. To extract the motion information from a sequence, we propose to track the superpixels of the lowest portions of human. To perform temporal alignment of videos, we propose to select the “best” walking cycle from the noisy motion information according to the intrinsic periodicity property of walking persons, that is fitted sinusoid in our implementation. To describe the video data in the selected walking cycle, we first divide the cycle into several segments according to the sinusoid, and then describe each segment by temporally aligned pooling. Extensive experimental results on the public datasets demonstrate the effectiveness of the proposed method compared with the state-of-the-art approaches. Changxin Gao, Jin Wang 0019, Leyuan Liu 0001, Jin-Gang Yu, Nong Sang |
ICIP | 5 |
| 2016 | A hyperspectral image restoration method based on analysis sparse filterabstractCosparse analysis model has shown its superior performance in image reconstruction. However, this analysis frame has not been exploited yet for hyperspectral image restoration task. An analysis operator learning method called GOAL (GeOmetric Analysis operator Learning) is applied for hyperspectral image. Considering the correlation of the hyperspectral bands, the hyperspectral images were cropped into cube cells to get training samples. To avoid the window effect by image patch strategy, the analysis sparse filter method which provides global support from local information of the image was adopted. The denoising experiments and inpainting experiments are implemented on the hyperspectral images. The results are compared with the state of the art method, which shows our method is robust and efficient. Chang Han, Nong Sang, Changxin Gao |
ICPR | 2 |
| 2016 | Contextual Similarity Regularized Metric Learning for person re-identificationabstractPerson re-identification, aiming to match a specific person among non-overlapping cameras, has attracted plenty of attention in recent years. It can be regarded as a visual retrieval task, namely given a query person image, ranking all gallery images according to their similarities to the query. Conventionally, this similarity function is learnt by forcing intra-distances to be small while inter-distances to be large, which are referred to as individual similarity constraints. In this paper, we propose to learn the similarity function by taking into account of both individual similarity constraints and contextual similarity constraints. The context of a query is defined as its k-nearest neighbors in the gallery. We argue that if two images are from the same person, apart from the visual likeness between them, denoted as the individual similarity, they should also possess similar k-nearest neighbors in the gallery, denoted as the contextual similarity. Motivated by this assumption, we propose a new Contextual Similarity Regularized Metric Learning (CSRML) method for person re-identification. The contextual similarity regularization term forces two images of the same person to share similar context. Both individual and contextual similarity constraints are encoded by a large margin logistic loss function and the final problem is solved by the stochastic gradient descent algorithm. Experiments on the challenging VIPeR and CUHK01 datasets show that our approach achieves very competitive performance. Jin Wang 0019, Junkang Zhu, Zheng Wang 0007, Changxin Gao, Nong Sang, Rui Huang 0001 |
ICPR | 5 |
| 2016 | Towards designing risk-based safe Laplacian Regularized Least Squares
Haitao Gan, Zhizeng Luo, Xugang Xi, Nong Sang, Rui Huang 0001 |
Expert Syst. Appl. | 5 |
| 2016 | Object localization with mid-level part detectors
Xiaoqin Kuang, Nong Sang, Changxin Gao |
Neurocomputing | 2 |
| 2016 | Completed local similarity pattern for color image recognition
Jing Li 0007, Nong Sang, Changxin Gao |
Neurocomputing | 2 |
| 2016 | Spatial multi-scale gradient orientation consistency for place instance and Scene category recognition
Changxin Gao, Nong Sang, Rui Huang 0001 |
Inf. Sci. | 2 |
| 2016 | Contrast-dependent surround suppression models for contour detection
Qiling Tang, Nong Sang, Haihua Liu |
Pattern Recognit. | 2 |
| 2016 | LEDTD: Local edge direction and texture descriptor for face recognition
Jing Li 0007, Nong Sang, Changxin Gao |
Signal Process. Image Commun. | 2 |
| 2016 | Similarity Learning with Top-heavy Ranking Loss for Person Re-identificationabstractPerson re-identification is the task of finding a person of interest across a network of cameras. In this paper, we propose a new similarity learning method for person re-identification. Conventional metric learning methods generally learn a linear transformation by employing sparse pairwise or triplet constraints. Since a lot of negative matching pairs or triplets are abandoned, the discriminative information is not fully exploited. Similarity learning methods with AUC loss can utilize all valid triplet constraints. However, the AUC loss has its own limitation by treating all false ranks occured at different positions equally. To address this limitation, we propose to extend the AUC loss to the top-heavy ranking loss by assigning large weights to top positions of the ranking list. Moreover, we introduce an explicit nonlinear transformation function for the original feature space and learn an inner product similarity under the structured output learning framework. Our approach achieves very promising results on the challenging VIPeR, CUHK Campus and PRID 450S datasets. Jin Wang 0019, Nong Sang, Zheng Wang 0007, Changxin Gao |
IEEE Signal Process. Lett. | 2 |
| 2016 | Hough Forest-based Association Framework with Occlusion Handling for Multi-Target TrackingabstractThis letter presents a novel multi-target tracking approach consisting of two parts. The first part is the detection based association to form global tracks. Short yet reliable tracklets are firstly generated. By effectively combining appearance and motion information, a Hough forest learning framework is constructed to obtain a more discriminative affinity model and produce longer association between tracklets. In the second part, in order to connect isolate detections for trajectory consistency, we present an appearance similarity model based on mutual occlusion reasoning. A novel fusion feature template is designed to accurately compute the matching score between each isolated detection and target. Experimental results show significant improvements of our method when compared with several state-of-the-art methods. Nong Sang, Jianhua Hou, Rui Huang 0001, Changxin Gao |
IEEE Signal Process. Lett. | 2 |
| 2016 | Multitarget Tracking Using Hough Forest Random FieldabstractThis paper presents a novel tracking-by-detection approach for multitarget tracking. There are two major steps in our framework: data association to form global tracklet association, followed by trajectory estimation to deal with the remaining gaps. In the first step, we formulate tracklet association as an inference problem in a Hough forest random field, which combines Hough forest and conditional random field and allows us to model both local and global tracklet relationships in one unified model. In the second step, we improve the reversible-jump Markov chain Monte Carlo particle filtering method with explicit mutual-occlusion reasoning to fill in the remaining gaps from the first step and increase the overall tracking precision. Extensive experiments have been conducted on five public data sets, and the performance is comparable to that of the state-of-the-art method, if not better. Nong Sang, Jianhua Hou, Rui Huang 0001, Changxin Gao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Accurate and robust facial expressions recognition by fusing multiple sparse representation based classifiers
Yan Ouyang, Nong Sang, Rui Huang 0001 |
Neurocomputing | 2 |
| 2015 | Text detection approach based on confidence map and context information
Nong Sang, Changxin Gao |
Neurocomputing | 2 |
| 2015 | Integral region-based covariance tracking with occlusion detection
Ruhan He, Nong Sang, Geli Bai, Jizi Li |
Multim. Tools Appl. | 3 |
| 2015 | Scene Text Identification by Leveraging Mid-level Patches and Context InformationabstractIn order to identify scene texts from the background interferences, many existing methods are developed just rely on exploring heuristic rules or designing sophisticated text-specific features. In this letter, we solve this problem from a different perspective, not only the low-level texture feature is adopted to describe the candidate regions, but also the mid-level patches and the context information are integrated for boosting the classification robustness. Experimental results on the three benchmark datasets show that the proposed approach has obtained the promising classification performance. Nong Sang, Changxin Gao |
IEEE Signal Process. Lett. | 2 |
| 2014 | Discovering distinctive action parts for action recognitionabstractRecent methods based on mid-level visual concepts have shown promising capability in human action recognition field. Automatically discovering semantic entities such as parts for an action class remains challenging. In this paper, we focus on discovering distinctive action parts for recognition of human actions by learning and selecting a small number of discriminative part detectors directly from training videos. We initially train a large collection of candidate Exemplar-LDA detectors from clusters obtained by clustering spatiotemporal patches in whitened space. A novel Coverage-Entropy curve is proposed as a means of measuring the representative and discriminative capabilities of part detectors, and used to select a set of compact and meaningful detectors out of the vast candidates. By integrating these mined detectors into “bag of parts” representation, our approach demonstrates state-of-the-art performance on the UCF50 dataset. Nong Sang, Changxin Gao, Xiaoqin Kuang |
ICIP | 2 |
| 2014 | Exemplar-based linear discriminant analysis for robust object trackingabstractTracking-by-detection has become an attractive tracking technique, which treats tracking as a category detection problem. However, the task in tracking is to search for a specific object, rather than an object category as in detection. In this paper, we propose a novel tracking framework based on exemplar detector rather than category detector. The proposed tracker is an ensemble of exemplar-based linear discriminant analysis (ELDA) detectors. Each detector is quite specific and discriminative, because it is trained by a single object instance and massive negatives. To improve its adaptivity, we update both object and background models. Experimental results on several challenging video sequences demonstrate the effectiveness and robustness of our tracking algorithm. Changxin Gao, Jin-Gang Yu, Rui Huang 0001, Nong Sang |
ICIP | 5 |
| 2014 | An online learned hough forest model for multi-target trackingabstractWe present an online learned framework for multiple target tracking in a crowded scene. The tracking problem is formulated as a detection-based progressive association task. Firstly, reliable tracklets are generated by low level constraints among detection responses. Then longer tracklets associations are generated based on online learned Hough forest framework which effectively combines motion and appearance information for discrimination between two tracklets. In online learning scene, the association is formulated as a MAP problem and training examples are collected based on spatial-temporal constraints. In order to alleviate the drifting problem of online learning, Hungarian algorithm is employed to modify associated errors and update the training set. The experimental results show the effectiveness of our approach. Nong Sang, Jianhua Hou |
ICIP | 2 |
| 2014 | Real-Time Tracking Combined with Object SegmentationabstractWe propose a new approach that integrates object tracking with object segmentation in a closed loop. The EM-like algorithm for color-histogram-based object tracking is modified to deal with the appearance models of the object and background represented by the Gaussian mixture models which are more efficient in RGB color space. It provides a rough object spatial model to guide segmentation. A five-layer region based graph cuts algorithm is developed to extract the accurate object region based on the object spatial model. It is effective even in cluttered background and runs more than 10 times as fast as Grab Cut. Then we can establish the appearance models of the object and background avoiding introducing errors and update them frame by frame without the problem of drift. The refined and adaptive models lead to robust tracking in return. Moreover, the motion of the object is estimated to produce a predicted object location in the new frame for tracking. A real-time robust tracking system is built based on the proposed approach and validated on a variety of challenging sequences. Nong Sang |
ICPR | 2 |
| 2014 | Max-min distance analysis by making a uniform distribution of class centers for dimensionality reduction
Guowan Shao, Nong Sang |
Neurocomputing | 2 |
| 2014 | Image segmentation using spectral clustering of Gaussian mixture models
Shan Zeng, Rui Huang 0001, Zhen Kang, Nong Sang |
Neurocomputing | 4 |
| 2013 | Medical Ultrasonography Denoising Using Sparse Coding ShrinkageabstractA locally adaptive shrinkage Bayesian estimate for medical ultrasonography denoising is proposed by exploiting the correlation among image sparse coding. The Laplacian distribution is used to model the coding coefficients. The paper deduces the MAP estimate formula and adaptive threshold. Simulation experiments are carried out to show the effectiveness of the new method. Results demonstrate that compared with classical denoising algorithms, the new method has increased peak signal-to-noise ratio (PSNR) and improved the quality of subjective visual effect. Our algorithm is also proved to be effective to the medical ultrasonography. Nong Sang, Haitao Gan, Zhiping Dan, Yanfei Chen, Hexing Ren |
ICIG | 2 |
| 2013 | A Transductive Transfer Learning Method for Ship Target RecognitionabstractShip target recognition in infrared image remains a difficult problem, due to the projection or silhouette of a three-dimensional ship target being variable in shape, orientation and scale to make its recognizability unstable. In this paper, a transductive transfer learning framework is proposed to solve the problem. Hu moments is firstly extracted as feature vectors of target. Then the transductive transfer learning method is used to find the common parameters between the feature spaces of the training ship samples and the detected ship targets, and transfer the similar knowledge from those data with different distributions. According to the experiment result in simulation infrared images, it shows that the ship targets can be recognized highly and reliably by our proposed framework. It demonstrates the robustness and effectiveness of our method for infrared images. Zhiping Dan, Nong Sang, Ruolin Wang, Yanfei Chen |
ICIG | 2 |
| 2013 | An Improved Self-Training for Face RecognitionabstractFace recognition has attracted considerable concerns in recent years. In practical applications, there are generally a small amount of labeled face images and a lot of unlabeled ones can be available. In this paper, we introduce a semi-supervised face recognition method where semi-supervised LDA (SDA) and Affinity Propagation (AP) are integrated into Self-training. SDA is employed to update the face subspace using labeled and unlabeled face images. And we employ AP to computer the templates which exist in the original face images. A series of experiments on three face datasets are carried out to evaluate the performance of our algorithm. Experimental results illustrate that our algorithm outperforms the other unsupervised, semi-supervised and supervised methods. Haitao Gan, Nong Sang, Zhiping Dan, Hexing Ren |
ICIG | 2 |
| 2013 | Spatial-Temporal Sparse Representation for Background ModelingabstractIn this paper, a sparse representation based background model is introduced for video surveillance. Inspired by the fact that spatial and temporal information are both important for foreground detection, a spatial-temporal image patch, namely brick, is used as atomic unit for online subspace learning and sparse representation. Furthermore, Random Projection emerged from Compressive Sensing theory is applied to reduce the dimension of bricks so as to speed up the algorithm. Experimental results show the effectiveness of the proposed method. Liangwei Jiang, Nong Sang |
ICIG | 3 |
| 2013 | Novel License Plate Detection Method for Complex ScenesabstractIn this paper, a novel license plate detection method is proposed. There are three key steps in our method, i.e. image preprocessing, license plate detection and license plate confirmation. First, the noises are removed and the diversities of license plate forms are unified through image preprocessing. And then, the license plates are detected roughly by using the cascade AdaBoost classifier. Finally, the gradient images are binarized, and the connected component analysis will be adopted to remove some false plates. Meanwhile, the offline trained Support Vector Machine (SVM) classifier is adopted to confirm the license plate candidates in further. The promising results of the proposed method is verified by experiments on a challenging database. Nong Sang, Ruolin Wang, Xiaoqin Kuang |
ICIG | 2 |
| 2013 | Learning to detect contours in natural images via biologically motivated schemesabstractA model for detecting contours in natural images is presented by combining the visual perceptual mechanisms and machine learning. The surround stimuli will enhance the response of the central stimulus if they can form a precise spatial configuration. On the other hand, surround inhibition will reduce the responses to homogeneous elements. Facilitation and inhibition activities in the primary visual cortex (V1) are used to enhance the well-organized structures and to reduce the non-meaningful distractors engendering from texture fields, respectively. We approach the task of facilitatory and inhibitory cue integration as a supervised learning problem using the logistic regression model. Our experiments demonstrate that the model can dramatically reduce texture edges and spurious contours, and meanwhile can to some extent avoid ground-truth contours missed by the detector. Qiling Tang, Nong Sang, Haihua Liu |
ICIP | 2 |
| 2013 | Semi-supervised Kernel Minimum Squared Error Based on Manifold Structure
Haitao Gan, Nong Sang |
ISNN (1) | 2 |
| 2013 | A Facial Expression Recognition Method by Fusing Multiple Sparse Representation Based Classifiers
Yan Ouyang, Nong Sang |
ISNN (1) | 2 |
| 2013 | Using clustering analysis to improve semi-supervised classification
Haitao Gan, Nong Sang, Rui Huang 0001, Xiaojun Tong, Zhiping Dan |
Neurocomputing | 2 |
| 2013 | A study on semi-supervised FCM algorithm
Shan Zeng, Xiaojun Tong, Nong Sang, Rui Huang 0001 |
Knowl. Inf. Syst. | 3 |
| 2013 | Multi-ring local binary patterns for rotation invariant texture classification
Yonggang He, Nong Sang |
Neural Comput. Appl. | 2 |
| 2013 | Multi-structure local binary patterns for texture classification
Yonggang He, Nong Sang, Changxin Gao |
Pattern Anal. Appl. | 2 |
| 2012 | Online Transfer Boosting for object tracking
Changxin Gao, Nong Sang, Rui Huang 0001 |
ICPR | 2 |
| 2012 | Fast image super resolution via local regression
Shuhang Gu, Nong Sang, Fan Ma |
ICPR | 2 |
| 2012 | Fractional-step max-min distance analysis for dimension reduction
Guowan Shao, Nong Sang |
ICPR | 2 |
| 2011 | Metrics for Objective Evaluation of Background Subtraction AlgorithmsabstractAlthough a large number of background subtraction (BS) algorithms have been proposed, relevant objective metrics for evaluating these algorithms are still lacking. In this paper, empirical discrepancy metrics, which quantify the spatial accuracy and temporal stability of estimated masks by taking into account the potential inaccuracy of reference masks, the location of the pixel errors relative to the border of reference masks as well as the type of errors, are presented for evaluating the performance of BS algorithms. To validate the proposed metrics, they are applied to tune the optimal parameters of LBP-based background subtraction algorithm, and the experimental results confirm the efficiency of them. Leyuan Liu 0001, Nong Sang |
ICIG | 2 |
| 2011 | Local Binary Pattern histogram based Texton learning for texture classificationabstractLocal Binary Pattern (LBP) and Texton are both widely used texture analysis techniques. In this paper we propose a patch-based texture classification method that takes advantage of both LBP and Texton. Unlike the traditional LBP methods that describe a texture with the occurrence of local binary patterns in the entire image, we compute the LBP histogram in a small region around each pixel to capture the local structure information. The texton learning method is then per- formed on these LBP histograms, resulting in a texture classification algorithm that outperforms the traditional LBP-based methods due to its preservation of local structure information. It also outperforms the traditional filtering-based texton methods due to its robustness to orientation and illumination. Experimental results on two benchmark databases validate the advantages of the proposed method. Yonggang He, Nong Sang, Rui Huang 0001 |
ICIP | 2 |
| 2011 | A Belief Propagation algorithm for bias field estimation and image segmentationabstractIntensity-based image segmentation is often plagued by the spatial intensity inhomogeneities (or non-uniformities) that are caused by the imperfection of the imaging devices and the varying operating conditions, also known as the bias field. We present a graphical model representation of the joint segmentation and bias field estimation problem and propose an iterative solver based on the Belief Propagation (BP) algorithm. The intractable joint inference problem of the original graphical model is decoupled into two MRF-MAP estimation problems and solved by a discrete-valued BP and a Gaussian BP, respectively and iteratively. We validate our method using both simulated and real data and show its connection to some of the classical filtering-based approaches. Rui Huang 0001, Nong Sang, Vladimir Pavlovic 0001, Dimitris N. Metaxas |
ICIP | 2 |
| 2011 | Image segmentation via coherent clustering in L*a*b* color space
Rui Huang 0001, Nong Sang, Dapeng Luo, Qiling Tang |
Pattern Recognit. Lett. | 2 |
| 2010 | Pyramid-Based Multi-structure Local Binary Pattern for Texture Classification
Yonggang He, Nong Sang, Changxin Gao |
ACCV (3) | 2 |
| 2010 | Saliency Based on Multi-scale Ratio of DissimilarityabstractRecently, many vision applications tend to utilize saliency maps derived from input images to guide them to focus on processing salient regions in images. In this paper, we propose a simple and effective method to quantify the saliency for each pixel in images. Specially, we define the saliency for a pixel in a ratio form, where the numerator measures the number of dissimilar pixels in its center-surround and the denominator measures the total number of pixels in its center-surround. The final saliency is obtained by combining these ratios of dissimilarity over multiple scales. For images, the saliency map generated by our method not only has a high quality in resolution also looks more reasonable. Finally, we apply our saliency map to extract the salient regions in images, and compare the performance with some state-of-the-art methods over an established ground-truth which contains 1000 images. Rui Huang 0001, Nong Sang, Leyuan Liu 0001, Qiling Tang |
ICPR | 2 |
| 2010 | Automatic Face Replacement in Video Based on 2D Morphable ModelabstractThis paper presents an automatic face replacement approach in video based on 2D morphable model. Our approach includes three main modules: face alignment, face morph, and face fusion. Given a source image and target video, the Active Shape Models (ASM) is adopted to source image and target frames for face alignment. Then the source face shape is warped to match the target face shape by a 2D morphable model. The color and lighting of source face are adjusted to keep consistent with those of target face, and seamlessly blended in the target face. Our approach is fully automatic without user interference, and generates natural and realistic results. Feng Min, Nong Sang, Zhefu Wang |
ICPR | 2 |
| 2010 | A Biologically-Inspired Top-Down Learning Model Based on Visual AttentionabstractA biologically-inspired top-down learning model based on visual attention is proposed in this paper. Low-level visual features are extracted from learning object itself and do not depend on the background information. All the features are expressed as a feature vector, which is looked as a random variable following a normal distribution. So every learning object is represented as the mean and standard deviation. All the learning objects are combined as an object class, which is represented as class's mean and class's standard deviation stored in long-term memory (LTM). Then the learned knowledge is used to find the similar location in an attended image. Experimental results indicate that: when the attended object doesn't always appear in the background similar to that in the learning objects or their combinations change hugely between learning images and attended images, our model is excellent to other two top-down visual attention models. Nong Sang, Longsheng Wei, Yuehuan Wang |
ICPR | 1 |
| 2010 | Motion Detection Based on Biological Correlation Model
Nong Sang, Yuehuan Wang, Qingqing Zheng |
ISNN (2) | 2 |
| 2010 | On selection and combination of weak learners in AdaBoost
Changxin Gao, Nong Sang, Qiling Tang |
Pattern Recognit. Lett. | 2 |
| 2009 | Face Recognition by Estimating Facial Distinctive Information Distribution
Bangyou Da, Nong Sang |
ACCV (3) | 2 |
| 2009 | Segmentation via Incremental Transductive LearningabstractIn this paper, we propose a novel unsupervised clustering method for feature space analysis. We combine mean shift with a transductive learning method, semi-supervised discriminant analysis (SDA), in an incremental learning scheme. We use mean shift clustering to generate the class label, and use SDA to do subspace selection. Both these steps are performed alternately. Our clustering result could maintain good spatial consistency for all data in feature space. On image segmentation, we directly apply our clustering method to the L*a*b* color feature space generated from superpixels, and set each pixel with the clustering label of its superpixel. We test our image segmentation method on Berkeley image data set. Rui Huang 0001, Nong Sang, Qiling Tang |
ICIG | 2 |
| 2009 | Feature Sharing Applied to Palmprint IdentificationabstractPalmprint identification is the means of recognizing an individual from the database using his or her palmprint features. This paper presents a palmprint recognition method by finding common features that can be shared across the classes (different persons), which is a welcome contrast to the existing approaches. At present almost all research is on low resolution images for civil and commercial applications, hence we trained classifiers working in similar feature spaces of the low dimensionality such as texture. In enrollment stage, we randomly select a certain number of samples each class as the training samples. After being preprocessed, the improved SIFT (Scale Invariant Feature Transformation) feature with SAX (Symbolic Aggregate approximation) representation is extracted from the image. The ECOC (Error Correct Output Code) matrix is obtained by layer joint boosting to select the best shared models and their shared features correspondingly. The sum of the weak classifiers which are constructed in each round of boosting turns into the strong classifier. In identification stage, query sample after being extracted features can be directly identified. The experimental results illustrate the effectiveness of the proposed approach. Nong Sang |
ICIG | 2 |
| 2009 | Interactive rotoscoping: Extracting and tracking object sketchabstractIn this paper, we discuss an integrated system for video rotoscoping extracting and tracking object sketch (structural shape) across video sequence. This system consists of two key components: object sketch computing and graph-based object tracking. Given a video clip, we first use a primal sketch algorithm to search bottom-up sketch proposals and additively pursue sketch strokes in each frame. User is allowed to edit the sketch in the beginning frame, such as adding strokes and removing the cluttered edges, and the refined sketch is saved as the template. A graph-based tracking method is then proposed for sequentially matching the template to following frames, and the template is kept update by geometric transformation. Once the matching is unsatisfied at one frame, the system is allowed the user interaction for correction. In the experiments, we apply this system on several videos and present the performance evaluation with comparison. Han Lv, Nong Sang |
ICIP | 4 |
| 2008 | Approximation of salient contours in cluttered scenesabstractThis paper proposes a new approach to describe the salient contours in cluttered scenes. No need to do the preprocessing, such as edge detection, we directly use a set of random straight line segments, as the intermediate level vision tokens, to approximate the salient contours. This line set is modeled by a stochastic framework, marked point process, in which the point denotes the center of lines, and the marker denotes the orientation and length of lines. Generic Gastalt factors of proximity and collinear continuity are embedded to constraint the geometrical inter-relations between lines. Different data likelihoods are used on synthetic and real images. Optimization is done by simulated annealing using Reversible Jump Markov chain Monte Carlo. Our results not only have a good approximation to the salient contours, also make other post-processing application more robust. Rui Huang 0001, Nong Sang, Qiling Tang |
ICPR | 2 |
| 2008 | An empirical study of facial components classification by integrating dimensionality reduction and clusteringabstractIn this paper we present an empirical study of facial components classification by integrating dimensionality reduction and unsupervised clustering. The proposed framework contains two iterative steps: 1) Fixing cluster labels, the facial samples are projected onto lower dimensional subspace through dimensionality reduction method; 2) Fixing the subspace, the clustering algorithm is performed to generate cluster labels. Through iterative steps, clusters are discovered in the lower dimensional subspaces to avoid the curse of dimensionality, while the subspaces are adaptively re-adjusted for global optimality. In order to achieve an effective and robust system, we compare the PCA with the LLE for dimensionality reduction, and compare K-means with LDA-guided K-means(LDA-Km) for unsupervised clustering. The quantitative experimental results prove the LLE and LDA-Km are superior for facial data on public Lotus Hill Institute(LHI) dataset. We also apply the presented study to improve the portrait sketching results. Feng Min, Nong Sang |
ICPR | 3 |
| 2007 | Knowledge-based adaptive thresholding segmentation of digital subtraction angiography images
Nong Sang, Weixue Peng, Tianxu Zhang |
Image Vis. Comput. | 1 |
| 2007 | Contour detection based on contextual influences
Qiling Tang, Nong Sang, Tianxu Zhang |
Image Vis. Comput. | 2 |
| 2007 | Extraction of salient contours from cluttered scenes
Qiling Tang, Nong Sang, Tianxu Zhang |
Pattern Recognit. | 2 |
| 2006 | Extraction of Salient Contours Via Excitatory-Inhibitory Interactions in the Visual Cortex
Qiling Tang, Nong Sang, Tianxu Zhang |
ACCV (2) | 2 |
| 2005 | A Neural Network Model for Extraction of Salient Contours
Qiling Tang, Nong Sang, Tianxu Zhang |
ISNN (2) | 2 |
| 2003 | Local entropy-based transition region extraction and thresholding
Chengxin Yan, Nong Sang, Tianxu Zhang |
Pattern Recognit. Lett. | 2 |
| 2001 | Segmentation of FLIR images by Hopfield neural network with edge constraint
Nong Sang, Tianxu Zhang |
Pattern Recognit. | 1 |
| 1996 | An effective method for identifying small objects on a complicated background
Tianxu Zhang, Nong Sang, Guoyou Wang |
Artif. Intell. Eng. | 2 |