Guiguang Ding

dblp:51/740 · DBLP profile ↗
← Back
207ranked-venue papers
13as first author
87since 2021 · last 2026
0000-0003-0137-9975ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 137 · 9 first-author · 50 since 2021Artificial intelligence and machine learning · 135 · 2 first-author · 67 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Computer networks · 4 · 1 since 2021Security and privacy · 4 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object Detection
abstract
Object detection in High-Resolution Wide (HRW) shots, or gigapixel images, presents unique challenges due to extreme object sparsity and vast scale variations. State-of-the-art methods like SparseFormer have pioneered sparse processing by selectively focusing on important regions, yet they apply a uniform computational model to all selected regions, overlooking their intrinsic complexity differences. This leads to a suboptimal trade-off between performance and efficiency. In this paper, we introduce GigaMoE, a novel backbone architecture that pioneers adaptive computation for this domain by replacing the standard Feed-Forward Networks (FFNs) with a Mixture-of-Experts (MoE) module. Our architecture first employs a shared expert to provide a robust feature baseline for all selected regions. Upon this foundation, our core innovation---a novel Sparsity-Guided Routing mechanism---insightfully repurposes importance scores from the sparse backbone to provide a "computational bonus,'' dynamically engaging a variable number of specialized experts based on content complexity. The entire system is trained efficiently via a loss-free load-balancing technique, eliminating the need for cumbersome auxiliary losses. Extensive experiments show that GigaMoE sets a new state-of-the-art on the PANDA benchmark, improving detection accuracy by 1.1% over SparseFormer while simultaneously reducing the computational cost (FLOPs) by a remarkable 32.3%.
Wenxi Li, Yuetong Wang, Chenyang Lyu, Haozhe Lin, Guiguang Ding
AAAI6
2026 Tracking and Segmenting Anything in Any Modality
abstract
Tracking and segmentation play essential roles in video understanding, providing basic positional information and temporal association of objects within video sequences. Despite their shared objective, existing approaches often tackle these tasks using specialized architectures or modality-specific parameters, limiting their generalization and scalability. Recent efforts have attempted to unify multiple tracking and segmentation sub-tasks from the perspectives of any modality input or multi-task inference. However, these approaches tend to overlook two critical challenges: the distributional gap across different modalities and the feature representation gap across tasks. These issues hinder effective cross-task and cross-modal knowledge sharing, ultimately constraining the development of a true generalist model. To address these limitations, we propose a universal tracking and segmentation framework named SATA, which unifies a broad spectrum of tracking and segmentation subtasks with any modality input. Specifically, a Decoupled Mixture-of-Expert (DeMoE) mechanism is presented to decouple the unified representation learning task into the modeling process of cross-modal shared knowledge and specific information, thus enabling the model to maintain flexibility while enhancing generalization. Additionally, we introduce a Task-aware Multi-object Tracking (TaMOT) pipeline to unify all the task outputs as a unified set of instances with calibrated ID information, thereby alleviating the degradation of task-specific knowledge during multi-task training. SATA demonstrates superior performance on 18 challenging tracking and segmentation benchmarks, offering a novel perspective for more generalizable video understanding.
Tianlu Zhang, Qiang Zhang 0020, Guiguang Ding, Jungong Han
AAAI3
2026 CMPF: Harmonizing Cross-Model Prior Fusion for Open-Vocabulary Segmentation
Sicheng Zhao, Xi Chen 0110, Hongxun Yao, Haosen Yang 0003, Yanhao Zhang 0001, Sheng Jin 0002, Xiatian Zhu, Haonan Lu, Kui Jiang, Guiguang Ding
Int. J. Comput. Vis.10
2026 RepAttn3D: Re-parameterizing 3D attention with spatiotemporal augmentation for video understanding
Xiusheng Lu, Lechao Cheng, Sicheng Zhao, Ying Zheng 0009, Yongheng Wang, Guiguang Ding, Mingli Song
Neural Networks6
2026 Fixing Background Misclassification in Few-Shot Object Detection via Product of Experts
abstract
Few-shot object detection (FSOD) poses a significant challenge due to the difficulty of learning robust and discriminative object representations under limited supervision. A widely adopted solution is the two-stage fine-tuning framework, wherein knowledge acquired from a large-scale base dataset is transferred to a novel dataset containing only a small number of labeled instances. However, this framework is prone to systematically misclassifying novel objects as background, primarily due to incorrect background label caused by the domain gap between base and novel datasets-an issue exacerbated by the sparse representation of novel categories. In this work, we show that this inherent weakness can be exploited by explicitly redefining the category structure and transferring the representations learned during the base training stage. Building on this insight, we propose a simple yet effective framework grounded in the Product of Experts (PoE) formulation, which estimates the joint distribution over background and novel categories by combining the unnormalized logits from independently trained classifiers. Notably, it does not require modifications of the base model or repetition of the base training phase. Furthermore, we introduce a strategy for identifying additional novel-category instances within the base dataset, which effectively augmenting the training set for fine-tuning. The resulting method is architecture-agnostic, imposes negligible overhead, and integrates seamlessly with existing two-stage fine-tuning pipelines. Extensive experiments on PASCAL VOC and COCO demonstrate that the proposed method yields consistent improvements across different baselines, achieving significant gains over state-of-the-art FSOD approaches.
Ding Sheng Ong, Yi Liu 0038, Changjing Shang, Guiguang Ding, Qiang Shen 0001, Jungong Han
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 CAIT: Triple-Win Compression Toward High Accuracy, Fast Inference, and Favorable Transferability for ViTs
abstract
Vision Transformers (ViTs) have emerged as state-of-the-art models for various vision tasks recently. However, their heavy computation costs remain daunting for resource-limited devices. To address this, researchers have dedicated themselves to compressing redundant information in ViTs for acceleration. However, existing approaches generally sparsely drop redundant image tokens by token pruning or brutally remove channels by channel pruning, leading to a sub-optimal balance between model performance and inference speed. Moreover, they struggle when transferring compressed models to downstream vision tasks that require the spatial structure of images, such as semantic segmentation. To tackle these issues, we propose CAIT, a joint compression method for ViTs that achieves a harmonious blend of high accuracy, fast inference speed, and favorable transferability to downstream tasks. Specifically, we introduce an asymmetric token merging (ATME) strategy to effectively integrate neighboring tokens. It can successfully compress redundant token information while preserving the spatial structure of images. On top of it, we further design a consistent dynamic channel pruning (CDCP) strategy to dynamically prune unimportant channels in ViTs. Thanks to CDCP, insignificant channels in multi-head self-attention modules of ViTs can be pruned uniformly, significantly enhancing the model compression. Extensive experiments on multiple benchmark datasets show that our proposed method can achieve state-of-the-art performance across various ViTs.
Hui Chen 0013, Zijia Lin, Sicheng Zhao, Jungong Han, Guiguang Ding
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Safeguarding the target: Enhancing multi-source domain adaptation through domain reorganization
Guiguang Ding, Fan Yang 0083
Pattern Recognit.2
2026 DINO-PCB: Two-stage vision foundation model pretraining and distillation for real-time circuit-board defect detection
Junjie Ke, Lihuo He, Jing Zhang 0037, Yuqi Ji, Hui Chen 0013, Jie Li 0001, Sicheng Zhao, Guiguang Ding, Xinbo Gao 0001
Pattern Recognit.9
2026 LLMI3D: MLLM-Based 3D Perception From a Single 2D Image
abstract
Recent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms. However, current 3D perception methods, especially specialized small models, exhibit poor generalization in open scenarios. On the other hand, multimodal large language models (MLLMs) excel in general capacity but underperform in 3D tasks, due to weak 3D local spatial object perception, poor text-based geometric numerical output, and inability to handle camera focal variations. To address these challenges, we develop LLMI3D, and propose the following solutions: Spatial-Enhanced Local Feature Mining for better 3D spatial feature extraction, 3D Query Token-Derived Info Decoding for precise geometric regression, and Geometry Projection-Based 3D Reasoning for handling camera focal length variations. We are the first to adapt an MLLM for image-based 3D perception. Additionally, we have constructed the IG3D dataset, which provides fine-grained descriptions and question-answer annotations. Extensive experiments demonstrate that our LLMI3D achieves state-of-the-art performance, outperforming other methods by a large margin. We will publicly release our code, models, and dataset.
Fan Yang 0083, Sicheng Zhao, Yanhao Zhang 0001, Hui Chen 0013, Haonan Lu, Jungong Han, Guiguang Ding
IEEE Trans. Multim.7
2025 Scaffold-BPE: Enhancing Byte Pair Encoding for Large Language Models with Simple and Effective Scaffold Token Removal
abstract
Byte Pair Encoding (BPE) serves as a foundation method for text tokenization in the Natural Language Processing (NLP) field. Despite its wide adoption, the original BPE algorithm harbors an inherent flaw: it inadvertently introduces a frequency imbalance for tokens in the text corpus. Since BPE iteratively merges the most frequent token pair in the text corpus to generate a new token and keeps all generated tokens in the vocabulary, it unavoidably holds tokens that primarily act as components of a longer token and appear infrequently on their own. We term such tokens as Scaffold Tokens. Due to their infrequent occurrences in the text corpus, Scaffold Tokens pose a learning imbalance issue. To address that issue, we propose Scaffold-BPE, which incorporates a dynamic scaffold token removal mechanism by parameter-free, computation-light, and easy-to-implement modifications to the original BPE method. This novel approach ensures the exclusion of low-frequency Scaffold Tokens from the token representations for given texts, thereby mitigating the issue of frequency imbalance and facilitating model training. On extensive experiments across language modeling and even machine translation, Scaffold-BPE consistently outperforms the original BPE, well demonstrating its effectiveness.
Haoran Lian, Yizhe Xiong, Jianwei Niu 0002, Shasha Mo, Zhenpeng Su, Zijia Lin, Hui Chen 0013, Jungong Han, Guiguang Ding
AAAI9
2025 Promptable Anomaly Segmentation with SAM Through Self-Perception Tuning
abstract
Segment Anything Model (SAM) has made great progress in anomaly segmentation tasks due to its impressive generalization ability. However, existing methods that directly apply SAM through prompting often overlook the domain shift issue, where SAM performs well on natural images but struggles in industrial scenarios. Parameter-Efficient Fine-Tuning (PEFT) offers a promising solution, but it may yield suboptimal performance by not adequately addressing the perception challenges during adaptation to anomaly images. In this paper, we propose a novel Self-Perception Tuning (SPT) method, aiming to enhance SAM's perception capability for anomaly segmentation. The SPT method incorporates a self-drafting tuning strategy, which generates an initial coarse draft of the anomaly mask, followed by a refinement process. Additionally, a visual-relation-aware adapter is introduced to improve the perception of discriminative relational information for mask generation. Extensive experimental results on several benchmark datasets demonstrate that our SPT method can significantly outperform baseline methods, validating its effectiveness.
Hui-Yue Yang, Hui Chen 0013, Kai Chen 0044, Zijia Lin, Yongliang Tang, Yuming Quan, Jungong Han, Guiguang Ding
AAAI10
2025 Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free Method
abstract
Processing long input remains a significant challenge for large language models (LLMs) due to the scarcity of large-scale long-context training data and the high computational cost of training models for extended context windows.In this paper, we propose Adaptive Grouped Positional Encoding (AdaGroPE), a training-free, plug-and-play method to enhance long-context understanding in existing LLMs.AdaGroPE progressively increases the reuse count of relative positions as the distance grows and dynamically adapts the positional encoding mapping to sequence length, thereby fully exploiting the range of pre-trained position embeddings.Its design is consistent with the principles of rotary position embedding (RoPE) and aligns with human perception of relative distance, enabling robust performance in realworld settings with variable-length inputs.Extensive experiments across various benchmarks demonstrate that our AdaGroPE consistently achieves state-of-the-art performance, surpassing baseline methods and even outperforming LLMs inherently designed for long-context processing on certain tasks.
Hui Chen 0013, Zijia Lin, Jungong Han, Guiguang Ding
ACL (1)6
2025 Breaking the Stage Barrier: A Novel Single-Stage Approach to Long Context Extension for Large Language Models
abstract
Recently, Large language models (LLMs) have revolutionized Natural Language Processing (NLP). Pretrained LLMs, due to limited training context size, struggle with handling long token sequences, limiting their performance on various downstream tasks. Current solutions toward long context modeling often employ multi-stage continual pertaining, which progressively increases the effective context length through several continual pretraining stages. However, those approaches require extensive manual tuning and human expertise. In this paper, we introduce a novel single-stage continual pretraining method, Head-Adaptive Rotary Position Embedding (HARPE), to equip LLMs with long context modeling capabilities while simplifying the training process. Our HARPE leverages different Rotary Position Embedding (RoPE) base frequency values across different attention heads and directly trains LLMs on the target context length. Extensive experiments on 4 language modeling benchmarks, including the latest RULER benchmark, demonstrate that HARPE excels in understanding and integrating long-context tasks with single-stage training, matching and even outperforming existing multi-stage methods. Our results highlight that HARPE successfully breaks the stage barrier for training LLMs with long context modeling capabilities.
Haoran Lian, Junmin Chen, Yizhe Xiong, Wenping Hu, Guiguang Ding, Hui Chen 0013, Jianwei Niu 0002, Zijia Lin, Di Zhang 0026
COLING6
2025 Bayesian Prompt Flow Learning for Zero-Shot Anomaly Detection
abstract
Recently, vision-language models (e.g. CLIP) have demonstrated remarkable performance in zero-shot anomaly detection (ZSAD). By leveraging auxiliary data during training, these models can directly perform cross-category anomaly detection on target datasets, such as detecting defects on industrial product surfaces or identifying tumors in organ tissues. Existing approaches typically construct text prompts through either manual design or the optimization of learnable prompt vectors. However, these methods face several challenges: 1) handcrafted prompts require extensive expert knowledge and trial-and-error; 2) single-form learnable prompts struggle to capture complex anomaly semantics; and 3) an unconstrained prompt space limits generalization to unseen categories. To address these issues, we propose Bayesian Prompt Flow Learning (Bayes-PFL), which models the prompt space as a learnable probability distribution from a Bayesian perspective. Specifically, a prompt flow module is designed to learn both image-specific and image-agnostic distributions, which are jointly utilized to regularize the text prompt space and improve the model's generalization on unseen categories. These learned distributions are then sampled to generate diverse text prompts, effectively covering the prompt space. Additionally, a residual cross-model attention (RCA) module is introduced to better align dynamic text embeddings with fine-grained image features. Extensive experiments on 15 industrial and medical datasets demonstrate our method's superior performance. The code is available at https://github.com/xiaozhen228/Bayes-PFL.
Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Qiyu Chen 0002, Zhengtao Zhang, Guiguang Ding
CVPR8
2025 DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval
abstract
The parameter-efficient adaptation of the image-text pre-training model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive understanding at the video level. Three key discrepancies emerge in the transfer from image-level to video-level: vision, language, and alignment. However, existing methods mainly focus on vision while neglecting language and alignment. In this paper, we propose Discrepancy Reduction in Vision, Language, and Alignment (DiscoVLA), which simultaneously mitigates all three discrepancies. Specifically, we introduce Image-Video Features Fusion to integrate image-level and video-level features, effectively tackling both vision and language discrepancies. Additionally, we generate pseudo image captions to learn fine-grained image-level alignment. To mitigate alignment discrepancies, we propose Image-To-Video Alignment Distillation, which leverages image-level alignment knowledge to enhance video-level alignment. Extensive experiments demonstrate the superiority of our DiscoVLA. In particular, on MSRVTT with CLIP (ViT-B/16), DiscoVLA outperforms previous methods by 2.2% R@1 and 7.5% R@sum. The code is available at https://github.com/LunarShen/DsicoVLA.
Leqi Shen, Guoqiang Gong, Tianxiang Hao 0001, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding
CVPR9
2025 LSNet: See Large, Focus Small
abstract
Vision network designs, including Convolutional Neural Networks and Vision Transformers, have significantly advanced the field of computer vision. Yet, their complex computations pose challenges for practical deployments, particularly in real-time applications. To tackle this issue, researchers have explored various lightweight and efficient network designs. However, existing lightweight models predominantly leverage self-attention mechanisms and convolutions for token mixing. This dependence brings limitations in effectiveness and efficiency in the perception and aggregation processes of lightweight networks, hindering the balance between performance and efficiency under limited computational budgets. In this paper, we draw inspiration from the dynamic heteroscale vision ability inherent in the efficient human vision system and propose a "See Large, Focus Small" strategy for lightweight vision network design. We introduce LS (Large-Small) convolution, which combines large-kernel perception and small-kernel aggregation. It can efficiently capture a wide range of perceptual information and achieve precise feature aggregation for dynamic and complex visual representations, thus enabling proficient processing of visual information. Based on LS convolution, we present LSNet, a new family of lightweight models. Extensive experiments demonstrate that LSNet achieves superior performance and efficiency over existing lightweight networks in various vision tasks. Codes and models are available at https://github.com/jameslahm/lsnet.
Hui Chen 0013, Zijia Lin, Jungong Han, Guiguang Ding
CVPR5
2025 HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator
abstract
AIGC images are prevalent across various fields, yet they frequently suffer from quality issues like artifacts and unnatural textures. Specialized models aim to predict defect region heatmaps but face two primary challenges: (1) lack of explainability, failing to provide reasons and analyses for subtle defects, and (2) inability to leverage common sense and logical reasoning, leading to poor generalization. Multimodal large language models (MLLMs) promise better comprehension and reasoning but face their own challenges: (1) difficulty in fine-grained defect localization due to the limitations in capturing tiny details, and (2) constraints in providing pixel-wise outputs necessary for precise heatmap generation. To address these challenges, we propose HEIE: a novel MLLM-Based Hierarchical Explainable Image Implausibility Evaluator. We introduce the CoT-Driven Explainable Trinity Evaluator, which integrates heatmaps, scores, and explanation outputs, using CoT to decompose complex tasks into subtasks of increasing difficulty and enhance interpretability. Our Adaptive Hierarchical Implausibility Mapper synergizes low-level image features with high-level mapper tokens from LLMs, enabling precise local-to-global hierarchical heatmap predictions through an uncertainty-based adaptive token approach. Moreover, we propose a new dataset: Expl-AIGI-Eval, designed to facilitate interpretable implausibility evaluation of AIGC images. Our method demonstrates state-of-the-art performance through extensive experiments. Our project is at https://yfthu.github.io/HEIE/.
Fan Yang 0083, Ru Zhen, Yanhao Zhang 0001, Haoxiang Chen 0007, Haonan Lu, Sicheng Zhao, Guiguang Ding
CVPR8
2025 DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs
abstract
Minxuan Lv, Zhenpeng Su, Leiyu Pan, Yizhe Xiong, Zijia Lin, Hui Chen, Wei Zhou, Jungong Han, Guiguang Ding, Wenwu Ou, Di Zhang, Kun Gai, Songlin Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Minxuan Lv, Zhenpeng Su, Leiyu Pan, Yizhe Xiong, Zijia Lin, Hui Chen 0013, Wei Zhou 0019, Jungong Han, Guiguang Ding, Wenwu Ou, Di Zhang 0026, Kun Gai, Songlin Hu 0001
EMNLP9
2025 Temporal Scaling Law for Large Language Models
abstract
Yizhe Xiong, Xiansheng Chen, Xin Ye, Hui Chen, Zijia Lin, Haoran Lian, Zhenpeng Su, Wei Huang, Jianwei Niu, Jungong Han, Guiguang Ding. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yizhe Xiong, Xiansheng Chen, Hui Chen 0013, Zijia Lin, Haoran Lian, Zhenpeng Su, Jianwei Niu 0002, Jungong Han, Guiguang Ding
EMNLP11
2025 LBPE: Long-token-first Tokenization to Improve Large Language Models
abstract
The prevalent use of Byte Pair Encoding (BPE) in Large Language Models (LLMs) facilitates robust handling of subword units and avoids issues of out-of-vocabulary words. Despite its success, a critical challenge persists: long tokens, rich in semantic information, have fewer occurrences in tokenized datasets compared to short tokens, which can result in imbalanced learning issue across different tokens. To address that, we propose LBPE, which prioritizes long tokens during the encoding process. LBPE generates tokens according to their descending order of token length rather than their ranks in the vocabulary, granting longer tokens higher priority during the encoding process. Consequently, LBPE smooths the frequency differences between short and long tokens, and thus mitigates the learning imbalance. Extensive experiments across diverse language modeling tasks demonstrate that LBPE consistently outperforms the original BPE, well demonstrating its effectiveness.
Haoran Lian, Yizhe Xiong, Zijia Lin, Jianwei Niu 0002, Shasha Mo, Hui Chen 0013, Guiguang Ding
ICASSP8
2025 DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup
abstract
Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised prompt learning or fine-tuning on seen classes. However, their cross-category generalization largely depends on prior knowledge of real seen anomaly samples. In this paper, we propose a novel framework, namely DictAS, which enables a unified model to detect visual anomalies in unseen object categories without any retraining on the target data, only employing a few normal reference images as visual prompts. The insight behind DictAS is to transfer dictionary lookup capabilities to the FSAS task for unseen classes via self-supervised learning, instead of merely memorizing the normal and abnormal feature patterns from the training set. Specifically, DictAS mainly consists of three components: (1) Dictionary Construction - to simulate the index and content of a real dictionary using features from normal reference images. (2) Dictionary Lookup - to retrieve queried region features from the dictionary via a sparse lookup strategy. When a query feature cannot be retrieved, it is classified as an anomaly. (3) Query Discrimination Regularization - to enhance anomaly discrimination by making abnormal features harder to retrieve from the dictionary. To achieve this, Contrastive Query Constraint and Text Alignment Constraint are further proposed. Extensive experiments on seven public industrial and medical datasets demonstrate that DictAS consistently outperforms state-of-the-art FSAS methods.
Zhen Qu, Xian Tao, Xinyi Gong, Shichen Qu, Fei Shen 0002, Zhengtao Zhang, Mukesh Prasad, Guiguang Ding
ICCV10
2025 YOLOE: Real-Time Seeing Anythi
Hui Chen 0013, Zijia Lin, Jungong Han, Guiguang Ding
ICCV6
2025 TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
abstract
Most text-video retrieval methods utilize the text-image pre-trained models like CLIP as a backbone. These methods process each sampled frame independently by the image encoder, resulting in high computational overhead and limiting practical deployment. Addressing this, we focus on efficient text-video retrieval by tackling two key challenges: 1. From the perspective of trainable parameters, current parameter-efficient fine-tuning methods incur high inference costs; 2. From the perspective of model complexity, current token compression methods are mainly designed for images to reduce spatial redundancy but overlook temporal redundancy in consecutive frames of a video. To tackle these challenges, we propose Temporal Token Merging (TempMe), a parameter-efficient and training-inference efficient text-video retrieval architecture that minimizes trainable parameters and model complexity. Specifically, we introduce a progressive multi-granularity framework. By gradually combining neighboring clips, we reduce spatio-temporal redundancy and enhance temporal modeling across different frames, leading to improved efficiency and performance. Extensive experiments validate the superiority of our TempMe. Compared to previous parameter-efficient text-video retrieval methods, TempMe achieves superior performance with just 0.50M trainable parameters. It significantly reduces output tokens by 95% and GFLOPs by 51%, while achieving a 1.8X speedup and a 4.4% R-Sum improvement. With full fine-tuning, TempMe achieves a significant 7.9% R-Sum improvement, trains 1.57X faster, and utilizes 75.2% GPU memory usage. The code is available at https://github.com/LunarShen/TempMe.
Leqi Shen, Tianxiang Hao 0001, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding
ICLR8
2025 Exploiting Position Information in Convolutional Kernels for Structural Re-parameterization
abstract
In order to boost the performance of a convolutional neural network (CNN), several approaches have shown the benefit of enhancing the spatial encoding of feature maps. However, few works paid attention to the positional properties of convolutional kernels. In this paper, we demonstrate that different kernel positions are of different importance, which depends on the task, dataset and architecture, and adaptively emphasizing the informative parts in convolutional kernels can lead to considerable improvement. Therefore, we propose a novel structural re-parameterization Position Boosting Convolution (PBConv) to exploit and enhance the position information in the convolutional kernel. PBConv consists of several concurrent small convolutional kernels, which can be equivalently converted to the original kernel and bring no extra inference cost. Different from existing structural re-parameterization methods, PBconv searches for the optimal re-parameterized structure by a fast heuristic algorithm based on the dispersion of kernel weights. Such heuristic search is efficient yet effective, well adapting the varying kernel weight distribution. As a result, PBConv can significantly improve the representational power of a model, especially its ability to extract fine-grained low-level features. Importantly, PBConv is orthogonal to procedural re-parameterization methods and can further boost performance based on them.
Tianxiang Hao 0001, Hui Chen 0013, Guiguang Ding
IJCAI3
2025 Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
abstract
Vision-language models (VLMs) exhibit remarkable zero-shot capabilities but struggle with distribution shifts in downstream tasks when labeled data is unavailable, which has motivated the development of Test-Time Adaptation (TTA) to improve VLMs' performance during inference without annotations. Among various TTA approaches, cache-based methods show promise by preserving historical knowledge from low-entropy samples in a dynamic cache and fostering efficient adaptation. However, these methods face two critical reliability challenges: (1) entropy often becomes unreliable under distribution shifts, causing error accumulation in the cache and degradation in adaptation performance; (2) the final predictions may be unreliable due to inflexible decision boundaries that fail to accommodate large downstream shifts. To address these challenges, we propose a Reliable Test-time Adaptation (ReTA) method that integrates two complementary strategies to enhance reliability from two perspectives. First, to mitigate the unreliability of entropy as a sample selection criterion for cache construction, we introduce Consistency-aware Entropy Reweighting (CER), which incorporates consistency constraints to weight entropy during cache updating. While conventional approaches rely solely on low entropy for cache prioritization and risk introducing noise, our method leverages predictive consistency to maintain a high-quality cache and facilitate more robust adaptation. Second, we present Diversity-driven Distribution Calibration (DDC), which models class-wise text embeddings as multivariate Gaussian distributions, enabling adaptive decision boundaries for more accurate predictions across visually diverse content. Extensive experiments demonstrate that ReTA consistently outperforms state-of-the-art methods, particularly under real-world distribution shifts.
Hui Chen 0013, Yizhe Xiong, Mengyao Lyu, Zijia Lin, Shuaicheng Niu, Sicheng Zhao, Jungong Han, Guiguang Ding
ACM Multimedia10
2025 CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts
abstract
Zhenpeng Su, Xing W, Zijia Lin, Yizhe Xiong, Minxuan Lv, Guangyuan Ma, Hui Chen, Songlin Hu, Guiguang Ding. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Zhenpeng Su, Xing Wu 0002, Zijia Lin, Yizhe Xiong, Minxuan Lv, Guangyuan Ma, Hui Chen 0013, Songlin Hu 0001, Guiguang Ding
NAACL (Long Papers)9
2025 Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided Decoding
abstract
Xinhao Xu, Hui Chen, Mengyao Lyu, Sicheng Zhao, Yizhe Xiong, Zijia Lin, Jungong Han, Guiguang Ding. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hui Chen 0013, Mengyao Lyu, Sicheng Zhao, Yizhe Xiong, Zijia Lin, Jungong Han, Guiguang Ding
NAACL (Long Papers)8
2025 FastVID: Dynamic Density Pruning for Fast Video Large Language Models
abstract
Video Large Language Models have demonstrated strong video understanding capabilities, yet their practical deployment is hindered by substantial inference costs caused by redundant video tokens. Existing pruning techniques fail to effectively exploit the spatiotemporal redundancy present in video data. To bridge this gap, we perform a systematic analysis of video redundancy from two perspectives: temporal context and visual context. Leveraging these insights, we propose Dynamic Density Pruning for Fast Video LLMs termed FastVID. Specifically, FastVID dynamically partitions videos into temporally ordered segments to preserve temporal structure and applies a density-based token pruning strategy to maintain essential spatial and temporal information. Our method significantly reduces computational overhead while maintaining temporal and visual integrity. Extensive evaluations show that FastVID achieves state-of-the-art performance across various short- and long-video benchmarks on leading Video LLMs, including LLaVA-OneVision, LLaVA-Video, Qwen2-VL, and Qwen2.5-VL. Notably, on LLaVA-OneVision-7B, FastVID effectively prunes $\textbf{90.3\%}$ of video tokens, reduces FLOPs to $\textbf{8.3\%}$, and accelerates the LLM prefill stage by $\textbf{7.1}\times$, while maintaining $\textbf{98.0\%}$ of the original accuracy. The code is available at https://github.com/LunarShen/FastVID.
Leqi Shen, Guoqiang Gong, Pengzhang Liu, Sicheng Zhao, Guiguang Ding
NeurIPS7
2025 PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
abstract
Recently, large vision-language models (LVLMs) have rapidly gained popularity for their strong generation and reasoning capabilities given diverse multimodal inputs. However, these models incur significant computational and memory overhead during inference, which greatly hinders the efficient deployment in practical scenarios. The extensive key-value (KV) cache, necessitated by the lengthy input and output sequences, notably contributes to the high inference cost. Based on this, recent works have investigated ways to reduce the KV cache size for higher efficiency. Although effective, they generally overlook the distinct importance distributions of KV vectors across layers and maintain the same cache size for each layer during the next token prediction. This results in the significant contextual information loss for certain layers, leading to notable performance decline. To address this, we present PrefixKV. It reframes the challenge of determining KV cache sizes for all layers into the task of searching for the optimal global prefix configuration. With an adaptive layer-wise KV retention recipe based on binary search, the maximum contextual information can thus be preserved in each layer, facilitating the generation. Extensive experiments demonstrate that our method achieves the state-of-the-art performance compared with others. It exhibits superior inference efficiency and generation quality trade-offs, showing promising potential for practical applications. Code is available at https://github.com/THU-MIG/PrefixKV.
Hui Chen 0013, Jianchao Tan, Zijia Lin, Jungong Han, Guiguang Ding
NeurIPS8
2025 AD2: Anomaly Detection During Training an Distillation-Based Anomaly Detection Model
Kai Chen 0044, Xiaowang Wang, Huiyue Yang, Hui Chen 0013, Yuwang Wang, Sicheng Zhao, Guiguang Ding
PRCV (12)7
2025 Correction: Multi-source-free Domain Adaptive Object Detection
Sicheng Zhao, Huizai Yao, Chuang Lin 0003, Yue Gao 0002, Guiguang Ding
Int. J. Comput. Vis.5
2025 Domain knowledge boosted adaptation: Leveraging vision-language models for multi-source domain adaptation
Juexiao Feng, Guiguang Ding
Neurocomputing3
2025 Hyper-YOLO: When Visual Object Detection Meets Hypergraph Computation
abstract
We introduce Hyper-YOLO, a new object detection method that integrates hypergraph computations to capture the complex high-order correlations among visual features. Traditional YOLO models, while powerful, have limitations in their neck designs that restrict the integration of cross-level features and the exploitation of high-order feature interrelationships. To address these challenges, we propose the Hypergraph Computation Empowered Semantic Collecting and Scattering (HGC-SCS) framework, which transposes visual feature maps into a semantic space and constructs a hypergraph for high-order message propagation. This enables the model to acquire both semantic and structural information, advancing beyond conventional feature-focused learning. Hyper-YOLO incorporates the proposed Mixed Aggregation Network (MANet) in its backbone for enhanced feature extraction and introduces the Hypergraph-Based Cross-Level and Cross-Position Representation Network (HyperC2Net) in its neck. HyperC2Net operates across five scales and breaks free from traditional grid structures, allowing for sophisticated high-order interactions across levels and positions. This synergy of components positions Hyper-YOLO as a state-of-the-art architecture in various scale models, as evidenced by its superior performance on the COCO dataset. Specifically, Hyper-YOLO-N significantly outperforms the advanced YOLOv8-N and YOLOv9-T with 12% and 9% improvements.
Yifan Feng 0001, Jiangang Huang, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Guiguang Ding, Rongrong Ji, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 SDRS: Sentiment-Aware Disentangled Representation Shifting for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) aims to leverage the complementary information from multiple modalities for affective understanding of user-generated videos. Existing methods mainly focused on designing sophisticated feature fusion strategies to integrate the separately extracted multimodal representations, ignoring the interference of the information irrelevant to sentiment. In this paper, we propose to disentangle the unimodal representations into sentiment-specific and sentiment-independent features, the former of which are fused for the MSA task. Specifically, we design a novel Sentiment-aware Disentangled Representation Shifting framework, termed SDRS, with two components.Interactive sentiment-aware representation disentanglementaims to extract sentiment-specific feature representations for each nonverbal modality by considering the contextual influence of other modalities with the newly developed cross-attention autoencoder.Attentive cross-modal representation shiftingtries to shift the textual representation in a latent token space using the nonverbal sentiment-specific representations after projection. The shifted representation is finally employed to fine-tune a pre-trained language model for multimodal sentiment analysis. Extensive experiments are conducted on three public benchmark datasets, i.e., CMU-MOSI, CMU-MOSEI, and CH-SIMS. The results demonstrate that the proposed SDRS framework not only obtains state-of-the-art results based solely on multimodal labels but also outperforms the methods that additionally require the labels of each modality.
Sicheng Zhao, Zhenhua Yang, Henglin Shi, Lingpengkun Meng, Bing Qin 0001, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding
IEEE Trans. Affect. Comput.9
2025 Toward Realistic Hierarchical Object Detection: Problem, Benchmark, and Solution
abstract
With the continuous advancement of deep learning, object detection has made remarkable progress in accurately identifying a wide range of object categories, even within increasingly complex scenes. However, as the number of categories grows, visual concepts naturally organize into a label hierarchy. We contend that existing hierarchical classification and detection methods predominantly prioritize fine-grained prediction, potentially leading to inconsistencies with realistic human perception. From this perspective, we investigate the Hierarchical Object Detection (HOD) problem to better align with real-world perception. To address the lack of benchmarks in the field, we build a large-scale HOD benchmark termed RHOD with open-source datasets, comprising 740 categories. To better align the hierarchical object detectors towards realistic perception, we propose a new evaluation metric named Hierarchical Average Precision (HAP). Furthermore, we present a novel hierarchical object detection method that includes two components, Tree Soft Labeling (TSL) and Hierarchical Extension and Suppression (HES). Our method mitigates the issue of overconfidence in fine-grained predictions, which has been prevalent in previous approaches. We evaluate a range of existing methods on the RHOD benchmark, including plain, hierarchical, and open-vocabulary models. Additionally, we perform comprehensive experiments to assess the performance of our proposed method. The experimental results show that our method achieves state-of-the-art performance on the RHOD benchmark.
Juexiao Feng, Yuhong Yang 0008, Mengyao Lyu, Tianxiang Hao 0001, Yi-Jie Huang, Yanchun Xie, Jungong Han, Liuyu Xiang, Guiguang Ding
IEEE Trans. Circuits Syst. Video Technol.10
2025 DAR-Prompt: Dynamic Regulation in Prompt Tuning for Multi-Label Zero-Shot Learning
abstract
Prompt tuning achieves superior performance across a wide range of tasks, including multi-label zero-shot classification. Existing approaches employ multiple prompts to acquire comprehensive knowledge from categories, demonstrating state-of-the-art performance and significant computational efficiency. However, two main challenges still exist in these methods that impede the full potential of generalization. First, the class imbalance is not carefully addressed. Despite some efforts to adopt re-weighted loss functions to alleviate the positive-negative imbalance, such strategies tend to exacerbate the class imbalance by over-suppression of labels with fewer samples and overfitting to dominant classes. Second, the multi-prompt methods neglect the interactions between prompts during parameter optimization, underestimating the potential of prompts and leading to suboptimal performance. To address these issues, we present a novel framework named Dynamic Regulation in Prompt Tuning (DAR-Prompt). DAR-Prompt introduces three dynamic components: semantic regulator and debiased regulator to address the class imbalance, along with contrastive gradient regularization to enhance feature separation through prompt interactions during the backward pass. Specifically, the semantic regulator generates class-adaptive thresholds to compensate for tail classes and mitigate over-suppression, while the debiased regulator focuses on learning biased classes by rectifying overconfident predictions. Moreover, we apply dynamic regularization to the gradient update directions of prompts to promote orthogonality, thereby enhancing feature distinctiveness. Extensive experiments on several benchmarks show that our method can achieve state-of-the-art performance, well demonstrating its effectiveness and superiority. Code is available at https://github.com/Evelyn1ywliang/DAR-Prompt.
Hui Chen 0013, Zijia Lin, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding
IEEE Trans. Image Process.10
2025 Source-Free Object Detection With Detection Transformer
abstract
Source-Free Object Detection (SFOD) enables knowledge transfer from a source domain to an unsupervised target domain for object detection without access to source data. Most existing SFOD approaches are either confined to conventional object detection (OD) models like Faster R-CNN or designed as general solutions without tailored adaptations for novel OD architectures, especially Detection Transformer (DETR). In this paper, we introduce Feature Reweighting ANd Contrastive Learning NetworK (FRANCK), a novel SFOD framework specifically designed to perform query-centric feature enhancement for DETRs. FRANCK comprises four key components: 1) an Objectness Score-based Sample Reweighting (OSSR) module that computes attention-based objectness scores on multi-scale encoder feature maps, reweighting the detection loss to emphasize less-recognized regions; 2) a Contrastive Learning with Matching-based Memory Bank (CMMB) module that integrates multi-level features into memory banks, enhancing class-wise contrastive learning; 3) an Uncertainty-weighted Query-fused Feature Distillation (UQFD) module that improves feature distillation through prediction quality reweighting and query feature fusion; and 4) an improved self-training pipeline with a Dynamic Teacher Updating Interval (DTUI) that optimizes pseudo-label quality. By leveraging these components, FRANCK effectively adapts a source-pre-trained DETR model to a target domain with enhanced robustness and generalization. Extensive experiments on several widely used benchmarks demonstrate that our method achieves state-of-the-art performance, highlighting its effectiveness and compatibility with DETR-based SFOD models.
Huizai Yao, Sicheng Zhao, Shuo Lu, Hui Chen 0013, Tengfei Xing, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding
IEEE Trans. Image Process.10
2025 Temporal Modeling With Frozen Vision-Language Foundation Models for Parameter-Efficient Text-Video Retrieval
abstract
Temporal modeling plays an important role in the effective adaption of the powerful pretrained text-image foundation model into text-video retrieval. However, existing methods often rely on additional heavy trainable modules, such as transformer or BiLSTM, which are inefficient. In contrast, we avoid introducing such heavy components by leveraging frozen foundation models. To this end, we propose temporal modeling with frozen vision-language foundation models (TFVL) to model the temporal dynamics with fixed encoders. Specifically, text encoder temporal modeling (TextTemp) and image encoder temporal modeling (ImageTemp) apply frozen text and image encoders within the video head and video backbone, respectively. TextTemp uses a frozen text encoder to interpret frame representations as "visual words" within a temporal "sentence," capturing temporal dependencies. On the other hand, ImageTemp uses a frozen image encoder to treat all frame tokens as a unified visual entity, learning spatiotemporal information. The total trainable parameters of our method, comprising a lightweight projection and several prompt tokens, are significantly fewer than those in other existing methods. We evaluate the effectiveness of our method on MSR-VTT, DiDeMo, ActivityNet, and LSMDC. Compared with full fine-tuning on MSR-VTT, our TFVL achieves an average 3.25% gain in R@1 with merely 0.35% of the parameters. Extensive experiments demonstrate that the proposed TFVL outperforms state-of-the-art methods with significantly fewer parameters.
Leqi Shen, Tianxiang Hao 0001, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding
IEEE Trans. Neural Networks Learn. Syst.8
2025 Spatio-Temporal Attention for Text-Video Retrieval
abstract
Text-video retrieval, a fundamental task for associating textual descriptions with video content, has become increasingly important in the video domain. Most existing methods focus on the single-modality features only considering the knowledge within individual video or text modalities, often neglecting cross-modal interactions. However, a text description corresponds to a specific spatio-temporal content within a video, involving a certain segment of a frame sequence and distinct sub-regions within these frames. Therefore, we focus on the text-conditioned video features to bridge the modality gap. In this article, we propose Spatio-Temporal Attention for video-text retrieval, termed STAttn, which utilizes textual information to focus on the spatio-temporal video content. Our final text-conditioned video features are generated from the text-related video frames and the text-related regions within these frames. First, we propose the Spatial Text-Attention Module (STAM) to learn the spatial information within video frames. STAM introduces the text-related salient patches to capture more fine-grained details. Second, we propose the Temporal Text-Attention Module (TTAM) to learn the temporal relationships between video frames. Temporal Triplet loss is proposed in TTAM to enhance the attention toward the text-related frames. Thus, the two modules learn the text-related spatio-temporal content from both intra-frame and inter-frame aspects. Extensive experiments on three benchmark datasets, MSRVTT, ActivityNet, and DiDeMo, demonstrate that our STAttn outperforms the state-of-the-art methods.
Leqi Shen, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Debiased Novel Category Discovering and Localization
abstract
In recent years, object detection in deep learning has experienced rapid development. However, most existing object detection models perform well only on closed-set datasets, ignoring a large number of potential objects whose categories are not defined in the training set. These objects are often identified as background or incorrectly classified as pre-defined categories by the detectors. In this paper, we focus on the challenging problem of Novel Class Discovery and Localization (NCDL), aiming to train detectors that can detect the categories present in the training data, while also actively discover, localize, and cluster new categories. We analyze existing NCDL methods and identify the core issue: object detectors tend to be biased towards seen objects, and this leads to the neglect of unseen targets. To address this issue, we first propose an Debiased Region Mining (DRM) approach that combines class-agnostic Region Proposal Network (RPN) and class-aware RPN in a complementary manner. Additionally, we suggest to improve the representation network through semi-supervised contrastive learning by leveraging unlabeled data. Finally, we adopt a simple and efficient mini-batch K-means clustering method for novel class discovery. We conduct extensive experiments on the NCDL benchmark, and the results demonstrate that the proposed DRM approach significantly outperforms previous methods, establishing a new state-of-the-art.
Juexiao Feng, Yuhong Yang 0008, Yanchun Xie, Yandong Guo, Liuyu Xiang, Guiguang Ding
AAAI9
2024 Geometry-Guided Domain Generalization for Monocular 3D Object Detection
abstract
Monocular 3D object detection (M3OD) is important for autonomous driving. However, existing deep learning-based methods easily suffer from performance degradation in real-world scenarios due to the substantial domain gap between training and testing. M3OD's domain gaps are complex, including camera intrinsic parameters, extrinsic parameters, image appearance, etc. Existing works primarily focus on the domain gaps of camera intrinsic parameters, ignoring other key factors. Moreover, at the feature level, conventional domain invariant learning methods generally cause the negative transfer issue, due to the ignorance of dependency between geometry tasks and domains. To tackle these issues, in this paper, we propose MonoGDG, a geometry-guided domain generalization framework for M3OD, which effectively addresses the domain gap at both camera and feature levels. Specifically, MonoGDG consists of two major components. One is geometry-based image reprojection, which mitigates the impact of camera discrepancy by unifying intrinsic parameters, randomizing camera orientations, and unifying the field of view range. The other is geometry-dependent feature disentanglement, which overcomes the negative transfer problems by incorporating domain-shared and domain-specific features. Additionally, we leverage a depth-disentangled domain discriminator and a domain-aware geometry regression attention mechanism to account for the geometry-domain dependency. Extensive experiments on multiple autonomous driving benchmarks demonstrate that our method achieves state-of-the-art performance in domain generalization for M3OD.
Fan Yang 0083, Hui Chen 0013, Sicheng Zhao, Chenghao Zhang 0006, Guiguang Ding
AAAI7
2024 One-dimensional Adapter to Rule Them All: Concepts, Diffusion Models and Erasing Applications
abstract
The prevalent use of commercial and open-source diffusion models (DMs) for text-to-image generation prompts risk mitigation to prevent undesired behaviors. Existing concept erasing methods in academia are all based on full parameter or specification-based fine-tuning, from which we observe the following issues: 1) Generation alteration towards erosion: Parameter drift during target elimination causes alterations and potential deformations across all generations, even eroding other concepts at varying degrees, which is more evident with multi-concept erased; 2) Transfer in-ability & deployment inefficiency: Previous model-specific erasure impedes the flexible combination of concepts and the training-free transfer towards other models, resulting in linear cost growth as the deployment scenarios increase. To achieve non-invasive, precise, customizable, and transferable elimination, we ground our erasing framework on one-dimensional adapters to erase multiple concepts from most DMs at once across versatile erasing applications. The concept-SemiPermeable structure is injected as a Membrane (SPM) into any DM to learn targeted erasing, and mean-time the alteration and erosion phenomenon is effectively mitigated via a novel Latent Anchoring fine-tuning strategy. Once obtained, SPMs can be flexibly combined and plug-and-play for other DMs without specific re-tuning, enabling timely and efficient adaptation to diverse scenarios. During generation, our Facilitated Transport mechanism dynamically regulates the permeability of each SPM to re-spond to different input prompts, further minimizing the impact on other concepts. Quantitative and qualitative results across ~40 concepts, 7 DMs and 4 erasing applications have demonstrated the superior erasing of SPM. Our code and pre-tuned SPMs are available on the project page https:/lyumengyao.github.io/projects/spm.
Mengyao Lyu, Yuhong Yang 0008, Haiwen Hong, Hui Chen 0013, Xuan Jin, Yuan He 0011, Hui Xue 0001, Jungong Han, Guiguang Ding
CVPR9
2024 Rep ViT: Revisiting Mobile CNN From ViT Perspective
abstract
Recently, lightweight Vision Transformers (ViTs) demon-strate superior performance and lower latency, compared with lightweight Convolutional Neural Networks (CNNs), on resource-constrained mobile devices. Researchers have discovered many structural connections be-tween lightweight ViTs and lightweight CNNs. However, the notable architectural disparities in the block structure, macro, and micro designs between them have not been adequately examined. In this study, we revisit the efficient design of lightweight CNNs from ViT perspective and emphasize their promising prospect for mobile devices. Specifically, we incrementally enhance the mobile-friendliness of a standard lightweight CNN, i.e., MobileNetV3, by integrating the efficient architectural designs of lightweight ViTs. This ends up with a new family of pure lightweight CNNs, namely RepViT. Extensive experiments show that RepViT outperforms existing state-of-the-art lightweight ViTs and exhibits favorable latency in various vision tasks. Notably, on ImageNet, RepViT achieves over 80% top-1 accuracy with 1.0 ms latency on an iPhone 12, which is the first time for a lightweight model, to the best of our knowledge. Besides, when RepViT meets SAM, our RepViT-SAM can achieve nearly 10x faster inference than the advanced MobileSAM. Codes and models are available at https://github.com/THU-MIG/RepViT.
Hui Chen 0013, Zijia Lin, Jungong Han, Guiguang Ding
CVPR5
2024 Context Enhancement with Reconstruction as Sequence for Unified Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection (AD) aims to train robust detection models using only normal samples, while can generalize well to unseen anomalies. Recent research focuses on a unified unsupervised AD setting in which only one model is trained for all classes, i.e., n-class-one-model paradigm. Feature-reconstruction-based methods achieve state-of-the-art performance in this scenario. However, existing methods often suffer from a lack of sufficient contextual awareness, thereby compromising the quality of the reconstruction. To address this issue, we introduce a novel Reconstruction as Sequence (RAS) method, which enhances the contextual correspondence during feature reconstruction from a sequence modeling perspective. In particular, based on the transformer technique, we integrate a specialized RASFormer block into RAS. This block enables the capture of spatial relationships among different image regions and enhances sequential dependencies throughout the reconstruction process. By incorporating the RASFormer block, our RAS method achieves superior contextual awareness capabilities, leading to remarkable performance. Experimental results show that our RAS significantly outperforms competing methods, well demonstrating the effectiveness and superiority of our method. Our code is available at https://github.com/Nothingtolose9979/RAS
Hui-Yue Yang, Hui Chen 0013, Zijia Lin, Kai Chen 0044, Jungong Han, Guiguang Ding
ECAI8
2024 Quantized Prompt for Efficient Generalization of Vision-Language Models
Tianxiang Hao 0001, Xiaohan Ding, Juexiao Feng, Yuhong Yang 0008, Hui Chen 0013, Guiguang Ding
ECCV (19)6
2024 Learn from the Learnt: Source-Free Active Domain Adaptation via Contrastive Sampling and Visual Persistence
Mengyao Lyu, Tianxiang Hao 0001, Hui Chen 0013, Zijia Lin, Jungong Han, Guiguang Ding
ECCV (1)7
2024 VCP-CLIP: A Visual Context Prompting Model for Zero-Shot Anomaly Segmentation
Zhen Qu, Xian Tao, Mukesh Prasad, Fei Shen 0002, Zhengtao Zhang, Xinyi Gong, Guiguang Ding
ECCV (69)7
2024 PYRA: Parallel Yielding Re-activation for Training-Inference Efficient Task Adaptation
Yizhe Xiong, Hui Chen 0013, Tianxiang Hao 0001, Zijia Lin, Jungong Han, Yuesong Zhang, Yongjun Bao, Guiguang Ding
ECCV (9)9
2024 Balanced Active Sampling for Person Re-identification
abstract
Active learning is attracting more and more attention in person re-identification (Re-ID), as it is promising in the scalability of Re-ID models to satisfy performance with reduced labeling cost. Active sampling of pair-wise images in Re-ID is a highly imbalanced problem, where negative pairs are the vast majority. To avoid sampled pairs being dominated by the negative relationship, previous works tend to sample pairs with confident positive relationships in various ways. However, it is a waste of the labeling budget as most sampled pairs will be positive and already have a very close distance. Thus, there is no significant improvement in the model performance. In this paper, we first argue that balanced sampling is the key to active learning for Re-ID. Along this line, we propose a naïve balanced sampling method based on the global estimation of the most confusing distance. It is further improved by the label-wise estimation and diversity measurement. We also formulate the training of Re-ID models as a constrained clustering problem, where labeled positive and negative pairs are as must-link and cannot-link. Then the model training is based on the pseudo labels. Extensive experiments on benchmarks evaluate the effectiveness and superiority of the proposed methods. Specifically, it achieves comparable performance with supervised counterparts with less than 0.1% pair-wise annotation, which significantly surpasses the state-of-the-art.
Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005
ICME3
2024 Camera Bias Regularization for Person Re-identification
abstract
Person re-identification (Re-ID) is to match persons captured by non-overlapping cameras. Due to the discrepancies between cameras caused by illumination, background, or viewpoint, the underlying difficulty for Re-ID is the camera bias problem, which leads to the large gap of within-identity features from different cameras. With limited cross-camera annotation, Re-ID models tend to learn camera-related features, instead of identity-related features. Consequently, Re-ID models suffer from poor transfer ability from seen to unseen domains. In this paper, we investigate the camera bias problem in both supervised and unsupervised learning. In particular, we propose a novel Camera Bias Regularization (CBR) term to reduce the feature distribution gap between cameras. The CBR works by simultaneously enlarging the distance of intra-camera distributions between positive and negative pairs, and reducing the distance of positive pairs’ distributions between intra-camera and cross-camera. In addition, a Cross-Camera (CC) clustering method is also designed for unsupervised learning, which puts more emphasis on cross-camera pairs than intra-camera ones during the clustering process. Extensive experiments are conducted to validate the effectiveness of the proposed CBR and CC. Specifically, with only a plain ResNet-50, it achieves 56.7% mAP and 40.7% mAP on the challenging MSMT17 dataset in supervised and unsupervised settings respectively, which surpasses most state-of-the-arts.
Leqi Shen, Guiguang Ding, Zhiheng Zhou 0001, Tianshi Xu, Xiaofeng Jin, Yuheng Huang 0005
ICME3
2024 X-ReID: Cross-Instance Transformer for Identity-Level Person Re-Identification
abstract
Currently, most existing person re-identification methods use instance-level features, which are extracted only from a single image. However, these instance-level features can easily ignore the discriminative information because the appearance of each identity varies greatly in different images. Thus, it is necessary to exploit identity-level features, which can be shared across different images of each identity. In this paper, we propose a novel training framework, named X-ReID, to promote instance-level features to identity-level features by employing cross-attention to incorporate information from one image to another of the same identity, thus more unified and discriminative pedestrian information can be obtained. Extensive experiments on benchmark datasets show the superiority of our method over existing works. Particularly, on the challenging MSMT17, our proposed method gains 1.1% mAP improvements when compared to the second place.
Leqi Shen, Sicheng Zhao, Zhelun Shen, Tianshi Xu, Guiguang Ding
ICME7
2024 TaD: A Plug-and-Play Task-Aware Decoding Method to Better Adapt LLMs on Downstream Tasks
Hui Chen 0013, Zijia Lin, Jungong Han, Lixing Gong, Yongjun Bao, Guiguang Ding
IJCAI8
2024 More is Better: Deep Domain Adaptation with Multiple Sources
Sicheng Zhao, Hui Chen 0013, Hu Huang 0009, Pengfei Xu 0013, Guiguang Ding
IJCAI5
2024 Multi-Label Learning with Block Diagonal Labels
abstract
Collecting large-scale multi-label data with full labels is difficult for real-world scenarios. Many existing studies have tried to address the issue of missing labels caused by annotation but ignored the difficulties encountered during the annotation process. We find that the high annotation workload can be attributed to two reasons: (1) Annotators are required to identify labels on widely varying visual concepts. (2) Exhaustively annotating the entire dataset with all the labels becomes notably difficult and time-consuming. In this paper, we propose a new setting, i.e. block diagonal labels, to reduce the workload on both sides. The numerous categories can be divided into different subsets based on semantics and relevance. Each annotator can only focus on its own subset of labels so that only a small set of highly relevant labels are required to be annotated per image. To deal with the issue of such missing labels, we introduce a simple yet effective method that does not require any prior knowledge of the dataset. In practice, we propose an Adaptive Pseudo-Labeling method to predict the unknown labels with less noise. Formal analysis is conducted to evaluate the superiority of our setting. Extensive experiments are conducted to verify the effectiveness of our method on multiple widely used benchmarks.
Leqi Shen, Sicheng Zhao, Hui Chen 0013, Jundong Zhou, Pengzhang Liu, Yongjun Bao, Guiguang Ding
ACM Multimedia8
2024 YOLOv10: Real-Time End-to-End Object Detection
abstract
Over the past years, YOLOs have emerged as the predominant paradigm in the field of real-time object detection owing to their effective balance between computational cost and detection performance. Researchers have explored the architectural designs, optimization objectives, data augmentation strategies, and others for YOLOs, achieving notable progress. However, the reliance on the non-maximum suppression (NMS) for post-processing hampers the end-to-end deployment of YOLOs and adversely impacts the inference latency. Besides, the design of various components in YOLOs lacks the comprehensive and thorough inspection, resulting in noticeable computational redundancy and limiting the model's capability. It renders the suboptimal efficiency, along with considerable potential for performance improvements. In this work, we aim to further advance the performance-efficiency boundary of YOLOs from both the post-processing and the model architecture. To this end, we first present the consistent dual assignments for NMS-free training of YOLOs, which brings the competitive performance and low inference latency simultaneously. Moreover, we introduce the holistic efficiency-accuracy driven model design strategy for YOLOs. We comprehensively optimize various components of YOLOs from both the efficiency and accuracy perspectives, which greatly reduces the computational overhead and enhances the capability. The outcome of our effort is a new generation of YOLO series for real-time end-to-end object detection, dubbed YOLOv10. Extensive experiments show that YOLOv10 achieves the state-of-the-art performance and efficiency across various model scales. For example, our YOLOv10-S is 1.8$\times$ faster than RT-DETR-R18 under the similar AP on COCO, meanwhile enjoying 2.8$\times$ smaller number of parameters and FLOPs. Compared with YOLOv9-C, YOLOv10-B has 46\% less latency and 25\% fewer parameters for the same performance. Code and models are available at https://github.com/THU-MIG/yolov10.
Hui Chen 0013, Kai Chen 0044, Zijia Lin, Jungong Han, Guiguang Ding
NeurIPS7
2024 Revisiting motion information for RGB-Event tracking with MOT philosophy
abstract
RGB-Event single object tracking (SOT) aims to leverage the merits of RGB and event data to achieve higher performance. However, existing frameworks focus on exploring complementary appearance information within multi-modal data, and struggle to address the association problem of targets and distractors in the temporal domain using motion information from the event stream. In this paper, we introduce the Multi-Object Tracking (MOT) philosophy into RGB-E SOT to keep track of targets as well as distractors by using both RGB and event data, thereby improving the robustness of the tracker. Specifically, an appearance model is employed to predict the initial candidates. Subsequently, the initially predicted tracking results, in combination with the RGB-E features, are encoded into appearance and motion embeddings, respectively. Furthermore, a Spatial-Temporal Transformer Encoder is proposed to model the spatial-temporal relationships and learn discriminative features for each candidate through guidance of the appearance-motion embeddings. Simultaneously, a Dual-Branch Transformer Decoder is designed to adopt such motion and appearance information for candidate matching, thus distinguishing between targets and distractors. The proposed method is evaluated on multiple benchmark datasets and achieves state-of-the-art performance on all the datasets tested.
Tianlu Zhang, Kurt Debattista, Qiang Zhang 0020, Guiguang Ding, Jungong Han
NeurIPS4
2024 Multi-source-free Domain Adaptive Object Detection
Sicheng Zhao, Huizai Yao, Chuang Lin 0003, Yue Gao 0002, Guiguang Ding
Int. J. Comput. Vis.5
2024 Confidence-Guided Centroids for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (ReID) aims to train a feature extractor for identity retrieval without exploiting identity labels. Due to the no-reference trust in imperfect clustering results, the learning is inevitably misled by unreliable pseudo labels. Albeit the pseudo label refinement has been investigated by previous works, they generally leverage auxiliary information such as camera IDs and body part predictions. This work explores the internal characteristics of clusters to refine pseudo labels. To this end, Confidence-Guided Centroids (CGC) are proposed to provide reliable cluster-wise prototypes for feature learning. Since samples with high confidence are exclusively involved in the formation of centroids, the identity information of low-confidence samples, i.e., boundary samples, are NOT likely to contribute to the corresponding centroid. Given the new centroids, the current learning scheme, where samples are forced to learn from their assigned centroids solely, is unwise. To remedy the situation, we propose to use Confidence-Guided pseudo Label (CGL), which enables samples to approach not only the originally assigned centroid but also other centroids that are potentially embedded with their identity information. Empowered by confidence-guided centroids and labels, our method yields comparable performance with, or even outperforms, state-of-the-art pseudo label refinement works that largely leverage auxiliary information.
Yunqi Miao, Jiankang Deng, Guiguang Ding, Jungong Han
IEEE Trans. Inf. Forensics Secur.3
2024 Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN
abstract
The existence of redundancy in convolutional neural networks (CNNs) enables us to remove some filters/channels with acceptable performance drops. However, the training objective of CNNs usually tends to minimize an accuracy-related loss function without any attention paid to the redundancy, making the redundancy distribute randomly on all the filters, such that removing any of them may trigger information loss and accuracy drop, necessitating a fine-tuning step for recovery. In this article, we propose to manipulate the redundancy during training to facilitate network pruning. To this end, we propose a novel centripetal SGD (C-SGD) to make some filters identical, resulting in ideal redundancy patterns, as such filters become purely redundant due to their duplicates, hence removing them does not harm the network. As shown on CIFAR and ImageNet, C-SGD delivers better performance because the redundancy is better organized, compared to the existing methods. The efficiency also characterizes C-SGD because it is as fast as regular SGD, requires no fine-tuning, and can be conducted simultaneously on all the layers even in very deep CNNs. Besides, C-SGD can improve the accuracy of CNNs by first training a model with the same architecture but wider layers and then squeezing it into the original width.
Tianxiang Hao 0001, Xiaohan Ding, Jungong Han, Guiguang Ding
IEEE Trans. Neural Networks Learn. Syst.5
2023 Box-Level Active Detection
abstract
Active learning selects informative samples for annotation within budget, which has proven efficient recently on object detection. However, the widely used active detection benchmarks conduct image-level evaluation, which is unrealistic in human workload estimation and biased towards crowded images. Furthermore, existing methods still perform image-level annotation, but equally scoring all targets within the same image incurs waste of budget and redundant labels. Having revealed above problems and limitations, we introduce a box-level active detection framework that controls a box-based budget per cycle, prioritizes informative targets and avoids redundancy for fair comparison and efficient application. Under the proposed box-level setting, we devise a novel pipeline, namely Complementary Pseudo Active Strategy (ComPAS). It exploits both human annotations and the model intelligence in a complementary fashion: an efficient input-end committee queries labels for informative objects only; meantime well-learned targets are identified by the model and compensated with pseudo-labels. ComPAS consistently outperforms 10 competitors under 4 settings in a unified codebase. With supervision from labeled data only, it achieves 100% supervised performance of VOC0712 with merely 19% box annotations. On the COCO dataset, it yields up to 4.3% mAP improvement over the second-best method. ComPAS also supports training with the unlabeled pool, where it surpasses 90% COCO supervised performance with 85% label reduction. Our source code is publicly available at https://github.com/lyumengyao/blad.
Mengyao Lyu, Jundong Zhou, Hui Chen 0013, Dongdong Yu, Yandong Guo, Liuyu Xiang, Guiguang Ding
CVPR10
2023 Confidence-based Visual Dispersal for Few-shot Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation aims to transfer knowledge from a fully-labeled source domain to an unlabeled target domain. However, in real-world scenarios, providing abundant labeled data even in the source domain can be infeasible due to the difficulty and high expense of annotation. To address this issue, recent works consider the Few-shot Unsupervised Domain Adaptation (FUDA) where only a few source samples are labeled, and conduct knowledge transfer via self-supervised learning methods. Yet existing methods generally overlook that the sparse label setting hinders learning reliable source knowledge for transfer. Additionally, the learning difficulty difference in target samples is different but ignored, leaving hard target samples poorly classified. To tackle both deficiencies, in this paper, we propose a novel Confidence-based Visual Dispersal Transfer learning method (C-VisDiT) for FUDA. Specifically, C-VisDiT consists of a cross-domain visual dispersal strategy that transfers only high-confidence source knowledge for model adaptation and an intra-domain visual dispersal strategy that guides the learning of hard target samples with easy ones. We conduct extensive experiments on Office-31, Office-Home, VisDA-C, and Domain- Net benchmark datasets and the results demonstrate that the proposed C-VisDiT significantly outperforms state-of- the-art FUDA methods. Our code is available at https://github.com/Bostoncake/C-VisDiT.
Yizhe Xiong, Hui Chen 0013, Zijia Lin, Sicheng Zhao, Guiguang Ding
ICCV5
2023 Re-parameterizing Your Optimizers rather than Architectures
Xiaohan Ding, Xiangyu Zhang 0005, Kaiqi Huang, Jungong Han, Guiguang Ding
ICLR6
2023 Consolidator: Mergable Adapter with Group Connections for Visual Adaptation
Tianxiang Hao 0001, Hui Chen 0013, Guiguang Ding
ICLR4
2023 Hierarchical Prompt Learning Using CLIP for Multi-label Classification with Single Positive Labels
abstract
Collecting full annotations to construct multi-label datasets is difficult and labor-consuming. As an effective solution to relieve the annotation burden, single positive multi-label learning (SPML) draws increasing attention from both academia and industry. It only annotates each image with one positive label, leaving other labels unobserved. Therefore, existing methods strive to explore the cue of unobserved labels to compensate for the insufficiency of label supervision. Though achieving promising performance, they generally consider labels independently, leaving out the inherent hierarchical semantic relationship among labels which reveals that labels can be clustered into groups. In this paper, we propose a hierarchical prompt learning method with a novel Hierarchical Semantic Prompt Network (HSPNet) to harness such hierarchical semantic relationships using a large-scale pretrained vision and language model, i.e., CLIP, for SPML. We first introduce a Hierarchical Conditional Prompt (HCP) strategy to grasp the hierarchical label-group dependency. Then we equip a Hierarchical Graph Convolutional Network (HGCN) to capture the high-order inter-label and inter-group dependencies. Comprehensive experiments and analyses on several benchmark datasets show that our method significantly outperforms the state-of-the-art methods, well demonstrating its superiority and effectiveness. Our code will be available at https://github.com/jameslahm/HSPNet.
Hui Chen 0013, Zijia Lin, Zixuan Ding, Pengzhang Liu, Yongjun Bao, Weipeng Yan, Guiguang Ding
ACM Multimedia8
2023 GPro3D: Deriving 3D BBox from ground plane in monocular 3D object detection
Fan Yang 0083, Hui Chen 0013, Guiguang Ding
Neurocomputing7
2023 Toward Label-Efficient Emotion and Sentiment Analysis
abstract
Emotion and sentiment play a central role in various human activities, such as perception, decision-making, social interaction, and logical reasoning. Developing artificial emotional intelligence (AEI) for machines is becoming a bottleneck in human–computer interaction. The first step of AEI is to recognize the emotion and sentiment that are conveyed in different affective signals. Traditional supervised emotion and sentiment analysis (ESA) methods, especially deep learning-based ones, usually require large-scale labeled training data. However, due to the essential subjectivity, complexity, uncertainty and ambiguity, and subtlety, collecting such annotations is expensive, time-consuming, and difficult in practice. In this article, we introduce label-efficient ESA from the computational perspective. First, we present a hierarchical taxonomy for label-efficient learning based on the availability of sample labels, emotion categories, and data domains during training. Second, for each of the seven paradigms, i.e., unsupervised, semisupervised, weakly supervised, low-shot, incremental, domain-adaptive, and domain-generalizable ESA, we give the definition, summarize existing methods, and present our views on the quantitative and qualitative comparison. Finally, we provide several promising real-world applications, followed by unsolved challenges and potential future directions.
Sicheng Zhao, Xiaopeng Hong, Jufeng Yang, Guiguang Ding
Proc. IEEE5
2023 Margin-aware rectified augmentation for long-tailed recognition
Liuyu Xiang, Jungong Han, Guiguang Ding
Pattern Recognit.3
2022 SECRET: Self-Consistent Pseudo Label Refinement for Unsupervised Domain Adaptive Person Re-identification
abstract
Unsupervised domain adaptive person re-identification aims at learning on an unlabeled target domain with only labeled data in source domain. Currently, the state-of-the-arts usually solve this problem by pseudo-label-based clustering and fine-tuning in target domain. However, the reason behind the noises of pseudo labels is not sufficiently explored, especially for the popular multi-branch models. We argue that the consistency between different feature spaces is the key to the pseudo labels’ quality. Then a SElf-Consistent pseudo label RefinEmenT method, termed as SECRET, is proposed to improve consistency by mutually refining the pseudo labels generated from different feature spaces. The proposed SECRET gradually encourages the improvement of pseudo labels’ quality during training process, which further leads to better cross-domain Re-ID performance. Extensive experiments on benchmark datasets show the superiority of our method. Specifically, our method outperforms the state-of-the-arts by 6.3% in terms of mAP on the challenging dataset MSMT17. In the purely unsupervised setting, our method also surpasses existing works by a large margin. Code is available at https://github.com/LunarShen/SECRET.
Leqi Shen, Guiguang Ding, Zhenhua Guo 0005
AAAI4
2022 ReMoNet: Recurrent Multi-Output Network for Efficient Video Denoising
abstract
While deep neural network-based video denoising methods have achieved promising results, it is still hard to deploy them on mobile devices due to their high computational cost and memory demands. This paper aims to develop a lightweight deep video denoising method that is friendly to resource-constrained mobile devices. Inspired by the facts that 1) consecutive video frames usually contain redundant temporal coherency, and 2) neural networks are usually over-parameterized, we propose a multi-input multi-output (MIMO) paradigm to process consecutive video frames within one-forward-pass. The basic idea is concretized to a novel architecture termed Recurrent Multi-output Network (ReMoNet), which consists of recurrent temporal fusion and temporal aggregation blocks and is further reinforced by similarity-based mutual distillation. We conduct extensive experiments on NVIDIA GPU and Qualcomm Snapdragon 888 mobile platform with Gaussian noise and simulated Image-Signal-Processor (ISP) noise. The experimental results show that ReMoNet is both effective and efficient on video denoising. Moreover, we show that ReMoNet is more robust under higher noise level scenarios.
Liuyu Xiang, Jundong Zhou, Jirui Liu, Zerun Wang, Haidong Huang, Jie Hu 0021, Jungong Han, Guiguang Ding
AAAI9
2022 Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNs
abstract
We revisit large kernel design in modern convolutional neural networks (CNNs). Inspired by recent advances in vision transformers (ViTs), in this paper, we demonstrate that using a few large convolutional kernels instead of a stack of small kernels could be a more powerful paradigm. We suggested five guidelines, e.g., applying re-parameterized large depthwise convolutions, to design efficient high-performance large-kernel CNNs. Following the guidelines, we propose RepLKNet, a pure CNN architecture whose kernel size is as large as 31×31, in contrast to commonly used 3×3. RepLKNet greatly closes the performance gap between CNNs and ViTs, e.g., achieving comparable or superior results than Swin Transformer on ImageNet and a few typical downstream tasks, with lower latency. RepLKNet also shows nice scalability to big data and large models, obtaining 87.8% top-1 accuracy on ImageNet and 56.0% mIoU on ADE20K, which is very competitive among the state-of-the-arts with similar model sizes. Our study further reveals that, in contrast to small-kernel CNNs, large-kernel CNNs have much larger effective receptive fields and higher shape bias rather than texture bias. Code & models at https://github.com/megvii-research/RepLKNet.
Xiaohan Ding, Xiangyu Zhang 0005, Jungong Han, Guiguang Ding
CVPR4
2022 RepMLPNet: Hierarchical Vision MLP with Re-parameterized Locality
abstract
Compared to convolutional layers, fully-connected (FC) layers are better at modeling the long-range dependencies but worse at capturing the local patterns, hence usually less favored for image recognition. In this paper, we propose a methodology, Locality Injection, to incorporate local priors into an FC layer via merging the trained parameters of a parallel conv kernel into the FC kernel. Locality Injection can be viewed as a novel Structural Re-parameterization method since it equivalently converts the structures via transforming the parameters. Based on that, we propose a multi-layer-perceptron (MLP) block named RepMLP Block, which uses three FC layers to extract features, and a novel architecture named RepMLPNet. The hierarchical design distinguishes RepMLPNet from the other concurrently proposed vision MLPs. As it produces feature maps of different levels, it qualifies as a backbone model for downstream tasks like semantic segmentation. Our results reveal that 1) Locality Injection is a general methodology for MLP models; 2) RepMLPNet has favorable accuracy-efficiency trade-off compared to the other MLPs; 3) RepMLPNet is the first MLP that seamlessly transfer to Cityscapes semantic segmentation. The code and models are available at https://github.com/DingXiaoH/RepMLP.
Xiaohan Ding, Xiangyu Zhang 0005, Jungong Han, Guiguang Ding
CVPR5
2022 TAGPerson: A Target-Aware Generation Pipeline for Person Re-identification
abstract
Nowadays, real data in person re-identification (ReID) task is facing privacy issues, e.g., the banned dataset DukeMTMC-ReID. Thus it becomes much harder to collect real data for ReID task. Meanwhile, the labor cost of labeling ReID data is still very high and further hinders the development of the ReID research. Therefore, many methods turn to generate synthetic images for ReID algorithms as alternatives instead of real images. However, there is an inevitable domain gap between synthetic and real images. In previous methods, the generation process is based on virtual scenes, and their synthetic training data can not be changed according to different target real scenes automatically. To handle this problem, we propose a novel Target-Aware Generation pipeline to produce synthetic person images, called TAGPerson. Specifically, it involves a parameterized rendering method, where the parameters are controllable and can be adjusted according to the target scenes. In TAGPerson, we extract information from target scenes and use them to control our parameterized rendering process to generate target-aware synthetic images, which would hold a smaller gap to the real images in the specific target domain. In our experiments, our target-aware synthetic images can achieve a much higher performance than the generalized synthetic images on MSMT17, i.e. 47.5% vs. 40.9% for rank-1 accuracy. We will release this toolkit for the ReID community to generate synthetic images at any desired taste. The code is available at: https://github.com/tagperson/tagperson-blender
Kai Chen 0044, Fan Wang 0019, Xiuyu Sun, Guiguang Ding
ACM Multimedia8
2022 Bidirectional difference locating and semantic consistency reasoning for change captioning
abstract
Change captioning is an emerging task to describe the changes between a pair of images. The difficulty in this task is to discover the differences between the two images. Recently, some methods have been proposed to address this problem. However, they all employ unidirectional difference localization to identify the changes. This can lead to ambiguity about the nature of the changes. Instead, we propose a framework with bidirectional difference localization and semantic consistency reasoning to describe the image changes. First, we locate the changes in the two images by capturing bidirectional differences. Then we design a decoder with spatial-channel attention to generate the change caption. Finally, we introduce semantic consistency reasoning to constrain our bidirectional difference localization module and spatial-channel attention module. Extensive experiments on three public data sets show that the performance of our proposed model outperforms the state-of-the-art change captioning models by a large margin.
Yaoqi Sun, Liang Li 0003, Tongyv Lu, Bolun Zheng, Chenggang Yan 0001, Yongjun Bao, Guiguang Ding, Gregory Slabaugh
Int. J. Intell. Syst.9
2022 Towards real-time object detection in GigaPixel-level video
Kai Chen 0044, Zerun Wang, Dahan Gong, Longlong Yu 0001, Guiguang Ding
Neurocomputing7
2022 Long-tailed visual recognition with deep models: A methodological survey and evaluation
Yu Fu 0006, Liuyu Xiang, Yumna Zahid, Guiguang Ding, Tao Mei 0001, Qiang Shen 0001, Jungong Han
Neurocomputing4
2022 Affective Image Content Analysis: Two Decades Review and New Perspectives
abstract
Images can convey rich semantics and induce various emotions in viewers. Recently, with the rapid advancement of emotional intelligence and the explosive growth of visual data, extensive research efforts have been dedicated to affective image content analysis (AICA). In this survey, we will comprehensively review the development of AICA in the recent two decades, especially focusing on the state-of-the-art methods with respect to three main challenges - the affective gap, perception subjectivity, and label noise and absence. We begin with an introduction to the key emotion representation models that have been widely employed in AICA and description of available datasets for performing evaluation with quantitative comparison of label noise and dataset bias. We then summarize and compare the representative approaches on (1) emotion feature extraction, including both handcrafted and deep features, (2) learning methods on dominant emotion recognition, personalized emotion prediction, emotion distribution learning, and learning from noisy data or few labels, and (3) AICA based applications. Finally, we discuss some challenges and promising research directions in the future, such as image content and context understanding, group emotion clustering, and viewer-image interaction.
Sicheng Zhao, Xingxu Yao, Jufeng Yang, Guoli Jia, Guiguang Ding, Tat-Seng Chua, Björn W. Schuller, Kurt Keutzer
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Emotional Semantics-Preserved and Feature-Aligned CycleGAN for Visual Emotion Adaptation
abstract
Thanks to large-scale labeled training data, deep neural networks (DNNs) have obtained remarkable success in many vision and multimedia tasks. However, because of the presence of domain shift, the learned knowledge of the well-trained DNNs cannot be well generalized to new domains or datasets that have few labels. Unsupervised domain adaptation (UDA) studies the problem of transferring models trained on one labeled source domain to another unlabeled target domain. In this article, we focus on UDA in visual emotion analysis for both emotion distribution learning and dominant emotion classification. Specifically, we design a novel end-to-end cycle-consistent adversarial model, called CycleEmotionGAN++. First, we generate an adapted domain to align the source and target domains on the pixel level by improving CycleGAN with a multiscale structured cycle-consistency loss. During the image translation, we propose a dynamic emotional semantic consistency loss to preserve the emotion labels of the source images. Second, we train a transferable task classifier on the adapted domain with feature-level alignment between the adapted and target domains. We conduct extensive UDA experiments on the Flickr-LDL and Twitter-LDL datasets for distribution learning and ArtPhoto and Flickr and Instagram datasets for emotion classification. The results demonstrate the significant improvements yielded by the proposed CycleEmotionGAN++ compared to state-of-the-art UDA approaches.
Sicheng Zhao, Xuanbai Chen, Xiangyu Yue 0001, Chuang Lin 0003, Pengfei Xu 0013, Ravi Krishna, Jufeng Yang, Guiguang Ding, Alberto L. Sangiovanni-Vincentelli, Kurt Keutzer
IEEE Trans. Cybern.8
2022 Personalized Image Aesthetics Assessment via Meta-Learning With Bilevel Gradient Optimization
abstract
Typical image aesthetics assessment (IAA) is modeled for the generic aesthetics perceived by an "average" user. However, such generic aesthetics models neglect the fact that users' aesthetic preferences vary significantly depending on their unique preferences. Therefore, it is essential to tackle the issue for personalized IAA (PIAA). Since PIAA is a typical small sample learning (SSL) problem, existing PIAA models are usually built by fine-tuning the well-established generic IAA (GIAA) models, which are regarded as prior knowledge. Nevertheless, this kind of prior knowledge based on "average aesthetics" fails to incarnate the aesthetic diversity of different people. In order to learn the shared prior knowledge when different people judge aesthetics, that is, learn how people judge image aesthetics, we propose a PIAA method based on meta-learning with bilevel gradient optimization (BLG-PIAA), which is trained using individual aesthetic data directly and generalizes to unknown users quickly. The proposed approach consists of two phases: 1) meta-training and 2) meta-testing. In meta-training, the aesthetics assessment of each user is regarded as a task, and the training set of each task is divided into two sets: 1) support set and 2) query set. Unlike traditional methods that train a GIAA model based on average aesthetics, we train an aesthetic meta-learner model by bilevel gradient updating from the support set to the query set using many users' PIAA tasks. In meta-testing, the aesthetic meta-learner model is fine-tuned using a small amount of aesthetic data of a target user to obtain the PIAA model. The experimental results show that the proposed method outperforms the state-of-the-art PIAA metrics, and the learned prior model of BLG-PIAA can be quickly adapted to unseen PIAA tasks.
Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, Guangming Shi
IEEE Trans. Cybern.5
2021 Automated Model Design and Benchmarking of Deep Learning Models for COVID-19 Detection with Chest CT Scans
abstract
The COVID-19 pandemic has spread globally for several months. Because its transmissibility and high pathogenicity seriously threaten people's lives, it is crucial to accurately and quickly detect COVID-19 infection. Many recent studies have shown that deep learning (DL) based solutions can help detect COVID-19 based on chest CT scans. However, most existing work focuses on 2D datasets, which may result in low quality models as the real CT scans are 3D images. Besides, the reported results span a broad spectrum on different datasets with a relatively unfair comparison. In this paper, we first use three state-of-the-art 3D models (ResNet3D101, DenseNet3D121, and MC3\_18) to establish the baseline performance on three publicly available chest CT scan datasets. Then we propose a differentiable neural architecture search (DNAS) framework to automatically search the 3D DL models for 3D chest CT scans classification and use the Gumbel Softmax technique to improve the search efficiency. We further exploit the Class Activation Mapping (CAM) technique on our models to provide the interpretability of the results. The experimental results show that our searched models (CovidNet3D) outperform the baseline human-designed models on three datasets with tens of times smaller model size and higher accuracy. Furthermore, the results also verify that CAM can be well applied in CovidNet3D for COVID-19 datasets to provide interpretability for medical diagnosis. Code: https://github.com/HKBU-HPML/CovidNet3D.
Xin He 0019, Xiaowen Chu 0001, Shaohuai Shi, Jiangping Tang, Xin Liu 0027, Chenggang Yan 0001, Jiyong Zhang 0001, Guiguang Ding
AAAI9
2021 Diverse Branch Block: Building a Convolution as an Inception-Like Unit
abstract
We propose a universal building block of Convolutional Neural Network (ConvNet) to improve the performance without any inference-time costs. The block is named Diverse Branch Block (DBB), which enhances the representational capacity of a single convolution by combining diverse branches of different scales and complexities to enrich the feature space, including sequences of convolutions, multiscale convolutions, and average pooling. After training, a DBB can be equivalently converted into a single conv layer for deployment. Unlike the advancements of novel ConvNet architectures, DBB complicates the training-time microstructure while maintaining the macro architecture, so that it can be used as a drop-in replacement for regular conv layers of any architecture. In this way, the model can be trained to reach a higher level of performance and then transformed into the original inference-time structure for inference. DBB improves ConvNets on image classification (up to 1.9% higher top-1 accuracy on ImageNet), object detection and semantic segmentation. The PyTorch code and models are released at https://github.com/DingXiaoH/DiverseBranchBlock.
Xiaohan Ding, Xiangyu Zhang 0005, Jungong Han, Guiguang Ding
CVPR4
2021 RepVGG: Making VGG-Style ConvNets Great Again
abstract
We present a simple but powerful architecture of convolutional neural network, which has a VGG-like inference-time body composed of nothing but a stack of 3 × 3 convolution and ReLU, while the training-time model has a multi-branch topology. Such decoupling of the training-time and inference-time architecture is realized by a structural re-parameterization technique so that the model is named RepVGG. On ImageNet, RepVGG reaches over 80% top-1 accuracy, which is the first time for a plain model, to the best of our knowledge. On NVIDIA 1080Ti GPU, RepVGG models run 83% faster than ResNet-50 or 101% faster than ResNet-101 with higher accuracy and show favorable accuracy-speed trade-off compared to the state-of-the-art models like EfficientNet and RegNet. The code and trained models are available at https://github.com/megvii-model/RepVGG.
Xiaohan Ding, Xiangyu Zhang 0005, Ningning Ma, Jungong Han, Guiguang Ding, Jian Sun 0001
CVPR5
2021 ResRep: Lossless CNN Pruning via Decoupling Remembering and Forgetting
abstract
We propose ResRep, a novel method for lossless channel pruning (a.k.a. filter pruning), which slims down a CNN by reducing the width (number of output channels) of convolutional layers. Inspired by the neurobiology research about the independence of remembering and forgetting, we propose to re-parameterize a CNN into the remembering parts and forgetting parts, where the former learn to maintain the performance and the latter learn to prune. Via training with regular SGD on the former but a novel update rule with penalty gradients on the latter, we realize structured sparsity. Then we equivalently merge the remembering and forgetting parts into the original architecture with narrower layers. In this sense, ResRep can be viewed as a successful application of Structural Re-parameterization. Such a methodology distinguishes ResRep from the traditional learning-based pruning paradigm that applies a penalty on parameters to produce sparsity, which may suppress the parameters essential for the remembering. ResRep slims down a standard ResNet-50 with 76.15% accuracy on ImageNet to a narrower one with only 45% FLOPs and no accuracy drop, which is the first to achieve lossless pruning with such a high compression ratio. The code and models are at https://github.com/DingXiaoH/ResRep.
Xiaohan Ding, Tianxiang Hao 0001, Jianchao Tan, Ji Liu 0002, Jungong Han, Guiguang Ding
ICCV7
2021 MADAN: Multi-source Adversarial Domain Aggregation Network for Domain Adaptation
Sicheng Zhao, Bo Li 0080, Pengfei Xu 0013, Xiangyu Yue 0001, Guiguang Ding, Kurt Keutzer
Int. J. Comput. Vis.5
2021 Deep image compression with multi-stage representation
Guiguang Ding, Jungong Han, Fan Li 0003
J. Vis. Commun. Image Represent.2
2021 Where to Prune: Using LSTM to Guide Data-Dependent Soft Pruning
abstract
While convolutional neural network (CNN) has achieved overwhelming success in various vision tasks, its heavy computational cost and storage overhead limit the practical use on mobile or embedded devices. Recently, compressing CNN models has attracted considerable attention, where pruning CNN filters, also known as the channel pruning, has generated great research popularity due to its high compression rate. In this paper, a new channel pruning framework is proposed, which can significantly reduce the computational complexity while maintaining sufficient model accuracy. Unlike most existing approaches that seek to-be-pruned filters layer by layer, we argue that choosing appropriate layers for pruning is more crucial, which can result in more complexity reduction but less performance drop. To this end, we utilize a long short-term memory (LSTM) to learn the hierarchical characteristics of a network and generate a global network pruning scheme. On top of it, we propose a data-dependent soft pruning method, dubbed Squeeze-Excitation-Pruning (SEP), which does not physically prune any filters but selectively excludes some kernels involved in calculating forward and backward propagations depending on the pruning scheme. Compared with the hard pruning, our soft pruning can better retain the capacity and knowledge of the baseline model. Experimental results demonstrate that our approach still achieves comparable accuracy even when reducing 70.1% Floating-point operation per second (FLOPs) for VGG and 47.5% for Resnet-56.
Guiguang Ding, Zizhou Jia, Jungong Han
IEEE Trans. Image Process.1
2021 Learning Transformation-Invariant Local Descriptors With Low-Coupling Binary Codes
abstract
Despite the great success achieved by prevailing binary local descriptors, they are still suffering from two problems: 1) vulnerable to the geometric transformations; 2) lack of an effective treatment to the highly-correlated bits that are generated by directly applying the scheme of image hashing. To tackle both limitations, we propose an unsupervised Transformation-invariant Binary Local Descriptor learning method (TBLD). Specifically, the transformation invariance of binary local descriptors is ensured by projecting the original patches and their transformed counterparts into an identical high-dimensional feature space and an identical low-dimensional descriptor space simultaneously. Meanwhile, it enforces the dissimilar image patches to have distinctive binary local descriptors. Moreover, to reduce high correlations between bits, we propose a bottom-up learning strategy, termed Adversarial Constraint Module, where low-coupling binary codes are introduced externally to guide the learning of binary local descriptors. With the aid of the Wasserstein loss, the framework is optimized to encourage the distribution of the generated binary local descriptors to mimic that of the introduced low-coupling binary codes, eventually making the former more low-coupling. Experimental results on three benchmark datasets well demonstrate the superiority of the proposed method over the state-of-the-art methods. The project page is available at https://github.com/yoqim/TBLD.
Yunqi Miao, Zijia Lin, Xiao Ma 0013, Guiguang Ding, Jungong Han
IEEE Trans. Image Process.4
2021 Dynamic Selective Network for RGB-D Salient Object Detection
abstract
RGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models.
Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding
IEEE Trans. Image Process.9
2020 Heterogeneous Transfer Learning with Weighted Instance-Correspondence Data
abstract
Instance-correspondence (IC) data are potent resources for heterogeneous transfer learning (HeTL) due to the capability of bridging the source and the target domains at the instance-level. To this end, people tend to use machine-generated IC data, because manually establishing IC data is expensive and primitive. However, existing IC data machine generators are not perfect and always produce the data that are not of high quality, thus hampering the performance of domain adaption. In this paper, instead of improving the IC data generator, which might not be an optimal way, we accept the fact that data quality variation does exist but find a better way to use the data. Specifically, we propose a novel heterogeneous transfer learning method named Transfer Learning with Weighted Correspondence (TLWC), which utilizes IC data to adapt the source domain to the target domain. Rather than treating IC data equally, TLWC can assign solid weights to each IC data pair depending on the quality of the data. We conduct extensive experiments on HeTL datasets and the state-of-the-art results verify the effectiveness of TLWC.
Xiaoming Jin, Guiguang Ding, Jungong Han, Jiyong Zhang 0001, Sicheng Zhao
AAAI3
2020 Shallow Feature Based Dense Attention Network for Crowd Counting
abstract
While the performance of crowd counting via deep learning has been improved dramatically in the recent years, it remains an ingrained problem due to cluttered backgrounds and varying scales of people within an image. In this paper, we propose a Shallow feature based Dense Attention Network (SDANet) for crowd counting from still images, which diminishes the impact of backgrounds via involving a shallow feature based attention model, and meanwhile, captures multi-scale information via densely connecting hierarchical image features. Specifically, inspired by the observation that backgrounds and human crowds generally have noticeably different responses in shallow features, we decide to build our attention model upon shallow-feature maps, which results in accurate background-pixel detection. Moreover, considering that the most representative features of people across different scales can appear in different layers of a feature extraction network, to better keep them all, we propose to densely connect hierarchical image features of different layers and subsequently encode them for estimating crowd density. Experimental results on three benchmark datasets clearly demonstrate the superiority of SDANet when dealing with different scenarios. Particularly, on the challenging UCF_CC_50 dataset, our method outperforms other existing methods by a large margin, as is evident from a remarkable 11.9% Mean Absolute Error (MAE) drop of our SDANet.
Yunqi Miao, Zijia Lin, Guiguang Ding, Jungong Han
AAAI3
2020 IMRAM: Iterative Matching With Recurrent Attention Memory for Cross-Modal Image-Text Retrieval
abstract
Enabling bi-directional retrieval of images and texts is important for understanding the correspondence between vision and language. Existing methods leverage the attention mechanism to explore such correspondence in a fine-grained manner. However, most of them consider all semantics equally and thus align them uniformly, regardless of their diverse complexities. In fact, semantics are diverse (i.e. involving different kinds of semantic concepts), and humans usually follow a latent structure to combine them into understandable languages. It may be difficult to optimally capture such sophisticated correspondences in existing methods. In this paper, to address such a deficiency, we propose an Iterative Matching with Recurrent Attention Memory (IMRAM) method, in which correspondences between images and texts are captured with multiple steps of alignments. Specifically, we introduce an iterative matching scheme to explore such fine-grained correspondence progressively. A memory distillation unit is used to refine alignment knowledge from early steps to later ones. Experiment results on three benchmark datasets, i.e. Flickr8K, Flickr30K, and MS COCO, show that our IMRAM achieves state-of-the-art performance, well demonstrating its effectiveness. Experiments on a practical business advertisement dataset, named KWAI-AD, further validates the applicability of our method in practical scenarios.
Hui Chen 0013, Guiguang Ding, Zijia Lin, Ji Liu 0002, Jungong Han
CVPR2
2020 PANDA: A Gigapixel-Level Human-Centric Video Dataset
abstract
We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide field-of-view (~1 square kilometer area) and high-resolution details (~gigapixel-level/frame). The scenes may contain 4k head counts with over 100× scale variation. PANDA provides enriched and hierarchical ground-truth annotations, including 15,974.6k bounding boxes, 111.8k fine-grained attribute labels, 12.7k trajectories, 2.2k groups and 2.9k interactions. We benchmark the human detection and tracking tasks. Due to the vast variance of pedestrian pose, scale, occlusion and trajectory, existing approaches are challenged by both accuracy and efficiency. Given the uniqueness of PANDA with both wide FoV and high resolution, a new task of interaction-aware group detection is introduced. We design a `global-to-local zoom-in' framework, where global trajectories and local interactions are simultaneously encoded, yielding promising results. We believe PANDA will contribute to the community of artificial intelligence and praxeology by understanding human behaviors and interactions in large-scale real-world scenes. PANDA Website: http://www.panda-dataset.com.
Xiya Zhang, Yinheng Zhu, Xiaoyun Yuan, Liuyu Xiang, Zerun Wang, Guiguang Ding, David J. Brady, Qionghai Dai, Lu Fang 0001
CVPR8
2020 Learning From Multiple Experts: Self-paced Knowledge Distillation for Long-Tailed Classification
Liuyu Xiang, Guiguang Ding, Jungong Han
ECCV (5)2
2020 Joint Optimization in Cached-Enabled Heterogeneous Network for Efficient Industrial IoT
abstract
In the era of industrial 4.0, industrial Internet of Things (IIoT) has brought essential changes to human society. For IIoT, communication in network can be defined as the basic condition for further development and integrated information exchange. In this way, cached-enabled heterogeneous industrial network is necessary to be optimized. In this paper, we consider the optimal geographical placement of contents in cache-enabled heterogeneous networks to minimize the total missing probability. And the probability represents that typical user cannot find requested file in the nearby base stations (BSs). In contract to existing works which only concern content placement, we jointly optimize content placement at BSs and activation densities of BSs of different tiers subject to the cache size limits and the constraint on the BSs energy consumption cost. In addition, the user distribution in this work is modeled by a homogeneous Poisson Point Process. We prove that the original optimization problem can be transformed to a convex problem. The convexity of the optimization problem allows us to apply the KKT conditions to derive useful analytical results of the optimal solution. Based on this, we propose a low-complexity near-optimal algorithm to find the approximated content placement probabilities. We further extend the optimization to heterogeneous networks with the user distribution modeled by the modified Cluster Process. Extensive simulation results show the superior performance of joint optimization of content placement and BSs activation densities compared to only optimizing content placement.
Chaofan Ma, Bin Jiang 0003, Guiguang Ding, Gan Zheng 0001, Huihui Wang 0001
IEEE J. Sel. Areas Commun.4
2020 Deep Transfer Learning for Image Emotion Analysis: Reducing Marginal and Joint Distribution Discrepancies Together
abstract
A lot of research attentions have been paid to image emotion analysis in recent years. Meanwhile, as convolutional neural networks (CNNs) have made great successful in computer vision, many researchers start to employ CNN to discriminate image emotions. However, the training procedure of CNNs depends on sufficient labeled data. Therefore, a CNN is hard to perform well in an image domain with scant labeled information. In this paper, we propose a deep transfer learning method for image emotion analysis. The method can leverage rich emotion knowledge from a source domain to the target domain. Our method reduces both marginal and joint domain distribution discrepancies at fully-connected layers. Through this way, we can effectively extract more transferable features and advance the performance of CNNs on poor-label emotion-image domains.
Guiguang Ding
Neural Process. Lett.2
2020 Correction to: Deep Transfer Learning for Image Emotion Analysis: Reducing Marginal and Joint Distribution Discrepancies Together
Guiguang Ding
Neural Process. Lett.2
2020 Discrete Probability Distribution Prediction of Image Emotions with Shared Sparse Learning
abstract
Computationally modelling the affective content of images has been extensively studied recently because of its wide applications in entertainment, advertisement, and education. Significant progress has been made on designing discriminative features to bridge the affective gap. Assuming that viewers can reach a consensus on the emotion of images, most existing works focused on assigning the dominant emotion category or the average dimension values to an image. However, the image emotions perceived by viewers are subjective by nature with the influence of personal and situational factors. In this paper, we propose a novel machine learning approach that characterizes the categorical image emotions as a discrete probability distribution (DPD). To associate emotion with the visual features extracted from images, we present shared sparse learning to learn the combination coefficients, with which the DPD of an unseen image is predicted by linearly combining the DPDs of the training images. Furthermore, we extend our method to the setup where multi-features are available and learn the optimal weights for each feature to reflect the importance of different features. Extensive experiments are carried out on Abstract, Emotion6 and IESN datasets and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Xin Zhao 0020, Youbao Tang, Jungong Han, Hongxun Yao, Qingming Huang
IEEE Trans. Affect. Comput.2
2020 Personality-Assisted Multi-Task Learning for Generic and Personalized Image Aesthetics Assessment
abstract
Traditional image aesthetics assessment (IAA) approaches mainly predict the average aesthetic score of an image. However, people tend to have different tastes on image aesthetics, which is mainly determined by their subjective preferences. As an important subjective trait, personality is believed to be a key factor in modeling individual's subjective preference. In this paper, we present a personality-assisted multi-task deep learning framework for both generic and personalized image aesthetics assessment. The proposed framework comprises two stages. In the first stage, a multi-task learning network with shared weights is proposed to predict the aesthetics distribution of an image and Big-Five (BF) personality traits of people who like the image. The generic aesthetics score of the image can be generated based on the predicted aesthetics distribution. In order to capture the common representation of generic image aesthetics and people's personality traits, a Siamese network is trained using aesthetics data and personality data jointly. In the second stage, based on the predicted personality traits and generic aesthetics of an image, an inter-task fusion is introduced to generate individual's personalized aesthetic scores on the image. The performance of the proposed method is evaluated using two public image aesthetics databases. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts in both generic and personalized IAA tasks.
Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Weisi Lin
IEEE Trans. Image Process.4
2020 On Aggregation of Unsupervised Deep Binary Descriptor With Weak Bits
abstract
Despite the thrilling success achieved by existing binary descriptors, most of them are still in the mire of three limitations: 1) vulnerable to the geometric transformations; 2) incapable of preserving the manifold structure when learning binary codes; 3) NO guarantee to find the true match if multiple candidates happen to have the same Hamming distance to a given query. All these together make the binary descriptor less effective, given large-scale visual recognition tasks. In this paper, we propose a novel learning-based feature descriptor, namely Unsupervised Deep Binary Descriptor (UDBD), which learns transformation invariant binary descriptors via projecting the original data and their transformed sets into a joint binary space. Moreover, we involve a ℓ2,1-norm loss term in the binary embedding process to gain simultaneously the robustness against data noises and less probability of mistakenly flipping bits of the binary descriptor, on top of it, a graph constraint is used to preserve the original manifold structure in the binary space. Furthermore, a weak bit mechanism is adopted to find the real match from candidates sharing the same minimum Hamming distance, thus enhancing matching performance. Extensive experimental results on public datasets show the superiority of UDBD in terms of matching and retrieval accuracy over state-of-the-arts.
Gengshen Wu, Zijia Lin, Guiguang Ding, Qiang Ni, Jungong Han
IEEE Trans. Image Process.3
2020 ACMNet: Adaptive Confidence Matching Network for Human Behavior Analysis via Cross-modal Retrieval
abstract
Cross-modality human behavior analysis has attracted much attention from both academia and industry. In this article, we focus on the cross-modality image-text retrieval problem for human behavior analysis, which can learn a common latent space for cross-modality data and thus benefit the understanding of human behavior with data from different modalities. Existing state-of-the-art cross-modality image-text retrieval models tend to be fine-grained region-word matching approaches, where they begin with measuring similarities for each image region or text word followed by aggregating them to estimate the global image-text similarity. However, it is observed that such fine-grained approaches often encounter the similarity bias problem, because they only consider matched text words for an image region or matched image regions for a text word for similarity calculation, but they totally ignore unmatched words/regions, which might still be salient enough to affect the global image-text similarity. In this article, we propose an Adaptive Confidence Matching Network (ACMNet), which is also a fine-grained matching approach, to effectively deal with such a similarity bias. Apart from calculating the local similarity for each region(/word) with its matched words(/regions), ACMNet also introduces a confidence score for the local similarity by leveraging the global text(/image) information, which is expected to help measure the semantic relatedness of the region(/word) to the whole text(/image). Moreover, ACMNet also incorporates the confidence scores together with the local similarities in estimating the global image-text similarity. To verify the effectiveness of ACMNet, we conduct extensive experiments and make comparisons with state-of-the-art methods on two benchmark datasets, i.e., Flickr30k and MS COCO. Experimental results show that the proposed ACMNet can outperform the state-of-the-art methods by a clear margin, which well demonstrates the effectiveness of the proposed ACMNet in human behavior analysis and the reasonableness of tackling the mentioned similarity bias issue.
Hui Chen 0013, Guiguang Ding, Zijia Lin, Sicheng Zhao, Xiaopeng Gu, Wenyuan Xu 0001, Jungong Han
ACM Trans. Multim. Comput. Commun. Appl.2
2019 GRN: Gated Relation Network to Enhance Convolutional Neural Network for Named Entity Recognition
abstract
The dominant approaches for named entity recognitionm (NER) mostly adopt complex recurrent neural networks (RNN), e.g., long-short-term-memory (LSTM). However, RNNs are limited by their recurrent nature in terms of computational efficiency. In contrast, convolutional neural networks (CNN) can fully exploit the GPU parallelism with their feedforward architectures. However, little attention has been paid to performing NER with CNNs, mainly owing to their difficulties in capturing the long-term context information in a sequence. In this paper, we propose a simple but effective CNN-based network for NER, i.e., gated relation network (GRN), which is more capable than common CNNs in capturing long-term context. Specifically, in GRN we firstly employ CNNs to explore the local context features of each word. Then we model the relations between words and use them as gates to fuse local context features into global ones for predicting labels. Without using recurrent layers that process a sentence in a sequential manner, our GRN allows computations to be performed in parallel across the entire sentence. Experiments on two benchmark NER datasets (i.e., CoNLL2003 and Ontonotes 5.0) show that, our proposed GRN can achieve state-of-the-art performance with or without external knowledge. It also enjoys lower time costs to train and test.
Hui Chen 0013, Zijia Lin, Guiguang Ding, Jianguang Lou, Yusen Zhang 0003, Börje Karlsson 0001
AAAI3
2019 Dual-View Ranking with Hardness Assessment for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL.
Guiguang Ding, Jungong Han, Xiaohan Ding, Sicheng Zhao, Zheng Wang 0001, Chenggang Yan 0001, Qionghai Dai
AAAI2
2019 Adaptive Region Embedding for Text Classification
abstract
Deep learning models such as convolutional neural networks and recurrent networks are widely applied in text classification. In spite of their great success, most deep learning models neglect the importance of modeling context information, which is crucial to understanding texts. In this work, we propose the Adaptive Region Embedding to learn context representation to improve text classification. Specifically, a metanetwork is learned to generate a context matrix for each region, and each word interacts with its corresponding context matrix to produce the regional representation for further classification. Compared to previous models that are designed to capture context information, our model contains less parameters and is more flexible. We extensively evaluate our method on 8 benchmark datasets for text classification. The experimental results prove that our method achieves state-of-the-art performances and effectively avoids word ambiguity.
Liuyu Xiang, Xiaoming Jin, Lan Yi, Guiguang Ding
AAAI4
2019 CycleEmotionGAN: Emotional Semantic Consistency Preserved CycleGAN for Adapting Image Emotions
abstract
Deep neural networks excel at learning from large-scale labeled training data, but cannot well generalize the learned knowledge to new domains or datasets. Domain adaptation studies how to transfer models trained on one labeled source domain to another sparsely labeled or unlabeled target domain. In this paper, we investigate the unsupervised domain adaptation (UDA) problem in image emotion classification. Specifically, we develop a novel cycle-consistent adversarial model, termed CycleEmotionGAN, by enforcing emotional semantic consistency while adapting images cycleconsistently. By alternately optimizing the CycleGAN loss, the emotional semantic consistency loss, and the target classification loss, CycleEmotionGAN can adapt source domain images to have similar distributions to the target domain without using aligned image pairs. Simultaneously, the annotation information of the source images is preserved. Extensive experiments are conducted on the ArtPhoto and FI datasets, and the results demonstrate that CycleEmotionGAN significantly outperforms the state-of-the-art UDA approaches.
Sicheng Zhao, Chuang Lin 0003, Pengfei Xu 0013, Sendong Zhao, Ravi Krishna, Guiguang Ding, Kurt Keutzer
AAAI7
2019 Recurrent Attention Model for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that many semantic pedestrian attributes to be recognised tend to show spatial locality and semantic correlations by which they can be grouped while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)’s super capability of learning context correlations and Attention Model’s capability of highlighting the region of interest on feature map, this paper proposes end-to-end Recurrent Convolutional (RC) and Recurrent Attention (RA) models, which are complementary to each other. RC model mines the correlations among different attribute groups with convolutional LSTM unit, while RA model takes advantage of the intra-group spatial locality and inter-group attention correlation to improve the performance of pedestrian attribute recognition. Our RA method combines the Recurrent Learning and Attention Model to highlight the spatial position on feature map and mine the attention correlations among different attribute groups to obtain more precise attention. Extensive empirical evidence shows that our recurrent model frameworks achieve state-of-the-art results, based on pedestrian attribute datasets, i.e. standard PETA and RAP datasets.
Xin Zhao 0020, Liufang Sang, Guiguang Ding, Jungong Han, Na Di, Chenggang Yan 0001
AAAI3
2019 Centripetal SGD for Pruning Very Deep Convolutional Networks With Complicated Structure
abstract
The redundancy is widely recognized in Convolutional Neural Networks (CNNs), which enables to remove some unimportant filters from convolutional layers so as to slim the network with acceptable performance drop. Inspired by the linearity of convolution, we seek to make some filters increasingly close and eventually identical for network slimming. To this end, we propose Centripetal SGD (C-SGD), a novel optimization method, which can train several filters to collapse into a single point in the parameter hyperspace. When the training is completed, the removal of the identical filters can trim the network with NO performance loss, thus no finetuning is needed. By doing so, we have partly solved an open problem of constrained filter pruning on CNNs with complicated structure, where some layers must be pruned following the others. Our experimental results on CIFAR-10 and ImageNet have justified the effectiveness of C-SGD-based filter pruning. Moreover, we have provided empirical evidences for the assumption that the redundancy in deep neural networks helps the convergence of training by showing that a redundant CNN trained using C-SGD outperforms a normally trained counterpart with the equivalent width.
Xiaohan Ding, Guiguang Ding, Jungong Han
CVPR2
2019 ACNet: Strengthening the Kernel Skeletons for Powerful CNN via Asymmetric Convolution Blocks
abstract
As designing appropriate Convolutional Neural Network (CNN) architecture in the context of a given application usually involves heavy human works or numerous GPU hours, the research community is soliciting the architecture-neutral CNN structures, which can be easily plugged into multiple mature architectures to improve the performance on our real-world applications. We propose Asymmetric Convolution Block (ACB), an architecture-neutral structure as a CNN building block, which uses 1D asymmetric convolutions to strengthen the square convolution kernels. For an off-the-shelf architecture, we replace the standard square-kernel convolutional layers with ACBs to construct an Asymmetric Convolutional Network (ACNet), which can be trained to reach a higher level of accuracy. After training, we equivalently convert the ACNet into the same original architecture, thus requiring no extra computations anymore. We have observed that ACNet can improve the performance of various models on CIFAR and ImageNet by a clear margin. Through further experiments, we attribute the effectiveness of ACB to its capability of enhancing the model's robustness to rotational distortions and strengthening the central skeleton parts of square convolution kernels.
Xiaohan Ding, Guiguang Ding, Jungong Han
ICCV3
2019 Towards Better Uncertainty Sampling: Active Learning with Multiple Views for Deep Convolutional Neural Network
abstract
Convolutional neural network (CNN) has been successfully applied to many fields, such as image classification and object detection. It relies on huge amount of data. However, labelling a large amount of data is expensive. Active learning is one of the approaches to alleviate the labelling effort. We propose a new active learning approach for CNN. Different from existing active learning algorithms for CNN, first, the active query strategy is measured from multiple views, not only the last output of CNN; second, multiple views are obtained from multiple hidden layers in CNN, not from other related data or models. We evaluate our approach on three widely used datasets: Fashion-MNIST, SVHN and CIFAR-10. Experimental results show that the proposed method outperforms baseline methods in image classification.
Xiaoming Jin, Guiguang Ding, Lan Yi, Chenggang Yan 0001
ICME3
2019 Personality Driven Multi-task Learning for Image Aesthetic Assessment
abstract
With the prevalence of convolutional neural networks (CNNs), assessing the aesthetics of an image has gained great advances recently. Individual users often have different aesthetic preferences on images, which we believe are mainly affected by their personality traits. However, most of the current aesthetics models predict a generic aesthetic score based on handcrafted and/or learned feature representations, which are unified and thus cannot reflect the individual differences during image aesthetic rating. In this paper, we propose an end-to-end personality driven multi-task deep learning model to address this problem. Firstly, both image aesthetics and personality traits are learned from the proposed multi-task model. Then the personality features are employed to modulate the aesthetics features, producing the optimal generic image aesthetics scores. The experimental results on two public databases show that the proposed method is superior to the state-of-the-art approaches.
Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Allen Tan
ICME4
2019 Approximated Oracle Filter Pruning for Destructive CNN Width Optimization
abstract
It is not easy to design and run Convolutional Neural Networks (CNNs) due to: 1) finding the optimal number of filters (i.e., the width) at each layer is tricky, given an architecture; and 2) the computational intensity of CNNs impedes the deployment on computationally limited devices. Oracle Pruning is designed to remove the unimportant filters from a well-trained CNN, which estimates the filters’ importance by ablating them in turn and evaluating the model, thus delivers high accuracy but suffers from intolerable time complexity, and requires a given resulting width but cannot automatically find it. To address these problems, we propose Approximated Oracle Filter Pruning (AOFP), which keeps searching for the least important filters in a binary search manner, makes pruning attempts by masking out filters randomly, accumulates the resulting errors, and finetunes the model via a multi-path framework. As AOFP enables simultaneous pruning on multiple layers, we can prune an existing very deep CNN with acceptable time cost, negligible accuracy drop, and no heuristic knowledge, or re-design a model which exerts higher accuracy and faster inference.
Xiaohan Ding, Guiguang Ding, Jungong Han, Chenggang Yan 0001
ICML2
2019 Zero-shot Learning with Many Classes by High-rank Deep Embedding Networks
abstract
Zero-shot learning (ZSL) is a recently emerging research topic which aims to build classification models for unseen classes with knowledge from auxiliary seen classes. Though many ZSL works have shown promising results on small-scale datasets by utilizing a bilinear compatibility function, the ZSL performance on large-scale datasets with many classes (say, ImageNet) is still unsatisfactory. We argue that the bilinear compatibility function is a low-rank approximation of the true compatibility function such that it is not expressive enough especially when there are a large number of classes because of the rank limitation. To address this issue, we propose a novel approach, termed as High-rank Deep Embedding Networks (GREEN), for ZSL with many classes. In particular, we propose a feature-dependent mixture of softmaxes as the image-class compatibility function, which is a simple extension of the bilinear compatibility function, but yields much better results. It utilizes a mixture of non-linear transformations with feature-dependent latent variables to approximate the true function in a high-rank way, which makes GREEN more expressive. Experiments on several datasets including ImageNet demonstrate GREEN significantly outperforms the state-of-the-art approaches.
Guiguang Ding, Jungong Han, Qionghai Dai
IJCAI2
2019 Landmark Selection for Zero-shot Learning
abstract
Zero-shot learning (ZSL) is an emerging research topic whose goal is to build recognition models for previously unseen classes. The basic idea of ZSL is based on heterogeneous feature matching which learns a compatibility function between image and class features using seen classes. The function is constructed based on one-vs-all training in which each class has only one class feature and many image features. Existing ZSL works mostly treat all image features equivalently. However, in this paper we argue that it is more reasonable to use some representative cross-domain data instead of all. Motivated by this idea, we propose a novel approach, termed as Landmark Selection(LAST) for ZSL. LAST is able to identify representative cross-domain features which further lead to better image-class compatibility function. Experiments on several ZSL datasets including ImageNet demonstrate the superiority of LAST to the state-of-the-arts.
Guiguang Ding, Jungong Han, Chenggang Yan 0001, Jiyong Zhang 0001, Qionghai Dai
IJCAI2
2019 Low Shot Box Correction for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) has been widely studied but the accuracy of state-of-art methods remains far lower than strongly supervised methods. One major reason for this huge gap is the incomplete box detection problem which arises because most previous WSOD models are structured on classification networks and therefore tend to recognize the most discriminative parts instead of complete bounding boxes. To solve this problem, we define a low-shot weakly supervised object detection task and propose a novel low-shot box correction network to address it. The proposed task enables to train object detectors on a large data set all of which have image-level annotations, but only a small portion or few shots have box annotations. Given the low-shot box annotations, we use a novel box correction network to transfer the incomplete boxes into complete ones. Extensive empirical evidence shows that our proposed method yields state-of-art detection accuracy under various settings on the PASCAL VOC benchmark.
Tianxiang Pan, Bin Wang 0021, Guiguang Ding, Jungong Han, Jun-Hai Yong
IJCAI3
2019 Incremental Few-Shot Learning for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition has received increasing attention due to its important role in video surveillance applications. However, most existing methods are designed for a fixed set of attributes. They are unable to handle the incremental few-shot learning scenario, i.e. adapting a well-trained model to newly added attributes with scarce data, which commonly exists in the real world. In this work, we present a meta learning based method to address this issue. The core of our framework is a meta architecture capable of disentangling multiple attribute information and generalizing rapidly to new coming attributes. By conducting extensive experiments on the benchmark dataset PETA and RAP under the incremental few-shot setting, we show that our method is able to perform the task with competitive performances and low resource requirements.
Liuyu Xiang, Xiaoming Jin, Guiguang Ding, Jungong Han, Leida Li
IJCAI3
2019 Cross-Modal Image-Text Retrieval with Semantic Consistency
abstract
Cross-modal image-text retrieval has been a long-standing challenge in the multimedia community. Existing methods explore various complicated embedding spaces to assess the semantic similarity between a given image-text pair, but consider no/little about the consistency across them. To remedy this situation, we introduce the idea of semantic consistency for learning various embedding spaces jointly. Specifically, similar to the previous works, we start by constructing two different embedding spaces, namely the image-grounded embedding space and the text-grounded embedding space. However, instead of learning these two embedding spaces separately, we incorporate a semantic consistency constraint in the common ranking objective function such that both embedding spaces can be learned simultaneously and benefit from each other to gain performance improvement. We conduct extensive experiments on three benchmark datasets, \ie Flickr8k, Flickr30k and MS COCO. Results show that our model outperforms the state-of-the-art models on all three datasets, which can well demonstrate the effectiveness and superiority of the introduction of semantic consistency. Our source code is released at: \urlhttps://github.com/HuiChen24/SemanticConsistency.
Hui Chen 0013, Guiguang Ding, Zijia Lin, Sicheng Zhao, Jungong Han
ACM Multimedia2
2019 PDANet: Polarity-consistent Deep Attention Network for Fine-grained Visual Emotion Regression
abstract
Existing methods on visual emotion analysis mainly focus on coarse-grained emotion classification, i.e. assigning an image with a dominant discrete emotion category. However, these methods cannot well reflect the complexity and subtlety of emotions. In this paper, we study the fine-grained regression problem of visual emotions based on convolutional neural networks (CNNs). Specifically, we develop a Polarity-consistent Deep Attention Network (PDANet), a novel network architecture that integrates attention into a CNN with an emotion polarity constraint. First, we propose to incorporate both spatial and channel-wise attentions into a CNN for visual emotion regression, which jointly considers the local spatial connectivity patterns along each channel and the interdependency between different channels. Second, we design a novel regression loss, i.e. polarity-consistent regression (PCR) loss, based on the weakly supervised emotion polarity to guide the attention generation. By optimizing the PCR loss, PDANet can generate a polarity preserved attention map and thus improve the emotion regression performance. Extensive experiments are conducted on the IAPS, NAPS, and EMOTIC datasets, and the results demonstrate that the proposed PDANet outperforms the state-of-the-art approaches by a large margin for fine-grained visual emotion regression. Our source code is released at: https://github.com/ZizhouJia/PDANet.
Sicheng Zhao, Zizhou Jia, Hui Chen 0013, Leida Li, Guiguang Ding, Kurt Keutzer
ACM Multimedia5
2019 Global Sparse Momentum SGD for Pruning Very Deep Neural Networks
abstract
Deep Neural Network (DNN) is powerful but computationally expensive and memory intensive, thus impeding its practical usage on resource-constrained front-end devices. DNN pruning is an approach for deep model compression, which aims at eliminating some parameters with tolerable performance degradation. In this paper, we propose a novel momentum-SGD-based optimization method to reduce the network complexity by on-the-fly pruning. Concretely, given a global compression ratio, we categorize all the parameters into two parts at each training iteration which are updated using different rules. In this way, we gradually zero out the redundant parameters, as we update them using only the ordinary weight decay but no gradients derived from the objective function. As a departure from prior methods that require heavy human works to tune the layer-wise sparsity ratios, prune by solving complicated non-differentiable problems or finetune the model after pruning, our method is characterized by 1) global compression that automatically finds the appropriate per-layer sparsity ratios; 2) end-to-end training; 3) no need for a time-consuming re-training process after pruning; and 4) superior capability to find better winning tickets which have won the initialization lottery.
Xiaohan Ding, Guiguang Ding, Xiangxin Zhou, Jungong Han, Ji Liu 0002
NeurIPS2
2019 Zero-shot multi-label learning via label factorisation
abstract
This study considers the zero‐shot learning problem under the multi‐label setting where each test sample is associated with multiple labels that are unseen in training data. The authors propose a novel learning framework based on label factorisation for this problem. Specifically, the authors’ framework takes three key issues into consideration and addresses them in a unified way. The first is knowledge transfer that utilises information from seen classes to build recognition models for unseen classes. The second is label correlation which means that labels which have different semantics may co‐occur frequently. This is an important issue in multi‐label learning. The authors propose to learn a shared latent space by label factorisation and use the label semantics as the decoding function, which can address both issues. The third is the predictability which requires the learned latent space to be strongly related to the visual features. It is guaranteed by incorporating a regression model into the learning framework. The authors derive two specific formulations from the general framework and propose the corresponding learning algorithms. The authors conducted extensive experiments on three multi‐label data sets. The results demonstrated the effectiveness.
Guiguang Ding, Jungong Han
IET Comput. Vis.3
2019 MEIAH: Mixing explicit and implicit formulation of attributes in binary representation for person re-identification
Liufang Sang, Xin Zhao 0020, Guiguang Ding
Multim. Tools Appl.3
2019 Optimized projection for hashing
Chaoqun Chu, Dahan Gong, Kai Chen 0044, Jungong Han, Guiguang Ding
Pattern Recognit. Lett.6
2019 Cyber-Physical Security Design in Multimedia Data Cache Resource Allocation for Industrial Networks
abstract
For cyber-physical industrial networks, more and more multimedia data is faced in high-speed information transmission. The explosive data brings more challenges to the architecture of modern industrial networks. In this way, cache resource allocation technology is necessary for practical applications. In order to design reasonable caching framework, how to predict the multimedia data request is an important issue. In order to keep efficient and reliable data transmission in wireless industrial networks, security design is also critical for existing cache resource allocation. Based on previous works, some promising technologies have been applied, such as heterogeneous ultradense networks, wireless edge caching, and web content popularity prediction. In this paper, we summarize these promising technologies and provide a useful guidance for security design in cyber-physical cache resource allocation system. Specially, we can divide the proposed method into three main steps. First of all, a spatio-temporal multimedia content prediction based on long short-term memory is proposed for accurate prediction on multimedia data request. After that, we make use of Zipf fitting for caching model. At last, the caching optimization considering security is put forward in this paper. Experimental results show the satisfied performance of our proposed algorithm and it has obvious potential application value in cyber-physical industrial networks with cache resource allocation technology.
Bin Jiang 0003, Guiguang Ding, Huihui Wang 0001
IEEE Trans. Ind. Informatics3
2019 DECODE: Deep Confidence Network for Robust Image Classification
abstract
Recent years have witnessed the success of deep convolutional neural networks for image classification and many related tasks. It should be pointed out that the existing training strategies assume that there is a clean dataset for model learning. In elaborately constructed benchmark datasets, deep network has yielded promising performance under the assumption. However, in real-world applications, it is burdensome and expensive to collect sufficient clean training samples. On the other hand, collecting noisy labeled samples is very economical and practical, especially with the rapidly increasing amount of visual data in the web. Unfortunately, the accuracy of current deep models may drop dramatically even with 5%-10% label noise. Therefore, enabling label noise resistant classification has become a crucial issue in the data driven deep learning approaches. In this paper, we propose a DEep COnfiDEnce network (DECODE) to address this issue. In particular, based on the distribution of mislabeled data, we adopt a confidence evaluation module that is able to determine the confidence that a sample is mislabeled. With the confidence, we further use a weighting strategy to assign different weights to different samples so that the model pays less attention to low confidence data, which is more likely to be noise. In this way, the deep model is more robust to label noise. DECODE is designed to be general, such that it can be easily combined with existing studies. We conduct extensive experiments on several datasets, and the results validate that DECODE can improve the accuracy of deep models trained with noisy data.
Guiguang Ding, Kai Chen 0044, Chaoqun Chu, Jungong Han, Qionghai Dai
IEEE Trans. Image Process.1
2019 Unsupervised Deep Video Hashing via Balanced Code for Large-Scale Video Retrieval
abstract
This paper proposes a deep hashing framework, namely Unsupervised Deep Video Hashing (UDVH), for largescale video similarity search with the aim to learn compact yet effective binary codes. Our UDVH produces the hash codes in a self-taught manner by jointly integrating discriminative video representation with optimal code learning, where an efficient alternating approach is adopted to optimize the objective function. The key differences from most existing video hashing methods lie in 1) UDVH is an unsupervised hashing method that generates hash codes by cooperatively utilizing feature clustering and a specifically-designed binarization with the original neighborhood structure preserved in the binary space; 2) a specific rotation is developed and applied onto video features such that the variance of each dimension can be balanced, thus facilitating the subsequent quantization step. Extensive experiments performed on three popular video datasets show that UDVH is overwhelmingly better than the state-of-the-arts in terms of various evaluation metrics, which makes it practical in real-world applications.
Gengshen Wu, Jungong Han, Li Liu 0004, Guiguang Ding, Qiang Ni, Ling Shao 0001
IEEE Trans. Image Process.5
2019 Personalized Emotion Recognition by Personality-Aware High-Order Learning of Physiological Signals
abstract
Due to the subjective responses of different subjects to physical stimuli, emotion recognition methodologies from physiological signals are increasingly becoming personalized. Existing works mainly focused on modeling the involved physiological corpus of each subject, without considering the psychological factors, such as interest and personality. The latent correlation among different subjects has also been rarely examined. In this article, we propose to investigate the influence of personality on emotional behavior in a hypergraph learning framework. Assuming that each vertex is a compound tuple (subject, stimuli), multi-modal hypergraphs can be constructed based on the personality correlation among different subjects and on the physiological correlation among corresponding stimuli. To reveal the different importance of vertices, hyperedges, and modalities, we learn the weights for each of them. As the hypergraphs connect different subjects on the compound vertices, the emotions of multiple subjects can be simultaneously recognized. In this way, the constructed hypergraphs are vertex-weighted multi-modal multi-task ones. The estimated factors, referred to as emotion relevance, are employed for emotion recognition. We carry out extensive experiments on the ASCERTAIN dataset and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art emotion recognition approaches.
Sicheng Zhao, Amir Gholami, Guiguang Ding, Yue Gao 0002, Jungong Han, Kurt Keutzer
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Temporal-Difference Learning With Sampling Baseline for Image Captioning
abstract
The existing methods for image captioning usually train the language model under the cross entropy loss, which results in the exposure bias and inconsistency of evaluation metric. Recent research has shown these two issues can be well addressed by policy gradient method in reinforcement learning domain attributable to its unique capability of directly optimizing the discrete and non-differentiable evaluation metric. In this paper, we utilize reinforcement learning method to train the image captioning model. Specifically, we train our image captioning model to maximize the overall reward of the sentences by adopting the temporal-difference (TD) learning method, which takes the correlation between temporally successive actions into account. In this way, we assign different values to different words in one sampled sentence by a discounted coefficient when back-propagating the gradient with the REINFORCE algorithm, enabling the correlation between actions to be learned. Besides, instead of estimating a "baseline" to normalize the rewards with another network, we utilize the reward of another Monte-Carlo sample as the "baseline" to avoid high variance. We show that our proposed method can improve the quality of generated captions and outperforms the state-of-the-art methods on the benchmark dataset MS COCO in terms of seven evaluation metrics.
Hui Chen 0013, Guiguang Ding, Sicheng Zhao, Jungong Han
AAAI2
2018 Auto-Balanced Filter Pruning for Efficient Convolutional Neural Networks
abstract
In recent years considerable research efforts have been devoted to compression techniques of convolutional neural networks (CNNs). Many works so far have focused on CNN connection pruning methods which produce sparse parameter tensors in convolutional or fully-connected layers. It has been demonstrated in several studies that even simple methods can effectively eliminate connections of a CNN. However, since these methods make parameter tensors just sparser but no smaller, the compression may not transfer directly to acceleration without support from specially designed hardware. In this paper, we propose an iterative approach named Auto-balanced Filter Pruning, where we pre-train the network in an innovative auto-balanced way to transfer the representational capacity of its convolutional layers to a fraction of the filters, prune the redundant ones, then re-train it to restore the accuracy. In this way, a smaller version of the original network is learned and the floating-point operations (FLOPs) are reduced. By applying this method on several common CNNs, we show that a large portion of the filters can be discarded without obvious accuracy drop, leading to significant reduction of computational burdens. Concretely, we reduce the inference cost of LeNet-5 on MNIST, VGG-16 and ResNet-56 on CIFAR-10 by 95.1%, 79.7% and 60.9%, respectively.
Xiaohan Ding, Guiguang Ding, Jungong Han, Sheng Tang
AAAI2
2018 Zero-Shot Learning With Attribute Selection
abstract
Zero-shot learning (ZSL) is regarded as an effective way to construct classification models for target classes which have no labeled samples available. The basic framework is to transfer knowledge from (different) auxiliary source classes having sufficient labeled samples with some attributes shared by target and source classes as bridge. Attributes play an important role in ZSL but they have not gained sufficient attention in recent years. Previous works mostly assume attributes are perfect and treat each attribute equally. However, as shown in this paper, different attributes have different properties, such as their class distribution, variance, and entropy, which may have considerable impact on ZSL accuracy if treated equally. Based on this observation, in this paper we propose to use a subset of attributes, instead of the whole set, for building ZSL models. The attribute selection is conducted by considering the information amount and predictability under a novel joint optimization framework. To our knowledge, this is the first work that notices the influence of attributes themselves and proposes to use a refined attribute set for ZSL. Since our approach focuses on selecting good attributes for ZSL, it can be combined to any attribute based ZSL approaches so as to augment their performance. Experiments on four ZSL benchmarks demonstrate that our approach can improve zero-shot classification accuracy and yield state-of-the-art results.
Guiguang Ding, Jungong Han, Sheng Tang
AAAI2
2018 On Trivial Solution and High Correlation Problems in Deep Supervised Hashing
abstract
Deep supervised hashing (DSH), which combines binary learning and convolutional neural network, has attracted considerable research interests and achieved promising performance for highly efficient image retrieval. In this paper, we show that the widely used loss functions, pair-wise loss and triplet loss, suffer from the trivial solution problem and usually lead to highly correlated bits in practice, limiting the performance of DSH. One important reason is that it is difficult to incorporate proper constraints into the loss functions under the mini-batch based optimization algorithm. To tackle these problems, we propose to adopt ensemble learning strategy for deep model training. We found out that this simple strategy is capable of effectively decorrelating different bits, making the hashcodes more informative. Moreover, it is very easy to parallelize the training and support incremental model learning, which are very useful for real-world applications but usually ignored by existing DSH approaches. Experiments on benchmarks demonstrate the proposed ensemble based DSH can improve the performance of DSH approaches significant.
Xin Zhao 0020, Guiguang Ding, Jungong Han
AAAI3
2018 Shadow Detection Using Robust Texture Learning
Tianxiang Pan, Bin Wang 0021, Guiguang Ding, Jun-Hai Yong
BMVC3
2018 Show, Observe and Tell: Attribute-driven Attention Model for Image Captioning
abstract
Despite the fact that attribute-based approaches and attention-based approaches have been proven to be effective in image captioning, most attribute-based approaches simply predict attributes independently without taking the co-occurrence dependencies among attributes into account. Besides, most attention-based captioning models directly leverage the feature map extracted from CNN, in which many features may be redundant in relation to the image content. In this paper, we focus on training a good attribute-inference model via the recurrent neural network (RNN) for image captioning, where the co-occurrence dependencies among attributes can be maintained. The uniqueness of our inference model lies in the usage of a RNN with the visual attention mechanism to \textit{observe} the image before generating captions. Additionally, it is noticed that compact and attribute-driven features will be more useful for the attention-based captioning model. To this end, we extract the context feature for each attribute, and guide the captioning model adaptively attend to these context features. We verify the effectiveness and superiority of the proposed approach over the other captioning approaches by conducting massive experiments and comparisons on MS COCO image captioning dataset.
Hui Chen 0013, Guiguang Ding, Zijia Lin, Sicheng Zhao, Jungong Han
IJCAI2
2018 Implicit Non-linear Similarity Scoring for Recognizing Unseen Classes
abstract
Recognizing unseen classes is an important task for real-world applications, due to: 1) it is common that some classes in reality have no labeled image exemplar for training; and 2) novel classes emerge rapidly. Recently, to address this task many zero-shot learning (ZSL) approaches have been proposed where explicit linear scores, like inner product score, are employed to measure the similarity between a class and an image. We argue that explicit linear scoring (ELS) seems too weak to capture complicated image-class correspondence. We propose a simple yet effective framework, called Implicit Non-linear Similarity Scoring (ICINESS). In particular, we train a scoring network which uses image and class features as input, fuses them by hidden layers, and outputs the similarity. Based on the universal approximation theorem, it can approximate the true similarity function between images and classes if a proper structure is used in an implicit non-linear way, which is more flexible and powerful. With ICINESS framework, we implement ZSL algorithms by shallow and deep networks, which yield consistently superior results.
Guiguang Ding, Jungong Han, Sicheng Zhao, Bin Wang 0021
IJCAI2
2018 Automatic Gating of Attributes in Deep Structure
abstract
Deep structure has been widely applied in a large variety of fields for its excellence of representing data. Attributes are a unique type of data descriptions that have been successfully utilized in numerous tasks to enhance performance. However, to introduce attributes into deep structure is complicated and challenging, because different layers in deep structure accommodate features of different abstraction levels, while different attributes may naturally represent the data in different abstraction levels. This demands adaptively and jointly modeling of attributes and deep structure by carefully examining their relationship. Different from existing works that treat attributes straightforwardly as the same level without considering their abstraction levels, we can make better use of attributes in deep structure by properly connecting them. In this paper, we move forward along this new direction by proposing a deep structure named Attribute Gated Deep Belief Network (AG-DBN) that includes a tunable attribute-layer gating mechanism and automatically learns the best way of connecting attributes to appropriate hidden layers. Experimental results on a manually-labeled subset of ImageNet, a-Yahoo and a-Pascal data set justify the superiority of AG-DBN against several baselines including CNN model and other AG-DBN variants. Specifically, it outperforms the CNN model, VGG19, by significantly reducing the classification error from 26.70% to 13.56% on a-Pascal.
Xiaoming Jin, Lan Yi, Guiguang Ding, Dou Shen
IJCAI5
2018 Unsupervised Deep Hashing via Binary Latent Factor Models for Large-scale Cross-modal Retrieval
abstract
Despite its great success, matrix factorization based cross-modality hashing suffers from two problems: 1) there is no engagement between feature learning and binarization; and 2) most existing methods impose the relaxation strategy by discarding the discrete constraints when learning the hash function, which usually yields suboptimal solutions. In this paper, we propose a novel multimodal hashing framework, referred as Unsupervised Deep Cross-Modal Hashing (UDCMH), for multimodal data search in a self-taught manner via integrating deep learning and matrix factorization with binary latent factor models. On one hand, our unsupervised deep learning framework enables the feature learning to be jointly optimized with the binarization. On the other hand, the hashing system based on the binary latent factor models can generate unified binary codes by solving a discrete-constrained objective function directly with no need for a relaxation step. Moreover, novel Laplacian constraints are incorporated into the objective function, which allow to preserve not only the nearest neighbors that are commonly considered in the literature but also the farthest neighbors of data, even if the semantic labels are not available. Extensive experiments on multiple datasets highlight the superiority of the proposed framework over several state-of-the-art baselines.
Gengshen Wu, Zijia Lin, Jungong Han, Li Liu 0004, Guiguang Ding, Baochang Zhang 0001, Jialie Shen 0001
IJCAI5
2018 Affective Image Content Analysis: A Comprehensive Survey
abstract
Images can convey rich semantics and induce strong emotions in viewers. Recently, with the explosive growth of visual data, extensive research efforts have been dedicated to affective image content analysis (AICA). In this paper, we review the state-of-the-art methods comprehensively with respect to two main challenges -- affective gap and perception subjectivity. We begin with an introduction to the key emotion representation models that have been widely employed in AICA. Available existing datasets for performing evaluation are briefly described. We then summarize and compare the representative approaches on emotion feature extraction, personalized emotion prediction, and emotion distribution learning. Finally, we discuss some future research directions.
Sicheng Zhao, Guiguang Ding, Qingming Huang, Tat-Seng Chua, Björn W. Schuller, Kurt Keutzer
IJCAI2
2018 Personality-Aware Personalized Emotion Recognition from Physiological Signals
abstract
Emotion recognition methodologies from physiological signals are increasingly becoming personalized, due to the subjective responses of different subjects to physical stimuli. Existing works mainly focused on modelling the involved physiological corpus of each subject, without considering the psychological factors. The latent correlation among different subjects has also been rarely examined. We propose to investigate the influence of personality on emotional behavior in a hypergraph learning framework. Assuming that each vertex is a compound tuple (subject, stimuli), multi-modal hypergraphs can be constructed based on the personality correlation among different subjects and on the physiological correlation among corresponding stimuli. To reveal the different importance of vertices, hyperedges, and modalities, we assign each of them with weights. The emotion relevance learned on the vertex-weighted multi-modal multi-task hypergraphs is employed for emotion recognition. We carry out extensive experiments on the ASCERTAIN dataset and the results demonstrate the superiority of the proposed method.
Sicheng Zhao, Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI2
2018 Grouping Attribute Recognition for Pedestrian with Joint Recurrent Learning
abstract
Pedestrian attributes recognition is to predict attribute labels of pedestrian from surveillance images, which is a very challenging task for computer vision due to poor imaging quality and small training dataset. It is observed that semantic pedestrian attributes to be recognised tend to show semantic or visual spatial correlation. Attributes can be grouped by the correlation while previous works mostly ignore this phenomenon. Inspired by Recurrent Neural Network (RNN)'s super capability of learning context correlations, this paper proposes an end-to-end Grouping Recurrent Learning (GRL) model that takes advantage of the intra-group mutual exclusion and inter-group correlation to improve the performance of pedestrian attribute recognition. Our GRL method starts with the detection of precise body region via Body Region Proposal followed by feature extraction from detected regions. These features, along with the semantic groups, are fed into RNN for recurrent grouping attribute recognition, where intra group correlations can be learned. Extensive empirical evidence shows that our GRL model achieves state-of-the-art results, based on pedestrian attribute datasets, i.e. standard PETA and RAP datasets.
Xin Zhao 0020, Liufang Sang, Guiguang Ding, Xiaoming Jin
IJCAI3
2018 Where to Prune: Using LSTM to Guide End-to-end Pruning
abstract
Recent years have witnessed the great success of convolutional neural networks (CNNs) in many related fields. However, its huge model size and computation complexity bring in difficulty when deploying CNNs in some scenarios, like embedded system with low computation power. To address this issue, many works have been proposed to prune filters in CNNs to reduce computation. However, they mainly focus on seeking which filters are unimportant in a layer and then prune filters layer by layer or globally. In this paper, we argue that the pruning order is also very significant for model pruning. We propose a novel approach to figure out which layers should be pruned in each step. First, we utilize a long short-term memory (LSTM) to learn the hierarchical characteristics of a network and generate a pruning decision for each layer, which is the main difference from previous works. Next, a channel-based method is adopted to evaluate the importance of filters in a to-be-pruned layer, followed by an accelerated recovery step. Experimental results demonstrate that our approach is capable of reducing 70.1% FLOPs for VGG and 47.5% for Resnet-56 with comparable accuracy. Also, the learning results seem to reveal the sensitivity of each network layer.
Guiguang Ding, Jungong Han, Bin Wang 0021
IJCAI2
2018 EmotionGAN: Unsupervised Domain Adaptation for Learning Discrete Probability Distributions of Image Emotions
abstract
Deep neural networks have performed well on various benchmark vision tasks with large-scale labeled training data; however, such training data is expensive and time-consuming to obtain. Due to domain shift or dataset bias, directly transferring models trained on a large-scale labeled source domain to another sparsely labeled or unlabeled target domain often results in poor performance. In this paper, we consider the domain adaptation problem in image emotion recognition. Specifically, we study how to adapt the discrete probability distributions of image emotions from a source domain to a target domain in an unsupervised manner. We develop a novel adversarial model for emotion distribution learning, termed EmotionGAN, which alternately optimizes the Generative Adversarial Network (GAN) loss, semantic consistency loss, and regression loss. The EmotionGAN model can adapt source domain images such that they appear as if they were drawn from the target domain, while preserving the annotation information. Extensive experiments are conducted on the FlickrLDL and TwitterLDL datasets, and the results demonstrate the superiority of the proposed method as compared to state-of-the-art approaches.
Sicheng Zhao, Xin Zhao 0020, Guiguang Ding, Kurt Keutzer
ACM Multimedia3
2018 Predicting Personalized Image Emotion Perceptions in Social Networks
abstract
Images can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning (RMTHG) is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches.
Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua
IEEE Trans. Affect. Comput.4
2018 Real-Time Multimedia Social Event Detection in Microblog
abstract
Detecting events from massive social media data in social networks can facilitate browsing, search, and monitoring of real-time events by corporations, governments, and users. The short, conversational, heterogeneous, and real-time characteristics of social media data bring great challenges for event detection. The existing event detection approaches rely mainly on textual information, while the visual content of microblogs and the intrinsic correlation among the heterogeneous data are scarcely explored. To deal with the above challenges, we propose a novel real-time event detection method by generating an intermediate semantic level from social multimedia data, named microblog clique (MC), which is able to explore the high correlations among different microblogs. Specifically, the proposed method comprises three stages. First, the heterogeneous data in microblogs is formulated in a hypergraph structure. Hypergraph cut is conducted to group the highly correlated microblogs with the same topics as the MCs, which can address the information inadequateness and data sparseness issues. Second, a bipartite graph is constructed based on the generated MCs and the transfer cut partition is performed to detect the events. Finally, for new incoming microblogs, incremental hypergraph is constructed based on the latest MCs to generate new MCs, which are classified by bipartite graph partition into existing events or new ones. Extensive experiments are conducted on the events in the Brand-Social-Net dataset and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches.
Sicheng Zhao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua
IEEE Trans. Cybern.3
2018 Robust Quantization for General Similarity Search
abstract
The recent years have witnessed the emerging of vector quantization (VQ) techniques for efficient similarity search. VQ partitions the feature space into a set of codewords and encodes data points as integer indices using the codewords. Then the distance between data points can be efficiently approximated by simple memory lookup operations. By the compact quantization, the storage cost, and searching complexity are significantly reduced, thereby facilitating efficient large-scale similarity search. However, the performance of several celebrated VQ approaches degrades significantly when dealing with noisy data. In addition, it can barely facilitate a wide range of applications as the distortion measurement only limits to ℓ2norm. To address the shortcomings of the squared Euclidean (ℓ2,2norm) loss function employed by the VQ approaches, in this paper, we propose a novel robust and general VQ framework, named RGVQ, to enhance both robustness and generalization of VQ approaches. Specifically, a ℓp,q-norm loss function is proposed to conduct the ℓp-norm similarity search, rather than the ℓ2norm search, and the q-th order loss is used to enhance the robustness. Despite the fact that changing the loss function to ℓp,qnorm makes VQ approaches more robust and generic, it brings us a challenge that a non-smooth and non-convex orthogonality constrained ℓp,q-norm function has to be minimized. To solve this problem, we propose a novel and efficient optimization scheme and specify it to VQ approaches and theoretically prove its convergence. Extensive experiments on benchmark data sets demonstrate that the proposed RGVQ is better than the original VQ for several approaches, especially when searching similarity in noisy data.
Guiguang Ding, Jungong Han
IEEE Trans. Image Process.2
2018 Real-Time Scalable Visual Tracking via Quadrangle Kernelized Correlation Filters
abstract
Correlation filter (CF) has been widely used in tracking tasks due to its simplicity and high efficiency. However, conventional CF-based trackers fail to handle the scale variation that occurs when the targeted object is moving, which is one of the most notable unsolved problems of visual object tracking. In this paper, we propose a scalable visual tracking algorithm based on kernelized correlation filters, referred to as quadrangle kernelized correlation filters (QKCF). Unlike existing complicated scalable trackers that either perform the correlation filtering operation multiple times or extract many candidate windows at various scales, our tracker intends to estimate the scale of the object based on the positions of its four corners, which can be detected using a new Gaussian training output matrix within one filtering process. After obtaining four peak values corresponding to the four corners, we measure the detection confidence of each part response by evaluating its spatial and temporal smoothness. On top of it, a weighted Bayesian inference framework is employed to estimate the final location and size of the bounding box from the response matrix, where the weights are synchronized with the calculated detection likelihoods. Experiments are performed on the OTB-100 data set and 16 benchmark sequences with significant scale variations. The results demonstrate the superiority of the proposed method in terms of both effectiveness and robustness, compared with the state-of-the-art methods.
Guiguang Ding, Wenshuo Chen, Sicheng Zhao, Jungong Han, Qiaoyan Liu
IEEE Trans. Intell. Transp. Syst.1
2018 End-to-End Feature-Aware Label Space Encoding for Multilabel Classification With Many Classes
abstract
To make the problem of multilabel classification with many classes more tractable, in recent years, academia has seen efforts devoted to performing label space dimension reduction (LSDR). Specifically, LSDR encodes high-dimensional label vectors into low-dimensional code vectors lying in a latent space, so as to train predictive models at much lower costs. With respect to the prediction, it performs classification for any unseen instance by recovering a label vector from its predicted code vector via a decoding process. In this paper, we propose a novel method, namely End-to-End Feature-aware label space Encoding (E2FE), to perform LSDR. Instead of requiring an encoding function like most previous works, E2FE directly learns a code matrix formed by code vectors of the training instances in an end-to-end manner. Another distinct property of E2FE is its feature awareness attributable to the fact that the code matrix is learned by jointly maximizing the recoverability of the label space and the predictability of the latent space. Based on the learned code matrix, E2FE further trains predictive models to map instance features into code vectors, and also learns a linear decoding matrix for efficiently recovering the label vector of any unseen instance from its predicted code vector. Theoretical analyses show that both the code matrix and the linear decoding matrix in E2FE can be efficiently learned. Moreover, similar to previous works, E2FE can be specified to learn an encoding function. And it can also be extended with kernel tricks to handle nonlinear correlations between the feature space and the latent space. Comprehensive experiments conducted on diverse benchmark data sets with many classes show consistent performance gains of E2FE over the state-of-the-art methods.
Zijia Lin, Guiguang Ding, Jungong Han, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.2
2017 Reference Based LSTM for Image Captioning
abstract
Image captioning is an important problem in artificial intelligence, related to both computer vision and natural language processing. There are two main problems in existing methods: in the training phase, it is difficult to find which parts of the captions are more essential to the image; in the caption generation phase, the objects or the scenes are sometimes misrecognized. In this paper, we consider the training images as the references and propose a Reference based Long Short Term Memory (R-LSTM) model, aiming to solve these two problems in one goal. When training the model, we assign different weights to different words, which enables the network to better learn the key information of the captions. When generating a caption, the consensus score is utilized to exploit the reference information of neighbor images, which might fix the misrecognition and make the descriptions more natural-sounding. The proposed R-LSTM model outperforms the state-of-the-art approaches on the benchmark dataset MS COCO and obtains top 2 position on 11 of the 14 metrics on the online test server.
Minghai Chen, Guiguang Ding, Sicheng Zhao, Hui Chen 0013, Qiang Liu 0016, Jungong Han
AAAI2
2017 Active Learning with Cross-Class Similarity Transfer
abstract
How to save labeling efforts for training supervised classifiers is an important research topic in machine learning community. Active learning (AL) and transfer learning (TL) are two useful tools to achieve this goal, and their combination, i.e., transfer active learning (T-AL) has also attracted considerable research interest. However, existing T-AL approaches consider to transfer knowledge from a source/auxiliary domain which has the same class labels as the target domain, but ignore the relationship among classes. In this paper, we investigate a more practical setting where the classes in source domain are related/similar to but different from the target domain classes. Specifically, we propose a novel cross-class T-AL approach to simultaneously transfer knowledge from source domain and actively annotate the most informative samples in target domain so that we can train satisfactory classifiers with as few labeled samples as possible. In particular, based on the class-class similarity and sample-sample similarity, we adopt a similarity propagation to find the source domain samples that can well capture the characteristics of a target class and then transfer the similar samples as the (pseudo) labeled data for the target class. In turn, the labeled and transferred samples are used to train classifiers and actively select new samples for annotation. Extensive experiments on three datasets demonstrate that the proposed approach outperforms significantly the state-of-the-art related approaches.
Guiguang Ding, Yue Gao 0002, Jungong Han
AAAI2
2017 Zero-Shot Recognition via Direct Classifier Learning with Transferred Samples and Pseudo Labels
abstract
As an interesting and emerging topic, zero-shot recognition (ZSR) makes it possible to train a recognition model by specifying the category's attributes when there are no labeled exemplars available. The fundamental idea for ZSR is to transfer knowledge from the abundant labeled data in different but related source classes via the class attributes. Conventional ZSR approaches adopt a two-step strategy in test stage, where the samples are projected into the attribute space in the first step, and then the recognition is carried out based on considering the relationship between samples and classes in the attribute space. Due to this intermediate transformation, information loss is unavoidable, thus degrading the performance of the overall system. Rather than following this two-step strategy, in this paper, we propose a novel one-step approach that is able to perform ZSR in the original feature space by using directly trained classifiers. To tackle the problem that no labeled samples of target classes are available, we propose to assign pseudo labels to samples based on the reliability and diversity, which in turn will be used to train the classifiers. Moreover, we adopt a robust SVM that accounts for the unreliability of pseudo labels. Extensive experiments on four datasets demonstrate consistent performance gains of our approach over the state-of-the-art two-step ZSR approaches.
Guiguang Ding, Jungong Han, Yue Gao 0002
AAAI2
2017 Fully Convolutional Neural Networks with Full-Scale-Features for Semantic Segmentation
abstract
In this work, we propose a novel method to involve full-scale-features into the fully convolutional neural networks (FCNs) for Semantic Segmentation. Current works on FCN has brought great advances in the task of semantic segmentation, but the receptive field, which represents region areas of input volume connected to any output neuron, limits the available information of output neuron's prediction accuracy. We investigate how to involve the full-scale or full-image features into FCNs to enrich the receptive field. Specially, the full-scale feature network (FFN) extends the full-connected network and makes an end-to-end unified training structure. It has two appealing properties. First, the introduction of full-scale-features is beneficial for prediction. We build a unified extracting network and explore several fusion functions for concatenating features. Amounts of experiments have been carried out to prove that full-scale-features makes fair accuracy raising. Second, FFN is applicable to many variants of FCN which could be regarded as a general strategy to improve the segmentation accuracy. Our proposed method is evaluated on PASCAL VOC 2012, and achieves a state-of-art result.
Tianxiang Pan, Bin Wang 0021, Guiguang Ding, Jun-Hai Yong
AAAI3
2017 From Zero-Shot Learning to Conventional Supervised Classification: Unseen Visual Data Synthesis
abstract
Robust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new classes is unattainable. In this paper, we propose a new Zero-shot learning (ZSL) framework that can synthesise visual features for unseen classes without acquiring real images. Using the proposed Unseen Visual Data Synthesis (UVDS) algorithm, semantic attributes are effectively utilised as an intermediate clue to synthesise unseen visual features at the training stage. Hereafter, ZSL recognition is converted into the conventional supervised problem, i.e. the synthesised visual features can be straightforwardly fed to typical classifiers such as SVM. On four benchmark datasets, we demonstrate the benefit of using synthesised unseen data. Extensive experimental results manifest that our proposed approach significantly improve the state-of-the-art results.
Yang Long 0001, Li Liu 0004, Ling Shao 0001, Fumin Shen, Guiguang Ding, Jungong Han
CVPR5
2017 SitNet: Discrete Similarity Transfer Network for Zero-shot Hashing
abstract
Hashing has been widely utilized for fast image retrieval recently. With semantic information as supervision, hashing approaches perform much better, especially when combined with deep convolution neural network(CNN). However, in practice, new concepts emerge every day, making collecting supervised information for re-training hashing model infeasible. In this paper, we propose a novel zero-shot hashing approach, called Discrete Similarity Transfer Network (SitNet), to preserve the semantic similarity between images from both ``seen'' concepts and new ``unseen'' concepts. Motivated by zero-shot learning, the semantic vectors of concepts are adopted to capture the similarity structures among classes, making the model trained with seen concepts generalize well for unseen ones benefiting from the transferability of the semantic vector space. We adopt a multi-task architecture to exploit the supervised information for seen concepts and the semantic vectors simultaneously. Moreover, a discrete hashing layer is integrated into the network for hashcode generating to avoid the information loss caused by real-value relaxation in training phase, which is a critical problem in existing works. Experiments on three benchmarks validate the superiority of SitNet to the state-of-the-arts.
Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI2
2017 Synthesizing Samples for Zero-shot Learning
abstract
Zero-shot learning (ZSL) is to construct recognition models for unseen target classes that have no labeled samples for training. It utilizes the class attributes or semantic vectors as side information and transfers supervision information from related source classes with abundant labeled samples. Existing ZSL approaches adopt an intermediary embedding space to measure the similarity between a sample and the attributes of a target class to perform zero-shot classification. However, this way may suffer from the information loss caused by the embedding process and the similarity measure cannot fully make use of the data distribution. In this paper, we propose a novel approach which turns the ZSL problem into a conventional supervised learning problem by synthesizing samples for the unseen classes. Firstly, the probability distribution of an unseen class is estimated by using the knowledge from seen classes and the class attributes. Secondly, the samples are synthesized based on the distribution for the unseen class. Finally, we can train any supervised classifiers based on the synthesized samples. Extensive experiments on benchmarks demonstrate the superiority of the proposed approach to the state-of-the-art ZSL approaches.
Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI2
2017 Unsupervised Deep Video Hashing with Balanced Rotation
abstract
Recently, hashing video contents for fast retrieval has received increasing attention due to the enormous growth of online videos. As the extension of image hashing techniques, traditional video hashing methods mainly focus on seeking the appropriate video features but pay little attention to how the video-specific features can be leveraged to achieve optimal binarization. In this paper, an end-to-end hashing framework, namely Unsupervised Deep Video Hashing (UDVH), is proposed, where feature extraction, balanced code learning and hash function learning are integrated and optimized in a self-taught manner. Particularly, distinguished from previous work, our framework enjoys two novelties: 1) an unsupervised hashing method that integrates the feature clustering and feature binarization, enabling the neighborhood structure to be preserved in the binary space; 2) a smart rotation applied to the video-specific features that are widely spread in the low-dimensional space such that the variance of dimensions can be balanced, thus generating more effective hash codes. Extensive experiments have been performed on two real-world datasets and the results demonstrate its superiority, compared to the state-of-the-art video hashing methods. To bootstrap further developments, the source code will be made publically available.
Gengshen Wu, Li Liu 0004, Guiguang Ding, Jungong Han, Jialie Shen 0001, Ling Shao 0001
IJCAI4
2017 Approximating Discrete Probability Distribution of Image Emotions by Multi-Modal Features Fusion
abstract
Existing works on image emotion recognition mainly assigned the dominant emotion category or average dimension values to an image based on the assumption that viewers can reach a consensus on the emotion of images. However, the image emotions perceived by viewers are subjective by nature and highly related to the personal and situational factors. On the other hand, image emotions can be conveyed by different features, such as semantics and aesthetics. In this paper, we propose a novel machine learning approach that formulates the categorical image emotions as a discrete probability distribution (DPD). To associate emotions with the extracted visual features, we present a weighted multi-modal shared sparse leaning to learn the combination coefficients, with which the DPD of an unseen image can be predicted by linearly integrating the DPDs of the training images. The representation abilities of different modalities are jointly explored and the optimal weight of each modality is automatically learned. Extensive experiments on three datasets verify the superiority of the proposed method, as compared to the state-of-the-art.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Jungong Han
IJCAI2
2017 TUCH: Turning Cross-view Hashing into Single-view Hashing via Generative Adversarial Nets
abstract
Cross-view retrieval, which focuses on searching images as response to text queries or vice versa, has received increasing attention recently. Cross-view hashing is to efficiently solve the cross-view retrieval problem with binary hash codes. Most existing works on cross-view hashing exploit multi-view embedding method to tackle this problem, which inevitably causes the information loss in both image and text domains. Inspired by the Generative Adversarial Nets (GANs), this paper presents a new model that is able to Turn Cross-view Hashing into single-view hashing (TUCH), thus enabling the information of image to be preserved as much as possible. TUCH is a novel deep architecture that integrates a language model network T for text feature extraction, a generator network G to generate fake images from text feature and a hashing network H for learning hashing functions to generate compact binary codes. Our architecture effectively unifies joint generative adversarial learning and cross-view hashing. Extensive empirical evidence shows that our TUCH approach achieves state-of-the-art results, especially on text to image retrieval, based on image-sentences datasets, i.e. standard IAPRTC-12 and large-scale Microsoft COCO.
Xin Zhao 0020, Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI2
2017 Learning Visual Emotion Distributions via Multi-Modal Features Fusion
abstract
Current image emotion recognition works mainly classified the images into one dominant emotion category, or regressed the images with average dimension values by assuming that the emotions perceived among different viewers highly accord with each other. However, due to the influence of various personal and situational factors, such as culture background and social interactions, different viewers may react totally different from the emotional perspective to the same image. In this paper, we propose to formulate the image emotion recognition task as a probability distribution learning problem. Motivated by the fact that image emotions can be conveyed through different visual features, such as aesthetics and semantics, we present a novel framework by fusing multi-modal features to tackle this problem. In detail, weighted multi-modal conditional probability neural network (WMMCPNN) is designed as the learning model to associate the visual features with emotion probabilities. By jointly exploring the complementarity and learning the optimal combination coefficients of different modality features, WMMCPNN could effectively utilize the representation ability of each uni-modal feature. We conduct extensive experiments on three publicly available benchmarks and the results demonstrate that the proposed method significantly outperforms the state-of-the-art approaches for emotion distribution prediction.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Jungong Han
ACM Multimedia2
2017 Attribute-based supervised deep learning model for action recognition
Kai Chen 0044, Guiguang Ding, Jungong Han
Frontiers Comput. Sci.2
2017 Large-scale image retrieval with Sparse Embedded Hashing
Guiguang Ding, Jile Zhou, Zijia Lin, Sicheng Zhao, Jungong Han
Neurocomputing1
2017 Accelerated Manhattan hashing via bit-remapping with location information
Wenshuo Chen, Guiguang Ding, Zijia Lin, Iyad Jafar, Jisheng Pei
Multim. Tools Appl.2
2017 Query expansion for object retrieval with active learning using BoW and CNN feature
Xin Zhao 0020, Guiguang Ding
Multim. Tools Appl.2
2017 Cross-View Retrieval via Probability-Based Semantics-Preserving Hashing
abstract
For efficiently retrieving nearest neighbors from large-scale multiview data, recently hashing methods are widely investigated, which can substantially improve query speeds. In this paper, we propose an effective probability-based semantics-preserving hashing (SePH) method to tackle the problem of cross-view retrieval. Considering the semantic consistency between views, SePH generates one unified hash code for all observed views of any instance. For training, SePH first transforms the given semantic affinities of training data into a probability distribution, and aims to approximate it with another one in Hamming space, via minimizing their Kullback-Leibler divergence. Specifically, the latter probability distribution is derived from all pair-wise Hamming distances between to-be-learnt hash codes of the training data. Then with learnt hash codes, any kind of predictive models like linear ridge regression, logistic regression, or kernel logistic regression, can be learnt as hash functions in each view for projecting the corresponding view-specific features into hash codes. As for out-of-sample extension, given any unseen instance, the learnt hash functions in its observed views can predict view-specific hash codes. Then by deriving or estimating the corresponding output probabilities with respect to the predicted view-specific hash codes, a novel probabilistic approach is further proposed to utilize them for determining a unified hash code. To evaluate the proposed SePH, we conduct extensive experiments on diverse benchmark datasets, and the experimental results demonstrate that SePH is reasonable and effective.
Zijia Lin, Guiguang Ding, Jungong Han, Jianmin Wang 0001
IEEE Trans. Cybern.2
2017 Zero-Shot Learning With Transferred Samples
abstract
By transferring knowledge from the abundant labeled samples of known source classes, zero-shot learning (ZSL) makes it possible to train recognition models for novel target classes that have no labeled samples. Conventional ZSL approaches usually adopt a two-step recognition strategy, in which the test sample is projected into an intermediary space in the first step, and then the recognition is carried out by considering the similarity between the sample and target classes in the intermediary space. Due to this redundant intermediate transformation, information loss is unavoidable, thus degrading the performance of overall system. Rather than adopting this two-step strategy, in this paper, we propose a novel one-step recognition framework that is able to perform recognition in the original feature space by using directly trained classifiers. To address the lack of labeled samples for training supervised classifiers for the target classes, we propose to transfer samples from source classes with pseudo labels assigned, in which the transferred samples are selected based on their transferability and diversity. Moreover, to account for the unreliability of pseudo labels of transferred samples, we modify the standard support vector machine formulation such that the unreliable positive samples can be recognized and suppressed in the training phase. The entire framework is fairly general with the possibility of further extensions to several common ZSL settings. Extensive experiments on four benchmark data sets demonstrate the superiority of the proposed framework, compared with the state-of-the-art approaches, in various settings.
Guiguang Ding, Jungong Han, Yue Gao 0002
IEEE Trans. Image Process.2
2017 Learning to Hash With Optimized Anchor Embedding for Scalable Retrieval
abstract
Sparse representation and image hashing are powerful tools for data representation and image retrieval respectively. The combinations of these two tools for scalable image retrieval, i.e., sparse hashing (SH) methods, have been proposed in recent years and the preliminary results are promising. The core of those methods is a scheme that can efficiently embed the (high-dimensional) image features into a low-dimensional Hamming space, while preserving the similarity between features. Existing SH methods mostly focus on finding better sparse representations of images in the hash space. We argue that the anchor set utilized in sparse representation is also crucial, which was unfortunately underestimated by the prior art. To this end, we propose a novel SH method that optimizes the integration of the anchors, such that the features can be better embedded and binarized, termed as Sparse Hashing with Optimized Anchor Embedding. The central idea is to push the anchors far from the axis while preserving their relative positions so as to generate similar hashcodes for neighboring features. We formulate this idea as an orthogonality constrained maximization problem and an efficient and novel optimization framework is systematically exploited. Extensive experiments on five benchmark image data sets demonstrate that our method outperforms several state-of-the-art related methods.
Guiguang Ding, Li Liu 0004, Jungong Han, Ling Shao 0001
IEEE Trans. Image Process.2
2017 Sequential Discrete Hashing for Scalable Cross-Modality Similarity Retrieval
abstract
With the dramatic development of the Internet, how to exploit large-scale retrieval techniques for multimodal web data has become one of the most popular but challenging problems in computer vision and multimedia. Recently, hashing methods are used for fast nearest neighbor search in large-scale data spaces, by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. Inspired by this, in this paper, we introduce a novel supervised cross-modality hashing framework, which can generate unified binary codes for instances represented in different modalities. Particularly, in the learning phase, each bit of a code can be sequentially learned with a discrete optimization scheme that jointly minimizes its empirical loss based on a boosting strategy. In a bitwise manner, hash functions are then learned for each modality, mapping the corresponding representations into unified hash codes. We regard this approach as cross-modality sequential discrete hashing (CSDH), which can effectively reduce the quantization errors arisen in the oversimplified rounding-off step and thus lead to high-quality binary codes. In the test phase, a simple fusion scheme is utilized to generate a unified hash code for final retrieval by merging the predicted hashing results of an unseen instance from different modalities. The proposed CSDH has been systematically evaluated on three standard data sets: Wiki, MIRFlickr, and NUS-WIDE, and the results show that our method significantly outperforms the state-of-the-art multimodality hashing techniques.
Li Liu 0004, Zijia Lin, Ling Shao 0001, Fumin Shen, Guiguang Ding, Jungong Han
IEEE Trans. Image Process.5
2017 Continuous Probability Distribution Prediction of Image Emotions via Multitask Shared Sparse Regression
abstract
Previous works on image emotion analysis mainly focused on predicting the dominant emotion category or the average dimension values of an image for affective image classification and regression. However, this is often insufficient in various real-world applications, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the continuous probability distribution of image emotions which are represented in dimensional valence-arousal space. We carried out large-scale statistical analysis on the constructed Image-Emotion-Social-Net dataset, on which we observed that the emotion distribution can be well-modeled by a Gaussian mixture model. This model is estimated by an expectation-maximization algorithm with specified initializations. Then, we extract commonly used emotion features at different levels for each image. Finally, we formalize the emotion distribution prediction task as a shared sparse regression (SSR) problem and extend it to multitask settings, named multitask shared sparse regression (MTSSR), to explore the latent information between different prediction tasks. SSR and MTSSR are optimized by iteratively reweighted least squares. Experiments are conducted on the Image-Emotion-Social-Net dataset with comparisons to three alternative baselines. The quantitative results demonstrate the superiority of the proposed method.
Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Guiguang Ding
IEEE Trans. Multim.5
2016 Transductive Zero-Shot Recognition via Shared Model Space Learning
abstract
Zero-shot Recognition (ZSR) is to learn recognition models for novel classes without labeled data. It is a challenging task and has drawn considerable attention in recent years. The basic idea is to transfer knowledge from seen classes via the shared attributes. This paper focus on the transductive ZSR, i.e., we have unlabeled data for novel classes. Instead of learning models for seen and novel classes separately as in existing works, we put forward a novel joint learning approach which learns the shared model space (SMS) for models such that the knowledge can be effectively transferred between classes using the attributes. An effective algorithm is proposed for optimization. We conduct comprehensive experiments on three benchmark datasets for ZSR. The results demonstrates that the proposed SMS can significantly outperform the state-of-the-art related approaches which validates its efficacy for the ZSR task.
Guiguang Ding, Xiaoming Jin, Jianmin Wang 0001
AAAI2
2016 Active Learning with Cross-Class Knowledge Transfer
abstract
When there are insufficient labeled samples for training a supervised model, we can adopt active learning to select the most informative samples for human labeling, or transfer learning to transfer knowledge from related labeled data source. Combining transfer learning with active learning has attracted much research interest in recent years. Most existing works follow the setting where the class labels in source domain are the same as the ones in target domain. In this paper, we focus on a more challenging cross-class setting where the class labels are totally different in two domains but related to each other in an intermediary attribute space, which is barely investigated before. We propose a novel and effective method that utilizes the attribute representation as the seed parameters to generate the classification models for classes. And we propose a joint learning framework that takes into account the knowledge from the related classes in source domain, and the information in the target domain. Besides, it is simple to perform uncertainty sampling, a fundamental technique for active learning, based on the framework. We conduct experiments on three benchmark datasets and the results demonstrate the efficacy of the proposed method.
Guiguang Ding, Xiaoming Jin
AAAI2
2016 Multi-Domain Active Learning for Recommendation
abstract
Recently, active learning has been applied to recommendation to deal with data sparsity on a single domain. In this paper, we propose an active learning strategy for recommendation to alleviate the data sparsity in a multi-domain scenario. Specifically, our proposed active learning strategy simultaneously consider both specific and independent knowledge over all domains. We use the expected entropy to measure the generalization error of the domain-specific knowledge and propose a variance-based strategy to measure the generalization error of the domain-independent knowledge. The proposed active learning strategy use a unified function to effectively combine these two measurements. We compare our strategy with five state-of-the-art baselines on five different multi-domain recommendation tasks, which are constituted by three real-world data sets. The experimental results show that our strategy performs significantly better than all the baselines and reduces human labeling efforts by at least 5.6%, 8.3%, 11.8%, 12.5% and 15.4% on the five tasks, respectively.
Xiaoming Jin, Lianghao Li, Guiguang Ding, Qiang Yang 0001
AAAI4
2016 Semi-Supervised Active Learning with Cross-Class Sample Transfer
Guiguang Ding, Yue Gao 0002, Jianmin Wang 0001
IJCAI2
2016 Robust Iterative Quantization for Efficient ℓp-norm Similarity Search
Guiguang Ding, Jungong Han, Xiaoming Jin
IJCAI2
2016 Image Representation Optimization Based on Locally Aggregated Descriptors
Shijiang Chen, Guiguang Ding, Chenxiao Li
PAKDD (2)2
2016 Large-Scale Cross-Modality Search via Collective Matrix Factorization Hashing
abstract
By transforming data into binary representation, i.e., Hashing, we can perform high-speed search with low storage cost, and thus, Hashing has collected increasing research interest in the recent years. Recently, how to generate Hashcode for multimodal data (e.g., images with textual tags, documents with photos, and so on) for large-scale cross-modality search (e.g., searching semantically related images in database for a document query) is an important research issue because of the fast growth of multimodal data in the Web. To address this issue, a novel framework for multimodal Hashing is proposed, termed as Collective Matrix Factorization Hashing (CMFH). The key idea of CMFH is to learn unified Hashcodes for different modalities of one multimodal instance in the shared latent semantic space in which different modalities can be effectively connected. Therefore, accurate cross-modality search is supported. Based on the general framework, we extend it in the unsupervised scenario where it tries to preserve the Euclidean structure, and in the supervised scenario where it fully exploits the label information of data. The corresponding theoretical analysis and the optimization algorithms are given. We conducted comprehensive experiments on three benchmark data sets for cross-modality search. The experimental results demonstrate that CMFH can significantly outperform several state-of-the-art cross-modality Hashing methods, which validates the effectiveness of the proposed CMFH.
Guiguang Ding, Jile Zhou, Yue Gao 0002
IEEE Trans. Image Process.1
2015 Learning Predictable and Discriminative Attributes for Visual Recognition
abstract
Utilizing attributes for visual recognition has attracted increasingly interest because attributes can effectively bridge the semantic gap between low-level visual features and high-level semantic labels. In this paper, we propose a novel method for learning predictable and discriminative attributes. Specifically, we require the learned attributes can be reliably predicted from visual features, and discover the inherent discriminative structure of data. In addition, we propose to exploit the intra-category locality of data to overcome the intra-category variance in visual data. We conduct extensive experiments on Animals with Attributes (AwA) and Caltech256 datasets, and the results demonstrate that the proposed method achieves state-of-the-art performance.
Guiguang Ding, Xiaoming Jin, Jianmin Wang 0001
AAAI2
2015 Gaussian Cardinality Restricted Boltzmann Machines
abstract
Restricted Boltzmann Machine (RBM) has been applied to a wide variety of tasks due to its advantage in feature extraction. Implementing sparsity constraint in the activated hidden units of RBM is an important improvement on RBM. The sparsity constraints in the existing methods are usually specified by users and are independent of the input data. However, the input data could be heterogeneous in content and thus naturally demand elastic and adaptive settings of the sparsity constraints. To solve this problem, we proposed a generalized model with adaptive sparsity constraint, named Gaussian Cardinality Restricted Boltzmann Machines (GC-RBM). In this model, the thresholds of hidden unit activations are decided by the input data and a given Gaussian distribution on the pre-training phase. We provide a principled method to train the GC-RBM with Gaussian prior. Experimental results on two real world data sets justify the effectiveness of the proposed method and its superiority over CaRBM in terms of classification accuracy.
Xiaoming Jin, Guiguang Ding, Dou Shen
AAAI3
2015 Semantics-preserving hashing for cross-view retrieval
abstract
With benefits of low storage costs and high query speeds, hashing methods are widely researched for efficiently retrieving large-scale data, which commonly contains multiple views, e.g. a news report with images, videos and texts. In this paper, we study the problem of cross-view retrieval and propose an effective Semantics-Preserving Hashing method, termed SePH. Given semantic affinities of training data as supervised information, SePH transforms them into a probability distribution and approximates it with to-be-learnt hash codes in Hamming space via minimizing the Kullback-Leibler divergence. Then kernel logistic regression with a sampling strategy is utilized to learn the nonlinear projections from features in each view to the learnt hash codes. And for any unseen instance, predicted hash codes and their corresponding output probabilities from observed views are utilized to determine its unified hash code, using a novel probabilistic approach. Extensive experiments conducted on three benchmark datasets well demonstrate the effectiveness and reasonableness of SePH.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001
CVPR2
2015 Robust Nonnegative Matrix Factorization with Discriminability for image representation
abstract
Due to its psychological and physiological interpretation of naturally occurring data, Nonnegative Matrix Factorization (NMF) has attracted considerable attention for learning effective representation for images. And its graph-regularized extensions have shown promising results by exploiting the low dimensional manifold structure of data. Actually, their performance can be further improved because they still suffer from several important problems, i.e., sensitivity to noise in data, trivial solution problem, and ignoring the discriminative information. In this paper, we propose a novel method, referred to as Robust Nonnegative Matrix Factorization with Discriminability (RNMFD), for image representation, which can effectively and simultaneously cope with problems mentioned above by imposing a sparse noise matrix for data reconstruction and approximate orthogonal constraints. We carried out extensive experiments on five benchmark image datasets and the results demonstrate the superiority of our RNMFD in comparison with several state-of-the-art methods.
Guiguang Ding, Jile Zhou
ICME2
2015 Distribution Regularized Nonnegative Matrix Factorization for Transfer Visual Feature Learning
abstract
Transfer visual feature learning (TVFL), which learns compact representations for images such that we can build accurate classifier for target domain by leveraging rich labeled data in the source domain, has attracted increasingly attention recently. Previous methods mainly focus on reducing the distribution difference between domains but ignore the intrinsic hidden semantics in data. In this paper, we put forward a novel method for TVFL, called Distribution Regularized Nonnegative Matrix Factorization (DRNMF). Specifically, we employ Nonnegative Matrix Factorization (NMF) to uncover the intrinsic information in visual data, and regularize it with geometrical distribution, marginal probability distribution and conditional probability distribution. Thus, DRNMF can discover the intrinsic information, preserve the manifold structure and reducing both marginal and conditional probability distribution difference simultaneously, which all perspectives above are important for TVFL. We also propose an effective and efficient algorithm for the optimization of DRNMF and theoretically prove the convergence. Extensive experiments on three types of cross-domain image classification tasks in comparison with several state-of-the-art methods demonstrate the superiority of our DRNMF, which validates its effectiveness.
Guiguang Ding, Qiang Liu 0016
ICMR2
2015 Robust and Discriminative Concept Factorization for Image Representation
abstract
Concept Factorization (CF), as a variant of Nonnegative Matrix Factorization (NMF), has been widely used for learning compact representation for images because of its psychological and physiological interpretation of naturally occurring data. And graph regularization has been incorporated into the objective function of CF to exploit the intrinsic low-dimensional manifold structure, leading to better performance. But some shortcomings are shared by existing CF methods. 1) The squared loss used to measure the data reconstruction quality is sensitive to noise in image data. 2) The graph regularization may lead to trivial solution and scale transfer problems for CF such that the learned representation is meaningless. 3) Existing methods mostly ignore the discriminative information in image data. In this paper, we propose a novel method, called Robust and Discriminative Concept Factorization (RDCF) for image representation. Specifically, RDCF explicitly considers the influence of noise by imposing a sparse error matrix, and exploits the discriminative information by approximate orthogonal constraints which can also lead to nontrivial solution. We propose an iterative multiplicative updating rule for the optimization of RDCF and prove the convergence. Experiments on 5 benchmark image datasets show that RDCF can significantly out-perform several state-of-the-art related methods, which validates the effectiveness of RDCF.
Guiguang Ding, Jile Zhou, Qiang Liu 0016
ICMR2
2015 Image auto-annotation via tag-dependent random search over range-constrained visual neighbours
Zijia Lin, Guiguang Ding, Mingqing Hu
Multim. Tools Appl.2
2014 Matrix Factorization Meets Cosine Similarity: Addressing Sparsity Problem in Collaborative Filtering Recommender System
Hailong Wen, Guiguang Ding
APWeb2
2014 Collective Matrix Factorization Hashing for Multimodal Data
abstract
Nearest neighbor search methods based on hashing have attracted considerable attention for effective and efficient large-scale similarity search in computer vision and information retrieval community. In this paper, we study the problems of learning hash functions in the context of multimodal data for cross-view similarity search. We put forward a novel hashing method, which is referred to Collective Matrix Factorization Hashing (CMFH). CMFH learns unified hash codes by collective matrix factorization with latent factor model from different modalities of one instance, which can not only supports cross-view search but also increases the search accuracy by merging multiple view information sources. We also prove that CMFH, a similarity-preserving hashing learning method, has upper and lower boundaries. Extensive experiments verify that CMFH significantly outperforms several state-of-the-art methods on three different datasets.
Guiguang Ding, Jile Zhou
CVPR1
2014 Transfer Joint Matching for Unsupervised Domain Adaptation
abstract
Visual domain adaptation, which learns an accurate classifier for a new domain using labeled images from an old domain, has shown promising value in computer vision yet still been a challenging problem. Most prior works have explored two learning strategies independently for domain adaptation: feature matching and instance reweighting. In this paper, we show that both strategies are important and inevitable when the domain difference is substantially large. We therefore put forward a novel Transfer Joint Matching (TJM) approach to model them in a unified optimization problem. Specifically, TJM aims to reduce the domain difference by jointly matching the features and reweighting the instances across domains in a principled dimensionality reduction procedure, and construct new feature representation that is invariant to both the distribution difference and the irrelevant instances. Comprehensive experimental results verify that TJM can significantly outperform competitive methods for cross-domain image recognition problems.
Mingsheng Long, Jianmin Wang 0001, Guiguang Ding, Jia-Guang Sun 0001, Philip S. Yu
CVPR3
2014 CosSimReg: An Effective Transfer Learning Method in Social Recommender System
Hailong Wen, Guiguang Ding, Qiang Liu 0016
ICIC (1)3
2014 Kernel-based supervised hashing for cross-view similarity search
abstract
Spectral-based hashing (SpH) is the most used method for cross-view hash function learning (CVHFL). However, the following three problems are shared by many existing SpH methods. Firstly, preserving intra- and inter-similarity simultaneously increases models' complexity significantly. Secondly, linear model applied in many SpH methods is hard to handle multimodal data in cross-view scenarios. Thirdly, to learn irrelevant multiple bits, SpH imposes orthogonality constraints which decreases the mapping quality substantially with the increase of bit number. To address these challenges, we propose a novel SpH method for CVHFL in this paper, referred to as Kernel-based Supervised Hashing for Cross-view Similarity Search (KSH-CV). We prove that the intra-adjacency matrix is redundant given inter-adjacency matrix. Then we define our objective function in a supervised and k-ernelized way which just needs to preserve inter-similarity. Furthermore a novel Adaboost algorithm, which minimizes exponential mapping loss function for cross-view similarity search, is derived to solve the objective function efficiently while avoiding orthogonality constraints. Extensive experiments verifies that KSH-CV can significantly outperform several state-of-the-art methods on three cross-view datasets.
Jile Zhou, Guiguang Ding, Qiang Liu 0016, XinPeng Dong
ICME2
2014 Multi-label Classification via Feature-aware Implicit Label Space Encoding
abstract
To tackle a multi-label classification problem with many classes, recently label space dimension reduction (LSDR) is proposed. It encodes the original label space to a low-dimensional latent space and uses a decoding process for recovery. In this paper, we propose a novel method termed FaIE to perform LSDR via Feature-aware Implicit label space Encoding. Unlike most previous work, the proposed FaIE makes no assumptions about the encoding process and directly learns a code matrix, i.e. the encoding result of some implicit encoding function, and a linear decoding matrix. To learn both matrices, FaIE jointly maximizes the recoverability of the original label space from the latent space, and the predictability of the latent space from the feature space, thus making itself feature-aware. FaIE can also be specified to learn an explicit encoding function, and extended with kernel tricks to handle non-linear correlations between the feature space and the latent space. Extensive experiments conducted on benchmark datasets well demonstrate its effectiveness.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001
ICML2
2014 Latent semantic sparse hashing for cross-modal similarity search
abstract
Similarity search methods based on hashing for effective and efficient cross-modal retrieval on large-scale multimedia databases with massive text and images have attracted considerable attention. The core problem of cross-modal hashing is how to effectively construct correlation between multi-modal representations which are heterogeneous intrinsically in the process of hash function learning. Analogous to Canonical Correlation Analysis (CCA), most existing cross-modal hash methods embed the heterogeneous data into a joint abstraction space by linear projections. However, these methods fail to bridge the semantic gap more effectively, and capture high-level latent semantic information which has been proved that it can lead to better performance for image retrieval. To address these challenges, in this paper, we propose a novel Latent Semantic Sparse Hashing (LSSH) to perform cross-modal similarity search by employing Sparse Coding and Matrix Factorization. In particular, LSSH uses Sparse Coding to capture the salient structures of images, and Matrix Factorization to learn the latent concepts from text. Then the learned latent semantic features are mapped to a joint abstraction space. Moreover, an iterative strategy is applied to derive optimal solutions efficiently, and it helps LSSH to explore the correlation between multi-modal representations efficiently and automatically. Finally, the unified hashcodes are generated through the high level abstraction space by quantization. Extensive experiments on three different datasets highlight the advantage of our method under cross-modal scenarios and show that LSSH significantly outperforms several state-of-the-art methods.
Jile Zhou, Guiguang Ding
SIGIR2
2014 Image tag completion via dual-view linear sparse reconstructions
Zijia Lin, Guiguang Ding, Mingqing Hu, Yunzhen Lin, Shuzhi Sam Ge
Comput. Vis. Image Underst.2
2014 Transfer Learning with Graph Co-Regularization
abstract
Transfer learning is established as an effective technology to leverage rich labeled data from some source domain to build an accurate classifier for the target domain. The basic assumption is that the input domains may share certain knowledge structure, which can be encoded into common latent factors and extracted by preserving important property of original data, e.g., statistical property and geometric structure. In this paper, we show that different properties of input data can be complementary to each other and exploring them simultaneously can make the learning model robust to the domain difference. We propose a general framework, referred to as Graph Co-Regularized Transfer Learning (GTL), where various matrix factorization models can be incorporated. Specifically, GTL aims to extract common latent factors for knowledge transfer by preserving the statistical property across domains, and simultaneously, refine the latent factors to alleviate negative transfer by preserving the geometric structure in each domain. Based on the framework, we propose two novel methods using NMF and NMTF, respectively. Extensive experiments verify that GTL can significantly outperform state-of-the-art learning methods on several public text and image datasets.
Mingsheng Long, Jianmin Wang 0001, Guiguang Ding, Dou Shen, Qiang Yang 0001
IEEE Trans. Knowl. Data Eng.3
2014 Adaptation Regularization: A General Framework for Transfer Learning
abstract
Domain transfer learning, which learns a target classifier using labeled data from a different distribution, has shown promising value in knowledge discovery yet still been a challenging problem. Most previous works designed adaptive classifiers by exploring two learning strategies independently: distribution adaptation and label propagation. In this paper, we propose a novel transfer learning framework, referred to as Adaptation Regularization based Transfer Learning (ARTL), to model them in a unified way based on the structural risk minimization principle and the regularization theory. Specifically, ARTL learns the adaptive classifier by simultaneously optimizing the structural risk functional, the joint distribution matching between domains, and the manifold consistency underlying marginal distribution. Based on the framework, we propose two novel methods using Regularized Least Squares (RLS) and Support Vector Machines (SVMs), respectively, and use the Representer theorem in reproducing kernel Hilbert space to derive corresponding solutions. Comprehensive experiments verify that ARTL can significantly outperform state-of-the-art learning methods on several public text and image datasets.
Mingsheng Long, Jianmin Wang 0001, Guiguang Ding, Sinno Jialin Pan, Philip S. Yu
IEEE Trans. Knowl. Data Eng.3
2013 Image Tag Completion via Image-Specific and Tag-Specific Linear Sparse Reconstructions
abstract
Though widely utilized for facilitating image management, user-provided image tags are usually incomplete and insufficient to describe the whole semantic content of corresponding images, resulting in performance degradations in tag-dependent applications and thus necessitating effective tag completion methods. In this paper, we propose a novel scheme denoted as LSR for automatic image tag completion via image-specific and tag-specific Linear Sparse Reconstructions. Given an incomplete initial tagging matrix with each row representing an image and each column representing a tag, LSR optimally reconstructs each image (i.e. row) and each tag (i.e. column) with remaining ones under constraints of sparsity, considering image-image similarity, image-tag association and tag-tag concurrence. Then both image-specific and tag-specific reconstruction values are normalized and merged for selecting missing related tags. Extensive experiments conducted on both benchmark dataset and web images well demonstrate the effectiveness of the proposed LSR.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001, Xiaojun Ye 0001
CVPR2
2013 Transfer Sparse Coding for Robust Image Representation
abstract
Sparse coding learns a set of basis functions such that each input signal can be well approximated by a linear combination of just a few of the bases. It has attracted increasing interest due to its state-of-the-art performance in BoW based image representation. However, when labeled and unlabeled images are sampled from different distributions, they may be quantized into different visual words of the codebook and encoded with different representations, which may severely degrade classification performance. In this paper, we propose a Transfer Sparse Coding (TSC) approach to construct robust sparse representations for classifying cross-distribution images accurately. Specifically, we aim to minimize the distribution divergence between the labeled and unlabeled images, and incorporate this criterion into the objective function of sparse coding to make the new representations robust to the distribution difference. Experiments show that TSC can significantly outperform state-of-the-art methods on three types of computer vision datasets.
Mingsheng Long, Guiguang Ding, Jianmin Wang 0001, Jia-Guang Sun 0001, Philip S. Yu
CVPR2
2013 Transfer Feature Learning with Joint Distribution Adaptation
abstract
Transfer learning is established as an effective technology in computer vision for leveraging rich labeled data in the source domain to build an accurate classifier for the target domain. However, most prior methods have not simultaneously reduced the difference in both the marginal distribution and conditional distribution between domains. In this paper, we put forward a novel transfer learning approach, referred to as Joint Distribution Adaptation (JDA). Specifically, JDA aims to jointly adapt both the marginal distribution and conditional distribution in a principled dimensionality reduction procedure, and construct new feature representation that is effective and robust for substantial distribution difference. Extensive experiments verify that JDA can significantly outperform several state-of-the-art methods on four types of cross-domain image classification problems.
Mingsheng Long, Jianmin Wang 0001, Guiguang Ding, Jia-Guang Sun 0001, Philip S. Yu
ICCV3
2013 Multi-source image auto-annotation
abstract
Though the field of image auto-annotation has been extensively researched, most previous work concentrated on the single-source problem, assuming that both labelled and unseen to-be-annotated images are from a single source (e.g. an identical website), while in practice they are generally collected from multiple sources (e.g. different websites). In that case, treating each source independently may suffer from the insufficiency of labelled data for model training, while merging with labelled images from other sources can bring risky biases to the source-specific model. In this paper, we propose a multi-task learning model to alleviate the multi-source image auto-annotation problem, with each task defined as performing auto-annotation for the corresponding source. Specifically, the proposed model trains annotation models for all sources in parallel with the introduction of inter-source structure regularizers and parameter constraints for sharing information and enhancing the overall performance. Experiments conducted on three different-source benchmark datasets and their combinations yield inspiring results and demonstrate that the proposed model can well utilize the shared information and relieve the risky biases.
Zijia Lin, Guiguang Ding, Mingqing Hu
ICIP2
2013 Twin Bridge Transfer Learning for Sparse Collaborative Filtering
Jiangfeng Shi, Mingsheng Long, Qiang Liu 0016, Guiguang Ding, Jianmin Wang 0001
PAKDD (1)4
2013 Generating virtual ratings from chinese reviews to augment online recommendations
abstract
Collaborative filtering (CF) recommenders based on User-Item rating matrix as explicitly obtained from end users have recently appeared promising in recommender systems. However, User-Item rating matrix is not always available or very sparse in some web applications, which has critical impact to the application of CF recommenders. In this article we aim to enhance the online recommender system by fusing virtual ratings as derived from user reviews. Specifically, taking into account of Chinese reviews' characteristics, we propose to fuse the self-supervised emotion-integrated sentiment classification results into CF recommenders, by which the User-Item Rating Matrix can be inferred by decomposing item reviews that users gave to the items. The main advantage of this approach is that it can extend CF recommenders to some web applications without user rating information. In the experiments, we have first identified the self-supervised sentiment classification's higher precision and recall by comparing it with traditional classification methods. Furthermore, the classification results, as behaving as virtual ratings, were incorporated into both user-based and item-based CF algorithms. We have also conducted an experiment to evaluate the proximity between the virtual and real ratings and clarified the effectiveness of the virtual ratings. The experimental results demonstrated the significant impact of virtual ratings on increasing system's recommendation accuracy in different data conditions (i.e., conditions with real ratings and without).
Weishi Zhang, Guiguang Ding, Li Chen 0009, Chunping Li
ACM Trans. Intell. Syst. Technol.2
2012 Transfer Learning with Graph Co-Regularization
abstract
Transfer learning proves to be effective for leveraging labeled data in the source domain to build an accurate classifier in the target domain. The basic assumption behind transfer learning is that the involved domains share some common latent factors. Previous methods usually explore these latent factors by optimizing two separate objective functions, i.e., either maximizing the empirical likelihood, or preserving the geometric structure. Actually, these two objective functions are complementary to each other and optimizing them simultaneously can make the solution smoother and further improve the accuracy of the final model. In this paper, we propose a novel approach called Graph co-regularized Transfer Learning (GTL) for this purpose, which integrates the two objective functions seamlessly into one unified optimization problem. Thereafter, we present an iterative algorithm for the optimization problem with rigorous analysis on convergence and complexity. Our empirical study on two open data sets validates that GTL can consistently improve the classification accuracy compared to the state-of-the-art transfer learning methods.
Mingsheng Long, Jianmin Wang 0001, Guiguang Ding, Dou Shen, Qiang Yang 0001
AAAI3
2012 Automatic image annotation using tag-related random search over visual neighbors
abstract
In this paper, we propose a novel image auto-annotation model using tag-related random search over range-constrained visual neighbors of the to-be-annotated image. The proposed model, termed as TagSearcher, observes that the annotating performances of many previous visual-neighbor-based models are generally sensitive to the quantity setting of visual neighbors, and the probabilities for visual neighbors to be selected is better to be tag-dependent, meaning that each candidate tag can have its own trustworthy part of visual neighbors for score prediction. And thus TagSearcher uses a constrained range rather than an identical and fixed number of visual neighbors for auto-annotation. By performing a novel tag-related random search process over the graphical model made up of range-constrained visual neighbors, TagSearcher can find the trustworthy part for each candidate tag, and further utilize both visual similarities and tag correlations for score prediction. With the range constraint for visual neighbors and the tag-related random search process, TagSearcher can not only achieve satisfactory annotating performances, but also reduce the performance sensitivity. Experiments conducted on benchmark Corel5k well demonstrate its rationality and effectiveness.
Zijia Lin, Guiguang Ding, Mingqing Hu, Jianmin Wang 0001, Jia-Guang Sun 0001
CIKM2
2012 Dual Transfer Learning
abstract
Transfer learning aims to leverage the knowledge in the source domain to facilitate the learning tasks in the target domain. It has attracted extensive research interests recently due to its effectiveness in a wide range of applications. The general idea of the existing methods is to utilize the common latent structure shared across domains as the bridge for knowledge transfer. These methods usually model the common latent structure by using either the marginal distribution or the conditional distribution. However, without exploring the duality between these two distributions, these single bridge methods may not achieve optimal capability of knowledge transfer. In this paper, we propose a novel approach, Dual Transfer Learning (DTL), which simultaneously learns the marginal and conditional distributions, and exploits the duality between them in a principled way. The key idea behind DTL is that learning one distribution can help to learn the other. This duality property leads to mutual reinforcement when adapting both distributions across domains to transfer knowledge. The proposed method is formulated as an optimization problem based on joint nonnegative matrix trifactorizations (NMTF). The two distributions are learned from the decomposed latent factors that exhibit the duality property. An efficient alternating minimization algorithm is developed to solve the optimization problem with convergence guarantee. Extensive experimental results demonstrate that DTL is more effective than alternative transfer learning methods.
Mingsheng Long, Jianmin Wang 0001, Guiguang Ding, Wei Cheng 0002, Xiang Zhang 0001, Wei Wang 0010
SDM3
2011 Image annotation based on recommendation model
abstract
In this paper, a novel approach based on recommendation model is proposed for automatic image annotation. For any to-be-annotated image, we first select some related images with tags from training dataset according to their visual similarity. And then we estimate the initial ratings for tags of the training images based on tag ranking method and construct a rating matrix. We also construct a trust matrix based on visual similarity with a k-NN strategy. Then a recommendation model is built on both matrices to rank candidate tags for the target image. The proposed approach is evaluated using two benchmark image datasets, and experimental results have indicated its effectiveness.
Zijia Lin, Guiguang Ding, Jianmin Wang 0001
SIGIR2
2010 Automatic semantic annotation of images based on Web data
abstract
Image annotation is a promising approach to bridging the semantic gap between low-level features and high-level concepts, and it can avoid the heavy manual labor. Most existing automatic image annotation approaches are based on supervised learning. They often encounter several problems, such as insufficiency of training data, lack of ability in dealing with new concept, and a limited number of semantic concepts. Web images are massive, rich information, customized etc. Therefore, Web data is a potential repository to provide a sufficient source for semantic annotation. In this paper, we proposed a novel image annotation method based on Web data, which aims to utilize Web data to perform automatic image annotation. Web data, collected from several image search engine, are first preprocessed, clustered and mined to construct a concept clustering model. And then candidate annotation terms are extracted through the model for query image. Afterwards, a rank algorithm is designed to filter out noise terms. Finally, an update phase is implemented to improve the whole method. Evaluations on benchmark image datasets have indicated the effectiveness of our proposal.
Guiguang Ding
IAS1
2010 Ring Fingerprint Based on Interest Points for Video Copy Detection
abstract
In the field of information security, assurance, copyright protection etc., content-based copy detection is more and more important, which consists of two technologies - fingerprint extraction and fingerprint matching. Fingerprints extracted from videos are mainly described as global video fingerprints and local video fingerprints. To make best use of the advantages and bypass the disadvantages of different video fingerprints, we propose the ring fingerprint based on interest points and ordinal measure in this paper. A last we examine the proposed method with lots of experiments and discuss the performance of approaches. The experimental results also demonstrate the effectiveness of the proposed method for video copy detection.
Guiguang Ding, Rongxian Nie
ISM1
2010 A visual word weighting scheme based on emerging itemsets for video annotation
Guiguang Ding, Jianmin Wang 0001
Inf. Process. Lett.1
2009 A New Fingerprint Sequences Matching Algorithm for Content-Based Copy Detection
abstract
Content-based copy detection (CBCD) is more and more important in the field of information security, assurance, copyright protection etc. The sequences matching, one of the key technologies in CBCD, is very critical for querying quickly and accurately. In this paper, we propose a novel fingerprint sequences matching algorithm based on the dynamic programming. This method is universal to all kinds of fingerprint sequences matching. We examined this method with experiments at last in this paper. The experimental results demonstrate it's effective in sequences matching.
Rongxian Nie, Guiguang Ding, Jianmin Wang 0001, Li Zhang 0065
IAS2
2009 Adaptive-Push Peer-to-Peer Video Streaming System Based on Dynamic Unstructured Topology
abstract
This paper presents an adaptive-push scheme which flexibly applies push method in P2P live streaming system based on unstructured topology. In former P2P streaming systems, push method is hardly adopted in loose topology although it transmits data faster than pull method because the peers are randomly distributed and the aimless pushing will incur great redundancy. To improve P2P live streaming system's transmitting ability, the proposed scheme improves a mechanical application of the push method in loose topology by designing an adaptive source peer selection strategy. Experiments prove that the proposed scheme is more efficient and robust.
Guiguang Ding, Jianmin Wang 0001
IAS2
2009 Distributed video coding based on part intracoding and soft side information estimation
Guiguang Ding
Multim. Tools Appl.1
2007 Multi-View Images Coding Based on Multiterminal Source Coding
abstract
In this paper, we proposed a multi-view images coding method based on multiterminal source coding (MSC). Due to separate encoding in MSC, our coding scheme can achieve good random access performance and the spatial redundancy can be exploited even if the encoders can not communicate with each other. Because of joint decoding, the compression performance of our scheme is promising, far better than that of separate encoding and decoding scheme. Compared to multi-view coding based on Wyner Ziv coding, our coding scheme is more flexible and can be easily extended to N views coding. There is no need to classify the view images into key images and Wyner Ziv images, images can be compressed in the same way and we can easily change the compression rate of each view to adapt to the resource conditions, like network bandwidth or storage. Experiment results show that the compression performance of our scheme is better than that of JPEG encoder and decoder scheme.
Qionghai Dai, Guiguang Ding
ICASSP (1)3
2007 A Distributed Video Coding Scheme Based on Denoising Techniques
Guiguang Ding
MMM (1)1
2005 Affine-Invariant Image Retrieval Based on Wavelet Interest Points
abstract
This paper presents an affine-in variant image retrieval approach based on wavelet-based detector, which uses the space-tree property of the transform coefficients to estimate the interest points. Meanwhile, in order to retrieve images compressed by wavelet algorithm such as JPEG2000, the detector only uses the partial bit-planes of the wavelet coefficients to detect the interest points. To provide affine-invariant image matching, annular color histogram, annular texture histogram and spatial cohesion based on interest points are presented to describe image features. A series of experiments based on an image database consisting of 1000 images are performed to confirm the effectiveness of our method
Guiguang Ding, Qionghai Dai, Wenli Xu
MMSP1
2005 Adaptive Key Frame Selection Wyner-Ziv Video Coding
abstract
In Wyner-Ziv video coding, efficient compression is achieved by exploiting source statistics at the decoder only, which is radically different from conventional video coding. The performance of a Wyner-Ziv video codec is greatly dependent on the quality of reconstructed side information, which is an estimation version of current frame. Therefore the correlation between Wyner-Ziv (WZ) frame and key frame implicitly affects the performance of a Wyner-Ziv video codec. In this paper, an adaptive key frame selection method is proposed. Firstly, a simply interest points detector is utilized to detect interest points of a video frame. Then, a kind of interest points based measure, which could represent the correlation between video frames, is performed. If the number of different interest points between two video frames is below a threshold, then it should be transmitted as a key frame, otherwise as a WZ frame. Our experimental results are promising. About 1 dB gain in the quality of reconstructed frames has been achieved
Guiguang Ding, Qionghai Dai, Yaguang Yin
MMSP2
2004 A line-diamond parallel search algorithm for block motion estimation
abstract
The widespread use of block matching motion estimation (BMME) in video coding is due to its effectiveness and simplicity of implementation. This paper presents a novel fast BMME algorithm called the line-diamond parallel search (LDPS). The algorithm is based on the following two properties: the special directionality of the SAD distribution and the characteristics of the center-biased motion vector distribution. In addition, in order to increase the speed of search, the parallel processing idea is used in LDPS. That is to say LDPS realizes the coarse orientation and the accurate search in the same step. Our experimental results show that not only the processing speed of the LDPS algorithm is much higher than that of other fast algorithms, but also its accuracy of motion compensation is as nearly good as that of full search (FS).
Guiguang Ding, Qionghai Dai
ICIG1