VLDB 2026 Research / reviewers in the wild / expert
Bing Li 0001
dblp:13/2692-1
· DBLP profile ↗
161ranked-venue papers
15as first author
104since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 100 · 10 first-author · 68 since 2021Graphics, computer vision, multimedia, augmented reality and games · 97 · 9 first-author · 56 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-authorSecurity and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New ThinkingabstractIn this paper, we rethink model agent behaviors from a geometric structure perspective in multi-agent reinforcement learning. Modeling agent behaviors is essential for understanding how agents interact and facilitating effective decisions. The key lies in capturing the dependencies and sequential relationships among agent decisions. Since each decision influences the subsequent choices, this forms a hierarchical and nested tree-like structure of interdependencies. While modeling tree-like data in Euclidean spaces could cause distortion, which results in a loss of agent decision structure information. Motivated by this, we reconsider model agent behaviors in hyperbolic space and propose the Hyperbolic Multi-Agent Representations (HMAR) method, which projects the agent behaviors into a Poincaré ball and leverages hyperbolic neural networks to learn agent policy representations. Additionally, we designed a contrastive loss function to train this network, minimizing the distance in feature space between different representations of the same agent while maximizing the distance between representations of distinct agents. Experimental results provide empirical evidence for the effectiveness of the HMAR method in cooperative and competitive environments, demonstrating the potential of hyperbolic agent representations for effective decision-making in multi-agent environments. Bohao Qu, Xiaofeng Cao 0002, Bing Li 0001, Menglin Zhang, Tuan-Anh Vu, Di Lin 0002, Qing Guo 0005 |
AAAI | 3 |
| 2026 | MMhops-R1: Multimodal Multi-hop ReasoningabstractThe ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-step reasoning, as existing benchmarks lack the complexity needed to evaluate and drive multi-hop abilities. To bridge this gap, we introduce MMhops, a novel, large-scale benchmark designed to systematically evaluate and foster multi-modal multi-hop reasoning. MMhops dataset comprises two challenging task formats, Bridging and Comparison, which necessitate that models dynamically construct complex reasoning chains by integrating external knowledge. To tackle the challenges posed by MMhops, we propose MMhops-R1, a novel multi-modal Retrieval-Augmented Generation (mRAG) framework for dynamic reasoning. Our framework utilizes reinforcement learning to optimize the model for autonomously planning reasoning paths, formulating targeted queries, and synthesizing multi-level information. Comprehensive experiments demonstrate that MMhops-R1 significantly outperforms strong baselines on MMhops, highlighting that dynamic planning and multi-modal knowledge integration are crucial for complex reasoning. Moreover, MMhops-R1 demonstrates strong generalization to tasks requiring fixed-hop reasoning, underscoring the robustness of our dynamic planning approach. Ziqi Zhang 0010, Zongyang Ma, Bing Li 0001, Chunfeng Yuan, Guangting Wang, Fengyun Rao, Ying Shan, Weiming Hu 0004 |
AAAI | 5 |
| 2026 | Dynamic collaborative evolutionary network: A novel spatio-temporal feature extraction framework for EEG emotion recognition
Shuaiqi Liu 0001, Zhihui Gu, Yanling An, Shuhuan Zhao, Bing Li 0001, Yudong Zhang 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Open-Tag: A Generative Framework for Open-World Multimodal Tagging
Ziqi Zhang 0010, Zongyang Ma, Peijin Wang, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004 |
Int. J. Comput. Vis. | 4 |
| 2026 | Multi-view and spatial-correlation interaction for multi-scale object detection
Yike Yang, Zhaohui Zhu, Zekun Li 0006, Peidong He, Ziqi Zhang 0010, Bing Li 0001 |
Multim. Syst. | 8 |
| 2026 | MSFI: Multi-timescale spatio-temporal features integration in spiking neural networks
Dengfeng Xue, Chunfeng Yuan, Man Yao, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004, Haoliang Sun, Zhetao Li |
Neural Networks | 8 |
| 2026 | Reinforcement Learning-Based Sequential Parameter Tuning for Image Signal ProcessingabstractHardware image signal processing (ISP) transforms RAW inputs into high-quality RGB images through a series of processing modules, each with numerous tunable parameters. Traditionally, these parameters are manually tuned by imaging experts, a time-consuming and subjective process. Recent deep learning approaches predict ISP parameters, but often treat the process as a black box and overlook the intrinsic relationships among ISP modules. To address these fundamental issues, we introduce a novel ISP parameter optimization model based on single-agent reinforcement learning (RL) (i.e., SARL-ISP), formulating the hardware ISP parameter tuning as a sequential optimization problem. During the optimization process, the agent updates ISP parameter tuning strategies for different tasks through interaction with the environment. In order to explore the influence of the sequential structure of hardware ISP modules and the coupling relationships among ISP parameters on the tuning process, we further propose a sequential ISP framework based on collaborative multi-agent RL (i.e., MARL-ISP). Specifically, the serialized parameter tuning module (SPTM) realistically simulates the process of manual prediction and module pipeline. Additionally, the feature selection module (FSM) facilitates the transmission and fusion of agent features, thereby selecting more appropriate feature inputs for downstream tasks. Extensive experiments across various tasks (e.g., object detection, instance segmentation) validate the effectiveness and efficiency of our models. Even with minimal training data, our models also outperform current state-of-the-art methods in both quantitative metrics and qualitative evaluations. Bing Li 0001, Congyan Lang, Zhikun Zhao, Juan Wang 0012, Weihua Xiong, Weiming Hu 0004, Long Cheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Dual -phase transformer with spectral-spatial synergy for hyperspectral image fusion
Shuaiqi Liu 0001, Ruixia Cai, Huanru Yue, Bing Li 0001 |
Pattern Recognit. | 4 |
| 2026 | Multi-modal face anti-spoofing via self-supervised learning
Yufan Liu 0001, Lai Jiang 0004, Shengxi Li, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jinlong Lin |
Pattern Recognit. Lett. | 6 |
| 2026 | Deepfake Detection via Exploring Degradation InconsistencyabstractThe detection of face forgery has become increasingly vital due to the severe security concerns posed by face manipulation techniques. While recent studies on forgery detection have demonstrated promising results when the training and testing samples come from the same domains, the problem remains challenging when attempting to extend the detector to unseen methods. In this work, we propose an innovative approach to enhance the generalization capability of forgery detection methods by exploring degradation inconsistency clues interspersed between the background and the manipulated face regions. Our motivation stems from the observation that digital photos undergo different degradation during acquisition and transmission, resulting in backgrounds and faces from different sources containing distinct degradation patterns in the forged faces. The proposed framework, termed the Degradation Consistency Learning Framework, integrates two core components: a data generation network that modulates degradation transformations to obtain tampered facial images, and a detection network that mines degradation inconsistency clues from both spatial and frequency domains. These two components are tightly coupled through adversarial training, forming a dynamic architecture akin to a Generative Adversarial Network (GAN). Experimental results on different benchmark and evaluation protocols (i.e., indataset and cross-dataset) have demonstrated the effectiveness of our method. Weiming Bai, Yufan Liu 0001, Aixi Zhang, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | Toward Benchmarking and Assessing the Safety and Robustness of Autonomous Driving on Safety-Critical ScenariosabstractAutonomous driving has made significant progress in both academia and industry, including performance improvements in perception tasks and the development of end-to-end autonomous driving systems. However, the safety and robustness assessment of autonomous driving has not received sufficient attention. Current evaluations of autonomous driving are typically conducted in natural driving scenarios. However, accidents often occur in edge cases, also known as safety-critical scenarios. These safety-critical scenarios are difficult to collect, and there is currently no clear definition of what constitutes a safety-critical scenario. In this work, we explore the safety and robustness of autonomous driving in safety-critical scenarios. First, we provide a definition of safety-critical scenarios, including static traffic scenarios such as adversarial attack scenarios and natural distribution shifts, as well as dynamic traffic scenarios such as accident scenarios. Then, we develop an autonomous driving test framework to comprehensively evaluate autonomous driving systems, encompassing not only the assessment of perception modules but also system-level evaluations. Our work systematically constructs a safety verification process for autonomous driving, providing technical support for the industry to establish standardized test framework. Jingzheng Li, Xianglong Liu 0001, Shikui Wei, Yufei Ge, Bing Li 0001, Qing Guo 0005, Xianqi Yang, Yanjun Pu, Qianren Mao, Jiakai Wang |
IEEE Trans. Image Process. | 6 |
| 2025 | Towards More Discriminative Feature Learning in SNNs with Temporal-Self-Erasing SupervisionabstractSpiking Neural Networks (SNNs) are biologically inspired models that process visual inputs over multiple time steps. However, they often struggle with limited feature discrimination along the temporal dimension due to inherent spatiotemporal invariance. This limitation arises from the redundant activation of certain regions and shared supervision for multiple time steps, constraining the network’s ability to adapt and learn diverse features. To address this challenge, we propose a novel Temporal-Self-Erasing (TSE) supervision method that dynamically adapts the learning regions of interest for different time steps. The TSE method operates by identifying highly activated regions from predictions across multiple time steps and adaptively suppressing them during model training, thereby encouraging the network to focus on less activated yet potentially informative regions. This approach not only enhances the feature discrimination capability of SNNs but also facilitates more effective multi-time-step inference by exploiting more semantic information. Experimental results on benchmark datasets demonstrate that our TSE method significantly improves the classification accuracy and robustness of SNNs. Wei Liu 0153, Li Yang 0014, Mingxuan Zhao, Dengfeng Xue, Shuxun Wang, Boyu Cai, Bing Li 0001, Weiming Hu 0004 |
AAAI | 9 |
| 2025 | Visual-Instructed Degradation Diffusion for All-in-One Image RestorationabstractImage restoration tasks like deblurring, denoising, and dehazing usually need distinct models for each degradation type, restricting their generalization in real-world scenarios with mixed or unknown degradations. In this work, we propose Defusion, a novel all-in-one image restoration framework that utilizes visual instruction-guided degradation diffusion. Unlike existing methods that rely on task-specific models or ambiguous text-based priors, Defusion constructs explicit visual instructions that align with the visual degradation patterns. These instructions are grounded by applying degradations to standardized visual elements, capturing intrinsic degradation features while agnostic to image semantics. Defusion then uses these visual instructions to guide a diffusion-based model that operates directly in the degradation space, where it reconstructs high-quality images by denoising the degradation effects with enhanced stability and generalizability. Comprehensive experiments demonstrate that Defusion outperforms state-of-the-art methods across diverse image restoration tasks, including complex and real-world degradations. Wenyang Luo, Haina Qin, Zewen Chen, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 8 |
| 2025 | Reversing Flow for Image RestorationabstractImage restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, which introduces inefficiency and complexity. In this work, we propose ResFlow, a novel image restoration framework that models the degradation process as a deterministic path using continuous normalizing flows. ResFlow augments the degradation process with an auxiliary process that disambiguates the uncertainty in HQ prediction to enable reversible modeling of the degradation process. ResFlow adopts entropy-preserving flow paths and learns the augmented degradation flow by matching the velocity field. ResFlow significantly improves the performance and speed of image restoration, completing the task in fewer than four sampling steps. Extensive experiments demonstrate that ResFlow achieves state-of-the-art results across various image restoration benchmarks, offering a practical and efficient solution for real-world applications. Haina Qin, Wenyang Luo, Jingdong Chen, Ming Yang 0007, Bing Li 0001, Weiming Hu 0004 |
CVPR | 7 |
| 2025 | VisionMath: Vision-Form Mathematical Problem-Solving
Zongyang Ma, Ziqi Zhang 0010, Zhongang Oi, Chunfeng Yuan, Shaojie Zhu, Chengxiang Zhuo, Bing Li 0001, Ye Liu 0002, Zang Li, Ying Shan, Weiming Hu 0004 |
ICCV | 8 |
| 2025 | Multimodal Large Language Model-Guided ISP Hyperparameter Optimization with Dynamic Preference Learning
Zhikun Zhao, Congyan Lang, Bing Li 0001, Juan Wang 0012 |
ICCV | 4 |
| 2025 | DeepTAGE: Deep Temporal-Aligned Gradient Enhancement for Optimizing Spiking Neural NetworksabstractSpiking Neural Networks (SNNs), with their biologically inspired spatio-temporal dynamics and spike-driven processing, are emerging as a promising low-power alternative to traditional Artificial Neural Networks (ANNs). However, the complex neuronal dynamics and non-differentiable spike communication mechanisms in SNNs present substantial challenges for efficient training. By analyzing the membrane potentials in spiking neurons, we found that their distributions can increasingly deviate from the firing threshold as time progresses, which tends to cause diminished backpropagation gradients and unbalanced optimization. To address these challenges, we propose Deep Temporal-Aligned Gradient Enhancement (DeepTAGE), a novel approach that improves optimization gradients in SNNs from both internal surrogate gradient functions and external supervision methods. Our DeepTAGE dynamically adjusts surrogate gradients in accordance with the membrane potential distribution across different time steps, enhancing their respective gradients in a temporal-aligned manner that promotes balanced training. Moreover, to mitigate issues of gradient vanishing or deviating during backpropagation, DeepTAGE incorporates deep supervision at both spatial (network stages) and temporal (time steps) levels to ensure more effective and robust network optimization. Importantly, our method can be seamlessly integrated into existing SNN architectures without imposing additional inference costs or requiring extra control modules. We validate the efficacy of DeepTAGE through extensive experiments on static benchmarks (CIFAR10, CIFAR100, and ImageNet-1k) and a neuromorphic dataset (DVS-CIFAR10), demonstrating significant performance improvements. Wei Liu 0153, Li Yang 0014, Mingxuan Zhao, Shuxun Wang, Bing Li 0001, Weiming Hu 0004 |
ICLR | 7 |
| 2025 | Noise-Optimized Distribution Distillation for Dataset Condensation
Tongfei Liu, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Chenguang Ma |
ACM Multimedia | 3 |
| 2025 | SynCL: A Synergistic Training Strategy with Instance-Aware Contrastive Learning for End-to-End Multi-Camera 3D TrackingabstractWhile existing query-based 3D end-to-end visual trackers integrate detection and tracking via the *tracking-by-attention* paradigm, these two chicken-and-egg tasks encounter optimization difficulties when sharing the same parameters. Our findings reveal that these difficulties arise due to two inherent constraints on the self-attention mechanism, i.e., over-deduplication for object queries and self-centric attention for track queries. In contrast, removing self-attention mechanism not only minimally impacts regression predictions of the tracker, but also tends to generate more latent candidate boxes. Based on these analyses, we present SynCL, a novel plug-and-play synergistic training strategy designed to co-facilitate multi-task learning for detection and tracking. Specifically, we propose a Task-specific Hybrid Matching module for a weight-shared cross-attention-based decoder that matches the targets of track queries with multiple object queries to exploit promising candidates overlooked by the self-attention mechanism and the bipartite matching. To flexibly select optimal candidates for the one-to-many matching, we also design a Dynamic Query Filtering module controlled by model training status. Moreover, we introduce Instance-aware Contrastive Learning to break through the barrier of self-centric attention for track queries, effectively bridging the gap between detection and tracking. Without additional inference costs, SynCL consistently delivers improvements in various benchmarks and achieves state-of-the-art performance with $58.9\%$ AMOTA on the nuScenes dataset. Code and raw results are available at <https://github.com/shubolin028/SynCL>. Shubo Lin, Yutong Kou, Zirui Wu, Shaoru Wang, Bing Li 0001, Weiming Hu 0004 |
NeurIPS | 5 |
| 2025 | MI-TRQR: Mutual Information-Based Temporal Redundancy Quantification and Reduction for Energy-Efficient Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) provide energy-efficient computation through event-driven processing. However, the shared weights across multiple timesteps lead to serious temporal feature redundancy, limiting both efficiency and performance. This issue is further aggravated when processing static images due to the duplicated input. To mitigate this problem, we propose a parameter-free and plug-and-play module named Mutual Information-based Temporal Redundancy Quantification and Reduction (MI-TRQR), constructing energy-efficient SNNs. Specifically, Mutual Information (MI) is properly introduced to quantify redundancy between discrete spike features at different timesteps on two spatial scales: pixel (local) and the entire spatial features (global). Based on the multi-scale redundancy quantification, we apply a probabilistic masking strategy to remove redundant spikes. The final representation is subsequently recalibrated to account for the spike removal. Extensive experimental results demonstrate that our MI-TRQR achieves sparser spiking firing, higher energy efficiency, and better performance concurrently with different SNN architectures in tasks of neuromorphic data classification, static data classification, and time-series forecasting. Notably, MI-TRQR increases accuracy by \textbf{1.7\%} on CIFAR10-DVS with 4 timesteps while reducing energy cost by \textbf{37.5\%}. Our codes are available at https://github.com/dfxue/MI-TRQR. Dengfeng Xue, Yifan Lu 0001, Chunfeng Yuan, Yufan Liu 0001, Wei Liu 0153, Man Yao, Li Yang 0014, Bing Li 0001, Stephen J. Maybank, Weiming Hu 0004, Zhetao Li |
NeurIPS | 10 |
| 2025 | DFWA-Net: Dual-Domain Feature-Enhanced With Wavelet Attention Network for SAR Ship DetectionabstractSynthetic aperture radar (SAR) is a high-resolution remote sensing technology widely employed for ground and sea surface target detection. However, due to the unique imaging mechanism and information representation of SAR images, conventional spatial-domain feature extraction methods often struggle to fully capture their discriminative features. To address this limitation, this letter introduces the wavelet domain as an additional feature extraction space and proposes a dual-domain feature-enhanced network based on wavelet attention for SAR ship detection. Specifically, two wavelet attention modules are designed to independently and jointly compute attention for high-frequency and low-frequency features in the wavelet domain. Meanwhile, an embedding grouping strategy is adopted to reduce computational costs while enhancing the model’s detailed perception and global understanding of ship targets. Furthermore, a dynamic domain fusion module is proposed to more effectively integrate wavelet-domain and spatial-domain information, enriching feature representation. Comprehensive experiments on two widely used SAR ship datasets demonstrate that the proposed method outperforms many other state-of-the-art detectors. The source code is available at https://github.com/Wenjing-Jiang-hbu/DFWA-Net. Shuaiqi Liu 0001, Wenjing Jiang, Bing Li 0001, Yudong Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2025 | Incomplete multi-view clustering via efficient anchor tensor recovery framework
Jintian Ji, Songhe Feng, Taotao Wei, Peiwu Lv, Bing Li 0001 |
Neural Networks | 7 |
| 2025 | Improved two-view interactional fuzzy learning based on mutual-rectification and knowledge-mergenceabstractNasopharyngeal carcinoma (NPC) is a malignant tumor that originates from the back of the nasal canal from above the soft palate to the upper larynx. Because the nasopharyngeal location is deeply hidden, it is often difficult for a single imaging means to clarify its complex adjacency. In addition, there exist some differences and uncertainties in its clinical manifestations. Although two-view fuzzy classifiers can effectively tap into the nasopharyngeal location for hidden information and exhibit good classification performance, existing fuzzy reasoning for predicting whether or not a nasopharyngeal cancer often stems from the inability to reuse the one-sided rules. Therefore, a novel two-view mutual rectification and knowledge mergence Takagi-Sugeno-Kang fuzzy classifier (TVRM-TFC) is proposed here to address the challenge of using imaging means to fine-tune the organ tissues. Firstly, Kullback-Leibler divergence (KLIC) is used to select important features from various imaging sections (i.e., pieces of knowledge). Secondly, the interpretable zero-order Takagi-Sugeno-Kang (TSK) fuzzy classifier is used as the basic training unit to simultaneously obtain satisfactory accuracies and concise linguistic interpretability. Thirdly, from the perspective of both imaging means and the organ, this study fine-tunes the information required for decision-making between different imaging means, so that the complementary advantages of the different views may improve the decision-making information and thus increase decision accuracies. Finally, the perspective of imaging technology and the organ are merged to capture decision-making knowledge. These decision-making advantages from different views are organically integrated to compensate information and further optimize the decision-making information. The merits of the proposed classifier are demonstrated through comparative experimental analysis on CT and MRI data. Ta Zhou, Wei Yan 0030, Zhengxin Xia, Shuihua Wang, Bing Li 0001, Weiping Ding 0001, Jing Cai 0001 |
Neural Networks | 6 |
| 2025 | Accelerated Self-Supervised Multi-Illumination Color Constancy With Hybrid Knowledge DistillationabstractColor constancy, the human visual system's ability to perceive consistent colors under varying illumination conditions, is crucial for accurate color perception. Recently, deep learning algorithms have been introduced into this task and have achieved remarkable achievements. However, existing methods are limited by the scale of current multi-illumination datasets and model size, hindering their ability to learn discriminative features effectively and their practical value for deployment in cameras. To overcome these limitations, this paper proposes a multi-illumination color constancy approach based on self-supervised learning and knowledge distillation. This approach includes three phases: self-supervised pre-training, supervised fine-tuning, and knowledge distillation. During the pre-training phase, we train Transformer-based and U-Net based encoders by two pretext tasks: light normalization task to learn lighting color contextual representation and grayscale colorization task to acquire objects' inherent color information. For the downstream color constancy task, we fine-tune the encoders and design a lightweight decoder to obtain better illumination distributions with fewer parameters. During the knowledge distillation phase, we introduce a hybrid knowledge distillation technique to align CNN features with those of Transformer and U-Net respectively. Our proposed method outperforms state-of-the-art techniques on multi-illumination and single-illumination benchmarks. Extensive ablation studies and visualizations confirm the effectiveness of our model. Ziyu Feng, Bing Li 0001, Congyan Lang, Zheming Xu, Haina Qin, Juan Wang 0012, Weihua Xiong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | FiGVCL: Fine-Grained Benchmark and Method for Video Copy LocalizationabstractContent-based video copy localization (VCL) aims to detect and locate copied segments in pairs of videos. VCL requires fine-grained video analysis to robustly identify copied segments that have been edited. Despite recent progress, the prohibitive cost of annotating copied segments and the lack of a fine-grained benchmark hinder the development of effective VCL systems. In this work, we annotate a new real-world dataset, FiGVCL, with challenging scenarios designed to evaluate VCL methods. FiGVCL is carefully annotated to preserve the temporal correspondences observed in copied segments. Moreover, we propose a novel fine-grained VCL benchmark metric based on temporal correspondences to improve discriminability. Finally, we design a simple but effective baseline model that uses fine-grained local embeddings for accurate copied segment localization. We also present an unsupervised training strategy that outperforms previous supervised VCL methods. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | iESTA: Instance-Enhanced Spatial-Temporal Alignment for Video Copy LocalizationabstractVideo copy Segment Localization (VSL) requires the identification of the temporal segments within a pair of videos that contain copied content. Current methods primarily focus on global temporal modeling, overlooking the complementarity of global semantic and local fine-grained features, which limits their effectiveness. Some related methods attempt to incorporate local spatial information but often disrupt spatial semantic structures, resulting in less accurate matching. To address these issues, we propose the Instance-Enhanced Spatial-Temporal Alignment Framework (iESTA), based on a proper representation granularity that integrates instance-level local features and semantic global features. Specifically, the Instance-relation Graph (IRG) is constructed to capture instance-level features and fine-grained interactions, preserving local information integrity and better representing the video feature space in a proper granularity. An instance-GNN structure is designed to refine these graph representations. For global features, we enhance the representation of semantic information, capturing temporal relationships within videos using a Transformer framework. Additionally, we design a Complementarity-perception Alignment Module (CAM) to effectively process and integrate complementary spatial-temporal information, producing accurate frame-to-frame alignment maps. Our approach also incorporates a differentiable Dynamic Time Warping (DTW) method to utilize latent temporal alignments as weak supervisory signals, improving the accuracy of the matching process. Experimental results indicate that our proposed iESTA outperforms state-of-the-art methods on both the small-scale dataset VCDB and the large-scale dataset VCSL. Xinmiao Ding, Jinming Lou, Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Task-Aware Attentional Dynamic Alignment for Few-Shot Compressed Video ClassificationabstractWe present a novel Task-aware Attentional Dynamic Alignment (TADA) framework for visual-based few-shot video classification (FSVC) that addresses two key challenges in this field: efficiency and nuanced spatio-temporal reasoning. Existing methods are often hindered by computationally expensive video decoding processes and neglect the temporal order of videos. In contrast, our method harnesses compressed domain data to extract rich spatio-temporal cues at a fraction of the cost of traditional video processing methods. Specifically, we propose an embedding module to extract informative features from compressed domain data while minimizing computational overheads. Furthermore, to exploit the temporal order of frames, we develop a prototypical ADA module to align and classify videos with an explicit temporal order constraint. Our framework also incorporates a contextual mixer to enrich video embeddings with task-specific context. Extensive experiments on multiple datasets demonstrate that TADA achieves state-of-the-art performance and outperforms existing methods in accuracy and efficiency. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | An Asymptotic Multiscale Symmetric Fusion Network for Hyperspectral and Multispectral Image FusionabstractDespite the high spectral resolution and abundant information of hyperspectral images (HSI), their spatial resolution is relatively low due to limitations in sensor technology. Sensors often need to sacrifice some spatial resolution to ensure accurate light energy measurement when pursuing high spectral resolution. This trade-off results in HSI’s inability to capture fine spatial details, thereby limiting its application in scenarios requiring high-precision spatial information. HSI and multispectral images (MSI) fusion is a commonly used technique for generating high-resolution HSI (HR-HSI). However, many deep learning-based HSI-MSI fusion algorithms ignore correlation and multi-scale information between input images. To address this issue, we propose an asymptotic multi-scale symmetric fusion network (AMSF-Net) for hyperspectral and multispectral image fusion. AMSF-Net consists of two parts: the multi-level feature fusion (MFF) module and the progressive cross-scale spatial perception (PCP) module. The MFF module uses multi-stream feature extraction branches to perform information interaction between HSI and MSI at the same scale layer by layer, compensating for the spatial details lacking in HSI and the spectral details absent in MSI. The PCP module combines the input and output features of MFF, utilizes multi-scale bidirectional strip convolution and deep convolution to further refine edge features, and reconstructs HR-HSI by learning the features of different expansion roll branches by connecting across scales. Comparative experiments with several state-of-the-art HSI-MSI fusion algorithms on four publicly available datasets, CAVE, Chikusei, Houston and WorldView-3 are conducted to validate the effectiveness and superiority of AMSF-Net. On the Chikusei dataset, improvements were 9.1%, 12.5%, and 5.1%, respectively, on the indicators RMSE, ERGAS, and SAM, compared to the suboptimal method. Shuaiqi Liu 0001, Tingting Shao, Bing Li 0001, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Set Prediction Guided by Semantic Concepts for Diverse Video CaptioningabstractDiverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated captions. Different from them, we formulate diverse captioning into a semantic-concept-guided set prediction (SCG-SP) problem by fitting the predicted caption set to the ground-truth set, where the set-level relationship is fully captured. Specifically, our set prediction consists of two synergistic tasks, i.e., caption generation and an auxiliary task of concept combination prediction providing extra semantic supervision. Each caption in the set is attached to a concept combination indicating the primary semantic content of the caption and facilitating element alignment in set prediction. Furthermore, we apply a diversity regularization term on concepts to encourage the model to generate semantically diverse captions with various concept combinations. These two tasks share multiple semantics-specific encodings as input, which are obtained by iterative interaction between visual features and conceptual queries. The correspondence between the generated captions and specific concept combinations further guarantees the interpretability of our model. Extensive experiments on benchmark datasets show that the proposed SCG-SP achieves state-of-the-art (SOTA) performance under both relevance and diversity metrics. Yifan Lu 0001, Ziqi Zhang 0010, Chunfeng Yuan, Yan Wang 0153, Bing Li 0001, Weiming Hu 0004 |
AAAI | 6 |
| 2024 | RL-SeqISP: Reinforcement Learning-Based Sequential Optimization for Image Signal ProcessingabstractHardware image signal processing (ISP), aiming at converting RAW inputs to RGB images, consists of a series of processing blocks, each with multiple parameters. Traditionally, ISP parameters are manually tuned in isolation by imaging experts according to application-specific quality and performance metrics, which is time-consuming and biased towards human perception due to complex interaction with the output image. Since the relationship between any single parameter’s variation and the output performance metric is a complex, non-linear function, optimizing such a large number of ISP parameters is challenging. To address this challenge, we propose a novel Sequential ISP parameter optimization model, called the RL-SeqISP model, which utilizes deep reinforcement learning to jointly optimize all ISP parameters for a variety of imaging applications. Concretely, inspired by the sequential tuning process of human experts, the proposed model can progressively enhance image quality by seamlessly integrating information from both the image feature space and the parameter space. Furthermore, a dynamic parameter optimization module is introduced to avoid ISP parameters getting stuck into local optima, which is able to more effectively guarantee the optimal parameters resulting from the sequential learning strategy. These merits of the RL-SeqISP model as well as its high efficiency are substantiated by comprehensive experiments on a wide range of downstream tasks, including two visual analysis tasks (instance segmentation and object detection), and image quality assessment (IQA), as compared with representative methods both quantitatively and qualitatively. In particular, even using only 10% of the training data, our model outperforms other SOTA methods by an average of 7% mAP on two visual analysis tasks. Zhikun Zhao, Congyan Lang, Mingxuan Cai, Longfei Han, Juan Wang 0012, Bing Li 0001 |
AAAI | 8 |
| 2024 | Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-trainingabstractIn vision-language pre-training (VLP), masked image modeling (MIM) has recently been introduced for fine-grained cross-modal alignment. However, in most existing methods, the reconstruction targets for MIM lack high-level semantics, and text is not sufficiently involved in masked modeling. These two drawbacks limit the effect of MIM in facilitating cross-modal semantic alignment. In this work, we propose a semantics-enhanced cross-modal MIM framework (SemMIM) for vision-language representation learning. Specifically, to provide more semantically meaningful supervision for MIM, we propose a local semantics enhancing approach, which harvest high-level semantics from global image features via self-supervised agreement learning and transfer them to local patch encodings by sharing the encoding space. Moreover, to achieve deep involvement of text during the entire MIM process, we propose a text-guided masking strategy and devise an efficient way of injecting textual information in both masked modeling and reconstruction target acquisition. Experimental results validate that our method improves the effectiveness of the MIM task in facilitating cross-modal semantic alignment. Compared to previous VLP models with similar model size and data scale, our SemMIM model achieves state-of-the-art or competitive performance on multiple downstream vision-language tasks. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Bing Li 0001, Weiming Hu 0004 |
LREC/COLING | 10 |
| 2024 | Unifying Latent and Lexicon Representations for Effective Video-Text RetrievalabstractIn video-text retrieval, most existing methods adopt the dual-encoder architecture for fast retrieval, which employs two individual encoders to extract global latent representations for videos and texts. However, they face challenges in capturing fine-grained semantic concepts. In this work, we propose the UNIFY framework, which learns lexicon representations to capture fine-grained semantics and combines the strengths of latent and lexicon representations for video-text retrieval. Specifically, we map videos and texts into a pre-defined lexicon space, where each dimension corresponds to a semantic concept. A two-stage semantics grounding approach is proposed to activate semantically relevant dimensions and suppress irrelevant dimensions. The learned lexicon representations can thus reflect fine-grained semantics of videos and texts. Furthermore, to leverage the complementarity between latent and lexicon representations, we propose a unified learning scheme to facilitate mutual learning via structure sharing and self-distillation. Experimental results show our UNIFY framework largely outperforms previous video-text retrieval methods, with 4.8% and 8.2% Recall@1 improvement on MSR-VTT and DiDeMo respectively. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Qinghao Ye, Chenliang Li 0003, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Bing Li 0001, Weiming Hu 0004 |
LREC/COLING | 10 |
| 2024 | How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?abstractDominant dual-encoder models enable efficient image-text retrieval but suffer from limited accuracy, while the cross-encoder models offer higher accuracy at the expense of efficiency. Distilling cross-modality matching knowledge from cross-encoder to dual-encoder provides a natural approach to harness their strengths. Thus, we investigate the following valuable question: how to make cross-encoder a good teacher for dual-encoder? Our findings are threefold: (1) Cross-modal similarity score distribution of cross-encoder is more concentrated, while the result of dual-encoder is nearly normal, making vanilla logit distillation less effective. However, ranking distillation remains practical, as it is not affected by the score distribution. (2) Only the relative order between hard negatives conveys valid knowledge, while the order information between easy negatives has little significance. (3) Maintaining the coordination between distillation loss and dual-encoder training loss is beneficial for knowledge transfer. Based on these findings, we propose a novel Contrastive Partial Ranking Distillation (CPRD) method, which implements the objective of mimicking relative order between hard negative samples with contrastive learning. This approach coordinates with the training of the dual-encoder, effectively transferring valid knowledge from the cross-encoder to the dual-encoder. Extensive experiments on image-text retrieval and ranking tasks show that our method surpasses other distillation methods and significantly improves the accuracy of dual-encoder. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Junfu Pu, Ying Shan, Xiaojuan Qi 0001, Weiming Hu 0004 |
CVPR | 6 |
| 2024 | PromptIQA: Boosting the Performance and Generalization for No-Reference Image Quality Assessment via Prompts
Zewen Chen, Haina Qin, Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Liang Wang 0001 |
ECCV (1) | 5 |
| 2024 | EA-VTR: Event-Aware Video-Text Retrieval
Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Bing Li 0001, Yingmin Luo, Xu Li 0015, Xiaojuan Qi 0001, Ying Shan, Weiming Hu 0004 |
ECCV (52) | 6 |
| 2024 | MIBench: Evaluating Multimodal Large Language Models over Multiple ImagesabstractHaowei Liu, Xi Zhang, Haiyang Xu, Yaya Shi, Chaoya Jiang, Ming Yan, Ji Zhang, Fei Huang, Chunfeng Yuan, Bing Li, Weiming Hu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haiyang Xu 0001, Yaya Shi, Chaoya Jiang, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
EMNLP | 10 |
| 2024 | Learn from Noise: Detecting Deepfakes via Regional Noise ConsistencyabstractFace forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Various methods primarily concentrate on the features specific to certain generation techniques, potentially resulting in overfitting to the distinctive fingerprint characteristics of those manipulation techniques, thus undermining their generalizability. In contrast, our investigation reveals a prevalent phenomenon wherein regional noise consistency is disrupted during the integration of synthesized faces into source images, regardless of specific manipulation techniques. Motivated by this observation, we introduce the Regional Noise Consistency Learning Framework (RNCL), a novel approach designed to discern manipulated faces. Central to RNCL are two pivotal modules: Noise Consistency Enhancement (NCE) and Pyramidal Noise Consistency Learning (PNCL). The NCE module facilitates channel-wise and spatial-wise feature enhancement by exploiting noise inconsistencies between facial and non-facial regions. Complementarily, the PNCL module constructs a noise consistency pyramid to analyze enhanced features across multiple scales, enabling adaptive multi-scale feature integration. Leveraging the NCE and PNCL modules, our framework effectively transforms noise information into useful forgery cues, significantly enhancing forgery detection performance. Experimental results demonstrate that our method achieves state-of-the-art performance on standard benchmarks. The code will be publicly available. Weiming Bai, Yufan Liu 0001, Bo Wang 0147, Chengwei Peng, Weiming Hu 0004, Bing Li 0001 |
IJCNN | 8 |
| 2024 | NFT1000: A Cross-Modal Dataset For Non-Fungible Token RetrievalabstractWith the rise of "Metaverse" and "Web 3.0", Non-Fungible Token (NFT) has emerged as a kind of pivotal digital asset, garnering significant attention. By the end of March 2024, more than 1.7 billion NFTs have been minted across various blockchain platforms. To effectively locate a desired NFT, conducting searches within a vast array of NFTs is essential. The challenge in NFT retrieval is heightened due to the high degree of similarity among different NFTs, regarding regional and semantic aspects. In this paper, we will introduce a benchmark dataset named "NFT Top1000 Visual-Text Dataset" (NFT1000), containing 7.56 million image-text pairs, and being collected from 1000 most famous PFP1 NFT collections2 by sales volume on the Ethereum blockchain. Based on this dataset and leveraging the CLIP series of pre-trained models as our foundation, we propose the dynamic masking fine-tuning scheme. This innovative approach results in a 7.4\% improvement in the top1 accuracy rate, while utilizing merely 13\% of the total training data (0.79 million vs. 6.1 million). We also propose a robust metric Comprehensive Variance Index (CVI) to assess the similarity and retrieval difficulty of visual-text pairs data. The dataset will be released as an open-source resource. For more details, please refer to: https://github.com/ShuxunoO/NFT-Net.git. Shuxun Wang, Yunfei Lei, Ziqi Zhang 0010, Wei Liu 0153, Li Yang 0014, Bing Li 0001, Weiming Hu 0004 |
ACM Multimedia | 7 |
| 2024 | VQ-Map: Bird's-Eye-View Map Layout Estimation in Tokenized Discrete Space via Vector QuantizationabstractBird's-eye-view (BEV) map layout estimation requires an accurate and full understanding of the semantics for the environmental elements around the ego car to make the results coherent and realistic. Due to the challenges posed by occlusion, unfavourable imaging conditions and low resolution, \emph{generating} the BEV semantic maps corresponding to corrupted or invalid areas in the perspective view (PV) is appealing very recently. \emph{The question is how to align the PV features with the generative models to facilitate the map estimation}. In this paper, we propose to utilize a generative model similar to the Vector Quantized-Variational AutoEncoder (VQ-VAE) to acquire prior knowledge for the high-level BEV semantics in the tokenized discrete space. Thanks to the obtained BEV tokens accompanied with a codebook embedding encapsulating the semantics for different BEV elements in the groundtruth maps, we are able to directly align the sparse backbone image features with the obtained BEV tokens from the discrete representation learning based on a specialized token decoder module, and finally generate high-quality BEV maps with the BEV codebook embedding serving as a bridge between PV and BEV. We evaluate the BEV map layout estimation performance of our model, termed VQ-Map, on both the nuScenes and Argoverse benchmarks, achieving 62.2/47.6 mean IoU for surround-view/monocular evaluation on nuScenes, as well as 73.4 IoU for monocular evaluation on Argoverse, which all set a new record for this map layout estimation task. The code and models are available on \url{https://github.com/Z1zyw/VQ-Map}. Fudong Ge, Guan Luo, Bing Li 0001, Zhaoxiang Zhang 0001, Haibin Ling, Weiming Hu 0004 |
NeurIPS | 5 |
| 2024 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006, Stephen J. Maybank |
Int. J. Comput. Vis. | 3 |
| 2024 | Joint Learning of Audio-Visual Saliency Prediction and Sound Source Localization on Multi-face Videos
Minglang Qiao, Yufan Liu 0001, Mai Xu, Xin Deng 0002, Bing Li 0001, Weiming Hu 0004, Ali Borji |
Int. J. Comput. Vis. | 5 |
| 2024 | DCFNet: Discriminant Correlation Filters Network for Visual Tracking
Weiming Hu 0004, Qiang Wang 0051, Bing Li 0001, Stephen J. Maybank |
J. Comput. Sci. Technol. | 4 |
| 2024 | DA-CapsNet: A multi-branch capsule network based on adversarial domain adaption for cross-subject EEG emotion recognition
Shuaiqi Liu 0001, Zeyao Wang, Yanling An, Bing Li 0001, Yudong Zhang 0001 |
Knowl. Based Syst. | 4 |
| 2024 | MAS-DGAT-Net: A dynamic graph attention network with multibranch feature extraction and staged fusion for EEG emotion recognition
Shuaiqi Liu 0001, Mingqi Jiang, Yanling An, Zhihui Gu, Bing Li 0001, Yudong Zhang 0001 |
Knowl. Based Syst. | 6 |
| 2024 | Recursive Least-Squares Estimator-Aided Online Learning for Visual TrackingabstractTracking visual objects from a single initial exemplar in the testing phase has been broadly cast as a one-/few-shot problem, i.e., one-shot learning for initial adaptation and few-shot learning for online adaptation. The recent few-shot online adaptation methods incorporate the prior knowledge from large amounts of annotated training data via complex meta-learning optimization in the offline phase. This helps the online deep trackers to achieve fast adaptation and reduce overfitting risk in tracking. In this paper, we propose a simple yet effective recursive least-squares estimator-aided online learning approach for few-shot online adaptation without requiring offline training. It allows an in-built memory retention mechanism for the model to remember the knowledge about the object seen before, and thus the seen data can be safely removed from training. This also bears certain similarities to the emerging continual learning field in preventing catastrophic forgetting. This mechanism enables us to unveil the power of modern online deep trackers without incurring too much extra computational cost. We evaluate our approach based on two networks in the online learning families for tracking, i.e., multi-layer perceptrons in RT-MDNet and convolutional neural networks in DiMP. The consistent improvements on several challenging tracking benchmarks demonstrate its effectiveness and efficiency. Yan Lu 0001, Xiaojuan Qi 0001, Yutong Kou, Bing Li 0001, Liang Li 0006, Weiming Hu 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Chinese Title Generation for Short Videos: Dataset, Metric and AlgorithmabstractPrevious work for video captioning aims to objectively describe the video content but the captions lack human interest and attractiveness, limiting its practical application scenarios. The intention of video title generation (video titling) is to produce attractive titles, but there is a lack of benchmarks. This work offers CREATE, the first large-scale Chinese shoRt vidEo retrievAl and Title gEneration dataset, to assist research and applications in video titling, video captioning, and video retrieval in Chinese. CREATE comprises a high-quality labeled 210 K dataset and two web-scale 3 M and 10 M pre-training datasets, covering 51 categories, 50K+ tags, 537K+ manually annotated titles and captions, and 10M+ short videos with original video information. This work presents ACTEr, a unique Attractiveness-Consensus-based Title Evaluation, to objectively evaluate the quality of video title generation. This metric measures the semantic correlation between the candidate (model-generated title) and references (manual-labeled titles) and introduces attractive consensus weights to assess the attractiveness and relevance of the video title. Accordingly, this work proposes a novel multi-modal ALignment WIth Generation model, ALWIG, as one strong baseline to aid future model development. With the help of a tag-driven video-text alignment module and a GPT-based generation module, this model achieves video titling, captioning, and retrieval simultaneously. We believe that the release of the CREATE dataset, ACTEr metric, and ALWIG model will encourage in-depth research on the analysis and creation of Chinese short videos. Ziqi Zhang 0010, Zongyang Ma, Chunfeng Yuan, Peijin Wang, Zhongang Qi, Chenglei Hao, Bing Li 0001, Ying Shan, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | DARTScore: DuAl-Reconstruction Transformer for Video Captioning EvaluationabstractVideo captioning evaluation aims at assessing the semantic consistency between video and candidate text, which should include measurement from two aspects: faithfulness (whether the information conveyed by candidate is correct w.r.t. video) and comprehensiveness (whether the main video content is covered by candidate). However, previous approaches have difficulty in evaluating faithfulness and comprehensiveness due to heavy reliance on references or heterogeneous of visual and textual data. In this paper, we propose a vision-involved evaluation metric based on a novel DuAl-Reconstruction Transformer, named DARTScore. DARTScore formulates the caption evaluation task as a dual-reconstruction problem to evaluate both faithfulness and comprehensiveness explicitly. Since the word in a candidate is usually related to several frames, DARTScore adaptively collects relevant frames to reconstruct the word and computes the reconstruction accuracy as faithfulness to inherently reflect whether the word information is contained in the video. In the inversive way, DARTScore reconstructs each frame with relevant words to evaluate comprehensiveness. By integrating fine-grained bidirectional reconstruction accuracies, DARTScore drills into each word in candidate and each frame in video to fully evaluate the semantic consistency. Furthermore, we collect and annotate two Chinese datasets with a large domain gap, named CRAETE-EVAL and VATEX-ZH-EVAL, to systematically evaluate existing metrics and fill the blank of Chinese video captioning evaluation. Experimental results show that DARTScore achieves higher correlation with human judgments, has lower reference reliance, and generalizes well to data from different domains. Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004, Xiaohu Qie |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | IterDepth: Iterative Residual Refinement for Outdoor Self-Supervised Multi-Frame Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has been a challenging task in computer vision for a long time, and it relies on only monocular or stereo video for its supervision. To address the challenge, we propose a novel multi-frame monocular depth estimation method called IterDepth, which is based on an iterative residual refinement network. IterDepth extracts depth features from consecutive frames and computes a 3D cost volume measuring the difference between current and previous features transformed by PoseCNN (pose estimation convolutional neural network). We reformulate depth prediction as a residual learning problem, revamping the dominating depth regression to enable high-accuracy multi-frame monocular depth estimation. Specifically, we design a gated recurrent depth fusion unit that seamlessly blends depth features from the cost volume, image features, and the depth prediction. The unit updates the hidden states and refines the depth map through iterative refinement, achieving more accurate predictions than existing methods. Our experiments on the KITTI dataset demonstrate that IterDepth is$7\times $faster in terms of FPS (frames per second) than the recent state-of-the-art DepthFormer model with competitive performance. We also test IterDepth on the Cityscapes dataset to showcase its generalization capability in other real-world environments. Moreover, IterDepth can balance accuracy and computational efficiency by adjusting the number of refinement iterations and performs competitively with other CNN-based monocular depth estimation approaches. Source code is available athttps://github.com/PCwenyue/IterDepth-TCSVT. Zhen Chen 0004, Congxuan Zhang, Weiming Hu 0004, Bing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | LG-DBNet: Local and Global Dual-Branch Network for SAR Image DenoisingabstractSynthetic aperture radar (SAR) tends to be seriously affected by speckle noise due to its inherent imaging characteristics, which brings great challenges to the high-level visualization task of SAR images. Therefore, speckle suppression plays a crucial role in remote sensing image processing. Attention-based SAR image denoising algorithms frequently struggle to capture rich feature information and face challenges in balancing the trade-off between denoising and preserving texture details. To solve the above problems, this paper constructs a local and global dual-branch network (LG-DBNet) for SAR image denoising. This network can effectively suppress speckle noise while fully retaining the detail information of the original image. Firstly, the shallow features are extracted through simple convolution. Then, a dual-branch structure constructed using different attention modules is used to extract deep features from SAR images. Specifically, one branch performs local deep feature extraction of an image through a hybrid attention module built by a convolutional neural network (CNN), while the other branch utilizes a superposition of self-attention mechanisms for global deep feature extraction of the image. Finally, the final denoised image is generated through global residual learning. LG-DBNet can effectively extract the local and global image information through the dual-branch structure, and further focus on the noise information, which can better retain the texture information of the image while effectively denoising. The experimental results show that compared with the state-of-the-art SAR image denoising algorithms, the proposed algorithm not only improves on various objective indexes, but also shows great advantages in the visual effect after denoising. Shuaiqi Liu 0001, Shikang Tian, Bing Li 0001, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Coarse-Super-Resolution-Fine Network (CoSF-Net): A Unified End-to-End Neural Network for 4D-MRI With Simultaneous Motion Estimation and Super-ResolutionabstractFour-dimensional magnetic resonance imaging (4D-MRI) is an emerging technique for tumor motion management in image-guided radiation therapy (IGRT). However, current 4D-MRI suffers from low spatial resolution and strong motion artifacts owing to the long acquisition time and patients' respiratory variations. If not managed properly, these limitations can adversely affect treatment planning and delivery in IGRT. In this study, we developed a novel deep learning framework called the coarse-super-resolution-fine network (CoSF-Net) to achieve simultaneous motion estimation and super-resolution within a unified model. We designed CoSF-Net by fully excavating the inherent properties of 4D-MRI, with consideration of limited and imperfectly matched training datasets. We conducted extensive experiments on multiple real patient datasets to assess the feasibility and robustness of the developed network. Compared with existing networks and three state-of-the-art conventional algorithms, CoSF-Net not only accurately estimated the deformable vector fields between the respiratory phases of 4D-MRI but also simultaneously improved the spatial resolution of 4D-MRI, enhancing anatomical features and producing 4D-MR images with high spatiotemporal resolution. Shaohua Zhi, Yinghui Wang 0003, Haonan Xiao, Ti Bai, Bing Li 0001, Yunsong Tang, Wen Li 0010, Tian Li 0012, Jing Cai 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | CATS: A Pragmatic Chinese Answer-to-Sequence Dataset with Large Scale and High QualityabstractLiang Li, Ruiying Geng, Chengyang Fang, Bing Li, Can Ma, Rongyu Cao, Binhua Li, Fei Huang, Yongbin Li. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Liang Li 0006, Ruiying Geng, Chengyang Fang, Bing Li 0001, Can Ma, Rongyu Cao, Binhua Li, Fei Huang 0002, Yongbin Li 0001 |
ACL (1) | 4 |
| 2023 | AUNet: Learning Relations Between Action Units for Face Forgery DetectionabstractFace forgery detection becomes increasingly crucial due to the serious security issues caused by face manipulation techniques. Recent studies in deepfake detection have yielded promising results when the training and testing face forgeries are from the same domain. However, the problem remains challenging when one tries to generalize the detector to forgeries created by unseen methods during training. Observing that face manipulation may alter the relation between different facial action units (AU), we propose the Action-Units Relation Learning framework to improve the generality of forgery detection. In specific, it consists of the Action Units Relation Transformer (ART) and the Tampered AU Prediction (TAP). The ART constructs the relation between different AUs with AU-agnostic Branch and AU-specific Branch, which complement each other and work together to exploit forgery clues. In the Tampered AU Prediction, we tamper AU-related regions at the image level and develop challenging pseudo samples at the feature level. The model is then trained to predict the tampered AU regions with the generated location-specific supervision. Experimental results demonstrate that our method can achieve state-of-the-art performance in both the in-dataset and cross-dataset evaluations. Weiming Bai, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
CVPR | 4 |
| 2023 | ViLEM: Visual-Language Error Modeling for Image-Text RetrievalabstractDominant pre-training works for image-text retrieval adopt “dual-encoder” architecture to enable high efficiency, where two encoders are used to extract image and text representations and contrastive learning is employed for global alignment. However, coarse-grained global alignment ignores detailed semantic associations between image and text. In this work, we propose a novel proxy task, named Visual-Language Error Modeling (ViLEM), to inject detailed image-text association into “dual-encoder” model by “proofreading” each word in the text against the corresponding image. Specifically, we first edit the image-paired text to automatically generate diverse plausible negative texts with pre-trained language models. ViLEM then enforces the model to discriminate the correctness of each word in the plausible negative texts and further correct the wrong words via resorting to image information. Further-more, we propose a multi-granularity interaction framework to perform ViLEM via interacting text features with both global and local image features, which associates local text semantics with both high-level visual context and multi-level local visual information. Our method surpasses state-of-the-art “dual-encoder” methods by a large margin on the image-text retrieval task and significantly improves discriminativeness to local textual semantics. Our model can also generalize well to video-text retrieval. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004, Xiaohu Qie |
CVPR | 7 |
| 2023 | Learning to Exploit the Sequence-Specific Prior Knowledge for Image Processing Pipelines OptimizationabstractThe hardware image signal processing (ISP) pipeline is the intermediate layer between the imaging sensor and the downstream application, processing the sensor signal into an RGB image. The ISP is less programmable and consists of a series of processing modules. Each processing module handles a subtask and contains a set of tunable hyperparameters. A large number of hyperparameters form a complex mapping with the ISP output. The industry typically relies on manual and time-consuming hyperparameter tuning by image experts, biased towards human perception. Recently, several automatic ISP hyperparameter optimization methods using downstream evaluation metrics come into sight. However, existing methods for ISP tuning treat the high-dimensional parameter space as a global space for optimization and prediction all at once without inducing the structure knowledge of ISP. To this end, we propose a sequential ISP hyperparameter prediction framework that utilizes the sequential relationship within ISP modules and the similarity among parameters to guide the model sequence process. We validate the proposed method on object detection, image segmentation, and image quality tasks. Haina Qin, Longfei Han, Weihua Xiong, Juan Wang 0012, Bing Li 0001, Weiming Hu 0004 |
CVPR | 6 |
| 2023 | Learning from the Raw Domain: Cross Modality Distillation for Compressed Video Action RecognitionabstractVideo action recognition is faced with the challenges of both huge computation burden and performance requirements. Using compressed domain data, which saves much decoding computation, is a possible solution. Unfortunately, existing compressed-domain-based (CD) methods fail to obtain high performance, compared with state-of-the-art (SOTA) raw-domain-based (RD) methods. In order to solve the problem, we propose a cross-modality knowledge distillation method to force the CD model to learn the knowledge from the RD model. In particular, spatial knowledge and temporal knowledge are first constructed to align feature space between the raw domain and the compressed domain. Then, an adaptively multi-path knowledge learning scheme is presented to help the CD model learn in a more efficient way. Experiments verify the effectiveness of the proposed method in large-scale and small-scale datasets. Yufan Liu 0001, Jiajiong Cao, Weiming Bai, Bing Li 0001, Weiming Hu 0004 |
ICASSP | 4 |
| 2023 | Order-Prompted Tag Sequence Generation for Video TaggingabstractVideo Tagging intends to infer multiple tags spanning relevant content for a given video. Typically, video tags are freely defined and uploaded by a variety of users, so they have two characteristics: abundant in quantity and disordered intra-video. It is difficult for the existing multilabel classification and generation methods to adapt directly to this task. This paper proposes a novel generative model, Order-Prompted Tag Sequence Generation (OP-TSG), according to the above characteristics. It regards video tagging as a tag sequence generation problem guided by sample-dependent order prompts. These prompts are semantically aligned with tags and enable to decouple tag generation order, making the model focus on modeling the tag dependencies. Moreover, the word-based generation strategy enables the model to generate novel tags. To verify the effectiveness and generalization of the proposed method, a Chinese video tagging benchmark CREATE-tagging, and an English image tagging benchmark Pexel-tagging are established. Extensive results show that OP-TSG is significantly superior to other methods, especially the results on rare tags improve by 3.3% and 3% over SOTA methods on CREATE-tagging and Pexel-tagging, and novel tags generated on CREATE-tagging exhibit a tag gain of 7.04%. Zongyang Ma, Ziqi Zhang 0010, Zhongang Qi, Yingmin Luo, Zekun Li 0006, Chunfeng Yuan, Bing Li 0001, Xiaohu Qie, Ying Shan, Weiming Hu 0004 |
ICCV | 8 |
| 2023 | SMM: Self-supervised Multi-Illumination Color Constancy Model with Multiple Pretext TasksabstractColor constancy is an important ability of the human visual system to perceive constant colors across different illumination. In this paper, we study a more practical yet challenging task, removing color cast by multiple spatial-varying illumination. Previous methods are limited by the scale of the current multi-illumination datasets, which hinders them from learning more discriminative features. Instead, we first propose a self-supervised multi-illumination color constancy model that leverages multiple pretext tasks to fully explore lighting color contextual information and inherent color information without using any manual annotations. During the pre-training phase, we train multiple Transformer-based encoders by learning multiple pretext tasks: (i) the local color distortion recovery task, which is carefully designed to learn lighting color contextual representation, and (ii) the colorization task, which is utilized to acquire inherent knowledge. In the downstream color constancy task, we fine-tune the encoders and design a lightweight decoder to obtain better illumination distributions with fewer parameters. Our lightweight architecture outperforms the state-of-the-art methods on the multi-illuminant benchmark (LSMI) and got robust performance on the single illuminant benchmark (NUS-8). Additionally, extensive ablation studies and visualization results demonstrate the effectiveness of integrating lighting color contextual and inherent color information learning in a self-supervised manner. Ziyu Feng, Zheming Xu, Haina Qin, Congyan Lang, Bing Li 0001, Weihua Xiong |
ACM Multimedia | 5 |
| 2023 | Learning Semantics-Grounded Vocabulary Representation for Video-Text RetrievalabstractPrevious dual-encoder pre-training methods for video-text retrieval employ contrastive learning for cross-modal alignment in a latent space. However, such learned latent spaces often result in modality gap problem [26]. In this paper, we introduce a novel SemVTR framework designed to learn semantics-grounded video-text representations in a vocabulary space, in which each dimension corresponds to a semantic concept represented by a word. The representation is obtained by grounding video and text into semantically-related dimensions with high activation values. As video-text pairs share grounded dimensions, their vocabulary representations are expected to cluster together and thus alleviate modality gap problem. So, the crux of our method lies in grounding video and text into vocabulary space. Specifically, we propose a Multi-Granularity Video Semantics Grounding approach and a Textual Semantics Preserving training strategy. The visualization illustrates that SemVTR obtains semantics-gronded vocabulary representation and also alleviates the modality gap problem. SemVTR significantly outperforms existing methods on four video-text retrieval benchmarks. Yaya Shi, Haiyang Xu 0001, Zongyang Ma, Qinghao Ye, Anwen Hu, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
ACM Multimedia | 11 |
| 2023 | ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingabstractRecently, the transformer has enabled the speed-oriented trackers to approach state-of-the-art (SOTA) performance with high-speed thanks to the smaller input size or the lighter feature extraction backbone, though they still substantially lag behind their corresponding performance-oriented versions. In this paper, we demonstrate that it is possible to narrow or even close this gap while achieving high tracking speed based on the smaller input size. To this end, we non-uniformly resize the cropped image to have a smaller input size while the resolution of the area where the target is more likely to appear is higher and vice versa. This enables us to solve the dilemma of attending to a larger visual field while retaining more raw information for the target despite a smaller input size. Our formulation for the non-uniform resizing can be efficiently solved through quadratic programming (QP) and naturally integrated into most of the crop-based local trackers. Comprehensive experiments on five challenging datasets based on two kinds of transformer trackers, \ie, OSTrack and TransT, demonstrate consistent improvements over them. In particular, applying our method to the speed-oriented version of OSTrack even outperforms its performance-oriented counterpart by 0.6\% AUC on TNL2K, while running 50\% faster and saving over 55\% MACs. Codes and models are available at https://github.com/Kou-99/ZoomTrack. Yutong Kou, Bing Li 0001, Gang Wang 0031, Weiming Hu 0004, Yizheng Wang, Liang Li 0006 |
NeurIPS | 3 |
| 2023 | Exploiting Contextual Objects and Relations for 3D Visual Groundingabstract3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information to distinguish target objects from complex 3D scenes. The absence of annotations for contextual objects and relations further exacerbates the difficulties. In this paper, we propose a novel model, CORE-3DVG, to address these challenges by explicitly learning about contextual objects and relations. Our method accomplishes 3D visual grounding via three sequential modular networks, including a text-guided object detection network, a relation matching network, and a target identification network. During training, we introduce a pseudo-label self-generation strategy and a weakly-supervised method to facilitate the learning of contextual objects and relations, respectively. The proposed techniques allow the networks to focus more effectively on referred objects within 3D scenes by understanding their context better. We validate our model on the challenging Nr3D, Sr3D, and ScanRefer datasets and demonstrate state-of-the-art performance. Our code will be public at https://github.com/yangli18/CORE-3DVG. Li Yang 0014, Chunfeng Yuan, Ziqi Zhang 0010, Zhongang Qi, Wei Liu 0153, Ying Shan, Bing Li 0001, Weiping Yang, Yan Wang 0153, Weiming Hu 0004 |
NeurIPS | 8 |
| 2023 | β-divergence NMF with biorthogonal regularization for data representation
Ruixue Yuan, Chengcai Leng, Bing Li 0001, Anup Basu |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Hierarchical Curriculum Learning for No-Reference Image Quality Assessment
Juan Wang 0012, Zewen Chen, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
Int. J. Comput. Vis. | 4 |
| 2023 | Robust dual-graph discriminative NMF for data classification
Chengcai Leng, Bing Li 0001, Licheng Jiao, Anup Basu |
Knowl. Based Syst. | 3 |
| 2023 | Face swapping detection based on identity spatial constraints with weighted frequency division
Zupeng Ai, Chengwei Peng, Zekun Li 0006, Bing Li 0001 |
Multim. Syst. | 5 |
| 2023 | Ranking-Based Color Constancy With Limited Training SamplesabstractComputational color constancy is an important component of Image Signal Processors (ISP) for white balancing in many imaging devices. Recently, deep convolutional neural networks (CNN) have been introduced for color constancy. They achieve prominent performance improvements comparing with those statistics or shallow learning-based methods. However, the need for a large number of training samples, a high computational cost and a huge model size make CNN-based methods unsuitable for deployment on low-resource ISPs for real-time applications. In order to overcome these limitations and to achieve comparable performance to CNN-based methods, an efficient method is defined for selecting the optimal simple statistics-based method (SM) for each image. To this end, we propose a novel ranking-based color constancy method (RCC) that formulates the selection of the optimal SM method as a label ranking problem. RCC designs a specific ranking loss function, and uses a low rank constraint to control the model complexity and a grouped sparse constraint for feature selection. Finally, we apply the RCC model to predict the order of the candidate SM methods for a test image, and then estimate its illumination using the predicted optimal SM method (or fusing the results estimated by the top k SM methods). Comprehensive experiment results show that the proposed RCC outperforms nearly all the shallow learning-based methods and achieves comparable performance to (sometimes even better performance than) deep CNN-based methods with only 1/2000 of the model size and training time. RCC also shows good robustness to limited training samples and good generalization crossing cameras. Furthermore, to remove the dependence on the ground truth illumination, we extend RCC to obtain a novel ranking-based method without ground truth illumination (RCC_NO) that learns the ranking model using simple partial binary preference annotations provided by untrained annotators rather than experts. RCC_NO also achieves better performance than the SM methods and most shallow learning-based methods with low costs of sample collection and illumination measurement. Bing Li 0001, Haina Qin, Weihua Xiong, Yangxi Li, Songhe Feng, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Learning to Explore Distillability and Sparsability: A Joint Framework for Model CompressionabstractDeep learning shows excellent performance usually at the expense of heavy computation. Recently, model compression has become a popular way of reducing the computation. Compression can be achieved using knowledge distillation or filter pruning. Knowledge distillation improves the accuracy of a lightweight network, while filter pruning removes redundant architecture in a cumbersome network. They are two different ways of achieving model compression, but few methods simultaneously consider both of them. In this paper, we revisit model compression and define two attributes of a model: distillability and sparsability, which reflect how much useful knowledge can be distilled and how many pruned ratios can be obtained, respectively. Guided by our observations and considering both accuracy and model size, a dynamically distillability-and-sparsability learning framework (DDSL) is introduced for model compression. DDSL consists of teacher, student and dean. Knowledge is distilled from the teacher to guide the student. The dean controls the training process by dynamically adjusting the distillation supervision and the sparsity supervision in a meta-learning framework. An alternating direction method of multiplier (ADMM)-based knowledge distillation-with-pruning (KDP) joint optimization algorithm is proposed to train the model. Extensive experimental results show that DDSL outperforms 24 state-of-the-art methods, including both knowledge distillation and filter pruning methods. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Self-Prior Guided Pixel Adversarial Networks for Blind Image InpaintingabstractBlind image inpainting involves two critical aspects, i.e., "where to inpaint" and "how to inpaint". Knowing "where to inpaint" can eliminate the interference arising from corrupted pixel values; a good "how to inpaint" strategy yields high-quality inpainted results robust to various corruptions. In existing methods, these two aspects usually lack explicit and separate consideration. This paper fully explores these two aspects and proposes a self-prior guided inpainting network (SIN). The self-priors are obtained by detecting semantic-discontinuous regions and by predicting global semantic structures of the input image. On the one hand, the self-priors are incorporated into the SIN, which enables the SIN to perceive valid context information from uncorrupted regions and to synthesize semantic-aware textures for corrupted regions. On the other hand, the self-priors are reformulated to provide a pixel-wise adversarial feedback and a high-level semantic structure feedback, which can promote the semantic continuity of inpainted images. Experimental results demonstrate that our method achieves state-of-the-art performance in metric scores and in visual quality. It has an advantage over many existing methods that assume "where to inpaint" is known in advance. Extensive experiments on a series of related image restoration tasks validate the effectiveness of our method in obtaining high-quality inpainting. Juan Wang 0012, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Multi-scale self-attention-based feature enhancement for detection of targets with small image sizes
Xingliang Hu, Bing Li 0001, Congxuan Zhang, Weiming Hu 0004 |
Pattern Recognit. Lett. | 3 |
| 2023 | Dynamic adjustment of hyperparameters for anchor-based detection of objects with large image size differences
Xinliang Hu, Da Teng, Bing Li 0001, Congxuan Zhang, Weiming Hu 0004 |
Pattern Recognit. Lett. | 4 |
| 2023 | Cosine Multilinear Principal Component Analysis for RecognitionabstractExisting two-dimensional principal component analysis methods can only handle second-order tensors (i.e., matrices). However, with the advancement of technology, tensors of order three and higher are gradually increasing. This brings new challenges to dimensionality reduction. Thus, a multilinear method called MPCA was proposed. Although MPCA can be applied to all tensors, using the square of the F-norm makes it very sensitive to outliers. Several two-dimensional methods, such as Angle 2DPCA, have good robustness but cannot be applied to all tensors. We extend the robust Angle 2DPCA method to a multilinear method and propose Cosine Multilinear Principal Component Analysis (CosMPCA) for tensor representation. Our CosMPCA method considers the relationship between the reconstruction error and projection scatter and selects the cosine metric. In addition, our method naturally uses the F-norm to reduce the impact of outliers. We introduce an iterative algorithm to solve CosMPCA. We provide detailed theoretical analysis in both the proposed method and the analysis of the algorithm. Experiments show that our method is robust to outliers and is suitable for tensors of any order. Chengcai Leng, Bing Li 0001, Anup Basu, Licheng Jiao |
IEEE Trans. Big Data | 3 |
| 2023 | Jointing Recurrent Across-Channel and Spatial Attention for Multi-Object Tracking With Block-Erasing Data AugmentationabstractAlthough deep-learning-based multi-object tracking (MOT) approaches have achieved remarkable performances in terms of accuracy and efficiency, the issue of object occlusions remains an open challenge for most one-shot MOT methods. To address the problem of object occlusions, in this paper we present a recurrent across-channel and spatial attention-based one-shot multi-object tracking method with block-erasing data augmentation. First, we construct a multiattention feature learning module, named RASFL, that combines recurrent across -channel attention with spatial attention. The RASFL extracts both the correlations of the feature channels and the differences of the spatial locations to improve the accuracy of the re-identification (Re-ID) task. Second, we adopt a block-erasing data augmentation strategy to handle object occlusions by using random pixel blocks to simulate occlusion cases during the network training process. This block-erasing data augmentation assists the network to be more robust under object occlusions. By integrating the proposed RASFL module and the block-erasing data augmentation strategy into a one-shot online MOT system, we build an accurate and robust MOT model called DcMOT. Finally, we run our method on the MOT16, MOT17 and MOT20 datasets to conduct a comprehensive comparison with some of the state-of-the-art MOT methods. The experimental results demonstrate that the proposed DcMOT model achieves a competitive performance in terms of both accuracy and efficiency; with especially good performances in the occlusion cases. Keyu Deng, Congxuan Zhang, Zhen Chen 0004, Weiming Hu 0004, Bing Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | TranSkeleton: Hierarchical Spatial-Temporal Transformer for Skeleton-Based Action RecognitionabstractIn skeleton-based action recognition, it has been a dominant paradigm to extract motion features with temporal convolution and model spatial correlations with graph convolution. However, it’s difficult for temporal convolution to capture long-range dependencies effectively. Meanwhile, commonly used multi-branch graph convolution leads to high complexity. In this paper, we propose TranSkeleton, a powerful Transformer framework which neatly unifies the spatial and temporal modeling of skeleton sequences. For temporal modeling, we propose a novel partition-aggregation temporal Transformer. It works with hierarchical temporal partition and aggregation, and can capture both long-range dependencies and subtle temporal structures effectively. A difference-aware aggregation approach is designed to reduce information loss during temporal aggregation. For spatial modeling, we propose a topology-aware spatial Transformer which utilizes the prior information of human body topology to facilitate spatial correlation modeling. Extensive experiments on two challenging benchmark datasets demonstrate that TranSkeleton notably outperforms the state of the arts. Yongcheng Liu, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Learning Video-Text Aligned Representations for Video CaptioningabstractVideo captioning requires that the model has the abilities of video understanding, video-text alignment, and text generation. Due to the semantic gap between vision and language, conducting video-text alignment is a crucial step to reduce the semantic gap, which maps the representations from the visual to the language domain. However, the existing methods often overlook this step, so the decoder has to directly take the visual representations as input, which increases the decoder’s workload and limits its ability to generate semantically correct captions. In this paper, we propose a video-text alignment module with a retrieval unit and an alignment unit to learn video-text aligned representations for video captioning. Specifically, we firstly propose a retrieval unit to retrieve sentences as additional input which is used as the semantic anchor between visual scene and language description. Then, we employ an alignment unit with the input of the video and retrieved sentences to conduct the video-text alignment. The representations of two modal inputs are aligned in a shared semantic space. The obtained video-text aligned representations are used to generate semantically correct captions. Moreover, retrieved sentences provide rich semantic concepts which are helpful for generating distinctive captions. Experiments on two public benchmarks, i.e., VATEX and MSR-VTT, demonstrate that our method outperforms state-of-the-art performances by a large margin. The qualitative analysis shows that our method generates correct and distinctive captions. Yaya Shi, Haiyang Xu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | One More Check: Making "Fake Background" Be Tracked AgainabstractThe one-shot multi-object tracking, which integrates object detection and ID embedding extraction into a unified network, has achieved groundbreaking results in recent years. However, current one-shot trackers solely rely on single-frame detections to predict candidate bounding boxes, which may be unreliable when facing disastrous visual degradation, e.g., motion blur, occlusions. Once a target bounding box is mistakenly classified as background by the detector, the temporal consistency of its corresponding tracklet will be no longer maintained. In this paper, we set out to restore the bounding boxes misclassified as ``fake background'' by proposing a re-check network. The re-check network innovatively expands the role of ID embedding from data association to motion forecasting by effectively propagating previous tracklets to the current frame with a small overhead. Note that the propagation results are yielded by an independent and efficient embedding search, preventing the model from over-relying on detection results. Eventually, it helps to reload the ``fake background'' and repair the broken tracklets. Building on a strong baseline CSTrack, we construct a new one-shot tracker and achieve favorable gains by 70.7 ➡ 76.4, 70.6 ➡ 76.3 MOTA on MOT16 and MOT17, respectively. It also reaches a new state-of-the-art MOTA and IDF1 performance. Code is released at https://github.com/JudasDie/SOTS. Bing Li 0001, Weiming Hu 0004 |
AAAI | 4 |
| 2022 | Teacher-Guided Learning for Blind Image Quality Assessment
Zewen Chen, Juan Wang 0012, Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004 |
ACCV (3) | 3 |
| 2022 | Cross-Architecture Knowledge Distillation
Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Weiming Hu 0004, Jingting Ding, Liang Li 0006 |
ACCV (5) | 3 |
| 2022 | EMScore: Evaluating Video Captioning via Coarse-Grained and Fine-Grained Embedding MatchingabstractCurrent metrics for video captioning are mostly based on the text-level comparison between reference and candidate captions. However, they have some insuperable drawbacks, e.g., they cannot handle videos without references, and they may result in biased evaluation due to the one-to-many nature of video-to-text and the neglect of visual relevance. From the human evaluator's viewpoint, a high-quality caption should be consistent with the provided video, but not necessarily be similar to the reference in literal or semantics. Inspired by human evaluation, we propose EMScore (Embedding Matching-based score), a novel reference-free metric for video captioning, which directly measures similarity between video and candidate captions. Benefiting from the recent development of large-scale pre-training models, we exploit a well pre-trained vision-language model to extract visual and linguistic embeddings for computing EMScore. Specifically, EMScore combines matching scores of both coarse-grained (video and caption) and fine-grained (frames and words) levels, which takes the overall understanding and detailed characteristics of the video into account. Furthermore, considering the potential information gain, EMScore can be flexibly extended to the conditions where human-labeled references are available. Last but not least, we collect VATEX-EVAL and ActivityNet-FOIl datasets to systematically evaluate the existing metrics. VATEX-EVAL experiments demonstrate that EMScore has higher human correlation and lower reference dependency. ActivityNet-FOIL experiment verifies that EMScore can effectively identify “hallucinating” captions. Code and datasets are available at https://github.com/shiyaya/emscore. Yaya Shi, Xu Yang 0001, Haiyang Xu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Zhengjun Zha |
CVPR | 5 |
| 2022 | Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningabstractVisual grounding is a task to locate the target indicated by a natural language expression. Existing methods extend the generic object detection framework to this problem. They base the visual grounding on the features from pre-generated proposals or anchors, and fuse these features with the text embeddings to locate the target mentioned by the text. However, modeling the visual features from these predefined locations may fail to fully exploit the visual context and attribute information in the text query, which limits their performance. In this paper, we propose a transformer-based framework for accurate visual grounding by establishing text-conditioned discriminative features and performing multi-stage cross-modal reasoning. Specifically, we develop a visual-linguistic verification module to focus the visual features on regions relevant to the textual descriptions while suppressing the unrelated areas. A language-guided feature encoder is also devised to aggregate the visual contexts of the target object to improve the object's distinctiveness. To retrieve the target from the encoded visual features, we further propose a multi-stage cross-modal decoder to iteratively speculate on the correlations between the image and text for accurate target localization. Extensive experiments on five widely used datasets validate the efficacy of our proposed components and demonstrate state-of-the-art performance. Li Yang 0014, Chunfeng Yuan, Wei Liu 0153, Bing Li 0001, Weiming Hu 0004 |
CVPR | 5 |
| 2022 | Attention-Aware Learning for Hyperparameter Prediction in Image Processing Pipelines
Haina Qin, Longfei Han, Juan Wang 0012, Congxuan Zhang, Bing Li 0001, Weiming Hu 0004 |
ECCV (19) | 6 |
| 2022 | Learnable Pixel Clustering Via Structure and Semantic Dual Constraints for Unsupervised Image SegmentationabstractUnsupervised image segmentation is a challenge task, since a high-quality segmented image should perceive not only local object structures but also certain semantics without any annotations. In this paper, we propose a novel encoder-decoder pixel clustering framework with dual constraints to incorporate local structure and global semantic information for guiding pixel feature learning in a self-supervised manner. On one hand, a Local Structure Constraint (LStC) is constructed based on fine-grained superpixels, which improves the boundary perception of pixel features by keeping intra-superpixel feature consistency and largening inter-superpixel feature distance. On the other hand, a new Global Semantic Constraint (GSeC) is proposed via adapting the mutual information maximization technique to the single-image setting, and it strengthens the global semantic perception of pixel features and thus improves the segmenting integrity of objects. Finally, based on the learned pixel features, a smoothing component is employed to achieve semantically meaningful pixel clustering. The experimental evaluation on BSDS500 and PASCAL Context datasets show the superiority of our method on region and boundary qualities. Bo Wang 0147, Shiang Wang, Chunfeng Yuan, Zhonghai Wu, Bing Li 0001, Weiming Hu 0004, Jeffrey Xiong |
ICIP | 5 |
| 2022 | Inter-Intra Cross-Modality Self-Supervised Video Representation Learning by Contrastive ClusteringabstractThis paper introduces an online self-supervised method that leverages inter- and intra-level variance for video representation learning. Most existing methods tend to focus on instance-level or inter-variance encoding but ignore the intra-variance existing in clips. The key observation to solving this problem is the underlying correlation between visual and audio, in which the distribution of flow patterns in feature space is diverse, but expresses complementary similar semantics. And in the semantic feature space, the horizontal dimension of the feature matrix could be regarded as cluster labels. These cluster labels should be consistent for different modalities of the same video clip. Based on this idea, we propose an end-to-end inter-intra cross-modality contrastive clustering scheme to simultaneously optimize the inter- and intra-level contrastive loss. Experiments show that our proposed approach is able to considerably outperform previous methods for self-supervised learning on HMDB51 and UCF101 when applied to video retrieval and action recognition tasks. Jiutong Wei, Guan Luo, Bing Li 0001, Weiming Hu 0004 |
ICPR | 3 |
| 2022 | Learning Target-aware Representation for Visual Tracking via Informative InteractionsabstractWe introduce a novel backbone architecture to improve target-perception ability of feature representation for tracking. Having observed de facto frameworks perform feature matching simply using the backbone outputs for target localization, there is no direct feedback from the matching module to the backbone network, especially the shallow layers. Concretely, only the matching module can directly access the target information, while the representation learning of candidate frame is blind to the reference target. Therefore, the accumulated target-irrelevant interference in shallow stages may degrade the feature quality of deeper layers. In this paper, we approach the problem by conducting multiple branch-wise interactions inside the Siamese-like backbone networks (InBN). The core of InBN is a general interaction modeler (GIM) that injects the target information to different stages of the backbone network, leading to better target-perception of candidate feature representation with negligible computation cost. The proposed GIM module and InBN mechanism are general and applicable to different backbone types including CNN and Transformer for improvements, as evidenced on multiple benchmarks. In particular, the CNN version improves the baseline with 3.2/6.9 absolute gains of SUC on LaSOT/TNL2K. The Transformer version obtains SUC of 65.7/52.0 on LaSOT/TNL2K, which are on par with recent SOTAs. Mingzhe Guo, Heng Fan 0001, Liping Jing, Yilin Lyu, Bing Li 0001, Weiming Hu 0004 |
IJCAI | 6 |
| 2022 | Long-Short Term Cross-Transformer in Compressed Domain for Few-Shot Video ClassificationabstractCompared with image few-shot learning, most of the existing few-shot video classification methods perform worse on feature matching, because they fail to sufficiently exploit the temporal information and relation. Specifically, frames are usually evenly sampled, which may miss important frames. On the other hand, the heuristic model simply encodes the equally treated frames in sequence, which results in the lack of both long-term and short-term temporal modeling and interaction. To alleviate these limitations, we take advantage of the compressed domain knowledge and propose a long-short term Cross-Transformer (LSTC) for few-shot video classification. For short terms, the motion vector (MV) contains temporal cues and reflects the importance of each frame. For long terms, a video can be natively divided into a sequence of GOPs (Group Of Picture). Using this compressed domain knowledge helps to obtain a more accurate spatial-temporal feature space. Consequently, we design the long-short term selection module, short-term module, and long-term module to comprise the LSTC. Long-short term selection is performed to select informative compressed domain data. Long/short-term modules are utilized to sufficiently exploit the temporal information so that the query and support can be well-matched by cross-attention. Experimental results show the superiority of our method on various datasets. Wenyang Luo, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao, Yangxi Li |
IJCAI | 3 |
| 2022 | Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action ClassificationabstractFor CNN-based visual action recognition, the accuracy may be increased if local key action regions are focused on. The task of self-attention is to focus on key features and ignore irrelevant information. So, self-attention is useful for action recognition. However, current self-attention methods usually ignore correlations among local feature vectors at spatial positions in CNN feature maps. In this paper, we propose an effective interaction-aware self-attention model which can extract information about the interactions between feature vectors to learn attention maps. Since the different layers in a network capture feature maps at different scales, we introduce a spatial pyramid with the feature maps at different layers for attention modeling. The multi-scale information is utilized to obtain more accurate attention scores. These attention scores are used to weight the local feature vectors of the feature maps and then calculate attentional feature maps. Since the number of feature maps input to the spatial pyramid attention layer is unrestricted, we easily extend this attention layer to a spatio-temporal version. Our model can be embedded in any general CNN to form a video-level end-to-end attention network for action recognition. Several methods are investigated to combine the RGB and flow streams to obtain accurate predictions of human actions. Experimental results show that our method achieves state-of-the-art results on the datasets UCF101, HMDB51, Kinetics-400, and untrimmed Charades. Weiming Hu 0004, Chunfeng Yuan, Bing Li 0001, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | SDTP: Semantic-Aware Decoupled Transformer Pyramid for Dense Image PredictionabstractAlthough transformer has achieved great progress on computer vision tasks, the scale variation in dense image prediction is still the key challenge. Few effective multi-scale techniques are applied in transformer and there are two main limitations in the current methods. On the one hand, self-attention module in vanilla transformer fails to sufficiently exploit the diversity of semantic information because of its rigid mechanism. On the other hand, it is difficult to build attention and interaction among different levels due to the heavy computational burden. To alleviate this problem, we first revisit multi-scale problem in dense prediction, verifying the significance of diverse semantic representation and multi-scale interaction, and exploring the adaptation of transformer to pyramidal structure. Inspired by these findings, we propose a novel Semantic-aware Decoupled Transformer Pyramid (SDTP) for dense image prediction, consisting of Intra-level Semantic Promotion (ISP), Cross-level Decoupled Interaction (CDI) and Attention Refinement Function (ARF). ISP explores the semantic diversity in different receptive space through more flexible self-attention strategy. CDI builds the global attention and interaction among different levels in decoupled space which also solves the problem of heavy computation. Besides, ARF is further added to refine the attention in transformer. Experimental results demonstrate the validity and generality of the proposed method, which outperforms the state-of-the-art by a significant margin in dense image prediction tasks. Furthermore, the proposed components are all plug-and-play, which can be embedded in other methods. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Bailan Feng, Kebin Wu, Chengwei Peng, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Simple and Strong Baseline for Universal Targeted Attacks on Siamese Visual TrackingabstractSiamese trackers are shown to be vulnerable to adversarial attacks recently. However, the existing attack methods craft the perturbations for each video independently, which comes at a non-negligible computational cost. In this paper, we show the existence of universal perturbations that can enable the targeted attack, e.g., forcing a tracker to follow the ground-truth trajectory with specified offsets, to be video-agnostic and free from inference in a network. Specifically, we attack a tracker by adding a universal translucent perturbation to the template image and adding afake target, i.e., a small universal adversarial patch, into the search images adhering to the predefined trajectory, so that the tracker outputs the location and size of thefake targetinstead of the real target. Our approach allows perturbing a novel video to come at no additional cost except the mere addition operations – and not require gradient optimization or network inference. Experimental results on several datasets demonstrate that our approach can effectively fool the Siamese trackers in a targeted attack manner. We show that the proposed perturbations are not only universal across videos, but also generalize well across different trackers. Such perturbations are therefore doubly universal, both with respect to the data and the network architectures. Our code is available athttps://github.com/lizhenbang56/SiamAttack. Zhenbang Li, Yaya Shi, Shaoru Wang, Bing Li 0001, Pengpeng Liang, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | MRDDANet: A Multiscale Residual Dense Dual Attention Network for SAR Image DenoisingabstractSynthetic aperture radar (SAR), due to its inherent characteristics, will produce speckle noise, which results in the deterioration of image quality, so the removal of speckle in SAR image is very important for the subsequent high-level image processing. In order to balance the relationship between denoising and texture preservation, we propose a multiscale residual dense dual attention network (MRDDANet) for SAR image denoising. This algorithm can effectively suppress the speckle while fully retaining the texture details of the image. In MRDDANet, shallow features are extracted from the noisy images by multiscale modules with different kernel sizes, and then, the extracted shallow features are mapped to the residual dense dual-attention network to obtain the deep features of SAR image. Finally, the final denoising image is generated through global residual learning. MRDDANet has advantages of both multiscale blocks and residual dense dual attention networks. The dense connection can fully extract features in the image, and the dual-channel attention enables MRDDANet to pay more attention to noise information, which is beneficial to remove noise and keep the details of the original image at the same time. Compared with state-of-the-art algorithms, the results of the experiment indicate that our method not only improves various objective indicators but also shows great advantages in visual effects. Shuaiqi Liu 0001, Luyao Zhang 0004, Bing Li 0001, Weiming Hu 0004, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | SSAU-Net: A Spectral-Spatial Attention-Based U-Net for Hyperspectral Image FusionabstractCompared with traditional remoting image, there is a large amount of spectral information in the hyperspectral image (HSI), which makes HSI better reflect the actual condition of surface features. However, due to the limitations of imaging conditions, HSI tends to have a lower spatial resolution. In order to overcome this issue, we propose a spectral-spatial attention-based U-Net named SSAU-Net for HSI and multispectral image (MSI) fusion. The SSAU-Net constructs a spectral-spatial attention module by a coordinate-attention (CA) module and an efficient pyramid split attention (ESPA) module, which can enhance the image’s spectral information and spatial information. Meanwhile, the proposed network fully extracts the shallow and deep features of the images, and finally generates high-resolution (HR) hyperspectral images. Compared with state-of-the-art HSI-MSI fusion methods, the experimental results verify that the proposed method has a better subjective and objective fusion effect. Shuaiqi Liu 0001, Shichong Zhang, Bing Li 0001, Weiming Hu 0004, Yudong Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Rethinking the Competition Between Detection and ReID in Multiobject TrackingabstractDue to balanced accuracy and speed, one-shot models which jointly learn detection and identification embeddings, have drawn great attention in multi-object tracking (MOT). However, the inherent differences and relations between detection and re-identification (ReID) are unconsciously overlooked because of treating them as two isolated tasks in the one-shot tracking paradigm. This leads to inferior performance compared with existing two-stage methods. In this paper, we first dissect the reasoning process for these two tasks, which reveals that the competition between them inevitably would destroy task-dependent representations learning. To tackle this problem, we propose a novel reciprocal network (REN) with a self-relation and cross-relation design so that to impel each branch to better learn task-dependent representations. The proposed model aims to alleviate the deleterious tasks competition, meanwhile improve the cooperation between detection and ReID. Furthermore, we introduce a scale-aware attention network (SAAN) that prevents semantic level misalignment to improve the association capability of ID embeddings. By integrating the two delicately designed networks into a one-shot online MOT system, we construct a strong MOT tracker, namely CSTrack. Our tracker achieves the state-of-the-art performance on MOT16, MOT17 and MOT20 datasets, without other bells and whistles. Moreover, CSTrack is efficient and runs at 16.4 FPS on a single modern GPU, and its lightweight version even runs at 34.6 FPS. The complete code has been released at https://github.com/JudasDie/SOTS. Bing Li 0001, Shuyuan Zhu, Weiming Hu 0004 |
IEEE Trans. Image Process. | 4 |
| 2022 | Narrowing the Gap: Improved Detector Training With Noisy Location AnnotationsabstractDeep learning methods require massive of annotated data for optimizing parameters. For example, datasets attached with accurate bounding box annotations are essential for modern object detection tasks. However, labeling with such pixel-wise accuracy is laborious and time-consuming, and elaborate labeling procedures are indispensable for reducing man-made noise, involving annotation review and acceptance testing. In this paper, we focus on the impact of noisy location annotations on the performance of object detection approaches and aim to, on the user side, reduce the adverse effect of the noise. First, noticeable performance degradation is experimentally observed for both one-stage and two-stage detectors when noise is introduced to the bounding box annotations. For instance, our synthesized noise results in performance decrease from 38.9% AP to 33.6% AP for FCOS detector on COCO test split, and 37.8%AP to 33.7%AP for Faster R-CNN. Second, a self-correction technique based on a Bayesian filter for prediction ensemble is proposed to better exploit the noisy location annotations following a Teacher-Student learning paradigm. Experiments for both synthesized and real-world scenarios consistently demonstrate the effectiveness of our approach, e.g., our method increases the degraded performance of the FCOS detector from 33.6% AP to 35.6% AP on COCO. Shaoru Wang, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 3 |
| 2022 | PDNet: Toward Better One-Stage Object Detection With Prediction DecouplingabstractRecent one-stage object detectors follow a per-pixel prediction approach that predicts both the object category scores and boundary positions from every single grid location. However, the most suitable positions for inferring different targets, i.e., the object category and boundaries, are generally different. Predicting all these targets from the same grid location thus may lead to sub-optimal results. In this paper, we analyze the suitable inference positions for object category and boundaries, and propose a prediction-target-decoupled detector named PDNet to establish a more flexible detection paradigm. Our PDNet with the prediction decoupling mechanism encodes different targets separately in different locations. A learnable prediction collection module is devised with two sets of dynamic points, i.e., dynamic boundary points and semantic points, to collect and aggregate the predictions from the favorable regions for localization and classification. We adopt a two-step strategy to learn these dynamic point positions, where the prior positions are estimated for different targets first, and the network further predicts residual offsets to the positions with better perceptions of the object properties. Extensive experiments on the MS COCO benchmark demonstrate the effectiveness and efficiency of our method. With a single ResNeXt-64x4d-101-DCN as the backbone, our detector achieves 50.1 AP with single-scale testing, which outperforms the state-of-the-art methods by an appreciable margin under the same experimental settings. Moreover, our detector is highly efficient as a one-stage framework. Our code will be public. Li Yang 0014, Shaoru Wang, Chunfeng Yuan, Ziqi Zhang 0010, Bing Li 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 6 |
| 2022 | 3DCANN: A Spatio-Temporal Convolution Attention Neural Network for EEG Emotion RecognitionabstractSince electroencephalogram (EEG) signals can truly reflect human emotional state, emotion recognition based on EEG has turned into a critical branch in the field of artificial intelligence. Aiming at the disparity of EEG signals in various emotional states, we propose a new deep learning model named three-dimension convolution attention neural network (3DCANN) for EEG emotion recognition in this paper. The 3DCANN model is composed of spatio-temporal feature extraction module and EEG channel attention weight learning module, which can extract the dynamic relation well among multi-channel EEG signals and the internal spatial relation of multi-channel EEG signals during continuous period time. In this model, the spatio-temporal features are fused with the weights of dual attention learning, and the fused features are input into the softmax classifier for emotion classification. In addition, we utilize SJTU Emotion EEG Dataset (SEED) to appraise the feasibility and effectiveness of the proposed algorithm. Finally, experimental results display that the 3DCANN method has superior performance over the state-of-the-art models in EEG emotion recognition. Shuaiqi Liu 0001, Xu Wang 0029, Bing Li 0001, Weiming Hu 0004, Yudong Zhang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | DPFPS: Dynamic and Progressive Filter Pruning for Compressing Convolutional Neural Networks from ScratchabstractFilter pruning is a commonly used method for compressing Convolutional Neural Networks (ConvNets), due to its friendly hardware supporting and flexibility. However, existing methods mostly need a cumbersome procedure, which brings many extra hyper-parameters and training epochs. This is because only using sparsity and pruning stages cannot obtain a satisfying performance. Besides, many works do not consider the difference of pruning ratio across different layers. To overcome these limitations, we propose a novel dynamic and progressive filter pruning (DPFPS) scheme that directly learns a structured sparsity network from Scratch. In particular, DPFPS imposes a new structured sparsity-inducing regularization specifically upon the expected pruning parameters in a dynamic sparsity manner. The dynamic sparsity scheme determines sparsity allocation ratios of different layers and a Taylor series based channel sensitivity criteria is presented to identify the expected pruning parameters. Moreover, we increase the structured sparsity-inducing penalty in a progressive manner. This helps the model to be sparse gradually instead of forcing the model to be sparse at the beginning. Our method solves the pruning ratio based optimization problem by an iterative soft-thresholding algorithm (ISTA) with dynamic sparsity. At the end of the training, we only need to remove the redundant parameters without other stages, such as fine-tuning. Extensive experimental results show that the proposed method is competitive with 11 state-of-the-art methods on both small-scale and large-scale datasets (i.e., CIFAR and ImageNet). Specifically, on ImageNet, we achieve a 44.97% pruning ratio of FLOPs by compressing ResNet-101, even with an increase of 0.12% Top-5 accuracy. Our pruned models and codes are released at https://github.com/taoxvzi/DPFPS. Xiaofeng Ruan, Yufan Liu 0001, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004 |
AAAI | 3 |
| 2021 | Open-Book Video Captioning With Retrieve-Copy-Generate NetworkabstractIn this paper, we convert traditional video captioning task into a new paradigm, i.e., Open-book Video Captioning, which generates natural language under the prompts of video-content-relevant sentences, not limited to the video itself. To address the open-book video captioning problem, we propose a novel Retrieve-Copy-Generate network, where a pluggable video-to-text retriever is constructed to retrieve sentences as hints from the training corpus effectively, and a copy-mechanism generator is introduced to extract expressions from multi-retrieved sentences dynamically. The two modules can be trained end-to-end or separately, which is flexible and extensible. Our framework co-ordinates the conventional retrieval-based methods with orthodox encoder-decoder methods, which can not only draw on the diverse expressions in the retrieved sentences but also generate natural and accurate content of the video. Extensive experiments on several benchmark datasets show that our proposed approach surpasses the state-of-the-art performance, indicating the effectiveness and promising of the proposed paradigm in the task of video captioning. Ziqi Zhang 0010, Zhongang Qi, Chunfeng Yuan, Ying Shan, Bing Li 0001, Weiming Hu 0004 |
CVPR | 5 |
| 2021 | Practical Face Swapping Detection Based on Identity Spatial ConstraintsabstractThe generalization of face swapping detectors against unseen face manipulation methods is important to practical applications. Most existing methods based on convolutional neural networks (CNN) simply map the facial images to real/fake binary labels and achieve high performance on the known forgeries, but they almost fail to detect new manipulation methods. In order to improve the generalization of face swapping detection, this work concentrates on a practical scenario to protect specific persons by proposing a novel face swapping detector requiring a reference image. To this end, we design a new detection framework based on identity spatial constraints (DISC), which consists of a backbone network and an identity semantic encoder (ISE). When inspecting an image of a particular person, the ISE utilizes a real facial image of that person as the reference to constrain the backbone to focus on the identity-related facial areas, so as to exploit the intrinsic discriminative clues to the forgery in the query image. Cross-dataset evaluations on five large-scale face forgery datasets show that DISC significantly improves the performance against unseen manipulation methods and is robust against the distortions. Compared to the existing detection methods, the AUC scores achieve 10%~40% performance improvements. Bo Wang 0147, Bing Li 0001, Weiming Hu 0004 |
IJCB | 3 |
| 2021 | Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionabstractGraph convolutional networks (GCNs) have been widely used and achieved remarkable results in skeleton-based action recognition. In GCNs, graph topology dominates feature aggregation and therefore is the key to extracting representative features. In this work, we propose a novel Channel-wise Topology Refinement Graph Convolution (CTR-GC) to dynamically learn different topologies and effectively aggregate joint features in different channels for skeleton-based action recognition. The proposed CTR-GC models channel-wise topologies through learning a shared topology as a generic prior for all channels and refining it with channel-specific correlations for each channel. Our refinement method introduces few extra parameters and significantly reduces the difficulty of modeling channel-wise topologies. Furthermore, via reformulating graph convolutions into a unified form, we find that CTR-GC relaxes strict constraints of graph convolutions, leading to stronger representation capability. Combining CTR-GC with temporal modeling modules, we develop a powerful graph convolutional network named CTR-GCN which notably outperforms state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120, and NW-UCLA datasets.1 Ziqi Zhang 0010, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004 |
ICCV | 4 |
| 2021 | Learn to Match: Automatic Matching Network Design for Visual TrackingabstractSiamese tracking has achieved groundbreaking performance in recent years, where the essence is the efficient matching operator cross-correlation and its variants. Besides the remarkable success, it is important to note that the heuristic matching network design relies heavily on expert experience. Moreover, we experimentally find that one sole matching operator is difficult to guarantee stable tracking in all challenging environments. Thus, in this work, we introduce six novel matching operators from the perspective of feature fusion instead of explicit similarity learning, namely Concatenation, Pointwise-Addition, Pairwise-Relation, FiLM, Simple-Transformer and Transductive-Guidance, to explore more feasibility on matching operator selection. The analyses reveal these operators’ selective adaptability on different environment degradation types, which inspires us to combine them to explore complementary features. To this end, we propose binary channel manipulation (BCM) to search for the optimal combination of these operators. BCM determines to retrain or discard one operator by learning its contribution to other tracking steps. By inserting the learned matching networks to a strong baseline tracker Ocean [47], our model achieves favorable gains by 67.2 → 71.4, 52.6 → 58.3, 70.3 → 76.0 success on OTB100, LaSOT, and TrackingNet, respectively. Notably, Our tracker, dubbed AutoMatch, uses less than half of training data/time than the baseline tracker, and runs at 50 FPS using PyTorch. Code and model are released at https://github.com/JudasDie/SOTS. Yihao Liu 0001, Xiao Wang 0014, Bing Li 0001, Weiming Hu 0004 |
ICCV | 4 |
| 2021 | DSIC: Dynamic Sample-Individualized Connector for Multi-Scale Object DetectionabstractAlthough object detection has reached a milestone recently, the scale variation is still the key challenge. Integrating multilevel features is presented to alleviate the problems, like Feature Pyramid Network (FPN) and its improvements. However, the specifically designed architectures and fixed data flow paths of these methods are not flexible for feature fusion, especially when fed with various samples. To overcome the limitations, we propose a Dynamic Sample-Individualized Connector (DSIC) for multi-scale object detection, which dynamically adjusts network connections to fit different samples. In particular, DSIC consists of two components: Intra-scale Selection Gate (ISG) and Cross-scale Selection Gate (CSG). With the help of the presented gate operator, ISG adaptively extracts proper multi-level features from backbone as the inputs of feature integration. CSG automatically activates informative data flow paths based on the extracted multi-level features. These two components are both plug-and-play and can be embedded in any backbone. Experimental results demonstrate that the proposed method outperforms the state-of-the- arts. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Yanan Miao |
ICME | 3 |
| 2021 | Adaptive Coarse-to-Fine Interactor for Multi-Scale Object DetectionabstractScale variation is one of the key challenges of object detection. Multi-level feature fusion is presented to alleviate the problems, e.g., Feature Pyramid Network (FPN) and its extended methods. However, the input features fed into these methods and the interaction among features from different levels are insufficient and rigid. To fully exploit the features of multi-scale objects and enhance the feature interaction, we propose a novel and effective framework called Adaptive Coarse-to-Fine Interactor (ACFI). Specifically, ACFI consists of three cascaded components: Multi-Resolution Fusion (MRF), Fine-Grained Interaction (FGI), and Edge-aware Enhancement (EAE). MRF adaptively extracts multi-level features from multi-resolution images and multi-stage features, and then these features are fed into FGI to have a fine-grained interaction utilizing bottom-up guidance. After that, EAE further refines the features obtained by FGI, and enhances the detailed edge information and suppresses the redundant noise. After the coarse-to-fine process, we can obtain powerful multiscale representations of various objects. Each component can be embedded into any backbones, separately. Experimental results show the superiority of our method and verify the effectiveness of each proposed module. Zekun Li 0006, Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004 |
IJCNN | 3 |
| 2021 | Web Objectionable Video Recognition Based on Deep Multi-Instance Learning With Representative Prototypes SelectionabstractTo protect underage people from accessing objectionable videos in the Internet, an effective objectionable video recognition algorithm is necessary for web filtering. Recently, the multi-instance learning has been introduced for objectionable video recognition and achieves impressive results. However, hand-crafted features as well as redundant and noisy frames in objectionable videos become an intractable problem that inevitably degrades the recognition performance. In this paper, we propose a novel representative prototype selection algorithm embedding deep multi-instance representation learning. In the proposed method, an improved convolutional neural network is designed for multimodal multi-instance feature learning and a self-expressive dictionary learning model based on sparse and low rank constraint is designed to select the representative prototypes from each subspace of instances. Then the bag-level feature is constructed via mapping the bag to the selected prototypes. Experiments on three objectionable video sets show the effectiveness of our method for objectionable video recognition. Xinmiao Ding, Bing Li 0001, Yangxi Li, Weihua Xiong, Weiming Hu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Robust Texture-Aware Computer-Generated Image Forensic: Benchmark and AlgorithmabstractWith advances in rendering techniques and generative adversarial networks, computer-generated (CG) images tend to be indistinguishable from photographic (PG) images. Revisiting previous works towards CG image forensic, we observed that existing datasets are constructed years ago and limited in both quantity and diversity. Besides, current algorithms only consider the global visual features for forensic, ignoring finer differences between CG and PG images. To mitigate these problems, we first contribute a Large-Scale CG images Benchmark (LSCGB), and then propose a simple yet strong baseline model to address the forensic task. On the one hand, the introduced benchmark has three superior properties, 1) large-scale: the benchmark contains 71168 CG and 71168 PG images with the corresponding expert-annotated labels. It is orders of magnitude bigger than previous datasets. 2) high diversity: we collect CG images from 4 different scenes generated by various rendering techniques. The PG images are varied in terms of image content, camera models, and photographer styles. 3) small bias: we carefully filter the collected images to ensure that the distributions of color, brightness, tone and saturation between CG and PG images are close. Furthermore, inspired by an empirical study on texture difference between CG and PG images, an effective texture-aware network is proposed to improve forensic accuracy. Concretely, we first strengthen texture information of multilevel features extracted from a backbone. Then, the relations among feature channels are explored by learning its gram matrix. Each feature channel represents a specific texture pattern. The gram matrix is thus able to embed the finer texture differences. Experimental results demonstrate that this baseline surpasses the existing methods. The benchmark is publically available at https://github.com/wmbai/LSCGB. Weiming Bai, Bing Li 0001, Yangxi Li, Congxuan Zhang, Weiming Hu 0004 |
IEEE Trans. Image Process. | 3 |
| 2021 | Multi-Scale Low-Discriminative Feature Reactivation for Weakly Supervised Object LocalizationabstractFor weakly supervised object localization (WSOL), how to avoid the network focusing only on some small discriminative parts is a main challenge needed to solve. The widely-used Class Activation Mapping (CAM) based paradigm usually employs Adversarial Learning (AL) strategy to search more object parts by constantly hiding discovered object features, but the adversarial process is difficult to control. In this paper, we propose a novel CAM-based framework with Multi-scale Low-Discriminative Feature Reactivation (mLDFR) for WSOL. The mLDFR framework reactivates the low-discriminative object parts via bottom-up continuous feature maps recalibration and multi-scale object category mapping. Compared with the AL-based methods, our method fully improves the localization power of the network without damaging the classification power and can perform multi-instance localization, which are hard to achieve under the AL-based framework. Moreover, the mLDFR framework is flexible, and can be built on the top of various classical CNN backbones. Experimental results demonstrate the superiority of our method. With VGG16 as backbone, we achieve 46.96% Cls-Loc top1 err and 66.12% CorLoc on ILSVRC2014, 38.07% Cls-Loc top1 err and 75.04% CorLoc on CUB200-2011, surpassing the state-of-the-arts by a large margin. Bo Wang 0147, Chunfeng Yuan, Bing Li 0001, Xinmiao Ding, Zeya Li, Ying Wu 0001, Weiming Hu 0004 |
IEEE Trans. Image Process. | 3 |
| 2021 | Toward Accurate Pixelwise Object Tracking via Attention RetrievalabstractPixelwise single object tracking is challenging due to the competition of running speeds and segmentation accuracy. Current state-of-the-art real-time approaches seamlessly connect tracking and segmentation by sharing computation of the backbone network, e.g., SiamMask and D3S fork a light branch from the tracking model to predict segmentation mask. Although efficient, directly reusing features from tracking networks may harm the segmentation accuracy, since background clutter in the backbone feature tends to introduce false positives in segmentation. To mitigate this problem, we propose a unified tracking-retrieval-segmentation framework consisting of an attention retrieval network (ARN) and an iterative feedback network (IFN). Instead of segmenting the target inside the bounding box, the proposed framework performs soft spatial constraints on backbone features to obtain an accurate global segmentation map. Concretely, in ARN, a look-up-table (LUT) is first built by sufficiently using the information of the first frame. By retrieving it, a target-aware attention map is generated to suppress the negative influence of background clutter. To ulteriorly refine the contour of the segmentation, IFN iteratively enhances the features at different resolutions by taking the predicted mask as feedback guidance. Our framework sets a new state of the art on the recent pixelwise tracking benchmark VOT2020 and runs at 40 fps. Notably, the proposed model surpasses SiamMask by 11.7/4.2/5.5 points on VOT2020, DAVIS2016, and DAVIS2017, respectively. Code is available at https://github.com/JudasDie/SOTS. Yufan Liu 0001, Bing Li 0001, Weiming Hu 0004, Houwen Peng |
IEEE Trans. Image Process. | 3 |
| 2021 | EDP: An Efficient Decomposition and Pruning Scheme for Convolutional Neural Network CompressionabstractModel compression methods have become popular in recent years, which aim to alleviate the heavy load of deep neural networks (DNNs) in real-world applications. However, most of the existing compression methods have two limitations: 1) they usually adopt a cumbersome process, including pretraining, training with a sparsity constraint, pruning/decomposition, and fine-tuning. Moreover, the last three stages are usually iterated multiple times. 2) The models are pretrained under explicit sparsity or low-rank assumptions, which are difficult to guarantee wide appropriateness. In this article, we propose an efficient decomposition and pruning (EDP) scheme via constructing a compressed-aware block that can automatically minimize the rank of the weight matrix and identify the redundant channels. Specifically, we embed the compressed-aware block by decomposing one network layer into two layers: a new weight matrix layer and a coefficient matrix layer. By imposing regularizers on the coefficient matrix, the new weight matrix learns to become a low-rank basis weight, and its corresponding channels become sparse. In this way, the proposed compressed-aware block simultaneously achieves low-rank decomposition and channel pruning by only one single data-driven training stage. Moreover, the network of architecture is further compressed and optimized by a novel Pruning & Merging (PM) module which prunes redundant channels and merges redundant decomposed layers. Experimental results (17 competitors) on different data sets and networks demonstrate that the proposed EDP achieves a high compression ratio with acceptable accuracy degradation and outperforms state-of-the-arts on compression rate, accuracy, inference time, and run-time memory. Xiaofeng Ruan, Yufan Liu 0001, Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Yangxi Li, Stephen J. Maybank |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Object Relational Graph With Teacher-Recommended Learning for Video CaptioningabstractTaking full advantage of the information from both vision and language is critical for the video captioning task. Existing models lack adequate visual representation due to the neglect of interaction between object, and sufficient training for content-related words due to long-tailed problems. In this paper, we propose a complete video captioning system including both a novel model and an effective training strategy. Specifically, we propose an object relational graph (ORG) based encoder, which captures more detailed interaction features to enrich visual representation. Meanwhile, we design a teacher-recommended learning (TRL) method to make full use of the successful external language model (ELM) to integrate the abundant linguistic knowledge into the caption model. The ELM generates more semantically similar word proposals which extend the groundtruth words used for training to deal with the long-tailed problem. Experimental evaluations on three benchmarks: MSVD, MSR-VTT and VATEX show the proposed ORG-TRL system achieves state-of-the-art performance. Extensive ablation studies and visualizations illustrate the effectiveness of our system. Ziqi Zhang 0010, Yaya Shi, Chunfeng Yuan, Bing Li 0001, Peijin Wang, Weiming Hu 0004, Zhengjun Zha |
CVPR | 4 |
| 2020 | Learning to Predict Salient Faces: A Novel Visual-Audio Saliency Model
Yufan Liu 0001, Minglang Qiao, Mai Xu, Bing Li 0001, Weiming Hu 0004, Ali Borji |
ECCV (20) | 4 |
| 2020 | Ocean: Object-Aware Anchor-Free Tracking
Houwen Peng, Jianlong Fu, Bing Li 0001, Weiming Hu 0004 |
ECCV (21) | 4 |
| 2020 | End-to-End Temporal Feature Aggregation for Siamese TrackersabstractWhile siamese networks have demonstrated the significant improvement on object tracking performances, how to utilize the temporal information in siamese trackers has not been widely studied yet. In this paper, we introduce a novel siamese tracking architecture equipped with a temporal aggregation module, which improves the per-frame features by aggregating temporal information from adjacent frames. This temporal fusion strategy enables the siamese trackers to handle poor object appearance like motion blur, occlusion, etc. Furthermore, we incorporate the adversarial dropout module in the siamese network for computing discriminative target features in an end-to-end-fashion. Comprehensive experiments demonstrate that the proposed tracker performs favorably against state-of-the-art trackers. Zhenbang Li, Qiang Wang 0051, Bing Li 0001, Weiming Hu 0004 |
ICIP | 4 |
| 2020 | Globally Spatial-Temporal Perception: a Long-Term Tracking SystemabstractAlthough siamese trackers have achieved superior performance, these kinds of approaches tend to favour the local search mechanism and are thus prone to accumulating inaccuracies of predicted positions, leading to tracking drift over time, especially in long-term tracking scenario. To solve these problems, we propose a siamese tracker in the spirit of the faster RCNN's two-stage detection paradigm. This new tracker is dedicated to reducing cumulative inaccuracies and improving robustness based on a global perception mechanism, which allows the target to be retrieved in time spatially over the whole image plane. Since the very deep network can be enabled for feature learning in this two-stage tracking framework, the power of discrimination is guaranteed. What's more, we also add a CNN-based trajectory prediction module exploiting the target's temporal motion information to mitigate the interference of distractors. These two spatial and temporal modules exploit both the high-level appearance information and complementary trajectory information to improve the tracking robustness. Comprehensive experiments demonstrate that the proposed Globally Spatial-Temporal Perception-based tracking system performs favorably against state-of-the-art trackers. Zhenbang Li, Qiang Wang 0051, Bing Li 0001, Weiming Hu 0004 |
ICIP | 4 |
| 2020 | Graph convolutional network with structure pooling and joint-wise channel attention for action recognition
Gaoqun Ma, Chunfeng Yuan, Bing Li 0001, Fangshi Wang, Weiming Hu 0004 |
Pattern Recognit. | 4 |
| 2020 | Manipulating Template Pixels for Model Adaptation of Siamese Visual TrackingabstractIn this letter, we show that the challenging model adaptation task in visual object tracking can be handled by simply manipulating pixels of the template image in Siamese networks. For a target that is not included in the offline training set, a slight modification of the template image pixels will improve the prediction result of the offline trained Siamese network. The popular adversarial example generation methods can be used to perform template pixel manipulation for model adaptation. Different from current template update methods, which aim to combine the target features from previous frames, we focus on the initial adaptation using target ground-truth in the first frame. Our model adaptation method is pluggable, in the sense that it does not alter the overall architecture of its base tracker. To our knowledge, this work is the first attempt to directly manipulating template pixels for model adaptation in Siamese-based trackers. Extensive experiments on recent benchmarks demonstrate that our method achieves better performance than some other state-of-the-art trackers. Our code is available at https://github.com/lizhenbang56/MTP. Zhenbang Li, Bing Li 0001, Liang Li 0006, Weiming Hu 0004 |
IEEE Signal Process. Lett. | 2 |
| 2020 | Multi-Cue Semi-Supervised Color Constancy With Limited Training SamplesabstractColor constancy is one of the fundamental tasks in computer vision. Many supervised methods, including recently proposed Convolutional Neural Networks (CNN)-based methods, have been proved to work well on this problem, but they often require a sufficient number of labeled data. However, it is expensive and time-consuming to collect a large number of labeled training images with accurately measured illumination. In order to reduce the dependence on labeled images and leverage unlabeled ones without measured illumination, we propose a novel semi-supervised framework with limited training samples for illumination estimation. Our key insight is that the images with similar features from different cues will share similar lighting conditions. Consequently, three graphs based on three visual cues, low-level RGB color distribution, mid-level initial illuminant estimates and high-level scene content, are constructed to represent the relationship among different images. Then a multi-cue semi-supervised color constancy method (MSCC) is proposed after integrating these three graphs into a unified model. Extensive experiments on benchmark datasets demonstrate that our proposed MSCC method outperforms nearly all the existing supervised methods with limited labeled samples. Even with no unlabeled samples, MSCC still obtains better performance and stableness than most supervised methods. Xinwei Huang, Bing Li 0001, Shuai Li 0001, Weihua Xiong, Xuanwu Yin, Weiming Hu 0004, Hong Qin 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Anisotropic Convolution for Image ClassificationabstractConvolutional neural networks are built upon simple but useful convolution modules. The traditional convolution has a limitation on feature extraction and object localization due to its fixed scale and geometric structure. Besides, the loss of spatial information also restricts the networks' performance and depth. To overcome these limitations, this paper proposes a novel anisotropic convolution by adding a scale factor and a shape factor into the traditional convolution. The anisotropic convolution augments the receptive fields flexibly and dynamically depending on the valid sizes of objects. In addition, the anisotropic convolution is a generalized convolution. The traditional convolution, dilated convolution and deformable convolution can be viewed as its special cases. Furthermore, in order to improve the training efficiency and avoid falling into a local optimum, this paper introduces a simplified implementation of the anisotropic convolution. The anisotropic convolution can be applied to arbitrary convolutional networks and the enhanced networks are called ACNs (anisotropic convolutional networks). Experimental results show that ACNs achieve better performance than many state-of-the-art methods and the baseline networks in tasks of image classification and object localization, especially in classification task of tiny images. Bing Li 0001, Chunfeng Yuan, Yangxi Li, Haohao Wu, Weiming Hu 0004, Fangshi Wang |
IEEE Trans. Image Process. | 2 |
| 2020 | Anomaly Detection Using Local Kernel Density Estimation and Context-Based RegressionabstractCurrent local density-based anomaly detection methods are limited in that the local density estimation and the neighborhood density estimation are not accurate enough for complex and large databases, and the detection performance depends on the size parameter of the neighborhood. In this paper, we propose a new kernel function to estimate samples' local densities and propose a weighted neighborhood density estimation to increase the robustness to changes in the neighborhood size. We further propose a local kernel regression estimator and a hierarchical strategy for combining information from the multiple scale neighborhoods to refine anomaly factors of samples. We apply our general anomaly detection method to image saliency detection by regarding salient pixels in objects as anomalies to the background regions. Local density estimation in the visual feature space and kernel-based saliency score propagation in the image enable the assignment of similar saliency values to homogenous object regions. Experimental results on several benchmark datasets demonstrate that our anomaly detection methods overall outperform several state-of-art anomaly detection methods. The effectiveness of our image saliency detection method is validated by comparison with several state-of-art saliency detection methods. Weiming Hu 0004, Bing Li 0001, Ou Wu 0001, Junping Du 0001, Stephen J. Maybank |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2019 | Knowledge Distillation via Instance Relationship GraphabstractThe key challenge of knowledge distillation is to extract general, moderate and sufficient knowledge from a teacher network to guide a student network. In this paper, a novel Instance Relationship Graph (IRG) is proposed for knowledge distillation. It models three kinds of knowledge, including instance features, instance relationships and feature space transformation, while the latter two kinds of knowledge are neglected by previous methods. Firstly, the IRG is constructed to model the distilled knowledge of one network layer, by considering instance features and instance relationships as vertexes and edges respectively. Secondly, an IRG transformation is proposed to models the feature space transformation across layers. It is more moderate than directly mimicking the features at intermediate layers. Finally, hint loss functions are designed to force a student's IRGs to mimic the structures of a teacher's IRGs. The proposed method effectively captures the knowledge along the whole network via IRGs, and thus shows stable convergence and strong robustness to different network architectures. In addition, the proposed method shows superior performance over existing methods on datasets of various scales. Yufan Liu 0001, Jiajiong Cao, Bing Li 0001, Chunfeng Yuan, Weiming Hu 0004, Yangxi Li, Yunqiang Duan |
CVPR | 3 |
| 2019 | Multimodal Semantic Attention Network for Video CaptioningabstractInspired by the fact that different modalities in videos carry complementary information, we propose a Multimodal Semantic Attention Network(MSAN), which is a new encoder-decoder framework incorporating multimodal semantic attributes for video captioning. In the encoding phase, we detect and generate multimodal semantic attributes by formulating it as a multi-label classification problem. Moreover, we add auxiliary classification loss to our model that can obtain more effective visual features and high-level multimodal semantic attribute distributions for sufficient video encoding. In the decoding phase, we extend each weight matrix of the conventional LSTM to an ensemble of attribute-dependent weight matrices, and employ attention mechanism to pay attention to different attributes at each time of the captioning process. We evaluate algorithm on two popular public benchmarks: MSVD and MSR-VTT, achieving competitive results with current state-of-the-art across six evaluation metrics. Bing Li 0001, Chunfeng Yuan, Zhengjun Zha, Weiming Hu 0004 |
ICME | 2 |
| 2019 | Asymmetric 3D Convolutional Neural Networks for action recognition
Hao Yang 0010, Chunfeng Yuan, Bing Li 0001, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank |
Pattern Recognit. | 3 |
| 2018 | Hierarchical Nonlinear Orthogonal Adaptive-Subspace Self-Organizing Map Based Feature Extraction for Human Action RecognitionabstractFeature extraction is a critical step in the task of action recognition. Hand-crafted features are often restricted because of their fixed forms and deep learning features are more effective but need large-scale labeled data for training. In this paper, we propose a new hierarchical Nonlinear Orthogonal Adaptive-Subspace Self-Organizing Map(NOASSOM) to adaptively and learn effective features from data without supervision. NOASSOM is extended from Adaptive-Subspace Self-Organizing Map (ASSOM) which only deals with linear data and is trained with supervision by the labeled data. Firstly, by adding a nonlinear orthogonal map layer, NOASSOM is able to handle the nonlinear input data and it avoids defining the specific form of the nonlinear orthogonal map by a kernel trick. Secondly, we modify loss function of ASSOM such that every input sample is used to train model individually. In this way, NOASSOM effectively learns the statistic patterns from data without supervision. Thirdly, we propose a hierarchical NOASSOM to extract more representative features. Finally, we apply the proposed hierarchical NOASSOM to efficiently describe the appearance and motion information around trajectories for action recognition. Experimental results on widely used datasets show that our method has superior performance than many state-of-the-art hand-crafted features and deep learning features based methods. Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Hao Yang 0010, Zhikang Fu |
AAAI | 3 |
| 2018 | Interaction-Aware Spatio-Temporal Pyramid Attention Networks for Action Classification
Chunfeng Yuan, Bing Li 0001, Yangxi Li, Weiming Hu 0004 |
ECCV (16) | 3 |
| 2017 | Spatio-Temporal Self-Organizing Map Deep Network for Dynamic Object Detection from Videos
Chunfeng Yuan, Bing Li 0001, Weiming Hu 0004, Stephen J. Maybank |
CVPR | 3 |
| 2017 | Multi-View Multi-Instance Learning Based on Joint Sparse Representation and Multi-View Dictionary LearningabstractIn multi-instance learning (MIL), the relations among instances in a bag convey important contextual information in many applications. Previous studies on MIL either ignore such relations or simply model them with a fixed graph structure so that the overall performance inevitably degrades in complex environments. To address this problem, this paper proposes a novel multi-view multi-instance learning algorithm (MIL) that combines multiple context structures in a bag into a unified framework. The novel aspects are: (i) we propose a sparse -graph model that can generate different graphs with different parameters to represent various context relations in a bag, (ii) we propose a multi-view joint sparse representation that integrates these graphs into a unified framework for bag classification, and (iii) we propose a multi-view dictionary learning algorithm to obtain a multi-view graph dictionary that considers cues from all views simultaneously to improve the discrimination of the MIL. Experiments and analyses in many practical applications prove the effectiveness of the M IL. Bing Li 0001, Chunfeng Yuan, Weihua Xiong, Weiming Hu 0004, Houwen Peng, Xinmiao Ding, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Salient Object Detection via Structured Matrix DecompositionabstractLow-rank recovery models have shown potential for salient object detection, where a matrix is decomposed into a low-rank matrix representing image background and a sparse matrix identifying salient objects. Two deficiencies, however, still exist. First, previous work typically assumes the elements in the sparse matrix are mutually independent, ignoring the spatial and pattern relations of image regions. Second, when the low-rank and sparse matrices are relatively coherent, e.g., when there are similarities between the salient objects and background or when the background is complicated, it is difficult for previous models to disentangle them. To address these problems, we propose a novel structured matrix decomposition model with two structural regularizations: (1) a tree-structured sparsity-inducing regularization that captures the image structure and enforces patches from the same object to have similar saliency values, and (2) a Laplacian regularization that enlarges the gaps between salient objects and the background in feature space. Furthermore, high-level priors are integrated to guide the matrix decomposition and boost the detection. We evaluate our model for salient object detection on five challenging datasets including single object, multiple objects and complex scene images, and show competitive results as compared with 24 state-of-the-art methods in terms of seven performance metrics. Houwen Peng, Bing Li 0001, Haibin Ling, Weiming Hu 0004, Weihua Xiong, Stephen J. Maybank |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2016 | Graph Based Skeleton Motion Representation and Similarity Measurement for Action Recognition
Chunfeng Yuan, Weiming Hu 0004, Bing Li 0001, Yanning Zhang 0001 |
ECCV (7) | 4 |
| 2016 | Bootstrapping deep feature hierarchy for pornographic image recognitionabstractAutomatically recognizing pornographic images from the Web is a vital step to purify Internet environment. Inspired by the rapid developments of deep learning models, we present a deep architecture of convolutional neural network (CNN) for high accuracy pornographic image recognition. The proposed architecture is built upon existing CNNs which accepts input images of different sizes and incorporates features from different hierarchy to perform prediction. To effectively train the model, we propose a two-stage training strategy to learn the model parameters from scratch and end-to-end. During the training procedure, we also employ a hard negative sampling strategy to further reduce the false positive rate of the model. Experimental results on a large dataset demonstrate good performance of the proposed model and the effectiveness of our training strategies, with a considerable improvement over some traditional methods using hand-crafted features and deep learning method using mainstream CNN architecture. Kai Li 0022, Junliang Xing, Bing Li 0001, Weiming Hu 0004 |
ICIP | 3 |
| 2016 | A Novel Emotional Saliency Map to Model Emotional Attention Mechanism
Xinmiao Ding, Lulu Huang, Bing Li 0001, Congyan Lang, Zhen Hua |
MMM (2) | 3 |
| 2016 | Multi-Cue Illumination Estimation via a Tree-Structured Group Joint Sparse Representation
Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Brian V. Funt, Junliang Xing |
Int. J. Comput. Vis. | 1 |
| 2016 | SubMIL: Discriminative subspaces for multi-instance learning
Jiazheng Yuan, Xiankai Huang, Hongzhe Liu 0001, Bing Li 0001, Weihua Xiong |
Neurocomputing | 4 |
| 2016 | Multi-Instance Multi-Label Learning Combining Hierarchical Context and its Application to Image AnnotationabstractIn image annotation, one image is often modeled as a bag of regions (“instances”) associated with multiple labels, which is a typical application of multi-instance multi-label learning (MIML). Although lots of research has shown that the interplay embedded among instances and labels can largely boost the image annotation accuracy, most existing MIML methods consider none or partial context cues. In this paper, we propose a novel context-aware MIML model to integrate the instance context and label context into a general framework. Specially, the instance context is constructed with multiple graphs, while the label context is built up through a linear combination of several common latent conceptions that link low level features and high level semantic labels. Comparison with other leading methods on several benchmark datasets in terms of image annotation shows that our proposed method can get better performance than the state-of-the-art approaches. Xinmiao Ding, Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Bo Wang 0147 |
IEEE Trans. Multim. | 2 |
| 2016 | Multi-Perspective Cost-Sensitive Context-Aware Multi-Instance Sparse Coding and Its Application to Sensitive Video RecognitionabstractWith the development of video-sharing websites, P2P, micro-blog, mobile WAP websites, and so on, sensitive videos can be more easily accessed. Effective sensitive video recognition is necessary for web content security. Among web sensitive videos, this paper focuses on violent and horror videos. Based on color emotion and color harmony theories, we extract visual emotional features from videos. A video is viewed as a bag and each shot in the video is represented by a key frame which is treated as an instance in the bag. Then, we combine multi-instance learning (MIL) with sparse coding to recognize violent and horror videos. The resulting MIL-based model can be updated online to adapt to changing web environments. We propose a cost-sensitive context-aware multi- instance sparse coding (MI-SC) method, in which the contextual structure of the key frames is modeled using a graph, and fusion between audio and visual features is carried out by extending the classic sparse coding into cost-sensitive sparse coding. We then propose a multi-perspective multi- instance joint sparse coding (MI-J-SC) method that handles each bag of instances from an independent perspective, a contextual perspective, and a holistic perspective. The experiments demonstrate that the features with an emotional meaning are effective for violent and horror video recognition, and our cost-sensitive context-aware MI-SC and multi-perspective MI-J-SC methods outperform the traditional MIL methods and the traditional SVM and KNN-based methods. Weiming Hu 0004, Xinmiao Ding, Bing Li 0001, Fangshi Wang, Stephen J. Maybank |
IEEE Trans. Multim. | 3 |
| 2016 | Multimodal Web Aesthetics Assessment Based on Structural SVM and Multitask Fusion LearningabstractThe overall visual attributes (e.g., aesthetics) of Web pages significantly influence user experience. A beautiful and well laid out Web page greatly facilitates user access and enhances the browsing experience. In this paper, a new method is proposed to learn an assessment model for the (visual) aesthetics of Web pages. First, multimodal features (structural, local visual, global visual, and functional) of a Web page that are known to significantly affect the aesthetics of a Web page are extracted to construct a feature vector. Second, the interuser disagreement of aesthetics is analyzed and novel aesthetic representations are obtained from the multiuser ratings of a page. A structural learning algorithm is proposed for the new aesthetic representations. Third, as a Web page's functional purpose also affects the perceived aesthetics, we divide Web pages into different types using functional features, and a soft multitask fusion learning strategy is introduced to train assessment models for pages with functional purposes. Experimental results show the effectiveness of our method: 1) the combination of structural, local, and global visual features outperforms existing state-of-the-art Web aesthetic features; 2) the proposed structural learning algorithm achieves good results for the new aesthetic representations; and 3) the proposed soft multitask fusion learning strategy improves the performances of aesthetics assessment models. Ou Wu 0001, Haiqiang Zuo, Weiming Hu 0004, Bing Li 0001 |
IEEE Trans. Multim. | 4 |
| 2015 | Predicting Image Memorability by Multi-view Adaptive RegressionabstractThe images we encounter throughout our lives make different impressions on us: Some are remembered at first glance, while others are forgotten. This phenomenon is caused by the intrinsic memorability of images revealed by recent studies [5,6]. In this paper, we address the issue of automatically estimating the memorability of images by proposing a novel multi-view adaptive regression (MAR) model. The MAR model provides an effective mapping of visual features to memorability scores by taking advantage of robust feature selection and multiple feature integration. It consists of three major components: an adaptive loss function, an adaptive regularization and a multi-view modeling strategy. Moreover, we design an alternating direction method (ADM) optimization algorithm to solve the proposed objective function. Experimental results on the MIT benchmark dataset show the superiority of the proposed model compared with existing image memorability prediction methods. Houwen Peng, Kai Li 0022, Bing Li 0001, Haibin Ling, Weihua Xiong, Weiming Hu 0004 |
ACM Multimedia | 3 |
| 2015 | Horror Image Recognition Based on Context-Aware Multi-Instance LearningabstractHorror content sharing on the Web is a growing phenomenon that can interfere with our daily life and affect the mental health of those involved. As an important form of expression, horror images have their own characteristics that can evoke extreme emotions. In this paper, we present a novel context-aware multi-instance learning (CMIL) algorithm for horror image recognition. The CMIL algorithm identifies horror images and picks out the regions that cause the sensation of horror in these horror images. It obtains contextual cues among adjacent regions in an image using a random walk on a contextual graph. Borrowing the strength of the fuzzy support vector machine (FSVM), we define a heuristic optimization procedure based on the FSVM to search for the optimal classifier for the CMIL. To improve the initialization of the CMIL, we propose a novel visual saliency model based on the tensor analysis. The average saliency value of each segmented region is set as its initial fuzzy membership in the CMIL. The advantage of the tensor-based visual saliency model is that it not only adaptively selects features, but also dynamically determines fusion weights for saliency value combination from different feature subspaces. The effectiveness of the proposed CMIL model is demonstrated by its use in horror image recognition on two large-scale image sets collected from the Internet. Bing Li 0001, Weihua Xiong, Ou Wu 0001, Weiming Hu 0004, Stephen J. Maybank, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2014 | RGBD Salient Object Detection: A Benchmark and Algorithms
Houwen Peng, Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Rongrong Ji |
ECCV (3) | 2 |
| 2014 | Hierarchical sparse representation based Multi-Instance Semi-Supervised Learning with application to image categorization
Songhe Feng, Weihua Xiong, Bing Li 0001, Congyan Lang, Xiankai Huang |
Signal Process. | 3 |
| 2014 | Evaluating Combinational Illumination Estimation Methods on Real-World ImagesabstractIllumination estimation is an important component of color constancy and automatic white balancing. A number of methods of combining illumination estimates obtained from multiple subordinate illumination estimation methods now appear in the literature. These combinational methods aim to provide better illumination estimates by fusing the information embedded in the subordinate solutions. The existing combinational methods are surveyed and analyzed here with the goals of determining: 1) the effectiveness of fusing illumination estimates from multiple subordinate methods; 2) the best method of combination; 3) the underlying factors that affect the performance of a combinational method; and 4) the effectiveness of combination for illumination estimation in multiple-illuminant scenes. The various combinational methods are categorized in terms of whether or not they require supervised training and whether or not they rely on high-level scene content cues (e.g., indoor versus outdoor). Extensive tests and enhanced analyzes using three data sets of real-world images are conducted. For consistency in testing, the images were labeled according to their high-level features (3D stages, indoor/outdoor) and this label data is made available on-line. The tests reveal that the trained combinational methods (direct combination by support vector regression in particular) clearly outperform both the non-combinational methods and those combinational methods based on scene content cues. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Brian V. Funt |
IEEE Trans. Image Process. | 1 |
| 2013 | Salient Object Detection via Low-Rank and Structured Sparse Matrix DecompositionabstractSalient object detection provides an alternative solution to various image semantic understanding tasks such as object recognition, adaptive compression and image retrieval. Recently, low-rank matrix recovery (LR) theory has been introduced into saliency detection, and achieves impressed results. However, the existing LR-based models neglect the underlying structure of images, and inevitably degrade the associated performance. In this paper, we propose a Low-rank and Structured sparse Matrix Decomposition (LSMD) model for salient object detection. In the model, a tree-structured sparsity-inducing norm regularization is firstly introduced to provide a hierarchical description of the image structure to ensure the completeness of the extracted salient object. The similarity of saliency values within the salient object is then guaranteed by the $\ell _\infty$-norm. Finally, high-level priors are integrated to guide the matrix decomposition and enhance the saliency detection. Experimental results on the largest public benchmark database show that our model outperforms existing LR-based approaches and other state-of-the-art methods, which verifies the effectiveness and robustness of the structure cues in our model. Houwen Peng, Bing Li 0001, Rongrong Ji, Weiming Hu 0004, Weihua Xiong, Congyan Lang |
AAAI | 2 |
| 2013 | Illumination Estimation Based on Bilayer Sparse CodingabstractComputational color constancy is a very important topic in computer vision and has attracted many researchers' attention. Recently, lots of research has shown the effects of using high level visual content cues for improving illumination estimation. However, nearly all the existing methods are essentially combinational strategies in which image's content analysis is only used to guide the combination or selection from a variety of individual illumination estimation methods. In this paper, we propose a novel bilayer sparse coding model for illumination estimation that considers image similarity in terms of both low level color distribution and high level image scene content simultaneously. For the purpose, the image's scene content information is integrated with its color distribution to obtain optimal illumination estimation model. The experimental results on real-world image sets show that our algorithm is superior to some prevailing illumination estimation methods, even better than some combinational methods. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Houwen Peng |
CVPR | 1 |
| 2013 | Robust Object Tracking with Online Multi-lifespan Dictionary LearningabstractRecently, sparse representation has been introduced for robust object tracking. By representing the object sparsely, i.e., using only a few templates via L1-norm minimization, these so-called L1-trackers exhibit promising tracking results. In this work, we address the object template building and updating problem in these L1-tracking approaches, which has not been fully studied. We propose to perform template updating, in a new perspective, as an online incremental dictionary learning problem, which is efficiently solved through an online optimization procedure. To guarantee the robustness and adaptability of the tracking algorithm, we also propose to build a multi-lifespan dictionary model. By building target dictionaries of different life spans, effective object observations can be obtained to deal with the well-known drifting problem in tracking and thus improve the tracking accuracy. We derive effective observation models both generatively and discriminatively based on the online multi-lifespan dictionary learning model and deploy them to the Bayesian sequential estimation framework to perform tracking. The proposed approach has been extensively evaluated on ten challenging video sequences. Experimental results demonstrate the effectiveness of the online learned templates, as well as the state-of-the-art tracking performance of the proposed approach. Junliang Xing, Bing Li 0001, Weiming Hu 0004, Shuicheng Yan |
ICCV | 3 |
| 2012 | Visual Saliency Map from Tensor AnalysisabstractModeling visual saliency map of an image provides important information for image semantic understanding in many applications. Most existing computational visual saliency models follow a bottom-up framework that generates independent saliency map in each selected visual feature space and combines them in a proper way. Two big challenges to be addressed explicitly in these methods are (1) which features should be extracted for all pixels of the input image and (2) how to dynamically determine importance of the saliency map generated in each feature space. In order to address these problems, we present a novel saliency map computational model based on tensor decomposition and reconstruction. Tensor representation and analysis not only explicitly represent image's color values but also imply two important relationships inherent to color image. One is reflecting spatial correlations between pixels and the other one is representing interplay between color channels. Therefore, saliency map generator based on the proposed model can adaptively find the most suitable features and their combinational coefficients for each pixel. Experiments on a synthetic image set and a real image set show that our method is superior or comparable to other prevailing saliency map models. Bing Li 0001, Weihua Xiong, Weiming Hu 0004 |
AAAI | 1 |
| 2012 | Horror Video Scene Recognition Based on Multi-view Multi-instance Learning
Xinmiao Ding, Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Zhenchong Wang |
ACCV (3) | 2 |
| 2012 | Context-aware horror video scene recognition via cost-sensitive sparse coding
Xinmiao Ding, Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Zhenchong Wang |
ICPR | 2 |
| 2012 | Towards relevance and saliency ranking of image tagsabstractSocial image tag ranking has emerged as an important research topic recently due to its potential application on web image search. This paper presents an adaptive all-season tag ranking algorithm which can handle the images with and without distinct object(s) using different tag ranking strategies. Firstly, based on saliency map derived from the visual attention model, a linear SVM is trained to pre-classify an image as attentive or non-attentive category by using the gray histogram descriptor on the corresponding saliency map. Then, an image with distinct object is processed by an attention-driven tag saliency ranking algorithm emphasizing distinct object. On the other hand, an image without distinct object is processed by the tag relevance ranking algorithm via the sparse representation based neighbor-voting strategy. Such adaptive ranking strategy can be regarded as taking full advantage of existing tag ranking paradigms. Experiments conducted on well-known image data sets demonstrate the effectiveness and efficiency of the proposed framework. Songhe Feng, Congyan Lang, Bing Li 0001 |
ACM Multimedia | 3 |
| 2012 | Scaring or pleasing: exploit emotional impact of an imageabstractAutomatic image emotion analysis has emerged as a hot topic due to its potential application on high-level image understanding. Considering the fact that the emotion evoked by an image is not only from its global appearance but also interplays among local regions, we propose a novel affective image classification system based on bilayer sparse representation (BSR). The BSR model contains two layers: The global sparse representation (GSR) is to define global similarities between a test image and all the training images; and the local sparse representation (LSR) is to define similarities of local regions' appearances and their co-occurrence between a test image and all the training images. The experiments on real data sets demonstrate that our system is effective on image emotion recognition. Bing Li 0001, Songhe Feng, Weihua Xiong, Weiming Hu 0004 |
ACM Multimedia | 1 |
| 2012 | Context-aware affective images classification based on bilayer sparse representationabstractIn image understanding, the automatic recognition of emotion in an image is becoming important from an applicative viewpoint. Considering the fact that the emotion evoked by an image is not only from its global appearance but also interplays among local regions, we propose a novel context-aware classification model based on bilayer sparse representation (BSR) that simultaneously takes the local context and global-local context into account. The BSR model contains two layers: global sparse representation (GSR) and local sparse representation (LSR). The GSR is to define global similarities between a test image and all training images; while the LSR is to define similarities of local regions' appearances and their co-occurrence between a test image and all training images. The experiments on two data sets demonstrate that our method is effective on affective images classification. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Xinmiao Ding |
ACM Multimedia | 1 |
| 2012 | An evidential reasoning based classification algorithm and its application for face recognition with class noise
Xiaodong Wang 0011, Fang Liu 0001, Licheng Jiao, Jingjing Yu 0001, Bing Li 0001, Jianrui Chen 0002, Jiao Wu 0002, Fanhua Shang |
Pattern Recognit. | 6 |
| 2012 | Efficient Clustering Aggregation Based on Data FragmentsabstractClustering aggregation, known as clustering ensembles, has emerged as a powerful technique for combining different clustering results to obtain a single better clustering. Existing clustering aggregation algorithms are applied directly to data points, in what is referred to as the point-based approach. The algorithms are inefficient if the number of data points is large. We define an efficient approach for clustering aggregation based on data fragments. In this fragment-based approach, a data fragment is any subset of the data that is not split by any of the clustering results. To establish the theoretical bases of the proposed approach, we prove that clustering aggregation can be performed directly on data fragments under two widely used goodness measures for clustering aggregation taken from the literature. Three new clustering aggregation algorithms are described. The experimental results obtained using several public data sets show that the new algorithms have lower computational complexity than three well-known existing point-based clustering aggregation algorithms (Agglomerative, Furthest, and LocalSearch); nevertheless, the new algorithms do not sacrifice the accuracy. Ou Wu 0001, Weiming Hu 0004, Stephen J. Maybank, Mingliang Zhu, Bing Li 0001 |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2011 | Evaluating combinational color constancy methods on real-world imagesabstractLight color estimation is crucial to the color constancy problem. Past decades have witnessed great progress in solving this problem. Contrary to traditional methods, many researchers propose a variety of combinational color constancy methods through applying different color constancy mathematical models on an image simultaneously and then give out a final estimation in diverse ways. Although many comprehensive evaluations or reviews about color constancy methods are available, few focus on combinational strategies. In this paper, we survey some prevailing combinational strategies systematically; divide them into three categories and compare them qualitatively on three real-world image data sets in terms of the angular error and the perceptual Euclidean distance. The experimental results show that combinational strategies with training procedure always produces better performance. Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Ou Wu 0001 |
CVPR | 1 |
| 2011 | Horror video scene recognition via Multiple-Instance learningabstractAlong with the ever-growing Web comes the proliferation of objectionable content, such as pornography, violence, horror information, etc. Horror videos, whose threat to childrens health is no less than pornographic video, are sometimes neglected by existing Web filtering tools. Consequently, an effective horror video filtering tool is necessary for preventing children from accessing these harmful horror videos. In this paper, by introducing color emotion and color harmony theories, we propose a horror video scenes recognition algorithm. Firstly, the video scenes are decomposed into a set of shots. Then we extract the visual features, audio features and emotional features of each shot, the video scene is viewed as a bag and each shot is treated as an instance of the corresponding bag. Finally, by combining the three features, the horror video scenes are recognized by the Multiple-Instance learning(MIL). According to the experimental results on diverse video scenes, the proposed scheme based on the emotional perception could effectively deal with the horror video scene recognition and promising results are achieved. Bing Li 0001, Weiming Hu 0004, Ou Wu 0001 |
ICASSP | 2 |
| 2011 | Context-Aware Multi-instance Learning Based on Hierarchical Sparse RepresentationabstractMulti-instance learning (MIL), a variant of supervised learning framework, has been applied in many applications. More recently, researchers focus on two important issues for MIL: Instances' contextual structures representation in the same bag and online MIL schemes. In this paper, we present an effective context-aware multi-instance learning technique using a hierarchical sparse representation (HSR-MIL) that addresses the two challenges simultaneously. We firstly construct the inner contextual structure among instances in the same bag based on a novel sparse ε-graph. We then propose a graph kernel based sparse bag classifier through a modified kernel sparse coding in higher-dimension feature space. At last, the HSR-MIL approach is extended to achieve online learning manner with an incremental kernel matrix update scheme. The experiments on several data sets demonstrate that our method has better performances and online learning ability. Bing Li 0001, Weihua Xiong, Weiming Hu 0004 |
ICDM | 1 |
| 2011 | Web Horror Image Recognition Based on Context-Aware Multi-instance LearningabstractAlong with the ever-growing Web, horror contents sharing in the Internet has interfered with our daily life and affected our, especially children's, health. Therefore horror image recognition is becoming more important for web objectionable content filtering. This paper presents a novel context-aware multi-instance learning (CMIL) model for this task. This work is distinguished by three key contributions. Firstly, the traditional multi-instance learning is extended to context-aware multi-instance learning model through integrating an undirected graph in each bag that represents contextual relationships among instances. Secondly, by introducing a novel energy function, a heuristic optimization algorithm based on Fuzzy Support Vector Machine (FSVM) is given out to find the optimal classifier on CMIL. Finally, the CMIL is applied to recognize horror images. Experimental results on an image set collected from the Internet show that the proposed method is effective on horror image recognition. Bing Li 0001, Weihua Xiong, Weiming Hu 0004 |
ICDM | 1 |
| 2011 | Evaluating the visual quality of web pages using a computational aesthetic approachabstractCurrent Web mining explores useful and valuable information (content) online for users. However, there is scant research on the overall visual aspect of Web pages, even though visual elements such as aesthetics significantly influence user experience. A beautiful and well-laid out Web page greatly facilitates users' accessing and enhances browsing experiences.We use "visual quality (VisQ)" to denote the aesthetics of Web pages. In this paper, a computational aesthetics approach is proposed to learn the evaluation model for the visual quality of Web pages. First, a Web page layout extraction algorithm (V-LBE) is introduced to partition a Web page into major layout blocks. Then, regarding a Web page as a semi-structured image, features (e.g., layout,visual complexity, colorfulness) known to significantly affect the visual quality of a Web page are extracted to construct a feature vector. We present a multi-cost-sensitive learning for visual quality classification and a multi-value regression for visual quality score assignment. Our experiments compare the extracted features and conclude that the Web page's layout visual features (LV) and text visual features (TV) are the primary affecting factors toward Web page's visual quality. The performance of the learned visual quality classifier is close to some persons'. The learned regression function also achieves promising results. Ou Wu 0001, Yunfei Chen 0002, Bing Li 0001, Weiming Hu 0004 |
WSDM | 3 |
| 2010 | Horror Image Recognition Based on Emotional Attention
Bing Li 0001, Weiming Hu 0004, Weihua Xiong, Ou Wu 0001, Wei Li 0034 |
ACCV (2) | 1 |
| 2010 | Occlusion Handling with ℓ1-Regularized Sparse Reconstruction
Wei Li 0034, Bing Li 0001, Xiaoqin Zhang 0002, Weiming Hu 0004, Hanzi Wang, Guan Luo |
ACCV (4) | 2 |
| 2010 | Top-Down Cues for Event Recognition
Li Li 0010, Chunfeng Yuan, Weiming Hu 0004, Bing Li 0001 |
ACCV (3) | 4 |
| 2010 | Group ranking with application to image retrievalabstractMany existing ranking-related information processing applications can be summarized into one theoretical problem called group ranking (GR). A simple average-ranking approach is usually applied to GR. Although the approach seems reasonable, no theoretical analysis about its intrinsic mechanism has been presented, increasing the difficulty of evaluating the ranking results. This study provides a formal analysis for GR. We first construct an objective function for the GR problem, and discover that each GR problem can be transformed into a rank aggregation problem whose objective function is proved to be equal to the objective function of GR. As a consequence, the average-ranking approach can be explained by two well-known rank aggregation techniques. We incorporate two other effective rank aggregation methods into the GR problem and obtain two new GR algorithms. We apply the GR algorithms into image retrieval to diversify the image search results returned by search engines. Experimental results show the effectiveness of the proposed GR algorithms. Ou Wu 0001, Weiming Hu 0004, Bing Li 0001 |
CIKM | 3 |
| 2010 | Horror movie scene recognition based on emotional perceptionabstractThe number of video clips available online is growing at a tremendous pace. Meanwhile, the video scenes of pornography, violence and horror permeate the whole Web. Horror videos, whose threat to children's health is no less than pornographic video, are sometimes neglected by existing Web filtering tools. Consequently, an effective horror video filtering tool is necessary for preventing children from accessing these horror videos. In this paper, by introducing color emotion and color harmony theories, we propose a horror video scene recognition algorithm. Firstly, the video scenes are decomposed into a set of shots. Then we extract the visual features, audio features and color emotion features of each shot. Finally, by combining the three features, the horror video scenes are recognized by the Support Vector Machine (SVM) classifier. According to the experimental results on diverse video scenes, the proposed scheme based on the emotional perception could deal effectively with the horror video scene recognition and promising results are achieved. Bing Li 0001, Weiming Hu 0004, Ou Wu 0001 |
ICIP | 2 |
| 2010 | Event Recognition Based on Top-Down Motion AttentionabstractHow to fuse static and dynamic information is a key issue in event analysis. In this paper, a top-down motion guided fusing method is proposed for recognizing events in an unconstrained news video. In the method, the static information is represented as a Bag-of-SIFT-features and motion information is employed to generate event specific attention map to direct the sampling of the interest points. We build class-specific motion histograms for each event so as to give more weight on the interest points that are discriminative to the corresponding event. Experimental results on TRECVID 2005 video corpus demonstrate that the proposed method can improve the mean average accuracy of recognition. Li Li 0010, Weiming Hu 0004, Bing Li 0001, Chunfeng Yuan, Pengfei Zhu 0001, Wanqing Li 0001 |
ICPR | 3 |
| 2010 | Identifying Multi-instance OutliersabstractThis paper studies a new data mining problem called multi-instance outlier identification. This problem arises in tasks where each sample consists of many alternative feature vectors (instances) that describe it. This paper defines the multi-instance outliers and analyzes the basic types of multi-instance outliers. Two general identification approaches are proposed based on the state-of-the-art (single-instance) outlier detector LOF (local outlier factor). One approach utilizes the underlying mechanism of the kernel method and plunges the set distance into LOF to detect the multi-instance outliers. The other approach takes each instance's neighborhood into account. Based on the two approaches, four concrete multi-instance outlier detectors are then introduced. We conduct experiments over four synthetic data collections and three real-world data collections (two Musk data sets [22, 23] and a hard-drive inspection data set [24]). The experimental results show that the proposed multi-instance outlier detectors are effective while the algorithms that ignore the multi-instance settings perform poorly. Especially, the results on the two Musk sets are consistent with the multi-instance learning results; the results on the hard-drive inspection data set demonstrate that multi-instance outlier identification is promising for real applications. Ou Wu 0001, Weiming Hu 0004, Bing Li 0001, Mingliang Zhu |
SDM | 4 |
| 2010 | Learning to evaluate the visual quality of web pagesabstractA beautiful and well-laid out Web page greatly facilitates users' accessing and enhances browsing experiences. We use "visual quality (VQ)" to denote the aesthetics of Web pages. In this paper, a computational aesthetics approach is proposed to learn the evaluation model for the visual quality of Web pages. First, a Web page layout extraction algorithm (V-LBE) is introduced to partition a Web page into major layout blocks. Then, regarding a Web page as a semi-structured image, features known to significantly affect the visual quality of a Web page are extracted to construct a feature vector. The experimental results show the initial success of our approach. Potential applications include Web search and Web design. Ou Wu 0001, Yunfei Chen 0002, Bing Li 0001, Weiming Hu 0004 |
WWW | 3 |
| 2010 | A supervised combination strategy for illumination chromaticity estimationabstractColor constancy is an important perceptual ability of humans to recover the color of objects invariant of light information. It is also necessary for a robust machine vision system. Until now, a number of color constancy algorithms have been proposed in the literature. In particular, the edge-based color constancy uses the edge of an image to estimate light color. It is shown to be a rich framework that can represent many existing illumination estimation solutions with various parameter settings. However, color constancy is an ill-posed problem; every algorithm is always given out under some assumptions and can only produce the best performance when these assumptions are satisfied. In this article, we have investigated a combination strategy relying on the Extreme Learning Machine (ELM) technique that integrates the output of edge-based color constancy with multiple parameters. Experiments on real image data sets show that the proposed method works better than most single-color constancy methods and even some current state-of-the-art color constancy combination strategies. Bing Li 0001, Weihua Xiong, De Xu, Hong Bao |
ACM Trans. Appl. Percept. | 1 |
| 2007 | Visual Perception Theory Guided Depth Motion Estimation
Bing Li 0001, De Xu, Songhe Feng, Fangshi Wang |
MMM (1) | 1 |