VLDB 2026 Research / reviewers in the wild / expert
Yahong Han
dblp:15/6265
· DBLP profile ↗
193ranked-venue papers
21as first author
89since 2021 · last 2026
0000-0003-2768-1398ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 132 · 15 first-author · 54 since 2021Artificial intelligence and machine learning · 66 · 5 first-author · 33 since 2021Databases, data management, data science and information retrieval · 14 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 8 since 2021Computer networks · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object DetectionabstractIn this paper, we focus on Single-Domain Generalized Object Detection (Single-DGOD), aiming to transfer a detector trained on one source domain to multiple unknown domains. Existing methods for Single-DGOD typically rely on discrete data augmentation or static perturbation methods to expand data diversity, thereby mitigating the lack of access to target domain data. However, in real-world scenarios such as changes in weather or lighting conditions, domain shifts often occur continuously and gradually. Discrete augmentations and static perturbations fail to effectively capture the dynamic variation of feature distributions, thereby limiting the model's ability to perceive fine-grained cross-domain differences. To this end, we propose a new method, i.e., Liquid Temporal Feature Evolution, which simulates the progressive evolution of features from the source domain to simulated latent distributions by incorporating temporal modeling and liquid neural network–driven parameter adjustment. Specifically, we introduce controllable Gaussian noise injection and multi-scale Gaussian blurring to simulate initial feature perturbations, followed by temporal modeling and a liquid parameter adjustment mechanism to generate adaptive modulation parameters, enabling a smooth and continuous adaptation across domains. By capturing progressive cross-domain feature evolution and dynamically regulating adaptation paths, our method bridges the source-unknown domain distribution gap, significantly boosting generalization and robustness to unseen shifts. Significant performance improvements on the Diverse Weather dataset and Real-to-Art benchmark demonstrate the superiority of our method. Yang Li 0251, Aming Wu, Yahong Han |
AAAI | 4 |
| 2026 | Transferable diffusion transformer for low-light image enhancement
Runhua Jiang, Jianbin Zhao, Kaikai Hao, Yahong Han |
Multim. Syst. | 7 |
| 2026 | Enhanced visual prompt meets low-light saliency detection
Nana Yu, Jie Wang 0095, Yahong Han |
Pattern Recognit. | 3 |
| 2026 | Low-Light Salient Object Detection via Representation DecouplingabstractIn low-light scenes, images often suffer low contrast, poor distinction between the target and the background, and loss of regional information. This makes it challenging for Salient Object Detection (SOD) algorithms to identify and locate the targets. Most existing methods address low-light SOD by first enhancing low-light images and then performing saliency detection. However, splitting these into two sub-tasks may result in negative transfer effects. Additionally, some methods attempt to integrate these two sub-tasks into an end-to-end framework. However, conflicts may arise during the training process due to the differing features required by low-light enhancement and saliency detection. To address this conflict, We propose a decoupled representation network called DRNet, whose core is the construction of a Visual Center Decoupler (VCD). The VCD decouples the learned representations into enhancement-specific and SOD-specific embeddings. This module provides an implicit way to balance the specific requirements of the two subtasks. On the one hand, the enhancement-specific embeddings use RGB illumination constraints and Local Binary Patterns (LBP) feature aggregation to constrain illumination and maintain texture feature stability. On the other hand, the SOD-specific embeddings utilize dynamic multi-scale convolutions to integrate fine-grained details required at multiple scales. Finally, to validate the performance advantages of DRNet, we conduct extensive comparative experiments between the proposed DRNet and existing single-modal methods. Comparative experiments with some representative bi-modal methods further demonstrate the merits of our proposed method. Nana Yu, Jie Wang 0095, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Fourier-KAN: Feature Distribution Decomposition and Recombination for Unknown-Domain Object DetectionabstractSingle-domain Generalized Object Detection (Single-DGOD) is recently proposed, aiming to transfer a detector to multiple unknown domains never seen during training. For this task, the challenge mainly lies in how to utilize the single data distribution from the source domain to generalize across multiple unknown domains with diverse data distributions. Accordingly, the challenge could be addressed by expanding the data distribution of the source domain. In this paper, we propose feature recombination from a frequency perspective to generate a series of recombined features that exhibit diversity in style and rich variation in content features. Specifically, we propose a new method, Fourier-KAN Feature Recombination, which utilizes the Fast Fourier Transform (FFT) to decompose features into amplitude and phase components. Then we apply the Kolmogorov-Arnold theorem to further decompose these components into linear combinations of multiple base distributions. Finally, through multi-level recombination, we generate a series of recombined features with diverse distributions, effectively emulating deep cross-domain variations in feature levels and strengthening the model's generalization ability to unknown domains. Our method demonstrates strong adaptability to both two-stage and single-stage detection frameworks. Experimental results show that on the Diverse Weather and Real-to-Art benchmarks, our approach not only achieves outstanding detection accuracy but also significantly enhances the model's generalization ability, all while maintaining excellent real-time performance. Our code is available at https://github.com/2490o/Fourier-KAN. Yang Li 0251, Aming Wu, Yahong Han |
IEEE Trans. Image Process. | 4 |
| 2026 | Reinforcement and Complementary Space Search for Multimodal Black-Box AttackabstractCollaborative attack, an adversarial attack method that could attack multiple modalities simultaneously, has attracted significant interest in the research community. Vision-Language Pre-training (VLP) models have shown their vulnerability to collaborative attacks, as they perform tasks through the interaction of images and texts. However, current collaborative attacks against VLP models are mainly focused on white-box and transfer-based, which require full or partial prior knowledge of victim VLP models. Besides that, VLP collaborative attacks iteratively optimize the adversarial example on one set of image-text pairs, which may ignore other potential attack targets and cause local overfitting. In this paper, we propose a reinforcement and complementary space search attack (RCSS-Attack) as a new decision-based attack for multimodal models. Based on Markov Decision Process (MDP), we adaptatively model the adversarial attack process as a policy search process, which could enlarge the search space. We also utilize step-level constraint optimization to design a complementary space search policy. Benefiting from the wide-range search capability of reinforcement learning, the RCSS-Attack uses the reward mechanism to select the optimal direction between different image-text pairs, which could relieve the overfitting problem. In the experiments, we evaluate the RCSS-Attack on 4 V+L tasks with 3 datasets and 4 metrics. Boyuan Zhang 0003, Yahong Han |
IEEE Trans. Multim. | 4 |
| 2025 | Visual Consensus Prompting for Co-Salient Object DetectionabstractExisting co-salient object detection (CoSOD) methods generally employ a three-stage architecture (i.e., encoding, consensus extraction & dispersion, and prediction) along with a typical full fine-tuning paradigm. Although they yield certain benefits, they exhibit two notable limitations: 1) This architecture relies on encoded features to facilitate consensus extraction, but the meticulously extracted consensus does not provide timely guidance to the encoding stage. 2) This paradigm involves globally updating all parameters of the model, which is parameter-inefficient and hinders the effective representation of knowledge within the foundation model for this task. Therefore, in this paper, we propose an interaction-effective and parameter-efficient concise architecture for the CoSOD task, addressing two key limitations. It introduces, for the first time, a parameter-efficient prompt tuning paradigm and seamlessly embeds consensus into the prompts to formulate task-specific Visual Consensus Prompts (VCP). Our VCP aims to induce the frozen foundation model to perform better on CoSOD tasks by formulating task-specific visual consensus prompts with minimized tunable parameters. Concretely, the primary insight of the purposeful Consensus Prompt Generator (CPG) is to enforce limited tunable parameters to focus on co-salient representations and generate consensus prompts. The formulated Consensus Prompt Disperser (CPD) leverages consensus prompts to form task-specific visual consensus prompts, thereby arousing the powerful potential of pre-trained models in addressing CoSOD tasks. Extensive experiments demonstrate that our concise VCP outperforms 13 cutting-edge full fine-tuning models, achieving the new state of the art (with 6.8% improvement in Fmmetrics on the most challenging CoCA dataset). Source code has been available at https://github.com/WJ-CV/VCP. Jie Wang 0095, Nana Yu, Yahong Han |
CVPR | 4 |
| 2025 | Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionabstractRecently, a task of Single-Domain Generalized Object Detection (Single-DGOD) is proposed, aiming to generalize a detector to multiple unknown domains never seen before during training. Due to the unavailability of target-domain data, some methods leverage the multimodal capabilities of vision-language models, using textual prompts to estimate cross-domain information, enhancing the model’s generalization capability. These methods typically use a single textual prompt, referred to as the one-step prompt method. However, when dealing with complex styles, such as the combination of rain and night, we observe that the performance of the one-step prompt method tends to be relatively weak. The reason may be that many scenes incorporate a single style and a combination of multiple styles. The one- step prompt method may not effectively synthesize combined information involving various styles. To address this limitation, we propose a new method, i.e., Style Evolving along Chain-of-Thought, which aims to progressively integrate and expand style information along the chain of thought, enabling the continual evolution of styles. Specifically, by progressively refining style descriptions and guiding the diverse evolution of styles, this method enhances the simulation of various style characteristics, enabling the model to learn and adapt to subtle differences more effectively. Additionally, it exposes the model to a broader range of style features with different data distributions, thereby enhancing its generalization capability in unseen domains. The significant performance gains over five adverse-weather scenarios and the Real to Art benchmark demonstrate the superiorities of our method. Our code is available at https://github.com/ZZ2490/SE-COT. Aming Wu, Yahong Han |
CVPR | 3 |
| 2025 | Process Adaptive Learning for Visual-Language Navigation
Chaoqi Gao, Boyuan Zhang 0003, Yahong Han |
ICANN (4) | 3 |
| 2025 | Coupling the Generator with Teacher for Effective Data-Free Knowledge Distillation
Xu Chen 0053, Yang Li 0251, Yahong Han, Guangquan Xu, Jialie Shen 0001 |
ICCV | 3 |
| 2025 | Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic ScenariosabstractIn practice, environments constantly change over time and space, posing significant challenges for object detectors trained based on a closed-set assumption, i.e., training and test data share the same distribution. To this end, continual test-time adaptation has attracted much attention, aiming to improve detectors' generalization by fine-tuning a few specific parameters, e.g., BatchNorm layers. However, based on a small number of test images, fine-tuning certain parameters may affect the representation ability of other fixed parameters, leading to performance degradation. Instead, we explore a new mechanism, i.e., converting the fine-tuning process to a specific-parameter generation. Particularly, we first design a dual-path LoRA-based domain-aware adapter that disentangles features into domain-invariant and domain-specific components, enabling efficient adaptation. Additionally, a conditional diffusion-based parameter generation mechanism is presented to synthesize the adapter's parameters based on the current environment, preventing the optimization from getting stuck in local optima. Finally, we propose a class-centered optimal transport alignment method to mitigate catastrophic forgetting. Extensive experiments conducted on various continuous domain adaptive object detection tasks demonstrate the effectiveness. Meanwhile, visualization results show that the representation extracted by the generated parameters can capture more object-related information and strengthen the generalization ability. Deng Li 0003, Aming Wu, Yang Li 0251, Yaowei Wang 0001, Yahong Han |
ICCV | 5 |
| 2025 | Unknown Text Learning for Clip-Based Few-Shot Open-Set Recognition
Qilong Wang 0001, Bing Cao 0002, Qinghua Hu, Yahong Han |
ICCV | 5 |
| 2025 | LBA: Multi-Scale Video Segment Sampling for Open-Ended Video Question Answering
Yahong Han |
ICIC (6) | 2 |
| 2025 | Object-Centric Feature Enrichment for Single-Domain Generalized Object DetectionabstractSingle-domain generalized object detection (S-DGOD) aims to train a model on a single source domain while ensuring robust performance across unseen target domains. Existing methods primarily focus on improving model generalization through image-space augmentations; however, such approaches often result in limited diversity in the generated data. Additionally, these methods struggle to avoid over-reliance on background or contextual cues, which hinders the model’s generalization to unseen domains. To address these challenges, we propose synthesizing latent features in the feature space. Specifically, we introduce a novel Object-Centric Feature Enrichment (OCFE) framework that enhances the diversity of data distribution in the feature space and reduces interference from background and non-object regions by augmenting object features. Furthermore, we apply Deep Consistency Projection (DCP) to the enriched object features, enabling the learning of domain-invariant representations. Experimental results demonstrate the efficacy and superiority of the proposed method. Shukuan Yuan, Yahong Han |
ICME | 3 |
| 2025 | Ex Pede Herculem, Predicting Global Actionness Curve from Local ClipsabstractDense multi-label action detection in untrimmed long videos is a formidable task, with end-to-end training particularly challenging due to computational constraints, typically involving separate stages of off-the-shelf feature extraction and subsequent global modeling for action prediction.Existing methods fail to optimize all modules jointly for better performance. We introduce FreETAD, a Frequency-based End-to-end Temporal Action Detection approach, which shifts the focus from local actionness scores to frequency component estimation. Using the short-term Fourier Transform, FreETAD reconstructs the global action curve seamlessly. With a DETR-like decoder and frequency-encoded vectors for queries, it enhances multi-scale time-frequency interactions. FreETAD leverages end-to-end training effectively, boosting the mAP by 1.5% on Charades and 2.7% on MultiTHUMOS. Xu Chen 0053, Yang Li 0251, Yahong Han, Jialie Shen 0001 |
ACM Multimedia | 3 |
| 2025 | Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and ReasoningabstractIn this paper, we focus on Novel Class Discovery for Point Cloud Segmentation (3D-NCD), aiming to learn a model that can segment unlabeled (novel) 3D classes using only the supervision from labeled (base) 3D classes. The key to this task is to setup the exact correlations between the point representations and their base class labels, as well as the representation correlations between the points from base and novel classes. A coarse or statistical correlation learning may lead to the confusion in novel class inference. lf we impose a causal relationship as a strong correlated constraint upon the learning process, the essential point cloud representations that accurately correspond to the classes should be uncovered. To this end, we introduce a structural causal model (SCM) to re-formalize the 3D-NCD problem and propose a new method, i.e., Joint Learning of Causal Representation and Reasoning. Specifically, we first analyze hidden confounders in the base class representations and the causal relationships between the base and novel classes through SCM. We devise a causal representation prototype that eliminates confounders to capture the causal representations of base classes. A graph structure is then used to model the causal relationships between the base classes' causal representation prototypes and the novel class prototypes, enabling causal reasoning from base to novel classes. Extensive experiments and visualization results on 3D and 2D NCD semantic segmentation demonstrate the superiorities of our method. Yang Li 0251, Aming Wu, Yahong Han |
NeurIPS | 4 |
| 2025 | Robust source-free domain adaptation with anti-adversarial samples training
Zhirui Wang 0001, Liu Yang 0010, Yahong Han |
Neurocomputing | 3 |
| 2025 | Prototype-guided cross-task knowledge distillationabstractRecently, large-scale pretrained models have revealed their benefits in various tasks. However, due to the enormous computation complexity and storage demands, it is challenging to apply large-scale models to real scenarios. Existing knowledge distillation methods require mainly the teacher model and the student model to share the same label space, which restricts their application in real scenarios. To alleviate the constraint of different label spaces, we propose a prototype-guided cross-task knowledge distillation (ProC-KD) method to migrate the intrinsic local-level object knowledge of the teacher network to various task scenarios. First, to better learn the generalized knowledge in cross-task scenarios, we present a prototype learning module to learn the invariant intrinsic local representation of objects from the teacher network. Second, for diverse downstream tasks, a task-adaptive feature augmentation module is proposed to enhance the student network features with the learned generalization prototype representations and guide the learning of the student network to improve its generalization ability. Experimental results on various visual tasks demonstrate the effectiveness of our approach for cross-task knowledge distillation scenarios. Deng Li 0003, Aming Wu, Yahong Han |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | A comprehensive survey of physical adversarial vulnerabilities in autonomous driving systemsabstractAutonomous driving systems (ADSs) have attracted wide attention in the machine learning communities. With the help of deep neural networks (DNNs), ADSs have shown both satisfactory performance under significant uncertainties in the environment and the ability to compensate for system failures without external intervention. However, the vulnerability of ADSs has raised concerns since DNNs have been proven vulnerable to adversarial attacks. In this paper, we present a comprehensive survey of current physical adversarial vulnerabilities in ADSs. We first divide the physical adversarial attack methods and defense methods by their restrictions of deployment into three scenarios: the real-world, simulator-based, and digital-world scenarios. Then, we consider the adversarial vulnerabilities that focus on various sensors in ADSs and separate them as camera-based, light detection and ranging (LiDAR) based, and multifusion-based attacks. Subsequently, we divide the attack tasks by traffic elements. For the physical defenses, we establish the taxonomy with reference to input image preprocessing, adversarial example detection, and model enhancement for the DNN models to achieve full coverage of the adversarial defenses. Based on the above survey, we finally discuss the challenges in this research field and provide further outlook on future directions. Boyuan Zhang 0003, Yang Zhai, Yahong Han, Qinghua Hu |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | Dynamic prototype-guided structural information maintaining for unsupervised domain adaptation
Deng Li 0003, Yahong Han |
Pattern Anal. Appl. | 4 |
| 2025 | Progressive expansion for semi-supervised bi-modal salient object detection
Jie Wang 0095, Nana Yu, Yahong Han |
Pattern Recognit. | 4 |
| 2025 | Information disentanglement for unsupervised domain adaptive Oracle Bone Inscriptions detection
Yongge Liu, Deng Li 0003, Xu Chen 0053, Runhua Jiang, Yahong Han |
Signal Process. Image Commun. | 6 |
| 2025 | Explicitly Disentangling and Exclusively Fusing for Semi-Supervised Bi-Modal Salient Object DetectionabstractBi-modal (RGB-T and RGB-D) salient object detection (SOD) aims to enhance detection performance by leveraging the complementary information between modalities. While significant progress has been made, two major limitations persist. Firstly, mainstream fully supervised methods come with a substantial burden of manual annotation, while weakly supervised or unsupervised methods struggle to achieve satisfactory performance. Secondly, the indiscriminate modeling of local detailed information (object edge) and global contextual information (object body) often results in predicted objects with incomplete edges or inconsistent internal representations. In this work, we propose a novel paradigm to effectively alleviate the above limitations. Specifically, we first enhance the consistency regularization strategy to build a basic semi-supervised architecture for the bi-modal SOD task, which ensures that the model can benefit from massive unlabeled samples while effectively alleviating the annotation burden. Secondly, to ensure detection performance (i.e., complete edges and consistent bodies), we disentangle the SOD task into two parallel sub-tasks: edge integrity fusion prediction and body consistency fusion prediction. Achieving these tasks involves two key steps: 1) the explicitly disentangling scheme decouples salient object features into edge and body features, and 2) the exclusively fusing scheme performs exclusive integrity or consistency fusion for each of them. Eventually, our approach demonstrates significant competitiveness compared to 26 fully supervised methods, while effectively alleviating 90% of the annotation burden. Furthermore, it holds a substantial advantage over 15 non-fully supervised methods. Jie Wang 0095, Xiangji Kong, Nana Yu, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Single-Group Generalized RGB and RGB-D Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) aims to segment the co-occurring salient objects in a given group of relevant images. Existing methods typically rely on extensive group training data to enhance the model’s CoSOD capabilities. However, fitting prior knowledge of the extensive group results in a significant performance gap between the seen and out-of-sample image groups. Relaxing such a fitting with fewer prior groups may improve the generalization ability of CoSOD while alleviating the annotation burdens. Hence, it is essential to explore the use of fewer groups during the training phase, such as using only single group, to pursue a highly generalized CoSOD model. We term this new setting as Sg-CoSOD, which aims to train a model using only a single group and effectively apply it to any unseen RGB and RGB-D CoSOD test groups. Towards Sg-CoSOD, it is important to ensure detection performance with limited data and release class dependency with only a single-group. Thus, we present a method, i.e., cross-excitation between saliency and ‘Co’, which decouples the CoSOD task into two parallel branches: ‘Co’ To Saliency (CTS) and Saliency To ‘Co’ (STC). The CTS branch focuses on mining group consensus to guide image co-saliency predictions, while the STC branch is dedicated to using saliency priors to motivate group consensus mining. Furthermore, we propose a Class-Agnostic Triplet (CAT) loss to constrain intra-group consensus while suppressing the model from acquiring class prior knowledge. Extensive experiments on RGB and RGB-D CoSOD tasks with multiple unknown groups show that our model has higher generalization capabilities (e.g., for large-scale datasets CoSOD3k and CoSal1k with multiple generalized groups, we obtain a gain of over 15% in$F_{m}$). Further experimental analyses also reveal that the proposed Sg-CoSOD paradigm has significant potential and promising prospects. Jie Wang 0095, Nana Yu, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | AdvNeRF: Generating 3D Adversarial Meshes With NeRF to Fool Driving VehiclesabstractAdversarial attacks on deep neural networks (DNNs) have raised significant concerns, particularly in safety-critical applications such as autonomous driving. Autonomous vehicles rely on both vision and LiDAR sensors to provide accurate 3D visual perception of their surroundings. However, adversarial vulnerabilities in these models pose several risks, as they can lead to misinterpretation of sensor data, ultimately endangering safety. While substantial research has been devoted to image-level adversarial attacks, these efforts are predominantly confined to 2D-pixel spaces, lacking physical realism and applicability in the 3D world. To address these limitations, we introduce AdvNeRF, a groundbreaking approach for generating 3D adversarial meshes that effectively target both vision and LiDAR models simultaneously. AdvNeRF is a Transferable Target Adversarial Attack that leverages Neural Radiance Fields (NeRF) to achieve its objectives. NeRF ensures the creation of high-quality adversarial objects and enhances attack performance by maintaining consistency across unseen viewpoints, making the adversarial examples robust from multiple angles. By integrating NeRF, our method represents a leap forward in improving the robustness and effectiveness of 3D adversarial attacks. Experimental results validate the superior performance of AdvNeRF, demonstrating its ability to degrade the accuracy of 3D object detectors under various conditions. These findings highlight the critical implications of AdvNeRF, emphasizing its potential to consistently undermine the perception systems of autonomous vehicles across different perspectives, thus marking an advancement in the field of adversarial attacks and 3D perception security. Boyuan Zhang 0003, Yahong Han, Qinghua Hu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Latent Factor Modeling With Expert Network for Multi-Behavior RecommendationabstractTraditional recommendation methods, which typically focus on modeling a single user behavior (e.g., purchase), often face severe data sparsity issues. Multi-behavior recommendation methods offer a promising solution by leveraging user data from diverse behaviors. However, most existing approaches entangle multiple behavioral factors, learning holistic but imprecise representations that fail to capture specific user intents. To address this issue, we propose a multi-behavior method by modeling latent factors with an expert network (MBLFE). In our approach, we design a gating expert network, where the expert network models all latent factors within the entire recommendation scenario, with each expert specializing in a specific latent factor. The gating network dynamically selects the optimal combination of experts for each user, enabling a more accurate representation of user preferences. To ensure independence among experts and factor consistency of a particular expert, we incorporate self-supervised learning during the training process. Furthermore, we enrich embeddings with multi-behavior data to provide the expert network with more comprehensive collaborative information for factor extraction. Extensive experiments on three real-world datasets demonstrate that our method significantly outperforms state-of-the-art baselines, validating its effectiveness. Mingshi Yan, Zhiyong Cheng 0001, Yahong Han, Meng Wang 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | WiViPose: A Video-Aided Wi-Fi Framework for Environment-Independent 3D Human Pose EstimationabstractThe inherent complexity of Wi-Fi signals makes video-aided Wi-Fi 3D pose estimation difficult. The challenges include the limited generalizability of the task across diverse environments, its significant signal heterogeneity, and its inadequate ability to analyze local and geometric information. To overcome these challenges, we introduce WiViPose, a video-aided Wi-Fi framework for 3D pose estimation, which attains enhanced cross-environment generalization through cross-layer optimization. Bilinear temporal-spectral fusion (BTSF) is initially used to fuse the time-domain and frequency-domain features derived from Wi-Fi. Video features are derived from a multiresolution convolutional pose machine and enhanced by local self-attention. Cross-modality data fusion is facilitated through an attention-based transformer, with the process further refined under a supervisory mechanism. WiViPose demonstrates effectiveness by achieving an average percentage of correct keypoints (PCK)@50 of 91.01% across three typical indoor environments. Lei Zhang 0024, Haoran Ning, Jiaxin Tang, Yaping Zhong, Yahong Han |
IEEE Trans. Multim. | 6 |
| 2025 | A Static-Dynamic Composition Framework for Efficient Action RecognitionabstractThe dynamic inference, which adaptively allocates computational budgets for different samples, is a prevalent approach for achieving efficient action recognition. Current studies primarily focus on a data-efficient regime that reduces spatial or temporal redundancy, or their combination, by selecting partial video data, such as clips, frames, or patches. However, these approaches often utilize fixed and computationally expensive networks. From a different perspective, this article introduces a novel model-efficient regime that addresses network redundancy by dynamically selecting a partial network in real time. Specifically, we acknowledge that different channels of the neural network inherently contain redundant semantics either spatially or temporally. Therefore, by decreasing the width of the network, we can enhance efficiency while compromising the feature capacity. To strike a balance between efficiency and capacity, we propose the static-dynamic composition (SDCOM) framework, which comprises a static network with a fixed width and a dynamic network with a flexible width. In this framework, the static network extracts the primary feature with essential semantics from the input frame and simultaneously evaluates the gap toward achieving a comprehensive feature representation. Based on these evaluation results, the dynamic network activates a minimal width to extract a supplementary feature that fills the identified gap. We optimize the dynamic feature extraction through the employment of the slimmable network mechanism and a novel meta-learning scheme introduced in this article. Empirical analysis reveals that by combining the primary feature with an extremely lightweight supplementary feature, we can accurately recognize a large majority of frames (76%~92%). As a result, our proposed SDCOM significantly enhances recognition efficiency. For instance, on ActivityNet, FCVID, and Mini-Kinetics datasets, SDCOM saves 90% of the baseline's floating point operations (FLOPs) while achieving comparable or superior accuracy when compared with state-of-the-art methods. Xu Chen 0053, Yahong Han, Xiaojun Chang, Yifan Sun 0003, Yi Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Semantic Prompt Enhancement for Semi-Supervised Low-Light Salient Object DetectionabstractMost existing salient object detection (SOD) models are designed based on data collected in well-lit scenes, which is entirely inadequate for low-light conditions. Although recent models are designed for low-light conditions, they still have limitations. First, they simply integrate features without considering the impact of low-light scenes and fail to enhance the contextual information around salient objects. Second, in extremely dark scenes, it is difficult for the human eye to distinguish between the foreground and background, posing significant challenges for data labeling. To address these issues, we design a brightness Retinex enhancer (BRE) tailored for low-light SOD tasks and, for the first time, explore performing low-light SOD within a semi-supervised framework. By using sparse labeled semantic prompts to augment a large amount of unlabeled data, we mitigate the annotation burden while avoiding ineffective labeling in low-light conditions. More specifically, we first use Retinex decomposition to filter out the influence of illumination, while the semantic features extracted by a large model serve as semantic prompts to assist in enhancement. In addition, we introduce a context-guided encoder (CGE) to improve the model's understanding of salient objects. Finally, both labeled and unlabeled data undergo joint consistency training between the shared decoder (SD) and the perturbation decoder. The semi-supervised model enhances low-light SOD performance while also alleviating the burden of data annotation. Extensive experiments demonstrate that, compared with state-of-the-art fully supervised SOD models, the proposed semi-supervised model achieves highly competitive results across multiple test datasets. Nana Yu, Jie Wang 0095, Yahong Han, Weiping Ding 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | User Invariant Preference Learning for Multi-Behavior RecommendationabstractIn multi-behavior recommendation scenarios, analyzing users’ diverse behaviors, such as click , purchase , and rating , enables a more comprehensive understanding of their interests, facilitating personalized and accurate recommendations. A fundamental assumption of multi-behavior recommendation methods is the existence of shared user preferences across behaviors, representing users’ intrinsic interests. Based on this assumption, existing approaches aim to integrate information from various behaviors to enrich user representations. However, they often overlook the presence of both commonalities and individualities in users’ multi-behavior preferences. These individualities reflect distinct aspects of preferences captured by different behaviors, where certain auxiliary behaviors may introduce noise, hindering the prediction of the target behavior. To address this issue, we propose a user invariant preference learning (UIPL) for multi-behavior recommendation, aiming to capture users’ intrinsic interests (referred to as invariant preferences) from multi-behavior interactions to mitigate the introduction of noise. Specifically, UIPL leverages the paradigm of invariant risk minimization to learn invariant preferences. To implement this, we employ a variational autoencoder (VAE) to extract users’ invariant preferences, replacing the standard reconstruction loss with an invariant risk minimization constraint. Additionally, we construct distinct environments by combining multi-behavior data to enhance robustness in learning these preferences. Finally, the learned invariant preferences are used to provide recommendations for the target behavior. Extensive experiments on four real-world datasets demonstrate that UIPL significantly outperforms current state-of-the-art methods. Mingshi Yan, Zhiyong Cheng 0001, Fan Liu 0008, Yingda Lyu, Yahong Han |
ACM Trans. Inf. Syst. | 5 |
| 2024 | Multi-Source Collaborative Gradient Discrepancy Minimization for Federated Domain GeneralizationabstractFederated Domain Generalization aims to learn a domain-invariant model from multiple decentralized source domains for deployment on unseen target domain. Due to privacy concerns, the data from different source domains are kept isolated, which poses challenges in bridging the domain gap. To address this issue, we propose a Multi-source Collaborative Gradient Discrepancy Minimization (MCGDM) method for federated domain generalization. Specifically, we propose intra-domain gradient matching between the original images and augmented images to avoid overfitting the domain-specific information within isolated domains. Additionally, we propose inter-domain gradient matching with the collaboration of other domains, which can further reduce the domain shift across decentralized domains. Combining intra-domain and inter-domain gradient matching, our method enables the learned model to generalize well on unseen domains. Furthermore, our method can be extended to the federated domain adaptation task by fine-tuning the target model on the pseudo-labeled target domain. The extensive experiments on federated domain generalization and adaptation indicate that our method outperforms the state-of-the-art methods significantly. Yikang Wei, Yahong Han |
AAAI | 2 |
| 2024 | Prompt-Driven Dynamic Object-Centric Learning for Single Domain GeneralizationabstractSingle-domain generalization aims to learn a model from single source domain data attaining generalized performance on other unseen target domains. Existing works primarily focus on improving the generalization ability of static networks. However, static networks are unable to dynamically adapt to the diverse variations in different image scenes, leading to limited generalization capability. Different scenes exhibit varying levels of complexity, and the complexity of images further varies significantly in crossdomain scenarios. In this paper, we propose a dynamic object-centric perception network based on prompt learning, aiming to adapt to the variations in image complexity. Specifically, we propose an object-centric gating module based on prompt learning to focus attention on the object-centric features guided by the various scene prompts. Then, with the object-centric gating masks, the dynamic selective module dynamically selects highly correlated feature regions in both spatial and channel dimensions enabling the model to adaptively perceive object-centric relevant features, thereby enhancing the generalization capability. Extensive experiments were conducted on single-domain generalization tasks in image classification and object detection. The experimental results demonstrate that our approach outperforms state-of-the-art methods, which validates the effectiveness and versatility of our proposed method. Deng Li 0003, Aming Wu, Yaowei Wang 0001, Yahong Han |
CVPR | 4 |
| 2024 | Symmetrical Two-Stream with Selective Sampling for Diversifying Video CaptionsabstractVideo captioning has garnered significant attention due to its promising potential in the field of cross-modal retrieval and analysis. However, the majority of existing methods tend to overlook the inconsistent quality of annotations in video-text pairs. Moreover, prevalent two-stream models tend to generate oversimplified captions, since these methods sacrifice the diversity feature in the original modality by learning the uniform representation of a shared subspace. In this paper, we introduce a weighting network for selective sampling on the training dataset, mitigating the interference of low-quality samples. Additionally, we adopt a symmetrical two-stream method to gain a better understanding of the shared cross-modal information for diversifying video captions. Experimental results on the MSVD and MSR-VTT have demonstrated excellent performance and effectiveness of our proposed method compared to existing methods across various metrics. Yahong Han |
ICME | 2 |
| 2024 | A Patch-wise Adversarial Denoising Could Enhance the Robustness of Adversarial TrainingabstractThe adversarial examples have demonstrated the vulnerability of machine learning models. While the data augmentation strategy has been a cornerstone in circumventing overfitting within the realm of standard training paradigms, the prior research has proved the limited efficiency of data augmentation in ameliorating overfitting concerns, specifically within the scope of adversarial training. In this work, we have shown that a data augmentation strategy could enhance the generalization of adversarial training. Our framework, Adversarial Denoising-based Training(ADT), first utilizes the adversarial denoising strategy to improve the generalization of adversarial training. Specifically, we use patch-wise adversarial denoising as a data augmentation strategy to boost the robustness of adversarial training. We have evaluated our method on the MNIST, CIFAR-10, Tiny-ImageNet, and GTSRB datasets. The result shows that our framework could successfully improve the adversarial robustness of models against both digital and physical adversarial examples. Shibin Liu, Boyuan Zhang 0003, Yang Zhai, Yahong Han |
ICME | 6 |
| 2024 | Improving Transferability of Adversarial Examples with Adversaries CompetitionabstractAdversarial attacks deceive deep neural networks by adding subtle perturbations to benign examples, threatening various applications. However, traditional methods lack effective transferability to unknown black-box networks. Addressing this, we introduce a new method, Patch-based Transfer Attack (PTA), comprising Sensitive Area Localization (SAL) and Robust Perturbations Generation (RPG). SAL identifies highly activated regions, while RPG employs adversarial fine-tuning (AFT) and adversarial feature mixing (AFM) to target robust features, enhancing transferability. AFT adapts the surrogate model to current adversarial examples, and AFM combines adversarial and clean example features, compelling the focus on robust features. Our comprehensive tests demonstrate PTA’s superior transferability, outperforming existing methods, especially when integrated with other techniques. Boyuan Zhang 0003, Yang Zhai, Yahong Han |
ICME | 6 |
| 2024 | Behavior-Contextualized Item Preference Modeling for Multi-Behavior RecommendationabstractIn recommender systems, multi-behavior methods have demonstrated their effectiveness in mitigating issues like data sparsity, a common challenge in traditional single-behavior recommendation approaches. These methods typically infer user preferences from various auxiliary behaviors and apply them to the target behavior for recommendations. However, this direct transfer can introduce noise to the target behavior in recommendation, due to variations in user attention across different behaviors. To address this issue, this paper introduces a novel approach, Behavior-Contextualized Item Preference Modeling (BCIPM), for multi-behavior recommendation. Our proposed Behavior-Contextualized Item Preference Network discerns and learns users' specific item preferences within each behavior. It then considers only those preferences relevant to the target behavior for final recommendations, significantly reducing noise from auxiliary behaviors. These auxiliary behaviors are utilized solely for training the network parameters, thereby refining the learning process without compromising the accuracy of the target behavior recommendations. To further enhance the effectiveness of BCIPM, we adopt a strategy of pre-training the initial embeddings. This step is crucial for enriching the item-aware preferences, particularly in scenarios where data related to the target behavior is sparse. Comprehensive experiments conducted on four real-world datasets demonstrate BCIPM's superior performance compared to several leading state-of-the-art models, validating the robustness and efficiency of our proposed approach. Mingshi Yan, Fan Liu 0008, Jing Sun 0012, Fuming Sun, Zhiyong Cheng 0001, Yahong Han |
SIGIR | 6 |
| 2024 | VADS: Visuo-Adaptive DualStrike attack on visual question answer
Boyuan Zhang 0003, Yahong Han, Qinghua Hu |
Comput. Vis. Image Underst. | 4 |
| 2024 | Linking unknown characters via oracle bone inscriptions retrieval
Xu Chen 0053, Bang Li, Yongge Liu, Runhua Jiang, Yahong Han |
Multim. Syst. | 6 |
| 2024 | Degradation-removed multiscale fusion for low-light salient object detection
Nana Yu, Jie Wang 0095, Yahong Han |
Pattern Recognit. | 5 |
| 2024 | Pseudo-label refinement via hierarchical contrastive learning for source-free unsupervised domain adaptation
Deng Li 0003, Jianguang Zhang, Kunhong Wu, Yahong Han |
Pattern Recognit. Lett. | 5 |
| 2024 | Weakly-Supervised Video Anomaly Detection With Snippet Anomalous AttentionabstractWith a focus on abnormal events contained within untrimmed videos, there is increasing interest among researchers in video anomaly detection. Among different video anomaly detection scenarios, weakly-supervised video anomaly detection poses a significant challenge as it lacks frame-wise labels during the training stage, only relying on video-level labels as coarse supervision. Previous methods have made attempts to either learn discriminative features in an end-to-end manner or employ a two-stage self-training strategy to generate snippet-level pseudo labels. However, both approaches have certain limitations. The former tends to overlook informative features at the snippet level, while the latter can be susceptible to noises. In this paper, we propose an Anomalous Attention mechanism for weakly-supervised anomaly detection to tackle the aforementioned problems. Our approach takes into account snippet-level encoded features without the supervision of pseudo labels. Specifically, our approach first generates snippet-level anomalous attention and then feeds it together with original anomaly scores into a Multi-branch Supervision Module. The module learns different areas of the video, including areas that are challenging to detect, and also assists the attention optimization. Experiments on benchmark datasets XD-Violence and UCF-Crime verify the effectiveness of our method. Besides, thanks to the proposed snippet-level attention, we obtain a more precise anomaly localization. Yidan Fan, Yongxin Yu, Wenhuan Lu, Yahong Han |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Generalizing to Out-of-Sample Degradations via Model ReprogrammingabstractExisting image restoration models are typically designed for specific tasks and struggle to generalize to out-of-sample degradations not encountered during training. While zero-shot methods can address this limitation by fine-tuning model parameters on testing samples, their effectiveness relies on predefined natural priors and physical models of specific degradations. Nevertheless, determining out-of-sample degradations faced in real-world scenarios is always impractical. As a result, it is more desirable to train restoration models with inherent generalization ability. To this end, this work introduces the Out-of-Sample Restoration (OSR) task, which aims to develop restoration models capable of handling out-of-sample degradations. An intuitive solution involves pre-translating out-of-sample degradations to known degradations of restoration models. However, directly translating them in the image space could lead to complex image translation issues. To address this issue, we propose a model reprogramming framework, which translates out-of-sample degradations by quantum mechanic and wave functions. Specifically, input images are decoupled as wave functions of amplitude and phase terms. The translation of out-of-sample degradation is performed by adapting the phase term. Meanwhile, the image content is maintained and enhanced in the amplitude term. By taking these two terms as inputs, restoration models are able to handle out-of-sample degradations without fine-tuning. Through extensive experiments across multiple evaluation cases, we demonstrate the effectiveness and flexibility of our proposed framework. Our codes are available at https://github.com/ddghjikle/Out-of-sample-restoration. Runhua Jiang, Yahong Han |
IEEE Trans. Image Process. | 2 |
| 2024 | Joint Correcting and Refinement for Balanced Low-Light Image EnhancementabstractLow-light image enhancement tasks demand an appropriate balance among brightness, color, and illumination. While existing methods often focus on one aspect of the image without considering how to pay attention to this balance, which will cause problems of color distortion and overexposure etc. This seriously affects both human visual perception and the performance of high-level visual models. In this work, a novel synergistic structure is proposed which can balance brightness, color, and illumination more effectively. Specifically, the proposed method, so-called Joint Correcting and Refinement Network (JCRNet), which mainly consists of three stages to balance brightness, color, and illumination of enhancement. Stage 1: we utilize a basic encoder-decoder and local supervision mechanism to extract local information and more comprehensive details for enhancement. Stage 2: cross-stage feature transmission and spatial feature transformation further facilitate color correction and feature refinement. Stage 3: we employ a dynamic illumination adjustment approach to embed residuals between predicted and ground truth images into the model, adaptively adjusting illumination balance. Extensive experiments demonstrate that the proposed method exhibits comprehensive performance advantages over 21 state-of-the-art methods on 9 benchmark datasets. Furthermore, a more persuasive experiment has been conducted to validate our approach the effectiveness in downstream visual tasks (e.g., saliency detection). Compared to several enhancement models, the proposed method effectively improves the segmentation results and quantitative metrics of saliency detection. Nana Yu, Yahong Han |
IEEE Trans. Multim. | 3 |
| 2024 | Camouflaged Object Segmentation Based on Matching-Recognition-Refinement NetworkabstractIn the biosphere, camouflaged objects take the advantage of visional wholeness by keeping the color and texture of the objects highly consistent with the background, thereby confusing the visual mechanism of other creatures and achieving a concealed effect. This is also the main reason why the task of camouflaged object detection is challenging. In this article, we break the visual wholeness and see through the camouflage from the perspective of matching the appropriate field of view. We propose a matching-recognition-refinement network (MRR-Net), which consists of two key modules, i.e., the visual field matching and recognition module (VFMRM) and the stepwise refinement module (SWRM). In the VFMRM, various feature receptive fields are used to match candidate areas of camouflaged objects of different sizes and shapes and adaptively activate and recognize the approximate area of the real camouflaged object. The SWRM then uses the features extracted by the backbone to gradually refine the camouflaged region obtained by VFMRM, thus yielding the complete camouflaged object. In addition, a more efficient deep supervision method is exploited, making the features from the backbone input into the SWRM more critical and not redundant. Extensive experimental results demonstrate that our MRR-Net runs in real-time (82.6 frames/s) and significantly outperforms 30 state-of-the-art models on three challenging datasets under three standard metrics. Furthermore, MRR-Net is applied to four downstream tasks of camouflaged object segmentation (COS), and the results validate its practical application value. Our code is publicly available at: https://github.com/XinyuYanTJU/MRR-Net. Xinyu Yan 0001, Meijun Sun, Yahong Han, Zheng Wang 0008 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | ATRA: Efficient adversarial training with high-robust area
Shibin Liu, Yahong Han |
Vis. Comput. | 2 |
| 2023 | Reliable and Interpretable Personalized Federated LearningabstractFederated learning can coordinate multiple users to participate in data training while ensuring data privacy. The collaboration of multiple agents allows for a natural connection between federated learning and collective intelligence. When there are large differences in data distribution among clients, it is crucial for federated learning to design a reliable client selection strategy and an interpretable client communication framework to better utilize group knowledge. Herein, a reliable personalized federated learning approach, termed RIPFL, is proposed and fully interpreted from the perspective of social learning. RIPFL reliably selects and divides the clients involved in training such that each client can use different amounts of social information and more effectively communicate with other clients. Simultaneously, the method effectively integrates personal information with the social information generated by the global model from the perspective of Bayesian decision rules and evidence theory, enabling individuals to grow better with the help of collective wisdom. An interpretable federated learning mind is well scalable, and the experimental results indicate that the proposed method has superior robustness and accuracy than other state-of-the-art federated learning algorithms. Zixuan Qin, Liu Yang 0010, Qilong Wang 0001, Yahong Han, Qinghua Hu |
CVPR | 4 |
| 2023 | Exploring Instance Relation for Decentralized Multi-Source Domain AdaptationabstractMulti-source domain adaptation aims to transfer knowledge from multiple labeled source domains to an unlabeled target domain and reduce the domain shift. Considering data privacy and storage cost, the data from different domains are isolated, which leads to the difficulty of domain adaptation. To reduce the domain shift on the decentralized source domains and target domain, we propose an instance relation consistency method for decentralized multi-source domain adaptation. Specifically, we utilize the models from other domains as bridges to conduct domain adaptation. We impose inter-domain instance relation consistency on the isolated source and target domain to transfer the semantic relation knowledge across different domain models. Meanwhile, we exploit intra-domain instance relation consistency to learn the intrinsic semantic relation across different data views. Experiments on three benchmarks indicate the effectiveness of our method for decentralized multi-source domain adaptation. Yikang Wei, Yahong Han |
ICASSP | 2 |
| 2023 | Discriminative and Contrastive Consistency for Semi-supervised Domain Adaptive Image ClassificationabstractWith sufficient source and limited target supervised information, semi-supervised domain adaptation (SSDA) aims to perform well on unlabeled target domain. Although various strategies have been proposed in SSDA field, they fail to fully exploit limited target labels and adequately explore domain-invariant knowledge. In this study, we propose a framework that first introduces consistent processing of augmented training data based on contrastive learning. Specifically, supervised contrastive learning is introduced to assist the classical cross-entropy iteration to make full use of the limited target labels. Additionally, traditional unsupervised contrastive learning and pseudo-labeling are utilized to further minimize the intra-domain discrepancy. Besides, an adversarial loss is then combined with a sharpening function to acquire a more certain category center that is domain-invariant. Experimental results on DomainNet, Office-Home, and Office show the effectiveness of our method. Particularly, for 1-shot case of Office-Home with AlexNet as backbone, our method outperforms the previous state-of-the-art by 5.6% in terms of mean accuracy. Yidan Fan, Wenhuan Lu, Yahong Han |
ICME | 3 |
| 2023 | Uncertainty-Aware Variate Decomposition for Self-supervised Blind Image DeblurringabstractBlind image deblurring remains challenging due to the ill-posed nature of the traditional blurring function. Although previous supervised methods have achieved great breakthrough with synthetic blurry-sharp image pairs, their generalization ability to real-world blurs is limited by the discrepancy between synthetic and real blurs. To overcome this limitation, unsupervised deblurring methods have been proposed by using natural priors or generative adversarial networks. However, natural priors are vulnerable to random blur artifacts, while generators of generative adversarial networks always produce inaccurate details and unrealistic colors. Consequently, previous methods easily suffer from slow convergence and poor performance. In this work, we propose to formulate the traditional blurring function as the composition of multiple variates, thus allowing us explicitly define characteristics of residual images between blurry and sharp images. We also propose a multi-step self-supervised deblurring framework to address the slow convergence issue. Our framework continuously decomposes and composes input images, thus utilizing the uncertainty of blur artifacts to obtain diverse pseudo blurry-sharp image pairs for self-supervised learning. This framework is more efficient than previous methods, as it does not rely on natural priors or GANs. Extensive comparisons demonstrate that the proposed framework outperforms state-of-the-art unsupervised methods on both dynamic scene, human-aware centric motion, real-world and out-of-focus deblurring datasets. The codes are available at https://github.com/ddghjikle/MM-2023-USDF. Runhua Jiang, Yahong Han |
ACM Multimedia | 2 |
| 2023 | OraclePoints: A Hybrid Neural Representation for Oracle CharacterabstractOracle Bone Inscriptions (OBI) are ancient hieroglyphs originated in China and are considered one of the most famous writing systems in the world. Up to now, thousands of OBIs have been discovered, which require deciphering by experts to understand their contents. Experts typically need to restore, classify, and compare each character with previous inscriptions. Although existing research can assist with one of these operations, their performance falls short of practical requirements. In this work, we propose the OraclePoints framework, which represents OBI images as hybrid neural representations comprising features of images and point sets. The image representation provides inscription appearance and character structure, while the point representation makes it easy and effective to distinguish characters and noises. In addition, we demonstrate that OraclePoints can be easily integrated with existing models in a plug-and-play manner. Comprehensive experiments demonstrate that the proposed hybrid neural representation framework supports a range of OBI tasks, including character image retrieval, recognition, and denoising. It is also demonstrated that OraclePoints is helpful for deciphering OBIs by linking ancient characters to modern Chinese characters. Our codes are available at https://ddghjikle.github.io/. Runhua Jiang, Yongge Liu, Boyuan Zhang 0003, Xu Chen 0053, Deng Li 0003, Yahong Han |
ACM Multimedia | 6 |
| 2023 | Saliency Prototype for RGB-D and RGB-T Salient Object DetectionabstractMost of the existing bi-modal (RGB-D or RGB-T) salient object detection methods attempt to integrate multimodality information through various fusion strategies. However, existing methods lack a clear definition of salient regions before feature fusion, which results in poor model robustness. To tackle this problem, we propose a novel prototype, the saliency prototype, which captures common characteristic information among salient objects. A prototype contains inherent characteristics information of multiple salient objects, which can be used for feature enhancement of various salient objects. By utilizing the saliency prototype, we provide a clearer definition of salient regions and enable the model to focus on these regions before feature fusion, avoiding the influence of complex backgrounds during the feature fusion stage. Additionally, we utilize the saliency prototypes to address the quality issue of auxiliary modality. Firstly, we apply the saliency prototypes obtained by the primary modality to perform semantic enhancement of the auxiliary modality. Secondly, we dynamically allocate weights for the auxiliary modality during the feature fusion stage in proportion to its quality. Thus, we develop a new bi-modal salient detection architecture Saliency Prototype Network (SPNet), which can be used for both RGB-D and RGB-T SOD. Extensive experimental results on RGB-D and RGB-T SOD datasets demonstrate the effectiveness of the proposed approach against the state-of-the-art. Our code is available at https://github.com/ZZ2490/SPNet. Jie Wang 0095, Yahong Han |
ACM Multimedia | 3 |
| 2023 | A Cross-modal and Redundancy-reduced Network for Weakly-Supervised Audio-Visual Violence DetectionabstractMultimodal learning using audio and visual information has improved Violence Detection tasks. However, previous studies overlook the gap between pre-trained networks and the final violence detection task, as well as the semantic inconsistency between audio and visual features. We consider task-irrelevant information caused by the former situation and semantic noise due to the latter as redundancy, negatively affecting overall detection performance. Besides, the prevailing visual modality-centric approach with audio features as guidance may be biased. We contend that both modalities are crucial in violence detection. To address these issues, we propose a Cross-modal and Redundancy-reduced Network for Weakly-Supervised Audio-Visual Violence Detection. Our framework integrates a relation-ware module with a bi-directional cross-modal attention mechanism to explore interactions between modalities. Then, we introduce a feature filter gate to reduce redundancy. Finally, a multi-branch classification module is proposed for better utilization of both modalities. Extensive experiments demonstrate the effectiveness of our approach, surpassing previous methods with state-of-the-art performance in violence detection. Yidan Fan, Yongxin Yu, Wenhuan Lu, Yahong Han |
MMAsia | 4 |
| 2023 | Domain-specific feature elimination: multi-source domain adaptation for image classification
Kunhong Wu, Fan Jia 0006, Yahong Han |
Frontiers Comput. Sci. | 3 |
| 2023 | Dynamic parameterized learning for unsupervised domain adaptationabstractUnsupervised domain adaptation enables neural networks to transfer from a labeled source domain to an unlabeled target domain by learning domain-invariant representations. Recent approaches achieve this by directly matching the marginal distributions of these two domains. Most of them, however, ignore exploration of the dynamic trade-off between domain alignment and semantic discrimination learning, thus rendering them susceptible to the problems of negative transfer and outlier samples. To address these issues, we introduce the dynamic parameterized learning framework. First, by exploring domain-level semantic knowledge, the dynamic alignment parameter is proposed, to adaptively adjust the optimization steps of domain alignment and semantic discrimination learning. Besides, for obtaining semantic-discriminative and domain-invariant representations, we propose to align training trajectories on both source and target domains. Comprehensive experiments are conducted to validate the effectiveness of the proposed methods, and extensive comparisons are conducted on seven datasets of three visual tasks to demonstrate their practicability. Runhua Jiang, Yahong Han |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2023 | Improving transferable adversarial attack for vision transformers via global attention and local drop
Yahong Han |
Multim. Syst. | 2 |
| 2023 | Weakly supervised anomaly detection with multi-level contextual modeling
Yongge Liu, Yahong Han |
Multim. Syst. | 4 |
| 2023 | Weighted progressive alignment for multi-source domain adaptation
Kunhong Wu, Yahong Han |
Multim. Syst. | 3 |
| 2023 | Query-Efficient Black-Box Adversarial Attack With Customized Iteration and SamplingabstractIt is a challenging task to fool an image classifier based on deep neural networks under the black-box setting where the target model can only be queried. Among existing black-box attacks, transfer-based methods tend to overfit the substitute model on parameter settings. Decision-based methods have low query efficiency due to fixed sampling and greedy search strategy. To alleviate the above problems, we present a new framework for query-efficient black-box adversarial attack by bridging transfer-based and decision-based attacks. We reveal the relationship between current noise and variance of sampling, the monotonicity of noise compression, and the influence of transition function on the decision-based attack. Guided by the new framework, we propose a black-box adversarial attack named Customized Iteration and Sampling Attack (CISA). CISA estimates the distance from nearby decision boundary to set the stepsize, and uses a dual-direction iterative trajectory to find the intermediate adversarial example. Based on the intermediate adversarial example, CISA conducts customized sampling according to the noise sensitivity of each pixel to further compress noise, and relaxes the state transition function to achieve higher query efficiency. Extensive experiments demonstrate CISA's advantage in query efficiency of black-box adversarial attacks. Yahong Han, Qinghua Hu, Yi Yang 0001, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Source-free and black-box domain adaptation via distributionally adversarial training
Kunhong Wu, Yahong Han, Yunfeng Shao 0001, Bingshuai Li, Fei Wu 0001 |
Pattern Recognit. | 3 |
| 2023 | Multi-Source Collaborative Contrastive Learning for Decentralized Domain AdaptationabstractUnsupervised multi-source domain adaptation aims to obtain a model working well on the unlabeled target domain by reducing the domain gap between the labeled source domains and the unlabeled target domain. Considering the data privacy and storage cost, data from multiple source domains and target domain are isolated and decentralized. This data decentralization scenario brings the difficulty of domain alignment for reducing the domain gap between the decentralized source domains and target domain, respectively. For conducting domain alignment under the data decentralization scenario, we propose Multi-source Collaborative Contrastive learning for decentralized Domain Adaptation (MCC-DA). The models from other domains are used as the bridge to reduce the domain gap. On the source domains and target domain, we penalize the inconsistency of data features extracted from the source domain models and target domain model by contrastive alignment. With the collaboration of source domain models and target domain model, the domain gap between decentralized source domains and target domain is reduced without accessing the data from other domains. The experiment results on multiple benchmarks indicate that our method can reduce the domain gap effectively and outperform the state-of-the-art methods significantly. Yikang Wei, Liu Yang 0010, Yahong Han, Qinghua Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Active and Compact Entropy Search for High-Dimensional Bayesian OptimizationabstractEntropy search and its derivative methods are one class of Bayesian Optimization methods that achieve active exploration of black-box functions. They maximize the information gain about the position in the input space where the black-box function gets the global optimum. However, existing entropy search methods suffer from harassment caused by high dimensional optimization problems. On the one hand, the computation for estimating entropies increases exponentially as dimensions increase, which limits the applicability of entropy search to high dimensional problems. On the other hand, many high-dimensional problems have the property that a large number of dimensions have little influence on the objective function, but currently there is no compress mechanism to exclude these redundant dimensions. In this work, we propose Active Compact Entropy Search (AcCES) to fix these defects. Under the guidance of historical evaluation, AcCES actively explores the prevalent inter-dimensional correlations by maximizing the linear or non-linear relationships that may exist between dimensions in the acquisition function, which is ignored by existing Bayesian Optimization methods. In order to build a more compact input space, redundant dimensions are compressed by exploiting inter-dimensional correlations. Experiments demonstrate that AcCES achieves higher query efficiency and optimal results than existing entropy search methods. Run Li, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Logic Rule Guided Attribution with Dynamic AblationabstractWith the increasing demands for understanding the internal behaviors of deep networks, Explainable AI (XAI) has been made remarkable progress in interpreting the model's decision. A family of attribution techniques has been proposed, highlighting whether the input pixels are responsible for the model's prediction. However, the existing attribution methods suffer from the lack of rule guidance and require further human interpretations. In this paper, we construct the 'if-then' logic rules that are sufficiently precise locally. Moreover, a novel rule-guided method, dynamic ablation (DA), is proposed to find a minimal bound sufficient in an input image to justify the network's prediction and aggregate iteratively to reach a complete attribution. Both qualitative and quantitative experiments are conducted to evaluate the proposed DA. We demonstrate the advantages of our method in providing clear and explicit explanations that are also easy for human experts to understand. Besides, through the attribution on a series of trained networks with different architectures, we show that more complex networks require less information to make a specific prediction. Jianqiao An, Yuandu Lai, Yahong Han |
AAAI | 3 |
| 2022 | Mining Valuable Source Domain Instances for Privacy-Preserving Domain Adaptive Object DetectionabstractConventional Domain Adaptive Object Detection (DAOD) usually learns a well-performed detector by aligning the source/target representations in hidden space, which is un-practical in real situations because of data privacy. Recently, some methods based on pseudo-label technique have been proposed for privacy-preserving DAOD. However, they bring new difficulties in evaluating the quality of these labels. In this paper, we propose a Mining Valuable source domain instances (MiVi) algorithm for privacy-preserving DAOD where data from different domains cannot be aggregated. Specifically, we explore a novel Instances Refinement strategy to mine valuable instance-level samples from the source data, which are close to the feature distribution of the target domain. To achieve this, we are the first to introduce self-supervised learning to guide the detector to learn the instance-level features in different domains. Then, by reserving labeled source domain instance-level samples with domain-invariant features and discarding the ones with source domain-specific features, we transform unsupervised DAOD into supervised object detection. Extensive experiments in four representative DAOD scenarios show that our method outperforms existing methods. Jianqiao An, Yahong Han |
ICME | 3 |
| 2022 | Multi-Granularity Semantic Clues Extraction for Video Question AnsweringabstractConventional methods for video question answering focus on encoding video and questions in sequence, and carefully de-signing multi-modal interactions to fuse multi-modal infor-mation. Although achieving promising results, these meth-ods mainly focus on the sequential structure of video and fail to explore the underlying hierarchical semantic structure of the video. In this work, we argue that video content is se-quential in time space is organized in a hierarchical structure in semantic space (e.g., object-action-scene). Corresponding to the complex video structure, the questions also involve multi-granularity queries in the video. To cope with queries of different granularities in videos, we propose an Object-to-Scene relational reasoning (O2SR) framework, which en-codes videos in multi-granularities in the semantic space under the guidance of questions. By modeling the local object relations and the global scene dependencies while encapsu-lating corresponding questions semantic into visual elements, the O2SR achieves better generalization on different types of questions. Experimental evaluations show our method achieves state-of-the-art performance on four datasets. Code will be released at: https://github.com/zophe98/O2SR. Yahong Han |
ICME | 2 |
| 2022 | Decision-based Black-box Attack Against Vision Transformers via Patch-wise Adversarial RemovalabstractVision transformers (ViTs) have demonstrated impressive performance and stronger adversarial robustness compared to Convolutional Neural Networks (CNNs). On the one hand, ViTs' focus on global interaction between individual patches reduces the local noise sensitivity of images. On the other hand, the neglect of noise sensitivity differences between image regions by existing decision-based attacks further compromises the efficiency of noise compression, especially for ViTs. Therefore, validating the black-box adversarial robustness of ViTs when the target model can only be queried still remains a challenging problem. In this paper, we theoretically analyze the limitations of existing decision-based attacks from the perspective of noise sensitivity difference between regions of the image, and propose a new decision-based black-box attack against ViTs, termed Patch-wise Adversarial Removal (PAR). PAR divides images into patches through a coarse-to-fine search process and compresses the noise on each patch separately. PAR records the noise magnitude and noise sensitivity of each patch and selects the patch with the highest query value for noise compression. In addition, PAR can be used as a noise initialization method for other decision-based attacks to improve the noise compression efficiency on both ViTs and CNNs without introducing additional calculations. Extensive experiments on three datasets demonstrate that PAR achieves a much lower noise magnitude with the same number of queries. Yahong Han, Yu-an Tan 0001, Xiaohui Kuang |
NeurIPS | 2 |
| 2022 | Unidirectional RGB-T salient object detection with intertwined driving of encoding and fusion
Jie Wang 0095, Kechen Song, Yanqi Bao, Yunhui Yan, Yahong Han |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Instance-sequence reasoning for video question answering
Yahong Han |
Frontiers Comput. Sci. | 2 |
| 2022 | Exploring uncertainty in regression neural networks for construction of prediction intervals
Yuandu Lai, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
Neurocomputing | 3 |
| 2022 | Dual collaboration for decentralized multi-source domain adaptationabstractThe goal of decentralized multi-source domain adaptation is to conduct unsupervised multi-source domain adaptation in a data decentralization scenario. The challenge of data decentralization is that the source domains and target domain lack cross-domain collaboration during training. On the unlabeled target domain, the target model needs to transfer supervision knowledge with the collaboration of source models, while the domain gap will lead to limited adaptation performance from source models. On the labeled source domain, the source model tends to overfit its domain data in the data decentralization scenario, which leads to the negative transfer problem. For these challenges, we propose dual collaboration for decentralized multi-source domain adaptation by training and aggregating the local source models and local target model in collaboration with each other. On the target domain, we train the local target model by distilling supervision knowledge and fully using the unlabeled target domain data to alleviate the domain shift problem with the collaboration of local source models. On the source domain, we regularize the local source models in collaboration with the local target model to overcome the negative transfer problem. This forms a dual collaboration between the decentralized source domains and target domain, which improves the domain adaptation performance under the data decentralization scenario. Extensive experiments indicate that our method outperforms the state-of-the-art methods by a large margin on standard multi-source domain adaptation datasets. Yikang Wei, Yahong Han |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2022 | Complementary spatiotemporal network for video question answering
Aming Wu, Yahong Han |
Multim. Syst. | 3 |
| 2022 | Multi-attribute object detection benchmark for smart city
Yaowei Wang 0001, Zhouxin Yang, Deng Li 0003, Yuandu Lai, Lihan Ouyang, Leyuan Fang, Yahong Han |
Multim. Syst. | 8 |
| 2022 | Instance-Invariant Domain Adaptive Object Detection Via Progressive DisentanglementabstractMost state-of-the-art methods of object detection suffer from poor generalization ability when the training and test data are from different domains. To address this problem, previous methods mainly explore to align distribution between source and target domains, which may neglect the impact of the domain-specific information existing in the aligned features. Besides, when transferring detection ability across different domains, it is important to extract the instance-level features that are domain-invariant. To this end, we explore to extract instance-invariant features by disentangling the domain-invariant features from the domain-specific features. Particularly, a progressive disentangled mechanism is proposed to decompose domain-invariant and domain-specific features, which consists of a base disentangled layer and a progressive disentangled layer. Then, with the help of Region Proposal Network (RPN), the instance-invariant features are extracted based on the output of the progressive disentangled layer. Finally, to enhance the disentangled ability, we design a detached optimization to train our model in an end-to-end fashion. Experimental results on four domain-shift scenes show our method is separately 2.3, 3.6, 4.0, and 2.0 percent higher than the baseline method. Meanwhile, visualization analysis demonstrates that our model owns well disentangled ability. Aming Wu, Yahong Han, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Effective full-scale detection for salient object based on condensing-and-filtering network
Xinyu Yan 0001, Meijun Sun, Yahong Han, Zheng Wang 0008, Qi Tian 0001 |
Pattern Recognit. | 3 |
| 2022 | Action Keypoint Network for Efficient Video RecognitionabstractReducing redundancy is crucial for improving the efficiency of video recognition models. An effective approach is to select informative content from the holistic video, yielding a popular family of dynamic video recognition methods. However, existing dynamic methods focus on either temporal or spatial selection independently while neglecting a reality that the redundancies are usually spatial and temporal, simultaneously. Moreover, their selected content is usually cropped with fixed shapes (e.g., temporally-cropped frames, spatially-cropped patches), while the realistic distribution of informative content can be much more diverse. With these two insights, this paper proposes to integrate temporal and spatial selection into an Action Keypoint Network (AK-Net). From different frames and positions, AK-Net selects some informative points scattered in arbitrary-shaped regions as a set of "action keypoints" and then transforms the video recognition into point cloud classification. More concretely, AK-Net has two steps, i.e., the keypoint selection and the point cloud classification. First, it inputs the video into a baseline network and outputs a feature map from an intermediate layer. We view each pixel on this feature map as a spatial-temporal point and select some informative keypoints using self-attention. Second, AK-Net devises a ranking criterion to arrange the keypoints into an ordered 1D sequence. Since the video is represented with a 1D sequence after the specified layer, AK-Net transforms the subsequent layers into a point cloud classification sub-net by compacting the original 2D convolutional kernels into 1D kernels. Consequentially, AK-Net brings two-fold benefits for efficiency: The keypoint selection step collects informative content within arbitrary shapes and increases the efficiency for modeling spatial-temporal dependencies, while the point cloud classification step further reduces the computational cost by compacting the convolutional kernels. Experimental results show that AK-Net can consistently improve the efficiency and performance of baseline methods on several video recognition benchmarks. Xu Chen 0053, Yahong Han, Yifan Sun 0003, Yi Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Curiosity-Driven Salient Object Detection With Fragment AttentionabstractRecent deep learning based salient object detection methods with attention mechanisms have made great success. However, existing attention mechanisms can be generally separated into two categories. One part chooses to calculate weights indiscriminately, which yields computational redundancy. While one part focuses randomly on a small part of the images, such as hard attention, resulting in incorrectness owing to insufficiently targeted selection of a subset of tokens. To alleviate these problems, we design a Curiosity-driven Network (CNet) and a Curiosity-driven Learning Algorithm (CLA) based on fragment attention (FA) mechanism newly defined in this paper. FA imitates the process of cognition perception driven by human curiosity, and divides the degree of curiosity into three levels, i.e. curious, a little curious and not curious. These three levels correspond to five saliency degrees, including salient and non-salient, likewise salient and likewise non-salient, completely uncertain. With more knowledge gained by the network, CLA transforms the curiosity degree of each pixel to yield enhanced detail-enriched saliency maps. In order to extract more context-aware information of potential salient objects and make a better foundation for CLA, a high-level feature extraction module (HFEM) is further proposed. Based on the much better high-level features extracted by HFEM, FA can classify the curiosity degree for each pixel more reasonably and accurately. Extensive experiments on five popular datasets clearly demonstrate that our method outperforms the state-of-the-art approaches without any pre-processing operations or post-processing operations. Zheng Wang 0008, Pengzhi Wang, Yahong Han, Meijun Sun, Qi Tian 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Universal-Prototype Enhancing for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) aims to strengthen the performance of novel object detection with few labeled samples. To alleviate the constraint of few samples, enhancing the generalization ability of learned features for novel objects plays a key role. Thus, the feature learning process of FSOD should focus more on intrinsical object characteristics, which are invariant under different visual changes and therefore are helpful for feature generalization. Unlike previous attempts of the meta-learning paradigm, in this paper, we explore how to enhance object features with intrinsical characteristics that are universal across different object categories. We propose a new prototype, namely universal prototype, that is learned from all object categories. Besides the advantage of characterizing invariant characteristics, the universal prototypes alleviate the impact of unbalanced object categories. After enhancing object features with the universal prototypes, we impose a consistency loss to maximize the agreement between the enhanced features and the original ones, which is beneficial for learning invariant object characteristics. Thus, we develop a new framework of few-shot object detection with universal prototypes (F SODup) that owns the merit of feature generalization towards novel objects. Experimental results on PASCAL VOC and MS COCO show the effectiveness of F SODup. Particularly, for the 1-shot case of VOC Split2, FSODupoutperforms the baseline by 6.8% in terms of mAP. Aming Wu, Yahong Han, Linchao Zhu, Yi Yang 0001 |
ICCV | 2 |
| 2021 | Vector-Decomposed Disentanglement for Domain-Invariant Object DetectionabstractTo improve the generalization of detectors, for domain adaptive object detection (DAOD), recent advances mainly explore aligning feature-level distributions between the source and single-target domain, which may neglect the impact of domain-specific information existing in the aligned features. Towards DAOD, it is important to extract domain-invariant object representations. To this end, in this paper, we try to disentangle domain-invariant representations from domain-specific representations. And we propose a novel disentangled method based on vector decomposition. Firstly, an extractor is devised to separate domain-invariant representations from the input, which are used for extracting object proposals. Secondly, domain-specific representations are introduced as the differences between the input and domain-invariant representations. Through the difference operation, the gap between the domain-specific and domain-invariant representations is enlarged, which promotes domain-invariant representations to contain more domain-irrelevant information. In the experiment, we separately evaluate our method on the single- and compound-target case. For the single-target case, experimental results of four domain-shift scenes show our method obtains a significant performance gain over baseline methods. Moreover, for the compound-target case (i.e., the target is a compound of two different domains without domain labels), our method outperforms baseline methods by around 4%, which demonstrates the effectiveness of our method. Aming Wu, Yahong Han, Linchao Zhu, Yi Yang 0001 |
ICCV | 3 |
| 2021 | Anomaly Detection with Prototype-Guided Discriminative Latent EmbeddingsabstractRecent efforts towards video anomaly detection (VAD) try to learn a deep autoencoder to describe normal event patterns with small reconstruction errors. The video inputs with large reconstruction errors are regarded as anomalies at the test time. However, these methods sometimes reconstruct abnormal inputs well because of the powerful generalization ability of deep autoencoder. To address this problem, we present a novel approach for anomaly detection, which utilizes discriminative prototypes of normal data to reconstruct video frames. In this way, the model will favor the reconstruction of normal events and distort the reconstruction of abnormal events. Specifically, we use a prototype-guided memory module to perform discriminative latent embedding. We introduce a new discriminative criterion for the memory module, as well as a loss function correspondingly, which can encourage memory items to record the representative embeddings of normal data, i.e. prototypes. Besides, we design a novel two-branch autoencoder, which is composed of a future frame prediction network and an RGB difference generation network that share the same encoder. The stacked RGB difference contains motion information just like optical flow, so our model can learn temporal regularity. We evaluate the effectiveness of our method on three benchmark datasets and experimental results demonstrate the proposed method outperforms the state-of-the-art. Yuandu Lai, Yahong Han, Yaowei Wang 0001 |
ICDM | 2 |
| 2021 | Adversarial Attack with KD-Tree Searching on Training Set
Xinghua Guo, Fan Jia 0006, Jianqiao An, Yahong Han |
ICIG (2) | 4 |
| 2021 | Free Adversarial Training with Layerwise Heuristic Learning
Benyu Dong, Yahong Han, Yuanzhang Li 0001, Xiaohui Kuang |
ICIG (2) | 4 |
| 2021 | Zero Knowledge Adversarial Defense Via Iterative Translation CycleabstractImage classification networks based on deep learning are found to be vulnerable to the carefully designed adversarial examples. The existing approaches are still far from solving this problem. In this paper, we propose iterative translation cycle GAN (ITC-GAN) to jointly optimize generators and discriminators, so as to defense against adversarial examples with zero knowledge of adversarial attacks. We train two generators to form an iterative translation cycle which changes deep features of the input images and then reconstructs them. As the cycle of image translation can be conducted iteratively, the noises of adversarial examples are gradually eliminated. We only use clean images to train the whole ITC-GAN, so our method is not coupled to specific attack methods and specific classifiers. We conduct experiments on MNIST, CIFAR10, and ImageNet50. The experimental results demonstrate that ITC-GAN is more robust and flexible than state-of-the-art methods in different adversarial settings. Fan Jia 0006, Yahong Han |
ICME | 3 |
| 2021 | Graph-in-Graph Contrastive Learning for Semi-Supervised AdaptationabstractSemi-supervised domain adaptation (SSDA) aims to adapt the model from the labeled source domain to the target domain including few labeled data. Extracting the general features is important to solve SSDA, which is beneficial to promote the model to adapt to the target domain. To this end, in this paper, we propose a novel framework to enhance the generalization of the model which improves the accuracy in the target domain. Particularly, we construct a new graph-in-graph component to model the internal relationship of the input feature, which is helpful for extracting rich and general features. In addition, for large amounts of unlabeled data in the target domain, we use the contrastive loss to optimize the network and extract general representations. We evaluate our framework on three benchmark datasets including Domain-Net, Office-Home, and Office. The extensive experimental results demonstrate the proposed method achieves state-of-the-art performance. Aming Wu, Yahong Han |
ICME | 3 |
| 2021 | Video-to-Image Casting: A Flatting Method for Video AnalysisabstractPrevious mainstream video analysis methods, especially 3D CNNs-based models, mainly aim to transfer frameworks from the image domain to the video domain, and they follow the regime which has been succeeded in image processing, i.e., large-scale benchmarks and deep networks. However, processing videos is still time-consuming due to the increased computational cost. In this paper, we propose to flat the video and construct a Spatio-temporal Image (STI), i.e., squeezing the temporal dimension into a spatial plane. To pursuit the video-level modeling and efficient architecture, we devise a Collective Convolution (CoConv) operation to replace the 2D convolution. With the holistic sampling strategy, this novel operation can extract the video-level spatio-temporal representation. Moreover, we ensure that each CoConv operation has the same number of parameters as the original 2D filter, thus we can utilize a 2D network equipped with CoConv to analyze videos without additional computations. To verify the effectiveness of our method for the general video analysis, we evaluate it on three typical tasks, i.e., supervised action recognition, self-supervised action recognition, and dynamic texture recognition. Extensive experimental results show that our method can achieve comparable or state-of-the-art performances on these benchmarks while using much fewer computations compared with its 3D counterpart. Xu Chen 0053, Chenqiang Gao, Feng Yang 0015, Yi Yang 0001, Yahong Han |
ACM Multimedia | 6 |
| 2021 | WAB'21: 1st Workshop on Multimodal Product Identification in Livestreaming and WAB ChallengeabstractProduct identification has become a very important component in the modern E-commerce shopping system. Consumers could enjoy watching livingstreaming and buying products that livestream hosts recommended. However, with hundreds of products presented in a livingstreaming video, finding the specific product could be laboursome for consumers. Hence, automatic product identification is desired in livingstreaming based E-commerce system. Compared with the image-based visual searching system, the complicated contents in the livestreaming videos make the identification even more challenging. To promote the research on product identification in livestreaming, we present the largest multimodal product retrieval dataset named "Watch and Buy" (WAB) and launch the multimodal product retrieval challenge. We hope this workshop could help researchers further advance the performance and applicability of livestreaming product identification in real-world systems. Yueting Zhuang, Guilin Wu, Yahong Han, Haihong Tang, Baoming Yan, Yi Yang 0001 |
ACM Multimedia | 4 |
| 2021 | Locating Visual Explanations for Video Question Answering
Xuanwei Chen, Xiaomeng Song, Yahong Han |
MMM (1) | 4 |
| 2021 | Visual commonsense reasoning with directional visual connectionsabstractTo boost research into cognition-level visual understanding, i.e., making an accurate inference based on a thorough understanding of visual details, visual commonsense reasoning (VCR) has been proposed. Compared with traditional visual question answering which requires models to select correct answers, VCR requires models to select not only the correct answers, but also the correct rationales. Recent research into human cognition has indicated that brain function or cognition can be considered as a global and dynamic integration of local neuron connectivity, which is helpful in solving specific cognition tasks. Inspired by this idea, we propose a directional connective network to achieve VCR by dynamically reorganizing the visual neuron connectivity that is contextualized using the meaning of questions and answers and leveraging the directional information to enhance the reasoning ability. Specifically, we first develop a GraphVLAD module to capture visual neuron connectivity to fully model visual content correlations. Then, a contextualization process is proposed to fuse sentence representations with visual neuron representations. Finally, based on the output of contextualized connectivity, we propose directional connectivity to infer answers and rationales, which includes a ReasonVLAD module. Experimental results on the VCR dataset and visualization analysis demonstrate the effectiveness of our method. Yahong Han, Aming Wu, Linchao Zhu, Yi Yang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2021 | An Evolutionary-Based Black-Box Attack to Deep Neural Network Classifiers
Yutian Zhou, Yu-an Tan 0001, Quanxin Zhang 0001, Xiaohui Kuang, Yahong Han |
Mob. Networks Appl. | 5 |
| 2021 | Deep multi-scale and multi-modal fusion for 3D object detection
Deng Li 0003, Yahong Han |
Pattern Recognit. Lett. | 3 |
| 2021 | Hierarchical Memory Decoder for Visual NarratingabstractVisual narrating focuses on generating semantic descriptions to summarize visual content of images or videos, e.g., visual captioning and visual storytelling. The challenge mainly lies in how to design a decoder to generate accurate descriptions matching visual content. Recent advances often employ a recurrent neural network (RNN), e.g., Long-Short Term Memory (LSTM), as the decoder. However, RNN is prone to diluting long-term information, which weakens its performance of capturing long-term dependencies. Recent work has demonstrated memory network (MemNet) owns the advantage of storing long-term information. However, as the decoder, it has not been well exploited for visual narrating. The reason partially comes from the difficulty of multi-modal sequential decoding with MemNet. In this article, we devise a novel memory decoder for visual narrating. Concretely, to obtain a better multi-modal representation, we first design a new multi-modal fusion method to fully merge visual and lexical information. Then, based on the fusion result, during decoding, we construct a MemNet-based decoder consisting of multiple memory layers. Particularly, in each layer, we employ a memory set to store previous decoding information and utilize an attention mechanism to adaptively select the information related to the current output. Meanwhile, we also employ a memory set to store the decoding output of each memory layer at the current time step and still utilize an attention mechanism to select the related information. Thus, this decoder alleviates dilution of long-term information. Meanwhile, the hierarchical architecture leverages the latent information of each layer, which is helpful for generating accurate descriptions. Experimental results on two tasks of visual narrating, i.e., video captioning and visual storytelling, show that our decoder could obtain superior results and outperform the performance of conventional RNN-based decoder. Aming Wu, Yahong Han, Zhou Zhao 0001, Yi Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringabstractThe dominant video question answering methods are based on fine-grained representation or model-specific attention mechanism. They usually process video and question separately, then feed the representations of different modalities into following late fusion networks. Although these methods use information of one modality to boost the other, they neglect to integrate correlations of both inter- and intra-modality in an uniform module. We propose a deep heterogeneous graph alignment network over the video shots and question words. Furthermore, we explore the network architecture from four steps: representation, fusion, alignment, and reasoning. Within our network, the inter- and intra-modality information can be aligned and interacted simultaneously over the heterogeneous graph and used for cross-modal reasoning. We evaluate our method on three benchmark datasets and conduct extensive ablation study to the effectiveness of the network architecture. Experiments show the network to be superior in quality. Pin Jiang, Yahong Han |
AAAI | 2 |
| 2020 | Multi-Speaker Video Dialog with Frame-Level Temporal Localization
Pin Jiang, Zhiyi Guo, Yahong Han, Zhou Zhao 0001 |
AAAI | 4 |
| 2020 | Polishing Decision-Based Adversarial Noise With a Customized SamplingabstractAs an effective black-box adversarial attack, decision-based methods polish adversarial noise by querying the target model. Among them, boundary attack is widely applied due to its powerful noise compression capability, especially when combined with transfer-based methods. Boundary attack splits the noise compression into several independent sampling processes, repeating each query with a constant sampling setting. In this paper, we demonstrate the advantage of using current noise and historical queries to customize the variance and mean of sampling in boundary attack to polish adversarial noise. We further reveal the relationship between the initial noise and the compressed noise in boundary attack. We propose Customized Adversarial Boundary (CAB) attack that uses the current noise to model the sensitivity of each pixel and polish adversarial noise of each image with a customized sampling setting. On the one hand, CAB uses current noise as a prior belief to customize the multivariate normal distribution. On the other hand, CAB keeps the new samplings away from historical failed queries to avoid similar mistakes. Experimental results measured on several image classification datasets emphasizes the validity of our method. Yahong Han, Qi Tian 0001 |
CVPR | 2 |
| 2020 | Extract and Merge: Superpixel Segmentation with Regional Attributes
Jianqiao An, Yahong Han, Meijun Sun, Qi Tian 0001 |
ECCV (30) | 3 |
| 2020 | Video Anomaly Detection Via Predictive Autoencoder With Gradient-Based AttentionabstractVideo anomaly detection is a challenging problem due to the ambiguity and diversity of anomalies in different scenes. In this paper, we present a novel framework to detect abnormal in surveillance videos. Inspired by the common deep reconstruction methods and deep prediction ones, we propose a new two-branch predictive autoencoder, including a reconstruction decoder and a prediction decoder, in which the prediction decoder is used to generate future frame and carry out anomaly detection by comparing the difference between predicted future frame and its ground truth. And the reconstruction decoder reconstructs the current frame, which can constrains the encoder to learn video representations better. Moreover, reconstruction decoder provides a gradient-based attention, which significantly helps the prediction decoder to generate higher quality future frame. Our method unifies reconstruction and prediction methods in an end-to-end framework, and it obtains impressive results with better predicted future frame on some publicly available datasets including CUHK Avenue and UCSD Pedestrian. Yuandu Lai, Yahong Han |
ICME | 3 |
| 2020 | Two-Way Feature-Aligned And Attention-Rectified Adversarial TrainingabstractAdversarial training increases robustness by augmenting training data with adversarial examples. However, vanilla adversarial training may be overfitting to certain adversarial attacks. Small perturbations in images bring in error which is gradually amplified when forwarded through the model so that the error leads to wrong classification. Besides, small perturbations will also distract classifier's attention to significant features that are relevant to the true label. In this paper, we propose a novel two-way feature-aligned and attention-rectified adversarial training (FAAR) to improve adversarial training (AT). FAAR utilizes two-way feature alignment and attention rectification to mitigate the problems mentioned above. FAAR effectively suppresses perturbations in lowlevel, high-level and global features by moving features of perturbed images towards those of clean images with twoway feature alignment. It also leads the model into focusing more on useful features which are correlated with true label through rectifying gradient-weighted attention. Besides, feature alignment activates attention rectification by reducing perturbations in high-level feature. Our proposed method FAAR surpasses other existing AT methods in three aspects. First, it pushes the model to keep invariant when dealing with different adversarial attacks and different magnitude of perturbations. Second, it can be applied to any convolution neural networks. Third, the training process is end-to-end. For experiments, FAAR shows promising defense performance on CIFAR-10 and ImageNet. Fan Jia 0006, Quanxin Zhang 0001, Yahong Han, Xiaohui Kuang, Yu-an Tan 0001 |
ICME | 4 |
| 2020 | Bidirectional Adversarial Training for Semi-Supervised Domain AdaptationabstractSemi-supervised domain adaptation (SSDA) is a novel branch of machine learning that scarce labeled target examples are available, compared with unsupervised domain adaptation. To make effective use of these additional data so as to bridge the domain gap, one possible way is to generate adversarial examples, which are images with additional perturbations, between the two domains and fill the domain gap. Adversarial training has been proven to be a powerful method for this purpose. However, the traditional adversarial training adds noises in arbitrary directions, which is inefficient to migrate between domains, or generate directional noises from the source to target domain and reverse. In this work, we devise a general bidirectional adversarial training method and employ gradient to guide adversarial examples across the domain gap, i.e., the Adaptive Adversarial Training (AAT) for source to target domain and Entropy-penalized Virtual Adversarial Training (E-VAT) for target to source domain. Particularly, we devise a Bidirectional Adversarial Training (BiAT) network to perform diverse adversarial trainings jointly. We evaluate the effectiveness of BiAT on three benchmark datasets and experimental results demonstrate the proposed method achieves the state-of-the-art. Pin Jiang, Aming Wu, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
IJCAI | 3 |
| 2020 | Multi-Modal fusion with multi-level attention for Visual Dialog
Jingping Zhang, Yahong Han |
Inf. Process. Manag. | 3 |
| 2020 | Adaptive iterative attack towards explainable adversarial robustness
Yahong Han, Quanxin Zhang 0001, Xiaohui Kuang |
Pattern Recognit. | 2 |
| 2020 | Sequence in sequence for video captioning
Huiyun Wang, Chongyang Gao, Yahong Han |
Pattern Recognit. Lett. | 3 |
| 2020 | Movie Question Answering via Textual Memory and Plot GraphabstractMovies provide us with a mass of visual content as well as attracting stories. Existing methods have illustrated that understanding movie stories through only visual content is still a hard problem. In this paper, for answering questions about movies, we introduce a new dataset called PlotGraphs, as external knowledge. The dataset contains massive graph-based information of movies. In addition, we put forward a model that can utilize movie clip, subtitle, and graph-based external knowledge. The model contains two main parts: a layered memory network (LMN) and a plot graph representation network (PGRN). In particular, the LMN can represent frame-level and clip-level movie content by the fixed word memory module and the adaptive subtitle memory module, respectively. And the plot graph representation network can represent the entire graph. We first extract words and sentences from the training movie subtitles and then the hierarchically formed movie representations, which are learned from LMN. At the same time, the PGRN can represent the semantic information and the relationships in the graph. We conduct extensive experiments on the MovieQA dataset and the PlotGraphs dataset. With only visual content as inputs, the LMN with frame-level representation obtains a large performance improvement. When incorporating subtitles into LMN to form the clip-level representation, we achieve the state-of-the-art performance on the online evaluation task of “Video+Subtitles.” After the integration of external knowledge, the performance of the model consisting of the LMN and the PGRN is further improved. The good performance successfully demonstrates that the external knowledge and the proposed model are effective for movie understanding. Yahong Han, Bo Wang 0011, Richang Hong, Fei Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Convolutional Reconstruction-to-Sequence for Video CaptioningabstractRecent advances towards video captioning mainly follow an encoder-decoder (sequence-to-sequence) framework and generate captions via a recurrent neural network (RNN). However, employing RNN as the decoder (generator) is prone to diluting long-term information, which weakens its ability to capture long-term dependencies. Recently, some work has demonstrated that the convolutional neural network (CNN) could be used to model sequential information. Though strengths in representation ability and computation efficiency, CNN has not been well exploited in video captioning. The reason partially comes from the difficulty of modeling multi-modal sequence with CNN. In this paper, we devise a novel CNN-based encoder-decoder framework for video captioning. Particularly, we first append inter-frame differences to each CNN-extracted frame feature to get a more discriminative representation; then with that as the input, we encode each frame to be a more compact feature by a one-layer convolutional mapping, which could be taken as a reconstruction network. In the decoding stage, we first fuse visual and lexical feature; then we stack multiple dilated convolutional layers to form a hierarchical decoder. As long-term dependencies could be captured by a shorter path along the hierarchical structure, the decoder could alleviate the loss of long-term information. Experiments on two benchmark datasets show that our method could obtain state-of-the-art performance. Aming Wu, Yahong Han, Yi Yang 0001, Qinghua Hu, Fei Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Discern Depth Under Foul Weather: Estimate PM2.5 for Depth InferenceabstractNowadays, haze is a common and serious problem and PM$_{2.5}$is a main measurement for air quality. Current methods estimate the level of primary pollutant with professional instruments, which is expensive and inconvenient. Moreover, with haze, the captured images will be unclear and are difficult to estimate the depth of the scene using passive methods. This article proposes a cheap, fast, and convenient PM$_{2.5}$estimation method that only need a captured image using daily-life devices, and further, discerns the depth of the scene using the estimated PM$_{2.5}$. We learn haze-relevant classified mapping via the hybrid convolutional neural network and combine the high-level features extracted from the convolutional layer with ground-truth PM$_{2.5}$to train support vector regression. The transmission map is computed using nonlocal sparse priors, and the depth map is inferred using the estimated PM$_{2.5}$value through the atmospheric scattering model. Experimental results demonstrate that our method achieves accurate PM$_{2.5}$estimation and depth inference. This could be very useful in many applications, for both clean and foul weather. Kun Li 0001, Yahong Han, Xibin Yue, Jing-Yu Yang 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Adaptive Sparse Confidence-Weighted Learning for Online Feature SelectionabstractIn this paper, we propose a new online feature selection algorithm for streaming data. We aim to focus on the following two problems which remain unaddressed in literature. First, most existing online feature selection algorithms merely utilize the first-order information of the data streams, regardless of the fact that second-order information explores the correlations between features and significantly improves the performance. Second, most online feature selection algorithms are based on the balanced data presumption, which is not true in many real-world applications. For example, in fraud detection, the number of positive examples are much less than negative examples because most cases are not fraud. The balanced assumption will make the selected features biased towards the majority class and fail to detect the fraud cases. We propose an Adaptive Sparse Confidence-Weighted (ASCW) algorithm to solve the aforementioned two problems. We first introduce an `0-norm constraint into the second-order confidence-weighted (CW) learning for feature selection. Then the original loss is substituted with a cost-sensitive loss function to address the imbalanced data issue. Furthermore, our algorithm maintains multiple sparse CW learner with the corresponding cost vector to dynamically select an optimal cost. We theoretically enhance the theory of sparse CW learning and analyze the performance behavior in F-measure. Empirical studies show the superior performance over the stateof-the-art online learning methods in the online-batch setting. Yanbin Liu 0003, Yan Yan 0006, Ling Chen 0006, Yahong Han, Yi Yang 0001 |
AAAI | 4 |
| 2019 | Curls & Whey: Boosting Black-Box Adversarial AttacksabstractImage classifiers based on deep neural networks suffer from harassment caused by adversarial examples. Two defects exist in black-box iterative attacks that generate adversarial examples by incrementally adjusting the noise-adding direction for each step. On the one hand, existing iterative attacks add noises monotonically along the direction of gradient ascent, resulting in a lack of diversity and adaptability of the generated iterative trajectories. On the other hand, it is trivial to perform adversarial attack by adding excessive noises, but currently there is no refinement mechanism to squeeze redundant noises. In this work, we propose Curls & Whey black-box attack to fix the above two defects. During Curls iteration, by combining gradient ascent and descent, we `curl' up iterative trajectories to integrate more diversity and transferability into adversarial examples. Curls iteration also alleviates the diminishing marginal effect in existing iterative attacks. The Whey optimization further squeezes the `whey' of noises by exploiting the robustness of adversarial perturbation. Extensive experiments on Imagenet and Tiny-Imagenet demonstrate that our approach achieves impressive decrease on noise magnitude in 12 norm. Curls & Whey attack also shows promising transferability against ensemble models as well as adversarially trained models. In addition, we extend our attack to the targeted misclassification, effectively reducing the difficulty of targeted attacks under black-box condition. Yahong Han |
CVPR | 3 |
| 2019 | 3D Shape Retrieval through Multilayer RBF Neural Networkabstract3D object retrieval involves more efforts mainly because major computer vision features are designed for 2D images, which is rarely applicable for 3D models. In this paper, we propose to retrieve the 3D models based on the implicit parameters learned from the radial base functions that represent the 3D objects. The radial base functions are learned from the RBF neural network. As deep neural networks can represent the data that is not linearly separable, we apply multiple layers' neural network to train the radial base functions. Our feature can be applied to recover the 3D objects, which proves the effectiveness of our features in representing the 3D objects. Furthermore, the dimensionality of the learned feature is scalable, which leads to memory efficiency. Experiments demonstrate the accuracy of our feature in 3D model retrieval. Guoyu Lu 0001, Yahong Han |
ICIP | 2 |
| 2019 | Multi-Timescale Context Encoding for Scene Parsing PredictionabstractPredicting the future is a crucial ability for intelligence systems. It is of great importance for many real-world applications, such as autonomous driving, which need scene parsing to understand the environment. Recent research has shown that predicting in semantic level is more effective than segmenting the predicted RGB frames. In order to label pixels in future frames correctly, the rich contextual dependencies should be exploit which existing methods paid less attention to. Therefore, we propose a novel network which catches both the short-term and long-term relations of observed frames for future scene parsing. Specifically, we introduce an attention mechanism to model semantic interdependencies between consecutive frames and a modified convolutional LSTM to model the correlations among all the input frames. Experiments validate that our approach outperforms other state-of-the-art methods on the large-scale Cityscapes dataset. Yahong Han |
ICME | 2 |
| 2019 | Visual Dialog with Targeted ObjectsabstractVisual Dialog aims to exchange information of the image between a questioner and an answerer through asking and answering questions alternately. To generate an accurate response to the target question requires to understand the visual information of the image according to the semantic information of the question and dialog history. However, existing methods paid attention to the visual feature of the whole image while ignoring that certain objects in the image always convey specific semantic. In this paper, we propose a novel visual dialog approach to focus on objects in the image which contain both visual information of attributes and semantic information of categories. Besides, we design several selection methods with different types of semantic guidance of the question and dialog history to select relevant objects. We evaluate the proposed approach on two datasets VisDial v0.9 and VisDial v1.0. And the performance improvement demonstrates the effectiveness of our approach. Yahong Han |
ICME | 2 |
| 2019 | Untargeted Adversarial Attack via Expanding the Semantic GapabstractRecent studies have demonstrated deep neural network-based image classifiers are vulnerable to adversarial examples. Although many existing methods could obtain outstanding attack performance, they often require certain information about the attacked model, e.g., the output category scores. Meanwhile, the optimization-based methods need many steps to generate adversarial examples. In practice, we could obtain the output label but the category scores. Besides, compared to those samples with large semantic category gaps, e.g., Panda and Gibbon, most existing methods are not easy to find adversarial examples on samples with small semantic category gaps, e.g., Tabby Cat and Egyptian Cat. Thus, we propose an untargeted adversarial attack method via expanding the semantic gap, which only relies on the output label. And we use the optimization-based method to generate adversarial examples. On five normally trained models and five state-of-the-art attack methods, extensive experiments show that our method is effective and obtains better attack performance. Aming Wu, Yahong Han, Quanxin Zhang 0001, Xiaohui Kuang |
ICME | 2 |
| 2019 | Video Interactive Captioning with Human PromptsabstractVideo captioning aims at generating a proper sentence to describe the video content. As a video often includes rich visual content and semantic details, different people may be interested in different views. Thus the generated sentence always fails to meet the ad hoc expectations. In this paper, we make a new attempt that, we launch a round of interaction between a human and a captioning agent. After generating an initial caption, the agent asks for a short prompt from the human as a clue of his expectation. Then, based on the prompt, the agent could generate a more accurate caption. We name this process a new task of video interactive captioning (ViCap). Taking a video and an initial caption as input, we devise the ViCap agent which consists of a video encoder, an initial caption encoder, and a refined caption generator. We show that the ViCap can be trained via a full supervision (with ground-truth) way or a weak supervision (with only prompts) way. For the evaluation of ViCap, we first extend the MSRVTT with interaction ground-truth. Experimental results not only show the prompts can help generate more accurate captions, but also demonstrate the good performance of the proposed method. Aming Wu, Yahong Han, Yi Yang 0001 |
IJCAI | 2 |
| 2019 | Hierarchical Variational Network for User-Diversified & Query-Focused Video SummarizationabstractThis paper focuses on the query-focused video summarization, which is an extended task of video summarization and aims to automatically generate user-oriented summary by highlighting frames/shots relevant to the query. This task is different from traditional video summarization in paying attention to users' subjectivity through queries. Diversity is a recognized important property in video summarization. However, existing methods only consider diversity as the dissimilarity between frames/shots which is far from user-oriented summarization. Users' different understandings of video should be an important source of diversity, reflected in the process of eliminating query-unrelated redundancy. To this end, this paper explores user-diversified & query-focused video summarization via a well-devised hierarchical variational network called HVN. HVN has three distinctive characteristics: (i) it has a hierarchical structure to model query-related long-range temporal dependency; (ii) it employs diverse attention mechanisms to encode query-related and context-important information and makes them balanced; (iii) it employs a multilevel self-attention module and a variational autoencoder module to add user-oriented diversity and stochastic factors. Experimental results demonstrate that HVN not only outperforms the state-of-the-arts but also improves the user-oriented diversity to some extent. Pin Jiang, Yahong Han |
ICMR | 2 |
| 2019 | Ranking Video Salient Object DetectionabstractVideo salient object detection has been attracting more and more research interests recently. However, the definition of salient objects in videos has been controversial all the time, which has become a critical bottleneck in video salient object detection. Specifically, the sequential information contained in videos results in a fact that objects have a relative saliency ranking between each other rather than specific saliency. This implies that simply distinguishing objects into salient or not-salient as usual could not represent the information about saliency comprehensively. To address this issue, 1) in this paper we propose a completely new definition for the salient objects in videos---ranking salient objects, which considers relative saliency ranking assisted with eye fixation points. 2) Based on this definition, a ranking video salient object dataset(RVSOD) is built. 3) Leveraging our RVSOD, a novel neural network called Synthesized Video Saliency Network (SVSNet) is constructed to detect both traditional salient objects and human eye movements in videos. Finally, a ranking saliency module (RSM) takes the results of SVSNet as input to generate the ranking saliency maps. We hope our approach will serve as a baseline and lead to a conceptually new research in the field of video saliency. Zheng Wang 0008, Xinyu Yan 0001, Yahong Han, Meijun Sun |
ACM Multimedia | 3 |
| 2019 | Connective Cognition Network for Directional Visual Commonsense ReasoningabstractVisual commonsense reasoning (VCR) has been introduced to boost research of cognition-level visual understanding, i.e., a thorough understanding of correlated details of the scene plus an inference with related commonsense knowledge. Recent studies on neuroscience have suggested that brain function or cognition can be described as a global and dynamic integration of local neuronal connectivity, which is context-sensitive to specific cognition tasks. Inspired by this idea, towards VCR, we propose a connective cognition network (CCN) to dynamically reorganize the visual neuron connectivity that is contextualized by the meaning of questions and answers. Concretely, we first develop visual neuron connectivity to fully model correlations of visual content. Then, a contextualization process is introduced to fuse the sentence representation with that of visual neurons. Finally, based on the output of contextualized connectivity, we propose directional connectivity to infer answers or rationales. Experimental results on the VCR dataset demonstrate the effectiveness of our method. Particularly, in $Q \to AR$ mode, our method is around 4\% higher than the state-of-the-art method. Aming Wu, Linchao Zhu, Yahong Han, Yi Yang 0001 |
NeurIPS | 3 |
| 2019 | Capturing the spatio-temporal continuity for video semantic segmentationabstractIn recent years, image semantic segmentation based on a convolutional neural network has achieved many advances. However, the development of video semantic segmentation is relatively slow. Directly applying the image segmentation algorithms to each video frame separately may ignore the temporal region continuity inherent in videos. In this study, the authors propose a novel deep neural network architecture with a newly devised spatio‐temporal continuity (STC) module for video semantic segmentation. Particularly, the architecture includes an encoding network, an STC module, and a decoding network. The encoding network is used to extract a high‐level feature map. The STC module then uses the high‐level feature map as input to extract the STC feature map. For decoding, they use four dilated convolutional layers to obtain more abstract representation and a deconvolutional layer to increase the size of the representation. Finally, they fuse the current feature representation and the previous feature representation and get the class probabilities. Thus, this architecture receives a sequence of consecutive video frames and outputs the segmentation result of the current frame. They extensively evaluate the proposed approach on the CamVid and KITTI datasets. Compared with other methods, the authors’ approach not only achieves competitive performance but also has lower complexity. Aming Wu, Yahong Han |
IET Image Process. | 3 |
| 2019 | DCT-CNN-based classification method for the Gongbi and Xieyi techniques of Chinese ink-wash paintings
Wei Jiang 0038, Zheng Wang 0008, Jesse S. Jin, Yahong Han, Meijun Sun |
Neurocomputing | 4 |
| 2019 | Detecting adversarial examples via prediction difference for deep neural networks
Qingjie Zhao, Xiaohui Kuang, Jianwei Zhang 0001, Yahong Han, Yu-an Tan 0001 |
Inf. Sci. | 6 |
| 2019 | Multi-cue fusion: Discriminative enhancing for person re-identification
Yongge Liu, Nan Song, Yahong Han |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Image captioning: from structural tetrad to translated sentences
Shubo Ma, Yahong Han |
Multim. Tools Appl. | 3 |
| 2019 | Semisupervised Regression With Optimized Rank for Matrix Data ClassificationabstractThere has been growing interest in developing more effective algorithms for matrix data classification. At present, most of the existing vector-based classifications involve vectorization process, which results in two main problems. First, the underlying structural information is disregarded. Second, the vectorization of a matrix incurs the creation of a vector with potentially very high dimensionality, which may lead to overfitting when the number of training data is small. To avoid such problems, we propose a new matrix-based regression algorithm for classification, in which the input matrices to be classified are directly used to learn two regression matrices for each order of the input matrix. To further explore the discrimination information, we add a joint ℓ2,1-norm on two regression matrices, which endows the algorithm optimized regression rank by uncovering common sparse columns in the two regression matrices. To further boost the classification performance, we incorporate a semisupervised learning process, which leverages both labeled and unlabeled data to enhance the training process. Experiments on public benchmark datasets show that our method outperforms a number of the existing state-of-the-art classification methods even when only few labeled training samples are provided. Jianguang Zhang, Jianmin Jiang, Yahong Han |
IEEE Trans. Cybern. | 3 |
| 2019 | Introduction to the Special Issue on the Cross-Media Analysis for Visual Question Answeringabstracteditorial Free Access Share on Introduction to the Special Issue on the Cross-Media Analysis for Visual Question Answering Editors: Richang Hong Hefei University of Technology, China Hefei University of Technology, ChinaView Profile , Yahong Han Tianjin University, China Tianjin University, ChinaView Profile , Tat-Seng Chua National University of Singapore, Singapore National University of Singapore, SingaporeView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 15Issue 2sApril 2019 Article No.: 48pp 1–3https://doi.org/10.1145/3337985Published:03 July 2019Publication History 1citation368DownloadsMetricsTotal Citations1Total Downloads368Last 12 Months72Last 6 weeks9 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteView all FormatsPDF Richang Hong, Yahong Han, Tat-Seng Chua |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Movie Question Answering: Remembering the Textual Cues for Layered Visual ContentsabstractMovies provide us with a mass of visual content as well as attracting stories. Existing methods have illustrated that understanding movie stories through only visual content is still a hard problem. In this paper, for answering questions about movies, we put forward a Layered Memory Network (LMN) that represents frame-level and clip-level movie content by the Static Word Memory module and the Dynamic Subtitle Memory module, respectively. Particularly, we firstly extract words and sentences from the training movie subtitles. Then the hierarchically formed movie representations, which are learned from LMN, not only encode the correspondence between words and visual content inside frames, but also encode the temporal alignment between sentences and frames inside movie clips. We also extend our LMN model into three variant frameworks to illustrate the good extendable capabilities. We conduct extensive experiments on the MovieQA dataset. With only visual content as inputs, LMN with frame-level representation obtains a large performance improvement. When incorporating subtitles into LMN to form the clip-level representation, we achieve the state-of-the-art performance on the online evaluation task of 'Video+Subtitles'. The good performance successfully demonstrates that the proposed framework of LMN is effective and the hierarchically formed movie representations have good potential for the applications of movie question answering. Bo Wang 0011, Youjiang Xu, Yahong Han, Richang Hong |
AAAI | 3 |
| 2018 | Image-Based PM2.5 Estimation and its Application on Depth EstimationabstractAir pollution is still a big threat to human health particularly for developing countries. It is highly demanding to measure air quality with daily-used devices such as smartphones. On the other hand, it is difficult to estimate the scene depth under the foul weather using traditional vision-based methods. This paper proposes an image-based method for PM2.5 estimation by capturing a single image. We extract high-level features based on convolutional neural network (CNN) and learn the mapping between the features and PM2.5 by support vector regression (SVR). Given a captured image, we can estimate the PM2.5 value in real time. With the estimated PM2.5, we can estimate the depth of scene using sparse prior and nonlocal bilateral kernel. Experimental results demonstrate that the proposed method achieves the same accuracy of PM2.5 estimation as commodity measurement devices, and estimates the accurate depth information that is even better than the “ground-truth” captured by a laser in the no-haze condition. Kun Li 0001, Yahong Han, Pufeng Du, Jing-Yu Yang 0002 |
ICASSP | 3 |
| 2018 | Schmidt: Image Augmentation for Black-Box Adversarial AttackabstractDespite achieving great success in multimedia analysis, especially in image recognition, deep neural networks (DNNs) can be easily fooled by maliciously crafted adversarial examples. Attacker who generates adversarial examples can even launch black-box adversarial attack by querying the target DNN model, without access to its internal structure or training set. In this work, we develop Schmidt Augmentation, an image augmentation method better probes decision boundaries of the black-box model. Schmidt Augmentation helps attackers achieve higher accuracy decrease on MNIST and CIFAR-10 datasets. We also shed light on the harshest circumstance that attacker only has access to samples of the target DNN by providing a labeling method based on semi-supervised learning instead of querying the target model. Yahong Han |
ICME | 2 |
| 2018 | Image-based Air Pollution Estimation Using Hybrid Convolutional Neural NetworkabstractAir pollution has a serious impact on our daily life, and how to quickly and easily measure the air pollution level without any expensive equipment is a quite challenging task. This paper proposes an air pollution estimation method using deep hybrid convolutional neural network from a single image, e.g., captured by a smartphone. The captured image is input to the main network, a very deep network, which solves the side effects of increased depth (degradation issues) by skip connection. This can improve network performance by simply increasing the depth of the network. Dark channel map is computed and fed into a secondary network to enrich the features with implicit representation. We have collected 1575 images of different scenes with different values of PM2.5to train the network in the end-to-end fusion mode. Experimental results on synthetic dataset and real captured dataset demonstrate that our method achieves excellent performance on classification of air pollution levels from a single captured image. Kun Li 0001, Yahong Han, Jing-Yu Yang 0002 |
ICPR | 3 |
| 2018 | Universal Perturbation Generation for Black-box Attack Using Evolutionary AlgorithmsabstractImage classifiers based on deep neural networks (DNNs) are vulnerable to tiny, imperceptible perturbations. Maliciously generated adversarial examples can exploit the instability of DNNs and mislead it into outputting a wrong classification result. Prior works showed the transferability of adversarial perturbations between models and between images. In this work, we shed light on the combination of source/target misclassification, black-box attack, and universal perturbation by employing improved evolutionary algorithms. We additionally find that the use of adversarial initialization enhances the efficiency of evolutionary algorithms finding universal perturbations. Experiments demonstrate impressive misclassification rates and surprising transferability for the proposed attack method using different models trained on CIFAR-10 and CIFAR-100 datasets. Our attach method also shows robustness against defensive measures like adversarial training. Sivy Wang, Yahong Han |
ICPR | 3 |
| 2018 | Multi-modal Circulant Fusion for Video-to-Language and BackwardabstractMulti-modal fusion has been widely involved in focuses of the modern artificial intelligence research, e.g., from visual content to languages and backward. Common-used multi-modal fusion methods mainly include element-wise product, element-wise sum, or even simply concatenation between different types of features, which are somewhat straightforward but lack in-depth analysis. Recent studies have shown fully exploiting interactions among elements of multi-modal features will lead to a further performance gain. In this paper, we put forward a new approach of multi-modal fusion, namely Multi-modal Circulant Fusion (MCF). Particularly, after reshaping feature vectors into circulant matrices, we define two types of interaction operations between vectors and matrices. As each row of the circulant matrix shifts one elements, with newly-defined interaction operations, we almost explore all possible interactions between vectors of different modalities. Moreover, as only regular operations are involved and defined a priori, MCF avoids increasing parameters or computational costs for multi-modal fusion. We evaluate MCF with tasks of video captioning and temporal activity localization via language (TALL). Experiments on MSVD and MSRVTT show our method obtains the state-of-the-art for video captioning. For TALL, by plugging into MCF, we achieve a performance gain of roughly 4.2% on TACoS. Aming Wu, Yahong Han |
IJCAI | 2 |
| 2018 | HeterStyle: A Heterogeneous Video Style Transfer ApplicationabstractVideo style transfer aims to synthesize a stylized video that preserves the content of a given video and is rendered in the style of a reference image.A key issue in video style transfer is how to balance video content preservation and reference style rendering, in order to avoid over-stylization with serious video content loss or under-stylization with unrecognized reference style. In this demonstration, we illustrate a novel video style transfer application, named HeterStyle, which can stylize different regions in the video with adaptive intensities.The core algorithm of HeterStyle application is our proposed heterogeneous video style transfer method, which minimizes a heterogeneous style transfer loss function considering content, style and temporal consistency in a Convolutional Neural Networks based optimization framework.With the HeterStyle application, a user can easily generate the stylized videos with good video content preservation and reference style rendering. Jingfan Guo, Tongwei Ren, Yahong Han, Lei Huang 0004, Gangshan Wu |
ACM Multimedia | 4 |
| 2018 | Explore Multi-Step Reasoning in Video Question AnsweringabstractVideo question answering (VideoQA) always involves visual reasoning. When answering questions composing of multiple logic correlations, models need to perform multi-step reasoning. In this paper, we formulate multi-step reasoning in VideoQA as a new task to answer compositional and logical structured questions based on video content. Existing VideoQA datasets are inadequate as benchmarks for the multi-step reasoning due to limitations such as lacking logical structure and having language biases. Thus we design a system to automatically generate a large-scale dataset, namely SVQA (Synthetic Video Question Answering). Compared with other VideoQA datasets, SVQA contains exclusively long and structured questions with various spatial and temporal relations between objects. More importantly, questions in SVQA can be decomposed into human readable logical tree or chain layouts, each node of which represents a sub-task requiring a reasoning operation such as comparison or arithmetic. Towards automatic question answering in SVQA, we develop a new VideoQA model. Particularly, we construct a new attention module, which contains spatial attention mechanism to address crucial and multiple logical sub-tasks embedded in questions, as well as a refined GRU called ta-GRU (temporal-attention GRU) to capture the long-term temporal dependency and gather complete visual cues. Experimental results show the capability of multi-step reasoning of SVQA and the effectiveness of our model when compared with other existing models. Xiaomeng Song, Yahong Han |
ACM Multimedia | 4 |
| 2018 | Spotting and Aggregating Salient Regions for Video CaptioningabstractTowards an interpretable video captioning process, we target to locate salient regions of video objects along with the sequentially uttering words. This paper proposes a new framework to automatically spot salient regions in each video frame and simultaneously learn a discriminative spatio-temporal representation for video captioning. First, in a Spot Module, we automatically learn the saliency value of each location to separate salient regions from video content as the foreground and the rest as background by two operations of 'hard separation' and 'soft separation', respectively. Then, in an Aggregate Module, to aggregate the foreground/background descriptors into a discriminative spatio-temporal representation, we devise a trainable video VLAD process to learn the aggregation parameters. Finally, we utilize the attention mechanism to decode the spatio-temporal representations of different regions into video descriptions. Experiments on two benchmark datasets demonstrate our method outperforms most of the state-of-the-art methods in terms of [email protected], METEOR and CIDEr metrics for the task of video captioning. Also examples demonstrate our method can successfully utter words to sequentially salient regions of video objects. Huiyun Wang, Youjiang Xu, Yahong Han |
ACM Multimedia | 3 |
| 2018 | Multi-task CNN Model for Action DetectionabstractAction detection is a challenging task since it requires locating actions of interest in both spatial and temporal. In this paper, a multi-task cnn model (MTCNN) which employs both spatial and temporal modules is proposed to solve this task. Specifically, the spatial module fuses appearance and motion information of frames which helps to regress the action bounding boxes in every frame more accurately, while the temporal module utilizes the 3D ConvNet which can effectively capture the temporal correlation between frames thus predict the time interval of action more precisely. Moreover, these two modules share information before their final outputs and are trained simultaneously. Experiments on UCF101-24 and J-HMDB-21 datasets demonstrate that our proposed pipeline outperforms most state-of-the-art methods. Yahong Han |
VCIP | 2 |
| 2018 | Guest Editorial: Spatio-temporal Feature Learning for Unconstrained Video Analysis
Yahong Han, Liqiang Nie, Fei Wu 0001 |
Multim. Tools Appl. | 1 |
| 2018 | Understanding the effective receptive field in semantic image segmentation
Yongge Liu, Jianzhuang Yu, Yahong Han |
Multim. Tools Appl. | 3 |
| 2018 | Discriminative multi-task multi-view feature selection and fusion for multimedia analysis
Ziwei Yang 0001, Huiyun Wang, Yahong Han, Xianglei Zhu |
Multim. Tools Appl. | 3 |
| 2018 | Distribution Sensitive Product QuantizationabstractProduct quantization (PQ) seems to have become the most efficient framework of performing approximate nearest neighbor (ANN) search for high-dimensional data. However, almost all existing PQ-based ANN techniques uniformly allocate precious bit budget to each subspace. This is not optimal, because data are often not evenly distributed among different subspaces. A better strategy is to achieve an improved balance between data distribution and bit budget within each subspace. Motivated by this observation, we propose to develop an optimized PQ (OPQ) technique, named distribution sensitive PQ (DSPQ) in this paper. The DSPQ dynamically analyzes and compares the data distribution based on a newly defined aggregate degree for high-dimensional data; whenever further optimization is feasible, resources such as memory and bits can be dynamically rearranged from one subspace to another. Our experimental results have shown that the strategy of bit rearrangement based on aggregate degree achieves modest improvements on most datasets. Moreover, our approach is orthogonal to the existing optimization strategy for PQ; therefore, it has been found that distribution sensitive OPQ can even outperform previous OPQ in the literature. Linhao Li, Qinghua Hu, Yahong Han, Xin Li 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Pooling the Convolutional Layers in Deep ConvNets for Video Action RecognitionabstractDeep ConvNets have shown their good performance in image classification tasks. However, there still remains problems in deep video representations for action recognition. On one hand, current video ConvNets are relatively shallow compared with image ConvNets, which limits their capability of capturing the complex video action information; on the other hand, temporal information of videos is not properly utilized to pool and encode the video sequences. Toward these issues, in this paper we utilize two state-of-the-art ConvNets, i.e., the very deep spatial net (VGGNet [1]) and the temporal net from Two-Stream ConvNets [2], for action representation. The convolutional layers and the proposed new layer, called frame-diff layer, are extracted and pooled with two temporal pooling strategies: Trajectory pooling and Line pooling. The pooled local descriptors are then encoded with vector of locally aggregated descriptors (VLAD) [3] to form the video representations. In order to verify the effectiveness of the proposed framework, we conduct experiments on UCF101 and HMDB51 data sets. It achieves accuracy of 92.08% on UCF101, which is the state-of-the-art, and the accuracy of 65.62% on HMDB51, which is comparable to the state-of-the-art. In addition, we propose the new Line pooling strategy, which can speed up the extraction of feature and achieve the comparable performance of the Trajectory pooling. Shichao Zhao, Yanbin Liu 0003, Yahong Han, Richang Hong, Qinghua Hu, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Sequential Video VLAD: Training the Aggregation Locally and TemporallyabstractAs characterizing videos simultaneously from spatial and temporal cues has been shown crucial for the video analysis, the combination of convolutional neural networks and recurrent neural networks, i.e., recurrent convolution networks (RCNs), should be a native framework for learning the spatio-temporal video features. In this paper, we develop a novel sequential vector of locally aggregated descriptor (VLAD) layer, named SeqVLAD, to combine a trainable VLAD encoding process and the RCNs architecture into a whole framework. In particular, sequential convolutional feature maps extracted from successive video frames are fed into the RCNs to learn soft spatio-temporal assignment parameters, so as to aggregate not only detailed spatial information in separate video frames but also fine motion information in successive video frames. Moreover, we improve the gated recurrent unit (GRU) of RCNs by sharing the input-to-hidden parameters and propose an improved GRU-RCN architecture named shared GRU-RCN (SGRU-RCN). Thus, our SGRU-RCN has a fewer parameters and a less possibility of overfitting. In experiments, we evaluate SeqVLAD with the tasks of video captioning and video action recognition. Experimental results on Microsoft Research Video Description Corpus, Montreal Video Annotation Dataset, UCF101, and HMDB51 demonstrate the effectiveness and good performance of our method. Youjiang Xu, Yahong Han, Richang Hong, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Top attention in line with time: A light-weight strategyabstractFor video representation, dense sampling along trajectories or optical flow stacking are both heavy-cost computations. This paper aims to develop a light-weight strategy which could skip the computations of optical flow and trajectories. Particularly, taking frames as inputs to a pre-trained ConvNet, we extract top layers as video feature maps. Instead of trajectory pooling, we directly pooled these feature maps in line with time, which is named Line Pooling. We utilize the proposed Line-pooled Deep-convolutional Descriptors (LDDs) to weight regions with high motion saliency, which turns out to pay attention to actions in line with time. Experiments on UCF101 and HMDB51 demonstrate the efficiency, effectiveness, and promising performance of our method. Youjiang Xu, Shichao Zhao, Yahong Han, Qinghua Hu, Fei Wu 0001 |
ICME | 3 |
| 2017 | Catching the Temporal Regions-of-Interest for Video CaptioningabstractAs a crucial challenge for video understanding, exploiting the spatial-temporal structure of video has attracted much attention recently, especially on video captioning. Inspired by the insight that people always focus on certain interested regions of video content, we propose a novel approach which will automatically focus on regions-of-interest and catch their temporal structures. In our approach, we utilize a specific attention model to adaptively select regions-of-interest for each video frame. Then a Dual Memory Recurrent Model (DMRM) is introduced to incorporate temporal structure of global features and regions-of-interest features in parallel, which will obtain rough understanding of video content and particular information of regions-of-interest. Since the attention model could not always catch the right interests, we additionally adopt semantic supervision to attend to interested regions more correctly. We evaluate our method for video captioning on two public benchmarks: the Microsoft Video Description Corpus (MSVD) and the Montreal Video Annotation Dataset (M-VAD). The experiments demonstrate that catching temporal regions-of-interest information really enhances the representation of input videos and our approach obtains the state-of-the-art results on popular evaluation metrics like BLEU-4, CIDEr, and METEOR. Ziwei Yang 0001, Yahong Han, Zheng Wang 0008 |
ACM Multimedia | 2 |
| 2017 | Multirate Multimodal Video CaptioningabstractAutomatically describing videos with natural language is a crucial challenge of video understanding. Compared to images, videos have specific spatial-temporal structure and various modality information. In this paper, we propose a Multirate Multimodal Approach for video captioning. Considering that the speed of motion in videos varies constantly, we utilize a Multirate GRU to capture temporal structure of videos. It encodes video frames with different intervals and has a strong ability to deal with motion speed variance. As videos contain different modality cues, we design a particular multimodal fusion method. By incorporating visual, motion, and topic information together, we construct a well-designed video representation. Then the video representation is fed into a RNN-based language model for generating natural language descriptions. We evaluate our approach for video captioning on "Microsoft Research - Video to Text" (MSR-VTT), a large-scale video benchmark for video understanding. And our approach gets great performance on the 2nd MSR Video to Language Challenge. Ziwei Yang 0001, Youjiang Xu, Huiyun Wang, Bo Wang 0011, Yahong Han |
ACM Multimedia | 5 |
| 2017 | Guest Editorial: Intermediate representation for vision and multimedia applications
Yan Yan 0002, Yahong Han, Petia Radeva, Qi Tian 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Semi-supervised tensor learning for image classification
Jianguang Zhang, Yahong Han, Jianmin Jiang |
Multim. Syst. | 2 |
| 2017 | Semi-Supervised Image-to-Video Adaptation for Video Action RecognitionabstractHuman action recognition has been well explored in applications of computer vision. Many successful action recognition methods have shown that action knowledge can be effectively learned from motion videos or still images. For the same action, the appropriate action knowledge learned from different types of media, e.g., videos or images, may be related. However, less effort has been made to improve the performance of action recognition in videos by adapting the action knowledge conveyed from images to videos. Most of the existing video action recognition methods suffer from the problem of lacking sufficient labeled training videos. In such cases, over-fitting would be a potential problem and the performance of action recognition is restrained. In this paper, we propose an adaptation method to enhance action recognition in videos by adapting knowledge from images. The adapted knowledge is utilized to learn the correlated action semantics by exploring the common components of both labeled videos and images. Meanwhile, we extend the adaptation method to a semi-supervised framework which can leverage both labeled and unlabeled videos. Thus, the over-fitting can be alleviated and the performance of action recognition is improved. Experiments on public benchmark datasets and real-world datasets show that our method outperforms several other state-of-the-art action recognition methods. Jianguang Zhang, Yahong Han, Jinhui Tang 0001, Qinghua Hu, Jianmin Jiang |
IEEE Trans. Cybern. | 2 |
| 2017 | Semisupervised Online Multikernel Similarity Learning for Image RetrievalabstractMetric learning plays a fundamental role in the fields of multimedia retrieval and pattern recognition. Recently, an online multikernel similarity (OMKS) learning method has been presented for content-based image retrieval (CBIR), which was shown to be promising for capturing the intrinsic nonlinear relations within multimodal features from large-scale data. However, the similarity function in this method is learned only from labeled images. In this paper, we present a new framework to exploit unlabeled images and develop a semisupervised OMKS algorithm. The proposed method is a multistage algorithm consisting of feature selection, selective ensemble learning, active sample selection, and triplet generation. The novel aspects of our work are the introduction of classification confidence to evaluate the labeling process and select the reliably labeled images to train the metric function, and a method for reliable triplet generation, where a new criterion for sample selection is used to improve the accuracy of label prediction for unlabeled images. Our proposed method offers advantages in challenging scenarios, in particular, for a small set of labeled images with high-dimensional features. Experimental results demonstrate the effectiveness of the proposed method as compared with several baseline methods. Jianqing Liang, Qinghua Hu, Wenwu Wang 0001, Yahong Han |
IEEE Trans. Multim. | 4 |
| 2016 | Describing images by feeding LSTM with structural wordsabstractGenerating semantic description draws increasing attention recently. Describing objects with adaptive adjunct words make the sentence more informative. In this paper, we focus on the generation of descriptions for images according to the structural words we have generated such as a tetrad of. We propose to use deep machine translation method to generate semantically meaningful descriptions. In particular, the description is composed of objects with appropriate adjunct words, corresponding activities and scene. We propose to use a multi-task method to generate structural words. Taking these words sequence as source language, we train a LSTM encoder-decoder machine translation model to output the target language. Experiments on the benchmark datasets demonstrate our method has better performance than state-of-the-art methods of image caption in terms of language generation metrics. Shubo Ma, Yahong Han |
ICME | 2 |
| 2016 | Large-Scale E-Commerce Image Retrieval with Top-Weighted Convolutional Neural NetworksabstractSeveral recent researches have shown that image features produced by Convolutional Neural Networks (CNNs) provide the state-of-the-art performance for image classification and retrieval. Moreover, some researchers have found that the features extracted from the deep convolutional layers of CNNs perform better than that from the fully-connected layers. Features extracted from the convolutional layers have a natural interpretation: descriptors of local image regions correspond well to the receptive fields of the particular features. In order to obtain both representative and discriminative descriptors for large-scale e-commerce image retrieval, we come up with a new feature extraction framework. At first, we propose the Top-Weight method to detect the interesting area of e-commerce images automatically. With the estimated weight, we then aggregate local deep features and produce high-quality global representation for e-commerce image retrieval. We have conducted experiments on an e-commerce dataset ALISC [1] released by Alibaba Group. Experimental results show that our method outperforms other deep learning based methods. Shichao Zhao, Youjiang Xu, Yahong Han |
ICMR | 3 |
| 2016 | Describing Images with Ontology-Aware Dictionary Learning
Chengyue Zhang, Yahong Han |
MMM (1) | 2 |
| 2016 | Guest editorial: Adaptation methods for multimedia analysis
Yahong Han, Yi Yang 0001 |
Neurocomputing | 1 |
| 2016 | Cluster structure preserving unsupervised feature selection for multi-view tasks
Yahong Han, Qinghua Hu |
Neurocomputing | 3 |
| 2016 | Hierarchical support vector machine based structural classification with fused hierarchies
Yahong Han, Quan Zou 0001, Qinghua Hu |
Neurocomputing | 2 |
| 2016 | Combining neighborhood separable subspaces for classification via sparsity regularized optimization
Pengfei Zhu 0001, Qinghua Hu, Yahong Han, Changqing Zhang 0002 |
Inf. Sci. | 3 |
| 2016 | Semi-supervised image clustering with multi-modal information
Jianqing Liang, Yahong Han, Qinghua Hu |
Multim. Syst. | 2 |
| 2016 | Semi-supervised feature selection via hierarchical regression for web image classification
Jianguang Zhang, Yahong Han, Jianmin Jiang |
Multim. Syst. | 3 |
| 2016 | Tucker decomposition-based tensor learning for human action recognition
Jianguang Zhang, Yahong Han, Jianmin Jiang |
Multim. Syst. | 2 |
| 2016 | Image attribute learning with ontology guided fused lasso
Zhiyong Feng 0002, Yahong Han |
Multim. Tools Appl. | 3 |
| 2016 | Sketch4Image: a novel framework for sketch-based image retrieval based on product quantization with coding residuals
Yahong Han, Jianwu Dang 0001 |
Multim. Tools Appl. | 2 |
| 2016 | Guest editorial: web multimedia semantic inference using multi-cues
Yahong Han, Yi Yang 0001, Xiaofang Zhou 0001 |
World Wide Web | 1 |
| 2015 | Discriminative multi-view feature selection and fusionabstractIn computer vision tasks such as action recognition and image classification, combining multiple visual feature sets is proven to be an effective strategy. However, simply combing these features may cause high dimensionality and lead to noises. Feature selection and fusion are common choices for multiple feature representation. In this paper, we propose a multi-view feature selection and fusion method which chooses and fuses discriminative features from multiple feature sets. For discriminative feature selection, we learn the selection matrix W by the minimization of the trace ratio objective function with ℓ2,1norm regularization. For multiple feature fusion, we incorporate local structures of each view in the Laplacian matrix. Since the Laplacian matrix is constructed in unsupervised manner and scaled category indicator matrix is solved iteratively, our work is fully unsupervised. Experimental results on four action recognition datasets and two large-scale image classification datasets demonstrate the effectiveness of multi-view feature selection and fusion. Yanbin Liu 0003, Binbing Liao, Yahong Han |
ICME | 3 |
| 2015 | Inferring Painting Style with Multi-Task Dictionary Learning
Gaowen Liu, Yan Yan 0002, Elisa Ricci 0001, Yi Yang 0001, Yahong Han, Stefan Winkler 0001, Nicu Sebe |
IJCAI | 5 |
| 2015 | Describing Images with Hierarchical Concepts and Object Class LocalizationabstractCurrent research into automatic generation of semantic descriptions centers mainly on improving the annotation accuracy for individual tag or attributes. In this paper, we focus on the generation of more informative descriptions for images. We proposes to generate layered, semantically meaningful descriptions and create summaries of key aspects of the data from the component detectors. In particular, the output descriptions include superclass, class, attributes, and the location of the area of the object which may interest users. We propose to integrate ROI (Region of Interest) identification and hierarchical semantic elements detection into a joint framework. The joint optimization of the ROI localizer and the hierarchical concept detection make them mutually beneficial and reciprocal. In this way, we create a discriminative image description generation framework based on a tightly coupled multi-layer optimization. The output descriptions contain richer information of the image content with layered contextual information, thereby enabling better management and usage of image data. Experiments on two public open benchmark datasets demonstrate that the proposed method obtains state of the art performance. Yahong Han |
ICMR | 1 |
| 2015 | Summarization-based Video Caption via Deep Neural NetworksabstractGenerating appropriate descriptions for visual content draws increasing attention recently, where the promising progresses were obtained owing to the breakthroughs in deep neural networks. Different from the traditional SVO (subject, verb, object) based methods, in this paper, we propose a novel framework of video caption via deep neural networks. For each frame, we extract visual features by a fine-tuned deep Convulutional Neural Networks (CNN), which are then fed into a Recurrent Neural Networks (RNN) to generate novel sentences descriptions for each frame. In order to obtain the most representative and high-quality descriptions for target video, a well-devised automatic summarization process is incorporated to reduce the noises by ranking on the sentence-sequence graph. Moreover, our framework owns the merit of describing out-of-sample videos by transferring knowledge from pre-captioned images. Experiments on the benchmark datasets demonstrate our method has better performance than the state-of-the-art methods of video caption in language generation metrics as well as SVO accuracy. Shubo Ma, Yahong Han |
ACM Multimedia | 3 |
| 2015 | Tensor rank selection for multimedia analysis
Jianguang Zhang, Yahong Han, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Image aesthetics enhancement using composition-based saliency detection
Handong Zhao, Jingjing Chen 0001, Yahong Han, Xiaochun Cao |
Multim. Syst. | 3 |
| 2015 | Guest Editorial: Ad Hoc Web Multimedia Analysis with Limited Supervision
Yahong Han, Yi Yang 0001, Jingdong Wang 0001 |
Multim. Tools Appl. | 1 |
| 2015 | An Object-Level High-Order Contextual Descriptor Based on Semantic, Spatial, and Scale CuesabstractContext has been playing an increasingly important role in areas such as object detection, scene understanding, and image segmentation. Although many different types of contextual cues have been successfully explored, most of them only consider the pair-wise relationship between objects or parts. Several models utilize the high-order relationship for encoding contextual information. However, they mainly use a single contextual cue. In this paper, we present a novel high-order contextual descriptor (HOOD) to measure the strength of interactions among objects within an image. Heterogeneous contextual cues like semantic, spatial, and scale contexts are jointly integrated into HOOD to define the high-order interactions. The strength of these interactions are inferred by applying Bayes' rule on the pure dependence of the involved objects. As a result, an object-level graph is constructed to represent the contextually consistent interactions. Moreover, we propose a HOOD based object localization framework to verify the effectiveness of HOOD. Experimental results on two benchmark datasets including SUN09 and PASCAL2007 show that our framework outperforms the state-of-the-art context based object localization methods. Finally, we apply HOOD on two multimedia applications: structured image retrieval and out-of-context object detection, which demonstrates the potential usages of HOOD. Xiaochun Cao, Xingxing Wei 0001, Yahong Han, Xiaowu Chen 0001 |
IEEE Trans. Cybern. | 3 |
| 2015 | Robust Face Clustering Via Tensor DecompositionabstractFace clustering is a key component either in image managements or video analysis. Wild human faces vary with the poses, expressions, and illumination changes. All kinds of noises, like block occlusions, random pixel corruptions, and various disguises may also destroy the consistency of faces referring to the same person. This motivates us to develop a robust face clustering algorithm that is less sensitive to these noises. To retain the underlying structured information within facial images, we use tensors to represent faces, and then accomplish the clustering task based on the tensor data. The proposed algorithm is called robust tensor clustering (RTC), which firstly finds a lower-rank approximation of the original tensor data using a L1 norm optimization function. Because L1 norm does not exaggerate the effect of noises compared with L2 norm, the minimization of the L1 norm approximation function makes RTC robust. Then, we compute high-order singular value decomposition of this approximate tensor to obtain the final clustering results. Different from traditional algorithms solving the approximation function with a greedy strategy, we utilize a nongreedy strategy to obtain a better solution. Experiments conducted on the benchmark facial datasets and gait sequences demonstrate that RTC has better performance than the state-of-the-art clustering algorithms and is more robust to noises. Xiaochun Cao, Xingxing Wei 0001, Yahong Han, Dongdai Lin |
IEEE Trans. Cybern. | 3 |
| 2015 | Compact and Discriminative Descriptor Inference Using Multi-CuesabstractFeature descriptors around local interest points are widely used in human action recognition both for images and videos. However, each kind of descriptors describes the local characteristics around the reference point only from one cue. To enhance the descriptive and discriminative ability from multiple cues, this paper proposes a descriptor learning framework to optimize the descriptors at the source by learning a projection from multiple descriptors' spaces to a new Euclidean space. In this space, multiple cues and characteristics of different descriptors are fused and complemented for each other. In order to make the new descriptor more discriminative, we learn the multi-cue projection by the minimization of the ratio of within-class scatter to between-class scatter, and therefore, the discriminative ability of the projected descriptor is enhanced. In the experiment, we evaluate our framework on the tasks of action recognition from still images and videos. Experimental results on two benchmark image and two benchmark video data sets demonstrate the effectiveness and better performance of our method. Yahong Han, Yi Yang 0001, Fei Wu 0001, Richang Hong |
IEEE Trans. Image Process. | 1 |
| 2015 | Semisupervised Feature Selection via Spline Regression for Video Semantic RecognitionabstractTo improve both the efficiency and accuracy of video semantic recognition, we can perform feature selection on the extracted video features to select a subset of features from the high-dimensional feature set for a compact and accurate video data representation. Provided the number of labeled videos is small, supervised feature selection could fail to identify the relevant features that are discriminative to target classes. In many applications, abundant unlabeled videos are easily accessible. This motivates us to develop semisupervised feature selection algorithms to better identify the relevant video features, which are discriminative to target classes by effectively exploiting the information underlying the huge amount of unlabeled video data. In this paper, we propose a framework of video semantic recognition by semisupervised feature selection via spline regression (S(2)FS(2)R) . Two scatter matrices are combined to capture both the discriminative information and the local geometry structure of labeled and unlabeled training videos: A within-class scatter matrix encoding discriminative information of labeled training videos and a spline scatter output from a local spline regression encoding data distribution. An l2,1 -norm is imposed as a regularization term on the transformation matrix to ensure it is sparse in rows, making it particularly suitable for feature selection. To efficiently solve S(2)FS(2)R , we develop an iterative algorithm and prove its convergency. In the experiments, three typical tasks of video semantic recognition, such as video concept detection, video classification, and human action recognition, are used to demonstrate that the proposed S(2)FS(2)R achieves better performance compared with the state-of-the-art methods. Yahong Han, Yi Yang 0001, Yan Yan 0006, Zhigang Ma, Nicu Sebe, Xiaofang Zhou 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2014 | Output Feature Augmented LassoabstractLasso simultaneously conducts variable selection and supervised regression. In this paper, we extend Lasso to multiple output prediction, which belongs to the categories of structured learning. Though structured learning makes use of both input and output simultaneously, the joint feature mapping in current framework of structured learning is usually application-specific. As a result, ad hoc heuristics have to be employed to design different joint feature mapping functions for different applications, which results in the lackness of generalization ability for multiple output prediction. To address this limitation, in this paper, we propose to augment Lasso with output by decoupling the joint feature mapping function of traditional structured learning. The contribution of this paper is three-fold: 1) The augmented Lasso conducts regression and variable selection on both the input and output features, and thus the learned model could fit an output with both the selected input variables and the other correlated outputs. 2) To be more general, we set up nonlinear dependencies among output variables by generalized Lasso. 3) Moreover, the Augmented Lagrangian Method (ALM) with Alternating Direction Minimizing (ADM) strategy is used to find the optimal model parameters. The extensive experimental results demonstrate the effectiveness of the proposed method. Changqing Zhang 0002, Yahong Han, Xiaojie Guo 0001, Xiaochun Cao |
ICDM | 2 |
| 2014 | Attribute prediction with long-range interactions via path codingabstractDue to the describable or human-nameable nature of visual attributes, the appropriate utilization of attributes has been receiving much attention in recent years in many applications. Motivated by the assumption that the long-range interactions between attributes can boost image understanding and classification, path coding is utilized in this paper to model the long-range interactions between attributes for the attribute prediction, we call it attribute prediction via a path coding penalty (abbreviated as AP2CP). AP2CP not only introduces structured sparsity penalties over paths on a directed acyclic graph, but also captures the intrinsical long-range dependent interactions between attributes. The proposed AP2CP can be efficiently solved by leveraging network flow optimization. The experiments show that the proposed AP2CP achieves a better performance in attribute prediction. Zhuhao Wang, Fei Wu 0001, Yahong Han, Jiebo Luo 0001, Qi Tian 0001, Yueting Zhuang |
ICIP | 3 |
| 2014 | Augmented Image Retrieval using Multi-order Object Layout with AttributesabstractIn image retrieval, users' search intention is usually specified by textual queries, exemplar images, concept maps, and even sketches, which can only express the search intention partially. These query strategies lack the abilities to indicate the Regions Of Interests (ROIs) and represent the spatial or semantic correlations among the ROIs, which results in the so-called semantic gap between users' search intention and images' low-level visual content. In this paper, we propose a novel image search method, which allows the users to indicate any number of Regions Of Interest (ROIs) within the query as well as utilize various semantic concepts and spatial relations to search images. Specifically, we firstly propose a structured descriptor to jointly represent the categories, attributes, and spatial relations among objects. Then, based on the defined descriptor, our method ranks the images in the database according to the matching scores w.r.t. the category, attribute, and spatial relations. We conduct the experiments on the aPascal and aYahoo datasets, and experimental results show the advantage of the proposed method compared to the state of the arts. Xiaochun Cao, Xingxing Wei 0001, Xiaojie Guo 0001, Yahong Han, Jinhui Tang 0001 |
ACM Multimedia | 4 |
| 2014 | What Can We Learn about Motion Videos from Still Images?abstractHuman action recognition from motion videos plays an important role in multimedia analysis. Different from the temporal cues of action series in motion videos, the motion tendency can also be revealed from the still images or key frames. Thus, if the action knowledge in related still images can be well adapted to the target motion videos, we would have a great chance to improve the performance of video action recognition. In this paper, we propose a framework of Still-to-Motion Adaptation (SMA) for human action recognition. Common visual features are extracted both from the related images and target videos' key frames, by which the gap between still images and videos are bridged. Meanwhile, to utilize the unlabeled training videos in target domain, we incorporate a semi-supervised process into our framework. By minimizing the difference of action prediction from still features and motion features, we formulate the still-to-motion adaptation into a joint optimization process. Experiments successfully demonstrate the effectiveness of the proposed framework and show the better performance of action recognition compared with the state-of-the-art methods. We also analyze the impact on the recognition results of target videos by knowledge adaptation from still images. Jianguang Zhang, Yahong Han, Jinhui Tang 0001, Qinghua Hu, Jianmin Jiang |
ACM Multimedia | 2 |
| 2014 | Image decomposing for inpainting using compressed sensing in DCT domain
Yahong Han, Jianwu Dang 0001 |
Frontiers Comput. Sci. | 2 |
| 2014 | Feature selection with spatial path coding for multimedia analysis
Yahong Han, Jingjing Chen 0001, Xiaochun Cao, Congfu Xu, Haoquan Shen |
Inf. Sci. | 1 |
| 2014 | Regularity Preserved Superpixels and SupervoxelsabstractMost existing superpixel algorithms ignore the spatial structure and regularity properties, which result in undesirable sizes and location relationships for the subsequent processing. In this paper, we introduce a new method to generate the regularity preserved superpixels. Starting from the lattice seeds, our method relocates them to the pixel with locally maximal edge magnitudes and treats them as the superpixel junctions. Then, the shortest path algorithm is employed to find the local optimal boundary connecting each adjacent junction pair. Thanks to the local constraints, our method obtains homogeneous superpixels with adjacency in lowly textured and uniform regions and simultaneously preserves the boundary adherence in the high contrast contents. Our method preserves the regularity property without significantly sacrificing the segmentation accuracy. Moreover, we extend this regular constraint for generating the supervoxels. Our method obtains the regular supervoxels, which preserves the structural relation on both spatial and temporal spaces of the video. Quantitative and qualitative experimental results on benchmark datasets demonstrate that our simple but effective method outperforms the existing regular superpixel methods. Huazhu Fu, Xiaochun Cao, Dai Tang, Yahong Han, Dong Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2014 | Augmenting Image Descriptions Using Structured Prediction OutputabstractThe need for richer descriptions of images arises in a wide spectrum of applications ranging from image understanding to image retrieval. While the Automatic Image Annotation (AIA) has been extensively studied, image descriptions with the output labels lack sufficient information. This paper proposes to augment image descriptions using structured prediction output. We define a hierarchical tree-structured semantic unit to describe images, from which we can obtain not only the class and subclass one image belongs to, but also the attributes one image has. After defining a new feature map function of structured SVM, we decompose the loss function into every node of the hierarchical tree-structured semantic unit and then predict the tree-structured semantic unit for testing images. In the experiments, we evaluate the performance of the proposed method on two open benchmark datasets and compare with the state-of-the-art methods. Experimental results show the better prediction performance of the proposed method and demonstrate the strength of augmenting image descriptions. Yahong Han, Xingxing Wei 0001, Xiaochun Cao, Yi Yang 0001, Xiaofang Zhou 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Image Attribute AdaptationabstractVisual attributes can be considered as a middle-level semantic cue that bridges the gap between low-level image features and high-level object classes. Thus, attributes have the advantage of transcending specific semantic categories or describing objects across categories. Since attributes are often human-nameable and domain specific, much work constructs attribute annotations ad hoc or take them from an application-dependent ontology. To facilitate other applications with attributes, it is necessary to develop methods which can adapt a well-defined set of attributes to novel images. In this paper, we propose a framework for image attribute adaptation. The goal is to automatically adapt the knowledge of attributes from a well-defined auxiliary image set to a target image set, thus assisting in predicting appropriate attributes for target images. In the proposed framework, we use a non-linear mapping function corresponding to multiple base kernels to map each training images of both the auxiliary and the target sets to a Reproducing Kernel Hilbert Space (RKHS), where we reduce the mismatch of data distributions between auxiliary and target images. In order to make use of un-labeled images, we incorporate a semi-supervised learning process. We also introduce a robust loss function into our framework to remove the shared irrelevance and noise of training images. Experiments on two couples of auxiliary-target image sets demonstrate that the proposed framework has better performance of predicting attributes for target testing images, compared to three baselines and two state-of-the-art domain adaptation methods. Yahong Han, Yi Yang 0001, Zhigang Ma, Haoquan Shen, Nicu Sebe, Xiaofang Zhou 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | Robust Tensor Clustering with Non-Greedy Maximization
Xiaochun Cao, Xingxing Wei 0001, Yahong Han, Yi Yang 0001, Dongdai Lin |
IJCAI | 3 |
| 2013 | Co-Regularized Ensemble for Feature Selection
Yahong Han, Yi Yang 0001, Xiaofang Zhou 0001 |
IJCAI | 1 |
| 2013 | Object coding on the semantic graph for scene classificationabstractIn the scene classification, a scene can be considered as a set of object cliques. Objects inside each clique have semantic correlations with each other, while two objects from different cliques are relatively independent. To utilize these correlations for better recognition performance, we propose a new method - Object Coding on the Semantic Graph to address the scene classification problem. We first exploit prior knowledge by making statistics on a large number of labeled images and calculating the dependency degree between objects. Then, a graph is built to model the semantic correlations between objects. This semantic graph captures semantics by treating the objects as vertices and the objects affinities as the weights of edges. By encoding this semantic knowledge into the semantic graph, object coding is conducted to automatically select a set of object cliques that have strongly semantic correlations to represent a specific scene. The experimental results show that the Object Coding on semantic graph can improve the classification accuracy. Jingjing Chen 0001, Yahong Han, Xiaochun Cao, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2013 | Unified Dictionary Learning and Region Tagging with Hierarchical Sparse Representation
Xiaochun Cao, Xingxing Wei 0001, Yahong Han, Yi Yang 0001, Nicu Sebe, Alex Hauptmann 0001 |
Comput. Vis. Image Underst. | 3 |
| 2013 | Image classification with manifold learning for out-of-sample data
Yahong Han, Zhongwen Xu, Zhigang Ma, Zi Huang |
Signal Process. | 1 |
| 2013 | Discovering Discriminative Graphlets for Aerial Image Categories RecognitionabstractRecognizing aerial image categories is useful for scene annotation and surveillance. Local features have been demonstrated to be robust to image transformations, including occlusions and clutters. However, the geometric property of an aerial image (i.e., the topology and relative displacement of local features), which is key to discriminating aerial image categories, cannot be effectively represented by state-of-the-art generic visual descriptors. To solve this problem, we propose a recognition model that mines graphlets from aerial images, where graphlets are small connected subgraphs reflecting both the geometric property and color/texture distribution of an aerial image. More specifically, each aerial image is decomposed into a set of basic components (e.g., road and playground) and a region adjacency graph (RAG) is accordingly constructed to model their spatial interactions. Aerial image categories recognition can subsequently be casted as RAG-to-RAG matching. Based on graph theory, RAG-to-RAG matching is conducted by comparing all their respective graphlets. Because the number of graphlets is huge, we derive a manifold embedding algorithm to measure different-sized graphlets, after which we select graphlets that have highly discriminative and low redundancy topologies. Through quantizing the selected graphlets from each aerial image into a feature vector, we use support vector machine to discriminate aerial image categories. Experimental results indicate that our method outperforms several state-of-the-art object/scene recognition models, and the visualized graphlets indicate that the discriminative patterns are discovered by our proposed approach. Yahong Han, Yi Yang 0001, Mingli Song, Shuicheng Yan, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Graph-guided sparse reconstruction for region taggingabstractMany of contextual correlations co-exist within the segmented regions among images, like the visual context and semantic context. The appropriate integration and utilization of such contexts are very important to boost the performance of region tagging. Inspired by the recent advances of sparse reconstruction methods, this paper proposes an approach, called Graph-Guided Sparse Reconstruction for Region Tagging (G2SRRT). The G2SRRT consists of two steps: sparse reconstruction for testing regions and tag propagation from training regions to testing regions. In G2SRRT, graph is conducted to flexibly model the contextual correlations among regions. To integrate the graph structure learned from training regions into the sparse reconstruction, we define a Graph-Guided Fusion (G2F) penalty over the graph to encourage the sparsity of differences between two reconstruction coefficients, which corresponds to the linked regions in the graph. Guided by this G2F penalty, the highly correlated regions tend to be jointly selected for the reconstruction, which results in a better performance of region tagging. Experiments on three open benchmark image datasets demonstrate the effectiveness of the proposed algorithm. Yahong Han, Fei Wu 0001, Jian Shao 0001, Qi Tian 0001, Yueting Zhuang |
CVPR | 1 |
| 2012 | Correlated attribute transfer with multi-task graph-guided fusionabstractDue to the describable or human-nameable nature of visual attributes, the attribute-based methods have been receiving much attentions in recent years in many applications. The advantages of the utilization of visual attributes are that they can be composed to create descriptions at various levels of specificity or they can be learned once and then applied to recognize new objects or categories. Therefore, attribute prediction becomes an essential problem to boost image understanding. This paper proposes an approach for correlated attribute transfer from a well-defined source image set to an uncontrolled target image set for attribute prediction. We call it correlated attribute transfer with multi-task graph-guided fusion (CAT-MtG2F). The novelty of CAT-MtG2F is to encourage highly correlated attributes to share a common set of relevant low-level features and transfer the learned common structure from the source image set to the target image set. The experiments show that the proposed CAT-MtG2F achieves better performance in attribute prediction. Yahong Han, Fei Wu 0001, Qi Tian 0001, Yueting Zhuang, Jiebo Luo 0001 |
ACM Multimedia | 1 |
| 2012 | Sparse Unsupervised Dimensionality Reduction for Multiple View DataabstractDifferent kinds of high-dimensional visual features can be extracted from a single image. Images can thus be treated as multiple view data when taking each type of extracted high-dimensional visual feature as a particular understanding of images. In this paper, we propose a framework of sparse unsupervised dimensionality reduction for multiple view data. The goal of our framework is to find a low-dimensional optimal consensus representation from multiple heterogeneous features by multiview learning. In this framework, we first learn low-dimensional patterns individually from each view, considering the specific statistical property of each view. We construct a low-dimensional optimal consensus representation from those learned patterns, the goal of which is to leverage the complementary nature of the multiple views. We formulate the construction of the low-dimensional consensus representation to approximate the matrix of patterns by means of a low-dimensional consensus base matrix and a loading matrix. To select the most discriminative features for the spectral embedding of multiple views, we propose to add anl1-norm into the loading matrix's columns and impose orthogonal constraints on the base matrix. We develop a new alternating algorithm, i.e., spectral sparse multiview embedding, to efficiently obtain the solution. Each row of the loading matrix encodes structured information corresponding to multiple patterns. In order to gain flexibility in sharing information across subsets of the views, we impose a novel structured sparsity-inducing norm penalty on the loading matrix's rows. This penalty makes the loading coefficients adaptively load shared information across subsets of the learned patterns. We call this method structured sparse multiview dimensionality reduction. Experiments on a toy benchmark image data set and two real-world Web image data sets demonstrate the effectiveness of the proposed algorithms. Yahong Han, Fei Wu 0001, Dacheng Tao, Jian Shao 0001, Yueting Zhuang, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2012 | Image Annotation by Input-Output Structural Grouping SparsityabstractAutomatic image annotation (AIA) is very important to image retrieval and image understanding. Two key issues in AIA are explored in detail in this paper, i.e., structured visual feature selection and the implementation of hierarchical correlated structures among multiple tags to boost the performance of image annotation. This paper simultaneously introduces an input and output structural grouping sparsity into a regularized regression model for image annotation. For input high-dimensional heterogeneous features such as color, texture, and shape, different kinds (groups) of features have different intrinsic discriminative power for the recognition of certain concepts. The proposed structured feature selection by structural grouping sparsity can be used not only to select group-of-features but also to conduct within-group selection. Hierarchical correlations among output labels are well represented by a tree structure, and therefore, the proposed tree-structured grouping sparsity can be used to boost the performance of multitag image annotation. In order to efficiently solve the proposed regression model, we relax the solving process as a framework of the bilayer regression model for multilabel boosting by the selection of heterogeneous features with structural grouping sparsity (Bi-MtBGS). The first-layer regression is to select the discriminative features for each label. The aim of the second-layer regression is to refine the feature selection model learned from the first layer, which can be taken as a multilabel boosting process. Extensive experiments on public benchmark image data sets and real-world image data sets demonstrate that the proposed approach has better performance of multitag image annotation and leads to a quite interpretable model for image understanding. Yahong Han, Fei Wu 0001, Qi Tian 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 1 |
| 2011 | Stable multi-label boosting for image annotation with structural feature selection
Yueting Zhuang, Yahong Han, Fei Wu 0001, JiaCheng Yang |
Sci. China Inf. Sci. | 2 |
| 2010 | Multi-Task Sparse Discriminant Analysis (MtSDA) with Overlapping CategoriesabstractMulti-task learning aims at combining information across tasks to boost prediction performance, especially when the number of training samples is small and the number of predictors is very large. In this paper, we first extend the Sparse Discriminate Analysis (SDA) of Clemmensen et al.. We call this Multi-task Sparse Discriminate Analysis (MtSDA). MtSDA formulates multi-label prediction as a quadratic optimization problem whereas SDA obtains single labels via a nearest class mean rule. Second, we propose a class of equicorrelation matrices to use in MtSDA which includes the identity matrix. MtSDA with both matrices are compared with singletask learning (SVM and LDA+SVM) and multi-task learning (HSML). The comparisons are made on real data sets in terms of AUC and F-measure. The data results show that MtSDA outperforms other methods substantially almost all the time and in some cases MtSDA with the equicorrelation matrix substantially outperforms MtSDA with identity matrix. Yahong Han, Fei Wu 0001, Jinzhu Jia, Yueting Zhuang, Bin Yu 0001 |
AAAI | 1 |
| 2010 | Multi-label boosting for image annotation by structural grouping sparsityabstractWe can obtain high-dimensional heterogenous features from real-world images to describe their various aspects of visual characteristics, such as color, texture and shape etc.Different kinds of heterogenous features have different intrinsic discriminative power for image understanding. The selection of groups of discriminative features for certain semantics is hence crucial to make the image understanding more interpretable. This paper formulates the multi-label image annotation as a regression model with a regularized penalty. We call it Multi-label Boosting by the selection of heterogeneous features with structural Grouping Sparsity (MtBGS). MtBGS induces a (structural ) sparse selection model to identify subgroups of homogenous features for predicting a certain label. Moreover, the correlations among multiple tags are utilized in MtBGS to boost the performance of multi-label annotation. Extensive experiments on public image datasets show that the proposed approach has better multi-label image annotation performance and leads to a quite interpretable model for image understanding. Fei Wu 0001, Yahong Han, Qi Tian 0001, Yueting Zhuang |
ACM Multimedia | 2 |
| 2010 | Multiple Hypergraph Clustering of Web Images by MiningWord2Image Correlations
Fei Wu 0001, Yahong Han, Yueting Zhuang |
J. Comput. Sci. Technol. | 2 |
| 2010 | Multiple hypergraph ranking for video concept detectionabstractThis paper tackles the problem of video concept detection using the multi-modality fusion method. Motivated by multi-view learning algorithms, multi-modality features of videos can be represented by multiple graphs. And the graph-based semi-supervised learning methods can be extended to multiple graphs to predict the semantic labels for unlabeled video data. However, traditional graphs represent only homogeneous pairwise linking relations, and therefore the high-order correlations inherent in videos, such as high-order visual similarities, are ignored. In this paper we represent heterogeneous features by multiple hypergraphs and then the high-order correlated samples can be associated with hyperedges. Furthermore, the multi-hypergraph ranking (MHR) algorithm is proposed by defining Markov random walk on each hypergraph and then forming the mixture Markov chains so as to perform transductive learning in multiple hypergraphs. In experiments on the TRECVID dataset, a triple-hypergraph consisting of visual, textual features and multiple labeled tags is constructed to predict concept labels for unlabeled video shots by the MHR framework. Experimental results show that our approach is effective. Yahong Han, Jian Shao 0001, Fei Wu 0001, Baogang Wei |
J. Zhejiang Univ. Sci. C | 1 |
| 2010 | Multi-Label Transfer Learning With Sparse RepresentationabstractDue to the visually polysemous barrier, videos and images may be annotated by multiple tags. Discovering the correlations among different tags can significantly help predicting precise labels for videos and images. Many of recent studies toward multi-label learning construct a linear subspace embedding with encoded multi-label information, such that data points sharing many common labels tend to be close to each other in the embedded subspace. Motivated by the advances of compressive sensing research, a sparse representation that selects a compact subset to describe the input data can be more discriminative. In this paper, we propose a sparse multi-label learning method to circumvent the visually polysemous barrier of multiple tags. Our approach learns a multi-label encoded sparse linear embedding space from a related dataset, and maps the target data into the learned new representation space to achieve better annotation performance. Instead of using l1-norm penalty (lasso) to induce sparse representation, we propose to formulate the multi-label learning as a penalized least squares optimization problem with elastic-net penalty. By casting the video concept detection and image annotation tasks into a sparse multi-label transfer learning framework in this paper, ridge regression, lasso, elastic net, and the multi-label extended sparse discriminant analysis methods are, respectively, well explored and compared. Yahong Han, Fei Wu 0001, Yueting Zhuang, Xiaofei He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | s-HITSc: an improved model and algorithm for topic distillation on the Web
Zhuoming Xu, Xiao Cao, Yisheng Dong, Yahong Han |
Soft Comput. | 4 |
| 2006 | s-HITSc: an improved model and algorithm for topic distillation on the Web
Zhuoming Xu, Xiao Cao, Yisheng Dong, Yahong Han |
Soft Comput. | 4 |