VLDB 2026 Research / reviewers in the wild / expert
Aming Wu
dblp:219/1674
· DBLP profile ↗
46ranked-venue papers
23as first author
39since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 13 first-author · 25 since 2021Artificial intelligence and machine learning · 30 · 17 first-author · 26 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Separating feature extraction to enhance fine-grained knowledge for cross-domain federated time series forecasting
Jianji Ren, Yun Xin, Aming Wu, Shan Zhao 0009, Yanan Li 0004 |
Expert Syst. Appl. | 5 |
| 2026 | Simulating Distribution Dynamics: Liquid Temporal Feature Evolution for Single-Domain Generalized Object DetectionabstractIn this paper, we focus on Single-Domain Generalized Object Detection (Single-DGOD), aiming to transfer a detector trained on one source domain to multiple unknown domains. Existing methods for Single-DGOD typically rely on discrete data augmentation or static perturbation methods to expand data diversity, thereby mitigating the lack of access to target domain data. However, in real-world scenarios such as changes in weather or lighting conditions, domain shifts often occur continuously and gradually. Discrete augmentations and static perturbations fail to effectively capture the dynamic variation of feature distributions, thereby limiting the model's ability to perceive fine-grained cross-domain differences. To this end, we propose a new method, i.e., Liquid Temporal Feature Evolution, which simulates the progressive evolution of features from the source domain to simulated latent distributions by incorporating temporal modeling and liquid neural network–driven parameter adjustment. Specifically, we introduce controllable Gaussian noise injection and multi-scale Gaussian blurring to simulate initial feature perturbations, followed by temporal modeling and a liquid parameter adjustment mechanism to generate adaptive modulation parameters, enabling a smooth and continuous adaptation across domains. By capturing progressive cross-domain feature evolution and dynamically regulating adaptation paths, our method bridges the source-unknown domain distribution gap, significantly boosting generalization and robustness to unseen shifts. Significant performance improvements on the Diverse Weather dataset and Real-to-Art benchmark demonstrate the superiority of our method. Yang Li 0251, Aming Wu, Yahong Han |
AAAI | 3 |
| 2026 | Fourier-KAN: Feature Distribution Decomposition and Recombination for Unknown-Domain Object DetectionabstractSingle-domain Generalized Object Detection (Single-DGOD) is recently proposed, aiming to transfer a detector to multiple unknown domains never seen during training. For this task, the challenge mainly lies in how to utilize the single data distribution from the source domain to generalize across multiple unknown domains with diverse data distributions. Accordingly, the challenge could be addressed by expanding the data distribution of the source domain. In this paper, we propose feature recombination from a frequency perspective to generate a series of recombined features that exhibit diversity in style and rich variation in content features. Specifically, we propose a new method, Fourier-KAN Feature Recombination, which utilizes the Fast Fourier Transform (FFT) to decompose features into amplitude and phase components. Then we apply the Kolmogorov-Arnold theorem to further decompose these components into linear combinations of multiple base distributions. Finally, through multi-level recombination, we generate a series of recombined features with diverse distributions, effectively emulating deep cross-domain variations in feature levels and strengthening the model's generalization ability to unknown domains. Our method demonstrates strong adaptability to both two-stage and single-stage detection frameworks. Experimental results show that on the Diverse Weather and Real-to-Art benchmarks, our approach not only achieves outstanding detection accuracy but also significantly enhances the model's generalization ability, all while maintaining excellent real-time performance. Our code is available at https://github.com/2490o/Fourier-KAN. Yang Li 0251, Aming Wu, Yahong Han |
IEEE Trans. Image Process. | 3 |
| 2025 | Reasoning Mamba: Hypergraph-Guided Region Relation Calculating for Weakly Supervised Affordance GroundingabstractThis paper pays attention to Weakly Supervised Affordance Grounding (WSAG) task that aims to train model to identify affordance regions using human-object interaction images and egocentric images without the need for costly pixel-level annotations. Most existing methods usually consider the affordance regions to be isolated and directly employ class activation maps to conduct localization, ignoring the relationships with other object components and weakening the performance. For example, a cup’s handle is combined with its body to achieve the pouring ability. Obviously, capturing the region relationships is beneficial for improving the localization accuracy of affordance regions. To this end, we first explore exploiting hypergraph to discover these relations and propose a Reasoning Mamba (R-Mamba) framework. We first extract feature embeddings from exocentric and egocentric images to construct the hypergraphs consisting of multiple vertices and hyperedges, which capture the in-context local region relationships between different visual components. Subsequently, we design a Hypergraph-Guided State Space (HSS) block to reorganize these local relationships from the global perspective. By this mechanism, the model could leverage the captured relationships to improve the localization accuracy of affordance regions. Extensive experiments and visualization analyses demonstrate the superiority of our method. Aming Wu, Muli Yang, Yukuan Min, Yihang Zhu, Cheng Deng 0002 |
CVPR | 2 |
| 2025 | Percept, Memory, and Imagine: World Feature Simulating for Open-Domain Unknown Object DetectionabstractTo accelerate the safe deployment of object detectors, we focus on reducing the impact of both covariate and semantic shifts. And we consider a realistic yet challenging scenario, namely Open-Domain Unknown Object Detection (ODU-OD), which aims to detect unknown objects in unseen target domains without accessing any auxiliary data. Towards ODU-OD, it is feasible to learn a robust discriminative boundary by synthesizing virtual features. Generally, perception, memory, and imagination are three essential capacities for human beings. Through multi-level perception and rich memory about known objects, the characteristics of unknown objects can be imagined sufficiently, enhancing the ability of discriminating known from unknown objects. Inspired by this idea, an approach of World Feature Simulation (WFS) is proposed, mainly consisting of a multi-level perception, memory recorder, and unknown-feature generator. Specifically, after extracting the features of the input, we separately employ a Mamba and Graph Network to obtain the global-level and connective-level representations. Next, a codebook containing multiple learnable codewords is defined to preserve fragmented memory of known objects. Meanwhile, we perform a modulated operation on the memory to form the imagination bank involving unknown characteristics. Finally, to alleviate the impact of lacking supervision data, based on the multi-level representation and imagination bank, a dedicated unknown-feature generator is designed to recurrently synthesize outlier features deviating from in-distribution (ID) objects. The significant performance gains on four different detection tasks demonstrate the superiorities of our method. The code will be released at https://github.com/AmingWu/WFS. Aming Wu, Cheng Deng 0002 |
CVPR | 1 |
| 2025 | Style Evolving along Chain-of-Thought for Unknown-Domain Object DetectionabstractRecently, a task of Single-Domain Generalized Object Detection (Single-DGOD) is proposed, aiming to generalize a detector to multiple unknown domains never seen before during training. Due to the unavailability of target-domain data, some methods leverage the multimodal capabilities of vision-language models, using textual prompts to estimate cross-domain information, enhancing the model’s generalization capability. These methods typically use a single textual prompt, referred to as the one-step prompt method. However, when dealing with complex styles, such as the combination of rain and night, we observe that the performance of the one-step prompt method tends to be relatively weak. The reason may be that many scenes incorporate a single style and a combination of multiple styles. The one- step prompt method may not effectively synthesize combined information involving various styles. To address this limitation, we propose a new method, i.e., Style Evolving along Chain-of-Thought, which aims to progressively integrate and expand style information along the chain of thought, enabling the continual evolution of styles. Specifically, by progressively refining style descriptions and guiding the diverse evolution of styles, this method enhances the simulation of various style characteristics, enabling the model to learn and adapt to subtle differences more effectively. Additionally, it exposes the model to a broader range of style features with different data distributions, thereby enhancing its generalization capability in unseen domains. The significant performance gains over five adverse-weather scenarios and the Real to Art benchmark demonstrate the superiorities of our method. Our code is available at https://github.com/ZZ2490/SE-COT. Aming Wu, Yahong Han |
CVPR | 2 |
| 2025 | Continual Adaptation: Environment-Conditional Parameter Generation for Object Detection in Dynamic ScenariosabstractIn practice, environments constantly change over time and space, posing significant challenges for object detectors trained based on a closed-set assumption, i.e., training and test data share the same distribution. To this end, continual test-time adaptation has attracted much attention, aiming to improve detectors' generalization by fine-tuning a few specific parameters, e.g., BatchNorm layers. However, based on a small number of test images, fine-tuning certain parameters may affect the representation ability of other fixed parameters, leading to performance degradation. Instead, we explore a new mechanism, i.e., converting the fine-tuning process to a specific-parameter generation. Particularly, we first design a dual-path LoRA-based domain-aware adapter that disentangles features into domain-invariant and domain-specific components, enabling efficient adaptation. Additionally, a conditional diffusion-based parameter generation mechanism is presented to synthesize the adapter's parameters based on the current environment, preventing the optimization from getting stuck in local optima. Finally, we propose a class-centered optimal transport alignment method to mitigate catastrophic forgetting. Extensive experiments conducted on various continuous domain adaptive object detection tasks demonstrate the effectiveness. Meanwhile, visualization results show that the representation extracted by the generated parameters can capture more object-related information and strengthen the generalization ability. Deng Li 0003, Aming Wu, Yang Li 0251, Yaowei Wang 0001, Yahong Han |
ICCV | 2 |
| 2025 | Vision-Language Interactive Relation Mining for Open-Vocabulary Scene Graph Generation
Yukuan Min, Muli Yang, Aming Wu, Cheng Deng 0002 |
ICCV | 5 |
| 2025 | VGMamba: Attribute-to-Location Clue Reasoning for Quantity-Agnostic 3D Visual Grounding
Yihang Zhu, Aming Wu, Cheng Deng 0002 |
ICCV | 4 |
| 2025 | CFD: Learning Generalized Molecular Representation via Concept-Enhanced Feedback DisentanglementabstractTo accelerate biochemical research, e.g., drug and protein discovery, molecular representation learning (MRL) has attracted much attention. However, most existing methods follow the closed-set assumption that training and testing data share identical distribution, which limits their generalization abilities in out-of-distribution (OOD) cases. In this paper, we explore designing a new disentangled mechanism for learning generalized molecular representation that exhibits robustness against distribution shifts. And an approach of Concept-Enhanced Feedback Disentanglement (CFD) is proposed, whose goal is to exploit the feedback mechanism to learn distribution-agnostic representation. Specifically, we first propose two dedicated variational encoders to separately decompose distribution-agnostic and spurious features. Then, a set of molecule-aware concepts are tapped to focus on invariant substructure characteristics. By fusing these concepts into the disentangled distribution-agnostic features, the generalization ability of the learned molecular representation could be further enhanced. Next, we execute iteratively the disentangled operations based on a feedback received from the previous output. Finally, based on the outputs of multiple feedback iterations, we construct a self-supervised objective to promote the variational encoders to possess the disentangled capability. In the experiments, our method is verified on multiple real-world molecular datasets. The significant performance gains over state-of-the-art baselines demonstrate that our method can effectively disentangle generalized molecular representation in the presence of various distribution shifts. The source code will be released at https://github.com/AmingWu/MoleculeCFD. Aming Wu |
ICLR | 1 |
| 2025 | Novel Class Discovery for Point Cloud Segmentation via Joint Learning of Causal Representation and ReasoningabstractIn this paper, we focus on Novel Class Discovery for Point Cloud Segmentation (3D-NCD), aiming to learn a model that can segment unlabeled (novel) 3D classes using only the supervision from labeled (base) 3D classes. The key to this task is to setup the exact correlations between the point representations and their base class labels, as well as the representation correlations between the points from base and novel classes. A coarse or statistical correlation learning may lead to the confusion in novel class inference. lf we impose a causal relationship as a strong correlated constraint upon the learning process, the essential point cloud representations that accurately correspond to the classes should be uncovered. To this end, we introduce a structural causal model (SCM) to re-formalize the 3D-NCD problem and propose a new method, i.e., Joint Learning of Causal Representation and Reasoning. Specifically, we first analyze hidden confounders in the base class representations and the causal relationships between the base and novel classes through SCM. We devise a causal representation prototype that eliminates confounders to capture the causal representations of base classes. A graph structure is then used to model the causal relationships between the base classes' causal representation prototypes and the novel class prototypes, enabling causal reasoning from base to novel classes. Extensive experiments and visualization results on 3D and 2D NCD semantic segmentation demonstrate the superiorities of our method. Yang Li 0251, Aming Wu, Yahong Han |
NeurIPS | 2 |
| 2025 | Prototype-guided cross-task knowledge distillationabstractRecently, large-scale pretrained models have revealed their benefits in various tasks. However, due to the enormous computation complexity and storage demands, it is challenging to apply large-scale models to real scenarios. Existing knowledge distillation methods require mainly the teacher model and the student model to share the same label space, which restricts their application in real scenarios. To alleviate the constraint of different label spaces, we propose a prototype-guided cross-task knowledge distillation (ProC-KD) method to migrate the intrinsic local-level object knowledge of the teacher network to various task scenarios. First, to better learn the generalized knowledge in cross-task scenarios, we present a prototype learning module to learn the invariant intrinsic local representation of objects from the teacher network. Second, for diverse downstream tasks, a task-adaptive feature augmentation module is proposed to enhance the student network features with the learned generalization prototype representations and guide the learning of the student network to improve its generalization ability. Experimental results on various visual tasks demonstrate the effectiveness of our approach for cross-task knowledge distillation scenarios. Deng Li 0003, Aming Wu, Yahong Han |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | Towards OOD Object Detection With Unknown-Concept Guided Feature DiffusionabstractIn general, learning plentiful knowledge corresponding to known objects is an important ability for humans. The unknown objects could be assumed to depart from the familiar knowledge. Inspired by this idea, we explore leveraging the extracted knowledge to reason a set of unknown concepts. And they could be used to address unsupervised out-of-distribution object detection (OOD-OD) that aims to detect unseen OOD objects without accessing any auxiliary OOD data during training. To this end, we propose a new approach, i.e., Unknown-Concept Guided Feature Diffusion (UCFD), including an object-related knowledge extractor and an unknown-concept guided diffusor for synthesizing virtual OOD features. Specifically, we define multiple learnable codewords to capture object-relevant visual knowledge from all object categories. To avoid the detection performance degradation of the in-distribution (ID) objects, these codewords are utilized to enhance object features. Next, an unknown-concept pool is constructed by mixing up these extracted codewords. Finally, to reduce the impact of lacking OOD data for supervision, we design an unknown-concept guided diffusor, which leverages the sampled unknown concepts from the pool to guide the reverse process to generate expected OOD features that deviate from the familiar knowledge. The significant performance gains on three different tasks demonstrate the superiorities of our method. Meanwhile, extensive visualization results show that our method could synthesize effective virtual OOD features. Aming Wu, Cheng Deng 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Progressive Invariant Causal Feature Learning for Single Domain GeneralizationabstractSingle domain generalization (SDG) aims to transfer models trained on a single source domain to multiple unseen target domains while against the unknown domain shifts. The main challenge lies in learning the domain-invariant features to mitigate the domain shift impact. To address this challenge, we reconsider SDG from a causal perspective to capture the domain-invariant features accurately. Specifically, we present a Progressive Invariant Causal Feature Learning (PICF) method that leverages front-door adjustment to gradually obtain the invariant causal features for SDG. First, we introduce a foreground feature filter, which removes object-irrelevant confounders in a cyclical manner to extract the object-related causal features. Subsequently, to further enhance the causal feature invariance, we propose to train with augmented causal features by combining them with randomly-sampled styles from the object-irrelevant feature distribution boundary. As a result, our model bridges the gap between one seen domain and multiple unseen ones by capturing the invariant causal features, which largely enhances the model's generalization ability in SDG. In experiments, our method can be plugged into multiple state-of-the-art methods, and the significant performance improvements on multiple datasets demonstrate the superiority of our method. In particular, on the PACS dataset, our method achieves an accuracy improvement of 4.7%. Muli Yang, Aming Wu, Cheng Deng 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | Memory-Enhanced Confidence Calibration for Class-Incremental Unsupervised Domain AdaptationabstractIn this paper, we focus on Class-Incremental Unsupervised Domain Adaptation (CI-UDA), where the labeled source domain already includes all classes, and the classes in the unlabeled target domain emerge sequentially over time. This task involves addressing two main challenges. The first is the domain gap between the labeled source data and the unlabeled target data, which leads to weak generalization performance. The second is the inconsistency between the source and target category spaces at each time step, which causes catastrophic forgetting during the testing stage. Previous methods focus solely on the alignment of similar samples from different domains, which overlooks the underlying causes of the domain gap/class distribution difference. To tackle the issue, we rethink this task from a causal perspective for the first time. We first build a structural causal graph to describe the CI-UDA problem. Based on the causal graph, we present Memory-Enhanced Confidence Calibration (MECC), which aims to improve confidence in the predicted results. In particular, we argue that the domain discrepancy caused by the different styles is prone to make the model produce less confident predictions and thus weakens the generalization and continual learning abilities. To this end, we first explore using the gram matrix to generate source-style target data, which is combined with the original data to jointly train the model and thereby reduce the domain-shift impact. Second, we utilize the model of the previous time step to select corresponding samples that are used to build a memory bank, which is instrumental in alleviating catastrophic forgetting. Extensive experimental results on multiple datasets demonstrate the superiority of our method. Jiaping Yu, Muli Yang, Aming Wu, Cheng Deng 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Prompt-Driven Dynamic Object-Centric Learning for Single Domain GeneralizationabstractSingle-domain generalization aims to learn a model from single source domain data attaining generalized performance on other unseen target domains. Existing works primarily focus on improving the generalization ability of static networks. However, static networks are unable to dynamically adapt to the diverse variations in different image scenes, leading to limited generalization capability. Different scenes exhibit varying levels of complexity, and the complexity of images further varies significantly in crossdomain scenarios. In this paper, we propose a dynamic object-centric perception network based on prompt learning, aiming to adapt to the variations in image complexity. Specifically, we propose an object-centric gating module based on prompt learning to focus attention on the object-centric features guided by the various scene prompts. Then, with the object-centric gating masks, the dynamic selective module dynamically selects highly correlated feature regions in both spatial and channel dimensions enabling the model to adaptively perceive object-centric relevant features, thereby enhancing the generalization capability. Extensive experiments were conducted on single-domain generalization tasks in image classification and object detection. The experimental results demonstrate that our approach outperforms state-of-the-art methods, which validates the effectiveness and versatility of our proposed method. Deng Li 0003, Aming Wu, Yaowei Wang 0001, Yahong Han |
CVPR | 2 |
| 2024 | Modulated Phase Diffusor: Content-Oriented Feature Synthesis for Detecting Unknown ObjectsabstractTo promote the safe deployment of object detectors, a task of unsupervised out-of-distribution object detection (OOD-OD) is recently proposed, aiming to detect unknown objects during training without reliance on any auxiliary OOD data. To alleviate the impact of lacking OOD data, for this task, one feasible solution is to exploit the known in-distribution (ID) data to synthesize proper OOD information for supervision, which strengthens detectors' discrimination. From the frequency perspective, since the phase generally reflects the content of the input, in this paper, we explore leveraging the phase of ID features to generate expected OOD features involving different content. And a method of Modulated Phase Diffusion (MPD) is proposed, containing a shared forward and two different reverse processes. Specifically, after calculating the phase of the extracted features, to prevent the rapid loss of content in the phase, the forward process gradually performs Gaussian Average on the phase instead of adding noise. The averaged phase and original amplitude are combined to obtain the features taken as the input of the reverse process. Next, one OOD branch is defined to synthesize virtual OOD features by continually enlarging the content discrepancy between the OOD features and original ones. Meanwhile, another modulated branch is designed to generate augmented features owning a similar phase as the original features by scaling and shifting the OOD branch. Both original and augmented features are used for training, enhancing the discrimination. Experimental results on OOD-OD, incremental object detection, and open-set object detection demonstrate the superiorities of our method. The source code will be released at https://github.com/AmingWu/MPD. Aming Wu |
ICLR | 1 |
| 2024 | Pre-LogMGAE: Identification of Log Anomalies Using a Pre-Trained Masked Graph AutoencoderabstractLog-based anomaly detection in software systems is becoming increasingly crucial for monitoring network operations and ensuring system security. Deep learning-based methods are widely used for large-scale log anomaly detection due to their capacity to learn complex features. However, current research predominantly treats original logs as simple sequences, ignoring their complex structure and dynamic dependency relationships. Additionally, these methods often rely on extensive labeled data or domain-specific vectors to represent logs for model training, which can be labor-intensive to label manually and ineffective across various domains within a system. To address these challenges, this paper proposes Pre-LogMGAE, a universal masked graph autoencoder (GAE) framework with contrastive learning for self-supervised pre-training for log anomaly detection. In contrast to graph or link reconstruction, Pre-LogMGAE focuses on node feature reconstruction using a masking strategy to reduce the impact of excessive redundant information. Furthermore, we introduce Graph Attention Networks (GAT) with the Gated Recurrent Unit (GRU) to incorporate sequence modeling, allowing for capturing long-term and short-term dependencies in log events. We include contrastive learning objectives in finetuning to extract diverse features and enhance the algorithm's robustness. Through an extensive evaluation of three real-world datasets and specific case studies with configuration error, Pre-LogMGAE demonstrates superior performance compared to the six baselines, including PCA, IM, DeepLog, LogRobust, LogBERT, and DeepTraLog. This superiority is evident in terms of precision, recall, F1 score, and time efficiency, highlighting Pre-LogMGAE's stability and reliability in anomaly detection. The study aims to improve anomaly detection capabilities in multi-source system logs, offering innovative technical support to enhance system security and reliability. Aming Wu, Young-Woo Kwon 0001 |
SRDS | 1 |
| 2024 | TIB: Detecting Unknown Objects via Two-Stream Information BottleneckabstractDetecting diverse objects, including ones never-seen-before during training, is critical for the safe application of object detectors. To this end, a task of unsupervised out-of-distribution object detection (OOD-OD) is proposed to detect unknown objects without the reliance on an auxiliary dataset. For this task, it is important to reduce the impact of lacking unknown data for supervision and leverage in-distribution (ID) data to improve the model's discrimination. In this paper, we propose a method of Two-Stream Information Bottleneck (TIB), consisting of a standard IB and a dedicated Reverse Information Bottleneck (RIB). Specifically, after extracting the features of an ID image, we first define a standard IB network to disentangle instance representations that are beneficial for localizing and recognizing objects. Meanwhile, we present RIB to obtain simulative OOD features to alleviate the impact of lacking unknown data. Different from standard IB aiming to extract task-relevant compact representations, RIB is to obtain task-irrelevant representations by reversing the optimization objective of the standard IB. Next, to further enhance the discrimination, a mixture of information bottlenecks is designed to sufficiently capture object-related information. Experimental results on OOD-OD, open-vocabulary object detection, incremental object detection, and open-set object detection show the superiorities of our method. Aming Wu, Cheng Deng 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Unsupervised Out-of-Distribution Object Detection via PCA-Driven Dynamic Prototype EnhancementabstractTo promote the application of object detectors in real scenes, out-of-distribution object detection (OOD-OD) is proposed to distinguish whether detected objects belong to the ones that are unseen during training or not. One of the key challenges is that detectors lack unknown data for supervision, and as a result, can produce overconfident detection results on OOD data. Thus, this task requires to synthesize OOD data for training, which achieves the goal of enhancing the ability of localizing and discriminating OOD objects. In this paper, we propose a novel method, i.e., PCA-Driven dynamic prototype enhancement, to explore exploiting Principal Component Analysis (PCA) to extract simulative OOD data for training and obtain dynamic prototypes that are related to the current input and are helpful for boosting the discrimination ability. Concretely, the last few principal components of the backbone features are utilized to calculate an OOD map that involves plentiful information that deviates from the correlation distribution of the input. The OOD map is further used to extract simulative OOD data for training, which alleviates the impact of lacking unknown data. Besides, for in-distribution (ID) data, the category-level semantic information of objects between the backbone features and the high-level features should be kept consistent. To this end, we utilize the residual principal components to extract dynamic prototypes that reflect the semantic information of the current backbone features. Next, we define a contrastive loss to leverage these prototypes to enlarge the semantic gap between the simulative OOD data and the features from the residual principal components, which improves the ability of discriminating OOD objects. In the experiments, we separately verify our method on OOD-OD and incremental object detection. The significant performance gains demonstrate the superiorities of our method. Aming Wu, Cheng Deng 0002, Wei Liu 0005 |
IEEE Trans. Image Process. | 1 |
| 2024 | Prototype-Decomposed Knowledge Distillation for Learning Generalized Federated RepresentationabstractFederated learning (FL) enables distributed clients to collaboratively learn a global model, suggesting its potential for use in improving data privacy in machine learning. However, although FL has made many advances, its performance usually suffers from degradation due to the impact of domain shift when the trained models are applied to unseen domains. To enhance the model's generalization ability, we focus on solving federated domain generalization, which aims to properly generalize a federated model trained based on multiple source domains belonging to different distributions to an unseen target domain. A novel approach, namely Prototype-Decomposed Knowledge Distillation (PDKD), is proposed herein. Concretely, we first aggregate the local class prototypes that are learned from different clients. Subsequently, Singular Value Decomposition (SVD) is employed to decompose the local prototypes to obtain discriminative and generalized global prototypes that contain rich category-related information. Finally, the global prototypes are sent back to all clients. We exploit knowledge distillation to encourage local client models to distill generalized knowledge from the global prototypes, which boosts the generalization ability. Extensive experiments on multiple datasets demonstrate the effectiveness of our method. In particular, when implemented on the Office dataset, our method outperforms FedAvg by around 13.5%, which shows that our method is instrumental in ameliorating the generalization ability of federated models. Aming Wu, Jiaping Yu, Cheng Deng 0002 |
IEEE Trans. Multim. | 1 |
| 2023 | AD-TIN: Edge Anomaly Detection for Temporal Interaction Networks using Multi-representation AttentionabstractAnomaly detection in temporal interaction networks (TINs) has become critical in network security, digital finance, and social networks. While recent studies based on Graph Neural Networks (GNNs) have yielded promising results, the existing methods are still limited by insufficient labels and noisy data, often ignoring the information filtering for unrelated user interactions. Therefore, this paper proposes a dynamic edge anomaly detection framework, AD-TIN, to address these challenges based on a multi-representation attention mechanism. It encodes graph structural information using a network information propagation module with neighbor sampling and graph diffusion. Furthermore, the network update module combines past node states with current structural features to capture the temporal information in potential user relationships, effectively mitigating the impact of noisy data. Extensive experiments on three real-world datasets demonstrate the robustness and efficacy of AD-TIN in addressing noise and unrelated interactions for edge anomaly detection. Aming Wu, Young-Woo Kwon 0001 |
ASONAM | 1 |
| 2023 | Discriminating Known from Unknown Objects via Structure-Enhanced Recurrent Variational AutoEncoderabstractDiscriminating known from unknown objects is an important essential ability for human beings. To simulate this ability, a task of unsupervised out-of-distribution object detection (OOD-OD) is proposed to detect the objects that are never-seen-before during model training, which is beneficial for promoting the safe deployment of object detectors. Due to lacking unknown data for supervision, for this task, the main challenge lies in how to leverage the known in-distribution (ID) data to improve the detector's discrimination ability. In this paper, we first propose a method of Structure-Enhanced Recurrent Variational AutoEncoder (SR-VAE), which mainly consists of two dedicated recurrent VAE branches. Specifically, to boost the performance of object localization, we explore utilizing the classical Laplacian of Gaussian (LoG) operator to enhance the structure information in the extracted low-level features. Meanwhile, we design a VAE branch that recurrently generates the augmentation of the classification features to strengthen the discrimination ability of the object classifier. Finally, to alleviate the impact of lacking unknown data, another cycle-consistent conditional VAE branch is proposed to synthesize virtual OOD features that deviate from the distribution of ID features, which improves the capability of distinguishing OOD objects. In the experiments, our method is evaluated on OOD-OD, open-vocabulary detection, and incremental object detection. The significant performance gains over baselines show the superiorities of our method. The code will be released at https://github.com/AmingWu/SR-VAE. Aming Wu, Cheng Deng 0002 |
CVPR | 1 |
| 2023 | Environment-Invariant Curriculum Relation Learning for Fine-Grained Scene Graph GenerationabstractThe scene graph generation (SGG) task is designed to identify the predicates based on the subject-object pairs. However, existing datasets generally include two imbalance cases: one is the class imbalance from the predicted predicates and another is the context imbalance from the given subject-object pairs, which presents significant challenges for SGG. Most existing methods focus on the imbalance of the predicted predicate while ignoring the imbalance of the subject-object pairs, which could not achieve satisfactory results. To address the two imbalance cases, we propose a novel Environment Invariant Curriculum Relation learning (EICR) method, which can be applied in a plug-and-play fashion to existing SGG methods. Concretely, to remove the imbalance of the subject-object pairs, we first construct different distribution environments for the subject-object pairs and learn a model invariant to the environment changes. Then, we construct a class-balanced curriculum learning strategy to balance the different environments to remove the predicate imbalance. Comprehensive experiments conducted on VG and GQA datasets demonstrate that our EICR framework can be taken as a general strategy for various SGG models, and achieve significant improvements. Yukuan Min, Aming Wu, Cheng Deng 0002 |
ICCV | 2 |
| 2023 | Deep Feature Deblurring Diffusion for Detecting Out-of-Distribution ObjectsabstractTo promote the safe application of detectors, a task of unsupervised out-of-distribution object detection (OOD-OD) is recently proposed, whose goal is to detect unseen OOD objects without accessing any auxiliary OOD data. For this task, the challenge mainly lies in how to only leverage the known in-distribution (ID) data to detect OOD objects accurately without affecting the detection of ID objects, which can be framed as the diffusion problem for deep feature synthesis. Accordingly, such challenge could be addressed by the forward and reverse processes in the diffusion model. In this paper, we propose a new approach of Deep Feature Deblurring Diffusion (DFDD), consisting of forward blurring and reverse deblurring processes. Specifically, the forward process gradually performs Gaussian Blur on the extracted features, which is instrumental in retaining sufficient input-relevant information. By this way, the forward process could synthesize virtual OOD features that are close to the classification boundary between ID and OOD objects, which improves the performance of detecting OOD objects. During the reverse process, based on the blurred features, a dedicated deblurring model is designed to continually recover the lost details in the forward process. Both the deblurred features and original features are taken as the input for training, strengthening the discrimination ability. In the experiments, our method is evaluated on OOD-OD, open-set object detection, and incremental object detection. The significant performance gains over baselines demonstrate the superiorities of our method. The source code will be made available at: https://github.com/AmingWu/DFDD-OOD. Aming Wu, Da Chen 0003, Cheng Deng 0002 |
ICCV | 1 |
| 2023 | A Decomposable Causal View of Compositional Zero-Shot LearningabstractComposing and recognizing novel concepts that are combinations of known concepts,i.e., compositional generalization, is one of the greatest power of human intelligence. With the development of artificial intelligence, it becomes increasingly appealing to build a vision system that can generalize to unknown compositions based on restricted known knowledge, which has so far remained a great challenge to our community. In fact, machines can be easily misled by superficial correlations in the data, disregarding the causal patterns that are crucial to generalization. In this paper, we rethink compositional generalization with a causal perspective, upon the context of Compositional Zero-Shot Learning (CZSL). We develop a simple yet strong approach based on our novelDecomposableCausal view (dubbed “DeCa”), by approximating the causal effect with the combination of three easy-to-learn components. Our proposedDeCa11Code is available onhttps://github.com/muliyangm/DeCa.is evaluated on two challenging CZSL benchmarks by recognizing unknown compositions of known concepts. Despite being simple in the design, our approach achieves consistent improvements over state-of-the-art baselines, demonstrating its superiority towards the goal of compositional generalization. Muli Yang, Aming Wu, Cheng Deng 0002 |
IEEE Trans. Multim. | 3 |
| 2022 | Single-Domain Generalized Object Detection in Urban Scene via Cyclic-Disentangled Self-DistillationabstractIn this paper, we are concerned with enhancing the generalization capability of object detectors. And we consider a realistic yet challenging scenario, namely Single-Domain Generalized Object Detection (Single-DGOD), which aims to learn an object detector that performs well on many unseen target domains with only one source domain for training. Towards Single-DGOD, it is important to extract domain-invariant representations (DIR) containing intrinsical object characteristics, which is beneficial for improving the robustness for unseen domains. Thus, we present a method, i.e., cyclic-disentangled self-distillation, to disentangle DIR from domain-specific representations without the supervision of domain-related annotations (e.g., domain labels). Concretely, a cyclic-disentangled module is first proposed to cyclically extract DIR from the input visual features. Through the cyclic operation, the disentangled ability can be promoted without the reliance on domain-related annotations. Then, taking the DIR as the teacher, we design a self-distillation module to further enhance the generalization ability. In the experiments, our method is evaluated in urban-scene object detection. Experimental results of five weather conditions show that our method obtains a significant performance gain over baseline methods. Particularly, for the night-sunny scene, our method outperforms baselines by 3%, which indicates that our method is instrumental in enhancing generalization ability. Data and code are available at https://github.com/AmingWu/Single-DgoD. Aming Wu, Cheng Deng 0002 |
CVPR | 1 |
| 2022 | Divide and Conquer: Compositional Experts for Generalized Novel Class DiscoveryabstractIn response to the explosively-increasing requirement of annotated data, Novel Class Discovery (NCD) has emerged as a promising alternative to automatically recognize unknown classes without any annotation. To this end, a model makes use of a base set to learn basic semantic discriminability that can be transferred to recognize novel classes. Most existing works handle the base and novel sets using separate objectives within a two-stage training paradigm. Despite showing competitive performance on novel classes, they fail to generalize to recognizing samples from both base and novel sets. In this paper, we focus on this generalized setting of NCD (GNCD), and propose to divide and conquer it with two groups of Compositional Experts (ComEx). Each group of experts is designed to characterize the whole dataset in a comprehensive yet complementary fashion. With their union, we can solve GNCD in an efficient end-to-end manner. We further look into the draw-back in current NCD methods, and propose to strengthen ComEx with global-to-local and local-to-local regularization. ComEx11Code: https://github.com/muliyangm/ComEx. is evaluated on four popular benchmarks, showing clear superiority towards the goal of GNCD. Muli Yang, Yuehua Zhu, Jiaping Yu, Aming Wu, Cheng Deng 0002 |
CVPR | 4 |
| 2022 | Comparison of Meta-Heuristic Algorithms for Task Scheduling in Distributed Stream ProcessingabstractWith the emergence of IoT and cloud computing, the demand for big data processing continues to rise. To expedite such big data processing, distributed stream processing systems (DSPS) are commonly used. However, because the rate of incoming messages to DSPS can vary depending on a stream application and execution environments such as like network conditions, it can be challenging to provide the necessary quality of services (QoS). Modern DSPS typically use heuristic or meta-heuristic algorithms to find near-optimal solutions to meet QoS requirements; however, it is still difficult to accomplish multiple QoS goals at once. In this paper, multiple meta-heuristic algorithms are evaluated to determine if they can simultaneously achieve multiple objectives, including response time and system failure. We implemented schedulers using various meta-heuristic algorithms operating within DSPS simulation environments. Then, we executed three stream applications utilizing various scheduling algorithms and demonstrated that meta-heuristic algorithms outperform a conventional algorithm. Dohan Kim 0004, Aming Wu, Young-Woo Kwon 0001 |
PRDC | 2 |
| 2022 | Complementary spatiotemporal network for video question answering
Aming Wu, Yahong Han |
Multim. Syst. | 2 |
| 2022 | Instance-Invariant Domain Adaptive Object Detection Via Progressive DisentanglementabstractMost state-of-the-art methods of object detection suffer from poor generalization ability when the training and test data are from different domains. To address this problem, previous methods mainly explore to align distribution between source and target domains, which may neglect the impact of the domain-specific information existing in the aligned features. Besides, when transferring detection ability across different domains, it is important to extract the instance-level features that are domain-invariant. To this end, we explore to extract instance-invariant features by disentangling the domain-invariant features from the domain-specific features. Particularly, a progressive disentangled mechanism is proposed to decompose domain-invariant and domain-specific features, which consists of a base disentangled layer and a progressive disentangled layer. Then, with the help of Region Proposal Network (RPN), the instance-invariant features are extracted based on the output of the progressive disentangled layer. Finally, to enhance the disentangled ability, we design a detached optimization to train our model in an end-to-end fashion. Experimental results on four domain-shift scenes show our method is separately 2.3, 3.6, 4.0, and 2.0 percent higher than the baseline method. Meanwhile, visualization analysis demonstrates that our model owns well disentangled ability. Aming Wu, Yahong Han, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | CrowdQuake+: Data-driven Earthquake Early Warning via IoT and Deep LearningabstractIn recent years, a low-cost micro-electro-mechanical systems (MEMS) acceleration sensor has been widely used for earthquake early warning (EEW). In our previous work, we introduced a networked earthquake detection system, CrowdQuake with three-hundred smartphones’ acceleration sensors and a deep-learning based earthquake detection model. For one year’s operation, CrowdQuake detected a series of earthquakes and collected various earthquake and non-earthquake data. Based on the successful operation of CrowdQuake, in this paper, we discuss how it can be expanded across the country by addressing the following challenges: (1) sensor deployments for highly dense network, (2) earthquake detection performance using a deep learning model, and (3) high performance and scalable system design for big data processing. The improved system is CrowdQuake+ which can deal with acceleration data sent from 8,000 IoT sensors and detect an earthquake in few seconds using a newly proposed detection model. Moreover, CrowdQuake+ stores all acceleration data sent from sensors and assesses their qualities by calculating noise levels. Then, the collected data are used for deep learning model training, so that its detection performance becomes more accurate. Aming Wu, Jangsoo Lee, Irshad Khan, Young-Woo Kwon 0001 |
IEEE BigData | 1 |
| 2021 | Universal-Prototype Enhancing for Few-Shot Object DetectionabstractFew-shot object detection (FSOD) aims to strengthen the performance of novel object detection with few labeled samples. To alleviate the constraint of few samples, enhancing the generalization ability of learned features for novel objects plays a key role. Thus, the feature learning process of FSOD should focus more on intrinsical object characteristics, which are invariant under different visual changes and therefore are helpful for feature generalization. Unlike previous attempts of the meta-learning paradigm, in this paper, we explore how to enhance object features with intrinsical characteristics that are universal across different object categories. We propose a new prototype, namely universal prototype, that is learned from all object categories. Besides the advantage of characterizing invariant characteristics, the universal prototypes alleviate the impact of unbalanced object categories. After enhancing object features with the universal prototypes, we impose a consistency loss to maximize the agreement between the enhanced features and the original ones, which is beneficial for learning invariant object characteristics. Thus, we develop a new framework of few-shot object detection with universal prototypes (F SODup) that owns the merit of feature generalization towards novel objects. Experimental results on PASCAL VOC and MS COCO show the effectiveness of F SODup. Particularly, for the 1-shot case of VOC Split2, FSODupoutperforms the baseline by 6.8% in terms of mAP. Aming Wu, Yahong Han, Linchao Zhu, Yi Yang 0001 |
ICCV | 1 |
| 2021 | Vector-Decomposed Disentanglement for Domain-Invariant Object DetectionabstractTo improve the generalization of detectors, for domain adaptive object detection (DAOD), recent advances mainly explore aligning feature-level distributions between the source and single-target domain, which may neglect the impact of domain-specific information existing in the aligned features. Towards DAOD, it is important to extract domain-invariant object representations. To this end, in this paper, we try to disentangle domain-invariant representations from domain-specific representations. And we propose a novel disentangled method based on vector decomposition. Firstly, an extractor is devised to separate domain-invariant representations from the input, which are used for extracting object proposals. Secondly, domain-specific representations are introduced as the differences between the input and domain-invariant representations. Through the difference operation, the gap between the domain-specific and domain-invariant representations is enlarged, which promotes domain-invariant representations to contain more domain-irrelevant information. In the experiment, we separately evaluate our method on the single- and compound-target case. For the single-target case, experimental results of four domain-shift scenes show our method obtains a significant performance gain over baseline methods. Moreover, for the compound-target case (i.e., the target is a compound of two different domains without domain labels), our method outperforms baseline methods by around 4%, which demonstrates the effectiveness of our method. Aming Wu, Yahong Han, Linchao Zhu, Yi Yang 0001 |
ICCV | 1 |
| 2021 | Graph-in-Graph Contrastive Learning for Semi-Supervised AdaptationabstractSemi-supervised domain adaptation (SSDA) aims to adapt the model from the labeled source domain to the target domain including few labeled data. Extracting the general features is important to solve SSDA, which is beneficial to promote the model to adapt to the target domain. To this end, in this paper, we propose a novel framework to enhance the generalization of the model which improves the accuracy in the target domain. Particularly, we construct a new graph-in-graph component to model the internal relationship of the input feature, which is helpful for extracting rich and general features. In addition, for large amounts of unlabeled data in the target domain, we use the contrastive loss to optimize the network and extract general representations. We evaluate our framework on three benchmark datasets including Domain-Net, Office-Home, and Office. The extensive experimental results demonstrate the proposed method achieves state-of-the-art performance. Aming Wu, Yahong Han |
ICME | 2 |
| 2021 | Domain-Smoothing Network for Zero-Shot Sketch-Based Image RetrievalabstractZero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is a novel cross-modal retrieval task, where abstract sketches are used as queries to retrieve natural images under zero-shot scenario. Most existing methods regard ZS-SBIR as a traditional classification problem and employ a cross-entropy or triplet-based loss to achieve retrieval, which neglect the problems of the domain gap between sketches and natural images and the large intra-class diversity in sketches. Toward this end, we propose a novel Domain-Smoothing Network (DSN) for ZS-SBIR. Specifically, a cross-modal contrastive method is proposed to learn generalized representations to smooth the domain gap by mining relations with additional augmented samples. Furthermore, a category-specific memory bank with sketch features is explored to reduce intra-class diversity in the sketch domain. Extensive experiments demonstrate that our approach notably outperforms the state-of-the-art methods in both Sketchy and TU-Berlin datasets. Hao Wang 0062, Jiexi Yan, Aming Wu, Cheng Deng 0002 |
IJCAI | 4 |
| 2021 | Generalized and Discriminative Few-Shot Object Detection via SVD-Dictionary EnhancementabstractFew-shot object detection (FSOD) aims to detect new objects based on few annotated samples. To alleviate the impact of few samples, enhancing the generalization and discrimination abilities of detectors on new objects plays an important role. In this paper, we explore employing Singular Value Decomposition (SVD) to boost both the generalization and discrimination abilities. In specific, we propose a novel method, namely, SVD-Dictionary enhancement, to build two separated spaces based on the sorted singular values. Concretely, the eigenvectors corresponding to larger singular values are used to build the generalization space in which localization is performed, as these eigenvectors generally suppress certain variations (e.g., the variation of styles) and contain intrinsical characteristics of objects. Meanwhile, since the eigenvectors corresponding to relatively smaller singular values may contain richer category-related information, we can utilize them to build the discrimination space in which classification is performed. Dictionary learning is further leveraged to capture high-level discriminative information from the discrimination space, which is beneficial for improving detection accuracy. In the experiments, we separately verify the effectiveness of our method on PASCAL VOC and COCO benchmarks. Particularly, for the 2-shot case in VOC split1, our method significantly outperforms the baseline by 6.2\%. Moreover, visualization analysis shows that our method is instrumental in doing FSOD. Aming Wu, Suqi Zhao, Cheng Deng 0002, Wei Liu 0005 |
NeurIPS | 1 |
| 2021 | Visual commonsense reasoning with directional visual connectionsabstractTo boost research into cognition-level visual understanding, i.e., making an accurate inference based on a thorough understanding of visual details, visual commonsense reasoning (VCR) has been proposed. Compared with traditional visual question answering which requires models to select correct answers, VCR requires models to select not only the correct answers, but also the correct rationales. Recent research into human cognition has indicated that brain function or cognition can be considered as a global and dynamic integration of local neuron connectivity, which is helpful in solving specific cognition tasks. Inspired by this idea, we propose a directional connective network to achieve VCR by dynamically reorganizing the visual neuron connectivity that is contextualized using the meaning of questions and answers and leveraging the directional information to enhance the reasoning ability. Specifically, we first develop a GraphVLAD module to capture visual neuron connectivity to fully model visual content correlations. Then, a contextualization process is proposed to fuse sentence representations with visual neuron representations. Finally, based on the output of contextualized connectivity, we propose directional connectivity to infer answers and rationales, which includes a ReasonVLAD module. Experimental results on the VCR dataset and visualization analysis demonstrate the effectiveness of our method. Yahong Han, Aming Wu, Linchao Zhu, Yi Yang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2021 | Hierarchical Memory Decoder for Visual NarratingabstractVisual narrating focuses on generating semantic descriptions to summarize visual content of images or videos, e.g., visual captioning and visual storytelling. The challenge mainly lies in how to design a decoder to generate accurate descriptions matching visual content. Recent advances often employ a recurrent neural network (RNN), e.g., Long-Short Term Memory (LSTM), as the decoder. However, RNN is prone to diluting long-term information, which weakens its performance of capturing long-term dependencies. Recent work has demonstrated memory network (MemNet) owns the advantage of storing long-term information. However, as the decoder, it has not been well exploited for visual narrating. The reason partially comes from the difficulty of multi-modal sequential decoding with MemNet. In this article, we devise a novel memory decoder for visual narrating. Concretely, to obtain a better multi-modal representation, we first design a new multi-modal fusion method to fully merge visual and lexical information. Then, based on the fusion result, during decoding, we construct a MemNet-based decoder consisting of multiple memory layers. Particularly, in each layer, we employ a memory set to store previous decoding information and utilize an attention mechanism to adaptively select the information related to the current output. Meanwhile, we also employ a memory set to store the decoding output of each memory layer at the current time step and still utilize an attention mechanism to select the related information. Thus, this decoder alleviates dilution of long-term information. Meanwhile, the hierarchical architecture leverages the latent information of each layer, which is helpful for generating accurate descriptions. Experimental results on two tasks of visual narrating, i.e., video captioning and visual storytelling, show that our decoder could obtain superior results and outperform the performance of conventional RNN-based decoder. Aming Wu, Yahong Han, Zhou Zhao 0001, Yi Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Bidirectional Adversarial Training for Semi-Supervised Domain AdaptationabstractSemi-supervised domain adaptation (SSDA) is a novel branch of machine learning that scarce labeled target examples are available, compared with unsupervised domain adaptation. To make effective use of these additional data so as to bridge the domain gap, one possible way is to generate adversarial examples, which are images with additional perturbations, between the two domains and fill the domain gap. Adversarial training has been proven to be a powerful method for this purpose. However, the traditional adversarial training adds noises in arbitrary directions, which is inefficient to migrate between domains, or generate directional noises from the source to target domain and reverse. In this work, we devise a general bidirectional adversarial training method and employ gradient to guide adversarial examples across the domain gap, i.e., the Adaptive Adversarial Training (AAT) for source to target domain and Entropy-penalized Virtual Adversarial Training (E-VAT) for target to source domain. Particularly, we devise a Bidirectional Adversarial Training (BiAT) network to perform diverse adversarial trainings jointly. We evaluate the effectiveness of BiAT on three benchmark datasets and experimental results demonstrate the proposed method achieves the state-of-the-art. Pin Jiang, Aming Wu, Yahong Han, Yunfeng Shao 0001, Meiyu Qi, Bingshuai Li |
IJCAI | 2 |
| 2020 | Convolutional Reconstruction-to-Sequence for Video CaptioningabstractRecent advances towards video captioning mainly follow an encoder-decoder (sequence-to-sequence) framework and generate captions via a recurrent neural network (RNN). However, employing RNN as the decoder (generator) is prone to diluting long-term information, which weakens its ability to capture long-term dependencies. Recently, some work has demonstrated that the convolutional neural network (CNN) could be used to model sequential information. Though strengths in representation ability and computation efficiency, CNN has not been well exploited in video captioning. The reason partially comes from the difficulty of modeling multi-modal sequence with CNN. In this paper, we devise a novel CNN-based encoder-decoder framework for video captioning. Particularly, we first append inter-frame differences to each CNN-extracted frame feature to get a more discriminative representation; then with that as the input, we encode each frame to be a more compact feature by a one-layer convolutional mapping, which could be taken as a reconstruction network. In the decoding stage, we first fuse visual and lexical feature; then we stack multiple dilated convolutional layers to form a hierarchical decoder. As long-term dependencies could be captured by a shorter path along the hierarchical structure, the decoder could alleviate the loss of long-term information. Experiments on two benchmark datasets show that our method could obtain state-of-the-art performance. Aming Wu, Yahong Han, Yi Yang 0001, Qinghua Hu, Fei Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Untargeted Adversarial Attack via Expanding the Semantic GapabstractRecent studies have demonstrated deep neural network-based image classifiers are vulnerable to adversarial examples. Although many existing methods could obtain outstanding attack performance, they often require certain information about the attacked model, e.g., the output category scores. Meanwhile, the optimization-based methods need many steps to generate adversarial examples. In practice, we could obtain the output label but the category scores. Besides, compared to those samples with large semantic category gaps, e.g., Panda and Gibbon, most existing methods are not easy to find adversarial examples on samples with small semantic category gaps, e.g., Tabby Cat and Egyptian Cat. Thus, we propose an untargeted adversarial attack method via expanding the semantic gap, which only relies on the output label. And we use the optimization-based method to generate adversarial examples. On five normally trained models and five state-of-the-art attack methods, extensive experiments show that our method is effective and obtains better attack performance. Aming Wu, Yahong Han, Quanxin Zhang 0001, Xiaohui Kuang |
ICME | 1 |
| 2019 | Video Interactive Captioning with Human PromptsabstractVideo captioning aims at generating a proper sentence to describe the video content. As a video often includes rich visual content and semantic details, different people may be interested in different views. Thus the generated sentence always fails to meet the ad hoc expectations. In this paper, we make a new attempt that, we launch a round of interaction between a human and a captioning agent. After generating an initial caption, the agent asks for a short prompt from the human as a clue of his expectation. Then, based on the prompt, the agent could generate a more accurate caption. We name this process a new task of video interactive captioning (ViCap). Taking a video and an initial caption as input, we devise the ViCap agent which consists of a video encoder, an initial caption encoder, and a refined caption generator. We show that the ViCap can be trained via a full supervision (with ground-truth) way or a weak supervision (with only prompts) way. For the evaluation of ViCap, we first extend the MSRVTT with interaction ground-truth. Experimental results not only show the prompts can help generate more accurate captions, but also demonstrate the good performance of the proposed method. Aming Wu, Yahong Han, Yi Yang 0001 |
IJCAI | 1 |
| 2019 | Connective Cognition Network for Directional Visual Commonsense ReasoningabstractVisual commonsense reasoning (VCR) has been introduced to boost research of cognition-level visual understanding, i.e., a thorough understanding of correlated details of the scene plus an inference with related commonsense knowledge. Recent studies on neuroscience have suggested that brain function or cognition can be described as a global and dynamic integration of local neuronal connectivity, which is context-sensitive to specific cognition tasks. Inspired by this idea, towards VCR, we propose a connective cognition network (CCN) to dynamically reorganize the visual neuron connectivity that is contextualized by the meaning of questions and answers. Concretely, we first develop visual neuron connectivity to fully model correlations of visual content. Then, a contextualization process is introduced to fuse the sentence representation with that of visual neurons. Finally, based on the output of contextualized connectivity, we propose directional connectivity to infer answers or rationales. Experimental results on the VCR dataset demonstrate the effectiveness of our method. Particularly, in $Q \to AR$ mode, our method is around 4\% higher than the state-of-the-art method. Aming Wu, Linchao Zhu, Yahong Han, Yi Yang 0001 |
NeurIPS | 1 |
| 2019 | Capturing the spatio-temporal continuity for video semantic segmentationabstractIn recent years, image semantic segmentation based on a convolutional neural network has achieved many advances. However, the development of video semantic segmentation is relatively slow. Directly applying the image segmentation algorithms to each video frame separately may ignore the temporal region continuity inherent in videos. In this study, the authors propose a novel deep neural network architecture with a newly devised spatio‐temporal continuity (STC) module for video semantic segmentation. Particularly, the architecture includes an encoding network, an STC module, and a decoding network. The encoding network is used to extract a high‐level feature map. The STC module then uses the high‐level feature map as input to extract the STC feature map. For decoding, they use four dilated convolutional layers to obtain more abstract representation and a deconvolutional layer to increase the size of the representation. Finally, they fuse the current feature representation and the previous feature representation and get the class probabilities. Thus, this architecture receives a sequence of consecutive video frames and outputs the segmentation result of the current frame. They extensively evaluate the proposed approach on the CamVid and KITTI datasets. Compared with other methods, the authors’ approach not only achieves competitive performance but also has lower complexity. Aming Wu, Yahong Han |
IET Image Process. | 2 |
| 2018 | Multi-modal Circulant Fusion for Video-to-Language and BackwardabstractMulti-modal fusion has been widely involved in focuses of the modern artificial intelligence research, e.g., from visual content to languages and backward. Common-used multi-modal fusion methods mainly include element-wise product, element-wise sum, or even simply concatenation between different types of features, which are somewhat straightforward but lack in-depth analysis. Recent studies have shown fully exploiting interactions among elements of multi-modal features will lead to a further performance gain. In this paper, we put forward a new approach of multi-modal fusion, namely Multi-modal Circulant Fusion (MCF). Particularly, after reshaping feature vectors into circulant matrices, we define two types of interaction operations between vectors and matrices. As each row of the circulant matrix shifts one elements, with newly-defined interaction operations, we almost explore all possible interactions between vectors of different modalities. Moreover, as only regular operations are involved and defined a priori, MCF avoids increasing parameters or computational costs for multi-modal fusion. We evaluate MCF with tasks of video captioning and temporal activity localization via language (TALL). Experiments on MSVD and MSRVTT show our method obtains the state-of-the-art for video captioning. For TALL, by plugging into MCF, we achieve a performance gain of roughly 4.2% on TACoS. Aming Wu, Yahong Han |
IJCAI | 1 |