VLDB 2026 Research / reviewers in the wild / expert
Peng Wang 0023
dblp:95/4442-23
· DBLP profile ↗
64ranked-venue papers
7as first author
41since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 5 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 40 · 6 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Static: Related Questions Retrieval Through Conversations in Community Question AnsweringabstractIn community question answering (cQA) platforms like Stack Overflow, related question retrieval is recognized as a fundamental task that allows users to retrieve related questions to answer user queries automatically. Although many traditional approaches have been proposed for investigating this research field, they mostly rely on static approaches and neglect the interaction property. We argue that the conversational way can well distinguish the fine-grained representations of questions and has great potential to improve the performance of question retrieval. In this paper, we propose a related question retrieval model through conversations, called TeCQR, to locate related questions in cQA. Specifically, we build conversations by utilizing tag-enhanced clarifying questions. In addition, we design a noise tolerance model that evaluates the semantic similarity between questions and tags, enabling the model to effectively handle noisy feedback. Moreover, the tag-enhanced two-stage offline training is proposed to fully exploit the mutual relationships among user queries, questions, and tags to learn their fine-grained representations. Based on the learned representations and contextual conversations, TeCQR incorporates conversational feedback by learning to ask tag-enhanced clarifying questions to retrieve related questions more effectively. Experimental results demonstrate that our model significantly outperforms state-of-the-art baselines. Xiao Ao, Jie Zou 0001, Yibiao Wei, Peng Wang 0023, Weikang Guo |
AAAI | 4 |
| 2026 | From Dialogue to Destination: Geography-Aware Large Language Models with Multimodal Fusion for Conversational RecommendationabstractConversational Recommender Systems (CRS) aim to provide personalized recommendations by interacting with users through natural language dialogue. However, in scenarios requiring deep geospatial awareness, existing methods, including those based on Large Language Models (LLMs), still face significant challenges in effectively fusing heterogeneous, multimodal geographic information with dynamic dialogue context. Simple fusion strategies struggle to resolve the asymmetric dependencies between dynamic user intent and static geographic context and fail to bridge the semantic gap between LLMs and structured geospatial data. To address these issues, we propose a framework for geography-aware CRS, named GeoCRS. Our core idea is to empower a frozen LLM with powerful geospatial reasoning capabilities by conditioning it on a dynamic, multimodal guidance signal generated by an external fusion architecture, all without altering the LLM's internal parameters. Specifically, we first design a hierarchical geographical encoder to uniformly represent heterogeneous geographic data. Subsequently, we introduce a contextual feature modulation module that asymmetrically injects the geographic context into the user's dialogue intent via a novel modulation mechanism to improve conversational recommendation via both geographic and dialogue context. Extensive experiments on public benchmark datasets demonstrate that our proposed GeoCRS significantly outperforms state-of-the-art baselines on the geography-aware conversational recommendation task. Chenxi Liu 0003, Jie Zou 0001, Cheng Long 0001, Chaoning Zhang, Peng Wang 0023, Yang Yang 0002 |
AAAI | 6 |
| 2026 | Seeing Beyond Illusion: Generalized and Efficient Mirror DetectionabstractReflective imaging enables the mirror imagings and physical entities to possess identical attributes, e.g., color and shape. Current mirror detection (MD) methods primarily rely on designing functional components to establish the correlation and disparities between the imagings and entities, thereby identifying the mirror regions. However, the exploration of extended scenes with dynamic content changes is rarely investigated. Therefore, we propose the MirrorSAM designed for MD based on the Segment Anything Model (SAM). Specifically, due to the varying reflections produced by mirrors in different positions and the complex visual space that interferes with localization, we design the hierarchical mixture of direction experts (HMDE) in the low-rank space to reduce biases towards entities in SAM and dynamically adjust experts based on the input scene. We observe differences in depth between mirrors and adjacent areas, and propose the depth token calibration (DTC), which introduces a learnable depth token to generate the depth map and serve as an error correction factor. We further formulate the selective pixel-prototype contrastive (SPPC) loss, selecting partially confusable samples to promote the decoupling of mirror and non-mirror representations. Extensive experiments conducted on four mirror benchmarks and two settings demonstrate that our approach surpasses state-of-the-art methods with few trainable parameters and FLOPs. We further extend to four transparent surface benchmarks to validate generalization. Mingfeng Zha, Guoqing Wang 0001, Tianyu Li 0003, Wei Dong 0010, Peng Wang 0023, Yang Yang 0002 |
AAAI | 5 |
| 2026 | Teacher Agent: A Knowledge Distillation-Free Framework for Rehearsal-Based Video Incremental Learning
Shengqin Jiang, Yaoyu Fang, Haokui Zhang, Qingshan Liu 0001, Yuankai Qi, Yang Yang 0002, Peng Wang 0023 |
Int. J. Comput. Vis. | 7 |
| 2025 | Efficient Adaptation of Pre-Trained Vision Transformer Underpinned by Approximately Orthogonal Fine-Tuning StrategyabstractA prevalent approach in Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViT) involves freezing the majority of the backbone parameters and solely learning low-rank adaptation weight matrices to accommodate downstream tasks. These low-rank matrices are commonly derived through the multiplication structure of down-projection and up-projection matrices, exemplified by methods such as LoRA and Adapter. In this work, we observe an approximate orthogonality among any two row or column vectors within any weight matrix of the backbone parameters; however, this property is absent in the vectors of the down/up-projection matrices. Approximate orthogonality implies a reduction in the upper bound of the model's generalization error, signifying that the model possesses enhanced generalization capability. If the fine-tuned down/up-projection matrices were to exhibit this same property as the pre-trained backbone matrices, could the generalization capability of fine-tuned ViTs be further augmented? To address this question, we propose an Approximately Orthogonal Fine-Tuning (AOFT) strategy for representing the low-rank weight matrices. This strategy employs a single learnable vector to generate a set of approximately orthogonal vectors, which form the down/up-projection matrices, thereby aligning the properties of these matrices with those of the backbone. Extensive experimental results demonstrate that our method achieves competitive performance across a range of downstream image classification tasks, confirming the efficacy of the enhanced generalization capability embedded in the down/up-projection matrices. Yiting Yang, Qingsen Yan, Haokui Zhang, Wei Dong 0010, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ICCV | 8 |
| 2025 | Implicit Counterfactual Learning for Audio-Visual SegmentationabstractAudio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and imbalances. To overcome this, we propose the implicit counterfactual framework (ICF) to achieve unbiased cross-modal understanding. Due to the lack of semantics, heterogeneous representations may lead to erroneous matches, especially in complex scenes with ambiguous visual content or interference from multiple audio sources. We introduce the multi-granularity implicit text (MIT) involving video-, segment- and frame-level as the bridge to establish the modality-shared space, reducing modality gaps and providing prior guidance. Visual content carries more information and typically dominates, thereby marginalizing audio features in the decision-making. To mitigate knowledge preference, we propose the semantic counterfactual (SC) to learn orthogonal representations in the latent space, generating diverse counterfactual samples, thus avoiding biases introduced by complex functional designs and explicit modifications of text structures or attributes. We further formulate the collaborative distribution-aware contrastive learning (CDCL), incorporating factual-counterfactual and inter-modality contrasts to align representations, promoting cohesion and decoupling. Extensive experiments on three public datasets validate that the proposed method achieves state-of-the-art performance. Mingfeng Zha, Tianyu Li 0003, Guoyin Wang 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ICCV | 4 |
| 2025 | InteractGuide: LLM-Enhanced Multimodal Reasoning for User-Centric Interaction Recommendations in AR-HRI AuthoringabstractAugmented Reality (AR) enhances Human-Robot Interaction (HRI) by offering diverse interaction methods. However, existing systems often fail to resolve the conflict between a user's implicit preferences and physical ergonomics, leading to suboptimal experiences. We introduce InteractGuide, a novel framework that, for the first time, uses a Large Language Model (LLM) as a central reasoning engine to dynamically balance these competing factors. Our system translates physiological signals into a symbolic ''Preference Memory'' that the LLM reasons over, alongside real-time ergonomic and contextual data, to provide personalized interaction recommendations. A 29-participant study confirms our architecture improves efficiency and experience compared to single-factor approaches, showing the potential of LLMs as reasoning engines for complex AR-HRI. This work presents a validated end-to-end architecture for user-centric interaction adaptation, demonstrating the potential of LLMs as reasoning engines in complex AR-HRI systems. Yunqiang Pei, Hongrong Yang, Guoqing Wang 0001, Peng Wang 0023, Chaoning Zhang, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 5 |
| 2025 | Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language ModelsabstractIn this work, we address the task of table image to LaTeX code generation, with the goal of automating the reconstruction of high-quality, publication-ready tables from visual inputs. A central challenge of this task lies in accurately handling complex tables—those with large sizes, deeply nested structures, and semantically rich or irregular cell content—where existing methods often fail. We begin with a comprehensive analysis, identifying key challenges and highlighting the limitations of current evaluation protocols. To overcome these issues, we propose a reinforced multimodal large language model (MLLM) framework, where a pre-trained MLLM is fine-tuned on a large-scale table-to-LaTeX dataset. To further improve generation quality, we introduce a dual-reward reinforcement learning strategy based on Group Relative Policy Optimization (GRPO). Unlike standard approaches that optimize purely over text outputs, our method incorporates both a structure-level reward on LaTeX code and a visual fidelity reward computed from rendered outputs, enabling direct optimization of the visual output quality.
We adopt a hybrid evaluation protocol combining TEDS-Structure and CW-SSIM, and show that our method achieves state-of-the-art performance, particularly on structurally complex tables, demonstrating the effectiveness and robustness of our approach. Jun Ling, Yao Qi, Shibo Zhou, Yanqin Huang, Yang Yang 0002, Heng Tao Shen, Peng Wang 0023 |
NeurIPS | 11 |
| 2025 | AttentionAR: AR Adaptation and Warning for Real-World Safety via Attention Modeling and MLLM Reasoning
Yunqiang Pei, Renming Huang, Mingfeng Zha, Guoqing Wang 0001, Peng Wang 0023, Qiao Kang, Yang Yang 0002, Heng Tao Shen |
UIST | 5 |
| 2025 | An Empirical Study on Training Paradigms for Deep Supervised Hashing
Yang Shen 0006, Peng Wang 0023, Xiu-Shen Wei, Yazhou Yao |
Int. J. Comput. Vis. | 2 |
| 2025 | Delving Deep into Simplicity Bias for Long-Tailed Image Recognition
Xiu-Shen Wei, Xuhao Sun, Yang Shen 0006, Peng Wang 0023 |
Int. J. Comput. Vis. | 4 |
| 2025 | New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM CollaborationabstractReferring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. It serves as an essential testing ground for Multimodal Large Language Models (MLLMs). To advance this field, we introduced a new REC dataset in our previous conference paper, characterized by two key features. First, it is designed with controllable difficulty levels, requiring multi-level fine-grained reasoning across object categories, attributes, and multi-hop relationships. Second, it incorporates negative text and images generated through fine-grained editing and augmentation, explicitly testing a model's ability to reject scenarios where the target object is absent-an often-overlooked yet critical challenge in existing datasets. In this extended work, we propose two new methods to tackle the challenges of fine-grained REC by combining the strengths of Specialist Models and MLLMs. The first method adaptively assigns simple cases to faster, lightweight models and reserves complex ones for powerful MLLMs, balancing accuracy and efficiency. The second method lets a specialist generate a set of possible object regions, and the MLLM selects the most plausible one using its reasoning ability. These collaborative strategies lead to significant improvements on our dataset and other challenging benchmarks. Our results show that combining specialized and general-purpose models offers a practical path toward solving complex real-world vision-language tasks. Xuzheng Yang, Junzhuo Liu 0002, Peng Wang 0023, Guoqing Wang 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Fine-Grained Alignment and Interaction for Video Grounding With Cross-Modal Semantic Hierarchical GraphabstractVideo grounding tasks have recently gained significant attention. However, existing methods failed to fully comprehend the semantics within queries and videos, often overlooking key content. Moreover, the lack of fine-grained cross-modal alignment and interaction to guide the semantic matching of complex texts and videos lead to inconsistent representational modeling. To address this issue, we propose a Semantic Hierarchical Grounding model, referred to as SHG, and design a cross-modal semantic hierarchical graph to achieve fine-grained semantic understanding. SHG decomposes both the query and each video moment into three levels: global, action, and element. This topology, ranging from global to local, establishes multigranularity intrinsic connections between the two modalities, fostering a comprehensive understanding of dynamic semantics and fine-grained cross-modal matching. Accordingly, to fully leverage the rich information within the cross-modal semantic hierarchical graph, we employ contrastive learning by seeking samples with the same action and element semantics, then achieve node-moment cross-modal hierarchical matching for global alignment. This approach can unearth fine-grained clues and align semantics across multiple granularities. Moreover, we combine the designed hierarchical graph interaction for coarse-to-fine fusion of text and video, thereby enabling highly accurate video grounding. Extensive experiments conducted on three challenging public datasets (ActivityNet-Captions, TACoS, and Charades-STA) demonstrate that the proposed approach outperforms state-of-the-art techniques, validating its effectiveness. Ran Ran 0001, Jiwei Wei, Shiyuan He, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Low-Rank Rescaled Vision Transformer Fine-Tuning: A Residual Design ApproachabstractParameter-efficient fine-tuning for pre-trained Vision Transformers aims to adeptly tailor a model to downstream tasks by learning a minimal set of new adaptation parameters while preserving the frozen majority of pre-trained parameters. Striking a balance between retaining the generalizable representation capacity of the pre-trained model and acquiring task-specific features poses a key challenge. Currently, there is a lack of focus on guiding this delicate trade-off. In this study, we approach the problem from the perspective of Singular Value Decomposition (SVD) of pre-trained parameter matrices, providing insights into the tuning dynamics of existing methods. Building upon this understanding, we propose a Residual-based Low-Rank Rescaling (RLRR) fine-tuning strategy. This strategy not only enhances flexibility in parameter tuning but also ensures that new parameters do not deviate excessively from the pre-trained model through a residual design. Extensive experiments demonstrate that our method achieves competitive performance across various downstream image classification tasks, all while maintaining comparable new parameters. We believe this work takes a step forward in offering a unified perspective for interpreting existing methods and serves as motivation for the development of new approaches that move closer to effectively considering the crucial trade-off mentioned above. Our code is available at https://github.com/zstarN70/RLRR.git. Wei Dong 0010, Bihui Chen, Dawei Yan 0001, Zhijun Lin, Qingsen Yan, Peng Wang 0023, Yang Yang 0002 |
CVPR | 7 |
| 2024 | Diffusion Models as Optimizers for Efficient Planning in Offline RL
Renming Huang, Yunqiang Pei, Guoqing Wang 0001, Yang Yang 0002, Peng Wang 0023, Heng Tao Shen |
ECCV (51) | 6 |
| 2024 | Region-Aware Distribution Contrast: A Novel Approach to Multi-task Partially Supervised Learning
Tianyu Li 0003, Guoqing Wang 0001, Peng Wang 0023, Yang Yang 0002, Jie Zou 0001 |
ECCV (51) | 4 |
| 2024 | FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression ComprehensionabstractReferring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and languageto-image grounding.Consequently, it serves as an ideal testing ground for Multi-modal Large Language Models (MLLMs).In pursuit of this goal, we have established a new REC dataset characterized by two key features: Firstly, it is designed with controllable varying levels of difficulty, necessitating multi-level fine-grained reasoning across object categories, attributes, and multi-hop relationships.Secondly, it includes negative text and images created through fine-grained editing and generation based on existing data, thereby testing the model's ability to correctly reject scenarios where the target object is not visible in the image-an essential aspect often overlooked in existing datasets and approaches.Utilizing this high-quality dataset, we conducted comprehensive evaluations of both state-ofthe-art specialist models and MLLMs.Our findings indicate that there remains a significant gap in achieving satisfactory grounding performance.We anticipate that our dataset will inspire new approaches to enhance visual reasoning and develop more advanced cross-modal interaction strategies, ultimately unlocking the full potential of MLLMs.Our code and the datasets are available at https: //github.com/liujunzhuo/FineCops-Ref. Junzhuo Liu 0002, Xuzheng Yang, Weiwei Li 0005, Peng Wang 0023 |
EMNLP | 4 |
| 2024 | Emotion Recognition in HMDs: A Multi-task Approach Using Physiological Signals and Occluded FacesabstractPrior research on emotion recognition in extended reality (XR) has faced challenges due to the occlusion of facial expressions by Head-Mounted Displays (HMDs). This limitation hinders accurate Facial Expression Recognition (FER), which is crucial for immersive user experiences. This study aims to overcome the occlusion challenge by integrating physiological signals with partially visible facial expressions to enhance emotion recognition in XR environments. We employed a multi-task approach, utilizing a feature-level fusion to fuse Electroencephalography (EEG) and Galvanic Skin Response (GSR) signals with occluded facial expressions. The model predicts valence and arousal simultaneously from both macro-and micro-expression. Our method demonstrated improved accuracy in emotion recognition under partial occlusion conditions. The integration of temporal physiological signals with other modalities significantly enhanced performance, particularly for half-face emotion recognition. The study presents a novel approach to emotion recognition in XR, addressing the limitations of facial occlusion by HMDs. The findings suggest that physiological signals are vital for interpreting emotions in occluded scenarios, offering potential for real-time applications and advancing social XR applications. Yunqiang Pei, Jialei Tang, Qihang Tang, Mingfeng Zha, Dongyu Xie, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 9 |
| 2024 | Improving Interaction Comfort in Authoring Task in AR-HRI through Dynamic Dual-Layer Interaction AdjustmentabstractPrevious research has demonstrated the potential of Augmented Reality in enhancing psychological comfort in Human-Robot Interaction (AR-HRI) through shared robot intent, enhanced visual feedback, and increased expressiveness and creativity in interaction methods. However, the challenge of selecting interaction methods that enhance physical comfort in varying scenarios remains. This study purposes a dynamic dual-layer interaction adjustment mechanism to improve user comfort and interaction efficiency. The mechanism comprises two models: an general layer model, grounded in ergonomics principles, identifies appropriate areas for various interaction methods; a individual layer model predicts user discomfort levels using physiological signals. Interaction methods are dynamically adjusted based on discomfort level changes, enabling the system to adapt to individual differences and dynamic changes, thereby reducing misjudgments and enhancing comfort management. The mechanism's success in authoring tasks validates its effectiveness, significantly advancing AR-HRI and fostering more comfortable and enhancing efficient human-centered interactions. Yunqiang Pei, Hongrong Yang, Qihang Tang, Jialei Tang, Guoqing Wang 0001, Zhitao Liu, Ning Xie 0003, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
ACM Multimedia | 10 |
| 2024 | Efficient Adaptation of Pre-trained Vision Transformer via Householder TransformationabstractA common strategy for Parameter-Efficient Fine-Tuning (PEFT) of pre-trained Vision Transformers (ViTs) involves adapting the model to downstream tasks by learning a low-rank adaptation matrix. This matrix is decomposed into a product of down-projection and up-projection matrices, with the bottleneck dimensionality being crucial for reducing the number of learnable parameters, as exemplified by prevalent methods like LoRA and Adapter. However, these low-rank strategies typically employ a fixed bottleneck dimensionality, which limits their flexibility in handling layer-wise variations. To address this limitation, we propose a novel PEFT approach inspired by Singular Value Decomposition (SVD) for representing the adaptation matrix. SVD decomposes a matrix into the product of a left unitary matrix, a diagonal matrix of scaling values, and a right unitary matrix. We utilize Householder transformations to construct orthogonal matrices that efficiently mimic the unitary matrices, requiring only a vector. The diagonal values are learned in a layer-wise manner, allowing them to flexibly capture the unique properties of each layer. This approach enables the generation of adaptation matrices with varying ranks across different layers, providing greater flexibility in adapting pre-trained models. Experiments on standard downstream vision tasks demonstrate that our method achieves promising fine-tuning performance. Wei Dong 0010, Yiting Yang, Zhijun Lin, Qingsen Yan, Haokui Zhang, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
NeurIPS | 8 |
| 2024 | Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot LearningabstractMeta-learning offers a promising avenue for few-shot learning (FSL), enabling models to glean a generalizable feature embedding through episodic training on synthetic FSL tasks in a source domain. Yet, in practical scenarios where the target task diverges from that in the source domain, meta-learning based method is susceptible to over-fitting. To overcome this, we introduce a novel framework, Meta-Exploiting Frequency Prior for Cross-Domain Few-Shot Learning, which is crafted to comprehensively exploit the cross-domain transferable image prior that each image can be decomposed into complementary low-frequency content details and high-frequency robust structural characteristics. Motivated by this insight, we propose to decompose each query image into its high-frequency and low-frequency components, and parallel incorporate them into the feature embedding network to enhance the final category prediction. More importantly, we introduce a feature reconstruction prior and a prediction consistency prior to separately encourage the consistency of the intermediate feature as well as the final category prediction between the original query image and its decomposed frequency components. This allows for collectively guiding the network's meta-learning process with the aim of learning generalizable image feature embeddings, while not introducing any extra computational cost in the inference phase. Our framework establishes new state-of-the-art results on multiple cross-domain few-shot learning benchmarks. Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Zhenghua Chen, Wei Wei 0008, Chen Ding 0002, Guosheng Lin, Yanning Zhang 0001 |
NeurIPS | 2 |
| 2024 | Self-Supervised Node Representation Learning via Node-to-Neighbourhood AlignmentabstractSelf-supervised node representation learning aims to learn node representations from unlabelled graphs that rival the supervised counterparts. The key towards learning informative node representations lies in how to effectively gain contextual information from the graph structure. In this work, we present simple-yet-effective self-supervised node representation learning via aligning the hidden representations of nodes and their neighbourhood. Our first idea achieves such node-to-neighbourhood alignment by directly maximizing the mutual information between their representations, which, we prove theoretically, plays the role of graph smoothing. Our framework is optimized via a surrogate contrastive loss and a Topology-Aware Positive Sampling (TAPS) strategy is proposed to sample positives by considering the structural dependencies between nodes, which enables offline positive selection. Considering the excessive memory overheads of contrastive learning, we further propose a negative-free solution, where the main contribution is a Graph Signal Decorrelation (GSD) constraint to avoid representation collapse and over-smoothing. The GSD constraint unifies some of the existing constraints and can be used to derive new implementations to combat representation collapse. By applying our methods on top of simple MLP-based node representation encoders, we learn node representations that achieve promising node classification performance on a set of graph-structured datasets from small- to large-scale. Wei Dong 0010, Dawei Yan 0001, Peng Wang 0023 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Meta-collaborative comparison for effective cross-domain few-shot learning
Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
Pattern Recognit. | 2 |
| 2024 | Question-Aware Global-Local Video Understanding Network for Audio-Visual Question AnsweringabstractAs a newly emerging task, audio-visual question answering (AVQA) has attracted research attention. Compared with traditional single-modality (e.g., audio or visual) QA tasks, it poses new challenges due to the higher complexity of feature extraction and fusion brought by the multimodal inputs. First, AVQA requires more comprehensive understanding of the scene which involves both audio and visual information; Second, in the presence of more information, feature extraction has to be better connected with a given question; Third, features from different modalities need to be sufficiently correlated and fused. To address this situation, this work proposes a novel framework for multimodal question answering task. It characterises an audiovisual scene at both global and local levels, and within each level, the features from different modalities are well fused. Furthermore, the given question is utilised to guide not only the feature extraction at the local level but also the final fusion of global and local features to predict the answer. Our framework provides a new perspective for audio-visual scene understanding through focusing on both general and specific representations as well as aggregating multimodalities by prioritizing question-related information. As experimentally demonstrated, our method significantly improves the existing audio-visual question answering performance, with the averaged absolute gain of 3.3% and 3.1% on MUSIC-AVQA and AVQA datasets, respectively. Moreover, the ablation study verifies the necessity and effectiveness of our design. Our code will be publicly released. Zailong Chen, Lei Wang 0001, Peng Wang 0023 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Improving Pre-Trained Model-Based Speech Emotion Recognition From a Low-Level Speech Feature PerspectiveabstractMulti-view speech emotion recognition (SER) based on the pre-trained model has gained attention in the last two years, which shows great potential in improving the model performance in speaker-independent scenarios. However, the existing work either relies on various fine-tuning methods or uses excessive feature views with complex fusion strategies, causing the increase of complexity with limited performance benefit. In this paper, we improve multi-view SER based on the pre-trained model from the perspective of a low-level speech feature. Specifically, we forgo fine-tuning the pre-trained model and instead focus on learning effective features hidden in the low-level speech feature mel-scale frequency cepstral coefficient (MFCC). We propose atwo-streampoolingchannelattention (TsPCA) module to discriminatively weight the channel dimensions of the features derived from MFCC. This module enables inter-channel interaction and learning of emotion sequence information across channels. Furthermore, we design a simple but effective feature view fusion strategy to learn robust representations. In the comparison experiments, our method achieves the WA and UA of 73.97%/74.69% and 74.61%/75.66% on the IEMOCAP dataset, 97.21% and 97.11% on the Emo-DB dataset, 77.08% and 77.34% on the RAVDESS dataset, and 74.38% and 71.43% on the SAVEE dataset. Extensive experiments on the four datasets demonstrate that our method consistently surpasses existing methods and achieves a new State-of-the-Art result. Jiwei Wei, Jie Zou 0001, Peng Wang 0023, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 4 |
| 2023 | Glocal Energy-based Learning for Few-Shot Open-Set RecognitionabstractFew-shot open-set recognition (FSOR) is a challenging task of great practical value. It aims to categorize a sample to one of the predefined, closed-set classes illustrated by few examples while being able to reject the sample from unknown classes. In this work, we approach the FSOR task by proposing a novel energy-based hybrid model. The model is composed of two branches, where a classification branch learns a metric to classify a sample to one of closed-set classes and the energy branch explicitly estimates the open-set probability. To achieve holistic detection of open-set samples, our model leverages both class-wise and pixel-wise features to learn a glocal energy-based score, in which a global energy score is learned using the class-wise features, while a local energy score is learned using the pixel-wise features. The model is enforced to assign large energy scores to samples that are deviated from the few-shot examples in either the class-wise features or the pixel-wise features, and to assign small energy scores otherwise. Experiments on three standard FSOR datasets show the superior performance of our model.11Code is available at https://github.com/00why00/Glocal Haoyu Wang 0016, Guansong Pang, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
CVPR | 3 |
| 2023 | Revisiting Prototypical Network for Cross Domain Few-Shot LearningabstractPrototypical Network is a popular few-shot solver that aims at establishing a feature metric generalizable to novel few-shot classification (FSC) tasks using deep neural networks. However, its performance drops dramatically when generalizing to the FSC tasks in new domains. In this study, we revisit this problem and argue that the devil lies in the simplicity bias pitfall in neural networks. In specific, the network tends to focus on some biased shortcut features (e.g., color, shape, etc.) that are exclusively sufficient to distinguish very few classes in the meta-training tasks within a pre-defined domain, but fail to generalize across domains as some desirable semantic features. To mitigate this problem, we propose a Local-global Distillation Prototypical Network (LDP-net). Different from the standard Prototypical Network, we establish a two-branch network to classify the query image and its random local crops, respectively. Then, knowledge distillation is conducted among these two branches to enforce their class affiliation consistency. The rationale behind is that since such global-local semantic relationship is expected to hold regardless of data domains, the local-global distillation is beneficial to exploit some cross-domain transferable semantic features for feature metric establishment. Moreover, such local-global semantic consistency is further enforced among different images of the same class to reduce the intra-class semantic variation of the resultant feature. In addition, we propose to update the local branch as Exponential Moving Average (EMA) over training episodes, which makes it possible to better distill cross-episode knowledge and further enhance the generalization performance. Experiments on eight cross-domain FSC benchmarks empirically clarify our argument and show the state-of-the-art results of LDP-net. Code is available in https://github.com/NWPUZhoufei/LDP-Net Fei Zhou 0008, Peng Wang 0023, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001 |
CVPR | 2 |
| 2023 | Efficient Adaptation of Large Vision Transformer via Adapter Re-ComposingabstractThe advent of high-capacity pre-trained models has revolutionized problem-solving in computer vision, shifting the focus from training task-specific models to adapting pre-trained models. Consequently, effectively adapting large pre-trained models to downstream tasks in an efficient manner has become a prominent research area. Existing solutions primarily concentrate on designing lightweight adapters and their interaction with pre-trained models, with the goal of minimizing the number of parameters requiring updates. In this study, we propose a novel Adapter Re-Composing (ARC) strategy that addresses efficient pre-trained model adaptation from a fresh perspective. Our approach considers the reusability of adaptation parameters and introduces a parameter-sharing scheme. Specifically, we leverage symmetric down-/up-projections to construct bottleneck operations, which are shared across layers. By learning low-dimensional re-scaling coefficients, we can effectively re-compose layer-adaptive adapters. This parameter-sharing strategy in adapter design allows us to further reduce the number of new parameters while maintaining satisfactory performance, thereby offering a promising approach to compress the adaptation cost. We conduct experiments on 24 downstream image classification tasks using various Vision Transformer variants to evaluate our method. The results demonstrate that our approach achieves compelling transfer learning performance with a reduced parameter count. Our code is available at https://github.com/DavidYanAnDe/ARC. Wei Dong 0010, Dawei Yan 0001, Zhijun Lin, Peng Wang 0023 |
NeurIPS | 4 |
| 2023 | Learning Graph Convolutional Networks for Multi-Label Recognition and ApplicationsabstractThe task of multi-label image recognition is to predict a set of object labels that present in an image. As objects normally co-occur in an image, it is desirable to model the label dependencies to improve the recognition performance. To capture and explore such important information, we propose graph convolutional networks (GCNs) based models for multi-label image recognition, where directed graphs are constructed over classes and information is propagated between classes to learn inter-dependent class-level representations. Following this idea, we design two particular models that approach multi-label classification from different views. In our first model, the prior knowledge about the class dependencies is integrated into classifier learning. Specifically, we propose Classifier Learning GCN (C-GCN) to map class-level semantic representations (e.g., word embeddings) into classifiers that maintain the inter-class topology. In our second model, we decompose the visual representation of an image into a set of label-aware features and propose prediction learning GCN (P-GCN) to encode such features into inter-dependent image-level prediction scores. Furthermore, we also present an effective correlation matrix construction approach to capture inter-class relationships and consequently guide information propagation among classes. Empirical results on generic multi-label image recognition demonstrate that both of the proposed models can obviously outperform other existing state-of-the-arts. Moreover, the proposed methods also show advantages in some other multi-label classification related applications. Xiu-Shen Wei, Peng Wang 0023, Yanwen Guo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Attribute-Aware Deep Hashing With Self-Consistency for Large-Scale Fine-Grained Image RetrievalabstractOur work focuses on tackling large-scale fine-grained image retrieval as ranking the images depicting the concept of interests (i.e., the same sub-category labels) highest based on the fine-grained details in the query. It is desirable to alleviate the challenges of both fine-grained nature of small inter-class variations with large intra-class variations and explosive growth of fine-grained data for such a practical task. In this paper, we propose attribute-aware hashing networks with self-consistency for generating attribute-aware hash codes to not only make the retrieval process efficient, but also establish explicit correspondences between hash codes and visual attributes. Specifically, based on the captured visual representations by attention, we develop an encoder-decoder structure network of a reconstruction task to unsupervisedly distill high-level attribute-specific vectors from the appearance-specific visual representations without attribute annotations. Our models are also equipped with a feature decorrelation constraint upon these attribute vectors to strengthen their representative abilities. Then, driven by preserving original entities' similarity, the required hash codes can be generated from these attribute-specific vectors and thus become attribute-aware. Furthermore, to combat simplicity bias in deep hashing, we consider the model design from the perspective of the self-consistency principle and propose to further enhance models' self-consistency by equipping an additional image reconstruction path. Comprehensive quantitative experiments under diverse empirical settings on six fine-grained retrieval datasets and two generic retrieval datasets show the superiority of our models over competing methods. Moreover, qualitative results demonstrate that not only the obtained hash codes can strongly correspond to certain kinds of crucial properties of fine-grained objects, but also our self-consistency designs can effectively overcome simplicity bias in fine-grained hashing. Xiu-Shen Wei, Yang Shen 0006, Xuhao Sun, Peng Wang 0023, Yuxin Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Node Representation Learning in Graph via Node-to-Neighbourhood Mutual Information MaximizationabstractThe key towards learning informative node representations in graphs lies in how to gain contextual information from the neighbourhood. In this work, we present a simple-yet-effective self-supervised node representation learning strategy via directly maximizing the mutual information between the hidden representations of nodes and their neighbourhood, which can be theoretically justified by its link to graph smoothing. Following InfoNCE, our framework is optimized via a surrogate contrastive loss, where the positive selection underpins the quality and efficiency of rep-resentation learning. To this end, we propose a topology-aware positive sampling strategy, which samples positives from the neighbourhood by considering the structural dependencies between nodes and thus enables positive selection upfront. In the extreme case when only one positive is sampled, we fully avoid expensive neighbourhood aggregation. Our methods achieve promising performance on various node classification datasets. It is also worth mentioning by applying our loss function to MLP based node encoders, our methods can be orders of faster than existing solutions. Our codes and supplementary materials are available at https://github.com/dongwei156/n2n. Wei Dong 0010, Junsheng Wu, ZongYuan Ge, Peng Wang 0023 |
CVPR | 5 |
| 2022 | RPC: a large-scale and fine-grained retail product checkout dataset
Xiu-Shen Wei, Quan Cui, Lei Yang 0048, Peng Wang 0023, Lingqiao Liu, Jian Yang 0003 |
Sci. China Inf. Sci. | 4 |
| 2022 | Industrial IoT Intrusion Detection via Evolutionary Cost-Sensitive Learning and Fog ComputingabstractCyber attacks and intrusions have become the major obstacles to the adoption of the Industrial Internet of Things (IIoT) in critical industries. Imbalanced data distribution is a common problem in IIoT environments that negatively influence machine learning-based intrusion detection systems (IDSs). To address this issue, we introduce EvolCostDeep, a hybrid model of stacked autoencoders (SAE) and convolutional neural networks (CNNs) with a new cost-dependent loss function. The loss function aims to optimize the model’s parameters, where the costs are determined using an evolutionary algorithm. The combination of evolutionary algorithms and deep learning (DL) on Big data hinders the scalability of IIoT IDSs. In this regard, a fog computing-enabled framework, called DeepIDSFog, is designed at the data level, where the master node shares the EvolCostDeep model with worker nodes. In each fog worker node, the EvolCostDeep is parallelized through one task-level and two model-level mechanisms. After aggregating detection outputs from worker nodes to the master, the result is passed to the cloud platform for mitigating attacks. A series of experiments is conducted on the ToN-IoT and UNSW-NB15 data sets to evaluate the performance of EvolCostDeep and DeepIDSFog. The results show that our frameworks can effectively handle both class imbalance problem and scalability of big IIoT traffic data compared with the other models. The averaged values of the EvolCostDeep for recall, precision, and$F1$-score on the data sets are of 93.3%, 97.6%, and 95.2%, respectively, which are higher than the compared methods. Also, the DeepIDSFog provides an average speedup of$38.7\times $over other comparing models. Akbar Telikani, Jun Shen 0001, Jie Yang 0009, Peng Wang 0023 |
IEEE Internet Things J. | 4 |
| 2022 | Improving performance and efficiency of Graph Neural Networks by injective aggregation
Wei Dong 0010, Junsheng Wu, Xinwan Zhang, Zongwen Bai, Peng Wang 0023, Marcin Wozniak |
Knowl. Based Syst. | 5 |
| 2022 | Anisotropic Convolutional Neural Networks for RGB-D Based Semantic Scene CompletionabstractSemantic scene completion (SSC) is a computer vision task aiming to simultaneously infer the occupancy and semantic labels for each voxel in a scene from partial information consisting of a depth image and/or a RGB image. As a voxel-wise labeling task, the key for SSC is how to effectively model the visual and geometrical variations to complete the scene. To this end, we propose the Anisotropic Network (AIC-Net), with novel convolutional modules that can model varying anisotropic receptive fields voxel-wisely in a computationally efficient manner. The basic idea to achieve such anisotropy is to decompose 3D convolution into three consecutive dimensional convolutions, and determine the dimension-wise kernels on the fly. One module, termed kernel-selection anisotropic (KSA) convolution, adaptively selects the optimal kernel sizes for each dimensional convolution from a set of candidate kernels, and the other module, termed kernel-modulation anisotropic (KMA) convolution, directly modulates a single convolutional kernel for each dimension to derive more flexible receptive field. By stacking multiple such anisotropic modules, the 3D context modeling capability and flexibility can be further enhanced. Moreover, we present a new end-to-end trainable framework to approach the SSC task avoiding the expensive TSDF pre-processing as in many existing methods. Extensive experiments on SSC benchmarks show the advantage of the proposed methods. Jie Li 0040, Peng Wang 0023, Kai Han 0001, Yu Liu 0029 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Relation Regularized Scene Graph GenerationabstractScene graph generation (SGG) is built on top of detected objects to predict object pairwise visual relations for describing the image content abstraction. Existing works have revealed that if the links between objects are given as prior knowledge, the performance of SGG is significantly improved. Inspired by this observation, in this article, we propose a relation regularized network (R2-Net), which can predict whether there is a relationship between two objects and encode this relation into object feature refinement and better SGG. Specifically, we first construct an affinity matrix among detected objects to represent the probability of a relationship between two objects. Graph convolution networks (GCNs) over this relation affinity matrix are then used as object encoders, producing relation-regularized representations of objects. With these relation-regularized features, our R2-Net can effectively refine object labels and generate scene graphs. Extensive experiments are conducted on the visual genome dataset for three SGG tasks (i.e., predicate classification, scene graph classification, and scene graph detection), demonstrating the effectiveness of our proposed method. Ablation studies also verify the key roles of our proposed components in performance improvement. Yuyu Guo 0001, Lianli Gao, Jingkuan Song, Peng Wang 0023, Nicu Sebe, Heng Tao Shen, Xuelong Li 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Continual Referring Expression Comprehension via Dual Modular MemorizationabstractReferring Expression Comprehension (REC) aims to localize an image region of a given object described by a natural-language expression. While promising performance has been demonstrated, existing REC algorithms make a strong assumption that training data feeding into a model are given upfront, which degrades its practicality for real-world scenarios. In this paper, we propose Continual Referring Expression Comprehension (CREC), a new setting for REC, where a model is learning on a stream of incoming tasks. In order to continuously improve the model on sequential tasks without forgetting prior learned knowledge and without repeatedly re-training from a scratch, we propose an effective baseline method named Dual Modular Memorization (DMM), which alleviates the problem of catastrophic forgetting by two memorization modules: Implicit-Memory and Explicit-Memory. Specifically, the former module aims to constrain drastic changes to important parameters learned on old tasks when learning a new task; while the latter module maintains a buffer pool to dynamically select and store representative samples of each seen task for future rehearsal. We create three benchmarks for the new CREC setting, by respectively re-splitting three widely-used REC datasets RefCOCO, RefCOCO+ and RefCOCOg into sequential tasks. Extensive experiments on the constructed benchmarks demonstrate that our DMM method significantly outperforms other alternatives, based on two popular REC backbones. We make the source code and benchmarks publicly available to foster future progress in this field: https://github.com/zackschen/DMM. Heng Tao Shen, Peng Wang 0023, Lianli Gao, Meng Wang 0001, Jingkuan Song |
IEEE Trans. Image Process. | 3 |
| 2021 | Contrastive Learning Based Hybrid Networks for Long-Tailed Image ClassificationabstractLearning discriminative image representations plays a vital role in long-tailed image classification because it can ease the classifier learning in imbalanced cases. Given the promising performance contrastive learning has shown recently in representation learning, in this work, we explore effective supervised contrastive learning strategies and tailor them to learn better image representations from imbalanced data in order to boost the classification accuracy thereon. Specifically, we propose a novel hybrid network structure being composed of a supervised contrastive loss to learn image representations and a cross-entropy loss to learn classifiers, where the learning is progressively transited from feature learning to the classifier learning to embody the idea that better features make better classifiers. We explore two variants of contrastive loss for feature learning, which vary in the forms but share a common idea of pulling the samples from the same class together in the normalized embedding space and pushing the samples from different classes apart. One of them is the recently proposed supervised contrastive (SC) loss, which is designed on top of the state-of-the-art unsupervised contrastive loss by incorporating positive samples from the same class. The other is a prototypical supervised contrastive (PSC) learning strategy which addresses the intensive memory consumption in standard SC loss and thus shows more promise under limited memory budget. Extensive experiments on three long-tailed classification datasets demonstrate the advantage of the proposed contrastive learning based hybrid networks in long-tailed classification. Peng Wang 0023, Kai Han 0001, Xiu-Shen Wei, Lei Zhang 0054, Lei Wang 0001 |
CVPR | 1 |
| 2021 | Inter-Modality Fusion Based Attention for Zero-Shot Cross-Modal RetrievalabstractZero-shot cross-modal retrieval (ZS-CMR) performs the task of cross-modal retrieval where the classes of test categories have a different scope than the training categories. It borrows the intuition from zero-shot learning which targets to transfer the knowledge inferred during the training phase for seen classes to the testing phase for unseen classes. It mimics the real-world scenario where new object categories are continuously populating the multi-media data corpus. Unlike existing ZS-CMR approaches which use generative adversarial networks (GANs) to generate more data, we propose Inter-Modality Fusion based Attention (IMFA) and a framework ZS_INN_FUSE (Zero-Shot cross-modal retrieval using INNer product with image-text FUSEd). It exploits the rich semantics of textual data as guidance to infer additional knowledge during the training phase. This is achieved by generating attention weights through the fusion of image and text modalities to focus on the important regions in an image. We carefully create a zero-shot split based on the large-scale MS-COCO and Flickr30k datasets to perform experiments. The results show that our method achieves improvement over the ZS-CMR baseline and self-attention mechanism, demonstrating the effectiveness of inter-modality fusion in a zero-shot scenario. Bela Chakraborty, Peng Wang 0023, Lei Wang 0001 |
ICIP | 2 |
| 2021 | Exploring the auxiliary learning for long-tailed visual recognition
Junjie Zhang 0002, Lingqiao Liu, Peng Wang 0023, Jian Zhang 0002 |
Neurocomputing | 3 |
| 2021 | D3D: Dual 3-D Convolutional Network for Real-Time Action RecognitionabstractThree-dimensional convolutional neural networks (3D CNNs) have been explored to learn spatio-temporal information for video-based human action recognition. Expensive computational cost and memory demand resulted from standard 3D CNNs, however, hinder their application in practical scenarios. In this article, we address the aforementioned limitations by proposing a novel dual 3-D convolutional network (D3DNet) with two complementary lightweight branches. A coarse branch maintains large temporal receptive field by a fast temporal downsampling strategy and simulates the expensive 3-D convolutions using a combination of more efficient spatial convolutions and temporal convolutions. Meanwhile, a fine branch progressively downsamples the video in the temporal domain and adopts 3-D convolutional units with reduced channel capacities to capture multiresolution spatio-temporal information. Instead of learning these two branches independently, a shallow spatiotemporal downsampling module is shared for these two branches for efficient low-level feature learning. Besides, lateral connections are learned to effectively fuse the information from the two branches at multiple stages. The proposed network makes good balance between inference speed and action recognition performance. Based on RGB information only, it achieves competing performance on five popular video-based action recognition datasets, with inference speed of 3200 FPS on a single NVIDIA GTX 2080Ti card. Shengqin Jiang, Yuankai Qi, Haokui Zhang, Zongwen Bai, Xiaobo Lu, Peng Wang 0023 |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | Pixel-Aware Deep Function-Mixture Network for Spectral Super-ResolutionabstractSpectral super-resolution (SSR) aims at generating a hyperspectral image (HSI) from a given RGB image. Recently, a promising direction is to learn a complicated mapping function from the RGB image to the HSI counterpart using a deep convolutional neural network. This essentially involves mapping the RGB context within a size-specific receptive field centered at each pixel to its spectrum in the HSI. The focus thereon is to appropriately determine the receptive field size and establish the mapping function from RGB context to the corresponding spectrum. Due to their differences in category or spatial position, pixels in HSIs often require different-sized receptive fields and distinct mapping functions. However, few efforts have been invested to explicitly exploit this prior.To address this problem, we propose a pixel-aware deep function-mixture network for SSR, which is composed of a new class of modules, termed function-mixture (FM) blocks. Each FM block is equipped with some basis functions, i.e., parallel subnets of different-sized receptive fields. Besides, it incorporates an extra subnet as a mixing function to generate pixel-wise weights, and then linearly mixes the outputs of all basis functions with those generated weights. This enables us to pixel-wisely determine the receptive field size and the mapping function. Moreover, we stack several such FM blocks to further increase the flexibility of the network in learning the pixel-wise mapping. To encourage feature reuse, intermediate features generated by the FM blocks are fused in late stage, which proves to be effective for boosting the SSR performance. Experimental results on three benchmark HSI datasets demonstrate the superiority of the proposed method. Lei Zhang 0054, Zhiqiang Lang, Peng Wang 0023, Wei Wei 0008, Shengcai Liao, Ling Shao 0001, Yanning Zhang 0001 |
AAAI | 3 |
| 2020 | V-PROM: A Benchmark for Visual Reasoning Using Visual Progressive MatricesabstractAdvances in machine learning have generated increasing enthusiasm for tasks that require high-level reasoning on top of perceptual capabilities, particularly over visual data. Such tasks include, for example, image captioning, visual question answering, and visual navigation. Their evaluation is however hindered by task-specific confounding factors and dataset biases. In parallel, the existing benchmarks for abstract reasoning are limited to synthetic stimuli (e.g. images of simple shapes) and do not capture the challenges of real-world data. We propose a new large-scale benchmark to evaluates abstract reasoning over real visual data. The test involves visual questions that require operations fundamental to many high-level vision tasks, such as comparisons of counts and logical operations on complex visual properties. The benchmark measures a method's ability to infer high-level relationships and to generalise them over image-based concepts. We provide multiple training/test splits that require controlled levels of generalization. We evaluate a range of deep learning architectures, and find that existing models, including those popular for vision-and-language tasks, are unable to solve seemingly-simple instances. Models using relational networks fare better but leave substantial room for improvement. Damien Teney, Peng Wang 0023, Jiewei Cao, Lingqiao Liu, Chunhua Shen, Anton van den Hengel |
AAAI | 2 |
| 2020 | Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionabstractReferring expression comprehension (REF) aims at identifying a particular object in a scene by a natural language expression. It requires joint reasoning over the textual and visual domains to solve the problem. Some popular referring expression datasets, however, fail to provide an ideal test bed for evaluating the reasoning ability of the models, mainly because 1) their expressions typically describe only some simple distinctive properties of the object and 2) their images contain limited distracting information. To bridge the gap, we propose a new dataset for visual reasoning in context of referring expression comprehension with two main features. First, we design a novel expression engine rendering various reasoning logics that can be flexibly combined with rich visual properties to generate expressions with varying compositionality. Second, to better exploit the full reasoning chain embodied in an expression, we propose a new test setting by adding additional distracting images containing objects sharing similar properties with the referent, thus minimising the success rate of reasoning-free cross-domain alignment. We evaluate several state-of-the-art REF models, but find none of them can achieve promising performance. A proposed modular hard mining strategy performs the best but still leaves substantial room for improvement. Zhenfang Chen, Peng Wang 0023, Lin Ma 0002, Kwan-Yee Kenneth Wong, Qi Wu 0001 |
CVPR | 2 |
| 2020 | Anisotropic Convolutional Networks for 3D Semantic Scene CompletionabstractAs a voxel-wise labeling task, semantic scene completion (SSC) tries to simultaneously infer the occupancy and semantic labels for a scene from a single depth and/or RGB image. The key challenge for SSC is how to effectively take advantage of the 3D context to model various objects or stuffs with severe variations in shapes, layouts, and visibility. To handle such variations, we propose a novel module called anisotropic convolution, which properties with flexibility and power impossible for the competing methods such as standard 3D convolution and some of its variations. In contrast to the standard 3D convolution that is limited to a fixed 3D receptive field, our module is capable of modeling the dimensional anisotropy voxel-wisely. The basic idea is to enable anisotropic 3D receptive field by decomposing a 3D convolution into three consecutive 1D convolutions, and the kernel size for each such 1D convolution is adaptively determined on the fly. By stacking multiple such anisotropic convolution modules, the voxel-wise modeling capability can be further enhanced while maintaining a controllable amount of model parameters. Extensive experiments on two SSC benchmarks, NYU-Depth-v2 and NYUCAD, show the superior performance of the proposed method. Jie Li 0040, Kai Han 0001, Peng Wang 0023, Yu Liu 0029, Xia Yuan |
CVPR | 3 |
| 2020 | Semi-supervised Crowd Counting via Self-training on Surrogate Tasks
Yan Liu 0015, Lingqiao Liu, Peng Wang 0023, Yinjie Lei |
ECCV (15) | 3 |
| 2020 | Adaptive Importance Learning for Improving Lightweight Image Super-Resolution Network
Lei Zhang 0054, Peng Wang 0023, Chunhua Shen, Lingqiao Liu, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
Int. J. Comput. Vis. | 2 |
| 2020 | Towards Effective Deep Embedding for Zero-Shot LearningabstractZero-shot learning (ZSL) can be formulated as a cross-domain matching problem: after being projected into a joint embedding space, a visual sample will match against all candidate class-level semantic descriptions and be assigned to the nearest class. In this process, the embedding space underpins the success of such matching and is crucial for ZSL. In this paper, we conduct an in-depth study on the construction of embedding space for ZSL and posit that an ideal embedding space should satisfy two criteria: intra-class compactness and inter-class separability. While the former encourages the embeddings of visual samples of one class to distribute tightly close to the semantic description embedding of this class, the latter requires embeddings from different classes to be well separated from each other. Towards this goal, we present a simple but effective two-branch network to simultaneously map semantic descriptions and visual samples into a joint space, on which visual embeddings are forced to regress to their class-level semantic embeddings and the embeddings crossing classes are required to be distinguishable by a trainable classifier. Furthermore, we extend our method to a transductive setting to better handle the model bias problem in ZSL (i.e., samples from unseen classes tend to be categorized into seen classes) with minimal extra supervision. Specifically, we propose a pseudo labeling strategy to progressively incorporate the testing samples into the training process and thus balance the model between seen and unseen classes. Experimental results on five standard ZSL datasets show the superior performance of the proposed method and its transductive extension. Lei Zhang 0054, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Wei Wei 0008, Yanning Zhang 0001, Anton van den Hengel |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Neighbourhood Watch: Referring Expression Comprehension via Language-Guided Graph Attention NetworksabstractThe task in referring expression comprehension is to localize the object instance in an image described by a referring expression phrased in natural language. As a language-to-vision matching task, the key to this problem is to learn a discriminative object feature that can adapt to the expression used. To avoid ambiguity, the expression normally tends to describe not only the properties of the referent itself, but also its relationships to its neighbourhood. To capture and exploit this important information we propose a graph-based, language-guided attention mechanism. Being composed of node attention component and edge attention component, the proposed graph attention mechanism explicitly represents inter-object relationships, and properties with a flexibility and power impossible with competing approaches. Furthermore, the proposed graph attention mechanism enables the comprehension decision to be visualizable and explainable. Experiments on three referring expression comprehension datasets show the advantage of the proposed approach. Peng Wang 0023, Qi Wu 0001, Jiewei Cao, Chunhua Shen, Lianli Gao, Anton van den Hengel |
CVPR | 1 |
| 2019 | Multi-Label Image Recognition With Graph Convolutional NetworksabstractThe task of multi-label image recognition is to predict a set of object labels that present in an image. As objects normally co-occur in an image, it is desirable to model the label dependencies to improve the recognition performance. To capture and explore such important dependencies, we propose a multi-label classification model based on Graph Convolutional Network (GCN). The model builds a directed graph over the object labels, where each node (label) is represented by word embeddings of a label, and GCN is learned to map this label graph into a set of inter-dependent object classifiers. These classifiers are applied to the image descriptors extracted by another sub-net, enabling the whole network to be end-to-end trainable. Furthermore, we propose a novel re-weighted scheme to create an effective label correlation matrix to guide information propagation among the nodes in GCN. Experiments on two multi-label image recognition datasets show that our approach obviously outperforms other existing state-of-the-art methods. In addition, visualization analyses reveal that the classifiers learned by our model maintain meaningful semantic topology. Xiu-Shen Wei, Peng Wang 0023, Yanwen Guo 0001 |
CVPR | 3 |
| 2019 | Adaptive Multi-Path Aggregation for Human DensePose Estimation in the WildabstractDense human pose "in the wild'' task aims to map all 2D pixels of the detected human body to a 3D surface by establishing surface correspondences, i.e., surface patch index and part-specific UV coordinates. It remains challenging especially under the condition of "in the wild'', where RGB images capture complex, real-world scenes with background, occlusions, scale variations, and postural diversity. In this paper, we propose an end-to-end deep Adaptive Multi-path Aggregation network (AMA-net) for Dense Human Pose Estimation. In the proposed framework, we address two main problems: 1) how to design a simple yet effective pipeline for supporting distinct sub-tasks (e.g., instance segmentation, body part segmentation, and UV estimation); and 2) how to equip this pipeline with the ability of handling "in the wild''. To solve these problems, we first extend FPN by adding a branch for mapping 2D pixels to a 3D surface in parallel with the existing branch for bounding box detection. Then, in AMA-net, we extract variable-sized object-level feature maps (e.g., 7×7, 14×14, and 28×28), named multi-path, from multi-layer feature maps, which capture rich information of objects and are then adaptively utilized in different tasks. AMA-net is simple to train and adds only a small overhead to FPN. We discover that aside from the deep feature map, Adaptive Multi-path Aggregation is of particular importance for improving the accuracy of dense human pose estimation "in the wild''. The experimental results on the challenging Dense-COCO dataset demonstrate that our approach sets a new record for Dense Human Pose Estimation task, and it significantly outperforms the state-of-the-art methods. Our code: \urlhttps://github.com/nobody-g/AMA-net. Yuyu Guo 0001, Lianli Gao, Jingkuan Song, Peng Wang 0023, Wuyuan Xie, Heng Tao Shen |
ACM Multimedia | 4 |
| 2019 | Order-aware convolutional pooling for video based action recognition
Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Heng Tao Shen |
Pattern Recognit. | 1 |
| 2019 | Unsupervised Domain Adaptation Using Robust Class-Wise MatchingabstractUnsupervised domain adaptation (DA) enables a classifier trained on data from one domain to be applied to data from another without labels. Given that the key to transferring a classifier across domains is to mitigate the data distribution mismatch for each class, most previous works completely or partially focus on global distribution matching across domains. The global data space, however, can be complicated, which makes modeling the global distribution difficult. To mitigate this problem, we present a novel unsupervised DA framework where the DA problem is addressed by proposing a robust class-wise matching strategy. Specifically, through minimizing a maximum mean discrepancy-based class-wise fisher discriminant across domains, this framework jointly optimizes two modules: a transferable feature learning module that reduces the distribution discrepancy between the same classes as well as increasing the distribution discrepancy between different classes across domains by a linear projection, and a robust classifier that exploits both the supervised information in source domain and the unsupervised low-rank property of target domain. In experiments on three DA benchmark data sets, the proposed framework shows the state-of-the-art performance. Lei Zhang 0054, Peng Wang 0023, Wei Wei 0008, Hao Lu 0003, Chunhua Shen, Anton van den Hengel, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Piecewise Classifier Mappings: Learning Fine-Grained Learners for Novel Categories With Few ExamplesabstractHumans are capable of learning a new fine-grained concept with very little supervision, e.g., few exemplary images for a species of bird, yet our best deep learning systems need hundreds or thousands of labeled examples. In this paper, we try to reduce this gap by studying the fine-grained image recognition problem in a challenging few-shot learning setting, termed few-shot fine-grained recognition (FSFG). The task of FSFG requires the learning systems to build classifiers for the novel fine-grained categories from few examples (only one or less than five). To solve this problem, we propose an end-to-end trainable deep network, which is inspired by the state-of-the-art fine-grained recognition model and is tailored for the FSFG task. Specifically, our network consists of a bilinear feature learning module and a classifier mapping module: while the former encodes the discriminative information of an exemplar image into a feature vector, the latter maps the intermediate feature into the decision boundary of the novel category. The key novelty of our model is a "piecewise mappings" function in the classifier mapping module, which generates the decision boundary via learning a set of more attainable sub-classifiers in a more parameter-economic way. We learn the exemplar-to-classifier mapping based on an auxiliary dataset in a meta-learning fashion, which is expected to be able to generalize to novel categories. By conducting comprehensive experiments on three fine-grained datasets, we demonstrate that the proposed method achieves superior performance over the competing baselines. Xiu-Shen Wei, Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Jianxin Wu 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Visual Question Answering With Memory-Augmented NetworksabstractIn this paper, we exploit memory-augmented neural networks to predict accurate answers to visual questions, even when those answers rarely occur in the training set. The memory network incorporates both internal and external memory blocks and selectively pays attention to each training exemplar. We show that memory-augmented neural networks are able to maintain a relatively long-term memory of scarce training exemplars, which is important for visual question answering due to the heavy-tailed distribution of answers in a general VQA setting. Experimental results in two large-scale benchmark datasets show the favorable performance of the proposed algorithm with the comparison to state of the art. Chao Ma 0004, Chunhua Shen, Anthony R. Dick, Qi Wu 0001, Peng Wang 0023, Anton van den Hengel, Ian D. Reid 0001 |
CVPR | 5 |
| 2018 | Two-Stream 3-D convNet Fusion for Action Recognition in Videos With Arbitrary Size and Lengthabstract3-D convolutional neural networks (3-D-convNets) have been very recently proposed for action recognition in videos, and promising results are achieved. However, existing 3-D-convNets has two “artificial” requirements that may reduce the quality of video analysis: 1) It requires a fixed-sized (e.g., 112 $\times$ 112) input video; and 2) most of the 3-D-convNets require a fixed-length input (i.e., video shots with fixed number of frames). To tackle these issues, we propose an end-to-end pipeline named Two-stream 3-D-convNet Fusion, which can recognize human actions in videos of arbitrary size and length using multiple features. Specifically, we decompose a video into spatial and temporal shots. By taking a sequence of shots as input, each stream is implemented using a spatial temporal pyramid pooling (STPP) convNet with a long short-term memory (LSTM) or CNN-E model, softmax scores of which are combined by a late fusion. We devise the STPP convNet to extract equal-dimensional descriptions for each variable-size shot, and we adopt the LSTM/CNN-E model to learn a global description for the input video using these time-varying descriptions. With these advantages, our method should improve all 3-D CNN-based video analysis methods. We empirically evaluate our method for action recognition in videos and the experimental results show that our method outperforms the state-of-the-art methods (both 2-D and 3-D based) on three standard benchmark datasets (UCF101, HMDB51 and ACT datasets). Xuanhan Wang, Lianli Gao, Peng Wang 0023, Xiaoshuai Sun, Xianglong Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2017 | Event Video Mashup: From Hundreds of Videos to Minutes of SkeletonabstractThe explosive growth of video content on the Web has been revolutionizing the way people share, exchange and perceive information, such as events. While an individual video usually concerns a specific aspect of an event, the videos that are uploaded by different users at different locations and times can embody different emphasis and compensate each other in describing the event. Combining these videos from different sources together can unveil a more complete picture of the event. Simply concatenating videos together is an intuitive solution, but it may degrade user experience since it is time-consuming and tedious to view those highly redundant, noisy and disorganized content. Therefore, we develop a novel approach, termed event video mashup (EVM), to automatically generate a unified short video from a collection of Web videos to describe the storyline of an event. We propose a submodular based content selection model that embodies both importance and diversity to depict the event from comprehensive aspects in an efficient way. Importantly, the video content is organized temporally and semantically conforming to the event evolution. We evaluate our approach on a real-world YouTube event dataset collected by ourselves. The extensive experimental results demonstrate the effectiveness of the proposed framework. Lianli Gao, Peng Wang 0023, Jingkuan Song, Zi Huang, Jie Shao 0001, Heng Tao Shen |
AAAI | 2 |
| 2017 | Multi-attention Network for One Shot LearningabstractOne-shot learning is a challenging problem where the aim is to recognize a class identified by a single training image. Given the practical importance of one-shot learning, it seems surprising that the rich information present in the class tag itself has largely been ignored. Most existing approaches restrict the use of the class tag to finding similar classes and transferring classifiers or metrics learned thereon. We demonstrate here, in contrast, that the class tag can inform one-shot learning as a guide to visual attention on the training image for creating the image representation. This is motivated by the fact that human beings can better interpret a training image if the class tag of the image is understood. Specifically, we design a neural network architecture which takes the semantic embedding of the class tag to generate attention maps and uses those attention maps to create the image features for one-shot learning. Note that unlike other applications, our task requires that the learned attention generator can be generalized to novel classes. We show that this can be realized by representing class tags with distributed word embeddings and learning the attention map generator from an auxiliary training set. Also, we design a multiple-attention scheme to extract richer information from the exemplar image and this leads to substantial performance improvement. Through comprehensive experiments, we show that the proposed approach leads to superior performance over the baseline methods. Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Zi Huang, Anton van den Hengel, Heng Tao Shen |
CVPR | 1 |
| 2017 | Compositional Model Based Fisher Vector Coding for Image ClassificationabstractDeriving from the gradient vector of a generative model of local features, Fisher vector coding (FVC) has been identified as an effective coding method for image classification. Most, if not all, FVC implementations employ the Gaussian mixture model (GMM) as the generative model for local features. However, the representative power of a GMM can be limited because it essentially assumes that local features can be characterized by a fixed number of feature prototypes, and the number of prototypes is usually small in FVC. To alleviate this limitation, in this work, we break the convention which assumes that a local feature is drawn from one of a few Gaussian distributions. Instead, we adopt a compositional mechanism which assumes that a local feature is drawn from a Gaussian distribution whose mean vector is composed as a linear combination of multiple key components, and the combination weight is a latent random variable. In doing so we greatly enhance the representative power of the generative model underlying FVC. To implement our idea, we design two particular generative models following this compositional approach. In our first model, the mean vector is sampled from the subspace spanned by a set of bases and the combination weight is drawn from a Laplace distribution. In our second model, we further assume that a local feature is composed of a discriminative part and a residual part. As a result, a local feature is generated by the linear combination of discriminative part bases and residual part bases. The decomposition of the discriminative and residual parts is achieved via the guidance of a pre-trained supervised coding method. By calculating the gradient vector of the proposed models, we derive two new Fisher vector coding strategies. The first is termed Sparse Coding-based Fisher Vector Coding (SCFVC) and can be used as the substitute of traditional GMM based FVC. The second is termed Hybrid Sparse Coding-based Fisher vector coding (HSCFVC) since it combines the merits of both pre-trained supervised coding methods and FVC. Using pre-trained Convolutional Neural Network (CNN) activations as local features, we experimentally demonstrate that the proposed methods are superior to traditional GMM based FVC and achieve state-of-the-art performance in various image classification tasks. Lingqiao Liu, Peng Wang 0023, Chunhua Shen, Lei Wang 0001, Anton van den Hengel, Chao Wang 0063, Heng Tao Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Temporal Pyramid Pooling-Based Convolutional Neural Network for Action RecognitionabstractEncouraged by the success of convolutional neural networks (CNNs) in image classification, recently much effort is spent on applying the CNNs to the video-based action recognition problems. One challenge is that a video contains a varying number of frames, which is incompatible to the standard input format of the CNNs. Existing methods handle this issue either by directly sampling a fixed number of frames or bypassing this issue by introducing a 3D convolutional layer, which conducts convolution in spatial-temporal domain. In this paper, we propose a novel network structure, which allows an arbitrary number of frames as the network input. The key to our solution is to introduce a module consisting of an encoding layer and a temporal pyramid pooling layer. The encoding layer maps the activation from the previous layers to a feature vector suitable for pooling, whereas the temporal pyramid pooling layer converts multiple frame-level activations into a fixed-length video-level representation. In addition, we adopt a feature concatenation layer that combines the appearance and motion information. Compared with the frame sampling strategy, our method avoids the risk of missing any important frames. Compared with the 3D convolutional method, which requires a huge video data set for network training, our model can be learned on a small target data set because we can leverage the off-the-shelf image-level CNN for model parameter initialization. Experiments on three challenging data sets, Hollywood2, HMDB51, and UCF101 demonstrate the effectiveness of the proposed network. Peng Wang 0023, Yuanzhouhan Cao, Chunhua Shen, Lingqiao Liu, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | What's Wrong with That Object? Identifying Images of Unusual Objects by Modelling the Detection Score DistributionabstractThis paper studies the challenging problem of identifying unusual instances of known objects in images within an "open world" setting. That is, we aim to find objects that are members of a known class, but which are not typical of that class. Thus the "unusual object" should be distinguished from both the "regular object" and the "other objects". Such unusual objects may be of interest in many applications such as surveillance or quality control. We propose to identify unusual objects by inspecting the distribution of object detection scores at multiple image regions. The key observation motivating our approach is that "regular object" images, "unusual object" images and "other objects" images exhibit different region-level scores in terms of both the score values and the spatial distributions. To model these distributions we propose to use Gaussian Processes (GP) to construct two separate generative models, one for the "regular object" and the other for the "other objects". More specifically, we design a new covariance function to simultaneously model the detection score at a single location and the score dependencies between multiple regions. We demonstrate that the proposed approach outperforms comparable methods on a new large dataset constructed for the purpose. Peng Wang 0023, Lingqiao Liu, Chunhua Shen, Zi Huang, Anton van den Hengel, Heng Tao Shen |
CVPR | 1 |
| 2016 | Quartet-net Learning for Visual Instance RetrievalabstractRecently, neuron activations extracted from a pre-trained convolutional neural network (CNN) show promising performance in various visual tasks. However, due to the domain and task bias, using the features generated from the model pre-trained for image classification as image representations for instance retrieval is problematic. In this paper, we propose quartet-net learning to improve the discriminative power of CNN features for instance retrieval. The general idea is to map the features into a space where the image similarity can be better evaluated. Our network differs from the traditional Siamese-net in two ways. First, we adopt a double-margin contrastive loss with a dynamic margin tuning strategy to train the network which leads to more robust performance. Second, we introduce in the mimic learning regularization to improve the generalization ability of the network by preventing it from overfitting to the training data. Catering for the network learning, we collect a large-scale dataset, namely GeoPair, which consists of 68k matching image pairs and 63k non-matching pairs. Experiments on several standard instance retrieval datasets demonstrate the effectiveness of our method. Jiewei Cao, Zi Huang, Peng Wang 0023, Chao Li 0034, Xiaoshuai Sun, Heng Tao Shen |
ACM Multimedia | 3 |
| 2016 | Web Video Event Recognition by Semantic Analysis From Ubiquitous DocumentsabstractIn recent years, the task of event recognition from videos has attracted increasing interest in multimedia area. While most of the existing research was mainly focused on exploring visual cues to handle relatively small-granular events, it is difficult to directly analyze video content without any prior knowledge. Therefore, synthesizing both the visual and semantic analysis is a natural way for video event understanding. In this paper, we study the problem of Web video event recognition, where Web videos often describe large-granular events and carry limited textual information. Key challenges include how to accurately represent event semantics from incomplete textual information and how to effectively explore the correlation between visual and textual cues for video event understanding. We propose a novel framework to perform complex event recognition from Web videos. In order to compensate the insufficient expressive power of visual cues, we construct an event knowledge base by deeply mining semantic information from ubiquitous Web documents. This event knowledge base is capable of describing each event with comprehensive semantics. By utilizing this base, the textual cues for a video can be significantly enriched. Furthermore, we introduce a two-view adaptive regression model, which explores the intrinsic correlation between the visual and textual cues of the videos to learn reliable classifiers. Extensive experiments on two real-world video data sets show the effectiveness of our proposed framework and prove that the event knowledge base indeed helps improve the performance of Web video event recognition. Litao Yu, Yang Yang 0002, Zi Huang, Peng Wang 0023, Jingkuan Song, Heng Tao Shen |
IEEE Trans. Image Process. | 4 |
| 2014 | WeMash: An Online System for Web Video MashupabstractThe explosive growth of video content on the Web has been revolutionizing the way people share, exchange and perceive information. At the same time, however, the massive amount of Web video content may degrade user experience in that it is tedious and time-consuming to view relevant videos uploaded by different users and grasp the gist of their content. In this demonstration, we demonstrate a system that takes Web videos related to a query event as the input and automatically generates a short mashup video that can well describe the essence of the query. Specifically, we first identify shot cliques composed of near-duplicate shots and infer their semantic meanings by resorting to relevant Web images. Then a shot clique selection algorithm is performed to select representative content according to the twofold importance of the shot cliques evaluated, namely content importance and semantic importance, with content diversity considered. Finally, the selected shot cliques are organized into a mashup video based on semantic and temporal cues. Peng Wang 0023, Yang Yang 0002, Zi Huang, Jiewei Cao, Heng Tao Shen |
ACM Multimedia | 1 |