VLDB 2026 Research / reviewers in the wild / expert
Yibing Zhan
dblp:142/8486
· DBLP profile ↗
132ranked-venue papers
9as first author
119since 2021 · last 2026
0000-0003-3180-0484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 84 · 2 first-author · 82 since 2021Graphics, computer vision, multimedia, augmented reality and games · 62 · 8 first-author · 51 since 2021Databases, data management, data science and information retrieval · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Sample Augmented Test-Time Adaptation for Personalized Intraoperative Hypotension PredictionabstractIntraoperative hypotension (IOH) poses significant surgical risks, but accurate prediction remains challenging due to patient-specific variability. While test-time adaptation (TTA) offers a promising approach for personalized prediction, the rarity of IOH events often leads to unreliable test-time training. To address this, we propose CSA-TTA, a novel cross-sample augmented test-time adaptation framework that enhances training by incorporating hypotension events from other individuals. Specifically, we first construct a cross-sample bank by segmenting historical data into hypotensive and non-hypotensive samples. Then, we introduce a coarse-to-fine retrieval strategy for building test-time training data: we initially apply K-Shape clustering to identify representative cluster centers and subsequently retrieve the top-K semantically similar samples based on the current patient signal. Additionally, we integrate both self-supervised masked reconstruction and retrospective sequence forecasting signals during training to enhance model adaptability to rapid and subtle intraoperative dynamics. We evaluate the proposed CSA-TTA on both the VitalDB dataset and a real-world in-hospital dataset by integrating it with state-of-the-art time series forecasting models, including TimesFM and UniTS. CSA-TTA consistently enhances performance across settings—for instance, on VitalDB, it improves Recall and F1 scores by +1.33% and +1.13%, respectively, under fine-tuning, and by +7.46% and +5.07% in zero-shot scenarios—demonstrating strong robustness and generalization. Kanxue Li, Yibing Zhan, Chongchong Qi, Baosheng Yu |
AAAI | 2 |
| 2026 | ProGMLP: A Progressive Framework for GNN-to-MLP Knowledge Distillation with Efficient Trade-offsabstractGNN-to-MLP (G2M) methods have emerged as a promising approach to accelerate Graph Neural Networks (GNNs) by distilling their knowledge into simpler Multi-Layer Perceptrons (MLPs). These methods bridge the gap between the expressive power of GNNs and the computational efficiency of MLPs, making them well-suited for resource-constrained environments. However, existing G2M methods are limited by their inability to flexibly adjust inference cost and accuracy dynamically, a critical requirement for real-world applications where computational resources and time constraints can vary significantly. To address this, we introduce a Progressive framework designed to offer flexible and on-demand trade-offs between inference cost and accuracy for GNN-to-MLP knowledge distillation (ProGMLP). ProGMLP employs a Progressive Training Structure (PTS), where multiple MLP students are trained in sequence, each building on the previous one. Furthermore, ProGMLP incorporates Progressive Knowledge Distillation (PKD) to iteratively refine the distillation process from GNNs to MLPs, and Progressive Mixup Augmentation (PMA) to enhance generalization by progressively generating harder mixed samples. Our approach is validated through comprehensive experiments on eight real-world graph datasets, demonstrating that ProGMLP maintains high accuracy while dynamically adapting to varying runtime scenarios, making it highly effective for deployment in diverse application settings. Weigang Lu 0001, Ziyu Guan, Wei Zhao 0019, Yaming Yang 0002, Yibing Zhan, Dapeng Tao |
AAAI | 7 |
| 2026 | CERA: Conflict-Explicit Reflective Agent for Multimodal Emotion ReasoningabstractMultimodal Emotion Recognition (MER) aims to understand complex human emotions by jointly analyzing visual and textual data. However, in real-world scenarios, emotional cues from different modalities often contain conflict information, such as a smiling face paired with negative text, which poses great challenges for existing multimodal language models (MLLMs). Existing emotion MLLMs and multimodal emotion benchmarks often overlook or even intentionally avoid scenarios involving multimodal emotion conflicts, limiting their ability to reason about complex and contradictory affective cues. By addressing this, we propose Conflict-Explicit Reflective Agent (CERA), a training-free, conflict-aware, and language-driven agentic framework for MER. The concept of CERA is to treat modality emotion conflicts as meaningful signals and resolve them via a three-stage perception–evaluation–reflection reasoning loop. Firstly, the agent’s conflict-perceptive emotion graph construction module builds emotion graphs from fine-grained cues to reveal conflicts, and progressively refines them through iterative updates. Secondly, a reward model evaluates these graphs and produces natural language feedback that identifies unresolved conflicts. Lastly, the language-driven conflict refinement module generates graph editing signals from the feedback without any parameter tuning, enabling the overall CERA to refine its reasoning without training. Extensive experiments on two multimodal emotion datasets, MAFW and CH-SIMS, demonstrate that CERA significantly outperforms state-of-the-art training-free methods in both recognition accuracy and conflict interpretability, providing an effective training-free solution for complex emotional reasoning. Kejun Liu, Chang Tang, Zhe Chen 0013, Yibing Zhan |
ICMR | 8 |
| 2026 | Improving zero-shot translation with the navigation ability-enhanced language tags
Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibin Lei, Yibing Zhan, Weifeng Liu 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Cross-modal attention fusion and label co-occurrence feature enhancement for multi-label postoperative adverse reaction prediction
Jierong Li, Yibing Zhan, Dapeng Tao, Chongchong Qi |
Expert Syst. Appl. | 3 |
| 2026 | Evaluating large language models for real-world perioperative clinical consultation
Yibing Zhan, Baosheng Yu, Pingbo Xu, Lijing Chen, Chong Zhang 0013, Chengli Zhou, Xiongbin Wang, Dapeng Tao |
Neurocomputing | 2 |
| 2026 | Towards alleviating hallucination in text-to-image retrieval for CLIP in zero-shot learning
Hanyao Wang, Yibing Zhan, Liu Liu 0014, Liang Ding 0006, Jun Yu 0002 |
Neurocomputing | 2 |
| 2026 | Focusing on pedestrians like human for clothes changing person re-identificationabstract• Based on the ensemble coding hypothesis in cognitive neuroscience, we achieve data augmentation by simulating human focus capability. • Our method is the first local detail learning data augmentation designed specifically for clothes changing person re-identification. • Our method achieves SOTA on three public datasets and outperforms various traditional data augmentation methods. Current approaches focus mainly on the design of networks to learn key identity features from local body components for clothes-changing person re-identification (CC-ReID). In this paper, we propose a humanoid focus-inspired image augmentation (HFIA) method, which is intuitive image processing rather than a sophisticated network architecture designed to enhance local nuances of pedestrian images. Based on pedestrian silhouettes, we roughly divide a pedestrian image into five body components, that is, head-shoulder, upper left torso, upper right torso, lower left torso, and lower right torso. The HFIA has two key designs to deal with these components: the central emphasis strategy (CES) and the component continuity processing (CCP). For each component, leveraging the natural tendency of human visual attention towards central regions, the CES constructs an enlargement grid, where the closer the center, the greater the enlargement. To maintain the continuity of assembly, the CCP performs an overall alignment of component centers, that is, all components share the same normalized vertical coordinate and the left and right torsos have mirrored horizontal coordinates. Furthermore, the CCP implements a smoothing post-processing to uniformly erase the discontinuity between the head-shoulder, upper left torso, and upper right torso. Experiments show the state-of-the-art performance of HFIA. Wenjie Pan, Jianqing Zhu, Xiaolin Cui, Huanqiang Zeng, Yibing Zhan |
Neural Networks | 5 |
| 2026 | Contrastive knowledge embedding with discriminative self-weighted sampling
Sheng Wan, Yibing Zhan, Shirui Pan, Jian Yang 0003, Chen Gong 0002 |
Neural Networks | 2 |
| 2026 | DA-MoE: Addressing depth-sensitivity in graph-level analysis through mixture of experts
Zelin Yao, Mukun Chen, Chuang Liu 0008, Xianke Meng, Yibing Zhan, Jia Wu 0001, Shirui Pan, Huiting Xu, Wenbin Hu 0001 |
Neural Networks | 5 |
| 2026 | Synergistic knowledge distillation via reciprocal and self learning
Renjie Huang, Jianping Gou, Yibing Zhan, Zhang Yi 0001 |
Pattern Recognit. | 6 |
| 2026 | Mind the data: Evaluating data quality sensitivity in medical LLMs
Xiaodong Han, Yibing Zhan, Baosheng Yu, Dapeng Tao |
Pattern Recognit. Lett. | 2 |
| 2025 | Modeling All Response Surfaces in One for Conditional Search SpacesabstractBayesian Optimization (BO) is a sample-efficient black-box optimizer commonly used in search spaces where hyperparameters are independent. However, in many practical AutoML scenarios, there will be dependencies among hyperparameters, forming a conditional search space, which can be partitioned into structurally distinct subspaces. The structure and dimensionality of hyperparameter configurations vary across these subspaces, challenging the application of BO. Some previous BO works have proposed solutions to develop multiple Gaussian Process models in these subspaces. However, these approaches tend to be inefficient as they require a substantial number of observations to guarantee each GP's performance and cannot capture relationships between hyperparameters across different subspaces. To address these issues, this paper proposes a novel approach to model the response surfaces of all subspaces in one, which can model the relationships between hyperparameters elegantly via a self-attention mechanism. Concretely, we design a structure-aware hyperparameter embedding to preserve the structural information. Then, we introduce an attention-based deep feature extractor, capable of projecting configurations with different structures from various subspaces into a unified feature space, where the response surfaces can be formulated using a single standard Gaussian Process. The empirical results on a simulation function, various real-world tasks, and HPO-B benchmark demonstrate that our proposed approach improves the efficacy and efficiency of BO within conditional search spaces. Wei Liu 0005, Chao Xue 0003, Yibing Zhan, Xiaoxing Wang, Weifeng Liu 0001, Dacheng Tao |
AAAI | 4 |
| 2025 | AGMixup: Adaptive Graph Mixup for Semi-supervised Node ClassificationabstractMixup is a data augmentation technique that enhances model generalization by interpolating between data points using a mixing ratio lambda in the image domain. Recently, the concept of mixup has been adapted to the graph domain through node-centric interpolations. However, these approaches often fail to address the complexity of interconnected relationships, potentially damaging the graph's natural topology and undermining node interactions. Furthermore, current graph mixup methods employ a one-size-fits-all strategy with a randomly sampled lambda for all mixup pairs, ignoring the diverse needs of different pairs. This paper proposes an Adaptive Graph Mixup (AGMixup) framework for semi-supervised node classification. AGMixup introduces a subgraph-centric approach, which treats each subgraph similarly to how images are handled in Euclidean domains, thus facilitating a more natural integration of mixup into graph-based learning. We also propose an adaptive mechanism to tune the mixing ratio lambda for diverse mixup pairs, guided by the contextual similarity and uncertainty of the involved subgraphs. Extensive experiments across seven datasets on semi-supervised node classification benchmarks demonstrate AGMixup's superiority over state-of-the-art graph mixup methods. Weigang Lu 0001, Ziyu Guan, Wei Zhao 0019, Yaming Yang 0002, Yibing Zhan, Yiheng Lu, Dapeng Tao |
AAAI | 5 |
| 2025 | Beyond Human Data: Aligning Multimodal Large Language Models by Iterative Self-EvolutionabstractHuman preference alignment can significantly enhance the capabilities of Multimodal Large Language Models (MLLMs). However, collecting high-quality preference data remains costly. One promising solution is the self-evolution strategy, where models are iteratively trained on data they generate. Current multimodal self-evolution techniques, nevertheless, still need human- or GPT-annotated data. Some methods even require extra models or ground truth answers to construct preference data. To overcome these limitations, we propose a novel multimodal self-evolution framework that empowers the model to autonomously generate high-quality questions and answers using only unannotated images. First, in the question generation phase, we implement an image-driven self-questioning mechanism. This approach allows the model to create questions and evaluate their relevance and answerability based on the image content. If a question is deemed irrelevant or unanswerable, the model regenerates it to ensure alignment with the image. This process establishes a solid foundation for subsequent answer generation and optimization. Second, while generating answers, we design an answer self-enhancement technique to boost the discriminative power of answers. We begin by captioning the images and then use the descriptions to enhance the generated answers. Additionally, we utilize corrupted images to generate rejected answers, thereby forming distinct preference pairs for effective optimization. Finally, in the optimization step, we incorporate an image content alignment loss function alongside the Direct Preference Optimization (DPO) loss to mitigate hallucinations. This function maximizes the likelihood of the above generated descriptions in order to constrain the model's attention to the image content. As a result, model can generate more accurate and reliable outputs. Experiments demonstrate that our framework is competitively compared with previous methods that utilize external information, paving the way for more efficient and scalable MLLMs. Wentao Tan, Qiong Cao, Yibing Zhan, Chao Xue 0003, Changxing Ding |
AAAI | 3 |
| 2025 | Improving Complex Reasoning over Knowledge Graph with Logic-Aware Curriculum TuningabstractAnswering complex queries over incomplete knowledge graphs (KGs) is a challenging job. Most previous works have focused on learning entity/relation embeddings and simulating first-order logic operators with various neural networks. However, they are bottlenecked by the inability to share world knowledge to improve logical reasoning, thus resulting in suboptimal performance. In this paper, we propose a complex reasoning schema over KG upon large language models (LLMs), containing a curriculum-based logical-aware instruction tuning framework, named LACT. Specifically, we augment the arbitrary first-order logical queries via binary tree decomposition, to stimulate the reasoning capability of LLMs. To address the difficulty gap among different types of complex queries, we design a simple and flexible logic-aware curriculum learning framework. Experiments across widely used datasets demonstrate that LACT has substantial improvements~(brings an average +5.5% MRR score) over advanced methods, achieving the new state-of-the-art. Tianle Xia, Liang Ding 0006, Guojia Wan, Yibing Zhan, Bo Du 0001, Dacheng Tao |
AAAI | 4 |
| 2025 | Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated ReasoningabstractRecently, increasing attention has been focused on improving the ability of Large Language Models (LLMs) to perform complex reasoning. Advanced methods, such as Chain-of-Thought (CoT) and its variants, are found to enhance their reasoning skills by designing suitable prompts or breaking down complex problems into more manageable sub-problems. However, little concentration has been put on exploring the reasoning process, i.e., we discovered that most methods resort to Direct Reasoning (DR) and disregard Indirect Reasoning (IR). This can make LLMs difficult to solve IR tasks, which are often encountered in the real world. To address this issue, we propose a Direct-Indirect Reasoning (DIR) method, which considers DR and IR as multiple parallel reasoning paths that are merged to derive the final answer. We stimulate LLMs to implement IR by crafting prompt templates incorporating the principles of contrapositive and contradiction. These templates trigger LLMs to assume the negation of the conclusion as true, combine it with the premises to deduce a conclusion, and utilize the logical equivalence of the contrapositive to enhance their comprehension of the rules used in the reasoning process. Our DIR method is simple yet effective and can be straightforwardly integrated with existing variants of CoT methods. Experimental results on four datasets related to logical reasoning and mathematic proof demonstrate that our DIR method, when combined with various baseline methods, significantly outperforms all the original methods. Yanfang Zhang 0001, Yiliu Sun, Yibing Zhan, Dapeng Tao, Dacheng Tao, Chen Gong 0002 |
COLING | 3 |
| 2025 | End-to-End HOI Reconstruction Transformer with Graph-based EncodingabstractWith the diversification of human-object interaction (HOI) applications and the success of capturing human meshes, HOI reconstruction has gained widespread attention. Existing mainstream HOI reconstruction methods often rely on explicitly modeling interactions between humans and objects. However, such a way leads to a natural conflict between 3D mesh reconstruction, which emphasizes global structure, and fine-grained contact reconstruction, which focuses on local details. To address the limitations of explicit modeling, we propose the End-to-End HOI Reconstruction Transformer with Graph-based Encoding (HOI-TG). It implicitly learns the interaction between humans and objects by leveraging self-attention mechanisms. Within the transformer architecture, we devise graph residual blocks to aggregate the topology among vertices of different spatial structures. This dual focus effectively balances global and local representations. Without bells and whistles, HOI-TG achieves state-of-the-art performance on BEHAVE and InterCap datasets. Particularly on the challenging InterCap dataset, our method improves the reconstruction results for human and object meshes by 8.9% and 8.6%, respectively. Zhenrong Wang, Sihan Ma, Maosheng Ye, Yibing Zhan, Dongjiang Li |
CVPR | 5 |
| 2025 | Self-Supervised Learning for Detecting AI-Generated Faces as AnomaliesabstractThe detection of AI-generated faces is commonly approached as a binary classification task. Nevertheless, the resulting detectors frequently struggle to adapt to novel AI face generators, which evolve rapidly. In this paper, we describe an anomaly detection method for AI-generated faces by leveraging self-supervised learning of camera-intrinsic and face-specific features purely from photographic face images. The success of our method lies in designing a pretext task that trains a feature extractor to rank four ordinal exchangeable image file format (EXIF) tags and classify artificially manipulated face images. Subsequently, we model the learned feature distribution of photographic face images using a Gaussian mixture model. Faces with low likelihoods are flagged as AI-generated. Both quantitative and qualitative experiments validate the effectiveness of our method. Our code is available at https://github.com/MZMMSEC/AIGFD_EXIF.git. Mian Zou, Baosheng Yu, Yibing Zhan, Kede Ma |
ICASSP | 3 |
| 2025 | Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
Mian Zou, Nan Zhong, Baosheng Yu, Yibing Zhan, Kede Ma |
ICCV | 4 |
| 2025 | SkipNode: On Alleviating Performance Degradation for Deep Graph Convolutional Networks (Extended Abstract)abstractGraph Convolutional Networks (GCNs) are powerful tools for learning representations in graph-structured data. However, their performance tends to degrade with increased model depth due to over-smoothing. Although previous studies attribute degradation to over-smoothing, this work identifies the mutually reinforcing effects of over-smoothing and gradient vanishing as the root cause. In this paper, we propose SkipNode, a plug-and-play module that mitigates degradation in deep GCNs. SkipNode introduces node-sampling in each convolutional layer to selectively skip convolutions, preventing over-smoothing by reducing the depth experienced by specific nodes and facilitating gradient backpropagation. We demonstrate both theoretically and experimentally that SkipNode effectively curtails over-smoothing and gradient vanishing, improving deep GCN performance across diverse tasks. Extensive evaluations show SkipNode's robustness and superior performance over state-of-the-art (SOTA) baselines, establishing it as a practical solution for training deep GCNs. Weigang Lu 0001, Yibing Zhan, Binbin Lin 0001, Ziyu Guan, Liu Liu 0014, Baosheng Yu, Wei Zhao 0019, Yaming Yang 0002, Dacheng Tao |
ICDE | 2 |
| 2025 | NoVo: Norm Voting off Hallucinations with Attention Heads in Large Language ModelsabstractHallucinations in Large Language Models (LLMs) remain a major obstacle, particularly in high-stakes applications where factual accuracy is critical. While representation editing and reading methods have made strides in reducing hallucinations, their heavy reliance on specialised tools and training on in-domain samples, makes them difficult to scale and prone to overfitting. This limits their accuracy gains and generalizability to diverse datasets. This paper presents a lightweight method, Norm Voting (NoVo), which harnesses the untapped potential of attention head norms to dramatically enhance factual accuracy in zero-shot multiple-choice questions (MCQs). NoVo begins by automatically selecting truth-correlated head norms with an efficient, inference-only algorithm using only 30 random samples, allowing NoVo to effortlessly scale to diverse datasets. Afterwards, selected head norms are employed in a simple voting algorithm, which yields significant gains in prediction accuracy. On TruthfulQA MC1, NoVo surpasses the current state-of-the-art and all previous methods by an astounding margin---at least 19 accuracy points. NoVo demonstrates exceptional generalization to 20 diverse datasets, with significant gains in over 90\% of them, far exceeding all current representation editing and reading methods. NoVo also reveals promising gains to finetuning strategies and building textual adversarial defence. NoVo's effectiveness with head norms opens new frontiers in LLM interpretability, robustness and reliability. Our code is available at: https://github.com/hozhengyi/novo Zheng Yi Ho, Siyuan Liang 0004, Sen Zhang 0006, Yibing Zhan, Dacheng Tao |
ICLR | 4 |
| 2025 | NT-FAN: A simple yet effective noise-tolerant few-shot adaptation network
Wenjing Yang 0002, Haoang Chi, Yibing Zhan, Xiaoguang Ren, Dapeng Tao, Long Lan |
Artif. Intell. | 3 |
| 2025 | Semantic segmentation in power grid scenarios using scale-transforming transformer
Wenjie Pan, Linhan Huang, Yuqing Fu, Jianqing Zhu, Yibing Zhan |
Appl. Intell. | 6 |
| 2025 | Sample-Cohesive Pose-Aware Contrastive Facial Representation LearningabstractAbstract Self-supervised facial representation learning (SFRL) methods, especially contrastive learning (CL) methods, have been increasingly popular due to their ability to perform face understanding without heavily relying on large-scale well-annotated datasets. However, analytically, current CL-based SFRL methods still perform unsatisfactorily in learning facial representations due to their tendency to learn pose-insensitive features, resulting in the loss of some useful pose details. This could be due to the inappropriate positive/negative pair selection within CL. To conquer this challenge, we propose a Pose-disentangled Contrastive Facial Representation Learning (PCFRL) framework to enhance pose awareness for SFRL. We achieve this by explicitly disentangling the pose-aware features from non-pose face-aware features and introducing appropriate sample calibration schemes for better CL with the disentangled features. In PCFRL, we first devise a pose-disentangled decoder with a delicately designed orthogonalizing regulation to perform the disentanglement; therefore, the learning on the pose-aware and non-pose face-aware features would not affect each other. Then, we introduce a false-negative pair calibration module to overcome the issue that the two types of disentangled features may not share the same negative pairs for CL. Our calibration employs a novel neighborhood-cohesive pair alignment method to identify pose and face false-negative pairs, respectively, and further help calibrate them to appropriate positive pairs. Lastly, we devise two calibrated CL losses, namely calibrated pose-aware and face-aware CL losses, for adaptively learning the calibrated pairs more effectively, ultimately enhancing the learning with the disentangled features and providing robust facial representations for various downstream tasks. In the experiments, we perform linear evaluations on four challenging downstream facial tasks with SFRL using our method, including facial expression recognition, face recognition, facial action unit detection, and head pose estimation. Experimental results show that PCFRL outperforms existing state-of-the-art methods by a substantial margin, demonstrating the importance of improving pose awareness for SFRL. Our evaluation code and model will be available at https://github.com/fulaoze/CV/tree/main . Yuanyuan Liu 0004, Shaoze Feng, Yibing Zhan, Dapeng Tao, Zijing Chen, Zhe Chen 0013 |
Int. J. Comput. Vis. | 4 |
| 2025 | Noise-Resistant Multimodal Transformer for Emotion Recognition
Yuanyuan Liu 0004, Haoyu Zhang 0001, Yibing Zhan, Zijing Chen, Guanghao Yin, Zhe Chen 0013 |
Int. J. Comput. Vis. | 3 |
| 2025 | G-NodeMixup: Enhancing graph neural networks reachability under extremely limited labels
Ziyu Guan, Beilei Ling, Weigang Lu 0001, Meng Yan 0013, Yaming Yang 0002, Wei Zhao 0019, Yibing Zhan, Dapeng Tao |
Neurocomputing | 7 |
| 2025 | Hypnos: A domain-specific large language model for anesthesiology
Zhonghai Wang, Yibing Zhan, Bohao Zhou, Chong Zhang 0013, Baosheng Yu, Liang Ding 0006, Weifeng Liu 0001 |
Neurocomputing | 3 |
| 2025 | Degradation-adaptive attack-robust self-supervised facial representation learning
Yuanyuan Liu 0004, Chang Tang, Kun Sun 0002, Yibing Zhan, Zhe Chen 0013 |
Neurocomputing | 5 |
| 2025 | Building accurate translation-tailored large language models with language-aware instruction tuningabstractLarge language models (LLMs) exhibit remarkable capabilities in various natural language processing tasks, such as machine translation. However, the large number of LLM parameters incurs significant costs during inference. Previous studies have attempted to train translation-tailored LLMs with moderately sized models by fine-tuning them on the translation data. Nevertheless, when performing translations in zero-shot directions that are absent from the fine-tuning data, the problem of ignoring instructions and thus producing translations in the wrong language (i.e., the off-target translation issue) remains unresolved. In this work, we design a two-stage fine-tuning algorithm to improve the instruction-following ability of translation-tailored LLMs, particularly for maintaining accurate translation directions. We first fine-tune LLMs on the translation data to elicit basic translation capabilities. At the second stage, we construct instruction-conflicting samples by randomly replacing the instructions with the incorrect ones. Then, we introduce an extra unlikelihood loss to reduce the probability assigned to those samples. Experiments on two benchmarks using the LLaMA 2 and LLaMA 3 models, spanning 16 zero-shot directions, demonstrate that, compared to the competitive baseline—translation-finetuned LLaMA, our method could effectively reduce the off-target translation ratio (up to −62.4 percentage points), thus improving translation quality (up to +9.7 bilingual evaluation understudy). Analysis shows that our method can preserve the model’s performance on other tasks, such as supervised translation and general tasks. Code is released at https://github.com/alphadl/LanguageAware_Tuning . Changtong Zan, Liang Ding 0006, Li Shen 0008, Yibing Zhan, Xinghao Yang, Weifeng Liu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2025 | Graph explicit pooling for graph-level representation learning
Chuang Liu 0008, Wenhang Yu, Kuang Gao, Xueqi Ma, Yibing Zhan, Jia Wu 0001, Wenbin Hu 0001, Bo Du 0001 |
Neural Networks | 5 |
| 2025 | On neural architecture search and hyperparameter optimization: A max-flow based approach
Chao Xue 0003, Xiaoxing Wang, Yibing Zhan, Junchi Yan, Chun-Guang Li |
Neural Networks | 4 |
| 2025 | Learning to Explore Sample RelationshipsabstractDespite the great success achieved, deep learning technologies usually suffer from data scarcity issues in real-world applications, where existing methods mainly explore sample relationships in a vanilla way from the perspectives of either the input or the loss function. In this paper, we propose a batch transformer module, BatchFormerV1, to equip deep neural networks themselves with the abilities to explore sample relationships in a learnable way. Basically, the proposed method enables data collaboration, e.g., head-class samples will also contribute to the learning of tail classes. Considering that exploring instance-level relationships has very limited impacts on dense prediction, we generalize and refer to the proposed module as BatchFormerV2, which further enables exploring sample relationships for pixel-/patch-level dense representations. In addition, to address the train-test inconsistency where a mini-batch of data samples are neither necessary nor desirable during inference, we also devise a two-stream training pipeline, i.e., a shared model is first jointly optimized with and without BatchFormerV2 which is then removed during testing. The proposed module is plug-and-play without requiring any extra inference cost. Lastly, we evaluate the proposed method on over ten popular datasets, including 1) different data scarcity settings such as long-tailed recognition, zero-shot learning, domain generalization, and contrastive learning; and 2) different visual recognition tasks ranging from image classification to object detection and panoptic segmentation. Zhi Hou, Baosheng Yu, Yibing Zhan, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Aligning Text-to-Image Diffusion Models With Constrained Reinforcement LearningabstractReward finetuning has emerged as a powerful technique for aligning diffusion models with specific downstream objectives or user preferences. However, current approaches suffer from a persistent challenge of reward overoptimization, where models exploit imperfect reward feedback at the expense of overall performance. In this work, we identify three key contributors to overoptimization: (1) a granularity mismatch between the multi-step diffusion process and sparse rewards; (2) a loss of plasticity that limits the model's ability to adapt and generalize; and (3) an overly narrow focus on a single reward objective that neglects complementary performance criteria. Accordingly, we introduce Constrained Diffusion Policy Optimization (CDPO), a novel reinforcement learning framework that addresses reward overoptimization from multiple angles. Firstly, CDPO tackles the granularity mismatch through a temporal policy optimization strategy that delivers step-specific rewards throughout the entire diffusion trajectory, thereby reducing the risk of overfitting to sparse final-step rewards. Then we incorporate a neuron reset strategy that selectively resets overactive neurons in the model, preventing overoptimization induced by plasticity loss. Finally, to avoid overfitting to a narrow reward objective, we integrate constrained reinforcement learning with auxiliary reward objectives serving as explicit constraints, ensuring a balanced optimization across diverse performance metrics. Ziyi Zhang 0001, Sen Zhang 0006, Li Shen 0008, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Distilling interaction knowledge for semi-supervised egocentric action recognition
Haoran Wang 0001, Baosheng Yu, Yibing Zhan, Dapeng Tao, Haibin Ling |
Pattern Recognit. | 4 |
| 2025 | Graph Convolutional Networks With Collaborative Feature Fusion for Sequential RecommendationabstractSequential recommendation seeks to understand user preferences based on their past actions and predict future interactions with items. Recently, several techniques for sequential recommendation have emerged, primarily leveraging graph convolutional networks (GCNs) for their ability to model relationships effectively. However, real-world scenarios often involve sparse interactions, where early and recent short-term preferences play distinct roles in the recommendation process. Consequently, vanilla GCNs struggle to effectively capture the explicit correlations between these early and recent short-term preferences. To address these challenges, we introduce a novel approach termed Graph Convolutional Networks with Collaborative Feature Fusion (COFF). Specifically, our method addresses the issue by initially dividing each user interaction sequence into two segments. We then construct two separate graphs for these segments, aiming to capture the user's early and recent short-term preferences independently. To obtain robust prediction, we employ multiple GCNs in a collaborative distillation manner, incorporating a feature fusion module to establish connections between the early and recent short-term preferences. This approach enables a more precise representation of user preferences. Experimental evaluations conducted on five popular sequential recommendation datasets demonstrate that our COFF model outperforms recent state-of-the-art methods in terms of recommendation accuracy. Jianping Gou, Youhui Cheng, Yibing Zhan, Baosheng Yu, Weihua Ou, Zhang Yi 0001 |
IEEE Trans. Big Data | 3 |
| 2025 | Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding ApproachabstractIn recent years, the multimedia forensics and security community has seen remarkable progress in multitask learning for DeepFake (i.e., face forgery) detection. The prevailing approach has been to frame DeepFake detection as a binary classification problem augmented by manipulation-oriented auxiliary tasks. This scheme focuses on learning features specific to face manipulations with limited generalizability. In this paper, we delve deeper into semantics-oriented multitask learning for DeepFake detection, capturing the relationships among face semantics via joint embedding. We first propose an automated dataset expansion technique that broadens current face forgery datasets to support semantics-oriented DeepFake detection tasks at both the global face attribute and local face region levels. Furthermore, we resort to the joint embedding of face images and labels (depicted by text descriptions) for prediction. This approach eliminates the need for manually setting task-agnostic and task-specific parameters, which is typically required when predicting multiple labels directly from images. In addition, we employ bi-level optimization to dynamically balance the fidelity loss weightings of various tasks, making the training process fully automated. Extensive experiments on six DeepFake datasets show that our method improves the generalizability of DeepFake detection and renders some degree of model interpretation by providing human-understandable explanations. Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Fuzzy-Assisted Contrastive Decoding Improving Code Generation of Large Language ModelsabstractLarge Language Models (LLMs) play a crucial role in intelligent code generation tasks. Most existing work focuses on pre-training or fine-tuning specialized code LLMs, e.g., CodeLlama. However, pre-training or fine-tuning a code LLM requires a vast corpus of data, significant computational resources, and considerable human effort. Compared to pre-training or fine-tuning LLMs, a simple and flexible method of contrastive decoding has garnered widespread attention to improve the text generation quality of LLMs. While contrastive decoding can indeed improve the text generation quality of LLMs, our research has found that directly using contrastive decoding: 1) introduces erroneous information into the logit distribution generated from normal prompts (i.e., user's input), particularly in the code generation of LLMs; 2) significantly impedes the inference and decoding time of LLMs. In this work, the limitations of using contrastive decoding directly are systematically highlighted, and a novel real-time fuzzy-assisted contrastive decoding (FCD) mechanism is proposed to improve the code generation quality of LLMs. The proposed FCD mechanism initially categorises prompts into high-quality and low-quality groups based on the results of the evaluator (i.e., unit test) before integrating the LLM. Next, feature values (e.g., standard deviation, peak value, etc.) related to the logit distribution of predicted tokens during the LLM's inference process for both high-quality and low-quality prompts are extracted. Finally, the extracted feature values are used to train the fuzzy neural network (i.e, fuzzy min-max neural network) offline, allowing for the prejudgement of the reliability of the logit distribution for normal prompt outputs. This prevents the direct use of erroneous information from contrastive decoding and improves the code generation quality of LLMs. Through extensive experiments, it has been demonstrated that the proposed FCD mechanism can significantly improve the code generation quality of LLMs through fuzzy-assisted contrastive decoding. Moreover, the FCD mechanism can also reduce the time required for inference and contrastive decoding. The code and data are publicly available on GitHub11https://github.com/LLMcodegen/Fuzzy_contrastive_decoding.and HuggingFace22https://huggingface.co/wangle123/Fuzzy_contrastive_decoding.. Shuai Wang 0011, Liang Ding 0006, Yibing Zhan, Yong Luo 0002, Shuai Liu 0002, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | PMTSeg: Prompt-Driven Multimodal Transformer for Task-Adapted Remote Sensing Image SegmentationabstractMultimodal remote sensing image segmentation (MRSIS) is important for intelligent remote sensing image (RS) interpretation, which encompasses three distinct tasks: semantic segmentation, instance segmentation, and panoptic segmentation. Existing methods typically address individual tasks with specialized models, limiting generalization and real-world applicability. Multi-task learning approaches have introduced separated task heads to unify tasks, yet we identify two key challenges when directly applying them to MRSIS: (1) the modality gap, arising from semantic discrepancies and granularity discrepancies across RS modalities, and (2) the task gap, due to varying preferences in learning different segmentation tasks. To overcome these challenges, we propose PMTSeg—a novel Prompt-driven Multimodal Transformer for task-adapted MRSIS. PMTSeg integrates three key components: (1) Task-common Multimodal Affinity Approximation (TMAA), (2) Task-common Multi-scale Semantic Fusion (TMSF), and (3) a unified Prompt-driven Segmentation Head (PSH). First, TMAA addresses the modality gap by approximating inter-modal affinity matrices, extracting task-common features across modalities and aligning semantic information. Then, TMSF further integrates these features using the scale-matched fusion at multiple scales to produce enriched, multi-scale task-common features. Moreover, to address the task gap, the PSH leverages task-adapted text prompts and task-adapted contrastive loss to model relationships across tasks, enabling adaptive optimization for robust and universal MRSIS performance. Extensive experiments on three MRSIS datasets—VALID, SEMCITY TOULOUSE, and UBCV2—demonstrate that PMTSeg significantly surpasses state-of-the-art methods in all three segmentation tasks, offering a unified and accurate solution to MRSIS. Kejun Liu, Xuesong Yan 0001, Yuanyuan Liu 0004, Chang Tang, Yibing Zhan, Wujie Zhou, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Parameter-Free Spectral-Spatial Optimization Algorithm for Semiblind Hyperspectral and Multispectral Image FusionabstractSemiblind fusion of hyperspectral images (HSIs) and multispectral images (MSIs) is a critical technique for generating high-resolution HSIs (HR-HSIs) without the need for complex point spread function (PSF) estimation. Despite the broad application potential of semiblind fusion algorithms, existing methods face three major challenges. First, as imaging technology advances, the spatial resolution gap between HSIs and MSIs continues to widen, making high-magnification fusion increasingly urgent. Second, especially for deep learning-based methods, existing algorithms require meticulous hyperparameter tuning to enhance fusion quality. Third, the complex fusion process impedes the speed of image fusion. To address these challenges, we propose a parameter-free spectral-spatial optimization algorithm specifically designed to handle high-magnification differences in the semiblind fusion of HSI and MSI. This method enables fast fusion using simple matrix operations and consists of three main steps: 1) rapidly computing an initial solution using the Moore-Penrose inverse of the spatial response function (SRF); 2) constructing spectral errors from the initial solution to effectively extract spectral information; and 3) reconstructing spatial details using MSI to form spatial errors, thereby accurately reconstructing HR-HSI for high-quality fusion. Comparative experiments on two simulated and three real datasets with state-of-the-art algorithms demonstrate that our proposed semiblind fusion method not only achieves$64\times $high-magnification fusion but also reduces the computation time to just 23.9%–49.5% of that required by the fastest competing methods. The code is available athttps://github.com/Long-ji/PFSSOA. Jialin Gui, Yuanxi Peng, Yibing Zhan, Jun Li 0094, Yulei Tian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | SCAWaveNet: A Spatial-Channel Attention-Based Network for Global Significant Wave Height RetrievalabstractRecent advancements in spaceborne GNSS missions have produced extensive global datasets, providing a robust basis for deep learning-based significant wave height (SWH) retrieval. While existing deep learning models predominantly utilize CYGNSS data with four-channel information, they often adopt single-channel inputs or simple channel concatenation without leveraging the benefits of cross-channel information interaction during training. To address this limitation, a novel spatial–channel attention-based network, namely SCAWaveNet, is proposed for SWH retrieval. Specifically, features from each channel of the DDMs are modeled as independent attention heads, enabling the fusion of spatial and channel-wise information. For auxiliary parameters, a lightweight attention mechanism is designed to assign weights along the spatial and channel dimensions. The final feature integrates both spatial and channel-level characteristics. Model performance is evaluated using four-channel CYGNSS data. Quantitative and qualitative experiments were conducted on CYGNSS-ERA5 test set, SCAWaveNet achieves an average RMSE of 0.438 m. Compared to state-of-the-art models, SCAWaveNet reduces RMSE by at least 3.52%. Furthermore, evaluations on WW3, Jason-3, and NDBC buoy data, as well as in wind speed, rainstorm, typhoon and noisy scenarios, further confirm the superiority of SCAWaveNet. The code is available at https://github.com/Clifx9908/SCAWaveNet. Chong Zhang 0013, Xichao Liu, Jinwei Bu, Yibing Zhan, Dapeng Tao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Semantic Contextualization of Face Forgery: A New Definition, Dataset, and Detection MethodabstractIn recent years, deep learning has greatly streamlined the process of manipulating photographic face images. Aware of the potential dangers, researchers have developed various tools to spot these counterfeits. Yet, none asks the fundamental question:What digital manipulations make a real photographic face image fake, while others do not? In this paper, we put face forgery in a semantic context and define thatcomputational methods that alter semantic face attributes to exceed human discrimination thresholds are sources of face forgery. Following our definition, we construct a large face forgery image dataset, where each image is associated with a set of labels organized in a hierarchical graph. Our dataset enables two new testing protocols to probe the generalizability of face forgery detectors. Moreover, we propose a semantics-oriented face forgery detection method that captures label relations and prioritizes the primary task (i.e., real or fake face detection). We show that the proposed dataset successfully exposes the weaknesses of current detectors as the test set and consistently improves their generalizability as the training set. Additionally, we demonstrate the superiority of our semantics-oriented method over traditional binary and multi-class classification-based detectors. Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Cps-STS: Bridging the Gap Between Content and Position for Coarse-Point-Supervised Scene Text SpotterabstractRecently, weakly supervised methods for scene text spotter are increasingly popular with researchers due to their potential to significantly reduce dataset annotation efforts. The latest progress in this field is text spotter based on single or multi-point annotations. However, this method struggles with the sensitivity of text recognition to the precise annotation location and fails to capture the relative positions and shapes of characters, leading to impaired recognition of texts with extensive rotations and flips. To address these challenges, this paper develops a novel method named Coarse-point-supervised Scene Text Spotter (Cps-STS). Cps-STS first utilizes a few approximate points as text location labels and introduces a learnable position modulation mechanism, easing the accuracy requirements for annotations and enhancing model robustness. Additionally, we incorporate a Spatial Compatibility Attention (SCA) module for text decoding to effectively utilize spatial data such as position and shape. This module fuses compound queries and global feature maps, serving as a bias in the SCA module to express text spatial morphology. In order to accurately locate and decode text content, we introduce features containing spatial morphology information and text content into the input features of the text decoder. By introducing features with spatial morphology information as bias terms into the text decoder, ablation experiments demonstrate that this operation enables the model to effectively identify and utilize the relationship between text content and position to enhance the recognition performance of our model. One significant advantage of Cps-STS is its ability to achieve full supervision-level performance with just a few imprecise coarse points at a low cost. Extensive experiments validate the effectiveness and superiority of Cps-STS over existing approaches. Weida Chen, Jie Jiang 0015, Linfei Wang, Huafeng Li 0001, Yibing Zhan, Dapeng Tao |
IEEE Trans. Multim. | 5 |
| 2025 | Facial Expression Recognition With Heatmap Neighbor Contrastive LearningabstractMany supervised learning-based facial expression recognition (FER) methods achieve good performance with the assistance of expression labels and a complex framework. However, there are inconsistent annotations in different expression datasets, making the above methods disadvantageous for new expression datasets or datasets with limited training data. The objective of this paper is to learn self-supervised facial expression features that enable the FER model not to rely on the annotation consistency of the different datasets. Most current self-supervised learning algorithms based on contrastive learning learn the representation by forcing different augmented views of the same image close in the embedding space, but they cannot cover all variances within a semantic class. We propose a heatmap neighbor contrastive learning (HNCL) method for FER. It treats the images corresponding to the heatmap nearest neighbors of expressions as other positives, providing more semantic variations than pre-defined augmented transformations. Therefore, our HNCL can learn better expression features covering more intra-class variances, improving the performance of the FER model based on self-supervised learning. After fine-tuning, HNCL with a simple framework achieves top-three performance on the in-the-lab datasets and even matches the performance of state-of-the-art supervised learning methods on the in-the-wild datasets. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Yibing Zhan, Dapeng Tao, Jun Wan 0005 |
IEEE Trans. Multim. | 5 |
| 2025 | SpliceMix: A Cross-Scale and Semantic Blending Augmentation Strategy for Multi-Label Image ClassificationabstractRecently, Mix-style data augmentation methods (e.g., Mixup and CutMix) have shown promising performance in various visual tasks. However, these methods are primarily designed for single-label images, ignoring the considerable discrepancies between single- and multi-label images,i.e., a multi-label image involves multiple co-occurred categories and fickle object scales. On the other hand, previous multi-label image classification (MLIC) methods tend to design elaborate models, bringing expensive computation. In this article, we introduce a simple but effective augmentation strategy for multi-label image classification, namely SpliceMix. The “splice” in our method is two-fold:1)Each mixed image is a splice of several downsampled images in the form of a grid, where the semantics of images attending to mixing are blended without object deficiencies for alleviating co-occurred bias;2)We splice mixed images and the original mini-batch to form a new SpliceMixed mini-batch, which allows an image with different scales to contribute to training together. Furthermore, such splice in our SpliceMixed mini-batch enables interactions between mixed images and original regular images. We also provide a simple and non-parametric extension based on consistency learning (SpliceMix-CL) to show the potential of extending our SpliceMix. Extensive experiments on various tasks demonstrate that only using SpliceMix with a baseline model (e.g., ResNet) achieves better performance than state-of-the-art methods. Moreover, the generalizability of our SpliceMix is further validated by the improvements in current MLIC methods when married with our SpliceMix. Lei Wang 0095, Yibing Zhan, Leilei Ma 0002, Dapeng Tao, Liang Ding 0006, Chen Gong 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | TD²-Net: Toward Denoising and Debiasing for Video Scene Graph GenerationabstractDynamic scene graph generation (SGG) focuses on detecting objects in a video and determining their pairwise relationships. Existing dynamic SGG methods usually suffer from several issues, including 1) Contextual noise, as some frames might contain occluded and blurred objects. 2) Label bias, primarily due to the high imbalance between a few positive relationship samples and numerous negative ones. Additionally, the distribution of relationships exhibits a long-tailed pattern. To address the above problems, in this paper, we introduce a network named TD2-Net that aims at denoising and debiasing for dynamic SGG. Specifically, we first propose a denoising spatio-temporal transformer module that enhances object representation with robust contextual information. This is achieved by designing a differentiable Top-K object selector that utilizes the gumbel-softmax sampling strategy to select the relevant neighborhood for each object. Second, we introduce an asymmetrical reweighting loss to relieve the issue of label bias. This loss function integrates asymmetry focusing factors and the volume of samples to adjust the weights assigned to individual samples. Systematic experimental results demonstrate the superiority of our proposed TD2-Net over existing state-of-the-art approaches on Action Genome databases. In more detail, TD2-Net outperforms the second-best competitors by 12.7% on mean-Recall@10 for predicate classification. Chong Shi, Yibing Zhan, Zuopeng Yang, Dacheng Tao |
AAAI | 3 |
| 2024 | Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion ModelsabstractDenoising Diffusion Probabilistic Models (DDPMs) have achieved significant success in generation tasks. Nevertheless, the exposure bias issue, i.e., the natural discrepancy between the training (the output of each step is calculated individually by a given input) and inference (the output of each step is calculated based on the input iteratively obtained based on the model), harms the performance of DDPMs. To our knowledge, few works have tried to tackle this issue by modifying the training process for DDPMs, but they still perform unsatisfactorily due to 1) partially modeling the discrepancy and 2) ignoring the prediction error accumulation. To address the above issues, in this paper, we propose a multi-step denoising scheduled sampling (MDSS) strategy to alleviate the exposure bias for DDPMs. Analyzing the formulations of the training and inference of DDPMs, MDSS 1) comprehensively considers the discrepancy influence of prediction errors on the output of the model (the Gaussian noise) and the output of the step (the calculated input signal of the next step), and 2) efficiently models the prediction error accumulation by using multiple iterations of a mathematical formulation initialized from one-step prediction error obtained from the model. The experimental results, compared with previous works, demonstrate that our approach is more effective in mitigating exposure bias in DDPM, DDIM, and DPM-solver. In particular, MDSS achieves an FID score of 3.86 in 100 sample steps of DDIM on the CIFAR-10 dataset, whereas the second best obtains 4.78. The code will be available on GitHub. Zhiyao Ren, Yibing Zhan, Liang Ding 0006, Gaoang Wang, Zhongyi Fan, Dacheng Tao |
AAAI | 2 |
| 2024 | Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReIDabstractText-to-image person re-identification (ReID) retrieves pedestrian images according to textual descriptions. Manually annotating textual descriptions is time-consuming, restricting the scale of existing datasets and therefore the generalization ability of ReID models. As a result, we study the transferable text-to-image ReID problem, where we train a model on our proposed large-scale database and directly deploy it to various datasets for evaluation. We obtain substantial training data via Multi-modal Large Language Models (MLLMs). Moreover, we identify and address two key challenges in utilizing the obtained textual descriptions. First, an MLLM tends to generate descriptions with similar structures, causing the model to overfit specific sentence patterns. Thus, we propose a novel method that uses MLLMs to caption images according to various templates. These templates are obtained using a multi-turn dialogue with a Large Language Model (LLM). Therefore, we can build a large-scale dataset with diverse textual descriptions. Second, an MLLM may produce incorrect descriptions. Hence, we introduce a novel method that automatically identifies words in a description that do not correspond with the image. This method is based on the similarity between one text and all patch token embeddings in the image. Then, we mask these words with a larger probability in the subsequent training epoch, alleviating the impact of noisy textual descriptions. The experimental results demonstrate that our methods significantly boost the direct transfer text-to-image ReID performance. Benefiting from the pre-trained model weights, we also achieve state-of-the-art performance in the traditional evaluation settings.https://github.com/WentaoTan/MLLM4Text-ReID Wentao Tan, Changxing Ding, Jiayu Jiang, Fei Wang 0032, Yibing Zhan, Dapeng Tao |
CVPR | 5 |
| 2024 | Parameter-Efficient Multi-Task Model Fusion with Partial LinearizationabstractLarge pre-trained models have enabled significant advances in machine learning and served as foundation components.
Model fusion methods, such as task arithmetic, have been proven to be powerful and scalable to incorporate fine-tuned weights from different tasks into a multi-task model.
However, efficiently fine-tuning large pre-trained models on multiple downstream tasks remains challenging, leading to inefficient multi-task model fusion.
In this work, we propose a novel method to improve multi-task fusion for parameter-efficient fine-tuning techniques like LoRA fine-tuning.
Specifically, our approach partially linearizes only the adapter modules and applies task arithmetic over the linearized adapters.
This allows us to leverage the the advantages of model fusion over linearized fine-tuning, while still performing fine-tuning and inference efficiently.
We demonstrate that our partial linearization technique enables a more effective fusion of multiple tasks into a single model, outperforming standard adapter tuning and task arithmetic alone.
Experimental results demonstrate the capabilities of our proposed partial linearization technique to effectively construct unified multi-task models via the fusion of fine-tuned task vectors.
We evaluate performance over an increasing number of tasks and find that our approach outperforms standard parameter-efficient fine-tuning techniques. The results highlight the benefits of partial linearization for scalable and efficient multi-task model fusion. Anke Tang, Li Shen 0008, Yong Luo 0002, Yibing Zhan, Han Hu 0003, Bo Du 0001, Yixin Chen 0001, Dacheng Tao |
ICLR | 4 |
| 2024 | A Dual-module Framework for Counterfactual Estimation over TimeabstractEfficiently and effectively estimating counterfactuals over time is crucial for optimizing treatment strategies. We present the Adversarial Counterfactual Temporal Inference Network (ACTIN), a novel framework with dual modules to enhance counterfactual estimation. The balancing module employs a distribution-based adversarial method to learn balanced representations, extending beyond the limitations of current classification-based methods to mitigate confounding bias across various treatment types. The integrating module adopts a novel Temporal Integration Predicting (TIP) strategy, which has a wider receptive field of treatments and balanced representations from the beginning to the current time for a more profound level of analysis. TIP goes beyond the established Direct Predicting (DP) strategy, which only relies on current treatments and representations, by empowering the integrating module to effectively capture long-range dependencies and temporal treatment interactions. ACTIN exceeds the confines of specific base models, and when implemented with simple base models, consistently delivers state-of-the-art performance and efficiency across both synthetic and real-world datasets. Xin Wang 0179, Shengfei Lyu, Lishan Yang 0004, Yibing Zhan, Huanhuan Chen 0001 |
ICML | 4 |
| 2024 | Confronting Reward Overoptimization for Diffusion Models: A Perspective of Inductive and Primacy BiasesabstractBridging the gap between diffusion models and human preferences is crucial for their integration into practical generative workflows. While optimizing downstream reward models has emerged as a promising alignment strategy, concerns arise regarding the risk of excessive optimization with learned reward models, which potentially compromises ground-truth performance. In this work, we confront the reward overoptimization problem in diffusion model alignment through the lenses of both inductive and primacy biases. We first identify a mismatch between current methods and the temporal inductive bias inherent in the multi-step denoising process of diffusion models, as a potential source of reward overoptimization. Then, we surprisingly discover that dormant neurons in our critic model act as a regularization against reward overoptimization while active neurons reflect primacy bias. Motivated by these observations, we propose Temporal Diffusion Policy Optimization with critic active neuron Reset (TDPO-R), a policy gradient algorithm that exploits the temporal inductive bias of diffusion models and mitigates the primacy bias stemming from active neurons. Empirical results demonstrate the superior efficacy of our methods in mitigating reward overoptimization. Code is avaliable at https://github.com/ZiyiZhang27/tdpo. Ziyi Zhang 0001, Sen Zhang 0006, Yibing Zhan, Yong Luo 0002, Yonggang Wen 0001, Dacheng Tao |
ICML | 3 |
| 2024 | Where to Mask: Structure-Guided Masking for Graph Masked Autoencoders
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001 |
IJCAI | 3 |
| 2024 | Gradformer: Graph Transformer with Exponential Decay
Chuang Liu 0008, Zelin Yao, Yibing Zhan, Xueqi Ma, Shirui Pan, Wenbin Hu 0001 |
IJCAI | 3 |
| 2024 | MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models
Kanxue Li, Baosheng Yu, Yibing Zhan, Qiong Cao, Li Shen 0008, Lusong Li, Dapeng Tao, Xiaodong He 0001 |
IJCAI | 4 |
| 2024 | Joint Input and Output Coordination for Class-Incremental Learning
Shuai Wang 0011, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Wei Yu 0004, Yonggang Wen 0001, Dacheng Tao |
IJCAI | 2 |
| 2024 | DreamBooth++: Boosting Subject-Driven Generation via Region-Level References PackingabstractDreamBooth has demonstrated significant potential in subject-driven text-to-image generation, especially in scenarios requiring precise preservation of a subject's appearance. However, it still suffers from inefficiency and requires extensive iterative training to customize concepts using a small set of reference images. To address these issues, we introduce DreamBooth++, a region-level training strategy designed to significantly improve the efficiency and effectiveness of learning specific subjects. In particular, our approach employs a region-level data re-formulation technique that packs a set of reference images into a single sample, significantly reducing computational costs. Moreover, we adapt convolution and self-attention layers to ensure their processings are restricted within individual regions. Thus their operational scope (i.e., receptive field) can be preserved within a single subject, avoiding generating multiple sub-images within a single image. Last but not least, we design a text-guided prior regularization between our model and the pretrained one to preserve the original semantic generation ability. Comprehensive experiments demonstrate that our training strategy not only accelerates the subject-learning process but also significantly boosts fidelity to both subject and prompts in subject-driven generation. Zhongyi Fan, Zixin Yin, Yibing Zhan, Heliang Zheng |
ACM Multimedia | 4 |
| 2024 | Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingabstractIn Video-based Facial Expression Recognition (V-FER), models are typically trained on closed-set datasets with a fixed number of known classes. However, these models struggle with unknown classes common in real-world scenarios. In this paper, we introduce a challenging Open-set Video-based Facial Expression Recognition (OV-FER) task, aiming to identify both known and new, unseen facial expressions. While existing approaches use large-scale vision-language models like CLIP to identify unseen classes, we argue that these methods may not adequately capture the subtle human expressions needed for OV-FER. To address this limitation, we propose a novel Human Expression-Sensitive Prompting (HESP) mechanism to significantly enhance CLIP's ability to model video-based facial expression details effectively. Our proposed HESP comprises three components: 1) a textual prompting module with learnable prompts to enhance CLIP's textual representation of both known and unknown emotions, 2) a visual prompting module that encodes temporal emotional information from video frames using expression-sensitive attention, equipping CLIP with a new visual modeling ability to extract emotion-rich information, and 3) an open-set multi-task learning scheme that promotes interaction between the textual and visual modules, improving the understanding of novel human emotions in video sequences. Extensive experiments conducted on four OV-FER task settings demonstrate that HESP can significantly boost CLIP's performance (a relative improvement of 17.93% on AUROC and 106.18% on OSCR) and outperform other state-of-the-art open-set video understanding methods by a large margin. Code is available at https://github.com/cosinehuang/HESP. Yuanyuan Liu 0004, Yibing Zhan, Zijing Chen, Zhe Chen 0013 |
ACM Multimedia | 4 |
| 2024 | An Efficient Multi-prior Hybrid Approach for Consistent 3D Generation from Single Images
Yichen Ouyang, Jiayi Ye, Wenhao Chai, Dapeng Tao, Yibing Zhan, Gaoang Wang |
MMAsia | 5 |
| 2024 | DIVOTrack: A Novel Dataset and Baseline Method for Cross-View Multi-Object Tracking in DIVerse Open Scenes
Shengyu Hao, Peiyuan Liu, Yibing Zhan, Kaixun Jin, Zuozhu Liu, Mingli Song, Jenq-Neng Hwang, Gaoang Wang |
Int. J. Comput. Vis. | 3 |
| 2024 | HAG-Former: A Temporal-Polarimetric Relationship Inference Network From Local to GlobalabstractMultitemporal polarimetric SAR (PolSAR) data can provide a unique insight into the temporal scattering characteristics of targets and highlight their dynamic changes over time, therefore supporting improved classification performance. Constrained by the complexities of satellite orbit control technology and the challenges associated with time-series PolSAR data acquisition, most prevailing methodologies rely solely on a single PolSAR image to tackle land coverage classification, inherently limiting their ability to generalize across diverse scenarios. To address this limitation, this work introduces a novel Hybrid Attention-GRU Transformer (HAG-Former) model, which harnesses the power of pixel-level temporal-polarimetric change analysis and captures the dynamic variations in polarization scattering properties, thereby enhancing classification robustness and versatility. In this approach, we seamlessly integrate a self-attention mechanism, a Gated Recurrent Unit (GRU), and a transformer encoder to delve into pixel-level changes in polarimetric features. Initially, the self-attention mechanism pinpoints crucial classification-aiding features, bolstering their significance. The weighted features are then fed into the GRU model, enhancing local temporal-polarimetric relationship insights. These relationships, coupled with significant features from the self-attention mechanism, are subsequently processed by the transformer encoder, unraveling global information. Furthermore, we employ a label smoothing loss function during training, mitigating the impact of sample imbalance on classification accuracy. To validate the effectiveness of our proposed methodology, we evaluated it on two benchmark datasets. The results demonstrate a notable enhancement in classification performance, achieving an overall accuracy improvement of 2.21% and 1.79% over the state-of-the-art. The code is available athttps://github.com/Thomasakun/HAGFormer. Carlos López-Martínez, Yibing Zhan, Dapeng Tao |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | IoUformer: Pseudo-IoU prediction with transformer for visual tracking
Huayue Cai, Long Lan, Jing Zhang 0037, Xiang Zhang 0008, Yibing Zhan, Zhigang Luo |
Neural Networks | 5 |
| 2024 | Finding core labels for maximizing generalization of graph neural networks
Sichao Fu, Xueqi Ma, Yibing Zhan, Fanyu You, Qinmu Peng, Tongliang Liu, James Bailey 0001, Danilo P. Mandic |
Neural Networks | 3 |
| 2024 | Exploring sparsity in graph transformers
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Liang Ding 0006, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001, Bo Du 0001 |
Neural Networks | 2 |
| 2024 | Free-Form Composition Networks for Egocentric Action RecognitionabstractEgocentric action recognition is gaining significant attention in the field of human action recognition. In this paper, we address data scarcity issue in egocentric action recognition from a compositional generalization perspective. To tackle this problem, we propose a free-form composition network (FFCN) that can simultaneously learn disentangled verb, preposition, and noun representations, and then use them to compose new samples in the feature space for rare classes of action videos. First, we use a graph to capture the spatial-temporal relations among different hand/object instances in each action video. We thus decompose each action into a set of verb and preposition spatial-temporal representations using the edge features in the graph. The temporal decomposition extracts verb and preposition representations from different video frames, while the spatial decomposition adaptively learns verb and preposition representations from action-related instances in each frame. With these spatial-temporal representations of verbs and prepositions, we can compose new samples for those rare classes in a free-form manner, which is not restricted to a rigid form of a verb and a noun. The proposed FFCN can directly generate new training data samples for rare classes, hence significantly improve action recognition performance. We evaluated our method on three popular egocentric action recognition datasets, Something-Something V2, H2O, and EPIC-KITCHENS-100, and the experimental results demonstrate the effectiveness of the proposed method for handling data scarcity problems, including long-tailed and few-shot egocentric action recognition. Haoran Wang 0001, Qinghua Cheng, Baosheng Yu, Yibing Zhan, Dapeng Tao, Liang Ding 0006, Haibin Ling |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | DeIoU: Toward Distinguishable Box Prediction in Densely Packed Object DetectionabstractThe Intersection over Union (IoU) has been widely employed in various stages of object detection owing to its ability to quantify the similarity between boxes objectively. However, in densely packed scenes full of crowded and small-sized objects, adjacent positive boxes often exhibit high levels of overlap. This overlap interference compromises the consistency between quality evaluation and confidence, leading to ambiguous box prediction within the previous IoU-based models. To address this issue, we design a novel learning paradigm tailored for Dense scenes based on IoU, called DeIoU. This approach effectively suppresses unnecessary overlap between predicted boxes and thereby enhances representation learning for non-salient objects. Specifically, it consists of a dense box regression loss${\mathcal {L}}_{DeIoU}$and a one-to-many (O2M) label matching strategy guided by DeIoU. These components focus on calibrating the position and shape prediction quality during the model training, learning distinguishable object features by penalizing overlap interference between neighboring boxes. Extensive experiments on four object detection datasets including SKU-110K, CrowdHuman, MS COCO 2017, and DIOR, demonstrate that our DeIoU-based learning strategy outperforms other state-of-the-art methods. Notably, the proposed method delivers a substantial improvement (average$1.3~{AP}$and$1.8~MR^{-2}$) across popular detectors on SKU-110K and CrowdHuman while exhibiting distinct competitiveness on small objects within natural scenes. Linfei Wang, Yibing Zhan, Long Lan, Dapeng Tao, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Joint Spatial-Spectral Optimization for the High-Magnification Fusion of Hyperspectral and Multispectral ImagesabstractThe fusion of hyperspectral and multispectral images is an important strategy for enhancing the spatial resolution of hyperspectral images. With the rapid advancement of multispectral imaging technology, the disparity in spatial resolution between multispectral and hyperspectral images is increasing. In certain scenarios, termed high-magnification, this difference can exceed$32\times $. Previous methods do not perform well under high-magnification fusion, and naturally, a challenge arises in achieving effective high-magnification super-resolution fusion. In light of the above analysis, this article introduces a novel algorithm for high-magnification super-resolution fusion of hyperspectral and multispectral images based on the joint optimization of spatial and spectral information. Specifically, our algorithm consists of three stages: 1) a fast preliminary fusion stage based on the Moore-Penrose inverse and singular value correlation priors for the rapid acquisition of preliminary solutions; 2) a joint spatial-spectral optimization stage where a coupled optimization framework is constructed to achieve integrated optimization of spatial and spectral information; and 3) an error backpropagation optimization stage where an effective error optimization term is introduced to further refine the fusion performance. We conducted extensive experiments on widely employed publicly available simulated datasets and real datasets. The experimental results unequivocally indicate that our proposed methodology consistently exhibits superior fusion performance compared with state-of-the-art methods, even under the condition of${\geq }60\times $super-resolution. Yibing Zhan, Zhengbin Pang, Tong Zhou 0008, Xueqiong Li, Long Lan, Yuanxi Peng |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Few-Shot Learning With Dynamic Graph Structure PreservingabstractIn recent years, few-shot learning has received increasing attention in the Internet of Things areas. Few-shot learning aims to distinguish unseen classes with a few labeled samples from each class. Most recently transductive few-shot studies highly rely on the static geometry distributions generated on the feature space during the label propagation process between unseen class instances. However, these recent methods fail to guarantee that the generated graph structure preserves the true distributions between data properly. In this article, we propose a novel dynamic graph structure preserving (DGSP) model for few-shot learning. Specifically, we formulate the objective function of DGSP by simultaneously considering the data correlations from the feature space and the label space to update the generated graph structure, which can reasonably revise the inappropriate or mistaken local geometry relationships. Then, we design an efficient alternating optimization algorithm to jointly learn the label prediction matrix and the optimal graph structure, the latter of which can be formulated as a linear programming problem. Moreover, our proposed DGSP can be easily combined with any backbone networks during the learning process. We conduct extensive experimental results across different benchmarks, backbones, and task settings, and our method achieves state-of-the-art performance compared with methods based on transductive few-shot learning. Sichao Fu, Qiong Cao, Yunwen Lei, Yibing Zhan, Xinge You |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | SkipNode: On Alleviating Performance Degradation for Deep Graph Convolutional NetworksabstractGraph Convolutional Networks (GCNs) suffer from performance degradation when models go deeper. However, earlier works only attributed the performance degeneration to over-smoothing. In this paper, we conduct theoretical and experimental analysis to explore the fundamental causes of performance degradation in deep GCNs: over-smoothing and gradient vanishing have a mutually reinforcing effect that causes the performance to deteriorate more quickly in deep GCNs. On the other hand, existing anti-over-smoothing methods all perform full convolutions up to the model depth. They could not well resist the exponential convergence of over-smoothing due to model depth increasing. In this work, we propose a simple yet effective plug-and-play module,SkipNode, to overcome the performance degradation of deep GCNs. It samples graph nodes in each convolutional layer to skip the convolution operation. In this way, both over-smoothing and gradient vanishing can be effectively suppressed since (1) not all nodes'features propagate through full layers and, (2) the gradient can be directly passed back through “skipped” nodes. We provide both theoretical analysis and empirical evaluation to demonstrate the efficacy ofSkipNodeand its superiority over SOTA baselines. Weigang Lu 0001, Yibing Zhan, Binbin Lin 0001, Ziyu Guan, Liu Liu 0014, Baosheng Yu, Wei Zhao 0019, Yaming Yang 0002, Dacheng Tao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Difference-Aware Distillation for Semantic SegmentationabstractIn recent years, various distillation methods for semantic segmentation have been proposed. However, these methods typically train the student model to imitate the intermediate features or logits of the teacher model directly, thereby overlooking the high-discrepancy regions learned by both models, particularly the differences in instance edges. In this paper, we introduce a novel approach, called Difference-aware Distillation, to address this limitation. Our proposed method detects the discrepancies among the teacher model and the student model in the logit space through two masking mechanisms (i.e., masking by logit differences with respect to the ground truth labels and masking by differences in the predictive class probabilities), and guides the student model to restore the teacher's features with the focus on these highly-discrepant regions, resulting in improved segmentation performance. With the features jointly masked by these two mechanisms, the student model learns to preserve the teacher's features via a feature generation module, thus achieving better representation. Our experimental evaluation on three datasets, Cityscapes, Pascal2012, and ADE20 K, demonstrates our proposed approach outperforms several baselines considered. Further visualization analysis confirms that our method effectively directs the student model's attention to the discrepancies, such as the edges of small objects and the interiors of large objects. Jianping Gou, Xiabin Zhou, Lan Du 0002, Yibing Zhan, Wu Chen 0005, Zhang Yi 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Textual Enhanced Adaptive Meta-Fusion for Few-Shot Visual RecognitionabstractFew-shot learning (FSL) is a challenging task that aims to train a classifier to recognize novel categories, where only a few annotated examples are available in each category. Recently, many FSL approaches have been proposed based on the meta-learning paradigm, which attempts to learn transferable knowledge from similar tasks by designing a meta-learner. However, most of these approaches only exploit the information from visual modality and do not utilize ones from additional modalities (e.g., textual description). Since the labeled examples in FSL are limited, increasing the information on the examples is a probable solution to improve the classification performance. This motivates us to propose a novel meta-learning method, termed textual enhanced adaptive meta-fusion FSL (TAMF-FSL), which leverages both the visual information from the visual image and semantic information from language supervision. Specifically, TAMF-FSL exploits the semantic information of textual description to improve the visual-based models. We first employ a text encoder to learn the semantic features of each visual category, and then design a modality alignment module and meta-fusion module to align and fuse the visual and semantic features for final prediction. Extensive experiments show that the proposed method outperforms many recent or competitive FSL counterparts on two popular datasets. Mengya Han, Yibing Zhan, Yong Luo 0002, Han Hu 0003, Kehua Su, Bo Du 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Bounding Box Vectorization for Oriented Object Detection With Tanimoto Coefficient RegressionabstractCurrent oriented object detection methods mainly utilize a vanilla coordinate-angle representation for bounding box regression, which usually suffers from inconsistency between the bounding box regression losses and prediction errors induced with respect to different rotation angles, aspect ratios, and scales. Therefore, although the existing oriented object detectors have achieved very good performances under coarse evaluation metrics such as AP50, their performance significantly degrades when using stricter evaluation metric such as AP75. To address the abovementioned issues, we propose a new regression method with bounding box vectorization that implicitly represents the shape and orientation of an object with a set of orthogonal vectors. By doing this, the proposed method delicately avoids the inconsistency issues encountered in oriented bounding box regression. During training, we introduce the Tanimoto coefficient to evaluate the similarity of the bounding box vector in a shape- and orientation-aware manner, and we refer to the proposed box-to-vector loss as the B2V loss. In addition to 2D object detection, the proposed method can be easily generalized to 3D scenarios involving orientation estimation, such as autonomous driving. We evaluate the proposed method through extensive experiments conducted on four popular oriented object detection datasets, including both 2D and 3D datasets, where the proposed method significantly outperforms the recently developed state-of-the-art methods when using a more accurate evaluation metric. Linfei Wang, Yibing Zhan, Wei Liu 0005, Baosheng Yu, Dapeng Tao |
IEEE Trans. Multim. | 2 |
| 2024 | Not All Instances Contribute Equally: Instance-Adaptive Class Representation Learning for Few-Shot Visual RecognitionabstractFew-shot visual recognition refers to recognize novel visual concepts from a few labeled instances. Many few-shot visual recognition methods adopt the metric-based meta-learning paradigm by comparing the query representation with class representations to predict the category of query instance. However, the current metric-based methods generally treat all instances equally and consequently often obtain biased class representation, considering not all instances are equally significant when summarizing the instance-level representations for the class-level representation. For example, some instances may contain unrepresentative information, such as too much background and information of unrelated concepts, which skew the results. To address the above issues, we propose a novel metric-based meta-learning framework termed instance-adaptive class representation learning network (ICRL-Net) for few-shot visual recognition. Specifically, we develop an adaptive instance revaluing network (AIRN) with the capability to address the biased representation issue when generating the class representation, by learning and assigning adaptive weights for different instances according to their relative significance in the support set of corresponding class. In addition, we design an improved bilinear instance representation and incorporate two novel structural losses, i.e., intraclass instance clustering loss and interclass representation distinguishing loss, to further regulate the instance revaluation process and refine the class representation. We conduct extensive experiments on four commonly adopted few-shot benchmarks: miniImageNet, tieredImageNet, CIFAR-FS, and FC100 datasets. The experimental results compared with the state-of-the-art approaches demonstrate the superiority of our ICRL-Net. Mengya Han, Yibing Zhan, Yong Luo 0002, Bo Du 0001, Han Hu 0003, Yonggang Wen 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Comprehensive Graph Gradual Pruning for Sparse Training in Graph Neural NetworksabstractGraph neural networks (GNNs) tend to suffer from high computation costs due to the exponentially increasing scale of graph data and a large number of model parameters, which restricts their utility in practical applications. To this end, some recent works focus on sparsifying GNNs (including graph structures and model parameters) with the lottery ticket hypothesis (LTH) to reduce inference costs while maintaining performance levels. However, the LTH-based methods suffer from two major drawbacks: 1) they require exhaustive and iterative training of dense models, resulting in an extremely large training computation cost, and 2) they only trim graph structures and model parameters but ignore the node feature dimension, where vast redundancy exists. To overcome the above limitations, we propose a comprehensive graph gradual pruning framework termed CGP. This is achieved by designing a during-training graph pruning paradigm to dynamically prune GNNs within one training process. Unlike LTH-based methods, the proposed CGP approach requires no retraining, which significantly reduces the computation costs. Furthermore, we design a cosparsifying strategy to comprehensively trim all the three core elements of GNNs: graph structures, node features, and model parameters. Next, to refine the pruning operation, we introduce a regrowth process into our CGP framework, to reestablish the pruned but important connections. The proposed CGP is evaluated over a node classification task across six GNN architectures, including shallow models [graph convolutional network (GCN) and graph attention network (GAT)], shallow-but-deep-propagation models [simple graph convolution (SGC) and approximate personalized propagation of neural predictions (APPNP)], and deep models [GCN via initial residual and identity mapping (GCNII) and residual GCN (ResGCN)], on a total of 14 real-world graph datasets, including large-scale graph datasets from the challenging Open Graph Benchmark (OGB). Experiments reveal that the proposed strategy greatly improves both training and inference efficiency while matching or even exceeding the accuracy of the existing methods. Chuang Liu 0008, Xueqi Ma, Yibing Zhan, Liang Ding 0006, Dapeng Tao, Bo Du 0001, Wenbin Hu 0001, Danilo P. Mandic |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Boosting Graph Contrastive Learning via Adaptive SamplingabstractContrastive learning (CL) is a prominent technique for self-supervised representation learning, which aims to contrast semantically similar (i.e., positive) and dissimilar (i.e., negative) pairs of examples under different augmented views. Recently, CL has provided unprecedented potential for learning expressive graph representations without external supervision. In graph CL, the negative nodes are typically uniformly sampled from augmented views to formulate the contrastive objective. However, this uniform negative sampling strategy limits the expressive power of contrastive models. To be specific, not all the negative nodes can provide sufficiently meaningful knowledge for effective contrastive representation learning. In addition, the negative nodes that are semantically similar to the anchor are undesirably repelled from it, leading to degraded model performance. To address these limitations, in this article, we devise an adaptive sampling strategy termed "AdaS. " The proposed AdaS framework can be trained to adaptively encode the importance of different negative nodes, so as to encourage learning from the most informative graph nodes. Meanwhile, an auxiliary polarization regularizer is proposed to suppress the adverse impacts of the false negatives and enhance the discrimination ability of AdaS. The experimental results on a variety of real-world datasets firmly verify the effectiveness of our AdaS in improving the performance of graph CL. Sheng Wan, Yibing Zhan, Shuo Chen 0003, Shirui Pan, Jian Yang 0003, Dacheng Tao, Chen Gong 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Attentional Composition Networks for Long-Tailed Human Action RecognitionabstractThe problem of long-tailed visual recognition has been receiving increasing research attention. However, the long-tailed distribution problem remains underexplored for video-based visual recognition. To address this issue, in this article we propose a compositional learning based solution for video-based human action recognition. Our method, named Attentional Composition Networks (ACN), first learns verb-like and preposition-like components, then shuffles these components to generate samples for the tail classes in the feature space to augment the data for the tail classes. Specifically, during training, we represent each action video by a graph that captures the spatial-temporal relations (edges) among detected human/object instances (nodes). Then, ACN utilizes the position information to decompose each action into a set of verb and preposition representations using the edge features in the graph. After that, the verb and preposition features from different videos are combined via an attention structure to synthesize feature representations for tail classes. This way, we can enrich the data for the tail classes and consequently improve the action recognition for these classes. To evaluate the compositional human action recognition, we further contribute a new human action recognition dataset, namely NEU-Interaction (NEU-I). Experimental results on both Something-Something V2 and the proposed NEU-I demonstrate the effectiveness of the proposed method for long-tailed, few-shot, and zero-shot problems in human action recognition. Source code and the NEU-I dataset are available at https://github.com/YajieW99/ACN . Haoran Wang 0001, Baosheng Yu, Yibing Zhan, Chunfeng Yuan, Wankou Yang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State TrackingabstractQingyue Wang, Liang Ding, Yanan Cao, Yibing Zhan, Zheng Lin, Shi Wang, Dacheng Tao, Li Guo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Qingyue Wang, Liang Ding 0006, Yanan Cao 0001, Yibing Zhan, Zheng Lin 0001, Shi Wang 0002, Dacheng Tao, Li Guo 0001 |
ACL (1) | 4 |
| 2023 | Pose-disentangled Contrastive Learning for Self-supervised Facial RepresentationabstractSelf-supervised facial representation has recently attracted increasing attention due to its ability to perform face understanding without relying on large-scale annotated datasets heavily. However, analytically, current contrastive-based self-supervised learning (SSL) still performs unsatisfactorily for learning facial representation. More specifically, existing contrastive learning (CL) tends to learn pose-invariant features that cannot depict the pose details of faces, compromising the learning performance. To conquer the above limitation of CL, we propose a novel Pose-disentangled Contrastive Learning (PCL) method for general self-supervised facial representation. Our PCL first devises a pose-disentangled decoder (PDD) with a delicately designed orthogonalizing regulation, which disentangles the pose-related features from the face-aware features; therefore, pose-related and other pose-unrelated facial information could be performed in individual subnetworks and do not affect each other's training. Furthermore, we introduce a pose-related contrastive learning scheme that learns pose-related information based on data augmentation of the same image, which would deliver more effective face-aware representation for various downstream tasks. We conducted linear evaluation on four challenging downstream facial understanding tasks, i.e., facial expression recognition, face recognition, AU detection and head pose estimation. Experimental results demonstrate that PCL significantly outperforms cuttingedge SSL methods. Our Code is available at https://github.com/DreamMr/PCL. Yuanyuan Liu 0004, Wenbin Wang 0001, Yibing Zhan, Shaoze Feng, Kejun Liu, Zhe Chen 0013 |
CVPR | 3 |
| 2023 | Token Contrast for Weakly-Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) using image-level labels typically utilizes Class Activation Map (CAM) to generate the pseudo labels. Limited by the local structure perception of CNN, CAM usually cannot identify the integral object regions. Though the recent Vision Transformer (ViT) can remedy this flaw, we observe it also brings the over-smoothing issue, i.e., the final patch tokens incline to be uniform. In this work, we propose Token Contrast (ToCo) to address this issue and further explore the virtue of ViT for WSSS. Firstly, motivated by the observation that intermediate layers in ViT can still retain semantic diversity, we designed a Patch Token Contrast module (PTC). PTC supervises the final patch tokens with the pseudo token relations derived from intermediate layers, allowing them to align the semantic regions and thus yield more accurate CAM. Secondly, to further differentiate the low-confidence regions in CAM, we devised a Class Token Contrast module (CTC) inspired by the fact that class tokens in ViT can capture high-level semantics. CTC facilitates the representation consistency between uncertain local regions and global objects by contrasting their class tokens. Experiments on the PASCAL VOC and MS COCO datasets show the proposed ToCo can remarkably surpass other single-stage competitors and achieve comparable performance with state-of-the-art multi-stage methods. Code is available at https://github.com/rulixiang/ToCo. Lixiang Ru, Heliang Zheng, Yibing Zhan, Bo Du 0001 |
CVPR | 3 |
| 2023 | Combating Noisy Labels with Sample Selection by Mining High-Discrepancy ExamplesabstractThe sample selection approach is popular in learning with noisy labels. The state-of-the-art methods train two deep networks simultaneously for sample selection, which aims to employ their different learning abilities. To prevent two networks from converging to a consensus, their divergence should be maintained. Prior work presents that the divergence can be kept by locating the disagreement data on which the prediction labels of the two networks are different. However, this procedure is sample-inefficient for generalization, which means that only a few clean examples can be utilized in training. In this paper, to address the issue, we propose a simple yet effective method called CoDis. In particular, we select possibly clean data that simultaneously have high-discrepancy prediction probabilities between two networks. As selected data have high discrepancies in probabilities, the divergence of two networks can be maintained by training on such data. In addition, the condition of high discrepancies is milder than disagreement, which allows more data to be considered for training, and makes our method more sample-efficient. Moreover, we show that the proposed method enables to mine hard clean examples to help generalization. Empirical results show that CoDis is superior to multiple baselines in the robustness of trained models. Xiaobo Xia, Bo Han 0003, Yibing Zhan, Jun Yu 0001, Mingming Gong, Chen Gong 0002, Tongliang Liu |
ICCV | 3 |
| 2023 | Gapformer: Graph Transformer with Graph Pooling for Node ClassificationabstractGraph Transformers (GTs) have proved their advantage in graph-level tasks. However, existing GTs still perform unsatisfactorily on the node classification task due to 1) the overwhelming unrelated information obtained from a vast number of irrelevant distant nodes and 2) the quadratic complexity regarding the number of nodes via the fully connected attention mechanism. In this paper, we present Gapformer, a method for node classification that deeply incorporates Graph Transformer with Graph Pooling. More specifically, Gapformer coarsens the large-scale nodes of a graph into a smaller number of pooling nodes via local or global graph pooling methods, and then computes the attention solely with the pooling nodes rather than all other nodes. In such a manner, the negative influence of the overwhelming unrelated nodes is mitigated while maintaining the long-range information, and the quadratic complexity is reduced to linear complexity with respect to the fixed number of pooling nodes. Extensive experiments on 13 node classification datasets, including homophilic and heterophilic graph datasets, demonstrate the competitive performance of Gapformer over existing Graph Neural Networks and GTs. Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Liang Ding 0006, Dapeng Tao, Jia Wu 0001, Wenbin Hu 0001 |
IJCAI | 2 |
| 2023 | Graph Pooling for Graph Neural Networks: Progress, Challenges, and OpportunitiesabstractGraph neural networks have emerged as a leading architecture for many graph-level tasks, such as graph classification and graph generation. As an essential component of the architecture, graph pooling is indispensable for obtaining a holistic graph-level representation of the whole graph. Although a great variety of methods have been proposed in this promising and fast-developing research field, to the best of our knowledge, little effort has been made to systematically summarize these works. To set the stage for the development of future works, in this paper, we attempt to fill this gap by providing a broad review of recent methods for graph pooling. Specifically, 1) we first propose a taxonomy of existing graph pooling methods with a mathematical summary for each category; 2) then, we provide an overview of the libraries related to graph pooling, including the commonly used datasets, model architectures for downstream tasks, and open-source implementations; 3) next, we further outline the applications that incorporate the idea of graph pooling in a variety of domains; 4) finally, we discuss certain critical challenges facing current studies and share our insights on future potential directions for research on the improvement of graph pooling. Chuang Liu 0008, Yibing Zhan, Jia Wu 0001, Bo Du 0001, Wenbin Hu 0001, Tongliang Liu, Dacheng Tao |
IJCAI | 2 |
| 2023 | Semantics-Enriched Cross-Modal Alignment for Complex-Query Video Moment RetrievalabstractVideo moment retrieval (VMR) aims to search for a video segment that matches the search intent in a query sentence, which has received increasing attention in recent years, due to its practical values in various fields. Existing efforts devoted to this interesting yet challenging task typically encode the query sentence and video segments into unstructured global representations for cross-modal interaction and fusion, which may fail to accurately capture the search intent in complex queries with multi-granularity semantics. Xiang Zhang 0008, Xun Yang 0001, Yibing Zhan, Long Lan, Jianfeng Dong, Hongzhou Wu |
ACM Multimedia | 4 |
| 2023 | CLNode: Curriculum Learning for Node ClassificationabstractNode classification is a fundamental graph-based task that aims to predict the classes of unlabeled nodes, for which Graph Neural Networks (GNNs) are the state-of-the-art methods. Current GNNs assume that nodes in the training set contribute equally during training. However, the quality of training nodes varies greatly, and the performance of GNNs could be harmed by two types of low-quality training nodes: (1) inter-class nodes situated near class boundaries that lack the typical characteristics of their corresponding classes. Because GNNs are data-driven approaches, training on these nodes could degrade the accuracy. (2) mislabeled nodes. In real-world graphs, nodes are often mislabeled, which can significantly degrade the robustness of GNNs. To mitigate the detrimental effect of the low-quality training nodes, we present CLNode, which employs a selective training strategy to train GNN based on the quality of nodes. Specifically, we first design a multi-perspective difficulty measurer to accurately measure the quality of training nodes. Then, based on the measured qualities, we employ a training scheduler that selects appropriate training nodes to train GNN in each epoch. To evaluate the effectiveness of CLNode, we conduct extensive experiments by incorporating it in six representative backbone GNNs. Experimental results on real-world networks demonstrate that CLNode is a general framework that can be combined with various GNNs to improve their accuracy and robustness. Xiaowen Wei, Xiuwen Gong, Yibing Zhan, Bo Du 0001, Yong Luo 0002, Wenbin Hu 0001 |
WSDM | 3 |
| 2023 | Multi-target Knowledge Distillation via Student Self-reflectionabstractAbstract Knowledge distillation is a simple yet effective technique for deep model compression, which aims to transfer the knowledge learned by a large teacher model to a small student model. To mimic how the teacher teaches the student, existing knowledge distillation methods mainly adapt an unidirectional knowledge transfer, where the knowledge extracted from different intermedicate layers of the teacher model is used to guide the student model. However, it turns out that the students can learn more effectively through multi-stage learning with a self-reflection in the real-world education scenario, which is nevertheless ignored by current knowledge distillation methods. Inspired by this, we devise a new knowledge distillation framework entitled multi-target knowledge distillation via student self-reflection or MTKD-SSR, which can not only enhance the teacher’s ability in unfolding the knowledge to be distilled, but also improve the student’s capacity of digesting the knowledge. Specifically, the proposed framework consists of three target knowledge distillation mechanisms: a stage-wise channel distillation (SCD), a stage-wise response distillation (SRD), and a cross-stage review distillation (CRD), where SCD and SRD transfer feature-based knowledge (i.e., channel features) and response-based knowledge (i.e., logits) at different stages, respectively; and CRD encourages the student model to conduct self-reflective learning after each stage by a self-distillation of the response-based knowledge. Experimental results on five popular visual recognition datasets, CIFAR-100, Market-1501, CUB200-2011, ImageNet, and Pascal VOC, demonstrate that the proposed framework significantly outperforms recent state-of-the-art knowledge distillation methods. Jianping Gou, Xiangshuo Xiong, Baosheng Yu, Lan Du 0002, Yibing Zhan, Dacheng Tao |
Int. J. Comput. Vis. | 5 |
| 2023 | Attribute-Image Person Re-identification via Modal-Consistent Metric Learning
Jianqing Zhu, Liu Liu 0014, Yibing Zhan, Xiaobin Zhu 0001, Huanqiang Zeng, Dacheng Tao |
Int. J. Comput. Vis. | 3 |
| 2023 | Metapath-fused heterogeneous graph network for molecular property prediction
Guojia Wan, Yibing Zhan, Bo Du 0001 |
Inf. Sci. | 3 |
| 2023 | KE-X: Towards subgraph explanations of knowledge graph embedding based on knowledge information gain
Guojia Wan, Yibing Zhan, Zengmao Wang, Liang Ding 0006, Zhigao Zheng 0001, Bo Du 0001 |
Knowl. Based Syst. | 3 |
| 2023 | On exploring node-feature and graph-structure diversities for node drop graph pooling
Chuang Liu 0008, Yibing Zhan, Baosheng Yu, Liu Liu 0014, Bo Du 0001, Wenbin Hu 0001, Tongliang Liu |
Neural Networks | 2 |
| 2023 | Graph structure reforming framework enhanced by commute time distance for graph classification
Wenhang Yu, Xueqi Ma, James Bailey 0001, Yibing Zhan, Jia Wu 0001, Bo Du 0001, Wenbin Hu 0001 |
Neural Networks | 4 |
| 2023 | Sparse-to-Dense Matching Network for Large-Scale LiDAR Point Cloud RegistrationabstractPoint cloud registration is a fundamental problem in 3D computer vision. Previous learning-based methods for LiDAR point cloud registration can be categorized into two schemes: dense-to-dense matching methods and sparse-to-sparse matching methods. However, for large-scale outdoor LiDAR point clouds, solving dense point correspondences is time-consuming, whereas sparse keypoint matching easily suffers from keypoint detection error. In this paper, we propose SDMNet, a novel Sparse-to-Dense Matching Network for large-scale outdoor LiDAR point cloud registration. Specifically, SDMNet performs registration in two sequential stages: sparse matching stage and local-dense matching stage. In the sparse matching stage, we sample a set of sparse points from the source point cloud and then match them to the dense target point cloud using a spatial consistency enhanced soft matching network and a robust outlier rejection module. Furthermore, a novel neighborhood matching module is developed to incorporate local neighborhood consensus, significantly improving performance. The local-dense matching stage is followed for fine-grained performance, where dense correspondences are efficiently obtained by performing point matching in local spatial neighborhoods of high-confidence sparse correspondences. Extensive experiments on three large-scale outdoor LiDAR point cloud datasets demonstrate that the proposed SDMNet achieves state-of-the-art performance with high efficiency. Fan Lu 0001, Guang Chen 0001, Yinlong Liu, Yibing Zhan, Zhijun Li 0001, Dacheng Tao, Changjun Jiang 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Expression snippet transformer for robust video-based facial expression recognition
Yuanyuan Liu 0004, Wenbin Wang 0001, Chuanxu Feng, Haoyu Zhang 0001, Zhe Chen 0013, Yibing Zhan |
Pattern Recognit. | 6 |
| 2022 | Resistance Training Using Prior Bias: Toward Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. To address this problem, we propose Resistance Training using Prior Bias (RTPB) for the scene graph generation. Specifically, RTPB uses a distributed-based prior bias to improve models' detecting ability on less frequent relationships during training, thus improving the model generalizability on tail categories. In addition, to further explore the contextual information of objects and relationships, we design a contextual encoding backbone network, termed as Dual Transformer (DTrans). We perform extensive experiments on a very popular benchmark, VG150, to demonstrate the effectiveness of our method for the unbiased scene graph generation. In specific, our RTPB achieves an improvement of over 10% under the mean recall when applied to current SGG methods. Furthermore, DTrans with RTPB outperforms nearly all state-of-the-art methods with a large margin. Code is available at https://github.com/ChCh1999/RTPB Yibing Zhan, Baosheng Yu, Liu Liu 0014, Yong Luo 0002, Bo Du 0001 |
AAAI | 2 |
| 2022 | HL-Net: Heterophily Learning Network for Scene Graph GenerationabstractScene graph generation (SGG) aims to detect objects and predict their pairwise relationships within an image. Current SGG methods typically utilize graph neural net-works (GNNs) to acquire context information between ob-jects/relationships. Despite their effectiveness, however, current SGG methods only assume scene graph homophily while ignoring heterophily. Accordingly, in this paper, we propose a novel Heterophily Learning Network (HL-Net) to comprehensively explore the homophily and heterophily be-tween objects/relationships in scene graphs. More specif-ically, HL-Net comprises the following 1) an adaptive reweighting transformer module, which adaptively inte-grates the information from different layers to exploit both the heterophily and homophily in objects; 2) a relation-ship feature propagation module that efficiently explores the connections between relationships by considering het-erophily in order to refine the relationship representation; 3) a heterophily-aware message-passing scheme to fur-ther distinguish the heterophily and homophily between ob-jects/relationships, thereby facilitating improved message passing in graphs. We conducted extensive experiments on two public datasets: Visual Genome (VG) and Open Images (OI). The experimental results demonstrate the superiority of our proposed HL-Net over existing state-of-the-art approaches. In more detail, HL-Net outperforms the second-best competitors by 2.1% on the VG datasetfor scene graph classification and 1.2% on the IO dataset for the final score. Code is available at https://github.com/simI3/HL-Net. Changxing Ding, Yibing Zhan, Zijian Li 0011, Dacheng Tao |
CVPR | 3 |
| 2022 | RU-Net: Regularized Unrolling Network for Scene Graph GenerationabstractScene graph generation (SGG) aims to detect objects and predict the relationships between each pair of objects. Existing SGG methods usually suffer from several issues, including 1) ambiguous object representations, as graph neural network-based message passing (GMP) modules are typically sensitive to spurious inter-node correlations, and 2) low diversity in relationship predictions due to severe class imbalance and a large number of missing annotations. To address both problems, in this paper, we propose a regu-larized unrolling network (RU-Net). We first study the relation between GMP and graph Laplacian denoising (GLD) from the perspective of the unrolling technique, determining that GMP can be formulated as a solver for GLD. Based on this observation, we propose an unrolled message passing module and introduce an fp-based graph regularization to suppress spurious connections between nodes. Second, we propose a group diversity enhancement module that pro-motes the prediction diversity of relationships via rank max-imization. Systematic experiments demonstrate that RU-Net is effective under a variety of settings and metrics. Fur-thermore, RU-Net achieves new state-of-the-arts on three popular databases: VG, VRD, and OI. Code is available at https://github.com/siml3/RU-Net. Changxing Ding, Jing Zhang 0037, Yibing Zhan, Dacheng Tao |
CVPR | 4 |
| 2022 | Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersabstractWeakly-supervised semantic segmentation (WSSS) with image-level labels is an important and challenging task. Due to the high training efficiency, end-to-end solutions for WSSS have received increasing attention from the community. However, current methods are mainly based on convolutional neural networks and fail to explore the global information properly, thus usually resulting in incomplete object regions. In this paper, to address the aforementioned problem, we introduce Transformers, which naturally integrate global information, to generate more integral initial pseudo labels for end-to-end WSSS. Motivated by the inherent consistency between the self-attention in Transformers and the semantic affinity, we propose an Affinity from Attention (AFA) module to learn semantic affinity from the multi-head self-attention (MHSA) in Transformers. The learned affinity is then leveraged to refine the initial pseudo labels for segmentation. In addition, to efficiently derive reliable affinity labels for supervising AFA and ensure the local consistency of pseudo labels, we devise a Pixel-Adaptive Refinement module that incorporates low-level image appearance information to refine the pseudo labels. We perform extensive experiments and our method achieves 66.0% and 38.9% mIoU on the PASCAL VOC 2012 and MS COCO 2014 datasets, respectively, significantly outperforming recent end-to-end methods and several multi-stage competitors. Code is available at https://github.com/rulixiang/afa. Lixiang Ru, Yibing Zhan, Baosheng Yu, Bo Du 0001 |
CVPR | 2 |
| 2022 | Contrastive Boundary Learning for Point Cloud SegmentationabstractPoint cloud segmentation is fundamental in understanding 3D environments. However, current 3D point cloud segmentation methods usually perform poorly on scene boundaries, which degenerates the overall segmentation performance. In this paper, we focus on the segmentation of scene boundaries. Accordingly, we first explore metrics to evaluate the segmentation performance on scene boundaries. To address the unsatisfactory performance on boundaries, we then propose a novel contrastive boundary learning (CBL) framework for point cloud segmentation. Specifically, the proposed CBL enhances feature discrimination between points across boundaries by contrasting their representations with the assistance of scene contexts at multiple scales. By applying CBL on three different baseline methods, we experimentally show that CBL consistently improves different baselines and assists them to achieve compelling performance on boundaries, as well as the overall performance, e.g. in mIoU. The experimental results demonstrate the effectiveness of our method and the importance of boundaries for 3D point cloud segmentation. Code and model will be made publicly available at https://github.com/LiyaoTang/contrastBoundary. Liyao Tang, Yibing Zhan, Zhe Chen 0013, Baosheng Yu, Dacheng Tao |
CVPR | 2 |
| 2022 | Learning Graph Neural Networks for Image Style Transfer
Yongcheng Jing, Yining Mao, Yiding Yang, Yibing Zhan, Mingli Song, Xinchao Wang, Dacheng Tao |
ECCV (7) | 4 |
| 2022 | Hierarchical Semi-supervised Contrastive Learning for Contamination-Resistant Anomaly Detection
Gaoang Wang, Yibing Zhan, Xinchao Wang, Mingli Song, Klara Nahrstedt |
ECCV (25) | 2 |
| 2022 | TASA: Deceiving Question Answering Models by Twin Answer Sentences AttackabstractWe present Twin Answer Sentences Attack (TASA), an adversarial attack method for question answering (QA) models that produces fluent and grammatical adversarial contexts while maintaining gold answers.Despite phenomenal progress on general adversarial attacks, few works have investigated the vulnerability and attack specifically for QA models.In this work, we first explore the biases in the existing models and discover that they mainly rely on keyword matching between the question and context, and ignore the relevant contextual relations for answer prediction.Based on two biases above, TASA attacks the target model in two folds: (1) lowering the model's confidence on the gold answer with a perturbed answer sentence; (2) misguiding the model towards a wrong answer with a distracting answer sentence.Equipped with designed beam search and filtering methods, TASA can generate more effective attacks than existing textual attack methods while sustaining the quality of contexts, in extensive experiments on five QA datasets and human evaluations. Yu Cao 0014, Dianqi Li, Tianyi Zhou 0001, Yibing Zhan, Dacheng Tao |
EMNLP | 6 |
| 2022 | Improving Adversarial Robustness via Mutual Information EstimationabstractDeep neural networks (DNNs) are found to be vulnerable to adversarial noise. They are typically misled by adversarial samples to make wrong predictions. To alleviate this negative effect, in this paper, we investigate the dependence between outputs of the target model and input adversarial samples from the perspective of information theory, and propose an adversarial defense method. Specifically, we first measure the dependence by estimating the mutual information (MI) between outputs and the natural patterns of inputs (called natural MI) and MI between outputs and the adversarial patterns of inputs (called adversarial MI), respectively. We find that adversarial samples usually have larger adversarial MI and smaller natural MI compared with those w.r.t. natural samples. Motivated by this observation, we propose to enhance the adversarial robustness by maximizing the natural MI and minimizing the adversarial MI during the training process. In this way, the target model is expected to pay more attention to the natural pattern that contains objective semantics. Empirical evaluations demonstrate that our method could effectively improve the adversarial accuracy against multiple attacks. Dawei Zhou 0004, Nannan Wang 0001, Xinbo Gao 0001, Bo Han 0003, Xiaoyu Wang 0002, Yibing Zhan, Tongliang Liu |
ICML | 6 |
| 2022 | MCFR'22: 1st Workshop on Multimedia Computing towards Fashion RecommendationabstractWith the proliferation of online shopping, fashion recommendation, which aims to provide suitable suggestions to support the consumer's purchase in e-commerce platforms, has gained increasing research attention from both academia and industry. Although existing efforts have achieved great progress, they focus on the visual modality, lacking the exploration of other modalities of items, e.g., the textual descriptions and attributes of items. Accordingly, this workshop targets calling for a coordinated effort to promote the multimedia computing towards fashion recommendation. This workshop will showcase the innovative methodologies and ideas on new yet challenging research problems, including (not limited to) fashion recommendation, interactive fashion recommendation, interactive garment retrieval, and outfit compatibility modeling. Xuemeng Song, Jingjing Chen 0001, Federico Becattini, Weili Guan, Yibing Zhan, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2022 | Knowledge Graph enhanced Multimodal Learning for Few-shot Visual RecognitionabstractFew-shot learning (FSL) aims to learn a classifier for novel classes with only a few labeled samples per category available. The mainstream FSL approaches fall in the meta-learning paradigm, where a meta-learner is used to learn transferable knowledge and generalize to new tasks. However, these approaches usually only leverage information from a single modality (e.g., visual image) and fail to explore the information from other modalities (e.g., the knowledge graph). Since the labeled samples are scarce in FSL, increasing the information for each example is a possible solution to improve the performance. This motivates us to develop a new meta-learning framework for few-shot visual recognition termed Knowledge Graph enhanced FSL (KGFSL), which combines the information from multiple modalities: 1) the visual information in images and 2) the rich semantics and structural information in a knowledge graph (KG). Specifically, KGFSL exploits the word embedding of the category and its relationship to other categories to improve the visual-based models. A graph convolutional network (GCN) is first introduced to learn the semantic embeddings for each node (a visual category) in KG. The visual and semantic embeddings are then aligned and combined for final prediction. Finally, the whole framework is trained in an end-to-end manner. We conduct extensive experiments on two widely-used FSL benchmarks: miniImageNet and tieredImageNet. Experimental results demonstrate the effectiveness of the multimodal information for few-shot learning, and our proposed method can significantly outperform the state-of-the-art approaches. Mengya Han, Yibing Zhan, Baosheng Yu, Yong Luo 0002, Bo Du 0001, Dacheng Tao |
MMSP | 2 |
| 2022 | Estimating Noise Transition Matrix with Label Correlations for Noisy Multi-Label LearningabstractIn label-noise learning, the noise transition matrix, bridging the class posterior for noisy and clean data, has been widely exploited to learn statistically consistent classifiers. The effectiveness of these algorithms relies heavily on estimating the transition matrix. Recently, the problem of label-noise learning in multi-label classification has received increasing attention, and these consistent algorithms can be applied in multi-label cases. However, the estimation of transition matrices in noisy multi-label learning has not been studied and remains challenging, since most of the existing estimators in noisy multi-class learning depend on the existence of anchor points and the accurate fitting of noisy class posterior. To address this problem, in this paper, we first study the identifiability problem of the class-dependent transition matrix in noisy multi-label learning, and then inspired by the identifiability results, we propose a new estimator by exploiting label correlations without neither anchor points nor accurate fitting of noisy class posterior. Specifically, we estimate the occurrence probability of two noisy labels to get noisy label correlations. Then, we perform sample selection to further extract information that implies clean label correlations, which is used to estimate the occurrence probability of one noisy label when a certain clean label appears. By utilizing the mismatch of label correlations implied in these occurrence probabilities, the transition matrix is identifiable, and can then be acquired by solving a simple bilinear decomposition problem. Empirical results demonstrate the effectiveness of our estimator to estimate the transition matrix with label correlations, leading to better classification performance. Source codes are available at https://github.com/tmllab/Multi-Label-T. Shikun Li, Xiaobo Xia, Hansong Zhang 0003, Yibing Zhan, Shiming Ge, Tongliang Liu |
NeurIPS | 4 |
| 2022 | Pluralistic Image Completion with Gaussian Mixture ModelsabstractPluralistic image completion focuses on generating both visually realistic and diverse results for image completion. Prior methods enjoy the empirical successes of this task. However, their used constraints for pluralistic image completion are argued to be not well interpretable and unsatisfactory from two aspects. First, the constraints for visual reality can be weakly correlated to the objective of image completion or even redundant. Second, the constraints for diversity are designed to be task-agnostic, which causes the constraints to not work well. In this paper, to address the issues, we propose an end-to-end probabilistic method. Specifically, we introduce a unified probabilistic graph model that represents the complex interactions in image completion. The entire procedure of image completion is then mathematically divided into several sub-procedures, which helps efficient enforcement of constraints. The sub-procedure directly related to pluralistic results is identified, where the interaction is established by a Gaussian mixture model (GMM). The inherent parameters of GMM are task-related, which are optimized adaptively during training, while the number of its primitives can control the diversity of results conveniently. We formally establish the effectiveness of our method and demonstrate it with comprehensive experiments. The implementationis available at https://github.com/tmllab/PICMM. Xiaobo Xia, Yewen Li, Yibing Zhan, Bo Han 0003, Tongliang Liu |
NeurIPS | 5 |
| 2022 | Masked Graph Auto-Encoder Constrained Graph Pooling
Chuang Liu 0008, Yibing Zhan, Xueqi Ma, Dapeng Tao, Bo Du 0001, Wenbin Hu 0001 |
ECML/PKDD (2) | 2 |
| 2022 | Where Does the Performance Improvement Come From?: - A Reproducibility Concern about Image-Text RetrievalabstractThis article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models. Due to the increase of multimodal data over the last decade, image-text retrieval has steadily become a major research direction in the field of information retrieval. Numerous researchers train and evaluate image-text retrieval algorithms using benchmark datasets such as MS-COCO and Flickr30k. Research in the past has mostly focused on performance, with multiple state-of-the-art methodologies being suggested in a variety of ways. According to their assertions, these techniques provide improved modality interactions and hence more precise multimodal representations. In contrast to previous works, we focus on the reproducibility of the approaches and the examination of the elements that lead to improved performance by pretrained and nonpretrained models in retrieving images and text. Jun Rao, Fei Wang 0032, Liang Ding 0006, Shuhan Qi, Yibing Zhan, Weifeng Liu 0001, Dacheng Tao |
SIGIR | 5 |
| 2022 | Dual-branch Density Ratio Estimation for Signed Network EmbeddingabstractSigned network embedding (SNE) has received considerable attention in recent years. A mainstream idea of SNE is to learn node representations by estimating the ratio of sampling densities. Though achieving promising performance, these methods based on density ratio estimation are limited to the issues of confusing sample, expected error, and fixed priori. To alleviate the above-mentioned issues, in this paper, we propose a novel dual-branch density ratio estimation (DDRE) architecture for SNE. Specifically, DDRE 1) consists of a dual-branch network, dealing with the confusing sample; 2) proposes the expected matrix factorization without sampling to avoid the expected error; and 3) devises an adaptive cross noise sampling to alleviate the fixed priori. We perform sign prediction and node classification experiments on four real-world and three artificial datasets, respectively. Extensive empirical results demonstrate that DDRE not only significantly outperforms the methods based on density ratio estimation but also achieves competitive performance compared with other types of methods such as graph likelihood, generative adversarial networks, and graph convolutional networks. Code is publicly available at https://github.com/WHU-SNA/DDRE. Pinghua Xu, Yibing Zhan, Liu Liu 0014, Baosheng Yu, Bo Du 0001, Jia Wu 0001, Wenbin Hu 0001 |
WWW | 2 |
| 2022 | Weakly-Supervised Semantic Segmentation with Visual Words Learning and Hybrid Pooling
Lixiang Ru, Bo Du 0001, Yibing Zhan, Chen Wu 0003 |
Int. J. Comput. Vis. | 3 |
| 2022 | An End-to-end Supervised Domain Adaptation Framework for Cross-Domain Change Detection
Wenjie Xuan, Yuhang Gan, Yibing Zhan, Juhua Liu, Bo Du 0001 |
Pattern Recognit. | 4 |
| 2022 | Multi-level graph learning network for hyperspectral image classification
Sheng Wan, Shirui Pan, Shengwei Zhong 0001, Jie Yang 0002, Jian Yang 0003, Yibing Zhan, Chen Gong 0002 |
Pattern Recognit. | 6 |
| 2022 | Deep relational self-Attention networks for scene graph generation
Zhou Yu 0001, Yibing Zhan |
Pattern Recognit. Lett. | 3 |
| 2022 | Divide-and-Conquer Predictor for Unbiased Scene Graph GenerationabstractScene Graph Generation (SGG) aims to detect the objects and their pairwise predicates in an image. Existing SGG methods mainly fulfil the challenging predicate prediction task that involves severe long-tailed data distribution with a single classifier. However, we argue that this may be enough to differentiate predicates that present obvious differences (e.g.,$on$and$near$), but not sufficient to distinguish similar predicates that only have subtle differences (e.g.,$on$and$standing~on$). Towards this end, we divide the predicate prediction into a few sub-tasks with a Divide-and-Conquer Predictor (DC-Predictor). Specifically, we first develop an offline pattern-predicate correlation mining algorithm to discover the similar predicates that share the same object interaction pattern. Based on that, we devise a general pattern classifier and a set of specific predicate classifiers for DC-Predictor. The former works on recognizing the pattern of a given object pair and routing it to the corresponding specific predicate classifier, while the latter aims to differentiate similar predicates in each specific pattern. In addition, we introduce the Bayesian Personalized Ranking loss in each specific predicate classifier to enhance the pairwise differentiation between head predicates and their similar ones. Experiments on VG150 and GQA datasets show the superiority of our model over state-of-the-art methods. Xianjing Han, Xingning Dong, Xuemeng Song, Tian Gan 0002, Yibing Zhan, Yan Yan 0002, Liqiang Nie |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | BiN-Flow: Bidirectional Normalizing Flow for Robust Image DehazingabstractImage dehazing aims to remove haze in images to improve their image quality. However, most image dehazing methods heavily depend on strict prior knowledge and paired training strategy, which would hinder generalization and performance when dealing with unseen scenes. In this paper, to address the above problem, we propose Bidirectional Normalizing Flow (BiN-Flow), which exploits no prior knowledge and constructs a neural network through weakly-paired training with better generalization for image dehazing. Specifically, BiN-Flow designs 1) Feature Frequency Decoupling (FFD) for mining the various texture details through multi-scale residual blocks and 2) Bidirectional Propagation Flow (BPF) for exploiting the one-to-many relationships between hazy and haze-free images using a sequence of invertible Flow. In addition, BiN-Flow constructs a reference mechanism (RM) that uses a small number of paired hazy and haze-free images and a large number of haze-free reference images for weakly-paired training. Essentially, the mutual relationships between hazy and haze-free images could be effectively learned to further improve the generalization and performance for image dehazing. We conduct extensive experiments on five commonly-used datasets to validate the BiN-Flow. The experimental results that BiN-Flow outperforms all state-of-the-art competitors demonstrate the capability and generalization of our BiN-Flow. Besides, our BiN-Flow could produce diverse dehazing images for the same image by considering restoration diversity. Yiqiang Wu, Dapeng Tao, Yibing Zhan, Chenyang Zhang 0003 |
IEEE Trans. Image Process. | 3 |
| 2021 | Deep Graph-neighbor Coherence Preserving Network for Unsupervised Cross-modal HashingabstractUnsupervised cross-modal hashing (UCMH) has become a hot topic recently. Current UCMH focuses on exploring data similarities. However, current UCMH methods calculate the similarity between two data, mainly relying on the two data's cross-modal features. These methods suffer from inaccurate similarity problems that result in a suboptimal retrieval Hamming space, because the cross-modal features between the data are not sufficient to describe the complex data relationships, such as situations where two data have different feature representations but share the inherent concepts. In this paper, we devise a deep graph-neighbor coherence preserving network (DGCPN). Specifically, DGCPN stems from graph models and explores graph-neighbor coherence by consolidating the information between data and their neighbors. DGCPN regulates comprehensive similarity preserving losses by exploiting three types of data similarities (i.e., the graph-neighbor coherence, the coexistent similarity, and the intra- and inter-modality consistency) and designs a half-real and half-binary optimization strategy to reduce the quantization errors during hashing. Essentially, DGCPN addresses the inaccurate similarity problem by exploring and exploiting the data's intrinsic relationships in a graph. We conduct extensive experiments on three public UCMH datasets. The experimental results demonstrate the superiority of DGCPN, e.g., by improving the mean average precision from 0.722 to 0.751 on MIRFlickr-25K using 64-bit hashing codes to retrieval texts from images. We will release the source code package and the trained model on https://github.com/Atmegal/DGCPN. Jun Yu 0002, Yibing Zhan, Dacheng Tao |
AAAI | 3 |
| 2021 | Not All Operations Contribute Equally: Hierarchical Operation-adaptive Predictor for Neural Architecture SearchabstractGraph-based predictors have recently shown promising results on neural architecture search (NAS). Despite their efficiency, current graph-based predictors treat all operations equally, resulting in biased topological knowledge of cell architectures. Intuitively, not all operations are equally significant during forwarding propagation when aggregating information from these operations to another operation. To address the above issue, we propose a Hierarchical Operation-adaptive Predictor (HOP) for NAS. HOP contains an operation-adaptive attention module (OAM) to capture the diverse knowledge between operations by learning the relative significance of operations in cell architectures during aggregation over iterations. In addition, a cell-hierarchical gated module (CGM) further refines and enriches the obtained topological knowledge of cell architectures, by integrating cell information from each iteration of OAM. The experimental results compared with state-of-the-art predictors demonstrate the capability of our proposed HOP. Ziye Chen, Yibing Zhan, Baosheng Yu, Mingming Gong, Bo Du 0001 |
ICCV | 2 |
| 2021 | A Question Answering System for Unstructured Table ImagesabstractQuestion answering over tables is a very popular semantic parsing task in natural language processing (NLP). However, few existing methods focus on table images, even though there are usually large-scale unstructured tables in practice (e.g., table images). Table parsing from images is nontrivial since it is closely related to not only NLP but also computer vision (CV) to parse the tabular structure from an image. In this demo, we present a question answering system for unstructured table images. The proposed system mainly consists of 1) a table recognizer to recognize the tabular structure from an image and 2) a table parser to generate the answer to a natural language question over the table. In addition, to train the model, we further provide table images and structure annotations for two widely used semantic parsing datasets. Specifically, the test set is used for this demo, from where the users can either choose from default questions or enter a new custom question. Wenyuan Xue, Wen Wang 0019, Qingyong Li, Baosheng Yu, Yibing Zhan, Dacheng Tao |
ACM Multimedia | 6 |
| 2021 | Collocation and Try-on Network: Whether an Outfit is CompatibleabstractWhether an outfit is compatible? Using machine learning methods to assess an outfit's compatibility, namely, fashion compatibility modeling (FCM), has recently become a popular yet challenging topic. However, current FCM studies still perform far from satisfactory, because they only consider the collocation compatibility modeling, while neglecting the natural human habits that people generally evaluate outfit compatibility from both the collocation (discrete assess) and the try-on (unified assess) perspectives. In light of the above analysis, we propose a Collocation and Try-On Network (CTO-Net) for FCM, combining both the collocation and try-on compatibilities. In particular, for the collocation perspective, we devise a disentangled graph learning scheme, where the collocation compatibility is disentangled into multiple fine-grained compatibilities between items; regarding the try-on perspective, we propose an integrated distillation learning scheme to unify all item information in the whole outfit to evaluate the compatibility based on the latent try-on representation. To further enhance the collocation and try-on compatibilities, we exploit the mutual learning strategy to obtain a more comprehensive judgment. Extensive experiments on the real-world dataset demonstrate that our CTO-Net significantly outperforms the state-of-the-art methods. In particular, compared with the competitive counterparts, our proposed CTO-Net significantly improves AUC accuracy from 83.2% to 87.8% and MRR from 15.4% to 21.8%. We have released our source codes and trained models to benefit other researchers.1 Xuemeng Song, Qingying Niu, Yibing Zhan, Liqiang Nie |
ACM Multimedia | 5 |
| 2021 | Contrastive Graph Poisson Networks: Semi-Supervised Learning with Extremely Limited LabelsabstractGraph Neural Networks (GNNs) have achieved remarkable performance in the task of semi-supervised node classification. However, most existing GNN models require sufficient labeled data for effective network training. Their performance can be seriously degraded when labels are extremely limited. To address this issue, we propose a new framework termed Contrastive Graph Poisson Networks (CGPN) for node classification under extremely limited labeled data. Specifically, our CGPN derives from variational inference; integrates a newly designed Graph Poisson Network (GPN) to effectively propagate the limited labels to the entire graph and a normal GNN, such as Graph Attention Network, that flexibly guides the propagation of GPN; applies a contrastive objective to further exploit the supervision information from the learning process of GPN and GNN models. Essentially, our CGPN can enhance the learning performance of GNNs under extremely limited labels by contrastively propagating the limited labels to the entire graph. We conducted extensive experiments on different types of datasets to demonstrate the superiority of CGPN. Sheng Wan, Yibing Zhan, Liu Liu 0014, Baosheng Yu, Shirui Pan, Chen Gong 0002 |
NeurIPS | 2 |
| 2021 | Comprehensive Linguistic-Visual Composition Network for Image RetrievalabstractComposing text and image for image retrieval (CTI-IR) is a new yet challenging task, for which the input query is not the conventional image or text but a composition, i.e., a reference image and its corresponding modification text. The key of CTI-IR lies in how to properly compose the multi-modal query to retrieve the target image. In a sense, pioneer studies mainly focus on composing the text with either the local visual descriptor or global feature of the reference image. However, they overlook the fact that the text modifications are indeed diverse, ranging from the concrete attribute changes, like "change it to long sleeves", to the abstract visual property adjustments, e.g., "change the style to professional". Thus, simply emphasizing the local or global feature of the reference image for the query composition is insufficient. In light of the above analysis, we propose a Comprehensive Linguistic-Visual Composition Network (CLVC-Net) for image retrieval. The core of CLVC-Net is that it designs two composition modules: fine-grained local-wise composition module and fine-grained global-wise composition module, targeting comprehensive multi-modal compositions. Additionally, a mutual enhancement module is designed to promote local-wise and global-wise composition processes by forcing them to share knowledge with each other. Extensive experiments conducted on three real-world datasets demonstrate the superiority of our CLVC-Net. We released the codes to benefit other researchers. Haokun Wen, Xuemeng Song, Xin Yang 0008, Yibing Zhan, Liqiang Nie |
SIGIR | 4 |
| 2020 | Graph Pattern Loss Based Diversified Attention Network For Cross-Modal RetrievalabstractCross-modal retrieval aims to enable flexible retrieval experience by combining multimedia data such as image, video, text, and audio. One core of unsupervised approaches is to dig the correlations among different object representations to complete satisfied retrieval performance without requiring expensive labels. In this paper, we propose a Graph Pattern Loss based Diversified Attention Network (GPLDAN) for unsupervised cross-modal retrieval to deeply analyze correlations among representations. First, we propose a diversified attention feature projector by considering the interaction between different representations to generate multiple representations of an instance. Then, we design a novel graph pattern loss to explore the correlations among different representations, in this graph all possible distances between different representations are considered. In addition, a modality classifier is added to explicitly declare the corresponding modalities of features before fusion and guide the network to enhance discrimination ability. We test GPLDAN on four public datasets. Compared with the state-of-the-art cross-modal retrieval methods, the experimental results demonstrate the performance and competitiveness of GPLDAN. Xueying Chen, Rong Zhang 0004, Yibing Zhan |
ICIP | 3 |
| 2020 | Relationship graph learning network for visual relationship detectionabstractVisual relationship detection aims to predict the relationships between detected object pairs. It is well believed that the correlations between image components (i.e., objects and relationships between objects) are significant considerations when predicting objects' relationships. However, most current visual relationship detection methods only exploited the correlations among objects, and the correlations among objects' relationships remained underexplored. This paper proposes a relationship graph learning network (RGLN) to explore the correlations among objects' relationships for visual relationship detection. Specifically, RGLN obtains image objects using an object detector, and then, every pair of objects constitutes a relationship proposal. All relationship proposals construct a relationship graph, in which the proposals are treated as nodes. Accordingly, RGLN designs bi-stream graph attention subnetworks to detect relationship proposals, in which one graph attention subnetwork analyzes correlations among relationships based on visual and spatial information, and the other analyzes correlations based on semantic and spatial information. Besides, RGLN exploits a relationship selection subnetwork to ignore redundant information of object pairs with no relationships. We conduct extensive experiments on two public datasets: the VRD and the VG datasets. The experimental results compared with the state-of-the-art demonstrate the competitiveness of RGLN. Jun Yu 0002, Yibing Zhan, Zhi Chen 0010 |
MMAsia | 3 |
| 2020 | Multi-task Compositional Network for Visual Relationship Detection
Yibing Zhan, Jun Yu 0002, Ting Yu 0016, Dacheng Tao |
Int. J. Comput. Vis. | 1 |
| 2019 | On Exploring Undetermined Relationships for Visual Relationship DetectionabstractIn visual relationship detection, human-notated relationships can be regarded as determinate relationships. However, there are still large amount of unlabeled data, such as object pairs with less significant relationships or even with no relationships. We refer to these unlabeled but potentially useful data as undetermined relationships. Although a vast body of literature exists, few methods exploit these undetermined relationships for visual relationship detection. In this paper, we explore the beneficial effect of undetermined relationships on visual relationship detection. We propose a novel multi-modal feature based undetermined relationship learning network (MF-URLN) and achieve great improvements in relationship detection. In detail, our MF-URLN automatically generates undetermined relationships by comparing object pairs with human-notated data according to a designed criterion. Then, the MF-URLN extracts and fuses features of object pairs from three complementary modals: visual, spatial, and linguistic modals. Further, the MF-URLN proposes two correlated subnetworks: one subnetwork decides the determinate confidence, and the other predicts the relationships. We evaluate the MF-URLN on two datasets: the Visual Relationship Detection (VRD) and the Visual Genome (VG) datasets. The experimental results compared with state-of-the-art methods verify the significant improvements made by the undetermined relationships, e.g., the top-50 relation detection recall improves from 19.5% to 23.9% on the VRD dataset. Yibing Zhan, Jun Yu 0002, Ting Yu 0016, Dacheng Tao |
CVPR | 1 |
| 2018 | Attention-Based Convolutional Neural Network for the Detection of Built-Up Areas in High-Resolution SAR ImagesabstractThe detection of built-up areas is essential for high-resolution Synthetic Aperture Rader (SAR) applications, such as urban planning and environmental monitoring. In this paper, we proposed an attention based convolutional neural network for the detection of built-up areas in SAR images. Our network composes of two parts. First part contains two branches, one is designed to obtain high detection rate and the other one is designed to obtain low false alarm rate. Second part aims to merge the advantages of the two detection results of first part by using attention mechanism. Experiments on TerraSAR-X high resolution SAR images over Beijing demonstrate the effectiveness of proposed method. Rong Zhang 0004, Yibing Zhan |
IGARSS | 3 |
| 2018 | Comprehensive Distance-Preserving Autoencoders for Cross-Modal RetrievalabstractIn this paper, we propose a novel method with comprehensive distance-preserving autoencoders (CDPAE) to address the problem of unsupervised cross-modal retrieval. Previous unsupervised methods rely primarily on pairwise distances of representations extracted from cross media spaces that co-occur and belong to the same objects. However, besides pairwise distances, the CDPAE also considers heterogeneous distances of representations extracted from cross media spaces as well as homogeneous distances of representations extracted from single media spaces that belong to different objects. The CDPAE consists of four components. First, denoising autoencoders are used to retain the information from the representations and to reduce the negative influence of redundant noises. Second, a comprehensive distance-preserving common space is proposed to explore the correlations among different representations. This aims to preserve the respective distances between the representations within the common space so that they are consistent with the distances in their original media spaces. Third, a novel joint loss function is defined to simultaneously calculate the reconstruction loss of the denoising autoencoders and the correlation loss of the comprehensive distance-preserving common space. Finally, an unsupervised cross-modal similarity measurement is proposed to further improve the retrieval performance. This is carried out by calculating the marginal probability of two media objects based on a kNN classifier. The CDPAE is tested on four public datasets with two cross-modal retrieval tasks: "query images by texts" and "query texts by images". Compared with eight state-of-the-art cross-modal retrieval methods, the experimental results demonstrate that the CDPAE outperforms all the unsupervised methods and performs competitively with the supervised methods. Yibing Zhan, Jun Yu 0002, Zhou Yu 0001, Rong Zhang 0004, Dacheng Tao, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2018 | No-Reference Image Sharpness Assessment Based on Maximum Gradient and Variability of GradientsabstractGradients are commonly used in image sharpness assessment methods. However, research has not fully addressed the direct relationship between gradients and the perceived sharpness. In this paper, we discover and validate through experiments that the maximum gradient is an effective indicator of the perceived image sharpness on a global or local scale. Based on these observations, we propose a novel and efficient no-reference image quality assessment (NR-IQA) method for blurry images. Our method uses two elements to predict the quality of blurry images: the maximum gradient and the variability of gradients. The maximum gradient represents the sharpest spot in an image, and the variability of gradients shows variations within the content of the image. According to the characteristics of human visual systems, these factors are significant for humans when judging the quality of blurry images. The method was tested using blurry image datasets from five public IQA databases. Compared with nine other state-of-the-art NR-IQA methods for blurry images, the experimental results demonstrate that our method is more consistent with humans' subjective evaluations. The MATLAB source code of our method is available at https://github.com/Atmegal/Sharpness-evaluation. Yibing Zhan, Rong Zhang 0004 |
IEEE Trans. Multim. | 1 |
| 2017 | No-Reference JPEG Image Quality Assessment Based on Blockiness and Luminance ChangeabstractWhen scoring the quality of JPEG images, the two main considerations for viewers are blocking artifacts and improper luminance changes, such as blur. In this letter, we first propose two measures to estimate the blockiness and the luminance change within individual blocks. Then, a no-reference image quality assessment (NR-IQA) method for JPEG images is proposed. Our method obtains the quality score by considering the blocking artifacts and the luminance changes from all nonoverlapping 8 × 8 blocks in one JPEG image. The proposed method has been tested on five public IQA databases and compared with five state-of-the-art NR-IQA methods for JPEG images. The experimental results show that our method is more consistent with subjective evaluations than the state-of-the-art NR-IQA methods. The MATLAB source code of our method is available at http://image.ustc.edu.cn/IQA.html. Yibing Zhan, Rong Zhang 0004 |
IEEE Signal Process. Lett. | 1 |
| 2017 | A Structural Variation Classification Model for Image Quality AssessmentabstractStructural information is critical for image quality assessment (IQA). In this paper, first, we propose a novel model of structural variations in images. The proposed model classifies the types of structural variation within images into four categories: slight deformations, additive impairments, detail losses, and confusing contents. This system of classification applies to most types of structural variations observed in practice. In this model, each pixel from the distorted images is classified according to its structural variation using fuzzy logic based on a set of structural features extracted from the images. Then, a novel IQA method based on these pixel classifications is proposed. This proposed method evaluates the image quality by combining two aspects: the distribution of different structural variations and the degree of structural differences. We test the proposed method using seven public databases. The experimental results indicate that our method is more consistent with the results of the subjective evaluation than were the nine other state-of-the-art IQA methods. The MATLAB source code of our method is available at http://image.ustc.edu.cn/IQA.html. Yibing Zhan, Rong Zhang 0004 |
IEEE Trans. Multim. | 1 |
| 2016 | A new haze image database with detailed air quality information and a novel no-reference image quality assessment method for haze imagesabstractIn this paper, we propose a new standard haze image database with nearly all kinds of haze situations. Our database includes haze-free images as well as different levels and situations of haze images, such as snowy and extremely serious haze images. Our database also records the related weather and air quality information. Moreover, the database offers a mean opinion score (MOS) for each image as the subjective evaluation of the haze severity. Our database provides ample haze images with various objective and subjective descriptions, and has significantly potential scientific value. Since haze is a reason leading to the degradation of the image quality, obtaining the quality of haze images is also necessary and meaningful. In this paper, a novel no-reference image quality assessment (IQA) method for haze images is also proposed. Our method analyzes the model of haze images to evaluate the quality. The experimental results on the haze database show that our method is consistent with the subjective evaluation. Yibing Zhan, Rong Zhang 0004 |
ICASSP | 1 |
| 2016 | A novel structural variation detection strategy for image quality assessmentabstractStructural information is critical in image quality assessment (IQA). Although existing objective IQA methods have achieved high consistency with subjective perception, detecting structural variation remains a difficult task. In this paper, we propose a novel structural variation detection strategy that is based on binary logic and inspired by the bag-of-words model. The proposed strategy detects structural variation by comparing the occurrences of structural features within the original and distorted images. In order to show the effectiveness of this strategy, this paper also proposes a novel and simple IQA method based on this strategy. The proposed method evaluates the image quality from two aspects: the structure distortion and the luminance distortion. The experimental results from four public databases show that the proposed method is highly congruous with subjective evaluation. The results also prove that the detection strategy is useful. Yibing Zhan, Rong Zhang 0004 |
ICIP | 1 |
| 2015 | Image quality assessment based on structure variance classificationabstractIn this paper, we find that the structure variance of images could be divided into four classifications, slight deformations, additive impairments, detail losses, and confusing contents, and what's more, for each classification, subjective evaluation is different. According this, we propose a novel image quality assessment (IQA) method based on structure variance classification. The proposed method classifies the structure variance of each patch into one of the four classifications using binary logic and then summarizes the areas of different classifications. To get more comprehensive evaluation, the proposed method also incorporates the measurements of differences between extracted features. Our method is tested on five public databases and compared with seven state-of-art methods. The experimental results demonstrate that our method can achieve higher consistency in relation to the subjective evaluation compared to the state-of-art IQA methods. Yibing Zhan, Rong Zhang 0004 |
ICASSP | 1 |
| 2014 | Image quality assessment using a SVD-based structural projection
Anzhou Hu, Rong Zhang 0004, Yibing Zhan |
Signal Process. Image Commun. | 4 |