EDBT 2026 Demo / reviewers in the wild / expert
Bing Su 0001
dblp:41/5270-1
· DBLP profile ↗
67ranked-venue papers
19as first author
48since 2021 · last 2026
0000-0001-8560-1910ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 16 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 8 first-author · 23 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uniformity Preserving Transfer for Visual Prompt Tuning under Long-tailed Distribution
Hao Chen 0102, Bin Qin 0001, Jiangmeng Li, Jindong Wang 0001, Bing Su 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | Spatio-Temporal Multi-Subgraph GCN for 3D Human Motion PredictionabstractHuman motion prediction (HMP) involves forecasting future human motion based on historical data. Graph Convolutional Networks (GCNs) have garnered widespread attention in this field for their proficiency in capturing relationships among joints in human motion. However, existing GCN-based methods tend to focus on either temporal-domain or spatial-domain features, or they combine spatio-temporal features without fully leveraging the complementarity and cross-dependency of these two features. In this paper, we propose the Spatial-Temporal Multi-Subgraph Graph Convolutional Network (STMS-GCN) to capture complex spatio-temporal dependencies in human motion. Specifically, we decouple the modeling of temporal and spatial dependencies, enabling cross-domain knowledge transfer at multiple scales through a spatio-temporal information consistency constraint mechanism. Besides, we utilize multiple subgraphs to extract richer motion information and enhance the learning associations of diverse subgraphs through a homogeneous information constraint mechanism. Extensive experiments on the standard HMP benchmarks demonstrate the superiority of our method. Jiexin Wang 0003, Yiju Guo, Bing Su 0001 |
ICASSP | 3 |
| 2025 | Temporal Dynamics Decoupling with Inverse Processing for Enhancing Human Motion PredictionabstractExploring the bridge between historical and future motion behaviors remains a central challenge in human motion prediction. While most existing methods incorporate a reconstruction task as an auxiliary task into the decoder, thereby improving the modeling of spatio-temporal dependencies, they overlook the potential conflicts between reconstruction and prediction tasks. In this paper, we propose a novel approach: Temporal Decoupling Decoding with Inverse Processing (TD2IP). Our method strategically separates reconstruction and prediction decoding processes, employing distinct decoders to decode the shared motion features into historical or future sequences. Additionally, inverse processing reverses motion information in the temporal dimension and reintroduces it into the model, leveraging the bidirectional temporal correlation of human motion behaviors. By alleviating the conflicts between reconstruction and prediction tasks and enhancing the association of historical and future information, TD2IP fosters a deeper understanding of motion patterns. Extensive experiments demonstrate the adaptability of our method within existing methods. Jiexin Wang 0003, Yiju Guo, Bing Su 0001 |
ICASSP | 3 |
| 2025 | Enhancing Reward Models for High-Quality Image Generation: Beyond Text-Image Alignment
Ying Ba, Yalong Bai, Wenyi Mo, Bing Su 0001, Ji-Rong Wen |
ICCV | 6 |
| 2025 | Regulatory DNA Sequence Design with Reinforcement Learningabstract$\textit{Cis}$-regulatory elements (CREs), such as promoters and enhancers, are relatively short DNA sequences that directly regulate gene expression. The fitness of CREs, measured by their ability to modulate gene expression, highly depends on the nucleotide sequences, especially specific motifs known as transcription factor binding sites (TFBSs). Designing high-fitness CREs is crucial for therapeutic and bioengineering applications. Current CRE design methods are limited by two major drawbacks: (1) they typically rely on iterative optimization strategies that modify existing sequences and are prone to local optima, and (2) they lack the guidance of biological prior knowledge in sequence optimization. In this paper, we address these limitations by proposing a generative approach that leverages reinforcement learning (RL) to fine-tune a pre-trained autoregressive (AR) model. Our method incorporates data-driven biological priors by deriving computational inference-based rewards that simulate the addition of activator TFBSs and removal of repressor TFBSs, which are then integrated into the RL process. We evaluate our method on promoter design tasks in two yeast media conditions and enhancer design tasks for three human cell types, demonstrating its ability to generate high-fitness CREs while maintaining sequence diversity. The code is available at https://github.com/yangzhao1230/TACO. Zhao Yang 0006, Bing Su 0001, Chuan Cao, Ji-Rong Wen |
ICLR | 2 |
| 2025 | DenoiseVAE: Learning Molecule-Adaptive Noise Distributions for Denoising-based 3D Molecular Pre-trainingabstractDenoising learning of 3D molecules learns molecular representations by imposing noises into the equilibrium conformation and predicting the added noises to recover the equilibrium conformation, which essentially captures the information of molecular force fields. Due to the specificity of Potential Energy Surfaces, the probabilities of physically reasonable noises for each atom in different molecules are different. However, existing methods apply the shared heuristic hand-crafted noise sampling strategy to all molecules, resulting in inaccurate force field learning. In this paper, we propose a novel 3D molecular pre-training method, namely DenoiseVAE, which employs a Noise Generator to acquire atom-specific noise distributions for different molecules. It utilizes the stochastic reparameterization technique to sample noisy conformations from the generated distributions, which are inputted into a Denoising Module for denoising. The Noise Generator and the Denoising Module are jointly learned in a manner conforming with the paradigm of Variational Auto Encoder. Consequently, the sampled noisy conformations can be more diverse, adaptive, and informative, and thus DenoiseVAE can learn representations that better reveal the molecular force fields. Extensive experiments show that DenoiseVAE outperforms the current state-of-the-art methods on various molecular property prediction tasks, demonstrating the effectiveness of it. Yurou Liu, Jiangmeng Li, Wenbing Huang 0001, Bing Su 0001 |
ICLR | 6 |
| 2025 | Enhancing Human Motion Prediction via Multi-range Decoupling Decoding with Gating-adjusting AggregationabstractExpressive representation of pose sequences is crucial for accurate motion modeling in human motion prediction (HMP). While recent deep learning-based methods have shown promise in learning motion representations, these methods tend to overlook the varying relevance and dependencies between historical information and future moments, with a stronger correlation for short-term predictions and weaker for distant future predictions. This limits the learning of motion representation and then hampers prediction performance. In this paper, we propose a novel approach called multi-range decoupling decoding with gating-adjusting aggregation (MD2GA), which leverages the temporal correlations to refine motion representation learning. This approach employs a two-stage strategy for HMP. In the first stage, a multi-range decoupling decoding adeptly adjusts feature learning by decoding the shared features into distinct future lengths, where different decoders offer diverse insights into motion patterns. In the second stage, a gating-adjusting aggregation dynamically combines the diverse insights guided by input motion data. Extensive experiments demonstrate that the proposed method can be easily integrated into other motion prediction methods and enhance their prediction performance. Jiexin Wang 0003, Wenwen Qiang, Zhao Yang 0006, Bing Su 0001 |
ICME | 4 |
| 2025 | SPACE: Your Genomic Profile Predictor is a Powerful DNA Foundation ModelabstractInspired by the success of unsupervised pre-training paradigms, researchers have applied these approaches to DNA pre-training. However, we argue that these approaches alone yield suboptimal results because pure DNA sequences lack sufficient information, since their functions are regulated by genomic profiles like chromatin accessibility. Here, we demonstrate that supervised training for genomic profile prediction serves as a more effective alternative to pure sequence pre-training. Furthermore, considering the multi-species and multi-profile nature of genomic profile prediction, we introduce our Species-Profile Adaptive Collaborative Experts (SPACE) that leverages Mixture of Experts (MoE) to better capture the relationships between DNA sequences across different species and genomic profiles, thereby learning more effective DNA representations. Through extensive experiments across various tasks, our model achieves state-of-the-art performance, establishing that DNA models trained with supervised genomic profiles serve as powerful DNA representation learners. Zhao Yang 0006, Jiwei Zhu, Bing Su 0001 |
ICML | 3 |
| 2025 | Rethinking the Bias of Foundation Model under Long-tailed DistributionabstractLong-tailed learning has garnered increasing attention due to its practical significance. Among the various approaches, the fine-tuning paradigm has gained considerable interest with the advent of foundation models. However, most existing methods primarily focus on leveraging knowledge from these models, overlooking the inherent biases introduced by the imbalanced training data they rely on. In this paper, we examine how such imbalances from pre-training affect long-tailed downstream tasks. Specifically, we find the imbalance biases inherited in foundation models on downstream task as parameter imbalance and data imbalance. During fine-tuning, we observe that parameter imbalance plays a more critical role, while data imbalance can be mitigated using existing re-balancing strategies. Moreover, we find that parameter imbalance cannot be effectively addressed by current re-balancing techniques, such as adjusting the logits, during training, unlike data imbalance. To tackle both imbalances simultaneously, we build our method on causal learning and view the incomplete semantic factor as the confounder, which brings spurious correlations between input samples and labels. To resolve the negative effects of this, we propose a novel backdoor adjustment method that learns the true causal effect between input samples and labels, rather than merely fitting the correlations in the data. Notably, we achieve an average performance increase of about 1.67% on each dataset. Bin Qin 0001, Jiangmeng Li, Hao Chen 0102, Bing Su 0001 |
ICML | 5 |
| 2025 | Large Language-Geometry Model: When LLM meets EquivarianceabstractAccurately predicting 3D structures and dynamics of physical systems is crucial in scientific applications. Existing approaches that rely on geometric Graph Neural Networks (GNNs) effectively enforce $\mathrm{E}(3)$-equivariance, but they often fail in leveraging extensive broader information. While direct application of Large Language Models (LLMs) can incorporate external knowledge, they lack the capability for spatial reasoning with guaranteed equivariance. In this paper, we propose EquiLLM, a novel framework for representing 3D physical systems that seamlessly integrates $\mathrm{E}(3)$-equivariance with LLM capabilities. Specifically, EquiLLM comprises four key components: geometry-aware prompting, an equivariant encoder, an LLM, and an equivariant adapter. Essentially, the LLM guided by the instructive prompt serves as a sophisticated invariant feature processor, while 3D directional information is exclusively handled by the equivariant encoder and adapter modules. Experimental results demonstrate that EquiLLM delivers significant improvements over previous methods across molecular dynamics simulation, human motion simulation, and antibody design, highlighting its promising generalizability. Zongzhao Li, Jiacheng Cen, Bing Su 0001, Tingyang Xu, Yu Rong 0001, Deli Zhao, Wenbing Huang 0001 |
ICML | 3 |
| 2025 | Uniform Attention Maps: Boosting Image Fidelity in Reconstruction and EditingabstractText-guided image generation and editing using diffusion models have achieved remarkable advancements. Among these, tuning-free methods have gained attention for their ability to perform edits without extensive model adjustments, offering simplicity and efficiency. However, existing tuning-free approaches often struggle with balancing fidelity and editing precision. Reconstruction errors in DDIM Inversion are partly attributed to the cross-attention mechanism in U-Net, which introduces misalignments during the inversion and reconstruction process. To address this, we analyze reconstruction from a structural perspective and propose a novel approach that replaces traditional cross-attention with uniform attention maps, significantly enhancing image reconstruction fidelity. Our method effectively minimizes distortions caused by varying text conditions during noise prediction. To complement this improvement, we introduce an adaptive mask-guided editing technique that integrates seamlessly with our reconstruction approach, ensuring consistency and accuracy in editing tasks. Experimental results demonstrate that our approach not only excels in achieving high-fidelity image reconstruction but also performs robustly in real image composition and editing scenarios. This study underscores the potential of uniform attention maps to enhance the fidelity and versatility of diffusion-based image processing methods. Code is available at https://github.com/Mowenyii/Uniform-Attention-Maps. Wenyi Mo, Yalong Bai, Bing Su 0001, Ji-Rong Wen |
WACV | 4 |
| 2025 | Continual Test-Time Adaptation for Single Image Defocus Deblurring via Causal Siamese Networks
Jiangmeng Li, Xiongxin Tang, Bing Su 0001, Fanjiang Xu, Hui Xiong 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Supporting vision-language model few-shot inference with confounder-pruned knowledge prompt
Jiangmeng Li, Wenyi Mo, Chuxiong Sun, Wenwen Qiang, Bing Su 0001, Changwen Zheng |
Neural Networks | 6 |
| 2024 | Dynamic Prompt Optimizing for Text-to-Image GenerationabstractText-to-image generative models, specifically those based on diffusion models like Imagen and Stable Diffusion, have made substantial advancements. Recently, there has been a surge of interest in the delicate refinement of text prompts. Users assign weights or alter the injection time steps of certain words in the text prompts to improve the quality of generated images. However, the success of fine-control prompts depends on the accuracy of the text prompts and the careful selection of weights and time steps, which requires significant manual intervention. To address this, we introduce the Prompt Auto-Editing (PAE) method. Besides refining the original prompts for image generation, we further employ an online reinforcement learning strategy to explore the weights and injection time steps of each word, leading to the dynamic fine-control prompts. The re-wardfunction during training encourages the model to consider aesthetic score, semantic consistency, and user prefer-ences. Experimental results demonstrate that our proposed method effectively improves the original prompts, generating visually more appealing images while maintaining semantic alignment. Code is available at this https URL. Wenyi Mo, Yalong Bai, Bing Su 0001, Ji-Rong Wen, Qing Yang 0033 |
CVPR | 4 |
| 2024 | Domain-Adaptive and Subgroup-Specific Cascaded Temperature Regression for Out-of-Distribution CalibrationabstractAlthough deep neural networks yield high classification accuracy given sufficient training data, their predictions are typically overconfident or under-confident, i.e., the prediction confidences cannot truly reflect the accuracy. Post-hoc calibration tackles this problem by calibrating the prediction confidences without re-training the classification model. However, current approaches assume congruence between test and validation data distributions, limiting their applicability to out-of-distribution scenarios. To this end, we propose a novel meta-set-based cascaded temperature regression method for post-hoc calibration. Our method tailors fine-grained scaling functions to distinct test sets by simulating various domain shifts through data augmentation on the validation set. We partition each meta-set into subgroups based on predicted category and confidence level, capturing diverse uncertainties. A regression network is then trained to derive category-specific and confidence-level-specific scaling, achieving calibration across meta-sets. Extensive experimental results on MNIST, CIFAR-10, and TinyImageNet demonstrate the effectiveness of the proposed method. Jiexin Wang 0003, Bing Su 0001 |
ICASSP | 3 |
| 2024 | Position-Aware Active Learning for Multi-Modal Entity AlignmentabstractMulti-Modal Entity Alignment (MMEA) aims to identify equivalent entities across different knowledge graphs by utilizing auxiliary modalities such as images. While MMEA has made significant progress, prevailing methods still heavily rely on abundant annotated entity pairs. Active learning seeks to alleviate the labeling burden or enhance model efficiency within fixed labeling capacity through careful sample selection. However, active learning for entity alignment in multimodal scenarios remains unexplored. In our view, it is crucial that data selected from different modalities should complement each other without redundancy or overlap; otherwise, the obtained data may prove a waste of labeling budgets. To achieve this goal, we propose a novel acquisition function leveraging Graph Neural Networks’ (GNNs) capability to aggregate information over multiple hops, prioritizing data distant from other modalities’ selections. Moreover, existing approaches employ data augmentation by selecting entity pairs whose inter-entity similarities of other modalities exceed a predefined threshold, but this augmentation strategy inadequately capitalizes on the available similarity information among entities. We can further enhance performance by integrating similarity matrices from different modalities. Consequently, our method achieves considerable improvements over existing active learning methods for entity alignment, as demonstrated by the experiments. Baogui Xu, Yafei Lu, Bing Su 0001, Xiaoran Yan |
ICASSP | 3 |
| 2024 | Unlocking the Power of Spatial and Temporal Information in Medical Multimodal Pre-trainingabstractMedical vision-language pre-training methods mainly leverage the correspondence between paired medical images and radiological reports. Although multi-view spatial images and temporal sequences of image-report pairs are available in off-the-shelf multi-modal medical datasets, most existing methods have not thoroughly tapped into such extensive supervision signals. In this paper, we introduce the Med-ST framework for fine-grained spatial and temporal modeling to exploit information from multiple spatial views of chest radiographs and temporal historical records. For spatial modeling, Med-ST employs the *Mixture of View Expert (MoVE)* architecture to integrate different visual features from both frontal and lateral views. To achieve a more comprehensive alignment, Med-ST not only establishes the global alignment between whole images and texts but also introduces modality-weighted local alignment between text tokens and spatial regions of images. For temporal modeling, we propose a novel cross-modal bidirectional cycle consistency objective by forward mapping classification (FMC) and reverse mapping regression (RMR). By perceiving temporal information from simple to complex, Med-ST can learn temporal semantics. Experimental results across four distinct tasks demonstrate the effectiveness of Med-ST, especially in temporal classification tasks. Our code and model are available at https://github.com/SVT-Yang/MedST. Jinxia Yang, Bing Su 0001, Wayne Xin Zhao, Ji-Rong Wen |
ICML | 2 |
| 2024 | Instance-Specific Semantic Augmentation for Long-Tailed Image ClassificationabstractRecent long-tailed classification methods generally adopt the two-stage pipeline and focus on learning the classifier to tackle the imbalanced data in the second stage via re-sampling or re-weighting, but the classifier is easily prone to overconfidence in head classes. Data augmentation is a natural way to tackle this issue. Existing augmentation methods either perform low-level transformations or apply the same semantic transformation for all instances. However, meaningful augmentations for different instances should be different. In this paper, we propose feature-level augmentation (FLA) and pixel-level augmentation (PLA) learning methods for long-tailed image classification. In the first stage, the feature space is learned from the original imbalanced data. In the second stage, we model the semantic within-class transformation range for each instance by a specific Gaussian distribution and design a semantic transformation generator (STG) to predict the distribution from the instance itself. We train STG by constructing ground-truth distributions for instances of head classes in the feature space. In the third stage, for FLA, we generate instance-specific transformations by STG to obtain feature augmentations of tail classes for fine-tuning the classifier. For PLA, we use STG to guide pixel-level augmentations for fine-tuning the backbone. The proposed augmentation strategy can be combined with different existing long-tail classification methods. Extensive experiments on five imbalanced datasets show the effectiveness of our method. Bing Su 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton SequencesabstractSelf-supervised learning has demonstrated remarkable capability in representation learning for skeleton-based action recognition. Existing methods mainly focus on applying global data augmentation to generate different views of the skeleton sequence for contrastive learning. However, due to the rich action clues in the skeleton sequences, existing methods may only take a global perspective to learn to discriminate different skeletons without thoroughly leveraging the local relationship between different skeleton joints and video frames, which is essential for real-world applications. In this work, we propose a Partial Spatio-Temporal Learning (PSTL) framework to exploit the local relationship from a partial skeleton sequences built by a unique spatio-temporal masking strategy. Specifically, we construct a negative-sample-free triplet steam structure that is composed of an anchor stream without any masking, a spatial masking stream with Central Spatial Masking (CSM), and a temporal masking stream with Motion Attention Temporal Masking (MATM). The feature cross-correlation matrix is measured between the anchor stream and the other two masking streams, respectively. (1) Central Spatial Masking discards selected joints from the feature calculation process, where the joints with a higher degree of centrality have a higher possibility of being selected. (2) Motion Attention Temporal Masking leverages the motion of action and remove frames that move faster with a higher possibility. Our method achieves state-of-the-art performance on NTURGB+D 60, NTURGB+D 120 and PKU-MMD under various downstream tasks. Furthermore, to simulate the real-world scenarios, a practical evaluation is performed where some skeleton joints are lost in downstream tasks.In contrast to previous methods that suffer from large performance drops, our PSTL can still achieve remarkable results under this challenging setting, validating the robustness of our method. Haodong Duan, Anyi Rao, Bing Su 0001, Jiaqi Wang 0003 |
AAAI | 4 |
| 2023 | Transfer Knowledge from Head to Tail: Uncertainty Calibration under Long-tailed DistributionabstractHow to estimate the uncertainty of a given model is a crucial problem. Current calibration techniques treat different classes equally and thus implicitly assume that the distribution of training data is balanced, but ignore the fact that real-world data often follows a long-tailed distribution. In this paper, we explore the problem of calibrating the model trained from a long-tailed distribution. Due to the difference between the imbalanced training distribution and balanced test distribution, existing calibration methods such as temperature scaling can not generalize well to this problem. Specific calibration methods for domain adaptation are also not applicable because they rely on unlabeled target domain instances which are not available. Models trained from a long-tailed distribution tend to be more over-confident to head classes. To this end, we propose a novel knowledge-transferring-based calibration method by estimating the importance weights for samples of tail classes to realize long-tailed calibration. Our method models the distribution of each class as a Gaussian distribution and views the source statistics of head classes as a prior to calibrate the target distributions of tail classes. We adaptively transfer knowledge from head classes to get the target probability density of tail classes. The importance weight is estimated by the ratio of the target probability density over the source probability density. Extensive experiments on CIFAR-10-LT, MNIST-LT, CIFAR-100-LT, and ImageNet-LT datasets demonstrate the effectiveness of our method. Bing Su 0001 |
CVPR | 2 |
| 2023 | Modeling Video as Stochastic Processes for Fine-Grained Video Representation LearningabstractA meaningful video is semantically coherent and changes smoothly. However, most existing fine-grained video representation learning methods learn frame-wise features by aligning frames across videos or exploring relevance between multiple views, neglecting the inherent dynamic process of each video. In this paper, we propose to learn video representations by modeling Video as Stochastic Processes (VSP) via a novel process-based contrastive learning framework, which aims to discriminate between video processes and simultaneously capture the temporal dynamics in the processes. Specifically, we enforce the embeddings of the frame sequence of interest to approximate a goal-oriented stochastic process, i.e., Brownian bridge, in the latent space via a process-based contrastive loss. To construct the Brownian bridge, we adapt specialized sampling strategies under different annotations for both self-supervised and weakly-supervised learning. Experimental results on four datasets show that VSP stands as a state-of-the-art method for various video understanding tasks, including phase progression, phase classification, and frame retrieval. Code is available at https://github.com/hengRUC/VSP. Daqing Liu, Bing Su 0001 |
CVPR | 4 |
| 2023 | Task-Sensitive Discriminative Mutual Attention Network for Few-Shot LearningabstractMany few-shot image classification methods focus on learning a fixed feature space from sufficient samples of seen classes that can be readily transferred to unseen classes. For different tasks, the feature space is either kept the same or only adjusted by generating attentions to query samples. However, the discriminative channels and spatial parts for comparing different query and support images in different tasks are usually different. In this paper, we propose a task-sensitive discriminative mutual attention (TDMA) network to produce task-and-sample-specific features. For each task, TDMA first generates a discriminative task embedding that encodes the inter-class separability and within-class scatter, and then employs the task embedding to enhance discriminative channels respective to this task. Given a specific query and different support images, TDMA further incorporates the task embedding and long-range dependencies to locate the discriminative parts in the spatial dimension. Experimental results on miniImageNet, tieredImageNet and FC100 datasets show the effectiveness of the proposed model. Baogui Xu, Chengjin Xu, Zhiwu Lu 0001, Bing Su 0001 |
ECAI | 4 |
| 2023 | Preformer: Predictive Transformer with Multi-Scale Segment-Wise Correlations for Long-Term Time Series ForecastingabstractIn long-term time series forecasting, most Transformer-based methods adopt the standard point-wise attention mechanism, which not only has high complexity but also cannot explicitly capture the predictive dependencies from contexts since the corresponding key and value are transformed from the same point. This paper proposes a predictive Transformer-based model called Preformer. Preformer introduces a novel efficient Multi-Scale Segment-Correlation mechanism that divides time series into segments and utilizes segment-wise correlation-based attention to replace point-wise attention. A multi-scale structure is developed to aggregate dependencies at different temporal scales and facilitate the selection of segment length. Preformer further designs a predictive paradigm for decoding, where the key and value come from two successive segments rather than the same segment. Experiments demonstrate that Preformer outperforms other Transformer-based models. The codes are available at https://github.com/ddz16/Preformer. Dazhao Du, Bing Su 0001, Zhewei Wei |
ICASSP | 2 |
| 2023 | Toward Auto-Evaluation With Confidence-Based Category Relation-Aware RegressionabstractAuto-evaluation aims to automatically evaluate a trained model on any test dataset without human annotations. Most existing methods utilize global statistics of features extracted by the model as the representation of a dataset. This ignores the influence of the classification head and loses category-wise confusion information of the model. However, ratios of instances assigned to different categories together with their confidence scores reflect how many instances in which categories are difficult for the model to classify, which contain significant indicators for both overall and category-wise performances. In this paper, we propose a Confidence-based Category Relation-aware Regression (C2R2) method. C2R2divides all instances in a meta-set into different categories according to their confidence scores and extracts the global representation from them. For each category, C2R2encodes its local confusion relations to other categories into a local representation. The overall and category-wise performances are regressed from global and local representations, respectively. Extensive experiments show the effectiveness of our method. Jiexin Wang 0003, Bing Su 0001 |
ICASSP | 3 |
| 2023 | Decaying Contrast for Fine-Grained Video Representation LearningabstractPrior contrast-based methods for video representation learning mainly focus on clip discrimination while ignoring the context and relationship of clips from the same video in the temporal dimension. As a consequence, the learned spatiotemporal representations for successive clips are inconsistent and hence perform poorly in fine-grained downstream tasks such as video fragment retrieval or localization. In this paper, we propose a decaying strategy to grasp the gradual evolution along the temporal dimension for fine-grained spatiotemporal representation learning, which consists of two novel contrastive losses. The external decaying contrastive loss is designed to increase the relative similarity of clips from the same video while the internal decaying contrastive loss aims to maintain the discrimination of clips. Experimental results show that the proposed decaying contrastive training approach achieves a significant improvement in the fine-grained video retrieval task on multiple benchmark datasets. Bing Su 0001 |
ICASSP | 2 |
| 2023 | Exploring Temporal Concurrency for Video-Language Representation LearningabstractPaired video and language data is naturally temporal concurrency, which requires the modeling of the temporal dynamics within each modality and the temporal alignment across modalities simultaneously. However, most existing video-language representation learning methods only focus on discrete semantic alignment that encourages aligned semantics to be close in the latent space, or temporal context dependency that captures short-range coherence, failing in building the temporal concurrency. In this paper, we propose to learn video-language representations by modeling video-language pairs as Temporal Concurrent Processes (TCP) via a process-wised distance metric learning framework. Specifically, we employ the soft Dynamic Time Warping (DTW) to measure the distance between two processes across modalities and then optimize the DTW costs. Meanwhile, we further introduce a regularization term that enforces the embeddings of each modality approximating a stochastic process to guarantee the inherent dynamics. Experimental results on three benchmarks demonstrate that TCP stands as a state-of-the-art method for various video-language understanding tasks, including paragraph-to-video retrieval, video moment retrieval, and video question-answering. Code is available at https: //github.com/hengRUC/TCP. Daqing Liu, Zezhong Lv, Bing Su 0001, Dacheng Tao |
ICCV | 4 |
| 2023 | Temporal-enhanced Cross-modality Fusion Network for Video Sentence GroundingabstractVideo sentence grounding aims to localize a segment semantically aligning to the given language query from a video. Most existing works simply interact video and query only once at a single early stage. Not only multi-level dependencies within videos are not explored since interactions act fixedly on a specific level, but also the guiding role of the query is neglected. To tackle these issues, we propose an efficient network namely Temporal-enhanced Cross-modality Fusion Network (TCFN). By directly modulating the temporal receptive field, TCFN captures multilevel temporal enhanced visual features effectively. Furthermore, TCFN explicitly exploits the query to interact with the temporal-enhanced features in multiple stages for better alignment. Benefiting from its succinct architecture, TCFN achieves competitive performance compared to state-of-the-art with a much lower computational cost. Experiments on three benchmark datasets verify the effectiveness of the proposed TCFN. Zezhong Lv, Bing Su 0001 |
ICME | 2 |
| 2023 | Counterfactual Cross-modality Reasoning for Weakly Supervised Video Moment LocalizationabstractVideo moment localization aims to retrieve the target segment of an untrimmed video according to the natural language query. Weakly supervised methods gains attention recently, as the precise temporal location of the target segment is not always available. However, one of the greatest challenges encountered by the weakly supervised method is implied in the mismatch between the video and language induced by the coarse temporal annotations. To refine the vision-language alignment, recent works contrast the cross-modality similarities driven by reconstructing masked queries between positive and negative video proposals. However, the reconstruction may be influenced by the latent spurious correlation between the unmasked and the masked parts, which distorts the restoring process and further degrades the efficacy of contrastive learning since the masked words are not completely reconstructed from the cross-modality knowledge. In this paper, we discover and mitigate this spurious correlation through a novel proposed counterfactual cross-modality reasoning method. Specifically, we first formulate query reconstruction as an aggregated causal effect of cross-modality and query knowledge. Then by introducing counterfactual cross-modality knowledge into this aggregation, the spurious impact of the unmasked part contributing to the reconstruction is explicitly modeled. Finally, by suppressing the unimodal effect of masked query, we can rectify the reconstructions of video proposals to perform reasonable contrastive learning. Extensive experimental evaluations demonstrate the effectiveness of our proposed method. The code is available at https://github.com/sLdZ0306/CCR https://github.com/sLdZ0306/CCR. Zezhong Lv, Bing Su 0001, Ji-Rong Wen |
ACM Multimedia | 2 |
| 2023 | Spatio-Temporal Branching for Motion Prediction using Motion IncrementsabstractHuman motion prediction (HMP) has emerged as a popular research topic due to its diverse applications. Traditional methods rely on hand-crafted features and machine learning techniques, which often struggle to model the complex dynamics of human motion. Recent deep learning-based methods have achieved success by learning spatio-temporal representations of motion, but these models often overlook the reliability of motion data. Additionally, the temporal and spatial dependencies of skeleton nodes are distinct. The temporal relationship captures motion information over time, while the spatial relationship describes body structure and the relationships between different nodes. In this paper, we propose a novel spatio-temporal branching network using incremental information for HMP, which decouples the learning of temporal-domain and spatial-domain features, extracts more motion information, and achieves complementary cross-domain knowledge learning through knowledge distillation. Our approach effectively reduces noise interference and provides more expressive information for characterizing motion by separately extracting temporal and spatial features. We evaluate our approach on standard HMP benchmarks and outperform state-of-the-art methods in terms of prediction accuracy. Code is available at https://github.com/JasonWang959/STPMP. Jiexin Wang 0003, Wenwen Qiang, Ying Ba, Bing Su 0001, Ji-Rong Wen |
ACM Multimedia | 5 |
| 2023 | Cross-Modal Graph Attention Network for Entity AlignmentabstractThe increasing popularity of multi-modal knowledge graphs (MMKGs) has led to a need for efficient entity alignment techniques that can exploit multi-modal information to integrate knowledge from different sources. GNN-based multi-modal entity alignment (MMEA) methods have achieved significant progress in entity alignment(EA) areas. However, these methods only rely on Graph Neural Networks (GNNs) to encode structural information, while ignoring visual and semantic modalities, which may lead to incomplete representation, thus how to integrate the visual and semantic information into GNN-based EA methods remains unexplored. In light of our insight that incorporating the message-passing mechanism of Graph Neural Networks to integrate multi-modal information is essential for fully exploiting the graph representation capability of GNN, we propose a novel Cross-modal Graph attention network for Entity Alignment (XGEA) that enables visual knowledge to interact with other views of the entity, including structural and literal information. We leverage the information from one modality as complementary relation information to compute the attention of another modality in the graph attention layers, enabling the learning of entity embedding by integrating multiple modalities. Moreover, the quantity of labeled data plays a crucial role in model performance, yet obtaining sufficient training data is expensive. To mitigate this issue, we use visual and semantic information to generate pseudo-pairs and propose a soft pseudo-labeling method for entity alignment to assign weights to the augmented training data to balance its quantity and quality. Extensive experiments show that our XGEA achieves superior performance consistently over the state-of-the-art MMEA baselines. Baogui Xu, Chengjin Xu, Bing Su 0001 |
ACM Multimedia | 3 |
| 2023 | Synthesizing Long-Term Human Motions with Diffusion Models via Coherent SamplingabstractText-to-motion generation has gained increasing attention, but most existing methods are limited to generating short-term motions that correspond to a single sentence describing a single action. However, when a text stream describes a sequence of continuous motions, the generated motions corresponding to each sentence may not be coherently linked. Existing long-term motion generation methods face two main issues. Firstly, they cannot directly generate coherent motions and require additional operations such as interpolation to process the generated actions. Secondly, they generate subsequent actions in an autoregressive manner without considering the influence of future actions on previous ones. To address these issues, we propose a novel approach that utilizes a past-conditioned diffusion model with two optional coherent sampling methods: Past Inpainting Sampling and Compositional Transition Sampling. Past Inpainting Sampling completes subsequent motions by treating previous motions as conditions, while Compositional Transition Sampling models the distribution of the transition as the composition of two adjacent motions guided by different text prompts. Our experimental results demonstrate that our proposed method is capable of generating compositional and coherent long-term 3D human motions controlled by a user-instructed long text stream. The code is available at https://github.com/yangzhao1230/PCMDM Zhao Yang 0006, Bing Su 0001, Ji-Rong Wen |
ACM Multimedia | 2 |
| 2023 | Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and MaximizationabstractZero-shot skeleton-based action recognition aims to recognize actions of unseen categories after training on data of seen categories. The key is to build the connection between visual and semantic space from seen to unseen classes. Previous studies have primarily focused on encoding sequences into a singular feature vector, with subsequent mapping the features to an identical anchor point within the embedded space. Their performance is hindered by 1) the ignorance of the global visual/semantic distribution alignment, which results in a limitation to capture the true interdependence between the two spaces. 2) the negligence of temporal information since the frame-wise features with rich action clues are directly pooled into a single feature vector. We propose a new zero-shot skeleton-based action recognition method via mutual information (MI) estimation and maximization. Specifically, 1) we maximize the MI between visual and semantic space for distribution alignment; 2) we leverage the temporal information for estimating the MI by encouraging MI to increase as more frames are observed. Extensive experiments on three large-scale skeleton action datasets confirm the effectiveness of our method. Wenwen Qiang, Anyi Rao, Ning Lin, Bing Su 0001, Jiaqi Wang 0003 |
ACM Multimedia | 5 |
| 2023 | Meta Attention-Generation Network for Cross-Granularity Few-Shot Learning
Wenwen Qiang, Jiangmeng Li, Bing Su 0001, Jianlong Fu, Hui Xiong 0001, Ji-Rong Wen |
Int. J. Comput. Vis. | 3 |
| 2023 | Discriminative Self-Paced Group-Metric Adaptation for Online Visual IdentificationabstractExisting solutions to instance-level visual identification usually aim to learn faithful and discriminative feature extractors from offline training data and directly use them for the unseen online testing data. However, their performance is largely limited due to the severe distribution shifting issue between training and testing samples. Therefore, we propose a novel online group-metric adaptation model to adapt the offline learned identification models for the online data by learning a series of metrics for all sharing-subsets. Each sharing-subset is obtained from the proposed novel frequent sharing-subset mining module and contains a group of testing samples that share strong visual similarity relationships to each other. Furthermore, to handle potentially large-scale testing samples, we introduce self-paced learning (SPL) to gradually include samples into adaptation from easy to difficult which elaborately simulates the learning principle of humans. Unlike existing online visual identification methods, our model simultaneously takes both the sample-specific discriminant and the set-based visual similarity among testing samples into consideration. Our method is generally suitable to any off-the-shelf offline learned visual identification baselines for online performance improvement which can be verified by extensive experiments on several widely-used visual identification benchmarks. Jiahuan Zhou, Bing Su 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Modeling Multiple Views via Implicitly Preserving Global Consistency and Local ComplementarityabstractWhile self-supervised learning techniques are often used to mine hidden knowledge from unlabeled data via modeling multiple views, it is unclear how to perform effective representation learning in a complex and inconsistent context. To this end, we propose a new multi-view self-supervised learning method, namelyconsistency and complementarity network(CoCoNet), to comprehensively learn global inter-view consistent and local cross-view complementarity-preserving representations from multiple views. To capture crucial common knowledge which is implicitly shared among views, CoCoNet employs a global consistency module that aligns the probabilistic distribution of views by utilizing an efficient discrepancy metric based on the generalized sliced Wasserstein distance. To incorporate cross-view complementary information, CoCoNet proposes a heuristic complementarity-aware contrastive learning approach, which extracts a complementarity-factor jointing cross-view discriminative knowledge and uses it as the contrast to guide the learning of view-specific encoders. Theoretically, the superiority of CoCoNet is verified by our information-theoretical-based analyses. Empirically, our thorough experimental results show that CoCoNet outperforms the state-of-the-art self-supervised methods by a significant margin, for instance, CoCoNet beats the best benchmark method by an average margin of 1.1% on ImageNet. Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001, Farid Razzak, Ji-Rong Wen, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Robust Local Preserving and Global Aligning Network for Adversarial Domain AdaptationabstractUnsupervised domain adaptation (UDA) requires source domain samples with clean ground truth labels during training. Accurately labeling a large number of source domain samples is time-consuming and laborious. An alternative is to utilize samples with noisy labels for training. However, training with noisy labels can greatly reduce the performance of UDA. In this paper, we address the problem that learning UDA models only with access to noisy labels and propose a novel method called robust local preserving and global aligning network (RLPGA). RLPGA improves the robustness of the label noise from two aspects. One is learning a classifier by a robust informative-theoretic-based loss function. The other is constructing two adjacency weight matrices and two negative weight matrices by the proposed local preserving module to preserve the local topology structures of input data. We conduct theoretical analysis on the robustness of the proposed RLPGA and prove that the robust informative-theoretic-based loss and the local preserving module are beneficial to reduce the empirical risk of the target domain. A series of empirical studies show the effectiveness of our proposed RLPGA. Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2022 | Optimal Partial Transport Based Sentence Selection for Long-form Document MatchingabstractOne typical approach to long-form document matching is first conducting alignment between cross-document sentence pairs, and then aggregating all of the sentence-level matching signals. However, this approach could be problematic because the alignment between documents is partial — despite two documents as a whole are well-matched, most of the sentences could still be dissimilar. Those dissimilar sentences lead to spurious sentence-level matching signals which may overwhelm the real ones, increasing the difficulties of learning the matching function. Therefore, accurately selecting the key sentences for document matching is becoming a challenging issue. To address the issue, we propose a novel matching approach that equips existing document matching models with an Optimal Partial Transport (OPT) based component, namely OPT-Match, which selects the sentences that play a major role in matching. Enjoying the partial transport properties of OPT, the selected key sentences can not only effectively enhance the matching accuracy, but also be explained as the rationales for the matching results. Extensive experiments on four publicly available datasets demonstrated that existing methods equipped with OPT-Match consistently outperformed the corresponding underlying methods. Evaluations also showed that the key sentences selected by OPT-Match were consistent with human-provided rationales. Weijie Yu 0003, Liang Pang 0001, Jun Xu 0001, Bing Su 0001, Zhenhua Dong, Ji-Rong Wen |
COLING | 4 |
| 2022 | Temporal Alignment Prediction for Supervised Representation Learning and Few-Shot Sequence Classification
Bing Su 0001, Ji-Rong Wen |
ICLR | 1 |
| 2022 | MetAug: Contrastive Learning via Meta Feature AugmentationabstractWhat matters for contrastive learning? We argue that contrastive learning heavily relies on informative features, or “hard” (positive or negative) features. Early works include more informative features by applying complex data augmentations and large batch size or memory bank, and recent works design elaborate sampling approaches to explore informative features. The key challenge toward exploring such features is that the source multi-view data is generated by applying random data augmentations, making it infeasible to always add useful information in the augmented data. Consequently, the informativeness of features learned from such augmented data is limited. In response, we propose to directly augment the features in latent space, thereby learning discriminative representations without a large amount of input data. We perform a meta learning technique to build the augmentation generator that updates its network parameters by considering the performance of the encoder. However, insufficient input data may lead the encoder to learn collapsed features and therefore malfunction the augmentation generator. A new margin-injected regularization is further added in the objective function to avoid the encoder learning a degenerate mapping. To contrast all features in one gradient back-propagation step, we adopt the proposed optimization-driven unified contrastive loss instead of the conventional contrastive loss. Empirically, our method achieves state-of-the-art results on several benchmark datasets. Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
ICML | 4 |
| 2022 | Interventional Contrastive Learning with Meta Semantic RegularizerabstractContrastive learning (CL)-based self-supervised learning models learn visual representations in a pairwise manner. Although the prevailing CL model has achieved great progress, in this paper, we uncover an ever-overlooked phenomenon: When the CL model is trained with full images, the performance tested in full images is better than that in foreground areas; when the CL model is trained with foreground areas, the performance tested in full images is worse than that in foreground areas. This observation reveals that backgrounds in images may interfere with the model learning semantic information and their influence has not been fully eliminated. To tackle this issue, we build a Structural Causal Model (SCM) to model the background as a confounder. We propose a backdoor adjustment-based regularization method, namely Interventional Contrastive Learning with Meta Semantic Regularizer (ICL-MSR), to perform causal intervention towards the proposed SCM. ICL-MSR can be incorporated into any existing CL methods to alleviate background distractions from representation learning. Theoretically, we prove that ICL-MSR achieves a tighter error bound. Empirically, our experiments on multiple benchmark datasets demonstrate that ICL-MSR is able to improve the performances of different state-of-the-art CL methods. Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
ICML | 4 |
| 2022 | Log-Polar Space Convolution LayersabstractConvolutional neural networks use regular quadrilateral convolution kernels to extract features. Since the number of parameters increases quadratically with the size of the convolution kernel, many popular models use small convolution kernels, resulting in small local receptive fields in lower layers. This paper proposes a novel log-polar space convolution (LPSC) layer, where the convolution kernel is elliptical and adaptively divides its local receptive field into different regions according to the relative directions and logarithmic distances. The local receptive field grows exponentially with the number of distance levels. Therefore, the proposed LPSC not only naturally encodes local spatial structures, but also greatly increases the single-layer receptive field while maintaining the number of parameters. We show that LPSC can be implemented with conventional convolution via log-polar space pooling and can be applied in any network architecture to replace conventional convolutions. Experiments on different tasks and datasets demonstrate the effectiveness of the proposed LPSC. Bing Su 0001, Ji-Rong Wen |
NeurIPS | 1 |
| 2022 | MetaMask: Revisiting Dimensional Confounder for Self-Supervised LearningabstractAs a successful approach to self-supervised learning, contrastive learning aims to learn invariant information shared among distortions of the input sample. While contrastive learning has yielded continuous advancements in sampling strategy and architecture design, it still remains two persistent defects: the interference of task-irrelevant information and sample inefficiency, which are related to the recurring existence of trivial constant solutions. From the perspective of dimensional analysis, we find out that the dimensional redundancy and dimensional confounder are the intrinsic issues behind the phenomena, and provide experimental evidence to support our viewpoint. We further propose a simple yet effective approach MetaMask, short for the dimensional Mask learned by Meta-learning, to learn representations against dimensional redundancy and confounder. MetaMask adopts the redundancy-reduction technique to tackle the dimensional redundancy issue and innovatively introduces a dimensional mask to reduce the gradient effects of specific dimensions containing the confounder, which is trained by employing a meta-learning paradigm with the objective of improving the performance of masked representations on a typical self-supervised task. We provide solid theoretical analyses to prove MetaMask can obtain tighter risk bounds for downstream classification compared to typical contrastive methods. Empirically, our method achieves state-of-the-art performance on various benchmarks. Jiangmeng Li, Wenwen Qiang, Wenyi Mo, Changwen Zheng, Bing Su 0001, Hui Xiong 0001 |
NeurIPS | 6 |
| 2022 | SemMAE: Semantic-Guided Masking for Learning Masked AutoencodersabstractRecently, significant progress has been made in masked image modeling to catch up to masked language modeling. However, unlike words in NLP, the lack of semantic decomposition of images still makes masked autoencoding (MAE) different between vision and language. In this paper, we explore a potential visual analogue of words, i.e., semantic parts, and we integrate semantic information into the training process of MAE by proposing a Semantic-Guided Masking strategy. Compared to widely adopted random masking, our masking strategy can gradually guide the network to learn various information, i.e., from intra-part patterns to inter-part relations. In particular, we achieve this in two steps. 1) Semantic part learning: we design a self-supervised part learning method to obtain semantic parts by leveraging and refining the multi-head attention of a ViT-based encoder. 2) Semantic-guided MAE (SemMAE) training: we design a masking strategy that varies from masking a portion of patches in each part to masking a portion of (whole) parts in an image. Extensive experiments on various vision tasks show that SemMAE can learn better image representation by integrating semantic information. In particular, SemMAE achieves 84.5% fine-tuning accuracy on ImageNet-1k, which outperforms the vanilla MAE by 1.4%. In the semantic segmentation and fine-grained recognition tasks, SemMAE also brings significant improvements and yields the state-of-the-art performance. Heliang Zheng, Daqing Liu, Bing Su 0001, Changwen Zheng |
NeurIPS | 5 |
| 2022 | Transductive distribution calibration for few-shot learning
Changwen Zheng, Bing Su 0001 |
Neurocomputing | 3 |
| 2022 | RHMC: Modeling consistent information from deep multiple views via Regularized and Hybrid Multiview Coding
Jiangmeng Li, Wenwen Qiang, Changwen Zheng, Bing Su 0001 |
Knowl. Based Syst. | 4 |
| 2022 | Learning Meta-Distance for Sequences by Learning a Ground Metric via Virtual Sequence RegressionabstractDistance between sequences is structural by nature because it needs to establish the temporal alignments among the temporally correlated vectors in sequences with varying lengths. Generally, distances for sequences heavily depend on the ground metric between the vectors in sequences to infer the alignments and hence can be viewed as meta-distances upon the ground metric. Learning such meta-distance from multi-dimensional sequences is appealing but challenging. We propose to learn the meta-distance through learning a ground metric for the vectors in sequences. The learning samples are sequences of vectors for which how the ground metric between vectors induces the meta-distance is given. The objective is that the meta-distance induced by the learned ground metric produces large values for sequences from different classes and small values for those from the same class. We formulate the ground metric as a parameter of the meta-distance and regress each sequence to an associated pre-generated virtual sequence w.r.t. the meta-distance, where the virtual sequences for sequences of different classes are well-separated. We develop general iterative solutions to learn both the Mahalanobis metric and the deep metric induced by a neural network for any ground-metric-based sequence distance. Experiments on several sequence datasets demonstrate the effectiveness and efficiency of the proposed methods. Bing Su 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Linear and Deep Order-Preserving Wasserstein Discriminant AnalysisabstractSupervised dimensionality reduction for sequence data learns a transformation that maps the observations in sequences onto a low-dimensional subspace by maximizing the separability of sequences in different classes. It is typically more challenging than conventional dimensionality reduction for static data, because measuring the separability of sequences involves non-linear procedures to manipulate the temporal structures. In this paper, we propose a linear method, called order-preserving Wasserstein discriminant analysis (OWDA), and its deep extension, namely DeepOWDA, to learn linear and non-linear discriminative subspace for sequence data, respectively. We construct novel separability measures between sequence classes based on the order-preserving Wasserstein (OPW) distance to capture the essential differences among their temporal structures. Specifically, for each class, we extract the OPW barycenter and construct the intra-class scatter as the dispersion of the training sequences around the barycenter. The inter-class distance is measured as the OPW distance between the corresponding barycenters. We learn the linear and non-linear transformations by maximizing the inter-class distance and minimizing the intra-class scatter. In this way, the proposed OWDA and DeepOWDA are able to concentrate on the distinctive differences among classes by lifting the geometric relations with temporal constraints. Experiments on four 3D action recognition datasets show the effectiveness of OWDA and DeepOWDA. Bing Su 0001, Jiahuan Zhou, Ji-Rong Wen, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Auxiliary task guided mean and covariance alignment network for adversarial domain adaptation
Wenwen Qiang, Jiangmeng Li, Changwen Zheng, Bing Su 0001 |
Knowl. Based Syst. | 4 |
| 2020 | Online Joint Multi-Metric Adaptation From Frequent Sharing-Subset Mining for Person Re-IdentificationabstractPerson Re-IDentification (P-RID), as an instance-level recognition problem, still remains challenging in computer vision community. Many P-RID works aim to learn faithful and discriminative features/metrics from offline training data and directly use them for the unseen online testing data. However, their performance is largely limited due to the severe data shifting issue between training and testing data. Therefore, we propose an online joint multi-metric adaptation model to adapt the offline learned P-RID models for the online data by learning a series of metrics for all the sharing-subsets. Each sharing-subset is obtained from the proposed novel frequent sharing-subset mining module and contains a group of testing samples which share strong visual similarity relationships to each other. Unlike existing online P-RID methods, our model simultaneously takes both the sample-specific discriminant and the set-based visual similarity among testing samples into consideration so that the adapted multiple metrics can refine the discriminant of all the given testing samples jointly via a multi-kernel late fusion framework. Our proposed model is generally suitable to any offline learned P-RID baselines for online boosting, the performance improvement by our model is not only verified by extensive experiments on several widely-used P-RID benchmarks (CUHK03, Market1501, DukeMTMC-reID and MSMT17) and state-of-the-art P-RID baselines but also guaranteed by the provided in-depth theoretical analyses. Jiahuan Zhou, Bing Su 0001, Ying Wu 0001 |
CVPR | 2 |
| 2020 | Learning Low-Dimensional Temporal Representations with Latent AlignmentsabstractLow-dimensional discriminative representations enhance machine learning methods in both performance and complexity. This has motivated supervised dimensionality reduction (DR), which transforms high-dimensional data into a discriminative subspace. Most DR methods require data to be i.i.d. However, in some domains, data naturally appear in sequences, where the observations are temporally correlated. We propose a DR method, namely, latent temporal linear discriminant analysis (LT-LDA), to learn low-dimensional temporal representations. We construct the separability among sequence classes by lifting the holistic temporal structures, which are established based on temporal alignments and may change in different subspaces. We jointly learn the subspace and the associated latent alignments by optimizing an objective that favors easily separable temporal structures. We show that this objective is connected to the inference of alignments and thus allows for an iterative solution. We provide both theoretical insight and empirical evaluations on several real-world sequence datasets to show the applicability of our method. Bing Su 0001, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Order-Preserving Wasserstein Discriminant AnalysisabstractSupervised dimensionality reduction for sequence data projects the observations in sequences onto a low-dimensional subspace to better separate different sequence classes. It is typically more challenging than conventional dimensionality reduction for static data, because measuring the separability of sequences involves non-linear procedures to manipulate the temporal structures. This paper presents a linear method, namely Order-preserving Wasserstein Discriminant Analysis (OWDA), which learns the projection by maximizing the inter-class distance and minimizing the intra-class scatter. For each class, OWDA extracts the order-preserving Wasserstein barycenter and constructs the intra-class scatter as the dispersion of the training sequences around the barycenter. The inter-class distance is measured as the order-preserving Wasserstein distance between the corresponding barycenters. OWDA is able to concentrate on the distinctive differences among classes by lifting the geometric relations with temporal constraints. Experiments show that OWDA achieves competitive results on three 3D action recognition datasets. Bing Su 0001, Jiahuan Zhou, Ying Wu 0001 |
ICCV | 1 |
| 2019 | Learning Distance for Sequences by Learning a Ground MetricabstractLearning distances that operate directly on multi-dimensional sequences is challenging because such distances are structural by nature and the vectors in sequences are not independent. Generally, distances for sequences heavily depend on the ground metric between the vectors in sequences. We propose to learn the distance for sequences through learning a ground Mahalanobis metric for the vectors in sequences. The learning samples are sequences of vectors for which how the ground metric between vectors induces the overall distance is given, and the objective is that the distance induced by the learned ground metric produces large values for sequences from different classes and small values for those from the same class. We formulate the metric as a parameter of the distance, bring closer each sequence to an associated virtual sequence w.r.t. the distance to reduce the number of constraints, and develop a general iterative solution for any ground-metric-based sequence distance. Experiments on several sequence datasets demonstrate the effectiveness and efficiency of our method. Bing Su 0001, Ying Wu 0001 |
ICML | 1 |
| 2019 | Order-Preserving Optimal Transport for Distances between SequencesabstractWe present new distance measures between sequences that can tackle local temporal distortion and periodic sequences with arbitrary starting points. Through viewing the instances of each sequence as empirical samples of an unknown distribution, we cast the calculations of distances between sequences as optimal transport problems. To preserve the inherent temporal relationships of the instances in sequences, we propose two methods through incorporating the temporal information into the spatial ground metric and concentrating the transport with two novel temporal regularization terms, respectively. The inverse difference moment regularization enforces local homogeneous structures in the transport, and the KL-divergence with a prior distribution regularization prevents transport between instances with far temporal positions. We show that the resulting problems can be efficiently solved by the matrix scaling algorithm. Extensive experiments on eight datasets with different classifiers and performance measures show the effectiveness and generality of the proposed distances. Bing Su 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Easy Identification From Better Constraints: Multi-Shot Person Re-Identification From Reference ConstraintsabstractMulti-shot person re-identification (MsP-RID) utilizes multiple images from the same person to facilitate identification. Considering the fact that motion information may not be discriminative nor reliable enough for MsP-RID, this paper is focused on handling the large variations in the visual appearances through learning discriminative visual metrics for identification. Existing metric learning-based methods usually exploit pair-wise or triple-wise similarity constraints, that generally demands intensive optimization in metric learning, or leads to degraded performances by using sub-optimal solutions. In addition, as the training data are significantly imbalanced, the learning can be largely dominated by the negative pairs and thus produces unstable and non-discriminative results. In this paper, we propose a novel type of similarity constraint. It assigns the sample points to a set of reference points to produce a linear number of reference constraints. Several optimal transport-based schemes for reference constraint generation are proposed and studied. Based on those constraints, by utilizing a typical regressive metric learning model, the closed-form solution of the learned metric can be easily obtained. Extensive experiments and comparative studies on several public MsP-RID benchmarks have validated the effectiveness of our method and its significant superiority over the state-of-the-art MsP-RID methods in terms of both identification accuracy and running speed. Jiahuan Zhou, Bing Su 0001, Ying Wu 0001 |
CVPR | 2 |
| 2018 | Feature Fusion Network for Scene Text DetectionabstractDetecting scene text in natural images is a challenging task due to various text scales, uneven lighting, burring and perspective distortion. Retrospectively, text detection features are extracted from a single scale which is not sufficient enough since text regions are various in enormous sizes. To address the issue, we design a Feature Fusion Network to concatenate lower and higher level features to detect text regions of multiple scales. Besides, as it may incur learning bias due to the difference between scene text and general object when using models pre-trained on general object datasets, we extend DenseNet as a base network during the feature extraction stage to help to start training without pre-training. Our model enables an end-to-end scene text detector which can detect text regions of various scales without a pre-trained model. It achieves state-of-the-art results on ICDAR 2013 and COCO-Text benchmarks. Chenqin Cai, Bing Su 0001 |
ICIP | 3 |
| 2018 | Spatiotemporal Pyramid Pooling in 3D Convolutional Neural Networks for Action RecognitionabstractDeep 3-dimensional convolutional networks (3D ConvNets) trained on large scale video datasets have achieved promising results on action recognition. This paper improves their performance by taking into account the spatiotemporal pyramid pooling. Specifically, we propose the spatiotemporal pyramid pooling layer to tackle the temporal variations of video sequences. Based on this layer, we develop a new network architecture, called STPP-net, by incorporating it with 3D ConvNets. The proposed network is robust to spatial and temporal variation of human actions and can generate a fixed-dimensional representation regardless of video size/scale. We show that our new network architecture outperforms the original 3D ConvNets by a large margin on three large-scale video classification/action recognition benchmarks including HMDB51, UCF101, and Kinetics. Bing Su 0001 |
ICIP | 3 |
| 2018 | Learning Low-Dimensional Temporal RepresentationsabstractLow-dimensional discriminative representations enhance machine learning methods in both performance and complexity, motivating supervised dimensionality reduction (DR) that transforms high-dimensional data to a discriminative subspace. Most DR methods require data to be i.i.d., however, in some domains, data naturally come in sequences, where the observations are temporally correlated. We propose a DR method called LT-LDA to learn low-dimensional temporal representations. We construct the separability among sequence classes by lifting the holistic temporal structures, which are established based on temporal alignments and may change in different subspaces. We jointly learn the subspace and the associated alignments by optimizing an objective which favors easily-separable temporal structures, and show that this objective is connected to the inference of alignments, thus allows an iterative solution. We provide both theoretical insight and empirical evaluation on real-world sequence datasets to show the interest of our method. Bing Su 0001, Ying Wu 0001 |
ICML | 1 |
| 2018 | Discriminative Dimensionality Reduction for Multi-Dimensional SequencesabstractSince the observables at particular time instants in a temporal sequence exhibit dependencies, they are not independent samples. Thus, it is not plausible to apply i.i.d. assumption-based dimensionality reduction methods to sequence data. This paper presents a novel supervised dimensionality reduction approach for sequence data, called Linear Sequence Discriminant Analysis (LSDA). It learns a linear discriminative projection of the feature vectors in sequences to a lower-dimensional subspace by maximizing the separability of the sequence classes such that the entire sequences are holistically discriminated. The sequence class separability is constructed based on the sequence statistics, and the use of different statistics produces different LSDA methods. This paper presents and compares two novel LSDA methods, namely M-LSDA and D-LSDA. M-LSDA extracts model-based statistics by exploiting the dynamical structure of the sequence classes, and D-LSDA extracts the distance-based statistics by computing the pairwise similarity of samples from the same sequence class. Extensive experiments on several different tasks have demonstrated the effectiveness and the general applicability of the proposed methods. Bing Su 0001, Xiaoqing Ding, Hao Wang 0005, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Heteroscedastic Max-Min Distance Analysis for Dimensionality ReductionabstractMax-min distance analysis (MMDA) performs dimensionality reduction by maximizing the minimum pairwise distance between classes in the latent subspace under the homoscedastic assumption, which can address the class separation problem caused by the Fisher criterion, but is incapable of tackling heteroscedastic data properly. In this paper, we propose two heteroscedastic MMDA (HMMDA) methods to employ the differences of class covariances. Whitened HMMDA (WHMMDA) extends MMDA by utilizing the Chernoff distance as the separability measure between classes in the whitened space. Orthogonal HMMDA (OHMMDA) incorporates the maximization of the minimal pairwise Chernoff distance and the minimization of class compactness into a trace quotient formulation with an orthogonal constraint of the transformation, which can be solved by bisection search. Two variants of OHMMDA further encode the margin information by using only neighboring samples to construct the intra-class and inter-class scatters. Experiments on several UCI datasets and two face databases demonstrate the effectiveness of the HMMDA methods. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Order-Preserving Wasserstein Distance for Sequence MatchingabstractWe present a new distance measure between sequences that can tackle local temporal distortion and periodic sequences with arbitrary starting points. Through viewing the instances of sequences as empirical samples of an unknown distribution, we cast the calculation of the distance between sequences as the optimal transport problem. To preserve the inherent temporal relationships of the instances in sequences, we smooth the optimal transport problem with two novel temporal regularization terms. The inverse difference moment regularization enforces transport with local homogeneous structures, and the KL-divergence with a prior distribution regularization prevents transport between instances with far temporal positions. We show that this problem can be efficiently optimized through the matrix scaling algorithm. Extensive experiments on different datasets with different classifiers show that the proposed distance outperforms the traditional DTW variants and the smoothed optimal transport distance without temporal regularization. Bing Su 0001, Gang Hua 0001 |
CVPR | 1 |
| 2017 | Discriminative Transformation for Multi-Dimensional Temporal SequencesabstractFeature space transformation techniques have been widely studied for dimensionality reduction in vector-based feature space. However, these techniques are inapplicable to sequence data because the features in the same sequence are not independent. In this paper, we propose a method called max-min inter-sequence distance analysis (MMSDA) to transform features in sequences into a low-dimensional subspace such that different sequence classes are holistically separated. To utilize the temporal dependencies, MMSDA first aligns features in sequences from the same class to an adapted number of temporal states, and then, constructs the sequence class separability based on the statistics of these ordered states. To learn the transformation, MMSDA formulates the objective of maximizing the minimal pairwise separability in the latent subspace as a semi-definite programming problem and provides a new tractable and effective solution with theoretical proofs by constraints unfolding and pruning, convex relaxation, and within-class scatter compression. Extensive experiments on different tasks have demonstrated the effectiveness of MMSDA. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Hao Wang 0005, Ying Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Unsupervised Hierarchical Dynamic Parsing and Encoding for Action RecognitionabstractGenerally, the evolution of an action is not uniform across the video, but exhibits quite complex rhythms and non-stationary dynamics. To model such non-uniform temporal dynamics, in this paper, we describe a novel hierarchical dynamic parsing and encoding method to capture both the locally smooth dynamics and globally drastic dynamic changes. It parses the dynamics of an action into different layers and encodes such multi-layer temporal information into a joint representation for action recognition. At the first layer, the action sequence is parsed in an unsupervised manner into several smooth-changing stages corresponding to different key poses or temporal structures by temporal clustering. The dynamics within each stage are encoded by mean-pooling or rank-pooling. At the second layer, the temporal information of the ordered dynamics extracted from the previous layer is encoded again by rank-pooling to form the overall representation. Extensive experiments on a gesture action data set (Chalearn Gesture) and three generic action data sets (Olympic Sports, Hollywood2, and UCF101) have demonstrated the effectiveness of the proposed method. Bing Su 0001, Jiahuan Zhou, Xiaoqing Ding, Ying Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Hierarchical Dynamic Parsing and Encoding for Action Recognition
Bing Su 0001, Jiahuan Zhou, Xiaoqing Ding, Hao Wang 0005, Ying Wu 0001 |
ECCV (4) | 1 |
| 2015 | Heteroscedastic max-min distance analysisabstractMany discriminant analysis methods such as LDA and HLDA actually maximize the average pairwise distances between classes, which often causes the class separation problem. Max-min distance analysis (MMDA) addresses this problem by maximizing the minimum pairwise distance in the latent subspace, but it is developed under the homoscedastic assumption. This paper proposes Heteroscedastic MMDA (HMMDA) methods that explore the discriminative information in the difference of intra-class scatters for dimensionality reduction. WHMMDA maximizes the minimal pairwise Chenoff distance in the whitened space. OHMMDA incorporates this objective and the minimization of class compactness into a trace quotient formulation and imposes an orthogonal constraint to the final transformation, which can be solved by a bisection search algorithm. Two variants of OHMMDA are further proposed to encode the margin information. Experiments on several UCI Machine Learning datasets and the Yale Face database demonstrate the effectiveness of the proposed HMMDA methods. Bing Su 0001, Xiaoqing Ding, Changsong Liu, Ying Wu 0001 |
CVPR | 1 |
| 2013 | Linear Sequence Discriminant Analysis: A Model-Based Dimensionality Reduction Method for Vector SequencesabstractDimensionality reduction for vectors in sequences is challenging since labels are attached to sequences as a whole. This paper presents a model-based dimensionality reduction method for vector sequences, namely linear sequence discriminant analysis (LSDA), which attempts to find a subspace in which sequences of the same class are projected together while those of different classes are projected as far as possible. For each sequence class, an HMM is built from states of which statistics are extracted. Means of these states are linked in order to form a mean sequence, and the variance of the sequence class is defined as the sum of all variances of component states. LSDA then learns a transformation by maximizing the separability between sequence classes and at the same time minimizing the within-sequence class scatter. DTW distance between mean sequences is used to measure the separability between sequence classes. We show that the optimization problem can be approximately transformed into an eigen decomposition problem. LDA can be seen as a special case of LSDA by considering non-sequential vectors as sequences of length one. The effectiveness of the proposed LSDA is demonstrated on two individual sequence datasets from UCI machine learning repository as well as two concatenate sequence datasets: APTI Arabic printed text database and IFN/ENIT Arabic handwriting database. Bing Su 0001, Xiaoqing Ding |
ICCV | 1 |
| 2013 | Cross-Language Sensitive Words Distribution Map: A Novel Recognition-Based Document Understanding Method for Uighur and TibetanabstractCross-language document recognition and understanding have urgent realistic needs and extensive application prospects. In this paper, we propose a novel recognition-based Uighur and Tibetan document understanding method, termed "cross-language sensitive words distribution map" (CSWDM). In our unified recognition-understanding framework, digital Uighur/Tibetan document images are first recognized using OCR technology, and then CSWDM labels the Chinese information of sensitive words on the recognized transcriptions or directly on the original digital images, thus the space location and occurrence frequency of these sensitive words can be intuitively represented. With such information, readers can roughly understand the theme and meaning of the cross-language documents. Bing Su 0001, Xiaoqing Ding, Liangrui Peng, Changsong Liu |
ICDAR | 1 |
| 2013 | A Novel Baseline-independent Feature Set for Arabic Handwriting RecognitionabstractHMM-based analytical methods have been widely used for Arabic handwriting recognition. A key factor influencing the performance of HMM-based systems is the features extracted from a sliding window. In this paper, we propose a novel baseline-independent feature set extracted from a wider sliding window to directly capture the contextual information. This feature set is a combination of center of mass based log-space distribution features and inverse percentile features. Center of mass based log-space distribution features use a normalized histogram to describe the distribution of foreground pixels in different direction and distances with respect to the center of mass. Experiments on the IFN/ENIT database demonstrate the effectiveness of the proposed feature set. Further, this feature set can be combined with some popular baseline-independent features to form a large feature set, which achieves comparable results with several state-of-the-art systems using a simple HMM-based architecture. Bing Su 0001, Xiaoqing Ding, Liangrui Peng, Changsong Liu |
ICDAR | 1 |