EDBT 2026 Demo / reviewers in the wild / expert
Shengeng Tang
dblp:246/3080
· DBLP profile ↗
29ranked-venue papers
5as first author
28since 2021 · last 2026
0000-0001-6313-2543ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 15 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating Controllable Generation via Hybrid-grained CacheabstractControllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid-Grained Cache (HGC) approach that reduces computational overhead by adopting cache strategies with different granularities at different computational stages. Specifically, (1) we use a coarse-grained cache (block-level) based on feature reuse to dynamically bypass redundant computations in encoder-decoder blocks between each step of model reasoning. (2) We design a fine-grained cache (prompt-level) that acts within a module, where the fine-grained cache reuses cross-attention maps within consecutive reasoning steps and extends them to the corresponding module computations of adjacent steps. These caches of different granularities can be seamlessly integrated into each computational link of the controllable generation process. We verify the effectiveness of HGC on four benchmark datasets, especially its advantages in balancing generation efficiency and visual quality. For example, on the COCO-Stuff segmentation benchmark, our HGC significantly reduces the computational cost (MACs) by 63% (from 18.22T → 6.70T↓), while keeping the loss of semantic fidelity (quantized performance degradation) within 1.5%. Huixia Ben, Shuo Wang 0008, Jinda Lu, Junxiang Qiu, Shengeng Tang, Yanbin Hao |
AAAI | 6 |
| 2026 | LinProVSR: Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech RecognitionabstractVisual Speech Recognition (VSR), commonly known as lipreading, enables the recognition of spoken text by analyzing lip visual features. Due to the subtlety of lip movements, its recognition is much harder than other motion recognition tasks. Existing VSR models face the challenge of viseme ambiguity when processing phonemes with similar pronunciations—multiple phonemes share similar viseme features, leading to a notable drop in lipreading accuracy. To address this issue, this study proposes a Linguistics-Knowledge Guided Progressive Disambiguation Network for Visual Speech Recognition(LinProVSR) framework. First, an ambiguous sample set is constructed based on linguistic knowledge to provide supervisory signals for the model's training. Then, a Progressive Contrastive Disambiguation Network (PCDN) is designed, which progressively enhances the model's ability to capture the subtle viseme differences corresponding to similar phonemes through viseme-phoneme contrastive disambiguation in the encoding stage and text contrastive disambiguation in the decoding stage. Furthermore, we pioneer the Ambiguous Word Error Rate (AWER) metric specifically for evaluating recognition of phonetically ambiguous text, and verify the effectiveness of the proposed method on multiple public datasets, achieving a significant breakthrough especially in distinguishing visually similar phonemes. Feng Xue 0002, Baochao Zhu, Wei Jia 0001, Shujie Li 0002, Yu Li 0053, Shengeng Tang, Dan Guo 0001 |
AAAI | 7 |
| 2026 | Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningabstractOpen-world 3D scene understanding is fundamentally challenging for vision and robotics, due to the constraints of closed-vocabulary supervision and static annotations. To address this, we propose a unified framework for Open-World 3D Scene Graph Generation with Retrieval-Augmented Reasoning, which enables generalizable and interactive 3D scene understanding. Our method integrates vision-language models with retrieval-based reasoning to support multimodal exploration and language-guided interaction. The framework comprises two key components: (1) a dynamic scene graph generation module that detects objects and infers semantic relationships without fixed label sets, and (2) a retrieval-augmented reasoning pipeline that encodes scene graphs into a vector database to support text/image-conditioned queries. We evaluate our method on 3DSSG and Replica benchmarks across four tasks—scene question answering, visual grounding, instance retrieval, and task planning—demonstrating robust generalization and superior performance in diverse environments. Our results highlight the effectiveness of combining open-vocabulary perception with retrieval-based reasoning for scalable 3D scene understanding. Fei Yu 0012, Shengeng Tang, Lechao Cheng |
AAAI | 3 |
| 2026 | Wi-CBR: Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior RecognitionabstractThe challenge in WiFi-based cross-domain Behavior Recognition lies in the significant interference of domain-specific signals on gesture variation. However, previous methods alleviate this interference by mapping the phase from multiple domains into a common feature space. If the Doppler Frequency Shift (DFS) signal is used to dynamically supplement the phase features to achieve better generalization, it enables the model to not only explore a wider feature space but also to avoid potential degradation of gesture semantic information. Specifically, we propose a novel Salient-aware Adaptive WiFi Sensing for Cross-domain Behavior Recognition (Wi-CBR), which constructs a dual-branch self-attention module that captures temporal features from phase information reflecting dynamic path length variations while extracting kinematic features from DFS correlated with motion velocity. Moreover, we design a Saliency Guidance Module that employs group attention mechanisms to mine critical activity features and utilizes gating mechanisms to optimize information entropy, facilitating feature fusion and enabling effective interaction between salient and non-salient behavioral characteristics. Extensive experiments on two large-scale public datasets (Widar3.0 and XRF55) demonstrate the superior performance of our method in both in-domain and cross-domain scenarios. Ruobei Zhang, Shengeng Tang, Xiang Zhang 0011, Jiabao Guo |
AAAI | 2 |
| 2026 | CFLip: Generalizing Lipreading to Unseen Speakers by Learning Common FeaturesabstractLipreading refers to translating the lip movements observed in a video of a speaker into corresponding textual outputs, providing a visual alternative to auditory communication for individuals who are deaf or hard of hearing. Existing lipreading methods typically independently learn the lip movements of each speaker. This results in the model being highly sensitive to the individual visual features (lip color/shape) of the speakers in the training set, hindering the generalization of lipreading models. Despite the obvious visual variations in the lips of different speakers, we claim that there are still inherent common features when they pronounce the same phoneme. We attempt to learn the common pronunciation features across different speakers, so as to achieve better generalization of lipreading model to unseen speakers. In this article, we propose a sentence-level lipreading framework based on Learning Common Features (CFLip), designed to extract common pronunciation features fromvideo pairs. Specifically, we first employ data augmentation strategy to generate pseudo videos that share labels but with different speakers by replacing frame segments in real videos. With thesevideo pairs, we designed a dual-stream network to learn commonality feature by minimized the distance between the features of different speakers pronouncing the same words via Generalization Loss. Extensive experiments on benchmark datasets demonstrate that the proposed CFLip can effectively generalize to unseen speakers. Yu Li 0053, Feng Xue 0002, Dan Guo 0001, Shengeng Tang, Shujie Li 0002, Richang Hong |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2025 | Patch-level Sounding Object Tracking for Audio-Visual Question AnsweringabstractAnswering questions related to audio-visual scenes, i.e., the AVQA task, is becoming increasingly popular. A critical challenge is accurately identifying and tracking sounding objects related to the question along the timeline. In this paper, we present a new Patch-level Sounding Object Tracking (PSOT) method. It begins with a Motion-driven Key Patch Tracking (M-KPT) module, which relies on visual motion information to identify salient visual patches with significant movements that are more likely to relate to sounding objects and questions. We measure the patch-wise motion intensity map between neighboring video frames and utilize it to construct and guide a motion-driven graph network. Meanwhile, we design a Sound-driven KPT (S-KPT) module to explicitly track sounding patches. This module also involves a graph network, with the adjacency matrix regularized by the audio-visual correspondence map. The M-KPT and S-KPT modules are performed in parallel for each temporal segment, allowing balanced tracking of salient and sounding objects. Based on the tracked patches, we further propose a Question-driven KPT (Q-KPT) module to retain patches highly relevant to the question, ensuring the model focuses on the most informative clues. The audio-visual-question features are updated during the processing of these modules, which are then aggregated for final answer prediction. Extensive experiments on standard datasets demonstrate the effectiveness of our method, achieving competitive performance even compared to recent large-scale pretraining-based approaches. Zhangbin Li, Jinxing Zhou, Jing Zhang 0052, Shengeng Tang, Kun Li 0008, Dan Guo 0001 |
AAAI | 4 |
| 2025 | PhysDiff: Physiology-based Dynamicity Disentangled Diffusion Model for Remote Physiological MeasurementabstractRecent works on remote PhotoPlethysmoGraphy (rPPG) estimation typically use techniques like CNNs and Transformers to encode implicit features from facial videos for prediction. These methods learn to directly map facial videos to the static values of rPPG signals, overlooking the inherent dynamic characteristics of rPPG sequence. Moreover, the rPPG signal is extremely weak and highly susceptible to interference from various sources of noise, including illumination conditions, head movements, and variations in skin tone. To address these limitations, we propose a Physiology-based dynamicity disentangled diffusion (PhysDiff) model particularly designed for robust rPPG estimation. PhysDiff leverages the diffusion model to learn the distribution of quasi-periodic rPPG signal and uses a dynamicity disentanglement strategy to capture two dynamic characteristics in temporal rPPG signal, i.e., trend and amplitude. This disentanglement is motivated by the underlying dynamic physiological processes of vasodilation and vasoconstriction, ensuring a more precise representation of the rPPG signal. The disentangled components are then used as pivotal conditions in the proposed spatial-temporal hybrid denoiser for rPPG reconstruction. Besides, we introduce a periodicity-based multi-hypothesis selection strategy in model inference, which compares the natural periodicity of multiple generated rPPG hypotheses and selects the most favorable one as the final prediction. Extensive experiments on four datasets demonstrate that our PhysDiff significantly outperforms prior methods on both intra-dataset and cross-dataset testing. Gaoji Su, Dan Guo 0001, Jinxing Zhou, Bin Hu 0001, Shengeng Tang, Meng Wang 0001 |
AAAI | 7 |
| 2025 | Sign-IDD: Iconicity Disentangled Diffusion for Sign Language ProductionabstractSign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit them, which overlooks the relative positional relationships among joints. To this end, we provide a new perspective, constraining joint associations and gesture details by modeling the limb bones to improve the accuracy and naturalness of the generated poses. In this work, we propose a pioneering iconicity disentangled diffusion framework, termed Sign-IDD, specifically designed for SLP. Sign-IDD incorporates a novel Iconicity Disentanglement (ID) module to bridge the gap between relative positions among joints. The ID module disentangles the conventional 3D joint representation into a 4D bone representation, comprising the 3D spatial direction vector and 1D spatial distance vector between adjacent joints. Additionally, an Attribute Controllable Diffusion (ACD) module is introduced to further constrain joint associations, in which the attribute separation layer aims to separate the bone direction and length attributes, and the attribute control layer is designed to guide the pose generation by leveraging the above attributes. The ACD module utilizes the gloss embeddings as semantic conditions and finally generates sign poses from noise embeddings. Extensive experiments on PHOENIX14T and USTC-CSL datasets validate the effectiveness of our method. Shengeng Tang, Dan Guo 0001, Yanyan Wei, Feng Li 0037, Richang Hong |
AAAI | 1 |
| 2025 | Dense Audio-Visual Event Localization Under Cross-Modal Consistency and Multi-Temporal Granularity CollaborationabstractIn the field of audio-visual learning, most research tasks focus exclusively on short videos. This paper focuses on the more practical Dense Audio-Visual Event Localization (DAVEL) task, advancing audio-visual scene understanding for longer, untrimmed videos. This task seeks to identify and temporally pinpoint all events simultaneously occurring in both audio and visual streams. Typically, each video encompasses dense events of multiple classes, which may overlap on the timeline, each exhibiting varied durations. Given these challenges, effectively exploiting the audio-visual relations and the temporal features encoded at various granularities becomes crucial. To address these challenges, we introduce a novel CCNet, comprising two core modules: the Cross-Modal Consistency Collaboration (CMCC) and the Multi-Temporal Granularity Collaboration (MTGC). Specifically, the CMCC module contains two branches: a cross-modal interaction branch and a temporal consistency-gated branch. The former branch facilitates the aggregation of consistent event semantics across modalities through the encoding of audio-visual relations, while the latter branch guides one modality's focus to pivotal event-relevant temporal areas as discerned in the other modality. The MTGC module includes a coarse-to-fine collaboration block and a fine-to-coarse collaboration block, providing bidirectional support among coarse- and fine-grained temporal features. Extensive experiments on the UnAV-100 dataset validate our module design, resulting in a new state-of-the-art performance in dense audio-visual event localization. Jinxing Zhou, Shengeng Tang, Xiaojun Chang, Dan Guo 0001 |
AAAI | 4 |
| 2025 | Discrete to Continuous: Generating Smooth Transition Poses from Sign Language ObservationsabstractGenerating continuous sign language videos from discrete segments is challenging due to the need for smooth transitions that preserve natural flow and meaning. Traditional approaches that simply concatenate isolated signs often result in abrupt transitions, disrupting video coherence. To address this, we propose a novel framework, Sign-D2C, that employs a conditional diffusion model to synthesize contextually smooth transition frames, enabling the seamless construction of continuous sign language sequences. Our approach transforms the unsupervised problem of transition frame generation into a supervised training task by simulating the absence of transition frames through random masking of segments in long-duration sign videos. The model learns to predict these masked frames by denoising Gaussian noise, conditioned on the surrounding sign observations, allowing it to handle complex, unstructured transitions. During inference, we apply a linearly interpolating padding strategy that initializes missing frames through interpolation between boundary frames, providing a stable foundation for iterative refinement by the diffusion model. Extensive experiments on the PHOENIX14T, USTCCSL100, and USTC-SLR500 datasets demonstrate the effectiveness of our method in producing continuous, natural sign language videos. Shengeng Tang, Lechao Cheng, Jingjing Wu 0001, Dan Guo 0001, Richang Hong |
CVPR | 1 |
| 2025 | EvEnhancer: Empowering Effectiveness, Efficiency and Generalizability for Continuous Space-Time Video Super-Resolution with EventsabstractContinuous space-time video super-resolution (C-STVSR) endeavors to upscale videos simultaneously at arbitrary spatial and temporal scales, which has recently garnered increasing interest. However, prevailing methods struggle to yield satisfactory videos at out-of-distribution spatial and temporal scales. On the other hand, event streams characterized by high temporal resolution and high dynamic range, exhibit compelling promise in vision tasks. This paper presents EvEnhancer, an innovative approach that marries the unique advantages of event streams to elevate effectiveness, efficiency, and generalizability for C-STVSR. Our approach hinges on two pivotal components: 1) Event-adapted synthesis capitalizes on the spatiotemporal correlations between frames and events to discern and learn long-term motion trajectories, enabling the adaptive interpolation and fusion of informative spatiotemporal features; 2) Local implicit video transformer integrates local implicit video neural function with cross-scale spatiotemporal attention to learn continuous video representations utilized to generate plausible videos at arbitrary resolutions and frame rates. Experiments show that EvEnhancer achieves superiority on synthetic and real-world datasets and preferable generalizability on out-of-distribution scales against state-of-the-art methods. Code is available at https://github.com/W-Shuoyan/EvEnhancer. Shuoyan Wei, Feng Li 0037, Shengeng Tang, Yao Zhao 0001, Huihui Bai 0001 |
CVPR | 3 |
| 2025 | Linguistics-Vision Monotonic Consistent Network for Sign Language ProductionabstractSign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP suffers huge challenges in linguistics-vision consistency. In this work, we propose a Transformer-based Linguistics-Vision Monotonic Consistent Network (LVMCN) for SLP, which constrains fine-grained cross-modal monotonic alignment and coarse-grained multimodal semantic consistency in language-visual cues through Cross-modal Semantic Aligner (CSA) and Multimodal Semantic Comparator (MSC). In the CSA, we constrain the implicit alignment between corresponding gloss and pose sequences by computing the cosine similarity association matrix between cross-modal feature sequences (i.e., the order consistency of fine-grained sign glosses and actions). As for MSC, we construct multimodal triplets based on paired and unpaired samples in batch data. By pulling closer the corresponding text-visual pairs and pushing apart the non-corresponding text-visual pairs, we constrain the semantic co-occurrence degree between corresponding gloss and pose sequences (i.e., the semantic consistency of coarse-grained textual sentences and sign videos). Extensive experiments on the popular PHOENIX14T benchmark show that the LVMCN outperforms the state-of-the-art. Shengeng Tang, Peipei Song, Shuo Wang 0008, Dan Guo 0001, Richang Hong |
ICASSP | 2 |
| 2025 | Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion RecognitionabstractSpeech Emotion Recognition (SER) plays a critical role in enhancing user experience within human-computer interaction. However, existing methods are overwhelmed by temporal domain analysis, overlooking the valuable envelope structures of the frequency domain that are equally important for robust emotion recognition. To overcome this limitation, we propose TF-Mamba, a novel multi-domain framework that captures emotional expressions in both temporal and frequency dimensions. Concretely, we propose a temporal-frequency mamba block to extract temporal- and frequency-aware emotional features, achieving an optimal balance between computational efficiency and model expressiveness. Besides, we design a Complex Metric-Distance Triplet (CMDT) loss to enable the model to capture representative emotional clues for SER. Extensive experiments on the IEMOCAP and MELD datasets show that TF-Mamba surpasses existing methods in terms of model size and latency, providing a more practical solution for future SER applications. Fei Wang 0067, Kun Li 0008, Yanyan Wei, Shengeng Tang, Shu Zhao 0005, Xiao Sun 0003 |
ICASSP | 5 |
| 2025 | Efficient Asymmetric Shared Low-Rank Adaptation Based on Selective Scanning Vision Mamba for Medical Imaging AnalysisabstractThe Vision Mamba has garnered significant attention for its efficiency and accuracy. Nonetheless, the potential of parameter-efficient tuning in Vision Mamba for medical image analysis remains under-explored. In this work, we investigate the position-sensitive selective scanning mechanism in Vision Mamba and propose a novel low-rank adaptation strategy tailored for planar traversal dependencies. Specifically, we first reshape input patches into sequences following four distinct traversal paths, enabling selective scanning from multiple spatial perspectives. To adapt these paths efficiently, each Visual State Space module is equipped with a shared down-projection matrix and four learnable low-rank up-projection matrices, one for each traversal direction. By combining the shared initial parameters with the direction-specific low-rank components, our approach provides a flexible yet compact way to fine-tune model weights. To further enhance training efficiency, we introduce an iterative optimization strategy that updates the low-rank parameters and task-specific head layers in alternating steps. This strategy encourages faster convergence and reduced computational overhead, making it particularly valuable for large-scale medical image tasks. Extensive experiments on classification, detection, and segmentation benchmarks confirm that our method surpasses existing parameter-efficient approaches. Ziqiu Dong, Yuetong Luo, Yankong Zhang, Shengeng Tang, Lechao Cheng |
ICIP | 5 |
| 2025 | Exploring Effective Unfolding Covering Prompt Tuning for Vision MambaabstractThe Vision Mamba proposed recently has emerged as a popular architecture solution known for its efficiency in computational resources. However, the core designation of parallelized selective scan operation poses a challenge when performing visual prompt tuning on downstream tasks. The difficulty arises from the conflict between the unordered insertion of tokens in visual prompt tuning and the varying importance of tokens at different positions in vision mamba. To alleviate this issue, we propose an Unfolding Covering prompt tuning strategy that effectively customizes downstream tasks. In this work, we explore several visual prompting strategies to further improve performance with limited data. Exhaustive experiments on general tasks like classification and detection have demonstrated the superiority of our approach. Mingwang Wu, Yuetong Luo, Yankong Zhang, Shengeng Tang, Lechao Cheng |
ICIP | 5 |
| 2025 | Navigating Semantic Drift in Task-Agnostic Class-Incremental LearningabstractClass-incremental learning (CIL) seeks to enable a model to sequentially learn new classes while retaining knowledge of previously learned ones. Balancing flexibility and stability remains a significant challenge, particularly when the task ID is unknown. To address this, our study reveals that the gap in feature distribution between novel and existing tasks is primarily driven by differences in mean and covariance moments. Building on this insight, we propose a novel semantic drift calibration method that incorporates mean shift compensation and covariance calibration. Specifically, we calculate each class's mean by averaging its sample embeddings and estimate task shifts using weighted embedding changes based on their proximity to the previous mean, effectively capturing mean shifts for all learned classes with each new task. We also apply Mahalanobis distance constraint for covariance calibration, aligning class-specific embedding covariances between old and current networks to mitigate the covariance shift. Additionally, we integrate a feature-level self-distillation approach to enhance generalization. Comprehensive experiments on commonly used datasets demonstrate the effectiveness of our approach. The source code is available at https://github.com/fwu11/MACIL.git. Fangwen Wu, Lechao Cheng, Shengeng Tang, Chaowei Fang, Dingwen Zhang, Meng Wang 0001 |
ICML | 3 |
| 2025 | Knowledge Swapping via Learning and UnlearningabstractWe introduce Knowledge Swapping, a novel task designed to selectively regulate knowledge of a pretrained model by enabling the forgetting of user-specified information, retaining essential knowledge, and acquiring new knowledge simultaneously. By delving into the analysis of knock-on feature hierarchy, we find that incremental learning typically progresses from low-level representations to higher-level semantics, whereas forgetting tends to occur in the opposite direction—starting from high-level semantics and moving down to low-level features. Building upon this, we propose to benchmark the knowledge swapping task with the strategy of Learning Before Forgetting. Comprehensive experiments on various tasks like image classification, object detection, and semantic segmentation validate the effectiveness of the proposed strategy. The source code is available at https://github.com/xingmingyu123456/KnowledgeSwapping. Mingyu Xing, Lechao Cheng, Shengeng Tang, Yaxiong Wang, Zhun Zhong, Meng Wang 0001 |
ICML | 3 |
| 2025 | Shaping a Stabilized Video by Mitigating Unintended Changes for Concept-Augmented Video EditingabstractText-driven video editing powered by generative diffusion models holds significant promise for applications spanning film production, advertising, and beyond. However, the limited expressiveness of pre-trained word embeddings often restricts nuanced edits, especially when targeting novel concepts with specific attributes. In this work, we present a novel Concept-Augmented Textual Inversion (CATI) framework that flexibly integrates new object information from user-provided concept videos. By fine-tuning only the V (Value) projection in attention via Low-Rank Adaptation (LoRA), our approach preserves the original attention distribution of the diffusion model while efficiently incorporating external concept knowledge. To further stabilize editing results and mitigate the issue of attention dispersion when prompt keywords are modified, we introduce a Dual Prior Supervision (DPS) mechanism. DPS supervises cross-attention between the source and target prompts, preventing undesired changes to non-target areas and improving the fidelity of novel concepts. Extensive evaluations demonstrate that our plug-and-play solution not only maintains spatial and temporal consistency but also outperforms state-of-the-art methods in generating lifelike and stable edited videos. The source code is publicly available at https://guomc9.github.io/STIVE-PAGE/. Mingce Guo, Jingxuan He 0001, Yufei Yin, Zhangye Wang, Shengeng Tang, Lechao Cheng |
IJCAI | 5 |
| 2025 | Mixture of Multimodal Adapters for Sentiment AnalysisabstractKezhou Chen, Shuo Wang, Huixia Ben, Shengeng Tang, Yanbin Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kezhou Chen, Shuo Wang 0008, Huixia Ben, Shengeng Tang, Yanbin Hao |
NAACL (Long Papers) | 4 |
| 2025 | Leveraging vision-language prompts for real-world image restoration and enhancement
Yanyan Wei, Yilin Zhang 0012, Kun Li 0008, Fei Wang 0067, Shengeng Tang, Zhao Zhang 0001 |
Comput. Vis. Image Underst. | 5 |
| 2025 | Alleviating Confirmation Bias in Learning with Noisy Labels via Two-Network CollaborationabstractDeep neural networks (DNNs) have achieved remarkable success in various computer vision tasks, e.g., image classification. However, most of the existing models depend heavily on annotated data, where label noise is inevitable. Training with such noisy data negatively impacts the generalization performance of DNNs. To this end, recent advances in learning with noisy labels (LNL) adopt the sample selection strategy that identifies clean samples from the noisy dataset to update DNNs, using semi-supervised learning where rejected samples are treated as unlabeled data. However, existing LNL methods often overlook the varying fitting difficulties of different classes, resulting in suboptimal sample selection and confirmation bias, and consequently, the errors accumulate during semi-supervised training. In this article, we propose a novel method, TNCollab, which aims at alleviating confirmation bias in both sample selection and semi-supervised training stages via two-network collaboration. Specifically, we introduce a class-adaptive threshold for sample selection to address the varying fitting difficulties across different classes. Additionally, we construct a hard set consisting of samples where the two networks disagree and introduce a noise-robust loss to extract potentially useful information while maintaining robustness against label noise. Furthermore, we propose a dual consistency loss to ensure consistent predictions between the networks across different augmented views of the same sample, facilitating mutual learning. Extensive experiments demonstrate that TNCollab achieves state-of-the-art performance on image classification and facial expression recognition tasks, particularly on CIFAR-10, CIFAR-100, WebVision, Clothing1M, Tiny-ImageNet, and RAF-DB datasets, showing improved visual understanding and generalization capabilities. Our codes are available at https://github.com/Delete12137/TNCollab . Peipei Song, Shengeng Tang, Dan Guo 0001, Xun Yang 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2025 | Gloss-driven Conditional Diffusion Models for Sign Language ProductionabstractSign Language Production (SLP) aims to convert text or audio sentences into sign language videos corresponding to their semantics, which is challenging due to the diversity and complexity of sign languages, and cross-modal semantic mapping issues. In this work, we propose a Gloss-driven Conditional Diffusion Model (GCDM) for SLP. The core of the GCDM is a diffusion model architecture, in which the sign gloss sequence is encoded by a Transformer-based encoder and input into the diffusion model as a semantic prior condition. In the process of sign pose generation, the textual semantic priors carried in the encoded gloss features are integrated into the embedded Gaussian noise via cross-attention. Subsequently, the model converts the fused features into sign language pose sequences through T-round denoising steps. During the training process, the model uses the ground-truth labels of sign poses as the starting point, generates Gaussian noise through T rounds of noise, and then performs T rounds of denoising to approximate the real sign language gestures. The entire process is constrained by the MAE loss function to ensure that the generated sign language gestures are as close as possible to the real labels. In the inference phase, the model directly randomly samples a set of Gaussian noise, generates multiple sign language gesture sequence hypotheses under the guidance of the gloss sequence, and outputs a high-confidence sign language gesture video by averaging multiple hypotheses. Experimental results on the Phoenix2014T dataset show that the proposed GCDM method achieves competitiveness in both quantitative performance and qualitative visualization. Shengeng Tang, Feng Xue 0002, Jingjing Wu 0001, Shuo Wang 0008, Richang Hong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Temporal Boundary Awareness Network for Repetitive Action CountingabstractRepetitive Action Counting (RAC) is a critical and challenging task in video analysis, aiming to count the number of repeated actions in videos accurately. Existing methods typically generate a Temporal Self-similarity Matrix (TSM) as an intermediate representation to predict the number of repetitive actions. While this simplifies the process, it often overlooks the variable lengths between action cycles and the phenomenon of motion interruptions. The period inconsistency problem caused by the change in the action period and the motion interruption problem resulting from the motion pause are the two main challenges that affect the accuracy of RAC in complex scenes. To address these challenges, we propose a novel framework. First, we construct a boundary-aware encoder equipped with a temporal pyramid structure to build multi-scale video features, capturing the period information of different lengths of repetitive actions to solve the period inconsistency problem. Next, a cycle and boundary attention module is followed by each layer in the pyramid to enhance these multi-scale features with periodic and event boundary information. Finally, we design a gated density estimator to generate the actionness score for each frame that reflects the probability of the corresponding time point being within the motion cycle. These scores are used to weight features to reduce the impact of noise frames without actions present and solve the motion interruption problem for better density prediction. Extensive experiments conducted on public datasets demonstrate the effectiveness of our method. The source code will be available at https://github.com/zqzhang2023/TBANRAC . Zhenqiang Zhang, Kun Li 0008, Shengeng Tang, Yanyan Wei, Fei Wang 0073, Jinxing Zhou, Dan Guo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Emotional Video Captioning With Vision-Based Emotion Interpretation NetworkabstractEffectively summarizing and re-expressing video content by natural languages in a more human-like fashion is one of the key topics in the field of multimedia content understanding. Despite good progress made in recent years, existing efforts usually overlooked the emotions in user-generated videos, thus making the generated sentence a bit boring and soulless. To fill the research gap, this paper presents a novel emotional video captioning framework in which we design a Vision-based Emotion Interpretation Network to effectively capture the emotions conveyed in videos and describe the visual content in both factual and emotional languages. Specifically, we first model the emotion distribution over an open psychological vocabulary to predict the emotional state of videos. Then, guided by the discovered emotional state, we incorporate visual context, textual context, and visual-textual relevance into an aggregated multimodal contextual vector to enhance video captioning. Furthermore, we optimize the network in a new emotion-fact coordinated way that involves two losses- Emotional Indication Loss and Factual Contrastive Loss, which penalize the error of emotion prediction and visual-textual factual relevance, respectively. In other words, we innovatively introduce emotional representation learning into an end-to-end video captioning network. Extensive experiments on public benchmark datasets, EmVidCap and EmVidCap-S, demonstrate that our method can significantly outperform the state-of-the-art methods by a large margin. Quantitative ablation studies and qualitative analyses clearly show that our method is able to effectively capture the emotions in videos and thus generate emotional language sentences to interpret the video content. Peipei Song, Dan Guo 0001, Xun Yang 0001, Shengeng Tang, Meng Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Intermediary-Generated Bridge Network for RGB-D Cross-Modal Re-IdentificationabstractRGB-D cross-modal person re-identification (re-id) targets at retrieving the person of interest across RGB and depth image modalities. To cope with the modal discrepancy, some existing methods generate an auxiliary mode with either inherent properties of input modes or extra deep networks. However, such useful intermediary role included in generated mode is often overlooked in these approaches, leading to insufficient exploitation of crucial bridge knowledge. By contrast, in this article, we propose a novel approach that constructs an intermediary mode through the constraints of self-supervised intermediary learning, which is freedom from modal prior knowledge and additional module parameters. We then design a bridge network to fully mine the intermediary role of generated modality through carrying out multi-modal integration and decomposition. For one thing, this network leverages a multi-modal transformer to integrate the information of three modes via fully exploiting their heterogeneous relations with the intermediary mode as the bridge. It conducts the identification consistency constraint to promote cross-modal associations. For another, it employs circle contrastive learning to decompose the cross-modal constraint process into several subprocedures, which provides the intermediate relay during pulling two original modalities closer. Experiments on two public datasets demonstrate that the proposed method exceeds the state-of-the-arts. The effectiveness of each component in this method is verified through numerous ablation studies. Additionally, we have demonstrated the generalization ability of the proposed method through experiments. Jingjing Wu 0001, Richang Hong, Shengeng Tang |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2023 | Emotion-Prior Awareness Network for Emotional Video CaptioningabstractEmotional video captioning (EVC) is an emerging task to describe the factual content with the inherent emotion expressed in a video. It is crucial for the EVC task to effectively perceive subtle and ambiguous visual emotion cues in the stage of caption generation. However, existing captioning methods usually overlooked the learning of emotions in user-generated videos, thus making the generated sentence a bit boring and soulless. Peipei Song, Dan Guo 0001, Xun Yang 0001, Shengeng Tang, Erkun Yang, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2022 | Gloss Semantic-Enhanced Network with Online Back-Translation for Sign Language ProductionabstractSign Language Production (SLP) aims to generate the visual appearance of sign language according to the spoken language, in which a key procedure is to translate sign Gloss to Pose (G2P). Existing G2P methods mainly focus on regression prediction of posture coordinates, namely closely fitting the ground truth. In this paper, we provide a new viewpoint: a Gloss semantic-Enhanced Network is proposed with Online Back-Translation (GEN-OBT) for G2P in the SLP task. Specifically, GEN-OBT consists of a gloss encoder, a pose decoder, and an online reverse gloss decoder. In the gloss encoder based on the transformer, we design a learnable gloss token without any prior knowledge of gloss, to explore the global contextual dependency of the entire gloss sequence. During sign pose generation, the gloss token is aggregated onto the existing generated poses as gloss guidance. Then, the aggregated features are interacted with the entire gloss embedding vectors to generate the next pose. Furthermore, we design a CTC-based reverse decoder to convert the generated poses backward into glosses, which guarantees the semantic consistency during the processes of gloss-to-pose and pose-to-gloss. Extensive experiments on the challenging PHOENIX14T benchmark demonstrate that the proposed GEN-OBT outperforms the state-of-the-art models. Visualization results further validate the interpretability of our method. Shengeng Tang, Richang Hong, Dan Guo 0001, Meng Wang 0001 |
ACM Multimedia | 1 |
| 2022 | Graph-Based Multimodal Sequential Embedding for Sign Language TranslationabstractSign language translation (SLT) is a challenging weakly supervised task without word-level annotations. An effective method of SLT is to leverage multimodal complementarity and to explore implicit temporal cues. In this work, we propose a graph-based multimodal sequential embedding network (MSeqGraph), in which multiple sequential modalities are densely correlated. Specifically, we build a graph structure to realize the intra-modal and inter-modal correlations. First, we design a graph embedding unit (GEU), which embeds a parallel convolution with channel-wise and temporal-wise learning into the graph convolution to learn the temporal cues in each modal sequence and cross-modal complementarity. Then, a hierarchical GEU stacker with a pooling-based skip connection is proposed. Unlike the state-of-the-art methods, to obtain a compact and informative representation of multimodal sequences, the GEU stacker gradually compresses the channel$d$with multi-modalities$m$rather than the temporal dimension$t$. Finally, we adopt the connectionist temporal decoding strategy to explore the entire video’s temporal transition and translate the sentence. Extensive experiments on the USTC-CSL and BOSTON-104 datasets demonstrate the effectiveness of the proposed method. Shengeng Tang, Dan Guo 0001, Richang Hong, Meng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | Connectionist Temporal Modeling of Video and Language: a Joint Model for Translation and Sign LabelingabstractOnline sign interpretation suffers from challenges presented by hybrid semantics learning among sequential variations of visual representations, sign linguistics, and textual grammars. This paper proposes a Connectionist Temporal Modeling (CTM) network for sentence translation and sign labeling. To acquire short-term temporal correlations, a Temporal Convolution Pyramid (TCP) module is performed on 2D CNN features to realize (2D+1D)=pseudo 3D' CNN features. CTM aligns the pseudo 3D' with the original 3D CNN clip features and fuses them. Next, we implement a connectionist decoding scheme for long-term sequential learning. Here, we embed dynamic programming into the decoding scheme, which learns temporal mapping among features, sign labels, and the generated sentence directly. The solution using dynamic programming to sign labeling is considered as pseudo labels. Finally, we utilize the pseudo supervision cues in an end-to-end framework. A joint objective function is designed to measure feature correlation, entropy regularization on sign labeling, and probability maximization on sentence decoding. The experimental results using the RWTH-PHOENIX-Weather and USTC-CSL datasets demonstrate the effectiveness of the proposed approach. Dan Guo 0001, Shengeng Tang, Meng Wang 0001 |
IJCAI | 2 |