Bin Sun 0001

dblp:01/5401-1 · DBLP profile ↗
← Back
44ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0002-7029-8784ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 10 since 2021
YearPublicationVenuePosition
2026 Edge-guided feature fusion for real-time infrared instance segmentation of power transmission lines
Zhuoliang Yang, Bin Sun 0001, Junyi Tao, Tianfu Gong
Neurocomputing3
2026 Mamba-enhanced local attention network for remote sensing image super-resolution
Luoxin Zhao, Bin Sun 0001, Liguo Liu, Xudong Kang, Shutao Li 0001
Neurocomputing2
2026 Adaptive Contrastive Learning for Semisupervised Facial Expression Recognition
abstract
Semisupervised deep facial expression recognition (FER) tries to learn better representations from both labeled and unlabeled data to avoid the huge manual labeling cost in the supervised methods. Due to the imbalanced distribution of facial expression data and the varying difficulty in recognizing different expressions, conventional semisupervised FER methods that directly discard low-confidence unlabeled samples may exacerbate the deficiency of the minority classes, and trap the model into the Matthew effect. Therefore, an adaptive contrastive learning-based semisupervised deep facial expression recognition framework is proposed to utilize both high and low confidence unlabeled samples by two different contrastive learning strategies according to an adaptive threshold. The threshold varies from different classes and iterations, considering both the difficulty of different classes and current learning state of the model during the training process. After the adaptive thresholding, the class-aware contrastive learning is applied to the high confidence samples, while the distance-aware contrastive learning is to the low confidence ones. Our proposed method achieves the state-of-the-art performance through extensive experiments on two widely-used datasets RAF-DB and AffectNet.
Bin Sun 0001, Meiqi Liao, Shutao Li 0001, Fuyan Ma
IEEE Trans. Comput. Soc. Syst.1
2026 SPEN: Sub-Pixel Position Error Estimation Network for Multi-Modal Image Matching
Maoqing Hu, Bin Sun 0001, Shutao Li 0001, Jiayi Ma 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 DrVideo: Document Retrieval Based Long Video Understanding
abstract
Most of the existing methods for video understanding primarily focus on videos only lasting tens of seconds, with limited exploration of techniques for handling long videos. The increased number of frames in long videos poses two main challenges: difficulty in locating key information and performing long-range reasoning. Thus, we propose DrVideo, a document-retrieval-based system designed for long video understanding. Our key idea is to convert the long-video understanding problem into a long-document understanding task so as to effectively leverage the power of large language models. Specifically, DrVideo first transforms a long video into a coarse text-based long document to initially retrieve key frames and then updates the documents with the augmented key frame information. It then employs an agent-based iterative loop to continuously search for missing information and augment the document until sufficient question-related information is gathered for making the final predictions in a chain-of-thought manner. Extensive experiments on long video benchmarks confirm the effectiveness of our method. DrVideo significantly outperforms existing LLM-based state-of-the-art methods on EgoSchema benchmark (3 minutes), MovieChat-1K benchmark (10 minutes), and the long split of Video-MME benchmark (average of 44 minutes). Code is available at https://github.com/Upper9527/DrVideo.
Ziyu Ma, Chenhui Gou, Hengcan Shi, Bin Sun 0001, Shutao Li 0001, Seyed Hamid Rezatofighi, Jianfei Cai 0001
CVPR4
2025 Multimodal Prompt Alignment for Facial Expression Recognition
abstract
Prompt learning has been widely adopted to efficiently adapt vision-language models (VLMs) like CLIP for various downstream tasks. Despite their success, current VLM-based facial expression recognition (FER) methods struggle to capture fine-grained textual-visual relationships, which are essential for distinguishing subtle differences between facial expressions. To address this challenge, we propose a multimodal prompt alignment framework for FER, called MPA-FER, that provides fine-grained semantic guidance to the learning process of prompted visual features, resulting in more precise and interpretable representations. Specifically, we introduce a multi-granularity hard prompt generation strategy that utilizes a large language model (LLM) like ChatGPT to generate detailed descriptions for each facial expression. The LLM-based external knowledge is injected into the soft prompts by minimizing the feature discrepancy between the soft prompts and the hard prompts. To preserve the generalization abilities of the pretrained CLIP model, our approach incorporates prototype-guided visual feature alignment, ensuring that the prompted visual features from the frozen image encoder align closely with class-specific prototypes. Additionally, we propose a cross-modal global-local alignment module that focuses on expression-relevant facial features, further improving the alignment between textual and visual features. Extensive experiments demonstrate our framework outperforms state-of-the-art methods on three FER benchmark datasets, while retaining the benefits of the pretrained model and minimizing computational costs.
Fuyan Ma, Bin Sun 0001, Shutao Li 0001
ICCV3
2025 Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic Guidance
abstract
Open-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding and conversational fluency. However, existing LLM-based dialogue systems often fall short in proactively understanding the user's chatting preferences and guiding conversations toward user-centered topics. This lack of user-oriented proactivity can lead users to feel unappreciated, reducing their satisfaction and willingness to continue the conversation in human-computer interactions. To address this issue, we propose a User-oriented Proactive Chatbot (UPC) to enhance the user-oriented proactivity. Specifically, we first construct a critic to evaluate this proactivity inspired by the LLM-as-a-judge strategy. Given the scarcity of high-quality training data, we then employ the critic to guide dialogues between the chatbot and user agents, generating a corpus with enhanced user-oriented proactivity. To ensure the diversity of the user backgrounds, we introduce the ISCO-800, a diverse user background dataset for constructing user agents. Moreover, considering the communication difficulty varies among users, we propose an iterative curriculum learning method that trains the chatbot from easy-to-communicate users to more challenging ones, thereby gradually enhancing its performance. Experiments demonstrate that our proposed training method is applicable to different LLMs, improving user-oriented proactivity and attractiveness in open-domain dialogues. Code and appendix are available at github.com/wang678/LLM-UPC.
Yufeng Wang 0004, Jinwu Hu, Ziteng Huang, Kunyang Lin, Zitian Zhang, Peihao Chen, Yu Hu 0004, Qianyue Wang, Zhu Liang Yu, Bin Sun 0001, Xiaofen Xing, Mingkui Tan
IJCAI10
2025 Directional-Semantic-Enhanced Visual Grounding for Remote Sensing Images
abstract
Visual grounding for remote sensing images (RSVG) is a fundamental vision-language task, which aims to locate the objects referred to by the natural language expression from the RS images. Natural language expressions often rely heavily on detailed directional information to describe target objects. Therefore, the thorough and effective utilization of spatial directional information is crucial for accurately locating the referred objects within complex RS images. However, most existing RSVG methods fail to fully leverage directional information, leading to suboptimal outcomes. This paper introduces a novel directional semantic enhanced method for RSVG, dubbed DSEVG. Specifically, we propose a scale-adaptive language-guided interaction (SLI) module that derives scale-specific language features through a hierarchical language adaptation mechanism. These scale-specific language features guide the visual backbone to extract visual features highly relevant to referring expressions. Furthermore, we present a directional semantic enhancement (DSE) module that implicitly enhances directional semantics in referring expressions by leveraging visual spatial information. It also employs an explicit spatial semantic alignment loss to supervise this process, generating localization prompts with enhanced directional representations. These prompts are injected into the queries of each decoder layer to guide the model in effectively utilizing directional semantics. Experimental results on the DIOR-RSVG and OPT-RSVG benchmark datasets validate the effectiveness of the proposed method and demonstrate state-of-the-art performance. Code is available at: https://github.com/WH231203/DSEVG.
Hu Guo, Bin Sun 0001, Shutao Li 0001, Chenglong Lei, Xiliang Li, Mingkui Tan
IEEE Trans. Geosci. Remote. Sens.2
2025 FMA-Net: Flow-Driven Motion-Aware Network for Multiobject Tracking in Satellite Videos
abstract
Satellite video multi-object tracking (MOT) is fundamentally challenged by extremely weak inter-frame displacement. Most existing methods capture motion at the image level, limiting the effective utilization of fine-grained motion information. To address this limitation, we propose a Flow-driven Motion-Aware Network(FMA-Net) that performs pixel-level motion modeling via the designed Flow-driven Motion Estimator (FME). The captured fine-grained motion is then fully exploited by the Motion-Aware Feature Fusion (MAF) module to guide inter-frame feature fusion, allowing the network to focus on truly dynamic targets during detection. The Motion-Aware Refinement (MAR) module leverages the motion priors to maintain identity consistency and reduce mismatches during association, improving tracking robustness. By jointly modeling and exploiting fine-grained motion cues throughout both detection and association, FMA-Net significantly enhances tracking accuracy in satellite video MOT. Extensive experiments on three satellite video MOT benchmarks demonstrate that FMA-Net achieves state-of-the-art performance across diverse object categories, validating its effectiveness and generalization capability. The code will be publicly available at: https://github.com/luweiqing/FMA-Net.
Weiqing Lu, Bin Sun 0001, Shutao Li 0001, Xiliang Li
IEEE Trans. Geosci. Remote. Sens.2
2025 Hierarchical Augmentation and Region-Aware Contrastive Learning for Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
Semi-supervised semantic segmentation has gained significant attention as a method to reduce the substantial expense associated with pixel-level labeling. The existing methods primarily rely on consistency regularization or self-training. Recent consistency regularization methods augment the input images with weak or strong augmentation (SA) to improve the performance. However, such simple augmentations are not sufficient to simulate the variations in remote sensing images. The self-training methods exclude the noisy pseudo labels by some selection from the unsupervised training process to obtain better performance. The selection may lead to semantic information loss and bias of latent distribution. To solve the above two problems, we propose hierarchical augmentation (HA) and region-aware contrastive (RC) learning, namely HARC, for remote sensing images. The HA strategy simulates three levels of remote sensing image variations, i.e., spatial variations, uniform spectral variations, and uneven spectral variations. It can significantly enhance the model’s capability to handle more intricate variations. The RC learning learns a class-wise feature distribution of all unlabeled samples instead of some screened unlabeled samples. It can eliminate semantic information loss and enhance the model’s resistance to noise from pseudo labels. Our method is evaluated on three public remote sensing datasets, and the experimental results demonstrate its superiority over state-of-the-art (SOTA) semi-supervised methods.
Bin Sun 0001, Shutao Li 0001, Yulong Hu
IEEE Trans. Geosci. Remote. Sens.2
2025 HLDD: Hierarchically Learned Detector and Descriptor for Robust Image Matching
abstract
Image matching is a critical task in computer vision research, focusing on aligning two or more images with similar features. Feature detection and description constitute the core of image matching. Handcrafted detectors are capable of obtaining distinctive points but these points may not be repeatable on the image pairs especially those with dramatic appearance changes. On the contrary, the learned detectors can extract a large number of repeatable points but many of them tend to be ambiguous points with low distinctiveness. Moreover, in the scenarios of dramatic appearance change, commonly used contrast or triplet loss in the training of descriptors employ the hard negative mining strategy, which may obtain overly challenging negative samples by global sampling, resulting in sluggish convergence or even overfitting. Those learned descriptors may not guarantee that the corresponding points enjoy larger similarities than unmatched ones, leading to inaccurate matches. To address those issues, we propose a hierarchically learned detector and descriptor (HLDD) for robust image matching, which contains three modules: a handcrafted-learned detector, a hierarchically learned descriptor, and a coarse-to-fine matching strategy. The handcrafted-learned detector integrates the advantages of handcrafted and learned detectors. It extracts distinctive feature points from a learned repeatability map robust to image changes and eliminates the ambiguous ones according to a learned distinctiveness map. The descriptor is trained by a proposed hierarchical triplet loss, which employs a dual window strategy. It can obtain the hardest negative samples in local windows, which are comparatively easier over global sampling, ensuring the effective training of descriptors. The coarse-to-fine matching strategy performs global and local mutual nearest neighbor matching on the coarse and fine descriptor maps respectively to improve the matching accuracy progressively. By comparing with other matching methods, experimental results demonstrate the superiority of the proposed method in the task of image matching, homography estimation, visual localization, and relative pose estimation. Moreover, ablation studies illustrate the effectiveness of the three proposed modules.
Maoqing Hu, Bin Sun 0001, Fuhua Zhang 0002, Shutao Li 0001
IEEE Trans. Image Process.2
2025 DiffCL: A Diffusion-Based Contrastive Learning Framework With Semantic Alignment for Multimodal Recommendations
abstract
Multimodal recommendation systems integrate diverse multimodal information into the feature representations of both items and users, thereby enabling a more comprehensive modeling of user preferences. However, existing methods are hindered by data sparsity and the inherent noise within multimodal data, which impedes the accurate capture of users' interest preferences. Additionally, discrepancies in the semantic representations of items across different modalities can adversely impact the prediction accuracy of recommendation models. To address these challenges, we introduce a novel diffusion-based contrastive learning (DiffCL) framework for multimodal recommendation. DiffCL employs a diffusion model (DM) to generate contrastive views that effectively mitigate the impact of noise during the contrastive learning phase. Furthermore, it improves semantic consistency across modalities by aligning distinct visual and textual semantic information through stable ID embeddings. Finally, the introduction of the item-item graph (I-I graph) enhances multimodal feature representations, thereby alleviating the adverse effects of data sparsity on the overall system performance. We conduct extensive experiments on three public datasets, and the results demonstrate the superiority and effectiveness of the DiffCL.
Qiya Song, Jiajun Hu, Lin Xiao 0002, Bin Sun 0001, Xieping Gao 0001, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.4
2025 Cascade Fusion and Correlation Enhancement for Knowledge Distillation
abstract
Knowledge distillation (KD) improves the performance of a compact student network by transferring learned knowledge from a cumbersome teacher network. In the existing approaches, the multiscale feature knowledge is transferred via densely connected paths, which increases the optimization difficulty. Moreover, correlations among the labels are neglected despite their capability to enhance the intraclass similarity of samples. To solve these issues, we propose cascade fusion and correlation enhancement for KD (CC-KD). The multiscale feature knowledge is transferred via much simpler paths, which are constructed by fusing features of different scales with cross-scale attention (CSA) in a cascade manner, thereby reducing the optimization difficulty. On the other hand, the relational knowledge of teacher logits is further enhanced by correlations of the corresponding labels, so that the student can produce more similar logits for the samples in the same category. Extensive experimental results on five public datasets (i.e., CIFAR100/10, ImageNet, RAF-DB, and FERPlus) indicate superior performance of the proposed method over several state-of-the-arts (SOTAs). More specifically, our method obtains an accuracy of 71.70% on ImageNet and achieves a new record of 90.20% on RAF-DB with fewer calculations and parameters.
Bin Sun 0001, Zuxiang Long, Ziyu Ma, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Detection Assisted Change Captioning for Remote Sensing Image
abstract
Remote sensing image change captioning is a crucial image interpretation technique that auto-generates language captions of differences between multi-temporal remote sensing images. Previous attention-based methods were difficult to generate accurate captions due to their inability to precisely locate crucial visual change areas. To address this challenge, this paper introduces a novel method that aims to leverage explicit visual change information to enhance its change description capabilities. Specifically, the proposed model comprises three key components: 1) the change-visual enhancement module leverages the change image containing object-level visual information to enhance the multi-temporal images at the image level; 2) the multi-temporal feature fusion module captures accurate visual change features through a meticulously designed feature fusion at feature level; 3) the caption generation module inputs the visual change features into transformer-based generator to produce desired captions of multi-temporal remote sensing images. Experimental results on LEVIR-CC dataset demonstrate that our method has achieved state-of-the-art performance.
Xiliang Li, Bin Sun 0001, Shutao Li 0001
IGARSS2
2024 Region-Aware Contrastive Learning For Semi-Supervised Semantic Segmentation of Remote Sensing Images
abstract
Semi-supervised semantic segmentation has attracted a lot of attention due to the high cost of obtaining pixel-level labels. To get a good grade for semi-supervised semantic segmentation, many methods based on self-training have been proposed. The existing self-training method achieves semi-supervised learning by identifying the samples with pseudo labels with a threshold of its confidence or a classifier. The samples with pseudo labels enrich the training dataset so that the semantic segmentation performance is improved. However, it inevitably introduces samples with wrong pseudo labels, which may harm the final performance. To tackle this problem, we proposed a new supervision to make the model trained well in the pseudo labels that never be filtered. In particular, the region-aware contrastive learning constructs a class-level feature space. Then the class-level feature space is subsequently used for supervising the training on unlabeled images. The experimental results on Vaihingen dataset show that our method outperforms the state-of-the-art semi-supervised methods.
Bin Sun 0001, Shutao Li 0001
IGARSS2
2024 Less is More: Adaptive Feature Selection and Fusion for Eye Contact Detection
abstract
Detecting eye contact is essential for embodied robots to engage in natural interactions with humans, enhancing the intuitiveness and comfort of these exchanges. However, eye contact detection often presents a significant challenge due to a variety of factors, such as low contrast and various forms of occlusions. Existing methods incorporate convolutional neural networks (CNNs) or Transformers to learn discriminative representations, but usually ignore the influence of noisy or less relevant regions in facial images. To address this gap, we propose the deep feature selection and fusion network (FSFNet) for eye contact detection in multi-party conversations. Our proposed method adaptively selects fine-grained visual features and reduces the impacts of irrelevant features. Specifically, we present a local feature selection scheme that leverages the attention scores to progressively concentrate on the most informative features. By integrating the carefully selected features into the multi-head self-attention module, we can maintain the superior properties of Transformers while simultaneously reducing the overall computational demands. We evaluate the proposed method on the official eye contact detection datasets, which achieves promising results of 0.8174 and 0.79 on the validation and test sets, respectively. We have made the source code publicly accessible in https://github.com/ma-hnu/FSFNet.
Fuyan Ma, Bin Sun 0001, Shutao Li 0001
ACM Multimedia3
2024 Large Language Models With Holistically Thought Could Be Better Doctors
Yixuan Weng, Bin Li 0083, Minjun Zhu, Bin Sun 0001, Shizhu He, Shengping Liu, Kang Liu 0001, Shutao Li 0001, Jun Zhao 0001
NLPCC (2)5
2024 Distinct but correct: generating diversified and entity-revised medical response
Bin Li 0083, Bin Sun 0001, Shutao Li 0001, Encheng Chen, Hongru Liu, Yixuan Weng, Yongping Bai, Meiling Hu
Sci. China Inf. Sci.2
2024 Towards Visual-Prompt Temporal Answer Grounding in Instructional Video
abstract
Temporal answer grounding in instructional video (TAGV) is a new task naturally derived from temporal sentence grounding in general video (TSGV). Given an untrimmed instructional video and a text question, this task aims at locating the frame span from the video that can semantically answer the question, i.e., visual answer. Existing methods tend to solve the TAGV problem with a visual span-based predictor, taking visual information to predict the start and end frames in the video. However, due to the weak correlations between the semantic features of the textual question and visual answer, current methods using the visual span-based predictor do not work well in the TAGV task. In this paper, we propose a visual-prompt text span localization (VPTSL) method, which introduces the timestamped subtitles for a text span-based predictor. Specifically, the visual prompt is a learnable feature embedding, which brings visual knowledge to the pre-trained language model. Meanwhile, the text span-based predictor learns joint semantic representations from the input text question, video subtitles, and visual prompt feature with the pre-trained language model. Thus, the TAGV is reformulated as the task of the visual-prompt subtitle span localization for the visual answer. Extensive experiments on five instructional video datasets, namely MedVidQA, TutorialVQA, VehicleVQA, CrossTalk and Coin, show that the proposed method outperforms several state-of-the-art (SOTA) methods by a large margin in terms of mIoU score, which demonstrates the effectiveness of the proposed visual prompt and text span-based predictor.
Shutao Li 0001, Bin Li 0083, Bin Sun 0001, Yixuan Weng
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Transformer-Augmented Network With Online Label Correction for Facial Expression Recognition
abstract
Facial expression recognition (FER) in the wild is extremely challenging due to occlusions, variant head poses under unconstrained conditions and incorrect annotations (e.g., label noise). In this paper, we aim to improve the performance of in-the-wild FER with Transformers and online label correction. Different from pure CNNs based methods, we propose a Transformer-augmented network (TAN) to dynamically capture the relationships within each facial patch and across the facial patches. Specifically, the TAN translates a number of facial patch images into a set of visual feature sequences by a backbone convolutional neural network. The intra-patch Transformer is subsequently utilized to capture the most discriminative features within each visual feature sequence. The position-disentangled attention mechanism of the intra-patch Transformer is proposed to better incorporate the positional information for feature sequences. Furthermore, we propose the inter-patch Transformer to model the dependencies across these feature sequences. More importantly, we present the online label correction (OLC) framework to correct suspicious hard labels and accumulate soft labels based on the predictions of the model, which strengthens the robustness of our model against label noise. We validate our method on several widely-used datasets (RAF-DB, FERPlus, AffectNet), realistic occlusion and pose variation datasets, and synthetic noisy datasets. Extensive experiments on these benchmarks demonstrate that the proposed method performs favorably against state-of-the-art methods. The source code will be made publicly available.
Fuyan Ma, Bin Sun 0001, Shutao Li 0001
IEEE Trans. Affect. Comput.2
2024 A Two-Stage Selective Fusion Framework for Joint Intent Detection and Slot Filling
abstract
Spoken language understanding (SLU) is the core of the speech-centric human-robot interaction system, which mainly involves intent detection and slot filling. The recent SLU research focuses on the joint modeling of the two tasks due to their correlation. Furthermore, the slot information consists of slot position and slot type. Although the slot types are semantically related to the intent, the slot positions of the same intent may vary a lot in different utterances due to the diversity of spoken language. Thus, the conventional one-stage slot filling task may introduce unrelated information for slot position prediction in the slot-intent interaction of the joint modeling. Therefore, we propose a novel two-stage selective fusion framework for joint intent detection and slot filling. Unlike the previous one-stage framework, the proposed framework decomposes the slot filling into two stages, i.e., the slot proposal and slot classification. The slot proposal network consisting of BERT and bidirectional long short-term memory (Bi-LSTM)-conditional random field (CRF) predicts the slot positions. Instead of the tokenwise fusion in the existing methods, the slot-intent feature fusion is only performed in the slot classification. A selective fusion mechanism is designed to facilitate the slot-intent interaction within each slot candidate for more accurate slot-type classification. Experiments on five standard benchmarks (i.e., ATIS, SNIPS, MixATIS, MixSNIPS, and DSTC4) show that the proposed framework achieves the best performance in comparison with several state-of-the-art methods.
Ziyu Ma, Bin Sun 0001, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Learning To Locate Visual Answer In Video Corpus Using Question
abstract
We introduce a new task, named video corpus visual answer localization (VCVAL), which aims to locate the visual answer in a large collection of untrimmed instructional videos using a natural language question. This task requires a range of skills - the interaction between vision and language, video retrieval, passage comprehension, and visual answer localization. In this paper, we propose a cross-modal contrastive global-span (CCGS) method for the VCVAL, jointly training the video corpus retrieval and visual answer localization subtasks with the global-span matrix. We have reconstructed a dataset named MedVidCQA, on which the VCVAL task is benchmarked. Experimental results show that the proposed method outperforms other competitive methods both in the video corpus retrieval and visual answer localization sub-tasks. Most importantly, we perform detailed analyses on extensive experiments, paving a new path for understanding the instructional videos, which ushers in further research1.
Bin Li 0083, Yixuan Weng, Bin Sun 0001, Shutao Li 0001
ICASSP3
2023 Logo-Former: Local-Global Spatio-Temporal Transformer for Dynamic Facial Expression Recognition
abstract
Previous methods for dynamic facial expression recognition (DFER) in the wild are mainly based on Convolutional Neural Networks (CNNs), whose local operations ignore the long-range dependencies in videos. Transformer-based methods for DFER can achieve better performances but result in higher FLOPs and computational costs. To solve these problems, the local-global spatio-temporal Transformer (LOGO-Former) is proposed to capture discriminative features within each frame and model contextual relationships among frames while balancing the complexity. Based on the priors that facial muscles move locally and facial expressions gradually change, we first restrict both the space attention and the time attention to a local window to capture local interactions among feature tokens. Furthermore, we perform the global attention by querying a token with features from each local window iteratively to obtain long-range information of the whole video sequence. In addition, we propose the compact loss regularization term to further encourage the learned features have the minimum intra-class distance and the maximum inter-class distance. Experiments on two in-the-wild dynamic facial expression datasets (i.e., DFEW and FERV39K) indicate that our method provides an effective way to make use of the spatial and temporal dependencies for DFER.
Fuyan Ma, Bin Sun 0001, Shutao Li 0001
ICASSP2
2023 Multi-scale Conformer Fusion Network for Multi-participant Behavior Analysis
abstract
Understanding and elucidating human behavior across diverse scenarios represents a pivotal research challenge in pursuing seamless human-computer interaction. However, previous research on multi-participant dialogues has mostly relied on proprietary datasets, which are not standardized and openly accessible. To propel advancements in this domain, the MultiMediate'23 Challenge presents two sub-challenges: Eye contact detection and Next speaker prediction, aiming to foster a comprehensive understanding of multi-participant behavior. To tackle these challenges, we propose a multi-scale conformer fusion network (MSCFN) for enhancing the perception of multi-participant group behaviors. The conformer block combines the strengths of transformers and convolution networks to facilitate the establishment of global and local contextual relationships between sequences. Then the output features from all Conformer blocks are concatenated to fusion multi-scale representations. Our proposed method was evaluated using the officially provided dataset, and it achieves the best and second best performance in next speaker prediction and gaze detection tasks of MultiMediate'23, respectively.
Qiya Song, Renwei Dian, Bin Sun 0001, Jie Xie 0002, Shutao Li 0001
ACM Multimedia3
2023 Overview of the NLPCC 2023 Shared Task: Chinese Medical Instructional Video Question Answering
Bin Li 0083, Yixuan Weng, Hu Guo, Bin Sun 0001, Shutao Li 0001, Mengyao Qi, Xufei Liu, Yuwei Han, Haiwen Liang, Shuting Gao
NLPCC (3)4
2023 Object and attribute recognition for product image with self-supervised learning
Yi Li 0075, Bin Sun 0001
Neurocomputing3
2023 Facial Expression Recognition With Visual Transformers and Attentional Selective Fusion
abstract
Facial Expression Recognition (FER) in the wild is extremely challenging due to occlusions, variant head poses, face deformation and motion blur under unconstrained conditions. Although substantial progresses have been made in automatic FER in the past few decades, previous studies were mainly designed for lab-controlled FER. Real-world occlusions, variant head poses and other issues definitely increase the difficulty of FER on account of these information-deficient regions and complex backgrounds. Different from previous pure CNNs based methods, we argue that it is feasible and practical to translate facial images into sequences of visual words and perform expression recognition from a global perspective. Therefore, we propose the Visual Transformers with Feature Fusion (VTFF) to tackle FER in the wild by two main steps. First, we propose the attentional selective fusion (ASF) for leveraging two kinds of feature maps generated by two-branch CNNs. The ASF captures discriminative information by fusing multiple features with the global-local attention. The fused feature maps are then flattened and projected into sequences of visual words. Second, inspired by the success of Transformers in natural language processing, we propose to model relationships between these visual words with the global self-attention. The proposed method is evaluated on three public in-the-wild facial expression datasets (RAF-DB, FERPlus and AffectNet). Under the same settings, extensive experiments demonstrate that our method shows superior performance over other methods, setting new state of the art on RAF-DB with 88.14%, FERPlus with 88.81% and AffectNet with 61.85%. The cross-dataset evaluation on CK+ shows the promising generalization capability of the proposed method.
Fuyan Ma, Bin Sun 0001, Shutao Li 0001
IEEE Trans. Affect. Comput.2
2023 Multimodal Sparse Transformer Network for Audio-Visual Speech Recognition
abstract
Automatic speech recognition (ASR) is the major human-machine interface in many intelligent systems, such as intelligent homes, autonomous driving, and servant robots. However, its performance usually significantly deteriorates in the presence of external noise, leading to limitations of its application scenes. The audio-visual speech recognition (AVSR) takes visual information as a complementary modality to enhance the performance of audio speech recognition effectively, particularly in noisy conditions. Recently, the transformer-based architectures have been used to model the audio and video sequences for the AVSR, which achieves a superior performance. However, its performance may be degraded in these architectures due to extracting irrelevant information while modeling long-term dependences. In addition, the motion feature is essential for capturing the spatio-temporal information within the lip region to best utilize visual sequences but has not been considered in the AVSR tasks. Therefore, we propose a multimodal sparse transformer network (MMST) in this article. The sparse self-attention mechanism can improve the concentration of attention on global information by selecting the most relevant parts wisely. Moreover, the motion features are seamlessly introduced into the MMST model. We subtly allow motion-modality information to flow into visual modality through the cross-modal attention module to enhance visual features, thereby further improving recognition performance. Extensive experiments conducted on different datasets validate that our proposed method outperforms several state-of-the-art methods in terms of the word error rate (WER).
Qiya Song, Bin Sun 0001, Shutao Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2022 TA-CNN: A Unified Network for Human Behavior Analysis in Multi-Person Conversations
abstract
Human behavior analysis in multi-person conversations has been one of the most important research issues for natural human-robot interaction. However, previous datasets and studies mainly focus on single-person behavior analysis, therefore, can hardly be generalized in real-world application scenarios. Fortunately, the MultiMediate'22 Challenge provides various video clips of multi-party conversations. In this paper, we present a unified network named TA-CNN for both sub-challenges. Our TA-CNN can not only model the spatio-temporal dependencies for eye contact detection, but also capture the group-level discriminative features for multi-label next speaker prediction. We empirically evaluate the performance of our method on the officially provided datasets. Our method achieves the state-of-the-art result of 0.7261 for eye contact detection in terms of accuracy and the UAR of 0.5965 for next speaker prediction on the corresponding test sets.
Fuyan Ma, Ziyu Ma, Bin Sun 0001, Shutao Li 0001
ACM Multimedia3
2022 Scene-Aware Prompt for Multi-modal Dialogue Understanding and Generation
Bin Li 0083, Yixuan Weng, Ziyu Ma, Bin Sun 0001, Shutao Li 0001
NLPCC (2)4
2022 Deep Fusion of Spectral-Spatial Priors for Cropland Segmentation in Remote Sensing Images
abstract
Cropland segmentation is one of the critical techniques in the agriculture Remote Sensing (RS). Although the Deep Learning (DL) methods have achieved remarkable performance in the natural vision, the cropland segmentation of RS images still suffers from cropland adhesion due to the interference from the surrounding environment and the cropland cover. To tackle this problem, this letter proposes a two stage DL method with spectral-spatial priors. In the first stage, the Multi-feature Extraction Module (MEM) is designed to predict the boundary, an important spatial prior of the cropland. In the second stage, the spatial prior is further fused with the spectral prior by MEMs to get accurate cropland prediction. To evaluate the effectiveness and robustness of the proposed method, we construct a data set called Jiaxiang Cropland Set (JCS) and propose a region level evaluation indicator namely the Plot Mean Intersection over Union (PMIoU). The experiment results on the JCS demonstrate that the proposed method is both qualitatively and quantitatively competitive compared with the state-of-the-art methods.
Laifeng Huang, Bin Sun 0001, Wei Sun 0029, Shutao Li 0001
IEEE Geosci. Remote. Sens. Lett.3
2022 Curvature Filters-Based Multiscale Feature Extraction for Hyperspectral Image Classification
abstract
Exploring fast and effective spectral-spatial feature extraction algorithms for hyperspectral image (HSI) classification is one of the most focus problems in current hyperspectral remote-sensing research. Generally, the size of homogeneous regions in HSIs is not consistent in real scenario and real scenario usually consist of ground objects of different scales. Multiscale strategy starts to be used to construct discriminative features at different scales for HSI classification in recent years. To efficiently characterize the multiscale spectral-spatial features of HSIs, a curvature filters-based multiscale feature extraction method with multiscale superpixel segmentation constraint is proposed. The proposed algorithm is composed of the following major stages. First, global multiscale spectral-spatial features are efficiently extracted via progressively curvature filtering and downsampling operations, which can be regarded as an image pyramid decomposition method. Next, a multiscale superpixel segmentation strategy is applied on the first layer of the image pyramid, and a weighted mean operation is applied within and among superpixels to extract the local multiscale spatial features (LMSFs). Finally, the global multiscale curvature features (GMCFs) and the superpixel segmentation-based LMSFs are fused to form the final multiscale spectral-spatial features for classification purposes. To verify the capabilities of the proposed method, comprehensive experiments are performed on five real hyperspectral datasets. Experimental results demonstrate that the proposed method can significantly improve the classification accuracies compared to several standard HSI feature extraction and classification methods, especially when the number of samples for training is limited.
Qiaobo Hao, Bin Sun 0001, Shutao Li 0001, Melba M. Crawford, Xudong Kang
IEEE Trans. Geosci. Remote. Sens.2
2022 Semisupervised Semantic Segmentation of Remote Sensing Images With Consistency Self-Training
abstract
Semisupervised semantic segmentation is an effective way to reduce the expensive manual annotation cost and take advantage of the unlabeled data for remote sensing (RS) image interpretation. Recent related research has mainly adopted two strategies: self-training and consistency regularization. Self-training tries to acquire accurate pseudo-labels to explicitly expand the train set. However, the existing methods cannot accurately identify false pseudo-labels, suffering from their negative impact on model optimization. The consistency regularization constrains the model by producing consistent predictions robust to the perturbations introduced in the sample or feature domain but requires a sufficient number of training data. Therefore, we propose a strategy for the semisupervised semantic segmentation of the RS images. The proposed model in the generative adversarial network (GAN) framework is optimized by consistency self-training, learning the distributions of both labeled and unlabeled data. The discriminator is optimized by accurate pixel-level training labels instead of the image-level ones, thereby assessing the confidence for the prediction of each pixel, which is then used to reweight the loss of the unlabeled data in self-training. The generator is optimized with the consistency constraint with respect to all random perturbations on the unlabeled data, which increases the sample diversity and prompts the model to learn the underlying distribution of the unlabeled data. Experimental results on the the large-scale and densely annotated Instance Segmentation in Aerial Images Dataset (iSAID) datasets and the International Society for Photogrammetry and Remote Sensing (ISPRS) datasets show that our framework outperforms several state-of-the-art semisupervised semantic segmentation methods.
Jiahao Li 0003, Bin Sun 0001, Shutao Li 0001, Xudong Kang
IEEE Trans. Geosci. Remote. Sens.2
2022 Polygon Structure-Guided Hyperspectral Image Classification With Single Sample for Strong Geometric Characteristics Scenes
abstract
Combining spectral and spatial information can significantly improve the classification performance of hyperspectral image (HSI). Currently, a lot of spectral–spatial HSI classification methods have been proposed. However, the task of HSI classification has remained challenging since the number of training samples is limited in real scenarios. In this article, we propose a novel HSI classification framework with single sample, in which the spectral self-similarity and spatial polygon structure information are fully combined to improve the classification performance. On the one hand, spectral self-similarity is used to expand training samples, which makes it possible to obtain sufficient samples with minimal cost. On the other hand, polygonal partition is introduced to acquire the geometrical structure of land covers in man-made environments. Specifically, the edge information of geometric objects is captured by polygonal partition, which can be utilized to constrain the spatial range of sample expansion and optimize the classification results. Experimental results on three real HSIs illustrate that the proposed method performs very well under small training sample size even when the number of samples is single per class.
Shuo Zhang 0027, Xudong Kang, Puhong Duan, Bin Sun 0001, Shutao Li 0001
IEEE Trans. Geosci. Remote. Sens.4
2021 Skip-connected network with gram matrix for product image retrieval
Yi Li 0075, Bin Sun 0001, Li-Jun Liu
Neurocomputing3
2020 Vehicle Detection with Partial Anchors in Remote Sensing Images
abstract
Vehicle detection in remote sensing(RS) images has been an active topic with the development of computer vision in recent years. However, directly applying conventional horizontal anchor-based detection methods in oriented vehicle detection often acquires poor performance. Although rotated anchors have been used to tackle this problem, this design leads to heavy computational cost because of thousands of rotated anchors generated in each level feature map. In this paper, we propose to detect vehicles with partial anchors, which greatly accelerates detection process. The novel Partial Anchors based Detection Network(PADeN) filter out redundant anchors with semantic information. To boost the performance of PADeN, the centerness mask branch is added into the network. The results demonstrate that PADeN significantly outperforms previous approaches in vehicle detection and achieves the mAP of 76.9%.
Fuyan Ma, Bin Sun 0001, Shutao Li 0001, Jun Sun 0004
IGARSS2
2020 Hyperspectral Images Denoising via Nonconvex Regularized Low-Rank and Sparse Matrix Decomposition
abstract
Hyperspectral images (HSIs) are often degraded by a mixture of various types of noise during the imaging process, including Gaussian noise, impulse noise, and stripes. Such complex noise could plague the subsequent HSIs processing. Generally, most HSI denoising methods formulate sparsity optimization problems with convex norm constraints, which over-penalize large entries of vectors, and may result in a biased solution. In this paper, a nonconvex regularized low-rank and sparse matrix decomposition (NonRLRS) method is proposed for HSI denoising, which can simultaneously remove the Gaussian noise, impulse noise, dead lines, and stripes. The NonRLRS aims to decompose the degraded HSI, expressed in a matrix form, into low-rank and sparse components with a robust formulation. To enhance the sparsity in both the intrinsic low-rank structure and the sparse corruptions, a novel nonconvex regularizer named as normalized ε -penalty, is presented, which can adaptively shrink each entry. In addition, an effective algorithm based on the majorization minimization (MM) is developed to solve the resulting nonconvex optimization problem. Specifically, the MM algorithm first substitutes the nonconvex objective function with the surrogate upper-bound in each iteration, and then minimizes the constructed surrogate function, which enables the nonconvex problem to be solved in the framework of reweighted technique. Experimental results on both simulated and real data demonstrate the effectiveness of the proposed method.
Ting Xie 0003, Shutao Li 0001, Bin Sun 0001
IEEE Trans. Image Process.3
2019 Sea-Land Segmentation for Harbour Images with Superpixel CRF
abstract
Sea land segmentation is an important technique in many remote sensing applications, such as coastline surveillance and near-shore ship detection. High resolution satellite optical imaging is able to capture the details in the coastal regions, introducing intra-class variance and interferences like waves, shadow and forestry regions. To overcome such variance and interferences, the high resolution optical satellite image is first over-segmented into super-pixels, i.e., homogeneous regions. Then the conditional random fields (CRFs) are adopted to model the relations of the super-pixels. The optimal labels of all the superpixels are determined by performing the loopy belief propagation on the CRFs. The final segmentation result is obtained by refining the pixelwise probability conditioned on the estimated superpixel labels with edge preserving filtering. Experimental results on Google Earth data of different coastal cities show the effectiveness of the proposed sea land segmentation method.
Bin Sun 0001, Shutao Li 0001, Jie Xie 0002
IGARSS1
2019 Hyperspectral Compressive Sensing Via Spatial-Spectral Total Variation Regularized Low-Rank Tensor Decomposition
abstract
Hyperspectral compressive sensing (HCS) is considered for reconstructing the hyperspectral image (HSI) from a few random sampled measurements. HCS is crucial for the onboard imaging systems to cut down the acquisition time and data storage volume, and simultaneously maintain image quality. In this paper, a spatial-spectral total variation (SSTV) regularized low-rank tensor decomposition (LRTD) method is proposed for HCS. Specifically, for the HSI, the tensor nuclear norm based LRTD is utilized to characterize the global correlation among all bands, and an anisotropic SSTV regularization is explored to describe the local spatial smooth structure and spectral correlation of adjacent bands. In addition, an efficient algorithm based on the alternative direction multiplier method is developed to solve the resulting optimization problem. Experimental results demonstrate that the proposed method is superior to the existing state-of-the-art ones.
Ting Xie 0003, Shutao Li 0001, Bin Sun 0001
IGARSS3
2019 Graph-matching-based character recognition for Chinese seal images
Bin Sun 0001, Shaojun Hua, Shutao Li 0001, Jun Sun 0004
Sci. China Inf. Sci.1
2017 Random-Walker-Based Collaborative Learning for Hyperspectral Image Classification
abstract
Active learning (AL) and semisupervised learning (SSL) are both promising solutions to hyperspectral image classification. Given a few initial labeled samples, this work combines AL and SSL in a novel manner, aiming to obtain more manually labeled and pseudolabeled samples and use them together with the initial labeled samples to improve the classification performance. First, based on a comparison of the segmentation and spectral-spatial classification results obtained by random walker (RW) and extended RW (ERW) algorithms, the unlabeled samples are separated into two different sets, i.e., low- and high-confidence unlabeled data sets. For the high-confidence unlabeled data, pseudolabeling is performed, which can ensure the correctness and informativeness of the pseudolabeled samples. For the low-confidence unlabeled data, AL is used to select samples. In this way, the samples which are more effective for improvement of classification performance can be labeled in only a few iterations. Finally, with the learned training set and the original hyperspectral image as inputs, the ERW classifier is used to obtain the final classification result. Experiments performed on three real hyperspectral data sets show that the proposed method can achieve competitive classification accuracy even with a very limited number of manually labeled samples.
Bin Sun 0001, Xudong Kang, Shutao Li 0001, Jón Atli Benediktsson
IEEE Trans. Geosci. Remote. Sens.1
2016 Blind Bleed-Through Removal for Scanned Historical Document Image With Conditional Random Fields
abstract
Scanned images of historical documents often suffer from bleed-through, which refers to the ink on one side seeping through the paper and appearing on the other side. In this paper, a new conditional random field (CRF)-based method is proposed to remove the bleed-through from the scanned images of historical images. The proposed method only requires the scanned image of one side, referred as a blind method. In general, the scanned historical document image is composed of three components: foreground, bleed-through, and background. By assuming Gaussian distributions of the three components, the proposed method establishes conditional probability distribution (CPD) models of the three components first. The parameters of the component CPD models are estimated based on an initial segmentation of the input image. Then, CRFs are used to capture the relations between observed pixels in the scanned image and the corresponding labels as well as the spatial relation between the adjacent labels. The belief propagation algorithm is used to calculate the probabilities of different labels for each pixel. Once the labeling is completed by choosing the most possible label for each pixel, the bleed-through component is removed from the input historical image by a random-filling inpainting algorithm. Experimental results on the real data set show that the proposed method preserves the foreground component very well and removes the bleed-through effectively.
Bin Sun 0001, Shutao Li 0001, Xiao-Ping Zhang 0002, Jun Sun 0004
IEEE Trans. Image Process.1
2015 Blind bleed-through removal for scanned historical document images with conditional random fields
abstract
Due to the quality of paper and long-time preservation, the ink on one side of the historical documents often seeps through and appears on the other side. In this paper, a new blind ink bleed-through removal method is proposed to deal with the scanned historical document images. The scanned historical document image generally consists of three components: foreground, bleed-through and background. In the proposed method, conditional probability distribution (CPD) models of the three components are firstly established by statistics. Then, conditional random fields (CRFs) are used to model the observed scanned image and the corresponding labels. For each input scanned image, parameters of the component-wise CPD models are estimated and belief propagation is performed on the CRFs model to determine the most possible labels. Once the bleed-through component is found, an inpainting algorithm is proposed to remove the ink bleed-through from the input historical image. Experimental results show that the proposed method preserves the foreground component very well and removes the bleed-through effectively.
Bin Sun 0001, Shutao Li 0001, Jun Sun 0004
ICASSP1
2014 Scanned Image Descreening With Image Redundancy and Adaptive Filtering
abstract
Currently, most electrophotographic printers use halftoning technique to print continuous tone images, so scanned images obtained from such hard copies are usually corrupted by screen like artifacts. In this paper, a new model of scanned halftone image is proposed to consider both printing distortions and halftone patterns. Based on this model, an adaptive filtering based descreening method is proposed to recover high quality contone images from the scanned images. Image redundancy based denoising algorithm is first adopted to reduce printing noise and attenuate distortions. Then, screen frequency of the scanned image and local gradient features are used for adaptive filtering. Basic contone estimate is obtained by filtering the denoised scanned image with an anisotropic Gaussian kernel, whose parameters are automatically adjusted with the screen frequency and local gradient information. Finally, an edge-preserving filter is used to further enhance the sharpness of edges to recover a high quality contone image. Experiments on real scanned images demonstrate that the proposed method can recover high quality contone images from the scanned images. Compared with the state-of-the-art methods, the proposed method produces very sharp edges and much cleaner smooth regions.
Bin Sun 0001, Shutao Li 0001, Jun Sun 0004
IEEE Trans. Image Process.1