VLDB 2026 Research / reviewers in the wild / expert
Shuyong Gao
dblp:283/7097
· DBLP profile ↗
29ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0002-8992-0756ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 23 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Commonality in Few: Few-Shot Multimodal Anomaly Detection via Hypergraph-Enhanced MemoryabstractFew-shot multimodal industrial anomaly detection is a critical yet underexplored task, offering the ability to quickly adapt to complex industrial scenarios. In few-shot settings, insufficient training samples often fail to cover the diverse patterns present in test samples. This challenge can be mitigated by extracting structural commonality from a small number of training samples. In this paper, we propose a novel few-shot unsupervised multimodal industrial anomaly detection method based on structural commonality, CIF (Commonality In Few). To extract intra-class structural information, we employ hypergraphs, which are capable of modeling higher-order correlations, to capture the structural commonality within training samples, and use a memory bank to store this intra-class structural prior. Firstly, we design a semantic-aware hypergraph construction module tailored for single-semantic industrial images, from which we extract common structures to guide the construction of the memory bank. Secondly, we use a training-free hypergraph message passing module to update the visual features of test samples, reducing the distribution gap between test features and features in the memory bank. We further propose a hyperedge-guided memory search module, which utilizes structural information to assist the memory search process and reduce the false positive rate. Experimental results on the MVTec 3D-AD dataset and the Eyecandies dataset show that our method outperforms the state-of-the-art (SOTA) methods in few-shot settings. Yuxuan Lin 0001, Hanjing Yan, Xuan Tong, Yang Chang, Huanzhen Wang, Ziheng Zhou 0005, Shuyong Gao, Yan Wang 0068 |
AAAI | 7 |
| 2026 | Hi-EF: Benchmarking Emotion Forecasting in Human-interactionabstractAffective Forecasting is an psychology task that involves predicting an individual's future emotional responses, often hampered by reliance on external factors leading to inaccuracies, and typically remains at a qualitative analysis stage. To address these challenges, we narrows the scope of Affective Forecasting by introducing the concept of Human-interaction-based Emotion Forecasting (EF). This task is set within the context of a two-party interaction, positing that an individual's emotions are significantly influenced by their interaction partner's emotional expressions and informational cues. This dynamic provides a structured perspective for exploring the patterns of emotional change, thereby enhancing the feasibility of emotion forecasting. Haoran Wang 0006, Xinji Mai, Zeng Tao, Junxiong Lin, Xuan Tong, Ivy Pan, Shaoqi Yan, Yan Wang 0068, Shuyong Gao |
AAAI | 9 |
| 2026 | LVOS: A Benchmark for Large-Scale Long-Term Video Object SegmentationabstractVideo object segmentation (VOS) aims to distinguish and track target objects in a video. Despite the excellent performance achieved by off-the-shelf VOS models, part of the existing VOS benchmarks mainly focuses on short-term videos, where objects remain visible most of the time. However, these benchmarks may not fully capture challenges encountered in practical applications, and the absence of long-term datasets restricts further investigation of VOS in realistic scenarios. Thus, we propose a novel benchmark named LVOS, comprising 720 videos with 296,401 frames and 407,945 high-quality annotations. Videos in LVOS last 1.14 minutes on average. Each video includes various attributes, especially challenges encountered in the wild, such as long-term reappearing and cross-temporal similar objects. Compared to previous benchmarks, our LVOS better reflects VOS models' performance in real scenarios. Based on LVOS, we evaluate 15 existing VOS models under 3 different settings and conduct a comprehensive analysis. On LVOS, these models suffer a large performance drop, highlighting the challenge of achieving precise tracking and segmentation in real-world scenarios. Attribute-based analysis indicates that one of the significant factors contributing to accuracy decline is the increased video length, interacting with complex challenges such as long-term reappearance, cross-temporal confusion, and occlusion, which emphasize LVOS's crucial role. We hope our LVOS can advance development of VOS in real scenes. Lingyi Hong, Zhongying Liu, Chenzhi Tan, Yuang Feng, Xinyu Zhou 0006, Pinxue Guo, Zhaoyu Chen 0001, Shuyong Gao, Wei Zhang 0016 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | ClickVOS: Click Video Object SegmentationabstractVideo Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame during inference or lack the flexibility to specify arbitrary objects of interest. To address these limitations, we propose the setting named Click Video Object Segmentation (ClickVOS) which segments objects of interest across the whole video according to a single click per object in the first frame. And we provide the extended datasets DAVIS-P and YouTubeVOS-P that with point annotations to support this task. ClickVOS is of significant practical applications and research implications due to its only 1-2 seconds interaction time for indicating an object, comparing annotating the mask of an object needs several minutes. However, ClickVOS also presents increased challenges. To address this task, we propose an end-to-end baseline approach named called Attention Before Segmentation (ABS), motivated by the attention process of humans. ABS utilizes the given point in the first frame to perceive the target object through a concise yet effective segmentation attention. Although the initial object mask is possibly inaccurate, in our ABS, as the video goes on, the initially imprecise object mask can self-heal instead of deteriorating due to error accumulation, which is attributed to our designed improvement memory that continuously records stable global object memory and updates detailed dense memory. In addition, we conduct various baseline explorations utilizing off-the-shelf algorithms from related fields, which could provide insights for the further exploration of ClickVOS. The experimental results demonstrate the superiority of the proposed ABS approach. Extended datasets and codes will be available at https://github.com/PinxueGuo/ClickVOS. Pinxue Guo, Lingyi Hong, Xinyu Zhou 0006, Shuyong Gao, Wanyun Li, Zhaoyu Chen 0001, Xiaoqiang Li 0002, Wei Zhang 0016 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Scoring, Remember, and Reference: Catching Camouflaged Objects in VideosabstractVideo Camouflaged Object Detection (VCOD) aims to segment objects whose appearances closely resemble their surroundings, posing a challenging and emerging task. Existing vision models often struggle in such scenarios due to the indistinguishable appearance of camouflaged objects and the insufficient exploitation of dynamic information in videos. To address these challenges, we propose an end-to-end VCOD framework inspired by human memory-recognition, which leverages historical video information by integrating memory reference frames for camouflaged sequence processing. Specifically, we design a dual-purpose decoder that simultaneously generates predicted masks and scores, enabling reference frame selection based on scores while introducing auxiliary supervision to enhance feature extraction.Furthermore, this study introduces a novel reference-guided multilevel asymmetric attention mechanism, effectively integrating long-term reference information with short-term motion cues for comprehensive feature extraction. By combining these modules, we develop the Scoring, Remember, and Reference (SRR) framework, which efficiently extracts information to locate targets and employs memory guidance to improve subsequent processing. With its optimized module design and effective utilization of video data, our model achieves significant performance improvements, surpassing existing approaches by 10% on benchmark datasets while requiring fewer parameters (54M) and only a single pass through the video. The code will be made publicly available. Yu'ang Feng, Shuyong Gao, Fuzhen Yan, Yicheng Song, Lingyi Hong |
ICCV | 2 |
| 2025 | General Compression Framework for Efficient Transformer Object TrackingabstractPrevious works have attempted to improve tracking efficiency through lightweight architecture design or knowledge distillation from teacher models to compact student trackers. However, these solutions often sacrifice accuracy for speed to a great extent, and also have the problems of complex training process and structural limitations. Thus, we propose a general model compression framework for efficient transformer object tracking, named CompressTracker, to reduce model size while preserving tracking accuracy. Our approach features a novel stage division strategy that segments the transformer layers of the teacher model into distinct stages to break the limitation of model structure. Additionally, we also design a unique replacement training technique that randomly substitutes specific stages in the student model with those from the teacher model, as opposed to training the student model in isolation. Replacement training enhances the student model's ability to replicate the teacher model's behavior and simplifies the training process. To further forcing student model to emulate teacher model, we incorporate prediction guidance and stage-wise feature mimicking to provide additional supervision during the teacher model's compression process. CompressTracker is structurally agnostic, making it compatible with any transformer architecture. We conduct a series of experiment to verify the effectiveness and generalizability of our CompressTracker. Our CompressTracker-SUTrack, compressed from SUTrack, retains about 99 performance on LaSOT (72.2 AUC) while achieves 2.42x speed up. Code is available at https://github.com/LingyiHongfd/CompressTracker. Lingyi Hong, Xinyu Zhou 0006, Shilin Yan, Pinxue Guo, Kaixun Jiang, Zhaoyu Chen 0001, Shuyong Gao, Xingdong Sheng, Wei Zhang 0016, Hong Lu 0001 |
ICCV | 8 |
| 2025 | HSS-IAD: A Heterogeneous Same-Sort Industrial Anomaly Detection DatasetabstractMulti-class Unsupervised Anomaly Detection algorithms (MUAD) are receiving increasing attention due to their relatively low deployment costs and improved training efficiency. However, the real-world effectiveness of MUAD methods is questioned due to limitations in current Industrial Anomaly Detection (IAD) datasets. These datasets contain numerous classes that are unlikely to be produced by the same factory and fail to cover multiple structures or appearances. Additionally, the defects do not reflect real-world characteristics. Therefore, we introduce the Heterogeneous Same-Sort Industrial Anomaly Detection (HSS-IAD) dataset, which contains 8,580 images of metallic-like industrial parts and precise anomaly annotations. These parts exhibit variations in structure and appearance, with subtle defects that closely resemble the base materials. We also provide foreground images for synthetic anomaly generation. Finally, we evaluate popular IAD methods on this dataset under multi-class and class-separated settings, demonstrating its potential to bridge the gap between existing datasets and real factory conditions. The dataset is available at https://github.com/Qiqigeww/HSS-IAD-Dataset. Qishan Wang 0002, Shuyong Gao, Jiawen Yu, Xuan Tong |
ICME | 2 |
| 2025 | Progressive Representation Learning for Weakly-Supervised Camouflaged Object Detection
Shuyong Gao, Yu'ang Feng, Chunyuan Chen, Xujun Wei, Yan Wang 0068 |
ACM Multimedia | 1 |
| 2025 | Observe finer to select better: Learning key frame extraction via semantic coherence for dynamic facial expression recognition in the wild
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Zeng Tao, Wei Song 0007, Qing Zhao 0007, Boyang Wang 0003, Haoran Wang 0006, Shuyong Gao |
Inf. Sci. | 9 |
| 2024 | Attention in Attention for PET-CT Modality Consensus Lung Tumor SegmentationabstractCombination of multi-modal PET-CT imaging for lung tumor segmentation is significant for clinical treatment. Existing methods have not fully considered the impact of noise in PET-CT on the multi-modal interaction. To address this, we propose a novel Attention in Attention Network (AiANet). AiANet can mutually learn multi-modal characteristics for segmentation through its cross-learning modules. Within the cross-learning module, we introduce two nested-attention blocks, namely Attention in Self-Attention (AiSA) and Attention in Cross-Attention (AiCA), for multi-scale feature enhancement and multi-modal feature interaction. Importantly, since traditional attention weights calculated solely or unilaterally based on PET or CT can be vulnerable to the inevitable noisy information, we embed a novel Attention in Attention (AiA) module into AiSA and AiCA. The AiA module can seek cross-modal consensus for attention weights to alleviate their noise. Experimental results on clinical PET-CT data of lung cancer demonstrate the superiority of our method. Yuzhou Zhao, Xinyu Zhou 0006, Haijing Guo, Shaoli Song, Shuyong Gao |
ICME | 7 |
| 2024 | Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution
Junxiong Lin, Zen Tao, Xuan Tong, Xinji Mai, Haoran Wang 0006, Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Jiawen Yu, Yuxuan Lin 0001, Shaoqi Yan, Shuyong Gao |
ACM Multimedia | 12 |
| 2024 | All rivers run into the sea: Unified Modality Brain-Inspired Emotional Central MechanismabstractIn the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emotions. Inspired by the process through which the human brain handles emotions and the theory of cross-modal plasticity, we propose UMBEnet, a brain-like unified modal affective processing network. The primary design of UMBEnet includes a Dual-Stream (DS) structure that fuses inherent prompts with a Prompt Pool and a Sparse Feature Fusion (SFF) module. The design of the Prompt Pool is aimed at integrating information from different modalities, while inherent prompts are intended to enhance the system's predictive guidance capabilities and effectively manage knowledge related to emotion classification. Moreover, considering the sparsity of effective information across different modalities, the SSF module aims to make full use of all available sensory data through the sparse integration of modality fusion prompts and inherent prompts, maintaining high adaptability and sensitivity to complex emotional states. Extensive experiments on the largest benchmark datasets in the Dynamic Facial Expression Recognition (DFER) field, including DFEW, FERV39k, and MAFW, have proven that UMBEnet consistently outperforms the current state-of-the-art methods. Notably, in scenarios of Modality Missingness and multimodal contexts, UMBEnet significantly surpasses the leading current methods, demonstrating outstanding performance and adaptability in tasks that involve complex emotional understanding with rich multimodal information. Code can be obtained at https://github.com/Xinji-Mai/UMBEnet. Xinji Mai, Junxiong Lin, Haoran Wang 0006, Zeng Tao, Yan Wang 0068, Shaoqi Yan, Xuan Tong, Jiawen Yu, Boyang Wang 0003, Ziheng Zhou 0005, Qing Zhao 0007, Shuyong Gao |
ACM Multimedia | 12 |
| 2024 | Empower smart cities with sampling-wise dynamic facial expression recognition via frame-sequence contrastive learning
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Qing Zhao 0007, Wei Song 0007, Zeng Tao, Haoran Wang 0006, Shuyong Gao |
Comput. Commun. | 9 |
| 2024 | MGR3Net: Multigranularity Region Relation Representation Network for Facial Expression Recognition in Affective RobotsabstractAutomatic facial expression recognition (FER) based on face images is essential for affective robots, which are designed for interactive companions and intelligent healthcare. Although existing DL-based FERs have made significant progress, an accurate FER model in robots is challenging due to the subtle differences in facial expressions across various scenarios. To address this issue, we propose a multigranularity region relation representation network (MGR3Net) to improve the robustness and generalization of FER via attention-guided global-local fusion. The MGR3Net is composed of three modules: multigranularity attention (MGA), holistic-regional feature extractor (HRFE), and hybrid feature fusion. In the MGA module, we first process each holistic cropped face image into three granularity of face regions from coarse to fine, which are four region-cropped faces,$2^{2}$face partitions, and$4^{2}$face partitions. Then, we propose the region attention relation cell to model the relationship between each region and the aggregated representation while preserving the spatial information of the local features. In the HRFE module, we align multigranularity features from the coarse space to the finer space and extract one holistic embedding and multiple region embeddings for each granularity. Finally, we use a hybrid-level fusion strategy to combine global-local features from the three granularities for final classification. Extensive experiments demonstrate that the MGR3Net outperforms the state-of-the-art methods evaluated on the in-the-lab datasets, in-the-wild datasets, and occlusion/pose-based sets. Yan Wang 0068, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jing Liu 0050, Dingkang Yang, Shuyong Gao |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | MSC-AD: A Multiscene Unsupervised Anomaly Detection Dataset for Small Defect Detection of Casting SurfaceabstractIntelligent detection of product surface defects in the industrial scene is the key to ensuring product quality. On general benchmarks, current unsupervised anomaly detection techniques have achieved significant success. When used in complex industrial environments (e.g., large industrial components with small defects), the model needs to be able to adapt to different imaging scenarios (e.g., illumination and resolution) and accurately detect and localize anomalies, but its performance is still far from satisfactory. Besides, the complex and unstable optical lighting environment for collecting such data poses major challenges in establishing unified benchmarks for optical lighting and imaging resolution in defect detection. To fill this gap, we build a standard imaging system-based multiscene unsupervised anomaly detection dataset, coined as MSC-AD. In particular, it provides 12 imaging scenes, i.e., a cross combination of low-to-high three illuminations and 150 × 150 to 600 × 600 four resolutions, in which six types of large casting surfaces with different structures include five kinds of small defects with sample-level and pixel-level precise ground truth. We systematically investigate representative baseline methods and empirical analysis on this dataset to obtain a number of interesting findings, e.g., how to detach from distinctly different imaging scenes, and how to distinguish between subtly normal–anomaly classes. To the best of our knowledge, MSC-AD is the first multi-illumination, multiresolution, multisurface, and multidefect dataset built in a standard imaging system. Qing Zhao 0007, Yan Wang 0068, Boyang Wang 0003, Junxiong Lin, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jiawen Yu, Shuyong Gao |
IEEE Trans. Ind. Informatics | 9 |
| 2023 | Plug-and-Play Feature Generation for Few-Shot Medical Image ClassificationabstractFew-shot learning (FSL) presents immense potential in enhancing model generalization and practicality for medical image classification with limited training data; however, it still faces the challenge of severe overfitting in classifier training due to distribution bias caused by the scarce training samples. To address the issue, we propose MedMFG, a flexible and lightweight plug-and-play method designed to generate sufficient class-distinctive features from limited samples. Specifically, MedMFG first re-represents the limited prototypes to assign higher weights for more important information features. Then, the prototypes are variationally generated into abundant effective features. Finally, the generated features and prototypes are together to train a more generalized classifier. Experiments demonstrate that MedMFG outperforms the previous state-of-the-art methods on cross-domain benchmarks involving the transition from natural images to medical images, as well as medical images with different lesions. Notably, our method achieves over 10% performance improvement compared to several baselines. Fusion experiments further validate the adaptability of MedMFG, as it seamlessly integrates into various backbones and baselines, consistently yielding improvements of over 2.9% across all results. Huifang Du, Xing Jia, Shuyong Gao, Yan Teng 0002, Haofen Wang |
BIBM | 4 |
| 2023 | TINYCOD: Tiny and Effective Model for Camouflaged Object DetectionabstractThis paper introduces an effective and tiny model for real-time Camouflaged Object Detection (COD) named Tiny-COD. It achieves high performance with very low costs (Parameters < 5M, FLOPs < 1.5G), which can be applied on mobile devices. Specifically, we introduce a simple but effective Adjacent Scale Features Fusion module (ASFF), which can significantly enhance the representation ability of features from a lightweight backbone. Besides, as the edge areas of the camouflaged object often blend into the background, we carefully design an Edge Area Focus module (EAF) to solve this problem. Experimental results on COD datasets prove that the proposed method achieves state-of-the-art performance compared with other methods. Haozhe Xing, Shuyong Gao, Hao Tang 0005, Tsui Qin Mok, Yanlan Kang |
ICASSP | 2 |
| 2023 | SimulFlow: Simultaneously Extracting Feature and Identifying Target for Unsupervised Video Object SegmentationabstractUnsupervised video object segmentation (UVOS) aims at detecting the primary objects in a given video sequence without any human interposing. Most existing methods rely on two-stream architectures that separately encode the appearance and motion information before fusing them to identify the target and generate object masks. However, this pipeline is computationally expensive and can lead to suboptimal performance due to the difficulty of fusing the two modalities properly. In this paper, we propose a novel UVOS model called SimulFlow that simultaneously performs feature extraction and target identification, enabling efficient and effective unsupervised video object segmentation. Concretely, we design a novel SimulFlow Attention mechanism to bridege the image and motion by utilizing the flexibility of attention operation, where coarse masks predicted from fused feature at each stage are used to constrain the attention operation within the mask area and exclude the impact of noise. Because of the bidirectional information flow between visual and optical flow features in SimulFlow Attention, no extra hand-designed fusing module is required and we only adopt a light decoder to obtain the final prediction. We evaluate our method on several benchmark datasets and achieve state-of-the-art results. Our proposed approach not only outperforms existing methods but also addresses the computational complexity and fusion difficulties caused by two-stream architectures. Our models achieve 87.4 ℐ&F on DAVIS-16 with the highest speed (63.7 FPS on a 3090) and the lowest parameters (13.7 M). Our SimulFlow also obtains competitive results on video salient object detection datasets. Lingyi Hong, Wei Zhang 0016, Shuyong Gao, Hong Lu 0001 |
ACM Multimedia | 3 |
| 2023 | Towards End-to-End Unsupervised Saliency Detection with Self-Supervised Top-Down ContextabstractUnsupervised salient object detection aims to detect salient objects without using supervision signals eliminating the tedious task of manually labeling salient objects. To improve training efficiency, end-to-end methods for USOD have been proposed as a promising alternative. However, current solutions rely heavily on noisy handcraft labels and fail to mine rich semantic information from deep features. In this paper, we propose a self-supervised end-to-end salient object detection framework via top-down context. Specifically, motivated by contrastive learning, we exploit the self-localization from the deepest feature to construct the location maps which are then leveraged to learn the most instructive segmentation guidance. Further considering the lack of detailed information in deepest features, we exploit the detail-boosting refiner module to enrich the location labels with details. Moreover, we observe that due to lack of supervision, current unsupervised saliency models tend to detect non-salient objects that are salient in some other samples of corresponding scenarios. To address this widespread issue, we design a novel Unsupervised Non-Salient Suppression (UNSS) method developing the ability to ignore non-salient objects. Extensive experiments on benchmark datasets demonstrate that our method achieves leading performance among the recent end-to-end methods and most of the multi-stage solutions. The code is available. Yicheng Song, Shuyong Gao, Haozhe Xing, Yiting Cheng 0001, Yan Wang 0068 |
ACM Multimedia | 2 |
| 2023 | Freq-HD: An Interpretable Frequency-based High-Dynamics Affective Clip Selection Method for in-the-Wild Facial Expression Recognition in VideosabstractThe in-the-wild dynamic facial expression recognition (DFER) has been challenging due to several high-dynamics factors such as limited dynamic expression-related frames and variable non-expression noise in facial expression sequences. To provide more expression-related clips for DFER models, we propose a novel and interpretable frequency-based method (Freq-HD) for high-dynamics affective clip selection. It can select clips containing pure expression changes from sequences and aid different DFER network structures in recognizing in-the-wild dynamic facial expressions more accurately and efficiently. We first design a novel spatial-temporal frequency analysis (STFA) module to compute the dynamics values of each clip by using sliding windows and spatial-temporal frequency analysis. Moreover, we propose a multi-band complementary selection (MBC) module to amend the inappropriate reaction of the dynamics values of different spatial frequency bands in STFA when expression-irrelevant noise occurs. Specifically, the MBC uses an ingenious mapping method to generate the inhibitory factors to complement and separate the dynamics of expressions and non-expressions in different frequency bands. The Freq-HD can select the most expression-correlated clips and the consisting frames, which could be incorporated into any existing DFER models. We extensively evaluate the Freq-HD on two in-the-wild datasets and four DFER baselines, showing that our method significantly improves the subsequent network performance while using fewer input frames and reducing computation cost. More ablation studies and visualization analysis provide further empirical evidence of the effectiveness of our method. Zeng Tao, Yan Wang 0068, Zhaoyu Chen 0001, Boyang Wang 0003, Shaoqi Yan, Kaixun Jiang, Shuyong Gao |
ACM Multimedia | 7 |
| 2023 | A Capture to Registration Framework for Realistic Image Super-Resolution in the Industry EnvironmentabstractThe acquisition and processing of visual data in industrial environments are of paramount importance. High-resolution (HR) images offer superior clarity and richer textural detail compared to low-resolution (LR) images. On the one hand, owing to the incorporation of richer information, HR images demonstrate substantially enhanced performance compared to LR images in downstream applications, such as anomaly detection. On the other hand, they provide valuable insights to designers and quality inspectors who require a detailed understanding of the images. Currently, the majority of research on super-resolution focuses on natural scenes such as cities and fields, however, the development of datasets for industrial scenes is still in its infancy. To address the image distortion in building realistic LR-HR image pairs in the industry environment, we design a capture to registration framework. It consists of the standard imaging system, physical calibration of the imaging system, as well as the rigid to elastic registration of the LR-HR image pairs. Thus, we build the first realistic industrial sence super-resolution dataset (IndSR), comprises of 50 sets of calibrated images with three scale factors and five typical defects. To benchmark IndSR, we employ quantitative, qualitative, and task-oriented studies to evaluate the representative super-resolution and anomaly detection methods. Besides, we systematically investigate and discuss the performances and results of the existing SISR methods to advance research in the field of super-resolution in industry environment. The IndSR dataset can be available from https://byw4ng.github.io/IndSR/. Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Junxiong Lin, Zeng Tao, Pinxue Guo, Zhaoyu Chen 0001, Kaixun Jiang, Shaoqi Yan, Shuyong Gao |
ACM Multimedia | 10 |
| 2023 | Go Closer to See Better: Camouflaged Object Detection via Object Area Amplification and Figure-Ground ConversionabstractCamouflaged Object Detection (COD) aims to detect objects well hidden in the environment. The main challenges of COD come from the high degree of texture and color overlapping between the objects and their surroundings. Inspired by that humans tend to go closer to the object and magnify it to recognize ambiguous objects more clearly, we propose a novel three-stage architecture called Search-Amplify-Recognize and design a network SARNet to address the challenges. Specifically, In the Search part, we utilize an attention-based backbone to locate the object. In the Amplify part, to obtain rich searched features and fine segmentation, we design Object Area Amplification modules (OAA) to perform cross-level and adjacent-level feature fusion and amplifying operations on feature maps. Besides, the OAA can be regarded as a simple and effective plug-in module to integrate and amplify the feature maps. The main components of the Recognize part are the Figure-Ground Conversion modules (FGC). The FGC modules alternately pay attention to the foreground and background to precisely separate the highly similar foreground and background. Extensive experiments on benchmark datasets show that our model outperforms other SOTA methods not only on COD tasks but also in COD downstream tasks, such as polyp segmentation and video camouflaged object detection. Source codes will be available athttps://github.com/Haozhe-Xing/SARNet. Haozhe Xing, Shuyong Gao, Yan Wang 0068, Xujun Wei, Hao Tang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Weakly-Supervised Salient Object Detection Using Point SupervisonabstractCurrent state-of-the-art saliency detection models rely heavily on large datasets of accurate pixel-wise annotations, but manually labeling pixels is time-consuming and labor-intensive. There are some weakly supervised methods developed for alleviating the problem, such as image label, bounding box label, and scribble label, while point label still has not been explored in this field. In this paper, we propose a novel weakly-supervised salient object detection method using point supervision. To infer the saliency map, we first design an adaptive masked flood filling algorithm to generate pseudo labels. Then we develop a transformer-based point-supervised saliency detection model to produce the first round of saliency maps. However, due to the sparseness of the label, the weakly supervised model tends to degenerate into a general foreground detection model. To address this issue, we propose a Non-Salient Suppression (NSS) method to optimize the erroneous saliency maps generated in the first round and leverage them for the second round of training. Moreover, we build a new point-supervised dataset (P-DUTS) by relabeling the DUTS dataset. In P-DUTS, there is only one labeled point for each salient object. Comprehensive experiments on five largest benchmark datasets demonstrate our method outperforms the previous state-of-the-art methods trained with the stronger supervision and even surpass several fully supervised state-of-the-art models. The code is available at: https://github.com/shuyonggao/PSOD. Shuyong Gao, Wei Zhang 0016, Yan Wang 0068, Yangji He |
AAAI | 1 |
| 2022 | FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in VideosabstractCurrent benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in real-world application-oriented scenes. For example, the “Happy” expression with high intensity in Talk-Show is more discriminating than the same expression with low intensity in Official-Event. To fill this gap, we build a large-scale multi-scene dataset, coined as FERV39k. We analyze the important ingredients of constructing such a novel dataset in three aspects: (1) multi-scene hierarchy and expression class, (2) generation of candidate video clips, (3) trusted manual labelling process. Based on these guidelines, we select 4 scenarios subdivided into 22 scenes, annotate 86k samples automatically obtained from 4k videos based on the well-designed workflow, and finally build 38,935 video clips labeled with 7 classic expressions. Experiment benchmarks on four kinds of baseline frame-works were also provided and further analysis on their performance across different scenes and some challenges for future research were given. Besides, we systematically investigate key components of DFER by ablation studies. The baseline framework and our project are available on https://github.com/wangyanckxx/FERV39k. Yan Wang 0068, Yixuan Sun, Zhongying Liu, Shuyong Gao, Wei Zhang 0016, Weifeng Ge |
CVPR | 5 |
| 2022 | Position-aware Joint Entity and Relation Extraction with Attention MechanismabstractNamed entity recognition and relation extraction are two important core subtasks of information extraction, which aim to identify named entities and extract relations between them. In recent years, span representation methods have received a lot of attention and are widely used to extract entities and corresponding relations from plain texts. Most recent works focus on how to obtain better span representations from pre-trained encoders, but ignore the negative impact of a large number of span candidates on slowing down the model performance. In our work, we propose a joint entity and relation extraction model with an attention mechanism and position-attentive markers. The attention score of each candidate span is calculated, and most of the candidate spans with low attention scores are pruned before being fed into the span classifier, thus achieving the goal of removing the most irrelevant spans. At the same time, in order to explore whether the position information can improve the performance of the model, we add position-attentive markers to the model. The experimental results show that our model is effective. With the same pre-trained encoder, our model achieves the new state-of-the-art on standard benchmarks (ACE05, CoNLL04 and SciERC), obtaining a 4.7%-17.8% absolute improvement in relation F1. Shuyong Gao, Haofen Wang |
IJCAI | 2 |
| 2022 | Weakly Supervised Video Salient Object Detection via Point SupervisionabstractFully supervised video salient object detection models have achieved excellent performance, yet obtaining pixel-by-pixel annotated datasets is laborious. Several works attempt to use scribble annotations to mitigate this problem, but point supervision as a more labor-saving annotation method (even the most labor-saving method among manual annotation methods for dense prediction), has not been explored. In this paper, we propose a strong baseline model based on point supervision. To infer saliency maps with temporal information, we mine inter-frame complementary information from short-term and long-term perspectives, respectively. Specifically, we propose a hybrid token attention module, which mixes optical flow and image information from orthogonal directions, adaptively highlighting critical optical flow information (channel dimension) and critical token information (spatial dimension). To exploit long-term cues, we develop the Long-term Cross-Frame Attention module (LCFA), which assists the current frame in inferring salient objects based on multi-frame tokens. Furthermore, we label two point-supervised datasets, P-DAVIS and P-DAVSOD, by relabeling the DAVIS and the DAVSOD dataset. Experiments on the six benchmark datasets illustrate our method outperforms the previous state-of-the-art weakly supervised methods and even is comparable with some fully supervised approaches. Our source code and datasets are available at: https://github.com/shuyonggao/PVSOD. Shuyong Gao, Haozhe Xing, Wei Zhang 0016, Yan Wang 0068 |
ACM Multimedia | 1 |
| 2022 | DPCNet: Dual Path Multi-Excitation Collaborative Network for Facial Expression Representation Learning in VideosabstractCurrent works of facial expression learning in video consume significant computational resources to learn spatial channel feature representations and temporal relationships. To mitigate this issue, we propose a Dual Path multi-excitation Collaborative Network (DPCNet) to learn the critical information for facial expression representation from fewer keyframes in videos. Specifically, the DPCNet learns the important regions and keyframes from a tuple of four view-grouped frames by multi-excitation modules and produces dual-path representations of one video with consistency under two regularization strategies. A spatial-frame excitation module and a channel-temporal aggregation module are introduced consecutively to learn spatial-frame representation and generate complementary channel-temporal aggregation, respectively. Moreover, we design a multi-frame regularization loss to enforce the representation of multiple frames in the dual view to be semantically coherent. To obtain consistent prediction probabilities from the dual path, we further propose a dual path regularization loss, aiming to minimize the divergence between the distributions of two-path embeddings. Extensive experiments and ablation studies show that the DPCNet can significantly improve the performance of video-based FER and achieve state-of-the-art results on the large-scale DFEW dataset. Yan Wang 0068, Yixuan Sun, Wei Song 0007, Shuyong Gao, Zhaoyu Chen 0001, Weifeng Ge |
ACM Multimedia | 4 |
| 2021 | Dual-Stream Network Based On Global Guidance for Salient Object DetectionabstractHigh-level features can help low-level features eliminate semantic ambiguity, which is crucial for obtaining the precise salient object. Some methods use high-level features to provide global guidance for some layers of the network. However, there remain several problems: (1) the global guidance has not been fully mined, which leads to its limited capacity; (2) the semantic gap between global guidance and low-level features is ignored, and simple merging methods will cause feature aliasing. To remedy the problems, we propose a dual-stream network based on global guidance with two plug-ins, global attention based multi-scale high-level feature extraction module (GAMS) to mine global guidance and scale adaptive global guidance module (SAGG) to seamlessly integrate the global guidance into each decoding layer. Comprehensive experiments on the five largest benchmark datasets demonstrate our method outperforms previous state-of-the-art methods by a large margin. Code is available at https://github.com/shuyonggao/DSGGN. Shuyong Gao, Wei Zhang 0016, Zhongwei Ji |
ICASSP | 1 |
| 2021 | Global Cognition and Local Perception Network for Blind Image Deblurring
Chuanfa Zhang, Wei Zhang 0016, Yiting Cheng 0001, Shuyong Gao |
MMM (1) | 5 |