EDBT 2026 Demo / reviewers in the wild / expert
Peng Chen 0008
dblp:27/7017-8
· DBLP profile ↗
78ranked-venue papers
2as first author
69since 2021 · last 2026
0000-0001-6122-0574ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 34 since 2021Artificial intelligence and machine learning · 24 · 19 since 2021Security and privacy · 11 · 11 since 2021Systems, architecture and hardware · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Scene-Aware Meta-learning Framework for Robust Photovoltaic Power Forecasting
Yihan Yu, Yuanjie Dang, Peng Chen 0008, Yilong Zhang 0001, Ronghua Liang |
ICIC (16) | 4 |
| 2026 | DDPT: Enhancing complex reasoning in large language models via distillation and dynamic prompt tuning
Ge Teng, Chen Shen 0003, Wenxiao Wang 0001, Sinan Fan, Liang Xie 0003, Xiang Tian 0002, Peng Chen 0008, Yaowu Chen, Jieping Ye |
Neurocomputing | 7 |
| 2026 | GCFL: Gray-Modality Conversion and Feature Learning for Visible-Infrared Person Re-Identification
Ruohong Huan, Mingzhen Wu, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2026 | A Dual Asynchronous-Synchronous Relation Graph Method for Sensor-Based Group Activity Recognition
Ai Bo 0001, Ruohong Huan, Peng Chen 0008, Ronghua Liang |
IEEE Internet Things J. | 4 |
| 2026 | FreTransLS: Frequency Transformer based large-scale group activity recognition model for sensor data
Ruohong Huan, Meijiao Cao, Yantong Zhou, Peng Chen 0008, Guodao Sun, Ronghua Liang |
Pervasive Mob. Comput. | 5 |
| 2026 | Learning Decoupled Features With Perceptual Distillation for Blind Image Quality AssessmentabstractExisting Blind Image Quality Assessment (BIQA) approaches typically employ subjective scores as optimization targets to train the model, aiming for results consistent with human judgments. Such judgments are derived from a comprehensive analysis of complex distortions and diverse semantics from images, whereas subjective scores represent the overall quality. This poses a significant challenge for a single model to learn diverse perceptual cues under weak supervision. To address this, we propose a Decoupled Feature Learning (DFL) framework that learns compact global content-aware and local distortion-aware features in a disentangled modeling for BIQA. Our key insight is to leverage global-local input pairs to decompose content-aware and distortion-aware cues entangled in distorted images, and aggregate decoupled perceptual features into a single network. We design a perceptual knowledge distillation strategy that progressively guides the student from fragmented representations to build local-to-global correspondences by distilling self-supervised semantic knowledge, while incorporating the Just-Noticeable-Difference (JND) model to highlight the transfer of perceptually sensitive content features. Finally, we introduce a local distortion-guided attention module to model synergistic effects of different perceptual features from the student for quality evaluation. Extensive experiments on eight benchmark datasets demonstrate the superior performance of the proposed model over the state-of-the-arts. In addition, the DFL framework is flexibly used to improve the perception ability of other Transformer variants. The code is released at https://github.com/JianjunXiang/DFT. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Weisi Lin |
IEEE Trans. Image Process. | 3 |
| 2026 | Performance Optimization Strategies for Data Transmission From Edge to Cloud: A ReviewabstractWith the rapid proliferation of IoT devices, the volume of generated data is growing at an unprecedented pace. Due to the limited resources of edge devices, a significant portion of this data must be transmitted to the cloud for in-depth processing, large-scale analysis, long-term storage, and archival purposes. Consequently, the performance has become a critical concern. While identifying prevailing challenges and research gaps in this domain requires a systematic review, such efforts remain largely absent from existing survey literature. This article addresses this gap by offering a structured review of recent optimization approaches. It begins by categorizing the literature into three main strategies: lossless transmission, lossy transmission, and hybrid approaches. In the context of lossless transmission, we analyze techniques such as data compression algorithms and incremental versus full synchronization mechanisms. For lossy strategies, we analyze approaches including lossy compression and predictive methods. In addition, we investigate hybrid strategies that integrate both lossless and lossy techniques to leverage their complementary advantages. Finally, we discuss the limitations of existing studies and highlight promising directions for future research in optimizing edge-to-cloud data transmission. Jian Liu 0053, Yangyang Lin, Ziguang Fu, Gexi Lin, Guodao Sun, Zhu Xiao, Yilong Zhang 0001, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Knowl. Data Eng. | 9 |
| 2025 | DiffGen: Optimizing I/O Trace Generation with Differentiated Modeling Techniques
Jian Liu 0053, Zhiyang Feng, Ziguang Fu, Guodao Sun, Yilong Zhang 0001, Nan Gao 0001, Ronghua Liang, Peng Chen 0008 |
ICA3PP (5) | 8 |
| 2025 | Improving Sidescan Sonar Performance Using Array Upsampling Beamforming Synthetic ApertureabstractSidescan sonar technology plays an important role in seabed topography detection and mapping, but the performance of mainstream matched filter technology in long-distance image and towing speed limits its application scenarios. In addition, due to the slow propagation speed of sound waves in water, the synthetic aperture method in radar is difficult to adapt directly to sonar. To address these problems, a synthetic aperture algorithm framework for sidescan sonar with array upsampling beamforming is proposed (AUB-SAS). Firstly, the array element spacing is reduced by upsampling the array element to facilitate beamforming, and the long-distance information acquisition capability is improved by beam focusing. Secondly, the calculated multi-beam is used for synthetic aperture processing to make up for the low pulse repetition rate caused by the slow sound speed, thereby increasing the towing speed. In addition, the short-time Fourier transform with high resolution in the lateral direction and the motion compensation in the azimuth direction are derived, which further improve the imaging quality. The imaging results of lake test data prove the effectiveness of the proposed algorithm, achieving a lateral detection distance of 250 meters while the towing speed reaches 9 knots, and the imaging quality is significantly better than the traditional method. Weibo Mao, Peng Chen 0008, Shihui Liang, Ronghua Liang |
ICASSP | 3 |
| 2025 | ZJUT-PAD : A New Fingerprint Presentation Attack Detection Database based on Optical Coherence TomographyabstractFingerprints, due to their uniqueness and stability, have become the most widely used biometric feature. Automated Fingerprint Recognition Systems (AFRS) have been applied in various scenarios for identity verification and access control. However, these systems have long faced serious threats from presentation attacks (PA), posing potential risks of privacy breaches and financial loss. Optical Coherence Tomography (OCT), as a non-invasive imaging technology, aligns with the demand for more secure and stable fingerprint recognition methods. By integrating OCT into AFRS, fingerprint recognition can be extended from traditional 2D surface fingerprints to OCT fingerprints containing 3D fingertip information. The rich and difficult-to-replicate structural details encoded in OCT fingerprints offer a highly promising solution for fingerprint presentation attack detection (PAD). Currently, publicly available OCT fingerprint datasets remain limited, and those specifically designed for PAD research are even rarer. This scarcity of data has significantly constrained research in OCT-based fingerprint PAD. To address this gap, a dedicated database for OCT fingerprint anti-spoofing, referred to as the ZJUT-PAD, has been designed and released by our research team. This database consists of 175 Presentation Attack Instruments (PAIs) made from 10 different materials, covering 35 distinct types. Each PAI was captured five times using two different OCT devices, resulting in a total of 1,750 PA instances. This database serves as a critical evaluation platform for OCT fingerprint PAD research, enabling a comprehensive assessment of performance, generalization capability, and cross-device robustness of PAD methods. Haixia Wang 0002, Haohao Sun, Yilong Zhang 0001, Peng Chen 0008, Zilan Pan |
IJCB | 6 |
| 2025 | Spatial Continuity-Aware OCT Fingerprint Reconstruction Using Iterative Feature EnhancementabstractOptical coherence tomography (OCT) is a non-invasive imaging technique capable of acquiring depth information up to 1-3mm beneath the skin surface, including the stratum corneum and viable epidermis regions. This technique allows for the reconstruction of internal and external fingerprint images from grayscale data. However, existing fingerprint extraction methods heavily rely on contour features and current 2D approaches overlook the spatial continuity of biometric features in OCT images. Therefore, this paper proposes a novel iterative algorithm for internal and external fingerprint extraction from OCT images. This algorithm incorporates the spatial continuity of OCT slice images and an iterative feature enhancement module during the prediction phase to improve segmentation continuity. Additionally, a soft label technique is employed to reduce contour dependence and mitigate interference from noise and anomaly interference. Qualitative and quantitative experiments demonstrate significant improvements in segmentation accuracy with higher fingerprint quality, validating the effectiveness of the proposed approach. Yilong Zhang 0001, Xuanbing Chen, Shengming Zhu, Haohao Sun, Haixia Wang 0002, Jian Liu 0053, Yuanjie Dang, Ronghua Liang, Peng Chen 0008 |
IJCB | 9 |
| 2025 | Dual Teacher with Dempster-Shafer Guidance for Decision Making in Semi-Supervised Small Object DetectionabstractSmall-scale object detection remains a major challenge in semi-supervised object detection (SSOD), particularly in medical image analysis. Conventional teacher models often struggle to accurately capture the features of low-contrast small lesions, leading to noisy pseudo-labels in both localization and classification, which introduces severe uncertainty and degrades detection performance. To address this issue, we propose Dual Teacher, a novel multimodal semi-supervised detection framework designed to enhance pseudo-label reliability and improve small-scale lesion detection. Specifically, we introduce two complementary teacher models: Hybrid-Scale Teacher, which exploits downsampled views to strengthen multi-scale feature learning, and Entropy-Based Multi-Modal Teacher, which leverages entropy maps to refine the quality of small-scale pseudo-labels. To effectively fuse predictions from both teachers and resolve conflicts, we propose a Dempster-Shafer-based Dual-Teacher pseudo-label fusion strategy that explicitly models uncertainty and optimizes classification confidence. Additionally, we introduce a class-adaptive threshold mechanism that dynamically adjusts pseudo-label selection based on dual-teacher predictions, further boosting the recall of small-scale lesions. Extensive experiments on the Dental Disease Dataset, ChestX-Det and M3FD demonstrate that our method consistently surpasses state-of-the-art SSOD approaches. Code is available at: https://github.com/z316910/Dual-Teacher.git. Nan Gao 0001, Junchao Zhu, Yilong Zhang 0001, Ronghua Liang, Guodao Sun, Peng Chen 0008 |
ACM Multimedia | 6 |
| 2025 | Multi-granularity semantic relational mapping for image caption
Nan Gao 0001, Renyuan Yao, Peng Chen 0008, Ronghua Liang, Guodao Sun, Jijun Tang |
Expert Syst. Appl. | 3 |
| 2025 | TWDT: Training-free word-level controllable diffusion model for text generation
Nan Gao 0001, Yangjie Lu, Peng Chen 0008, Guodao Sun, Ronghua Liang, Yilong Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2025 | CFMW: Cross-Modality Fusion Mamba for Robust Object Detection Under Adverse Weather
Binjia Zhou, You Yao, Jiacheng Lin, Kailun Yang 0001, Peng Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | SonarPoint: Weak-Heterogeneity Awareness Object Detection Network for 3D Sonar Point CloudabstractUnderwater target detection is primarily achieved through two methods: optical imaging and underwater sonar. 3D sonar, as the most advanced underwater detection technology, is characterized by strong penetration and long scanning distance, making it more suitable for tasks such as deep-sea exploration, murky water detection, and long-distance target identification. However, acquiring underwater sonar images is challenging, and there is no open-source 3D sonar dataset. Traditional three-dimensional target detection methods typically require highquality data and face significant challenges when dealing with weak heterogeneous sonar point clouds caused by high noise, low resolution, and occlusions. To address the aforementioned issues, we first propose a novel fuzzy decoupling module that differs from traditional foreground-background segmentation. This module simultaneously extracts valuable information about the target and its surrounding environment, mitigating the reduction in heterogeneity caused by noise and sonar side lobes. To achieve efficient fusion and capture global information after fuzzy decoupling, a multi-hop Mamba seamless adaptive decoupling point is introduced. It effectively enhances the connection between the two decoupled parts. To address missing and occlusion problems, a second-stage refinement based on Markov prediction is proposed. This low-cost approach, in contrast to using the original point cloud for contour and detail completion, enriches target boundary information. To validate our method, we have designed a practical 3D sonar imaging system and tested it through lake-based experiments. We have collected extensive raw data from Qiandao Lake and conducted annotation work. Through qualitative and quantitative experiments, our method outperforms the most advanced methods by 11.4%. Tiancheng Cai, Peng Chen 0008, Weibo Mao, Yingtian Hu, Yilong Zhang 0001, Yuanjie Dang, Ronghua Liang, Xiang Tian 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Orthogonal View-Based Attention Network for Layer Segmentation of 3D OCT FingerprintsabstractRecently, optical coherence tomography (OCT) has been used to noninvasively image the 3D structure of fingertip skin at high resolution. Unlike traditional 2D sensors (e.g., infrared light or capacitive technologies), the friction ridge information in 3D OCT fingerprint measurements requires reconstruction through layer segmentation. Accurate layer segmentation is helpful for fingerprint recognition and antispoofing applications. OCT volumes contain information corresponding to different directions that naturally provide complementary views. Inspired by this fact, we propose a novel orthogonal view-based attention network called OVA-Net, which exploits orthogonal views to learn the complementary information implied in the 3D fingerprint structure. Specifically, 3D convolutions and an A-line-based attention module are proposed in the B-scan view to model the long short-term intraslice correlations, whereas their counterparts in the C-scan view aim to model interslice correlations. An optical flow-based attention module is also proposed in the B-scan view to extract correlations between B-scans, which complements the interslice correlation learned in the C-scan view. Features from orthogonal views are progressively incorporated into a fusion pipeline for 3D layer segmentation. The effectiveness of OVA-Net is comprehensively evaluated in terms of layer segmentation accuracy, fingerprint reconstruction quality, and recognition performance. Yipeng Liu 0002, Zhanqing Li, Jiajin Qi, Hangtao Yu, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Test-Time Image Reconstruction for Cross-Device OCT Fingerprint ExtractionabstractOptical coherence tomography (OCT) technology enables imaging of 3D fingerprint structures. Extracting surface and internal fingerprints for identity recognition is possible by processing OCT images with layer segmentation and contour extraction. However, due to domain shift effects, OCT fingerprint extraction models often struggle to perform well across different devices. In this paper, a cross-device OCT fingerprint extraction method based on test-time image reconstruction is proposed. This method simultaneously trains layer segmentation and image reconstruction tasks during training. Additionally, a contour classification task is integrated to ensure the continuity and robustness of the contour extraction results. During the testing phase, image reconstruction is performed on test images, and the shared modules are updated to adapt the layer segmentation and contour classification network to the test domain. The result with the minimum inconsistency during the testing phase is selected as the final prediction. Experiments and comparisons are performed in terms of the distance between the ground truth and the extracted contours. Yipeng Liu 0002, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Supervised Enhancement for Fingertip OCT Images Based on Paired Dataset Generation StrategyabstractOptical Coherence Tomography (OCT) is a high-resolution, non-invasive imaging technology increasingly used for biometric data collection from fingertips. OCT captures volume data up to 3mm below the skin surface in the form of a series of B-scan images, enabling the reconstruction of internal fingerprints (IF) and internal sweat pores (ISP), thereby enhancing the security of biometric recognition. Despite the advantages, OCT images suffer from speckle noise and tissue discontinuity, making the extraction of subcutaneous biometric features challenging. Traditional hardware and software-based enhancement methods often result in over-smoothing and structural loss. Recent advancements in deep learning (DL) offer promising alternatives, with supervised DL methods showing efficacy when trained with high-quality paired datasets. However, the absence of ground-truth (GT) data makes it impossible to apply these models. This study proposes a novel supervised enhancement method for fingertip OCT images, with a paired dataset generation strategy. An OCT few-shot GAN and a Quality Estimation Module are proposed and incorporated into the strategy to realize translation from minimal GT manual augmentation to high-quality paired dataset, effectively addressing the challenge of data scarcity. A Fast Supervised Enhancement GAN (FSE-GAN) is proposed thereafter to perform simultaneous speckle noise reduction and tissue structure restoration, facilitating accurate extraction of internal fingerprints and sweat pores. Experiments demonstrate that the enhanced images significantly simplify IF and ISP extraction while achieving outstanding result quality. Qingran Miao, Haixia Wang 0002, Jianru Zhou, Yilong Zhang 0001, Peng Chen 0008, Ronghua Liang, Yuanjing Feng |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | CRM-NAS: A Structure-Adaptive and Attention-Based Approach for Fingerprint Reconstruction From Noisy OCT DataabstractAs essential biometric features, fingerprints have been widely utilized in various security domains. However, the performance of conventional Automated Fingerprint Identification Systems (AFISs) is limited by the quality of the external fingerprint (EF), particularly in cases involving damaged or deformed prints. Using the internal fingerprint (IF) acquired by Optical Coherence Tomography (OCT) to address these limitations has emerged as a promising method. IFs can compensate for and restore missing ridge pattern features in degraded EFs, thereby improving the overall recognition accuracy of AFIS. However, the reconstruction of IF was significantly constrained by speckle noise in OCT images, making the accurate extraction of finger tissue contours complex and computationally intensive. To improve the applicability of OCT fingerprint, this paper proposes a Neural Architecture Search (NAS)-based OCT fingertip internal contour regression network, denoted as CRM-NAS. The CRM-NAS employs a NAS-based internal feature extraction module (NAS-IEM) to adaptively optimize the network architecture and complexity with noisy OCT fingertip data, facilitating the effective capture of global internal contour features. Furthermore, an attention-based contour regression module (Att-CRM) is introduced to refine local contour details by leveraging multi-scale intermediate features extracted from different network layers and to enable the generation of continuous and accurate internal contours. Experimental results demonstrate that CRM-NAS not only outperforms existing methods in terms of contour extraction accuracy, fingerprint reconstruction quality, and verification performance, but also maintains a relatively compact parameter size. Haohao Sun, Sihan Lan, Haixia Wang 0002, Yilong Zhang 0001, Yipeng Liu 0002, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | A Fingerprint Quality Driven Transformer-CNN Hybrid Model for External and Internal Fingerprint FusionabstractAdvancements in internal fingerprint extraction technology have made the fusion of external and internal fingerprints possible. It offers a viable solution to the problem of degraded performance in Automatic Fingerprint Identification System (AFIS) caused by epidermal abrasion and aging. Traditional fusion methods focus on information maximization. But for fingerprint, features like wrinkles and scars often yield high gradient variation information. It is detrimental to generating high-quality fingerprint. To address this, we propose a novel quality driven fusion method for external and internal fingerprints. It comprises several components. Firstly, there is a lightweight and efficient Transformer-CNN hybrid model. Secondly, it includes a closed-loop quality driven fusion mechanism. This mechanism is equipped with a quality prediction module, Weighted Complementary Fusion (WCF), and quality feedback. Thirdly, there is a jointly optimized combined loss function, which is accompanied by an asynchronous cross-training strategy. Unlike traditional paradigms, we change the optimization objective. It is shifted from information maximization to quality maximization, which is more appropriate for fingerprint. Experimental evaluations have been conducted, covering aspects such as fingerprint quality, matching performance, and network model ablation. The method we proposed demonstrates superiority in terms of quality score and matching performance. It outperforms both traditional and state-of-the-art approaches. It gives a new research path to boost fingerprint identification performance in identity security authentication. Haixia Wang 0002, Yilong Zhang 0001, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | A Spatial-Aware Temporal Modeling Network for Imitation Learning-Based Drone NavigationabstractImitation learning-based drone autonomous navigation has attracted significant attention due to the ability of leveraging deep neural networks to learn the control policy from human pilot demonstrations. However, most current studies generate the control command using only a single image, overlooking the semantic information embedded in the sequential input images. While some reinforcement learning-based methods have explored the temporal modeling of sequential input images, they often overlook the spatial relations between frames and vectorize 2D information of each image into a 1D feature. In this paper, we propose a novel imitation learning-based method, termed the spatial-aware temporal modeling network (SATMN), for autonomous drone navigation using sequential images as input. Specifically, we introduce a spatial-temporal-separated modeling mechanism to extract low-resolution spatial features from original images and then perceive spatial-temporal relations among these 2D features. SATMN preserves the spatial information of each 2D image feature during temporal modeling and enables real-time onboard computing on a drone. To validate the effectiveness of the proposed method, we design a compact quadrotor platform capable of autonomous navigation using SATMN, entirely powered by onboard computing devices. Comprehensive and reproducible experiments on public datasets demonstrate the superior performance of our method compared to existing approaches. Tianwei Yu, Yuanjie Dang, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | MulDeF: A Model-Agnostic Debiasing Framework for Robust Multimodal Sentiment AnalysisabstractIn recent years, multimodal sentiment analysis (MSA) has gained prominence with the proliferation of social media. However, prior studies have often disregarded the possibility of spurious correlations between multimodal data and sentiment labels. Neglecting these factors often results in significant performance degradation, hampering the model's ability to generalize in out-of-distribution (OOD) scenarios. To gain a comprehensive understanding of multimodal knowledge and enhance the model's generalization across diverse distribution scenarios, we present the Multimodal Debiasing Framework (MulDeF). This model-agnostic framework addresses label bias through causal intervention and tackles multimodal biases using counterfactual reasoning. During the training phase, MulDeF rectifies multimodal representations through frontdoor adjustment in causal intervention, effectively eliminating label bias. In order to model conditional expectation calculations within the context of frontdoor adjustment, we introduce multimodal causal attention (MCA). In the inference phase, it employs counterfactual reasoning to eliminate multimodal biases. To further refine our debiasing strategy, we categorize multimodal biases into two distinct types: nonverbal bias and verbal bias. Nonverbal bias is addressed at the utterance level, involving the establishment of unimodal models for audio and visual modalities to estimate their biases concerning sentiment labels. Conversely, verbal bias mitigation occurs at the word level. Here, we mask “harmless” words to generate corresponding counterfactual texts, which are then assessed by the text model to identify word-level bias. Experimental results validate the effectiveness of MulDeF, showcasing its superior performance in OOD settings compared to state-of-the-art methods, while also achieving competitive results in independent and identically distributed (IID) settings. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Multim. | 3 |
| 2025 | Learning Video Salient Object Detection Progressively From Unlabeled VideosabstractRecently, deep learning-based video salient object detection (VSOD) has achieved some breakthroughs, but these methods rely on expensive annotated videos with pixel-wise annotations or weak annotations. In this paper, based on the similarities and differences between VSOD and image salient object detection (SOD), we propose a novel VSOD method via a progressive framework that locates and segments salient objects in sequence without utilizing any video annotation. To efficiently use the knowledge learned in the SOD dataset for VSOD efficiently, we introduce dynamic saliency to compensate for the lack of motion information of SOD during the locating process while maintaining the same fine segmenting process. Specifically, we utilize the coarse locating model trained on the image dataset, to identify frames with both static and dynamic saliency. Locating results of these frames are selected as spatiotemporal location labels. Moreover, by tracking salient objects in adjacent frames, the number of spatiotemporal location labels is increased. On the basis of these location labels, a two-stream locating network with an optical flow branch is proposed to capture salient objects in videos. The results with respect to five public benchmarks demonstrate that our method outperforms the state-of-the-art weakly and unsupervised methods. Binwei Xu, Qiuping Jiang, Haoran Liang 0001, Dingwen Zhang, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 6 |
| 2025 | Generate anomalies from normal: a partial pseudo-anomaly augmented approach for video anomaly detection
Yuanjie Dang, Jiangyun Chen, Peng Chen 0008, Nan Gao 0001, Ruohong Huan |
Vis. Comput. | 3 |
| 2024 | ACPNet: Enhancing Small-Scale Dieases Detection in Panoramic X-raysabstractDeep learning-based disease detection can automatically identify dental diseases in panoramic X-rays and improve the accuracy and efficiency of doctors’ diagnoses. However, due to the complex data distribution of panoramic oral X-rays, significant scale differences among lesions, and the presence of many small-scale diseases, automated disease detection in panoramic oral X-rays faces considerable challenges. To alleviate the aforementioned issues, we propose ACPNet, which introduces a novel two-stage approach for detecting small-scale dental diseases in panoramic X-rays using the Contextual Attention Alignment Network (CAAN) and the Point-to-Patch Module (PTPM). To the best of our knowledge, we are the first to explore the detection of small-scale dental diseases in panoramic X-rays under limited sample conditions. Specifically, CAAN integrates deformable convolution with the global attention mechanism of transformer attention, enabling the model to more accurately extract small target foreground features in the complex background of panoramic X-rays. PTPM employs key point detection and cascade dynamic patches to adjust the bounding boxes of lesions, ensuring that small-scale diseases have sufficient high-quality proposals, thereby enhancing detector performance. Additionally, we collected a dataset containing 1157 instances of dental diseases to validate the effectiveness of our algorithm. Extensive experiments demonstrate that ACPNet achieves state-of-the-art performance, highlighting its superiority over baseline and other detection methods. Nan Gao 0001, Junchao Zhu, Peng Chen 0008, Jijun Tang, Ronghua Liang |
BIBM | 3 |
| 2024 | Graph-Guided Multi-view Text Classification: Advanced Solutions for Fast Inference
Nan Gao 0001, Peng Chen 0008 |
ICANN (5) | 3 |
| 2024 | Task-Agnostic Self-Distillation for Few-Shot Action Recognition
Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Nan Gao 0001, Ruohong Huan, Xiaofei He 0001 |
IJCAI | 3 |
| 2024 | Semantic-Aware and Quality-Aware Interaction Network for Blind Video Quality AssessmentabstractCurrent state-of-the-art video quality assessment (VQA) models typically integrate various perceptual features to comprehensively represent video quality degradation. These models either directly concatenate features or fuse different perceptual scores while ignoring the domain gaps between cross-aware features, thus failing to adequately learn the correlations and interactions between different perceptual features. To this end, we analyze the independent effects and information gaps of quality-and semantic-aware features on video quality. Based on an analysis of the spatial and temporal differences between two aware features, we propose a semantic-Aware and quality-Aware Interaction Network (A2INet) for blind VQA. For spatial gaps, we introduce a cross-aware guided interaction module to enhance the interaction between semantic-and quality-aware features in a local-to-global manner. Considering temporal discrepancies, we design a cross-aware temporal modeling module to further perceive temporal content variation and quality saliency information, and perceptual features are regressed into quality score by a temporal network and a temporal pooling. Extensive experiments on six benchmark VQA datasets show that our model achieves state-of-the-art performance, and ablation studies further validate the effectiveness of each module. We also present a simple video sampling strategy to balance the effectiveness and efficiency of the model. The code for the proposed method will be released at https://github.com/JianjunXiang/A2INet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan, Nan Gao 0001 |
ACM Multimedia | 3 |
| 2024 | Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionabstractTemporal relation modeling is one of the core aspects of few-shot action recognition. Most previous works mainly focus on temporal relation modeling based on coarse-level actions, without considering the atomic action details and fine-grained temporal information. This oversight represents a significant limitation in this task. Specifically, coarse-level temporal relation modeling can make the few-shot models overfit in high-discrepancy temporal context, and ignore the low-discrepancy but high-semantic relevance action details in the video. To address these issues, we propose a saliency-guided fine-grained temporal mask learning method that models the temporal atomic action relation for few-shot action recognition in a finer manner. First, to model the comprehensive temporal relations of video instances, we design a temporal mask learning architecture to automatically search for the best matching of each atomic action snippet. Next, to exploit the low-discrepancy atomic action features, we introduce a saliency-guided temporal mask module to adaptively locate and excavate the atomic action information. After that, the few-shot predictions can be obtained by feeding the embedded rich temporal-relation features to a common feature matcher. Extensive experimental results on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Yuanjie Dang, Peng Chen 0008, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 3 |
| 2024 | Focus on Subtle Actions: Semantic and Saliency Knowledge Co-Propagation Method for Weakly-Supervised Temporal Action Localization
Yuanjie Dang, Haoyu Shou, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Yilong Zhang 0001 |
PRCV (10) | 3 |
| 2024 | TLCSFI: A Pose-Guided Person Re-Identification Method with Two-Level Channel-Spatial Feature IntegrationabstractPerson re-identification methods currently encounter challenges in feature learning, primarily due to difficulties in expressing the correlation between local features and integrating global and local features effectively. To address these issues, a pose-guided person re-identification method with Two-Level Channel–Spatial Feature Integration (TLCSFI) is proposed. In TLCSFI, a two-level integration mechanism is implemented. At the first level, TLCSFI integrates the spatial information from local features to generate fine-grained spatial features. At the second level, the fine-grained spatial feature and the coarse-grained channel feature are integrated together to complete channel–spatial feature integration. In the method, a Pose-based Spatial Feature Integration (PSFI) module is introduced to generate the pose union feature, which calculates intra-body affinity to guide the integration of spatial information among local pose feature maps. Then, a Channel and Spatial Union Feature Integration (CSUFI) module is proposed to efficiently integrate the channel information of the global feature and the spatial information of the pose union feature. Two individual networks are designed to extract channel and spatial information, respectively, in CSUFI, which are then weighted and integrated. Experiments are conducted on three publicly available datasets to evaluate TLCSFI, and the experimental results demonstrate its competitive performance. Ruohong Huan, Nan Gao 0001, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2024 | Learning Reliable Dense Pseudo-Labels for Point-Level Weakly-Supervised Action LocalizationabstractAbstract Point-level weakly-supervised temporal action localization aims to accurately recognize and localize action segments in untrimmed videos, using only point-level annotations during training. Current methods primarily focus on mining sparse pseudo-labels and generating dense pseudo-labels. However, due to the sparsity of point-level labels and the impact of scene information on action representations, the reliability of dense pseudo-label methods still remains an issue. In this paper, we propose a point-level weakly-supervised temporal action localization method based on local representation enhancement and global temporal optimization. This method comprises two modules that enhance the representation capacity of action features and improve the reliability of class activation sequence classification, thereby enhancing the reliability of dense pseudo-labels and strengthening the model’s capability for completeness learning. Specifically, we first generate representative features of actions using pseudo-label feature and calculate weights based on the feature similarity between representative features of actions and segments features to adjust class activation sequence. Additionally, we maintain the fixed-length queues for annotated segments and design a action contrastive learning framework between videos. The experimental results demonstrate that our modules indeed enhance the model’s capability for comprehensive learning, particularly achieving state-of-the-art results at high IoU thresholds. Yuanjie Dang, Guozhu Zheng, Peng Chen 0008, Nan Gao 0001, Ruohong Huan, Ronghua Liang |
Neural Process. Lett. | 3 |
| 2024 | ZJUT-EIFD: A Synchronously Collected External and Internal Fingerprint DatabaseabstractExternal fingerprints (EFs) based only on epidermal information are vulnerable to spoofing attacks and non-ideal skin conditions. To solve such shortcomings, internal fingerprints (IFs) collected using optical coherence tomography (OCT) have been proposed and widely researched. However, the development of IF is limited by the lack of in-depth researches on the IF and the EF-IF interoperability, which is partially caused by the lack of public OCT database. The obvious gap in the applications of EF and IF recognition motivated us to design and publish a comprehensive fingerprint database containing both traditional EFs and OCT IFs, denoted as ZJUT-EIFD. To the best of our knowledge, ZJUT-EIFD is the first public database that combines OCT and total internal reflection (TIR) via synchronous acquisition, with 399 different fingers from 60 subjects. In this article, the composition of the database, the quality of EFs and IFs, and the verification performance of different types of fingerprints were detailed. In addition, potential application directions of ZJUT-EIFD were demonstrated. ZJUT-EIFD can serve benchmarks and interoperability tests for EF-IF research, which will promote the research and development of EF and IF. Haohao Sun, Haixia Wang 0002, Yilong Zhang 0001, Ronghua Liang, Peng Chen 0008, Jianjiang Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | TriSAT: Trimodal Representation Learning for Multimodal Sentiment AnalysisabstractTransformer-based multimodal sentiment analysis frameworks commonly facilitate cross-modal interactions between two modalities through the attention mechanism. However, such interactions prove inadequate when dealing with three or more modalities, leading to increased computational complexity and network redundancy. To address this challenge, this paper introduces a novel framework, Trimodal representations for Sentiment Analysis from Transformers (TriSAT), tailored for multimodal sentiment analysis. TriSAT incorporates a trimodal transformer featuring a module called Trimodal Multi-Head Attention (TMHA). TMHA considers language as the primary modality, combines information from language, video, and audio using a single computation, and analyzes sentiment from a trimodal perspective. This approach significantly reduces the computational complexity while delivering high performance. Moreover, we propose Attraction-Repulsion (AR) loss and Trimodal Supervised Contrastive (TSC) loss to further enhance sentiment analysis performance. We conduct experiments on three public datasets to evaluate TriSAT's performance, which consistently demonstrates its competitiveness compared to state-of-the-art approaches. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | A Distributed and Parallel Accelerator Design for 3-D Acoustic Imaging on FPGA-Based Systemsabstract3-D imaging sonar is crucial in the exploration of marine resources, and the development of portable device with high imaging quality and high real-time performance is the general trend. However, traditional framework methods are limited by the huge amount of computation brought by high-quality imaging, making it difficult to implement in engineering. To address this issue, we develop 3-D real-time sonar system in an algorithm-hardware co-designed way. An ultrawideband distributed and parallel subarray beamforming algorithm (UWBDPS) is proposed for 3-D acoustic imaging. This is a multi-stage array time-frequency beamforming method under a distributed parallel computing architecture. Based on this, we propose field-programmable gate array (FPGA)-based accelerator. It divides a large sonar receiving planar array into multiple parallel subarrays, and complete the beamforming in two stages, which can reduce the calculation load and speeds up 3-D imaging. For engineering implementation, we optimized the sparseness of the planar transducer array, with a sparse rate as high as 97.7%. The experimental results show that the calculation amount of the proposed UWB-DPS algorithm is reduced to 1/5.7 of the traditional framework algorithm, the imaging performance is effectively improved, and the FPGA-based accelerator outperforms the CPU software implementation by 935×. Weibo Mao, Peng Chen 0008, Yingtian Hu, Haoran Liang 0001, Yuanjie Dang, Ronghua Liang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | A Visual Representation-Guided Framework With Global Affinity for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) methods have made considerable progress in performance, yet these models rely heavily on expensive pixel-wise labels. Recently, to achieve a trade-off between labeling burden and performance, scribble-based SOD methods have attracted increasing attention. Previous scribble-based models directly implement the SOD task only based on SOD training data with limited information, it is extremely difficult for them to understand the image and further achieve a superior SOD task. In this paper, we propose a simple yet effective framework guided by general visual representations with rich contextual semantic knowledge for scribble-based SOD. These general visual representations are generated by self-supervised learning based on large-scale unlabeled datasets. Our framework consists of a task-related encoder, a general visual module, and an information integration module to efficiently combine the general visual representations with task-related features to perform the SOD task based on understanding the contextual connections of images. Meanwhile, we propose a novel global semantic affinity loss to guide the model to perceive the global structure of the salient objects. Experimental results on five public benchmark datasets demonstrate that our method, which only utilizes scribble annotations without introducing any extra label, outperforms the state-of-theart weakly supervised SOD methods. Specifically, it outperforms the previous best scribble-based method on all datasets with an average gain of 5.5% for max f-measure, 5.8% for mean f-measure, 24% for MAE, and 3.1% for E-measure. Moreover, our method achieves comparable or even superior performance to the state-of-the-art fully supervised models. Binwei Xu, Haoran Liang 0001, Weihua Gong, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A Wavelet-Based Memory Autoencoder for Noncontact Fingerprint Presentation Attack DetectionabstractFingerprint presentation attack detection (FPAD) is essential in fingerprint identification systems. Noncontact methods such as fingerprint biometrics are becoming popular because they are not affected by skin conditions and there are no hygiene issues. However, most of the existing noncontact FPAD methods are supervised methods with poor generalizability and poor performance during events such as unseen presentation attacks (PAs). Moreover, easily overlooked frequency domain information contributes to the fingerprint antispoofing task. Therefore, we propose a wavelet-based memory-augmented autoencoder that fully utilizes the frequency domain information. Specifically, the model first decomposes the input image into high- and low-frequency information and extracts features separately. Subsequently, we propose a frequency complementary connection (FCC) module to realize the fusion and complementation of frequency domain information at the feature level. Moreover, a memory distance expansion loss is proposed to keep the memory module diverse. Experiments are conducted to verify the effectiveness of the method. The code of our model is available onhttps://github.com/SuperIOyht/WaveMemAE. Yipeng Liu 0002, Hangtao Yu, Haonan Fang, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | A multi-stage recognizer for nested named entity with weakly labeled data
Nan Gao 0001, Bowei Yang, Peng Chen 0008, Li Ping Qian 0001 |
J. Supercomput. | 3 |
| 2024 | Multi-Level Objective Alignment Transformer for Fine-Grained Oral Panoramic X-Ray Report GenerationabstractAutomatically generated oral panoramic X-ray report is highly beneficial for improving the efficiency of dental diagnosis. However, recent solutions adopt holistic methods, resulting in a cursory description of the oral condition. This may lead to reports lacking details, such as specific sites or lesion contours. Therefore, we propose a Multi-Level objective Alignment Transformer(MLAT) network, which integrates all tooth and disease objects into a positional alignment graph to extract fine-grained object-level features. Specifically, we introduce a novel Object-Level Collaborative Encoder (OLCE) module, which uses a positional alignment graph to construct object relationships. OLCE enhances object-level feature extraction by eliminating interference information between pathologically unrelated objects. In addition, we build a high-quality panoramic X-ray image-report dataset consisting of 562 sets of images and reports labeled by 13 experienced dental specialists. Experiments on the collected dataset show that the proposed MLAT significantly outperforms the state-of-the-art baselines by more than 5% in 4 different metrics, including BLEUs, Meteor, Rouge, and BERTScore. Nan Gao 0001, Renyuan Yao, Ronghua Liang, Peng Chen 0008, Tianshuang Liu, Yuanjie Dang |
IEEE Trans. Multim. | 4 |
| 2024 | UniMF: A Unified Multimodal Framework for Multimodal Sentiment Analysis in Missing Modalities and Unaligned Multimodal SequencesabstractIn current multimodal sentiment analysis, aligned and complete multimodal sequences are often crucial. Obtaining complete multimodal data in the real world presents various challenges, and aligning multimodal sequences often requires a significant amount of effort. Unfortunately, most multimodal sentiment analysis methods fail when dealing with missing modalities or unaligned multimodal sequences. To tackle these two challenges simultaneously in a simple and lightweight manner, we present the Unified Multimodal Framework (UniMF). The primary components of UniMF comprise two distinct modules. The first module, Translation Module, translates missing modalities using information from existing modalities. The second module, Prediction Module, uses the attention mechanism to fuse the multimodal information and generate predictions. To enhance the translation performance of the Translation Module, we introduce the Multimodal Generation Mask (MGM) and utilize it to construct the Multimodal Generation Transformer (MGT). The MGT can generate the missing modality while focusing on information from existing modalities. Furthermore, we introduce the Multimodal Understanding Transformer (MUT) in the Prediction Module, which includes the Multimodal Understanding Mask (MUM) and a unique sequence,MultiModalSequence(MMSeq), representing a unified multimodality. To assess the performance of UniMF, we perform experiments on four multimodal sentiment datasets, and UniMF attains competitive or state-of-the-art outcomes with fewer learnable parameters. Furthermore, the experimental outcomes signify that UniMF, supported by MGT and MUT - two transformers utilizing special attention mechanisms, can efficiently manage both generating task of missing modalities and understanding task of unaligned multimodal sequences. Ruohong Huan, Guowei Zhong, Peng Chen 0008, Ronghua Liang |
IEEE Trans. Multim. | 3 |
| 2024 | Pseudo Light Field Image and 4D Wavelet-Transform-Based Reduced-Reference Light Field Image Quality AssessmentabstractReduced-reference light field image (LFI) quality assessment (RR LFIQA) automatically assesses image quality with only partial information about the reference LFI is available. Existing RR LFIQA has difficulty extracting effective RR information and perceptual features to represent the LFI quality. In this article, we propose an RR LFIQA model based on pseudo LFI (PLFI) and four-dimensional (4D) wavelet transform. To extract RR information related to LFI perceptual quality, a PLFI is created as the RR information of the LFI using a view synthesis algorithm. Considering that the high-dimensional characteristics of the PLFI, 4D wavelet transform is used to decompose the original and distorted PLFIs. The 4D wavelet transform essentially performs a continuous 1D wavelet transform for the 4D signal to enable the local 4D structure of the PLFIs to be characterized effectively in the 4D wavelet domain. A novel spatial-angular weighting strategy is proposed to describe the importance of each location for quality evaluation, to further improve the performance of the proposed method. Experimental results on four benchmark datasets show that the proposed model performs better than the representative 2DIQA and LFIQA models. Jianjun Xiang, Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Gangyi Jiang |
IEEE Trans. Multim. | 2 |
| 2024 | Synthesize Boundaries: A Boundary-Aware Self-Consistent Framework for Weakly Supervised Salient Object DetectionabstractFully supervised salient object detection (SOD) has made considerable progress based on expensive and time-consuming data with pixel-wise annotations. Recently, to relieve the labeling burden while maintaining performance, some scribble-based SOD methods have been proposed. However, learning precise boundary details from scribble annotations that lack edge information is still difficult. In this article, we propose to learn precise boundaries from our designed synthetic images and labels without introducing any extra auxiliary data. The synthetic image creates boundary information by inserting synthetic concave regions that simulate the real concave regions of salient objects. Furthermore, we propose a novel self-consistent framework that consists of a global integral branch (GIB) and a boundary-aware branch (BAB) to train a saliency detector. GIB aims to identify integral salient objects, whose input is the original image. BAB aims to help predict accurate boundaries, whose input is the synthetic image. These two branches are connected through a self-consistent loss to guide the saliency detector to predict precise boundaries while identifying salient objects. Experimental results on five benchmarks demonstrate that our method outperforms the state-of-the-art weakly supervised SOD methods and further narrows the gap with the fully supervised methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IEEE Trans. Multim. | 4 |
| 2024 | Discriminative Action Snippet Propagation Network for Weakly Supervised Temporal Action LocalizationabstractWeakly supervised temporal action localization (WTAL) aims to classify and localize actions in untrimmed videos with only video-level labels. Recent studies have attempted to obtain more accurate temporal boundaries by exploiting latent action instances in ambiguous snippets or propagating representative action features. However, empirically handcrafted ambiguous snippet extraction and the imprecise alignment of representative snippet propagation lead to challenges in modeling the completeness of actions for these methods. In this article, we propose a Discriminative Action Snippet Propagation Network (DASP-Net) to accurately discover ambiguous snippets in videos and propagate discriminative instance-level features throughout the video for improving action completeness. Specifically, we introduce a novel discriminative feature propagation module for capturing the global contextual attention and propagating the action concept across the whole video by perceiving the discriminative action snippets with instance information from the same video. Simultaneously, we incorporate denoised pseudo-labels as supervision, where we correct the controversial prediction based on the feature space distribution during training, thereby alleviating false detection caused by noise background features. Furthermore, we design an ambiguous feature mining module, which maximizes the feature affinity information of action and background in ambiguous snippets to generate more accurate latent action and background snippets and learns more precise action instance boundaries through contrastive learning of action and background snippets. Extensive experiments show that DASP-Net achieves state-of-the-art results on THUMOS14 and ActivityNet1.2 datasets. Yuanjie Dang, Chunxia Huang, Peng Chen 0008, Nan Gao 0001, Ronghua Liang, Ruohong Huan |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | BTCN: Bridging the Gap Between Pre-trained and Downstream Models for Endoscopic Caries DetectionabstractAlthough deep learning has been widely applied in the field of dental caries detection, there are still certain challenges that need to be addressed. The limitations of sharing the same backbone between the pre-trained model and the downstream model hinder the feature alignment capability of self-supervised learning (SSL) during the fine-tuning stage, leading to incomplete transfer from the pre-trained model to the downstream model. To address this challenge, we introduce an SSL pre-trained model called Bi-branches Transformer CNN Network (BTCN). BTCN adopts a parallel structure combining the CNN and Transformer branches. This parallel structure allows the pre-trained model to capture additional global representations, which helps alleviate feature differences during fine-tuning and better adapt to downstream detection models. Additionally, to further enhance the fusion quality of the bi-branches encoder, we introduced the Multi-layer Supervision Strategy (MSS) to increase the supervision on features at different layers. To validate the effectiveness of our approach, we collected a dedicated dataset for caries detection, comprising 1039 endoscopic images of dental caries. Through extensive experimental research, our results demonstrate the effectiveness of the proposed BTCN and MSS, showing significant improvements compared to the current state-of-the-art methods. Nan Gao 0001, Peng Chen 0008, Yukai Li, Jijun Tang, Ronghua Liang, Tianshuang Liu |
BIBM | 2 |
| 2023 | STAN: Spatio-Temporal Alignment Network for No-Reference Video Quality Assessment
Zhengyi Yang 0008, Yuanjie Dang, Jianjun Xiang, Peng Chen 0008 |
ICANN (3) | 4 |
| 2023 | Spatial-angular Quality-aware Representation Learning for Blind Light Field Image Quality AssessmentabstractBlind light field image quality assessment (BLFIQA) remains a challenging task in deep learning due to the unique spatial-angular structure of light field images (LFIs) and the lack of large-scale labeled data for training. In this work, we propose a novel BLFIQA method using spatial-angular quality-aware representation learning in a self-supervised learning manner. Visual content and distortion type are important factors affecting the perceived quality of LFIs. In our observation, the band-pass transform maps of LFIs with the same distortion type exhibit similar Gaussian distributions. Thus, we learn spatial-angular quality-aware representations by minimizing the distance in the embedding space between the luminance map and the band-pass transform map of the same LFI. To implement spatial-angular quality-aware representations of LFI, we also build a large-scale unlabeled dataset containing 40k distorted LFIs with different distortion types and visual content. Further, we propose a fusion-separation-fusion network (FSFNet) to extract features for representing the intrinsic spatial-angular structure of the LFI. After pre-training on the unlabeled dataset using the proposed self-supervised learning, the FSFNet is employed for downstream BLFIQA tasks and achieves good performance. Experimental results show that our proposed method outperforms seventeen state-of-the-art models on the Win5-LID, NBU-LF1.0 and LFDD datasets, and achieves 3.78%, 6.61% and 4.06% SRCC improvements, respectively. The code and dataset will be publicly available in https://github.com/JianjunXiang/SSL_and_FSFNet. Jianjun Xiang, Yuanjie Dang, Peng Chen 0008, Ronghua Liang, Ruohong Huan |
ACM Multimedia | 3 |
| 2023 | Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionabstractFew-shot action recognition (FSAR) aims to classify unseen query actions into categories represented by a few labeled support videos. Most current FSAR methods adopt the frame-level matching mechanism that requires continuous actions to be represented by a fixed number of frame features. However, this could compromise the completeness of the contextual video information and make it difficult to handle video features of varying frame sampling speeds. In this paper, we propose a multi-speed global contextual subspace matching (MGCSM) method that generates global contextual action subspace representations from videos containing different numbers of frames to preserve contextual semantic information. Specifically, we propose to obtain the scale-agnostic information of embedding video features using a global contextual aggregation (GCA) module and then generate the discriminative action subspace representation with an action subspace generation (ASG) module. Furthermore, we introduce a multi-speed subspace matching (MSM) mechanism that generates a multi-speed classification score by integrating the similarities between query videos and support subspaces of varying sampling speeds. The proposed method is embedding-agnostic and can be combined with most mainstream embedding networks without model re-designs. Comprehensive and reproducible experiments on standard datasets demonstrate our method's superior performance compared to existing state-of-the-art methods. Tianwei Yu, Peng Chen 0008, Yuanjie Dang, Ruohong Huan, Ronghua Liang |
ACM Multimedia | 2 |
| 2023 | NTAM: A New Transition-Based Attention Model for Nested Named Entity Recognition
Nan Gao 0001, Bowei Yang, Peng Chen 0008 |
NLPCC (2) | 4 |
| 2023 | HPAN: A Hybrid Pose Attention Network for Person Re-Identification
Ruohong Huan, Tianya Chen, Ziwei Zhan, Peng Chen 0008, Ronghua Liang |
PRCV (12) | 4 |
| 2023 | Anti-spoofing study on palm biometric features
Haixia Wang 0002, Lixun Su, Hongxiang Zeng, Peng Chen 0008, Ronghua Liang, Yilong Zhang 0001 |
Expert Syst. Appl. | 4 |
| 2023 | SS-Norm: Spectral-spatial normalization for single-domain generalization with application to retinal vessel segmentationabstractAbstract Retinal vessel segmentation is an important computer vision task for eye retinopathy diagnosis. In the real scenarios, most datasets of source domain and target domain have distribution deviation, and the model often fails to generate accurate segmentation results due to the lack of data variation in single‐source domain, which damages the generalization ability to unseen target domains and may mislead doctors or artificial intelligence model in the following diseases diagnosis. Feature normalization is one feasible solution which can standardize data into uniform and stable distribution without additional data. However, the existing methods like batch normalization, uniform the data by global parameters. This leads to insufficient representation of important semantic information in the local region. To address this problem, the authors propose the spectral‐spatial normalization (SS‐Norm) module to enhance the generalization ability of the model. More specifically, the authors perform a discrete cosine transform (DCT) to decompose the feature into multiple frequency components and to analyze the semantic contribution degree of each component. By learning a spectral vector, the authors reweight the frequency components of features and therefore normalize the distribution in the spectral domain. Extensive experiments on six datasets prove the effectiveness of the authors’ methods. Yipeng Liu 0002, Dongxu Zeng, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 4 |
| 2023 | DOMOPT: A Detection-Based Online Multi-Object Pedestrian Tracking Network for VideosabstractDue to the problem of low tracking accuracy and weak tracking stability of current multi-object pedestrian tracking algorithms in complex scenes for videos, a Detection-based Online Multi-Object Pedestrian Tracking (DOMOPT) network is proposed. First, a Multi-Level Feature Fusion (MLFF) pedestrian detection network is proposed based on the Center and Scale Prediction (CSP) algorithm. The pyramid convolutional neural network is used as the backbone to enhance the feature extraction capability for small objects. The shallow features and deep features at multiple levels are integrated to fully obtain the position and semantic information to further improve the detection performance for small objects. Then, on the basis of Joint Detection and Embedding (JDE) architecture, a Multi-Branch Pedestrian Appearance (MBPA) feature extraction network is proposed and added into the pedestrian detection network to extract the appearance feature vector corresponding to each pedestrian. The pedestrian appearance feature extraction is treated as a classification task jointly training with the pedestrian detection task, using the multi-task learning strategy. Experimental results show that the proposed network has better tracking accuracy and stability compared with state-of-the-art algorithms. Ruohong Huan, Shuaishuai Zheng, Chaojie Xie, Peng Chen 0008, Ronghua Liang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2023 | MLFFCSP: a new anti-occlusion pedestrian detection network with multi-level feature fusion for small targets
Ruohong Huan, Chaojie Xie, Ronghua Liang, Peng Chen 0008 |
Multim. Tools Appl. | 5 |
| 2023 | Boosting Short Text Classification by Solving the OOV ProblemabstractIn the field of natural language processing, text classification has received a lot of attention. Compared with long texts, short texts have fewer words and lack contextual semantic information. Existing approaches enrich short text information by linking the external knowledge graph, but they ignore the out-of-vocabulary (OOV) problem during entity linking, especially when dealing with domain-oriented data, which has some rare words or domain-specific nouns. In this paper, to alleviate the OOV problem caused by linking the external knowledge graph(KG), we propose a domain knowledge graph and entity complementation strategy to improve the performance of short text classification. Specifically, the external knowledge graph is used to enrich the information of short texts. The self-build domain knowledge graph is used to solve the problem of entities failing to link to the external knowledge graph. Finally, we conduct experiments on various datasets: 1. a labeled Chinese electronic domain dataset; 2. an open-source dataset to test the performance of our algorithm in different data distribution scenarios. The results demonstrate our dual knowledge graph model outperforms the state-of-the-art short text classification methods, especially when the OOV problem is severe. Nan Gao 0001, Peng Chen 0008, Jijun Tang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | End-to-End Surface and Internal Fingerprint Reconstruction From Optical Coherence Tomography Based on Contour RegressionabstractOptical coherence tomography (OCT), as a non-destructive and high-resolution imaging technique, has been used to collect 3D fingertip data, which contains surface and internal fingerprints. Methods have been proposed for OCT fingerprint reconstruction. However, these methods have complex processing flow and are time consuming. In this paper, an end-to-end convolutional neural network based surface and internal fingerprint reconstruction method is proposed. A simple yet effective contour regression module is proposed and integrated in the network for direct estimation of contours of stratum corneum and viable epidermis junction from noisy OCT volume data, thus greatly simplify the processing flow. The proposed network further integrates multi-task learning with conventional segmentation task as auxiliary task and contour regression task as main task to facilitate the feature extraction and improve the robustness of the network. Depthwise separable convolution is adapted to a light-weight network for network computation complexity reduction. To the best of our knowledge, it is the first time that an end-to-end method is proposed for surface and internal fingerprint extraction from noisy OCT volume data. Experiments and comparisons are carried out in terms of contour estimation accuracy, fingerprint quality, fingerprint matching performance and computation efficiency. Compared with conventional method, the proposed method utilizes only 6% of original network parameters and 0.7% of original computation time, but achieves comparably results. Fingerprint by depth proves the accuracy and robustness of contour regression than pixel-wise layer segmentation. The proposed method is noise-insensitive, process-simple and time-efficient for OCT fingerprint reconstruction, which is significant for real time application in Automated Fingerprint Recognition Systems. Baojin Ding, Haixia Wang 0002, Ronghua Liang, Yilong Zhang 0001, Peng Chen 0008 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2023 | A New Approach in Automated Fingerprint Presentation Attack Detection Using Optical Coherence TomographyabstractPresentation attack detection (PAD) is a critical component of automated fingerprint recognition systems (AFRSs). However, existing PAD technologies based on optical coherence tomography (OCT) mainly rely on local information, ignoring the global continuity and correlation of physiological structures. Furthermore, the lack of appropriate presentation attack instruments (PAIs) that cater to the unique OCT characteristics leads to the insufficient evaluation of PAD. The identification features, including external fingerprint (EF), internal fingerprint (IF), and subcutaneous sweat pore (SSP), provide valuable information about the intrinsic connections of physiological structures. Such intrinsic connections hold potential clues for PAD. Building upon this premise, this paper proposed a novel PAD method based on three OCT hand-crafted features: EF-IF self-matching score (SMS), SSP number (SN), and SSP coincidence rate (SCR). These simple yet effective PAD features offer a more precise and detailed description of the internal physiological structure, enabling accurate distinction between presentation attack (PA) and bona-fide. The proposed method achieves a 4% Equal Error Rate (EER), significantly outperforming other existing PAD methods. Additionally, the cross-device experiment demonstrates the generalization capability of the proposed method on both our dataset and the public OCT dataset. Haohao Sun, Yilong Zhang 0001, Peng Chen 0008, Haixia Wang 0002, Yipeng Liu 0002, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | Path-Analysis-Based Reinforcement Learning Algorithm for Imitation FilmingabstractImitation filming has been applied to autonomous filming by mimicking human operators. To imitate the operation of cameramen when filming multiple human actions, existing methods plan the camera motion through time series prediction or train multiple models to handle a particular style in a specific situation. As a result, these methods require various settings to adapt to different scenarios. In this work, we overcome such limitations and propose an end-to-end imitation learning framework for drone cinematography systems. The framework consists of two main components: (1) an efficient motion feature extraction module for generating a compact motion feature space, (2) a path-analysis-based reinforcement learning (PABRL) algorithm for imitating multiple filming styles from demonstrations and incorporating aesthetical features for improved perspective shots. Our PABRL method is based on the actor–critic network, which regards multiple human motion variables, camera translations, and image composition as inputs and then outputs an aesthetical filming strategy related to the subject motion. In addition, we propose an attention mechanism and a long–short-term rewarding function to enhance the motion feature space and the integrity of the generated trajectory, respectively. Extensive experimental results in simulated and real outdoor environments demonstrate that compared with state-of-the-art methods, our method can achieve 69.8% higher performance in terms of trajectory planning accuracy while successfully incorporating aesthetical features into the captured videos. Yuanjie Dang, Chong Huang 0005, Peng Chen 0008, Ronghua Liang, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 3 |
| 2022 | CFN: A coarse-to-fine network for eye fixation predictionabstractAbstract Many image‐to‐image computer vision approaches have made great progress by an end‐to‐end framework with the encoder–decoder architecture. However, the same image‐to‐image eye fixation prediction task is not the same as those computer vision tasks in that it focuses more on salient regions rather than precise predictions for every pixel. Thus, it is not appropriate to directly apply the end‐to‐end encoder–decoder to the eye fixation prediction task. In addition, although high‐level feature is important, the contribution of low‐level feature should also be kept and balanced in computational model. Nevertheless, some low‐level features that attract attention are easily neglected while transiting through the deep network. Therefore, the effective way to integrate low‐level and high‐level features for improving eye fixation prediction performance is still a challenging task. In this paper, a coarse‐to‐fine network (CFN) that encompasses two pathways with different training strategies are proposed: coarse perceiving network (CFN‐Coarse) can be a simple encoder network or any of the existing pretrained network to capture the distribution of salient regions and generate high‐quality feature maps; fine integrating network (CFN‐Fine) uses fixed parameters from the CFN‐Coarse and combines features from deep to shallow in the deconvolution process by adding skip connections between down‐sampling and up‐sampling paths to efficiently integrate deep and shallow features. The saliency map obtained by the method is evaluated over 6 standard benchmark datasets, namely SALICON, MIT1003, MIT300, Toronto, OSIE, and SUN500. The results demonstrate that the method can surpass the state‐of‐the‐art accuracy of eye fixation prediction and achieves the competitive performance to date under most evaluation metrics on SALICON Saliency Prediction Challenge (LSUN2017). Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
IET Image Process. | 4 |
| 2022 | One-Shot Imitation Drone Filming of Human Motion VideosabstractImitation learning has recently been applied to mimic the operation of a cameraman in existing autonomous camera systems. To imitate a certain demonstration video, existing methods require users to collect a significant number of training videos with a similar filming style. Because the trained model is style-specific, it is challenging to generalize the model to imitate other videos with a different filming style. To address this problem, we propose a framework that we term "one-shot imitation filming", which can imitate a filming style by "seeing" only a single demonstration video of the target style without style-specific model training. This is achieved by two key enabling techniques: 1) filming style feature extraction, which encodes sequential cinematic characteristics of a variable-length video clip into a fixed-length feature vector; and 2) camera motion prediction, which dynamically plans the camera trajectory to reproduce the filming style of the demo video. We implemented the approach with a deep neural network and deployed it on a 6 degrees of freedom (DOF) drone system by first predicting the future camera motions, and then converting them into the drone's control commands via an odometer. Our experimental results on comprehensive datasets and showcases exhibit that the proposed approach achieves significant improvements over conventional baselines, and our approach can mimic the footage of an unseen style with high fidelity. Chong Huang 0005, Yuanjie Dang, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Locate Globally, Segment Locally: A Progressive Architecture With Knowledge Review Network for Salient Object DetectionabstractSalient object location and segmentation are two different tasks in salient object detection (SOD). The former aims to globally find the most attractive objects in an image, whereas the latter can be achieved only using local regions that contain salient objects. However, previous methods mainly accomplish the two tasks simultaneously in a simple end-to-end manner, which leads to the ignorance of the differences between them. We assume that the human vision system orderly locates and segments objects, so we propose a novel progressive architecture with knowledge review network (PA-KRN) for SOD. It consists of three parts. (1) A coarse locating module (CLM) that uses body-attention label locates rough areas containing salient objects without boundary details. (2) An attention-based sampler highlights salient object regions with high resolution based on body-attention maps. (3) A fine segmenting module (FSM) finely segments salient objects. The networks applied in CLM and FSM are mainly based on our proposed knowledge review network (KRN) that utilizes the finest feature maps to reintegrate all previous layers, which can make up for the important information that is continuously diluted in the top-down path. Experiments on five benchmarks demonstrate that our single KRN can outperform state-of-the-art methods. Furthermore, our PA-KRN performs better and substantially surpasses the aforementioned methods. Binwei Xu, Haoran Liang 0001, Ronghua Liang, Peng Chen 0008 |
AAAI | 4 |
| 2021 | Subcutaneous sweat pore estimation from optical coherence tomographyabstractAbstract Abstract Sweat pore, one of the level 3 features of fingerprint, has attracted much attention in fingerprint recognition. Traditional sweat pores on surface fingerprint are unclear or blurred when fingers are stained or damaged. Subcutaneous sweat pores, as cross section of the sweat glands, are resistant to external interferences. With 3D fingertip information measured by optical coherence tomography (OCT), the subcutaneous sweat pore estimation from OCT volume data is investigated. First, an adaptive subcutaneous pore image reconstruction method is proposed. It utilizes the skin surface and viable epidermis junction as reference and realizes depth‐adaptive pore image reconstruction. Second, a dilated U‐Net combining the U‐Net with dilated convolution is proposed for subcutaneous sweat pore extraction, which can prevent information loss of sweat pores caused by downsampling. To the best knowledge, it is the first time that subcutaneous sweat pore extraction is investigated and proposed. Experiments on subcutaneous pore image reconstruction and sweat pore extraction are both conducted. The qualitative and quantitative results show that the proposed adaptive method performs better in subcutaneous pore image reconstruction compared with the fix‐depth method, and the dilated U‐Net outperforms other methods on subcutaneous sweat pore extraction. Baojin Ding, Haixia Wang 0002, Peng Chen 0008, Yilong Zhang 0001, Ronghua Liang, Yipeng Liu 0002 |
IET Image Process. | 3 |
| 2021 | Blood vessel and background separation for retinal image quality assessmentabstractAbstract Retinal image analysis has become an intuitive and standard aided diagnostic technique for eye diseases. The good image quality is essential support for doctors to provide timely and accurate disease diagnosis. This paper proposes an end‐to‐end learning based method for evaluating the retinal image quality. First, blood vessels of the input image are segmented by U‐Net, and the fundus image is divided into two parts: blood vessels and background. Then, we design a dual branch network module which extracts global features that influence the image quality and suppress the interference of blood vessels and local textures to achieve better performance. The proposed module can be embedded in various advanced network structures. The experimental results show the more efficient convergence rate for the network with the module. The best network accuracy rate is 85.83%, the AUC is 0.9296, and the F1‐score is 0.7967 on the collected local dataset. Additionally, the model generalization is tested on the public DRIMDB dataset. The accuracy, AUC, and F1‐score reach 97.89%, 0.9978, and 0.9688, respectively. Compared with the state‐of‐the‐art networks, the performance of the proposed method is proven to be accurate and effective for retinal image quality assessment. Yipeng Liu 0002, Yajun Lv, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 6 |
| 2021 | Feature pyramid U-Net for retinal vessel segmentationabstractAbstract The retinal vessel is the only microvascular network that can be directly and non‐invasively observed in humans. Cardiovascular and cerebrovascular diseases, such as diabetes, hypertension, can lead to structural changes of the retinal microvascular network. Therefore, it is of great significance to study effective retinal vessel segmentation methods and assist doctors in early diagnoses with quantitative results for vascular networks. In this study, we propose a novel convolutional neural network named feature pyramid U‐Net (FPU‐Net) that extracts multiscale representations by constructing two feature pyramids both on the encoder and the decoder of U‐Net. In this representation, objects features with different size like micro‐vessels and pathology will be fused for better vessel segmentation. The experimental results show that compared with state‐of‐the‐art methods, FPU‐Net is superior in terms of accuracy, sensitivity, F1‐score, and area under the curve and capable of stronger domain generalisation across different datasets. Yipeng Liu 0002, Xue Rui, Zhanqing Li, Dongxu Zeng, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 6 |
| 2021 | Multiscale ensemble of convolutional neural networks for skin lesion classificationabstractAbstract Early detection and treatment of skin cancer can considerably reduce the patient mortality rates. Convolutional neural network (CNN) has been widely applied in the field of computer aided diagnosis. However, for skin lesions, the inconsistent size of lesion regions in dermatoscope images hinders the convolutional neural network precise discrimination. To solve this problem, multiscale ensemble of convolutional neural networks called MECNN is proposed, which involves three branches with different lesion scales as the model input. The first branch locates the lesion region outline by identifying the largest local response point. Then, MECNN reduces the search area of the lesion region and divides the outline into two scales used as the input for the other two branches. A global loss function is defined to control the learning objectives of the three branches and MECNN fuses the branches output as the final classification result. The proposed model is evaluated on the public HAM10000 dataset and achieves a higher classification accuracy than the comparative state‐of‐the‐art methods. Yipeng Liu 0002, Zhanqing Li, Peng Chen 0008, Ronghua Liang |
IET Image Process. | 6 |
| 2021 | Video multimodal emotion recognition based on Bi-GRU and attention fusion
Ruohong Huan, Jia Shu, Shenglin Bao, Ronghua Liang, Peng Chen 0008, Kaikai Chi |
Multim. Tools Appl. | 5 |
| 2021 | A hybrid CNN and BLSTM network for human complex activity recognition with multi-feature fusion
Ruohong Huan, Ziwei Zhan, Luoqi Ge, Kaikai Chi, Peng Chen 0008, Ronghua Liang |
Multim. Tools Appl. | 5 |
| 2021 | Surface and Internal Fingerprint Reconstruction From Optical Coherence Tomography Through Convolutional Neural NetworkabstractOptical coherence tomography (OCT), as a non-destructive and high-resolution fingerprint acquisition technology, is robust against poor skin conditions and resistant to spoof attacks. It measures fingertip information on and beneath skin as 3D volume data, containing the surface fingerprint, internal fingerprint and sweat glands. Various methods have been proposed to extract internal fingerprints, which ignore the inter-slice dependence and often require manually selected parameters. In this article, a modified U-Net that combines residual learning, bidirectional convolutional long short-term memory and hybrid dilated convolution (denoted as BCL-U Net) for OCT volume data segmentation and two fingerprint reconstruction approaches are proposed. To the best of our knowledge, it is the first time that simultaneous and automatic extraction is performed for surface fingerprint, internal fingerprint and sweat gland. The proposed BCL-U Net utilizes the spatial dependence in OCT volume data and deals with segmentation of objects with diverse sizes to achieve accurate extraction. Comparisons have been performed to demonstrate the advantages of the proposed method. A thorough evaluation of the recognition abilities of internal and surface fingerprints is conducted using a dataset significantly larger than previous studies. Four databases containing internal and surface fingerprints are generated from 1572 OCT volume data by the proposed method. The internal fingerprint matching experiment has achieved a lowest equal error rate (EER) of 0.95%. Mixed internal and surface fingerprint matching experiment is also performed and achieves an EER of 3.67%, verifying the consistency of the internal and surface fingerprints. The matching experiments for fingers under poor skin conditions show a 2.47% EER of internal fingerprints that is much lower than that of surface fingerprints, which proves the advantage of internal fingerprints and indicates the potential of the internal fingerprints to supplement or replace the surface fingerprints for some specific applications. Baojin Ding, Haixia Wang 0002, Peng Chen 0008, Yilong Zhang 0001, Zhenhua Guo 0001, Jianjiang Feng, Ronghua Liang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2021 | Fast Depth Prediction and Obstacle Avoidance on a Monocular Drone Using Probabilistic Convolutional Neural NetworkabstractRecent studies employ advanced deep convolutional neural networks (CNNs) for monocular depth perception, which can hardly run efficiently on small drones that rely on low/middle-grade GPU(e.g. TX2 and 1050Ti) for computation. In addition, the methods which can effectively and efficiently produce probabilistic depth prediction with a measure of model confidence have not been well studied. The lack of such a method could yield erroneous, sometimes fatal, decisions in drone applications (e.g. selecting a waypoint in a region with a large depth yet a low estimation confidence). This paper presents a real-time onboard approach for monocular depth prediction and obstacle avoidance with a lightweight probabilistic CNN (pCNN), which will be ideal for use in a lightweight energy-efficient drone. For each video frame, our pCNN can efficiently predict its depth map and the corresponding confidence. The accuracy of our lightweight pCNN is greatly boosted by integrating sparse depth estimation from a visual odometry into the network for guiding dense depth and confidence inference. The estimated depth map is transformed into Ego Dynamic Space (EDS) by embedding both dynamic motion constraints of a drone and the confidence values into the spatial depth map. Traversable waypoints are automatically computed in EDS based on which appropriate control inputs for the drone are produced. Extensive experimental results on public datasets demonstrate that our depth prediction method runs at 12Hz and 45Hz on TX2 and 1050Ti GPU respectively, which is 1.8X~5.6X faster than the state-of-the-art methods and achieves better depth estimation accuracy. We also conducted experiments of obstacle avoidance in both simulated and real environments to demonstrate the superiority of our method to the baseline methods. Xin Yang 0008, Yuanjie Dang, Hongcheng Luo, Yuesheng Tang, Chunyuan Liao, Peng Chen 0008, Kwang-Ting Cheng |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2019 | Learning to Film From Professional Human Motion VideosabstractWe investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or adopt ad-hoc solutions without carefully considering the impact of video contents and previous camera motions on the future camera motions. As a result, they can hardly achieve satisfactory results in our drone cinematography task. In this study, we propose a learning-based framework which incorporates the video contents and previous camera motions to predict the future camera motions that enable the capture of professional videos. Specifically, the inputs of our framework are video contents which are represented using subject-related feature based on 2D skeleton and scene-related features extracted from background RGB images, and camera motions which are represented using optical flows. The correlation between the inputs and output future camera motions are learned via a sequence-to-sequence convolutional long short-term memory (Seq2Seq ConvLSTM) network from a large set of video clips. We deploy our approach to a real drone cinematography system by first predicting the future camera motions, and then converting them to the drone's control commands via an odometer. Our experimental results on extensive datasets and showcases exhibit significant improvements in our approach over conventional baselines and our approach can successfully mimic the footage of a professional cameraman. Chong Huang 0005, David Chuan-En Lin, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
CVPR | 5 |
| 2019 | Learning to Capture a Film-Look Video with a Camera DroneabstractThe development of intelligent drones has simplified aerial filming and provided smarter assistant tools for users to capture a film-look footage. Existing methods of autonomous aerial filming either specify predefined camera movements for a drone to capture a footage, or employ heuristic approaches for camera motion planning. However, both predefined movements and heuristically planned motions are hardly able to provide cinematic footages for various dynamic scenarios. In this paper, we propose a data-driven learning-based approach, which can imitate a professional cameraman's intention for capturing a film-look aerial footage of a single subject in real-time. We model the decision-making process of the cameraman with two steps: 1) we train a network to predict the future image composition and camera position, and 2) our system then generates control commands to achieve the desired shot framing. At the system level, we deploy our algorithm on the limited resources of a drone and demonstrate the feasibility of running automatic filming onboard in real-time. Our experiments show how our data-driven planning approach achieves film-look footages and successfully mimics the work of a professional cameraman. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
ICRA | 4 |
| 2018 | ACT: An Autonomous Drone Cinematography System for Action ScenesabstractDrones are enabling new forms of cinematography. Aerial filming via drones in action scenes is difficult because it requires users to understand the dynamic scenarios and operate the drone and camera simultaneously. Existing systems allow the user to manually specify the shots and guide the drone to capture footage, while none of them employ aesthetic objectives to automate aerial filming in action scenes. Meanwhile, these drone cinematography systems depend on the external motion capture systems to perceive the human action, which is limited to the indoor environment. In this paper, we propose an Autonomous CinemaTography system “ACT” on the drone platform to address the above the challenges. To our knowledge, this is the first drone camera system which can autonomously capture cinematic shots of action scenes based on limb movements in both indoor and outdoor environments. Our system includes the following novelties. First, we propose an efficient method to extract 3D skeleton points via a stereo camera. Second, we design a real-time dynamical camera planning strategy that fulfills the aesthetic objectives for filming and respects the physical limits of a drone. At the system level, we integrate cameras and GPUs into the limited space of a drone and demonstrate the feasibility of running the entire cinematography system onboard in real-time. Experimental results in both simulation and real-world scenarios demonstrate that our cinematography system “ACT” can capture more expressive video footage of human action than that of a state-of-the-art drone camera system. Chong Huang 0005, Fei Gao 0011, Jie Pan 0004, Weihao Qiu, Peng Chen 0008, Xin Yang 0008, Shaojie Shen, Kwang-Ting Cheng |
ICRA | 6 |
| 2018 | Through-the-Lens Drone FilmingabstractAerial filming in action scenes using a drone is difficult for inexperienced flyers because manipulating a remote controller and meeting the desired image composition are two independent, while concurrent, tasks. Existing systems attempt to utilize wearable GPS-based or infrared-based sensors to track the human movement and to assist in capturing footage. However, these sensors work only in either indoor (infrared-based) or outdoor environments (GPS-based), but not both. In this paper, we introduce a novel drone filming system which integrates monocular 3D human pose estimation and localization into a drone platform to remove the constraints imposed by wearable-sensor-based solutions. Meanwhile, given the estimated position, we propose a novel drone control system, called “through-the-lens drone filming”, to allow a cameraman to conveniently control the drone by manipulating a 3D model in the preview, which closes the gap between the flight control and the viewpoint design. Our system includes two key enabling techniques: 1) subject localization based on visual-inertial fusion, and 2) through-the-lens camera planning. This is the first drone camera system which allows users to capture human actions by manipulating the camera in a virtual environment. From the drone hardware, we integrate a gimbal camera and two GPUs into the limited space of a drone and demonstrate the feasibility of running the entire system onboard with insignificant delays, which are sufficient for filming in our real-time application. Experimental results, in both simulation and real-world scenarios, demonstrate that our techniques can greatly ease camera control and capture better videos. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 4 |
| 2018 | Real-Time Object Tracking on a Drone With Multi-Inertial Sensing DataabstractReal-time object tracking on a drone under a dynamic environment has been a challenging issue for many years, with existing approaches using off-line calculation or powerful computation units on board. This paper presents a new lightweight real-time onboard object tracking approach with multi-inertial sensing data, wherein a highly energy-efficient drone is built based on the Snapdragon flight board of Qualcomm. The flight board uses a digital signal processor core of the Snapdragon 801 processor to realize PX4 autopilot, an open-source autopilot system oriented toward inexpensive autonomous aircraft. It also uses an ARM core to realize Linux, robot operating systems, open-source computer vision library, and related algorithms. A lightweight moving object detection algorithm is proposed that extracts feature points in the video frame using the oriented FAST and rotated binary robust independent elementary features algorithm and adapts a local difference binary algorithm to construct the image binary descriptors. The K-nearest neighbor method is then used to match the image descriptors. Finally, an object tracking method is proposed that fuses inertial measurement unit data, global positioning system data, and the moving object detection results to calculate the relative position between coordinate systems of the object and the drone. All the algorithms are run on the Qualcomm platform in real time. Experimental results demonstrate the superior performance of our method over the state-of-the-art visual tracking method. Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Wei Zhu 0006, Xiaofei He 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | REDBEE: A visual-inertial drone system for real-time moving object detectionabstractAerial surveillance and monitoring demand both real-time and robust motion detection from a moving camera. Most existing techniques for drones involve sending a video data streams back to a ground station with a high-end desktop computer or server. These methods share one major drawback: data transmission is subjected to considerable delay and possible corruption. Onboard computation can not only overcome the data corruption problem but also increase the range of motion. Unfortunately, due to limited weight-bearing capacity, equipping drones with computing hardware of high processing capability is not feasible. Therefore, developing a motion detection system with real-time performance and high accuracy for drones with limited computing power is highly desirable. In this paper, we propose a visual-inertial drone system for real-time motion detection, namely REDBEE, that helps overcome challenges in shooting scenes with strong parallax and dynamic background. REDBEE, which can run on the state-of-the-art commercial low-power application processor (e.g. Snapdragon Flight board used for our prototype drone), achieves real-time performance with high detection accuracy. The REDBEE system overcomes obstacles in shooting scenes with strong parallax through an inertial-aided dual-plane homography estimation; it solves the issues in shooting scenes with dynamic background by distinguishing the moving targets through a probabilistic model based on spatial, temporal, and entropy consistency. The experiments are presented which demonstrate that our system obtains greater accuracy when detecting moving targets in outdoor environments than the state-of-the-art real-time onboard detection systems. Chong Huang 0005, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 2 |
| 2014 | Fast macroblock encoding algorithm based on rate-distortion activity for multiview video coding
Wei Zhu 0006, Yayu Zheng, Peng Chen 0008, Jie Feng 0010 |
Signal Process. Image Commun. | 3 |
| 2011 | An Adaptive Inter Mode Decision for Multiview Video CodingabstractMultiview video coding (MVC) plays an important role in 3D video system, while the huge computational complexity blocks its applications. This paper proposes an adaptive Inter mode decision algorithm to reduce the complexity of MVC. First, the selection of Inter modes is determined by using the textural region type of macro block (MB). Then, the estimation of small size Inter modes (Inter16×8, Inter8×16, and Inter8×8) is decided based on the motion homogenization of MB, which is predicted by utilizing the motion estimation results of Inter16×16 mode. Finally, the complexity of Inter8×8 mode estimation is progressively reduced by employing rate-distortion (RD) costs of estimated modes. As compared to the full mode decision in MVC reference software, the proposed algorithm achieved 71% encoding time saving on average with 0.026 dB peak signal-to-noise ratio loss and 0.74% bit rate increase. Wei Zhu 0006, Peng Chen 0008, Yayu Zheng, Jie Feng 0010 |
ISM | 2 |
| 2010 | Optimized simulated annealing algorithm for thinning and weighting large planar arraysabstractThis paper proposes an optimized simulated annealing (SA) algorithm for thinning and weighting large planar arrays in 3D underwater sonar imaging systems. The optimized algorithm has been developed for use in designing a 2D planar array (a rectangular grid with a circular boundary) with a fixed side-lobe peak and a fixed current taper ratio under a narrow-band excitation. Four extensions of the SA algorithm and the procedure for the optimized SA algorithm are described. Two examples of planar arrays are used to assess the efficiency of the optimized method. The proposed method achieves a similar beam pattern performance with fewer active transducers and faster convergence ability than previous SA algorithms. Peng Chen 0008, Bin-jian Shen, Li-sheng Zhou, Yaowu Chen |
J. Zhejiang Univ. Sci. C | 1 |