VLDB 2026 Research / reviewers in the wild / expert
Chenglizhao Chen
dblp:161/2811
· DBLP profile ↗
94ranked-venue papers
23as first author
78since 2021 · last 2026
0000-0001-9982-5667ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 19 first-author · 48 since 2021Artificial intelligence and machine learning · 38 · 4 first-author · 35 since 2021Computer networks · 5 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AerialMind: Towards Referring Multi-Object Tracking in UAV ScenariosabstractReferring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research remains mostly confined to ground-level scenarios, which constrains their ability to capture broad-scale scene contexts and perform comprehensive tracking and path planning. In contrast, Unmanned Aerial Vehicles (UAVs) leverage their expansive aerial perspectives and superior maneuverability to enable wide-area surveillance. Moreover, UAVs have emerged as critical platforms for Embodied Intelligence, which has given rise to an unprecedented demand for intelligent aerial systems capable of natural language interaction. To this end, we introduce AerialMind, the first large-scale RMOT benchmark in UAV scenarios, which aims to bridge this research gap. To facilitate its construction, we develop an innovative semi-automated collaborative agent-based labeling assistant (COALA) framework that significantly reduces labor costs while maintaining annotation quality. Furthermore, we propose HawkEyeTrack (HETrack), a novel method that collaboratively enhances vision-language representation learning and improves the perception of UAV scenarios. Comprehensive experiments validated the challenging nature of our dataset and the effectiveness of our method. Chenglizhao Chen, Shaofeng Liang, Runwei Guan, Xiaolou Sun, Haocheng Zhao, Haiyun Jiang, Tao Huang 0008, Henghui Ding, Qing-Long Han |
AAAI | 1 |
| 2026 | Lithology identification from missing well log data: a multi-constraint guided self-supervised diffusion frameworkabstractAccurate lithology identification from well log data is fundamental for reservoir characterization, yet its practical application is often impeded by the dual challenges of incomplete well log data and scarce labeled samples. Existing data-driven methods often address these issues separately or require extensive complete labeled data, limiting their effectiveness in low-resource geological exploration scenarios. To address this compound challenge, this paper proposes a multi-constraint guided self-supervised diffusion framework, designated as MCG-SSDF, based on a generative pre-training paradigm. The framework first employs a novel multi-constraint guided diffusion model, MCG-Diff, for unsupervised pre-training on large-scale unlabeled data. This model integrates a Mamba backbone with three geological priors, namely a global structural constraint, a petrophysical correlation constraint, and a geological morphological constraint, to enhance the fidelity and physical plausibility of well log data generation and imputation. Subsequently, a parameter-efficient fine-tuning strategy is applied, where only a lightweight classification head is trained on a small set of labeled, incomplete samples, enabling the framework to perform end-to-end lithology identification directly from incomplete data during the inference stage. Systematic evaluations on two oilfield datasets validate the framework across deterministic accuracy and uncertainty quantification. Results reveal that MCG-SSDF outperforms baselines, particularly under extreme label scarcity. Ablation studies and attribution analyses further confirm the synergistic contributions and decision reliability of integrated geological constraints. This methodology provides a robust solution for high-precision lithology identification in low-resource environments. Qingwei Pang, Chenglizhao Chen |
Adv. Eng. Informatics | 2 |
| 2026 | Weakly supervised visual-auditory fixation prediction with multigranularity perception
Guotao Wang 0004, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, Qinping Zhao |
Sci. China Inf. Sci. | 2 |
| 2026 | Dual-domain parallel attention-driven underwater image enhancement method
Tianmeng Sun, Haiyuan Cui, Jieru Chi, Guowei Yang 0002, Chenglizhao Chen |
Eng. Appl. Artif. Intell. | 7 |
| 2026 | PM-SSL: Physics-guided multimodal self-supervised learning for lithology identification with scarce labels
Qingwei Pang, Qingyu Pang, Chenglizhao Chen |
Neurocomputing | 3 |
| 2026 | Fine-grained tensor completion for incomplete multi-view clustering
Chong Peng 0001, Chundan Liu, Yongyong Chen, Zhao Kang 0001, Junyu Dong, Guiyuan Jiang, Chenglizhao Chen |
Pattern Recognit. | 8 |
| 2026 | Adapting Visual Trackers to Dynamic View Transitions With Shift-View Prompt TuningabstractVisual tracking is essential across numerous video analysis applications, surveillance systems, entertainment, and autonomous applications. However, most conventional state-of-the- art visual trackers are designed for constant-view scenarios with fixed camera viewpoints, and they only achieve satisfactory performance under stable visual features scenarios. In reality, visual tracking often encounters shift-view scenarios (e.g., sports broadcasting, ground-aerial surveillance), where cameras’ dynamic view transitions between ground and aerial views. These shifts lead to large variations in target scale and environmental complexity, resulting in inconsistent visual features that ultimately degrade the robustness of conventional visual trackers. Although developing a dedicated tracker for such shift-view scenarios is possible, it requires expensive temporal and computational costs. To address this challenge, we propose Shift-view Prompt Tuning, a cost-efficient method that enables conventional trackers to handle dynamic view transitions. We use sample pairs from different view datasets as prompts to guide the tracker’s adaptation. By embedding distinctive visual information from these prompts into training samples, we help the tracker learn about dynamic view transitions without requiring it to be relearned from scratch. This approach seamlessly transforms any constant-view trackers into shift-view trackers. Our extensive experiments on 14 datasets with 3 different view types show that our approach significantly enhances tracking performance. This advancement extends the application scope of current trackers and offers a robust solution for multimedia content production, sports analytics, and security monitoring in video analysis systems. Chenglizhao Chen, Shaofeng Liang, Luming Li, Mengke Song, Xu Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Domain Adaptation for Cross-View Localization via Multi-Teacher Knowledge DistillationabstractCross-view fine-grained localization estimates a ground camera’s pixel-level coordinates in aerial images by analyzing visual correspondences between views. Recent studies have made significant progress in this task, but when the models trained in a source area are directly applied to a new target area, their localization performance often suffers significant degradation due to the domain gap between the two areas. Moreover, obtaining accurate ground truth (GT) for the target area to retrain the models is prohibitively expensive. To adapt the localization model to the target area, this article proposes a weakly supervised learning approach based on multi-teacher knowledge distillation. This approach utilizes multiple pre-trained teacher models to make predictions for the target area and employs a learning-free cross-view instance matching and view alignment (CVMA) module to evaluate the quality of predicted coordinates from geometric, semantic, and visual perspectives. Based on the evaluation results, the best prediction is selected as pseudo-GT, and potential anomalous training samples are filtered out. The CVMA module also functions as a learning-free fine-grained localization method, achieving performance comparable to some learning-based methods. Our approach is validated on the VIGOR benchmark using three state-of-the-art models, and experimental results show that our method significantly improves the localization performance of models in the target area. Chenglizhao Chen, Qianxi Yuan, Shaofeng Liang, Mengke Song, Xinyu Liu 0029 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2026 | Channel-Wise Contribution Assessment for RGB-D Salient Object DetectionabstractIn RGB-D salient object detection (SOD), a common approach to improving accuracy is by using a dual-stream architecture to combine RGB and depth data. However, the effectiveness of depth information varies depending on the scene. In scenarios where depth maps provide a limited contribution, integrating them with RGB can be challenging and sometimes even detrimental to performance rather than enhancing it. Conventional RGB-D SOD methods often lack precision in assessing depth map quality, neglecting to account for the distinct contributions of various regions within the map, often resulting in a suboptimal fusion of regions where depth information is minimally beneficial or irrelevant. To address these issues, this article presents a novel channel-wise contribution assessment method that precisely evaluates the contributions of both the RGB and depth channels. By employing a controlled perturbation process to challenge the saliency detection model with specific, manageable disturbances, we are able to measure how much RGB and depth information each contributes to the final saliency map. Based on this analysis, we have developed a novel routing-style fusion of modality that dynamically adjusts the integration of the two modalities. This approach significantly lessens the negative impact of regions where depth data have a low, no, or even detrimental contribution, leading to a more effective and balanced fusion of RGB and depth information. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently achieves competitive performance and improves the robustness of RGB-D salient object detection across diverse and challenging scenarios. Chenglizhao Chen, Mengke Song, Xinyu Liu 0029, Wenfeng Song |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Fine-Grained Perception in Panoramic Scenes: A Novel Task, Dataset, and Method for Object Importance RankingabstractExisting Salient Object Ranking (SOR) aims to infer ranking of salient objects based on their saliency degree. However, it tends to only focus on salient objects while neglecting non-salient ones. This coarse-grained ranking limits the performance of downstream tasks. For instance, in image retrieval tasks, focusing solely on the relationship between salient objects is insufficient for achieving fine-grained scene analysis, which may result in retrieved results that do not satisfy user requirements. High-quality retrieval requires fine-grained analysis, making it essential to rank non-salient objects. Based on this need, we propose a new task: Fine-grained Object Importance Ranking in 360 Scenes (FOIR-360), which focus on predicting the relative importance of "ALL objects'' at the instance-level. Our task takes into account all objects, allowing us to refine the original "coarse-grained'' to a "fine-grained'' level. Currently, the main challenge for this new task is the lack of supervised data for model training or even for model testing. Therefore, we propose a novel weakly supervised method to address the shortage of datasets. Furthermore, to the best of our knowledge, there is no existing suitable annotation protocol for this new task. The main reason is that annotating fine-grained rankings is extremely difficult, especially in panoramic scenes that contain numerous instances where even humans are unable to determine which one is more important than others. As the first attempt, we introduce a new annotation protocol designed to highlight the ranking of objects that are non-salient yet still important. Based on this protocol, we construct the first fine-grained 360Rank dataset. In summary, all these new task, weakly supervised method, annotation protocol, and dataset have the potential to drive advancements in the field. Chenglizhao Chen, Xu Yu 0001 |
AAAI | 2 |
| 2025 | SharpEdge: High-quality data-driven monocular depth estimation for enhanced boundary precision
Mengke Song, Luming Li, Xu Yu 0001, Chenglizhao Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Multi-domain masked reconstruction self-supervised learning for lithology identification using well-logging data
Qingwei Pang, Chenglizhao Chen |
Knowl. Based Syst. | 2 |
| 2025 | Dual-level semantic collaboration and inference network for medical image report generation
Junsan Zhang, Yuxue Liu, Ming-Wen Shao, Chenglizhao Chen, Zixuan Wang 0012, Yao Wan 0001, Philip S. Yu |
Knowl. Based Syst. | 4 |
| 2025 | Prior-based bi-encoder transformer for underwater image enhancement
Jinqiang Yan, Haiyuan Cui, Jieru Chi, Guowei Yang 0002, Chenglizhao Chen |
Multim. Syst. | 7 |
| 2025 | Saliency-Free and Aesthetic-Aware Panoramic Video NavigationabstractMost of the existing panoramic video navigation approaches are saliency-driven, whereby off-the-shelf saliency detection tools are directly employed to aid the navigation approaches in localizing video content that should be incorporated into the navigation path. In view of the dilemma faced by our research community, we rethink if the "saliency clues" are really appropriate to serve the panoramic video navigation task. According to our in-depth investigation, we argue that using "saliency clues" cannot generate a satisfying navigation path, failing to well represent the given panoramic video, and the views in the navigation path are also low aesthetics. In this paper, we present a brand-new navigation paradigm. Although our model is still trained on eye-fixations, our methodology can additionally enable the trained model to perceive the "meaningful" degree of the given panoramic video content. Outwardly, the proposed new approach is saliency-free, but inwardly, it is developed from saliency but biasing more to be "meaningful-driven"; thus, it can generate a navigation path with more appropriate content coverage. Besides, this paper is the first attempt to devise an unsupervised learning scheme to ensure all localized meaningful views in the navigation path have high aesthetics. Thus, the navigation path generated by our approach can also bring users an enjoyable watching experience. As a new topic in its infancy, we have devised a series of quantitative evaluation schemes, including objective verifications and subjective user studies. All these innovative attempts would have great potential to inspire and promote this research field in the near future. Chenglizhao Chen, Guangxiao Ma, Wenfeng Song, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | WinDB: HMD-Free and Distortion-Free Panoptic Video Fixation LearningabstractTo date, the widely adopted way to perform fixation collection in panoptic video is based on a head-mounted display (HMD), where users' fixations are collected while wearing a HMD to explore the given panoptic scene freely. However, this widely-used data collection method is insufficient for training deep models to accurately predict which regions in a given panoptic are most important when it contains intermittent salient events. The main reason is that there always exist "blind zooms" when using HMD to collect fixations since the users cannot keep spinning their heads to explore the entire panoptic scene all the time. Consequently, the collected fixations tend to be trapped in some local views, leaving the remaining areas to be the "blind zooms". Therefore, fixation data collected using HMD-based methods that accumulate local views cannot accurately represent the overall global importance - the main purpose of fixations - of complex panoptic scenes. To conquer, this paper introduces the auxiliary window with a dynamic blurring (WinDB) fixation collection approach for panoptic video, which doesn't need HMD and is able to well reflect the regional-wise importance degree. Using our WinDB approach, we have released a new PanopticVideo-300 dataset, containing 300 panoptic clips covering over 225 categories. Specifically, since using WinDB to collect fixations is blind zoom free, there exists frequent and intensive "fixation shifting" - a very special phenomenon that has long been overlooked by the previous research - in our new set. Thus, we present an effective fixation shifting network (FishNet) to conquer it. All these new fixation collection tool, dataset, and network could be very potential to open a new age for fixation-related research and applications in 360o environments. Guotao Wang 0004, Chenglizhao Chen, Aimin Hao, Hong Qin 0001, Deng-Ping Fan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Pixel is All You Need: Adversarial Spatio-Temporal Ensemble Active Learning for Salient Object DetectionabstractAlthough weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial spatio-temporal ensemble active learning. Our contributions are four-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed spatio-temporal ensemble strategy not only achieves outstanding performance but significantly reduces the model's computational cost. 3) Our proposed relationship-aware diversity sampling can conquer oversampling while boosting model performance. 4) We provide theoretical proof for the existence of such a point-labeled dataset. Experimental results show that our approach can find such a point-labeled dataset, where a saliency model trained on it obtained 98%-99% performance of its fully-supervised version with only ten annotated points per image. Wei Wang 0169, Yacong Li, Fengmao Lv, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Omni-Directional View Person Re-Identification Through 3D Human ReconstructionabstractPerson re-identification (ReID) aims to identify the same individual across different cameras. Most existing researches focus on horizontal perspectives, where cameras and individuals are positioned at similar heights. However, in real-word applications, cameras are usually mounted at varying heights (e.g., either high-view or low-view) to achieve a broader field of view. Hence, some studies have explored high-view ReID, yet these rely heavily on manually annotating large datasets, which is extremely time-consuming and not publicly available. To improve, we propose a “controllable” data generation protocol that automatically generates omni-directional view data. This protocol can extend any common ReID dataset into an extensive omni-directional view one. By upgrading existing ReID SOTAs with the enhanced data, they can be made to handle ReID tasks with varying camera angles. B.t.w., to verify the effectiveness, we still need “real” data for testing. Thus, we constructed a small testing dataset containing diverse camera angles. Extensive quantitative results demonstrate that our solution is generic and can be applied to any SOTA ReID to achieve extensive performance promotions,e.g., 3%-12% improvement in mAP. Chenglizhao Chen, Chaoying Bai, Xu Yu 0001 |
IEEE Signal Process. Lett. | 1 |
| 2025 | Unveiling Context-Related Anomalies: Knowledge Graph Empowered Decoupling of Scene and Action for Human-Related Video Anomaly DetectionabstractVideo anomaly detection methods are mainly classified into two categories based on their primary feature types: appearance-based and action-based. Appearance-based methods rely on low-level visual features like color, texture, and shape, learning patterns specific to training scenes. While effective in familiar settings, they struggle with unknown or altered scenes due to poor generalization and limited understanding of action-scene relationships. In contrast, action-based methods focus on detecting action anomalies but often overlook contextual scene associations, leading to misjudgments (e.g., running on a street being deemed normal without considering scene context). To overcome these limitations, we propose a novel decoupling-based anomaly detection architecture (DecoAD). Its core lies in the decoupling and interweaving of scenes and actions, enabling explicit modeling of their complex relationships. By reconstructing these interactions using knowledge graphs, DecoAD achieves a deeper understanding of behaviors and contexts. This design ensures strong performance in both known and unknown scenarios, significantly enhancing generalization. To evaluate its effectiveness in dynamic scenes and its ability to handle scene-related anomalies, we introduce UFSR, the first video anomaly detection dataset featuring dynamic scenes and scene-related anomalies. DecoAD supports fully-supervised, weakly-supervised, and unsupervised settings, improving AUC on UBnormal by 1.1%, 3.1%, and 2.1% in fully-supervised, weakly-supervised, and unsupervised settings, and on UFSR by 1.2% and 8.2% in weakly-supervised and unsupervised settings. The source code and datasets are available at:https://github.com/liuxy3366/DecoAD. Chenglizhao Chen, Xinyu Liu 0029, Mengke Song, Luming Li, Shaojiang Yuan, Xu Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | UNI-IQA: A Unified Approach for Mutual Promotion of Natural and Screen Content Image Quality AssessmentabstractTo date, the image quality assessment (IQA) research field has mainly focused on natural images (NIs)-based IQA and screen content images (SCIs)-based IQA. Usually, these two research branches are quite independent due to the large differences between NIs and SCIs, where NIs, captured by cameras directly, contain pictorial information solely, yet, SCIs, synthesized or GPU-rendered, have pictures and textures. Moreover, the distortion types are also different, and subjective scores of different datasets assigned by participants are usually not well aligned. So, due to the above-mentioned “domain shifts” and “dataset misalignments”, our research community has widely believed that it could be very difficult to achieve joint mutual promotions between NIs- and SCIs-based IQA. In this paper, we argue that despite the “differences”, there still are some “common characteristics” — our human visual system perceives the “pictures” in both SCIs and NIs almost the same way. Thus, we can still achieve mutual performance promotion if we can appropriately use the “common characteristics” between SCIs and NIs. Our key idea is to devise a “content-aware” data switch, which, from the perspective of input’s contents (i.e., pictures or textures), aims at letting the model automatically enhance the commonness and compress the discrepancies between the two tasks. Notice that none of the existing fusion schemes can reach this goal since they are actually content-unaware, degenerating the “mutual interactions” into “mutual interferences”. This paper is the first attempt to achieve full end-to-end “mutual interactions” between NIs- and SCIs-based IQA. Using the proposed switch, we are also the first to achieve solid mutual promotions for the two tasks, reaching new SOTA results. Mengke Song, Chenglizhao Chen, Wenfeng Song, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Pushing the Boundaries of Salient Object Detection: A Denoising-Driven ApproachabstractSalient Object Detection (SOD) aims to identify the most attention-grabbing regions in an image and focuses on distinguishing salient objects from their backgrounds. Current SOD methods primarily use a discriminative approach, which works well for clear images but struggles in complex scenes with similar colors and textures between objects and backgrounds. To address these limitations, we introduce the diffusion-based salient object detection model (DiffSOD), which leverages a noise-to-image denoising process within a diffusion framework, enhancing saliency detection in both RGB and RGB-D images. Unlike conventional fusion-based SOD methods that directly merge RGB and depth information, we treat RGB and depth as distinct conditions, i.e., the appearance condition and the structure condition, respectively. These conditions serve as controls within the diffusion UNet architecture, guiding the denoising process. To facilitate this guidance, we employ two specialized control adapters: the appearance control adapter and the structure control adapter. Moreover, conventional denoising UNet models may struggle when handling low-quality depth maps, potentially introducing detrimental cues into the denoising process. To mitigate the impact of low-quality depth maps, we introduce a quality-aware filter. This filter selectively processes only high-quality depth data, ensuring that the denoising process is based on reliable information. Comparative evaluations on benchmark datasets have shown that DiffSOD substantially surpasses existing RGB and RGB-D saliency detection methods, improving average performance by 1.5% and 1.2% respectively, thus setting a new benchmark for diffusion-based dense prediction models in visual saliency detection. Mengke Song, Luming Li, Xu Yu 0001, Chenglizhao Chen |
IEEE Trans. Image Process. | 4 |
| 2025 | Adapting Generic RGB-D Salient Object Detection for Specific Traffic ScenariosabstractExisting RGB-D salient object detection (SOD) models are primarily trained on general-purpose datasets, which may lead to domain shift issues when applied directly to new, specific scenes, such as stereo traffic datasets. Though “large-scale datasets (COME15K and ReDweb-S)” have been released, they only partially address the domain shift problem. From the perspective of data augmentation, this paper presents a novel solution, which follows a weakly-supervised way to adapt generic RGB-D SOD models for specific scenarios, with a focus on traffic scene imagery. Our key idea is to equip plain videos (specific scenarios, i.e., traffic scenes) with newly estimated saliency informative depth maps and pseudo-SOD GTs, enabling them to support the retraining of existing RGB-D SOD models for meeting the requirements of these specific scenes. To achieve this, we offer a fresh perspective on how depth information can be leveraged in the SOD task and introduce a new paradigm for extracting intrinsic information from optical flows derived from videos to refine RGB-D SOD models. Our method achieves a 1.2% improvement in F-measure on RGB-D datasets and a 27% enhancement on real-world street view datasets compared to baseline models. These results demonstrate the effectiveness of our approach in enhancing model adaptability for traffic scene imagery, even with limited target domain data. Codes, datasets, and results are available at https://github.com/MengkeSong/AGSS. Chenglizhao Chen, Mengke Song, Chong Peng 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Tropical Cyclone Image Super-Resolution via Multimodality FusionabstractThe traditional super-resolution dataset construction using artificial down-sampling techniques can result in information loss, insufficient diversity, and non-uniqueness. Furthermore, existing methods for image super-resolution are limited to single-modal images and cannot accommodate the complexities of multimodal images. This is problematic because diverse modal data requires individualized model design and training, which can hinder the exploitation of complementary relationships among multimodal data. In this article, we have addressed these issues by undertaking a two-step solution approach. In the first step, we constructed a super-resolution dataset that utilized remote-sensing images of tropical cyclones in “real cases.” This dataset comprises HR–LR image pairs originating from multiple sensors of varying satellite sources, resulting in multimodal data. However, the HR–LR image pairs suffer from an additional misalignment issue. Thus, in the second step, we designed a super-resolution network based on MAT to address the misalignment problem in multimodal environment. After numerous ablation experiments and comparison experiments, we have shown that our model is effective, with an improvement of 50% over the original baseline model, and an increase varying between 20% and 50% compared to other common super-resolution models. We have made our source code and data publicly available online at https://github.com/kleenY/MMTCSR . Tao Song 0001, Fan Meng 0008, Xin Li 0244, Handan Sun, Chenglizhao Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | A Comprehensive Survey on 3D Single-View Object ReconstructionabstractSingle-view 3D object reconstruction (SVOR) aims to recover the 3D shape of an object from a single 2D image. Despite advances in deep learning (DL), challenges such as incomplete image information, scarce 3D data annotation, and highly variable object shapes still limit the performance of SVOR. Meanwhile, with the rapid development of novel view synthesis (NVS) techniques, the SVOR field has received significant advancements. However, existing reviews have not comprehensively covered the rapid developments in NVS-based approaches. This article aims to fill this gap by highlighting the latest progress in SVOR, particularly advancements related to NVS-based methods. Additionally, we observed discrepancies between existing quality evaluation metrics in SVOR and human visual perception. This is because some critical object parts are essential to consider during the evaluation. For example, when reconstructing airplanes, critical parts like the empennage and wings are often overlooked in evaluation metrics due to their smaller size compared to the fuselage. Consequently, poor reconstruction of these parts may not significantly affect overall evaluation scores. To address this issue, we propose a more comprehensive evaluation method that reflects human visual perception accurately. To achieve this, we introduce a weighted evaluation method that considers part saliency and proposes a novel technique for automatically perceiving reconstruction discrepancies. This study effectively enhances the accuracy and consistency of evaluations through these approaches, offering new insights and methodologies, filling a void in the existing literature, and providing valuable contributions to both research and practical applications in SVOR. Chenglizhao Chen, Ziyue Xue, Longyan Yang, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | Fine-Grained Bipartite Concept Factorization for ClusteringabstractIn this paper, we propose a novel concept factorization method that seeks factor matrices using a cross-order positive semi-definite neighbor graph, which provides comprehensive and complementary neighbor information of the data. The factor matrices are learned with bipartite graph partitioning, which exploits explicit cluster structure of the data and is more geared towards clustering application. We develop an effective and efficient optimization algorithm for our method, and provide elegant theoretical results about the convergence. Extensive experimental results confirm the effectiveness of the proposed method. Chong Peng 0001, Pengfei Zhang 0016, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
CVPR | 5 |
| 2024 | Arbitrary Motion Style Transfer with Multi-Condition Motion Latent Diffusion ModelabstractComputer animation's quest to bridge content and style has historically been a challenging venture, with previous efforts often leaning toward one at the expense of the other. This paper tackles the inherent challenge of content-style duality, ensuring a harmonious fusion where the core narrative of the content is both preserved and elevated through stylistic enhancements. We propose a novel Multi-condition Motion Latent Diffusion Model (MCM-LDM) for Arbitrary Motion Style Transfer (AMST). Our MCM-LDM significantly emphasizes preserving trajectories, recognizing their fundamental role in defining the essence and fluidity of motion content. Our MCM-LDM's cornerstone lies in its ability first to disentangle and then intricately weave together motion's tripartite components: motion trajectory, motion content, and motion style. The critical insight of MCM-LDM is to embed multiple conditions with distinct priorities. The content channel serves as the primary flow, guiding the overall structure and movement, while the trajectory and style channels act as auxiliary components and synchronize with the primary one dynamically. This mechanism ensures that multi-conditions can seamlessly integrate into the main flow, enhancing the overall animation without overshadowing the core content. Empirical evaluations underscore the model's proficiency in achieving fluid and authentic motion style transfers, setting a new benchmark in the realm of computer animation. The source code and model are available at https://github.com/XingliangJin/MCM-LDM.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou, Hong Qin 0001 |
CVPR | 4 |
| 2024 | HOIAnimator: Generating Text-Prompt Human-Object Animations Using Novel Perceptive Diffusion ModelsabstractTo date, the quest to rapidly and effectively produce human-object interaction (HOI) animations directly from textual descriptions stands at the forefront of computer vision research. The underlying challenge demands both a discriminating interpretation of language and a comprehen-sive physics-centric model supporting real-world dynamics. To ameliorate, this paper advocates HOIAnimator, a novel and interactive diffusion model with perception ability and also ingeniously crafted to revolutionize the animation of complex interactions from linguistic narratives. The effectiveness of our model is anchored in two ground-breaking innovations: (1) Our Perceptive Diffusion Models (PDM) brings together two types of models: one focused on hu-man movements and the other on objects. This combination allows for animations where humans and objects move in concert with each other, making the overall motion more realistic. Additionally, we propose a Perceptive Message Passing (PMP) mechanism to enhance the communication bridging the two models, ensuring that the animations are smooth and unified; (2) We devise an Interaction Contact Field (ICF), a sophisticated model that implicitly captures the essence of HOls. Beyond mere predictive contact points, the ICF assesses the proximity of human and object to their respective environment, informed by a probabilistic distribution of interactions learned throughout the denoising phase. Our comprehensive evaluation showcases HOlani-mator's superior ability to produce dynamic, context-aware animations that surpass existing benchmarks in text-driven animation synthesis. Wenfeng Song, Shuai Li 0001, Yang Gao 0032, Aimin Hao, Xia Hau, Chenglizhao Chen, Hong Qin 0001 |
CVPR | 7 |
| 2024 | Cross-View Diversity Embedded Consensus Learning for Multi-View Clustering
Chong Peng 0001, Kai Zhang 0008, Yongyong Chen, Chenglizhao Chen, Qiang Shawn Cheng |
IJCAI | 4 |
| 2024 | Instance-Level Data Augmentation for Multi-Person Pose Estimation: Improving Recognition of Individuals at Different ScalesabstractIn the realm of multi-person pose estimation, bottom-up approaches often tackle the task of identifying human keypoints for individuals at various scales within a given image. However, in practical scenarios, algorithms tend to perform better in recognizing larger individuals. This is primarily attributed to the increased pixel count and richer feature information available. Conversely, recognizing smaller-scale individuals poses a notably more challenging task.To address this challenge, we propose an instance-level data augmentation strategy. This strategy involves applying transformations to individual instances rather than the entire image. Its primary objectives are to enhance dataset diversity, refine the distribution of different human scale samples in the training data, and augment the representation of medium-sized human instances in the training set. The goal of this augmentation strategy is to empower the model to better recognize finer details.Our extensive experiments, conducted on the HigherHRNet benchmark model, demonstrate the effectiveness of our approach in improving accuracy, particularly in the recognition of mediumsized individuals. Importantly, these improvements are achieved without introducing additional model complexity or requiring additional image collection. Yangqi Liu, Guodong Wang 0001, Chenglizhao Chen |
IJCNN | 3 |
| 2024 | Instance-aware image dehazing
Qingqing Chao, Jinqiang Yan, Tianmeng Sun, Silong Li, Jieru Chi, Guowei Yang 0002, Chenglizhao Chen |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | A novel bi-stream network for image dehazing
Qiaoyu Ma, Guowei Yang 0002, Chenglizhao Chen |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Semantic and style based multiple reference learning for artistic and general image aesthetic assessment
Tengfei Shi, Chenglizhao Chen, Aimin Hao |
Neurocomputing | 2 |
| 2024 | Fine-Grained Essential Tensor Learning for Robust Multi-View Spectral ClusteringabstractMulti-view subspace clustering (MVSC) has drawn significant attention in recent study. In this paper, we propose a novel approach to MVSC. First, the new method is capable of preserving high-order neighbor information of the data, which provides essential and complicated underlying relationships of the data that is not straightforwardly preserved by the first-order neighbors. Second, we design log-based nonconvex approximations to both tensor rank and tensor sparsity, which are effective and more accurate than the convex approximations. For the associated shrinkage problems, we provide elegant theoretical results for the closed-form solutions, for which the convergence is guaranteed by theoretical analysis. Moreover, the new approximations have some interesting properties of shrinkage effects, which are guaranteed by elegant theoretical results. Extensive experimental results confirm the effectiveness of the proposed method. Chong Peng 0001, Kehan Kang, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
IEEE Trans. Image Process. | 5 |
| 2024 | Rethinking Object Saliency Ranking: A Novel Whole-Flow Processing ParadigmabstractExisting salient object detection methods are capable of predicting binary maps that highlight visually salient regions. However, these methods are limited in their ability to differentiate the relative importance of multiple objects and the relationships among them, which can lead to errors and reduced accuracy in downstream tasks that depend on the relative importance of multiple objects. To conquer, this paper proposes a new paradigm for saliency ranking, which aims to completely focus on ranking salient objects by their "importance order". While previous works have shown promising performance, they still face ill-posed problems. First, the saliency ranking ground truth (GT) orders generation methods are unreasonable since determining the correct ranking order is not well-defined, resulting in false alarms. Second, training a ranking model remains challenging because most saliency ranking methods follow the multi-task paradigm, leading to conflicts and trade-offs among different tasks. Third, existing regression-based saliency ranking methods are complex for saliency ranking models due to their reliance on instance mask-based saliency ranking orders. These methods require a significant amount of data to perform accurately and can be challenging to implement effectively. To solve these problems, this paper conducts an in-depth analysis of the causes and proposes a whole-flow processing paradigm of saliency ranking task from the perspective of "GT data generation", "network structure design" and "training protocol". The proposed approach outperforms existing state-of-the-art methods on the widely-used SALICON set, as demonstrated by extensive experiments with fair and reasonable comparisons. The saliency ranking task is still in its infancy, and our proposed unified framework can serve as a fundamental strategy to guide future work. The code and data will be available at https://github.com/MengkeSong/Saliency-Ranking-Paradigm. Mengke Song, Dunquan Wu, Wenfeng Song, Chenglizhao Chen |
IEEE Trans. Image Process. | 5 |
| 2024 | Improving Image Aesthetic Assessment via Multiple Image Joint LearningabstractImage Aesthetic Assessment (IAA) is an emerging paradigm that predicts aesthetic score as the popular aesthetic taste for an image. Previous IAA approaches take a single image as input to predict the aesthetic score of the image. However, we discover that most existing IAA methods fail dramatically to predict the images with a large variance of aesthetic voting distribution. Motivated by the practice that people consider similar experiences to improve the consistence of the voting result, we present a novel Multiple Image joint Learning Network (MILNet) to mimic this natural process. Our novelty is mainly three-fold: (a) Semantic-based retrieval method that constructs aesthetic similarity (the similarity of aesthetic attribution) to select reference images; (b) Graph network reasoning that initializes and updates the weight of intrinsic relationships among multiple images; (c) Adaptive Earth Mover’s Distance (AdaEMD) loss function that adjusts weight for easy and hard instances to mitigate unbalanced distribution of aesthetic datasets. Our evaluation with the benchmark AVA and TAD datasets demonstrates that the proposed MILNet outperforms state-of-the-art IAA methods. The code is available at https://github.com/flyingbird93/MILNet . Tengfei Shi, Chenglizhao Chen, Aimin Hao, Yuming Fang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object DetectionabstractAlthough weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image. Wei Wang 0169, Qing Xia 0002, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
AAAI | 5 |
| 2023 | RGB and LUT based Cross Attention Network for Image Enhancement
Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao |
BMVC | 2 |
| 2023 | Lightweight Portrait Segmentation Via Edge-Optimized AttentionabstractWith the outbreak of COVID-19 around the world, the frequency of video conferencing at home is increasing. Therefore, a segmentation architecture that can quickly carry out close-range portrait segmentation has become a current need. However, the current portrait segmentation architectures cannot meet the requirements of lightweight and edge-friendly. We built architecture with 0.06G FLOPs and 0.02M parameters to overcome this phenomenon. This lightweight architecture can be better embedded and run on mobile devices that only support CPU computing. Our network achieves an FPS of 39.02 on CPU, which is more than three times faster than other networks. In addition, we pay special attention to the enhancement of edge features. The independent edge feature enhancement is embedded, and the edge-optimized attention mechanism (EOAM) is designed to collect specific edge areas for the bottom features and the high-level features in the process of feature fusion. Codes and results are publicly available at https://github.com/XinyueZhangqdu/ESPS. Xinyue Zhang 0009, Guodong Wang 0001, Chenglizhao Chen |
ICASSP | 4 |
| 2023 | Joint Probability Distribution Regression for Image CroppingabstractImage cropping aims at locating a candidate (rectangle region) with the highest aesthetic quality in professional photography. One solution of the previous methods is to generate a large number of candidates and then filter them, which leads to low efficiency. Another idea directly regresses the candidate coordinates to speed up but ignores the aesthetic subjectivity of the candidate’s evaluation, limiting the model’s performance. In this paper, we present an Aesthetic and Composition joint Probability Distribution regression Network (ACPD-Net) to explicitly investigate the process of generating the candidate with a joint probability distribution paradigm to improve the performance of cropping results in an efficient way. The joint probability distribution paradigm between location and size branch can identify the subjective aesthetic region and satisfy the objective composition rules in an end-to-end manner. Our method has been tested on the FCDB and FLMS datasets, which shows the superiority of ACPD-Net. The code is available at https://github.com/flyingbird93/ACPD-Net. Tengfei Shi, Chenglizhao Chen, Yuanbo He, Wenfeng Song, Aimin Hao |
ICIP | 2 |
| 2023 | Modality Profile - A New Critical Aspect to be Considered When Generating RGB-D Salient Object Detection Training SetabstractIt is widely acknowledged that selecting appropriate training data is crucial for obtaining good results in real-world testing, more so than utilizing complex network architectures. However, in the field of RGB-D SOD research, researchers have primarily focused on enhancing network architectures and have given less consideration to the choice of training and testing datasets, which may not translate well in practical applications. This paper aims to address an existing issue - how can we automatically generate a data-driven RGB-D SOD training dataset? We propose that in addition to scene similarity, the concept of "modality profile'' should be taken into account. The term "modality profile'' refers to the complementary status of modalities within a given dataset. A training dataset with a modality profile similar to the test dataset can significantly improve performance. To address this, we present a viable solution for automatically generating a training dataset with any desired modality profile in a weakly supervised manner. Our method also provides high-quality pseudo-GTs for all RGB-D images obtained from the web, making it suitable for training RGB-D SOD models. Extensive quantitative evaluations demonstrate the significance of the proposed "modality profile'' and confirm the superiority of the newly constructed training set guided by our "modality profile''. All codes, datasets, and results are available at this link. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
ACM Multimedia | 3 |
| 2023 | Saliency Driven Monocular Depth Estimation Based on Multi-scale Graph Convolutional Network
Dunquan Wu, Chenglizhao Chen |
PRCV (9) | 2 |
| 2023 | RFA-Net: Residual feature attention network for fine-grained image inpainting
Shengrui Zang, Zhenhua Ai, Jieru Chi, Guowei Yang 0002, Chenglizhao Chen |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Texture-aware gray-scale image colorization using a bistream generative adversarial network with multi scale attention structure
Shengrui Zang, Zhenhua Ai, Jieru Chi, Guowei Yang 0002, Chenglizhao Chen |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Multi-view inter-modality representation with progressive fusion for image-text matching
Jie Wu 0033, Leiquan Wang, Chenglizhao Chen, Jing Lu 0013, Chunlei Wu |
Neurocomputing | 3 |
| 2023 | Global and local similarity learning in multi-kernel space for nonnegative matrix factorization
Chong Peng 0001, Xingrong Hou, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
Knowl. Based Syst. | 5 |
| 2023 | Lossless segmentation of cardiac medical images by a resolution consistent network with nondamage data preprocessing
Chenglizhao Chen, Jingyang Gao |
Multim. Tools Appl. | 2 |
| 2023 | Consensus Low-Rank Multi-View Subspace Clustering With Cross-View Diversity PreservingabstractMulti-view subspace clustering has drawn significant attentions in recent years, which significantly improves learning performance of the single-view methods. In this letter, we propose a novel multi-view subspace clustering method, which learns a consensus representation with auto-weighted local neighboring transition probability matrix fusion and preserves cross-view diversity with a matrix-induced term. The new model is convex and thus admits efficient optimization. The effectiveness is confirmed by extensive experiments. Kehan Kang, Chenglizhao Chen, Chong Peng 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | Essential Low-Rank Sample Learning for Group-Aware Subspace ClusteringabstractIn this letter, we proposea novel subspace clustering method, named Es$^{3}$SC, that learns essential samples for low-dimensional representation construction. The essential samples are expected to retain key features and better estimate the example-wise similarities, which are more geared to seeking the representation matrix (RM). Moreover, the RM is enforced to have block-diagonal structural property, which directly reveals grouping structure of the data and is essentially desired by clustering application. Experimental results show that the Es$^{3}$SC is effective in both clustering and essential feature recovery. Fusheng Wang 0011, Chenglizhao Chen, Chong Peng 0001 |
IEEE Signal Process. Lett. | 2 |
| 2023 | A Comprehensive Survey on Video Saliency Detection With Auditory Information: The Audio-Visual Consistency Perceptual is the Key!abstractVideo saliency detection (VSD) aims at fast locating the most attractive objects/things/patterns in a given video clip. Existing VSD-related works have mainly relied on the visual system but paid less attention to the audio aspect. In contrast, our audio system is the most vital complementary part of our visual system. Also, audio-visual saliency detection (AVSD), one of the most representative research topics for mimicking human perceptual mechanisms, is currently in its infancy, and none of the existing survey papers have touched on it, especially from the perspective of saliency detection. Thus, the ultimate goal of this paper is to provide an extensive review to bridge the gap between audio-visual fusion and saliency detection. In addition, as another highlight of this review, we have provided a deep insight into key factors that could directly determine AVSD deep models’ performances. We claim that the audio-visual consistency degree (AVC) — a long-overlooked issue, can directly influence the effectiveness of using audio to benefit its visual counterpart when performing saliency detection. Moreover, to make the AVC issue more practical and valuable for future followers, we have newly equipped almost all existing publicly available AVSD datasets with additional frame-wise AVC labels. Based on these upgraded datasets, we have conducted extensive quantitative evaluations to ground our claim on the importance of AVC in the AVSD task. In a word, our ideas and new sets serve as a convenient platform with preliminaries and guidelines, all of which can potentially facilitate future works in further promoting state-of-the-art (SOTA) performance. Chenglizhao Chen, Mengke Song, Wenfeng Song, Li Guo 0016, Muwei Jian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | An Intelligent Virtual Standard Patient for Medical Students Training Based on Oral Knowledge GraphabstractVirtual standard patient (VSP) is in high demand for medical students' diagnosis ability training in an efficient manner. Different from the traditional conversation system in medical dialogue generation, VSP needs a novel conversation paradigm to act as the patient instead of the doctor. However, existing conversation techniques still have limited ability in terms of generation of symptoms exhibited by patients with the personalized and knowledge-centered expressions. To alleviate these problems, we propose to construct a novel oral knowledge graph, which sufficiently provides medical clues of the certain disease. Accordingly, the VSP could accurately interact with the dentists for their underlying intention and express the symptoms characters in a natural style. To efficiently retrieve the related disease clues, the symptoms descriptions of the oral diseases are encoded into the oral knowledge graph, which could well organize the disease-centered symptom entities and speaking styles. Moreover, to transfer the common sense knowledge from existing large scale of medical knowledge graph to the specific oral knowledge graph, a coupled pre-trained Bert models is further designed to learn the related medical knowledge from coarse-level to fine-level hierarchically. Finally, a series of well-designed personalized templates are proposed to generate plausible and realistic answers in condition of the certain disease. We also conduct extensive user studies to demonstrate that the VSP satisfies the medical students' diagnosis practice requirement in terms of naturalness, realism, and topic relevance. Wenfeng Song, Xia Hou, Shuai Li 0001, Chenglizhao Chen, Danyang Gao, Xian'e Wang, Yuzhe Sun, Jianxia Hou, Aimin Hao |
IEEE Trans. Multim. | 4 |
| 2023 | FineStyle: Semantic-Aware Fine-Grained Motion Style Transfer with Dual Interactive-Flow FusionabstractWe present FineStyle, a novel framework for motion style transfer that generates expressive human animations with specific styles for virtual reality and vision fields. It incorporates semantic awareness, which improves motion representation and allows for precise and stylish animation generation. Existing methods for motion style transfer have all failed to consider the semantic meaning behind the motion, resulting in limited controls over the generated human animations. To improve, FineStyle introduces a new cross-modality fusion module called Dual Interactive-Flow Fusion (DIFF). As the first attempt, DIFF integrates motion style features and semantic flows, producing semantic-aware style codes for fine-grained motion style transfer. FineStyle uses an innovative two-stage semantic guidance approach that leverages semantic clues to enhance the discriminative power of both semantic and style features. At an early stage, a semantic-guided encoder introduces distinct semantic clues into the style flow. Then, at a fine stage, both flows are further fused interactively, selecting the matched and critical clues from both flows. Extensive experiments demonstrate that FineStyle outperforms state-of-the-art methods in visual quality and controllability. By considering the semantic meaning behind motion style patterns, FineStyle allows for more precise control over motion styles. Source code and model are available on https://github.com/XingliangJin/Fine-Style.git. Wenfeng Song, Xingliang Jin, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Xia Hou |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | Mask-Guided Self-Distillation For Visual TrackingabstractRecently, Siamese-based visual trackers has been dramatically improved on performance, however, excessive parameters of tracking networks make models seriously hinder the practical deployment on edge devices. In this work, we present mask-guided self-distillation(MGSD) to compress the models of Siamese-based visual trackers, which enables Siamese-based visual trackers to capture crucial knowledge for effecting the performance of tracking. Specifically, MGSD consists of mask-guided semantic features self-distillation and decoupled tracking-head self-distillation, which discards the redundant convolution parameters in feature-dependency and task-dependency. It is worth noting that the proposed method compress exponentially the models of Siamese-based trackers under the condition of keeping even improving the performance of tracking. Extensive comprehensive experiments on tracking benchmarks including OTB2015, UAV123, LaSOT demonstrate that our proposed method can compress model size of Siamese-based visual trackers to 50% while achieving state-of-the-art performance. Our source code is available at: https://github.com/xl0312/MGSD. Luming Li, Chenglizhao Chen, Xiaowei Zhang 0003 |
ICME | 2 |
| 2022 | Boundary-Guided Camouflaged Object DetectionabstractCamouflaged object detection (COD), segmenting objects that are elegantly blended into their surroundings, is a valuable yet challenging task. Existing deep-learning methods often fall into the difficulty of accurately identifying the camouflaged object with complete and fine object structure. To this end, in this paper, we propose a novel boundary-guided network (BGNet) for camouflaged object detection. Our method explores valuable and extra object-related edge semantics to guide representation learning of COD, which forces the model to generate features that highlight object structure, thereby promoting camouflaged object detection of accurate boundary localization. Extensive experiments on three challenging benchmark datasets demonstrate that our BGNet significantly outperforms the existing 18 state-of-the-art methods under four widely-used evaluation metrics. Our code is publicly available at: https://github.com/thograce/BGNet. Shuo Wang 0010, Chenglizhao Chen, Tian-Zhu Xiang |
IJCAI | 3 |
| 2022 | Synthetic Data Supervised Salient Object DetectionabstractAlthough deep salient object detection (SOD) has achieved remarkable progress, deep SOD models are extremely data-hungry, requiring large-scale pixel-wise annotations to deliver such promising results. In this paper, we propose a novel yet effective method for SOD, coined SODGAN, which can generate infinite high-quality image-mask pairs requiring only a few labeled data, and these synthesized pairs can replace the human-labeled DUTS-TR to train any off-the-shelf SOD model. Its contribution is three-fold. 1) Our proposed diffusion embedding network can address the manifold mismatch and is tractable for the latent code generation, better matching with the ImageNet latent space. 2) For the first time, our proposed few-shot saliency mask generator can synthesize infinite accurate image synchronized saliency masks with a few labeled data. 3) Our proposed quality-aware discriminator can select highquality synthesized image-mask pairs from noisy synthetic data pool, improving the quality of synthetic data. For the first time, our SODGAN tackles SOD with synthetic data directly generated from the generative model, which opens up a new research paradigm for SOD. Extensive experimental results show that the saliency model trained on synthetic data can achieve $98.4%$ F-measure of the saliency model trained on the DUTS-TR. Moreover, our approach achieves a new SOTA performance in semi/weakly-supervised methods, and even outperforms several fully-supervised SOTA methods. Code is available at https://github.com/wuzhenyubuaa/SODGAN Wei Wang 0169, Tengfei Shi, Chenglizhao Chen, Aimin Hao, Shuo Li 0001 |
ACM Multimedia | 5 |
| 2022 | CFA-Net: Cross-Level Feature Fusion and Aggregation Network for Salient Object Detection
Huiqi Li, Chenglizhao Chen |
PRCV (4) | 5 |
| 2022 | Two-dimensional semi-nonnegative matrix factorization for clustering
Chong Peng 0001, Chenglizhao Chen, Zhao Kang 0001, Qiang Shawn Cheng |
Inf. Sci. | 3 |
| 2022 | Log-based sparse nonnegative matrix factorization for data representation
Chong Peng 0001, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
Knowl. Based Syst. | 5 |
| 2022 | Preserving bilateral view structural information for subspace clustering
Chong Peng 0001, Yongyong Chen, Chenglizhao Chen, Zhao Kang 0001, Li Guo 0016, Qiang Shawn Cheng |
Knowl. Based Syst. | 5 |
| 2022 | Recursive multi-model complementary deep fusion for robust salient object detection via parallel sub-networks
Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
Pattern Recognit. | 3 |
| 2022 | A Novel Video Salient Object Detection Method via Semisupervised Motion Quality PerceptionabstractPrevious video salient object detection (VSOD) approaches have mainly focused on the perspective of network design for achieving performance improvements. However, with the recent slowdown in the development of deep learning techniques, it might become increasingly difficult to anticipate another breakthrough solely via complex networks. Therefore, this paper proposes a universal learning scheme to obtain a further 3% performance improvement for all state-of-the-art (SOTA) VSOD models. The major highlight of our method is that we propose the ‘motion quality’, a new concept for mining video frames from the ‘buffered’ testing video stream for constructing a fine-tuning set. By using our approach, all frames in this set can all well-detect their salient object by the ‘target SOTA model’ — the one we want to improve. Thus, the VSOD results of the mined set, which were previously derived by the target SOTA model, can be directly applied as pseudolearning objectives to fine-tune a completely new spatial model that has been pretrained on the widely used DAVIS-TR set. Since some spatial scenes in the buffered testing video stream are shown, the fine-tuned spatial model can perform very well for the remaining unseen testing frames, outperforming the target SOTA model significantly. Although offline model fine tuning requires additional time costs, the performance gain can still benefit scenarios without speed requirements. Moreover, its semisupervised methodology might have considerable potential to inspire the VSOD community in the future. Chenglizhao Chen, Chong Peng 0001, Guodong Wang 0001, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | A Novel Long-Term Iterative Mining Scheme for Video Salient Object DetectionabstractThe existing state-of-the-art (SOTA) video salient object detection (VSOD) models have widely followed short-term methodology, which dynamically determines the balance between spatial and temporal saliency fusion by solely considering the current consecutive limited frames. However, the short-term methodology has one critical limitation, which conflicts with the real mechanism of our visual system — a typical long-term methodology. As a result, failure cases keep showing up in the results of the current SOTA models, and the short-term methodology becomes the major technical bottleneck. To solve this problem, this paper proposes a novel VSOD approach, which performs VSOD in a complete long-term way. Our approach converts the sequential VSOD, a sequential task, to a data mining problem, i.e., decomposing the input video sequence to object proposals in advance and then mining salient object proposals as much as possible in an easy-to-hard way. Since all object proposals are simultaneously available, the proposed approach is a complete long-term approach, which can alleviate some difficulties rooted in conventional short-term approaches. In addition, we devised an online updating scheme that can grasp the most representative and trustworthy pattern profile of the salient objects, outputting framewise saliency maps with rich details and smoothing both spatially and temporally. The proposed approach outperforms almost all SOTA models on five widely used benchmark datasets. Chenglizhao Chen, Hengsen Wang, Yuming Fang 0001, Chong Peng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Hyperspectral Image Denoising Using Nonconvex Local Low-Rank and Sparse Separation With Spatial-Spectral Total Variation RegularizationabstractIn this paper, we propose a novel nonconvex approach to robust principal component analysis for HSI denoising, which focuses on simultaneously developing more accurate approximations to both rank and column-wise sparsity for the low-rank and sparse components, respectively. In particular, the new method adopts the log-determinant rank approximation and a novell2,lognorm, to restrict the local low-rank or column-wisely sparse properties for the component matrices, respectively. For thel2,log-regularized shrinkage problem, we develop an efficient, closed-form solution, which is namedl2,log-shrinkage operator. The new regularization and the corresponding operator can be generally used in other problems that require column-wise sparsity. Moreover, we impose the spatial-spectral total variation regularization in the log-based nonconvex RPCA model, which enhances the global piece-wise smoothness and spectral consistency from the spatial and spectral views in the recovered HSI. Extensive experiments on both simulated and real HSIs demonstrate the effectiveness of the proposed method in denoising HSIs. Chong Peng 0001, Kehan Kang, Yongyong Chen, Xinxing Wu, Andrew Cheng, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2022 | Improving RGB-D Salient Object Detection via Modality-Aware DecoderabstractMost existing RGB-D salient object detection (SOD) methods are primarily focusing on cross-modal and cross-level saliency fusion, which has been proved to be efficient and effective. However, these methods still have a critical limitation, i.e., their fusion patterns - typically the combination of selective characteristics and its variations, are too highly dependent on the network's non-linear adaptability. In such methods, the balances between RGB and D (Depth) are formulated individually considering the intermediate feature slices, but the relation at the modality level may not be learned properly. The optimal RGB-D combinations differ depending on the RGB-D scenarios, and the exact complementary status is frequently determined by multiple modality-level factors, such as D quality, the complexity of the RGB scene, and degree of harmony between them. Therefore, given the existing approaches, it may be difficult for them to achieve further performance breakthroughs, as their methodologies belong to some methods that are somewhat less modality sensitive. To conquer this problem, this paper presents the Modality-aware Decoder (MaD). The critical technical innovations include a series of feature embedding, modality reasoning, and feature back-projecting and collecting strategies, all of which upgrade the widely-used multi-scale and multi-level decoding process to be modality-aware. Our MaD achieves competitive performance over other state-of-the-art (SOTA) models without using any fancy tricks in the decoder's design. Codes and results will be publicly available at https://github.com/MengkeSong/MaD. Mengke Song, Wenfeng Song, Guowei Yang 0002, Chenglizhao Chen |
IEEE Trans. Image Process. | 4 |
| 2022 | Salient Object Detection via Dynamic Scale RoutingabstractRecent research advances in salient object detection (SOD) could largely be attributed to ever-stronger multi-scale feature representation empowered by the deep learning technologies. The existing SOD deep models extract multi-scale features via the off-the-shelf encoders and combine them smartly via various delicate decoders. However, the kernel sizes in this commonly-used thread are usually "fixed". In our new experiments, we have observed that kernels of small size are preferable in scenarios containing tiny salient objects. In contrast, large kernel sizes could perform better for images with large salient objects. Inspired by this observation, we advocate the "dynamic" scale routing (as a brand-new idea) in this paper. It will result in a generic plug-in that could directly fit the existing feature backbone. This paper's key technical innovations are two-fold. First, instead of using the vanilla convolution with fixed kernel sizes for the encoder design, we propose the dynamic pyramid convolution (DPConv), which dynamically selects the best-suited kernel sizes w.r.t. the given input. Second, we provide a self-adaptive bidirectional decoder design to accommodate the DPConv-based encoder best. The most significant highlight is its capability of routing between feature scales and their dynamic collection, making the inference process scale-aware. As a result, this paper continues to enhance the current SOTA performance. Both the code and dataset are publicly available at https://github.com/wuzhenyubuaa/DPNet. Shuai Li 0001, Chenglizhao Chen, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 3 |
| 2022 | Deeper Look at Image Salient Object Detection: Bi-Stream Network With a Small Training DatasetabstractCompared with the conventional hand-crafted approaches, the deep learning based ISOD (image salient object detection) models have achieved tremendous performance improvements by training exquisitely crafted fancy networks over large-scale training sets. However, do we really need large-scale training set for ISOD? In this article, we provide a deeper insight into the interrelationship between the ISOD performance and the training data. To alleviate the conventional demands for large-scale training data, we provide a feasible way to construct a novel small-scale training set, which only contains 4 K images. To take full advantage of this new set, we propose a novel bi-stream network consisting of two different feature backbones. Benefit from the proposed gate control unit, this bi-stream network is able to achieve complementary fusion status for its subbranches. To our best knowledge, this is the first attempt to use a small-scale training set to compete with other large-scale ones; nevertheless, our method can still achieve the leading SOTA performance on all tested benchmark datasets. Both the code and dataset are publicly available athttps://github.com/wuzhenyubuaa/TSNet. Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | From Semantic Categories to Fixations: A Novel Weakly-Supervised Visual-Auditory Saliency Detection ApproachabstractThanks to the rapid advances in the deep learning techniques and the wide availability of large-scale training sets, the performances of video saliency detection models have been improving steadily and significantly. However, the deep learning based visual-audio fixation prediction is still in its infancy. At present, only a few visual-audio sequences have been furnished with real fixations being recorded in the real visual-audio environment. Hence, it would be neither efficiency nor necessary to re-collect real fixations under the same visual-audio circumstance. To address the problem, this paper advocate a novel approach in a weakly-supervised manner to alleviating the demand of large-scale training sets for visual-audio model training. By using the video category tags only, we propose the selective class activation mapping (SCAM), which follows a coarse-to-fine strategy to select the most discriminative regions in the spatial-temporal-audio circumstance. Moreover, these regions exhibit high consistency with the real human-eye fixations, which could subsequently be employed as the pseudo GTs to train a new spatial-temporal-audio (STA) network. Without resorting to any real fixation, the performance of our STA network is comparable to that of the fully supervised ones. Our code and results are publicly available at https://github.com/guotaowang/STANet. Guotao Wang 0004, Chenglizhao Chen, Deng-Ping Fan, Aimin Hao, Hong Qin 0001 |
CVPR | 2 |
| 2021 | Mutual Graph Learning for Camouflaged Object DetectionabstractAutomatically detecting/segmenting object(s) that blend in with their surroundings is difficult for current models. A major challenge is that the intrinsic similarities between such foreground objects and background surroundings make the features extracted by deep model indistinguishable. To overcome this challenge, an ideal model should be able to seek valuable, extra clues from the given scene and incorporate them into a joint learning framework for representation co-enhancement. With this inspiration, we design a novel Mutual Graph Learning (MGL) model, which generalizes the idea of conventional mutual learning from regular grids to the graph domain. Specifically, MGL decouples an image into two task-specific feature maps — one for roughly locating the target and the other for accurately capturing its boundary details — and fully exploits the mutual benefits by recurrently reasoning their high-order relations through graphs. Importantly, in contrast to most mutual learning approaches that use a shared function to model all between-task interactions, MGL is equipped with typed functions for handling different complementary relations to maximize information interactions. Experiments on challenging datasets, including CHAMELEON, CAMO and COD10K, demonstrate the effectiveness of our MGL with superior performance to existing state-of-the-art methods. Code is available at https://github.com/fanyang587/MGL. Qiang Zhai, Xin Li 0079, Fan Yang 0054, Chenglizhao Chen, Hong Cheng 0002, Deng-Ping Fan |
CVPR | 4 |
| 2021 | Depth quality-aware selective saliency fusion for RGB-D image salient object detection
Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
Neurocomputing | 3 |
| 2021 | Improving video anomaly detection performance by mining useful data from unseen video frames
Renzhi Wu, Shuai Li 0001, Chenglizhao Chen, Aimin Hao |
Neurocomputing | 3 |
| 2021 | Nonnegative matrix factorization with local similarity learning
Chong Peng 0001, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
Inf. Sci. | 4 |
| 2021 | Learning discriminative representation for image classification
Chong Peng 0001, Zhao Kang 0001, Yongyong Chen, Chenglizhao Chen, Qiang Shawn Cheng |
Knowl. Based Syst. | 6 |
| 2021 | Kernel two-dimensional ridge regression for subspace clustering
Chong Peng 0001, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
Pattern Recognit. | 4 |
| 2021 | A Plug-and-Play Scheme to Adapt Image Saliency Deep Model for Video DataabstractWith the rapid development of deep learning techniques, image saliency deep models trained solely by spatial information have occasionally achieved detection performance for video data comparable to that of the models trained by both spatial and temporal information. However, due to the lesser consideration of temporal information, the image saliency deep models may become fragile in the video sequences dominated by temporal information. Thus, the most recent video saliency detection approaches have adopted the network architecture starting with a spatial deep model that is followed by an elaborately designed temporal deep model. However, such methods easily encounter the performance bottleneck arising from the single stream learning methodology, so the overall detection performance is largely determined by the spatial deep model. In sharp contrast to the current mainstream methods, this paper proposes a novel plug-and-play scheme to weakly retrain a pretrained image saliency deep model for video data by using the newly sensed and coded temporal information. Thus, the retrained image saliency deep model will be able to maintain temporal saliency awareness, achieving much improved detection performance. Moreover, our method is simple yet effective for adapting any off-the-shelf pre-trained image saliency deep model to obtain high-quality video saliency detection. Additionally, both the data and source code of our method are publicly available. Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Exploring Rich and Efficient Spatial Temporal Interactions for Real-Time Video Salient Object DetectionabstractWe have witnessed a growing interest in video salient object detection (VSOD) techniques in today's computer vision applications. In contrast with temporal information (which is still considered a rather unstable source thus far), the spatial information is more stable and ubiquitous, thus it could influence our vision system more. As a result, the current main-stream VSOD approaches have inferred and obtained their saliency primarily from the spatial perspective, still treating temporal information as subordinate. Although the aforementioned methodology of focusing on the spatial aspect is effective in achieving a numeric performance gain, it still has two critical limitations. First, to ensure the dominance by the spatial information, its temporal counterpart remains inadequately used, though in some complex video scenes, the temporal information may represent the only reliable data source, which is critical to derive the correct VSOD. Second, both spatial and temporal saliency cues are often computed independently in advance and then integrated later on, while the interactions between them are omitted completely, resulting in saliency cues with limited quality. To combat these challenges, this paper advocates a novel spatiotemporal network, where the key innovation is the design of its temporal unit. Compared with other existing competitors (e.g., convLSTM), the proposed temporal unit exhibits an extremely lightweight design that does not degrade its strong ability to sense temporal information. Furthermore, it fully enables the computation of temporal saliency cues that interact with their spatial counterparts, ultimately boosting the overall VSOD performance and realizing its full potential towards mutual performance improvement for each. The proposed method is easy to implement yet still effective, achieving high-quality VSOD at 50 FPS in real-time applications. Chenglizhao Chen, Guotao Wang 0004, Chong Peng 0001, Yuming Fang 0001, Dingwen Zhang, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Depth-Quality-Aware Salient Object DetectionabstractThe existing fusion-based RGB-D salient object detection methods usually adopt the bistream structure to strike a balance in the fusion trade-off between RGB and depth (D). While the D quality usually varies among the scenes, the state-of-the-art bistream approaches are depth-quality-unaware, resulting in substantial difficulties in achieving complementary fusion status between RGB and D and leading to poor fusion results for low-quality D. Thus, this paper attempts to integrate a novel depth-quality-aware subnet into the classic bistream structure in order to assess the depth quality prior to conducting the selective RGB-D fusion. Compared to the SOTA bistream methods, the major advantage of our method is its ability to lessen the importance of the low-quality, no-contribution, or even negative-contribution D regions during RGB-D fusion, achieving a much improved complementary status between RGB and D. Our source code and data are available online at https://github.com/qdu1995/DQSD. Chenglizhao Chen, Jipeng Wei, Chong Peng 0001, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Rethinking Image Salient Object Detection: Object-Level Semantic Saliency Reranking First, Pixelwise Saliency Refinement LaterabstractHuman attention is an interactive activity between our visual system and our brain, using both low-level visual stimulus and high-level semantic information. Previous image salient object detection (SOD) studies conduct their saliency predictions via a multitask methodology in which pixelwise saliency regression and segmentation-like saliency refinement are conducted simultaneously. However, this multitask methodology has one critical limitation: the semantic information embedded in feature backbones might be degenerated during the training process. Our visual attention is determined mainly by semantic information, which is evidenced by our tendency to pay more attention to semantically salient regions even if these regions are not the most perceptually salient at first glance. This fact clearly contradicts the widely used multitask methodology mentioned above. To address this issue, this paper divides the SOD problem into two sequential steps. First, we devise a lightweight, weakly supervised deep network to coarsely locate the semantically salient regions. Next, as a postprocessing refinement, we selectively fuse multiple off-the-shelf deep models on the semantically salient regions identified by the previous step to formulate a pixelwise saliency map. Compared with the state-of-the-art (SOTA) models that focus on learning the pixelwise saliency in single images using only perceptual clues, our method aims at investigating the object-level semantic ranks between multiple images, of which the methodology is more consistent with the human attention mechanism. Our method is simple yet effective, and it is the first attempt to consider salient object detection as mainly an object-level semantic reranking problem. Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Data-Level Recombination and Lightweight Fusion Scheme for RGB-D Salient Object DetectionabstractExisting RGB-D salient object detection methods treat depth information as an independent component to complement RGB and widely follow the bistream parallel network architecture. To selectively fuse the CNN features extracted from both RGB and depth as a final result, the state-of-the-art (SOTA) bistream networks usually consist of two independent subbranches: one subbranch is used for RGB saliency, and the other aims for depth saliency. However, depth saliency is persistently inferior to the RGB saliency because the RGB component is intrinsically more informative than the depth component. The bistream architecture easily biases its subsequent fusion procedure to the RGB subbranch, leading to a performance bottleneck. In this paper, we propose a novel data-level recombination strategy to fuse RGB with D (depth) before deep feature extraction, where we cyclically convert the original 4-dimensional RGB-D into DGB, RDB and RGD. Then, a newly lightweight designed triple-stream network is applied over these novel formulated data to achieve an optimal channel-wise complementary fusion status between the RGB and D, achieving a new SOTA performance. Xuehao Wang, Shuai Li 0001, Chenglizhao Chen, Yuming Fang 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Full-reference Screen Content Image Quality Assessment by Fusing Multilevel Structure SimilarityabstractScreen content images (SCIs) usually comprise various content types with sharp edges, in which artifacts or distortions can be effectively sensed by a vanilla structure similarity measurement in a full-reference manner. Nonetheless, almost all of the current state-of-the-art (SOTA) structure similarity metrics are “locally” formulated in a single-level manner, while the true human visual system (HVS) follows the multilevel manner; such mismatch could eventually prevent these metrics from achieving reliable quality assessment. To ameliorate this issue, this article advocates a novel solution to measure structure similarity “globally” from the perspective of sparse representation. To perform multilevel quality assessment in accordance with the real HVS, the abovementioned global metric will be integrated with the conventional local ones by resorting to the newly devised selective deep fusion network. To validate its efficacy and effectiveness, we have compared our method with 12 SOTA methods over two widely used large-scale public SCI datasets, and the quantitative results indicate that our method yields significantly higher consistency with subjective quality scores than the current leading works. Both the source code and data are also publicly available to gain widespread acceptance and facilitate new advancement and validation. Chenglizhao Chen, Hongmeng Zhao, Huan Yang 0001, Chong Peng 0001, Hong Qin 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Robust principal component analysis: A factorization-based approach with linear complexity
Chong Peng 0001, Yongyong Chen, Zhao Kang 0001, Chenglizhao Chen, Qiang Shawn Cheng |
Inf. Sci. | 4 |
| 2020 | Accurate image super-resolution using dense connections and dimension reduction network
Guodong Wang 0001, Chenglizhao Chen, Zhenkuan Pan 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Multi-scale dilated convolution of convolutional neural network for crowd counting
Guodong Wang 0001, Chenglizhao Chen, Zhenkuan Pan 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Improved Robust Video Saliency Detection Based on Long-Term Spatial-Temporal InformationabstractThis paper proposes to utilize supervised deep convolutional neural networks to take full advantage of the long-term spatial-temporal information in order to improve the video saliency detection performance. The conventional methods, which use the temporally neighbored frames solely, could easily encounter transient failure cases when the spatial-temporal saliency clues are less-trustworthy for a long period. To tackle the aforementioned limitation, we plan to identify those beyond-scope frames with trustworthy long-term saliency clues first and then align it with the current problem domain for an improved video saliency detection. Chenglizhao Chen, Guotao Wang 0004, Chong Peng 0001, Xiaowei Zhang 0003, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Improved Saliency Detection in RGB-D Images Using Two-Phase Depth Estimation and Selective Deep FusionabstractTo solve the saliency detection problem in RGB-D images, the depth information plays a critical role in distinguishing salient objects or foregrounds from cluttered backgrounds. As the complementary component to color information, the depth quality directly dictates the subsequent saliency detection performance. However, due to artifacts and the limitation of depth acquisition devices, the quality of the obtained depth varies tremendously across different scenarios. Consequently, conventional selective fusion-based RGB-D saliency detection methods may result in a degraded detection performance in cases containing salient objects with low color contrast coupled with a low depth quality. To solve this problem, we make our initial attempt to estimate additional high-quality depth information, which is denoted by Depth+. Serving as a complement to the original depth, Depth+ will be fed into our newly designed selective fusion network to boost the detection performance. To achieve this aim, we first retrieve a small group of images that are similar to the given input, and then the inter-image, nonlocal correspondences are built accordingly. Thus, by using these inter-image correspondences, the overall depth can be coarsely estimated by utilizing our newly designed depth-transferring strategy. Next, we build fine-grained, object-level correspondences coupled with a saliency prior to further improve the depth quality of the previous estimation. Compared to the original depth, our newly estimated Depth+ is potentially more informative for detection improvement. Finally, we feed both the original depth and the newly estimated Depth+ into our selective deep fusion network, whose key novelty is to achieve an optimal complementary balance to make better decisions toward improving saliency boundaries. Chenglizhao Chen, Jipeng Wei, Chong Peng 0001, Hong Qin 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Accurate and Robust Video Saliency Detection via Self-Paced DiffusionabstractConventional video saliency detection methods frequently follow the common bottom-up thread to estimate video saliency within the short-term fashion. As a result, such methods can not avoid the obstinate accumulation of errors when the collected low-level clues are constantly ill-detected. Also, being noticed that a portion of video frames, which are not nearby the current video frame over the time axis, may potentially benefit the saliency detection in the current video frame. Thus, we propose to solve the aforementioned problem using our newly-designed key frame strategy (KFS), whose core rationale is to utilize both the spatial-temporal coherency of the salient foregrounds and the objectness prior (i.e., how likely it is for an object proposal to contain an object of any class) to reveal the valuable long-term information. We could utilize all this newly-revealed long-term information to guide our subsequent “self-paced” saliency diffusion, which enables each key frame itself to determine its diffusion range and diffusion strength to correct those ill-detected video frames. At the algorithmic level, we first divide a video sequence into short-term frame batches, and the object proposals are obtained in a frame-wise manner. Then, for each object proposal, we utilize a pre-trained deep saliency model to obtain high-dimensional features in order to represent the spatial contrast. Since the contrast computation within multiple neighbored video frames (i.e., the non-local manner) is relatively insensitive to the appearance variation, those object proposals with high-quality low-level saliency estimation frequently exhibit strong similarity over the temporal scale. Next, the long-term common consistency (e.g., appearance models/movement patterns) of the salient foregrounds could be explicitly revealed via similarity analysis accordingly. We further boost the detection accuracy via long-term information guided saliency diffusion in a self-paced manner. We have conducted extensive experiments to compare our method with 16 state-of-the-art methods over 4 largest public available benchmarks, and all results demonstrate the superiority of our method in terms of both accuracy and robustness. Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Salient Object Detection via Multiple Instance Joint Re-LearningabstractIn recent years deep neural networks have been widely applied to visual saliency detection tasks with remarkable detection performance improvements. As for the salient object detection in single image, the automatically computed convolutional features frequently demonstrate high discriminative power to distinguish salient foregrounds from its non-salient surroundings in most cases. Yet, the obstinate feature conflicts still persist, which naturally gives rise to the learning ambiguity, arriving at massive failure detections. To solve such problem, we propose to jointly re-learn common consistency of inter-image saliency and then use it to boost the detection performance. Its core rationale is to utilize the easy-to-detect cases to re-boost much harder ones. Compared with the conventional methods, which focus on their problem domain within the single image scope, our method attempts to utilize those beyond-scope information to facilitate the current salient object detection. To validate our new approach, we have conducted a comprehensive quantitative comparisons between our approach and 13 state-of-the-art methods over 5 publicly available benchmarks, and all the results suggest the advantage of our approach in terms of accuracy, reliability, and versatility. Guangxiao Ma, Chenglizhao Chen, Shuai Li 0001, Chong Peng 0001, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Stage-wise Salient Object Detection in 360° Omnidirectional Image via Object-level Semantical Saliency RankingabstractThe 2D image based salient object detection (SOD) has been extensively explored, while the 360° omnidirectional image based SOD has received less research attention and there exist three major bottlenecks that are limiting its performance. Firstly, the currently available training data is insufficient for the training of 360° SOD deep model. Secondly, the visual distortions in 360° omnidirectional images usually result in large feature gap between 360° images and 2D images; consequently, the widely used stage-wise training-a widely-used solution to alleviate the training data shortage problem, becomes infeasible when conducing SOD in 360° omnidirectional images. Thirdly, the existing 360° SOD approach has followed a multi-task methodology that performs salient object localization and segmentation-like saliency refinement at the same time, being faced with extremely large problem domain, making the training data shortage dilemma even worse. To tackle all these issues, this paper divides the 360° SOD into a multi-staqe task, the key rationale of which is to decompose the original complex problem domain into sequential easy sub problems that only demand for small-scale training data. Meanwhile, we learn how to rank the "object-level semantical saliency", aiming to locate salient viewpoints and objects accurately. Specifically, to alleviate the training data shortage problem, we have released a novel dataset named 360-SSOD, containing 1,105 360° omnidirectional images with manually annotated object-level saliency ground truth, whose semantical distribution is more balanced than that of the existing dataset. Also, we have compared the proposed method with 13 SOTA methods, and all quantitative results have demonstrated the performance superiority. Guangxiao Ma, Shuai Li 0001, Chenglizhao Chen, Aimin Hao, Hong Qin 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2019 | RES-PCA: A Scalable Approach to Recovering Low-Rank MatricesabstractRobust principal component analysis (RPCA) has drawn significant attentions due to its powerful capability in recovering low-rank matrices as well as successful appplications in various real world problems. The current state-of-the-art algorithms usually need to solve singular value decomposition of large matrices, which generally has at least a quadratic or even cubic complexity. This drawback has limited the application of RPCA in solving real world problems. To combat this drawback, in this paper we propose a new type of RPCA method, RES-PCA, which is linearly efficient and scalable in both data size and dimension. For comparison purpose, AltProj, an existing scalable approach to RPCA requires the precise knowlwdge of the true rank; otherwise, it may fail to recover low-rank matrices. By contrast, our method works with or without knowing the true rank; even when both methods work, our method is faster. Extensive experiments have been performed and testified to the effectiveness of proposed method quantitatively and in visual quality, which suggests that our method is suitable to be employed as a light-weight, scalable component for RPCA in any application pipelines. Chong Peng 0001, Chenglizhao Chen, Zhao Kang 0001, Qiang Shawn Cheng |
CVPR | 2 |
| 2019 | Multi-scale dilated convolution of convolutional neural network for image denoising
Guodong Wang 0001, Chenglizhao Chen, Zhenkuan Pan 0001 |
Multim. Tools Appl. | 3 |
| 2018 | A Novel Bottom-Up Saliency Detection Method for Video With Dynamic BackgroundabstractAfter years of extensive studies, the salient motion detection problem has gained plausible performance improvement that was primarily propelled by the rapid development of self-adaptive top-down modeling techniques. Nevertheless, almost all the conventional solutions are still not robust enough to handle video sequences captured by hand-hold cameras. This is mainly due to the absence of the position alignment information that is indispensable for top-down background modeling. In contrast, the bottom-up video saliency detection methods, though achieving excellent salient motion detection in either stationary or nonstationary videos, still have rather poor detection performance in scenarios with massive dynamic background. In this letter, we explore a bottom-up saliency framework by introducing a novel spatial-temporal regional filter method to handle the dynamic background problem. Our key rationale is to assign large saliency value to those regions with stable spatial-temporal coherency while eliminating irregular, repeating dynamic background. As far as we know, this is the first work to address the dynamic background problem from the perspective of the bottom-up video saliency. We conduct massive quantitative evaluations over public available benchmarks to validate the effectiveness and robustness of our method. Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE Signal Process. Lett. | 1 |
| 2018 | Bilevel Feature Learning for Video Saliency DetectionabstractThis paper advocates a novel learning solution to the modeling of long-term spatial-temporal saliency consistency in order to boost the accuracy for video saliency detection. Conventional methods typically utilize the “slack” spatial-temporal model to locally ensure the smoothness of the computed video saliency, yet they could easily encounter the performance tradeoff dilemma (i.e., detection' accuracy and integrity). In contrast, our novel approach proposes the bilevel learning strategy to globally exploit the saliency consistency while overcoming the aforementioned difficulty. Our method first starts with the contrast computation of low-level saliency clues in a frame-wise manner. Then, based on such obtained saliency clues, we devise a novel bilevel Markov Random Field (bMRF) solution to conduct semantic labelling, which can explicitly indicates both the salient salient foregrounds and nonsalient nearby surroundings with high confidence while shrinking the low confidence remains. In such a way, the spatial-temporal consistency constraint is embedded intrinsically into the above explicit semantic labels, and we prevent the performance tradeoff problem from occurring. Next, based on those semantic labels made by our bMRF method, we further propose learning multiple nonlinear feature transformations to enlarge the feature margin between the salient foregrounds and the non-salient nearby surroundings, whose key rationale is to resort to long-term common consistencies to enforce the spatial-temporal smoothness. Thus, we can utilize these learned non-linear feature transformations to simultaneously suppress those short-term false-alarms and correct those hollow effects. To validate our new approach, we conduct extensive experiments on five publicly available benchmarks, and make comprehensive, quantitative evaluations between our method and 17 state-of-the-art techniques. All of the results demonstrate our method's advantages in terms of accuracy, reliability, robustness, and versatility. Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Zhenkuan Pan 0001, Guowei Yang 0002 |
IEEE Trans. Multim. | 1 |
| 2017 | Video Saliency Detection via Spatial-Temporal Fusion and Low-Rank Coherency DiffusionabstractThis paper advocates a novel video saliency detection method based on the spatial-temporal saliency fusion and low-rank coherency guided saliency diffusion. In sharp contrast to the conventional methods, which conduct saliency detection locally in a frame-by-frame way and could easily give rise to incorrect low-level saliency map, in order to overcome the existing difficulties, this paper proposes to fuse the color saliency based on global motion clues in a batch-wise fashion. And we also propose low-rank coherency guided spatial-temporal saliency diffusion to guarantee the temporal smoothness of saliency maps. Meanwhile, a series of saliency boosting strategies are designed to further improve the saliency accuracy. First, the original long-term video sequence is equally segmented into many short-term frame batches, and the motion clues of the individual video batch are integrated and diffused temporally to facilitate the computation of color saliency. Then, based on the obtained saliency clues, inter-batch saliency priors are modeled to guide the low-level saliency fusion. After that, both the raw color information and the fused low-level saliency are regarded as the low-rank coherency clues, which are employed to guide the spatial-temporal saliency diffusion with the help of an additional permutation matrix serving as the alternative rank selection strategy. Thus, it could guarantee the robustness of the saliency map's temporal consistence, and further boost the accuracy of the computed saliency map. Moreover, we conduct extensive experiments on five public available benchmarks, and make comprehensive, quantitative evaluations between our method and 16 state-of-the-art techniques. All the results demonstrate the superiority of our method in accuracy, reliability, robustness, and versatility. Chenglizhao Chen, Shuai Li 0001, Yongguang Wang, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 1 |
| 2016 | Robust salient motion detection in non-stationary videos via novel integrated strategies of spatio-temporal coherency clues and low-rank analysis
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 1 |
| 2015 | Real-time and robust object tracking in video via low-rank coherency analysis in feature space
Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
Pattern Recognit. | 1 |
| 2015 | Structure-Sensitive Saliency Detection via Multilevel Rank Analysis in Intrinsic Feature SpaceabstractThis paper advocates a novel multiscale, structure-sensitive saliency detection method, which can distinguish multilevel, reliable saliency from various natural pictures in a robust and versatile way. One key challenge for saliency detection is to guarantee the entire salient object being characterized differently from nonsalient background. To tackle this, our strategy is to design a structure-aware descriptor based on the intrinsic biharmonic distance metric. One benefit of introducing this descriptor is its ability to simultaneously integrate local and global structure information, which is extremely valuable for separating the salient object from nonsalient background in a multiscale sense. Upon devising such powerful shape descriptor, the remaining challenge is to capture the saliency to make sure that salient subparts actually stand out among all possible candidates. Toward this goal, we conduct multilevel low-rank and sparse analysis in the intrinsic feature space spanned by the shape descriptors defined on over-segmented super-pixels. Since the low-rank property emphasizes much more on stronger similarities among super-pixels, we naturally obtain a scale space along the rank dimension in this way. Multiscale saliency can be obtained by simply computing differences among the low-rank components across the rank scale. We conduct extensive experiments on some public benchmarks, and make comprehensive, quantitative evaluation between our method and existing state-of-the-art techniques. All the results demonstrate the superiority of our method in accuracy, reliability, robustness, and versatility. Chenglizhao Chen, Shuai Li 0001, Hong Qin 0001, Aimin Hao |
IEEE Trans. Image Process. | 1 |