VLDB 2026 Research / reviewers in the wild / expert
Lv Tang
dblp:243/6697
· DBLP profile ↗
33ranked-venue papers
12as first author
30since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 9 first-author · 21 since 2021Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | ReVerSeg: Training-free reason-and-verify video segmentation for traffic scenes
Senyun Kuang, Shuo Feng 0002, Lv Tang, Yintao Wei |
Expert Syst. Appl. | 5 |
| 2026 | UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World ModelabstractVision-and-Language Navigation (VLN) requires agents to autonomously navigate complex environments via visual images and natural language instructions—remains highly challenging. Recent research on enhancing language-guided navigation reasoning using pre-trained large language models (LLMs) has shown promising prospects. However, the reasoning of such methods is limited to the linguistic modality, lacking visual reasoning capabilities. Moreover, existing reasoning modules are optimized separately from navigation policies, leading to incompatibility and potential conflicts in optimization objectives. To tackle these challenges, we introduce UNeMo, a novel framework designed for the collaborative optimization of visual state reasoning and navigational decision-making. It introduces a Multimodal World Model (MWM) that takes visual features, language instructions, and navigational actions as inputs to jointly predict subsequent visual states, enabling cross-modal reasoning. Via a Hierarchical Prediction-Feedback (HPN) mechanism, MWM collaborates with navigation policies: the first layer generates actions using current vision-and-language features; MWM then infers post-action visual states to guide the second layer’s fine-grained decisions. This forms a dynamic bidirectional promotion mechanism where MWM reasoning optimizes navigation policies, while policy decisions feedback to improve MWM’s reasoning accuracy. Experiments on R2R and REVERIE datasets show UNeMo outperforms state-of-the-art methods by 2.1% and 0.7% in navigation accuracy for unseen scenes, validating its effectiveness. Changxin Huang, Lv Tang, Zhaohuan Zhan, Lisha Yu, Runhao Zeng, Zun Liu |
AAAI | 2 |
| 2026 | Disentangled self-supervised video camouflaged object detection and salient object detection
Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li |
Neural Networks | 2 |
| 2026 | UAR-NVC: A Unified Autoregressive Framework for Memory-Efficient Neural Video CompressionabstractImplicit Neural Representations (INRs) have demonstrated significant potential in video compression by representing videos as neural networks. However, as the number of frames increases, the memory consumption for training and inference increases substantially, posing challenges in resource-constrained scenarios. Inspired by the success of traditional video compression frameworks, which process video frame by frame and can efficiently compress long videos, we adopt this modeling strategy for INRs to decrease memory consumption, while aiming to unify the frameworks from the perspective of timeline-based autoregressive modeling. In this work, we present a novel understanding of INR models from an autoregressive (AR) perspective and introduce a Unified AutoRegressive Framework for memory-efficient Neural Video Compression (UAR-NVC). UAR-NVC integrates timeline-based and INR-based neural video compression under a unified autoregressive paradigm. It partitions videos into several clips and processes each clip using a different INR model instance, leveraging the advantages of both compression frameworks while allowing seamless adaptation to either in form. To further reduce temporal redundancy between clips, we treat the corresponding model parameters as proxies for these clips, and design two modules to optimize the initialization, training, and compression of these model parameters. In special, the Residual Quantization and Entropy Constraint (RQEC) module dynamically balances the reconstruction quality of the current clip and the newly introduced bitrate cost using the previously optimized parameters as conditioning. In addition, the Interpolation-based Initialization (II) module flexibly adjusts the degree of reference used during the initialization of neighboring video clips, based on their correlation. UAR-NVC supports adjustable latencies by varying the clip length. Extensive experimental results demonstrate that UAR-NVC, with its flexible video clip setting, can adapt to resource-constrained environments and significantly improve performance compared to different baseline models. The project page: https://wj-inf.github.io/UAR-NVC-page/. Jia Wang 0018, Xinfeng Zhang 0001, Gai Zhang, Lv Tang, Li Zhang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | HANeRV: Hierarchically Adaptive Neural Representation for Video CompressionabstractRecent advances in video compression introduce implicit neural representation (INR) based methods, which effectively capture global dependencies and characteristics of entire video sequences. Unlike traditional and deep learning based approaches, INR-based methods optimize network parameters from a global perspective, resulting in superior compression potential. However, most current INR methods utilize a fixed and uniform network architecture across all frames, limiting their adaptability to dynamic variations within and between video sequences. This often leads to suboptimal compression outcomes as these methods struggle to capture the distinct nuances and transitions in video content. To overcome these challenges, we propose Hierarchically Adaptive Neural Representation for Video Compression (HANeRV), an innovative INR-based video compression network that adaptively conducts structure optimisation based on the specific content of each video sequence. To better capture dynamic information across video sequences, we propose a dynamic architecture-level adjustment (DAA). Furthermore, to enhance the capture of dynamics between frames within a sequence, we implement a dynamic frame-level adjustment (DFA). Finally, to effectively capture spatial structural information within video frames, thereby enhancing the detail restoration capabilities of HANeRV, we devise a structure level hierarchical structural adaptation (HSA). Experimental results show that HANeRV achieves state-of-the-art performance among INR-based video compression methods and surpasses the H.266/VVC (x266, medium preset) anchor on diverse datasets. Lv Tang, Xinfeng Zhang 0001, Li Zhang 0006, Siwei Ma 0001, Qingming Huang |
IEEE Trans. Image Process. | 1 |
| 2025 | Boosting Vision State Space Model with Fractal ScanningabstractRecently, foundational models have significantly advanced in different tasks, accompanied by Transformer as the general backbone. However, Transformer's quadratic complexity poses challenges for handling longer sequences and higher resolution images, which may limit foundational models further development. To alleviate this issue, various efficient State Space Models (SSMs) like Mamba have emerged, initially matching Transformer performance and gradually surpassing it. To improve the performance of SSMs in computer vision tasks, one crucial viewpoint is effective serialization of images. Existing vision Mambas, which rely on a linear scanning mechanism, often struggle to capture complex spatial relationships in 2D images. This results in feature loss during serialization and negatively impacts model performance. To overcome this limitation, we propose the use of fractal scanning curves for image serialization to enhance the Mambas’ ability to accurately model complex spatial dependencies. Additionally, unlike existing vision Mambas, which are designed with various curve scanning directions that increase the complexity, contradicting the original intent of Mamba to enhance model performance. We novelty introduce the Fractal Fusion Pathway (FFP) for our FractalMamba, which can enhance its performance efficiently. Extensive experiments underscore the superiority of our proposed FractalMamba. Haoke Xiao, Lv Tang, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0130 |
AAAI | 2 |
| 2025 | Towards Training-Free Open-World Segmentation via Image Prompt Foundation Models
Lv Tang, Peng-Tao Jiang, Haoke Xiao, Bo Li 0115 |
Int. J. Comput. Vis. | 1 |
| 2025 | GCVIF: Pioneering Explainable Domain-Shared Representation Learning for Fault Signal Detection in Multiple Working States SimultaneouslyabstractWith the rapid development of sensor systems brought about by the industrial Internet of Things process, the need to unsupervisedly detect fault signals under multiple working conditions simultaneously has led to the emergence of multitarget domain adaptation (MTDA). This advancement is confronted by two primary challenges on domain adaptation: 1) fault feature extraction prefers target domains akin to the source domain, often sidelining others and 2) the learned features’ lack of interpretability. To address these, this article proposes the generalized-Gaussian cyclic variational inference framework (GCVIF). This framework engages the generalized-Gaussian cyclostationary distribution to initially capture the non-Gaussian and nonstationary attributes of fault signals, with an extended likelihood ratio test proposed to estimate distribution parameters. Leveraging these estimated distributions as priors, the generalized-Gaussian cyclic variational autoencoder is then developed to infer domain-shared representations. The process is steered by a specialized domain-shared representation learning principle, focusing on compact representation in the encoder and cyclostationary structure reconstruction in the decoder. Remarkably, extensive fault detection trials affirm that leveraging distributions estimated unsupervised as priors enables unbiased feature extraction, and the inference of domain-shared representations is inherently aligned with fault cyclostationary pulses simulation on the signal time domain, ensuring their direct mechanistic explainability. Lv Tang, Tan Chin-Hon, Tielin Shi, Jianping Xuan, Yu-Chao Cheng |
IEEE Internet Things J. | 2 |
| 2025 | Blending-Target Domain Adaptation for Intelligent Fault Recognition With Minimum Cycle Spiking Encoding and Adversarial AttackabstractThe generalization performance of intelligent fault recognition models is tied to the assumption of identical distribution. Domain adaptation allows the source model to be extended to single or multitarget domains with distribution shifts. However, the reliable transfer of multitarget domain adaptation (MTDA) is inseparable from domain annotation. In this article, we consider a more pragmatic but challenging MTDA setting where domain labels are absent. This setting threatens most existing methods due to the elusive gaps and agnostic affiliation. We propose a systematic approach to the new setting. First, the category semantic destruction and self-supervised clustering are used to estimate domain labels. Second, the attack features are constructed to consolidate adaptation by gradient alignment and classifier robustness. Extensive experiments demonstrate that the new setting is quite challenging for existing methods, while the proposed method outperforms the existing methods and effectively suppresses transfer preference. Lv Tang, Shaochen Li, Jianping Xuan, Tielin Shi |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | UVC: A Unified Deep Video Compression FrameworkabstractRecently, many works have applied deep learning techniques to video compression tasks, achieving promising results and advancing the field of Deep Learning-Based Video Compression (DLVC). However, the architecture design of the existing DLVC is rigid and limited in terms of flexibility. Specifically, different networks must be designed for different scenarios, such as delay-constrained scenario or non-delay-constrained scenario. Frequent switching between networks would reduce the speed of modern deep learning platforms and increase the maintenance costs. To address this problem, we propose a Unified Video Compression (UVC) framework that can be freely switched to different application scenarios without changing the network architecture. Our proposed UVC framework is based on the explicit-compression and implicit-generation perspective, which contains two sub-networks—the Explicit Reference Frame Compression Network (ERFCN) and the Implicit Reference Frame Generation Network (IRFGN). The aim of ERFCN is to compress the current frame with the help of the reference frame. To improve the performance of ERFCN, we first introduce the Transformer in this network, which can fully remove the spatial redundancy of the input image and is beneficial for the following inter-prediction process. We also develop a novel long-range motion estimation module for inter-prediction to generate motion vectors based on global motion information between two frames, which can handle long-range complex motion relations. The aim of IRFGN is to capture the temporal relationship between forward and backward reconstructed frames and synthesize a high-quality implicit reference frame for the current frame. To achieve this, we design the split spatial-temporal attention and multi-scale prediction module. We conduct extensive experiments on three widely used video compression databases (HEVC, UVG, and MCL-JVC), and the results demonstrate the superiority of our approach over other related DLVC methods. Lv Tang, Xinfeng Zhang 0001, Li Zhang 0006 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | ASAM: Boosting Segment Anything Model with Adversarial TuningabstractIn the evolving landscape of computer vision, foundation models have emerged as pivotal tools, exhibiting ex-ceptional adaptability to a myriad of tasks. Among these, the Segment Anything Model (SAM) by Meta AI has distin-guished itself in image segmentation. However, SAM, like its counterparts, encounters limitations in specific niche ap-plications, prompting a quest for enhancement strategies that do not compromise its inherent capabilities. This pa-per introduces ASAM, a novel methodology that amplifies SAM's performance through adversarial tuning. We har-ness the potential of natural adversarial examples, inspired by their successful implementation in natural language pro-cessing. By utilizing a stable diffusion model, we augment a subset (1%) of the SA-1B dataset, generating adversar-ial instances that are more representative of natural variations rather than conventional imperceptible perturbations. Our approach maintains the photorealism of adversarial ex-amples and ensures alignment with original mask annotations, thereby preserving the integrity of the segmentation task. The fine-tuned ASAM demonstrates significant im-provements across a diverse range of segmentation tasks without necessitating additional data or architectural mod-ifications. The results of our extensive evaluations confirm that ASAM establishes new benchmarks in segmentation tasks, thereby contributing to the advancement of foundational models in computer vision. Our project page is in https://asam2024.github.io/. Bo Li 0115, Haoke Xiao, Lv Tang |
CVPR | 3 |
| 2024 | Zero-Shot Co-Salient Object Detection FrameworkabstractCo-salient Object Detection (CoSOD) endeavors to replicate the human visual system’s capacity to recognize common and salient objects within a collection of images. Despite recent advancements in deep learning models, these models still rely on training with well-annotated CoSOD datasets. The exploration of training-free zero-shot CoSOD frameworks has been limited. In this paper, taking inspiration from the zero-shot transfer capabilities of foundational computer vision models, we introduce the first zero-shot CoSOD framework that harnesses these models without any training process. To achieve this, we introduce two novel components in our proposed framework: the group prompt generation (GPG) module and the co-saliency map generation (CMP) module. We evaluate the framework’s performance on widely-used datasets and observe impressive results. Our approach surpasses existing unsupervised methods and even outperforms fully supervised methods developed before 2020, while remaining competitive with some fully supervised methods developed before 2022. Haoke Xiao, Lv Tang, Bo Li 0115, Zhiming Luo, Shaozi Li |
ICASSP | 2 |
| 2024 | VQNeRV: Vector Quantization Neural Representation for Video CompressionabstractThe application of Implicit Neural Representations (INR) for video compression represents an evolving area of research. Despite its potential, conventional CNN-based INR methods often encounter difficulties in modeling complex spatiotemporal scenes. This is primarily due to the inherent limitations of CNNs in extracting intricate information, particularly when it comes to capturing temporal dynamics. Consequently, the integration of spatiotemporal contextual information becomes imperative to enhance INR's capacity for scene modeling. In response to these challenges, this paper introduces a novel approach, termed Vector Quantization Neural Representation (VQNeRV), specifically designed for video compression. Our methodology unfolds in three distinct stages: Firstly, we establish a CNN-based INR equipped with spatial embeddings to model the video content. Secondly, a Vector Quantization (VQ) encoder is then utilized to distill spatiotemporal features from the video. These features are subsequently amalgamated with the spatial embeddings through cross-attention mechanisms, facilitating a comprehensive feature fusion. Finally, an INR decoder, leveraging the combined features embodying spatiotemporal contextual information, reconstructs the video. Experimental results demonstrate that our method shows a much better rate-distortion performance compared to the VVC. Gai Zhang, Lv Tang, Xinfeng Zhang 0001 |
ISCAS | 2 |
| 2024 | Chain of Visual Perception: Harnessing Multimodal Large Language Models for Zero-shot Camouflaged Object DetectionabstractIn this paper, we introduce a novel multimodal camo-perceptive framework (MMCPF) aimed at handling zero-shot Camouflaged Object Detection (COD) by leveraging the powerful capabilities of Multimodal Large Language Models (MLLMs). Recognizing the inherent limitations of current COD methodologies, which predominantly rely on supervised learning models demanding extensive and accurately annotated datasets, resulting in weak generalization, our research proposes a zero-shot MMCPF that circumvents these challenges. Although MLLMs hold significant potential for broad applications, their effectiveness in COD is hindered and they would make misinterpretations of camouflaged objects. To address this challenge, we further propose a strategic enhancement called the Chain of Visual Perception (CoVP), which significantly improves the perceptual capabilities of MLLMs in camouflaged scenes by leveraging both linguistic and visual cues more effectively. We validate the effectiveness of MMCPF on five widely used COD datasets, containing CAMO, COD10K, NC4K, MoCA-Mask and OVCamo. Experiments show that MMCPF can outperform all existing state-of-the-art zero-shot COD methods, and achieve competitive performance compared to weakly-supervised and fully-supervised methods, which demonstrates the potential of MMCPF. The Github link of this paper is https://github.com/luckybird1994/MMCPF. Lv Tang, Peng-Tao Jiang, Hao Zhang 0063, Jinwei Chen 0003, Bo Li 0115 |
ACM Multimedia | 1 |
| 2024 | Open set transfer learning for bearing defect recognition based on selective momentum contrast and dual adversarial structure
Shaochen Li, Jianping Xuan, Zisheng Wang, Lv Tang, Tielin Shi |
Adv. Eng. Informatics | 5 |
| 2024 | From Composited to Real-World: Transformer-Based Natural Image MattingabstractThe task of image matting is an active research area in computer vision, and various trimap-free methods have been proposed to improve its performance. However, these methods do not consider the gap between composited and real-world images, resulting in limited generalization ability. To address this issue, we propose a domain alignment (DA) module that consists of local region-wise alignment (LRA) and global harmonious alignment (GHA). The LRA aligns the most diverse pixels in the transparent regions of the foreground between composited and real images. On the other hand, the GHA aligns the global image harmonization for both composited and real images, which helps the network choose the appropriate semantics for real harmonious images. Additionally, we design a transformer-based network with dynamic attention pruning (DAP) mechanism to accurately locate domain-sensitive regions, allowing the DA module to work more effectively. Furthermore, we introduce a new dataset, the Real-world Matting Dataset (RM-1k), to advance the real-world matting task. Our proposed method is evaluated on two composited benchmarks (Composite-1k and Distinctions-646) and two real-world datasets (AIM-500 and RM-1k), and the results show that our method achieves robust performance on both composited and real-world images. Lv Tang, Yijie Zhong 0001, Bo Li 0115 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | High Efficiency Deep-learning Based Video CompressionabstractAlthough deep learning technique has achieved significant improvement on image compression, but its advantages are not fully explored in video compression, which leads to the performance of deep-learning-based video compression (DLVC) is obviously inferior to that of hybrid video coding framework. In this article, we proposed a novel network to improve the performance of DLVC from its most important modules, including Motion Process (MP), Residual Compression (RC), and Frame Reconstruction (FR). In MP, we design a split second-order attention and multi-scale feature extraction module to fully remove the warping artifacts from multi-scale feature space and pixel space, which can help reduce the distortion in the following process. In RC, we propose a channel selection mechanism to gradually drop redundant information while preserving informative channels for a better rate-distortion performance. Finally, in FR, we introduce a residual multi-scale recurrent network to improve the quality of the current reconstructed frame by progressively exploiting temporal context information between it and its several previous reconstructed frames. Extensive experiments are conducted on the three widely used video compression datasets (HEVC, UVG, and MCL-JVC), and the performance demonstrates the superiority of our proposed approach over the state-of-the-art methods. Lv Tang, Xinfeng Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Unified and Scalable Deep Image Compression Framework for Human and MachineabstractImage compression aims to minimize the amount of data in image representation while maintaining a certain visual quality for humans, which is an essential technique for storage and transmission. Recently, along with the development of computer vision, machines have become another primary receiver for images and require compressed images at a certain quality level, which may be different from that of human vision. In many scenarios, compressed images should serve both human and machine vision tasks, but few compression methods are designed for both goals simultaneously. In this article, we propose a unified and scalable deep image compression (USDIC) framework that jointly optimizes the image quality according to human and machine vision in an end-to-end style. For the encoder, we propose an information splitting mechanism (ISM) to separate images into semantic and visual features, which mainly aims at machine analysis and human viewing tasks. For the decoder, we design a scalable decoding architecture. The encoded semantic feature is first decoded for machine analysis tasks, and the image is decoded and reconstructed further by leveraging the decoded semantic features. Herein, to further remove the redundancy between the semantic and visual features of images, we propose a scalable entropy model (SEM) with a joint optimization strategy to reconstruct the image using the two kinds of decoded features. Extensive experimental results show that the proposed USDIC achieves much better performance on the image analysis task while maintaining competitive performance on the traditional image reconstruction task compared with popular image compression methods. Gai Zhang, Xinfeng Zhang 0001, Lv Tang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Scene Matters: Model-based Deep Video CompressionabstractVideo compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal prediction theory to enhance compression performance by designing high efficient intra and inter prediction strategies and compressing video frames one by one. In this paper, we propose a novel model-based video compression (MVC) framework that regards scenes as the fundamental units for video sequences. Our proposed MVC directly models the intensity variation of the entire video sequence in one scene, seeking non-redundant representations instead of reducing redundancy through spatio-temporal predictions. To achieve this, we employ implicit neural representation as our basic modeling architecture. To improve the efficiency of video modeling, we first propose context-related spatial positional embedding and frequency domain supervision in spatial context enhancement. For temporal correlation capturing, we design the scene flow constrain mechanism and temporal contrastive loss. Extensive experimental results demonstrate that our method achieves up to a 20% bitrate reduction compared to the latest video coding standard H.266 and is more efficient in decoding than existing video coding strategies. Lv Tang, Xinfeng Zhang 0001, Gai Zhang, Xiaoqi Ma |
ICCV | 1 |
| 2023 | HM-PCGC: A Human-Machine Balanced Point Cloud Geometry Compression SchemeabstractPoint cloud compression has various purposes in different application scenarios, such as requiring visual fidelity in human vision tasks and pursuing semantic fidelity in machine vision tasks. To accommodate these diverse requirements, we propose a Human-Machine balanced point cloud geometry compression scheme (HM-PCGC) which considers the characteristics of various tasks. Our proposed scheme starts from a pre-trained, lightweight point cloud compression backbone and employs a Learned Semantic Mining module to aggregate multi-tasks features. By leveraging the aggregated features, HM-PCGC is able to retain the geometry and semantic properties of the point clouds during compression. To better balance between the signal distortion and semantic distortion, we integrated a multi-task learning mechanism during the training phase. Our approach is extensively evaluated and analyzed, and the results demonstrate its superiority over state-of-the-art traditional and deep learning based point cloud codecs for both signal reconstruction and machine vision tasks. Xiaoqi Ma, Yingzhan Xu, Xinfeng Zhang 0001, Lv Tang, Kai Zhang 0007, Li Zhang 0006 |
ICIP | 4 |
| 2023 | Transfer reinforcement learning method with multi-label learning for compound fault recognition
Zisheng Wang, Lv Tang, Tielin Shi, Jianping Xuan |
Adv. Eng. Informatics | 3 |
| 2023 | Enhanced Quantified Local Implicit Neural Representation for Image CompressionabstractRecently, implicit neural representation (INR) has been applied to image compression. However, the rate-distortion performance of most existing INR-based image compression methods is still obviously inferior to the state-of-the-art image compression methods. In this paper, we propose an Enhanced Quantified Local Implicit Neural Representation (EQLINR) for image compression by enhancing the utilization of local relationships of INR and narrow the quantization gap between training and encoding to further improve the performance of INR-based image compression. Our framework consists of latent representation and the corresponding implicit neural network consisting of MLP and CNN, which can transform the latent representation into the image space. To enhance local relationships utilization, we design a local enhancement module (LEM) consisted of CNN to capture the neighborhood relationships of the reconstructed image from MLP. Furthermore, to mitigate the performance loss caused by quantization of latent representation, we employ an enhanced quantization scheme (EQS) in our training process. We use uniform noise for network initialization and then use Stochastic Gumbel Annealing (SGA) with dynamic temperature regulation as a proxy function for quantization during training. Extensive experimental results demonstrate that our approach significantly the compression performance of INR-based image compression, and even better than BPG Gai Zhang, Xinfeng Zhang 0001, Lv Tang |
IEEE Signal Process. Lett. | 3 |
| 2022 | Detecting Camouflaged Object in Frequency DomainabstractCamouflaged object detection (COD) aims to identify objects that are perfectly embedded in their environment, which has various downstream applications in fields such as medicine, art, and agriculture. However, it is an extremely challenging task to spot camouflaged objects with the perception ability of human eyes. Hence, we claim that the goal of COD task is not just to mimic the human visual ability in a single RGB domain, but to go beyond the human biological vision. We then introduce the frequency domain as an additional clue to better detect camouflaged objects from backgrounds. To well involve the frequency clues into the CNN models, we present a powerful network with two special components. We first design a novel frequency enhancement module (FEM) to dig clues of camouflaged objects in the frequency domain. It contains the offline discrete cosine transform followed by the learnable enhancement. Then we use a feature alignment to fuse the features from RGB domain and frequency domain. Moreover, to further make full use of the frequency information, we propose the high-order relation module (HOR) to handle the rich fusion feature. Comprehensive experiments on three widely-used COD datasets show the proposed method significantly outperforms other state-of-the-art methods by a large margin. Yijie Zhong 0001, Bo Li 0115, Lv Tang, Senyun Kuang, Shuang Wu 0001, Shouhong Ding |
CVPR | 3 |
| 2022 | Rethinking Two-B-Real Net for Real-Time Salient Object DetectionabstractExploring a fast and accurate salient object detection (SOD) model is a promising research area. TBRS [1] has been proposed a two-branch network for real-time SOD. However, its principle of adding an extra path to encode spatial information is time-consuming. And its backbone is borrowed from image classification tasks, may be inefficient for SOD due to the deficiency of task-specific design. To handle these problems, we propose a novel and efficient structure named short-range concatenate module (SRCM) by removing structure redundancy. Specifically, we gradually reduce the dimension of feature maps and use the aggregation of them for image representation, which forms the basic module of SRCM network. Moreover, we propose an efficient detail guidance branch (DBG) to further encode detail structural information in low-level stages instead of the time-consuming perceptual branch used in TBRS. Finally, low-level features and high-level features are fused by the feature projection module (FPM). Extensive evaluations and analysis demonstrate that our proposed algorithm achieves the leading accuracy performance with real-time speed (216fps). We hope that our series of works can motivate future research for real-time SOD task. Senyun Kuang, Shijin Meng, Lv Tang, Bo Li 0115 |
ICASSP | 4 |
| 2022 | Multitarget domain adaptation with transferable hyperbolic prototypes for intelligent fault diagnosis
Lv Tang, Jianping Xuan, Tielin Shi |
Knowl. Based Syst. | 1 |
| 2022 | Re-Thinking the Relations in Co-Saliency DetectionabstractCo-salient object detection (CoSOD) aims to detect common salient objects sharing the same attributes in an image group. The key issue of CoSOD is how to model the inter-saliency relations within an image group. The major limitation of previous methods is that they pre-define the group-to-one relations within an image group. In this paper, we propose a new concept of structural inter-saliency relations and solve the CoSOD with deep reinforcement learning framework. Firstly, we design a semantic relation graph (SRG) to model the structural inter-saliency relations. Then the feature selecting agent (FS-agent) aims to select the informative features, which can help the SRG effectively model structural inter-saliency relations. Finally, relation updating agent (RU-agent) progressively updates the SRG to focus on the co-salient relations like human decision-making process. Extensive experiments on co-saliency datasets show that because of well modeling inter-saliency relations in image group, our proposed method achieves superior performance compared to the state-of-the-art methods. We hope that this paper can motivate future research for visual co-analysis tasks. Lv Tang, Bo Li 0115, Senyun Kuang, Mofei Song, Shouhong Ding |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Toward Stable Co-Saliency Detection and Object Co-SegmentationabstractIn this paper, we present a novel model for simultaneous stable co-saliency detection (CoSOD) and object co-segmentation (CoSEG). To detect co-saliency (segmentation) accurately, the core problem is to well model inter-image relations between an image group. Some methods design sophisticated modules, such as recurrent neural network (RNN), to address this problem. However, order-sensitive problem is the major drawback of RNN, which heavily affects the stability of proposed CoSOD (CoSEG) model. In this paper, inspired by RNN-based model, we first propose a multi-path stable recurrent unit (MSRU), containing dummy orders mechanisms (DOM) and recurrent unit (RU). Our proposed MSRU not only helps CoSOD (CoSEG) model captures robust inter-image relations, but also reduces order-sensitivity, resulting in a more stable inference and training process. Moreover, we design a cross-order contrastive loss (COCL) that can further address order-sensitive problem by pulling close the feature embedding generated from different input orders. We validate our model on five widely used CoSOD datasets (CoCA, CoSOD3k, Cosal2015, iCoseg and MSRC), and three widely used datasets (Internet, iCoseg and PASCAL-VOC) for object co-segmentation, the performance demonstrates the superiority of the proposed approach as compared to the state-of-the-art (SOTA) methods. Bo Li 0115, Lv Tang, Senyun Kuang, Mofei Song, Shouhong Ding |
IEEE Trans. Image Process. | 2 |
| 2021 | Highly Efficient Natural Image Matting
Yijie Zhong 0001, Bo Li 0115, Lv Tang, Hao Tang 0005, Shouhong Ding |
BMVC | 3 |
| 2021 | Fast: Feature Aggregation for Detecting Salient Object in Real-TimeabstractThis paper introduces a method named FAST for real-time salient object detection with an extremely efficient CNN architecture. Our proposed network starts from a single lightweight backbone and aggregates discriminative features through network-level and phase-level respectively. Based on the multi-scale feature propagation, FAST substantially reduces the number of parameters, but still obtains sufficient receptive field and enhances the model learning ability, which strikes a balance between the speed and performance. To better preserve object boundaries, we also explore the complementary between salient object information and edge information within our lightweight architecture. Extensive evaluations and analysis demonstrate that the proposed algorithm achieves the leading accuracy performance with real-time speed (186fps) which is significantly faster than the existing state-of-the-art methods. Lv Tang, Bo Li 0115, Yanliang Wu, Shouhong Ding |
ICASSP | 1 |
| 2021 | Disentangled High Quality Salient Object DetectionabstractAiming at discovering and locating most distinctive objects from visual scenes, salient object detection (SOD) plays an essential role in various computer vision systems. Coming to the era of high resolution, SOD methods are facing new challenges. The major limitation of previous methods is that they try to identify the salient regions and estimate the accurate objects boundaries simultaneously with a single regression task at low-resolution. This practice ignores the inherent difference between the two difficult problems, resulting in poor detection quality. In this paper, we propose a novel deep learning framework for high-resolution SOD task, which disentangles the task into a low-resolution saliency classification network (LRSCN) and a high-resolution refinement network (HRRN). As a pixel-wise classification task, LRSCN is designed to capture sufficient semantics at low-resolution to identify the definite salient, background and uncertain image regions. HRRN is a regression task, which aims at accurately refining the saliency value of pixels in the uncertain region to preserve a clear object boundary at high-resolution with limited GPU memory. It is worth noting that by introducing uncertainty into the training process, our HRRN can well address the high-resolution refinement task without using any high-resolution training data. Extensive experiments on high-resolution saliency datasets as well as some widely used saliency benchmarks show that the proposed method achieves superior performance compared to the state-of-the-art methods. Lv Tang, Bo Li 0115, Yijie Zhong 0001, Shouhong Ding, Mofei Song |
ICCV | 1 |
| 2020 | CLASS: Cross-Level Attention and Supervision for Salient Objects Detection
Lv Tang, Bo Li 0061 |
ACCV (3) | 1 |
| 2019 | Two-B-real Net: Two-branch Network for Real-time Salient Object DetectionabstractAs a hot topic in computer vision, recent researches on salient object detection (SOD) have focused on using the over-designed deep convolutional neural networks (CNNs) to improve the detection accuracy. However, these complex architectures constraint themselves to low speed and drag them on wide-ranging applications. In this paper, we simplify the over-designed networks and propose the Two-Branch Network for Real-time Salient Object Detection (Two-B-Real Net). Particularly, the Perceptual Branch and the Objectness Branch in our network can efficiently capture detailed information and distinctive objectness simultaneously. And we also design novel attention mechanisms to guide the network to focus on most saliency-related features and generate more accurate results. Extensive evaluations show that the proposed algorithm achieves the leading accuracy performance with real-time speed (125fps) which is significantly faster than the existing methods. Bo Li 0061, Zhengxing Sun, Lv Tang, Anqi Hu |
ICASSP | 3 |
| 2019 | Detecting Robust Co-Saliency with Recurrent Co-Attention Neural NetworkabstractEffective feature representations which should not only express the images individual properties, but also reflect the interaction among group images are essentially crucial for robust co-saliency detection. This paper proposes a novel deep learning co-saliency detection approach which simultaneously learns single image properties and robust group feature in a recurrent manner. Specifically, our network first extracts the semantic features of each image. Then, a specially designed Recurrent Co-Attention Unit (RCAU) will explore all images in the group recurrently to generate the final group representation using the co-attention between images, and meanwhile suppresses noisy information. The group feature which contains complementary synergetic information is later merged with the single image features which express the unique properties to infer robust co-saliency. We also propose a novel co-perceptual loss to make full use of interactive relationships of whole images in the training group as the supervision in our end-to-end training process. Extensive experimental results demonstrate the superiority of our approach in comparison with the state-of-the-art methods. Bo Li 0061, Zhengxing Sun, Lv Tang, Yunhan Sun |
IJCAI | 3 |